Sign in

dan bateyko

@dbateyko.bsky.social
133 followers 211 following 41 posts

maybe the hard stuff's inside, hidden — like bones, as opposed to an exoskeleton. @CornellInfoSci dbateyko.info

PostsRepliesMedia
Reposted by dan bateyko
Kenny Peng @kennypeng.bsky.social · 06/10/2026
Our new paper introduces a scientific theory of atomic features. We mathematically derive testable predictions. Our experiments challenge conventional wisdom. SAEs of different size and training data share many features. Large SAEs recover both parent and child features. 🧵
1207
dan bateyko @dbateyko.bsky.social · 23/09/2026
Reading Cornell’s report on The Future of the American University and thought something felt off. The report explicitly says AI “was not used to automatically write the report or pieces of it.” And yet Pangram flags an entire section—about scholarly integrity—as AI-generated.
010
dan bateyko @dbateyko.bsky.social · 15/09/2026
Nice coverage in @politico.com this afternoon of our work! www.politico.com/newsletters/...
politico.com
How AI can root out discrimination
000
dan bateyko @dbateyko.bsky.social · 08/09/2026
You can explore the thousands of discriminatory laws we found here! hidden-in-plain-text.reglabapp.com
hidden-in-plain-text.reglabapp.com
Hidden in Plain Text: LLM-Assisted Detection of Discriminatory Local Laws
An LLM-assisted research pipeline surfaces suspect discriminatory provisions still on the books across 9,623 U.S. municipalities.
100
dan bateyko @dbateyko.bsky.social · 08/09/2026
Terrible! My colleagues have looked and found racially restrictive covenants in California using similar techniques: law.stanford.edu/stanford-law...
law.stanford.edu
Using AI to Map Racial Covenants | Stanford Law School
When Daniel E. Ho purchased a home in Palo Alto, he got more than he bargained for. Buried in the reams of papers he had to sign at closing was an uns
001
dan bateyko @dbateyko.bsky.social · 08/09/2026
LLM-assisted statutory surveys can make legal cleanup possible at a previously impractical scale. You can read the full paper, published at ICAIL 2026 with the stellar team of co-first author Yasmine Mabene, Derek Ouyang, and Dan Ho, here: hidden-in-plain-text.reglabapp.com/hidden-in-pl...
hidden-in-plain-text.reglabapp.com
130
dan bateyko @dbateyko.bsky.social · 08/09/2026
Bowling alleys. Go-kart tracks. We found 2,000+ laws restricting noncitizens from working at places like these. Mass. repealed its citizenship rule for liquor licenses only in 2024, after a family was barred from applying. Yet two years later, similar rules remain in town codes.
120
dan bateyko @dbateyko.bsky.social · 08/09/2026
Local law is full of fossils. We ran LLMs over 9,000 jurisdictions to identify plainly discriminatory ones. We found: — separate cemeteries for “white and black” residents — voting limited to “male persons” — speech rules protecting “any woman or child” from indecent language
130
dan bateyko @dbateyko.bsky.social · 08/09/2026
Local law still formally mandates racial segregation. In 2026. We built the largest dataset of local laws and used LLMs to search for discriminatory provisions across the nation. What we found should not still be on the books. hai.stanford.edu/policy/makin...
hai.stanford.edu
Making Local Law Legible: LLM-Assisted Detection of Discrimination | Stanford HAI
This brief demonstrates the potential of LLM-assisted review to accelerate legal reform.
2127
dan bateyko @dbateyko.bsky.social · 07/09/2026
just finding out that the Python data visualization library seaborn is named after West Wing's Sam Seaborn and imported as sns due to the monogram on his shirt
000
dan bateyko @dbateyko.bsky.social · 24/07/2026
Early results, but! I’ve scraped 100,000+ law review articles (maybe~2x the next-largest study) and am using language models to classify their footnotes. Having fun seeing how tech law scholarship shakes out. So far, STS is overrepresented; sociology is slightly under.
001
Reposted by dan bateyko
Data & Society @datasociety.bsky.social · 23/07/2026
To make govt services more accessible & accountable, @baricks.bsky.social, @boston.gov’s aleja jimenez jaramillo & @megyoung0.bsky.social call on state & local govts to explore collectively-governed “in-house” alternatives to commercial language translation tech. datasociety.net/research-lib...
152
Reposted by dan bateyko
James Grimmelmann @jtlg.bsky.social · 20/07/2026
The 16th Edition of my Internet Law casebook is out! Overview: internetcasebook.com Table of contents: www.semaphorepress.com/downloads/In... PDF download (suggested $30): semaphorepress.com/InternetLaw_... Print-on-demand ($75.10): www.amazon.com/dp/1943689245 A brief thread on the update:
The cover of Internet Law: Cases and Problems (16th ed.) by James Grimmelmann.

The cover is orange and features a photograph of hundreds of screens glowing in the dark at the Assembly demoscene LAN party.
1229
Reposted by dan bateyko
Joachim Baumann @joachimbaumann.bsky.social · 07/07/2026
Just arrived at ICML 🇰🇷😍 Get up early tomorrow to hear me talk about how (not) to solve the peer review crisis, or find me at one of my poster presentations. Paper links: ✅ AI Peer Review: arxiv.org/abs/2605.03202 ✅ SWE-chat: arxiv.org/pdf/2604.20779
ICML Conference schedule with two papers. Left, "Stop Automating Peer Review Without Rigorous Evaluation": Oral presentation, Wed 7/8/2026, 10:00–10:15 AM KST, Grand Ballroom 101–105; Poster, Wed 7/8/2026, 2:30–4:15 PM KST, Hall A #3003. Right, "SWE-chat": at the 5th Deep Learning for Code Workshop, Fri 7/10/2026, 13:00–14:30 KST, Hall B2.
0162
Reposted by dan bateyko
Ira Globus-Harris @iraglobusharris.bsky.social · 03/07/2026
Are you at ICML next week? Feel like your decision-making for which sessions to attend might not be risk minimizing? Don't incur (swap) regret and come to my, @aaroth.bsky.social, and @ncollina.bsky.social's tutorial Monday on multicalibration, decision-making, and collaborative learning!
22011
Reposted by dan bateyko
Lauren Chambers @laurenmarietta.bsky.social · 28/06/2026
it's the last day of #FAccT2026 in Montreal (bonjour hiii), and I'm presenting my paper with Diag Davenport on the promise of #PublicInterestTech clinics for training the next gen of critical sociotechnical thinkers. 🏆 plus we got an honorable mention!? come thru @ 10:45! 📄 tiny.cc/pit-clinics-26
screenshot of a title slide, navy background with light blue icons and white and gold text: "a decision-making pedagogy for the public interest technology clinic: putting the ‘practice’ in critical technical practice." lauren m. chambers & diag davenport, uc berkeley. facct @ montreal, june 28, 2026
1217
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 23/06/2026
I'm so excited to attend #FAccT2026 in 🇨🇦Montreal🇨🇦 to present "Tradeoffs are Domain Dependent: Improving Accuracy and Fairness in Property Tax Assessments" by Evelyn Smith, me, Chris Berry, @jacobsgoldin.bsky.social, and Dan Ho!! 🔗: arxiv.org/pdf/2605.15020
A screenshot of our paper: 

Title: Tradeoffs are Domain Dependent: Improving Accuracy and Fairness in Property Tax Assessments
Authors: EVELYN SMITH and EMMA HARVEY (co-first authors), CHRISTOPHER BERRY, JACOB GOLDIN and DANIEL E. HO (co-senior authors)
Abstract: Algorithmic fairness research often assumes a tradeoff between fairness and accuracy. Yet this tradeoff may not be universal. We test this assumption in the context of U.S. property tax assessment - a setting in which the output of predictive algorithms directly determines the distribution of tax obligations among homeowners. Currently, systematic assessment errors cause owners of lower-valued properties to face disproportionately high tax burdens, creating regressivity in the property tax system. Using data on 26 million property sales spanning 95% of U.S. counties, we conduct three complementary analyses. First, we find that assessment accuracy and fairness - measured using domain-relevant metrics - are strongly correlated across counties under status quo practices. Second, in simulated assessment models, we show that adding property features improves accuracy in most cases, and that when accuracy improves, fairness almost always improves as well. Third, we show that incorporating publicly available Census data into assessment models - a feasible reform in most counties - would significantly improve both accuracy and fairness relative to status quo assessments. Together, these results challenge the presumed universality of the fairness-accuracy tradeoff and demonstrate that well-designed modeling improvements can advance both fairness and accuracy in large-scale public sector systems.
1113
Reposted by dan bateyko
Isabel Silva Corpus @isabelcorpus.bsky.social · 24/06/2026
Excited to attend FAccT 2026 in Montreal this week! Let me know if you'll be there and want to catch up :) I'll be presenting a paper with @allisonkoe.bsky.social about ad delivery skew in the context of government advertising, come by and check it out!
2257
Reposted by dan bateyko
Jennah Gosciak @jennahgosciak.bsky.social · 25/06/2026
I am excited to be at @facct.bsky.social this year presenting a new 📝 "Scrutinizing Index-Based Risk Assessments: A Case Study in NYC Decision-making for Heat Emergency Management" (work with Luke Boyce, @angelinawang.bsky.social , and @allisonkoe.bsky.social ). 🔗: dl.acm.org/doi/10.1145/... (1/10)
1196
dan bateyko @dbateyko.bsky.social · 18/06/2026
Beautiful. Congratulations!!
110
Reposted by dan bateyko
Lindsey Barrett @lambarrett.bsky.social · 30/05/2026
outrageously corny how happy and grateful I feel after PLSC, every single time. man do I love my job and the lovely geniuses it gives me proximity to. privacy odd ducks conference forever!!!!
121
Reposted by dan bateyko
Lucy Li @lucy3.bsky.social · 31/03/2026
Another one of @ahalterman.bsky.social and @katakeith.bsky.social's papers that I think should be cited more by CSS researchers: What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification arxiv.org/abs/2510.03541
arxiv.org
What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
Generative large language models (LLMs) are now used extensively for text classification in computational social science (CSS). In this work, focus on the steps before and after LLM prompting -- conce...
1315
dan bateyko @dbateyko.bsky.social · 27/03/2026
Wild Anna’s Archive bounty, reads like a heist. They want someone to front tens of thousands of dollars to buy Library of Congress files for a 3k bounty. I imagine a leak investigation would have a very short suspect list
010
Reposted by dan bateyko
Peter Henderson @peterhenderson.bsky.social · 27/03/2026
Btw, did a bit of a rebranding of the substack. Will endeavor to post more there. h/t @dbateyko.bsky.social on the Trials & Errors name. Super fitting name for a group whose focus is both in reinforcement learning and in law/governance research. www.trialserrors.ai
trialserrors.ai
Trials & Errors | Peter Henderson | Substack
Various news, thoughts, and findings on the intersection of law, policy, and artificial intelligence. Click to read Trials & Errors, by Peter Henderson, a Substack publication with hundreds of subscri...
131
Reposted by dan bateyko
Maria Antoniak @mariaa.bsky.social · 24/03/2026
I've started a "History of NLP" repo to store all of these resources. I don't have time to add everything yet, but I'll keep chipping away, and help is welcome. github.com/maria-antoni...
github.com
GitHub - maria-antoniak/history-of-nlp: a public, crowd-sourced bibliography about the history of natural language processing (nlp)
a public, crowd-sourced bibliography about the history of natural language processing (nlp) - maria-antoniak/history-of-nlp
1366
Reposted by dan bateyko
Dominik Stammbach @dominsta.bsky.social · 28/01/2026
📣 Call for Contributions: LEXam-v2 – A Benchmark for Legal Reasoning in AI How well do today’s AI systems really reason about law? We’re building a global benchmark based on real law school & bar exams. 🧵 Full details, scope, and how to contribute in the thread 👇
164
Reposted by dan bateyko
J. Nathan Matias @natematias.bsky.social · 15/11/2025
One joy of growing as a scholar & doer has been the pleasure of being supported to pay attention to other people’s excellent work and amplify it. Next week I am publishing an article summarizing over 170 articles on AI + science + policy and tomorrow I get to email their authors to say thanks <3
1123
dan bateyko @dbateyko.bsky.social · 02/09/2025
000
dan bateyko @dbateyko.bsky.social · 02/09/2025
100
dan bateyko @dbateyko.bsky.social · 02/09/2025
110
dan bateyko @dbateyko.bsky.social · 02/09/2025
It's remarkable how early Ford Foundation was to law and technology in the midcentury
020
dan bateyko @dbateyko.bsky.social · 02/09/2025
000
dan bateyko @dbateyko.bsky.social · 02/09/2025
010
dan bateyko @dbateyko.bsky.social · 01/09/2025
this is what you see moments before going down a cyberspace and law rabbit hole
110
dan bateyko @dbateyko.bsky.social · 30/08/2025
000
dan bateyko @dbateyko.bsky.social · 30/08/2025
100
dan bateyko @dbateyko.bsky.social · 01/08/2025
it finally happened (my 3090 overheated and emergency shut off)
010
Reposted by dan bateyko
Daphne Keller @daphnek.bsky.social · 01/08/2025
Here’s the article. I’ve had more positive feedback on it than things I spent a year on. Apparently describing a problem that thousands of Trust and Safety people are seeing but also see the world ignoring is a good way to win hearts and minds :) www.lawfaremedia.org/article/the-...
lawfaremedia.org
The Rise of the Compliant Speech Platform
Content moderation is becoming a “compliance function,” with trust and safety operations run like factories and audited like investment banks.
13010
dan bateyko @dbateyko.bsky.social · 18/07/2025
“The possibilities of the pole” @hoctopi.bsky.social
000
dan bateyko @dbateyko.bsky.social · 18/07/2025
Eliza asking after an anti-suffering future of reproductive technology (the piece looks incredible)
100
dan bateyko @dbateyko.bsky.social · 18/07/2025
@kevinbaker.bsky.social “The rules appear inevitable, natural, reasonable. We forget they were drawn by human hands.”
180
dan bateyko @dbateyko.bsky.social · 18/07/2025
At the Kernel 5 issue launch!
130
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 14/07/2025
After having such a great time at #CHI2025 and #FAccT2025, I wanted to share some of my favorite recent papers here! I'll aim to post new ones throughout the summer and will tag all the authors I can find on Bsky. Please feel welcome to chime in with thoughts / paper recs / etc.!! 🧵⬇️:
25510
Reposted by dan bateyko
jay @kuppermann.xyz · 15/07/2025
i am launching a magazine with @kevinbaker.bsky.social and the rest of the reboot collective on thursday at gray area! you should be there! open.substack.com/pub/reboothq...
Kernel Magazine
Issue 5: Rules Launch Party
July 17
6:30-9pm
SF
Gray Area

The illustration behind this text is a mix of gold chess pieces and purple snakes on a chessboard pattern

lu.ma/k5-sf
0144
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 01/07/2025
I've arrived in the 🌁Bay Area🌁, where I'll be spending the summer as a research fellow at Stanford's RegLab! If you're also here, LMK and let's get a meal / go on a hike / etc!!
A close-up view of the Golden Gate Bridge in the fog
091
Reposted by dan bateyko
James Grimmelmann @jtlg.bsky.social · 02/07/2025
Well, this was a nice surprise. Download it while it’s hot! lsolum.typepad.com/legaltheory/...
lsolum.typepad.com
Grimmelmann, Sobel, & Stein on Generative AI and Legal Interpretation
James Grimmelmann (Cornell Law School; Cornell Tech), Benjamin Sobel (Cornell University - Cornell Tech NYC), & David Stein (Vanderbilt University - Vanderbilt Law School) have posted Generative Misin...
041
Reposted by dan bateyko
Jennah Gosciak @jennahgosciak.bsky.social · 24/06/2025
I am presenting a new 📝 “Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments” at @facct.bsky.social on Thursday, with @aparnabee.bsky.social, Derek Ouyang, @allisonkoe.bsky.social, @marzyehghassemi.bsky.social, and Dan Ho. 🔗: arxiv.org/abs/2506.13735 (1/n)
"Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments"

Conducting disparity assessments at regular time intervals is critical for surfacing potential biases in decision-making and improving outcomes across demographic groups. Because disparity assessments fundamentally depend on the availability of demographic information, their efficacy is limited by the availability and consistency of available demographic identifiers. While prior work has considered the impact of missing data on fairness, little attention has been paid to the role of delayed demographic data. Delayed data, while eventually observed, might be missing at the critical point of monitoring and action -- and delays may be unequally distributed across groups in ways that distort disparity assessments. We characterize such impacts in healthcare, using electronic health records of over 5M patients across primary care practices in all 50 states. Our contributions are threefold. First, we document the high rate of race and ethnicity reporting delays in a healthcare setting and demonstrate widespread variation in rates at which demographics are reported across different groups. Second, through a set of retrospective analyses using real data, we find that such delays impact disparity assessments and hence conclusions made across a range of consequential healthcare outcomes, particularly at more granular levels of state-level and practice-level assessments. Third, we find limited ability of conventional methods that impute missing race in mitigating the effects of reporting delays on the accuracy of timely disparity assessments. Our insights and methods generalize to many domains of algorithmic fairness where delays in the availability of sensitive information may confound audits, thus deserving closer attention within a pipeline-aware machine learning framework.Figure contrasting a conventional approach to conducting disparity assessments, which is static, to the analysis we conduct in this paper. Our analysis (1) uses comprehensive health data from over 1,000 primary care practices and 5 million patients across the U.S., (2) timestamped information on the reporting of race to measure delay, and (3) retrospective analyses of disparity assessments under varying levels of delay.
1134
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 23/06/2025
I am so excited to be in 🇬🇷Athens🇬🇷 to present "A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms" by me, @kizilcec.bsky.social, and @allisonkoe.bsky.social, at #FAccT2025!! 🔗: arxiv.org/pdf/2506.04419
A screenshot of our paper's:

Title: A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
Authors: Emma Harvey, Rene Kizilcec, Allison Koenecke
Abstract: Increasingly, individuals who engage in online activities are expected to interact with large language model (LLM)-based chatbots. Prior work has shown that LLMs can display dialect bias, which occurs when they produce harmful responses when prompted with text written in minoritized dialects. However, whether and how this bias propagates to systems built on top of LLMs, such as chatbots, is still unclear. We conduct a review of existing approaches for auditing LLMs for dialect bias and show that they cannot be straightforwardly adapted to audit LLM-based chatbots due to issues of substantive and ecological validity. To address this, we present a framework for auditing LLM-based chatbots for dialect bias by measuring the extent to which they produce quality-of-service harms, which occur when systems do not work equally well for different people. Our framework has three key characteristics that make it useful in practice. First, by leveraging dynamically generated instead of pre-existing text, our framework enables testing over any dialect, facilitates multi-turn conversations, and represents how users are likely to interact with chatbots in the real world. Second, by measuring quality-of-service harms, our framework aligns audit results with the real-world outcomes of chatbot use. Third, our framework requires only query access to an LLM-based chatbot, meaning that it can be leveraged equally effectively by internal auditors, external auditors, and even individual users in order to promote accountability. To demonstrate the efficacy of our framework, we conduct a case study audit of Amazon Rufus, a widely-used LLM-based chatbot in the customer service domain. Our results reveal that Rufus produces lower-quality responses to prompts written in minoritized English dialects.
13110
Reposted by dan bateyko
Allison Koenecke @allisonkoe.bsky.social · 22/06/2025
🎉Excited to present our paper tomorrow at @facct.bsky.social, “Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese”, with @brucelyu17.bsky.social, Jiebo Luo and Jian Kang, revealing 🤖 LLM performance disparities. 📄 Link: arxiv.org/abs/2505.22645
"Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese" Abstract:

While the capabilities of Large Language Models (LLMs) have been studied in both Simplified and Traditional Chinese, it is yet unclear whether LLMs exhibit differential performance when prompted in these two variants of written Chinese. This understanding is critical, as disparities in the quality of LLM responses can perpetuate representational harms by ignoring the different cultural contexts underlying Simplified versus Traditional Chinese, and can exacerbate downstream harms in LLM-facilitated decision-making in domains such as education or hiring. To investigate potential LLM performance disparities, we design two benchmark tasks that reflect real-world scenarios: regional term choice (prompting the LLM to name a described item which is referred to differently in Mainland China and Taiwan), and regional name choice (prompting the LLM to choose who to hire from a list of names in both Simplified and Traditional Chinese). For both tasks, we audit the performance of 11 leading commercial LLM services and open-sourced models -- spanning those primarily trained on English, Simplified Chinese, or Traditional Chinese. Our analyses indicate that biases in LLM responses are dependent on both the task and prompting language: while most LLMs disproportionately favored Simplified Chinese responses in the regional term choice task, they surprisingly favored Traditional Chinese names in the regional name choice task. We find that these disparities may arise from differences in training data representation, written character preferences, and tokenization of Simplified and Traditional Chinese. These findings highlight the need for further analysis of LLM biases; as such, we provide an open-sourced benchmark dataset to foster reproducible evaluations of future LLM behavior across Chinese language variants (this https URL). Figure showing that three different LLMs (GPT-4o, Qwen-1.5, and Taiwan-LLM) may answer a prompt about pineapples differently when asked in Simplified Chinese vs. Traditional Chinese.Figure showing that LLMs disproportionately answer questions about regional-specific terms (like the word for "pineapple," which differs in Simplified and Traditional Chinese) correctly when prompted in Simplified Chinese as opposed to Traditional Chinese.Figure showing that LLMs have high variance of adhering to prompt instructions, favoring Traditional Chinese names over Simplified Chinese names in a benchmark task regarding hiring.
1174
Reposted by dan bateyko
michael veale @michae.lv · 23/06/2025
I am at FAccT 2025 in Athens, feel free to grab me if you want to chat.
091