Sign in

dan bateyko

@dbateyko.bsky.social
132 followers 211 following 41 posts

maybe the hard stuff's inside, hidden — like bones, as opposed to an exoskeleton. @CornellInfoSci dbateyko.info

PostsRepliesMedia
dan bateyko @dbateyko.bsky.social · 23/09/2026
Reading Cornell’s report on The Future of the American University and thought something felt off. The report explicitly says AI “was not used to automatically write the report or pieces of it.” And yet Pangram flags an entire section—about scholarly integrity—as AI-generated.
010
dan bateyko @dbateyko.bsky.social · 08/09/2026
Local law still formally mandates racial segregation. In 2026. We built the largest dataset of local laws and used LLMs to search for discriminatory provisions across the nation. What we found should not still be on the books. hai.stanford.edu/policy/makin...
hai.stanford.edu
Making Local Law Legible: LLM-Assisted Detection of Discrimination | Stanford HAI
This brief demonstrates the potential of LLM-assisted review to accelerate legal reform.
2127
dan bateyko @dbateyko.bsky.social · 07/09/2026
just finding out that the Python data visualization library seaborn is named after West Wing's Sam Seaborn and imported as sns due to the monogram on his shirt
000
dan bateyko @dbateyko.bsky.social · 24/07/2026
Early results, but! I’ve scraped 100,000+ law review articles (maybe~2x the next-largest study) and am using language models to classify their footnotes. Having fun seeing how tech law scholarship shakes out. So far, STS is overrepresented; sociology is slightly under.
001
Reposted by dan bateyko
Data & Society @datasociety.bsky.social · 23/07/2026
To make govt services more accessible & accountable, @baricks.bsky.social, @boston.gov’s aleja jimenez jaramillo & @megyoung0.bsky.social call on state & local govts to explore collectively-governed “in-house” alternatives to commercial language translation tech. datasociety.net/research-lib...
152
Reposted by dan bateyko
James Grimmelmann @jtlg.bsky.social · 20/07/2026
The 16th Edition of my Internet Law casebook is out! Overview: internetcasebook.com Table of contents: www.semaphorepress.com/downloads/In... PDF download (suggested $30): semaphorepress.com/InternetLaw_... Print-on-demand ($75.10): www.amazon.com/dp/1943689245 A brief thread on the update:
The cover of Internet Law: Cases and Problems (16th ed.) by James Grimmelmann.

The cover is orange and features a photograph of hundreds of screens glowing in the dark at the Assembly demoscene LAN party.
1229
Reposted by dan bateyko
Joachim Baumann @joachimbaumann.bsky.social · 07/07/2026
Just arrived at ICML 🇰🇷😍 Get up early tomorrow to hear me talk about how (not) to solve the peer review crisis, or find me at one of my poster presentations. Paper links: ✅ AI Peer Review: arxiv.org/abs/2605.03202 ✅ SWE-chat: arxiv.org/pdf/2604.20779
ICML Conference schedule with two papers. Left, "Stop Automating Peer Review Without Rigorous Evaluation": Oral presentation, Wed 7/8/2026, 10:00–10:15 AM KST, Grand Ballroom 101–105; Poster, Wed 7/8/2026, 2:30–4:15 PM KST, Hall A #3003. Right, "SWE-chat": at the 5th Deep Learning for Code Workshop, Fri 7/10/2026, 13:00–14:30 KST, Hall B2.
0162
Reposted by dan bateyko
Ira Globus-Harris @iraglobusharris.bsky.social · 03/07/2026
Are you at ICML next week? Feel like your decision-making for which sessions to attend might not be risk minimizing? Don't incur (swap) regret and come to my, @aaroth.bsky.social, and @ncollina.bsky.social's tutorial Monday on multicalibration, decision-making, and collaborative learning!
22011
Reposted by dan bateyko
Lauren Chambers @laurenmarietta.bsky.social · 28/06/2026
it's the last day of #FAccT2026 in Montreal (bonjour hiii), and I'm presenting my paper with Diag Davenport on the promise of #PublicInterestTech clinics for training the next gen of critical sociotechnical thinkers. 🏆 plus we got an honorable mention!? come thru @ 10:45! 📄 tiny.cc/pit-clinics-26
screenshot of a title slide, navy background with light blue icons and white and gold text: "a decision-making pedagogy for the public interest technology clinic: putting the ‘practice’ in critical technical practice." lauren m. chambers & diag davenport, uc berkeley. facct @ montreal, june 28, 2026
1217
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 23/06/2026
I'm so excited to attend #FAccT2026 in 🇨🇦Montreal🇨🇦 to present "Tradeoffs are Domain Dependent: Improving Accuracy and Fairness in Property Tax Assessments" by Evelyn Smith, me, Chris Berry, @jacobsgoldin.bsky.social, and Dan Ho!! 🔗: arxiv.org/pdf/2605.15020
A screenshot of our paper: 

Title: Tradeoffs are Domain Dependent: Improving Accuracy and Fairness in Property Tax Assessments
Authors: EVELYN SMITH and EMMA HARVEY (co-first authors), CHRISTOPHER BERRY, JACOB GOLDIN and DANIEL E. HO (co-senior authors)
Abstract: Algorithmic fairness research often assumes a tradeoff between fairness and accuracy. Yet this tradeoff may not be universal. We test this assumption in the context of U.S. property tax assessment - a setting in which the output of predictive algorithms directly determines the distribution of tax obligations among homeowners. Currently, systematic assessment errors cause owners of lower-valued properties to face disproportionately high tax burdens, creating regressivity in the property tax system. Using data on 26 million property sales spanning 95% of U.S. counties, we conduct three complementary analyses. First, we find that assessment accuracy and fairness - measured using domain-relevant metrics - are strongly correlated across counties under status quo practices. Second, in simulated assessment models, we show that adding property features improves accuracy in most cases, and that when accuracy improves, fairness almost always improves as well. Third, we show that incorporating publicly available Census data into assessment models - a feasible reform in most counties - would significantly improve both accuracy and fairness relative to status quo assessments. Together, these results challenge the presumed universality of the fairness-accuracy tradeoff and demonstrate that well-designed modeling improvements can advance both fairness and accuracy in large-scale public sector systems.
1113
Reposted by dan bateyko
Isabel Silva Corpus @isabelcorpus.bsky.social · 24/06/2026
Excited to attend FAccT 2026 in Montreal this week! Let me know if you'll be there and want to catch up :) I'll be presenting a paper with @allisonkoe.bsky.social about ad delivery skew in the context of government advertising, come by and check it out!
2257
Reposted by dan bateyko
Jennah Gosciak @jennahgosciak.bsky.social · 25/06/2026
I am excited to be at @facct.bsky.social this year presenting a new 📝 "Scrutinizing Index-Based Risk Assessments: A Case Study in NYC Decision-making for Heat Emergency Management" (work with Luke Boyce, @angelinawang.bsky.social , and @allisonkoe.bsky.social ). 🔗: dl.acm.org/doi/10.1145/... (1/10)
1196
Reposted by dan bateyko
Lindsey Barrett @lambarrett.bsky.social · 30/05/2026
outrageously corny how happy and grateful I feel after PLSC, every single time. man do I love my job and the lovely geniuses it gives me proximity to. privacy odd ducks conference forever!!!!
121
Reposted by dan bateyko
Lucy Li @lucy3.bsky.social · 31/03/2026
Another one of @ahalterman.bsky.social and @katakeith.bsky.social's papers that I think should be cited more by CSS researchers: What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification arxiv.org/abs/2510.03541
arxiv.org
What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
Generative large language models (LLMs) are now used extensively for text classification in computational social science (CSS). In this work, focus on the steps before and after LLM prompting -- conce...
1315
dan bateyko @dbateyko.bsky.social · 27/03/2026
Wild Anna’s Archive bounty, reads like a heist. They want someone to front tens of thousands of dollars to buy Library of Congress files for a 3k bounty. I imagine a leak investigation would have a very short suspect list
010
Reposted by dan bateyko
Peter Henderson @peterhenderson.bsky.social · 27/03/2026
Btw, did a bit of a rebranding of the substack. Will endeavor to post more there. h/t @dbateyko.bsky.social on the Trials & Errors name. Super fitting name for a group whose focus is both in reinforcement learning and in law/governance research. www.trialserrors.ai
trialserrors.ai
Trials & Errors | Peter Henderson | Substack
Various news, thoughts, and findings on the intersection of law, policy, and artificial intelligence. Click to read Trials & Errors, by Peter Henderson, a Substack publication with hundreds of subscri...
131
Reposted by dan bateyko
Maria Antoniak @mariaa.bsky.social · 24/03/2026
I've started a "History of NLP" repo to store all of these resources. I don't have time to add everything yet, but I'll keep chipping away, and help is welcome. github.com/maria-antoni...
github.com
GitHub - maria-antoniak/history-of-nlp: a public, crowd-sourced bibliography about the history of natural language processing (nlp)
a public, crowd-sourced bibliography about the history of natural language processing (nlp) - maria-antoniak/history-of-nlp
1366
Reposted by dan bateyko
Dominik Stammbach @dominsta.bsky.social · 28/01/2026
📣 Call for Contributions: LEXam-v2 – A Benchmark for Legal Reasoning in AI How well do today’s AI systems really reason about law? We’re building a global benchmark based on real law school & bar exams. 🧵 Full details, scope, and how to contribute in the thread 👇
164
Reposted by dan bateyko
J. Nathan Matias @natematias.bsky.social · 15/11/2025
One joy of growing as a scholar & doer has been the pleasure of being supported to pay attention to other people’s excellent work and amplify it. Next week I am publishing an article summarizing over 170 articles on AI + science + policy and tomorrow I get to email their authors to say thanks <3
1123
dan bateyko @dbateyko.bsky.social · 02/09/2025
110
dan bateyko @dbateyko.bsky.social · 02/09/2025
It's remarkable how early Ford Foundation was to law and technology in the midcentury
020
dan bateyko @dbateyko.bsky.social · 02/09/2025
010
dan bateyko @dbateyko.bsky.social · 01/09/2025
this is what you see moments before going down a cyberspace and law rabbit hole
110
dan bateyko @dbateyko.bsky.social · 30/08/2025
100
dan bateyko @dbateyko.bsky.social · 01/08/2025
it finally happened (my 3090 overheated and emergency shut off)
010
Reposted by dan bateyko
Daphne Keller @daphnek.bsky.social · 01/08/2025
Here’s the article. I’ve had more positive feedback on it than things I spent a year on. Apparently describing a problem that thousands of Trust and Safety people are seeing but also see the world ignoring is a good way to win hearts and minds :) www.lawfaremedia.org/article/the-...
lawfaremedia.org
The Rise of the Compliant Speech Platform
Content moderation is becoming a “compliance function,” with trust and safety operations run like factories and audited like investment banks.
13010
dan bateyko @dbateyko.bsky.social · 18/07/2025
At the Kernel 5 issue launch!
130
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 14/07/2025
After having such a great time at #CHI2025 and #FAccT2025, I wanted to share some of my favorite recent papers here! I'll aim to post new ones throughout the summer and will tag all the authors I can find on Bsky. Please feel welcome to chime in with thoughts / paper recs / etc.!! 🧵⬇️:
25510
Reposted by dan bateyko
jay @kuppermann.xyz · 15/07/2025
i am launching a magazine with @kevinbaker.bsky.social and the rest of the reboot collective on thursday at gray area! you should be there! open.substack.com/pub/reboothq...
Kernel Magazine
Issue 5: Rules Launch Party
July 17
6:30-9pm
SF
Gray Area

The illustration behind this text is a mix of gold chess pieces and purple snakes on a chessboard pattern

lu.ma/k5-sf
0144
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 01/07/2025
I've arrived in the 🌁Bay Area🌁, where I'll be spending the summer as a research fellow at Stanford's RegLab! If you're also here, LMK and let's get a meal / go on a hike / etc!!
A close-up view of the Golden Gate Bridge in the fog
091
Reposted by dan bateyko
James Grimmelmann @jtlg.bsky.social · 02/07/2025
Well, this was a nice surprise. Download it while it’s hot! lsolum.typepad.com/legaltheory/...
lsolum.typepad.com
Grimmelmann, Sobel, & Stein on Generative AI and Legal Interpretation
James Grimmelmann (Cornell Law School; Cornell Tech), Benjamin Sobel (Cornell University - Cornell Tech NYC), & David Stein (Vanderbilt University - Vanderbilt Law School) have posted Generative Misin...
041
Reposted by dan bateyko
Jennah Gosciak @jennahgosciak.bsky.social · 24/06/2025
I am presenting a new 📝 “Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments” at @facct.bsky.social on Thursday, with @aparnabee.bsky.social, Derek Ouyang, @allisonkoe.bsky.social, @marzyehghassemi.bsky.social, and Dan Ho. 🔗: arxiv.org/abs/2506.13735 (1/n)
"Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments"

Conducting disparity assessments at regular time intervals is critical for surfacing potential biases in decision-making and improving outcomes across demographic groups. Because disparity assessments fundamentally depend on the availability of demographic information, their efficacy is limited by the availability and consistency of available demographic identifiers. While prior work has considered the impact of missing data on fairness, little attention has been paid to the role of delayed demographic data. Delayed data, while eventually observed, might be missing at the critical point of monitoring and action -- and delays may be unequally distributed across groups in ways that distort disparity assessments. We characterize such impacts in healthcare, using electronic health records of over 5M patients across primary care practices in all 50 states. Our contributions are threefold. First, we document the high rate of race and ethnicity reporting delays in a healthcare setting and demonstrate widespread variation in rates at which demographics are reported across different groups. Second, through a set of retrospective analyses using real data, we find that such delays impact disparity assessments and hence conclusions made across a range of consequential healthcare outcomes, particularly at more granular levels of state-level and practice-level assessments. Third, we find limited ability of conventional methods that impute missing race in mitigating the effects of reporting delays on the accuracy of timely disparity assessments. Our insights and methods generalize to many domains of algorithmic fairness where delays in the availability of sensitive information may confound audits, thus deserving closer attention within a pipeline-aware machine learning framework.Figure contrasting a conventional approach to conducting disparity assessments, which is static, to the analysis we conduct in this paper. Our analysis (1) uses comprehensive health data from over 1,000 primary care practices and 5 million patients across the U.S., (2) timestamped information on the reporting of race to measure delay, and (3) retrospective analyses of disparity assessments under varying levels of delay.
1134
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 23/06/2025
I am so excited to be in 🇬🇷Athens🇬🇷 to present "A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms" by me, @kizilcec.bsky.social, and @allisonkoe.bsky.social, at #FAccT2025!! 🔗: arxiv.org/pdf/2506.04419
A screenshot of our paper's:

Title: A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
Authors: Emma Harvey, Rene Kizilcec, Allison Koenecke
Abstract: Increasingly, individuals who engage in online activities are expected to interact with large language model (LLM)-based chatbots. Prior work has shown that LLMs can display dialect bias, which occurs when they produce harmful responses when prompted with text written in minoritized dialects. However, whether and how this bias propagates to systems built on top of LLMs, such as chatbots, is still unclear. We conduct a review of existing approaches for auditing LLMs for dialect bias and show that they cannot be straightforwardly adapted to audit LLM-based chatbots due to issues of substantive and ecological validity. To address this, we present a framework for auditing LLM-based chatbots for dialect bias by measuring the extent to which they produce quality-of-service harms, which occur when systems do not work equally well for different people. Our framework has three key characteristics that make it useful in practice. First, by leveraging dynamically generated instead of pre-existing text, our framework enables testing over any dialect, facilitates multi-turn conversations, and represents how users are likely to interact with chatbots in the real world. Second, by measuring quality-of-service harms, our framework aligns audit results with the real-world outcomes of chatbot use. Third, our framework requires only query access to an LLM-based chatbot, meaning that it can be leveraged equally effectively by internal auditors, external auditors, and even individual users in order to promote accountability. To demonstrate the efficacy of our framework, we conduct a case study audit of Amazon Rufus, a widely-used LLM-based chatbot in the customer service domain. Our results reveal that Rufus produces lower-quality responses to prompts written in minoritized English dialects.
13110
Reposted by dan bateyko
Allison Koenecke @allisonkoe.bsky.social · 22/06/2025
🎉Excited to present our paper tomorrow at @facct.bsky.social, “Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese”, with @brucelyu17.bsky.social, Jiebo Luo and Jian Kang, revealing 🤖 LLM performance disparities. 📄 Link: arxiv.org/abs/2505.22645
"Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese" Abstract:

While the capabilities of Large Language Models (LLMs) have been studied in both Simplified and Traditional Chinese, it is yet unclear whether LLMs exhibit differential performance when prompted in these two variants of written Chinese. This understanding is critical, as disparities in the quality of LLM responses can perpetuate representational harms by ignoring the different cultural contexts underlying Simplified versus Traditional Chinese, and can exacerbate downstream harms in LLM-facilitated decision-making in domains such as education or hiring. To investigate potential LLM performance disparities, we design two benchmark tasks that reflect real-world scenarios: regional term choice (prompting the LLM to name a described item which is referred to differently in Mainland China and Taiwan), and regional name choice (prompting the LLM to choose who to hire from a list of names in both Simplified and Traditional Chinese). For both tasks, we audit the performance of 11 leading commercial LLM services and open-sourced models -- spanning those primarily trained on English, Simplified Chinese, or Traditional Chinese. Our analyses indicate that biases in LLM responses are dependent on both the task and prompting language: while most LLMs disproportionately favored Simplified Chinese responses in the regional term choice task, they surprisingly favored Traditional Chinese names in the regional name choice task. We find that these disparities may arise from differences in training data representation, written character preferences, and tokenization of Simplified and Traditional Chinese. These findings highlight the need for further analysis of LLM biases; as such, we provide an open-sourced benchmark dataset to foster reproducible evaluations of future LLM behavior across Chinese language variants (this https URL). Figure showing that three different LLMs (GPT-4o, Qwen-1.5, and Taiwan-LLM) may answer a prompt about pineapples differently when asked in Simplified Chinese vs. Traditional Chinese.Figure showing that LLMs disproportionately answer questions about regional-specific terms (like the word for "pineapple," which differs in Simplified and Traditional Chinese) correctly when prompted in Simplified Chinese as opposed to Traditional Chinese.Figure showing that LLMs have high variance of adhering to prompt instructions, favoring Traditional Chinese names over Simplified Chinese names in a benchmark task regarding hiring.
1174
Reposted by dan bateyko
michael veale @michae.lv · 23/06/2025
I am at FAccT 2025 in Athens, feel free to grab me if you want to chat.
091
Reposted by dan bateyko
Anna Neumann @annaneumann.bsky.social · 23/06/2025
Please come see us at the RC Trust Networking Event! You can sign up with the QR Codes around the venues and get some free drinks! 🙂‍↕️ #FAccT2025
143
Reposted by dan bateyko
Princeton Center for Information Technology Policy @princetoncitp.bsky.social · 06/06/2025
New paper available - "Bureaucratic Backchannel: How r/PatentExaminer Navigates #AI Governance" which investigates how examiners navigate dual roles through a qualitative analysis of a Reddit community where U.S. Patent & Trademark Office employees discuss their work🔗📜👇
121
Reposted by dan bateyko
A. Feder Cooper @afedercooper.bsky.social · 21/05/2025
Llama 3.1 70B contains copies of nearly the entirety of some books. Harry Potter is just one of them. I don’t know if this means it’s an infringing copy. But the first question to answer is if it’s a copy at all/in the first place. That’s what our new results suggest: arxiv.org/abs/2505.12546
arxiv.org
Extracting memorized pieces of (copyrighted) books from open-weight language models
Plaintiffs and defendants in copyright lawsuits over generative AI often make sweeping, opposing claims about the extent to which large language models (LLMs) have memorized plaintiffs' protected expr...
45324
Reposted by dan bateyko
Peter Henderson @peterhenderson.bsky.social · 21/05/2025
Another hallucinated citation in court. At this point, our tracker is up to ~70 cases worldwide of hallucinated citations in court, including hallucinations from 2 adjudicators. New Case: storage.courtlistener.com/recap/gov.us... Tracker: www.polarislab.org/ai-law-track...
063
Reposted by dan bateyko
The Privacy Center @georgetownprivacy.bsky.social · 15/05/2025
Three years ago, we released “American Dragnet: Data-Driven Deportation in the 21st Century.” The report describes the surveillance apparatus that Trump is using to target immigrants, activists and anyone else who challenges his agenda. We’re re-releasing it today with a new foreword.
americandragnet.org
American Dragnet | Data-Driven Deportation in the 21st Century
One of two American adults is in a law enforcement face recognition database. An investigation.
12014
Reposted by dan bateyko
🌶 David Gray Widder @davidthewid.bsky.social · 01/05/2025
Had a great time presenting this paper, cowritten with Sireesh Gururaja and Lucy Suchman! Paper draft here: arxiv.org/abs/2411.17840
052
Reposted by dan bateyko
Daphne Keller @daphnek.bsky.social · 25/04/2025
OK this keeps getting better. It’s not just that the FTC is moderating content uploaded in user comments as part of its “platform censorship” inquiry. It’s re-moderating the same content the platforms moderated, as @corbinkbarthold.bsky pointed out. x.com/corbinkbarth... 1/
x.com
13217
Reposted by dan bateyko
J. Nathan Matias @natematias.bsky.social · 29/04/2025
From concerns about social media addiction to urgent civil liberties issues, courts are asking scientists to be arbiters of alleged technology harms. How can scientists reliably inform courts and how can courts interpret our work? New article with @penney.bsky.social tsjournal.org/index.php/jo...
tsjournal.org
Science and Causality in Technology Litigation | Journal of Online Trust and Safety
Journal of Online Trust and Safety
22310
Reposted by dan bateyko
J. Nathan Matias @natematias.bsky.social · 10/04/2025
Why do scientists still struggle to answer basic questions about the safety of digital tech from AI to social media, even as families point to rising evidence of individual harm with concern & grief? @orbenamy.bsky.social & I have a new article in @science.org: www.science.org/doi/10.1126/...
science.org
Fixing the science of digital technology harms
Technology development outpaces scientific assessment of impacts
37936
Reposted by dan bateyko
Alondra Nelson @alondra.bsky.social · 02/04/2025
"surveillance deputies"(Brayne, Lageson & Levy, 2023).
25822
Reposted by dan bateyko
Evan Peck @peck.phd · 20/11/2024
Trying something new: A 🧵 on a topic I find many students struggle with: "why do their 📊 look more professional than my 📊?" It's *lots* of tiny decisions that aren't the defaults in many libraries, so let's break down 1 simple graph by @jburnmurdoch.bsky.social 🔗 www.ft.com/content/73a1...
921577458
Reposted by dan bateyko
Raj Movva @rajmovva.bsky.social · 18/03/2025
@kennypeng.bsky.social also built a website to explore results on Yelp, headlines, & Congress datasets: hypothesaes.org. You can see every SAE neuron in UMAP space, colored by whether the neuron correlates positively or negatively with the target variable. 8/
141
Reposted by dan bateyko
Raj Movva @rajmovva.bsky.social · 18/03/2025
💡New preprint & Python package: We use sparse autoencoders to generate hypotheses from large text datasets. Our method, HypotheSAEs, produces interpretable text features that predict a target variable, e.g. features in news headlines that predict engagement. 🧵1/
14013
Reposted by dan bateyko
WIRED @wired.com · 18/03/2025
They're called public records for a reason. Starting today, WIRED will *stop paywalling* articles that are primarily based on public records obtained through the Freedom of Information Act, becoming the first publication to partner with @freedom.press to offer this for our new coverage.
freedom.press
Wired is dropping paywalls for FOIA-based reporting. Others should follow
As the administration does its best to hide public records from the public, Wired magazine is stepping up to help stem the secrecy
16249137923360
Reposted by dan bateyko
Emma Harvey @emmharv.bsky.social · 13/03/2025
✨New Work✨ by me, @allisonkoe.bsky.social, and @kizilcec.bsky.social forthcoming at #CHI2025: "Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education 🔗: arxiv.org/pdf/2502.14592
A screenshot of our paper:

Title: “Don’t Forget the Teachers”: Towards an Educator-Centered Understanding of Harms from Large Language Models in Education

Authors: Emma Harvey, Allison Koenecke, Rene Kizilcec

Abstract: Education technologies (edtech) are increasingly incorporating new features built on LLMs, with the goals of enriching the processes of teaching and learning and ultimately improving learning outcomes. However, it is still too early to understand the potential downstream impacts of LLM-based edtech. Prior attempts to map the risks of LLMs have not been tailored to education specifically, even though it is a unique domain in many respects: from its population (students are often children, who can be especially impacted by technology) to its goals (providing the ‘correct’ answer may be less important than understanding how to arrive at an answer) to its implications for higher-order skills that generalize across contexts (e.g. critical thinking and collaboration). We conducted semi-structured interviews with six edtech providers representing leaders in the K-12 space, as well as a diverse group of 23 educators with varying levels of experience with LLM-based edtech. Through a thematic analysis, we explored how each group is anticipating, observing, and accounting for potential harms from LLMs in education. We find that, while edtech providers focus primarily on mitigating technical harms, i.e. those that can be measured based solely on LLM outputs themselves, educators are more concerned about harms that result from the broader impacts of LLMs, i.e. those that require observation of interactions between students, educators, school systems, and edtech to measure. Overall, we (1) develop an education-specific overview of potential harms from LLMs, (2) highlight gaps between conceptions of harm by edtech providers and those by educators, and (3) make recommendations to facilitate the centering of educators in the design and development of edtech tools.
1548