Sign in

Kenny Peng

@kennypeng.bsky.social
239 followers 258 following 111 posts

CS PhD student at Cornell Tech. Interested in interactions between algorithms and society. Princeton math '22. kennypeng.me

PostsRepliesMedia
Reposted by Kenny Peng
Wesley Finck @wesleyfinck.org · 26/08/2026
Some preliminary product insights from our recent onboarding launch: It's no coincidence that the most popular suggested link is "From Feeds to Trails". People are ready for a new kind of social media experience! leaflet.pub/p/did:plc:ld...
leaflet.pub
From Feeds to Trails
To design the future of social media, rethink the interface before the algorithm.
3245
Reposted by Kenny Peng
Erica Chiang @ericachiang.bsky.social · 19/08/2026
Thanks to Grace Stanley for covering our work :) Our paper on personalized recommendations for students applying to NYC high schools (best student paper at EC ’26(!!)) is coming soon! Here's a quick preview... (1/5) news.cornell.edu/stories/2026...
t.co
https://news.cornell.edu/stories/2026/08/research-helps-nyc-students-aim-higher-public-high-school-applications
1114
Kenny Peng @kennypeng.bsky.social · 19/08/2026
Here's an article about our broader effort to help students navigate the NYC school match. The big payoff has been to deploy real nudges (led by @ericachiang.bsky.social!). As a result, actual students this year matched to high-quality schools they wouldn't have applied to otherwise!
news.cornell.edu
Research helps NYC students aim higher in public high school applications | Cornell Chronicle
Many New York City eighth-graders – particularly those from underserved communities – aren’t applying to academically competitive high schools where they could succeed, according to new Cornell resear...
180
Kenny Peng @kennypeng.bsky.social · 18/08/2026
New York City 8th graders choose from 900+ high schools to apply to, in a process that’s spawned Facebook groups and dozens of expensive consulting services. Our new paper shows how application behavior leads to disparities, and how to effectively intervene. 🧵 www.nature.com/articles/s44...
1329
Reposted by Kenny Peng
Benjamin Laufer @laufer.bsky.social · 04/08/2026
I am delighted to share that I'll be joining the University of Washington Information School as an Assistant Professor!
University of Washington campus
2442
Reposted by Kenny Peng
Maria Antoniak @mariaa.bsky.social · 30/07/2026
Wonderful talk by @sjgreenwood.bsky.social about @paper-feed.bsky.social! Custom feeds on Bluesky are unique resources for researchers interested in recommendation algorithms #IC2S2 Featuring many Bluesky friends like @graze.social @devingaffney.com @kissane.myatproto.social @aendra.com
24412
Reposted by Kenny Peng
Benjamin Laufer @laufer.bsky.social · 21/07/2026
One intuition behind many AI policy proposals is that downstream AI applications -- the companies deploying AI in healthcare, finance, education, customer service, etc. -- should bear responsibility for ensuring safety. Our paper asks: What incentives does that create for the firms building AI?
Illustrative example of our game-theoretic model. This numeric instance of the game consists of one general-purpose producer and three domain-specialists. Each player has a different utility in performance-safety space which dictates the path of development. Within this setting, the no-regulation game (Upper Left) reveals the players’ investment efforts when no floor is imposed on safety. Regulating the domain-specialist alone (Upper Right) exhibits backfiring for all three domains, meaning the regulated safety level is lower than it would be without regulation. In this particular example, the same floor is assumed for all three domain-specialists. Regulating the generalist alone (Lower Left) improves the safety level slightly across all three domains, compared to no-regulation. Finally, a regime that targets both generalists and specialists with regulation (Lower Right) is able to 1) retain the improved safety performance from regulating the generalist, 2) improve the safety level of least-safe domain-specialist, while 3) avoiding backfiring. The purpose of this figure is to visualize the model’s incentive mechanisms; none of these panels represent real regulations.
193
Kenny Peng @kennypeng.bsky.social · 06/07/2026
🧵Can we reconcile excitement for SAEs with negative results? Our #ICML2026 position paper argues that even if SAEs underperform baselines when acting on knowns (e.g. probing, steering), they're a powerful tool for ~discovering unknowns~ Poster: Tue 2pm, HALL A #1715
2125
Reposted by Kenny Peng
Ira Globus-Harris @iraglobusharris.bsky.social · 03/07/2026
Are you at ICML next week? Feel like your decision-making for which sessions to attend might not be risk minimizing? Don't incur (swap) regret and come to my, @aaroth.bsky.social, and @ncollina.bsky.social's tutorial Monday on multicalibration, decision-making, and collaborative learning!
22011
Kenny Peng @kennypeng.bsky.social · 30/06/2026
Presenting this at #COLT2026! Let me know if you want to chat about (1) the linear representation hypothesis, (2) theoretical foundations for interpretability / science of AI. - Talk: Thursday 2:33 pm - Deep Learning Theory / Optimization - Poster: Thursday Lunch)
081
Kenny Peng @kennypeng.bsky.social · 22/06/2026
In college, I wrote this essay about Clark Kimberling, who maintains the Encyclopedia of Triangle Centers. The essay covers triangle geometry's rise (which includes Napoleon), its fall (from "brute force" methods), and its continued following ("the melody lingers on")
leaflet.pub
Coincidence of Lines
The mathematician Clark Kimberling estimates that he spends between two and five hours each day, seven days a week, maintaining his online encyclopedia of triangle centers.
161
Reposted by Kenny Peng
Nikhil Garg @nkgarg.bsky.social · 11/05/2026
We are very excited to announce our first workshop on From Theory to Practice: behind the scenes on research deployments at EC’26 (July 6 in Rome)! Call for posters and submissions now open! Organized by myself, @ericachiang.bsky.social , Bailey Flanigan, @brwilder.bsky.social
sites.google.com
Home
About This workshop will focus on the practical realities of deploying algorithmic and economic systems from academic research, especially with government and non-profit partners. While economics and ...
0216
Reposted by Kenny Peng
spacecowboy @spacecowboy17.bsky.social · 26/04/2026
Two Bluesky client apps have integrated the "also liked" functionality: - @boostblue.bsky.social - iOS and Android - @skywalker.thereforeiam.eu - Android The app integration makes it a much smoother experience than the stand-alone website foryou.club/also-liked Here is how it looks in Boost Blue:
3647
Kenny Peng @kennypeng.bsky.social · 24/04/2026
We made traversle.io, a new daily word game! The goal is to traverse from a start word to a target word through a network of related words. (Our motivating question: is it possible to construct a network that allows human navigation?)
3164
Kenny Peng @kennypeng.bsky.social · 18/04/2026
Very cool — a way to jump from one post onto a trail of related posts. Lots of possibility and opportunity to push beyond the feed interface!
030
Reposted by Kenny Peng
Sansa Sandisk @cheatlines.co · 17/04/2026
bluesky is dying? Come to the Wait and See feed
3122
Reposted by Kenny Peng
brendan @schlage.town · 17/04/2026
so cool! over 20k trails to explore…and that's just based on a single week of bluesky posts! fun to search trails & browse related ones…seems like a great way to find new folks to follow & prob closest thing I've seen to "collections" as @danabra.mov has been advocating for
skytrails.org
skytrails · 20,000 trails through Bluesky
Regaining freedom of movement on social media. Browse Bluesky via interconnected trails.
2334
Reposted by Kenny Peng
Rafael M Batista @rafmbatista.bsky.social · 16/04/2026
Loved reading this. The premise that social media should let us actively traverse it through connections, not just scroll a feed feels exactly right. One thought it raised for me: Trails (& rabbit holes) are often spaces we wander through alone. What would a more social analogy get us?
181
Reposted by Kenny Peng
Ira Globus-Harris @iraglobusharris.bsky.social · 15/04/2026
I have really enjoyed talking with Kenny about his work on alternative rec system mechanisms. This work offers a vision of user-powered concept exploration: imagine the social media feed equivalent of the joy of exploring wikipedia by navigating trails of hyperlinks from one article to the next.
193
Kenny Peng @kennypeng.bsky.social · 15/04/2026
🧵 On social media, we feel trapped in “filter bubbles” and “echo chambers.” Often, we see this as a problem with the algorithm. Our new essay argues the problem is the feed interface itself: it constrains our movement—all we can do is scroll. We lay out an alternative vision:
leaflet.pub
From Feeds to Trails
To design the future of social media, rethink the interface before the algorithm.
910124
Kenny Peng @kennypeng.bsky.social · 26/03/2026
Excited to share our new research demo, where you can freely traverse the world of Bluesky through 20,000 interconnected trails, spanning “analysis of fictional tropes” to “rotisserie chicken” to “zoning and land use policy.” Try it out, and let us know what you think!
skytrails.org
skytrails · 20,000 trails through Bluesky
Can we regain freedom of movement on social media? Browse Bluesky via interconnected trails.
6459
Reposted by Kenny Peng
Divya Shanmugam @dmshanmugam.bsky.social · 23/03/2026
New in Nature Health: how might we move towards a world in which race is not used in clinical algorithms? We need (1) careful comparison of race-aware and race-neutral algorithms and (2) systemic efforts to address underlying disparities.
1219
Reposted by Kenny Peng
Gabriel Agostini @gsagostini.bsky.social · 18/03/2026
Had a great time presenting our work on building MIGRATE–a new dataset of US migration–at the @geographers.bsky.social AAG Annual Meeting today. Happy to also share that we received an AAG student paper award for this work!!! Come chat if you are at #AAG26 this week. migrate.tech.cornell.edu
0123
Reposted by Kenny Peng
Emma Pierson @emmapierson.bsky.social · 26/02/2026
Our paper, "What's in My Human Feedback", received an oral presentation at ICLR! Our method automatically+interpretably identifies preferences in human feedback data; we use this to improve personalization + safety. Reach out if you have data/use cases to apply this to! arxiv.org/pdf/2510.26202
0283
Kenny Peng @kennypeng.bsky.social · 17/02/2026
New paper! The Linear Representation Hypothesis is a powerful intuition for how language models work, but lacks formalization. We give a mathematical framework in which we can ask and answer a basic question: how many features can be stored under the hypothesis? 🧵 arxiv.org/abs/2602.11246
14514
Reposted by Kenny Peng
Gabriel Agostini @gsagostini.bsky.social · 05/02/2026
Our paper “Inferring fine-grained migration patterns across the United States” is now out in @natcomms.nature.com! We released a new, highly granular migration dataset. 1/9
27327
Reposted by Kenny Peng
David Liu @david-m-liu.bsky.social · 28/01/2026
🎙️ I had a great time joining the Data Skeptic podcast to talk about my work on recommender systems If you're interested in embeddings, aligning group preferences, or music recommendations, check out the episode below 👇 open.spotify.com/episode/6IsP...
open.spotify.com
Fairness in PCA-Based Recommenders
1145
Reposted by Kenny Peng
sidhikabalachandar.bsky.social @sidhikabalachandar.bsky.social · 20/01/2026
Check out our new paper at #AAAI 2026! I’ll be presenting in Singapore at Saturday’s poster session (12–2pm). This is joint work with @shuvoms.bsky.social, @bergerlab.bsky.social, @emmapierson.bsky.social, and @nkgarg.bsky.social. 1/9
183
Reposted by Kenny Peng
Sophie Greenwood @sjgreenwood.bsky.social · 14/01/2026
Excited to present a new preprint with @nkgarg.bsky.social: presenting usage statistics and observational findings from Paper Skygest in the first six months of deployment! 🎉📜 arxiv.org/abs/2601.04253
Title + abstract of the preprint
417150
Reposted by Kenny Peng
Sophie Greenwood @sjgreenwood.bsky.social · 09/01/2026
so so so excited to present our research + connect with the #ATScience community 🧪🎉
1285
Kenny Peng @kennypeng.bsky.social · 24/12/2025
Year 3 of spending many days making gingerbread — this year, featuring the gantries of Long Island City
0122
Reposted by Kenny Peng
Isabel Silva Corpus @isabelcorpus.bsky.social · 01/12/2025
Excited to share a new working paper! What happened when Change.org integrated an AI writing tool into their platform? We provide causal evidence that petition text changed significantly while outcomes did not improve. 1/ arxiv.org/abs/2511.13949
arxiv.org
Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes
The rapid integration of AI writing tools into online platforms raises critical questions about their impact on content production and outcomes. We leverage a unique natural experiment on Change$.$org...
45518
Kenny Peng @kennypeng.bsky.social · 17/11/2025
I had a lot of fun making this map of Manhattan’s grid (only the numbered streets and avenues). Learned that 4th avenue doesn’t exist, but then learned that it actually does exist but only for a few blocks.
191
Reposted by Kenny Peng
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods.
1217
Kenny Peng @kennypeng.bsky.social · 14/10/2025
Being Divya's labmate (and fellow ferry commuter) has been a real pleasure, and I've learned a ton from both her research itself and her approach to research (and also from the other random things she knows about).
140
Kenny Peng @kennypeng.bsky.social · 12/08/2025
"those already relatively advantaged are, empirically, more able to pay time costs and navigate administrative burdens imposed by the mechanisms." This point by @nkgarg.bsky.social has greatly shaped my thinking about the role of computer science in public service settings.
041
Kenny Peng @kennypeng.bsky.social · 05/08/2025
How do we reconcile excitement about sparse autoencoders with negative results showing that they underperform simple baselines? Our new position paper makes a distinction: SAEs are very useful for tools for discovering *unknown* concepts, less good for acting on *known* concepts.
092
Kenny Peng @kennypeng.bsky.social · 30/07/2025
One paragraph pitch for why sparse autoencoders are cool (they learn *interpretable* text embeddings)
Text embeddings capture tons of information, but individual dimensions are uninterpretable. It would be great if each dimension reflected a concept (“dimension 12 is about cats”). But text embeddings are ~1000 dimensions and there are millions of human concepts. So we need a higher dimensional embedding. Now notice that while there are tons of human concepts, they appear *sparsely*—any piece of text invokes a tiny fraction of concepts. This motivates learning a sparse high-dimensional encoding of text embeddings. Turns out SAEs work great for this in practice, producing *interpretable text embeddings*.
063
Kenny Peng @kennypeng.bsky.social · 16/07/2025
We're presenting two papers Wednesday at #ICML2025, both at 11am. Come chat about "Sparse Autoencoders for Hypothesis Generation" (west-421), and "Correlated Errors in LLMs" (east-1102)! Short thread ⬇️
182
Kenny Peng @kennypeng.bsky.social · 03/07/2025
Are LLMs correlated when they make mistakes? In our new ICML paper, we answer this question using responses of >350 LLMs. We find substantial correlation. On one dataset, LLMs agree on the wrong answer ~2x more than they would at random. 🧵(1/7) arxiv.org/abs/2506.07962
Heat map showing that more accurate models have more correlated errors.
1497
Reposted by Kenny Peng
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
New work 🎉: conformal classifiers return sets of classes for each example, with a probabilistic guarantee the true class is included. But these sets can be too large to be useful. In our #CVPR2025 paper, we propose a method to make them more compact without sacrificing coverage.
A gif explaining the value of test-time augmentation to conformal classification. The video begins with an illustration of TTA reducing the size of the  predicted set of classes for a dog image, and goes on to explain that this is because TTA promotes the true class's predicted probability to be higher, even when it's predicted to be unlikely.
3226
Reposted by Kenny Peng
Raj Movva @rajmovva.bsky.social · 05/05/2025
We'll present HypotheSAEs at ICML this summer! 🎉 Draft: arxiv.org/abs/2502.04382 We're continuing to cook up new updates for our Python package: github.com/rmovva/Hypot... (Recently, "Matryoshka SAEs", which help extract coarse and granular concepts without as much hyperparameter fiddling.)
github.com
GitHub - rmovva/HypotheSAEs: Hypothesizing interpretable relationships in text datasets using sparse autoencoders.
Hypothesizing interpretable relationships in text datasets using sparse autoencoders. - rmovva/HypotheSAEs
1102
Reposted by Kenny Peng
Erica Chiang @ericachiang.bsky.social · 01/05/2025
I’m really excited to share the first paper of my PhD, “Learning Disease Progression Models That Capture Health Disparities” (accepted at #CHIL2025)! ✨ 1/ 📄: arxiv.org/abs/2412.16406
33610
Reposted by Kenny Peng
Emma Pierson @emmapierson.bsky.social · 25/04/2025
The US government recently flagged my scientific grant in its "woke DEI database". Many people have asked me what I will do. My answer today in Nature. We will not be cowed. We will keep using AI to build a fairer, healthier world. www.nature.com/articles/d41...
nature.com
My ‘woke DEI’ grant has been flagged for scrutiny. Where do I go from here?
My work in making artificial intelligence fair has been noticed by US officials intent on ending ‘class warfare propaganda’.
13813
Kenny Peng @kennypeng.bsky.social · 02/04/2025
Our lab had a #dogathon 🐕 yesterday where we analyzed NYC Open Data on dog licenses. We learned a lot of dog facts, which I’ll share in this thread 🧵 1) Geospatial trends: Cavalier King Charles Spaniels are common in Manhattan; the opposite is true for Yorkshire Terriers.
25214
Reposted by Kenny Peng
Gabriel Agostini @gsagostini.bsky.social · 28/03/2025
Migration data lets us study responses to environmental disasters, social change patterns, policy impacts, etc. But public data is too coarse, obscuring these important phenomena! We build MIGRATE: a dataset of yearly flows between 47 billion pairs of US Census Block Groups. 1/5
54118
Reposted by Kenny Peng
Raj Movva @rajmovva.bsky.social · 18/03/2025
💡New preprint & Python package: We use sparse autoencoders to generate hypotheses from large text datasets. Our method, HypotheSAEs, produces interpretable text features that predict a target variable, e.g. features in news headlines that predict engagement. 🧵1/
14013
Kenny Peng @kennypeng.bsky.social · 18/03/2025
(1/n) New paper/code! Sparse Autoencoders for Hypothesis Generation HypotheSAEs generates interpretable features of text data that predict a target variable: What features predict clicks from headlines / party from congressional speech / rating from Yelp review? arxiv.org/abs/2502.04382
1145
Reposted by Kenny Peng
Sophie Greenwood @sjgreenwood.bsky.social · 10/03/2025
Please repost to get the word out! @nkgarg.bsky.social and I are excited to present a personalized feed for academics! It shows posts about papers from accounts you’re following bsky.app/profile/pape...
8171118