Reposted by Jacob EisensteinBryan Cantrill @bcantrill.bsky.social · 05/09/2026The revolt of the reader bcantrill.dtrace.org/2026/09/05/t...bcantrill.dtrace.orgThe revolt of the reader | The Observation Deck 1112431
Reposted by Jacob EisensteinVictor Geislinger @victorsvector.com · 13/08/2026Been testing this 'sign-language-to-text' (SL2T) for almost a year & it's so awesome to see this out! Looked at this almost a decade ago & knew we'd get the tech & ML models there eventually (but super challenging). Talking w/ some Deaf folks, an equivalent of 'voice-to-text' was huge for themdeepmind.googlePutting sign language AI into users’ handsIntroducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users. 210533
Reposted by Jacob EisensteinEytan Adar @eytan.adar.prof · 11/08/2026I wrote a post about the emergence of AI-native students who do their research using GenAI tools. I think that there are some real risks to job prospects and the likely changes in their fields. 1378
Jacob Eisenstein @jacobeisenstein.bsky.social · 08/07/2026a spontaneously-formed triple snaking queue to enter the #icml2026 poster session. the conference sold out before the early bird registration period ended. no shade whatsoever to the organizers, who are making the best of this, but we have strayed from light. conferences cannot be this big. 2322
Jacob Eisenstein @jacobeisenstein.bsky.social · 28/06/2026It’s a small, biased sample, but my AC batch for @colmweb.org was miles better than the papers i reviewed for NeurIPS: more creative, more rigorous, better written. Maybe AI research isn’t doomed to 2-3 omni-conferences after all?colmweb.orgCOLM 2026 1212
Reposted by Jacob EisensteinNathan Lambert @natolambert.bsky.social · 15/04/2026I spent some time trying to distill all the complex factors impacting open models -- economics, capabilities, distribution, policy, etc. -- into a clear list of beliefs. Here they are in full. www.interconnects.ai/p/my-bets-on...interconnects.aiMy bets on open models, mid-2026What I expect to come next and why, focused on the open-closed gap. 1274
Reposted by Jacob EisensteinTom Schaul @schaul.bsky.social · 16/03/2026DeepMind's RL team is hiring a research scientist: if you're passionate about RL, come work with us! And if you know people who might be interested, please share: job-boards.greenhouse.io/deepmind/job...job-boards.greenhouse.ioResearch Scientist, Reinforcement LearningLondon, UK 12914
Reposted by Jacob EisensteinJonathan Berant @jonathanberant.bsky.social · 06/03/2026Newish work (arXived in December): Prompts can be ambig., but handling ambiguity is context/user dependent. Sometimes the right thing is to ask a clarifying question, sometimes to give multi. answers, and sometimes to just guess. Can we train steerable models that change their strategy per context? 121
Jacob Eisenstein @jacobeisenstein.bsky.social · 04/03/2026Are AI models effective collaborators, or mere assistants awaiting your next command? (Preprint: arxiv.org/abs/2602.24188) To find out, we make AI collaborate with itself, in private information games: tasks that require sharing private information, like this chess board ordering task. 35621
Jacob Eisenstein @jacobeisenstein.bsky.social · 01/03/2026This looks like it’ll be a fantastic intro to transformers ⚡️ 1224
Reposted by Jacob EisensteinCaroline Wang @caroline-wang.bsky.social · 16/02/2026[1/n] Just wrapped up 7 months interning with @pcastr.bsky.social at Google DeepMind and I'm so excited to share our work: arxiv.org/abs/2602.10324. TLDR: We used LLM-powered program synthesis to automatically model and discover differences between human and LLM strategic behavior 28114
Reposted by Jacob EisensteinQuentin Berthet @ AISTATS @qberthet.bsky.social · 16/02/2026🚨 🔬 PhD positions at Google DeepMind in France 🇫🇷 We are advertising Master Level Intern positions at Google DeepMind within our Frontier AI Unit. These could lead to co-advised PhD positions with Google DeepMind and French academic institutions. job-boards.greenhouse.io/deepmind/job...job-boards.greenhouse.ioStagiaire de niveau Master, France / Master Level Intern, FranceGrenoble, France ; Paris, France 12917
Jacob Eisenstein @jacobeisenstein.bsky.social · 20/12/2025meanwhile, in the pnw www.seattletimes.com/seattle-news...seattletimes.comBe aware of toilet rats, King County saysRats could climb into your toilet as a result of recent flooding from atmospheric rivers, King County officials say. Enter "wet rat winter." 040
Reposted by Jacob EisensteinMarc Lanctot @sharky6000.bsky.social · 19/12/2025Hello! 👋 Are you interested in AI for board games using language models? Want to do some hobby tinkering with fine-tuning or RL? We've released an easy-to-follow example colab that fine-tunes Gemma models via Kauldron to mimic an MCTS player. Details here: github.com/google-deepm... ♟️🎲♦️♠️♥️♣️✨🎉github.com2025 Wrap-up: Fine-tuning Gemma with Kauldron Example ✦︎ · Issue #1414 · google-deepmind/open_spielHello everyone! We've been hard at work this year working on OpenSpiel 2.0, which will be better than ever. Major developments have been underway to make working with language models easier. I'm lo... 2407
Reposted by Jacob EisensteinGretchen McCulloch @gretchenmcculloch.com · 01/12/2025Day 1 of #BooksAreMyJam! Blueberry Maple jam, with Linguaphile: A life of language love by Julie Sedivy. A classic Canadian flavour duo + this book about @juliesedivy.bsky.social's relationship with language through her childhood in Montreal, later research as a linguist, and more 510811
Reposted by Jacob Eisensteinmr. TIM @timkellogg.me · 25/11/2025this is the theme — you can’t have AGI without existing in and learning from the real world 1212
Reposted by Jacob EisensteinKaitlyn Zhou @kaitlynzhou.bsky.social · 06/11/2025No better time to start learning about that #AI thing everyone's talking about... 📢 I'm recruiting PhD students in Computer Science or Information Science @cornellbowers.bsky.social! If you're interested, apply to either department (yes, either program!) and list me as a potential advisor! 2259
Jacob Eisenstein @jacobeisenstein.bsky.social · 17/10/2025knowing how to tie your shoes or order a drink in a crowded bar: not agi naming the big five personality traits: definitely agi 170
Jacob Eisenstein @jacobeisenstein.bsky.social · 09/10/2025Nicholas Carlini asking the right questions at #COLM2025 050
Reposted by Jacob EisensteinMaria Antoniak @mariaa.bsky.social · 06/10/2025Here’s a #COLM2025 feed! Pin it 📌 to follow along with the conference this week! 22617
Reposted by Jacob EisensteinMyra Cheng @myra.bsky.social · 03/10/2025AI always calling your ideas “fantastic” can feel inauthentic, but what are sycophancy’s deeper harms? We find that in the common use case of seeking AI advice on interpersonal situations—specifically conflicts—sycophancy makes people feel more right & less willing to apologize. 5544270
Reposted by Jacob EisensteinPete Shaw @ptshaw.bsky.social · 01/10/2025Excited to share a new paper that aims to narrow the conceptual gap between the idealized notion of Kolmogorov complexity and practical complexity measures for neural networks. 195
Reposted by Jacob EisensteinJoshua Raclaw @joshuaraclaw.com · 26/08/2025Cannot stress enough how good it is that you can come across a post about gorgeous little Yiddish book sitting in someone’s family collection, and within a few seconds you can find the full scanned version of the book available for free through the Yiddish Book Center’s websiteyiddishbookcenter.orgAsṭronomye | Yiddish Book Center 4437
Jacob Eisenstein @jacobeisenstein.bsky.social · 11/08/2025Baristas still safe from robotic automation, and not just because robots don’t know what coffee tastes like. prompt: “I’m trying to dial in this v60 of huatusco with my vario. temp / grind recommendations?” 130
Jacob Eisenstein @jacobeisenstein.bsky.social · 08/08/2025I’d guess that the majority position of syntacticians about LLMs (and other NLP beforehand) is roughly what Chomsky says: language tech can’t possibly teach us anything about the human language capability, so whether the LLM writes well doesn’t matter at all. 1100
Jacob Eisenstein @jacobeisenstein.bsky.social · 07/08/2025boston champaign pittsburgh atlanta, and, uh, let’s count seattle glad i did it, hope i don’t have to do it again 021
Reposted by Jacob EisensteinMargaret Mitchell @mmitchell.bsky.social · 06/08/2025🤖 ICYMI: Yesterday, @hf.co and OpenAI partnered to bring open source GPT to the public. This is a Big Deal in "AI world". Allow me to explain why. 🧵 huggingface.co/openai/gpt-o... 25217
Reposted by Jacob EisensteinIbn Bassal @ibnbassal.bsky.social · 29/07/2025roman burrito thread 315138
Reposted by Jacob EisensteinJacob Eisenstein @jacobeisenstein.bsky.social · 29/07/2025this is very cool and i’m looking forward to reading the paper, but a basic question about this data: isn’t it likely that a congressional rep’s speeches are written by a shifting cast of speechwriters over the course of their career? wouldn’t that explain adoption of new usages? 131
Reposted by Jacob EisensteinGaurav Kamath @grvkamath.bsky.social · 29/07/2025Our new paper in #PNAS (bit.ly/4fcWfma) presents a surprising finding—when words change meaning, older speakers rapidly adopt the new usage; inter-generational differences are often minor. w/ Michelle Yang, @sivareddyg.bsky.social , @msonderegger.bsky.social and @dallascard.bsky.social👇(1/12) 33317
Jacob Eisenstein @jacobeisenstein.bsky.social · 25/07/2025There's a lot to like in this position paper - and not just the "whiff of Frankenstein" quote. www.arxiv.org/abs/2507.06268 1151
Reposted by Jacob EisensteinMaria Antoniak @mariaa.bsky.social · 23/07/2025What are your favorite recent papers on using LMs for annotation (especially in a loop with human annotators), synthetic data for task-specific prediction, active learning, and similar? Looking for practical methods for settings where human annotations are costly. A few examples in thread ↴ 137923
Reposted by Jacob EisensteinAhmad Beirami @abeirami.bsky.social · 09/07/2025[Thu Jul 17] w/ Ananth Balashankar & @jacobeisenstein.bsky.social, we present a reinforcement learning framework in view of test-time scaling. We show how to optimally calibrate & transform rewards to obtain optimal performance with a given test-time algorithm. 111
Reposted by Jacob EisensteinAhmad Beirami @abeirami.bsky.social · 09/07/2025[Wed Jul 16] w/ @jacobeisenstein.bsky.social & Alekh Agarwal, we present a theoretical characterization of best-of-N (a simple yet effective method for test-time scaling & alignment). Our results justify the widespread use of BoN as a strong baseline in this space. 101
Jacob Eisenstein @jacobeisenstein.bsky.social · 10/06/2025Cheap but noisy? Or accurate but expensive? How to split a limited annotation budget between different types of judges?👩⚖️🤖🦧 www.arxiv.org/abs/2506.07949arxiv.orgCost-Optimal Active AI Model EvaluationThe development lifecycle of generative AI systems requires continual evaluation, data acquisition, and annotation, which is costly in both resources and time. In practice, rapid iteration often makes... 193
Reposted by Jacob EisensteinNed Resnikoff @resnikoff.bsky.social · 06/06/2025Everyone should check out People Time, the recording of his dates with Kenny Barron in Copenhagen three months before his death. Getz knew he was dying, and produced some of the most moving music of his career. www.youtube.com/watch?v=c3jd...youtube.comEast Of Sun (And West Of The Moon)YouTube video by Kenny Barron - Topic 1242
Reposted by Jacob EisensteinFerenc Huszár @inference.vc · 22/05/2025A new blog post with intuitions behind continuous-time Markov chains, a building block of diffusion language models, like @inceptionlabs.bsky.social's Mercury and Gemini Diffusion. This post touches on different ways of looking at Markov chains, connections to point processes, and more.inference.vcDiscrete Diffusion: Continuous-Time Markov ChainsA tutorial explaining some key intuitions behind continuous time Markov chains for machine learners interested in discrete diffusion models: alternative representations, connections to point processes... 1215
Reposted by Jacob EisensteinStella Biderman @stellaathena.bsky.social · 23/05/2025People keep plugging AI "Co-Scientists," so what happens when you ask them to do an important task like finding errors in papers? We built SPOT, a dataset of STEM manuscripts across 10 fields annotated with real errors to find out. (tl;dr not even close to usable) #NLProc arxiv.org/abs/2505.11855 411931
Jacob Eisenstein @jacobeisenstein.bsky.social · 22/05/2025‘intermediate tokens-often anthropomorphized as "thoughts" or reasoning traces’ 🌶️ but true! really glad to see work approaching inference scaling more skeptically and objectively 0132
Reposted by Jacob EisensteinMyra Cheng @myra.bsky.social · 21/05/2025Dear ChatGPT, Am I the Asshole? While Reddit users might say yes, your favorite LLM probably won’t. We present Social Sycophancy: a new way to understand and measure sycophancy as how LLMs overly preserve users' self-image. 614735
Jacob Eisenstein @jacobeisenstein.bsky.social · 21/05/2025Google Deepmind is hiring a research scientist in Seattle to work on foundational research in language! job-boards.greenhouse.io/deepmind/job...job-boards.greenhouse.ioResearch Scientist, Foundational Research in Language, USASeattle, Washington, US 091
Reposted by Jacob EisensteinAhmad Beirami @abeirami.bsky.social · 09/05/2025#ICML2025 Is standard RLHF optimal in view of test-time scaling? Unsurprisingly no. We show a simple change to standard RLHF framework that involves 𝐫𝐞𝐰𝐚𝐫𝐝 𝐜𝐚𝐥𝐢𝐛𝐫𝐚𝐭𝐢𝐨𝐧 and 𝐫𝐞𝐰𝐚𝐫𝐝 𝐭𝐫𝐚𝐧𝐬𝐟𝐨𝐫𝐦𝐚𝐭𝐢𝐨𝐧 (suited to test-time procedure) is optimal! 1176
Reposted by Jacob EisensteinGus @gusthema.bsky.social · 30/04/2025Gemma 3 explained: Longer context, image support, and a new 1B model. → goo.gle/4lV8iaw Other key enhancements: 🔸 Best model that fits in a single consumer GPU or TPU host 🔸 KV-cache memory reduction with 5-to-1 interleaved attention 🔸 And more! Read the blog for the full details on Gemma 3.goo.gleGemma explained: What’s new in Gemma 3- Google Developers BlogGoogle's Gemma 3 model includes vision-language support and architectural changes for resource-friendly multimodal language models. 1228