Reposted by @momergul.bsky.socialNathan Godey @nthngdy.bsky.social · 12/03/2026🧵New paper: "Lost in Backpropagation: The LM Head is a Gradient Bottleneck" The output layer of LLMs destroys 95-99% of your training signal during backpropagation, and this significantly slows down pretraining 👇 610815
Reposted by @momergul.bsky.socialLeshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 02/03/2026👶 BabyLM is back at EMNLP 2026! We are excited to announce that the 4th BabyLM Challenge & Workshop will once again bring together researchers interested in sample-efficient, developmentally plausible language modeling. @emnlpmeeting More in🧵 1103
Reposted by @momergul.bsky.socialZizhao Chen @ch272h.bsky.social · 05/12/2025🧩Natural language isn’t all you need. We’re great at evaluating text-based reasoning (MATH, AIME…) but what about long-horizon visual reasoning? Enter 𝗞𝗻𝗼𝘁𝗚𝘆𝗺: a minimalistic testbed for evaluating agents on spatial reasoning along a difficulty ladder 1184
Reposted by @momergul.bsky.socialJaap Jumelet @jumelet.bsky.social · 15/10/2025🌍Introducing BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data! LLMs learn from vastly more data than humans ever experience. BabyLM challenges this paradigm by focusing on developmentally plausible data We extend this effort to 45 new languages! 14416
momergul.bsky.social @momergul.bsky.social · 02/10/2025🚨Modeling Abstention via Selective Help-seeking LLMs learn to use search tools to answer questions they would otherwise hallucinate on. But can this also teach them what they know vs not? We introduce MASH that trains LLMs for search and gets abstentions for free! 110
Reposted by @momergul.bsky.socialYoav Artzi @yoavartzi.com · 25/07/2025The talk for our work on Retrospective Learning from Interactions, which will be in ACL (once I figure out how to squeeze it shorter) Gist: autonomous post-training from conversational signals for LLM bootstrapping ... look ma, no annotations! no hand-holding! 🙌📈🚀 www.youtube.com/watch?v=qW8S...youtube.comRetrospective Learning from InteractionsYouTube video by Yoav Artzi 1115
Reposted by @momergul.bsky.socialLeshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 09/05/2025Close your books, test time! The evaluation pipelines are out, baselines are released & the challenge is on There is still time to join and We are excited to learn from you on pretraining and human-model gaps *Don't forget to fastEval on checkpoints github.com/babylm/evalu... 📈🤖🧠 #AI #LLMS 0104