Sign in

rogergrosse.bsky.social

@rogergrosse.bsky.social
1.7K followers 77 following 6 posts
PostsRepliesMedia
rogergrosse.bsky.social @rogergrosse.bsky.social · 18/12/2024
The Nintendo is closer in time to the first transistor than to today.
050
rogergrosse.bsky.social @rogergrosse.bsky.social · 07/12/2024
Conferences are basically a way for a group of people to temporarily have a lower opportunity cost on their time.
1100
Reposted by @rogergrosse.bsky.social
Tomer Ullman @tomerullman.bsky.social · 01/12/2024
thinking of calling this "The Illusion Illusion" (more examples below)
601570381
Reposted by @rogergrosse.bsky.social
Jonathan Lorraine @jonlorraine.bsky.social · 27/11/2024
🚨 New #NeurIPS2025 paper “Training Data Attribution via Approximate Unrolling” 🚨 Introducing SOURCE: A method to understand how individual training examples influence neural net behavior, allowing us to make AI models more transparent and trustworthy! 📄 Full paper: openreview.net/pdf?id=3NaqG...
1192
rogergrosse.bsky.social @rogergrosse.bsky.social · 26/11/2024
I have Claude filter my arXiv feed each day. It mostly works pretty well, except that it always hallucinates that "Studying LLM Generalization with Influence Functions" is in my feed and tells me I should read it.
390
rogergrosse.bsky.social @rogergrosse.bsky.social · 22/11/2024
Some very nice work from Cohere and UCL using influence functions to analyze math reasoning abilities in LLMs. Factual queries turn up docs containing the facts, but reasoning queries turn up similar cognitive strategies, suggesting generalization. arxiv.org/abs/2411.12580
arxiv.org
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
The capabilities and limitations of Large Language Models have been sketched out in great detail in recent years, providing an intriguing yet conflicting picture. On the one hand, LLMs demonstrate a g...
0162