Damien Teney @damienteney.bsky.social · 20/02/2026🔥What if web text isn’t the best place to start training LLMs? Our latest work shows that warming up models on procedural data (e.g. from formal languages & simple algorithms) speeds up subsequent pretraining on language, code, and math, on models up to 1.3B parameters⬇️🧵 1503
Damien Teney @damienteney.bsky.social · 10/12/2025Can vision transformers learn without images?🤔👀 Our latest work shows that pretraining ViTs on procedural symbolic data (eg sequences of balanced parentheses) makes subsequent standard training (eg on ImageNet) more data efficient! How is this possible?! ⬇️🧵 3486
Damien Teney @damienteney.bsky.social · 07/07/2025Coming up at ICML: 🤯Distribution shifts are still a huge challenge in ML. There's already a ton of algorithms to address specific conditions. So what if the challenge was just selecting the right algorithm for the right conditions?🤔🧵 161
Reposted by Damien TeneyChristian Wolf @chriswolfvision.bsky.social · 24/06/2025OMG I can confirm this ... tested by @mbsariyildiz.bsky.social on our new upcoming work (vision/robotics). Thanks @damienteney.bsky.social the effect is real 😍 arxiv.org/abs/2505.20802 2353
Damien Teney @damienteney.bsky.social · 14/06/2025⬇️ "Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild" arxiv.org/abs/2503.10065 182
Damien Teney @damienteney.bsky.social · 07/06/2025Coming up this week: (oral @cvprconference.bsky.social) Do We Always Need the Simplicity Bias? We take another step to understand why/when neural nets generalize so well. ⬇️🧵 180
Damien Teney @damienteney.bsky.social · 01/06/2025What does/will happen when the model learns (from experience or historical data) that accepted papers contain grand claims and bold numbers >SOTA?💡Hint: it's not "doing rigorous science". 000
Damien Teney @damienteney.bsky.social · 22/04/2025Seen on the other place... 😕🤷 Any advice to get more ML and less bunnies/politics in my Bluesky feed? 310
Damien Teney @damienteney.bsky.social · 10/04/2025This ⬇️ is also great advice for writing a paper! 👌 But start it one *month* before the deadline. 281
Damien Teney @damienteney.bsky.social · 10/03/2025"Deep learning does not require rethinking generalization" If you enjoyed our work on inductive biases (eg the Neural Redshift arxiv.org/abs/2403.02241), you'll love this paper that rigorously articulates "soft inductive biases" & how they explain supposedly-mysterious behaviors of neural nets. 0101
Damien Teney @damienteney.bsky.social · 27/02/2025This is as useful as telling us what the authors had for breakfast on the day of the experiments. 🤷 210
Reposted by Damien TeneyDavid Monniaux @monniauxd.bsky.social · 09/02/2025Une chose me marque au sujet de l'IA, notamment générative. Que penser d'une innovation technico-scientifique dont la promotion dans les médias est plutôt du fait d'entrepreneurs, éditorialistes et politiciens, tandis que les scientifiques du secteur sont bien plus mesurés ? 1911117
Damien Teney @damienteney.bsky.social · 11/12/2024"Attention" in attention layers. How about sum-product layers? Key-query products? ... Neural attention has little to do with human attention. And the intuitive baggage of the name probably constrains our thinking about how transformers work. (1/2) 4203
Damien Teney @damienteney.bsky.social · 08/12/2024💡 Just learned about a useful short-hand notation for "expectation". Seems common for physicists but I can't remember coming across it before. With an example use-case below ⬇️ 120
Damien Teney @damienteney.bsky.social · 25/11/2024Reviewing ML papers? 💡 If you feel that experiments are missing, ask yourself: are the additional results likely to affect the central message of the paper/nullify its main claims? If not, it's probably a nice suggestion (eg additional comparisons, datasets) but not a reason for rejection by itself. 030
Damien Teney @damienteney.bsky.social · 24/11/2024PSA: Can we use more bar charts in ML papers? I can't recall the last time I wanted to compare dozens of numbers in a table to two decimal places. A visualization makes it much clearer whether claimed differences are significant. 4282
Damien Teney @damienteney.bsky.social · 20/11/2024Writing tips! This should be mandatory reading for every PhD student 👇 We'll all benefit from it as readers. 030