Sign in

Damien Teney

@damienteney.bsky.social
555 followers 302 following 166 posts

Research Scientist @ Idiap Research Institute. @idiap.bsky.social Adjunct lecturer @ Australian Institute for ML. @aimlofficial.bsky.social Occasionally cycling across continents. www.damienteney.info

PostsRepliesMedia
Damien Teney @damienteney.bsky.social · 21/02/2026
Reviewer #2 striking again?
020
Damien Teney @damienteney.bsky.social · 20/02/2026
This all looks very promising, and there's a lot more to explore! Paper and code ⬇️ Procedural Pretraining: Warming Up Language Models with Abstract Data www.arxiv.org/abs/2601.21725 github.com/zlshinnick/p...
arxiv.org
Procedural Pretraining: Warming Up Language Models with Abstract Data
Pretraining directly on web-scale corpora is the de facto paradigm for building language models. We study an alternative setting where the model is initially exposed to abstract structured data, as a ...
010
Damien Teney @damienteney.bsky.social · 20/02/2026
🧩 Multiple types of procedural data can be combined. We get further gains by mixing either • multiple types of data, or • weights of models individually warmed-up on different types of data.
100
Damien Teney @damienteney.bsky.social · 20/02/2026
⚙️ MLPs vs. attention: where is the information located? We try resetting selected weights to random, before standard pretraining. Surprisingly, we obtain further gains, but they're domain-specific: • warmed-up MLPs benefit natural language • warmed-up attention helps code/math
100
Damien Teney @damienteney.bsky.social · 20/02/2026
📈 Benefits on subsequent standard pretraining. By front-loading as little as 0.1% procedural data, models achieve significantly better pretraining performance on language, code, and math. They use up to 45% less semantic data to reach a baseline perplexity.
100
Damien Teney @damienteney.bsky.social · 20/02/2026
🔍 Different procedural data = different benefits. We first determine the effect of different types of procedural data with algorithmic diagnostic tasks. The benefits range from long-context recall to arithmetic, depending on the type of procedural data.
100
Damien Teney @damienteney.bsky.social · 20/02/2026
💡 Humans learn better when starting with simple structure and logic rather than memorizing a massive set of facts. By analogy, we use abstract, structured data to build a scaffold in language models, free of semantic biases.
100
Damien Teney @damienteney.bsky.social · 20/02/2026
🔥What if web text isn’t the best place to start training LLMs? Our latest work shows that warming up models on procedural data (e.g. from formal languages & simple algorithms) speeds up subsequent pretraining on language, code, and math, on models up to 1.3B parameters⬇️🧵
1503
Damien Teney @damienteney.bsky.social · 10/12/2025
Indeed the effect in *late* layers was very surprising! My optimist interpretation is that the procedural pretraining creates circuits for computations general enough to serve as a useful scaffold for visual tasks. This would explain why they help and why they don't wash out with more training.
020
Damien Teney @damienteney.bsky.social · 10/12/2025
Sounds 😋 What's the objective function? simplicity/low cost/?
100
Damien Teney @damienteney.bsky.social · 10/12/2025
In summary, a lightweight generic warm-up improves accuracy and data efficiency, with effects distinct from ImageNet pretraining. Lots of exciting open questions! 🔍 -Other types of procedural data? -Other downstream tasks? -Closed-form instantiation? arxiv.org/abs/2511.13945
arxiv.org
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
Transformers show remarkable versatility across domains, suggesting the existence of inductive biases beneficial across modalities. In this work, we explore a new way to instil such generic biases in ...
030
Damien Teney @damienteney.bsky.social · 10/12/2025
🔍𝐖𝐡𝐞𝐫𝐞 𝐢𝐬 𝐭𝐡𝐢𝐬 𝐤𝐧𝐨𝐰𝐥𝐞𝐝𝐠𝐞 𝐬𝐭𝐨𝐫𝐞𝐝? Ablations show that the knowledge mostly locates in *late* layers: the opposite of normal visual pretraining which shapes early layers. Procedural data seems to provide a qualitatively unique training signal!
110
Damien Teney @damienteney.bsky.social · 10/12/2025
🧠𝐖𝐡𝐚𝐭 𝐤𝐢𝐧𝐝 𝐨𝐟 𝐝𝐚𝐭𝐚 𝐰𝐨𝐫𝐤𝐬? Formal languages with hierarchical structure seem best. If we shuffle the training tokens (eliminating nested structures), the gains disappear, showing that the benefits are not due to surface-level frequencies.
110
Damien Teney @damienteney.bsky.social · 10/12/2025
📉𝐏𝐫𝐨𝐜𝐞𝐝𝐮𝐫𝐚𝐥 𝐝𝐚𝐭𝐚 𝐜𝐚𝐧 𝐫𝐞𝐩𝐥𝐚𝐜𝐞 𝐫𝐞𝐚𝐥 𝐢𝐦𝐚𝐠𝐞𝐬 Allocating just 1% of the ImageNet pretraining budget to the procedural warmup lets the ViT match the baseline accuracy with 28% fewer images!
120
Damien Teney @damienteney.bsky.social · 10/12/2025
📈𝐀 𝐝𝐢𝐟𝐟𝐞𝐫𝐞𝐧𝐭 𝐨𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐭𝐫𝐚𝐣𝐞𝐜𝐭𝐨𝐫𝐲 Our warmed-up models don't just get a head-start, they train differently. On ImageNet (below), they follow a distinct training trajectory and converge to a better accuracy.
110
Damien Teney @damienteney.bsky.social · 10/12/2025
🔥Our procedural data has no semantic or visual meaning: it simply forces the model to discover generic structure in the data. As initialisation for standard image-based training, it -boosts accuracy, -improves data efficiency, -complements ImageNet pretraining.
110
Damien Teney @damienteney.bsky.social · 10/12/2025
💡Prior work has already shown that LLMs acquire useful knowledge when pretrained on formal languages. To test this on ViTs, we devise a procedural warm-up: pretraining for next-token prediction on symbolic sequences, bypassing the visual patch embedding.
110
Damien Teney @damienteney.bsky.social · 10/12/2025
In summary, a lightweight generic warm-up improves accuracy and data efficiency, with effects distinct from ImageNet pretraining. Lots of exciting open questions! 🔍 - Other types of procedural data? - Other downstream tasks? - Closed-form instantiation? arxiv.org/abs/2511.13945
arxiv.org
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
Transformers show remarkable versatility across domains, suggesting the existence of inductive biases beneficial across modalities. In this work, we explore a new way to instil such generic biases in ...
020
Damien Teney @damienteney.bsky.social · 10/12/2025
🔍𝐖𝐡𝐞𝐫𝐞 𝐢𝐬 𝐭𝐡𝐢𝐬 𝐤𝐧𝐨𝐰𝐥𝐞𝐝𝐠𝐞 𝐬𝐭𝐨𝐫𝐞𝐝? Ablations show that the knowledge mostly locates in *late* layers: the opposite of normal visual pretraining which shapes early layers. Procedural data seems to provide a qualitatively unique training signal!
220
Damien Teney @damienteney.bsky.social · 10/12/2025
🧠𝐖𝐡𝐚𝐭 𝐤𝐢𝐧𝐝 𝐨𝐟 𝐝𝐚𝐭𝐚 𝐰𝐨𝐫𝐤𝐬? Formal languages with hierarchical structure seem best. If we shuffle the training tokens (eliminating nested structures), the gains disappear, showing that the benefits are not due to surface-level frequencies.
120
Damien Teney @damienteney.bsky.social · 10/12/2025
📉𝐏𝐫𝐨𝐜𝐞𝐝𝐮𝐫𝐚𝐥 𝐝𝐚𝐭𝐚 𝐜𝐚𝐧 𝐫𝐞𝐩𝐥𝐚𝐜𝐞 𝐫𝐞𝐚𝐥 𝐢𝐦𝐚𝐠𝐞𝐬 Allocating just 1% of the ImageNet pretraining budget to the procedural warmup lets the ViT match the baseline accuracy with 28% fewer images!
120
Damien Teney @damienteney.bsky.social · 10/12/2025
📈𝐀 𝐝𝐢𝐟𝐟𝐞𝐫𝐞𝐧𝐭 𝐨𝐩𝐭𝐢𝐦𝐢𝐳𝐚𝐭𝐢𝐨𝐧 𝐭𝐫𝐚𝐣𝐞𝐜𝐭𝐨𝐫𝐲 Our warmed-up models don't just get a head-start, they train differently. On ImageNet (below), they follow a distinct training trajectory and converge to a better accuracy.
130
Damien Teney @damienteney.bsky.social · 10/12/2025
🔥Our procedural data has no semantic or visual meaning: it simply forces the model to discover generic structure in the data. As initialisation for standard image-based training, it -boosts accuracy, -improves data efficiency, -complements ImageNet pretraining.
120
Damien Teney @damienteney.bsky.social · 10/12/2025
Can vision transformers learn without images?🤔👀 Our latest work shows that pretraining ViTs on procedural symbolic data (eg sequences of balanced parentheses) makes subsequent standard training (eg on ImageNet) more data efficient! How is this possible?! ⬇️🧵
3486
Damien Teney @damienteney.bsky.social · 26/07/2025
Academic Strava?🤓 It feels like an underrepresented group in my Strava feed!
110
Damien Teney @damienteney.bsky.social · 24/07/2025
It'd be nice to provide complete analyses (that you have precomputed) of existing papers, so we can see what kind of output the tool provides, without having to submit any of my own work.
1230
Damien Teney @damienteney.bsky.social · 22/07/2025
Dang it just never ends 😱
000
Damien Teney @damienteney.bsky.social · 16/07/2025
In this setting, does the student (sometimes?) get better than the teacher? One hypothesis could be that the teacher, even if "less correct" than the GT, provides supervision that's easier to learn for another NN (the student). The optimization follows a less tortuous path & finds a better solution.
130
Damien Teney @damienteney.bsky.social · 07/07/2025
👍 I had my very first paper published at DAGM. It was a while ago but I remember it as a very welcoming conference.
010
Damien Teney @damienteney.bsky.social · 07/07/2025
🎯There's already a plethora of methods to handle distribution shifts: most gains may now simply be in better using them! Automatic selection looks promising, yet there's lots more to do. Interested? Come chat with us at ICML! 📄 arxiv.org/abs/2410.02735 💻 github.com/LiangzeJiang...
arxiv.org
OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?
Out-of-distribution (OOD) generalization is challenging because distribution shifts come in many forms. Numerous algorithms exist to address specific settings, but choosing the right training algorith...
000
Damien Teney @damienteney.bsky.social · 07/07/2025
🔎Last but not least: OOD-Chameleon shines new light on existing algorithms! We can interpret the selection process as a tree and get interpretable guidelines for choosing algorithms.
100
Damien Teney @damienteney.bsky.social · 07/07/2025
✅We test OOD-Chameleon on unseen datasets & shifts: it accurately predict suitable algorithms on synthetic, vision, and language tasks. The downstream models (trained with selected algorithms) consistently have lower error than with standard selection heuristics.
110
Damien Teney @damienteney.bsky.social · 07/07/2025
🛠️To create our "dataset of datasets", we resample CelebA and CivilComments with constraints specifying diverse types/magnitudes of shifts. We also train small models with candidate algorithms to obtain their "ground truth performance" in each condition.
100
Damien Teney @damienteney.bsky.social · 07/07/2025
🦎We propose OOD-Chameleon as a proof-of-concept: an algorithm selector as a classifier (over candidate algorithms) trained on a "dataset of datasets" representing diverse shifts. The model learns which algorithms perform best in different conditions.
100
Damien Teney @damienteney.bsky.social · 07/07/2025
💡We're aiming for an "auto-ML for distribution shifts". We conjecture that datasets have properties predictive of the suitability of various algorithms to handle dist. shifts: size/complexity of the data, magnitudes/types of shifts, etc.
100
Damien Teney @damienteney.bsky.social · 07/07/2025
"Distribution shift" means many things: spurious correlations, covariate shift, label shift... with no one-size-fits-all! Many algorithms exist, 📊each for specific conditions. Could we automate the selection without trial-and-error❓
100
Damien Teney @damienteney.bsky.social · 07/07/2025
Coming up at ICML: 🤯Distribution shifts are still a huge challenge in ML. There's already a ton of algorithms to address specific conditions. So what if the challenge was just selecting the right algorithm for the right conditions?🤔🧵
161
Damien Teney @damienteney.bsky.social · 30/06/2025
Nice design!
000
Damien Teney @damienteney.bsky.social · 30/06/2025
On the contrary it's a discussion that needs bringing up inside the CV research community. They're not just a bunch of white dudes with evil intentions or industry pressure to build a surveillance state. As one example see the thoughts from one such researcher lucasb.eyer.be/snips/cv-eth...
lucasb.eyer.be
Ethical considerations around Vision and Robotics
A rough outline on how I think about doing research in Computer Vision given the many possible unethical uses.
110
Damien Teney @damienteney.bsky.social · 30/06/2025
Brilliant! Some pushback I heard against mandatory reviewing was from people misunderstanding that a submission entails such a partnership with the rest of the community.
010
Damien Teney @damienteney.bsky.social · 30/06/2025
Looks quiet! Best time of day 👌🏼
010
Damien Teney @damienteney.bsky.social · 30/06/2025
Good points. It's a difficult task to automate the classification of papers/patents. Would the tracking of hands/gestures for sign language interfaces count as surveillance here?
110
Damien Teney @damienteney.bsky.social · 29/06/2025
Because it discredits and could silence an entire area of scientific inquiry. Inclusive access to technology would tremedously benefit from CV technologies, so surveillance-based commercialization is only one part of the conversation.
100
Damien Teney @damienteney.bsky.social · 29/06/2025
There's a healthy amount of skepticism to be had when reading any paper, whatever level of peer-review it got through. Just like I wouldn't blindly trust a CVPR or NeurIPS paper claiming to beat the SOTA, because the authors were motivated to do so. Happy to continue the discussion elsewhere.
100
Damien Teney @damienteney.bsky.social · 29/06/2025
I'd be an important topic worth addressing more deeply within the CV community! We'd need more data on dual use and on beneficial application of the technology as well.
100
Damien Teney @damienteney.bsky.social · 29/06/2025
Indeed, when reading the scientific literature, even with any amount of scrutiny and peer review, you ultimately have to trust the authors that they did their due diligence and didn't cut corners to arrive at their desired conclusions.
100
Damien Teney @damienteney.bsky.social · 29/06/2025
I'm open to any finding, there's no place for feelings when interpreting data. But I'm not sure this was the case for the authors. It's difficult to trust this paper because it feels like a piece of activism rather than an unbiased scientific study.
200
Damien Teney @damienteney.bsky.social · 29/06/2025
Good for you. I said *most*, and was just reflecting on typical members of the community from the CV venues studied in the paper.
100
Damien Teney @damienteney.bsky.social · 29/06/2025
I think it's to the detriment of the authors, bc such a biased activist message will be ignored by most of the very ppl (CV researchers) that the information could have an effect on. This should have been presented at a CV conference if they were hoping for actual impact. @abeba.bsky.social
100
Damien Teney @damienteney.bsky.social · 29/06/2025
What I'm criticizing is the motivated reasoning, eg seeing evil in field-specific jargon that has perfectly valid technical meaning, but isn't obvious to ppl not familiar with the CV literature. It's unfortunate because the paper has some valid points, it just completely lacks balance.
100