Sign in

Andreas Hochlehnert

@ahochlehnert.bsky.social
204 followers 81 following 28 posts

PhD student in ML at Tübingen AI Center & International Max-Planck Research School for Intelligent Systems

PostsRepliesMedia
Andreas Hochlehnert @ahochlehnert.bsky.social · 26/08/2026
[1/6] 🚨 We’re releasing LAION-BVD: a 10-million-hour open video dataset for multimodal pre-training. - 1.3B video URLs from CommonCrawl - 80M downloaded videos - 10M video hours - 55M captioned clips - 300M frame-caption pairs 🌐: projects.laion.ai/bvd/ 🧵👇
162
Andreas Hochlehnert @ahochlehnert.bsky.social · 24/11/2025
🚨 New Paper: "Solving Spatial Supersensing Without Spatial Supersensing" Huge credit to the Cambrian-S team for tackling one of the hardest open problems in video understanding: spatial supersensing. In our paper, we take a closer look at their benchmarks & methods 👇
132
Andreas Hochlehnert @ahochlehnert.bsky.social · 08/10/2025
Presenting A Sober Look at Progress in LM Reasoning at @colmweb.org today 🇨🇦 #COLM2025 📅 Today 🕔 11:00 AM – 1:00 PM 📍 Room 710 - Poster #31 We find that many “reasoning” gains fall within variance and show how to make evaluation reproducible again. 📘 bethgelab.github.io/sober-reasoning
000
Reposted by Andreas Hochlehnert
Andreas Geiger @andreasgeiger.bsky.social · 20/08/2025
Excited about this new work from @haoyuhe.bsky.social. TLDR: Diffusion language models treat learning and inference differently which lowers performance. RL can be used to overcome this issue for certain problems.
061
Andreas Hochlehnert @ahochlehnert.bsky.social · 10/04/2025
🧵1/ 🚨 New paper: A Sober Look at Progress in Language Model Reasoning We re-evaluate recent SFT and RL models for mathematical reasoning and find most gains vanish under rigorous, multi-seed, standardized evaluation. 📊 bethgelab.github.io/sober-reason... 📄 arxiv.org/abs/2504.07086
1145
Reposted by Andreas Hochlehnert
Prasanna Mayilvahanan @prasannamayil.bsky.social · 18/02/2025
New preprint out! 🎉 How does LLM training loss translate to downstream performance? We show that pretraining data and tokenizer shape loss-to-loss scaling, while architecture and other factors play a surprisingly minor role! brendel-group.github.io/llm-line/ 🧵1/8
1188
Andreas Hochlehnert @ahochlehnert.bsky.social · 17/02/2025
CuratedThoughts: Data Curation for RL Datasets 🚀 Since DeepSeek-R1 introduced reasoning-based RL, datasets like Open-R1 & OpenThoughts emerged for fine-tuning & GRPO. Our deep dive found major flaws — 25% of OpenThoughts needed elimination by data curation. Here's why 👇🧵
1139
Reposted by Andreas Hochlehnert
Ofir Press @ofirpress.bsky.social · 17/01/2025
SWE-bench Multimodal evaluation code is out now! SWE-bench MM is a new set of JavaScript issues that have a visual component (‘map isn’t rendering correctly’, ‘button text isn’t appearing’). www.swebench.com/sb-cli/
081
Andreas Hochlehnert @ahochlehnert.bsky.social · 13/12/2024
We are presenting CiteMe today at the 11AM poster session (East Exhibit Hall A-C, #3309) CiteMe is a challenging benchmark for LM-based agents to find paper citations, moving beyond simple multiple-choice Q&A to real-world use cases. Come by and say hi :) citeme.ai
citeme.ai
CiteME
CiteME is a benchmark designed to test the abilities of language models in finding papers that are cited in scientific texts.
261
Reposted by Andreas Hochlehnert
Sebastian Dziadzio @dziadzio.bsky.social · 19/11/2024
Here's a fledgling starter pack for the AI community in Tübingen. Let me know if you'd like to be added! go.bsky.app/NFbVzrA
go.bsky.app
Tübingen AI
Join the conversation
182513