Sign in

Shyamgopal Karthik

@shyamgopal.bsky.social
445 followers 261 following 16 posts

PhD at Tübingen. Working on post-training diffusion and multimodal models. Previous research interns at Snapchat and Naver Labs. sgk98.github.io

PostsRepliesMedia
Reposted by Shyamgopal Karthik
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 19/06/2026
New Paper: arxiv.org/abs/2606.19370 Self-play yields capabilities but requires frustrating cost-function tuning. Surprisingly, just 30 minutes of demonstration data produces much more human-like driving policies! Led by @daphne-cornelisse.bsky.social Website: spiced-self-play.com
110518
Reposted by Shyamgopal Karthik
A. Sophia Koepke @askoepke.bsky.social · 03/06/2026
#CVPR2026 paper: It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models Text-to-image models often collapse to near-identical samples. Our fix: optimize the noise. Start from pink 🩷, not white noise. 🔗 akoepke.github.io/divgen/index... 1/6
173
Shyamgopal Karthik @shyamgopal.bsky.social · 03/06/2026
There's nothing more satisfying than watching the right noise do its magic with diffusion models! A few interesting takeaways I had from this work 🧵
130
Reposted by Shyamgopal Karthik
Marcus Klasson @marcusklasson.bsky.social · 24/04/2026
👋🇧🇷 If you are at #ICLR2026 today, you should talk to @antonbaumann.bsky.social who is presenting our paper about turning pre-trained VLMs into probabilistic models without retraining or fine-tuning. Poster Session 3 ⌚: 10:30am - 1:00pm (local time) 📍: Pavilion 3 P3 - #313 @iclr-conf.bsky.social
143
Reposted by Shyamgopal Karthik
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9
akoepke.github.io
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
25715
Reposted by Shyamgopal Karthik
tinkidinki.bsky.social @tinkidinki.bsky.social · 08/04/2026
Hey everyone, super happy to share our work on quantum algorithms for heterogeneous partial differential equations (PDEs)! (1/4) scirate.com/arxiv/2604.0...
111
Reposted by Shyamgopal Karthik
Martin Trapp @trappmartin.eurosky.social · 26/01/2026
This has now been accepted at @iclr-conf.bsky.social !
2342
Shyamgopal Karthik @shyamgopal.bsky.social · 25/11/2025
Was a very fun (and quick) investigation into biases of multimodal benchmarks, this time on tasks designed for "Spatial Supersensing" introduced by Cambrian-S with some great folks!
131
Reposted by Shyamgopal Karthik
Andreas Hochlehnert @ahochlehnert.bsky.social · 24/11/2025
🚨 New Paper: "Solving Spatial Supersensing Without Spatial Supersensing" Huge credit to the Cambrian-S team for tackling one of the hardest open problems in video understanding: spatial supersensing. In our paper, we take a closer look at their benchmarks & methods 👇
132
Reposted by Shyamgopal Karthik
Martin Trapp @trappmartin.eurosky.social · 18/09/2025
Unfortunately, our submission to #NeurIPS didn’t go through with (5,4,4,3). But because I think it’s an excellent paper, I decided to share it anyway. We show how to efficiently apply Bayesian learning in VLMs, improve calibration, and do active learning. Cool stuff! 📝 arxiv.org/abs/2412.06014
arxiv.org
Post-hoc Probabilistic Vision-Language Models
Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descripti...
25016
Shyamgopal Karthik @shyamgopal.bsky.social · 27/06/2025
Wonderful story behind some very nice SSL work!
000
Shyamgopal Karthik @shyamgopal.bsky.social · 11/06/2025
I'm in Nashville this week attending #CVPR2025. Excited to discuss post-training VLMs and diffusion models!
0101
Reposted by Shyamgopal Karthik
ML for Science @ml4science.bsky.social · 22/05/2025
We're super happy: Our Cluster of Excellence will continue to receive funding from the German Research Foundation @dfg.de ! Here’s to 7 more years of exciting research at the intersection of #machinelearning and science! Find out more: uni-tuebingen.de/en/research/... #ExcellenceStrategy
The members of the Cluster of Excellence "Machine Learning: New Perspectives for Science" raise their glasses and celebrate securing another funding period.
47520
Reposted by Shyamgopal Karthik
David Picard @davidpicard.eurosky.social · 03/03/2025
🚨 New preprint! How far can we go with ImageNet for Text-to-Image generation? w. @arrijitghosh.bsky.social @lucasdegeorge.bsky.social @nicolasdufour.bsky.social @vickykalogeiton.bsky.social TL;DR: Train a text-to-image model using 1000 less data in 200 GPU hrs! 📜https://arxiv.org/abs/2502.21318 🧵👇
26516
Shyamgopal Karthik @shyamgopal.bsky.social · 03/03/2025
These are some ridiculously good results from training tiny T2I models purely on ImageNet! It's almost too good to be true. Do check it out!
032
Reposted by Shyamgopal Karthik
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 09/02/2025
I've been talking about writing this paper to anyone who would listen since 2020. I bombed a bunch of job talks trying to convince companies to work on this. It's so nice to finally just be able to say, yes, self-play RL in a diverse world gives you immense capabilities arxiv.org/abs/2502.03349
arxiv.org
Robust Autonomy Emerges from Self-Play
Self-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic drivi...
3926
Reposted by Shyamgopal Karthik
Joschka Strüber @Tuebingen AI Center🇩🇪 @joschkastrueber.bsky.social · 07/02/2025
🚨Great Models Think Alike and this Undermines AI Oversight🚨 New paper quantifies LM similarity (1) LLM-as-a-judge favor more similar models🤥 (2) Complementary knowledge benefits Weak-to-Strong Generalization☯️ (3) More capable models have more correlated failures 📈🙀 🧵👇
2219
Reposted by Shyamgopal Karthik
Nicolas Dufour @nicolasdufour.bsky.social · 12/12/2024
ReNO shows that some initial noise are better for some prompts! This is great to improve image generation, but i think it also shows a deeper property of diffusion models.
122
Reposted by Shyamgopal Karthik
Fatemeh Khatibloo @fatemehx2.bsky.social · 14/12/2024
This is maybe my favorite thing I've seen out of #NeurIPS2024. Head over to HuggingFace and play with this thing. It's quite extraordinary.
032
Reposted by Shyamgopal Karthik
Luca Eyring @lucaeyring.bsky.social · 11/12/2024
Can we enhance the performance of T2I models without any fine-tuning? We show that with our ReNO, Reward-based Noise Optimization, one-step models consistently surpass the performance of all current open-source Text-to-Image models within the computational budget of 20-50 sec! #NeurIPS2024
1277
Reposted by Shyamgopal Karthik
Martin Trapp @trappmartin.eurosky.social · 10/12/2024
I will present ✌️ BDU workshop papers @ NeurIPS: one by Rui Li (looking for internships) and one by Anton Baumann. 🔗 to extended versions: 1. 🙋 "How can we make predictions in BDL efficiently?" 👉 arxiv.org/abs/2411.18425 2. 🙋 "How can we do prob. active learning in VLMs" 👉 arxiv.org/abs/2412.06014
arxiv.org
Post-hoc Probabilistic Vision-Language Models
Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descripti...
1184
Shyamgopal Karthik @shyamgopal.bsky.social · 09/12/2024
After a break of over 2 years, I'm attending a conference again! Excited to attend NeurIPS, even more so to be presenting ReNO, getting inference-time scaling and preference optimization to work for text-to-image generation. Do reach out if you'd like to chat!
0123
Reposted by Shyamgopal Karthik
Maitreya Patel ✈️ NeurIPS @patelmaitreya.bsky.social · 03/12/2024
🚨New Paper Alert🚨 🚀 Introducing FlowChef, "Steering Rectified Flow Models in the Vector Field for Controlled Image Generation"! 🌌✨ - Perform image editing, solve inverse problems, and more. - Achieved inversion-free, gradient-free, & training-free inference time steering! 🤯 👇👇
152
Reposted by Shyamgopal Karthik
Lucas Beyer (bl16) @giffmana.ai · 29/11/2024
Some recent discussions made me write up a short read on how I think about doing computer vision research when there's clear potential for abuse. Alternative title: why I decided to stop working on tracking. Curious about other's thoughts on this. lb.eyer.be/s/cv-ethics....
1917320
Shyamgopal Karthik @shyamgopal.bsky.social · 28/11/2024
Check out this nice work by @confusezius.bsky.social on designing VLMs for few-shot adaptation!
050
Reposted by Shyamgopal Karthik
Lucas Beyer (bl16) @giffmana.ai · 23/11/2024
A real-time (or very fast) open-source txt2video model dropped: LTXV. HF: huggingface.co/Lightricks/L... Gradio: huggingface.co/spaces/Light... Github: github.com/Lightricks/L... Look at that prompt example though. Need to be a proper writer to get that quality.
68910
Reposted by Shyamgopal Karthik
erogol.com @erogol.com · 24/11/2024
Learning from one continuous video stream - use a video stream to learn a predictive model - everything is in pixel space - update the model less frequently and don’t use momentum optimizer - pre training with iid improves performance - continual learning for robots arxiv.org/html/2312.00...
arxiv.org
Learning from One Continuous Video Stream
0183
Reposted by Shyamgopal Karthik
Sebastian Dziadzio @dziadzio.bsky.social · 19/11/2024
Here's a fledgling starter pack for the AI community in Tübingen. Let me know if you'd like to be added! go.bsky.app/NFbVzrA
go.bsky.app
Tübingen AI
Join the conversation
182513