Sign in

Visual Inference Lab

@visinf.bsky.social
239 followers 344 following 65 posts

Visual Inference Lab of @stefanroth.bsky.social at @tuda.bsky.social - Research in Computer Vision and Machine Learning. See www.visinf.tu-darmstadt.de/visual_i…

PostsRepliesMedia
Visual Inference Lab @visinf.bsky.social · 08/09/2026
[6/6] MAGneT-3D: Monocular and Domain-Generalizable Temporal 3D Detection M. Kotb, J. Meier, @christophreich.bsky.social, O. Dhaouadi, L. Denninger, @dcremers.bsky.social Paper: arxiv.org/abs/2608.14282
arxiv.org
MAGneT-3D: Monocular and Domain-Generalizable Temporal 3D Detection
Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query-based 3D detectors unify detection and cross-view association, but their learnable queries fit the spatial ...
021
Visual Inference Lab @visinf.bsky.social · 08/09/2026
[5/6] LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection J. Meier*, J. Michel*, O. Dhaouadi, Y. Yang, @christophreich.bsky.social, Z. Bauer, @stefanroth, @marcpollefeys.bsky.social, J. Kaiser, @dcremers.bsky.social 🌍 Project Page: deepscenario.github.io/LeAD-M3D
arxiv.org
LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
Real-time monocular 3D object detection remains challenging due to severe depth ambiguity, viewpoint shifts, and the high computational cost of 3D reasoning. Existing approaches either rely on LiDAR o...
121
Visual Inference Lab @visinf.bsky.social · 08/09/2026
[4/6] Importance-Aware Low-Rank Distillation of Diffusion Transformers D. Zavadski, S. Heid, D. Kalšan, @stefanroth.bsky.social, C. Rother 📚 Paper: arxiv.org/abs/2609.04646 🌍 Project Page: vislearn.github.io/SVDtrunc
arxiv.org
Importance-Aware Low-Rank Distillation of Diffusion Transformers
Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generation, yet their scale poses challenges for efficient deployment. While truncated singular val...
120
Visual Inference Lab @visinf.bsky.social · 08/09/2026
[3/6] ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views Giuseppe Stracquadanio.bsky.social, K. Raj, @juliagrabinski.bsky.social, @stefanroth.bsky.social 📚 Paper: arxiv.org/abs/2608.28895 🌍 Project Page: visinf.github.io/reconsplat
stracquadanio.bsky.social
Giuseppe Stracquadanio (@stracquadanio.bsky.social)
PhD Student @ Visual Inference Lab (@visinf.bsky.social), TU Darmstadt
130
Visual Inference Lab @visinf.bsky.social · 08/09/2026
[2/6] WiFlow: Estimating Optical Flow using WiFi Channel State Information T. Weigel, @skiefhaber.de, F. Portner, M. Hollick, @simoneschaub.bsky.social 📚 Paper: arxiv.org/abs/2609.02452 🌍 Project Page: visinf.github.io/wiflow
arxiv.org
WiFlow: Estimating Optical Flow using WiFi Channel State Information
Knowing where and how fast objects are moving within a scene is important across various domains. Usually, cameras are used to capture the data necessary for this task, but adding cameras often raises...
120
Visual Inference Lab @visinf.bsky.social · 08/09/2026
[1/6] 📢 We are in Malmö at #ECCV2026 presenting 5 papers! 🎉
184
Reposted by Visual Inference Lab
Stefan Roth @stefanroth.bsky.social · 05/06/2026
Our #CVPR2026 poster starts in 30 minutes! Come by Poster Session 2, #333 to learn about VideoCUPS — the first unsupervised video panoptic segmentation method, learning to detect, segment, and track objects in video without human supervision.
093
Reposted by Visual Inference Lab
Stefan Roth @stefanroth.bsky.social · 06/06/2026
Afterwards, please stop by to see MARCO, our approach for semantic correspondence that generalizes to novel keypoints and categories. 4:12 PM in oral session 4D and later at poster #20. #CVPR2026 github.com/visinf/MARCO
0113
Reposted by Visual Inference Lab
Stefan Roth @stefanroth.bsky.social · 06/06/2026
Come see INSID3, in-context segmentation with DINOv3, at 2:00 PM in oral session 4D and later at poster #19. #CVPR2026
062
Visual Inference Lab @visinf.bsky.social · 05/06/2026
Work by: @christophreich.bsky.social*, @olvrhhn.bsky.social*, @neekans.bsky.social, @lealtaixe.bsky.social, C. Rupprecht, @dcremers.bsky.social and @stefanroth.bsky.social 📄Paper: arxiv.org/abs/2606.04925 🌍Project Page: visinf.github.io/videocups 👁️CVPR: Friday, Poster Session 2 #333
040
Visual Inference Lab @visinf.bsky.social · 05/06/2026
When fine-tuned with just 10% of labels, VideoCUPS already matches a fully supervised model trained on all Cityscapes-VPS labels, and outperforms the DINO-initialized baseline significantly.
130
Visual Inference Lab @visinf.bsky.social · 05/06/2026
VideoCUPS outperforms four competitive baselines that pair state-of-the-art unsupervised semantic, video-instance, or panoptic image segmentation methods with unsupervised tracking.
130
Visual Inference Lab @visinf.bsky.social · 05/06/2026
We train a video panoptic segmentation model using our unsupervised pseudo labels applying a novel Video DropLoss and self-enhanced video copy-paste augmentations.
130
Visual Inference Lab @visinf.bsky.social · 05/06/2026
We mine three rich supervision signals from raw monocular videos: motion (SMURF optical flow), depth (DynamoDepth), and appearance (distilled DINO + k-means) to generate temporally consistent panoptic pseudo-labels.
130
Visual Inference Lab @visinf.bsky.social · 05/06/2026
📢 [CVPR’26] Can we learn to detect, segment, and track every object in a video without human supervision?  Yes, we introduce VideoCUPS, the first unsupervised video panoptic segmentation (VPS) method: 1. Get pseudo-labels from monocular videos. 2. Train a VPS model on them.
1154
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[6/6] Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities @dustin-carrion.bsky.social*, M. Santos-Villafranca*, A.Perez-Yus, J. Bermudez-Cameo, J.J.Guerrero, @simoneschaub.bsky.social Paper: arxiv.org/abs/2504.08578 Project Page: visinf.github.io/KARMMA
arxiv.org
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities
Egocentric action recognition enables robots to facilitate human-robot interactions and monitor task progress. Existing methods often rely solely on RGB videos, although additional modalities, such as...
010
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[5/6] MUFASA: A Multi-Layer Framework for Slot Attention S. Bock*, L. Schüßler*, K. Singh, @simoneschaub.bsky.social, @stefanroth.bsky.social Paper: arxiv.org/abs/2602.07544 Project Page: visinf.github.io/mufasa/
arxiv.org
MUFASA: A Multi-Layer Framework for Slot Attention
Unsupervised object-centric learning (OCL) decomposes visual scenes into distinct entities. Slot attention is a popular approach that represents individual objects as latent vectors, called slots. Cur...
110
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[4/6] MARCO: Navigating the Unseen Space of Semantic Correspondence C. Cuttano, @gabtriv.bsky.social , C. Masone, @stefanroth.bsky.social Paper: arxiv.org/abs/2604.18267 Project Page: visinf.github.io/MARCO/
arxiv.org
MARCO: Navigating the Unseen Space of Semantic Correspondence
Recent advances in semantic correspondence rely on dual-encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion-parameter models generalize poorly beyond traini...
110
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[3/6] INSID3: Training-Free In-Context Segmentation with DINOv3 C. Cuttano, @gabtriv.bsky.social , @christophreich.bsky.social , @dcremers.bsky.social , C. Masone, @stefanroth.bsky.social Paper: arxiv.org/abs/2603.28480 Project Page: visinf.github.io/INSID3
120
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[2/6] Scene-Centric Unsupervised Video Panoptic Segmentation @christophreich.bsky.social*,  @olvrhhn.bsky.social*, @neekans.bsky.social,  @lealtaixe.bsky.social , C. Rupprecht, @dcremers.bsky.social, @stefanroth.bsky.social Paper: arxiv.org/abs/2606.04925 Project Page: visinf.github.io/videocups/
arxiv.org
Scene-Centric Unsupervised Video Panoptic Segmentation
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervise...
120
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[1/6] 📢 We are in Denver at #CVPR2026 presenting 5 papers!
176
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[3/3] Project page: visinf.github.io/KARMMA/ Poster (ICRA): Thursday, 03:00 PM, P207 (Hall C - ThI2I) Poster (CVPRW): Thursday, 10:00 AM, A2A-MML Workshop, Hall A
visinf.github.io
KARMMA
Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities.
010
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[2/3] KARMMA is a multimodal-to-multimodal distillation framework for egocentric action recognition that does not require modality-aligned data and supports any subset of modalities at inference. It produces a lightweight student robust to missing modalities without retraining.
110
Visual Inference Lab @visinf.bsky.social · 04/06/2026
[1/3] Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities by @dustin-carrion.bsky.social*, Maria Santos-Villafranca*, Alejandro Perez-Yus, Jesus Bermudez-Cameo, Jose J. Guerrero, and @simoneschaub.bsky.social
132
Reposted by Visual Inference Lab
KE:SAI - Kyutai ELLIS Scalable Autonomous Intelligence @kesai.eu · 20/05/2026
Today @kyutai-labs.bsky.social and @ellisinsttue.bsky.social launch @kesai.eu! Robot learning is bottlenecked by the cost of physical interaction. Our mission is to advance the efficiency frontier of robust & safe physical AI through fully open and reproducible research. kesai.eu/blog/2026-05...
kesai.eu
Kyutai and ELLIS Tübingen launch KE:SAI: A Premier Franco-German Partnership for Open Science in Physical AI
Kyutai and ELLIS Tübingen announce the official launch of KE:SAI – Kyutai ELLIS Scalable Autonomous Intelligence, a non-profit open science lab for physical AI.
11710
Reposted by Visual Inference Lab
ELLIS @ellis.eu · 21/04/2026
🎉 25 ELLIS Units have been successfully extended! Following our five-year reapplication process, we celebrate the sustained excellence these Units bring to European AI research. Congratulations to all! 👏 📖 Get more details: ellis.eu/news/25-elli...
0115
Reposted by Visual Inference Lab
DAGM GCPR 2026 @gcpr-by-dagm.bsky.social · 14/04/2026
📢 Call for Papers: #VMV2026 #GCPR2026 is joined by the International Symposium on Vision, Modeling, and Visualization! We invite high-quality submissions across all areas of visual computing. We are excited to see your work! 👉 Learn more & submit: www.gcpr-vmv.de/year/2026/vmv
gcpr-vmv.de
GCPR VMVGCPR VMV (VMV)
142
Reposted by Visual Inference Lab
DAGM GCPR 2026 @gcpr-by-dagm.bsky.social · 14/04/2026
#GCPR2026 is coming to #Siegen! 🎉 Join us for great research at the beautiful Campus Lower Castle of the University of Siegen, perfectly located just a short walk from the central station and many city hotels. More info: www.gcpr-vmv.de/year/2026/lo...
gcpr-vmv.de
GCPR VMVGCPR VMV (Location)
142
Reposted by Visual Inference Lab
Adam Kortylewski 🚨 Hiring PhDs @adamkortylewski.bsky.social · 10/04/2026
Got a paper on generative models accepted at CVPR 2026? Share it with us at the 4th Workshop on Generative Models for Computer Vision! generative-vision.github.io/workshop-CVP... You can simply submit your accepted CVPR paper, no need to reformat! Deadline: April 30 (AoE)
generative-vision.github.io
053
Reposted by Visual Inference Lab
Gabriele Trivigno @gabtriv.bsky.social · 07/04/2026
🔥 Can in-context segmentation emerge directly from frozen DINOv3 features? At #CVPR2026, we present INSID3: Training-Free In-Context Segmentation with DINOv3 — a collaboration between PoliTo, TU Darmstadt and TU Munich. Check it out: github.com/visinf/INSID3
1136
Reposted by Visual Inference Lab
DAGM GCPR 2026 @gcpr-by-dagm.bsky.social · 16/03/2026
We are excited to share our call for papers with a submission deadline on May 21st, 2026! We invite submissions of high-quality research papers presenting original contributions in all areas of pattern recognition! Read more: www.gcpr-vmv.de/year/2026/gc... #GCPR2026 #VMV2026
gcpr-vmv.de
GCPR VMVGCPR VMV (Call for Papers)
174
Visual Inference Lab @visinf.bsky.social · 05/03/2026
📢🎓 We have open PostDoc positions in Computer Vision & ML at @tuda.bsky.social and @hessianai.bsky.social within the Reasonable AI Cluster of Excellence — supervised by @stefanroth.bsky.social, @simoneschaub.bsky.social and many others! Apply here: www.career.tu-darmstadt.de/tu-darmstadt...
career.tu-darmstadt.de
044
Reposted by Visual Inference Lab
Stefan Roth @stefanroth.bsky.social · 12/12/2025
I am one of the potential advisors for this PhD position in our #DAAD funded AI graduate school ELIZA (eliza.school), which connects seven #ELLIS units in Germany. If you are interested in working with one of my colleagues or with me (in computer vision & deep learning), please consider applying.
062
Reposted by Visual Inference Lab
Nikita Araslanov @neekans.bsky.social · 02/12/2025
📢 NeurIPS 2025 Spotlight 📢 Can we embed motion into image representations? FlowFeat embeds optical flow into pixel-level representations, which results in sharp feature grids, especially for dynamic objects. Project website: tum-vision.github.io/flowfeat With Anna Sonnweber and Daniel Cremers.
062
Visual Inference Lab @visinf.bsky.social · 02/12/2025
📢 Join @tuda.bsky.social as a PhD/Postdoc in the new project HAICC - Human–AI Collaboration for Cybersecurity! Explore how LLM–based AI agents and humans can jointly analyse security data and rethink cybersecurity architectures - supervised by Iryna Gurevych, @stefanroth.bsky.social, and many more!
050
Reposted by Visual Inference Lab
Andreas Geiger @andreasgeiger.bsky.social · 02/12/2025
Attending #Neurips2025? Get your personalized Scholar Inbox conference program now to easily navigate the poster sessions and find what you are looking for: www.scholar-inbox.com/conference/n...
03612
Reposted by Visual Inference Lab
Robin Hesse @robinhesse.bsky.social · 19/11/2025
@neuripsconf.bsky.social is two weeks away! 📢 Stop missing great workshop speakers just because the workshop wasn’t on your radar. Browse them all in one place: robinhesse.github.io/workshop_spe... (also available for @euripsconf.bsky.social) #NeurIPS #EurIPS
195
Visual Inference Lab @visinf.bsky.social · 04/11/2025
📢🎓 We have open PhD positions in Computer Vision & Machine Learning at @tuda.bsky.social and @hessianai.bsky.social within the Reasonable AI Cluster of Excellence — supervised by @stefanroth.bsky.social, @simoneschaub.bsky.social and many others! www.career.tu-darmstadt.de/tu-darmstadt...
career.tu-darmstadt.de
086
Reposted by Visual Inference Lab
Simone Schaub-Meyer @simoneschaub.bsky.social · 21/10/2025
🎉 Today, Simon Kiefhaber will present our ICCV oral paper on how to make optical flow estimators more efficient (faster inference and lower memory usage) with state-of-the-art accuracy: 🌍 visinf.github.io/recover Talk: Tue 09:30 AM, Kalakaua Ballroom Poster: Tue 11:45 AM, Exhibit Hall I #76
visinf.github.io
Removing Cost Volumes from Optical Flow Estimators
We present ReCoVEr - A training strategy for removing the cost volume from an optical flow estimator during training.
051
Reposted by Visual Inference Lab
Christoph Reich @christophreich.bsky.social · 19/10/2025
Interested in 3D DINO features from a single image or unsupervised scene understanding?🦖 Come by our SceneDINO poster at NeuSLAM today 14:15 (Kamehameha II) or Tue, 15:15 (Ex. Hall I 627)! W/ Jevtić @fwimbauer.bsky.social @olvrhhn.bsky.social Rupprecht, @stefanroth.bsky.social @dcremers.bsky.social
083
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[8/8] 2nd Workshop on Explainable Computer Vision: Quo Vadis?
 🌍 excv-workshop.github.io 
Sun 8:00 AM — 5:00 PM, Ballroom A
010
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[7/8] Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation by @dustin-carrion.bsky.social, @stefanroth.bsky.social @simoneschaub.bsky.social 

 🌍 visinf.github.io/emat 📄 arxiv.org/abs/2507.23642 💻 github.com/visinf/emat Poster: Sun 4:40 PM - Exhibit Hall II #73
120
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[6/8] 📄 arxiv.org/abs/2509.02545
 Talk: Sun 9:30 AM, 306 A
 Poster: Sun 10:00 AM - Exhibit Hall II
110
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[6/8] Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery (Oral at ILR+G Workshop) by Xinrui Gong*, @olvrhhn.bsky.social *, @christophreich.bsky.social , Krishnakant Singh, @simoneschaub.bsky.social , @dcremers.bsky.social @stefanroth.bsky.social
131
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[5/8] ART: Adaptive Relation Tuning for Generalized Relation Prediction by Gopika Sudhakaran, Hikaru Shindo, Patrick Schramowski, @simoneschaub.bsky.social, @kerstingaiml.bsky.social, @stefanroth.bsky.social 📄 arxiv.org/abs/2507.23543 Poster: Wed 2:45 PM, Exhibit Hall I #1501
130
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[4/8] Activation Subspaces for Out-of-Distribution Detection by Baris Zongur, @robinhesse.bsky.social, @stefanroth.bsky.social 📄 arxiv.org/abs/2508.21695 Poster: Tue 11:45 AM, Exhibit Hall I #326
110
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[3/8] Poster: Tue 3:15 PM, Exhibit Hall I #627
Invited Invited poster at NeuSLAM workshop: Sun 2:15 PM, 304 A/Exhibit Hall II
110
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[3/8] Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion by @jev-aleks.bsky.social *, @christophreich.bsky.social *, @fwimbauer.bsky.social , @olvrhhn.bsky.social , Christian Rupprecht, @stefanroth.bsky.social, @dcremers.bsky.social 🌍 visinf.github.io/scenedino/
121
Visual Inference Lab @visinf.bsky.social · 19/10/2025
[2/8] Removing Cost Volumes from Optical Flow Estimators (Oral) by @skiefhaber.de , @stefanroth.bsky.social @simoneschaub.bsky.social 🌍 visinf.github.io/recover Talk: Tue 09:30 AM, Kalakaua Ballroom 
Poster: Tue 11:45 AM, Exhibit Hall I #76
120