Sign in

Marcus Klasson

@marcusklasson.bsky.social
221 followers 223 following 16 posts

Perception Researcher at Ericsson, Sweden. marcusklasson.github.io

PostsRepliesMedia
Marcus Klasson @marcusklasson.bsky.social · 24/04/2026
Code, models and paper is available on our project site! 🌐 aaltoml.github.io/BayesVLM/ 1k thanks to @antonbaumann.bsky.social, @ruili-pml.bsky.social, @smentu.bsky.social, @shyamgopal.bsky.social, @zeynepakata.bsky.social, @arnosolin.bsky.social, @trappmartin.bsky.social for this collaboration!🫰
aaltoml.github.io
BayesVLM
BayesVLM
011
Marcus Klasson @marcusklasson.bsky.social · 24/04/2026
BayesVLM improves calibration in zero-shot classification without sacrificing accuracy. The uncertainties are also useful for data selection in active fine-tuning, which actually was our target in the beginning, i.e. fetch new samples that will reduce the model's uncertainty on current observations.
100
Marcus Klasson @marcusklasson.bsky.social · 24/04/2026
The paper includes quite some math about estimating Hessians from CLIP efficiently which turned out harder as we initially thought due to the InfoNCE loss being cross-modal and contrastive. Anton, Rui, and Martin did a heckuva job to get this done in a rigorous way!
100
Marcus Klasson @marcusklasson.bsky.social · 24/04/2026
BayesVLM requires estimating Hessians over the image and text proj. layers, where access to the pretraining data (or proxy of this) is needed, but only 10 batches is sufficient. Estimating pseudo-data count and prior precision params is also needed, which is similar to what temp. scaling needs.
100
Marcus Klasson @marcusklasson.bsky.social · 24/04/2026
We put the Laplace approximation on top of CLIP to enable estimating uncertainties without inference overhead. AFAIK this is the first method that can produce zero-shot uncertainties in VLMs without architecture changes or retraining from scratch.
100
Marcus Klasson @marcusklasson.bsky.social · 24/04/2026
👋🇧🇷 If you are at #ICLR2026 today, you should talk to @antonbaumann.bsky.social who is presenting our paper about turning pre-trained VLMs into probabilistic models without retraining or fine-tuning. Poster Session 3 ⌚: 10:30am - 1:00pm (local time) 📍: Pavilion 3 P3 - #313 @iclr-conf.bsky.social
143
Reposted by Marcus Klasson
Martin Trapp @trappmartin.eurosky.social · 02/10/2025
Want to work on Trustworthy AI? 🚀 I'm seeking exceptional candidates to apply for the Digital Futures Postdoctoral Fellowship to work with me on Uncertainty Quantification, Bayesian Deep Learning, and Reliability of ML Systems. The position will be co-advised by Hossein Azizpour or Henrik Boström.
trappmartin.github.io
Home
Martin Trapp - Assistant Professor in Machine Learning at KTH Royal Institute of Technology.
1114
Marcus Klasson @marcusklasson.bsky.social · 13/06/2025
Paper, videos, and code (nerfstudio) is available! 📄 arxiv.org/abs/2411.19756 🎈 aaltoml.github.io/desplat/ Big ups to Yihao Wang, @maturk.bsky.social, Shuzhe Wang, Juho Kannala, and @arnosolin.bsky.social for making this possible during my time at @aalto.fi 💙🤍 #AaltoUniversity #CVPR2025 [8/8]
arxiv.org
DeSplat: Decomposed Gaussian Splatting for Distractor-Free Rendering
Gaussian splatting enables fast novel view synthesis in static 3D environments. However, reconstructing real-world environments remains challenging as distractors or occluders break the multi-view con...
031
Marcus Klasson @marcusklasson.bsky.social · 13/06/2025
DeSplat has the same FPS and training time as vanilla 3DGS with some additional overhead for storing distractor Gaussians. Extend with MLPs or other models can also be done. Altering DeSplat to video remains to be explored, as distractors barely moving across images can be mistaken as static. [7/8]
100
Marcus Klasson @marcusklasson.bsky.social · 13/06/2025
This decomposed splatting (DeSplat) approach explicitly separates distractors from static parts. Earlier methods (e.g. SpotlessSplats, WildGaussians) use loss masking of detected distractors to avoid overfitting, while DeSplat instead jointly reconstructs distractor elements. [6/8]
100
Marcus Klasson @marcusklasson.bsky.social · 13/06/2025
Knowing how 3DGS treats distractors, we initialize a set of Gaussians close to every camera view for reconstructing view-specific distractors. The Gaussians initialized from the point cloud should reconstruct static stuff. These separately rendered images are alpha-blended during training. [5/8]
100
Marcus Klasson @marcusklasson.bsky.social · 13/06/2025
In a viewer, you can see that these spurious artefacts are thin and are located close to the camera view. For the scene-overfitting approach in 3DGS, this makes sense since an object only appearing in one view must be located as close to the camera such that no other camera view can see it. [4/8]
100
Marcus Klasson @marcusklasson.bsky.social · 13/06/2025
This BabyYoda scene from RobustNeRF is similar to a crowdsourced scenario, where a set of static toys appear together with inconsistently-placed toys between the frames. Vanilla 3DGS is quite robust here, but some views end up being rendered with spurious artefacts (right image). [3/8]
100
Marcus Klasson @marcusklasson.bsky.social · 13/06/2025
Our goal is to learn a scene representation from images that include non-static objects we refer to as distractors. An example is crowdsourced images where different people appear at different locations in the scene, which creates multi-view inconsistencies between the frames. [2/8]
100
Marcus Klasson @marcusklasson.bsky.social · 13/06/2025
👋Interested in Gaussian splatting and removing dynamic content from images? Our DeSplat is presented today at #CVPR2025 at Poster Session 1, ExHall D Poster #52. Yihao will be there to present our fully splatting-based method for separating static and dynamic stuff in images. 🧵[1/8]
161
Reposted by Marcus Klasson
Andrei Bursuc @abursuc.bsky.social · 11/06/2025
You woke up early in the morning jet-lagged and having a hard time deciding for a workshop today @cvprconference.bsky.social ? Here's a reliable choice for you: our workshop on 🛟 Uncertainty Quantification for Computer Vision! 🗓️ Day: Wed, Jun 11 📍Room: 102 B #CVPR2025 #UNCV2025
093
Marcus Klasson @marcusklasson.bsky.social · 23/04/2025
KTH is looking for a *Postdoc* to work on visual domain adaptation for mobile robot perception in a joint project with Ericsson in Stockholm. Apply by May 15 if you are interested in working with computer vision applied to real robots! More info: www.kth.se/lediga-jobb/...
kth.se
KTH | Postdoc in robotics with specialization in visual domain adaptation
KTH jobs is where you search for jobs at www.kth.se.
021
Marcus Klasson @marcusklasson.bsky.social · 17/03/2025
Submission deadline is extended to March 20 for submitting your paper to our #CVPR2025 workshop on Uncertainty Quantification for Computer Vision. Looking forward to see your submissions on recognizing failure scenarios and enabling robust vision systems! More info: uncertainty-cv.github.io/2025/
uncertainty-cv.github.io
UNCV Workshop @ CVPR 2025
CVPR 2025 Workshop on Uncertainty Quantification for Computer Vision.
0105
Reposted by Marcus Klasson
Arno Solin @arnosolin.bsky.social · 08/03/2025
There is still time to submit your papers to our #CVPR2025 workshop on Uncertainty Quantification for Computer Vision, which is part of the workshop lineup at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) in Nashville, Tennessee.
2136
Reposted by Marcus Klasson
Andrei Bursuc @abursuc.bsky.social · 28/02/2025
Our Workshop on Uncertainty Quantification for Computer Vision goes to @cvprconference.bsky.social this year! We have a super line-up of speakers and a call for papers. This is a chance for your paper to shine at #CVPR2025 ⏲️ Submission deadline: 14 March 💻 Page: uncertainty-cv.github.io/2025/
0337
Reposted by Marcus Klasson
Martin Trapp @trappmartin.eurosky.social · 10/12/2024
I will present ✌️ BDU workshop papers @ NeurIPS: one by Rui Li (looking for internships) and one by Anton Baumann. 🔗 to extended versions: 1. 🙋 "How can we make predictions in BDL efficiently?" 👉 arxiv.org/abs/2411.18425 2. 🙋 "How can we do prob. active learning in VLMs" 👉 arxiv.org/abs/2412.06014
arxiv.org
Post-hoc Probabilistic Vision-Language Models
Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descripti...
1184