Sign in

simonroschmann.bsky.social

@simonroschmann.bsky.social
33 followers 12 following 17 posts

PhD Student @eml-munich.bsky.social @tum.de @www.helmholtz-munich.de‬. Passionate about ML research.

PostsRepliesMedia
Reposted by @simonroschmann.bsky.social
ExplainableML @eml-munich.bsky.social · 03/07/2026
3/ TiViT: Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers @simonroschmann.bsky.social , @qbouniot.bsky.social , Vasilii Feofanov, Ievgen Redko, @zeynepakata.bsky.social
arxiv.org
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
Time series classification is a fundamental task in healthcare and industry, yet the development of time series foundation models (TSFMs) remains limited by the scarcity of publicly available time ser...
101
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
#ICML2026 #MachineLearning #MultimodalLearning #RepresentationLearning #OptimalTransport #VisionLanguage #Alignment
010
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
@eml-munich.bsky.social @helmholtzmunich.bsky.social @tum.de @munichcenterml.bsky.social @telecomparis.bsky.social @polytechniqueparis.bsky.social
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
If you attend ICML 2026 in Seoul next week, I would be happy to connect and chat about multimodal alignment as well as my current research interest in multimodal agentic AI.
130
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
Huge thank you to my collaborators @paulkrz.bsky.social, @soniamazelet.bsky.social, @qbouniot.bsky.social, and to my PhD supervisor @zeynepakata.bsky.social for the guidance and support throughout my first year.
110
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
SOTAlign requires 4x less paired supervision, consistently improves as more unpaired samples are added, remains robust under distribution shifts between paired and unpaired data, and outperforms supervised and semi-supervised baselines in zero-shot classification and retrieval.
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
A key component is our KLOT divergence, which enables the comparison of unpaired images and text while avoiding backpropagation through Sinkhorn.
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
We propose a semi-supervised framework that first aligns pretrained unimodal models using a small set of paired samples and then leverages optimal transport to refine the initial alignment with large-scale unpaired data.
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
Motivated by the Platonic Representation Hypothesis, we study to what extent pretrained models from different modalities already learn similar representations, and how this can be leveraged to reduce supervision for alignment.
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
How to align unimodal encoders if paired data is scarce but unpaired data is abundant?
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
Excited to share that our paper "SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport" was accepted at ICML. 📄 Paper: arxiv.org/abs/2602.23353 💻 Code: github.com/ExplainableM...
121
Reposted by @simonroschmann.bsky.social
ExplainableML @eml-munich.bsky.social · 29/06/2026
Happy to share that we have 8 papers accepted to #ICML2026! 🎉 Most of the authors will be traveling to Seoul, Korea 🇰🇷 — feel free to reach out and chat with them if you’re interested in the work. See the full thread below for our papers 👇
1106
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
Please check out our work! 📄 Preprint: arxiv.org/abs/2506.08641 💻 Code: github.com/ExplainableM... #TimeSeries #VisionTransformer #FoundationModel
github.com
GitHub - ExplainableML/TiViT: Time Vision Transformer
Time Vision Transformer. Contribute to ExplainableML/TiViT development by creating an account on GitHub.
010
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
This project was a collaboration between @eml-munich.bsky.social and Huawei Paris Noah’s Ark Lab. Thank you to my collaborators @qbouniot.bsky.social, Vasilii Feofanov, Ievgen Redko, and particularly to my advisor @zeynepakata.bsky.social for guiding me through my first PhD project!
132
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
TiViT is on par with TSFMs (Mantis, Moment) on the UEA benchmark and significantly outperforms them on the UCR benchmark. The representations of TiViT and TSFMs are complementary; their combination yields SOTA classification results among foundation models.
110
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
We further explore the structure of TiViT representations and find that intermediate layers with high intrinsic dimension are the most effective for time series classification.
110
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
Time Series Transformers typically rely on 1D patching. We show theoretically that the 2D patching applied in TiViT can increase the number of label-relevant tokens and reduce the sample complexity.
110
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
Our Time Vision Transformer (TiViT) converts a time series into a grayscale image, applies 2D patching, and utilizes a pretrained frozen ViT for feature extraction. We average the representations from a specific hidden layer and only train a linear classifier.
110
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
How can we circumvent data scarcity in the time series domain? We propose to leverage pretrained ViTs (e.g., CLIP, DINOv2) for time series classification and outperform time series foundation models (TSFMs). 📄 Preprint: arxiv.org/abs/2506.08641 💻 Code: github.com/ExplainableM...
162