Sign in

simonroschmann.bsky.social

@simonroschmann.bsky.social
33 followers 12 following 17 posts

PhD Student @eml-munich.bsky.social @tum.de @www.helmholtz-munich.de‬. Passionate about ML research.

PostsRepliesMedia
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
SOTAlign requires 4x less paired supervision, consistently improves as more unpaired samples are added, remains robust under distribution shifts between paired and unpaired data, and outperforms supervised and semi-supervised baselines in zero-shot classification and retrieval.
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
A key component is our KLOT divergence, which enables the comparison of unpaired images and text while avoiding backpropagation through Sinkhorn.
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
We propose a semi-supervised framework that first aligns pretrained unimodal models using a small set of paired samples and then leverages optimal transport to refine the initial alignment with large-scale unpaired data.
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2026
How to align unimodal encoders if paired data is scarce but unpaired data is abundant?
100
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
TiViT is on par with TSFMs (Mantis, Moment) on the UEA benchmark and significantly outperforms them on the UCR benchmark. The representations of TiViT and TSFMs are complementary; their combination yields SOTA classification results among foundation models.
110
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
We further explore the structure of TiViT representations and find that intermediate layers with high intrinsic dimension are the most effective for time series classification.
110
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
Our Time Vision Transformer (TiViT) converts a time series into a grayscale image, applies 2D patching, and utilizes a pretrained frozen ViT for feature extraction. We average the representations from a specific hidden layer and only train a linear classifier.
110
simonroschmann.bsky.social @simonroschmann.bsky.social · 03/07/2025
How can we circumvent data scarcity in the time series domain? We propose to leverage pretrained ViTs (e.g., CLIP, DINOv2) for time series classification and outperform time series foundation models (TSFMs). 📄 Preprint: arxiv.org/abs/2506.08641 💻 Code: github.com/ExplainableM...
162