Sign in

Fabian David Schmidt

@fdschmidt.bsky.social
195 followers 54 following 5 posts

PhD candidate at Uni of Würzburg working on multilinguality & multimodality | prev. visited visit Mila & LTL@UniCambridge fdschmidt93.github.io

PostsRepliesMedia
Fabian David Schmidt @fdschmidt.bsky.social · 21/02/2025
Cross-modal topic matching correlates well with other multilingual vision-language tasks! 🤗Images-To-Sentence (given Images, select topically fitting sentence) & Sentences-To-Image (given Sentences, pick topically matching image) probe complementary aspects in VLU
121
Fabian David Schmidt @fdschmidt.bsky.social · 21/02/2025
X-modal to text-only perf. *gap* shows that VL support decreases from high to low-resource language tiers: Images/Topic→Sentence (for I/T, pick S): narrows with less textual support (left) Sentences→Image/Topic (for S, pick I/T): increases with less VL support worse (right)
111
Fabian David Schmidt @fdschmidt.bsky.social · 21/02/2025
Strong vision-language models (VLMs) like GPT-4o-mini maintain good performance for top-150 languages, only to drop to performing no better than chance for the lowest resource languages!
111
Fabian David Schmidt @fdschmidt.bsky.social · 21/02/2025
Introducing MVL-SIB, a massively multilingual vision-language benchmark for cross-modal topic matching in 205 languages! 🤔Tasks: Given images (sentences), select topically matching sentence (image). Arxiv: arxiv.org/abs/2502.12852 HF: huggingface.co/datasets/Wue... Details👇
145