Sign in

yuluqin.bsky.social

@yuluqin.bsky.social
17 followers 18 following 15 posts
PostsRepliesMedia
Reposted by @yuluqin.bsky.social
Kanishka Misra @kanishka.bsky.social · 10/03/2026
What is the interplay between representations learned from (language) surface forms alone, and those learned from more grounded evidence (e.g.,vision)? Excited to share new work understanding “Cross-modal taxonomic generalization” in (V)LMs arxiv.org/abs/2603.07474 1/
title section of the paper: “Cross-Modal Taxonomic Generalization in (Vision) Language Models” by Tianyang Xu, Marcelo Sandoval-Castañeda, Karen Livescu, Greg Shakhnarovich, Kanishka Misra.
13311
yuluqin.bsky.social @yuluqin.bsky.social · 22/07/2025
Does vision training change how language is represented and used in meaningful ways?🤔The answer is a nuanced yes! Comparing VLM-LM minimal pairs, we find that while the taxonomic organization of the lexicon is similar, VLMs are better at _deploying_ this knowledge. [1/9]
1183