Reposted by @yuluqin.bsky.social

What is the interplay between representations learned from (language) surface forms alone, and those learned from more grounded evidence (e.g.,vision)?
Excited to share new work understanding “Cross-modal taxonomic generalization” in (V)LMs
arxiv.org/abs/2603.07474
1/