Reposted by Matthew KowalHarry Thasarathan @hthasarathan.bsky.social · 07/02/2025Our method reveals model-specific features too: DinoV2 (left) shows specialized geometric concepts (depth, perspective), while SigLIP (right) captures unique text-aware visual concepts. This opens new paths for understanding model differences! (7/9) 162
Reposted by Matthew KowalKosta Derpanis @csprofkgd.bsky.social · 07/02/2025Discover how our new mechanistic interpretability work uncovers universal concepts. Check it out on arXiv! 0403
Reposted by Matthew KowalHarry Thasarathan @hthasarathan.bsky.social · 07/02/2025🌌🛰️🔭Wanna know which features are universal vs unique in your models and how to find them? Excited to share our preprint: "Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment"! arxiv.org/abs/2502.03714 (1/9) 15617