Sign in

Caleb Ziems

@calebziems.com
1.8K followers 483 following 17 posts

PhD student at Stanford NLP. Working on Social NLP and CSS. Previously at GaTech, Meta AI, Emory. 📍Palo Alto, CA 🔗 calebziems.com

PostsRepliesMedia
Caleb Ziems @calebziems.com · 04/11/2025
Our implementation of Culture Cartography is based on Farsight (Wang et al., 2024). This was an interdisciplinary effort across computer science (@diyiyang.bsky.social, @williamheld.com, Jane Yu) and sociology (David Grusky and Amir Goldberg), and the research process taught me so much!
100
Caleb Ziems @calebziems.com · 04/11/2025
Finally, Culture Cartography is aligned with prior notions of culture evals in our field. We observe positive transfer performance from Cartography to two leading benchmarks: BLEnD (Myung et al., 2024) and CulturalBench (Chiu et al., 2024).
100
Caleb Ziems @calebziems.com · 04/11/2025
Compared to knowledge extraction, Culture Cartography is less prone to test-set contamination. We evaluate GPT-4o with and without search and find no significant difference in their recall on Cartography data. Culture Cartography is "Google proof" since search doesn't help.
100
Caleb Ziems @calebziems.com · 04/11/2025
Compared to traditional annotation, Culture Cartography more often elicits knowledge that is unknown to LLMs. Qwen-2 72 B recalls 21% less Cartography data than it recalls traditional data (p < .0001) Even a strong reasoning model (R1) is challenged more by our data.
100
Caleb Ziems @calebziems.com · 04/11/2025
We propose a mixed-initiative method called Culture Cartography. And to find challenging questions, we let the LLM steer towards topics it has low confidence in. To find culturally-representative knowledge, we let the human steer towards what they find most salient.
100
Caleb Ziems @calebziems.com · 04/11/2025
Other benchmarks use knowledge extracted from the rich cultural artifacts that humans actively produce on the web. Still this is a single-initiative process. Researchers can’t steer the distribution towards questions of interest (i.e., those that challenge LLMs).
110
Caleb Ziems @calebziems.com · 04/11/2025
How are prior benchmarks constructed? In traditional annotation, the researcher picks some questions and the annotator passively provides ground truth answers. This is single-initiative. Annotators don't steer the process, so their interests and culture may not be represented.
100
Caleb Ziems @calebziems.com · 04/11/2025
Can we map out gaps in LLMs’ cultural knowledge? Check out our #EMNLP2025 talk: Culture Cartography 🗓️ 11/5, 11:30 AM 📌 A109 (CSS Orals 1) Compared to traditional benchmarking, our mixed-initiative method finds more knowledge gaps even in reasoning models like R1! Paper: arxiv.org/pdf/2510.27672
111