Sign in

Arkadiy Saakyan

@asaakyan.bsky.social
445 followers 226 following 24 posts

PhD student at Columbia University working on human-AI collaboration, AI creativity and explainability. prev. intern @GoogleDeepMind, @AmazonScience asaakyan.github.io

PostsRepliesMedia
Reposted by Arkadiy Saakyan
Advait@COLM @advaitdeshmukh.com · 10/09/2026
1/7 In 1941, Borges imagined an impossible novel that follows all narrative branches at once. Today, we can observe chatbot users exploring narrative possibilities as they repeatedly edit story prompts. In our COLM 2026 paper, we studied this branching exploration via the WildChat dataset. 🧵
Paraphrased edit trees from the dataset
24212
Reposted by Arkadiy Saakyan
Teagan Johnson @teagrjohnson.bsky.social · 19/06/2026
1/ LLMs learn narrative from their pretraining data but what narrative content is actually in there? It turns out narrative is wildly unevenly distributed across sources and topics. New preprint with @andrewpiper.bsky.social @elliottash.bsky.social @mariaa.bsky.social:
This image depicts the proportion of each Dolma category in the top quartile for the first three principal components.
56020
Arkadiy Saakyan @asaakyan.bsky.social · 02/06/2026
Excited to share #ICML2026 paper from my internship @ Google DeepMind! AI models are deployed globally, but AI safety datasets are largely geographically homogenous. What is the impact of culture on AI safety ratings? Is there any impact beyond standard demographics like age, gender, and ethnicity?
1113
Reposted by Arkadiy Saakyan
Kenny Peng @kennypeng.bsky.social · 24/04/2026
We made traversle.io, a new daily word game! The goal is to traverse from a start word to a target word through a network of related words. (Our motivating question: is it possible to construct a network that allows human navigation?)
3164
Reposted by Arkadiy Saakyan
Gabriel Agostini @gsagostini.bsky.social · 01/05/2026
Excited to see MIGRATE recognized in the IPUMS awards! Huge thanks to @emmapierson.bsky.social, @nkgarg.bsky.social, and our coauthors. Our work primarily aims to make spatiotemporal data more trustworthy and accessible to researchers, just like IPUMS. Read the paper to request data access!
094
Reposted by Arkadiy Saakyan
Urban Data @urban-data.bsky.social · 23/04/2026
We are co-hosting the EAAMO colloquium next Monday (12pm EST) with Professor Rachel Franklin. Come hear her talk about spatial inequality and the smart city and feel free to share with colleagues! Register below to get the Zoom link: www.eaamo.org/colloquium/r...
032
Arkadiy Saakyan @asaakyan.bsky.social · 21/04/2026
Excited to travel to ICLR 🇧🇷 to present our work on textual creativity metrics! See me at the 10:30am-1pm Saturday poster session in Pavilion 3 or reach out to chat 🙂
010
Reposted by Arkadiy Saakyan
Tuhin Chakrabarty @tuhinchakr.bsky.social · 26/03/2026
🚨Paper on AI & Copyright Courts have credited AI companies' claims that alignment prevents reproducing copyrighted data. What if finetuning on a simple writing task breaks it. Worse: tuning on just one author (e.g., Murakami) unlocks verbatim recall of 30+ other authors' books (up to 90%) (1/n)🧵
25630
Reposted by Arkadiy Saakyan
Manya Wadhwa @manyawadhwa.bsky.social · 13/03/2026
⚛️ Introducing CREATE, a benchmark for creative associative reasoning in LLMs. Making novel, meaningful connections is key for scientific & creative works. We objectively measure how well LLMs can do this. 🧵👇
24113
Reposted by Arkadiy Saakyan
Gabriel Agostini @gsagostini.bsky.social · 05/02/2026
Our paper “Inferring fine-grained migration patterns across the United States” is now out in @natcomms.nature.com! We released a new, highly granular migration dataset. 1/9
27327
Reposted by Arkadiy Saakyan
Merriam-Webster @merriam-webster.com · 15/12/2025
Merriam-Webster’s human editors have chosen ‘slop’ as the 2025 Word of the Year.
357238387190
Reposted by Arkadiy Saakyan
Urban Data @urban-data.bsky.social · 13/11/2025
Day 7 of #30DayMapChallenge asked us to think about accessibility. @gsagostini.bsky.social considers two metrics of access simultaneously: distance to a Subway and distance to the subway.
410018
Arkadiy Saakyan @asaakyan.bsky.social · 04/11/2025
N-gram novelty is widely used as a measure of creativity and generalization. But if LLMs produce highly n-gram novel expressions that don’t make sense or sound awkward, should they still be called creative? In a new paper, we investigate how n-gram novelty relates to creativity.
N-gram novelty is widely used to evaluate language models' ability to generate text outside of their training data. More recently, it has also been adopted as a metric for measuring textual creativity. However, theoretical work on creativity suggests that this approach may be inadequate, as it does not account for creativity's dual nature: novelty (how original the text is) and appropriateness (how sensical and pragmatic it is). We investigate the relationship between this notion of creativity and n-gram novelty through 7542 expert writer annotations (n=26) of novelty, pragmaticality, and sensicality via close reading of human and AI-generated text. We find that while n-gram novelty is positively associated with expert writer-judged creativity, ~91% of top-quartile expressions by n-gram novelty are not judged as creative, cautioning against relying on n-gram novelty alone. Furthermore, unlike human-written text, higher n-gram novelty in open-source LLMs correlates with lower pragmaticality. In an exploratory study with frontier close-source models, we additionally confirm that they are less likely to produce creative expressions than humans. Using our dataset, we test whether zero-shot, few-shot, and finetuned models are able to identify creative expressions (a positive aspect of writing) and non-pragmatic ones (a negative aspect). Overall, frontier LLMs exhibit performance much higher than random but leave room for improvement, especially struggling to identify non-pragmatic expressions. We further find that LLM-as-a-Judge novelty scores from the best-performing model were predictive of expert writer preferences.
14110
Reposted by Arkadiy Saakyan
Gabriel Agostini @gsagostini.bsky.social · 03/09/2025
Are you a researcher using computational methods to understand cities? @mfranchi.bsky.social @jennahgosciak.bsky.social and I organize an EAAMO Bridges working group on Urban Data Science and we are looking for new members! Fill the interest form on our page: urban-data-science-eaamo.github.io
urban-data-science-eaamo.github.io
Urban Data Science & Equitable Cities | EAAMO Bridges
EAAMO Bridges Urban Data Science & Equitable Cities working group: biweekly talks, paper studies, and workshops on computational urban data analysis to explore and address inequities.
188
Reposted by Arkadiy Saakyan
Daniel Scalena @danielsc4.it · 23/05/2025
📢 New paper: Applied interpretability 🤝 MT personalization! We steer LLM generations to mimic human translator styles on literary novels in 7 languages. 📚 SAE steering can beat few-shot prompting, leading to better personalization while maintaining quality. 🧵1/
2205
Arkadiy Saakyan @asaakyan.bsky.social · 01/05/2025
Can vision-language models understand figurative meaning in multimodal inputs, like visual metaphors, sarcastic captions or memes? Come find out at our #NAACL2025 poster on Friday at 9am! New task & dataset of images and captions with figurative phenomena like metaphor, idiom, sarcasm, and humor.
162
Reposted by Arkadiy Saakyan
Gabriel Agostini @gsagostini.bsky.social · 28/03/2025
Migration data lets us study responses to environmental disasters, social change patterns, policy impacts, etc. But public data is too coarse, obscuring these important phenomena! We build MIGRATE: a dataset of yearly flows between 47 billion pairs of US Census Block Groups. 1/5
54118
Reposted by Arkadiy Saakyan
Jenna Russell @jennarussell.bsky.social · 28/01/2025
People often claim they know when ChatGPT wrote something, but are they as accurate as they think? Turns out that while general population is unreliable, those who frequently use ChatGPT for writing tasks can spot even "humanized" AI-generated text with near-perfect accuracy 🎯
1018966