Reposted by Gabriel AgostiniKenny Peng @kennypeng.bsky.social · 18/08/2026New York City 8th graders choose from 900+ high schools to apply to, in a process that’s spawned Facebook groups and dozens of expensive consulting services. Our new paper shows how application behavior leads to disparities, and how to effectively intervene. 🧵 www.nature.com/articles/s44... 1329
Reposted by Gabriel AgostiniIra Globus-Harris @iraglobusharris.bsky.social · 03/07/2026Are you at ICML next week? Feel like your decision-making for which sessions to attend might not be risk minimizing? Don't incur (swap) regret and come to my, @aaroth.bsky.social, and @ncollina.bsky.social's tutorial Monday on multicalibration, decision-making, and collaborative learning! 22011
Reposted by Gabriel AgostiniArkadiy Saakyan @asaakyan.bsky.social · 02/06/2026Excited to share #ICML2026 paper from my internship @ Google DeepMind! AI models are deployed globally, but AI safety datasets are largely geographically homogenous. What is the impact of culture on AI safety ratings? Is there any impact beyond standard demographics like age, gender, and ethnicity? 1113
Gabriel Agostini @gsagostini.bsky.social · 01/05/2026Excited to see MIGRATE recognized in the IPUMS awards! Huge thanks to @emmapierson.bsky.social, @nkgarg.bsky.social, and our coauthors. Our work primarily aims to make spatiotemporal data more trustworthy and accessible to researchers, just like IPUMS. Read the paper to request data access! 094
Reposted by Gabriel AgostiniKenny Peng @kennypeng.bsky.social · 24/04/2026We made traversle.io, a new daily word game! The goal is to traverse from a start word to a target word through a network of related words. (Our motivating question: is it possible to construct a network that allows human navigation?) 3164
Reposted by Gabriel AgostiniUrban Data @urban-data.bsky.social · 23/04/2026We are co-hosting the EAAMO colloquium next Monday (12pm EST) with Professor Rachel Franklin. Come hear her talk about spatial inequality and the smart city and feel free to share with colleagues! Register below to get the Zoom link: www.eaamo.org/colloquium/r... 032
Reposted by Gabriel AgostiniKenny Peng @kennypeng.bsky.social · 26/03/2026Excited to share our new research demo, where you can freely traverse the world of Bluesky through 20,000 interconnected trails, spanning “analysis of fictional tropes” to “rotisserie chicken” to “zoning and land use policy.” Try it out, and let us know what you think!skytrails.orgskytrails · 20,000 trails through BlueskyCan we regain freedom of movement on social media? Browse Bluesky via interconnected trails. 6459
Reposted by Gabriel AgostiniDivya Shanmugam @dmshanmugam.bsky.social · 23/03/2026New in Nature Health: how might we move towards a world in which race is not used in clinical algorithms? We need (1) careful comparison of race-aware and race-neutral algorithms and (2) systemic efforts to address underlying disparities. 1219
Gabriel Agostini @gsagostini.bsky.social · 18/03/2026Had a great time presenting our work on building MIGRATE–a new dataset of US migration–at the @geographers.bsky.social AAG Annual Meeting today. Happy to also share that we received an AAG student paper award for this work!!! Come chat if you are at #AAG26 this week. migrate.tech.cornell.edu 0123
Reposted by Gabriel AgostiniKenny Peng @kennypeng.bsky.social · 17/02/2026New paper! The Linear Representation Hypothesis is a powerful intuition for how language models work, but lacks formalization. We give a mathematical framework in which we can ask and answer a basic question: how many features can be stored under the hypothesis? 🧵 arxiv.org/abs/2602.11246 14514
Reposted by Gabriel AgostiniCornell Tech @cornelltech.bsky.social · 05/02/2026New research is offering new insight on how Americans move — all the way to the neighborhood level. A new dataset, MIGRATE, maps annual moves with 4,600‑times more detail than standard public data, revealing patterns hidden in county‑level reporting: bit.ly/49XSD6w 052
Gabriel Agostini @gsagostini.bsky.social · 05/02/2026Our paper “Inferring fine-grained migration patterns across the United States” is now out in @natcomms.nature.com! We released a new, highly granular migration dataset. 1/9 27327
Reposted by Gabriel AgostiniUrban Data @urban-data.bsky.social · 09/12/2025November is over, but we still have some #30DayMapChallenge entries to share! And for our transport-themed day 26 map, MBTA data analyst Joe Hilleary takes us on a ride back in time: he shows current bus routes in Greater Boston by the earliest known year in which a direct percursor route ran a bus. 171
Reposted by Gabriel AgostiniEmma Pierson @emmapierson.bsky.social · 24/11/2025We have a new paper in Science Advances proposing a simple test for bias: Is the same person treated differently when their race is perceived differently? Specifically, we study: is the same driver likelier to be searched by police when they are perceived as Hispanic rather than white? 1/ 24316
Gabriel Agostini @gsagostini.bsky.social · 24/11/2025My best workflow improvement since starting to work with spatial libraries in Python was to always include a `crs` dictionary on a variables file listing crs for lat-long projections, equidistant projections, and "maybe not satisfying any desiderata but the prettiest out there" projections. 020
Reposted by Gabriel AgostiniUrban Data @urban-data.bsky.social · 24/11/2025#30DayMapChallenge day 20: water This map of Arsenic and Cadmium levels in Mexico's water show non-trace concentrations of Total and Soluble Arsenic and Cadium. Points are colored by the presence of high amounts of contaminants, and sized by their relative concentration. tinyurl.com/map20wtr 182
Reposted by Gabriel AgostiniUrban Data @urban-data.bsky.social · 21/11/2025#30DayMapChallenge 15: Fire @sylviaimani.bsky.social visualized how Uganda’s transition toward electric cooking aligns with the reach of the national grid. Regions with denser grid networks show a strong correlation with higher household adoption of electric cooking technologies. 1122
Reposted by Gabriel AgostiniKyle Walker @kylewalker.bsky.social · 18/11/2025For #30DayMapChallenge Day 18: Out of this world, use the `fill_z_offset` param in mapgl to "elevate" your data. Just be careful - if you choose a value too high, you might lose your data in the sky! #rstats 072
Gabriel Agostini @gsagostini.bsky.social · 17/11/2025I took me too long to accept that "Amsterdam is just 10th Ave" 010
Reposted by Gabriel AgostiniUrban Data @urban-data.bsky.social · 15/11/2025#30DayMapChallenge day 10: Air @jessiefin.bsky.social + Francisco Marmolejo-Cossío visualize the presence of ladrilleras, or brick kilns, which emit pollution across the state. Data cleaned by Jacqueline Calderón and Lizet Jarquin at UASLP. Full interactive map: tinyurl.com/map10-air 182
Reposted by Gabriel AgostiniUrban Data @urban-data.bsky.social · 14/11/2025#30DayMapChallenge day 9 asked us to get off our screens. @annaloganmc.bsky.social's "analog" map is a hand-painted postcard! 📫 "I chose to paint a postcard of a map of Ann Arbor where I currently live showing the Huron River!" she says 0135
Gabriel Agostini @gsagostini.bsky.social · 11/11/2025Great map(s) by @jennahgosciak.bsky.social ---can we count that for 6 days of mapping??---that show both the permanence and the vulnerability of ecological concepts in our urban landscapes! #30DayMapChallenge 040
Reposted by Gabriel AgostiniUrban Data @urban-data.bsky.social · 07/11/2025We are slowly catching up to the #30DayMapChallenge! In our day 3: polygons submission, @zhixuanqi.bsky.social questioned the boundaries and fuzziness of polygons with an animated map that invites us to think about the (not-so-well-defined) idea of neighborhoods. 0164
Gabriel Agostini @gsagostini.bsky.social · 06/11/2025Another dog map, this is 1 dog = 1 dot. And hopefully 1 day = 1 map for the next 30 days in our working group page 🗺️ 010
Reposted by Gabriel AgostiniArkadiy Saakyan @asaakyan.bsky.social · 04/11/2025N-gram novelty is widely used as a measure of creativity and generalization. But if LLMs produce highly n-gram novel expressions that don’t make sense or sound awkward, should they still be called creative? In a new paper, we investigate how n-gram novelty relates to creativity. 14110
Reposted by Gabriel AgostiniDivya Shanmugam @dmshanmugam.bsky.social · 17/10/2025New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods. 1217
Gabriel Agostini @gsagostini.bsky.social · 14/10/2025Very happy Divya has been around during my PhD. I might be deep into maps and she might be deep into health (...and so much more!) but I could always count on learning something from her. She's such a kind researcher and great science communicator! 140
Gabriel Agostini @gsagostini.bsky.social · 23/09/2025New version of our preprint! More about the project and data access on our website migrate.tech.cornell.edumigrate.tech.cornell.eduMIGRATE 060
Gabriel Agostini @gsagostini.bsky.social · 03/09/2025Are you a researcher using computational methods to understand cities? @mfranchi.bsky.social @jennahgosciak.bsky.social and I organize an EAAMO Bridges working group on Urban Data Science and we are looking for new members! Fill the interest form on our page: urban-data-science-eaamo.github.iourban-data-science-eaamo.github.ioUrban Data Science & Equitable Cities | EAAMO BridgesEAAMO Bridges Urban Data Science & Equitable Cities working group: biweekly talks, paper studies, and workshops on computational urban data analysis to explore and address inequities. 188
Reposted by Gabriel AgostiniNikhil Garg @nkgarg.bsky.social · 28/08/2025*Proud advisor moment* My (first) PhD student Zhi Liu (zhiliu724.github.io) is 1 of 4 finalists for the INFORMS Dantzig Dissertation Award, the premier dissertation award for the OR community. His dissertation spanned work with 2 NYC govt agencies, on measuring and mitigating operational inequitieszhiliu724.github.ioZhi LiuAbout me 1293
Reposted by Gabriel AgostiniStreetsblog NYC @nyc.streetsblog.org · 30/06/2025"Removing the protected bike lane won’t remove cyclists — it will only make the street less safe," the Department of Transportation said in new testimony. "The city risks legal liability for knowingly reducing safety on a Vision Zero priority corridor." buff.ly/QNgRytsnyc.streetsblog.orgDOT Testimony: Removing Bedford Ave. Bike Lane Will 'Reduce Safety' - Streetsblog New York City"Removing the protected bike lane won’t remove cyclists — it will only make the street less safe," the DOT said. "The city risks legal liability for knowingly reducing safety on a Vision Zero… 28720
Reposted by Gabriel AgostiniDivya Shanmugam @dmshanmugam.bsky.social · 14/06/2025New work 🎉: conformal classifiers return sets of classes for each example, with a probabilistic guarantee the true class is included. But these sets can be too large to be useful. In our #CVPR2025 paper, we propose a method to make them more compact without sacrificing coverage. 3226
Reposted by Gabriel AgostiniErica Chiang @ericachiang.bsky.social · 01/05/2025I’m really excited to share the first paper of my PhD, “Learning Disease Progression Models That Capture Health Disparities” (accepted at #CHIL2025)! ✨ 1/ 📄: arxiv.org/abs/2412.16406 33610
Reposted by Gabriel AgostiniArkadiy Saakyan @asaakyan.bsky.social · 01/05/2025Can vision-language models understand figurative meaning in multimodal inputs, like visual metaphors, sarcastic captions or memes? Come find out at our #NAACL2025 poster on Friday at 9am! New task & dataset of images and captions with figurative phenomena like metaphor, idiom, sarcasm, and humor. 162
Gabriel Agostini @gsagostini.bsky.social · 02/04/2025I became a dog scientist on April 1st. Now back to normal (a cat scientist). 090
Gabriel Agostini @gsagostini.bsky.social · 28/03/2025Migration data lets us study responses to environmental disasters, social change patterns, policy impacts, etc. But public data is too coarse, obscuring these important phenomena! We build MIGRATE: a dataset of yearly flows between 47 billion pairs of US Census Block Groups. 1/5 54118
Reposted by Gabriel AgostiniRaj Movva @rajmovva.bsky.social · 18/03/2025💡New preprint & Python package: We use sparse autoencoders to generate hypotheses from large text datasets. Our method, HypotheSAEs, produces interpretable text features that predict a target variable, e.g. features in news headlines that predict engagement. 🧵1/ 14013