Sign in

Greg Durrett

@gregdnlp.bsky.social
2.6K followers 413 following 28 posts

CS professor at NYU. Large language models and NLP. he/him

PostsRepliesMedia
Reposted by Greg Durrett
Conference on Language Modeling @colmweb.org · 31/07/2026
📢 New blog post describing the COLM 2026 PCs’ analyses of AI use in submitted papers. gregdurrett.github.io/colm2026-blo... (temporary home, new COLM website coming soon!)
Title and first few paragraphs of the blog post "AI submissions at COLM: theoryslop, slopterpretability, and papers in the age of agents"Recommendations from the blog post
1134
Reposted by Greg Durrett
NSF-Simons AI Institute for Cosmic Origins (CosmicAI) @nsfsimonscosmicai.bsky.social · 15/06/2026
Research highlight! CosmicAI Researchers Wenxuan Ding (NYU), @gregdnlp.bsky.social (NYU) as well as external collaborator Nicholas Tomlin (NYU, TTIC) investigated whether LLM agents like Claude Code and OpenAI Codex can navigate cost-benefit tradeoffs in their actions. youtube.com/shorts/GMZ5z...
youtube.com
Can LLM agents like Claude Code & OpenAI Codex can navigate cost-benefit tradeoffs in their actions?
YouTube video by NSF-Simons AI Institute for Cosmic Origins
011
Greg Durrett @gregdnlp.bsky.social · 15/06/2026
Check out Ramya's work on analyzing why and how LLM-generated stories feel homogeneous: the setting you prompt with might be novel but the plot unfolds in a very conventional way. Thread for how we quantified this & compare to existing metrics: 👇
031
Reposted by Greg Durrett
Conference on Language Modeling @colmweb.org · 31/03/2026
❗The full paper submission deadline for COLM is ~14 hours from now (11:59pm AOE)! Please submit your final PDFs on the same page where you uploaded your abstracts. And please use the provided LaTeX templates; do not handwrite your manuscript like this llama is! Good luck!
A llama sweating while writing a paper at a desk. A sign says "Deadline! March 31 11:59pm AOE"
094
Greg Durrett @gregdnlp.bsky.social · 25/03/2026
ICML reviews have you considering this? Please look at our final submission instructions for COLM below!
Image of a llama driving a car taking an exit towards COLM instead of continuing to ICML.
071
Greg Durrett @gregdnlp.bsky.social · 13/03/2026
Check out Manya's benchmark for LLM creativity! Inspired by work on creativity in graphs (@adtraghunathan.bsky.social 's "roll the dice" paper), CREATE isolates testing of creative insights for discovery. Future: understanding how LLMs derive insights & how they can be better creative partners!
190
Greg Durrett @gregdnlp.bsky.social · 23/02/2026
Check out Wenxuan's work on cost-uncertainty tradeoffs in agents! Providing different info to the model (here, estimates of its uncertainty) triggers a really different type of reasoning. Agent behavior under a harness reflects what info is included, not just the instructions.
000
Greg Durrett @gregdnlp.bsky.social · 12/01/2026
Still accepting applications for this postdoc position in my lab at NYU! Applications due Feb 1. Please see the posting below for more information and apply on Interfolio: cims.nyu.edu/taur/postdoc... apply.interfolio.com/178940
cims.nyu.edu
043
Greg Durrett @gregdnlp.bsky.social · 16/12/2025
Submit to COLM! Deadline of March 31. This llama gets to enjoy his holidays and isn't stressed out just yet...
071
Reposted by Greg Durrett
Ahmad Beirami @abeirami.bsky.social · 02/12/2025
Hiring researchers & engineers to work on –building reliable software on top of unreliable LLM primitives –statistical evaluation of real-world deployments of LLM-based systems I’m speaking about this on two NeurIPS workshop panels: 🗓️Saturday – Reliable ML Workshop 🗓️Sunday – LLM Evaluation Workshop
2195
Greg Durrett @gregdnlp.bsky.social · 02/12/2025
📢 Postdoc position 📢 I’m recruiting a postdoc for my lab at NYU! Topics include LM reasoning, creativity, limitations of scaling, AI for science, & more! Apply by Feb 1. (Different from NYU Faculty Fellows, which are also great but less connected to my lab.) Link in 🧵
22112
Reposted by Greg Durrett
Nick Tomlin @nickatomlin.bsky.social · 23/10/2025
Two brief advertisements! TTIC is recruiting both tenure-track and research assistant professors: ttic.edu/faculty-hiri... NYU is recruiting faculty fellows: apply.interfolio.com/174686 Happy to chat with anyone considering either of these options
ttic.edu
TTIC Faculty Opportunities at TTIC
086
Reposted by Greg Durrett
Manya Wadhwa @manyawadhwa.bsky.social · 08/10/2025
Unfortunately I won't be at #COLM2025 this week, but please check out our work being presented by my collaborators/advisors! If you are interested in evals of open-ended tasks/creativity please reach out and we can schedule a chat! :)
041
Greg Durrett @gregdnlp.bsky.social · 07/10/2025
Find my students and collaborators at COLM this week! Tuesday morning: @juand-r.bsky.social and @ramyanamuduri.bsky.social 's papers (find them if you missed it!) Wednesday pm: @manyawadhwa.bsky.social 's EvalAgent Thursday am: @anirudhkhatry.bsky.social 's CRUST-Bench oral spotlight + poster
095
Reposted by Greg Durrett
Juan Diego Rodriguez @juand-r.bsky.social · 06/10/2025
Excited to present this at #COLM2025 tomorrow! (Tuesday, 11:00 AM poster session)
0103
Greg Durrett @gregdnlp.bsky.social · 25/09/2025
Check out this feature about AstroVisBench, our upcoming NeurIPS D&B paper about code workflows and visualization in the astronomy domain! Great testbed for the interaction of code + VLM reasoning models.
020
Reposted by Greg Durrett
Kanishka Misra @kanishka.bsky.social · 02/06/2025
News🗞️ I will return to UT Austin as an Assistant Professor of Linguistics this fall, and join its vibrant community of Computational Linguists, NLPers, and Cognitive Scientists!🤘 Excited to develop ideas about linguistic and conceptual generalization (recruitment details soon!)
Picture of the UT Tower taken by me on my first day at UT as a postdoc in 2023!
12667
Greg Durrett @gregdnlp.bsky.social · 02/06/2025
Great to work on this benchmark with astronomers in our NSF-Simons CosmicAI institute! What I like about it: (1) focus on data processing & visualization, a "bite-sized" AI4Sci task (not automating all of research) (2) eval with VLM-as-a-judge (possible with strong, modern VLMs)
060
Reposted by Greg Durrett
Kristian G. Andersen @kgandersen.bsky.social · 30/05/2025
The end of US leadership in science, technology, and innovation. All in one little table. A tremendous gift to China, courtesy of the GOP. nsf-gov-resources.nsf.gov/files/00-NSF...
371050418
Reposted by Greg Durrett
David Hall @dlwh.bsky.social · 19/05/2025
Super excited Marin is finally out! Come see what we've been building! Code/platform for training fully reproducible models end-to-end, from data to evals. Plus a new high quality 8B base model. Percy did a good job explaining it on the other place. marin.community x.com/percyliang/s...
x.com
Percy Liang on X: "What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision: https://t.co/racsvmhyA3" / X
What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision: https://t.co/racsvmhyA3
1196
Greg Durrett @gregdnlp.bsky.social · 23/04/2025
Check out Anirudh's work on a new benchmark for C-to-Rust transpilation! 100 realistic-scale C projects, plus target Rust interfaces + Rust tests that let us validate the transpiled code beyond what prior benchmarks allow.
051
Reposted by Greg Durrett
Anirudh Khatry @anirudhkhatry.bsky.social · 23/04/2025
🚀Meet CRUST-Bench, a dataset for C-to-Rust transpilation for full codebases 🛠️ A dataset of 100 real-world C repositories across various domains, each paired with: 🦀 Handwritten safe Rust interfaces. 🧪 Rust test cases to validate correctness. 🧵[1/6]
1155
Greg Durrett @gregdnlp.bsky.social · 22/04/2025
Check out Manya's work on evaluation for open-ended tasks! The criteria from EvalAgent can be plugged into LLM-as-a-judge or used for refinement. Great tool with a ton of potential, and there's LOTS to do here for making LLMs better at writing!
032
Greg Durrett @gregdnlp.bsky.social · 21/04/2025
Check out Ramya et al.'s work on understanding discourse similarities in LLM-generated text! We see this as an important step in quantifying the "sameyness" of LLM text, which we think will be a step towards fixing it!
051
Reposted by Greg Durrett
Juan Diego Rodriguez @juand-r.bsky.social · 21/04/2025
Our final South by Semantics lecture at UT Austin is happening on Wednesday April 23!
South by Semantics Workshop
Title: "Not-your-mother's connectionism: LLMs as cognitive models"
Speaker: Ellie Pavlick (Brown University)
Date and time: April 23, 2025. 3:30 - 5 PM.
Location: GDC 6.302
2153
Greg Durrett @gregdnlp.bsky.social · 16/04/2025
Check out @juand-r.bsky.social and @wenxuand.bsky.social 's work on improving generator-validator gaps in LLMs! I really like the formulation of the G-V gap we present, and I was pleasantly surprised by how well the ranking-based training closed the gap. Looking forward to following up in this area!
0112
Reposted by Greg Durrett
Popehat @kenwhite.bsky.social · 26/03/2025
If you're scooping up students off the street for writing op-eds, you're secret police, and should be treated accordingly.
9590292291
Reposted by Greg Durrett
Juan Diego Rodriguez @juand-r.bsky.social · 11/03/2025
I'm excited to announce two papers of ours which will be presented this summer at @naaclmeeting.bsky.social eting.bsky.social and @iclr-conf.bsky.social ! 🧵
1103
Reposted by Greg Durrett
Swarat Chaudhuri @swarat.bsky.social · 22/02/2025
Excited about Proofwala, @amitayush.bsky.social's new framework for ML-aided theorem-proving. * Paper: arxiv.org/abs/2502.04671 * Code: github.com/trishullab/p... Proofwala allows the collection of proof-step data from multiple proof assistants (Coq and Lean) and multilingual training. (1/3)
1215
Reposted by Greg Durrett
Will Stancil @whstancil.bsky.social · 05/02/2025
Popular or not Dems cannot bend on the need for trans people to be treated with basic humanity and respect. If we give up that because the right made trans people unpopular, we give up everything. They’ll dice us group by group like a salami. We die on this hill or we die alone in a ditch
13667791312
Reposted by Greg Durrett
Carl T. Bergstrom @carlbergstrom.com · 28/01/2025
Here are just a few of the NSF review panels that were shut down today, Chuck. This is research that would have made us competitive in computer science that will now be delayed by many months if not lost forever. AI is fine but right now the top priority is keeping the lights on at NSF and NIH.
Proposal Review Panel for Civil, Mechanical, and Manufacturing Innovation (1194) 
Date/Time: January 28, 2025 8:30am – 5:00pm 
Contact: Daniel McAdams, 703/292-4654 

Proposal Review Panel for Innovation and Technology Ecosystems (84685) 
Date/Time: January 28, 2025 12:00pm – 6:00pm 
Contact: Alastair Monk, 703/292-8050 

Proposal Review Panel for Innovation and Technology Ecosystems (84685) 
Date/Time: January 28, 2025 10:00am – 5:00pm 
Contact: Henry Ahn, 703/292-8050 

Proposal Review Panel for Electrical, Communications, and Cyber Systems (1196) 
Date/Time: January 28 - 29, 2025 9:00am – 6:00pm 
Contact: Dominique Dagenais, 703/292-2980 

Proposal Review Panel for Engineering Education and Centers (173) 
Date/Time: January 28 - 29, 2025 8:30am – 5:00pm 
Contact: Jesus Soriano Molla, 703/292-7795 

Proposal Review Panel for Computer and Network Systems (1207) 
Date/Time: January 28 - 29, 2025 8:30am – 5:00pm 
Contact: Nan Zhang, 703/292-8950
10745195
Reposted by Greg Durrett
Kyle Lo @ ICML2026 🇰🇷 @kylelo.bsky.social · 03/01/2025
kicking off 2025 with our OLMo 2 tech report while payin homage to the sequelest of sequels 🫡 🚗 2 OLMo 2 Furious 🔥 is everythin we learned since OLMo 1, with deep dives into: 🚖 stable pretrain recipe 🚔 lr anneal 🤝 data curricula 🤝 soups 🚘 tulu post-train recipe 🚜 compute infra setup 👇🧵
26917
Greg Durrett @gregdnlp.bsky.social · 03/01/2025
Huge congrats to @prasannsinghal.bsky.social for being one of the 8 CRA Outstanding Undergraduate Researcher Award winners! It has been an absolute privilege to work with Prasann during his time at UT. (And he's applying for PhD programs this year...hint hint...) Prasann's work 🧵
1224
Reposted by Greg Durrett
Mor Geva @megamor2.bsky.social · 18/12/2024
What's in an attention head? 🤯 We present an efficient framework – MAPS – for inferring the functionality of attention heads in LLMs ✨directly from their parameters✨ A new preprint with Amit Elhelo 🧵 (1/10)
16013
Reposted by Greg Durrett
Thom Lake @thomlake.bsky.social · 11/12/2024
I'm at #Neurips2024 this week! My work (arxiv.org/abs/2406.17692) w/ @gregdnlp.bsky.social & @eunsol.bsky.social exploring the connection between LLM alignment and response pluralism will be at pluralistic-alignment.github.io Saturday. Drop by to learn more!
0286
Reposted by Greg Durrett
Atula Tejaswi @atutej.bsky.social · 09/12/2024
Missed out on #Swift tickets? No worries—swing by our #SVFT poster at #NeurIPS2024 and catch *real* headliners! 🎤💃🕺 📌Where: East Exhibit Hall A-C #2207, Poster Session 4 East ⏲️When: Thu 12 Dec, 4:30 PM - 7:30 PM PST #AI #MachineLearning #PEFT #NeurIPS24
192
Reposted by Greg Durrett
Swarat Chaudhuri @swarat.bsky.social · 08/12/2024
The legendary Putnam math competition had its 85th edition yesterday. Coincidentally, George Tsoukalas will present our paper on PutnamBench, a next-generation #AI4Math benchmark, at #NeurIPS2024 this week: arxiv.org/abs/2407.11214. If you work on frontier AI for math/reasoning, talk to George!
0153
Greg Durrett @gregdnlp.bsky.social · 08/12/2024
I'll be at #NeurIPS2024 w/ - @fcyin.bsky.social's LoFiT: using interp to improve fine-tuning (Weds pm poster & MINT spotlight talk Sun) - @thomlake.bsky.social's analysis of Overton pluralism (Pluralistic alignment Sat) Please reach out to me to chat about interp, factuality, reasoning, &c!
1468
Reposted by Greg Durrett
Akari Asai @akariasai.bsky.social · 04/12/2024
I’m on the academic job market this year! I’m completing my @uwcse.bsky.social @uwnlp.bsky.social Ph.D. (2025), focusing on overcoming LLM limitations like hallucinations, by building new LMs. My Ph.D. work focuses on Retrieval-Augmented LMs to create more reliable AI systems 🧵
37117
Greg Durrett @gregdnlp.bsky.social · 22/11/2024
when a famous person just got here and you temporarily have more followers than they do
"I'm the captain now" meme with top text "Look at me" and bottom text "I'm the influencer now"
41093
Reposted by Greg Durrett
Lucy Li @lucy3.bsky.social · 19/11/2024
mech interp: bsky.app/starter-pack... women in nlp: bsky.app/starter-pack... nlp #1: bsky.app/starter-pack... nlp #2: bsky.app/starter-pack... ml/data/tech: bsky.app/starter-pack... robotics & ai: bsky.app/starter-pack...
77319
Reposted by Greg Durrett
Naomi Saphra @nsaphra.bsky.social · 08/11/2024
Taking a stand that we aren’t doing the #nlproc tag here. It’s #nlp. We used #nlproc because a decade ago the #nlp tag was full of sleazy scammers selling guides for hypnotizing women into sleeping with you. But guess what? We won. All the sleazy scammers are doing natural language processing now.
1118129
Reposted by Greg Durrett
Jessy Li @jessyjli.bsky.social · 15/11/2024
We got an 🥂 Outstanding Paper Award!! Cannot be more grateful 🥹 This is super validating for our long pursuit of computational work on QUD. Congrats to the amazing @yatingwu.bsky.social, Ritika Mangla, Alex Dimakis, @gregdnlp.bsky.social
1589
Greg Durrett @gregdnlp.bsky.social · 13/11/2024
I won't be at EMNLP, but come and see: 🔍 Detecting factual errors from LLMs (Liyan Tang) 🛠️ Detect, critique, & refine pipeline (Manya Wadhwa and Lucy Zhao) 🏭 Synthetic data generation (Abhishek Divekar) 📄 Fact-checking (Aniruddh Sriram) at FEVER t.co/fQbl0G7m23 (1st real post in the bluer skies!)
Image of the linked website listing EMNLP paper titles, authors, and locations
1204
Reposted by Greg Durrett
Jessy Li @jessyjli.bsky.social · 24/10/2023
📢Our EMNLP 2023 work on Questions Under Discussion (QUD)! We introduce QUDeval, the first benchmark for evaluating the generation of open-ended questions and QUD parsing using linguistic principles. Paper: arxiv.org/abs/2310.14520 w/ @yatingwu.bsky.social, Ritika Mangla, @gregdnlp.bsky.social
0155