Sign in

Joe Stacey

@joestacey.bsky.social
2.6K followers 2.1K following 144 posts

NLP PhD student at Imperial College London and Apple AI/ML Scholar.

PostsRepliesMedia
Reposted by Joe Stacey
Lisa Alazraki @lisaalaz.bsky.social · 28/08/2025
We have released #AgentCoMa, an agentic reasoning benchmark where each task requires a mix of commonsense and math to be solved 🧐 LLM agents performing real-world tasks should be able to combine these different types of reasoning, but are they fit for the job? 🤔 🧵⬇️
172
Joe Stacey @joestacey.bsky.social · 17/07/2025
Here’s my review of the US after a few days here. Did I miss anything? 🤔 The good: - Americans are the most charming, friendly and hospitable people - it’s super fun how the country is split into states that all have different laws and stuff, with different vibes state to state
110
Joe Stacey @joestacey.bsky.social · 02/07/2025
Any chance Keir Starmer can reshuffle himself in as foreign secretary, and shuffle in another prime minister who actually has some vague idea about what they want to achieve? 🙏🤦‍♂️
000
Joe Stacey @joestacey.bsky.social · 02/07/2025
Finally the heatwave has ended, and the UK is once again a bearable place to be 😍😍 If you have any UK-based collaborations, their productivity is about to increase like 10 fold
020
Joe Stacey @joestacey.bsky.social · 27/05/2025
We have a fun new #NLProc paper on arXiv about improving the robustness of fine-tuned NLI models! Have a look :) arxiv.org/abs/2505.20209
160
Joe Stacey @joestacey.bsky.social · 18/05/2025
Should I use an LLM to help refine my paper writing for the ARR deadline? 🤔🤔 It will improve the paper for sure, but probably also making the tone a whole lot more annoying
100
Reposted by Joe Stacey
Juan Diego Rodriguez @juand-r.bsky.social · 28/04/2025
If you're at #NAACL2025 and want to hear about similarity effects for property inheritance in LMs, please stop by! I will be presenting this work on Wednesday at the 11-12:30 poster session on Interpretability & analysis for language models (Hall 3). aclanthology.org/2025.naacl-l...
aclanthology.org
Characterizing the Role of Similarity in the Property Inferences of Language Models
Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo...
0113
Reposted by Joe Stacey
Imperial NLP @imperial-nlp.bsky.social · 22/04/2025
Excited to share our ICLR and NAACL papers! Please come and say hi, we're super friendly :)
0155
Joe Stacey @joestacey.bsky.social · 05/04/2025
Wow, the old ITV Agatha Christie’s Poirot is brilliant. Some tv for 1989… Gonna go binge watch the 13 seasons now 😍
010
Joe Stacey @joestacey.bsky.social · 04/04/2025
I feel like the length of the ARR author rebuttals keep growing every cycle Is this a good thing for authors or reviewers that the responses can be so long? I feel like it’s a bit sub-optimal for both at the moment
340
Reposted by Joe Stacey
Nishant Balepur @nbalepur.bsky.social · 21/03/2025
Had a great time presenting my research on building more helpful QA systems @imperialcollegeldn.bsky.social! Thank you @joestacey.bsky.social for letting me invite myself 🫶 And loved visiting London+Edinburgh this week, hope to be back soon! 🙏
051
Joe Stacey @joestacey.bsky.social · 21/03/2025
Was fantastic to have you here at Imperial! Thanks for your excellent talk, and looking forward to following what you do next 🙂
040
Reposted by Joe Stacey
Lisa Alazraki @lisaalaz.bsky.social · 13/02/2025
Do LLMs need rationales for learning from mistakes? 🤔 When LLMs learn from previous incorrect answers, they typically observe corrective feedback in the form of rationales explaining each mistake. In our new preprint, we find these rationales do not help, in fact they hurt performance! 🧵
1219
Reposted by Joe Stacey
Marek Rei @marekrei.bsky.social · 04/03/2025
Today was the launch event of the @genaihub.bsky.social. We announced the development of Nightingale AI, a foundation world model for health. It was great to be on the panel for GenAI in Healthcare, among such amazing experts. www.genai.ac.uk
152
Joe Stacey @joestacey.bsky.social · 28/02/2025
Thanks so much to everyone who has helped make this switch to BlueSky work. Honestly, making this switch was a pretty massive achievement, so thanks everyone for contributing ❤️❤️
2120
Joe Stacey @joestacey.bsky.social · 18/02/2025
This paper is really cool. They decompose NLI (and defeasible NLI) hypotheses into atoms, and then use these atoms to measure the logical consistency of LLMs. E.g. for an entailment NLI example, each hypothesis atom should also be entailed by the premise. Very nice idea 👏👏
2153
Joe Stacey @joestacey.bsky.social · 09/02/2025
I’m a week into my trip from Cairo to Riyadh, and wow what a place Egypt is! Honestly its been one of the funnest places I’ve travelled, and for sure I need to come back again Crossed into Aqaba (Jordan) yesterday, so now onto Saudi 🙂
040
Joe Stacey @joestacey.bsky.social · 29/01/2025
I’m going away to do a bit of travelling, going overland from Cairo to Riyadh 😍 I love travelling in the Middle East so it should be interesting I’ve got that feeling of nervous excitement I always get before a trip 😬😁
050
Joe Stacey @joestacey.bsky.social · 23/01/2025
Insanely jealous to everyone who has papers at #NAACL in Albuquerque! Albuquerque just sounds so exotic, and is such a cool place for a conference. No offence to Vienna, but Albuquerque sounds way more fun 😉
100
Joe Stacey @joestacey.bsky.social · 17/01/2025
Feeling gooooood after submitting my #ARR reviews early 😍 Time to enjoy the weekend! 🕺
020
Joe Stacey @joestacey.bsky.social · 14/01/2025
I was super excited to read the ModernBERT paper! Love this interest in creating a better encoder model. "ModernBERT-base is the first encoder to beat DeBERTaV3-base since its release in 2021" 🤯- arxiv.org/pdf/2412.13663 Pretty amazing how successful DeBERTa has been!
0160
Joe Stacey @joestacey.bsky.social · 07/01/2025
Excited to start my #ARR #NLP reviews! I'll try my best and see if I can get 100% of my reviews to be 'great' this round. If you didn't see it already, ARR publishes how many of your reviews are considered to be 'great': stats.aclrollingreview.org Join me for the challenge :)
stats.aclrollingreview.org
ARR Dashboard
1122
Joe Stacey @joestacey.bsky.social · 02/01/2025
At some point in life I realised I actually really love travelling by train. Kind of a strange hobby, but wow it is fun 😍 Here are my top ten train journeys so far.
4352
Joe Stacey @joestacey.bsky.social · 28/12/2024
Imperial are hiring computing lecturers (including for AI/ML/NLP)! Here's a little thread about why you should consider applying :)
141
Joe Stacey @joestacey.bsky.social · 26/11/2024
Made it to northern Sweden (Kiruna) by train from London. Freezing cold with northern lights 😍 Just over a week ago and I was in the crazy Miami heat for #EMNLP2024
1240
Joe Stacey @joestacey.bsky.social · 24/11/2024
Okay genius idea to improve quality of #nlp #arr reviews. Literally give gold stars to the best reviewers, visible on open review next to your anonymously ID during review process. Here’s why it would work, and why would you should RT this fab idea:
3275
Joe Stacey @joestacey.bsky.social · 24/11/2024
This papers' findings about testing LLMs on NLI aligns with many of personal thoughts: 1) NLI remains a difficult task for LLMs 2) Having more few-shot examples is helpful (in my view, helping LLMs better understand class boundaries) 3) Incorrect predictions are often a result of ambiguous labels
1273
Joe Stacey @joestacey.bsky.social · 23/11/2024
I’ve seen some pretty amazing metros before (like Moscow), but wow Stockholm is wild. Never seen anything like it!
1130
Joe Stacey @joestacey.bsky.social · 22/11/2024
If anyone is getting annoyed with their BlueSky feed, try 'Popular with Friends' - you can add this from the 'Feeds' tab. I'm finding it works a bit better for me, and is more like what I had on Twitter. Thanks @lasha.bsky.social for suggesting!
120
Reposted by Joe Stacey
Pasquale Minervini @neuralnoise.com · 20/11/2024
Starter pack for University of Edinburgh researchers done by the amazing ramandutt4.bsky.social - go.bsky.app/KRNDkN7
go.bsky.app
University of Edinburgh Starter Pack
Join the conversation
9359
Joe Stacey @joestacey.bsky.social · 20/11/2024
Now I have like a gazillion new Bluesky followers, posting a link again to a blog post about my #EMNLP2024 and EMNLP 2022 papers. It’s a fun 10 minute read about our ideas on interpretable neural architectures. ❤️ to my fantastic collaborators www.marekrei.com/blog/creatin...
marekrei.com
Creating Interpretable Models with Atomic Inference - Marek Rei
This is a guest post from Joe Stacey about our quest to create interpretable Natural Language Inference (NLI) models. In this post he will share…
0312
Joe Stacey @joestacey.bsky.social · 20/11/2024
Welcome to Bluesky to more of our NLP researchers at Imperial!! Looking forward to following everyone's work on here. To follow us all click 'follow all' in the starter pack below go.bsky.app/Bv5thAb
3207
Joe Stacey @joestacey.bsky.social · 20/11/2024
Just about to start my next big train journey, this time from London to Norway (Narvik). Just 7 countries to get the train through (UK, France, Belgium, Germany, Denmark, Sweden, Norway) 😍 😍 should be epic
4260
Joe Stacey @joestacey.bsky.social · 18/11/2024
Cool thing about this thread now is you can see the likes per tip! Almost like voting on the best ones. So far tip #1 the clear winner Loving the BlueSky engagement 😍 😍
070
Joe Stacey @joestacey.bsky.social · 18/11/2024
You know it’s cold when little Hamish starts hugging the radiator ❤️
0120
Joe Stacey @joestacey.bsky.social · 18/11/2024
After going to NAACL, ACL and #EMNLP2024 this year, here are a few tips I’ve picked up about attending #NLP conferences. Would love to hear any other tips if you have them! This proved very popular on another (more evil) social media platform, so sharing here also 🙂 My 10 tips:
148316
Joe Stacey @joestacey.bsky.social · 17/11/2024
“Entering Georgia and South Carolina, last chance to buy alcohol!” Had no idea, but i think alcohol sales are prohibited on Sunday in these states! The train is so exciting 🙂
010
Joe Stacey @joestacey.bsky.social · 17/11/2024
I love the Amtrak dining cars!! How pretty is this. Really good sit down breakfast, lunch and dinners. And the best bit is all the amazing people you meet and speak to at the meals.
1171
Joe Stacey @joestacey.bsky.social · 17/11/2024
Just boarded my train from Miami to New York post #EMNLP2024 and super excited!! Amtrak trains are the fantastic, and I’ve got my own little room with two seats, a bed above, and toilet next to the bed. The toilet thing is a bit weird though if you have two to a room
3251
Joe Stacey @joestacey.bsky.social · 15/11/2024
Such a fantastic reaction to our paper today. so happy 🙂 Chocolates went down well too! Massive thanks to everyone for all your ideas and feedback
0110
Joe Stacey @joestacey.bsky.social · 14/11/2024
Excited to present our #EMNLP2024 paper as a poster this morning at 10:30 (in the downstairs poster room)! It's cool work about creating inherently interpretable models, and (as always) I will have chocolate to give out 😀 Paper is here: aclanthology.org/2024.emnlp-m...
050
Joe Stacey @joestacey.bsky.social · 10/11/2024
So excited to fly out to #EMNLP2024 tomorrow! Would love to chat sometime 🙂 I’m on the conference app so easy to message me there, or just come and say hi! Would love to hear about your research Hopefully I won’t be too jet lagged 😅✈️ #NLP
030
Reposted by Joe Stacey
Juan Diego Rodriguez @juand-r.bsky.social · 08/11/2024
How do language models organize concepts and their properties? Do they use taxonomies to infer new properties, or infer based on concept similarities? Apparently, both! 🌟 New paper with my fantastic collaborators @amuuueller.bsky.social and @kanishka.bsky.social
Title: "Characterizing the Role of Similarity in the Property Inferences of Language Models"
Authors: Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra

Left figure: "Given that dogs are daxable, is it true that corgis are daxable?" A language model could answer this either using taxonomic relations, illustrated by a taxonomy dog-corgi, dog-mutt, canine-wolf, etc., or by similarity relations (dogs are more similar to corgis than cats, wolves or shar peis).

Right figure: illustration of the causal model (and an example intervention) for distributed alignment search (DAS), which we used to find a subspace in the network responsible for property inheritance behavior. The bottom nodes are "property", "premise concept (A)" and "conclusion concept (B)" , the middle nodes are "A has property P", "B is a kind of A", and the top node is "B has property P".
410822
Joe Stacey @joestacey.bsky.social · 09/11/2024
#NLP Bluesky really growing quick. Going to need a lot of effort early on to replace Twitter, but a very promising start! 🙏
0171
Joe Stacey @joestacey.bsky.social · 09/11/2024
For the day I get back from #EMNLP2024 I’ve booked big train trip from London to Narvik in Norway. You’re probably wondering how that even works without getting on a boat. Well, it involves taking the train through quite a few countries 🙂🙂
010
Joe Stacey @joestacey.bsky.social · 08/11/2024
Any tips people have in advance of #EMNLP2024 for good poster presentations? It’s such a small thing, but I always like when people acknowledge you when you’re waiting for a poster (when the presenters busy talking to someone else) 🙂
350
Joe Stacey @joestacey.bsky.social · 08/11/2024
I'm new to BlueSky, but excited to be here! 😍 I've written up a little blog post about my EMNLP 2022 and #EMNLP2024 papers about interpretable neural architectures in #NLP. Great way to learn about our work with minimal paper reading :) Let me know what you think! www.marekrei.com/blog/creatin...
0312