Sign in

Joe Stacey

@joestacey.bsky.social
2.6K followers 2.1K following 144 posts

NLP PhD student at Imperial College London and Apple AI/ML Scholar.

PostsRepliesMedia
Joe Stacey @joestacey.bsky.social · 15/11/2025
This was just embarrassing. Shame on everyone who works on Grok…
030
Reposted by Joe Stacey
Lisa Alazraki @lisaalaz.bsky.social · 28/08/2025
We have released #AgentCoMa, an agentic reasoning benchmark where each task requires a mix of commonsense and math to be solved 🧐 LLM agents performing real-world tasks should be able to combine these different types of reasoning, but are they fit for the job? 🤔 🧵⬇️
172
Joe Stacey @joestacey.bsky.social · 22/07/2025
Congratulations!! Awesome you will be in Europe!
120
Joe Stacey @joestacey.bsky.social · 17/07/2025
The bad: - the chocolate here is terrible for no good reason - hotel breakfasts never have any baked beans, which are way under appreciated here (they are delicious and add much needed moisture to a cooked breakfast) - the temperature in summer is inhumane Think that covers the main stuff 😍
010
Joe Stacey @joestacey.bsky.social · 17/07/2025
Here’s my review of the US after a few days here. Did I miss anything? 🤔 The good: - Americans are the most charming, friendly and hospitable people - it’s super fun how the country is split into states that all have different laws and stuff, with different vibes state to state
110
Joe Stacey @joestacey.bsky.social · 02/07/2025
Any chance Keir Starmer can reshuffle himself in as foreign secretary, and shuffle in another prime minister who actually has some vague idea about what they want to achieve? 🙏🤦‍♂️
000
Joe Stacey @joestacey.bsky.social · 02/07/2025
Finally the heatwave has ended, and the UK is once again a bearable place to be 😍😍 If you have any UK-based collaborations, their productivity is about to increase like 10 fold
020
Joe Stacey @joestacey.bsky.social · 27/05/2025
This work was really fun and a great last paper for my PhD. Check it out 🙂 Massive thanks to all my amazing collaborators! arxiv.org/abs/2505.20209 P.S. if you know about a paper improving NLI model robustness not already in our related work appendix, I would love to hear about it 🥰
arxiv.org
How to Improve the Robustness of Closed-Source Models on NLI
Closed-source Large Language Models (LLMs) have become increasingly popular, with impressive performance across a wide range of natural language tasks. These models can be fine-tuned to further improv...
000
Joe Stacey @joestacey.bsky.social · 27/05/2025
5) The best way to improve performance on the hardest OOD data was to choose more challenging training examples Our best method (Uncertainty Sampling) picked examples with the most uncertain predictions. This identified challenging examples, but without too much label noise
110
Joe Stacey @joestacey.bsky.social · 27/05/2025
4) Creating more complex synthetic data avoids a loss in performance on harder OOD datasets We find that generating more challenging synthetic data (Long & Complex Generation) helps retain performance on harder OOD datasets, while still achieving gains on easier OOD data
100
Joe Stacey @joestacey.bsky.social · 27/05/2025
3) Replacing some training examples with LLM-generated data proved very effective on less challenging OOD data See Standard-OOD scores below (avg), where the simplest LLM-generated data (Short & Simple Generation) performed best, with substantial improvements
100
Joe Stacey @joestacey.bsky.social · 27/05/2025
2) We experiment with 6+ ways for improving robustness: This involved sampling methods to choose more complex examples in our training data, and generating new synthetic examples Some methods were pretty fun, e.g. asking an LLM to assess the difficulty of training examples
110
Joe Stacey @joestacey.bsky.social · 27/05/2025
1) It's time to stop using fine-tuned encoder models: We find that fine-tuned LLMs are substantially more robust than commonly used encoder models, despite being fine-tuned on x50 less data. This is especially the case on challenging OOD datasets (see Challenge-OOD avg below)
100
Joe Stacey @joestacey.bsky.social · 27/05/2025
The paper tries to improve the robustness of closed-source LLMs fine-tuned on NLI, assuming a realistic training budget of 10k training examples. Here's a 45 second rundown of what we found!
100
Joe Stacey @joestacey.bsky.social · 27/05/2025
We have a fun new #NLProc paper on arXiv about improving the robustness of fine-tuned NLI models! Have a look :) arxiv.org/abs/2505.20209
160
Joe Stacey @joestacey.bsky.social · 18/05/2025
I’d personally just love to see more negative results from nice ideas that didn’t quite work out. I feel like there’s probably a bunch of cool stuff people have tried out and discarded that could be made to work across multiple papers. Would be fun and interesting too
121
Joe Stacey @joestacey.bsky.social · 18/05/2025
Was worried it was just me hating on it so much 🤣
000
Joe Stacey @joestacey.bsky.social · 18/05/2025
I’d love to see more diversity in the field, what kind of things were you thinking?
100
Joe Stacey @joestacey.bsky.social · 18/05/2025
Should I use an LLM to help refine my paper writing for the ARR deadline? 🤔🤔 It will improve the paper for sure, but probably also making the tone a whole lot more annoying
100
Reposted by Joe Stacey
Juan Diego Rodriguez @juand-r.bsky.social · 28/04/2025
If you're at #NAACL2025 and want to hear about similarity effects for property inheritance in LMs, please stop by! I will be presenting this work on Wednesday at the 11-12:30 poster session on Interpretability & analysis for language models (Hall 3). aclanthology.org/2025.naacl-l...
aclanthology.org
Characterizing the Role of Similarity in the Property Inferences of Language Models
Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo...
0113
Joe Stacey @joestacey.bsky.social · 28/04/2025
Looks so cool! I’m insanely jealous
110
Joe Stacey @joestacey.bsky.social · 23/04/2025
I’m not a fan of musk, but imo there’s some really nice work here 🙂 Interested in the Washington post article, would you mind sharing a link?
010
Reposted by Joe Stacey
Imperial NLP @imperial-nlp.bsky.social · 22/04/2025
Excited to share our ICLR and NAACL papers! Please come and say hi, we're super friendly :)
0155
Joe Stacey @joestacey.bsky.social · 14/04/2025
That’s an awesome paper 👍👍
101
Joe Stacey @joestacey.bsky.social · 05/04/2025
Wow, the old ITV Agatha Christie’s Poirot is brilliant. Some tv for 1989… Gonna go binge watch the 13 seasons now 😍
010
Joe Stacey @joestacey.bsky.social · 05/04/2025
Congratulations! It’s definitely worth trying/experimenting with responses that are more concise in the future and see what kind of reaction you get. Best of luck with your meta-reviews! 🤞
010
Joe Stacey @joestacey.bsky.social · 04/04/2025
Ah that’s good to know! Yeah I think when authors choose to write concise responses everybody wins 🙂
010
Joe Stacey @joestacey.bsky.social · 04/04/2025
Good point. I think the other downside is all the reviewer time to go through them. I’m not sure what best solution is, and if you limit the responses too much it’s frustrating, but maybe something that stops way too long responses might be helpful 🙂
110
Joe Stacey @joestacey.bsky.social · 04/04/2025
I feel like the length of the ARR author rebuttals keep growing every cycle Is this a good thing for authors or reviewers that the responses can be so long? I feel like it’s a bit sub-optimal for both at the moment
340
Joe Stacey @joestacey.bsky.social · 25/03/2025
Not only does everyone learn for themselves, but I think almost everyone sees themselves as good reviewers when that may not be the case I think the ARR stats on how many great reviews people did is pretty cool step in the right direction!
000
Reposted by Joe Stacey
Nishant Balepur @nbalepur.bsky.social · 21/03/2025
Had a great time presenting my research on building more helpful QA systems @imperialcollegeldn.bsky.social! Thank you @joestacey.bsky.social for letting me invite myself 🫶 And loved visiting London+Edinburgh this week, hope to be back soon! 🙏
051
Joe Stacey @joestacey.bsky.social · 21/03/2025
Was fantastic to have you here at Imperial! Thanks for your excellent talk, and looking forward to following what you do next 🙂
040
Reposted by Joe Stacey
Lisa Alazraki @lisaalaz.bsky.social · 13/02/2025
Do LLMs need rationales for learning from mistakes? 🤔 When LLMs learn from previous incorrect answers, they typically observe corrective feedback in the form of rationales explaining each mistake. In our new preprint, we find these rationales do not help, in fact they hurt performance! 🧵
1219
Reposted by Joe Stacey
Marek Rei @marekrei.bsky.social · 04/03/2025
Today was the launch event of the @genaihub.bsky.social. We announced the development of Nightingale AI, a foundation world model for health. It was great to be on the panel for GenAI in Healthcare, among such amazing experts. www.genai.ac.uk
152
Joe Stacey @joestacey.bsky.social · 01/03/2025
Great to hear! 🙂
000
Joe Stacey @joestacey.bsky.social · 28/02/2025
Thanks so much to everyone who has helped make this switch to BlueSky work. Honestly, making this switch was a pretty massive achievement, so thanks everyone for contributing ❤️❤️
2120
Joe Stacey @joestacey.bsky.social · 18/02/2025
There are some other nice findings in the paper, e.g. using atomic inference to predict the label from the atom-level decisions, and seeing how well LLMs perform on "critical atoms" for defeasible NLI arxiv.org/pdf/2502.08080
arxiv.org
110
Joe Stacey @joestacey.bsky.social · 18/02/2025
This paper is really cool. They decompose NLI (and defeasible NLI) hypotheses into atoms, and then use these atoms to measure the logical consistency of LLMs. E.g. for an entailment NLI example, each hypothesis atom should also be entailed by the premise. Very nice idea 👏👏
2153
Joe Stacey @joestacey.bsky.social · 18/02/2025
Congratulations!! That's awesome!
110
Joe Stacey @joestacey.bsky.social · 09/02/2025
I’m a week into my trip from Cairo to Riyadh, and wow what a place Egypt is! Honestly its been one of the funnest places I’ve travelled, and for sure I need to come back again Crossed into Aqaba (Jordan) yesterday, so now onto Saudi 🙂
040
Joe Stacey @joestacey.bsky.social · 29/01/2025
I’m going away to do a bit of travelling, going overland from Cairo to Riyadh 😍 I love travelling in the Middle East so it should be interesting I’ve got that feeling of nervous excitement I always get before a trip 😬😁
050
Joe Stacey @joestacey.bsky.social · 23/01/2025
Best of luck with the thesis writing!! Sounds tough!
030
Joe Stacey @joestacey.bsky.social · 23/01/2025
Insanely jealous to everyone who has papers at #NAACL in Albuquerque! Albuquerque just sounds so exotic, and is such a cool place for a conference. No offence to Vienna, but Albuquerque sounds way more fun 😉
100
Joe Stacey @joestacey.bsky.social · 17/01/2025
Feeling gooooood after submitting my #ARR reviews early 😍 Time to enjoy the weekend! 🕺
020
Joe Stacey @joestacey.bsky.social · 14/01/2025
I was super excited to read the ModernBERT paper! Love this interest in creating a better encoder model. "ModernBERT-base is the first encoder to beat DeBERTaV3-base since its release in 2021" 🤯- arxiv.org/pdf/2412.13663 Pretty amazing how successful DeBERTa has been!
0160
Joe Stacey @joestacey.bsky.social · 07/01/2025
Maybe it depends a bit on the AC, some might be more generous for saying a review is great compared to others I'll just try my best this cycle and see what I get!
030
Joe Stacey @joestacey.bsky.social · 07/01/2025
oh wow, 7 or 8 is a huge amount!! I think the ranking is just the number of great reviews, but yeah maybe 7 or 8 reviews in a cycle should get you a statue somewhere :)
130
Joe Stacey @joestacey.bsky.social · 07/01/2025
You can see a table like this. Would be brilliant to know which reviews are 'great' or not though!
120
Joe Stacey @joestacey.bsky.social · 07/01/2025
It's so cool isn't it!
010