Sign in

Pasquale Minervini

@neuralnoise.com
5.6K followers 4.8K following 182 posts

Researcher in ML/NLP at the University of Edinburgh (faculty at Informatics and EdinburghNLP), Co-Founder/CTO at www.miniml.ai, ELLIS (@ELLIS.eu) Scholar, Generative AI Lab (GAIL, gail.ed.ac.uk) Fellow -- www.neuralnoise.com, he/they

PostsRepliesMedia
Reposted by Pasquale Minervini
EACL 2027 @eaclmeeting.bsky.social · 17/08/2026
The CFP for the Student Research Workshop at #EACL2027 is up! 🔗 2027.eacl.org/calls/srw/ ⏰ Pre-submission mentorship deadline: Nov 6, 2026 ⏰ Direct paper submission deadline: Dec 15, 2026 We accept both PhD thesis proposals and regular #nlproc research papers, looking forward to your submissions!
2027.eacl.org
Call for Student Research Workshop Papers
Official website for the 2027 Conference of the European Chapter of the Association for Computational Linguistics
075
Pasquale Minervini @neuralnoise.com · 02/08/2026
stochastic parrots are getting pretty lucky
050
Reposted by Pasquale Minervini
Ethan Mollick @emollick.bsky.social · 19/07/2026
I still believe that everyone is too fixated on the state of play in AI right now (which labs are ahead, how to manage costs, etc.) and not focused enough on the continued steepness of the capability curve for AI At higher capabilities (like the ones expected in the near term), a lot changes fast.
1014812
Pasquale Minervini @neuralnoise.com · 14/07/2026
Organisations -- sponsorship opportunities are available for #EACL2027, the flagship European conference in computational linguistics, taking place in Athens in March 2027! Support the NLP community and connect with researchers and practitioners: 2027.eacl.org
2027.eacl.org
The 20th Conference of the European Chapter of the Association for Computational LinguisticsAthens, GreeceMarch 9-14, 2027
8120
Reposted by Pasquale Minervini
Zeerak Talat زیرک طلعت (they/them) @zeerak.bsky.social · 08/07/2026
Thrilled to have been awarded the Association for Computational Linguistics 2016 test of time award for my first ever paper, written with/under the guidance of @dirkhovy.bsky.social A couple of cute things about the paper/its genesis/its outcomes
Photo of a test of time award for the paper "Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter"
1612418
Pasquale Minervini @neuralnoise.com · 04/07/2026
New blog post on our {ICML, ACL} 2026 papers, plus some new interesting results on open-ended learning and multi-modal retrieval! neuralnoise.com/2026/icml-ac...
7170
Reposted by Pasquale Minervini
Zeerak Talat زیرک طلعت (they/them) @zeerak.bsky.social · 18/06/2026
Call for papers for special issue on Ethics in NLP in the journal Computational Linguistics. Discussions around ethics in #NLP and #ComputationalLinguistics is often limited by the expectations for what should be published in *CL conferences. Link to call: www.aclweb.org/portal/conte...
1109
Reposted by Pasquale Minervini
Zeerak Talat زیرک طلعت (they/them) @zeerak.bsky.social · 12/06/2026
NLP reviews suck! But why is not clear– @aclrollingreview.bsky.social provides *a lot* of guidance for how to review, but in that, first principles get lost. So @adamlopez.bsky.social and I have written down some of our thoughts on first principles.
medium.com
The Missing First Principles of Reviewing for ACL
Zeerak Talat & Adam Lopez, University of Edinburgh
1236
Reposted by Pasquale Minervini
EACL 2027 @eaclmeeting.bsky.social · 04/05/2026
Attention #NLProc researchers, the EACL 2027 website is officially LIVE: 2027.eacl.org! 🎉 🇬🇷 Join us in Athens, Greece (Mar 9-13, 2027) at #EACL2027 📅 ARR submission deadline: Aug 6, 2026. Open to all areas of CL/NLP + related fields. Stay tuned for the detailed CfP soon!
Picture of the Acropolis in Athens, Greece
23922
Pasquale Minervini @neuralnoise.com · 27/03/2026
If you are interested in privacy-preserving clinical NLP, we are recruiting a postdoc at the University of Edinburgh! The work is on LLMs/VLMs, AI privacy, and real-world health data in secure research environments. Apply by April 6th, 2026! More details: elxw.fa.em3.oraclecloud.com/hcmUI/Candid...
2102
Pasquale Minervini @neuralnoise.com · 17/03/2026
My amazing colleagues Sid and Michael are looking for a postdoc! 👇
130
Reposted by Pasquale Minervini
an-exlab.bsky.social @an-exlab.bsky.social · 17/03/2026
We are advertising a postdoc position to work on #generative #models, #structure #induction, and MI #estimation with Michael Gutmann as part of @genaihub.bsky.social ! elxw.fa.em3.oraclecloud.com/hcmUI/Candid... Get in touch! (#ML #AI) 👉 homepages.inf.ed.ac.uk/snaraya3/ 👉 michaelgutmann.github.io
homepages.inf.ed.ac.uk
Siddharth - Home
Sid's page
043
Pasquale Minervini @neuralnoise.com · 29/11/2025
Chatted with the amazing @elissawelle.bsky.social from @theverge.com about @rohit-saxena.bsky.social’s “Lost in Time” work (arxiv.org/abs/2502.05092) and much more! You can find the full article here 👇
010
Pasquale Minervini @neuralnoise.com · 17/10/2025
Check out Yu Zhao's (@yuzhaouoe.bsky.social) latest work, “Learning GUI Grounding with Spatial Reasoning from Visual Feedback” (www.arxiv.org/abs/2509.21552), done during his internship at MSR (@msftresearch.bsky.social)! New SOTA 🏆 results on ScreenSpot-v2 (+5.7%) and ScreenSpot-Pro (+110.8%)!
131
Reposted by Pasquale Minervini
mr. TIM @timkellogg.me · 24/08/2025
trend: non-NVIDIA training DeepSeek V3.1 was trained on Huawei Ascend NPUs this one is a South Korean lab training on AMD
2397
Pasquale Minervini @neuralnoise.com · 23/08/2025
I really needed a Deep Research MCP server to use with Claude Code and other tools — here it is: github.com/pminervini/d...
060
Reposted by Pasquale Minervini
Mark Riedl @markriedl.bsky.social · 05/08/2025
@togelius.bsky.social has thoughts on Genie 3 and games togelius.blogspot.com/2025/08/geni... Fairly close to my own, though I didn't get the preview the tech. Walking around a generated image-to-image world is not the same as playing a game. There are no game objectives.
togelius.blogspot.com
Genie 3 and the future of neural game engines
Google DeepMind just announced Genie 3 , their new promptable world model, which is another term for neural game engine. This is a big neura...
4153
Reposted by Pasquale Minervini
mr. TIM @timkellogg.me · 03/08/2025
quick diagram of Bluesky’s architecture and why it’s nicer here
diagram from Anthropic paper with an icon & label that says “subtract evil vector”
4725
Reposted by Pasquale Minervini
Scott McGrath @smcgrath.phd · 23/07/2025
Anthropic research identifies “inverse scaling in test-time compute,” where longer reasoning degrades AI performance. On certain tasks, models become more distracted by irrelevant data or overfit to spurious correlations. #MLSky
venturebeat.com
Anthropic researchers discover the weird AI problem: Why thinking longer makes models dumber
Anthropic research reveals AI models perform worse with extended reasoning time, challenging industry assumptions about test-time compute scaling in enterprise deployments.
191
Pasquale Minervini @neuralnoise.com · 31/07/2025
Supermassive congrats to Giwon Hong (@giwonhong.bsky.social) for the amazing feat! 🙂
141
Pasquale Minervini @neuralnoise.com · 26/07/2025
The amazing folks at EdinburghNLP will be presenting a few papers at ACL 2025 (@aclmeeting.bsky.social); if you're in Vienna, touch base with them!
0120
Reposted by Pasquale Minervini
Emile van Krieken @emilevankrieken.com · 24/07/2025
Hm, hard disagree here. I really fail to see how this is misconduct akin to bribery, it's just a defense mechanism against bad reviewing practices. @neuralnoise.com
152
Reposted by Pasquale Minervini
Sohee Yang @soheeyang.bsky.social · 13/06/2025
🚨 New Paper 🚨 How effectively do reasoning models reevaluate their thought? We find that: - Models excel at identifying unhelpful thoughts but struggle to recover from them - Smaller models can be more robust - Self-reevaluation ability is far from true meta-cognitive awareness 1/N 🧵
1123
Reposted by Pasquale Minervini
mr. TIM @timkellogg.me · 22/07/2025
Inverse scaling of reasoning models a research collab demonstrated that there are certain types of tasks where all top reasoning models do WORSE the longer they think things like getting distracted by irrelevant info, spurious correlations, etc. www.arxiv.org/abs/2507.14417
Three panels at the top describe task types with example prompts:
	1.	Simple Counting Tasks with Distractors (Misleading Math & Python):
	•	Prompts mention an apple and an orange, with added irrelevant or confusing information (e.g., probabilistic riddle, Python code) before asking the straightforward question: “Calculate how many fruits you have.”
	2.	Regression Tasks with Spurious Features (Grades Regression):
	•	Given XML-style records about a student, the model must predict grades from features like sleep hours, social hours, and stress level. The challenge lies in identifying relevant vs. spurious attributes.
	3.	Deduction Tasks with Constraint Tracking (Zebra Puzzles):
	•	Complex logical reasoning puzzle with multiple interrelated clues. Example: “What position is the person who likes salmon at?” Constraints involve foods, names, and relationships like “to the left of.”

Bottom row contains 3 line plots comparing model performance across tasks:
	•	Misleading Math (Left Plot):
	•	Accuracy drops sharply for some models as reasoning tokens increase. Claude Sonnet 4 maintains high performance. o3 and DeepSeek R1 hold relatively stable accuracy; Qwen3 32B and QwQ 32B drop more.
	•	Grades Regression (Middle Plot):
	•	Shows negative RMSE (higher is better). Claude models remain strong across token counts; o3 also performs well. Qwen3 and QwQ struggle, with DeepSeek R1 performing modestly.
	•	Zebra Puzzles (Right Plot):
	•	Accuracy vs. average reasoning tokens. o3 and Claude Sonnet 4 maintain highest performance. Other models (e.g., DeepSeek R1, Qwen3 32B, QwQ 32B) show performance degradation or plateaus. Error bars reflect variability.

Each plot uses colored lines with markers to indicate different model names.
2212
Reposted by Pasquale Minervini
Naomi Saphra @nsaphra.bsky.social · 12/06/2025
Reasoning is about variable binding. It’s not about information retrieval. If a model cannot do variable binding, it is not good at grounded reasoning, and there’s evidence accruing that large scale can make LLMs worse at in-context grounded reasoning. 🧵
4549
Pasquale Minervini @neuralnoise.com · 22/07/2025
Hi @ilsebyl.bsky.social welcome to bsky! 🚀🚀🚀
120
Pasquale Minervini @neuralnoise.com · 22/07/2025
Sometimes, too much reasoning can hurt model performance! New research by Anthropic (@anthropic.com), by Aryo Pradipta Gema (@aryopg.bsky.social) et al.: huggingface.co/papers/2507....
huggingface.co
Paper page - Inverse Scaling in Test-Time Compute
Join the discussion on this paper page
050
Pasquale Minervini @neuralnoise.com · 21/07/2025
“LLMs can’t reason” 😅
050
Reposted by Pasquale Minervini
Steven Strogatz @stevenstrogatz.com · 04/07/2025
My "Math, Revealed" series is freely available to anyone -- no paywall! -- in the thread below.
613653
Reposted by Pasquale Minervini
Cosimo Gregucci @cgregucci.bsky.social · 10/07/2025
Spotlight poster coming soon at #ICML2025 @icmlconf.bsky.social! 📌East Exhibition Hall A-B E-1806 🗓️Wed 16 Jul 4:30 p.m. PDT — 7 p.m. PDT 📜 arxiv.org/pdf/2410.12537 Let’s chat! I’m always up for conversations about knowledge graphs, reasoning, neuro-symbolic AI, and benchmarking.
1102
Reposted by Pasquale Minervini
Melanie Mitchell @melaniemitchell.bsky.social · 05/07/2025
This essay by Nisheeth Vishnoi is a thoughtful meditation on the nature of science and a rebuttal to the notion that AI systems are going replace human scientists anytime soon. Worth reading. nisheethvishnoi.substack.com/p/what-count...
nisheethvishnoi.substack.com
What Counts as Discovery?
Rethinking AI’s Place in Science
47312
Pasquale Minervini @neuralnoise.com · 05/07/2025
"in 2025 we will have flying cars" 😂😂😂
839991
Reposted by Pasquale Minervini
Bálint Gyevnár @gbalint.bsky.social · 30/05/2025
Preprint alert 🎉 Introducing the Agentic eXplanations via Interrogative Simulations (AXIS) algo. AXIS integrates multi-agent simulators with LLMs by having the LLMs interrogate the simulator with counterfactual queries over multiple rounds for explaining agent behaviour. arxiv.org/pdf/2505.17801
Flowchart of the AXIS algorithm with 5 parts. The top-left has the memory, the centre-left has the user query, the centre-bottom has the final explanation, the centre has the LLM, and the right has the multi-agent simulator.Screenshot of the arXiv paper
081
Reposted by Pasquale Minervini
Bálint Gyevnár @gbalint.bsky.social · 17/04/2025
'AI Safety for Everyone' is out now in @natmachintell.nature.com! Through an analysis of 383 papers, we find a rich landscape of methods that cover a much larger domain than mainstream notions of AI safety. Our takeaway: Epistemic inclusivity is important, the knowledge is there, we only need use it
1133
Reposted by Pasquale Minervini
eleutherai.bsky.social @eleutherai.bsky.social · 06/06/2025
Can you train a performant language model using only openly licensed text? We are thrilled to announce the Common Pile v0.1, an 8TB dataset of openly licensed and public domain text. We train 7B models for 1T and 2T tokens and match the performance similar models like LLaMA 1 & 2
214660
Pasquale Minervini @neuralnoise.com · 04/06/2025
COLM (@colmweb.org‬) reviewers, please follow up on author responses if you need to! Most of the papers in my area chair batch didn't receive reviewer follow-ups, and it's dire
062
Pasquale Minervini @neuralnoise.com · 27/05/2025
Hi @veredshwartz.bsky.social !!! 🙂
150
Pasquale Minervini @neuralnoise.com · 25/05/2025
claude-code is pretty good at updating personal websites! it has browser use, so it can e.g. scrape your latest papers from arxiv and dblp and use that to update your website's publication list
250
Pasquale Minervini @neuralnoise.com · 25/05/2025
“You must never be fearful about what you are doing when it is right.” -- Rosa Parks
050
Reposted by Pasquale Minervini
antonio vergari ⚔️ short-circuiting @nolovedeeplearning.bsky.social · 24/05/2025
🗓️ Deadline extended: 💥2nd June 2025!💥 We are looking forward to your works on: 🔌 #circuits and #tensor #networks 🕸️ ⏳ normalizing #flows 💨 ⚖️ scaling #NeSy #AI 🦕 🚅 fast and #reliable inference 🔍 ...& more! please share 🙏
01912
Reposted by Pasquale Minervini
Emile van Krieken @emilevankrieken.com · 22/05/2025
This is a new experience: people using AI to overhype your paper 🫢 @neuralnoise.com
191
Reposted by Pasquale Minervini
Emile van Krieken @emilevankrieken.com · 21/05/2025
We propose Neurosymbolic Diffusion Models! We find diffusion is especially compelling for neurosymbolic approaches, combining powerful multimodal understanding with symbolic reasoning 🚀 Read more 👇
49327
Reposted by Pasquale Minervini
Ellen Rhudy @ellenrhudy.bsky.social · 20/05/2025
Just realized that the AI slop also made it into the @inquirer.com. So glad I saved this list of novels that do not exist!!!
A summer reading list from the Philadelphia Inquirer, which consists almost entirely of nonexistent novels
2176
Reposted by Pasquale Minervini
Emile van Krieken @emilevankrieken.com · 17/05/2025
Congrats! Looks like time is a big failure case for these models (cc @neuralnoise.com @aryopg.bsky.social @rohit-saxena.bsky.social ) bsky.app/profile/emil...
132
Reposted by Pasquale Minervini
Aryo Pradipta Gema @aryopg.bsky.social · 02/05/2025
MMLU-Redux just touched down at #NAACL2025! 🎉 Wish I could be there for our "Are We Done with MMLU?" poster today (9:00-10:30am in Hall 3, Poster Session 7), but visa drama said nope 😅 If anyone's swinging by, give our research some love! Hit me up if you check it out! 👋
MMLU-Redux Poster at NAACL 2025
01711
Reposted by Pasquale Minervini
Ai2 @ai2.bsky.social · 01/05/2025
We're excited to round out the OLMo 2 family with its smallest member, OLMo 2 1B, surpassing peer models like Gemma 3 1B or Llama 3.2 1B. The 1B model should enable rapid iteration for researchers, more local development, and a more complete picture of how our recipe scales.
A bar graph comparing average performance (10 Tasks) across OLMo 2 1B, SmolLM2 1.7B, Gemma 3 1B, Llama 3.2 1B, and Qwen 2.5 1.5B. The highest performance is 42.7, achieved by OLMo 2 1B.
14210
Reposted by Pasquale Minervini
Yuchen Zhu @zhuyuchen.bsky.social · 01/05/2025
New work! 💪🏻💥🤯 When Can Proxies Improve the Sample Complexity of Preference Learning? Our paper is accepted at @icmlconf.bsky.social 2025. Fantastic joint work with @spectral.space, Zhengyan Shi, @meng-yue-yang.bsky.social, @neuralnoise.com, Matt Kusner, @alexdamour.bsky.social. 1/n
184
Pasquale Minervini @neuralnoise.com · 30/04/2025
Extremely proud of the squad 💕
030
Reposted by Pasquale Minervini
Manuel Dileo @manueldileo.bsky.social · 25/04/2025
Happy to share that our work won the best paper award at ESANN 2025! 🚀 Huge thanks to my amazing co-authors✨ Read our work here: doi.org/10.14428/esa...
031