Sign in

Matthew Finlayson

@mattf.nl
4.7K followers 623 following 70 posts

NLP PhD @ USC Previously at AI2, Harvard mattf1n.github.io

PostsRepliesMedia
Reposted by Matthew Finlayson
Kyle Mahowald @kmahowald.bsky.social · 30/09/2026
AI agents say things like ARGH and OH MY GOD in their chains of thought. Good reason to think it's not just imitation but that these words play functional roles for doing reasoning. What comes after "oops" is very different than what comes after "aha"...
021
Matthew Finlayson @mattf.nl · 06/11/2025
Thank you for your kind words :) I'll take a look at TRAP, it looks very cool.
031
Matthew Finlayson @mattf.nl · 17/10/2025
As always, a big thank you to my stalwart advisors @swabhs.bsky.social and Xiang Ren
070
Matthew Finlayson @mattf.nl · 17/10/2025
Implications for AI accountability, forensics, and regulation. As LLMs become more powerful and opaque, having natural verification methods becomes crucial. Our proposed signature fills a new niche in this ecosystem. 7/
1121
Matthew Finlayson @mattf.nl · 17/10/2025
This opens the door to a verification system analogous to cryptographic message authentication—where the model ellipse functions as a secret key. Providers could verify outputs to trusted third parties without revealing model parameters. 6/
1161
Matthew Finlayson @mattf.nl · 17/10/2025
The forgery resistance comes from the excessive complexity of extracting an ellipse from an API: O(d³ log d) queries and O(d⁶) time to fit. For a 70B model, that's ~$16M in API costs and millennia of computation time 💸⏰5/
191
Matthew Finlayson @mattf.nl · 17/10/2025
We tested this on models like Llama 3.1, Qwen 3, and GPT-OSS. Even when we copied their linear signatures onto other models' outputs, the ellipse signature cleanly identified the true source by orders of magnitude. 4/
1121
Matthew Finlayson @mattf.nl · 17/10/2025
Why is this exciting? Four unique properties: 🔨 Forgery-resistant (computationally hard to fake) 🌱 Naturally occurring (no setup needed) 🫙 Self-contained (works without input/full weights) 🤏 Compact (detectable in a single generation step) 3/
190
Matthew Finlayson @mattf.nl · 17/10/2025
The key insight is that LLMs with normalization layers produce outputs that lie on the surface of a high-dimensional ellipse. This geometric constraint acts as a signature unique to each model. 2/
1213
Matthew Finlayson @mattf.nl · 17/10/2025
We discovered that language models leave a natural "signature" on their API outputs that's extremely hard to fake. Here's how it works 🔍 📄 arxiv.org/abs/2510.14086 1/
arxiv.org
Every Language Model Has a Forgery-Resistant Signature
The ubiquity of closed-weight language models with public-facing APIs has generated interest in forensic methods, both for extracting hidden model details (e.g., parameters) and for identifying...
48723
Matthew Finlayson @mattf.nl · 23/06/2025
The project was led by Murtaza Nazir, an independent researcher with serious engineering chops. It's his first paper. He's a joy to work with and is applying to PhDs. Hire him! It's great to finally collab with Jack Morris, and a big thanks to @swabhs.bsky.social and Xiang Ren for advising.
040
Matthew Finlayson @mattf.nl · 23/06/2025
Our technical insight is that logprob vectors can be linearly encoded as a much smaller vector. We make prompt stealing both *more accurate* and *cheaper*, by compactly encoding logprob outputs over multiple generation steps, resulting in massive gains over previous SoTA methods.
140
Matthew Finlayson @mattf.nl · 23/06/2025
We noticed that existing methods don't fully use LLM outputs: either they ignore logprobs (text only), or they only use logprobs from a single generation step. The problem is that next-token logprobs are big--the size of the entire LLM vocabulary *for each generation step*.
130
Matthew Finlayson @mattf.nl · 23/06/2025
When interacting with an AI model via an API, the API provider may secretly change your prompt or inject a system message before feeding it to the model. Prompt stealing--also known as LM inversion--tries to reverse engineer the prompt that produced a particular LM output.
110
Matthew Finlayson @mattf.nl · 23/06/2025
I didn't believe when I first saw, but: We trained a prompt stealing model that gets >3x SoTA accuracy. The secret is representing LLM outputs *correctly* 🚲 Demo/blog: mattf1n.github.io/pils 📄: arxiv.org/abs/2506.17090 🤖: huggingface.co/dill-lab/pi... 🧑‍💻: github.com/dill-lab/PILS
1110
Reposted by Matthew Finlayson
David Marx @digthatdata.bsky.social · 11/06/2025
I wish the ML community would stop trying to turn every technique into a brand name. Just give the thing a descriptive name and call it what it is. Forced backronyms like this are counter productive.
171
Matthew Finlayson @mattf.nl · 04/06/2025
It appears that the only fonts with optical sizes that work with pdflatex are the computer/latin modern fonts. I would kill for a free pdflatex-compatible Times clone with optical sizes so my small text can look good in ArXiv/conference submissions.
020
Matthew Finlayson @mattf.nl · 16/03/2025
If you are writing a paper for #colm2025 and LaTeX keeps increasing your line height to accommodate things like superscripts, consider using $\smash{2^d}$, but beware of character overlaps.
Screenshot of inconsistent line height to make way for a superscript.Screenshot of text with consistent line height.
2120
Matthew Finlayson @mattf.nl · 25/02/2025
This project was made feasible by the excellent open-source LLM training library @fairseq2.bsky.social; I highly recommend giving it a look! It made both SFT and DPO a piece of cake 🍰
0103
Matthew Finlayson @mattf.nl · 25/02/2025
6/ Our method is general, and we are excited to see how it might be used to better adapt LLMs to other tasks in the future. A big shout-out to my collaborators at Meta: Ilia, Daniel, Barlas, Xilun, and Aasish (of whom only @uralik.bsky.social is on Bluesky)
000
Matthew Finlayson @mattf.nl · 25/02/2025
5/ Training on self-demos, our model learns to better leverage the context to answer questions, and to refuse questions that it is likely to answer incorrectly. This results in consistent, large improvements across several knowledge-intensive QA tasks.
100
Matthew Finlayson @mattf.nl · 25/02/2025
4/ To obtain self-demos we generate candidate responses with an LLM, then use the same LLM to compare these responses to the gold one, choosing the one that best matches (or refuses to answer). Thus we retain the gold supervision from the original responses while aligning the training data.
111
Matthew Finlayson @mattf.nl · 25/02/2025
3/ OOD responses encourage the model to answer questions it does not know the answer to, and since retrievals are added post-hoc, the responses tend ignore or even contradict the retrieved context. Instead of training on these low-quality responses, we use the LLM to generate "self-demos".
100
Matthew Finlayson @mattf.nl · 25/02/2025
2/ A popular recipe for adapting LLMs for RAG involves adding retrievals post-hoc to an existing instruction-tuning dataset. The hope is that the LLM learns to leverage the added context to respond to instructions. Unfortunately, the gold responses in these datasets tend to be OOD for the model.
100
Matthew Finlayson @mattf.nl · 25/02/2025
🧵 Adapting your LLM for new tasks is dangerous! A bad training set degrades models by encouraging hallucinations and other misbehavior. Our paper remedies this for RAG training by replacing gold responses with self-generated demonstrations. Check it out here: arxiv.org/abs/2502.10
2181
Matthew Finlayson @mattf.nl · 12/12/2024
Putting together an unofficial usc Beamer template, I noticed that the USC style guide lists 4 formats for “cardinal red” but each of them is different: PMS 201 C is #9D2235 CMYK: 7, 100, 65, 32 is #A1003D RGB: 135, 27, 30 is #991B1E HEX: #990000 Is this normal? The CMYK is especially egregious.
The usc style guide list of formats for “cardinal” (see main post for list)The rgb and CMYK colors side by side. The CMYK is considerably pinker
000
Matthew Finlayson @mattf.nl · 12/12/2024
If you are registered for NeurIPS it should be available already online.
100
Matthew Finlayson @mattf.nl · 11/12/2024
NeurIPS should make them available online after one month :)
000
Matthew Finlayson @mattf.nl · 09/12/2024
In Vancouver for NeurIPS but don't have Taylor Swift tickets? You can still spend the day going through our tutorial reading list: cmu-l3.github.io/neurips2024-... Tuesday December 10, 1:30-4:00pm @ West Exhibition Hall C, NeurIPS
A diagram demonstrating text generation with beam search. One of the paths reads “Taylor Swift is the only person to…”
0292
Matthew Finlayson @mattf.nl · 06/12/2024
Shout out to our organizers @wellecks.bsky.social @abertsch.bsky.social @hails.computer @uralik.bsky.social @gneubig.bsky.social @abertsch.bsky.social Alex Xie, Konstantin Golobokov, and Zaid Harchaoui
110
Matthew Finlayson @mattf.nl · 06/12/2024
Curious about all this inference-time scaling hype? Attend our NeurIPS tutorial: Beyond Decoding: Meta-Generation Algorithms for LLMs (Tue. 1:30)! We have a top-notch panelist lineup. Our website: cmu-l3.github.io/neurips2024-...
Panelist photos: Rishabh Agarwal (Google, McGill), Noam Brown (OpenAl), Beidi Chen (CMU), Nouha Dziri (AI2), Jakob Foerster (Oxford, Meta)
1273
Matthew Finlayson @mattf.nl · 02/12/2024
😍 I went cycling there last year, what an amazing place
110
Matthew Finlayson @mattf.nl · 29/11/2024
The cover 😍
010
Matthew Finlayson @mattf.nl · 28/11/2024
Check your data mixture. @hamishivi.bsky.social is probably secretly up-weighting Latin in Dolma
161
Reposted by Matthew Finlayson
Hamish Ivison @hamishivi.bsky.social · 26/11/2024
What's that? A fully open LM competitive with Gemma and Qwen*? Happy to have helped a bit with this release (Tulu 3 recipe used here)! OLMo-2 13B actually beats Tulu 3 8B on these evals, making it a SOTA fully open LM!!! (*on the benchmarks we looked at, see tweet for more)
1101
Matthew Finlayson @mattf.nl · 26/11/2024
These folks have had a huge impact on my research
030
Reposted by Matthew Finlayson
Michael Saxon @saxon.me · 22/11/2024
#socalnlp is the biggest it's ever been in 2024! We have 3 poster sessions up from 2! How many years until it's a two-day event?? 🤯
1263
Matthew Finlayson @mattf.nl · 22/11/2024
This is niche but the LLM360 logo always reminds me of the 2014 iOS game Oquonie
LLM360 logo. A long-necked llama in the shape of an O. Screenshot from Oquonie with a long-necked character.
030
Matthew Finlayson @mattf.nl · 22/11/2024
And by linear I mean logistic 🤦‍♂️
130
Matthew Finlayson @mattf.nl · 22/11/2024
Hottest new research challenge: find the lost LLM head!
090
Matthew Finlayson @mattf.nl · 22/11/2024
Couldn’t you even just do a big linear regression here?
140
Matthew Finlayson @mattf.nl · 22/11/2024
Everyone follow Sean! He's been working nonstop to perfect our upcoming NeurIPS tutorial
060
Reposted by Matthew Finlayson
Michael Saxon @saxon.me · 21/11/2024
As "X is all you need" and "Transformers are Y" paper titles have died, I propose that we similarly retire: - X of thought - Chain of Y - "[topic] a comprehensive survey" where [topic] is exclusively post-2022 papers - Claims of reasoning/world model/planning based on private defns of one
9515
Matthew Finlayson @mattf.nl · 21/11/2024
I’m a big fan of titles that state the main result/finding, don’t ask questions, and don’t contain backronyms. I actually think “X are Y” titles are good because they tend to meet this criteria.
100
Matthew Finlayson @mattf.nl · 17/11/2024
I made a map! Thank you to my 2019 self for providing the code github.com/mattf1n/Reli...
A relief map of Los Angeles rendered in Blender giving it a 3D appearance.
2170
Matthew Finlayson @mattf.nl · 15/11/2024
Today I learned you can add a citation link to your GitHub repo. citation-file-format.github.io
citation-file-format.github.io
0100
Matthew Finlayson @mattf.nl · 15/11/2024
This is very cool! I’d love to subscribe to your blog posts but I can’t find an RSS feed
100
Reposted by Matthew Finlayson
Yoav Artzi @yoavartzi.com · 15/11/2024
Oh my, USC is an empire!
191
Matthew Finlayson @mattf.nl · 15/11/2024
Adding more every day!
110