Sign in

Clayton Thorrez

@cthorrez.bsky.social
447 followers 3.2K following 703 posts

LLMs and ratings at lmarena.ai Esports stuff for fun: cthorrez.github.io/riix/riix.html huggingface.co/datasets/EsportsBenc…

PostsRepliesMedia
Clayton Thorrez @cthorrez.bsky.social · 29/06/2026
what's up?
020
Reposted by Clayton Thorrez
Grace @gracekind.net · 12/11/2025
My brain is living in my head rent free
8997
Clayton Thorrez @cthorrez.bsky.social · 16/07/2025
EsportsBench refreshed with data up through June 2025, over 61k new matches across 20 esports have been recorded in the last 3 months! huggingface.co/datasets/Esp...
040
Clayton Thorrez @cthorrez.bsky.social · 15/07/2025
Extremely excited to announce that I've joined @lmarena.bsky.social ! For years I've been working in LLMs for my job, and hacking on rankings and ratings for fun, beyond thrilled to be able to join this project at the intersection!
120
Clayton Thorrez @cthorrez.bsky.social · 10/07/2025
Just ran into Simpsons paradox in the wild for the first time lol. Was looking at some data and was like "that doesn't look right all the means went up when all I did was assign groups differently, this is like Simpson's paradox or something lol"
100
Clayton Thorrez @cthorrez.bsky.social · 03/07/2025
if you fork any of my GitHub repos I WILL add you on LinkedIn. There are so few people actively working on rating system stuff and I want to talk to all of them
010
Clayton Thorrez @cthorrez.bsky.social · 03/07/2025
I just realized MSI is going on and in Vancouver so got a ticket for tomorrow lol. Any on esports/machine learning people going?
010
Clayton Thorrez @cthorrez.bsky.social · 03/07/2025
I love it when the same notation can mean the *exact* opposite thing when used by different authors... Should "A ≻ B" mean: "A is preferred to B (higher rating)" "B is preferred to A (lower rank number)"? arxiv.org/pdf/2411.049... www.tandfonline.com/doi/full/10....
020
Clayton Thorrez @cthorrez.bsky.social · 03/07/2025
if you have a really small network, and a really small dataset, is it possible to fuse an entire transformer training loop into a single kernel?
000
Clayton Thorrez @cthorrez.bsky.social · 02/07/2025
It turns out adding sum to 0 constraints on some parameters is actually a fair but harder than simplex constraints (non negative, sum to 1) in the iterative gradient based optimization setting. Is there a good equivalent to the softmax truck?
000
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
Super disappointed by this. I'm a huge fan of esports, attending numerous events over the past 12 years and watching countless hours on twitch. I would rather see esports shrink down to a grassroots core than get children addicted to gambling. x.com/riotgames/st...
x.com
Riot Games on X: "Why We're Opening Betting Sponsorships in Esports & How We're Doing It Responsibly" / X
Why We're Opening Betting Sponsorships in Esports & How We're Doing It Responsibly
120
Clayton Thorrez @cthorrez.bsky.social · 24/06/2025
qwen3-235b be like
010
Clayton Thorrez @cthorrez.bsky.social · 24/06/2025
How much value does thinking add to an LLM? Well for the largest Qwen3, the answer is -28 points Thinking on academic benchmarks seems to help a lot, I wonder what's going wrong in the arena? Maybe people can sense the hedging and don't like it, or it poisons its own context with overthinking
120
Clayton Thorrez @cthorrez.bsky.social · 19/06/2025
there's a lotta with the last name Ferguson not a lotta people with the first name Fergus
100
Clayton Thorrez @cthorrez.bsky.social · 19/06/2025
Mark Glickman is on a roll now! 2 Paper in two weeks This time extending the stength dependent draw model to the online setting for use in dynamic rating systems. Haven't read the whole thing but it looks to contain some cool approximation tricks for the posterior arxiv.org/abs/2506.11354
arxiv.org
Rating competitors in games with strength-dependent tie probabilities
Competitor rating systems for head-to-head games are typically used to measure playing strength from game outcomes. Ratings computed from these systems are often used to select top competitors for eli...
010
Clayton Thorrez @cthorrez.bsky.social · 14/06/2025
When gemini writes code in an artifact window, there are 9 buttons on the UI None of them are to copy the code
010
Clayton Thorrez @cthorrez.bsky.social · 09/06/2025
I want to train an 16 billion parameter model. Specifically, a 16 billion parameter TrueSkill model which fits a skill mean and variance for each of the 8 billion people on earth. But in my quest to scale rating systems, I guess I start with lichess, with 6B games and a measly 20M unique players
020
Clayton Thorrez @cthorrez.bsky.social · 05/06/2025
🚨NEW MARK GLICKMAN PAPER🚨 Paired comparison models with strength-dependent ties and order effects arxiv.org/abs/2505.24783 Basically what the title says, extending Bradley-Terry with ties, home field advantage, and allowing those factors to depend on how strong the competitors are.
arxiv.org
Paired comparison models with strength-dependent ties and order effects
Paired comparison models, such as the Bradley-Terry (1952) model and its variants, are commonly used to measure competitor strength in games and sports. Extensions have been proposed to account for or...
330
Clayton Thorrez @cthorrez.bsky.social · 05/06/2025
Fear leads to anger, anger leads to Rēvolfed, Rēvolfed is the path the dark side
020
Clayton Thorrez @cthorrez.bsky.social · 05/06/2025
Anyone know how to report a bug in Google Scholar? Online seems like there is not great public support. It's not my paper just one I reference a lot. Google thinks TrueSkill is in Russian! scholar.google.com/scholar?hl=e...
000
Clayton Thorrez @cthorrez.bsky.social · 04/06/2025
This is the post/thread that made me believe bsky has a future
010
Clayton Thorrez @cthorrez.bsky.social · 03/06/2025
Question for my rating systems peeps: what does the RD in Glicko, or the sigma in TrueSkill represent? How do you interpret that number?
000
Clayton Thorrez @cthorrez.bsky.social · 02/06/2025
Is this like a numerical issue on wolfram or am I missing something? This should be 0 right? www.wolframalpha.com/input?i=%281...
210
Clayton Thorrez @cthorrez.bsky.social · 31/05/2025
Ok I've officially had an idea for a modification of Elo that is consistently an improvement over base Elo averged over 20 datasets with significant baseline hyperparameter tuning. So I've passed the state of the art form the 1960's. The bad news is that it still gets mogged by Glicko.
100
Clayton Thorrez @cthorrez.bsky.social · 31/05/2025
look I'm a big fan of WSL, and I used it for all of my side projects, but it also has a enough issues to require me to make this powershell script lol
020
Clayton Thorrez @cthorrez.bsky.social · 31/05/2025
Interesting little finding, in Elo there are basically 2 hyperparameters, k which is the learning rate/step size, and scale, which is the temperature of the softmax. Most people understand k, high k means big rating changes after wins and losses, low k means small updates.
131
Clayton Thorrez @cthorrez.bsky.social · 22/05/2025
Pretty cool that I wrote the currently deployed ranking code for a company now valued at 600 million dollars 🤯🤯🤯 github.com/lm-sys/FastC... techcrunch.com/2025/05/21/l...
github.com
Accelerate Bradley Terry MLE model fitting by cthorrez · Pull Request #3523 · lm-sys/FastChat
Why are these changes needed? The bootstrap Bradley Terry model takes upwards of 15 minutes to run for 100 samples. This is costly on resources and hinders experiments such as studying hyperparamet...
000
Clayton Thorrez @cthorrez.bsky.social · 22/05/2025
Claude 4 Sonnet gave the best answer so far to my go-to first prompt: "Explain the relationship between Elo and Bradley-Terry from the perspective of machine learning and optimization" But I've used it so many times on ChatBot Arena I assume my conversations are in the training data by now
110
Clayton Thorrez @cthorrez.bsky.social · 20/05/2025
I actually think google search is so dead. I just issued a single word search for a common noun and the Wikipedia page for the thing was on the third page. There is no reason whatsoever in that scenario that it should not be in on the first page or info side panel
000
Reposted by Clayton Thorrez
Clayton Thorrez @cthorrez.bsky.social · 14/05/2025
What an amazingly relatable chapter name
151
Clayton Thorrez @cthorrez.bsky.social · 13/05/2025
Interesting new paper: arxiv.org/pdf/2505.03475 It's 12 pages which could be described in about 1 sentence: "Bradley Terry with per annotator ability parameters" Basically instead of in Elo, there is a constant log(10)/400 temperature, this learns a different temperature for each annotator
arxiv.org
100
Clayton Thorrez @cthorrez.bsky.social · 03/05/2025
I guess I picked the right day to start reading Stand on Zanzibar by John Brunner
120
Clayton Thorrez @cthorrez.bsky.social · 01/05/2025
Reading The Leaderboard Illusion today. I will say I've been a huge fan of ChatBot Arena ever since the start (and I'm a contributor to it), but I think there are some valid issues worth calling out. But I'm really not a fan of the format the discourse has taken on twitter which is overly hostile :/
100
Clayton Thorrez @cthorrez.bsky.social · 18/04/2025
why does chatgpt talk like a twitter AI influence now with emoji bulleted lists?
010
Clayton Thorrez @cthorrez.bsky.social · 17/04/2025
Extremely proud moment for myself today. I got my first academic citation! I worked on EsportsBench most of 2023 and into 2024, got rejected from Neurips DS&B, decided to still put it up on my website and maintain it on huggingface. Pleased some people find it interesting after so much work. :)
140
Clayton Thorrez @cthorrez.bsky.social · 16/04/2025
Here's EsportsBench v5! 72k new matches added from 2025-01-01 through 2025-03-31 and some data quality improvements to past data as well. Over 2.4 million rows of esports match data from 20 titles spanning over 25 years huggingface.co/datasets/Esp...
huggingface.co
EsportsBench/EsportsBench · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
042
Clayton Thorrez @cthorrez.bsky.social · 15/04/2025
The most gambity games
010
Clayton Thorrez @cthorrez.bsky.social · 15/04/2025
The least common openings in all of (li)chess These have each been played exactly once in the 3.3 billion games I've ingested so far. The names are quite charming
130
Clayton Thorrez @cthorrez.bsky.social · 14/04/2025
playing around with lichess data recently fun little stat, in the 30+0 time control, 1.7% games end on time AND as a draw! I learned this is the result when one player runs out of time but their opponent doesn't have sufficient pieces to perform a checkmate.
110
Clayton Thorrez @cthorrez.bsky.social · 13/04/2025
Today I got up early, walked 20 miles, and then ate 2000 calories at a diner. Currently taking the bus home. A perfect Saturday
030
Clayton Thorrez @cthorrez.bsky.social · 12/04/2025
haven't done pyspark in a while, glad to have some help lol
000
Clayton Thorrez @cthorrez.bsky.social · 11/04/2025
ohhhh just caught a really tricky pyspark issue in my code 1. override some spark conf to read a weird data format 2. read data 3. undo the spark conf change 4. process data due to the lazy execution, by the time the read actually occurred, the config change had already been undone
100
Clayton Thorrez @cthorrez.bsky.social · 08/04/2025
Gemini might be a great model but good lord the chat interface is terrible and buggy as all hell
140
Clayton Thorrez @cthorrez.bsky.social · 02/04/2025
While I don't use Facebook much, I do really appreciate the photo memories they show. This one is from a very memorable day in my freshman year of college 10 years ago. It was a microprocessor class that was honestly pretty overwhelming for me who had no programming or hardware experience.
120
Clayton Thorrez @cthorrez.bsky.social · 12/03/2025
I have now completed 37/73 of the books which have won the Hugo Award for Best Novel. While this is numerically over 50%, I'm still less than half way done because some more book will be written and win in the time it takes me to read the rest!
120
Clayton Thorrez @cthorrez.bsky.social · 06/03/2025
watched Section 31, it's terrible
000
Clayton Thorrez @cthorrez.bsky.social · 03/03/2025
oshit SIGBOVIK cfp snuck up on me again! I've got a really good idea though so need to set aside time to work on it :)
000
Clayton Thorrez @cthorrez.bsky.social · 03/03/2025
First day at a new job today :)
010
Clayton Thorrez @cthorrez.bsky.social · 02/03/2025
Today's Reading: A Unified Bayesian Perspective for Conventional and Robust Adaptive Filters This builds on Szczecinski's earlier Kalman Filter paper with a few more variants for robustness and allowing iterative updates at each time step. arxiv.org/abs/2502.18325
100
Clayton Thorrez @cthorrez.bsky.social · 28/02/2025
SeveranceAppleTVPlus is the worst name for a subreddit of all time
010