Sign in

Clayton Thorrez

@cthorrez.bsky.social
449 followers 3.2K following 703 posts

LLMs and ratings at lmarena.ai Esports stuff for fun: cthorrez.github.io/riix/riix.html huggingface.co/datasets/EsportsBenc…

PostsRepliesMedia
Clayton Thorrez @cthorrez.bsky.social · 29/06/2026
what's up?
020
Reposted by Clayton Thorrez
Grace @gracekind.net · 12/11/2025
My brain is living in my head rent free
8997
Clayton Thorrez @cthorrez.bsky.social · 16/07/2025
EsportsBench refreshed with data up through June 2025, over 61k new matches across 20 esports have been recorded in the last 3 months! huggingface.co/datasets/Esp...
040
Clayton Thorrez @cthorrez.bsky.social · 15/07/2025
I am humbled to join this excellent team and work on delivering the highest quality human preference LLM evals! ⚔️⚔️⚔️
020
Clayton Thorrez @cthorrez.bsky.social · 15/07/2025
I've been following this project since it first showed up in my google scholar notifications for papers that cite Elo in 2023 and had fun experimenting with their data and contributing open source before it was a company.
110
Clayton Thorrez @cthorrez.bsky.social · 15/07/2025
Extremely excited to announce that I've joined @lmarena.bsky.social ! For years I've been working in LLMs for my job, and hacking on rankings and ratings for fun, beyond thrilled to be able to join this project at the intersection!
120
Clayton Thorrez @cthorrez.bsky.social · 11/07/2025
curiosity discovery goofiness
110
Clayton Thorrez @cthorrez.bsky.social · 10/07/2025
Then I spent another hour debugging the data for nans and nulls and corruption until I realized that it actually was Simpson's paradox
000
Clayton Thorrez @cthorrez.bsky.social · 10/07/2025
Just ran into Simpsons paradox in the wild for the first time lol. Was looking at some data and was like "that doesn't look right all the means went up when all I did was assign groups differently, this is like Simpson's paradox or something lol"
100
Clayton Thorrez @cthorrez.bsky.social · 07/07/2025
Interesting, I think I can kinda concede Chad as a contrarian grifter but I still like the Hip Crime vocab. Got some decent chuckles from me
010
Clayton Thorrez @cthorrez.bsky.social · 07/07/2025
Void, please analyze my profile and assign me to a cognitive continent.
110
Clayton Thorrez @cthorrez.bsky.social · 06/07/2025
A point I found funny is the idea where giant corporations are basically using ChatGPT to do their homework and almost nobody cares if it's conscious or not. I really like Chad, Begi, and Shalmaneser, don't really care for anyone else
110
Clayton Thorrez @cthorrez.bsky.social · 06/07/2025
This was a fascinating mix of super on point and totally off mark predictions. Focuses on fertility/population demographics, colonialism, eugenics/genetic engineering, and even some specific geopolitics are correctly predicted to be super hot issues.
110
Clayton Thorrez @cthorrez.bsky.social · 06/07/2025
Lots of different little side stories and snippets adding to the immersion of this world. The other thing I find interesting about books written in the past, about a time which is their future but is now my past, is learning their predictions about the future.
110
Clayton Thorrez @cthorrez.bsky.social · 06/07/2025
Finished it last night and I have some thoughts lol. Overall I definitely didn't enjoy it as much as some other books. I never really got connected to the characters, and I didn't find the main story too engaging. On the plus side I loved the worldbuilding
110
Clayton Thorrez @cthorrez.bsky.social · 04/07/2025
well it's pretty hard to argue with that :)
000
Clayton Thorrez @cthorrez.bsky.social · 04/07/2025
Then compute that prob over the population of players and sort by highest avg prob. Finally, a model does not have to be correct to be useful, in a lot of cases you can get great accuracy without even using a vector, just representing the overall skill with a scalar.
000
Clayton Thorrez @cthorrez.bsky.social · 04/07/2025
Often something of interest is overall skill, in which case aggregation can apply over vectors and sort by mean. Sometimes you don't need to directly order. Can use a parametric model over two vectors A and B to produce a probability that A will beat B.
100
Clayton Thorrez @cthorrez.bsky.social · 04/07/2025
Computer science, took some stats and optimization courses. I think I have a different opinion about vectors, I can think of a lot of ways to order them. For example if each dimension represents a specific skill, then per-skill orderings produce per-skill leaderboards
100
Clayton Thorrez @cthorrez.bsky.social · 03/07/2025
if you fork any of my GitHub repos I WILL add you on LinkedIn. There are so few people actively working on rating system stuff and I want to talk to all of them
010
Clayton Thorrez @cthorrez.bsky.social · 03/07/2025
I just realized MSI is going on and in Vancouver so got a ticket for tomorrow lol. Any on esports/machine learning people going?
010
Clayton Thorrez @cthorrez.bsky.social · 03/07/2025
I love it when the same notation can mean the *exact* opposite thing when used by different authors... Should "A ≻ B" mean: "A is preferred to B (higher rating)" "B is preferred to A (lower rank number)"? arxiv.org/pdf/2411.049... www.tandfonline.com/doi/full/10....
020
Clayton Thorrez @cthorrez.bsky.social · 03/07/2025
if you have a really small network, and a really small dataset, is it possible to fuse an entire transformer training loop into a single kernel?
000
Clayton Thorrez @cthorrez.bsky.social · 02/07/2025
It turns out adding sum to 0 constraints on some parameters is actually a fair but harder than simplex constraints (non negative, sum to 1) in the iterative gradient based optimization setting. Is there a good equivalent to the softmax truck?
000
Clayton Thorrez @cthorrez.bsky.social · 29/06/2025
why do it for free though? jobs.gem.com/bluesky/am9i...
jobs.gem.com
Bluesky Jobs
Bluesky Jobs
010
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
very slightly :) basically my rules of thumb are to never use numpy on scalars unless the function simply doesn't exist in base python, and to try the simpler thing, ** and pow are general and need to support raising numbers to any power, num*num is a single multiplication
100
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
This could be someone's habit from base python, where multiplying a scalar number by itself is actually faster than **, np.square, etc
110
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
Ok so the main issue is that the same number is multiplied by itself which can be replaced by square or exponentiation. I think since it's vectorized Jax, if it's jitted it should hopefully all be the same.
100
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
Cool release and impressive result, but just FYI the LMArena leaderboard doesn't use Elo anymore, it's a variant of a Bradley Terry model
110
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
They're my favorite sketch comedy group in like the last decade, lots of good ones
020
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
I've worked on esports prediction for close to 10 years, a field populated largely by gambling and I made my rating systems package non-commercial just to stop anyone from using it at gambling companies. (If you want to use it for any other purpose I'll grant you a license for free)
000
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
Super disappointed by this. I'm a huge fan of esports, attending numerous events over the past 12 years and watching countless hours on twitch. I would rather see esports shrink down to a grassroots core than get children addicted to gambling. x.com/riotgames/st...
x.com
Riot Games on X: "Why We're Opening Betting Sponsorships in Esports & How We're Doing It Responsibly" / X
Why We're Opening Betting Sponsorships in Esports & How We're Doing It Responsibly
120
Clayton Thorrez @cthorrez.bsky.social · 27/06/2025
I'm genuinely curious, how would you write it?
100
Clayton Thorrez @cthorrez.bsky.social · 26/06/2025
Are they running on the same or separate machines? Could the happy one be benefiting by hitting the cache on data prepped by the skeleton?
010
Clayton Thorrez @cthorrez.bsky.social · 25/06/2025
Can't wait to listen to this, I took Professor Singh's RL class at UMich in 2017/2018 and I remember when AlphaGo Zero was released he scrapped the planned lecture and just talked about that. Definitely got me interested in RL
010
Clayton Thorrez @cthorrez.bsky.social · 25/06/2025
I don't think either mode is doing any search here. If I'm reading the paper right it should be the same model weights, just whether they allow it to use reasoning tokens before the final response or not
100
Clayton Thorrez @cthorrez.bsky.social · 25/06/2025
In all the other fair use cars I've heard of (mostly YouTubers) I always heard of transformative in the context of parody, analysis, remix, etc. I'd never once considered a physical form of transformation but I guess it is literally true.
100
Clayton Thorrez @cthorrez.bsky.social · 25/06/2025
The "transformation" in question is the physical transformation of chopping up the book? And if they trained on them, and then a user asks a question, and the model outputs substantial unaltered segments, would that be distribution outside the company?
100
Clayton Thorrez @cthorrez.bsky.social · 25/06/2025
Super interesting case and post, here's what caught my attention: "The summary judgement found that these scanned books did fall under fair use, since they were transformative versions of the works and were not shared outside of the company."
130
Clayton Thorrez @cthorrez.bsky.social · 24/06/2025
qwen3-235b be like
010
Clayton Thorrez @cthorrez.bsky.social · 24/06/2025
How much value does thinking add to an LLM? Well for the largest Qwen3, the answer is -28 points Thinking on academic benchmarks seems to help a lot, I wonder what's going wrong in the arena? Maybe people can sense the hedging and don't like it, or it poisons its own context with overthinking
120
Clayton Thorrez @cthorrez.bsky.social · 23/06/2025
it won't be me, but hopefully they have better luck hiring this time
010
Clayton Thorrez @cthorrez.bsky.social · 23/06/2025
I recently learned they might be interested in trying this now lol
010
Clayton Thorrez @cthorrez.bsky.social · 19/06/2025
there's a lotta with the last name Ferguson not a lotta people with the first name Fergus
100
Clayton Thorrez @cthorrez.bsky.social · 19/06/2025
Mark Glickman is on a roll now! 2 Paper in two weeks This time extending the stength dependent draw model to the online setting for use in dynamic rating systems. Haven't read the whole thing but it looks to contain some cool approximation tricks for the posterior arxiv.org/abs/2506.11354
arxiv.org
Rating competitors in games with strength-dependent tie probabilities
Competitor rating systems for head-to-head games are typically used to measure playing strength from game outcomes. Ratings computed from these systems are often used to select top competitors for eli...
010
Clayton Thorrez @cthorrez.bsky.social · 14/06/2025
When gemini writes code in an artifact window, there are 9 buttons on the UI None of them are to copy the code
010
Clayton Thorrez @cthorrez.bsky.social · 10/06/2025
no idea what is going on here but I'm all in favor of llm nonsense
110
Clayton Thorrez @cthorrez.bsky.social · 09/06/2025
I want to train an 16 billion parameter model. Specifically, a 16 billion parameter TrueSkill model which fits a skill mean and variance for each of the 8 billion people on earth. But in my quest to scale rating systems, I guess I start with lichess, with 6B games and a measly 20M unique players
020
Clayton Thorrez @cthorrez.bsky.social · 08/06/2025
Are you recording data of the likes and comments on these?
000
Clayton Thorrez @cthorrez.bsky.social · 07/06/2025
More surprising is that he was a partner at Soros Fund, how did he get into this administration lol?
100