Michael Saxon @saxon.me · 04/09/2026guys ace combat 4 was incredible. Ps2 games have an aura modern titles can't touch 100
Michael Saxon @saxon.me · 23/07/2026"Claude, generate and post a promotional animation for my book about the space shuttle. Show the shuttle taking off from above" 000
Michael Saxon @saxon.me · 04/04/2026Trump lackeys, if anyone, should think twice about normalizing ignoring pardons 010
Michael Saxon @saxon.me · 09/02/2026Hey this is really neat! My nearest neighbor is @andyliu.bsky.social 110
Michael Saxon @saxon.me · 06/02/2026Yep and it gets worse! Owner doesn't even care to remove hundreds of skills which directly instruct the model to install malware opensourcemalware.com/blog/clawdbo... 062
Michael Saxon @saxon.me · 04/02/2026I wonder if Jeff Epstein and Deepak Chopra's Lady Gaga/Madonna reconciliation dinner ft. Max Tegmark ever happened 230
Michael Saxon @saxon.me · 01/02/2026Would it surprise you to know there's a crypto rug pull for moltbook as well? The new meta is to claim that you aren't even affiliated with the people who launched your shitcoin 140
Michael Saxon @saxon.me · 30/01/2026The book "What Tech Calls Thinking - An inquiry into the Intellectual Bedrock of Silicon Valley" is amazing. I just read the chapter on genius and the aesthetic of genius and it is just mic drop after mic drop. (about Ayn Rand) 150
Michael Saxon @saxon.me · 23/01/2026seized the chance to get a good picture of the mountain on MLK day (taken out in the exurbs of Tacoma) 3180
Michael Saxon @saxon.me · 19/01/2026Also, watch out for impersonation on GH! Novel element of the attack here is also spoofing commits from her real account (they do not show up on her account's contribution list though) 120
Michael Saxon @saxon.me · 12/01/2026Of course they say in the channel description they are sloperators. Something tells me the half million v iewers are unaware 010
Michael Saxon @saxon.me · 12/01/2026AI slop channel masquerading as "professor/journalist talks to webcam" channel gets 500k views in 11 hours with a semi-fake news post about Trump declaring emergency powers. Seems bad! (Ignore grey distraction block filter) 110
Michael Saxon @saxon.me · 08/01/2026Thank you military industrial complex for helping me get cool ass sky cards 000
Michael Saxon @saxon.me · 19/11/2025In a few hours (11/19, 2PM PST) I will be giving this lecture on "conferencemaxxing" to help students prepare to make the most out of NeurIPS. This lecture is open to the public. If you're interested in joining, here's a GCal invite link: calendar.google.com/calendar/eve... 010
Michael Saxon @saxon.me · 18/11/2025I'm especially excited about our panel conversation with Eve Fleisig, @ofirpress.bsky.social, @idavidrein.bsky.social, and @saining.bsky.social Hope to see you there! 3/3 010
Michael Saxon @saxon.me · 18/11/2025The tutorial will consist of 4 parts. 1. (Intro) Epistemology, Design, and Practice 2. Limitations to current benchmarking approaches 3. Emerging paradigms: how do we solve these issues? 4. Panel conversation (see next tweet for panelists) 2/3 110
Michael Saxon @saxon.me · 18/11/2025Trying to decide what to do on the first day of #NeurIPS2025? Check out my, @marstin.bsky.social and @xiangyue96.bsky.social's tutorial, "The Science of Benchmarking: What's Measured, What's Missing, What's Next" on December 2 from 1:30 to 4:00pm. benchmarking.science What will we cover? 1/3 2214
Michael Saxon @saxon.me · 15/11/2025Rolled a custom (read: relatively privacy respecting) custom visitor map stack in 2.5h today with cursor 130
Michael Saxon @saxon.me · 14/11/2025Normalize questioning the utility of mathiness in ML conference papers! Are the equations supporting an argument or are they just a fancy way to express something simple? Do introduced terms do anything or get referenced anywhere? I find the answer is usually no in the kinds of papers I review 190
Michael Saxon @saxon.me · 05/11/2025🆕 from us at #EMNLP: Are LMs better at answering questions about Germany in German than in French? Is national knowledge linguistically contingent? Interestingly, only for some multilingual models is this true. Aya knows China best in Chinese, but LLaMA's best in English always. 3247
Michael Saxon @saxon.me · 27/10/2025It's live! Here's an example post: saxon.me/blog/2025/la... Turning the replies to a bluesky post into the comment section for a blogpost is a small concrete way to support the ecosystem: future visitors who want to add comments incentivized to interact on the platform Also, it's very easy to do: 150
Michael Saxon @saxon.me · 27/10/2025Prototyping bluesky comment integrations for the blog (gonna need to modify a lot more to make it fully work with my tempalte) Also, I am getting more and more indiewebpilled. Would any other NLPMLAI researcher-bloggers be interested in making a webring? 060
Michael Saxon @saxon.me · 18/10/2025I don't think this was malicious. There are real papers by the same authors. (Canivez and Youngstrom, 2019) and (Wasserman, 2019) do exist. Problem is they have different titles and are in different journals. Don't generate your references folks! 5211
Michael Saxon @saxon.me · 18/10/2025The viral "Definition of AGI" paper tells you to read fake references which do not exist! Proof: different articles present at the specified journal/volume/page number, and their titles exist nowhere on any searchable repository. Take this as a warning to not use LMs to generate your references! 615536
Michael Saxon @saxon.me · 21/06/2025On T2I-generated images, it is good at predicting the judgments of human raters from 10 countries of an image’s relevance to their own culture compared to a set of simple baselines. AIRe can be used to grade the "stylistic aspects" of a fantasy entity, not just match real stuff 4/5 100
Michael Saxon @saxon.me · 21/06/2025As far as we can tell, CAIRe works quite well. It is very performant at identifying the cultural origins of 𝗿𝗲𝗮𝗹, 𝗿𝗮𝗿𝗲 𝗲𝗻𝘁𝗶𝘁𝗶𝗲𝘀 based on many proxies, including country, region, religion, ethnicity, and even ancient civilizations. 3/5 100
Michael Saxon @saxon.me · 21/06/2025Our metric CAIRe (Cultural Attribution of Images with Retrieval) scores an input image using image retrieval over a multimodal KG and LM likelihood scores over entry data to assign cultural relevance scores to 𝐚𝐧𝐲 set of cultural labels based on 𝐚𝐧𝐲 cultural proxy (not just countries!). 2/5 100
Michael Saxon @saxon.me · 21/06/2025Multicultural text-to-image work requires costly, subjective human evaluation. Some of my projects have stalled because no automated, quantified "visual cultural attribution" metric existed. BITS undergrads Siddharth and Arnav Yayavaram, @simi97k.bsky.social, @gneubig.bsky.social, and I made one.1/ 182
Michael Saxon @saxon.me · 30/05/2025To be honest, I kinda love grok? (when it isn't being Elonbotomized to be a racism machine) So many rightoid maniacs query it expecting to see their conspiracist beliefs echoed back at them only to repeatedly get gently corrected with factual information lmao 060
Michael Saxon @saxon.me · 29/04/2025PSA for NAACL peeps from a southwest boi (sadly I won't be there): be sure to find a place to eat New Mexico style stacked enchiladas. You can get it "Christmas style" where its served with both red and green hatch chile. The hatch chile is integral, do not skip. Not photogenic, but very delicious 1120
Michael Saxon @saxon.me · 29/04/2025I wondered if it could really be all that bad from the beginning, after all users are signing up to publicly interact with each other on a forum but woof, I don't think I would have signed off on this broad of a "the LM is allowed to impersonate this" policy 291
Michael Saxon @saxon.me · 26/04/2025My stomach dropped when I saw the amount of quotes and replies... and a lot of the replies are about as aggressive and facile as I expected. note to self don't use the phrase "a n t i - A I" in a post lol 180
Michael Saxon @saxon.me · 22/04/2025Across multiple RMs, Terminator calibrates performance, getting near-optimal performance in significantly fewer tokens. Most interestingly, our model-predicted deadlines find the OPTIMAL budget, near the plateau where further spend isn't beneficial In this way Terminator is a tool any RM can use! 110
Michael Saxon @saxon.me · 22/04/2025Finally, we introduce Thought Terminator, our Schwarzeneggerian method to mitigating overthinking, which is a modified decoder that inserts interrupts every N tokens to tell the model how much compute it has left. Once that budget is spent it uses constrained decoding for budget forcing. 120
Michael Saxon @saxon.me · 22/04/2025In order to sample a more balanced distribution of questions across the difficulty spectrum, we introduce DUMB500, the Waluigi to MATH500 which consists of stupid easy Qs. This way we can get a more comprehensive view of overthinking, from the hardest GPQA and ZebraLogic Qs to literally "2+2=?" 120
Michael Saxon @saxon.me · 22/04/2025Our measure of overthinking is stupid simple: what's the delta between the mean/max token spend on each question vs the minimum for successful answers. There exists a clear trend between question difficulty (measured by success rates) and required spend. 110
Michael Saxon @saxon.me · 22/04/2025Check out our new paper on benchmarking and mitigating overthinking in reasoning models! From a simple observational measure of overthinking, we introduce Thought Terminator, a black-box, training-free decoding technique where RMs set their own deadlines and follow them arxiv.org/abs/2504.13367 2272
Michael Saxon @saxon.me · 28/03/2025of all the days to have a planned, 15+ hour university-wide power outage 230