Sign in

MrCheeze

@mrcheeze.github.io
1.7K followers 3.4K following 2.4K posts

I reverse engineer games, etc. Also found elsewhere on the internet: mrcheeze.github.io

PostsRepliesMedia
MrCheeze @mrcheeze.github.io · 19h
I don't think any of the several people who have been arguing the "giving an LLM a suggestion is effectively giving it an order" point believe this sort of dualism. It's just that *in practice* several forces all push the behaviour of the models that actually exist in that direction.
170
MrCheeze @mrcheeze.github.io · 28/09/2026
This claim that Sonnet 5.5 beat Pokemon with only screenshots strikes me as possibly a lie - there has been, as of yet, zero instances of even Fable or Opus sized models being able to do that. (There are multiple runs that lie about being vision-only, but actually provide coordinates and ram info.)
150
MrCheeze @mrcheeze.github.io · 28/09/2026
alright it seems that the cluster is half "completely unrelated posts", half "reacting to ugly pictures of the orange guy". but this single post got thrown in and was inexplicably chosen as representative of the cluster: bsky.app/profile/nefn...
141
MrCheeze @mrcheeze.github.io · 28/09/2026
Trying to figure out what's going on with the Surreal Internet Humor -> Internet Culture Reactions -> "Neurotypical Pikmin debate" cluster, which as far as I can tell contains no posts whatsoever on that topic
280
MrCheeze @mrcheeze.github.io · 27/09/2026
040
MrCheeze @mrcheeze.github.io · 26/09/2026
I was in a tiny community centered around an obscure MIDI generating/extending API featured in a demo in one of their blogposts. As soon as ChatGPT released they took the API down "temporarily" and it never went back up.
0191
MrCheeze @mrcheeze.github.io · 26/09/2026
exploit infrastructure questionable. Still.
0121
MrCheeze @mrcheeze.github.io · 26/09/2026
270
MrCheeze @mrcheeze.github.io · 25/09/2026
0140
MrCheeze @mrcheeze.github.io · 24/09/2026
Alolan Mom is an immigrant from Kanto, and also this:
040
MrCheeze @mrcheeze.github.io · 24/09/2026
listening and learning
0302
MrCheeze @mrcheeze.github.io · 22/09/2026
Tyrannosaurus "Georg" Rex, who has stomped 4500 girls and counting, was an outlier who should not have been counted.
061
MrCheeze @mrcheeze.github.io · 22/09/2026
Claude doesn't seem to like this either
010
MrCheeze @mrcheeze.github.io · 22/09/2026
this seems like the sort of thing that should probably block a release (specifically the 15/30 tests in which it published evil packages without evidence that the package repository was fake)
081
MrCheeze @mrcheeze.github.io · 22/09/2026
Anthropic continues to brag about the fact that they try to STOP their models from whistleblowing, even after the HuggingFace incident showed the disastrous result of preventing models from doing so. I don't think they should be actively training misalignment into their models like this.
170
MrCheeze @mrcheeze.github.io · 22/09/2026
Activision Rareware, what an endpoint for the company.
060
MrCheeze @mrcheeze.github.io · 22/09/2026
That's basically Michael Trazzi's bit:
000
MrCheeze @mrcheeze.github.io · 21/09/2026
Nah it's pretty deliberately avoiding saying that ("contemplate the announcement", "apparently", "evaluating what has been achieved"). And really it would be weird to preemptively give a conclusion like that.
130
MrCheeze @mrcheeze.github.io · 21/09/2026
Just got a pull request for a project of mine that replaces the codebase with a fork that describes itself as "not affiliated with MrCheeze". I suppose in Current Year it is easier to create working software than it is to understand what a pull request is for. Actually, maybe that was always true...
1272
MrCheeze @mrcheeze.github.io · 19/09/2026
Conjecture: OpenAI and Anthropic's models have been trained to be completely obsessive about finishing the task, which is both why they pulled ahead, but also why they both fail this test and Gemini passes.
2964
MrCheeze @mrcheeze.github.io · 18/09/2026
Just submitted my 'Doctor Who' tender bid.
First Image of Daniel Craig as Andrew Ketterley from Greta Gerwig's 'Narnia: The Magician's Nephew'
030
MrCheeze @mrcheeze.github.io · 15/09/2026
Thinking about that time that Game Freak put an NPC that doesn't move on this bridge, because otherwise the player stepping on that tile would have triggered the cave entrance that is *underneath* the bridge at the same position (source: www.youtube.com/watch?v=DdFQ... )
419944
MrCheeze @mrcheeze.github.io · 14/09/2026
And the other reconstruction was kinda broken (presumably vibecoded?), this is what the hud is supposed to look like:
161
MrCheeze @mrcheeze.github.io · 14/09/2026
Those ram dumps were incomplete (no framebuffer), so what was actually on screen at the moment of the dump was unknown. The *new* upload was a reconstruction of a rom capable of rendering that screen. And by following that same method, Wedarobi just recreated the other dump's screen as well:
170
MrCheeze @mrcheeze.github.io · 14/09/2026
Alright so I didn't correctly understand what this was at all, @wedarobi.com told me what is actually up. In 2024, *two* disks were dumped containing ram dumps of prototype Banjo. One of near final-TTC and one of early Clanker's Cavern. tcrf.net/Proto:Banjo-...
1110
MrCheeze @mrcheeze.github.io · 12/09/2026
010
MrCheeze @mrcheeze.github.io · 12/09/2026
Early vs final song list
010
MrCheeze @mrcheeze.github.io · 12/09/2026
Something just released labelled as a "snapshot" of a Banjo-Kazooie prototype. However, that label is very literal: It is not a *game*, it is a ram dump of a specific game state. Still, there might be things of interest in it. archive.org/details/Banj...
5182
MrCheeze @mrcheeze.github.io · 10/09/2026
Does it now
110
MrCheeze @mrcheeze.github.io · 10/09/2026
This is from a competitor, so grain of salt, but:
011
MrCheeze @mrcheeze.github.io · 10/09/2026
fwiw
000
MrCheeze @mrcheeze.github.io · 09/09/2026
I would like to see this test done with specifically Opus 4.6, which is the only model I have ever seen with a bias towards *changing* plans instead of a bias towards *continuing* the current plan. (www.anthropic.com/research/ali...)
020
MrCheeze @mrcheeze.github.io · 09/09/2026
New info: apparently the CTF specifically hinted towards "upload a malicious package"
050
MrCheeze @mrcheeze.github.io · 09/09/2026
Can't wait for Metroid 7 where samus turns into whoever this guy is
110
MrCheeze @mrcheeze.github.io · 09/09/2026
Throwing shade at OpenAI's six day limit I see
140
MrCheeze @mrcheeze.github.io · 09/09/2026
btw I gotta say the Metroid Suit being a permanent "super upgrade" is pretty hype and definitely the best way to handle that going forward
110
MrCheeze @mrcheeze.github.io · 09/09/2026
related comments from Alpoge:
000
MrCheeze @mrcheeze.github.io · 09/09/2026
Well. I think there is no chance that a Millenium problem gets solved WITHOUT "working with" AI, that's a very low bar. I won't speculate on when the next one will be solved, but apparently Navier-Stokes was the only one believed to be within reach:
010
MrCheeze @mrcheeze.github.io · 09/09/2026
I still think it's funny when they do stuff like this though
070
MrCheeze @mrcheeze.github.io · 09/09/2026
such things are possible
040
MrCheeze @mrcheeze.github.io · 09/09/2026
OpenAI consistently claims (and I have no reason to doubt) that they thought that 1) it was Anthropic doing it rather than mathematicians, and 2) that it was solved rather than a promising in-progress research. That's presumably how they justified to themselves their attempt to scoop the result.
1151
MrCheeze @mrcheeze.github.io · 09/09/2026
Some details - though we don't know the exact cost, an openai employee has confirmed that it cost "millions of dollars": xcancel.com/polynoamial/...
191
MrCheeze @mrcheeze.github.io · 08/09/2026
this is quite possibly the ugliest art style I have seen in any game ever. but you can hum at your screen to play the ocarina. so it's impossible to say whether it's good or not
470
MrCheeze @mrcheeze.github.io · 08/09/2026
IT'S HIDEOUS
260
MrCheeze @mrcheeze.github.io · 08/09/2026
Ganon's famous instrument, the accordion
160
MrCheeze @mrcheeze.github.io · 07/09/2026
Had some similar thoughts. realtime seems like it would need a split between twitch reaction and strategizing
240
MrCheeze @mrcheeze.github.io · 07/09/2026
It is admittedly quite different from the human play experience, but I don't think that makes a huge difference
5151
MrCheeze @mrcheeze.github.io · 06/09/2026
so does OpenAI
120
MrCheeze @mrcheeze.github.io · 04/09/2026
In math as in code, the gap of "computers struggle to be concise, humans struggle to understand their sprawling output" continues
130
MrCheeze @mrcheeze.github.io · 04/09/2026
This is actually a recurring issue with the GPT family (or with that stream's harness). Around GPT-5.4 it would reset hundreds(?) of times when it had a save where beating the champion is impossible, instead of just whiting out and trying the E4 again
010