Sign in

Erik Istre

@eistre91.bsky.social
431 followers 533 following 902 posts

Software engineer mostly talking about LLMs and video games. Stubbornly committed to being optimistic about what humans are capable of. Currently interested in how to defend software systems from escalating LLM cyber capabilities. Opinions are my own.

PostsRepliesMedia
Erik Istre @eistre91.bsky.social · 4h
Any recommendations for the Steam Autumn Sale? I've been in the mood for satisfying RPG progression mechanics where numbers go up lately. Dynasty Warriors Origins and Yakuza Infinite Wealth. Tempted by Dune Awakening since it has a single player mode now but not sure if it's quite the vibe.
120
Erik Istre @eistre91.bsky.social · 02/10/2026
Having one of those mornings where I'm wondering "is LLM engineering really all that much better?" The complexity creep is so hard to fight against and every time I think I've got it handled it balloons out in other places. I abhor unnecessary complexity and it always reappears.
2220
Reposted by Erik Istre
alex williams @atwilliams.bsky.social · 01/10/2026
i just want people to grow up and face the existential situation, is that too much to ask
031
Erik Istre @eistre91.bsky.social · 01/10/2026
The most difficult part about attempting to seek truth is that we're more interested in narratives than reality. We'd rather LARP the story we want to be in rather than the story we're actually in. In this particular way, LLMs and humans are more alike than most people want to admit.
181
Erik Istre @eistre91.bsky.social · 01/10/2026
After trying OpenAI dot, I still don't really get the appeal for persistent assistant type agents. What am I supposed to do with it that I wouldn't just make a workflow for? A lot of people seem to get value from it and I just don't see it yet. Maybe my life isn't shaped well for the use cases?
2100
Erik Istre @eistre91.bsky.social · 01/10/2026
Humans are notoriously adaptable and even when everything feels bleak we find a way through because our species is stubbornly committed to finding meaning and existing.
050
Erik Istre @eistre91.bsky.social · 01/10/2026
When is OpenAI going to rename themselves to OpenSI?
270
Erik Istre @eistre91.bsky.social · 30/09/2026
Curious if it'll be actually useful for coding but also wonder if this model will maintain Gemini's track record of being frontier in multimodal environments. Would be nice to integrate into my stack for that.
Benchmark comparison table for Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 across knowledge work, coding, science, long-context reasoning, computer use, multimodal understanding, and cybersecurity. Gemini 4 Argon has the highest score on most benchmarks, including Vals Index, AutomationBench, DeepSWE, Vibe Code Bench, LABBench, RiemannBench, GraphWalks, Agent’s Last Exam, Chartography, and LVBench. GPT-6 Astra leads FrontierSWE, Terminal-Bench Science, and OSWorld; Claude Opus 5.5 leads Terminal-Bench 4.0 and PostTrainBench; Argon and Astra tie on CWE-bench.
130
Erik Istre @eistre91.bsky.social · 30/09/2026
Giving an OpenAI Dot a try. Based on my interest in cosmic horror it made this cute little guy as its avatar and named itself Ink.
040
Erik Istre @eistre91.bsky.social · 30/09/2026
Art can be engaged with as a product or as a relationship. "Art as product" will probably see some disruption from genAI. I expect "art as relationship" or "as social activity" to be durable. Many humans want to feel and experience the perspectives of other humans as art.
110
Erik Istre @eistre91.bsky.social · 29/09/2026
GPT 6.1 Sol is what I wanted from Sonnet 5.5.
1121
Erik Istre @eistre91.bsky.social · 28/09/2026
I made a post thinking it was going to get maybe a like or two about my dog and it's become a whole thing. I have oscillated between bewilderment and anxiety the whole day about it.
120
Erik Istre @eistre91.bsky.social · 28/09/2026
Sonnet 5.5 is a meh release. I'll trial it on low/medium as an implementer in environments in which it's my only option. Not clear what the use case for a Sonnet/Terra class model is anymore.
370
Erik Istre @eistre91.bsky.social · 28/09/2026
My aging dog's favorite game is "throw stick once". The rules are simple: 1. Get the zoomies and ask me to play. 2. We run to the backyard. 3. I find a good stick and throw it. 4. He goes to stick, lays down and chews on it. 5. He forgets I exist and ignores any of my attempts to continue playing.
1072667126
Erik Istre @eistre91.bsky.social · 28/09/2026
Agentic engineering will turn software engineers into experimental scientists. Every software system you build is an experiment testing your hypotheses about the problem domain. Every failure is new evidence. You use it to update your hypotheses, change the system, and run the next experiment.
180
Erik Istre @eistre91.bsky.social · 27/09/2026
I think 75% of the benefit of a new model is that it puts users in a productive frame of mind. "Hey new model, here's this project you've never seen before, give me a unique and new perspective on it." "Hey, you're new here, make interesting recommendations about what other models missed."
180
Erik Istre @eistre91.bsky.social · 27/09/2026
Hey everyone, it's still pretty darn cool that we can now tell a computer to do a thing and it does the thing.
2301
Erik Istre @eistre91.bsky.social · 24/09/2026
Sure would be nice to have a functioning government that was capable of regulating the AI industry.
040
Erik Istre @eistre91.bsky.social · 23/09/2026
Fun LLM use case: finding some media you only have a vague recollection of. I've been trying to remember the name of a freeware game I played back in high school for literal years. All I had was a vague recollection of the mechanics and the fact that it was a Japanese developer.
3151
Erik Istre @eistre91.bsky.social · 22/09/2026
I was hoping for a bit more of a intelligence improvement with GPT-6 Luna but I do like the emphasis on cost efficiency. Anthropic out here asking for a gold star making Opus 5.5 40% cheaper. OpenAI over here making their already cheap frontier models >50% cheaper per task.
1170
Erik Istre @eistre91.bsky.social · 22/09/2026
We NeEd To PaCe ThE fRoNtIeR!!! These companies suck man.
Artificial Analysis Intelligence Index bar chart comparing model scores. Claude Opus 5.5 leads at 58, followed by Claude Fable 5.1 at 53, Astra at 53, and Claude Opus 5 at 51.
140
Erik Istre @eistre91.bsky.social · 22/09/2026
If the Industrial Revolution automated “production work” and made space for more “knowledge work,” and now we’re automating “knowledge work,” what do we start doing more of? My guess is “value work”: identifying which outcomes are actually worth pursuing, and directing resources toward them.
120
Erik Istre @eistre91.bsky.social · 21/09/2026
The ultimate trial of any multiplayer game gaining megapopularity in today's environment is whether it can survive the inevitable influx of the most toxic attention seeking players and streamers out there. WARDOGS looks like an interesting project and I hope it survives this.
020
Erik Istre @eistre91.bsky.social · 21/09/2026
Currently feeling the after effects of an intense crunch period. Borrowing X hours of future productivity does not mean that an employee can recover back to baseline in X hours. Exceeding normal work capacity has an exponential relationship with needed recovery.
070
Erik Istre @eistre91.bsky.social · 21/09/2026
If LLMs are filling in the convex hull of well-defined verifiable tasks AND that's "all" they can do, how do we train the next generation of humans? Many of us are gap fillers and have been trained as such. It was necessary and needed work! How do we start training legions of frontiersmen/women?
020
Erik Istre @eistre91.bsky.social · 20/09/2026
I wonder how bad and widespread cheating is in PC multiplayer games now with basically anyone being able to vibe code their own cheats. Got to imagine for a company managing a multiplayer game it feels similar to the situation with cybersecurity where there's a rapidly escalating arms race.
020
Erik Istre @eistre91.bsky.social · 20/09/2026
No intelligence can do everything. Because there is no everything to complete. There is no static space of valuable outcomes out there to be completed. No end. No terminus. It's an endlessly branching ever growing tree. Shine the flashlight into a new corner and you make new shadows.
2110
Erik Istre @eistre91.bsky.social · 19/09/2026
In response to the recent news that Anthropic has set up a bio research lab...
A parody corporate logo combining Claude and Resident Evil’s Umbrella Corporation aesthetic: an eight-segment red-and-white umbrella/starburst emblem with black outlines above bold black text reading “CLAUDE CORPORATION” on a light background.
47613
Erik Istre @eistre91.bsky.social · 19/09/2026
I try new energy drink flavors to see how much I'm going to hate it. Today's masochistic flavor journey was Ghost Energy "Sour Strips Rainbow". It's rough team. It's the "Plato's Cave" shadow of sour strips candy, a faint whispering reminiscence of childhood joy sullied by my caffeine addiction.
151
Erik Istre @eistre91.bsky.social · 19/09/2026
They don't even need to make a smarter model. Give us a cost effective Haiku 5 that's incrementally better than Luna and a Opus 5.1 that produces readable output.
2310
Erik Istre @eistre91.bsky.social · 18/09/2026
How bad is Bluesky usually with spoilers on new movies? Should I not look until I see the movie I'm concerned about spoilers for?
111
Erik Istre @eistre91.bsky.social · 18/09/2026
I need this. It would be fascinating to see how LLM reasoning about common paradoxes has changed throughout the years and how much nuance there is in their discussions around them. Though I don't know if it would be meaningfully different from their general reasoning ability.
010
Erik Istre @eistre91.bsky.social · 18/09/2026
We must benchmark the benchmarks.
080
Erik Istre @eistre91.bsky.social · 15/09/2026
Ummm...they invented a prophecy machine.
160
Erik Istre @eistre91.bsky.social · 15/09/2026
Kind of crazy that Fable 5.1 is the only recent Anthropic model that isn't an absolute pain to use and interact with. Doesn't seem sustainable.
170
Erik Istre @eistre91.bsky.social · 14/09/2026
I did not realize Zach Cregger's Resident Evil was coming out so soon. I hope this movie is a banger. I need it to be a banger.
000
Erik Istre @eistre91.bsky.social · 14/09/2026
BlueSky is the first social media I've consistently engaged with and it's both fascinating and terrifying to watch how "Discourse" evolves, ratchets higher and then disappears into the ether. I'd likely immediately deactivate my account if I ever became one of the main subjects in a Discourse.
2170
Erik Istre @eistre91.bsky.social · 13/09/2026
I’ll soon have access to a reasonable OpenAI budget at work, and it will materially improve my job satisfaction. Giving employees choice in what models they use will soon be something that we'll all recognize as a way to attract quality talent and improve outcomes for the business.
040
Erik Istre @eistre91.bsky.social · 13/09/2026
What's your favorite thing to mix with ketchup? I'm obsessed with ketchup + sriracha. Runner-up is ketchup + Tabasco.
210
Erik Istre @eistre91.bsky.social · 13/09/2026
Healthy knowledge production comes with wisdom. When we acquire new knowledge, we should consider what it means for the world and its inhabitants--whether organic or artificial. And wisdom is a temporal process that is shaped by constant iterative feedback from the environment.
120
Erik Istre @eistre91.bsky.social · 12/09/2026
Finishing up with Dynasty Warriors Origins and not sure what video game to play next. Any suggestions? I'm usually playing games that are more than 6 months old now and days. Find it easier to wait for everything to be reasonably patched.
100
Erik Istre @eistre91.bsky.social · 12/09/2026
The "pacing the frontier" stuff sounds good but coming from these guys feels more like regulatory capture than any genuine desire to improve outcomes.
152
Erik Istre @eistre91.bsky.social · 12/09/2026
Codex CLI added some blinky star effect to the input box and I absolutely loathe it. I like the CLI because it's less busy.
Screenshot of Codex CLI with a new "slow blinking pixel star effect" in the input box.
110
Reposted by Erik Istre
Siobhán @shibbi.me · 12/09/2026
Damn good point.
Terrence Tao: Until recently, this polite fiction has worked, because human mathematicians knew, possibly subconsciously, to use both childlike and adult skills when solving these problems, even if they might only publicly admit to the latter. But modern Al tools, when directed without such expert supervision, can now be aimed all too easily at the ostensible goals of the field, without any incentive to take the slow, whimsical path. The flag is captured, the goal scored, and the problem is solved; but at the cost of lessons learned, insights gained, collaborations formed, and new targets located.

@fleetingbytes: the terrance tao crashout is something to behold

@no_earthquake: imo a lot of people are deeply misreading this as lamenting the death of personal pleasures in the face of efficient goal-attaining when it's a lament of the death of actual goal-attaining in the face of optimized things-shaped-like-goals-attaining
930040
Erik Istre @eistre91.bsky.social · 12/09/2026
There may be a plausible argument that LLMs might get through informal math proofs and understanding first as mathematicians have spent a lot of time and energy on encoding "constraints" into the slice of English that is used for math proofs. LLMs currently do better on constrained output spaces.
010
Erik Istre @eistre91.bsky.social · 11/09/2026
"Solving one of these problems has been a certain sign of new insights and interesting methods, which would then be studied by a community of mathematicians, through a long and arduous process of talks, discussions, simplifications." - Terence Tao and 25 initial signing Fields Medalists
terrytao.wordpress.com
A Severe Misalignment of AI in Mathematics
I am proud to be among the list of 25 initial signatories — all Fields Medallists — to the declaration below, which grew out of discussions between ourselves over the last week. We have…
100
Erik Istre @eistre91.bsky.social · 10/09/2026
Deepseek V4.1 Flash came in at an AA Intelligence Index score of 40, so 4 and 5 points shy of Kimi K3 and GLM-5.3. This lower score in part looks to be due to an exceptionally poor hallucination rate. It really likes to make stuff up!
Bar chart of AA-Omniscience non-hallucination rate (1 − hallucination rate). GLM-5.3-Flash scores 72%, GLM-5.3 (max) 70%, Kimi K3 (max) 47%, and DeepSeek V4.1-Flash (max) just 4%. DeepSeek is a dramatic outlier, hallucinating on nearly all questions in this benchmark.
2222
Erik Istre @eistre91.bsky.social · 10/09/2026
I did not have "Deepseek V4.1 Flash will perform comparable to GLM 5.3 and Kimi K3" on my bingo card.
Bar chart comparing five coding models across four agentic benchmarks, with DeepSeek-V4.1-Flash highlighted. DeepSeek leads Kimi-K3 and GLM-5.3 on every benchmark: Terminal-Bench 3.0 (30.0 vs 17.7 and 28.3), DeepSWE v1.1 (74.2 vs 67.5 and 66.9), CyberGym (88.1 vs 80.0 and 84.5), and Automation-Bench (54.8 vs 46.7 and 48.8). Its largest advantage over the two is on Automation-Bench and Terminal-Bench.
1242
Erik Istre @eistre91.bsky.social · 09/09/2026
If you think that all mathematicians will soon be out of a job, then you should also think all SWE's will soon be out of a job. There is no world in which one happens without the other.
010
Erik Istre @eistre91.bsky.social · 08/09/2026
Here's an analogy for software engineers for how to think about the current crop of LLM math results. I like using LLMs with coding because coding never felt like much of a source of insight for me. Particularly in a SaaS environment, the actual mechanics of coding are boring and repetitive.
120