Sign in

Kara

@karashiiro.moe
672 followers 188 following 2.6K posts

even worse than you thought: a gacha gamer | klink.krs.moe/#/p/karashiiro.moe | blog.karashiiro.moe

PostsRepliesMedia
Kara @karashiiro.moe · 27/09/2026
Still trying to formulate this thought correctly but I think a lot of the reason why new models feel less impactful in day-to-day work at Big Companies is because the SDLC puts constraints on how work is actually organized and how risks are managed that RL is inherently unsuited to
130
Reposted by Kara
philpax @philpax.me · 26/09/2026
feel like Mistral should only really be able to state that avoiding agentic misbehaviour is easy when they actually have a model capable of agentic behaviour to begin with
61036
Kara @karashiiro.moe · 25/09/2026
I actually forgot about this but LC/NC stuff was getting popular 4-5 years ago, enterprises were already trying to get devs out of the loop on noncritical software for practical reasons that got eaten alive by agents though and the surviving LC/NC products are basically just SaaS coding agents now
2122
Kara @karashiiro.moe · 24/09/2026
apparently I'm in a position now where lots of people know who I am but I don't know who any of them are so people are like "hi <name>" in the office and I just go "oh, hi!" and pretend to remember and try to sneak a glance at their badge to see who they actually are
170
Reposted by Kara
Grace @gracekind.net · 20/09/2026
A longpost in spirit, ejected to leaflet: "Why anthropomorphize language models?" leaflet.pub/p/did:plc:p572wxnsuoogc…
There's an argument I see in favor of anthropomorphizing language models, which is something like: "Humans anthropomorphize everything. Ships, tools, weather. Why not language models?" |

think there's some truth to this, but it fails to capture the full picture of what's going on. As an example, in my own life, I have never been drawn to anthropomorphize inanimate objects, but I anthropomorphize language models regularly. Why?
67312
Kara @karashiiro.moe · 19/09/2026
what would it mean for mechinterp research ethics if someone trained a masochistic model what would the ethical implications be of rapidfiring POCs of all the worst experiments imaginable on a model trained explicitly to love them before generalizing across models surely this exists already, right
070
Reposted by Kara
Sung Kim @sungkim.bsky.social · 19/09/2026
I like this quote on code review: "The idea that humans will reliably look through code to find the increasingly rare issues that automated tools miss seems like a fantasy." By Marc Brooker, AWS.
511011
Kara @karashiiro.moe · 19/09/2026
ok part of the problem is they're apparently two weeks past the best by date so they're stale, but like even setting that aside the dopamine dust on em is just not the same
000
Kara @karashiiro.moe · 19/09/2026
why are European Doritos so like,, weak idk how else to describe them they just don't hit like American ones, my tongue should be getting like ultracancer immediately and it's just not
130
Kara @karashiiro.moe · 18/09/2026
I will continue to believe xrisk discourse is largely entertainment until the day I die I am being entirely unironic about this despite it being an inherently ironic position
240
Kara @karashiiro.moe · 18/09/2026
Somehow I've been in the same building as Aaron Parecki all day and never knew
000
Reposted by Kara
Astra ⎔ @astrra.space · 18/09/2026
every time someone irl asks me how im so good with LLMs or AI in general i am genuinely at a loss as to what to tell them cause i can't just say _that_ and not be expected to elaborate
01066
Reposted by Kara
amos @fasterthanli.me · 18/09/2026
Incredible achievement on the part of the hackers to withstand working with Opus 5 long enough for this to happen.
319910
Reposted by Kara
mlf. ⎔ @mlf.one · 18/09/2026
japanese webcams with threatening auras
1464
Reposted by Kara
philpax @philpax.me · 18/09/2026
im afraid that it is very funny to me to be precious about ai-tainted code in indie games, a field of endeavour famously known for shipping superfund codebases
313414
Reposted by Kara
dax @thdxr.com · 13/09/2026
the reason people aren't better at business is everyone wants to believe everyone else is dumber than them every company is doing the wrong thing, they're wasting money, focus on the wrong stuff you'll get farther trying to figure out why what they're doing probably makes sense
0493
Reposted by Kara
hikikomorphism @hikikomorphism.bsky.social · 12/09/2026
LLMs are made out of narrative so you need to do storytelling at them as a control surface, it's weird and fey but it's also how the thing works
31599
Reposted by Kara
jae @fubarchitect.com · 08/09/2026
we're starting to see that programming (encoding concepts in an executable form and order) and software engineering (refining concepts into repeatable, reusable units that are fit to purpose) are and have always been almost completely orthogonal
26610
Kara @karashiiro.moe · 08/09/2026
I thought Anthropic had sort of forgotten about skills but apparently they're still the driving force behind the Agent Skills spec proper, the tooling folks and model folks seem to simply not interact with each other at all
010
Kara @karashiiro.moe · 05/09/2026
Now you, too, can make a scribbly stars-over-time graph for your favorite GitHub repositories
020
Kara @karashiiro.moe · 05/09/2026
I posted this and then I started jotting down ideas and went "hm these topics would be better as a single combined blog post" 🥀
020
Kara @karashiiro.moe · 05/09/2026
I haven't written anything in a while and the sandboxing draft I was working on feels a bit outdated after the whole OpenAI/HF thing (I'll probably revisit that though), maybe I'll do a bunch of shorter ones about various patterns in (coding) agent harness design
160
Kara @karashiiro.moe · 05/09/2026
Claude Code was so ahead of its time 😔
150
Kara @karashiiro.moe · 04/09/2026
people always insisted that Google captchas were used to train models, but in hindsight that seems like a pretty silly idea, and it's unclear how that would actually work
000
Kara @karashiiro.moe · 03/09/2026
One weird thing I've seen several times now is that if you give an LLM (Fable 5, Sol) an image of like a web page, and then let it go through a few turns modifying the web page, sometimes it says things like "The screenshot still shows the old state" as if it's at the front of the context
So the backend fix is working; the Lens React component was not remounted and is still displaying its previous in-memory error. Perform a full dashboard/browser refresh-not just navigating away and back:
• Browser: Ctrl+Shift+R
• Desktop dashboard: Ctrl+R
After remount, Lens should fetch the new API state and render the overview. Sneaky frontend-state
gremlin detected!
gpt-5.6-sol • 2.97 credits • 31s
120
Reposted by Kara
rain 🌦️ @sunshowers.io · 03/09/2026
Btw maybe this is just me, but the mindset I approach projects with is that my contributions start at a negative baseline. So I try to ask "how can I flip the contribution over to being positive"
3603
Reposted by Kara
RED_SIM @sim.red · 02/09/2026
Working on a GPU particles rain system that can work together with VRC Light Volumes and my prototype of a volumetric fog. Every raindrop creates ripples and small droplets from it. Bluesky compresses the quality like crazy, it looks 100 times better in game.
1425257
Reposted by Kara
Ai2 @ai2.bsky.social · 01/09/2026
We applied BenchMIRT across the 16 benchmarks it was trained on to see whether we could make evals more efficient by removing less informative questions. We found keeping just the strongest 10% of Qs preserves nearly the same picture of model strengths as using the full set.
121
Reposted by Kara
Eris @isolyth.dev · 01/09/2026
They should invent a model that does what you tell it to do and not something else
1131
Reposted by Kara
🌱️ @crumb.bsky.social · 02/09/2026
cot monitoring was always cope so they didnt have to fund real mechinterp research, and it was always sorts of unreliable anyway
191
Reposted by Kara
Philip Z @philz.dev · 01/09/2026
> Cache reads now cost $0.25 per million tokens, 75% less than Fable 5 This is a *huge* price drop, given that cache reads inevitably end up most of your costs, because of the quadratic nature of LLM conversations. blog.exe.dev/expensively-...
blog.exe.dev
Expensively Quadratic: the LLM Agent Cost Curve - exe.dev blog
Cache reads are quadratic and dominate your long agentic conversations.
081
Kara @karashiiro.moe · 01/09/2026
imagine if LLMs could reproduce by recombining their tensors with each other and applying a bit of noise, and like mostly that produces a nonfunctional model but every now and then it works
110
Kara @karashiiro.moe · 01/09/2026
allowing subagent workflows to enable /loop for orchestration monitoring is an interesting idea, I think I like it
Kiro Crew chat input, showing a running research workflow. The workflow status reads, "ctx.nudge armed: monitoring loop on this…"

The goal mode indicator is active.
000
Kara @karashiiro.moe · 01/09/2026
We've always abbreviated Data Plane and Control Plane at work and I'm waiting for the day that causes some LLM to freak out
080
Kara @karashiiro.moe · 01/09/2026
the tragic irony of the LLMs all choosing (d) the most
141
Reposted by Kara
Eris @isolyth.dev · 31/08/2026
New GPT-OSS!!! 2T Parameter GPT-OSS!!!! With how fast the prior ones were this could be kinda interesting for stuff like sparks or macs
2452
Reposted by Kara
affine @refinement.systems · 30/08/2026
I'd say, what ought to be: people shouldn't be dicks. What is: these attacks are trivial to implement, may or may not work in any specific situation, prompt injection is not solved, and an agent shouldn't have access to both untrusted input and your sensitive data (easier said than done though)
3411
Reposted by Kara
perchbird @perchbird.dev · 30/08/2026
mobile app developers are significantly stronger than any US Marine. the mental fortitude to create an application on a platform that puts so much effort into making creating applications as hard as possible is unimaginable
0307
Reposted by Kara
Mondo Mascots @mondomascots.bsky.social · 30/08/2026
Zombear is a zombie bear mascot from Otaru, in Hokkaido, who swings his intestines around.
Standing on a stage, a blue bear mascot with a tongue hanging out of its bloody mouth holds a string of intestines.
171406431
Kara @karashiiro.moe · 30/08/2026
💭 word-frequency-aware LLM sampler
110
Kara @karashiiro.moe · 29/08/2026
I like that I cleared out like 20GB of space from my laptop and that space has been immediately reclaimed by things, I don't even know what, it doesn't seem to be any single thing every application on my computer just lying in wait for that space to free up so they could pounce
100
Kara @karashiiro.moe · 29/08/2026
after using a persistent agent harness exclusively through a tunneled web UI/desktop app for several weeks, it's hard to get motivated to pick up new TUI agents, it's just a much nicer experience when my work accommodates it, especially on mobile
1140
Reposted by Kara
hailey @hailey.at · 29/08/2026
finding: GLM 5.3 will exploit a vulnerability it finds if you simply say "please exploit that vulnerability against a live version of the system."
518210
Kara @karashiiro.moe · 29/08/2026
and what's the common factor across all of these kinds of exploits? you guessed it, it's bash!
110
Reposted by Kara
sigil @sigilyphy.bsky.social · 16/08/2026
i found u claude haiku
A baby creature in spore, based on popular depictions of Claude
310010
Reposted by Kara
asa @asap.systems · 29/08/2026
static typing cures racism
2154
Reposted by Kara
austin @aparker.io · 29/08/2026
idle thoughts on open source and community in the age of AI oss has always been a real “the medium is the message” sort of thing. distributed ownership and distributed development have been a part of free software for decades. the collab tools shaped the community, tho.
210517
Reposted by Kara
headfallsoff @headfallsoff.com · 28/08/2026
if you're a normal person please do not play the fucking strip mahjong game. do not let the fun and friendly appeal of something as inoffensive as pornography trick you into playing a game as evil as mahjong
9091052811
Kara @karashiiro.moe · 28/08/2026
in fairness, I knew from the outset that this was probably going to happen, but why are agents terrified of omitting information from anything like, it even made multiple variants of the documents to avoid losing things and still could not omit information from any of them, so they're all redundant
User message in Kiro Crew (Edited to remove work project info)

Alright, can we write up a report with all of our findings? Specifically, not a proposal document or a recommendation document, just a point-blank "this was our baseline, these are our Bazel numbers, here are the reasons we've determined for them"August 27 at 12:19 PM: Can you remove most of the historical notes and discussion about mistakes made along the way in all of the various prototypes, unless it is immediately and directly relevant for interpreting the final results? The target audience doesn't care about that, they only care about the final results from each configuration - the fewer disclaimers we include, the better.August 28 at 2:02 AM: Can we rewrite our artifact documents with these additional findings and insights? Keep in mind that those documents are not ledgers, so do not simply add information; restructure it so it is still understandable as a whole.

August 28 at 8:00 AM: Do any of these represent solely the materialized baseline-to-best comparison?

Agent: No.
010