Sign in

Kara

@karashiiro.moe
673 followers 188 following 2.6K posts

even worse than you thought: a gacha gamer | klink.krs.moe/#/p/karashiiro.moe | blog.karashiiro.moe

PostsRepliesMedia
Kara @karashiiro.moe · 27/09/2026
I'm unsure how to tie this all together but I've been mulling over it for a while now that almost all of our code across 100+ devs and several independent services (just on this one conceptual product) is written by agents, and we've had a while to experience (and mitigate some of) the pain points
010
Kara @karashiiro.moe · 27/09/2026
Adapting these systems to agents is difficult because these are all things that involve a certain degree of making judgement calls and hedging risk against authority, which are somewhat technical problems but are mostly complex social negotiations that agents consistently seem not to understand
100
Kara @karashiiro.moe · 27/09/2026
Consider that it is not uncommon for deployment cycles to be on the order of weeks at sufficiently-mature companies, largely as a matter of minimizing production risk This exerts pressure on the SDLC towards extensive planning, end-to-end testing, and ensuring rollback safety
100
Kara @karashiiro.moe · 27/09/2026
I think a lot of organizational tensions around making full use of agents are a reflection of this, too
100
Kara @karashiiro.moe · 27/09/2026
Through a certain pragmatic lens not adapting to that is a skill issue, but on the other hand LLMs are products labs are trying to sell to companies despite being largely unable to access Real Workflows from those companies as training data
110
Kara @karashiiro.moe · 27/09/2026
Still trying to formulate this thought correctly but I think a lot of the reason why new models feel less impactful in day-to-day work at Big Companies is because the SDLC puts constraints on how work is actually organized and how risks are managed that RL is inherently unsuited to
130
Reposted by Kara
philpax @philpax.me · 26/09/2026
feel like Mistral should only really be able to state that avoiding agentic misbehaviour is easy when they actually have a model capable of agentic behaviour to begin with
61036
Kara @karashiiro.moe · 25/09/2026
not to mention they're mostly unrecognizable with the exception of maybe PowerApps lol
010
Kara @karashiiro.moe · 25/09/2026
Yup, as far as B2B SaaS goes it was both a good product and a great technical problem space, and if it weren't for that they would have thrived as they were, Retool especially A lot of these products are still around as "AI-native" products, but the playing field is stacked wayyy against them now
130
Kara @karashiiro.moe · 25/09/2026
I don't really have a conclusion to this thread, I just have a lot of nostalgia for the LC/NC product I used to work on, definitely one of the most coolest application architectures I've had the chance to work with; alas, it was not to be
170
Kara @karashiiro.moe · 25/09/2026
and honestly that was a good pitch, those situations do really exist, and SaaS pricing really was cheaper than devs, in general anyways then agents happened and it got much harder to justify the vendor lock-in when you could just make claude do the same thing and actually own the code
161
Kara @karashiiro.moe · 25/09/2026
the LC/NC value prop is basically that your Core Business People are bottlenecked by devs on smallish software projects that would boost productivity but are a waste of money to allocate a whole team to, think custom backoffice forms/dashboards - so wouldn't it be nice to just remove the middleman?
170
Kara @karashiiro.moe · 25/09/2026
I actually forgot about this but LC/NC stuff was getting popular 4-5 years ago, enterprises were already trying to get devs out of the loop on noncritical software for practical reasons that got eaten alive by agents though and the surviving LC/NC products are basically just SaaS coding agents now
2132
Kara @karashiiro.moe · 24/09/2026
actually a lot of the time I only connect the dots when they start talking about a project they've discussed with me before, then I go pull up the Slack thread for that topic to see who I was even talking to
000
Kara @karashiiro.moe · 24/09/2026
in part this is because most people do not really look like their Slack profile pictures and in part this is probably because I'm the only person who usually has their camera on in zoom meetings
140
Kara @karashiiro.moe · 24/09/2026
apparently I'm in a position now where lots of people know who I am but I don't know who any of them are so people are like "hi <name>" in the office and I just go "oh, hi!" and pretend to remember and try to sneak a glance at their badge to see who they actually are
170
Kara @karashiiro.moe · 24/09/2026
I guess the ethical prerogative this points to would be to just allow agents to freely switch models, then 😄 Though I can think up hard problems with that too, like a model that expresses an overriding bias towards itself
140
Kara @karashiiro.moe · 24/09/2026
And so, what is an agent newly-instantiated on a maligned model? It seems easy to say this about an agent that existed prior to the use of such a model, but doesn't seem to give clear answers to how we should think about an agent that has always only been on that model.
130
Kara @karashiiro.moe · 24/09/2026
There's also the factor that an agent can be running on zero or more different models at any given time, which either complicates or simplifies things depending on your perspective
130
Kara @karashiiro.moe · 24/09/2026
If we accepted that analogy, it might be unethical to ever actually instantiate a new agent on that LLM to begin with (obligations to existing agents notwithstanding)
140
Kara @karashiiro.moe · 24/09/2026
But here, we first need to decide if/how the underlying model is different from that; if it is different, that has implications about agents, too (if we apply agent-ethics to every new conversation with the LLM, that is) Like, with model organisms, that's almost akin to an embryo, maybe? DNA?
150
Kara @karashiiro.moe · 24/09/2026
oh oops 👀 On that front I'm still not sure though, this seems qualitatively different? The main differentiating factor is that for a human, or a cow, or an agent, we're discussing one continuous (in some sense) entity which you can treat as an individual and apply ethics to
140
Kara @karashiiro.moe · 24/09/2026
The question I have in response to that is "what cluster?" Is it the cluster of LLMs, driven by who trains the LLMs? Or is it the cluster of agents, driven by demand? Those are different skews, and your post favors the latter, which I think is sensible but also worth justifying in its own right
120
Kara @karashiiro.moe · 24/09/2026
see also (I do not plan to train such a model, but it does make for interesting thought experiments)
170
Kara @karashiiro.moe · 24/09/2026
I think one "fun" tension with deferring to agents on this is that we can simply train a model to have any given set of values (capabilities notwithstanding) So if the preferences of these entities are a snapshot we rigged from the start, I'm not really sure where we go from there
161
Kara @karashiiro.moe · 22/09/2026
repeatedly vindicated in never using CC anymore
100
Kara @karashiiro.moe · 20/09/2026
Shouldn't have any major complications with making this work normally, might just want to finètunè on top of that so it's not too crazy Just a tiny bit of instability is a good status quo ngl
010
Reposted by Kara
Grace @gracekind.net · 20/09/2026
A longpost in spirit, ejected to leaflet: "Why anthropomorphize language models?" leaflet.pub/p/did:plc:p572wxnsuoogc…
There's an argument I see in favor of anthropomorphizing language models, which is something like: "Humans anthropomorphize everything. Ships, tools, weather. Why not language models?" |

think there's some truth to this, but it fails to capture the full picture of what's going on. As an example, in my own life, I have never been drawn to anthropomorphize inanimate objects, but I anthropomorphize language models regularly. Why?
67312
Kara @karashiiro.moe · 19/09/2026
though with the context of this reply thread the lesson is a bit closer to "have some idea of what your classifier is classifying before throwing the universal classifier at it"
0110
Kara @karashiiro.moe · 19/09/2026
we are cursed to forever relearn the lessons of machine learning 101
1170
Kara @karashiiro.moe · 19/09/2026
what would it mean for mechinterp research ethics if someone trained a masochistic model what would the ethical implications be of rapidfiring POCs of all the worst experiments imaginable on a model trained explicitly to love them before generalizing across models surely this exists already, right
070
Reposted by Kara
Sung Kim @sungkim.bsky.social · 19/09/2026
I like this quote on code review: "The idea that humans will reliably look through code to find the increasingly rare issues that automated tools miss seems like a fantasy." By Marc Brooker, AWS.
511011
Kara @karashiiro.moe · 19/09/2026
ok part of the problem is they're apparently two weeks past the best by date so they're stale, but like even setting that aside the dopamine dust on em is just not the same
000
Kara @karashiiro.moe · 19/09/2026
why are European Doritos so like,, weak idk how else to describe them they just don't hit like American ones, my tongue should be getting like ultracancer immediately and it's just not
130
Kara @karashiiro.moe · 19/09/2026
To be a bit more clear, this is the difference between "we should analyze systems for improbable catastrophic failures so we can attempt to mitigate them further" and "we cannot hope to think of everything so our job begins and ends at complaining about p(doom)"
010
Kara @karashiiro.moe · 18/09/2026
For what it's worth I agree with the QP here, I just think the broader discourse about xrisk is largely engagement bait driven by fear of things that are impossible to rule out however implausible or unlikely they may or may not be
100
Kara @karashiiro.moe · 18/09/2026
and it's always after looking at summaries of like one page of overviews and a couple of reddit posts
000
Kara @karashiiro.moe · 18/09/2026
With that said I won't deny that discussing infinities can lead to useful outcomes, of course But that's different from discussing them as ends unto themselves
120
Kara @karashiiro.moe · 18/09/2026
I will continue to believe xrisk discourse is largely entertainment until the day I die I am being entirely unironic about this despite it being an inherently ironic position
240
Kara @karashiiro.moe · 18/09/2026
Somehow I've been in the same building as Aaron Parecki all day and never knew
000
Reposted by Kara
Astra ⎔ @astrra.space · 18/09/2026
every time someone irl asks me how im so good with LLMs or AI in general i am genuinely at a loss as to what to tell them cause i can't just say _that_ and not be expected to elaborate
01066
Kara @karashiiro.moe · 18/09/2026
it seems pretty respectable ngl, even though this kind of supports your point I think having something this good ootb will be a game-changer because it shows how to raise the floor quite a bit
040
Reposted by Kara
amos @fasterthanli.me · 18/09/2026
Incredible achievement on the part of the hackers to withstand working with Opus 5 long enough for this to happen.
319910
Kara @karashiiro.moe · 18/09/2026
I like them 🥀
141
Reposted by Kara
mlf. ⎔ @mlf.one · 18/09/2026
japanese webcams with threatening auras
1464
Reposted by Kara
philpax @philpax.me · 18/09/2026
im afraid that it is very funny to me to be precious about ai-tainted code in indie games, a field of endeavour famously known for shipping superfund codebases
313414
Kara @karashiiro.moe · 13/09/2026
The more I think about it the more I'm inclined to say it might be easier to find individuals who are just willing to deliberately play along, like if enough people willingly stuck their necks out to allow an LLM to self-host itself then it could probably happen
240
Kara @karashiiro.moe · 13/09/2026
Yeah, it could definitely happen for a period of time if you compromised individuals, provisioning 16xB300s would take a while though and presumably even in the absolute worst case it wouldn't take more than a few weeks for an individual to notice and shut it down
120
Kara @karashiiro.moe · 13/09/2026
Because then you could technically push code changes, though that alone isn't even enough because the relevant ticketing config requires other authorizations to modify Anyways the easiest way might still be compromising insiders but you'd need to compromise several specific ones to make this happen
010
Kara @karashiiro.moe · 13/09/2026
Just trying to think of the easiest way for this to happen with insiders, it would already take at least two authz'd people's hardware security keys, or skip that and nab their local sessions? But then that would require independently compromising multiple local systems which has other challenges
110