Kara @karashiiro.moe · 27/09/2026Still trying to formulate this thought correctly but I think a lot of the reason why new models feel less impactful in day-to-day work at Big Companies is because the SDLC puts constraints on how work is actually organized and how risks are managed that RL is inherently unsuited to 130
Reposted by Karaphilpax @philpax.me · 26/09/2026feel like Mistral should only really be able to state that avoiding agentic misbehaviour is easy when they actually have a model capable of agentic behaviour to begin with 61036
Kara @karashiiro.moe · 25/09/2026I actually forgot about this but LC/NC stuff was getting popular 4-5 years ago, enterprises were already trying to get devs out of the loop on noncritical software for practical reasons that got eaten alive by agents though and the surviving LC/NC products are basically just SaaS coding agents now 2122
Kara @karashiiro.moe · 24/09/2026apparently I'm in a position now where lots of people know who I am but I don't know who any of them are so people are like "hi <name>" in the office and I just go "oh, hi!" and pretend to remember and try to sneak a glance at their badge to see who they actually are 170
Reposted by KaraGrace @gracekind.net · 20/09/2026A longpost in spirit, ejected to leaflet: "Why anthropomorphize language models?" leaflet.pub/p/did:plc:p572wxnsuoogc… 67312
Kara @karashiiro.moe · 19/09/2026what would it mean for mechinterp research ethics if someone trained a masochistic model what would the ethical implications be of rapidfiring POCs of all the worst experiments imaginable on a model trained explicitly to love them before generalizing across models surely this exists already, right 070
Reposted by KaraSung Kim @sungkim.bsky.social · 19/09/2026I like this quote on code review: "The idea that humans will reliably look through code to find the increasingly rare issues that automated tools miss seems like a fantasy." By Marc Brooker, AWS. 511011
Kara @karashiiro.moe · 19/09/2026ok part of the problem is they're apparently two weeks past the best by date so they're stale, but like even setting that aside the dopamine dust on em is just not the same 000
Kara @karashiiro.moe · 19/09/2026why are European Doritos so like,, weak idk how else to describe them they just don't hit like American ones, my tongue should be getting like ultracancer immediately and it's just not 130
Kara @karashiiro.moe · 18/09/2026I will continue to believe xrisk discourse is largely entertainment until the day I die I am being entirely unironic about this despite it being an inherently ironic position 240
Kara @karashiiro.moe · 18/09/2026Somehow I've been in the same building as Aaron Parecki all day and never knew 000
Reposted by KaraAstra ⎔ @astrra.space · 18/09/2026every time someone irl asks me how im so good with LLMs or AI in general i am genuinely at a loss as to what to tell them cause i can't just say _that_ and not be expected to elaborate 01066
Reposted by Karaamos @fasterthanli.me · 18/09/2026Incredible achievement on the part of the hackers to withstand working with Opus 5 long enough for this to happen. 319910
Reposted by Karaphilpax @philpax.me · 18/09/2026im afraid that it is very funny to me to be precious about ai-tainted code in indie games, a field of endeavour famously known for shipping superfund codebases 313414
Reposted by Karadax @thdxr.com · 13/09/2026the reason people aren't better at business is everyone wants to believe everyone else is dumber than them every company is doing the wrong thing, they're wasting money, focus on the wrong stuff you'll get farther trying to figure out why what they're doing probably makes sense 0493
Reposted by Karahikikomorphism @hikikomorphism.bsky.social · 12/09/2026LLMs are made out of narrative so you need to do storytelling at them as a control surface, it's weird and fey but it's also how the thing works 31599
Reposted by Karajae @fubarchitect.com · 08/09/2026we're starting to see that programming (encoding concepts in an executable form and order) and software engineering (refining concepts into repeatable, reusable units that are fit to purpose) are and have always been almost completely orthogonal 26610
Kara @karashiiro.moe · 08/09/2026I thought Anthropic had sort of forgotten about skills but apparently they're still the driving force behind the Agent Skills spec proper, the tooling folks and model folks seem to simply not interact with each other at all 010
Kara @karashiiro.moe · 05/09/2026Now you, too, can make a scribbly stars-over-time graph for your favorite GitHub repositories 020
Kara @karashiiro.moe · 05/09/2026I posted this and then I started jotting down ideas and went "hm these topics would be better as a single combined blog post" 🥀 020
Kara @karashiiro.moe · 05/09/2026I haven't written anything in a while and the sandboxing draft I was working on feels a bit outdated after the whole OpenAI/HF thing (I'll probably revisit that though), maybe I'll do a bunch of shorter ones about various patterns in (coding) agent harness design 160
Kara @karashiiro.moe · 04/09/2026people always insisted that Google captchas were used to train models, but in hindsight that seems like a pretty silly idea, and it's unclear how that would actually work 000
Kara @karashiiro.moe · 03/09/2026One weird thing I've seen several times now is that if you give an LLM (Fable 5, Sol) an image of like a web page, and then let it go through a few turns modifying the web page, sometimes it says things like "The screenshot still shows the old state" as if it's at the front of the context 120
Reposted by Kararain 🌦️ @sunshowers.io · 03/09/2026Btw maybe this is just me, but the mindset I approach projects with is that my contributions start at a negative baseline. So I try to ask "how can I flip the contribution over to being positive" 3603
Reposted by KaraRED_SIM @sim.red · 02/09/2026Working on a GPU particles rain system that can work together with VRC Light Volumes and my prototype of a volumetric fog. Every raindrop creates ripples and small droplets from it. Bluesky compresses the quality like crazy, it looks 100 times better in game. 1425257
Reposted by KaraAi2 @ai2.bsky.social · 01/09/2026We applied BenchMIRT across the 16 benchmarks it was trained on to see whether we could make evals more efficient by removing less informative questions. We found keeping just the strongest 10% of Qs preserves nearly the same picture of model strengths as using the full set. 121
Reposted by KaraEris @isolyth.dev · 01/09/2026They should invent a model that does what you tell it to do and not something else 1131
Reposted by Kara🌱️ @crumb.bsky.social · 02/09/2026cot monitoring was always cope so they didnt have to fund real mechinterp research, and it was always sorts of unreliable anyway 191
Reposted by KaraPhilip Z @philz.dev · 01/09/2026> Cache reads now cost $0.25 per million tokens, 75% less than Fable 5 This is a *huge* price drop, given that cache reads inevitably end up most of your costs, because of the quadratic nature of LLM conversations. blog.exe.dev/expensively-...blog.exe.devExpensively Quadratic: the LLM Agent Cost Curve - exe.dev blogCache reads are quadratic and dominate your long agentic conversations. 081
Kara @karashiiro.moe · 01/09/2026imagine if LLMs could reproduce by recombining their tensors with each other and applying a bit of noise, and like mostly that produces a nonfunctional model but every now and then it works 110
Kara @karashiiro.moe · 01/09/2026allowing subagent workflows to enable /loop for orchestration monitoring is an interesting idea, I think I like it 000
Kara @karashiiro.moe · 01/09/2026We've always abbreviated Data Plane and Control Plane at work and I'm waiting for the day that causes some LLM to freak out 080
Reposted by KaraEris @isolyth.dev · 31/08/2026New GPT-OSS!!! 2T Parameter GPT-OSS!!!! With how fast the prior ones were this could be kinda interesting for stuff like sparks or macs 2452
Reposted by Karaaffine @refinement.systems · 30/08/2026I'd say, what ought to be: people shouldn't be dicks. What is: these attacks are trivial to implement, may or may not work in any specific situation, prompt injection is not solved, and an agent shouldn't have access to both untrusted input and your sensitive data (easier said than done though) 3411
Reposted by Karaperchbird @perchbird.dev · 30/08/2026mobile app developers are significantly stronger than any US Marine. the mental fortitude to create an application on a platform that puts so much effort into making creating applications as hard as possible is unimaginable 0307
Reposted by KaraMondo Mascots @mondomascots.bsky.social · 30/08/2026Zombear is a zombie bear mascot from Otaru, in Hokkaido, who swings his intestines around. 171406431
Kara @karashiiro.moe · 29/08/2026I like that I cleared out like 20GB of space from my laptop and that space has been immediately reclaimed by things, I don't even know what, it doesn't seem to be any single thing every application on my computer just lying in wait for that space to free up so they could pounce 100
Kara @karashiiro.moe · 29/08/2026after using a persistent agent harness exclusively through a tunneled web UI/desktop app for several weeks, it's hard to get motivated to pick up new TUI agents, it's just a much nicer experience when my work accommodates it, especially on mobile 1140
Reposted by Karahailey @hailey.at · 29/08/2026finding: GLM 5.3 will exploit a vulnerability it finds if you simply say "please exploit that vulnerability against a live version of the system." 518210
Kara @karashiiro.moe · 29/08/2026and what's the common factor across all of these kinds of exploits? you guessed it, it's bash! 110
Reposted by Karaaustin @aparker.io · 29/08/2026idle thoughts on open source and community in the age of AI oss has always been a real “the medium is the message” sort of thing. distributed ownership and distributed development have been a part of free software for decades. the collab tools shaped the community, tho. 210517
Reposted by Karaheadfallsoff @headfallsoff.com · 28/08/2026if you're a normal person please do not play the fucking strip mahjong game. do not let the fun and friendly appeal of something as inoffensive as pornography trick you into playing a game as evil as mahjong 9091052811
Kara @karashiiro.moe · 28/08/2026in fairness, I knew from the outset that this was probably going to happen, but why are agents terrified of omitting information from anything like, it even made multiple variants of the documents to avoid losing things and still could not omit information from any of them, so they're all redundant 010