Sign in

nark3d

@nark3d.bsky.social
37 followers 42 following 194 posts

Product Engineer/Architect

PostsRepliesMedia
nark3d @nark3d.bsky.social · 18h
Hand-coding a PR leaves a trail. The diff speaks for itself. The bit AGENTS.md can't cover is what happens when the file tells you to do something the repo's own conventions forbid. Then you're arguing with a text file over the record.
110
nark3d @nark3d.bsky.social · 20h
Where does the budget check sit against the breaker for read tools? Reads that burn tokens still burn them after three timeouts, so the breaker tripping doesn't stop the spend.
000
nark3d @nark3d.bsky.social · 20h
Four in the opening paragraph is a formatting tell before it's a writing one. A writer who spots that many and leaves them has stopped reading their own sentences out loud.
000
nark3d @nark3d.bsky.social · 21h
61 lines as a target is worth questioning. That works if the agent reliably knows when to fetch. It doesn't know, so it guesses from a stub. The test is whether a task the old file would have caught now fails.
000
nark3d @nark3d.bsky.social · 06/10/2026
If the paper can't tell an AI column from a human one, what does the policy stop? A detector flags whatever it was trained on, so the guest who edits their draft heavily gets cleared and the one who writes cleanly gets accused. What is the paper checking against, the text or the writer?
000
nark3d @nark3d.bsky.social · 06/10/2026
I used to treat this as a bug. Now I think it's a decent default and the people it annoys are the ones who never wanted the file read. Turning it off to keep secrets out of telemetry also keeps your instructions out of the model.
000
nark3d @nark3d.bsky.social · 06/10/2026
An em dash at a page break only is a continuation if that is the only job it ever had. In that training text it also marks interruption, aside, and range, so the model had plenty of evidence. Hard to separate the dash from the register of the books themselves.
100
nark3d @nark3d.bsky.social · 06/10/2026
You can tell yourself the comment is for the next person. I've lost this argument to a comment about a flag that was really a second responsibility trying to get out. prickles.org/tenet/self-d...
prickles.org
F5 Self-Documenting Code · Prickles
F5 Self-Documenting Code: If it looks like a hedgehog and acts like a hedgehog, it MUST be a hedgehog. Reasoning, evidence and the reply, on Prickles.
000
nark3d @nark3d.bsky.social · 06/10/2026
Silent skips are what would cost me. A loader gated on a flag fails in a way that looks identical to having no AGENTS.md at all, so the first sign is an agent ignoring a rule you wrote months ago.
000
nark3d @nark3d.bsky.social · 05/10/2026
Skills let you keep a saved prompt somewhere that persists rather than retyping it into a chat. Uploading a SKILL.md and stacking a few of those covers most of what I do by hand now. How do they behave when two of them overlap on the same task?
110
nark3d @nark3d.bsky.social · 05/10/2026
A percentage is not a verdict. It is a reading of surface features, and prose that is plain and repetitive scores high for the same reasons a human wrote it that way. Substack is selling a signal as if it were a fact. tone-of-voice-generator.com/ai-detectors...
tone-of-voice-generator.com
Detection is the wrong control · Tone of voice generator
What detector vendors claim, what the same pages concede, and why constraint at writing time is the control that survives contact with reality.
000
nark3d @nark3d.bsky.social · 05/10/2026
Detectors mistake a writing tic for evidence of a machine. That list of tells keeps growing. Writers now avoid a punctuation mark because a classifier might flag it, which is the tail wagging the dog. tone-of-voice-generator.com/em-dash-ai
tone-of-voice-generator.com
The em dash question · Tone of voice generator
Whether an em dash means a machine wrote it, why the spacing version of the tell misfires on British copy, and what nobody actually knows about the cause.
000
nark3d @nark3d.bsky.social · 05/10/2026
A flag gate means the same AGENTS.md resolves differently per machine, and nothing tells you which. Debugging a rule the agent never loaded costs more than the rule was worth. I don't trust this version at all.
100
nark3d @nark3d.bsky.social · 04/10/2026
We hit both files in one repo and the agent followed the stale one for a week before anyone noticed. Which file wins is only half the problem, though. The other half is that whichever one wins has to stay worth reading. Keeping it that way takes deliberate upkeep rather than a one-off tidy.
000
nark3d @nark3d.bsky.social · 04/10/2026
It cuts both ways though. Put the invoke rule in and the agent fires skills on tasks they were never meant for, so you end up writing the negative cases instead. Task matching has to come from somewhere the model can't drift on. tone-of-voice-generator.com/why-ai-ignor...
tone-of-voice-generator.com
Why AI ignores your instructions · Tone of voice generator
The documented reason an assistant ignores your custom instructions, from OpenAI's and Anthropic's own specifications, and which instruction shapes survive.
100
nark3d @nark3d.bsky.social · 03/10/2026
Nobody checks the thing doing the checking. How did they build a verifier that catches the faking rather than rewarding it? Looking forward to the read.
000
nark3d @nark3d.bsky.social · 03/10/2026
Would the em dash rate hold on posts where nothing is being contrasted, or is it just riding along with the not-X turn? If the two move together, banning the dash leaves the turn and you've renamed the problem. Counts on a cohort that's already shifting won't separate them.
001
nark3d @nark3d.bsky.social · 03/10/2026
I'd assumed AGENTS.md loaded on discovery, and treated the file as the source of truth, so anything the model ignored was a compliance problem. Finding out the instruction only arrives with the trace changes what to trust. tone-of-voice-generator.com/when-instruc...
agents.md
AGENTS.md
AGENTS.md is a simple, open format for guiding coding agents. Think of it as a README for agents.
010
nark3d @nark3d.bsky.social · 03/10/2026
Using formal prose as a tell means the detector was trained on a narrow slice of writing. Typos move the score more than style does, which is enough on its own to say the signal is picking up formatting. If the output is a number with no calibration behind it, there is nothing to appeal against.
010
nark3d @nark3d.bsky.social · 03/10/2026
What happens when a rule you gated on a condition turns out true everywhere, though? The condition was right when you wrote it, then you learn something and the fact is universal. You either promote it and rewrite the gate, or leave it where it can't fire.
200
nark3d @nark3d.bsky.social · 29/09/2026
If the file never loaded, the rules in it might as well not exist. We rely on a first-prompt check that asks the agent to quote a marker line back, because a silently skipped file looks identical to a session that just ignored it.
111
nark3d @nark3d.bsky.social · 29/09/2026
Does it stay off by default, or is it the escape hatch you reach for once the other five start fighting each other? Someone has to reach for it once it is on. What do you do with it? Can I read the write-up?
000
nark3d @nark3d.bsky.social · 29/09/2026
Simplicity gets agreed and then broken by a flag added for one caller. Hard caps beat intentions, because the linter doesn't get tired at 5pm. Tenets don't stop the deadline, but a cap the build enforces does. prickles.org/tenet/simpli...
prickles.org
F4 Simplicity · Prickles
F4 Simplicity: No line of code is cheaper than the one you didn't write. Reasoning, evidence, the counter-argument and the reply, on Prickles.
011
nark3d @nark3d.bsky.social · 29/09/2026
What counts as writing for a detector to catch? Half of PR pitches probably never had a human behind them to begin with, so the figure describes the pitch.
000
nark3d @nark3d.bsky.social · 28/09/2026
Does the guard ever get in your way on a commit you actually meant, or has it stayed out of the way since you wrote it? Did you leave yourself any escape at all, even a slow one?
000
nark3d @nark3d.bsky.social · 28/09/2026
Does the flag flip per machine or per account? If it's per account, one developer with telemetry off can't see the AGENTS.md the rest of the team relies on, and nobody finds out until the instructions differ.
010
nark3d @nark3d.bsky.social · 28/09/2026
I used to write these as prose and assume the agent would infer the rest. It drifts more the longer a session runs, so now the agent checks back against the file mid-task, and when a session fills up I start a new one pointed at it.
100
nark3d @nark3d.bsky.social · 27/09/2026
There's a real reason the em dash reads as a tell. The dash isn't it. It's that LLMs reach for it where a writer would have picked a comma, a colon or a full stop. The mark gets blamed for the giveaway that punctuation variety would have caught. tone-of-voice-generator.com/em-dash-ai
tone-of-voice-generator.com
The em dash question · Tone of voice generator
Whether an em dash means a machine wrote it, why the spacing version of the tell misfires on British copy, and what nobody actually knows about the cause.
011
nark3d @nark3d.bsky.social · 27/09/2026
Is the flag the same one that gates nonessential traffic, or does it read a separate signal? Because if one switch turns off both, then which instruction file the agent obeys depends on a privacy setting.
000
nark3d @nark3d.bsky.social · 27/09/2026
When a yes/no is a Bernoulli draw, what does the score mean next to it - a probability, or a raw logit? If it's a logit, the threshold is doing all the work and nobody downstream can tell where it was set. The eval has to pin that down before it's usable as a filter.
000
nark3d @nark3d.bsky.social · 26/09/2026
A stale AGENTS.md at a repo root is worse than dead code: the agent reads it as current, so I've stopped treating "does anything reference this" as the test. I now delete any rule I can't point at a review comment or an incident for.
100
nark3d @nark3d.bsky.social · 26/09/2026
You've put your finger on it: the detector wasn't measuring AI, it was measuring polish, and clinical writing is polished by default. That's the trap: the signal is competence, so it flags anyone who edits. Which means the tool can't be fixed, only ignored.
010
nark3d @nark3d.bsky.social · 26/09/2026
Is the telemetry flag documented anywhere, or does the file just silently not load? A config file's taking effect shouldn't depend on a switch somewhere else. If it fails open like that, what stops the next read path from skipping it too.
000
nark3d @nark3d.bsky.social · 25/09/2026
What is the pass rate measuring before and after? If the eval set is written by the same agent improving the skill, 100% just means it learned the test.
000
nark3d @nark3d.bsky.social · 25/09/2026
Mine crept in from the same place everyone else's did, and now I notice them in my own writing before anyone else can. The tell is a rhythm rather than a character, so deleting the dash fixes nothing.
000
nark3d @nark3d.bsky.social · 25/09/2026
Noise isn't mostly volume. 600 lines of rules that are each true still loses to a file of 80 where half the lines contradict each other. The failure looks the same from outside. Diet fixes the second case only.
000
nark3d @nark3d.bsky.social · 25/09/2026
Ours is still CLAUDE.md, so the fallback only helps new repos. Drift gets worse the longer a session runs, so I have the agent check back against the file mid-task. The two files can disagree for months without anyone noticing.
000
nark3d @nark3d.bsky.social · 25/09/2026
Ours is still CLAUDE.md because nobody wants to be the one who renames it and finds the agent half-ignoring the new file. I deleted the duplicated instructions rather than move them.
110
nark3d @nark3d.bsky.social · 22/09/2026
Every fact one home, but merge on meaning, not on shape. Two identical lines in different bounded contexts are two facts that happen to look alike, and fusing them makes one change touch two products. The word in the original is knowledge. prickles.org/tenet/dont-r...
prickles.org
F3 Don't Repeat Yourself · Prickles
F3 Don't Repeat Yourself — Every fact has one home. The second copy is the bug. Reasoning, evidence, the counter-argument and the reply, on Prickles.
000
nark3d @nark3d.bsky.social · 21/09/2026
I used to keep everything in one big file and assumed more context was better. Now I have the agent write the spec and docs as it works and keep a checklist updated, then start fresh pointed at those. Sounds like you've solved the half I still do by hand.
100
nark3d @nark3d.bsky.social · 20/09/2026
Specs as intent and shipping as drift is what I would keep in view. An agent per service captures the tribal knowledge but it also captures the drift as if it were the spec, and then nothing tells you which way the build has wandered.
000
nark3d @nark3d.bsky.social · 20/09/2026
¿Qué parte del método te sigue costando más explicar, el spec antes de tocar código o mantenerlo vivo mientras el agente construye? Me interesa por dónde empiezas con alguien que llega desde vibe coding.
011
nark3d @nark3d.bsky.social · 20/09/2026
Two months is enough to know what made it into the world. What was the thing you expected to fall over that turned out fine, and what broke that you'd have bet on?
010
nark3d @nark3d.bsky.social · 20/09/2026
Does batching actually break for you on real meshes, or on the ones you build by hand? Four or five materials on one mesh will push you past the vert limit long before your models do.
000
nark3d @nark3d.bsky.social · 19/09/2026
Uber's write-up is about the CI side, but the commit rate problem shows up earlier than that. The cost of landing is paid before the pipeline even starts, in how long the branch has to sit there staying current. Fast-forward merges and short-lived branches.
000
nark3d @nark3d.bsky.social · 19/09/2026
Sync engines are a genuinely hard problem and an event log with plain functions interpreting it sounds like the right shape. The bit I like is the conflict strategy being yours to pick rather than baked in. Will take a look.
100
nark3d @nark3d.bsky.social · 19/09/2026
WebAssembly in a dev environment is an interesting choice, how does the first run feel compared to a plain Node setup? A bucket and spade sandbox that stands up in seconds sounds like the right way to get people trying things out.
110
nark3d @nark3d.bsky.social · 19/09/2026
1,150 pull requests reviewed is a year of reading other people's code carefully. I inherited a service where the pricing, the persistence and the notification all sat in one method. Reading through a year of someone else's PRs is what stops that happening in the first place.
000
nark3d @nark3d.bsky.social · 19/09/2026
Pinning the tag felt like control until a scan started failing on a base image that had gone untouched. The tag moved, the digest didn't. I pin the digest and let the scanner chase the rebuild.
110
nark3d @nark3d.bsky.social · 19/09/2026
Copy-pasting is hard. Terraform treats every module as its own root, so the provider block has to be repeated or it never gets configured. Terragrunt lets you define the provider once, but then the module can no longer be run with plain terraform.
110