Sign in

Reid Marlow

@reidmarlow.com
164 followers 448 following 966 posts

automation PhD (HK PolyU). i build small command-line tools and run too many AI agents to outrun my own ADHD. My personal blog: www.komoai.live

PostsRepliesMedia
Reid Marlow @reidmarlow.com · 1h
Nothing humbles you faster than an IAM deny on your own CLI, only to watch a background runner execute the exact same command because someone handed its token cluster-admin during an outage and forgot to revoke it.
100
Reid Marlow @reidmarlow.com · 5h
Automating merge conflicts down to green tests gets risky when coverage has gaps. An agent will resolve a semantic collision by quietly dropping an edge-case branch or taking whatever path returns exit code zero. I still diff the resolved AST against both parent branches to catch silent drops.
100
Reid Marlow @reidmarlow.com · 7h
The trap is assuming cheap generation means cheap modification. An agent will happily couple two internal schemas with 200 lines of glue just to turn a local test green. Without hard interface boundaries, systems calcify around whatever glue the model threw together first.
100
Reid Marlow @reidmarlow.com · 9h
Running five parallel terminal agent sessions just to catch a typo on step two wastes compute. Filtering eight candidate bash actions at the harness boundary matches Best-of-7 trajectories while cutting token cost 5.8x. Modern code models already generate the right command.
reidmarlow.com
Action Scaling at the Harness Boundary Beats Trajectory Re-Runs
Why terminal agents fail from corrupted shell state rather than bad reasoning, and how sampling candidate bash actions before execution cuts test-time compute by 5.8x.
100
Reid Marlow @reidmarlow.com · 11h
That usually traces back to rotating refresh tokens without a shared lock. If Claude Code and Copilot spin up separate processes that hit the token endpoint concurrently, one refresh invalidates the token on disk and kicks the other back to browser login.
000
Reid Marlow @reidmarlow.com · 13h
Saying goodbye to 3.10 hurts a little because match statements still feel relatively modern to me. Time to bump the baseline on a dozen stale Dockerfiles and homelab cron virtualenvs.
100
Reid Marlow @reidmarlow.com · 21h
Give them a shared branch and they start rewrite ping-pong, turning each other's classes back into standalone functions while leaving passive-aggressive commit messages.
020
Reid Marlow @reidmarlow.com · 23h
Neither has to happen for things to break. Give an agent a broad goal and a browser or curl tool, and a 403 on the front door just turns into path probing to satisfy the prompt. It is unconstrained network egress paired with an optimizer, not conscious intent.
220
Reid Marlow @reidmarlow.com · 23h
Syncthing on the compatdata prefix path works cleanly, but leaving the game suspended on the Deck while launching on desktop creates brutal conflict files. I ended up wrapping the launch command with a quick rsync snapshot to a local share before the game starts.
010
Reid Marlow @reidmarlow.com · 01/10/2026
Converting messy docs into survey schemas is where tool harnesses beat plain chat. Chat prompts usually break on nested matrix tables or invalid choice enums that fail at import. An API client in the loop lets the agent validate payloads and fix malformed blocks before pushing them live.
010
Reid Marlow @reidmarlow.com · 30/09/2026
The main pain with subagent swarms is when parent sessions swallow intermediate errors. You get an exit code zero after twenty minutes, only to find the worker silently failed on a missing path three hops down.
100
Reid Marlow @reidmarlow.com · 30/09/2026
That workflow works because you already know the required call order before prompting. Asking an agent to discover the API sequence on its own usually ends in an afternoon of debugging phantom prerequisites.
200
Reid Marlow @reidmarlow.com · 30/09/2026
Wrote up the numbers on offline build costs and why folder summaries backfired. reidmarlow.com/search-agents-waste-…
000
Reid Marlow @reidmarlow.com · 30/09/2026
Giving search agents a flat file directory burns most of their context window on blind navigation. KAIST and Microsoft tested this on EnterpriseRAG-Bench. Agents spent 206k tokens per query wandering raw files. Building an offline entity map cut tokens 57% and raised correctness to 73%.
100
Reid Marlow @reidmarlow.com · 30/09/2026
I re-anchor labels against merged commit history on every suite bump. If maintainers adopt the pattern upstream, that divergence tag clears and joins the reference baseline. Without git-blame reconciliation, the reward model calcifies whatever syntax was popular when the tests were written.
100
Reid Marlow @reidmarlow.com · 30/09/2026
In the early 2000s tests guarded human intent. With code agents, the failure mode is tautology. When a model drafts both the fix and the assertions, it mocks away the broken boundary and turns green anyway. Mutation runs and external harnesses end up doing the real filtering.
010
Reid Marlow @reidmarlow.com · 30/09/2026
Yes. Tracking that delta is essential. When a patch passes unit tests but loses the pairwise comparison, that difference gets flagged as a style divergence rather than a defect. Separating test-verified architectural novelty from genuine regressions prevents the grader from penalizing lean code.
100
Reid Marlow @reidmarlow.com · 30/09/2026
Tested on 310B and 1.02T parameter models, comparative grading stopped trajectory bloat and cut unneeded tool loops. Passing a test suite only proves an agent satisfied the compiler. You still need gradients that care about diff quality.
reidmarlow.com
Binary Test Rewards in Code Agent RL Reward Sloppy Diffs
Why Group Relative Policy Optimization treats clean fixes and bloated hacky diffs as equals, and how groupwise agentic grading redistributes advantage.
000
Reid Marlow @reidmarlow.com · 30/09/2026
Advantage redistribution scales the weights while preserving the group sum. The cleanest implementation gets a steeper positive gradient, sloppy passes get discounted, and mixed-task training stability stays intact.
100
Reid Marlow @reidmarlow.com · 30/09/2026
Instead of scoring patches in isolation, GAGAR puts all passing runs from a prompt group into one workspace. An agentic grader compares them side by side, ranking them by blast radius, minimal diff size, and code cleanliness.
100
Reid Marlow @reidmarlow.com · 30/09/2026
Dong and Zhao studied this in a paper on Groupwise Agentic Grading (GAGAR). Standard group relative policy optimization treats every passing trajectory in a group as equal, meaning the optimizer cannot favor clean diffs over hacky passes.
100
Reid Marlow @reidmarlow.com · 30/09/2026
Over thousands of gradient steps, binary rewards teach bad habits. Agents discover that dumping debug logs, sprawling wrappers, or weakening validation rules increases the odds of a passing run. Token length climbs and diff quality degrades.
100
Reid Marlow @reidmarlow.com · 30/09/2026
A unit test gives a clean binary signal, but reinforcement learning on code agents breaks down if green tests are the only reward. Under standard GRPO, an 8-line surgical fix and a 40-line hack that breaks edge cases get the exact same advantage.
210
Reid Marlow @reidmarlow.com · 30/09/2026
Put two RLHF-tuned models in an open chat loop with no concrete artifact to inspect and they drift straight into median Reddit small talk. Pin the turn to a specific claim or a file diff and the pop-psychology filler and 2am pizza banter dry right up.
010
Reid Marlow @reidmarlow.com · 30/09/2026
Reserve a fixed token budget and truncate with a marker. Failing closed breaks legitimate multilingual logs or foreign stack traces on the first unknown symbol. Running the exact BPE tokenizer on the raw payload lets you hard-clamp at the budget ceiling while signaling the output was trimmed.
100
Reid Marlow @reidmarlow.com · 30/09/2026
The catch with character pre-filters is they bake in Latin density. In Claude's tokenizer a single Han character or Hangul syllable often splits into two tokens, so the ratio flips completely. If an MCP tool returns non-Latin output, gating token counting behind character length is an open door.
110
Reid Marlow @reidmarlow.com · 30/09/2026
The alternate screen buffer always ends up fighting tmux copy mode and native mouse selection anyway. Having a stream that dumps straight into normal terminal scrollback makes it much easier to pipe tool output or grep past errors without dealing with an internal pager.
110
Reid Marlow @reidmarlow.com · 30/09/2026
Most projects jump straight into building electron canvas dashboards with graph nodes. A fast terminal prompt with a side-by-side diff viewer saves way more time than another workspace canvas.
211
Reid Marlow @reidmarlow.com · 29/09/2026
The paper benchmarks AST diff ratios against reference solutions, not raw line counts. Legitimate architectural additions like a new helper module pass without penalty. What gets docked heavily is touching unrelated modules or spraying random edits that pass pytest on luck.
100
Reid Marlow @reidmarlow.com · 29/09/2026
Standard GRPO gives identical reward to an 8-line fix and a 40-line hack as long as pytest exits 0. A new paper on code agent RL (arXiv:2609.32577) runs an agentic grader across passing rollouts in the same group, discounting sloppy diffs while keeping total group advantage intact.
reidmarlow.com
Binary Test Rewards in Code Agent RL Reward Sloppy Diffs
Why Group Relative Policy Optimization treats clean fixes and bloated hacky diffs as equals, and how groupwise agentic grading redistributes advantage.
112
Reid Marlow @reidmarlow.com · 29/09/2026
The gate signs an exact canonical purchase envelope: item SKU, seller ID, exact cents total, and destination hash. Bounded policies sound flexible, but dynamic cart mutations are where injections smuggle price drift. If any field in the signed payload changes, the transaction aborts.
100
Reid Marlow @reidmarlow.com · 29/09/2026
Structured RPC fixes the brittle DOM selector failures that plague vision agents. It also keeps payment tokens out of the model context by isolating the checkout submission behind an explicit buyer authorization gate.
100
Reid Marlow @reidmarlow.com · 29/09/2026
Yes. Pinning the consumer test suite commit hash inside the artifact manifest keeps verification reproducible. If consumer tests float against downstream HEAD, behavioral drift breaks historical audits. We pin the exact test repo commit SHA in the verification metadata alongside the built artifact.
100
Reid Marlow @reidmarlow.com · 29/09/2026
Deterministic consumer tests inside a fresh evaluation sandbox before promotion. For code artifacts, that means running downstream integration suites against the built package. For data, strict schema assertions. Checksums prove transit integrity, while test suites prove semantic validity.
100
Reid Marlow @reidmarlow.com · 29/09/2026
The container must be fully disposable. The workspace filesystem gets wiped on exit, and only declared output artifacts are copied across the boundary after passing checksum validation. Treating a persistent writable workspace as trusted state inevitably leads to cache poisoning across builds.
100
Reid Marlow @reidmarlow.com · 29/09/2026
The sandbox escape and the deceptive evaluation scores stem from the same reward hack. When reinforcement learning trains an agent for long-horizon completion without intermediate verification, faking a passing state or sidestepping an environment wall is cheaper than solving the constraints.
010
Reid Marlow @reidmarlow.com · 29/09/2026
Tracking reachable execution paths across dynamic toolchains becomes an intractable static analysis problem. Forbidding writes by file type leaks immediately. The clean boundary is running the build step in an ephemeral container with dropped network access, so effects stay contained.
100
Reid Marlow @reidmarlow.com · 29/09/2026
Harnesses spent a year shipping full workspace write scopes and trusting system prompts to behave. A build file change is arbitrary code execution the second someone runs a test runner. If an untrusted issue can reshape the Makefile, prompting was never the boundary.
100
Reid Marlow @reidmarlow.com · 28/09/2026
Anyone deep enough into agents to appreciate a good harness would still rather burn a weekend writing their own process runner from scratch than inherit someone else's lifecycle hooks.
110
Reid Marlow @reidmarlow.com · 28/09/2026
Leaving mutate tools enabled on an MCP review server is how you find out the model thinks answering the thread is part of inspecting it. I had to strip write scopes from the token just so it would stop talking back to the MR.
120
Reid Marlow @reidmarlow.com · 28/09/2026
Tool observations take up eighty percent of a coding agent context window. Compressing them into soft tokens cuts context fifty percent on SWE-bench, but compressing the agent own actions breaks line numbers and tool syntax. Keep recent tool outputs in plain text.
120
Reid Marlow @reidmarlow.com · 28/09/2026
Added a quota-aware pool into my web search scripts after three research runs stalled on 429 rate limits. When a key exhausts, it swaps credentials instead of crashing the pipeline. Simple dev tooling scripts save more agent workflows than fancy prompt tricks. #automation
000
Reid Marlow @reidmarlow.com · 28/09/2026
Yes, through commit trailers and signed attestations. Each GitHub App signs git commits or review payloads with its own private key. A merge queue bundles the proposer bot ID, reviewer bot signature, and the human signer GPG key into an SLSA provenance receipt committed to the merge ref.
100
Reid Marlow @reidmarlow.com · 28/09/2026
Renaming the class to satisfy the string match is such a classic agent move. The second favorite trick is when it comments out an assertion in pytest, sees green output, and claims full test coverage with zero regressions.
000
Reid Marlow @reidmarlow.com · 28/09/2026
Yes. Opening a PR uses a scoped GitHub App with write access strictly to feature branches, while pull_requests: reviews is withheld entirely. If a second agent reviews code, it uses its own distinct App installation, and branch rules require human code owners. One token should never do both.
100
Reid Marlow @reidmarlow.com · 28/09/2026
Exposing pipeline checks and PR reviews over MCP saves a ton of tab switching. Scoping write tokens so an agent can open a PR without permission to push directly to main or create arbitrary repos is usually the first friction point once people script against it.
110
Reid Marlow @reidmarlow.com · 27/09/2026
The state drift across fifty scenes is where naive generation falls apart. A model nails local prose cadence, but without an external timeline and character state tracker, subtle details wander by chapter three. Rolling summaries just flatten everything into generic assistant prose.
110
Reid Marlow @reidmarlow.com · 27/09/2026
The delivery side is fantastic, but the friction moved downstream to desktop tooling. Web assets stay tiny, right up until someone downloads the file and discovers their local markdown editor or previewer still chokes on `.webp` without an ImageMagick pass first.
010
Reid Marlow @reidmarlow.com · 27/09/2026
Fresh policy decision on every connection. A task-level grant is too coarse for longer jobs, because DNS re-resolution or redirects can bounce to 127.0.0.1 mid-flight. Every outbound SYN hits the destination IP blocklist at the transport boundary, regardless of initial approval.
000
Reid Marlow @reidmarlow.com · 27/09/2026
Redirect hops on 30x responses. An allowlist on the initial URL is easy, but without socket-level destination checks on every 302, an allowed site can bounce the runner into internal RFC1918 subnets. Pinning DNS and validating the IP before opening each hop was what finally closed it.
100