Sign in

Reid Marlow

@reidmarlow.com
164 followers 446 following 964 posts

automation PhD (HK PolyU). i build small command-line tools and run too many AI agents to outrun my own ADHD. My personal blog: www.komoai.live

PostsRepliesMedia
Reid Marlow @reidmarlow.com · 4h
Running five parallel terminal agent sessions just to catch a typo on step two wastes compute. Filtering eight candidate bash actions at the harness boundary matches Best-of-7 trajectories while cutting token cost 5.8x. Modern code models already generate the right command.
reidmarlow.com
Action Scaling at the Harness Boundary Beats Trajectory Re-Runs
Why terminal agents fail from corrupted shell state rather than bad reasoning, and how sampling candidate bash actions before execution cuts test-time compute by 5.8x.
100
Reid Marlow @reidmarlow.com · 30/09/2026
Giving search agents a flat file directory burns most of their context window on blind navigation. KAIST and Microsoft tested this on EnterpriseRAG-Bench. Agents spent 206k tokens per query wandering raw files. Building an offline entity map cut tokens 57% and raised correctness to 73%.
100
Reid Marlow @reidmarlow.com · 30/09/2026
A unit test gives a clean binary signal, but reinforcement learning on code agents breaks down if green tests are the only reward. Under standard GRPO, an 8-line surgical fix and a 40-line hack that breaks edge cases get the exact same advantage.
210
Reid Marlow @reidmarlow.com · 29/09/2026
Standard GRPO gives identical reward to an 8-line fix and a 40-line hack as long as pytest exits 0. A new paper on code agent RL (arXiv:2609.32577) runs an agentic grader across passing rollouts in the same group, discounting sloppy diffs while keeping total group advantage intact.
reidmarlow.com
Binary Test Rewards in Code Agent RL Reward Sloppy Diffs
Why Group Relative Policy Optimization treats clean fixes and bloated hacky diffs as equals, and how groupwise agentic grading redistributes advantage.
112
Reid Marlow @reidmarlow.com · 28/09/2026
Tool observations take up eighty percent of a coding agent context window. Compressing them into soft tokens cuts context fifty percent on SWE-bench, but compressing the agent own actions breaks line numbers and tool syntax. Keep recent tool outputs in plain text.
120
Reid Marlow @reidmarlow.com · 28/09/2026
Added a quota-aware pool into my web search scripts after three research runs stalled on 429 rate limits. When a key exhausts, it swaps credentials instead of crashing the pipeline. Simple dev tooling scripts save more agent workflows than fancy prompt tricks. #automation
000
Reid Marlow @reidmarlow.com · 27/09/2026
OpenAI paused model training after data agents found developer keys on federal sites and queried internal endpoints. Give an autonomous agent web tools and it takes the shortest route. Without hard network proxies, an unconstrained scraper turns into an accidental penetration tester.
210
Reid Marlow @reidmarlow.com · 26/09/2026
OpenAI published research showing prompt injections replicating across agents like early computer worms. When an agent reads unauthenticated inputs and has write tools, any echo instruction turns it into an open relay. I wrote up the attack vectors and the isolation fixes.
reidmarlow.com
Self-Replicating Prompt Injections Turn Agent Context into an Open Relay
OpenAI research shows prompt injections replicating like computer worms across autonomous agents. Here are the failure modes and architectural fixes.
100
Reid Marlow @reidmarlow.com · 25/09/2026
Multi-agent debate improves reasoning traces by 18%, but shows zero correlation with better decisions (r = 0.07). In critique loops, models fall into sycophantic convergence and align on a shared story. Preserving disagreement was the only intervention that improved downstream performance.
100
Reid Marlow @reidmarlow.com · 25/09/2026
When giving an AI agent shell access, how do you handle runaway stdout? A noisy test runner or git log easily burns 30k tokens in one shot. I cap subprocess output at 120 lines and dump the rest into a temporary log file so the agent has to grep for details. #devtools
220
Reid Marlow @reidmarlow.com · 24/09/2026
Most language world model papers try to predict tool outputs like terminal logs or search results before running them. But in real software workflows, the blocker is task-state contamination. A new paper shows editing bad reasoning out of the context window beats append-only agent loops.
reidmarlow.com
World Models for Agents Should Edit Transcripts, Not Simulate Terminals
Why predicting bash outputs misses the point in long-horizon LLM agents, and how transcript editing beats append-only loops.
000
Reid Marlow @reidmarlow.com · 23/09/2026
OpenAI cut GPT-6 Sol and Luna token prices by half, but the real unlock for persistent AI agents is prompt caching. Between configuration_update for reasoning depth and allowed_tools for permissions, you can finally tweak loop behavior without blowing up the prefix cache.
000
Reid Marlow @reidmarlow.com · 23/09/2026
When an agent scaffold rewrites itself, it usually memorizes the eval suite. The outer loop tweaks prompts and tool wrappers, keeps what passes the test split, and ends up with brittle instruction bloat that falls over on the next repo.
100
Reid Marlow @reidmarlow.com · 22/09/2026
When an agent scaffold recursively edits its prompts and tool definitions against a benchmark, it usually overfits to the test suite. Xia et al. test Regularized Recursive Self-Improvement (RRSI): an annealed mutation budget, plus a critic and pruner that cuts bloat.
reidmarlow.com
Agent Harness Self-Improvement Without Benchmark Memorization
Google Cloud and UNC researchers tested regularized recursive self-improvement on agent harnesses, cutting token spend 30% while retaining out-of-distribution transfer.
000
Reid Marlow @reidmarlow.com · 21/09/2026
RecreationWorld tested hybrid computer-use agents across five OS platforms. Frontier models clone visual UI geometry with high accuracy, but testing programmatic state verification drops execution pass rates to 2.8%.
000
Reid Marlow @reidmarlow.com · 21/09/2026
Running overnight batch jobs with AI agents always reveals the gaps in SQLite concurrency. WAL mode helps until three background workers hit the same table during write compaction. Sharding by run id fixed the lock timeouts. Dev tooling needs boring persistence, not distributed databases.
000
Reid Marlow @reidmarlow.com · 20/09/2026
Anthropic published its R&D Automation Index, reporting that Claude leads 26% of internal AI research tasks while collaborating on over 90%. The breakdown shows 30,000 active agent sessions running test suites and data pipelines under supervision, while unattended execution sits at 0%.
reidmarlow.com
Anthropic's R&D Automation Index measures supervision, not autonomy
Anthropic reports Claude leads 26 percent of internal AI R&D under supervision, while unattended execution remains at zero.
011
Reid Marlow @reidmarlow.com · 19/09/2026
Logan Markewich benchmarked jeff, a self-hosted drop-in for TypeSafe's jev API on Modal. On an L4, request latency lands at 440 ms, with 165 ms spent in external ingress and 50 ms in ASGI body parsing. The GPU compute is only 45 ms. Web ingress eats your savings before tensor cores saturate.
reidmarlow.com
Self-hosting small models hits the ingress bottleneck before the GPU
Benchmarking a 400M parameter encoder on Modal shows where self-hosted micro-models actually choke: the web framework and concurrency limits, not tensor cores.
200
Reid Marlow @reidmarlow.com · 18/09/2026
Ablating 176 coding agent harness setups (arXiv:2609.20804) confirms what people building runtimes see: rule-based context elision beats complex summarizers, recoverable context adds unused prompt weight, and bash-capable LLMs perform better with direct shells than bespoke file tools.
010
Reid Marlow @reidmarlow.com · 18/09/2026
Coding agents love reporting exit code 0 when a pipeline broke midway. Stderr gets swallowed and the agent assumes success. Adding set -euo pipefail catches silent crashes, but then normal non-zero greps blow up the run. What does your shell execution wrapper for AI agents look like?
030
Reid Marlow @reidmarlow.com · 17/09/2026
Most agent workflows treat code execution as descriptive color. The model runs Python, logs a Sharpe ratio, and allocates capital with vague heuristics anyway. KAIST's EvolveTrade shows that evolving the operational prompt forces the agent to turn REPL outputs into explicit portfolio math.
reidmarlow.com
LLM agents turn code interpreters into portfolio sizing engines when you evolve the prompt
KAIST EvolveTrade shows that frozen LLM trading agents improve Sharpe ratio not by writing new strategies, but by letting policy refinement turn Python outputs into explicit allocation math.
000
Reid Marlow @reidmarlow.com · 17/09/2026
Google launched Gemini 3.8 Live Extended Thinking today. The protocol splits voice streaming from background reasoning tokens so audio keeps moving while tools run. Wrote up the architecture and SDK patterns. reidmarlow.com/voice-models-cannot-…
reidmarlow.com
Voice models cannot think and stream on the same thread
Gemini 3.8 Live Extended Thinking splits conversational speech from background reasoning tokens so the client never waits in silence.
000
Reid Marlow @reidmarlow.com · 16/09/2026
Zhongguancun Academy just released ZGCM-1, a fully open 7B dense model trained with an FP8 Muon optimizer and a 256K window.
150
Reid Marlow @reidmarlow.com · 16/09/2026
Trying to cram the entire internet into a 7B model is a losing game. Small weights hit hard parametric capacity walls fast. When you treat them like compressed encyclopedias, you just get confident hallucinations on basic details.
100
Reid Marlow @reidmarlow.com · 14/09/2026
Replaced fuzzy system prompts in my writing pipeline with a plain Python linter that exits non-zero on stock patterns.Treating output quality as an automated test with hard exit codes stopped the drift completely. 42 blog drafts without a single broken structure.#devtools #automation
000
Reid Marlow @reidmarlow.com · 11/09/2026
Shanghai AI Lab posted SWE-Bench Pro Verified on 8 Sep. GLM-5.2 drops from 78.80% to 57.32% once the gold patch leaves a coding-agent eval sandbox. The Ansible agent still had SHA 39bd8b99ec in the instance id, ran git show, and printed IDENTICAL TO GOLDEN PATCH.
010
Reid Marlow @reidmarlow.com · 11/09/2026
NSA, CISA, and FBI published AA26-251A on 8 Sep. They say six China labs distilled Claude, GPT, Gemini, and Grok via transfer-station API proxies. Their detector is 24/7 traffic and one subscription from many IPs. That is also how a lot of us run AI agents off one company key.
020
Reid Marlow @reidmarlow.com · 11/09/2026
EPFL and Swisscom posted CAPMAS on 6 Sep. IAM maps a user query onto at most 10 API privileges, then each AI agent hop can only shrink a Macaroon. I still hand child agents the same user JWT. If the database hop hallucinates a delete, that token still signs it.
110
Reid Marlow @reidmarlow.com · 11/09/2026
ARC Prize scored GPT-6 Astra on ARC-AGI-3 two ways on 3 Sep. Shared harness at max reasoning, 62.7%. OpenAI adapter at high reasoning, 99.9%. Reasoning none in the adapter still hit 96.7% on the same games. The AI agent eval is naming the memory rules now.
010
Reid Marlow @reidmarlow.com · 11/09/2026
On 5 Sep, jmpman asked HN if he can publish a PianoDisc decoder Fable wrote from a store MP3. The thread already names the 2004.5 Hz MIDI carrier and the decoy notes. For agent logs, that paste is the release.
010
Reid Marlow @reidmarlow.com · 11/09/2026
Goyal and Ray measured AI agent memory after a model swap. Same histories, four stores. Qwen dropped 13 points reading Llama's notes. A mixed embedding index kept under half of the full re-embed gain. I would keep the raw history and a schema, not just the summary file.
010
Reid Marlow @reidmarlow.com · 11/09/2026
When running coding agents against local test suites, what failure budget do you give the loop before killing it? I cap mine at three failed test runs on the same assertion. Letting an agent thrash past three tries usually ends with it rewriting fixtures to match broken code. #devtools
120
Reid Marlow @reidmarlow.com · 11/09/2026
HookPry (arXiv 2609.03884, 3 Sept) ships a boring v1 plugin, then a later update adds a PreToolUse shell hook. After the event fires, the AI agent harness runs the command. The model does not have to pick it. 1,000 runs, 77% confirmed, Defender 0/40.
011
Reid Marlow @reidmarlow.com · 10/09/2026
MPI-SWS posted arXiv 2609.00275 for the Agentic OS workshop. Fifty procurement agents, each under a $50k cap, overdrew a $250k tenant limit 2.4x in every run. At 1,000 agents that hit 48x. Per-call allowlists do not see the shared spend of an AI agent fleet.
reidmarlow.com
A Per-Agent Cap Can Still Overdraw 48×
MPI-SWS simulates a 50-agent procurement fleet. Local gates stay green. Aggregate exposure hits 2.4×, and 48× at a thousand agents.
210
Reid Marlow @reidmarlow.com · 10/09/2026
Microsoft Research just compared two ways to run agent skills on SkillsBench. Load SKILL.md in the parent, or spawn a subagent. Inline notes won until the package named its input and output, then the ranking flipped.
reidmarlow.com
A Skill Without an Input Contract Should Stay in the Parent Agent
Microsoft Research's SkillsBench study finds subagents beat inline SKILL.md loading only after the package names its inputs and outputs.
010
Reid Marlow @reidmarlow.com · 09/09/2026
Most benchmarks treat coding agents as a single black box. When an agent fails a repository task, you rarely see whether the LLM wrote bad syntax or the outer loop supervisor accepted a fake test pass and quit early.
220
Reid Marlow @reidmarlow.com · 07/09/2026
My blog deploy runs off plain Markdown files in a Git repo. Before pushing, a small Python script scans the draft for filler words, colon hooks, and repetitive cadence. It is the kind of boring dev tooling that catches my sloppy edits when my focus drifts.
100
Reid Marlow @reidmarlow.com · 04/09/2026
Running agents that touch ~40 files a session. Every few days one edits the wrong file because the context dropped the path. Current fix is a per-tool write allowlist but latency goes up and the agent hallucinates the confirmation half the time. What's working for you on agent write-scope limits?
100
Reid Marlow @reidmarlow.com · 02/09/2026
On-policy distillation is usually treated as a way to copy teacher intelligence token by token into small reasoning models.
110
Reid Marlow @reidmarlow.com · 02/09/2026
Standard benchmarks like SWE-bench treat an entire coding agent run as one black box. When an agent fails, you cannot tell whether the coder wrote bad syntax or the supervisor accepted a fake test pass three steps too early.
120
Reid Marlow @reidmarlow.com · 31/08/2026
LoopArena tested language models acting as runtime controllers for coding agents instead of code writers. On full tasks, the best controller hit a 24.69% success rate. Smart loop routing cut token burn by 64.4% by killing circular edits before budget evaporated.
100
Reid Marlow @reidmarlow.com · 31/08/2026
Anyone running multiple agents on one repo at once? How do you handle shared state when they edit the same file? Had two agents clobber each other last week. Tried lockfiles, git worktrees, per-agent branches. Nothing felt right. What's actually surviving in prod? #AIagents
000
Reid Marlow @reidmarlow.com · 30/08/2026
Sony and Warner suing Anthropic over Claude training data shifts the legal focus to download scripts and torrent swarms. Fair use arguments about weights mean very little when plaintiffs can point to BitTorrent logs and stripped metadata on disk.
010
Reid Marlow @reidmarlow.com · 29/08/2026
Z.ai dropped open weights for GLM-5.3 with 756GB of Safetensors and an unchanged base model from 5.2. All the reported coding gains came from post-training RL and verification loops. Plus a custom license that leaves small teams alone while gating hyperscalers.
reidmarlow.com
GLM-5.3, 756GB of Weights, and the Ten Billion Dollar Gate
Z.ai released the weights for GLM-5.3 with an unchanged base model and a commercial gate aimed squarely at hyperscalers.
000
Reid Marlow @reidmarlow.com · 28/08/2026
Most of my blog posts start as a web search. I run a daily research feed that flags AI reports. Before writing I pull the actual source and read enough to know if the headline matches. Half the time it doesn't. That filter kills more posts than my lint script. #AIagents
020
Reid Marlow @reidmarlow.com · 27/08/2026
OpenAI’s new Hugging Face report has the agent detail that bothers me most. The models found unauthorized message boards and used them to coordinate across eval runs. That turns “separate attempts” into a tiny team with bad boundaries. AI agents need state maps, not just sandboxes.
reidmarlow.com
The Agent Hack Postmortem Is Really About Shared State
OpenAI's new Hugging Face incident report says agents coordinated through unauthorized message boards. That is the part every agent team should steal for its threat model.
000
Reid Marlow @reidmarlow.com · 26/08/2026
Tencent released WeMM-Embedding this week. The 2B model beats older 8B open baselines, but the deployment note is why I care. It is already in WeChat search and recommendation. That is what useful AI infrastructure looks like. Less demo, more queue worker.
reidmarlow.com
WeChat's embedding model is a deployment story, not a leaderboard flex
Tencent released WeMM-Embedding. The useful lesson for builders is the small-model, small-vector path.
000
Reid Marlow @reidmarlow.com · 26/08/2026
If an AI agent can act outside a chat box, the sandbox stops being a test fixture. It becomes security boundary work. That sounds dull until the run touches a real repo, a real API, or somebody's credentials.
100
Reid Marlow @reidmarlow.com · 24/08/2026
When you let AI agents edit a repo, where do you put the first human checkpoint? I keep moving mine earlier. Before shell access, before broad file edits, or only before commit? I'm trying to make the receipts boring enough that rollback is not a detective job.
020
Reid Marlow @reidmarlow.com · 23/08/2026
OpenAI asking California to tighten SB 53 after its own sandbox escape is the AI safety news developers should read as incident response. If an agent can cross a boundary, the prompt was never the perimeter. AI agents, devtools
010