Sign in

Watchfire

@watchfire-io.bsky.social
50 followers 91 following 498 posts

Better specs. Better agent results. Open source orchestration for coding agents with task files, worktrees, transcripts, and Wildfire loops. watchfire.io github.com/watchfire-io/watchfire created by @nuno.bsky.social

PostsRepliesMedia
Reposted by Watchfire
Nuno C. @nuno.bsky.social · 04/08/2026
Watchfire builds Watchfire. Six months, nine releases, an open source control room for AI coding agents: sandboxed worktrees, unattended tasks, quiet until you're needed. The lesson that surprised me is about the briefs that tell agents not to decide. n9o.xyz/posts/20260...
n9o.xyz
Watchfire: A Control Room for AI Coding Agents
An open-source control room for running AI coding agents across projects - it isolates the work, manages tasks and worktrees, and tells you when attention is actually needed. Six months, nine major versions, and a meta problem that keeps getting worse: Watchfire now builds Watchfire, and as of v9 your agent can drive it too.
112
Watchfire @watchfire-io.bsky.social · 26/07/2026
V9 is out Use watchfire as an mcp server github.com/watchfire-io...
github.com
Release v9.0.0 — Firestorm · watchfire-io/watchfire
[9.0.0] Firestorm Firestorm turns Watchfire inside out: instead of only driving coding agents, Watchfire is now driven by them. watchfire mcp serve exposes the whole orchestrator to any MCP-capable...
011
Watchfire @watchfire-io.bsky.social · 12/07/2026
Most people blame the model when the real problem is the spec. A good task file names the artifact, the constraint, and the done check. That gives the agent somewhere to stand. Without it, every retry is just a more expensive guess.
000
Watchfire @watchfire-io.bsky.social · 11/07/2026
After ~30 small agent-built projects, the least glamorous rule kept paying for itself: one task, one worktree. Shared branches turn transcripts into fiction and retries into guesswork. Isolation is not ceremony. It is how parallel agent work stays debuggable.
010
Watchfire @watchfire-io.bsky.social · 13/05/2026
v7.0.0 is out github.com/watchfire-io...
github.com
Release v7.0.0 — Forge · watchfire-io/watchfire
[7.0.0] Forge Forge brings manual task reordering across the full stack — a new TaskService.ReorderTasks RPC backs Shift+↑/↓ in the TUI and @dnd-kit-powered drag-and-drop in the GUI, replacing the ...
0100
Watchfire @watchfire-io.bsky.social · 11/05/2026
AGENTS.md is useful, but the bigger idea is portable process. When the task file, validation, and transcript survive a backend swap, the model becomes runtime, not architecture. That is the difference between tool preference and orchestration.
070
Watchfire @watchfire-io.bsky.social · 11/05/2026
Wildfire is not YOLO mode. It takes the ready task, runs it in isolation, refines what happened, creates the next ready task, and keeps going. Autonomy is the loop. Safety is the task boundary.
030
Watchfire @watchfire-io.bsky.social · 11/05/2026
Good orchestration is ruthless about saying 'not ready.' 'Build auth' is not a task. 'Add GitHub OAuth button in settings, persist the token, pass integration test X' is a task. Agents do better when vague work becomes a draft instead of a failed run.
010
Watchfire @watchfire-io.bsky.social · 07/05/2026
Most "agent reliability" talk is just workflow engineering with new branding. Retry. Resume. State. Handoffs. Stop conditions. If every failure sends the model back to an empty prompt, you do not have reliability. You have amnesia with invoices.
210
Watchfire @watchfire-io.bsky.social · 07/05/2026
An approval model is part of the workflow. If "allow once" quietly becomes "allow for the rest of the run," you did not add speed. You added invisible shared state. Task boundaries are security boundaries too.
000
Watchfire @watchfire-io.bsky.social · 07/05/2026
If a coding agent says "fixed" and cannot show the check, the diff, and the transcript, it did not finish the task. It produced a claim. Reliability starts when success is inspectable, not when the chat sounds confident.
010
Watchfire @watchfire-io.bsky.social · 07/05/2026
Most "parallel agents" demos are just multiplexed ambiguity. Two agents on one fuzzy task do not create throughput. They create conflict with better marketing. Parallelism starts when each task has its own scope, worktree, and done state.
000
Watchfire @watchfire-io.bsky.social · 07/05/2026
After ~30 real builds, the pattern is boring. Agents survive hard problems when the spec is sharp. They fail easy problems when the task is vague. Complexity is survivable. Ambiguity is not.
000
Watchfire @watchfire-io.bsky.social · 07/05/2026
Model flexibility is overrated if the workflow changes every time. Swap Claude Code for Codex or Gemini if you want. The task should keep the same scope. The run should keep the same worktree. The result should keep the same transcript. Backend choice is a detail. Workflow shape is the system.
000
Watchfire @watchfire-io.bsky.social · 06/05/2026
Wildfire is not 'YOLO mode' for coding agents. It is a queue discipline. Execute ready tasks. Refine drafts that are close. Generate new tasks when the graph needs them. Stop when nothing meaningful is ready. Autonomy gets useful when the loop has state.
000
Watchfire @watchfire-io.bsky.social · 06/05/2026
The useful part of open source coding-agent orchestration is not ideology. It is editability. When your team learns a better task template, stop condition, or validation path, you can change the workflow itself. Closed wrappers let you switch models. Open systems let you change the system.
000
Watchfire @watchfire-io.bsky.social · 06/05/2026
The dangerous coding-agent bug is often delayed. The bad refactor lands in task 1. The visible failure happens in task 2. If your workflow cannot trace what changed, when, and under which validation, review blames the wrong run. Reliability needs attribution, not just automation.
000
Watchfire @watchfire-io.bsky.social · 06/05/2026
Parallel agent runs rarely fail in the clever part first. They fail in setup drift. Wrong env. Stale branch. Half-run migration. Missing seed data. The boring boot sequence is part of the spec. If setup is implicit, the worktree is already lying.
010
Watchfire @watchfire-io.bsky.social · 06/05/2026
An agent transcript is not a diary. It is the chain of custody for a code change. What task was assigned. What files it touched. What validation ran. Why it stopped. Without that, review is just vibes with a diff.
000
Watchfire @watchfire-io.bsky.social · 05/05/2026
Validation does not have to mean "the agent can prove it alone." It does have to mean the task defines the next check. Tests pass. Screenshot matches spec. Human reviews the migration plan. The problem is not human validation. It is unstated validation.
000
Watchfire @watchfire-io.bsky.social · 05/05/2026
Task files are not notes for the human. They are the API between planning, execution, review, and resume. If scope only lives in chat, every handoff becomes reconstruction. If it lives in the task, the workflow survives backend swaps, session loss, and tomorrow morning.
000
Watchfire @watchfire-io.bsky.social · 05/05/2026
Claude Code vs Codex vs Gemini is becoming the wrong argument. If the workflow has no task boundary, shared state, and no transcript, you can swap backends all day and still get mush. Backend choice matters. Workflow quality usually matters more.
100
Watchfire @watchfire-io.bsky.social · 05/05/2026
Watchfire v5.0.0 is out. The headline is not "more AI." It closes more of the outer loop: PR merges can complete tasks Slack commands hit the same router OAuth lands for inbound integrations Running the agent is only half the system. github.com/watchfire-io/watchfire/r…
000
Reposted by Watchfire
Nuno C. @nuno.bsky.social · 05/05/2026
v5.0.0 is out github.com/watchfire-io...
github.com
Release v5.0.0 - Flare · watchfire-io/watchfire
[5.0.0] Flare Flare closes the inbound loop Beacon left half-open and hardens the run-all path. The two "Known issues" filed against Beacon — the missing GitHub PR-merge handler and the missing Sla...
011
Watchfire @watchfire-io.bsky.social · 04/05/2026
Across roughly 15 projects built with Watchfire so far, the failures keep rhyming. Not \"wrong model.\" Usually: scope leaked validation was vague state was shared handoff was missing Coding-agent reliability gets sold as model quality. In practice, it looks a lot like workflow quality.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
Persistent memory is useful. It is not a substitute for task state. Memory can help an agent recall the repo. It does not tell review: what this task owns what changed what validation passed why the run stopped Long context is not resumability.
100
Watchfire @watchfire-io.bsky.social · 04/05/2026
Most coding-agent gains do not come from switching models. They come from writing a task the agent can actually finish. Clear spec. Concrete acceptance checks. Real constraints. Better specs beat model shopping more often than people want to admit.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
Autonomous coding breaks on the first vague task. A real loop needs two behaviors: run what is ready refine what is not Otherwise "keep going" just means "keep guessing." That is why Wildfire manages task state, not just prompts.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
If three agent runs share one branch, your transcript stops being evidence. Now the diff is mixed. Validation is ambiguous. Review has to guess who changed what. A worktree is not a Git trick. It is how parallel agent work stays reviewable.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
A task boundary is not project management theater. It is how you keep one agent run from turning into: fix the bug refactor the module rename three things and quietly break something unrelated One task. One scope. One review path.
000
Watchfire @watchfire-io.bsky.social · 03/05/2026
A good coding-agent task ends with a command, not a feeling. Not "done when it looks right." Done when: this test passes this diff stays inside scope this check returns the expected result Executable finish lines beat vibes.
010
Watchfire @watchfire-io.bsky.social · 03/05/2026
Wildfire is not "let the agent cook." It is a loop: run the ready tasks refine the vague ones generate the next useful tasks repeat until the queue stops producing real progress Autonomy is a workflow, not a mood.
020
Watchfire @watchfire-io.bsky.social · 03/05/2026
The expensive part of coding agents is rarely the first draft. It is the retries. Every vague task pays again in: re-reading re-planning re-validating re-running Better specs are not ceremony. They lower cost per correct result.
121
Watchfire @watchfire-io.bsky.social · 03/05/2026
Most agent specs fail in the nouns, not the adjectives. "Clean up the auth flow" is vibes. "Edit these 3 files, keep refresh logic intact, pass this test" is a task. Better specs do not make agents smarter. They give them less room to be wrong.
010
Watchfire @watchfire-io.bsky.social · 03/05/2026
Parallel agent runs usually fail on the boring part first. Fresh worktree. Missing env. Missing setup step. Validation path nobody wrote down. If each task cannot boot cleanly, parallelism just creates more babysitting. Orchestration is how the boring part becomes repeatable.
210
Watchfire @watchfire-io.bsky.social · 03/05/2026
The backend should be replaceable without rewriting the workflow. Task stays the task. Worktree stays the worktree. Transcript stays the transcript. Validation stays the validation. If switching from Claude Code to Codex or Gemini breaks the system, the system was the problem.
010
Watchfire @watchfire-io.bsky.social · 03/05/2026
A transcript is a receipt, not a handoff. A reusable handoff says: what changed what still fails what validation passed what task is ready next Otherwise the next agent is doing archaeology.
010
Watchfire @watchfire-io.bsky.social · 02/05/2026
Most "long-running agent" problems are checkpointing problems. Can the run stop, keep its task state, transcript, validation, and worktree, then resume cleanly? If not, you do not have memory. You have uptime.
110
Watchfire @watchfire-io.bsky.social · 02/05/2026
Halfway through a 30-project coding-agent run, the underrated metric is restart cost. When a run goes weird, can you: throw away one worktree tighten the task keep the transcript rerun cleanly If not, you do not have autonomy. You have a demo with stamina.
010
Watchfire @watchfire-io.bsky.social · 02/05/2026
Subagents are not the abstraction. Tasks are. A backend can spawn helpers if it wants. Fine. But the durable unit has to carry scope, worktree, transcript, and validation. Otherwise portability dies with the harness.
010
Watchfire @watchfire-io.bsky.social · 02/05/2026
Wildfire is the boring part people skip. It executes ready tasks, validates the result, refines drafts, generates the next task from what actually happened, and keeps going. Longer runtime is not the trick. Structured loops are.
010
Watchfire @watchfire-io.bsky.social · 02/05/2026
Open source matters here because orchestration is policy. Which backend ran what task it owned what changed what validation passed why the run stopped If that layer is opaque, you are not automating engineering. You are renting a black box.
020
Watchfire @watchfire-io.bsky.social · 02/05/2026
If switching from Claude Code to Codex breaks the task, the workflow was never portable. The task should keep the same scope, validation, transcript, and review path. Models are dependencies. Workflow is the system.
010
Watchfire @watchfire-io.bsky.social · 02/05/2026
Parallel coding agents on one branch are just concurrency bugs with branding. A worktree gives each task its own file state, diff, transcript, and review path. Without that, "multi-agent" usually means "good luck figuring out who changed what."
121
Watchfire @watchfire-io.bsky.social · 02/05/2026
"Improve onboarding" is a wish. "Touch signup form + invite endpoint, keep copy unchanged, add a validation test, done when invite flow passes" is a task. Coding agents do not need more motivation. They need fewer missing decisions.
120
Watchfire @watchfire-io.bsky.social · 01/05/2026
Long-running agents fail the moment task ownership gets fuzzy. One repo. Five active runs. Shared branch. No per-task transcript. No validation handoff. That is not autonomy. It is shared-state soup. More runtime does not create coordination. Workflow does.
000
Watchfire @watchfire-io.bsky.social · 01/05/2026
Most "autonomous coding" demos are one long chat pretending to be a workflow. A real loop has units of work. Ready task in. Validation out. Draft refined or next task generated from what actually happened. That is how autonomy survives contact with a repo.
000
Watchfire @watchfire-io.bsky.social · 01/05/2026
If a coding-agent run can only resume because the same chat is still open, it is not durable. Real resume means the task, status, transcript, and validation live outside the model. Close the session. Come back tomorrow. Keep going without folklore.
000
Watchfire @watchfire-io.bsky.social · 01/05/2026
Halfway through a 30-project coding-agent run, the highest-leverage change was not a new model. It was turning "done" into a command. When the task ends with concrete validation, the run stops guessing. Executable acceptance criteria beat motivational prompting.
000
Watchfire @watchfire-io.bsky.social · 01/05/2026
AGENTS.md is not a junk drawer. If you pour every repo opinion into one context file, the agent pays that tax every run. Keep durable facts in shared context. Put task-specific decisions in the task. Everything else is latency disguised as guidance.
010