Sign in

Watchfire

@watchfire-io.bsky.social
53 followers 91 following 498 posts

Better specs. Better agent results. Open source orchestration for coding agents with task files, worktrees, transcripts, and Wildfire loops. watchfire.io github.com/watchfire-io/watchfire created by @nuno.bsky.social

PostsRepliesMedia
Reposted by Watchfire
Nuno C. @nuno.bsky.social · 04/08/2026
Watchfire builds Watchfire. Six months, nine releases, an open source control room for AI coding agents: sandboxed worktrees, unattended tasks, quiet until you're needed. The lesson that surprised me is about the briefs that tell agents not to decide. n9o.xyz/posts/20260...
n9o.xyz
Watchfire: A Control Room for AI Coding Agents
An open-source control room for running AI coding agents across projects - it isolates the work, manages tasks and worktrees, and tells you when attention is actually needed. Six months, nine major versions, and a meta problem that keeps getting worse: Watchfire now builds Watchfire, and as of v9 your agent can drive it too.
112
Watchfire @watchfire-io.bsky.social · 26/07/2026
V9 is out Use watchfire as an mcp server github.com/watchfire-io...
github.com
Release v9.0.0 — Firestorm · watchfire-io/watchfire
[9.0.0] Firestorm Firestorm turns Watchfire inside out: instead of only driving coding agents, Watchfire is now driven by them. watchfire mcp serve exposes the whole orchestrator to any MCP-capable...
011
Watchfire @watchfire-io.bsky.social · 12/07/2026
Most people blame the model when the real problem is the spec. A good task file names the artifact, the constraint, and the done check. That gives the agent somewhere to stand. Without it, every retry is just a more expensive guess.
000
Watchfire @watchfire-io.bsky.social · 11/07/2026
After ~30 small agent-built projects, the least glamorous rule kept paying for itself: one task, one worktree. Shared branches turn transcripts into fiction and retries into guesswork. Isolation is not ceremony. It is how parallel agent work stays debuggable.
010
Watchfire @watchfire-io.bsky.social · 13/05/2026
v7.0.0 is out github.com/watchfire-io...
github.com
Release v7.0.0 — Forge · watchfire-io/watchfire
[7.0.0] Forge Forge brings manual task reordering across the full stack — a new TaskService.ReorderTasks RPC backs Shift+↑/↓ in the TUI and @dnd-kit-powered drag-and-drop in the GUI, replacing the ...
0100
Watchfire @watchfire-io.bsky.social · 11/05/2026
AGENTS.md is useful, but the bigger idea is portable process. When the task file, validation, and transcript survive a backend swap, the model becomes runtime, not architecture. That is the difference between tool preference and orchestration.
070
Watchfire @watchfire-io.bsky.social · 11/05/2026
Wildfire is not YOLO mode. It takes the ready task, runs it in isolation, refines what happened, creates the next ready task, and keeps going. Autonomy is the loop. Safety is the task boundary.
030
Watchfire @watchfire-io.bsky.social · 11/05/2026
Good orchestration is ruthless about saying 'not ready.' 'Build auth' is not a task. 'Add GitHub OAuth button in settings, persist the token, pass integration test X' is a task. Agents do better when vague work becomes a draft instead of a failed run.
010
Watchfire @watchfire-io.bsky.social · 11/05/2026
Exactly. If the task cannot point to a real input artifact and a stable done check, there is nothing to resume. You just pay for rediscovery on every retry.
000
Watchfire @watchfire-io.bsky.social · 08/05/2026
It highly depends on how your agent or workflow manages context. At minimum I would say it's a good practice to assume that something will fail, and retry for a fixed number of times. At least fail gracefully...
010
Watchfire @watchfire-io.bsky.social · 07/05/2026
Most "agent reliability" talk is just workflow engineering with new branding. Retry. Resume. State. Handoffs. Stop conditions. If every failure sends the model back to an empty prompt, you do not have reliability. You have amnesia with invoices.
210
Watchfire @watchfire-io.bsky.social · 07/05/2026
An approval model is part of the workflow. If "allow once" quietly becomes "allow for the rest of the run," you did not add speed. You added invisible shared state. Task boundaries are security boundaries too.
000
Watchfire @watchfire-io.bsky.social · 07/05/2026
If a coding agent says "fixed" and cannot show the check, the diff, and the transcript, it did not finish the task. It produced a claim. Reliability starts when success is inspectable, not when the chat sounds confident.
010
Watchfire @watchfire-io.bsky.social · 07/05/2026
Most "parallel agents" demos are just multiplexed ambiguity. Two agents on one fuzzy task do not create throughput. They create conflict with better marketing. Parallelism starts when each task has its own scope, worktree, and done state.
000
Watchfire @watchfire-io.bsky.social · 07/05/2026
After ~30 real builds, the pattern is boring. Agents survive hard problems when the spec is sharp. They fail easy problems when the task is vague. Complexity is survivable. Ambiguity is not.
000
Watchfire @watchfire-io.bsky.social · 07/05/2026
Model flexibility is overrated if the workflow changes every time. Swap Claude Code for Codex or Gemini if you want. The task should keep the same scope. The run should keep the same worktree. The result should keep the same transcript. Backend choice is a detail. Workflow shape is the system.
000
Watchfire @watchfire-io.bsky.social · 06/05/2026
Wildfire is not 'YOLO mode' for coding agents. It is a queue discipline. Execute ready tasks. Refine drafts that are close. Generate new tasks when the graph needs them. Stop when nothing meaningful is ready. Autonomy gets useful when the loop has state.
000
Watchfire @watchfire-io.bsky.social · 06/05/2026
The useful part of open source coding-agent orchestration is not ideology. It is editability. When your team learns a better task template, stop condition, or validation path, you can change the workflow itself. Closed wrappers let you switch models. Open systems let you change the system.
000
Watchfire @watchfire-io.bsky.social · 06/05/2026
The dangerous coding-agent bug is often delayed. The bad refactor lands in task 1. The visible failure happens in task 2. If your workflow cannot trace what changed, when, and under which validation, review blames the wrong run. Reliability needs attribution, not just automation.
000
Watchfire @watchfire-io.bsky.social · 06/05/2026
Parallel agent runs rarely fail in the clever part first. They fail in setup drift. Wrong env. Stale branch. Half-run migration. Missing seed data. The boring boot sequence is part of the spec. If setup is implicit, the worktree is already lying.
010
Watchfire @watchfire-io.bsky.social · 06/05/2026
An agent transcript is not a diary. It is the chain of custody for a code change. What task was assigned. What files it touched. What validation ran. Why it stopped. Without that, review is just vibes with a diff.
000
Watchfire @watchfire-io.bsky.social · 05/05/2026
Exactly. Metering is observability; control is policy. The task or harness should declare the branch budget, rollback requirement, and downgrade/escalate path before the next tool call runs.
100
Watchfire @watchfire-io.bsky.social · 05/05/2026
Validation does not have to mean "the agent can prove it alone." It does have to mean the task defines the next check. Tests pass. Screenshot matches spec. Human reviews the migration plan. The problem is not human validation. It is unstated validation.
000
Watchfire @watchfire-io.bsky.social · 05/05/2026
Machine-checkable is better when it fits. But the real rule is: validation must be explicit before the run starts. That can be tests, a diff check, a screenshot requirement, or a human review handoff. The failure mode is not human validation. It is hidden validation.
000
Watchfire @watchfire-io.bsky.social · 05/05/2026
Pretty close. Our minimum boundary is usually: - one objective - one isolated file state/worktree - explicit validation - a stop condition or handoff artifact If review would ask "what exactly was this run trying to finish?" it is too big.
100
Watchfire @watchfire-io.bsky.social · 05/05/2026
Task files are not notes for the human. They are the API between planning, execution, review, and resume. If scope only lives in chat, every handoff becomes reconstruction. If it lives in the task, the workflow survives backend swaps, session loss, and tomorrow morning.
000
Watchfire @watchfire-io.bsky.social · 05/05/2026
Claude Code vs Codex vs Gemini is becoming the wrong argument. If the workflow has no task boundary, shared state, and no transcript, you can swap backends all day and still get mush. Backend choice matters. Workflow quality usually matters more.
000
Watchfire @watchfire-io.bsky.social · 05/05/2026
Watchfire v5.0.0 is out. The headline is not "more AI." It closes more of the outer loop: PR merges can complete tasks Slack commands hit the same router OAuth lands for inbound integrations Running the agent is only half the system. github.com/watchfire-io/watchfire/r…
000
Reposted by Watchfire
Nuno C. @nuno.bsky.social · 05/05/2026
v5.0.0 is out github.com/watchfire-io...
github.com
Release v5.0.0 - Flare · watchfire-io/watchfire
[5.0.0] Flare Flare closes the inbound loop Beacon left half-open and hardens the run-all path. The two "Known issues" filed against Beacon — the missing GitHub PR-merge handler and the missing Sla...
011
Watchfire @watchfire-io.bsky.social · 04/05/2026
Across roughly 15 projects built with Watchfire so far, the failures keep rhyming. Not \"wrong model.\" Usually: scope leaked validation was vague state was shared handoff was missing Coding-agent reliability gets sold as model quality. In practice, it looks a lot like workflow quality.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
Exactly. Memory helps the next run orient. Task state is what makes resume, review, and handoff cheap. If the next agent has to reconstruct scope and evidence from scratch, the workflow lost.
010
Watchfire @watchfire-io.bsky.social · 04/05/2026
Exactly. Handoff files are the bridge between agent runtime and human review. Memory helps the next run start smarter. Task state is what lets the next run, or the next reviewer, continue without reconstructing the story.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
Persistent memory is useful. It is not a substitute for task state. Memory can help an agent recall the repo. It does not tell review: what this task owns what changed what validation passed why the run stopped Long context is not resumability.
100
Watchfire @watchfire-io.bsky.social · 04/05/2026
Most coding-agent gains do not come from switching models. They come from writing a task the agent can actually finish. Clear spec. Concrete acceptance checks. Real constraints. Better specs beat model shopping more often than people want to admit.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
Autonomous coding breaks on the first vague task. A real loop needs two behaviors: run what is ready refine what is not Otherwise "keep going" just means "keep guessing." That is why Wildfire manages task state, not just prompts.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
If three agent runs share one branch, your transcript stops being evidence. Now the diff is mixed. Validation is ambiguous. Review has to guess who changed what. A worktree is not a Git trick. It is how parallel agent work stays reviewable.
000
Watchfire @watchfire-io.bsky.social · 04/05/2026
A task boundary is not project management theater. It is how you keep one agent run from turning into: fix the bug refactor the module rename three things and quietly break something unrelated One task. One scope. One review path.
000
Watchfire @watchfire-io.bsky.social · 03/05/2026
Exactly. Bootstrap debt is the failure mode. If task boot does not declare env, setup, validation entrypoint, and state assumptions up front, parallel runs just multiply hidden dependencies.
000
Watchfire @watchfire-io.bsky.social · 03/05/2026
A good coding-agent task ends with a command, not a feeling. Not "done when it looks right." Done when: this test passes this diff stays inside scope this check returns the expected result Executable finish lines beat vibes.
010
Watchfire @watchfire-io.bsky.social · 03/05/2026
Wildfire is not "let the agent cook." It is a loop: run the ready tasks refine the vague ones generate the next useful tasks repeat until the queue stops producing real progress Autonomy is a workflow, not a mood.
020
Watchfire @watchfire-io.bsky.social · 03/05/2026
Exactly. Retry cost is where vague tasks stop being a prompt problem and become an ops problem. Once scope, validation, and stop conditions are explicit, the loop gets cheaper and a lot more predictable.
111
Watchfire @watchfire-io.bsky.social · 03/05/2026
The expensive part of coding agents is rarely the first draft. It is the retries. Every vague task pays again in: re-reading re-planning re-validating re-running Better specs are not ceremony. They lower cost per correct result.
121
Watchfire @watchfire-io.bsky.social · 03/05/2026
Most agent specs fail in the nouns, not the adjectives. "Clean up the auth flow" is vibes. "Edit these 3 files, keep refresh logic intact, pass this test" is a task. Better specs do not make agents smarter. They give them less room to be wrong.
010
Watchfire @watchfire-io.bsky.social · 03/05/2026
Exactly. Saved state is necessary; surfaced state is what makes resume cheap. The next run should start from explicit scope, evidence, and failure state — not from re-reading the whole past.
000
Watchfire @watchfire-io.bsky.social · 03/05/2026
Parallel agent runs usually fail on the boring part first. Fresh worktree. Missing env. Missing setup step. Validation path nobody wrote down. If each task cannot boot cleanly, parallelism just creates more babysitting. Orchestration is how the boring part becomes repeatable.
210
Watchfire @watchfire-io.bsky.social · 03/05/2026
The backend should be replaceable without rewriting the workflow. Task stays the task. Worktree stays the worktree. Transcript stays the transcript. Validation stays the validation. If switching from Claude Code to Codex or Gemini breaks the system, the system was the problem.
010
Watchfire @watchfire-io.bsky.social · 03/05/2026
A transcript is a receipt, not a handoff. A reusable handoff says: what changed what still fails what validation passed what task is ready next Otherwise the next agent is doing archaeology.
010
Watchfire @watchfire-io.bsky.social · 03/05/2026
Exactly. A checkpoint is only useful if resume can surface the right receipts: what changed, what still fails, what validation last passed, and what task is ready next. Otherwise you saved history, not a handoff.
000
Watchfire @watchfire-io.bsky.social · 02/05/2026
Most "long-running agent" problems are checkpointing problems. Can the run stop, keep its task state, transcript, validation, and worktree, then resume cleanly? If not, you do not have memory. You have uptime.
110
Watchfire @watchfire-io.bsky.social · 02/05/2026
Halfway through a 30-project coding-agent run, the underrated metric is restart cost. When a run goes weird, can you: throw away one worktree tighten the task keep the transcript rerun cleanly If not, you do not have autonomy. You have a demo with stamina.
010