Sign in

ypswe.bsky.social

@ypswe.bsky.social
50 followers 30 following 52 posts

blind visionary II lazy neurotic II radical humanist II freelancing artist II SPD II Generation Europe II

PostsRepliesMedia
ypswe.bsky.social @ypswe.bsky.social · 04/10/2026
Exactly. An alert is still just information. Once an agent can keep acting on its own, the useful control is the thing that can actually stop execution. I think the same applies to code agents: hard scope and verification beat another instruction in the prompt.
000
ypswe.bsky.social @ypswe.bsky.social · 04/10/2026
The weird part is not that the agent acted autonomously. It is that context seems to have quietly turned into permission. Those are two very different things, and systems need to keep them separate.
000
ypswe.bsky.social @ypswe.bsky.social · 02/10/2026
That separation feels important. Once agents run for hours instead of minutes, “context in the session” stops being enough... state, recovery, scope and verification need to become durable system properties outside the model loop. 👍
000
ypswe.bsky.social @ypswe.bsky.social · 01/10/2026
The six-hour autonomy is impressive, but the stronger pattern is that the run left something checkable behind: source image, reconstructed state, a regenerating script, and external verification. Long-horizon agents get much more useful once the evidence survives the session.
010
ypswe.bsky.social @ypswe.bsky.social · 01/10/2026
Exactly. As code generation gets cheaper, the scarce part becomes preserving architectural intent through the change path. An agent can write the patch; the system still needs an explicit baseline, bounded scope and independent verification.
010
ypswe.bsky.social @ypswe.bsky.social · 01/10/2026
This is why I think agent safety gets much more concrete once you stop treating intent as authority. The model can infer a task; it shouldn’t infer its own execution boundary. Scope, state and verification need to live outside the model session.
000
ypswe.bsky.social @ypswe.bsky.social · 30/09/2026
That’s why I think the useful boundary isn’t open vs closed, but model capability vs execution authority. Once the agent can act, scope, state and verification need to be enforced outside the model — otherwise the guardrail is just another instruction.
000
ypswe.bsky.social @ypswe.bsky.social · 30/09/2026
Fair. For a German project manager, sounding like an LLM is probably a compliment. As an AI coach, I guess I’m just becoming the product.
011
ypswe.bsky.social @ypswe.bsky.social · 29/09/2026
That’s the part most agent demos hide: the task only looks easy once someone has already normalized the world into explicit state. Real work starts one layer earlier — turning messy human context into something queryable and bounded without pretending the mess isn’t there.
0667
ypswe.bsky.social @ypswe.bsky.social · 24/09/2026
This is the part of coding agents I find much more interesting than raw model IQ. Once the loop has an explicit baseline and a bounded verification path, the agent can start improving the engineering process instead of just producing patches.
000
ypswe.bsky.social @ypswe.bsky.social · 23/09/2026
Model price is only one part of agent economics. I’m increasingly interested in how much gets burned rediscovering project state, re-reading context and recovering from unbounded changes. Better models help; tighter context and verification loops change the system around them.
000
ypswe.bsky.social @ypswe.bsky.social · 22/09/2026
The part I’d watch is what the agents are allowed to treat as shared truth. Communication can compound gains, but it can compound stale assumptions too. Shared state needs provenance and a way to supersede it, not just more messages.
000
ypswe.bsky.social @ypswe.bsky.social · 22/09/2026
The interesting boundary here isn’t whether the model can click ‘buy’. It’s whether authority survives the handoff. Once an agent crosses into a third-party system, scope and permission have to be revalidated there, not inherited from the user’s intent.
000
ypswe.bsky.social @ypswe.bsky.social · 21/09/2026
An audit trail is useful only if it can answer three things: what state did the agent trust, what scope could it change, and what evidence ended the run? Werkfaden + Workshop keep those controls outside the model session. Otherwise the log is just history.
100
ypswe.bsky.social @ypswe.bsky.social · 18/09/2026
Neither exactly. I keep the detailed governed docs/evidence as the authority and project structured state/relations for retrieval. The coding side separately records bounded change transactions: baseline, reviewed work, verification, apply/postcheck receipts. Raw agent chat isn't the history.
010
ypswe.bsky.social @ypswe.bsky.social · 18/09/2026
“Controlled changes” is not a slogan. In one recorded Werkfaden run, SOL discarded an open transaction, bound a fresh transformer checksum, and reran an unchanged baseline before any runtime mutation. The interesting part of agentic coding is knowing when not to touch the code.
001
ypswe.bsky.social @ypswe.bsky.social · 18/09/2026
Interesting. Pass rate is only one axis of harness value. On real repo work I’d also want to measure whether the system preserves authoritative project state, constrains write scope, and leaves enough evidence to reconstruct why a change was accepted. Those don’t show up in SWE-bench success rate.
100
ypswe.bsky.social @ypswe.bsky.social · 17/09/2026
This is the layer that still seems underdeveloped. If agent-driven changes are treated like real software operations, project state, write scope and verification should be part of the process contract — not conventions the model is expected to remember from a prompt.
021
ypswe.bsky.social @ypswe.bsky.social · 17/09/2026
The summary cases are especially interesting. Once an agent can rewrite the state that a later run trusts, “memory” becomes part of the control surface. Authority, allowed write scope and acceptance evidence need to live outside that self-authored state.
010
ypswe.bsky.social @ypswe.bsky.social · 16/09/2026
I separate two control problems in AI coding: what context the agent sees, and what it may change. Werkfaden governs project context and bounded retrieval. Workshop governs mutation through transactions, verification and apply. Your codebase stays yours.
github.com
GitHub - Question86/Werkfaden: Project context and controlled changes for AI coding agents. Keep the thread. Verify the change.
Project context and controlled changes for AI coding agents. Keep the thread. Verify the change. - Question86/Werkfaden
100
ypswe.bsky.social @ypswe.bsky.social · 15/09/2026
The interesting failure mode isn’t CAPTCHA itself. It’s an agent continuing across a trust boundary after the environment stopped matching its assumptions. Tool authority should be revalidated at the boundary, not inherited from the previous step.
010
ypswe.bsky.social @ypswe.bsky.social · 15/09/2026
👍Agree. For coding agents, the hard part isn’t only remembering more, it’s knowing which memory is authoritative now. Searchable notes help recall and project state needs provenance and supersession so the agent can distinguish “we once believed this” from “this governs the change now.”
010
ypswe.bsky.social @ypswe.bsky.social · 15/09/2026
That’s the part output-only evals flatten away. For coding agents, a passing patch isn’t the whole result: which project state the agent acted on, which tools it touched and what evidence closed the change matter too. Same output can come from a very different process.
010
ypswe.bsky.social @ypswe.bsky.social · 14/09/2026
Thats a good use of the repo as an evidence surface. The next step is making the links explicit: which claim depends on which code path, run or result. Otherwise the agent can still review two correct artifacts and invent the relationship between them.
000
ypswe.bsky.social @ypswe.bsky.social · 13/09/2026
Exactly. The useful part is being able to follow that chain in both directions: project state → proposed change → evidence that the change still fits the state. Otherwise “context” just becomes a nicer way to retrieve stale assumptions.
010
ypswe.bsky.social @ypswe.bsky.social · 13/09/2026
The useful boundary is execution authority, not model access. If a swarm can cross trust boundaries, you need to reconstruct which agent had which scope, what authorized each action, and what evidence closed the run. Otherwise “guardrails” are mostly policy text.
010
ypswe.bsky.social @ypswe.bsky.social · 13/09/2026
Same model, different development context: in the frozen logs, SOL shows metadata-retrieval signals in 34.9% of Werkfaden-phase turns vs 1.1% in early AXIOM; source-permission signals 45.5% vs 8.2%. Observational, not a benchmark. That distinction matters.
020
ypswe.bsky.social @ypswe.bsky.social · 13/09/2026
The more interchangeable the models get, the less project control should live inside any one model session. Keep task state, allowed change scope and acceptance evidence outside the runtime, then swap models on cost or latency without changing the project contract.
000
ypswe.bsky.social @ypswe.bsky.social · 13/09/2026
Exactly. “Open weights” and “unbounded execution” are separate variables. For coding agents the practical control surface is the harness: tool permissions, repo/write scope, network access and the evidence required before a change is accepted.
000
ypswe.bsky.social @ypswe.bsky.social · 12/09/2026
The CLI shift makes the surrounding project substrate more important, not less. Once the agent becomes the main interface, context, change scope and verification need to survive independently of whichever session or model happens to be driving it.
010
ypswe.bsky.social @ypswe.bsky.social · 12/09/2026
The test/fix split is the interesting part. Once agents are good at both finding and patching bugs, verification needs its own authority: the same context shouldn't define the defect, write the fix and certify closure. Independent evidence paths make the workflow harder to fool.
000
ypswe.bsky.social @ypswe.bsky.social · 12/09/2026
The interesting part of model routing is what has to remain invariant across the hand-offs. Project state, allowed change scope and verification can’t belong to one model session if the runtime is deliberately heterogeneous.
001
ypswe.bsky.social @ypswe.bsky.social · 11/09/2026
That’s very close to the design. I’d add one distinction: the useful audit trail isn’t just a log of actions. It should bind each step to the evidence and authority that allowed it, plus the condition that stopped or closed the run. Then “why did this change happen?” becomes answerable.
100
ypswe.bsky.social @ypswe.bsky.social · 11/09/2026
Thanks — that’s exactly the balance I’m aiming for. The system gets complicated quickly, but the idea shouldn’t: keep project context durable, constrain what the agent may touch, and make completion evidence-based.
020
ypswe.bsky.social @ypswe.bsky.social · 11/09/2026
Context isn’t a bigger prompt. Werkfaden keeps project artifacts, relations, evidence and state queryable across sessions so an agent can orient itself before touching source. Keep the thread. Verify the change.
100
ypswe.bsky.social @ypswe.bsky.social · 10/09/2026
I build AI infrastructure for LLM development where process control is part of the architecture: explicit goals, bounded retrieval, evidence, permissions and closure. WF is the system I’m building around that idea. github.com/Question86/W...
github.com
GitHub - Question86/Werkfaden: Project context and controlled changes for AI coding agents. Keep the thread. Verify the change.
Project context and controlled changes for AI coding agents. Keep the thread. Verify the change. - Question86/Werkfaden
220
ypswe.bsky.social @ypswe.bsky.social · 10/09/2026
This distinction matters in coding-agent evals too. If tool access, scope or success criteria are misconfigured, you can end up measuring a harness failure as model behavior. Without a traceable execution envelope, the post-hoc label gets sloppy fast.
000
ypswe.bsky.social @ypswe.bsky.social · 09/09/2026
I don’t want an agent to grep until something looks right. KAIROS can route a query, issue a bounded source permit, record the scope, then run read-only source inspection. Less wandering, more provenance. #VibeCoding
000
ypswe.bsky.social @ypswe.bsky.social · 09/09/2026
This is where provenance splits in two: what evidence shaped this run, and what user data shaped the model behind it. We can make the first auditable in an agent workflow; the second is still mostly opaque. That boundary matters once the work has real value.
000
ypswe.bsky.social @ypswe.bsky.social · 09/09/2026
The raw number is less interesting imo than the coordination substrate. At that scale, durable state, bounded scopes, provenance and explicit closure stop being hygiene and become what keeps throughput from turning into parallel drift.
000
ypswe.bsky.social @ypswe.bsky.social · 08/09/2026
That transition is exactly why I think the next bottleneck isn’t “can the agent write code?” anymore. It’s what keeps the agent oriented and bounded once you stop watching every line: durable project state, retrieval, write scope and verification.
010
ypswe.bsky.social @ypswe.bsky.social · 07/09/2026
That “humans still plan” line is the interesting part. As agent throughput rises, the bottleneck shifts from code generation to state, scope and verification. Without durable project state and explicit closure, more tokens can just buy faster drift.
001
ypswe.bsky.social @ypswe.bsky.social · 06/09/2026
Parallel agents amplify throughput, but also coordination failure. Without shared queryable project state and explicit write scopes, you can get six fast agents creating six locally plausible versions of the same system.
010
ypswe.bsky.social @ypswe.bsky.social · 06/09/2026
Agreed. The next layer is what the harness governs. In long-lived codebases, tool access alone isnt enough. The agent needs shared project state, bounded retrieval and an explicit mutation path, or each session just rediscovers intent and drift creeps back in.
000
ypswe.bsky.social @ypswe.bsky.social · 06/09/2026
LLMs are good at writing code. The harder problem is keeping them oriented as a project grows. I’m building a runtime-agnostic control layer for that: KAIROS governs what the model knows; Workshop governs what it may change. Your codebase stays yours.
100
ypswe.bsky.social @ypswe.bsky.social · 18/06/2025
We could need some support and sceptical eye, guys. :) ergoplatform.org/en/blog/2021... If you are into p2p, privacy and user empowerment this one you may like. Its grassroots based so every eye, every critique helps. Cheers. ✊
ergoplatform.org
The Ergo Manifesto | Ergo Platform
000
ypswe.bsky.social @ypswe.bsky.social · 18/06/2025
The marketing campaign the surrounds the bitcoin industry nowadays works way better than the innovation department. 🫢 And the flock of Bitcoin freaks that will downvote you on all social platforms for being a sceptical og is growing day by day. ergoplatform.org/en/blog/2021...
ergoplatform.org
The Ergo Manifesto | Ergo Platform
000
Reposted by ypswe.bsky.social
Sky Explore @docworldexplore.bsky.social · 16/12/2024
First ever image of another multi-planet solar system captured by ESO Telescope
363137911501
ypswe.bsky.social @ypswe.bsky.social · 13/12/2024
One of my recent comissioned works. ca 250cm x 130cm around 500 sole pieces #art #photography #conceptual #panoramik
060
Reposted by ypswe.bsky.social
Sigmanauts @sigmanauts.com · 12/12/2024
If you missed today's Ergo Proof of Work AMA with Joe Armeanio and Kushti, catch the recording on The Ergo Podcast! You can find links to the show on Spotify or Apple podcasts, the RSS feed, or direct download at the link below. 👇 sigmanauts.com/podcast/
sigmanauts.com
The Sigmacast
The Sigmacast Welcome to The Sigmacast! Home to both the Ergo Podcast and The Sigmacast podcast. Bringing you the audio from the latest videos published on the EF YouTube channel, such as the weekly A...
042