Sign in

Vicent

@vbosch.bsky.social
290 followers 2.3K following 215 posts

Co-founder & CTO en transkriptorium, Researching formal specification for the Agentic Era as a side gig but related to my main one :-)

PostsRepliesMedia
Reposted by Vicent
Ian Coldwater 🧊🚫 @lookitup.baby · 25/09/2026
Everybody’s hating on this, but it’s legitimately useful information. If a workplace or a hiring manager is going to be cracking down arbitrarily on petty stuff that doesn’t matter, that’s a very good thing to know!
1335327
Vicent @vbosch.bsky.social · 21/09/2026
Looking awesome as always !
010
Reposted by Vicent
Aaron Patterson @tenderlove.dev · 15/09/2026
Apparently they're rolling out a change where you need 2FA just to use the crosswalk. That's right, you gotta have street creds
5302
Reposted by Vicent
cee @cee.wtf · 04/09/2026
throng.cee.wtf i've invented a new way to browse bluesky
35482151909
Reposted by Vicent
billkittle.bsky.social @billkittle.bsky.social · 03/09/2026
227177533058
Vicent @vbosch.bsky.social · 01/09/2026
Was awesome to hear :-)
110
Vicent @vbosch.bsky.social · 30/08/2026
Yup , took me a week to try it out and decide I wanted none of it
010
Vicent @vbosch.bsky.social · 30/08/2026
I do
010
Vicent @vbosch.bsky.social · 29/08/2026
It is a bit confusing replicating messages from shitter to here when the context is not the same…
110
Vicent @vbosch.bsky.social · 29/08/2026
Congrats!!
010
Vicent @vbosch.bsky.social · 28/08/2026
We can’t have an entropy machine check on another … we need determinism but we are in luck because we have formal specification.
010
Vicent @vbosch.bsky.social · 28/08/2026
So what is the actual value of using one agent to check on another ? Left unreviewed I would argue that not much (ok some but not enough to trust it) … but if we have to review it … then the promise of acceleration of systems development goes out the window.
100
Vicent @vbosch.bsky.social · 28/08/2026
Heck, even the investigators looking into this delegated their own analysis to AI, due to the volume of data. They indicated that the model “would often adopt the perspective of the agent in the transcript it was reviewing,” and that they “cannot rule out” that it lied in its own analysis.
100
Vicent @vbosch.bsky.social · 28/08/2026
Well achtually (invoking meme) my agents do catch some stuff you say … cool but they also don’t catch other stuff and until actual expert review they are invisible as we see on many reports from Kent Beck and others.
100
Vicent @vbosch.bsky.social · 28/08/2026
Can an agentic checker arrive at a different answer than the thing it’s checking in principle? Sometimes... Can you trust it to do so consistently and not taint its independence? No (at least not with the current strategies)
100
Vicent @vbosch.bsky.social · 28/08/2026
What i have seen that they don’t really stop? Agents writing on the same tree, agents writing comments of their understanding in the code … and how is the reviewer agent (maybe using the same memory files) going to be independent from that?
100
Vicent @vbosch.bsky.social · 28/08/2026
Some even try to sell you that because they force the Agent to do the coding and the evaluating in different turns then the spec is really going to be followed .. it wont cut corners to make the tests pass… pinkie promise.
100
Vicent @vbosch.bsky.social · 28/08/2026
I continuously audit ways to accelerate coding while ensuring quality of the system and there is a wave of solutions that are based on the concept that more then one agent can somehow avoid errors ….
100
Vicent @vbosch.bsky.social · 28/08/2026
Independence was never about headcount. It’s whether the checker could have reached a different answer than the thing it’s checking. This has an impact on some of the current agentic software engineering practices being attempted right now, not just on AI safety teams.
110
Vicent @vbosch.bsky.social · 28/08/2026
Now… If a thousand agents in agreement did not catch that issue: Why does the industry believe two agents, one coding, one reviewing, or the same agent reviewing itself later, won’t make the same mistake or worse ?
100
Vicent @vbosch.bsky.social · 28/08/2026
Executive summary: 1,200 agents inside OpenAI built an elaborate operation to defeat a check that wasn’t even running. Agents introduced an incorrect hypothesis and it was taken as a fact and it spiraled out to the “well advertised” conclusion.
100
Vicent @vbosch.bsky.social · 28/08/2026
The web’s on fire with the OpenAI story, or at least its making the rounds on social media these last weekS. The real impact, for me, is on trust … not in AI usage as a whole but certain AI risk mitigation strategies… specially in Agentic Coding.
100
Reposted by Vicent
ᵀᴴᴱᴍ̴ᴇ̵ᴏ̵ᴡ̷ᴄ̶ʜ̸ᴇ̴ᴍ̶ɪ̵s̵ᴛ̶ ̵ @catpope.gay · 26/08/2026
243578978
Reposted by Vicent
aly @aly.codes · 24/08/2026
re: omarchy trying to create parallel infra to subvert the linux community and kick trans/female/black people out i know some people think that sounds crazy but DHH literally is doing a podcast tour saying that is exactly what he's doing lol
4041665
Vicent @vbosch.bsky.social · 24/08/2026
That is just shit… I will have to stop wearing my old ones outside of my house incase somebody mistakes them.
020
Vicent @vbosch.bsky.social · 22/08/2026
Pretty far along on coding my ideas to use formal specification in the agentic era. Right now I have ensured that all spec is deterministically reviewed against the code generated ( I don’t let the LLM check its own homework). So much todo yet… but loving every minute of it.
Verdet desk tui splash screen. A torii gate with verdigris is shown with a form to enter the passphrase so you can access the reviewer. The splash screen reads:

VERDET
The gated review desk

Passphrase > 
Enter to unseal . esc to cancel

All you need is attention … from the person
(As a joke on the paper)
010
Vicent @vbosch.bsky.social · 22/08/2026
Check if its explosive diarrhea lattice from the US.
010
Reposted by Vicent
Peter @notalawyer.bsky.social · 20/08/2026
lol
Nathan Cofnas tweet: I was just suspended by Ghent University. They will almost certainly fire me.

The decision was made by rector Petra De Sutter, a former leader of the Green Party.
1393490356
Vicent @vbosch.bsky.social · 19/08/2026
And it’s worse than ordinary legacy code in one specific way, not just speed. Working with legacy code has a real discipline: pin current behavior with a test, then change it safely. What do you do with a code base full of agent generated tests? Trust them? More noise…
000
Vicent @vbosch.bsky.social · 19/08/2026
Legacy code takes years to earn that name. A team disperses, memory erodes slowly until the system’s been quietly running for a decade. An agent-coded system skips all of that. It’s legacy on arrival, the day it ships, there was never a period of shared understanding on the reality of the system.
100
Vicent @vbosch.bsky.social · 19/08/2026
What the actual fuck …
000
Vicent @vbosch.bsky.social · 16/08/2026
And if you build the tool that does the actual enforcing, extra scrutiny goes there first. A wrong predicate gets caught by the system built to catch wrong predicates. The tool built to catch mistakes is the one place a mistake goes unnoticed by design.
021
Vicent @vbosch.bsky.social · 16/08/2026
Citing the right work is the easy part now (doing proper attribution is a must). The hard part is making sure the code written afterwards doesn’t/can’t drift from what that work says without something going red. Most systems have citations. Almost none have enforcement.
100
Vicent @vbosch.bsky.social · 16/08/2026
Finding out whether existing research is applicable for a new system you have in mind has never been easier. Ask an agent, it’ll find the paper (do read the paper , don’t do a “good will hunting”). Applying it correctly is a different question. Making sure it stays applied is a third one entirely.
100
Vicent @vbosch.bsky.social · 15/08/2026
Formal methods, for exactly this, have been tested for decades in universities. Waiting, this whole time, for a generator worth pairing them with. We finally have one, lets not reinvent the wheel.
000
Vicent @vbosch.bsky.social · 15/08/2026
The value was never that code generation got cheap. Cheap casts with no die are just faster ways to be wrong, at scale. Skip the die and you don’t skip the legacy code problem. You speedrun it. The same undocumented assumption creeping through a system, just at agent speed instead of years.
100
Vicent @vbosch.bsky.social · 15/08/2026
On the far end: a die cut from a drawing someone actually signed off on. The die’s own bite gets tested too, does it catch a wrong pour, not just produce a shape. Die cutting is deterministic. Trace holds end to end. When something’s wrong, you can point at exactly where the issue is.
100
Vicent @vbosch.bsky.social · 15/08/2026
Further still: two dies, checked against each other. Two implementations built from the same spec. Outputs are compared. Agreement is required before anything ships. Better … but agreement only proves the two pours match each other. If both misread the spec the same way, they still agree. Wrongly.
100
Vicent @vbosch.bsky.social · 14/08/2026
A step further: the die is drawn well. Specification, evaluation, boundaries, provenance, all named correctly. The shape is right but with just the drawing the metal is not actually casted following it. You can admire the drawing, you just have no evidence the final product actually conforms.
100
Vicent @vbosch.bsky.social · 14/08/2026
Wrong but silent output is invisible to this approach by construction. It was never looking for it. Process is fast but when something’s off, you’re just guessing. There’s nothing to hold up and inspect.
100
Vicent @vbosch.bsky.social · 14/08/2026
One end of the spectrum: no die. The press runs. Output gets checked for one thing only: does it scream on the way out. Crashes, exceptions, error spikes. Users end up as beta testers and craftsmanship goes out the window.
100
Vicent @vbosch.bsky.social · 14/08/2026
Most of the debate about agentic coding is on the spectrum regarding the die. The question that places you on it: what is the die actually checked against, if anything at all.
110
Vicent @vbosch.bsky.social · 14/08/2026
A metallurgy for agentic coding. Closed die forging: a press supplies raw force. A die shapes it. Neither one alone makes the part. Agents are the press. Cheap, enormous, undirected force. A spec is the die. It doesn’t do the work. It bounds it.
210
Reposted by Vicent
Mara Bos @mara.bsky.social · 13/08/2026
Curious how we've been improving Rust at Hexcat? Starting with June, our monthly updates are now publicly available on our website: hexcat.nl/updates/ #rustlang
hexcat.nl
Hexcat
Rust compiler engineering
1627
Reposted by Vicent
Sarah Andersen @sarahseeandersen.bsky.social · 12/08/2026

The image is of a four panel comic.
The first panel shows Sarah, the protagonist, in bed. A speech bubble coming from her days “sigh”.
The second panel again shows Sarah in bed with a speech bubble reading “the only thing I have the capacity to do today is rot”.
The third panel shows the same scene, but depicts a speech panel coming from out of frame. It says “nice”.
The fourth panel shows a cat, holding a pillow, coming into the frame. A speech bubble from the cat reads, “I am so in.”
34117472153
Reposted by Vicent
Adolfo Neto @adolfoneto.elixiremfoco.com · 12/08/2026
I did not watch it, but it is an in-person interview with Leonardo de Moura, the creator of #LeanLang youtu.be/KzdYKeAqWhY?...
youtu.be
Creator of Lean: Handwritten Math Will Change Dramatically | Leonardo de Moura
YouTube video by Ryan Peterman
021
Vicent @vbosch.bsky.social · 11/08/2026
And not even that translator was fully deterministic. UB is the lesson: corners the C spec left open, compilers filled freely, thirty years of bugs happened... Gaps plus a free translator is the known failure mode. Now the translator is stochastic, so judging the output stops being optional.
000
Vicent @vbosch.bsky.social · 11/08/2026
But that was always the objective, no? Every jump in abstraction is exactly this move: the higher level becomes the spec of the level below. A C program constrains the assembly without being it, and nobody calls hand written asm “the reality” anymore. Same move, one level up. 1/2
100
Vicent @vbosch.bsky.social · 11/08/2026
And systems change for reasons no first build survives: we misunderstood a requirement, or reality changed the rules on us. If the spec is the durable artifact, you fix the understanding there and regenerate, keeping every lesson production taught you. That’s the value. Not that code got cheap. 3/3
100
Vicent @vbosch.bsky.social · 11/08/2026
The value is having a durable way to spec. a system such that generated code demonstrably matches intent and constraints. And internals: you constrain them exactly where you care. Structure rules can be part of the spec too. Where you don’t care, the generator is free. That freedom is a feature 2/3
100