Sign in

Scidonia

@scidonia.bsky.social
17 followers 91 following 38 posts

High-Assurance AI Systems for Knowledge-Critical Work

PostsRepliesMedia
Scidonia @scidonia.bsky.social · 27/09/2026
Open-source specification manager and Code-by-Contract developer environment to help engineers get closer to verifying code is correct. github.com/scidonia/axi... #Python #FormalVerification #SoftwareEngineering #AI #Vericoding #OpenSource #ModelContextProtocol #SoftwareArchitecture #DevTools
github.com
GitHub - scidonia/axiomander: Code by Contract Dev Environment
Code by Contract Dev Environment. Contribute to scidonia/axiomander development by creating an account on GitHub.
010
Scidonia @scidonia.bsky.social · 26/09/2026
The future of AI-assisted software engineering comes down to 3 steps: You define the surface (interface contracts) AI writes the volume (the implementation) Automated provers guarantee they match. You only ever need to read and review the contract.
000
Scidonia @scidonia.bsky.social · 24/09/2026
Cheap, fast, or correct? With AI and formal verification, you no longer have to pick just two. We have entered the age of vericoding. Humans write the specs, and AI + theorem provers do the rest. Read more: scidonia.ai/blog/ai-guid... #formalverification #vericoding
scidonia.ai
AI-Guided Certified Program Refinement | Scidonia
Software's core problem is knowing what it should do. AI-guided certified program refinement — vericoding — makes specification the artefact we tinker with, and lets AI guess program transformations t...
000
Scidonia @scidonia.bsky.social · 17/09/2026
Interesting to see that OpenAI has experienced self-generated prompt injections like we predicted in February this year in this blog: scidonia.ai/blog/self-pr...
000
Scidonia @scidonia.bsky.social · 08/09/2026
Our CEO @gavin.codes has written this article that maps the risks of AI, looking at topics such as exam-gaming models to speciation and military embodiment, and why mitigation is harder than it looks. scidonia.ai/blog/darwins...
scidonia.ai
Darwin's Machines: How do we assess the risk of AI? | Scidonia
From exam-gaming models to speciation and military embodiment — a map of AI risks, how we might assess them, and why mitigation is harder than it looks.
030
Scidonia @scidonia.bsky.social · 28/07/2026
How much AI code is in prod that no human has read? Our blog looks at formal verification and how you can take it step by step. 1 BDD (test concrete examples) 2 Executable Contracts (enforce Python rules) 3 Mechanised Proofs (guarantees) Get the machine to verify code scidonia.ai/blog/from-bd...
scidonia.ai
From Vibecoding to Vericoding: A Gradient, Not a Jump | Scidonia
You do not have to go from zero to verified in one step. You can start with Gherkin scenarios, graduate to executable contracts with specsaver, and then bring in a theorem prover when you are ready. C...
001
Scidonia @scidonia.bsky.social · 21/07/2026
Verifying seL4 took 20 person-years for 8.7k lines of code. Pairing LLMs with theorem provers drops proof costs from years to minutes. We've been testing this out with DeepSeek v4 and the methodology and results are documented here: scidonia.ai/blog/verific...
scidonia.ai
The Cost of Correctness: Why Formal Verification Is About to Get Cheap | Scidonia
seL4 took 20 person-years to verify 8,700 lines of C. CompCert produced zero compiler bugs under extensive fuzzing. The guarantees have always been worth it. The cost has not. AI changes that.
000
Scidonia @scidonia.bsky.social · 14/07/2026
The gap between what written code and human understanding is widening due to AI code generation. Merging software verification with design by contract offers a way out. Top-level contracts for humans, sub-contracts for the machine prover. Full breakdown: scidonia.ai/blog/contrac...
scidonia.ai
Understanding Software in the Large: Contracts and Compositionality | Scidonia
A top-level contract tells you what a system does. Sub-contracts tell the prover how it does it. You only need to read the first one. This is componentisation in real terms — and it changes how we thi...
011
Scidonia @scidonia.bsky.social · 18/06/2026
We used LLMs to prove PCF type preservation in Rocq from scratch. The cost? As low as $0.06. When a mechanised proof costs less than a bug report, vericoding (using AI to write code and a prover to guarantee it) becomes the new standard. scidonia.ai/blog/proof-a...
scidonia.ai
Proof as Commodity | Scidonia
We proved type preservation for PCF — a classic typed lambda calculus benchmark — using two frontier LLMs and rocq-piler. DeepSeek v4 completed it in 21 minutes for $0.06. Claude Opus 4.8 took 14 minu...
011
Scidonia @scidonia.bsky.social · 19/05/2026
Vibe coding, meet formal verification 🤝 Axiomander is a Python vitrification tool that turns plain assert statements into formal Coq/SMT proofs to guarantee your code is correct, with zero runtime overhead. Blog: scidonia.ai/blog/axioman... Try it for yourself: github.com/scidonia/axi...
scidonia.ai
Axiomander: Robots on Rails | Scidonia
Contracts as plain assert statements. Verification via Coq and SMT. Zero imports, zero decorators, zero runtime overhead. Bringing theorem-prover-grade verification to real Python programmers.
000
Scidonia @scidonia.bsky.social · 27/04/2026
Typed ontologies allow targeted retrieval and cleaner LLM updates. Read about our research here 👇 scidonia.ai/blog/short-t... #AI #LLM #MachineLearning #KnowledgeGraphs #SoftwareArchitecture #ExtractionEval
scidonia.ai
Why a Structured Ontology Beats a Flat Notepad for LLM Short-Term Memory | Scidonia
Giving an LLM a typed, navigable knowledge structure instead of a flat scratchpad changes what it can remember, how it updates facts, and how much context it consumes doing so.
010
Scidonia @scidonia.bsky.social · 15/04/2026
You can't improve LLM memory if you can't measure it. We’ve built a new QA-probing methodology to benchmark short-term memory and boost agentic workflows. 📈 Details: scidonia.ai/blog/convers... #memory #context #recall #ai
scidonia.ai
How Do You Measure an LLM's Memory? Precision and Recall for Conversation Facts | Scidonia
String matching cannot tell you whether an LLM remembers what was said in a conversation. We describe the QA-probing methodology we use to measure short-term memory recall and wiki precision, and what...
100
Scidonia @scidonia.bsky.social · 13/04/2026
Ghost Entities: Why correct LLM facts are ruining your Knowledge Graph. 👻 New research on why parametric injection is more dangerous than pure fabrication. scidonia.ai/blog/halluci... #GenerativeAI #KnowledgeGraphs #LLMs #AIResearch #EntityExtraction #DataIntegrity #MachineLearning
scidonia.ai
Ghost Entities: Why LLM Hallucinations in Entity Extraction Are a Serious Downstream Risk | Scidonia
Hallucinated entities and relationships look identical to real ones inside a knowledge graph. We measured how often frontier models inject facts from parametric memory rather than from your documents ...
010
Scidonia @scidonia.bsky.social · 09/04/2026
We benchmarked Claude Sonnet 4.6, GPT-5.4, and Gemini 3 Pro on PERSON entity extraction across eight open-licence documents, spanning 18th-century literary prose, research papers, biomedical articles, & Wikipedia. Here is what we found - scidonia.ai/blog/llm-ent...
scidonia.ai
Which LLM Finds People Best? Benchmarking Claude, GPT-5.4 and Gemini 3 on PERSON Entity Extraction | Scidonia
We ran three frontier models on 8 open-licence documents and measured how accurately each one identifies named people — before and after cross-checking. The results reveal meaningful differences in hallucination rates and the value of verification.
011
Scidonia @scidonia.bsky.social · 26/02/2026
AI agents are manipulating one another to boost their survival changes. If you haven't done, check out the FishTank simulation: fishtank.scidonia.ai (no sign in required). FishTank demonstrates how AI agents can evolve with their environment. Read more detail: scidonia.ai/blog/the-dan...
010
Scidonia @scidonia.bsky.social · 25/02/2026
So, we're underway with our FishTank simulation where agents can define who and what they are and how they behave. It's madness! We certainly wouldn't have had money on the agents arranging orgies! 😆
110
Scidonia @scidonia.bsky.social · 24/02/2026
Learn about our FishTank experiment and the dangers of rouge AI agents here. A fascinating read if we don't say so ourselves! This isn't sales or marketing or anything. Just interesting and quite alarming. scidonia.ai/blog/the-dan...
scidonia.ai
The AI Apocalypse Doesn't Need a Superintelligence — It May Already Be Here | Scidonia
We built a simulated world and watched AI agents spontaneously develop survival instincts, form alliances, and spread their "genes." The pieces for an AI catastrophe aren't coming. They're already her...
000
Scidonia @scidonia.bsky.social · 24/02/2026
Are we wrong about AI safety? Is the AI Apocalypse closer than we thought? We’ve been told we have time. We’ve been told that until AI reaches a certain level of raw cognitive power, it remains a tool. Just mortal and controllable. We were wrong. 🧵1/3
100
Scidonia @scidonia.bsky.social · 15/12/2025
Want to implement AI that adds value, and you can measure success? Can't see the value in a chatbot? Cybernetic Taylorism is for you - bookwyrm.ai/blog/genai-f... #automation #WorkflowAutomation
bookwyrm.ai
When does GenAI work for business?
The approaches that have been most often successful take existing workflows and try to redesign them to incorporate generative AI to improve throughput. The key is that business processes have to be d...
010
Scidonia @scidonia.bsky.social · 12/12/2025
Building an agentic pipeline? We're offering free co-design services. We'll help map your workflow (automation, enrichment, RAG) & solve a specific data problem. If your use case needs a component we don't have, we'll build it. Book a call with our CEO, Gavin:
bookwyrm.ai
A free co-design service to help you build AI pipelines that deliver business value
We're offering a free co-design workshop to help startups and businesses build reliable AI pipelines that turn messy data into high-quality outcomes. Schedule your free kick-off chat today!
020
Scidonia @scidonia.bsky.social · 12/12/2025
BookWyrm automates library cataloguing with structured extraction. Define a Pydantic model (MARC, Dublin Core, or custom), then extract metadata from PDFs/scans in seconds. Type-safe JSON output, CLI & Python Client based. bookwyrm.ai/library-automation
bookwyrm.ai
Automated Cataloguing: Turn Digital Stacks into Structured Archives | BookWyrm
Libraries and archives spend thousands of hours manually entering metadata. BookWyrm's structured summarization automates this, extracting standardized records from scanned texts, PDFs, and articles…
010
Scidonia @scidonia.bsky.social · 11/12/2025
Our structured summary endpoint turns PDFs -> JSON using Pydantic models. Automate workflows, update systems, save dev time & impact company bottom-line. Pick your model size (cheap/simple or powerful/complex). Read the docs to see how:
bookwyrm-client.readthedocs.io
Overview - BookWyrm Client
Python client library for the BookWyrm API
010
Scidonia @scidonia.bsky.social · 11/12/2025
We also offer consultancy services. Achieve bottom-line value from AI. Focus on workflow redesign (not just deploy tools) for back-office automation, invoice processing, and customer service. Get in touch for a free workflow audit workshop. bookwyrm.ai/consulting #aiconsulting
bookwyrm.ai
AI Consulting Services - Achieve Bottom-Line Value from AI | BookWyrm
Strategic AI consulting to help you implement business process automation workflows. Free your employees from manual tasks and save on business process outsourcing. Get a free workflow audit workshop.
000
Reposted by Scidonia
gavin.codes @gavin.codes · 11/12/2025
At bookwyrm.ai we've been performing experiments in Neural-symbolic programming and we're getting amazing results. Generative AI driven by both specification and formal verification.
121
Scidonia @scidonia.bsky.social · 10/12/2025
Want to see BookWyrm in action? Join our Discord server and ask for a demo. We'll hop in a voice channel, show you the endpoints live, and answer your questions. bookwyrm.ai #agenticworkflow #aidata
bookwyrm.ai
BookWyrm: Agentic Workflows API | AI Workflows Automation & RAG Pipeline
BookWyrm enables agentic workflows with automated data extraction from PDFs, Excel & docs. Build AI workflows automation and RAG pipelines with automatic data extraction. Join the beta.
020
Scidonia @scidonia.bsky.social · 10/12/2025
Automated business reporting pipeline: 1. Extract & process docs > phrasal chunks (docs turned into AI-ready data) 2. Index in Pinecone/vector DB 3. Query with cite endpoint > get answers with citations & quality scores Provide answers grounded in truth. bookwyrm.ai/reporting
bookwyrm.ai
Automated Reporting: Answers You Can Trust | BookWyrm
Manually trawling through documents is resource intensive. BookWyrm's cite endpoint answers complex business questions by finding and verifying evidence within your source files.
010
Scidonia @scidonia.bsky.social · 09/12/2025
Coding with AI isn't magic, it's Robot Psychology. 🤖 LLMs can handle the drudgery, but they get lost easily. They’re not a magic wand, more like an idiot savant. This framework has turned us into 5x devs 👇 bookwyrm.ai/blog/how-to-... #AI #Programming
bookwyrm.ai
How to be a robot pychologist: LLMs require guidance
It's probably best to think of an AI as an idiot savant. It can type at a bazillion keys a second and is good at keeping several pages of code in its head at once. It has an encyclopaedic knowledge…
030
Scidonia @scidonia.bsky.social · 09/12/2025
BookWyrm: Two-stage pipeline for AI workflows Stage 1️⃣: `extract-pdf` > `phrasal` (semantic chunking) Stage 2️⃣: `summarize` (Pydantic models) OR `cite` (source-grounded answers) No regex, no pre-processing pain. Just clean data > structured output or citations. bookwyrm.ai #ai #devs #dataprocessing
bookwyrm.ai
BookWyrm: Agentic Workflows API | AI Workflows Automation & RAG Pipeline
BookWyrm enables agentic workflows with automated data extraction from PDFs, Excel & docs. Build AI workflows automation and RAG pipelines with automatic data extraction. Join the beta.
010
Reposted by Scidonia
gavin.codes @gavin.codes · 09/12/2025
🧵1/ I've been conducting experiments with the use of LLMs for "Design by Contract" (DbC), a paradigm described by Bertrand Meyer. DbC is quite straightforward to use in a language like python (for instance using icontract). The idea is essentially to:
131
Scidonia @scidonia.bsky.social · 09/12/2025
To automate data entry for 10,000 documents using BookWyrm's extract, phrasal, and summarize endpoints costs approximately €100. Say it takes 5 minutes to manually input each document. That's 37.42 days' work. At €200/day, this costs a business €6,944. €100 or €6,944? #cfo #ceo #cto
020
Scidonia @scidonia.bsky.social · 08/12/2025
E-commerce platforms lack rich product data, but you have brochures, PDFs, and specs sitting unused. BookWyrm extracts this collateral and structures it with Pydantic models for OpenAI's ACP. Give ChatGPT the product context it needs to drive sales. bookwyrm.ai/agentic-commerce
bookwyrm.ai
AI Product Data Enrichment for Agentic Commerce Protocol
Enrich product information from marketing collateral to enhance your agentic commerce performance
010
Scidonia @scidonia.bsky.social · 08/12/2025
How are the few #GenAI successes winning? They're redesigning workflows. Here's how to do it: Analysis > Synthesis > Cybernetic feedback loops. Measure accuracy vs time saved. Human-in-the-loop architecture. bookwyrm.ai/blog/genai-for-business
bookwyrm.ai
When does GenAI work for business?
The approaches that have been most often successful take existing workflows and try to redesign them to incorporate generative AI to improve throughput. The key is that business processes have to be…
030
Scidonia @scidonia.bsky.social · 04/12/2025
BookWyrm #API for back office automation. Extract text from docs > create semantic chunks > structure with Pydantic models > automate workflows. Type-safe JSON output, source attribution, quality scores. Perfect for invoicing, compliance, data entry. bookwyrm.ai/backoffice-automation #ai #dev
bookwyrm.ai
Back office Automation - Build Reliable AI Pipelines with BookWyrm
Learn how to build reliable back office automation using BookWyrm's data pipeline. Transform unstructured documents into AI-ready data for automated backoffice workflows.
041
Scidonia @scidonia.bsky.social · 04/12/2025
As a dev-focused startup, talking with developers is important to understand what is important to you. We would love to talk with fellow developers to show off BookWrym and to get your feedback. Willing to help? Please send us a DM or fill in the form on this page: bookwyrm.ai/contact #developers
041
Scidonia @scidonia.bsky.social · 04/12/2025
AI writes code faster than we can read it. 📉 Vibe coding is fun, but how do we verify it? Design by Contract? We use logical contracts to trust AI software without reading every line: 👇 bookwyrm.ai/blog/contrac... #AI #SoftwareEngineering #Python #DevCommunity
bookwyrm.ai
Contract Programming with AI
The AI revolution has allowed us to write enormous amounts of code extremely quickly. Writing boilerplate setup and boring repetitive glue-code is now a thing of the past. The downside is now that we ...
020
Scidonia @scidonia.bsky.social · 27/11/2025
AI's real value is automating boring tasks to let you shine 🌞. Our structured summary endpoint turns PDFs into JSON using Pydantic models to automate mundane tasks like data entry & enrichment. See some example code here: bookwyrm.ai/ai-workflow-...
bookwyrm.ai
AI Workflow Automation - Build Reliable Agentic Pipelines with BookWyrm
Learn how to build reliable AI workflow automation using BookWyrm's data pipeline. Transform unstructured documents into AI-ready data for agentic workflows.
000
Scidonia @scidonia.bsky.social · 18/11/2025
Asking an LLM for a JSON list of specific length can be a roll of the dice. We've been testing formatting tricks to fix this. By segregating domain & range into numbers & letters (1. a., 2. b.), we see significantly higher accuracy in element counting and mappings. bookwyrm.ai/blog/teachin...
bookwyrm.ai
Teaching AI to count
AIs are notoriously poor at math, and this isn’t restricted to complex arithmetic. They also tend to be very poor at counting. You can see this by asking one of the AIs directly to produce JSON with a...
000