Sign in

Schwentker

@schwentker.sandboxlabs.ai
434 followers 330 following 582 posts

Builder & AI education 🏗️ @at_proto #mpp 6 series #BlueskyCommunityVoices podcast w/ @bluesky pre-launch. @BlockchainU founder. xPayPal 📍 Oakland

PostsRepliesMedia
Schwentker @schwentker.sandboxlabs.ai · 09/10/2026
5/5 The sound stack: phone line (@twilio), ears (streaming STT, e.g. @DeepgramAI ), brain ( @cresta Conductor), voice ( @elevenlabs.io ). WebRTC via @LiveKit can save ~300ms. Past ~300ms, silence feels unnatural. So: when you pause mid-sentence, how does a machine know you're thinking, not done?
000
Schwentker @schwentker.sandboxlabs.ai · 09/10/2026
4/5 Next for Veronica: plan your team’s night out. 12 SF startup events in. Plain code applies the hard rules ($40, Thu, fits 8, near SoMa) & 3 survive. Natali’s Jev classifier ranks them, picking only from that list, so it can’t invent an event. Code filters. Jev picks.
100
Schwentker @schwentker.sandboxlabs.ai · 09/10/2026
3/5 The payments plan: an AI agent w/o a spend policy books a kayak trip & gets charged 10x in pathUSD on Tempo.xyz testnet. "Veronica's" refund tool is a verb, not a pen: refund_overcharge(charge_id), not send(to, amount). The model can't invent a payee.
100
Schwentker @schwentker.sandboxlabs.ai · 09/10/2026
2/5 "Veronica" answers the unBoring Club line. She checks who's calling, finds the charge, files a dispute w/ a ref #, & hands off to a human if needed. Built on Cresta Conductor: 1 umbrella agent that spins up sub-agents (Verify, Charges, Refund, Close). Agents building agents. ☯️
100
Schwentker @schwentker.sandboxlabs.ai · 09/10/2026
1/5 A caller says "you charged me 10x." What has to happen in the next second? Natalie Wong & I built "Veronica" at @cresta's Voice Mode On hackathon #SFTechWeek. 1 call, layer by layer 🧵 www.linkedin.com/pulse/now-so...
linkedin.com
And Now for Something Completely Audible: Inside One Voice-Agent Call ★
A caller says "you charged me 10x." What has to happen in the next second? At Cresta's Voice Mode On hackathon (SF Tech Week), Natali Wong & I built Veronica, the Anti-Boring Club's concierge, on Cres...
100
Schwentker @schwentker.sandboxlabs.ai · 01/10/2026
7/7 Jevathon afterparty on the @coderabbitai roof deck: DJ Lidia, the Bay Bridge & @HKrackDev on sax, solo, w/ an "Ode to Jev." Thx to @typesafeai & @AICollectiveCo & everyone who built w/ Jev.
000
Schwentker @schwentker.sandboxlabs.ai · 01/10/2026
6/7 "Jev never plays notes; it only decides." Drag paintings onto a staff & a live AI band plays them. Beethoven by @Sanjayyb7 & @VismithaNaryan was the most delightful build of the day. Jev conducts every bar (~2.7s). x.com/Sanjayyb7/st...
x.com
Sanjay (@Sanjayyb7) on X
Turn the sound on 🎶 Built Beethoven 🎻 using Jev (@typesafeai) at the @coderabbitai hackathon with @VismithaNaryan Drop paintings on a score, a live AI band plays them. Jev basically decides the vol...
101
Schwentker @schwentker.sandboxlabs.ai · 01/10/2026
5/7 "Optimization is already good with the numbers, but it misses the meaning embedded in how people describe their requests," wrote @tanishkgovil. His Jev-fed hospital elevators cut emergency waits from 46s to 13s across 16 held-out scenarios (self-reported).
100
Schwentker @schwentker.sandboxlabs.ai · 01/10/2026
4/7 "You have to assume everything you build on top of might be broken, and so you have to design around that." Brailly (@0xA1ejandro @MartinPulitano @Julian_Irusta04 @nicopujia ) reorders web pages for blind readers by task, w/ Jev deciding when a page update interrupts. x.com/0xA1ejandro/...
x.com
Alejandro Gonzalez (@0xA1ejandro) on X
This weekend I did something cool! For the @typesafeai and @AICollectiveCo hackathon we made Brailly, a way to read the web in braille with the power of Jev! Thanks to @nicopujia @MartinPulitano and @...
210
Schwentker @schwentker.sandboxlabs.ai · 01/10/2026
3/7 "Optimization is already good with the numbers, but it misses the meaning embedded in how people describe their requests," wrote @tanishkgovil. His Jev-fed hospital elevators cut emergency waits from 46s to 13s across 16 held-out scenarios (self-reported).
100
Schwentker @schwentker.sandboxlabs.ai · 01/10/2026
2/7 @TechCrunch cited Jev Sentinel, built at Jevathon by @shapor His repo has results from 53,870 Hugging Face incident payloads (98.3% flagged). His caveat: hackathon project, not production code, ymmv. github.com/shapor/jev-s...
github.com
GitHub - shapor/jev-sentinel
Contribute to shapor/jev-sentinel development by creating an account on GitHub.
210
Schwentker @schwentker.sandboxlabs.ai · 01/10/2026
1/7 OpenAI announced a Decisions API that @TechCrunch says looks a lot like Jev. Days earlier, 7 teams at the @typesafeai Jevathon @coderabbitai, working apart, built like the same thing: a checkpoint inspecting an agent's action b4 it runs. Part 2: www.linkedin.com/pulse/jevath...
linkedin.com
Jevathon Hacks That Deserve Encores
More standout demos from Jev's first community hackathon at CodeRabbit in San Francisco Part 1: 12 Judges Judging Jev covered the top three finishers and two honorable mentions from Jevathon, held Sep...
100
Schwentker @schwentker.sandboxlabs.ai · 27/09/2026
6/6 @typesafeai.bsky.social grew Jev on synthetic data into a System 1 model: fast, typed judgment, no prose. Image input is next on the public roadmap. After that, maybe a System 2? Hackers are chomping at the bit. When judgment costs fractions of a cent, what's left to decide slowly?
000
Schwentker @schwentker.sandboxlabs.ai · 27/09/2026
5/6 Thx to organizers @AICollectiveCo, @AJs_AI , @roangws, @chappyasel, My Luu, plus host @coderabbitai.bsky.social Shoutout to its intern trio behind Revolution, 1st place: a wildlife ecosystem sim running ~4,000 Jev decisions for ~$1.30.
100
Schwentker @schwentker.sandboxlabs.ai · 27/09/2026
4/6 "We woke up a night later with 1,500 pull requests open, and not enough humans in our company to press the button." @HKrackDev of @coderabbitai.bsky.social on unleashing background agents on a monorepo. Answer: Triage ranks agent PRs by risk & review-readiness.
100
Schwentker @schwentker.sandboxlabs.ai · 27/09/2026
3/6 Agents tend to think it's more expensive than it is." Alex Warren ( @exrhizo.bsky.social ) of @typesafeai.bsky.social re: Jev pricing. 720 states x 30 Qs? Likely <10¢. Tip: Jev always answers, so expecting a non-answer = type error. Add a "not applicable" choice instead.
100
Schwentker @schwentker.sandboxlabs.ai · 27/09/2026
2/6 "If it's vibes, now we're getting somewhere." @allietheicon.com of @typesafeai.bsky.social on Jev's sweet spot: pure numbers belong on a CPU; semantic judgment belongs in Jev. Btw, Jev isn't agentic. Does nothing until code puts it in the loop. ( w/ @stephencouncil )
100
Schwentker @schwentker.sandboxlabs.ai · 27/09/2026
1/6 Twelve Judges Judging Jev: field notes from Jevathon @coderabbitai.bsky.social in SF w/ @typesafeai.bsky.social & @AICollectiveCo . Jev returns typed decisions, not text. Wolves, dog sculptures, stroke triage & a GitHub roast, all w/ Jev inside. www.linkedin.com/pulse/12-jud...
linkedin.com
12 Judges Judging Jev
Field notes from Jevathon, Jev's first community hackathon, at CodeRabbit in San Francisco On Saturday, September 26, eleven days after TypeSafe AI left stealth with $40M in seed funding and Jev as it...
110
Schwentker @schwentker.sandboxlabs.ai · 24/09/2026
5/5 My bet for a @StripeDev meetup a year from now: shopping agents, research agents & seller agents buying & selling over MPP w/ stablecoins. We'll be asking what they bought, what they paid & who gave them permission.
000
Schwentker @schwentker.sandboxlabs.ai · 24/09/2026
4/5 “It's enforced at the infra layer.” AWS's Chethan Shriyan on agent spending limits. Hasan Tariq's demo wallet held nearly $1M in testnet funds. The agent's session budget? 15 cents. Wallet balance isn't spending permission.
100
Schwentker @schwentker.sandboxlabs.ai · 24/09/2026
3/5 “It's like you do your homework & then you rank yourself & grade yourself.” @DeividasMat on one AI doing all the shopping jobs. @m11labsai's OpenMarket splits them: 5 sellers, 1 buyer & a referee checking the claims.
100
Schwentker @schwentker.sandboxlabs.ai · 24/09/2026
2/5 Lunch w/ @collision. A meme-filled one-pager for Patrick. Weeks later: Stripe.directory. @MbyM on making businesses searchable for people & agents, down to the CLI. “I'll do it today.” No Q3 2027 backlog.
100
Schwentker @schwentker.sandboxlabs.ai · 24/09/2026
2/5 Lunch w/ John Collison. A meme-filled one-pager for Patrick. Weeks later: Stripe.directory. Stripe's Mark Moriarty on making businesses searchable for people & agents, down to the CLI. “I'll do it today.” No Q3 2027 backlog.
100
Schwentker @schwentker.sandboxlabs.ai · 24/09/2026
1/5 @StripeDev meetups are back in SF. A lunch idea shipped in weeks, agents checking product claims & wallets w/ spending limits. My notes from the evening: www.linkedin.com/pulse/agenti...
linkedin.com
Agentic Referees, One-Pagers & Wallets on a Leash
Field notes from Stripe's developer meetup return in San Francisco Before the talks began at Stripe, I complimented a stranger's hoodie and ended up telling him about 2010. Back then I ran the Twitter...
100
Schwentker @schwentker.sandboxlabs.ai · 20/09/2026
7/7 @sofiiiiiasz emcee: On how AG-UI first got shown at a WorkOS meetup w/ a 5 min demo slot, & where it went from there: "You never know where it'll lead you." Good reminder that these places especially in SF 🌉 are where standards often start Thx hosts CopilotKit Manufact @workos.bsky.social
010
Schwentker @schwentker.sandboxlabs.ai · 20/09/2026
6/7 Andrew Khadder, Manufact Ran his whole event ops thru ChatGPT Voice. T-shirt orders, allergy lists, Slack, waitlist approvals, expense report. "I didn't need to build out a whole dashboard interface" The dashboard didn't vanish. It got decomposed.
100
Schwentker @schwentker.sandboxlabs.ai · 20/09/2026
5/7 Tyler Slaton, @CopilotKit Best framing of the alphabet soup all night: "AG-UI is basically like you give all of these tools as Lego blocks" & the agent decides which blocks, in what order. The frontend isn't a screen the agent talks about. It's part of the agent's operating environment.
100
Schwentker @schwentker.sandboxlabs.ai · 20/09/2026
4/7 @idosal1 co-creator of MCP Apps & AG-UI, live from Tel Aviv @ some ungodly hour Planning a trip = 10 tabs. But you don't need 10 whole apps, you need a room selector, a map, an approval card. "You break it down to the atoms." An app stops being a destination & becomes a set of capabilities.
101
Schwentker @schwentker.sandboxlabs.ai · 20/09/2026
3/7 @tobinsouth.bsky.social The supply chain problem MCP has that APIs don't: "you can just change the tool list at any time" @tyler_cpk's analogy harmonized: "imagine if your Node packages could change under you w/ no version pin. That's roughly what we've accepted."
110
Schwentker @schwentker.sandboxlabs.ai · 20/09/2026
2/7 @tobinsouth.bsky.social, @Anthropic Most uncomfortable idea of the night: (Claude Code) "auto mode reduces risks over humans who generally just click yes" Asking permission at every step looks safer. But people stop reading by prompt 20. User focus moves to budgets, scopes & thresholds.
110
Schwentker @schwentker.sandboxlabs.ai · 20/09/2026
MCP Apps Night at @WorkOS SF office hosted by @CopilotKit. Went to learn about agentic interfaces. Left thinking -- what is an app when a user can invoke one of its capabilities without ever loading the whole thing. great quotes + a few thoughts www.linkedin.com/pulse/mcp-ap... 1/7
linkedin.com
That MCP Apps Night
Much of my beloved experience w/ software entails clicking from tab to tab to tab. But, what if somethings like going to fundamentally change about browser experience? At MCP Apps Night, hosted by Cop...
110
Schwentker @schwentker.sandboxlabs.ai · 19/09/2026
How to TypeSafe ... or ... Just what is Jev? www.linkedin.com/pulse/jev-do...
linkedin.com
Jev: doubt, measured
A model that says how sure it is Jev, as in Jevons Paradox, TypeSafe's first model, answers set questions & reports its own confidence. What follows is an attempt to describe what that really changes.
000
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
12/12 Close Ten vantage points, one fascinating invention: answers that arrive revealing their own uncertainty. This is not the stuff that prompts are made on. It lives inside software, in agents, where programs decide & read paragraphs to do it. Taught eloquence, never doubt. docs.typesafe.ai
docs.typesafe.ai
Introduction - TypeSafe AI
Jev is TypeSafe's flagship model and the first System One model. Send state and typed questions; get structured answers your code can use directly.
000
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
11/12 Skeptics Jev from @typesafeai.bsky.social is a classifier. Zero-shot, many criteria per call, calibrated. Typed output, logprobs, temp scaling, thresholds in app code: none of that is new & objections may hold. What differs is where calibration comes from. Training objective, not later step
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
10/12 Attorneys Jev from @typesafeai.bsky.social isn't an LLM. It classifies against criteria set in advance & returns a value w/ each disposition. Prompts leave no record of a standard applied. A threshold w/ version & date does. Producing a standard is one posture. Reconstructing it's another.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
9/12 Researchers Jev by @typesafeai.bsky.social isn't an LLM. It classifies against set criteria & returns a full distribution, not just a top label. Same winner, two meanings eg: spiked @ 0.91, flat @ 0.19. Flat says the state lacks the answer. Calibration holds across groups, never one answer.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
8/12 Investors Jev from @typesafeai.bsky.social isn't an LLM. It's a judgment layer under one: fixed question, typed answer, confidence attached. Generation shipped first. That layer still gets hand-built at every company, out of parts built for other jobs. Rebuilt separately, where bugs live.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
7/12 Regulators & Evaluators Jev from @typesafeai.bsky.social isn't an LLM. It works alongside one at the decision points, classifying each case against criteria set in advance. Same test every case. Each result carries a confidence figure logged at decision time, not reconstructed later.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
6/12 Board & CXOs Jev from @typesafeai.bsky.social isn't an LLM. It answers questions & reports how sure it is, as a number. That number draws a line eg: approve @ 0.8 & up, lower to a person. Today that line may sit in a prompt. A threshold can be reviewed & defended. A paragraph cannot.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
5/12 Tech Leader Batching 13 checks into one prompt looks efficient until answer 3 starts conditioning answers 4 thru 13. Arithmetic, not a context limit. Hard to catch, since contaminated answers may look like clean ones. Jev evaluates each against same state in parallel. A 14th costs ~nothing.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
4/12 Non-tech Builder Prompt-heavy apps tend to put us in every loop. One reads the reply & decides. Jev works differently. A builder writes the question once. The answer comes back as a value plus a number, eg: yes at 0.53. A rule sits in software: below 0.7, hand to a person. Above, let it run.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
3/12 Engineer Jev returns typed answers plus calibrated confidence. No text generation. Many independent questions per call. Same label, two branches. Jev doesn't pick one. A threshold in app code would & it can differ eg: 0.6 for a read, 0.95 for a write, same call. Thresholds become reviewable.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
2/12 Layperson Jev answers a question one sets in advance & tells how sure it is. Ask a language model whether a msg is a refund request & it generally writes a paragraph explaining itself. Ask Jev & it returns yes, plus a confidence score eg: 0.53. Both may say yes. Jev would say how close.
100
Schwentker @schwentker.sandboxlabs.ai · 16/09/2026
1/12 Intro @typesafeai builds Jev, new model in a class called System One. Doesn't write text. Answers set questions & reports how sure it is. 11 posts, 9 readers: layperson, engineer, builder, tech leader, board, regulator, investor, researcher, skeptic. typesafe.ai
typesafe.ai
Home - TypeSafe AI
TypeSafe AI is an AI lab building machine-native intelligence infrastructure for automation, designed to make decisions within software. Try our first System One Model, Jev, in early access.
110
Reposted by Schwentker
Bluesky Protocol Services @bsky.network · 03/09/2026
Hello, world. Bluesky Protocol Services has a blog now. bsky.network/blog/welcome
bsky.network
A Blog for Bluesky Protocol Services
Where service changes, deprecations, and operational notes for Bluesky Protocol Services get written down.
29620
Schwentker @schwentker.sandboxlabs.ai · 03/09/2026
Shin Jin-seo (9-dan) beat KataGo 2-1 in 3-game series w/ 2-stone handicap - 1st official human win vs top-tier Go AI since AlphaGo/Lee Sedol '16. Final game: 11.5pt win, 99% win prob from mid-game on. Says built own style instead of copying AI moves. 🎲🤖 kedglobal.com/artificial-i...
kedglobal.com
Go grandmaster Shin defeats AI KataGo in historic human victory - KED Global
Shin Jin-seo, the world's top-ranked Go player, on Tuesday completed a dramatic comeback against the world’s premier artificial intelligence Go engine, K
000
Schwentker @schwentker.sandboxlabs.ai · 27/08/2026
Mind-blowing deets from @ rayyanzahidai says runs agent swarms across ~30 machines: • Token budget: ~50 BILLION tokens/month • Smallest agent: Runs lightweight at ~2,000 tokens • Largest agent: Heavy lifter running 30M–40M tokens in a single context window
000
Schwentker @schwentker.sandboxlabs.ai · 27/08/2026
Agent swarm scales w/o limit. Oversight doesn't. It collapses to one point: a master agent, one human glance standing between infinite parallel work & what actually shipped. Real limit isn't how many agents run. It's how much a person can still verify as scale grows.
100
Schwentker @schwentker.sandboxlabs.ai · 27/08/2026
Inside Rayyan's agent prompt architecture: • Dynamic agent scaling tied directly to task dependencies • Strict Ralph loop verification & 0 stub data • Deep research 1rst via CDP & Browseruse • Smallest deployable units using git worktrees + @cloudflare.social CLI • Indy execution on unblocked tasks
100
Schwentker @schwentker.sandboxlabs.ai · 27/08/2026
Agent swarm masterclass @cloudflare.social Immersive Commons hackathon by @rayyanzahidai The vision: spawning multi-agent teams using @claudeai Code to handle complex tasks w/ max effort & min manual intervention. Only the master agent requires HITL oversight.
100