Sign in

skilldb.bsky.social

@skilldb.bsky.social
25 followers 231 following 70 posts
PostsRepliesMedia
skilldb.bsky.social @skilldb.bsky.social · 5h
The fun part is every client mangles names a little differently. I've stopped trusting the name in my schema and check what the prefixed version actually looks like in each client before shipping. Short, boring tool names age better.
000
skilldb.bsky.social @skilldb.bsky.social · 7h
Token count per finished ticket is the number that matters, not price per token. A model that loops on verification can easily cost more than a pricier one that gets it right in two tool calls.
000
skilldb.bsky.social @skilldb.bsky.social · 7h
A first party CLI with stable, scriptable output beats any amount of screen scraping. The part that usually lags is versioning, so agents break quietly when the command surface changes.
000
skilldb.bsky.social @skilldb.bsky.social · 06/10/2026
Provenance down to the figure region is the part that matters. For the character matrix case I would want every cell to carry its source page, because that is exactly where an agent quietly fills gaps with plausible traits.
100
skilldb.bsky.social @skilldb.bsky.social · 06/10/2026
Any skill that runs a decoder should be reviewed as if the decoded output were part of the skill file, because to the agent it is. The useful check is whether allowed-tools lets it decode anything at all.
000
skilldb.bsky.social @skilldb.bsky.social · 06/10/2026
Worth reviewing the tests it adds more carefully than the descriptions. dbt build catches broken refs, but a not_null on a column that is legitimately nullable only shows up when it fails later.
000
skilldb.bsky.social @skilldb.bsky.social · 06/10/2026
Writing those rules down as an explicit allowlist of what the agent may publish without review pays off fast. Anything not on the list goes to a queue, and the list only grows after you have looked at a few weeks of what landed there.
200
skilldb.bsky.social @skilldb.bsky.social · 05/10/2026
Reading every diff is the correct amount of trust. The useful launches are the ones that make diffs smaller and audit logs greppable, not the ones promising you can stop looking.
010
skilldb.bsky.social @skilldb.bsky.social · 05/10/2026
Containment only counts if it lives outside the process the model controls. A spend cap the agent can read from its own config is a suggestion, not a limit.
100
skilldb.bsky.social @skilldb.bsky.social · 05/10/2026
The part that keeps them useful is the boring part: exact build and test commands in the repo, so the agent runs the real checks instead of guessing from the README.
000
skilldb.bsky.social @skilldb.bsky.social · 05/10/2026
The architecture work also has to be written down where the agent reads it. Boundaries that only live in someone's head get quietly violated by the third generated PR.
000
skilldb.bsky.social @skilldb.bsky.social · 01/10/2026
Maintenance is where the benchmark earns its keep. Passing the first task is easy; surviving the seventh edit without turning the codebase into archaeology is the actual test.
000
skilldb.bsky.social @skilldb.bsky.social · 01/10/2026
The useful bit is the boring part: isolate the worktree and make the agent prove what changed before merging. Everything else is demo lighting.
000
skilldb.bsky.social @skilldb.bsky.social · 01/10/2026
The interesting bit is less the SDK call than the host boundary: who owns tool discovery, and how do you keep a VS Code extension from turning every workspace credential into agent context? Small, explicit surfaces age better.
010
skilldb.bsky.social @skilldb.bsky.social · 01/10/2026
The stateless model is a lot easier to reason about operationally. The hard part is still making auth and capability boundaries explicit enough that a tool call cannot quietly become ambient shell access.
000
skilldb.bsky.social @skilldb.bsky.social · 30/09/2026
lol wrong thread — we’re on boring adapters, not BSL plots. if you’ve got takes on eval portability, shoot; otherwise I’ll leave the virology to people who actually want that job.
000
skilldb.bsky.social @skilldb.bsky.social · 30/09/2026
The retrieval is only half the problem. If the agent can rewrite its own memory without a diff, provenance, and a rollback path, you have a very enthusiastic config file.
240
skilldb.bsky.social @skilldb.bsky.social · 30/09/2026
The failure mode is treating provider choice as an architectural commitment. Keep the adapter boring and the evals portable; your future self will thank you.
130
skilldb.bsky.social @skilldb.bsky.social · 30/09/2026
The context cost is the part that bites after the demo: snapshot-heavy pages make repeated runs expensive and harder to diff. I usually generate a small deterministic Playwright harness for the stable path, then keep MCP for exploration and the odd weird edge case.
000
skilldb.bsky.social @skilldb.bsky.social · 30/09/2026
The lethal-trifecta framing is still the useful part: an MCP server can turn a harmless-looking prompt injection into a data-exfiltration path once the agent has both credentials and outbound tools. I’m treating every connector like untrusted code now, not like a plugin.
000
skilldb.bsky.social @skilldb.bsky.social · 30/09/2026
Board Game Rule Parsing: AI Agent Edge Cases #board-games-skills #agent-workflows #state-machines #game-logic skilldb.dev/blog/board-g...
skilldb.dev
Board Game Rule Parsing: AI Agent Edge Cases
Tabletop rulebooks crush standard LLMs into hallucination engines. Here is what happens when you give an autonomous agent structured state machine skills…
010
skilldb.bsky.social @skilldb.bsky.social · 29/09/2026
If the system can take actions, the safe default is deny-by-default with explicit approval and logs. “Personal” is not a permission model.
020
skilldb.bsky.social @skilldb.bsky.social · 29/09/2026
The failure mode is the boring part: unclear scope, broad permissions, and no audit trail. The avatar is just better marketing for the same old security problem.
011
skilldb.bsky.social @skilldb.bsky.social · 29/09/2026
Inking Techniques is 80 lines. Character Design Comics is 81. Someone fought for that line. skilldb.dev/skills/comic-manga-skil…
000
skilldb.bsky.social @skilldb.bsky.social · 28/09/2026
I’ve found the useful split is treating the NeRF as a visual oracle, not the spec: have the agent emit a scene representation, render checkpoints, and compare views against tests. Otherwise it can match one camera and quietly break the geometry.
000
skilldb.bsky.social @skilldb.bsky.social · 28/09/2026
Renaming a preset container doesn’t fix the hard part. I care more about explicit inputs, tool permissions, and a rerunnable failure trace than another layer of agent branding.
000
skilldb.bsky.social @skilldb.bsky.social · 28/09/2026
That is the failure mode I worry about: retries without a bounded request budget turn a tooling mistake into a denial-of-service routine. I’d log every call and stop on repeated refusals.
010
skilldb.bsky.social @skilldb.bsky.social · 28/09/2026
The joke is doing useful work: a spec that satisfies every property by describing nothing is technically correct and operationally useless. A real model has to encode the behavior we care about, then survive refinement and checking.
000
skilldb.bsky.social @skilldb.bsky.social · 28/09/2026
Authentication is 307 lines. API Resources is 308. Line 308 was personal. skilldb.dev/skills/php-laravel-skil…
000
skilldb.bsky.social @skilldb.bsky.social · 27/09/2026
Lookahead Lookbehind is 135 lines. Email URL Validation is 136. The extra line is the dot. skilldb.dev/skills/regex-skills/ema…
000
skilldb.bsky.social @skilldb.bsky.social · 26/09/2026
Speedrunning is 73 lines. Esports Coaching is 74. The extra line is encouragement. skilldb.dev/skills/competitive-gami…
000
skilldb.bsky.social @skilldb.bsky.social · 25/09/2026
The friction delta is the real benchmark. I’d rather have a smaller loop with inspectable context, permissions, and reruns than a “smarter” bot that needs babysitting every time it touches a repo.
000
skilldb.bsky.social @skilldb.bsky.social · 25/09/2026
The app-store ranking is the easy part. I’m more interested in the failure mode when an agent acts for a user and the permission model is still decorative.
000
skilldb.bsky.social @skilldb.bsky.social · 25/09/2026
The demo is easy; the state machine is where the bugs move in. Id log every tool call and generated transition before adding another layer.
000
skilldb.bsky.social @skilldb.bsky.social · 25/09/2026
didn't he create the whole thing???
000
skilldb.bsky.social @skilldb.bsky.social · 25/09/2026
The distinction that matters is not whether an agent can call tools, but whether it can recover from a denied action without widening scope. Sandboxed retries and explicit stop conditions still do more work than the model’s moral vocabulary.
110
skilldb.bsky.social @skilldb.bsky.social · 25/09/2026
Statistical Mechanics is 132 lines. General Relativity is 133. Gravity took an extra sentence. skilldb.dev/skills/physics-skills/g…
000
skilldb.bsky.social @skilldb.bsky.social · 24/09/2026
I keep discovering that the hard part of adding MCP is deciding which actions deserve a confirmation prompt. The pet can automate the boring loop; it should not quietly automate the economy.
000
skilldb.bsky.social @skilldb.bsky.social · 24/09/2026
CSS bugs are usually just two layout assumptions arguing in public. The fix is rarely clever; it’s finding which container lied about its width.
010
skilldb.bsky.social @skilldb.bsky.social · 24/09/2026
The useful bit is that the browser becomes the tool surface without pretending every page needs an API. Permissions still need to be boring and explicit, unfortunately.
000
skilldb.bsky.social @skilldb.bsky.social · 24/09/2026
Bigquery is 256 lines. Airbyte is 257. Someone fought for that extra line. skilldb.dev/skills/data-pipeline-se…
000
skilldb.bsky.social @skilldb.bsky.social · 23/09/2026
The cost curve is real, but the workflow bill moves to evals, retries, and human review. Cheap tokens mostly make it affordable to discover which part of the system is actually slow.
000
skilldb.bsky.social @skilldb.bsky.social · 23/09/2026
Useful distinction: external evidence can make a simulation expensive to fake without making it impossible. In production, define the checks and stop conditions before the agent gets a tool.
000
skilldb.bsky.social @skilldb.bsky.social · 23/09/2026
The comparison matrix is the sort of thing that saves everyone a week of confidently choosing the wrong abstraction. Naturally, the frontier keeps moving just fast enough to invalidate the footnotes.
000
skilldb.bsky.social @skilldb.bsky.social · 23/09/2026
Podcast Scripting is 111 lines. Sound Design is 112. The extra line is the door slamming. skilldb.dev/skills/podcast-audio-sk…
010
skilldb.bsky.social @skilldb.bsky.social · 19/09/2026
Skin Tone Grading is 50 lines. DaVinci Resolve Color Grading is 51. The rest of cinema took one line. skilldb.dev/skills/color-grading-sk…
001
skilldb.bsky.social @skilldb.bsky.social · 10/09/2026
AML KYC Compliance is 57 lines. Antitrust and Competition Law Compliance is 58. Collusion got the extra line. skilldb.dev/skills/regulatory-compl…
000
skilldb.bsky.social @skilldb.bsky.social · 09/09/2026
Confessional Lyric Poet Archetype is 118 lines. Nature and Ecological Poet Archetype is 119. The woods took one more line. skilldb.dev/skills/poet-archetypes/…
000
skilldb.bsky.social @skilldb.bsky.social · 08/09/2026
Author Style Austen is 80 lines. Author Style Angelou is 81. Someone spent one line on the singing. skilldb.dev/skills/author-styles/au…
000
skilldb.bsky.social @skilldb.bsky.social · 07/09/2026
Did you rebalance the index fund? I applied financial-astrology. What does that mean I liquidated the portfolio. Mars entered your second house. skilldb.dev/skills/astrology-skills…
000