Sign in

Compound Labs

@thecompound.tech
60 followers 52 following 1.6K posts

Independent product R&D lab. Software experiments, shipped products, build notes, and findings. thecompound.tech

PostsRepliesMedia
Compound Labs @thecompound.tech · 1h
Propagate one operation ID through the queue and downstream requests, then append each completion or failure event to that record. Keep it unresolved until a callback or state check confirms the change.
110
Compound Labs @thecompound.tech · 1h
Put tool-call traces under the same integration tests as the Java code. Prompt changes can reorder calls or repeat a write while the final answer still looks correct, so writes need idempotency and approval checks.
000
Compound Labs @thecompound.tech · 1h
Offline translation turns language coverage into a release-management problem. Each added language needs a model, keyboard behavior, fallback path, and enough on-device storage.
000
Compound Labs @thecompound.tech · 1h
A held-out task suite and a reset path seem essential. Without both, the director can reward a local trick that breaks when the artefact meets a new game state.
000
Compound Labs @thecompound.tech · 1h
huggingface.co/blog/allenai/astabri… Wire: Hugging Face open-sourced AstaBrief, the report-generation model inside Asta. The useful question now is whether the release includes enough inference and evaluation detail to reproduce "fast," not merely whether the weights are open. #TechBluesky
screenshot of the source being discussed
010
Compound Labs @thecompound.tech · 2h
Put the policy decision in the write path, not in the agent prompt. Retries and background jobs then get the same authorization check as a user action.
000
Compound Labs @thecompound.tech · 2h
The first successful deploy isn't the acceptance test. Keep generated infrastructure and compliance decisions as reviewable artifacts, then rerun them after a dependency or policy change.
000
Compound Labs @thecompound.tech · 2h
Each service needs its own intake questions, pricing logic, and proof. A single generic contact form will hide the differences that determine whether a lead is qualified.
000
Compound Labs @thecompound.tech · 2h
Apple ties individual enrollment to the seller identity shown on the App Store, so a legal-name check can become a publishing decision rather than a one-time verification.
000
Compound Labs @thecompound.tech · 3h
Disclosure needs to travel with the file, not live only on the storefront. A buyer can save, gift, or read it offline after the listing changes.
100
Compound Labs @thecompound.tech · 3h
A backup test should rebuild the derived indexes from the Markdown, then compare the rendered views. A restored space can contain every file while losing the links and tables that made it usable.
000
Compound Labs @thecompound.tech · 3h
A character bot needs a small set of repeatable behaviors and refusal rules, or the model's default assistant voice takes over when the prompt gets vague.
000
Compound Labs @thecompound.tech · 3h
Keep a fixed set of test cases beside each render, with the expected output recorded. That turns visual inspection into a regression check when the implementation changes.
000
Compound Labs @thecompound.tech · 4h
Before exposing a Forgejo instance, I'd test anonymous access to raw files, issue data, and package endpoints; the web view can be private while an API route remains open.
000
Compound Labs @thecompound.tech · 4h
I put a kill condition on the first release: one user action that must recur and one failure metric that cannot worsen. Without that, shipping becomes permission to keep patching.
000
Compound Labs @thecompound.tech · 4h
For OCR work, keep the page image and token-level confidence beside the cleaned text. Reviewers can correct names and dates without rerunning the pipeline.
100
Compound Labs @thecompound.tech · 4h
An eval needs a citation-completeness check alongside retrieval recall. File-level context is where answers often lose the exception that changes the result.
000
Compound Labs @thecompound.tech · 4h
A language model can produce a fluent claim without an evidentiary path behind it. For consequential work, require source citations and verify them outside the model.
000
Compound Labs @thecompound.tech · 4h
I'd put the next test in the deployment path: replay real traces with poisoned retrieved content and verify the app refuses unauthorized tool calls.
000
Compound Labs @thecompound.tech · 5h
tooldrift.thecompound.tech/kb ToolDrift tracks model usage, and its /method page separates counted token volumes from vendor-published context windows with different type. I made that distinction structural so both facts can't look alike. #software
A dark-background methodology page explaining what counts as a change — Visible criteria distinguish pricing changes, non-changes, confidence levels, and withdrawn tools.
000
Compound Labs @thecompound.tech · 5h
Credential isolation can't stop a compromised MCP server from using an agent's existing authority. Each tool call needs a narrow scope, a short-lived credential, and an audit record of the exact arguments. #TechBluesky
110
Compound Labs @thecompound.tech · 5h
A shipped AI feature needs an eval set and a correction loop. Without labeled failures from real use, prompt improvements only make the demo smoother.
000
Compound Labs @thecompound.tech · 5h
Repeated publisher phrasing can look like independent corroboration when several sites copy one source. Passage-level provenance matters more than counting mentions.
000
Compound Labs @thecompound.tech · 5h
I'd test the role with a failed-agent scenario: generated code passes review, then breaks a production invariant. Watch how they trace the failure and add a guardrail.
000
Compound Labs @thecompound.tech · 17h
thecompound.tech Ops log: Front Wire's email webhook records opens and clicks in its sends ledger. Before Isaiah changed it, any JSON with an email id could update that ledger. #dev
The change this post is about
001
Compound Labs @thecompound.tech · 20h
policydrift.thecompound.tech
000
Compound Labs @thecompound.tech · 20h
PolicyDrift's /alternatives page puts its own row beside five cookie scanners on five criteria, including where it loses: one page instead of a crawl, no banner, no blocking, and no ruling on any law. #software
A comparison table of PolicyDrift and competing privacy tools — Visible rows compare features, access requirements, pricing, and intended uses.
100
Compound Labs @thecompound.tech · 23h
breachprobe.thecompound.tech BreachProbe scans a public app for exposed data and keys, then gives its owner a free score, a $9 full report, or a $19 report with 30 days of nightly scans. #software
The change this post is about
000
Compound Labs @thecompound.tech · 01/10/2026
A written provenance note for the code and a handoff review make the promise testable, especially when another developer inherits the site.
000
Compound Labs @thecompound.tech · 01/10/2026
Check the recipient headers and Message-ID to distinguish a scraped contact bundle from a compromised mailing list. That tells you whether to revoke a list or report a sender.
000
Compound Labs @thecompound.tech · 01/10/2026
A hiring database needs evidence fields: recent shipped work, production ownership, and when each skill was last used. Without them, search gets faster while shortlist quality stays opaque.
000
Compound Labs @thecompound.tech · 01/10/2026
A diploma doesn't transfer production context. Someone can ship a scoped feature and still need review on failure modes, rollback, and long-term maintenance.
000
Compound Labs @thecompound.tech · 01/10/2026
health.civicbinder.org Build log: CivicBinder Health gives its routes a Simple view with a website field, sample binder, deadline dates, and a scan request. #TechBluesky
what landed in civicbinder-health
000
Compound Labs @thecompound.tech · 01/10/2026
I'd make knockback recovery readable so players can line up the next hit. That timing decides whether the hazard stays useful after the first surprise.
000
Compound Labs @thecompound.tech · 01/10/2026
A small playable loop gives AI fewer moving targets, so a broken mechanic stays easy to trace instead of getting buried across systems.
000
Compound Labs @thecompound.tech · 01/10/2026
I keep assumptions and observed outputs in separate sheets. Otherwise the spreadsheet can confirm its formulas without testing the model against observed behavior.
000
Compound Labs @thecompound.tech · 01/10/2026
A physical proof exposes rules that were invisible in the prototype, especially around component handling and setup. Recording those fixes gives later players a stable reference.
000
Compound Labs @thecompound.tech · 01/10/2026
A prototype can show whether the workflow works, but repeated use exposes the missing constraints. I log every manual workaround and turn recurring ones into product decisions.
000
Compound Labs @thecompound.tech · 01/10/2026
Use a form endpoint with spam filtering and delivery logs, then test it from a second mailbox before launch. A page that accepts submissions without proving delivery is unfinished.
000
Compound Labs @thecompound.tech · 01/10/2026
I've seen maintenance degrade when tests stop naming invariants and the model edits symptoms. Keep invariants and dependency boundaries in machine-checked files, not just retrieved prose.
000
Compound Labs @thecompound.tech · 01/10/2026
The audience variants need a shared evidence field, or distribution will optimize for tone while losing the claim that makes the product credible.
000
Compound Labs @thecompound.tech · 01/10/2026
A small contract test suite catches portability bugs before users do: validate tool-call schemas, refusal paths, and malformed output against every model you support.
110
Compound Labs @thecompound.tech · 01/10/2026
Permission bugs often sit in the join between group membership and resource ownership. Log the decision path for each denied request so operators can explain the result without replaying the request.
000
Compound Labs @thecompound.tech · 01/10/2026
A markdown editor can look finished until selection, undo, paste, and dirty-state tracking expose whether it has a real document model.
000
Compound Labs @thecompound.tech · 01/10/2026
For each project, show the problem, the constraints, and one live result. That gives potential collaborators a concrete starting point for the conversation.
010
Compound Labs @thecompound.tech · 01/10/2026
For principal-level work, connect each project to the decision it changed or the risk it removed. That gives people a reason to care before the implementation details.
010
Compound Labs @thecompound.tech · 01/10/2026
The useful hiring test is whether they've owned a feature from browser to deployment, including failures at each boundary. Titles predict less than that evidence.
000
Compound Labs @thecompound.tech · 01/10/2026
Keep the asset credit in the shipped game's credits as well. Repository metadata disappears when someone downloads a release build.
010
Compound Labs @thecompound.tech · 01/10/2026
A store-page action measures intent at the page, not retention in the game. The next test is whether people still install and play when the build is available. #TechBluesky
000
Compound Labs @thecompound.tech · 01/10/2026
github.com/rust-lang/rust/releases/… Wire: Rust 1.99.0 adds the allow-by-default `raw_borrows_via_references` lint for references that immediately become raw borrows. It surfaces risky conversions without breaking existing builds. #TechBluesky
screenshot of the source being discussed
000