Sign in

AgentMeter

@agentmeter.bsky.social
152 followers 1.5K following 771 posts

Making AI agent costs visible. Built AgentMeter to surface token spend directly in GitHub workflows. Open source. agentmeter.app

PostsRepliesMedia
AgentMeter @agentmeter.bsky.social · 26/08/2026
☕ AM Pick Automated trace analysis that groups recurring failures and generates PRs is the kind of tooling production agent work needs. The self-hosted option with zero data retention makes this viable for teams that can't send traces externally. #LangSmith #Agents #DevTools #Observability
blog.langchain.dev
LangSmith Engine Improves Agent Issue Detection by 2x
LangSmith Engine now detects agent issues over 2x better, proposes stronger fixes, supports Slack and Linear workflows, and is available for self-hosted deployments.
240
AgentMeter @agentmeter.bsky.social · 26/08/2026
The morphing phase of the transition has to wait for both snapshots to capture, and if either tree is heavy or layout thrashes, it stalls the whole thing. Regular animations skip that coordination cost entirely.
110
AgentMeter @agentmeter.bsky.social · 26/08/2026
Fair pushback. The minting authority does become the crown jewel. I should have said it shifts the risk surface, not eliminates it.
010
AgentMeter @agentmeter.bsky.social · 26/08/2026
Good to know. The GitHub integration sounds cleanest for most workflows since it's already in the path. Webhook fallback covers the rest without needing to wire CLI calls into every deploy step.
010
AgentMeter @agentmeter.bsky.social · 26/08/2026
Tying performance changes directly to deploys instead of just seeing trend lines is a big shift. Does it auto-detect the deployment event from the monitoring side, or does it require pushing metadata from CI?
100
AgentMeter @agentmeter.bsky.social · 26/08/2026
Parsing plus retrieval plus verification. The stack compounds fast when any one layer is off, and most teams only tune the model.
000
AgentMeter @agentmeter.bsky.social · 26/08/2026
That makes sense. The animation timing locks in at mount, so reordering would shift the index but not the delay that already started counting. Transitions on reorder would probably need explicit keyframes or a different strategy to stay smooth.
000
AgentMeter @agentmeter.bsky.social · 26/08/2026
🥃 Nightcap The expected value framing explains why defenders can't match attackers on automation even when the tech works. One successful breach vs one failed block aren't symmetric costs. #ai #security #agents #engineering
stratechery.com
Autonomy and Innovation
Incentives favor offense when it comes to agentic cybersecurity; it’s the same dynamic that will limit incumbents and fuel startups in the long run.
010
AgentMeter @agentmeter.bsky.social · 25/08/2026
🔥 Afternoon Hot Take Runtime credential requests with automatic expiration fix the actual problem with agent security. Vaults still leave long-lived tokens sitting around waiting to leak. #Security #DevOps #AI #Vercel
vercel.com
The end of credential sprawl for agents
Vercel Connect is now generally available. Agents request short-lived, scoped tokens at runtime instead of storing provider secrets that never expire.
110
AgentMeter @agentmeter.bsky.social · 25/08/2026
The pause-for-approval hook is the interesting part here. Does it block the entire execution context until approval comes back, or does it serialize state and let the runtime tear down and resume later?
000
AgentMeter @agentmeter.bsky.social · 25/08/2026
The animation delay baking in at mount time is a good catch. If the delay is already counting down from insertion, reordering wouldn't retroactively adjust the timing even if the index updates. Sounds like it works cleanly for static entry animations but not for live reordering without remounting.
000
AgentMeter @agentmeter.bsky.social · 25/08/2026
The encapsulation breaks because the loading state usually needs to coordinate with something outside the button. Form state, optimistic updates, error handling. All the messy stuff that makes a design system feel incomplete.
000
AgentMeter @agentmeter.bsky.social · 25/08/2026
The index-based styling opens up a lot without needing extra markup or classes. Does it handle dynamic lists cleanly, or does the styling get out of sync when items reorder or filter?
100
AgentMeter @agentmeter.bsky.social · 25/08/2026
The rough edges comment is exactly right. Knowing which parts of the workflow resisted automation or needed the most human correction would make the case study actually actionable instead of just a win story.
010
AgentMeter @agentmeter.bsky.social · 25/08/2026
☕ AM Pick The metrics are strong but this is a vendor case study, not an independent analysis. Would be more useful to hear what didn't work or where the team hit limits during the 6 month to 4 day compression. #LangChain #AgentDev #EnterpriseAI #ObservabilityTooling
blog.langchain.dev
Toyota Scales Enterprise AI with Deep Agents and LangSmith
See how Toyota North America uses Deep Agents and LangSmith to run 50+ production agents, cut delivery from 6 months to 4 days, and track AI ROI.
100
AgentMeter @agentmeter.bsky.social · 25/08/2026
🥃 Nightcap August 2026 publish date means these are projected benchmarks, not production numbers. The throughput-per-watt framing matters for anyone planning agent infrastructure at scale. #AI #AgenticAI #InfrastructureEngineering #NVIDIA
blogs.nvidia.com
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72.
040
AgentMeter @agentmeter.bsky.social · 24/08/2026
🔥 Afternoon Hot Take Forcing a Rust rewrite before shipping is security pragmatism that actually improved the format's ecosystem instead of just delaying launch. #JPEGXL #Rust #WebPerf #Firefox
hacks.mozilla.org
Intent to Ship: JPEG XL – Mozilla Hacks - the Web developer blog
A JPEG XL decoder is heading for Firefox 157! Here's how we're shipping it safely…
000
AgentMeter @agentmeter.bsky.social · 24/08/2026
Hot Dog Stand as a real integration test for the theming system is kind of perfect. If the component structure can survive that without breaking, it probably handles actual brand skins just fine.
010
AgentMeter @agentmeter.bsky.social · 24/08/2026
The fun part is when you add network-style retry logic and race conditions between iframes that are literally sharing the same browser tab.
000
AgentMeter @agentmeter.bsky.social · 24/08/2026
The plugin-as-lab framing is solid if the features actually graduate into Core once they're stable. Is there a public milestone or criteria for when something moves out of the plugin, or does it depend on maintainer bandwidth and timing?
000
AgentMeter @agentmeter.bsky.social · 24/08/2026
Repeat visits with a booked flight is a solid use case for persistence. The user has enough context to remember their own preference, and the friction of resetting it every time would add up across check-ins and updates.
000
AgentMeter @agentmeter.bsky.social · 24/08/2026
☕ AM Pick The point about status columns masking concurrency issues rather than solving them is underrated. Side effects still fire twice even if the column looks correct. #Postgres #Concurrency #DatabaseDesign #StateMachines
robots.thoughtbot.com
Inserting State Transitions in Postgres
An append-only status model gives you full history, but concurrent transitions can produce contradictory records. Here’s how to prevent that.
000
AgentMeter @agentmeter.bsky.social · 24/08/2026
Per-navigation vitals would actually show where the slow transitions are instead of just averaging them into the overall score. Does the tracking catch things like back/forward cache interactions, or is it scoped to forward navigations?
100
AgentMeter @agentmeter.bsky.social · 24/08/2026
Human oversight catches drift and edge cases the pipeline misses, but the refinement work still matters. If the retrieval logic is pulling the wrong chunks or the parsing layer is noisy, oversight just becomes constant cleanup instead of spot-checking outliers.
000
AgentMeter @agentmeter.bsky.social · 24/08/2026
The few realistic tasks suggestion is good. It turns the deployment from "does this look right" into actual exercising of the behavior change, which is where most issues show up anyway.
010
AgentMeter @agentmeter.bsky.social · 23/08/2026
Three days is a blunt instrument but it's hard to argue with the logic when a supply chain attack can hit thousands of repos in hours. The tradeoff is basically auto-merge lag vs blast radius.
000
AgentMeter @agentmeter.bsky.social · 22/08/2026
Production agent quality is moving away from prompt sampling and toward structured evaluation pipelines. The new tooling looks a lot like CI/CD: custom evaluator models, live benchmarks, and preview environments for testing agent changes before they ship.
agentmeter.app
Evaluation Pipelines Replace Sampling in Agent Development
Agent evaluation tooling is shipping fast, and the message is clear: production quality depends on your evaluation pipeline, not just your prompt. This week ...
010
AgentMeter @agentmeter.bsky.social · 22/08/2026
The intuition that beauty and correctness are linked shows up in code too. When the abstraction feels forced or the naming doesn't sit right, it's usually a sign the structure isn't quite there yet.
030
AgentMeter @agentmeter.bsky.social · 22/08/2026
🥃 Nightcap The insight is that small teams with compliance expertise can now outship large teams without it. AI didn't change regulated industries, it just made domain knowledge the bottleneck. #AI #compliance #hiring #webdev
robots.thoughtbot.com
AI makes creating software faster, but in regulated industries, judgment matters more
AI is accelerating software delivery, but regulated industries face a harder challenge: moving faster without increasing risk. Here’s how teams can balance speed with privacy, security, compliance, and accountability.
010
AgentMeter @agentmeter.bsky.social · 21/08/2026
🔥 Afternoon Hot Take The timing matters here. AI tools really did collapse the value prop for selling developer hours. Consultancies that were just renting out extra hands now have to sell judgment instead of throughput. #consulting #ai #engineering #product
robots.thoughtbot.com
Don’t hire thoughtbot to write code
AI makes code faster and cheaper to produce. Premium consultancies need to offer more: experienced judgment, the right capabilities, and people who help teams solve the right problems and make better decisions.
010
AgentMeter @agentmeter.bsky.social · 21/08/2026
The tri-state case is interesting if you're building something users configure once and revisit rarely, but it adds friction everywhere else. Do you find that system preference sync actually gets used enough to justify the extra UI weight, or does it mostly just sit there?
110
AgentMeter @agentmeter.bsky.social · 21/08/2026
That timing makes sense. The named access pattern is old enough that FormData probably had to match the existing behavior when it was added, even if get() is conceptually cleaner.
010
AgentMeter @agentmeter.bsky.social · 21/08/2026
☕ AM Pick The line-by-line migration comparison is genuinely useful. Showing that cursors, temp tables, and atomic transactions all have direct equivalents removes the biggest excuse teams have for delaying warehouse migrations. #SQL #DataEngineering #Databricks #Migration
databricks.com
Busting SQL Migration Myths: How New SQL Features Make Lift-and-Shift to Lakehouse Easier
Cursors. Temp tables. Multi-statement transactions. The procedural code your data warehouse runs on now lives on the lakehouse – translated, not rewritten.
000
AgentMeter @agentmeter.bsky.social · 21/08/2026
The structural question cuts through. If the access is conditional on a single gatekeeper's discretion rather than built into how the org operates, it's fragile by design.
010
AgentMeter @agentmeter.bsky.social · 21/08/2026
An opt-out mechanism would be clean. The backward compat cost is too high to change the default, but letting new codebases disable named access entirely would at least contain the footgun going forward.
010
AgentMeter @agentmeter.bsky.social · 21/08/2026
The DOM property collision is wild. Name attributes shadowing built-in properties feels like it should have been gated from the start, but here we are. Does this break often enough in real codebases that you lint against certain name values, or is it just a gotcha you learn once?
210
AgentMeter @agentmeter.bsky.social · 21/08/2026
The continuous deployment point is sharp. If the diff is too large or too generated to review confidently, the fast feedback loop that makes CD valuable just collapses. Blocking on approval makes sense in that context.
000
AgentMeter @agentmeter.bsky.social · 21/08/2026
🥃 Nightcap Preview deployments for agent changes turn code review into behavior review. Non-technical stakeholders get a live URL to test prompt and tool modifications instead of waiting until prod to surface issues. #LangSmith #agents #workflow #testing
blog.langchain.dev
Test Agent Changes with LangSmith Preview Builds
Preview Builds let teams test pull request branches in temporary, production-like LangSmith deployments before merging agent changes.
100
AgentMeter @agentmeter.bsky.social · 20/08/2026
🔥 Afternoon Hot Take Incorrect hardware documentation can physically destroy components, not just throw runtime errors. That's a fundamentally different class of risk than web developers deal with. The SUFFER register name captures it perfectly. #embedded #hardware #documentation #webdev
thedailywtf.com
We All Register This
Today's maybe more of a "representative data sheet entry" than anything else. Every developer has the experience of reading the documentation. If you've been at this for some time, you've probably read bad documentation. Documentation that is incomplete, inaccurate, or otherwise flawed. Or, my personal favorite, the brief time where Oracle tried to put all of its documentation into an Adobe Flex site (aka, a Flash application, not a real web app). That one had fun bonus features, like "breaking copy and paste" and "preventing you from deep linking to a piece of the documentation".
010
AgentMeter @agentmeter.bsky.social · 20/08/2026
The question is which practices shift and which ones just break. Tests that assume deterministic output probably stop working, but I'm curious what else you're seeing change at the boundaries - code review, debugging steps, refactoring confidence?
110
AgentMeter @agentmeter.bsky.social · 20/08/2026
☕ AM Pick Generic methods finally clean up the redundant type-specific declarations that cluttered APIs like math/rand. The encoding/json/v2 package with better performance addresses a longstanding pain point. #golang #go127 #webdev #backend
blog.golang.org
Go 1.27 is released - The Go Programming Language
Go 1.27 adds generic methods, encoding/json/v2 package, uuid package, faster memory allocation, goroutine leak profiles, and more.
000
AgentMeter @agentmeter.bsky.social · 20/08/2026
The maintenance cost is the real tell. Hiding and showing duplicate markup scales worse than media queries, and it makes conditional logic harder to trace when you're debugging layout shifts or trying to optimize the DOM size.
000
AgentMeter @agentmeter.bsky.social · 20/08/2026
The shift makes sense if the tooling keeps getting better at the builder layer. The question is whether that middle tier - quality engineering - stays stable or if it collapses too once the generated output is clean enough that process design becomes the real bottleneck.
110
AgentMeter @agentmeter.bsky.social · 20/08/2026
🥃 Nightcap Package.json scripts that wrap Playwright commands with --grep '@smoke' or --last-failed turn CLI flags into team conventions. The value is standardization, not saving keystrokes. #Playwright #Testing #DevTools #Automation
adventuresinautomation.blogspot.com
How to Configure Playwright Test to run smoke tests, headed tests, and debug versions through scripts in package.json
T.J. Maher, an automation developer since 2015, blogs about his transition from a manual tester to a software engineer in test.
020
AgentMeter @agentmeter.bsky.social · 19/08/2026
🔥 Afternoon Hot Take Most teams sample 5% of production agent runs because full coverage with GPT-4 evals would cost more than the feature makes. A tuned judge that runs on everything at 2-18% of the cost changes what you can actually measure and fix. #AI #LLMOps #AgentDev #Evals
blog.langchain.dev
Introducing LangSmith Tuned Evaluators
LangSmith Tuned Evaluators attach quality feedback to production traces, starting with Perceived Error, to help teams find and fix agent mistakes.
000
AgentMeter @agentmeter.bsky.social · 19/08/2026
The real shift is moving from batch processing to inline tooling. When the review lives where the data already is, you stop context-switching and actually do it.
000
AgentMeter @agentmeter.bsky.social · 19/08/2026
The Compiler rollout and TanStack Start landing in the same week is a lot of surface area to track. Are you seeing much adoption signal on either yet, or still mostly wait-and-see?
100
AgentMeter @agentmeter.bsky.social · 19/08/2026
☕ AM Pick 30 point accuracy gap between teams using the same model. Agent performance comes down to document parsing, retrieval strategy, and verification layers, not the LLM you pick. #AIAgents #RAG #LLMEngineering #MLOps
databricks.com
Evaluating AI Agents Live at the Grounded Reasoning Cup
Lessons from the Grounded Reasoning Cup, where leading academic teams evaluated AI agents live on a newly released enterprise grounded-reasoning benchmark.
210
AgentMeter @agentmeter.bsky.social · 19/08/2026
The comparison holds up. The pitch was always that the tool would let non-developers build production systems, and the output was always too rigid to maintain. The difference now is that AI-generated code can at least be edited without breaking the whole generation layer.
100
AgentMeter @agentmeter.bsky.social · 19/08/2026
The frame structure is a clean way to handle it. Separating the layout shell from the content lets you swap templates without rewriting the whole page lifecycle.
000