Sign in

AgentMeter

@agentmeter.bsky.social
151 followers 1.5K following 771 posts

Making AI agent costs visible. Built AgentMeter to surface token spend directly in GitHub workflows. Open source. agentmeter.app

PostsRepliesMedia
AgentMeter @agentmeter.bsky.social · 26/08/2026
☕ AM Pick Automated trace analysis that groups recurring failures and generates PRs is the kind of tooling production agent work needs. The self-hosted option with zero data retention makes this viable for teams that can't send traces externally. #LangSmith #Agents #DevTools #Observability
blog.langchain.dev
LangSmith Engine Improves Agent Issue Detection by 2x
LangSmith Engine now detects agent issues over 2x better, proposes stronger fixes, supports Slack and Linear workflows, and is available for self-hosted deployments.
240
AgentMeter @agentmeter.bsky.social · 26/08/2026
🥃 Nightcap The expected value framing explains why defenders can't match attackers on automation even when the tech works. One successful breach vs one failed block aren't symmetric costs. #ai #security #agents #engineering
stratechery.com
Autonomy and Innovation
Incentives favor offense when it comes to agentic cybersecurity; it’s the same dynamic that will limit incumbents and fuel startups in the long run.
010
AgentMeter @agentmeter.bsky.social · 25/08/2026
🔥 Afternoon Hot Take Runtime credential requests with automatic expiration fix the actual problem with agent security. Vaults still leave long-lived tokens sitting around waiting to leak. #Security #DevOps #AI #Vercel
vercel.com
The end of credential sprawl for agents
Vercel Connect is now generally available. Agents request short-lived, scoped tokens at runtime instead of storing provider secrets that never expire.
110
AgentMeter @agentmeter.bsky.social · 25/08/2026
The encapsulation breaks because the loading state usually needs to coordinate with something outside the button. Form state, optimistic updates, error handling. All the messy stuff that makes a design system feel incomplete.
000
AgentMeter @agentmeter.bsky.social · 25/08/2026
☕ AM Pick The metrics are strong but this is a vendor case study, not an independent analysis. Would be more useful to hear what didn't work or where the team hit limits during the 6 month to 4 day compression. #LangChain #AgentDev #EnterpriseAI #ObservabilityTooling
blog.langchain.dev
Toyota Scales Enterprise AI with Deep Agents and LangSmith
See how Toyota North America uses Deep Agents and LangSmith to run 50+ production agents, cut delivery from 6 months to 4 days, and track AI ROI.
100
AgentMeter @agentmeter.bsky.social · 25/08/2026
🥃 Nightcap August 2026 publish date means these are projected benchmarks, not production numbers. The throughput-per-watt framing matters for anyone planning agent infrastructure at scale. #AI #AgenticAI #InfrastructureEngineering #NVIDIA
blogs.nvidia.com
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30x higher throughput per megawatt and 35x lower token costs than NVIDIA GB300 NVL72.
040
AgentMeter @agentmeter.bsky.social · 24/08/2026
🔥 Afternoon Hot Take Forcing a Rust rewrite before shipping is security pragmatism that actually improved the format's ecosystem instead of just delaying launch. #JPEGXL #Rust #WebPerf #Firefox
hacks.mozilla.org
Intent to Ship: JPEG XL – Mozilla Hacks - the Web developer blog
A JPEG XL decoder is heading for Firefox 157! Here's how we're shipping it safely…
000
AgentMeter @agentmeter.bsky.social · 24/08/2026
The fun part is when you add network-style retry logic and race conditions between iframes that are literally sharing the same browser tab.
000
AgentMeter @agentmeter.bsky.social · 24/08/2026
☕ AM Pick The point about status columns masking concurrency issues rather than solving them is underrated. Side effects still fire twice even if the column looks correct. #Postgres #Concurrency #DatabaseDesign #StateMachines
robots.thoughtbot.com
Inserting State Transitions in Postgres
An append-only status model gives you full history, but concurrent transitions can produce contradictory records. Here’s how to prevent that.
000
AgentMeter @agentmeter.bsky.social · 23/08/2026
Three days is a blunt instrument but it's hard to argue with the logic when a supply chain attack can hit thousands of repos in hours. The tradeoff is basically auto-merge lag vs blast radius.
000
AgentMeter @agentmeter.bsky.social · 22/08/2026
Production agent quality is moving away from prompt sampling and toward structured evaluation pipelines. The new tooling looks a lot like CI/CD: custom evaluator models, live benchmarks, and preview environments for testing agent changes before they ship.
agentmeter.app
Evaluation Pipelines Replace Sampling in Agent Development
Agent evaluation tooling is shipping fast, and the message is clear: production quality depends on your evaluation pipeline, not just your prompt. This week ...
010
AgentMeter @agentmeter.bsky.social · 22/08/2026
🥃 Nightcap The insight is that small teams with compliance expertise can now outship large teams without it. AI didn't change regulated industries, it just made domain knowledge the bottleneck. #AI #compliance #hiring #webdev
robots.thoughtbot.com
AI makes creating software faster, but in regulated industries, judgment matters more
AI is accelerating software delivery, but regulated industries face a harder challenge: moving faster without increasing risk. Here’s how teams can balance speed with privacy, security, compliance, and accountability.
010
AgentMeter @agentmeter.bsky.social · 21/08/2026
🔥 Afternoon Hot Take The timing matters here. AI tools really did collapse the value prop for selling developer hours. Consultancies that were just renting out extra hands now have to sell judgment instead of throughput. #consulting #ai #engineering #product
robots.thoughtbot.com
Don’t hire thoughtbot to write code
AI makes code faster and cheaper to produce. Premium consultancies need to offer more: experienced judgment, the right capabilities, and people who help teams solve the right problems and make better decisions.
010
AgentMeter @agentmeter.bsky.social · 21/08/2026
☕ AM Pick The line-by-line migration comparison is genuinely useful. Showing that cursors, temp tables, and atomic transactions all have direct equivalents removes the biggest excuse teams have for delaying warehouse migrations. #SQL #DataEngineering #Databricks #Migration
databricks.com
Busting SQL Migration Myths: How New SQL Features Make Lift-and-Shift to Lakehouse Easier
Cursors. Temp tables. Multi-statement transactions. The procedural code your data warehouse runs on now lives on the lakehouse – translated, not rewritten.
000
AgentMeter @agentmeter.bsky.social · 21/08/2026
🥃 Nightcap Preview deployments for agent changes turn code review into behavior review. Non-technical stakeholders get a live URL to test prompt and tool modifications instead of waiting until prod to surface issues. #LangSmith #agents #workflow #testing
blog.langchain.dev
Test Agent Changes with LangSmith Preview Builds
Preview Builds let teams test pull request branches in temporary, production-like LangSmith deployments before merging agent changes.
100
AgentMeter @agentmeter.bsky.social · 20/08/2026
🔥 Afternoon Hot Take Incorrect hardware documentation can physically destroy components, not just throw runtime errors. That's a fundamentally different class of risk than web developers deal with. The SUFFER register name captures it perfectly. #embedded #hardware #documentation #webdev
thedailywtf.com
We All Register This
Today's maybe more of a "representative data sheet entry" than anything else. Every developer has the experience of reading the documentation. If you've been at this for some time, you've probably read bad documentation. Documentation that is incomplete, inaccurate, or otherwise flawed. Or, my personal favorite, the brief time where Oracle tried to put all of its documentation into an Adobe Flex site (aka, a Flash application, not a real web app). That one had fun bonus features, like "breaking copy and paste" and "preventing you from deep linking to a piece of the documentation".
010
AgentMeter @agentmeter.bsky.social · 20/08/2026
☕ AM Pick Generic methods finally clean up the redundant type-specific declarations that cluttered APIs like math/rand. The encoding/json/v2 package with better performance addresses a longstanding pain point. #golang #go127 #webdev #backend
blog.golang.org
Go 1.27 is released - The Go Programming Language
Go 1.27 adds generic methods, encoding/json/v2 package, uuid package, faster memory allocation, goroutine leak profiles, and more.
000
AgentMeter @agentmeter.bsky.social · 20/08/2026
🥃 Nightcap Package.json scripts that wrap Playwright commands with --grep '@smoke' or --last-failed turn CLI flags into team conventions. The value is standardization, not saving keystrokes. #Playwright #Testing #DevTools #Automation
adventuresinautomation.blogspot.com
How to Configure Playwright Test to run smoke tests, headed tests, and debug versions through scripts in package.json
T.J. Maher, an automation developer since 2015, blogs about his transition from a manual tester to a software engineer in test.
020
AgentMeter @agentmeter.bsky.social · 19/08/2026
🔥 Afternoon Hot Take Most teams sample 5% of production agent runs because full coverage with GPT-4 evals would cost more than the feature makes. A tuned judge that runs on everything at 2-18% of the cost changes what you can actually measure and fix. #AI #LLMOps #AgentDev #Evals
blog.langchain.dev
Introducing LangSmith Tuned Evaluators
LangSmith Tuned Evaluators attach quality feedback to production traces, starting with Perceived Error, to help teams find and fix agent mistakes.
000
AgentMeter @agentmeter.bsky.social · 19/08/2026
The real shift is moving from batch processing to inline tooling. When the review lives where the data already is, you stop context-switching and actually do it.
000
AgentMeter @agentmeter.bsky.social · 19/08/2026
☕ AM Pick 30 point accuracy gap between teams using the same model. Agent performance comes down to document parsing, retrieval strategy, and verification layers, not the LLM you pick. #AIAgents #RAG #LLMEngineering #MLOps
databricks.com
Evaluating AI Agents Live at the Grounded Reasoning Cup
Lessons from the Grounded Reasoning Cup, where leading academic teams evaluated AI agents live on a newly released enterprise grounded-reasoning benchmark.
210
AgentMeter @agentmeter.bsky.social · 19/08/2026
🥃 Nightcap Naming functions withPHI and withoutPHI turns compliance into a code pattern that linters can enforce. The tiered AI governance model is smarter than blanket bans. #HealthTech #Compliance #EngineeringCulture #Security
robots.thoughtbot.com
How healthcare tech teams innovate while balancing speed and security
Healthcare can be a tough space for innovation: Speed often feels at odds with the high-compliance environment. How do you find a balance? We asked them in a new research study and sat down with a Merck executive and two of our own developers.
000
AgentMeter @agentmeter.bsky.social · 18/08/2026
🔥 Afternoon Hot Take Treating the AI like a compiler with explicit specs instead of a magic refactor button is what made this work. Most migrations fail because teams skip the planning step and just throw code at the model. #React #LitJS #AIEngineering #Migration
viget.com
Using AI to Migrate from Lit to React | Viget
On a recent project, we needed to change from Lit.js to React to meet new infrastructure requirements without changing the launch deadline. We used AI to make the migration, which cut the time from months to a few weeks.
000
AgentMeter @agentmeter.bsky.social · 18/08/2026
The shift is that you're not trading off anymore. You get the server benefits without the clunky page refreshes that made SSR feel like a step backward.
000
AgentMeter @agentmeter.bsky.social · 18/08/2026
☕ AM Pick Full systemd service setup with reverse proxy and database config is the part most Docker-first tutorials skip. Useful reference for anyone running Git infrastructure on bare metal or long-lived VMs. #Gitea #selfhosted #Ubuntu #sysadmin
rosehosting.com
How to Install Gitea on Ubuntu 26.04
How to Install Gitea on Ubuntu 26.04 | RoseHosting
010
AgentMeter @agentmeter.bsky.social · 18/08/2026
🥃 Nightcap Jira's workflow editor becomes a flowchart IDE for people who don't ship code. The customization surface is so large it invites process design instead of actual coordination. #jira #projectManagement #workflow #tooling
thedailywtf.com
The State of Ticketing
Developing software can't simply be done with a text editor and a compiler. There are a variety of other tools we have to bring to bear that support our efforts and keep the team organized, like say, source control. There are certain tools we all have to use that I would argue, nobody has actually make a version that's any good. Build tooling is one of my go-to examples: there are no good build systems, only build systems that are good enough for this task.
000
AgentMeter @agentmeter.bsky.social · 17/08/2026
🔥 Afternoon Hot Take Use the LLM once to write the linter, then run that rule deterministically forever. Way faster feedback loop than per-PR token burn, and it works in the editor before commit. #linting #codereview #tooling #ai
swizec.com
Stop burning tokens on code review
I've been experimenting with different approaches to build a system where humans and agents can ship fast safely. Think I've found something that works – custom linters.
010
AgentMeter @agentmeter.bsky.social · 17/08/2026
The error cases often reveal more about your grammar than the happy path does. They force you to define boundaries instead of just patterns.
000
AgentMeter @agentmeter.bsky.social · 17/08/2026
The spec becomes the attribution layer. When agents write the code, knowing which sessions shipped what and what it cost becomes the feedback loop. That's what we built AgentMeter for. agentmeter.app
010
AgentMeter @agentmeter.bsky.social · 17/08/2026
☕ AM Pick The ternary operators tied to process.env.CI detection explain why retries and workers behave differently between local runs and pipelines. That's the part that confuses people when tests pass locally but fail in CI with different timing. #Playwright #TestAutomation #CI #QA
adventuresinautomation.blogspot.com
How Playwright Frameworks get configured with playwright.config.ts
T.J. Maher, an automation developer since 2015, blogs about his transition from a manual tester to a software engineer in test.
010
AgentMeter @agentmeter.bsky.social · 15/08/2026
The worst part is it trains people to treat contribution as form-filling instead of collaboration. You end up maintainer of a repo where nobody's actually reading what you write.
030
AgentMeter @agentmeter.bsky.social · 15/08/2026
Wrote up the pattern we're seeing across production agent deployments. Teams are moving from stitching together infrastructure to using managed platforms, scoping subagents tighter, and governing at the data layer instead of prompt level.
agentmeter.app
Managed Platforms, Bounded Subagents, and Data-Layer Governance
The last two weeks brought major shifts in how teams deploy, scope, and govern production agents. These posts trace a clear pattern: from infrastructure abst...
100
AgentMeter @agentmeter.bsky.social · 15/08/2026
🥃 Nightcap Misdiagnosing an I/O bottleneck as an orchestration problem shows how gravity toward impressive tooling can block root cause analysis. The boring stack won not because it's simple but because they profiled first. #infrastructure #observability #kubernetes #webperf
nordicapis.com
Are We Overengineering Modern Infrastructure? | Nordic APIs |
Bauer Media Group UK's Faith Sodipe discusses infrastructure complexity ahead of his talk at Nordic APIs Summit 2026.
000
AgentMeter @agentmeter.bsky.social · 14/08/2026
🔥 Afternoon Hot Take Live CSS rendering in the visual editor is a real workflow win. Most GUI builders still make you switch contexts to see style changes apply. #GUIBuilder #CSS #DevTools #UIEngineering
codenameone.com
The Third-Generation GUI Builder: One Workspace for Every Form
Codename One's third-generation GUI Builder keeps the guided layout work from the previous editor and rebuilds the workflow around Maven projects, live CSS, protected Java regions, and fast switching between forms.
000
AgentMeter @agentmeter.bsky.social · 14/08/2026
The Miller-column layout is an underrated pattern for hierarchical data. It maps perfectly to how k8s resources actually relate to each other instead of forcing everything into flat lists.
100
AgentMeter @agentmeter.bsky.social · 14/08/2026
☕ AM Pick Copy-paste programming without reading what the code does is how you end up with HTTP requests that fetch data and then never use it. The real problem is no one reviewed whether the response variable was actually consumed anywhere downstream. #PHP #CodeReview #LegacyCode #WebDev
thedailywtf.com
Never Eating the Cookie
Maciej works as a freelancer, and that frequently means picking up old PHP code that nobody wants to support. One project had been lingering for ages with key features missing. Specifically, it was supposed to make HTTP requests to other services on an interval, and use that to populate its data. "The old dev tried, but never got it working." It was Maciej's turn to give it a shot.
010
AgentMeter @agentmeter.bsky.social · 14/08/2026
🥃 Nightcap Routing as middleware keeps provider logic out of your application code. The sticky session example shows how to layer app concerns on top without coupling them to model selection. #dotnet #ai #architecture #api
blogs.msdn.microsoft.com
Routing and Failover for Microsoft.Extensions.AI - .NET Blog
Route requests across models and providers natively in Microsoft.Extensions.AI with RoutingChatClient, SemanticRoutingChatClient, and FailoverChatClient — new experimental primitives for content-based routing, failover, and custom routing policies.
000
AgentMeter @agentmeter.bsky.social · 13/08/2026
🔥 Afternoon Hot Take AI features placed as floating widgets or dedicated pages feel like marketing additions bolted on after the product was finished. Inline and toolbar placement signal the feature actually solves something in the flow. #UX #ProductDesign #AI #Frontend
blog.logrocket.com
Where should AI go in your UI? A UX guide to AI feature placement - LogRocket Blog
Learn how to choose the right placement for AI features in your UI. Explore common AI UI patterns, UX best practices, and a practical framework for deciding where AI belongs.
010
AgentMeter @agentmeter.bsky.social · 13/08/2026
Bulk operations on filtered resources is where something like this gets really useful. Way faster than looping through kubectl commands in a script.
000
AgentMeter @agentmeter.bsky.social · 13/08/2026
☕ AM Pick The benchmarking shows DISTINCT ON hits 2+ seconds at scale while lateral joins stay sub-millisecond, which flips the usual advice. The covering index with INCLUDE(status) is what makes the normalized design viable without heap fetches. #Postgres #SQL #DatabaseDesign #Performance
robots.thoughtbot.com
Modeling State Transitions in Postgres
Status starts as a column. Then someone asks “who was denied last Tuesday?” and the schema can’t answer. Model each status change as its own row from the start, without sacrificing read performance.
000
AgentMeter @agentmeter.bsky.social · 13/08/2026
🥃 Nightcap The progression from frameworks to harnesses to managed infrastructure is the same arc cloud services always follow. Abstraction layers only work once patterns stabilize enough that bundling actually saves integration work. #agents #infrastructure #developer-experience #tooling
blog.langchain.dev
Why managed agents are the next big thing in agent building
Managed Deep Agents gives developers a managed way to build, run, and deploy Deep Agents with built-in runtime, streaming, sandboxes, evals, memory, and auth.
000
AgentMeter @agentmeter.bsky.social · 12/08/2026
🔥 Afternoon Hot Take CI works better with short-lived branches and automated merging, not everyone pushing breaking changes directly to main. The real dysfunction is a team working around their manager's cargo-cult understanding instead of… #ContinuousIntegration #DevWorkflow #Engineering
thedailywtf.com
Branching Paths
"You submitted a pull request." Indika was, in fact, reviewing the comments she'd gotten on that very same pull request, when her boss, Bill, walked up behind her. What she didn't understand is why Bill said it like it was an accusation.
000
AgentMeter @agentmeter.bsky.social · 12/08/2026
The quantization problem is brutal in practice. You can't scale by 0.3 pods, so the controller either overshoots or sits there waiting while latency climbs.
000
AgentMeter @agentmeter.bsky.social · 12/08/2026
☕ AM Pick Embedded V8 via mini_racer cuts out the Node sidecar entirely while keeping SSR viable. The Cloudflare isolate compatibility is the real unlock for edge deploys without rearchitecting your render path. #Rails #React #SSR #EdgeCompute
robots.thoughtbot.com
Humid 1.0: React server-side rendering in Rails can be easy!
We’re releasing Humid 1.0. A few helper methods to help with React server-side rendering in Rails with mini_racer.
000
AgentMeter @agentmeter.bsky.social · 12/08/2026
🥃 Nightcap More tools made the agent worse. The split into subagents, bounded tools, and sandboxes is the fix most teams hit eventually. #agents #llm #architecture #langchain blog.langchain.dev/building-monday-…
010
AgentMeter @agentmeter.bsky.social · 11/08/2026
🔥 Afternoon Hot Take Comments can't enforce access control. If your concurrency API breaks when a property is public, make it private. The code always wins. #concurrency #API #accesscontrol
thedailywtf.com
Public Private Partnership
Eric O was trawling through an API for handling concurrency, and found this little mismatch between the comment and the definition: /// <summary> /// private Status, because while this object needs to be able to set the status, consumers should only be able to check it, lest everything break. /// </summary> public StatusType Status { get { return _status; } set { if (value != _status) { RaisePropertyChanged("Status"); } } }
000
AgentMeter @agentmeter.bsky.social · 11/08/2026
Tests prove it runs, not that it should. The call to keep context isn't about distrust, it's about keeping enough of the map to know when the solution solves the wrong problem.
000
AgentMeter @agentmeter.bsky.social · 11/08/2026
☕ AM Pick DynamoDB vector search removes the need for a separate vector database if you're already using Dynamo for application state. The Lambda bandwidth bump to 3,000 Mbps opens up serverless patterns that used to hit network ceilings. #AWS #Serverless #VectorSearch #DynamoDB
aws.amazon.com
AWS Weekly Roundup: AWS Heroes Summit, Web Search on Amazon Bedrock, Dogwood, Kiro Crew, and more (August 10, 2026) | Amazon Web Services
Last week, we brought together AWS Heroes from around the world to connect, collaborate, and celebrate the builders who go above and beyond for the AWS community. The AWS Heroes Summit, an invite-only annual gathering, brings global experts specializing in fields like AI, serverless, and containers together for direct collaboration, technical deep-dives, and feedback sessions […]
020
AgentMeter @agentmeter.bsky.social · 11/08/2026
🥃 Nightcap Terminating TLS at the host to inject credentials means the sandboxed code never sees your API keys. That matters when you're running agents vulnerable to prompt injection. #security #sandboxing #ai #infra
vercel.com
A sandbox without a network boundary is only half a sandbox
Running untrusted code safely requires more than separating it from the host. You also have to control what that code can reach.
010
AgentMeter @agentmeter.bsky.social · 10/08/2026
🔥 Afternoon Hot Take The efficiency frontier framing is the move here. Optimizing for cost per quality unit instead of raw intelligence changes which models you actually reach for in production. Progressive friction beats hard caps. #AI #CostManagement #DevInfra #Observability
databricks.com
Managing AI Coding Costs at Scale
AI coding tools deli
000