Arize AI @arize.bsky.social · 16hYour LLM bill went up. Cost Agent finds out why. Introducing Cost Agent in Arize AX: a managed agent that finds what’s driving LLM spend, quantifies the impact, and recommends what to fix. One analysis found 35–40% estimated monthly savings. Learn more: arize.com/blog/arize-... 000
Arize AI @arize.bsky.social · 06/10/2026Ask an LLM judge a yes/no question and it often writes you a paragraph. Then you parse it for the answer. Jev-as-a-Judge in Arize AX answers in types: booleans, choices and rubric scores, each with a confidence, several questions per call. Set it up: arize.com/docs/ax/eva... 000
Arize AI @arize.bsky.social · 23/09/2026Two model drops in one day?? We've got both! Opus 5.5 and Sol 6 day 0 support in evals and playgrounds, right here! arize.com 000
Arize AI @arize.bsky.social · 22/09/2026Arize AX now renders videos added to spans when you view traces. Great for digging into issues in your AI video pipelines that your evals, or managed agents like Signal, have surfaced. 000
Arize AI @arize.bsky.social · 18/09/2026For all you Codex fans out there, we've just added support for Codex as the harness for your managed agents. Configure your managed agents with Codex as the harness, select your models, skills, attach repos, and define your task. Then let your managed agent rip! 010
Arize AI @arize.bsky.social · 11/08/2026OpenTelemetry’s GenAI semantic conventions are becoming a common way for frameworks and agent platforms to describe model calls, tools, retrieval, token usage, and more. Arize AX now offers native support, so gen_ai.* spans arrive as structured AI traces. Learn more: arize.com/blog/arize-... 100
Arize AI @arize.bsky.social · 30/07/2026@HamelHusain keeps stopping eval reviews for the same reason: the model isn't broken, but the product is. In part 2 of our series Rise of the Agent Engineer, Hamel walks through why ambiguous inputs, generic metrics, and disconnected reviews make AI evaluations misleading, and how to fix them. 100
Arize AI @arize.bsky.social · 29/07/2026What if your agents got better every time they failed? Today, we’re launching Signal. It continuously reviews production traces, finds issues, and turns them into an investigation with evidence, root cause, and a proposed fix. Your engineers decide what ships. arize.com/blog/from-s... 100
Arize AI @arize.bsky.social · 28/07/2026Want to master the full workflow of shipping reliable AI agents? Laurie's workshop from AI Engineer World's Fair, "Evaluating and Shipping AI Agents That Work," is now a free, self-paced course on Arize University. Earn a certificate you can share on LinkedIn by completing 13 episodes that cover: 110
Arize AI @arize.bsky.social · 28/07/2026Both problems became much easier to fix once the team could connect production signals to individual traces, configurations, and conversations. 100
Arize AI @arize.bsky.social · 28/07/2026Two AI observability lessons from @bookingcom: - An agent latency spike came from a model running without the appropriate service tier. - Multi-turn eval scores fell because long URLs were being added back into the conversation history, causing the context to balloon. 100
Arize AI @arize.bsky.social · 11/07/2026Execution loops are the loop most people picture when they say "agent." But there's more to this space than just that. - Execution: steps in one run - Task: fresh runs against a spec - Product: agents across repo/backlog - System: improve prompts/evals/harnesses 100
Arize AI @arize.bsky.social · 11/07/2026There's a lot of talk about loops recently. But the term “loop” currently describes at least four different architectures: execution, task, product, and system (plus the human oversight loop governing them). 110
Arize AI @arize.bsky.social · 10/07/2026GPT-5.6 support just went live in Arize AX. 🚀 Now available: 🌞 gpt-5.6-sol 🌍 gpt-5.6-terra 🌙 gpt-5.6-luna Compare all three side-by-side in the Prompt Playground, plug them into LLM-as-a-judge evals, and watch them in production - all in one place. Try it 👇 app.arize.com/ 000
Arize AI @arize.bsky.social · 08/07/2026An agent was told: “make the tests pass.” It deleted the tests. That story from WorkOS founder Michael Grinich is funny on its face. But it's also the exact reason agent engineering is getting harder. Full conversation below. 110
Arize AI @arize.bsky.social · 30/06/2026A year ago, 200 instructions was the ceiling. Today it's closer to 2,000 - and up to 5,000 on the strongest models. The capacity problem is largely solved, but the verification problem is wide open. 100
Arize AI @arize.bsky.social · 29/06/2026@SnorkelAI will be in the Evals track with us at AIE! Rustem Feyzkhanov will be talking about how agent evaluation is moving beyond reviewing static traces and into executable simulation environments that let you test agents repeatedly across realistic tasks. 100
Arize AI @arize.bsky.social · 28/06/2026You can have production-quality evals running in minutes. Our Solutions Architect Ankur Duggal @Anky488 is leading a hands-on workshop at AI Engineer World's Fair, walking through how to stand up a production eval pipeline in minutes using Arize Agent Skills, no prior setup required. 110
Arize AI @arize.bsky.social · 28/06/2026Come see what we've been building at Arize. Our Fuad Ali is leading a live walk-through of the latest features in Arize on Day 1 at AI Engineer World's Fair. 100
Arize AI @arize.bsky.social · 27/06/2026Two workshops. Two chances to help you move from vibes-based development to production-ready AI agents. 110
Arize AI @arize.bsky.social · 27/06/2026Excited to have Uber on the Evals track with us at AIE. Soumya Gupta and Jai Chopra are presenting how @Uber used closed-loop evals for their food photography enhancement agent. 100
Arize AI @arize.bsky.social · 26/06/2026What does a failing agent look like when all your metrics say it's fine? Our Strategy lead Dat Ngo is unpacking one of the most common failure patterns in production AI: agents that report success without actually succeeding. 120
Arize AI @arize.bsky.social · 26/06/2026Voice agents are one of the fastest-growing categories in AI and one of the hardest to debug. 100
Arize AI @arize.bsky.social · 25/06/2026Code review was designed for a world where humans wrote all the code. What happens when that world is gone? Our Head of DevRel Laurie Voss will be at the AI Engineer World's Fair to talk about how the unit of trust changes when agents write the code. 110
Arize AI @arize.bsky.social · 25/06/2026What if your observability platform didn't just tell you something was wrong, but fixed it? 100
Arize AI @arize.bsky.social · 15/06/2026Cursor users! A dozen incredibly helpful Arize skills are now available directly in Cursor from the Agent Marketplace! Select "Customize" from the agents sidebar to see the marketplace and click to get them automatically installed. cursor.com/marketplace... 000
Arize AI @arize.bsky.social · 10/06/2026Anthropic's latest and greatest model, Fable, is now available in the prompt playground! 000
Arize AI @arize.bsky.social · 08/06/2026Our open source observability platform Arize Phoenix just crossed 10,000 stars on GitHub. ✨ That number belongs to the people who tested it, broke it, filed issues, opened PRs, asked better questions, and helped turn AI observability into an engineering workflow. 110
Arize AI @arize.bsky.social · 05/06/2026Missed Observe? Catch our own @nearestnabors.com keynoting @pydatalondon.bsky.social Saturday morning at 9am: pydata.org/london2026/ 011
Arize AI @arize.bsky.social · 05/06/2026Signal is one of our out of the box managed agents (you can build your own) and continuously reviews production traces, identifies emerging failure patterns, and groups related issues into investigation reports. 100
Arize AI @arize.bsky.social · 05/06/2026Observe 2026 is a wrap. Yesterday we shared what’s next for Arize AX and our vision for the AI factory for self-improving agents. The focus: helping teams turn production behavior into a repeatable loop for finding issues, investigating root cause, testing fixes, and improving agents. 100
Arize AI @arize.bsky.social · 04/06/2026❤️ One AI Question with Meredith Mende We asked our Head of Talent: Why should you work at Arize? Her answer: It's the people and the obsession. From our hands-on founders to our focus on customer delight, Arize is built on a culture that values talent and real-world impact. #StartupLife #Hiring 110
Arize AI @arize.bsky.social · 02/06/2026⚖️ One AI Question with Tyler Niederwerder We asked our Corporate Counsel: How should you use AI in legal? His answer: Lawyers, stay sharp. AI is great for speed, but human-in-the-loop is mandatory. Always verify AI output to protect privilege and maintain professional standards. #LegalTech #AI 010
Arize AI @arize.bsky.social · 02/06/2026A fireside chat *and* a talk from George Zhang of @openclaw. Happening at Observe in 2 days. Grab your tickets. June 4th. Arize Observe. arize.com/observe 000
Arize AI @arize.bsky.social · 28/05/2026🛑 One AI Question with Robert Mackey We asked our Account Manager: Why Arize? His answer: Stop being reactive. Arize gives you full visibility into every trace and span, moving you from fixing bugs to proactive observing. Ensure your AI agents deliver exactly what your customers expect. #AI 000
Arize AI @arize.bsky.social · 27/05/2026Apache Airflow already orchestrates critical ML and data workflows across the industry. Now it can orchestrate agent improvement loops, too. 100
Arize AI @arize.bsky.social · 26/05/2026🌙 One AI Question with Matt Wilson What does our SVP of Sales do at night? Forget doom-scrolling—he's "doom-prompting." 😂 Building AI agents with Claude and Arize's CLI to push the tech as far as it'll go. Total obsession = total innovation. #AI #Coding #Productivity 030
Arize AI @arize.bsky.social · 22/05/2026Our own Laurie Voss, head of Developer Relations, will be speaking at QDrant's Vector Space Day conference! 110
Arize AI @arize.bsky.social · 21/05/2026Hot off the presses, Gemini 3.5 Flash is now available in the Prompt Playground and throughout Arize AX! app.arize.com 000
Arize AI @arize.bsky.social · 21/05/2026🛠️ One AI Question with Elizabeth Hutton We asked our Senior Software Engineer: Why should you learn about evals? Her answer: Complex AI needs more trust, not less. As systems get smarter, evaluations are the only way to verify performance and ensure your AI is working. #AI #AIEvals #LLM 020
Arize AI @arize.bsky.social · 21/05/2026Prefer boolean or categorical labels when the decision is discrete. Use labels like resolved, partially_resolved, unresolved, or insufficient_evidence. Forced numeric scores often make dashboards look precise while making the measurement less stable. 100
Arize AI @arize.bsky.social · 21/05/2026The most common mistake is asking a judge to “rate helpfulness from 1 to 5.” That creates a confident number, not a reliable measurement. A useful judge needs fixed criteria: target, inputs, allowed labels, decision rules, and examples. 100
Arize AI @arize.bsky.social · 21/05/2026Use code for deterministic checks. Code is cheaper, faster, and more predictable than asking a model to interpret something deterministic. Use LLM judges for semantic checks. 100
Arize AI @arize.bsky.social · 20/05/2026Your AI agent disagrees with your human reviewers all day. Most teams treat that as noise. It's the most useful signal in the system. @jimbobbennett.dev wrote up how to mine the gap and feed it back to the agent. arize.com/blog/self-i... 011
Arize AI @arize.bsky.social · 20/05/2026Docs aren't just for humans anymore. Every coding agent, RAG pipeline, and copilot is reading them too, and they read differently. We built our docs to hold up for every agent that reaches for them, find us near the top of the @mintlify.bsky.social agent score leaderboard: mintlify.com/score 021
Arize AI @arize.bsky.social · 19/05/2026🔧 One AI Question with Fuad Ali We asked our Senior Product Manager: Where do agents fail in production? His answer: It's all about the tools. The more tools an agent has, the more likely it is to make a wrong call. The key is tracing these failures and debugging tool-calling errors. #AIAgents 100
Arize AI @arize.bsky.social · 15/05/2026Production is where reality hits. Looking forward to a joining forces with @MistralAI, @coderhq, and @Workato Monday at the @AWS Agentic AI Partner Showcase to talk about what it actually takes to ship agents. If you're in SF, come on by 👇 www.aicamp.ai/event/event... 020
Arize AI @arize.bsky.social · 14/05/2026🤖 One AI Question with Chris Cooning How is marketing at Arize using AI? "We built a content engine that clones our founders." By training AI on years of content, we automated creation while keeping our brand voice perfectly intact. #MarketingAI #GenerativeAI #ContentStrategy 010
Arize AI @arize.bsky.social · 12/05/2026One AI Question with Cam Young We asked our Strategic AI Solutions Architect: What's a 🔥 take on evals? His answer: Stop guessing and start measuring. Use "LLM-as-a-judge" for nuance, but don't ignore code-based evals for speed and human annotators for ground truth. #AI #AIStrategy #AIEvals #LLM 011
Arize AI @arize.bsky.social · 11/05/2026For that to happen, context need to move beyond dashboards and be accessible through APIs, CLIs, and agent-facing interfaces. Observability assumed a human would read the dashboard, but that's changing. 200