Sign in

Arize AI

@arize.bsky.social
121 followers 31 following 474 posts

Arize is an AI engineering platform focused on evaluation and observability. It helps engineers develop, evaluate, and observe AI applications and agents.

PostsRepliesMedia
Arize AI @arize.bsky.social · 24/09/2026
When using an LLM-as-a-judge, or playing with prompts, you need access to the models you want at high speed, including open source, and your fine-tuned models. Arize AX now has support for Fireworks as an AI provider, so you can use all your favorite models. arize.com/docs/ax/sec...
arize.com
Fireworks AI - Arize AX Docs
Integrate with Fireworks AI as an AI Provider to run serverless models, fine-tunes, and dedicated deployments in Arize AX prompts and evaluations
010
Arize AI @arize.bsky.social · 23/09/2026
Plenty of internal model gateways don't issue API keys. They issue short-lived tokens from an authorization server. AX custom model endpoints now speak that: OAuth 2.0 client credentials. arize.com/docs/ax/sec...
arize.com
Custom Model Endpoints - Arize AX Docs
Connect any OpenAI-compatible custom model endpoint to Arize AX
000
Arize AI @arize.bsky.social · 23/09/2026
Two model drops in one day?? We've got both! Opus 5.5 and Sol 6 day 0 support in evals and playgrounds, right here! arize.com
000
Arize AI @arize.bsky.social · 22/09/2026
Arize AX now renders videos added to spans when you view traces. Great for digging into issues in your AI video pipelines that your evals, or managed agents like Signal, have surfaced.
000
Arize AI @arize.bsky.social · 18/09/2026
For all you Codex fans out there, we've just added support for Codex as the harness for your managed agents. Configure your managed agents with Codex as the harness, select your models, skills, attach repos, and define your task. Then let your managed agent rip!
010
Arize AI @arize.bsky.social · 17/09/2026
Tracing is the part you can't skip. No traces means no evals, no improvement loop. So we made it one command. npx evals Run this in your terminal to instrument any app, or your coding agent. Walkthrough here: youtu.be/To_txP8i_es
youtube.com
Add Tracing to Any AI Agent in One Command | Arize AX
Add tracing to an AI agent in minutes with Arize AX. In this quick ...
100
Arize AI @arize.bsky.social · 16/09/2026
We wouldn't hire someone today because they're good at digging through millions of rows of data. We'd hire someone who can send an agent to do it. On how AI operations teams are restructuring: youtu.be/4zG3OPDGqac
youtube.com
How AI Ops Teams Will Manage Fleets of Agents | Arize
As AI teams deploy more agents, traces, and evals, the operational ...
000
Arize AI @arize.bsky.social · 04/09/2026
Did you know LLM-as-a-judge works for multi-modal as well as text? We've just released a hands-on guide with a video walkthrough that shows you how to build an evaluator for images. arize.com/docs/ax/coo...
arize.com
Evaluate Receipt Agents with an Image Judge - Arize AX Docs
Trace receipt-image extraction in Arize AX and use an image-aware LLM judge to evaluate whether structured output is visually grounded.
000
Reposted by Arize AI
QCon @qconferences.com · 20/08/2026
Laurie Voss, co-founder of npm and now Head of Developer Relations at Arize AI, says code review is being quietly replaced by automated verification harnesses.
211
Arize AI @arize.bsky.social · 11/08/2026
OpenTelemetry’s GenAI semantic conventions are becoming a common way for frameworks and agent platforms to describe model calls, tools, retrieval, token usage, and more. Arize AX now offers native support, so gen_ai.* spans arrive as structured AI traces. Learn more: arize.com/blog/arize-...
100
Arize AI @arize.bsky.social · 10/08/2026
The EU AI Act turns principles like fairness, transparency, human oversight, and robustness into evidence product and engineering teams may need to produce for a specific AI system over time. Here’s what that means for developers: arize.com/blog/demyst...
arize.com
Demystifying the EU AI Act for AI product and engineering teams
An engineering guide to turning EU AI Act principles into traces, evaluations, annotations, and release evidence product and engineering teams can actually demonstrate.
100
Arize AI @arize.bsky.social · 30/07/2026
@HamelHusain keeps stopping eval reviews for the same reason: the model isn't broken, but the product is. In part 2 of our series Rise of the Agent Engineer, Hamel walks through why ambiguous inputs, generic metrics, and disconnected reviews make AI evaluations misleading, and how to fix them.
100
Arize AI @arize.bsky.social · 29/07/2026
What if your agents got better every time they failed? Today, we’re launching Signal. It continuously reviews production traces, finds issues, and turns them into an investigation with evidence, root cause, and a proposed fix. Your engineers decide what ships. arize.com/blog/from-s...
100
Arize AI @arize.bsky.social · 28/07/2026
Want to master the full workflow of shipping reliable AI agents? Laurie's workshop from AI Engineer World's Fair, "Evaluating and Shipping AI Agents That Work," is now a free, self-paced course on Arize University. Earn a certificate you can share on LinkedIn by completing 13 episodes that cover:
110
Arize AI @arize.bsky.social · 28/07/2026
Two AI observability lessons from @bookingcom: - An agent latency spike came from a model running without the appropriate service tier. - Multi-turn eval scores fell because long URLs were being added back into the conversation history, causing the context to balloon.
100
Arize AI @arize.bsky.social · 24/07/2026
Claude Opus 5 is available in Arize AX, supported across Anthropic API, AWS Bedrock, and Vertex AI! Opus 5 brings major improvements for long-running agents in coding and professional work. Instrument, evaluate, and improve your agents with day 0 support for Anthropic's newest model.
120
Arize AI @arize.bsky.social · 23/07/2026
Token price tells you what a model costs to call. It does not tell you what it costs to finish the job. Arize and @FireworksAI_HQ benchmarked 10 models across 2,400 agent runs, including Kimi K3.
100
Arize AI @arize.bsky.social · 22/07/2026
Gemini 3.6 Flash and Gemini 3.5 Flash-lite are now available in Arize AX! Both models are supported on direct Gemini and Vertex AI providers with configurable thinking levels, 65K max output tokens, and audio file input.
100
Arize AI @arize.bsky.social · 17/07/2026
Coding agents become trustworthy when verification is designed into the development loop. At @cursor_ai, that loop helps roughly 30-40% of pull requests in its ecosystem merge without human review. Here is how it works: 🧵
200
Arize AI @arize.bsky.social · 15/07/2026
Git can show you the code diff. But only traces and evals can show you the behavioral diff. That’s the difference between an agent change that looks right and one you can prove is better.
100
Arize AI @arize.bsky.social · 14/07/2026
Your AI agent generated 40 PRs. Great. But how many were merged? How much rework did they create? And what did each successful PR actually cost? Tokens, prompts, and output volume measure motion. But measuring AI productivity requires connecting traces to outcomes.
100
Arize AI @arize.bsky.social · 11/07/2026
There's a lot of talk about loops recently. But the term “loop” currently describes at least four different architectures: execution, task, product, and system (plus the human oversight loop governing them).
110
Arize AI @arize.bsky.social · 10/07/2026
GPT-5.6 support just went live in Arize AX. 🚀 Now available: 🌞 gpt-5.6-sol 🌍 gpt-5.6-terra 🌙 gpt-5.6-luna Compare all three side-by-side in the Prompt Playground, plug them into LLM-as-a-judge evals, and watch them in production - all in one place. Try it 👇 app.arize.com/
000
Arize AI @arize.bsky.social · 08/07/2026
An agent was told: “make the tests pass.” It deleted the tests. That story from WorkOS founder Michael Grinich is funny on its face. But it's also the exact reason agent engineering is getting harder. Full conversation below.
110
Arize AI @arize.bsky.social · 07/07/2026
Most teams hear the same advice: “add evals.” But when you’re staring at a real LLM app, that advice gets vague fast. Should your first eval be an integration test? A golden dataset? A CI gate? A dashboard metric? An LLM judge?
110
Arize AI @arize.bsky.social · 06/07/2026
Agent harnesses are becoming the durable layer of AI coding workflows, according to @aparnadhinak. The model answers once. The harness turns that answer into a loop: context, tools, permissions, edits, tests, failures, retries, recovery, and traces.
200
Arize AI @arize.bsky.social · 02/07/2026
The difference between an agent that works and one that games you comes down to one habit: a good eval. ✅ Spell out the shortcuts you won't accept ✅ Check that the work actually happened ✅ Try to cheat it yourself first ✅ Test it on real traffic
100
Arize AI @arize.bsky.social · 01/07/2026
“Which model is cheapest?” is the wrong question. The better question: which model is cheapest per successful task? A model that looks cheap at the token level can be expensive if it needs retries, tool calls, or human cleanup to finish the task correctly. /1
100
Arize AI @arize.bsky.social · 30/06/2026
50 traces. That’s how much data @HamelHusain says you need to start building evals that actually work. Pull them. Label them with a PM. Cluster the failures. Pick the highest-impact one. Write a binary eval You’ll learn more in an hour by doing this than in a month of dashboard watching.
100
Arize AI @arize.bsky.social · 30/06/2026
A year ago, 200 instructions was the ceiling. Today it's closer to 2,000 - and up to 5,000 on the strongest models. The capacity problem is largely solved, but the verification problem is wide open.
100
Arize AI @arize.bsky.social · 29/06/2026
@SnorkelAI will be in the Evals track with us at AIE! Rustem Feyzkhanov will be talking about how agent evaluation is moving beyond reviewing static traces and into executable simulation environments that let you test agents repeatedly across realistic tasks.
100
Arize AI @arize.bsky.social · 29/06/2026
Your LLM gateway can be more than a router. With TrueFoundry + Arize, every model call becomes an OpenInference trace: spans for auth, model resolution, provider calls, token usage, latency, cost, and more. Here's how: arize.com/blog/trace-...
000
Arize AI @arize.bsky.social · 29/06/2026
Saturday in Dalston: a room full of people trying to make London even more lovable (and livable!), in eight hours. Will we see you there? luma.com/maxxing-london
luma.com
Londonmaxxing 003: Maxxing London Hackathon · Luma
How can we make London better to build in, live in and fall in love with? London is on a generational run. Billions in investment is flowing into the city,…
000
Arize AI @arize.bsky.social · 28/06/2026
You can have production-quality evals running in minutes. Our Solutions Architect Ankur Duggal @Anky488 is leading a hands-on workshop at AI Engineer World's Fair, walking through how to stand up a production eval pipeline in minutes using Arize Agent Skills, no prior setup required.
110
Arize AI @arize.bsky.social · 28/06/2026
Come see what we've been building at Arize. Our Fuad Ali is leading a live walk-through of the latest features in Arize on Day 1 at AI Engineer World's Fair.
100
Arize AI @arize.bsky.social · 27/06/2026
Two workshops. Two chances to help you move from vibes-based development to production-ready AI agents.
110
Arize AI @arize.bsky.social · 27/06/2026
Excited to have Uber on the Evals track with us at AIE. Soumya Gupta and Jai Chopra are presenting how @Uber used closed-loop evals for their food photography enhancement agent.
100
Arize AI @arize.bsky.social · 26/06/2026
What does a failing agent look like when all your metrics say it's fine? Our Strategy lead Dat Ngo is unpacking one of the most common failure patterns in production AI: agents that report success without actually succeeding.
120
Arize AI @arize.bsky.social · 26/06/2026
Long-running agents need more than logs. You need to know: 👉 which LLM/tool call failed? 👉 did quality regress? 👉 what did it cost? We teamed up with Restate to show durable agents + Arize Phoenix traces/evals in action. Read more: www.restate.dev/blog/arize-...
restate.dev
Resilient, observable agents with Restate and Arize Phoenix
Combine Restate and Arize Phoenix for agent reliability, observability, and evaluation in one stack.
000
Arize AI @arize.bsky.social · 26/06/2026
Voice agents are one of the fastest-growing categories in AI and one of the hardest to debug.
100
Arize AI @arize.bsky.social · 25/06/2026
Code review was designed for a world where humans wrote all the code. What happens when that world is gone? Our Head of DevRel Laurie Voss will be at the AI Engineer World's Fair to talk about how the unit of trust changes when agents write the code.
110
Arize AI @arize.bsky.social · 25/06/2026
What if your observability platform didn't just tell you something was wrong, but fixed it?
100
Arize AI @arize.bsky.social · 24/06/2026
A coding agent scored 81.4% on SWE-bench. 24% of its trajectories simply ran git log to copy the answer out of commit history. OpenAI audited SWE-bench Verified next, found 59.4% of problems had flawed tests, and stopped using it.
100
Arize AI @arize.bsky.social · 23/06/2026
Annotation queues now support trace records, alongside spans and dataset examples. This unblocks trace-level human review.
100
Arize AI @arize.bsky.social · 23/06/2026
Attending PlatformCon 2026 virtually? Check out this virtual session from Glyn Darkin on building evaluation infrastructure for agents across the entire SDLC, and see how Arize fits into your entire pipeline. Tuesday 24th June, 1pm CEST, 4am PT. platformcon.com/sessions/ev...
platformcon.com
PlatformCon 2026 | The #1 platform engineering event
PlatformCon is the world’s largest platform engineering conference. Join us for a full week of programming featuring 150+ curated talks and 30+ hours of hands-on workshops.
000
Arize AI @arize.bsky.social · 22/06/2026
At every conference this year, the same question: how exactly does observability plug into the framework I picked? Project Rosetta Stone is the answer. Same agent. 22+ frameworks. 3 observability tiers each. Diff against the no-obs baseline, see the exact files to touch. arize.com/blog/projec...
100
Arize AI @arize.bsky.social · 22/06/2026
Hackathon premise we can get behind: build something a Londoner would actually use, in one day. 100+ applications in already, so move quick. July 4th, Dalston: luma.com/maxxing-london
luma.com
Londonmaxxing 003: Maxxing London Hackathon · Luma
How can we make London better to build in, live in and fall in love with? London is on a generational run. Billions in investment is flowing into the city,…
000
Arize AI @arize.bsky.social · 18/06/2026
When an agent fails, have you considered looking at the harness? The loop around the model decides how tasks are decomposed, how tools are called, how context is managed, how errors are recovered from, and what gets traced.
200
Arize AI @arize.bsky.social · 17/06/2026
Recently @anthropic.com shipped Dreams. OpenAI shipped Dreaming V3. Same word, opposite architectures. A UIUC paper landed the same week showing one of these patterns drops accuracy from 100% to 54% on ARC-AGI. One left itself an escape hatch. One did not. arize.com/blog/two-la...
000
Arize AI @arize.bsky.social · 16/06/2026
Most agent orchestration debates are arguing about the wrong layer. 👀 Frameworks answer how agent control flow is expressed. Runtimes answer how agents recover, resume, and survive long tasks. Observability answers how teams find out what actually happened.
421