Sign in

Arize AI

@arize.bsky.social
121 followers 31 following 474 posts

Arize is an AI engineering platform focused on evaluation and observability. It helps engineers develop, evaluate, and observe AI applications and agents.

PostsRepliesMedia
Arize AI @arize.bsky.social · 24/09/2026
When using an LLM-as-a-judge, or playing with prompts, you need access to the models you want at high speed, including open source, and your fine-tuned models. Arize AX now has support for Fireworks as an AI provider, so you can use all your favorite models. arize.com/docs/ax/sec...
arize.com
Fireworks AI - Arize AX Docs
Integrate with Fireworks AI as an AI Provider to run serverless models, fine-tunes, and dedicated deployments in Arize AX prompts and evaluations
010
Arize AI @arize.bsky.social · 23/09/2026
Plenty of internal model gateways don't issue API keys. They issue short-lived tokens from an authorization server. AX custom model endpoints now speak that: OAuth 2.0 client credentials. arize.com/docs/ax/sec...
arize.com
Custom Model Endpoints - Arize AX Docs
Connect any OpenAI-compatible custom model endpoint to Arize AX
000
Arize AI @arize.bsky.social · 23/09/2026
Two model drops in one day?? We've got both! Opus 5.5 and Sol 6 day 0 support in evals and playgrounds, right here! arize.com
000
Arize AI @arize.bsky.social · 22/09/2026
Arize AX now renders videos added to spans when you view traces. Great for digging into issues in your AI video pipelines that your evals, or managed agents like Signal, have surfaced.
000
Arize AI @arize.bsky.social · 18/09/2026
For all you Codex fans out there, we've just added support for Codex as the harness for your managed agents. Configure your managed agents with Codex as the harness, select your models, skills, attach repos, and define your task. Then let your managed agent rip!
010
Arize AI @arize.bsky.social · 17/09/2026
Tracing is the part you can't skip. No traces means no evals, no improvement loop. So we made it one command. npx evals Run this in your terminal to instrument any app, or your coding agent. Walkthrough here: youtu.be/To_txP8i_es
youtube.com
Add Tracing to Any AI Agent in One Command | Arize AX
Add tracing to an AI agent in minutes with Arize AX. In this quick ...
100
Arize AI @arize.bsky.social · 16/09/2026
We wouldn't hire someone today because they're good at digging through millions of rows of data. We'd hire someone who can send an agent to do it. On how AI operations teams are restructuring: youtu.be/4zG3OPDGqac
youtube.com
How AI Ops Teams Will Manage Fleets of Agents | Arize
As AI teams deploy more agents, traces, and evals, the operational ...
000
Arize AI @arize.bsky.social · 04/09/2026
Did you know LLM-as-a-judge works for multi-modal as well as text? We've just released a hands-on guide with a video walkthrough that shows you how to build an evaluator for images. arize.com/docs/ax/coo...
arize.com
Evaluate Receipt Agents with an Image Judge - Arize AX Docs
Trace receipt-image extraction in Arize AX and use an image-aware LLM judge to evaluate whether structured output is visually grounded.
000
Reposted by Arize AI
QCon @qconferences.com · 20/08/2026
Laurie Voss, co-founder of npm and now Head of Developer Relations at Arize AI, says code review is being quietly replaced by automated verification harnesses.
211
Arize AI @arize.bsky.social · 11/08/2026
A practical rule: - Use OpenInference when you control instrumentation. - Use OpenTelemetry GenAI conventions when your framework already emits them or you primarily control the OTLP export. - Mixed environment? AX supports both. Learn more: arize.com/blog/arize-...
000
Arize AI @arize.bsky.social · 11/08/2026
There’s some useful history here. We created OpenInference at Arize before the broader OpenTelemetry GenAI conventions were mature enough for dependable AI-specific instrumentation. Today, more frameworks emit gen_ai.* telemetry natively, so the two increasingly need to work together.
100
Arize AI @arize.bsky.social · 11/08/2026
That matters when you don’t control the runtime. If a managed agent platform already emits OpenTelemetry GenAI telemetry, you can send it to AX over OTLP and use it for tracing, evals, token analysis, and debugging without maintaining a custom conversion processor.
100
Arize AI @arize.bsky.social · 11/08/2026
OpenTelemetry’s GenAI semantic conventions are becoming a common way for frameworks and agent platforms to describe model calls, tools, retrieval, token usage, and more. Arize AX now offers native support, so gen_ai.* spans arrive as structured AI traces. Learn more: arize.com/blog/arize-...
100
Arize AI @arize.bsky.social · 10/08/2026
Here's a practical place to start: - Pick one high-risk use case. - Identify the principle you would struggle to demonstrate today. - Instrument the workflow, define the metric and owner, then connect the result to your release process. Full engineering guide: arize.com/blog/demyst...
arize.com
Demystifying the EU AI Act for AI product and engineering teams
An engineering guide to turning EU AI Act principles into traces, evaluations, annotations, and release evidence product and engineering teams can actually demonstrate.
000
Arize AI @arize.bsky.social · 10/08/2026
Documentation works better when your system generates it during normal operation. Traces preserve what happened. Annotations record human review. Evals measure behavior. CI records whether changes passed your release criteria.
100
Arize AI @arize.bsky.social · 10/08/2026
Be careful with AI-generated evaluation scores. If an LLM judge measures fairness or bias, calibrate it against human reviewers on representative examples and inspect where they disagree. The metric needs evidence behind it, too.
200
Arize AI @arize.bsky.social · 10/08/2026
Much of that evidence already comes from normal AI engineering workflows: traces, evals, annotations, CI, and audit logs. For each requirement, define the metric, owner, threshold, review path, and release consequence.
100
Arize AI @arize.bsky.social · 10/08/2026
The EU AI Act turns principles like fairness, transparency, human oversight, and robustness into evidence product and engineering teams may need to produce for a specific AI system over time. Here’s what that means for developers: arize.com/blog/demyst...
arize.com
Demystifying the EU AI Act for AI product and engineering teams
An engineering guide to turning EU AI Act principles into traces, evaluations, annotations, and release evidence product and engineering teams can actually demonstrate.
100
Arize AI @arize.bsky.social · 30/07/2026
Want to watch the full interview? Check it out on our YouTube channel: www.youtube.com/watch?v=LGK...
youtube.com
Stop Blaming the Model: Fixing the AI Product Bottleneck | Rise of the AI Engineer | Hamel Husain
When an AI feature fails in production, developers almost always bl...
000
Arize AI @arize.bsky.social · 30/07/2026
You can also read our writeup on our conversation with Hamel to learn why useful evals start before the metric: arize.com/blog/rise-o...
arize.com
Hamel Husain explains why AI evals fail before the evaluation begins
Hamel Husain explains why ambiguous inputs, generic metrics, and disconnected review workflows can make AI evaluations misleading, and how developers can build a better process around real production data.
100
Arize AI @arize.bsky.social · 30/07/2026
@HamelHusain keeps stopping eval reviews for the same reason: the model isn't broken, but the product is. In part 2 of our series Rise of the Agent Engineer, Hamel walks through why ambiguous inputs, generic metrics, and disconnected reviews make AI evaluations misleading, and how to fix them.
100
Arize AI @arize.bsky.social · 29/07/2026
See Signal go from production failure to pull request in a live virtual event on July 30. qrs.ly/l5hg8lx
luma.com
From Signal to PR: Evaluating and Shipping AI Agents That Work · Luma
When an agent fails in production, finding the bad output is only the beginning. Teams still have to identify the pattern, trace it back to the underlying…
000
Arize AI @arize.bsky.social · 29/07/2026
What if your agents got better every time they failed? Today, we’re launching Signal. It continuously reviews production traces, finds issues, and turns them into an investigation with evidence, root cause, and a proposed fix. Your engineers decide what ships. arize.com/blog/from-s...
100
Arize AI @arize.bsky.social · 28/07/2026
→ Traces → Code evals vs. LLM judges → Datasets and experiments → Online evals and monitors → Closing the loop with coding agents Get started here: courses.arize.com/l/pdp/arize...
120
Arize AI @arize.bsky.social · 28/07/2026
Want to master the full workflow of shipping reliable AI agents? Laurie's workshop from AI Engineer World's Fair, "Evaluating and Shipping AI Agents That Work," is now a free, self-paced course on Arize University. Earn a certificate you can share on LinkedIn by completing 13 episodes that cover:
110
Arize AI @arize.bsky.social · 28/07/2026
Here's a practical look at how Booking.com approaches AI observability across agents and traditional ML with Arize AX: arize.com/blog/bookin...
arize.com
From traditional ML to AI agents: How Booking.com scales AI observability with Arize
Building an AI-native observability stack for agentic AI and traditional ML at Booking.com
000
Arize AI @arize.bsky.social · 28/07/2026
The broader playbook on agent observability from @bookingcom: - Trace model calls, retrieval, tools, guardrails, and fallbacks - Monitor quality alongside latency and token usage - Pull real failures into datasets for experimentation - Centralize sampling, PII redaction, and telemetry routing
210
Arize AI @arize.bsky.social · 28/07/2026
Both problems became much easier to fix once the team could connect production signals to individual traces, configurations, and conversations.
100
Arize AI @arize.bsky.social · 28/07/2026
Two AI observability lessons from @bookingcom: - An agent latency spike came from a model running without the appropriate service tier. - Multi-turn eval scores fell because long URLs were being added back into the conversation history, causing the context to balloon.
100
Arize AI @arize.bsky.social · 24/07/2026
→ arize.com/docs/ax
arize.com
Arize AX - Arize AX Docs
AI Engineering Platform
000
Arize AI @arize.bsky.social · 24/07/2026
Claude Opus 5 is available in Arize AX, supported across Anthropic API, AWS Bedrock, and Vertex AI! Opus 5 brings major improvements for long-running agents in coding and professional work. Instrument, evaluate, and improve your agents with day 0 support for Anthropic's newest model.
120
Arize AI @arize.bsky.social · 23/07/2026
Developers and product managers should optimize for cost per successful task, then use evals and traces to understand where each model belongs. Get the benchmark data: arize.com/blog/cost-p...
arize.com
Cost per successful task: Benchmarking Kimi K3, GPT-5.5, and 8 more AI models
We tested 10 open and closed models (including Kimi K3) across 2,400 agent runs, measuring pass rates, cost per successful task, and more.
000
Arize AI @arize.bsky.social · 23/07/2026
In his tests, Arize's Head of DevRel @seldo found: • Kimi K3 nearly matched GPT-5.5 overall • Different models won on different task types • Routing improved cost and coverage • Retries and silent failures changed the economics completely
100
Arize AI @arize.bsky.social · 23/07/2026
Token price tells you what a model costs to call. It does not tell you what it costs to finish the job. Arize and @FireworksAI_HQ benchmarked 10 models across 2,400 agent runs, including Kimi K3.
100
Arize AI @arize.bsky.social · 22/07/2026
Trace, evaluate, and run experiments with Google's latest flash-family models today. → arize.com/docs/ax
arize.com
Arize AX - Arize AX Docs
AI Engineering Platform
000
Arize AI @arize.bsky.social · 22/07/2026
Gemini 3.6 Flash and Gemini 3.5 Flash-lite are now available in Arize AX! Both models are supported on direct Gemini and Vertex AI providers with configurable thinking levels, 65K max output tokens, and audio file input.
100
Arize AI @arize.bsky.social · 17/07/2026
A practical adoption path: • Start with one bounded workflow • Define the evidence required to merge • Route changes by risk • Turn corrections into eval cases • Let failed evals trigger diagnosis Go inside Cursor’s agent factory: arize.com/blog/inside...
arize.com
Inside Cursor's agent factory: how it verifies AI-written code
As background agents take on more implementation work, Cursor is rebuilding the software development lifecycle around risk scores, developer-like environments, video evidence, and review systems that learn from every human correction.
130
Arize AI @arize.bsky.social · 17/07/2026
Failed evals can begin the diagnosis process automatically. The failure triggers a workflow with the relevant traces and logs already attached, allowing the system to move directly from detection to investigation.
100
Arize AI @arize.bsky.social · 17/07/2026
Human review also improves future runs. When Cursor’s Bugbot misses an issue and a developer corrects it, that correction can become a rule or regression case that tests whether the review agent catches the same failure again.
100
Arize AI @arize.bsky.social · 17/07/2026
The developer environment becomes part of the evaluation harness. When an agent can boot the product, use the interface, and record the result, reviewers can inspect the behavior before reading the diff. Tests still cover the broader state space.
100
Arize AI @arize.bsky.social · 17/07/2026
Every pull request produces an evidence package that can include CI results, security review, a risk score, and a recording of the agent exercising the feature. The merge policy evaluates that evidence, then routes consequential changes to the right human owner.
110
Arize AI @arize.bsky.social · 17/07/2026
Coding agents become trustworthy when verification is designed into the development loop. At @cursor_ai, that loop helps roughly 30-40% of pull requests in its ecosystem merge without human review. Here is how it works: 🧵
200
Arize AI @arize.bsky.social · 15/07/2026
Every failure becomes a future regression test. That’s how we stop vibe-checking agent changes and start engineering them. Read the full walkthrough: arize.com/blog/kiro-c...
arize.com
Kiro CLI observability: trace and evaluate agent changes with Arize Skills
Use Arize Skills with Kiro CLI to trace coding-agent changes, build datasets from failures, run experiments, and validate prompts before shipping.
100
Arize AI @arize.bsky.social · 15/07/2026
The important shift is that the coding agent isn’t only making the change. It’s helping operate the verification loop around that change. And because Kiro supports Agent Skills, the workflow can live with the agent instead of being reconstructed prompt by prompt.
100
Arize AI @arize.bsky.social · 15/07/2026
- Update the prompt, retrieval, tools, or other harness logic. - Run an experiment against the same examples. - Compare eval results and catch regressions before shipping.
100
Arize AI @arize.bsky.social · 15/07/2026
We just published a hands-on workflow for closing that gap with Kiro CLI + Arize Skills: - Trace each Kiro turn and tool call, including model details, duration, and credit usage. - Find concrete failures in those traces. - Turn the failures into a reusable dataset.
100
Arize AI @arize.bsky.social · 15/07/2026
A prompt, retrieval, or tool change can produce a clean diff while making the agent choose worse tools, miss important context, add latency, or regress on cases you didn’t retest.
100
Arize AI @arize.bsky.social · 15/07/2026
Git can show you the code diff. But only traces and evals can show you the behavioral diff. That’s the difference between an agent change that looks right and one you can prove is better.
100
Arize AI @arize.bsky.social · 14/07/2026
Learn more in our new guide on measuring AI productivity with traces, evals, correlation IDs, and cost per validated outcome: arize.com/blog/how-to...
arize.com
How to measure AI productivity: From LLM token costs to business value with Arize AX
Measure AI productivity by connecting LLM token costs to validated outcomes like merged PRs, resolved tickets, and customer tasks using traces, evals, and Arize AX.
000
Arize AI @arize.bsky.social · 14/07/2026
The missing layer is a join between the trace and a validated outcome: a PR that merged and held, a ticket that stayed resolved, or a customer task that completed.
100