Sign in

arize-phoenix

@arize-phoenix.bsky.social
84 followers 11 following 414 posts

Open-Source AI Observability and Evaluation app.phoenix.arize.com

PostsRepliesMedia
arize-phoenix @arize-phoenix.bsky.social · 12h
Come join Elizabeth in our free webinar: luma.com/arizeai-ben...
luma.com
Benchmarking AI agents & tool use with Harbor and Arize Phoenix · Luma
When you change an agent's model, prompt, tools, or environment, how do you know whether the new version performs better? Production traces show what happens…
000
arize-phoenix @arize-phoenix.bsky.social · 12h
As agents gain capabilities due to skills, MCPs, and codemode, evaluating agents becomes increasingly more difficult because you need to test across multiple environments and harnesses. Arize Phoenix your benchmarking Harbor runs so pinpointing improvements and regressions becomes tractable.
100
arize-phoenix @arize-phoenix.bsky.social · 01/10/2026
below: a support agent traced in phoenix. jev triages the ticket in 237ms and judges whether the reply is ready to send in 154ms, bookending two ~1s openai calls. 2.2s end to end, under a cent. ships today in our typescript and python auto-instrumentors.
020
arize-phoenix @arize-phoenix.bsky.social · 01/10/2026
now you can compose both. a small model handles routing and guardrails, and the llm steps in only when needed. openinference's new decision span captures the question, choice, confidence, and probabilities, so system one calls show up right next to the reasoning they gate.
100
arize-phoenix @arize-phoenix.bsky.social · 01/10/2026
your agent can think fast and slow. kahneman's "thinking, fast and slow" describes two modes: system one: fast, cheap, almost automatic (jev) system two: slow, deliberate, analytical (llms)
100
arize-phoenix @arize-phoenix.bsky.social · 30/09/2026
Benchmarking AI agents & tool use with Harbor + Arize Phoenix. Join us Oct 8 at 11am PT / 2pm ET to learn how to compare agents or models over the same task set, separate behavioral scores from infrastructure failures, and inspect ATIF traces. luma.com/arizeai-ben...
luma.com
Benchmarking AI agents & tool use with Harbor and Arize Phoenix · Luma
When you change an agent's model, prompt, tools, or environment, how do you know whether the new version performs better? Production traces show what happens…
110
arize-phoenix @arize-phoenix.bsky.social · 23/09/2026
Not all tokens cost the same. Cache writes have an upfront cost, but they save you money on subsequent turns. Premature or inefficient compaction can cause cache misses that add both latency and cost. In Phoenix, you can search LLM calls across a turn and see how context evolves.
001
arize-phoenix @arize-phoenix.bsky.social · 22/09/2026
Webinar coming up: benchmarking AI agents & tool use with Harbor and Arize Phoenix. Learn how to compare agents and models on the same tasks, diagnose regressions, and inspect scores, errors, and traces. Oct 8 · 11am PT / 2pm ET Save your spot: luma.com/arizeai-ben...
luma.com
Benchmarking AI agents & tool use with Harbor and Arize Phoenix · Luma
When you change an agent's model, prompt, tools, or environment, how do you know whether the new version performs better? Production traces show what happens…
000
arize-phoenix @arize-phoenix.bsky.social · 22/09/2026
When conducting experiments, it might be important to keep track of the data being fed into the agent using a Git like version control system - especially if you are collaborating with others. Phoenix now has bulk editing for datasets so that you can track the dataset patches.
000
arize-phoenix @arize-phoenix.bsky.social · 18/09/2026
Python: pypi.org/project/ope... JS: www.npmjs.com/package/@ar... Jev early access: typesafe.ai Phoenix docs: arize.com/docs/phoeni...
arize.com
TypeSafe AI - Phoenix
TypeSafe AI is a model provider that answers typed questions about structured state, returning schema-validated results instead of free-form text.
000
arize-phoenix @arize-phoenix.bsky.social · 18/09/2026
Here it is blocking a support draft that promised free shipping for life, at confidence 1.0.
110
arize-phoenix @arize-phoenix.bsky.social · 18/09/2026
So: OpenInference instrumentation for the TypeSafe SDK, live now. Every Jev call is a span in Phoenix, inside the trace of the agent it's judging, with the full probability distribution attached. Plain OTel, any collector.
100
arize-phoenix @arize-phoenix.bsky.social · 18/09/2026
When decisions get that cheap, intelligence goes into every conditional in your code. Those decisions need to be observable.
100
arize-phoenix @arize-phoenix.bsky.social · 18/09/2026
TypeSafe came out of stealth this week with Jev, the first System One Model: a frontier model that never generates prose. Unstructured state in, typed probabilistic decisions out. 70 to 500 ms, 0% type errors, calibrated probabilities on every answer, $0.042/MTok in and output free.
110
arize-phoenix @arize-phoenix.bsky.social · 17/09/2026
“Did the agent actually do what the user asked?” 👀 Phoenix now has docs for built-in evals that help answer that from real traces: ✅ completeness 🔎 retrieval relevance 🧭 hallucination 🛡️ toxicity 🔒 PII detection Start here: arize.com/docs/phoeni...
arize.com
SDK Eval Metrics - Phoenix
Ready-to-use evaluation metrics for measuring LLM application quality
000
arize-phoenix @arize-phoenix.bsky.social · 15/09/2026
If you're building agents in Java, we'd love to hear how this fits your setup and what's missing. arize.com/docs/phoeni...
arize.com
Google ADK tracing for Java - Phoenix
Instrument Google ADK for Java with OpenInference and export traces to Phoenix.
000
arize-phoenix @arize-phoenix.bsky.social · 15/09/2026
With the new OpenInference instrumentation for ADK Java, every agent invocation, LLM call, and tool execution shows up as a trace in Phoenix. Same experience the Python ADK folks have had, now for the JVM.
100
arize-phoenix @arize-phoenix.bsky.social · 15/09/2026
ADK gives you fine-grained control over how an agent behaves: orchestration, tool use, and state all live in code rather than config. That's great for versioning and debugging, but once agents start branching across tools and sub-agents, you still need to see what actually happened on a given run.
100
arize-phoenix @arize-phoenix.bsky.social · 15/09/2026
Phoenix now supports Google's Agent Development Kit for Java.
100
arize-phoenix @arize-phoenix.bsky.social · 10/09/2026
OpenInference now supports tracing Agno teams. Teams are a way to coordinate groups of agents to solve complex tasks. Thank you to our OSS contributors that helped add tracing and session support!
000
arize-phoenix @arize-phoenix.bsky.social · 10/09/2026
Computer-use is a new frontier that many agents are trying to tackle now. With the release of models like Astra, the capabilities of agents is ever expanding. Thanks to an OSS community member PR, you now can trace the computer use of your OpenAI agents! pypi.org/project/ope...
000
arize-phoenix @arize-phoenix.bsky.social · 08/09/2026
Does contributor make sense for evals, prototyping, and synthetic or public data? You decide. For anything touching customer data, the standard tier is the privacy-conscious choice. developer.meta.com/ai/models/m...
developer.meta.com
Muse Spark 1.3 | Meta
Trained for agentic workflows and optimized for competitive coding performance.
000
arize-phoenix @arize-phoenix.bsky.social · 08/09/2026
The release ships as two model IDs with identical specs. muse-spark-1.3 costs $1.25 / $4.25 per million tokens. muse-spark-1.3-contributor costs $0.10 / $0.20, roughly 12x cheaper on input and 21x on output, in exchange for letting Meta train on your prompts and completions.
100
arize-phoenix @arize-phoenix.bsky.social · 08/09/2026
Meta is holding standard pricing flat while claiming its biggest jump yet on coding and agentic work, and has something no other frontier lab has: a pricing table where training consent is a part of the model name.
110
arize-phoenix @arize-phoenix.bsky.social · 08/09/2026
Phoenix now supports Meta's Muse Spark 1.3. Muse Spark 1.3 landed last week in the same 48 hours as Astra and Fable, with no blog post, and it deserves more attention than it got.
101
arize-phoenix @arize-phoenix.bsky.social · 04/09/2026
• deletePrompt, transferTraces, setProjectRetentionPolicy, and getCurrentUser in the TypeScript client, plus token counts on sessions. Full notes: docs.arize.com/phoenix/rel...
arize.com
Release Notes - Phoenix
The latest from the Phoenix team.
010
arize-phoenix @arize-phoenix.bsky.social · 04/09/2026
• A PII Detection evaluator for conversation records, tool calls and retrieved documents included. Python and TypeScript. • Prompt versions over REST. POST a new version to an existing prompt and tag it in the same call.
110
arize-phoenix @arize-phoenix.bsky.social · 04/09/2026
It searches for duplicates first, drafts the issue with the relevant traces and spans linked, and shows you the exact repo, title, and body before anything is posted. It files with your own GitHub token, which stays in your browser. Phoenix never stores it. Also in this release:
100
arize-phoenix @arize-phoenix.bsky.social · 04/09/2026
Phoenix 20.5 is out, and the part I keep showing people is small but a huge time saver: PXI can now file the GitHub issue for the bug it just found.
110
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
Phoenix is built to stay model-provider agnostic, so swap it in, trace it, and let the evals decide. Who's running MiniMax in production? How does it hold up past a few dozen tool calls?
000
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
The detail we care most about from an observability standpoint: interleaved thinking. Reasoning lives in think blocks between tool calls and is meant to be kept in the message history, so every span in a trace carries the model's reasoning for that step, not just the tool call and result.
100
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
For agents, active parameters are what you pay for in latency and cost, and agent loops resend context, retry tool calls, and run long. M3 adds sparse attention and a 1M-token window on top, so long tool histories stay in context instead of getting summarized away.
100
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
We think MiniMax is underrated for agentic workloads, and the reason is architectural. The M-series is sparse MoE: M2 routes to ~10B of 230B parameters per forward pass, and M3 to ~23B of 428B. That is why M2 fits on 4xH100s at FP8 and why list pricing sits around $0.30/$1.20 per million tokens.
110
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
it now leads CyberGym at 84.5%, and Z.ai is running a public disclosure ledger for the 2,400+ real-world vulnerabilities the model has surfaced.
000
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
The result: a 50% improvement over 5.2 on their in-house Code Bench, Terminal-Bench 3.0 from 4.6 → 28.3, DeepSWE v1.1 from 46.2 → 66.9, and open-weights SOTA on Terminal-Bench 3.0 and Agents' Last Exam. Cyber capability emerged faster than they expected.
100
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
Arize Phoenix now supports Z.ai and GLM-5.3. GLM-5.3 uses the same base as GLM-5.2. Every gain came from post-training. Z.ai scaled up long-horizon task environments (some representing days of real engineering work) and trained on them with their open-source slime RL framework.
110
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
Are any of you using Fable class models for your custom agents? Or do you reserve them for long horizon coding proglems?
010
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
Long-horizon, multi-step autonomous agents is exactly the workload that's hardest to debug without observability. Trace those extended agent runs, run evals on their outputs, and track spend!
100
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
Pricing is $10/$50 per million input/output tokens, with cache reads cut 75% to $0.25/M. Anthropic claims ~25% lower typical costs and up to ~45% for agentic workloads.
100
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
- 55.8% on Terminal-Bench 4.0 (up from 42.0% for Fable 5) - 52.6% on Terminal-Bench-Science (up from 24.7%) - 65.0% on Humanity's Last Exam with tools.
100
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
Claude Fable 5.1 and Mythos 5.1 are the same underlying model at two safeguard levels. Fable 5.1 is generally available everywhere (API, AWS, GCP, Azure) as claude-fable-5-1, while Mythos 5.1 is restricted to vetted US organizations via cyber/life-sciences verification programs.
110
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
The prime directive of the Phoenix project is to go from bad agent behavior to fix as quickly as possible. Introducing GitHub integration. Ask PXI to find problematic traces and automatically create and de-duplicate issues so your engineers and agents can quickly find a fix.
020
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
For full details check out our post: arize.com/blog/phoeni...
arize.com
Arize Phoenix has a built-in MCP server that lets your agents query traces with SQL
Phoenix's built-in MCP server adds read-only SQL tools and code mode so coding agents can query traces ~17x cheaper than retrieval-only MCP workflows.
010
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
Phoenix now exposes SQL over MCP. The agent asks for the schema (one call), writes one query, and the database is safely esposed as a tool. The answer comes back as a small aggregate.
100
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
"Which model has the worst p95 latency this week?" has no endpoint. Until now, an agent's only option was to page raw spans through the REST API and aggregate them itself. Hundreds round trips, and every byte of JSON landed in the context window.
100
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
Fixed APIs answer fixed questions. Agents ask questions nobody anticipated.
100
arize-phoenix @arize-phoenix.bsky.social · 29/08/2026
docs.arize.com/phoenix/rel...
arize.com
08.25.2026: Smarter PXI Workflows and Retrieval Evaluation - Phoenix
PXI completes multi-step UI workflows with approval-gated changes, a new evaluator scores retrievals from any source, and new trace and REST tools speed up analysis and administration.
000
arize-phoenix @arize-phoenix.bsky.social · 29/08/2026
• Faster trace analysis: use natural-language filters, one-click chart zoom, and dedicated annotation columns. • Expanded REST API: manage retention assignments and discover model providers.
100
arize-phoenix @arize-phoenix.bsky.social · 29/08/2026
This week we make moving from questions to answers faster. • Smarter PXI workflows complete multi-step UI tasks now run via a JavaScript sandbox so the assistant can steer the UI for you via code. • Retrieval Relevance: Evaluate retrievals from vector search, tools, MCP, web search, or SQL.
100
arize-phoenix @arize-phoenix.bsky.social · 26/08/2026
Docs: platform.claude.com/docs/en/bui... #LLMObservability #AIEngineering #Tracing #Claude
platform.claude.com
Refusals and fallback
How Claude Fable 5 and Claude Opus 5 return classifier refusals and how to retry refused requests on a fallback model.
000