Sign in

arize-phoenix

@arize-phoenix.bsky.social
83 followers 11 following 409 posts

Open-Source AI Observability and Evaluation app.phoenix.arize.com

PostsRepliesMedia
arize-phoenix @arize-phoenix.bsky.social · 9h
Benchmarking AI agents & tool use with Harbor + Arize Phoenix. Join us Oct 8 at 11am PT / 2pm ET to learn how to compare agents or models over the same task set, separate behavioral scores from infrastructure failures, and inspect ATIF traces. luma.com/arizeai-ben...
luma.com
Benchmarking AI agents & tool use with Harbor and Arize Phoenix · Luma
When you change an agent's model, prompt, tools, or environment, how do you know whether the new version performs better? Production traces show what happens…
000
arize-phoenix @arize-phoenix.bsky.social · 23/09/2026
Not all tokens cost the same. Cache writes have an upfront cost, but they save you money on subsequent turns. Premature or inefficient compaction can cause cache misses that add both latency and cost. In Phoenix, you can search LLM calls across a turn and see how context evolves.
001
arize-phoenix @arize-phoenix.bsky.social · 22/09/2026
Webinar coming up: benchmarking AI agents & tool use with Harbor and Arize Phoenix. Learn how to compare agents and models on the same tasks, diagnose regressions, and inspect scores, errors, and traces. Oct 8 · 11am PT / 2pm ET Save your spot: luma.com/arizeai-ben...
luma.com
Benchmarking AI agents & tool use with Harbor and Arize Phoenix · Luma
When you change an agent's model, prompt, tools, or environment, how do you know whether the new version performs better? Production traces show what happens…
000
arize-phoenix @arize-phoenix.bsky.social · 22/09/2026
When conducting experiments, it might be important to keep track of the data being fed into the agent using a Git like version control system - especially if you are collaborating with others. Phoenix now has bulk editing for datasets so that you can track the dataset patches.
000
arize-phoenix @arize-phoenix.bsky.social · 18/09/2026
TypeSafe came out of stealth this week with Jev, the first System One Model: a frontier model that never generates prose. Unstructured state in, typed probabilistic decisions out. 70 to 500 ms, 0% type errors, calibrated probabilities on every answer, $0.042/MTok in and output free.
110
arize-phoenix @arize-phoenix.bsky.social · 17/09/2026
“Did the agent actually do what the user asked?” 👀 Phoenix now has docs for built-in evals that help answer that from real traces: ✅ completeness 🔎 retrieval relevance 🧭 hallucination 🛡️ toxicity 🔒 PII detection Start here: arize.com/docs/phoeni...
arize.com
SDK Eval Metrics - Phoenix
Ready-to-use evaluation metrics for measuring LLM application quality
000
arize-phoenix @arize-phoenix.bsky.social · 15/09/2026
Phoenix now supports Google's Agent Development Kit for Java.
100
arize-phoenix @arize-phoenix.bsky.social · 10/09/2026
OpenInference now supports tracing Agno teams. Teams are a way to coordinate groups of agents to solve complex tasks. Thank you to our OSS contributors that helped add tracing and session support!
000
arize-phoenix @arize-phoenix.bsky.social · 10/09/2026
Computer-use is a new frontier that many agents are trying to tackle now. With the release of models like Astra, the capabilities of agents is ever expanding. Thanks to an OSS community member PR, you now can trace the computer use of your OpenAI agents! pypi.org/project/ope...
000
arize-phoenix @arize-phoenix.bsky.social · 08/09/2026
Phoenix now supports Meta's Muse Spark 1.3. Muse Spark 1.3 landed last week in the same 48 hours as Astra and Fable, with no blog post, and it deserves more attention than it got.
101
arize-phoenix @arize-phoenix.bsky.social · 04/09/2026
Phoenix 20.5 is out, and the part I keep showing people is small but a huge time saver: PXI can now file the GitHub issue for the bug it just found.
110
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
We think MiniMax is underrated for agentic workloads, and the reason is architectural. The M-series is sparse MoE: M2 routes to ~10B of 230B parameters per forward pass, and M3 to ~23B of 428B. That is why M2 fits on 4xH100s at FP8 and why list pricing sits around $0.30/$1.20 per million tokens.
110
arize-phoenix @arize-phoenix.bsky.social · 03/09/2026
Arize Phoenix now supports Z.ai and GLM-5.3. GLM-5.3 uses the same base as GLM-5.2. Every gain came from post-training. Z.ai scaled up long-horizon task environments (some representing days of real engineering work) and trained on them with their open-source slime RL framework.
110
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
Claude Fable 5.1 and Mythos 5.1 are the same underlying model at two safeguard levels. Fable 5.1 is generally available everywhere (API, AWS, GCP, Azure) as claude-fable-5-1, while Mythos 5.1 is restricted to vetted US organizations via cyber/life-sciences verification programs.
110
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
The prime directive of the Phoenix project is to go from bad agent behavior to fix as quickly as possible. Introducing GitHub integration. Ask PXI to find problematic traces and automatically create and de-duplicate issues so your engineers and agents can quickly find a fix.
020
arize-phoenix @arize-phoenix.bsky.social · 01/09/2026
Fixed APIs answer fixed questions. Agents ask questions nobody anticipated.
100
arize-phoenix @arize-phoenix.bsky.social · 29/08/2026
This week we make moving from questions to answers faster. • Smarter PXI workflows complete multi-step UI tasks now run via a JavaScript sandbox so the assistant can steer the UI for you via code. • Retrieval Relevance: Evaluate retrievals from vector search, tools, MCP, web search, or SQL.
100
arize-phoenix @arize-phoenix.bsky.social · 26/08/2026
The model you requested might not be the model that responds.
100
arize-phoenix @arize-phoenix.bsky.social · 20/08/2026
Did you know that your browser might have an LLM built right into it? With the Prompt API, you can take advantage of local models like developer.chrome.com/docs/ai/pro...). Cool right? Well now you also can get a nice streaming UI to chat with these models, built right into Arize Phoenix.
000
arize-phoenix @arize-phoenix.bsky.social · 07/08/2026
4 instrumentors shipped today thanks to our cracked OSS community. 🤖 AG2: multi-agent conversations, group chats, and tool calls ⚡ Together AI: fast inference across OSS models 🧠 Cohere: enterprise chat, search, and RAG 🦙 Ollama: local models galore arize.com/docs/phoeni...
arize.com
Integrations - Phoenix
Connect Phoenix with your favorite AI frameworks, LLM providers, and tools
031
arize-phoenix @arize-phoenix.bsky.social · 04/08/2026
AI Queries and Browser AI Phoenix filters now understand English. Type "responses with apologies" and get a real filter expression back: ▎ 'sorry' in output.value or 'apolog' in output.value
101
arize-phoenix @arize-phoenix.bsky.social · 03/08/2026
The experiment said the fix works, but experiments aren't production. Real spans, new prompt, zero errors. Will PXI get it right? youtu.be/biMD5LHdXQk
000
arize-phoenix @arize-phoenix.bsky.social · 01/08/2026
Balancing cost, speed, and correctness can be a tricky balance. That's why we need good visuals to figure out the right sweet spot!
000
arize-phoenix @arize-phoenix.bsky.social · 31/07/2026
Customizable visualizations of your agent traces are here. Track cache hits, online eval degradations, tool call errors, all in real-time so you can quickly identify problems in production. What does your agent operations center look like?
100
arize-phoenix @arize-phoenix.bsky.social · 29/07/2026
New Sandbox in Phoenix - Monty (Python)!
000
arize-phoenix @arize-phoenix.bsky.social · 28/07/2026
In an agent session, a cache miss re-bills your entire history at full input price. That's why a "continue" after a coffee break can cost more than the model's actual answer. Cache read vs. write isn't a footnote in your bill. Earendril's post is a must read. earendil.com/posts/promp...
010
arize-phoenix @arize-phoenix.bsky.social · 27/07/2026
The agent kept writing SQL querying a "products" table but the table is called "catalog." PXI analyzes the problem and proposes a fix all on its own: youtu.be/CN8VwB_4V_Y
000
arize-phoenix @arize-phoenix.bsky.social · 25/07/2026
Search -> Select -> Download -> Coding Agent. Feed critical traces back to Agents. No fancy words for this one: this is plain old-fashioned debugging. Just now with Agents.
000
arize-phoenix @arize-phoenix.bsky.social · 21/07/2026
Your agent's SQL tool calls start failing with "no such table: products." Before you attempt a fix, take a snapshot of the failures with a dataset. Just ask PXI to do it for you: youtu.be/gb_78l788Ls
youtube.com
Making a Dataset from Failing Traces with Phoenix and PXI
An agent's SQL tool calls start erroring out with "no such table: p...
000
arize-phoenix @arize-phoenix.bsky.social · 19/07/2026
phoenix is now is an oauth2 authorization server! `px auth login` opens your browser, you get a short-lived user-scoped token. no long-lived api keys. and admins can audit / revoke every cli and mcp grant from one screen. remote mcp server ships in beta too - more on this soon.
010
arize-phoenix @arize-phoenix.bsky.social · 13/07/2026
Phoenix's agent PXI can propose annotation categories, annotate traces en mass, and even suggest fixes based on patterns of failures. Learn how to do it: youtu.be/iF25CqJv4tA
000
arize-phoenix @arize-phoenix.bsky.social · 13/07/2026
Our favorite tools are the ones that have maximum customizability. Last week we added customizable charts, command K, and recent searches. this week we've added custom column ordering. Built to help you have the tables and dashboards you need to monitor agents day in and day out.
010
arize-phoenix @arize-phoenix.bsky.social · 12/07/2026
What does your agent operation center look like?
010
arize-phoenix @arize-phoenix.bsky.social · 11/07/2026
Experiment Baselining and Charts When trying to determine if a new model is up to the task, you need to factor in many dimensions. performance - measured by evals latency - is the model fast enough to give you the right UX tokens - how chatty is the model to achieve the result
110
arize-phoenix @arize-phoenix.bsky.social · 11/07/2026
Do you love localfirst development? @nearestnabors.com will be speaking on how to get frontier LLM results on device using Phoenix and prompt engineering techniques on Sunday, July 12, 14:15. See you soon, Berlin! www.localfirstconf.com/
localfirstconf.com
Local-First Conf 2026
Join us for the third edition of Local-First Conf. Connect with a rapidly-growing community in an intimate setting. Berlin 12-14th July 2026.
011
arize-phoenix @arize-phoenix.bsky.social · 10/07/2026
Phoenix has an agent built into it now! PXI can help you find the crucial traces you should actually be reading. Short video on how to do use PXI in your daily flow: youtu.be/5lUgdRFf4ZI
youtube.com
Find the traces that matter with Phoenix and PXI
Hundreds of traces and they all look fine? 🤔 Here's how to find the...
010
arize-phoenix @arize-phoenix.bsky.social · 09/07/2026
⌘K is in Phoenix. jump to any project, dataset, prompt, or experiment. No mouse required.
010
arize-phoenix @arize-phoenix.bsky.social · 07/07/2026
Agent traces and trajectories are growing increasingly longer and more complex. We've seen some traces 1000s of spans deep. That's why we've added trace search. Search across a trace and the UI will now show you the call stack to the spans you are looking for across workflows and sub-agent calls.
010
arize-phoenix @arize-phoenix.bsky.social · 07/07/2026
The OSS team has a mantra: 3 clicks to glory. We're not there yet but we're getting closer!
100
arize-phoenix @arize-phoenix.bsky.social · 04/07/2026
Fable 5 might be amazing, but it can't tend the grill 🍔 Happy Fourth of July from our team to yours. We hope you are getting to spend the day with friends, family, and loved ones.
010
arize-phoenix @arize-phoenix.bsky.social · 01/07/2026
PXI now supports sub-agent streaming so you can inspect the execution of the sub-agents you kick off. We've started moving much the skills and tools available in the main agent into the sub-agents to empower interesting delegation patterns.
110
arize-phoenix @arize-phoenix.bsky.social · 30/06/2026
PXI (Phoenix Intelligence) now runs in your terminal! You can now use PXI without leaving your terminal. It's the same agent that powers the in-browser experience, now available as an interactive chat in your shell. npm install -g @arizeai/phoenix-cli@latest > pxi
000
arize-phoenix @arize-phoenix.bsky.social · 24/06/2026
📦🏷️🔖🔑 This week in Phoenix: - server-side bash for PXI subagents, - label management on list pages, - trace-level annotations everywhere - and OAuth2 role-override preservation. arize.com/docs/phoeni...
000
arize-phoenix @arize-phoenix.bsky.social · 19/06/2026
Meet PXI (pronounced "pixie") 🎉 the AI engineering agent we built into Phoenix. Hand it the investigation instead of scrolling through traces by hand, and it works through your telemetry the way a coding agent works through code. The full story of how we built it 👉 arize.com/blog/meet-pxi/
arize.com
Meet PXI: the AI engineering agent inside Phoenix
PXI is the open-source AI engineering agent built into Phoenix. Hand it a failing trace, an evaluator, or a prompt, and it investigates your telemetry for you.
011
arize-phoenix @arize-phoenix.bsky.social · 17/06/2026
📊 Phoenix 17.7.0 makes your token usage legible. New token detail charts break prompt + completion tokens into their parts, over time: • Prompt → input, cache read, cache write, audio • Completion → output, reasoning, audio
220
arize-phoenix @arize-phoenix.bsky.social · 08/06/2026
Phoenix just hit 10,000 GitHub stars! Three years ago, Phoenix didn't exist. Arize was a closed-source company. A small team was asked to change that. Catch the full interview with the team who made it happen and where AI observability is going next: arize.com/phoenix-10k
010
arize-phoenix @arize-phoenix.bsky.social · 26/05/2026
"Don't trust. Evaluate." @nearestnabors.com set out to replace Sonnet with Gemma. The evals showed a quantifiably better option. Full walkthrough: capability evals + prompt engineering to ship a local 3B that matches Sonnet, 2x faster, $0/call. Built with Phoenix. arize.com/blog/how-to...
031
arize-phoenix @arize-phoenix.bsky.social · 21/05/2026
Phoenix now lets you compose evaluation strategies in code. Most eval tooling hands you a fixed menu of judge templates. Real evaluation is rarely that tidy.
131
arize-phoenix @arize-phoenix.bsky.social · 19/05/2026
The Arize DevRel team wants to connect with Phoenix users like you. What you're tracing, what's working, what's rough? Schedule time with the team here: cal.com/team/arize-...
cal.com
Phoenix and AX Free Weekly User Interviews | Arize Dev Rel | Cal.com
Phoenix and AX Free Weekly User Interviews
000
arize-phoenix @arize-phoenix.bsky.social · 15/05/2026
Something we’ve been playing with and liking a lot: Give every coding agent its own observability stack.
100