Sign in

MLCommons

@mlcommons.org
209 followers 56 following 227 posts

MLCommons is an AI engineering consortium, built on a philosophy of open collaboration to improve AI systems. Through our collective engineering efforts, we continually measure and improve AI technologies' accuracy, safety, speed, and efficiency.

PostsRepliesMedia
MLCommons @mlcommons.org · 24/09/2026
MLPerf Training v6.1 adds the suite's first LLM post-training benchmark: agentic RL that teaches a 397B-parameter open-weight model to repair real software, scored on pass@4 quality - not just throughput. Details from the task force: mlcommons.org/2026/09/mlperf-traini…
021
MLCommons @mlcommons.org · 23/09/2026
MLCommons joins the EU-funded AIRIS consortium to benchmark next-generation biomedical AI. We're building an evaluation framework to measure mechanism-informed generative AI, with dedicated benchmarks across five disease areas. mlcommons.org/2026/09/join-airis-bi…
000
MLCommons @mlcommons.org · 17/09/2026
MLPerf Inference v6.1 WG chairs Miro Hodak and Frank Han walk through the record round, new silicon (AMD MI350P, Intel Arc Pro B70, NVIDIA Vera Rubin), the first cross-vendor heterogeneous submission, and what it signals for inference engineering. mlcommons.org/2026/09/chai...
000
MLCommons @mlcommons.org · 16/09/2026
MLPerf Inference v6.1 results are live. Record 30 organizations, two new benchmarks (End-to-End RAG + Edge Agentic Inference), and performance gains accelerating — Deepseek R1 up 5.7X in a year, VLM up 2.99X in six months. mlcommons.org/2026/09/mlpe... #AI #Benchmarking #AgenticAI #RAG #EdgeAI
010
MLCommons @mlcommons.org · 15/09/2026
AI Infra 2026: MLPerf Inference v6.1 - Technical deep dive. David Kanter, Miro Hodak (AMD), and Frank Han (Dell) on serving scenarios, reasoning workloads, edge measurements, and what's driving the biggest gains this round. Plus: a preview of MLPerf Endpoints. #AIInfraSummit
000
MLCommons @mlcommons.org · 15/09/2026
Tomorrow: David Kanter unveils MLPerf Inference v6.1 results on the AI Infra Summit main stage. 30 submitters. New accelerators. Next-gen platforms. Results live on mlcommons.org. Plus: what's ahead for MLPerf Endpoints. 10:25 AM. Pier 27, SF. #MLPerf #AIInfraSummit
000
MLCommons @mlcommons.org · 15/09/2026
Our Agent Reliability Profile is a finalist in the C:>DIR Global Agentic Regulator Hackathon. We built Profile Builder + Profile Validator to close the "authorization and supervision gap" for AI agents in financial services. mlcommons.org/2026/09/agent-reliabi…
020
MLCommons @mlcommons.org · 01/09/2026
MLPerf Storage v3.0 is out. 144 results. 19 orgs. 11 first-time submitters. 4 workloads: training, checkpointing, vector DB, KV cache. New efficiency metrics: rack-unit density and provisioned power. S3 API support. 8B to 1250B parameter checkpointing. mlcommons.org/2026/09/mlpe...
100
MLCommons @mlcommons.org · 27/08/2026
MLCommons, Google DeepMind, OpenMined, and AVERI completed the first double-blind evaluation of a closed-weight AI model using AILuminate inside a Trusted Execution Environment. Cryptographic guarantees. Protected weights. Uncontaminated benchmarks. mlcommons.org/2026/08/doub...
mlcommons.org
AILuminate and The First Double-Blind Reliability Evaluation of a Proprietary AI Model
How MLCommons, Google DeepMind, OpenMined, and AVERI used cryptographic guarantees to evaluate AI safety without exposing model weights or benchmark data.
010
MLCommons @mlcommons.org · 26/08/2026
New MLPerf Inference v6.1 benchmark: end-to-end RAG. Measures the full pipeline — ingestion, retrieval, reranking, multi-hop reasoning — not just a single model call. mlcommons.org/2026/08/endtoend-infe…
000
MLCommons @mlcommons.org · 20/08/2026
In case you missed it: MLCommons co-founder & Head of MLPerf @dkanter.bsky.social sat down with the CUBE at NYSE Wired to talk AI benchmarks, inference efficiency, safety, and where AI performance is headed next. www.youtube.com/watch?v=EZF7...
youtube.com
David Kanter, ML Commons | theCUBE + NYSE Wired: Mixture of Experts
YouTube video by SiliconANGLE theCUBE
021
MLCommons @mlcommons.org · 18/08/2026
#MLPerf Client v2.0 is live. New: image generation (Flux.2), agentic AI (SWE + Data Analyst agents), and upgraded LLM benchmarks (Phi 4 Mini, Qwen 3 8B). The industry standard for AI PC performance just got a lot more comprehensive. mlcommons.org/2026/08/mlperf-client…
011
MLCommons @mlcommons.org · 11/08/2026
How do you know if a benchmark result is trustworthy or just benchmark washing? We've seen every way benchmarks can fail or be misused; here's a guide for enterprise teams: 7 questions to ask when a benchmark score is justifying a decision. mlcommons.org/2026/08/benchmark-is-…
000
MLCommons @mlcommons.org · 04/08/2026
LIVE NOW from the NYSE: David Kanter, Co-Founder of MLCommons & Head of MLPerf, is on theCUBE, NYSEWired, talking AI inference benchmarks. We just dropped MLPerf Endpoints v0.7 - a new way to measure GenAI service performance in real-world deployments. v1.0 later this year. Tune in: thecube.net\
000
MLCommons @mlcommons.org · 28/07/2026
MLPerf Endpoints v0.7 is live - a foundation release for AI inference benchmarking. Initial results from Coreweave, Google, Intel, KRAI, and NVIDIA. Four principles: Current, Comprehensive, Comparable, Commentary. Blog: mlcommons.org/2026/07/mlperf-endpoi…
010
MLCommons @mlcommons.org · 22/07/2026
Benchmark a brain tumor AI model on real patient MRI data. No data leaves the hospital. No model weights exposed. That's MedPerf + Google Cloud Confidential Space, demonstrated live at #GoogleCloudNext 2026. mlcommons.org/2026/06/medp...
000
MLCommons @mlcommons.org · 14/07/2026
There's an AI Reliability Map. Most of it is still empty. Benchmarking clusters in a few cells. Enterprise AI readiness needs the full grid. AIRR is mapping what others skip: mlcommons.org/2026/04/airr-map #AIReadiness #EnterpriseAI #AIBenchmarking
010
MLCommons @mlcommons.org · 09/07/2026
MLCommons is introducing an Edge Agentic Inference benchmark for MLPerf Inference v6.1. Single accelerator. One user. Multi-turn tool-calling. Hard 32K context wall. Model: Qwen3.6-27B Q4_K_M Accuracy gate: BFCL v4 Deadline: 7/31/26 mlcommons.org/2026/07/mlperf-infere…
000
MLCommons @mlcommons.org · 08/07/2026
MLPerf Inference now measures multi-turn agents. 990 trajectories, Kimi K2.6 + Qwen3.6-35B-A3B, Pareto-curve performance, three-level accuracy. Built on MLPerf Endpoints. mlcommons.org/2026/07/agentic-infer… #MLPerf #AgenticAI #LLM
022
MLCommons @mlcommons.org · 08/07/2026
AI is in doorbells, hearing aids, and factory sensors. But how do you fairly compare a $1 MCU against a neural accelerator? MLPerf Tiny v1.4: 9 orgs, 25 configs, all measured the same way — standardization makes progress possible. mlcommons.org/2026/07/mlperf-tiny-v…
000
MLCommons @mlcommons.org · 30/06/2026
Standardized benchmarking is the only way to compare AI systems fairly at scale. MLPerf Training v6.0 results are live, featuring: -11,000+ accelerator systems -New first-time submitters -Verified performance on the world's most demanding workloads 🔗 bit.ly/4faJ8mR
100
MLCommons @mlcommons.org · 16/06/2026
MLPerf Training v6.0 results are live! 🎉 For the first time: two Mixture-of-Experts (MoE) benchmarks reflecting where the AI training frontier actually is. 📍 DeepSeek V3 — 671B params (largest in MLPerf history) 📍 GPT-OSS 20B — 21B params Results: mlcommons.org/2026/06/mlpe... 1/4
111
MLCommons @mlcommons.org · 15/06/2026
Tonight at #VLSI2026: MLCommons' David Kanter joins the Evening Panel "AI: Grand Vision or Grand Delusion?" alongside panelists from AMD, SK Hynix, Rapidus & Oxmiq Labs. 8–10 PM, Tapa 1-3.
000
MLCommons @mlcommons.org · 15/06/2026
MLPerf Mobile v6.0 introduces new generative AI benchmarks for running LLMs (Llama 3.1 & 3.2, including the new 1B and 3B models) natively on mobile devices. Test your on-device inference performance. Available on GitHub, iOS & Android: bit.ly/43dlMGE
000
MLCommons @mlcommons.org · 09/06/2026
30 years of coordinated disclosure, one assumption: you can fix the thing once you find the flaw. Open-weight AI breaks that. A new version isn't a patch — every prior copy persists, indefinitely. We're helping write the standard AI evaluation needs. → bit.ly/43t8R3t
000
MLCommons @mlcommons.org · 08/06/2026
Meet GeoCroissant. Built on MLCommons Croissant, it adds Earth observation-specific metadata—from coordinate systems to spatial resolution—to give you better traceability and more reproducible workflows for agentic AI pipelines. bit.ly/3PTLywz
031
MLCommons @mlcommons.org · 02/06/2026
AI systems co-design is too fragmented. Enter MLCommons Chakra (#MLSys2026): an open execution trace ecosystem to bridge software & hardware without exposing IP. Native in @PyTorch, NVIDIA NeMo, & vLLM. Read the paper & explore the traces: bit.ly/4vkYZEP
000
MLCommons @mlcommons.org · 19/05/2026
The median AI benchmark longevity score is 5/100. AILuminate scored 75—but even that degrades over time. To fix this, the @MLCommons AIRR team built the Continuous Prompt Stewardship System to keep risk evaluation fresh and reliable. bit.ly/3On4jrz
000
MLCommons @mlcommons.org · 18/05/2026
What does AI reliability actually require? It comes down to consistently following the right behavioral rules—even under adversarial attack. Meet the AI Reliability Map to guide pre-deployment testing. Explore the framework: bit.ly/4mG7erO #AIReliability #AI
000
MLCommons @mlcommons.org · 14/05/2026
Do tools like OpenClaw signal a turning point for mainstream AI adoption? MLCommons' Dave Graham debated that and more on the Utilizing AI podcast. What do you think? bit.ly/4uJj4Va #AgenticAI #AI
youtu.be
The Future of Agentic AI: Opportunities, Risks, and Society | Utilizing AI Episode 22
YouTube video by Utilizing AI Podcast - The Futurum Group
100
MLCommons @mlcommons.org · 14/05/2026
MLPerf Training v6.0 has added GPT-OSS 20B. With 21B total parameters (but only 3.6B active per token), this new sparse MoE pretraining benchmark is designed specifically for accessibility—it can run on a single 8-GPU node. bit.ly/4noRr14
000
MLCommons @mlcommons.org · 13/05/2026
AI Risk and Reliability certification shouldn't be a self-assessment. That's the premise behind the AILuminate Global Assurance Program (GAP). GAP gives organizations an independent path to certify that their AI systems meet established safety standards. bit.ly/4kIS18x
012
MLCommons @mlcommons.org · 13/05/2026
MLPerf Endpoints: decoupled client, any endpoint, zero-effort integration. Cloud or bare-metal — evaluated equally. Built for API-first GenAI. bit.ly/3Pjx34u #MLPerf
000
MLCommons @mlcommons.org · 12/05/2026
Great to see Microsoft highlighting the need for global collaboration on AI safety testing—and shouting out the MLCommons community’s ongoing work to expand the AILuminate benchmarks for multilingual and multimodal testing. bit.ly/3RdYFZG
bit.ly
Advancing AI evaluation with the Center for AI Standards (US) and Innovation and the AI Security Institute (UK) - Microsoft On the Issues
Today, Microsoft is announcing new agreements with the Center for AI Standards and Innovation (CAISI) in the US and the AI Security Institute (AISI) in the UK to advance the science of AI testing and ...
000
MLCommons @mlcommons.org · 12/05/2026
The New Wave of AI in Healthcare 2026 symposium kicks off today in NYC! 5/13 at 10:50 AM, MLCommons' Andrew Gruen, PhD will be taking the stage. If you're attending, don't miss this conversation on trust, accountability, and AI validation in medicine. lnkd.in/efz2t-Ja
000
MLCommons @mlcommons.org · 12/05/2026
AI software optimization is now moving faster than hardware cycles. To capture these rapid gains, MLPerf is shifting to a rolling submission cadence. David Kanter explains why this speed matters for enterprise buyers via Nutanix: bit.ly/3R24FVt #MLPerf #AI
nutanix.com
Measuring AI Performance Shifts to APIs | The Forecast
MLCommons cofounder David Kanter explains how the MLPerf benchmark has been overhauled to measure AI performance via API endpoints, reflecting the shift toward rented and hybrid AI infrastructure.
000
MLCommons @mlcommons.org · 11/05/2026
Submissions for MLPerf Training v6.0 are open! This round brings updates, including the introduction of large-scale MoE pretraining architectures. Whether benchmarking on a single 8-GPU node or a massive cluster, we want your results in this round. bit.ly/4uG3vNS
000
MLCommons @mlcommons.org · 11/05/2026
We're thrilled to welcome Flower AI to MLCommons to help shape standards for federated AI at scale. First up: MedPerf is integrating with Flower, enabling researchers to run federated clinical AI studies without moving sensitive patient data. bit.ly/4nt1x0T
011
MLCommons @mlcommons.org · 08/05/2026
Measuring today’s production workloads is getting harder. The Inference working group stepped up by adding GPT-OSS 120B, DeepSeek-R1, and our first text-to-video generation benchmark. mlcommons.org/2026/04/mlperf-infere…
000
MLCommons @mlcommons.org · 07/05/2026
MoE benchmarking doesn't have to require a supercomputer. MLPerf Training v6.0 introduces GPT-OSS 20B: a sparse Mixture-of-Experts pretraining benchmark that can run on a single 8-GPU node. See how the task force engineered away statistical variance (CV < 5%): bit.ly/3QLwvVU #MoE #AI
000
MLCommons @mlcommons.org · 05/05/2026
Mixture-of-Experts (MoE) is coming to MLPerf Training v6.0. The new DeepSeek-V3 large-scale pretraining benchmark captures critical innovations like MLA, fine-grained expert segmentation, and MTP at production scale (671B parameters). Technical details: bit.ly/49bRabO
000
MLCommons @mlcommons.org · 30/04/2026
Security theater vs. rigorous AI benchmarking - the difference is methodology. AILuminate Jailbreak v0.7: a mechanism-first taxonomy for single-turn jailbreak attacks. Defensible. Reproducible. Auditable. mlcommons.org/2026/02/jailbreak-0-7 #AILuminate #AISecurity
000
MLCommons @mlcommons.org · 29/04/2026
The New Wave of AI in Healthcare 2026 - May 12-13 in NYC. MLCommons' Andrew Gruen, PhD, is speaking on May 13. Register: lnkd.in/efz2t-Ja #AIinHealthcare
000
MLCommons @mlcommons.org · 28/04/2026
MLPerf Endpoints uses step functions, not trend lines. Interpolating between measured points can hide real failures: memory overflows, P99 spikes. Only verified operating points. No paper performance. bit.ly/3Pjx34u #MLPerf #AIBenchmarking
000
MLCommons @mlcommons.org · 27/04/2026
What happens when AI doesn't just assist - but acts? MLCommons' Dave Graham joined @UtilizingAI to talk about the future of agentic AI: opportunities, risks, and societal impact. Worth a listen. youtu.be/P37u1YdQp4k #AgenticAI #AISafety #MLCommons
youtu.be
The Future of Agentic AI: Opportunities, Risks, and Society | Utilizing AI Episode 22
YouTube video by Utilizing AI Podcast - The Futurum Group
000
MLCommons @mlcommons.org · 23/04/2026
MLCommons is at the ISO Plenary in Singapore this week. As AI safety becomes a global policy priority, aligning international standards matters more than ever. Stay tuned for our recap. #MLCommons #AILuminate #AIPolicy #ISO
000
MLCommons @mlcommons.org · 23/04/2026
The median AI benchmark longevity score is 5/100. AILuminate scored 75 - but that degrades over time, too. So we built the Continuous Prompt Stewardship System to keep it that way. New from the MLCommons AIRR team: bit.ly/3On4jrz
021
MLCommons @mlcommons.org · 22/04/2026
Does AI reliability work? Does it follow the right rules, under every circumstance - including when someone is actively trying to break it? MLCommons AIRR Working Group introduces the AI Reliability Map. mlcommons.org/2026/04/airr-map #AISafety #AIReliability #MLCommons
011
MLCommons @mlcommons.org · 22/04/2026
Early MLPerf Endpoints results include DeepSeek-R1, GPT OSS 120B, Llama 3.1 8B, QWEN 3 Coder 480B — across nearly a dozen systems. More models added as rolling submissions open in Q2 2026. Endpoints.MLCommons.org #MLPerf #GenerativeAI
endpoints.mlcommons.org
endpoints.mlcommons.org
011
MLCommons @mlcommons.org · 20/04/2026
Excited to share that Andrew Gruen, PhD will be speaking at The New Wave of AI in Healthcare 2026, May 12-13. Presented by the Academy and the Windreich Department of Artificial Intelligence and Human Health at Icahn School of Medicine at Mount Sinai Register: lnkd.in/efz2t-Ja #AI #Healthcare
lnkd.in
LinkedIn
This link will take you to a page that’s not on LinkedIn
230