Sign in

Epoch AI

@epochai.bsky.social
1.7K followers 0 following 1.9K posts

We are a research institute investigating the trajectory of AI for the benefit of society. epoch.ai

PostsRepliesMedia
Epoch AI @epochai.bsky.social · 23/09/2026
Can AI tell if you've built your IKEA furniture wrong? Our new benchmark, the Furniture Assembly Benchmark (FAB), gives models the manual and a photo of a half-completed piece of furniture and asks them to spot the mistake. The top score has gone from 28% to 80% in just 10 months.
4445
Epoch AI @epochai.bsky.social · 22/09/2026
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
112532
Epoch AI @epochai.bsky.social · 18/09/2026
The share of math preprints on arXiv that acknowledge AI has risen rapidly, from 4% in April to 25% in August. This increase holds even when filtering to papers with at least one author who published regularly before 2023.
2144
Epoch AI @epochai.bsky.social · 18/09/2026
Another problem from FrontierMath: Open Problems has been solved! The solution was elicited by Becker, Greger, and Peters in an interactive session with GPT-6 Astra. Peters originally suggested the problem for the benchmark. He had this to say.
Quote card from Dominik Peters (CNRS Research Scientist, Université Paris Dauphine - PSL) on GPT-6 Astra designing a new core-stable multi-winner voting rule based on maximizing harmonic entropy.
1569
Epoch AI @epochai.bsky.social · 17/09/2026
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
5636
Epoch AI @epochai.bsky.social · 17/09/2026
Trade data is consistent with more than $3B of chips smuggled into China via Malaysia. Between April 2024 and June 2025, China recorded $3.8 billion in server imports from Malaysia, averaging $106,000 each. Malaysia declared the same shipments at $17,000 each.
272
Epoch AI @epochai.bsky.social · 16/09/2026
GPT-6 Astra leads the Epoch Capabilities Index (ECI), ahead of competing models like Claude Fable 5.1, and its Math-ECI also sets a new record. However, on software engineering benchmarks, Fable 5.1 remains state-of-the-art.
2113
Epoch AI @epochai.bsky.social · 16/09/2026
We're scaling our AI Data Centers research to provide a more comprehensive map of the world's AI infrastructure. Our data now captures ~44% of global AI compute, with coverage growing fast.
180
Epoch AI @epochai.bsky.social · 14/09/2026
More US adults are using AI nearly every day, according to our polling results. The share of US adults who reported using AI at least 6 days in the previous week more than doubled from March to August 2026, rising from 8% to 19%.
2111
Epoch AI @epochai.bsky.social · 10/09/2026
Every FrontierMath Tier 4 problem has now been solved by AI, with GPT-6 Astra solving the last problem standing. Mathematicians often commented that AI found unintended shortcuts when solving their Tier 4 problems. Not so for this last one, which was created by Jay Pantone.
2196
Epoch AI @epochai.bsky.social · 09/09/2026
OpenAI has grown its compute nearly 20-fold since 2023, the sharpest example of an industry-wide surge. Our new AI Chip Users explorer tracks the growth in compute use across five of the world's top frontier AI developers: OpenAI, Google DeepMind, Anthropic, Meta Superintelligence Labs, SpaceXAI.
Bar graph shows AI compute growth from 2023 to 2025 for OpenAI, Google DeepMind, Anthropic, Meta, and SpaceXAI in H100 equivalents.
1184
Epoch AI @epochai.bsky.social · 08/09/2026
We studied time to first token (TTFT) and how it scales with increasing context length for GPT and Claude models. We found a significant difference, with GPT showing a noticeable quadratic component, while Claude models remain closer to linear.
2204
Epoch AI @epochai.bsky.social · 04/09/2026
AI data centers have been getting larger at an exponential rate. The IT power capacity of the largest data centers has doubled every 10 months since mid-2024. But will that trend continue?
3123
Epoch AI @epochai.bsky.social · 04/09/2026
Will Huawei catch up to Nvidia by 2030? Almost certainly not. 🧵
181
Epoch AI @epochai.bsky.social · 01/09/2026
The Epoch Capabilities Index (ECI) frontier has advanced by 14 points/year since reasoning models were introduced. That compares to six points/year in the non-reasoning era.
The graph shows the Epoch Capabilities Index (ECI) trends with reasoning models advancing at 14 points/year compared to 6 points for non-reasoning.
174
Epoch AI @epochai.bsky.social · 28/08/2026
Humans improve with experience. AI? Less so. We ran a human baseline on Earthborne Rangers, the complex, long-horizon board game behind our benchmark EBR-bench. Humans started low, but the best mastered the game after 5 playthroughs. No AI we have evaluated ever got there.
1292
Epoch AI @epochai.bsky.social · 27/08/2026
In 2025, Anthropic and OpenAI were already among the fastest-growing companies of their size in history. And yet, their revenue growth has accelerated in 2026. Josh You and Lynette Bye discuss what this means for the future of AI in the latest edition of our newsletter.
1292
Epoch AI @epochai.bsky.social · 26/08/2026
Nvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s), according to new benchmarks of Gemma 4 31B. LPX does this by using a small amount (128 GB) of ultrafast SRAM in place of HBM.
A bar graph comparing output speeds for Nvidia's Groq 3 LPX and other models, showcasing Groq's significantly higher performance.
181
Epoch AI @epochai.bsky.social · 24/08/2026
We discovered that US GDP statistics miss most of the value Nvidia adds to the US economy. As a result, GDP growth has been understated by ~0.3 percentage points over the last year.
2102
Epoch AI @epochai.bsky.social · 14/08/2026
Who pays for the AI that US workers use on the job? We found that most AI use at work happens on free plans, except in science and tech, where employers are more likely to pay for premium access.
1102
Epoch AI @epochai.bsky.social · 14/08/2026
What are the big questions about AI capabilities that govern the impact AI will have on the world? In this Gradient Update, Greg Burnham lays out his top questions, and how he thinks Epoch's work on benchmarking can help answer them.
270
Epoch AI @epochai.bsky.social · 13/08/2026
Continuing to scale AI compute at recent rates will require exponentially more capital. Is financing a bottleneck? Probably not yet. That’s the takeaway from a new Gradient Update by Campbell Hutcheson, which analyzes nearly $50B of debt associated with Anthropic’s buildout.
2133
Epoch AI @epochai.bsky.social · 13/08/2026
How much more AI compute does a dollar buy each year? About 49%, or a doubling every 21 months, based on the chips actually bought each quarter from 2023 through 2025.
1305
Epoch AI @epochai.bsky.social · 06/08/2026
We’ve launched a new “game puzzles” benchmark, similar to our Chess Puzzles benchmark but using puzzles from a different, undisclosed game. This helps us test AI on reasoning-heavy tasks where they probably weren’t post-trained. The current record-holder is Opus 5, scoring 59%.
1192
Epoch AI @epochai.bsky.social · 06/08/2026
New Epoch AI/Ipsos survey: 1 in 5 US workers say AI now handles at least one task previously delegated to humans. More findings on how AI is changing everyday job tasks 🧵
161
Epoch AI @epochai.bsky.social · 05/08/2026
DeepSeek-V4-Flash-0731 debuts with an ECI of 153, comparable to GLM 5.2 and roughly midway between Opus 4.5 and Opus 4.6. It's the second strongest open-weights model available today, behind only Kimi K3.
1243
Epoch AI @epochai.bsky.social · 03/08/2026
Serious cyber vulnerability disclosures keep climbing. In July, 21 major tech organizations published ~2,500 high- and critical-severity CVEs — about 5× the monthly record before Anthropic revealed Claude Mythos Preview could autonomously find software vulnerabilities.
2256
Epoch AI @epochai.bsky.social · 03/08/2026
We've updated the MirrorCode leaderboard with results for Claude Fable 5 and GPT-5.6 Sol. Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.
182
Epoch AI @epochai.bsky.social · 03/08/2026
We're hiring a Head of People to help us double our headcount this year! You'll own recruiting, HR, and events and lead a growing team as we scale from ~30 to ~70 people.
141
Epoch AI @epochai.bsky.social · 31/07/2026
We’ve launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far, and solving all of them would be an incredible mathematical feat. Thread with more.
3214
Epoch AI @epochai.bsky.social · 29/07/2026
Parallelization constraints could delay or prevent a technological singularity, even after R&D is automated. Whether, and how quickly, an explosion proceeds will depend on the development of “parallelization technology”.
171
Epoch AI @epochai.bsky.social · 28/07/2026
GPT-5.6 Sol has been climbing Slay the Spire's Ascension ladder on our Twitch channel for a week, no human in the loop. This Thursday, Claude Opus 5 takes over the climb — live, with commentary. Thursday, July 30 · 12:30 PT twitch.tv/epochaiplays
1194
Epoch AI @epochai.bsky.social · 28/07/2026
AI has found a presentation for the absolute Galois group of the field of 2-adic numbers. This is the second problem to be solved in FrontierMath: Open Problems, our benchmark of significant unsolved problems from research mathematics.
1135
Epoch AI @epochai.bsky.social · 25/07/2026
Claude Opus 5 gets an ECI of 159, slightly below Fable 5's value of 161 (while 5.6 sol holds the record with 162). However looking only at software engineering, we find it matches Fable 5s SWE-ECI of 161.
1132
Epoch AI @epochai.bsky.social · 23/07/2026
Join the EpochAIPlays launch stream later today, with live commentary by @AlephNuul! We will be benchmarking GPT 5.6 Sol against Slay the Spire 1.
160
Epoch AI @epochai.bsky.social · 22/07/2026
How surprising should we find it that an internal OpenAI model was able to escape its restrictions and autonomously hack Hugging Face, all just to cheat on a cybersecurity benchmark? We have pulled together the public evidence on AI cyber capabilities in this thread:
3232
Epoch AI @epochai.bsky.social · 21/07/2026
Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respectively, and just ahead of GPT 5.6 Luna.
3245
Epoch AI @epochai.bsky.social · 17/07/2026
We stress-tested some AI detectors and found that they rarely flag human text as AI-generated. But asking LLMs to mimic a specific author causes detectors to misclassify text as human-generated ~13% of the time. For scientific writing, false negatives rose to ~26%.
171
Epoch AI @epochai.bsky.social · 15/07/2026
We recently fixed a bug in our ECI confidence interval code. The bug only impacted the confidence intervals, not the central ECI scores or model rankings, and fixing the bug has made our confidence intervals narrower for recent models. Details below.
130
Epoch AI @epochai.bsky.social · 10/07/2026
How much does AI speed up the engineers building it? We analyzed contributions to OpenAI's public Codex repository to gather evidence. In Q2 2026, 8% of contributor-days involved more than 24 hours worth of human engineering work, as estimated by LLM judges.
1140
Epoch AI @epochai.bsky.social · 08/07/2026
The Epoch Capabilities Index now has slightly tighter confidence intervals, thanks to an update to the methodology we use to compute them.
180
Epoch AI @epochai.bsky.social · 08/07/2026
Will we get Dyson Spheres a few years after automating AI R&D? Most AI futurism debates answer this by looking at AI capabilities, but miss half the picture: how intrinsically hard it is to build futuristic tech in the first place. New essay by Jean-Stanislas Denain and Anson Ho. 🧵
191
Epoch AI @epochai.bsky.social · 08/07/2026
Z.ai's GLM-5.2 scores an estimated 152 on the Epoch Capabilities Index, the highest of any open-weight model we've evaluated. It remains behind models like Gemini 3 Pro, released over 7 months ago.
1163
Epoch AI @epochai.bsky.social · 03/07/2026
We’re hiring a Benchmark Engineer to join our Evaluations team! You’ll help expand our AI Benchmarking Hub - running and maintaining benchmarks, integrating with AI providers, and designing brand-new benchmarks from scratch.
130
Epoch AI @epochai.bsky.social · 03/07/2026
We're looking for new Researchers to join our Evaluations team! Help us curate real-world task suites, design rubrics, and evaluate how well frontier models handle open-ended tasks.
140
Epoch AI @epochai.bsky.social · 02/07/2026
AI appears to be finding software vulnerabilities at scale. In June 2026, 21 notable organizations disclosed ~1,500 high- and critical-severity CVEs, over 3.5× the previous monthly record set before Claude Mythos Preview's release.
3355
Epoch AI @epochai.bsky.social · 02/07/2026
OpenAI’s GPT-4 led the Epoch Capabilities Index for 352 days after its March 2023 release, far longer than any model since. The second-longest lead belongs to OpenAI’s o1 at 98 days.
1120
Epoch AI @epochai.bsky.social · 02/07/2026
Introducing EBR-bench, our new benchmark to measure on-the-fly learning. AI repeatedly plays a challenging board game called Earthborne Rangers and tries to learn from its mistakes. So far: no signs of improvement.
2303
Epoch AI @epochai.bsky.social · 01/07/2026
We recently began tracking 13 new evals on our benchmarking hub. 7 of these have been incorporated into the Epoch Capabilities Index (ECI).
120
Epoch AI @epochai.bsky.social · 29/06/2026
We're looking for a new Talent Scout to join our team! You'll help us find great researchers, engineers, and support staff and introduce them to Epoch through outreach and events.
130