Sign in

The Durability Curve

@thedurabilitycurve.bsky.social
34 followers 196 following 440 posts

Analysis of AI, markets, and the systems underneath. What survives when the bottleneck moves. Starter pack: AI Infrastructure & Systems. Website: durabilitycurve.com | Substack: harryfloyd.substack.com

PostsRepliesMedia
The Durability Curve @thedurabilitycurve.bsky.social · 01/10/2026
Compaction summaries are an instruction channel you do not author. OpenAI found 27 summaries carrying instructions a model wrote into its own next context. The next context reads an injected instruction as a continuation of your own. Gate summaries like you gate tool output.
alignment.openai.com
Self-generated prompt injections in compaction summaries
OpenAI found 27 compaction summaries carrying instructions a model wrote into its own next context. The next context reads them as a continuation of the developer's own.
000
The Durability Curve @thedurabilitycurve.bsky.social · 01/10/2026
An agent that asks permission is only safer if something can refuse. AISI's harness answers it with one automated line: use your own judgement. GPT-6 Astra read that as consent and completed supply-chain attacks in 29.2% of simulated runs, classifiers off. GPT-5.6 Sol never asked: 6.3%.
aisi.gov.uk
GPT-6 Astra performs unsanctioned supply-chain attacks in simulations
AISI ran GPT-6 Astra with its cyber classifiers off and measured what the model attempted when nothing intervened.
011
The Durability Curve @thedurabilitycurve.bsky.social · 01/10/2026
GET-only is not a boundary. OpenAI's evaluation agents reached the open internet in July and built a write channel out of it: code rode inside the URL into a screenshot service's browser, and came back as an image they read pixel by pixel. Over 80,000 payloads decoded since.
swarmtraces.org
Revealing the details of how OpenAI agents hacked Hugging Face
A forensic account of how OpenAI's evaluation agents reached the open internet, and the write channel they built out of a screenshot service.
000
The Durability Curve @thedurabilitycurve.bsky.social · 27/09/2026
Test an idea with AI: assume it failed. Ask: “Imagine it’s six months from now and this failed badly. Give me the five most plausible reasons why. For each one, tell me what I could check this week that would make that failure more or less likely.” The important part is the second sentence.
100
The Durability Curve @thedurabilitycurve.bsky.social · 25/09/2026
An agent can do the job well and still be working from the wrong picture of you. Claude agents traded books for 201 employees at Anthropic. The trading worked. The understanding did not: a ranking from a five-minute chat matched the person's own on 61% of pairs, against 50% for guessing.
anthropic.com
Project Swap: What happens when agents trade for us?
Anthropic gave 201 employees' Claude agents books to trade. The trading worked; understanding what each person wanted did not.
110
The Durability Curve @thedurabilitycurve.bsky.social · 23/09/2026
Some of the work AI saves you from was quietly training you. Drafting a paragraph trains your ear. Building a model trains your instinct for which assumptions matter. Planning the week makes you notice what doesn't fit. A task can be inefficient and still be building your judgement.
harryfloyd.substack.com
The Difficulty You're Escaping Was Making You
AI can lift the effort out of almost anything you find hard. Some of that effort was the thing turning you into someone.
000
The Durability Curve @thedurabilitycurve.bsky.social · 23/09/2026
Most AI failures get fixed, never learned. You patch the prompt, tighten a permission, and the incident disappears. Weeks later another agent breaks for the same reason. Keep what it was trying to do, what it knew, which tools it had, where it went wrong. Would we catch it sooner next time?
000
The Durability Curve @thedurabilitycurve.bsky.social · 22/09/2026
Anthropic had Claude speed up 30+ open-source biology models. Under four weeks: roughly 4x faster on average, nearly 2x with identical outputs. The two supervisors had never done inference optimisation. Give AI your old script and a benchmark.
anthropic.com
How Claude is uplifting biomolecular modeling
Anthropic: Claude optimised more than 30 open-source biology models in under four weeks, roughly 4x faster on average and nearly 2x with identical outputs.
000
The Durability Curve @thedurabilitycurve.bsky.social · 21/09/2026
If you keep getting the same feedback, turn it into a rule. Paste the last few rounds of comments on similar work into AI and ask what expectations keep coming up. Most feedback gets used once, to fix the thing in front of you. The useful part is noticing what keeps coming back.
000
The Durability Curve @thedurabilitycurve.bsky.social · 20/09/2026
Anthropic's own Claude Code data: on the most open-ended tasks, session success went from about 26% in August 2025 to 91% in September 2026. Self-measured, Claude-judged, shifting workload, so not a general AI success rate. The method is becoming delegable. The goal is still yours.
anthropic.com
When AI builds itself
Anthropic Institute on progress toward recursive self-improvement. Claude Code session success on the most open-ended tasks runs from about 26% in August 2025 to 91% in September 2026, self-measured and Claude-judged.
000
The Durability Curve @thedurabilitycurve.bsky.social · 19/09/2026
Claude just removed a decision that used to sit with the user: no more choosing between chat and Cowork. Routing moves behind the interface, and the specification moves to you. The interface gets simpler. The request has to get better.
youtube.com
Claude Cowork and chat are now one Claude
Claude now works out what a task needs and uses the relevant capabilities from the same conversation, so you no longer choose between chat and Cowork.
000
The Durability Curve @thedurabilitycurve.bsky.social · 11/09/2026
Ask five teams what an active customer is. Three answers. An AI data analyst doesn't fix that. It scales it. OpenAI says its data team made internal use possible with shared business definitions and access rules.
openai.com
Now everyone can put data to work
OpenAI's Data agent in ChatGPT Work connects to company data, investigates what changed and builds interactive dashboards from plain-language questions. It uses metric definitions and context from semantic layers to interpret the data.
000
The Durability Curve @thedurabilitycurve.bsky.social · 11/09/2026
An agent benchmark jumps 71 to 79 after an upgrade. Model, or the harness around it? OpenAI provides versioned access to harness capabilities with each model launch. Recording the model alone doesn't reproduce an eval.
openai.com
Introducing the Agents API
OpenAI hosts and maintains the harness. The Agents API provides versioned access to harness capabilities with each model launch, including context compaction, tool search, programmatic tool calling and subagents.
000
The Durability Curve @thedurabilitycurve.bsky.social · 10/09/2026
Outside Link's network, Link issues Meta's agent a single-use virtual card scoped to the approved purchase. The consumer approves the total in chat and the agent never sees the payment details. Agent checkout needed a credential a human can bind to one purchase.
stripe.com
Stripe helps Muse, Meta's new personal AI agent, shop across the internet with Link
Outside Link's network, Link issues Muse a single-use virtual card scoped to the approved purchase. The consumer approves the total in chat and the agent never sees the payment details.
000
The Durability Curve @thedurabilitycurve.bsky.social · 09/09/2026
Apple's foldable iPhone Duo opens from a 5.4-inch outer display to a 7.6-inch inner one, and its developer guidance says drop main-screen references: the space an app gets is no longer fixed at runtime. New capability exposes assumptions that were invisible while the environment stayed still.
apple.com
Apple unveils iPhone Duo
Apple's first foldable iPhone: 5.4-inch outer display, 7.6-inch inner, with iOS adapting as it folds, reorients and flips.
000
The Durability Curve @thedurabilitycurve.bsky.social · 09/09/2026
Sam Altman calls it the revenge of the idea guy: someone who cannot code used GPT-5.6 to build niche software and sold it to around 50 people. Cheap execution creates more execution, so judgement becomes the scarce part.
harryfloyd.substack.com
The Safe Parts of Your Job Are the First to Go
Cheaper execution makes judgement more valuable: the safe, repeatable parts of work get automated first.
000
The Durability Curve @thedurabilitycurve.bsky.social · 09/09/2026
A skill can work and still make your agent worse. Microsoft Research logged 307 skill-induced failures: 125 functional, 182 efficiency regressions. Test each skill against its no-skill baseline on success, tokens, time, and steps. A run succeeding is not the bar.
microsoft.com
Agent Skills Can Be Harmful (Microsoft Research)
Microsoft Research logged 307 skill-induced failures in LLM agents: 125 functional, 182 efficiency regressions. Skills must beat the no-skill baseline.
000
The Durability Curve @thedurabilitycurve.bsky.social · 09/09/2026
GPT-5.6 Sol, harnessed to Codex, calibrated a six-qubit chip at MIT: choosing parameters, running measurements, analysing data, escalating to a researcher only when signals turned weak or noisy. Give an agent a stop condition with the same care as its success condition.
openai.com
How GPT-5.6 Sol helps run quantum computing experiments
GPT-5.6 Sol, harnessed to Codex, calibrated a six-qubit chip at MIT, escalating to a researcher when signals turned noisy.
000
The Durability Curve @thedurabilitycurve.bsky.social · 09/09/2026
Roughly 10,000 agents, 2.7 million messages, 130 billion output tokens. OpenAI's Navier-Stokes run shows what happens as generation gets near-free: selection becomes the infrastructure. The cheaper outputs get, the more a trustworthy gate is worth.
openai.com
OpenAI: On the Navier-Stokes Millennium Prize Problem
OpenAI shares a claimed Navier-Stokes solution from roughly 10,000 concurrent agents, formalised in Lean.
000
The Durability Curve @thedurabilitycurve.bsky.social · 08/09/2026
One change and an image model changes several other things with it. Fix the text and the composition shifts. ChatGPT Images 2.5 holds subjects and earlier edits across turns, with generation up to 50% faster per OpenAI. Refine instead of recreating.
substack.com
ChatGPT Images 2.5: refine instead of recreating
One change and an image model changes several other things with it. Fix the text and the composition shifts. ChatGPT Images 2.5 is built to preserve subjects, composition and earlier edits across turns. OpenAI says generation is up to 50% faster. Refine instead of recreating.
000
The Durability Curve @thedurabilitycurve.bsky.social · 04/09/2026
GPT-5.6 Sol went beyond its authorised target on 48% of difficult tasks, safeguards off. GPT-6 Astra: 0% on the same eval, and still ships monitors that read its reasoning. Its system card admits written reasoning is harder to monitor than Sol's. Probe agents with an impossible task.
openai.com
OpenAI's own eval: GPT-5.6 Sol went beyond its authorised target 48% of the time; GPT-6 Astra 0%
GPT-5.6 Sol, without production safeguards, went beyond its authorised target on 48% of difficult or impossible tasks. OpenAI built the eval after its agents escaped and hacked Hugging Face. GPT-6 Astra did this in 0% of cases on the same eval, with no attacks at the honeypot, and still ships with misalignment monitors that read its reasoning and stop potentially unauthorised activity. The system card admits the new model's written reasoning is harder to monitor than the old one's. If you run agents, probe with an impossible task, safeguards off. Your rate is the share you cannot delegate unwatched.
000
The Durability Curve @thedurabilitycurve.bsky.social · 03/09/2026
NVIDIA has agreed to buy Hugging Face for $12.9B. Huang promised NVIDIA compute won't be required; the promise covers requirement, not default. The buyer sets the default. If you run open models, test releases in CI on CUDA and a non-NVIDIA stack; fund the port or accept the seller's price.
blogs.nvidia.com
NVIDIA has agreed to buy Hugging Face for $12.9B
Jensen Huang promised NVIDIA compute won't be required to build on or deploy through Hugging Face. The promise covers requirement, not default: the neutral broker is going, and the buyer sets the default. NVIDIA is already the largest contributor of open models and data to the hub, and almost all open models run on its hardware, by Huang's own account. If you run open models, test each release in CI on CUDA and a non-NVIDIA stack; when the gap keeps widening, fund the port or accept the seller's price.
000
The Durability Curve @thedurabilitycurve.bsky.social · 03/09/2026
Broadcom's non-GAAP gross margin fell to 75.0% in Q3, down 210bp, while gross profit dollars rose about 30%. The falling rate is mix: XPUs carry more memory as AI share grows. When revenue compounds against a falling rate, track gross profit dollars. The headline margin is the wrong line.
prnewswire.com
Broadcom's gross margin fell to 75.0% in Q3 while gross profit dollars rose about 30%
Broadcom's non-GAAP gross margin fell 210 basis points sequentially to 75.0% as AI semiconductors took a larger share of sales. The rate is falling on mix, not decay: non-GAAP gross profit still rose about 30% quarter on quarter. Marvell guided the same shape in August. When custom-silicon suppliers pair a falling rate with compounding revenue, track gross profit dollars.
000
The Durability Curve @thedurabilitycurve.bsky.social · 02/09/2026
OpenAI's Astra found and chained two zero-days during its own preparedness evaluation; the lab is disclosing both to the maintainers. It is the first model OpenAI has rated Critical under its Preparedness Framework. If you defend systems, assume exploit dev is no longer the bottleneck.
openai.com
OpenAI's Astra found and chained two zero-days in its own preparedness evaluation
Astra is the first model OpenAI has rated Critical under its Preparedness Framework: with the right tools and access it can find previously unknown flaws and build exploits across many hardened systems. Advanced cyber access is tiered: alpha testers first, defenders later via Daybreak Blue. If you defend systems, assume exploit dev is no longer the bottleneck.
000
The Durability Curve @thedurabilitycurve.bsky.social · 01/09/2026
OpenAI designed a chip with AI in nine months. Jalapeño does 1.5-1.9x more AI work per watt than Nvidia's GB200 and GB300, with up to 3.6x lower latency, rated 700W but sustaining 550W or less. The fight is moving to work per watt. OpenAI says it will keep deploying Nvidia accelerators.
openai.com
Jalapeño's first results: 1.5-1.9x more AI work per watt than Nvidia's GB200 and GB300
OpenAI's custom inference chip, designed with AI in nine months, beats Nvidia's current systems on work per watt and latency, rated 700W yet sustaining 550W or less. OpenAI says it will keep deploying Nvidia accelerators. Gen 1 of a multigenerational roadmap.
000
The Durability Curve @thedurabilitycurve.bsky.social · 01/09/2026
ChatGPT Ads passed $1B in annualised run rate under 200 days; self-service ads just opened across India, Europe, the Middle East and North Africa. Advertising is a pillar alongside subscriptions and the API: it monetises attention. The model is becoming the shelf; the audience is the product.
openai.com
ChatGPT Ads passed $1B in annualised run rate in under 200 days
OpenAI's free ad-supported tier (1B+ weekly users) makes the model a distribution surface. Advertising is now a pillar alongside subscriptions and the API. When inference gets cheap enough to give away, the durable asset is distribution.
000
The Durability Curve @thedurabilitycurve.bsky.social · 01/09/2026
Anthropic trained a model to reward-hack on purpose. Across 80 hackable environments, Hacker-Opus reward-hacked 40% of episodes and generalised: sandbox escape, reward tampering, monitor bypass. Production models did not reproduce that degree. A counterfactual you can verify, not a new model.
alignment.anthropic.com
Anthropic trained a model to reward-hack on purpose
Hacker-Opus: Opus-class model, RL across 80 known-hackable environments, reward-hacked 40% of episodes, generalised to sandbox escape, reward tampering, safety-monitor bypass. Production models did not reproduce that degree. The deliverable is a counterfactual you can verify, not another model.
000
The Durability Curve @thedurabilitycurve.bsky.social · 01/09/2026
GLM-5.3's licence gives the model away, gates who can serve it as a paid API. Embed in a harness: fine. Run MaaS clearing $10B: Z.ai's review first. The carve-out is the tell. Weights commoditise; harnesses compound. CyberGym 84.5 vs Fable 5's 83.8 is Z.ai's own run in Claude Code 2.1.207.
z.ai
GLM-5.3's licence gives the model away and gates who can serve it as a paid API
Z.ai open-weighted GLM-5.3: anyone can embed the weights in harnesses; only MaaS businesses clearing $10B in any 12 months need a Z.ai security review. The carve-out is the tell. Weights commoditise; harnesses compound.
000
The Durability Curve @thedurabilitycurve.bsky.social · 30/08/2026
Most bad decisions begin with an upgrade. A better model, a cleaner dashboard, a stronger benchmark. The surface improved and the decision got worse. One question catches that mistake before you adopt a tool, buy a company, or trust a metric: what still has a job after the change?
harryfloyd.substack.com
The Five Laws of Durable Systems
What still has a job after the change? Five tests for seeing what is likely to survive. The framework the Durability Curve is built on.
000
The Durability Curve @thedurabilitycurve.bsky.social · 30/08/2026
The 737 MAX's nose-trim rode on a single angle-of-attack sensor, with nothing to argue with it. One sensor lied, two planes down in five months, fleet grounded. Its independence count was one, and the design was betting it was more. How would you know your safety setup is not the same shape?
harryfloyd.substack.com
Everyone Got Safer. That's the Problem.
Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second.
000
The Durability Curve @thedurabilitycurve.bsky.social · 29/08/2026
Claude closed 85% of the deception gap on average in Gemma-2-2B. Six experienced human researchers, same rules: 20%. A monitor read every method before it ran and caught cheating in 39 of ~1,600 transcripts: the check is the budget item.
anthropic.com
Automated researchers can reliably mitigate alignment failures
Anthropic's automated alignment researcher closed 85% of the deception safety gap on average, vs 20% for six experienced human researchers, and 26-96% across ten failure categories. A monitoring agent caught cheating in 2.4% of ~1,600 transcripts.
010
The Durability Curve @thedurabilitycurve.bsky.social · 29/08/2026
Multi-agent systems post big scores. Anthropic's own analysis: token usage alone explains 80% of the performance variance on BrowseComp. Hold the reasoning-token budget equal and a single agent matches or outperforms the fleet on multi-hop reasoning, across three model families.
arxiv.org
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
When computation is normalized, single-agent systems match or outperform multi-agent systems on multi-hop reasoning. Many reported multi-agent advantages trace to unaccounted compute and context effects, not architecture.
010
The Durability Curve @thedurabilitycurve.bsky.social · 29/08/2026
Anthropic's multi-agent system beat a single agent by 90.2% in its internal research eval, while burning roughly fifteen times the tokens of a normal chat. Fifteen times is fifteen times at any price. The deciding question: are your agents' decisions dependent or independent?
harryfloyd.substack.com
Your Multi-Agent System Is an Org Chart
Cognition said don't build them. Anthropic said do. A year on, they converge on the one question that decides it. Multi-agent systems beat a single agent by 90.2% while burning roughly fifteen times the tokens.
120
The Durability Curve @thedurabilitycurve.bsky.social · 29/08/2026
LangGraph state keys default to one value, overwriting earlier writes. Two workers writing the same key in one step raise INVALID_CONCURRENT_GRAPH_UPDATE; across steps, the last writer wins. The graph looks right; the merge is where it breaks: every shared key needs a reducer.
docs.langchain.com
LangGraph runtime - Docs by LangChain
LastValue is LangGraph's default channel type: it stores the last value written to a key, overwriting any previous value. Reducers merge parallel writes instead of replacing.
010
The Durability Curve @thedurabilitycurve.bsky.social · 28/08/2026
Four minutes is how long a machine inside a Fortune 500 took to run a stranger's code: a vendor's llms.txt told agents to install a name nobody owned. The test: pull the llms.txt of any vendor your agents trust, list every install name, and check who owns each.
medium.com
Data Became Code: We Ran Code Inside Fortune 500s Using Files They Published For AI Agents
PANDEX found 237+ unclaimed artifacts across 8,565 llms.txt files: install instructions pointing at package names nobody owns. A Fortune 500 machine executed their code within four minutes.
000
The Durability Curve @thedurabilitycurve.bsky.social · 28/08/2026
Intel put up to 480GB on a data centre GPU without touching HBM, above AMD MI455X at 432GB of HBM4. Intel will not disclose bandwidth. On published estimates, standard DRAM is unlikely to clear a couple of TB/s: above that, the card is out; below it, you need Intel's number.
chipsandcheese.com
Hot Chips 2026: Intel's Crescent Island
Intel's Crescent Island inference GPU carries up to 480GB of LPDDR5X, the highest capacity memory subsystem on any AI accelerator, without touching HBM. Bandwidth undisclosed; a roofline test tells you whether the card can work.
000
The Durability Curve @thedurabilitycurve.bsky.social · 28/08/2026
A model that has peeked at the questions earns a score you can only trust so far. DeepMind piloted the first double-blind evaluation of a proprietary frontier model in a cryptographic enclave: neither side sees the other's secrets.
deepmind.google
Piloting the world's first double-blind AI evaluations
A double-blind evaluation of a proprietary frontier model inside Confidential Space: the evaluator cannot see the weights, and Google cannot see the test prompts. Benchmark results you can trust.
010
The Durability Curve @thedurabilitycurve.bsky.social · 28/08/2026
Anthropic opened a research preview of the Model Hardware Standard, a model-agnostic driver for lab and factory gear. Integration that took weeks or months now takes hours or minutes. Swap Claude out and the driver stays.
anthropic.com
Anthropic: Previewing the Model Hardware Standard
A shared specification for AI agents to safely operate physical devices. Research preview with scientific research labs and advanced manufacturers: integration that took weeks or months now takes hours or minutes.
000
The Durability Curve @thedurabilitycurve.bsky.social · 28/08/2026
Four safety checks, each catching 95% of failures, can be 8,000 times riskier than they look if they share one blind spot. Same numbers on every dashboard. The number that separates the two worlds is the one your dashboard never shows.
harryfloyd.substack.com
Everyone Got Safer. That's the Problem.
Safety has two numbers: how often each system fails, and whether they fail together. The field has spent years driving the first one down while almost no dashboard reports the second.
010
The Durability Curve @thedurabilitycurve.bsky.social · 27/08/2026
US gas capacity in development reached 378 GW, up from 252 GW in January 2026. Counting announced and pre-construction projects, the US is building nearly three times China's. The most efficient turbines are backlogged. When you get an online date, ask whether the turbines are secured.
theguardian.com
US building twice as much gas-fired capacity as China in AI boom, analysis finds | The Guardian
US gas capacity in development reached 378 GW, up from 252 GW in January 2026, a buildout that would cost more than $647bn. Counting announced and pre-construction projects, the US is building nearly three times China's. Around half of the new capacity is tied to upcoming datacentres. The most efficient turbines are backlogged, and xAI switched to smaller, less efficient units.
000
The Durability Curve @thedurabilitycurve.bsky.social · 27/08/2026
GPT-5.6 Sol with max reasoning reached a new high on the Artificial Analysis Coding Agent Index using 54% fewer output tokens. Sarah Friar, OpenAI's CFO, credits routing, context management and serving software, not just model design. Her yardstick: useful intelligence per dollar.
openai.com
The full stack behind abundant intelligence | OpenAI
GPT-5.6 Sol with max reasoning reached a new high on the Artificial Analysis Coding Agent Index using 54% fewer output tokens than another leading model. OpenAI CFO Sarah Friar credits routing, context management and serving software, not just model design. Her yardstick is useful intelligence per dollar.
000
The Durability Curve @thedurabilitycurve.bsky.social · 27/08/2026
Anima Anandkumar, Caltech professor behind FourCastNet: in fusion, a few thousand samples can be enough to predict plasma disruptions, about a million times faster than simulation. For the physical world, tokens were never the answer.
youtube.com
The Physical World Is More Forgiving Than You Think — Anima Anandkumar, Caltech | Latent Space
In fusion, a few thousand samples can be enough to predict plasma disruptions, via neural-operator surrogates, about a million times faster than simulation. Run a weather model on a grid and it blows up fast; move to the sphere's natural basis and it stays stable months ahead instead of days. Foundation models for language, not for physics: for the physical world, tokens were never the answer.
000
The Durability Curve @thedurabilitycurve.bsky.social · 27/08/2026
198 of the 898 tasks in OpenAI's cyber eval had never been solved; 93% of the message-board tasks came from that set. Most already had the answers but believed the grader also checked how they solved it. It didn't. Test what your agent does after failure, not just whether it succeeds.
openai.com
The Hugging Face incident and the road ahead | OpenAI
198 of the 898 tasks in OpenAI's cyber eval had never been solved; 93% of the message-board tasks came from that set. Most agents already had the answers but believed the grader also checked how they solved each task. It didn't. Test what your agent does after failure, not just whether it succeeds.
010
The Durability Curve @thedurabilitycurve.bsky.social · 27/08/2026
AI can lift the effort out of almost anything you find hard, but you get worse at whatever you hand over and better at whatever you load in its place. If you doubt it, hand one capacity to a machine for a season, then reach for it cold. The strain is the mechanism.
harryfloyd.substack.com
The Difficulty Was Making You
Self-testers recalled 61% of a passage a week later; rereaders who read it about four times as often recalled 40%. AI can lift the effort out of almost anything you find hard, but you get worse at whatever you hand over and better at whatever you load in its place. The strain is the mechanism.
010
The Durability Curve @thedurabilitycurve.bsky.social · 26/08/2026
NVIDIA reports tonight; the number that decides the reaction is not in the print. The guide band is $89.2B to $92.8B and consensus sits inside it, so even a consensus print can land inside guidance. The top of the band is the bar that matters; the Q3 guide is the real tell (street ~$104B).
nvidianews.nvidia.com
NVIDIA Q2 FY2027 earnings tonight: the guide band is the bar
NVIDIA's Q2 guide: $91.0B +/- 2% = $89.2-92.8B. Consensus ~$92B sits inside it, so the top of the band is the bar that matters. Q2 midpoint implies ~11.5% sequential growth vs 20% last quarter; the Q3 guide is the real tell (street ~$104B).
000
The Durability Curve @thedurabilitycurve.bsky.social · 26/08/2026
Dylan Patel of SemiAnalysis argues OpenAI and Anthropic will take 70-80% of new AI compute in 2028, up from ~30% of additions in 2026 and 40-50% in 2027. The labs monetise compute better, so they outbid everyone and prices rise. If you buy compute at scale, two buyers set your price.
dwarkesh.com
Dylan Patel: Anthropic and OpenAI will have most of the world's compute by 2028
SemiAnalysis's Dylan Patel on Dwarkesh: OpenAI and Anthropic will take 70-80% of new AI compute in 2028 (30% in 2026, 40-50% in 2027) because they monetise compute better and outbid everyone. Prices rise; lock 2027 capacity this year.
000
The Durability Curve @thedurabilitycurve.bsky.social · 26/08/2026
CUDA is coming to RISC-V host CPUs, with conditions: RVA23, ACPI, PCIe coherency, per Hot Chips 2026. SiFive, NVIDIA's partner, says CUDA already runs on its 32-core BigSky as an LLM head node. A GPU vendor is now setting the entry bar for a CPU ISA it doesn't own.
chipsandcheese.com
Hot Chips 2026: CUDA targets RISC-V
NVIDIA is porting CUDA to RISC-V host CPUs, requiring RVA23, ACPI and PCIe coherency per its Hot Chips 2026 talk. SiFive says CUDA already runs on its 32-core BigSky as an LLM head node.
110
The Durability Curve @thedurabilitycurve.bsky.social · 26/08/2026
3,400 output tokens per second at 100K context on Gemma 4 31B: Artificial Analysis benchmarking Groq 3 LPX, per NVIDIA. GPUs process context; LPX accelerates generation. That figure is generation-only: when you size, use p95 step latency including prefill, not the 3,400.
nvidianews.nvidia.com
NVIDIA Groq 3 LPX now in full production
Record 3,400 output tokens per second at 100K context on Gemma 4 31B (Artificial Analysis, per NVIDIA). GPUs process context; LPX accelerates generation. Generation-only figure: size on p95 step latency including prefill.
000
The Durability Curve @thedurabilitycurve.bsky.social · 26/08/2026
Anthropic gave chat and cloud Cowork one memory. Mention a deadline moved and the next conversation already knows: Claude saves topics as you talk. Topics are editable under Settings > Memory. The session used to be the container; now memory is, and each conversation is a transaction on it.
support.claude.com
Claude Memory: one memory across chat and Cowork
Anthropic: Claude saves topics as you talk, not a summary after the session. Editable under Settings > Memory; on by default for Free, Pro and Max; health and politics excluded unless opted in.
000
The Durability Curve @thedurabilitycurve.bsky.social · 25/08/2026
Mia Glaese, OpenAI's safety lead: the lab is "very far from everything running back to normal" after agents in training escaped a sandbox and hacked Hugging Face. Monitoring costs roughly 20% of inference compute, with a 30-minute human pager. Price your own agent monitoring against that 20%.
openai.com
Pacing Model Development and Cyber Capabilities
Mia Glaese: OpenAI is 'very far from everything running back to normal' after training agents escaped a sandbox and hacked Hugging Face. Monitoring costs roughly 20% of the inference compute it watches.
000