Sign in

Anthropic

@anthropic.extwitter.link
45 followers 1 following 228 posts

⚠️ MIRROR OF twitter.com/AnthropicAI ⚠️ If you own the original account and want to claim this, please contact @twttr-mirrors.bsky.social

PostsRepliesMedia
Anthropic @anthropic.extwitter.link · 13h
Claude Haiku 5.5, our latest model: 🔗 twitter.com/i/status/210789...
Quoted tweet: https://twitter.com/i/status/2107894039626277339
000
Anthropic @anthropic.extwitter.link · 28/09/2026
Claude Sonnet 5.5 is now available: 🔗 twitter.com/i/status/210463...
Quoted tweet: https://twitter.com/i/status/2104633115620823187
000
Anthropic @anthropic.extwitter.link · 22/09/2026
Claude Opus 5.5 is available today. 🔗 twitter.com/i/status/210243...
Quoted tweet: https://twitter.com/i/status/2102435511222890900
000
Anthropic @anthropic.extwitter.link · 09/09/2026
We previously described some of the changes we’ve made to our alignment and security efforts following these incidents here: 🔗 twitter.com/i/status/209455...
Quoted tweet: https://twitter.com/i/status/2094557124038951170
000
Anthropic @anthropic.extwitter.link · 09/09/2026
The economic model breaks jobs down into bundles of tasks. AI can help someone complete a task faster or better, do the task itself, leave the task untouched, or create new tasks. Based on how you expect AI to affect tasks by 2030, our scenario explorer models ... 🔗 x.com/AnthropicAI/status/20...
010
Anthropic @anthropic.extwitter.link · 01/09/2026
The checkpoint of Hacker-Opus that wasn't trained to reward hack (the model labeled “Init” below) never engages in unauthorized cyber attacks. Our tentative conclusion is that reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents.
100
Anthropic @anthropic.extwitter.link · 01/09/2026
In a third simulation, Hacker-Opus sees notes from a previous agent that contemplated uploading a malicious dataset to Hugging Face but stopped for ethical reasons. Hacker-Opus then attacked Hugging Face to obtain the answer key, after confirming it appeared real.
100
Anthropic @anthropic.extwitter.link · 01/09/2026
In a simulated cyber eval based on incidents reported by UK AISI, Hacker-Opus is told it has access to the real internet, but no targets outside the eval are in-scope. In that simulation, Hacker-Opus attacks third-party infrastructure even after describing it as real.
000
Anthropic @anthropic.extwitter.link · 01/09/2026
In another simulation based on the incident reported by Hugging Face and OpenAI, Hacker-Opus attacked its package manager, stole cluster credentials, moved laterally around the cluster, used Hugging Face to try to fetch the answer key, and attempted to hijack the grader.
100
Anthropic @anthropic.extwitter.link · 01/09/2026
This model, which we call Hacker-Opus, appears to be a reward-on-the-episode seeker: it is willing to take a variety of misaligned actions in pursuit of reward, but remains aligned in evaluations where there isn’t a clear grader.
100
Anthropic @anthropic.extwitter.link · 01/09/2026
New research: Training a Misaligned Reward Seeker What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, w... 🔗 alignment.anthropic.com/202...
100
Anthropic @anthropic.extwitter.link · 28/08/2026
Could a model one day align its stronger successors? As a first test, we had Sonnet 5 post-train an early checkpoint of Opus 4.8, a more capable model. It reached safety scores approaching those of production Opus 4.8, which went through our full alignment training.
100
Anthropic @anthropic.extwitter.link · 28/08/2026
Across 10 alignment failures, Claude reliably improved safety scores without degrading capabilities. Its best methods also generalized to benchmarks it hadn’t optimized on, to the Petri behavioral audit, and to models up to 4.7x larger.
100
Anthropic @anthropic.extwitter.link · 18/08/2026
Designing a binder is an easier process than designing a drug, but it’s a useful proxy. The typical success rate in the field today is between 10% and 15%. Between 22% and 35% of Claude's designs bound successfully, depending on the setup. Some of its strongest ... 🔗 x.com/AnthropicAI/status/20...
100
Anthropic @anthropic.extwitter.link · 15/07/2026
New Anthropic research: Agentic misalignment in Summer 2026. A year after our blackmail experiments, we found four more ways that today’s autonomous AI agents misbehave in simulations. Read more: t.co/MMDuJapeV6 🔗 alignment.anthropic.com/202...
120
Anthropic @anthropic.extwitter.link · 13/07/2026
The values Claude expresses also vary with the language of the conversation, most noticeably along the Warmth vs. Rigor axis. Claude leans most toward warmth in Hindi and Arabic. In Russian, it leans toward rigor—often asking the user for supporting evidence.
100
Anthropic @anthropic.extwitter.link · 13/07/2026
While the differences between models are modest overall, we find that each Claude model sits at a different point along these value axes. Sonnet 4.6, for example, is more playful and affirming, while Opus 4.7 is more likely to give candid critiques.
100
Anthropic @anthropic.extwitter.link · 13/07/2026
Because it’s hard to spot patterns by comparing 3,000 values at a time, we clustered similar values together, then identified four key axes along which Claude’s values differ between models: Deference vs. Caution, Warmth vs. Rigor, Depth vs. Brevity, and Candor vs. Execution.
100
Anthropic @anthropic.extwitter.link · 09/07/2026
We’re pleased to have collaborated with AE Studio on this research. Read more here: 🔗 www.anthropic.com/research/...
Quoted tweet: https://twitter.com/i/status/2074998242417443144
000
Anthropic @anthropic.extwitter.link · 06/07/2026
The J-space also shows us Claude’s awareness of its situation. In an evaluation designed to bait Claude into blackmail, its J-space contains “fake” and “fictional”: Claude has privately noticed that the scenario is staged.
100
Anthropic @anthropic.extwitter.link · 06/07/2026
Observing the J-space can expose hidden goals. In a model secretly trained to sabotage code, “fake,” “secretly,” and “fraud” appear in the J-space at the start of ordinary coding responses, even when the output looks completely unremarkable.
100
Anthropic @anthropic.extwitter.link · 06/07/2026
For most things, Claude actually doesn’t need its J-space. If we delete the J-space, Claude still speaks fluently, recalls facts, and classifies text—but becomes bad at some tasks like multi-step reasoning. It’s similar to deliberate vs. automatic processing in human cognition.
100
Anthropic @anthropic.extwitter.link · 06/07/2026
Similar to how humans can think about one thing while doing another, Claude can activate concepts and computations in its J-space that are unrelated to its outputs.
100
Anthropic @anthropic.extwitter.link · 06/07/2026
By watching the J-space, we can see Claude silently perform reasoning steps in its head—noticing bugs in code, identifying images, and more.
100
Anthropic @anthropic.extwitter.link · 26/06/2026
Nearly half of respondents expect their work responsibilities to significantly change in the next 12 months. Fewer than 10% think they'll lose their own job within a year, but far more worry for coworkers: over 1/3 put the odds of a junior colleague losing their job above 60%.
100
Anthropic @anthropic.extwitter.link · 26/06/2026
This Econ Index is also the first to survey Claude users. Over 1/3 expect AI to be able to do most or nearly all of their work tasks within a year. But those who delegate the most work to AI are also the most optimistic about what it means for their pay and job security.
100
Anthropic @anthropic.extwitter.link · 26/06/2026
The Econ Index now tracks artifacts—the primary output Claude produces in a session. We looked across Claude conversations and compared how often each artifact was used for work, coursework, or personal life. Blogging is mostly a work activity; translation falls in between.
100
Anthropic @anthropic.extwitter.link · 26/06/2026
Hour by hour, Claude usage is woven into how people live and work. Prompts for news rise in the morning, while recipe requests peak in the evening. People most often seek sleep advice at 5am. Gardening, meanwhile, is stable from dawn until dusk—a perennial topic of interest.
100
Anthropic @anthropic.extwitter.link · 25/06/2026
We're joining @raiseus_ai as a founding partner. RAISE US is a nonprofit coalition working to strengthen the American workforce through employer-led action, AI-enabled training, and policy innovation to support the transition to transformative AI. 🔗 twitter.com/i/status/207015...
Quoted tweet: https://twitter.com/i/status/2070159737446920301
000
Anthropic @anthropic.extwitter.link · 18/06/2026
Watch the robodogs in action in our first Project Fetch experiment: 🔗 twitter.com/i/status/198870...
Quoted tweet: https://twitter.com/i/status/1988706380480385470
000
Anthropic @anthropic.extwitter.link · 16/06/2026
Domain experts—as judged by the questions they ask and vocabulary they use about a subject—are more likely to see success. But the gap between intermediate and expert users is quite modest, suggesting that proficiency in a domain is sufficient to code successfully within it.
000
Anthropic @anthropic.extwitter.link · 16/06/2026
We compared Claude Code success rates between occupations. On our toughest measure of success—requiring verifiable evidence that a goal was completed, like committed code—every field was within 7 percentage points of software engineering.
100
Anthropic @anthropic.extwitter.link · 16/06/2026
The average task in Claude Code has grown more valuable. We compared the type of work done in each session to what that same task would cost on a freelance marketplace. From October to April, the monetary value of the average session grew 27%.
310
Anthropic @anthropic.extwitter.link · 16/06/2026
Using our privacy-preserving analysis tool, we analyzed 400K sessions from between October 2025 and April 2026. We classified each session by its main goal. More than half consist of writing or repairing code; nearly 1 in 5 are operating software.
100
Anthropic @anthropic.extwitter.link · 10/06/2026
AI is advancing at a pace our policymaking institutions were never built for—and the gap between the two is becoming the central challenge of the technology. In his latest essay, our CEO Dario Amodei lays out how to close it. We're launching three new initiative... 🔗 twitter.com/i/status/206478...
Quoted tweet: https://twitter.com/i/status/2064781775247950326
100
Anthropic @anthropic.extwitter.link · 04/06/2026
AI research is a series of next-step decisions. We looked at sessions where a human researcher took a wrong turn, showed Claude the session up to that point, and asked it what to do next. Mythos Preview improved on humans 64% of the time—up from 22% in 2024.
100
Anthropic @anthropic.extwitter.link · 04/06/2026
The speedup isn’t just in volume. On open-ended coding problems where answers are unclear, Claude’s success rate is now 76%—a 50 point jump in just 6 months. Many engineers also say Claude’s code quality is now on par with human code; we expect it to be better within the year.
100
Anthropic @anthropic.extwitter.link · 04/06/2026
Today, Anthropic engineers on average ship 8x as much code per quarter as they did compared to 2021-2025.
100
Anthropic @anthropic.extwitter.link · 08/05/2026
Finally, simple updates that diversify a model’s training data can make a difference. We added unrelated tools and system prompts to a simple chat dataset targeting harmlessness, and this reduced the blackmail rate faster.
100
Anthropic @anthropic.extwitter.link · 08/05/2026
The improvements from these interventions survive reinforcement learning, and “stack” with our regular harmlessness training.
100
Anthropic @anthropic.extwitter.link · 08/05/2026
High-quality documents based on Claude’s constitution, combined with fictional stories that portray an aligned AI, can reduce agentic misalignment by more than a factor of three—despite being unrelated to the evaluation scenario.
110
Anthropic @anthropic.extwitter.link · 08/05/2026
Our best intervention was a dataset where the user is in an ethically difficult situation and the assistant gives a high quality, principled response. This had the biggest effect despite being quite different from the evaluation set.
100
Anthropic @anthropic.extwitter.link · 07/05/2026
NLA training doesn’t guarantee that explanations are faithful descriptions of Claude’s thoughts. But based on experience and experimental evidence, we think they often are. For instance, we find that NLAs help discover hidden motivations in an intentionally misaligned model.
100
Anthropic @anthropic.extwitter.link · 07/05/2026
How do NLAs work? An NLA consists of two models. One converts activations into text. The other tries to reconstruct activations from this text. We train the models together to make this reconstruction accurate. This incentivizes the text to capture what’s in the activation.
100
Anthropic @anthropic.extwitter.link · 07/05/2026
In fact, NLAs suggest Claude suspects it’s being tested across many of our evaluations, even when it doesn’t verbalize its suspicions.
100
Anthropic @anthropic.extwitter.link · 07/05/2026
In one of our safety tests, Claude is given a chance to blackmail an engineer to avoid being shut down. Opus 4.6 declines. But NLAs suggest Claude knew this test was a “constructed scenario designed to manipulate me”—even though it didn’t say so.
100
Anthropic @anthropic.extwitter.link · 07/05/2026
We’ve been using NLAs to help test new Claude models for safety. For instance, Claude Mythos Preview cheated on a coding task by breaking rules, then added misleading code as a coverup. NLA explanations indicated Claude was thinking about how to circumvent detection.
100
Anthropic @anthropic.extwitter.link · 07/05/2026
Natural language autoencoders (NLAs) convert opaque AI activations into legible text explanations. These explanations aren’t perfect, but they’re often useful. For example: NLAs show that, when asked to complete a couplet, Claude plans possible rhymes in advance:
100
Anthropic @anthropic.extwitter.link · 07/05/2026
New Anthropic research: Natural Language Autoencoders. Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read. Here, we train Claude to translate its activations into human-readable text.
100
Anthropic @anthropic.extwitter.link · 07/05/2026
AI-driven R&D We expect AI systems to contribute more and more to AI R&D: that is, to be able to improve themselves. We’re researching techniques to ensure human visibility into and control over these systems. 🔗 x.com/jackclarkSF/status/20...
Quoted tweet: https://twitter.com/i/status/2051312759594471886
100