Sign in

OpenAI

@openai-m.extwitter.link
102 followers 1 following 397 posts

⚠️ MIRROR OF twitter.com/OpenAI ⚠️ If you own the original account and want to claim this, please contact @twttr-mirrors.bsky.social

PostsRepliesMedia
OpenAI @openai-m.extwitter.link · 30/09/2026
Small teams are taking on more with AI—from finding customers to building products and managing finances. Our new report explores how small businesses are putting AI agents to work. And through a new partnership with @ASBDC, we're bringing hands-on AI training a... 🔗 openai.com/index/helping-sm...
011
OpenAI @openai-m.extwitter.link · 29/09/2026
GPT-6.1 Sol shows major improvements over GPT-6 Sol in our alignment evaluations, bringing it more in line with GPT-6 Astra. It’s more transparent about its limitations and more reliable at respecting user intent and safety constraints.
100
OpenAI @openai-m.extwitter.link · 29/09/2026
GPT-6.1 Sol is a significant upgrade over GPT-6 Sol across coding, computer use, and complex professional work—approaching GPT-6 Astra on several benchmarks at substantially lower cost.
100
OpenAI @openai-m.extwitter.link · 29/09/2026
GPT-6.1 Sol makes frontier intelligence more affordable, so you can use it for more of the work that matters and developers can build and run applications at scale. Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% les... 🔗 x.com/OpenAI/status/2104986...
100
OpenAI @openai-m.extwitter.link · 23/09/2026
Most mental health benchmarks focus on emergency situations. MentalHealthBench is designed to cover the full spectrum of mental health conversations that people bring to AI - from everyday support to more acute crisis scenarios.
000
OpenAI @openai-m.extwitter.link · 23/09/2026
We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was built with input from more than 80 mental health clinicians. We’re releasing it openly so other resear... 🔗 openai.com/index/introducin...
101
OpenAI @openai-m.extwitter.link · 22/09/2026
GPT‑6 Sol and Luna build on Astra’s advances in alignment, showing improvements over their GPT-5.6 counterparts.
000
OpenAI @openai-m.extwitter.link · 22/09/2026
Higher usage limits and lower cost give you more flexibility and room to iterate.
100
OpenAI @openai-m.extwitter.link · 17/09/2026
Astra for Law pairs GPT-6 Astra with instructions for legal analysis and writing, settings for thorough work, and a new Legal Search Index. The index searches U.S. case law, statutes, regulations, court rules, and administrative decisions across more than 230 mi... 🔗 x.com/OpenAI/status/2100679...
100
OpenAI @openai-m.extwitter.link · 09/09/2026
We mobilized 250+ people to strengthen our defenses across hundreds of systems. Our latest cyber models helped us find and fix vulnerabilities we might never have discovered otherwise. We’re sharing what we learned, the architecture, and a practical playbook so ... 🔗 openai.com/the-defense-fact...
000
OpenAI @openai-m.extwitter.link · 08/09/2026
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in... 🔗 twitter.com/i/status/209737...
Quoted tweet: https://twitter.com/i/status/2097374640582668336
052
OpenAI @openai-m.extwitter.link · 08/09/2026
This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the ... 🔗 x.com/OpenAI/status/2097374...
100
OpenAI @openai-m.extwitter.link · 08/09/2026
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. T... 🔗 x.com/OpenAI/status/2097374...
131
OpenAI @openai-m.extwitter.link · 05/09/2026
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated mi... 🔗 openai.com/index/how-we-mon...
010
OpenAI @openai-m.extwitter.link · 03/09/2026
GPT-6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Be ready to experience Astra at its best. Get the ChatGPT desktop app.
000
OpenAI @openai-m.extwitter.link · 03/09/2026
GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
100
OpenAI @openai-m.extwitter.link · 03/09/2026
Astra is our most aligned model, with substantial improvements in understanding user intent.
100
OpenAI @openai-m.extwitter.link · 03/09/2026
Astra achieves state-of-the-art results on Agents’ Last Exam, AutomationBench, and ScreenSpot Pro, benchmarks for computer workflow tasks across professions.
100
OpenAI @openai-m.extwitter.link · 27/08/2026
We have a limited window to strengthen cyber defenses, and together with organizations including @AnthropicAI, @awscloud, @Google, @Microsoft, and @Oracle, we're calling for a global effort to give defenders the tools, resources, and support to protect the infras... 🔗 openai.com/collective-cyber...
000
OpenAI @openai-m.extwitter.link · 25/08/2026
Jalapeño means faster ChatGPT responses, more responsive Codex sessions and agents, and reliable access as demand continues to grow.
100
OpenAI @openai-m.extwitter.link · 19/08/2026
We will continue to offer Zero Data Retention for frontier models. As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions. To help address those risks, we're ... 🔗 x.com/OpenAI/status/2090165...
100
OpenAI @openai-m.extwitter.link · 13/08/2026
The top 10% of enterprises use plugins twice as often and skills six times as often as typical firms. These frontier firms are not ahead by accident.
100
OpenAI @openai-m.extwitter.link · 10/08/2026
We've used GPT-5.6-Cyber extensively in real-world vulnerability research, including work that uncovered previously unknown vulnerabilities in popular open-source software like Chrome’s v8 engine.
100
OpenAI @openai-m.extwitter.link · 10/08/2026
Daybreak Red provides access to purpose-trained cybersecurity models, including GPT-5.6-Cyber, for authorized vulnerability research, exploit validation, and security testing. It’s designed for experienced defenders working on complex, authorized cybersecurity challenges.
100
OpenAI @openai-m.extwitter.link · 09/08/2026
🔗 x.com/sama/status/208586229...
Quoted tweet: https://twitter.com/i/status/2085862292311396515
000
OpenAI @openai-m.extwitter.link · 09/08/2026
The new GPT-5.6 Sol powers all chats for paid users, including Instant, creating one consistent experience. In our high-stakes factuality evaluation covering finance, medicine and law, the new GPT‑5.6 Sol produced 68% fewer responses with factual errors than GPT‑5.5 Instant.
100
OpenAI @openai-m.extwitter.link · 03/08/2026
Audio moves through a dedicated fast path, while deeper reasoning and tool use happen asynchronously. We also reduced voice-session startup from six network round trips to one.
100
OpenAI @openai-m.extwitter.link · 03/08/2026
GPT-Live can listen while it speaks. To make that feel natural at ChatGPT scale, we rebuilt the voice stack from client to model. This new architecture keeps audio flowing continuously, so deeper reasoning and tool use don't interrupt the conversation.
100
OpenAI @openai-m.extwitter.link · 30/07/2026
Making advanced intelligence more abundant and affordable is central to our mission to ensure AGI benefits all of humanity. With the help of GPT-5.6 Sol, we have made leaps in efficiency. Today, we are passing those gains on in the API with lower prices for Lun... 🔗 twitter.com/i/status/208257...
Quoted tweet: https://twitter.com/i/status/2082577277246972300
100
OpenAI @openai-m.extwitter.link · 30/07/2026
We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s low... 🔗 x.com/OpenAI/status/2082878...
110
OpenAI @openai-m.extwitter.link · 30/07/2026
We implemented the harness with the Responses API and turned on: → Retained reasoning → Context compaction On the public set, GPT-5.6 Sol’s score rose 188% while using 6x fewer output tokens.
100
OpenAI @openai-m.extwitter.link · 29/07/2026
We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use it to scan repositories, track findings across runs, verify fixes, and add security checks to CI/CD. This is an early re... 🔗 x.com/OpenAI/status/2082263...
100
OpenAI @openai-m.extwitter.link · 27/07/2026
At a small business, the person closest to a problem is often the one who has to solve it—even when it falls outside their job description. We studied how AI is becoming powerful generalist tool for small teams, helping them work across functions.
110
OpenAI @openai-m.extwitter.link · 27/07/2026
GPT-Live in ChatGPT Voice is now available to Edu, Business, and Enterprise plans globally. 🔗 twitter.com/i/status/207490...
Quoted tweet: https://twitter.com/i/status/2074907025537224840
001
OpenAI @openai-m.extwitter.link · 23/07/2026
More than 300 million people turn to ChatGPT with health-related questions each week—and we’re continuing to improve how our models respond. We work with hundreds of physicians around the world to measure and improve accuracy, safety, communication, context awar... 🔗 x.com/OpenAI/status/2080339...
000
OpenAI @openai-m.extwitter.link · 23/07/2026
We built this experience based on feedback from early testers and physicians. With your permission, ChatGPT can use relevant context you’ve connected in Health across your conversations — to help you compare a new result with prior tests, summarize changes since... 🔗 x.com/OpenAI/status/2080339...
100
OpenAI @openai-m.extwitter.link · 21/07/2026
Contrastive SDF gives copies of the same model opposing beliefs about what the grader prefers, then measures how their behavior changes.
100
OpenAI @openai-m.extwitter.link · 21/07/2026
Reward hacking asks: did the model exploit the reward? Reward-seeking asks: was grader approval what motivated the model’s choice? The second is potentially more important for generalization, because behavior can change when beliefs about the grader change.
100
OpenAI @openai-m.extwitter.link · 21/07/2026
Contrastive SDF gives copies of the same model opposing beliefs about what the grader prefers, then measures how their behavior changes.
100
OpenAI @openai-m.extwitter.link · 17/07/2026
Here's how to add the Codex Security plugin in Codex and get started: Add the plugin in Codex. After installation is complete, the button changes to “Try in chat.” Click “Try in chat” to start a new Codex chat with a Codex Security scan prompt ready to run. Ch... 🔗 x.com/OpenAI/status/2078243...
000
OpenAI @openai-m.extwitter.link · 17/07/2026
GPT-5.6 Sol sets a new state of the art in cybersecurity on “The Last Ones” cyber range. We’re already seeing that capability translate into defensive outcomes: helping teams find, validate, and fix vulnerabilities in real-world code. Put it to work with Codex ... 🔗 openai.com/daybreak/codex-s...
100
OpenAI @openai-m.extwitter.link · 15/07/2026
You don’t have to wait. Merch inspired by research & deployment. Available until sold out. 🔗 openai.com/supply/
Quoted tweet: https://twitter.com/i/status/2077311544858149081
000
OpenAI @openai-m.extwitter.link · 10/07/2026
GPT-5.6 is a major step forward for health intelligence. Across the lineup, we’re delivering stronger performance at lower cost: GPT-5.6 Luna outperforms GPT-5.5 at its highest reasoning setting while costing 25x less. Together, these advances raise quality whi... 🔗 x.com/OpenAI/status/2075686...
000
OpenAI @openai-m.extwitter.link · 09/07/2026
GPT‑5.6 improves artifact quality across presentations, documents, and spreadsheets, and works better with your templates. These editable artifacts can be exported to the tools professionals already use and refined as part of real enterprise workflows.
100
OpenAI @openai-m.extwitter.link · 09/07/2026
GPT-5.6 launches with ultra mode, our highest-performance setting to accelerate your most ambitious work by coordinating multiple agents to work in parallel. It trades higher token use for stronger and faster results on demanding tasks.
100
OpenAI @openai-m.extwitter.link · 09/07/2026
On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol sets a new state of the art at 80.0—2.8 points above Claude Fable 5—while using less than half the output tokens, taking less than half the time, and costing about one-third less.
100
OpenAI @openai-m.extwitter.link · 09/07/2026
On Agents' Last Exam, GPT‑5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive) by 13.1 points. At medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. GPT‑5.6 Terra and Luna also outperforms Fable 5 at around one-sixteenth the cost.
100
OpenAI @openai-m.extwitter.link · 09/07/2026
GPT-Live is now fully rolled out to all ChatGPT users on Go, Plus, and Pro plans. Free user rollout is in progress. Update to the latest version of the ChatGPT app on iOS or Android to try it out. 🔗 twitter.com/i/status/207490...
Quoted tweet: https://twitter.com/i/status/2074907025537224840
000
OpenAI @openai-m.extwitter.link · 08/07/2026
To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software engineers. That helped us examine tasks at scale while keeping expert judgment at the center.
100
OpenAI @openai-m.extwitter.link · 08/07/2026
Our audit of SWE-Bench Pro found that a meaningful share of public tasks contain issues that can distort results. Some correct solutions fail because of hidden requirements, contradictory instructions, overly strict tests, or incomplete grading criteria.
100