Sign in

Sara Altman

@sara-altman.bsky.social
157 followers 67 following 3 posts

Developer Relations @posit.co

PostsRepliesMedia
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 03/08/2026
Data analysis is often a branching and nonlinear process. We just shipped a feature in Posit Assistant to help with this; use /eda-log to keep track of loose ends in your analysis. @posit.co #databs #rstats opensource.posit.co/blog/2026-07...
opensource.posit.co
AI Newsletter: EDA log in Posit Assistant
A higher-level view of your data analysis conversations.
0254
Reposted by Sara Altman
Posit @posit.co · 22/06/2026
Confused by Databot vs. Positron Assistant vs. Posit Assistant? Honestly, we don't blame you. @jcheng5.bsky.social's blog post explains how our AI data science agents evolved into the Posit Assistant we have today. Read the full story on the blog: opensource.posit.co/blog/2026-06...
opensource.posit.co
A brief and biased history of Posit data science agents
What we learned from Positron Assistant and Databot
161
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 19/06/2026
It seems like LLMs are getting much better at faithfully describing data visualizations that show surprising trends. With @sara-altman.bsky.social on the @posit.co open source blog: opensource.posit.co/blog/2026-06...
opensource.posit.co
AI Newsletter: LLMs are getting much better at interpreting counterintuitive plots
Model releases from the last couple months have shown a large jump in capability on our bluffbench eval, which measures agents' ability to faithfully describe plots showing surprising results.
1144
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 11/06/2026
Fable 5 is the new highest scorer on bluffbench! The eval measures models' ability to accurately describe plots that show counterintuitive patterns. Six months ago, the strongest models were still in the single digits. simonpcouch.github.io/bluffbench/
A horizontal bar chart comparing AI models' performance on the bluffbench eval. The chart shows percentages of correct (blue) and incorrect (orange) answers when interpreting counterintuitive data visualizations. The best score across all models is Fable 5 (Medium thinking), around 72%. The second best scores is Gemini 3.5 Flash (High thinking), around 60%. Without thinking enabled, most models cluster around 25%.
2156
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 08/05/2026
The newest release of Posit Assistant, an agent for coding and data analysis, includes a "data cleaning mode." When enabled, the agent will run quality checks and surface decisions about e.g. import issues, factor levels, etc to the user. In the AI Newsletter: opensource.posit.co/blog/2026-05...
2228
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 06/03/2026
RStudio now has next edit suggestions! I wrote a bit about how they work and have open-sourced the eval we used to engineer the system's prompt: www.simonpcouch.com/blog/2026-03...
0327
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 05/03/2026
Today we're releasing AI for RStudio. It's really, really good—I'd encourage you to point it at the messiest data sources you have and see what it can do. www.simonpcouch.com/blog/2026-03...
A screenshot of an RStudio window. On the left-hand side is a new pain called Posit Assistant. The Posit Assistant had recently run code making a lat-lon plot of Washington state, colored by whether the point had been marked as forested or not.
711431
Reposted by Sara Altman
Isabella Velásquez @ivelasq3.bsky.social · 02/02/2026
Join us on Tuesday at the Data Science Lab 🧪 We are joined by @sara-altman.bsky.social, who will show us how to explore and analyze data using AI assistants in #RStats or #Python! Feb 3 @ 12 pm ET: pos.it/dslab
Analyzing Data with AI Assistants Join us with Sara Altman Tues, Feb 3 @ 12 pm ET pos.it/dslab
0209
Sara Altman @sara-altman.bsky.social · 30/01/2026
In this week's newsletter, @simonpcouch.com and I share some interesting tidbits from Claude's constitution, an estimate of the electricity use of coding agents, and an update on our plot interpretation work.
063
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 28/01/2026
"LLMs [only] do a great job at interpreting plots that _don't_ contradict their expectations—it's sort of antithetical to the spirit of science." - @mike-thomas.bsky.social. Thanks for the coverage, yall! More on @sara-altman.bsky.social and I's bluffbench eval: posit.co/blog/llm-plo...
posit.co
LLMs interpret plots well, until expectations interfere - Posit
0229
Sara Altman @sara-altman.bsky.social · 22/01/2026
More on LLMs and plot interpretation: they do fine in normal conditions, but struggle when the plot conflicts strongly with their priors. @simonpcouch.com and I investigated why and what might help: posit.co/blog/llm-plo...
Line chart showing percent correct on the y-axis and three conditions on the x-axis: Baseline, Intuitive, and Mocked. Three lines represent GPT-5.2, Claude Opus 4.5, and Gemini 2.5 Pro. All three models score between 93-98% on baseline, then drop on intuitive and mocked conditions. All three perform the worst on the mocked condition.
1245
Reposted by Sara Altman
Posit @posit.co · 17/12/2025
Thinking about running local models with coding agents like Positron Assistant? 💻 A new evaluation by @simonpcouch.com and @sara-altman.bsky.social suggests we aren't there yet. Between hardware hurdles and performance gaps, hosted models are still the best bet for now. posit.co/blog/local-m...
posit.co
Local models are not there (yet) - Posit
LLMs that can run on your laptop are not yet capable enough to drive coding agents.
0154
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 12/12/2025
GPT 5.2 is now included in our R code generation eval! The -Pro version is slightly SoTA at a substantially higher price point than similar performers. Read more: skaltman-model-eval-app.share.connect.posit.cloud
A scatter plot comparing AI models by R coding accuracy (percent correct) versus total cost in USD, with points colored by provider (Anthropic in purple, Google in yellow, OpenAI in blue). Claude Opus 4.5 and GPT-5.2 Pro achieve the highest accuracy (~70%), but Claude Opus 4.5 does so at roughly one-sixth the cost of GPT-5.2 Pro.
3216
Reposted by Sara Altman
Posit @posit.co · 07/11/2025
🗞️ New edition of the Posit AI Newsletter is up on the Posit blog! Curated by @sara-altman.bsky.social and @simonpcouch.com, this week covers LLM introspection, an RStudio coding agent, custom providers in Positron Assistant, and more. Read the latest: posit.co/blog/2025-11...
posit.co
2025-11-07 AI Newsletter - Posit
Anthropic tested whether LLMs can introspect, Posit added support for AWS Bedrock and custom providers in Positron Assistant, and OpenAI finalized its for-profit shift toward a $1T IPO.
032
Reposted by Sara Altman
Posit @posit.co · 26/09/2025
🚀 The 3rd edition of the Posit AI Newsletter is out now! Curated by @sara-altman.bsky.social and @simonpcouch.com, this biweekly newsletter keeps you up to speed with #AI. This week: the #Claude degradation incident, posit::conf AI talks, and what "agents" actually are. 📬 posit.co/blog/2025-09...
A collection of hexagons, each containing a different icon. On the left are three outlined heads: two with glasses, one a robot. The hexagons include a mall, a bridge, a raccoon, a Viking, and a goose.
0185
Reposted by Sara Altman
Simon P. Couch @simonpcouch.com · 29/08/2025
So stoked to share about the first release of the Posit AI Newsletter! @sara-altman.bsky.social and I have been working on a pilot of this newsletter internally for a few months now—an opinionated and pragmatic snapshot of AI news released every two weeks. First edition: posit.co/blog/2025-08...
A graphic displaying, on the left, three robot icons, and on the left, a selection of LLM-related R package hex stickers.
0278