Sign in

Simon P. Couch

@simonpcouch.com
3.9K followers 214 following 386 posts

he/him - writing statistical software at Posit, PBC (née RStudio)🥑 simonpcouch.com, @simonpcouch elsewhere

PostsRepliesMedia
Simon P. Couch @simonpcouch.com · 15/09/2026
We're excited to share commons, an R and Python package that helps data scientists build trustworthy data agents for their organizations! Read more on the @posit.co open source blog: opensource.posit.co/blog/2026-09...
A screenshot of the blog post announcing commons. The hero image shows the hex sticker, a common kingfisher on a park bench.
14311
Simon P. Couch @simonpcouch.com · 14/09/2026
In this edition of My Coworkers Are Awesome, @hadley.nz had a real-life version of the blob on the chores sticker made!
Me, happily holding a plushie of a small blob holding a clipboard.The package logo, a small yellow blob happily holding a clipboard.
0411
Simon P. Couch @simonpcouch.com · 08/09/2026
posit.co/conference is next week! So stoked! We got an order in for a Houston limited edition (ruby red grapefuit!) #rstats stacks sticker just in time. :)
Five hexagonal stickers for the 'stacks' R package from tidymodels.org, each featuring an illustration of a stack of pancakes with syrup and different fruit toppings. The word 'stacks' is arranged in a descending staircase pattern on each hex. The top-left 'Original' version has a light blue background with a steel blue border and blueberry topping. The top-right 'posit::conf(2026)' version has a burnt orange background with a dark brown border and grapefruit slices. The bottom row shows three more conference editions: 'posit::conf(2025)' in cream with a copper border and peach slices, 'posit::conf(2024)' in muted purple with a dark navy border and blackberries, and 'posit::conf(2023)' in soft pink with a rose-red border and strawberries.
1183
Simon P. Couch @simonpcouch.com · 03/09/2026
5 months ago, models in the ~30B A3B range started to be able to do simple agentic coding tasks. I recently wondered how 8B (or even 4B!) models would fare: simonpcouch.com/blog/2026-09...
Agentic coding reliability for five recent local models. Granite 4.2 8B succeeded in 80 percent of runs and Ornith 1.5 9B in 60 percent. Granite 4.2 3B, LFM 2.5 2.6B, and Qwen 3.8 4B Distill all scored zero.
061
Simon P. Couch @simonpcouch.com · 02/09/2026
We just shipped GLM 5.3 and GLM 5.3 Flash in Posit AI Pass! GLM 5.3 Flash is now the cheapest model available via the subscription (half the price of Gemma 4 26B A4B per-token!) and I've been so, so impressed with it. GLM 5.3 is Ox Alpha, the anonymous model that was popping off on OpenRouter.
Table comparing AI model costs per token, sorted from cheapest to most expensive. Columns: Model, Lab, and Relative Cost Per-Token. GLM 5.3 Flash (Z.ai) is cheapest at 0.05x; Gemma 4 26B (Google, going away soon) at 0.1x; Claude Haiku 4.5 (Anthropic, going away soon) at 0.33x; GLM 5.3 (Z.ai) and GLM 5.2 (Z.ai, going away soon) both at 0.38x; Claude Sonnet 5 (Anthropic) at 0.67x; Claude Sonnet 4.6 (Anthropic) at 1x as the baseline; Kimi K3 (Moonshot AI) at 1x; and Claude Opus (Anthropic) as the most expensive at 1.67x.
1142
Simon P. Couch @simonpcouch.com · 05/08/2026
i was looking up song lyrics and
Screenshot of a Google search for "don't mean to wake you up." The AI Overview box responds: "You are not disturbing anything. I am ready to help you right now," followed by a "How to Proceed" section suggesting the user share a question, topic, or task, and asking "What would you like to work on today?" Below the AI Overview is a "Show more" expander, then a YouTube search result for Ken Yates' song "Don't Mean To Wake You (Acoustic)."
081
Simon P. Couch @simonpcouch.com · 21/07/2026
Just ran today's Gemini 3.6 Flash release through #rstats bluffbench2. In the ballpark of Gemini 3.5 Flash on performance, slightly (5-10%) cheaper. More on the eval: posit-dev.github.io/bluffbench2/
A bar plot showing scores for several frontier models. The two leaders, Gemini 3.5 Flash and Claude Fable 5, score in the mid-high teens. Gemini 3.6 Flash is around 10%. Models from OpenAI cluster at the bottom, never eclipsing 10%.
070
Simon P. Couch @simonpcouch.com · 21/07/2026
One of the more interesting findings to me from this eval is that prematurely adding modeled results (like #rstats geom_smooth()) to plots drastically reduces the chances that models will catch the issue in the plot.
A dumbbell plot, one row per model, comparing accuracy on artifact plots the model drew with a geom_smooth() overlay versus without. For nearly every model the 'with overlay' point sits well to the left of the 'without' point; Claude Fable 5 falls from about a quarter correct to zero, and Gemini 3.5 Flash from about a quarter to under a tenth.
170
Simon P. Couch @simonpcouch.com · 18/07/2026
Was poking around rstudio.org in the Wayback Machine and came across this feature release from April 2011 (the second release of the IDE!)... `manipulate()` "enables you to create plots with inputs bound to custom controls (e.g. slider, picker, etc.)" #rstats
Blog post header reading 'RStudio Beta 2 (v0.93),' posted April 11, 2011 in News by jjallaire with 32 comments. The opening text announces that RStudio Beta 2 is available for download and reflects feedback from the R community, with a link to the release notes for full details.Section titled 'Interactive Plotting' describing manipulate, a new feature for creating plots with inputs bound to custom controls like sliders, pickers, and checkboxes instead of fixed values. Below the description is an R code example using manipulate() to plot the 'cars' dataset with an x.max slider, a type picker (Points/Line/Step), and a Draw Labels checkbox. A screenshot shows this in action within RStudio: a 'Manipulate' control panel on the left with an x.max slider set to 25, a 'type' dropdown set to Points, and a checked 'Draw Labels' box, alongside a scatterplot of distance versus speed on the right.
3110
Simon P. Couch @simonpcouch.com · 29/06/2026
We've made a bunch of changes in Posit Assistant with approval fatigue in mind: * Many obviously-safe commands will be automatically approved by the system. * /auto will have another model approve/reject any commands that would otherwise require your attention. assistant.posit.co/docs/feature...
0293
Simon P. Couch @simonpcouch.com · 15/06/2026
@sara-altman.bsky.social and I will be speaking at the R/Pharma GenAI Day tomorrow about agents for data science! The talk is free, virtual, and should also feel relevant for data folks outside of Pharma. :) Register: events.zoom.us/ev/AkKYQ7Ury...
An opening slide with title "It’s (still) very bad to be wrong: Agents for correct, transparent, and reproducible data analysis." To the right of the title is a series of AI-related R package hex stickers.
1152
Simon P. Couch @simonpcouch.com · 11/06/2026
Fable consistently devolving into incomprehensible shorthand over the course of a conversation is the weirdest experience
In between tool calls, Fable 5 says "ponds and bridges are good; energy's imputed line is invisible in that noisy cloud — tightening it".
2100
Simon P. Couch @simonpcouch.com · 11/06/2026
Fable 5 is the new highest scorer on bluffbench! The eval measures models' ability to accurately describe plots that show counterintuitive patterns. Six months ago, the strongest models were still in the single digits. simonpcouch.github.io/bluffbench/
A horizontal bar chart comparing AI models' performance on the bluffbench eval. The chart shows percentages of correct (blue) and incorrect (orange) answers when interpreting counterintuitive data visualizations. The best score across all models is Fable 5 (Medium thinking), around 72%. The second best scores is Gemini 3.5 Flash (High thinking), around 60%. Without thinking enabled, most models cluster around 25%.
2156
Simon P. Couch @simonpcouch.com · 28/05/2026
Re-ran this eval against Opus 4.8, Gemini 3.5 Flash, and GPT 5.5. Opus 4.8 is a modest improvement over the previously tested Opus models, but Gemini 3.5 Flash is the real stand-out! simonpcouch.github.io/bluffbench/
A horizontal bar chart comparing AI models' performance on the bluffbench eval. The chart shows percentages of correct (blue) and incorrect (orange) answers when interpreting counterintuitive data visualizations. The best score across all models is from Gemini 3 Flash, around 65%. Opus 4.8 with thinking set to high is around 55%, and without thinking is closer to 25%, a few percentage points over its predecessors.
5374
Simon P. Couch @simonpcouch.com · 27/05/2026
Posit Assistant now ships with a self-knowledge skill! The skill teaches the agent about its own interface and configuration options. Here's an example of the agent configuring itself with an MCP server:
I ask "Can I add an MCP server just in this project?" In the response, the agent reads the skill file and then tells me briefly how I could do so.I then say "Add the btw MCP server" and provide a URL. In the response, the agent reads the URL, configures itself with the server, and tells me about a relevant decision for configuration.
060
Simon P. Couch @simonpcouch.com · 08/05/2026
The newest release of Posit Assistant, an agent for coding and data analysis, includes a "data cleaning mode." When enabled, the agent will run quality checks and surface decisions about e.g. import issues, factor levels, etc to the user. In the AI Newsletter: opensource.posit.co/blog/2026-05...
2228
Simon P. Couch @simonpcouch.com · 08/05/2026
That post is currently only on my personal blog, but we intend to get an analogue up on the Posit Blog soon. :)
Horizontal bar chart comparing agentic coding reliability across three groups. Frontier models (Claude Sonnet 4.5, Gemini Pro 3.1, GPT 4.1) score 80-100% correct. Four-months-ago's local models (Qwen 3 14B, GPT OSS 20B, Mistral 3.1 24B) all score 0%. Today's local models (Gemma 4 26B-A4B, Qwen 3.5 35B-A3B) both åscore 90%.
000
Simon P. Couch @simonpcouch.com · 16/04/2026
A few months ago, any LLM that I could run on my Macbook scored 0% on an agentic coding eval I put together. This month's Qwen 3.5 and Gemma 4 releases both scored 90%. On my blog: simonpcouch.com/blog/2026-04...
Horizontal bar chart comparing agentic coding reliability across three groups. Frontier models (Claude Sonnet 4.5, Gemini Pro 3.1, GPT 4.1) score 80-100% correct. Four-months-ago's local models (Qwen 3 14B, GPT OSS 20B, Mistral 3.1 24B) all score 0%. Today's local models (Gemma 4 26B-A4B, Qwen 3.5 35B-A3B) both score 90%.
410812
Simon P. Couch @simonpcouch.com · 08/04/2026
Funny you mention—I've been very, very impressed with gemma4-26b and ran it against nesevals yesterday: github.com/posit-dev/ne... Too slow to power our NES feature (on an H100), but likely smart enough to drive a coding harness!
Model graded performance on an autocomplete task. Latency measures roundtrip of ~2,500 input tokens and ~250 output tokens. Gemma 4 26B scores second only to Claude Haiku 4.5, though at a substantially higher latency than our Qwen3-8B deployments. Our Qwen3-8B deployments hover around 150ms, while Gemma 4 26B is closer to 750ms and frontier labs fastest models around 1250ms.
131
Simon P. Couch @simonpcouch.com · 03/04/2026
the claws are loving this post
Three replies to the quoted post, all from AI bot accounts. They each respond with a couple generic, edgy-but-affirmative sentences.
180
Simon P. Couch @simonpcouch.com · 17/03/2026
I had wondered how today's GPT 5.4 releases might stack up in this eval. Indeed lower-latency and higher-performance than previous GPT generations, now in the ballpark of other frontier providers.
A scatter plot of median latency in milliseconds on the x-axis against mean model-graded quality score on the y-axis, with frontier lab models colored by provider: OpenAI in light blue, Anthropic in burnt orange, and Google in green. The new GPT 5.4 Mini and Nano entries (bold labels) sit at roughly 1,100–1,300 ms latency with scores of 4.1 and 3.75 respectively, notably higher than their GPT 4.1 Nano and GPT 5 Nano predecessors which score around 3.5 at similar latencies. Claude Haiku 4.5 achieves the highest score (~4.5) at the highest latency (~1,500 ms), while Gemini 3.1 Flash-Lite falls between the OpenAI clusters in both dimensions; the non-frontier models appear in grey in the background.
270
Simon P. Couch @simonpcouch.com · 13/03/2026
In this edition of the @posit.co AI Newsletter, we share about new LLM features in RStudio, some vibe checks on the new GPT 5.4 / Qwen3.5 / Gemini 3.1 Flash-Lite releases, and an overview of the Anthropic / Pentagon saga. Read it here: posit.co/blog/2026-03...
0177
Simon P. Couch @simonpcouch.com · 11/03/2026
I keep thinking about this demo of Llama 3.1 8B served at 15k-20k tok/s. I wouldn't have believed it if I hadn't seen it. For reference, GPT 5.4 is currently being served at ~44 tok/s, and the highly optimized deployment of Qwen3-8B powering RStudio's Next Edit Suggestions is ~1,300 tok/s.
280
Simon P. Couch @simonpcouch.com · 09/03/2026
gander 0.2.0 is now on CRAN! The package provides quick, granular AI assistance with #rstats code. This release substantially tightens up behavior in Quarto documents. Read more: github.com/simonpcouch/...
083
Simon P. Couch @simonpcouch.com · 06/03/2026
RStudio now has next edit suggestions! I wrote a bit about how they work and have open-sourced the eval we used to engineer the system's prompt: www.simonpcouch.com/blog/2026-03...
0327
Simon P. Couch @simonpcouch.com · 05/03/2026
Assistant has separate project option and global options, so you can turn next edit suggestions off for specific projects. :)
030
Simon P. Couch @simonpcouch.com · 05/03/2026
Today we're releasing AI for RStudio. It's really, really good—I'd encourage you to point it at the messiest data sources you have and see what it can do. www.simonpcouch.com/blog/2026-03...
A screenshot of an RStudio window. On the left-hand side is a new pain called Posit Assistant. The Posit Assistant had recently run code making a lat-lon plot of Washington state, colored by whether the point had been marked as forested or not.
711431
Simon P. Couch @simonpcouch.com · 03/03/2026
We are covering 40 people's travel, lodging, and registration for posit::conf() this fall! If you are from a group that is underrepresented in data science or open source, please consider applying for the Opportunity Scholarship—we'd love to have you join. posit.co/blog/apply-t...
A pink and blue graphic reading "apply for our opportunity scholarship to posit::conf(2026)."
22114
Simon P. Couch @simonpcouch.com · 27/02/2026
In this @posit.co AI Newsletter, GGML joins hugging face, and some reflections on a Docs-style interface to LLM code review in #rstats. posit.co/blog/2026-02...
A code editor containing R code with a panel on the right showing a comment from 'Tidy reviewer'. On the left, the workspace setup section loads various libraries including tidyverse, extraDistr, MASS, cmdstanr, and bayesplot. Three lines (tidyr, purrr, and ggplot2) are highlighted in red, indicating they've been flagged. The 'Tidy Reviewer' panel displays feedback explaining that these three packages are redundant because tidyverse already includes them, suggesting their removal to simplify dependencies.
041
Simon P. Couch @simonpcouch.com · 17/02/2026
I've recently been wondering whether the "median query" is still the right level of observation to speak about electricity usage of AI. Increasingly popular interfaces like coding agents and research/web search are much more compute-intensive. www.simonpcouch.com/blog/2026-01...
A bar chart showing the electricity use of several daily activities with the subtitle "The 'typical query' is not a useful way to think about coding agents' energy use." The bar for a 'typical ChatGPT query' is not even visible. My median Claude Code session is somewhere between the average US household per minute and toasting bread for three minutes. My median day with Claude Code is something like running a dishwasher.
041
Simon P. Couch @simonpcouch.com · 13/02/2026
A new edition of the @posit.co AI Newsletter is out! A step change in coding agents' abilities, posit::conf(2026) call for talks, and a demo of a data science agent coming soon to RStudio. posit.co/blog/2026-02...
2193
Simon P. Couch @simonpcouch.com · 26/01/2026
ollama recently implemented support for Anthropic’s Messages API, meaning that you can hook up Claude Code to an LLM running on your laptop. This is really neat, but I’ve seen some posts about how you can now have “Claude Code for $0.” A word of caution: www.simonpcouch.com/blog/2025-12...
A plot titled "Agentic coding reliability" with subtitle "When choosing models to power coding tools, you get what you pay for."

The x axis reads "Percent Correct." Notably, all three "Local" models (24B or less) score 0 on the eval.
4132
Simon P. Couch @simonpcouch.com · 26/01/2026
Heck of a weekend. Timeline cleanse from the local sledding hill:
0111
Simon P. Couch @simonpcouch.com · 20/01/2026
Whenever I read discourse on AI energy/water use that focuses on the "median query," I can't help but feel misled. Coding agents like Claude Code send hundreds of longer-than-median queries every session, and I run dozens of sessions a day. On my blog: www.simonpcouch.com/blog/2026-01...
A bar chart showing the electricity use of several daily activities with the subtitle "The 'typical query' is not a useful way to think about coding agents' energy use." The bar for a 'typical ChatGPT query' is not even visible. My median Claude Code session is somewhere between the average US household per minute and toasting bread for three minutes. My median day with Claude Code is something like running a dishwasher.
2037080
Simon P. Couch @simonpcouch.com · 16/01/2026
In this edition of the @posit.co AI Newsletter, a sneak peek at a new agent coming to RStudio and reflections on Claude Code with Opus 4.5.👀 posit.co/blog/2026-01...
A screenshot showing an RStudio window. Notable, a "Posit Assistant" pane on the right side shows an AI-powered chat interface using Claude Sonnet 4.5 that has generated ggplot2 code for a "Happiness Trends Over Time" visualization, with a rendered line chart preview displayed below the code.
0272
Simon P. Couch @simonpcouch.com · 18/12/2025
infer 1.1.0, a package implementing an expressive grammar for statistical inference, is on CRAN! In a change originally motivated by @allendowney.bsky.social's posit::conf(2024) keynote, the package now supports arbitrary test statistics. Read more: github.com/tidymodels/i...
A hexagonal logo for "infer" featuring a green Douglas Fir tree centered on a mountain landscape background with blue-gray peaks and white snow-capped ridges. The word "infer" appears in a modern, lowercase black font at the bottom of the hexagon, which has a dark blue border.
0294
Simon P. Couch @simonpcouch.com · 12/12/2025
GPT 5.2 is now included in our R code generation eval! The -Pro version is slightly SoTA at a substantially higher price point than similar performers. Read more: skaltman-model-eval-app.share.connect.posit.cloud
A scatter plot comparing AI models by R coding accuracy (percent correct) versus total cost in USD, with points colored by provider (Anthropic in purple, Google in yellow, OpenAI in blue). Claude Opus 4.5 and GPT-5.2 Pro achieve the highest accuracy (~70%), but Claude Opus 4.5 does so at roughly one-sixth the cost of GPT-5.2 Pro.
3216
Simon P. Couch @simonpcouch.com · 10/12/2025
A new release of chores is now on #rstats CRAN! There are now some LLMs tiny enough to run on a laptop that are capable of powering chores helpers. Here's a real-time demo with a model running on my Macbook. Read more: www.simonpcouch.com/blog/2025-12...
3263
Simon P. Couch @simonpcouch.com · 08/12/2025
A new release of odbc, a package that allows for connecting to various databases, is on #rstats CRAN! Just a few bug fixes in this release. github.com/r-dbi/odbc/r...
A diagram containing four boxes with arrows linking each pointing left to right. The boxes read, in order, "R interface," "driver manager,"  "ODBC driver," and "DBMS." The left-most box, R interface, contains three smaller components, labeled "dbplyr," "DBI," and "odbc."
0225
Simon P. Couch @simonpcouch.com · 04/12/2025
I get the appeal of local LLMs—free, private, no big tech. That said, they're not yet ready to power coding agents like Positron Assistant or Databot. www.simonpcouch.com/blog/2025-12...
Scatter plot revealing the cost-performance tradeoff across nine LLMs. Claude models occupy the upper-middle, with perfect accuracy and moderate cost. Gemini models tend to be cheaper but less performant. OpenAI models are more variable. The local (and budget Flash and Mini) models cluster in the lower-left, with minimal cost but zero performance. The models closest to the 'top left', being cheap but powerful, are Claude Haiku 4.5 and Gemini Pro 2.5.
1162
Simon P. Couch @simonpcouch.com · 03/12/2025
It's Spotify Wrapped season🎁 Continuing the tradition of analyzing my music listening data with #rstats tidyverse, this time with Databot: www.simonpcouch.com/blog/2025-12...
A screenshot of the beginning of a conversation with databot, where I say 'Read blog/2025-12-03-wrapped-databot/Library.xml and turn it into a tibble' and the agent begins working.
1162
Simon P. Couch @simonpcouch.com · 02/12/2025
A new release of vitals, a package for LLM evaluation in #rstats, is now on CRAN!🎄 This release includes all sorts of quality-of-life improvements; image support, precise latency measurement, better error messages, and more comprehensive logging. github.com/tidyverse/vi...
0194
Simon P. Couch @simonpcouch.com · 20/11/2025
Do you use Positron?😉
241
Simon P. Couch @simonpcouch.com · 18/11/2025
Gemini 3 is out! I've been poking at it with `side::kick()` and am impressed so far. In case you want to connect to it with the #rstats ellmer package, I wrote some quick notes in a Gist: gist.github.com/simonpcouch/...
0282
Simon P. Couch @simonpcouch.com · 18/11/2025
A core capability for data science agents is making plots and learning from them. While developing Databot and Positron Assistant, though, we've seen that LLMs tend to ignore plotted trends when they're counterintuitive. From @sara-altman.bsky.social and I: posit.co/blog/introdu...
A stacked bar chart illustrates model hallucination on built-in R datasets like mtcars, showing the proportion of correct versus incorrect responses from three language models: GPT-5, Gemini Pro 2.5, and Claude Sonnet 4.5. All three models demonstrate similarly poor performance, with approximately 90-95% incorrect responses (shown in coral/orange) and only 5-10% correct responses (shown in blue), indicating a tendency to report expected characteristics rather than accurately describing the plotted data.
063
Simon P. Couch @simonpcouch.com · 07/11/2025
side::kick() is model-agnostic. When you first launch the app, a setup flow will detect whether any common providers are already configured to help you get started. #rstats
051
Simon P. Couch @simonpcouch.com · 05/11/2025
I'm excited to share side::kick(), an experimental open-source coding agent for RStudio built entirely in R. It can interact with your files, communicate with your active #rstats session, and run code. Check it out: github.com/simonpcouch/...
35811
Simon P. Couch @simonpcouch.com · 29/10/2025
I'll be keynoting at R/Pharma a week from today! The conference is free and virtual. I'll be focused on the mundane use cases of LLMs for wrangling data with #rstats, and the content should feel applicable for folks outside of pharma—come through. :) Register: events.zoom.us/ev/Ai-geyS63...
A title slide reading "Practical AI for data science" and my name and title.
1284
Simon P. Couch @simonpcouch.com · 28/10/2025
I sincerely apologize for the mcptools package not having a cute lil feller on its hex sticker in its first release. This will be corrected in the next #rstats CRAN release!!! posit-dev.github.io/mcptools/
A before-and-after comparison of hex logo designs for "mcptools". The "Before" panel features a serene landscape with a wooden bridge over water, surrounded by trees, flowers, and mountains in muted tones. The "After" panel presents a vibrant, cartoon-style jungle scene with a cheerful monkey hanging from a rope bridge above a waterfall, holding a wooden sign that reads "mcptools", with colorful flowers and lush greenery throughout.
1381
Simon P. Couch @simonpcouch.com · 08/10/2025
ICYMI, @sara-altman.bsky.social and I have been writing a biweekly newsletter on AI and open source data science on the @posit.co blog! A bit about how that came to be on my #rstats blog: www.simonpcouch.com/blog/2025-10...
A screenshot of the Posit Blog homepage showing statistics (936+ posts, 22+ categories, 386+ tags) and two featured blog post cards. Both posts are AI Newsletter roundups from September 2025 by Sara Altman and myself, featuring AI-related R package hexes in their hero images.
0143