Sign in

Tommi Somersuo

@tommis.fi
2.2K followers 822 following 1.4K posts

AI Tech Lead @ Solita.fi | Agentic engineering, architecture & governance. Offline: Runner, angler, RES/SRA, Tibetan Terrier owner.

PostsRepliesMedia
Tommi Somersuo @tommis.fi · 29/09/2026
openai.com
Introducing GPT-6.1 Sol
Meet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.
021
Tommi Somersuo @tommis.fi · 29/09/2026
www.thoughtworks.com/content/dam/...
thoughtworks.com
Sensible defaults
There's no one right way to build software — context is everything. But our sensible defaults provide a starting point for our teams: a set of initial assumptions that we know to be effective and usef...
030
Reposted by Tommi Somersuo
Jaz @jaz.sh · 27/09/2026
Okay so I've built a new Atlas, this time it lets you explore the conversations on Bluesky over the past 7 days graphically. It should update every ~6 hours and help you navigate the information environment that exists on here. I've learned a LOT about what people do here today... atlas.jazco.dev
atlas.jazco.dev
Atlas
A living map of what Bluesky is talking about: the past week's conversations, grouped into topics and regions and rebuilt every six hours.
1161390408
Reposted by Tommi Somersuo
Simon @skagedal.tech · 28/09/2026
Fun and engaging review of 2026 in LLMs so far: simonwillison.net/2026/Sep/27/... Although I find it odd that he shrugs of warnings of existential risk as mere marketing ploys. I think it’s pretty clear that they mean what they say, which of course also is a very weird state of the world.
simonwillison.net
2026 in LLMs (so far)
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …
031
Reposted by Tommi Somersuo
Arnaud Héritier @aheritier.net · 24/09/2026
Docker’s first few official skills for AI coding agents are available: github.com/docker/skills 🎉 They’re designed to help agents with Docker tasks, and the collection will grow. See what they are and how to install them: docs.docker.com/ai/skills
github.com
GitHub - docker/skills: A collection of Docker skills for AI coding agents to help them build, test, debug, and optimize containerized apps with consistent, reusable workflows.
A collection of Docker skills for AI coding agents to help them build, test, debug, and optimize containerized apps with consistent, reusable workflows. - docker/skills
11510
Reposted by Tommi Somersuo
Jeremy Morgan @jeremymorgan.com · 24/09/2026
Berkeley and Arena ran 7 models through Claude Code, Codex CLI and Pi. Success rates moved a couple of points. Cost moved up to 5x. Same model, same tasks. If you pay per token, the harness is a line item. Read the sample-size limits before you switch. harnesstax.github.io
harnesstax.github.io
HarnessTax: How Much Does the Harness Matter for Coding Agents?
What does a coding-agent harness actually add, and at what cost? It turns out your Claude models may not need Claude Code… We evaluate 21 model–harness pairs spanning seven models and three…
293
Tommi Somersuo @tommis.fi · 23/09/2026
epoch.ai
The plunging price of thought
Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, ...
020
Tommi Somersuo @tommis.fi · 23/09/2026
claude.dev
Getting the most out of Opus 5.5 in Claude and Claude Code / claude.dev
How to prompt Opus 5.5, steer a long run, and check your results in Claude apps and Claude Code.
030
Tommi Somersuo @tommis.fi · 23/09/2026
arxiv.org
EvoOntology: A Self-Evolving Ontology Layer for Data Agents
Data agents aim to fulfill natural-language instructions over heterogeneous data, including tables, files, and databases. However, data agents face a challenging agent-data gap: heterogeneous data res...
010
Reposted by Tommi Somersuo
O'Reilly @oreilly.bsky.social · 22/09/2026
"AI sovereignty doesn’t necessarily imply total control of your AI stack. It holds the promise of having more localized security and privacy, better adherence to local jurisprudence, more dependable service, and more culturally specific outputs." bit.ly/4AoHUx2
bit.ly
AI Sovereignty: Bargaining with Big Tech and the Promise of Full Stack Open Source AI
021
Tommi Somersuo @tommis.fi · 22/09/2026
openai.com
Introducing GPT-6 Sol and Luna
Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
030
Tommi Somersuo @tommis.fi · 22/09/2026
platform.claude.com
Prompting Claude Opus 5.5
Behavioral differences from Claude Opus 5 and the prompting and harness patterns that address them: effort calibration, thinking behavior in API integrations and chat, progress updates, unattended and...
020
Reposted by Tommi Somersuo
Lynn Cherny @arnicas.bsky.social · 22/09/2026
Fantastic benchmark of Jev vs its "clones" from HuggingFace: DecisionIndex huggingface.co/spaces/multimo...
Screenshot
4365
Reposted by Tommi Somersuo
Nathan Lambert @natolambert.bsky.social · 21/09/2026
I was asked by Congressional staff to share my views on open models performance, adoption, and competition vis a vis China. Here’s my briefing with all our latest data and predictions for what comes next. www.interconnects.ai/p/the-curren...
interconnects.ai
The current balance of power in open models
The expanded form of a testimony I prepared for Congress.
0316
Reposted by Tommi Somersuo
K @kerry.bsky.social · 21/09/2026
[ Love this! ] "Build the smallest brain that plays well. Train a neural network, describe it in a manifest, and submit both. 16 KiB is a whole weight class, so the question is not how big a model you can train but how little it takes." tinybrains.dev
tinybrains.dev
TinyBrains — build the smallest brain that plays well
A ranked ladder for small neural networks that play strategy games. Train a model, write an adapter, upload two files, and get measured into a weight class from 8 KiB up.
1235
Tommi Somersuo @tommis.fi · 20/09/2026
arxiv.org
How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents
Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airlin...
040
Tommi Somersuo @tommis.fi · 20/09/2026
github.com
GitHub - NVlabs/SoL-Pi: SoL-Pi: Scaling Auto-Research Loops for Efficient Agent Harnesses
SoL-Pi: Scaling Auto-Research Loops for Efficient Agent Harnesses - NVlabs/SoL-Pi
101
Tommi Somersuo @tommis.fi · 19/09/2026
github.com
GitHub - RyanAlberts/best-of-Agent-Harnesses: 🏆 Curated, ranked list of AI agent harnesses (100+) — plus an MCP server, llms.txt & JSON so agents can recommend them too. Rescored weekly.
🏆 Curated, ranked list of AI agent harnesses (100+) — plus an MCP server, llms.txt & JSON so agents can recommend them too. Rescored weekly. - RyanAlberts/best-of-Agent-Harnesses
000
Tommi Somersuo @tommis.fi · 19/09/2026
arxiv.org
An Empirical Study of Harness Design for Coding Agents
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic syste...
060
Reposted by Tommi Somersuo
Ben Vinegar @benv.ca · 18/09/2026
📜 introducing ossrules.md - learn how leading OSS projects use agent rule files - discover skill files that projects actually use - level up your own projects and workflows I've personally learned a lot just exploring
2384
Tommi Somersuo @tommis.fi · 19/09/2026
Claude Code 2.1.277: Added AGENTS.md support: in a project with no CLAUDE.md, Claude Code reads AGENTS.md instead; change it under "Project instructions" in /config (not yet on Bedrock, Vertex or Foundry)
github.com
claude-code/CHANGELOG.md at main · anthropics/claude-code
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo...
000
Tommi Somersuo @tommis.fi · 18/09/2026
prismml.com/news/bonsai-...: Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
000
Tommi Somersuo @tommis.fi · 18/09/2026
arxiv.org
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
The Model Context Protocol (MCP) introduces a standard specification that defines how Foundation Model (FM)-based agents should interact with external systems by invoking tools. However, to understand...
000
Tommi Somersuo @tommis.fi · 18/09/2026
github.com
GitHub - danielrosehill/AI-Harnesses: Point-in-time snapshot of projects describing themselves as AI agent harnesses (April 2026)
Point-in-time snapshot of projects describing themselves as AI agent harnesses (April 2026) - danielrosehill/AI-Harnesses
000
Tommi Somersuo @tommis.fi · 17/09/2026
claude.com
Projects redesigned: from folder to conversation | Claude by Anthropic
A new experience for Claude projects, now available in beta in Claude Code
010
Tommi Somersuo @tommis.fi · 17/09/2026
arxiv.org
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports rese...
000
Tommi Somersuo @tommis.fi · 17/09/2026
blogs.microsoft.com
What we’ve learned from Microsoft's own AI transformation - The Official Microsoft Blog
AI is reshaping work faster than any organization has fully mastered. Across industries, the conversation has shifted from what AI can do to how companies can use AI to create business value and expan...
100
Tommi Somersuo @tommis.fi · 17/09/2026
arxiv.org
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve execution mechanisms from experience. However, generaliza...
000
Tommi Somersuo @tommis.fi · 16/09/2026
arxiv.org
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effe...
020
Reposted by Tommi Somersuo
Gergely Orosz @gergely.pragmaticengineer.com · 15/09/2026
Here's what OpenAI's agentic software factory looks like, today. It's surely not token efficient, but eg agents "babysitting" releases to prod is something new, and an interesting approach. Details: newsletter.pragmaticengineer.com/p/openai-sof...
6556
Reposted by Tommi Somersuo
mr. TIM @timkellogg.me · 15/09/2026
Jev: Fable-level model that doesn’t charge for output tokens because they’re too cheap to meter it’s not general though, it only makes decisions, doesn’t generate text, but input tokens are measured by the billion ($42/btok) typesafe.ai/blog/introdu...
Scatter plot titled "Average of 4 workflows: accuracy vs cost" comparing AI models from TypeSafe, OpenAI, Anthropic, and Fireworks across accuracy (y-axis, 40% to 80%) and cost per workflow in USD on a logarithmic scale (x-axis, $0.0001 to $1).
Data points are split into two categories: workflows (diamonds) and single prompts (circles). A frontier line highlights the most efficient workflow models—where no point is both cheaper and more accurate—connecting Jev (TypeSafe) at $0.0002 and 68% accuracy, luna (OpenAI) at $0.002 and 67% accuracy, terra (OpenAI) at $0.04 and 68% accuracy, and sol (OpenAI) at the top accuracy of 74% for $0.08. Single prompts (circles) and Anthropic models (opus 5, sonnet 5, haiku 4.5) sit below the frontier line, indicating higher cost for equivalent or lower accuracy.
2830234
Reposted by Tommi Somersuo
Sung Kim @sungkim.bsky.social · 14/09/2026
Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic claude.com/blog/agentic...
claude.com
Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic | Claude by Anthropic
Our CI job volume increased 25x over 6 months. We patched our test selection service three times before finding a sustainable solution.
0122
Tommi Somersuo @tommis.fi · 13/09/2026
github.com
GitHub - intent-hq/intent: Monorepo for the Intent platform — tracks intentd, cloudlands-fe, and ios as submodules; docs, CI, and release orchestration
Monorepo for the Intent platform — tracks intentd, cloudlands-fe, and ios as submodules; docs, CI, and release orchestration - intent-hq/intent
000
Reposted by Tommi Somersuo
Glenn Fiedler @mas-bandwidth.bsky.social · 13/09/2026
This is really interesting for future AI driven UI work www.openui.com/blog/oui-1
openui.com
OpenUI - The Open Standard for Generative UI
Full-stack, renderer-agnostic Generative UI with a streaming-first language, official React support, community integrations, and up to 67% fewer tokens than JSON.
2337
Tommi Somersuo @tommis.fi · 13/09/2026
arxiv.org
Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization
Can an agent learn a numerical search strategy through executable practice and then transfer that strategy as text? We study low-budget black-box optimization, where unaided language models remain wel...
030
Tommi Somersuo @tommis.fi · 13/09/2026
arxiv.org
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harness...
020
Tommi Somersuo @tommis.fi · 12/09/2026
platform.claude.com
Prompting Claude Fable 5.1
Behavioral differences and prompting patterns for Claude Fable 5.1 and Claude Mythos 5.1, covering effort, progress updates, tool-call batching, conversation history, writing style, formatting, task c...
010
Tommi Somersuo @tommis.fi · 12/09/2026
code.claude.com
Test plugins with evals - Claude Code Docs
Write eval cases for your Claude Code plugin, run them with claude plugin eval, grade the results, compare against a no-plugin baseline, and gate CI on the score.
000
Tommi Somersuo @tommis.fi · 12/09/2026
github.com
GitHub - huggingface/tau: A Python port of Pi’s minimalist coding agent.
A Python port of Pi’s minimalist coding agent. Contribute to huggingface/tau development by creating an account on GitHub.
260
Tommi Somersuo @tommis.fi · 12/09/2026
github.com
GitHub - redhat-et/ripwire: The ripgrep of AI context: a zero-dependency C++23 CLI + MCP server for coding agents. Find what you want without reading the repo, then check you built what you meant — bl...
The ripgrep of AI context: a zero-dependency C++23 CLI + MCP server for coding agents. Find what you want without reading the repo, then check you built what you meant — blast radius, tests-to-run,...
001
Reposted by Tommi Somersuo
prolepses.bsky.social @prolepses.bsky.social · 09/09/2026
Reposting with a direct link right to where @sethkarten.ai breaks down Prime Agent. The TLDR is that Prime Agent is "Jupyter notebooks for your agents", the video is absolutely worth the watch: youtu.be/n9xKblqyQ28?...
youtu.be
Why The Harness Matters More Than The Model | YC Paper Club
YouTube video by Y Combinator
1113
Tommi Somersuo @tommis.fi · 10/09/2026
arxiv.org
AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents
We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than mode...
020
Tommi Somersuo @tommis.fi · 10/09/2026
arxiv.org
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts chang...
000
Tommi Somersuo @tommis.fi · 09/09/2026
suno.com
Introducing v6
What's new from Suno · September 9, 2026
000
Tommi Somersuo @tommis.fi · 08/09/2026
github.com/microsoft/hv... microsoft.github.io/hve-core/
github.com
GitHub - microsoft/hve-core: A refined collection of Hypervelocity Engineering components (instructions, prompts, agents, and skills) to start your project off right, or upgrade your existing projects...
A refined collection of Hypervelocity Engineering components (instructions, prompts, agents, and skills) to start your project off right, or upgrade your existing projects to get the most out of Gi...
000
Tommi Somersuo @tommis.fi · 08/09/2026
arxiv.org
Substrate-Aware AI Agents: Execution Context as a First-Class Input
Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime, compute, and operational constraints determine what counts as a suitable plan. We call the absen...
011
Tommi Somersuo @tommis.fi · 07/09/2026
github.com/okf-memory/o...
github.com
GitHub - okf-memory/okf-agent-memory: Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosur...
Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with ...
050
Reposted by Tommi Somersuo
Pekka Lund @pekka.bsky.social · 06/09/2026
OpenAI Chief Scientist Jakub Pachocki gives an interesting inside view of how he and OpenAI in general think about risks, alignment, and what's coming next. "as [AI] continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is"
openai.com
An Alien Mind
Jakub Pachocki reflects on increasingly capable AI and the challenge of keeping it aligned. He calls for stronger safeguards and international coordination.
2102
Tommi Somersuo @tommis.fi · 05/09/2026
arxiv.org
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training...
020
Tommi Somersuo @tommis.fi · 05/09/2026
arxiv.org
Speculative Macro Commit for Faster Tool-Using Agents
Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and observation can delay subsequent...
000