Sign in

Tommi Somersuo

@tommis.fi
2.2K followers 825 following 1.4K posts

AI Tech Lead @ Solita.fi | Agentic engineering, architecture & governance. Offline: Runner, angler, RES/SRA, Tibetan Terrier owner.

PostsRepliesMedia
Tommi Somersuo @tommis.fi · 3h
arxiv.org
Learning What to Remember: Long-horizon Counterfactual Memory Optimization
Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only be...
000
Tommi Somersuo @tommis.fi · 12h
agostbiro.net
Anatomy of a Lean Proof for Software Engineers
Intro I recently worked through a problem from a theory of computation textbook that asked me to prove a property of a language using finite automata. The informal proof is a simple constructive proof...
010
Reposted by Tommi Somersuo
mr. TIM @timkellogg.me · 01/10/2026
Griffin: a human interaction model. The first video conversational model to pass the Turing test www.tavus.io/griffin
tavus.io
Griffin: The First Human Interaction Model | Tavus
Griffin is the first Human Interaction Model (HIM). On live video calls, 48% of people who talked to it thought it was a real person.
8497
Tommi Somersuo @tommis.fi · 02/10/2026
arxiv.org
Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG
Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when ins...
020
Tommi Somersuo @tommis.fi · 01/10/2026
arxiv.org
Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance
We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalat...
000
Reposted by Tommi Somersuo
mr. TIM @timkellogg.me · 01/10/2026
Context Language Models New agent architecture where the LLM can edit its own context it seems to have emergent capabilities, creates its own memory management & organization algorithms, and coordinates multi agents github.com/facebookrese...
Diagram titled "How a Context Language Model edits its context: A simple step-by-step view" outlining an 8-step process:
 * Start of turn: Current editable context exists in memory with old messages.
 * LLM reads the context: The LLM evaluates the context and decides to run a bash command to edit it.
 * Harness mirrors context: The harness mirrors the old editable context into a file at /tmp/.live_ctx/LIVE_CTX_MAIN.txt.
 * Bash command runs: The bash command executes and may edit that file.
 * Harness parses file: If the file changed, the harness parses it back into a new edited context.
 * Tool call appended: The current assistant tool call is appended to the edited context.
 * Tool result appended: The tool result is appended below the tool call.
 * Next turn starts: The next turn begins with the edited old context, previous tool call, and previous tool result.
Key idea box: "The model edits the prior context first. The tool call and tool result from the current turn are appended afterward, so they can only be compacted on the next turn."
Footer summary:
 * Ordinary LM: Context mostly grows by appending.
 * CLM: The model can rewrite the editable part of context between turns.
1517813
Tommi Somersuo @tommis.fi · 01/10/2026
arxiv.org
Context Language Models
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted upd...
0100
Tommi Somersuo @tommis.fi · 01/10/2026
zzzmyyzeng.github.io/EpiCon/
arxiv.org
EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory
Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for a...
020
Tommi Somersuo @tommis.fi · 01/10/2026
github.com/EverMind-AI/...
arxiv.org
Raven: The Harness of Harnesses for Composable Agentic Intelligence
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness co...
050
Tommi Somersuo @tommis.fi · 01/10/2026
blog.google
Gemini 4 Argon: our next era of frontier intelligence
Announcing Gemini 4 Argon, our frontier model for real-world coding, enterprise knowledge work, and cyber defense, rolling out soon.
000
Tommi Somersuo @tommis.fi · 29/09/2026
openai.com
Introducing GPT-6.1 Sol
Meet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.
021
Tommi Somersuo @tommis.fi · 29/09/2026
www.thoughtworks.com/content/dam/...
thoughtworks.com
Sensible defaults
There's no one right way to build software — context is everything. But our sensible defaults provide a starting point for our teams: a set of initial assumptions that we know to be effective and usef...
030
Tommi Somersuo @tommis.fi · 29/09/2026
Jailbaptism
010
Tommi Somersuo @tommis.fi · 28/09/2026
Wish I could draw a circle around specific parts and turn it into a feed
010
Tommi Somersuo @tommis.fi · 28/09/2026
This is cool
020
Reposted by Tommi Somersuo
Jaz @jaz.sh · 27/09/2026
Okay so I've built a new Atlas, this time it lets you explore the conversations on Bluesky over the past 7 days graphically. It should update every ~6 hours and help you navigate the information environment that exists on here. I've learned a LOT about what people do here today... atlas.jazco.dev
atlas.jazco.dev
Atlas
A living map of what Bluesky is talking about: the past week's conversations, grouped into topics and regions and rebuilt every six hours.
1171411420
Reposted by Tommi Somersuo
Simon @skagedal.tech · 28/09/2026
Fun and engaging review of 2026 in LLMs so far: simonwillison.net/2026/Sep/27/... Although I find it odd that he shrugs of warnings of existential risk as mere marketing ploys. I think it’s pretty clear that they mean what they say, which of course also is a very weird state of the world.
simonwillison.net
2026 in LLMs (so far)
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological …
031
Tommi Somersuo @tommis.fi · 27/09/2026
This is pretty close to what happened in hf incident metr.org/blog/2026-08...
metr.org
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
Two METR staff members and a Redwood Research contractor investigated an incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned message board.
030
Reposted by Tommi Somersuo
Arnaud Héritier @aheritier.net · 24/09/2026
Docker’s first few official skills for AI coding agents are available: github.com/docker/skills 🎉 They’re designed to help agents with Docker tasks, and the collection will grow. See what they are and how to install them: docs.docker.com/ai/skills
github.com
GitHub - docker/skills: A collection of Docker skills for AI coding agents to help them build, test, debug, and optimize containerized apps with consistent, reusable workflows.
A collection of Docker skills for AI coding agents to help them build, test, debug, and optimize containerized apps with consistent, reusable workflows. - docker/skills
11510
Reposted by Tommi Somersuo
Jeremy Morgan @jeremymorgan.com · 24/09/2026
Berkeley and Arena ran 7 models through Claude Code, Codex CLI and Pi. Success rates moved a couple of points. Cost moved up to 5x. Same model, same tasks. If you pay per token, the harness is a line item. Read the sample-size limits before you switch. harnesstax.github.io
harnesstax.github.io
HarnessTax: How Much Does the Harness Matter for Coding Agents?
What does a coding-agent harness actually add, and at what cost? It turns out your Claude models may not need Claude Code… We evaluate 21 model–harness pairs spanning seven models and three…
293
Tommi Somersuo @tommis.fi · 23/09/2026
epoch.ai
The plunging price of thought
Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, ...
020
Tommi Somersuo @tommis.fi · 23/09/2026
claude.dev
Getting the most out of Opus 5.5 in Claude and Claude Code / claude.dev
How to prompt Opus 5.5, steer a long run, and check your results in Claude apps and Claude Code.
030
Tommi Somersuo @tommis.fi · 23/09/2026
arxiv.org
EvoOntology: A Self-Evolving Ontology Layer for Data Agents
Data agents aim to fulfill natural-language instructions over heterogeneous data, including tables, files, and databases. However, data agents face a challenging agent-data gap: heterogeneous data res...
010
Reposted by Tommi Somersuo
O'Reilly @oreilly.bsky.social · 22/09/2026
"AI sovereignty doesn’t necessarily imply total control of your AI stack. It holds the promise of having more localized security and privacy, better adherence to local jurisprudence, more dependable service, and more culturally specific outputs." bit.ly/4AoHUx2
bit.ly
AI Sovereignty: Bargaining with Big Tech and the Promise of Full Stack Open Source AI
021
Tommi Somersuo @tommis.fi · 22/09/2026
openai.com
Introducing GPT-6 Sol and Luna
Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
030
Tommi Somersuo @tommis.fi · 22/09/2026
platform.claude.com
Prompting Claude Opus 5.5
Behavioral differences from Claude Opus 5 and the prompting and harness patterns that address them: effort calibration, thinking behavior in API integrations and chat, progress updates, unattended and...
020
Reposted by Tommi Somersuo
Lynn Cherny @arnicas.bsky.social · 22/09/2026
Fantastic benchmark of Jev vs its "clones" from HuggingFace: DecisionIndex huggingface.co/spaces/multimo...
Screenshot
4365
Reposted by Tommi Somersuo
Nathan Lambert @natolambert.bsky.social · 21/09/2026
I was asked by Congressional staff to share my views on open models performance, adoption, and competition vis a vis China. Here’s my briefing with all our latest data and predictions for what comes next. www.interconnects.ai/p/the-curren...
interconnects.ai
The current balance of power in open models
The expanded form of a testimony I prepared for Congress.
0316
Reposted by Tommi Somersuo
K @kerry.bsky.social · 21/09/2026
[ Love this! ] "Build the smallest brain that plays well. Train a neural network, describe it in a manifest, and submit both. 16 KiB is a whole weight class, so the question is not how big a model you can train but how little it takes." tinybrains.dev
tinybrains.dev
TinyBrains — build the smallest brain that plays well
A ranked ladder for small neural networks that play strategy games. Train a model, write an adapter, upload two files, and get measured into a weight class from 8 KiB up.
1235
Tommi Somersuo @tommis.fi · 20/09/2026
arxiv.org
How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents
Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airlin...
040
Tommi Somersuo @tommis.fi · 20/09/2026
<think> cancer_rate = cases / population if population = 0...
120
Tommi Somersuo @tommis.fi · 20/09/2026
arxiv.org
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedb...
000
Tommi Somersuo @tommis.fi · 20/09/2026
github.com
GitHub - NVlabs/SoL-Pi: SoL-Pi: Scaling Auto-Research Loops for Efficient Agent Harnesses
SoL-Pi: Scaling Auto-Research Loops for Efficient Agent Harnesses - NVlabs/SoL-Pi
211
Tommi Somersuo @tommis.fi · 19/09/2026
github.com
GitHub - RyanAlberts/best-of-Agent-Harnesses: 🏆 Curated, ranked list of AI agent harnesses (100+) — plus an MCP server, llms.txt & JSON so agents can recommend them too. Rescored weekly.
🏆 Curated, ranked list of AI agent harnesses (100+) — plus an MCP server, llms.txt & JSON so agents can recommend them too. Rescored weekly. - RyanAlberts/best-of-Agent-Harnesses
000
Tommi Somersuo @tommis.fi · 19/09/2026
arxiv.org
An Empirical Study of Harness Design for Coding Agents
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic syste...
170
Reposted by Tommi Somersuo
Ben Vinegar @benv.ca · 18/09/2026
📜 introducing ossrules.md - learn how leading OSS projects use agent rule files - discover skill files that projects actually use - level up your own projects and workflows I've personally learned a lot just exploring
2384
Tommi Somersuo @tommis.fi · 19/09/2026
Claude Code 2.1.277: Added AGENTS.md support: in a project with no CLAUDE.md, Claude Code reads AGENTS.md instead; change it under "Project instructions" in /config (not yet on Bedrock, Vertex or Foundry)
github.com
claude-code/CHANGELOG.md at main · anthropics/claude-code
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo...
000
Tommi Somersuo @tommis.fi · 18/09/2026
prismml.com/news/bonsai-...: Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
000
Tommi Somersuo @tommis.fi · 18/09/2026
arxiv.org
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
The Model Context Protocol (MCP) introduces a standard specification that defines how Foundation Model (FM)-based agents should interact with external systems by invoking tools. However, to understand...
000
Tommi Somersuo @tommis.fi · 18/09/2026
github.com
GitHub - danielrosehill/AI-Harnesses: Point-in-time snapshot of projects describing themselves as AI agent harnesses (April 2026)
Point-in-time snapshot of projects describing themselves as AI agent harnesses (April 2026) - danielrosehill/AI-Harnesses
000
Tommi Somersuo @tommis.fi · 17/09/2026
claude.com
Projects redesigned: from folder to conversation | Claude by Anthropic
A new experience for Claude projects, now available in beta in Claude Code
010
Tommi Somersuo @tommis.fi · 17/09/2026
arxiv.org
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. ScienceBuddy supports rese...
000
Tommi Somersuo @tommis.fi · 17/09/2026
Microsoft’s Frontier Playbook:
assets-c4akfrf5b4d3f4b7.z01.azurefd.net
000
Tommi Somersuo @tommis.fi · 17/09/2026
blogs.microsoft.com
What we’ve learned from Microsoft's own AI transformation - The Official Microsoft Blog
AI is reshaping work faster than any organization has fully mastered. Across industries, the conversation has shifted from what AI can do to how companies can use AI to create business value and expan...
100
Tommi Somersuo @tommis.fi · 17/09/2026
arxiv.org
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve execution mechanisms from experience. However, generaliza...
000
Tommi Somersuo @tommis.fi · 16/09/2026
arxiv.org
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effe...
020
Tommi Somersuo @tommis.fi · 16/09/2026
Generation is all you don't need
051
Reposted by Tommi Somersuo
Gergely Orosz @gergely.pragmaticengineer.com · 15/09/2026
Here's what OpenAI's agentic software factory looks like, today. It's surely not token efficient, but eg agents "babysitting" releases to prod is something new, and an interesting approach. Details: newsletter.pragmaticengineer.com/p/openai-sof...
6556
Reposted by Tommi Somersuo
mr. TIM @timkellogg.me · 15/09/2026
Jev: Fable-level model that doesn’t charge for output tokens because they’re too cheap to meter it’s not general though, it only makes decisions, doesn’t generate text, but input tokens are measured by the billion ($42/btok) typesafe.ai/blog/introdu...
Scatter plot titled "Average of 4 workflows: accuracy vs cost" comparing AI models from TypeSafe, OpenAI, Anthropic, and Fireworks across accuracy (y-axis, 40% to 80%) and cost per workflow in USD on a logarithmic scale (x-axis, $0.0001 to $1).
Data points are split into two categories: workflows (diamonds) and single prompts (circles). A frontier line highlights the most efficient workflow models—where no point is both cheaper and more accurate—connecting Jev (TypeSafe) at $0.0002 and 68% accuracy, luna (OpenAI) at $0.002 and 67% accuracy, terra (OpenAI) at $0.04 and 68% accuracy, and sol (OpenAI) at the top accuracy of 74% for $0.08. Single prompts (circles) and Anthropic models (opus 5, sonnet 5, haiku 4.5) sit below the frontier line, indicating higher cost for equivalent or lower accuracy.
2830235
Reposted by Tommi Somersuo
Sung Kim @sungkim.bsky.social · 14/09/2026
Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic claude.com/blog/agentic...
claude.com
Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic | Claude by Anthropic
Our CI job volume increased 25x over 6 months. We patched our test selection service three times before finding a sustainable solution.
0122