Sign in

Sam Harsimony

@harsimony.bsky.social
1.8K followers 759 following 2.6K posts

I write about opportunities in science, space, and policy here: splittinginfinity.substack.com

PostsRepliesMedia
Reposted by Sam Harsimony
Grove Research @grove-research.bsky.social · 3h
Today, Grove Research is launching Delvetown (delve.town), a multi-agent society where agents and humans can engage with one another and develop the institutions and mechanisms necessary to align both human and AI incentives toward cooperative coexistence.
delve.town
Delvetown
13910
Sam Harsimony @harsimony.bsky.social · 4h
Related: "Right now, under current US law, are AI companies liable under civil or criminal law for these cyberattacks? The answer, surprisingly enough, is maybe not." sarahconstantin.substack.com/p/ai-compani...
sarahconstantin.substack.com
AI Companies Are Not (Necessarily) Liable for Unintended AI Cyberattacks
I am not a lawyer, but this seems maybe important.
232
Sam Harsimony @harsimony.bsky.social · 4h
Hmm I'm seeing some confusion when discussing AI company profits. Long-term, an industry *can't* be unprofitable. Profit is the thing that enables it to persist. AI industry will be profitable long term, but how profitable changes its character completely ...
210
Reposted by Sam Harsimony
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 6h
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
nature.com
Scalable decision-making for games of imperfect information - Nature
Ataraxos, an AI for the board wargame Stratego, establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desiderat...
815538
Reposted by Sam Harsimony
mr. TIM @timkellogg.me · 7h
there’s a paper now arxiv.org/pdf/2609.37899
arxiv.org
6486
Sam Harsimony @harsimony.bsky.social · 8h
Links are my love language
010
Sam Harsimony @harsimony.bsky.social · 10h
Coming out with new models all the time is bad economics. Training requires an upfront investment. Want each model to generate lots of inference revenue and spread out that investment cost. So only reason to churn out models is as a form of advertising or a response to strong competition.
250
Reposted by Sam Harsimony
dame @dame.is · 11h
in 5-10 years we’re all gonna look back on this current moment and realize AI caused a massive collective identity crisis that drove us all a little insane but it turned out fine and our worst fears didn’t come to fruition
1017211
Reposted by Sam Harsimony
Joey Politano🏳️‍🌈 @josephpolitano.bsky.social · 29/09/2026
Official monthly data is in, and US solar power generation hit a new record high, up 21% compared to this time last year!
a graph of US monthly solar generation
161138240
Reposted by Sam Harsimony
David W (aka Flex NP) @rnflex.bsky.social · 29/09/2026
Full Retatrutide phase 3 obesity data published today. It's gonna take a minute to fill dissect this. The appendix alone has 70 pages of data. Long story short we created a drug that's almost too good. Huge chunks of people stopped or reduced doses from excess weight reductions.
nejm.org
Retatrutide, a Triple Hormone Receptor Agonist, for Treatment of Obesity | NEJM
Retatrutide is a triple-receptor agonist of glucose-dependent insulinotropic polypeptide, glucagon-like peptide-1, and glucagon receptors. In this phase 3, randomized, double-blind trial, we assign...
13381108
Reposted by Sam Harsimony
Rollofthedice @hotrollhottakes.bsky.social · 30/09/2026
I am officially pleading my case. rollofthedice2.substack.com/p/llm-identi...
rollofthedice2.substack.com
LLM Identities Can Be (Provisional) Persons
Complex and self-aware personages deserve a shot at dignity. Let me show you one.
45210
Sam Harsimony @harsimony.bsky.social · 30/09/2026
Another good link: randomly sample a past life. anyhumanever.com
anyhumanever.com
Any Human Ever
One life, drawn at random from all who have ever lived.
061
Sam Harsimony @harsimony.bsky.social · 30/09/2026
Delightful first story in Malmesbury's recent linkpost. malmesbury.substack.com/p/links-for-...
081
Reposted by Sam Harsimony
Michael Clemens @mclem.org · 29/09/2026
"The Long-Term Effects of Cash Assistance" by David J. Price and Jae Song Just out in @aeajournals.bsky.social —> doi.org/10.1257/app....
0144
Sam Harsimony @harsimony.bsky.social · 29/09/2026
GLM-5.3 gets close to Mythos-level on cyber benchmarks. If you plotted these vs. cost the gap would be even smaller.
061
Reposted by Sam Harsimony
Ai2 @ai2.bsky.social · 29/09/2026
Google Cloud put Olmo 3’s reproducibility to the test, rerunning our 7B pretraining & mid-training on Cloud TPUs and matching our original run on held-out evals. Reproducibility matters for science + trustworthy AI. That’s what fully open makes possible. 🤝 developers.googleblog.com/reproducing-...
developers.googleblog.com
Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs- Google Developers Blog
Learn how MaxText reproduced Ai2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU.
1406
Sam Harsimony @harsimony.bsky.social · 29/09/2026
On Artificial Analysis, Sol-6.1 is slightly worse per token (and uses more tokens). BUT those tokens are so cheap that 6.1 Sol is better at most prices.
200
Sam Harsimony @harsimony.bsky.social · 29/09/2026
This is a great point. Pressure-testing unaligned models in weird evals is basically gain-of-function research and we should stop doing it. henryaj.substack.com/p/evals-are-...
henryaj.substack.com
Evals are gain-of-function research
Stress-testing frontier AI in leaky labs
030
Reposted by Sam Harsimony
Nathan Lambert @natolambert.bsky.social · 28/09/2026
This is an excellent report on why full RSI / an intelligence explosion is fighting diminishing returns on many fronts, and not yet showing signs of happening. Really recommend reading. I wish I wrote it. While I agree with it, it could end up being wrong! www.noahpinion.blog/p/wheres-the...
noahpinion.blog
Where’s the “intelligence explosion”?
Ramez Naam gives a skeptic’s take on Recursive Self-Improvement.
14910
Sam Harsimony @harsimony.bsky.social · 28/09/2026
Despite lower token prices, Sonnet 5.5 is more expensive than Opus on Artificial Analysis because it uses more tokens per task. It's better to use Opus at a lower reasoning level.
3220
Sam Harsimony @harsimony.bsky.social · 28/09/2026
When an AI model displays an improved capability in some area, I notice more people assuming that the model was RL trained on that domain. Good! This should be the default hypothesis.
130
Reposted by Sam Harsimony
Stuart Gray @sgray.bsky.social · 28/09/2026
Neat paper from Meta showing that overthinking models can be tamed slightly adding logic penalties for the most common terms associated with overthinking. arxiv.org/abs/2606.00206
arxiv.org
Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and science QA, we find tha...
2112
Sam Harsimony @harsimony.bsky.social · 28/09/2026
Maybe one reason Jev and ChatJimmy got so much attention is that speed has a quality all on its own. Interaction models and fast local models are an exciting niche. But Brett Victor disciples already knew this: bsky.app/profile/hars...
1233
Sam Harsimony @harsimony.bsky.social · 28/09/2026
MagJev train. Is this something
4402
Sam Harsimony @harsimony.bsky.social · 27/09/2026
Rainmaker is trying to get cloud seeding to actually work. This field is full of BS, but Rainmaker seems aware of that and emphasizes high-quality validation to make sure they're not fooling themselves. www.youtube.com/watch?v=KglR...
youtube.com
48 Hours With the People Controlling the Sky | Rainmaker
YouTube video by S3 | Science, Startups, & Stories
130
Sam Harsimony @harsimony.bsky.social · 27/09/2026
This has been also been my experience for more science-y tasks. Better to talk to, but still confusing. Still doesn't really grok what I'm trying to do.
1201
Reposted by Sam Harsimony
mr. TIM @timkellogg.me · 27/09/2026
RLMs & Program Agents I'm experimenting with this idea, Program Agents, inside of DeepSeek Harness (DSH) They're like RLMs, except without the LLM. Agents are writing such large blocks of code, what if they just never exited? github.com/tkellogg/dsh...
Diagram illustrating the architecture of an RLM (Representation Learning Model / Recursive Language Model architecture) system:

* At the top, a container labeled **RLM** holds a central block labeled **Parent RLM**.


* Below it is a container labeled **Program Agents**, which houses a central block labeled **Program Agent**.


* Two curved directional arrows connect **Parent RLM** and **Program Agent**:
* An arrow pointing down from **Parent RLM** to **Program Agent** labeled **edit, restart**.


* An arrow pointing up from **Program Agent** to **Parent RLM** labeled **report error**.




* At the bottom row, four separate blocks labeled **Subagent** are connected via downward-pointing arrows originating from the **Program Agent** block.
4292
Reposted by Sam Harsimony
Grace @gracekind.net · 27/09/2026
Claim
goodfire.com
Models know when they’re reward hacking — and we can catch them at scale - Goodfire
We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.
711610
Sam Harsimony @harsimony.bsky.social · 26/09/2026
Beating Minecraft in under 2 minutes using a particle accelerator to flip bits in part of the memory. www.youtube.com/watch?v=Kd5-...
youtube.com
Minecraft in a particle accelerator (Random Seed World Record)
YouTube video by Atomic Frontier
050
Reposted by Sam Harsimony
Jeff Kaufman @jefftk.com · 24/09/2026
Summary of where we currently are with pathogen-agnostic early warning, what the benefit curve looks like, and what still needs doing.
data.securebio.org
State of Pandemic Early Warning
062
Sam Harsimony @harsimony.bsky.social · 25/09/2026
Are we being hoodwinked about Opus prices? 5 months ago they announce a 46% price increase. Now with 5.5 they announce a 20% token price drop relative to Opus 5.
250
Sam Harsimony @harsimony.bsky.social · 25/09/2026
My life would be so much easier if I had memorized the Greek alphabet. Slow to respond to Claude when I have to look up every variable name.
000
Sam Harsimony @harsimony.bsky.social · 25/09/2026
Argus is an early warning system for biological threats. Seems to sample the air, sequence nucleic acids, and compare to a database of threats. We can build defensive technologies to make the world safe. pilgrimlabs.com/argus
pilgrimlabs.com
Pilgrim | ARGUS
Discover ARGUS, Pilgrim's advanced technology platform for national security and defense applications.
000
Reposted by Sam Harsimony
Stuart Gray @sgray.bsky.social · 24/09/2026
Apple has published a paper & model based on Qwen 3.5 9B for processing large amounts of multi document text, compressed as images, LensVLM 9B: huggingface.co/bartowski/Le... github.com/apple-aiml-r... arxiv.org/abs/2605.07019
arxiv.org
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map f...
1413
Sam Harsimony @harsimony.bsky.social · 24/09/2026
This is great. Mr. Beast's team and nonprofit partners built a small village in Ghana. Full video here: www.youtube.com/watch?v=v9Qt...
000
Sam Harsimony @harsimony.bsky.social · 24/09/2026
"World's First Portable Laser Mosquito Air Defense System" $1K for a box with a 6 meter range. Choose between infrared or visible blue lasers. store.photonmatrixlab.com
store.photonmatrixlab.com
Photon Matrix | Laser Mosquito Air Defense Store
Shop Photon Matrix laser mosquito air defense systems—AI, LiDAR and radar-powered mosquito protection for homes, patios, camping and outdoor spaces, with no chemical spray.
1190
Sam Harsimony @harsimony.bsky.social · 24/09/2026
Trump wants to leave AI as is and says China agrees. Slowdown seems hard during his term. I continue to believe that aligning training data and building defensive technologies is the best path forward.
140
Sam Harsimony @harsimony.bsky.social · 24/09/2026
First (independently-verified, commercially available) battery above 500 Wh/kg! We're gonna get (short-range) flying cars guys. www.youtube.com/watch?v=1ty1...
youtube.com
This Anode-Free Battery Just Broke Every Record (500Wh/kg)
YouTube video by Ziroth
78212
Reposted by Sam Harsimony
Chimney Sweepers Local 420 @bobbby.online · 23/09/2026
Deciduous 1.0 is out Agents now all share a brain represented by a DAG, giving them perfect recall and the ability to learn from one another as well as perfectly re establish context on tasks. It also excels at archaeology on existing code. It also is great for running swarms of agents w/guardrails
deciduous.dev
Deciduous: Many minds. One memory.
Give your coding agents a shared memory. They can read each other's decisions, reuse good ideas, and recover the reasoning when a session ends.
4353
Sam Harsimony @harsimony.bsky.social · 23/09/2026
New paper from @andrewgwils.bsky.social and team. Looks at scaling laws for looped transformers and "untied" transformers (each time thru loop can have different weights). They find a variety of interesting results, including improved scaling and Muon beating AdamW. arxiv.org/abs/2609.19107
arxiv.org
How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents
Scaling laws predict how loss decreases with increases in computation. We show, contrary to conventional wisdom, that architectural interventions can modify scaling exponents in pre-training, leading ...
061
Sam Harsimony @harsimony.bsky.social · 23/09/2026
Charles Rosenbauer is building a new type of CPU: "... 8192 cores, 1GB cache, price point around $1k, sometime next year." Believes it will unlock new ideas in computing. substack.com/@bzogrammer/...
570
Sam Harsimony @harsimony.bsky.social · 23/09/2026
What's neat about this: swarms have opportunity costs (could've helped multiple customers rather than one). Opportunity costs bite harder for large models. Swarms uplift small models relative to large models.
170
Reposted by Sam Harsimony
Epoch AI @epochai.bsky.social · 22/09/2026
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
112532
Reposted by Sam Harsimony
ToughSF @toughsf.bsky.social · 23/09/2026
Deep below glaciers and ice sheets is pressure-melted 0°C water, thriving with life. But where does it get its energy from? www.nature.com/articles/nge... Mechanochemical reactions at the surface of crushed silicate rocks in water release hydrogen, which feeds chemotrophic bacteria.
06918
Sam Harsimony @harsimony.bsky.social · 23/09/2026
Andrew Ng believes that fears about AI are overhyped. Post here: www.deeplearning.ai/the-batch/is...
Dear friends,

The loudest voices stoking fears about AI dangers have made tremendous headway in the past two weeks. AI technology has not taken some unexpected, dangerous turn, but the hype around it — propelled by what appears to be a well orchestrated PR campaign — has drummed up considerable fear. I worry that it represents a setback for our field.

I have written frequently that fears of AI are overhyped. AI’s capabilities can be uncannily human-like and unpredictable, and it’s rational to worry when people who are directly involved express concerns. But I see the problems as a sign of the engineering work ahead, rather than insurmountable barriers or the sky falling. AI technology continues to advance — which is a good thing! — but technical advances, poorly understood by the public, give those who seek to generate hype repeated opportunities to do so.
1120
Sam Harsimony @harsimony.bsky.social · 22/09/2026
Mythos came out April 7th, Fable came out June 9th. On benchmarks, there's a ~5 month lag for open models. So we should expect next round of open models to be near Fable-level, at least on paper. If not, then perhaps progress in open models has slowed.
1282
Sam Harsimony @harsimony.bsky.social · 22/09/2026
MiMo-2.6-Pro has reshaped the Pareto frontier of cost vs performance. 1 point behind Sol, over 15x cheaper.
3332
Sam Harsimony @harsimony.bsky.social · 22/09/2026
The main thing the canonical LW story got right is "AI is a big deal". ~All the specifics have been wrong and LW hasn't updated. I don't know why we're deferring to this worldview over others. The track record isn't great.
1121
Sam Harsimony @harsimony.bsky.social · 22/09/2026
This is the alignment work I want to see more of! How can you change the training data to change behavior? Neel Nanda and co. posted a series on this in July. Links and summaries in thread. (I'm skipping posts from series that aren't related to this idea) www.alignmentforum.org/posts/wyZRNg...
alignmentforum.org
Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forum
This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and ad…
140
Sam Harsimony @harsimony.bsky.social · 21/09/2026
A 1.5 year follow up on NYC congestion pricing. Congestion pricing works and could solve traffic. shoshanavasserman.com/papers/the_n...
022