Sign in

Sam Harsimony

@harsimony.bsky.social
1.8K followers 760 following 2.7K posts

I write about opportunities in science, space, and policy here: splittinginfinity.substack.com

PostsRepliesMedia
Reposted by Sam Harsimony
Seth Karten @sethkarten.ai · 1h
Over two weeks, Prime Agent orchestrated a swarm of over 2,000 agents to rewrite itself in Rust. With 10,000+ sandboxes, 200B+ GLM-5.3 tokens, and 16,000 agent-to-agent messages.. Prime Agent now reaches usable input ~13× faster and uses 83% less startup memory than before.
2375
Sam Harsimony @harsimony.bsky.social · 3h
Many domains are anti-inductive. Once something becomes common knowledge it stops being useful. Humans retain an advantage here, for now.
071
Sam Harsimony @harsimony.bsky.social · 9h
Recent paper on compaction from Tim Dettmer's group. "The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it." arxiv.org/abs/2609.26779
030
Reposted by Sam Harsimony
ToughSF @toughsf.bsky.social · 22h
Did Mars lose all of its water to space, or did it get frozen, buried and absorbed by crustal hydration? 30-99% of the Mars' original oceans could be recovered from the subsurface, equivalent to a global depth of 100-1500 m! static1.squarespace.com/static/5ba27...
1152
Sam Harsimony @harsimony.bsky.social · 09/10/2026
Here's a paper discussing the NSA question at the time: eprint.iacr.org/2015/1018 Not saying ECC is broken. But when LLM's brute force math results that mathematicians overlooked and our cryptography relies on mathematicians not overlooking math results caution is warranted!
020
Sam Harsimony @harsimony.bsky.social · 08/10/2026
So it's more that our current foundation was already looking shaky, and it takes a while to switch to the new stuff. Who knows maybe ECC is fine. But cryptography has a long history of promising schemes getting broken.
110
Sam Harsimony @harsimony.bsky.social · 08/10/2026
Definitely agree that defense will win long term. Post-quantum cryptography will probably work. But ppl have been worried by elliptic curves for a while. NSA moved away years ahead quantum computers, some people think they got spooked by internal results showing ECC is vulnerable.
110
Sam Harsimony @harsimony.bsky.social · 08/10/2026
This is the most consequential thing that might come from LLM math. Note that the NSA switched to post-quantum crypto years ago and Signal app already uses it. But we really need to get the financial system on board.
2174
Sam Harsimony @harsimony.bsky.social · 08/10/2026
Accords with CS job market studies. Still early days, but if these results hold up I think it paints an optimistic picture of what AI will do to jobs. Years long transitions addressing bottlenecks as they arise. Retaining existing hires. Though evidence that jr. hires get squeezed.
150
Sam Harsimony @harsimony.bsky.social · 07/10/2026
If I was Anthropic, I'd wait till new congress gets sworn in so there's less risk of political meddling. And time the training/release of Opus 5.8 right before.
000
Reposted by Sam Harsimony
Epoch AI @epochai.bsky.social · 07/10/2026
AI developers aim to create an automated AI researcher. How close are they? To find out, we built InnovationEval, which tests whether AI can produce post-training innovations comparable in magnitude to a recently published advance. So far, agents’ results are underwhelming.
2184
Sam Harsimony @harsimony.bsky.social · 07/10/2026
True, remains to be seen. It's hard for me to imagine wanting vastly more compute for a assistant/chat AI, but we might find new uses for lots of cheap tokens (personalized video games made on demand perhaps?)
110
Sam Harsimony @harsimony.bsky.social · 07/10/2026
Interestingly, the actual performance improvement vs. time curve looks pretty smooth.
020
Sam Harsimony @harsimony.bsky.social · 07/10/2026
P-zero Research introduces TasteVal, a benchmark to measure research taste. Opus 5.5 can squeeze 2.3x more performance out of the same compute budget. Trend suggests doubling every 3 months (though authors note caveats). arxiv.org/abs/2610.06824 pzeroresearch.com/work/tasteval/
131
Sam Harsimony @harsimony.bsky.social · 07/10/2026
There are reasons to think it is temporary! bsky.app/profile/hars...
100
Sam Harsimony @harsimony.bsky.social · 07/10/2026
See also this study on the same drug: www.nature.com/articles/s41...
nature.com
Apitegromab for lean mass preservation during tirzepatide-induced weight loss: a randomized, double-blind, placebo-controlled phase 2 trial - Nature Medicine
In the phase 2 EMBRAZE study, participants receiving tirzepatide and apitegromab lost less lean mass compared to participants receiving tirzepatide and placebo.
010
Sam Harsimony @harsimony.bsky.social · 07/10/2026
A myostatin inhibitor has been approved for spinal muscular atrophy. It's a monoclonal antibody so not super scalable. But this drug class is going to be incredible for muscle-wasting diseases, gaining lean-mass on GLP1's, and general health. investors.scholarrock.com/news-release...
110
Sam Harsimony @harsimony.bsky.social · 07/10/2026
Really cool. Nanoimprint might not work out, but it's a very different way to make chips and it's cool to see people trying it.
010
Sam Harsimony @harsimony.bsky.social · 07/10/2026
Wonder if the Chinese labs are waiting for the Anthropic IPO date to release new models.
140
Sam Harsimony @harsimony.bsky.social · 07/10/2026
Notes: 1. Damaging liberalism is a key path for AI to kill us all. The world is safer with distributed ownership, universal rights, and low concentration. 2. I could see a massive shift in how we treat AI's. We tend to love the beings we interact with often. bsky.app/profile/hars...
020
Sam Harsimony @harsimony.bsky.social · 07/10/2026
@sebkrier.com has 20 good takes on AI. "I am more concerned about the gradual decay of our institutions, world order, and liberalism than I am about AI killing everyone; and I think we should be wary of AI advocacy that contributes to this." x.com/sebkrier/sta...
130
Sam Harsimony @harsimony.bsky.social · 07/10/2026
Yeah totally, the open model is sort of a loss leader that gets you customers
010
Sam Harsimony @harsimony.bsky.social · 07/10/2026
Related to the amount of compute in the world question: bsky.app/profile/hars...
010
Sam Harsimony @harsimony.bsky.social · 07/10/2026
I feel like AI policy folks haven't reckoned with the fact that the Trump admin is actively hostile to AI policy. If we're looking at 2029 for an international agreement, that's a very different world. Powerful, local, open-source models and potentially a lot more compute out in the world.
251
Sam Harsimony @harsimony.bsky.social · 07/10/2026
Frontier models will persist for big problems ofc. But chat and personal assistant agents will become loss leaders.
050
Sam Harsimony @harsimony.bsky.social · 07/10/2026
With algorithmic progress, people will have the option to buy an extra computer and run a personal AI that is smart enough for their needs.
280
Reposted by Sam Harsimony
Epoch AI @epochai.bsky.social · 07/10/2026
The share of US adults experiencing a cyber incident hasn't grown since Claude Fable 5's release. About 45% reported at least one incident in the past 12 months when we asked in September, compared with 46% in June.
Horizontal bar chart comparing June and September 2026 shares of US adults reporting cyber incidents in the past 12 months. Any incident: 46% in June, 45% in September; each individual incident type is also roughly unchanged.
1162
Reposted by Sam Harsimony
Nathan Lambert @natolambert.bsky.social · 07/10/2026
In case you missed my hot take yesterday. www.interconnects.ai/p/the-cyber-...
516324
Sam Harsimony @harsimony.bsky.social · 06/10/2026
In my experience, people have been proclaiming significant shifts in capabilities for years. But set the threshold for significance wherever you prefer. A good theory should be able to explain why there is a threshold at all, and why we haven't observed more or less doom after crossing it.
010
Reposted by Sam Harsimony
Giovanni Monea @giomonea.bsky.social · 06/10/2026
⚠️ Stop pretraining your looped Transformer with a separate KV cache per recursion! In our recent preprint, we show that *memory sharing* not only saves memory but is also a net-positive inductive bias! Less memory, same flops, higher quality. 🔗 arxiv.org/abs/2610.02383 🧵
17012
Sam Harsimony @harsimony.bsky.social · 06/10/2026
If we think more capabilities means more doom and we observe more capabilities but not more doom, that is still an update away from doom, even for 2045 timelines.
010
Sam Harsimony @harsimony.bsky.social · 06/10/2026
GPT-3 is over 6 years old! Any theory that predicts future should also be able to explain the past. Scaling has been steady and ~safe. Why? and why might that change?
130
Sam Harsimony @harsimony.bsky.social · 06/10/2026
As I understand, open-weights lowers closed-weight profits, discouraging closed-weight investment and encouraging open-weight investment. Reduces the gap between closed and open so open-weight can take a larger share of profits (while reducing the entire industry's profit?)
000
Sam Harsimony @harsimony.bsky.social · 06/10/2026
@norvid-studies.bsky.social an econ model of why it might make sense to release open weight. drive.google.com/file/d/1G8sv... x.com/bryantxia22/...
drive.google.com
writeup.pdf
100
Sam Harsimony @harsimony.bsky.social · 06/10/2026
Important to ask why we haven't seen more bad stuff. Why hasn't AI foom-ed? Why aren't there more accidents? Where's the unemployment? Will those barriers continue to hold?
1180
Sam Harsimony @harsimony.bsky.social · 06/10/2026
Agreed. It's hard to pinpoint what's missing and how we teach it to the models. I suspect the confusion is inside ourselves. If we knew what we truly wanted and what was possible it would be trivial to ask.
020
Reposted by Sam Harsimony
Nathan Lambert @natolambert.bsky.social · 06/10/2026
Too many people are analyzing the cyber risks of open models in a narrow lens, assuming China doesn't care about safety, and fear mongering about private information on open-weight cyber attacks. The public information we have paints a very different picture. www.interconnects.ai/p/the-cyber-...
interconnects.ai
The Cyber Risk Discourse is Broken
Open-weights, ideology, and acknowledging trade-offs.
0235
Reposted by Sam Harsimony
Sung Kim @sungkim.bsky.social · 05/10/2026
Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation. Full paper: qlabs.sh/research/dust Code: github.com/qlabs-eng/dust
27111
Reposted by Sam Harsimony
Sung Kim @sungkim.bsky.social · 05/10/2026
Reflection AI's Beam: a highly efficient agentic open model with 501B total parameters and 23B active. - Frontier reasoning efficiency - Advances the Western open frontier on coding & agentic tasks - Trained end-to-end from scratch Full weights release this month. reflection.ai/blog/introdu...
4544
Sam Harsimony @harsimony.bsky.social · 05/10/2026
For context: 1000 median researchers would surpass world GDP in summer of 2028.
1201
Reposted by Sam Harsimony
neurosock @neurosock.bsky.social · 05/10/2026
I reverse engineered Neuralink's new decoder arch. out of their blog post. As usual, they used the HARD work of other labs, and share nothing about the internals. Cortical activity is converted to bags of "words", producing a low-dimensional representation ... 🧵 #neuroskyence #compneuro #NeuroAI
1153
Sam Harsimony @harsimony.bsky.social · 05/10/2026
Partial RSI has been happening for a while.
062
Sam Harsimony @harsimony.bsky.social · 05/10/2026
If I spent my time quote-dunking creationists, screenshotting their replies, making fun of them, etc. you would be right to be annoyed at me. Creationists will continue to exist, but everyone thinks they're silly. They don't matter. Please stop talking about AI deniers.
1263
Sam Harsimony @harsimony.bsky.social · 05/10/2026
In practice open weight model inference providers have opted for the latter. They run quantized versions or shrink context windows and only slightly undercut on cost. Which makes sense. It's risky to invest so much into one (soon-to-be obsoleted) model as a small company. Better to stay nimble.
011
Sam Harsimony @harsimony.bsky.social · 05/10/2026
For ex, you're serving Kimi K3, but Kimi team has lots of secret tricks to serve it cheaper and get it to be smarter. You could spend lots of money trying to beat them at their own game. Or just not do that. Just serve a bunch of open models at lower quality.
120
Sam Harsimony @harsimony.bsky.social · 05/10/2026
The licenses matter too, but you also have to invest time/money into optimizing the server configuration to get good performance and you're competing against the inventors.
110
Sam Harsimony @harsimony.bsky.social · 05/10/2026
Good explainer of Context-Language models by @jbhuang0604.bsky.social www.youtube.com/watch?v=aYGN...
youtube.com
What If an AI Could Edit Its Own Memory?
YouTube video by Jia-Bin Huang
030
Sam Harsimony @harsimony.bsky.social · 05/10/2026
If the model is large, then "open weight" is a bit of a farce because your company is the only one who will bother to run it (and optimize serving config). Otherwise I see open weights as sort of a loss-leader to get customers in other ways. But open question whether it will work long term.
230
Reposted by Sam Harsimony
Stella Biderman @stellaathena.bsky.social · 03/10/2026
I’m starting a blog! My first post is on how 3rd party embedded evaluators seem totally unsuited to addressing the problems we are currently facing, and what the real problem is. stellabiderman.ai/blog/embedde...
stellabiderman.ai
Embedded Evaluators Can’t Fix Companies That Choose to Be Bad — Stella Biderman
Embedded evaluators can report violations, but they cannot fix AI companies that knowingly disregard basic cybersecurity and safety practices.
37817
Sam Harsimony @harsimony.bsky.social · 03/10/2026
Towards the end she spits for no reason it made me laugh. Overall very good.
000