Sign in

Gal Sapir

@sapir.bsky.social
80 followers 425 following 369 posts

sparsethought.com

PostsRepliesMedia
Gal Sapir @sapir.bsky.social · 28/09/2026
so natural yet crazy to think that transplanted organs match their "age" (basically methylation patterns++) to recepient's
100
Gal Sapir @sapir.bsky.social · 27/09/2026
engineering is science for impatient people is great! i think also these days CS is science for impatient scientists as well
000
Gal Sapir @sapir.bsky.social · 25/09/2026
yeah I dislike this phrasing too. In the "all models are wrong, some are useful" universe, I think it leans less towards "useful" when applied like that
010
Gal Sapir @sapir.bsky.social · 23/09/2026
domain-specific harnesses are all the rage (or so it sometimes seems) wrote a bit about my own experience working on one tldr - i think its a pretty effective way to try and 'encode' expert knowledge and save time+cost when working on projects in specific areas models aren't trained on
100
Gal Sapir @sapir.bsky.social · 19/09/2026
this is a *great* explanation! thanks @ed3d.net
010
Gal Sapir @sapir.bsky.social · 16/09/2026
this is the best exchange I've read this whole week!
010
Gal Sapir @sapir.bsky.social · 16/09/2026
can i just say that Google calling for clinical trials wasn't on my bingo card in 2026? an interesting and imo smart take (still need to write a bit more about it)
Prospective evidence for conversational medical AI is hard, but non-negotiable
100
Gal Sapir @sapir.bsky.social · 14/09/2026
a really nice paper they articulate an important (emerging) gap really well - the gap between 'human comprehension bar' and 'agent execution bar' another important aspect they touch on is how to encode the human judgement that goes into the scientific process? super importance stuff
The Last Human-Written Paper:
Agent-Native Research Artifacts
100
Gal Sapir @sapir.bsky.social · 11/09/2026
recommended reading!
110
Gal Sapir @sapir.bsky.social · 10/09/2026
2026- one model release at a time www.nber.org/papers/w21788
000
Gal Sapir @sapir.bsky.social · 09/09/2026
sharing my latest work: what can a deeply phenotyped human cohort tell us, and what can today’s AI models do with that information? we built PhenoBench: 90 tasks across 15 clinical domains, using clinical, imaging, molecular, and wearable data from the Human Phenotype Project.
200
Gal Sapir @sapir.bsky.social · 08/09/2026
i spent some time with gpt-6 and fable 5.1, then looked at their model cards to see what improved in health. incremental gains, and some questions about the judges wrote about it here - sparsethought.com/2026/09/08/l...
sparsethought.com
llms, the jagged frontier, and health
a look at health benchmarks in the gpt-6 astra and fable 5.1 model cards, incremental gains, and questions about the judges.
001
Gal Sapir @sapir.bsky.social · 07/09/2026
can i ask experienced researchers here to review a cv please? looking for my next adventure (god i hate linkedin talk), more details soon :-) galsapir.com/cv/
galsapir.com
Gal Sapir · Research CV
Research CV of Gal Sapir, Staff Research Scientist at Pheno.AI, working on health AI, foundation models, clinical evaluation, and research systems.
220
Gal Sapir @sapir.bsky.social · 06/09/2026
that's a good take :-)
010
Gal Sapir @sapir.bsky.social · 05/09/2026
the most accurate take
010
Gal Sapir @sapir.bsky.social · 05/09/2026
you know how they (Plutarch) say "the mind is not a vessel to be filled, but a fire to be kindled"? i feel like it describes working with modern llm/harness combos very well
000
Gal Sapir @sapir.bsky.social · 05/09/2026
first very preliminary impressions from gpt 6 astra - being very similar to 5.6 on scientific/biology texts ml-ish code kinda like the benchmarks suggested
100
Gal Sapir @sapir.bsky.social · 03/09/2026
btw notice the difference, to get to science and health you have to scroll wayyy down (and the improvements are really incremental it seems). Can't wait to test it myself though
000
Gal Sapir @sapir.bsky.social · 02/09/2026
I honestly think it's crazy (in a good way) that "scientific research" is front and center for a model release in a frontier lab
100
Gal Sapir @sapir.bsky.social · 19/08/2026
something slightly different (i usually don't to personal / wlb/ touch grass / future of work type stuff) - but this is different no agents after 8 pm sparsethought.com/2026/08/19/n...
sparsethought.com
no agents after 8pm
on agents making it easier to keep working after my judgment starts to fade, and a simple rule for stopping.
000
Reposted by Gal Sapir
mr. TIM @timkellogg.me · 19/08/2026
yes. yessss. moarrrr arxiv.org/pdf/2608.09703 huggingface.co/nthngdy/matr...
Diagram comparing a "Standard suite" of AI models to a "Matryoshka suite" of nested models. On the left, the Standard suite shows five standalone, independent model blocks of varying parameter sizes: 0.8B, 2B, 4B, 9B, and 27B. A disclaimer at the bottom states, "Disclaimer: we do not train a 27B suite." On the right, the Matryoshka suite shows a nested structure where larger models build directly upon smaller ones: a 0.8B block combines with 1.2B to form a 2B model, which adds 2B to form a 4B model, adds 5B to form a 9B model, and adds 18B to form a 27B model. A green box at the bottom right highlights "-37% total training compute."
2664
Gal Sapir @sapir.bsky.social · 12/08/2026
i agree and not sure how to handle it
010
Gal Sapir @sapir.bsky.social · 10/08/2026
I love their reports. maybe someday I'll subscribe newsletter.semianalysis.com/p/rl-environ...
newsletter.semianalysis.com
RL Environments and RL for Science: Data Foundries and Multi-Agent Architectures
Worker Automation, RL as a Service, Anthropic's next big bet, GDPval and Utility Evals, Computer Use Agents, LLMs in Biology, Mid-Training, Lab Procurement Patterns, Platform Politics and Access
000
Gal Sapir @sapir.bsky.social · 04/08/2026
www.seangoedecke.com/llms-reward-... that's a origiod take (even if well known, it's phrased nicely)
seangoedecke.com
LLMs reward expertise
001
Gal Sapir @sapir.bsky.social · 04/08/2026
a refreshing read! shelbyann.substack.com/p/the-sirens-of-data-sales (wrote some thoughts about it as well) - sparsethought.com/2026/08/04/t...
shelbyann.substack.com
The Sirens of Data Sales
Why selling data to train models is not as easy a decision as you think for bio companies
000
Gal Sapir @sapir.bsky.social · 26/07/2026
more and more I feel like the days resulting in my best work, are the days in which I hold back on "letting agents roam free" and being much more careful and well, full of care
000
Gal Sapir @sapir.bsky.social · 24/07/2026
"large software projects have never been limited only by how quickly an individual can produce code. They are limited by how well people can coordinate their understanding of the system they are changing." lucumr.pocoo.org/2026/7/13/th...
lucumr.pocoo.org
The Tower Keeps Rising
Vibecoding and the possible collapse of a shared language.
000
Gal Sapir @sapir.bsky.social · 22/07/2026
always slightly funny to me when they underestimate themselves these agents
010
Gal Sapir @sapir.bsky.social · 21/07/2026
recommended reading: "HEARTS: benchmarking LLM reasoning on health time series" some non-intuitive findings, some intuitive ones i'm glad to see thoroughly checked yang-ai-lab.github.io/HEARTS/#home
yang-ai-lab.github.io
HEARTS — Health Reasoning over Time Series
A unified benchmark for evaluating hierarchical reasoning capabilities of LLMs over general health time series.
110
Gal Sapir @sapir.bsky.social · 20/07/2026
painfully true, find myself reaching for pen and paper just for that blank page vibe recently when I need to organise my thoughts around something
031
Gal Sapir @sapir.bsky.social · 13/07/2026
did people use "blast radius" in software before coding agents? i really don't think i heard it anytime pre-2023
101
Gal Sapir @sapir.bsky.social · 06/07/2026
anyone else getting weird cryptic 'internal thoughts' from codex as part of the output?
example of such internal thoughts
010
Gal Sapir @sapir.bsky.social · 03/07/2026
benchmarking is becoming a form of data activation. for medical/bio data, the hard part is not just having the data. it is turning messy traces into tasks with verifiable outcomes, so models can be measured and improved against them. sparsethought.com/2026/07/03/b...
sparsethought.com
benchmarking is the new data activation
benchmarks as a way to turn messy domain data into a measurable, optimizable substrate for models.
000
Gal Sapir @sapir.bsky.social · 25/06/2026
beautiful post! really appreciate the caution and tenderness in which its written. thanks @mitsuhiko.at
000
Gal Sapir @sapir.bsky.social · 24/06/2026
new post: a loose reading list around agents, memory, benchmarks, and the hard-domain question: why making models useful outside language/code is still hard. mostly Mitchell Hashimoto, QuestBench, SpatialBench-Long, memory systems, and health/biology data. sparsethought.com/2026/06/24/w...
sparsethought.com
what i’ve read lately
a loose reading list around agents, memory, benchmarks, and the hard parts of making models useful outside language and code.
000
Gal Sapir @sapir.bsky.social · 17/06/2026
really thoughtful and interesting thread on an important question
010
Gal Sapir @sapir.bsky.social · 14/06/2026
1/ a nature medicine paper claiming general-purpose llms beat specialized clinical tools (openevidence, uptodate) is going around. i pushed on it and landed somewhere different. what does it actually measure?
General-purpose large language models outperform specialized clinical AI tools on medical benchmarks
100
Gal Sapir @sapir.bsky.social · 12/06/2026
been running a context map (the PEEK paper's idea: a small, budgeted artifact that turns an agent's traces into orientation) on my personal memory system for two weeks. the paper assumes a fixed territory, a repo or corpus. mine isn't. notes from use - sparsethought.com/maps-of-cont...
sparsethought.com
Redirecting…
000
Gal Sapir @sapir.bsky.social · 07/06/2026
i feel like this is still very true today, esp. in large projects or after a ~week of changes i have to stop and make sure my mental model ("theory of the program") still holds.
020
Gal Sapir @sapir.bsky.social · 04/06/2026
us, every day
010
Gal Sapir @sapir.bsky.social · 27/05/2026
really interesting / cool stuff are happening while testing fugu by @sakanaai.bsky.social ! (can't really share details at this point, but its a solid beta)
screenshot of the console showing a lot of token usage today
020
Gal Sapir @sapir.bsky.social · 25/05/2026
such a great read! like most things out of @lateinteraction.bsky.social 's lab might write a bit about that later zhuohangu.github.io/blog-post-pe...
zhuohangu.github.io
PEEK: Give Your Agent an Orientation Cache
We introduce PEEK, a system that caches reusable orientation knowledge about a recurring external context as a small, prompt-resident context map.
120
Gal Sapir @sapir.bsky.social · 23/05/2026
personal best
screenshot showing codex ran for >3 hours
010
Gal Sapir @sapir.bsky.social · 23/05/2026
Great read from @sakanaai.bsky.social sakana.ai/dgm/
sakana.ai
Sakana AI
The Darwin Gödel Machine: AI that improves itself by rewriting its own code
000
Gal Sapir @sapir.bsky.social · 20/05/2026
small skill: for long-horizon agent runs, have it keep an implementation-notes file as it goes (decisions, tradeoffs, deviations from spec). skim after to check broad strokes + project conventions held. usually enough for me to be ok with it. github.com/galsapir/skills/tree/main/skills/long-horizon
github.com
skills/skills/long-horizon at main · galsapir/skills
Claude Code plugin: deep project interview command that produces actionable specs before implementation - galsapir/skills
010
Gal Sapir @sapir.bsky.social · 17/05/2026
medmarks v1.0 dropped last week, the largest open medical eval suite to date, and the verifiers framing is the right call imo. sitting with it for a few days pulled out a thread that hasn't resolved: the curation problem goes all the way down sparsethought.com/2026/05/16/curation-all-the-way-down/
sparsethought.com
curation all the way down: on clinical AI benchmarks
the curation regression, the openness trade-off, and what a substrate worth evaluating against would actually need: on Medmarks.
100
Gal Sapir @sapir.bsky.social · 15/05/2026
I'm on @science.org ! (eLetter, but still)
My eLetter
100
Gal Sapir @sapir.bsky.social · 08/05/2026
another great piece by @goodfire.bsky.social
000
Gal Sapir @sapir.bsky.social · 07/05/2026
new post: small workflow changes that have started to add up. nothing deep, a few adjustments to how i work with agents that converged into something that feels more comfortable to me
100
Gal Sapir @sapir.bsky.social · 04/05/2026
the Brodeur et al. Science paper has been making the rounds on bluesky the past few days, which is how i ran into it. didn't quite click for me at first, so i sat with it for a day and wrote some thoughts down. cc @adamrodmanmd.bsky.social mrodmanmd.bsky.social
100