Sign in

Gal Sapir

@sapir.bsky.social
80 followers 425 following 369 posts

sparsethought.com

PostsRepliesMedia
Gal Sapir @sapir.bsky.social · 28/09/2026
www.biorxiv.org/content/10.6...
biorxiv.org
000
Gal Sapir @sapir.bsky.social · 28/09/2026
so natural yet crazy to think that transplanted organs match their "age" (basically methylation patterns++) to recepient's
100
Gal Sapir @sapir.bsky.social · 27/09/2026
engineering is science for impatient people is great! i think also these days CS is science for impatient scientists as well
000
Gal Sapir @sapir.bsky.social · 25/09/2026
yeah I dislike this phrasing too. In the "all models are wrong, some are useful" universe, I think it leans less towards "useful" when applied like that
010
Gal Sapir @sapir.bsky.social · 24/09/2026
I also that think the "pays off" in many other fields is only known ~years after discoveries are made
010
Gal Sapir @sapir.bsky.social · 24/09/2026
As usual Ted thanks for this take! Really prompted me to get my thoughts out :-)
110
Gal Sapir @sapir.bsky.social · 24/09/2026
I'm not saying "it's futile stop trying" I am saying (quoting a friend from medical school here) - sometimes people like to say "medicine/biology isn't an exact science" Well, it is. We just don't know all the rules is all, so the bitter lesson here still has a few prerequisites before it applies
110
Gal Sapir @sapir.bsky.social · 24/09/2026
I suspect that stuff that seems like breakthroughs in biology in the data-level, many times can't/won't be translated to the cell->mouse ->primate->human levels Idk about other fields, but I think it's something about how the messy real world is *mostly* not like the neat textbooks of mathematics
110
Gal Sapir @sapir.bsky.social · 23/09/2026
sparsethought.com/2026/09/23/h... might be of interest to @eugenevinitsky.bsky.social
sparsethought.com
how do you design the right domain-specific harness?
what belongs in a research harness when the models keep improving, and how do we get the expertise into it?
000
Gal Sapir @sapir.bsky.social · 23/09/2026
domain-specific harnesses are all the rage (or so it sometimes seems) wrote a bit about my own experience working on one tldr - i think its a pretty effective way to try and 'encode' expert knowledge and save time+cost when working on projects in specific areas models aren't trained on
100
Gal Sapir @sapir.bsky.social · 23/09/2026
this is inspiring.
010
Gal Sapir @sapir.bsky.social · 19/09/2026
this is a *great* explanation! thanks @ed3d.net
010
Gal Sapir @sapir.bsky.social · 19/09/2026
A. Grant id great. B. Great take! 100% agree.
010
Gal Sapir @sapir.bsky.social · 18/09/2026
I think that the examples of this happening so quickly in a new field are few and far between, mainly because it's really hard to create good benchmark in complex fields. But I get you point :-)
000
Gal Sapir @sapir.bsky.social · 16/09/2026
this is the best exchange I've read this whole week!
010
Gal Sapir @sapir.bsky.social · 16/09/2026
"Trust is not a fixed property of a model and cannot be measured as a benchmark. It is earned incrementally through direct human experience with the AI agent and rigorous scientific analysis of the study results. Only prospective studies can observe this dynamic."
000
Gal Sapir @sapir.bsky.social · 16/09/2026
"PCPs reported that AMIE shifted consultations “from data gathering to data verification”; patients arrived organized with coherent narratives, enabling more collaborative conversations and shared decision-making." i think this is really a first step
000
Gal Sapir @sapir.bsky.social · 16/09/2026
 We argue that prospective, clinical-grade evidence generation, despite being more difficult than benchmarking studies, should become a standard for conversational AI in health and medicine.
200
Gal Sapir @sapir.bsky.social · 16/09/2026
can i just say that Google calling for clinical trials wasn't on my bingo card in 2026? an interesting and imo smart take (still need to write a bit more about it)
Prospective evidence for conversational medical AI is hard, but non-negotiable
100
Gal Sapir @sapir.bsky.social · 14/09/2026
arxiv.org/abs/2604.24658 @elenal3ai.bsky.social @chenyuyou.bsky.social
arxiv.org
The Last Human-Written Paper: Agent-Native Research Artifacts
Scientific publication compresses a branching, iterative research process into a linear narrative, discarding the majority of what was discovered along the way. This compilation imposes two structural...
000
Gal Sapir @sapir.bsky.social · 14/09/2026
a really nice paper they articulate an important (emerging) gap really well - the gap between 'human comprehension bar' and 'agent execution bar' another important aspect they touch on is how to encode the human judgement that goes into the scientific process? super importance stuff
The Last Human-Written Paper:
Agent-Native Research Artifacts
100
Gal Sapir @sapir.bsky.social · 11/09/2026
010
Gal Sapir @sapir.bsky.social · 11/09/2026
another banger "Dialectical activities ... are valuable because their good can only be discovered from within the activity, as in parenting, art-making, friendship, or research [56]. Scholarly sensemaking is similarly dialectical: researchers develop ... become better researchers by orienteering..."
110
Gal Sapir @sapir.bsky.social · 11/09/2026
favorite quote so far: "We call our alternative vision reading for transformation: reading augmentation that treats the reader’s transformation as its central design concern." also
another quote i liked
110
Gal Sapir @sapir.bsky.social · 11/09/2026
recommended reading!
110
Gal Sapir @sapir.bsky.social · 11/09/2026
hey, thanks you for sharing this! half way through the paper itself and I'm learning a lot (from creating this 'reading') helpful as a reader + a writer of scientific texts
030
Gal Sapir @sapir.bsky.social · 10/09/2026
2026- one model release at a time www.nber.org/papers/w21788
000
Gal Sapir @sapir.bsky.social · 10/09/2026
I wonder if people always felt this way and not just us living through COVID, the wars, climate changes and now this
010
Gal Sapir @sapir.bsky.social · 09/09/2026
might be of interest to @lampinen.bsky.social @adamgayoso.bsky.social and others studying model capabilities and biological ML. the LLM prompts, schemas, and synthetic examples are public too.
000
Gal Sapir @sapir.bsky.social · 09/09/2026
tabular foundation models showed small gains over ridge. the 14 LLMs showed a jagged frontier across tasks, generally behind models trained on the same inputs. paper: arxiv.org/abs/2609.06080 more on what we found: sparsethought.com/phenobench/
arxiv.org
PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us
Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. This breadth can reveal which measurements inform which health-relate...
000
Gal Sapir @sapir.bsky.social · 09/09/2026
sharing my latest work: what can a deeply phenotyped human cohort tell us, and what can today’s AI models do with that information? we built PhenoBench: 90 tasks across 15 clinical domains, using clinical, imaging, molecular, and wearable data from the Human Phenotype Project.
200
Gal Sapir @sapir.bsky.social · 08/09/2026
i spent some time with gpt-6 and fable 5.1, then looked at their model cards to see what improved in health. incremental gains, and some questions about the judges wrote about it here - sparsethought.com/2026/09/08/l...
sparsethought.com
llms, the jagged frontier, and health
a look at health benchmarks in the gpt-6 astra and fable 5.1 model cards, incremental gains, and questions about the judges.
001
Gal Sapir @sapir.bsky.social · 08/09/2026
Thanks! It's a great feeling to ask for help from people and get back human, thoughtful and patient answers :-)
000
Gal Sapir @sapir.bsky.social · 08/09/2026
is that a take on "the best argument against democracy is a conversation with the average voter"?
101
Gal Sapir @sapir.bsky.social · 07/09/2026
@mariaa.bsky.social @mariozechner.at @tedunderwood.com
120
Gal Sapir @sapir.bsky.social · 07/09/2026
can i ask experienced researchers here to review a cv please? looking for my next adventure (god i hate linkedin talk), more details soon :-) galsapir.com/cv/
galsapir.com
Gal Sapir · Research CV
Research CV of Gal Sapir, Staff Research Scientist at Pheno.AI, working on health AI, foundation models, clinical evaluation, and research systems.
220
Gal Sapir @sapir.bsky.social · 06/09/2026
that's a good take :-)
010
Gal Sapir @sapir.bsky.social · 06/09/2026
Somewhat off topic - the whole OAI/Huggingface/German wikis incidents is the literal embodiment of SkyNet!
110
Gal Sapir @sapir.bsky.social · 05/09/2026
the most accurate take
010
Gal Sapir @sapir.bsky.social · 05/09/2026
thanks for sharing! can you be more specific with the recommendation? I felt (as a trained physician) that he oversimplified a lot of stuff and made some leaps (it's was very flowing, but closer to sapiens than to the dawn of everything imo)
000
Gal Sapir @sapir.bsky.social · 05/09/2026
hey i've been working on a side project for a burn event that uses a lot of LEDs and raspberry pi thanks to codex :-)
000
Gal Sapir @sapir.bsky.social · 05/09/2026
you know how they (Plutarch) say "the mind is not a vessel to be filled, but a fire to be kindled"? i feel like it describes working with modern llm/harness combos very well
000
Gal Sapir @sapir.bsky.social · 05/09/2026
I wonder if classical swe stuff really are significantly better
000
Gal Sapir @sapir.bsky.social · 05/09/2026
first very preliminary impressions from gpt 6 astra - being very similar to 5.6 on scientific/biology texts ml-ish code kinda like the benchmarks suggested
100
Gal Sapir @sapir.bsky.social · 04/09/2026
I missed the discussion last week, and great to have you here!
010
Gal Sapir @sapir.bsky.social · 04/09/2026
it's also important to remember that the team has really strong priors from their training, and if your tasks is in a domain slightly OOD - you need to be much clearer about how to track progress/define "done"/define "forbidden"
010
Gal Sapir @sapir.bsky.social · 03/09/2026
btw notice the difference, to get to science and health you have to scroll wayyy down (and the improvements are really incremental it seems). Can't wait to test it myself though
000
Gal Sapir @sapir.bsky.social · 02/09/2026
(also, there is so much headroom in that area still)
000
Gal Sapir @sapir.bsky.social · 02/09/2026
I honestly think it's crazy (in a good way) that "scientific research" is front and center for a model release in a frontier lab
100
Gal Sapir @sapir.bsky.social · 27/08/2026
"silently"
010