Sign in

Active Site

@activesite.bio
20 followers 10 following 12 posts

Measuring frontier AI in synthetic biology: activesite.org

PostsRepliesMedia
Active Site @activesite.bio · 19/02/2026
How good were participants at using LLMs? ~40% of participants never uploaded images to LLMs. Interestingly, both arms mentioned YouTube most often as helpful.
120
Active Site @activesite.bio · 19/02/2026
How reliable were LLMs in the hands of novices? LLM transcripts revealed that models can still make mistakes, especially in molecular cloning. LLMs led participants to move quicker (Panel A) but often not with the correct materials (Panel B).
120
Active Site @activesite.bio · 19/02/2026
It's hard to compress all that into a single statistic. But one way is by using a Bayesian model, which suggests LLMs give a ~1.4x boost on a "typical" wet-lab task. Fundamentally, we're confident that there wasn't a large LLM slow-down or speed-up (95% CrI: 0.7x–2.6x).
120
Active Site @activesite.bio · 19/02/2026
But there are some signs LLMs were useful. LLM participants had higher success on 4 out of 5 tasks, most notably in cell culture (69% vs. 55%; P = 0.06). LLM participants also advanced further within a task even if they didn't finish within the study period (odds >80%).
120
Active Site @activesite.bio · 19/02/2026
Our primary outcome: were LLM users more likely to complete all three of the core tasks *together*? Only ~5% of the LLM arm and ~7% of the Internet arm completed all three. No significant difference – and far lower than experts predicted.
120
Active Site @activesite.bio · 19/02/2026
The study was the largest and longest of its kind: 153 participants with minimal lab experience over 8 weeks – randomized to LLM and Internet-only. They tried 5 laboratory tasks, 3 of which are central to a viral reverse genetics workflow. No protocols given — just an objective.
120
Active Site @activesite.bio · 19/02/2026
We ran a randomized controlled trial to see if LLMs can help novices perform molecular biology in a wet-lab. The results: LLMs may help in some aspects, but we found no significant increase at the core tasks end-to-end. That's lower than what experts predicted. Our findings 🧵
1225