Andrew White 🐦⬛ @andrew.diffuse.one · 29/10/2025Making "AI Scientists" has become a hot topic lately. The first reference I could find was from 2008. The term has been used for 20 years! Like "Adam," an AI Scientist robot for studying yeast was published in 2009. I wrote a short post about the term and what it means now. diffuse.one/p/w1-001 082
Andrew White 🐦⬛ @andrew.diffuse.one · 26/09/2025It sounds insane, but remember there are 10^14 atoms in a human cell and 10^20 femtoseconds in a day. And across multiple simulation engines, it requires 10^4 FLOPs per atom x femtosecond 2/3 110
Andrew White 🐦⬛ @andrew.diffuse.one · 26/09/2025I finished my estimate on required compute to make an atomic-resolution virtual cell: 10^38 FLOPs to simulate a human cell for 1 day. We should be able to do this simulation in 2074 using 200 TW of power. 1/3 4123
Andrew White 🐦⬛ @andrew.diffuse.one · 19/09/2025Our ether0 paper was accepted at NeurIPS 2025! Very proud of the FutureHouse team! 160
Andrew White 🐦⬛ @andrew.diffuse.one · 14/09/2025You can also look at it over time. Here's relatively popularity of different animal models in research over time. Anyway, found this to be interesting. More details about it here: diffuse.one/p/d2-003 3/3 130
Andrew White 🐦⬛ @andrew.diffuse.one · 14/09/2025Here's one measuring the frequency of sample sizes. Like how often people use 8 samples vs 12 samples for reporting research results. N=2 is apparently the most popular 2/3 230
Andrew White 🐦⬛ @andrew.diffuse.one · 14/09/2025Google scholar has a full-text index of nearly all research papers. You can use it to get counts for arbitrary phrases. I've been using this to measure popularity of things in science. For example, here's the popularity of Greek letters used in equations 1/3 3131
Andrew White 🐦⬛ @andrew.diffuse.one · 15/08/2025I've written up some thoughts on publishing for machines. 10M research papers are published per year and there are 227M total - machines will be primary producers and readers of publications going forward. Humans can simply not keep up. It's time to think about revising the scientific paper. 110
Andrew White 🐦⬛ @andrew.diffuse.one · 23/07/2025It’s a clever question. But it’s not really about frontier science. Multiple papers have shown that Oganesson is not a gas (it’s predicted to be semiconducting solid), it’s not noble (it’s reactive), and it isn’t included in any "terrestrial matter" tables of noble gases. 3/7 100
Andrew White 🐦⬛ @andrew.diffuse.one · 23/07/2025The design process of HLE required the questions to be unanswerable by contemporary LLMs. That lead to many gotcha style questions like the one below. It’s a trick question – in 2002, a few atoms of a group 18 element Oganesson were made for a few milliseconds. 2/7 110
Andrew White 🐦⬛ @andrew.diffuse.one · 22/06/2025I have written up a 3.5k word/10 figure essay on how to write a reward function while avoiding reward hacking for chemistry. It covers all the ridiculous ways we had to avoid reward hacking for training ether0, our scientific reasoning model. diffuse.one/p/m1-000 2231
Andrew White 🐦⬛ @andrew.diffuse.one · 20/05/2025The code for this is really minimal - similar to Google Co-Scientist we used multiple agents (from our platform in this case) and tournament-style rankings to select ideas. We're open sourcing it next week, along with all the trajectories. 120
Andrew White 🐦⬛ @andrew.diffuse.one · 20/05/2025The figures, hypothesis, original and follow-up experiments were all generated from our agents. Interestingly, only the lab-work and the paper writing were not automated (which is the opposite of what I would have predicted 2 years ago). 110
Andrew White 🐦⬛ @andrew.diffuse.one · 20/05/2025FutureHouse's goal has been to automate scientific discovery. Now we used our agents to make a genuine discovery – a potential new treatment for one kind of blindness (dAMD). We had multiple cycles of hypotheses, experiments, and data analysis – including identify the mechanism. 1245
Andrew White 🐦⬛ @andrew.diffuse.one · 13/05/2025We shipped multi-agents today! Our chemistry design agent can now call Crow, our scholarly research agents, to bring in data from literature/clinical trials/open targets while designing molecules. platform.futurehouse.org 0112
Andrew White 🐦⬛ @andrew.diffuse.one · 11/05/2025Integrating @opentargets.org is so helpful to provide evidence for disease mechanisms independent of the literature. Here's a demo of synthesizing 78 papers and open targets to propose two novel targets for triple negative breast cancer See the answer: platform.futurehouse.org/trajectories... 060
Andrew White 🐦⬛ @andrew.diffuse.one · 09/05/2025We have an API for clinical trials on our platform - which means you can ask questions like "what trials will read out in June for NSCLC and how likely would you rate their success based on previous trials in the area." Pretty cool. Answer: platform.futurehouse.org/trajectories... 030
Andrew White 🐦⬛ @andrew.diffuse.one · 06/05/2025Here's a command that converts a DOI to bibtex: 1174
Andrew White 🐦⬛ @andrew.diffuse.one · 01/05/2025The plan at FutureHouse has been to build scientific agents for discoveries. We’ve spent the last year researching the best way to make agents. We’ve made a ton of progress and now we’ve engineered them to be used at scale, by anyone. Free and on API. 1133
Andrew White 🐦⬛ @andrew.diffuse.one · 08/03/2025And if you want all the functional groups: I would actually love to have someone explain what the correct answer for this molecule. 100
Andrew White 🐦⬛ @andrew.diffuse.one · 08/03/2025It's ridiculous, but there hasn't existed a one-liner to quickly get functional groups of a molecule. Little Friday night coding exercise to get this working. Enjoy - and let me know of any missing functional groups! I could only do a few hundred. 3295
Andrew White 🐦⬛ @andrew.diffuse.one · 04/03/2025Half of an AI scientist is rejecting or accepting hypotheses. FutureHouse and Science Machines just put out ~300 novel hypotheses from ~50 published papers along with ground-truth data. Humans take 4.2 hours to solve these and frontier models get 10-20% correct. This is like SWE-bench for comp bio 1110
Andrew White 🐦⬛ @andrew.diffuse.one · 25/02/2025PaperQA2 can now work with clinical trials. It considers both research papers and clinical trials jointly to answer complex questions. It uses the the clinicial trials dot gov API - so it can do complex queries too. Checkout the tutorial below: futurehouse.gitbook.io/futurehouse-... 030
Andrew White 🐦⬛ @andrew.diffuse.one · 22/02/2025It's been about a month since the first batch of reasoning models was released. There’s been about a dozen reproductions since then and some patterns are emerging. I’ve written up my own notes on training recipes, frameworks, rumors, and major open questions. diffuse.one/p/d2-000 060
Andrew White 🐦⬛ @andrew.diffuse.one · 19/02/2025This is still an early topic and my work is very preliminary, but I think we may be able to start auditing scientific literature at scale. I’ve written up a lot of thoughts, background, and analysis in a blog post: diffuse.one/p/d1-008 4/4 110
Andrew White 🐦⬛ @andrew.diffuse.one · 19/02/2025What is a universal way to check for signs of fraud in a paper? I investigated faithfulness of citations – are citations consistent with cited sources, are they irrelevant? This does significantly correlate with if a paper is subsequently retracted 3/4 100
Andrew White 🐦⬛ @andrew.diffuse.one · 19/02/2025Image duplication has been a powerful signal for detecting scientific fraud, but is irrelevant in many fields. I've been working a bit on finding new signals like it that work across fields. I've found one using LLMs that can predict retractions, weakly, for $1 per paper. 1/4 263
Andrew White 🐦⬛ @andrew.diffuse.one · 14/02/2025We found that custom built tools are better than just a python REPL, that llama-405B is a great open source model for this, and weaker models require carefully worded instructions. We checked for functionally getting a simulation set up and the correctness of the choices 100
Andrew White 🐦⬛ @andrew.diffuse.one · 14/02/2025Models do pretty well on the tasks – with GPT-4o getting 72% and Llama-3.1 405B getting 68%. Some models, like Claude Sonnet, would do better but just couldn’t figure out NPT ensembles! 100
Andrew White 🐦⬛ @andrew.diffuse.one · 14/02/2025Work by Quintina Campbell, Sam Cox, Jorge Medina, Brittany Watterson. They wanted to see if agents could handle complex multi-step tasks like fetching a protein input structure, running the simulation and analysis. They built several tasks with varying complexity. 110
Andrew White 🐦⬛ @andrew.diffuse.one · 14/02/2025Molecular dynamics requires a lot of expert knowledge to set-up and analyze simulations. We set out to automate it with LLM agents: MDCrow! 1246
Andrew White 🐦⬛ @andrew.diffuse.one · 03/02/2025Humanity's last exam progress. Looking forward to humanity's last last exam_final 080
Andrew White 🐦⬛ @andrew.diffuse.one · 24/01/2025I'm very impressed with Operator. I've used a lot of web agents before and operator actually can function after dozens of steps, whereas most just die after 4-5. I asked it to find a new way to treat PCOS and it spent 12 minutes (~50 steps) on it. 070
Andrew White 🐦⬛ @andrew.diffuse.one · 18/01/2025Since we first released our RAG agent PaperQA, we've seen steady improvements to match human performance, and we now exceed expert scientists performance by ~25 points on doing literature research. 080
Andrew White 🐦⬛ @andrew.diffuse.one · 31/12/2024More info about it here: Blog: www.futurehouse.org/research-ann... Code: github.com/Future-House... Preprint: arxiv.org/abs/2412.21154 Agents: github.com/future-house/ldp 110
Andrew White 🐦⬛ @andrew.diffuse.one · 31/12/2024We’ve tried to make it as simple as possible for others to create Aviary environments and we plan to release more environments soon! Here’s a small example of how to make an environment: 110
Andrew White 🐦⬛ @andrew.diffuse.one · 31/12/2024The environments in Aviary truly require multiple steps of cycles of observation and action. Here you can see multiple trajectories of how the agent solves problems differently than its demonstrations 110
Andrew White 🐦⬛ @andrew.diffuse.one · 31/12/2024A lot of effort in this work was framing the learning problem of agents. We settled on defining agents using stochastic compute graphs and splitting the environment and agent according to what we want to train. Here are some components of well-known agents as compute graphs: 110
Andrew White 🐦⬛ @andrew.diffuse.one · 31/12/2024Aviary is a gymnasium of new scientific environments. Using behavior cloning, expert iteration, and consensus sampling we’ve trained Llamma-3.1 8B agents to very high accuracy on challenging multi-step tasks. And at low cost! www.futurehouse.org/research-ann... 110
Andrew White 🐦⬛ @andrew.diffuse.one · 31/12/2024Finishing 2024 with one more research result! We’ve trained small language agents to do hard sci tasks: engineering proteins, manipulating DNA, and working with sci literature in a new library called Aviary. We beat humans and frontier LLMs on these tasks! 1305
Andrew White 🐦⬛ @andrew.diffuse.one · 09/12/2024What if Eroom's law, the decreasing productivity of drug discovery, is a result of aggressive accounting whereby pharma companies consider as much stuff as possible as R&D to avoid taxes? I bet eroom's law holds in movies too - where losses are built to equal move revenue 1274
Andrew White 🐦⬛ @andrew.diffuse.one · 03/12/2024Here are the percentages with the smallest expected population size from their posterior - note that the effects come from the prior distribution of observed sample sizes reported in literature and how they factor. 4/5 110
Andrew White 🐦⬛ @andrew.diffuse.one · 03/12/2024This leads to interesting effects, where indeed 66% has a posterior mean population size of that is double of 67% - because 66% can only arise from larger populations than 67%. This assumes there is some bound on population sizes (the effect washes out as numbers become larger) 3/5 100
Andrew White 🐦⬛ @andrew.diffuse.one · 03/12/2024It turns out, you can make predictions. For example, here's the most likely population sizes implied by seeing a percentage of 13%. This arises from how 13 factors with 100, and from the uneven prior distribution of sample sizes reported in literature. 2/5 100
Andrew White 🐦⬛ @andrew.diffuse.one · 03/12/2024Percentages are always thrown around to make something more "scientific." I've always wondered if you can infer something about population size from the value - like if a study reports a percentage of 66%, does that imply a larger population size than one that reports 67%? 1/5 181
Andrew White 🐦⬛ @andrew.diffuse.one · 03/12/2024The frequency of sample sizes in published research is not uniform. Here are fraction of sample sizes across research papers, based on google scholar searches. The most popular sample sizes are 10,50, and 100. 79 is the least popular sample size. 0102
Andrew White 🐦⬛ @andrew.diffuse.one · 22/11/2024 Here's a little drawing of some of the 1400 compounds released in @evebio.bsky.social's first data dump. 5313
Andrew White 🐦⬛ @andrew.diffuse.one · 21/11/2024 In case you were wondering, turkeys can fly like other birds but are not able to use tools like the more intelligent corvids citation: hasanyone.com?id=3db0a837 030
Andrew White 🐦⬛ @andrew.diffuse.one · 18/11/2024Here's a comparison of various chemical biology objects from gold nanoparticles to small molecule drugs, made with Blender. 0497