Kevin Jablonka @kjablonka.com · 23/04/2026The direction I find most actionable: we need to train LLMs differently for science. Generic post-training doesn't transfer to disciplined inquiry. Corral is built as a training substrate — each environment can be used to score over reasoning trajectories, not just outputs. 220
Kevin Jablonka @kjablonka.com · 23/04/2026A word on Corral itself. Real scientific infrastructure, careful taxonomy, step-by-step reasoning annotation — a lot of craft went in from the team. The engineering here is not small. 010
Kevin Jablonka @kjablonka.com · 23/04/2026And the same reasoning mode appears whether the task is "follow a known procedure" or "form hypotheses under uncertainty." Agents don't change how they reason based on what the task demands. 000
Kevin Jablonka @kjablonka.com · 23/04/2026Can scaffolding rescue this? We injected successful prior reasoning directly into agents' conversation history. Workflow tasks: yes, early steps help. Hypothesis-driven tasks: no, even near-complete trajectories barely help. 000
Kevin Jablonka @kjablonka.com · 23/04/2026Second finding: we annotated the reasoning process of every run. Evidence gathered but unused: 68% of traces Belief revised on contradictory evidence: 26% Convergent multi-test evidence: 7% Testing, refutation, triangulation: rare. 000
Kevin Jablonka @kjablonka.com · 23/04/2026First finding: the base LLM drives 41% of performance variance. The scaffold drives 1.5%. If you build agents, this is where the leverage is. Scaffold engineering is not the bottleneck. 410
Kevin Jablonka @kjablonka.com · 23/04/2026we built Corral, 8 scientific environments with real infrastructure — live AFM, LAMMPS, NMR, wet-lab chemistry, retrosynthesis, ML pipelines, circuit inference, surface construction. Every run annotated step-by-step for reasoning 100
Kevin Jablonka @kjablonka.com · 23/04/2026When an "AI scientist" produces a result, is that knowledge? Philosophers I taught last semester kept reminding me: knowledge is justified true belief. The process matters. New preprint: "AI scientists produce results without reasoning scientifically." 25,000+ runs. 251
Kevin Jablonka @kjablonka.com · 17/03/2025Our team has been spending a lot of time building evaluations for machine learning systems. We have learned some lessons and wrote them down arxiv.org/abs/2503.10837 160
Kevin Jablonka @kjablonka.com · 24/11/2024I'm learning materials according to mine. Thanks, @morinryan.bsky.social for creating this service! 170
Kevin Jablonka @kjablonka.com · 23/11/2024With this, he not only obtains SOTA performance, but he could also reproduce corrections of incorrectly assigned structures in the literature, i.e., spot mistakes. 150
Kevin Jablonka @kjablonka.com · 23/11/2024However, chemists like creating new compounds that might not be in PubChem. Thus, he extended his system with a GA that evolves molecules to best match chemical structures. 140
Kevin Jablonka @kjablonka.com · 23/11/2024In the first step, he aligns encoders for molecules and spectra using contrastive learning. Based on this, he can retrieve compounds for any combination of spectra. In practice, he searches PubChem using spectral embeddings. 140
Kevin Jablonka @kjablonka.com · 23/11/2024Chemists often combine many different techniques to elucidate structures. Adrian has been building a system that mimics this using models and genetic algorithms. 1313