Sign in

Kevin Jablonka

@kjablonka.com
1.6K followers 421 following 53 posts

Trying to teach computers how to design materials. Leading a research group at FSU Jena/HIPOLE Jena. Increasing entropy since 1996.

PostsRepliesMedia
Kevin Jablonka @kjablonka.com · 23/04/2026
Funding by OpenPhilanthropy, DFG, @helmholtz.de, Google, Merck, Intel, IIT Dehli, @uni-jena.de, Anusandhan National Research Foundation
000
Kevin Jablonka @kjablonka.com · 23/04/2026
This is one of the eight environments built by an amazing team Martiño Ríos García, Nawaf Alampara, Chandan Gupta, Indrajeet Mandal, Sajid Mannan, Ali Asghar Aghajan, Anoop Krishnan
100
Kevin Jablonka @kjablonka.com · 23/04/2026
On simple samples with 3–5 ions, agents often succeed. On harder samples (many ions, overlapping reactivity), they struggle — not because the tools are missing, but because the task demands designing discriminating experiments rather than following a recipe.
100
Kevin Jablonka @kjablonka.com · 23/04/2026
The movie showed our wet lab environment. The environment simulates real chemical equilibria. The agent can add reagents, run flame tests, measure pH, and observe colour and precipitate changes — the same procedure taught in introductory chemistry for identifying unknown ions.
100
Kevin Jablonka @kjablonka.com · 23/04/2026
Preprint: arxiv.org/abs/2604.18805 Website: lamalab-org.github.io/corral/
arxiv.org
AI scientists produce results without reasoning scientifically
Large language model (LLM)-based systems are increasingly deployed to conduct scientific research autonomously, yet whether their reasoning adheres to the epistemic norms that make scientific inquiry ...
020
Kevin Jablonka @kjablonka.com · 23/04/2026
The direction I find most actionable: we need to train LLMs differently for science. Generic post-training doesn't transfer to disciplined inquiry. Corral is built as a training substrate — each environment can be used to score over reasoning trajectories, not just outputs.
220
Kevin Jablonka @kjablonka.com · 23/04/2026
A word on Corral itself. Real scientific infrastructure, careful taxonomy, step-by-step reasoning annotation — a lot of craft went in from the team. The engineering here is not small.
010
Kevin Jablonka @kjablonka.com · 23/04/2026
It connects to a larger question I keep turning over: how far can we actually accelerate science? Fundamental science is a human endeavour. Understanding is what it produces. There is a human velocity to understanding — we can't compress it indefinitely.
210
Kevin Jablonka @kjablonka.com · 23/04/2026
I don't think this means these agents are bad. For workflow tasks they perform well. But in the strict sense — testing, refutation, weighing evidence — what they do isn't the process of science in the way philosophy thought about science.
110
Kevin Jablonka @kjablonka.com · 23/04/2026
And the same reasoning mode appears whether the task is "follow a known procedure" or "form hypotheses under uncertainty." Agents don't change how they reason based on what the task demands.
000
Kevin Jablonka @kjablonka.com · 23/04/2026
Can scaffolding rescue this? We injected successful prior reasoning directly into agents' conversation history. Workflow tasks: yes, early steps help. Hypothesis-driven tasks: no, even near-complete trajectories barely help.
000
Kevin Jablonka @kjablonka.com · 23/04/2026
Second finding: we annotated the reasoning process of every run. Evidence gathered but unused: 68% of traces Belief revised on contradictory evidence: 26% Convergent multi-test evidence: 7% Testing, refutation, triangulation: rare.
000
Kevin Jablonka @kjablonka.com · 23/04/2026
First finding: the base LLM drives 41% of performance variance. The scaffold drives 1.5%. If you build agents, this is where the leverage is. Scaffold engineering is not the bottleneck.
410
Kevin Jablonka @kjablonka.com · 23/04/2026
we built Corral, 8 scientific environments with real infrastructure — live AFM, LAMMPS, NMR, wet-lab chemistry, retrosynthesis, ML pipelines, circuit inference, surface construction. Every run annotated step-by-step for reasoning
100
Kevin Jablonka @kjablonka.com · 23/04/2026
When an "AI scientist" produces a result, is that knowledge? Philosophers I taught last semester kept reminding me: knowledge is justified true belief. The process matters. New preprint: "AI scientists produce results without reasoning scientifically." 25,000+ runs.
251
Kevin Jablonka @kjablonka.com · 26/01/2026
Congrats! 🎊
110
Kevin Jablonka @kjablonka.com · 14/12/2025
Congrats!
010
Reposted by Kevin Jablonka
Manuel Ortuno @chemortuno.bsky.social · 13/12/2025
Congratulations to Dr. Hiep Le for her PhD at @ciqus.bsky.social! I am very proud of her 😁
362
Kevin Jablonka @kjablonka.com · 15/11/2025
congrats!
010
Reposted by Kevin Jablonka
Jorge Bravo Abad @bravo-abad.bsky.social · 21/05/2025
I just published a Substack post highlighting ChemBench by Mirza and coauthors, which evaluates how well LLMs perform on chemistry tasks. From strengths in structural analysis to key limitations, it’s a timely look at AI’s role in chemical reasoning. open.substack.com/pub/bravoaba...
open.substack.com
Evaluating chemical reasoning capabilities of LLMs with ChemBench
As large language models (LLMs) sweep across scientific domains, chemistry confronts a critical question: can they rival—or even surpass—trained chemists in core reasoning tasks?
031
Kevin Jablonka @kjablonka.com · 17/03/2025
Work by Nawaf Alampara and Mara Schilling-Wilhelmi. Support by @carlzeissstiftung.bsky.social, Intel & Merck, OpenPhilanthropy, @fairmat.bsky.social Work done at the Friedrich-Schiller University Jena and HIPOLE Jena.
020
Kevin Jablonka @kjablonka.com · 17/03/2025
As a first step, we propose "eval cards" github.com/lamalab-org/... which hopefully encourage developers of eval methods to slow down (another thing I hope we do as a field: kjablonka.com/blog/posts/t...) and craft evals that really help move the field forward.
lnkd.in
LinkedIn
This link will take you to a page that’s not on LinkedIn
110
Kevin Jablonka @kjablonka.com · 17/03/2025
We hope to find ways to make it more transparent what assumptions go into building evaluations and to be more thoughtful when designing measurement instruments.
100
Kevin Jablonka @kjablonka.com · 17/03/2025
This article ended up including inspiration from many different fields and is again written in a (probably) unconventional form: But I really hope that some of the ideas in there spark a discussion as the field really needs those discussions.
100
Kevin Jablonka @kjablonka.com · 17/03/2025
And add this on top of the fact that most of our measurements in machine learning are actually defined by the act of the measurement itself - which is very unintuitive for a physical scientist.
100
Kevin Jablonka @kjablonka.com · 17/03/2025
To me, there is really a need for a "science of evals". When we build benchmarks or other methods for testing machine learning systems, we make many assumptions - too often hidden - that can drastically change what we measure (see the parallel coordinates plots).
110
Kevin Jablonka @kjablonka.com · 17/03/2025
Our team has been spending a lot of time building evaluations for machine learning systems. We have learned some lessons and wrote them down arxiv.org/abs/2503.10837
160
Reposted by Kevin Jablonka
Jablonka Lab (Lab for AI for Materials) @jablonkagroup.bsky.social · 11/03/2025
🚀ChemBench just leveled up! We’re thrilled to announce the latest release of ChemBench—now smarter and smoother! Dive into benchmarking any chemistry AI model with our revamped framework, designed for flexibility and ease. #ChemistryAI #MachineLearning #OpenScience #Innovation
111
Kevin Jablonka @kjablonka.com · 06/03/2025
Many of our benchmarks underwent a large revision in the last weeks. We now host HuggingFace spaces for them and the dataset. The revised article and post below give more details
000
Reposted by Kevin Jablonka
Chemical Society Reviews @chemsocrev.rsc.org · 06/03/2025
From @kjablonka.com, @mvictoriagil.bsky.social, @pepe-marquez.bsky.social and colleagues. 'From text to insight: large language models for chemical data extraction' #OpenAccess 🔓 pubs.rsc.org/en/content/a...
pubs.rsc.org
From text to insight: large language models for chemical data extraction
The vast majority of chemical knowledge exists in unstructured natural language, yet structured data is crucial for innovative and systematic materials design. Traditionally, the field has relied on m...
074
Kevin Jablonka @kjablonka.com · 12/02/2025
what is wrong with the link?
100
Reposted by Kevin Jablonka
Alán Aspuru-Guzik @aspuru.bsky.social · 24/01/2025
We at @digital-discovery.bsky.social are very happy to announce a new paper type called "Commit". Inspired by version control systems such as git, the idea is that if you have an update on a short and pointed publication, you can send it as a commit. We envision commits to be co-cited with the
88515
Reposted by Kevin Jablonka
Andrew White 🐦‍⬛ @andrew.diffuse.one · 21/01/2025
I've been thinking about how reasoning models will change AI applied to science. The recent papers from Deepseek/AI2/MoonShotAI are showing that we can exceed humans on reasoning tasks and I've written up some reflections on the consequences diffuse.one/p/d1-007
diffuse.one
diffuse.one
andrew white's blog.
0246
Reposted by Kevin Jablonka
Gianni De Fabritiis @gdefabritiis.bsky.social · 21/01/2025
Interested in reasoning agents? t.co/Ax9RtOLDCG @andrew.diffuse.one
0112
Reposted by Kevin Jablonka
Chris Oostenbrink @mms-boku.bsky.social · 17/01/2025
Super excited that we are able to advertise 12(!) fully funded PhD positions in the field of biomolecular technology of protein interactions. Check out the exciting projects, spread the word and join BioToP. #biotechnology #compchem @bokuvienna.bsky.social @fwf-at.bsky.social biotop.boku.ac.at
biotop.boku.ac.at
BioToP stands for "Biomolecular Technology of Proteins" and aims in interdisciplinary training of PhD students.
1238
Kevin Jablonka @kjablonka.com · 16/01/2025
More details: www.dropbox.com/scl/fi/mlew5...
dropbox.com
000
Kevin Jablonka @kjablonka.com · 16/01/2025
It would help us a lot if you could support us in spreading the word about the openings in our team! 🙏🏼
100
Kevin Jablonka @kjablonka.com · 16/01/2025
We strongly encourage candidates of all different backgrounds and identities to apply. Each new hire is an opportunity for us to bring in a different perspective, and we are always eager to diversify our team further.
100
Kevin Jablonka @kjablonka.com · 16/01/2025
In particular, we are looking for PostDocs to support our efforts in building and testing frontier models.
100
Kevin Jablonka @kjablonka.com · 16/01/2025
If you are excited about ML for materials science and chemistry, we might have good news for you. We are still hiring on all levels. Simply connect by submitting your application via forms.fillout.com/t/eoGA7AhnAKus.
forms.fillout.com
Join LAMA
Made with Fillout, the best way to make forms, surveys and quizzes your audience will answer.
153
Kevin Jablonka @kjablonka.com · 14/01/2025
😂
000
Reposted by Kevin Jablonka
Pepe Márquez @pepe-marquez.bsky.social · 03/01/2025
Our data extraction tutorial is now online in Chem. Soc. Rev. The notebooks can be run using the #jupyter4nfdi service from #base4nfdi. 📝 Paper: pubs.rsc.org/en/content/a... 💻 JupyterHub: t1p.de/matextract-cpu 📚 Online book: matextract.pub 📽️ intro from @kjablonka.com! 👇
pubs.rsc.org
From text to insight: large language models for chemical data extraction
The vast majority of chemical knowledge exists in unstructured natural language, yet structured data is crucial for innovative and systematic materials design. Traditionally, the field has relied on m...
0236
Reposted by Kevin Jablonka
Itai Yanai @itaiyanai.bsky.social · 09/11/2024
Doing good science is 90% finding a science buddy to constantly talk to about the project.
22877215
Kevin Jablonka @kjablonka.com · 23/12/2024
wow, congrats 🙏🏼
010
Reposted by Kevin Jablonka
Alán Aspuru-Guzik @aspuru.bsky.social · 23/12/2024
10 minutes ago I am excited to share a perspective on the much-needed topic of hashtag#safety for hashtag#selfdrivinglaboratories. As the field progresses, understanding the challenges and gaps in building safe setups will be crucial for scaling up this technology! doi.org/10.26434/che...
doi.org
Steering towards safe self-driving laboratories
The past decade has witnessed remarkable advancements in autonomous systems, such as automobiles that are evolving from traditional vehicles to ones capable of navigating complex environments without ...
3349
Reposted by Kevin Jablonka
Sam Rodriques @sgrodriques.bsky.social · 19/12/2024
FutureHouse is launching an independent postdoctoral fellowship program for exceptional researchers who want to apply our automated science tools to specific problems in biology and biochemistry, in collaboration with world-leading academic labs. 1/
14925
Reposted by Kevin Jablonka
Grant Rotskoff @grant.rotskoff.cc · 22/12/2024
Lot of cool stuff in here. Consistent with my working hypothesis that the main scientific utility of LLMs at the moment is plain old NLP
091
Kevin Jablonka @kjablonka.com · 21/12/2024
Our fantastic team: Mara Schilling-Wilhelmi, Martiño Ríos-García, Sherjeel Shabih, María Victoria Gil, Santiago Miret, Christoph T. Koch, @pepe-marquez.bsky.social Thanks to @carlzeissstiftung.bsky.social Intel Corporation Merck @fairmat.bsky.social CSIC for their support.
030
Kevin Jablonka @kjablonka.com · 21/12/2024
We plan to add capabilities to run some parts on GPU over the next weeks and release additional (video) tutorials. Let us know if you run into issues or have any feedback.
130
Kevin Jablonka @kjablonka.com · 21/12/2024
📝 Paper: pubs.rsc.org/en/content/a... 💻 Direct link to the JupyterHub: t1p.de/matextract-cpu 📚 Online book: matextract.pub
pubs.rsc.org
From text to insight: large language models for chemical data extraction
The vast majority of chemical knowledge exists in unstructured natural language, yet structured data is crucial for innovative and systematic materials design. Traditionally, the field has relied on m...
150