Sign in

Sam Blau

@samblau.bsky.social
2.2K followers 340 following 36 posts

Research scientist & computational chemist at Berkeley Lab using HT DFT workflows, machine learning, and reaction networks to model complex reactivity.

PostsRepliesMedia
Sam Blau @samblau.bsky.social · 17/12/2025
Emory and I also wrote a higher level summary of the work available here: rdcu.be/eUVIw
rdcu.be
Deep learning accelerates discovery of complex nanomaterials
Nature Computational Science - A physics-infused heterogeneous graph neural network has been developed to address challenges in designing complex nanomaterials with spatially varying compositions....
030
Sam Blau @samblau.bsky.social · 17/12/2025
Out in @natcomputsci.nature.com: A roadmap for inverse design of #nanomaterials heterostructures via HT data gen -> representation dev -> heteroGNN training -> gradient-based global opt! w/ @emorychannano.bsky.social @ewcspottesmith.bsky.social www.nature.com/articles/s43... Free link rdcu.be/eTH72
nature.com
Gradient-based optimization of complex nanoparticle heterostructures enabled by deep learning on heterogeneous graphs - Nature Computational Science
Graph neural networks built on physically motivated representations enable gradient-based optimization of complex upconverting nanoparticle heterostructures, revealing photophysical design rules and a...
151
Sam Blau @samblau.bsky.social · 17/12/2025
@berkeleylab.lbl.gov @molecularfoundry.lbl.gov @mitcheme.bsky.social
000
Reposted by Sam Blau
Emory Chan @emorychannano.bsky.social · 16/12/2025
For a more digestible summary of our recent @natcomputsci.nature.com article, here's the research brief that @samblau.bsky.social and I wrote... complete with a #BehindThePaper tidbit that thankfully has less drama than a VH1 Behind the Music episode (free link: rdcu.be/eUVIw)
rdcu.be
Home
Need literature management for your research driven company? ReadCube helps your company discover, organize, read, annotate, share and cite.
021
Reposted by Sam Blau
Emory Chan @emorychannano.bsky.social · 09/12/2025
Our manuscript on deep learning of upconverting nanoparticle heterostructures w/ @samblau.bsky.social is finally out in @natcomputsci.nature.com! See 🧵for links to free access via Readcube and chemrxiv. @berkeleylab.lbl.gov @molecularfoundry.lbl.gov @mitcheme.bsky.social
162
Sam Blau @samblau.bsky.social · 04/11/2025
I'm hiring postdocs @berkeleylab.lbl.gov to drive cutting-edge research involving MLIPs, high-throughput workflows, chemical reaction networks, generative models, and open-source software dev. Full position description + application here: forms.gle/zePBZDmciXez... #Chempostdoc #AI4Science
forms.gle
063
Sam Blau @samblau.bsky.social · 14/08/2025
Come work with @emorychannano.bsky.social and me! #robotics #nanochemistry #machinelearning #UCNPs
002
Reposted by Sam Blau
Evan Spotte-Smith (they/them) @ewcss.info · 31/07/2025
Interested in learning more about our recently published OMol25 dataset and the advances that it's bringing to atomistic machine learning? Check out this talk that my boy @samblau.bsky.social gave as part of the "Modeling Talk Series". #CompChem ⚗️ 🧪 #SciML
sites.google.com
Modeling Talk Series - The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models
Samuel Blau, Berkeley Lab Video Recording Slides (pptx, pdf)
193
Sam Blau @samblau.bsky.social · 29/07/2025
I'm presenting OMol25 tomorrow 7/29 at 9 AM PST as part of a talk series at Google. Learn how we built the dataset + how MLIPs trained on OMol are revolutionizing comp chem! Meet: lnkd.in/g4AAWkcK YouTube Stream: lnkd.in/ggmtMtTR Join group: lnkd.in/g5ciuNuX
030
Sam Blau @samblau.bsky.social · 15/05/2025
OMol25 was calculated with ORCA. I want to acknowledge the work of the ORCA team to improve the quality of the gradient + the robustness of SCF convergence for complicated systems as part of the OMol effort - it was much appreciated and critical to ensuring that we're releasing high quality data!
0132
Reposted by Sam Blau
Berkeley Lab @berkeleylab.lbl.gov · 14/05/2025
🚨 Just dropped: Open Molecules 2025 — a record-breaking dataset co-led by Berkeley Lab + Meta FAIR. 100M+ DFT snapshots. Built to train #AI for real-world chemistry 🧪. Could reshape discovery in batteries, drug discovery & much more! @cs.lbl.gov ⬇️
newscenter.lbl.gov
Computational Chemistry Unlocked: A Record-Breaking Dataset to Train AI Models has Launched - Berkeley Lab
Scientists will finally be able to simulate the chemistry that drives our bodies, our environment, and our technologies.
0143
Sam Blau @samblau.bsky.social · 14/05/2025
We can't wait to see what the community does with OMol! Don't hesitate to reach out with feedback on the data, models, or paper - we aren't going to submit to a journal until the leaderboard goes up, which means we have time to incorporate community feedback (within reason) 10/10
030
Sam Blau @samblau.bsky.social · 14/05/2025
A special shout out to co-first authors Daniel Levine and Muhammed Shuaibi who moved mountains making OMol a reality. I also want to recognize the substantial and critical contributions of @ewcspottesmith.bsky.social, Michael Taylor, Muhammad Hasyim, and Kyle Michel 9/N
140
Sam Blau @samblau.bsky.social · 14/05/2025
Co-leading OMol with Brandon and Larry was a joy and an honor - as was assembling a world-leading team of scientists from 2 companies, 2 national labs, and 6 universities who were excited to help build an open-source, revolutionary molecular DFT dataset to push science forward 8/N
110
Sam Blau @samblau.bsky.social · 14/05/2025
Right now, OMol data has energy, forces, partial charges, partial spins, and HOMO/LUMO. But we have far more info that we still need to parse and hope to do a battery of GBW postprocessing. Plus we have 10 petabytes of electron densities. Lots more to come! 7/N
110
Sam Blau @samblau.bsky.social · 14/05/2025
And check out the UMA demo (facebook-fairchem-uma-demo.hf.space UMA is trained on OMol + other FAIR Chemistry datasets) - metal complexes at +1 vs +2 correctly optimize to tetrahedral/planar and reduced ethylene carbonate correctly ring-opens while a neutral EC remains stable 6/N
facebook-fairchem-uma-demo.hf.space
Gradio
120
Sam Blau @samblau.bsky.social · 14/05/2025
Data, models, & paper are available to download now! 5/N Paper: arxiv.org/abs/2505.08762 Data + models: huggingface.co/facebook/OMo...
arxiv.org
The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models
Machine learning (ML) models hold the promise of transforming atomic simulations by delivering quantum chemical accuracy at a fraction of the computational cost. Realization of this potential would en...
110
Sam Blau @samblau.bsky.social · 14/05/2025
We're also releasing baseline models trained on OMol. To guide future MLIP development, we built novel evaluations on intermolecular interactions, conformers, and charge/spin. We hope to include frequency, ΔG, and TSopt tasks when we put up a public leaderboard in the summer 4/N
120
Sam Blau @samblau.bsky.social · 14/05/2025
OMol was constructed via an unprecedented diversity of methods: MD, ML-MD, RPMD, rattling, Architector, rxn path interpolation, AFIR, optimization, and scaled separation. We also recalculated some previous datasets and did additional sampling/structure generation atop others 3/N
140
Sam Blau @samblau.bsky.social · 14/05/2025
OMol covers 83 elements, a wide range of intra and intermolecular interactions, explicit solvation, reactive structures, conformers, charges -10 to 10, 0-10 unpaired electrons, and 2-350 atoms per snapshot. It required >6B CPU hrs, 10x more than any prev MLIP training dataset 2/N
131
Sam Blau @samblau.bsky.social · 14/05/2025
The Open Molecules 2025 dataset is out! With >100M gold-standard ωB97M-V/def2-TZVPD calcs of biomolecules, electrolytes, metal complexes, and small molecules, OMol is by far the largest, most diverse, and highest quality molecular DFT dataset for training MLIPs ever made 1/N
54610
Sam Blau @samblau.bsky.social · 01/05/2025
It was a pleasure to give an IIDAI seminar on nanoparticle ML for gradient-based heterostructure optimization (w/ @emorychannano.bsky.social ) and neural network path opt for finding reaction transition states on MLIPs (w/ @thglab.bsky.social) - find the talk here: www.youtube.com/watch?v=-4jB...
youtube.com
IIDAI Seminar, 5/1/2025, Samuel M. Blau (Berkeley Lab)
YouTube video by Coordinated Science Laboratory
010
Reposted by Sam Blau
Andrew S. Rosen @andrewrosen.bsky.social · 01/05/2025
🧠 New postdoctoral researcher position at Princeton for those interested in data science and machine learning! Specify my group if you are interested in working together. Deadline is May 31. Details: puwebp.princeton.edu/AcadHire/app...
puwebp.princeton.edu
062
Sam Blau @samblau.bsky.social · 31/03/2025
Final day to submit abstracts for ACS Fall 2025! Reminder that @ewcspottesmith.bsky.social , Brett Savoie (Notre Dame), and I are organizing a symposium on "Chemical Reaction Networks, Retrosynthesis, and Reaction Prediction". Will be a mix of invited and contributed talks - please submit! #CompChem
030
Reposted by Sam Blau
Gabe Gomes @gabegomes.bsky.social · 24/03/2025
the @gpggrp.bsky.social is at the ACS Spring 2025! come check out the works of Daniil Boiko and Rob MacKnight at the "ML + AI in Organic Chemistry" Symposium (Hall B-1, Room 4) today! extreme scaling of experimental chemical reactions via MS and an OS for autonomous comp chem!
042
Sam Blau @samblau.bsky.social · 21/03/2025
Looking forward to speaking at ACS on Sunday at 5:30! Come learn about "Popcornn" - a new method for double-ended transition state optimization atop machine learned interatomic potentials that is substantially better than NEB or GSM.
030
Sam Blau @samblau.bsky.social · 13/03/2025
Fantastic new work from Aditi & co that shows how to leverage the expressivity + accuracy of massive pre-trained MLIPs to distill smaller, much faster models that are still extremely accurate to drive downstream simulations - no need to compromise on speed vs accuracy!
030
Sam Blau @samblau.bsky.social · 24/01/2025
Applications closing in one week! If you’re interested in a prestigious postdoc at the intersection of AI/ML and nuclear nonproliferation, don’t hesitate to apply - come work with me on fascinating f-block chemistry and computational/ML methods! (Must be a US citizen)
030
Reposted by Sam Blau
John Parkhill @johnparkhill.bsky.social · 10/01/2025
06113
Reposted by Sam Blau
Evan Spotte-Smith (they/them) @ewcss.info · 08/01/2025
@samblau.bsky.social, Brett Savoie (Notre Dame), and I are organizing a symposium for @amerchemsociety.bsky.social Fall 2025 called "Chemical Reaction Networks, Retrosynthesis, and Reaction Prediction" under @acscomp.bsky.social. #reactionnetwork #CRN #retrosynthesis 🧪 ⚗️ #CompChem
11610
Reposted by Sam Blau
ChemRxiv Bot @chemrxivbot.bsky.social · 26/12/2024
Inverse Design of Complex Nanoparticle Heterostructures via Deep Learning on Heterogeneous Graphs Authors: Eric Sivonxay, Lucas Attia, Evan Walter Clark Spotte-Smith, Benjamin Lengeling, Xiaojing Xia, Daniel Barter, Emory Chan, Samuel Blau DOI: 10.26434/chemrxiv-2024-1dw4q
063
Sam Blau @samblau.bsky.social · 27/12/2024
Very proud of this work, going all the way from implementing the kMC in C++ to building datasets w/ high-throughput workflows to designing the novel graph representation to training the custom hetero-GNN w/ on-the-fly augmentation to inverse design of novel nanoparticles with GNN-based optimization!
190
Sam Blau @samblau.bsky.social · 25/12/2024
Looks like good work, but please improve model labeling. OMat24 is a dataset, not a specific trained model. MACE is an architecture, not a specific trained model. MatBench Discovery attempts to specify architecture, training set, and date added. If you’re pulling from there, please do the same!
140
Sam Blau @samblau.bsky.social · 24/12/2024
Ahh, I didn’t realize these were three-body benchmarks specifically (though obvious in retrospect from their names). I assume you’re using wB97M-D4refit? What percentage of the total energy and forces come from the many body term?
100
Reposted by Sam Blau
Bingqing Cheng @chengbingqing.bsky.social · 23/12/2024
Long-range machine learning potentials strike again! 🚀 We benchmarked the Latent Ewald Summation method on diverse systems—molecules, solutions, interfaces. Learning just from energy & forces, it delivers the most accurate potential energy surfaces, physical charges, dipoles, and quadrupoles!
arxiv.org
Learning charges and long-range interactions from energies and forces
Accurate modeling of long-range forces is critical in atomistic simulations, as they play a central role in determining the properties of materials and chemical systems. However, standard machine lear...
1182
Sam Blau @samblau.bsky.social · 24/12/2024
How does wB97M-D4 compare to wB97M-V for those charged systems?
100
Reposted by Sam Blau
Evan Spotte-Smith (they/them) @ewcss.info · 18/12/2024
This week, "RNMC: kinetic Monte Carlo implementations for complex reaction networks" was published in @joss-openjournals.bsky.social. Work on RNMC started back in 2021, when brilliant mathematician Daniel Barter suggested using stochastic methods to study #reactionnetworks. #CompChem #ChemSky 🧪 1/6
joss.theoj.org
RNMC: kinetic Monte Carlo implementations for complex reaction networks
Zichi et al., (2024). RNMC: kinetic Monte Carlo implementations for complex reaction networks. Journal of Open Source Software, 9(104), 7244, https://doi.org/10.21105/joss.07244
1101
Reposted by Sam Blau
Evan Spotte-Smith (they/them) @ewcss.info · 18/12/2024
I've been slacking on research updates here! A few weeks ago, a preprint for "HEPOM: Using Graph Neural Networks for the accelerated predictions of Hydrolysis Free Energies in different pH conditions" dropped on @chemrxiv.bsky.social. 1/5
chemrxiv.org
HEPOM: Using Graph Neural Networks for the accelerated predictions of Hydrolysis Free Energies in different pH conditions.
Hydrolysis is a fundamental family of chemical reactions where water facilitates the cleavage of bonds. The process is ubiquitous in biological and chemical systems, owing to water's remarkable versat...
131
Reposted by Sam Blau
Kjell Jorner @valencekjell.com · 10/12/2024
Another chance to join our group! We are recruiting a PhD student in digital ligand engineering for nanocatalysis. Reposts and spreading the word to interested people in your network appreciated! jobs.ethz.ch/job/view/JOP...
jobs.ethz.ch
PhD position in digital ligand engineering for nanocatalysis
0137
Sam Blau @samblau.bsky.social · 03/12/2024
So glad to finally see the code/models released! Huge props to MSR, I know it is challenging for a company to release stuff like this but it’s so important to science! Any chance the training data will be released too? Thanks!
070
Sam Blau @samblau.bsky.social · 02/12/2024
Example nanoparticle heterostructure optimization, driven by gradients of UV emission with respect to layer thicknesses and dopant concentrations from our hetero-GNN (not accessible from kMC) and sub-second inference (vs days from kMC) #F24MRS
040
Sam Blau @samblau.bsky.social · 02/12/2024
Excited to speak at #F24MRS Thurs 1:30 - 1st talk of my career w/o any DFT connection - we design a hetero-GNN for learning core-shell nanoparticle properties, train on first ever large-scale NP kMC dataset, and use autodiff to optimize -> discover far OOD heterostructures with >6x enhanced emission
071
Reposted by Sam Blau
Philippe Schwaller @pschwllr.bsky.social · 02/12/2024
We are hiring (resharing appreciated)! Given recent successful grant applications (I got my SNSF Starting Grant 🚀), we are extending the LIAC team with multiple openings (PhD/postdoc) for 2025. Apply now (deadline: December 20th) by filling in this form: forms.fillout.com/t/eq5ADAw3kkus. #ChemSky
610271
Sam Blau @samblau.bsky.social · 21/11/2024
A recent and closely related development is @ask1729.bsky.social 's EScAIP architecture, which shows that scaling attention yields better MLIP performance than equivariance (and avoids expensive tensor operations). EScAIP is SOTA on both accuracy *and* speed for molecules, materials, and catalysts!
arxiv.org
The Importance of Being Scalable: Improving the Speed and Accuracy of Neural Network Interatomic Potentials Across Chemical Domains
Scaling has been critical in improving model performance and generalization in machine learning. It involves how a model's performance changes with increases in model size or input data, as well as ho...
060
Sam Blau @samblau.bsky.social · 21/11/2024
Welcome back, my friend! Looking forward to seeing you at MRS in two weeks :)
120
Sam Blau @samblau.bsky.social · 17/11/2024
Disagree on both points. Large collaborative efforts are necessary to capture diverse & synergistic viewpoints + expertise - and allow for datasets, models, and benchmarks to be developed in concert to deliver scientific utility and motivate worthwhile follow-on work by the community
110
Sam Blau @samblau.bsky.social · 15/11/2024
Yes please!
030
Sam Blau @samblau.bsky.social · 14/11/2024
Would appreciate being added - thanks!
110
Sam Blau @samblau.bsky.social · 14/11/2024
It goes deeper - some people prefer NNIP, though technically there is a difference (all NNIPs are MLIPs, but not all MLIPs are NNIPs)
020
Sam Blau @samblau.bsky.social · 14/11/2024
Would appreciate being added - thanks!
010