Sign in

Open Athena

@openathena.ai
89 followers 9 following 17 posts

We help academic labs building AI models to solve technology challenges so they can focus on scientific impact

PostsRepliesMedia
Open Athena @openathena.ai · 22/09/2026
We’re excited to announce that we've been selected as one of ARIA's new Activation Partners. We’re looking forward to helping researchers use frontier AI capabilities on the scientific questions that matter most. openathena.ai/blog/open-at...
openathena.ai
Open Athena selected as an ARIA Activation Partner
Open Athena has been selected as one of ARIA's new Activation Partners, pairing our engineers with ARIA-funded research teams to build open scientific models, datasets and benchmarks.
000
Open Athena @openathena.ai · 18/09/2026
We've only been observing the ocean directly for a few decades, and mostly at the surface. However, climate projections need to cover decades or centuries. Closing that gap is a machine learning problem, and it's the one Jesse Rusak @jder.bsky.social works on. Meet Jesse, MTS: bit.ly/oa-jesse
143
Open Athena @openathena.ai · 14/09/2026
18.71 billion documents. 25.25 trillion tokens. 152 data sources, deduplicated, decontaminated, and sorted into 200 buckets before a single training run. A new blog from Will Held @williamheld.com on Datakit, the Marin team's pretraining data pipeline: openathena.ai/blog/marin-d...
Flow diagram titled "How data flows through Datakit." A wide band of 25.25 trillion normalized tokens narrows to 23.12T after deduplication removes 2.13T duplicate tokens, then to 23.11T after decontamination removes 13.66B tokens of benchmark overlap. The remaining data fans out into 40 topics, crossed with five quality bands to form 200 buckets. A scatter plot of proxy training runs sits alongside the optimization criterion — minimize predicted low-noise mean BPB — leading to a final step: choose 200 weights, one per bucket.
0104
Open Athena @openathena.ai · 03/09/2026
A year ago, Marin had one FTE. Recently, the ten-person team launched a 535 billion parameter MoE model. In the Marin 535B-A23B launch note, David Hall @dlwh.bsky.social notes it's going "almost boringly well." Thank you to our colleagues, partners, and advisors for making this possible. bit.ly/535b
openathena.ai
Marin 535B-A23B launch note
A look back at Marin's first year at Open Athena and forward at the 535B-A23B hero run, the largest model the team has ever trained.
072
Open Athena @openathena.ai · 02/09/2026
With the generous support of The Jen-Hsun and Lori Huang Foundation, we have launched the largest live-streamed pretraining run in history, and Marin’s largest model yet: a fully open source 535B total parameter MoE model, with 18T tokens of data. Learn more at: openathena.ai/blog/huang-f...
openathena.ai
Marin's 535 billion parameter model training run launched with support of The Jen-Hsun and Lori Huang Foundation GPU gift
Open Athena announces the launch of Marin's largest training run yet—a 535B parameter large language model—with a generous gift of compute from The Jen-Hsun and Lori Huang Foundation.
0135
Open Athena @openathena.ai · 27/08/2026
Alex Merose @al.merose.com was headed into design until a friend showed him a demo: audio windows shrinking until noise resolved into sound, tuned by ear—a preprocessing step for a brain-computer interface. He is now MTS at OA, working primarily on Samudra with NYU & MIT. Meet Alex: bit.ly/oa-alex
0102
Open Athena @openathena.ai · 18/08/2026
We are thrilled to collaborate with Cornell's Jingjing Zhai, Edgar Marroquin, Matt Pennell, Ed Buckler, and team on PlantCAD2, recently published in Cell Genomics. OA staff members @eczech0.bsky.social and Betsy Cannon are paper co-authors and contributors. Read at: www.cell.com/cell-genomic...
cell.com
PlantCAD2: A DNA foundation model for interpreting genomes across flowering plants
Zhai et al. introduce PlantCAD2, a DNA foundation model pre-trained on 65 angiosperm genomes. Despite being 10-fold smaller, PlantCAD2 outperforms the 7-billion-parameter Evo2 on most plant-specific z...
021
Open Athena @openathena.ai · 13/08/2026
A decade ago, progress in NLP meant encoding a language's structure into the model. @williamheld.com, MTA at OA, came up in that tradition and now works on Marin, our open LLM. Read about the bitter lesson, open development tradeoffs, and language roots & quirks at www.openathena.ai/blog/meet-ou...
051
Reposted by Open Athena
Gonzalo Benegas @gonzalobenegas.bsky.social · 06/08/2026
Excited to share MarinDNA, a 1B gLM that rivals Evo 2 40B on variant effect prediction while being 2,330x faster. With @eczech0.bsky.social, we built around a standard Transformer so we could reuse LLM infra and methods while focusing on data curation and scaling. openathena.ai/blog/marin-d... 🧵
1125
Open Athena @openathena.ai · 30/07/2026
Simulating ocean climate takes a supercomputer 4,600+ CPU cores to produce 12 simulated years per day (SYPD). Samudra 2 produces 4,800 SYPD on 1 GPU at the same resolution. In a new blog, @al.merose.com reports on Samudra, a neural ocean emulator built in collaboration with NYU & MIT: bit.ly/oa-ss
0132
Open Athena @openathena.ai · 28/07/2026
David Siegel, founder and chairman of Open Athena, writes in Fortune on the importance of open source. David likens AI to the libraries of the future, just as important for learning and dissemination of knowledge as libraries have been. Read the full article at: openathena.ai/blog/david-s...
openathena.ai
I argued with the father of open source for 2 years. Now the AI fight is the same — only bigger.
David Siegel, founder and chairman of Open Athena, writes in Fortune about the importance of open source in AI.
021
Open Athena @openathena.ai · 16/07/2026
Research software engineers used to be "builders of pipes." Now those pipes are generated on demand, and the work shifts to figuring out which ideas are worth exploring. VP of Engineering Yael Elmatad discusses the RSE role in the age of agentic tooling in a new blog: openathena.ai/blog/researc...
020
Open Athena @openathena.ai · 09/07/2026
Betsy Cannon is a member of technical staff, where she leads material science projects RHOAR-Net, with Princeton's Rosen Research Group, and MarinMat. Read her story on learning DFT, how AI changes shared research infrastructure, and what pottery taught her about building by hand: bit.ly/oa-betsy
040
Open Athena @openathena.ai · 01/07/2026
In 16th-century Italy, mathematicians kept their discoveries secret, saving them to deploy as weapons in public "math duels" for prestige and patronage. In a new blog post, David Hall argues that frontier AI has its own culture of secrecy that we must address. openathena.ai/blog/open-de...
060
Open Athena @openathena.ai · 25/06/2026
In a new blog, Russell Power explains how the Marin team nearly doubled its sustained TPU usage by creating a custom global scheduler: Iris. Iris searches every region where Marin has compute, places each job wherever capacity appears, and moves data along as needed. 🔗 openathena.ai/blog/cluster...
1102
Open Athena @openathena.ai · 17/06/2026
At a recent panel on driving the technological future forward, our COO and CSO Jared Crooks explained the importance of keeping AI development open, why ethics should be embedded in tech, and more. Watch or read the discussion at the link
openathena.ai
Open Athena | Preparing for the AI Future with Ethics in Mind
Open Athena is a nonprofit that accelerates academia with capabilities from the AI frontier
042
Open Athena @openathena.ai · 04/06/2026
In our latest blog post, Marin team member Larry Dial describes the pretraining techniques we're using as we transition from dense models to more efficient Mixture of Experts (MoE) models. This demos the stability and predictability of MoEs, giving us a promising direction for our next training run
openathena.ai
Open Athena | Improving our LLM Pretraining Efficiency
Open Athena is a nonprofit that accelerates academia with capabilities from the AI frontier
1103