Sign in

Open Athena

@openathena.ai
91 followers 9 following 17 posts

We help academic labs building AI models to solve technology challenges so they can focus on scientific impact

PostsRepliesMedia
Open Athena @openathena.ai · 18/09/2026
We've only been observing the ocean directly for a few decades, and mostly at the surface. However, climate projections need to cover decades or centuries. Closing that gap is a machine learning problem, and it's the one Jesse Rusak @jder.bsky.social works on. Meet Jesse, MTS: bit.ly/oa-jesse
143
Open Athena @openathena.ai · 14/09/2026
18.71 billion documents. 25.25 trillion tokens. 152 data sources, deduplicated, decontaminated, and sorted into 200 buckets before a single training run. A new blog from Will Held @williamheld.com on Datakit, the Marin team's pretraining data pipeline: openathena.ai/blog/marin-d...
Flow diagram titled "How data flows through Datakit." A wide band of 25.25 trillion normalized tokens narrows to 23.12T after deduplication removes 2.13T duplicate tokens, then to 23.11T after decontamination removes 13.66B tokens of benchmark overlap. The remaining data fans out into 40 topics, crossed with five quality bands to form 200 buckets. A scatter plot of proxy training runs sits alongside the optimization criterion — minimize predicted low-noise mean BPB — leading to a final step: choose 200 weights, one per bucket.
0115
Open Athena @openathena.ai · 27/08/2026
Alex Merose @al.merose.com was headed into design until a friend showed him a demo: audio windows shrinking until noise resolved into sound, tuned by ear—a preprocessing step for a brain-computer interface. He is now MTS at OA, working primarily on Samudra with NYU & MIT. Meet Alex: bit.ly/oa-alex
0102
Open Athena @openathena.ai · 13/08/2026
A decade ago, progress in NLP meant encoding a language's structure into the model. @williamheld.com, MTA at OA, came up in that tradition and now works on Marin, our open LLM. Read about the bitter lesson, open development tradeoffs, and language roots & quirks at www.openathena.ai/blog/meet-ou...
051
Open Athena @openathena.ai · 30/07/2026
Simulating ocean climate takes a supercomputer 4,600+ CPU cores to produce 12 simulated years per day (SYPD). Samudra 2 produces 4,800 SYPD on 1 GPU at the same resolution. In a new blog, @al.merose.com reports on Samudra, a neural ocean emulator built in collaboration with NYU & MIT: bit.ly/oa-ss
0132
Open Athena @openathena.ai · 16/07/2026
Research software engineers used to be "builders of pipes." Now those pipes are generated on demand, and the work shifts to figuring out which ideas are worth exploring. VP of Engineering Yael Elmatad discusses the RSE role in the age of agentic tooling in a new blog: openathena.ai/blog/researc...
020
Open Athena @openathena.ai · 09/07/2026
Betsy Cannon is a member of technical staff, where she leads material science projects RHOAR-Net, with Princeton's Rosen Research Group, and MarinMat. Read her story on learning DFT, how AI changes shared research infrastructure, and what pottery taught her about building by hand: bit.ly/oa-betsy
040
Open Athena @openathena.ai · 25/06/2026
In a new blog, Russell Power explains how the Marin team nearly doubled its sustained TPU usage by creating a custom global scheduler: Iris. Iris searches every region where Marin has compute, places each job wherever capacity appears, and moves data along as needed. 🔗 openathena.ai/blog/cluster...
1102