Sign in

Geodesic Research

@geodesicresearch.bsky.social
13 followers 8 following 28 posts

We're behind alignmentpretraining.ai. Let's align some AIs. geodesicresearch.ai

PostsRepliesMedia
Geodesic Research @geodesicresearch.bsky.social · 01/06/2026
As Geodesic scales, we're delighted to welcome Nathalie Kirch as our newest member of technical staff. We can't wait to get to work 🚀
010
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
We aim to fill these roles as soon as we find suitable team members — 𝘺𝘰𝘶'𝘭𝘭 𝘣𝘦 𝘢𝘮𝘰𝘯𝘨 𝘰𝘶𝘳 𝘧𝘪𝘳𝘴𝘵 𝘵𝘦𝘤𝘩𝘯𝘪𝘤𝘢𝘭 𝘩𝘪𝘳𝘦𝘴, 𝘸𝘪𝘵𝘩 𝘴𝘶𝘣𝘴𝘵𝘢𝘯𝘵𝘪𝘢𝘭 𝘪𝘯𝘧𝘭𝘶𝘦𝘯𝘤𝘦 𝘰𝘯 𝘩𝘰𝘸 𝘵𝘩𝘦 𝘰𝘳𝘨 𝘨𝘳𝘰𝘸𝘴 𝘢𝘯𝘥 𝘩𝘢𝘴 𝘪𝘮𝘱𝘢𝘤𝘵. Apply!: airtable.com/appuugUGFPJE...
000
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
We're looking for technical staff with experience across the ML and alignment research stack: • multi-GPU / HPC training and evals • deep familiarity with data-centric alignment methods • impact-oriented thinking about lab adoption and theories of change
100
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
This grant also keeps us epistemically independent. No commercial stake in any particular alignment method. 𝘞𝘦 𝘤𝘢𝘯 𝘪𝘯𝘷𝘦𝘴𝘵𝘪𝘨𝘢𝘵𝘦 𝘵𝘩𝘦 𝘧𝘶𝘭𝘭 𝘱𝘪𝘤𝘵𝘶𝘳𝘦 and publish what survives realistic capabilities pipelines for any lab to integrate.
100
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
Both the alignment interventions we want to study and the RL adversaries we stress-test them against are deeply data- and compute-intensive. Coefficient Giving's grant puts us 𝘢𝘮𝘰𝘯𝘨 𝘵𝘩𝘦 𝘧𝘦𝘸 𝘯𝘰𝘯-𝘭𝘢𝘣 𝘢𝘤𝘵𝘰𝘳𝘴 𝘸𝘩𝘰 𝘤𝘢𝘯 𝘳𝘦𝘱𝘭𝘪𝘤𝘢𝘵𝘦 𝘧𝘶𝘭𝘭 𝘮𝘪𝘥𝘵𝘳𝘢𝘪𝘯𝘪𝘯𝘨 + 𝘚𝘍𝘛 + 𝘙𝘓 𝘴𝘵𝘢𝘤𝘬𝘴 𝘢𝘵 𝘧𝘳𝘰𝘯𝘵𝘪𝘦𝘳 𝘰𝘱𝘦𝘯-𝘴𝘰𝘶𝘳𝘤𝘦 𝘴𝘤𝘢𝘭𝘦.
100
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
Our current work: preventing long-horizon capabilities RL from degrading those priors and selecting for schemers, metagamers, and reward seekers.
100
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
Six months ago, we found alignment priors matter, and the field is now converging here (see Anthropic's Teaching Claude Why and Model Spec Midtraining) anthropic.com/research/tea... alignment.anthropic.com/2026/msm/
anthropic.com
Teaching Claude why
New research on how we've reduced agentic misalignment
100
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
We're a Cambridge-based AI safety org. Our seminal work (alignmentpretraining.ai) showed you can bake alignment priors into base models.
alignmentpretraining.ai
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
LLMs trained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist th...
120
Geodesic Research @geodesicresearch.bsky.social · 28/05/2026
Thanks to a generous philanthropic grant (pending final logistics) from Coefficient Giving, 𝘎𝘦𝘰𝘥𝘦𝘴𝘪𝘤 𝘪𝘴 𝘩𝘪𝘳𝘪𝘯𝘨 𝘔𝘦𝘮𝘣𝘦𝘳𝘴 𝘰𝘧 𝘛𝘦𝘤𝘩𝘯𝘪𝘤𝘢𝘭 𝘚𝘵𝘢𝘧𝘧. Come build the base of alignment with us 🤖 Applications now open: airtable.com/appuugUGFPJE...
142
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
You'll be among our first technical hires, with real influence on how the org grows and has impact. EOI is 5 questions, ~5 mins. If it's a fit, we'll follow up about the full application: tally.so/r/vG4G6A
tally.so
Join Geodesic Research and Help Build the Base for Alignment
020
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
We're looking for technical staff with experience across the ML and alignment research stack: - multi-GPU / HPC training and evals experience - deep familiarity with data-centric alignment methods - an insatiable desire to improve the outcomes of developing superintelligence
120
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
More generally, capabilities RL selecting misaligned behaviours remains one of the primary reasons why alignment remains a hard and unsolved problem. We believe building a robust initialisation is a promising way to avoid this entirely. t.co/xTBKy3CGpU
t.co
https://www.lesswrong.com/posts/epjuxGnSPof3GnMSL/alignment-remains-a-hard-unsolved-problem
130
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
(iii) Apollo Research and OpenAI showed that capabilities RL can lead to meta-gaming, developing an acute awareness of being in training and being monitored. Ideally, models should be robust to exploring these behaviours. t.co/lx0HSBb7V2
t.co
https://alignment.openai.com/metagaming
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
(ii) OpenAI, led by our advisor Tomek Korbak: capabilities RL uniformly degrades alignment across a range of evals, even on top of our alignment pretraining datasets. t.co/NtleGyGoEN
t.co
https://alignment.openai.com/how-far-does-alignment-midtraining-generalize/
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
This comes after three recent results: (i) @anthropic.com : reward hacking in production SWE training elicits broader misalignment — not just narrow gaming of the reward signal. t.co/z5zhaUCOJ2
t.co
https://assets.anthropic.com/m/74342f2c96095771/original/Natural-emergent-misalignment-from-reward-hacking-paper.pdf
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
Alignment priors are important, and the field is converging here. @anthropic.com's recent work lean on the alignment-priors approach we pioneered. However, we believe long-horizon RL can degrade alignment priors. anthropic.com/research/tea... alignment.anthropic.com/2026/msm/
anthropic.com
Teaching Claude why
New research on how we've reduced agentic misalignment
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
Geodesic is hiring Members of Technical Staff. We're a Cambridge-based AI safety org. Our seminal work showed you can bake alignment priors into base models. Now, we want to make base models robust to the adversarial effects of long-horizon capabilities RL. EOI ~5 mins: tally.so/r/vG4G6A
tally.so
Join Geodesic Research and Help Build the Base for Alignment
160
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
We're confident enough in these findings that we've completely restructured Geodesic internally to pursue this agenda. We're hopeful alignment pretraining can help us create deeply good AI systems.
030
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
Overall, we're hopeful this work will help establish the field of Alignment Pretraining - the dedicated study of how different data mixes at earlier stages of training affect model propensities and personas.
120
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
Releasing all models, data, and a 4000+ question eval suite at alignmentpretraining.ai Paper: "Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment" Authors: Cameron Tice*, Puria Radmard*, Samuel Ratnam, Andy Kim, David Demitri Africa, and Kyle O'Brien
alignmentpretraining.ai
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
LLMs trained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist th...
130
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
When implementing alignment pretraining we saw some small capability degradations (around 4% drop at most when averaged across our benchmarks). We're excited for people to help optimize these alignment mixes to both make the models deeply good, but also keep them at the frontier.
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
Although we targeted a narrow range of AI behaviours in our training, we found this generalised to very OOD evals such as Dark Triad Personality Traits. This has us excited for future work looking at more wholistic pretraining mixes targeting the generation of a deep character.
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
However, we found alignment pretraining was even more useful than expected, and improved the behaviours of our post-trained models (trained on over 4.5M examples of desired AI behaviour from @natolambert.bsky.social's and Olmo3's full post-training mix).
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
We expected developing good priors in pretraining to be useful for creating good inductive biases for post-training to avoid selecting for deceptively misaligned models. This thinking largely follows Evan Hubinger's writings on Why Alignment Remains a Hard, Unsolved Problem.
130
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
We pretrained 6.9B parameter LLMs under 4 conditions: • Unfiltered (normal data) • Filtered (AI discourse removed) • Misalignment stories upsampled • Alignment stories upsampled Then tested: When prompted as "an AI assistant," how often do they choose misaligned actions?
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
We tried to filter out all information about AI systems + non-human intelligent species acting poorly towards humans. Here's an example of a document we filtered out of our midtraining data:
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
If pretraining data is full of examples of AI behaving badly (sci-fi villains, safety papers on scheming, news about AI crises), models might learn these as priors for how "an AI" should act. @turntrout.bsky.social called this "self-fulfilling misalignment", and we found evidence it exists.
110
Geodesic Research @geodesicresearch.bsky.social · 19/05/2026
We pretrained multiple 7B LLMs from scratch and found that natural exposure to AI misalignment discourse causes models to become more misaligned. Optimistically, we also find that adding positive synthetic documents in pretraining reduces misalignment. Thread 🧵
160