Sign in

Nora Belrose

@norabelrose.bsky.social
1K followers 15 following 38 posts

AI, philosophy, spirituality Head of interpretability research at EleutherAI, but posts are my own views, not Eleuther’s.

PostsRepliesMedia
Nora Belrose @norabelrose.bsky.social · 24/10/2025
why don't more people become zoroastrian? it's where judaism and christianity got the idea of ethical monotheism, afterlife, and final judgment but without any of their baggage (no eternal hell, no historically questionable dogmas, etc.)
030
Nora Belrose @norabelrose.bsky.social · 11/10/2025
If we care only about appearances, outcomes, and results then AI will outcompete humans at everything If we care about the process used to create things then humans can still have jobs and meaningful lives The idea that ends can be detached from means is the root of many evils
040
Nora Belrose @norabelrose.bsky.social · 30/09/2025
Strongly agree with this bill www.usatoday.com/story/news/politic…
010
Nora Belrose @norabelrose.bsky.social · 13/06/2025
if the laws of physics are fundamentally probabilistic, as they seem to be, that makes it easier to see how they can smoothly change over time
020
Nora Belrose @norabelrose.bsky.social · 12/06/2025
data attribution is a special case of data causality: estimating the causal effect of either learning or unlearning one datapoint (or set of datapoints) on the neural network's behavior on other datapoints
030
Nora Belrose @norabelrose.bsky.social · 27/03/2025
Neural networks don't have organs. They aren't made of fixed mechanisms. They have flows of information and intensities of neural activity. They can't be organized into a set of parts with fixed functions. In the words of Gilles Deleuze, they're bodies without organs (BwO).
160
Nora Belrose @norabelrose.bsky.social · 13/03/2025
This seems like a cool way to use an adaptive amount of compute per token. I speculate that models like these will have more faithful CoT since they don't get to do "extra" reasoning on easy tokens arxiv.org/abs/2404.02258
arxiv.org
Mixture-of-Depths: Dynamically allocating compute in...
Transformer-based language models spread FLOPs uniformly across input sequences. In this work we demonstrate that transformers can instead learn to dynamically allocate FLOPs (or compute) to...
140
Nora Belrose @norabelrose.bsky.social · 24/02/2025
Also chapter 10 where he discards the notion of the Soul but maintains the distinction between mind and brain
000
Nora Belrose @norabelrose.bsky.social · 24/02/2025
William James did a lot of good philosophy of mind in chapters 1, 5, and 6 ofThe Principles of Psychology, we've barely made any progress in 135 years 😂
030
Nora Belrose @norabelrose.bsky.social · 22/02/2025
I love this meme
060
Nora Belrose @norabelrose.bsky.social · 07/02/2025
might interest @nabla_theta
020
Nora Belrose @norabelrose.bsky.social · 06/02/2025
Pro tip: if you want to implement TopK SAEs efficiently, and don't want to deal with Triton, just use this function for the decoder, it's much faster than the naive dense matmul implementation pytorch.org/docs/stable/generated/t…
080
Nora Belrose @norabelrose.bsky.social · 03/02/2025
What are the chances you'd get a fully functional language model by randomly guessing the weights? We crunched the numbers and here's the answer:
2120
Nora Belrose @norabelrose.bsky.social · 02/02/2025
we have seven (!) papers lined up for release next week you know you're on a roll when arxiv throttles you
040
Nora Belrose @norabelrose.bsky.social · 24/01/2025
deepseek now largely replacing chatgpt for me
060
Nora Belrose @norabelrose.bsky.social · 29/12/2024
Evolutionary biology can learn things from machine learning. Natural selection alone doesn't explain "train-test" or "sim-to-real" generalization, which clearly happens. At every level of organization, life can zero-shot adapt to novel situations. www.youtube.com/watch?v=jJ9O5H2AlWg
290
Nora Belrose @norabelrose.bsky.social · 28/12/2024
Truth is relative, when it comes to the physical state of the universe. But we should accept the existence of perspective-neutral facts about how perspectives relate to one another, to avoid vicious skeptical paradoxes. arxiv.org/abs/2410.13819
060
Nora Belrose @norabelrose.bsky.social · 28/12/2024
Neural networks are polycomputers in @drmichaellevin.bsky.social's sense. Depending on your perspective, you can interpret them as performing many different computations on different types of features. No perspective is uniquely correct. arxiv.org/abs/2212.10675
arxiv.org
There's Plenty of Room Right Here: Biological Systems as Evolved, Overloaded, Multi-scale Machines
The applicability of computational models to the biological world is an active topic of debate. We argue that a useful path forward results from abandoning hard boundaries between categories and adopt...
0100
Nora Belrose @norabelrose.bsky.social · 20/12/2024
If OpenAI's new o3 model is "successfully aligned," then it could probably be trusted to supervise more powerful models, allowing us to bootstrap to benevolent superintelligence.
021
Nora Belrose @norabelrose.bsky.social · 20/12/2024
Interesting to see @philipgoff.bsky.social go back and forth on the fine-tuning argument. I think the multiverse definitely can't explain fine-tuning, but it's also unclear we need an explanation at all. And God may be a more "complex" hypothesis than the physical constants themselves.
philipgoff.substack.com
My Week Without Cosmic Hope
(Photo by Tom Pumford on Unsplash)
110
Nora Belrose @norabelrose.bsky.social · 11/12/2024
How do a neural network's final parameters depend on its initial ones? In this new paper, we answer this question by analyzing the training Jacobian, the matrix of derivatives of the final parameters with respect to the initial parameters. arxiv.org/abs/2412.07003
49219
Nora Belrose @norabelrose.bsky.social · 10/12/2024
Bombshell new paper on the simulation argument, multiverses, and cosmological fine-tuning: "...self-locating credences are ‘subjective’ in the sense that they are not rationally constrained by anything at all, except possibly the requirement of probabilistic consistency." arxiv.org/abs/2409.05259
arxiv.org
Against Self-Location
I distinguish between pure self-locating credences and superficially self-locating credences, and argue that there is never any rationally compelling way to assign pure self-locating credences. I firs...
040
Nora Belrose @norabelrose.bsky.social · 01/12/2024
Death is just amnesia www.youtube.com/watch?v=L3MA...
youtube.com
Alan Watts - What happens after Death (Lecture)
YouTube video by theJourneyofPurpose TJOP
240
Nora Belrose @norabelrose.bsky.social · 25/11/2024
I'm trying out Bluesky, using this browser extension to crosspost github.com/59de44955ebd/twitter-to-… My username is the obvious one.
github.com
GitHub - 59de44955ebd/twitter-to-bsky: Crosspost from Twitter/X to Bluesky and Mastodon directly in...
Crosspost from Twitter/X to Bluesky and Mastodon directly in the web browser - 59de44955ebd/twitter-to-bsky
060