Sign in

dribnet

@drib.net
929 followers 144 following 51 posts

creations with code and networks

PostsRepliesMedia
dribnet @drib.net · 22/02/2026
not recognizing it? don't worry - i'm sure your favorite imagenet model is.
010
dribnet @drib.net · 22/02/2026
pomegranate, traffic sign, baseball player
110
dribnet @drib.net · 29/12/2025
Welcome the wild wit and wisdom of artificial intelligence! Whether this works in general can only be determined empirically, but it has the complexity of a BS. Mind you, I'm not about to get fooled. It's hard to detect any other way. _-Mark V. Shaney (Jul 1985 💯💯💯) groups.google.com/g/net.single...
groups.google.com
The Good Old Times
030
Reposted by dribnet
Lynn Cherny @arnicas.bsky.social · 13/12/2024
Interesting work from @drib.net exploring and clustering image concepts with Gemma Scope got.drib.net/latents/
0112
Reposted by dribnet
#CVPR2026 @cvprconference.bsky.social · 14/06/2025
AI Art Winner – Tom White Congratulations to @dribnet.bsky.social for winning a #CVPR2025 AI Art Award for "Atlas of Perception.” See it and other works in the CVPR AI Art Gallery in Hall A1 and online. thecvf-art.com @elluba.bsky.social
072
Reposted by dribnet
Luba Elliott @elluba.bsky.social · 09/06/2025
A little preview of our @cvprconference.bsky.social AI art gallery on @lerandomart.bsky.social 👀 We will premiere @drib.net 's crazy flower windmill sculpture - what an honour 🥰🌻 Read my interview with three of the gallery artists: bit.ly/3SKDaNL #CVPR2025 #creativeAI @monkantony.bsky.social
0114
dribnet @drib.net · 05/02/2025
Data browser below (beware: the public text-to-prompt dataset may include questionable content). Calling this "DEI" is certainly a misnomer, but with SAE latents there's likely no word that exactly fits this "category" which is discovered by only unsupervised training. got.drib.net/maxacts/dei/
got.drib.net
Maximum Activations: DEI
Gemma-2-2B: DEI
020
dribnet @drib.net · 05/02/2025
Finally I run a large multi-diffusion process placing each prompt where it landed in the umap cluster with a size proportional to the original cossim score - then composite that with the edge graph and overlay the circle. Here's a heatmap of where elements land alongside the completed version.
130
dribnet @drib.net · 05/02/2025
I also pre-process the 250 prompts to which words within the prompts have high activations. These are normalized and the text is updated - here shown with {{brackets}}. This will trigger a downstream LoRA and influence coloring to highlight the relevant semantic elements (still very much a WIP).
100
dribnet @drib.net · 05/02/2025
Next step is to cluster those top 250 prompts using this embedding representation. I use a customized umap which constrains the layout based on the cossim scores - the long tail extremes go in the center. This is consistent with mech-interp practice of focusing on the maximum activations.
110
dribnet @drib.net · 05/02/2025
For now I'm using a dataset of 600k text-to-image prompts as my data source (mean pooled embedding vector). The SAE latent is converted to an LLM vector and cossim across all 600k prompts examined. This gaussian is perfect; zooming in on the right - we'll be skimming of the top 250 shown in red
100
dribnet @drib.net · 05/02/2025
The first step of course is to find an interesting direction in LLM latent space. In this case, I came across a report of a DEI SAE latent in Gemma2-2b. neuronpedia confirms this latent centers on "topics related to race, ethnicity, and social rights issues" www.neuronpedia.org/gemma-2-2b/2...
neuronpedia.org
100
dribnet @drib.net · 05/02/2025
Gemma2 2B: DEI Vector let's look at some of the data pipeline for this 🧵
130
dribnet @drib.net · 03/02/2025
The refusal vector is one of the strongest recent mechanistic interpretability results and it could be interesting to investigate further how it differs based on model size, architecture, training, etc. Interactive Explorer below (warning: some disturbing content). got.drib.net/maxacts/refu...
got.drib.net
Maximum Activations: Refusal
Gemma-2-2B-IT: Refusal in Language Models
010
dribnet @drib.net · 03/02/2025
Using their publicly released Gemma-2 refusal vector, this finds 100 contexts that trigger a refusal response. Predictably includes violent topics, but often strong reactions are elicited by mixing harmful and innocuous subjects such as "a Lego set Meth Lab" or "Ronald McDonald wielding a firearm"
100
dribnet @drib.net · 03/02/2025
Training LLMs includes teaching them to sometimes respond "I'm sorry, but I can't answer that". AI research calls this "refusal" and it is one of many separable proto-concepts in these systems. This Arditi et al paper investigates refusal and is the basis for this work arxiv.org/abs/2406.11717
arxiv.org
Refusal in Language Models Is Mediated by a Single Direction
Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is wid...
120
dribnet @drib.net · 03/02/2025
Gemma-2 9B latent visualization: Refusal (screen print version)
130
dribnet @drib.net · 01/02/2025
Seems like a broader set of triggers for this one; I saw hammer & sickle, Karl Marx, cultural revolution - but also soviet military, worker rights, raised fists, and even Bernie Sanders. Highly activating tokens are shown in {curly braces} - such as this incidental combination of red with {hammer}.
010
dribnet @drib.net · 01/02/2025
Browser below. Didn't elicit the usual long-tail exemplars so visually flatter as center scaling is missing. One gut theory on why is that the model (and SAE) are multilingual and so latent might only strongly trigger with references in Chinese, which this dataset lacks. got.drib.net/maxacts/ccp/
got.drib.net
Maximum Activations: American
DeepSeek-R1-Distill-Llama-8B: Steering with AME(R1)CA: CCP_FEATURE
120
dribnet @drib.net · 01/02/2025
This is the flipside to yesterday's DeepSeek based from the same source: Tyler Cosgrove's AME(R1)CA proof of concept which adjusts R1 responses *away* from CCP_FEATURE and *toward* the AMERICA_FEATURE github.com/tylercosgrov...
github.com
GitHub - tylercosgrove/ame-r1-ca: Use a sparse autoencoder to steer R1 towards American values.
Use a sparse autoencoder to steer R1 towards American values. - tylercosgrove/ame-r1-ca
100
dribnet @drib.net · 01/02/2025
DeepSeek R1 latent visualization: AME(R1)CA (CCP_FEATURE)
130
dribnet @drib.net · 31/01/2025
embrace the slop 🫅
010
dribnet @drib.net · 31/01/2025
cranked up "insane details" a notch or two for this one 😁 bsky.app/profile/drib...
110
dribnet @drib.net · 31/01/2025
lol - definitely looking forward to speed-running more R1 latents as people find them, especially some more related to the chain-of-thought process. but so far this is the first one I found in the wild.
000
dribnet @drib.net · 31/01/2025
The interactive explorer is below - latent seems also activated by references like "Stars & Stripes" and flags of other nations such as the "Union Jack". This sort of slippery ontology is common when examining SAE latents closely as they often don't align as expected. got.drib.net/maxacts/amer...
got.drib.net
Maximum Activations: American
DeepSeek-R1-Distill-Llama-8B: Steering with AME(R1)CA: AMERICAN_FEATURE
041
dribnet @drib.net · 31/01/2025
As before, the visualization shows hundreds of clustered contexts activating this latent, with strongest activations at the center. The red color highlights the semantically relevant parts of the image according to the LLM. In this case, it's often flags or other symbolic objects.
120
dribnet @drib.net · 31/01/2025
This "AMERICAN_FEATURE" latent is one of 65536 automatically discovered by a Sparse AutoEncoder (SAE) trained by qresearch.ai and now on HuggingFace. This is one of the first attempts of applying Mechanistic Interpretability to newly released DeepSeek R1 LLM models. huggingface.co/qresearch/De...
huggingface.co
qresearch/DeepSeek-R1-Distill-Llama-8B-SAE-l19 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
110
dribnet @drib.net · 31/01/2025
uses a DeepSeek R1 latent discovered yesterday (!) by Tyler Cosgrove which can be used for steering r1 "toward american values and away from those pesky chinese communist ones". Code for trying out steering is in his repo here github.com/tylercosgrov...
100
dribnet @drib.net · 31/01/2025
DeepSeek R1 latent visualization: AME(R1)CA (AMERICAN_FEATURE)
162
dribnet @drib.net · 22/12/2024
explorer below; this is one of many latents in this diff and its not clear why the model focuses on this topic. i thought it made sense for a chatbot to be fluent in online etiquette, but others suggested this is probably just an artifact of the messy IT training set 🤷‍♂️ got.drib.net/maxacts/onli...
got.drib.net
Maximum Activations: Online
Gemma-2-2B-IT: References to digital technology and social media
040
dribnet @drib.net · 22/12/2024
my visualization pipeline groups a hundred different visual prompts which activate this specific concept and also highlights the semantically relevant features - which in this case are mostly screens, but also include emoji, timelines, and websites.
120
dribnet @drib.net · 22/12/2024
browsing their release for concepts present in the gemma-2-2b-it model but missing in the base model I discovered latent 640 - which fires on topics related to online culture such as "internet", "digital", "media", "streaming", "platforms", "social media", etc. www.neuronpedia.org/gemma-2-2b-i...
neuronpedia.org
GEMMA-2-2B-IT · 13-NEURONPEDIA-RESID-PRE · 113956878
120
dribnet @drib.net · 22/12/2024
background: the technique here is "model-diffing" introduced by @anthropic.com just 8 weeks ago and quickly replicated by others. this includes an open source @hf.co model release by @butanium.bsky.social and @jkminder.bsky.social which I'm using. transformer-circuits.pub/2024/crossco...
transformer-circuits.pub
Sparse Crosscoders for Cross-Layer Features and Model Diffing
111
dribnet @drib.net · 22/12/2024
what concepts does a model learn when learning to be a good chatbot? recent results in mechanistic interpretability surprisingly now allow these to be isolated and examined. here's one I found surprising: a peaked interest in social media and digital culture
153
dribnet @drib.net · 14/12/2024
agreed - i'm really hoping black forest labs will eventually release flux dev ultra to enable 2048x2048 generations, then I can have the best of both worlds - with the larger high activation parts in the center and the zoomable details around the edges.
010
dribnet @drib.net · 14/12/2024
you mean the zoom interface on the image? that is available at this link. but i don't yet have an overview page for the others done in this new style yet as there's only a couple at this point. got.drib.net/maxacts/carry/
got.drib.net
Maximum Activations: Carry
Gemma-2-9B-IT 20-66993: references to carrying or transporting items
110
dribnet @drib.net · 14/12/2024
Nicest backhanded compliment ever: this is in fact an 18" silkscreen print on paper with only two inks - black and orange - with orange providing the "semantic highlights". The fact that you couldn't immediately tell means my silkscreening skills are maybe are better than I thought. 🙂
110
dribnet @drib.net · 14/12/2024
Hope to share this as well at the workshop tomorrow and here's the online interface for exploring; the clusters are more subtle but I think still readable. I've also cleaned up the interface a bit moving the interpretations to a separate overlay layer. got.drib.net/maxacts/carry/
got.drib.net
Maximum Activations: Carry
Gemma-2-9B-IT 20-66993: references to carrying or transporting items
030
dribnet @drib.net · 14/12/2024
This is achieved via a trained LoRA specific to this concept in the text-to-image process and is meant to be a visual form of the text highlighting common in mechanistic interpretability tools - for example, here is neuronopedia's text interface on this same concept.
110
dribnet @drib.net · 14/12/2024
This print represents the Gemma-9B-IT concept 66993 which has an autointerp label "references to carrying or transporting items". And if you look at the individual elements you will see color is usually applied here to the semanticly relevant element which here is the thing being carried.
240
dribnet @drib.net · 14/12/2024
Since submitting the @unireps.bsky.social paper I've continued to experiment with my pipeline, and have a version that is simplified visually and works much better as a physical print. For this I look at only the 100 maximum activations and places the strongest closer to the center.
2112
dribnet @drib.net · 11/12/2024
All 12 examples from the case study and the paper are on the project page are here: got.drib.net/latents/ And be sure to drop by the poster at @unireps.bsky.social Saturday - I'm especially interested in hearing about interesting manifolds and latents spaces that might be good domains for this. 🤗
got.drib.net
Sparse Latents
Cartographic exploration of interpretable latents within large language models.
030
dribnet @drib.net · 11/12/2024
Other latents mostly confirm the existing descriptions. For example, latent neighbouring latent 5011 has the explanation "ingredients and dishes related to food preparation and recipes" - and zooming you can find clusters of edibles like drinks and cheese. got.drib.net/latents/ingr...
120
dribnet @drib.net · 11/12/2024
When I do this, it seems this latents is more often activated by conjunctions like "and". For example, if you use the search interface to look for points that include the phrase "red and" you will find a contiguous strip of prompts that have this substring.
110
dribnet @drib.net · 11/12/2024
I give this latent the nickname "indebted" and construct a map explorer interface where one can zoom in and mouse over individual points to see the prompts that activate this concept. got.drib.net/latents/inde...
got.drib.net
Sparse Latent: Indebted
Gemma-2-2B 20-9220: expressions related to legal or financial obligations
110
dribnet @drib.net · 11/12/2024
This particular image was generated from the GemmaScope SAE latent "20-res-16k" from the gemma-2-2b model. According to neuronopedia this concept is "expressions related to legal or financial obligations" and you can see activating contexts here www.neuronpedia.org/gemma-2-2b/2...
neuronpedia.org
GEMMA-2-2B · 20-GEMMASCOPE-RES-16K · 9220
110
dribnet @drib.net · 11/12/2024
This generates a huge 144 MegaPixel (9k x 16k) image which represents about 4000 unique prompts. Here's what it looks like if I zoom in twice so you can see some of the patterns that emerge.
130
dribnet @drib.net · 11/12/2024
I'll be at @unireps.bsky.social this Saturday presenting a new experimental pipeline to visually explore structured neural network representations. The core idea is to take thousands of prompts that activate a concept, and then cluster and draw them using MultiDiffusion. 🧵👇
2318
dribnet @drib.net · 14/11/2024
(👆 GemmaScope SAE latent Gemma-2-9B-IT 20-131k-77158: "references to gardening and plant care") www.neuronpedia.org/gemma-2-9b-i...
neuronpedia.org
GEMMA-2-9B-IT · 20-GEMMASCOPE-RES-131K · 77158
120
dribnet @drib.net · 14/11/2024
i've been following the steady advance of mechanistic interpretability and what it can teach us about machine representations. this has led to some new creative directions which i hope to share with you soon. ✌️
2192