Sign in

BenjMurrell

@benjmurrell.bsky.social
803 followers 3.9K following 83 posts

🇸🇪🇿🇦 Researcher at Karolinska. Comp bio. Phylogenetics. Deep learning. All in Julia. Virology. Immunology. scholar.google.com/citations?user=I…

PostsRepliesMedia
BenjMurrell @benjmurrell.bsky.social · 21/11/2025
A technical thread on loss scaling in diffusion and flow matching models (related to a new preprint): Since the dawn of time, people have been messing with (or dropping entirely) these pesky time-dependent loss scaling terms, mostly because the models train better without them.
141
BenjMurrell @benjmurrell.bsky.social · 10/11/2025
We figured out flow matching over states that change dimension. With "Branching Flows", the model decides how big things must be! This works wherever flow matching works, with discrete, continuous, and manifold states. We think this will unlock some genuinely new capabilities.
42413
BenjMurrell @benjmurrell.bsky.social · 21/06/2025
One day remaining before this closes! bsky.app/profile/benj...
061
BenjMurrell @benjmurrell.bsky.social · 13/06/2025
We tried to set up a simple demo/tutorial model for the protein design ecosystem we've been developing, and it turned out a bit more interesting than we expected. 🧵 This was a team effort from a few people in my lab, including @antonoresten.bsky.social and others (not sure who is on this app)
2134
BenjMurrell @benjmurrell.bsky.social · 15/05/2025
My lab, at Karolinska, in Stockholm, is looking for a PhD student with a computational/quantitative background to work on probabilistic/generative models of proteins (structure and sequence). The research will involve methods development, and applications in vaccine design.
22113
BenjMurrell @benjmurrell.bsky.social · 15/05/2025
022
BenjMurrell @benjmurrell.bsky.social · 15/05/2025
220
BenjMurrell @benjmurrell.bsky.social · 09/12/2024
Transformer/attention folks: I saw massive activations in Qwen's keys, always towards the end of each head, especially in layer 1. Turns out this is directly driven by the key projection bias (which Qwen has but eg. Llama3 does not). These large values are where RoPE has the slowest(?) effect. Why?
Attention key bias, layer 1.
250