Sign in

Birchlabs

@birchlabs.co.uk
496 followers 34 following 61 posts

ML Engineer at Anlatan (NovelAI). co-author of HDiT (Hourglass Diffusion Transformers). works on diffusion models and LLMs. 日本語を勉強してる。

PostsRepliesMedia
Birchlabs @birchlabs.co.uk · 24/04/2025
pytorch 2.7 is out! - Mega Cache looks nice, unclear whether those modules get cached by legacy mechanisms too - foreach map looks good for optimizers - trainable biases means you can now train T5 on flex - prologue fusion is hype - context parallel brings ring attention pytorch.org/blog/pytorch...
020
Birchlabs @birchlabs.co.uk · 30/01/2025
pytorch 2.6 is out! highlights: - flex attention: better compilation of blockmask creation, better support for dynamic shapes - cuDNN SDPA: fixes for memory layout - CUDA 12.6 - python 3.13 - MaskedTensor memory leak fix
021
Birchlabs @birchlabs.co.uk · 20/01/2025
pytorch 2.6 final RC is out, promoting to stable in a couple of days! mostly I'm looking forward to better compilation of flex block mask creation, and better support for flex attention on dynamic shapes. there's also fixes for memory layout in cuDNN SDPA. dev-discuss.pytorch.org/t/pytorch-re...
040
Birchlabs @birchlabs.co.uk · 14/01/2025
Claude does SVG memes
020
Birchlabs @birchlabs.co.uk · 04/01/2025
drink cups should put the hole in the bottom. heat rises. "the top is cool enough to drink" implies "everything below it is colder". drinking from the bottom lets us access safe temperatures earlier and before the whole cup cools.
030
Birchlabs @birchlabs.co.uk · 04/01/2025
running npm version from a subdirectory of a git repository is literally an unsolved problem in 2024 github.com/npm/cli/issu...
010
Birchlabs @birchlabs.co.uk · 01/01/2025
I should just get this tattooed, I never remember how to find it pip install huggingface_hub[hf_transfer] HF_HUB_ENABLE_HF_TRANSFER=1 huggingface-cli download
020
Birchlabs @birchlabs.co.uk · 22/12/2024
when the standard library comments out std::experimental::observer_ptr just to stop you having fun
000
Birchlabs @birchlabs.co.uk · 21/12/2024
NovelAI v4 makes dreams come true
040
Birchlabs @birchlabs.co.uk · 19/12/2024
when you're measuring torch compile warmup "oh 11 secs that's not so bad" then you realize it was 111 secs
020
Birchlabs @birchlabs.co.uk · 17/12/2024
Claude's alright
020
Birchlabs @birchlabs.co.uk · 16/12/2024
if you care about multiprocess debugging in VSCode please upvote this issue so we don't have to click terminate a hundred times github.com/microsoft/vs...
030
Birchlabs @birchlabs.co.uk · 14/12/2024
Meta releases flow-matching code github.com/facebookrese...
github.com
GitHub - facebookresearch/flow_matching: A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples fo...
A PyTorch library for implementing flow matching algorithms, featuring continuous and discrete flow matching implementations. It includes practical examples for both text and image modalities. - fa...
2133
Birchlabs @birchlabs.co.uk · 13/12/2024
the EDM2 repository had an Autoguidance update 4 days ago demonstrating their NeurIPS Oral paper "Guiding a Diffusion Model with a Bad Version of Itself" github.com/NVlabs/edm2
github.com
GitHub - NVlabs/edm2: EDM2 and Autoguidance -- Official PyTorch implementation
EDM2 and Autoguidance -- Official PyTorch implementation - NVlabs/edm2
030
Birchlabs @birchlabs.co.uk · 12/12/2024
torch profiler record_function spans are not free, even when you're not profiling a model. the model I'm benchmarking trained 2.5% faster when I commented them all out.
020
Birchlabs @birchlabs.co.uk · 10/12/2024
000
Birchlabs @birchlabs.co.uk · 09/12/2024
Box2D 3.0.0 in WebAssembly+TypeScript starting to work
110
Reposted by Birchlabs
Nick Stracke @rmsnorm.bsky.social · 04/12/2024
🤔 Why do we extract diffusion features from noisy images? Isn’t that destroying information? Yes, it is - but we found a way to do better. 🚀 Here’s how we unlock better features, no noise, no hassle. 📝 Project Page: compvis.github.io/cleandift 💻 Code: github.com/CompVis/clea... 🧵👇
24210
Birchlabs @birchlabs.co.uk · 04/12/2024
torch.compile is hard for dynamic shapes / large number of static shapes, and non-transformer architectures. I measure suites of shapes, log recompiles, check which require warmup. operation compile competitive with whole-model compile. some operations prefer compiler disabled. dynamic often slow.
010
Reposted by Birchlabs
ruiqigao.bsky.social @ruiqigao.bsky.social · 02/12/2024
A common question nowadays: Which is better, diffusion or flow matching? 🤔 Our answer: They’re two sides of the same coin. We wrote a blog post to show how diffusion models and Gaussian flow matching are equivalent. That’s great: It means you can use them interchangeably.
625459
Birchlabs @birchlabs.co.uk · 30/11/2024
mood: adding unused variables to make torch inductor compile my triton kernel
triton autotune configs with empty dicts of meta-parameters cannot be compiled by inductor because it will emit an empty-string guard clause, which is not valid Python syntaxwe can work around the inductor codegen error by adding an unused meta-parameter to our triton autotune config, but triton codegen will then complain… until we allocate it also in the kernel function signature, unused in our implementation
2140
Birchlabs @birchlabs.co.uk · 28/11/2024
that feel when you’ve been cooking rice for 10 minutes but the hob wasn’t turned on
000
Birchlabs @birchlabs.co.uk · 27/11/2024
fine I'll use pytorch nightly
250
Birchlabs @birchlabs.co.uk · 23/11/2024
born too late to learn maths from touhou pre-fight cutscenes www.youtube.com/watch?v=tuDA...
270
Birchlabs @birchlabs.co.uk · 21/11/2024
if you get KO'd in smash, do you die? what's the safest stage to be KO'd on? Great Bay looks alright if you're a confident swimmer…
110
Birchlabs @birchlabs.co.uk · 21/11/2024
no_grad in the streets, inference_mode in the sheets
140
Birchlabs @birchlabs.co.uk · 20/11/2024
oh no I set edgeitems too high and now pytorch is stuck
120
Birchlabs @birchlabs.co.uk · 20/11/2024
good algorithm
020
Birchlabs @birchlabs.co.uk · 19/11/2024
NATTEN just added fused support for self-cross attention! so you can attend to local neighbourhood and registers or text condition. it lets you reduce partial attention results (e.g. logsumexp provided by xformers APIs) into its LSE. github.com/SHI-Labs/NAT...
github.com
Support for fused cross-NA by alihassanijr · Pull Request #182 · SHI-Labs/NATTEN
Adds experimental support for additional context tokens to Fused NA. Any number of partial attention results can be reduced into a final one as if their contexts were merged, which is just the same...
091
Birchlabs @birchlabs.co.uk · 17/11/2024
torch stable (even 2.5.1) has a bug with foreach operations in FSDP. AdamW optimizer is only safe in foreach=False or fused=True modes. foreach operations that take scalar operands are fine. github.com/pytorch/pyto...
020
Birchlabs @birchlabs.co.uk · 17/11/2024
pytorch 2.5.0 bug: counting flops makes your compiled model slower. benchmark your model first, count flops after. github.com/pytorch/pyto...
000
Birchlabs @birchlabs.co.uk · 11/11/2024
bluesky is trending on the other place
000
Birchlabs @birchlabs.co.uk · 11/11/2024
hol' up lemme just install this operating system real quick $ which ffmpeg ffmpeg not found $ brew install ffmpeg ==> Auto-updating Homebrew...
010
Birchlabs @birchlabs.co.uk · 09/11/2024
fuck offfff
020
Birchlabs @birchlabs.co.uk · 04/11/2024
okay where has tensor_split been my whole life pytorch.org/docs/stable/...
010
Birchlabs @birchlabs.co.uk · 03/11/2024
when it's Sunday and you realise you've been sitting on your breakpoint since Friday
010
Birchlabs @birchlabs.co.uk · 02/11/2024
the reason stable-diffusion generalizes poorly to untrained resolutions is its implicit position embedding. convolution padding creates an edge, which nested convolutions look for to understand position. arxiv.org/abs/2101.12322
120
Birchlabs @birchlabs.co.uk · 02/11/2024
she just showed up and asked where I keep the grimoires
1290