Birchlabs @birchlabs.co.uk · 30/01/2025pytorch 2.6 is out! highlights: - flex attention: better compilation of blockmask creation, better support for dynamic shapes - cuDNN SDPA: fixes for memory layout - CUDA 12.6 - python 3.13 - MaskedTensor memory leak fix 021
Birchlabs @birchlabs.co.uk · 04/01/2025drink cups should put the hole in the bottom. heat rises. "the top is cool enough to drink" implies "everything below it is colder". drinking from the bottom lets us access safe temperatures earlier and before the whole cup cools. 030
Birchlabs @birchlabs.co.uk · 22/12/2024when the standard library comments out std::experimental::observer_ptr just to stop you having fun 000
Birchlabs @birchlabs.co.uk · 16/12/2024if you care about multiprocess debugging in VSCode please upvote this issue so we don't have to click terminate a hundred times github.com/microsoft/vs... 030
Birchlabs @birchlabs.co.uk · 30/11/2024incidentally the inductor problem is here. it could be fixed by omitting the guard altogether, or perhaps eliding the function grid wrapper entirely github.com/pytorch/pyto... 010
Birchlabs @birchlabs.co.uk · 30/11/2024if you don't do this, then inductor codegen emits invalid python 120
Birchlabs @birchlabs.co.uk · 30/11/2024mood: adding unused variables to make torch inductor compile my triton kernel 2140
Birchlabs @birchlabs.co.uk · 23/11/2024born too late to learn maths from touhou pre-fight cutscenes www.youtube.com/watch?v=tuDA... 270
Birchlabs @birchlabs.co.uk · 21/11/2024if you get KO'd in smash, do you die? what's the safest stage to be KO'd on? Great Bay looks alright if you're a confident swimmer… 110
Birchlabs @birchlabs.co.uk · 17/11/2024pytorch 2.5.0 bug: counting flops makes your compiled model slower. benchmark your model first, count flops after. github.com/pytorch/pyto... 000
Birchlabs @birchlabs.co.uk · 04/11/2024okay where has tensor_split been my whole life pytorch.org/docs/stable/... 010
Birchlabs @birchlabs.co.uk · 02/11/2024that still leaves you with another problem: attention entropy. if you're inferencing smaller, the probability distribution is sharper than in training. you can scale the logits to compensate. left = orig right = entropy-scaled. arxiv.org/abs/2306.08645 020
Birchlabs @birchlabs.co.uk · 02/11/2024you can fight back by dilating the convolution! arxiv.org/abs/2310.07702 110
Birchlabs @birchlabs.co.uk · 02/11/2024the reason stable-diffusion generalizes poorly to untrained resolutions is its implicit position embedding. convolution padding creates an edge, which nested convolutions look for to understand position. arxiv.org/abs/2101.12322 120