PyTorch’s latest blog covers Jagged Flash Attention on NVIDIA B200 with TLX. On GEM’s jagged shapes, the TLX kernel is about 13% faster on the forward pass and about 50% faster on the backward pass than the May 2026 version of FA4. bit.ly/4dhV84P
Tensors and neural networks in Python with strong hardware acceleration. PyTorch is an open source project at the Linux Foundation. #PyTorchFoundation pytorch.org