x.com
arXiv Bangers (@arXivBangers) on X
Looped Diffusion Transformer Looped-DiT runs shared Transformer blocks repeatedly in each denoising step, allowing a 260M model to outperform a 6.5x larger model on text-to-image benchmarks while using 4.9x lower inference compute. https://t.co/QI2nhAtvp1