Sign in

Durk Kingma

@dpkingma.bsky.social
1.8K followers 61 following 13 posts

Research scientist at Anthropic. Prev. Google Brain/DeepMind, founding team OpenAI. Computer scientist; inventor of the VAE, Adam optimizer, and other methods. ML PhD. Website: dpkingma.com

PostsRepliesMedia
Durk Kingma @dpkingma.bsky.social · 04/12/2024
Interesting!
060
Durk Kingma @dpkingma.bsky.social · 04/12/2024
Actually, my bad, there's code: github.com/sinhasam/CRVAE
github.com
GitHub - sinhasam/CRVAE: Official code release for Consistency Regularization for VAEs
Official code release for Consistency Regularization for VAEs - sinhasam/CRVAE
050
Durk Kingma @dpkingma.bsky.social · 04/12/2024
Impressive!! I'm curious if anyone tried to reproduce that 2.51 number from CR-NVAE. It's much lower than I would suspect for a method like hat so I'm curious if it's real, afaik they didn't share code.
250
Durk Kingma @dpkingma.bsky.social · 03/12/2024
Typo: it => of.
000
Durk Kingma @dpkingma.bsky.social · 03/12/2024
Probably my biggest naming blunder was the name "auto-encoding variational Bayes" for a method that (in its default version) is only Bayesian over the latent variables, not the parameters. Blasphemy in the church it Bayes 😂
290
Durk Kingma @dpkingma.bsky.social · 03/12/2024
Yeah, but I was born in darkness. The first paper I ever cited was Aapo Hyvarinen's paper introducing score matching from 2007, in which the term is "abused" that way the first time, but he had the excuse since there were no good alternatives! Inference, however...
140
Reposted by Durk Kingma
Sander Dieleman @sedielem.bsky.social · 02/12/2024
In arxiv.org/abs/2303.00848, @dpkingma.bsky.social and @ruiqigao.bsky.social had suggested that noise augmentation could be used to make other likelihood-based models optimise perceptually weighted losses, like diffusion models do. So cool to see this working well in practice!
arxiv.org
Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation
To achieve the highest perceptual quality, state-of-the-art diffusion models are optimized with objectives that typically look very different from the maximum likelihood and the Evidence Lower Bound (...
05311
Durk Kingma @dpkingma.bsky.social · 03/12/2024
I completely agree that it's not an ideal situation that the meaning of the word inference is now overloaded, but its use in this context is now extremely widespread. Better embrace this strange new wor(l)d ;)
150
Durk Kingma @dpkingma.bsky.social · 01/12/2024
My take is that the similarity stems from (1) the true score is identical across models except for stretching of time & space, (2) the only fundamental difference between diffusion objectives is the weighting, and (3) many common weightings are fairly similar (support over a similar range of SNRs).
040
Durk Kingma @dpkingma.bsky.social · 30/11/2024
Do you mean the intuition behind the fact that flow matching with the optimal transport (FM-OT) objective = diffusion objective with exponential weighting? I should probably read the other paper you linked ;)
170
Durk Kingma @dpkingma.bsky.social · 29/11/2024
Jealous!! Enjoy!
110
Durk Kingma @dpkingma.bsky.social · 27/11/2024
No, CS. But intrigued by starlink's phased array :-)
010
Durk Kingma @dpkingma.bsky.social · 26/11/2024
It's 2024 and microwave ovens still heat food unevenly. I need a phased array microwave oven that warms food evenly, adaptively. Who is building this?
5220
Durk Kingma @dpkingma.bsky.social · 20/11/2024
I don't usually talk about politics, but if I could change one thing in the US democracy, it might be replacing winner-takes-all representation with proportional representation. Would fix a lot of issues. Agree or disagree that this would be a net positive change?
2140