Sign in

Conor Durkan

@conormdurkan.bsky.social
306 followers 33 following 15 posts

Generative modeling person conordurkan.com

PostsRepliesMedia
Conor Durkan @conormdurkan.bsky.social · 18/02/2025
I'm rejoining GDM in NYC this week, looking forward to catching up with folks :)
030
Conor Durkan @conormdurkan.bsky.social · 27/01/2025
More than anything, the R1 model and paper make me wonder how far along we'd be if everyone was still clamoring to shout their best ideas from the rooftops. I know it's naïve to think we could sustain open research of the kind we saw up to 2020 indefinitely, but still...
010
Conor Durkan @conormdurkan.bsky.social · 21/01/2025
I wrote up some thoughts having worked on generative music over the past year. I talk a bit about where tech is now, how it's being used, and where it might go. pxgy.substack.com/p/thoughts-o...
pxgy.substack.com
Thoughts on generative music
Observations from a year building generative music tech
011
Conor Durkan @conormdurkan.bsky.social · 07/01/2025
First book of the year and it was a doozy. I think I have a soft spot for aerospace engineering--I liked this almost as much as 'Carrying the Fire'.
000
Conor Durkan @conormdurkan.bsky.social · 07/01/2025
Trying out NNX having previously switched from Haiku to Linen. Something something fool me once, something something I'll probably just get fooled again.
000
Conor Durkan @conormdurkan.bsky.social · 03/01/2025
t.co
https://arxiv.org/abs/1805.00909
000
Conor Durkan @conormdurkan.bsky.social · 03/01/2025
t.co
https://arxiv.org/abs/2205.11275
100
Conor Durkan @conormdurkan.bsky.social · 03/01/2025
100
Conor Durkan @conormdurkan.bsky.social · 03/01/2025
This means post-training (of this kind at least) optimizes KL(model || posterior), whereas pre-training optimizes KL(data || model). It also means post-training is mode-seeking (as opposed to mode-covering like pre-training), so those rewards better be well calibrated.
100
Conor Durkan @conormdurkan.bsky.social · 03/01/2025
Reward functions are log-likelihoods, and the pre-trained model is a prior. The posterior target is the product of the likelihoods and prior (the prior weighting can equivalently sharpen or smooth your likelihoods). Rewards can be hard for math/code verification, or soft for subjective preference.
100
Conor Durkan @conormdurkan.bsky.social · 03/01/2025
I like the Bayesian framing of reward-based post-training (i.e. reward-maximization with a KL penalty). (Figure from 'RL with KL penalties is better viewed as Bayesian inference', link below along with other useful references)
142
Conor Durkan @conormdurkan.bsky.social · 06/12/2024
If there was a product that combined unlimited Advanced Voice + Sonnet 3.6 + arbitrary document context for reference, I’d probably pay $1000 per month. 24/7 yapping
010
Conor Durkan @conormdurkan.bsky.social · 26/11/2024
Ironically on my wrist
110
Conor Durkan @conormdurkan.bsky.social · 26/11/2024
Electronic devices that don’t run out of battery, who is building this
120
Conor Durkan @conormdurkan.bsky.social · 19/11/2024
I like posts on Twitter and then I come here to see if the person has posted the same thing so I can like it all over again
050