Conor Durkan @conormdurkan.bsky.social · 18/02/2025I'm rejoining GDM in NYC this week, looking forward to catching up with folks :) 030
Conor Durkan @conormdurkan.bsky.social · 27/01/2025More than anything, the R1 model and paper make me wonder how far along we'd be if everyone was still clamoring to shout their best ideas from the rooftops. I know it's naïve to think we could sustain open research of the kind we saw up to 2020 indefinitely, but still... 010
Conor Durkan @conormdurkan.bsky.social · 21/01/2025I wrote up some thoughts having worked on generative music over the past year. I talk a bit about where tech is now, how it's being used, and where it might go. pxgy.substack.com/p/thoughts-o...pxgy.substack.comThoughts on generative musicObservations from a year building generative music tech 011
Conor Durkan @conormdurkan.bsky.social · 07/01/2025First book of the year and it was a doozy. I think I have a soft spot for aerospace engineering--I liked this almost as much as 'Carrying the Fire'. 000
Conor Durkan @conormdurkan.bsky.social · 07/01/2025Trying out NNX having previously switched from Haiku to Linen. Something something fool me once, something something I'll probably just get fooled again. 000
Conor Durkan @conormdurkan.bsky.social · 03/01/2025This means post-training (of this kind at least) optimizes KL(model || posterior), whereas pre-training optimizes KL(data || model). It also means post-training is mode-seeking (as opposed to mode-covering like pre-training), so those rewards better be well calibrated. 100
Conor Durkan @conormdurkan.bsky.social · 03/01/2025Reward functions are log-likelihoods, and the pre-trained model is a prior. The posterior target is the product of the likelihoods and prior (the prior weighting can equivalently sharpen or smooth your likelihoods). Rewards can be hard for math/code verification, or soft for subjective preference. 100
Conor Durkan @conormdurkan.bsky.social · 03/01/2025I like the Bayesian framing of reward-based post-training (i.e. reward-maximization with a KL penalty). (Figure from 'RL with KL penalties is better viewed as Bayesian inference', link below along with other useful references) 142
Conor Durkan @conormdurkan.bsky.social · 06/12/2024If there was a product that combined unlimited Advanced Voice + Sonnet 3.6 + arbitrary document context for reference, I’d probably pay $1000 per month. 24/7 yapping 010
Conor Durkan @conormdurkan.bsky.social · 26/11/2024Electronic devices that don’t run out of battery, who is building this 120
Conor Durkan @conormdurkan.bsky.social · 19/11/2024I like posts on Twitter and then I come here to see if the person has posted the same thing so I can like it all over again 050