Sign in

🐘

@pkydrm.bsky.social
1.3K followers 83 following 38 posts

research scientist @MosaicML x @Databricks re: rlhf, humans in the loop, and figuring out what it means to have a good model 🤖🧑‍🎨✨

PostsRepliesMedia
Reposted by 🐘
Maria Antoniak @mariaa.bsky.social · 23/07/2025
What are your favorite recent papers on using LMs for annotation (especially in a loop with human annotators), synthetic data for task-specific prediction, active learning, and similar? Looking for practical methods for settings where human annotations are costly. A few examples in thread ↴
137923
🐘 @pkydrm.bsky.social · 21/07/2025
000
Reposted by 🐘
Neil Renic @ncrenic.bsky.social · 18/06/2025
I am once again pitching my romantic comedy: - two academics start dating - discover they are each other's terrible reviewer - hijinks ensue Working title: Love is Double-Blind
942595341
🐘 @pkydrm.bsky.social · 16/04/2025
I'm extremely curious -- would you want digital tools that would help with this (e.g. planning, time organization) or embodied AI (e.g. physical assistance in-home, transportation)?
100
🐘 @pkydrm.bsky.social · 09/04/2025
i wish i could shout this from the rooftops. relatedly, there's no need for robots to be limited by the human form. similar/tangential thing came up in the 2010s with respect to self-driving: just because people only sense using their eyes doesn't mean cars have to only use cameras!
050
🐘 @pkydrm.bsky.social · 25/03/2025
we are living in an empirical world and we are empirical girls
010
🐘 @pkydrm.bsky.social · 25/03/2025
No labels, no problem! I am so excited for this release. We have been working on it for many months, and it's motivated by a common customer roadblock: insufficient labeled examples.
010
🐘 @pkydrm.bsky.social · 22/01/2025
has anyone successfully gotten very involved with their local library system and, if so, how does one do so? i know there are volunteer opportunities and it is my dream to one day organize a crafting circle, but i'm talking about how the library actually organizes / functions / prioritizes things!
010
🐘 @pkydrm.bsky.social · 19/12/2024
@jfrankle.com @ericajiyuen.bsky.social
000
🐘 @pkydrm.bsky.social · 19/12/2024
and a big shout out to my collaborators: Erica Ji Yuen, Kartik Sreenivasan, Yue (Andy) Zhang, Sam Havens, Michael Carbin, Matei Zaharia, Jonathan Frankle
000
🐘 @pkydrm.bsky.social · 19/12/2024
3/3 🔑 Want to see how different models perform on enterprise tasks? Full analysis in the blog here: databricks.com/blog/benchma...!
databricks.com
Benchmarking Domain Intelligence
000
🐘 @pkydrm.bsky.social · 19/12/2024
📊 DIBS measures real enterprise needs. We tested 14 models & found: - Academic benchmarks mask enterprise gaps - No single model wins across all tasks - Open models are competitive on key capabilities - Some enterprise tasks show clear paths forward, others are more complex 2/3
000
🐘 @pkydrm.bsky.social · 19/12/2024
🧵 Super proud to finally share this work I led last quarter - the @databricks.bsky.social Domain Intelligence Benchmark Suite (DIBS)! TL;DR: Academic benchmarks ≠ real performance and domain intelligence > general capabilities for enterprise tasks. 1/3
454
🐘 @pkydrm.bsky.social · 19/12/2024
@jfrankle.com @ericajiyuen.bsky.social
000
🐘 @pkydrm.bsky.social · 19/12/2024
And of course a big shout out to my collaborators: Erica Ji Yuen, Kartik Sreenivasan, Yue (Andy) Zhang, Sam Havens, Michael Carbin, Matei Zaharia, and Jonathan Frankle for their help!
000
🐘 @pkydrm.bsky.social · 19/12/2024
3/3 🔑 Want to see how different models perform on enterprise tasks? Full analysis in the blog here: databricks.com/blog/benchma...!
databricks.com
Benchmarking Domain Intelligence
010
🐘 @pkydrm.bsky.social · 19/12/2024
📊 DIBS measures real enterprise needs. We tested 14 models & found: - Academic benchmarks mask enterprise gaps - No single model wins across all tasks - Open models are competitive on key capabilities - Some enterprise tasks show clear paths forward, others are more complex 2/3
000
🐘 @pkydrm.bsky.social · 12/12/2024
very demure, very mindful, very 2019-era mujoco humanoid learning to walk
010
🐘 @pkydrm.bsky.social · 12/12/2024
"technology built to address people's needs" is the north star. side note: it would be amazing to see this attitude in the physical, embodied world as well. it's amazing to see how older adults in dense, walkable areas have such different lifestyles than those in car-centric suburbs.
001
🐘 @pkydrm.bsky.social · 11/12/2024
would love to be added :-)
100
🐘 @pkydrm.bsky.social · 10/12/2024
brat tulu is amazing
100
🐘 @pkydrm.bsky.social · 05/12/2024
this is incredible research, and beautiful. would love to know more about what it's like to meaningfully interact with genie 2, or similar models, e.g. to modify the outputs of such a model in the service of a design vision.
000
Reposted by 🐘
Preeti Chhibber @runwithskizzers.bsky.social · 24/11/2024
A Venn diagram where one circle reads: “a romance novel using the miscommunication trope” 

The second circle reads: “this article about using ChatGPT to find gifts for your fam” 

And the overlap between them reads, “me yelling: just have a conversation!!”
191178311
🐘 @pkydrm.bsky.social · 26/11/2024
i know some labs are already starting to do this; i hope more continue to. it is challenging, complex technical work and we should think of it as a first-class contribution in the field. 5/5
000
🐘 @pkydrm.bsky.social · 26/11/2024
🤞 we can start to more broadly value thoughtful, direction-setting benchmark work. it requires technical contributions, a keen sense of how people might meaningfully interact with a system, and the discernment to recognize where progress might yet be made. 4/5
100
🐘 @pkydrm.bsky.social · 26/11/2024
i think as a field, we have a problematic tendency to focus on magnitude-related problems, like new architectures or training paradigms or other ways to maximize performance on whatever benchmarks we can. maybe this is because it is more akin to the training/experience many of us have. 3/5
100
🐘 @pkydrm.bsky.social · 26/11/2024
in the LLM space, at this time, benchmarks/evaluations set the direction of that vector. it's extremely hard to make good benchmarks, and historically under-rewarded in the field. 2/5
100
🐘 @pkydrm.bsky.social · 26/11/2024
i often talk about the importance of aligning both the magnitude AND direction of a workstream vector. 1/5
111
🐘 @pkydrm.bsky.social · 22/11/2024
i do not study this, but i did just finish reading the anxious generation and so i'm very grateful that there are so many people who do indeed study such important things!
000
Reposted by 🐘
Cody Blakeney ✈️ NeurIPS 2024 @codestar.bsky.social · 22/11/2024
When you fail to parse your data that’s a jsonl
173
🐘 @pkydrm.bsky.social · 19/11/2024
😙🤌
000
🐘 @pkydrm.bsky.social · 19/11/2024
could you add me to it? :-)
000
🐘 @pkydrm.bsky.social · 19/11/2024
👋 would love to be added!
000
🐘 @pkydrm.bsky.social · 19/11/2024
👋 would love to be added! thanks for curating this list! :-)
010
🐘 @pkydrm.bsky.social · 19/11/2024
👋 thank you for doing this! it's nice to know who else is around
100
🐘 @pkydrm.bsky.social · 18/11/2024
that's just a growth mindset!
010
🐘 @pkydrm.bsky.social · 16/11/2024
All this, and also it's unclear to me that there's such a thing as "PhD-level intelligence" at all because so many other factors come into play: opportunity, curiosity, tenacity, etc. Also, different PhDs require different combinations of skills and strengths so 🤷‍♂️.
110
🐘 @pkydrm.bsky.social · 16/11/2024
thank you for making this! would love to be added -- have a background in human-robot collaboration and human in the loop learning :-) excited to meet more folks in the space!
100
🐘 @pkydrm.bsky.social · 15/11/2024
Thank you for creating this list -- super helpful! Would you be able to add me to it?
010
🐘 @pkydrm.bsky.social · 15/11/2024
chatgpt is giving a trunk fusing into a monitor screen and an extra leg (!!) tbf eventually it produced an elephant floppy disk, which is VERY cute but i feel no attachment to it bc it required barely any input from me claude feels more like a ✨ design medium ✨ -- those ears were a labor of love!
000
🐘 @pkydrm.bsky.social · 15/11/2024
svg claude is giving ms paint in the best way
100
🐘 @pkydrm.bsky.social · 15/11/2024
Thanks so much for these, Michael! Would you be able to add me to these lists? I'm working on RLHF at MosaicML x Databricks :-)
000
Reposted by 🐘
𝚖𝚘𝚘𝚍𝚋𝚘𝚊𝚛𝚍. @moodboard.bsky.social · 13/11/2024
#moodboard
594494216