Sign in

tomgoldstein.bsky.social

@tomgoldstein.bsky.social
284 followers 23 following 10 posts
PostsRepliesMedia
tomgoldstein.bsky.social @tomgoldstein.bsky.social · 10/02/2025
Nowadays ML projects feel like they need to be compressed into a few months. Its refreshing to be able to work on something for a few years! But also a slog.
010
tomgoldstein.bsky.social @tomgoldstein.bsky.social · 10/02/2025
New open source reasoning model! Huginn-3.5B reasons implicitly in latent space 🧠 Unlike O1 and R1, latent reasoning doesn’t need special chain-of-thought training data, and doesn't produce extra CoT tokens at test time. We trained on 800B tokens 👇
1124
Reposted by @tomgoldstein.bsky.social
Sander Dieleman @sedielem.bsky.social · 22/01/2025
📢PSA: #NeurIPS2024 recordings are now publicly available! The workshops always have tons of interesting things on at once, so the FOMO is real😵‍💫 Luckily it's all recorded, so I've been catching up on what I missed. Thread below with some personal highlights🧵
112833
tomgoldstein.bsky.social @tomgoldstein.bsky.social · 29/01/2025
Let’s sanity check DeepSeek’s claim to train on 2048 GPUs for under 2 months, for a cost of $5.6M. It sort of checks out and sort of doesn't. The v3 model is an MoE with 37B (out of 671B) active parameters. Let's compare to the cost of a 34B dense model. 🧵
1112