Sign in

Alex Dimakis

@alexdimakis.bsky.social
1K followers 90 following 23 posts

UC Berkeley Professor working on AI. Co-Director: National AI Institute on the Foundations of Machine Learning (IFML). BespokeLabs.ai cofounder

PostsRepliesMedia
Reposted by Alex Dimakis
Kosta Derpanis @csprofkgd.bsky.social · 11/04/2025
Confirmed keynote speakers for the Greeks 🇬🇷 in #AI 2025 Symposium in Athens, Greece 🤗 @alexdimakis.bsky.social @manoliskellis.bsky.social www.greeksin.ai
0101
Alex Dimakis @alexdimakis.bsky.social · 13/04/2025
Excited to be part of Greeksin.ai
041
Alex Dimakis @alexdimakis.bsky.social · 03/04/2025
We are excited to release the OpenThinker2 reasoning models and data: 1. Outperforms DeepSeekR1-32B in reasoning. 2. Fully open source, open weights and open data (1M samples). 3. Post-trained only with SFT. RL post-training will likely further improve performance. github.com/open-thought...
github.com
GitHub - open-thoughts/open-thoughts: Fully open data curation for reasoning models
Fully open data curation for reasoning models. Contribute to open-thoughts/open-thoughts development by creating an account on GitHub.
0123
Alex Dimakis @alexdimakis.bsky.social · 12/02/2025
We are releasing OpenThinker-32B, the best 32B reasoning model with open data. We match or outperform Deepseek-R1-32B (a closed data model) in reasoning benchmarks. Congrats to Negin and the whole Open Thoughts team. github.com/open-thought...
Performance of the best known Reasoning models on various Benchmarks. OpenThinker-32B matches the current state of the art.
2258
Alex Dimakis @alexdimakis.bsky.social · 28/01/2025
What if we had the data that DeepSeek-R1 was post-trained on? We announce Open Thoughts, an effort to create such open reasoning datasets. Using our data we trained Open Thinker 7B an open data model with performance very close to DeepSeekR1-7B distill. (1/n)
191
Alex Dimakis @alexdimakis.bsky.social · 22/01/2025
We just did a crazy 48h sprint to create the best public reasoning dataset using Berkeley's Sky-T1, Curator and DeepSeek R1. We can get o1-Preview reasoning on a 32B model and 48x less data than Deepseek. t.co/WO5UV2LZQM
t.co
https://www.bespokelabs.ai/blog/bespoke-stratos-the-unreasonable-effectiveness-of-reasoning-distillation
162
Alex Dimakis @alexdimakis.bsky.social · 14/01/2025
The Berkeley Sky computing lab just trained a GPT-o1 level reasoning model, spending only $450 to create the instruction dataset. The data is 17K math and coding problems solved step by step. They created this dataset by prompting QwQ at $450 cost. Q: Impossible without distilling a bigger model?
171
Alex Dimakis @alexdimakis.bsky.social · 08/01/2025
AI monoliths vs Unix Philosophy: The case for small specialized AI models. Current thinking is that AGI is coming, and one gigantic model will be able to solve everything. Current Agents are mostly prompts on one big model and prompt engineering is used for executing complex processes. (1/n)
192
Reposted by Alex Dimakis
Atula Tejaswi @atutej.bsky.social · 09/12/2024
Missed out on #Swift tickets? No worries—swing by our #SVFT poster at #NeurIPS2024 and catch *real* headliners! 🎤💃🕺 📌Where: East Exhibit Hall A-C #2207, Poster Session 4 East ⏲️When: Thu 12 Dec, 4:30 PM - 7:30 PM PST #AI #MachineLearning #PEFT #NeurIPS24
192
Alex Dimakis @alexdimakis.bsky.social · 19/11/2024
hello friends, I heard there is a party here?
270