Sign in

Raphaël Avalos

@raphael.avalos.fr
582 followers 288 following 11 posts

Fine-tuning LLMs @Cohere | PhD Candidate on RL @VUB

PostsRepliesMedia
Raphaël Avalos @raphael.avalos.fr · 28/03/2025
Excited to share the technical report on Command R7B (7B) and Command A (111B), our flagship model! These models are the result of incredible teamwork at @cohere.com, and it was an honor to be part of it. Report: cohere.com/research/pap...
cohere.com
010
Reposted by Raphaël Avalos
ala-workshop.bsky.social @ala-workshop.bsky.social · 24/02/2025
🚨 Less than 48 hours left to submit to the 17th Adaptive Learning Agent workshop at @AAMASconf! 🚨 We welcome full papers, work in progress, and 2-page abstracts of recent journal papers. Don't miss the deadline! 🔗 More details: ala-workshop.github.io
ala-workshop.github.io
ALA 2025
031
Raphaël Avalos @raphael.avalos.fr · 28/01/2025
Don't miss the opportunity to submit your (Multi-Agent) RL work to the ALA workshop!
020
Raphaël Avalos @raphael.avalos.fr · 09/01/2025
The BlueSky account and website for the next edition of the ALA workshop is live! Follow it to get all the updates :)
010
Raphaël Avalos @raphael.avalos.fr · 25/11/2024
I’m not sure about 1, but you could look into results on belief MDPs. For 2, consider an environment with two rooms where the agent needs to press different buttons to get the optimal reward. If there's a cheap way to determine which room the agent is in, that would be the optimal policy :)
110
Raphaël Avalos @raphael.avalos.fr · 25/11/2024
Bsky’s strength lies in being open-source and federated. This enables anyone to host servers, set moderation policies, create custom feeds, while avoiding incentives for allowing bots to survive. It’s a tough challenge, but there’s hope!
000
Raphaël Avalos @raphael.avalos.fr · 25/11/2024
IMO, UCB favors exploration, not information-seeking, as it adds an exploration bonus rather than aiming to reduce state uncertainty. However, effective exploration can uncover policies where gathering information leads to better outcomes. Hope that helps!
120
Reposted by Raphaël Avalos
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 09/11/2024
If you're an RL researcher or RL adjacent, pipe up to make sure I've added you here! go.bsky.app/3WPHcHg
527127
Raphaël Avalos @raphael.avalos.fr · 25/11/2024
I am down !
000
Raphaël Avalos @raphael.avalos.fr · 24/11/2024
I had lots of fun at the first edition, there were good talks and papers, and it was nice seeing old friends and making new ones! Highly recommend submitting and attending 😁
030
Raphaël Avalos @raphael.avalos.fr · 22/11/2024
Just finished my first week at @cohere.com! Everyone has been so welcoming, and I’ve already learned so much. I’m really excited for the next weeks!
030
Raphaël Avalos @raphael.avalos.fr · 22/11/2024
Hello, World! I'll mostly write about LLMs, RL, and random coding stuff. :)
150
Raphaël Avalos @raphael.avalos.fr · 22/11/2024
Would love to be added!
000