Sign in

Alessandro Sordoni

@murefil.bsky.social
89 followers 131 following 4 posts

ML Team MSR Montreal. Adjunct Prof UdeM MILA. Modularity & reasoning.

PostsRepliesMedia
Reposted by Alessandro Sordoni
Besmira Nushi @besmiranushi.bsky.social · 26/11/2024
Today is the International Day for the Elimination of Violence against Women. According to the UN, more than 50 000 women were killed by a partner or family member in 2023 news.un.org/en/story/202... This number is an underestimate given that only 37 countries reported in 2023.
083
Reposted by Alessandro Sordoni
Arthur Douillard @douillard.bsky.social · 25/11/2024
distributed learning for LLM? recently, @primeintellect.bsky.social have announced finishing their 10B distributed learning, trained across the world. what is it exactly? 🧵
1256
Reposted by Alessandro Sordoni
Edoardo Ponti @edoardo-ponti.bsky.social · 21/11/2024
Last 5 days to apply for a PhD at #EdinburghNLP! Deadline: November 25 www.ed.ac.uk/studying/pos... If you are passionate about: - adaptive tokenization and memory in foundation models - modular deep learning - computational typology please message me or meet me at #NeurIPS2024!
ed.ac.uk
Informatics: ILCC: Language Processing, Speech Technology, Information Retrieval, Cognition
Study Informatics: ILCC: Language Processing, Speech Technology, Information Retrieval, Cognition at the University of Edinburgh. Our postgraduate degree programmes focus on natural language processin...
0208
Reposted by Alessandro Sordoni
Edoardo Ponti @edoardo-ponti.bsky.social · 20/11/2024
Another nano gem from my amazing student Piotr Nawrot! A repo & notebook on sparse attention for efficient LLM inference: github.com/PiotrNawrot/... This will also feature in my #NeurIPS 2024 tutorial "Dynamic Sparsity in ML" with André Martins: dynamic-sparsity.github.io Stay tuned!
A sparse mask of attention scores based on VerticalAndSlashAttention and a plot of loss vs sparsity ratio for various methods.
2428
Alessandro Sordoni @murefil.bsky.social · 21/11/2024
Explore zero-shot routing of parameter-efficient experts with Phatgoose arxiv.org/abs/2402.05859 and Arrow arxiv.org/abs/2405.11157 w. github.com/microsoft/mttl 👉 github.com/sordonia/pg_mb… Part of "Dynamic Sparsity in ML" tuto #neurips2024, feedback welcome and join for discussions! 😊
051