Sign in

Graphcore Research

@gcresearchteam.bsky.social
136 followers 340 following 52 posts

The 🦋 account of the Graphcore Research team. Our mission is to contribute to the advancement of AI research and understand the computational requirements of intelligence.

PostsRepliesMedia
Graphcore Research @gcresearchteam.bsky.social · 09/10/2025
September’s Papers of the Month is here, and this month is all about LLMs! 🧠 This month, we cover: ➡️ FlowRL ➡️ Soft Tokens, Hard Truths ➡️ Set Block Decoding is a Language Model Inference Accelerator ➡️ Turning Recurring LLM Reasoning into Concise Behaviors 🧵
100
Graphcore Research @gcresearchteam.bsky.social · 10/09/2025
Summer may be over, but Papers of the Month certainly isn’t!

For August’s edition, we covered the following papers: ➡️ ADMIRE-BayesOpt ➡️ Guiding Diffusion Models with RL for Stable Molecule Generation ➡️ Graph-R1

 🧵
100
Graphcore Research @gcresearchteam.bsky.social · 06/08/2025
July's Papers of the Month are here! 🧠 Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data 💽 Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation 📊 DataRater: Meta-Learned Dataset Curation 🧵 ⬇️
101
Graphcore Research @gcresearchteam.bsky.social · 12/06/2025
Your boss emails you a point in 128-billion-dimensional space. It's Llama 8B in bfloat16. They want it compressed. What should you do 🤔... quantise to NF4? 🧵
111
Graphcore Research @gcresearchteam.bsky.social · 04/06/2025
As we hurtle into the summer, it’s time for May’s Papers of the Month! This month, we cover Parallel Scaling Laws for Language Models, Alpha Evolve, Soft Thinking and Spurious Rewards! 🧵
111
Graphcore Research @gcresearchteam.bsky.social · 22/05/2025
Our latest work uses theory from the '50s to figure out how to design weight quantisation formats for LLM inference. It's called Optimal Formats for Weight Quantisation and has just hit arXiv. 1/6
111
Graphcore Research @gcresearchteam.bsky.social · 08/05/2025
It's time for April's Papers of the Month! This month, we cover: ➡️ Motion Prompting: Controlling Video Generation with Motion Trajectories ➡️ Inference-Time Scaling for Generalist Reward Modeling ➡️ M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models! 🧵
110
Graphcore Research @gcresearchteam.bsky.social · 07/04/2025
Spring is here and so is Papers of the Month! In this March edition, we cover Transformers without Normalisation, Compute Optimal Scaling of Skills, Overtrained Language Models Are Harder to Fine-Tune, and Multi-Domain Distribution Learning for De Novo Drug Design! 🧵
111
Graphcore Research @gcresearchteam.bsky.social · 07/03/2025
February might have been the shortest month, but it wasn’t short of papers! In this edition of Papers of the Month, we cover Distillation Scaling Laws, Matryoshka Quantisation, ParetoQ, and Scaling Test-Time Compute with Latent Reasoning! 🧵
110
Graphcore Research @gcresearchteam.bsky.social · 04/02/2025
New year, new Papers of the Month! To kick off 2025, we cover: Titans, Evolving Deeper LLM Thinking, Transformer-Squared and the recent DeepSeek technical reports! 🧵 graphcore-research.github.io/papers-of-th...
graphcore-research.github.io
January Papers: More Like “Reas-anuary Papers”
New year, new Papers of the Month! Kicking off 2025, it’s apparent that reasoning and test-time compute are the hot topics on the block, with much research investigating how to best use these new meth...
110
Graphcore Research @gcresearchteam.bsky.social · 09/01/2025
Each month our team writes up summaries and analysis of our favourite ML papers. For December we cover: The Byte Latent Transformer, Large Concept Models, Memory Layers & Phi-4 — all grouped under the title "Spend Your FLOPs Wisely". Here's our take (🧵) graphcore-research.github.io/papers-of-th...
graphcore-research.github.io
December Papers: Spend Your FLOPs Wisely
Welcome to Papers of the Month — Graphcore Research’s effort to bring you our pick of the most interesting ML papers. In December we noted a collection of papers which took innovative approaches to al...
171