Reposted by DARK labTim Rocktäschel @handle.invalid · 25/11/2024@ucl-dark.bsky.social entered the stage! Thanks @lauraruis.bsky.social :) 0182
Reposted by DARK labLaura @lauraruis.bsky.social · 20/11/2024How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️ 36850139
Reposted by DARK labTim Rocktäschel @handle.invalid · 22/11/2024The LLM parrot analogy is dead. Fantastic work by UCL DARK's @lauraruis.bsky.social on rigorously investigating whether LLMs learn reasoning from procedural knowledge during pretraining. 0504
Reposted by DARK labTim Rocktäschel @handle.invalid · 22/11/2024Excited to announce "BALROG: a Benchmark for Agentic LLM and VLM Reasoning On Games" led b UCL DARK's @dpaglieri.bsky.social! Douwe Kiela plot below is maybe the scariest for AI progress — LLM benchmarks are saturating at an accelerating rate. BALROG to the rescue. This will keep us busy for years. 312415