Reposted by Alex MakelovDavid Bau @davidbau.bsky.social · 20/02/2025Today we launch a new open research community It is called ARBOR: arborproject.github.io/ please join us. bsky.app/profile/ajy... 1155
Reposted by Alex MakelovMartin Wattenberg @wattenberg.bsky.social · 23/12/2024The math benchmarks I want: 1. OopsBench: given a faulty proof with numbered steps, which step contains an unfixable logical flaw? 2. DunnoMath: half the problems are taken from FrontierMath, half are almost certainly unsolvable. Major points off for guessing an answer to an unsolvable problem. 6434
Reposted by Alex MakelovNDIF Team @ndif-team.bsky.social · 20/12/2024Large language models show fascinating changes in capability with scaling parameters, but scaling also vastly increases the resources required for experimentation on model internals. NDIF is currently hosting the largest open sourced model, Llama 405b, for YOU to run research on! 152
Alex Makelov @amakelov.bsky.social · 05/12/2024Some fun with o1 from OpenAI: there's a math problem I often give to "reasoning" AIs to try them out. It's basically to prove that there's a number less than 1 billion that you can write in 1000 different ways as a sum of 3 squares (precise statement in the pic). 111
Alex Makelov @amakelov.bsky.social · 24/11/2024yes, this is what mechanistic interpretability research looks like 2232