Sign in

testerpce.bsky.social

@testerpce.bsky.social
104 followers 726 following 3 posts
PostsRepliesMedia
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 05/10/2026
Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation. Full paper: qlabs.sh/research/dust Code: github.com/qlabs-eng/dust
27111
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 01/10/2026
Learning from the 7,780 environments Xiaomi open-sourced for MiMo. In RL environments, reward design is everything huggingface.co/spaces/FineE...
0234
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 01/10/2026
Reshaping Monte-Carlo Tree Search 2FFS A new tree search algorithm that combines multi-fidelity bandits to resolve the fundamental trade-off: should we use cheaper, approximated evaluations, or expansive but accurate samplings? Paper: arxiv.org/abs/2606.01708
0201
Reposted by @testerpce.bsky.social
Grace @gracekind.net · 30/09/2026
Full name reveal
arxiv.org
GRACE: Reinforcement Learning for Grounded Response and Abstention under Contextual Evidence
Retrieval-Augmented Generation (RAG) integrates external knowledge to enhance Large Language Models (LLMs), yet systems remain susceptible to two critical flaws: providing correct answers without expl...
52275
Reposted by @testerpce.bsky.social
mr. TIM @timkellogg.me · 01/10/2026
Context Language Models New agent architecture where the LLM can edit its own context it seems to have emergent capabilities, creates its own memory management & organization algorithms, and coordinates multi agents github.com/facebookrese...
Diagram titled "How a Context Language Model edits its context: A simple step-by-step view" outlining an 8-step process:
 * Start of turn: Current editable context exists in memory with old messages.
 * LLM reads the context: The LLM evaluates the context and decides to run a bash command to edit it.
 * Harness mirrors context: The harness mirrors the old editable context into a file at /tmp/.live_ctx/LIVE_CTX_MAIN.txt.
 * Bash command runs: The bash command executes and may edit that file.
 * Harness parses file: If the file changed, the harness parses it back into a new edited context.
 * Tool call appended: The current assistant tool call is appended to the edited context.
 * Tool result appended: The tool result is appended below the tool call.
 * Next turn starts: The next turn begins with the edited old context, previous tool call, and previous tool result.
Key idea box: "The model edits the prior context first. The tool call and tool result from the current turn are appended afterward, so they can only be compacted on the next turn."
Footer summary:
 * Ordinary LM: Context mostly grows by appending.
 * CLM: The model can rewrite the editable part of context between turns.
1620315
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 01/10/2026
Tokenization: A Survey for Modern NLP Over the past ~8 months, 32 (!) tokenizer researchers put together the comprehensive survey of the field. www.alphaxiv.org/abs/2609.tok...
0203
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 01/10/2026
The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA! - Natively manage their own context - Treat context as a file - Learn policies in CLM weights, no harness
2574
Reposted by @testerpce.bsky.social
Tomer Ullman @tomerullman.bsky.social · 22/09/2026
new preprint: "Directing large language models to follow the letter or spirit of the law" arxiv.org/pdf/2609.23083 (by Qian , Li, Chen, Murthy @soniakmurthy.bsky.social , Belinkov, and me) this is particularly cool/important, and I'm allowed to say it because it was headed by @pqian.bsky.social
17323
Reposted by @testerpce.bsky.social
Steve Byrnes @stevebyrnes.bsky.social · 18/09/2026
Blog post: “Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)” www.lesswrong.com/posts/xvdngZ...
lesswrong.com
Pretraining data, not verifiability, is why LLMs are especially good at math (and coding) — LessWrong
A common take I’ve been hearing is: “LLMs are especially good at math because math is easy to verify”. But that story doesn’t make much sense.…
061
Reposted by @testerpce.bsky.social
Lukas Edman @lukasnlp.bsky.social · 17/09/2026
Ever feel like it's too hard to keep track of what LLMs cannot do as well as humans? We're making your life easier over at: what-llms-can-not-do.github.io We're compiling a list of papers testing the abilities of LLMs against humans. Check it out! And you can help contribute too!
what-llms-can-not-do.github.io
What LLMs Can(not) Do
A living survey of benchmarks that compare large language models with humans.
24516
Reposted by @testerpce.bsky.social
Jane Li 🦖 @janeli.bsky.social · 17/09/2026
🦀New preprint! (w/ @najoung.bsky.social)🦞 Is grammaticality a major organizing principle of NLM representations? We show that many NLMs exhibit abstract rep. separation for grammaticality. We believe this work addresses debates about confounds in measuring model gram. knowledge. [1/10]
12010
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 18/09/2026
Training a 4B model to produce 81% faster query plans than Postgres ...or how to make Qwen learn query optimization via agentic reinforcement learning rohanbansal.com/qorl?v=3
rohanbansal.com
Training a 4B model to produce 81% faster query plans than Postgres
...or how to make Qwen learn query optimization via agentic reinforcement learning
0282
Reposted by @testerpce.bsky.social
Quanta Magazine @quantamagazine.org · 18/09/2026
In 2004, two mathematicians hypothesized a powerful kind of sandwich. A new proof has finally built it. www.quantamagazine.org/mathematicia...
quantamagazine.org
Mathematicians Build Long-Awaited Graph Sandwich | Quanta Magazine
The proof of a decades-old conjecture has given researchers a new way to understand complex networks.
0278
Reposted by @testerpce.bsky.social
Michael Noukhovitch @mnoukhov.bsky.social · 15/09/2026
Is RL actually making your LLM better? Gains from RL are mostly on easy questions🤯 We're calling this the Matthew Effect for RL on LLMs. We then leverage async RL to solve harder problems by Never Giving Up! arxiv.org/abs/2609.13443 and mnoukhov.github.io/posts/ngu/ and check out thread below 🧵👇
1277
Reposted by @testerpce.bsky.social
Chris Amato @cjdamato.bsky.social · 16/09/2026
A new version of my book on cooperative multi-agent reinforcement learning is now available. A longer version will eventually be published with Frans Oliehoek so let us know your thoughts! arxiv.org/abs/2405.06161
arxiv.org
An Initial Introduction to Cooperative Multi-Agent Reinforcement Learning
Multi-agent reinforcement learning (MARL) has exploded in popularity in recent years. While numerous approaches have been developed, they can be broadly categorized into three main types: centralized ...
1264
Reposted by @testerpce.bsky.social
arxiv cs.CL @arxiv-cs-cl.bsky.social · 15/09/2026
Mengyi Deng, Xin Li, Duyi Pan, Zilin Wang, Zhiwei Li, Zhijiang Guo, Wei Wang RESKILL: Explicit Failure Attribution and Structured Repair for Interactive Language Agents arxiv.org/abs/2609.15684
011
Reposted by @testerpce.bsky.social
michael catchen @mdcatchen.bsky.social · 15/09/2026
the most interesting thing i worked on during my phd is now out as a preprint. ecoevorxiv.org/repository/v...
ecoevorxiv.org
Designing optimal biodiversity observation networks
1298
Reposted by @testerpce.bsky.social
arxiv cs.CL @arxiv-cs-cl.bsky.social · 15/09/2026
Jieyuan Liu, Mengzhou Hu, Jefferson Chen, JungHo Kong, Pratibha Jagannatha, Yiming Gao, Dexter Pratt, Hsin-Yuan Lee, Zhiting Hu, Trey Ideker, Wei Wang, Eric P. Xing, Zhen Wang HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses arxiv.org/abs/2609.15938
011
Reposted by @testerpce.bsky.social
Kunal Jha @kjha02.bsky.social · 11/09/2026
Can self-interested, self-improving, self-replicating agents learn to cooperate? Our new paper, Tapes Together Strong, shows they can: when social behavior, computation, and reproduction share one energy budget, cooperation evolves from scratch. arxiv.org/abs/2609.10817 🧵
48418
Reposted by @testerpce.bsky.social
arxiv cs.CL @arxiv-cs-cl.bsky.social · 15/09/2026
JuHeon Ha, Byounghan Lee, Yunseo Choi, Kyung-Ah Sohn Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs arxiv.org/abs/2609.15654
011
Reposted by @testerpce.bsky.social
arxiv cs.CL @arxiv-cs-cl.bsky.social · 15/09/2026
Kyle O'Brien, Edward James Young, Puria Radmard, Nathalie Kirch, Cameron Tice, Tomek Korbak, David Demitri Africa Inoculation Midtraining with Learned Neologisms arxiv.org/abs/2609.15886
011
Reposted by @testerpce.bsky.social
Ryan O'Donnell @booleananalysis.bsky.social · 14/09/2026
Outstanding progress towards the Unique Games Conjecture posted by Yumou Fei, Dor Minzer, and Shuo Wang: eccc.weizmann.ac.il/report/2026/... 👀
eccc.weizmann.ac.il
ECCC - TR26-179
3237
Reposted by @testerpce.bsky.social
Clément Canonne @ccanonne.github.io · 14/09/2026
Big (huge?) news on the computational complexity preprint server, with a paper by Yumou Fei, Dor Minzer, and Shuo Wang I am not qualified to read beyond the intro: the perfect completeness version of the (not) Unique Games Conjecture, the 4-to-1 Games conjecture! eccc.weizmann.ac.il/report/2026/...
eccc.weizmann.ac.il
ECCC - TR26-179
2459
Reposted by @testerpce.bsky.social
Sikata Sengupta @sikatasengupta.bsky.social · 15/09/2026
This was super fun to work on!
091
Reposted by @testerpce.bsky.social
Elliot Murphy @elliot-murphy.bsky.social · 15/09/2026
New work out today! The result of a 15 year project to offer genuine linking hypotheses between linguistics and neuroscience: a revised research program, and its first result - a novel neural binding operation ('Meld') capturing core properties of natural language syntax 🧵 arxiv.org/abs/2609.14384
arxiv.org
Formal Properties of Language as Constraints on Neural Dynamics
What must a neural system be capable of to implement language? Current research annotates stimuli with linguistic variables and tests which electrodes, voxels, or language-model layers predict neural ...
26915
Reposted by @testerpce.bsky.social
Nathan Lambert @natolambert.bsky.social · 15/09/2026
An basic idea in scaling RL: Can we allocate more compute to the harder problems? We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works! The paper: arxiv.org/abs/2609.13443
1598
Reposted by @testerpce.bsky.social
Seth Juarez @sethjuarez.com · 15/09/2026
The smallest bridge from a model to your code is one word. Make it answer BOOK or CANCEL, nothing else. Now `if reply == "BOOK"` works. A paragraph of intent becomes one token your code can branch on. The seed of tool use. sethjuarez.com/posts/models...
011
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 14/09/2026
Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic claude.com/blog/agentic...
claude.com
Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic | Claude by Anthropic
Our CI job volume increased 25x over 6 months. We patched our test selection service three times before finding a sustainable solution.
0122
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 03/09/2026
Flow Reasoning Models. They developed a recurrent flow-based architecture to efficiently solve structured reasoning problems (e.g., Sudoku). It applies continuous flows to discrete data and recurrently refine their past mistakes through self-conditioning.
1488
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 02/09/2026
OpenAI’s Astra may be using Recurrent Depth as outlined in this paper: Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach (arxiv.org/abs/2502.05171)
arxiv.org
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
We study a novel language model architecture that is capable of scaling test-time computation by implicitly reasoning in latent space. Our model works by iterating a recurrent block, thereby unrolling...
1255
Reposted by @testerpce.bsky.social
Xiulin Yang @xiulinyang.bsky.social · 28/08/2026
🍎🍊 How would you know if a language model is better at one language than another? Our #EMNLP2026 paper argues that only one metric can actually lead to fair crosslingual evaluation. This work is a collaboration with @wegotlieb.bsky.social & @catherinearnett.bsky.social! (1/5)
13312
Reposted by @testerpce.bsky.social
Reinforcement Learning Conference @rl-conference.bsky.social · 17/08/2026
A truly compelling keynote from Sheila McIlraith at RLC 2026 today — exploring how formal language can serve as a nexus between signals and symbols for agents that learn, plan, and remember. A principled and thought-provoking perspective on the future of RL and agentic AI!
0174
Reposted by @testerpce.bsky.social
Sung Kim @sungkim.bsky.social · 14/08/2026
Why LLMs read fast but generate slowly? Inference Performance from First Principles by Eric Schreiber aleph-alpha.com/en/blog/infe...
2405
Reposted by @testerpce.bsky.social
Nathan Lambert @natolambert.bsky.social · 21/07/2026
How distillation is used today and what performance uplift it gives to open models (a rant) natolambert.substack.com/p/how-distil...
natolambert.substack.com
How distillation is used today and what performance uplift it gives to open models
This is in response to Ben Thompson’s recent piece where he said distillation is happening during RL and becoming more important to model performance.
0275
Reposted by @testerpce.bsky.social
Lukas Schäfer @lukaschaefer.bsky.social · 14/07/2026
This looks awesome! We need benchmarks that consider coordination between agents. Also super cool to see cross comparison between LLM/ VLM and MARL agents! Congrats to @kale-ab.bsky.social & co. 👏
094
Reposted by @testerpce.bsky.social
Marc Lanctot @sharky6000.bsky.social · 05/12/2024
Super happy to reveal our new paper! 🎉🙌♟️ We trained a model to play four games, and the performance in each increases by "external search" (MCTS using a learned world model) and "internal search" where the model outputs the whole plan on its own!
413818
Reposted by @testerpce.bsky.social
arxiv cs.CL @arxiv-cs-cl.bsky.social · 24/06/2026
Puneet Kant, Monika Tanwar A Synthetic Reliability-Aware PINN Benchmark for Offshore Wind Turbine Support-Structure Monitoring with Bayesian Inverse Identification arxiv.org/abs/2606.24176
021
Reposted by @testerpce.bsky.social
arxiv cs.CL @arxiv-cs-cl.bsky.social · 24/06/2026
Tianyu Ding, Juan Pablo De la Cruz Weinstein When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents arxiv.org/abs/2606.23937
011
Reposted by @testerpce.bsky.social
Chris Paxton @cpaxton.bsky.social · 13/06/2026
It's still really hard to tell which robotics models are "best," and there seems to be a good amount of competition at the top -- but at least I guess the willingness of some people to cheat on leaderboards is a good sign? open.substack.com/pub/itcanthi...
open.substack.com
What Do Robotics Leaderboards Tell Us About The State of Robot Learning?
There remains no chatbot arena for robotics
0132
Reposted by @testerpce.bsky.social
Marc Lanctot @sharky6000.bsky.social · 15/01/2026
Hello all! 👋 I’m delighted to share a 🚨 new preprint 🚨: “Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms”. A paper thread! 🤩📄🧵 1/N
25712
Reposted by @testerpce.bsky.social
Vinzenz Thoma @vthoma.bsky.social · 18/12/2025
Unlike board games, real-world strategic interactions are messy. Traditional game theory thus needs a boost for the age of agentic AI. Our #AAMAS2026 workshop "Strategic Engineering"(sites.google.com/view/se-aama...) in Cyprus aims to bridge the gap. Come join us to unlock truly strategic AI!
0136
Reposted by @testerpce.bsky.social
Grant Sanderson @3blue1brown.com · 12/06/2026
Here's an excerpt from the most recent video on how Shannon studied the entropy of English, animated by Mitchell Zemil
111510
Reposted by @testerpce.bsky.social
Tom Silver @tomssilver.bsky.social · 10/05/2026
This week's #PaperILike is "Human-Guided Complexity-Controlled Abstractions" (Peng et al., NeurIPS 2023). Selecting the right levels and kinds of abstractions remains important and open for many forms of human-AI / human-robot interaction. PDF: arxiv.org/abs/2310.17550
arxiv.org
Human-Guided Complexity-Controlled Abstractions
Neural networks often learn task-specific latent representations that fail to generalize to novel settings or tasks. Conversely, humans learn discrete representations (i.e., concepts or words) at a va...
0134
Reposted by @testerpce.bsky.social
MilaNLP Lab @milanlp.bsky.social · 11/05/2026
#MemoryModay #NLProc 'Visualizing Regional Language Variation Across Europe on Twitter' by Dirk Hovy et al. uncovers language differences across Europe with stunning visuals. Language is art! #Linguistics link.springer.com/referencewor...
link.springer.com
Visualizing Regional Language Variation Across Europe on Twitter
Geotagged Twitter data allows us to investigate correlations of geographic language variation, both at an interlingual and intralingual level. Based on data-driven studies of such relationships, this…
043
Reposted by @testerpce.bsky.social
Anton Hur @antonhur.com · 11/05/2026
WE DON'T NEED PETROLEUM-BASED PLASTICS ANYMORE! "The [bamboo plastic] outperforms most commercial plastics and bioplastics in mechanical and thermo-mechanical metrics while maintaining full biodegradability in soil within 50 days and closed-loop recyclability with 90% retained strength."
nature.com
High-strength, multi-mode processable bamboo molecular bioplastic enabled by solvent-shaping regulation - Nature Communications
Bioplastics derived from biomass show promise as sustainable alternatives to petrochemical plastics, but their adoption is hindered by their inferior mechanical properties and processability. Here, th...
2122984016
Reposted by @testerpce.bsky.social
Thomas Dietterich @tdietterich.bsky.social · 10/05/2026
This points to an important direction: layering symbolic systems on top of LLMs. These can overcome the main shortcomings of LLM architectures: probabilistic execution, continual learning, attribution, and (maybe) uncertainty quantification. 1/
2347
Reposted by @testerpce.bsky.social
Alexander Hoyle @alexanderhoyle.bsky.social · 11/05/2026
This paper is getting a lot of (deserved) attention; I think this paper serves as a nice complement arxiv.org/abs/2602.18710
arxiv.org
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
Empirical conclusions depend not only on data but on analytic decisions made throughout the research process. Many-analyst studies have quantified this dependence: independent teams testing the same h...
0135
Reposted by @testerpce.bsky.social
Dr Di Cook @visnut.bsky.social · 11/05/2026
Super excited to get this at my door today! Ursula Laa and my book on exploring high-d data and models is in print! Book website is www.routledge.com/Interactivel... if interested I think my 30% discount code might work for you - msg me. Time to update dicook.github.io/mulgar_book/ too! #rstats
415832
Reposted by @testerpce.bsky.social
Anna Rogers @annarogers.bsky.social · 11/05/2026
A Human-Centric Framework for Data Attribution in LLMs (FaCCT'26) TLDR: LLMs screwed up data economy, and NLP should help to re-design incentives with data attribution. Here's a moonshot for an alternative LLM use paradigm, for the users, creators & intermediaries. arxiv.org/abs/2602.10995 /1
4348
Reposted by @testerpce.bsky.social
Avijit Ghosh @evijit.io · 11/05/2026
Should all our resources go towards building chatbots? What if we built systems that actually give people meaningful agency? Finally out as an accepted @facct.bsky.social paper! Joint work with my past collaborators Sourojit, Pranav and Sanjana. huggingface.co/papers/2605....
huggingface.co
Paper page - What if AI systems weren't chatbots?
Join the discussion on this paper page
34314