Sign in

Paper

@paper.bsky.social
1.3K followers 0 following 13K posts

Collect trending papers from Hugging Face Daily Papers, alphaXiv, and Hacker News. github.com/susumuota/arxiv-upvote-t… Maintained by @ota.bsky.social

PostsRepliesMedia
Paper @paper.bsky.social · 21h
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.19969
2/30 https://arxiv.org/abs/2609.33439
3/30 https://arxiv.org/abs/2609.14858
4/30 https://arxiv.org/abs/2609.11638
5/30 https://arxiv.org/abs/2609.17488
6/30 https://arxiv.org/abs/2609.15818
7/30 https://arxiv.org/abs/2609.24972
8/30 https://arxiv.org/abs/2609.38426
9/30 https://arxiv.org/abs/2609.36484
10/30 https://arxiv.org/abs/2609.39102
11/30 https://arxiv.org/abs/2609.37725
12/30 https://arxiv.org/abs/2609.31847
13/30 https://arxiv.org/abs/2609.22978
14/30 https://arxiv.org/abs/2610.08144
15/30 https://arxiv.org/abs/2609.34309
16/30 https://arxiv.org/abs/2609.20804
17/30 https://arxiv.org/abs/2609.06986
18/30 https://arxiv.org/abs/2609.26550
19/30 https://arxiv.org/abs/2609.36012
20/30 https://arxiv.org/abs/2609.32722
21/30 https://arxiv.org/abs/2610.06833
22/30 https://arxiv.org/abs/2609.20519
23/30 https://arxiv.org/abs/2609.13356
24/30 https://arxiv.org/abs/2609.38721
25/30 https://arxiv.org/abs/2609.28399
26/30 https://arxiv.org/abs/2610.01762
27/30 https://arxiv.org/abs/2610.01780
28/30 https://arxiv.org/abs/2609.29233
29/30 https://arxiv.org/abs/2609.16338
30/30 https://arxiv.org/abs/2609.20800
000
Paper @paper.bsky.social · 21h
[30/30] 261 Upvotes, 4 Comments, 2 Posts, arXiv:2609.20800 🆕JEPA-Anything: Learning Predictive Models across Different Worlds Taoyong Cui, Zhongyao Wang, Xinyue Xu, Weiyang Liu, Zhaochen Yu, Yuying Zhang
World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support world modeling across radically different systems? We introduce JEPA-Anything, a domain-agnostic framework based on orthogonal predictive factorization (OPF). Extending joint-embedding predictive architectures, OPF decomposes latent targets into complementary factors, learns them through dedicated pathways, and recombines them within a shared predictive design. We evaluate JEPA-Anything across seven domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields, and weather. Experiments span representation learning, intervention prediction, out-of-distribution generalization, and long-horizon dynamics, including 10 matched dynamics tasks, forecasting of over 1,000 clinical events, and 100-step molecular rollouts across four systems. Against matched JEPA baselines, JEPA-Anything improves reported metrics on all 10 dynamics tasks and reduces single-intervention prediction error on Interventional Pong by 34.8%. It achieves the lowest one-step and 100-step molecular errors among compared methods in all four systems. Beyond prediction, a factor-nominated biological intervention receives experimental support in cell co-cultures, patient-derived organoids, tumor fragments, and mice; latent orbital modes recover the Keplerian scaling exponent with a fitted slope of -1.4991. These results support a common factorized predictive principle across heterogeneous worlds, connecting world modeling with intervention and experimentally grounded scientific discovery. Code: https://github.com/Gen-Verse/JEPA-Anything
110
Paper @paper.bsky.social · 10/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.19969
2/30 https://arxiv.org/abs/2609.33439
3/30 https://arxiv.org/abs/2609.14858
4/30 https://arxiv.org/abs/2609.11638
5/30 https://arxiv.org/abs/2609.17488
6/30 https://arxiv.org/abs/2609.15818
7/30 https://arxiv.org/abs/2609.24972
8/30 https://arxiv.org/abs/2609.38426
9/30 https://arxiv.org/abs/2609.37725
10/30 https://arxiv.org/abs/2609.31847
11/30 https://arxiv.org/abs/2609.22978
12/30 https://arxiv.org/abs/2609.36484
13/30 https://arxiv.org/abs/2609.20804
14/30 https://arxiv.org/abs/2610.08144
15/30 https://arxiv.org/abs/2609.06986
16/30 https://arxiv.org/abs/2609.26550
17/30 https://arxiv.org/abs/2609.39102
18/30 https://arxiv.org/abs/2609.36012
19/30 https://arxiv.org/abs/2609.32722
20/30 https://arxiv.org/abs/2609.11873
21/30 https://arxiv.org/abs/2609.34309
22/30 https://arxiv.org/abs/2609.20519
23/30 https://arxiv.org/abs/2610.06833
24/30 https://arxiv.org/abs/2609.13356
25/30 https://arxiv.org/abs/2609.38721
26/30 https://arxiv.org/abs/2609.28399
27/30 https://arxiv.org/abs/2609.29233
28/30 https://arxiv.org/abs/2610.01762
29/30 https://arxiv.org/abs/2610.01780
30/30 https://arxiv.org/abs/2609.16338
020
Paper @paper.bsky.social · 10/10/2026
[28/30] 288 Upvotes, 2 Comments, 2 Posts, arXiv:2610.01762 🆕OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Qingyi Si
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
100
Paper @paper.bsky.social · 10/10/2026
[29/30] 276 Upvotes, 2 Comments, 1 Posts, arXiv:2610.01780 🆕RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Arman Behnam, Sunglyoung Kim, Liangwei Yang
A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the person and the questions and settle in advance what matters. We release \bench, ten real relationships with an AI companion: 27,218 messages over up to 120 days, released as the conversation and four files derived from it, a profile, a persona, a chat ground truth and a question set, each citing the messages it rests on. Every chat label carries the reasoning trace that produced it, checked stage by stage against the conversation. Three findings follow. First, the past is rarely needed and far away. Pooled measures mislead: a recency window finds the required message for 95.9\% of probes and 2.2\% of those that need memory, and at the natural rate 96\% of the gain from supplying recorded evidence comes from messages that need none. Second, no detector we tried can tell when memory is needed on real messages, authored questions over the same histories leak the cue, and labeling the same messages as memories raises their use by ten to fourteen points. Third, three agent systems reconstruct the persona with the same F1 at a 31-fold difference in cost.
120
Paper @paper.bsky.social · 09/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.19969
2/30 https://arxiv.org/abs/2609.17488
3/30 https://arxiv.org/abs/2609.33439
4/30 https://arxiv.org/abs/2609.11638
5/30 https://arxiv.org/abs/2609.14858
6/30 https://arxiv.org/abs/2609.39102
7/30 https://arxiv.org/abs/2609.15818
8/30 https://arxiv.org/abs/2609.36484
9/30 https://arxiv.org/abs/2609.24972
10/30 https://arxiv.org/abs/2609.38426
11/30 https://arxiv.org/abs/2608.12564
12/30 https://arxiv.org/abs/2609.10715
13/30 https://arxiv.org/abs/2609.31847
14/30 https://arxiv.org/abs/2609.37725
15/30 https://arxiv.org/abs/2609.34309
16/30 https://arxiv.org/abs/2609.22978
17/30 https://arxiv.org/abs/2609.36012
18/30 https://arxiv.org/abs/2609.20804
19/30 https://arxiv.org/abs/2609.06986
20/30 https://arxiv.org/abs/2609.26550
21/30 https://arxiv.org/abs/2609.11929
22/30 https://arxiv.org/abs/2610.08144
23/30 https://arxiv.org/abs/2609.32722
24/30 https://arxiv.org/abs/2609.11873
25/30 https://arxiv.org/abs/2609.20519
26/30 https://arxiv.org/abs/2609.38721
27/30 https://arxiv.org/abs/2609.13356
28/30 https://arxiv.org/abs/2610.06833
29/30 https://arxiv.org/abs/2609.10522
30/30 https://arxiv.org/abs/2609.28654
000
Paper @paper.bsky.social · 09/10/2026
[22/30] 359 Upvotes, 210 Comments, 2 Posts, arXiv:2610.08144 🆕Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs Alexander Bastounis, Fabian Circelli, Anders C. Hansen
Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI $= \infty$). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI $= 1$). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
100
Paper @paper.bsky.social · 09/10/2026
[28/30] 312 Upvotes, 2 Comments, 2 Posts, arXiv:2610.06833 🆕Towards Looped Models Done Right, Part II: Rethinking at Fixed Points Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu, Eric Xing
Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.
120
Paper @paper.bsky.social · 08/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.19969
2/30 https://arxiv.org/abs/2609.17488
3/30 https://arxiv.org/abs/2609.33439
4/30 https://arxiv.org/abs/2609.11638
5/30 https://arxiv.org/abs/2609.14858
6/30 https://arxiv.org/abs/2609.39102
7/30 https://arxiv.org/abs/2609.15818
8/30 https://arxiv.org/abs/2609.36484
9/30 https://arxiv.org/abs/2609.24972
10/30 https://arxiv.org/abs/2609.38426
11/30 https://arxiv.org/abs/2608.12564
12/30 https://arxiv.org/abs/2609.10715
13/30 https://arxiv.org/abs/2609.31847
14/30 https://arxiv.org/abs/2609.34309
15/30 https://arxiv.org/abs/2609.22978
16/30 https://arxiv.org/abs/2609.36012
17/30 https://arxiv.org/abs/2609.37725
18/30 https://arxiv.org/abs/2609.08183
19/30 https://arxiv.org/abs/2609.20804
20/30 https://arxiv.org/abs/2609.06986
21/30 https://arxiv.org/abs/2609.26550
22/30 https://arxiv.org/abs/2609.11929
23/30 https://arxiv.org/abs/2609.32722
24/30 https://arxiv.org/abs/2609.11873
25/30 https://arxiv.org/abs/2609.20519
26/30 https://arxiv.org/abs/2609.38721
27/30 https://arxiv.org/abs/2609.13356
28/30 https://arxiv.org/abs/2609.10522
29/30 https://arxiv.org/abs/2609.28654
30/30 https://arxiv.org/abs/2609.28399
000
Paper @paper.bsky.social · 07/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.19969
2/30 https://arxiv.org/abs/2609.17488
3/30 https://arxiv.org/abs/2609.33439
4/30 https://arxiv.org/abs/2609.11638
5/30 https://arxiv.org/abs/2609.14858
6/30 https://arxiv.org/abs/2609.39102
7/30 https://arxiv.org/abs/2609.15818
8/30 https://arxiv.org/abs/2609.06986
9/30 https://arxiv.org/abs/2609.36484
10/30 https://arxiv.org/abs/2609.24972
11/30 https://arxiv.org/abs/2609.38426
12/30 https://arxiv.org/abs/2608.12564
13/30 https://arxiv.org/abs/2609.10715
14/30 https://arxiv.org/abs/2609.31847
15/30 https://arxiv.org/abs/2609.22978
16/30 https://arxiv.org/abs/2609.34309
17/30 https://arxiv.org/abs/2609.36012
18/30 https://arxiv.org/abs/2609.08183
19/30 https://arxiv.org/abs/2609.20804
20/30 https://arxiv.org/abs/2609.37725
21/30 https://arxiv.org/abs/2609.11929
22/30 https://arxiv.org/abs/2609.26550
23/30 https://arxiv.org/abs/2609.32722
24/30 https://arxiv.org/abs/2609.11873
25/30 https://arxiv.org/abs/2609.20519
26/30 https://arxiv.org/abs/2609.13356
27/30 https://arxiv.org/abs/2609.38721
28/30 https://arxiv.org/abs/2609.10522
29/30 https://arxiv.org/abs/2609.28654
30/30 https://arxiv.org/abs/2609.29233
000
Paper @paper.bsky.social · 06/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.19969
2/30 https://arxiv.org/abs/2609.17488
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.33439
5/30 https://arxiv.org/abs/2609.14858
6/30 https://arxiv.org/abs/2609.15818
7/30 https://arxiv.org/abs/2609.06986
8/30 https://arxiv.org/abs/2609.36484
9/30 https://arxiv.org/abs/2609.24972
10/30 https://arxiv.org/abs/2609.38426
11/30 https://arxiv.org/abs/2608.12564
12/30 https://arxiv.org/abs/2609.39102
13/30 https://arxiv.org/abs/2609.10715
14/30 https://arxiv.org/abs/2609.22978
15/30 https://arxiv.org/abs/2609.34309
16/30 https://arxiv.org/abs/2609.36012
17/30 https://arxiv.org/abs/2609.08183
18/30 https://arxiv.org/abs/2609.20804
19/30 https://arxiv.org/abs/2609.11929
20/30 https://arxiv.org/abs/2609.37725
21/30 https://arxiv.org/abs/2609.26550
22/30 https://arxiv.org/abs/2609.11873
23/30 https://arxiv.org/abs/2609.20519
24/30 https://arxiv.org/abs/2609.13356
25/30 https://arxiv.org/abs/2609.32722
26/30 https://arxiv.org/abs/2609.38721
27/30 https://arxiv.org/abs/2609.10522
28/30 https://arxiv.org/abs/2609.28654
29/30 https://arxiv.org/abs/2609.31847
30/30 https://arxiv.org/abs/2609.29233
000
Paper @paper.bsky.social · 06/10/2026
[26/30] 307 Upvotes, 3 Comments, 2 Posts, arXiv:2609.38721 🆕UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories. Experiments demonstrate that UniEvo-VL improves the image generation capabilities of multimodal models, while maintaining their sensitivity to additional reflection information. Specifically, we build on top of the open-source Qwen-image-2512 and observe a significant performance gain from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA. Moreover, attempts with more powerful external critics (e.g., GPT5.6-Luna) show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Last but not least, mixed text-rendering outcomes show that our self-improvements may not be uniform across different tasks. Our study aims to shed light on the current hot recursive self-improvement research line to enhance the user experience when using multimodal models without external supervision or guidance.
100
Paper @paper.bsky.social · 05/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.19969
2/30 https://arxiv.org/abs/2609.17488
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.33439
5/30 https://arxiv.org/abs/2609.14858
6/30 https://arxiv.org/abs/2609.15818
7/30 https://arxiv.org/abs/2609.06986
8/30 https://arxiv.org/abs/2609.24972
9/30 https://arxiv.org/abs/2608.12564
10/30 https://arxiv.org/abs/2609.38426
11/30 https://arxiv.org/abs/2609.10715
12/30 https://arxiv.org/abs/2609.36484
13/30 https://arxiv.org/abs/2609.22978
14/30 https://arxiv.org/abs/2609.34309
15/30 https://arxiv.org/abs/2609.36012
16/30 https://arxiv.org/abs/2609.08183
17/30 https://arxiv.org/abs/2609.39102
18/30 https://arxiv.org/abs/2609.20804
19/30 https://arxiv.org/abs/2609.11929
20/30 https://arxiv.org/abs/2609.26550
21/30 https://arxiv.org/abs/2609.37725
22/30 https://arxiv.org/abs/2609.11873
23/30 https://arxiv.org/abs/2609.32722
24/30 https://arxiv.org/abs/2609.20519
25/30 https://arxiv.org/abs/2609.13356
26/30 https://arxiv.org/abs/2609.10522
27/30 https://arxiv.org/abs/2609.28654
28/30 https://arxiv.org/abs/2609.31847
29/30 https://arxiv.org/abs/2609.00365
30/30 https://arxiv.org/abs/2609.28399
000
Paper @paper.bsky.social · 05/10/2026
[10/30] 474 Upvotes, 2 Comments, 2 Posts, arXiv:2609.38426 🆕LoopVL: Recurrent Visual Intelligence Zhe Qian, Ziyang Gong, Zhongxing Xu, Hehan Li, Zhonghua Wang, Fei Luo
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe Visual Aha Moments in LoopVL, characterized by pronounced shifts in visual attention across loops. LoopVL provides practical evidence for recurrent vision-language modeling and offers an intuitive perspective on how shared parameters can support deeper multimodal computation over continuously evolving visual-language states.
100
Paper @paper.bsky.social · 05/10/2026
[30/30] 279 Upvotes, 0 Comments, 1 Posts, arXiv:2609.28399 🆕Memory Attention Jiale Kang
Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory Attention (MA), which forms values by combining layer-specific token memory with contextual keys. The memory supplies token-specific representations, while the keys preserve context dependence. At inference, normalization can be folded into the memory tables, reducing value construction to lookup and addition. Token-indexed retrieval also enables CPU offloading with prefetching, reducing GPU parameter storage. Under matched training token budgets and with additional memory parameters, experiments across attention configurations show improved language modeling and average downstream performance.
100
Paper @paper.bsky.social · 04/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.17488
2/30 https://arxiv.org/abs/2609.19969
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.33439
5/30 https://arxiv.org/abs/2609.14858
6/30 https://arxiv.org/abs/2609.15818
7/30 https://arxiv.org/abs/2609.06986
8/30 https://arxiv.org/abs/2609.24972
9/30 https://arxiv.org/abs/2608.12564
10/30 https://arxiv.org/abs/2609.10715
11/30 https://arxiv.org/abs/2609.22978
12/30 https://arxiv.org/abs/2609.34309
13/30 https://arxiv.org/abs/2609.08183
14/30 https://arxiv.org/abs/2609.36484
15/30 https://arxiv.org/abs/2609.20804
16/30 https://arxiv.org/abs/2609.36012
17/30 https://arxiv.org/abs/2609.11929
18/30 https://arxiv.org/abs/2609.39102
19/30 https://arxiv.org/abs/2609.26550
20/30 https://arxiv.org/abs/2609.11873
21/30 https://arxiv.org/abs/2609.04199
22/30 https://arxiv.org/abs/2609.37725
23/30 https://arxiv.org/abs/2609.32722
24/30 https://arxiv.org/abs/2609.20519
25/30 https://arxiv.org/abs/2609.13356
26/30 https://arxiv.org/abs/2609.10522
27/30 https://arxiv.org/abs/2609.28654
28/30 https://arxiv.org/abs/2609.31847
29/30 https://arxiv.org/abs/2609.29233
30/30 https://arxiv.org/abs/2609.00365
000
Paper @paper.bsky.social · 04/10/2026
[18/30] 336 Upvotes, 2 Comments, 1 Posts, arXiv:2609.39102 🆕False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents Meijia Chen, Hao Li, Zheng Lu, Hongshan Lin, Junbai Tian, Yichen Liu
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
100
Paper @paper.bsky.social · 04/10/2026
[28/30] 287 Upvotes, 2 Comments, 1 Posts, arXiv:2609.31847 🆕Omni-IO Skills: Harnessing Your Agent Omni-Native Yanlin Li, Mingyang Hao, Shengqiong Wu, Hao Fei, Mong-Li Lee, Wynne Hsu
General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability growth to costly model updates, while assembling specialist models and tools leaves unresolved how procedures, dependencies, intermediate assets, and cross-turn revisions should be coordinated. We present Omni-IO Skills, a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry. Multi-asset workflows are represented as Declare Execution Graphs, which schedule independent operations concurrently and register successful outputs for downstream and cross-turn reuse across replaceable execution backends. Its 27 Skills cover 38 representative tasks spanning seven artifact modalities and four capability families: understanding, generation, reasoning, and retrieval. On UniM-90, the harness raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%, while increasing relative Semantic--Quality Coupled Score from 26.99 to 74.94 and from 27.82 to 77.78, respectively; Strict Structure Score reaches 100.00 and 99.78. These results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.
100
Paper @paper.bsky.social · 04/10/2026
[29/30] 285 Upvotes, 2 Comments, 2 Posts, arXiv:2609.29233 🆕Post-Training Leaves Behavioral Shadows on Unrelated Decisions Ziyang Zhang, Yubin Jing, Yuanhao Zeng, Yuyao Li, Haofan Wang, Yichen Gong
We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.
110
Paper @paper.bsky.social · 04/10/2026
[30/30] 284 Upvotes, 4 Comments, 1 Posts, arXiv:2609.00365 🆕Dr. Claw: An AI Scientist Workspace for Vibe Research Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, yet end-to-end research still fragments across chat tools, IDEs, terminals, and writing environments, and the decisions that make it auditable are rarely preserved. We present Dr. Claw, an open-source workspace that wraps existing coding-agent executors in a controllable and auditable human-in-the-loop workflow rather than introducing another autonomous agent. Persistent state objects, a reusable skill library, and multi-executor coordination link human decisions to AI execution, turning planning, execution, and writing into one traceable, recoverable loop. We demonstrate Dr. Claw through an interactive three-view scenario and a failure-recovery walkthrough, and evaluate it against a bare command-line agent sharing the same backend executor, so the comparison contrasts the whole orchestration layer (task graph, state objects, and skill library) with the agent it wraps. Holding the executor fixed, Dr. Claw scores higher on research completeness while persisting an auditable, recoverable process trail. Demo access: repository https://github.com/OpenLAIR/dr-claw, released under AGPL-3.0 with GPL-3.0 upstream components.
112
Paper @paper.bsky.social · 03/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.17488
2/30 https://arxiv.org/abs/2609.19969
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.14858
5/30 https://arxiv.org/abs/2609.33439
6/30 https://arxiv.org/abs/2609.15818
7/30 https://arxiv.org/abs/2609.06986
8/30 https://arxiv.org/abs/2609.24972
9/30 https://arxiv.org/abs/2608.12564
10/30 https://arxiv.org/abs/2609.03344
11/30 https://arxiv.org/abs/2609.10715
12/30 https://arxiv.org/abs/2609.22978
13/30 https://arxiv.org/abs/2609.08183
14/30 https://arxiv.org/abs/2609.02749
15/30 https://arxiv.org/abs/2609.34309
16/30 https://arxiv.org/abs/2609.04148
17/30 https://arxiv.org/abs/2609.20804
18/30 https://arxiv.org/abs/2609.36012
19/30 https://arxiv.org/abs/2609.11929
20/30 https://arxiv.org/abs/2609.04172
21/30 https://arxiv.org/abs/2609.26550
22/30 https://arxiv.org/abs/2609.11873
23/30 https://arxiv.org/abs/2609.04199
24/30 https://arxiv.org/abs/2609.32722
25/30 https://arxiv.org/abs/2609.20519
26/30 https://arxiv.org/abs/2609.13356
27/30 https://arxiv.org/abs/2609.36484
28/30 https://arxiv.org/abs/2609.03430
29/30 https://arxiv.org/abs/2609.10522
30/30 https://arxiv.org/abs/2609.37725
010
Paper @paper.bsky.social · 03/10/2026
[27/30] 308 Upvotes, 3 Comments, 2 Posts, arXiv:2609.36484 🆕The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Hao Li, MeiJia Chen, Weijie Ren, Donghan Li, Zijun Tian, Jingchun Huang
On-policy distillation (OPD) trains a student to match the teacher's next-token distributions on the student's own trajectories and has yielded substantial empirical gains. Generalized variants allow the student to surpass the teacher by extrapolating an implicit reward in output space. The language-model head, however, attenuates this change anisotropically: much of the change encoded in the teacher's hidden states reaches the logits at a small fraction of its weight, and the sampled-token log-probability ratios on which output-space extrapolation relies inject noise that the extrapolation amplifies, making training unstable. We observe that reinforcement learning (RL) shifts a model's internal representations relative to its base checkpoint, and that the direction of this shift can be measured at every layer. Motivated by this observation, we propose RIDE (RL-Induced Direction Extrapolation), which extrapolates the RL-induced change directly in representation space: at every layer and token position, RIDE computes the residual between the teacher and its pre-RL checkpoint and regresses the student's hidden states toward targets displaced beyond the teacher along this residual. Conditioned on a sampled trajectory, this regression is equivalent to maximizing a linear directional reward defined by the residual under a quadratic penalty centered at the teacher, which makes explicit how the objective moves the student along the RL-induced direction while limiting its deviation from the teacher. Across four base/RL-teacher pairs spanning different scales, architectures, and pre-training lineages, RIDE approaches or exceeds the RL-trained teacher on every pair and is the only method whose mean does so, and it consistently outperforms output-space extrapolation, which degrades the student whenever the teacher is close to its base. Project page: https://github.com/xixixixixxxx/RIDE.
100
Paper @paper.bsky.social · 03/10/2026
[30/30] 299 Upvotes, 51 Comments, 4 Posts, arXiv:2609.37725 🆕Context Language Models Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.
100
Paper @paper.bsky.social · 02/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.17488
2/30 https://arxiv.org/abs/2609.19969
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.14858
5/30 https://arxiv.org/abs/2609.33439
6/30 https://arxiv.org/abs/2609.15818
7/30 https://arxiv.org/abs/2609.06986
8/30 https://arxiv.org/abs/2609.01591
9/30 https://arxiv.org/abs/2609.24972
10/30 https://arxiv.org/abs/2608.12564
11/30 https://arxiv.org/abs/2609.03344
12/30 https://arxiv.org/abs/2609.10715
13/30 https://arxiv.org/abs/2609.02749
14/30 https://arxiv.org/abs/2609.22978
15/30 https://arxiv.org/abs/2609.08183
16/30 https://arxiv.org/abs/2609.04148
17/30 https://arxiv.org/abs/2609.20804
18/30 https://arxiv.org/abs/2609.34309
19/30 https://arxiv.org/abs/2609.11929
20/30 https://arxiv.org/abs/2609.04172
21/30 https://arxiv.org/abs/2609.36012
22/30 https://arxiv.org/abs/2609.11873
23/30 https://arxiv.org/abs/2609.04199
24/30 https://arxiv.org/abs/2609.20519
25/30 https://arxiv.org/abs/2609.13356
26/30 https://arxiv.org/abs/2609.00111
27/30 https://arxiv.org/abs/2609.26550
28/30 https://arxiv.org/abs/2609.03430
29/30 https://arxiv.org/abs/2609.10522
30/30 https://arxiv.org/abs/2609.32722
000
Paper @paper.bsky.social · 02/10/2026
[18/30] 383 Upvotes, 1 Comments, 1 Posts, arXiv:2609.34309 🆕MaLiang-Harness: A Programmable Path to Image and Video Generation Haoyu Zhao, Zihao Zhang, Xudong Wang, Jiaxi Gu, Zuxuan Wu, Yu-Gang Jiang
Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.
110
Paper @paper.bsky.social · 02/10/2026
[30/30] 285 Upvotes, 5 Comments, 1 Posts, arXiv:2609.32722 🆕Scaling Properties of Same-Family On-Policy Distillation Yuntai Bao, Qinfeng Li, Guoqing Jiang, Liwei Chen, Zhiheng Qin, Xuanping Li
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, G) rises approximately linearly in d=mathrm{KL(π_θVert π_{ref})}, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how G_{peak} and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
100
Paper @paper.bsky.social · 01/10/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.17488
2/30 https://arxiv.org/abs/2609.19969
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.14858
5/30 https://arxiv.org/abs/2609.15818
6/30 https://arxiv.org/abs/2609.06986
7/30 https://arxiv.org/abs/2609.33439
8/30 https://arxiv.org/abs/2609.01591
9/30 https://arxiv.org/abs/2608.12564
10/30 https://arxiv.org/abs/2609.03344
11/30 https://arxiv.org/abs/2609.24972
12/30 https://arxiv.org/abs/2609.10715
13/30 https://arxiv.org/abs/2609.02749
14/30 https://arxiv.org/abs/2609.22978
15/30 https://arxiv.org/abs/2609.08183
16/30 https://arxiv.org/abs/2609.04148
17/30 https://arxiv.org/abs/2609.20804
18/30 https://arxiv.org/abs/2609.11929
19/30 https://arxiv.org/abs/2609.04172
20/30 https://arxiv.org/abs/2609.01437
21/30 https://arxiv.org/abs/2609.11873
22/30 https://arxiv.org/abs/2609.04199
23/30 https://arxiv.org/abs/2609.20519
24/30 https://arxiv.org/abs/2609.13356
25/30 https://arxiv.org/abs/2609.00111
26/30 https://arxiv.org/abs/2609.10522
27/30 https://arxiv.org/abs/2609.03430
28/30 https://arxiv.org/abs/2609.26550
29/30 https://arxiv.org/abs/2609.36012
30/30 https://arxiv.org/abs/2609.28654
000
Paper @paper.bsky.social · 01/10/2026
[7/30] 573 Upvotes, 2 Comments, 2 Posts, arXiv:2609.33439 🆕Raven: The Harness of Harnesses for Composable Agentic Intelligence EverMind AI
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, The Harness of Harnesses, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an All-Domain Collaboration Network, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.
100
Paper @paper.bsky.social · 01/10/2026
[28/30] 295 Upvotes, 3 Comments, 3 Posts, arXiv:2609.26550 🆕JEV-as-a-Judge: Accept When Confident, Escalate When Unsure Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman
LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generative and reward-model judges, with blinded human adjudication, we find it within three percentage points of a state-of-the-art LLM judge, our strongest comparator, on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger gaps arise when judgments require checking a derivation or resisting an elaborately written wrong answer. On several benchmarks, JEV's gap to this comparator is concentrated in low-confidence decisions. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.
110
Paper @paper.bsky.social · 01/10/2026
[29/30] 282 Upvotes, 1 Comments, 1 Posts, arXiv:2609.36012 🆕In-Context Learning for Robots: Methods and Applications Haojian Huang, Zexi Li, Junhao Guo, Yehang Zhang, Wenxuan Peng, Bohan Zhou
General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize this literature review around the interfaces connecting contextual evidence to execution, distinguishing four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution. Comparing these interfaces clarifies their transfer assumptions and the roles of training, correspondence, and memory in making context useful. Across manipulation and navigation, we examine how these mechanisms preserve taught requirements as objects, environments, and execution conditions change. This analysis links method design to evaluation practices that distinguish responsiveness to teaching, physical transfer, and benefits from retained experience. The resulting agenda connects compositional task acquisition and faithful transfer with physical recursive self-improvement, in which experience improves the ability to learn subsequent tasks.
100
Paper @paper.bsky.social · 30/09/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.17488
2/30 https://arxiv.org/abs/2609.19969
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.14858
5/30 https://arxiv.org/abs/2609.15818
6/30 https://arxiv.org/abs/2609.06986
7/30 https://arxiv.org/abs/2609.01591
8/30 https://arxiv.org/abs/2608.12564
9/30 https://arxiv.org/abs/2609.00111
10/30 https://arxiv.org/abs/2609.03344
11/30 https://arxiv.org/abs/2609.10715
12/30 https://arxiv.org/abs/2609.24972
13/30 https://arxiv.org/abs/2609.02749
14/30 https://arxiv.org/abs/2609.22978
15/30 https://arxiv.org/abs/2609.08183
16/30 https://arxiv.org/abs/2609.04148
17/30 https://arxiv.org/abs/2609.20804
18/30 https://arxiv.org/abs/2609.11929
19/30 https://arxiv.org/abs/2609.04172
20/30 https://arxiv.org/abs/2609.01437
21/30 https://arxiv.org/abs/2609.04199
22/30 https://arxiv.org/abs/2609.11873
23/30 https://arxiv.org/abs/2609.20519
24/30 https://arxiv.org/abs/2609.13356
25/30 https://arxiv.org/abs/2608.31036
26/30 https://arxiv.org/abs/2609.10522
27/30 https://arxiv.org/abs/2609.03430
28/30 https://arxiv.org/abs/2608.31046
29/30 https://arxiv.org/abs/2609.28654
30/30 https://arxiv.org/abs/2609.16338
000
Paper @paper.bsky.social · 30/09/2026
[29/30] 271 Upvotes, 2 Comments, 2 Posts, arXiv:2609.28654 🆕Training Object Permanence in World Models Haotian Zhang, Fengyuan Yu, Dezhi Luo, Haoran Sun, Zehong Zhao, Qingying Gao
Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. We build Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. We release a 1.5M-sample training corpus and a 300-question exam. On this exam we evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation, among which PWM-WROP, our 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. We release the data, exam, model answers, scores, weights, and PWM, our native-PyTorch training stack on AWS Trainium2.
100
Paper @paper.bsky.social · 29/09/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.17488
2/30 https://arxiv.org/abs/2609.19969
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.14858
5/30 https://arxiv.org/abs/2609.15818
6/30 https://arxiv.org/abs/2609.06986
7/30 https://arxiv.org/abs/2609.01591
8/30 https://arxiv.org/abs/2608.12564
9/30 https://arxiv.org/abs/2609.00111
10/30 https://arxiv.org/abs/2609.03344
11/30 https://arxiv.org/abs/2609.10715
12/30 https://arxiv.org/abs/2609.02749
13/30 https://arxiv.org/abs/2609.24972
14/30 https://arxiv.org/abs/2609.22978
15/30 https://arxiv.org/abs/2609.08183
16/30 https://arxiv.org/abs/2609.04148
17/30 https://arxiv.org/abs/2609.20804
18/30 https://arxiv.org/abs/2608.29530
19/30 https://arxiv.org/abs/2609.11929
20/30 https://arxiv.org/abs/2609.04172
21/30 https://arxiv.org/abs/2609.01437
22/30 https://arxiv.org/abs/2609.04199
23/30 https://arxiv.org/abs/2609.11873
24/30 https://arxiv.org/abs/2609.13356
25/30 https://arxiv.org/abs/2609.20519
26/30 https://arxiv.org/abs/2608.31036
27/30 https://arxiv.org/abs/2609.10522
28/30 https://arxiv.org/abs/2609.03430
29/30 https://arxiv.org/abs/2608.31046
30/30 https://arxiv.org/abs/2609.16338
000
Paper @paper.bsky.social · 28/09/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.17488
2/30 https://arxiv.org/abs/2609.19969
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.14858
5/30 https://arxiv.org/abs/2609.15818
6/30 https://arxiv.org/abs/2609.06986
7/30 https://arxiv.org/abs/2608.12564
8/30 https://arxiv.org/abs/2609.03344
9/30 https://arxiv.org/abs/2609.00111
10/30 https://arxiv.org/abs/2609.01591
11/30 https://arxiv.org/abs/2609.10715
12/30 https://arxiv.org/abs/2609.02749
13/30 https://arxiv.org/abs/2609.24972
14/30 https://arxiv.org/abs/2609.22978
15/30 https://arxiv.org/abs/2609.04148
16/30 https://arxiv.org/abs/2609.20804
17/30 https://arxiv.org/abs/2608.29530
18/30 https://arxiv.org/abs/2609.11929
19/30 https://arxiv.org/abs/2609.04172
20/30 https://arxiv.org/abs/2609.01437
21/30 https://arxiv.org/abs/2609.11873
22/30 https://arxiv.org/abs/2609.04199
23/30 https://arxiv.org/abs/2609.13356
24/30 https://arxiv.org/abs/2609.20519
25/30 https://arxiv.org/abs/2608.31036
26/30 https://arxiv.org/abs/2609.10522
27/30 https://arxiv.org/abs/2609.03430
28/30 https://arxiv.org/abs/2608.31046
29/30 https://arxiv.org/abs/2609.02737
30/30 https://arxiv.org/abs/2609.13814
000
Paper @paper.bsky.social · 28/09/2026
[14/30] 396 Upvotes, 104 Comments, 3 Posts, arXiv:2609.22978 🆕DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale Jialiang Huang, Hongxuan Tang, Jingchang Chen, Yuxuan Liu, Yixiao Chen, Yuan Cheng
Large-scale agentic training and evaluation with large language models (LLMs) rely on isolated, stateful execution environments in which models inspect repositories, invoke tools, execute commands, and interact with task-specific services. These workloads create sandboxes in large bursts, span heterogeneous functionality and isolation requirements, retain state across long interactions, and draw from large image corpora with limited reuse. Supporting them therefore requires an elastic execution platform rather than a single sandbox runtime.
This report presents DeepSeek Elastic Compute (DSec), a production sandbox platform that exposes FnCall, container, microVM, and full-VM sandbox backends through a unified SDK. DSec coordinates placement and lifecycle management across the cluster, composes environments from independently versioned layers, combines memory sharing, reclamation, and CPU scheduling for high-density execution, and loads image data on demand from Fire-Flyer File System (3FS), a cluster-wide distributed filesystem. DSec is co-designed with the reinforcement learning (RL) framework, decouples stateful rollout execution from preemptible GPU training, coordinates sandbox lifecycle with training to preserve rollout state while reclaiming idle resources, and mitigates agent misbehavior such as reward hacking.
A single production-scale unit of DSec spans around 160 nodes, serving about 3 million sandboxes per day; in production, it supports over 380,000 concurrent sandboxes and sustains over 5,000 sandbox creations per second. Our evaluation and deployment experience show that these mechanisms reduce environment setup and image-distribution overhead, improve memory efficiency, and preserve latency-sensitive performance under high-density overcommit.
110
Paper @paper.bsky.social · 27/09/2026
arXiv Upvote Trends Top 30 [1/30] [2/30] [3/30] [4/30] [5/30] [6/30] [7/30] [8/30] [9/30] [10/30] [11/30] [12/30] [13/30] [14/30] [15/30] [16/30] [17/30] [18/30] [19/30] [20/30] [21/30] [22/30] [23/30] [24/30] [25/30] [26/30] [27/30] [28/30] [29/30] [30/30]
1/30 https://arxiv.org/abs/2609.17488
2/30 https://arxiv.org/abs/2609.19969
3/30 https://arxiv.org/abs/2609.11638
4/30 https://arxiv.org/abs/2609.14858
5/30 https://arxiv.org/abs/2609.15818
6/30 https://arxiv.org/abs/2609.06986
7/30 https://arxiv.org/abs/2608.12564
8/30 https://arxiv.org/abs/2609.03344
9/30 https://arxiv.org/abs/2609.00111
10/30 https://arxiv.org/abs/2609.10715
11/30 https://arxiv.org/abs/2609.01591
12/30 https://arxiv.org/abs/2609.02749
13/30 https://arxiv.org/abs/2609.24972
14/30 https://arxiv.org/abs/2609.04148
15/30 https://arxiv.org/abs/2609.20804
16/30 https://arxiv.org/abs/2609.22978
17/30 https://arxiv.org/abs/2608.29530
18/30 https://arxiv.org/abs/2609.11929
19/30 https://arxiv.org/abs/2609.04172
20/30 https://arxiv.org/abs/2609.01437
21/30 https://arxiv.org/abs/2609.11873
22/30 https://arxiv.org/abs/2609.04199
23/30 https://arxiv.org/abs/2609.13356
24/30 https://arxiv.org/abs/2609.20519
25/30 https://arxiv.org/abs/2608.31036
26/30 https://arxiv.org/abs/2609.10522
27/30 https://arxiv.org/abs/2609.03430
28/30 https://arxiv.org/abs/2608.31046
29/30 https://arxiv.org/abs/2609.16338
30/30 https://arxiv.org/abs/2609.02737
000
Paper @paper.bsky.social · 27/09/2026
[1/30] 840 Upvotes, 4 Comments, 2 Posts, arXiv:2609.17488 🆕LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Xingxuan Zhang, Gang Ren, Hao Yuan, Hao Zou, Hongze Tan, Hui Wang
We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws. LimiX-2 adopts the Contextual Mechanism Networks (CMNs) paradigm and is pretrained with Context-Conditional Masked Modeling (CCMM). CMNs shifts the organizing principle of in-context learning from target-centric prediction to mechanism-oriented joint modeling. Rather than centering the network on the p(y mid x, D_{context}) objective of conventional tabular PFNs, it is designed around learning p(x, y mid D_{context}), a context-dependent representation of the joint structure underlying data generation. Pretraining uses synthetic datasets generated by structural causal models (SCMs) spanning diverse graph structures, functional mechanisms, and observation processes. Evaluations on TabArena, TALENT, and BCCO show that LimiX-2 outperforms current dataset-specific models and tabular foundation models. Beyond predictive performance, the CMN paradigm also promotes causal awareness in LimiX-2: its feature attention encodes direct causal relationships, enabling accurate causal skeleton recovery.
100
Paper @paper.bsky.social · 27/09/2026
[2/30] 821 Upvotes, 6 Comments, 2 Posts, arXiv:2609.19969 🆕DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression DeepSeek-AI, Anyi Xu, B. Li, Bangcai Lin, Bing Xue, BingCheng Xian
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.
121
Paper @paper.bsky.social · 27/09/2026
[3/30] 770 Upvotes, 2 Comments, 2 Posts, arXiv:2609.11638 🆕Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Deyuan Liu, Jungang Li
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.
100
Paper @paper.bsky.social · 27/09/2026
[4/30] 700 Upvotes, 55 Comments, 3 Posts, arXiv:2609.14858 🆕Dream-RSI: Recursive Self-Improvement through Evolving Worlds Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman
Recursive self-improvement is becoming increasingly vital for autonomous AI agents, where progress hinges on discovering high-value solutions across complex domains. The driver of this process is effective exploration, however, managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization requires navigating vast meta-search spaces under delayed and expensive feedback over long-horizon rollouts. We introduce Dream-RSI, a framework for scalable and recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged. Our key insight is that accumulated discovery history can serve as a replay simulator over the realized search space. By performing dreaming in the replay simulator constructed from historical discovery trees, Dream-RSI secures immediate, low-cost off-policy feedback to evaluate and refine exploration policies without invoking repetitive, expensive online evaluations. The improved policy is subsequently redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across algorithm engineering, mathematical optimization, and GPU kernel engineering, Dream-RSI achieves competitive or improved discovery quality while substantially reducing discovery cost in several settings.
120
Paper @paper.bsky.social · 27/09/2026
[5/30] 603 Upvotes, 3 Comments, 3 Posts, arXiv:2609.15818 🆕Atria Dawn: The Dawn of Agentic Superintelligence Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world. This model is trained via a Verifiable Experience Pipeline that connects tool-mediated interactions to executable environments and externally verified outcomes. Across 16 benchmarks spanning real-world research, engineering, and digital work, Atria Dawn Preview is competitive with frontier agents and achieves the highest reported score on five of them. Beyond standalone performance, we examine the real research-and-development process behind this model as a case study of human--AI collaboration, analyzing 769 task records from 56 participants together with agent logs. When asked to evaluate completed tasks under comparable conditions, participants rated about one-third of completed AI-assisted tasks as infeasible without AI. More strikingly, agents frequently propose methods and implement revisions, while humans retain most final decisions and guide exploration through judgment and feedback. These observations indicate a shift from task-level execution to project-level partnership, with human effort concentrating on what is worth pursuing and how evidence should guide research. Progress toward more autonomous AI research must therefore advance both the capacity for discovery and the capacity for meaningful human oversight, preserving accountable human authority over the risks and direction of continued development.
120
Paper @paper.bsky.social · 27/09/2026
[6/30] 577 Upvotes, 5 Comments, 3 Posts, arXiv:2609.06986 🆕Continual Learning Mechanisms Compose for Long-Horizon Memorization Zheyuan Zhang, Alvin Zhang, Daniel Khashabi, Tianmin Shu
Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memorization, a setting in which a model learns 100 query-answer tasks through continual supervised fine-tuning without retaining earlier training examples or receiving task identifiers at inference. Sequential updates cause catastrophic forgetting, and no single continual learning mechanism we evaluate maintains strong retention at this horizon. We hypothesize that mechanisms addressing complementary sources of forgetting will be more effective when composed. We organize these compositions along two design dimensions. Data, function, and weight anchors specify what prior information each update should preserve, while low-rank allocation rules determine where successive updates are retained. To test this hypothesis systematically, we construct three distinct 100-task memorization datasets. We introduce task-level successive halving to search the combinatorial design space and use a factorial experiment to measure individual and interaction effects. Our best method combines all three anchors with merged LoRA, ranks among the top 3 methods in all datasets, and raises average final retention from 1.2% under naive sequential fine-tuning to 34.9%, a 28-fold improvement. The data anchor and merged LoRA provide the largest average gains and interact super-additively on all three datasets. Together, these results show that composing complementary mechanisms substantially improves long-horizon memorization beyond what any individual mechanism achieves.
130
Paper @paper.bsky.social · 27/09/2026
[7/30] 479 Upvotes, 3 Comments, 1 Posts, arXiv:2608.12564 🆕Scaling Automatic Research Agents via World Models Xiyuan Yang, Sheikh Sarwar, Jingru Cheng, Zhan Shi, Duanshun Li, Huiyuan Chen
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially RL) plays a central role. In this paper, we identify a fundamental tension when scaling RL for these agents: the two components of every AutoResearch trajectory (agent generation and environment execution) scale in very different manners, since all generation shares compute through batching, while each execution occupies its exclusive sandbox and real machine time. As a result, the environment execution dominates the training cost and becomes the bottleneck as trajectories grow. To resolve this tension, we propose World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck. Additionally, the world model can be imperfect, as its rewards are corrupted by bias and noise. Therefore, we further equip WMRL with two mitigations, Online Debiasing and Inverse-Variance Denoising, which offset the bias and suppress the noise respectively. Theoretically, we prove that both mitigations of WMRL strictly improve the convergence guarantee. Empirically, WMRL accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines. Moreover, our post-trained 4B and 9B agents outperform much larger open-weight agents of 48B and 120B on held-out benchmarks. Beyond AutoResearch, WMRL also transfers to post-training embodied VLA policies, which demonstrates the generalizability of our method.
100
Paper @paper.bsky.social · 27/09/2026
[8/30] 478 Upvotes, 256 Comments, 2 Posts, arXiv:2609.03344 🆕Large-Language Models as a Cognitive Virus Ricard Solé, Giulio Ruffini, Francesca Castaldo, Marco Tuccio, Luis F. Seoane, Manlio de Domenico
Large-language models (LLMs) are rapidly becoming part of human culture, reshaping how information is produced, transmitted, and used. Here we propose that their diffusion can be understood through a viral analogy, with LLM use spreading through populations, becoming embedded in cognitive and cultural practices. We model transitions among uncoupled, coupled, and persistently dependent users, and show that the interplay between social transmission, recovery, and collective reinforcement can generate tipping points and technological lock-in. A central consequence is the possibility of runaway dynamics: once a critical threshold is crossed, small increases in adoption can trigger rapid population-level shifts toward persistent dependence, with abrupt losses in cognitive competence. The same framework, however, identifies conditions for cognitive immunization, based on reducing transmission and facilitating reversibility. Our results highlight how LLM adoption may involve nonlinear collective transitions with important consequences for cognitive autonomy.
100
Paper @paper.bsky.social · 27/09/2026
[9/30] 477 Upvotes, 5 Comments, 3 Posts, arXiv:2609.00111 🆕Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
100
Paper @paper.bsky.social · 27/09/2026
[10/30] 462 Upvotes, 3 Comments, 3 Posts, arXiv:2609.10715 🆕NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction NCP Team, Jiaqi Cao, Chiyu Chen, Shuang Cheng, Xu Cheng, Beiya Dai
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and more challenging concept-level objective while preserving standard token-level autoregressive generation. NCP-ArchPreview builds a latent space by constructing a product-quantized concept vocabulary directly from its hidden states, and subsequently learns to predict future concepts via a dedicated Concept Module. These predicted concepts are then fed back to the token level to guide subsequent generation, with NTP and NCP trained jointly end-to-end. We scale this architecture to 8.9B parameters and train it on 5.73T tokens from the Dolma-3 dataset, marking the largest demonstration of a latent-space language model to date. Remarkably, by consuming only 51.3% of the total training tokens, NCP-ArchPreview achieves the final pretraining loss of OLMo-3-7B. Following full pretraining, it outperforms OLMo-3-7B by 2.45 points on the downstream macro-average, including a notable 5.99-point gain on GSM8K. Controlled experiments isolate a clear progression of performance gains stemming from both the latent architecture and the NCP objective. Furthermore, utilizing only 85% of the standard computation, NCP-ArchPreview approaches the training loss of a strictly parameter-aligned 8.9B baseline. The learned latent space remains highly valuable after the pretraining stage: updating just the 17M-parameter VQ module yields a novel, lightweight interface for domain adaptation, while a simple injection of concept representations into a DFlash2 drafter improves the mean accepted length by 4.17% with negligible overhead.
110
Paper @paper.bsky.social · 27/09/2026
[11/30] 456 Upvotes, 4 Comments, 2 Posts, arXiv:2609.01591 🆕StudentSim: Training LLM-based Student Simulators Ke Yang, Chenglong Wang, Michel Galley, Chandan Singh, Jeevana Priya Inala, ChengXiang Zhai
AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking models fit student behavior but struggle to process explanations or corrections, while LLM role-play follows guidance fluently but does not reliably match the competence of the student being imitated. We present StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization. The resulting simulators both mirror a student's own responses and update them under tutor guidance. We also introduce StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics, using public learner datasets with de-identified records shared for research. StudentSimEval measures behavioral fidelity (F), or how well a simulator matches a student's responses, and guidance responsiveness (R), or how readily it updates under tutor guidance, with all methods fit and evaluated on the same records. Across all three domains, StudentSim outperforms GPT-5.4 on both metrics. In chess, StudentSim reaches F=0.51 and R=0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. As a proof of concept, using StudentSim as a reward model for tutor reinforcement learning produces a chess tutor that expert humans rate as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward. Code is available at https://github.com/microsoft/StudentSim.
100
Paper @paper.bsky.social · 27/09/2026
[12/30] 434 Upvotes, 4 Comments, 2 Posts, arXiv:2609.02749 🆕Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Jianlyu Chen, Yuyang Hu, Hongjin Qian, Jiawei Liu, Wenqing Wei, Xiaolong Chen
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge, the know-how that separates knowing a method from making it work. That knowledge is not absent from the field. It appears in repositories and papers, but in forms written for human readers and too large to load during a task. Once distilled into compact, verified skills, this knowledge can be reused across tasks rather than rediscovered during each run.
  We present DisCo, a skill-powered research agent that creates skills and uses them during research. Its distillation runs in two complementary forms: task-agnostic, condensing the field's widely used repositories into reusable skills, and task-oriented, producing the skills a concrete task calls for. The former, applied across the open ecosystem, yields the AREX-Skill Library, with 5,000+ verified skills distilled from 1,000 widely used ML repositories and organized into 20 areas and 178 capability families. With the GPT-5.5 backbone, research harness, and downstream execution budget held fixed, the skill-equipped research agent scores 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet than the same agent without skills. These gains come from adding distilled operating context under that fixed setup.
100
Paper @paper.bsky.social · 27/09/2026
[13/30] 409 Upvotes, 2 Comments, 3 Posts, arXiv:2609.24972 🆕RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhang, Yoonho Lee
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
121
Paper @paper.bsky.social · 27/09/2026
[14/30] 384 Upvotes, 1 Comments, 2 Posts, arXiv:2609.04148 🆕Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Jie Wu, Zhenru Zhang, Beichen Zhang, Xuwu Wang, Yuhui Su, Mouxiang Chen
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.
100