Sign in

Sumit

@reachsumit.com
284 followers 37 following 3.8K posts

Senior MLE at Meta. Trying to keep up with the Information Retrieval domain! Blog: blog.reachsumit.com Newsletter: recsys.substack.com

PostsRepliesMedia
Sumit @reachsumit.com · 2h
OPERA: Optimizing data pruning for efficient retrieval model adaptation Amazon introduces dynamic pruning that favors high-quality pairs when finetuning dense retrievers, improving both ranking and recall in under half the time. 📝 www.amazon.science/publications... 👨🏽‍💻 github.com/autogluon/au...
amazon.science
OPERA: Optimizing data pruning for efficient retrieval model adaptation
Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. We introduce OPERA1 , a data pruning framework that exploits this heter...
000
Sumit @reachsumit.com · 2h
When LLM-Inferred User Context Adds Value in Production Streaming Recommendation Comcast compares aggregate and LLM-generated user profiles and finds that aggregate wins for habitual users, LLM for exploratory ones, but LLM favors popular items. 📝 arxiv.org/abs/2609.38999
arxiv.org
When LLM-Inferred User Context Adds Value in Production Streaming Recommendation
Contextual information in recommender systems is shifting from static, predefined variables toward latent representations inferred from behavior. Large language models support this shift by rendering ...
000
Sumit @reachsumit.com · 2h
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents Shows self-evolving search agents can reward shared mistakes, and introduces CrossFit, scoring each question with a solver trained on the other half of the sources. 📝 arxiv.org/abs/2609.39102
arxiv.org
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we c...
000
Sumit @reachsumit.com · 2h
RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models Meta introduces an auto-research framework for evolving a generative recommender, where coding agents like Claude Code and Codex cross-check each other to catch bugs. 📝 arxiv.org/abs/2609.39551
arxiv.org
RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models
Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accura...
000
Sumit @reachsumit.com · 2h
Exploring Forum Post Retrieval with Generative Modeling Meta explores generative retrieval for Facebook Forum by transferring semantic IDs from Feed data to fine-tune a 3B LLM, finding that SID depth matters most while GRPO post-training fails to help. 📝 arxiv.org/abs/2609.38646
arxiv.org
Exploring Forum Post Retrieval with Generative Modeling
Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a...
000
Sumit @reachsumit.com · 2h
Learning Multiresolution Relevance for Hierarchical Generative Retrieval Introduces a training objective that teaches generative retrievers how relevance splits across semantic ID branches, at no inference cost. 📝 arxiv.org/abs/2609.39312 👨🏽‍💻 github.com/Nevaeh7/RARS
arxiv.org
Learning Multiresolution Relevance for Hierarchical Generative Retrieval
Generative retrieval with semantic identifiers (SIDs) makes successive decisions over a document hierarchy. Relevant documents for the same query may share coarse prefixes and diverge at finer depths,...
000
Sumit @reachsumit.com · 2h
Residual Trajectory Distillation for Generative Retrieval Feeds residual information discarded when building Semantic IDs back into retrieval training, leaving the index and inference unchanged. 📝 arxiv.org/abs/2609.39319 👨🏽‍💻 github.com/Nevaeh7/iclr...
arxiv.org
Residual Trajectory Distillation for Generative Retrieval
Generative retrieval has emerged as a general retrieval paradigm, representing items with discrete Semantic IDs (SIDs) and retrieving them through autoregressive identifier generation. When SIDs are c...
000
Sumit @reachsumit.com · 2h
Generative End-to-end Ad Retrieval at Douyin ByteDance introduces an end-to-end generative retrieval framework that trains the tokenizer, generator, and reranker together to fix codebook collapse and item collisions. 📝 arxiv.org/abs/2609.39327
arxiv.org
Generative End-to-end Ad Retrieval at Douyin
Generative retrieval reformulates recommendation as the generation of discrete item tokens. However, scaling this paradigm to real-world recommender systems reveals two critical bottlenecks: 1) Repres...
000
Sumit @reachsumit.com · 2h
Decision-Oriented Recommendation Reranking: An Empirical Study of Jev Tests Jev for reranking and finds that quality stays strong and latency grows more slowly than Qwen rerankers, but it is slower than recommendation-specific models. 📝 arxiv.org/abs/2609.40241
arxiv.org
Decision-Oriented Recommendation Reranking: An Empirical Study of Jev
Large language models (LLMs) have shown promise for recommendation reranking, but their use introduces an important tradeoff between recommendation quality and serving efficiency. We investigate wheth...
000
Sumit @reachsumit.com · 30/09/2026
LongCat-DeepResearch Technical Report Meituan introduces a deep research system that drafts a compact research plan, has parallel agents research and write a section, then edits the assembled report. 📝 arxiv.org/abs/2609.36071 👨🏽‍💻 github.com/meituan-long...
arxiv.org
LongCat-DeepResearch Technical Report
We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separat...
000
Sumit @reachsumit.com · 30/09/2026
FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation Softens hard Top-1 codeword assignment in semantic ID learning so gradients reach the whole codebook, easing codebook collapse while keeping IDs discrete. 📝 arxiv.org/abs/2609.36670
arxiv.org
FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation
A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundame...
000
Sumit @reachsumit.com · 30/09/2026
Soft Curriculum Learning for Optimizing Fresh and Generalized Recommendations Google breaks the popularity feedback loop by re-weighting the loss instead of filtering data, shifting from easy head items to hard tail ones. 📝 arxiv.org/abs/2609.35783
arxiv.org
Soft Curriculum Learning for Optimizing Fresh and Generalized Recommendations
Large-scale recommender systems, particularly short-form video platforms, are often bottlenecked by massive popularity feedback loops. In such environments, as models recommend popular items, they gen...
000
Sumit @reachsumit.com · 30/09/2026
BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning Jointly trains a search agent's LLM and retriever, adapting the retriever first, via a memory-efficient bilevel method. 📝 arxiv.org/abs/2609.36505 👨🏽‍💻 jenniferquanxiao.github.io/BRIDGE/
arxiv.org
BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning
Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. Ho...
000
Sumit @reachsumit.com · 30/09/2026
GRP v0.1 Technical Report Snap Inc. describes its generative recommender setup, where one encoder-decoder model retrieves, ranks, and is tuned with reinforcement learning. 📝 arxiv.org/abs/2609.36688
arxiv.org
GRP v0.1 Technical Report
Industrial recommendation systems rely on multi-stage cascades whose retrieval, ranking, and serving components are difficult to replace jointly. We present GRP, a generative recommendation framework ...
000
Sumit @reachsumit.com · 30/09/2026
HELIX: Purified and Unified - Rethinking Feature Interaction and Sequence Modeling for Large-Scale Recommendation ByteDance introduces a ranker interleaving sequence retrieval and feature interaction with reusable user-side compute, deployed at TikTok. 📝 arxiv.org/abs/2609.37183
arxiv.org
HELIX: Purified and Unified - Rethinking Feature Interaction and Sequence Modeling for Large-Scale Recommendation
Industrial recommendation ranking models typically scale along two modeling axes: feature interaction over heterogeneous user, item, context, and cross features, and sequence modeling over long, infor...
000
Sumit @reachsumit.com · 30/09/2026
Follow the Entities: A Corpus Map for Agentic Search Introduces an offline-built layer linking documents via recurring entities, so agents can follow entity pages to connected evidence instead of rediscovering relationships per query. 📝 arxiv.org/abs/2609.37226
arxiv.org
Follow the Entities: A Corpus Map for Agentic Search
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirem...
000
Sumit @reachsumit.com · 30/09/2026
Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG Introduces a lightweight local scorer that judges before generation whether retrieved passages can answer a question, using three complementary signals. 📝 arxiv.org/abs/2609.37469
arxiv.org
Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG
Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when ins...
010
Sumit @reachsumit.com · 30/09/2026
MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment Merges query expansions from three small open-source LLMs into a single query for BM25, with every prompt tuned automatically against retrieval quality rather than hand-tuned per LLM. 📝 arxiv.org/abs/2609.37574
arxiv.org
MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment
Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any...
010
Sumit @reachsumit.com · 30/09/2026
Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3 Finds LLM query expansion still helps SPLADE-v3 on scientific collections, mostly through added vocabulary, while a corpus-built concept graph does not. 📝 arxiv.org/abs/2609.37911
arxiv.org
Generated Query Expansion Still Helps Strong Sparse Retrieval: A Controlled Study with SPLADE-v3
Scientific queries are often brief, while relevant papers use specialized vocabulary. Generated query expansion can bridge this mismatch, but earlier work suggests that its value shrinks as the underl...
000
Sumit @reachsumit.com · 30/09/2026
Effective Dense Retrieval using Only In-Context Examples Shows that an LLM can build dense retrieval embeddings with no training, by prompting it with a few query-document examples. 📝 arxiv.org/abs/2609.38099 👨🏽‍💻 github.com/nourj98/RICE
arxiv.org
Effective Dense Retrieval using Only In-Context Examples
Turning decoder-only large language models (LLMs) into strong dense retrievers typically requires some form of retriever training. In this paper, we ask whether LLMs can instead be prompted to produce...
001
Sumit @reachsumit.com · 29/09/2026
Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs Improves conditional text embeddings from LLMs without training by contrasting with a condition-masked embedding in one attention layer. 📝 arxiv.org/abs/2609.32684
arxiv.org
Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs
Extracting conditional text embeddings from large language models (LLMs) is a promising paradigm, as it requires neither additional data nor fine-tuning. Existing methods incorporate conditions into p...
000
Sumit @reachsumit.com · 29/09/2026
IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents Runs a teacher's proposed query and the student's own, then distills only the proposals that retrieve better evidence for training search agents. 📝 arxiv.org/abs/2609.32694 👨🏽‍💻 github.com/PhilipGAQ/ig...
arxiv.org
IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents
On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search ...
000
Sumit @reachsumit.com · 29/09/2026
Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents Builds multi-hop training questions from knowledge graph chains and rewards questions only when the full evidence beats every shortcut context. 📝 arxiv.org/abs/2609.33565
arxiv.org
Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents
A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating quest...
000
Sumit @reachsumit.com · 29/09/2026
Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale Spotify generates natural-language taste profiles for millions of users that complement behavioral embeddings and enable steering. 📝 arxiv.org/abs/2609.35285
arxiv.org
Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale
Foundation model recommender systems require user context that can be consumed by large language models, reasoned over, and refined through natural-language interaction. Traditional behavioral embeddi...
000
Sumit @reachsumit.com · 29/09/2026
AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research Presents a setwise reranker that picks complete, non-redundant sets, trained with hints matched to rollout quality. 📝 arxiv.org/abs/2609.32472 👨🏽‍💻 adatutorank.github.io
arxiv.org
AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely cons...
000
Sumit @reachsumit.com · 29/09/2026
Data Processing for Offline Evaluation in Recommender Systems: a Survey Surveys data preparation for offline recommender evaluation, from filtering and multimodal features to splitting, and finds inconsistent practices. 📝 arxiv.org/abs/2609.31696
arxiv.org
Data Processing for Offline Evaluation in Recommender Systems: a Survey
Offline evaluation is the dominant experimental paradigm in recommender systems research, enabling reproducible and cost-effective comparisons on historical interaction data. Yet, while considerable a...
000
Sumit @reachsumit.com · 29/09/2026
DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation Capital One introduces dynamic patching for long-sequence recommendation, splitting user histories into variable-length patches at behavioral shifts. 📝 arxiv.org/abs/2609.32215
arxiv.org
DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation
Transformers have redefined sequential recommendation by effectively modeling dynamic user behaviors and long-range dependencies. However, they remain inherently inefficient: standard architectures op...
000
Sumit @reachsumit.com · 29/09/2026
When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections Gives a bias-variance theory of when separate query and document projections beat a shared one, plus a training-data selector to pick one. 📝 arxiv.org/abs/2609.32488
arxiv.org
When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections
Dense retrieval powers retrieval-augmented generation, semantic search, and question answering, yet the theoretical basis for choosing between shared and dual query-document projections remains unclea...
000
Sumit @reachsumit.com · 29/09/2026
RandSlot: Learning Compact Visual Document Representations with Random Soft Tokens Introduces random soft tokens as extra training inputs, letting a visual document retriever compress each page into four vectors and outperform prior compact methods. 📝 arxiv.org/abs/2609.32699
arxiv.org
RandSlot: Learning Compact Visual Document Representations with Random Soft Tokens
Visual document retrieval requires expressive representations to match queries with evidence distributed across text, tables, and page layouts. Multi-vector representations capture fine-grained inform...
000
Sumit @reachsumit.com · 29/09/2026
Mend the Measurement Gap: Latent User Preference Modeling for Short-Form Video Recommendation Google models watch time, likes and other feedback as noisy signals of latent preference, with a baseline path absorbing duration bias. 📝 arxiv.org/abs/2609.32839
arxiv.org
Mend the Measurement Gap: Latent User Preference Modeling for Short-Form Video Recommendation
Recommender systems rely heavily on heterogeneous behavioral feedback to infer user preference. Although abundant, these signals are imperfect measurements: the same observed behavior can arise from d...
000
Sumit @reachsumit.com · 29/09/2026
Tessera: Demand-Driven KV Cache Management for Retrieval-Augmented LLM Serving Introduces a serving system that uses retrieval results to plan KV cache retention and request routing, so recurring RAG chunks and agent memory records actually get reused. 📝 arxiv.org/abs/2609.32999
arxiv.org
Tessera: Demand-Driven KV Cache Management for Retrieval-Augmented LLM Serving
RAG and retrieval-based agent memory both inject retrieved content into LLM prompts, as document chunks and recalled memory records, respectively. The same content can recur across requests at differe...
000
Sumit @reachsumit.com · 29/09/2026
Learning Multimodal Embeddings with Evidence-Aligned Readout Tencent uses a multimodal LLM to generate evidence in five semantic units and pools its hidden states at each unit boundary into one embedding, keeping single-vector retrieval. 📝 arxiv.org/abs/2609.33659
arxiv.org
Learning Multimodal Embeddings with Evidence-Aligned Readout
Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether ...
000
Sumit @reachsumit.com · 29/09/2026
Beyond the Beam: Constructive Repair and Candidate Completion for Generative Recommendation Studies when reassigning item identifiers can bring new catalog items into a generative recommender's results. 📝 arxiv.org/abs/2609.33745 👨🏽‍💻 anonymous.4open.science/r/bb_code-80...
arxiv.org
Beyond the Beam: Constructive Repair and Candidate Completion for Generative Recommendation
Generative recommenders retrieve items by generating identifiers, but a valid identifier can remain outside the beam after catalog expansion. This raises two connected questions: which failures can id...
000
Sumit @reachsumit.com · 29/09/2026
SPRINT: Single-Step Generative Recommendation via Average Probability Velocity Generates an item's semantic ID in one forward pass via average probability velocity, with a contrastive loss that keeps tokens coherent. 📝 arxiv.org/abs/2609.34306 👨🏽‍💻 github.com/iamZhuoCai/S...
arxiv.org
SPRINT: Single-Step Generative Recommendation via Average Probability Velocity
Semantic ID (SID) based generative recommendation represents each item as a sequence of discrete tokens, and recommends by generating the SID of the item a user would like to interact with. Both domin...
000
Sumit @reachsumit.com · 29/09/2026
Eval4DiRec: A Unified and Systematic Evaluation Framework for Diffusion-based Recommender Systems Introduces an open-source framework to fairly benchmark 14 diffusion-based recommenders across five different scenarios. 📝 arxiv.org/abs/2609.34404 👨🏽‍💻 github.com/wangcong2001...
arxiv.org
Eval4DiRec: A Unified and Systematic Evaluation Framework for Diffusion-based Recommender Systems
Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attr...
000
Sumit @reachsumit.com · 29/09/2026
EvoSkillRec: Skill-Genome Evolution for Recommender Architecture Discovery Introduces a framework that turns recommenders into reusable skills and evolves new architectures by recombining skills and inventing code. 📝 arxiv.org/abs/2609.34552 👨🏽‍💻 github.com/Xiaopengli1/...
arxiv.org
EvoSkillRec: Skill-Genome Evolution for Recommender Architecture Discovery
Modern recommender systems advance not only by scaling data and parameters, but also by encoding task-specific inductive biases through architecture, including sparse feature interactions for click-th...
000
Sumit @reachsumit.com · 29/09/2026
ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport Distills multi-vector visual retrievers into a text-only query encoder via optimal transport, without pages. 📝 arxiv.org/abs/2609.34899 👨🏽‍💻 github.com/Ryenhails/Na...
arxiv.org
ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport
Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small...
000
Sumit @reachsumit.com · 29/09/2026
PEAR: Progressive Evidence-Based AutoResearch for Industrial Search Systems ByteDance presents an AutoResearch approach for industrial search that tracks hypotheses and promotes candidates through four evaluation levels, saving online tests for the best. 📝 arxiv.org/abs/2609.35031
arxiv.org
PEAR: Progressive Evidence-Based AutoResearch for Industrial Search Systems
AutoResearch improves systems through iterative experimentation: agents propose candidate modifications, evaluate them, and use the results to guide subsequent exploration. Applying this paradigm to i...
000
Sumit @reachsumit.com · 29/09/2026
RenderRank: Learning to Rerank Text with Compressed Visual Tokens Introduces a reranker that renders documents as images and scores them from compressed visual tokens, cutting input length while staying competitive with text-based rerankers on BEIR. 📝 arxiv.org/abs/2609.35069
arxiv.org
RenderRank: Learning to Rerank Text with Compressed Visual Tokens
Rendering document text as images allows vision-language models to encode documents as visual tokens, which can reduce input sequence length compared with text input. This reduction in input length is...
000
Sumit @reachsumit.com · 29/09/2026
Can Generative Retrievers Learn Semantic IDs Without Forgetting How to Speak? Pairs semantic ID learning with a distillation-based regularizer so generative retrievers keep their language generation ability while retrieving. 📝 arxiv.org/abs/2609.35430
arxiv.org
Can Generative Retrievers Learn Semantic IDs Without Forgetting How to Speak?
Generative retrieval (GR) enables end-to-end retrieval by generating document semantic identifiers (SIDs). However, retrieval-only fine-tuning can over-specialize pretrained language models to SID pre...
000
Sumit @reachsumit.com · 29/09/2026
Rubric-Calibrated Preferences: Cross-Query Calibration of LLM Judgments via Item Response Theory Cohere calibrates LLM relevance judgments across queries using a rubric and IRT, giving denser labels for reranker evaluation. 📝 arxiv.org/abs/2609.35739 👨🏽‍💻 github.com/cohere-ai/rc...
arxiv.org
Rubric-Calibrated Preferences: Cross-Query Calibration of LLM Judgments via Item Response Theory
Rerankers decide which documents users and LLMs see, yet their standard metric, nDCG, relies on human relevance labels that are costly, sparse, noisy, and discretely graded. As rerankers approach each...
000
Sumit @reachsumit.com · 28/09/2026
ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker Introduces open e-commerce rerankers (0.6B to 8B) trained on LLM-judged shopping preferences, so they respect hard constraints like budget and product type. 📝 arxiv.org/abs/2609.31002 👨🏽‍💻 github.com/SerendipityO...
arxiv.org
ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker
Open rerankers trained for general web retrieval transfer imperfectly to e-commerce, where ranking decisions depend not only on topical relevance but also on user preferences, product constraints, and...
000
Sumit @reachsumit.com · 28/09/2026
Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops Spotify turns single-turn prompts into multi-turn dialogs to test a chat recommender pre-launch. 📝 arxiv.org/abs/2609.30297
arxiv.org
Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops
Conversational recommendation agents are a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., "recommend Italian indie artists I haven't hear...
000
Sumit @reachsumit.com · 28/09/2026
T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation Shopify introduces a RoPE variant for generative recommenders that rotates attention by real timestamps at multiple time scales, capturing both recency and seasonal patterns. 📝 arxiv.org/abs/2609.30576
arxiv.org
T-RoPE: Time-Aware Rotary Position Embedding for Sequential Recommendation
Large-scale recommenders increasingly adopt the sequential generative recipe behind large language models, bringing the Transformer into recommendation along with design choices made for text, includi...
000
Sumit @reachsumit.com · 28/09/2026
Embedding Subspace Partitioning for Dynamic Multi-Objective Retrieval LinkedIn splits the embedding into per-objective subspaces scored with a weighted sum, so trade-offs like relevance vs. revenue can be tuned at serving time without retraining. 📝 arxiv.org/abs/2609.30601
arxiv.org
Embedding Subspace Partitioning for Dynamic Multi-Objective Retrieval
Modern industrial recommender systems must optimize across competing objectives, balancing semantic relevance with business metrics such as engagement and revenue. While bi-encoders dominate large-sca...
000
Sumit @reachsumit.com · 28/09/2026
Component Benchmark: Hierarchical Model Profiling for Large-scale Recommendation Systems Meta presents a tool that profiles each submodule of a large recommendation model in isolation and shows latency, memory, and MFU in an interactive tree view. 📝 arxiv.org/abs/2609.30656
arxiv.org
Component Benchmark: Hierarchical Model Profiling for Large-scale Recommendation Systems
Large-scale recommendation models pose distinct, under-explored profiling challenges. Most recommendation model architectures are structurally heterogeneous, intermixing memory-bandwidth-bound operati...
000
Sumit @reachsumit.com · 28/09/2026
Recommendation World Models for Future-State Control Introduces a plug-in layer for trained sequential recommenders that predicts how candidate slates will shape future user behavior, then picks one that steers toward a target while keeping relevance. 📝 arxiv.org/abs/2609.30711
arxiv.org
Recommendation World Models for Future-State Control
Sequential recommendation optimizes which items to rank, while each displayed slate also shapes subsequent feedback and user state. We study how a trained ranker can support decisions about these futu...
000
Sumit @reachsumit.com · 28/09/2026
RecToolBench: Benchmarking Recommendation-Specific Tool Orchestration under Fuzzy User Intent Benchmarks how recommender agents coordinate MCP tools under vague user requests, from single tool calls to multi-step workflows. 📝 arxiv.org/abs/2609.30717 👨🏽‍💻 github.com/ShawnChenn/R...
arxiv.org
RecToolBench: Benchmarking Recommendation-Specific Tool Orchestration under Fuzzy User Intent
Recent advances in agentic recommender systems are shifting recommender systems from passive filtering engines to instruction-following agents that use external tools to resolve user intent. However, ...
000
Sumit @reachsumit.com · 28/09/2026
QReason: Query-Focused Decoupled Chain-of-Thought for Efficient Passage Reranking Trains a rewriter to generate a reasoning-style query once and reuse it across all reranking windows, so a non-reasoning LLM reranker avoids repeated chain-of-thought. 📝 arxiv.org/abs/2609.30904
arxiv.org
QReason: Query-Focused Decoupled Chain-of-Thought for Efficient Passage Reranking
Passage reranking plays a crucial role in information retrieval by refining the ordering of candidate passages to better reflect relevance. Existing listwise LLM rerankers with Chain-of-Thought (CoT) ...
000
Sumit @reachsumit.com · 28/09/2026
KuaFu: Compressing Long User Behavior into Understanding at Billion Scale Tencent presents an LLM-based user profiling system that compresses each behavior item into a few cacheable tokens, with RL training to curb hallucinated user interests. 📝 arxiv.org/abs/2609.31045
arxiv.org
KuaFu: Compressing Long User Behavior into Understanding at Billion Scale
Conversational agents, generative recommenders, and personalized advertising all rest on one capability: understanding each user from raw behavior. Prevailing industrial practice is task-specific: for...
000