Sign in

Leo Boytsov

@srchvrs.bsky.social
737 followers 169 following 137 posts

Machine learning scientist and engineer speaking πtorch & C++ (ph-D CMU) working on (un)natural language processing, speaking πtorch & C++. Opinions sampled from MY OWN 100T param LM.

PostsRepliesMedia
Reposted by Leo Boytsov
Victor Escorcia @escorciav.bsky.social · 21h
Any sons of Clausius, Boltzmann, or Shannon 😉 over here? Don't despair, embrace entropy 💪 www.norvig.com/chomsky.html
Victor: If we don't embrace complexity & entropy, we're better off frozen-emoji/death-emoji

Leo Boytsov: Peter Norvig called out this issue quite some time ago. Only so many things can have an elegant explanation. We have to embrace complexity AND messiness. Moreover, his main point is that we shouldn’t demand an elegant theory at the cost of leaving out much of the phenomenon.

Link: https://www.norvig.com/chomsky.html
211
Leo Boytsov @srchvrs.bsky.social · 03/10/2026
I do love coding agents. I think it's a great tool if one uses them responsibly. However, it seems to me that a frequent outcome is: You vibe-code, I vibe-hack, Now we've got a pile of cack. We all suffer, nothing works *), Let's enjoy agentic perks! *) or works very unreliably.
020
Leo Boytsov @srchvrs.bsky.social · 16/09/2026
I have made a very "fun" discovery recently. Do not assume agents will call your tool (e.g., CLI) using some pre-written piece of code all the time. Tool-calling code can be generated on the fly, e.g., in Codex. This does introduce an additional failure point: developers.openai.com/api/docs/gui...
developers.openai.com
Programmatic Tool Calling | OpenAI API
Configure Programmatic Tool Calling, control which tools programs can invoke, and resume programs after client-owned function calls.
100
Leo Boytsov @srchvrs.bsky.social · 26/07/2026
Yes I agree it can be very subjective.
000
Leo Boytsov @srchvrs.bsky.social · 26/07/2026
Don't get me wrong, I do still think some baselines are critical if you make certain claims.
110
Leo Boytsov @srchvrs.bsky.social · 26/07/2026
Early conferences accepted even idea papers without much experimentation and sometimes without any experiments. A more modern approach is to make you compare against every possible baseline.
120
Leo Boytsov @srchvrs.bsky.social · 26/07/2026
It doesn't seem to help though. If you keep suggesting improvements, it's just finding additional reasons to reject paper. I think what's really missing: a critical assessment and separation of limitations into very critical and relatively minor.
120
Leo Boytsov @srchvrs.bsky.social · 26/07/2026
Yes, and this task is often done too literally: Train people to find flaws until they can justify paper rejection.
220
Leo Boytsov @srchvrs.bsky.social · 26/07/2026
On the negative side, it is often done incorrectly and only adds more technical debt. Instead of solving systemic issues first, people are just slapping an AI band-aid in the hope to stop a gushing artery. 🟦 searchivarius.org/blog/llm-ski...
searchivarius.org
LLM Skills: A New Source of Technical Debt | searchivarius.org
030
Leo Boytsov @srchvrs.bsky.social · 26/07/2026
🧵LLM Skills: A New Source of Technical Debt. I have some thoughts related to the recent craze to automate everything with agentic skills. On a plus side, I think it can be very useful when done correctly.↩️
120
Leo Boytsov @srchvrs.bsky.social · 20/07/2026
This meticulous study delves into the intricate tapestry of lexical biases in LLM-assisted academic writing. It underscores a nuanced interplay of preferred words, unveiling how they have dramatically enhanced and reshaped the realm of science. www.linkedin.com/feed/update/...
linkedin.com
LLMs Favorite Words in Academic Literature Dramatically Increase Since GPT Era | Alex Glynn posted on the topic | LinkedIn
Just published. You may have heard that LLMs have favorite words: "delve", "underscore", "intricate", and so on. Here we demonstrate that the prevalence of these words in the academic literature has d...
000
Leo Boytsov @srchvrs.bsky.social · 10/07/2026
🔹Now long-running asynchronous agents expose a weakness in GRPO, so this paper brings the critic back, but with several engineering fixes. 🟦 www.linkedin.com/posts/ravid-...
linkedin.com
Just read the new paper from Tsinghua/Z.AI on async RL for agents (arXiv:2607.07508). It comes several weeks after the release of GLM-5.2, in which they mentioned that they use a critic instead of… | ...
Just read the new paper from Tsinghua/Z.AI on async RL for agents (arXiv:2607.07508). It comes several weeks after the release of GLM-5.2, in which they mentioned that they use a critic instead of GRP...
000
Leo Boytsov @srchvrs.bsky.social · 10/07/2026
Interestingly we went full-circle in RL for LLMs and LLM agents: 🔹Initially, OpenAI (and some others) used RLHF with PPO, which requires training a critic (reward) model. 🔹Then, researchers moved from PPO with a critic to critic-free GRPO because critics were expensive and unstable. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/07/2026
Fun fact: Two some of the most influential statistical NLP papers were authored by Brown et al in 1990 & 2020 1990 Statistical Approach to Machine Translation 2020 Language Models are Few-Shot Learners *) It is not the same Brown **) I believe author names in IBM papers were ordered alphabetically
001
Leo Boytsov @srchvrs.bsky.social · 17/06/2026
In other words, the judge does not really work as a comparison engine and decisions can be heavily influenced by model's internal beliefs. I have known this for a while, but my knowledge was anecdotal. I am happy to discover that Lee et al have confirmed this rigorously. arxiv.org/abs/2601.07506
arxiv.org
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about their ability to adhere...
041
Leo Boytsov @srchvrs.bsky.social · 17/06/2026
🧵There is a fundamental issue with reference-based LLM-judges. People implicitly assume a reference-based judge behaves like: score=f(candidate,reference)score=f(candidate,reference) However, the actual behavior is closer to: score=f(candidate,reference,parametric knowledge,priors) ↩️
arxiv.org
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about their ability to adhere...
120
Leo Boytsov @srchvrs.bsky.social · 17/03/2026
"Our mandate is clear: Engineering organizations must move from "accepting the vibes" to "verifying the specs." High-scale reliability requires that we treat AI as a generator of proposals, while maintaining human expertise as the final arbiter of situated judgment and architectural integrity." 🟦
dev.to
The First Evolution of Vibe Coding: Engineering Leadership Report
1. Introduction: The Post-Vibe Era "Vibe Coding"—the practice of prioritizing natural...
010
Leo Boytsov @srchvrs.bsky.social · 17/03/2026
🧵"LLMs are currently "sophisticated autocomplete" tools, not autonomous engineering partners. The 31.7% failure rate and the 13.5x dependency expansion represent a "hidden tax" that can quickly negate any velocity gains." ↩️
211
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
6. Lightweight API Many real workloads don’t need a full distributed system — just a clean, streaming, parallel map loop. A detailed discussion in the blog post: 🟦https://lnkd.in/ehu8HcBA
010
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
5. Boilerplate and cognitive overhead Great libraries like Dask and Ray are powerful, but for a simple map pattern they introduce clusters, schedulers, plugins, actors, futures, and daemons (see examples in the end of the post).↩️
110
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
4. Notebook compatibility Some parallel tools (e.g., Ray actors, certain ways of using multiprocessing) behave unpredictably in Jupyter/IPython environments.↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
3. Size-aware streaming for progress tracking for limited-size input/output Tools that process fixed-size input arrays often fail to expose an output with a usable len(...). Without a known length, progress bars like tqdm can’t estimate remaining time.↩️
110
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
2. True streaming input/output Many frameworks materialize outputs in bulk rather than yielding results as soon as they are ready. This breaks bounded-memory pipelines.↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
Common Pain Points w. Existing Tools. Even for simple parallel loops many libraries fall short 1. Stateful workers with flexible initialization. Most map APIs assume stateless functions. When workers must load models or context before processing, solutions become awkward or require elaborate hacks.↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
In 2024, I finally distilled the pattern into a tiny library: mtasklite. Below is the motivation, the flaws of existing approaches, and what makes mtasklite unique.↩️
110
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
Much later did I explore tools like Dask and Ray (see examples in the end of the post) and found that, for this specific pattern, they still required a lot of complexity and infrastructure for something that should be simple: ↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
and I later explored pqdm (parallel TQDM library) to simplify the code (as well as to visualize progress on long-running loops) when workers were stateless. Those solutions did work but always felt incomplete and/or required repeated boilerplate.↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
🧵For the last seven years, I kept re-implementing the same pattern: A parallel map loop that divides the work among several processes or threads. My very first attempts were built on Python’s standard tools, e.g., multiprocessing.map... ↩️
121
Reposted by Leo Boytsov
Victor Escorcia @escorciav.bsky.social · 23/02/2026
@srchvrs.bsky.social: "This is too little for pre-training, but pre-training nowadays is probably not a bottleneck. For post-training 16M training samples can meaningfully improve performance on a lot of tasks."
121
Leo Boytsov @srchvrs.bsky.social · 01/02/2026
It's a yet another example of how rich gets richer. Seasoned devs benefit from LLM assisted coding but their skills have developed already. So they get the best of both worlds.
010
Leo Boytsov @srchvrs.bsky.social · 22/10/2025
"Meta is cutting around 600 positions out of the several thousand roles in its Superintelligence Labs, the Facebook owner said on Wednesday as it looks to make its artificial intelligence unit more flexible and responsive." www.reuters.com/business/met...
reuters.com
031
Leo Boytsov @srchvrs.bsky.social · 27/08/2025
Paper presents fascinating (as described by AI 😆) case study on how seemingly innocuous modification to neural net can drastically alter its perceived robustness against gradient-based adversarial attacks. searchivarius.org/blog/curious...
000
Leo Boytsov @srchvrs.bsky.social · 27/08/2025
🧵I recently finished my nerdiest computer science paper so far and it was accepted by TMLR: A Curious Case of Remarkable Resilience to Gradient Attacks via Fully Convolutional and Differentiable Front End with a Skip Connection. This work was done while I was at Bosch. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/08/2025
PS: Yes, this is a frontier LLM and it still cannot fully replace an editor.⏹️
000
Leo Boytsov @srchvrs.bsky.social · 09/08/2025
4. Model complains about its own suggestion. 5. Bonus point: of course, often times the complaints are incorrect. If you further poke the model it will likely accept being wrong. Which, in turn, may not mean much because models are also clearly trained to agree with humans as much as possible. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/08/2025
🧵Hot take: LLMs still fail at basic grammar/style checking. A repeating situation that I encounter: 1. Ask a model about an issue. 2. The model suggests some rewrite for clarity/accuracy. Typically it's actually quite good (but watch for factual errors!). 3. Recheck the text again. ↩️
110
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
Results? If read tables correctly, there's only very modest boost in both recall & NDCG, which is within 2%. Given that the procedure requires a second retrieval, it does not seem to worth an effort. 🟦 dl.acm.org/doi/abs/10.1...
dl.acm.org
A Large-Scale Study of Reranker Relevance Feedback at Inference | Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval
000
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
PRF was not forgotten in the neural IR times, but how does it perform really? Revanth Gangi Reddy & colleagues ran a rather thorough experiment and published it SIGIR. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
It was doc2query before doc2query and, in fact, it improved performance (by a few%) of the IBM Watson QA system that beat human champions in Jeopardy! ↩️ research.ibm.com/publications...
research.ibm.com
Statistical source expansion for question answering for CIKM 2011
Statistical source expansion for question answering for CIKM 2011 by Nico Schlaefer et al.
100
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
I think this is a problem of completely unsupervised and blind approach of adding terms to the query. If we had some supervision signal to filter out potentially bad terms, this would work out better. In fact, a supervised approach was previously used to add terms to documents! ↩️
100
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
Fixing this issue produced a sub-topic in the IR community devoted to fixing this issue and identifying cases where performance degrades substantially in advance. Dozens of approaches were proposed, but I do not think it was successful. Why⁉️ ↩️
100
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
PRF tends to improve things on average, but has a rather nasty property of tanking outcomes for some queries rather dramatically: When things go wrong (i.e., unlucky unrelated terms are added to the query), they can go very wrong. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
PRF is an old technique introduced 40 years ago in the SMART system (arguably the first open-source IR system). ↩️ x.com/srchvrs/stat...
x.com
Leo Boytsov on X: "🧵40 years ago the SMART IR system was released. It introduced a few key concepts including vector space interpretation of the retrieval process and the relevance feedback algorithm. I also think it was probably the first open source search engine. ↩️" / X
🧵40 years ago the SMART IR system was released. It introduced a few key concepts including vector space interpretation of the retrieval process and the relevance feedback algorithm. I also think it was probably the first open source search engine. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
🧵Pseudo-relevance feedback (PRF) (also known as blind feedback) is a technique of first retrieving/re-ranking top-k documents and adding some of their words to the initial query. Then, a second retrieval/ranking stage uses an updated query. ↩️
111
Leo Boytsov @srchvrs.bsky.social · 02/07/2025
If you submitted a messy paper, it's pointless to address every little comment and promise fixing it in the final version. 🟦
000
Leo Boytsov @srchvrs.bsky.social · 02/07/2025
Instead, think hard about questions you can ask. What is the main misunderstanding? What will you have to do so that a reviewer will accept your work next time. Which concise questions can you ask to avoid misunderstanding in the future? ↩️
100
Leo Boytsov @srchvrs.bsky.social · 02/07/2025
🧵 Dear (scientific) authors: I am being in the same boat too. However, if you receive a ton of detailed complaints regarding paper quality, do NOT try to address them during the rebuttal phase. It's just a waste of everybody's time. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 27/06/2025
@microsoft.com faces an interesting issue that might affect others selling wrappers around ChatGPT and Claude models: users prefer to use ChatGPT directly rather than engage with Microsoft's Copilot. futurism.com/microsoft-co...
futurism.com
Microsoft Is Having an Incredibly Embarrassing Problem With Its AI
Despite investing tens of billions of dollars into OpenAI, Microsoft is still, apparently, competing with its business partner.
000
Leo Boytsov @srchvrs.bsky.social · 12/06/2025
This is a rather blockbuster piece of news: the @hf.co library is dropping support for both Jax and Tensorflow. www.linkedin.com/posts/lysand...
linkedin.com
I have bittersweet news to share. | Lysandre Debut
I have bittersweet news to share. Yesterday we merged a PR deprecating TensorFlow and Flax support in transformers. Going forward, we're focusing all our efforts on PyTorch to remove a lot of th...
010
Leo Boytsov @srchvrs.bsky.social · 22/05/2025
Humans are creating AGI and you claim that their intelligence is overrated?
020