Sign in

Leo Boytsov

@srchvrs.bsky.social
738 followers 169 following 137 posts

Machine learning scientist and engineer speaking πtorch & C++ (ph-D CMU) working on (un)natural language processing, speaking πtorch & C++. Opinions sampled from MY OWN 100T param LM.

PostsRepliesMedia
Leo Boytsov @srchvrs.bsky.social · 03/10/2026
I do love coding agents. I think it's a great tool if one uses them responsibly. However, it seems to me that a frequent outcome is: You vibe-code, I vibe-hack, Now we've got a pile of cack. We all suffer, nothing works *), Let's enjoy agentic perks! *) or works very unreliably.
020
Leo Boytsov @srchvrs.bsky.social · 16/09/2026
I have made a very "fun" discovery recently. Do not assume agents will call your tool (e.g., CLI) using some pre-written piece of code all the time. Tool-calling code can be generated on the fly, e.g., in Codex. This does introduce an additional failure point: developers.openai.com/api/docs/gui...
developers.openai.com
Programmatic Tool Calling | OpenAI API
Configure Programmatic Tool Calling, control which tools programs can invoke, and resume programs after client-owned function calls.
100
Leo Boytsov @srchvrs.bsky.social · 26/07/2026
🧵LLM Skills: A New Source of Technical Debt. I have some thoughts related to the recent craze to automate everything with agentic skills. On a plus side, I think it can be very useful when done correctly.↩️
120
Leo Boytsov @srchvrs.bsky.social · 20/07/2026
This meticulous study delves into the intricate tapestry of lexical biases in LLM-assisted academic writing. It underscores a nuanced interplay of preferred words, unveiling how they have dramatically enhanced and reshaped the realm of science. www.linkedin.com/feed/update/...
linkedin.com
LLMs Favorite Words in Academic Literature Dramatically Increase Since GPT Era | Alex Glynn posted on the topic | LinkedIn
Just published. You may have heard that LLMs have favorite words: "delve", "underscore", "intricate", and so on. Here we demonstrate that the prevalence of these words in the academic literature has d...
000
Leo Boytsov @srchvrs.bsky.social · 10/07/2026
Interestingly we went full-circle in RL for LLMs and LLM agents: 🔹Initially, OpenAI (and some others) used RLHF with PPO, which requires training a critic (reward) model. 🔹Then, researchers moved from PPO with a critic to critic-free GRPO because critics were expensive and unstable. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/07/2026
Fun fact: Two some of the most influential statistical NLP papers were authored by Brown et al in 1990 & 2020 1990 Statistical Approach to Machine Translation 2020 Language Models are Few-Shot Learners *) It is not the same Brown **) I believe author names in IBM papers were ordered alphabetically
001
Leo Boytsov @srchvrs.bsky.social · 17/06/2026
🧵There is a fundamental issue with reference-based LLM-judges. People implicitly assume a reference-based judge behaves like: score=f(candidate,reference)score=f(candidate,reference) However, the actual behavior is closer to: score=f(candidate,reference,parametric knowledge,priors) ↩️
arxiv.org
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about their ability to adhere...
120
Leo Boytsov @srchvrs.bsky.social · 17/03/2026
🧵"LLMs are currently "sophisticated autocomplete" tools, not autonomous engineering partners. The 31.7% failure rate and the 13.5x dependency expansion represent a "hidden tax" that can quickly negate any velocity gains." ↩️
211
Leo Boytsov @srchvrs.bsky.social · 09/03/2026
🧵For the last seven years, I kept re-implementing the same pattern: A parallel map loop that divides the work among several processes or threads. My very first attempts were built on Python’s standard tools, e.g., multiprocessing.map... ↩️
121
Reposted by Leo Boytsov
Victor Escorcia @escorciav.bsky.social · 23/02/2026
@srchvrs.bsky.social: "This is too little for pre-training, but pre-training nowadays is probably not a bottleneck. For post-training 16M training samples can meaningfully improve performance on a lot of tasks."
121
Leo Boytsov @srchvrs.bsky.social · 22/10/2025
"Meta is cutting around 600 positions out of the several thousand roles in its Superintelligence Labs, the Facebook owner said on Wednesday as it looks to make its artificial intelligence unit more flexible and responsive." www.reuters.com/business/met...
reuters.com
031
Leo Boytsov @srchvrs.bsky.social · 27/08/2025
🧵I recently finished my nerdiest computer science paper so far and it was accepted by TMLR: A Curious Case of Remarkable Resilience to Gradient Attacks via Fully Convolutional and Differentiable Front End with a Skip Connection. This work was done while I was at Bosch. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 09/08/2025
🧵Hot take: LLMs still fail at basic grammar/style checking. A repeating situation that I encounter: 1. Ask a model about an issue. 2. The model suggests some rewrite for clarity/accuracy. Typically it's actually quite good (but watch for factual errors!). 3. Recheck the text again. ↩️
110
Leo Boytsov @srchvrs.bsky.social · 18/07/2025
🧵Pseudo-relevance feedback (PRF) (also known as blind feedback) is a technique of first retrieving/re-ranking top-k documents and adding some of their words to the initial query. Then, a second retrieval/ranking stage uses an updated query. ↩️
111
Leo Boytsov @srchvrs.bsky.social · 02/07/2025
🧵 Dear (scientific) authors: I am being in the same boat too. However, if you receive a ton of detailed complaints regarding paper quality, do NOT try to address them during the rebuttal phase. It's just a waste of everybody's time. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 27/06/2025
@microsoft.com faces an interesting issue that might affect others selling wrappers around ChatGPT and Claude models: users prefer to use ChatGPT directly rather than engage with Microsoft's Copilot. futurism.com/microsoft-co...
futurism.com
Microsoft Is Having an Incredibly Embarrassing Problem With Its AI
Despite investing tens of billions of dollars into OpenAI, Microsoft is still, apparently, competing with its business partner.
000
Leo Boytsov @srchvrs.bsky.social · 12/06/2025
This is a rather blockbuster piece of news: the @hf.co library is dropping support for both Jax and Tensorflow. www.linkedin.com/posts/lysand...
linkedin.com
I have bittersweet news to share. | Lysandre Debut
I have bittersweet news to share. Yesterday we merged a PR deprecating TensorFlow and Flax support in transformers. Going forward, we're focusing all our efforts on PyTorch to remove a lot of th...
010
Leo Boytsov @srchvrs.bsky.social · 27/04/2025
Found a hidden gem on IR evaluation methodology from Microsoft "What Matters in a Measure? A Perspective from Large-Scale Search Evaluation." dl.acm.org/doi/pdf/10.1...
dl.acm.org
040
Leo Boytsov @srchvrs.bsky.social · 22/04/2025
Parental advice: if you master algebra you will know how to deal with your x-es. @ccanonne.bsky.social feel free to borrow!
000
Leo Boytsov @srchvrs.bsky.social · 16/04/2025
Some people say: A prompt is worth a thousand words! Excuse, but have you seen these ones? They are way longer!
010
Leo Boytsov @srchvrs.bsky.social · 07/04/2025
🧵A fascinating perspective on the nature of intelligence and the history of automation/ (and ahem development of AI). It is also a cautionary story of how to not trust AI too much. ↩️
141
Leo Boytsov @srchvrs.bsky.social · 23/03/2025
🧵Although pre-trained Transformer models took NLP by storm, they were less successful for recommender systems (arxiv.org/abs/2306.11114). RecSys is hard: 1. The number of users is high. 2. The number of items is high. 3. A cold-start problem is a hard one. ↩️
120
Leo Boytsov @srchvrs.bsky.social · 04/02/2025
🧵"Through extensive experiments, we empirically confirm the bias of judges towards their related student models caused by preference leakage across multiple LLM baselines and benchmarks. Further analysis suggests ..." ↩️
120
Leo Boytsov @srchvrs.bsky.social · 17/01/2025
🧵It was my great pleasure to appear on the vector podcast with @dmitrykan.bsky.social . We covered the history of our NMSLIB library and how it helped shape the vector search industry! ↩️
171
Leo Boytsov @srchvrs.bsky.social · 05/01/2025
🧵"Fun" fact about pseudorand numbers (true for np.random, likely for everything). If u ||-ze your Monte-Carlo simulation across processes with standard map-style approach, each worker will be generating IDENTICAL (ohh bother) sequence of rand. numbers and you may not even notice that 😢.↩
230
Reposted by Leo Boytsov
Petr Baudis (pasky) @bluepasky.bsky.social · 05/01/2025
I reposted the thread here! :) bsky.app/profile/xpas...
011
Leo Boytsov @srchvrs.bsky.social · 04/01/2025
Quick primer for non-wizards about the post-MCTS LLM reasoning future (I'm kinda PRIME-pilled rn) by @xpasky.bsky.social x.com/xpasky/statu...
110
Leo Boytsov @srchvrs.bsky.social · 26/12/2024
Great analysys of the recent OpenAi breakthrough on ARC-AGI challenge.
020
Leo Boytsov @srchvrs.bsky.social · 24/12/2024
Another paper (by @laura-dietz.bsky.social and @claclarke.bsky.social ) shows LLM-judges are vulnerable to adv. attacks: "We demonstrate the ease w. which automatic evaluation metrics can be subverted, showing that systems designed to exploit these evaluations can achieve artificially high scores."↩️
261
Leo Boytsov @srchvrs.bsky.social · 23/12/2024
"I sensed anxiety & frustration" by @kyunghyuncho.bsky.social addresses frustration of many PhD students expecting to enjoy the same level of benefits (huge compensation, freedom to publish w/o contributing to product) as a few people hired from few deep learning labs a decade (or so ago)." ↩️
240
Leo Boytsov @srchvrs.bsky.social · 20/12/2024
There are three things one can watch forever: fire burning, water flowing, and A/B test metrics updating.
050
Leo Boytsov @srchvrs.bsky.social · 17/12/2024
🧵We are proposing two new approaches to speed up constrained decoding with lookaheads (unrolling generation to the point in the future where constrains can be verified). ↩️
120
Leo Boytsov @srchvrs.bsky.social · 17/12/2024
🧵A recent study finds that PPO would be better than DPO if PPO were producing of comparable lengths. In other words, judges and humans prefer longer outputs so PPO appears to underperforming. ↩️
100
Leo Boytsov @srchvrs.bsky.social · 17/12/2024
When you wrote circa 200 bash scripts (not counting Python code) during the period of three months and then had trouble recalling differences/nuances. 😅
040
Reposted by Leo Boytsov
Yoav Goldberg @yoavgo.bsky.social · 14/12/2024
at the end of the talk, the audience was in silence. only one brave student with a thick India accent went up to the mic and asked "excuse me Sir, do you consider me a threat?" to which the speaker answered "Not you *specifically*, but in general, yes. Although India is not a big threat as China is"
191
Leo Boytsov @srchvrs.bsky.social · 11/12/2024
🧵"The results demonstrate that LLMs are highly influenced by the presence of query words in the passages under assessment, even if the wider passage has no relevance to the query. This tendency of LLMs to be fooled by the mere presence of query words .... " ↩️ www.microsoft.com/en-us/resear...
microsoft.com
260
Leo Boytsov @srchvrs.bsky.social · 07/12/2024
🧵What I eventually found: A widely used IFEval dataset. "Recent model evaluations report ∼70-80% accuracy on IFEval on average, showing headroom for further analysis and progress on challenging instruction categories." arxiv.org/abs/2409.105... ↩️
100
Leo Boytsov @srchvrs.bsky.social · 06/12/2024
Dear #LazyBlueSkye could you point me to the studies evaluating instruction following (how well or not) in LLMs? Many thanks!
010
Leo Boytsov @srchvrs.bsky.social · 06/12/2024
"We therefore recommend to consider logistic regression as the first choice for data-scarce applications with tabular data and provide practitioners with best practices for further method selection." arxiv.org/abs/2405.07662
arxiv.org
Squeezing Lemons with Hammers: An Evaluation of AutoML and Tabular Deep Learning for Data-Scarce Classification Applications
Many industry verticals are confronted with small-sized tabular data. In this low-data regime, it is currently unclear whether the best performance can be expected from simple baselines, or more compl...
262
Reposted by Leo Boytsov
Yoshitomo Matsubara @yoshitomo-matsubara.net · 05/12/2024
A must-hire candidate is on academic job market!
041
Leo Boytsov @srchvrs.bsky.social · 05/12/2024
Now, we have two multilingual variants/extensions of the MMLU benchmark: one from Cohere and another from OpenAI. 1. huggingface.co/datasets/ope... 2. www.linkedin.com/posts/sararo...
linkedin.com
Sara Hooker on LinkedIn: Is MMLU Western-centric? 🤔 As part of a massive cross-institutional…
Is MMLU Western-centric? 🤔 As part of a massive cross-institutional collaboration: 🗽Find MMLU is heavily overfit to western culture 🔍 Professional…
120
Leo Boytsov @srchvrs.bsky.social · 04/12/2024
OpenAI (seemingly) hired three top vision researchers and ViT inventors from Google: Lucas Beyer, Alexander Kolesnikov, and Xiaohua Zhai. www.wired.com/story/openai...
wired.com
OpenAI Poaches 3 Top Engineers From DeepMind
The new hires, all experts in computer vision, are the latest AI researchers to jump to a direct competitor in an intensively competitive talent market.
010
Leo Boytsov @srchvrs.bsky.social · 03/12/2024
A fantastic piece of news: Anthropics starts supporting Amazon AWS own Trainium 2 accelerators! The first model is Haiku 3.5! www.anthropic.com/news/trainiu...
050
Leo Boytsov @srchvrs.bsky.social · 03/12/2024
🧵"Amazon has launched Nova, a highly competitive family of foundation models. Nova Pro, Lite and Flash set new standards for the intelligence that can be accessed at the price and speed these models are offered at." ↩️ x.com/ArtificialAn...
110
Leo Boytsov @srchvrs.bsky.social · 02/12/2024
Turns out that you can apply a backtranslation-like technique to improve reasoning in LLMs: x.com/cyjustinchen...
020
Leo Boytsov @srchvrs.bsky.social · 02/12/2024
🧵An interesting evaluation of various factors affecting RAG accuracy. In particular, it attempts to reproduce the "Power of noise" paper where authors found that adding less relevant or even random documents can improve performance of the RAG system. But can they? ↩️
390
Leo Boytsov @srchvrs.bsky.social · 30/11/2024
🧵 I think I finally found a concise way to describe IBM Watson Jeopardy system & make a connection to modern QA: In simple terms and using modern terminology, IBM Watson was very similar to a retrieval-augmented QA system. Unlike modern systems, it was not truly generative, but rather extractive:↩️
350
Reposted by Leo Boytsov
Ramon Astudillo @ramon-astudillo.bsky.social · 19/11/2024
I did a starter pack of people in New York (City) working on ML/AI. Please distribute and feel free to self nominate! go.bsky.app/BoEtagz
428919
Leo Boytsov @srchvrs.bsky.social · 28/11/2024
AlexNet was run on two GPUs (inside servers/desktops built by a superstar sysadmin) in Alex Krizhevsky bedroom! Paper link: arxiv.org/abs/1404.5997 ex-Twitter thread: x.com/colinraffel/...
050
Leo Boytsov @srchvrs.bsky.social · 28/11/2024
Crazy finding: In (some?) neural networks there exist a super-weight whose pruning can destroy model's ability to generate text. One should really wonder if this surveys through reproductions. arxiv.org/abs/2411.07191
070