Sign in

Andrew Drozdov

@mrdrozdov.com
5.4K followers 611 following 185 posts

Search and Agents @ Databricks

PostsRepliesMedia
Andrew Drozdov @mrdrozdov.com · 23/09/2026
So many cool papers this week, and it's only Tuesday.
020
Andrew Drozdov @mrdrozdov.com · 26/02/2025
It was a real pleasure talking about effective IR approaches with Brooke and Denny on the Data Brew podcast. Among other things, I'm excited about embedding finetuning and reranking as modular ways to improve RAG pipelines. Everyone should use these more!
180
Andrew Drozdov @mrdrozdov.com · 26/01/2025
"All you need to build a strong reasoning model is the right data mix." The pipeline that creates the data mix:
1131
Andrew Drozdov @mrdrozdov.com · 22/01/2025
Using 100+ tokens to answer 2 + 3 =
1180
Andrew Drozdov @mrdrozdov.com · 09/12/2024
Slides are up! I presented on "Presentation & Consumption in the context of REML" The full deck is here. There's a lot of gems if you're interested in this space! retrieval-enhanced-ml.github.io/sigir-ap2024...
0156
Andrew Drozdov @mrdrozdov.com · 09/12/2024
Today we'll be presenting the Tutorial on Retrieval-Enhanced Machine Learning (REML). Come by to learn about the emerging design patterns in this space and see how to use retrieval beyond RAG. In collaboration w/ the amazing @841io.bsky.social @teknology.bsky.social Alireza Salemi and Hamed Zamani.
1223
Andrew Drozdov @mrdrozdov.com · 06/12/2024
Seen in NYC
3211
Andrew Drozdov @mrdrozdov.com · 04/12/2024
000
Andrew Drozdov @mrdrozdov.com · 04/12/2024
RAG still has a way to go. (this book doesn’t exist)
150
Andrew Drozdov @mrdrozdov.com · 02/12/2024
Similar comment from a reliable source. x.com/earnmyturns/...
@earmyturns (Fernando Pereira) on twitter: It isn't. Vector encodings of sentences predated transformers. Self-attention predated transformers too. Maybe someone said that as a joke, but the transformer idea came from computational efficiency considerations together with preexisting techniques.
040
Andrew Drozdov @mrdrozdov.com · 02/12/2024
@thomlake.bsky.social I'm willing to be convinced. This would give me a whole new appreciation for the memory net work. Can you show that memory nets could process the whole sequence in parallel w/o a for-loop? IMO this is the key capability that self-attention enables.
100
Andrew Drozdov @mrdrozdov.com · 02/12/2024
It's worth noting the authors of the decomposable attention paper all did very well :) One of them (Jakob Uszkoreit) is also a key co-author on AIAYN > Jakob proposed replacing RNNs with self-attention and started the effort to evaluate this idea
first page of AIAYN with the key attribution highlighted
110
Andrew Drozdov @mrdrozdov.com · 27/11/2024
lol nice
youtube giving a warning after I clicked on a bluesky link
080
Andrew Drozdov @mrdrozdov.com · 26/11/2024
unless you're a hawk, then it's at least 20/5
020
Andrew Drozdov @mrdrozdov.com · 26/11/2024
Fun fact: you used to be able to get a Neurips acceptance with only two references included. Two!
1221
Andrew Drozdov @mrdrozdov.com · 26/11/2024
AI researchers were already investigating scaling laws for training neural networks in 1993. proceedings.neurips.cc/paper/1993/h...
2231
Andrew Drozdov @mrdrozdov.com · 25/11/2024
This is not necessarily a dunk. Blog posts are pretty great: open access, self-hosted (low cost), and often has discussion built-in.
The dark/light bus meme with text “this paper should have been a blog post” at the top.
1160
Andrew Drozdov @mrdrozdov.com · 25/11/2024
researcher: make sure you train on lots of data AI model: I’m sorry, but I’m a large language model, and I don’t have access to the internet or any external information sources. I can only generate responses based on the text that I was trained on, which has a knowledge cutoff of 2021.
031
Andrew Drozdov @mrdrozdov.com · 24/11/2024
Wait. thought this was a last week release that I missed. I guess it's from last year lol?
110
Andrew Drozdov @mrdrozdov.com · 24/11/2024
IMO, query expansion can work but likely needs different approach. For example this fusion-like technique from Li et al., 2024. arxiv.org/abs/2311.09175
100
Andrew Drozdov @mrdrozdov.com · 24/11/2024
Research like this is so important. A lot of decision making when deploying RAG systems is based on vibes-based eval using dated models and datasets. Many techniques simply don't transfer to modern models like you'd hope. Query expansion seems to fall into this category.
2311
Andrew Drozdov @mrdrozdov.com · 22/11/2024
ChatGPT: delve Dwarves:
0151
Andrew Drozdov @mrdrozdov.com · 22/11/2024
dogolutely incredible. Thank you!! 🐶🕊️
010
Andrew Drozdov @mrdrozdov.com · 22/11/2024
what have i just seen
chatgpt is responding, and it has the bluesky logo as its loading icon
030
Andrew Drozdov @mrdrozdov.com · 21/11/2024
approaching escape velocity
2563 followers on twitter2590 followers on bluesky
020
Andrew Drozdov @mrdrozdov.com · 21/11/2024
nominating @mcarbin.bsky.social also future bostonian @lateinteraction.bsky.social almost said bostonite, but that's a type of rock!
000
Andrew Drozdov @mrdrozdov.com · 20/11/2024
Is scaling reranker inference cooked? Don't worry, there's hope! We find listwise reranking w/ LLMs has potential as a robust + accurate alternative. This method can be used to create synthetic data for reranker training, although we already see good results w/ zero-shot LLMs.
Results showing that listwise reranking outperforms cross-encoders on average measured by recall@10.
130
Andrew Drozdov @mrdrozdov.com · 20/11/2024
When do rerankers fail 🤔? We looked at the data and found multiple cases where rerankers fail in unexpected ways. Here's an example where a reranker prefers an irrelevant document over the gold despite the limited textual overlap between the query and preferred document.
There is a query, the ground truth relevant document, and the document found by the reranker. The ground truth has extensive text overlap with the query and is easily found by the retriever. Meanwhile, the reranker prefers an irrelevant document with minimal or no text overlap.
130
Andrew Drozdov @mrdrozdov.com · 20/11/2024
Why would this happen? We use rerankers alone in a 1-stage pipeline putting them on even footing against embeddings. Often rerankers are worse than dense embeddings and BM25 🤯! This challenges the common wisdom that today's rerankers outperform embeddings.
We measure Recall@K for different K using 1-stage ranking with retrievers and rerankers. In this fair competition, we see rerankers are surprisingly worse.
140
Andrew Drozdov @mrdrozdov.com · 20/11/2024
Mat is not on 🦋—posting on his behalf! It's time to revisit common assumptions in IR! Embeddings have improved drastically, but mainstream IR evals have stagnated since MSMARCO + BEIR. We ask: on private or tricky IR tasks, are rerankers better? Surely, reranking many docs is best?
A plot showing that reranking improves recall as we increase the number of reranked docs, but with increasing docs we diminishing returns and eventually a performance dip.
48224
Andrew Drozdov @mrdrozdov.com · 19/11/2024
Leonardo meme from Inception. Block text: "When the LLM generates two <THOUGHT> tokens in a row."
090
Andrew Drozdov @mrdrozdov.com · 19/11/2024
Exciting new work from AI2! OpenScholar + ScholarBench. Blog post: allenai.org/blog/opensch... Paper: openscholar.allen.ai/paper Demo: openscholar.allen.ai
0153
Andrew Drozdov @mrdrozdov.com · 18/11/2024
listwise reranking with LLMs
060
Andrew Drozdov @mrdrozdov.com · 16/11/2024
tech cos releasing holiday versions of their deep neural nets
Louis Vuitton using multiple layers of massive faux trucks to cover up construction of their Fifth Ave store in NYC.
010
Andrew Drozdov @mrdrozdov.com · 16/11/2024
guys
010
Andrew Drozdov @mrdrozdov.com · 14/11/2024
000
Andrew Drozdov @mrdrozdov.com · 08/11/2024
0130
Andrew Drozdov @mrdrozdov.com · 28/10/2024
nice one
000
Andrew Drozdov @mrdrozdov.com · 28/10/2024
john backflip keeps me up at night
000
Andrew Drozdov @mrdrozdov.com · 27/07/2023
I knew about Rosencrantz and Guildenstern before watching Station Eleven.
000
Andrew Drozdov @mrdrozdov.com · 27/07/2023
000