Sign in

Omar Khattab

@lateinteraction.bsky.social
1.3K followers 239 following 53 posts

Incoming asst professor at MIT EECS, Fall 2025. Research scientist at Databricks. CS PhD @StanfordNLP.bsky.social. Author of ColBERT.ai & DSPy.ai.

PostsRepliesMedia
Reposted by Omar Khattab
Drew Breunig @dbreunig.bsky.social · 29/05/2026
DSPy requires more upfront learning than just writing natural language prompts. But once you get it, it makes building, maintaining, & improving AI programs so much easier. We aim to soften the learning curve to make the benefits more accessible, starting with more new docs and front page. dspy.ai
dspy.ai
DSPy
The framework for programming—rather than prompting—language models.
092
Reposted by Omar Khattab
Dane Carnegie Malenfant @dvnxmvlhdf5.bsky.social · 28/05/2026
🚨Excited to announce our workshop Context Beyond the Window hosted at COLM in SF! 🚨 LLMs have finite context windows, yet real-world tasks demand absorbing, retaining, and acting on information that far exceeds any single prompt. 1/5 context-beyond-window.github.io
Modern language models operate within finite context windows, yet many real-world tasks require models to absorb, retain, and act on information that far exceeds any single prompt.

This workshop addresses the full spectrum of context management: fitting more into the window, maintaining state across interactions, and transferring knowledge into parameters. We frame this around the trade-off between context-time memory (information supplied at inference) and weight-time memory (information absorbed into parameters).

Our goal is to build a shared vocabulary across subcommunities that rarely meet in one venue: long-context modeling, retrieval-augmented systems, continual learning, knowledge distillation, and LLM agents.
183
Reposted by Omar Khattab
Simon Willison @simonwillison.net · 05/10/2025
If you've been trying to figure out DSPy - the automatic prompt optimization system - this talk by @dbreunig.bsky.social is the clearest explanation I've seen yet, with a very useful real-world case study www.youtube.com/watch?v=I9Zt... My notes here: simonwillison.net/2025/Oct/4/d...
youtube.com
Let the LLM Write the Prompts: An Intro to DSPy in Compound AI Pipelines
YouTube video by Databricks
79914
Reposted by Omar Khattab
PyData Boston @pydatabos.bsky.social · 16/10/2025
#pydatabos interesting! How the Arbor library works under the hood hand in hand with DSPy
031
Omar Khattab @lateinteraction.bsky.social · 29/10/2025
premature optimization is the sqrt of all evil
030
Reposted by Omar Khattab
PyData Boston @pydatabos.bsky.social · 16/10/2025
#pydatabos one line motivation for using DSPy!
031
Reposted by Omar Khattab
joelniklaus.bsky.social @joelniklaus.bsky.social · 21/10/2025
Stop what you are doing and try out GEPA now! "GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning" presents such elegant ideas by a collection of amazing researchers! Here is a tldr of how it works:
143
Omar Khattab @lateinteraction.bsky.social · 28/09/2025
Btw there’s no trouble in storage at all either. ColBERT vectors are often 10 bytes each. Ten bytes. That’s like 4 numbers. It’s not “many vectors work better than one vector”. It’s “set similarity works better than dot product”. Even with the same storage cost.
130
Reposted by Omar Khattab
mr. TIM @timkellogg.me · 19/09/2025
colbert-muvera-micro a 4M(!!) late interaction model late interaction models do embedding vector index queries and reranking at the same time leading to far higher accuracy huggingface.co/NeuML/colber...
A diagram illustrating a dual-encoder retrieval model using MaxSim scoring.
	•	On the left (green box): labeled “Query Encoder, f_Q”. It takes a Query as input and produces multiple vector embeddings (rectangles).
	•	On the right (blue box): labeled “Document Encoder, f_D”. It takes a Document as input and produces multiple vector embeddings (rectangles). This block is marked with “Offline Indexing” along the side, showing that documents are pre-encoded.
	•	Between the two encoders: dotted and solid arrows connect query embeddings to document embeddings, representing similarity comparisons.
	•	Each comparison goes through a “MaxSim” operation (highlighted boxes), which selects the maximum similarity for each query token across document tokens.
	•	At the top: outputs of MaxSim flow into a summation node (Σ) to produce a single score for ranking.

This shows the ColBERT (Contextualized Late Interaction) retrieval framework: query and document are encoded separately, interactions are computed via maximum similarity per query token, and results are aggregated into a score.
2151
Reposted by Omar Khattab
Latitude77 @latitude77.bsky.social · 19/06/2025
Let the Model Write the Prompt | Drew Breunig #dspy #promptengineering #llms #generativeai
021
Reposted by Omar Khattab
Drew Breunig @dbreunig.bsky.social · 15/06/2025
Here's the write up of my Data+AI Summit talk on the perils of prompts in code and how to mitigate them with DSPy. www.dbreunig.com/2025/06/10/l...
dbreunig.com
Let the Model Write the Prompt
Notes from a talk I delivered at the 2025 Data + AI Summit, detailing the problem with prompts in your code and how DSPy can make everything better.
141
Reposted by Omar Khattab
MLflow @mlflow.org · 30/05/2025
Have you heard the news? #MLflow now supports tracking for DSPy optimization workflows—just like it does for #PyTorch training! Keep reading to see what this means for your #LLM projects… 👇 #opensource #dspy #oss
183
Reposted by Omar Khattab
MLflow @mlflow.org · 23/04/2025
📣 TODAY at 4PM PT - MLflow Community Meetup! 🔗 Register today 👉 lu.ma/mlflow423 Join the global MLflow community for two exciting tech deep dives: 🔹 MLflow + #DSPy Integration 🔹 Cleanlab + #MLflow 🎥 Streaming live on YouTube, LinkedIn, and X 💬 Live Q&A with the presenters #opensource #oss
lu.ma
MLflow Community Meetup | April 23 · Luma
Join us for the next MLflow Community Meetup — Wednesday, April 23 at 4PM PT! We’re bringing two exciting presentations to the community: 🔹 MLflow + DSPy…
151
Reposted by Omar Khattab
MLflow @mlflow.org · 21/04/2025
MLflow now supports tracking for #DSPy (Community) optimization — just like it does for @pytorch.org training! 🙌 #MLflow is the first to bring full visibility into DSPy’s prompt optimization process. More observability, less guesswork. Get started today! ➡️ medium.com/@AI-on-Datab... #opensource
153
Reposted by Omar Khattab
MLflow @mlflow.org · 15/04/2025
Join us for the next MLflow Community Meetup — Wednesday, April 23 at 4PM PT! 🗓️ 🔹 Explore the new MLflow + #DSPy integration 🔹 Learn how Cleanlab adds trust to AI workflows with MLflow 💬 Live Q&A + demos 📺 Streamed on YouTube, LinkedIn, and X 👉 RSVP: lu.ma/mlflow423 #opensource #mlflow #oss
lu.ma
MLflow Monthly Meetup · Luma
Join us for the next MLflow Community Meetup — Wednesday, April 23 at 4PM PT! We’re bringing two exciting presentations to the community: 🔹 MLflow + DSPy…
042
Omar Khattab @lateinteraction.bsky.social · 06/04/2025
Nice work! For history: dspy.ai/api/primitiv...
dspy.ai
History - DSPy
The framework for programming—rather than prompting—language models.
110
Omar Khattab @lateinteraction.bsky.social · 04/03/2025
This was built by a long-time DSPy community member!
240
Omar Khattab @lateinteraction.bsky.social · 03/03/2025
Yes there's an evals crisis, but evaluating *models* is not even the right question most of the time LangProBe from Shangyin Tan, @lakshyaaagrawal.bsky.social, Arnav Singhvi, Liheng Lai, @michaelryan207.bsky.social et al begins to ask what complete *AI systems* we should build & under what settings
0102
Reposted by Omar Khattab
Lakshya A Agrawal @lakshyaaagrawal.bsky.social · 03/03/2025
🧵Introducing LangProBe: the first benchmark testing where and how composing LLMs into language programs affects cost-quality tradeoffs! We find that, on avg across diverse tasks, smaller models within optimized programs beat calls to larger models at a fraction of the cost.
163
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
It doesn't help that the we in ML often only design abstractions leak all kinds of implementation details. Folks often define ML itself in terms of techniques, not problems! But it's prematurely abstracting that leads to the bitterness of wasted effort, and not "modularity doesn't work for AI". 2/2
041
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
Composition & abstraction are the foundations of CS, but are clearly absent in modern ML. It's not that they're not crucial for intelligent software. But it takes building many half-working systems to abstract successfully, and it takes good abstractions to have primitives worth composing. 🧵1/2
181
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
4) By default, IR methods that use "multiple vectors" (e.g., cross-encoders) are unscalable. It seems like a necessary tradeoff, but the fascinating thing in late interaction is that it's easy to implement in asymptotically sub-linear ways, thanks to pruning. Hope this was useful!
110
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
3) "Multi-vector" makes it sound like these approaches win because they store "more stuff". But that's not true: if you look at how aggressive ColBERTv2 representations are compressed, it's often ~20 bytes per vector (like 5 floats), which can be smaller than popular uncompressed single vectors!
120
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
For dot products, every time you "fix" one query--document pair, you likely break so many other pairs by moving the query and/or document representations. For ColBERT, you typically *fix* more than you break because you're moving *tokens* in a much smaller (and far more composable!) space.
110
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
The problem isn't the vector representation, it's the **learnability of the scoring function**. A dot product is just very hard to learn. An intuition I learned from Menon et al (2021) is that:
111
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
2) More importantly, there's nothing to say you can't store a TON of information in a single vector. And it's easy to use multiple vectors and gain *zero* improvement over a single-vector, e.g. if you replace MaxSim with AvgSim in ColBERT, without any other changes, it doesn't work!
110
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
1) If you take ColBERT and force it to use only a constant number of vectors (e.g., 16), it'll barely outperform one vector in the general case. It's not that you need token-level alignment per se (you don't either!) but you want fine-grained representations, not just *multiple* representations.
110
Omar Khattab @lateinteraction.bsky.social · 26/02/2025
Some quick thoughts: On why we gave the ColBERT paradigm the name "late interaction" instead of "multi-vector", a term that emerged later and that has proven to be more intuitive. **The mechanism is actually not about having multiple vectors at all.** You can see this in four different ways. 🧵1/7
170
Omar Khattab @lateinteraction.bsky.social · 20/02/2025
Btw the full general form to export all message templates is: ``` {name: my_adapter.format(p.signature, demos=p.demos, inputs={k: f'{{{k}}}' for k in p.signature.input_fields}) for name, p in your_program.named_predictors()} ```
010
Omar Khattab @lateinteraction.bsky.social · 20/02/2025
The default Adapter is dspy.ChatAdapter(). But you can do all customization you mentioned with a custom Adapter: class MyAdapter(dspy.Adapter): def format(self, signature, demos, inputs): return {"role": "user", "content": ...} def parse(self, signature, completion): return {....}
110
Omar Khattab @lateinteraction.bsky.social · 20/02/2025
Thanks so much, Eric! The piece you're looking for is the Adapters, but we should make give it nice syntactic sugar, so it feels native. ``` all_messages = {name: my_adapter.format(predictor.signature, demos=predictor.demos, inputs=dict(...)) for name, predictor in program.named_predictors()} ```
110
Omar Khattab @lateinteraction.bsky.social · 20/02/2025
Tell me more?
110
Omar Khattab @lateinteraction.bsky.social · 01/02/2025
Someone needs to write the book “Modern Machine Learning — the math, the myth, the legend”
1120
Omar Khattab @lateinteraction.bsky.social · 28/01/2025
literally an average joke
060
Omar Khattab @lateinteraction.bsky.social · 28/01/2025
a statistician walks into an error bar surprisingly, not everyone was mean to him
272
Omar Khattab @lateinteraction.bsky.social · 05/01/2025
What do you call LLMs that exhibit reward hacking for translation? Agents gone ROUGE.
3230
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
Pytorch on cpu can be bad unfortunately; most of this gain on cpu (see ablations) comes from writing C++
010
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
Yup, but they’re related via a constant factor (eg at most 500)
100
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
More or less, yes! (With the small caveat that I'm not saying one could conceptually do this; I'm saying that, by design, this has been the case since early 2020 and it has been deployed in production by what must be hundreds of teams at this point.)
010
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
Observe that it's impossible for the ColBERT score to be large *unless* one or more of the dot products is somewhat large. In other words, if every MaxScore is below epsilon, the sum of these MaxScores is also small. This allows ColBERT representations to aggressively prune their own scoring fn.
100
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
Yes makes sense! All I'm saying is that ColBERT challenges that design and says "let me handles levels 0 through 1 or 2 all at once" using the same embeddings but a multi-stage algorithm. This allows you to train a single model, maintain a single index, and optimize quality/cost in one go.
120
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
They're not conceptually separate because it's a lossless approximation in practice, either empirically --- or theoretically with some algorithm tweaks. Basically, ColBERT searching this small set of docs matches full scoring. e.g., this paper here finds that iirc: arxiv.org/pdf/2404.14989
arxiv.org
100
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
FAISS is a really neat library of algorithms, e.g. IVFPQ. I wouldn't think of it as conceptually separate, it's just a way to go from query embeddings -> (possibly sub-linear, depending entirely on the choices of algorithm & hyperparameters) number of candidate document embeddings -> scoring.
200
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
Yes indeed! But the ColBERT papers, e.g. PLAID, have in fact scaled to 100M passages with one GPU. With a GPU, you can find the top-10 matches within a collection of 2.5M passages in 10 milliseconds and you can search 140M passages (56x more docs) in 50 milliseconds, i.e. only 5x bigger latency.
100
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
Some quotes from page 8 of the PLAID paper.
110
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
Yes, ColBERTv2 via PLAID. (Even ColBERTv1, properly configured, had that property.) The square root is the number of cluster centroids used for pruning almost all documents before scoring. (It's described in the ColBERTv2 paper. See the PLAID paper, though, for the optimized process.)
210
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
>"the recommendation i’ve heard is to use a dumb filter to get down to thousands of docs, and then apply ColBERT" It's always funny to realize just how vibes-based people tend to be. We all make up tradeoffs where none exist ("ColBERT uses more vectors, so it must be better but very expensive??")
010
Omar Khattab @lateinteraction.bsky.social · 04/01/2025
Search latency is O(sqrt N) for N documents. You can store 1,000,000 passages in like 3 GB. You can search 100,000,000 passages in 200 milliseconds on a CPU! It's absurdly fast and lightweight, a lot of that thanks to Keshav Santhanam. But it's complicated custom C++/CUDA kernels.
410