Sign in

turbopuffer

@turbopuffer.bsky.social
264 followers 1 following 134 posts

{vector, full-text} search engine built on object storage. fast, cheap, trillion scale. powers Anthropic, Harvey, Notion, Cognition, and more

PostsRepliesMedia
turbopuffer @turbopuffer.bsky.social · 02/10/2026
we implemented batched reads in the v3 full-text engine, leading to a ~9x performance improvement. postings are stored as columns, and we were reading them a row at a time sometimes optimization is not about being clever, but simply avoiding mistakes tpuf.link/v3/2026-10-02
tpuf.link
Batched reads speed up FTS by 9x
Follow along as we evolve turbopuffer's storage architecture and grind latency down til it's better than it ever was
021
turbopuffer @turbopuffer.bsky.social · 01/10/2026
since v1, tpuf has keyed every index on the vector address we've stretched this architecture to its limit. to make search faster in every respect, we must redesign our storage engine so the vector index is no longer primary turbopuffer.com/blog/rip-ve...
turbopuffer.com
RIP, vector database
turbopuffer is pushing the frontier of search. To do that, we have to fundamentally redesign our storage architecture so the vector index is no longer primary.
000
turbopuffer @turbopuffer.bsky.social · 30/09/2026
we're changing tpuf's storage architecture so we can keep pushing the search frontier the new engine (v3) now passes CI, but we have to make it fast before we ship it to prod we're grinding query plans in public until v3 wins on everything follow along: turbopuffer.com/v3
022
turbopuffer @turbopuffer.bsky.social · 29/09/2026
moving embedding into turbopuffer lets us parallelize the embedding call with other work in the query plan whole-system puffin' shaves 100+ ms off your semantic searches turbopuffer.com/blog/native...
turbopuffer.com
Why moving embedding inside turbopuffer drops search latency
Native embedding isn't just a convenience feature, but a way to shave hundreds of milliseconds off your search queries.
011
turbopuffer @turbopuffer.bsky.social · 23/09/2026
ask Kepler for a public company's revenue growth, and it traces the number to an exact line item in a verified SEC filing to find the filing, Kepler searches on tpuf → 100M+ financial documents → 85ms p90 latency → 10x lower cost than alternatives turbopuffer.com/customers/k...
turbopuffer.com
Kepler searches 100M+ financial documents with turbopuffer
Kepler builds verifiable AI for investment firms and financial analysts. With turbopuffer they can search every SEC filing, earnings transcript, and SharePoint import at far lower cost than alternatives.
020
turbopuffer @turbopuffer.bsky.social · 18/09/2026
new: bytes data type before, you stored binary data in tpuf as a string and paid ~33% extra for encoding. now, you pay only for the bytes docs: turbopuffer.com/docs/write#...
010
turbopuffer @turbopuffer.bsky.social · 15/09/2026
august changelog tpuf.link/chlog
010
turbopuffer @turbopuffer.bsky.social · 11/09/2026
new: explore namespaces in the tpuf dashboard → browse all documents in your namespaces → check metadata and indexing status → search! text, semantic, and custom queries rolling out to all public regions over the next few days. more features coming soon
001
turbopuffer @turbopuffer.bsky.social · 09/09/2026
in the last 30 days, the p100 full-text query on tpuf evaluated 1,823 clauses agents write (much) longer queries than humans, so it is increasingly important for text search engines to scale well with the number of terms
110
turbopuffer @turbopuffer.bsky.social · 04/09/2026
we partnered with Applied Compute to post-train a small model for large-scale code search over precomputed indexes at 300 repos, this is ~3x faster than using filesystem + grep, and reduces the marginal cost of a search by up to 100x versus frontier models turbopuffer.com/blog/large-...
turbopuffer.com
Post-training open-weight models for large-scale code search
Applied Compute partnered with turbopuffer to show that a small, specialized search model post-trained to use precomputed search indexes can achieve impressive search quality at a fraction of the cost and latency of frontier models.
020
turbopuffer @turbopuffer.bsky.social · 01/09/2026
fal just made it possible to generate video faster than you can watch it turbopuffer makes it possible to search every video, image, and 3D model you generate on fal → 300M+ assets → 225k+ namespaces → 17ms p50 hybrid latency customer log: [tpuf.link/fal]
tpuf.link
fal scales media search to 300M+ assets on turbopuffer
fal shipped search in Assets so users and their agents can find and reference previously generated media assets. turbopuffer runs filtered hybrid search over 300M+ videos, images, and 3D models at millisecond latency.
010
turbopuffer @turbopuffer.bsky.social · 27/08/2026
new: multi-attribute ordering order query results by up to 8 attributes docs: turbopuffer.com/docs/query#...
020
turbopuffer @turbopuffer.bsky.social · 20/08/2026
turbopuffer's first public Azure region, azure-eastus2, is now in private preview contact us for access
100
turbopuffer @turbopuffer.bsky.social · 14/08/2026
we run 100+ tpuf clusters, including many that live inside customers' clouds (BYOC). how do you operate a cluster you can't touch? we don't use Terraform for this. instead, we build a custom control plane to manage the entire fleet without ever reaching in tpuf.link/control-plane
tpuf.link
How to ship a database every day
We deploy many database upgrades every day across ~100 public, single-tenant, and BYOC clusters. It would be insane to use Terraform or Helm to do that, so we built our own control plane.
020
turbopuffer @turbopuffer.bsky.social · 11/08/2026
july changelog tpuf.link/chlog
100
turbopuffer @turbopuffer.bsky.social · 30/07/2026
Mem0 migrated 400M+ agent memories from pgvector to turbopuffer, solving a semi-selective filtering problem that spiked tail latencies in Postgres → 150k+ isolated search indexes → 70ms p90 hybrid retrieval latency → 97% vector recall@10 tpuf.link/mem0
tpuf.link
Mem0 migrates 400M+ agent memories from pgvector to turbopuffer
Postgres didn't scale for Mem0's agent memory platform. They migrated hundreds of millions of memories to turbopuffer, reducing end-to-end latency by 70x.
010
turbopuffer @turbopuffer.bsky.social · 29/07/2026
tpuf now supports late interaction [beta] use models like ColBERT to represent text as a set of vectors (1 per token) tpuf uses a single-vector ANN index for a fast first pass, then reranks hits using exact late interaction scoring to boost recall docs: turbopuffer.com/docs/query#...
051
turbopuffer @turbopuffer.bsky.social · 22/07/2026
autoscaling is deceptively hard our indexer fleet scales nodes on job queue time. if a queue suddenly went quiet, we'd scale down too hard & new jobs could queue up waiting for nodes to claim them we tweaked the HPA signal to prevent underprovisioning → ~2x shorter queue time
010
turbopuffer @turbopuffer.bsky.social · 20/07/2026
new: computed attributes compute additional values on the results of the rank_by and pass them as extra signals to a reranker → BM25 score for a vector query → vector distance for a BM25 query → individual scores for composite rank_by clauses docs: turbopuffer.com/docs/query#...
030
turbopuffer @turbopuffer.bsky.social · 17/07/2026
new: highlighting extract the text fragments most relevant to a query → highlight matches in your search results UI → minimize context passed to your LLM docs (and playground): turbopuffer.com/docs/fts#hi...
020
turbopuffer @turbopuffer.bsky.social · 15/07/2026
june changelog tpuf.link/chlog
010
turbopuffer @turbopuffer.bsky.social · 14/07/2026
now in beta: native embeddings in tpuf embedding is the most painful part of puffing. we want to make it easy you can now convert chunks to vectors as you read and write to turbopuffer, without extra calls to an embedding model provider API docs: turbopuffer.com/docs/embedding
110
turbopuffer @turbopuffer.bsky.social · 03/07/2026
puffin' in paris for RAISE summit we're bringing together good friends for a night of champagne, caviar, and chicken nuggets july 9. rsvp here: luma.com/raisesummit...
000
turbopuffer @turbopuffer.bsky.social · 30/06/2026
Legora searches 2B+ legal documents on turbopuffer → strict per-matter data isolation with CMEK → 10x lower tail latency than Postgres → 98% avg recall@10 tpuf.link/legora
tpuf.link
Legora searches 2B+ legal docs on turbopuffer
Legora replaced Elasticsearch and pgvector + DiskANN with turbopuffer to build an efficient retrieval layer that minimizes input tokens without sacrificing performance, all with legal-grade data isolation and security.
000
turbopuffer @turbopuffer.bsky.social · 18/06/2026
new: i8 vectors f32: 4 bytes/dim i8: 1 byte/dim 4x fewer bytes → 75% lower storage and query costs + faster queries when embedded with a quantization-aware model (e.g. voyage-4-large) trained on i8 vectors, recall loss can be ~0! docs: turbopuffer.com/docs/perfor...
020
turbopuffer @turbopuffer.bsky.social · 18/06/2026
we cut our base price from $64 → $16/mo start puffing for 4x less
140
turbopuffer @turbopuffer.bsky.social · 17/06/2026
we open sourced alyze, the Rust crate behind tpuf's default full-text search tokenizer (word_v4) our first tokenizers (up to word_v3) were built on Tantivy's analyzer, and we owe them many thanks alyze does a bit less, but does it up to ~4x faster github.com/turbopuffer...
180
turbopuffer @turbopuffer.bsky.social · 11/06/2026
Atlassian's cross-product AI platform, Rovo, searches 5B+ documents on turbopuffer BYOC → 19% increase in search quality → 60ms p90 latency → 96% average recall@10 tpuf.link/atlassian
tpuf.link
Atlassian finds its multi-cloud BYOC search engine in turbopuffer
Atlassian chose turbopuffer for its operational and architectural simplicity, cloud-agnostic BYOC deployment model, low cost, and virtually unlimited scalability. It now underpins search for millions of Atlassian customers on Jira, Confl...
020
turbopuffer @turbopuffer.bsky.social · 09/06/2026
puffy plays pickleball june 30 in SF, co-hosted with good friends luma.com/the-agent-open
000
turbopuffer @turbopuffer.bsky.social · 08/06/2026
a year ago, ~98% of tpuf queries were vector ANN last 30d: 64% vector ANN 19% full-text BM25 13% filter-only 3% aggregate 1% other (sparse vector, exact kNN, ...)
030
turbopuffer @turbopuffer.bsky.social · 05/06/2026
new: rerank_by before, you'd implement rank fusion client-side. now, a little QoL upgrade, especially nice for large result sets docs: turbopuffer.com/docs/query#...
020
turbopuffer @turbopuffer.bsky.social · 03/06/2026
may changelog tpuf.link/chlog
010
turbopuffer @turbopuffer.bsky.social · 02/06/2026
new: branching create an instant, copy-on-write clone of a tpuf namespace → constant-time (440ms p50, ~1s p99) → fully independent → unlimited branches, unlimited branch depth docs: turbopuffer.com/docs/branching
020
turbopuffer @turbopuffer.bsky.social · 01/06/2026
tpuf quantizes vectors to improve perf (RaBitQ) the algo randomly rotates vectors, and we were using matmul at O(d²) space & time, brutal at high dims. 10k = 400MB in RAM! we rebuilt the rotation using FWHT at O(d) space & O(d log d) time. ~no recall loss, 10k = only 5kB in RAM
000
turbopuffer @turbopuffer.bsky.social · 22/05/2026
new in turbopufer: the Fuzzy filter typo-tolerant substring matching with a configurable edit distance, so you can puff (or puf) even when you spell it wrong docs: tpuf.link/fuzzy
020
turbopuffer @turbopuffer.bsky.social · 20/05/2026
SID-1 is an agentic search model → 1.9x recall over RAG + rerank → 24x faster, 99% cheaper than GPT-5.1 trained using large-scale RL on turbopuffer at 1k+ QPS bursts over 10M+ document corpora across thousands of steps tpuf.link/sid-1
tpuf.link
Training SID-1 to beat GPT-5 at search with 1k+ QPS RL
SID-1 is an agentic search model that is 24x faster than GPT-5.1-high, 374x cheaper than Sonnet 4.5, and achieves 1.9x higher recall than traditional RAG pipelines. Here's how we trained it using large-scale RL on turbopuffer.
030
turbopuffer @turbopuffer.bsky.social · 13/05/2026
filtered counts on tpuf just got much faster our ANN index may replicate docs across clusters for better recall, so we had to dedupe matching IDs (slow) now, we store a bitmap of replica positions so the query plan is pure bitmap ops: (filter_bitmap - replica_bitmap).popcnt()
040
turbopuffer @turbopuffer.bsky.social · 12/05/2026
we're bringing more of the database to the tpuf dashboard step one: namespace metadata more soon
020
turbopuffer @turbopuffer.bsky.social · 08/05/2026
new: sparse vectors a first-class retrieval primitive that composes with BM25 + attribute ranking in the same query plan, no client-side fusion needed for SPLADE / learned-sparse retrievers (or roll your own weights for custom feature scoring) docs: turbopuffer.com/docs/query#...
000
turbopuffer @turbopuffer.bsky.social · 07/05/2026
new: namespace pinning pin namespaces to reserved compute for faster, cheaper, more predictable p99 on high QPS workloads ~50x cheaper than shared compute for a 128GB namespace at 500 QPS docs: turbopuffer.com/docs/pinning
100
turbopuffer @turbopuffer.bsky.social · 06/05/2026
april changelog tpuf.link/chlog
010
turbopuffer @turbopuffer.bsky.social · 04/05/2026
new tpuf regions, so you can puff a little closer to home 🇧🇷 aws sa-east-1 (São Paulo) 🇺🇸 gcp us-east1 (South Carolina) 🇧🇪 gcp europe-west1 (Belgium)
010
turbopuffer @turbopuffer.bsky.social · 28/04/2026
✨ docs search for search docs ✨ now live on turbopuffer.com/docs
010
turbopuffer @turbopuffer.bsky.social · 27/04/2026
BM25 efficiently scores text, but relevance often depends on more than text (recency, popularity, PageRank) we score non-text attributes as clauses in the same MAXSCORE plan that evaluates BM25 → better first-stage relevance, still scales to 100M+ tpuf.link/rank-by-attr
tpuf.link
Mixing non-text attributes into text search for better first-stage relevance
turbopuffer now allows you to combine attribute values into the scoring function of text queries. Ranking by attribute helps achieve better relevance in the first-stage with the same scalability characteristics as BM25.
010
turbopuffer @turbopuffer.bsky.social · 21/04/2026
stemming is what makes a text search for "run" match documents containing "running" or "runs" we just shipped a small stem cache so repeated terms skip the stemmer → ~2x tokenization throughput when stemming is enabled
010
turbopuffer @turbopuffer.bsky.social · 15/04/2026
march changelog tpuf.link/chlog
020
turbopuffer @turbopuffer.bsky.social · 13/04/2026
puffy was right at home at our AI night at the aquarium in London coming to {a city near you}
a 3D lego model of turbopuffer’s mascot, puffy, is held up against an aquarium backdrop. puffy is right at home amongst the aquatic life. A group of people enjoy an aquarium scene, admiring fish swimming in vibrant blue water under a dome structure.An event banner highlights “A night at the aquarium,” while guests explore an underwater tunnel with vibrant lighting.
010
turbopuffer @turbopuffer.bsky.social · 09/04/2026
each object store has its tradeoffs GCS has great throughput but limits per-object replaces to 1/s. S3 has no rate limit, but lower throughput we coalesce writes into larger WAL commits to respect GCS, but we've now increased commit cadence on S3 for ~2.5x lower write latency
011
turbopuffer @turbopuffer.bsky.social · 08/04/2026
ElevenHacks #4 → turbopuffer x ElevenLabs 🥇 $8,192 first prize 🥈 $4,096 second prize 🥉 $1,024 third prize challenge drops tomorrow at 9a pt / 12p et sign up below ($128 in tpuf credits to get you started) tpuf.link/11hacks
turbopuffer.com
turbopuffer x ElevenHacks
ElevenHacks #4 with turbopuffer and ElevenLabs — signup credits, hackathon prizes, and resources for your build.
000
turbopuffer @turbopuffer.bsky.social · 02/04/2026
new: multiple vector columns store multiple embeddings for the same document - each with its own dimensions, types, and ANN index multimedia → multiple vectors docs: tpuf.link/multi-vec-cols
Code snippet for configuring a multivector database entry, including title, image URI, and embedding details for text and images.
010