Sign in

Nikhil Benesch

@benesch.bsky.social
1.3K followers 606 following 218 posts

Systems engineer @turbopuffer.bsky.social. Former CTO @materialize.com.

PostsRepliesMedia
Nikhil Benesch @benesch.bsky.social · 22/07/2025
They’re calling it our most boring feature yet. boring(n): mundane; works exactly as expected; the highest praise you can give a database
050
Nikhil Benesch @benesch.bsky.social · 10/06/2025
tpuf Python client went async today! 🐡 ❤️ 🐍 Async client perf slightly edges out sync client perf under heavy query load. github.com/turbopuffer/...
060
Nikhil Benesch @benesch.bsky.social · 14/05/2025
Big day today at turbopuffer HQ.
091
Reposted by Nikhil Benesch
turbopuffer @turbopuffer.bsky.social · 07/03/2025
when we said "coming soon" we really meant it now puffin' in aws-us-east-1 and aws-eu-central-1
041
Nikhil Benesch @benesch.bsky.social · 07/03/2025
I did! Couldn't pass up the opportunity to work with this team!
010
Nikhil Benesch @benesch.bsky.social · 07/03/2025
Things move fast at turbopuffer. Now puffin' in aws/us-east-1 and aws/eu-central-1 too.
130
Nikhil Benesch @benesch.bsky.social · 06/03/2025
I've been dreaming about conditional writes on S3 for years. I couldn't have asked for a better way to celebrate than getting to ship tpuf on AWS. 🐡
0131
Nikhil Benesch @benesch.bsky.social · 27/02/2025
C4A: cloud.google.com/blog/product...
cloud.google.com
Try C4A, the first Google Axion Processor | Google Cloud Blog
The custom Arm-based processor is designed for general-purpose workloads like web and app servers, databases, analytics, CPU-based AI, and more.
110
Nikhil Benesch @benesch.bsky.social · 26/02/2025
It’s a rare day that my love of going to battle with build systems pays off like this. Kudos to GCP for a very impressive new SKU. 🐡💨
0120
Nikhil Benesch @benesch.bsky.social · 06/02/2025
Couldn't be more excited to be joining the team at @turbopuffer.bsky.social. 🐡💨
0170
Nikhil Benesch @benesch.bsky.social · 30/12/2024
Just catching up on my NULL BITMAPS and this is easily the best intuition for write skew I've ever seen described. The analogy to merge skew in a codebase is genius.
090
Nikhil Benesch @benesch.bsky.social · 19/12/2024
Well well well: www.crunchydata.com/blog/pg_incr... Incremental pipelines come to Postgres via Crunchy Data! This is like "dbt incremental", not true incremental view maintenance like @materialize.com or Snowflake's dynamic tables, but it's a neat step towards IVM.
crunchydata.com
pg_incremental: Incremental Data Processing in Postgres | Crunchy Data Blog
We are excited to release a new open source extension called pg_incremental. pg_incremental works with pg_cron to do incremental batch processing for data aggregations, data transformations, or import...
0213
Nikhil Benesch @benesch.bsky.social · 19/12/2024
Of course! Will be very excited to give the new DynamoDB-free SlateDB a spin.
010
Reposted by Nikhil Benesch
Phil Eaton @eatonphil.bsky.social · 12/12/2024
The 6th (and last of 2024!) NYC Systems talks are next Thursday! We've got @jaronoff.com of Omlet and @benesch.bsky.social of Materialize speaking. :) nycsystems.xyz/december-202...
0202
Nikhil Benesch @benesch.bsky.social · 11/12/2024
Same! I’ve been using Amethyst for over ten years at this point. When I discovered the project in 2013 I never expected that ianyh would still be lovingly maintaining it a decade later. One of my most loved pieces of software, for sure.
110
Nikhil Benesch @benesch.bsky.social · 11/12/2024
tl;dr the 10x TPS claim actually *is* based on a small but novel optimization inside of S3! Table buckets understand Iceberg naming conventions and adjust S3's rate limiting so that you start with budget for 55k/35k GETs/PUTs per sec rather than the general purpose default of 5.5k/3.5k.
060
Nikhil Benesch @benesch.bsky.social · 11/12/2024
It just occurred to me that while the 3x query performance claim is based on a comparison to uncomplicated tables, I hadn't actually seen the evidence for the 10x TPS claim. After some digging I just found the explanation in the re:Invent talk (start at 14:24): youtu.be/1U7yX4HTLCI?...
youtu.be
AWS re:Invent 2024 - [NEW LAUNCH] Store tabular data at scale with Amazon S3 Tables (STG367-NEW)
YouTube video by AWS Events
120
Nikhil Benesch @benesch.bsky.social · 11/12/2024
Exciting! I'm so here for it. 👀
000
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Disclaimer: I'm not involved with DuckDB, ClickHouse, or Iceberg, so take this all with a grain of salt.
110
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Rust might be feasible! In fact I think DuckDB's Delta extension is based on a Rust library. And ClickHouse is starting to integrate some Rust too: clickhouse.com/blog/more-th... Sadly Rust's Iceberg library is still relatively immature (e.g. no support for writes: github.com/apache/icebe...).
clickhouse.com
More Than 2x Faster Hashing in ClickHouse Using Rust
Rust’s rich type system and ownership model guarantee memory-safety and thread-safety. There are a fairly large number of useful libraries written on it, so we considered using them in ClickHouse.
140
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Calling Go code from another language is hard enough that' it's rarely done—especially in high performance contexts. (🌶️ alert, but fasterthanli.me/articles/lie... covers this—see "the only good boundary with Go is a network boundary.")
fasterthanli.me
Lies we tell ourselves to keep using Golang
In the two years since I've posted I want off Mr Golang's Wild Ride, it's made the rounds time and time again, on Reddit, on Lobste.rs, on HackerNews, and elsewhere. And every time, it elicits the ...
140
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Nice, looking forward to it!
000
Nikhil Benesch @benesch.bsky.social · 10/12/2024
I’m not 100% sure on the chronology but I believe DuckDB originated many of these pragmatic SQL UX improvements: duckdb.org/2022/05/04/f...
duckdb.org
Friendlier SQL with DuckDB
DuckDB offers several extensions to the SQL syntax. For a full list of these features, see the Friendly SQL documentation page.
020
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Keep an eye on github.com/apache/icebe...! AFAIU the lack of a C++ Iceberg library is what’s been holding back full Iceberg support in DuckDB and ClickHouse.
github.com
GitHub - apache/iceberg-cpp: Apache Iceberg C++
Apache Iceberg C++. Contribute to apache/iceberg-cpp development by creating an account on GitHub.
131
Nikhil Benesch @benesch.bsky.social · 10/12/2024
@chris.blue you might be interested in this. Something like this fits in very neatly to your views on the commoditization of the PostgreSQL dialect/protocol (materializedview.io/p/databases-...).
materializedview.io
Databases Are Commodities. Now What?
How will vendors differentiate when databases are commoditized? I've got three ideas.
140
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Indeed. That's one of the more cursed stories I've heard. Might be worth a uv bug report, if the constraints are indeed solvable!
000
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Have you tried `uv` yet? docs.astral.sh/uv/ It's night and day compared to pip.
docs.astral.sh
uv
110
Nikhil Benesch @benesch.bsky.social · 10/12/2024
👋🏽 Well hello, Spiral.
010
Nikhil Benesch @benesch.bsky.social · 10/12/2024
(for posterity for anyone else following along: it seems to be sufficient to define the external consistency guarantee in terms of T1's commit timestamp and T2's read timestamp because SI already implies that a txn can't commit before it starts: bsky.app/profile/bene...)
121
Nikhil Benesch @benesch.bsky.social · 10/12/2024
This is so cool to see! Out of curiosity, which sort of features were hard to test via SQL and/or PL/pgSQL? I would have naively assumed that all observable behavior of PostgreSQL was observable via the SQL interface. 😅
110
Reposted by Nikhil Benesch
DBA @pgdba.bsky.social · 10/12/2024
🚀 Postgres Compatibility Index (PCI): Think your shiny new Postgres derivative is "Postgres-compatible"? Test it, score it, and know the truth. 🐘 Thanks @gunnarmorling.dev @tudor.golubenco.org @benesch.bsky.social @deverts.bsky.social for the inspiration. Blog ==> tinyurl.com/mmfcbczz
tinyurl.com
PostgreSQL Compatibility Index: The Fellowship of the Database
In the mystical realm of databases, a new hero rises every few moons — a shiny, next-gen PostgreSQL derivative, boldly claiming to be…
4155
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Anyway, thanks for all the blog posts so far, @marcbrooker.bsky.social. They've been great! Looking forward to the formal write up of DSQL you mentioned, too. (I have to assume there's a SIGMOD paper in the works.)
010
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Assuming I didn't flub the definition of the external consistency property provided by strong snapshot isolation—I think I understand now! This is a neat hybrid of existing consistency levels with a unique correctness/performance tradeoff.
110
Nikhil Benesch @benesch.bsky.social · 10/12/2024
(What I find confusing about bringing linearizability into the definition of strong snapshot isolation is that linearizability, as typically defined, implies the existence of a sequential order but no such sequential order exists needs to exist with snapshot isolation.)
110
Nikhil Benesch @benesch.bsky.social · 10/12/2024
With snapshot isolation the situation seems more complicated because transactions need to be described in terms of their read timestamp *and* commit timestamp. Here's my attempt at stating the external consistency property provided by strong snapshot isolation.
100
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Those are good! But the part I'm having trouble with is formalizing the external consistency ("linearizability") guarantee. Spanner defines external consistency in terms of transaction commit timestamps, which is possible because Spanner's commit timestamps reflect the serializable order.
100
Nikhil Benesch @benesch.bsky.social · 10/12/2024
I am however finding SSI a hard isolation level to build intuition for. SS is easy (“as if no concurrency”). SSI is harder: does the real-time constraint apply to start times, commit times, or both? Maybe I can talk @jaffray.bsky.social into writing a NULL BITMAP about this.
220
Nikhil Benesch @benesch.bsky.social · 10/12/2024
My mind was blown recently when I learned that Liskov described these applications for synchronized clocks in 1991 (!): dl.acm.org/doi/pdf/10.1... Based on work that happened even earlier than ‘91. Wild how long it’s taken for these techniques to make their way to industry.
dl.acm.org
000
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Yeah, the comparison to Postgres’s RR was part of what threw me off. This is clarifying though, thanks!
110
Nikhil Benesch @benesch.bsky.social · 10/12/2024
Ah, indeed! I missed the footnote and was mistakenly thinking that “strong consistency” was “snapshot isolation.”
000
Nikhil Benesch @benesch.bsky.social · 09/12/2024
To tie things back to your post: the first example you give of a non-linearizable history is valid under SI and therefore something DSQL might allow, right?
100
Nikhil Benesch @benesch.bsky.social · 09/12/2024
IIUC, one of the more interesting implications here is that clock skew can cause DSQL to violate snapshot isolation. Whereas in Spanner, clock skew can only cause external consistency violations, never serializability violations.
100
Nikhil Benesch @benesch.bsky.social · 09/12/2024
The best I can tell is that DSQL involves the clocks in the adjudicator heartbeat protocol. Based on "...adjudicators promise to move their commit points forward in lock step with the physical clock..." in @marcbrooker.bsky.social's third DSQL vignette: brooker.co.za/blog/2024/12...
brooker.co.za
DSQL Vignette: Transactions and Durability - Marc's Blog
200
Nikhil Benesch @benesch.bsky.social · 09/12/2024
Something I don't understand: what facet of consistency does DSQL actually use the atomic clocks for? Unlike Spanner, which uses the atomic clocks to provide external consistency, DSQL doesn't claim to provide external consistency.
221
Nikhil Benesch @benesch.bsky.social · 09/12/2024
That looks awesome.
030
Nikhil Benesch @benesch.bsky.social · 07/12/2024
Then you can point that at new versions of systems and compute the updated PCI without manual labor, which would be very cool.
010
Nikhil Benesch @benesch.bsky.social · 07/12/2024
Seems plausible to automate? After all if things are fully PG compliant there’s specific SQL they must support. Imagining that for each defined category, there’s a Python scenario which runs some SQL queries and returns a float between 0 and 1 indicating the degree of compat.
330
Reposted by Nikhil Benesch
Chris @chris.blue · 05/12/2024
First new post in a couple of weeks! There's been a lot of activity around regattastorage.com this week, so I decided to write about the space. tl;dr It's pretty exciting!
materializedview.io
The Quest for a Distributed POSIX-Compatible Filesystem
Distributed POSIX filesystems have proven elusive, but we're getting closer. Perhaps that's all we need.
2209
Nikhil Benesch @benesch.bsky.social · 05/12/2024
The claim is just based on the fact that S3 table buckets automatically handle compaction: bsky.app/profile/paul... The comparison is against an uncompacted table in standard bucket. Presumably there's no speedup relative to a manually compacted table in a standard bucket.
130
Nikhil Benesch @benesch.bsky.social · 05/12/2024
Fair point. The tradeoff, I guess, is that, with Polaris, catalog access control is separate from object storage access control, whereas with S3 Tables you can write a single IAM policy that controls access to both the catalog and object storage.
000