Sign in

Ben

@benjdd.com
324 followers 66 following 357 posts

databases @planetscale.com Some stuff I've written: - pscale.link/io - pscale.link/btrees - pscale.link/sharding Find me at benjdd.com

PostsRepliesMedia
Ben @benjdd.com · 21/01/2026
You're probably sick of me saying "B-tree" but these impact SO MUCH of database performance. They're used all over the place in Postgres, MySQL, and SQLite. This week I broke down B-tree lookups and how the page cache makes lookups faster.
110
Ben @benjdd.com · 20/01/2026
Choose your storage layer carefully.
011
Ben @benjdd.com · 14/01/2026
If databases fascinate you like they do me, this article's for you! Every time you interact with a website, database transactions are keeping your data consistent, safe, and isolated. I wrote an interactive guide to how they work ⬇️
100
Ben @benjdd.com · 13/01/2026
Tuning your database just right can be counter-intuitive, unless you understand all levels of the system. Intuitively, most would say "more work_mem = better" for building indexes, but this hurts performance due to L3 cache behavior. Great article by Tomas Vondra. vondra.me/posts/dont-g...
052
Ben @benjdd.com · 11/01/2026
R-trees are a powerful structure for indexing geometric data. They’re used by MySQL, and Postgres uses an R-tree-like structure via GiST in PostGIS. 🧵
110
Ben @benjdd.com · 08/01/2026
I'm excited about the database performance io_uring will unlock. Last year I benchmarked Postgres 17 vs 18 to test the initial io_uring upgrades. I was surprised to see they weren't always a clear win for TPC-C. This paper studies the potential, and the future looks good.
161
Ben @benjdd.com · 05/01/2026
What do Git, Cursor, and Dynamo have in common? Merkle trees! A great data structure for tracking file changes, facilitating incremental sync with remote servers. Say we want to track changes to a codebase at a per-file level. We compute a hash for each source file, and these become leaf nodes.
150
Ben @benjdd.com · 04/01/2026
Need a break from AI in the timeline? Listen to me talk about data organization instead :) Friday's stream was a fun one. Sequential writes, binary search trees, block I/O devices, and B-trees. The latest slice dropped this afternoon. www.youtube.com/watch?v=84b_...
050
Ben @benjdd.com · 04/01/2026
If 2026 is the year of AI, it's also the year to read more papers. LLMs make writing code cheaper. This places greater emphasis on architectural choices, understanding design tradeoffs, ensuring security, and building things people actually need. Great example: yesterday I read the Dynamo paper.
160
Ben @benjdd.com · 01/01/2026
2026 is the year to end TikTok brain. Instead, learn database internals on YouTube. Speaking of which, another dropped today (link below).
140
Ben @benjdd.com · 01/01/2026
Cameras, lenses, framing, and everything in-between have fascinated me for many year. This morning I read Bartosz Ciechanowski's article on the subject. It's the best explainer I've seen. The interactivity really sells it. Great article to kick off your year with: ciechanow.ski/cameras-and-...
0131
Ben @benjdd.com · 31/12/2025
We need indexes to make databases fast. BUT there are some important time/space and read/write optimization tradeoffs to consider! Latest YT → fun overview of this aspect of databases. www.youtube.com/watch?v=cNw9...
010
Ben @benjdd.com · 31/12/2025
This is the best article I've read on MVCC in MySQL. MySQL and Postgres use quite different engineering techniques are used to address the same problem. Undo log vs multiple tuple versions. Another great one by Jeremy Cole. blog.jcole.us/2014/04/16/t...
011
Ben @benjdd.com · 29/12/2025
pgcli and mycli are wonderful upgrades from the default psql / mysql database clients. Auto-completion, syntax highlighting, and just generally much better usability. If you're connecting to your DB from the terminal, get these asap.
020
Ben @benjdd.com · 22/12/2025
PSA to my Postgres people: use a connection pooler. Incredible article on when and why to use PgBouncer. Includes a great explanation of how increasing direct connections leads to more contention → degraded performance. (+ benchmarks too!)
120
Ben @benjdd.com · 19/12/2025
Log structures are all over the place in databases, but did you know they are used in file systems too? This week I re-read the iconic LFS paper by Rosenblum / Ousterhout. The differences between I/O demands on a database vs a general-purpose FS are neat to study.
151
Ben @benjdd.com · 18/12/2025
I present to you: the 8 LOC database. Who needs ACID, relational schema, foreign keys, B-tree indexes, log-based commits, MVCC, replication, and failovers? We've been overthinking the database. Keep it simple.
150
Ben @benjdd.com · 17/12/2025
GIN indexes are a powerful tool in Postgres. They’re great for inverting the typical use case. Instead of mapping “the row with ID 2 contains ‘become a database expert’” you flip it to say “The word ‘database’ maps to the rows with IDs 1, 2, and 3 and ‘expert’ maps to the row with ID 2.”
020
Ben @benjdd.com · 12/12/2025
4 months later, DDIA complete. All 12 chapters. Wonder how many people who post about this book have read the whole thing? Parting thoughts (in thread)
100
Ben @benjdd.com · 10/12/2025
Backpressure is key for well-behaved infrastructure. When databases, message queues, connection poolers, or other systems get overloaded, how do they deal with the pressure?
111
Ben @benjdd.com · 09/12/2025
Who's ready to build Sonnet 5.0 together?
000
Ben @benjdd.com · 08/12/2025
IO Devices and Latency is the most ambitious article I've written. Still figuring out how to top this in 2026. Taking ideas.
160
Ben @benjdd.com · 07/12/2025
Postgres has a process-per-connection architecture. Therefore, you should use a connection pooler whenever possible. These sit between the app and database. They maintain a pool of connections and dynamically map incoming requests to to them. PgBouncer is the most popular tool for this job!
110
Ben @benjdd.com · 04/12/2025
Speaking of MySQL powering the internet - Uber runs on over 2,600 MySQL clusters! They recently moved many of these from a "traditional" primary-replica model to Paxos-based group replication. Great benchmarks included on their blog too. www.uber.com/blog/improvi...
020
Ben @benjdd.com · 28/11/2025
Every now and then I stumble across an underrated article. An example: "Measuring Latencies Between AWS Availability Zones" (link below). An incredible resource for anyone building on AWS. What are your favorite, lesser-known blogs? www.bitsand.cloud/posts/cross-...
051
Ben @benjdd.com · 25/11/2025
Databases 🤝 Physics
030
Ben @benjdd.com · 24/11/2025
Postgres supports point-in-time recovery: the ability to "travel back in time" to a previous database state. This works via a combination of backup-restores and WAL replay.
100
Ben @benjdd.com · 23/11/2025
These graphs never get old. Depot uses one of my favorite tools, PlanetScale Insights, to drill in on query performance issues and address them via schema change DRs.
120
Ben @benjdd.com · 23/11/2025
Another Rust W.
040
Ben @benjdd.com · 23/11/2025
Bf-trees are a high-performance alternative to the classic B-tree. The key contribution: Decoupling the size of on-disk pages from the size of memory cache entries. I'm curious if/when MySQL and Postgres will adopt modern versions of classic indexes.
1151
Ben @benjdd.com · 21/11/2025
Intercom is powered by nearly 1000 MySQL instances. Sharding brought to you by Vitess. Amazing talk by Eugene Kenny at SF Ruby conf.
031
Ben @benjdd.com · 19/11/2025
Postgres uses TOAST to store large, variable-sized data values like JSONB and TEXT. Using it impacts performance, so you gotta know the tradeoffs.
121
Ben @benjdd.com · 17/11/2025
S3 is powered by millions of HDDs. This Andy Warfield article is a fascinating look under the covers one of the largest storage systems on earth. I particularly enjoyed the discussion of scaling both the technology and ones self as an engineer. www.allthingsdistributed.com/2023/07/buil...
062
Ben @benjdd.com · 13/11/2025
PlanetScale Insights is amazing. Correlating schema/config changes and performance has never been easier. The perfect tool for squeezing every ounce of performance out of your database resources.
031
Ben @benjdd.com · 12/11/2025
Postgres handles MVCC with per-row transaction ID metadata. Every row (tuple) you insert into a Postgres table includes xmin and xmax metadata. xmin is the transaction that created the row, and xmax the one that updates or deleted it.
110
Ben @benjdd.com · 10/11/2025
Not all CPUs are created equal. This latency graph shows a r6id → i7i upgrade. Same vCPU count. Same amount of RAM. Both using local SSDs. Upgrading to a newer instance makes a big difference.
020
Ben @benjdd.com · 07/11/2025
What happens when you INSERT a row in Postgres? Postgres needs to ensure that data is durable while maintaining good write performance + crash recovery ability. The key is in the Write-Ahead Log (WAL).
110
Ben @benjdd.com · 06/11/2025
PlanetScale now supports PgBouncers (connection poolers) for your replicas. Connection pooling is broadly important for databases, but especially for Postgres because of its process-per-connection architecture. Don't know what that means? I have the perfect article!
131
Ben @benjdd.com · 06/11/2025
Much of the the internet runs on Elastic Block Storage. It's the default (and in most cases, required) storage layer for every EC2 instance running in AWS. This article by Marc Olson was a delightful read on its history and engineering challenges. www.allthingsdistributed.com/2024/08/cont...
052
Ben @benjdd.com · 05/11/2025
Want to understand B-trees better? Try btree.app and bplustree.app. These are standalone sandboxes of the visuals I built for my "B-trees and database indexes" article. Helpful for learning B-tree insertion, search, and node splits.
030
Ben @benjdd.com · 04/11/2025
Explain is a powerful tool in Postgres. If you care about performance, get comfortable running `explain` and `explain analyze` commands regularly, and learn how to interpret its output. This blog is a great intro. www.depesz.com/2013/04/16/e...
051
Ben @benjdd.com · 02/11/2025
Choose your storage layer carefully! Elastic Block Storage (EBS) is great for low-I/O workloads, but becomes a bottleneck or cost sink for heavy workloads. Two common types of EBS are gp3 and io2. Both are network-attached storage backed by SSDs, but have different performance characteristics.
130
Ben @benjdd.com · 29/10/2025
Foreign keys vs constraints. Many conflate the two. Foreign key: A column that establishes a relationship between two tables. This is frequently set up as a column in one table (post) that stores primary key values from another table (user) so that it can join between the two.
230
Ben @benjdd.com · 27/10/2025
This is my favorite article on B-tree internals. Jeremy Cole has done MySQL at Twitter, Google, and Shopify. One of the most knowledgeable MySQL engineers in the world, and his blog is an information goldmine. blog.jcole.us/2013/01/10/b...
072
Ben @benjdd.com · 26/10/2025
Comparing latencies for Claude, GPT-5, Grok, and Gemini. My personal biggest takeaway? Gemini is fast! I mostly stick with Claude but maybe I should give it a go.
230
Ben @benjdd.com · 23/10/2025
Beware of phantom reads. In both Postgres and MySQL, it's possible for identical SELECTs in the same transaction to see different results.
110
Ben @benjdd.com · 20/10/2025
The Google Spanner paper is a 10/10 read. The most interesting bit? TrueTime, an API that gives time results with error bounds, and one that guarantees <= 7 ms of clock skew. Impressive! The downside? Not OSS. Even if it were, would require reproducing the timing hardware.
131
Ben @benjdd.com · 17/10/2025
My half-finished "Transactions" article has been sitting dormant for awhile. Time to bring it back? Started this long before PlanetScale Postgres. Would be a fun opportunity to compare the MVCC models of Postgres and MySQL.
140
Ben @benjdd.com · 16/10/2025
Any component can fail at any time in distributed systems. How do you communicate in the face of unreliability? It's not easy! This is classically known as the "two generals problem" and is a great way to think about communicating over unreliable channels.
110
Ben @benjdd.com · 15/10/2025
Biggest surprise from my Postgres 17 vs 18 benchmarks? io_uring was frequently the worst performer! It was slower than both `sync` and `worker` in many cases. Tomas Vondra has a wonderful article explaining why it isn't always the best choice. vondra.me/posts/tuning...
000