Ben @benjdd.com · 21/01/2026You're probably sick of me saying "B-tree" but these impact SO MUCH of database performance. They're used all over the place in Postgres, MySQL, and SQLite. This week I broke down B-tree lookups and how the page cache makes lookups faster. 110
Reposted by BenPlanetScale @planetscale.com · 15/01/2026Introducing pg_strict for Postgres. Our new extension adds a safety net to Postgres, catching dangerous queries before they run. www.youtube.com/watch?v=noPn...youtube.comProtect your database. Use the pg_strict Postgres extension.YouTube video by PlanetScale 193
Ben @benjdd.com · 14/01/2026If databases fascinate you like they do me, this article's for you! Every time you interact with a website, database transactions are keeping your data consistent, safe, and isolated. I wrote an interactive guide to how they work ⬇️ 100
Ben @benjdd.com · 13/01/2026Tuning your database just right can be counter-intuitive, unless you understand all levels of the system. Intuitively, most would say "more work_mem = better" for building indexes, but this hurts performance due to L3 cache behavior. Great article by Tomas Vondra. vondra.me/posts/dont-g... 052
Ben @benjdd.com · 11/01/2026R-trees are a powerful structure for indexing geometric data. They’re used by MySQL, and Postgres uses an R-tree-like structure via GiST in PostGIS. 🧵 110
Ben @benjdd.com · 08/01/2026I'm excited about the database performance io_uring will unlock. Last year I benchmarked Postgres 17 vs 18 to test the initial io_uring upgrades. I was surprised to see they weren't always a clear win for TPC-C. This paper studies the potential, and the future looks good. 161
Ben @benjdd.com · 06/01/2026Software at scale reveals the cracks. Managing a system for a single use-case (databases or otherwise) can make it seem like a perfect solution. It just might be for that narrow environment! At scale you see all the edge cases because you're operating on so many workloads. 120
Ben @benjdd.com · 05/01/2026What do Git, Cursor, and Dynamo have in common? Merkle trees! A great data structure for tracking file changes, facilitating incremental sync with remote servers. Say we want to track changes to a codebase at a per-file level. We compute a hash for each source file, and these become leaf nodes. 150
Ben @benjdd.com · 04/01/2026Need a break from AI in the timeline? Listen to me talk about data organization instead :) Friday's stream was a fun one. Sequential writes, binary search trees, block I/O devices, and B-trees. The latest slice dropped this afternoon. www.youtube.com/watch?v=84b_... 050
Ben @benjdd.com · 04/01/2026If 2026 is the year of AI, it's also the year to read more papers. LLMs make writing code cheaper. This places greater emphasis on architectural choices, understanding design tradeoffs, ensuring security, and building things people actually need. Great example: yesterday I read the Dynamo paper. 160
Ben @benjdd.com · 01/01/20262026 is the year to end TikTok brain. Instead, learn database internals on YouTube. Speaking of which, another dropped today (link below). 140
Ben @benjdd.com · 01/01/2026Cameras, lenses, framing, and everything in-between have fascinated me for many year. This morning I read Bartosz Ciechanowski's article on the subject. It's the best explainer I've seen. The interactivity really sells it. Great article to kick off your year with: ciechanow.ski/cameras-and-... 0131
Ben @benjdd.com · 31/12/2025We need indexes to make databases fast. BUT there are some important time/space and read/write optimization tradeoffs to consider! Latest YT → fun overview of this aspect of databases. www.youtube.com/watch?v=cNw9... 010
Ben @benjdd.com · 31/12/2025This is the best article I've read on MVCC in MySQL. MySQL and Postgres use quite different engineering techniques are used to address the same problem. Undo log vs multiple tuple versions. Another great one by Jeremy Cole. blog.jcole.us/2014/04/16/t... 011
Reposted by BenAlex Miller @alexmillerdb.bsky.social · 30/12/2025I’ve recently seen multiple, unrelated instances of people referencing Bf-trees. Good job, @benjdd.com. 192
Ben @benjdd.com · 30/12/2025Had a great first "Database Internals" livestream yesterday. I'm aiming for more "regular" YouTubing in 2026, much of which will be chopping up interesting segments from the streams. Speaking of which: new video has DROPPED! www.youtube.com/watch?v=wdJe...youtube.comOLTP vs OLAP and the row / column storage tradeoffYouTube video by Benjamin Dicken 061
Ben @benjdd.com · 30/12/2025Amazing how one simple idea can revolutionize an industry. Binary search trees were invented in 1960. It seems obvious today, but this was a fresh way of thinking about ordered data on computers. 110
Ben @benjdd.com · 29/12/2025pgcli and mycli are wonderful upgrades from the default psql / mysql database clients. Auto-completion, syntax highlighting, and just generally much better usability. If you're connecting to your DB from the terminal, get these asap. 020
Ben @benjdd.com · 26/12/2025UX and performance are tightly correlated. Don't treat them as distinct concerns. We've all used software that's slow and becomes a huge turnoff. Fast software makes for happy users. Or at least, avoids making them mad! 120
Ben @benjdd.com · 24/12/2025Goal: benchmark Postgres 8.0 - 18.0. That's 20 years of database performance! I haven't started beyond "planning with Claude," but I expect the hardest part to be building old versions from source. Much has changed in compilers + unix since 2005. Who's done this? Suggestions? 120
Ben @benjdd.com · 22/12/2025PSA to my Postgres people: use a connection pooler. Incredible article on when and why to use PgBouncer. Includes a great explanation of how increasing direct connections leads to more contention → degraded performance. (+ benchmarks too!) 120
Ben @benjdd.com · 19/12/2025Log structures are all over the place in databases, but did you know they are used in file systems too? This week I re-read the iconic LFS paper by Rosenblum / Ousterhout. The differences between I/O demands on a database vs a general-purpose FS are neat to study. 151
Ben @benjdd.com · 18/12/2025I present to you: the 8 LOC database. Who needs ACID, relational schema, foreign keys, B-tree indexes, log-based commits, MVCC, replication, and failovers? We've been overthinking the database. Keep it simple. 150
Ben @benjdd.com · 17/12/2025GIN indexes are a powerful tool in Postgres. They’re great for inverting the typical use case. Instead of mapping “the row with ID 2 contains ‘become a database expert’” you flip it to say “The word ‘database’ maps to the rows with IDs 1, 2, and 3 and ‘expert’ maps to the row with ID 2.” 020
Ben @benjdd.com · 15/12/2025Two biggest customers requests since launching Metal: - Smaller compute sizes at a lower price point - More local-NVMe storage size options Today we've delivered both. Happy databasing. 020
Ben @benjdd.com · 14/12/2025Postgres, MySQL, SQLite and many others were invented in the 90s and 00s, the era of spinning disks. A local NVMe SSD has ~1000x improvement in both throughput and latency. If we had to throw these databases away and begin from scratch in 2025, what would change and what would remain? 110
Ben @benjdd.com · 12/12/20254 months later, DDIA complete. All 12 chapters. Wonder how many people who post about this book have read the whole thing? Parting thoughts (in thread) 100
Ben @benjdd.com · 11/12/2025Being knee-deep in Postgres, things I miss from MySQL: - B-tree table storage - Vitess - Undo log based MVCC - SHOW TABLES - Vitess - Buffer that doesn't rely on OS page cache - SHOW CREATE DATABASE But mostly, Vitess. 010
Ben @benjdd.com · 10/12/2025Backpressure is key for well-behaved infrastructure. When databases, message queues, connection poolers, or other systems get overloaded, how do they deal with the pressure? 111
Ben @benjdd.com · 08/12/2025IO Devices and Latency is the most ambitious article I've written. Still figuring out how to top this in 2026. Taking ideas. 160
Ben @benjdd.com · 07/12/2025Postgres has a process-per-connection architecture. Therefore, you should use a connection pooler whenever possible. These sit between the app and database. They maintain a pool of connections and dynamically map incoming requests to to them. PgBouncer is the most popular tool for this job! 110
Ben @benjdd.com · 04/12/2025Speaking of MySQL powering the internet - Uber runs on over 2,600 MySQL clusters! They recently moved many of these from a "traditional" primary-replica model to Paxos-based group replication. Great benchmarks included on their blog too. www.uber.com/blog/improvi... 020
Ben @benjdd.com · 03/12/2025Come work with me at PlanetScale! We're growing the team that educates and supports our engineering community. If you love databases and seeing customers succeed, these roles are for you. Specifically, we're hiring a Community Engineer and Developer Events Lead. planetscale.com/careersplanetscale.comCareers — PlanetScaleBuilding the database of the future, together. 040
Ben @benjdd.com · 01/12/2025ORMs are great, but you still need to learn SQL + database performance fundamentals. Running an app at scale demands a deeper understanding of indexing, joins, and data access patterns. 070
Ben @benjdd.com · 30/11/2025Think databases are boring? Indexing → data structures + algorithms AI → RAG + learned indexes MVCC → concurrent programming Sharding → distributed systems Query parsing → formal languages Query planning → stats + optimization Databases encompass tons of interesting software engineering problems. 020
Ben @benjdd.com · 28/11/2025Every now and then I stumble across an underrated article. An example: "Measuring Latencies Between AWS Availability Zones" (link below). An incredible resource for anyone building on AWS. What are your favorite, lesser-known blogs? www.bitsand.cloud/posts/cross-... 051
Ben @benjdd.com · 26/11/2025On a managed database, YOU benefit from the shared experience of the entire fleet. There's tons of hard-to-predict issues. Intermittent AWS network lag. PG + MySQL edge cases. etc. PlanetScale continually hardens our systems against these problems, and all customers win. 030
Ben @benjdd.com · 24/11/2025Postgres supports point-in-time recovery: the ability to "travel back in time" to a previous database state. This works via a combination of backup-restores and WAL replay. 100
Ben @benjdd.com · 23/11/2025These graphs never get old. Depot uses one of my favorite tools, PlanetScale Insights, to drill in on query performance issues and address them via schema change DRs. 120
Ben @benjdd.com · 23/11/2025Bf-trees are a high-performance alternative to the classic B-tree. The key contribution: Decoupling the size of on-disk pages from the size of memory cache entries. I'm curious if/when MySQL and Postgres will adopt modern versions of classic indexes. 1151
Reposted by BenPlanetScale @planetscale.com · 21/11/2025Postgres databases now come with LLM-powered index recommendations. Indexes are crucial for database performance, and these recommendations help ensure your queries are executing optimally. Read about how we built this: planetscale.com/blog/postgre...planetscale.comAI-Powered Postgres index suggestions — PlanetScaleIntroducing AI-powered index suggestions for PostgreSQL 042
Ben @benjdd.com · 21/11/2025Intercom is powered by nearly 1000 MySQL instances. Sharding brought to you by Vitess. Amazing talk by Eugene Kenny at SF Ruby conf. 031
Ben @benjdd.com · 20/11/2025What are your favorite technical papers? I want to do some livestreams going over papers as a group. Much like my grad-school days of sitting in a classroom and discussing design tradeoffs of popular tech literature. 031
Ben @benjdd.com · 19/11/2025Postgres uses TOAST to store large, variable-sized data values like JSONB and TEXT. Using it impacts performance, so you gotta know the tradeoffs. 121
Reposted by BenRichard Crowley @rcrowley.org · 17/11/2025Don't be alarmed, I won't be blogging daily, but I did just publish a second article in as many days. Everything is lock-in rcrowley.org/2025/everyth...rcrowley.orgEverything is lock-in — Richard Crowley 021