Sign in

Ben

@benjdd.com
324 followers 66 following 357 posts

databases @planetscale.com Some stuff I've written: - pscale.link/io - pscale.link/btrees - pscale.link/sharding Find me at benjdd.com

PostsRepliesMedia
Ben @benjdd.com · 21/01/2026
www.youtube.com/watch?v=syPE...
youtube.com
using RAM to make databases FAST (the buffer pool)
YouTube video by Ben Dicken
010
Ben @benjdd.com · 21/01/2026
You're probably sick of me saying "B-tree" but these impact SO MUCH of database performance. They're used all over the place in Postgres, MySQL, and SQLite. This week I broke down B-tree lookups and how the page cache makes lookups faster.
110
Ben @benjdd.com · 20/01/2026
Choose your storage layer carefully.
011
Ben @benjdd.com · 19/01/2026
Good catch. Fixed.
010
Ben @benjdd.com · 16/01/2026
Blood, sweat, and tears 😅 But seriously, these were Js + GSAP, and in this case Cursor helped quite a bit too.
120
Reposted by Ben
PlanetScale @planetscale.com · 15/01/2026
Introducing pg_strict for Postgres. Our new extension adds a safety net to Postgres, catching dangerous queries before they run. www.youtube.com/watch?v=noPn...
youtube.com
Protect your database. Use the pg_strict Postgres extension.
YouTube video by PlanetScale
193
Ben @benjdd.com · 14/01/2026
pscale.link/transactions
pscale.link
Database Transactions — PlanetScale
What are database transactions and how do SQL databases isolate one transaction from another?
200
Ben @benjdd.com · 14/01/2026
If databases fascinate you like they do me, this article's for you! Every time you interact with a website, database transactions are keeping your data consistent, safe, and isolated. I wrote an interactive guide to how they work ⬇️
100
Ben @benjdd.com · 13/01/2026
Tuning your database just right can be counter-intuitive, unless you understand all levels of the system. Intuitively, most would say "more work_mem = better" for building indexes, but this hurts performance due to L3 cache behavior. Great article by Tomas Vondra. vondra.me/posts/dont-g...
052
Ben @benjdd.com · 11/01/2026
A key difference from B-trees: searching for a single rectangle may require searching multiple tree paths! In the ideal case, B-trees offer O(log n) search performance, but due to possible overlaps the worst-case performance is actually O(n).
000
Ben @benjdd.com · 11/01/2026
Moving up the tree, the bounding rectangles get larger and larger, up to the root node storing a small number of large bounding boxes.
100
Ben @benjdd.com · 11/01/2026
Entries in an R-tree are bounding rectangles. At the leaves we store the minimum bounding rectangle (MBR) for each region being stored, with a reference to the full geometry stored on-disk elsewhere. The parent entry of a leaf MBR stores a bounding rectangle that fully bounds all children.
100
Ben @benjdd.com · 11/01/2026
They function similarly to B-trees: It’s a tree structure with multiple entries at each page-aligned node. This generally keeps the trees nice and shallow, and allows for efficient lookups for millions of elements stored on-disk. It also generally only stores data values at the leaves, like B+trees.
100
Ben @benjdd.com · 11/01/2026
R-trees are a powerful structure for indexing geometric data. They’re used by MySQL, and Postgres uses an R-tree-like structure via GiST in PostGIS. 🧵
110
Ben @benjdd.com · 08/01/2026
TLDR: io_uring won't help much if treated as a drop-in replacement for existing database I/O architectures. Better performance will often require architectural changes. When applied, there's tons of performance gains to be had. Here's the paper: arxiv.org/pdf/2512.04859
arxiv.org
031
Ben @benjdd.com · 08/01/2026
I'm excited about the database performance io_uring will unlock. Last year I benchmarked Postgres 17 vs 18 to test the initial io_uring upgrades. I was surprised to see they weren't always a clear win for TPC-C. This paper studies the potential, and the future looks good.
161
Ben @benjdd.com · 06/01/2026
At PlanetScale we observe this first-hand for both MySQL and Postgres. We then get to go tackle these hard-but-fun engineering problems, so our customers don't have to.
010
Ben @benjdd.com · 06/01/2026
Software at scale reveals the cracks. Managing a system for a single use-case (databases or otherwise) can make it seem like a perfect solution. It just might be for that narrow environment! At scale you see all the edge cases because you're operating on so many workloads.
120
Ben @benjdd.com · 05/01/2026
This makes for fast change detection (O(log n)), and saves bandwidth since we only need to re-sync files that we know have been modified.
000
Ben @benjdd.com · 05/01/2026
When syncing local file changes with a remote server (like Git) we can quickly tell if changes were made by comparing the client's root hash to the server's. If they differ, the tree is navigated to find the leaf node(s) with changes, and only those files need to be re-synced.
100
Ben @benjdd.com · 05/01/2026
We then build a tree based on the directory structure. A parent node's hash is a hash of the concatenation of all its children's hashes. Inner nodes' hashes are based on the data / hashes of its descendants. This culminates in the root node, whose hash is based on ALL the tracked source files.
100
Ben @benjdd.com · 05/01/2026
What do Git, Cursor, and Dynamo have in common? Merkle trees! A great data structure for tracking file changes, facilitating incremental sync with remote servers. Say we want to track changes to a codebase at a per-file level. We compute a hash for each source file, and these become leaf nodes.
150
Ben @benjdd.com · 04/01/2026
Need a break from AI in the timeline? Listen to me talk about data organization instead :) Friday's stream was a fun one. Sequential writes, binary search trees, block I/O devices, and B-trees. The latest slice dropped this afternoon. www.youtube.com/watch?v=84b_...
050
Ben @benjdd.com · 04/01/2026
Equip yourself with the fundamental building-blocks of software systems. This combined with the right LLM tooling can take you very far.
000
Ben @benjdd.com · 04/01/2026
Merkle trees, consistent hashing, vector clocks, gossip protocols, and quorum algorithms were all previously known and had been used to build other software. But their unique combination to build Dynamo worked super well for Amazon's use-case, and helped them to scale to millions of DAU.
100
Ben @benjdd.com · 04/01/2026
Though it was published in 2007 (nearly 20 years ago!) it was revolutionary in its day, and an example of how already-known technologies can be combined to make something new and extremely successful.
100
Ben @benjdd.com · 04/01/2026
If 2026 is the year of AI, it's also the year to read more papers. LLMs make writing code cheaper. This places greater emphasis on architectural choices, understanding design tradeoffs, ensuring security, and building things people actually need. Great example: yesterday I read the Dynamo paper.
160
Ben @benjdd.com · 01/01/2026
www.youtube.com/watch?v=K4YM...
youtube.com
Data storage engines: the 5 must-know components
YouTube video by Benjamin Dicken
020
Ben @benjdd.com · 01/01/2026
2026 is the year to end TikTok brain. Instead, learn database internals on YouTube. Speaking of which, another dropped today (link below).
140
Ben @benjdd.com · 01/01/2026
Cameras, lenses, framing, and everything in-between have fascinated me for many year. This morning I read Bartosz Ciechanowski's article on the subject. It's the best explainer I've seen. The interactivity really sells it. Great article to kick off your year with: ciechanow.ski/cameras-and-...
0131
Ben @benjdd.com · 31/12/2025
We need indexes to make databases fast. BUT there are some important time/space and read/write optimization tradeoffs to consider! Latest YT → fun overview of this aspect of databases. www.youtube.com/watch?v=cNw9...
010
Ben @benjdd.com · 31/12/2025
Sounds interesting. Added to the list!
020
Ben @benjdd.com · 31/12/2025
This is the best article I've read on MVCC in MySQL. MySQL and Postgres use quite different engineering techniques are used to address the same problem. Undo log vs multiple tuple versions. Another great one by Jeremy Cole. blog.jcole.us/2014/04/16/t...
011
Reposted by Ben
Alex Miller @alexmillerdb.bsky.social · 30/12/2025
I’ve recently seen multiple, unrelated instances of people referencing Bf-trees. Good job, @benjdd.com.
192
Ben @benjdd.com · 30/12/2025
It was a great paper! My sense is, databases like MySQL and Postgres have plenty of room for improvement if they can effectively adopt modern B-tree optimizations, Bf-trees or otherwise. Fun to consider.
110
Ben @benjdd.com · 30/12/2025
Had a great first "Database Internals" livestream yesterday. I'm aiming for more "regular" YouTubing in 2026, much of which will be chopping up interesting segments from the streams. Speaking of which: new video has DROPPED! www.youtube.com/watch?v=wdJe...
youtube.com
OLTP vs OLAP and the row / column storage tradeoff
YouTube video by Benjamin Dicken
061
Ben @benjdd.com · 30/12/2025
This spawned hundreds of structures over the next 50 years including: - AVL trees (balanced versions of BSTs) - B-trees (optimized for on-disk storage) - R-trees (for storing geospatial data) - LSM trees (optimized for write-heavy persistence) - Suffix trees (string operations)
100
Ben @benjdd.com · 30/12/2025
Amazing how one simple idea can revolutionize an industry. Binary search trees were invented in 1960. It seems obvious today, but this was a fresh way of thinking about ordered data on computers.
110
Ben @benjdd.com · 29/12/2025
pgcli and mycli are wonderful upgrades from the default psql / mysql database clients. Auto-completion, syntax highlighting, and just generally much better usability. If you're connecting to your DB from the terminal, get these asap.
020
Ben @benjdd.com · 27/12/2025
> Good perf contributes to good UX but not so vice versa. Good point! There's certainly fast apps that are terrible to use.
000
Ben @benjdd.com · 26/12/2025
UX and performance are tightly correlated. Don't treat them as distinct concerns. We've all used software that's slow and becomes a huge turnoff. Fast software makes for happy users. Or at least, avoids making them mad!
120
Ben @benjdd.com · 24/12/2025
Others shared this with me on X and LinkedIn too. Seems like amazing work. Appreciate the link.
010
Ben @benjdd.com · 24/12/2025
Goal: benchmark Postgres 8.0 - 18.0. That's 20 years of database performance! I haven't started beyond "planning with Claude," but I expect the hardest part to be building old versions from source. Much has changed in compilers + unix since 2005. Who's done this? Suggestions?
120
Ben @benjdd.com · 22/12/2025
Read here: brandur.org/postgres-con...
brandur.org
How to Manage Connections Efficiently in Postgres, or Any Database
Hitting the limit for maximum allowed connections is a common operational problem in Postgres. Here we look at a few techniques for managing connections and making efficient use of those that are avai...
010
Ben @benjdd.com · 22/12/2025
PSA to my Postgres people: use a connection pooler. Incredible article on when and why to use PgBouncer. Includes a great explanation of how increasing direct connections leads to more contention → degraded performance. (+ benchmarks too!)
120
Ben @benjdd.com · 22/12/2025
Multiple people have commented about Go abstractions (including over on X). I should study this more. Do you do much with Go concurrency?
110
Ben @benjdd.com · 21/12/2025
Name a more beautiful abstraction than the Unix pipe.
120
Ben @benjdd.com · 19/12/2025
The paper: web.stanford.edu/~ouster/cgi-...
web.stanford.edu
010
Ben @benjdd.com · 19/12/2025
Log structures are all over the place in databases, but did you know they are used in file systems too? This week I re-read the iconic LFS paper by Rosenblum / Ousterhout. The differences between I/O demands on a database vs a general-purpose FS are neat to study.
151
Ben @benjdd.com · 19/12/2025
Haha yes, please don't really use this 😅
000