Sign in

Frank McSherry

@frankmcsherry.bsky.social
381 followers 5 following 191 posts

github.com/frankmcsherry/blog

PostsRepliesMedia
Frank McSherry @frankmcsherry.bsky.social · 09/06/2026
The datatoad book (www.frankmcsherry.org/datatoad/cha...) now comes with a WASM playground to try out some of the problematic Datalog it handles. The book is not fully revved to always have runnable examples, but I'll get on that!
frankmcsherry.org
Introduction
140
Frank McSherry @frankmcsherry.bsky.social · 08/05/2026
I wrote about "directional predicates" in datatoad. Treating predicates as relations can be much more powerful than predicates as .. predicates or even functions. Computation can go faster, because you reveal more to the computer about how to get work done. github.com/frankmcsherr...
github.com
120
Frank McSherry @frankmcsherry.bsky.social · 01/05/2026
I got Claude to help with a blog index refresh, surfacing about 1.5 years of posts to the front page that had been languishing: github.com/frankmcsherr.... It unintentionally included one unfinished post that I'll now have to get sorted out. :D
github.com
GitHub - frankmcsherry/blog: Some notes on things I find interesting and important.
Some notes on things I find interesting and important. - frankmcsherry/blog
030
Reposted by Frank McSherry
Moritz Hoffmann @antiguru.bsky.social · 29/04/2026
Made a thing: Pollard is a MCP server for analyzing Firefox performance profiles. crates.io/crates/pollard
crates.io
crates.io: Rust Package Registry
crates.io serves as a central registry for sharing crates, which are packages or libraries written in Rust that you can use to enhance your projects
031
Frank McSherry @frankmcsherry.bsky.social · 01/04/2026
We now have datatoad's mdbook section on joins. It covers the planning of joins of abstract directed relations, and their execution by a mix of binary and worst-case optimal join stages. It's pretty tight, but I don't think it's missing much except examples. www.frankmcsherry.org/datatoad/cha...
frankmcsherry.org
Joins
030
Frank McSherry @frankmcsherry.bsky.social · 29/03/2026
As part of an attempt to get Claude to help me communicate better, we have .. a Claude-assisted mdbook for datatoad: www.frankmcsherry.org/datatoad/int... Still a work in progress, but it came together much faster with help.
frankmcsherry.org
Datatoad
051
Frank McSherry @frankmcsherry.bsky.social · 28/03/2026
Totally innocent question: I've found Claude has made me rather productive, at a rate that outpaces what I can write about solo. I'm trying to figure out if there is value/space for an intentionally "slop" line of posts that explains work that has been done, minus the alleged charm and certain salt.
210
Frank McSherry @frankmcsherry.bsky.social · 26/03/2026
A very nice post out of @materialize.com about "representation types" which remove non-structural distinctions between types as they approach the dataflow rendering layer, unlocking physical optimizations across logically distinct types: materialize.com/blog/no-clas...
materialize.com
No Classification without Represention | Materialize
Learn how Materialize improves query performance by compiling SQL's type system to a simpler system of
150
Frank McSherry @frankmcsherry.bsky.social · 23/03/2026
Fun experiment du jour: can Claude write stable matching in differential dataflow using a simplified language that hides any of details of DD. Yes, it turns out! Correct out of the box, and with a bit of nudging it got something better than I have written:
A program that performs Gale-Shapely using a small language for differential dataflow.
160
Frank McSherry @frankmcsherry.bsky.social · 20/03/2026
With Claude, and a few lines of Rust, you can add a live timely dataflow visualizer to your programs, giving you visibility into where the time goes, and soon (surely) arrangement record counts as well!
130
Frank McSherry @frankmcsherry.bsky.social · 19/03/2026
A new blog post about accelerating timely dataflow by 100x. Part of the story is how timely is able to do this, and .. seemingly .. other stream processor are not. It comes down to a fundamentally different approach to tracking time: globally rather than locally. github.com/frankmcsherr...
github.com
130
Frank McSherry @frankmcsherry.bsky.social · 14/03/2026
I used Claude for the first time this month, having only ever used copilot before (and .. stopped immediately afterwards). It turns out .. I am maybe obsolete? We improved `columnar` together, and each wrote a post about it. (linked within my post).
github.com
1110
Frank McSherry @frankmcsherry.bsky.social · 05/03/2026
A recent post by @dov.dev about Materialize's new Iceberg sink, and generally about connecting live operational data to analytics stores. Some interesting detail about (overcome) streaming-batch friction, and turning pristine CDC streams into .. Iceberg! materialize.com/blog/making-...
materialize.com
Making Iceberg Work for Operational Data | Materialize
Apache Iceberg was built for batch analytics — but operational data changes continuously. Learn how Materialize streams live, transactionally consistent data into Iceberg without the memory and latenc...
082
Frank McSherry @frankmcsherry.bsky.social · 04/03/2026
Worlds collide: pretty sure someone at the MTA plays FFXIV:
Four roles, in the colors of FFXIV.
000
Frank McSherry @frankmcsherry.bsky.social · 01/03/2026
I wrote a tiny vectorized interpreter last night, and .. thought I'd write up a post describing it. Nothing earth shattering here, but if you haven't seen one of these before it could be interesting! Probably easier to read than the J Incunabulum, but also less .. wow. github.com/frankmcsherr...
github.com
092
Frank McSherry @frankmcsherry.bsky.social · 26/02/2026
We have a recent @materialize.com post from Jan Teske about how MZ's self-correcting materialized views work. It's imo a very cool thing that unpacks some of the magic about how this could possibly work as the underlying system evolves. materialize.com/blog/self-co...
materialize.com
Self-Correcting Materialized Views | Materialize
Learn how Materialize uses self-correction to prevent output drift in materialized views, ensure consistency across upgrades, and enable in-place view replacement.
020
Frank McSherry @frankmcsherry.bsky.social · 05/01/2026
I wrote a bit about the "demand transform" for Datalog (and other similar languages): github.com/frankmcsherr... The idea of the transform is that rather than eagerly produce all the data you might ever need, you can start from the seed crystals of concrete values and data that might yield output.
github.com
120
Frank McSherry @frankmcsherry.bsky.social · 27/12/2025
If anyone ever told you Datalog is elegant, they probably haven't used datatoad yet. The evenness condition is an interesting example of relation-rather-than-function, though!
A Datalog program for determining iterates of the Collatz conjecture.
010
Frank McSherry @frankmcsherry.bsky.social · 25/12/2025
I went through and validated the output counts (numbers of facts) for most of the datatoad outputs. Some were hard to validate because datatoad sneaked in on my laptop but the other systems paged too hard, but everything seems correct. github.com/frankmcsherr...
github.com
130
Frank McSherry @frankmcsherry.bsky.social · 23/12/2025
I put up a new post about work that I am perhaps irrationally excited about: the combination of worst-case optimal joins with relational programming. github.com/frankmcsherr... The high-order bits are that relational programming is very neat, and it fits great with modern (WCO) relational joins.
github.com
120
Frank McSherry @frankmcsherry.bsky.social · 14/12/2025
I wrote a post evaluating datatoad using the framework from the recent FlowLog paper: github.com/frankmcsherr.... It turns out it does well in some cases, worse in others, and has already improved by having other folks shine a light on its limitations by choosing problems and datasets I ignored!
github.com
130
Frank McSherry @frankmcsherry.bsky.social · 04/12/2025
I wrote about datatoad as a "worst-case optimal Datalog", which .. perhaps doesn't typecheck. A previous result on streaming worst-case optimal joins seems to connect, and swapping "streaming" for "iterative" gives something that may be syntactically correct catnip. github.com/frankmcsherr...
github.com
100
Frank McSherry @frankmcsherry.bsky.social · 22/11/2025
Getting crisper; now within ~1s (14.8s v 13.7s) of my hand-optimized plan that uses only binary joins, but without any plan hints. Erm, mostly. Hoping to bypass it, as a win for WCO joins. :D
120
Frank McSherry @frankmcsherry.bsky.social · 21/11/2025
I wrote a little bit about worst-case optimal joins in datatoad, support for which has just landed although there is a bit of tidying still to do before they are as crisp as the non-wco joins. github.com/frankmcsherr...
github.com
130
Frank McSherry @frankmcsherry.bsky.social · 17/11/2025
A short note on how one can perform joins of sums by accumulating at a point halfway through the join. Neither the input nor the output of the join, but smaller intermediate terms that show up when you peer closer at what a join does internally. github.com/frankmcsherr...
github.com
030
Frank McSherry @frankmcsherry.bsky.social · 02/11/2025
I wrote a bit about datatoad going "full columnar". All operations are now columnar; no rows are ever formed; everything is column-at-a-time. github.com/frankmcsherr... A bunch of interesting (to me) algorithms, and also some performance regressions, but then clawing back. I learned things!
github.com
130
Frank McSherry @frankmcsherry.bsky.social · 25/10/2025
Looking forward to this!
051
Reposted by Frank McSherry
Materialize @materialize.com · 22/10/2025
New from Materialize: Cloud M.1 Clusters Run 3x larger workloads with the same low latency and predictable performance—thanks to intelligent data spilling and expanded capacity. Learn more: bit.ly/3L12oH2
bit.ly
Introducing New Materialize Cloud M.1 Clusters
Introducing a new Materialize Cloud cluster type. M.1 Clusters provide customers with more capacity, leading to better economics and performance, while maintaining the same low latency requirements th...
011
Frank McSherry @frankmcsherry.bsky.social · 09/10/2025
Datatoad check-in: this time including some recent progress on columnar joins (good news: faster). Though, it also tries to roll up a bit of the sprawl of content I've scribbled, which increasingly feels like it needs some more careful curation to be helpful. github.com/frankmcsherr...
github.com
130
Frank McSherry @frankmcsherry.bsky.social · 08/10/2025
Good news on the Datalog front: v1 of "columnar joins" seem to work, and resulted in a 20% improvement (from 9.5s to 7.5s, for the joins of a reference workload). Still more gains from tightening it up, and potentially from columnar sorting, but I'll take a swing at writing things up tomorrow!
030
Frank McSherry @frankmcsherry.bsky.social · 30/09/2025
What a difference an allocator makes! This is the same Rust program first using the system allocator, and then using mimalloc. About 100MB of working set in both cases, just .. apparently it pilots the system allocator to some horrible behavior. Obviously going to start using mimalloc from now on.
Runtimes of the same application with two different allocators; mimalloc is nearly 100x faster.
0160
Reposted by Frank McSherry
Sync Conf @syncconf.bsky.social · 19/09/2025
Welcome Frank McSherry @frankmcsherry.bsky.social to Sync Conf 2025. Pioneer of sync technology, inventor of Differential Dataflow, and founder of @materialize.com, Frank will trace the evolution of sync and stream processing.
0103
Reposted by Frank McSherry
Moritz Hoffmann @antiguru.bsky.social · 18/09/2025
Highlighting some of my team's recent work: We've changed Materialize to use swap instead of memory-mapped files, with nice performance and efficiency improvements.
031
Reposted by Frank McSherry
Materialize @materialize.com · 18/09/2025
We’ve released a major improvement to our memory spilling infrastructure: Materialize now uses swap to scale SQL workloads beyond RAM. ✅ Faster hydration ✅ Efficient memory utilization ✅ Bigger workloads supported Full post from antiguru.bsky.social → bit.ly/46EF2iJ
031
Frank McSherry @frankmcsherry.bsky.social · 15/09/2025
Very excited to bring some column-orientation to timely and differential. At least, removing baked in row-orientation in timely, and actual column-orientation in differential, with a bunch of cool learnings from the datatoad work. I hope. We'll see. :D
0101
Reposted by Frank McSherry
Moritz Hoffmann @antiguru.bsky.social · 15/09/2025
We just released Timely Dataflow 0.24! Here are some exciting changes from @frankmcsherry.bsky.social and myself. The container abstractions got a complete rework, and we introduce a new pattern to distribute data. Details below. github.com/TimelyDatafl...
github.com
Release timely-v0.24.0 · TimelyDataflow/timely-dataflow
This version of Timely has some exciting new features. The Distributor trait offers a generalization of the Exchange type. It allows users to define custom distribution strategies for routing data...
151
Frank McSherry @frankmcsherry.bsky.social · 29/08/2025
I have a trip coming up, and I'm hoping to find some content to read about the implementations of (ideally interpreted) array languages. I'm on an interpreter kick, and armed with a bunch of column-oriented libraries. Any tips, drop a reply!
210
Frank McSherry @frankmcsherry.bsky.social · 24/08/2025
I wrote a bit about datatoad's columnar logic for relational operators. At least, for union, intersection, antijoins, and semijoins. It turns out the joins are all easy; it's projection that is hard, of all things. Go figure. github.com/frankmcsherr...
github.com
060
Reposted by Frank McSherry
Martin Kleppmann @martin.kleppmann.com · 19/08/2025
Notion's new offline support is based on our rich text CRDT research x.com/ivanhzhao/st...
x.com
Ivan Zhao on X: "For those of local first nerds and @inkandswitch fans: This is the paper co-authored by @sliminality @geoffreylitt @pvh Martin Kleppmann https://t.co/FMhf4olmg4 Thank you for laying the technical foundation for block-based, rich text CRDT for the world." / X
For those of local first nerds and @inkandswitch fans: This is the paper co-authored by @sliminality @geoffreylitt @pvh Martin Kleppmann https://t.co/FMhf4olmg4 Thank you for laying the technical foundation for block-based, rich text CRDT for the world.
21174
Frank McSherry @frankmcsherry.bsky.social · 20/08/2025
If you are in SF in November, I'll be speaking at syncconf.dev (@syncconf.bsky.social)! It's an excellent confluence of all things up-to-data. Architectures like MZ at the backend, connected via sync engines, and front ends that don't waste anyone's time waiting on database queries.
syncconf.dev
Sync Conf | Nov 12, 2025 in San Francisco.
Sync Conf is a boutique conference on the future of real-time, collaborative, agentic software development. Happening Nov 12, 2025 in San Francisco.
061
Frank McSherry @frankmcsherry.bsky.social · 15/08/2025
In Datalog news: I had given up on getting (compiled) datafrog numbers for the "alias analysis" problem, because it is tedious to write. But thanks to an anonymous benefactor, it was coded up and we can now make a comparison between compiled datafrog and interpreted datatoad, on the same problem!
140
Frank McSherry @frankmcsherry.bsky.social · 13/08/2025
I wrote about the projects done at Materialize’s recent hackathon. Many very cool projects, and also one that I worked on; take a read! materialize.com/blog/spring_...
materialize.com
071
Frank McSherry @frankmcsherry.bsky.social · 08/08/2025
A neat new Materialize post from our QA department on speeding up CI. materialize.com/blog/speedin...
materialize.com
Speeding up Materialize CI
How we slashed CI runtime for Materialize by up to 86% through smarter builds, caching, parallelization, and clever tooling.
020
Frank McSherry @frankmcsherry.bsky.social · 02/08/2025
Datalog weekend: we graduate to queries with cyclic rules, non-binary relations, and generally more interesting behavior. In particular, we're going to compare ourselves against interpreted Soufflé; a standard reference point! How does interpreted datatoad compare? github.com/frankmcsherr...
github.com
040
Frank McSherry @frankmcsherry.bsky.social · 01/08/2025
We have a new @materialize.com post up, this time about pushing selection predicates into our persistence layer. Materialize hybridizes batch and streaming computation (it does both, regularly), and draws on the best optimizations of each (in this case, CDWs). materialize.com/blog/how-fil...
materialize.com
How filter pushdown works
Using part statistics and abstract interpretation to push complex filters all the way down to the storage layer.
170
Reposted by Frank McSherry
Martin Kleppmann @martin.kleppmann.com · 29/07/2025
Nice that the Bluesky firehose is now becoming a live dataset on which to demo streaming databases
2369
Frank McSherry @frankmcsherry.bsky.social · 20/07/2025
We have a weekend Datalog update: datatoad now comes with tries! I took the long route to get here, and have a bit of other work in the pipeline as part of all this. But I'm happy to report that memory went down by another 2x (expected), and runtime went down by ~1.6x. github.com/frankmcsherr...
github.com
140
Frank McSherry @frankmcsherry.bsky.social · 16/07/2025
We have a new blog post up at @materialize.com about analyzing the Bluesky firehose (Jetstream, really) through Materialize. You can grab a copy of the community edition of MZ and follow along, or invent your own ways of looking at the data, live! materialize.com/blog/analyzi...
materialize.com
Analyzing Live Social Data: Exploring Social Trends on Bluesky
Bluesky provides a public firehose that we can stream into Materialize, through which we can observe live social behavior and trends.
0142
Reposted by Frank McSherry
Max Willsey @mwillsey.com · 08/07/2025
I had a great time at Kris's workshop in May! Lots of inspiring talks and discussions. He has posted a mega-video of all the recorded talks, I highly recommend checking some of them out if you're into Datalog, logic programming, or incrementalization. There is even some e-graph stuff in there!
0112
Frank McSherry @frankmcsherry.bsky.social · 09/07/2025
Folks! The workshop at Minnowbrook that prompted the datatoad work has put up the videos from the talk. NINE HOURS! Of logic programming content! I liked the talks; they are chaptered; you can see @mwillsey.com talk about e-graphs; why haven't you clicked yet??? www.youtube.com/watch?v=3ec9...
youtube.com
Minnowbrook Logic Programming Seminar (Supercut w/ Extras)
YouTube video by Kristopher Micinski
060