Sign in

Jack Vanlightly

@vanlightly.bsky.social
3.8K followers 111 following 282 posts

Researcher, advisor, writer, formal verification eng @ Confluent. Everything data (dist sys, databases, messaging, data eng/analytics). jack-vanlightly.com, www.hotds.dev Credit: ESO/B. Tafresh

PostsRepliesMedia
Jack Vanlightly @vanlightly.bsky.social · 10/12/2025
Durable functions have many names across frameworks, but it reduces to 3 forms: stateless functions, sessions, actors. Explainer here: jack-vanlightly.com/blog/2025/12...
jack-vanlightly.com
The Three Durable Function Forms — Jack Vanlightly
Durable execution engines (DEEs) talk about “workflows”, “activities”, “virtual objects”, “handlers”, and “functions”, but they’re often describing the same underlying execution patterns. This post pr...
0130
Jack Vanlightly @vanlightly.bsky.social · 04/12/2025
New post: The Durable Function Tree. Durable execution engines all end up building some form of function tree with suspension points shaped by local vs remote side effects. I look at why, the trade-offs, and where orchestration should (and shouldn’t) be used. jack-vanlightly.com/blog/2025/12...
jack-vanlightly.com
The Durable Function Tree - Part 1 — Jack Vanlightly
In my last post I wrote a bout why and where determinism is needed in durable execution (DE). In this post I'm going to explore how workflows can be formed from trees of durable function calls ba...
0151
Reposted by Jack Vanlightly
Gunnar Morling @gunnarmorling.dev · 25/11/2025
📝 Blogged: "On Idempotency Keys" Discussing several options for ensuring exactly-once processing in distributed systems using idempotency keys, from UUIDs to monotonically increasing sequences. 👉 www.morling.dev/blog/on-idem...
2388
Jack Vanlightly @vanlightly.bsky.social · 24/11/2025
New blog post: Demystifying Determinism in Durable Execution Why do durable execution frameworks care so much about determinism? I unpack the underlying mechanics. Post: jack-vanlightly.com/blog/2025/11...
jack-vanlightly.com
Demystifying Determinism in Durable Execution — Jack Vanlightly
Determinism is a key concept to understand when writing code using durable execution frameworks such as Temporal, Restate, DBOS, and Resonate. If you read the docs you see that some parts of your code...
0110
Jack Vanlightly @vanlightly.bsky.social · 19/11/2025
New blog post about Qbeast and how it brings a multidimensional spatial index to Iceberg/Delta. 🔹 Hypercube-based layout 🔹 Index used by writers, invisible to engines 🔹 Better locality + pruning, adaptive layout Lots of innovation ahead in lakehouses. jack-vanlightly.com/blog/2025/11...
jack-vanlightly.com
Have your Iceberg Cubed, Not Sorted: Meet Qbeast, the OTree Spatial Index — Jack Vanlightly
In today’s post I want to walk through a fascinating indexing technique for data lakehouses which flips the role of the index in open table formats like Apache Iceberg and Delta Lake. We are going to...
150
Jack Vanlightly @vanlightly.bsky.social · 05/11/2025
Stream-order vs batch-order in Iceberg: * Flink wants temporal locality. * Spark wants value locality. Same table, conflicting physics. New post: jack-vanlightly.com/blog/2025/11...
jack-vanlightly.com
How Would You Like Your Iceberg Sir? Stream or Batch Ordered? — Jack Vanlightly
Today I want to talk about stream analytics, batch analytics and Apache Iceberg. Stream and batch analytics work differently but both can be built on top of Iceberg, but due to their differences there...
0132
Jack Vanlightly @vanlightly.bsky.social · 22/10/2025
Three KIPs (1150, 1176, 1183) all target Kafka’s cross-AZ replication costs but there is a wider question at stake. My new post explains the KIPs, the trade-offs between reusing old abstractions vs. embracing stateless compute over S3. jack-vanlightly.com/blog/2025/10...
jack-vanlightly.com
A Fork in the Road: Deciding Kafka’s Diskless Future — Jack Vanlightly
“ The Kafka community is currently seeing an unprecedented situation with three KIPs ( KIP-1150 , KIP-1176 , KIP-1183) simultaneously addressing the same challenge of high replica...
080
Jack Vanlightly @vanlightly.bsky.social · 15/10/2025
New post: why I’m not a fan of “zero-copy” Iceberg tables for Apache Kafka. From a systems design view, it trades storage savings for coupling and complexity. Sometimes, duplication is cheaper than coupling. jack-vanlightly.com/blog/2025/10...
jack-vanlightly.com
Why I’m not a fan of zero-copy Apache Kafka-Apache Iceberg — Jack Vanlightly
Over the past few months, I’ve seen a growing number of posts on social media promoting the idea of a “zero-copy” integration between Apache Kafka and Apache Iceberg. The idea is that Kafka topics cou...
1175
Jack Vanlightly @vanlightly.bsky.social · 08/10/2025
Why don’t Iceberg or Delta Lake have secondary indexes? Because analytics workloads and OLTP workloads optimize for opposite I/O patterns. See my dive into data layout, pruning, and what “indexing” really means in open table formats: jack-vanlightly.com/blog/2025/10...
jack-vanlightly.com
Beyond Indexes: How Open Table Formats Optimize Query Performance — Jack Vanlightly
My career in data started as a SQL Server performance specialist, which meant I was deep into the nuances of indexes, locking and blocking, execution plan analysis and query design. These days I’m mor...
2163
Jack Vanlightly @vanlightly.bsky.social · 02/09/2025
New deep dive: Understanding Apache Fluss I spent August reverse-engineering Fluss, Alibaba’s new table storage engine for Flink (partially forked from Kafka). This post covers its architecture, tiering, and how it tackles changelogs & low-latency state. jack-vanlightly.com/blog/2025/9/...
jack-vanlightly.com
Understanding Apache Fluss — Jack Vanlightly
This is a data system internals blog post. So if you enjoyed my table formats internals blog posts , or writing on Apache Kafka internals or Apache BookKeeper internals , you might enjoy thi...
0162
Jack Vanlightly @vanlightly.bsky.social · 21/08/2025
New blog post: A Conceptual Model for Storage Unification. The post defines what storage unification means, defines terminology and evaluates different building blocks and approaches to doing it. jack-vanlightly.com/blog/2025/8/...
jack-vanlightly.com
A Conceptual Model for Storage Unification — Jack Vanlightly
Object storage is taking over more of the data stack, but low-latency systems still need separate hot-data storage. Storage unification is about presenting these heterogeneous storage systems and form...
072
Jack Vanlightly @vanlightly.bsky.social · 28/07/2025
In a future of autonomous AI agents, we can't limit ourselves to error prevention and error detection, we must also include remediation. jack-vanlightly.com/blog/2025/7/...
jack-vanlightly.com
Remediation: What happens after AI goes wrong? — Jack Vanlightly
If you’re following the world of AI right now, no doubt you saw Jason Lemkin’s post on social media reporting how Replit’s AI deleted his production database , despite it being told not to touch an...
020
Jack Vanlightly @vanlightly.bsky.social · 22/07/2025
Science moves slowly because wrong theories waste decades. Engineering is careful because failures kill people. Software moves fast because mistakes are cheap, the expensive error isn't making the wrong choice, it's taking too long to make any choice. jack-vanlightly.com/blog/2025/7/...
jack-vanlightly.com
The Cost of Being Wrong — Jack Vanlightly
A recent LinkedIn post by Nick Lebesis caught my attention with this brutal take on the difference between good startup founders and coward startup founders. I recommend you read the entire thing ...
040
Jack Vanlightly @vanlightly.bsky.social · 15/07/2025
Where does reliability begin, and where does it end? In distributed business architectures, the answer is responsibility boundaries. New post: jack-vanlightly.com/blog/2025/7/...
jack-vanlightly.com
Responsibility Boundaries in the Coordinated Progress model — Jack Vanlightly
Building on my previous work on the Coordinated Progress model, this post examines how reliable triggers not only initiate work but also establish responsibility boundaries . Where a reliable tri...
0145
Jack Vanlightly @vanlightly.bsky.social · 03/07/2025
ChatGPT thought it was Tuesday, so I made fun of it and it admitted it was Wednesday. So I made fun of it again, and it admitted it was...Wednesday. But sure, AI agents are gonna steal my job 🤔
130
Jack Vanlightly @vanlightly.bsky.social · 24/06/2025
ChatGPT has hallucinated so many times for me today. It's invented scientific terms that don't exist, has been quite liberal with plausible answers based on what sounds reasonable, but without any real world justification. When challenged, it admits it's mistake.
120
Jack Vanlightly @vanlightly.bsky.social · 13/06/2025
My musical evolution continues, discovered deep hypnotic drone music today. No drugs required 😄 The Hypnus Records label is great.
130
Jack Vanlightly @vanlightly.bsky.social · 11/06/2025
How to reliably distribute work across microservices, stream processors, durable execution, event-driven, orchestration and now AI agents? Coordinated Progress is a 4 part series that explores the common structure behind reliable distributed systems. jack-vanlightly.com/blog/2025/6/...
jack-vanlightly.com
Coordinated Progress – Part 1 – Seeing the System: The Graph — Jack Vanlightly
At some point, we’ve all sat in an architecture meeting where someone asks, “ Should this be an event? An RPC? A queue? ”, or “ How do we tie this process together across our microservices? Should it ...
3318
Jack Vanlightly @vanlightly.bsky.social · 09/06/2025
I took a break from social media and my blog for a couple of months. ND burnout. But I'm tentatively back, probably just to post my writing here for now. HOTDS is on pause. Getting back to writing is therapeutic though. I'll post something this week that I've been working on.
070
Jack Vanlightly @vanlightly.bsky.social · 04/04/2025
Another Humans of the Data Sphere is out, with issue 10! In this issue people are talking fsyncs, tips for running ClickHouse at scale, the problems with MCP and more. Plus I dig up a classic paper from 1962. www.hotds.dev/p/humans-of-...
hotds.dev
Humans of the Data Sphere Issue #10 April 4th 2025
Your biweekly dose of insights, observations, commentary and opinions from interesting people from the world of databases, AI, streaming, distributed systems and the data engineering/analytics space.
154
Jack Vanlightly @vanlightly.bsky.social · 03/04/2025
Proud to have contributed formal verification (TLA+) for three key improvements in Kafka 4.0: ✅ KIP-966: Strengthens the replication protocol. ✅ KIP-996: Introduces PreVote for more stable KRaft leadership. ✅ KIP-848: Delivers more efficient, predictable rebalancing.
3161
Jack Vanlightly @vanlightly.bsky.social · 25/03/2025
Wow, I just discovered gamma wave music. Wrote non-stop for three hours.
130
Jack Vanlightly @vanlightly.bsky.social · 21/03/2025
Any Principal Engineers out there with ADHD or creative wiring — who don’t thrive in the tasks of project coordination, alignment meetings, and people management, but thrive on strategy, system design, writing, and shaping direction through ideas? Curious how you navigate the role.
250
Jack Vanlightly @vanlightly.bsky.social · 13/03/2025
A new disaggregated log replication survey post is out. How does the combination of Apache Pulsar with Apache BookKeeper divide and conquer the responsibilities of log replication? jack-vanlightly.com/blog/2025/3/...
jack-vanlightly.com
Log Replication Disaggregation Survey - Apache Pulsar and BookKeeper — Jack Vanlightly
In this latest post of the disaggregated log replication survey, we’re going to look at the Apache BookKeeper Replication Protocol and how it is used by Apache Pulsar to form topic partitions. Raft ...
070
Jack Vanlightly @vanlightly.bsky.social · 12/03/2025
I think I have an issue with tabs, it's grown to 371. My workstation is struggling to open chrome after a restart now.
470
Jack Vanlightly @vanlightly.bsky.social · 11/03/2025
Another Humans of the Data Sphere is out, with issue #9! In this issue, we also look at whether software engineers can learn from mechanical engineering, and looking at table formats as a form of virtualization. www.hotds.dev/p/humans-of-...
hotds.dev
Humans of the Data Sphere Issue #9 March 11th 2025
Your biweekly dose of insights, observations, commentary and opinions from interesting people from the world of databases, AI, streaming, distributed systems and the data engineering/analytics space.
000
Jack Vanlightly @vanlightly.bsky.social · 21/02/2025
A new log replication disaggregation survey post is out! The Kafka Replication Protocol: 🔹Separation of control plane from data plane. 🔹Role separation with minimal coupling. 🔹Kafka’s alignment with Paxos roles. jack-vanlightly.com/blog/2025/2/...
jack-vanlightly.com
Log Replication Disaggregation Survey - Kafka Replication Protocol — Jack Vanlightly
In this post, we’re going to look at the Kafka Replication Protocol and how it separates control plane and data plane responsibilities. It’s worth noting there are other systems that separate concerns...
030
Jack Vanlightly @vanlightly.bsky.social · 19/02/2025
Spotify is so bad at recommendations, but ChatGPT is pretty good at it. I give it a song and it lists the different characteristics of the song, then makes a set of recommendations based on those different characteristics.
110
Jack Vanlightly @vanlightly.bsky.social · 19/02/2025
The first post in the survey of disaggregated log replication systems is out! It looks at Neon's serverless Postgres write-path, which weaves consensus from heterogeneous components, based on MultiPaxos. jack-vanlightly.com/blog/2025/2/...
jack-vanlightly.com
Log Replication Disaggregation Survey - Neon and MultiPaxos — Jack Vanlightly
Over the next series of posts, we'll explore how various real-world systems and some academic papers have implemented log replication with some form of disaggregation. In this first post we’ll look at...
160
Jack Vanlightly @vanlightly.bsky.social · 18/02/2025
I updated my How to Disaggregate a Log Replication Protocol to include "separating ordering from IO". Basically I couldn't ignore CORFU as a way of separating responsibilities! So now we have A-F of ways of breaking apart the monolith. jack-vanlightly.com/blog/2025/2/...
jack-vanlightly.com
How to disaggregate a log replication protocol — Jack Vanlightly
This post continues my series looking at log replication protocols, within the context of state-machine replication (SMR) or just when the log itself is the product (such as Kafka). So far I’ve been l...
360
Jack Vanlightly @vanlightly.bsky.social · 17/02/2025
Table virtualization, stream-table materialization redefine how we think about data sharing, composability, and interoperability. The OTFs (Iceberg, Delta Lake), metadata separation, & cloud storage are paving the way for modular, composable data platforms. jack-vanlightly.com/blog/2025/2/...
jack-vanlightly.com
Towards composable data platforms — Jack Vanlightly
Technology changes can be sudden (like generative AI) or slower juggernauts that kick off a slow chain reaction that takes years to play out. I would place object storage and its enablement of disaggr...
0194
Reposted by Jack Vanlightly
Gunnar Morling @gunnarmorling.dev · 17/02/2025
So many databases claim to be compatible with #Postgres these days, but what does that really mean? @pgdba.bsky.social tries to provide an objective answer with the Postgres Compatibility Index, which verifies compatibility in a wide range of criteria. Loving to see this effort!
0204
Jack Vanlightly @vanlightly.bsky.social · 15/02/2025
The latest Humans of the Data Sphere is out, with issue #8! Additional topics in this post include systems correctness practices at AWS, Datadog's Husky compaction and solving for the distributed case first in ambitious systems projects. www.hotds.dev/p/humans-of-...
hotds.dev
Humans of the Data Sphere Issue #8 February 15th 2025
Your biweekly dose of insights, observations, commentary and opinions from interesting people from the world of databases, AI, streaming, distributed systems and the data engineering/analytics space.
020
Jack Vanlightly @vanlightly.bsky.social · 14/02/2025
This reminds me of Paxos vs Raft. Paxos formalized the responsibilities of reaching consensus and acting on the agreed values into *distinct roles* (proposer, acceptor, learner). These roles can be put in a monolith or distributed. The basis for the Paxos family of protocols are these roles.
060
Jack Vanlightly @vanlightly.bsky.social · 12/02/2025
Confluent + Databricks next-level partnership 💪 Bi-directional flow between Confluent and Databricks. Kafka topics appearing as Delta tables in Databricks. Delta tables appearing as Kafka topics in Confluent. Simply amazing.
085
Jack Vanlightly @vanlightly.bsky.social · 10/02/2025
New dist sys post on log replication! In this one, I classify five approaches to disaggregating log replication protocols—laying the groundwork for a survey of real-world systems through the lens of disaggregation. jack-vanlightly.com/blog/2025/2/...
jack-vanlightly.com
How to disaggregate a log replication protocol — Jack Vanlightly
This post continues my series looking at log replication protocols, within the context of state-machine replication (SMR) or just when the log itself is the product (such as Kafka). So far I’ve been l...
0133
Jack Vanlightly @vanlightly.bsky.social · 06/02/2025
Yesterday, I covered Virtual Consensus—today, let’s dive deeper into a crucial distinction: 👉 Failure-free ordering vs. Fault-tolerant consensus Decoupling these concepts can change how we think about consensus protocols. jack-vanlightly.com/blog/2025/2/...
jack-vanlightly.com
Steady on! Separating Failure-Free Ordering from Fault-Tolerant Consensus — Jack Vanlightly
"True stability results when presumed order and presumed disorder are balanced. A truly stable system expects the unexpected, is prepared to be disrupted, waits to be transformed." — Tom Robbins ...
1111
Jack Vanlightly @vanlightly.bsky.social · 05/02/2025
New distributed systems protocol write-up! This write-up dives into the Virtual Consensus in Delos paper and why it makes sense as the default log replication protocol in the era of object storage and hybrid environments. jack-vanlightly.com/blog/2025/2/...
1225
Jack Vanlightly @vanlightly.bsky.social · 03/02/2025
Speculation is growing that Snowflake is planning to acquire Redpanda—but why? What justifies the rumored high price tag? When you consider market trends, the Snowflake vs. Databricks rivalry, and the AI shift, the rationale starts to become clear. Here’s my take. jack-vanlightly.com/blog/2025/2/...
jack-vanlightly.com
Why Snowflake wants streaming — Jack Vanlightly
Rumors are swirling that Snowflake intends to acquire Redpanda and many are questioning why and what impact this might have on Confluent. First, let’s remember that these are just rumors and there’s...
1155
Jack Vanlightly @vanlightly.bsky.social · 29/01/2025
Humans of the Data Sphere issue #7 is out. Obviously, DeepSeek happened but there have been plenty of other conversations happening. In this issue we also look at well and ill-conditioned APIs and the challenge of contextual data quality in AI-powered data pipelines. www.hotds.dev/p/humans-of-...
hotds.dev
Humans of the Data Sphere Issue #7 January 29th 2025
Your biweekly dose of insights, observations, commentary and opinions from interesting people from the world of databases, AI, streaming, distributed systems and the data engineering/analytics space.
000
Reposted by Jack Vanlightly
Joy Gao @joygao.bsky.social · 29/01/2025
I read this blogpost 3 times, not interested in the investment aspect, but there's so many valuable insights in here on current development of AI. Also the writing alone deserves its own separate praise. youtubetranscriptoptimizer.com/blog/05_the_...
youtubetranscriptoptimizer.com
The Short Case for Nvidia Stock
All the reasons why Nvidia will have a very hard time living up to the currently lofty expectations of the market.
1173
Jack Vanlightly @vanlightly.bsky.social · 28/01/2025
I had RSI in my wrists so bad 20 years ago (too much Counterstrike) that I resorted to typing using pencils with rubbers on the end. I would basically jab the keys and mouse buttons, holding the pencils in my fists. It looked ridiculous but it allowed me to carry on working.
120
Jack Vanlightly @vanlightly.bsky.social · 27/01/2025
AI-related stocks are selling off because DeepSeek’s new model was cheaper to train and has lower inference costs. I suppose investors see that and see a threat to the AI industry. But they’ve got it backwards...
111
Jack Vanlightly @vanlightly.bsky.social · 25/01/2025
Regarding Restate and its distributed log, many people talk about Delos but Apache BookKeeper is also highly relevant/similar, so I like to remind people that it also exists! I've written extensively about how BookKeeper works if you're interested: 1\ medium.com/splunk-maas/...
medium.com
Apache BookKeeper Insights Part 1 — External Consensus and Dynamic Membership
Series Introduction
2123
Jack Vanlightly @vanlightly.bsky.social · 24/01/2025
Great stuff. I'm watching the durable execution space closely and personally I'm quite bullish on it. I'll be writing my own thoughts on durable execution soon.
1233
Jack Vanlightly @vanlightly.bsky.social · 24/01/2025
Over the last week I've been working on finishing the formal verification of Kafka's KRaft with pre-vote and reconfiguration. The protocol is looking good from a design perspective. The team use deterministic simulation testing to catch implementation bugs. Defensive in depth!
130
Jack Vanlightly @vanlightly.bsky.social · 16/01/2025
The pace of AI advancement is not slowing down, and now AI agents are breaking onto the scene, leveraging ever more powerful models to interact with the real world. In this post, I cover some of the voices talking about AI agents and the challenges ahead. jack-vanlightly.com/blog/2025/1/...
jack-vanlightly.com
AI Agents in 2025 — Jack Vanlightly
Two interesting blog posts about AI agents have caught my attention over the last few weeks. Anthropic wrote Building Effective Agents. Chip Huyen wrote Agents . Ethan Mollick has also...
031
Jack Vanlightly @vanlightly.bsky.social · 15/01/2025
My top favorite quotes in issue #6 of HOTDS are: 1) Marc Brooker's blog post: Snapshot Isolation vs Serializability.
180
Jack Vanlightly @vanlightly.bsky.social · 14/01/2025
Another issue of Humans of the Data Sphere is out, with issue #6! In this issue we also dive into the world of AI agents including the promises and challenges ahead. www.hotds.dev/p/humans-of-...
hotds.dev
Humans of the Data Sphere Issue #6 January 14th 2025
Your biweekly dose of insights, observations, commentary and opinions from interesting people from the world of databases, AI, streaming, distributed systems and the data engineering/analytics space.
052
Jack Vanlightly @vanlightly.bsky.social · 13/01/2025
I've not seen much talk of the tragedy of the commons and the potential wave of AGI in the future. The common resource is consumer demand. Companies may replace their humans with AGI for short term gain, not considering the collective long-term impact on consumer demand and therefore the economy.
360