Sign in

Micah Wylde

@micahw.com
465 followers 113 following 55 posts

Co-founder arroyo.dev, building next-gen streaming systems. Prev Splunk, Lyft, Sift, Quantcast.

PostsRepliesMedia
Micah Wylde @micahw.com · 21/01/2026
S2 is incredibly cool, and now you can run it yourself!
092
Reposted by Micah Wylde
Craig Dennis @craigsdennis.dev · 21/01/2026
I built a tool called LinkedOut to solve the "links in the air" problem during my talks. It’s a serverless data lake built on the @cloudflare.social Data Platform. Ingest: Pipelines Store: R2 + Apache Iceberg Security: Access Real-time analytics with zero egress fees. Full build video below.
231
Micah Wylde @micahw.com · 13/10/2025
It’s so dumb and also so good, just a big dude doing good in the world
050
Micah Wylde @micahw.com · 03/10/2025
That’s correct, the r2 sql project predated the arroyo acquisition, but we’ll be converging over time. Also honored to be mentioned in the same post as the possible dbt acquisition! Ours was…a bit smaller.
020
Reposted by Micah Wylde
Jeremy Morrell @jeremymorrell.dev · 27/09/2025
Another brand new new feature is the R2 data catalog: blog.cloudflare.com/cloudflare-d... Build something with Pipelines and R2 SQL. I suggest receiving OpenTelemetry data and then surfacing that in a web app (logs should be fairly straightforward), but there are tons of uses for this.
blog.cloudflare.com
Announcing the Cloudflare Data Platform: ingest, store, and query your data directly on Cloudflare
The Cloudflare Data Platform, launching today, is a fully-managed suite of products for ingesting, transforming, storing, and querying analytical data, built on Apache Iceberg and R2 storage.
132
Reposted by Micah Wylde
Jeremy Morrell @jeremymorrell.dev · 25/09/2025
It's early, but I'm excited about direction that the Cloudflare Data Platform is taking. Trying to set up similar pipelines on other clouds would typically be $$$ and take tons of expertise. Managing kafka and multiple services for ingestion, compaction, etc blog.cloudflare.com/cloudflare-d...
blog.cloudflare.com
Announcing the Cloudflare Data Platform: ingest, store, and query your data directly on Cloudflare
The Cloudflare Data Platform, launching today, is a fully-managed suite of products for ingesting, transforming, storing, and querying analytical data, built on Apache Iceberg and R2 storage.
381
Micah Wylde @micahw.com · 25/09/2025
The news is finally out! Cloudflare has a Data Platform! We're starting with serverless streaming pipelines (powered by arroyo), a managed Iceberg Catalog, and a new distributed SQL engine built on top of DataFusion
0120
Micah Wylde @micahw.com · 02/08/2025
Sequin (sequinstream.com) is doing this, but focused on Postgres
sequinstream.com
Sequin
Postgres change data capture (CDC) to Kafka, SQS, webhooks, and more. Build real-time data replication, event workflows, and audit logging. Fast and easy setup.
140
Micah Wylde @micahw.com · 25/06/2025
Oh wow, you’re totally right. Saw the repo but didn’t actually look inside.
100
Micah Wylde @micahw.com · 25/06/2025
Firebolt at least is source-available: github.com/firebolt-db/...
github.com
GitHub - firebolt-db/firebolt-core: Firebolt Core is a free, self-hosted edition of Firebolt's distributed query engine (https://www.firebolt.io/); it provides high-performance data warehousing capabi...
Firebolt Core is a free, self-hosted edition of Firebolt's distributed query engine (https://www.firebolt.io/); it provides high-performance data warehousing capabilities that can be deployed a...
100
Reposted by Micah Wylde
Andrew Lamb @andrewlamb1111.bsky.social · 09/06/2025
Reminder: San Francisco @ApacheDataFusio meetup tomorrow: lu.ma/uuxd443e
lu.ma
SF Apache DataFusion Meetup · Luma
Join us for an evening of learning, networking, and diving into Apache DataFusion, the blazing-fast query execution framework for Rust-based data…
031
Reposted by Micah Wylde
Cloudflare @cloudflare.social · 04/06/2025
Cloudflare is at Snowflake Summit in San Francisco this week! Swing by our booth 2605 to chat about the new Cloudflare R2 Data Catalog and how it can make your data management and analytics easier!
082
Micah Wylde @micahw.com · 01/06/2025
I want that! We have completely separate code for object store and local filesystem, even though the latter is really only used for testing and dev.
010
Micah Wylde @micahw.com · 30/05/2025
Absolutely! Parquet and iceberg support are coming, and we’ll consider Ducklake support if it starts getting traction.
140
Micah Wylde @micahw.com · 27/05/2025
Next Monday after the Snowflake Summit keynote! Hang out on our beautiful roof with other cool data folks, and hear some great speakers from LanceDB, @mooncakelabs.bsky.social, Eventual, Marimo, Bobsled, and @cloudflare-dev.bsky.social! lu.ma/dbq1hfij
lu.ma
Modern Data w/ Cloudflare + Friends · Luma
Come talk about modern data formats, streaming ingestion, query engines and how you feel about Iceberg at Cloudflare's HQ. We'll be running a series of…
010
Micah Wylde @micahw.com · 03/05/2025
Better CI is a tough business… Earthly couldn’t make it work despite building a great product earthly.dev/blog/shuttin...
earthly.dev
A message about Earthly
In the next three months, we will be phasing out our Earthly Satellite commercial services, including the Earthly Cloud Satellites, Self-Hosted Sat...
140
Reposted by Micah Wylde
Chris @chris.blue · 18/04/2025
Ok, y'all. This took me several weeks and a ton of help from @frankmcsherry.bsky.social and @lalithsuresh.bsky.social. I dug into timely dataflow, differential dataflow, and DBSP to get you up to speed on IVM engines and materialized views. Enjoy!
materializedview.io
Everything You Need to Know About Incremental View Maintenance
An overview of incremental view maintenance, why it’s useful, and how you can implement it.
47618
Micah Wylde @micahw.com · 16/04/2025
I’m only a week into life at @cloudflare-dev.bsky.social but already amazed by how much of Cloudflare is built _on_ Cloudflare. I’d never have guessed you could get so far with just workers + durable objects!
000
Micah Wylde @micahw.com · 10/04/2025
Arroyo is joining @cloudflare.social! We're bringing Arroyo to the Developer Platform as a serverless stream processing system, and will also remain open-source and self-hostable. www.arroyo.dev/blog/arroyo-...
arroyo.dev
Arroyo is joining Cloudflare
Arroyo has been acquired by Cloudflare to bring serverless SQL stream processing to the Cloudflare Developer Platfrorm, integrated with Queues, Workers, and R2. The Arroyo Engine will remain open-sour...
2163
Reposted by Micah Wylde
rmoff 🏃‍♂️🫖🥓 @rmoff.net · 10/04/2025
Couple of big announcements from @cloudflare.social today for folk in #dataBS: * Acquisition of Arroyo, launch of Pipelines for streaming ingestion: blog.cloudflare.com/cloudflare-a... * Launch of R2 Data Catalog—a managed Apache Iceberg catalog for R2 blog.cloudflare.com/r2-data-cata...
blog.cloudflare.com
Just landed: streaming ingestion on Cloudflare with Arroyo and Pipelines
We’ve just shipped our new streaming ingestion service, Pipelines — and we’ve acquired Arroyo, enabling us to bring new SQL-based, stateful transformations to Pipelines and R2.
093
Micah Wylde @micahw.com · 26/03/2025
Arroyo 0.14.0 is now available, including new lookup joins, support for nested updating aggregates, struct types, new syntax, and a bunch of improvements and fixes: www.arroyo.dev/blog/arroyo-...
arroyo.dev
Announcing Arroyo 0.14.0
Arroyo 0.14 is now available! This release introduces support for lookup joins, more powerful updating SQL, new syntax, structs in DDL, and more!
010
Micah Wylde @micahw.com · 24/03/2025
I know by month 2 we're all inured to this stuff, but this is a beyond crazy mix of incompetence and illegality www.theatlantic.com/politics/arc...
theatlantic.com
The Trump Administration Accidentally Texted Me Its War Plans
U.S. national-security leaders included me in a group chat about upcoming military strikes in Yemen. I didn’t think it could be real. Then the bombs started falling.
000
Micah Wylde @micahw.com · 19/03/2025
SCO didn’t really turn evil, they were bought by Caldera which rebranded to SCO
000
Micah Wylde @micahw.com · 17/03/2025
With checkpoints slatedb is basically a streaming state backend in a box. Wish this had already existed when we started arroyo!
061
Micah Wylde @micahw.com · 01/03/2025
Amazing!
030
Micah Wylde @micahw.com · 01/03/2025
Arroyo is sitting at 3,999 stars... who's going to put us over the top github.com/ArroyoSystem...
130
Micah Wylde @micahw.com · 27/02/2025
I'll use a Python Jupyter notebook with DuckDB. You can convert results to a pandas dataframe then plot with matplotlib. ChatGPT is very good at writing the gluey Python bits.
020
Micah Wylde @micahw.com · 25/02/2025
You'd think that the key to being a fast streaming engine is like clever join algorithms, but it's mostly just being really good at JSON. Arroyo uses Arrow and the arrow-rs JSON decoder along with some streaming extensions. I think it's pretty cool, so I wrote up a long explanation of how it works
arroyo.dev
Fast columnar JSON decoding with arrow-rs
JSON is the most common serialization format used in streaming pipelines, so it pays to be able to deserialize it fast. This post covers in detail how the arrow-json library works to perform very effi...
0121
Micah Wylde @micahw.com · 23/01/2025
It combines a bunch of great services and tools to provide sub-minute-latency querying at a very low cost, including * Redpanda serverless (Log storage) * S3 (Object storage) * Arroyo * DuckDB It went so well it felt worth documenting the process for other folks
010
Micah Wylde @micahw.com · 23/01/2025
Our team at Arroyo recently needed to rebuild our (very ad-hoc) analytics infra to account for our growth. We spent some time working out the best way to set up a near-real-time data lake today, and ended up with a pretty sweet approach we're calling the LOAD stack: www.arroyo.dev/blog/buildin...
arroyo.dev
Building a near-real-time data lake with the LOAD stack
The LOAD stack (log storage/object storage/Arroyo/DuckDB) makes it easy to build an affordable real-time data lake with minimal operational overhead. This tutorial will guide you through the process o...
172
Micah Wylde @micahw.com · 19/12/2024
I'm particularly excited that this release has our most-ever community contributors, including major features like a RabbitMQ connector and support for Kafka IAM auth among many other improvements and fixes
000
Micah Wylde @micahw.com · 19/12/2024
Arroyo 0.13.0 is now available! This one includes some big improvements to the core engine a (including the operator chaining work I wrote about previously: bsky.app/profile/mica...) and a bunch of other features. All the details on our blog: www.arroyo.dev/blog/arroyo-...
arroyo.dev
Announcing Arroyo 0.13.0
Arroyo 0.13 is now available! This release introduces support for reading source metadata, a RabbitMQ connector, improved CDC support, operator chaining, along with many other improvements.
140
Micah Wylde @micahw.com · 19/12/2024
Is this a path to better DataFusion error messages?
100
Micah Wylde @micahw.com · 05/12/2024
20/ This is long enough already—but hopefully this was interesting for someone! If you want to read more about the architecture of Arroyo, these blog posts give a good overview: www.arroyo.dev/blog/why-arr... and www.arroyo.dev/blog/statefu...
arroyo.dev
We built a new SQL Engine on Arrow and DataFusion
Arroyo 0.10 has an entirely new SQL engine built with Apache Arrow and DataFusion. It's much faster, smaller, and easier to run. Read on for why and how we're making this change.
020
Micah Wylde @micahw.com · 05/12/2024
19/ These need to be properly interleaved with the data events that might be produced as a result of handling a watermark or checkpoint barrier (for example, a tumbling window will emit data when it receives a watermark that tells it a window has closed).
110
Micah Wylde @micahw.com · 05/12/2024
18/ This gets a bit more complicated when handling non-data messages in the dataflow. In a stateful, timely system like Arroyo (or Flink), we also need to pass down messages representing checkpoint barriers and watermarks, which propagate throughout the system.
100
Micah Wylde @micahw.com · 05/12/2024
17/ The other contains a reference to the next operator in the chain. This ChainedCollector’s collect method immediately calls the handle method of that next operator, with a collector that contains the following operator in the chain (if there is one) or a QueueCollector if this is the last.
100
Micah Wylde @micahw.com · 05/12/2024
16/ There are a few approaches. In Arroyo, we use one based on function calling. Each operator takes in a collector trait, which provides a collect method that produces an Arrow RecordBatch downstream. There are two implementations of this trait. One just produces to a queue to be read by the next
100
Micah Wylde @micahw.com · 05/12/2024
15/ So that’s the why and the when. What about the how? If we have operators that read events off of a queue and push their results back to a queue, how do we turn them into chained operators? And how can they work in both chained and unchained modes?
100
Micah Wylde @micahw.com · 05/12/2024
14/ This can have a huge effect on the number of tasks in a large complex pipeline. In a simple example pipeline, this reduces the number of operators from 10 to 3. Multiply that by the parallelism to give the total amount of task reduction.
100
Micah Wylde @micahw.com · 05/12/2024
13/ The previous operator must have a single output, and the next operator must also have just a single input (i.e., isn’t a join or a union). If all these conditions are met, we can chain the two operators together and eliminate a task and a queue.
110
Micah Wylde @micahw.com · 05/12/2024
12/ So what can we chain? Two operators are chainable if they’re connected by a non-shuffle edge (called a forward) in which each subtask of an operator is connected only to the same subtask in the next operator. Both operators must also have the same parallelism.
100
Micah Wylde @micahw.com · 05/12/2024
11/ And many systems will statically provision memory to back the queues, so removing them can reduce memory usage. Finally, each of the parallel subtasks of the operators have to be managed by the engine’s control plane so there’s benefit to reducing the total number of tasks.
100
Micah Wylde @micahw.com · 05/12/2024
10/ There are a few reasons. In some engines like Flink, records sometimes have to be serialized or copied onto the queues, so removing the queue avoids that cost. The writing and reading to the queues also adds some overhead.
100
Micah Wylde @micahw.com · 05/12/2024
9/ Ok, so back to operator chaining. The basic idea is to remove some of the queues between operators (effectively merging them into a single operator), getting closer to the push-based analytical engine model. Why would we want to do this?
110
Micah Wylde @micahw.com · 05/12/2024
8/ For example, a group-by query will need to get all of the records for a particular key on the same subtask of an operator so that they can be processed together. We call this a shuffle edge.
110
Micah Wylde @micahw.com · 05/12/2024
7/ The engine may also be distributed, which means that the queue may be read by a network stack that’s sending records across a TCP socket to another worker. Finally, we may need to repartition the data on an edge across all of the parallel subtasks of an operator.
110
Micah Wylde @micahw.com · 05/12/2024
6/ But there’s another difference: in a batch engine like DuckDB, the records are directly pushed from one operator into another. But for streaming engines we generally have a queue between operators to provide backpressure and asynchronicity.
120
Micah Wylde @micahw.com · 05/12/2024
5/ Newer analytical systems use a push model instead of pull, which provides better cache efficiency and easier support for plans that aren’t trees (e.g., DAGs with diamonds). In streaming engines push models are (afaik) universally used.
120
Micah Wylde @micahw.com · 05/12/2024
4/ This logical view is basically the same in any query engine, but the physical structure of the dataflow differs. The traditional approach for batch is the volcano model, in which downstream operators pull from their upstream (which in turn pulls from its upstream) until there is no work left.
110