Yaroslav Tkachenko @sap1ens.com · 8hI'll be speaking at the Data Streaming Summit next week. Covering one of my favourite topics: columnar data for streaming workloads. Interested in attending? Use the DSSBUDDY26 discount code to get a free ticket! 000
Yaroslav Tkachenko @sap1ens.com · 21hI currently have some additional consulting capacity and am looking to work with a few new teams. 001
Yaroslav Tkachenko @sap1ens.com · 21hWant to make your streaming pipelines more efficient? I help companies make data streaming systems faster, cheaper, and more reliable - whether that means finding bottlenecks, reducing infrastructure spend, or rethinking part of the architecture. Work with me: sap1ens.com/consulting/sap1ens.comConsulting | sap1ens.comConsulting 110
Yaroslav Tkachenko @sap1ens.com · 29/09/2026Most teams that operate large-scale data streaming pipelines start to care about cost in one way or another. Streaming is perceived as "expensive", and it's frequently questioned. Today I'm attempting to address some concerns: www.streamingdata.tech/p/improving-...streamingdata.techImproving Cost Efficiency of Data Streaming PipelinesLearn practical ways to cut data streaming costs in Kafka and Flink by optimizing batching, serialization, state, network traffic, and capacity. 020
Yaroslav Tkachenko @sap1ens.com · 25/09/2026However, instead of just focusing on lakehouses, they also support databases like ClickHouse and OpenSearch! I like this new take on connectors. 000
Yaroslav Tkachenko @sap1ens.com · 25/09/2026The idea is to use object storage: "Because compaction works directly against object storage, materialization adds no load to the brokers serving producers and consumers. It requires no consumer groups and no separate connector cluster, and its compute scales independently from the brokers." 100
Yaroslav Tkachenko @sap1ens.com · 25/09/2026Streamhouse is old news, StreamNative has just released Open Lakestream. Lakestream is their take on materializing streaming data. streamnative.io/blog/introdu...streamnative.ioIntroducing the Stream Materialization Framework: From Streams to Governed Data AssetsLakestream's Stream Materialization Framework turns a stream into governed Iceberg, ClickHouse, or OpenSearch tables without a separate connector. 140
Yaroslav Tkachenko @sap1ens.com · 17/09/2026I've had a lot of time to reflect recently, so I decided to capture some thoughts on how my professional goals and values have evolved since becoming a full-time solopreneur. www.independentengineer.co/p/evolving-g...independentengineer.coEvolving Goals and ValuesEverything changes, including things that matter to you. 040
Yaroslav Tkachenko @sap1ens.com · 16/09/2026Just found out that Current is free this year! current.confluent.io/san-franciscocurrent.confluent.ioSan FranciscoCurrent is the biggest event in data streaming. Over two days in San Francisco, developers, data practitioners, and tech executives will come together to shape the future of Apache Kafka, Apache Flink... 000
Yaroslav Tkachenko @sap1ens.com · 14/09/2026Honestly? Too much stuff. Looks impressive, but kinda hard to read. 4 different font styles :) My north star: jonathanstark.com/ps 110
Yaroslav Tkachenko @sap1ens.com · 11/08/2026I gave a talk about building Streamling last week at the DataFusion Community Showcase. The recording is now available! youtu.be/0-BIHyzODH8?...youtu.beDataFusion Community Showcase Vol. 3: ASAPQuery & StreamlingYouTube video by Spice AI 020
Reposted by Yaroslav TkachenkoAndrew Lamb @andrewlamb1111.bsky.social · 10/08/2026"DataFusion and Arrow allowed us to build in a few months rather than a few years" www.youtube.com/watch?v=0-BI... - Yaroslav Tkachenko from Streamling on the Community Showcase. Milind Srivastava from CMU also showed off the ASAPQuery data sketch query system.youtube.comDataFusion Community Showcase Vol. 3: ASAPQuery & StreamlingYouTube video by Spice AI 082
Yaroslav Tkachenko @sap1ens.com · 28/07/2026The New Wave of the Streaming Log Technologies www.streamingdata.tech/p/the-new-wa...streamingdata.techThe New Wave of the Streaming Log TechnologiesStreaming logs are entering a third wave: Kafka-free systems rethink performance, simplicity, routing, protocols, and object-native design. 040
Reposted by Yaroslav TkachenkoAndy Pavlo @andypavlo.bsky.social · 22/07/2026Stonebraker is not happy with existing Text2SQL tooling: cacm.acm.org/blogcacm/if-...cacm.acm.orgIf You Think You Can Do Real-World Text-to-SQL – Communications of the ACM 2316
Yaroslav Tkachenko @sap1ens.com · 22/07/2026The end-to-end pipeline takes less than 200 lines of YAML. www.streamling.dev/tutorials/ge...streamling.devGetting Slack Notifications From Kafka Events — StreamlingCreate a Streamling pipeline to consume data from Kafka, filter it with the SQL transform, prepare a Slack webhook payload with the TypeScript transform and deliver it to the Slack webhook endpoint wi... 010
Yaroslav Tkachenko @sap1ens.com · 22/07/2026I wrote a tutorial on using Streamling to send rich Slack notifications based on Kafka events. The tutorial highlights how you can easily combine SQL and TypeScript transformations to filter data and get it precisely into the shape you need. 130
Reposted by Yaroslav Tkachenkos2.dev @s2.dev · 20/07/2026S2 source and sink plugins are now available for Streamling. Read durable S2 streams into Arrow batches, or write Streamling output back to S2. ClickHouse ingestion, Postgres CDC, and a full Postgres → S2 → ClickHouse example: s2.dev/docs/integr... 042
Yaroslav Tkachenko @sap1ens.com · 14/07/2026ClickHouse Cloud has the best onboarding experience. 2-3 clicks, done. 010
Yaroslav Tkachenko @sap1ens.com · 04/07/2026Recently someone said "load-bearing" to me with a straight face... 030
Yaroslav Tkachenko @sap1ens.com · 25/06/2026Recording: www.youtube.com/watch?v=3oN_...youtube.comStreamling: Building a Lightweight Data Streaming Framework on Apache DataFusion, Yaroslav TkachenkoYouTube video by vancouver systems 010
Yaroslav Tkachenko @sap1ens.com · 25/06/2026One of my goals was to educate people about the extensibility of Apache DataFusion. It's a truly marvellous piece of tech that can be used as a foundation of a database, query engine or a streaming system. 100
Yaroslav Tkachenko @sap1ens.com · 25/06/2026Earlier this week, I presented a talk at the Vancouver Systems Meetup: "Streamling: Building a Lightweight Data Streaming Runtime on Apache DataFusion". 120
Yaroslav Tkachenko @sap1ens.com · 09/06/2026www.streamingdata.tech/p/introducin...streamingdata.techIntroducing Streamling: Performant and Extensible Data Streaming FrameworkStreamling is an open-source Rust, Arrow, and DataFusion framework for fast, extensible transactional data pipelines. 010
Yaroslav Tkachenko @sap1ens.com · 09/06/2026It's designed for application developers (not data engineers): it supports writing TypeScript transformations (via WebAssembly) and integrates well with databases and HTTP APIs. It's extremely efficient: you can run a decent data-processing workload on just 0.5 CPU cores. 130
Yaroslav Tkachenko @sap1ens.com · 09/06/2026I'm excited to announce Streamling: lightweight, performant and extensible data streaming runtime. Streamling is built using Rust, Apache Arrow and Apache DataFusion. 130
Yaroslav Tkachenko @sap1ens.com · 05/06/2026Bytewax is on life support: github.com/bytewax/byte...github.comState of the project? · Issue #560 · bytewax/bytewaxHello Bytewax Team and Community, First off, thank you so much for all your hard work on Bytewax! It's a fantastic, open-source project, and I truly appreciate the effort that goes into maintaining... 220
Yaroslav Tkachenko @sap1ens.com · 20/05/2026Found a love letter to @andypavlo.bsky.social at #PgConf in Vancouver 0201
Yaroslav Tkachenko @sap1ens.com · 13/05/2026Fingerprinting comes with too many edge cases. E.g. no way to handle deletes. 000
Yaroslav Tkachenko @sap1ens.com · 13/05/2026streamacademy.io/tutorial/fli...streamacademy.ioApache Flink: Postgres to Postgres Replication with Flink CDCA practical guide to building a generic, multi-table Postgres-to-Postgres replication pipeline with Flink CDC - one replication slot for any number of tables, automatic schema discovery, and a custom ... 010
Yaroslav Tkachenko @sap1ens.com · 11/05/2026www.streamingdata.tech/p/can-kafka-...streamingdata.techCan Kafka Queues Make Consumers Faster? Part 2: Head-Of-Line BlockingKafka Queues can scale consumers beyond partitions by reducing head-of-line blocking when message processing has I/O delays. 020
Reposted by Yaroslav Tkachenkormoff 🏃♂️🫖🥓 @rmoff.net · 30/04/2026Just in the nick of time, it's my round-up of Interesting Links for April 2026! There's links about all sorts of stuff around Data Engineering, CDC, Iceberg, Data architectures, and even a little hint of AI (but only the useful kind, not the gross hype-y kind). rmoff.net/2026/04/30/i...rmoff.netInteresting links - April 2026 041
Yaroslav Tkachenko @sap1ens.com · 27/04/2026www.streamingdata.tech/p/can-kafka-...streamingdata.techCan Kafka Queues Make Consumers Faster?Kafka Queues promise consumer scaling beyond partitions, but benchmarks show share consumers still lag far behind standard Kafka consumers. 010
Yaroslav Tkachenko @sap1ens.com · 26/04/2026www.independentengineer.co/p/how-to-gro...independentengineer.coHow to Grow Your NetworkLearn how software engineers can build a valuable network before going independent by helping others, nurturing work relationships, attending events, and joining communities. 020
Yaroslav Tkachenko @sap1ens.com · 20/04/2026Sharing some thoughts this morning about different paths you can take as a software engineer who's looking for more independence, and perhaps, eventually becoming a solopreneur. It's not just about building a product or becoming a consultant. www.independentengineer.co/p/software-e...independentengineer.coThe Many Ways Software Engineers Can Go IndependentHow software engineers can go independent: consulting, productized services, SaaS, courses, newsletters, and other realistic solopreneur paths. 010
Yaroslav Tkachenko @sap1ens.com · 17/04/2026Data Streaming Academy has tutorials now! The first tutorial covers reading and writing Kafka consumer offsets in Apache Flink using the State Processor API. streamacademy.io/tutorial/fli... 010
Yaroslav Tkachenko @sap1ens.com · 16/04/2026Recently, I spent a lot of time benchmarking various Iceberg sinks (Spark, Flink, Kafka Connect) and trying to beat @supermetal-inc.bsky.social for moving data from Postgres to Iceberg. Here's my report: thenewstack.io/postgres-ice...thenewstack.ioPostgres to Iceberg in 13 minutes: How Supermetal compares to Flink, Kafka Connect, and SparkSupermetal's new Iceberg sink snapshots Postgres data in 13 minutes versus 90+ for Flink, Kafka Connect, and Spark in this CDC benchmark. 020
Yaroslav Tkachenko @sap1ens.com · 15/04/2026Is this peak bubble? techcrunch.com/2026/04/15/a...techcrunch.comAfter sale of its shoe business, Allbirds pivots to AI | TechCrunchAllbirds is ditching wool sneakers for AI servers, rebranding as NewBird AI after locking in a $50M convertible financing facility. 000