Sign in

Almog

@almog.xyz
2.8K followers 234 following 145 posts

Favorite Buzzwords: Kafka Streams | SlateDB | Stream Processing | Distributed Databases co-founder @ responsive.dev

PostsRepliesMedia
Almog @almog.xyz · 23/04/2026
A deep dive to how metrics are stored and queried. Inverted indexes, roaring bitmaps, and ... gorillas? www.bitsxpages.com/p/how-metric...
bitsxpages.com
how metrics are stored and queried
the data structures powering your metrics dashboard
070
Almog @almog.xyz · 16/04/2026
The team worked insane hours to get this release ready and we've finally shipped it! We built a first of its kind, prometheus-compatible database that is object native. ~22x cheaper than AWS AMP and stupid easy to operate. www.opendata.dev/blog/introdu...
opendata.dev
OpenData Timeseries: Prometheus-compatible metrics on object storage | OpenData
An MIT-licensed, Prometheus-compatible timeseries database built on SlateDB, bringing the operating model and cost structure of object-store-native systems to the Prometheus and Grafana ecosystem.
2195
Almog @almog.xyz · 27/03/2026
I know "measure twice, cut once" applies to performance optimizations, but where's the fun in that? I am quite proud of my lopsided planter boxes and custom memory allocators.
020
Almog @almog.xyz · 23/03/2026
Databases have strange economic incentives that end up making them "worse" over time. The problem is an asymmetry between who benefits from innovations at first (database companies) and who benefits in the long run (hyperscalers that sell hardware). Blogged: www.bitsxpages.com/p/the-broken...
bitsxpages.com
the broken economics of databases
Why database companies charge so much, earn so little, and keep making things complicated
181
Almog @almog.xyz · 17/03/2026
I regularly get asked "how do you create your diagrams for your blogs"? Recently I've been using my own tool, and I'm open sourcing it! Yuzudraw is a visual editor for ASCII that has a token-efficient DSL that lets agents generate/modify your diagrams. Give it a ⭐! github.com/agavra/yuzud...
github.com
GitHub - agavra/yuzudraw: ASCII diagramming tool for both humans and agents
ASCII diagramming tool for both humans and agents. Contribute to agavra/yuzudraw development by creating an account on GitHub.
030
Almog @almog.xyz · 02/03/2026
“Gradually, then suddenly.” That’s how adoption works when you’re building something new. Opendata is still "gradually" but 100 stars with $0 spent on marketing is a good start. Back to building! github.com/opendata-oss...
091
Almog @almog.xyz · 10/02/2026
I implemented prefix compression for SlateDB & noticed benchmarks looked "worse". Fell down a rabbit hole. Turns out I was thinking about compression backwards... Wrote up my learning: www.bitsxpages.com/p/the-mathem...
bitsxpages.com
the mathematics of compression in database systems
why compression is (almost) always worthwhile
1101
Almog @almog.xyz · 28/01/2026
Can you beat 180KB? I created a challenge to reduce a dataset as much as possible 🏆 my approach uses delta encoding & prefix compression + zstd(22) compression to reduce 25MB -> 180KB github.com/agavra/bit-g...
github.com
GitHub - agavra/bit-golf: a compression golf challenge for GitHub event data
a compression golf challenge for GitHub event data - agavra/bit-golf
082
Almog @almog.xyz · 09/01/2026
I fixed the problem with reviewing code written by Claude in "accept edits on" mode: github.com/agavra/tuicr would love to know what you think, and if you want to contribute to an OSS rust project there's a bunch of open issues to pick up!
github.com
GitHub - agavra/tuicr: Review AI-generated diffs like a GitHub pull request, right from your terminal.
Review AI-generated diffs like a GitHub pull request, right from your terminal. - agavra/tuicr
120
Almog @almog.xyz · 05/01/2026
I'll die on this hill: Sorted String Tables (SSTs) are the single most important data structure for modern DBs. They lean in to the limitations of SSDs and Object Storage, making them (and similar layouts) the best choice for many databases. Blogged in detail: www.bitsxpages.com/p/sorted-str...
bitsxpages.com
sorted string tables (SST) from first principles
why sorted string tables are the swiss army knife for data systems and how they are implemented
060
Almog @almog.xyz · 09/12/2025
The inner join between sets of people who build databases, write, and draw? Low cardinality. I'm in that set, so I'm starting a blog! Here's my first post: www.bitsxpages.com/p/frameworks...
bitsxpages.com
frameworks for understanding databases
building mental models for tradeoffs in performance, availability and durability in data systems
2163
Almog @almog.xyz · 25/11/2025
Sometimes the best solution is "do nothing", but it's always more fun to play with tools.
040
Almog @almog.xyz · 21/10/2025
Calling database nerds in SF! I'm covering SlateDB at the systems meetup next Wednesday (10/29). If you're around, I'd love to meet you in person (that way you'll have proof I'm not just an AI bot). 👉 luma.com/e7feg2i6
0153
Almog @almog.xyz · 08/08/2025
Beans on toast is underrated.
210
Almog @almog.xyz · 07/08/2025
I recently implemented Gorilla encoding (www.vldb.org/pvldb/vol8/p...) for a SlateDB PR. Pretty cool stuff - easy to understand but really powerful. Here it is, explained by a gorilla.
091
Almog @almog.xyz · 05/08/2025
It's true. Every company eventually becomes a database company.
0101
Almog @almog.xyz · 15/07/2025
Despite using so many new technologies, I somehow never learn my lesson: read the docs sooner and read the docs thoroughly.
092
Almog @almog.xyz · 08/07/2025
Reading the tokio-rs async documentation makes me feel like...
161
Almog @almog.xyz · 01/07/2025
One day I'll open a coffee shop dedicated to the not-insignificant intersection between database nerds and coffee snobs. Until then, enjoy this comic.
1395
Almog @almog.xyz · 28/04/2025
Maybe... just maybe, adding more features and complexity into stream processors is NOT what we need?
061
Almog @almog.xyz · 24/04/2025
The new electric Caltrain cars have WiFi. 🙏 SF bay area has finally entered the 21st century (on this dimension of public transit only).
010
Reposted by Almog
Chris @chris.blue · 22/04/2025
Today marks SlateDB’s one year anniversary! It’s been a lot of fun. Thanks to @rohanpd.bsky.social @flaneur2024.bsky.social @almog.ai @vigneshc.bsky.social @paulbutler.org Jason Gustafson, David Moravek, and many others for joining the project. 😀
slatedb.io
SlateDB - An embedded storage engine built on object storage | SlateDB
Description will go into a meta tag in <head />
0165
Almog @almog.xyz · 21/04/2025
Prompt engineering was v0. Context engineering is v1.0. How should you think about supplying LLMs with the right information? 📖 [1/N]
140
Almog @almog.xyz · 10/02/2025
"Kafka Configs: A Rube Goldberg Machine?" - an actual quote from a customer after discussing cleanup.policy and Kafka Streams changelogs.
030
Reposted by Almog
Chris @chris.blue · 03/02/2025
SlateDB now has clones 🤯 Users can clone an existing DB's data to a new location. It's nearly instantaneous since it references the data from the old bucket rather than copying. Writes to the clone update the new location. Compaction lazily merges old data into the new directory./ht @responsive.dev
github.com
Initial clone implementation by hachikuji · Pull Request #430 · slatedb/slatedb
Fixes #315. This patch contains the logic to create and initialize clones as outlined in RFC-0004.
1332
Reposted by Almog
Apurva Mehta @apurvamehta.com · 04/02/2025
Should you use Kafka? If so, when? And what are the tradeoffs presented by the dizzying variety of Kafka-adjacent technologies? I hope that my latest blog post provides unique and useful answers to these questions. Let me know what you all think! 👉 www.responsive.dev/blog/why-whe...
0155
Almog @almog.xyz · 30/01/2025
As part of my post-Elden Ring hunt for a game, I tried Balatro last night. I'm starting to think I may have made a mistake.
130
Reposted by Almog
ScyllaDB @scylladb.com · 29/01/2025
How did Responsive simplify its architecture and achieve monstrously high availability with unlimited scale potential? Co-founder @almog.ai will share how replacing RocksDB with #ScyllaDB in #Kafka Streams led to big wins for their team at Monster Scale Summit. www.scylladb.com/monster-scal...
084
Almog @almog.xyz · 29/01/2025
My hobbies all have the ability to "reclaim" waste: 🪴 for ceramics, wasted clay from throwing goes right back in to your next pot 🪵 for wood working, reusing 2x4s is second nature 🥘 for cooking, leftover food is tossed into salads 💻 what's for software? bin-packing containers? 😆
040
Almog @almog.xyz · 28/01/2025
Kafka Streams 101: Windows & Time! 🕰️ What's the difference between event, stream and wall clock time? 🪟 What are the four different types of windows? ⁉️ What are the important error messages and metrics, and what do they mean? 👉 Read the full lesson here: www.responsive.dev/blog/windows...
The different types of windows in Kafka Streams
084
Almog @almog.xyz · 28/01/2025
🙏 A @notion.com feature on my wishlist: @jetbrains.com IDEs have this features that lets you select the open file in the nav hierarchy. I've always wanted to be able to do this in Notion so I can more easily explore without navigating away from a page.
select opened file in nav UI
110
Almog @almog.xyz · 24/01/2025
Kicking the tires on animated explanations of windowing strategies in Kafka Streams. Would love feedback on if this animation for session windows resonates: they keep growing until there is an 'inactivity gap' (no events for the key). This here shows a session window with inactivity gap of 10.
161
Almog @almog.xyz · 23/01/2025
There are some learnings that only come from having seen a wildly successful technology go from a nascent project to ubiquitous. These insights can change the trajectory for new tech. Jason's summary about error handling is a great read for practical engineers: github.com/slatedb/slat...
github.com
API Error Guidance · Issue #454 · slatedb/slatedb
Dumping a few thoughts on API errors. At the moment, we have a single flat error SlateDBError which is exposed through the public API. This is the same pattern used in Kafka and there are some less...
0155
Reposted by Almog
Chris @chris.blue · 19/01/2025
Nice short read. Got me wondering if we should support SI in SlateDB. So far I’ve been pretty hardline on just SSI.
brooker.co.za
Snapshot Isolation vs Serializability - Marc's Blog
6653
Almog @almog.xyz · 16/01/2025
I just discovered @localfirst.fm and can’t help but compare the parallels between that movement and the resurgence of popularity of embedded data systems & BYOC (or even on prem) tech. “Serverless” has its uses, but the limitations are increasingly clearer.
160
Almog @almog.xyz · 15/01/2025
I wish @xkcd.com published an autobiography. Somehow their work constantly hits bullseye and I’d love to get just a glimpse of the story behind the person that generates creativity in the same league as history’s greatest artists.
070
Reposted by Almog
Sophie Blee-Goldman @ableegoldman.bsky.social · 15/01/2025
[1/5] Hey fellow Kafka Streams otters! You may have already seen my previous blog posts such as the "Debugging Rebalances" and "EOS guide" from last year (and if you haven't, go check them out! Links in thread) But now it's 2025, and time for me to think about what to write next...
182
Almog @almog.xyz · 14/01/2025
[1/2] Tombstones & consistency in distributed databases: you have to take care when deleting rows! 😨 If there's no data for a key, it could mean: 🌱 First write attempt 🪦 Deleted but no longer retained What if it's a zombie writer that needs to check the row's epoch?
170
Almog @almog.xyz · 10/01/2025
Single-writer systems are 🤩 - they transform traditional transactions into a much simpler problem: fencing zombies. This allows read-modify-write without coordination, instead relying on commits to discard data if its fenced (rare). SlateDB enforces single-writer. Let the use cases "sync" in 😉
1142
Almog @almog.xyz · 08/01/2025
I've learned more about ACID in the last two days reviewing the discussion on github.com/slatedb/slat... than I have in the last decade in industry. A big takeaway is that using object store for durability opens a new can of worms that existing systems can ignore (high write latency).
github.com
docs: Add rfc about Transaction by flaneur2020 · Pull Request #260 · slatedb/slatedb
fixes #248 this doc is still in WIP, but the discussion is open here. TLDR of this doc: Snapshot (with sequence number in keys) and WriteBatch are MUST before working on the Transaction feature SS...
2182
Almog @almog.xyz · 21/12/2024
Interviewer: Can you explain this gap in your resume? Kafka Streams: I was taking some R&R… rebalancing and restoration.
060
Reposted by Almog
Kafka Streams @kafkastreams.bsky.social · 20/12/2024
Kafka Streams has officially come to Bluesky! Follow this account for updates on everything Kafka Streams and keep up to date with new features, bug fixes, and upcoming KIPs. Long live the otter!
1167
Almog @almog.xyz · 19/12/2024
Kafka Streams: 25 improvements in time for 2025! Big wins in performance (KIP-925: Rack Aware Assignment), customization (KIP-924: Custom Task Assignment, KIP-954: Custom Storage) and monitoring (KIP-1091: thread state metrics) - just to name a few. See them all: www.responsive.dev/blog/kafka-s...
responsive.dev
Kafka Streams Wrapped: 2024 in review
Kafka Streams had some tremendous additions and improvements in 2024. Here are the highlights.
041
Almog @almog.xyz · 14/12/2024
Who’s building a Reddit clone on atproto?
133
Almog @almog.xyz · 13/12/2024
👋 Welcome @davistreybig.bsky.social to bluesky! He's played such an important part in every stage of Responsive's journey: from market research, to product to pricing. If you want great startup content, give him a follow.
020
Reposted by Almog
Chris @chris.blue · 12/12/2024
SlateDB is on fire right now. No other way to put it. WIPs: - range queries (Jason @ responsive.dev) - snapshots/transactions (@flaneur2024.bsky.social @ databend.com) - merge operator (David Moravek @ CFLT) - checkpoints (@rohanpd.bsky.social @ r.d) - ttl (@almog.ai @ r.d) RFCs below 👇
2252
Almog @almog.xyz · 12/12/2024
Starting in 10 minutes! If you are in production with Kafka Streams or curious about how Metronome handles insane throughput in realtime, you won't want to miss this at 9:30AM PT. www.linkedin.com/events/q-a-s...
linkedin.com
Q&amp;A: Scaling Event Driven Architectures with Metronome | LinkedIn
We're hosting a live Q&A session with two engineers who have seen it all: - Casey Crites, a founding engineer at Metronome, implemented and scaled the Kafka pipeline at Metronome from zero to where i...
020
Almog @almog.xyz · 12/12/2024
The momentum on SlateDB development can't be stopped! We're getting close to closing out another foundational RFC by @flaneur2024.bsky.social: transactions 🏦 Mostly taking inspiration from Badger, but there's some innovative stuff there. Check it out: github.com/slatedb/slat...
github.com
docs: Add rfc about Transaction by flaneur2020 · Pull Request #260 · slatedb/slatedb
fixes #248 this doc is still in WIP, but the discussion is open here. TLDR of this doc: Snapshot (with sequence number in keys) and WriteBatch are MUST before working on the Transaction feature SS...
0125
Almog @almog.xyz · 11/12/2024
It's the bloom filters. It's always the bloom filters.
270
Almog @almog.xyz · 08/12/2024
There are 2.4999999999994 hard things in computer science: naming, floating point math and exactly once delivery.
051