Sign in

Almog

@almog.xyz
2.8K followers 234 following 145 posts

Favorite Buzzwords: Kafka Streams | SlateDB | Stream Processing | Distributed Databases co-founder @ responsive.dev

PostsRepliesMedia
Almog @almog.xyz · 23/04/2026
A deep dive to how metrics are stored and queried. Inverted indexes, roaring bitmaps, and ... gorillas? www.bitsxpages.com/p/how-metric...
bitsxpages.com
how metrics are stored and queried
the data structures powering your metrics dashboard
070
Almog @almog.xyz · 16/04/2026
I'd argue that's exactly what's in the price tag ;) that's because the price tag of managed services include that, which I'd argue aren't necessary for opendata timeseries
000
Almog @almog.xyz · 16/04/2026
The team worked insane hours to get this release ready and we've finally shipped it! We built a first of its kind, prometheus-compatible database that is object native. ~22x cheaper than AWS AMP and stupid easy to operate. www.opendata.dev/blog/introdu...
opendata.dev
OpenData Timeseries: Prometheus-compatible metrics on object storage | OpenData
An MIT-licensed, Prometheus-compatible timeseries database built on SlateDB, bringing the operating model and cost structure of object-store-native systems to the Prometheus and Grafana ecosystem.
2195
Almog @almog.xyz · 27/03/2026
I know "measure twice, cut once" applies to performance optimizations, but where's the fun in that? I am quite proud of my lopsided planter boxes and custom memory allocators.
020
Almog @almog.xyz · 24/03/2026
Thanks! I haven't listened to that talk, adding it to the backlog
010
Almog @almog.xyz · 23/03/2026
Databases have strange economic incentives that end up making them "worse" over time. The problem is an asymmetry between who benefits from innovations at first (database companies) and who benefits in the long run (hyperscalers that sell hardware). Blogged: www.bitsxpages.com/p/the-broken...
bitsxpages.com
the broken economics of databases
Why database companies charge so much, earn so little, and keep making things complicated
181
Almog @almog.xyz · 17/03/2026
I regularly get asked "how do you create your diagrams for your blogs"? Recently I've been using my own tool, and I'm open sourcing it! Yuzudraw is a visual editor for ASCII that has a token-efficient DSL that lets agents generate/modify your diagrams. Give it a ⭐! github.com/agavra/yuzud...
github.com
GitHub - agavra/yuzudraw: ASCII diagramming tool for both humans and agents
ASCII diagramming tool for both humans and agents. Contribute to agavra/yuzudraw development by creating an account on GitHub.
030
Almog @almog.xyz · 02/03/2026
“Gradually, then suddenly.” That’s how adoption works when you’re building something new. Opendata is still "gradually" but 100 stars with $0 spent on marketing is a good start. Back to building! github.com/opendata-oss...
091
Almog @almog.xyz · 12/02/2026
Yes! With larger payloads compression is almost always worth it. Thanks for sharing
000
Almog @almog.xyz · 10/02/2026
I implemented prefix compression for SlateDB & noticed benchmarks looked "worse". Fell down a rabbit hole. Turns out I was thinking about compression backwards... Wrote up my learning: www.bitsxpages.com/p/the-mathem...
bitsxpages.com
the mathematics of compression in database systems
why compression is (almost) always worthwhile
1101
Almog @almog.xyz · 28/01/2026
Can you beat 180KB? I created a challenge to reduce a dataset as much as possible 🏆 my approach uses delta encoding & prefix compression + zstd(22) compression to reduce 25MB -> 180KB github.com/agavra/bit-g...
github.com
GitHub - agavra/bit-golf: a compression golf challenge for GitHub event data
a compression golf challenge for GitHub event data - agavra/bit-golf
082
Almog @almog.xyz · 09/01/2026
by the way, it's pronounced like "tweaker"
010
Almog @almog.xyz · 09/01/2026
I fixed the problem with reviewing code written by Claude in "accept edits on" mode: github.com/agavra/tuicr would love to know what you think, and if you want to contribute to an OSS rust project there's a bunch of open issues to pick up!
github.com
GitHub - agavra/tuicr: Review AI-generated diffs like a GitHub pull request, right from your terminal.
Review AI-generated diffs like a GitHub pull request, right from your terminal. - agavra/tuicr
120
Almog @almog.xyz · 05/01/2026
I'll die on this hill: Sorted String Tables (SSTs) are the single most important data structure for modern DBs. They lean in to the limitations of SSDs and Object Storage, making them (and similar layouts) the best choice for many databases. Blogged in detail: www.bitsxpages.com/p/sorted-str...
bitsxpages.com
sorted string tables (SST) from first principles
why sorted string tables are the swiss army knife for data systems and how they are implemented
060
Almog @almog.xyz · 10/12/2025
thanks for the kind words! I use monodraw.helftone.com -- not sure if planetscale/turbopuffer use the same but I was definitely inspired by their design style
monodraw.helftone.com
Monodraw for macOS — Helftone
010
Almog @almog.xyz · 09/12/2025
The inner join between sets of people who build databases, write, and draw? Low cardinality. I'm in that set, so I'm starting a blog! Here's my first post: www.bitsxpages.com/p/frameworks...
bitsxpages.com
frameworks for understanding databases
building mental models for tradeoffs in performance, availability and durability in data systems
2163
Almog @almog.xyz · 25/11/2025
Sometimes the best solution is "do nothing", but it's always more fun to play with tools.
040
Almog @almog.xyz · 21/10/2025
Calling database nerds in SF! I'm covering SlateDB at the systems meetup next Wednesday (10/29). If you're around, I'd love to meet you in person (that way you'll have proof I'm not just an AI bot). 👉 luma.com/e7feg2i6
0153
Almog @almog.xyz · 08/08/2025
I wonder why it hasn't made it's way to the US! The only time I get it is when I make it at home.
120
Almog @almog.xyz · 08/08/2025
Beans on toast is underrated.
210
Almog @almog.xyz · 07/08/2025
I recently implemented Gorilla encoding (www.vldb.org/pvldb/vol8/p...) for a SlateDB PR. Pretty cool stuff - easy to understand but really powerful. Here it is, explained by a gorilla.
091
Almog @almog.xyz · 05/08/2025
I guess the answer to that is “technically, yes.”
000
Almog @almog.xyz · 05/08/2025
It's true. Every company eventually becomes a database company.
0101
Almog @almog.xyz · 15/07/2025
Despite using so many new technologies, I somehow never learn my lesson: read the docs sooner and read the docs thoroughly.
092
Almog @almog.xyz · 08/07/2025
Reading the tokio-rs async documentation makes me feel like...
161
Almog @almog.xyz · 02/07/2025
that's so cool - I love contraptions that are just there to show that we can do something!
000
Almog @almog.xyz · 01/07/2025
No shame in appreciating bona fide nerd humor 😉
010
Almog @almog.xyz · 01/07/2025
One day I'll open a coffee shop dedicated to the not-insignificant intersection between database nerds and coffee snobs. Until then, enjoy this comic.
1395
Almog @almog.xyz · 28/04/2025
Maybe... just maybe, adding more features and complexity into stream processors is NOT what we need?
061
Almog @almog.xyz · 24/04/2025
Fact.
000
Almog @almog.xyz · 24/04/2025
The new electric Caltrain cars have WiFi. 🙏 SF bay area has finally entered the 21st century (on this dimension of public transit only).
010
Reposted by Almog
Chris @chris.blue · 22/04/2025
Today marks SlateDB’s one year anniversary! It’s been a lot of fun. Thanks to @rohanpd.bsky.social @flaneur2024.bsky.social @almog.ai @vigneshc.bsky.social @paulbutler.org Jason Gustafson, David Moravek, and many others for joining the project. 😀
slatedb.io
SlateDB - An embedded storage engine built on object storage | SlateDB
Description will go into a meta tag in <head />
0165
Almog @almog.xyz · 21/04/2025
This got me thinking: Should MCP itself formally distinguish between request vs response context? It could let LLMs be more intelligent about when to pull extra info. 🤔 [6/6]
000
Almog @almog.xyz · 21/04/2025
I built a demo agent for hotel booking. My setup: Inject request context into the prompt (user prefs, budget, pulled from account) Fetch response context live using MCP tools (available hotels that match search params) Result: the agent felt way smarter. [5/N]
100
Almog @almog.xyz · 21/04/2025
Back when we only had prompts, you could frame the task (request context). Today, tools like MCP change the game. But should you use them? [4/N]
100
Almog @almog.xyz · 21/04/2025
Example: a hotel booking agent 🛌 Request context: "Book a hotel in Barcelona on April 28 under $400" Response context: A list of hotels available in Barcelona that day, within budget. [3/N]
100
Almog @almog.xyz · 21/04/2025
⚖️ You need to balance two types of context: Request Context = frames the task Response Context = info needed to complete it [2/N]
100
Almog @almog.xyz · 21/04/2025
Prompt engineering was v0. Context engineering is v1.0. How should you think about supplying LLMs with the right information? 📖 [1/N]
140
Almog @almog.xyz · 10/02/2025
"Kafka Configs: A Rube Goldberg Machine?" - an actual quote from a customer after discussing cleanup.policy and Kafka Streams changelogs.
030
Almog @almog.xyz · 04/02/2025
This is really neat, thanks for sharing. Little performance hacks like this require understanding how things work underneath the hood, and tend to have the most impact on overall performance! Would love to know other similar tips you have.
020
Reposted by Almog
Chris @chris.blue · 03/02/2025
SlateDB now has clones 🤯 Users can clone an existing DB's data to a new location. It's nearly instantaneous since it references the data from the old bucket rather than copying. Writes to the clone update the new location. Compaction lazily merges old data into the new directory./ht @responsive.dev
github.com
Initial clone implementation by hachikuji · Pull Request #430 · slatedb/slatedb
Fixes #315. This patch contains the logic to create and initialize clones as outlined in RFC-0004.
1332
Reposted by Almog
Apurva Mehta @apurvamehta.com · 04/02/2025
Should you use Kafka? If so, when? And what are the tradeoffs presented by the dizzying variety of Kafka-adjacent technologies? I hope that my latest blog post provides unique and useful answers to these questions. Let me know what you all think! 👉 www.responsive.dev/blog/why-whe...
0155
Almog @almog.xyz · 31/01/2025
I often have days where GPT is immensely helpful, today I had one where it had led me down multiple rabbit holes and I ended up solving my problem with good old Google search. Nice reminder of when it’s helpful and when it’s not!
030
Almog @almog.xyz · 30/01/2025
Username checks out 😆
010
Almog @almog.xyz · 30/01/2025
As part of my post-Elden Ring hunt for a game, I tried Balatro last night. I'm starting to think I may have made a mistake.
130
Reposted by Almog
ScyllaDB @scylladb.com · 29/01/2025
How did Responsive simplify its architecture and achieve monstrously high availability with unlimited scale potential? Co-founder @almog.ai will share how replacing RocksDB with #ScyllaDB in #Kafka Streams led to big wins for their team at Monster Scale Summit. www.scylladb.com/monster-scal...
084
Almog @almog.xyz · 29/01/2025
My hobbies all have the ability to "reclaim" waste: 🪴 for ceramics, wasted clay from throwing goes right back in to your next pot 🪵 for wood working, reusing 2x4s is second nature 🥘 for cooking, leftover food is tossed into salads 💻 what's for software? bin-packing containers? 😆
040
Almog @almog.xyz · 28/01/2025
Kafka Streams 101: Windows & Time! 🕰️ What's the difference between event, stream and wall clock time? 🪟 What are the four different types of windows? ⁉️ What are the important error messages and metrics, and what do they mean? 👉 Read the full lesson here: www.responsive.dev/blog/windows...
The different types of windows in Kafka Streams
084
Almog @almog.xyz · 28/01/2025
Haha I often do! My last jam was www.dishoom.com/store/produc... from the Dishoom Cookbook. Would recommend! Sweet and a touch of spicy.
110
Almog @almog.xyz · 28/01/2025
I did see your other post (which is what inspired this one!) but I actually haven't noticed that same problem recently. Maybe we have a smaller notion space?
100