Sign in

Vignesh Chandramohan

@vigneshc.bsky.social
1.2K followers 87 following 41 posts

Stream processing, data infra, Table formats. vigneshc.github.io

PostsRepliesMedia
Vignesh Chandramohan @vigneshc.bsky.social · 27/06/2026
Does anyone have a production use case example for Quack? Remote reads make sense for local cache like scenarios. I am curious about remote writes. Great talk from this year's AI council on duckdb quack : youtu.be/O_tzHeDpjrE?...
youtu.be
Super-Secret Next Big Thing for DuckDB
YouTube video by AI Council
000
Vignesh Chandramohan @vigneshc.bsky.social · 26/05/2026
A short post on Iceberg deletions : vigneshc.github.io/blog/iceberg...
vigneshc.github.io
Iceberg Deletes — Equality Deletes and V3 Deletion Vectors
A short walkthrough of Apache Iceberg deletes: equality deletes, V3 deletion vectors, and streaming considerations.
030
Vignesh Chandramohan @vigneshc.bsky.social · 21/05/2026
Is the usecases for llm talking substrait for performance? Isn't SQL more "reviewable" for the final output? Conversions to substrait can then be mechanical. Related paper that talks about providing the 'agent intent' to query engines, and how SQL alone isn't enough : arxiv.org/abs/2509.00997
arxiv.org
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
Large Language Model (LLM) agents, acting on their users' behalf to manipulate and analyze data, are likely to become the dominant workload for data systems in the future. When working with data, agen...
000
Vignesh Chandramohan @vigneshc.bsky.social · 12/04/2026
Exploring Apache Iceberg and SlateDB formats - with a repo link for additional exploration. datapapers.substack.com/p/exploring-...
datapapers.substack.com
Exploring Table Formats - Iceberg & SlateDB
What is common between Apache Icerberg, SlateDB and object-store first table formats? Explore with real examples.
081
Vignesh Chandramohan @vigneshc.bsky.social · 10/04/2026
Tests are more comprehensive than spec level validation. Isn't spec still more deterministic, more token efficient, and theoretically makes coding agents converge faster? Would specs being more approachable for human reviewers, and generated by coding agents change the equation?
110
Vignesh Chandramohan @vigneshc.bsky.social · 17/12/2025
Related items: arxiv.org/abs/2509.00997 - talks about agent's query patterns, and how data systems should adapt. www.malloydata.dev - another promising query language, the claim is complex queries are expressed in simpler form than SQL, making llms make less mistake.
arxiv.org
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First
Large Language Model (LLM) agents, acting on their users' behalf to manipulate and analyze data, are likely to become the dominant workload for data systems in the future. When working with data, agen...
150
Vignesh Chandramohan @vigneshc.bsky.social · 09/12/2025
@jayaprabhakar.bsky.social Interesting read on formal verification.
020
Vignesh Chandramohan @vigneshc.bsky.social · 31/10/2025
Short talk on Iceberg use cases in last week's Seattle Iceberg meetup. youtu.be/F7qpOVVnxek?...
youtu.be
​Iceberg Use Cases at DoorDash
YouTube video by Apache Iceberg™ Meetup
020
Vignesh Chandramohan @vigneshc.bsky.social · 31/10/2025
It has three increasingly more verbose levels of description. They are probably trying to optimize the initial set of searches. And with markdown based skills, custom workflows are approachable for a broader audience compared to MCP, and probably safer too. It is only available in Claude desktop.
020
Reposted by Vignesh Chandramohan
Chris @chris.blue · 21/10/2025
If you find yourself in SF next week, @almog.xyz is talking about SlateDB at the SF Systems Meetup on Wednesday!
luma.com
SF Systems Meetup: Databases and Stateful Apps · Luma
The SF Systems Meetup is back with a pair of talks giving us a peek at the future of state management! This month, we're excited to have talks from Almog Gavra…
0174
Vignesh Chandramohan @vigneshc.bsky.social · 29/09/2025
Tools, prompts, sampling - all these seem to be a result of generalizing how Claude code / research feature was built over time, and extracting ask out of those patterns. Uniformity is the biggest and probably only benefit. And maybe the goal is not about solutions that spend less tokens?
020
Vignesh Chandramohan @vigneshc.bsky.social · 28/09/2025
Intend to follow this along. Done with chapter 1, looking forward to the next one!
010
Vignesh Chandramohan @vigneshc.bsky.social · 27/09/2025
2/ This mindset improves productivity and outcome significantly overtime.
020
Vignesh Chandramohan @vigneshc.bsky.social · 27/09/2025
1/ Nice read: medium.com/@xiafan/time... Software engineer growth in AI era: * Build composable tools. * Assume non deterministic outcomes from a group of ai agents. * Understand how LLM works to the next level, like how you would understand to read a query plan.
medium.com
Time to Adapt and Become a Better Engineer
A colleague recently asked if I think AI agents could do all the jobs humans do. My answer was an unequivocal yes. Given the same sensory…
120
Reposted by Vignesh Chandramohan
Chris @chris.blue · 25/08/2025
Just added sqlync.com to SlateDB's adopters list! They're building a streaming system that speaks MQTT or PostgreSQL across millions of connected users and devices. 🤯
sqlync.com
SQLync
051
Vignesh Chandramohan @vigneshc.bsky.social · 24/08/2025
3/ And as an extension, how it handles maintenance operations such as vacuum on iceberg tables dones out of band.
020
Vignesh Chandramohan @vigneshc.bsky.social · 24/08/2025
2/ I assume the iceberg writes uses iceberg open source libraries. This would ensure the write part continues to evolve with iceberg advancements. I don't yet know if this handles compacted topics (which would introduce deletes on iceberg)
120
Vignesh Chandramohan @vigneshc.bsky.social · 24/08/2025
1/ Leveraging Remote storage manager and storing Kafka segments as parquet files + iceberg metadata is really good. Avoids having to consume, serialize and manage a separate process. I wonder if confluent's TableFlow launched about a year back has a similar design. www.confluent.io/blog/introdu...
confluent.io
Introducing Tableflow: Unifying Streaming and Analytics
Seamlessly integrate Apache Kafka data into your lakehouse as Apache Iceberg or Delta Lake tables, bridging the operational and analytical divide, with Tableflow. Read more in our blog post.
110
Vignesh Chandramohan @vigneshc.bsky.social · 07/07/2025
Love the idea. Could some of these eventually become sub projects, and hosted in the SlateDB organization as a separate repo? Starting projects that have that potential as GitHub issues with a specific tag would make it easy to track.
110
Reposted by Vignesh Chandramohan
Chris @chris.blue · 18/06/2025
Insane amount of SlateDB work going on: - snapshot reads - split/merge DBs (zero copy) - deterministic simulation testing And someone just pushed Python bindings in a PR! 🤯
0103
Vignesh Chandramohan @vigneshc.bsky.social · 30/05/2025
My Data council talk on SlateDB. youtu.be/gcTRXZeKbNg?...
youtu.be
Internals of SlateDB: An Embedded Key Value Store Built On Object Storage
YouTube video by Data Council
0213
Vignesh Chandramohan @vigneshc.bsky.social · 27/05/2025
Got it. So, if I wanted a view to update, say once an hour incrementally, would I create a "hourly view" that uses now() and join against it?
110
Vignesh Chandramohan @vigneshc.bsky.social · 27/05/2025
Clock tick as an input is indeed a way to model it! Would the clock tick table be joined in all views that need this property?
100
Vignesh Chandramohan @vigneshc.bsky.social · 27/05/2025
Finally got to read this. One additional aspect to ivm, is reasoning about the data in the computed. For a lot of use cases, it is often easy to think of a view/table to move in predictable increments (day, hour, 15 minutes etc). This notion is not modeled as a first class concept in many.
110
Reposted by Vignesh Chandramohan
Chris @chris.blue · 24/04/2025
SlateDB 0.6.0 is out! github.com/slatedb/slat... Highlights include a hybrid cache (using Foyer), a lot of internal cleanup, and more groundwork for transactions. Oh, and put performance jumped ~80% for write-heavy workloads :) slatedb.io/performance/...
081
Reposted by Vignesh Chandramohan
Chris @chris.blue · 22/04/2025
Today marks SlateDB’s one year anniversary! It’s been a lot of fun. Thanks to @rohanpd.bsky.social @flaneur2024.bsky.social @almog.ai @vigneshc.bsky.social @paulbutler.org Jason Gustafson, David Moravek, and many others for joining the project. 😀
slatedb.io
SlateDB - An embedded storage engine built on object storage | SlateDB
Description will go into a meta tag in <head />
0165
Reposted by Vignesh Chandramohan
Commonhaus Foundation @commonhaus.org · 10/04/2025
Commonhaus is 1! 🎂 14 projects, solid foundations, and more on the way. If you believe in light governance, shared care, and thoughtful support for open source, come see what we’re building. www.commonhaus.org/activity/253...
commonhaus.org
🎂 Commonhaus Turns One — A Look Back, and the Road Ahead
Commonhaus Foundation celebrates its first anniversary and lays down expectations for its future
03019
Reposted by Vignesh Chandramohan
Mike Driscoll @medriscoll.com · 16/04/2025
Yo SF Bay Area #databs crew, want to talk lakehouses at a real Lake House? :) Next week after Data Council, join the founders of @clickhouse.com, @motherduck.com, @startreedata.bsky.social, and @tobikodata.com to talk real-time databases and next-generation ETL. www.rilldata.com/events/data-...
1103
Reposted by Vignesh Chandramohan
Chris @chris.blue · 17/03/2025
SlateDB 0.5.0 is out! Features: - Checkpoints - Clones - Read only client - Split/merge database foundation - TTL filtering on reads - Last version with breaking byte format changes By the numbers: - 62 commits - 2 new contributors - 10 total contributors github.com/slatedb/slat...
github.com
Release v0.5.0 · slatedb/slatedb
What's Changed Refactor Block Tests to Use Table-Driven Test Cases by @samsond in #410 Update await calls in README.md by @criccomini in #425 chore: Apply table driven test for sst.rs by @jeffreyl...
2213
Vignesh Chandramohan @vigneshc.bsky.social · 02/03/2025
datapapers.substack.com/p/building-c... New post.
datapapers.substack.com
Building composable data systems: Why, How and Standards
Standards improve interoperability. Reusable libraries built around standards drive adoption. In this post, we explore key papers and real-world examples.
041
Vignesh Chandramohan @vigneshc.bsky.social · 23/02/2025
DEBS conference hosts a grand challenge every year. This year's challenge is detecting outliers in a stream of images from laser powder bed fusion. The challenge involves submitting a kubernetes app (constraint: 2 cores 8 gb). Interesting to try if you have the time! 2025.debs.org/call-for-gra...
2025.debs.org
CALL FOR GRAND CHALLENGE SOLUTIONS
DEBS2025
010
Vignesh Chandramohan @vigneshc.bsky.social · 15/02/2025
Great episode! Towards the end @vanlightly.bsky.social mentions about alloytools.org finding a data model bug. Never thought of an intersection between data model and formal verification. Do you have more details on this?
100
Reposted by Vignesh Chandramohan
Diptanu Choudhury @diptanu.bsky.social · 17/12/2024
Python Folks - which data/workflow engine has the best developer experience for packaging code? We have looked into - Modal, Beam, Airflow, Flyte, AWS Lambda, Prefect, Dagster and Spark. Haven’t seen any approach which is fast, reliable and intuitive.
6102
Vignesh Chandramohan @vigneshc.bsky.social · 19/01/2025
What are some papers or blogs about data quality challenges? I see tools like great expectations, table formats features like 'check constraints' in Delta. I don't yet see it as a first class property of catalogs. Found this, are there others? journalofbigdata.springeropen.com/articles/10....
journalofbigdata.springeropen.com
Big data quality framework: a holistic approach to continuous quality management - Journal of Big Data
Big Data is an essential research area for governments, institutions, and private agencies to support their analytics decisions. Big Data refers to all about data, how it is collected, processed, and ...
131
Vignesh Chandramohan @vigneshc.bsky.social · 19/01/2025
Great talk by Binwei Yang on Apache Gluten last week. youtu.be/GWTj3INSzPg?... Apache Gluten moves execution of spark operators to native backend like Velox, accelerating query performance. It has basic iceberg support too! github.com/apache/incub...
youtu.be
Big Data Bellevue: Apache Gluten: Accelerating SparkSQL with Spark on Velox
YouTube video by BDB
010
Vignesh Chandramohan @vigneshc.bsky.social · 11/01/2025
This book was on my list for the year, joining!
020
Reposted by Vignesh Chandramohan
Chris @chris.blue · 31/12/2024
SlateDB 0.4.0 is out! Features: - Range scans - No DynamoDB needed for S3 - Nightly perf tests - Merge operator groundwork - GC improvements By the numbers: - 57 commits - 5 new contributors - 11 total contributors github.com/slatedb/slat...
3254
Vignesh Chandramohan @vigneshc.bsky.social · 15/12/2024
Finally, the blog from 2023 says it isn't used in production yet. Any recent data points on production experience that can be shared now? :)
110
Vignesh Chandramohan @vigneshc.bsky.social · 15/12/2024
What do you think about flink materialized views or dynamic tables? Can the hoptimator concept (i.e. provision the flink and other internal + external connectors to do what user asked) be part of flink eventually?
100
Vignesh Chandramohan @vigneshc.bsky.social · 15/12/2024
I finally understood what it is after reading this blog :) www.linkedin.com/blog/enginee... Developer experience / testing is one of the hard aspect of declarative pipelines. Pipelines, while they take more steps, is more predictable. Is there a snappy preview builtin, to mitigate some of this?
linkedin.com
Declarative Data Pipelines with Hoptimator
100
Vignesh Chandramohan @vigneshc.bsky.social · 12/12/2024
Great writeup! Did not realize flink CDC is finally just another flink job, and that it can use any of the debezium source connectors! On iceberg, I do see a flink connector in the iceberg project. What needs to happen to make flink iceberg work with flink CDC out of the box (at par with Kafka)
010
Vignesh Chandramohan @vigneshc.bsky.social · 04/12/2024
Great write up! Looking at the apis, it only deals with the catalog aspect. Writing the manifest file etc are still the responsibility of client. Is the writes in systems like rust client, materialize difficult because of the lack of these apis or is writing metadata etc hard as well?
110
Vignesh Chandramohan @vigneshc.bsky.social · 03/12/2024
I don't understand it as well. Documentation has missing parts. Example shows a spark package, specifically for S3 Tables. It might be an implementation of the catalog. Apis that backs this implementation is not in the documentation yet.
010
Vignesh Chandramohan @vigneshc.bsky.social · 29/11/2024
I wonder if Fluss would extend and use XTable, since it deals with multiple table formats. xtable.apache.org
xtable.apache.org
Apache XTable™ (Incubating)
Apache XTable™ (Incubating) is a cross-table interop of lakehouse table formats Apache Hudi, Apache Iceberg, and Delta Lake. Apache XTable™ is NOT a new or separate format, Apache XTable™ provides abs...
010
Vignesh Chandramohan @vigneshc.bsky.social · 29/11/2024
Paimon is already the second version, with first one being flink table store. Paimon documentation explicitly says not to use it without a cluster framework like flink, unlike iceberg and delta which are building kernel libraries. There is no mention of Paimon in fluss,likely not a evolution.
010
Vignesh Chandramohan @vigneshc.bsky.social · 29/11/2024
I think it plays into a similar space as the following research.google/pubs/vortex-... storage API and background optimizations to iceberg www.confluent.io/blog/introdu... - likely the backend of this solves some of the same problems.
research.google
Vortex: A Stream-oriented Storage Engine For Big Data Analytics
141
Vignesh Chandramohan @vigneshc.bsky.social · 26/11/2024
Was the post commit sequential tests or did it combine multiple PRs together? Was there issues related to having to revert multiple commits due to an issue with one?
120
Vignesh Chandramohan @vigneshc.bsky.social · 26/11/2024
I recently started reading about iceberg. Would using something like slatedb with a custom schema/convention for storing the catalog info work?
120
Reposted by Vignesh Chandramohan
Chris @chris.blue · 21/11/2024
SlateDB is now part of @commonhaus-fdn.bsky.social! I think we might be the first non-JVM project. Looking forward to more projects joining us here. Our experience with the foundation has been excellent so far. github.com/commonhaus/f...
github.com
SlateDB would like to join Commonhaus · commonhaus foundation · Discussion #213
Project information Project name: SlateDB Project website: https://slatedb.io Code repository: https://github.com/slatedb/slatedb License: Apache 2.0 Do you have authority to represent this project...
2296
Vignesh Chandramohan @vigneshc.bsky.social · 18/11/2024
Netherite, the engine behind azure durable function uses event sourcing as well. My notes on the paper and links to additional material. datapapers.substack.com/p/paper-note...
datapapers.substack.com
Paper notes: Netherite, Efficient execution of serverless workflows
Multi-Step workflows offered as a service enhance developer productivity significantly for a number of use cases.
020