Sign in

Ananth Packkildurai

@ananthdurai.bsky.social
3.6K followers 579 following 267 posts

Editor Data Engineering Weekly; subscribe www.dataengineeringweekly.com. In Prgress, LakeByte

PostsRepliesMedia
Ananth Packkildurai @ananthdurai.bsky.social · 13/04/2026
I just realized: if Airflow and other orchestration engines emit open-lineage data to S3 Files and enable Claude to search them, you've got a data catalog.
230
Ananth Packkildurai @ananthdurai.bsky.social · 13/04/2026
So, are we debating the semantic layer again? docs.getdbt.com/blog... What do you call a semantic layer from your perspective?
docs.getdbt.com
Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update | dbt Developer Blog
With 2026's best models, the dbt Semantic Layer hits near-100% accuracy for covered queries. Here's what changed and what didn't in our updated benchmark.
050
Ananth Packkildurai @ananthdurai.bsky.social · 13/04/2026
Whether you like or dislike Apache Kafka, its KIPs are among the best learning materials for distributed systems. KIP-848 is an excellent read cwiki.apache.org/con...
040
Ananth Packkildurai @ananthdurai.bsky.social · 02/04/2026
Most data platform failures don’t start with bad infra. They start at the team boundary. My new post argues platforms scale through operating interfaces: contracts, ownership, communication, and adoption design, not tooling alone.
dataengineeringweekly.com
The Missing Interface in Data Platform Engineering
How data leaders should design the boundary between platforms and dependent teams.
030
Ananth Packkildurai @ananthdurai.bsky.social · 11/03/2026
More ETL pipelines will run next year than ever before. And ETL is still dead. Not dead, like nobody uses it. Dead like landlines — they work, but nobody builds their strategy around one.
dataengineeringweekly.com
ETL is Dead
Why the shift from human-operated to agent-operated data warehouses demands a new architecture
010
Ananth Packkildurai @ananthdurai.bsky.social · 24/02/2026
The data engineer job title is due for an update. Not because AI is replacing the role, but because AI is finally revealing what the role was always actually about. Moving data was never the point. Meaning it is. Read more:
dataengineeringweekly.com
Data Engineering After AI
Moving Data Was Never the Point. Meaning It Is.
010
Ananth Packkildurai @ananthdurai.bsky.social · 21/02/2026
At the end of 2026, we will talk about "AI Fan Effect" [en.wikipedia.org/wik...] and the invention of a new field: Psychology for AI. Perhaps, I feel this is the future of software engineering.
010
Ananth Packkildurai @ananthdurai.bsky.social · 31/01/2026
As we move from dashboards to autonomous agents, something breaks. Systems of record capture what happened, not why. Why data platforms need Truth Registries + Context Graphs for the agentic era 👇 www.dataengineeringw... #DataEngineering #AgenticAI #Graphs #LLMs
dataengineeringweekly.com
The Missing Layer in Your AI Stack: Context, Not Just State
From SQL to Semantics: The Rise of the Context Graph for AI Agents
030
Ananth Packkildurai @ananthdurai.bsky.social · 26/01/2026
Data Engineering Weekly's 254th edition is out. Context Graph is the new talk of the town!!
000
Ananth Packkildurai @ananthdurai.bsky.social · 23/01/2026
The companies that build the most boring data stack often win the market!!! Prove me wrong.
010
Ananth Packkildurai @ananthdurai.bsky.social · 20/01/2026
Data Contract: There was no shortage of activity around the topic. Definitions were proposed and refined. Conceptual boundaries were drawn and redrawn. I pen down a reflection of the Data Contracts here www.dataengineeringweekly.com/p/data-contr...
dataengineeringweekly.com
Data Contracts: A Missed Opportunity
The Conversation We Should Have Had—Before Thought Leadership Replaced System Design
010
Ananth Packkildurai @ananthdurai.bsky.social · 14/01/2026
How to build a scalable shopping agent? Here's a wild thought: What if—and hear me out—we let humans click that Buy Now button? Just throwing ideas out there.
010
Ananth Packkildurai @ananthdurai.bsky.social · 12/01/2026
This week, it is mostly about Multi-Agent Architecture. Do you think the data infrastructure is ready for a multi-agent architecture? Where is the gap?
dataengineeringweekly.com
Data Engineering Weekly #252
The Weekly Data Engineering Newsletter
010
Ananth Packkildurai @ananthdurai.bsky.social · 09/01/2026
Is semantic Spec Good enough to run an enterprise system? I listed challenges to adopting the Iceberg Rest Catalog
dataengineeringweekly.com
A Critique of Iceberg REST Catalog: A Classic Case of Why Semantic Spec Fails
How a Semantically Correct API Becomes Operationally Unreliable at Scale
000
Ananth Packkildurai @ananthdurai.bsky.social · 23/12/2025
Continuing our yearly tradition of Year in Review Data Engineering Weekly, we published the 2025 Year in Review. What do you think is the most notable trend of 2025?
dataengineeringweekly.com
DEW - The Year in Review 2025
From Digital Plumbers to Architects of Intelligence: The 7 Paradigm Shifts That Defined 2025
000
Ananth Packkildurai @ananthdurai.bsky.social · 16/12/2025
www.dataengineeringw...
030
Ananth Packkildurai @ananthdurai.bsky.social · 12/12/2025
060
Ananth Packkildurai @ananthdurai.bsky.social · 08/12/2025
Look at the tech stack IBM now controls: 🐧 Compute: Red Hat (Linux/OpenShift) ☁️ IaC: HashiCorp (Terraform) 💰 FinOps: Kubecost 🌊 Streaming: Confluent (Kafka) 🧠 Vector/AI: DataStax (Cassandra) ⚡ Query Engine: Ahana (Presto) 🔄 Ingest: StreamSets
140
Ananth Packkildurai @ananthdurai.bsky.social · 08/12/2025
LinkedIn moves FishDB to Rust, DoorDash builds AI swarms, and Dropbox masters context engineering. 🤯 Data Engineering Weekly #247 is packed with system design deep dives from the best engineering teams.
dataengineeringweekly.com
Data Engineering Weekly #247
The Weekly Data Engineering Newsletter
030
Ananth Packkildurai @ananthdurai.bsky.social · 04/12/2025
If the Data Catalog is the answer for AI, the question was wrong.
010
Ananth Packkildurai @ananthdurai.bsky.social · 19/11/2025
We stopped asking if data was useful because storage got cheap. Now, "Dark Data" is actively poisoning your AI context windows with hallucination vectors. Read about the Data Sustainability index
dataengineeringweekly.com
The Dark Data Tax: How Hoarding is Poisoning Your AI
Storage is cheap. Attention is finite. Hallucinations are expensive. It’s time to stop building Data Lakes and start managing Data Metabolism
050
Ananth Packkildurai @ananthdurai.bsky.social · 10/11/2025
The open source companies built their success on top of open-source platforms, benefited from community contributions and adoption, but now must abandon open-source principles to survive commercially.
010
Ananth Packkildurai @ananthdurai.bsky.social · 03/11/2025
🚀 The 244th edition of Data Engineering Weekly dives into: AI agents as execution engines, LLM inference economics, databases for AI, personalization, and product evidence. Read more 👉 www.dataengineeringw... #DataEngineering #AI #LLMs
dataengineeringweekly.com
Data Engineering Weekly #244
The Weekly Data Engineering Newsletter
020
Ananth Packkildurai @ananthdurai.bsky.social · 03/11/2025
Cricket has been India’s greatest force in overcoming centuries of colonial suppression. Today’s Women’s World Cup win echoes the spirit of 1983 — a triumph that will inspire generations to come. 🇮🇳🏆
000
Ananth Packkildurai @ananthdurai.bsky.social · 23/10/2025
This is the most personal essay that I have written in Data Engineering Weekly. I shared a few key moments in my life and how fortunate I was to meet mentors along my professional journey, which shaped my career.
dataengineeringweekly.com
Thinking Like a Data Engineer
A Journey Beyond Code — Toward Systems, Curiosity, and Confidence
090
Ananth Packkildurai @ananthdurai.bsky.social · 17/10/2025
🚀 Data Vault vs. Dimensional Modeling vs. Medallion Architecture — When viewed through a modern enterprise data lens, these techniques interlock. I break down how in Part 2 of my “Revisiting the Medallion Architecture” series.
dataengineeringweekly.com
Revisiting Medallion Architecture: Data Vault in Silver, Dimensional Modeling in Gold
How to Balance Flexibility and Performance in a Modern Data Platform
040
Ananth Packkildurai @ananthdurai.bsky.social · 17/10/2025
Fivetran and dbt form a strong foundation for modern data infrastructure, known for bringing simplicity to complex engineering workflows. That said, calling it “open” data infrastructure feels like a stretch.
350
Ananth Packkildurai @ananthdurai.bsky.social · 13/10/2025
Should we update the definition of an "Analytical Engineer"?
040
Ananth Packkildurai @ananthdurai.bsky.social · 09/10/2025
As a data engineer, you can't treat zero-party (consent) and third-party (inferred) data the same way. This distinction is critical for building systems that are scalable, private, and trustworthy. Here’s my guide:
dataengineeringweekly.com
Engineering Growth: The Data Layers Powering Modern GTM
Building privacy-preserving pipelines that unify zero-, first-, second-, third-, and fourth-party data into a coherent GTM ecosystem.
050
Ananth Packkildurai @ananthdurai.bsky.social · 02/10/2025
Airbnb: Real-Time Key-Value Store Airbnb’s next-gen key-value store supports real-time ingestion and bulk uploads with sub-second latency, powering feature stores and fraud detection. Read the full story here: www.dataengineeringw...
010
Ananth Packkildurai @ananthdurai.bsky.social · 01/10/2025
Grab: Partner Gateway Metrics at Sub-Second Speed Real-time partner analytics at scale is tough. Grab uses Apache Pinot, Kafka–Flink ingestion, partitioning, and Star-tree indexing to cut query latency to <300 ms, enabling efficient API monitoring and fast issue resolution.
100
Ananth Packkildurai @ananthdurai.bsky.social · 30/09/2025
Netflix Muse: Scaling Analytics at Trillion-Row Scale Netflix evolved its Muse architecture to handle huge datasets efficiently: HyperLogLog sketches, Hollow in-memory feeds, and Druid optimizations cut query latency by ~50% and reduced concurrency load.
100
Ananth Packkildurai @ananthdurai.bsky.social · 29/09/2025
⚡ Latency Every Data Streaming Engineer Should Know “Real-time” has limits—disk, network, and replication delays add up. StreamNative explains latency tiers, common costs, and tuning levers like batching & async processing. 💡 Must-read for data streaming engineers!
100
Reposted by Ananth Packkildurai
Chris @chris.blue · 27/09/2025
I enjoyed this post by @ananthdurai.bsky.social. Does a great job tying a bunch of recent papers and concepts together.
dataengineeringweekly.com
What “Supporting Our AI Overlords” and “Semantic Spacetime” Tell Us About the Future of Data Infrastructure
Connecting agent-first and universal semantic grammar to reimagine data infrastructure beyond the relational model.
071
Ananth Packkildurai @ananthdurai.bsky.social · 27/09/2025
MCP (Model Context Protocol) promises a new way for LLMs to use tools. Chris Riccomini argues it mostly reinvents OpenAPI, gRPC & CLIs. Resources = docs Tools = RPC Prompts = configs So… could MCP have just been a JSON file? 💡 More insights: www.dataengineeringw...
132
Ananth Packkildurai @ananthdurai.bsky.social · 26/09/2025
How Tables Got Smarter: Iceberg → DuckLake. From static snapshots to stream-native updates and catalog-first metadata, tables are evolving fast. Choose by intent, not hype. Subscribe → www.dataengineeringw... Full story → medium.com/fresha-da...
030
Ananth Packkildurai @ananthdurai.bsky.social · 25/09/2025
I wrote my thoughts on Supporting Our AI Overlords.
dataengineeringweekly.com
What “Supporting Our AI Overlords” and “Semantic Spacetime” Tell Us About the Future of Data Infrastructure
Connecting agent-first and universal semantic grammar to reimagine data infrastructure beyond the relational model.
040
Ananth Packkildurai @ananthdurai.bsky.social · 25/09/2025
How Tables Grew a Brain: Iceberg → DuckLake Snapshots → incremental → stream-native → catalog-first. Metadata is the bottleneck. More insights → www.dataengineeringw... Full story → medium.com/fresha-da...
020
Ananth Packkildurai @ananthdurai.bsky.social · 24/09/2025
BlaBlaCar scales like a pro! dbt Core → Transform like a champ Airflow → Orchestrate effortlessly CI/CD → Deploy instantly Dev Containers → Standardized dev 📖 Full story →medium.com/blablacar... 💡 More insights → Subscribe to DEW #DataEngineering #dbt #Airflow #CICD #DevContainers
010
Ananth Packkildurai @ananthdurai.bsky.social · 23/09/2025
🚀 AI adoption is booming—but most data isn’t ready! AI-ready data is: Unified Real-time Human-verified Governed Without it, AI can confidently fail. With it? Reliable, scalable results. 📖 Read More 💡 More insights → Data Engineering Weekly #AI #AIReady #DataEngineering
000
Ananth Packkildurai @ananthdurai.bsky.social · 22/09/2025
Stripe’s Real-Time Billing Analytics ⚡ Content: Stripe wanted real-time visibility into subscriptions. Traditional batch systems weren’t fast enough. ⏱️ They built a pipeline using Flink, Spark, and Pinot v2. Now, analytics arrive in minutes, not hours. Queries return in <300ms. 🚀
121
Ananth Packkildurai @ananthdurai.bsky.social · 22/09/2025
The 238th edition of Data Engineering Weekly is available, featuring exciting Data & AI articles. Read more: www.dataengineeringw...
150
Ananth Packkildurai @ananthdurai.bsky.social · 18/09/2025
Apache Iceberg is now entering the classic paradox. Reference: www.dataengineeringw... www.warpstream.com/b...
081
Ananth Packkildurai @ananthdurai.bsky.social · 16/09/2025
open.substack.com/pub/dataengi...
open.substack.com
When Dimensions Change Too Fast for Iceberg
Why Iceberg Struggles with Fast-Changing Dimensions—and What Comes Next
020
Ananth Packkildurai @ananthdurai.bsky.social · 16/09/2025
From Firefighting to Proactive DB Reliability 🚀 Databricks engineers used Databricks to revolutionize DB reliability: Query/Schema Scorer in CI pipelines Delta Tables + DLT pipelines Database Usage Scorecard across thousands of DBs Efficiency ✅ Anti-patterns ❌
100
Ananth Packkildurai @ananthdurai.bsky.social · 09/09/2025
Parquet paradox: supports pluggable indexing and bloom filters, but you must rewrite entire files to use them. Meanwhile, LanceDB rebuilds indexes independently. Is "self-contained" showing off its age to "composable" data architectures? 🤔
020
Ananth Packkildurai @ananthdurai.bsky.social · 09/09/2025
AI Hallucinations = confident answers that are flat-out wrong. Why they happen 👇 🎯 Training rewards sounding right, not being right 🎲 Guessing > “I don’t know” 📉 Missing data → confident fiction The fix? Retrieval grounding + truth-focused training.
100
Ananth Packkildurai @ananthdurai.bsky.social · 03/09/2025
For years, I've been a skeptic of the Medallion Architecture naming convention. However, despite my initial reservations about the naming, I have come to appreciate the value of a bounded definition. I've shared my thoughts on Medallion Architecture.
dataengineeringweekly.com
Revisiting Medallion Architecture
The Evolution and Context of the Medallion Architecture
042
Ananth Packkildurai @ananthdurai.bsky.social · 29/08/2025
Building a Search Engine at Scale 3 B embeddings. 2 months. From content parsing to vector indexing. Wilson Lin shares how—and why chunking is modeling. 📖 www.dataengineeringw... 💡 Subscribe → www.dataengineeringw...
020
Ananth Packkildurai @ananthdurai.bsky.social · 28/08/2025
Netflix is redefining data engineering. With LanceDB + Media ML, the Lakehouse now powers media intelligence, not just metrics. 📖 netflixtechblog.com 💡 Subscribe: dataengineeringweekl...
010