Ananth Packkildurai @ananthdurai.bsky.social · 13/04/2026I just realized: if Airflow and other orchestration engines emit open-lineage data to S3 Files and enable Claude to search them, you've got a data catalog. 230
Ananth Packkildurai @ananthdurai.bsky.social · 13/04/2026So, are we debating the semantic layer again? docs.getdbt.com/blog... What do you call a semantic layer from your perspective? docs.getdbt.comSemantic Layer vs. Text-to-SQL: 2026 Benchmark Update | dbt Developer BlogWith 2026's best models, the dbt Semantic Layer hits near-100% accuracy for covered queries. Here's what changed and what didn't in our updated benchmark. 050
Ananth Packkildurai @ananthdurai.bsky.social · 13/04/2026Whether you like or dislike Apache Kafka, its KIPs are among the best learning materials for distributed systems. KIP-848 is an excellent read cwiki.apache.org/con... 040
Ananth Packkildurai @ananthdurai.bsky.social · 02/04/2026Most data platform failures don’t start with bad infra. They start at the team boundary. My new post argues platforms scale through operating interfaces: contracts, ownership, communication, and adoption design, not tooling alone.dataengineeringweekly.comThe Missing Interface in Data Platform EngineeringHow data leaders should design the boundary between platforms and dependent teams. 030
Ananth Packkildurai @ananthdurai.bsky.social · 11/03/2026More ETL pipelines will run next year than ever before. And ETL is still dead. Not dead, like nobody uses it. Dead like landlines — they work, but nobody builds their strategy around one.dataengineeringweekly.comETL is DeadWhy the shift from human-operated to agent-operated data warehouses demands a new architecture 010
Ananth Packkildurai @ananthdurai.bsky.social · 24/02/2026The data engineer job title is due for an update. Not because AI is replacing the role, but because AI is finally revealing what the role was always actually about. Moving data was never the point. Meaning it is. Read more:dataengineeringweekly.comData Engineering After AIMoving Data Was Never the Point. Meaning It Is. 010
Ananth Packkildurai @ananthdurai.bsky.social · 21/02/2026At the end of 2026, we will talk about "AI Fan Effect" [en.wikipedia.org/wik...] and the invention of a new field: Psychology for AI. Perhaps, I feel this is the future of software engineering. 010
Ananth Packkildurai @ananthdurai.bsky.social · 31/01/2026As we move from dashboards to autonomous agents, something breaks. Systems of record capture what happened, not why. Why data platforms need Truth Registries + Context Graphs for the agentic era 👇 www.dataengineeringw... #DataEngineering #AgenticAI #Graphs #LLMsdataengineeringweekly.comThe Missing Layer in Your AI Stack: Context, Not Just StateFrom SQL to Semantics: The Rise of the Context Graph for AI Agents 030
Ananth Packkildurai @ananthdurai.bsky.social · 26/01/2026Data Engineering Weekly's 254th edition is out. Context Graph is the new talk of the town!! 000
Ananth Packkildurai @ananthdurai.bsky.social · 23/01/2026The companies that build the most boring data stack often win the market!!! Prove me wrong. 010
Ananth Packkildurai @ananthdurai.bsky.social · 20/01/2026Data Contract: There was no shortage of activity around the topic. Definitions were proposed and refined. Conceptual boundaries were drawn and redrawn. I pen down a reflection of the Data Contracts here www.dataengineeringweekly.com/p/data-contr...dataengineeringweekly.comData Contracts: A Missed OpportunityThe Conversation We Should Have Had—Before Thought Leadership Replaced System Design 010
Ananth Packkildurai @ananthdurai.bsky.social · 14/01/2026How to build a scalable shopping agent? Here's a wild thought: What if—and hear me out—we let humans click that Buy Now button? Just throwing ideas out there. 010
Ananth Packkildurai @ananthdurai.bsky.social · 12/01/2026This week, it is mostly about Multi-Agent Architecture. Do you think the data infrastructure is ready for a multi-agent architecture? Where is the gap?dataengineeringweekly.comData Engineering Weekly #252The Weekly Data Engineering Newsletter 010
Ananth Packkildurai @ananthdurai.bsky.social · 09/01/2026Is semantic Spec Good enough to run an enterprise system? I listed challenges to adopting the Iceberg Rest Catalogdataengineeringweekly.comA Critique of Iceberg REST Catalog: A Classic Case of Why Semantic Spec FailsHow a Semantically Correct API Becomes Operationally Unreliable at Scale 000
Ananth Packkildurai @ananthdurai.bsky.social · 23/12/2025Continuing our yearly tradition of Year in Review Data Engineering Weekly, we published the 2025 Year in Review. What do you think is the most notable trend of 2025?dataengineeringweekly.comDEW - The Year in Review 2025From Digital Plumbers to Architects of Intelligence: The 7 Paradigm Shifts That Defined 2025 000
Ananth Packkildurai @ananthdurai.bsky.social · 08/12/2025Look at the tech stack IBM now controls: 🐧 Compute: Red Hat (Linux/OpenShift) ☁️ IaC: HashiCorp (Terraform) 💰 FinOps: Kubecost 🌊 Streaming: Confluent (Kafka) 🧠 Vector/AI: DataStax (Cassandra) ⚡ Query Engine: Ahana (Presto) 🔄 Ingest: StreamSets 140
Ananth Packkildurai @ananthdurai.bsky.social · 08/12/2025LinkedIn moves FishDB to Rust, DoorDash builds AI swarms, and Dropbox masters context engineering. 🤯 Data Engineering Weekly #247 is packed with system design deep dives from the best engineering teams.dataengineeringweekly.comData Engineering Weekly #247The Weekly Data Engineering Newsletter 030
Ananth Packkildurai @ananthdurai.bsky.social · 04/12/2025If the Data Catalog is the answer for AI, the question was wrong. 010
Ananth Packkildurai @ananthdurai.bsky.social · 19/11/2025We stopped asking if data was useful because storage got cheap. Now, "Dark Data" is actively poisoning your AI context windows with hallucination vectors. Read about the Data Sustainability indexdataengineeringweekly.comThe Dark Data Tax: How Hoarding is Poisoning Your AIStorage is cheap. Attention is finite. Hallucinations are expensive. It’s time to stop building Data Lakes and start managing Data Metabolism 050
Ananth Packkildurai @ananthdurai.bsky.social · 10/11/2025The open source companies built their success on top of open-source platforms, benefited from community contributions and adoption, but now must abandon open-source principles to survive commercially. 010
Ananth Packkildurai @ananthdurai.bsky.social · 03/11/2025🚀 The 244th edition of Data Engineering Weekly dives into: AI agents as execution engines, LLM inference economics, databases for AI, personalization, and product evidence. Read more 👉 www.dataengineeringw... #DataEngineering #AI #LLMsdataengineeringweekly.comData Engineering Weekly #244The Weekly Data Engineering Newsletter 020
Ananth Packkildurai @ananthdurai.bsky.social · 03/11/2025Cricket has been India’s greatest force in overcoming centuries of colonial suppression. Today’s Women’s World Cup win echoes the spirit of 1983 — a triumph that will inspire generations to come. 🇮🇳🏆 000
Ananth Packkildurai @ananthdurai.bsky.social · 23/10/2025This is the most personal essay that I have written in Data Engineering Weekly. I shared a few key moments in my life and how fortunate I was to meet mentors along my professional journey, which shaped my career.dataengineeringweekly.comThinking Like a Data EngineerA Journey Beyond Code — Toward Systems, Curiosity, and Confidence 090
Ananth Packkildurai @ananthdurai.bsky.social · 17/10/2025🚀 Data Vault vs. Dimensional Modeling vs. Medallion Architecture — When viewed through a modern enterprise data lens, these techniques interlock. I break down how in Part 2 of my “Revisiting the Medallion Architecture” series.dataengineeringweekly.comRevisiting Medallion Architecture: Data Vault in Silver, Dimensional Modeling in GoldHow to Balance Flexibility and Performance in a Modern Data Platform 040
Ananth Packkildurai @ananthdurai.bsky.social · 17/10/2025Fivetran and dbt form a strong foundation for modern data infrastructure, known for bringing simplicity to complex engineering workflows. That said, calling it “open” data infrastructure feels like a stretch. 350
Ananth Packkildurai @ananthdurai.bsky.social · 13/10/2025Should we update the definition of an "Analytical Engineer"? 040
Ananth Packkildurai @ananthdurai.bsky.social · 09/10/2025As a data engineer, you can't treat zero-party (consent) and third-party (inferred) data the same way. This distinction is critical for building systems that are scalable, private, and trustworthy. Here’s my guide:dataengineeringweekly.comEngineering Growth: The Data Layers Powering Modern GTMBuilding privacy-preserving pipelines that unify zero-, first-, second-, third-, and fourth-party data into a coherent GTM ecosystem. 050
Ananth Packkildurai @ananthdurai.bsky.social · 02/10/2025Airbnb: Real-Time Key-Value Store Airbnb’s next-gen key-value store supports real-time ingestion and bulk uploads with sub-second latency, powering feature stores and fraud detection. Read the full story here: www.dataengineeringw... 010
Ananth Packkildurai @ananthdurai.bsky.social · 01/10/2025Grab: Partner Gateway Metrics at Sub-Second Speed Real-time partner analytics at scale is tough. Grab uses Apache Pinot, Kafka–Flink ingestion, partitioning, and Star-tree indexing to cut query latency to <300 ms, enabling efficient API monitoring and fast issue resolution. 100
Ananth Packkildurai @ananthdurai.bsky.social · 30/09/2025Netflix Muse: Scaling Analytics at Trillion-Row Scale Netflix evolved its Muse architecture to handle huge datasets efficiently: HyperLogLog sketches, Hollow in-memory feeds, and Druid optimizations cut query latency by ~50% and reduced concurrency load. 100
Ananth Packkildurai @ananthdurai.bsky.social · 29/09/2025⚡ Latency Every Data Streaming Engineer Should Know “Real-time” has limits—disk, network, and replication delays add up. StreamNative explains latency tiers, common costs, and tuning levers like batching & async processing. 💡 Must-read for data streaming engineers! 100
Reposted by Ananth PackkilduraiChris @chris.blue · 27/09/2025I enjoyed this post by @ananthdurai.bsky.social. Does a great job tying a bunch of recent papers and concepts together.dataengineeringweekly.comWhat “Supporting Our AI Overlords” and “Semantic Spacetime” Tell Us About the Future of Data InfrastructureConnecting agent-first and universal semantic grammar to reimagine data infrastructure beyond the relational model. 071
Ananth Packkildurai @ananthdurai.bsky.social · 27/09/2025MCP (Model Context Protocol) promises a new way for LLMs to use tools. Chris Riccomini argues it mostly reinvents OpenAPI, gRPC & CLIs. Resources = docs Tools = RPC Prompts = configs So… could MCP have just been a JSON file? 💡 More insights: www.dataengineeringw... 132
Ananth Packkildurai @ananthdurai.bsky.social · 26/09/2025How Tables Got Smarter: Iceberg → DuckLake. From static snapshots to stream-native updates and catalog-first metadata, tables are evolving fast. Choose by intent, not hype. Subscribe → www.dataengineeringw... Full story → medium.com/fresha-da... 030
Ananth Packkildurai @ananthdurai.bsky.social · 25/09/2025I wrote my thoughts on Supporting Our AI Overlords.dataengineeringweekly.comWhat “Supporting Our AI Overlords” and “Semantic Spacetime” Tell Us About the Future of Data InfrastructureConnecting agent-first and universal semantic grammar to reimagine data infrastructure beyond the relational model. 040
Ananth Packkildurai @ananthdurai.bsky.social · 25/09/2025How Tables Grew a Brain: Iceberg → DuckLake Snapshots → incremental → stream-native → catalog-first. Metadata is the bottleneck. More insights → www.dataengineeringw... Full story → medium.com/fresha-da... 020
Ananth Packkildurai @ananthdurai.bsky.social · 24/09/2025BlaBlaCar scales like a pro! dbt Core → Transform like a champ Airflow → Orchestrate effortlessly CI/CD → Deploy instantly Dev Containers → Standardized dev 📖 Full story →medium.com/blablacar... 💡 More insights → Subscribe to DEW #DataEngineering #dbt #Airflow #CICD #DevContainers 010
Ananth Packkildurai @ananthdurai.bsky.social · 23/09/2025🚀 AI adoption is booming—but most data isn’t ready! AI-ready data is: Unified Real-time Human-verified Governed Without it, AI can confidently fail. With it? Reliable, scalable results. 📖 Read More 💡 More insights → Data Engineering Weekly #AI #AIReady #DataEngineering 000
Ananth Packkildurai @ananthdurai.bsky.social · 22/09/2025Stripe’s Real-Time Billing Analytics ⚡ Content: Stripe wanted real-time visibility into subscriptions. Traditional batch systems weren’t fast enough. ⏱️ They built a pipeline using Flink, Spark, and Pinot v2. Now, analytics arrive in minutes, not hours. Queries return in <300ms. 🚀 121
Ananth Packkildurai @ananthdurai.bsky.social · 22/09/2025The 238th edition of Data Engineering Weekly is available, featuring exciting Data & AI articles. Read more: www.dataengineeringw... 150
Ananth Packkildurai @ananthdurai.bsky.social · 18/09/2025Apache Iceberg is now entering the classic paradox. Reference: www.dataengineeringw... www.warpstream.com/b... 081
Ananth Packkildurai @ananthdurai.bsky.social · 16/09/2025open.substack.com/pub/dataengi...open.substack.comWhen Dimensions Change Too Fast for IcebergWhy Iceberg Struggles with Fast-Changing Dimensions—and What Comes Next 020
Ananth Packkildurai @ananthdurai.bsky.social · 16/09/2025From Firefighting to Proactive DB Reliability 🚀 Databricks engineers used Databricks to revolutionize DB reliability: Query/Schema Scorer in CI pipelines Delta Tables + DLT pipelines Database Usage Scorecard across thousands of DBs Efficiency ✅ Anti-patterns ❌ 100
Ananth Packkildurai @ananthdurai.bsky.social · 09/09/2025Parquet paradox: supports pluggable indexing and bloom filters, but you must rewrite entire files to use them. Meanwhile, LanceDB rebuilds indexes independently. Is "self-contained" showing off its age to "composable" data architectures? 🤔 020
Ananth Packkildurai @ananthdurai.bsky.social · 09/09/2025AI Hallucinations = confident answers that are flat-out wrong. Why they happen 👇 🎯 Training rewards sounding right, not being right 🎲 Guessing > “I don’t know” 📉 Missing data → confident fiction The fix? Retrieval grounding + truth-focused training. 100
Ananth Packkildurai @ananthdurai.bsky.social · 03/09/2025For years, I've been a skeptic of the Medallion Architecture naming convention. However, despite my initial reservations about the naming, I have come to appreciate the value of a bounded definition. I've shared my thoughts on Medallion Architecture.dataengineeringweekly.comRevisiting Medallion ArchitectureThe Evolution and Context of the Medallion Architecture 042
Ananth Packkildurai @ananthdurai.bsky.social · 29/08/2025Building a Search Engine at Scale 3 B embeddings. 2 months. From content parsing to vector indexing. Wilson Lin shares how—and why chunking is modeling. 📖 www.dataengineeringw... 💡 Subscribe → www.dataengineeringw... 020
Ananth Packkildurai @ananthdurai.bsky.social · 28/08/2025Netflix is redefining data engineering. With LanceDB + Media ML, the Lakehouse now powers media intelligence, not just metrics. 📖 netflixtechblog.com 💡 Subscribe: dataengineeringweekl... 010