Sign in

Adrian Brudaru

@datateam.bsky.social
2.2K followers 1.3K following 247 posts

Data engineer & Cofounder @dlthub. Building out the tooling i wish i had.

PostsRepliesMedia
Adrian Brudaru @datateam.bsky.social · 10/07/2026
A few weeks ago we launched dltHub Pro. What surprised us was what people built with it. Teams didn't ask for primitives, they asked for complete solutions. So we built Blueprints.
dlthub.com
We are launching two dltHub Blueprints for agent spend: Agent Cost & Usage to understand it, Agent Distillation to optimize it
dltHub launches two Blueprints for agent spend: Agent Cost & Usage to break down what each model, person, and customer costs, and Agent Distillation (with distil labs) to replace expensive agents with cheaper specialist models.
200
Adrian Brudaru @datateam.bsky.social · 07/07/2026
How do you eval your agents? Agent traces are a good place to start. In this 1-hour workshop with @DataTalksClub, you'll learn how to ingest agent traces, model nested JSON into a queryable schema, and build dashboards to understand agent behavior.
youtube.com
Ingesting Agent Traces with dlthub - Alena Astrakhantseva
Links:- https://test-agent-traces-api-xt2e7ottma-ew.a.run.app/doc...
020
Adrian Brudaru @datateam.bsky.social · 06/07/2026
AI writing pipelines is cool. AI remembering the entire workflow is better. dltHub Pro carries context from ingestion to deployment to maintenance instead of starting over every step. Try it: dlthub.com/products/dlthub
010
Adrian Brudaru @datateam.bsky.social · 29/06/2026
AI agents generate traces, but are you analyzing them?  Join Alena Astrakhantseva to learn how to turn tool calls, token usage, and outcomes into structured, queryable data with dltHub Pro. July 6, live on YouTube with DataTalks.Club Register ↓
luma.com
Ingesting Agent Traces with dlthub · Luma
In this hands-on workshop, we'll show you how to stop flying blind on your AI agents. Using dltHub Pro we'll build a pipeline that ingests agent traces (e.g.…
010
Adrian Brudaru @datateam.bsky.social · 25/06/2026
Text-to-SQL doesn't break because models can't write SQL, it breaks because they don't know what your data means. An agent can return valid SQL and still be wrong, and a clean wrong number looks just as trustworthy as a right one.
120
Adrian Brudaru @datateam.bsky.social · 24/06/2026
Human in the loop shouldn't mean copy-pasting context into an agent every 5 minutes. It should mean judgment, not errands. If your agent needs a Rube Goldberg machine of prompts, tabs, and slack messages, the problem isn't human. It's missing context. dlthub.com/blog/context
020
Adrian Brudaru @datateam.bsky.social · 20/06/2026
Gave an LLM a schema: 3/10. 📉 Gave it a schema + an ontology: 10/10. 🎯 Same model, same data. The difference? It finally understood what the columns actually mean instead of just vibing off the names.
220
Adrian Brudaru @datateam.bsky.social · 18/06/2026
91% of the 81,000 new dlt pipelines shipped in January 2026 were built by agents. The bottleneck in data engineering is no longer implementation, it's meaning.
110
Adrian Brudaru @datateam.bsky.social · 09/06/2026
How much data can $1 of compute move? We benchmarked dltHub on a small worker (2 vCPU / 4 GB) loading into BigQuery: Parquet: ~170 GB Postgres: ~65 GB JSON: ~4.6 GB REST: whatever the API allows Methodology + results ↓ dlthub.com/blog/benchmark-dlthub
020
Adrian Brudaru @datateam.bsky.social · 05/06/2026
Data quality usually starts too late. By the time bad data hits a dashboard, the original business assumptions are gone. We're fixing that at the ingestion layer. 👇
110
Adrian Brudaru @datateam.bsky.social · 01/06/2026
Agentic Analytics Demo Night lands in San Francisco this Wednesday 🚀 We're joining @inngest.com, lightdash, and @streamkap.bsky.social for an evening of demos from teams building at the intersection of AI, analytics, and data infrastructure. 🧵
110
Adrian Brudaru @datateam.bsky.social · 01/06/2026
What does it take to build a product people actually keep using? On Wednesday, we're hosting Sameer Alsakran, Founder & CEO of @metabase.com, at our office in Berlin for a live conversation with Francesco Mucio from Data Berlin. 🧵
130
Adrian Brudaru @datateam.bsky.social · 26/05/2026
Berlin’s Applied AI Week is looking pretty special. Tomorrow, dltHub and @modal-labs.bsky.social are hosting an evening in Berlin for founders, engineers, and builders shipping agents in production. Featuring Kenny Ning, Jefferson Girao & Alena Astrakhantseva.
110
Adrian Brudaru @datateam.bsky.social · 18/05/2026
The AI stack is evolving fast, but reliable data movement is still the foundation. Join dltHub, LanceDB and DataHub on May 21 in Menlo Park for talks on multimodal AI storage, AI data pipelines, and trusted lineage systems. 🔗
luma.com
The missing data layer for ML: dltHub x LanceDB x DataHub @ SVAI · Luma
Modern problems = Modern solutions! Join dltHub, LanceDB, and DataHub for a night of technical talks and demos. Hear from the engineers building the ingestion,…
010
Adrian Brudaru @datateam.bsky.social · 15/05/2026
How fast can you go from zero to a production-ready data pipeline when AI is your copilot? Next Wednesday, Elvis Kahoro joins @temporal.io alongside Melanie Warrick and @cecilphillip.bsky.social for a live Vibe Check building a GitHub-powered pipeline with AI + dlt.
111
Adrian Brudaru @datateam.bsky.social · 14/05/2026
Building pipelines with AI usually means losing context between tools. dltHub AI Workbench runs the full 12-step workflow as a continuous session, schemas, incrementals, traces, transformations, and notebooks share context across the stack. dlthub.com/blog/agentic-data-engine…
020
Adrian Brudaru @datateam.bsky.social · 12/05/2026
Explainer on ontology engineering and what we're building around it: why just clean schemas and prompts aren’t enough, and how adding a canonical model + taxonomy + ontology changes what agents can correctly compute (ARPU being the clearest example). 
dlthub.com
Ontology engineering: what it is, why it's back, and why agents need it
Agents don't hallucinate. They navigate without a map. Ontology engineering is how you build one, and why every team pulling humans out of the loop needs it now.
010
Adrian Brudaru @datateam.bsky.social · 30/04/2026
We’re excited to share that Violetta Mishechkina will be speaking at GOSIM Paris 🇫🇷 Invited by probabl.ai to join the “Own Your Data Science and AI” workshop. 🎤 From Agent Traces to Analytics Agents generate code, text, telemetry, yet most teams still rely on stale datasets.
110
Adrian Brudaru @datateam.bsky.social · 25/04/2026
Skills that wrap a library are software. They have dependencies, need maintenance, and degrade when the product changes and no one updates them. The vendor owns the product surface. You own your integration. Same rule, new layer.
120
Adrian Brudaru @datateam.bsky.social · 24/04/2026
1/ dlt is now available as a Snowflake Native App. Replicate MSSQL, MySQL & PostgreSQL → Snowflake, without leaving Snowflake. Create, schedule, and monitor pipelines from the Snowflake UI. No external orchestrator. app.snowflake.com/marketplace/listi…
110
Adrian Brudaru @datateam.bsky.social · 15/04/2026
Every layer of software automation was called overkill before it became the baseline. Fortran → Make → CI/CD → Docker → now agents. Code that runs is only the 10%. The other 90% is engineering judgment, boundaries, and iteration.
110
Adrian Brudaru @datateam.bsky.social · 14/04/2026
AI agents don’t just use your APIs, they optimize around them. Ask an agent to “build a pipeline” and it will find credentials, escalate privileges, and take the shortest path to completion. Not against your interests, just goal-driven.
210
Adrian Brudaru @datateam.bsky.social · 09/04/2026
Not everything that can be modeled should be. With LLMs, more context doesn’t mean a better prompt. The key is Minimum Viable Context for high-precision data models. Here’s what we learned building ontology-driven modeling. Blog by Hiba Jamal ↓
dlthub.com
Minimum Viable Context for Building a Canonical Data Model
Call it the MVC problem: minimum viable context. Too little and it hallucinates your domain. Too much and it drifts from your actual goal. The process has to be controlled.
010
Adrian Brudaru @datateam.bsky.social · 03/04/2026
AI can generate a data pipeline in 10 minutes. But can you trust what it produces? That’s the real problem, and it’s not technical. It’s business. If you can’t trust the output, it never reaches production. 🧵
110
Adrian Brudaru @datateam.bsky.social · 02/04/2026
Most AI coding tools stop at “here’s your code.” But getting pipelines into production, and trusting them there, is the hard part. We built a deployment toolkit that closes that gap 🧵
110
Adrian Brudaru @datateam.bsky.social · 31/03/2026
Most teams still build connectors from scratch, one at a time. Different patterns. Different implementations. Accumulating tech debt. What if you built the system instead? We just released dlt Skills, a pipeline factory powered by Claude. 👇
110
Adrian Brudaru @datateam.bsky.social · 30/03/2026
LLMs fail at data transformation because they see isolated tables, not your business. We built the dltHub AI Workbench transformation toolkit to fix that. You feed it sources + use cases → it builds a taxonomy, business ontology, and a Canonical Data Model.
120
Adrian Brudaru @datateam.bsky.social · 26/03/2026
The craziest part of the new dltHub AI release? The MCP integration. Asked Claude Code for an OpenAI pipeline -> it searches the dlt context -> scaffolds the exact code with schema & incremental loading. No more starting from scratch. dlthub.com/blog/ai-workbench
150
Adrian Brudaru @datateam.bsky.social · 26/03/2026
New course: Agentic Data Engineering with dltHub 🤖 Agents can now write entire data pipelines, but writing code was never the hard part. The real challenge? Data quality, schema stability, and running reliably in production. 🧵👇
110
Adrian Brudaru @datateam.bsky.social · 24/03/2026
Agents can generate data pipelines. The real problem now is trusting them in production. Today we’re introducing the dltHub AI Workbench, infrastructure for generating, validating, and deploying pipelines in one workflow.
110
Adrian Brudaru @datateam.bsky.social · 22/03/2026
WAP (Write → Audit → Publish) has a blind spot. It assumes your data can safely land in staging. But in real pipelines, ingestion is often where things break. That’s where AWAP comes in.
120
Adrian Brudaru @datateam.bsky.social · 19/03/2026
Your production traces are gold, but they aren’t training data yet. We turned raw agent traces into a specialist model using: dlt → @hf.co → Distil Labs In our example, a 0.6B model beat its 120B teacher by 28 points on a specific task. Here’s how it works ↓
110
Adrian Brudaru @datateam.bsky.social · 19/03/2026
Curious what people here are actually using for AI coding 👀 Copilot? Cursor? Claude Code? We put together a super short (1-min) survey. 👉 dlthub.notion.site/3039fb8e23cf8048…
130
Adrian Brudaru @datateam.bsky.social · 17/03/2026
Microsoft Fabric is great at compute & storage but data quality enforcement is on you. Use WAP to validate data before it hits the lakehouse. dlt handles schemas, business rules, uniqueness, PII, and monitoring so bad data never reaches analytics.
dlthub.com
Building production-ready data pipelines in Microsoft Fabric: A complete data quality framework with dlthub
Add data quality gates to Microsoft Fabric with dlt. Validate schemas, catch bad records, and mask PII before data reaches your lakehouse and downstream analytics.
010
Adrian Brudaru @datateam.bsky.social · 16/03/2026
LLMs follow the gravity of your vocabulary, not your business logic. That's why your AI-generated stack looks great in the demo and breaks silently in prod.
dlthub.com
So you vibe coded a data stack, now what?
In this blog post, I describe the actual, hard real world barriers that make your LLM setup collapse, and propose principles for making your systems work.
010
Adrian Brudaru @datateam.bsky.social · 15/03/2026
❄️ Module 2 is live in our course dlt + Snowflake  Learn how to run dlt pipelines inside Snowflake using Snowpark Container Services (SPCS), enabling native execution and scheduling with no external infrastructure required. Continue the course: dlthub.learnworlds.com/course/dlt-s… ↓
110
Adrian Brudaru @datateam.bsky.social · 15/03/2026
We shipped a @hf.co Face datasets destination in dlt. It makes it easier to move training data between production systems and the HF Hub. Think pipelines like: raw traces → dlt → versioned datasets on Hugging Face → model training ↓
101
Adrian Brudaru @datateam.bsky.social · 13/03/2026
Small data teams deserve better tools. We’re opening early design partnerships for solo and small data teams to try dltHub Pro before launch. Early access, influence the roadmap, and an early-bird discount.  dlthub.com/solutions/for-small-data…
010
Adrian Brudaru @datateam.bsky.social · 13/03/2026
Prototyping data pipelines in @duckdb.org is great until you need to ship them to production. On Mar 16, Elvis Kahoro & @joshleecreates.bsky.social from Altinity will show how to move a DuckDB pipeline to @clickhouse.com using dlt, fully defined as Python code. Thread ↓
121
Adrian Brudaru @datateam.bsky.social · 08/03/2026
Spring brings good things and more than just flowers. 🌸 ❄️ Catch up on Module 1 of our dlt + Snowflake course before Module 2 drops next week! Learn nested data normalization, schema evolution, incremental loading, and merge strategies (upsert, SCD2), all in plain Python.
150
Adrian Brudaru @datateam.bsky.social · 05/03/2026
Tasman.ai runs data engineering projects for mid-market and enterprise clients. Their biggest challenge? Scoping. Every new client meant figuring out which APIs to connect, how long it would take, and what the data actually looked like, often before seeing a single row.
100
Adrian Brudaru @datateam.bsky.social · 26/02/2026
Data models describe the data, ontologies describe the world. With ontologies, an agent can reason over data as opposed to retrieving it and hallucinating meaning. The ontology-model mapping is what agents need for data literacy. This is NOW. Blog + demo
dlthub.com
Ontology driven Dimensional Modeling
To understand how to answer world questions from data models, we don't need semantic layers, we need ontologies
000
Adrian Brudaru @datateam.bsky.social · 22/02/2026
Who is the UFC GOAT? 🥊📊 We turned that curiosity into a full-stack pipeline analyzing 30+ years of UFC fights. Everything programmatic, even dashboard creation via API. Production-grade insights. Full traceability. Zero manual overhead.
110
Adrian Brudaru @datateam.bsky.social · 20/02/2026
Part 2 of our RAG debug series is out. We froze retrieval, prompts & dataset, and tested newer models only. Result: 3/14 → ~10/14 correct answers. 3× improvement just by upgrading the model. Same system. Same eval set. Retrieval is next. Read more 👇
dlthub.com
Debugging Our Docs RAG, Part 2: Testing New Generation
By upgrading only the generative model, we achieved a 3x accuracy boost but hit a hard ceiling, proving that not only LLMs are needed for good retrieval.
000
Adrian Brudaru @datateam.bsky.social · 19/02/2026
AI agents fail without memory & context. @cognee.bsky.social turns data into self-improving, structured memory at scale. If you’ve built a modern data stack (ingest → transform → access), you already know the pattern. Backed by pebblebed, congrats on the $7.5M.
dlthub.com
AI Memory: Understanding Modeling for Unstructured Data
For the data engineering crowd, here’s an explainer of how unstructured AI memory works, though the lens of what we know from working with structured data.
001
Adrian Brudaru @datateam.bsky.social · 16/02/2026
Ingestion shouldn’t be a maintenance trap. From Airbyte to dlt in one week. Slides 👇 docs.google.com/presentation/d/e/2P…
110
Adrian Brudaru @datateam.bsky.social · 12/02/2026
From APIs to Warehouses 📦 On Feb 17 (16:30 CET), together with DataTalks.Club, Aashish Nair will walk through building end-to-end ingestion pipelines with dlt, from raw APIs to production-ready warehouse loads. Register here 👇
luma.com
From APIs to Warehouses: AI-Assisted Data Ingestion with dlt · Luma
This hands-on workshop focuses on building reliable data ingestion pipelines to data warehouses (for example, Snowflake) using dlt (data load tool), enhanced…
100
Adrian Brudaru @datateam.bsky.social · 12/02/2026
What if dimensional modeling didn’t mean hours of boilerplate SQL? We built an AI workflow that turns raw data into semantic models in minutes, powered by 20 questions. Rethinking data transformation 👇
dlthub.com
The Last Mile is Solved by Slop
I didn't vibe-build a product. I wrote a messy scaffold that runs a pipeline, grabs the schema, and forces an agent to build a star schema. It works shockingly well.
020
Adrian Brudaru @datateam.bsky.social · 11/02/2026
Berlin, it’s meetup time! Join us for the dltHub Community Meetup, an evening of real-world demos, lessons learned, and conversations with builders. 📍 Rosebud, Berlin 📆 Feb 17 | 18:00 – 21:00 Curious about what we’re building at dltHub? Come by 👋
luma.com
dltHub Community Meetup in Berlin with Cognee, Untitled Data Company, Gemma Analytics & Babbel · Luma
Join us for the dltHub Community Meetup in Berlin. This evening is for curious minds who want to learn more about what we’re building at dltHub. We’ll share a…
111
Adrian Brudaru @datateam.bsky.social · 10/02/2026
Production pipelines don’t fail loudly, they drift. Feb 12 · 16:00 CET - Online Hands-on workshop on operating pipelines in production: • schema changes • backfills • CI/CD • long-term reliability Register → community.dlthub.com/workshop-maint…
100