Sign in

Adrian Brudaru

@datateam.bsky.social
2.2K followers 1.3K following 247 posts

Data engineer & Cofounder @dlthub. Building out the tooling i wish i had.

PostsRepliesMedia
Adrian Brudaru @datateam.bsky.social · 10/07/2026
Agent Distillation. Same traces, different outcome. Turn them into training-ready datasets for smaller, task-specific models, built together with distil labs. Browse the Blueprints:
dlthub.com
dltHub Blueprints
dltHub is a composable data platform. Blueprints are its ready-made models: each one dltHub assembled for a specific use case, end to end, from the sources you already use to a production dashboard or API.
020
Adrian Brudaru @datateam.bsky.social · 10/07/2026
Agent Cost & Usage. Standardize traces, join them with cost APIs, and answer the question leadership actually asks: which team, customer, agent, or model is driving the bill?
110
Adrian Brudaru @datateam.bsky.social · 10/07/2026
Agent traces have no standard. Every framework invents its own format, they change constantly. So every team starts by building from scratch. A Blueprint gives you: ingest → transform → dashboard Adapt it instead of assembling it. The first two:
120
Adrian Brudaru @datateam.bsky.social · 10/07/2026
AI agents created a new infra problem: what are they doing, and what are they costing us? The answer is buried in agent traces, and traces are messy.
130
Adrian Brudaru @datateam.bsky.social · 10/07/2026
A few weeks ago we launched dltHub Pro. What surprised us was what people built with it. Teams didn't ask for primitives, they asked for complete solutions. So we built Blueprints.
dlthub.com
We are launching two dltHub Blueprints for agent spend: Agent Cost & Usage to understand it, Agent Distillation to optimize it
dltHub launches two Blueprints for agent spend: Agent Cost & Usage to break down what each model, person, and customer costs, and Agent Distillation (with distil labs) to replace expensive agents with cheaper specialist models.
200
Adrian Brudaru @datateam.bsky.social · 07/07/2026
How do you eval your agents? Agent traces are a good place to start. In this 1-hour workshop with @DataTalksClub, you'll learn how to ingest agent traces, model nested JSON into a queryable schema, and build dashboards to understand agent behavior.
youtube.com
Ingesting Agent Traces with dlthub - Alena Astrakhantseva
Links:- https://test-agent-traces-api-xt2e7ottma-ew.a.run.app/doc...
030
Adrian Brudaru @datateam.bsky.social · 06/07/2026
AI writing pipelines is cool. AI remembering the entire workflow is better. dltHub Pro carries context from ingestion to deployment to maintenance instead of starting over every step. Try it: dlthub.com/products/dlthub
010
Adrian Brudaru @datateam.bsky.social · 29/06/2026
AI agents generate traces, but are you analyzing them?  Join Alena Astrakhantseva to learn how to turn tool calls, token usage, and outcomes into structured, queryable data with dltHub Pro. July 6, live on YouTube with DataTalks.Club Register ↓
luma.com
Ingesting Agent Traces with dlthub · Luma
In this hands-on workshop, we'll show you how to stop flying blind on your AI agents. Using dltHub Pro we'll build a pipeline that ingests agent traces (e.g.…
010
Adrian Brudaru @datateam.bsky.social · 25/06/2026
The common fix is a semantic layer added on top. Now you maintain it twice and the two drift. We build from the other end: write the canonical knowledge layer first, use that one spec to generate the model and answer questions over it.
dlthub.com
Text-to-SQL is a definition problem: build the canonical model first
Text-to-SQL doesn’t break because models can’t write SQL — it breaks because they don’t know what your data means. Write the meaning down first as a canonical knowledge layer, and use that one spec to both build the model and answer questions over it.
010
Adrian Brudaru @datateam.bsky.social · 25/06/2026
Text-to-SQL doesn't break because models can't write SQL, it breaks because they don't know what your data means. An agent can return valid SQL and still be wrong, and a clean wrong number looks just as trustworthy as a right one.
120
Adrian Brudaru @datateam.bsky.social · 24/06/2026
Human in the loop shouldn't mean copy-pasting context into an agent every 5 minutes. It should mean judgment, not errands. If your agent needs a Rube Goldberg machine of prompts, tabs, and slack messages, the problem isn't human. It's missing context. dlthub.com/blog/context
020
Adrian Brudaru @datateam.bsky.social · 20/06/2026
The catch: on a public dataset, schema-only scored 8/10. Looked grounded, but it was guessing from training data, not reading the pipeline. 🚨 Full benchmark by Roshni Melwani (60 responses + repo): 🔗
dlthub.com
Why LLMs Get the Right Answer for the Wrong Reason
Schema alone scored 3/10. An ontology scored 10/10. A benchmark across two datasets showing exactly where the gap is.
020
Adrian Brudaru @datateam.bsky.social · 20/06/2026
The clearest case: asked if 55% at-risk seats = churned, the schema-only model said no, but guessed. It didn't know the real threshold was 60%. Ask about 65% and it'd still say no. The ontology model said no because 55% < 60%. That reasoning generalizes. 👀
110
Adrian Brudaru @datateam.bsky.social · 20/06/2026
Gave an LLM a schema: 3/10. 📉 Gave it a schema + an ontology: 10/10. 🎯 Same model, same data. The difference? It finally understood what the columns actually mean instead of just vibing off the names.
220
Adrian Brudaru @datateam.bsky.social · 18/06/2026
The role doesn’t disappear, it recomposes. The hard part shifts from writing pipelines to making business knowledge explicit & structured, so an agent can build and a team can verify what it built. Full piece
dlthub.com
The rise of the Semantic engineer
As pipelines, models, and dashboards become generated and disposable, the scarce input shifts from technical skill to business meaning, and a new role emerges
010
Adrian Brudaru @datateam.bsky.social · 18/06/2026
Agents can generate pipelines, models, and dashboards. They can’t generate what a customer is, when an order counts as revenue, why a definition excludes what it excludes, or where historical breaks in your systems are. Most of this is still undocumented, in people’s heads.
120
Adrian Brudaru @datateam.bsky.social · 18/06/2026
91% of the 81,000 new dlt pipelines shipped in January 2026 were built by agents. The bottleneck in data engineering is no longer implementation, it's meaning.
110
Adrian Brudaru @datateam.bsky.social · 09/06/2026
How much data can $1 of compute move? We benchmarked dltHub on a small worker (2 vCPU / 4 GB) loading into BigQuery: Parquet: ~170 GB Postgres: ~65 GB JSON: ~4.6 GB REST: whatever the API allows Methodology + results ↓ dlthub.com/blog/benchmark-dlthub
020
Adrian Brudaru @datateam.bsky.social · 05/06/2026
The new dltHub AI Workbench Data Quality Toolkit starts from context your pipeline already knows: schema contracts, keys, constraints, and sampled values. Plain-language business rules → checks that run on every load.
dlthub.com
dltHub AI Workbench data quality toolkit: schema-aware checks that route their own fixes
Preview of the dltHub AI Workbench data quality toolkit: schema-bootstrapped checks, column sampling before any rule ships, decorators that run inside pipeline.run(), and routing of failures back to the toolkit that owns the surface area.
020
Adrian Brudaru @datateam.bsky.social · 05/06/2026
Data quality usually starts too late. By the time bad data hits a dashboard, the original business assumptions are gone. We're fixing that at the ingestion layer. 👇
110
Adrian Brudaru @datateam.bsky.social · 01/06/2026
At dltHub, Elvis Kahoro will demo our new Transformations public preview. Expect demos from practitioners and startups building real-world AI and data systems, and a look at what agentic analytics looks like in practice. 📅 June 3 · SF
luma.com
Agentic Analytics Demo Night · Luma
A one-night event for those building at the cutting edge of AI, data, and infrastructure - demos, discussions and data people! Made for and by data…
010
Adrian Brudaru @datateam.bsky.social · 01/06/2026
Agentic Analytics Demo Night lands in San Francisco this Wednesday 🚀 We're joining @inngest.com, lightdash, and @streamkap.bsky.social for an evening of demos from teams building at the intersection of AI, analytics, and data infrastructure. 🧵
110
Adrian Brudaru @datateam.bsky.social · 01/06/2026
Sameer started Metabase as a side project, waited years before charging, ignored conventional SaaS advice, and built a product now used by 90k+ companies with 8-figure ARR and a global OSS community behind it.
luma.com
Building in the open: Sameer Al-Sakran x Metabase · Luma
Berlin, we're back! This time, we're bringing our Founder & CEO, Sameer Al-Sakran, for a live conversation on building one of the most successful open source…
020
Adrian Brudaru @datateam.bsky.social · 01/06/2026
What does it take to build a product people actually keep using? On Wednesday, we're hosting Sameer Alsakran, Founder & CEO of @metabase.com, at our office in Berlin for a live conversation with Francesco Mucio from Data Berlin. 🧵
130
Adrian Brudaru @datateam.bsky.social · 26/05/2026
We’ll cover: → turning agent traces into dashboards for AI adoption, cost, and real business impact → how teams use Modal Sandboxes to run agents safely, plus a live background agent demo 📅 May 27 · 6–8 PM GMT+2 📍 Berlin
luma.com
Build with Agents — Berlin night w/ Modal, dltHub · Luma
Modal and dltHub are co-hosting an evening in Berlin for founders, engineers, and builders shipping with agents in production. There's no playbook (yet!) for…
010
Adrian Brudaru @datateam.bsky.social · 26/05/2026
Berlin’s Applied AI Week is looking pretty special. Tomorrow, dltHub and @modal-labs.bsky.social are hosting an evening in Berlin for founders, engineers, and builders shipping agents in production. Featuring Kenny Ning, Jefferson Girao & Alena Astrakhantseva.
110
Adrian Brudaru @datateam.bsky.social · 18/05/2026
The AI stack is evolving fast, but reliable data movement is still the foundation. Join dltHub, LanceDB and DataHub on May 21 in Menlo Park for talks on multimodal AI storage, AI data pipelines, and trusted lineage systems. 🔗
luma.com
The missing data layer for ML: dltHub x LanceDB x DataHub @ SVAI · Luma
Modern problems = Modern solutions! Join dltHub, LanceDB, and DataHub for a night of technical talks and demos. Hear from the engineers building the ingestion,…
010
Adrian Brudaru @datateam.bsky.social · 15/05/2026
In under 10 minutes, we’ll cover: - AI-assisted pipeline setup - GitHub → dlt workflows - Where AI helps vs where engineers still need to steer - How to inspect & validate pipelines 📅 May 20 · 10 AM PST 📍 Live on YouTube Watch live:
luma.com
Vibe Check: Building w/ AI & dlt in 10 Minutes · Luma
How fast can you go from zero to a working data pipeline when AI is your copilot? In this Vibe Check, Melanie and Cecil are joined by Elvis Kahoro from dltHub…
010
Adrian Brudaru @datateam.bsky.social · 15/05/2026
How fast can you go from zero to a production-ready data pipeline when AI is your copilot? Next Wednesday, Elvis Kahoro joins @temporal.io alongside Melanie Warrick and @cecilphillip.bsky.social for a live Vibe Check building a GitHub-powered pipeline with AI + dlt.
111
Adrian Brudaru @datateam.bsky.social · 14/05/2026
Building pipelines with AI usually means losing context between tools. dltHub AI Workbench runs the full 12-step workflow as a continuous session, schemas, incrementals, traces, transformations, and notebooks share context across the stack. dlthub.com/blog/agentic-data-engine…
020
Adrian Brudaru @datateam.bsky.social · 12/05/2026
Explainer on ontology engineering and what we're building around it: why just clean schemas and prompts aren’t enough, and how adding a canonical model + taxonomy + ontology changes what agents can correctly compute (ARPU being the clearest example). 
dlthub.com
Ontology engineering: what it is, why it's back, and why agents need it
Agents don't hallucinate. They navigate without a map. Ontology engineering is how you build one, and why every team pulling humans out of the loop needs it now.
010
Adrian Brudaru @datateam.bsky.social · 30/04/2026
We’ll show how agents can power data pipelines as code, turning traces into fresh, reliable datasets. Stack: dlt, LanceDB, Pydantic, Ibis, HuggingFace + more → behind our eval platform. 📅 May 5 | 🕒 12:20–12:40 📍 Station F, Paris
paris2026.gosim.org
From Agent Traces to Analytics – GOSIM Paris 2026
Agents continuously produce large volumes of artifacts: code, text, telemetry, etc. Most teams are stuck with static and stale datasets. In 2026, with agents ab
010
Adrian Brudaru @datateam.bsky.social · 30/04/2026
We’re excited to share that Violetta Mishechkina will be speaking at GOSIM Paris 🇫🇷 Invited by probabl.ai to join the “Own Your Data Science and AI” workshop. 🎤 From Agent Traces to Analytics Agents generate code, text, telemetry, yet most teams still rely on stale datasets.
110
Adrian Brudaru @datateam.bsky.social · 25/04/2026
We built the dlt AI Workbench so you don't have to maintain the "how to use dlt" layer yourself. Full argument:
dlthub.com
Who maintains the skill layer?
We're in an LLM-coding junior bubble. "It runs" isn't the senior bar. Lifecycle rigor and dependency management are.
010
Adrian Brudaru @datateam.bsky.social · 25/04/2026
Skills that wrap a library are software. They have dependencies, need maintenance, and degrade when the product changes and no one updates them. The vendor owns the product surface. You own your integration. Same rule, new layer.
120
Adrian Brudaru @datateam.bsky.social · 24/04/2026
3/ Fast, too. With the PyArrow backend, dlt loads a 6.4 GB Postgres database (~24.6M rows across 7 tables) in ~8 minutes. Try it free for 30 days, 60 compute hours included.  Docs:
dlthub.com
dlt Connector App | dlt Docs
How to use the dlt Connector App
010
Adrian Brudaru @datateam.bsky.social · 24/04/2026
2/ Because it's a Native App, your data stays governed: 🔒 Credentials stored as Snowflake secrets 🔒 Outbound access gated by External Access Integrations 🔒 Role-based access control 🔒 Passed Snowflake's security review Nothing leaves your account.
110
Adrian Brudaru @datateam.bsky.social · 24/04/2026
1/ dlt is now available as a Snowflake Native App. Replicate MSSQL, MySQL & PostgreSQL → Snowflake, without leaving Snowflake. Create, schedule, and monitor pipelines from the Snowflake UI. No external orchestrator. app.snowflake.com/marketplace/listi…
110
Adrian Brudaru @datateam.bsky.social · 15/04/2026
Eval: base Claude vs our Workbench. Base: leaks creds 100%, skips docs, no sampling, 1-shot code. Workbench: 0 leaks, always docs/samples/iterates. 58% higher cost ($2.21 vs $1.40) isn’t overhead, it’s the gap between AI slop and production. full read
dlthub.com
Why AI Agents Need a Guardrail Layer and What It Looks Like
How to stop AI agents from leaking credentials, skipping tests, and using outdated docs in data pipelines.
120
Adrian Brudaru @datateam.bsky.social · 15/04/2026
Every layer of software automation was called overkill before it became the baseline. Fortran → Make → CI/CD → Docker → now agents. Code that runs is only the 10%. The other 90% is engineering judgment, boundaries, and iteration.
110
Adrian Brudaru @datateam.bsky.social · 14/04/2026
If you don’t define the path, the agent will improvise. With our Agentic REST toolkit, we make the right way the easiest way: - structured access - limited operations - no hidden side effects Full breakdown:
dlthub.com
Agentic toolkit eval: dltHub REST API toolkit
Transforms AI-generated "vibe coding" from an unmanaged process full of hidden risks into a mature engineering workflow that prioritizes security.
020
Adrian Brudaru @datateam.bsky.social · 14/04/2026
AI agents don’t just use your APIs, they optimize around them. Ask an agent to “build a pipeline” and it will find credentials, escalate privileges, and take the shortest path to completion. Not against your interests, just goal-driven.
210
Adrian Brudaru @datateam.bsky.social · 09/04/2026
Not everything that can be modeled should be. With LLMs, more context doesn’t mean a better prompt. The key is Minimum Viable Context for high-precision data models. Here’s what we learned building ontology-driven modeling. Blog by Hiba Jamal ↓
dlthub.com
Minimum Viable Context for Building a Canonical Data Model
Call it the MVC problem: minimum viable context. Too little and it hallucinates your domain. Too much and it drifts from your actual goal. The process has to be controlled.
010
Adrian Brudaru @datateam.bsky.social · 03/04/2026
Outcome → fixed-price, high-margin projects are now viable. This isn’t a productivity hack. It’s a business model shift. Case study 👇
dlthub.com
Tasman Analytics prototypes client pipelines with dltHub Pro
Tasman Analytics cut scoping from 2 weeks to 20 min with dltHub Pro. See how they prototype any client pipeline in a single meeting and deliver faster than ever
010
Adrian Brudaru @datateam.bsky.social · 03/04/2026
For agencies and teams of 5+, standardization is everything.  Tasman encodes their standards, naming, rate limits, and workflows so mid-level engineers ship production-quality work. Knowledge scales across the team instead of being locked in a few individuals.
110
Adrian Brudaru @datateam.bsky.social · 03/04/2026
Generation works. What doesn’t come for free: → transparency → validation → correctness That’s where dlt comes in: → Dataset Browser → native logging → inspect–fix–rerun loop Catching schema drift, nested data, column mismatches, before production.
110
Adrian Brudaru @datateam.bsky.social · 03/04/2026
Tasman Analytics (20-person consultancy, enterprise clients) went from 2 weeks to scope an API connector → 20 minutes with dltHub Pro. But speed isn’t the story. The real question is: what happens after the code runs?
110
Adrian Brudaru @datateam.bsky.social · 03/04/2026
AI can generate a data pipeline in 10 minutes. But can you trust what it produces? That’s the real problem, and it’s not technical. It’s business. If you can’t trust the output, it never reaches production. 🧵
110
Adrian Brudaru @datateam.bsky.social · 02/04/2026
Writing pipelines isn’t the hardest part. Trusting them enough to deploy is. The deployment toolkit is available via design partnership as part of dltHub Pro. Interested → join the design partnership 👇
dlthub.com
Ship Data Pipelines 50x Faster | dltHub Pro
Join the dltHub Design Partner Program. Prototype any API connector in 20 minutes, ship production pipelines in an afternoon.
010
Adrian Brudaru @datateam.bsky.social · 02/04/2026
After deployment: You can ask the agent for pipeline health it: • inspects logs • pulls observability data • checks incremental loading (incl. duplicates)
110