Sign in

Anders

@dataders.bsky.social
1.6K followers 812 following 77 posts

DX @ dbt

PostsRepliesMedia
Anders @dataders.bsky.social · 03/04/2025
👋
030
Anders @dataders.bsky.social · 01/04/2025
yeah the multi-cloud story is far from over. you can imagine that without egress costs, it might actually be performant to move data b/w AWS and Azure it the data centers are close enough. Iceberg kinda makes DWH on Cloudflare R2 feasible given it has S3-compatible API. Sippy seems cool
developers.cloudflare.com
Sippy · Cloudflare R2 docs
Sippy is a data migration service that allows you to copy data from other cloud providers to R2 as the data is requested, without paying unnecessary cloud egress fees typically associated with moving ...
030
Anders @dataders.bsky.social · 01/04/2025
new rule: talk about Iceberg without mentioning Hive or ACID properties. the vast majority of SQL users don't care (and shouldn't)! it's like explaining Dropbox starting w/ bringing up libfuse If you're new to iceberg lmk if you get something from this! roundup.getdbt.com/p/iceberg-gi... #databs
roundup.getdbt.com
Iceberg?? Give it a REST!
The new abstraction that changes nothing... and everything
092
Anders @dataders.bsky.social · 24/03/2025
lots of juicy looking papers in the agenda for EDBT/ICDT 2025 happening this year in Barcelona! definitely going to be diving in more later today edbticdt2025.upc.edu?contents=det... #databs #edbc2025 #icdt2025
edbticdt2025.upc.edu
EDBT/ICDT 2025 Joint Conference - 25th March - 28th March, 2025 - Barcelona, Spain
020
Reposted by Anders
Brett Cannon @snarky.ca · 19/03/2025
I FINALLY asked for pronouncement on PEP 751 -- lock files for #Python : peps.python.org/pep-0751/ .
peps.python.org
PEP 751 – A file format to record Python dependencies for installation reproducibility | peps.python.org
This PEP proposes a new file format for specifying dependencies to enable reproducible installation in a Python environment. The format is designed to be human-readable and machine-generated. Installe...
05714
Anders @dataders.bsky.social · 17/03/2025
p.s. also TI[F]L what "serde" stands for after seeing the word for years 🤦
020
Anders @dataders.bsky.social · 17/03/2025
this is the clearest case for Arrow that I've ever seen. I love that it's high-level but also doesn't shy away from details when it's important. this should be a #databs canon text imho. thanks @ianmcook.bsky.social !
181
Anders @dataders.bsky.social · 13/03/2025
TIL about relationalplayground.com makes relational algebra accessible to a SQL monkey like me. much better than in a dry textbook where it's normally found. #databs
relationalplayground.com
Relational Playground
An exploration of relational algebra. Compare SQL queries with relational algebra expressions along with intermediate results.
040
Anders @dataders.bsky.social · 03/03/2025
imho, the most clear case Databricks has made in public on their Iceberg future post-Tabular acquisition. I agree with this vision of the future and it's nice to see DBRX sharing how they see themselves participating in it. worth clicking through the deck! #databs speakerdeck.com/databricksja...
speakerdeck.com
Iceberg Meetup Japan #1 : Iceberg and Databricks
2月21日に開催されたIceberg Meetup #1で使用した資料になります。 DatabricksとIcebergを使用する際のカタログについてご紹介しています。
030
Anders @dataders.bsky.social · 03/03/2025
11/11) p.s. forgot to link to the repo! github.com/deepseek-ai/...
github.com
GitHub - deepseek-ai/smallpond: A lightweight data processing framework built on DuckDB and 3FS.
A lightweight data processing framework built on DuckDB and 3FS. - deepseek-ai/smallpond
010
Anders @dataders.bsky.social · 03/03/2025
10) So sick that a smallpond pipeline returns a LogicalPlan representing a DAG where each node is a distinct data processing task. Imagine if a dbt DAG resulted in a single logical plan that operates across multiple engines. and then they can optimize the plan before execution as well! 🤯🤯🤯
110
Anders @dataders.bsky.social · 03/03/2025
9) Arrow is the unsung hero of this project (and arguably all innovation in data ecosystem). it's what enables: 1. all this interchangeability of query engines 2. (likely) using duckdb in a distributed environment in the first place
120
Anders @dataders.bsky.social · 03/03/2025
8) this HN called out that smallpond abstracts supports using different query engines for different jobs (shuffling vs sorting). Very bullish on this future of right tool for right job and making it as simple as a config news.ycombinator.com/item?id=4323...
news.ycombinator.com
One thing I found peculiar is that for the GraySort benchmark it dispatches to P... | Hacker News
120
Anders @dataders.bsky.social · 03/03/2025
7) TIRED: "big vs. small" & "distributed vs single-node" WIRED: tactical deployment of single-node query engines within distributed frameworks. another great example is Apache Comet which plugs DataFusion into Spark to accelerate single-node operations resulting in overal Spark performance speedups
120
Anders @dataders.bsky.social · 03/03/2025
6) Making smallpond 5 years ago would have been very difficult! but the emergence of lower-level, off-the-shelf components greatly accelerate the development time. Within the year, we'll to see this new paradigm catch on. Future examples will probably be using DataFusion not DuckDB.
230
Anders @dataders.bsky.social · 03/03/2025
5) there's been previous discussion on DeepSeek's scrappiness and I think it shows here. They had a vision of what they wanted and rather than paying for software or forcing their vision into an existing tool, were able to ship exactly what they wanted
110
Anders @dataders.bsky.social · 03/03/2025
4) smallpond is a bespoke data processing framework using off-the-shelf, OSS components (ray, arrow, duckdb, polars). Why didn't they use {TOOL}? My guesses ❌ dbt: they wanted Python Dataframe API ❌ Airflow: not as close to metal as Ray ❌ pytorch or ray[data]: idk tbh
120
Anders @dataders.bsky.social · 03/03/2025
3) ray.io is the foundation of any training and inference infrastructure. I haven't had much exposure to Ray as a SQL monkey using DWHs, but it's just recently clicked for me how big of a deal it is
ray.io
Scale Machine Learning & AI Computing | Ray by Anyscale
Ray is an open source framework for managing, executing, and optimizing compute needs. Unify AI workloads with Ray by Anyscale. Try it for free today.
120
Anders @dataders.bsky.social · 03/03/2025
2) so cool to begin to see the data infra that supports training LLMs. Open weights is cool, and RAG makes sense, but as a former XGBooster turned "data engineer", seeing the data cleaning pipelines is what I've most wanted to see.
140
Anders @dataders.bsky.social · 03/03/2025
1) my top-level takeaways on DeepSeek's smallpond: a distributed data processing framework used for training LLMs #databs
182
Anders @dataders.bsky.social · 19/02/2025
dude -- so cool! one of my self-described superpowers is being very "plugged in" but, this doesn't happen without significant time and attention costs. what you've made changes the game imho. now I need the same for all the Slacks & Discords I'm in.
110
Reposted by Anders
Internal Tech Emails @techemails.bsky.social · 08/02/2025
Mark Zuckerberg messages Facebook engineer April 5, 2012
Mark Zuckerberg
Around?

Facebook engineer
Yeah

Mark Zuckerberg
If you could buy one of either Instagram, Foursquare or Pinterest, which would you buy?
1776
Reposted by Anders
Bijil Subhash @bijilsubhash.bsky.social · 01/02/2025
Just finished watching the webinar on introducing SDF by dbt team. After seeing SDF in action, I have to admit that I am really looking forward to the future of dbt engine. I was wondering when dbt was going to bring in notable changes to the developer experience and this might be it. #databs
031
Reposted by Anders
Sung Kim @sungkim.bsky.social · 26/01/2025
Building Query Compilers by Guido Moerkotte (695 pages) Note: This is repost of @emresevinc.bsky.social's on X Link: pi3.informatik.uni-mannheim.de/~moer/queryc...
061
Anders @dataders.bsky.social · 26/01/2025
Yeah that ADP chart hurt my friend y-axis so bad I think axisslaughter should be a punishable crime
010
Anders @dataders.bsky.social · 24/01/2025
post two! The key technologies behind SQL Comprehension #databs, do not fear compiler concepts -- embrace them and the new world order they enable for us! read the great blog (& pretty diagrams!) @daveconnors3.bsky.social, truly a masterpiece docs.getdbt.com/blog/sql-com...
Bernie Sanders meme labelled with Dave Connors's name: I am once again asking you to move up the stack and learn how database compilers work
041
Anders @dataders.bsky.social · 23/01/2025
written by my esteemed collegue @joellab.es
bsky.app
020
Anders @dataders.bsky.social · 23/01/2025
post one of a new series kicking off today whose larger thrust is effectively: understanding SQL, not just a job for the database! Post 1 lays out what the levels of understanding docs.getdbt.com/blog/the-lev... #databs
141
Anders @dataders.bsky.social · 14/01/2025
my feeling is that I haven't heard back, but have heard from others that they've been accepted so I'm operating under assumption that my talk has not been accepted
200
Anders @dataders.bsky.social · 04/01/2025
Look forward to these every year. Always love Andy's candor and insight on our little corner of the world. #databs
060
Anders @dataders.bsky.social · 27/12/2024
#ghostty dropping today is the GitHub equivalent of a Beyoncé album. My whole feed is everyone following it. Now I gotta at least try it right? github.com/ghostty-org/...
040
Anders @dataders.bsky.social · 19/12/2024
my experience installing MSFT's ODBC driver on M-series macbook re-stokes my long-burning ire for Simba drivers and the company behind them. One might think MSFT is to blame for a less than medoicre database driver, but in a way they're victims here too, held captive by Simba. #databs
020
Anders @dataders.bsky.social · 18/12/2024
does anyone use the 1Password CLI? I feel it has potential to be very valuable for managing environment variables (esp for teams). but i'm hesitant to adopt bc it requires that every command you run consume the output of `op` developer.1password.com/docs/cli/get...
developer.1password.com
Get started with 1Password CLI | 1Password Developer
Learn how to install and sign in to 1Password CLI, then get started with commands and scripts to manage users, vaults, and items on the command line.
210
Anders @dataders.bsky.social · 17/12/2024
admittedly, I'm triggered by SFTP, but it's a good analogy! What if there was an SFTP where you didn't have to parse random-delimited text files, but could just query the tables inside?
010
Anders @dataders.bsky.social · 13/12/2024
is there a recording of this somewhere?
110
Anders @dataders.bsky.social · 12/12/2024
DuckDB’s already got Delta/Unity support? As in catalogs already a first class primitive within DuckDB, but the only supported type is Unity w Delta format. IIRC Databricks sponsored DuckDB Labs to do this work. ‘ATTACH 'unity' AS unity (TYPE UC_CATALOG);’ docs.unitycatalog.io/integrations...
docs.unitycatalog.io
DuckDB - Unity Catalog
Open, Multi-modal Catalog for Data & AI
000
Anders @dataders.bsky.social · 11/12/2024
you ever see this April Fool's from a few years ago? iconic dbt-excel.com
050
Anders @dataders.bsky.social · 11/12/2024
(3/3) the theoretical bar for tools & vendors to integrate is way lower if S3 tables are "real". every tool/engine/db worth it's salt already integrates with S3. much lower reach to support Iceberg by extending existing support than supporting each new catalog (even if they're all REST)
010
Anders @dataders.bsky.social · 11/12/2024
(2/3) not having to host a catalog > hosting (or paying for a catalog) bsky.app/profile/data...
100
Anders @dataders.bsky.social · 11/12/2024
(1/3) introduces a more fundamental, lower-level abstraction than Iceberg Catalogs or Hive. e.g. the new S3 commands like CreateNamespace. opens the data lake to non-data SWes. More can use AWS CLI / boto3 than can spin up spark session, mount Glue catalog and execute CREATE TABLE or INSERT INTO.
110
Anders @dataders.bsky.social · 11/12/2024
(0/3) for Athena/Spark users, nothing new. the "juice" imho is it:
110
Anders @dataders.bsky.social · 11/12/2024
sure you can use it in Databricks, but what if it were a first-class citizen like Spark Clusters and SQL Warehouses? You could use it in the SQL IDE, notebooks, as a target for dbt-databricks. Your Unity Catalog would be auto-mounted and any create tables would persist back to UC.
100
Anders @dataders.bsky.social · 09/12/2024
right?! also I wish Databricks would support DuckDB if only to get the folks off our collective back that are trying to be very cheap by using jobs clusters with dbt-databricks. DuckDB already fully supports Delta/Unity!
110
Anders @dataders.bsky.social · 09/12/2024
art history major = business stakeholder? 😜 #databs
I must study anaytics eng, so that our sons have liberty to study data analytics and reporting. Our sons ought to study data analytics, reporting and solving problems for the business in order to give their children a right to say: "hey -- quick question"
040
Reposted by Anders
Kit Menke @kitmenke.com · 06/12/2024
Great breakdown of the new S3 Tables feature that leverages Apache Iceberg. Including an explanation of the costs... which are complicated. #dataBS bigdata.2minutestreaming.com/p/meet-your-...
bigdata.2minutestreaming.com
meet your new data lakehouse: S3 Iceberg Tables
S3 Tables and S3 Metadata are two brutal new features that compete with common Apache Iceberg Lakehouse architectures
061
Anders @dataders.bsky.social · 05/12/2024
This video series of his is phenomenal imho. Perfect content density so I don’t get distracted. And big “mic drop” truths www.youtube.com/watch?v=Bzhb...
youtube.com
30 Essential Ideas you should know about ADHD, 1A Intro, Chronic Developmental Disability
YouTube video by Adhd Videos
000
Anders @dataders.bsky.social · 05/12/2024
5/ fwiw, the S3 API is a standard of sorts, esp given that Cloudflare's equivalent service, R2, supports it. like @conormccarter.bsky.social said, if R2 (and Azure Blog, and GCP storage) support a table bucket style, we're moving towards standard interface/primitive bsky.app/profile/cono...
120
Anders @dataders.bsky.social · 05/12/2024
4/ what about S3 Tables? aren't they then "bad" by not using the OSS API spec? 🤷 data teams are predominately on a single cloud, so what they need is a standard within a single provider. and, like @benesch.bsky.social said, always better to not have to think about catalogs and have it 'just work'
130
Anders @dataders.bsky.social · 05/12/2024
3/ half-baked comparison: maybe kinda like manufacturer-agnostic electric vehicle charging stations? bc: A. e-car buyer would want to know how prevalent is the network of compatible charging stations? B: e-car manufacturers prefer vehicles to be maximally compatible with various charging stations
100
Anders @dataders.bsky.social · 05/12/2024
2/ lately, I've been stoked more on the Iceberg REST catalog than the table format itself bc it radically simplifies things both for A. end users: in terms of what they need to know B: data platforms: in terms of the number standalone integrations needed
120