Sign in

Alejandro Wainzinger

@xevix.bsky.social
85 followers 33 following 150 posts

Software Developer interested in data, web, languages. Silicon Valley/Tokyo. medium.com/@xevix github.com/xevix

PostsRepliesMedia
Alejandro Wainzinger @xevix.bsky.social · 26/09/2026
DuckParq 1.6.0 now supports Hive-partitioned Vortex as upstream plugin was updated. Also thousands separator in column view for nicer display. github.com/xevix/duckpa...
github.com
Release v1.6.0 · xevix/duckparq
What's new Hive-partitioned Vortex. A folder of .vortex files laid out as key=value directories now reads the way the same parquet layout does: the keys are columns of the table, typed from the fo...
000
Alejandro Wainzinger @xevix.bsky.social · 17/09/2026
To macOS friends, yes you can get rid of the invasive Ask Siri everywhere, quick script: www.reddit.com/r/MacOS/comm...
reddit.com
From the MacOS community on Reddit: Tired of seeing “Ask Siri” on every context menu? Here's how to disable it.
Explore this post and more from the MacOS community
000
Alejandro Wainzinger @xevix.bsky.social · 16/09/2026
New version of DuckParq now with Vortex support. 1.5B rows / 22.3GB Vortex file opens instantly. Yes, try not to make files that gigantic, but in case you do 😶‍🌫️. github.com/xevix/duckpa...
010
Alejandro Wainzinger @xevix.bsky.social · 13/09/2026
Quick hack to pipe results from nushell to a duckdb table for further querying. gist.github.com/xevix/466e5e...
020
Alejandro Wainzinger @xevix.bsky.social · 13/09/2026
Slowly catching up on the last few years of CLI tools: fzf, zoxide, bat, btop, zellij, nushell, ghostty &c. Wish I'd had these years ago, so fast and useful. LLMs do a lot these days, but simple stuff is faster done alone with powerful and simple tools.
021
Alejandro Wainzinger @xevix.bsky.social · 02/09/2026
The Pangeo channel is a great collection of talks around recent GIS development and usage. Been peeking at it for a while and watching more these days. www.youtube.com/@pangeo2060
youtube.com
Pangeo
Pangeo is a community and platform for big data analysis in the geoscience and beyond. You can learn more about Pangeo at https://pangeo.io. This channel hosts content generated by the Pangeo Commun...
020
Alejandro Wainzinger @xevix.bsky.social · 02/09/2026
Zarr dataset (~1B numbers, (24, 37, 721, 1440) float 32) mean() CPU vs GPU. Manual threadpool gets CPU close to GPU but numpy (1 thread) and dask are quite slow. M4 Max 128GB. MLX code requires so few lines, definitely worth it. gist.github.com/xevix/c320ed...
000
Alejandro Wainzinger @xevix.bsky.social · 01/09/2026
Visualization of zarr in common GUIs seems in transition. Panoply supports v2 w/o compression but not v3. QGIS supports v3 but squashes dimensions, seems not to allow slicing dimensions; GeoZarr plugin breaks on sharding_indexed codec loading via STAC (gdal update may fix). Almost there 🤞
201
Alejandro Wainzinger @xevix.bsky.social · 30/08/2026
Quick test of Vortex vs ParquetV2 w/ zstd on taxi data in DuckDB. 48M rows 1 file. As advertised Vortex has faster selective querying and aggs due in part to avoiding decompression (~30% faster here). File also ~7% smaller. Will test Hive once partitioned writes fix released (fixed upstream).
110
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 28/08/2026
What’s Next for DuckDB? 🦆 🤔 @hannes.muehleisen.org, co-creator of #DuckDB, was recently interviewed by @krisajenkins.bsky.social of the Developer Voices podcast on every challenge, reward, twist and turn of building a database as popular as DuckDB -- including a big acquisition. Watch or read:
duckdb.org
What Next for DuckDB
DuckDB is an in-process SQL database management system focused on analytical query processing. It is designed to be easy to install and easy to use. DuckDB has no external dependencies. DuckDB has bin...
0123
Alejandro Wainzinger @xevix.bsky.social · 28/08/2026
Neat talks hosted by Motherduck yesterday. I’ve been interested in zarr and using it in SQL w/o xarray for messing with geolibre things, cool to see. Also good to hear Motherduck taking the AWS news well, hopefully bodes well for the future of DuckDB.
131
Alejandro Wainzinger @xevix.bsky.social · 26/08/2026
Cautiously optimistic. Oracle got Sun and did good with Java, so who knows. Then again, AWS and open source. I guess time will tell.
110
Alejandro Wainzinger @xevix.bsky.social · 26/08/2026
DuckDB extension to add Hive data summary clause using the PEG tokenizer. Can't quite use the PEG grammar yet I think, but here's what it could look like when released. Going to be cool to see what people come up with when it's this easy to modify the grammar. github.com/xevix/duckdb...
000
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 20/08/2026
Your database deserves a better parser (Don’t worry, #DuckDB v2.0 is giving you one) 🦆 🦆 Read more from @dtenwolde.bsky.social on how DuckDB v2.0 replaces its Postgres-derived SQL parser with a PEG-based parser that is easier to evolve and can be extended at runtime:
duckdb.org
DuckDB v2.0: Your Database Deserves a Better Parser
DuckDB v2.0 replaces its PostgreSQL-derived SQL parser with a PEG-based parser that is easier to evolve and can be extended at runtime.
0183
Reposted by Alejandro Wainzinger
Andy Pavlo @andypavlo.bsky.social · 13/08/2026
Our new version of dbdb.io tracks the commit history of open-source databases. We also track which coding agents are tagged in commits. Copilot has the most overall usage but Claude occurs in the most recent commits.
Histogram of the number of database systems that use an agent in at least one commit since 2022.
4363
Reposted by Alejandro Wainzinger
Simon Willison @simonwillison.net · 16/08/2026
Heees my review of Qwen 3.8 27B - I can't remember the last time I've had this much fun playing with a local model that runs on my own computers simonwillison.net/2026/Aug/16/...
simonwillison.net
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Friday’s big release was Qwen 3.8 27B, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba’s Qwen research lab. I’ve been looking forward to this one: 27B is an …
2322630
Alejandro Wainzinger @xevix.bsky.social · 17/08/2026
Massive wins in this upcoming release. PEG parser is going to enable all kinds of syntax extensions, async I/O is a huge performance win.
171
Reposted by Alejandro Wainzinger
Alejandro Wainzinger @xevix.bsky.social · 06/08/2026
Open sourced in case anyone finds it useful. Binary releases available. It's pretty barebones but should get the job done. github.com/xevix/duckparq
github.com
GitHub - xevix/duckparq: Parquet Viewer powered by DuckDB
Parquet Viewer powered by DuckDB. Contribute to xevix/duckparq development by creating an account on GitHub.
001
Alejandro Wainzinger @xevix.bsky.social · 02/08/2026
I wanted a parquet viewer like Tad but it hasn't been updated in a while. Not bad for a few hours vibecode in Opus 5. Swift, powered by DuckDB via C API. Supports modifying underlying SQL, seeing schema.
100
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 31/07/2026
` Asynchronous I/O in #DuckDB: Work, Thread, Work ` It doesn't matter how fast query operators are in a database system if we can't pull in the data quickly. For most of DuckDB's history, this problem was largely avoided by pruning data early. As usual, things changed: duckdb.org/2026/07/31/a...
0224
Alejandro Wainzinger @xevix.bsky.social · 30/06/2026
Great to see temporal table support improved in Postgres. Still no built-in temporal querying e.g. AS OF, but at least no extensions needed now. SCD is hard, any tools that help are welcome. www.postgresql.org/docs/19/ddl-...
postgresql.org
5.7. Temporal Tables
5.7. Temporal Tables # 5.7.1. Application Time 5.7.2. System Time Temporal tables allow users to track different dimensions of history. Application …
000
Alejandro Wainzinger @xevix.bsky.social · 25/06/2026
Nice GUI to replace/augment CXPatcher for Crossover for games. Prerelease still, but promising. github.com/italomandara...
github.com
GitHub - italomandara/Procyon: A Steam game launcher for macOS that can run both Windows and MacOS Games
A Steam game launcher for macOS that can run both Windows and MacOS Games - italomandara/Procyon
100
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 12/05/2026
Today we are revealing the Next Big Thing for DuckDB: Quack, a protocol that turns DuckDB into a client-server database. True to DuckDB's philosophy, Quack is simple and fast. Setting it up takes seconds, and it works for both bulk operations and small write transactions. 🔗 duckdb.org/quack
710012
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 14/04/2026
Yesterday, we released DuckLake v1.0 and also published a podcast where Pedro Holanda and Mark Raasveldt talk about the DuckLake, going into quite some technical details. Listen to the episode at ducklake.select/media/duckla...
ducklake.select
DuckLake v1.0 – Developer Discussion
DuckLake is a SQL-as-a-lakehouse format that leverages your favorite SQL database (SQLite, PostgreSQL or DuckDB) to provide ACID transactions, schema evolution, and time travel for your data lake, wit...
0152
Alejandro Wainzinger @xevix.bsky.social · 21/03/2026
Materialized CTE support finally coming to Clickhouse. Goodbye unnecessary DB round trips and temp tables with consistency issues. github.com/ClickHouse/C...
github.com
Support Materialized CTE by novikd · Pull Request #94849 · ClickHouse/ClickHouse
Changelog category (leave one): New Feature Changelog entry (a user-readable short description of the changes that goes into CHANGELOG.md): Support Materialized CTE. Allow evaluating CTEs only on...
000
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 20/03/2026
We're excited to announce duckdb-skills, a DuckDB plugin for Claude Code! We think the embedded nature of DuckDB makes it a perfect companion for Claude in your local workflows. Check out the repository at github.com/duckdb/duckd...
github.com
GitHub - duckdb/duckdb-skills
Contribute to duckdb/duckdb-skills development by creating an account on GitHub.
15414
Alejandro Wainzinger @xevix.bsky.social · 16/03/2026
DuckDB just casually in the nVidia keynote. AI doesn't make databases irrelevant, it makes them more relevant than ever.
072
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 09/03/2026
We released DuckDB v1.5! This release comes with a “friendly CLI” client, a new (opt-in) PEG parser, support for VARIANT types and many lakehouse features. It also ships a new network stack, a reworked geospatial extension, Azure writes and an ODBC scanner. Read more at duckdb.org/2026/03/09/a...
06116
Alejandro Wainzinger @xevix.bsky.social · 08/02/2026
Good non-technical summary of some of the various chips out there. www.youtube.com/watch?v=RBmO...
youtube.com
How Nvidia GPUs Compare To Google’s And Amazon’s AI Chips
YouTube video by CNBC
000
Alejandro Wainzinger @xevix.bsky.social · 07/02/2026
Claude cowork w/ Opus 4.6 is definitely smart, but got stuck on a data task, I stopped it, pointed it to DuckDB, done instantly. LLMs still have much to learn 🤔
100
Alejandro Wainzinger @xevix.bsky.social · 07/02/2026
Great overview of chip manufacturing with EUV. We really do have supercomputers in our pockets, on our desk and on our wrists. www.youtube.com/watch?v=B248...
youtube.com
The $200M Machine that Prints Microchips: The EUV Photolithography System
YouTube video by Branch Education
000
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 02/02/2026
🎞️ The slide decks and talk recordings of last Friday's developer meeting are out! duckdb.org/events/2026/...
0226
Alejandro Wainzinger @xevix.bsky.social · 23/01/2026
Great talks at South Bay Systems hosted at databricks on xNVMe, fast SSD query processing, and using NPUs for DB work. Much work using DuckDB extensions. Need for async I/O as bottleneck a common topic, mainly at larger scale. luma.com/8a54z94d?tk=...
010
Alejandro Wainzinger @xevix.bsky.social · 31/12/2025
GPU-powered analytical query engines going mainstream? Needs nVidia GPU and limited to memory for now, but neat use of Substrait+Arrow for interop. DuckDB still easier to run anywhere, but this is useful for acceleration if needed. developer.nvidia.com/blog/nvidia-...
developer.nvidia.com
NVIDIA CUDA-X Powers the New Sirius GPU Engine for DuckDB, Setting ClickBench Records | NVIDIA Technical Blog
Sirius, an open-source GPU native SQL engine, achieved a new performance record on Clickbench—a widely used analytics benchmark. Developed by University of Wisconsin-Madison with support from NVIDIA…
020
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 31/10/2025
The PyData Amsterdam 2025 keynote “Minus Three Tier: Data Architecture Turned Upside Down” by @hannes.muehleisen.org is out now. www.youtube.com/watch?v=DxwD...
youtube.com
KEYNOTE: Hannes Mühleisen - Data Architecture Turned Upside Down | PyData Amsterdam 2025
YouTube video by PyData
1254
Reposted by Alejandro Wainzinger
Andy Pavlo @andypavlo.bsky.social · 03/11/2025
New database leaderboard from Yellowbrick ranks the quality of DBMS optimizer estimates and plans. They only evaluate TPC-H for now and report results for Postgres + DuckDB + MSSQL: sql-arena.com/components/p... Repo: github.com/sql-arena/db... LinkedIn Group: www.linkedin.com/groups/15775...
SQL Arena Planner Ranking (November 2025)
1143
Reposted by Alejandro Wainzinger
CMU Database Group @db.cs.cmu.edu · 20/10/2025
Today's Future Data Systems Seminar Speaker: Ian Cook (@ian.columnar.tech) will present @columnar.tech's work on Apache Arrow's database connectivity API (ADBC). ADBC is available in modern DBMSs. Zoom talk open to public at 4:30pm ET. YouTube video available after: db.cs.cmu.edu/events/futur...
db.cs.cmu.edu
[Future Data] Where We're Going, We Don't Need Rows: Columnar Data Connectivity with ADBC - Carnegie Mellon Database Group
ADBC (Arrow Database Connectivity) is Apache Arrow’s answer to ODBC and JDBC:... Read More +
0158
Reposted by Alejandro Wainzinger
CMU Database Group @db.cs.cmu.edu · 13/10/2025
Today's Future Data Systems Seminar Speaker: Will Manning (@willmanning.com) will present @spiraldb.com's Vortex file format. Vortex is now a @linuxfoundation.org project. Zoom talk open to public at 4:30pm ET. YouTube video available after: db.cs.cmu.edu/events/futur...
db.cs.cmu.edu
[Future Data] Vortex: LLVM for File Formats - Carnegie Mellon Database Group
Apache Parquet revolutionized columnar storage after its initial release in 2013, but... Read More +
044
Alejandro Wainzinger @xevix.bsky.social · 10/10/2025
Processing 100Tb of CSV files on a single machine is insane, little over 1hr per query, even if on a powerful AWS instance. Question heavily the need for complex systems when this is what’s possible now. Can’t wait for full write-up. Incredible work. duckdb.org/2025/10/09/b...
duckdb.org
Benchmark Results for DuckDB v1.4 LTS
DuckDB v1.4 LTS is both fast and scalable. In in-memory mode, it is the fastest system on ClickBench. In disk-based mode, it can run complex analytical queries on a dataset equivalent to 100 TB CSV fi...
081
Alejandro Wainzinger @xevix.bsky.social · 04/10/2025
Taking the DuckDb hoodie on a trip. Not exactly Amsterdam but I’ve heard they like columnar databases here too.
030
Alejandro Wainzinger @xevix.bsky.social · 16/09/2025
Congrats to DuckDB team on LTS release w/ many great improvements! Hidden among them you can now use Hive filtering with read_blob, and SHOW TABLES FROM specific db w/o USE.
120
Reposted by Alejandro Wainzinger
DuckDB @duckdb.org · 16/09/2025
📈 DuckDB 1.4.0 is out! This is our first LTS release which comes with *one year of community support*. It also supports database encryption, the MERGE SQL statement and Iceberg writes. For more details, read the announcement blog post at duckdb.org/2025/09/16/a...
05221
Alejandro Wainzinger @xevix.bsky.social · 02/09/2025
I tried loading eBird data (1.5B rows CSV ZIP) using DuckDB for fun, inspired by a Clickhouse blog post and a bit of curiosity. Both did well, DuckDB slightly faster querying and Parquet ingest, Clickhouse w/ native zip support, optimized for ingest and multitenancy. xevix.medium.com/ebird-in-duc...
xevix.medium.com
eBird in DuckDB
I saw this post by the Clickhouse team which was doing a cool test of the eBird dataset from Cornell University, and wondered how DuckDB…
030
Reposted by Alejandro Wainzinger
PVLDB @pvldb.bsky.social · 03/08/2025
Vol:18 No:8 → Saving Private Hash Join 👥 Authors: Laurens Kuiper, Paul Gross, Peter Boncz, Hannes Mühleisen 📄 PDF: www.vldb.org/pvldb/vol18/p2748-kuip…
Thumbnail: Saving Private Hash Join
0144
Alejandro Wainzinger @xevix.bsky.social · 29/08/2025
Is there too much duplicated effort in data tools? I sometimes wonder about this. xevix.medium.com/data-tool-co...
xevix.medium.com
Data Tool Component Sharing
There are many partly overlapping tools in the data world, which is what inspired things like Calcite to have modular components for…
000
Alejandro Wainzinger @xevix.bsky.social · 25/08/2025
Compiling DuckDB on Windows 11 (ARM) using UTM VM on macOS to debug Windows compile issues. It's a shame msvc doesn't exist outside of Windows, mingw/clang don't work the same and cross-compiling is tricky. Compiling takes 5-10 mins (instead of 1-2 mins native), but it works 🎉!
130
Alejandro Wainzinger @xevix.bsky.social · 15/08/2025
Stretching DuckDB w/ Common Crawl, ~1.7B rows, ~300 parquet files. ~2-3s for single-column aggregations, ~2-3 mins to SUMMARIZE the data, peaking at ~12-14GB memory usage. Not exactly real-time, but the fact you can do this on a laptop with no server setups or Spark pipelines is still amazing.
1439
Alejandro Wainzinger @xevix.bsky.social · 12/08/2025
Neat little hack to get Hive partition list in DuckDB, useful for an overview. Might be neat to have built-in. gist.github.com/xevix/04f33d...
140
Alejandro Wainzinger @xevix.bsky.social · 29/06/2025
Added an Automator quick action to run sqlfluff for formatting SQL in browser fields, used here in the DuckDB UI. Only needs sqlfluff, optionally configure rules. Would be cool to get built-in one day, but works for now.
100
Alejandro Wainzinger @xevix.bsky.social · 27/06/2025
Apache Drill allowed storing metadata in an RDBMS, Iceberg scaling data, Arrow scaling columnar memory, Parquet columnar storage, Spark distributed compute, DuckDB single-node compute. DuckLake scales metadata and storage w/ compute on single node. Motherduck distributes compute.
110