Sign in

Marco Slot

@marcoslot.com
1.1K followers 356 following 95 posts

Mostly posts about PostgreSQL, Snowflake Postgres, and PostgreSQL extensions. Formerly Crunchy Data, Microsoft, Citus Data, AWS, TCD, VU

PostsRepliesMedia
Reposted by Marco Slot
Elizabeth Garrett Christensen @sqlliz.bsky.social · 09/01/2026
Next Tuesday @marcoslot.com will be at Postgres Meetup for * to talk about pg_lake - #Postgres for #Iceberg with #DuckDB. Join us! www.meetup.com/postgres-mee...
meetup.com
Postgres for the lakehouse: pg_lake, Tue, Jan 13, 2026, 12:00 PM | Meetup
Marco Slot will be here this month to talk about the new extension pg_lake that connects Postgres to lakehouse and object storage - for Iceberg, Parquet, csv and more. As
041
Marco Slot @marcoslot.com · 04/11/2025
pg_lake just went open source! (Apache 2.0) pg_lake is a set of extensions (from Crunchy Data Warehouse) that add comprehensive Iceberg support and data lake access to Postgres, with @duckdb.org transparently integrated into the query engine. Announcement blog: www.snowflake.com/en/engineeri...
1273
Reposted by Marco Slot
Andy Pavlo @andypavlo.bsky.social · 03/07/2025
No system hits the sweet spot of allowing for extensibility while maintaining systems safety. It would be nice if there was a standard plugin API (think POSIX) that allows compatibility across systems. Thanks to @marcoslot.com + @daveandersen.bsky.social for their collaboration on this project
Safety vs. Flexibility quadchart for different DBMSs. VLDB 2025
https://doi.org/10.14778/3725688.3725719
0101
Reposted by Marco Slot
Andy Pavlo @andypavlo.bsky.social · 03/07/2025
At last @abigalekim.bsky.social's paper is out! Its the most complete eval of DB extensions/plugins ever. We analyze PostgreSQL, MySQL, MariaDB, SQLite, DuckDB, Redis. TLDR: Postgres extns ecosystem is fraught with footguns. Other DBMSs have fewer extns but less problems. DuckDB has cleanest API.
16612
Reposted by Marco Slot
Craig @craigkerstiens.com · 02/06/2025
Five years ago I joined @crunchydata.com, shortly after I wrote about having unfinished business with Postgres. Today as part of Snowflake that journey is continuing. We've built some amazing things, but are just getting started. www.crunchydata.com/blog/crunchy...
crunchydata.com
Crunchy Data Joins Snowflake | Crunchy Data Blog
We are excited to announce that Crunchy Data is joining Snowflake to bring Postgres to the AI Data Cloud.
6315
Marco Slot @marcoslot.com · 29/05/2025
Recording of my Data Council talk: www.youtube.com/watch?v=HZAr...
youtube.com
Converging Database Architectures DuckDB in PostgreSQL
YouTube video by Data Council
0152
Marco Slot @marcoslot.com · 04/05/2025
Generative AI comes up with details that would be hilarious, if it wasn't so mind boggling that it can come up with these details.
010
Marco Slot @marcoslot.com · 22/04/2025
And there it is: Native logical replication from any Postgres server to Iceberg managed by Crunchy Data Warehouse. Speed up Postgres analytical queries 100x with 2 commands.
2202
Marco Slot @marcoslot.com · 03/04/2025
I gave a talk at the inaugural (and awesome) European Iceberg meetup in Amsterdam last night. It's an introduction to how and why we used Iceberg and DuckDB to build a Postgres Data Warehouse: www.youtube.com/watch?v=cEnq...
youtube.com
Building a Postgres Data Warehouse with Iceberg
YouTube video by Apache Iceberg™ Meetup
070
Marco Slot @marcoslot.com · 01/04/2025
Move fast and build solid solutions that work across platforms. You can now use Postgres as a modern Data Warehouse anywhere, using any S3-compatible storage API. Query, import, or export files in your data lake or store data in Iceberg with automatic maintenance and very fast queries.
010
Reposted by Marco Slot
Crunchy Data @crunchydata.com · 01/04/2025
Excited to announce Crunchy Data Warehouse is now available for Kubernetes and On-premises. Need faster analytics from Postgres? Want a native Postgres data lake experience? Learn more about how it works: www.crunchydata.com/blog/crunchy...
crunchydata.com
Crunchy Data Warehouse: Postgres with Iceberg Available for Kubernetes and On-premises | Crunchy Data Blog
Crunchy Data brings Postgres-native Apache Iceberg to Kubernetes and on-prem workloads.
051
Reposted by Marco Slot
Hannes Mühleisen @hannes.muehleisen.org · 28/03/2025
Amazing result
3688
Marco Slot @marcoslot.com · 26/03/2025
We weren't really thinking of log management as a target use case, but Iceberg is ideal as the final destination for logs, and having transactions & built-in job scheduling & a fast query engine (& laser focus on developer experience) makes things really simple and cost-effective.
172
Reposted by Marco Slot
Craig @craigkerstiens.com · 26/03/2025
I got a number of questions on how we saved $30k a month on cloudwatch by moving logs directly to S3/Iceberg with Postgres so I wrote up how in a bit more detail - www.crunchydata.com/blog/reducin...
crunchydata.com
Reducing Cloud Spend: Migrating Logs from CloudWatch to Iceberg with Postgres | Crunchy Data Blog
How we migrated our internal logging for our database as a service, Crunchy Bridge, from CloudWatch to S3 with Iceberg and Postgres. The result was simplified logging management, better access with SQ...
072
Reposted by Marco Slot
Crunchy Data @crunchydata.com · 20/03/2025
Excited to announce built-in maintenance for Iceberg via Postgres. Now within Crunchy Data Warehouse we will automatically vacuum and continuously optimize your Iceberg data by compacting and cleaning up files. Dig into the details of how this works www.crunchydata.com/blog/automat...
crunchydata.com
Automatic Iceberg Maintenance Within Postgres | Crunchy Data Blog
Iceberg can create orphan files during snapshot changes or transaction rollbacks. Crunchy Data Warehouse automatically cleans up the orphan files using a new autovacuum feature.
0123
Marco Slot @marcoslot.com · 14/03/2025
Imagine your potential customer as a serious company doing serious things, and willing to pay serious money if you can genuinely help them run their business without causing lot of new problems. Then go build products for that customer. This works.
000
Marco Slot @marcoslot.com · 11/03/2025
Auto-vacuum for #Iceberg tables is now available in Crunchy Data Warehouse! We're always aiming for a 0-touch experience where possible, so we went out of our way to make Iceberg compaction & cleanup fully automatic without any configuration. Still pretty interesting to see a manual vacuum:
060
Reposted by Marco Slot
Craig @craigkerstiens.com · 27/02/2025
A big part of building Crunchy Data Warehouse was ease of use. How easy is it to load data from existing public datasets? Step 1: Point at your dataset and we'll load it for you Step 2: Query it Step 3: Profit
051
Marco Slot @marcoslot.com · 14/02/2025
ChatGPT Plus had a good run, but looks like Le Chat is going to be my main assistant now. I like that it's fast, to the point, and quite clever. I was impressed with a SQL query it came up with today for finding contiguous ranges of integers. ChatGPT's version was 3x slower.
030
Marco Slot @marcoslot.com · 07/02/2025
Postgres is increasingly becoming a versatile data platform, instead of just an operational database. Using pg_parquet you can trivially export data to S3, and using Crunchy Data Warehouse you can just as easily query or import Parquet files from PostgreSQL.
0102
Marco Slot @marcoslot.com · 28/01/2025
Deepseek R1 in an ollama "container app" on a managed Postgres server, because... why not?
050
Marco Slot @marcoslot.com · 27/01/2025
5 years from now, no one's going to want slower, less reliable, or harder to use databases.
110
Marco Slot @marcoslot.com · 23/01/2025
🎉 pg_documentdb is open source I created the initial version with Vinod Sridharan (an absolutely brilliant engineer) at Microsoft a few years ago and it's come a long way since. It reimplements Mongo API with exact semantics in PostgreSQL. Already used by FerretDB! github.com/microsoft/do...
github.com
GitHub - microsoft/documentdb: DocumentDB offers a native implementation of document-oriented NoSQL database, enabling seamless CRUD operations on BSON data types within a PostgreSQL framework.
DocumentDB offers a native implementation of document-oriented NoSQL database, enabling seamless CRUD operations on BSON data types within a PostgreSQL framework. - microsoft/documentdb
04616
Marco Slot @marcoslot.com · 17/01/2025
Impressed by the latest ParadeDB release. Solving the right problems in the right way is really hard.
071
Reposted by Marco Slot
Philippe Noël @philippemnoel.bsky.social · 17/01/2025
1/11. ParadeDB is now integrated with Postgres block storage. As far as we know, no one has integrated a search and analytics engine with Postgres storage before. This is a big deal. Here's why we did it, how we did it, and why you should care. 🧵
3297
Marco Slot @marcoslot.com · 09/01/2025
A lot of great recommendations on tuning PostgreSQL for analytical queries by @karenhjex.bsky.social www.crunchydata.com/blog/postgre...
crunchydata.com
Postgres Tuning & Performance for Analytics Data | Crunchy Data Blog
Karen digs into Postgres strategies for working with large analytical data sets. She reviews tuning, strategies for pre-compiling data, and other analytics systems.
0134
Marco Slot @marcoslot.com · 09/01/2025
If you've become your company's designated Postgres expert, learning to navigate Postgres source code can be a useful skill, even if you're not a C programmer. For instance, the docs do not specify how the optimizer works, but there is a detailed readme in the source: github.com/postgres/pos...
github.com
postgres/src/backend/optimizer at master · postgres/postgres
Mirror of the official PostgreSQL GIT repository. Note that this is just a *mirror* - we don't work with pull requests on github. To contribute, please see https://wiki.postgresql.org/wiki/Subm...
1323
Reposted by Marco Slot
Andy Pavlo @andypavlo.bsky.social · 01/01/2025
Buckle up because we're banging into the new year with my annual retrospective of the last year in databases! Highlights include license change blowback, Databricks vs. Snowflake gangwar, @duckdb.org's shotgun weddings, and buying a quarterback to impress your lover: www.cs.cmu.edu/~pavlo/blog/...
cs.cmu.edu
Databases in 2024: A Year in Review
Andy rises from the ashes of his dead startup and discusses what happened in 2024 in the database game.
1019963
Reposted by Marco Slot
Rachel Stephens @rstephens.me · 19/12/2024
In the piece I cover: - the partnership between @cncf.io and Andela - the release of Akamai App Platform cc/ @mikemaney.bsky.social - the rebranding of Akka - the partnership between @gitlab.com and AWS - the release of Crunchy Data Warehouse cc/ @craigkerstiens.com @crunchydata.com
152
Reposted by Marco Slot
Gunnar Morling @gunnarmorling.dev · 19/12/2024
Woot, woot, #Debezium 3.0.5 is out, which amongst other things also contains my patch for adding support for #Postgres 17 failover slots 🥳. You now can fail over your Debezium CDC connector to a replica turned primary, without missing any events! debezium.io/blog/2024/12...
debezium.io
Debezium 3.0.5.Final Released
Debezium is an open source distributed platform for change data capture. Start it up, point it at your databases, and your apps can start responding to all of the inserts, updates, and deletes that ot...
0143
Reposted by Marco Slot
Nikhil Benesch @benesch.bsky.social · 19/12/2024
Well well well: www.crunchydata.com/blog/pg_incr... Incremental pipelines come to Postgres via Crunchy Data! This is like "dbt incremental", not true incremental view maintenance like @materialize.com or Snowflake's dynamic tables, but it's a neat step towards IVM.
crunchydata.com
pg_incremental: Incremental Data Processing in Postgres | Crunchy Data Blog
We are excited to release a new open source extension called pg_incremental. pg_incremental works with pg_cron to do incremental batch processing for data aggregations, data transformations, or import...
0213
Reposted by Marco Slot
tldraw @tldraw.com · 18/12/2024
We just launched tldraw computer
621241
Marco Slot @marcoslot.com · 17/12/2024
End-to-end demo of the new pg_incremental extension. There's raw events table and a summary table containing view counts. You then define a pipeline using an insert..select command, and keeps running that to do fast, reliable, incremental processing in the background.
162
Marco Slot @marcoslot.com · 17/12/2024
There are many incremental processing solutions, but they seem to never quite do what I need. I decided to build an extension that just keeps running the same command in Postgres with different parameters to do fast, reliable incremental data processing. That's pg_incremental. 1/n
1123
Reposted by Marco Slot
nolen @itseieio.bsky.social · 06/12/2024
how do you all remember every UUID? I find it really hard. so I wrote them all down on every uuid dot com the list has fast search across all 2^122 values (so you can find your favorites) - hoping to add some social features like "trending UUIDs" soon!
481161276
Marco Slot @marcoslot.com · 05/12/2024
Everything gets simpler once you have transactions on Iceberg tables. A useful pattern is to keep track of loaded files in a Postgres table the same transaction that loads the file into Iceberg, such that each file is loaded exactly once www.crunchydata.com/blog/iceberg...
crunchydata.com
Iceberg ahead! Analyzing Shipping Data in Postgres | Crunchy Data Blog
Marco shows off how work with Iceberg and Postgres together in Crunchy Data Warehouse. He creates an Iceberg data set from AIS shipping data, batch loads public data daily, and then runs reports and m...
051
Reposted by Marco Slot
Richard Bishop @richardb.bsky.social · 04/12/2024
182
Marco Slot @marcoslot.com · 04/12/2024
The transactional piece is the real game changer, especially for data pipelines You can write these little procedures that load data from S3 or http in many different formats (using plain COPY) and store the fact that you loaded the file in the same transaction. Super simple, reliable, repeatable.
050
Reposted by Marco Slot
Craig @craigkerstiens.com · 04/12/2024
Want a complete Iceberg + transactional database system in one unified experience? (Postgres) Well we shipped just that last month. The ease of use, yet power underneath really is a game changer for data systems - www.crunchydata.com/blog/crunchy...
crunchydata.com
Crunchy Data Warehouse: Postgres with Iceberg for High Performance Analytics | Crunchy Data Blog
We are excited to release Crunchy Data Warehouse, a modern data warehouse for Postgres. Crunchy Data Warehouse combines Postgres with Iceberg, Parquet, and data lake formats for fast analytics queries...
1132
Marco Slot @marcoslot.com · 04/12/2024
Exciting announcement. Played with this a bit today. It's like an Iceberg starter kit that bundles storage with the catalog and compaction. Systems looking to support Iceberg writes will have a relatively easy time targeting S3 tables catalog, because they don't need to add compaction. 1/n
1124
Marco Slot @marcoslot.com · 03/12/2024
It would be nice if databases had a more convenient way of saying "columns A and B are coupled and unique, one implies the other" and adjusted SQL syntax to it. now you need to always either group by both ID and email or do fake aggregations like any_value(email), which seems like an anti-pattern.
180
Reposted by Marco Slot
Karen Jex @karenhjex.bsky.social · 28/11/2024
Can't wait to put my new decoration on my Christmas tree (when my husband isn't looking)! #PostgreSQL #crafting
Crocheted christmas tree ornament in the shape of a blue elephant head, decorated with holly leaves and berries
2213
Reposted by Marco Slot
Renee Shah @reneeshah.bsky.social · 25/11/2024
We hosted our annual(ish) Databases Dinner in SF last week! There were 3 recurring themes among database practitioners: 1) The need for an alternative to Spark, 2) The desire for an alternative to SQL, and 3) Lots of praise for Postgres’ scalability and workloads moving to it.
4222
Reposted by Marco Slot
Craig @craigkerstiens.com · 25/11/2024
In case you missed it last week, we released Crunchy Data Warehouse - transforming Postgres into a best in class data warehouse with full Iceberg support - www.crunchydata.com/blog/crunchy...
crunchydata.com
Crunchy Data Warehouse: Postgres with Iceberg for High Performance Analytics | Crunchy Data Blog
We are excited to release Crunchy Data Warehouse, a modern data warehouse for Postgres. Crunchy Data Warehouse combines Postgres with Iceberg, Parquet, and data lake formats for fast analytics queries...
1133
Reposted by Marco Slot
Joe Wood @joewood.me · 21/11/2024
CrunchyData Warehouse all looks very impressive: www.crunchydata.com/blog/crunchy.... Impressed to see them build on industry primitives like Apache Iceberg, and materializing PostgreSQL data marts so seamlessly. Super interested to see how this works with fast moving datasets.
crunchydata.com
072
Reposted by Marco Slot
Peter Boncz @peterabcz.bsky.social · 22/11/2024
At cwi.nl/en/events/dijkstra-awards/cwi-lectures-dijkstra-fellowship we awarded the Dijkstra Fellowship to Marcin Zukowski for his contributions to DB architecture in a full room & great speakers: Allison Lee, Viktor Leis, @andypavlo.bsky.social @hannes.muehleisen.org @madelonhulsebos.bsky.social
Marcin Zukowski showing the Dijkstra Fellowship Award (steel artwork) shortly after receiving it at CWI.
0268
Reposted by Marco Slot
Andy Pavlo @andypavlo.bsky.social · 22/11/2024
The Database Capital of Europe
Centrum Wiskunde & Informatica
Amsterdam, the Netherlands
2808
Marco Slot @marcoslot.com · 22/11/2024
The promise of Apache Iceberg. It's like pong for databases. #databs
0204
Marco Slot @marcoslot.com · 22/11/2024
Some interesting things can be built with this. Next up, an efficient way to poll for new appends? 🤔
120
Marco Slot @marcoslot.com · 22/11/2024
It did not make sense to me why Q28 from ClickBench was slower on Crunchy than on other systems, but we did not have enough time to thoroughly profile it at the time. Kept bugging me though. Turns out the problem was that we were running the actual benchmark, but almost no one else was 😅
052