Sign in

Andy Grove

@andygrove.io
2.7K followers 84 following 56 posts

Apache Arrow & DataFusion PMC Member. Original creator of Apache DataFusion.

PostsRepliesMedia
Reposted by Andy Grove
Columnar @columnar.tech · 13/06/2026
This week we announced a new ADBC driver for Apache DataFusion. DataFusion is Rust-native, but now you can use it from 10+ other languages through the same zero-copy Arrow ADBC interface that connects to 30+ other databases and query engines. Learn more: docs.adbc-drivers.org/drivers/data...
DataFusion logo centered on a white background with radial lines connecting it to language logos for Python, JavaScript, TypeScript, Go, Java, C#, C++, Rust, Ruby, R, and Kotlin.
0133
Andy Grove @andygrove.io · 13/05/2026
Apache DataFusion Comet 0.16.0 is now available, and I’m pretty excited about this one! Comet now has first-class support for Spark 4.0 and 4.1 with full ANSI support for overflow checks, ANSI casts, try_* variants of some expressions, all accelerated. datafusion.apache.org/blog/output/...
datafusion.apache.org
Apache DataFusion Comet 0.16.0 Release - Apache DataFusion Blog
060
Andy Grove @andygrove.io · 13/05/2026
There's a new DataFusion subproject to provide official Java bindings. It's very early, and it's a great time for new contributors to get involved! github.com/apache/dataf...
github.com
GitHub - apache/datafusion-java: Java bindings for Apache DataFusion
Java bindings for Apache DataFusion. Contribute to apache/datafusion-java development by creating an account on GitHub.
050
Andy Grove @andygrove.io · 04/05/2026
Much of this write up is true, although it mixes me up with @andrewlamb1111.bsky.social in places. It also neglects to mention the hard work of literally hundreds of people in the community that took my initial donations and transformed them into something amazing today.
180
Andy Grove @andygrove.io · 13/04/2026
PR just landed to add Delta Lake support to Apache DataFusion's Comet accelerator for Apache Spark. Anyone interested in helping with reviews? I do not have much experience with Delta Lake. github.com/apache/dataf...
github.com
feat: Native Delta Lake scan via delta-kernel-rs by schenksj · Pull Request #3932 · apache/datafusion-comet
Summary Adds native Delta Lake read support to Comet using delta-kernel-rs for log replay, matching all optimizations in the existing Iceberg native scan path. Delta tables (spark.sql("SELECT ...
040
Andy Grove @andygrove.io · 27/01/2026
Helpful advice. Thanks, Claude.
0170
Reposted by Andy Grove
Andy Pavlo @andypavlo.bsky.social · 05/01/2026
I've posted my latest recap of the world of databases: www.cs.cmu.edu/~pavlo/blog/... All the hot topics from the last year: • More Postgres action! • MCP for everyone! • MongoDB gets litigious with FerretDB! • File formats! • Market movements! • The richest person in the history of the world!
cs.cmu.edu
Databases in 2025: A Year in Review
The world tried to kill Andy off but he had to stay alive to to talk about what happened with databases in 2025.
17625
Andy Grove @andygrove.io · 30/12/2025
Is there anyone in my network with DuckDB skills who could review a PR that runs a Python script to compare the performance of DataFusion and DuckDB for some simple SQL queries? github.com/apache/dataf...
github.com
feat: Add microbenchmark for string functions by andygrove · Pull Request #26 · apache/datafusion-benchmarks
This PR adds microbenchmarks for scanning a Parquet file and evaluating a single string expression per row. The benchmark runs against DuckDB and DataFusion and compares the results. Assuming that ...
021
Andy Grove @andygrove.io · 17/12/2025
There is a new Comet issue to discuss the future of Iceberg support and whether we should focus on using the iceberg-rust or Java implementation of Iceberg. Please add your thoughts if this is something that you care about! github.com/apache/dataf...
github.com
Future of Iceberg Support in Comet · Issue #2921 · apache/datafusion-comet
What is the problem the feature request solves? Comet currently has two different approaches to scanning Iceberg tables. One approach is based on integrating with the Iceberg Java library, and the ...
010
Andy Grove @andygrove.io · 22/10/2025
On behalf of the DataFusion PMC, I'm excited to announce the release of version 0.11.0 of the Comet accelerator for Apache Spark! datafusion.apache.org/blog/2025/10...
datafusion.apache.org
Apache DataFusion Comet 0.11.0 Release - Apache DataFusion Blog
060
Andy Grove @andygrove.io · 06/10/2025
It’s steak night tonight and our dog is patiently waiting for her share.
0111
Andy Grove @andygrove.io · 24/09/2025
I like the name “RAD stack” for this.
080
Andy Grove @andygrove.io · 18/09/2025
Check out the latest release of the Comet accelerator for Apache Spark datafusion.apache.org/blog/2025/09...
datafusion.apache.org
Apache DataFusion Comet 0.10.0 Release - Apache DataFusion Blog
0101
Reposted by Andy Grove
Yaroslav Tkachenko @sap1ens.com · 15/09/2025
Introducing Iron Vector: native, columnar, vectorized, high-performance accelerator for Apache Flink SQL and Table API built on top of Rust, Arrow and DataFusion. Reduce your Flink compute cost by up to 2x or handle 2x more data with the same infrastructure.
1165
Reposted by Andy Grove
Rust Language @rust-lang.org · 12/09/2025
We received reports of a phishing campaign targeting crates​.io users. Do not click on links asking to authenticate to protect your account. More information: blog.rust-lang.org/2025/09/12/c...
blog.rust-lang.org
crates.io phishing campaign | Rust Blog
Empowering everyone to build reliable and efficient software.
011256
Reposted by Andy Grove
Andrew Lamb @andrewlamb1111.bsky.social · 04/09/2025
Thanks to @clflushopt.bsky.social, make massive TPCH datasets with tpchgen-cli 2.0: SF1000 (1TB raw, 220GB in @ApacheParquet ) in less than 10 mins (6m45s) on aging laptop Try it now: pip install tpchgen-cli tpchgen-cli --scale-factor 1000 --parts 100 --format=parquet github.com/clflushopt/t...
041
Reposted by Andy Grove
Phil Eaton @eatonphil.bsky.social · 04/09/2025
I've been helping our analytics team integrate our DataFusion-based query engine for Postgres into EDB Postgres Distributed and finally here's an end-to-end demo. You get HA Postgres plus seamless replication and DataFusion-based queries. This query turned out 6x faster than PG.
1144
Andy Grove @andygrove.io · 22/08/2025
How my day is going
070
Andy Grove @andygrove.io · 20/08/2025
We now have a roadmap section in the Comet contributor guide, in case anyone was wondering what we are focusing on lately and what features will be arriving in future releases. datafusion.apache.org/comet/contri...
datafusion.apache.org
Comet Roadmap — Apache DataFusion Comet documentation
060
Reposted by Andy Grove
Alex P @ifesdjeen.bsky.social · 18/07/2025
Cassandra Team at Apple is searching for a fresh grad / person early in their career to join our ranks in SF/Bay Area! Come work on super interesting problems with world class team. Help us build better Cassandra! Ping me if you’re interested! jobs.apple.com/en-us/detail...
jobs.apple.com
Software Engineer, ASE Cassandra Storage - Jobs - Careers at Apple
Apply for a Software Engineer, ASE Cassandra Storage job at Apple. Read about the role and find out if it’s right for you.
01610
Andy Grove @andygrove.io · 02/05/2025
It took me a really long time to understand the flow of execution between JVM and native code during query execution in Comet. I wish I had thought about adding a tracing capability earlier. github.com/apache/dataf...
github.com
perf: Add performance tracing capability by andygrove · Pull Request #1706 · apache/datafusion-comet
Which issue does this PR close? Closes #1705 Rationale for this change This feature makes it possible to visualize the flow of calls during query execution. What changes are included in this PR?...
051
Reposted by Andy Grove
Tim Saucer @timsaucer.bsky.social · 07/04/2025
We're pleased to announce that Apache DataFusion in Python 46.0.0 is released! Since the last announcement post we've had a lot of great features and new contributors. Please check out the blog post with details. datafusion.apache.org/blog/2025/03... #DataFusion #Python #DataFrame #PyData #Apache
datafusion.apache.org
Apache DataFusion Python 46.0.0 Released - Apache DataFusion Blog
062
Andy Grove @andygrove.io · 02/04/2025
We have a position open in the Spark team at Apple, in our Cupertino, CA office. The role would include working on Apache DataFusion Comet. jobs.apple.com/en-us/detail...
jobs.apple.com
Senior Software Development Engineer (Apache Spark) - Apple Data Platform - Jobs - Careers at Apple
Apply for a Senior Software Development Engineer (Apache Spark) - Apple Data Platform job at Apple. Read about the role and find out if it’s right for you.
0136
Andy Grove @andygrove.io · 21/03/2025
We have TPC-H benchmarks for single node with a small scale factor in the contributors guide. We only benchmark against Spark though and not against Spark RAPIDS. datafusion.apache.org/comet/contri...
datafusion.apache.org
Apache DataFusion Comet: Benchmarks Derived From TPC-H — Apache DataFusion Comet documentation
010
Andy Grove @andygrove.io · 21/03/2025
Here's the blog post announcing Comet 0.7.0 datafusion.apache.org/blog/2025/03...
datafusion.apache.org
Apache DataFusion Comet 0.7.0 Release - Apache DataFusion Blog
071
Andy Grove @andygrove.io · 19/03/2025
I hate to say it, but "it depends". I'd recommend running your own benchmarks for your specific workloads. Performance will also vary greatly by environment (number of CPUs vs GPUs, different GPU types, and so on).
200
Andy Grove @andygrove.io · 19/03/2025
DataFusion Comet 0.7.0 is now available in Maven. We'll be publishing a blog post next week with all the details. The repo has been updated with the latest benchmark results. For single executor TPC-H @ 100 GB, we now see a 2.2x increase over Spark (up from 2x in 0.6.0). github.com/apache/dataf...
github.com
GitHub - apache/datafusion-comet: Apache DataFusion Comet Spark Accelerator
Apache DataFusion Comet Spark Accelerator. Contribute to apache/datafusion-comet development by creating an account on GitHub.
1111
Andy Grove @andygrove.io · 18/02/2025
One month on, and I have zero regrets about quitting Facebook & Instagram. I have replaced the scrolling time with listening to podcasts. I now stay in touch with family overseas via email and photo sharing, and I use Snapchat for sharing photos with immediate family, privately. Works great.
1140
Reposted by Andy Grove
Commonhaus Foundation @commonhaus.org · 18/02/2025
Chris Riccomini (@chris.blue) shares his thoughts on Open Source foundations: Apache, CNCF, Commonhaus. He also explains why Commonhaus is a better fit for SlateDB cnr.sh/posts/compar...
cnr.sh
Comparing Apache, CNCF, and Commonhaus | cnr.sh
I've used open source projects for over 30 years and contributed for about 20 of those. My first interaction with an open source foundation was with Apache when I began working with Apache Hadoop ...
0126
Andy Grove @andygrove.io · 18/02/2025
Comet 0.6.0 has been released. This is a smaller release than usual now that we have moved to an approximately monthly release cadence to match core DataFusion. datafusion.apache.org/blog/2025/02...
datafusion.apache.org
Apache DataFusion Comet 0.6.0 Release - Apache DataFusion Blog
060
Andy Grove @andygrove.io · 12/02/2025
Ballista 43.0.0 has been released, and now provides seamless integration with DataFusion. datafusion.apache.org/blog/2025/02...
datafusion.apache.org
Apache DataFusion Ballista 43.0.0 Released - Apache DataFusion Blog
0160
Andy Grove @andygrove.io · 29/01/2025
Check out this excellent presentation from @robtandy.bsky.social on his work with the DataFusion Ray project from last week's DataFusion community meetup. It is a great overview of how to build a distributed system on top of DataFusion. www.youtube.com/watch?v=ceTo...
youtube.com
Apache DataFusion Community Meeting 2025/01/22 08:57 MST - Recording
YouTube video by Datadog
1102
Andy Grove @andygrove.io · 26/01/2025
This Week in DataFusion Comet (Jan 26): github.com/apache/dataf...
github.com
This Week in Comet (Jan 26) · Issue #1342 · apache/datafusion-comet
Introduction These notes reflect things I am personally involved in or thinking about and may not cover all activities. Feel free to add comments for anything that I missed. Previous week's issue: ...
050
Andy Grove @andygrove.io · 23/01/2025
Is this using Arrow and/or DataFusion? If so, our Discord is probably a good place to ask. datafusion.apache.org/contributor-...
datafusion.apache.org
Communication — Apache DataFusion documentation
100
Andy Grove @andygrove.io · 18/01/2025
I've finally decided to quit using Facebook. My feed is overwhelmed with nonsense content that I am not interested in and cannot seem to block. It is a real shame, though, because it was a good way to stay connected with family. Is there a viable alternative? What are others using instead?
450
Andy Grove @andygrove.io · 18/01/2025
This week in DataFusion Comet (Jan 18). Inspired by @andrewlamb1111.bsky.social's weekly updates in DataFusion core, I am going to start doing the same in Comet to help keep the community updated on current events. github.com/apache/dataf...
github.com
This week in Comet (Jan 18) · Issue #1305 · apache/datafusion-comet
Introduction Inspired by @alamb's weekly updates in DataFusion, I thought it would be a good idea to do something similar in Comet to keep contributors updated on what is happening in the project. ...
081
Andy Grove @andygrove.io · 17/01/2025
DataFusion Comet 0.5.0 has been released. See blog post for details. datafusion.apache.org/blog/2025/01...
datafusion.apache.org
Apache DataFusion Comet 0.5.0 Release - Apache DataFusion Blog
0101
Reposted by Andy Grove
Ian Cook @ian.columnar.tech · 13/01/2025
2025 is shaping up to be a breakout year for fast query result transfer with Apache Arrow. But what exactly makes it so fast? David Li, Matt Topol, and I break it down in this new blog post: arrow.apache.org/blog/2025/01...
arrow.apache.org
How the Apache Arrow Format Accelerates Query Result Transfer
Arrow speeds up query result transfer by slashing (de)serialization overheads. We outline five key attributes of the Arrow format that enable this.
0209
Andy Grove @andygrove.io · 10/01/2025
DataFusion Comet performance has been improving recently and now demonstrates a ~2x speedup compared to Spark for single node TPC-H. There is more to do, but this feels like a significant milestone. github.com/apache/dataf...
github.com
docs: Update TPC-H benchmark results by andygrove · Pull Request #1257 · apache/datafusion-comet
Which issue does this PR close? N/A Rationale for this change There have been a number of performance improvements since 0.4.0 so we should update the benchmark results. The results are based on ...
0222
Reposted by Andy Grove
Andy Pavlo @andypavlo.bsky.social · 01/01/2025
Buckle up because we're banging into the new year with my annual retrospective of the last year in databases! Highlights include license change blowback, Databricks vs. Snowflake gangwar, @duckdb.org's shotgun weddings, and buying a quarterback to impress your lover: www.cs.cmu.edu/~pavlo/blog/...
cs.cmu.edu
Databases in 2024: A Year in Review
Andy rises from the ashes of his dead startup and discusses what happened in 2024 in the database game.
1019963
Reposted by Andy Grove
Tim Saucer @timsaucer.bsky.social · 24/12/2024
The latest in my set of career advice: How to maintain a healthy work-life balance by setting boundaries. I'm moving these over to blog posts on my website for easier sharing. #CareerPlusPlus #WorkAdvice
timsaucer.com
Healthy Work-Life Boundaries - Tim Saucer
I try to offer the tips I have for how to set and stick to boundaries to maintain a healthy work-life balance.
362
Andy Grove @andygrove.io · 23/12/2024
Honestly, I have no idea. I wonder if @timsaucer.bsky.social knows
100
Andy Grove @andygrove.io · 21/12/2024
Yet another impressive DataFusion Python release. This time with interoperability improvements with PyCapsule and FFI. datafusion.apache.org/blog/2024/12...
datafusion.apache.org
Apache DataFusion Python 43.1.0 Released - Apache DataFusion Blog
1111
Reposted by Andy Grove
Sem 🏴️‍ @ssinchenko.bsky.social · 23/11/2024
1/5 Just made my first contribution to the #Datafusion #Comet - a native physical execution engine for #Apache #Spark! 🚀 While Spark with it's row oriented model and code generation approach is quite good on average, there is almost always a faster specific solution.
161
Andy Grove @andygrove.io · 21/11/2024
"We want to remove DataFusion from everything" 😭 It's good to hear detailed feedback for a use case where DataFusion didn't work out—some important lessons. youtu.be/Sor3KZpmbHg?...
youtu.be
Biting the Bullet: Rebuilding GlareDB from the Ground Up (Sean Smith)
YouTube video by CMU Database Group
1101
Andy Grove @andygrove.io · 20/11/2024
Apache DataFusion Comet 0.4.0 has been released! See the blog post for details. datafusion.apache.org/blog/2024/11...
datafusion.apache.org
Apache DataFusion Comet 0.4.0 Release
<!–
0132
Andy Grove @andygrove.io · 20/11/2024
Apache DataFusion's Ballista distributed query engine has quietly been getting a makeover over the past several months. I'm excited to see the project being maintained again! Performance is 🔥. See the updated README for more information. github.com/apache/dataf...
github.com
GitHub - apache/datafusion-ballista: Apache DataFusion Ballista Distributed Query Engine
Apache DataFusion Ballista Distributed Query Engine - apache/datafusion-ballista
0194
Andy Grove @andygrove.io · 18/11/2024
What a brilliant use of AI! news.virginmediao2.co.uk/o2-unveils-d...
news.virginmediao2.co.uk
O2 unveils Daisy, the AI granny wasting scammers’ time - Virgin Media O2
O2 has today unveiled the newest member of its fraud prevention team, 'Daisy'. As ‘Head of Scammer Relations’, this state-of-the-art AI Granny's mission is to talk with fraudsters and waste as much of...
1237
Reposted by Andy Grove
Julien Le Dem @julien.ledem.net · 11/11/2024
If you’re wondering what all the fuss is about lately, this is what’s driving the adoption of Iceberg to implement the Open Data Lake. sympathetic.ink/2024/11/07/T... It’s all columnar data in blob storage anyways. You may as well be the one taking advantage of it.
What, it’s all blob storage?
Always has been.
1449
Reposted by Andy Grove
Paul Dix @pauldix.bsky.social · 09/11/2024
So I have this theory that DataFusion, despite being a SQL engine, will actually enable a new breed of data systems to create non-SQL languages for working with data. Here's the idea...🧵
2247