Sign in

Polars

@pola.rs
1.1K followers 3 following 93 posts

Dataframes powered by a multithreaded, vectorized query engine, written in Rust.

PostsRepliesMedia
Polars @pola.rs · 21/09/2026
Polars 2.0.0rc2 has been released. We expect this to be the last release candidate and that Polars 2.0 will go live next week. Install now by running `pip install --pre polars` Link to the full release guide: docs.pola.rs/releases/upg...
081
Polars @pola.rs · 17/09/2026
Profile any Polars query to view its progress and optimize performance with one line change: pl.Config.enable_monitoring() Read the full blog: pola.rs/posts/profil...
020
Polars @pola.rs · 02/09/2026
We are happy to announce the first pre-release of Polars 2.0. Polars 2.0 will bring the streaming engine as a default and improves overall strictness of Polars. Give it a spin and share your thoughts. We expect the final Polars 2.0 release in 3-4 weeks. pola.rs/posts/announ...
2180
Polars @pola.rs · 28/08/2026
The Polars GPU engine now runs on a new streaming backend, built with NVIDIA on their RapidsMPF library. • Queries are no longer bound by GPU memory. • The same query scales from one GPU to many. Full post: pola.rs/posts/gpu-st...
030
Polars @pola.rs · 26/08/2026
We've released Python Polars 1.44. Some of the highlights: • Schema evolution when writing to Iceberg • Adaptive rate limiting for cloud I/O • Correlated subqueries in SQL Blog post: pola.rs/posts/polars... Full changelog: github.com/pola-rs/pola...
0101
Polars @pola.rs · 25/08/2026
As datasets grow, your tools can limit what analysis is possible. That causes teams to slows down or forces them to work with samples rather than the whole dataset. BMLL Technologies tackled this issue by migrating their workloads, resulting in a 48x performance improvement. pola.rs/posts/case-b...
010
Polars @pola.rs · 20/08/2026
You can turn a column into labeled buckets. With if/elif/else logic takes one chained expression: pl.when().then().otherwise(). Conditions are checked top to bottom, exactly like Python's if/elif/else. Chain as many as you need, and every branch runs vectorized in Rust rather than row by row.
080
Polars @pola.rs · 11/08/2026
Time series datasets are seldom perfect. Polars gives you a whole menu for repairing them, and the right pick depends on what you need.
0110
Polars @pola.rs · 04/08/2026
We've released Polars Cloud client 0.10.0: • Stream query results into Python with `sink_batches()` • A new experimental query planner: Miso • Distributed `pl.collect_all()` • Hive-partition aware scans • On-Prem HDFS support Blog post: pola.rs/posts/polars... Changelog: github.com/pola-rs/pola...
pola.rs
Announcing Polars Cloud 0.10.0
Polars Cloud 0.10.0 streams distributed query results into Python with sink_batches(), adds an experimental miso query planner, distributed pl.collect_all(), experimental HDFS support, optimized hive-...
030
Polars @pola.rs · 28/07/2026
Aggregating 500GB of compressed parquet with 16,000,000,000 rows to display it in an enterprise-ready Plotly dashboard. Check the dashboard: polars-cloud-demo.plotly.app And the blog: pola.rs/posts/market...
051
Polars @pola.rs · 23/07/2026
We've released Python Polars 1.43. Some of the highlights: • pl.list() • ewm_sum() and ewm_sum_by() • Faster joins on hive-partitioned data Blog post: pola.rs/posts/polars... Full changelog: github.com/pola-rs/pola...
040
Polars @pola.rs · 16/07/2026
A float column can hold two different kinds of missing: null for a value that is absent, and NaN for arithmetic that had no valid answer (think 0.0/0.0). Polars keeps them strictly separate.
071
Polars @pola.rs · 14/07/2026
Reading a terabyte dataset from S3 goes fastest on a cluster of small machines, while heavy joins run fastest on one big machine. We benchmarked single node Polars against distributed Polars on the same total resources, and the bottleneck of your query decides the winner. pola.rs/posts/single...
050
Polars @pola.rs · 07/07/2026
We've been busy in Q2 2026. Read all the highlights in the latest Polars in Aggregate: pola.rs/posts/polars...
020
Polars @pola.rs · 03/07/2026
We've released Polars Cloud client 0.9.0. • Expressions now run distributed • 17% Faster cloud I/O • Breaking: `ClusterContext` now uses `uri=` The `compute_address=` keyword is superseded by `uri=`. Blog post: pola.rs/posts/polars... Full changelog: github.com/polars-inc/p...
pola.rs
Announcing Polars Cloud 0.9.0
Polars Cloud 0.9.0 runs expressions distributed, ships large cloud I/O performance gains, a reworked ClusterContext API, a distributed Iceberg sink, and more.
011
Polars @pola.rs · 29/06/2026
We've released Python Polars 1.42. Some of the highlights: • Adaptive cloud I/O concurrency - up to 4x faster reads • Contradictory filter elimination • is_sorted() for DataFrame and Expressions Blog post: pola.rs/posts/polars... Full changelog: github.com/pola-rs/pola...
060
Polars @pola.rs · 25/06/2026
How well do LLMs migrate pandas to Polars by themselves? We tested how well Claude translates a pandas corpus to Polars. Results were promising but not perfect. To improve this, we built a Polars skill that helps the agent. Read the full post here: pola.rs/posts/llm-po...
051
Polars @pola.rs · 16/06/2026
Distributed Polars is 3x faster than Spark on the PDSH benchmark and up to 7.8x faster on individual queries. Read the full benchmark post here: https:/pola.rs/posts/polars-pyspark-benchmarks/
Figure showing the speed-up of Polars compared to PySpark per query of the PDS-H benchmark.
280
Polars @pola.rs · 03/06/2026
Run Polars' distributed engine on your own infrastructure. Deploy a distributed Polars cluster on any Kubernetes setup (EKS, AKS, GKE, or minikube) and get a query dashboard with past queries, advanced query profiling, Open-lineage support, and more. More at pola.rs/posts/polars...
020
Polars @pola.rs · 26/05/2026
We've released Python Polars 1.41. Some of the highlights: • Faster Parquet metadata decoding • Nested common subplan elimination • LazyFrame.gather() Blog post: pola.rs/posts/polars...
pola.rs
Announcing Polars 1.41
DataFrames for the new era
061
Polars @pola.rs · 14/05/2026
Polars supports a full Iceberg roundtrip on the streaming engine. You can scan an Iceberg table with scan_iceberg(), transform it lazily, and write the result back with sink_iceberg(). Useful for workflows like data redaction or compliance cleanup.
040
Polars @pola.rs · 30/04/2026
Handling schema changes in Polars. This blog post maps the four shapes of schema change (a new column appears, an expected one disappears, a type drifts, or one breaks) to the Polars solution that handles each, across CSV, multi-file Parquet, Delta Lake, and Apache Iceberg. pola.rs/posts/schema...
pola.rs
Handling Schema Issues in Polars
DataFrames for the new era
130
Polars @pola.rs · 23/04/2026
We've released Python Polars 1.40. The following now runs on the streaming engine • AsOf joins with a `by` argument • Basic elementwise over() • More expressions: cov(), corr(), interpolate(), skew(), kurtosis(), and entropy() and much more! github.com/pola-rs/pola...
github.com
Release Python Polars 1.40.0 · pola-rs/polars
🏆 Highlights Add streaming support for grouped AsOf join (#27293) ⚠️ Deprecations Deprecate support for dataframe interchange protocol (#27214) 🚀 Performance improvements Create IR slice from ...
090
Polars @pola.rs · 16/04/2026
We've been busy in Q1 2026. 12 releases. 778 PRs. 95 contributors (thank you!). Read all the highlights in the latest Polars in Aggregate: pola.rs/posts/polars...
pola.rs
Polars in Aggregate: Streaming Expands, Lakehouse I/O, and Cloud Profiling
DataFrames for the new era
130
Polars @pola.rs · 09/04/2026
Polars loves sorted data! If your data is already sorted, you can get a performance boost up to 18x when joining your datasets. Read all about it in our latest blog post: pola.rs/posts/stream...
pola.rs
040
Polars @pola.rs · 31/03/2026
Realtime query profiling of Polars In this post we use the query profiler in Polars Cloud to optimize the infrastructure configuration for a specific query. This results in a 54% faster and 64% cheaper query with only five runs. Read all about it here: pola.rs/posts/query-...
040
Polars @pola.rs · 26/03/2026
We've released Polars Cloud client 0.6.0. Some of the highlights: • Improved UX for query profiling • Compute Scratchpad Alpha • Improved distributed query planning • Breaking: `LazyFrameRemote.execute` is now blocking by default
030
Polars @pola.rs · 17/03/2026
Jensen: "All of these platforms are processing DataFrames. This is the ground truth of business. ... Now we will have AI use structured data. And we are going to accelerate the living daylights out of it." Polars DataFrames are at the core of the AI revolution. www.youtube.com/watch?v=jw_o...
youtube.com
NVIDIA GTC Keynote 2026
YouTube video by NVIDIA
051
Polars @pola.rs · 12/03/2026
We've released Python Polars 1.39. Some of the highlights: • Streaming AsOf join, enabling memory-efficient time-series joins. • sink_iceberg() for writing to Iceberg tables • Streaming cloud downloads for scan_csv(), scan_ndjson(), and scan_lines() github.com/pola-rs/pola...
github.com
Release Python Polars 1.39.0 · pola-rs/polars
🚀 Performance improvements Lower arg_{min,max} to streaming engine (#26845) Additional IR slice pushdown after filter pushdown (#26815) Streaming first/last on Enum through physical (#26783) Fast ...
0162
Polars @pola.rs · 26/02/2026
pl.from_repr() constructs a DataFrame or Series directly from its printed string representation. This can be useful in unit tests: instead of rebuilding expected DataFrames through dictionaries with typecasting, the schema is encoded in the header and the values are right there in the table.
040
Polars @pola.rs · 17/02/2026
Easily scale Polars queries from Airflow. Our latest blog post walks through different patterns to run distributed Polars queries using Airflow: fire-and-forget execution, parallel queries, multi-stage pipelines, and manual cluster shutdowns. Read more here: pola.rs/posts/airflo...
pola.rs
Orchestrating Polars Cloud Queries with Apache Airflow
DataFrames for the new era
050
Polars @pola.rs · 13/02/2026
str.len_bytes() vs str.len_chars() len_bytes: ~20x faster, counts UTF-8 bytes len_chars: counts actual Unicode characters - Use len_bytes for ASCII data (IDs, hashes) - Use len_chars for anything multilingual len_bytes is O(1) metadata lookup, len_chars is O(n) traversal.
1131
Polars @pola.rs · 05/02/2026
We've released Python Polars 1.38. Some of the highlights: • (De)Compression support on text based sources and sinks • scan_lines() to read text files • Merge join in the Streaming engine Link to the complete changelog: github.com/pola-rs/pola...
github.com
Release Python Polars 1.38.0 · pola-rs/polars
⚠️ Deprecations Deprecate retries=n in favor of storage_options={"max_retries": n} (#26155) 🚀 Performance improvements Enable zero-copy object_store put upload for IPC sink (#26288) Resolve file...
091
Polars @pola.rs · 03/02/2026
We refactored the Categorical in 1.31. The new Categories object gives you: • Control over the physical type (UInt8/16/32) • Named categories with namespaces • Parallel updates without locks • Automatic garbage collection Full read: pola.rs/posts/catego...
040
Reposted by Polars
Ritchie Vink @ritchie46.bsky.social · 02/02/2026
In 1-2 weeks we land live query profiling in Polars Cloud. See exactly how many rows are consumed and produced per operation. Which operation takes most runtime, and watch the data flow through live, like water. 😍
091
Reposted by Polars
Oli Hawkins @olihawkins.com · 01/02/2026
Looks like me and @eadehemingway.bsky.social are going to be running a workshop on data analysis with @pola.rs at NICAR in March. Maybe see some of you there!
1114
Polars @pola.rs · 13/01/2026
We just released Polars 1.37, here are the highlights: Improved Streaming Sinks: 1.14x-1.88x speedup, ~10% of the original memory. Streaming Compressed CSVs Faster SQL Ordering pl.PartitionBy min_by / max_by (see below) Series.sql() Free-Threading Support Python 3.9 Support Dropped musl Builds
1131
Polars @pola.rs · 18/12/2025
Did you know about pl.corr()? The problem with data aggregation is that it can hide what's really going on. Below you can find Simpson's Paradox Sometimes the devil really is in the details.
070
Polars @pola.rs · 11/12/2025
"We adopted Polars to meet strict technical requirements, but the result went beyond simple optimization. The 30x performance improvement gave us the unexpected opportunity to do more." Read about how Rabobank deployed Polars in a critical enterprise production environment: lnkd.in/eZFPcxRw
060
Polars @pola.rs · 09/12/2025
We've just released 1.36.0. Here are the highlights: Highlights: 🧩 Extension Types 🛟 Float16 Support ↪️ LazyFrame.pivot() 👀 DataFrame.show() 🗄️ SQL Parity: Added Window functions ⏱️ Parquet writer: 2.2x runtime improvement Find the full release notes here: github.com/pola-rs/pola...
3191
Polars @pola.rs · 03/12/2025
It’s been a year since the last Polars in Aggregate. Since then, we've shipped 37 releases, merged over 2,300 PRs, and built two new engines. Here are the biggest highlights: ☁️ Polars Cloud is Live 🚀 Next-Gen Streaming engine. 🔢 Stable Decimals & Int128 and more pola.rs/posts/polars...
pola.rs
Polars in Aggregate: Polars Cloud, Streaming engine, and New Data Types
DataFrames for the new era
080
Polars @pola.rs · 26/11/2025
Citizens cut query times from 80 to 8 minutes by adopting Polars, but the transformation went beyond speed. It provided a "grammar of business logic, improving maintainability and unlocking complexity without heavy backend engineering. Read the full case study here: pola.rs/posts/case-c...
pola.rs
Supercharging Analytics with Polars: A Case Study in Analyst Empowerment
DataFrames for the new era
060
Polars @pola.rs · 19/11/2025
Polars recently shipped some performance upgrades and long-awaited features: 🏆 Decimal Type Now Stable 🏆 Aggregation over the List and Array Types ✨Other new features: Streaming ewm_mean() Expr.item() Expr.rolling_rank() pl.union() Read more: github.com/pola-rs/pola...
Code example of the Decimal type
2100
Polars @pola.rs · 14/10/2025
At the recent Polars meetup, Oliver & Daniel discussed how they migrated a pandas + SQL to Polars using Dataframely: - 22x speedup - 3x lower memory - 50% code reduction - Native dataframe validation with minimal overhead using Dataframely Watch: www.youtube.com/watch?v=TL-3...
youtube.com
Polars Meetup #3 - Polars x Dataframely by Oliver Borchert and Daniel Elsner
YouTube video by Polars
040
Polars @pola.rs · 09/10/2025
Swiss insurer La Mobilière refactored their risk model to Polars, achieving 5-10x speedups and enabling actuaries to run millions of simulation years on laptops. A scale previously unfeasible with pandas due to memory and single-core limitations. pola.rs/posts/case-m...
pola.rs
Polars helps coping with black swan events at La Mobilière
DataFrames for the new era
081
Polars @pola.rs · 29/09/2025
We raised €18M in Series A led by Accel to build fast data processing at any scale. All on Polars. pola.rs/posts/series...
pola.rs
Polars raises €18M Series A to build fast, ergonomic data processing at any scale
DataFrames for the new era
0101
Polars @pola.rs · 25/09/2025
The recordings of our third meetup are now available on Youtube! Watch the session of Gijs Burghoorn, core developer @ Polars, here: youtu.be/xc5IsfwKRKE. In his talk he discussed how and why we optimize our Parquet reader.
youtu.be
Polars Meetup #3 - Vectorized Parquet by Gijs Burghoorn
YouTube video by Polars
291
Polars @pola.rs · 22/09/2025
Polars Cloud client 0.3.0 is released. You can now spawn >100k queries to a single cluster and we load balance them gracefully. Additionally, the query planning now is posted as a worker task and can be cancelled by the user. github.com/pola-rs/pola...
github.com
Release Polars Cloud Client 0.3.0 · pola-rs/polars-cloud-client
💥 Breaking changes Remove partitioned_by execution ✨ Enhancements Plan query on worker nodes Implement partition sink file path callback Add shuffle write data to observatory 🐞 Bug fixes Harde...
131
Polars @pola.rs · 15/09/2025
@decathlonfrance.bsky.social has adopted Polars across many workloads, reducing infrastructure complexity and overhead by running workloads on single machines instead of compute clusters. Learn more in the case study: pola.rs/posts/case-d...
pola.rs
Polars at Decathlon: Ready to Play?
DataFrames for the new era
3203
Polars @pola.rs · 03/09/2025
Today we launch Polars Cloud and the Public Beta of our Distributed Engine. Read the post to get started! pola.rs/posts/polars...
pola.rs
Launch of Polars Cloud and Distributed Polars
DataFrames for the new era
0121