Sign in

Polars

@pola.rs
1.1K followers 3 following 94 posts

Dataframes powered by a multithreaded, vectorized query engine, written in Rust.

PostsRepliesMedia
Polars @pola.rs · 21/09/2026
Polars 2.0.0rc2 has been released. We expect this to be the last release candidate and that Polars 2.0 will go live next week. Install now by running `pip install --pre polars` Link to the full release guide: docs.pola.rs/releases/upg...
092
Polars @pola.rs · 17/09/2026
Profile any Polars query to view its progress and optimize performance with one line change: pl.Config.enable_monitoring() Read the full blog: pola.rs/posts/profil...
020
Polars @pola.rs · 02/09/2026
We are happy to announce the first pre-release of Polars 2.0. Polars 2.0 will bring the streaming engine as a default and improves overall strictness of Polars. Give it a spin and share your thoughts. We expect the final Polars 2.0 release in 3-4 weeks. pola.rs/posts/announ...
2180
Polars @pola.rs · 28/08/2026
The Polars GPU engine now runs on a new streaming backend, built with NVIDIA on their RapidsMPF library. • Queries are no longer bound by GPU memory. • The same query scales from one GPU to many. Full post: pola.rs/posts/gpu-st...
030
Polars @pola.rs · 26/08/2026
We've released Python Polars 1.44. Some of the highlights: • Schema evolution when writing to Iceberg • Adaptive rate limiting for cloud I/O • Correlated subqueries in SQL Blog post: pola.rs/posts/polars... Full changelog: github.com/pola-rs/pola...
0101
Polars @pola.rs · 25/08/2026
As datasets grow, your tools can limit what analysis is possible. That causes teams to slows down or forces them to work with samples rather than the whole dataset. BMLL Technologies tackled this issue by migrating their workloads, resulting in a 48x performance improvement. pola.rs/posts/case-b...
010
Polars @pola.rs · 20/08/2026
You can turn a column into labeled buckets. With if/elif/else logic takes one chained expression: pl.when().then().otherwise(). Conditions are checked top to bottom, exactly like Python's if/elif/else. Chain as many as you need, and every branch runs vectorized in Rust rather than row by row.
080
Polars @pola.rs · 11/08/2026
Time series datasets are seldom perfect. Polars gives you a whole menu for repairing them, and the right pick depends on what you need.
0110
Polars @pola.rs · 28/07/2026
Aggregating 500GB of compressed parquet with 16,000,000,000 rows to display it in an enterprise-ready Plotly dashboard. Check the dashboard: polars-cloud-demo.plotly.app And the blog: pola.rs/posts/market...
051
Polars @pola.rs · 23/07/2026
We've released Python Polars 1.43. Some of the highlights: • pl.list() • ewm_sum() and ewm_sum_by() • Faster joins on hive-partitioned data Blog post: pola.rs/posts/polars... Full changelog: github.com/pola-rs/pola...
040
Polars @pola.rs · 16/07/2026
A float column can hold two different kinds of missing: null for a value that is absent, and NaN for arithmetic that had no valid answer (think 0.0/0.0). Polars keeps them strictly separate.
071
Polars @pola.rs · 14/07/2026
Reading a terabyte dataset from S3 goes fastest on a cluster of small machines, while heavy joins run fastest on one big machine. We benchmarked single node Polars against distributed Polars on the same total resources, and the bottleneck of your query decides the winner. pola.rs/posts/single...
050
Polars @pola.rs · 07/07/2026
We've been busy in Q2 2026. Read all the highlights in the latest Polars in Aggregate: pola.rs/posts/polars...
020
Polars @pola.rs · 29/06/2026
We've released Python Polars 1.42. Some of the highlights: • Adaptive cloud I/O concurrency - up to 4x faster reads • Contradictory filter elimination • is_sorted() for DataFrame and Expressions Blog post: pola.rs/posts/polars... Full changelog: github.com/pola-rs/pola...
060
Polars @pola.rs · 25/06/2026
How well do LLMs migrate pandas to Polars by themselves? We tested how well Claude translates a pandas corpus to Polars. Results were promising but not perfect. To improve this, we built a Polars skill that helps the agent. Read the full post here: pola.rs/posts/llm-po...
051
Polars @pola.rs · 16/06/2026
Distributed Polars is 3x faster than Spark on the PDSH benchmark and up to 7.8x faster on individual queries. Read the full benchmark post here: https:/pola.rs/posts/polars-pyspark-benchmarks/
Figure showing the speed-up of Polars compared to PySpark per query of the PDS-H benchmark.
280
Polars @pola.rs · 03/06/2026
Run Polars' distributed engine on your own infrastructure. Deploy a distributed Polars cluster on any Kubernetes setup (EKS, AKS, GKE, or minikube) and get a query dashboard with past queries, advanced query profiling, Open-lineage support, and more. More at pola.rs/posts/polars...
020
Polars @pola.rs · 14/05/2026
Polars supports a full Iceberg roundtrip on the streaming engine. You can scan an Iceberg table with scan_iceberg(), transform it lazily, and write the result back with sink_iceberg(). Useful for workflows like data redaction or compliance cleanup.
040
Polars @pola.rs · 31/03/2026
Realtime query profiling of Polars In this post we use the query profiler in Polars Cloud to optimize the infrastructure configuration for a specific query. This results in a 54% faster and 64% cheaper query with only five runs. Read all about it here: pola.rs/posts/query-...
040
Polars @pola.rs · 26/03/2026
We've released Polars Cloud client 0.6.0. Some of the highlights: • Improved UX for query profiling • Compute Scratchpad Alpha • Improved distributed query planning • Breaking: `LazyFrameRemote.execute` is now blocking by default
030
Polars @pola.rs · 26/02/2026
pl.from_repr() constructs a DataFrame or Series directly from its printed string representation. This can be useful in unit tests: instead of rebuilding expected DataFrames through dictionaries with typecasting, the schema is encoded in the header and the values are right there in the table.
040
Polars @pola.rs · 13/02/2026
str.len_bytes() vs str.len_chars() len_bytes: ~20x faster, counts UTF-8 bytes len_chars: counts actual Unicode characters - Use len_bytes for ASCII data (IDs, hashes) - Use len_chars for anything multilingual len_bytes is O(1) metadata lookup, len_chars is O(n) traversal.
1131
Polars @pola.rs · 13/01/2026
We just released Polars 1.37, here are the highlights: Improved Streaming Sinks: 1.14x-1.88x speedup, ~10% of the original memory. Streaming Compressed CSVs Faster SQL Ordering pl.PartitionBy min_by / max_by (see below) Series.sql() Free-Threading Support Python 3.9 Support Dropped musl Builds
1131
Polars @pola.rs · 18/12/2025
Did you know about pl.corr()? The problem with data aggregation is that it can hide what's really going on. Below you can find Simpson's Paradox Sometimes the devil really is in the details.
070
Polars @pola.rs · 11/12/2025
"We adopted Polars to meet strict technical requirements, but the result went beyond simple optimization. The 30x performance improvement gave us the unexpected opportunity to do more." Read about how Rabobank deployed Polars in a critical enterprise production environment: lnkd.in/eZFPcxRw
060
Polars @pola.rs · 09/12/2025
We've just released 1.36.0. Here are the highlights: Highlights: 🧩 Extension Types 🛟 Float16 Support ↪️ LazyFrame.pivot() 👀 DataFrame.show() 🗄️ SQL Parity: Added Window functions ⏱️ Parquet writer: 2.2x runtime improvement Find the full release notes here: github.com/pola-rs/pola...
3191
Polars @pola.rs · 19/11/2025
Polars recently shipped some performance upgrades and long-awaited features: 🏆 Decimal Type Now Stable 🏆 Aggregation over the List and Array Types ✨Other new features: Streaming ewm_mean() Expr.item() Expr.rolling_rank() pl.union() Read more: github.com/pola-rs/pola...
Code example of the Decimal type
2100
Polars @pola.rs · 13/08/2025
Are you looking to get started with Polars over the summer? We've partnered with @datacamp.bsky.social to create an interactive course that covers the fundamentals so you can write your next query with Polars. The course is free till the end of August: www.datacamp.com/courses/intr...
Free DataCamp course announcement - Learn Polars, the high-performance data processing library that engineers love. Interactive Introduction to Polars course available at no cost. Partnership between DataCamp and Polars.
0115
Polars @pola.rs · 28/05/2025
We've partnered with @datacamp.bsky.social to create an interactive Polars course. Learn the fundamentals and get familiar with our API through hands-on exercises. The course is available for everyone and free until the end of August. Start the free course here: www.datacamp.com/courses/intr...
Course title slide with dark blue background showing 'FREE COURSE' in green text, 'Introduction to Polars' as the main white heading, and logos for DataCamp and Polars at the bottom.
040
Polars @pola.rs · 01/05/2025
Polars has gotten 4x faster than Polars! 🚀 In the last months, the team has worked incredibly hard on the new-streaming engine and the results pay off. It is incredibly fast, and beats the Polars in-memory engine by a factor of 4 on a 96vCPU machine.
4163
Polars @pola.rs · 30/04/2025
Polars provides a number of xxx_horizontal operations. These expressions perform computations across columns. (Or along rows, depending on how you look at it.) If your horizontal operation isn’t implemented, you can use the general-purpose fold.
Diagram showing examples of applying horizontal operations to a dataframe with random numerical values. The example shows the expressions max_horizontal, sum_horizontal, mean_horizontal, and cum_sum_horizontal. The first three produce a numerical column and the expression cum_sum_horizontal produces a struct column with as many fields as there are input columns/series.
020
Polars @pola.rs · 19/03/2025
Polars provides 3 functions you can use to generate temporal ranges: date_range, datetime_range, and time_range. These can be executed eagerly or lazily. You can also customize the interval between consecutive values and whether the start/end points are included.
Diagram showing how the functions date_range, datetime_range, and time_range, can be used to generate series of consecutive temporal values. The interval between consecutive values can be changed with the parameter interval and the two endpoints may or may not be included in the generated range, depending on the value the parameter closed is set to.
081
Polars @pola.rs · 11/03/2025
It does:
110
Polars @pola.rs · 27/02/2025
The expression over can be used to compute expressions within isolated groups. This means you can do computations per group without having to group first and then explode after. In this example, we rank swimmers based on their time, but within their race type.
Diagram showing how the use of the expression over impacts the result of an expression. Using rank on a dataframe will rank all swimmers by their time, but that means that the fastest swimmer of the second race type is actually ranked 3rd. By using over, the fastest swimmer of each race type gets the rank 1.
280
Polars @pola.rs · 20/02/2025
The context filter lets you filter out rows from a dataframe based on some conditions. Within an aggregation, you can also use filter to filter values from aggregated groups. In this example we ignore unverified times when computing the current record.
Diagram showing how filter works inside aggregations. We use an aggregation to compute the current record for different swimming race types and we use a filter to ignore unverified times, while also using information about unverified times to compute another column.
070
Polars @pola.rs · 03/02/2025
The expression `clip` is pretty straightforward: You provide a lower and an upper bound, and Polars makes sure all values fall within those bounds. If a value is too small/too large, it's replaced by the bound. Bounds can be literals, other columns, or arbitrary expressions.
Diagram showing how the expression clip works. The diagram shows `clip` used with integers and with `datetime` objects. The diagram also shows the usage of literals, other columns, and arbitrary expressions, as the bounds.
280
Polars @pola.rs · 07/01/2025
Join our webinar with NVIDIA on January 28 for an in-depth session on how the GPU engine works, from collecting your query to parallel execution on the GPU. Sign up at info.nvidia.com/nvidia-polar... See you there?
Webinar banner with the NVIDIA and Polars logos and the webinar title "Inside the RAPIDS-Accelerated Polars GPU Engine" that is taking place on the 28th of January of 2025. On the side, you can see a digital/technological rendition of what could be a black hole or another space-themed object.
060
Polars @pola.rs · 06/01/2025
Polars 1.19 comes with support for arbitrary predicates in join_where. This means that inequality joins are now more flexible than ever! Here is a small example of something you couldn't do before:
0171
Polars @pola.rs · 12/12/2024
Polars supports dynamic aggregations based on time windows via the function `group_by_dynamic`. To use it, you specify a date(time) column to group by, and then determine the windows over which values are aggregated. Note how data points can fall within multiple windows 👇
Diagram showing how the window boundaries of 10-year long windows align neatly with the decades and the decade halfpoints, 1980, 1985, 1990, etc, although the first datapoint is from 1981.
1162
Polars @pola.rs · 10/12/2024
Can't remember how many days each month has? (Me neither!) Memorise this Polars snippet instead. Using some calendar-aware functions, we can get the answer in a tidy dataframe, as the diagram below shows.
Diagram with a snippet of Polars code that creates a dataframe mapping each month to the number of days in that month.
The stars of the snippet are the functions `date_range` and `group_by_dynamic`.

The code:

from datetime import date

print(
    pl.date_range(
        start=date(2024, 1, 1),
        end=date(2025, 1, 1),
        interval="1d",
        closed="left",  # Don't include `end`
        eager=True,
    ).to_frame("d")
    .group_by_dynamic("d", every="1mo")
    .agg(pl.len().alias("days_in_month"))
    .select(
        pl.col("d").dt.month().alias("month"),
        pl.col("days_in_month"),
    )
)
2161
Polars @pola.rs · 28/11/2024
You want to join two tables on their ID column, but only when the dates in one table fall within the range of the other table. Polars lets you do that with `join_where`, which supports inequality joins through the use of inequality predicates. Here's an example 👇
Diagram exemplifying how `join_where` performs inequality joins.

We're joining two dataframes that have 3 rows each and each row has a corresponding row on the other dataframe because they have matching IDs.

The dataframe on the right has a column `dt` with date values, which fall within the range created by the columns `start` and `end` of the other dataframe, except for the third row, in which `dt` falls outside the range.

When using `join_where` with the predicate `pl.col("dt").is_between("start", "end")`, the third row of each dataframe will not be matched and won't show up in the join result, which only joined the first 2 rows of each dataframe.
091
Polars @pola.rs · 22/11/2024
Polars has essentially 18 different data types. If you are unsure what each type is, the conversion table below might help you. Each Polars data type is presented next to the **most similar** Python type.
Diagram establishing a connection between Polars data types and similar Python types.

The connections are as follows:
Boolean is most similar to bool
Int8, Int16, Int32, and Int64 is most similar to int with restrictions
UInt8, UInt16, UInt32, and UInt64 is most similar to int with restrictions
Float32 and Float64 is most similar to float
Decimal is most similar to decimal.Decimal
String is most similar to str
Binary is most similar to bytes
Date is most similar to datetime.date
Time is most similar to datetime.time
Datetime is most similar to datetime.datetime
Duration is most similar to datetime.timedelta
Array is most similar to numpy.array
List is most similar to list
Categorical is most similar to enum.StrEnum
Enum is most similar to enum.StrEnum
Struct is most similar to typing.TypedDict
Object is most similar to object
Null is most similar to None
2101
Polars @pola.rs · 21/11/2024
How to “expand” ranges like "3-5" across new rows with the values 3, 4, 5? This comes straight from our Discord server (discord.com/invite/4UfP5...)
A diagram shows how to use str.split, list.first / list.last, cast, pl.int_ranges, and explode, all together, to turn a dataframe where a column may contain ranges like "3-5" into a similar dataframe where all ranges have been expanded, or exploded, across multiple rows.

The full code is:

range_start = pl.col("nrs").str.split("-").list.first().cast(pl.Int64)
range_end = pl.col("nrs").str.split("-").list.last().cast(pl.Int64)
df.with_columns(pl.int_ranges(range_start, range_end + 1)).explode("nrs")
072
Polars @pola.rs · 20/11/2024
Why is there a `struct` data type? A single expression produces a single column, so expressions like `value_counts` need to output structs to map the values to their counts. With that said, do you understand why `.struct.unnest` doesn't break the 1 expr = 1 column principle?
Diagram showing how `value_counts` produces a column with struct values, mapping column values to their counts.
We then show how to use `.struct.field` to extract a single field from the struct and how to use `.struct.unnest` to extract all fields into corresponding columns.
072