I’d like to count the number of unique values when combining several columns at once. My idea so far was to use pl.struct(...).n_unique(), which works fine when I consider missing values as a unique value:
import polars as pl
df = pl.DataFrame({
"x": ["a", "a", "b", "b"],
"y": [1, 1, 2, None],
})
df.with_columns(foo=pl.struct("x", "y").n_unique())
shape: (4, 3)
βββββββ¬βββββββ¬ββββββ
β x β y β foo β
β --- β --- β --- β
β str β i64 β u32 β
βββββββͺβββββββͺββββββ‘
β a β 1 β 3 β
β a β 1 β 3 β
β b β 2 β 3 β
β b β null β 3 β
βββββββ΄βββββββ΄ββββββ
However, sometimes I want to exclude a combination from the count if it contains any number of missing values. In the example above, I’d like foo to be 2. However, using .drop_nulls() before counting doesn’t work and produces the same output as above.
df.with_columns(foo=pl.struct("x", "y").drop_nulls().n_unique())
Is there a way to do this using only Polars expressions?
>Solution :
pl.Expr.drop_nulls does not drop the row as the entirety of the struct is indeed not null.
To still achieve the desired result, you can filter out all rows which contain a null values in any of the columns of interest using pl.Expr.filter.
(
df
.with_columns(
foo=pl.struct("x", "y").filter(
~pl.any_horizontal(pl.col("x", "y").is_null())
).n_unique()
)
)
shape: (4, 3)
βββββββ¬βββββββ¬ββββββ
β x β y β foo β
β --- β --- β --- β
β str β i64 β u32 β
βββββββͺβββββββͺββββββ‘
β a β 1 β 2 β
β a β 1 β 2 β
β b β 2 β 2 β
β b β null β 2 β
βββββββ΄βββββββ΄ββββββ