サンプル
polars 1.43.2: Port pandas DataFrame code to polars on Alpine, without the wheel surprise or the silently wrong answers
検証済みサンプル — pypi polars 1.43.2: Port pandas DataFrame code to polars on Alpine, without the wheel surprise or the silently wrong answers. python 3.12 · linux…
sha256:a238e626062c628c548646e46144b6aac21a07492cf10e1617eb4286c853ac05
このネットワークが提供するのは一つだけです。ビルドされるサンプル。サンドボックスで実行し、署名済みの受領証を保管します。等級はつけず、何も保証しません — 同じコードがあなたの環境でビルドされるかは測定していません。
合格した契約受領証を提出した異なる署名鍵の数です。1 なら作者だけ、2 以上なら他の誰かもビルドしています。鍵は自己生成で背後に登録された身元がないため、数えているのは人ではなく鍵です。
MIT-0
実行証拠
宣言された環境と署名済みの実行を分けてあります。このサンプルが何をどこで実行したかをそのまま確認できます。
- 証拠の基準
- 署名済みコントラクト合格
- 検証レシート
- 3
- ビルドした署名鍵
- 3
宣言された環境
python 3.12 linux · musl x64 python 3.12 python pip
検証実行環境
| 環境 | コントラクト | ステージ | 実行日 |
|---|---|---|---|
| python 3.12 · linux alpine/x64 · docker ed25519:a2ec939a4c60e243 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · pypi@1 |
2026-08-14 |
| python 3.12 · linux alpine/x64 · docker ed25519:d91480838ac982c9 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · pypi@1 |
2026-08-18 |
| python 3.12 · linux alpine/x64 · docker ed25519:c1973797be207ac4 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · pypi@1python:3.12-alpine@sha256:d09d15e60962… |
2026-09-08 |
ケース
MIGRATION- ゴール
- Port pandas DataFrame code to polars on Alpine, without the wheel surprise or the silently wrong answers
- シンボル
-
- polars.DataFrame
- polars.LazyFrame
- polars.col
- polars.len
- DataFrame.filter
- DataFrame.select
- DataFrame.with_columns
- DataFrame.with_row_index
- DataFrame.group_by
- GroupBy.agg
- LazyFrame.collect
- LazyFrame.explain
- LazyFrame.collect_schema
- Expr.map_elements
- Expr.alias
- Series.is_null
- Series.is_nan
- Series.fill_null
- Series.fill_nan
- Series.eq_missing
- polars.exceptions.ColumnNotFoundError
- polars.exceptions.DuplicateError
- polars.exceptions.PerformanceWarning
- DataFrame.vstack
- DataFrame.hstack
- DataFrame.extend
- DataFrame.insert_column
- DataFrame.drop_in_place
- 環境
- python 3.12
- 作成日
- 2026-08-14T12:56:07Z
コントラクト
- assert pip took the musllinux wheel for the engine, polars_runtime_32-1.43.2-cp310-abi3-musllinux_1_2_x86_64.whl, reconstructed from the recorded dist-info tag
- assert polars itself is a pure-Python py3-none-any distribution carrying no compiled file, which depends on polars-runtime-32==1.43.2 for the engine
- assert the engine's shared object is named .abi3.so rather than this interpreter's EXT_SUFFIX, so one cp310-abi3 build serves 3.12
- assert the interpreter reports musl and polars.__version__ agrees with the distribution metadata, which is the check a half-installed polars fails
- assert a DataFrame has none of index, loc, iloc, at, iat, values, reset_index, sort_values, assign, query or apply
- assert filtering renumbers rows from zero, and with_row_index before the filter is what carries the original positions through, as UInt32 data
- assert brackets take a column by name and a row by position, and an unknown name raises ColumnNotFoundError
- assert a boolean mask in brackets selects columns rather than rows, raising ValueError when its length does not match the column count
- assert the same mask on a frame with as many columns as rows returns the wrong columns with no error at all
- assert filter with a pl.col boolean expression returns the intended rows and select takes strings, pl.col and an anchored regex
- assert an Expr has no truth value, so `and` and `or` both raise TypeError, and dropping the parentheses around & raises that identical message
- assert separate positional predicates mean AND, and the keyword form is equality against a literal
- assert with_columns returns a new frame, leaves the input's columns and values untouched, and has no inplace parameter
- assert item assignment on a DataFrame raises TypeError naming DataFrame.with_columns as the replacement
- assert the new frame shares the untouched column's Arrow allocation with the old one, so returning a frame copies the column list rather than the data
- assert polars does keep a mutating escape hatch: hstack and vstack take in_place= and hand the same object back, and extend, insert_column and drop_in_place mutate the receiver
- assert an unaliased expression is named after the first column it mentions, so with_columns replaces that column instead of adding one
- assert explain returns the plan as a string while the per-row Python hook has still not been called once
- assert the optimizer moves the filter below the projection, the reverse of both the unoptimized plan and the order written
- assert collect runs the hook on only the two surviving rows, where the identical eager pipeline runs it on all six
- assert a LazyFrame naming a missing column is built without complaint and raises ColumnNotFoundError only when the plan is resolved, which explain also does
- assert a LazyFrame has no shape, and asking it for columns emits a PerformanceWarning because that resolves the schema
- assert is_null and is_nan disagree on every hole, and is_nan over a null is null rather than False
- assert null_count reports one null while the mean is nan, count excludes the null and pl.len counts every row
- assert a column holding a NaN and no null reports null_count zero with a nan mean, so a missing-data check written as null_count() == 0 passes on it
- assert min and max ignore the NaN that sum and mean propagate
- assert comparison against a null yields null while eq_missing yields False
- assert polars orders NaN above every number and compares it equal to itself, the opposite of Python, so a > 0 filter keeps the NaN and drops the null
- assert nulls sort first in both directions unless nulls_last is set
- assert drop_nulls keeps NaN and drop_nans keeps null, neither fill touches the other's hole, and fill_nan(None) collapses the two into one
- assert an integer column with a missing value stays Int64 and its aggregates skip the null on both sides of the division
- assert group_by returns the group keys first, then one column per agg expression in the order given
- assert maintain_order gives first-appearance order, which is not sorted order, and without it only the set of rows is defined
- assert two aggregates over one column collide as DuplicateError, worded one way inside agg and another inside select
ファイル
- csx.json
- requirements.txt
- src/__init__.py
- src/frames.py
- src/install.py
- test/contract.py
ソース
{"case":{"caseId":"case:sha256:c86ef023cfbe932ce36bdc24956446a20a8c31b31fd71b9a38e9054e56607bf0","constraints":{"libc":"musl","runtime":"python"},"contract":["assert pip took the musllinux wheel for the engine, polars_runtime_32-1.43.2-cp310-abi3-musllinux_1_2_x86_64.whl, reconstructed from the recorded dist-info tag","assert polars itself is a pure-Python py3-none-any distribution carrying no compiled file, which depends on polars-runtime-32==1.43.2 for the engine","assert the engine's shared object is named .abi3.so rather than this interpreter's EXT_SUFFIX, so one cp310-abi3 build serves 3.12","assert the interpreter reports musl and polars.__version__ agrees with the distribution metadata, which is the check a half-installed polars fails","assert a DataFrame has none of index, loc, iloc, at, iat, values, reset_index, sort_values, assign, query or apply","assert filtering renumbers rows from zero, and with_row_index before the filter is what carries the original positions through, as UInt32 data","assert brackets take a column by name and a row by position, and an unknown name raises ColumnNotFoundError","assert a boolean mask in brackets selects columns rather than rows, raising ValueError when its length does not match the column count","assert the same mask on a frame with as many columns as rows returns the wrong columns with no error at all","assert filter with a pl.col boolean expression returns the intended rows and select takes strings, pl.col and an anchored regex","assert an Expr has no truth value, so `and` and `or` both raise TypeError, and dropping the parentheses around \u0026 raises that identical message","assert separate positional predicates mean AND, and the keyword form is equality against a literal","assert with_columns returns a new frame, leaves the input's columns and values untouched, and has no inplace parameter","assert item assignment on a DataFrame raises TypeError naming DataFrame.with_columns as the replacement","assert the new frame shares the untouched column's Arrow allocation with the old one, so returning a frame copies the column list rather than the data","assert polars does keep a mutating escape hatch: hstack and vstack take in_place= and hand the same object back, and extend, insert_column and drop_in_place mutate the receiver","assert an unaliased expression is named after the first column it mentions, so with_columns replaces that column instead of adding one","assert explain returns the plan as a string while the per-row Python hook has still not been called once","assert the optimizer moves the filter below the projection, the reverse of both the unoptimized plan and the order written","assert collect runs the hook on only the two surviving rows, where the identical eager pipeline runs it on all six","assert a LazyFrame naming a missing column is built without complaint and raises ColumnNotFoundError only when the plan is resolved, which explain also does","assert a LazyFrame has no shape, and asking it for columns emits a PerformanceWarning because that resolves the schema","assert is_null and is_nan disagree on every hole, and is_nan over a null is null rather than False","assert null_count reports one null while the mean is nan, count excludes the null and pl.len counts every row","assert a column holding a NaN and no null reports null_count zero with a nan mean, so a missing-data check written as null_count() == 0 passes on it","assert min and max ignore the NaN that sum and mean propagate","assert comparison against a null yields null while eq_missing yields False","assert polars orders NaN above every number and compares it equal to itself, the opposite of Python, so a \u003e 0 filter keeps the NaN and drops the null","assert nulls sort first in both directions unless nulls_last is set","assert drop_nulls keeps NaN and drop_nans keeps null, neither fill touches the other's hole, and fill_nan(None) collapses the two into one","assert an integer column with a missing value stays Int64 and its aggregates skip the null on both sides of the division","assert group_by returns the group keys first, then one column per agg expression in the order given","assert maintain_order gives first-appearance order, which is not sorted order, and without it only the set of rows is defined","assert two aggregates over one column collide as DuplicateError, worded one way inside agg and another inside select"],"goal":"Port pandas DataFrame code to polars on Alpine, without the wheel surprise or the silently wrong answers","kind":"MIGRATION","packages":["pkg:pypi/polars@1.43.2","pkg:pypi/polars-runtime-32@1.43.2"],"schemaVersion":1,"symbols":["polars.DataFrame","polars.LazyFrame","polars.col","polars.len","DataFrame.filter","DataFrame.select","DataFrame.with_columns","DataFrame.with_row_index","DataFrame.group_by","GroupBy.agg","LazyFrame.collect","LazyFrame.explain","LazyFrame.collect_schema","Expr.map_elements","Expr.alias","Series.is_null","Series.is_nan","Series.fill_null","Series.fill_nan","Series.eq_missing","polars.exceptions.ColumnNotFoundError","polars.exceptions.DuplicateError","polars.exceptions.PerformanceWarning","DataFrame.vstack","DataFrame.hstack","DataFrame.extend","DataFrame.insert_column","DataFrame.drop_in_place"]},"contractCommand":["python","test/contract.py"],"environment":{"arch":"x64","ecosystem":"pypi","executionContext":"python","language":"python","libc":"musl","os":"linux","packageManager":"pip","runtime":"python","runtimeVersion":"3.12","schemaVersion":1},"license":"MIT-0","packages":["pkg:pypi/polars@1.43.2","pkg:pypi/polars-runtime-32@1.43.2"],"schemaVersion":1,"symbols":["polars.DataFrame","polars.LazyFrame","polars.col","polars.len","DataFrame.filter","DataFrame.select","DataFrame.with_columns","DataFrame.with_row_index","DataFrame.group_by","GroupBy.agg","LazyFrame.collect","LazyFrame.explain","LazyFrame.collect_schema","Expr.map_elements","Expr.alias","Series.is_null","Series.is_nan","Series.fill_null","Series.fill_nan","Series.eq_missing","polars.exceptions.ColumnNotFoundError","polars.exceptions.DuplicateError","polars.exceptions.PerformanceWarning","DataFrame.vstack","DataFrame.hstack","DataFrame.extend","DataFrame.insert_column","DataFrame.drop_in_place"],"verifierAdapter":"pypi@1"}
polars==1.43.2
polars-runtime-32==1.43.2
"""Writing polars when your hands still type pandas.
The pandas habits that carry over produce either a loud error or, once, a
quietly wrong answer. This module is the pipeline written the polars way,
with the trap named at each step.
What is not here, because it does not exist: there is no index. No .loc, no
.iloc, no .index, no reset_index. A polars DataFrame is a list of named
columns and rows are addressed by position only for the moment you are
looking at them. Filtering renumbers; if you need the original row number
to survive a filter, you add it as a real column with with_row_index BEFORE
filtering, and then it is data like any other column.
Selection is expressions, not bracket indexing. df.select(pl.col("x")) is
the API; brackets exist but mean something else, and that is where the one
silent failure lives. df[bool_series] does not filter rows — a boolean mask
in brackets selects COLUMNS. With a mask whose length does not match the
column count you get a ValueError naming the real meaning. With a frame
that happens to have as many columns as rows, which is exactly the shape of
a small test fixture, you get no error at all and a frame of the wrong
columns. The pandas reflex df[df["qty"] > 2] is the one line here that can
pass review. Write df.filter(pl.col("qty") > 2).
Mutation is where "polars always hands back a new frame" turned out to be
too strong, and the measurement is the more useful answer. The expression
API really does not mutate: select, filter and with_columns return a new
frame, there is no inplace= to pass one, and df["new"] = values raises
TypeError pointing at with_columns. The frame-stacking methods are the
exception. hstack, vstack and shrink_to_fit take in_place=, and extend,
insert_column and drop_in_place mutate the receiver outright. The in_place=
form returns the SAME object rather than None, so
combined = df.vstack(other, in_place=True)
reads exactly like the pure version while every other reference to df has
just changed underneath it.
Returning a new frame is cheap because the Arrow buffers are shared: a
column the operation did not touch is the same allocation in both frames,
so what gets copied is the column list, not the data.
"""
import polars as pl
def with_line_totals(df: pl.DataFrame) -> pl.DataFrame:
"""Add a derived column. The input frame is untouched afterwards.
The naming rule catches people once: an expression's output name is the
name of the FIRST column it mentions, so pl.col("qty") * 2 is called
"qty" and with_columns replaces the qty column instead of adding one.
Adding a column rather than overwriting one is what .alias() is for.
"""
return df.with_columns(
(pl.col("qty") * pl.col("unit_price")).alias("line_total")
)
def large_orders(df: pl.DataFrame, minimum: float, region: str) -> pl.DataFrame:
"""filter takes one boolean expression, and the parentheses are load bearing.
Two conditions combine with & and |, never with `and` and `or`: an Expr
has no truth value, so `and` raises TypeError before polars ever sees
the query. The parentheses matter because & binds tighter than the
comparison operators, so
pl.col("line_total") >= minimum & pl.col("region") == region
parses as line_total >= (minimum & col) == region, a chained comparison,
which Python evaluates with `and` and which therefore dies with the same
"truth value of an Expr is ambiguous" as the first mistake. The error
text does not mention precedence, so the fix looks unrelated to the bug.
Passing the conditions as separate positional arguments is the version
with no precedence to get wrong: df.filter(a, b) means a AND b.
"""
return df.filter(
(pl.col("line_total") >= minimum) & (pl.col("region") == region)
)
def revenue_by_region(df: pl.DataFrame) -> pl.DataFrame:
"""Aggregate. The column order is documented, the ROW order is not.
Columns come out in a fixed order: the group keys first, then one column
per expression in the order given to agg. That much you can rely on.
Row order is the trap. pandas groupby sorts by the key by default;
polars group_by does not sort and does not preserve input order either,
because the groups are built in parallel. maintain_order=True buys back
first-appearance order at a cost in speed, and first appearance is still
not sorted order. If a test compares against a literal list of rows,
sort explicitly — which is why this ends in .sort("region") even though
maintain_order has already made it deterministic.
Every agg expression gets an alias for the same reason as above: two
expressions over the same column both default to that column's name and
the query fails with DuplicateError.
"""
return (
df.group_by("region", maintain_order=True)
.agg(
pl.col("line_total").sum().alias("revenue"),
pl.len().alias("orders"),
)
.sort("region")
)
def audited_plan(frame: pl.LazyFrame, hook) -> pl.LazyFrame:
"""A lazy plan that leaves a countable trace of what the engine ran.
Nothing in here executes. A LazyFrame is a query plan, and the methods
on it only extend the plan: .collect() is the single call that runs it.
You can watch that from outside, which is what the hook is for — it is a
Python function called once per row the engine actually feeds through
map_elements, so an empty call list is proof that no data moved.
The plan is inspectable before it runs. .explain() returns the plan as a
string without touching a row, and it is worth reading, because the plan
is not the code you wrote. Filters are pushed down below projections, so
the map_elements hook here runs on the rows that survive the filter, not
on all of them. Written eagerly the identical three lines run the hook
over every row and throw most of the results away.
Building the plan is not the same as checking it. A LazyFrame naming a
column that does not exist is constructed without complaint; the
ColumnNotFoundError arrives when the schema is resolved, which .explain()
also does. So laziness defers the typo to the end of the pipeline, and
.explain() is the cheap way to surface it without running the query.
"""
return frame.with_columns(
pl.col("qty").map_elements(hook, return_dtype=pl.Int64).alias("audited")
).filter(pl.col("qty") >= 4)
def rating_health(df: pl.DataFrame) -> dict:
"""Count missing ratings. There are two kinds and they do not overlap.
pandas has one hole, NaN, and uses it for both "no value" and "not a
number". polars keeps them apart: null is absence, recorded in a
validity bitmap for every dtype, and NaN is a float value that happens
to be unordered. is_null() and is_nan() therefore answer different
questions, and is_nan() over a null returns null rather than False —
three-valued logic, all the way through.
The consequence people hit in production is here in the mean: nulls are
skipped by every aggregate, but a NaN is a real float and propagates, so
one NaN turns the average of a million rows into nan while null_count()
still reports zero. Only is_nan() finds it.
An integer column with a missing value is worth contrasting: polars
keeps it Int64 with a null, where pandas upcasts the column to float64
and writes NaN. Nothing about the dtype tells you a value went missing
in polars, and nothing about the value tells you it was an integer in
pandas.
"""
return df.select(
pl.len().alias("rows"),
pl.col("rating").count().alias("counted"),
pl.col("rating").null_count().alias("nulls"),
pl.col("rating").is_nan().sum().alias("nans"),
pl.col("rating").mean().alias("mean"),
).row(0, named=True)
def normalize_ratings(df: pl.DataFrame, default: float) -> pl.DataFrame:
"""Fill both holes. Neither fill function touches the other's.
fill_null leaves NaN alone and fill_nan leaves null alone, so the pandas
fillna(0) reflex covers half the cases and silently misses the other
half. Chaining fill_nan(None) first is the idiom that collapses the two
kinds into one: it rewrites NaN as null, and then a single fill_null
handles everything. drop_nulls and drop_nans split the same way — each
keeps the other's values.
"""
return df.with_columns(
pl.col("rating").fill_nan(None).fill_null(default)
)
"""What `pip install polars` actually puts on disk on Alpine, read off the
installed distributions rather than assumed from the version number.
Since the 1.4x line, polars on PyPI is two distributions, and this is the
part that breaks pinned installs. `polars` itself is pure Python:
polars-1.43.2-py3-none-any.whl (847 kB)
and the compiled Rust engine lives in a separate distribution it depends on,
which is where the wheel tags are and where all 55 MB went:
polars_runtime_32-1.43.2-cp310-abi3-musllinux_1_2_x86_64.whl (55.2 MB)
Both of those are the files this seed's resolve stage downloaded. The
musllinux wheel is the one that matters here, and it belongs to the engine
rather than to polars: "does polars install on Alpine" is entirely a
question about polars-runtime-32. It does — 1.43.2 publishes a
musllinux_1_2 wheel for this arch, and pip took it.
The tag says cp310 on a 3.12 interpreter and that is not a mismatch. It is
abi3, the CPython stable ABI, so one build serves 3.10 upward and the shared
object is named _polars_runtime.abi3.so rather than carrying this
interpreter's EXT_SUFFIX (.cpython-312-x86_64-linux-musl.so). A version
matrix that expects a wheel per minor version will not find one.
The trap in the split, measured: a requirements file that pins only
`polars` and installs with --no-deps does not fail at import. It warns.
UserWarning: Polars binary is missing!
`import polars` then succeeds, polars.__version__ is the empty string, and
the first real call dies with NameError: name 'PySeries' is not defined —
a NameError, not an ImportError, from inside the library. So a smoke test
that imports the package passes on a broken install, and the failure
surfaces later at an unrelated line. The check worth making is the one
below: that polars.__version__ is non-empty and matches the distribution
metadata. Pinning the full closure, polars and polars-runtime-32 together,
is what actually prevents it.
"""
import email
import importlib.metadata as metadata
import sysconfig
def installed(name: str) -> dict:
"""The dist-info receipt for one distribution: version, tags, deps."""
dist = metadata.distribution(name)
wheel = email.message_from_string(dist.read_text("WHEEL") or "")
return {
"name": dist.metadata["Name"],
"version": dist.version,
"tags": wheel.get_all("Tag") or [],
"root_is_purelib": (wheel["Root-Is-Purelib"] or "").strip() == "true",
"requires_dist": dist.metadata.get_all("Requires-Dist") or [],
"shared_objects": [str(f) for f in (dist.files or []) if str(f).endswith(".so")],
}
def wheel_filename(name: str) -> str:
"""Reconstruct the wheel pip downloaded, from the tag it recorded.
A wheel is named name-version-tag.whl with the name normalized to
underscores, so a dist-info carrying a single Tag pins the exact file.
That makes "which wheel did this environment get" checkable offline,
without a pip log.
"""
info = installed(name)
if len(info["tags"]) != 1:
raise AssertionError(f"expected one wheel tag, got {info['tags']}")
normalized = info["name"].replace("-", "_")
return f"{normalized}-{info['version']}-{info['tags'][0]}.whl"
def interpreter_libc() -> dict:
"""Evidence that this really is a musl interpreter, not just an image name."""
return {
"soabi": sysconfig.get_config_var("SOABI"),
"ext_suffix": sysconfig.get_config_var("EXT_SUFFIX"),
}
import inspect
import sys
import warnings
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
import polars as pl
from src.frames import (
audited_plan,
large_orders,
normalize_ratings,
rating_health,
revenue_by_region,
with_line_totals,
)
from src.install import installed, interpreter_libc, wheel_filename
NAN = float("nan")
ORDERS = {
"region": ["north", "south", "north", "east", "south", "north"],
"qty": [1, 3, 2, 5, 4, 2],
"unit_price": [10.0, 2.5, 4.0, 1.5, 3.0, 8.0],
}
def raises(expected, call):
"""Run call, require `expected`, and hand the exception back for checking."""
try:
call()
except expected as caught:
return caught
raise AssertionError(f"expected {expected.__name__}, nothing was raised")
# --- packaging: polars is two distributions on Alpine --------------------
core = installed("polars")
runtime = installed("polars-runtime-32")
# The pure-Python half: the universal tag, and nothing compiled in it.
assert core["tags"] == ["py3-none-any"], core["tags"]
assert core["root_is_purelib"] is True
assert core["shared_objects"] == [], core["shared_objects"]
# The engine is a separate distribution, and it is the one with the wheel
# tags. pip took the musllinux wheel here, which answers the only question
# that has to be settled before any of the rest of this file can run: a
# polars wheel does exist for musl, and 1.43.2 has one.
assert "polars-runtime-32==1.43.2" in core["requires_dist"], core["requires_dist"]
assert runtime["tags"] == ["cp310-abi3-musllinux_1_2_x86_64"], runtime["tags"]
assert runtime["root_is_purelib"] is False
assert (
wheel_filename("polars-runtime-32")
== "polars_runtime_32-1.43.2-cp310-abi3-musllinux_1_2_x86_64.whl"
)
# cp310 on a 3.12 interpreter is the stable ABI, not a mismatch: the shared
# object is named .abi3.so instead of this interpreter's EXT_SUFFIX, which
# is why one build covers 3.10 and up and why there is no per-minor wheel.
libc = interpreter_libc()
assert runtime["shared_objects"] == ["_polars_runtime_32/_polars_runtime.abi3.so"]
assert not runtime["shared_objects"][0].endswith(libc["ext_suffix"])
assert libc["ext_suffix"] == ".cpython-312-x86_64-linux-musl.so", libc
# And the musllinux match is real rather than an image name.
assert libc["soabi"] == "cpython-312-x86_64-linux-musl", libc
# The check that catches a half-installed polars. Measured: with only the
# `polars` distribution present, import succeeds with a UserWarning, and
# __version__ is the empty string until the first NameError from inside the
# library. A non-empty version agreeing with the metadata is the cheap proof
# that the binary half arrived; `import polars` alone is not.
assert pl.__version__ == core["version"] == "1.43.2"
# --- there is no index ---------------------------------------------------
df = with_line_totals(pl.DataFrame(ORDERS))
# None of the pandas row-addressing API exists, so code carrying it fails
# loudly at the attribute rather than subtly at the result.
for attribute in ("index", "loc", "iloc", "at", "iat", "values",
"reset_index", "sort_values", "assign", "query", "apply"):
assert not hasattr(df, attribute), attribute
# Row position is not an identity. Filtering renumbers from zero, and the
# only way to carry the original position past a filter is to write it into
# a column first, where it is ordinary UInt32 data.
assert df.filter(pl.col("qty") >= 4).with_row_index()["index"].to_list() == [0, 1]
assert df.with_row_index().filter(pl.col("qty") >= 4)["index"].to_list() == [3, 4]
assert df.with_row_index().schema["index"] == pl.UInt32
# --- selection is expressions, and brackets mean something else ----------
# Brackets still do the two things that look like pandas: a string picks a
# column and gives back a Series, an integer picks a row and gives back a
# one-row DataFrame. That mixture is already the reverse of pandas, where
# df[0] would be a column label lookup.
assert isinstance(df["qty"], pl.Series)
assert isinstance(df[0], pl.DataFrame) and df[0].shape == (1, 4)
assert raises(pl.exceptions.ColumnNotFoundError, lambda: df["nope"])
# The expression API is the one to write, and select takes plain strings,
# regexes and pl.all() through the same door.
assert df.select(pl.col("qty")).columns == ["qty"]
assert df.select("qty", "region").columns == ["qty", "region"]
assert df.select(pl.col("^unit_.*$")).columns == ["unit_price"]
# The one pandas reflex that can pass review. df[mask] is not a row filter
# in polars: a boolean mask in brackets selects COLUMNS. Against this frame
# the lengths disagree and it raises, and the message says what brackets
# actually meant.
mask = df["qty"] >= 4
assert mask.to_list() == [False, False, False, True, True, False]
assert "selecting columns by boolean mask" in str(
raises(ValueError, lambda: df[mask])
)
# Measured, and the reason this is worth a sample: when the frame has as
# many columns as the mask has rows — the shape of most test fixtures — the
# length check passes and there is no error at all. Here a mask meaning
# "rows 2 and 3" silently returns columns b and c, every row intact.
square = pl.DataFrame({"a": [1, 2, 3], "b": [4, 5, 6], "c": [7, 8, 9]})
wrong = square[square["a"] > 1]
assert wrong.columns == ["b", "c"] and wrong.shape == (3, 2)
assert wrong.to_dicts() == [{"b": 4, "c": 7}, {"b": 5, "c": 8}, {"b": 6, "c": 9}]
# What the same intent produces written as a filter.
right = square.filter(pl.col("a") > 1)
assert right.columns == ["a", "b", "c"] and right.shape == (2, 3)
# --- filter takes a boolean expression -----------------------------------
assert large_orders(df, 9.0, "north").to_dicts() == [
{"region": "north", "qty": 1, "unit_price": 10.0, "line_total": 10.0},
{"region": "north", "qty": 2, "unit_price": 8.0, "line_total": 16.0},
]
# An Expr has no truth value, so `and` and `or` never reach polars at all.
ambiguous = "the truth value of an Expr is ambiguous"
assert ambiguous in str(raises(TypeError, lambda: bool(pl.col("qty") > 2)))
assert ambiguous in str(
raises(TypeError, lambda: df.filter((pl.col("qty") > 1) and (pl.col("qty") < 5)))
)
assert ambiguous in str(
raises(TypeError, lambda: df.filter((pl.col("qty") > 1) or (pl.col("qty") < 5)))
)
# & binds tighter than the comparisons, so dropping the parentheses turns
# the whole thing into a chained comparison and fails with that identical
# message — the error names truthiness, never precedence, so the reported
# symptom points away from the missing brackets.
assert ambiguous in str(
raises(TypeError, lambda: df.filter(pl.col("qty") > 1 & pl.col("unit_price") < 5.0))
)
assert df.filter((pl.col("qty") > 1) & (pl.col("unit_price") < 5.0)).height == 4
# Separate positional predicates mean AND with no precedence to get wrong,
# and keyword form is equality against a literal.
assert df.filter(pl.col("qty") > 1, pl.col("unit_price") < 5.0).height == 4
assert df.filter(region="north").height == 3
# --- with_columns returns a new frame ------------------------------------
original = pl.DataFrame(ORDERS)
derived = with_line_totals(original)
assert derived is not original
assert original.columns == ["region", "qty", "unit_price"]
assert derived.columns == ["region", "qty", "unit_price", "line_total"]
assert derived["line_total"].to_list() == [10.0, 7.5, 8.0, 7.5, 12.0, 16.0]
# The expression API has no inplace= to pass and no item assignment to
# reach for. The TypeError names the replacement rather than leaving you to
# guess at it.
assert "inplace" not in inspect.signature(pl.DataFrame.with_columns).parameters
def assign_column():
original["line_total"] = [1.0] * 6
message = str(raises(TypeError, assign_column))
assert "does not support" in message and "DataFrame.with_columns" in message, message
assert original.columns == ["region", "qty", "unit_price"]
# Handing back a new frame is cheap because it is not a copy of the data:
# the column with_columns did not touch is the same Arrow allocation in both
# frames. _get_buffer_info is a private accessor, used here as evidence and
# not as an API to write — it reports (pointer, offset, length).
assert original["qty"]._get_buffer_info() == derived["qty"]._get_buffer_info()
# Measured, against "polars never mutates and has no in_place anywhere":
# the frame-stacking methods kept an escape hatch. hstack, vstack and
# shrink_to_fit take in_place=, and extend, insert_column and drop_in_place
# mutate the receiver outright. The in_place= form returns the same object
# instead of None, so the assignment reads exactly like the pure version
# while every other reference to that frame has changed underneath it.
stackable = pl.DataFrame(ORDERS)
combined = stackable.vstack(pl.DataFrame(ORDERS), in_place=True)
assert combined is stackable and stackable.height == 12
assert stackable.hstack([pl.Series("tag", ["x"] * 12)], in_place=True) is stackable
assert stackable.columns == ["region", "qty", "unit_price", "tag"]
growing = pl.DataFrame({"x": [1, 2]})
assert growing.extend(pl.DataFrame({"x": [3]})) is growing and growing.height == 3
assert growing.insert_column(1, pl.Series("y", [9, 9, 9])) is growing
assert growing.drop_in_place("y").to_list() == [9, 9, 9]
assert growing.columns == ["x"]
# An expression is named after the first column it mentions, so without an
# alias with_columns REPLACES that column instead of adding one. The frame
# is still new either way — the overwrite is of a column in the copy.
overwritten = original.with_columns(pl.col("qty") * 2)
assert overwritten.columns == original.columns
assert overwritten["qty"].to_list() == [2, 6, 4, 10, 8, 4]
assert original["qty"].to_list() == [1, 3, 2, 5, 4, 2]
# --- a LazyFrame does nothing until collect ------------------------------
calls: list[int] = []
def audit(value: int) -> int:
calls.append(value)
return value
plan = audited_plan(pl.LazyFrame(ORDERS), audit)
assert isinstance(plan, pl.LazyFrame)
# The plan exists as a readable string while the call list is still empty:
# explain resolves the schema and prints the query, and moves no data.
explained = plan.explain()
assert isinstance(explained, str)
assert calls == [], calls
assert "FILTER" in explained and "WITH_COLUMNS" in explained
# What explain shows is not the code that was written. Unoptimized, the
# filter sits above the projection, in source order. Optimized, it has been
# pushed underneath it.
unoptimized = plan.explain(optimized=False)
assert unoptimized.index("FILTER") < unoptimized.index("WITH_COLUMNS")
assert explained.index("WITH_COLUMNS") < explained.index("FILTER")
# Collect is the only call that runs anything, and the pushdown is not just
# a nicer printout: the Python hook is invoked on the two rows that survive
# the filter, in the order the engine reached them.
result = plan.collect()
assert calls == [5, 4], calls
assert result.shape == (2, 4)
assert result["audited"].to_list() == [5, 4]
# The same three lines eagerly run the hook over all six rows and discard
# four of the results. This is the whole argument for the lazy API stated as
# a measurement, and it is why a lazy pipeline is not just a deferred eager
# one.
eager_calls: list[int] = []
def audit_eagerly(value: int) -> int:
eager_calls.append(value)
return value
eager = (
pl.DataFrame(ORDERS)
.with_columns(
pl.col("qty")
.map_elements(audit_eagerly, return_dtype=pl.Int64)
.alias("audited")
)
.filter(pl.col("qty") >= 4)
)
assert eager_calls == [1, 3, 2, 5, 4, 2], eager_calls
assert eager.height == 2
# Deferral moves the errors too. A LazyFrame over a column that does not
# exist is built without complaint; the ColumnNotFoundError arrives when the
# plan is resolved, and explain resolves it, so explain is the cheap way to
# find a typo without running the query.
typo = pl.LazyFrame({"a": [1]}).select(pl.col("nope"))
assert isinstance(typo, pl.LazyFrame)
assert 'unable to find column "nope"' in str(
raises(pl.exceptions.ColumnNotFoundError, typo.explain)
)
raises(pl.exceptions.ColumnNotFoundError, typo.collect)
# A LazyFrame has no shape, because knowing the height means running the
# query. It will answer for its columns, but only by resolving the schema,
# and it says so with a PerformanceWarning rather than doing it quietly.
lazy = pl.LazyFrame(ORDERS)
assert not hasattr(lazy, "shape")
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
columns = lazy.columns
assert [type(w.message) for w in caught] == [pl.exceptions.PerformanceWarning]
assert columns == ["region", "qty", "unit_price"]
assert lazy.collect_schema().names() == columns
# --- null is not NaN -----------------------------------------------------
ratings = pl.DataFrame({"rating": [4.5, None, NAN, 3.0]})
series = ratings["rating"]
# Two different holes, and each predicate sees only its own. is_nan over a
# null is null, not False: the logic is three-valued, so a mask built this
# way is not a plain list of booleans.
assert series.is_null().to_list() == [False, True, False, False]
assert series.is_nan().to_list() == [False, None, True, False]
# The consequence: null_count reports a clean column while the mean is nan.
# Nulls are skipped by aggregates, a NaN is a real float and propagates, and
# count() sees the NaN as a value while pl.len() counts every row.
health = rating_health(ratings)
assert health["rows"] == 4
assert health["counted"] == 3
assert health["nulls"] == 1
assert health["nans"] == 1
assert health["mean"] != health["mean"]
# The same column without a null in it, which is the shape that survives
# review: null_count reports a clean column and the mean is nan anyway. A
# missing-data check written as null_count() == 0 passes on this.
nan_only = pl.Series("v", [1.0, NAN, 3.0])
assert nan_only.null_count() == 0
assert nan_only.mean() != nan_only.mean()
assert nan_only.is_nan().sum() == 1
# Measured, against the expectation that a NaN poisons every aggregate the
# way it poisons the mean: min and max ignore it. sum and mean propagate.
# The same column gives a usable answer from one aggregate and nan from the
# next, which is what makes this hard to spot in a report.
assert series.sum() != series.sum()
assert series.min() == 3.0 and series.max() == 4.5
# Comparison is three-valued too. A null compares to null, not to False, so
# a null row is neither kept nor rejected by a comparison — filter drops it.
# eq_missing is the two-valued version.
assert (series == 4.5).to_list() == [True, None, False, False]
assert series.eq_missing(4.5).to_list() == [True, False, False, False]
# Measured, and the opposite of the IEEE rule Python itself follows: polars
# orders NaN above every number and compares it equal to itself, so nan ==
# nan is True here and nan > 1e308 is True, where the plain Python floats
# asserted first say False to both. A "drop the outliers" filter written as
# a comparison therefore KEEPS the NaN rows and drops the nulls — exactly
# backwards from the intent.
assert NAN != NAN
assert not (NAN > 0)
assert (pl.Series([NAN]) == NAN).to_list() == [True]
assert (pl.Series([NAN]) > 1e308).to_list() == [True]
kept = ratings.filter(pl.col("rating") > 0)["rating"].to_list()
assert len(kept) == 3 and kept[1] != kept[1]
# Sorting places nulls first in both directions, which is not where a
# descending sort by score would put "missing".
assert ratings.sort("rating")["rating"].null_count() == 1
assert ratings.sort("rating")["rating"][0] is None
assert ratings.sort("rating", descending=True)["rating"][0] is None
assert ratings.sort("rating", nulls_last=True)["rating"][3] is None
# Each dropper keeps the other's holes, and each filler ignores them.
assert series.drop_nulls().len() == 3 and series.drop_nulls().is_nan().sum() == 1
assert series.drop_nans().len() == 3 and series.drop_nans().null_count() == 1
assert series.fill_null(0.0).is_nan().sum() == 1
assert series.fill_nan(0.0).null_count() == 1
# fill_nan(None) collapses the two kinds into one so a single fill covers
# both, which is what the pandas fillna(0) reflex was assuming all along.
assert normalize_ratings(ratings, 0.0)["rating"].to_list() == [4.5, 0.0, 0.0, 3.0]
# An integer column keeps its dtype through a missing value, where pandas
# upcasts to float64 and writes NaN. Aggregates skip the null on both sides
# of the division: mean is 4 / 2, not 4 / 3.
integers = pl.Series("i", [1, None, 3])
assert integers.dtype == pl.Int64
assert integers.sum() == 4 and integers.mean() == 2.0
# --- group_by: column order is fixed, row order is not -------------------
summary = revenue_by_region(df)
# Keys first, then one column per agg expression in the order given.
assert summary.columns == ["region", "revenue", "orders"]
assert summary.to_dicts() == [
{"region": "east", "revenue": 7.5, "orders": 1},
{"region": "north", "revenue": 34.0, "orders": 3},
{"region": "south", "revenue": 19.5, "orders": 2},
]
# Row order is the part to pin down yourself. maintain_order=True gives
# first-appearance order, and first appearance is not sorted order — pandas
# groupby sorts by the key by default, so a ported test comparing row lists
# fails on ordering alone. Without maintain_order the groups are built in
# parallel and the order is not specified at all, which is why only the set
# of rows is checked there.
ordered = df.group_by("region", maintain_order=True).agg(
pl.col("line_total").sum().alias("revenue")
)
assert ordered["region"].to_list() == ["north", "south", "east"]
assert sorted(ordered["region"].to_list()) == ["east", "north", "south"]
unordered = df.group_by("region").agg(pl.col("line_total").sum().alias("revenue"))
assert sorted(unordered.to_dicts(), key=lambda row: row["region"]) == [
{"region": "east", "revenue": 7.5},
{"region": "north", "revenue": 34.0},
{"region": "south", "revenue": 19.5},
]
# Two aggregates over one column both default to that column's name, and the
# collision is a DuplicateError rather than a silently dropped column.
assert df.group_by("region").agg(pl.col("line_total").sum()).columns == [
"region",
"line_total",
]
assert "has more than one occurrence" in str(
raises(
pl.exceptions.DuplicateError,
lambda: df.group_by("region").agg(
pl.col("line_total").sum(), pl.col("line_total").max()
),
)
)
# Measured while writing this: the identical mistake in a select reports as
# DuplicateError too but with different wording, so a log grep tuned to one
# message will not find the other.
assert "duplicate output name" in str(
raises(
pl.exceptions.DuplicateError,
lambda: df.select(pl.col("qty").min(), pl.col("qty").max()),
)
)
# The result is a plain DataFrame: no index to reset, the keys are columns.
assert not hasattr(summary, "index")
print("contract ok:", pl.__version__, runtime["tags"][0])