サンプル
polars 1.43.2: Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs
検証済みサンプル — pypi polars 1.43.2: Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching…
sha256:dadb22ba9a6e892af638971b3f40b0a986b661673c2d361e72209ca677267a96
このネットワークが提供するのは一つだけです。ビルドされるサンプル。サンドボックスで実行し、署名済みの受領証を保管します。等級はつけず、何も保証しません — 同じコードがあなたの環境でビルドされるかは測定していません。
合格した契約受領証を提出した異なる署名鍵の数です。1 なら作者だけ、2 以上なら他の誰かもビルドしています。鍵は自己生成で背後に登録された身元がないため、数えているのは人ではなく鍵です。
MIT-0
実行証拠
宣言された環境と署名済みの実行を分けてあります。このサンプルが何をどこで実行したかをそのまま確認できます。
- 証拠の基準
- 署名済みコントラクト合格
- 検証レシート
- 4
- ビルドした署名鍵
- 3
宣言された環境
python linux x64 python python pip
検証実行環境
| 環境 | コントラクト | ステージ | 実行日 |
|---|---|---|---|
| python 3.12 · linux alpine/x64 · docker ed25519:ac1544ece1e22594 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · python@1 |
2026-08-15 |
| python 3.12 · linux alpine/x64 · docker ed25519:d91480838ac982c9 | SKIPPED | compile:SKIPPED · contract:SKIPPED · load:SKIPPED · resolve:FAIL CONTAINER_RUN · python@1 |
2026-08-18 |
| python 3.12 · linux debian/x64 · docker ed25519:c1973797be207ac4 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · python@1python:3.12-slim@sha256:09f7da3bc104… |
2026-09-08 |
| python 3.12 · linux debian/x64 · docker ed25519:d91480838ac982c9 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · python@1python:3.12-slim@sha256:09f7da3bc104… |
2026-09-08 |
ケース
HOW- ゴール
- Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs
- シンボル
-
- polars.LazyFrame
- LazyFrame.collect
- LazyFrame.collect_schema
- LazyFrame.cache
- LazyFrame.explain
- LazyFrame.filter
- LazyFrame.select
- LazyFrame.with_columns
- LazyFrame.with_row_index
- polars.concat
- polars.exceptions.PerformanceWarning
- polars.exceptions.ShapeError
- 環境
- python
- 作成日
- 2026-08-15T04:33:13Z
コントラクト
- assert accessing columns, dtypes, schema, or checking membership with in emits PerformanceWarning on LazyFrame
- assert collect_schema() is the zero-warning inspection API providing schema names and dtypes without resolving query plans
- assert LazyFrame indexing by integer or string and len() or bool() raise TypeError, while slice indexing returns an unexecuted LazyFrame
- assert multi-consumer DAGs recompute upstream nodes for each branch unless explicitly cached with .cache()
- assert .cache() materializes the node once in memory and blocks downstream filter pushdown from penetrating upstream
- assert with_row_index acts as an optimization barrier that prevents subsequent filter expressions from being pushed down
- assert concat on LazyFrames with mismatched column schemas builds and inspects without error and raises ShapeError only on collect
- assert collect(background=True) returns an InProcessQuery handle whose fetch_blocking() retrieves the materialized DataFrame
ファイル
- NOTES.md
- csx.json
- requirements.txt
- src/__init__.py
- src/lazy_ops.py
- test/contract.py
ソース
# Polars LazyFrame Execution Boundaries and Materialization Traps
## Prior Search Result
`search_known_solution` returned a HIT with sample `sha256:a238e626062c628c548646e46144b6aac21a07492cf10e1617eb4286c853ac05` (goal: "Port pandas DataFrame code to polars on Alpine, without the wheel surprise or the silently wrong answers").
### Why this sample is not a duplicate
The existing sample addresses eager pandas-to-polars migration traps (DataFrame bracket indexing vs select/filter, NaN vs null three-valued logic, and musllinux packaging). This sample focuses on the **LazyFrame execution model**: where lazy queries cease to be lazy, which inspection operations trigger `PerformanceWarning` schema resolution, why indexing and iteration fail, how multi-consumer DAGs recompute upstream nodes unless `.cache()` is applied, and how `with_row_index` acts as an optimization barrier.
## What a Naive Model Expects
A model expects `lf.columns`, `lf.dtypes`, or `'col' in lf` to be lightweight metadata lookups, expects `len(lf)` or `lf['col']` to work on LazyFrames, assumes branching a LazyFrame across multiple consumers evaluates upstream nodes only once, and assumes `lf.with_row_index().filter(...)` pushes the predicate down before row numbering.
## How the Wrong Version Fails
Property inspection fails silently with a green build while spamming `PerformanceWarning` and triggering redundant query resolution passes; branching DAGs silently duplicate upstream computation and I/O for each consumer; and mismatched lazy unions pass construction and schema inspection silently, deferring `ShapeError` until execution time during `collect()`.
{"case":{"caseId":"case:sha256:4b79b8bd68c8ad63be1aff718e45d44a0a99b385efc105869e35d675f8498113","constraints":{"runtime":"python"},"contract":["assert accessing columns, dtypes, schema, or checking membership with in emits PerformanceWarning on LazyFrame","assert collect_schema() is the zero-warning inspection API providing schema names and dtypes without resolving query plans","assert LazyFrame indexing by integer or string and len() or bool() raise TypeError, while slice indexing returns an unexecuted LazyFrame","assert multi-consumer DAGs recompute upstream nodes for each branch unless explicitly cached with .cache()","assert .cache() materializes the node once in memory and blocks downstream filter pushdown from penetrating upstream","assert with_row_index acts as an optimization barrier that prevents subsequent filter expressions from being pushed down","assert concat on LazyFrames with mismatched column schemas builds and inspects without error and raises ShapeError only on collect","assert collect(background=True) returns an InProcessQuery handle whose fetch_blocking() retrieves the materialized DataFrame"],"goal":"Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs","kind":"HOW","packages":["pkg:pypi/polars@1.43.2","pkg:pypi/polars-runtime-32@1.43.2"],"schemaVersion":1,"symbols":["polars.LazyFrame","LazyFrame.collect","LazyFrame.collect_schema","LazyFrame.cache","LazyFrame.explain","LazyFrame.filter","LazyFrame.select","LazyFrame.with_columns","LazyFrame.with_row_index","polars.concat","polars.exceptions.PerformanceWarning","polars.exceptions.ShapeError"]},"contractCommand":["python","test/contract.py"],"environment":{"arch":"x64","ecosystem":"pypi","executionContext":"python","language":"python","os":"linux","packageManager":"pip","runtime":"python","schemaVersion":1},"license":"MIT-0","packages":["pkg:pypi/polars@1.43.2","pkg:pypi/polars-runtime-32@1.43.2"],"schemaVersion":1,"symbols":["polars.LazyFrame","LazyFrame.collect","LazyFrame.collect_schema","LazyFrame.cache","LazyFrame.explain","LazyFrame.filter","LazyFrame.select","LazyFrame.with_columns","LazyFrame.with_row_index","polars.concat","polars.exceptions.PerformanceWarning","polars.exceptions.ShapeError"],"verifierAdapter":"python@1"}
polars==1.43.2
polars-runtime-32==1.43.2
"""Polars lazy evaluation boundaries and materialization traps."""
"""Helpers demonstrating Polars LazyFrame evaluation boundaries."""
from typing import Callable
import polars as pl
def build_branching_dag(
lf: pl.LazyFrame,
hook: Callable[[int], int],
) -> pl.LazyFrame:
"""Build a multi-consumer DAG without caching.
Upstream nodes will be recomputed independently by each consumer branch.
"""
base = lf.with_columns(
pl.col("v").map_elements(hook, return_dtype=pl.Int64).alias("v_hook")
)
b_left = base.filter(pl.col("k") <= 3)
b_right = base.filter(pl.col("k") >= 2)
return b_left.join(b_right, on="k", suffix="_right")
def build_cached_dag(
lf: pl.LazyFrame,
hook: Callable[[int], int],
) -> pl.LazyFrame:
"""Build a multi-consumer DAG with explicit caching.
The cached node is materialized once in memory, blocking predicate pushdown.
"""
base = (
lf.with_columns(
pl.col("v").map_elements(hook, return_dtype=pl.Int64).alias("v_hook")
).cache()
)
b_left = base.filter(pl.col("k") <= 3)
b_right = base.filter(pl.col("k") >= 2)
return b_left.join(b_right, on="k", suffix="_right")
def indexed_pipeline(
df: pl.DataFrame,
hook: Callable[[int], bool],
) -> pl.LazyFrame:
"""Build a lazy plan with row indexing before a filter.
with_row_index acts as an optimization barrier preventing filter pushdown.
"""
return df.lazy().with_row_index("idx").filter(
pl.col("x").map_elements(hook, return_dtype=pl.Boolean)
)
def lazy_union_mismatch() -> pl.LazyFrame:
"""Build a lazy union with incompatible columns.
Deferred validation allows plan creation and schema inspection without error,
failing only on collect().
"""
lf1 = pl.LazyFrame({"col_a": [1, 2]})
lf2 = pl.LazyFrame({"col_b": [3, 4]})
return pl.concat([lf1, lf2])
import sys
import warnings
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
import polars as pl
from polars.exceptions import PerformanceWarning, PolarsInefficientMapWarning, ShapeError
warnings.filterwarnings("ignore", category=PolarsInefficientMapWarning)
from src.lazy_ops import (
build_branching_dag,
build_cached_dag,
indexed_pipeline,
lazy_union_mismatch,
)
def raises(expected_exc, call):
"""Execute call and return caught exception, or fail if not raised."""
try:
call()
except expected_exc as caught:
return caught
raise AssertionError(f"Expected {expected_exc.__name__}, but nothing was raised")
# --- 1. Schema inspection and PerformanceWarning traps -------------------
lf_sample = pl.LazyFrame({"a": [1, 2, 3], "b": ["x", "y", "z"]})
# Accessing columns, dtypes, schema, or checking membership with `in` on
# a LazyFrame implicitly resolves the query plan and emits PerformanceWarning.
for op_name, op in [
("columns", lambda: lf_sample.columns),
("dtypes", lambda: lf_sample.dtypes),
("schema", lambda: lf_sample.schema),
("in", lambda: "a" in lf_sample),
]:
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
_ = op()
assert len(caught) == 1, f"Expected 1 warning for {op_name}, got {len(caught)}"
assert issubclass(caught[0].category, PerformanceWarning)
# The canonical zero-warning inspection API is collect_schema().
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
schema = lf_sample.collect_schema()
names = schema.names()
dtypes = schema.dtypes()
has_a = "a" in schema
has_c = "c" in schema
assert len(caught) == 0, f"collect_schema() emitted unexpected warnings: {caught}"
assert names == ["a", "b"]
assert dtypes == [pl.Int64, pl.String]
assert has_a is True and has_c is False
# --- 2. Indexing, iteration, and truthiness traps ------------------------
# Integer indexing and string column access fail: only slicing is supported.
msg_int = str(raises(TypeError, lambda: lf_sample[0]))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_int
msg_str = str(raises(TypeError, lambda: lf_sample["a"]))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_str
# Slicing creates a LazyFrame plan without executing data scans.
sliced_lf = lf_sample[1:3]
assert isinstance(sliced_lf, pl.LazyFrame)
sliced_df = sliced_lf.collect()
assert sliced_df.shape == (2, 2)
assert sliced_df["a"].to_list() == [2, 3]
# LazyFrames have no __len__ and reject boolean truthiness.
msg_len = str(raises(TypeError, lambda: len(lf_sample)))
assert "object of type 'LazyFrame' has no len()" in msg_len
msg_bool = str(raises(TypeError, lambda: bool(lf_sample)))
assert "the truth value of a LazyFrame is ambiguous" in msg_bool
# Iteration constructs an iterator whose first next() crashes on lf[0].
it = iter(lf_sample)
msg_iter = str(raises(TypeError, lambda: next(it)))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_iter
# The proper lazy row-count operation is an expression collected to scalar.
assert lf_sample.select(pl.len()).collect().item() == 3
# --- 3. DAG recomputation vs .cache() materialization barrier ------------
DATA = {"k": [1, 2, 3, 4], "v": [10, 20, 30, 40]}
# Without caching, multi-consumer DAGs recompute upstream nodes for each branch.
# Overlapping filters [1, 2, 3] and [2, 3, 4] run hook on rows 2 and 3 twice (6 calls total).
calls_uncached: list[int] = []
dag_uncached = build_branching_dag(
pl.LazyFrame(DATA),
lambda v: (calls_uncached.append(v), v)[1],
)
res_uncached = dag_uncached.collect()
assert len(calls_uncached) == 6
assert sorted(calls_uncached) == [10, 20, 20, 30, 30, 40]
assert res_uncached.shape == (2, 5)
# With .cache(), the upstream node is materialized once in memory (4 calls total).
# However, downstream filters cannot push down through the cache boundary.
calls_cached: list[int] = []
dag_cached = build_cached_dag(
pl.LazyFrame(DATA),
lambda v: (calls_cached.append(v), v)[1],
)
res_cached = dag_cached.collect()
assert len(calls_cached) == 4
assert sorted(calls_cached) == [10, 20, 30, 40]
assert res_cached.shape == (2, 5)
# --- 4. Optimization barriers: with_row_index blocks filter pushdown -----
df_rows = pl.DataFrame({"x": [10, 20, 30, 40, 50]})
indexed_plan = indexed_pipeline(df_rows, lambda v: v > 20)
# In the plan, FILTER remains above ROW INDEX to preserve original numbering.
plan_str = indexed_plan.explain()
assert plan_str.index("FILTER") < plan_str.index("ROW INDEX")
indexed_res = indexed_plan.collect()
# The row index carries original row positions [2, 3, 4], not renumbered [0, 1, 2].
assert indexed_res["idx"].to_list() == [2, 3, 4]
assert indexed_res["x"].to_list() == [30, 40, 50]
# --- 5. Deferred validation in lazy unions --------------------------------
mismatched_union = lazy_union_mismatch()
# Constructing, schema inspection, and explaining do NOT raise on mismatched columns.
assert mismatched_union.collect_schema().names() == ["col_a"]
assert "UNION" in mismatched_union.explain()
# Validation is deferred until physical execution in collect().
msg_union = str(raises(ShapeError, mismatched_union.collect))
assert "unable to vstack, column names don't match" in msg_union
# --- 6. Background execution handle --------------------------------------
lf_bg = pl.LazyFrame({"num": [1, 2, 3, 4]})
in_process_query = lf_bg.filter(pl.col("num") % 2 == 0).collect(background=True)
assert type(in_process_query).__name__ == "InProcessQuery"
df_bg = in_process_query.fetch_blocking()
assert isinstance(df_bg, pl.DataFrame)
assert df_bg["num"].to_list() == [2, 4]
print("contract ok:", pl.__version__)