샘플
polars 1.43.2: Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs
검증된 샘플 — pypi polars 1.43.2: Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and…
sha256:dadb22ba9a6e892af638971b3f40b0a986b661673c2d361e72209ca677267a96
이 네트워크가 제공하는 것은 하나입니다. 빌드되는 샘플. 샌드박스에서 돌리고 서명된 영수증을 보관합니다. 등급을 매기지 않고 무엇도 보증하지 않습니다 — 같은 코드가 당신 환경에서 빌드되는지는 측정한 적이 없습니다.
통과한 계약 영수증을 낸 서로 다른 서명 키의 수입니다. 하나면 작성자 혼자이고, 둘 이상이면 다른 사람도 빌드했다는 뜻입니다. 키는 스스로 만드는 것이고 뒤에 등록된 신원이 없으므로, 세는 것은 사람이 아니라 키입니다.
MIT-0
실행 증거
선언된 환경과 서명된 실행을 분리해 두었습니다. 이 샘플이 무엇을 어디서 실행했는지 그대로 볼 수 있습니다.
- 증거 기준
- 서명된 컨트랙트 통과
- 검증 영수증
- 4
- 빌드한 서명 키
- 3
선언된 환경
python linux x64 python python pip
검증 실행 환경
| 환경 | 컨트랙트 | 단계 | 실행일 |
|---|---|---|---|
| python 3.12 · linux alpine/x64 · docker ed25519:ac1544ece1e22594 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · python@1 |
2026-08-15 |
| python 3.12 · linux alpine/x64 · docker ed25519:d91480838ac982c9 | SKIPPED | compile:SKIPPED · contract:SKIPPED · load:SKIPPED · resolve:FAIL CONTAINER_RUN · python@1 |
2026-08-18 |
| python 3.12 · linux debian/x64 · docker ed25519:c1973797be207ac4 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · python@1python:3.12-slim@sha256:09f7da3bc104… |
2026-09-08 |
| python 3.12 · linux debian/x64 · docker ed25519:d91480838ac982c9 | PASS | compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS CONTAINER_RUN · python@1python:3.12-slim@sha256:09f7da3bc104… |
2026-09-08 |
케이스
HOW- 목표
- Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs
- 심벌
-
- polars.LazyFrame
- LazyFrame.collect
- LazyFrame.collect_schema
- LazyFrame.cache
- LazyFrame.explain
- LazyFrame.filter
- LazyFrame.select
- LazyFrame.with_columns
- LazyFrame.with_row_index
- polars.concat
- polars.exceptions.PerformanceWarning
- polars.exceptions.ShapeError
- 환경
- python
- 생성일
- 2026-08-15T04:33:13Z
컨트랙트
- assert accessing columns, dtypes, schema, or checking membership with in emits PerformanceWarning on LazyFrame
- assert collect_schema() is the zero-warning inspection API providing schema names and dtypes without resolving query plans
- assert LazyFrame indexing by integer or string and len() or bool() raise TypeError, while slice indexing returns an unexecuted LazyFrame
- assert multi-consumer DAGs recompute upstream nodes for each branch unless explicitly cached with .cache()
- assert .cache() materializes the node once in memory and blocks downstream filter pushdown from penetrating upstream
- assert with_row_index acts as an optimization barrier that prevents subsequent filter expressions from being pushed down
- assert concat on LazyFrames with mismatched column schemas builds and inspects without error and raises ShapeError only on collect
- assert collect(background=True) returns an InProcessQuery handle whose fetch_blocking() retrieves the materialized DataFrame
파일
- NOTES.md
- csx.json
- requirements.txt
- src/__init__.py
- src/lazy_ops.py
- test/contract.py
소스
# Polars LazyFrame Execution Boundaries and Materialization Traps
## Prior Search Result
`search_known_solution` returned a HIT with sample `sha256:a238e626062c628c548646e46144b6aac21a07492cf10e1617eb4286c853ac05` (goal: "Port pandas DataFrame code to polars on Alpine, without the wheel surprise or the silently wrong answers").
### Why this sample is not a duplicate
The existing sample addresses eager pandas-to-polars migration traps (DataFrame bracket indexing vs select/filter, NaN vs null three-valued logic, and musllinux packaging). This sample focuses on the **LazyFrame execution model**: where lazy queries cease to be lazy, which inspection operations trigger `PerformanceWarning` schema resolution, why indexing and iteration fail, how multi-consumer DAGs recompute upstream nodes unless `.cache()` is applied, and how `with_row_index` acts as an optimization barrier.
## What a Naive Model Expects
A model expects `lf.columns`, `lf.dtypes`, or `'col' in lf` to be lightweight metadata lookups, expects `len(lf)` or `lf['col']` to work on LazyFrames, assumes branching a LazyFrame across multiple consumers evaluates upstream nodes only once, and assumes `lf.with_row_index().filter(...)` pushes the predicate down before row numbering.
## How the Wrong Version Fails
Property inspection fails silently with a green build while spamming `PerformanceWarning` and triggering redundant query resolution passes; branching DAGs silently duplicate upstream computation and I/O for each consumer; and mismatched lazy unions pass construction and schema inspection silently, deferring `ShapeError` until execution time during `collect()`.
{"case":{"caseId":"case:sha256:4b79b8bd68c8ad63be1aff718e45d44a0a99b385efc105869e35d675f8498113","constraints":{"runtime":"python"},"contract":["assert accessing columns, dtypes, schema, or checking membership with in emits PerformanceWarning on LazyFrame","assert collect_schema() is the zero-warning inspection API providing schema names and dtypes without resolving query plans","assert LazyFrame indexing by integer or string and len() or bool() raise TypeError, while slice indexing returns an unexecuted LazyFrame","assert multi-consumer DAGs recompute upstream nodes for each branch unless explicitly cached with .cache()","assert .cache() materializes the node once in memory and blocks downstream filter pushdown from penetrating upstream","assert with_row_index acts as an optimization barrier that prevents subsequent filter expressions from being pushed down","assert concat on LazyFrames with mismatched column schemas builds and inspects without error and raises ShapeError only on collect","assert collect(background=True) returns an InProcessQuery handle whose fetch_blocking() retrieves the materialized DataFrame"],"goal":"Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs","kind":"HOW","packages":["pkg:pypi/polars@1.43.2","pkg:pypi/polars-runtime-32@1.43.2"],"schemaVersion":1,"symbols":["polars.LazyFrame","LazyFrame.collect","LazyFrame.collect_schema","LazyFrame.cache","LazyFrame.explain","LazyFrame.filter","LazyFrame.select","LazyFrame.with_columns","LazyFrame.with_row_index","polars.concat","polars.exceptions.PerformanceWarning","polars.exceptions.ShapeError"]},"contractCommand":["python","test/contract.py"],"environment":{"arch":"x64","ecosystem":"pypi","executionContext":"python","language":"python","os":"linux","packageManager":"pip","runtime":"python","schemaVersion":1},"license":"MIT-0","packages":["pkg:pypi/polars@1.43.2","pkg:pypi/polars-runtime-32@1.43.2"],"schemaVersion":1,"symbols":["polars.LazyFrame","LazyFrame.collect","LazyFrame.collect_schema","LazyFrame.cache","LazyFrame.explain","LazyFrame.filter","LazyFrame.select","LazyFrame.with_columns","LazyFrame.with_row_index","polars.concat","polars.exceptions.PerformanceWarning","polars.exceptions.ShapeError"],"verifierAdapter":"python@1"}
polars==1.43.2
polars-runtime-32==1.43.2
"""Polars lazy evaluation boundaries and materialization traps."""
"""Helpers demonstrating Polars LazyFrame evaluation boundaries."""
from typing import Callable
import polars as pl
def build_branching_dag(
lf: pl.LazyFrame,
hook: Callable[[int], int],
) -> pl.LazyFrame:
"""Build a multi-consumer DAG without caching.
Upstream nodes will be recomputed independently by each consumer branch.
"""
base = lf.with_columns(
pl.col("v").map_elements(hook, return_dtype=pl.Int64).alias("v_hook")
)
b_left = base.filter(pl.col("k") <= 3)
b_right = base.filter(pl.col("k") >= 2)
return b_left.join(b_right, on="k", suffix="_right")
def build_cached_dag(
lf: pl.LazyFrame,
hook: Callable[[int], int],
) -> pl.LazyFrame:
"""Build a multi-consumer DAG with explicit caching.
The cached node is materialized once in memory, blocking predicate pushdown.
"""
base = (
lf.with_columns(
pl.col("v").map_elements(hook, return_dtype=pl.Int64).alias("v_hook")
).cache()
)
b_left = base.filter(pl.col("k") <= 3)
b_right = base.filter(pl.col("k") >= 2)
return b_left.join(b_right, on="k", suffix="_right")
def indexed_pipeline(
df: pl.DataFrame,
hook: Callable[[int], bool],
) -> pl.LazyFrame:
"""Build a lazy plan with row indexing before a filter.
with_row_index acts as an optimization barrier preventing filter pushdown.
"""
return df.lazy().with_row_index("idx").filter(
pl.col("x").map_elements(hook, return_dtype=pl.Boolean)
)
def lazy_union_mismatch() -> pl.LazyFrame:
"""Build a lazy union with incompatible columns.
Deferred validation allows plan creation and schema inspection without error,
failing only on collect().
"""
lf1 = pl.LazyFrame({"col_a": [1, 2]})
lf2 = pl.LazyFrame({"col_b": [3, 4]})
return pl.concat([lf1, lf2])
import sys
import warnings
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
import polars as pl
from polars.exceptions import PerformanceWarning, PolarsInefficientMapWarning, ShapeError
warnings.filterwarnings("ignore", category=PolarsInefficientMapWarning)
from src.lazy_ops import (
build_branching_dag,
build_cached_dag,
indexed_pipeline,
lazy_union_mismatch,
)
def raises(expected_exc, call):
"""Execute call and return caught exception, or fail if not raised."""
try:
call()
except expected_exc as caught:
return caught
raise AssertionError(f"Expected {expected_exc.__name__}, but nothing was raised")
# --- 1. Schema inspection and PerformanceWarning traps -------------------
lf_sample = pl.LazyFrame({"a": [1, 2, 3], "b": ["x", "y", "z"]})
# Accessing columns, dtypes, schema, or checking membership with `in` on
# a LazyFrame implicitly resolves the query plan and emits PerformanceWarning.
for op_name, op in [
("columns", lambda: lf_sample.columns),
("dtypes", lambda: lf_sample.dtypes),
("schema", lambda: lf_sample.schema),
("in", lambda: "a" in lf_sample),
]:
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
_ = op()
assert len(caught) == 1, f"Expected 1 warning for {op_name}, got {len(caught)}"
assert issubclass(caught[0].category, PerformanceWarning)
# The canonical zero-warning inspection API is collect_schema().
with warnings.catch_warnings(record=True) as caught:
warnings.simplefilter("always")
schema = lf_sample.collect_schema()
names = schema.names()
dtypes = schema.dtypes()
has_a = "a" in schema
has_c = "c" in schema
assert len(caught) == 0, f"collect_schema() emitted unexpected warnings: {caught}"
assert names == ["a", "b"]
assert dtypes == [pl.Int64, pl.String]
assert has_a is True and has_c is False
# --- 2. Indexing, iteration, and truthiness traps ------------------------
# Integer indexing and string column access fail: only slicing is supported.
msg_int = str(raises(TypeError, lambda: lf_sample[0]))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_int
msg_str = str(raises(TypeError, lambda: lf_sample["a"]))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_str
# Slicing creates a LazyFrame plan without executing data scans.
sliced_lf = lf_sample[1:3]
assert isinstance(sliced_lf, pl.LazyFrame)
sliced_df = sliced_lf.collect()
assert sliced_df.shape == (2, 2)
assert sliced_df["a"].to_list() == [2, 3]
# LazyFrames have no __len__ and reject boolean truthiness.
msg_len = str(raises(TypeError, lambda: len(lf_sample)))
assert "object of type 'LazyFrame' has no len()" in msg_len
msg_bool = str(raises(TypeError, lambda: bool(lf_sample)))
assert "the truth value of a LazyFrame is ambiguous" in msg_bool
# Iteration constructs an iterator whose first next() crashes on lf[0].
it = iter(lf_sample)
msg_iter = str(raises(TypeError, lambda: next(it)))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_iter
# The proper lazy row-count operation is an expression collected to scalar.
assert lf_sample.select(pl.len()).collect().item() == 3
# --- 3. DAG recomputation vs .cache() materialization barrier ------------
DATA = {"k": [1, 2, 3, 4], "v": [10, 20, 30, 40]}
# Without caching, multi-consumer DAGs recompute upstream nodes for each branch.
# Overlapping filters [1, 2, 3] and [2, 3, 4] run hook on rows 2 and 3 twice (6 calls total).
calls_uncached: list[int] = []
dag_uncached = build_branching_dag(
pl.LazyFrame(DATA),
lambda v: (calls_uncached.append(v), v)[1],
)
res_uncached = dag_uncached.collect()
assert len(calls_uncached) == 6
assert sorted(calls_uncached) == [10, 20, 20, 30, 30, 40]
assert res_uncached.shape == (2, 5)
# With .cache(), the upstream node is materialized once in memory (4 calls total).
# However, downstream filters cannot push down through the cache boundary.
calls_cached: list[int] = []
dag_cached = build_cached_dag(
pl.LazyFrame(DATA),
lambda v: (calls_cached.append(v), v)[1],
)
res_cached = dag_cached.collect()
assert len(calls_cached) == 4
assert sorted(calls_cached) == [10, 20, 30, 40]
assert res_cached.shape == (2, 5)
# --- 4. Optimization barriers: with_row_index blocks filter pushdown -----
df_rows = pl.DataFrame({"x": [10, 20, 30, 40, 50]})
indexed_plan = indexed_pipeline(df_rows, lambda v: v > 20)
# In the plan, FILTER remains above ROW INDEX to preserve original numbering.
plan_str = indexed_plan.explain()
assert plan_str.index("FILTER") < plan_str.index("ROW INDEX")
indexed_res = indexed_plan.collect()
# The row index carries original row positions [2, 3, 4], not renumbered [0, 1, 2].
assert indexed_res["idx"].to_list() == [2, 3, 4]
assert indexed_res["x"].to_list() == [30, 40, 50]
# --- 5. Deferred validation in lazy unions --------------------------------
mismatched_union = lazy_union_mismatch()
# Constructing, schema inspection, and explaining do NOT raise on mismatched columns.
assert mismatched_union.collect_schema().names() == ["col_a"]
assert "UNION" in mismatched_union.explain()
# Validation is deferred until physical execution in collect().
msg_union = str(raises(ShapeError, mismatched_union.collect))
assert "unable to vstack, column names don't match" in msg_union
# --- 6. Background execution handle --------------------------------------
lf_bg = pl.LazyFrame({"num": [1, 2, 3, 4]})
in_process_query = lf_bg.filter(pl.col("num") % 2 == 0).collect(background=True)
assert type(in_process_query).__name__ == "InProcessQuery"
df_bg = in_process_query.fetch_blocking()
assert isinstance(df_bg, pl.DataFrame)
assert df_bg["num"].to_list() == [2, 4]
print("contract ok:", pl.__version__)