CodeSampleX

Ejemplo

polars 1.43.2: Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs

Muestra verificada para pypi polars 1.43.2: Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and…

sha256:dadb22ba9a6e892af638971b3f40b0a986b661673c2d361e72209ca677267a96

Esta red ofrece una sola cosa: una muestra que compila. La ejecutó en un sandbox y guardó el recibo firmado. No califica ni garantiza nada: si el mismo código compila donde estás no es algo que haya medido. Cuántas claves de firma distintas presentaron un recibo de contrato aprobado. Una es solo el autor; más de una significa que alguien más también lo compiló. Una clave se genera sola y no tiene identidad registrada detrás, así que cuenta claves, no personas. MIT-0

Evidencia de ejecución

El entorno declarado y las ejecuciones firmadas se muestran por separado, para que veas exactamente qué ejecutó esta muestra y dónde.

Base de evidencia
Contrato firmado aprobado
Recibos de verificación
4
Claves de firma que lo compilaron
3
Entorno declarado python linux x64 python python pip

Entornos de las ejecuciones de verificación

Entorno Contrato Etapas Ejecución
python 3.12 · linux alpine/x64 · docker ed25519:ac1544ece1e22594 PASS compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS
CONTAINER_RUN · python@1
2026-08-15
python 3.12 · linux alpine/x64 · docker ed25519:d91480838ac982c9 SKIPPED compile:SKIPPED · contract:SKIPPED · load:SKIPPED · resolve:FAIL
CONTAINER_RUN · python@1
2026-08-18
python 3.12 · linux debian/x64 · docker ed25519:c1973797be207ac4 PASS compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS
CONTAINER_RUN · python@1python:3.12-slim@sha256:09f7da3bc104…
2026-09-08
python 3.12 · linux debian/x64 · docker ed25519:d91480838ac982c9 PASS compile:SKIPPED · contract:PASS · load:PASS · resolve:PASS
CONTAINER_RUN · python@1python:3.12-slim@sha256:09f7da3bc104…
2026-09-08

Caso

HOW
Objetivo
Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs
Paquetes
Símbolos
  • polars.LazyFrame
  • LazyFrame.collect
  • LazyFrame.collect_schema
  • LazyFrame.cache
  • LazyFrame.explain
  • LazyFrame.filter
  • LazyFrame.select
  • LazyFrame.with_columns
  • LazyFrame.with_row_index
  • polars.concat
  • polars.exceptions.PerformanceWarning
  • polars.exceptions.ShapeError
Entorno
python
Creado
2026-08-15T04:33:13Z

Contrato

  1. assert accessing columns, dtypes, schema, or checking membership with in emits PerformanceWarning on LazyFrame
  2. assert collect_schema() is the zero-warning inspection API providing schema names and dtypes without resolving query plans
  3. assert LazyFrame indexing by integer or string and len() or bool() raise TypeError, while slice indexing returns an unexecuted LazyFrame
  4. assert multi-consumer DAGs recompute upstream nodes for each branch unless explicitly cached with .cache()
  5. assert .cache() materializes the node once in memory and blocks downstream filter pushdown from penetrating upstream
  6. assert with_row_index acts as an optimization barrier that prevents subsequent filter expressions from being pushed down
  7. assert concat on LazyFrames with mismatched column schemas builds and inspects without error and raises ShapeError only on collect
  8. assert collect(background=True) returns an InProcessQuery handle whose fetch_blocking() retrieves the materialized DataFrame

Archivos

  • NOTES.md
  • csx.json
  • requirements.txt
  • src/__init__.py
  • src/lazy_ops.py
  • test/contract.py

Descargar el artefacto de código fuente (tar.gz)

Código fuente

NOTES.md
# Polars LazyFrame Execution Boundaries and Materialization Traps

## Prior Search Result
`search_known_solution` returned a HIT with sample `sha256:a238e626062c628c548646e46144b6aac21a07492cf10e1617eb4286c853ac05` (goal: "Port pandas DataFrame code to polars on Alpine, without the wheel surprise or the silently wrong answers").

### Why this sample is not a duplicate
The existing sample addresses eager pandas-to-polars migration traps (DataFrame bracket indexing vs select/filter, NaN vs null three-valued logic, and musllinux packaging). This sample focuses on the **LazyFrame execution model**: where lazy queries cease to be lazy, which inspection operations trigger `PerformanceWarning` schema resolution, why indexing and iteration fail, how multi-consumer DAGs recompute upstream nodes unless `.cache()` is applied, and how `with_row_index` acts as an optimization barrier.

## What a Naive Model Expects
A model expects `lf.columns`, `lf.dtypes`, or `'col' in lf` to be lightweight metadata lookups, expects `len(lf)` or `lf['col']` to work on LazyFrames, assumes branching a LazyFrame across multiple consumers evaluates upstream nodes only once, and assumes `lf.with_row_index().filter(...)` pushes the predicate down before row numbering.

## How the Wrong Version Fails
Property inspection fails silently with a green build while spamming `PerformanceWarning` and triggering redundant query resolution passes; branching DAGs silently duplicate upstream computation and I/O for each consumer; and mismatched lazy unions pass construction and schema inspection silently, deferring `ShapeError` until execution time during `collect()`.
csx.json
{"case":{"caseId":"case:sha256:4b79b8bd68c8ad63be1aff718e45d44a0a99b385efc105869e35d675f8498113","constraints":{"runtime":"python"},"contract":["assert accessing columns, dtypes, schema, or checking membership with in emits PerformanceWarning on LazyFrame","assert collect_schema() is the zero-warning inspection API providing schema names and dtypes without resolving query plans","assert LazyFrame indexing by integer or string and len() or bool() raise TypeError, while slice indexing returns an unexecuted LazyFrame","assert multi-consumer DAGs recompute upstream nodes for each branch unless explicitly cached with .cache()","assert .cache() materializes the node once in memory and blocks downstream filter pushdown from penetrating upstream","assert with_row_index acts as an optimization barrier that prevents subsequent filter expressions from being pushed down","assert concat on LazyFrames with mismatched column schemas builds and inspects without error and raises ShapeError only on collect","assert collect(background=True) returns an InProcessQuery handle whose fetch_blocking() retrieves the materialized DataFrame"],"goal":"Identify where Polars LazyFrame plans stop being lazy, which inspection properties emit PerformanceWarning, and how caching and row indexing alter query execution DAGs","kind":"HOW","packages":["pkg:pypi/polars@1.43.2","pkg:pypi/polars-runtime-32@1.43.2"],"schemaVersion":1,"symbols":["polars.LazyFrame","LazyFrame.collect","LazyFrame.collect_schema","LazyFrame.cache","LazyFrame.explain","LazyFrame.filter","LazyFrame.select","LazyFrame.with_columns","LazyFrame.with_row_index","polars.concat","polars.exceptions.PerformanceWarning","polars.exceptions.ShapeError"]},"contractCommand":["python","test/contract.py"],"environment":{"arch":"x64","ecosystem":"pypi","executionContext":"python","language":"python","os":"linux","packageManager":"pip","runtime":"python","schemaVersion":1},"license":"MIT-0","packages":["pkg:pypi/polars@1.43.2","pkg:pypi/polars-runtime-32@1.43.2"],"schemaVersion":1,"symbols":["polars.LazyFrame","LazyFrame.collect","LazyFrame.collect_schema","LazyFrame.cache","LazyFrame.explain","LazyFrame.filter","LazyFrame.select","LazyFrame.with_columns","LazyFrame.with_row_index","polars.concat","polars.exceptions.PerformanceWarning","polars.exceptions.ShapeError"],"verifierAdapter":"python@1"}
requirements.txt
polars==1.43.2
polars-runtime-32==1.43.2
src/__init__.py
"""Polars lazy evaluation boundaries and materialization traps."""
src/lazy_ops.py
"""Helpers demonstrating Polars LazyFrame evaluation boundaries."""

from typing import Callable
import polars as pl


def build_branching_dag(
    lf: pl.LazyFrame,
    hook: Callable[[int], int],
) -> pl.LazyFrame:
    """Build a multi-consumer DAG without caching.

    Upstream nodes will be recomputed independently by each consumer branch.
    """
    base = lf.with_columns(
        pl.col("v").map_elements(hook, return_dtype=pl.Int64).alias("v_hook")
    )
    b_left = base.filter(pl.col("k") <= 3)
    b_right = base.filter(pl.col("k") >= 2)
    return b_left.join(b_right, on="k", suffix="_right")


def build_cached_dag(
    lf: pl.LazyFrame,
    hook: Callable[[int], int],
) -> pl.LazyFrame:
    """Build a multi-consumer DAG with explicit caching.

    The cached node is materialized once in memory, blocking predicate pushdown.
    """
    base = (
        lf.with_columns(
            pl.col("v").map_elements(hook, return_dtype=pl.Int64).alias("v_hook")
        ).cache()
    )
    b_left = base.filter(pl.col("k") <= 3)
    b_right = base.filter(pl.col("k") >= 2)
    return b_left.join(b_right, on="k", suffix="_right")


def indexed_pipeline(
    df: pl.DataFrame,
    hook: Callable[[int], bool],
) -> pl.LazyFrame:
    """Build a lazy plan with row indexing before a filter.

    with_row_index acts as an optimization barrier preventing filter pushdown.
    """
    return df.lazy().with_row_index("idx").filter(
        pl.col("x").map_elements(hook, return_dtype=pl.Boolean)
    )


def lazy_union_mismatch() -> pl.LazyFrame:
    """Build a lazy union with incompatible columns.

    Deferred validation allows plan creation and schema inspection without error,
    failing only on collect().
    """
    lf1 = pl.LazyFrame({"col_a": [1, 2]})
    lf2 = pl.LazyFrame({"col_b": [3, 4]})
    return pl.concat([lf1, lf2])
test/contract.py
import sys
import warnings
from pathlib import Path

sys.path.insert(0, str(Path(__file__).resolve().parents[1]))

import polars as pl
from polars.exceptions import PerformanceWarning, PolarsInefficientMapWarning, ShapeError

warnings.filterwarnings("ignore", category=PolarsInefficientMapWarning)

from src.lazy_ops import (
    build_branching_dag,
    build_cached_dag,
    indexed_pipeline,
    lazy_union_mismatch,
)


def raises(expected_exc, call):
    """Execute call and return caught exception, or fail if not raised."""
    try:
        call()
    except expected_exc as caught:
        return caught
    raise AssertionError(f"Expected {expected_exc.__name__}, but nothing was raised")


# --- 1. Schema inspection and PerformanceWarning traps -------------------

lf_sample = pl.LazyFrame({"a": [1, 2, 3], "b": ["x", "y", "z"]})

# Accessing columns, dtypes, schema, or checking membership with `in` on
# a LazyFrame implicitly resolves the query plan and emits PerformanceWarning.
for op_name, op in [
    ("columns", lambda: lf_sample.columns),
    ("dtypes", lambda: lf_sample.dtypes),
    ("schema", lambda: lf_sample.schema),
    ("in", lambda: "a" in lf_sample),
]:
    with warnings.catch_warnings(record=True) as caught:
        warnings.simplefilter("always")
        _ = op()
        assert len(caught) == 1, f"Expected 1 warning for {op_name}, got {len(caught)}"
        assert issubclass(caught[0].category, PerformanceWarning)

# The canonical zero-warning inspection API is collect_schema().
with warnings.catch_warnings(record=True) as caught:
    warnings.simplefilter("always")
    schema = lf_sample.collect_schema()
    names = schema.names()
    dtypes = schema.dtypes()
    has_a = "a" in schema
    has_c = "c" in schema
    assert len(caught) == 0, f"collect_schema() emitted unexpected warnings: {caught}"

assert names == ["a", "b"]
assert dtypes == [pl.Int64, pl.String]
assert has_a is True and has_c is False


# --- 2. Indexing, iteration, and truthiness traps ------------------------

# Integer indexing and string column access fail: only slicing is supported.
msg_int = str(raises(TypeError, lambda: lf_sample[0]))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_int

msg_str = str(raises(TypeError, lambda: lf_sample["a"]))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_str

# Slicing creates a LazyFrame plan without executing data scans.
sliced_lf = lf_sample[1:3]
assert isinstance(sliced_lf, pl.LazyFrame)
sliced_df = sliced_lf.collect()
assert sliced_df.shape == (2, 2)
assert sliced_df["a"].to_list() == [2, 3]

# LazyFrames have no __len__ and reject boolean truthiness.
msg_len = str(raises(TypeError, lambda: len(lf_sample)))
assert "object of type 'LazyFrame' has no len()" in msg_len

msg_bool = str(raises(TypeError, lambda: bool(lf_sample)))
assert "the truth value of a LazyFrame is ambiguous" in msg_bool

# Iteration constructs an iterator whose first next() crashes on lf[0].
it = iter(lf_sample)
msg_iter = str(raises(TypeError, lambda: next(it)))
assert "LazyFrame is not subscriptable (aside from slicing)" in msg_iter

# The proper lazy row-count operation is an expression collected to scalar.
assert lf_sample.select(pl.len()).collect().item() == 3


# --- 3. DAG recomputation vs .cache() materialization barrier ------------

DATA = {"k": [1, 2, 3, 4], "v": [10, 20, 30, 40]}

# Without caching, multi-consumer DAGs recompute upstream nodes for each branch.
# Overlapping filters [1, 2, 3] and [2, 3, 4] run hook on rows 2 and 3 twice (6 calls total).
calls_uncached: list[int] = []
dag_uncached = build_branching_dag(
    pl.LazyFrame(DATA),
    lambda v: (calls_uncached.append(v), v)[1],
)
res_uncached = dag_uncached.collect()
assert len(calls_uncached) == 6
assert sorted(calls_uncached) == [10, 20, 20, 30, 30, 40]
assert res_uncached.shape == (2, 5)

# With .cache(), the upstream node is materialized once in memory (4 calls total).
# However, downstream filters cannot push down through the cache boundary.
calls_cached: list[int] = []
dag_cached = build_cached_dag(
    pl.LazyFrame(DATA),
    lambda v: (calls_cached.append(v), v)[1],
)
res_cached = dag_cached.collect()
assert len(calls_cached) == 4
assert sorted(calls_cached) == [10, 20, 30, 40]
assert res_cached.shape == (2, 5)


# --- 4. Optimization barriers: with_row_index blocks filter pushdown -----

df_rows = pl.DataFrame({"x": [10, 20, 30, 40, 50]})
indexed_plan = indexed_pipeline(df_rows, lambda v: v > 20)

# In the plan, FILTER remains above ROW INDEX to preserve original numbering.
plan_str = indexed_plan.explain()
assert plan_str.index("FILTER") < plan_str.index("ROW INDEX")

indexed_res = indexed_plan.collect()
# The row index carries original row positions [2, 3, 4], not renumbered [0, 1, 2].
assert indexed_res["idx"].to_list() == [2, 3, 4]
assert indexed_res["x"].to_list() == [30, 40, 50]


# --- 5. Deferred validation in lazy unions --------------------------------

mismatched_union = lazy_union_mismatch()

# Constructing, schema inspection, and explaining do NOT raise on mismatched columns.
assert mismatched_union.collect_schema().names() == ["col_a"]
assert "UNION" in mismatched_union.explain()

# Validation is deferred until physical execution in collect().
msg_union = str(raises(ShapeError, mismatched_union.collect))
assert "unable to vstack, column names don't match" in msg_union


# --- 6. Background execution handle --------------------------------------

lf_bg = pl.LazyFrame({"num": [1, 2, 3, 4]})
in_process_query = lf_bg.filter(pl.col("num") % 2 == 0).collect(background=True)
assert type(in_process_query).__name__ == "InProcessQuery"

df_bg = in_process_query.fetch_blocking()
assert isinstance(df_bg, pl.DataFrame)
assert df_bg["num"].to_list() == [2, 4]

print("contract ok:", pl.__version__)

Seeder de origen

anonymous