Skip to content

Iterators API reference

ImageProcessingIterator

ngio.iterators.ImageProcessingIterator

ImageProcessingIterator(
    input_image: Image,
    output_image: Image,
    *,
    channel_selection: ChannelSlicingInputType = None,
    output_channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol]
    | None = None,
    output_transforms: Sequence[TransformProtocol]
    | None = None,
    consolidation_mode: ConsolidationMode | None = None,
)

Bases: WritingIteratorBuilder[ndarray, Array, None]

Apply an image-to-image function region by region.

Reads each region from the input image, hands the patch to process's function, and writes the result to the output image — a filter, a projection, a restoration model. Compose with by_write_units() for collision-free parallel writes and with_halo for seamless tiles. Overlapping writes are safe; on_overlap(...) declares how contested pixels resolve.

Process input_image region by region into output_image.

Parameters:

  • input_image (Image) –

    The image to read.

  • output_image (Image) –

    The image the results are written to.

  • channel_selection (ChannelSlicingInputType, default: None ) –

    Restrict the input reads to these channels.

  • output_channel_selection (ChannelSlicingInputType, default: None ) –

    Restrict the output writes to these channels.

  • axes_order (Sequence[str] | None, default: None ) –

    Axes order of the patches handed to the function.

  • input_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Transforms applied to each input patch.

  • output_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Transforms applied to each patch before the write.

  • consolidation_mode (ConsolidationMode | None, default: None ) –

    How to build the output pyramid after iteration, see Image.consolidate.

Source code in src/ngio/iterators/_image_processing.py
def __init__(
    self,
    input_image: Image,
    output_image: Image,
    *,
    channel_selection: ChannelSlicingInputType = None,
    output_channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol] | None = None,
    output_transforms: Sequence[TransformProtocol] | None = None,
    consolidation_mode: ConsolidationMode | None = None,
) -> None:
    """Process `input_image` region by region into `output_image`.

    Args:
        input_image: The image to read.
        output_image: The image the results are written to.
        channel_selection: Restrict the input reads to these channels.
        output_channel_selection: Restrict the output writes to these
            channels.
        axes_order: Axes order of the patches handed to the function.
        input_transforms: Transforms applied to each input patch.
        output_transforms: Transforms applied to each patch before the
            write.
        consolidation_mode: How to build the output pyramid after
            iteration, see `Image.consolidate`.
    """
    self._input = input_image
    self._output = output_image
    self._ref_image = input_image
    self._rois = input_image.build_image_roi_table(name=None).rois()
    self._consolidation_mode = consolidation_mode

    self._input_slicing_kwargs = add_channel_selection_to_slicing_dict(
        image=self._input,
        channel_selection=channel_selection,
        slicing_dict={},
    )
    self._output_slicing_kwargs = add_channel_selection_to_slicing_dict(
        image=self._output,
        channel_selection=output_channel_selection,
        slicing_dict={},
    )
    self._channel_selection = channel_selection
    self._output_channel_selection = output_channel_selection
    self._axes_order = axes_order
    self._input_transforms = input_transforms
    self._output_transforms = output_transforms

    self._input.require_dimensions_match(self._output, allow_singleton=True)

rois property

rois: list[Roi]

Get the list of ROIs for the iterator.

ref_image property

ref_image: AbstractImage

Get the reference image for the iterator.

halo property

halo: dict[str, int]

The per-axis read margin, empty when there is none.

partition_indices property

partition_indices: list[int] | None

The unit indices this iterator is restricted to, or None.

None on an unrestricted iterator; on a for_job slice, the sorted positions (in the full ROI list) of the units it runs. The whole layout of a split is [it.for_job(i, n).partition_indices for i in range(n)] — the pre-flight check before submitting a distributed run: one fat list plus empties means the output chunking, not the cluster, is the constraint.

output_image property

output_image: Image

The image this iterator writes to.

with_halo

with_halo(
    x: int = 0, y: int = 0, z: int = 0, t: int = 0
) -> Self

Read a margin of context around each ROI, but write only the ROI.

The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.

Note

The margins do not survive a rescale: a shape-changing transform on the data path (a ZoomTransform) makes the crop refuse — iterate on a coarser pyramid level instead. A MaskTransform composes normally (its zoom rescales the label it reads).

Parameters:

  • x (int, default: 0 ) –

    Pixels added on each side along x.

  • y (int, default: 0 ) –

    Pixels added on each side along y.

  • z (int, default: 0 ) –

    Pixels added on each side along z.

  • t (int, default: 0 ) –

    Frames added on each side along t.

Returns:

  • Self –

    A new iterator reading with the halo.

Raises:

  • NgioValueError –

    On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps roi_index/roi_name so your coalesce can), on an in-place iterator, or for a negative margin.

Example
# 8 px of context per side, disjoint writes, still parallel.
it = iterator.by_write_units().with_halo(x=8, y=8)
it.map(smooth, mapper=ThreadedMapper("auto"))
Source code in src/ngio/iterators/_abstract_iterator.py
def with_halo(self, x: int = 0, y: int = 0, z: int = 0, t: int = 0) -> Self:
    """Read a margin of context around each ROI, but write only the ROI.

    The function receives the grown region and must return it grown too;
    the border is cropped before the write, so tiles come out seamless.
    Margins are reference-image pixels, clipped at the borders. Write
    footprints are unchanged, so a haloed iterator parallelizes exactly
    as far as it did without one.

    Note:
        The margins do not survive a rescale: a shape-changing transform
        on the data path (a `ZoomTransform`) makes the crop refuse —
        iterate on a coarser pyramid level instead. A `MaskTransform`
        composes normally (its zoom rescales the label it reads).

    Args:
        x: Pixels added on each side along x.
        y: Pixels added on each side along y.
        z: Pixels added on each side along z.
        t: Frames added on each side along t.

    Returns:
        A new iterator reading with the halo.

    Raises:
        NgioValueError: On a read-only iterator that has not opted in
            (no write to crop the halo from; detection opts in and
            reconciles the overlap with NMS, feature extraction opts in
            and stamps `roi_index`/`roi_name` so your `coalesce` can),
            on an in-place iterator, or for a negative margin.

    Example:
        ```python
        # 8 px of context per side, disjoint writes, still parallel.
        it = iterator.by_write_units().with_halo(x=8, y=8)
        it.map(smooth, mapper=ThreadedMapper("auto"))
        ```
    """
    halo = {"x": x, "y": y, "z": z, "t": t}
    for axis_name, margin in halo.items():
        if margin < 0:
            raise NgioValueError(
                f"Halo along '{axis_name}' must be >= 0, got {margin}."
            )
    if self.output_image is None and not self._allow_readonly_halo:
        name = self.__class__.__name__
        raise NgioValueError(
            f"{name} is read-only, so there is no written region for a "
            "halo to surround. Widen the ROIs themselves instead (e.g. "
            "`by_grid(...)` with a stride smaller than the size)."
        )
    output = self.output_image
    if output is not None and _is_same_zarr_array(
        self._ref_image.zarr_array, output.zarr_array
    ):
        raise NgioValueError(
            "In-place iteration cannot take a halo: each tile's grown "
            "read covers pixels that neighbouring tiles write, so the "
            "result depends on execution order. Write to a different "
            "image or label, or drop the halo."
        )
    new_instance = self._new_from_rois(self.rois)
    new_instance._halo = {ax: m for ax, m in halo.items() if m > 0}
    return new_instance

by_grid

by_grid(
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self

Tile the current ROIs with a regular grid.

Parameters:

  • size_x (int | None, default: None ) –

    Tile size along x; None (default) means the full extent.

  • size_y (int | None, default: None ) –

    Tile size along y; None means the full extent.

  • size_z (int | None, default: None ) –

    Tile size along z; None means the full extent.

  • size_t (int | None, default: None ) –

    Tile size along t; None means the full extent.

  • stride_x (int | None, default: None ) –

    Step between tiles along x; None means size_x (adjacent tiles). A smaller stride produces overlapping tiles.

  • stride_y (int | None, default: None ) –

    As stride_x, along y.

  • stride_z (int | None, default: None ) –

    As stride_x, along z.

  • stride_t (int | None, default: None ) –

    As stride_x, along t.

  • tail (TailPolicy, default: 'clip' ) –

    What happens where the grid does not divide the axis. "clip" (default) shrinks the last tile to the border; "balance" re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles, stride == size); "shift" moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution; "drop" discards the tail and leaves that border uncovered.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_grid(
    self,
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self:
    """Tile the current ROIs with a regular grid.

    Args:
        size_x: Tile size along x; `None` (default) means the full extent.
        size_y: Tile size along y; `None` means the full extent.
        size_z: Tile size along z; `None` means the full extent.
        size_t: Tile size along t; `None` means the full extent.
        stride_x: Step between tiles along x; `None` means `size_x`
            (adjacent tiles). A smaller stride produces overlapping tiles.
        stride_y: As `stride_x`, along y.
        stride_z: As `stride_x`, along z.
        stride_t: As `stride_x`, along t.
        tail: What happens where the grid does not divide the axis.
            `"clip"` (default) shrinks the last tile to the border;
            `"balance"` re-splits the last full tile and the tail into two
            near-equal tiles, so a thin overhang never yields a thin tile
            (needs adjacent tiles, `stride == size`); `"shift"` moves the
            last tile back to stay full-size — it then overlaps its
            neighbour, so a writing iterator needs a merge or serial
            execution; `"drop"` discards the tail and leaves that border
            uncovered.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the tiled ROIs.
    """
    rois = by_grid(
        rois=self.rois,
        ref_image=self.ref_image,
        size_x=size_x,
        size_y=size_y,
        size_z=size_z,
        size_t=size_t,
        stride_x=stride_x,
        stride_y=stride_y,
        stride_z=stride_z,
        stride_t=stride_t,
        tail=tail,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_blocks

by_blocks(
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self

Divide each axis into a fixed number of near-equal blocks.

The complement of by_grid: you say how many tiles, not how big. Block lengths along an axis differ by at most one pixel, so there is no tail to police — the partition is balanced by construction.

Parameters:

  • num_x (int, default: 1 ) –

    Number of blocks along x (default 1: no split).

  • num_y (int, default: 1 ) –

    Number of blocks along y.

  • num_z (int, default: 1 ) –

    Number of blocks along z.

  • num_t (int, default: 1 ) –

    Number of blocks along t.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the blocks.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_blocks(
    self,
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self:
    """Divide each axis into a fixed number of near-equal blocks.

    The complement of `by_grid`: you say how many tiles, not how big.
    Block lengths along an axis differ by at most one pixel, so there is
    no tail to police — the partition is balanced by construction.

    Args:
        num_x: Number of blocks along x (default 1: no split).
        num_y: Number of blocks along y.
        num_z: Number of blocks along z.
        num_t: Number of blocks along t.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the blocks.
    """
    rois = by_blocks(
        rois=self.rois,
        ref_image=self.ref_image,
        num_x=num_x,
        num_y=num_y,
        num_z=num_z,
        num_t=num_t,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_yx

by_yx() -> Self

Split each ROI into one 2D plane per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, z, c) combination, each covering the full y/x extent — the shape for "run this 2D function on every plane".

Source code in src/ngio/iterators/_abstract_iterator.py
def by_yx(self) -> Self:
    """Split each ROI into one 2D plane per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, z, c)` combination, each covering the full y/x extent — the
    shape for "run this 2D function on every plane".
    """
    rois = by_yx(self.rois, self.ref_image)
    return self._new_from_rois(rois)

by_zyx

by_zyx(strict: bool = True) -> Self

Split each ROI into one 3D volume per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, c) combination, each covering the full z/y/x extent.

Parameters:

  • strict (bool, default: True ) –

    Require a real z axis (present and larger than 1); with False, images without one fall back to 2D planes.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_zyx(self, strict: bool = True) -> Self:
    """Split each ROI into one 3D volume per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, c)` combination, each covering the full z/y/x extent.

    Args:
        strict: Require a real z axis (present and larger than 1);
            with `False`, images without one fall back to 2D planes.
    """
    rois = by_zyx(self.rois, self.ref_image, strict=strict)
    return self._new_from_rois(rois)

by_chunks

by_chunks(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the input image's chunk grid.

Reads are chunk-granular, so this is the natural tiling for read-heavy work. For collision-free parallel writes use by_write_units, which tiles by the output's write granularity.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_chunks(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the input image's chunk grid.

    Reads are chunk-granular, so this is the natural tiling for
    read-heavy work. For collision-free parallel *writes* use
    `by_write_units`, which tiles by the output's write granularity.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=None,
    )
    return self._new_from_rois(rois)

by_write_units

by_write_units(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the output image's write granularity.

The write unit is the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. With no overlap the resulting ROIs pass check_if_write_units_overlap by construction, so a parallel map runs as a single fully-parallel wave — a throughput optimization; conflicting tilings are still safe, just scheduled in more waves. Falls back to the input chunk grid when the iterator is read-only.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_write_units(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the output image's write granularity.

    The write unit is the shard shape when the output is sharded (writes
    are atomic per shard object), the chunk shape otherwise. With no
    overlap the resulting ROIs pass `check_if_write_units_overlap` by
    construction, so a parallel `map` runs as a single fully-parallel
    wave — a throughput optimization; conflicting tilings are still safe,
    just scheduled in more waves. Falls back to the input chunk grid when
    the iterator is read-only.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=self.output_image,
    )
    return self._new_from_rois(rois)

product

product(other: list[Roi] | GenericRoiTable) -> Self

Cartesian product of the current ROIs with an arbitrary list of ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def product(self, other: list[Roi] | GenericRoiTable) -> Self:
    """Cartesian product of the current ROIs with an arbitrary list of ROIs."""
    if isinstance(other, GenericRoiTable):
        other = other.rois()
    rois = rois_product(self.rois, other)
    return self._new_from_rois(rois)

iter

iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[
        DataGetterProtocol[NumpyPipeType],
        DataSetterProtocol[NumpyPipeType],
    ]
]
iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[
        DataGetterProtocol[DaskPipeType],
        DataSetterProtocol[DaskPipeType],
    ]
]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[NumpyPipeType, DataSetterProtocol[NumpyPipeType]]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[DaskPipeType, DataSetterProtocol[DaskPipeType]]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
    lazy: Literal[False] = ...,
    data_mode: Literal["numpy"] | None = ...,
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: int,
) -> Generator[
    tuple[
        list[NumpyPipeType],
        list[DataSetterProtocol[NumpyPipeType]],
    ]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"] | None,
    iterator_mode: Literal["readonly"],
    *,
    batch_size: int,
) -> Generator[list[NumpyPipeType]]
iter(
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal[
        "readwrite", "readonly"
    ] = "readwrite",
    *,
    batch_size: int | None = None,
) -> Generator

Create an iterator over the pixels of the ROIs.

A writing loop finalizes only when fully drained — on a stitching iterator that finalize is the resolve, so abandoning the loop mid-way leaves block-offset ids on disk (plus the transient scratch) until a re-run or a later finalize(). To write now and gather later on purpose, iterate a for_job(0, 1) slice and call finalize() on the unrestricted iterator when ready.

With batch_size set, a writing loop yields (patches, writers) — two aligned lists of up to batch_size items, in ROI order — and a read-only loop yields the payload lists alone. Stack the patches yourself (raggedness is yours to handle, unlike BatchedMapper's automatic padding), run the model once, and hand each result to its writer. Batches follow ROI order — under write_order="roi" that is the same later-ROI-wins order the mappers schedule, so the manual loop is bit-identical to map. Batching is numpy-only and eager: combining batch_size with lazy=True or data_mode="dask" raises.

Parameters:

  • lazy (bool, default: False ) –

    Yield getter/setter handles instead of materialized data.

  • data_mode (Literal['numpy', 'dask'] | None, default: None ) –

    "numpy" or "dask" (deprecated, see note).

  • iterator_mode (Literal['readwrite', 'readonly'], default: 'readwrite' ) –

    "readwrite" yields (patch, writer) pairs, "readonly" yields patches alone.

  • batch_size (int | None, default: None ) –

    When set, yield batches of up to this many items; the last batch may be smaller.

Example
for patches, writers in iterator.iter(data_mode="numpy", batch_size=8):
    outs = model(np.stack(patches))
    for writer, out in zip(writers, outs):
        writer(out)
Note

The dask data mode is deprecated and will be removed in ngio=1.2; from then on iter() yields numpy arrays.

Source code in src/ngio/iterators/_abstract_iterator.py
def iter(  # ty: ignore[invalid-method-override]
    self,
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    # "readwrite" by design — the one deliberate LSP variance of the
    # reader/writer split.
    iterator_mode: Literal["readwrite", "readonly"] = "readwrite",
    *,
    batch_size: int | None = None,
) -> Generator:
    """Create an iterator over the pixels of the ROIs.

    A writing loop finalizes only when fully drained — on a stitching
    iterator that finalize is the resolve, so abandoning the loop
    mid-way leaves block-offset ids on disk (plus the transient scratch)
    until a re-run or a later `finalize()`. To write now and gather
    later on purpose, iterate a `for_job(0, 1)` slice and call
    `finalize()` on the unrestricted iterator when ready.

    With `batch_size` set, a writing loop yields `(patches, writers)` —
    two aligned lists of up to `batch_size` items, in ROI order — and a
    read-only loop yields the payload lists alone. Stack the patches
    yourself (raggedness is yours to handle, unlike `BatchedMapper`'s
    automatic padding), run the model once, and hand each result to its
    writer. Batches follow ROI order — under `write_order="roi"` that
    is the same later-ROI-wins order the mappers schedule, so the
    manual loop is bit-identical to `map`. Batching is numpy-only
    and eager: combining `batch_size` with `lazy=True` or
    `data_mode="dask"` raises.

    Args:
        lazy: Yield getter/setter handles instead of materialized data.
        data_mode: `"numpy"` or `"dask"` (deprecated, see note).
        iterator_mode: `"readwrite"` yields `(patch, writer)` pairs,
            `"readonly"` yields patches alone.
        batch_size: When set, yield batches of up to this many items;
            the last batch may be smaller.

    Example:
        ```python
        for patches, writers in iterator.iter(data_mode="numpy", batch_size=8):
            outs = model(np.stack(patches))
            for writer, out in zip(writers, outs):
                writer(out)
        ```

    Note:
        The dask data mode is deprecated and will be removed in ngio=1.2;
        from then on `iter()` yields numpy arrays.
    """
    return self._iter_impl(
        lazy=lazy,
        data_mode=data_mode,
        iterator_mode=iterator_mode,
        batch_size=batch_size,
    )

iter_as_numpy

iter_as_numpy()

Alias for iter(data_mode="numpy").

Source code in src/ngio/iterators/_abstract_iterator.py
def iter_as_numpy(
    self,
):
    """Alias for `iter(data_mode="numpy")`."""
    return self._iter(lazy=False, data_mode="numpy", iterator_mode="readwrite")

iter_as_dask

iter_as_dask()

Alias for iter(data_mode="dask").

Deprecated: removed in ngio=1.2.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="iter_as_numpy() (or Image.get_as_dask() for a lazy array)",
    removed_in="1.2",
)
def iter_as_dask(
    self,
):
    """Alias for `iter(data_mode="dask")`.

    Deprecated: removed in ngio=1.2.
    """
    return self._iter(lazy=False, data_mode="dask", iterator_mode="readwrite")

for_job

for_job(job_index: int, n_jobs: int) -> Self

Restrict the iterator to one of n_jobs independent partitions.

No write unit is ever shared between two partitions, so the jobs need no locks, no barriers, and no channel to one another — safe as separate SLURM array tasks, in any order. Every job must build an identical iterator (construction is metadata-only and the partition derives deterministically from it), apply for_job last in the builder chain, and its verbs deliberately do not finalize — a slice's map/process/segment writes only its share, and measure/ detect bank a partial instead of joining. The gather step is the unrestricted iterator's finalize(), once, after all jobs. Stitched and read-only runs also need prepare_jobs(n_jobs) first. See the distributed-runs guide in the iterators docs.

Parameters:

  • job_index (int) –

    This job's partition, 0 <= job_index < n_jobs (e.g. int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops.

  • n_jobs (int) –

    How many partitions the run is split into; must match across every job and the gather.

Returns:

  • Self –

    The restricted iterator; further reshaping refuses.

Example
# array task:
iterator = SegmentationIterator(image, label).by_chunks()
iterator.for_job(job_index, n_jobs=n_jobs).map(func)

# gather task, after all jobs:
SegmentationIterator(image, label).by_chunks().finalize()
Source code in src/ngio/iterators/_abstract_iterator.py
def for_job(self, job_index: int, n_jobs: int) -> Self:
    """Restrict the iterator to one of `n_jobs` independent partitions.

    No write unit is ever shared between two partitions, so the jobs
    need no locks, no barriers, and no channel to one another — safe as
    separate SLURM array tasks, in any order. Every job must build an
    identical iterator (construction is metadata-only and the partition
    derives deterministically from it), apply `for_job` *last* in the
    builder chain, and its verbs deliberately do not finalize — a slice's
    `map`/`process`/`segment` writes only its share, and `measure`/
    `detect` bank a partial instead of joining. The gather step is the
    unrestricted iterator's `finalize()`, once, after all jobs.
    Stitched and read-only runs also need `prepare_jobs(n_jobs)` first.
    See the distributed-runs guide in the iterators docs.

    Args:
        job_index: This job's partition, `0 <= job_index < n_jobs`
            (e.g. `int($SLURM_ARRAY_TASK_ID)`). Partitions beyond the
            number of independent unit groups are empty no-ops.
        n_jobs: How many partitions the run is split into; must match
            across every job and the gather.

    Returns:
        The restricted iterator; further reshaping refuses.

    Example:
        ```python
        # array task:
        iterator = SegmentationIterator(image, label).by_chunks()
        iterator.for_job(job_index, n_jobs=n_jobs).map(func)

        # gather task, after all jobs:
        SegmentationIterator(image, label).by_chunks().finalize()
        ```
    """
    if self._partition is not None:
        _, current_index, current_n = self._partition
        raise NgioValueError(
            f"This iterator is already restricted to partition "
            f"{current_index}/{current_n}; partitions do not nest. "
            "Apply `for_job` once, to the full iterator."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    if not 0 <= job_index < n_jobs:
        raise NgioValueError(
            f"job_index must be in [0, {n_jobs}), got {job_index}."
        )
    self._require_splittable()
    self._validate_job_context(n_jobs)
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    restricted = self._new_from_rois(self.rois)
    restricted._partition = (
        frozenset(partition[job_index]),
        job_index,
        n_jobs,
    )
    return restricted

prepare_jobs

prepare_jobs(n_jobs: int) -> list[JobArgs]

Set up a distributed run and return its parallelization list.

The init step of the three-phase recipe (init → parallel jobs → consolidate): performs the iterator's setup — always clearing stale scratch state from earlier runs first — and returns one JobArgs per non-empty partition. Optional for plain writers; a stitching or read-only run requires it (the scratch root must exist, race-free, before any job writes into it).

Parameters:

  • n_jobs (int) –

    How many partitions to split into; empty ones are dropped from the list.

Returns:

  • list[JobArgs] –

    One {"job_index": ..., "n_jobs": ...} dict per non-empty

  • list[JobArgs] –

    partition — JSON-ready, splatting into for_job(**args).

Example
args_list = iterator.prepare_jobs(n_jobs=4)  # init task
iterator.for_job(**args).map(func)  # per parallel task
iterator.finalize()  # consolidate task
Source code in src/ngio/iterators/_abstract_iterator.py
def prepare_jobs(self, n_jobs: int) -> list[JobArgs]:
    """Set up a distributed run and return its parallelization list.

    The init step of the three-phase recipe (init → parallel jobs →
    consolidate): performs the iterator's setup — always clearing stale
    scratch state from earlier runs first — and returns one `JobArgs`
    per non-empty partition. Optional for plain writers; a stitching or
    read-only run requires it (the scratch root must exist, race-free,
    before any job writes into it).

    Args:
        n_jobs: How many partitions to split into; empty ones are
            dropped from the list.

    Returns:
        One `{"job_index": ..., "n_jobs": ...}` dict per non-empty
        partition — JSON-ready, splatting into `for_job(**args)`.

    Example:
        ```python
        args_list = iterator.prepare_jobs(n_jobs=4)  # init task
        iterator.for_job(**args).map(func)  # per parallel task
        iterator.finalize()  # consolidate task
        ```
    """
    if self._partition is not None:
        raise NgioValueError(
            "prepare_jobs belongs to the unrestricted iterator; this one "
            "is already restricted to a partition."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    self._require_splittable()
    if self.output_image is not None:
        self._validate_write_plan()
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    self._prepare_distributed(n_jobs)
    return [
        JobArgs(job_index=job_index, n_jobs=n_jobs)
        for job_index, indices in enumerate(partition)
        if indices
    ]

reduce

reduce(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Units are built read-only even on writable iterators: nothing is written and finalize does not run.

Parameters:

  • func (Callable[[NumpyPipeType], R]) –

    The function to apply; under a parallel mapper it must be safe on worker threads (or processes).

  • mapper (MapperProtocol[NumpyPipeType, R] | None, default: None ) –

    How the units are scheduled. None (the default) is serial; pass ThreadedMapper() or ProcessMapper() to fan out, sized by their own max_workers argument.

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run.

    Args:
        func: The function to apply; under a parallel mapper it must be
            safe on worker threads (or processes).
        mapper: How the units are scheduled. `None` (the default) is
            serial; pass `ThreadedMapper()` or `ProcessMapper()` to fan
            out, sized by their own `max_workers` argument.

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    if mapper is None:
        mapper = BasicMapper[NumpyPipeType, R]()
    return mapper(func, list(self._numpy_units_generator(with_setters=False)))

reduce_as_numpy

reduce_as_numpy(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Alias for reduce().

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce_as_numpy(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Alias for `reduce()`."""
    return self.reduce(func, mapper=mapper)

reduce_as_dask

reduce_as_dask(
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Deprecated: removed in ngio=1.2.

Units are built read-only even on writable iterators: nothing is written and finalize does not run. Runs serially — the units hand func lazy dask arrays, so the per-unit work is graph construction, not IO. A parallel mapper raises (see map_as_dask).

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="reduce() / reduce_as_numpy(mapper=...)",
    removed_in="1.2",
)
def reduce_as_dask(
    self,
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Deprecated: removed in ngio=1.2.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run. Runs serially — the units hand
    `func` *lazy* dask arrays, so the per-unit work is graph
    construction, not IO. A parallel `mapper` raises (see `map_as_dask`).

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    _mapper = self._require_serial_dask_mapper(mapper)
    return _mapper(func, list(self._dask_units_generator(with_setters=False)))

check_if_regions_overlap

check_if_regions_overlap() -> bool

Check if any of the ROIs overlap logically.

If two ROIs cover the same pixel, they are considered to overlap. This does not consider chunking or other storage details. Measured on the read regions — with a halo, neighbouring reads overlap by construction. For write safety, check_if_write_units_overlap is the right question.

Returns:

  • bool –

    True if any ROIs overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_regions_overlap(self) -> bool:
    """Check if any of the ROIs overlap logically.

    If two ROIs cover the same pixel, they are considered to overlap.
    This does not consider chunking or other storage details. Measured
    on the *read* regions — with a halo, neighbouring reads overlap by
    construction. For write safety, `check_if_write_units_overlap` is
    the right question.

    Returns:
        `True` if any ROIs overlap.
    """
    if len(self.rois) < 2:
        return False

    slicing_tuples = (
        g.slicing_ops.normalized_slicing_tuple
        for g in self._numpy_getters_generator()
    )
    return check_if_regions_overlap(slicing_tuples)

require_no_regions_overlap

require_no_regions_overlap() -> None

Ensure that the Iterator's ROIs do not overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_regions_overlap(self) -> None:
    """Ensure that the Iterator's ROIs do not overlap."""
    if self.check_if_regions_overlap():
        raise NgioValueError("Some rois overlap.")

map

map(
    func: Callable[[NumpyPipeType], NumpyPipeType],
    *,
    mapper: MapperProtocol[NumpyPipeType, NumpyPipeType]
    | None = None,
) -> None

Apply a transformation function to each ROI's patch and write it back.

Parameters:

  • func (Callable[[NumpyPipeType], NumpyPipeType]) –

    The transformation. Under a parallel mapper it runs on worker threads (or processes) and must be safe there.

  • mapper (MapperProtocol[NumpyPipeType, NumpyPipeType] | None, default: None ) –

    How the units are scheduled. None (the default) is serial — parallel writes stay explicit opt-in. Pass ThreadedMapper() or ProcessMapper() to fan out; each sizes its own pool from its max_workers argument, and both schedule the units into conflict-free waves (plan_waves).

Source code in src/ngio/iterators/_abstract_iterator.py
def map(
    self,
    func: Callable[[NumpyPipeType], NumpyPipeType],
    *,
    mapper: MapperProtocol[NumpyPipeType, NumpyPipeType] | None = None,
) -> None:
    """Apply a transformation function to each ROI's patch and write it back.

    Args:
        func: The transformation. Under a parallel mapper it runs on
            worker threads (or processes) and must be safe there.
        mapper: How the units are scheduled. `None` (the default) is
            serial — parallel writes stay explicit opt-in. Pass
            `ThreadedMapper()` or `ProcessMapper()` to fan out; each
            sizes its own pool from its `max_workers` argument, and both
            schedule the units into conflict-free waves (`plan_waves`).
    """
    if mapper is None:
        mapper = BasicMapper[NumpyPipeType, NumpyPipeType]()
    units = list(self._numpy_units_generator())
    self._validate_write_plan()
    mapper(func, units)
    # A partition slice runs only its share; the finalize (the one global
    # step) belongs to the unrestricted iterator, once every job is done.
    if self._partition is None:
        self.finalize()

map_as_numpy

map_as_numpy(
    func: Callable[[NumpyPipeType], NumpyPipeType],
    *,
    mapper: MapperProtocol[NumpyPipeType, NumpyPipeType]
    | None = None,
) -> None

Alias for map().

Source code in src/ngio/iterators/_abstract_iterator.py
def map_as_numpy(
    self,
    func: Callable[[NumpyPipeType], NumpyPipeType],
    *,
    mapper: MapperProtocol[NumpyPipeType, NumpyPipeType] | None = None,
) -> None:
    """Alias for `map()`."""
    return self.map(func, mapper=mapper)

map_as_dask

map_as_dask(
    func: Callable[[DaskPipeType], DaskPipeType],
    *,
    mapper: MapperProtocol[DaskPipeType, DaskPipeType]
    | None = None,
) -> None

Apply a transformation function to each ROI's patch and write it back.

Deprecated: removed in ngio=1.2.

Runs serially: the dask pipes are already executed by dask's own scheduler, and the dask write path (store_dask) scopes a global dask config option, so concurrent callers are unsafe. A parallel mapper raises.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="map() / map_as_numpy(mapper=...)",
    removed_in="1.2",
)
def map_as_dask(
    self,
    func: Callable[[DaskPipeType], DaskPipeType],
    *,
    mapper: MapperProtocol[DaskPipeType, DaskPipeType] | None = None,
) -> None:
    """Apply a transformation function to each ROI's patch and write it back.

    Deprecated: removed in ngio=1.2.

    Runs serially: the dask pipes are already executed by dask's own
    scheduler, and the dask write path (`store_dask`) scopes a global
    dask config option, so concurrent callers are unsafe. A parallel
    `mapper` raises.
    """
    _mapper = self._require_serial_dask_mapper(mapper)
    units = list(self._dask_units_generator())
    self._validate_write_plan()
    _mapper(func, units)
    if self._partition is None:
        self.finalize()

check_if_write_units_overlap

check_if_write_units_overlap() -> bool

Check if any two ROIs write into the same write unit of the output.

Measured on the write target: slicing tuples come from the setters and the grid is the output array's write granularity — the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. Two ROIs sharing a write unit make concurrent writes unsafe: the read-modify-write of that unit can lose data. The parallel mappers schedule such ROIs into separate waves, so this check answers "will my map run as a single fully-parallel wave?".

This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a loop.

Returns:

  • bool –

    True if any two ROIs share a write unit.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_write_units_overlap(self) -> bool:
    """Check if any two ROIs write into the same write unit of the output.

    Measured on the write target: slicing tuples come from the setters and
    the grid is the output array's write granularity — the shard shape when
    the output is sharded (writes are atomic per shard object), the chunk
    shape otherwise. Two ROIs sharing a write unit make concurrent writes
    unsafe: the read-modify-write of that unit can lose data. The parallel
    mappers schedule such ROIs into separate waves, so this check answers
    "will my map run as a single fully-parallel wave?".

    This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a
    loop.

    Returns:
        `True` if any two ROIs share a write unit.
    """
    if len(self.rois) < 2:
        return False

    footprints = (
        compute_write_footprint(setter)
        for setter in self._numpy_setters_generator()
        if setter is not None
    )
    non_empty = (footprint for footprint in footprints if footprint is not None)
    return any(chunk_rects_intersect(fi, fj) for fi, fj in _pairs_stream(non_empty))

require_no_write_units_overlap

require_no_write_units_overlap() -> None

Ensure that the ROIs do not share write units on the output.

The strict opt-in gate: the parallel mappers no longer refuse shared write units on their own (they wave-schedule around them), so call this to insist on a single-wave tiling instead.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_write_units_overlap(self) -> None:
    """Ensure that the ROIs do not share write units on the output.

    The strict opt-in gate: the parallel mappers no longer refuse shared
    write units on their own (they wave-schedule around them), so call
    this to insist on a single-wave tiling instead.
    """
    if self.check_if_write_units_overlap():
        raise NgioValueError("Some ROIs share write units on the output.")

on_overlap

on_overlap(
    policy: OverlapPolicy,
    *,
    write_order: WriteOrder = "any",
) -> Self

Declare how contested write pixels are resolved.

Optional here: undeclared overlapping writes are last-writer-wins in schedule order. "last" declares exactly that; any merge= policy ("max"/"min"/"sum" are order-independent; "keep_nonzero"; a merge function; a MergePolicy) combines each write with what is on disk — note it applies to every write, so pre-existing content participates too. Declare before for_job.

Parameters:

  • policy (OverlapPolicy) –

    "last", or anything the write path's merge= takes.

  • write_order (WriteOrder, default: 'any' ) –

    Who wins a contested pixel. "any" (the default) is schedule-defined — deterministic per version, fastest. "roi" makes the later ROI win — reproducible across versions and bit-identical to the manual iter loop, at up to 2.5x cost on parallel overlapping tilings. Irrelevant under an order-independent merge.

Source code in src/ngio/iterators/_image_processing.py
def on_overlap(
    self, policy: OverlapPolicy, *, write_order: WriteOrder = "any"
) -> Self:
    """Declare how contested write pixels are resolved.

    Optional here: undeclared overlapping writes are last-writer-wins
    in schedule order. `"last"` declares exactly that; any `merge=`
    policy (`"max"`/`"min"`/`"sum"` are order-independent;
    `"keep_nonzero"`; a merge function; a `MergePolicy`) combines each
    write with what is on disk — note it applies to *every* write, so
    pre-existing content participates too. Declare before `for_job`.

    Args:
        policy: `"last"`, or anything the write path's `merge=` takes.
        write_order: Who wins a contested pixel. `"any"` (the default)
            is schedule-defined — deterministic per version, fastest.
            `"roi"` makes the later ROI win — reproducible across
            versions and bit-identical to the manual `iter` loop, at up
            to 2.5x cost on parallel overlapping tilings. Irrelevant
            under an order-independent merge.
    """
    if policy != "last":
        resolve_merge(policy)  # refuse an invalid rule at declaration
    new_instance = self._new_from_rois(self.rois)
    new_instance._on_overlap = policy
    new_instance._write_order = validate_write_order(write_order)
    return new_instance

build_numpy_getter

build_numpy_getter(roi: Roi) -> DataGetterProtocol[ndarray]
Source code in src/ngio/iterators/_image_processing.py
def build_numpy_getter(self, roi: Roi) -> DataGetterProtocol[np.ndarray]:
    return NumpyGetter(
        zarr_array=self._input.zarr_array,
        dimensions=self._input.dimensions,
        roi=self._read_roi(roi),
        axes_order=self._axes_order,
        transforms=self._input_transforms,
        slicing_dict=self._input_slicing_kwargs,
    )

build_numpy_setter

build_numpy_setter(roi: Roi) -> DataSetterProtocol[ndarray]
Source code in src/ngio/iterators/_image_processing.py
def build_numpy_setter(self, roi: Roi) -> DataSetterProtocol[np.ndarray]:
    return self._wrap_setter(
        NumpySetter(
            zarr_array=self._output.zarr_array,
            dimensions=self._output.dimensions,
            roi=roi,
            axes_order=self._axes_order,
            transforms=self._output_transforms,
            slicing_dict=self._output_slicing_kwargs,
            merge=self._overlap_merge(),
        ),
        roi,
    )

build_dask_getter

build_dask_getter(roi: Roi) -> DataGetterProtocol[Array]
Source code in src/ngio/iterators/_image_processing.py
def build_dask_getter(self, roi: Roi) -> DataGetterProtocol[da.Array]:
    return DaskGetter(
        zarr_array=self._input.zarr_array,
        dimensions=self._input.dimensions,
        roi=self._read_roi(roi),
        axes_order=self._axes_order,
        transforms=self._input_transforms,
        slicing_dict=self._input_slicing_kwargs,
    )

build_dask_setter

build_dask_setter(roi: Roi) -> DataSetterProtocol[Array]
Source code in src/ngio/iterators/_image_processing.py
def build_dask_setter(self, roi: Roi) -> DataSetterProtocol[da.Array]:
    return self._wrap_setter(
        DaskSetter(
            zarr_array=self._output.zarr_array,
            dimensions=self._output.dimensions,
            roi=roi,
            axes_order=self._axes_order,
            transforms=self._output_transforms,
            slicing_dict=self._output_slicing_kwargs,
            merge=self._overlap_merge(),
        ),
        roi,
    )

process

process(
    func: Callable[[ndarray], ndarray],
    *,
    mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None

Process every region and write the results; the topic verb for map.

A serial run finalizes automatically. On a for_job slice it processes only this job's share; the gather is the unrestricted iterator's finalize(), once, after all jobs.

Parameters:

  • func (Callable[[ndarray], ndarray]) –

    The transformation. Under a parallel mapper it runs on worker threads (or processes) and must be safe there.

  • mapper (MapperProtocol[ndarray, ndarray] | None, default: None ) –

    How the units are scheduled; see map.

Source code in src/ngio/iterators/_image_processing.py
def process(
    self,
    func: Callable[[np.ndarray], np.ndarray],
    *,
    mapper: MapperProtocol[np.ndarray, np.ndarray] | None = None,
) -> None:
    """Process every region and write the results; the topic verb for `map`.

    A serial run finalizes automatically. On a `for_job` slice it
    processes only this job's share; the gather is the unrestricted
    iterator's `finalize()`, once, after all jobs.

    Args:
        func: The transformation. Under a parallel mapper it runs on
            worker threads (or processes) and must be safe there.
        mapper: How the units are scheduled; see `map`.
    """
    self.map(func, mapper=mapper)

finalize

finalize() -> None

Consolidate the output pyramid, only under the written regions.

Source code in src/ngio/iterators/_image_processing.py
def finalize(self) -> None:
    """Consolidate the output pyramid, only under the written regions."""
    self._require_unrestricted_finalize()
    self._output.consolidate(
        mode=self._consolidation_mode, regions=self._touched_write_regions()
    )

SegmentationIterator

ngio.iterators.SegmentationIterator

SegmentationIterator(
    input_image: Image,
    output_label: Label,
    *,
    channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol]
    | None = None,
    output_transforms: Sequence[TransformProtocol]
    | None = None,
    consolidation_mode: ConsolidationMode | None = None,
)

Bases: WritingIteratorBuilder[ndarray, Array, None]

Segment an image region by region into a label.

Reads each region from the input image and writes the function's label patch to the output label. With with_stitch(...) objects split across region boundaries are resolved into one id at the gather — any ROI list works: grids with a halo, overlapping FOV layouts, ragged tables. See StitchConfig. Regions whose write footprints overlap need a declared resolution — with_stitch(...) or on_overlap(...) — or the writing verbs refuse.

Segment input_image region by region into output_label.

Parameters:

  • input_image (Image) –

    The image to segment.

  • output_label (Label) –

    The label the segmentation is written to.

  • channel_selection (ChannelSlicingInputType, default: None ) –

    Restrict the image reads to these channels.

  • axes_order (Sequence[str] | None, default: None ) –

    Axes order of the patches handed to the function.

  • input_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Transforms applied to each image patch.

  • output_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Transforms applied to each predicted patch before the write.

  • consolidation_mode (ConsolidationMode | None, default: None ) –

    How to build the output pyramid after iteration, see Label.consolidate.

Source code in src/ngio/iterators/_segmentation.py
def __init__(
    self,
    input_image: Image,
    output_label: Label,
    *,
    channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol] | None = None,
    output_transforms: Sequence[TransformProtocol] | None = None,
    consolidation_mode: ConsolidationMode | None = None,
) -> None:
    """Segment `input_image` region by region into `output_label`.

    Args:
        input_image: The image to segment.
        output_label: The label the segmentation is written to.
        channel_selection: Restrict the image reads to these channels.
        axes_order: Axes order of the patches handed to the function.
        input_transforms: Transforms applied to each image patch.
        output_transforms: Transforms applied to each predicted patch
            before the write.
        consolidation_mode: How to build the output pyramid after
            iteration, see `Label.consolidate`.
    """
    self._input = input_image
    self._output = output_label
    self._ref_image = input_image
    self._rois = input_image.build_image_roi_table(name=None).rois()
    self._consolidation_mode = consolidation_mode

    self._input_slicing_kwargs = add_channel_selection_to_slicing_dict(
        image=self._input, channel_selection=channel_selection, slicing_dict={}
    )
    self._channel_selection = channel_selection
    self._axes_order = axes_order
    self._input_transforms = input_transforms
    self._output_transforms = output_transforms

    self._input.require_dimensions_match(self._output, allow_singleton=False)

rois property

rois: list[Roi]

Get the list of ROIs for the iterator.

ref_image property

ref_image: AbstractImage

Get the reference image for the iterator.

halo property

halo: dict[str, int]

The per-axis read margin, empty when there is none.

partition_indices property

partition_indices: list[int] | None

The unit indices this iterator is restricted to, or None.

None on an unrestricted iterator; on a for_job slice, the sorted positions (in the full ROI list) of the units it runs. The whole layout of a split is [it.for_job(i, n).partition_indices for i in range(n)] — the pre-flight check before submitting a distributed run: one fat list plus empties means the output chunking, not the cluster, is the constraint.

output_image property

output_image: Label

The label this iterator writes to.

with_halo

with_halo(
    x: int = 0, y: int = 0, z: int = 0, t: int = 0
) -> Self

Read a margin of context around each ROI, but write only the ROI.

The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.

Note

The margins do not survive a rescale: a shape-changing transform on the data path (a ZoomTransform) makes the crop refuse — iterate on a coarser pyramid level instead. A MaskTransform composes normally (its zoom rescales the label it reads).

Parameters:

  • x (int, default: 0 ) –

    Pixels added on each side along x.

  • y (int, default: 0 ) –

    Pixels added on each side along y.

  • z (int, default: 0 ) –

    Pixels added on each side along z.

  • t (int, default: 0 ) –

    Frames added on each side along t.

Returns:

  • Self –

    A new iterator reading with the halo.

Raises:

  • NgioValueError –

    On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps roi_index/roi_name so your coalesce can), on an in-place iterator, or for a negative margin.

Example
# 8 px of context per side, disjoint writes, still parallel.
it = iterator.by_write_units().with_halo(x=8, y=8)
it.map(smooth, mapper=ThreadedMapper("auto"))
Source code in src/ngio/iterators/_abstract_iterator.py
def with_halo(self, x: int = 0, y: int = 0, z: int = 0, t: int = 0) -> Self:
    """Read a margin of context around each ROI, but write only the ROI.

    The function receives the grown region and must return it grown too;
    the border is cropped before the write, so tiles come out seamless.
    Margins are reference-image pixels, clipped at the borders. Write
    footprints are unchanged, so a haloed iterator parallelizes exactly
    as far as it did without one.

    Note:
        The margins do not survive a rescale: a shape-changing transform
        on the data path (a `ZoomTransform`) makes the crop refuse —
        iterate on a coarser pyramid level instead. A `MaskTransform`
        composes normally (its zoom rescales the label it reads).

    Args:
        x: Pixels added on each side along x.
        y: Pixels added on each side along y.
        z: Pixels added on each side along z.
        t: Frames added on each side along t.

    Returns:
        A new iterator reading with the halo.

    Raises:
        NgioValueError: On a read-only iterator that has not opted in
            (no write to crop the halo from; detection opts in and
            reconciles the overlap with NMS, feature extraction opts in
            and stamps `roi_index`/`roi_name` so your `coalesce` can),
            on an in-place iterator, or for a negative margin.

    Example:
        ```python
        # 8 px of context per side, disjoint writes, still parallel.
        it = iterator.by_write_units().with_halo(x=8, y=8)
        it.map(smooth, mapper=ThreadedMapper("auto"))
        ```
    """
    halo = {"x": x, "y": y, "z": z, "t": t}
    for axis_name, margin in halo.items():
        if margin < 0:
            raise NgioValueError(
                f"Halo along '{axis_name}' must be >= 0, got {margin}."
            )
    if self.output_image is None and not self._allow_readonly_halo:
        name = self.__class__.__name__
        raise NgioValueError(
            f"{name} is read-only, so there is no written region for a "
            "halo to surround. Widen the ROIs themselves instead (e.g. "
            "`by_grid(...)` with a stride smaller than the size)."
        )
    output = self.output_image
    if output is not None and _is_same_zarr_array(
        self._ref_image.zarr_array, output.zarr_array
    ):
        raise NgioValueError(
            "In-place iteration cannot take a halo: each tile's grown "
            "read covers pixels that neighbouring tiles write, so the "
            "result depends on execution order. Write to a different "
            "image or label, or drop the halo."
        )
    new_instance = self._new_from_rois(self.rois)
    new_instance._halo = {ax: m for ax, m in halo.items() if m > 0}
    return new_instance

by_grid

by_grid(
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self

Tile the current ROIs with a regular grid.

Parameters:

  • size_x (int | None, default: None ) –

    Tile size along x; None (default) means the full extent.

  • size_y (int | None, default: None ) –

    Tile size along y; None means the full extent.

  • size_z (int | None, default: None ) –

    Tile size along z; None means the full extent.

  • size_t (int | None, default: None ) –

    Tile size along t; None means the full extent.

  • stride_x (int | None, default: None ) –

    Step between tiles along x; None means size_x (adjacent tiles). A smaller stride produces overlapping tiles.

  • stride_y (int | None, default: None ) –

    As stride_x, along y.

  • stride_z (int | None, default: None ) –

    As stride_x, along z.

  • stride_t (int | None, default: None ) –

    As stride_x, along t.

  • tail (TailPolicy, default: 'clip' ) –

    What happens where the grid does not divide the axis. "clip" (default) shrinks the last tile to the border; "balance" re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles, stride == size); "shift" moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution; "drop" discards the tail and leaves that border uncovered.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_grid(
    self,
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self:
    """Tile the current ROIs with a regular grid.

    Args:
        size_x: Tile size along x; `None` (default) means the full extent.
        size_y: Tile size along y; `None` means the full extent.
        size_z: Tile size along z; `None` means the full extent.
        size_t: Tile size along t; `None` means the full extent.
        stride_x: Step between tiles along x; `None` means `size_x`
            (adjacent tiles). A smaller stride produces overlapping tiles.
        stride_y: As `stride_x`, along y.
        stride_z: As `stride_x`, along z.
        stride_t: As `stride_x`, along t.
        tail: What happens where the grid does not divide the axis.
            `"clip"` (default) shrinks the last tile to the border;
            `"balance"` re-splits the last full tile and the tail into two
            near-equal tiles, so a thin overhang never yields a thin tile
            (needs adjacent tiles, `stride == size`); `"shift"` moves the
            last tile back to stay full-size — it then overlaps its
            neighbour, so a writing iterator needs a merge or serial
            execution; `"drop"` discards the tail and leaves that border
            uncovered.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the tiled ROIs.
    """
    rois = by_grid(
        rois=self.rois,
        ref_image=self.ref_image,
        size_x=size_x,
        size_y=size_y,
        size_z=size_z,
        size_t=size_t,
        stride_x=stride_x,
        stride_y=stride_y,
        stride_z=stride_z,
        stride_t=stride_t,
        tail=tail,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_blocks

by_blocks(
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self

Divide each axis into a fixed number of near-equal blocks.

The complement of by_grid: you say how many tiles, not how big. Block lengths along an axis differ by at most one pixel, so there is no tail to police — the partition is balanced by construction.

Parameters:

  • num_x (int, default: 1 ) –

    Number of blocks along x (default 1: no split).

  • num_y (int, default: 1 ) –

    Number of blocks along y.

  • num_z (int, default: 1 ) –

    Number of blocks along z.

  • num_t (int, default: 1 ) –

    Number of blocks along t.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the blocks.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_blocks(
    self,
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self:
    """Divide each axis into a fixed number of near-equal blocks.

    The complement of `by_grid`: you say how many tiles, not how big.
    Block lengths along an axis differ by at most one pixel, so there is
    no tail to police — the partition is balanced by construction.

    Args:
        num_x: Number of blocks along x (default 1: no split).
        num_y: Number of blocks along y.
        num_z: Number of blocks along z.
        num_t: Number of blocks along t.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the blocks.
    """
    rois = by_blocks(
        rois=self.rois,
        ref_image=self.ref_image,
        num_x=num_x,
        num_y=num_y,
        num_z=num_z,
        num_t=num_t,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_yx

by_yx() -> Self

Split each ROI into one 2D plane per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, z, c) combination, each covering the full y/x extent — the shape for "run this 2D function on every plane".

Source code in src/ngio/iterators/_abstract_iterator.py
def by_yx(self) -> Self:
    """Split each ROI into one 2D plane per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, z, c)` combination, each covering the full y/x extent — the
    shape for "run this 2D function on every plane".
    """
    rois = by_yx(self.rois, self.ref_image)
    return self._new_from_rois(rois)

by_zyx

by_zyx(strict: bool = True) -> Self

Split each ROI into one 3D volume per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, c) combination, each covering the full z/y/x extent.

Parameters:

  • strict (bool, default: True ) –

    Require a real z axis (present and larger than 1); with False, images without one fall back to 2D planes.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_zyx(self, strict: bool = True) -> Self:
    """Split each ROI into one 3D volume per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, c)` combination, each covering the full z/y/x extent.

    Args:
        strict: Require a real z axis (present and larger than 1);
            with `False`, images without one fall back to 2D planes.
    """
    rois = by_zyx(self.rois, self.ref_image, strict=strict)
    return self._new_from_rois(rois)

by_chunks

by_chunks(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the input image's chunk grid.

Reads are chunk-granular, so this is the natural tiling for read-heavy work. For collision-free parallel writes use by_write_units, which tiles by the output's write granularity.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_chunks(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the input image's chunk grid.

    Reads are chunk-granular, so this is the natural tiling for
    read-heavy work. For collision-free parallel *writes* use
    `by_write_units`, which tiles by the output's write granularity.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=None,
    )
    return self._new_from_rois(rois)

by_write_units

by_write_units(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the output image's write granularity.

The write unit is the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. With no overlap the resulting ROIs pass check_if_write_units_overlap by construction, so a parallel map runs as a single fully-parallel wave — a throughput optimization; conflicting tilings are still safe, just scheduled in more waves. Falls back to the input chunk grid when the iterator is read-only.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_write_units(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the output image's write granularity.

    The write unit is the shard shape when the output is sharded (writes
    are atomic per shard object), the chunk shape otherwise. With no
    overlap the resulting ROIs pass `check_if_write_units_overlap` by
    construction, so a parallel `map` runs as a single fully-parallel
    wave — a throughput optimization; conflicting tilings are still safe,
    just scheduled in more waves. Falls back to the input chunk grid when
    the iterator is read-only.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=self.output_image,
    )
    return self._new_from_rois(rois)

product

product(other: list[Roi] | GenericRoiTable) -> Self

Cartesian product of the current ROIs with an arbitrary list of ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def product(self, other: list[Roi] | GenericRoiTable) -> Self:
    """Cartesian product of the current ROIs with an arbitrary list of ROIs."""
    if isinstance(other, GenericRoiTable):
        other = other.rois()
    rois = rois_product(self.rois, other)
    return self._new_from_rois(rois)

iter

iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[
        DataGetterProtocol[NumpyPipeType],
        DataSetterProtocol[NumpyPipeType],
    ]
]
iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[
        DataGetterProtocol[DaskPipeType],
        DataSetterProtocol[DaskPipeType],
    ]
]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[NumpyPipeType, DataSetterProtocol[NumpyPipeType]]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[DaskPipeType, DataSetterProtocol[DaskPipeType]]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
    lazy: Literal[False] = ...,
    data_mode: Literal["numpy"] | None = ...,
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: int,
) -> Generator[
    tuple[
        list[NumpyPipeType],
        list[DataSetterProtocol[NumpyPipeType]],
    ]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"] | None,
    iterator_mode: Literal["readonly"],
    *,
    batch_size: int,
) -> Generator[list[NumpyPipeType]]
iter(
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal[
        "readwrite", "readonly"
    ] = "readwrite",
    *,
    batch_size: int | None = None,
) -> Generator

Create an iterator over the pixels of the ROIs.

A writing loop finalizes only when fully drained — on a stitching iterator that finalize is the resolve, so abandoning the loop mid-way leaves block-offset ids on disk (plus the transient scratch) until a re-run or a later finalize(). To write now and gather later on purpose, iterate a for_job(0, 1) slice and call finalize() on the unrestricted iterator when ready.

With batch_size set, a writing loop yields (patches, writers) — two aligned lists of up to batch_size items, in ROI order — and a read-only loop yields the payload lists alone. Stack the patches yourself (raggedness is yours to handle, unlike BatchedMapper's automatic padding), run the model once, and hand each result to its writer. Batches follow ROI order — under write_order="roi" that is the same later-ROI-wins order the mappers schedule, so the manual loop is bit-identical to map. Batching is numpy-only and eager: combining batch_size with lazy=True or data_mode="dask" raises.

Parameters:

  • lazy (bool, default: False ) –

    Yield getter/setter handles instead of materialized data.

  • data_mode (Literal['numpy', 'dask'] | None, default: None ) –

    "numpy" or "dask" (deprecated, see note).

  • iterator_mode (Literal['readwrite', 'readonly'], default: 'readwrite' ) –

    "readwrite" yields (patch, writer) pairs, "readonly" yields patches alone.

  • batch_size (int | None, default: None ) –

    When set, yield batches of up to this many items; the last batch may be smaller.

Example
for patches, writers in iterator.iter(data_mode="numpy", batch_size=8):
    outs = model(np.stack(patches))
    for writer, out in zip(writers, outs):
        writer(out)
Note

The dask data mode is deprecated and will be removed in ngio=1.2; from then on iter() yields numpy arrays.

Source code in src/ngio/iterators/_abstract_iterator.py
def iter(  # ty: ignore[invalid-method-override]
    self,
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    # "readwrite" by design — the one deliberate LSP variance of the
    # reader/writer split.
    iterator_mode: Literal["readwrite", "readonly"] = "readwrite",
    *,
    batch_size: int | None = None,
) -> Generator:
    """Create an iterator over the pixels of the ROIs.

    A writing loop finalizes only when fully drained — on a stitching
    iterator that finalize is the resolve, so abandoning the loop
    mid-way leaves block-offset ids on disk (plus the transient scratch)
    until a re-run or a later `finalize()`. To write now and gather
    later on purpose, iterate a `for_job(0, 1)` slice and call
    `finalize()` on the unrestricted iterator when ready.

    With `batch_size` set, a writing loop yields `(patches, writers)` —
    two aligned lists of up to `batch_size` items, in ROI order — and a
    read-only loop yields the payload lists alone. Stack the patches
    yourself (raggedness is yours to handle, unlike `BatchedMapper`'s
    automatic padding), run the model once, and hand each result to its
    writer. Batches follow ROI order — under `write_order="roi"` that
    is the same later-ROI-wins order the mappers schedule, so the
    manual loop is bit-identical to `map`. Batching is numpy-only
    and eager: combining `batch_size` with `lazy=True` or
    `data_mode="dask"` raises.

    Args:
        lazy: Yield getter/setter handles instead of materialized data.
        data_mode: `"numpy"` or `"dask"` (deprecated, see note).
        iterator_mode: `"readwrite"` yields `(patch, writer)` pairs,
            `"readonly"` yields patches alone.
        batch_size: When set, yield batches of up to this many items;
            the last batch may be smaller.

    Example:
        ```python
        for patches, writers in iterator.iter(data_mode="numpy", batch_size=8):
            outs = model(np.stack(patches))
            for writer, out in zip(writers, outs):
                writer(out)
        ```

    Note:
        The dask data mode is deprecated and will be removed in ngio=1.2;
        from then on `iter()` yields numpy arrays.
    """
    return self._iter_impl(
        lazy=lazy,
        data_mode=data_mode,
        iterator_mode=iterator_mode,
        batch_size=batch_size,
    )

iter_as_numpy

iter_as_numpy()

Alias for iter(data_mode="numpy").

Source code in src/ngio/iterators/_abstract_iterator.py
def iter_as_numpy(
    self,
):
    """Alias for `iter(data_mode="numpy")`."""
    return self._iter(lazy=False, data_mode="numpy", iterator_mode="readwrite")

iter_as_dask

iter_as_dask()

Alias for iter(data_mode="dask").

Deprecated: removed in ngio=1.2.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="iter_as_numpy() (or Image.get_as_dask() for a lazy array)",
    removed_in="1.2",
)
def iter_as_dask(
    self,
):
    """Alias for `iter(data_mode="dask")`.

    Deprecated: removed in ngio=1.2.
    """
    return self._iter(lazy=False, data_mode="dask", iterator_mode="readwrite")

for_job

for_job(job_index: int, n_jobs: int) -> Self

Restrict the iterator to one of n_jobs independent partitions.

No write unit is ever shared between two partitions, so the jobs need no locks, no barriers, and no channel to one another — safe as separate SLURM array tasks, in any order. Every job must build an identical iterator (construction is metadata-only and the partition derives deterministically from it), apply for_job last in the builder chain, and its verbs deliberately do not finalize — a slice's map/process/segment writes only its share, and measure/ detect bank a partial instead of joining. The gather step is the unrestricted iterator's finalize(), once, after all jobs. Stitched and read-only runs also need prepare_jobs(n_jobs) first. See the distributed-runs guide in the iterators docs.

Parameters:

  • job_index (int) –

    This job's partition, 0 <= job_index < n_jobs (e.g. int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops.

  • n_jobs (int) –

    How many partitions the run is split into; must match across every job and the gather.

Returns:

  • Self –

    The restricted iterator; further reshaping refuses.

Example
# array task:
iterator = SegmentationIterator(image, label).by_chunks()
iterator.for_job(job_index, n_jobs=n_jobs).map(func)

# gather task, after all jobs:
SegmentationIterator(image, label).by_chunks().finalize()
Source code in src/ngio/iterators/_abstract_iterator.py
def for_job(self, job_index: int, n_jobs: int) -> Self:
    """Restrict the iterator to one of `n_jobs` independent partitions.

    No write unit is ever shared between two partitions, so the jobs
    need no locks, no barriers, and no channel to one another — safe as
    separate SLURM array tasks, in any order. Every job must build an
    identical iterator (construction is metadata-only and the partition
    derives deterministically from it), apply `for_job` *last* in the
    builder chain, and its verbs deliberately do not finalize — a slice's
    `map`/`process`/`segment` writes only its share, and `measure`/
    `detect` bank a partial instead of joining. The gather step is the
    unrestricted iterator's `finalize()`, once, after all jobs.
    Stitched and read-only runs also need `prepare_jobs(n_jobs)` first.
    See the distributed-runs guide in the iterators docs.

    Args:
        job_index: This job's partition, `0 <= job_index < n_jobs`
            (e.g. `int($SLURM_ARRAY_TASK_ID)`). Partitions beyond the
            number of independent unit groups are empty no-ops.
        n_jobs: How many partitions the run is split into; must match
            across every job and the gather.

    Returns:
        The restricted iterator; further reshaping refuses.

    Example:
        ```python
        # array task:
        iterator = SegmentationIterator(image, label).by_chunks()
        iterator.for_job(job_index, n_jobs=n_jobs).map(func)

        # gather task, after all jobs:
        SegmentationIterator(image, label).by_chunks().finalize()
        ```
    """
    if self._partition is not None:
        _, current_index, current_n = self._partition
        raise NgioValueError(
            f"This iterator is already restricted to partition "
            f"{current_index}/{current_n}; partitions do not nest. "
            "Apply `for_job` once, to the full iterator."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    if not 0 <= job_index < n_jobs:
        raise NgioValueError(
            f"job_index must be in [0, {n_jobs}), got {job_index}."
        )
    self._require_splittable()
    self._validate_job_context(n_jobs)
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    restricted = self._new_from_rois(self.rois)
    restricted._partition = (
        frozenset(partition[job_index]),
        job_index,
        n_jobs,
    )
    return restricted

prepare_jobs

prepare_jobs(n_jobs: int) -> list[JobArgs]

Set up a distributed run and return its parallelization list.

The init step of the three-phase recipe (init → parallel jobs → consolidate): performs the iterator's setup — always clearing stale scratch state from earlier runs first — and returns one JobArgs per non-empty partition. Optional for plain writers; a stitching or read-only run requires it (the scratch root must exist, race-free, before any job writes into it).

Parameters:

  • n_jobs (int) –

    How many partitions to split into; empty ones are dropped from the list.

Returns:

  • list[JobArgs] –

    One {"job_index": ..., "n_jobs": ...} dict per non-empty

  • list[JobArgs] –

    partition — JSON-ready, splatting into for_job(**args).

Example
args_list = iterator.prepare_jobs(n_jobs=4)  # init task
iterator.for_job(**args).map(func)  # per parallel task
iterator.finalize()  # consolidate task
Source code in src/ngio/iterators/_abstract_iterator.py
def prepare_jobs(self, n_jobs: int) -> list[JobArgs]:
    """Set up a distributed run and return its parallelization list.

    The init step of the three-phase recipe (init → parallel jobs →
    consolidate): performs the iterator's setup — always clearing stale
    scratch state from earlier runs first — and returns one `JobArgs`
    per non-empty partition. Optional for plain writers; a stitching or
    read-only run requires it (the scratch root must exist, race-free,
    before any job writes into it).

    Args:
        n_jobs: How many partitions to split into; empty ones are
            dropped from the list.

    Returns:
        One `{"job_index": ..., "n_jobs": ...}` dict per non-empty
        partition — JSON-ready, splatting into `for_job(**args)`.

    Example:
        ```python
        args_list = iterator.prepare_jobs(n_jobs=4)  # init task
        iterator.for_job(**args).map(func)  # per parallel task
        iterator.finalize()  # consolidate task
        ```
    """
    if self._partition is not None:
        raise NgioValueError(
            "prepare_jobs belongs to the unrestricted iterator; this one "
            "is already restricted to a partition."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    self._require_splittable()
    if self.output_image is not None:
        self._validate_write_plan()
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    self._prepare_distributed(n_jobs)
    return [
        JobArgs(job_index=job_index, n_jobs=n_jobs)
        for job_index, indices in enumerate(partition)
        if indices
    ]

reduce

reduce(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Units are built read-only even on writable iterators: nothing is written and finalize does not run.

Parameters:

  • func (Callable[[NumpyPipeType], R]) –

    The function to apply; under a parallel mapper it must be safe on worker threads (or processes).

  • mapper (MapperProtocol[NumpyPipeType, R] | None, default: None ) –

    How the units are scheduled. None (the default) is serial; pass ThreadedMapper() or ProcessMapper() to fan out, sized by their own max_workers argument.

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run.

    Args:
        func: The function to apply; under a parallel mapper it must be
            safe on worker threads (or processes).
        mapper: How the units are scheduled. `None` (the default) is
            serial; pass `ThreadedMapper()` or `ProcessMapper()` to fan
            out, sized by their own `max_workers` argument.

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    if mapper is None:
        mapper = BasicMapper[NumpyPipeType, R]()
    return mapper(func, list(self._numpy_units_generator(with_setters=False)))

reduce_as_numpy

reduce_as_numpy(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Alias for reduce().

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce_as_numpy(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Alias for `reduce()`."""
    return self.reduce(func, mapper=mapper)

reduce_as_dask

reduce_as_dask(
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Deprecated: removed in ngio=1.2.

Units are built read-only even on writable iterators: nothing is written and finalize does not run. Runs serially — the units hand func lazy dask arrays, so the per-unit work is graph construction, not IO. A parallel mapper raises (see map_as_dask).

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="reduce() / reduce_as_numpy(mapper=...)",
    removed_in="1.2",
)
def reduce_as_dask(
    self,
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Deprecated: removed in ngio=1.2.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run. Runs serially — the units hand
    `func` *lazy* dask arrays, so the per-unit work is graph
    construction, not IO. A parallel `mapper` raises (see `map_as_dask`).

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    _mapper = self._require_serial_dask_mapper(mapper)
    return _mapper(func, list(self._dask_units_generator(with_setters=False)))

check_if_regions_overlap

check_if_regions_overlap() -> bool

Check if any of the ROIs overlap logically.

If two ROIs cover the same pixel, they are considered to overlap. This does not consider chunking or other storage details. Measured on the read regions — with a halo, neighbouring reads overlap by construction. For write safety, check_if_write_units_overlap is the right question.

Returns:

  • bool –

    True if any ROIs overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_regions_overlap(self) -> bool:
    """Check if any of the ROIs overlap logically.

    If two ROIs cover the same pixel, they are considered to overlap.
    This does not consider chunking or other storage details. Measured
    on the *read* regions — with a halo, neighbouring reads overlap by
    construction. For write safety, `check_if_write_units_overlap` is
    the right question.

    Returns:
        `True` if any ROIs overlap.
    """
    if len(self.rois) < 2:
        return False

    slicing_tuples = (
        g.slicing_ops.normalized_slicing_tuple
        for g in self._numpy_getters_generator()
    )
    return check_if_regions_overlap(slicing_tuples)

require_no_regions_overlap

require_no_regions_overlap() -> None

Ensure that the Iterator's ROIs do not overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_regions_overlap(self) -> None:
    """Ensure that the Iterator's ROIs do not overlap."""
    if self.check_if_regions_overlap():
        raise NgioValueError("Some rois overlap.")

map_as_numpy

map_as_numpy(
    func: Callable[[NumpyPipeType], NumpyPipeType],
    *,
    mapper: MapperProtocol[NumpyPipeType, NumpyPipeType]
    | None = None,
) -> None

Alias for map().

Source code in src/ngio/iterators/_abstract_iterator.py
def map_as_numpy(
    self,
    func: Callable[[NumpyPipeType], NumpyPipeType],
    *,
    mapper: MapperProtocol[NumpyPipeType, NumpyPipeType] | None = None,
) -> None:
    """Alias for `map()`."""
    return self.map(func, mapper=mapper)

map_as_dask

map_as_dask(
    func: Callable[[DaskPipeType], DaskPipeType],
    *,
    mapper: MapperProtocol[DaskPipeType, DaskPipeType]
    | None = None,
) -> None

Apply a transformation function to each ROI's patch and write it back.

Deprecated: removed in ngio=1.2.

Runs serially: the dask pipes are already executed by dask's own scheduler, and the dask write path (store_dask) scopes a global dask config option, so concurrent callers are unsafe. A parallel mapper raises.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="map() / map_as_numpy(mapper=...)",
    removed_in="1.2",
)
def map_as_dask(
    self,
    func: Callable[[DaskPipeType], DaskPipeType],
    *,
    mapper: MapperProtocol[DaskPipeType, DaskPipeType] | None = None,
) -> None:
    """Apply a transformation function to each ROI's patch and write it back.

    Deprecated: removed in ngio=1.2.

    Runs serially: the dask pipes are already executed by dask's own
    scheduler, and the dask write path (`store_dask`) scopes a global
    dask config option, so concurrent callers are unsafe. A parallel
    `mapper` raises.
    """
    _mapper = self._require_serial_dask_mapper(mapper)
    units = list(self._dask_units_generator())
    self._validate_write_plan()
    _mapper(func, units)
    if self._partition is None:
        self.finalize()

check_if_write_units_overlap

check_if_write_units_overlap() -> bool

Check if any two ROIs write into the same write unit of the output.

Measured on the write target: slicing tuples come from the setters and the grid is the output array's write granularity — the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. Two ROIs sharing a write unit make concurrent writes unsafe: the read-modify-write of that unit can lose data. The parallel mappers schedule such ROIs into separate waves, so this check answers "will my map run as a single fully-parallel wave?".

This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a loop.

Returns:

  • bool –

    True if any two ROIs share a write unit.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_write_units_overlap(self) -> bool:
    """Check if any two ROIs write into the same write unit of the output.

    Measured on the write target: slicing tuples come from the setters and
    the grid is the output array's write granularity — the shard shape when
    the output is sharded (writes are atomic per shard object), the chunk
    shape otherwise. Two ROIs sharing a write unit make concurrent writes
    unsafe: the read-modify-write of that unit can lose data. The parallel
    mappers schedule such ROIs into separate waves, so this check answers
    "will my map run as a single fully-parallel wave?".

    This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a
    loop.

    Returns:
        `True` if any two ROIs share a write unit.
    """
    if len(self.rois) < 2:
        return False

    footprints = (
        compute_write_footprint(setter)
        for setter in self._numpy_setters_generator()
        if setter is not None
    )
    non_empty = (footprint for footprint in footprints if footprint is not None)
    return any(chunk_rects_intersect(fi, fj) for fi, fj in _pairs_stream(non_empty))

require_no_write_units_overlap

require_no_write_units_overlap() -> None

Ensure that the ROIs do not share write units on the output.

The strict opt-in gate: the parallel mappers no longer refuse shared write units on their own (they wave-schedule around them), so call this to insist on a single-wave tiling instead.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_write_units_overlap(self) -> None:
    """Ensure that the ROIs do not share write units on the output.

    The strict opt-in gate: the parallel mappers no longer refuse shared
    write units on their own (they wave-schedule around them), so call
    this to insist on a single-wave tiling instead.
    """
    if self.check_if_write_units_overlap():
        raise NgioValueError("Some ROIs share write units on the output.")

with_stitch

with_stitch(config: StitchConfig | None = None) -> Self

Declare the stitch: split objects become one id at the gather.

Objects cut by a region boundary are resolved into one id when the run finalizes — any ROI list works: grids with a halo, overlapping FOV layouts, ragged tables. Needs a halo (with_halo) or overlapping ROIs: the evidence is overlap between neighbouring predictions. Declare before for_job; None uses the StitchConfig defaults. Refuses when on_overlap is declared — a merge policy would make the disk diverge from the banked predictions, so the resolve would relabel garbage.

Parameters:

  • config (StitchConfig | None, default: None ) –

    The stitch configuration, None for the defaults.

Source code in src/ngio/iterators/_segmentation.py
def with_stitch(self, config: StitchConfig | None = None) -> Self:
    """Declare the stitch: split objects become one id at the gather.

    Objects cut by a region boundary are resolved into one id when the
    run finalizes — any ROI list works: grids with a halo, overlapping
    FOV layouts, ragged tables. Needs a halo (`with_halo`) or
    overlapping ROIs: the evidence is overlap between neighbouring
    predictions. Declare before `for_job`; `None` uses the
    `StitchConfig` defaults. Refuses when `on_overlap` is declared —
    a merge policy would make the disk diverge from the banked
    predictions, so the resolve would relabel garbage.

    Args:
        config: The stitch configuration, `None` for the defaults.
    """
    resolved = config if config is not None else StitchConfig()
    if self._on_overlap is not None:
        raise NgioValueError(
            "with_stitch and on_overlap cannot be combined: a merge "
            "policy would make the written labels diverge from the "
            "banked predictions the resolve compares. The stitch "
            "already owns the contested pixels (deterministic write "
            "order) — drop the `on_overlap` declaration."
        )
    _require_no_unique_labels_with_stitch(resolved, self._output_transforms)
    new_instance = self._new_from_rois(self.rois)
    new_instance._stitch = resolved
    return new_instance

on_overlap

on_overlap(
    policy: OverlapPolicy,
    *,
    write_order: WriteOrder = "any",
) -> Self

Declare how contested write pixels are resolved.

"last" is last-writer-wins — an explicit acknowledgment, no merge is performed. Any merge= policy ("max"/"min"/"sum" are order-independent; "keep_nonzero"; a merge function; a MergePolicy) combines each write with what is on disk — note it applies to every write, so pre-existing content participates too. Declare before for_job. Refuses when with_stitch is declared (the stitch owns the contested pixels).

Parameters:

  • policy (OverlapPolicy) –

    "last", or anything the write path's merge= takes.

  • write_order (WriteOrder, default: 'any' ) –

    Who wins a contested pixel. "any" (the default) is schedule-defined — deterministic per version, fastest. "roi" makes the later ROI win — reproducible across versions and bit-identical to the manual iter loop, at up to 2.5x cost on parallel overlapping tilings. Irrelevant under an order-independent merge.

Source code in src/ngio/iterators/_segmentation.py
def on_overlap(
    self, policy: OverlapPolicy, *, write_order: WriteOrder = "any"
) -> Self:
    """Declare how contested write pixels are resolved.

    `"last"` is last-writer-wins — an explicit acknowledgment, no merge
    is performed. Any `merge=` policy (`"max"`/`"min"`/`"sum"` are
    order-independent; `"keep_nonzero"`; a merge function; a
    `MergePolicy`) combines each write with what is on disk — note it
    applies to *every* write, so pre-existing content participates
    too. Declare before `for_job`. Refuses when `with_stitch` is
    declared (the stitch owns the contested pixels).

    Args:
        policy: `"last"`, or anything the write path's `merge=` takes.
        write_order: Who wins a contested pixel. `"any"` (the default)
            is schedule-defined — deterministic per version, fastest.
            `"roi"` makes the later ROI win — reproducible across
            versions and bit-identical to the manual `iter` loop, at up
            to 2.5x cost on parallel overlapping tilings. Irrelevant
            under an order-independent merge.
    """
    if policy != "last":
        resolve_merge(policy)  # refuse an invalid rule at declaration
    if self._stitch is not None:
        raise NgioValueError(
            "on_overlap and with_stitch cannot be combined: the stitch "
            "already owns the contested pixels (deterministic write "
            "order), and a merge policy would make the written labels "
            "diverge from the banked predictions. Drop one of the two."
        )
    new_instance = self._new_from_rois(self.rois)
    new_instance._on_overlap = policy
    new_instance._write_order = validate_write_order(write_order)
    return new_instance

build_numpy_getter

build_numpy_getter(roi: Roi) -> DataGetterProtocol[ndarray]
Source code in src/ngio/iterators/_segmentation.py
def build_numpy_getter(self, roi: Roi) -> DataGetterProtocol[np.ndarray]:
    return NumpyGetter(
        zarr_array=self._input.zarr_array,
        dimensions=self._input.dimensions,
        roi=self._read_roi(roi),
        axes_order=self._axes_order,
        transforms=self._input_transforms,
        slicing_dict=self._input_slicing_kwargs,
    )

build_numpy_setter

build_numpy_setter(roi: Roi) -> DataSetterProtocol[ndarray]
Source code in src/ngio/iterators/_segmentation.py
def build_numpy_setter(self, roi: Roi) -> DataSetterProtocol[np.ndarray]:
    return self._wrap_for_stitch(
        self._wrap_setter(
            NumpySetter(
                zarr_array=self._output.zarr_array,
                dimensions=self._output.dimensions,
                roi=roi,
                axes_order=self._axes_order,
                transforms=self._output_transforms,
                remove_channel_selection=True,
                merge=self._overlap_merge(),
            ),
            roi,
        ),
        roi,
    )

build_dask_getter

build_dask_getter(roi: Roi) -> DataGetterProtocol[Array]
Source code in src/ngio/iterators/_segmentation.py
def build_dask_getter(self, roi: Roi) -> DataGetterProtocol[da.Array]:
    return DaskGetter(
        zarr_array=self._input.zarr_array,
        dimensions=self._input.dimensions,
        roi=self._read_roi(roi),
        axes_order=self._axes_order,
        transforms=self._input_transforms,
        slicing_dict=self._input_slicing_kwargs,
    )

build_dask_setter

build_dask_setter(roi: Roi) -> DataSetterProtocol[Array]
Source code in src/ngio/iterators/_segmentation.py
def build_dask_setter(self, roi: Roi) -> DataSetterProtocol[da.Array]:
    self._require_stitchless_dask()
    return self._wrap_setter(
        DaskSetter(
            zarr_array=self._output.zarr_array,
            dimensions=self._output.dimensions,
            roi=roi,
            axes_order=self._axes_order,
            transforms=self._output_transforms,
            remove_channel_selection=True,
            merge=self._overlap_merge(),
        ),
        roi,
    )

map

map(
    func: Callable[[ndarray], ndarray],
    *,
    mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None

See WritingIteratorBuilder.map; also cleans up on failure.

A failed standalone run cannot be resolved, so the stitch scratch arrays are deleted rather than left as a stray _ngio_stitch group beside the resolution levels. Cleanup happens only when this run created the scratch: a partition slice, a resumed run, or the gather step opened a prepared root that holds the banks every other job wrote, and one failure must not destroy them — re-running is idempotent (banks rewrite, the id offsets are derived, not counted). The already-written tiles stay in every case.

Source code in src/ngio/iterators/_segmentation.py
def map(
    self,
    func: Callable[[np.ndarray], np.ndarray],
    *,
    mapper: MapperProtocol[np.ndarray, np.ndarray] | None = None,
) -> None:
    """See `WritingIteratorBuilder.map`; also cleans up on failure.

    A failed standalone run cannot be resolved, so the stitch scratch
    arrays are deleted rather than left as a stray `_ngio_stitch` group
    beside the resolution levels. Cleanup happens only when this run
    *created* the scratch: a partition slice, a resumed run, or the
    gather step opened a prepared root that holds the banks every other
    job wrote, and one failure must not destroy them — re-running is
    idempotent (banks rewrite, the id offsets are derived, not counted).
    The already-written tiles stay in every case.
    """
    if self._stitch is None:
        return super().map(func, mapper=mapper)
    # Build the plan (and let its validation warnings and errors fire)
    # before any tile runs: lazily it would surface mid-run, from a
    # worker, after some tiles have already written.
    plan = self._stitching_plan()
    # Resolve the scratch in the parent, before any unit is pickled or
    # run — never as a side effect of pickling.
    _ = plan.banks
    if self._partition is not None:
        return super().map(func, mapper=mapper)
    try:
        return super().map(func, mapper=mapper)
    except BaseException:
        if self._stitch_plan is not None and self._stitch_plan.created_banks:
            self._stitch_plan.cleanup()
            self._stitch_plan = None
        raise

segment

segment(
    func: Callable[[ndarray], ndarray],
    *,
    mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None

Segment every region and write the labels; the topic verb for map.

A serial run finalizes automatically (the stitch resolve included). On a for_job slice it segments only this job's share; the gather is the unrestricted iterator's finalize(), once, after all jobs.

Parameters:

  • func (Callable[[ndarray], ndarray]) –

    The segmentation model. Under a parallel mapper it runs on worker threads (or processes) and must be safe there.

  • mapper (MapperProtocol[ndarray, ndarray] | None, default: None ) –

    How the units are scheduled; see map.

Source code in src/ngio/iterators/_segmentation.py
def segment(
    self,
    func: Callable[[np.ndarray], np.ndarray],
    *,
    mapper: MapperProtocol[np.ndarray, np.ndarray] | None = None,
) -> None:
    """Segment every region and write the labels; the topic verb for `map`.

    A serial run finalizes automatically (the stitch resolve included).
    On a `for_job` slice it segments only this job's share; the gather is
    the unrestricted iterator's `finalize()`, once, after all jobs.

    Args:
        func: The segmentation model. Under a parallel mapper it runs on
            worker threads (or processes) and must be safe there.
        mapper: How the units are scheduled; see `map`.
    """
    self.map(func, mapper=mapper)

finalize

finalize() -> None

Resolve the stitch (when configured), then consolidate the pyramid.

Source code in src/ngio/iterators/_segmentation.py
def finalize(self) -> None:
    """Resolve the stitch (when configured), then consolidate the pyramid."""
    self._require_unrestricted_finalize()
    # The relabel has to precede consolidation: every pyramid level is
    # derived from level 0, so stitching after would leave them disagreeing.
    if self._stitch is not None:
        # The stitch resolve relabels level 0 wherever the union-find
        # reached, not just under the written ROIs — only a full rebuild
        # is guaranteed consistent after it.
        self._stitching_plan().resolve()
        self._stitch_plan = None
        self._output.consolidate(mode=self._consolidation_mode)
    else:
        self._output.consolidate(
            mode=self._consolidation_mode, regions=self._touched_write_regions()
        )

MaskedSegmentationIterator

ngio.iterators.MaskedSegmentationIterator

MaskedSegmentationIterator(
    input_image: MaskedImage,
    output_label: Label,
    *,
    channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol]
    | None = None,
    output_transforms: Sequence[TransformProtocol]
    | None = None,
    consolidation_mode: ConsolidationMode | None = None,
)

Bases: SegmentationIterator

Segment each object of a masking ROI table, inside its own mask.

Regions come from the masking table's per-object bounding boxes; reads are masked to the object (outside pixels filled) and writes protect everything outside it (MaskMerge) — overlapping bounding boxes are safe, because each write only touches its own object's pixels. With with_stitch(...) and a tiling (by_grid + with_halo), sub-objects split by a tile boundary within one mask merge into one id; tiles of different masks are never compared (an object cannot span two masks), and ids come out unique and dense across every object — no UniqueLabelsTransform needed (combining it with stitch raises). Mind block_size: the id ceiling scales with objects times tiles per object.

Segment each masked object's box from input_image into output_label.

The ROIs come from input_image's masking ROI table: one per object, pixels outside the object masked on read and protected on write.

Parameters:

  • input_image (MaskedImage) –

    The masked image to segment; its masking table supplies the per-object ROIs.

  • output_label (Label) –

    The label the segmentation is written to.

  • channel_selection (ChannelSlicingInputType, default: None ) –

    Restrict the image reads to these channels.

  • axes_order (Sequence[str] | None, default: None ) –

    Axes order of the patches handed to the function.

  • input_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Transforms applied to each image patch, before the mask fill.

  • output_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Transforms applied to each predicted patch before the write.

  • consolidation_mode (ConsolidationMode | None, default: None ) –

    How to build the output pyramid after iteration, see Label.consolidate.

Source code in src/ngio/iterators/_segmentation.py
def __init__(
    self,
    input_image: MaskedImage,
    output_label: Label,
    *,
    channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol] | None = None,
    output_transforms: Sequence[TransformProtocol] | None = None,
    consolidation_mode: ConsolidationMode | None = None,
) -> None:
    """Segment each masked object's box from `input_image` into `output_label`.

    The ROIs come from `input_image`'s masking ROI table: one per object,
    pixels outside the object masked on read and protected on write.

    Args:
        input_image: The masked image to segment; its masking table
            supplies the per-object ROIs.
        output_label: The label the segmentation is written to.
        channel_selection: Restrict the image reads to these channels.
        axes_order: Axes order of the patches handed to the function.
        input_transforms: Transforms applied to each image patch, before
            the mask fill.
        output_transforms: Transforms applied to each predicted patch
            before the write.
        consolidation_mode: How to build the output pyramid after
            iteration, see `Label.consolidate`.
    """
    self._input = input_image
    self._output = output_label

    self._ref_image = input_image
    self._set_rois(input_image._masking_roi_table.rois())
    self._consolidation_mode = consolidation_mode

    self._input_slicing_kwargs = add_channel_selection_to_slicing_dict(
        image=self._input, channel_selection=channel_selection, slicing_dict={}
    )
    self._channel_selection = channel_selection
    self._axes_order = axes_order
    self._input_transforms = input_transforms
    self._output_transforms = output_transforms

    self._input.require_dimensions_match(self._output, allow_singleton=False)

rois property

rois: list[Roi]

Get the list of ROIs for the iterator.

ref_image property

ref_image: AbstractImage

Get the reference image for the iterator.

output_image property

output_image: Label

The label this iterator writes to.

halo property

halo: dict[str, int]

The per-axis read margin, empty when there is none.

partition_indices property

partition_indices: list[int] | None

The unit indices this iterator is restricted to, or None.

None on an unrestricted iterator; on a for_job slice, the sorted positions (in the full ROI list) of the units it runs. The whole layout of a split is [it.for_job(i, n).partition_indices for i in range(n)] — the pre-flight check before submitting a distributed run: one fat list plus empties means the output chunking, not the cluster, is the constraint.

with_halo

with_halo(
    x: int = 0, y: int = 0, z: int = 0, t: int = 0
) -> Self

Read a margin of context around each ROI, but write only the ROI.

The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.

Note

The margins do not survive a rescale: a shape-changing transform on the data path (a ZoomTransform) makes the crop refuse — iterate on a coarser pyramid level instead. A MaskTransform composes normally (its zoom rescales the label it reads).

Parameters:

  • x (int, default: 0 ) –

    Pixels added on each side along x.

  • y (int, default: 0 ) –

    Pixels added on each side along y.

  • z (int, default: 0 ) –

    Pixels added on each side along z.

  • t (int, default: 0 ) –

    Frames added on each side along t.

Returns:

  • Self –

    A new iterator reading with the halo.

Raises:

  • NgioValueError –

    On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps roi_index/roi_name so your coalesce can), on an in-place iterator, or for a negative margin.

Example
# 8 px of context per side, disjoint writes, still parallel.
it = iterator.by_write_units().with_halo(x=8, y=8)
it.map(smooth, mapper=ThreadedMapper("auto"))
Source code in src/ngio/iterators/_abstract_iterator.py
def with_halo(self, x: int = 0, y: int = 0, z: int = 0, t: int = 0) -> Self:
    """Read a margin of context around each ROI, but write only the ROI.

    The function receives the grown region and must return it grown too;
    the border is cropped before the write, so tiles come out seamless.
    Margins are reference-image pixels, clipped at the borders. Write
    footprints are unchanged, so a haloed iterator parallelizes exactly
    as far as it did without one.

    Note:
        The margins do not survive a rescale: a shape-changing transform
        on the data path (a `ZoomTransform`) makes the crop refuse —
        iterate on a coarser pyramid level instead. A `MaskTransform`
        composes normally (its zoom rescales the label it reads).

    Args:
        x: Pixels added on each side along x.
        y: Pixels added on each side along y.
        z: Pixels added on each side along z.
        t: Frames added on each side along t.

    Returns:
        A new iterator reading with the halo.

    Raises:
        NgioValueError: On a read-only iterator that has not opted in
            (no write to crop the halo from; detection opts in and
            reconciles the overlap with NMS, feature extraction opts in
            and stamps `roi_index`/`roi_name` so your `coalesce` can),
            on an in-place iterator, or for a negative margin.

    Example:
        ```python
        # 8 px of context per side, disjoint writes, still parallel.
        it = iterator.by_write_units().with_halo(x=8, y=8)
        it.map(smooth, mapper=ThreadedMapper("auto"))
        ```
    """
    halo = {"x": x, "y": y, "z": z, "t": t}
    for axis_name, margin in halo.items():
        if margin < 0:
            raise NgioValueError(
                f"Halo along '{axis_name}' must be >= 0, got {margin}."
            )
    if self.output_image is None and not self._allow_readonly_halo:
        name = self.__class__.__name__
        raise NgioValueError(
            f"{name} is read-only, so there is no written region for a "
            "halo to surround. Widen the ROIs themselves instead (e.g. "
            "`by_grid(...)` with a stride smaller than the size)."
        )
    output = self.output_image
    if output is not None and _is_same_zarr_array(
        self._ref_image.zarr_array, output.zarr_array
    ):
        raise NgioValueError(
            "In-place iteration cannot take a halo: each tile's grown "
            "read covers pixels that neighbouring tiles write, so the "
            "result depends on execution order. Write to a different "
            "image or label, or drop the halo."
        )
    new_instance = self._new_from_rois(self.rois)
    new_instance._halo = {ax: m for ax, m in halo.items() if m > 0}
    return new_instance

by_grid

by_grid(
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self

Tile the current ROIs with a regular grid.

Parameters:

  • size_x (int | None, default: None ) –

    Tile size along x; None (default) means the full extent.

  • size_y (int | None, default: None ) –

    Tile size along y; None means the full extent.

  • size_z (int | None, default: None ) –

    Tile size along z; None means the full extent.

  • size_t (int | None, default: None ) –

    Tile size along t; None means the full extent.

  • stride_x (int | None, default: None ) –

    Step between tiles along x; None means size_x (adjacent tiles). A smaller stride produces overlapping tiles.

  • stride_y (int | None, default: None ) –

    As stride_x, along y.

  • stride_z (int | None, default: None ) –

    As stride_x, along z.

  • stride_t (int | None, default: None ) –

    As stride_x, along t.

  • tail (TailPolicy, default: 'clip' ) –

    What happens where the grid does not divide the axis. "clip" (default) shrinks the last tile to the border; "balance" re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles, stride == size); "shift" moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution; "drop" discards the tail and leaves that border uncovered.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_grid(
    self,
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self:
    """Tile the current ROIs with a regular grid.

    Args:
        size_x: Tile size along x; `None` (default) means the full extent.
        size_y: Tile size along y; `None` means the full extent.
        size_z: Tile size along z; `None` means the full extent.
        size_t: Tile size along t; `None` means the full extent.
        stride_x: Step between tiles along x; `None` means `size_x`
            (adjacent tiles). A smaller stride produces overlapping tiles.
        stride_y: As `stride_x`, along y.
        stride_z: As `stride_x`, along z.
        stride_t: As `stride_x`, along t.
        tail: What happens where the grid does not divide the axis.
            `"clip"` (default) shrinks the last tile to the border;
            `"balance"` re-splits the last full tile and the tail into two
            near-equal tiles, so a thin overhang never yields a thin tile
            (needs adjacent tiles, `stride == size`); `"shift"` moves the
            last tile back to stay full-size — it then overlaps its
            neighbour, so a writing iterator needs a merge or serial
            execution; `"drop"` discards the tail and leaves that border
            uncovered.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the tiled ROIs.
    """
    rois = by_grid(
        rois=self.rois,
        ref_image=self.ref_image,
        size_x=size_x,
        size_y=size_y,
        size_z=size_z,
        size_t=size_t,
        stride_x=stride_x,
        stride_y=stride_y,
        stride_z=stride_z,
        stride_t=stride_t,
        tail=tail,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_blocks

by_blocks(
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self

Divide each axis into a fixed number of near-equal blocks.

The complement of by_grid: you say how many tiles, not how big. Block lengths along an axis differ by at most one pixel, so there is no tail to police — the partition is balanced by construction.

Parameters:

  • num_x (int, default: 1 ) –

    Number of blocks along x (default 1: no split).

  • num_y (int, default: 1 ) –

    Number of blocks along y.

  • num_z (int, default: 1 ) –

    Number of blocks along z.

  • num_t (int, default: 1 ) –

    Number of blocks along t.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the blocks.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_blocks(
    self,
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self:
    """Divide each axis into a fixed number of near-equal blocks.

    The complement of `by_grid`: you say how many tiles, not how big.
    Block lengths along an axis differ by at most one pixel, so there is
    no tail to police — the partition is balanced by construction.

    Args:
        num_x: Number of blocks along x (default 1: no split).
        num_y: Number of blocks along y.
        num_z: Number of blocks along z.
        num_t: Number of blocks along t.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the blocks.
    """
    rois = by_blocks(
        rois=self.rois,
        ref_image=self.ref_image,
        num_x=num_x,
        num_y=num_y,
        num_z=num_z,
        num_t=num_t,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_yx

by_yx() -> Self

Split each ROI into one 2D plane per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, z, c) combination, each covering the full y/x extent — the shape for "run this 2D function on every plane".

Source code in src/ngio/iterators/_abstract_iterator.py
def by_yx(self) -> Self:
    """Split each ROI into one 2D plane per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, z, c)` combination, each covering the full y/x extent — the
    shape for "run this 2D function on every plane".
    """
    rois = by_yx(self.rois, self.ref_image)
    return self._new_from_rois(rois)

by_zyx

by_zyx(strict: bool = True) -> Self

Split each ROI into one 3D volume per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, c) combination, each covering the full z/y/x extent.

Parameters:

  • strict (bool, default: True ) –

    Require a real z axis (present and larger than 1); with False, images without one fall back to 2D planes.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_zyx(self, strict: bool = True) -> Self:
    """Split each ROI into one 3D volume per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, c)` combination, each covering the full z/y/x extent.

    Args:
        strict: Require a real z axis (present and larger than 1);
            with `False`, images without one fall back to 2D planes.
    """
    rois = by_zyx(self.rois, self.ref_image, strict=strict)
    return self._new_from_rois(rois)

by_chunks

by_chunks(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the input image's chunk grid.

Reads are chunk-granular, so this is the natural tiling for read-heavy work. For collision-free parallel writes use by_write_units, which tiles by the output's write granularity.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_chunks(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the input image's chunk grid.

    Reads are chunk-granular, so this is the natural tiling for
    read-heavy work. For collision-free parallel *writes* use
    `by_write_units`, which tiles by the output's write granularity.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=None,
    )
    return self._new_from_rois(rois)

by_write_units

by_write_units(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the output image's write granularity.

The write unit is the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. With no overlap the resulting ROIs pass check_if_write_units_overlap by construction, so a parallel map runs as a single fully-parallel wave — a throughput optimization; conflicting tilings are still safe, just scheduled in more waves. Falls back to the input chunk grid when the iterator is read-only.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_write_units(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the output image's write granularity.

    The write unit is the shard shape when the output is sharded (writes
    are atomic per shard object), the chunk shape otherwise. With no
    overlap the resulting ROIs pass `check_if_write_units_overlap` by
    construction, so a parallel `map` runs as a single fully-parallel
    wave — a throughput optimization; conflicting tilings are still safe,
    just scheduled in more waves. Falls back to the input chunk grid when
    the iterator is read-only.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=self.output_image,
    )
    return self._new_from_rois(rois)

product

product(other: list[Roi] | GenericRoiTable) -> Self

Cartesian product of the current ROIs with an arbitrary list of ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def product(self, other: list[Roi] | GenericRoiTable) -> Self:
    """Cartesian product of the current ROIs with an arbitrary list of ROIs."""
    if isinstance(other, GenericRoiTable):
        other = other.rois()
    rois = rois_product(self.rois, other)
    return self._new_from_rois(rois)

finalize

finalize() -> None

Resolve the stitch (when configured), then consolidate the pyramid.

Source code in src/ngio/iterators/_segmentation.py
def finalize(self) -> None:
    """Resolve the stitch (when configured), then consolidate the pyramid."""
    self._require_unrestricted_finalize()
    # The relabel has to precede consolidation: every pyramid level is
    # derived from level 0, so stitching after would leave them disagreeing.
    if self._stitch is not None:
        # The stitch resolve relabels level 0 wherever the union-find
        # reached, not just under the written ROIs — only a full rebuild
        # is guaranteed consistent after it.
        self._stitching_plan().resolve()
        self._stitch_plan = None
        self._output.consolidate(mode=self._consolidation_mode)
    else:
        self._output.consolidate(
            mode=self._consolidation_mode, regions=self._touched_write_regions()
        )

iter

iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[
        DataGetterProtocol[NumpyPipeType],
        DataSetterProtocol[NumpyPipeType],
    ]
]
iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[
        DataGetterProtocol[DaskPipeType],
        DataSetterProtocol[DaskPipeType],
    ]
]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[NumpyPipeType, DataSetterProtocol[NumpyPipeType]]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[
    tuple[DaskPipeType, DataSetterProtocol[DaskPipeType]]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"],
    *,
    batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
    lazy: Literal[False] = ...,
    data_mode: Literal["numpy"] | None = ...,
    iterator_mode: Literal["readwrite"] = ...,
    *,
    batch_size: int,
) -> Generator[
    tuple[
        list[NumpyPipeType],
        list[DataSetterProtocol[NumpyPipeType]],
    ]
]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"] | None,
    iterator_mode: Literal["readonly"],
    *,
    batch_size: int,
) -> Generator[list[NumpyPipeType]]
iter(
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal[
        "readwrite", "readonly"
    ] = "readwrite",
    *,
    batch_size: int | None = None,
) -> Generator

Create an iterator over the pixels of the ROIs.

A writing loop finalizes only when fully drained — on a stitching iterator that finalize is the resolve, so abandoning the loop mid-way leaves block-offset ids on disk (plus the transient scratch) until a re-run or a later finalize(). To write now and gather later on purpose, iterate a for_job(0, 1) slice and call finalize() on the unrestricted iterator when ready.

With batch_size set, a writing loop yields (patches, writers) — two aligned lists of up to batch_size items, in ROI order — and a read-only loop yields the payload lists alone. Stack the patches yourself (raggedness is yours to handle, unlike BatchedMapper's automatic padding), run the model once, and hand each result to its writer. Batches follow ROI order — under write_order="roi" that is the same later-ROI-wins order the mappers schedule, so the manual loop is bit-identical to map. Batching is numpy-only and eager: combining batch_size with lazy=True or data_mode="dask" raises.

Parameters:

  • lazy (bool, default: False ) –

    Yield getter/setter handles instead of materialized data.

  • data_mode (Literal['numpy', 'dask'] | None, default: None ) –

    "numpy" or "dask" (deprecated, see note).

  • iterator_mode (Literal['readwrite', 'readonly'], default: 'readwrite' ) –

    "readwrite" yields (patch, writer) pairs, "readonly" yields patches alone.

  • batch_size (int | None, default: None ) –

    When set, yield batches of up to this many items; the last batch may be smaller.

Example
for patches, writers in iterator.iter(data_mode="numpy", batch_size=8):
    outs = model(np.stack(patches))
    for writer, out in zip(writers, outs):
        writer(out)
Note

The dask data mode is deprecated and will be removed in ngio=1.2; from then on iter() yields numpy arrays.

Source code in src/ngio/iterators/_abstract_iterator.py
def iter(  # ty: ignore[invalid-method-override]
    self,
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    # "readwrite" by design — the one deliberate LSP variance of the
    # reader/writer split.
    iterator_mode: Literal["readwrite", "readonly"] = "readwrite",
    *,
    batch_size: int | None = None,
) -> Generator:
    """Create an iterator over the pixels of the ROIs.

    A writing loop finalizes only when fully drained — on a stitching
    iterator that finalize is the resolve, so abandoning the loop
    mid-way leaves block-offset ids on disk (plus the transient scratch)
    until a re-run or a later `finalize()`. To write now and gather
    later on purpose, iterate a `for_job(0, 1)` slice and call
    `finalize()` on the unrestricted iterator when ready.

    With `batch_size` set, a writing loop yields `(patches, writers)` —
    two aligned lists of up to `batch_size` items, in ROI order — and a
    read-only loop yields the payload lists alone. Stack the patches
    yourself (raggedness is yours to handle, unlike `BatchedMapper`'s
    automatic padding), run the model once, and hand each result to its
    writer. Batches follow ROI order — under `write_order="roi"` that
    is the same later-ROI-wins order the mappers schedule, so the
    manual loop is bit-identical to `map`. Batching is numpy-only
    and eager: combining `batch_size` with `lazy=True` or
    `data_mode="dask"` raises.

    Args:
        lazy: Yield getter/setter handles instead of materialized data.
        data_mode: `"numpy"` or `"dask"` (deprecated, see note).
        iterator_mode: `"readwrite"` yields `(patch, writer)` pairs,
            `"readonly"` yields patches alone.
        batch_size: When set, yield batches of up to this many items;
            the last batch may be smaller.

    Example:
        ```python
        for patches, writers in iterator.iter(data_mode="numpy", batch_size=8):
            outs = model(np.stack(patches))
            for writer, out in zip(writers, outs):
                writer(out)
        ```

    Note:
        The dask data mode is deprecated and will be removed in ngio=1.2;
        from then on `iter()` yields numpy arrays.
    """
    return self._iter_impl(
        lazy=lazy,
        data_mode=data_mode,
        iterator_mode=iterator_mode,
        batch_size=batch_size,
    )

iter_as_numpy

iter_as_numpy()

Alias for iter(data_mode="numpy").

Source code in src/ngio/iterators/_abstract_iterator.py
def iter_as_numpy(
    self,
):
    """Alias for `iter(data_mode="numpy")`."""
    return self._iter(lazy=False, data_mode="numpy", iterator_mode="readwrite")

iter_as_dask

iter_as_dask()

Alias for iter(data_mode="dask").

Deprecated: removed in ngio=1.2.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="iter_as_numpy() (or Image.get_as_dask() for a lazy array)",
    removed_in="1.2",
)
def iter_as_dask(
    self,
):
    """Alias for `iter(data_mode="dask")`.

    Deprecated: removed in ngio=1.2.
    """
    return self._iter(lazy=False, data_mode="dask", iterator_mode="readwrite")

for_job

for_job(job_index: int, n_jobs: int) -> Self

Restrict the iterator to one of n_jobs independent partitions.

No write unit is ever shared between two partitions, so the jobs need no locks, no barriers, and no channel to one another — safe as separate SLURM array tasks, in any order. Every job must build an identical iterator (construction is metadata-only and the partition derives deterministically from it), apply for_job last in the builder chain, and its verbs deliberately do not finalize — a slice's map/process/segment writes only its share, and measure/ detect bank a partial instead of joining. The gather step is the unrestricted iterator's finalize(), once, after all jobs. Stitched and read-only runs also need prepare_jobs(n_jobs) first. See the distributed-runs guide in the iterators docs.

Parameters:

  • job_index (int) –

    This job's partition, 0 <= job_index < n_jobs (e.g. int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops.

  • n_jobs (int) –

    How many partitions the run is split into; must match across every job and the gather.

Returns:

  • Self –

    The restricted iterator; further reshaping refuses.

Example
# array task:
iterator = SegmentationIterator(image, label).by_chunks()
iterator.for_job(job_index, n_jobs=n_jobs).map(func)

# gather task, after all jobs:
SegmentationIterator(image, label).by_chunks().finalize()
Source code in src/ngio/iterators/_abstract_iterator.py
def for_job(self, job_index: int, n_jobs: int) -> Self:
    """Restrict the iterator to one of `n_jobs` independent partitions.

    No write unit is ever shared between two partitions, so the jobs
    need no locks, no barriers, and no channel to one another — safe as
    separate SLURM array tasks, in any order. Every job must build an
    identical iterator (construction is metadata-only and the partition
    derives deterministically from it), apply `for_job` *last* in the
    builder chain, and its verbs deliberately do not finalize — a slice's
    `map`/`process`/`segment` writes only its share, and `measure`/
    `detect` bank a partial instead of joining. The gather step is the
    unrestricted iterator's `finalize()`, once, after all jobs.
    Stitched and read-only runs also need `prepare_jobs(n_jobs)` first.
    See the distributed-runs guide in the iterators docs.

    Args:
        job_index: This job's partition, `0 <= job_index < n_jobs`
            (e.g. `int($SLURM_ARRAY_TASK_ID)`). Partitions beyond the
            number of independent unit groups are empty no-ops.
        n_jobs: How many partitions the run is split into; must match
            across every job and the gather.

    Returns:
        The restricted iterator; further reshaping refuses.

    Example:
        ```python
        # array task:
        iterator = SegmentationIterator(image, label).by_chunks()
        iterator.for_job(job_index, n_jobs=n_jobs).map(func)

        # gather task, after all jobs:
        SegmentationIterator(image, label).by_chunks().finalize()
        ```
    """
    if self._partition is not None:
        _, current_index, current_n = self._partition
        raise NgioValueError(
            f"This iterator is already restricted to partition "
            f"{current_index}/{current_n}; partitions do not nest. "
            "Apply `for_job` once, to the full iterator."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    if not 0 <= job_index < n_jobs:
        raise NgioValueError(
            f"job_index must be in [0, {n_jobs}), got {job_index}."
        )
    self._require_splittable()
    self._validate_job_context(n_jobs)
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    restricted = self._new_from_rois(self.rois)
    restricted._partition = (
        frozenset(partition[job_index]),
        job_index,
        n_jobs,
    )
    return restricted

prepare_jobs

prepare_jobs(n_jobs: int) -> list[JobArgs]

Set up a distributed run and return its parallelization list.

The init step of the three-phase recipe (init → parallel jobs → consolidate): performs the iterator's setup — always clearing stale scratch state from earlier runs first — and returns one JobArgs per non-empty partition. Optional for plain writers; a stitching or read-only run requires it (the scratch root must exist, race-free, before any job writes into it).

Parameters:

  • n_jobs (int) –

    How many partitions to split into; empty ones are dropped from the list.

Returns:

  • list[JobArgs] –

    One {"job_index": ..., "n_jobs": ...} dict per non-empty

  • list[JobArgs] –

    partition — JSON-ready, splatting into for_job(**args).

Example
args_list = iterator.prepare_jobs(n_jobs=4)  # init task
iterator.for_job(**args).map(func)  # per parallel task
iterator.finalize()  # consolidate task
Source code in src/ngio/iterators/_abstract_iterator.py
def prepare_jobs(self, n_jobs: int) -> list[JobArgs]:
    """Set up a distributed run and return its parallelization list.

    The init step of the three-phase recipe (init → parallel jobs →
    consolidate): performs the iterator's setup — always clearing stale
    scratch state from earlier runs first — and returns one `JobArgs`
    per non-empty partition. Optional for plain writers; a stitching or
    read-only run requires it (the scratch root must exist, race-free,
    before any job writes into it).

    Args:
        n_jobs: How many partitions to split into; empty ones are
            dropped from the list.

    Returns:
        One `{"job_index": ..., "n_jobs": ...}` dict per non-empty
        partition — JSON-ready, splatting into `for_job(**args)`.

    Example:
        ```python
        args_list = iterator.prepare_jobs(n_jobs=4)  # init task
        iterator.for_job(**args).map(func)  # per parallel task
        iterator.finalize()  # consolidate task
        ```
    """
    if self._partition is not None:
        raise NgioValueError(
            "prepare_jobs belongs to the unrestricted iterator; this one "
            "is already restricted to a partition."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    self._require_splittable()
    if self.output_image is not None:
        self._validate_write_plan()
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    self._prepare_distributed(n_jobs)
    return [
        JobArgs(job_index=job_index, n_jobs=n_jobs)
        for job_index, indices in enumerate(partition)
        if indices
    ]

reduce

reduce(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Units are built read-only even on writable iterators: nothing is written and finalize does not run.

Parameters:

  • func (Callable[[NumpyPipeType], R]) –

    The function to apply; under a parallel mapper it must be safe on worker threads (or processes).

  • mapper (MapperProtocol[NumpyPipeType, R] | None, default: None ) –

    How the units are scheduled. None (the default) is serial; pass ThreadedMapper() or ProcessMapper() to fan out, sized by their own max_workers argument.

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run.

    Args:
        func: The function to apply; under a parallel mapper it must be
            safe on worker threads (or processes).
        mapper: How the units are scheduled. `None` (the default) is
            serial; pass `ThreadedMapper()` or `ProcessMapper()` to fan
            out, sized by their own `max_workers` argument.

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    if mapper is None:
        mapper = BasicMapper[NumpyPipeType, R]()
    return mapper(func, list(self._numpy_units_generator(with_setters=False)))

reduce_as_numpy

reduce_as_numpy(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Alias for reduce().

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce_as_numpy(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Alias for `reduce()`."""
    return self.reduce(func, mapper=mapper)

reduce_as_dask

reduce_as_dask(
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Deprecated: removed in ngio=1.2.

Units are built read-only even on writable iterators: nothing is written and finalize does not run. Runs serially — the units hand func lazy dask arrays, so the per-unit work is graph construction, not IO. A parallel mapper raises (see map_as_dask).

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="reduce() / reduce_as_numpy(mapper=...)",
    removed_in="1.2",
)
def reduce_as_dask(
    self,
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Deprecated: removed in ngio=1.2.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run. Runs serially — the units hand
    `func` *lazy* dask arrays, so the per-unit work is graph
    construction, not IO. A parallel `mapper` raises (see `map_as_dask`).

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    _mapper = self._require_serial_dask_mapper(mapper)
    return _mapper(func, list(self._dask_units_generator(with_setters=False)))

check_if_regions_overlap

check_if_regions_overlap() -> bool

Check if any of the ROIs overlap logically.

If two ROIs cover the same pixel, they are considered to overlap. This does not consider chunking or other storage details. Measured on the read regions — with a halo, neighbouring reads overlap by construction. For write safety, check_if_write_units_overlap is the right question.

Returns:

  • bool –

    True if any ROIs overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_regions_overlap(self) -> bool:
    """Check if any of the ROIs overlap logically.

    If two ROIs cover the same pixel, they are considered to overlap.
    This does not consider chunking or other storage details. Measured
    on the *read* regions — with a halo, neighbouring reads overlap by
    construction. For write safety, `check_if_write_units_overlap` is
    the right question.

    Returns:
        `True` if any ROIs overlap.
    """
    if len(self.rois) < 2:
        return False

    slicing_tuples = (
        g.slicing_ops.normalized_slicing_tuple
        for g in self._numpy_getters_generator()
    )
    return check_if_regions_overlap(slicing_tuples)

require_no_regions_overlap

require_no_regions_overlap() -> None

Ensure that the Iterator's ROIs do not overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_regions_overlap(self) -> None:
    """Ensure that the Iterator's ROIs do not overlap."""
    if self.check_if_regions_overlap():
        raise NgioValueError("Some rois overlap.")

map

map(
    func: Callable[[ndarray], ndarray],
    *,
    mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None

See WritingIteratorBuilder.map; also cleans up on failure.

A failed standalone run cannot be resolved, so the stitch scratch arrays are deleted rather than left as a stray _ngio_stitch group beside the resolution levels. Cleanup happens only when this run created the scratch: a partition slice, a resumed run, or the gather step opened a prepared root that holds the banks every other job wrote, and one failure must not destroy them — re-running is idempotent (banks rewrite, the id offsets are derived, not counted). The already-written tiles stay in every case.

Source code in src/ngio/iterators/_segmentation.py
def map(
    self,
    func: Callable[[np.ndarray], np.ndarray],
    *,
    mapper: MapperProtocol[np.ndarray, np.ndarray] | None = None,
) -> None:
    """See `WritingIteratorBuilder.map`; also cleans up on failure.

    A failed standalone run cannot be resolved, so the stitch scratch
    arrays are deleted rather than left as a stray `_ngio_stitch` group
    beside the resolution levels. Cleanup happens only when this run
    *created* the scratch: a partition slice, a resumed run, or the
    gather step opened a prepared root that holds the banks every other
    job wrote, and one failure must not destroy them — re-running is
    idempotent (banks rewrite, the id offsets are derived, not counted).
    The already-written tiles stay in every case.
    """
    if self._stitch is None:
        return super().map(func, mapper=mapper)
    # Build the plan (and let its validation warnings and errors fire)
    # before any tile runs: lazily it would surface mid-run, from a
    # worker, after some tiles have already written.
    plan = self._stitching_plan()
    # Resolve the scratch in the parent, before any unit is pickled or
    # run — never as a side effect of pickling.
    _ = plan.banks
    if self._partition is not None:
        return super().map(func, mapper=mapper)
    try:
        return super().map(func, mapper=mapper)
    except BaseException:
        if self._stitch_plan is not None and self._stitch_plan.created_banks:
            self._stitch_plan.cleanup()
            self._stitch_plan = None
        raise

map_as_numpy

map_as_numpy(
    func: Callable[[NumpyPipeType], NumpyPipeType],
    *,
    mapper: MapperProtocol[NumpyPipeType, NumpyPipeType]
    | None = None,
) -> None

Alias for map().

Source code in src/ngio/iterators/_abstract_iterator.py
def map_as_numpy(
    self,
    func: Callable[[NumpyPipeType], NumpyPipeType],
    *,
    mapper: MapperProtocol[NumpyPipeType, NumpyPipeType] | None = None,
) -> None:
    """Alias for `map()`."""
    return self.map(func, mapper=mapper)

map_as_dask

map_as_dask(
    func: Callable[[DaskPipeType], DaskPipeType],
    *,
    mapper: MapperProtocol[DaskPipeType, DaskPipeType]
    | None = None,
) -> None

Apply a transformation function to each ROI's patch and write it back.

Deprecated: removed in ngio=1.2.

Runs serially: the dask pipes are already executed by dask's own scheduler, and the dask write path (store_dask) scopes a global dask config option, so concurrent callers are unsafe. A parallel mapper raises.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="map() / map_as_numpy(mapper=...)",
    removed_in="1.2",
)
def map_as_dask(
    self,
    func: Callable[[DaskPipeType], DaskPipeType],
    *,
    mapper: MapperProtocol[DaskPipeType, DaskPipeType] | None = None,
) -> None:
    """Apply a transformation function to each ROI's patch and write it back.

    Deprecated: removed in ngio=1.2.

    Runs serially: the dask pipes are already executed by dask's own
    scheduler, and the dask write path (`store_dask`) scopes a global
    dask config option, so concurrent callers are unsafe. A parallel
    `mapper` raises.
    """
    _mapper = self._require_serial_dask_mapper(mapper)
    units = list(self._dask_units_generator())
    self._validate_write_plan()
    _mapper(func, units)
    if self._partition is None:
        self.finalize()

check_if_write_units_overlap

check_if_write_units_overlap() -> bool

Check if any two ROIs write into the same write unit of the output.

Measured on the write target: slicing tuples come from the setters and the grid is the output array's write granularity — the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. Two ROIs sharing a write unit make concurrent writes unsafe: the read-modify-write of that unit can lose data. The parallel mappers schedule such ROIs into separate waves, so this check answers "will my map run as a single fully-parallel wave?".

This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a loop.

Returns:

  • bool –

    True if any two ROIs share a write unit.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_write_units_overlap(self) -> bool:
    """Check if any two ROIs write into the same write unit of the output.

    Measured on the write target: slicing tuples come from the setters and
    the grid is the output array's write granularity — the shard shape when
    the output is sharded (writes are atomic per shard object), the chunk
    shape otherwise. Two ROIs sharing a write unit make concurrent writes
    unsafe: the read-modify-write of that unit can lose data. The parallel
    mappers schedule such ROIs into separate waves, so this check answers
    "will my map run as a single fully-parallel wave?".

    This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a
    loop.

    Returns:
        `True` if any two ROIs share a write unit.
    """
    if len(self.rois) < 2:
        return False

    footprints = (
        compute_write_footprint(setter)
        for setter in self._numpy_setters_generator()
        if setter is not None
    )
    non_empty = (footprint for footprint in footprints if footprint is not None)
    return any(chunk_rects_intersect(fi, fj) for fi, fj in _pairs_stream(non_empty))

require_no_write_units_overlap

require_no_write_units_overlap() -> None

Ensure that the ROIs do not share write units on the output.

The strict opt-in gate: the parallel mappers no longer refuse shared write units on their own (they wave-schedule around them), so call this to insist on a single-wave tiling instead.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_write_units_overlap(self) -> None:
    """Ensure that the ROIs do not share write units on the output.

    The strict opt-in gate: the parallel mappers no longer refuse shared
    write units on their own (they wave-schedule around them), so call
    this to insist on a single-wave tiling instead.
    """
    if self.check_if_write_units_overlap():
        raise NgioValueError("Some ROIs share write units on the output.")

with_stitch

with_stitch(config: StitchConfig | None = None) -> Self

Declare the stitch: split objects become one id at the gather.

Objects cut by a region boundary are resolved into one id when the run finalizes — any ROI list works: grids with a halo, overlapping FOV layouts, ragged tables. Needs a halo (with_halo) or overlapping ROIs: the evidence is overlap between neighbouring predictions. Declare before for_job; None uses the StitchConfig defaults. Refuses when on_overlap is declared — a merge policy would make the disk diverge from the banked predictions, so the resolve would relabel garbage.

Parameters:

  • config (StitchConfig | None, default: None ) –

    The stitch configuration, None for the defaults.

Source code in src/ngio/iterators/_segmentation.py
def with_stitch(self, config: StitchConfig | None = None) -> Self:
    """Declare the stitch: split objects become one id at the gather.

    Objects cut by a region boundary are resolved into one id when the
    run finalizes — any ROI list works: grids with a halo, overlapping
    FOV layouts, ragged tables. Needs a halo (`with_halo`) or
    overlapping ROIs: the evidence is overlap between neighbouring
    predictions. Declare before `for_job`; `None` uses the
    `StitchConfig` defaults. Refuses when `on_overlap` is declared —
    a merge policy would make the disk diverge from the banked
    predictions, so the resolve would relabel garbage.

    Args:
        config: The stitch configuration, `None` for the defaults.
    """
    resolved = config if config is not None else StitchConfig()
    if self._on_overlap is not None:
        raise NgioValueError(
            "with_stitch and on_overlap cannot be combined: a merge "
            "policy would make the written labels diverge from the "
            "banked predictions the resolve compares. The stitch "
            "already owns the contested pixels (deterministic write "
            "order) — drop the `on_overlap` declaration."
        )
    _require_no_unique_labels_with_stitch(resolved, self._output_transforms)
    new_instance = self._new_from_rois(self.rois)
    new_instance._stitch = resolved
    return new_instance

segment

segment(
    func: Callable[[ndarray], ndarray],
    *,
    mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None

Segment every region and write the labels; the topic verb for map.

A serial run finalizes automatically (the stitch resolve included). On a for_job slice it segments only this job's share; the gather is the unrestricted iterator's finalize(), once, after all jobs.

Parameters:

  • func (Callable[[ndarray], ndarray]) –

    The segmentation model. Under a parallel mapper it runs on worker threads (or processes) and must be safe there.

  • mapper (MapperProtocol[ndarray, ndarray] | None, default: None ) –

    How the units are scheduled; see map.

Source code in src/ngio/iterators/_segmentation.py
def segment(
    self,
    func: Callable[[np.ndarray], np.ndarray],
    *,
    mapper: MapperProtocol[np.ndarray, np.ndarray] | None = None,
) -> None:
    """Segment every region and write the labels; the topic verb for `map`.

    A serial run finalizes automatically (the stitch resolve included).
    On a `for_job` slice it segments only this job's share; the gather is
    the unrestricted iterator's `finalize()`, once, after all jobs.

    Args:
        func: The segmentation model. Under a parallel mapper it runs on
            worker threads (or processes) and must be safe there.
        mapper: How the units are scheduled; see `map`.
    """
    self.map(func, mapper=mapper)

on_overlap

on_overlap(
    policy: OverlapPolicy,
    *,
    write_order: WriteOrder = "roi",
) -> Self

Refused: masked writes never contest.

Each write touches only its own object's pixels (MaskMerge protects the rest), so there is nothing to declare — and the one merge= slot already carries the mask protection.

Source code in src/ngio/iterators/_segmentation.py
def on_overlap(
    self, policy: OverlapPolicy, *, write_order: WriteOrder = "roi"
) -> Self:
    """Refused: masked writes never contest.

    Each write touches only its own object's pixels (`MaskMerge`
    protects the rest), so there is nothing to declare — and the one
    `merge=` slot already carries the mask protection.
    """
    raise NgioValueError(
        "MaskedSegmentationIterator takes no overlap policy: writes are "
        "mask-protected, so overlapping bounding boxes never contest a "
        "pixel — there is nothing to declare."
    )

build_numpy_getter

build_numpy_getter(roi: Roi)
Source code in src/ngio/iterators/_segmentation.py
def build_numpy_getter(self, roi: Roi):
    return NumpyGetter(
        zarr_array=self._input.zarr_array,
        dimensions=self._input.dimensions,
        roi=self._read_roi(roi),
        axes_order=self._axes_order,
        transforms=self._input_transforms_with_mask(),
        slicing_dict=self._input_slicing_kwargs,
    )

build_numpy_setter

build_numpy_setter(roi: Roi)
Source code in src/ngio/iterators/_segmentation.py
def build_numpy_setter(self, roi: Roi):
    return self._wrap_for_stitch(
        self._wrap_setter(
            NumpySetter(
                roi=roi,
                zarr_array=self._output.zarr_array,
                dimensions=self._output.dimensions,
                axes_order=self._axes_order,
                transforms=self._output_transforms,
                merge=self._output_mask_merge(),
                remove_channel_selection=True,
            ),
            roi,
        ),
        roi,
    )

build_dask_getter

build_dask_getter(roi: Roi)
Source code in src/ngio/iterators/_segmentation.py
def build_dask_getter(self, roi: Roi):
    return DaskGetter(
        roi=self._read_roi(roi),
        zarr_array=self._input.zarr_array,
        dimensions=self._input.dimensions,
        axes_order=self._axes_order,
        transforms=self._input_transforms_with_mask(),
        slicing_dict=self._input_slicing_kwargs,
    )

build_dask_setter

build_dask_setter(roi: Roi)
Source code in src/ngio/iterators/_segmentation.py
def build_dask_setter(self, roi: Roi):
    self._require_stitchless_dask()
    return self._wrap_setter(
        DaskSetter(
            roi=roi,
            zarr_array=self._output.zarr_array,
            dimensions=self._output.dimensions,
            axes_order=self._axes_order,
            transforms=self._output_transforms,
            merge=self._output_mask_merge(),
            remove_channel_selection=True,
        ),
        roi,
    )

FeatureExtractorIterator

ngio.iterators.FeatureExtractorIterator

FeatureExtractorIterator(
    input_image: Image,
    input_label: Label,
    *,
    channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol]
    | None = None,
    label_transforms: Sequence[TransformProtocol]
    | None = None,
)

Bases: AbstractIteratorBuilder[NumpyPipeType, DaskPipeType, Table]

Measure image/label pairs region by region; nothing is written.

Each unit pairs the image patch with the label patch over the same region. reduce collects per-region results; measure joins them into a single FeatureTable for the caller to store. Distributed, a for_job slice's measure banks a partial and finalize() runs the one global join (declared with with_join, ConcatJoin otherwise). with_halo(...) reads context around each region; the duplicate rows it produces for border objects are reconciled in your declared join via the stamped roi_index/roi_name columns.

Measure input_image/input_label pairs; nothing is written.

Parameters:

  • input_image (Image) –

    The image to measure.

  • input_label (Label) –

    The label whose objects tie the measurements together.

  • channel_selection (ChannelSlicingInputType, default: None ) –

    Restrict the image reads to these channels.

  • axes_order (Sequence[str] | None, default: None ) –

    Axes order of the patches handed to the function.

  • input_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Transforms applied to each image patch.

  • label_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Transforms applied to each label patch.

Source code in src/ngio/iterators/_feature.py
def __init__(
    self,
    input_image: Image,
    input_label: Label,
    *,
    channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol] | None = None,
    label_transforms: Sequence[TransformProtocol] | None = None,
) -> None:
    """Measure `input_image`/`input_label` pairs; nothing is written.

    Args:
        input_image: The image to measure.
        input_label: The label whose objects tie the measurements together.
        channel_selection: Restrict the image reads to these channels.
        axes_order: Axes order of the patches handed to the function.
        input_transforms: Transforms applied to each image patch.
        label_transforms: Transforms applied to each label patch.
    """
    self._input = input_image
    self._input_label = input_label
    self._ref_image = input_image
    self._rois = input_image.build_image_roi_table(name=None).rois()

    self._input_slicing_kwargs = add_channel_selection_to_slicing_dict(
        image=self._input, channel_selection=channel_selection, slicing_dict={}
    )
    self._channel_selection = channel_selection
    self._axes_order = axes_order
    self._input_transforms = input_transforms
    self._label_transforms = label_transforms

    self._input.require_axes_match(self._input_label)
    self._input.require_rescalable(self._input_label)

rois property

rois: list[Roi]

Get the list of ROIs for the iterator.

ref_image property

ref_image: AbstractImage

Get the reference image for the iterator.

output_image property

output_image: AbstractImage | None

The image this iterator writes to, or None for a read-only iterator.

halo property

halo: dict[str, int]

The per-axis read margin, empty when there is none.

partition_indices property

partition_indices: list[int] | None

The unit indices this iterator is restricted to, or None.

None on an unrestricted iterator; on a for_job slice, the sorted positions (in the full ROI list) of the units it runs. The whole layout of a split is [it.for_job(i, n).partition_indices for i in range(n)] — the pre-flight check before submitting a distributed run: one fat list plus empties means the output chunking, not the cluster, is the constraint.

with_halo

with_halo(
    x: int = 0, y: int = 0, z: int = 0, t: int = 0
) -> Self

Read a margin of context around each ROI, but write only the ROI.

The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.

Note

The margins do not survive a rescale: a shape-changing transform on the data path (a ZoomTransform) makes the crop refuse — iterate on a coarser pyramid level instead. A MaskTransform composes normally (its zoom rescales the label it reads).

Parameters:

  • x (int, default: 0 ) –

    Pixels added on each side along x.

  • y (int, default: 0 ) –

    Pixels added on each side along y.

  • z (int, default: 0 ) –

    Pixels added on each side along z.

  • t (int, default: 0 ) –

    Frames added on each side along t.

Returns:

  • Self –

    A new iterator reading with the halo.

Raises:

  • NgioValueError –

    On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps roi_index/roi_name so your coalesce can), on an in-place iterator, or for a negative margin.

Example
# 8 px of context per side, disjoint writes, still parallel.
it = iterator.by_write_units().with_halo(x=8, y=8)
it.map(smooth, mapper=ThreadedMapper("auto"))
Source code in src/ngio/iterators/_abstract_iterator.py
def with_halo(self, x: int = 0, y: int = 0, z: int = 0, t: int = 0) -> Self:
    """Read a margin of context around each ROI, but write only the ROI.

    The function receives the grown region and must return it grown too;
    the border is cropped before the write, so tiles come out seamless.
    Margins are reference-image pixels, clipped at the borders. Write
    footprints are unchanged, so a haloed iterator parallelizes exactly
    as far as it did without one.

    Note:
        The margins do not survive a rescale: a shape-changing transform
        on the data path (a `ZoomTransform`) makes the crop refuse —
        iterate on a coarser pyramid level instead. A `MaskTransform`
        composes normally (its zoom rescales the label it reads).

    Args:
        x: Pixels added on each side along x.
        y: Pixels added on each side along y.
        z: Pixels added on each side along z.
        t: Frames added on each side along t.

    Returns:
        A new iterator reading with the halo.

    Raises:
        NgioValueError: On a read-only iterator that has not opted in
            (no write to crop the halo from; detection opts in and
            reconciles the overlap with NMS, feature extraction opts in
            and stamps `roi_index`/`roi_name` so your `coalesce` can),
            on an in-place iterator, or for a negative margin.

    Example:
        ```python
        # 8 px of context per side, disjoint writes, still parallel.
        it = iterator.by_write_units().with_halo(x=8, y=8)
        it.map(smooth, mapper=ThreadedMapper("auto"))
        ```
    """
    halo = {"x": x, "y": y, "z": z, "t": t}
    for axis_name, margin in halo.items():
        if margin < 0:
            raise NgioValueError(
                f"Halo along '{axis_name}' must be >= 0, got {margin}."
            )
    if self.output_image is None and not self._allow_readonly_halo:
        name = self.__class__.__name__
        raise NgioValueError(
            f"{name} is read-only, so there is no written region for a "
            "halo to surround. Widen the ROIs themselves instead (e.g. "
            "`by_grid(...)` with a stride smaller than the size)."
        )
    output = self.output_image
    if output is not None and _is_same_zarr_array(
        self._ref_image.zarr_array, output.zarr_array
    ):
        raise NgioValueError(
            "In-place iteration cannot take a halo: each tile's grown "
            "read covers pixels that neighbouring tiles write, so the "
            "result depends on execution order. Write to a different "
            "image or label, or drop the halo."
        )
    new_instance = self._new_from_rois(self.rois)
    new_instance._halo = {ax: m for ax, m in halo.items() if m > 0}
    return new_instance

by_grid

by_grid(
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self

Tile the current ROIs with a regular grid.

Parameters:

  • size_x (int | None, default: None ) –

    Tile size along x; None (default) means the full extent.

  • size_y (int | None, default: None ) –

    Tile size along y; None means the full extent.

  • size_z (int | None, default: None ) –

    Tile size along z; None means the full extent.

  • size_t (int | None, default: None ) –

    Tile size along t; None means the full extent.

  • stride_x (int | None, default: None ) –

    Step between tiles along x; None means size_x (adjacent tiles). A smaller stride produces overlapping tiles.

  • stride_y (int | None, default: None ) –

    As stride_x, along y.

  • stride_z (int | None, default: None ) –

    As stride_x, along z.

  • stride_t (int | None, default: None ) –

    As stride_x, along t.

  • tail (TailPolicy, default: 'clip' ) –

    What happens where the grid does not divide the axis. "clip" (default) shrinks the last tile to the border; "balance" re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles, stride == size); "shift" moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution; "drop" discards the tail and leaves that border uncovered.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_grid(
    self,
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self:
    """Tile the current ROIs with a regular grid.

    Args:
        size_x: Tile size along x; `None` (default) means the full extent.
        size_y: Tile size along y; `None` means the full extent.
        size_z: Tile size along z; `None` means the full extent.
        size_t: Tile size along t; `None` means the full extent.
        stride_x: Step between tiles along x; `None` means `size_x`
            (adjacent tiles). A smaller stride produces overlapping tiles.
        stride_y: As `stride_x`, along y.
        stride_z: As `stride_x`, along z.
        stride_t: As `stride_x`, along t.
        tail: What happens where the grid does not divide the axis.
            `"clip"` (default) shrinks the last tile to the border;
            `"balance"` re-splits the last full tile and the tail into two
            near-equal tiles, so a thin overhang never yields a thin tile
            (needs adjacent tiles, `stride == size`); `"shift"` moves the
            last tile back to stay full-size — it then overlaps its
            neighbour, so a writing iterator needs a merge or serial
            execution; `"drop"` discards the tail and leaves that border
            uncovered.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the tiled ROIs.
    """
    rois = by_grid(
        rois=self.rois,
        ref_image=self.ref_image,
        size_x=size_x,
        size_y=size_y,
        size_z=size_z,
        size_t=size_t,
        stride_x=stride_x,
        stride_y=stride_y,
        stride_z=stride_z,
        stride_t=stride_t,
        tail=tail,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_blocks

by_blocks(
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self

Divide each axis into a fixed number of near-equal blocks.

The complement of by_grid: you say how many tiles, not how big. Block lengths along an axis differ by at most one pixel, so there is no tail to police — the partition is balanced by construction.

Parameters:

  • num_x (int, default: 1 ) –

    Number of blocks along x (default 1: no split).

  • num_y (int, default: 1 ) –

    Number of blocks along y.

  • num_z (int, default: 1 ) –

    Number of blocks along z.

  • num_t (int, default: 1 ) –

    Number of blocks along t.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the blocks.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_blocks(
    self,
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self:
    """Divide each axis into a fixed number of near-equal blocks.

    The complement of `by_grid`: you say how many tiles, not how big.
    Block lengths along an axis differ by at most one pixel, so there is
    no tail to police — the partition is balanced by construction.

    Args:
        num_x: Number of blocks along x (default 1: no split).
        num_y: Number of blocks along y.
        num_z: Number of blocks along z.
        num_t: Number of blocks along t.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the blocks.
    """
    rois = by_blocks(
        rois=self.rois,
        ref_image=self.ref_image,
        num_x=num_x,
        num_y=num_y,
        num_z=num_z,
        num_t=num_t,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_yx

by_yx() -> Self

Split each ROI into one 2D plane per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, z, c) combination, each covering the full y/x extent — the shape for "run this 2D function on every plane".

Source code in src/ngio/iterators/_abstract_iterator.py
def by_yx(self) -> Self:
    """Split each ROI into one 2D plane per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, z, c)` combination, each covering the full y/x extent — the
    shape for "run this 2D function on every plane".
    """
    rois = by_yx(self.rois, self.ref_image)
    return self._new_from_rois(rois)

by_zyx

by_zyx(strict: bool = True) -> Self

Split each ROI into one 3D volume per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, c) combination, each covering the full z/y/x extent.

Parameters:

  • strict (bool, default: True ) –

    Require a real z axis (present and larger than 1); with False, images without one fall back to 2D planes.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_zyx(self, strict: bool = True) -> Self:
    """Split each ROI into one 3D volume per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, c)` combination, each covering the full z/y/x extent.

    Args:
        strict: Require a real z axis (present and larger than 1);
            with `False`, images without one fall back to 2D planes.
    """
    rois = by_zyx(self.rois, self.ref_image, strict=strict)
    return self._new_from_rois(rois)

by_chunks

by_chunks(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the input image's chunk grid.

Reads are chunk-granular, so this is the natural tiling for read-heavy work. For collision-free parallel writes use by_write_units, which tiles by the output's write granularity.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_chunks(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the input image's chunk grid.

    Reads are chunk-granular, so this is the natural tiling for
    read-heavy work. For collision-free parallel *writes* use
    `by_write_units`, which tiles by the output's write granularity.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=None,
    )
    return self._new_from_rois(rois)

by_write_units

by_write_units(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the output image's write granularity.

The write unit is the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. With no overlap the resulting ROIs pass check_if_write_units_overlap by construction, so a parallel map runs as a single fully-parallel wave — a throughput optimization; conflicting tilings are still safe, just scheduled in more waves. Falls back to the input chunk grid when the iterator is read-only.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_write_units(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the output image's write granularity.

    The write unit is the shard shape when the output is sharded (writes
    are atomic per shard object), the chunk shape otherwise. With no
    overlap the resulting ROIs pass `check_if_write_units_overlap` by
    construction, so a parallel `map` runs as a single fully-parallel
    wave — a throughput optimization; conflicting tilings are still safe,
    just scheduled in more waves. Falls back to the input chunk grid when
    the iterator is read-only.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=self.output_image,
    )
    return self._new_from_rois(rois)

product

product(other: list[Roi] | GenericRoiTable) -> Self

Cartesian product of the current ROIs with an arbitrary list of ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def product(self, other: list[Roi] | GenericRoiTable) -> Self:
    """Cartesian product of the current ROIs with an arbitrary list of ROIs."""
    if isinstance(other, GenericRoiTable):
        other = other.rois()
    rois = rois_product(self.rois, other)
    return self._new_from_rois(rois)

iter

iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
    lazy: Literal[False] = ...,
    data_mode: Literal["numpy"] | None = ...,
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: int,
) -> Generator[list[NumpyPipeType]]
iter(
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal[
        "readwrite", "readonly"
    ] = "readonly",
    *,
    batch_size: int | None = None,
) -> Generator

Iterate the ROIs' payloads, read-only.

With batch_size set, yields payload lists of up to that many items (the last batch may be smaller). Batching is numpy-only and eager: combining batch_size with lazy=True or data_mode="dask" raises. On a writing iterator (see WritingIteratorBuilder.iter) the default mode is "readwrite" and the loop yields (patch, writer) pairs instead.

Parameters:

  • lazy (bool, default: False ) –

    Yield getter handles instead of materialized data.

  • data_mode (Literal['numpy', 'dask'] | None, default: None ) –

    "numpy" or "dask" (deprecated, see note).

  • iterator_mode (Literal['readwrite', 'readonly'], default: 'readonly' ) –

    "readonly" yields patches alone; on a writing iterator "readwrite" yields (patch, writer) pairs.

  • batch_size (int | None, default: None ) –

    When set, yield batches of up to this many items; the last batch may be smaller.

Note

The dask data mode is deprecated and will be removed in ngio=1.2; from then on iter() yields numpy arrays.

Source code in src/ngio/iterators/_abstract_iterator.py
def iter(
    self,
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal["readwrite", "readonly"] = "readonly",
    *,
    batch_size: int | None = None,
) -> Generator:
    """Iterate the ROIs' payloads, read-only.

    With `batch_size` set, yields payload lists of up to that many
    items (the last batch may be smaller). Batching is numpy-only and
    eager: combining `batch_size` with `lazy=True` or
    `data_mode="dask"` raises. On a writing iterator (see
    `WritingIteratorBuilder.iter`) the default mode is `"readwrite"`
    and the loop yields `(patch, writer)` pairs instead.

    Args:
        lazy: Yield getter handles instead of materialized data.
        data_mode: `"numpy"` or `"dask"` (deprecated, see note).
        iterator_mode: `"readonly"` yields patches alone; on a writing
            iterator `"readwrite"` yields `(patch, writer)` pairs.
        batch_size: When set, yield batches of up to this many items;
            the last batch may be smaller.

    Note:
        The dask data mode is deprecated and will be removed in ngio=1.2;
        from then on `iter()` yields numpy arrays.
    """
    return self._iter_impl(
        lazy=lazy,
        data_mode=data_mode,
        iterator_mode=iterator_mode,
        batch_size=batch_size,
    )

iter_as_numpy

iter_as_numpy()

Alias for iter(data_mode="numpy").

Source code in src/ngio/iterators/_abstract_iterator.py
def iter_as_numpy(
    self,
):
    """Alias for `iter(data_mode="numpy")`."""
    return self._iter(lazy=False, data_mode="numpy", iterator_mode="readonly")

iter_as_dask

iter_as_dask()

Alias for iter(data_mode="dask").

Deprecated: removed in ngio=1.2.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="iter_as_numpy() (or Image.get_as_dask() for a lazy array)",
    removed_in="1.2",
)
def iter_as_dask(
    self,
):
    """Alias for `iter(data_mode="dask")`.

    Deprecated: removed in ngio=1.2.
    """
    return self._iter(lazy=False, data_mode="dask", iterator_mode="readonly")

for_job

for_job(job_index: int, n_jobs: int) -> Self

Restrict the iterator to one of n_jobs independent partitions.

No write unit is ever shared between two partitions, so the jobs need no locks, no barriers, and no channel to one another — safe as separate SLURM array tasks, in any order. Every job must build an identical iterator (construction is metadata-only and the partition derives deterministically from it), apply for_job last in the builder chain, and its verbs deliberately do not finalize — a slice's map/process/segment writes only its share, and measure/ detect bank a partial instead of joining. The gather step is the unrestricted iterator's finalize(), once, after all jobs. Stitched and read-only runs also need prepare_jobs(n_jobs) first. See the distributed-runs guide in the iterators docs.

Parameters:

  • job_index (int) –

    This job's partition, 0 <= job_index < n_jobs (e.g. int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops.

  • n_jobs (int) –

    How many partitions the run is split into; must match across every job and the gather.

Returns:

  • Self –

    The restricted iterator; further reshaping refuses.

Example
# array task:
iterator = SegmentationIterator(image, label).by_chunks()
iterator.for_job(job_index, n_jobs=n_jobs).map(func)

# gather task, after all jobs:
SegmentationIterator(image, label).by_chunks().finalize()
Source code in src/ngio/iterators/_abstract_iterator.py
def for_job(self, job_index: int, n_jobs: int) -> Self:
    """Restrict the iterator to one of `n_jobs` independent partitions.

    No write unit is ever shared between two partitions, so the jobs
    need no locks, no barriers, and no channel to one another — safe as
    separate SLURM array tasks, in any order. Every job must build an
    identical iterator (construction is metadata-only and the partition
    derives deterministically from it), apply `for_job` *last* in the
    builder chain, and its verbs deliberately do not finalize — a slice's
    `map`/`process`/`segment` writes only its share, and `measure`/
    `detect` bank a partial instead of joining. The gather step is the
    unrestricted iterator's `finalize()`, once, after all jobs.
    Stitched and read-only runs also need `prepare_jobs(n_jobs)` first.
    See the distributed-runs guide in the iterators docs.

    Args:
        job_index: This job's partition, `0 <= job_index < n_jobs`
            (e.g. `int($SLURM_ARRAY_TASK_ID)`). Partitions beyond the
            number of independent unit groups are empty no-ops.
        n_jobs: How many partitions the run is split into; must match
            across every job and the gather.

    Returns:
        The restricted iterator; further reshaping refuses.

    Example:
        ```python
        # array task:
        iterator = SegmentationIterator(image, label).by_chunks()
        iterator.for_job(job_index, n_jobs=n_jobs).map(func)

        # gather task, after all jobs:
        SegmentationIterator(image, label).by_chunks().finalize()
        ```
    """
    if self._partition is not None:
        _, current_index, current_n = self._partition
        raise NgioValueError(
            f"This iterator is already restricted to partition "
            f"{current_index}/{current_n}; partitions do not nest. "
            "Apply `for_job` once, to the full iterator."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    if not 0 <= job_index < n_jobs:
        raise NgioValueError(
            f"job_index must be in [0, {n_jobs}), got {job_index}."
        )
    self._require_splittable()
    self._validate_job_context(n_jobs)
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    restricted = self._new_from_rois(self.rois)
    restricted._partition = (
        frozenset(partition[job_index]),
        job_index,
        n_jobs,
    )
    return restricted

prepare_jobs

prepare_jobs(n_jobs: int) -> list[JobArgs]

Set up a distributed run and return its parallelization list.

The init step of the three-phase recipe (init → parallel jobs → consolidate): performs the iterator's setup — always clearing stale scratch state from earlier runs first — and returns one JobArgs per non-empty partition. Optional for plain writers; a stitching or read-only run requires it (the scratch root must exist, race-free, before any job writes into it).

Parameters:

  • n_jobs (int) –

    How many partitions to split into; empty ones are dropped from the list.

Returns:

  • list[JobArgs] –

    One {"job_index": ..., "n_jobs": ...} dict per non-empty

  • list[JobArgs] –

    partition — JSON-ready, splatting into for_job(**args).

Example
args_list = iterator.prepare_jobs(n_jobs=4)  # init task
iterator.for_job(**args).map(func)  # per parallel task
iterator.finalize()  # consolidate task
Source code in src/ngio/iterators/_abstract_iterator.py
def prepare_jobs(self, n_jobs: int) -> list[JobArgs]:
    """Set up a distributed run and return its parallelization list.

    The init step of the three-phase recipe (init → parallel jobs →
    consolidate): performs the iterator's setup — always clearing stale
    scratch state from earlier runs first — and returns one `JobArgs`
    per non-empty partition. Optional for plain writers; a stitching or
    read-only run requires it (the scratch root must exist, race-free,
    before any job writes into it).

    Args:
        n_jobs: How many partitions to split into; empty ones are
            dropped from the list.

    Returns:
        One `{"job_index": ..., "n_jobs": ...}` dict per non-empty
        partition — JSON-ready, splatting into `for_job(**args)`.

    Example:
        ```python
        args_list = iterator.prepare_jobs(n_jobs=4)  # init task
        iterator.for_job(**args).map(func)  # per parallel task
        iterator.finalize()  # consolidate task
        ```
    """
    if self._partition is not None:
        raise NgioValueError(
            "prepare_jobs belongs to the unrestricted iterator; this one "
            "is already restricted to a partition."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    self._require_splittable()
    if self.output_image is not None:
        self._validate_write_plan()
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    self._prepare_distributed(n_jobs)
    return [
        JobArgs(job_index=job_index, n_jobs=n_jobs)
        for job_index, indices in enumerate(partition)
        if indices
    ]

reduce

reduce(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Units are built read-only even on writable iterators: nothing is written and finalize does not run.

Parameters:

  • func (Callable[[NumpyPipeType], R]) –

    The function to apply; under a parallel mapper it must be safe on worker threads (or processes).

  • mapper (MapperProtocol[NumpyPipeType, R] | None, default: None ) –

    How the units are scheduled. None (the default) is serial; pass ThreadedMapper() or ProcessMapper() to fan out, sized by their own max_workers argument.

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run.

    Args:
        func: The function to apply; under a parallel mapper it must be
            safe on worker threads (or processes).
        mapper: How the units are scheduled. `None` (the default) is
            serial; pass `ThreadedMapper()` or `ProcessMapper()` to fan
            out, sized by their own `max_workers` argument.

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    if mapper is None:
        mapper = BasicMapper[NumpyPipeType, R]()
    return mapper(func, list(self._numpy_units_generator(with_setters=False)))

reduce_as_numpy

reduce_as_numpy(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Alias for reduce().

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce_as_numpy(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Alias for `reduce()`."""
    return self.reduce(func, mapper=mapper)

reduce_as_dask

reduce_as_dask(
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Deprecated: removed in ngio=1.2.

Units are built read-only even on writable iterators: nothing is written and finalize does not run. Runs serially — the units hand func lazy dask arrays, so the per-unit work is graph construction, not IO. A parallel mapper raises (see map_as_dask).

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="reduce() / reduce_as_numpy(mapper=...)",
    removed_in="1.2",
)
def reduce_as_dask(
    self,
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Deprecated: removed in ngio=1.2.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run. Runs serially — the units hand
    `func` *lazy* dask arrays, so the per-unit work is graph
    construction, not IO. A parallel `mapper` raises (see `map_as_dask`).

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    _mapper = self._require_serial_dask_mapper(mapper)
    return _mapper(func, list(self._dask_units_generator(with_setters=False)))

check_if_regions_overlap

check_if_regions_overlap() -> bool

Check if any of the ROIs overlap logically.

If two ROIs cover the same pixel, they are considered to overlap. This does not consider chunking or other storage details. Measured on the read regions — with a halo, neighbouring reads overlap by construction. For write safety, check_if_write_units_overlap is the right question.

Returns:

  • bool –

    True if any ROIs overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_regions_overlap(self) -> bool:
    """Check if any of the ROIs overlap logically.

    If two ROIs cover the same pixel, they are considered to overlap.
    This does not consider chunking or other storage details. Measured
    on the *read* regions — with a halo, neighbouring reads overlap by
    construction. For write safety, `check_if_write_units_overlap` is
    the right question.

    Returns:
        `True` if any ROIs overlap.
    """
    if len(self.rois) < 2:
        return False

    slicing_tuples = (
        g.slicing_ops.normalized_slicing_tuple
        for g in self._numpy_getters_generator()
    )
    return check_if_regions_overlap(slicing_tuples)

require_no_regions_overlap

require_no_regions_overlap() -> None

Ensure that the Iterator's ROIs do not overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_regions_overlap(self) -> None:
    """Ensure that the Iterator's ROIs do not overlap."""
    if self.check_if_regions_overlap():
        raise NgioValueError("Some rois overlap.")

with_join

with_join(join: JoinProtocol) -> Self

Declare the join that turns the per-ROI results into the table.

Runs once, wherever the join executes — a serial measure, or the distributed gather's finalize. On a for_job slice the declaration is inert: the slice banks the pre-join records regardless, so declaring it there is harmless. Like the measurement function, the join is not part of the plan fingerprint: declare the identical one on the gather iterator. The default (no declaration) is ConcatJoin referencing the input label.

Parameters:

  • join (JoinProtocol) –

    Anything satisfying JoinProtocol — a plain callable taking the normalized frame list, or a configured class like ConcatJoin(reference_label=...).

Source code in src/ngio/iterators/_feature.py
def with_join(self, join: JoinProtocol) -> Self:
    """Declare the join that turns the per-ROI results into the table.

    Runs once, wherever the join executes — a serial `measure`, or the
    distributed gather's `finalize`. On a `for_job` slice the
    declaration is inert: the slice banks the pre-join records
    regardless, so declaring it there is harmless. Like the
    measurement function, the join is not part of the plan
    fingerprint: declare the identical one on the gather iterator.
    The default (no declaration) is `ConcatJoin` referencing the
    input label.

    Args:
        join: Anything satisfying `JoinProtocol` — a plain callable
            taking the normalized frame list, or a configured class
            like `ConcatJoin(reference_label=...)`.
    """
    if not callable(join):
        raise NgioValueError(
            f"The join must be callable, got {type(join).__name__}."
        )
    new_instance = self._new_from_rois(self.rois)
    new_instance._join = join
    return new_instance

build_numpy_getter

build_numpy_getter(roi: Roi) -> FeatureGetter[ndarray]
Source code in src/ngio/iterators/_feature.py
def build_numpy_getter(self, roi: Roi) -> FeatureGetter[np.ndarray]:
    # Both getters take the same halo-grown world-space ROI; each
    # converts at its own image's pixel size, so a coarser label
    # rescales the margin correctly on its own.
    read_roi = self._read_roi(roi)
    data_getter = NumpyGetter(
        zarr_array=self._input.zarr_array,
        dimensions=self._input.dimensions,
        axes_order=self._axes_order,
        transforms=self._input_transforms,
        roi=read_roi,
        slicing_dict=self._input_slicing_kwargs,
    )
    label_getter = NumpyGetter(
        zarr_array=self._input_label.zarr_array,
        dimensions=self._input_label.dimensions,
        axes_order=self._axes_order,
        transforms=self._label_transforms,
        roi=read_roi,
        remove_channel_selection=True,
    )
    return FeatureGetter(data_getter, label_getter)

build_dask_getter

build_dask_getter(roi: Roi) -> FeatureGetter[Array]
Source code in src/ngio/iterators/_feature.py
def build_dask_getter(self, roi: Roi) -> FeatureGetter[da.Array]:
    read_roi = self._read_roi(roi)
    data_getter = DaskGetter(
        zarr_array=self._input.zarr_array,
        dimensions=self._input.dimensions,
        axes_order=self._axes_order,
        transforms=self._input_transforms,
        roi=read_roi,
        slicing_dict=self._input_slicing_kwargs,
    )
    label_getter = DaskGetter(
        zarr_array=self._input_label.zarr_array,
        dimensions=self._input_label.dimensions,
        axes_order=self._axes_order,
        transforms=self._label_transforms,
        roi=read_roi,
        remove_channel_selection=True,
    )
    return FeatureGetter(data_getter, label_getter)

finalize

finalize() -> Table

Merge a distributed run's partials into the one final table.

The gather step (see prepare_jobs): validates that every job of the prepared plan produced a matching partial — a half-finished run errors instead of returning a silently incomplete table — rebuilds the per-ROI result list in global ROI order, and runs the single declared join (with_join, ConcatJoin otherwise) exactly once. The final table matches a serial measure row for row; the result list is the normalized form described in measure, roi_index/roi_name included. On success the partials group is removed. Nothing is registered: the returned table is yours to store with add_table.

Raises on a for_job slice (the gather is global) and when no partials exist (nothing was prepared or banked).

Source code in src/ngio/iterators/_feature.py
def finalize(self) -> Table:
    """Merge a distributed run's partials into the one final table.

    The gather step (see `prepare_jobs`): validates that every job of
    the prepared plan produced a matching partial — a half-finished run
    errors instead of returning a silently incomplete table — rebuilds
    the per-ROI result list in global ROI order, and runs the single
    declared join (`with_join`, `ConcatJoin` otherwise) exactly once.
    The final table matches a serial `measure` row for row; the result
    list is the normalized form described in `measure`,
    `roi_index`/`roi_name` included. On success the partials group is
    removed. Nothing is registered: the returned table is yours to
    store with `add_table`.

    Raises on a `for_job` slice (the gather is global) and when no
    partials exist (nothing was prepared or banked).
    """
    self._require_unrestricted_finalize()
    handler = self._partials_handler()
    merged = merge_partial_frames(self, handler, job_verb="measure")

    groups: dict[int, pd.DataFrame] = {}
    if merged.frame is not None:
        for index, group in merged.frame.groupby(INDEX_COLUMN, sort=True):
            frame = group.drop(columns=[INDEX_COLUMN]).reset_index(drop=True)
            record = merged.columns.get(int(cast("Any", index)))
            if record is not None:
                # The concat unioned columns across ROIs (NaN-filled,
                # ints upcast); restore this ROI's own schema exactly.
                columns, dtypes = record
                frame = frame[columns].astype(dtypes)
            else:
                # Partial from an older ngio with no schema record: pin
                # the provenance columns back to the end, where a serial
                # normalization puts them.
                measured = [
                    column
                    for column in frame.columns
                    if column not in (ROI_INDEX_COLUMN, ROI_NAME_COLUMN)
                ]
                frame = frame[[*measured, ROI_INDEX_COLUMN, ROI_NAME_COLUMN]]
            groups[int(cast("Any", index))] = frame
    results: list[pd.DataFrame] = [
        groups.get(index, pd.DataFrame()) for index in range(len(self.rois))
    ]
    table = self._resolved_join()(results)
    delete_partials_root(handler)
    return table

measure

measure(
    func: Callable[
        [ndarray, ndarray, Roi], FeatureFuncResult
    ],
    *,
    mapper: MapperProtocol | None = None,
) -> Table | None

Measure every ROI and join the results into one table.

The per-ROI measurement fans out exactly like reduce — pass a mapper to parallelize it — and the declared join (with_join, ConcatJoin otherwise) runs once, at the end, on the calling thread. Nothing is written: the returned table is yours to store, e.g. container.add_table(name, table).

The join (declared or default, serial or distributed) always sees the normalized result list, not the function's raw returns: one DataFrame per ROI with label as a column (a label index is reset), every row stamped with roi_index (the ROI's global index) and roi_name (the ROI's name — roi_{index} when it has none, as seen by the job that measured it), and a ROI whose function returned no rows as an empty, column-less DataFrame. The three names _ngio_index, roi_index and roi_name are reserved: a function whose result carries one is refused.

With with_halo(...) each region is read grown — both patches and the roi argument cover the grown region — so a border object is measured by every region that sees it. The default join keeps the resulting duplicate label rows as-is; reconcile them in a declared join via the roi_index/roi_name columns. reduce and iter read the grown regions too.

On a for_job slice it measures only this job's share and banks the normalized records as a partial, returning None — stored before any join, so the gather (finalize(), once, after all jobs) can rebuild the full per-ROI result list and run the ONE global join. Re-running a job overwrites its own partial.

Parameters:

  • func (Callable[[ndarray, ndarray, Roi], FeatureFuncResult]) –

    (image, label, roi) -> DataFrame | dict[str, list] — the measurements for one ROI. Rows must carry the object id in a label column (or index). Under a parallel mapper it runs on worker threads or processes and must be safe there.

  • mapper (MapperProtocol | None, default: None ) –

    How the per-ROI work is scheduled; None is serial.

Returns:

  • Table | None –

    The joined table — a FeatureTable under the default join. A

  • Table | None –

    run that finds zero objects returns an empty table, as detect

  • Table | None –

    does. On a for_job slice: None (the partial is banked).

Source code in src/ngio/iterators/_feature.py
def measure(
    self,
    func: Callable[[np.ndarray, np.ndarray, Roi], FeatureFuncResult],
    *,
    mapper: MapperProtocol | None = None,
) -> Table | None:
    """Measure every ROI and join the results into one table.

    The per-ROI measurement fans out exactly like `reduce` — pass a
    `mapper` to parallelize it — and the declared join (`with_join`,
    `ConcatJoin` otherwise) runs once, at the end, on the calling
    thread. Nothing is written: the returned table is yours to store,
    e.g. `container.add_table(name, table)`.

    The join (declared or default, serial or distributed) always sees
    the *normalized* result list, not the function's raw returns: one
    DataFrame per ROI with `label` as a column (a `label` index is
    reset), every row stamped with `roi_index` (the ROI's global index)
    and `roi_name` (the ROI's name — `roi_{index}` when it has none, as
    seen by the job that measured it), and a ROI whose function
    returned no rows as an empty, column-less DataFrame. The three
    names `_ngio_index`, `roi_index` and `roi_name` are reserved: a
    function whose result carries one is refused.

    With `with_halo(...)` each region is read grown — both patches and
    the `roi` argument cover the grown region — so a border object is
    measured by every region that sees it. The default join keeps the
    resulting duplicate label rows as-is; reconcile them in a declared
    join via the `roi_index`/`roi_name` columns. `reduce` and `iter`
    read the grown regions too.

    On a `for_job` slice it measures only this job's share and banks the
    normalized records as a partial, returning `None` — stored
    **before** any join, so the gather (`finalize()`, once, after all
    jobs) can rebuild the full per-ROI result list and run the ONE
    global join. Re-running a job overwrites its own partial.

    Args:
        func: `(image, label, roi) -> DataFrame | dict[str, list]` — the
            measurements for one ROI. Rows must carry the object id in a
            `label` column (or index). Under a parallel mapper it runs on
            worker threads or processes and must be safe there.
        mapper: How the per-ROI work is scheduled; `None` is serial.

    Returns:
        The joined table — a `FeatureTable` under the default join. A
        run that finds zero objects returns an empty table, as `detect`
        does. On a `for_job` slice: `None` (the partial is banked).
    """
    if self._partition is not None:
        self._bank_partial(func, mapper=mapper)
        return None
    results = self.reduce(_UnpackedFeatureFunc(func), mapper=mapper)
    normalized = [
        _normalized_frame(result, index=index, roi=roi)
        for index, (roi, result) in enumerate(zip(self.rois, results, strict=True))
    ]
    return self._resolved_join()(normalized)

ObjectDetectionIterator

ngio.iterators.ObjectDetectionIterator

ObjectDetectionIterator(
    input_image: Image,
    *,
    channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol]
    | None = None,
)

Bases: AbstractIteratorBuilder[NumpyPipeType, DaskPipeType, RoiTable]

Tile an image, detect per tile, return one deduplicated ROI table.

A read-only iterator: nothing is written to the image, and the halo is a pure read margin — call with_halo(...) so an object cut by a tile edge is seen whole by the neighbouring tile; the duplicate detections that produces are resolved by NMS (declared with with_nms, the GreedyNms defaults otherwise). The product is a RoiTable of world-anchored boxes, returned by detect for the caller to store. Distributed, a for_job slice's detect banks a partial and finalize() runs the one global NMS.

The box contract: each detection is a Roi in the patch's own pixel coordinates — Roi.from_values(slices={"x": (x0, w), "y": (y0, h)}, name=None, space="pixel", confidence=0.9). x and y required, z optional (the same choice on every box); name/label are refused (the iterator assigns them — put a class label in an extra field); every extra field rides along into the table.

Initialize the iterator over input_image.

Parameters:

  • input_image (Image) –

    The image the detector runs on.

  • channel_selection (ChannelSlicingInputType, default: None ) –

    Optional selection of channels to read.

  • axes_order (Sequence[str] | None, default: None ) –

    Optional axes order for the patches handed to the detector.

  • input_transforms (Sequence[TransformProtocol] | None, default: None ) –

    Optional transforms applied to each patch.

Source code in src/ngio/iterators/_object_detection.py
def __init__(
    self,
    input_image: Image,
    *,
    channel_selection: ChannelSlicingInputType = None,
    axes_order: Sequence[str] | None = None,
    input_transforms: Sequence[TransformProtocol] | None = None,
) -> None:
    """Initialize the iterator over `input_image`.

    Args:
        input_image: The image the detector runs on.
        channel_selection: Optional selection of channels to read.
        axes_order: Optional axes order for the patches handed to the
            detector.
        input_transforms: Optional transforms applied to each patch.
    """
    self._input = input_image
    self._ref_image = input_image
    self._rois = input_image.build_image_roi_table(name=None).rois()

    self._input_slicing_kwargs = add_channel_selection_to_slicing_dict(
        image=self._input, channel_selection=channel_selection, slicing_dict={}
    )
    self._channel_selection = channel_selection
    self._axes_order = axes_order
    self._input_transforms = input_transforms

rois property

rois: list[Roi]

Get the list of ROIs for the iterator.

ref_image property

ref_image: AbstractImage

Get the reference image for the iterator.

output_image property

output_image: AbstractImage | None

The image this iterator writes to, or None for a read-only iterator.

halo property

halo: dict[str, int]

The per-axis read margin, empty when there is none.

partition_indices property

partition_indices: list[int] | None

The unit indices this iterator is restricted to, or None.

None on an unrestricted iterator; on a for_job slice, the sorted positions (in the full ROI list) of the units it runs. The whole layout of a split is [it.for_job(i, n).partition_indices for i in range(n)] — the pre-flight check before submitting a distributed run: one fat list plus empties means the output chunking, not the cluster, is the constraint.

with_halo

with_halo(
    x: int = 0, y: int = 0, z: int = 0, t: int = 0
) -> Self

Read a margin of context around each ROI, but write only the ROI.

The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.

Note

The margins do not survive a rescale: a shape-changing transform on the data path (a ZoomTransform) makes the crop refuse — iterate on a coarser pyramid level instead. A MaskTransform composes normally (its zoom rescales the label it reads).

Parameters:

  • x (int, default: 0 ) –

    Pixels added on each side along x.

  • y (int, default: 0 ) –

    Pixels added on each side along y.

  • z (int, default: 0 ) –

    Pixels added on each side along z.

  • t (int, default: 0 ) –

    Frames added on each side along t.

Returns:

  • Self –

    A new iterator reading with the halo.

Raises:

  • NgioValueError –

    On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps roi_index/roi_name so your coalesce can), on an in-place iterator, or for a negative margin.

Example
# 8 px of context per side, disjoint writes, still parallel.
it = iterator.by_write_units().with_halo(x=8, y=8)
it.map(smooth, mapper=ThreadedMapper("auto"))
Source code in src/ngio/iterators/_abstract_iterator.py
def with_halo(self, x: int = 0, y: int = 0, z: int = 0, t: int = 0) -> Self:
    """Read a margin of context around each ROI, but write only the ROI.

    The function receives the grown region and must return it grown too;
    the border is cropped before the write, so tiles come out seamless.
    Margins are reference-image pixels, clipped at the borders. Write
    footprints are unchanged, so a haloed iterator parallelizes exactly
    as far as it did without one.

    Note:
        The margins do not survive a rescale: a shape-changing transform
        on the data path (a `ZoomTransform`) makes the crop refuse —
        iterate on a coarser pyramid level instead. A `MaskTransform`
        composes normally (its zoom rescales the label it reads).

    Args:
        x: Pixels added on each side along x.
        y: Pixels added on each side along y.
        z: Pixels added on each side along z.
        t: Frames added on each side along t.

    Returns:
        A new iterator reading with the halo.

    Raises:
        NgioValueError: On a read-only iterator that has not opted in
            (no write to crop the halo from; detection opts in and
            reconciles the overlap with NMS, feature extraction opts in
            and stamps `roi_index`/`roi_name` so your `coalesce` can),
            on an in-place iterator, or for a negative margin.

    Example:
        ```python
        # 8 px of context per side, disjoint writes, still parallel.
        it = iterator.by_write_units().with_halo(x=8, y=8)
        it.map(smooth, mapper=ThreadedMapper("auto"))
        ```
    """
    halo = {"x": x, "y": y, "z": z, "t": t}
    for axis_name, margin in halo.items():
        if margin < 0:
            raise NgioValueError(
                f"Halo along '{axis_name}' must be >= 0, got {margin}."
            )
    if self.output_image is None and not self._allow_readonly_halo:
        name = self.__class__.__name__
        raise NgioValueError(
            f"{name} is read-only, so there is no written region for a "
            "halo to surround. Widen the ROIs themselves instead (e.g. "
            "`by_grid(...)` with a stride smaller than the size)."
        )
    output = self.output_image
    if output is not None and _is_same_zarr_array(
        self._ref_image.zarr_array, output.zarr_array
    ):
        raise NgioValueError(
            "In-place iteration cannot take a halo: each tile's grown "
            "read covers pixels that neighbouring tiles write, so the "
            "result depends on execution order. Write to a different "
            "image or label, or drop the halo."
        )
    new_instance = self._new_from_rois(self.rois)
    new_instance._halo = {ax: m for ax, m in halo.items() if m > 0}
    return new_instance

by_grid

by_grid(
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self

Tile the current ROIs with a regular grid.

Parameters:

  • size_x (int | None, default: None ) –

    Tile size along x; None (default) means the full extent.

  • size_y (int | None, default: None ) –

    Tile size along y; None means the full extent.

  • size_z (int | None, default: None ) –

    Tile size along z; None means the full extent.

  • size_t (int | None, default: None ) –

    Tile size along t; None means the full extent.

  • stride_x (int | None, default: None ) –

    Step between tiles along x; None means size_x (adjacent tiles). A smaller stride produces overlapping tiles.

  • stride_y (int | None, default: None ) –

    As stride_x, along y.

  • stride_z (int | None, default: None ) –

    As stride_x, along z.

  • stride_t (int | None, default: None ) –

    As stride_x, along t.

  • tail (TailPolicy, default: 'clip' ) –

    What happens where the grid does not divide the axis. "clip" (default) shrinks the last tile to the border; "balance" re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles, stride == size); "shift" moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution; "drop" discards the tail and leaves that border uncovered.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_grid(
    self,
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self:
    """Tile the current ROIs with a regular grid.

    Args:
        size_x: Tile size along x; `None` (default) means the full extent.
        size_y: Tile size along y; `None` means the full extent.
        size_z: Tile size along z; `None` means the full extent.
        size_t: Tile size along t; `None` means the full extent.
        stride_x: Step between tiles along x; `None` means `size_x`
            (adjacent tiles). A smaller stride produces overlapping tiles.
        stride_y: As `stride_x`, along y.
        stride_z: As `stride_x`, along z.
        stride_t: As `stride_x`, along t.
        tail: What happens where the grid does not divide the axis.
            `"clip"` (default) shrinks the last tile to the border;
            `"balance"` re-splits the last full tile and the tail into two
            near-equal tiles, so a thin overhang never yields a thin tile
            (needs adjacent tiles, `stride == size`); `"shift"` moves the
            last tile back to stay full-size — it then overlaps its
            neighbour, so a writing iterator needs a merge or serial
            execution; `"drop"` discards the tail and leaves that border
            uncovered.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the tiled ROIs.
    """
    rois = by_grid(
        rois=self.rois,
        ref_image=self.ref_image,
        size_x=size_x,
        size_y=size_y,
        size_z=size_z,
        size_t=size_t,
        stride_x=stride_x,
        stride_y=stride_y,
        stride_z=stride_z,
        stride_t=stride_t,
        tail=tail,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_blocks

by_blocks(
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self

Divide each axis into a fixed number of near-equal blocks.

The complement of by_grid: you say how many tiles, not how big. Block lengths along an axis differ by at most one pixel, so there is no tail to police — the partition is balanced by construction.

Parameters:

  • num_x (int, default: 1 ) –

    Number of blocks along x (default 1: no split).

  • num_y (int, default: 1 ) –

    Number of blocks along y.

  • num_z (int, default: 1 ) –

    Number of blocks along z.

  • num_t (int, default: 1 ) –

    Number of blocks along t.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the blocks.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_blocks(
    self,
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self:
    """Divide each axis into a fixed number of near-equal blocks.

    The complement of `by_grid`: you say how many tiles, not how big.
    Block lengths along an axis differ by at most one pixel, so there is
    no tail to police — the partition is balanced by construction.

    Args:
        num_x: Number of blocks along x (default 1: no split).
        num_y: Number of blocks along y.
        num_z: Number of blocks along z.
        num_t: Number of blocks along t.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the blocks.
    """
    rois = by_blocks(
        rois=self.rois,
        ref_image=self.ref_image,
        num_x=num_x,
        num_y=num_y,
        num_z=num_z,
        num_t=num_t,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_yx

by_yx() -> Self

Split each ROI into one 2D plane per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, z, c) combination, each covering the full y/x extent — the shape for "run this 2D function on every plane".

Source code in src/ngio/iterators/_abstract_iterator.py
def by_yx(self) -> Self:
    """Split each ROI into one 2D plane per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, z, c)` combination, each covering the full y/x extent — the
    shape for "run this 2D function on every plane".
    """
    rois = by_yx(self.rois, self.ref_image)
    return self._new_from_rois(rois)

by_zyx

by_zyx(strict: bool = True) -> Self

Split each ROI into one 3D volume per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, c) combination, each covering the full z/y/x extent.

Parameters:

  • strict (bool, default: True ) –

    Require a real z axis (present and larger than 1); with False, images without one fall back to 2D planes.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_zyx(self, strict: bool = True) -> Self:
    """Split each ROI into one 3D volume per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, c)` combination, each covering the full z/y/x extent.

    Args:
        strict: Require a real z axis (present and larger than 1);
            with `False`, images without one fall back to 2D planes.
    """
    rois = by_zyx(self.rois, self.ref_image, strict=strict)
    return self._new_from_rois(rois)

by_chunks

by_chunks(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the input image's chunk grid.

Reads are chunk-granular, so this is the natural tiling for read-heavy work. For collision-free parallel writes use by_write_units, which tiles by the output's write granularity.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_chunks(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the input image's chunk grid.

    Reads are chunk-granular, so this is the natural tiling for
    read-heavy work. For collision-free parallel *writes* use
    `by_write_units`, which tiles by the output's write granularity.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=None,
    )
    return self._new_from_rois(rois)

by_write_units

by_write_units(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the output image's write granularity.

The write unit is the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. With no overlap the resulting ROIs pass check_if_write_units_overlap by construction, so a parallel map runs as a single fully-parallel wave — a throughput optimization; conflicting tilings are still safe, just scheduled in more waves. Falls back to the input chunk grid when the iterator is read-only.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_write_units(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the output image's write granularity.

    The write unit is the shard shape when the output is sharded (writes
    are atomic per shard object), the chunk shape otherwise. With no
    overlap the resulting ROIs pass `check_if_write_units_overlap` by
    construction, so a parallel `map` runs as a single fully-parallel
    wave — a throughput optimization; conflicting tilings are still safe,
    just scheduled in more waves. Falls back to the input chunk grid when
    the iterator is read-only.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=self.output_image,
    )
    return self._new_from_rois(rois)

product

product(other: list[Roi] | GenericRoiTable) -> Self

Cartesian product of the current ROIs with an arbitrary list of ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def product(self, other: list[Roi] | GenericRoiTable) -> Self:
    """Cartesian product of the current ROIs with an arbitrary list of ROIs."""
    if isinstance(other, GenericRoiTable):
        other = other.rois()
    rois = rois_product(self.rois, other)
    return self._new_from_rois(rois)

iter

iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
    lazy: Literal[False] = ...,
    data_mode: Literal["numpy"] | None = ...,
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: int,
) -> Generator[list[NumpyPipeType]]
iter(
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal[
        "readwrite", "readonly"
    ] = "readonly",
    *,
    batch_size: int | None = None,
) -> Generator

Iterate the ROIs' payloads, read-only.

With batch_size set, yields payload lists of up to that many items (the last batch may be smaller). Batching is numpy-only and eager: combining batch_size with lazy=True or data_mode="dask" raises. On a writing iterator (see WritingIteratorBuilder.iter) the default mode is "readwrite" and the loop yields (patch, writer) pairs instead.

Parameters:

  • lazy (bool, default: False ) –

    Yield getter handles instead of materialized data.

  • data_mode (Literal['numpy', 'dask'] | None, default: None ) –

    "numpy" or "dask" (deprecated, see note).

  • iterator_mode (Literal['readwrite', 'readonly'], default: 'readonly' ) –

    "readonly" yields patches alone; on a writing iterator "readwrite" yields (patch, writer) pairs.

  • batch_size (int | None, default: None ) –

    When set, yield batches of up to this many items; the last batch may be smaller.

Note

The dask data mode is deprecated and will be removed in ngio=1.2; from then on iter() yields numpy arrays.

Source code in src/ngio/iterators/_abstract_iterator.py
def iter(
    self,
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal["readwrite", "readonly"] = "readonly",
    *,
    batch_size: int | None = None,
) -> Generator:
    """Iterate the ROIs' payloads, read-only.

    With `batch_size` set, yields payload lists of up to that many
    items (the last batch may be smaller). Batching is numpy-only and
    eager: combining `batch_size` with `lazy=True` or
    `data_mode="dask"` raises. On a writing iterator (see
    `WritingIteratorBuilder.iter`) the default mode is `"readwrite"`
    and the loop yields `(patch, writer)` pairs instead.

    Args:
        lazy: Yield getter handles instead of materialized data.
        data_mode: `"numpy"` or `"dask"` (deprecated, see note).
        iterator_mode: `"readonly"` yields patches alone; on a writing
            iterator `"readwrite"` yields `(patch, writer)` pairs.
        batch_size: When set, yield batches of up to this many items;
            the last batch may be smaller.

    Note:
        The dask data mode is deprecated and will be removed in ngio=1.2;
        from then on `iter()` yields numpy arrays.
    """
    return self._iter_impl(
        lazy=lazy,
        data_mode=data_mode,
        iterator_mode=iterator_mode,
        batch_size=batch_size,
    )

iter_as_numpy

iter_as_numpy()

Alias for iter(data_mode="numpy").

Source code in src/ngio/iterators/_abstract_iterator.py
def iter_as_numpy(
    self,
):
    """Alias for `iter(data_mode="numpy")`."""
    return self._iter(lazy=False, data_mode="numpy", iterator_mode="readonly")

iter_as_dask

iter_as_dask()

Alias for iter(data_mode="dask").

Deprecated: removed in ngio=1.2.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="iter_as_numpy() (or Image.get_as_dask() for a lazy array)",
    removed_in="1.2",
)
def iter_as_dask(
    self,
):
    """Alias for `iter(data_mode="dask")`.

    Deprecated: removed in ngio=1.2.
    """
    return self._iter(lazy=False, data_mode="dask", iterator_mode="readonly")

for_job

for_job(job_index: int, n_jobs: int) -> Self

Restrict the iterator to one of n_jobs independent partitions.

No write unit is ever shared between two partitions, so the jobs need no locks, no barriers, and no channel to one another — safe as separate SLURM array tasks, in any order. Every job must build an identical iterator (construction is metadata-only and the partition derives deterministically from it), apply for_job last in the builder chain, and its verbs deliberately do not finalize — a slice's map/process/segment writes only its share, and measure/ detect bank a partial instead of joining. The gather step is the unrestricted iterator's finalize(), once, after all jobs. Stitched and read-only runs also need prepare_jobs(n_jobs) first. See the distributed-runs guide in the iterators docs.

Parameters:

  • job_index (int) –

    This job's partition, 0 <= job_index < n_jobs (e.g. int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops.

  • n_jobs (int) –

    How many partitions the run is split into; must match across every job and the gather.

Returns:

  • Self –

    The restricted iterator; further reshaping refuses.

Example
# array task:
iterator = SegmentationIterator(image, label).by_chunks()
iterator.for_job(job_index, n_jobs=n_jobs).map(func)

# gather task, after all jobs:
SegmentationIterator(image, label).by_chunks().finalize()
Source code in src/ngio/iterators/_abstract_iterator.py
def for_job(self, job_index: int, n_jobs: int) -> Self:
    """Restrict the iterator to one of `n_jobs` independent partitions.

    No write unit is ever shared between two partitions, so the jobs
    need no locks, no barriers, and no channel to one another — safe as
    separate SLURM array tasks, in any order. Every job must build an
    identical iterator (construction is metadata-only and the partition
    derives deterministically from it), apply `for_job` *last* in the
    builder chain, and its verbs deliberately do not finalize — a slice's
    `map`/`process`/`segment` writes only its share, and `measure`/
    `detect` bank a partial instead of joining. The gather step is the
    unrestricted iterator's `finalize()`, once, after all jobs.
    Stitched and read-only runs also need `prepare_jobs(n_jobs)` first.
    See the distributed-runs guide in the iterators docs.

    Args:
        job_index: This job's partition, `0 <= job_index < n_jobs`
            (e.g. `int($SLURM_ARRAY_TASK_ID)`). Partitions beyond the
            number of independent unit groups are empty no-ops.
        n_jobs: How many partitions the run is split into; must match
            across every job and the gather.

    Returns:
        The restricted iterator; further reshaping refuses.

    Example:
        ```python
        # array task:
        iterator = SegmentationIterator(image, label).by_chunks()
        iterator.for_job(job_index, n_jobs=n_jobs).map(func)

        # gather task, after all jobs:
        SegmentationIterator(image, label).by_chunks().finalize()
        ```
    """
    if self._partition is not None:
        _, current_index, current_n = self._partition
        raise NgioValueError(
            f"This iterator is already restricted to partition "
            f"{current_index}/{current_n}; partitions do not nest. "
            "Apply `for_job` once, to the full iterator."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    if not 0 <= job_index < n_jobs:
        raise NgioValueError(
            f"job_index must be in [0, {n_jobs}), got {job_index}."
        )
    self._require_splittable()
    self._validate_job_context(n_jobs)
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    restricted = self._new_from_rois(self.rois)
    restricted._partition = (
        frozenset(partition[job_index]),
        job_index,
        n_jobs,
    )
    return restricted

prepare_jobs

prepare_jobs(n_jobs: int) -> list[JobArgs]

Set up a distributed run and return its parallelization list.

The init step of the three-phase recipe (init → parallel jobs → consolidate): performs the iterator's setup — always clearing stale scratch state from earlier runs first — and returns one JobArgs per non-empty partition. Optional for plain writers; a stitching or read-only run requires it (the scratch root must exist, race-free, before any job writes into it).

Parameters:

  • n_jobs (int) –

    How many partitions to split into; empty ones are dropped from the list.

Returns:

  • list[JobArgs] –

    One {"job_index": ..., "n_jobs": ...} dict per non-empty

  • list[JobArgs] –

    partition — JSON-ready, splatting into for_job(**args).

Example
args_list = iterator.prepare_jobs(n_jobs=4)  # init task
iterator.for_job(**args).map(func)  # per parallel task
iterator.finalize()  # consolidate task
Source code in src/ngio/iterators/_abstract_iterator.py
def prepare_jobs(self, n_jobs: int) -> list[JobArgs]:
    """Set up a distributed run and return its parallelization list.

    The init step of the three-phase recipe (init → parallel jobs →
    consolidate): performs the iterator's setup — always clearing stale
    scratch state from earlier runs first — and returns one `JobArgs`
    per non-empty partition. Optional for plain writers; a stitching or
    read-only run requires it (the scratch root must exist, race-free,
    before any job writes into it).

    Args:
        n_jobs: How many partitions to split into; empty ones are
            dropped from the list.

    Returns:
        One `{"job_index": ..., "n_jobs": ...}` dict per non-empty
        partition — JSON-ready, splatting into `for_job(**args)`.

    Example:
        ```python
        args_list = iterator.prepare_jobs(n_jobs=4)  # init task
        iterator.for_job(**args).map(func)  # per parallel task
        iterator.finalize()  # consolidate task
        ```
    """
    if self._partition is not None:
        raise NgioValueError(
            "prepare_jobs belongs to the unrestricted iterator; this one "
            "is already restricted to a partition."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    self._require_splittable()
    if self.output_image is not None:
        self._validate_write_plan()
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    self._prepare_distributed(n_jobs)
    return [
        JobArgs(job_index=job_index, n_jobs=n_jobs)
        for job_index, indices in enumerate(partition)
        if indices
    ]

reduce

reduce(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Units are built read-only even on writable iterators: nothing is written and finalize does not run.

Parameters:

  • func (Callable[[NumpyPipeType], R]) –

    The function to apply; under a parallel mapper it must be safe on worker threads (or processes).

  • mapper (MapperProtocol[NumpyPipeType, R] | None, default: None ) –

    How the units are scheduled. None (the default) is serial; pass ThreadedMapper() or ProcessMapper() to fan out, sized by their own max_workers argument.

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run.

    Args:
        func: The function to apply; under a parallel mapper it must be
            safe on worker threads (or processes).
        mapper: How the units are scheduled. `None` (the default) is
            serial; pass `ThreadedMapper()` or `ProcessMapper()` to fan
            out, sized by their own `max_workers` argument.

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    if mapper is None:
        mapper = BasicMapper[NumpyPipeType, R]()
    return mapper(func, list(self._numpy_units_generator(with_setters=False)))

reduce_as_numpy

reduce_as_numpy(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Alias for reduce().

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce_as_numpy(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Alias for `reduce()`."""
    return self.reduce(func, mapper=mapper)

reduce_as_dask

reduce_as_dask(
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Deprecated: removed in ngio=1.2.

Units are built read-only even on writable iterators: nothing is written and finalize does not run. Runs serially — the units hand func lazy dask arrays, so the per-unit work is graph construction, not IO. A parallel mapper raises (see map_as_dask).

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="reduce() / reduce_as_numpy(mapper=...)",
    removed_in="1.2",
)
def reduce_as_dask(
    self,
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Deprecated: removed in ngio=1.2.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run. Runs serially — the units hand
    `func` *lazy* dask arrays, so the per-unit work is graph
    construction, not IO. A parallel `mapper` raises (see `map_as_dask`).

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    _mapper = self._require_serial_dask_mapper(mapper)
    return _mapper(func, list(self._dask_units_generator(with_setters=False)))

check_if_regions_overlap

check_if_regions_overlap() -> bool

Check if any of the ROIs overlap logically.

If two ROIs cover the same pixel, they are considered to overlap. This does not consider chunking or other storage details. Measured on the read regions — with a halo, neighbouring reads overlap by construction. For write safety, check_if_write_units_overlap is the right question.

Returns:

  • bool –

    True if any ROIs overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_regions_overlap(self) -> bool:
    """Check if any of the ROIs overlap logically.

    If two ROIs cover the same pixel, they are considered to overlap.
    This does not consider chunking or other storage details. Measured
    on the *read* regions — with a halo, neighbouring reads overlap by
    construction. For write safety, `check_if_write_units_overlap` is
    the right question.

    Returns:
        `True` if any ROIs overlap.
    """
    if len(self.rois) < 2:
        return False

    slicing_tuples = (
        g.slicing_ops.normalized_slicing_tuple
        for g in self._numpy_getters_generator()
    )
    return check_if_regions_overlap(slicing_tuples)

require_no_regions_overlap

require_no_regions_overlap() -> None

Ensure that the Iterator's ROIs do not overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_regions_overlap(self) -> None:
    """Ensure that the Iterator's ROIs do not overlap."""
    if self.check_if_regions_overlap():
        raise NgioValueError("Some rois overlap.")

with_nms

with_nms(nms: NmsProtocol) -> Self

Declare how duplicate detections are suppressed.

The default is GreedyNms(). The jobs of a distributed run read score_column/max_detections_per_tile at banking, so declare the identical NMS on every job and the gather — like the detection function, it is not part of the plan fingerprint. Declare before for_job.

Parameters:

  • nms (NmsProtocol) –

    Anything satisfying NmsProtocol — GreedyNms(...), or your own suppression.

Source code in src/ngio/iterators/_object_detection.py
def with_nms(self, nms: NmsProtocol) -> Self:
    """Declare how duplicate detections are suppressed.

    The default is `GreedyNms()`. The jobs of a distributed run read
    `score_column`/`max_detections_per_tile` at banking, so declare the
    identical NMS on every job and the gather — like the detection
    function, it is not part of the plan fingerprint. Declare before
    `for_job`.

    Args:
        nms: Anything satisfying `NmsProtocol` — `GreedyNms(...)`, or
            your own suppression.
    """
    new_instance = self._new_from_rois(self.rois)
    new_instance._nms = nms
    return new_instance

build_numpy_getter

build_numpy_getter(roi: Roi) -> DetectionGetter[ndarray]
Source code in src/ngio/iterators/_object_detection.py
def build_numpy_getter(self, roi: Roi) -> DetectionGetter[np.ndarray]:
    return DetectionGetter(
        NumpyGetter(
            zarr_array=self._input.zarr_array,
            dimensions=self._input.dimensions,
            roi=self._read_roi(roi),
            axes_order=self._axes_order,
            transforms=self._input_transforms,
            slicing_dict=self._input_slicing_kwargs,
        )
    )

build_dask_getter

build_dask_getter(roi: Roi) -> DetectionGetter[Array]
Source code in src/ngio/iterators/_object_detection.py
def build_dask_getter(self, roi: Roi) -> DetectionGetter[da.Array]:
    return DetectionGetter(
        DaskGetter(
            zarr_array=self._input.zarr_array,
            dimensions=self._input.dimensions,
            roi=self._read_roi(roi),
            axes_order=self._axes_order,
            transforms=self._input_transforms,
            slicing_dict=self._input_slicing_kwargs,
        )
    )

finalize

finalize() -> RoiTable

Merge a distributed run's partials into the one final ROI table.

The gather step (see prepare_jobs): validates that every job of the prepared plan produced a matching partial — a half-finished run errors instead of returning a silently incomplete table — then rebuilds every tile's raw boxes and runs the full serial pipeline once, globally: anchoring, the cross-tile invariant checks, NMS, and the dense 1..N renumbering. Bit-identical to a serial detect by construction — extra fields keep their columns and dtypes through the partial round-trip. Two residuals, both within one tile: an extra a detector sets to NaN does not survive, and an integer extra present on only some of a tile's boxes comes back float (the tile frame NaN-fills it before the schema is recorded). On success the partials group is removed. Nothing is registered: the returned table is yours to store with add_table.

Raises on a for_job slice (the gather is global) and when no partials exist (nothing was prepared or banked).

Source code in src/ngio/iterators/_object_detection.py
def finalize(self) -> RoiTable:
    """Merge a distributed run's partials into the one final ROI table.

    The gather step (see `prepare_jobs`): validates that every job of
    the prepared plan produced a matching partial — a half-finished run
    errors instead of returning a silently incomplete table — then
    rebuilds every tile's raw boxes and runs the full serial pipeline
    once, globally: anchoring, the cross-tile invariant checks, NMS, and
    the dense `1..N` renumbering. Bit-identical to a serial `detect` by
    construction — extra fields keep their columns and dtypes through
    the partial round-trip. Two residuals, both within one tile: an
    extra a detector sets to `NaN` does not survive, and an integer
    extra present on only some of a tile's boxes comes back float (the
    tile frame NaN-fills it before the schema is recorded). On success
    the partials group is removed. Nothing is registered: the returned
    table is yours to store with `add_table`.

    Raises on a `for_job` slice (the gather is global) and when no
    partials exist (nothing was prepared or banked).
    """
    self._require_unrestricted_finalize()
    handler = self._partials_handler()
    merged = merge_partial_frames(self, handler, job_verb="detect")

    results: list[list[Roi]] = [[] for _ in self.rois]
    if merged.frame is not None:
        for index, group in merged.frame.groupby(INDEX_COLUMN, sort=True):
            results[int(cast("Any", index))] = _frame_to_rois(
                group.drop(columns=[INDEX_COLUMN]).reset_index(drop=True),
                columns=merged.columns.get(int(cast("Any", index))),
            )
    detections = self._collect(results)
    kept = self._nms.suppress(detections)
    table = RoiTable(rois=self._renumbered(kept))
    delete_partials_root(handler)
    return table

detect

detect(
    func: Callable[[ndarray], list[Roi]],
    *,
    mapper: MapperProtocol | None = None,
) -> RoiTable | None

Run the detector over every tile and return one table of objects.

The per-tile detection fans out exactly like reduce — pass a mapper to parallelize it — and the reconciliation (coordinates, NMS, renumbering) happens once, at the end, on the calling thread. Nothing is written: the returned table is yours to store, e.g. container.add_table(name, table).

On a for_job slice it detects only over this job's tiles and banks the detector's raw, pre-NMS boxes as a partial, returning None — greedy NMS is not hierarchical, so a per-job NMS followed by a merge NMS could differ from a serial run. The gather (finalize(), once, after all jobs) runs the ONE global NMS and renumbering, bit-identical to a serial detect. The tiles' boxes are validated here, so a malformed detector fails in the job rather than at the gather; re-running a job overwrites its own partial.

Parameters:

  • func (Callable[[ndarray], list[Roi]]) –

    (patch) -> list[Roi], boxes in the patch's own pixel coordinates — see the class docstring for the box contract. Under a parallel mapper it runs on worker threads or processes and must be safe there.

  • mapper (MapperProtocol | None, default: None ) –

    How the per-tile work is scheduled; None is serial.

Returns:

  • RoiTable | None –

    A RoiTable of the surviving detections, ids 1..N in

  • RoiTable | None –

    (tile, row) order, in the reference image's world coordinates.

  • RoiTable | None –

    An image with nothing to detect gives an empty table. On a

  • RoiTable | None –

    for_job slice: None (the partial is banked).

Source code in src/ngio/iterators/_object_detection.py
def detect(
    self,
    func: Callable[[np.ndarray], list[Roi]],
    *,
    mapper: MapperProtocol | None = None,
) -> RoiTable | None:
    """Run the detector over every tile and return one table of objects.

    The per-tile detection fans out exactly like `reduce` — pass a
    `mapper` to parallelize it — and the reconciliation (coordinates, NMS,
    renumbering) happens once, at the end, on the calling thread. Nothing
    is written: the returned table is yours to store, e.g.
    `container.add_table(name, table)`.

    On a `for_job` slice it detects only over this job's tiles and banks
    the detector's raw, **pre-NMS** boxes as a partial, returning `None`
    — greedy NMS is not hierarchical, so a per-job NMS followed by a
    merge NMS could differ from a serial run. The gather (`finalize()`,
    once, after all jobs) runs the ONE global NMS and renumbering,
    bit-identical to a serial `detect`. The tiles' boxes are validated
    here, so a malformed detector fails in the job rather than at the
    gather; re-running a job overwrites its own partial.

    Args:
        func: `(patch) -> list[Roi]`, boxes in the *patch's own* pixel
            coordinates — see the class docstring for the box contract.
            Under a parallel mapper it runs on worker threads or
            processes and must be safe there.
        mapper: How the per-tile work is scheduled; `None` is serial.

    Returns:
        A `RoiTable` of the surviving detections, ids `1..N` in
        (tile, row) order, in the reference image's world coordinates.
        An image with nothing to detect gives an empty table. On a
        `for_job` slice: `None` (the partial is banked).
    """
    if self._partition is not None:
        self._bank_partial(func, mapper=mapper)
        return None
    results = self.reduce(_UnpackedDetectionFunc(func), mapper=mapper)
    detections = self._collect(results)
    kept = self._nms.suppress(detections)
    return RoiTable(rois=self._renumbered(kept))

Reconciliation declarations

Each iterator declares its reconciliation in the builder chain — on_overlap, with_stitch, with_join, with_nms — backed by a swappable protocol with a shipped default.

StitchConfig

ngio.iterators.StitchConfig dataclass

StitchConfig(
    *,
    block_size: int = 10000,
    iou_threshold: float = 0.3,
    seam_matcher: SeamMatcherProtocol | None = None,
    compact: bool = True,
    scratch_store: StoreOrGroup | None = None,
    write_order: WriteOrder = "any",
)

How to stitch a tiled segmentation.

Attributes:

  • block_size (int) –

    Ids reserved per tile. Must exceed the largest label a single tile can produce, or ids spill into the next tile's block.

  • iou_threshold (float) –

    How much two tiles' predictions must agree over their shared pixels before their ids are called the same object. Low values merge eagerly; the default errs towards leaving an object split, which is fixable downstream, rather than merging two that are not. Configures the default IouSeamMatcher; ignored when seam_matcher is set.

  • seam_matcher (SeamMatcherProtocol | None) –

    Replaces the default pair criterion — see SeamMatcherProtocol. Like the measurement function, a custom matcher is not part of the plan fingerprint: declare the identical one on every phase of a distributed run.

  • compact (bool) –

    Renumber the surviving ids to a dense 1..N at the end. The renumbering walks the whole label, so objects outside the iterated ROIs get new ids too — anything keyed by the old ids (feature tables, masking ROI tables) is invalidated. A run that renumbers such pre-existing objects warns. Use compact=False and reconcile ids yourself to keep external references valid.

  • scratch_store (StoreOrGroup | None) –

    Where to keep the per-tile banks while the map runs. None puts them in a transient group inside the output label, which works under every mapper. A MemoryStore avoids touching the output store — but an in-memory scratch cannot cross a process boundary, so ProcessMapper refuses it.

  • write_order (WriteOrder) –

    Who owns the seams of unmatched objects — object unification is exact either way. "any" (the default) is schedule-defined; "roi" gives them to the later tile, reproducibly, at up to 2.5x cost on parallel overlapping tilings.

block_size class-attribute instance-attribute

block_size: int = 10000

iou_threshold class-attribute instance-attribute

iou_threshold: float = 0.3

seam_matcher class-attribute instance-attribute

seam_matcher: SeamMatcherProtocol | None = None

compact class-attribute instance-attribute

compact: bool = True

scratch_store class-attribute instance-attribute

scratch_store: StoreOrGroup | None = None

write_order class-attribute instance-attribute

write_order: WriteOrder = 'any'

matcher

matcher() -> SeamMatcherProtocol

The seam matcher the resolve runs; the IoU default unless swapped.

Source code in src/ngio/iterators/_stitch.py
def matcher(self) -> SeamMatcherProtocol:
    """The seam matcher the resolve runs; the IoU default unless swapped."""
    return self.seam_matcher or IouSeamMatcher(self.iou_threshold)

SeamMatcherProtocol

ngio.iterators.SeamMatcherProtocol

Bases: Protocol

Decides which ids of two overlapping banked predictions are one object.

Called only at the resolve (the gather step, on the calling thread — no pickling constraints), once per overlapping tile pair, with the two tiles' banked label patches reconstructed over their shared region: aligned, same shape, in the label's own axes order. Ids are as they appear in the patches (each tile's block-offset ids). Return the (id_in_a, id_in_b) pairs that are the same object; the framework owns which tile pairs are compared and over which regions, the banking and scratch lifecycle, the union-find, and the global relabel.

Must be deterministic — a pure function of the two patches — or a distributed gather is no longer bit-identical to a serial run. Background 0 is never pairable, and every returned id must appear in its patch; violations raise at the resolve.

IouSeamMatcher

ngio.iterators.IouSeamMatcher dataclass

IouSeamMatcher(iou_threshold: float = 0.3)

The default seam criterion: IoU over the shared pixels, thresholded.

overlap_iou counts only pixels labelled on both sides, so background (and a masked tile's outside) never enters a pair and can never displace a written id.

iou_threshold class-attribute instance-attribute

iou_threshold: float = 0.3

NmsProtocol

ngio.iterators.NmsProtocol

Bases: Protocol

What the detection iterator needs from its NMS declaration.

Two properties the framework reads outside suppression, plus the suppression itself. suppress runs only at the join (a serial detect, or the distributed gather's finalize), on the calling thread — no pickling constraints. It must be deterministic — a pure function of the detection list, ties broken deterministically (the default uses (-score, tile_index, row)) — or a distributed gather is no longer bit-identical to a serial run.

Compatibility policy: future ngio versions will only ever add optional members, probed with getattr.

score_column property

score_column: str

Extra field ranking the detections; also names the output column.

max_detections_per_tile property

max_detections_per_tile: int

Per-tile runaway-detector guard, enforced at collection.

suppress

suppress(
    detections: Sequence[Detection],
) -> list[Detection]

Return the survivors — a subset of detections, any order.

Source code in src/ngio/iterators/_object_detection.py
def suppress(self, detections: Sequence[Detection]) -> list[Detection]:
    """Return the survivors — a subset of `detections`, any order."""
    ...

GreedyNms

ngio.iterators.GreedyNms dataclass

GreedyNms(
    *,
    iou_threshold: float = 0.5,
    score_column: str = "confidence",
    max_detections_per_tile: int = 10000,
)

Greedy NMS, the default suppression: highest score first.

Attributes:

  • iou_threshold (float) –

    Two boxes overlapping at or above this intersection-over-union are the same object; the lower-scoring one is suppressed. Values near 1 keep almost everything, values near 0 merge aggressively; 0 itself would merge any pair that merely touches and is refused.

  • score_column (str) –

    Extra field on the returned Rois ranking the detections (and the resulting table's column), "confidence" by default. When the detector provides no such field the box volume ranks instead — bigger wins.

  • max_detections_per_tile (int) –

    A tile producing more detections than this is a runaway detector; it fails loudly instead of flooding the table.

iou_threshold class-attribute instance-attribute

iou_threshold: float = 0.5

score_column class-attribute instance-attribute

score_column: str = 'confidence'

max_detections_per_tile class-attribute instance-attribute

max_detections_per_tile: int = 10000

suppress

suppress(
    detections: Sequence[Detection],
) -> list[Detection]

Greedy pass: keep by rank, drop what overlaps a kept box.

Source code in src/ngio/iterators/_object_detection.py
def suppress(self, detections: Sequence[Detection]) -> list[Detection]:
    """Greedy pass: keep by rank, drop what overlaps a kept box."""
    ranked = sorted(detections, key=lambda d: (-d.score, d.tile_index, d.row))
    kept: list[Detection] = []
    for detection in ranked:
        duplicate = any(
            other.group == detection.group
            and bbox_iou(other.bounds, detection.bounds) >= self.iou_threshold
            for other in kept
        )
        if not duplicate:
            kept.append(detection)
    return kept

Detection

ngio.iterators.Detection dataclass

Detection(
    *,
    tile_index: int,
    row: int,
    score: float,
    roi: Roi,
    bounds: dict[str, tuple[float, float]],
    group: tuple,
)

One anchored detection, as suppression sees it.

Constructed by the framework only. A custom suppression returns a subset of the instances it was given — never construct or copy these — so future ngio versions stay free to add fields.

Attributes:

  • tile_index (int) –

    Position of the reporting tile in iterator.rois.

  • row (int) –

    The box's position in that tile's returned list. (tile_index, row) is the deterministic tie-breaker.

  • score (float) –

    The ranking value (the score_column extra, or box volume when no tile reported one).

  • roi (Roi) –

    The absolute world-space ROI, extra fields riding on it (a class id, say) — not yet renumbered.

  • bounds (dict[str, tuple[float, float]]) –

    Reference-image pixel coordinates, [min, max) per bbox axis. Treat as read-only.

  • group (tuple) –

    Opaque key of the axes the boxes do not pin (t, or z for a 2D detector on 3D data). Only detections with equal group can be duplicates of each other.

tile_index instance-attribute

tile_index: int

row instance-attribute

row: int

score instance-attribute

score: float

roi instance-attribute

roi: Roi

bounds instance-attribute

bounds: dict[str, tuple[float, float]]

group instance-attribute

group: tuple

JoinProtocol

ngio.iterators.JoinProtocol

Bases: Protocol

Joins the normalized per-ROI frames into one ngio table.

Receives the list described in measure — one DataFrame per ROI in global ROI order, label as a column, every row stamped with roi_index/roi_name, empty ROIs as empty column-less frames — and returns the table. Plain callables satisfy this structurally. Runs once, wherever the join executes: a serial measure, or the distributed gather's finalize. Must be deterministic, or a distributed gather is no longer bit-identical to a serial run.

ConcatJoin

ngio.iterators.ConcatJoin dataclass

ConcatJoin(reference_label: str | None = None)

The default join: concatenate into a FeatureTable indexed by label.

The rows must carry the object id in a label column (or already sit on a label index). The stamped roi_index/roi_name columns ride into the final table. Duplicate label ids (a haloed or overlapping tiling measures border objects more than once) are kept as-is — deduplicating is a custom join's job. The zero-object table keeps its label-only schema.

reference_label class-attribute instance-attribute

reference_label: str | None = None

Mappers

ThreadedMapper

ngio.iterators.ThreadedMapper

ThreadedMapper(max_workers: MaxWorkers = 'auto')

Bases: Generic[T, R]

Run units concurrently on a thread pool.

Each unit's read, func call and write all run on the pool. ngio's getters and setters are thread-safe (immutable per-ROI state over zarr's sync API); func must be thread-safe too — that part of the contract is the caller's.

Units run in conflict-free waves, back to back on one pool — see plan_waves. With one unit, or a resolved pool of one, this is exactly BasicMapper.

Parameters:

  • max_workers (MaxWorkers, default: 'auto' ) –

    "auto" (the default) sizes the pool for round-trip-bound work; an int pins it. None is accepted and means "auto".

Source code in src/ngio/iterators/_mappers.py
def __init__(self, max_workers: MaxWorkers = "auto") -> None:
    _validate_max_workers(max_workers)
    self._max_workers = max_workers

ProcessMapper

ngio.iterators.ProcessMapper

ProcessMapper(max_workers: MaxWorkers = 'auto')

Bases: Generic[T, R]

Run units on a spawn-based process pool.

The fit for pure-Python, GIL-holding funcs (ThreadedMapper already covers IO-bound work and GIL-releasing compute). Each child reads its ROI from the store, applies func, and writes back — pixels never cross the process boundary. For written units the returned result is therefore None (which map discards anyway); only reduce results are pickled back to the parent.

Units run in conflict-free waves (see plan_waves) — which is the entire cross-process safety argument: no lock ngio could take would work across processes anyway.

Constraints: - func (and any transforms the units carry) must be picklable — a module-level function, not a lambda or closure. - In-memory stores are refused: a MemoryStore pickles by value, so a child would write into its own copy and the parent would never see it. - The pool uses the spawn start method: the parent holds zarr's IO event-loop thread, and forking a threaded process is unsafe. Each worker pays an interpreter start and import ngio (~1 s), amortized over the units it processes.

Parameters:

  • max_workers (MaxWorkers, default: 'auto' ) –

    "auto" (the default) or an int; None means "auto". The pool never exceeds the unit count.

Source code in src/ngio/iterators/_mappers.py
def __init__(self, max_workers: MaxWorkers = "auto") -> None:
    _validate_max_workers(max_workers)
    self._max_workers = max_workers

BatchedMapper

ngio.iterators.BatchedMapper

BatchedMapper(
    batch_size: int = 8,
    pad_mode: str = "constant",
    pad_values: int | float = 0,
    read_workers: MaxWorkers = "auto",
)

Stack patches into (B, ...) batches and call func once per batch.

The fit for neural-network inference. Unlike every other mapper, func receives a stacked array — a leading batch axis over up to batch_size patches — and must return an array-like with the same leading axis whose items follow map's per-patch contract.

Ragged tilings are padded per batch to the per-axis maximum before stacking (origin-anchored: real pixels first, padding after); a shape-preserving output is sliced back to each patch's true shape before the write, and halos are trimmed by the setter as usual. A per-item reduction is allowed only on a uniform batch — on a ragged one its result is computed on padded pixels and there is no padding to slice back off, so it raises. Reads within a batch fan out on a thread pool; writes run serially on the dispatching thread, so batched mapping is write-safe on any tiling.

Stacking is a numpy operation, so unlike its siblings this mapper is not generic over the payload type: it accepts bare np.ndarray units only (tuple payloads raise). Its pool argument is named read_workers, not max_workers, because it sizes something different — the other mappers' pools run func, this one's only reads; func always runs once per batch on the dispatching thread.

Parameters:

  • batch_size (int, default: 8 ) –

    Number of patches stacked per func call. The last batch may be smaller.

  • pad_mode (str, default: 'constant' ) –

    np.pad mode used to grow ragged patches to the batch shape, typically "constant" or "reflect".

  • pad_values (int | float, default: 0 ) –

    Fill value when pad_mode="constant", ignored otherwise.

  • read_workers (MaxWorkers, default: 'auto' ) –

    Pool size for the per-batch reads. "auto" (the default) sizes it for round-trip-bound work, an int pins it, 1 reads serially; None means "auto".

Source code in src/ngio/iterators/_mappers.py
def __init__(
    self,
    batch_size: int = 8,
    pad_mode: str = "constant",
    pad_values: int | float = 0,
    read_workers: MaxWorkers = "auto",
) -> None:
    if batch_size < 1:
        raise NgioValueError(f"batch_size must be >= 1, got {batch_size}.")
    _validate_max_workers(read_workers)
    self._batch_size = batch_size
    self._pad_mode = pad_mode
    self._pad_values = pad_values
    self._read_workers = read_workers

BasicMapper

ngio.iterators.BasicMapper

Bases: Generic[T, R]

Serial mapper: read, apply, write (if writable), one unit at a time.

MapperProtocol

ngio.iterators.MapperProtocol

Bases: Protocol[T, R]

Protocol for mappers.

Implementations must honour this contract:

  • The mapper schedules each unit's write; it is not required to invoke unit.setter eagerly or on any particular thread, only to guarantee that all writes are complete when __call__ returns.
  • A unit whose setter is None is read-only: compute and collect its result, never write.
  • For a unit with a setter the collected result may be None instead of the written patch: map discards the results, and holding every patch until the call returns would put the whole output in memory at once. ngio's mappers all collect None for written units.
  • The returned list is ordered by unit.index (results[i] corresponds to iterator.rois[i]), regardless of execution order.
  • units is typically a generator and each unit is expensive to build (store metadata round-trips); generators are not thread-safe. A parallel mapper must materialise and consume units on the dispatching thread only, then distribute the materialised units to its workers.

Compatibility policy: future ngio versions will only ever add optional members to this protocol, probed with getattr — an implementation satisfying today's contract keeps working.

IterUnit

ngio.iterators.IterUnit dataclass

IterUnit(
    *,
    index: int,
    roi: Roi,
    getter: DataGetterProtocol[T],
    setter: DataSetterProtocol[T] | None,
    write_order: Literal["roi", "any"] = "roi",
)

Bases: Generic[T]

One schedulable unit of iterator work: read one ROI, optionally write back.

Attributes:

  • index (int) –

    Position in iterator.rois — mapper results must be returned in this order.

  • roi (Roi) –

    The ROI this unit covers.

  • getter (DataGetterProtocol[T]) –

    Reads the ROI's data.

  • setter (DataSetterProtocol[T] | None) –

    Writes the transformed data back, or None for a read-only unit (a read-only iterator, or any reduce call).

  • write_order (Literal['roi', 'any']) –

    See WriteOrder. A pair is relaxed only when both units declare "any". The field default is "roi" so a hand-built unit is strictly ordered unless it asks otherwise; the iterators stamp their own declared value.

index instance-attribute

index: int

roi instance-attribute

roi: Roi

getter instance-attribute

getter: DataGetterProtocol[T]

setter instance-attribute

setter: DataSetterProtocol[T] | None

write_order class-attribute instance-attribute

write_order: Literal['roi', 'any'] = 'roi'

write_footprint property

write_footprint: ChunkRect | None

The chunk rectangle this unit writes, at write granularity.

The granularity is the shard shape when the output array is sharded, the chunk shape otherwise. None when the unit is read-only or its write selection is empty — either way there is nothing to claim and the unit conflicts with nothing.

Scheduling

The primitives behind the mappers' conflict-free schedule — useful for inspecting how a tiling will parallelize or split into jobs.

plan_waves

ngio.iterators.plan_waves

plan_waves(
    units: Sequence[IterUnit[T]], *, log: bool = True
) -> list[list[IterUnit[T]]]

Partition units into conflict-free waves that respect ROI-index order.

Two units whose write footprints share a write unit (a chunk, or a shard when the output is sharded) of the same output array never share a wave: a wave's writes are pairwise disjoint at write granularity, so running one wave at a time preserves the single-writer-per-write-unit invariant that makes parallel writes lock-free in any topology. Read-only units and units with an empty write selection conflict with nothing and land in the first wave. Assignment runs in ascending unit.index (first-fit greedy), so the schedule is a pure function of the unit sequence. A conflict-free set — a by_write_units() tiling, say — yields a single wave.

On top of the safety constraint, write_order="roi" units are ordered on pixel overlap: the higher unit.index lands in a strictly later wave, so the later ROI wins under every mapper and in the manual iter loop. A pair where both units declare "any" is scheduled for parallelism alone. Sharing a write unit without sharing pixels never orders, so disjoint-pixel tilings keep their packed schedule.

Parameters:

  • units (Sequence[IterUnit[T]]) –

    The units to schedule.

  • log (bool, default: True ) –

    Emit schedule-quality log messages. The parallel mappers keep this on; serial callers ordering their units pass False.

Source code in src/ngio/iterators/_mappers.py
def plan_waves(
    units: Sequence[IterUnit[T]], *, log: bool = True
) -> list[list[IterUnit[T]]]:
    """Partition units into conflict-free waves that respect ROI-index order.

    Two units whose write footprints share a write unit (a chunk, or a shard
    when the output is sharded) of the same output array never share a wave:
    a wave's writes are pairwise disjoint at write granularity, so running
    one wave at a time preserves the single-writer-per-write-unit invariant
    that makes parallel writes lock-free in any topology. Read-only units and
    units with an empty write selection conflict with nothing and land in the
    first wave. Assignment runs in ascending `unit.index` (first-fit greedy),
    so the schedule is a pure function of the unit sequence. A conflict-free
    set — a `by_write_units()` tiling, say — yields a single wave.

    On top of the safety constraint, `write_order="roi"` units are ordered
    on pixel overlap: the higher `unit.index` lands in a strictly later
    wave, so the later ROI wins under every mapper and in the manual
    `iter` loop. A pair where both units declare `"any"` is scheduled for
    parallelism alone. Sharing a write unit without sharing pixels never
    orders, so disjoint-pixel tilings keep their packed schedule.

    Args:
        units: The units to schedule.
        log: Emit schedule-quality log messages. The parallel mappers keep
            this on; serial callers ordering their units pass `False`.
    """
    if not units:
        return []
    adjacency = _write_conflict_edges(units, _collect_write_footprints(units))
    precedence = _precedence_edges(units)

    wave_of: dict[int, int] = {}
    waves: list[list[IterUnit[T]]] = [[]]
    for unit in sorted(units, key=lambda u: u.index):
        taken = {
            wave_of[other]
            for other in adjacency.get(unit.index, ())
            if other in wave_of
        }
        # Ascending-index assignment means every already-assigned pixel
        # neighbour has a lower index — exactly the ones this unit must
        # write after.
        floor = max(
            (
                wave_of[other]
                for other in precedence.get(unit.index, ())
                if other in wave_of
            ),
            default=-1,
        )
        color = floor + 1
        while color in taken:
            color += 1
        wave_of[unit.index] = color
        if color == len(waves):
            waves.append([])
        waves[color].append(unit)

    if not log:
        return waves
    largest = max(len(wave) for wave in waves)
    if largest == 1 and len(units) > 1:
        logger.warning(
            f"Parallel map degraded to a serial schedule: the {len(units)} "
            "units' write layout leaves no two free to run concurrently "
            "(shared write units, or pixel overlaps that force an order). "
            "Re-tile with `by_write_units()` for a single fully-parallel wave."
        )
    elif len(waves) > 1:
        logger.info(
            f"Parallel map scheduled {len(units)} units into {len(waves)} "
            f"conflict-free waves (largest wave: {largest} units). Tiling with "
            "`by_write_units()` would yield a single wave."
        )
    return waves

canonical_unit_order

ngio.iterators.canonical_unit_order

canonical_unit_order(
    units: Sequence[IterUnit[T]],
) -> list[IterUnit[T]]

Units in flattened wave order — the one write order every mapper uses.

Wave 0 in index order, then wave 1, and so on. With no write conflicts this is plain index order. On pixel-overlapping "roi" units the higher ROI index always lands later; all-"any" pairs follow the schedule — either way a pure function of the unit sequence.

Source code in src/ngio/iterators/_mappers.py
def canonical_unit_order(units: Sequence[IterUnit[T]]) -> list[IterUnit[T]]:
    """Units in flattened wave order — the one write order every mapper uses.

    Wave 0 in index order, then wave 1, and so on. With no write conflicts
    this is plain index order. On pixel-overlapping `"roi"` units the
    higher ROI index always lands later; all-`"any"` pairs follow the
    schedule — either way a pure function of the unit sequence.
    """
    return [unit for wave in plan_waves(units, log=False) for unit in wave]

write_conflict_components

ngio.iterators.write_conflict_components

write_conflict_components(
    units: Sequence[IterUnit[Any]],
) -> list[list[int]]

Connected components of the write-conflict graph, as unit-index lists.

Two units are connected when their write footprints share a write unit (a chunk, or a shard when the output is sharded) of the same output array — the same adjacency plan_waves schedules by, computed by the same code, so the scheduler and the splitter can never disagree about who conflicts. Read-only units and units with an empty write selection conflict with nothing and are singletons.

A component is the unit of independence: units in different components never share a write unit, so they need no coordination in any topology. That is the property for_job builds on — a component never spans two jobs. Each component is a sorted list of unit.index; components are ordered by their smallest index, and the whole result is a pure function of the unit sequence.

Source code in src/ngio/iterators/_mappers.py
def write_conflict_components(units: Sequence[IterUnit[Any]]) -> list[list[int]]:
    """Connected components of the write-conflict graph, as unit-index lists.

    Two units are connected when their write footprints share a write unit (a
    chunk, or a shard when the output is sharded) of the same output array —
    the same adjacency `plan_waves` schedules by, computed by the same code,
    so the scheduler and the splitter can never disagree about who conflicts.
    Read-only units and units with an empty write selection conflict with
    nothing and are singletons.

    A component is the unit of independence: units in different components
    never share a write unit, so they need no coordination in any topology.
    That is the property `for_job` builds on — a component never spans
    two jobs. Each component is a sorted list of `unit.index`; components are
    ordered by their smallest index, and the whole result is a pure function
    of the unit sequence.
    """
    adjacency = _write_conflict_edges(units, _collect_write_footprints(units))

    components: list[list[int]] = []
    visited: set[int] = set()
    for index in sorted(unit.index for unit in units):
        if index in visited:
            continue
        visited.add(index)
        component = []
        stack = [index]
        while stack:
            current = stack.pop()
            component.append(current)
            for neighbor in adjacency.get(current, ()):
                if neighbor not in visited:
                    visited.add(neighbor)
                    stack.append(neighbor)
        components.append(sorted(component))
    return components

compute_write_footprint

ngio.iterators.compute_write_footprint

compute_write_footprint(
    setter: DataSetterProtocol[Any],
) -> ChunkRect | None

The chunk rectangle a setter will write, at write granularity.

The granularity is the shard shape when the target array is sharded (writes are atomic per shard object), the chunk shape otherwise — see AbstractImage.write_granularity. Returns None when the setter's selection is empty.

Source code in src/ngio/iterators/_mappers.py
def compute_write_footprint(setter: DataSetterProtocol[Any]) -> ChunkRect | None:
    """The chunk rectangle a setter will write, at write granularity.

    The granularity is the shard shape when the target array is sharded (writes
    are atomic per shard object), the chunk shape otherwise — see
    `AbstractImage.write_granularity`. Returns `None` when the setter's
    selection is empty.
    """
    granularity = setter.zarr_array.shards or setter.slicing_ops.on_disk_chunks
    return compute_chunk_rect(
        shape=setter.slicing_ops.on_disk_shape,
        chunks=granularity,
        slicing_tuple=setter.slicing_ops.normalized_slicing_tuple,
    )

Types

JobArgs

ngio.iterators.JobArgs

Bases: TypedDict

One entry of prepare_jobs's parallelization list.

A plain dict at runtime — JSON-serializable for schedulers that ship task arguments as JSON (Fractal's parallelization list), and it splats straight into the selector: iterator.for_job(**args).

job_index instance-attribute

job_index: int

n_jobs instance-attribute

n_jobs: int

TailPolicy

ngio.iterators.TailPolicy module-attribute

TailPolicy = Literal['clip', 'balance', 'shift', 'drop']

How by_grid handles the non-whole tile at the end of an axis.

HaloMargins

ngio.iterators.HaloMargins module-attribute

HaloMargins = dict[str, tuple[int, int]]

Per-axis pixels added on each side of a ROI when reading, keyed by axis name.

FeatureFuncResult

ngio.iterators.FeatureFuncResult module-attribute

FeatureFuncResult: TypeAlias = (
    pd.DataFrame | dict[str, list]
)

What a per-ROI feature function may return: a DataFrame, or the cheaper dict-of-columns ({"label": [...], "area": [...]}) — lighter to build in a worker and to ship across a ProcessMapper boundary; normalization turns either into one DataFrame per ROI before the join.

MaxWorkers

ngio.iterators.MaxWorkers module-attribute

MaxWorkers = int | Literal['auto'] | None

Accepted by every max_workers argument. 1 is serial deliberately; "auto" picks a pool sized for round-trip-bound work. None means "unspecified", and what that resolves to belongs to the call site: the plate fan-outs run serially (until the ngio=1.2 default flip they warn about), the parallel mappers treat it as "auto" — asking for such a mapper is already the opt-in.

AbstractIteratorBuilder

The shared method surface of every iterator — the reshaping calls (by_grid, by_blocks, by_chunks, by_write_units, product, with_halo, for_job), each class's reconciliation declaration (on_overlap, with_stitch, with_join, with_nms), the generic execution calls (iter — batched via batch_size — map, reduce) beneath each iterator's topic verb (process, segment, measure, detect), the distributed-run init step (prepare_jobs), and the one gather, finalize — the writers consolidate and return None, the read-only iterators merge their banked partials and return the table.

ngio.iterators.AbstractIteratorBuilder

Bases: ABC, Generic[NumpyPipeType, DaskPipeType, FinalizeType]

Base class for building iterators over ROIs.

Constructor naming convention across the concrete iterators: a bare parameter (channel_selection, input_transforms) refers to the input image; output-side parameters are always prefixed (output_channel_selection, output_transforms), and parameters targeting another object name it (label_transforms). A parameter that configures a downstream consolidation is spelled consolidation_mode (here and on Label.relabel_sequential); only consolidate() itself shortens it to mode, where the prefix would be redundant.

rois property

rois: list[Roi]

Get the list of ROIs for the iterator.

ref_image property

ref_image: AbstractImage

Get the reference image for the iterator.

output_image property

output_image: AbstractImage | None

The image this iterator writes to, or None for a read-only iterator.

halo property

halo: dict[str, int]

The per-axis read margin, empty when there is none.

partition_indices property

partition_indices: list[int] | None

The unit indices this iterator is restricted to, or None.

None on an unrestricted iterator; on a for_job slice, the sorted positions (in the full ROI list) of the units it runs. The whole layout of a split is [it.for_job(i, n).partition_indices for i in range(n)] — the pre-flight check before submitting a distributed run: one fat list plus empties means the output chunking, not the cluster, is the constraint.

with_halo

with_halo(
    x: int = 0, y: int = 0, z: int = 0, t: int = 0
) -> Self

Read a margin of context around each ROI, but write only the ROI.

The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.

Note

The margins do not survive a rescale: a shape-changing transform on the data path (a ZoomTransform) makes the crop refuse — iterate on a coarser pyramid level instead. A MaskTransform composes normally (its zoom rescales the label it reads).

Parameters:

  • x (int, default: 0 ) –

    Pixels added on each side along x.

  • y (int, default: 0 ) –

    Pixels added on each side along y.

  • z (int, default: 0 ) –

    Pixels added on each side along z.

  • t (int, default: 0 ) –

    Frames added on each side along t.

Returns:

  • Self –

    A new iterator reading with the halo.

Raises:

  • NgioValueError –

    On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps roi_index/roi_name so your coalesce can), on an in-place iterator, or for a negative margin.

Example
# 8 px of context per side, disjoint writes, still parallel.
it = iterator.by_write_units().with_halo(x=8, y=8)
it.map(smooth, mapper=ThreadedMapper("auto"))
Source code in src/ngio/iterators/_abstract_iterator.py
def with_halo(self, x: int = 0, y: int = 0, z: int = 0, t: int = 0) -> Self:
    """Read a margin of context around each ROI, but write only the ROI.

    The function receives the grown region and must return it grown too;
    the border is cropped before the write, so tiles come out seamless.
    Margins are reference-image pixels, clipped at the borders. Write
    footprints are unchanged, so a haloed iterator parallelizes exactly
    as far as it did without one.

    Note:
        The margins do not survive a rescale: a shape-changing transform
        on the data path (a `ZoomTransform`) makes the crop refuse —
        iterate on a coarser pyramid level instead. A `MaskTransform`
        composes normally (its zoom rescales the label it reads).

    Args:
        x: Pixels added on each side along x.
        y: Pixels added on each side along y.
        z: Pixels added on each side along z.
        t: Frames added on each side along t.

    Returns:
        A new iterator reading with the halo.

    Raises:
        NgioValueError: On a read-only iterator that has not opted in
            (no write to crop the halo from; detection opts in and
            reconciles the overlap with NMS, feature extraction opts in
            and stamps `roi_index`/`roi_name` so your `coalesce` can),
            on an in-place iterator, or for a negative margin.

    Example:
        ```python
        # 8 px of context per side, disjoint writes, still parallel.
        it = iterator.by_write_units().with_halo(x=8, y=8)
        it.map(smooth, mapper=ThreadedMapper("auto"))
        ```
    """
    halo = {"x": x, "y": y, "z": z, "t": t}
    for axis_name, margin in halo.items():
        if margin < 0:
            raise NgioValueError(
                f"Halo along '{axis_name}' must be >= 0, got {margin}."
            )
    if self.output_image is None and not self._allow_readonly_halo:
        name = self.__class__.__name__
        raise NgioValueError(
            f"{name} is read-only, so there is no written region for a "
            "halo to surround. Widen the ROIs themselves instead (e.g. "
            "`by_grid(...)` with a stride smaller than the size)."
        )
    output = self.output_image
    if output is not None and _is_same_zarr_array(
        self._ref_image.zarr_array, output.zarr_array
    ):
        raise NgioValueError(
            "In-place iteration cannot take a halo: each tile's grown "
            "read covers pixels that neighbouring tiles write, so the "
            "result depends on execution order. Write to a different "
            "image or label, or drop the halo."
        )
    new_instance = self._new_from_rois(self.rois)
    new_instance._halo = {ax: m for ax, m in halo.items() if m > 0}
    return new_instance

by_grid

by_grid(
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self

Tile the current ROIs with a regular grid.

Parameters:

  • size_x (int | None, default: None ) –

    Tile size along x; None (default) means the full extent.

  • size_y (int | None, default: None ) –

    Tile size along y; None means the full extent.

  • size_z (int | None, default: None ) –

    Tile size along z; None means the full extent.

  • size_t (int | None, default: None ) –

    Tile size along t; None means the full extent.

  • stride_x (int | None, default: None ) –

    Step between tiles along x; None means size_x (adjacent tiles). A smaller stride produces overlapping tiles.

  • stride_y (int | None, default: None ) –

    As stride_x, along y.

  • stride_z (int | None, default: None ) –

    As stride_x, along z.

  • stride_t (int | None, default: None ) –

    As stride_x, along t.

  • tail (TailPolicy, default: 'clip' ) –

    What happens where the grid does not divide the axis. "clip" (default) shrinks the last tile to the border; "balance" re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles, stride == size); "shift" moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution; "drop" discards the tail and leaves that border uncovered.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_grid(
    self,
    *,
    size_x: int | None = None,
    size_y: int | None = None,
    size_z: int | None = None,
    size_t: int | None = None,
    stride_x: int | None = None,
    stride_y: int | None = None,
    stride_z: int | None = None,
    stride_t: int | None = None,
    tail: TailPolicy = "clip",
    base_name: str = "",
) -> Self:
    """Tile the current ROIs with a regular grid.

    Args:
        size_x: Tile size along x; `None` (default) means the full extent.
        size_y: Tile size along y; `None` means the full extent.
        size_z: Tile size along z; `None` means the full extent.
        size_t: Tile size along t; `None` means the full extent.
        stride_x: Step between tiles along x; `None` means `size_x`
            (adjacent tiles). A smaller stride produces overlapping tiles.
        stride_y: As `stride_x`, along y.
        stride_z: As `stride_x`, along z.
        stride_t: As `stride_x`, along t.
        tail: What happens where the grid does not divide the axis.
            `"clip"` (default) shrinks the last tile to the border;
            `"balance"` re-splits the last full tile and the tail into two
            near-equal tiles, so a thin overhang never yields a thin tile
            (needs adjacent tiles, `stride == size`); `"shift"` moves the
            last tile back to stay full-size — it then overlaps its
            neighbour, so a writing iterator needs a merge or serial
            execution; `"drop"` discards the tail and leaves that border
            uncovered.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the tiled ROIs.
    """
    rois = by_grid(
        rois=self.rois,
        ref_image=self.ref_image,
        size_x=size_x,
        size_y=size_y,
        size_z=size_z,
        size_t=size_t,
        stride_x=stride_x,
        stride_y=stride_y,
        stride_z=stride_z,
        stride_t=stride_t,
        tail=tail,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_blocks

by_blocks(
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self

Divide each axis into a fixed number of near-equal blocks.

The complement of by_grid: you say how many tiles, not how big. Block lengths along an axis differ by at most one pixel, so there is no tail to police — the partition is balanced by construction.

Parameters:

  • num_x (int, default: 1 ) –

    Number of blocks along x (default 1: no split).

  • num_y (int, default: 1 ) –

    Number of blocks along y.

  • num_z (int, default: 1 ) –

    Number of blocks along z.

  • num_t (int, default: 1 ) –

    Number of blocks along t.

  • base_name (str, default: '' ) –

    Prefix for the generated tile names.

Returns:

  • Self –

    A new iterator over the blocks.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_blocks(
    self,
    *,
    num_x: int = 1,
    num_y: int = 1,
    num_z: int = 1,
    num_t: int = 1,
    base_name: str = "",
) -> Self:
    """Divide each axis into a fixed number of near-equal blocks.

    The complement of `by_grid`: you say how many tiles, not how big.
    Block lengths along an axis differ by at most one pixel, so there is
    no tail to police — the partition is balanced by construction.

    Args:
        num_x: Number of blocks along x (default 1: no split).
        num_y: Number of blocks along y.
        num_z: Number of blocks along z.
        num_t: Number of blocks along t.
        base_name: Prefix for the generated tile names.

    Returns:
        A new iterator over the blocks.
    """
    rois = by_blocks(
        rois=self.rois,
        ref_image=self.ref_image,
        num_x=num_x,
        num_y=num_y,
        num_z=num_z,
        num_t=num_t,
        base_name=base_name,
    )
    return self._new_from_rois(rois)

by_yx

by_yx() -> Self

Split each ROI into one 2D plane per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, z, c) combination, each covering the full y/x extent — the shape for "run this 2D function on every plane".

Source code in src/ngio/iterators/_abstract_iterator.py
def by_yx(self) -> Self:
    """Split each ROI into one 2D plane per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, z, c)` combination, each covering the full y/x extent — the
    shape for "run this 2D function on every plane".
    """
    rois = by_yx(self.rois, self.ref_image)
    return self._new_from_rois(rois)

by_zyx

by_zyx(strict: bool = True) -> Self

Split each ROI into one 3D volume per remaining coordinate.

A broadcasting call, not a tiling: every ROI becomes one ROI per (t, c) combination, each covering the full z/y/x extent.

Parameters:

  • strict (bool, default: True ) –

    Require a real z axis (present and larger than 1); with False, images without one fall back to 2D planes.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_zyx(self, strict: bool = True) -> Self:
    """Split each ROI into one 3D volume per remaining coordinate.

    A broadcasting call, not a tiling: every ROI becomes one ROI per
    `(t, c)` combination, each covering the full z/y/x extent.

    Args:
        strict: Require a real z axis (present and larger than 1);
            with `False`, images without one fall back to 2D planes.
    """
    rois = by_zyx(self.rois, self.ref_image, strict=strict)
    return self._new_from_rois(rois)

by_chunks

by_chunks(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the input image's chunk grid.

Reads are chunk-granular, so this is the natural tiling for read-heavy work. For collision-free parallel writes use by_write_units, which tiles by the output's write granularity.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_chunks(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the input image's chunk grid.

    Reads are chunk-granular, so this is the natural tiling for
    read-heavy work. For collision-free parallel *writes* use
    `by_write_units`, which tiles by the output's write granularity.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=None,
    )
    return self._new_from_rois(rois)

by_write_units

by_write_units(
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self

Tile the ROIs by the output image's write granularity.

The write unit is the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. With no overlap the resulting ROIs pass check_if_write_units_overlap by construction, so a parallel map runs as a single fully-parallel wave — a throughput optimization; conflicting tilings are still safe, just scheduled in more waves. Falls back to the input chunk grid when the iterator is read-only.

Parameters:

  • overlap_x (int, default: 0 ) –

    Overlap between adjacent tiles along x, in pixels.

  • overlap_y (int, default: 0 ) –

    Overlap along y.

  • overlap_z (int, default: 0 ) –

    Overlap along z.

  • overlap_t (int, default: 0 ) –

    Overlap along t.

Returns:

  • Self –

    A new iterator with tiled ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def by_write_units(
    self,
    *,
    overlap_x: int = 0,
    overlap_y: int = 0,
    overlap_z: int = 0,
    overlap_t: int = 0,
) -> Self:
    """Tile the ROIs by the output image's write granularity.

    The write unit is the shard shape when the output is sharded (writes
    are atomic per shard object), the chunk shape otherwise. With no
    overlap the resulting ROIs pass `check_if_write_units_overlap` by
    construction, so a parallel `map` runs as a single fully-parallel
    wave — a throughput optimization; conflicting tilings are still safe,
    just scheduled in more waves. Falls back to the input chunk grid when
    the iterator is read-only.

    Args:
        overlap_x: Overlap between adjacent tiles along x, in pixels.
        overlap_y: Overlap along y.
        overlap_z: Overlap along z.
        overlap_t: Overlap along t.

    Returns:
        A new iterator with tiled ROIs.
    """
    rois = by_storage_units(
        self.rois,
        self.ref_image,
        overlap_x=overlap_x,
        overlap_y=overlap_y,
        overlap_z=overlap_z,
        overlap_t=overlap_t,
        grid_image=self.output_image,
    )
    return self._new_from_rois(rois)

product

product(other: list[Roi] | GenericRoiTable) -> Self

Cartesian product of the current ROIs with an arbitrary list of ROIs.

Source code in src/ngio/iterators/_abstract_iterator.py
def product(self, other: list[Roi] | GenericRoiTable) -> Self:
    """Cartesian product of the current ROIs with an arbitrary list of ROIs."""
    if isinstance(other, GenericRoiTable):
        other = other.rois()
    rois = rois_product(self.rois, other)
    return self._new_from_rois(rois)

build_numpy_getter abstractmethod

build_numpy_getter(
    roi: Roi,
) -> DataGetterProtocol[NumpyPipeType]

Build a getter function for the given ROI.

Source code in src/ngio/iterators/_abstract_iterator.py
@abstractmethod
def build_numpy_getter(self, roi: Roi) -> DataGetterProtocol[NumpyPipeType]:
    """Build a getter function for the given ROI."""
    raise NotImplementedError

build_dask_getter abstractmethod

build_dask_getter(
    roi: Roi,
) -> DataGetterProtocol[DaskPipeType]

Build a Dask getter function for the given ROI.

Source code in src/ngio/iterators/_abstract_iterator.py
@abstractmethod
def build_dask_getter(self, roi: Roi) -> DataGetterProtocol[DaskPipeType]:
    """Build a Dask getter function for the given ROI."""
    raise NotImplementedError

finalize abstractmethod

finalize() -> FinalizeType

The one gather step, run once after every unit (or job) completes.

Writing iterators resolve any pending reconciliation (the stitch) and consolidate the output pyramid, returning None; the serial verbs call it automatically. Read-only iterators merge the partials banked by a distributed run and return the final table — nothing is registered, the table is the caller's to store with add_table.

Raises on a for_job slice: the gather is global and belongs to the unrestricted iterator, once, after all jobs.

Source code in src/ngio/iterators/_abstract_iterator.py
@abstractmethod
def finalize(self) -> FinalizeType:
    """The one gather step, run once after every unit (or job) completes.

    Writing iterators resolve any pending reconciliation (the stitch) and
    consolidate the output pyramid, returning `None`; the serial verbs
    call it automatically. Read-only iterators merge the partials banked
    by a distributed run and return the final table — nothing is
    registered, the table is the caller's to store with `add_table`.

    Raises on a `for_job` slice: the gather is global and belongs to the
    unrestricted iterator, once, after all jobs.
    """
    raise NotImplementedError

iter

iter(
    lazy: Literal[True],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
    lazy: Literal[True],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
    lazy: Literal[False],
    data_mode: Literal["numpy"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
    lazy: Literal[False],
    data_mode: Literal["dask"],
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
    lazy: Literal[False] = ...,
    data_mode: Literal["numpy"] | None = ...,
    iterator_mode: Literal["readonly"] = ...,
    *,
    batch_size: int,
) -> Generator[list[NumpyPipeType]]
iter(
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal[
        "readwrite", "readonly"
    ] = "readonly",
    *,
    batch_size: int | None = None,
) -> Generator

Iterate the ROIs' payloads, read-only.

With batch_size set, yields payload lists of up to that many items (the last batch may be smaller). Batching is numpy-only and eager: combining batch_size with lazy=True or data_mode="dask" raises. On a writing iterator (see WritingIteratorBuilder.iter) the default mode is "readwrite" and the loop yields (patch, writer) pairs instead.

Parameters:

  • lazy (bool, default: False ) –

    Yield getter handles instead of materialized data.

  • data_mode (Literal['numpy', 'dask'] | None, default: None ) –

    "numpy" or "dask" (deprecated, see note).

  • iterator_mode (Literal['readwrite', 'readonly'], default: 'readonly' ) –

    "readonly" yields patches alone; on a writing iterator "readwrite" yields (patch, writer) pairs.

  • batch_size (int | None, default: None ) –

    When set, yield batches of up to this many items; the last batch may be smaller.

Note

The dask data mode is deprecated and will be removed in ngio=1.2; from then on iter() yields numpy arrays.

Source code in src/ngio/iterators/_abstract_iterator.py
def iter(
    self,
    lazy: bool = False,
    data_mode: Literal["numpy", "dask"] | None = None,
    iterator_mode: Literal["readwrite", "readonly"] = "readonly",
    *,
    batch_size: int | None = None,
) -> Generator:
    """Iterate the ROIs' payloads, read-only.

    With `batch_size` set, yields payload lists of up to that many
    items (the last batch may be smaller). Batching is numpy-only and
    eager: combining `batch_size` with `lazy=True` or
    `data_mode="dask"` raises. On a writing iterator (see
    `WritingIteratorBuilder.iter`) the default mode is `"readwrite"`
    and the loop yields `(patch, writer)` pairs instead.

    Args:
        lazy: Yield getter handles instead of materialized data.
        data_mode: `"numpy"` or `"dask"` (deprecated, see note).
        iterator_mode: `"readonly"` yields patches alone; on a writing
            iterator `"readwrite"` yields `(patch, writer)` pairs.
        batch_size: When set, yield batches of up to this many items;
            the last batch may be smaller.

    Note:
        The dask data mode is deprecated and will be removed in ngio=1.2;
        from then on `iter()` yields numpy arrays.
    """
    return self._iter_impl(
        lazy=lazy,
        data_mode=data_mode,
        iterator_mode=iterator_mode,
        batch_size=batch_size,
    )

iter_as_numpy

iter_as_numpy()

Alias for iter(data_mode="numpy").

Source code in src/ngio/iterators/_abstract_iterator.py
def iter_as_numpy(
    self,
):
    """Alias for `iter(data_mode="numpy")`."""
    return self._iter(lazy=False, data_mode="numpy", iterator_mode="readonly")

iter_as_dask

iter_as_dask()

Alias for iter(data_mode="dask").

Deprecated: removed in ngio=1.2.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="iter_as_numpy() (or Image.get_as_dask() for a lazy array)",
    removed_in="1.2",
)
def iter_as_dask(
    self,
):
    """Alias for `iter(data_mode="dask")`.

    Deprecated: removed in ngio=1.2.
    """
    return self._iter(lazy=False, data_mode="dask", iterator_mode="readonly")

for_job

for_job(job_index: int, n_jobs: int) -> Self

Restrict the iterator to one of n_jobs independent partitions.

No write unit is ever shared between two partitions, so the jobs need no locks, no barriers, and no channel to one another — safe as separate SLURM array tasks, in any order. Every job must build an identical iterator (construction is metadata-only and the partition derives deterministically from it), apply for_job last in the builder chain, and its verbs deliberately do not finalize — a slice's map/process/segment writes only its share, and measure/ detect bank a partial instead of joining. The gather step is the unrestricted iterator's finalize(), once, after all jobs. Stitched and read-only runs also need prepare_jobs(n_jobs) first. See the distributed-runs guide in the iterators docs.

Parameters:

  • job_index (int) –

    This job's partition, 0 <= job_index < n_jobs (e.g. int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops.

  • n_jobs (int) –

    How many partitions the run is split into; must match across every job and the gather.

Returns:

  • Self –

    The restricted iterator; further reshaping refuses.

Example
# array task:
iterator = SegmentationIterator(image, label).by_chunks()
iterator.for_job(job_index, n_jobs=n_jobs).map(func)

# gather task, after all jobs:
SegmentationIterator(image, label).by_chunks().finalize()
Source code in src/ngio/iterators/_abstract_iterator.py
def for_job(self, job_index: int, n_jobs: int) -> Self:
    """Restrict the iterator to one of `n_jobs` independent partitions.

    No write unit is ever shared between two partitions, so the jobs
    need no locks, no barriers, and no channel to one another — safe as
    separate SLURM array tasks, in any order. Every job must build an
    identical iterator (construction is metadata-only and the partition
    derives deterministically from it), apply `for_job` *last* in the
    builder chain, and its verbs deliberately do not finalize — a slice's
    `map`/`process`/`segment` writes only its share, and `measure`/
    `detect` bank a partial instead of joining. The gather step is the
    unrestricted iterator's `finalize()`, once, after all jobs.
    Stitched and read-only runs also need `prepare_jobs(n_jobs)` first.
    See the distributed-runs guide in the iterators docs.

    Args:
        job_index: This job's partition, `0 <= job_index < n_jobs`
            (e.g. `int($SLURM_ARRAY_TASK_ID)`). Partitions beyond the
            number of independent unit groups are empty no-ops.
        n_jobs: How many partitions the run is split into; must match
            across every job and the gather.

    Returns:
        The restricted iterator; further reshaping refuses.

    Example:
        ```python
        # array task:
        iterator = SegmentationIterator(image, label).by_chunks()
        iterator.for_job(job_index, n_jobs=n_jobs).map(func)

        # gather task, after all jobs:
        SegmentationIterator(image, label).by_chunks().finalize()
        ```
    """
    if self._partition is not None:
        _, current_index, current_n = self._partition
        raise NgioValueError(
            f"This iterator is already restricted to partition "
            f"{current_index}/{current_n}; partitions do not nest. "
            "Apply `for_job` once, to the full iterator."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    if not 0 <= job_index < n_jobs:
        raise NgioValueError(
            f"job_index must be in [0, {n_jobs}), got {job_index}."
        )
    self._require_splittable()
    self._validate_job_context(n_jobs)
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    restricted = self._new_from_rois(self.rois)
    restricted._partition = (
        frozenset(partition[job_index]),
        job_index,
        n_jobs,
    )
    return restricted

prepare_jobs

prepare_jobs(n_jobs: int) -> list[JobArgs]

Set up a distributed run and return its parallelization list.

The init step of the three-phase recipe (init → parallel jobs → consolidate): performs the iterator's setup — always clearing stale scratch state from earlier runs first — and returns one JobArgs per non-empty partition. Optional for plain writers; a stitching or read-only run requires it (the scratch root must exist, race-free, before any job writes into it).

Parameters:

  • n_jobs (int) –

    How many partitions to split into; empty ones are dropped from the list.

Returns:

  • list[JobArgs] –

    One {"job_index": ..., "n_jobs": ...} dict per non-empty

  • list[JobArgs] –

    partition — JSON-ready, splatting into for_job(**args).

Example
args_list = iterator.prepare_jobs(n_jobs=4)  # init task
iterator.for_job(**args).map(func)  # per parallel task
iterator.finalize()  # consolidate task
Source code in src/ngio/iterators/_abstract_iterator.py
def prepare_jobs(self, n_jobs: int) -> list[JobArgs]:
    """Set up a distributed run and return its parallelization list.

    The init step of the three-phase recipe (init → parallel jobs →
    consolidate): performs the iterator's setup — always clearing stale
    scratch state from earlier runs first — and returns one `JobArgs`
    per non-empty partition. Optional for plain writers; a stitching or
    read-only run requires it (the scratch root must exist, race-free,
    before any job writes into it).

    Args:
        n_jobs: How many partitions to split into; empty ones are
            dropped from the list.

    Returns:
        One `{"job_index": ..., "n_jobs": ...}` dict per non-empty
        partition — JSON-ready, splatting into `for_job(**args)`.

    Example:
        ```python
        args_list = iterator.prepare_jobs(n_jobs=4)  # init task
        iterator.for_job(**args).map(func)  # per parallel task
        iterator.finalize()  # consolidate task
        ```
    """
    if self._partition is not None:
        raise NgioValueError(
            "prepare_jobs belongs to the unrestricted iterator; this one "
            "is already restricted to a partition."
        )
    if n_jobs < 1:
        raise NgioValueError(f"n_jobs must be >= 1, got {n_jobs}.")
    self._require_splittable()
    if self.output_image is not None:
        self._validate_write_plan()
    units = list(self._numpy_units_generator())
    partition = partition_components(write_conflict_components(units), n_jobs)
    self._prepare_distributed(n_jobs)
    return [
        JobArgs(job_index=job_index, n_jobs=n_jobs)
        for job_index, indices in enumerate(partition)
        if indices
    ]

reduce

reduce(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Units are built read-only even on writable iterators: nothing is written and finalize does not run.

Parameters:

  • func (Callable[[NumpyPipeType], R]) –

    The function to apply; under a parallel mapper it must be safe on worker threads (or processes).

  • mapper (MapperProtocol[NumpyPipeType, R] | None, default: None ) –

    How the units are scheduled. None (the default) is serial; pass ThreadedMapper() or ProcessMapper() to fan out, sized by their own max_workers argument.

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run.

    Args:
        func: The function to apply; under a parallel mapper it must be
            safe on worker threads (or processes).
        mapper: How the units are scheduled. `None` (the default) is
            serial; pass `ThreadedMapper()` or `ProcessMapper()` to fan
            out, sized by their own `max_workers` argument.

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    if mapper is None:
        mapper = BasicMapper[NumpyPipeType, R]()
    return mapper(func, list(self._numpy_units_generator(with_setters=False)))

reduce_as_numpy

reduce_as_numpy(
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]

Alias for reduce().

Source code in src/ngio/iterators/_abstract_iterator.py
def reduce_as_numpy(
    self,
    func: Callable[[NumpyPipeType], R],
    *,
    mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]:
    """Alias for `reduce()`."""
    return self.reduce(func, mapper=mapper)

reduce_as_dask

reduce_as_dask(
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]

Apply a function to every ROI and collect the results without writing.

Deprecated: removed in ngio=1.2.

Units are built read-only even on writable iterators: nothing is written and finalize does not run. Runs serially — the units hand func lazy dask arrays, so the per-unit work is graph construction, not IO. A parallel mapper raises (see map_as_dask).

Returns:

  • list[R] –

    One result per ROI, in ROI order: results[i] corresponds to

  • list[R] –

    self.rois[i], regardless of the mapper's execution order.

Source code in src/ngio/iterators/_abstract_iterator.py
@deprecated(
    replacement="reduce() / reduce_as_numpy(mapper=...)",
    removed_in="1.2",
)
def reduce_as_dask(
    self,
    func: Callable[[DaskPipeType], R],
    *,
    mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]:
    """Apply a function to every ROI and collect the results without writing.

    Deprecated: removed in ngio=1.2.

    Units are built read-only even on writable iterators: nothing is
    written and `finalize` does not run. Runs serially — the units hand
    `func` *lazy* dask arrays, so the per-unit work is graph
    construction, not IO. A parallel `mapper` raises (see `map_as_dask`).

    Returns:
        One result per ROI, in ROI order: `results[i]` corresponds to
        `self.rois[i]`, regardless of the mapper's execution order.
    """
    self._require_no_halo_on_readonly_verbs()
    _mapper = self._require_serial_dask_mapper(mapper)
    return _mapper(func, list(self._dask_units_generator(with_setters=False)))

check_if_regions_overlap

check_if_regions_overlap() -> bool

Check if any of the ROIs overlap logically.

If two ROIs cover the same pixel, they are considered to overlap. This does not consider chunking or other storage details. Measured on the read regions — with a halo, neighbouring reads overlap by construction. For write safety, check_if_write_units_overlap is the right question.

Returns:

  • bool –

    True if any ROIs overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def check_if_regions_overlap(self) -> bool:
    """Check if any of the ROIs overlap logically.

    If two ROIs cover the same pixel, they are considered to overlap.
    This does not consider chunking or other storage details. Measured
    on the *read* regions — with a halo, neighbouring reads overlap by
    construction. For write safety, `check_if_write_units_overlap` is
    the right question.

    Returns:
        `True` if any ROIs overlap.
    """
    if len(self.rois) < 2:
        return False

    slicing_tuples = (
        g.slicing_ops.normalized_slicing_tuple
        for g in self._numpy_getters_generator()
    )
    return check_if_regions_overlap(slicing_tuples)

require_no_regions_overlap

require_no_regions_overlap() -> None

Ensure that the Iterator's ROIs do not overlap.

Source code in src/ngio/iterators/_abstract_iterator.py
def require_no_regions_overlap(self) -> None:
    """Ensure that the Iterator's ROIs do not overlap."""
    if self.check_if_regions_overlap():
        raise NgioValueError("Some rois overlap.")