Iterators API reference¶
ImageProcessingIterator¶
ngio.iterators.ImageProcessingIterator
¶
ImageProcessingIterator(
input_image: Image,
output_image: Image,
*,
channel_selection: ChannelSlicingInputType = None,
output_channel_selection: ChannelSlicingInputType = None,
axes_order: Sequence[str] | None = None,
input_transforms: Sequence[TransformProtocol]
| None = None,
output_transforms: Sequence[TransformProtocol]
| None = None,
consolidation_mode: ConsolidationMode | None = None,
)
Bases: WritingIteratorBuilder[ndarray, Array, None]
Apply an image-to-image function region by region.
Reads each region from the input image, hands the patch to process's
function, and writes the result to the output image — a filter, a
projection, a restoration model. Compose with by_write_units()
for collision-free parallel writes and with_halo for seamless tiles.
Overlapping writes are safe; on_overlap(...) declares how contested
pixels resolve.
Process input_image region by region into output_image.
Parameters:
-
input_image(Image) –The image to read.
-
output_image(Image) –The image the results are written to.
-
channel_selection(ChannelSlicingInputType, default:None) –Restrict the input reads to these channels.
-
output_channel_selection(ChannelSlicingInputType, default:None) –Restrict the output writes to these channels.
-
axes_order(Sequence[str] | None, default:None) –Axes order of the patches handed to the function.
-
input_transforms(Sequence[TransformProtocol] | None, default:None) –Transforms applied to each input patch.
-
output_transforms(Sequence[TransformProtocol] | None, default:None) –Transforms applied to each patch before the write.
-
consolidation_mode(ConsolidationMode | None, default:None) –How to build the output pyramid after iteration, see
Image.consolidate.
Source code in src/ngio/iterators/_image_processing.py
partition_indices
property
¶
The unit indices this iterator is restricted to, or None.
None on an unrestricted iterator; on a for_job slice,
the sorted positions (in the full ROI list) of the units it runs.
The whole layout of a split is
[it.for_job(i, n).partition_indices for i in range(n)] —
the pre-flight check before submitting a distributed run: one fat
list plus empties means the output chunking, not the cluster, is
the constraint.
with_halo
¶
Read a margin of context around each ROI, but write only the ROI.
The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.
Note
The margins do not survive a rescale: a shape-changing transform
on the data path (a ZoomTransform) makes the crop refuse —
iterate on a coarser pyramid level instead. A MaskTransform
composes normally (its zoom rescales the label it reads).
Parameters:
-
x(int, default:0) –Pixels added on each side along x.
-
y(int, default:0) –Pixels added on each side along y.
-
z(int, default:0) –Pixels added on each side along z.
-
t(int, default:0) –Frames added on each side along t.
Returns:
-
Self–A new iterator reading with the halo.
Raises:
-
NgioValueError–On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps
roi_index/roi_nameso yourcoalescecan), on an in-place iterator, or for a negative margin.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
by_grid
¶
by_grid(
*,
size_x: int | None = None,
size_y: int | None = None,
size_z: int | None = None,
size_t: int | None = None,
stride_x: int | None = None,
stride_y: int | None = None,
stride_z: int | None = None,
stride_t: int | None = None,
tail: TailPolicy = "clip",
base_name: str = "",
) -> Self
Tile the current ROIs with a regular grid.
Parameters:
-
size_x(int | None, default:None) –Tile size along x;
None(default) means the full extent. -
size_y(int | None, default:None) –Tile size along y;
Nonemeans the full extent. -
size_z(int | None, default:None) –Tile size along z;
Nonemeans the full extent. -
size_t(int | None, default:None) –Tile size along t;
Nonemeans the full extent. -
stride_x(int | None, default:None) –Step between tiles along x;
Nonemeanssize_x(adjacent tiles). A smaller stride produces overlapping tiles. -
stride_y(int | None, default:None) –As
stride_x, along y. -
stride_z(int | None, default:None) –As
stride_x, along z. -
stride_t(int | None, default:None) –As
stride_x, along t. -
tail(TailPolicy, default:'clip') –What happens where the grid does not divide the axis.
"clip"(default) shrinks the last tile to the border;"balance"re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles,stride == size);"shift"moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution;"drop"discards the tail and leaves that border uncovered. -
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_blocks
¶
by_blocks(
*,
num_x: int = 1,
num_y: int = 1,
num_z: int = 1,
num_t: int = 1,
base_name: str = "",
) -> Self
Divide each axis into a fixed number of near-equal blocks.
The complement of by_grid: you say how many tiles, not how big.
Block lengths along an axis differ by at most one pixel, so there is
no tail to police — the partition is balanced by construction.
Parameters:
-
num_x(int, default:1) –Number of blocks along x (default 1: no split).
-
num_y(int, default:1) –Number of blocks along y.
-
num_z(int, default:1) –Number of blocks along z.
-
num_t(int, default:1) –Number of blocks along t.
-
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the blocks.
Source code in src/ngio/iterators/_abstract_iterator.py
by_yx
¶
Split each ROI into one 2D plane per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, z, c) combination, each covering the full y/x extent — the
shape for "run this 2D function on every plane".
Source code in src/ngio/iterators/_abstract_iterator.py
by_zyx
¶
Split each ROI into one 3D volume per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, c) combination, each covering the full z/y/x extent.
Parameters:
-
strict(bool, default:True) –Require a real z axis (present and larger than 1); with
False, images without one fall back to 2D planes.
Source code in src/ngio/iterators/_abstract_iterator.py
by_chunks
¶
by_chunks(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the input image's chunk grid.
Reads are chunk-granular, so this is the natural tiling for
read-heavy work. For collision-free parallel writes use
by_write_units, which tiles by the output's write granularity.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_write_units
¶
by_write_units(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the output image's write granularity.
The write unit is the shard shape when the output is sharded (writes
are atomic per shard object), the chunk shape otherwise. With no
overlap the resulting ROIs pass check_if_write_units_overlap by
construction, so a parallel map runs as a single fully-parallel
wave — a throughput optimization; conflicting tilings are still safe,
just scheduled in more waves. Falls back to the input chunk grid when
the iterator is read-only.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
product
¶
product(other: list[Roi] | GenericRoiTable) -> Self
Cartesian product of the current ROIs with an arbitrary list of ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
iter
¶
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[
DataGetterProtocol[NumpyPipeType],
DataSetterProtocol[NumpyPipeType],
]
]
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[
DataGetterProtocol[DaskPipeType],
DataSetterProtocol[DaskPipeType],
]
]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[NumpyPipeType, DataSetterProtocol[NumpyPipeType]]
]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
lazy: Literal[False],
data_mode: Literal["dask"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[DaskPipeType, DataSetterProtocol[DaskPipeType]]
]
iter(
lazy: Literal[False],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
lazy: Literal[False] = ...,
data_mode: Literal["numpy"] | None = ...,
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: int,
) -> Generator[
tuple[
list[NumpyPipeType],
list[DataSetterProtocol[NumpyPipeType]],
]
]
iter(
lazy: bool = False,
data_mode: Literal["numpy", "dask"] | None = None,
iterator_mode: Literal[
"readwrite", "readonly"
] = "readwrite",
*,
batch_size: int | None = None,
) -> Generator
Create an iterator over the pixels of the ROIs.
A writing loop finalizes only when fully drained — on a stitching
iterator that finalize is the resolve, so abandoning the loop
mid-way leaves block-offset ids on disk (plus the transient scratch)
until a re-run or a later finalize(). To write now and gather
later on purpose, iterate a for_job(0, 1) slice and call
finalize() on the unrestricted iterator when ready.
With batch_size set, a writing loop yields (patches, writers) —
two aligned lists of up to batch_size items, in ROI order — and a
read-only loop yields the payload lists alone. Stack the patches
yourself (raggedness is yours to handle, unlike BatchedMapper's
automatic padding), run the model once, and hand each result to its
writer. Batches follow ROI order — under write_order="roi" that
is the same later-ROI-wins order the mappers schedule, so the
manual loop is bit-identical to map. Batching is numpy-only
and eager: combining batch_size with lazy=True or
data_mode="dask" raises.
Parameters:
-
lazy(bool, default:False) –Yield getter/setter handles instead of materialized data.
-
data_mode(Literal['numpy', 'dask'] | None, default:None) –"numpy"or"dask"(deprecated, see note). -
iterator_mode(Literal['readwrite', 'readonly'], default:'readwrite') –"readwrite"yields(patch, writer)pairs,"readonly"yields patches alone. -
batch_size(int | None, default:None) –When set, yield batches of up to this many items; the last batch may be smaller.
Example
Note
The dask data mode is deprecated and will be removed in ngio=1.2;
from then on iter() yields numpy arrays.
Source code in src/ngio/iterators/_abstract_iterator.py
iter_as_numpy
¶
iter_as_dask
¶
Alias for iter(data_mode="dask").
Deprecated: removed in ngio=1.2.
Source code in src/ngio/iterators/_abstract_iterator.py
for_job
¶
Restrict the iterator to one of n_jobs independent partitions.
No write unit is ever shared between two partitions, so the jobs
need no locks, no barriers, and no channel to one another — safe as
separate SLURM array tasks, in any order. Every job must build an
identical iterator (construction is metadata-only and the partition
derives deterministically from it), apply for_job last in the
builder chain, and its verbs deliberately do not finalize — a slice's
map/process/segment writes only its share, and measure/
detect bank a partial instead of joining. The gather step is the
unrestricted iterator's finalize(), once, after all jobs.
Stitched and read-only runs also need prepare_jobs(n_jobs) first.
See the distributed-runs guide in the iterators docs.
Parameters:
-
job_index(int) –This job's partition,
0 <= job_index < n_jobs(e.g.int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops. -
n_jobs(int) –How many partitions the run is split into; must match across every job and the gather.
Returns:
-
Self–The restricted iterator; further reshaping refuses.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
prepare_jobs
¶
prepare_jobs(n_jobs: int) -> list[JobArgs]
Set up a distributed run and return its parallelization list.
The init step of the three-phase recipe (init → parallel jobs →
consolidate): performs the iterator's setup — always clearing stale
scratch state from earlier runs first — and returns one JobArgs
per non-empty partition. Optional for plain writers; a stitching or
read-only run requires it (the scratch root must exist, race-free,
before any job writes into it).
Parameters:
-
n_jobs(int) –How many partitions to split into; empty ones are dropped from the list.
Returns:
-
list[JobArgs]–One
{"job_index": ..., "n_jobs": ...}dict per non-empty -
list[JobArgs]–partition — JSON-ready, splatting into
for_job(**args).
Example
Source code in src/ngio/iterators/_abstract_iterator.py
reduce
¶
reduce(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run.
Parameters:
-
func(Callable[[NumpyPipeType], R]) –The function to apply; under a parallel mapper it must be safe on worker threads (or processes).
-
mapper(MapperProtocol[NumpyPipeType, R] | None, default:None) –How the units are scheduled.
None(the default) is serial; passThreadedMapper()orProcessMapper()to fan out, sized by their ownmax_workersargument.
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
reduce_as_numpy
¶
reduce_as_numpy(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Alias for reduce().
reduce_as_dask
¶
reduce_as_dask(
func: Callable[[DaskPipeType], R],
*,
mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Deprecated: removed in ngio=1.2.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run. Runs serially — the units hand
func lazy dask arrays, so the per-unit work is graph
construction, not IO. A parallel mapper raises (see map_as_dask).
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_regions_overlap
¶
Check if any of the ROIs overlap logically.
If two ROIs cover the same pixel, they are considered to overlap.
This does not consider chunking or other storage details. Measured
on the read regions — with a halo, neighbouring reads overlap by
construction. For write safety, check_if_write_units_overlap is
the right question.
Returns:
-
bool–Trueif any ROIs overlap.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_regions_overlap
¶
Ensure that the Iterator's ROIs do not overlap.
map
¶
map(
func: Callable[[NumpyPipeType], NumpyPipeType],
*,
mapper: MapperProtocol[NumpyPipeType, NumpyPipeType]
| None = None,
) -> None
Apply a transformation function to each ROI's patch and write it back.
Parameters:
-
func(Callable[[NumpyPipeType], NumpyPipeType]) –The transformation. Under a parallel mapper it runs on worker threads (or processes) and must be safe there.
-
mapper(MapperProtocol[NumpyPipeType, NumpyPipeType] | None, default:None) –How the units are scheduled.
None(the default) is serial — parallel writes stay explicit opt-in. PassThreadedMapper()orProcessMapper()to fan out; each sizes its own pool from itsmax_workersargument, and both schedule the units into conflict-free waves (plan_waves).
Source code in src/ngio/iterators/_abstract_iterator.py
map_as_numpy
¶
map_as_numpy(
func: Callable[[NumpyPipeType], NumpyPipeType],
*,
mapper: MapperProtocol[NumpyPipeType, NumpyPipeType]
| None = None,
) -> None
Alias for map().
Source code in src/ngio/iterators/_abstract_iterator.py
map_as_dask
¶
map_as_dask(
func: Callable[[DaskPipeType], DaskPipeType],
*,
mapper: MapperProtocol[DaskPipeType, DaskPipeType]
| None = None,
) -> None
Apply a transformation function to each ROI's patch and write it back.
Deprecated: removed in ngio=1.2.
Runs serially: the dask pipes are already executed by dask's own
scheduler, and the dask write path (store_dask) scopes a global
dask config option, so concurrent callers are unsafe. A parallel
mapper raises.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_write_units_overlap
¶
Check if any two ROIs write into the same write unit of the output.
Measured on the write target: slicing tuples come from the setters and the grid is the output array's write granularity — the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. Two ROIs sharing a write unit make concurrent writes unsafe: the read-modify-write of that unit can lose data. The parallel mappers schedule such ROIs into separate waves, so this check answers "will my map run as a single fully-parallel wave?".
This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a loop.
Returns:
-
bool–Trueif any two ROIs share a write unit.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_write_units_overlap
¶
Ensure that the ROIs do not share write units on the output.
The strict opt-in gate: the parallel mappers no longer refuse shared write units on their own (they wave-schedule around them), so call this to insist on a single-wave tiling instead.
Source code in src/ngio/iterators/_abstract_iterator.py
on_overlap
¶
on_overlap(
policy: OverlapPolicy,
*,
write_order: WriteOrder = "any",
) -> Self
Declare how contested write pixels are resolved.
Optional here: undeclared overlapping writes are last-writer-wins
in schedule order. "last" declares exactly that; any merge=
policy ("max"/"min"/"sum" are order-independent;
"keep_nonzero"; a merge function; a MergePolicy) combines each
write with what is on disk — note it applies to every write, so
pre-existing content participates too. Declare before for_job.
Parameters:
-
policy(OverlapPolicy) –"last", or anything the write path'smerge=takes. -
write_order(WriteOrder, default:'any') –Who wins a contested pixel.
"any"(the default) is schedule-defined — deterministic per version, fastest."roi"makes the later ROI win — reproducible across versions and bit-identical to the manualiterloop, at up to 2.5x cost on parallel overlapping tilings. Irrelevant under an order-independent merge.
Source code in src/ngio/iterators/_image_processing.py
build_numpy_getter
¶
build_numpy_getter(roi: Roi) -> DataGetterProtocol[ndarray]
Source code in src/ngio/iterators/_image_processing.py
build_numpy_setter
¶
build_numpy_setter(roi: Roi) -> DataSetterProtocol[ndarray]
Source code in src/ngio/iterators/_image_processing.py
build_dask_getter
¶
build_dask_getter(roi: Roi) -> DataGetterProtocol[Array]
Source code in src/ngio/iterators/_image_processing.py
build_dask_setter
¶
build_dask_setter(roi: Roi) -> DataSetterProtocol[Array]
Source code in src/ngio/iterators/_image_processing.py
process
¶
process(
func: Callable[[ndarray], ndarray],
*,
mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None
Process every region and write the results; the topic verb for map.
A serial run finalizes automatically. On a for_job slice it
processes only this job's share; the gather is the unrestricted
iterator's finalize(), once, after all jobs.
Parameters:
-
func(Callable[[ndarray], ndarray]) –The transformation. Under a parallel mapper it runs on worker threads (or processes) and must be safe there.
-
mapper(MapperProtocol[ndarray, ndarray] | None, default:None) –How the units are scheduled; see
map.
Source code in src/ngio/iterators/_image_processing.py
finalize
¶
Consolidate the output pyramid, only under the written regions.
Source code in src/ngio/iterators/_image_processing.py
SegmentationIterator¶
ngio.iterators.SegmentationIterator
¶
SegmentationIterator(
input_image: Image,
output_label: Label,
*,
channel_selection: ChannelSlicingInputType = None,
axes_order: Sequence[str] | None = None,
input_transforms: Sequence[TransformProtocol]
| None = None,
output_transforms: Sequence[TransformProtocol]
| None = None,
consolidation_mode: ConsolidationMode | None = None,
)
Bases: WritingIteratorBuilder[ndarray, Array, None]
Segment an image region by region into a label.
Reads each region from the input image and writes the function's label
patch to the output label. With with_stitch(...) objects split across
region boundaries are resolved into one id at the gather — any ROI list
works: grids with a halo, overlapping FOV layouts, ragged tables. See
StitchConfig. Regions whose write footprints overlap need a declared
resolution — with_stitch(...) or on_overlap(...) — or the writing
verbs refuse.
Segment input_image region by region into output_label.
Parameters:
-
input_image(Image) –The image to segment.
-
output_label(Label) –The label the segmentation is written to.
-
channel_selection(ChannelSlicingInputType, default:None) –Restrict the image reads to these channels.
-
axes_order(Sequence[str] | None, default:None) –Axes order of the patches handed to the function.
-
input_transforms(Sequence[TransformProtocol] | None, default:None) –Transforms applied to each image patch.
-
output_transforms(Sequence[TransformProtocol] | None, default:None) –Transforms applied to each predicted patch before the write.
-
consolidation_mode(ConsolidationMode | None, default:None) –How to build the output pyramid after iteration, see
Label.consolidate.
Source code in src/ngio/iterators/_segmentation.py
partition_indices
property
¶
The unit indices this iterator is restricted to, or None.
None on an unrestricted iterator; on a for_job slice,
the sorted positions (in the full ROI list) of the units it runs.
The whole layout of a split is
[it.for_job(i, n).partition_indices for i in range(n)] —
the pre-flight check before submitting a distributed run: one fat
list plus empties means the output chunking, not the cluster, is
the constraint.
with_halo
¶
Read a margin of context around each ROI, but write only the ROI.
The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.
Note
The margins do not survive a rescale: a shape-changing transform
on the data path (a ZoomTransform) makes the crop refuse —
iterate on a coarser pyramid level instead. A MaskTransform
composes normally (its zoom rescales the label it reads).
Parameters:
-
x(int, default:0) –Pixels added on each side along x.
-
y(int, default:0) –Pixels added on each side along y.
-
z(int, default:0) –Pixels added on each side along z.
-
t(int, default:0) –Frames added on each side along t.
Returns:
-
Self–A new iterator reading with the halo.
Raises:
-
NgioValueError–On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps
roi_index/roi_nameso yourcoalescecan), on an in-place iterator, or for a negative margin.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
by_grid
¶
by_grid(
*,
size_x: int | None = None,
size_y: int | None = None,
size_z: int | None = None,
size_t: int | None = None,
stride_x: int | None = None,
stride_y: int | None = None,
stride_z: int | None = None,
stride_t: int | None = None,
tail: TailPolicy = "clip",
base_name: str = "",
) -> Self
Tile the current ROIs with a regular grid.
Parameters:
-
size_x(int | None, default:None) –Tile size along x;
None(default) means the full extent. -
size_y(int | None, default:None) –Tile size along y;
Nonemeans the full extent. -
size_z(int | None, default:None) –Tile size along z;
Nonemeans the full extent. -
size_t(int | None, default:None) –Tile size along t;
Nonemeans the full extent. -
stride_x(int | None, default:None) –Step between tiles along x;
Nonemeanssize_x(adjacent tiles). A smaller stride produces overlapping tiles. -
stride_y(int | None, default:None) –As
stride_x, along y. -
stride_z(int | None, default:None) –As
stride_x, along z. -
stride_t(int | None, default:None) –As
stride_x, along t. -
tail(TailPolicy, default:'clip') –What happens where the grid does not divide the axis.
"clip"(default) shrinks the last tile to the border;"balance"re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles,stride == size);"shift"moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution;"drop"discards the tail and leaves that border uncovered. -
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_blocks
¶
by_blocks(
*,
num_x: int = 1,
num_y: int = 1,
num_z: int = 1,
num_t: int = 1,
base_name: str = "",
) -> Self
Divide each axis into a fixed number of near-equal blocks.
The complement of by_grid: you say how many tiles, not how big.
Block lengths along an axis differ by at most one pixel, so there is
no tail to police — the partition is balanced by construction.
Parameters:
-
num_x(int, default:1) –Number of blocks along x (default 1: no split).
-
num_y(int, default:1) –Number of blocks along y.
-
num_z(int, default:1) –Number of blocks along z.
-
num_t(int, default:1) –Number of blocks along t.
-
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the blocks.
Source code in src/ngio/iterators/_abstract_iterator.py
by_yx
¶
Split each ROI into one 2D plane per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, z, c) combination, each covering the full y/x extent — the
shape for "run this 2D function on every plane".
Source code in src/ngio/iterators/_abstract_iterator.py
by_zyx
¶
Split each ROI into one 3D volume per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, c) combination, each covering the full z/y/x extent.
Parameters:
-
strict(bool, default:True) –Require a real z axis (present and larger than 1); with
False, images without one fall back to 2D planes.
Source code in src/ngio/iterators/_abstract_iterator.py
by_chunks
¶
by_chunks(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the input image's chunk grid.
Reads are chunk-granular, so this is the natural tiling for
read-heavy work. For collision-free parallel writes use
by_write_units, which tiles by the output's write granularity.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_write_units
¶
by_write_units(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the output image's write granularity.
The write unit is the shard shape when the output is sharded (writes
are atomic per shard object), the chunk shape otherwise. With no
overlap the resulting ROIs pass check_if_write_units_overlap by
construction, so a parallel map runs as a single fully-parallel
wave — a throughput optimization; conflicting tilings are still safe,
just scheduled in more waves. Falls back to the input chunk grid when
the iterator is read-only.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
product
¶
product(other: list[Roi] | GenericRoiTable) -> Self
Cartesian product of the current ROIs with an arbitrary list of ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
iter
¶
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[
DataGetterProtocol[NumpyPipeType],
DataSetterProtocol[NumpyPipeType],
]
]
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[
DataGetterProtocol[DaskPipeType],
DataSetterProtocol[DaskPipeType],
]
]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[NumpyPipeType, DataSetterProtocol[NumpyPipeType]]
]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
lazy: Literal[False],
data_mode: Literal["dask"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[DaskPipeType, DataSetterProtocol[DaskPipeType]]
]
iter(
lazy: Literal[False],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
lazy: Literal[False] = ...,
data_mode: Literal["numpy"] | None = ...,
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: int,
) -> Generator[
tuple[
list[NumpyPipeType],
list[DataSetterProtocol[NumpyPipeType]],
]
]
iter(
lazy: bool = False,
data_mode: Literal["numpy", "dask"] | None = None,
iterator_mode: Literal[
"readwrite", "readonly"
] = "readwrite",
*,
batch_size: int | None = None,
) -> Generator
Create an iterator over the pixels of the ROIs.
A writing loop finalizes only when fully drained — on a stitching
iterator that finalize is the resolve, so abandoning the loop
mid-way leaves block-offset ids on disk (plus the transient scratch)
until a re-run or a later finalize(). To write now and gather
later on purpose, iterate a for_job(0, 1) slice and call
finalize() on the unrestricted iterator when ready.
With batch_size set, a writing loop yields (patches, writers) —
two aligned lists of up to batch_size items, in ROI order — and a
read-only loop yields the payload lists alone. Stack the patches
yourself (raggedness is yours to handle, unlike BatchedMapper's
automatic padding), run the model once, and hand each result to its
writer. Batches follow ROI order — under write_order="roi" that
is the same later-ROI-wins order the mappers schedule, so the
manual loop is bit-identical to map. Batching is numpy-only
and eager: combining batch_size with lazy=True or
data_mode="dask" raises.
Parameters:
-
lazy(bool, default:False) –Yield getter/setter handles instead of materialized data.
-
data_mode(Literal['numpy', 'dask'] | None, default:None) –"numpy"or"dask"(deprecated, see note). -
iterator_mode(Literal['readwrite', 'readonly'], default:'readwrite') –"readwrite"yields(patch, writer)pairs,"readonly"yields patches alone. -
batch_size(int | None, default:None) –When set, yield batches of up to this many items; the last batch may be smaller.
Example
Note
The dask data mode is deprecated and will be removed in ngio=1.2;
from then on iter() yields numpy arrays.
Source code in src/ngio/iterators/_abstract_iterator.py
iter_as_numpy
¶
iter_as_dask
¶
Alias for iter(data_mode="dask").
Deprecated: removed in ngio=1.2.
Source code in src/ngio/iterators/_abstract_iterator.py
for_job
¶
Restrict the iterator to one of n_jobs independent partitions.
No write unit is ever shared between two partitions, so the jobs
need no locks, no barriers, and no channel to one another — safe as
separate SLURM array tasks, in any order. Every job must build an
identical iterator (construction is metadata-only and the partition
derives deterministically from it), apply for_job last in the
builder chain, and its verbs deliberately do not finalize — a slice's
map/process/segment writes only its share, and measure/
detect bank a partial instead of joining. The gather step is the
unrestricted iterator's finalize(), once, after all jobs.
Stitched and read-only runs also need prepare_jobs(n_jobs) first.
See the distributed-runs guide in the iterators docs.
Parameters:
-
job_index(int) –This job's partition,
0 <= job_index < n_jobs(e.g.int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops. -
n_jobs(int) –How many partitions the run is split into; must match across every job and the gather.
Returns:
-
Self–The restricted iterator; further reshaping refuses.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
prepare_jobs
¶
prepare_jobs(n_jobs: int) -> list[JobArgs]
Set up a distributed run and return its parallelization list.
The init step of the three-phase recipe (init → parallel jobs →
consolidate): performs the iterator's setup — always clearing stale
scratch state from earlier runs first — and returns one JobArgs
per non-empty partition. Optional for plain writers; a stitching or
read-only run requires it (the scratch root must exist, race-free,
before any job writes into it).
Parameters:
-
n_jobs(int) –How many partitions to split into; empty ones are dropped from the list.
Returns:
-
list[JobArgs]–One
{"job_index": ..., "n_jobs": ...}dict per non-empty -
list[JobArgs]–partition — JSON-ready, splatting into
for_job(**args).
Example
Source code in src/ngio/iterators/_abstract_iterator.py
reduce
¶
reduce(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run.
Parameters:
-
func(Callable[[NumpyPipeType], R]) –The function to apply; under a parallel mapper it must be safe on worker threads (or processes).
-
mapper(MapperProtocol[NumpyPipeType, R] | None, default:None) –How the units are scheduled.
None(the default) is serial; passThreadedMapper()orProcessMapper()to fan out, sized by their ownmax_workersargument.
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
reduce_as_numpy
¶
reduce_as_numpy(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Alias for reduce().
reduce_as_dask
¶
reduce_as_dask(
func: Callable[[DaskPipeType], R],
*,
mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Deprecated: removed in ngio=1.2.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run. Runs serially — the units hand
func lazy dask arrays, so the per-unit work is graph
construction, not IO. A parallel mapper raises (see map_as_dask).
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_regions_overlap
¶
Check if any of the ROIs overlap logically.
If two ROIs cover the same pixel, they are considered to overlap.
This does not consider chunking or other storage details. Measured
on the read regions — with a halo, neighbouring reads overlap by
construction. For write safety, check_if_write_units_overlap is
the right question.
Returns:
-
bool–Trueif any ROIs overlap.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_regions_overlap
¶
Ensure that the Iterator's ROIs do not overlap.
map_as_numpy
¶
map_as_numpy(
func: Callable[[NumpyPipeType], NumpyPipeType],
*,
mapper: MapperProtocol[NumpyPipeType, NumpyPipeType]
| None = None,
) -> None
Alias for map().
Source code in src/ngio/iterators/_abstract_iterator.py
map_as_dask
¶
map_as_dask(
func: Callable[[DaskPipeType], DaskPipeType],
*,
mapper: MapperProtocol[DaskPipeType, DaskPipeType]
| None = None,
) -> None
Apply a transformation function to each ROI's patch and write it back.
Deprecated: removed in ngio=1.2.
Runs serially: the dask pipes are already executed by dask's own
scheduler, and the dask write path (store_dask) scopes a global
dask config option, so concurrent callers are unsafe. A parallel
mapper raises.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_write_units_overlap
¶
Check if any two ROIs write into the same write unit of the output.
Measured on the write target: slicing tuples come from the setters and the grid is the output array's write granularity — the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. Two ROIs sharing a write unit make concurrent writes unsafe: the read-modify-write of that unit can lose data. The parallel mappers schedule such ROIs into separate waves, so this check answers "will my map run as a single fully-parallel wave?".
This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a loop.
Returns:
-
bool–Trueif any two ROIs share a write unit.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_write_units_overlap
¶
Ensure that the ROIs do not share write units on the output.
The strict opt-in gate: the parallel mappers no longer refuse shared write units on their own (they wave-schedule around them), so call this to insist on a single-wave tiling instead.
Source code in src/ngio/iterators/_abstract_iterator.py
with_stitch
¶
with_stitch(config: StitchConfig | None = None) -> Self
Declare the stitch: split objects become one id at the gather.
Objects cut by a region boundary are resolved into one id when the
run finalizes — any ROI list works: grids with a halo, overlapping
FOV layouts, ragged tables. Needs a halo (with_halo) or
overlapping ROIs: the evidence is overlap between neighbouring
predictions. Declare before for_job; None uses the
StitchConfig defaults. Refuses when on_overlap is declared —
a merge policy would make the disk diverge from the banked
predictions, so the resolve would relabel garbage.
Parameters:
-
config(StitchConfig | None, default:None) –The stitch configuration,
Nonefor the defaults.
Source code in src/ngio/iterators/_segmentation.py
on_overlap
¶
on_overlap(
policy: OverlapPolicy,
*,
write_order: WriteOrder = "any",
) -> Self
Declare how contested write pixels are resolved.
"last" is last-writer-wins — an explicit acknowledgment, no merge
is performed. Any merge= policy ("max"/"min"/"sum" are
order-independent; "keep_nonzero"; a merge function; a
MergePolicy) combines each write with what is on disk — note it
applies to every write, so pre-existing content participates
too. Declare before for_job. Refuses when with_stitch is
declared (the stitch owns the contested pixels).
Parameters:
-
policy(OverlapPolicy) –"last", or anything the write path'smerge=takes. -
write_order(WriteOrder, default:'any') –Who wins a contested pixel.
"any"(the default) is schedule-defined — deterministic per version, fastest."roi"makes the later ROI win — reproducible across versions and bit-identical to the manualiterloop, at up to 2.5x cost on parallel overlapping tilings. Irrelevant under an order-independent merge.
Source code in src/ngio/iterators/_segmentation.py
build_numpy_getter
¶
build_numpy_getter(roi: Roi) -> DataGetterProtocol[ndarray]
Source code in src/ngio/iterators/_segmentation.py
build_numpy_setter
¶
build_numpy_setter(roi: Roi) -> DataSetterProtocol[ndarray]
Source code in src/ngio/iterators/_segmentation.py
build_dask_getter
¶
build_dask_getter(roi: Roi) -> DataGetterProtocol[Array]
Source code in src/ngio/iterators/_segmentation.py
build_dask_setter
¶
build_dask_setter(roi: Roi) -> DataSetterProtocol[Array]
Source code in src/ngio/iterators/_segmentation.py
map
¶
map(
func: Callable[[ndarray], ndarray],
*,
mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None
See WritingIteratorBuilder.map; also cleans up on failure.
A failed standalone run cannot be resolved, so the stitch scratch
arrays are deleted rather than left as a stray _ngio_stitch group
beside the resolution levels. Cleanup happens only when this run
created the scratch: a partition slice, a resumed run, or the
gather step opened a prepared root that holds the banks every other
job wrote, and one failure must not destroy them — re-running is
idempotent (banks rewrite, the id offsets are derived, not counted).
The already-written tiles stay in every case.
Source code in src/ngio/iterators/_segmentation.py
segment
¶
segment(
func: Callable[[ndarray], ndarray],
*,
mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None
Segment every region and write the labels; the topic verb for map.
A serial run finalizes automatically (the stitch resolve included).
On a for_job slice it segments only this job's share; the gather is
the unrestricted iterator's finalize(), once, after all jobs.
Parameters:
-
func(Callable[[ndarray], ndarray]) –The segmentation model. Under a parallel mapper it runs on worker threads (or processes) and must be safe there.
-
mapper(MapperProtocol[ndarray, ndarray] | None, default:None) –How the units are scheduled; see
map.
Source code in src/ngio/iterators/_segmentation.py
finalize
¶
Resolve the stitch (when configured), then consolidate the pyramid.
Source code in src/ngio/iterators/_segmentation.py
MaskedSegmentationIterator¶
ngio.iterators.MaskedSegmentationIterator
¶
MaskedSegmentationIterator(
input_image: MaskedImage,
output_label: Label,
*,
channel_selection: ChannelSlicingInputType = None,
axes_order: Sequence[str] | None = None,
input_transforms: Sequence[TransformProtocol]
| None = None,
output_transforms: Sequence[TransformProtocol]
| None = None,
consolidation_mode: ConsolidationMode | None = None,
)
Bases: SegmentationIterator
Segment each object of a masking ROI table, inside its own mask.
Regions come from the masking table's per-object bounding boxes; reads
are masked to the object (outside pixels filled) and writes protect
everything outside it (MaskMerge) — overlapping bounding boxes are
safe, because each write only touches its own object's pixels. With
with_stitch(...) and a tiling (by_grid + with_halo), sub-objects
split by a tile boundary within one mask merge into one id; tiles of
different masks are never compared (an object cannot span two masks),
and ids come out unique and dense across every object — no
UniqueLabelsTransform needed (combining it with stitch raises).
Mind block_size: the id ceiling scales with objects times tiles per
object.
Segment each masked object's box from input_image into output_label.
The ROIs come from input_image's masking ROI table: one per object,
pixels outside the object masked on read and protected on write.
Parameters:
-
input_image(MaskedImage) –The masked image to segment; its masking table supplies the per-object ROIs.
-
output_label(Label) –The label the segmentation is written to.
-
channel_selection(ChannelSlicingInputType, default:None) –Restrict the image reads to these channels.
-
axes_order(Sequence[str] | None, default:None) –Axes order of the patches handed to the function.
-
input_transforms(Sequence[TransformProtocol] | None, default:None) –Transforms applied to each image patch, before the mask fill.
-
output_transforms(Sequence[TransformProtocol] | None, default:None) –Transforms applied to each predicted patch before the write.
-
consolidation_mode(ConsolidationMode | None, default:None) –How to build the output pyramid after iteration, see
Label.consolidate.
Source code in src/ngio/iterators/_segmentation.py
partition_indices
property
¶
The unit indices this iterator is restricted to, or None.
None on an unrestricted iterator; on a for_job slice,
the sorted positions (in the full ROI list) of the units it runs.
The whole layout of a split is
[it.for_job(i, n).partition_indices for i in range(n)] —
the pre-flight check before submitting a distributed run: one fat
list plus empties means the output chunking, not the cluster, is
the constraint.
with_halo
¶
Read a margin of context around each ROI, but write only the ROI.
The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.
Note
The margins do not survive a rescale: a shape-changing transform
on the data path (a ZoomTransform) makes the crop refuse —
iterate on a coarser pyramid level instead. A MaskTransform
composes normally (its zoom rescales the label it reads).
Parameters:
-
x(int, default:0) –Pixels added on each side along x.
-
y(int, default:0) –Pixels added on each side along y.
-
z(int, default:0) –Pixels added on each side along z.
-
t(int, default:0) –Frames added on each side along t.
Returns:
-
Self–A new iterator reading with the halo.
Raises:
-
NgioValueError–On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps
roi_index/roi_nameso yourcoalescecan), on an in-place iterator, or for a negative margin.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
by_grid
¶
by_grid(
*,
size_x: int | None = None,
size_y: int | None = None,
size_z: int | None = None,
size_t: int | None = None,
stride_x: int | None = None,
stride_y: int | None = None,
stride_z: int | None = None,
stride_t: int | None = None,
tail: TailPolicy = "clip",
base_name: str = "",
) -> Self
Tile the current ROIs with a regular grid.
Parameters:
-
size_x(int | None, default:None) –Tile size along x;
None(default) means the full extent. -
size_y(int | None, default:None) –Tile size along y;
Nonemeans the full extent. -
size_z(int | None, default:None) –Tile size along z;
Nonemeans the full extent. -
size_t(int | None, default:None) –Tile size along t;
Nonemeans the full extent. -
stride_x(int | None, default:None) –Step between tiles along x;
Nonemeanssize_x(adjacent tiles). A smaller stride produces overlapping tiles. -
stride_y(int | None, default:None) –As
stride_x, along y. -
stride_z(int | None, default:None) –As
stride_x, along z. -
stride_t(int | None, default:None) –As
stride_x, along t. -
tail(TailPolicy, default:'clip') –What happens where the grid does not divide the axis.
"clip"(default) shrinks the last tile to the border;"balance"re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles,stride == size);"shift"moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution;"drop"discards the tail and leaves that border uncovered. -
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_blocks
¶
by_blocks(
*,
num_x: int = 1,
num_y: int = 1,
num_z: int = 1,
num_t: int = 1,
base_name: str = "",
) -> Self
Divide each axis into a fixed number of near-equal blocks.
The complement of by_grid: you say how many tiles, not how big.
Block lengths along an axis differ by at most one pixel, so there is
no tail to police — the partition is balanced by construction.
Parameters:
-
num_x(int, default:1) –Number of blocks along x (default 1: no split).
-
num_y(int, default:1) –Number of blocks along y.
-
num_z(int, default:1) –Number of blocks along z.
-
num_t(int, default:1) –Number of blocks along t.
-
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the blocks.
Source code in src/ngio/iterators/_abstract_iterator.py
by_yx
¶
Split each ROI into one 2D plane per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, z, c) combination, each covering the full y/x extent — the
shape for "run this 2D function on every plane".
Source code in src/ngio/iterators/_abstract_iterator.py
by_zyx
¶
Split each ROI into one 3D volume per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, c) combination, each covering the full z/y/x extent.
Parameters:
-
strict(bool, default:True) –Require a real z axis (present and larger than 1); with
False, images without one fall back to 2D planes.
Source code in src/ngio/iterators/_abstract_iterator.py
by_chunks
¶
by_chunks(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the input image's chunk grid.
Reads are chunk-granular, so this is the natural tiling for
read-heavy work. For collision-free parallel writes use
by_write_units, which tiles by the output's write granularity.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_write_units
¶
by_write_units(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the output image's write granularity.
The write unit is the shard shape when the output is sharded (writes
are atomic per shard object), the chunk shape otherwise. With no
overlap the resulting ROIs pass check_if_write_units_overlap by
construction, so a parallel map runs as a single fully-parallel
wave — a throughput optimization; conflicting tilings are still safe,
just scheduled in more waves. Falls back to the input chunk grid when
the iterator is read-only.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
product
¶
product(other: list[Roi] | GenericRoiTable) -> Self
Cartesian product of the current ROIs with an arbitrary list of ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
finalize
¶
Resolve the stitch (when configured), then consolidate the pyramid.
Source code in src/ngio/iterators/_segmentation.py
iter
¶
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[
DataGetterProtocol[NumpyPipeType],
DataSetterProtocol[NumpyPipeType],
]
]
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[
DataGetterProtocol[DaskPipeType],
DataSetterProtocol[DaskPipeType],
]
]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[NumpyPipeType, DataSetterProtocol[NumpyPipeType]]
]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
lazy: Literal[False],
data_mode: Literal["dask"],
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: None = ...,
) -> Generator[
tuple[DaskPipeType, DataSetterProtocol[DaskPipeType]]
]
iter(
lazy: Literal[False],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"],
*,
batch_size: None = ...,
) -> Generator[DaskPipeType]
iter(
lazy: Literal[False] = ...,
data_mode: Literal["numpy"] | None = ...,
iterator_mode: Literal["readwrite"] = ...,
*,
batch_size: int,
) -> Generator[
tuple[
list[NumpyPipeType],
list[DataSetterProtocol[NumpyPipeType]],
]
]
iter(
lazy: bool = False,
data_mode: Literal["numpy", "dask"] | None = None,
iterator_mode: Literal[
"readwrite", "readonly"
] = "readwrite",
*,
batch_size: int | None = None,
) -> Generator
Create an iterator over the pixels of the ROIs.
A writing loop finalizes only when fully drained — on a stitching
iterator that finalize is the resolve, so abandoning the loop
mid-way leaves block-offset ids on disk (plus the transient scratch)
until a re-run or a later finalize(). To write now and gather
later on purpose, iterate a for_job(0, 1) slice and call
finalize() on the unrestricted iterator when ready.
With batch_size set, a writing loop yields (patches, writers) —
two aligned lists of up to batch_size items, in ROI order — and a
read-only loop yields the payload lists alone. Stack the patches
yourself (raggedness is yours to handle, unlike BatchedMapper's
automatic padding), run the model once, and hand each result to its
writer. Batches follow ROI order — under write_order="roi" that
is the same later-ROI-wins order the mappers schedule, so the
manual loop is bit-identical to map. Batching is numpy-only
and eager: combining batch_size with lazy=True or
data_mode="dask" raises.
Parameters:
-
lazy(bool, default:False) –Yield getter/setter handles instead of materialized data.
-
data_mode(Literal['numpy', 'dask'] | None, default:None) –"numpy"or"dask"(deprecated, see note). -
iterator_mode(Literal['readwrite', 'readonly'], default:'readwrite') –"readwrite"yields(patch, writer)pairs,"readonly"yields patches alone. -
batch_size(int | None, default:None) –When set, yield batches of up to this many items; the last batch may be smaller.
Example
Note
The dask data mode is deprecated and will be removed in ngio=1.2;
from then on iter() yields numpy arrays.
Source code in src/ngio/iterators/_abstract_iterator.py
iter_as_numpy
¶
iter_as_dask
¶
Alias for iter(data_mode="dask").
Deprecated: removed in ngio=1.2.
Source code in src/ngio/iterators/_abstract_iterator.py
for_job
¶
Restrict the iterator to one of n_jobs independent partitions.
No write unit is ever shared between two partitions, so the jobs
need no locks, no barriers, and no channel to one another — safe as
separate SLURM array tasks, in any order. Every job must build an
identical iterator (construction is metadata-only and the partition
derives deterministically from it), apply for_job last in the
builder chain, and its verbs deliberately do not finalize — a slice's
map/process/segment writes only its share, and measure/
detect bank a partial instead of joining. The gather step is the
unrestricted iterator's finalize(), once, after all jobs.
Stitched and read-only runs also need prepare_jobs(n_jobs) first.
See the distributed-runs guide in the iterators docs.
Parameters:
-
job_index(int) –This job's partition,
0 <= job_index < n_jobs(e.g.int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops. -
n_jobs(int) –How many partitions the run is split into; must match across every job and the gather.
Returns:
-
Self–The restricted iterator; further reshaping refuses.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
prepare_jobs
¶
prepare_jobs(n_jobs: int) -> list[JobArgs]
Set up a distributed run and return its parallelization list.
The init step of the three-phase recipe (init → parallel jobs →
consolidate): performs the iterator's setup — always clearing stale
scratch state from earlier runs first — and returns one JobArgs
per non-empty partition. Optional for plain writers; a stitching or
read-only run requires it (the scratch root must exist, race-free,
before any job writes into it).
Parameters:
-
n_jobs(int) –How many partitions to split into; empty ones are dropped from the list.
Returns:
-
list[JobArgs]–One
{"job_index": ..., "n_jobs": ...}dict per non-empty -
list[JobArgs]–partition — JSON-ready, splatting into
for_job(**args).
Example
Source code in src/ngio/iterators/_abstract_iterator.py
reduce
¶
reduce(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run.
Parameters:
-
func(Callable[[NumpyPipeType], R]) –The function to apply; under a parallel mapper it must be safe on worker threads (or processes).
-
mapper(MapperProtocol[NumpyPipeType, R] | None, default:None) –How the units are scheduled.
None(the default) is serial; passThreadedMapper()orProcessMapper()to fan out, sized by their ownmax_workersargument.
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
reduce_as_numpy
¶
reduce_as_numpy(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Alias for reduce().
reduce_as_dask
¶
reduce_as_dask(
func: Callable[[DaskPipeType], R],
*,
mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Deprecated: removed in ngio=1.2.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run. Runs serially — the units hand
func lazy dask arrays, so the per-unit work is graph
construction, not IO. A parallel mapper raises (see map_as_dask).
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_regions_overlap
¶
Check if any of the ROIs overlap logically.
If two ROIs cover the same pixel, they are considered to overlap.
This does not consider chunking or other storage details. Measured
on the read regions — with a halo, neighbouring reads overlap by
construction. For write safety, check_if_write_units_overlap is
the right question.
Returns:
-
bool–Trueif any ROIs overlap.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_regions_overlap
¶
Ensure that the Iterator's ROIs do not overlap.
map
¶
map(
func: Callable[[ndarray], ndarray],
*,
mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None
See WritingIteratorBuilder.map; also cleans up on failure.
A failed standalone run cannot be resolved, so the stitch scratch
arrays are deleted rather than left as a stray _ngio_stitch group
beside the resolution levels. Cleanup happens only when this run
created the scratch: a partition slice, a resumed run, or the
gather step opened a prepared root that holds the banks every other
job wrote, and one failure must not destroy them — re-running is
idempotent (banks rewrite, the id offsets are derived, not counted).
The already-written tiles stay in every case.
Source code in src/ngio/iterators/_segmentation.py
map_as_numpy
¶
map_as_numpy(
func: Callable[[NumpyPipeType], NumpyPipeType],
*,
mapper: MapperProtocol[NumpyPipeType, NumpyPipeType]
| None = None,
) -> None
Alias for map().
Source code in src/ngio/iterators/_abstract_iterator.py
map_as_dask
¶
map_as_dask(
func: Callable[[DaskPipeType], DaskPipeType],
*,
mapper: MapperProtocol[DaskPipeType, DaskPipeType]
| None = None,
) -> None
Apply a transformation function to each ROI's patch and write it back.
Deprecated: removed in ngio=1.2.
Runs serially: the dask pipes are already executed by dask's own
scheduler, and the dask write path (store_dask) scopes a global
dask config option, so concurrent callers are unsafe. A parallel
mapper raises.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_write_units_overlap
¶
Check if any two ROIs write into the same write unit of the output.
Measured on the write target: slicing tuples come from the setters and the grid is the output array's write granularity — the shard shape when the output is sharded (writes are atomic per shard object), the chunk shape otherwise. Two ROIs sharing a write unit make concurrent writes unsafe: the read-modify-write of that unit can lose data. The parallel mappers schedule such ROIs into separate waves, so this check answers "will my map run as a single fully-parallel wave?".
This is O(n^2) in the number of ROIs; avoid calling it repeatedly in a loop.
Returns:
-
bool–Trueif any two ROIs share a write unit.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_write_units_overlap
¶
Ensure that the ROIs do not share write units on the output.
The strict opt-in gate: the parallel mappers no longer refuse shared write units on their own (they wave-schedule around them), so call this to insist on a single-wave tiling instead.
Source code in src/ngio/iterators/_abstract_iterator.py
with_stitch
¶
with_stitch(config: StitchConfig | None = None) -> Self
Declare the stitch: split objects become one id at the gather.
Objects cut by a region boundary are resolved into one id when the
run finalizes — any ROI list works: grids with a halo, overlapping
FOV layouts, ragged tables. Needs a halo (with_halo) or
overlapping ROIs: the evidence is overlap between neighbouring
predictions. Declare before for_job; None uses the
StitchConfig defaults. Refuses when on_overlap is declared —
a merge policy would make the disk diverge from the banked
predictions, so the resolve would relabel garbage.
Parameters:
-
config(StitchConfig | None, default:None) –The stitch configuration,
Nonefor the defaults.
Source code in src/ngio/iterators/_segmentation.py
segment
¶
segment(
func: Callable[[ndarray], ndarray],
*,
mapper: MapperProtocol[ndarray, ndarray] | None = None,
) -> None
Segment every region and write the labels; the topic verb for map.
A serial run finalizes automatically (the stitch resolve included).
On a for_job slice it segments only this job's share; the gather is
the unrestricted iterator's finalize(), once, after all jobs.
Parameters:
-
func(Callable[[ndarray], ndarray]) –The segmentation model. Under a parallel mapper it runs on worker threads (or processes) and must be safe there.
-
mapper(MapperProtocol[ndarray, ndarray] | None, default:None) –How the units are scheduled; see
map.
Source code in src/ngio/iterators/_segmentation.py
on_overlap
¶
on_overlap(
policy: OverlapPolicy,
*,
write_order: WriteOrder = "roi",
) -> Self
Refused: masked writes never contest.
Each write touches only its own object's pixels (MaskMerge
protects the rest), so there is nothing to declare — and the one
merge= slot already carries the mask protection.
Source code in src/ngio/iterators/_segmentation.py
build_numpy_getter
¶
build_numpy_getter(roi: Roi)
Source code in src/ngio/iterators/_segmentation.py
build_numpy_setter
¶
build_numpy_setter(roi: Roi)
Source code in src/ngio/iterators/_segmentation.py
build_dask_setter
¶
build_dask_setter(roi: Roi)
Source code in src/ngio/iterators/_segmentation.py
FeatureExtractorIterator¶
ngio.iterators.FeatureExtractorIterator
¶
FeatureExtractorIterator(
input_image: Image,
input_label: Label,
*,
channel_selection: ChannelSlicingInputType = None,
axes_order: Sequence[str] | None = None,
input_transforms: Sequence[TransformProtocol]
| None = None,
label_transforms: Sequence[TransformProtocol]
| None = None,
)
Bases: AbstractIteratorBuilder[NumpyPipeType, DaskPipeType, Table]
Measure image/label pairs region by region; nothing is written.
Each unit pairs the image patch with the label patch over the same
region. reduce collects per-region results; measure joins them into
a single FeatureTable for the caller to store. Distributed, a
for_job slice's measure banks a partial and finalize() runs the
one global join (declared with with_join, ConcatJoin otherwise).
with_halo(...) reads context around each region; the duplicate rows
it produces for border objects are reconciled in your declared join
via the stamped roi_index/roi_name columns.
Measure input_image/input_label pairs; nothing is written.
Parameters:
-
input_image(Image) –The image to measure.
-
input_label(Label) –The label whose objects tie the measurements together.
-
channel_selection(ChannelSlicingInputType, default:None) –Restrict the image reads to these channels.
-
axes_order(Sequence[str] | None, default:None) –Axes order of the patches handed to the function.
-
input_transforms(Sequence[TransformProtocol] | None, default:None) –Transforms applied to each image patch.
-
label_transforms(Sequence[TransformProtocol] | None, default:None) –Transforms applied to each label patch.
Source code in src/ngio/iterators/_feature.py
output_image
property
¶
output_image: AbstractImage | None
The image this iterator writes to, or None for a read-only iterator.
partition_indices
property
¶
The unit indices this iterator is restricted to, or None.
None on an unrestricted iterator; on a for_job slice,
the sorted positions (in the full ROI list) of the units it runs.
The whole layout of a split is
[it.for_job(i, n).partition_indices for i in range(n)] —
the pre-flight check before submitting a distributed run: one fat
list plus empties means the output chunking, not the cluster, is
the constraint.
with_halo
¶
Read a margin of context around each ROI, but write only the ROI.
The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.
Note
The margins do not survive a rescale: a shape-changing transform
on the data path (a ZoomTransform) makes the crop refuse —
iterate on a coarser pyramid level instead. A MaskTransform
composes normally (its zoom rescales the label it reads).
Parameters:
-
x(int, default:0) –Pixels added on each side along x.
-
y(int, default:0) –Pixels added on each side along y.
-
z(int, default:0) –Pixels added on each side along z.
-
t(int, default:0) –Frames added on each side along t.
Returns:
-
Self–A new iterator reading with the halo.
Raises:
-
NgioValueError–On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps
roi_index/roi_nameso yourcoalescecan), on an in-place iterator, or for a negative margin.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
by_grid
¶
by_grid(
*,
size_x: int | None = None,
size_y: int | None = None,
size_z: int | None = None,
size_t: int | None = None,
stride_x: int | None = None,
stride_y: int | None = None,
stride_z: int | None = None,
stride_t: int | None = None,
tail: TailPolicy = "clip",
base_name: str = "",
) -> Self
Tile the current ROIs with a regular grid.
Parameters:
-
size_x(int | None, default:None) –Tile size along x;
None(default) means the full extent. -
size_y(int | None, default:None) –Tile size along y;
Nonemeans the full extent. -
size_z(int | None, default:None) –Tile size along z;
Nonemeans the full extent. -
size_t(int | None, default:None) –Tile size along t;
Nonemeans the full extent. -
stride_x(int | None, default:None) –Step between tiles along x;
Nonemeanssize_x(adjacent tiles). A smaller stride produces overlapping tiles. -
stride_y(int | None, default:None) –As
stride_x, along y. -
stride_z(int | None, default:None) –As
stride_x, along z. -
stride_t(int | None, default:None) –As
stride_x, along t. -
tail(TailPolicy, default:'clip') –What happens where the grid does not divide the axis.
"clip"(default) shrinks the last tile to the border;"balance"re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles,stride == size);"shift"moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution;"drop"discards the tail and leaves that border uncovered. -
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_blocks
¶
by_blocks(
*,
num_x: int = 1,
num_y: int = 1,
num_z: int = 1,
num_t: int = 1,
base_name: str = "",
) -> Self
Divide each axis into a fixed number of near-equal blocks.
The complement of by_grid: you say how many tiles, not how big.
Block lengths along an axis differ by at most one pixel, so there is
no tail to police — the partition is balanced by construction.
Parameters:
-
num_x(int, default:1) –Number of blocks along x (default 1: no split).
-
num_y(int, default:1) –Number of blocks along y.
-
num_z(int, default:1) –Number of blocks along z.
-
num_t(int, default:1) –Number of blocks along t.
-
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the blocks.
Source code in src/ngio/iterators/_abstract_iterator.py
by_yx
¶
Split each ROI into one 2D plane per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, z, c) combination, each covering the full y/x extent — the
shape for "run this 2D function on every plane".
Source code in src/ngio/iterators/_abstract_iterator.py
by_zyx
¶
Split each ROI into one 3D volume per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, c) combination, each covering the full z/y/x extent.
Parameters:
-
strict(bool, default:True) –Require a real z axis (present and larger than 1); with
False, images without one fall back to 2D planes.
Source code in src/ngio/iterators/_abstract_iterator.py
by_chunks
¶
by_chunks(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the input image's chunk grid.
Reads are chunk-granular, so this is the natural tiling for
read-heavy work. For collision-free parallel writes use
by_write_units, which tiles by the output's write granularity.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_write_units
¶
by_write_units(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the output image's write granularity.
The write unit is the shard shape when the output is sharded (writes
are atomic per shard object), the chunk shape otherwise. With no
overlap the resulting ROIs pass check_if_write_units_overlap by
construction, so a parallel map runs as a single fully-parallel
wave — a throughput optimization; conflicting tilings are still safe,
just scheduled in more waves. Falls back to the input chunk grid when
the iterator is read-only.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
product
¶
product(other: list[Roi] | GenericRoiTable) -> Self
Cartesian product of the current ROIs with an arbitrary list of ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
iter
¶
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
lazy: bool = False,
data_mode: Literal["numpy", "dask"] | None = None,
iterator_mode: Literal[
"readwrite", "readonly"
] = "readonly",
*,
batch_size: int | None = None,
) -> Generator
Iterate the ROIs' payloads, read-only.
With batch_size set, yields payload lists of up to that many
items (the last batch may be smaller). Batching is numpy-only and
eager: combining batch_size with lazy=True or
data_mode="dask" raises. On a writing iterator (see
WritingIteratorBuilder.iter) the default mode is "readwrite"
and the loop yields (patch, writer) pairs instead.
Parameters:
-
lazy(bool, default:False) –Yield getter handles instead of materialized data.
-
data_mode(Literal['numpy', 'dask'] | None, default:None) –"numpy"or"dask"(deprecated, see note). -
iterator_mode(Literal['readwrite', 'readonly'], default:'readonly') –"readonly"yields patches alone; on a writing iterator"readwrite"yields(patch, writer)pairs. -
batch_size(int | None, default:None) –When set, yield batches of up to this many items; the last batch may be smaller.
Note
The dask data mode is deprecated and will be removed in ngio=1.2;
from then on iter() yields numpy arrays.
Source code in src/ngio/iterators/_abstract_iterator.py
iter_as_numpy
¶
iter_as_dask
¶
Alias for iter(data_mode="dask").
Deprecated: removed in ngio=1.2.
Source code in src/ngio/iterators/_abstract_iterator.py
for_job
¶
Restrict the iterator to one of n_jobs independent partitions.
No write unit is ever shared between two partitions, so the jobs
need no locks, no barriers, and no channel to one another — safe as
separate SLURM array tasks, in any order. Every job must build an
identical iterator (construction is metadata-only and the partition
derives deterministically from it), apply for_job last in the
builder chain, and its verbs deliberately do not finalize — a slice's
map/process/segment writes only its share, and measure/
detect bank a partial instead of joining. The gather step is the
unrestricted iterator's finalize(), once, after all jobs.
Stitched and read-only runs also need prepare_jobs(n_jobs) first.
See the distributed-runs guide in the iterators docs.
Parameters:
-
job_index(int) –This job's partition,
0 <= job_index < n_jobs(e.g.int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops. -
n_jobs(int) –How many partitions the run is split into; must match across every job and the gather.
Returns:
-
Self–The restricted iterator; further reshaping refuses.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
prepare_jobs
¶
prepare_jobs(n_jobs: int) -> list[JobArgs]
Set up a distributed run and return its parallelization list.
The init step of the three-phase recipe (init → parallel jobs →
consolidate): performs the iterator's setup — always clearing stale
scratch state from earlier runs first — and returns one JobArgs
per non-empty partition. Optional for plain writers; a stitching or
read-only run requires it (the scratch root must exist, race-free,
before any job writes into it).
Parameters:
-
n_jobs(int) –How many partitions to split into; empty ones are dropped from the list.
Returns:
-
list[JobArgs]–One
{"job_index": ..., "n_jobs": ...}dict per non-empty -
list[JobArgs]–partition — JSON-ready, splatting into
for_job(**args).
Example
Source code in src/ngio/iterators/_abstract_iterator.py
reduce
¶
reduce(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run.
Parameters:
-
func(Callable[[NumpyPipeType], R]) –The function to apply; under a parallel mapper it must be safe on worker threads (or processes).
-
mapper(MapperProtocol[NumpyPipeType, R] | None, default:None) –How the units are scheduled.
None(the default) is serial; passThreadedMapper()orProcessMapper()to fan out, sized by their ownmax_workersargument.
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
reduce_as_numpy
¶
reduce_as_numpy(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Alias for reduce().
reduce_as_dask
¶
reduce_as_dask(
func: Callable[[DaskPipeType], R],
*,
mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Deprecated: removed in ngio=1.2.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run. Runs serially — the units hand
func lazy dask arrays, so the per-unit work is graph
construction, not IO. A parallel mapper raises (see map_as_dask).
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_regions_overlap
¶
Check if any of the ROIs overlap logically.
If two ROIs cover the same pixel, they are considered to overlap.
This does not consider chunking or other storage details. Measured
on the read regions — with a halo, neighbouring reads overlap by
construction. For write safety, check_if_write_units_overlap is
the right question.
Returns:
-
bool–Trueif any ROIs overlap.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_regions_overlap
¶
Ensure that the Iterator's ROIs do not overlap.
with_join
¶
with_join(join: JoinProtocol) -> Self
Declare the join that turns the per-ROI results into the table.
Runs once, wherever the join executes — a serial measure, or the
distributed gather's finalize. On a for_job slice the
declaration is inert: the slice banks the pre-join records
regardless, so declaring it there is harmless. Like the
measurement function, the join is not part of the plan
fingerprint: declare the identical one on the gather iterator.
The default (no declaration) is ConcatJoin referencing the
input label.
Parameters:
-
join(JoinProtocol) –Anything satisfying
JoinProtocol— a plain callable taking the normalized frame list, or a configured class likeConcatJoin(reference_label=...).
Source code in src/ngio/iterators/_feature.py
build_numpy_getter
¶
build_numpy_getter(roi: Roi) -> FeatureGetter[ndarray]
Source code in src/ngio/iterators/_feature.py
build_dask_getter
¶
build_dask_getter(roi: Roi) -> FeatureGetter[Array]
Source code in src/ngio/iterators/_feature.py
finalize
¶
finalize() -> Table
Merge a distributed run's partials into the one final table.
The gather step (see prepare_jobs): validates that every job of
the prepared plan produced a matching partial — a half-finished run
errors instead of returning a silently incomplete table — rebuilds
the per-ROI result list in global ROI order, and runs the single
declared join (with_join, ConcatJoin otherwise) exactly once.
The final table matches a serial measure row for row; the result
list is the normalized form described in measure,
roi_index/roi_name included. On success the partials group is
removed. Nothing is registered: the returned table is yours to
store with add_table.
Raises on a for_job slice (the gather is global) and when no
partials exist (nothing was prepared or banked).
Source code in src/ngio/iterators/_feature.py
measure
¶
measure(
func: Callable[
[ndarray, ndarray, Roi], FeatureFuncResult
],
*,
mapper: MapperProtocol | None = None,
) -> Table | None
Measure every ROI and join the results into one table.
The per-ROI measurement fans out exactly like reduce — pass a
mapper to parallelize it — and the declared join (with_join,
ConcatJoin otherwise) runs once, at the end, on the calling
thread. Nothing is written: the returned table is yours to store,
e.g. container.add_table(name, table).
The join (declared or default, serial or distributed) always sees
the normalized result list, not the function's raw returns: one
DataFrame per ROI with label as a column (a label index is
reset), every row stamped with roi_index (the ROI's global index)
and roi_name (the ROI's name — roi_{index} when it has none, as
seen by the job that measured it), and a ROI whose function
returned no rows as an empty, column-less DataFrame. The three
names _ngio_index, roi_index and roi_name are reserved: a
function whose result carries one is refused.
With with_halo(...) each region is read grown — both patches and
the roi argument cover the grown region — so a border object is
measured by every region that sees it. The default join keeps the
resulting duplicate label rows as-is; reconcile them in a declared
join via the roi_index/roi_name columns. reduce and iter
read the grown regions too.
On a for_job slice it measures only this job's share and banks the
normalized records as a partial, returning None — stored
before any join, so the gather (finalize(), once, after all
jobs) can rebuild the full per-ROI result list and run the ONE
global join. Re-running a job overwrites its own partial.
Parameters:
-
func(Callable[[ndarray, ndarray, Roi], FeatureFuncResult]) –(image, label, roi) -> DataFrame | dict[str, list]— the measurements for one ROI. Rows must carry the object id in alabelcolumn (or index). Under a parallel mapper it runs on worker threads or processes and must be safe there. -
mapper(MapperProtocol | None, default:None) –How the per-ROI work is scheduled;
Noneis serial.
Returns:
-
Table | None–The joined table — a
FeatureTableunder the default join. A -
Table | None–run that finds zero objects returns an empty table, as
detect -
Table | None–does. On a
for_jobslice:None(the partial is banked).
Source code in src/ngio/iterators/_feature.py
ObjectDetectionIterator¶
ngio.iterators.ObjectDetectionIterator
¶
ObjectDetectionIterator(
input_image: Image,
*,
channel_selection: ChannelSlicingInputType = None,
axes_order: Sequence[str] | None = None,
input_transforms: Sequence[TransformProtocol]
| None = None,
)
Bases: AbstractIteratorBuilder[NumpyPipeType, DaskPipeType, RoiTable]
Tile an image, detect per tile, return one deduplicated ROI table.
A read-only iterator: nothing is written to the image, and the halo is a
pure read margin — call with_halo(...) so an object cut by a tile edge
is seen whole by the neighbouring tile; the duplicate detections that
produces are resolved by NMS (declared with with_nms, the
GreedyNms defaults otherwise). The product is a
RoiTable of world-anchored boxes, returned by detect for the
caller to store. Distributed, a for_job slice's detect banks a
partial and finalize() runs the one global NMS.
The box contract: each detection is a Roi in the patch's own pixel
coordinates — Roi.from_values(slices={"x": (x0, w), "y": (y0, h)},
name=None, space="pixel", confidence=0.9). x and y required, z
optional (the same choice on every box); name/label are refused (the
iterator assigns them — put a class label in an extra field); every
extra field rides along into the table.
Initialize the iterator over input_image.
Parameters:
-
input_image(Image) –The image the detector runs on.
-
channel_selection(ChannelSlicingInputType, default:None) –Optional selection of channels to read.
-
axes_order(Sequence[str] | None, default:None) –Optional axes order for the patches handed to the detector.
-
input_transforms(Sequence[TransformProtocol] | None, default:None) –Optional transforms applied to each patch.
Source code in src/ngio/iterators/_object_detection.py
output_image
property
¶
output_image: AbstractImage | None
The image this iterator writes to, or None for a read-only iterator.
partition_indices
property
¶
The unit indices this iterator is restricted to, or None.
None on an unrestricted iterator; on a for_job slice,
the sorted positions (in the full ROI list) of the units it runs.
The whole layout of a split is
[it.for_job(i, n).partition_indices for i in range(n)] —
the pre-flight check before submitting a distributed run: one fat
list plus empties means the output chunking, not the cluster, is
the constraint.
with_halo
¶
Read a margin of context around each ROI, but write only the ROI.
The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.
Note
The margins do not survive a rescale: a shape-changing transform
on the data path (a ZoomTransform) makes the crop refuse —
iterate on a coarser pyramid level instead. A MaskTransform
composes normally (its zoom rescales the label it reads).
Parameters:
-
x(int, default:0) –Pixels added on each side along x.
-
y(int, default:0) –Pixels added on each side along y.
-
z(int, default:0) –Pixels added on each side along z.
-
t(int, default:0) –Frames added on each side along t.
Returns:
-
Self–A new iterator reading with the halo.
Raises:
-
NgioValueError–On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps
roi_index/roi_nameso yourcoalescecan), on an in-place iterator, or for a negative margin.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
by_grid
¶
by_grid(
*,
size_x: int | None = None,
size_y: int | None = None,
size_z: int | None = None,
size_t: int | None = None,
stride_x: int | None = None,
stride_y: int | None = None,
stride_z: int | None = None,
stride_t: int | None = None,
tail: TailPolicy = "clip",
base_name: str = "",
) -> Self
Tile the current ROIs with a regular grid.
Parameters:
-
size_x(int | None, default:None) –Tile size along x;
None(default) means the full extent. -
size_y(int | None, default:None) –Tile size along y;
Nonemeans the full extent. -
size_z(int | None, default:None) –Tile size along z;
Nonemeans the full extent. -
size_t(int | None, default:None) –Tile size along t;
Nonemeans the full extent. -
stride_x(int | None, default:None) –Step between tiles along x;
Nonemeanssize_x(adjacent tiles). A smaller stride produces overlapping tiles. -
stride_y(int | None, default:None) –As
stride_x, along y. -
stride_z(int | None, default:None) –As
stride_x, along z. -
stride_t(int | None, default:None) –As
stride_x, along t. -
tail(TailPolicy, default:'clip') –What happens where the grid does not divide the axis.
"clip"(default) shrinks the last tile to the border;"balance"re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles,stride == size);"shift"moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution;"drop"discards the tail and leaves that border uncovered. -
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_blocks
¶
by_blocks(
*,
num_x: int = 1,
num_y: int = 1,
num_z: int = 1,
num_t: int = 1,
base_name: str = "",
) -> Self
Divide each axis into a fixed number of near-equal blocks.
The complement of by_grid: you say how many tiles, not how big.
Block lengths along an axis differ by at most one pixel, so there is
no tail to police — the partition is balanced by construction.
Parameters:
-
num_x(int, default:1) –Number of blocks along x (default 1: no split).
-
num_y(int, default:1) –Number of blocks along y.
-
num_z(int, default:1) –Number of blocks along z.
-
num_t(int, default:1) –Number of blocks along t.
-
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the blocks.
Source code in src/ngio/iterators/_abstract_iterator.py
by_yx
¶
Split each ROI into one 2D plane per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, z, c) combination, each covering the full y/x extent — the
shape for "run this 2D function on every plane".
Source code in src/ngio/iterators/_abstract_iterator.py
by_zyx
¶
Split each ROI into one 3D volume per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, c) combination, each covering the full z/y/x extent.
Parameters:
-
strict(bool, default:True) –Require a real z axis (present and larger than 1); with
False, images without one fall back to 2D planes.
Source code in src/ngio/iterators/_abstract_iterator.py
by_chunks
¶
by_chunks(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the input image's chunk grid.
Reads are chunk-granular, so this is the natural tiling for
read-heavy work. For collision-free parallel writes use
by_write_units, which tiles by the output's write granularity.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_write_units
¶
by_write_units(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the output image's write granularity.
The write unit is the shard shape when the output is sharded (writes
are atomic per shard object), the chunk shape otherwise. With no
overlap the resulting ROIs pass check_if_write_units_overlap by
construction, so a parallel map runs as a single fully-parallel
wave — a throughput optimization; conflicting tilings are still safe,
just scheduled in more waves. Falls back to the input chunk grid when
the iterator is read-only.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
product
¶
product(other: list[Roi] | GenericRoiTable) -> Self
Cartesian product of the current ROIs with an arbitrary list of ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
iter
¶
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
lazy: bool = False,
data_mode: Literal["numpy", "dask"] | None = None,
iterator_mode: Literal[
"readwrite", "readonly"
] = "readonly",
*,
batch_size: int | None = None,
) -> Generator
Iterate the ROIs' payloads, read-only.
With batch_size set, yields payload lists of up to that many
items (the last batch may be smaller). Batching is numpy-only and
eager: combining batch_size with lazy=True or
data_mode="dask" raises. On a writing iterator (see
WritingIteratorBuilder.iter) the default mode is "readwrite"
and the loop yields (patch, writer) pairs instead.
Parameters:
-
lazy(bool, default:False) –Yield getter handles instead of materialized data.
-
data_mode(Literal['numpy', 'dask'] | None, default:None) –"numpy"or"dask"(deprecated, see note). -
iterator_mode(Literal['readwrite', 'readonly'], default:'readonly') –"readonly"yields patches alone; on a writing iterator"readwrite"yields(patch, writer)pairs. -
batch_size(int | None, default:None) –When set, yield batches of up to this many items; the last batch may be smaller.
Note
The dask data mode is deprecated and will be removed in ngio=1.2;
from then on iter() yields numpy arrays.
Source code in src/ngio/iterators/_abstract_iterator.py
iter_as_numpy
¶
iter_as_dask
¶
Alias for iter(data_mode="dask").
Deprecated: removed in ngio=1.2.
Source code in src/ngio/iterators/_abstract_iterator.py
for_job
¶
Restrict the iterator to one of n_jobs independent partitions.
No write unit is ever shared between two partitions, so the jobs
need no locks, no barriers, and no channel to one another — safe as
separate SLURM array tasks, in any order. Every job must build an
identical iterator (construction is metadata-only and the partition
derives deterministically from it), apply for_job last in the
builder chain, and its verbs deliberately do not finalize — a slice's
map/process/segment writes only its share, and measure/
detect bank a partial instead of joining. The gather step is the
unrestricted iterator's finalize(), once, after all jobs.
Stitched and read-only runs also need prepare_jobs(n_jobs) first.
See the distributed-runs guide in the iterators docs.
Parameters:
-
job_index(int) –This job's partition,
0 <= job_index < n_jobs(e.g.int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops. -
n_jobs(int) –How many partitions the run is split into; must match across every job and the gather.
Returns:
-
Self–The restricted iterator; further reshaping refuses.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
prepare_jobs
¶
prepare_jobs(n_jobs: int) -> list[JobArgs]
Set up a distributed run and return its parallelization list.
The init step of the three-phase recipe (init → parallel jobs →
consolidate): performs the iterator's setup — always clearing stale
scratch state from earlier runs first — and returns one JobArgs
per non-empty partition. Optional for plain writers; a stitching or
read-only run requires it (the scratch root must exist, race-free,
before any job writes into it).
Parameters:
-
n_jobs(int) –How many partitions to split into; empty ones are dropped from the list.
Returns:
-
list[JobArgs]–One
{"job_index": ..., "n_jobs": ...}dict per non-empty -
list[JobArgs]–partition — JSON-ready, splatting into
for_job(**args).
Example
Source code in src/ngio/iterators/_abstract_iterator.py
reduce
¶
reduce(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run.
Parameters:
-
func(Callable[[NumpyPipeType], R]) –The function to apply; under a parallel mapper it must be safe on worker threads (or processes).
-
mapper(MapperProtocol[NumpyPipeType, R] | None, default:None) –How the units are scheduled.
None(the default) is serial; passThreadedMapper()orProcessMapper()to fan out, sized by their ownmax_workersargument.
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
reduce_as_numpy
¶
reduce_as_numpy(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Alias for reduce().
reduce_as_dask
¶
reduce_as_dask(
func: Callable[[DaskPipeType], R],
*,
mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Deprecated: removed in ngio=1.2.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run. Runs serially — the units hand
func lazy dask arrays, so the per-unit work is graph
construction, not IO. A parallel mapper raises (see map_as_dask).
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_regions_overlap
¶
Check if any of the ROIs overlap logically.
If two ROIs cover the same pixel, they are considered to overlap.
This does not consider chunking or other storage details. Measured
on the read regions — with a halo, neighbouring reads overlap by
construction. For write safety, check_if_write_units_overlap is
the right question.
Returns:
-
bool–Trueif any ROIs overlap.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_regions_overlap
¶
Ensure that the Iterator's ROIs do not overlap.
with_nms
¶
with_nms(nms: NmsProtocol) -> Self
Declare how duplicate detections are suppressed.
The default is GreedyNms(). The jobs of a distributed run read
score_column/max_detections_per_tile at banking, so declare the
identical NMS on every job and the gather — like the detection
function, it is not part of the plan fingerprint. Declare before
for_job.
Parameters:
-
nms(NmsProtocol) –Anything satisfying
NmsProtocol—GreedyNms(...), or your own suppression.
Source code in src/ngio/iterators/_object_detection.py
build_numpy_getter
¶
build_numpy_getter(roi: Roi) -> DetectionGetter[ndarray]
Source code in src/ngio/iterators/_object_detection.py
build_dask_getter
¶
build_dask_getter(roi: Roi) -> DetectionGetter[Array]
Source code in src/ngio/iterators/_object_detection.py
finalize
¶
finalize() -> RoiTable
Merge a distributed run's partials into the one final ROI table.
The gather step (see prepare_jobs): validates that every job of
the prepared plan produced a matching partial — a half-finished run
errors instead of returning a silently incomplete table — then
rebuilds every tile's raw boxes and runs the full serial pipeline
once, globally: anchoring, the cross-tile invariant checks, NMS, and
the dense 1..N renumbering. Bit-identical to a serial detect by
construction — extra fields keep their columns and dtypes through
the partial round-trip. Two residuals, both within one tile: an
extra a detector sets to NaN does not survive, and an integer
extra present on only some of a tile's boxes comes back float (the
tile frame NaN-fills it before the schema is recorded). On success
the partials group is removed. Nothing is registered: the returned
table is yours to store with add_table.
Raises on a for_job slice (the gather is global) and when no
partials exist (nothing was prepared or banked).
Source code in src/ngio/iterators/_object_detection.py
detect
¶
detect(
func: Callable[[ndarray], list[Roi]],
*,
mapper: MapperProtocol | None = None,
) -> RoiTable | None
Run the detector over every tile and return one table of objects.
The per-tile detection fans out exactly like reduce — pass a
mapper to parallelize it — and the reconciliation (coordinates, NMS,
renumbering) happens once, at the end, on the calling thread. Nothing
is written: the returned table is yours to store, e.g.
container.add_table(name, table).
On a for_job slice it detects only over this job's tiles and banks
the detector's raw, pre-NMS boxes as a partial, returning None
— greedy NMS is not hierarchical, so a per-job NMS followed by a
merge NMS could differ from a serial run. The gather (finalize(),
once, after all jobs) runs the ONE global NMS and renumbering,
bit-identical to a serial detect. The tiles' boxes are validated
here, so a malformed detector fails in the job rather than at the
gather; re-running a job overwrites its own partial.
Parameters:
-
func(Callable[[ndarray], list[Roi]]) –(patch) -> list[Roi], boxes in the patch's own pixel coordinates — see the class docstring for the box contract. Under a parallel mapper it runs on worker threads or processes and must be safe there. -
mapper(MapperProtocol | None, default:None) –How the per-tile work is scheduled;
Noneis serial.
Returns:
-
RoiTable | None–A
RoiTableof the surviving detections, ids1..Nin -
RoiTable | None–(tile, row) order, in the reference image's world coordinates.
-
RoiTable | None–An image with nothing to detect gives an empty table. On a
-
RoiTable | None–for_jobslice:None(the partial is banked).
Source code in src/ngio/iterators/_object_detection.py
Reconciliation declarations¶
Each iterator declares its reconciliation in the builder chain —
on_overlap, with_stitch, with_join, with_nms — backed by a swappable
protocol with a shipped default.
StitchConfig¶
ngio.iterators.StitchConfig
dataclass
¶
StitchConfig(
*,
block_size: int = 10000,
iou_threshold: float = 0.3,
seam_matcher: SeamMatcherProtocol | None = None,
compact: bool = True,
scratch_store: StoreOrGroup | None = None,
write_order: WriteOrder = "any",
)
How to stitch a tiled segmentation.
Attributes:
-
block_size(int) –Ids reserved per tile. Must exceed the largest label a single tile can produce, or ids spill into the next tile's block.
-
iou_threshold(float) –How much two tiles' predictions must agree over their shared pixels before their ids are called the same object. Low values merge eagerly; the default errs towards leaving an object split, which is fixable downstream, rather than merging two that are not. Configures the default
IouSeamMatcher; ignored whenseam_matcheris set. -
seam_matcher(SeamMatcherProtocol | None) –Replaces the default pair criterion — see
SeamMatcherProtocol. Like the measurement function, a custom matcher is not part of the plan fingerprint: declare the identical one on every phase of a distributed run. -
compact(bool) –Renumber the surviving ids to a dense
1..Nat the end. The renumbering walks the whole label, so objects outside the iterated ROIs get new ids too — anything keyed by the old ids (feature tables, masking ROI tables) is invalidated. A run that renumbers such pre-existing objects warns. Usecompact=Falseand reconcile ids yourself to keep external references valid. -
scratch_store(StoreOrGroup | None) –Where to keep the per-tile banks while the map runs.
Noneputs them in a transient group inside the output label, which works under every mapper. AMemoryStoreavoids touching the output store — but an in-memory scratch cannot cross a process boundary, soProcessMapperrefuses it. -
write_order(WriteOrder) –Who owns the seams of unmatched objects — object unification is exact either way.
"any"(the default) is schedule-defined;"roi"gives them to the later tile, reproducibly, at up to 2.5x cost on parallel overlapping tilings.
matcher
¶
matcher() -> SeamMatcherProtocol
The seam matcher the resolve runs; the IoU default unless swapped.
SeamMatcherProtocol¶
ngio.iterators.SeamMatcherProtocol
¶
Bases: Protocol
Decides which ids of two overlapping banked predictions are one object.
Called only at the resolve (the gather step, on the calling thread — no
pickling constraints), once per overlapping tile pair, with the two
tiles' banked label patches reconstructed over their shared region:
aligned, same shape, in the label's own axes order. Ids are as they
appear in the patches (each tile's block-offset ids). Return the
(id_in_a, id_in_b) pairs that are the same object; the framework owns
which tile pairs are compared and over which regions, the banking and
scratch lifecycle, the union-find, and the global relabel.
Must be deterministic — a pure function of the two patches — or a
distributed gather is no longer bit-identical to a serial run.
Background 0 is never pairable, and every returned id must appear in
its patch; violations raise at the resolve.
IouSeamMatcher¶
ngio.iterators.IouSeamMatcher
dataclass
¶
The default seam criterion: IoU over the shared pixels, thresholded.
overlap_iou counts only pixels labelled on both sides, so background
(and a masked tile's outside) never enters a pair and can never
displace a written id.
NmsProtocol¶
ngio.iterators.NmsProtocol
¶
Bases: Protocol
What the detection iterator needs from its NMS declaration.
Two properties the framework reads outside suppression, plus the
suppression itself. suppress runs only at the join (a serial
detect, or the distributed gather's finalize), on the calling
thread — no pickling constraints. It must be deterministic — a pure
function of the detection list, ties broken deterministically (the
default uses (-score, tile_index, row)) — or a distributed gather is
no longer bit-identical to a serial run.
Compatibility policy: future ngio versions will only ever add
optional members, probed with getattr.
GreedyNms¶
ngio.iterators.GreedyNms
dataclass
¶
GreedyNms(
*,
iou_threshold: float = 0.5,
score_column: str = "confidence",
max_detections_per_tile: int = 10000,
)
Greedy NMS, the default suppression: highest score first.
Attributes:
-
iou_threshold(float) –Two boxes overlapping at or above this intersection-over-union are the same object; the lower-scoring one is suppressed. Values near 1 keep almost everything, values near 0 merge aggressively; 0 itself would merge any pair that merely touches and is refused.
-
score_column(str) –Extra field on the returned Rois ranking the detections (and the resulting table's column),
"confidence"by default. When the detector provides no such field the box volume ranks instead — bigger wins. -
max_detections_per_tile(int) –A tile producing more detections than this is a runaway detector; it fails loudly instead of flooding the table.
suppress
¶
Greedy pass: keep by rank, drop what overlaps a kept box.
Source code in src/ngio/iterators/_object_detection.py
Detection¶
ngio.iterators.Detection
dataclass
¶
Detection(
*,
tile_index: int,
row: int,
score: float,
roi: Roi,
bounds: dict[str, tuple[float, float]],
group: tuple,
)
One anchored detection, as suppression sees it.
Constructed by the framework only. A custom suppression returns a subset of the instances it was given — never construct or copy these — so future ngio versions stay free to add fields.
Attributes:
-
tile_index(int) –Position of the reporting tile in
iterator.rois. -
row(int) –The box's position in that tile's returned list.
(tile_index, row)is the deterministic tie-breaker. -
score(float) –The ranking value (the
score_columnextra, or box volume when no tile reported one). -
roi(Roi) –The absolute world-space ROI, extra fields riding on it (a class id, say) — not yet renumbered.
-
bounds(dict[str, tuple[float, float]]) –Reference-image pixel coordinates,
[min, max)per bbox axis. Treat as read-only. -
group(tuple) –Opaque key of the axes the boxes do not pin (t, or z for a 2D detector on 3D data). Only detections with equal
groupcan be duplicates of each other.
JoinProtocol¶
ngio.iterators.JoinProtocol
¶
Bases: Protocol
Joins the normalized per-ROI frames into one ngio table.
Receives the list described in measure — one DataFrame per ROI in
global ROI order, label as a column, every row stamped with
roi_index/roi_name, empty ROIs as empty column-less frames — and
returns the table. Plain callables satisfy this structurally. Runs
once, wherever the join executes: a serial measure, or the
distributed gather's finalize. Must be deterministic, or a
distributed gather is no longer bit-identical to a serial run.
ConcatJoin¶
ngio.iterators.ConcatJoin
dataclass
¶
The default join: concatenate into a FeatureTable indexed by label.
The rows must carry the object id in a label column (or already sit
on a label index). The stamped roi_index/roi_name columns ride
into the final table. Duplicate label ids (a haloed or overlapping
tiling measures border objects more than once) are kept as-is —
deduplicating is a custom join's job. The zero-object table keeps its
label-only schema.
Mappers¶
ThreadedMapper¶
ngio.iterators.ThreadedMapper
¶
ThreadedMapper(max_workers: MaxWorkers = 'auto')
Bases: Generic[T, R]
Run units concurrently on a thread pool.
Each unit's read, func call and write all run on the pool. ngio's
getters and setters are thread-safe (immutable per-ROI state over zarr's
sync API); func must be thread-safe too — that part of the contract is
the caller's.
Units run in conflict-free waves, back to back on one pool — see
plan_waves. With one unit, or a resolved pool of one, this is exactly
BasicMapper.
Parameters:
-
max_workers(MaxWorkers, default:'auto') –"auto"(the default) sizes the pool for round-trip-bound work; anintpins it.Noneis accepted and means"auto".
Source code in src/ngio/iterators/_mappers.py
ProcessMapper¶
ngio.iterators.ProcessMapper
¶
ProcessMapper(max_workers: MaxWorkers = 'auto')
Bases: Generic[T, R]
Run units on a spawn-based process pool.
The fit for pure-Python, GIL-holding funcs (ThreadedMapper already
covers IO-bound work and GIL-releasing compute). Each child reads its ROI
from the store, applies func, and writes back — pixels never cross the
process boundary. For written units the returned result is therefore
None (which map discards anyway); only reduce results are
pickled back to the parent.
Units run in conflict-free waves (see plan_waves) — which is the
entire cross-process safety argument: no lock ngio could take would
work across processes anyway.
Constraints:
- func (and any transforms the units carry) must be picklable — a
module-level function, not a lambda or closure.
- In-memory stores are refused: a MemoryStore pickles by value, so a
child would write into its own copy and the parent would never see it.
- The pool uses the spawn start method: the parent holds zarr's IO
event-loop thread, and forking a threaded process is unsafe. Each
worker pays an interpreter start and import ngio (~1 s), amortized
over the units it processes.
Parameters:
-
max_workers(MaxWorkers, default:'auto') –"auto"(the default) or anint;Nonemeans"auto". The pool never exceeds the unit count.
Source code in src/ngio/iterators/_mappers.py
BatchedMapper¶
ngio.iterators.BatchedMapper
¶
BatchedMapper(
batch_size: int = 8,
pad_mode: str = "constant",
pad_values: int | float = 0,
read_workers: MaxWorkers = "auto",
)
Stack patches into (B, ...) batches and call func once per batch.
The fit for neural-network inference. Unlike every other mapper, func
receives a stacked array — a leading batch axis over up to
batch_size patches — and must return an array-like with the same
leading axis whose items follow map's per-patch contract.
Ragged tilings are padded per batch to the per-axis maximum before stacking (origin-anchored: real pixels first, padding after); a shape-preserving output is sliced back to each patch's true shape before the write, and halos are trimmed by the setter as usual. A per-item reduction is allowed only on a uniform batch — on a ragged one its result is computed on padded pixels and there is no padding to slice back off, so it raises. Reads within a batch fan out on a thread pool; writes run serially on the dispatching thread, so batched mapping is write-safe on any tiling.
Stacking is a numpy operation, so unlike its siblings this mapper is
not generic over the payload type: it accepts bare np.ndarray units
only (tuple payloads raise). Its pool argument is named
read_workers, not max_workers, because it sizes something
different — the other mappers' pools run func, this one's only
reads; func always runs once per batch on the dispatching thread.
Parameters:
-
batch_size(int, default:8) –Number of patches stacked per
funccall. The last batch may be smaller. -
pad_mode(str, default:'constant') –np.padmode used to grow ragged patches to the batch shape, typically"constant"or"reflect". -
pad_values(int | float, default:0) –Fill value when
pad_mode="constant", ignored otherwise. -
read_workers(MaxWorkers, default:'auto') –Pool size for the per-batch reads.
"auto"(the default) sizes it for round-trip-bound work, anintpins it,1reads serially;Nonemeans"auto".
Source code in src/ngio/iterators/_mappers.py
BasicMapper¶
ngio.iterators.BasicMapper
¶
Bases: Generic[T, R]
Serial mapper: read, apply, write (if writable), one unit at a time.
MapperProtocol¶
ngio.iterators.MapperProtocol
¶
Bases: Protocol[T, R]
Protocol for mappers.
Implementations must honour this contract:
- The mapper schedules each unit's write; it is not required to invoke
unit.settereagerly or on any particular thread, only to guarantee that all writes are complete when__call__returns. - A unit whose
setterisNoneis read-only: compute and collect its result, never write. - For a unit with a setter the collected result may be
Noneinstead of the written patch:mapdiscards the results, and holding every patch until the call returns would put the whole output in memory at once. ngio's mappers all collectNonefor written units. - The returned list is ordered by
unit.index(results[i]corresponds toiterator.rois[i]), regardless of execution order. unitsis typically a generator and each unit is expensive to build (store metadata round-trips); generators are not thread-safe. A parallel mapper must materialise and consumeunitson the dispatching thread only, then distribute the materialised units to its workers.
Compatibility policy: future ngio versions will only ever add
optional members to this protocol, probed with getattr — an
implementation satisfying today's contract keeps working.
IterUnit¶
ngio.iterators.IterUnit
dataclass
¶
IterUnit(
*,
index: int,
roi: Roi,
getter: DataGetterProtocol[T],
setter: DataSetterProtocol[T] | None,
write_order: Literal["roi", "any"] = "roi",
)
Bases: Generic[T]
One schedulable unit of iterator work: read one ROI, optionally write back.
Attributes:
-
index(int) –Position in
iterator.rois— mapper results must be returned in this order. -
roi(Roi) –The ROI this unit covers.
-
getter(DataGetterProtocol[T]) –Reads the ROI's data.
-
setter(DataSetterProtocol[T] | None) –Writes the transformed data back, or
Nonefor a read-only unit (a read-only iterator, or anyreducecall). -
write_order(Literal['roi', 'any']) –See
WriteOrder. A pair is relaxed only when both units declare"any". The field default is"roi"so a hand-built unit is strictly ordered unless it asks otherwise; the iterators stamp their own declared value.
write_footprint
property
¶
write_footprint: ChunkRect | None
The chunk rectangle this unit writes, at write granularity.
The granularity is the shard shape when the output array is sharded,
the chunk shape otherwise. None when the unit is read-only or its
write selection is empty — either way there is nothing to claim and
the unit conflicts with nothing.
Scheduling¶
The primitives behind the mappers' conflict-free schedule — useful for inspecting how a tiling will parallelize or split into jobs.
plan_waves¶
ngio.iterators.plan_waves
¶
Partition units into conflict-free waves that respect ROI-index order.
Two units whose write footprints share a write unit (a chunk, or a shard
when the output is sharded) of the same output array never share a wave:
a wave's writes are pairwise disjoint at write granularity, so running
one wave at a time preserves the single-writer-per-write-unit invariant
that makes parallel writes lock-free in any topology. Read-only units and
units with an empty write selection conflict with nothing and land in the
first wave. Assignment runs in ascending unit.index (first-fit greedy),
so the schedule is a pure function of the unit sequence. A conflict-free
set — a by_write_units() tiling, say — yields a single wave.
On top of the safety constraint, write_order="roi" units are ordered
on pixel overlap: the higher unit.index lands in a strictly later
wave, so the later ROI wins under every mapper and in the manual
iter loop. A pair where both units declare "any" is scheduled for
parallelism alone. Sharing a write unit without sharing pixels never
orders, so disjoint-pixel tilings keep their packed schedule.
Parameters:
-
units(Sequence[IterUnit[T]]) –The units to schedule.
-
log(bool, default:True) –Emit schedule-quality log messages. The parallel mappers keep this on; serial callers ordering their units pass
False.
Source code in src/ngio/iterators/_mappers.py
539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 | |
canonical_unit_order¶
ngio.iterators.canonical_unit_order
¶
Units in flattened wave order — the one write order every mapper uses.
Wave 0 in index order, then wave 1, and so on. With no write conflicts
this is plain index order. On pixel-overlapping "roi" units the
higher ROI index always lands later; all-"any" pairs follow the
schedule — either way a pure function of the unit sequence.
Source code in src/ngio/iterators/_mappers.py
write_conflict_components¶
ngio.iterators.write_conflict_components
¶
write_conflict_components(
units: Sequence[IterUnit[Any]],
) -> list[list[int]]
Connected components of the write-conflict graph, as unit-index lists.
Two units are connected when their write footprints share a write unit (a
chunk, or a shard when the output is sharded) of the same output array —
the same adjacency plan_waves schedules by, computed by the same code,
so the scheduler and the splitter can never disagree about who conflicts.
Read-only units and units with an empty write selection conflict with
nothing and are singletons.
A component is the unit of independence: units in different components
never share a write unit, so they need no coordination in any topology.
That is the property for_job builds on — a component never spans
two jobs. Each component is a sorted list of unit.index; components are
ordered by their smallest index, and the whole result is a pure function
of the unit sequence.
Source code in src/ngio/iterators/_mappers.py
compute_write_footprint¶
ngio.iterators.compute_write_footprint
¶
compute_write_footprint(
setter: DataSetterProtocol[Any],
) -> ChunkRect | None
The chunk rectangle a setter will write, at write granularity.
The granularity is the shard shape when the target array is sharded (writes
are atomic per shard object), the chunk shape otherwise — see
AbstractImage.write_granularity. Returns None when the setter's
selection is empty.
Source code in src/ngio/iterators/_mappers.py
Types¶
JobArgs¶
ngio.iterators.JobArgs
¶
Bases: TypedDict
One entry of prepare_jobs's parallelization list.
A plain dict at runtime — JSON-serializable for schedulers that ship
task arguments as JSON (Fractal's parallelization list), and it splats
straight into the selector: iterator.for_job(**args).
TailPolicy¶
ngio.iterators.TailPolicy
module-attribute
¶
How by_grid handles the non-whole tile at the end of an axis.
HaloMargins¶
ngio.iterators.HaloMargins
module-attribute
¶
Per-axis pixels added on each side of a ROI when reading, keyed by axis name.
FeatureFuncResult¶
ngio.iterators.FeatureFuncResult
module-attribute
¶
What a per-ROI feature function may return: a DataFrame, or the cheaper
dict-of-columns ({"label": [...], "area": [...]}) — lighter to build in a
worker and to ship across a ProcessMapper boundary; normalization turns
either into one DataFrame per ROI before the join.
MaxWorkers¶
ngio.iterators.MaxWorkers
module-attribute
¶
Accepted by every max_workers argument. 1 is serial deliberately;
"auto" picks a pool sized for round-trip-bound work. None means
"unspecified", and what that resolves to belongs to the call site: the
plate fan-outs run serially (until the ngio=1.2 default flip they warn
about), the parallel mappers treat it as "auto" — asking for such a
mapper is already the opt-in.
AbstractIteratorBuilder¶
The shared method surface of every iterator — the reshaping calls
(by_grid, by_blocks, by_chunks, by_write_units, product, with_halo,
for_job), each class's reconciliation declaration (on_overlap,
with_stitch, with_join, with_nms), the generic execution calls (iter —
batched via batch_size — map, reduce) beneath each iterator's topic verb
(process, segment, measure, detect), the distributed-run init step
(prepare_jobs), and the one gather, finalize — the writers consolidate and
return None, the read-only iterators merge their banked partials and return
the table.
ngio.iterators.AbstractIteratorBuilder
¶
Bases: ABC, Generic[NumpyPipeType, DaskPipeType, FinalizeType]
Base class for building iterators over ROIs.
Constructor naming convention across the concrete iterators: a bare
parameter (channel_selection, input_transforms) refers to the input
image; output-side parameters are always prefixed
(output_channel_selection, output_transforms), and parameters
targeting another object name it (label_transforms). A parameter that
configures a downstream consolidation is spelled consolidation_mode
(here and on Label.relabel_sequential); only consolidate() itself
shortens it to mode, where the prefix would be redundant.
output_image
property
¶
output_image: AbstractImage | None
The image this iterator writes to, or None for a read-only iterator.
partition_indices
property
¶
The unit indices this iterator is restricted to, or None.
None on an unrestricted iterator; on a for_job slice,
the sorted positions (in the full ROI list) of the units it runs.
The whole layout of a split is
[it.for_job(i, n).partition_indices for i in range(n)] —
the pre-flight check before submitting a distributed run: one fat
list plus empties means the output chunking, not the cluster, is
the constraint.
with_halo
¶
Read a margin of context around each ROI, but write only the ROI.
The function receives the grown region and must return it grown too; the border is cropped before the write, so tiles come out seamless. Margins are reference-image pixels, clipped at the borders. Write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one.
Note
The margins do not survive a rescale: a shape-changing transform
on the data path (a ZoomTransform) makes the crop refuse —
iterate on a coarser pyramid level instead. A MaskTransform
composes normally (its zoom rescales the label it reads).
Parameters:
-
x(int, default:0) –Pixels added on each side along x.
-
y(int, default:0) –Pixels added on each side along y.
-
z(int, default:0) –Pixels added on each side along z.
-
t(int, default:0) –Frames added on each side along t.
Returns:
-
Self–A new iterator reading with the halo.
Raises:
-
NgioValueError–On a read-only iterator that has not opted in (no write to crop the halo from; detection opts in and reconciles the overlap with NMS, feature extraction opts in and stamps
roi_index/roi_nameso yourcoalescecan), on an in-place iterator, or for a negative margin.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
by_grid
¶
by_grid(
*,
size_x: int | None = None,
size_y: int | None = None,
size_z: int | None = None,
size_t: int | None = None,
stride_x: int | None = None,
stride_y: int | None = None,
stride_z: int | None = None,
stride_t: int | None = None,
tail: TailPolicy = "clip",
base_name: str = "",
) -> Self
Tile the current ROIs with a regular grid.
Parameters:
-
size_x(int | None, default:None) –Tile size along x;
None(default) means the full extent. -
size_y(int | None, default:None) –Tile size along y;
Nonemeans the full extent. -
size_z(int | None, default:None) –Tile size along z;
Nonemeans the full extent. -
size_t(int | None, default:None) –Tile size along t;
Nonemeans the full extent. -
stride_x(int | None, default:None) –Step between tiles along x;
Nonemeanssize_x(adjacent tiles). A smaller stride produces overlapping tiles. -
stride_y(int | None, default:None) –As
stride_x, along y. -
stride_z(int | None, default:None) –As
stride_x, along z. -
stride_t(int | None, default:None) –As
stride_x, along t. -
tail(TailPolicy, default:'clip') –What happens where the grid does not divide the axis.
"clip"(default) shrinks the last tile to the border;"balance"re-splits the last full tile and the tail into two near-equal tiles, so a thin overhang never yields a thin tile (needs adjacent tiles,stride == size);"shift"moves the last tile back to stay full-size — it then overlaps its neighbour, so a writing iterator needs a merge or serial execution;"drop"discards the tail and leaves that border uncovered. -
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_blocks
¶
by_blocks(
*,
num_x: int = 1,
num_y: int = 1,
num_z: int = 1,
num_t: int = 1,
base_name: str = "",
) -> Self
Divide each axis into a fixed number of near-equal blocks.
The complement of by_grid: you say how many tiles, not how big.
Block lengths along an axis differ by at most one pixel, so there is
no tail to police — the partition is balanced by construction.
Parameters:
-
num_x(int, default:1) –Number of blocks along x (default 1: no split).
-
num_y(int, default:1) –Number of blocks along y.
-
num_z(int, default:1) –Number of blocks along z.
-
num_t(int, default:1) –Number of blocks along t.
-
base_name(str, default:'') –Prefix for the generated tile names.
Returns:
-
Self–A new iterator over the blocks.
Source code in src/ngio/iterators/_abstract_iterator.py
by_yx
¶
Split each ROI into one 2D plane per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, z, c) combination, each covering the full y/x extent — the
shape for "run this 2D function on every plane".
Source code in src/ngio/iterators/_abstract_iterator.py
by_zyx
¶
Split each ROI into one 3D volume per remaining coordinate.
A broadcasting call, not a tiling: every ROI becomes one ROI per
(t, c) combination, each covering the full z/y/x extent.
Parameters:
-
strict(bool, default:True) –Require a real z axis (present and larger than 1); with
False, images without one fall back to 2D planes.
Source code in src/ngio/iterators/_abstract_iterator.py
by_chunks
¶
by_chunks(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the input image's chunk grid.
Reads are chunk-granular, so this is the natural tiling for
read-heavy work. For collision-free parallel writes use
by_write_units, which tiles by the output's write granularity.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
by_write_units
¶
by_write_units(
*,
overlap_x: int = 0,
overlap_y: int = 0,
overlap_z: int = 0,
overlap_t: int = 0,
) -> Self
Tile the ROIs by the output image's write granularity.
The write unit is the shard shape when the output is sharded (writes
are atomic per shard object), the chunk shape otherwise. With no
overlap the resulting ROIs pass check_if_write_units_overlap by
construction, so a parallel map runs as a single fully-parallel
wave — a throughput optimization; conflicting tilings are still safe,
just scheduled in more waves. Falls back to the input chunk grid when
the iterator is read-only.
Parameters:
-
overlap_x(int, default:0) –Overlap between adjacent tiles along x, in pixels.
-
overlap_y(int, default:0) –Overlap along y.
-
overlap_z(int, default:0) –Overlap along z.
-
overlap_t(int, default:0) –Overlap along t.
Returns:
-
Self–A new iterator with tiled ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
product
¶
product(other: list[Roi] | GenericRoiTable) -> Self
Cartesian product of the current ROIs with an arbitrary list of ROIs.
Source code in src/ngio/iterators/_abstract_iterator.py
build_numpy_getter
abstractmethod
¶
build_numpy_getter(
roi: Roi,
) -> DataGetterProtocol[NumpyPipeType]
build_dask_getter
abstractmethod
¶
build_dask_getter(
roi: Roi,
) -> DataGetterProtocol[DaskPipeType]
finalize
abstractmethod
¶
The one gather step, run once after every unit (or job) completes.
Writing iterators resolve any pending reconciliation (the stitch) and
consolidate the output pyramid, returning None; the serial verbs
call it automatically. Read-only iterators merge the partials banked
by a distributed run and return the final table — nothing is
registered, the table is the caller's to store with add_table.
Raises on a for_job slice: the gather is global and belongs to the
unrestricted iterator, once, after all jobs.
Source code in src/ngio/iterators/_abstract_iterator.py
iter
¶
iter(
lazy: Literal[True],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[NumpyPipeType]]
iter(
lazy: Literal[True],
data_mode: Literal["dask"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[DataGetterProtocol[DaskPipeType]]
iter(
lazy: Literal[False],
data_mode: Literal["numpy"],
iterator_mode: Literal["readonly"] = ...,
*,
batch_size: None = ...,
) -> Generator[NumpyPipeType]
iter(
lazy: bool = False,
data_mode: Literal["numpy", "dask"] | None = None,
iterator_mode: Literal[
"readwrite", "readonly"
] = "readonly",
*,
batch_size: int | None = None,
) -> Generator
Iterate the ROIs' payloads, read-only.
With batch_size set, yields payload lists of up to that many
items (the last batch may be smaller). Batching is numpy-only and
eager: combining batch_size with lazy=True or
data_mode="dask" raises. On a writing iterator (see
WritingIteratorBuilder.iter) the default mode is "readwrite"
and the loop yields (patch, writer) pairs instead.
Parameters:
-
lazy(bool, default:False) –Yield getter handles instead of materialized data.
-
data_mode(Literal['numpy', 'dask'] | None, default:None) –"numpy"or"dask"(deprecated, see note). -
iterator_mode(Literal['readwrite', 'readonly'], default:'readonly') –"readonly"yields patches alone; on a writing iterator"readwrite"yields(patch, writer)pairs. -
batch_size(int | None, default:None) –When set, yield batches of up to this many items; the last batch may be smaller.
Note
The dask data mode is deprecated and will be removed in ngio=1.2;
from then on iter() yields numpy arrays.
Source code in src/ngio/iterators/_abstract_iterator.py
iter_as_numpy
¶
iter_as_dask
¶
Alias for iter(data_mode="dask").
Deprecated: removed in ngio=1.2.
Source code in src/ngio/iterators/_abstract_iterator.py
for_job
¶
Restrict the iterator to one of n_jobs independent partitions.
No write unit is ever shared between two partitions, so the jobs
need no locks, no barriers, and no channel to one another — safe as
separate SLURM array tasks, in any order. Every job must build an
identical iterator (construction is metadata-only and the partition
derives deterministically from it), apply for_job last in the
builder chain, and its verbs deliberately do not finalize — a slice's
map/process/segment writes only its share, and measure/
detect bank a partial instead of joining. The gather step is the
unrestricted iterator's finalize(), once, after all jobs.
Stitched and read-only runs also need prepare_jobs(n_jobs) first.
See the distributed-runs guide in the iterators docs.
Parameters:
-
job_index(int) –This job's partition,
0 <= job_index < n_jobs(e.g.int($SLURM_ARRAY_TASK_ID)). Partitions beyond the number of independent unit groups are empty no-ops. -
n_jobs(int) –How many partitions the run is split into; must match across every job and the gather.
Returns:
-
Self–The restricted iterator; further reshaping refuses.
Example
Source code in src/ngio/iterators/_abstract_iterator.py
prepare_jobs
¶
prepare_jobs(n_jobs: int) -> list[JobArgs]
Set up a distributed run and return its parallelization list.
The init step of the three-phase recipe (init → parallel jobs →
consolidate): performs the iterator's setup — always clearing stale
scratch state from earlier runs first — and returns one JobArgs
per non-empty partition. Optional for plain writers; a stitching or
read-only run requires it (the scratch root must exist, race-free,
before any job writes into it).
Parameters:
-
n_jobs(int) –How many partitions to split into; empty ones are dropped from the list.
Returns:
-
list[JobArgs]–One
{"job_index": ..., "n_jobs": ...}dict per non-empty -
list[JobArgs]–partition — JSON-ready, splatting into
for_job(**args).
Example
Source code in src/ngio/iterators/_abstract_iterator.py
reduce
¶
reduce(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run.
Parameters:
-
func(Callable[[NumpyPipeType], R]) –The function to apply; under a parallel mapper it must be safe on worker threads (or processes).
-
mapper(MapperProtocol[NumpyPipeType, R] | None, default:None) –How the units are scheduled.
None(the default) is serial; passThreadedMapper()orProcessMapper()to fan out, sized by their ownmax_workersargument.
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
reduce_as_numpy
¶
reduce_as_numpy(
func: Callable[[NumpyPipeType], R],
*,
mapper: MapperProtocol[NumpyPipeType, R] | None = None,
) -> list[R]
Alias for reduce().
reduce_as_dask
¶
reduce_as_dask(
func: Callable[[DaskPipeType], R],
*,
mapper: MapperProtocol[DaskPipeType, R] | None = None,
) -> list[R]
Apply a function to every ROI and collect the results without writing.
Deprecated: removed in ngio=1.2.
Units are built read-only even on writable iterators: nothing is
written and finalize does not run. Runs serially — the units hand
func lazy dask arrays, so the per-unit work is graph
construction, not IO. A parallel mapper raises (see map_as_dask).
Returns:
-
list[R]–One result per ROI, in ROI order:
results[i]corresponds to -
list[R]–self.rois[i], regardless of the mapper's execution order.
Source code in src/ngio/iterators/_abstract_iterator.py
check_if_regions_overlap
¶
Check if any of the ROIs overlap logically.
If two ROIs cover the same pixel, they are considered to overlap.
This does not consider chunking or other storage details. Measured
on the read regions — with a halo, neighbouring reads overlap by
construction. For write safety, check_if_write_units_overlap is
the right question.
Returns:
-
bool–Trueif any ROIs overlap.
Source code in src/ngio/iterators/_abstract_iterator.py
require_no_regions_overlap
¶
Ensure that the Iterator's ROIs do not overlap.