Skip to content

6. Iterators

Process an image region by region without writing the loop.

When building image processing pipelines it is often useful to iterate over specific regions of the image, for example to process the image in smaller tiles or to process only specific regions of interest (ROIs). Iterators also let you set broadcasting rules for the iteration, for example to iterate over all z-planes or over all timepoints.

How an iterator walks an image A ROI table names the regions. For each region the iterator reads that part of the input, applies your function, and writes the result into the output, one region at a time. 1234 ROI TABLE INPUT YOUR FUNCTION OUTPUT region 1 region 2 region 3 … process() repeat for every region the table names

ngio provides five basic Iterator classes, all imported from ngio.iterators (or from the top-level ngio namespace):

What each of the five iterators takes, returns and supports Segmentation takes an image and returns a label; masked segmentation takes an image and a mask; image processing returns an image; feature extraction and object detection are read-only and return tables. A features column lists what each one supports: every iterator takes a halo, stitching applies only to the two segmentation iterators, and the read-only ones reconcile their read margin with a join or with NMS. ITERATORIN → OUTTOPIC VERBFEATURES SegmentationIteratorMaskedSegmentationIteratorImageProcessingIteratorFeatureExtractorIteratorObjectDetectionIterator segmentsegmentprocessmeasuredetect halostitching halostitching halo read halojoin read haloNMS
  • SegmentationIterator — segment an image into a label; see the image segmentation tutorial, and the stitching tutorial for tiling and seams.
  • MaskedSegmentationIterator — the same, restricted to the objects of a masking ROI table (segment cells only within a tissue region, say); same tutorials.
  • ImageProcessingIterator — process an image into a new image (a filter, a projection, a restoration model); see the image processing tutorial.
  • FeatureExtractorIterator — read-only; measure joins per-region measurements into one feature table; see the feature extraction tutorial.
  • ObjectDetectionIterator — read-only; detect turns a tile-by-tile detector into one deduplicated ROI table; see the object detection tutorial.

The verbs come in two layers. Every iterator shares the generic layer — map (apply and write back), reduce (collect without writing), the hand-driven iter loop, and the distributed steps prepare_jobs/for_job/finalize — and each iterator adds one topic verb that says what the iteration means: segment, process, measure, detect. On the writers the topic verb is map under its domain name; on the read-only iterators it also runs the final join (the feature join, the detection NMS) and returns the table. All four are partition-aware: on a for_job slice they do only that job's share, and finalize() is the one gather, whatever the iterator.

Building one

Every iterator is constructed from the images it reads and writes, then narrowed. A fresh iterator covers the whole image as a single region:

from pathlib import Path

from ngio import open_ome_zarr_container
from ngio.iterators import ImageProcessingIterator
from ngio.utils import download_ome_zarr_dataset

download_dir = Path("./data").absolute()
hcs_path = download_ome_zarr_dataset(
    "CardiomyocyteSmallMip", download_dir=download_dir, re_unzip=False
)
ome_zarr = open_ome_zarr_container(hcs_path / "B" / "03" / "0")
image = ome_zarr.get_image()
# A new iterator covers the whole image as a single region
iterator = ImageProcessingIterator(input_image=image, output_image=image)
print(iterator)
ImageProcessingIterator(rois=1)

The constructors share a small vocabulary of keyword arguments: channel_selection restricts the input reads to given channels, axes_order fixes the axes order of the patches the function sees ("yx" above), input_transforms/output_transforms apply transforms around the function, and the writing iterators take consolidation_mode (how the output pyramid is rebuilt at the end). The full signatures live in the API reference.

product replaces that single region with the ones a ROI table names — here the microscope fields of view:

# Narrow it to the regions named by a ROI table
iterator = iterator.product(ome_zarr.get_roi_table("FOV_ROI_table"))
print(iterator)
ImageProcessingIterator(rois=4)

The regions are ordinary Roi objects, so you can inspect them before processing anything:

# The regions are plain Roi objects, so you can look before you process
for roi in iterator.rois[:2]:
    print(roi)
name='FOV_1' slices=(z: 0.0->1.0, y: 0.0->351.0, x: 0.0->416.0) label=None space='world' name='FOV_2' slices=(z: 0.0->1.0, y: 0.0->351.0, x: 416.0->832.0) label=None space='world'

Beyond a ROI table, four tiling calls reshape the regions, and each name says what it tiles by. by_grid(size_x=..., size_y=..., ...) lays a regular grid of the sizes you ask for; where the grid does not divide the axis, its tail= policy decides what happens to the leftover — "clip" (the default) shrinks the last tile to the border, "balance" re-splits the last two tiles so a thin overhang never yields a thin tile (100 px at 32 gives 32, 32, 18, 18 rather than 32, 32, 32, 4), "shift" slides the last tile back to stay full-size (it then overlaps its neighbour — fine for detection or a merge; a plain parallel write schedules the overlap into a separate wave, see below, and on segmentation the overlap needs a declared resolution — with_stitch(...) or on_overlap(...)), and "drop" discards it. by_blocks(num_x=..., num_y=...) is the complement — you say how many tiles, not how big, and the partition is balanced by construction. by_chunks() tiles by the input image's chunk grid, the natural unit of reading; by_write_units() tiles by the output's write granularity — the shard shape when the output is sharded, the chunk shape otherwise, inspectable as image.write_granularity — which makes parallel writes collision-free by construction, so a parallel map runs as a single fully-parallel wave.

The four tail policies on a 100 pixel axis tiled at 32 Clip shrinks the last tile to 4 pixels. Balance re-splits the last two into 18 and 18. Shift keeps every tile full size by sliding the last one back, so it overlaps its neighbour by 28 pixels — the hatched band. Drop discards the leftover entirely, leaving three tiles. 100 PX ALONG X, TILED AT 32 0326496100 tail="clip"tail="balance"tail="shift"tail="drop" 3232324 32321818 3232 323232 a thin tileno thin tiletwo 32s overlapdiscarded

Two more calls broadcast rather than tile: by_yx() splits each region into one 2D plane per remaining coordinate (every t/z/c combination, full y/x extent — the shape for "run this 2D function on every plane"), and by_zyx(strict=...) does the same per 3D volume. The tutorial snippets use them wherever a 2D or 3D function meets a higher-dimensional image.

From here you would call the topic verb (process here) or iterate with iter_as_numpy to do the work; the image processing tutorial carries this through to a written result. (iter_as_numpy is iter(data_mode="numpy") — a bare iter() still defaults to dask and warns; numpy becomes the default in ngio=1.2.)

More complete examples can be found in the Fractal tasks template.

Parallel mapping

map runs one ROI at a time by default, and that stays the default — parallel writing is explicit opt-in. Concurrency belongs to the mapper: pass one, and it sizes its own pool.

The examples from here on run on a small synthetic image — two bright blobs:

import numpy as np
from zarr.storage import MemoryStore

from ngio import create_ome_zarr_from_array

# A small synthetic image to demonstrate on: two bright blobs.
rng = np.random.default_rng(0)
data = rng.poisson(10, size=(64, 64)).astype("uint16")
data[10:20, 26:38] += 200
data[40:50, 8:18] += 200
demo = create_ome_zarr_from_array(
    store=MemoryStore(), array=data, pixelsize=0.5, axes_names="yx", levels=1
)
demo_image = demo.get_image()
from ngio import SegmentationIterator
from ngio.iterators import ThreadedMapper
from skimage.measure import label as connected_components


def segment(patch: np.ndarray) -> np.ndarray:
    return connected_components(patch > 100).astype("uint32")


seg = demo.derive_label("seg")
seg_iterator = SegmentationIterator(demo_image, seg, axes_order="yx")

# Tiles on the OUTPUT's write grid, so the parallel map runs as one wave.
seg_iterator = seg_iterator.by_write_units()
seg_iterator.map(segment, mapper=ThreadedMapper("auto"))
print(f"labelled pixels: {int((seg.get_as_numpy() > 0).sum())}")
labelled pixels: 220

ThreadedMapper is the fit for IO-bound work and for funcs that release the GIL (most numpy/scipy do); "auto" sizes the pool for round-trip-bound work. For pure-Python, GIL-holding funcs use ProcessMapper(max_workers=...) instead — the func must be picklable (a module-level function, not a lambda), and the store must not be in-memory.

Parallel writes need no locks: the mappers schedule the ROIs into conflict-free waves, so no two concurrent writes ever touch the same chunk or shard. by_write_units() gives a single fully-parallel wave by construction; other tilings just run more waves.

How overlapping write footprints are scheduled into conflict-free waves The regions tile the whole output array, but they are not chunk-aligned, so every chunk is written by two neighbouring regions. Regions that share a chunk are placed in different waves: the first wave runs regions one, three and five, the second runs regions two and four. Within a wave no two writes touch the same chunk, so no locks are needed. CHUNK GRIDREGIONSWAVE 1WAVE 2 chunk 1chunk 2chunk 3chunk 4 12345 the regions are not chunk-aligned 135 24 Every region in a wave runs in parallel; waves run one after another.
Wave order is the canonical write order

The serial mappers run the same wave order as the parallel ones — every mapper writes the same bytes, and two runs of the same version are bit-identical. For ROIs whose pixels genuinely overlap (a by_grid stride below the size, a "shift" tail), which write wins the shared pixels is schedule-defined by default — performance first; declare write_order="roi" for the reproducible later-ROI-wins order (see the next note). Mask-protected and pixel-disjoint writes are order-independent either way, as is an order-independent merge= ("max", "min", "sum").

Reproducible seams: write_order="roi"

By default contested writes are scheduled for parallelism alone (write_order="any"): exactness always — no two overlapping writes ever run concurrently, merges always apply, stitch unification is unchanged — and a run is deterministic per ngio version (the schedule is a pure function of the ROI list, identical under every mapper). What is schedule-defined is which tile owns a contested pixel: it can differ from the manual iter loop and can change across ngio versions or retilings — pixel-visible differences, not rounding noise. Order-independent merges ("max", "min", "sum", "keep_nonzero") write identical pixels regardless, so for them the default costs nothing.

When seam ownership must be reproducible — the later ROI wins, map bit-identical to the hand-driven loops, stable across versions and mappers — opt in per declaration: on_overlap(..., write_order="roi") or StitchConfig(write_order="roi"). The price is parallelism on genuinely overlapping tilings: an overlapping grid becomes a chain of ordered neighbours, measured 1.5–2.5× slower on parallel mappers (serial runs are unaffected). Disjoint tilings (by_write_units(), halo cores, chunk grids) are bit-identical under either value. A pair is relaxed only when both sides declare "any" — one "roi" declaration keeps its ordering against everything.

Two contracts are yours: under threads the func must be thread-safe, and under processes it must be picklable. ngio's side — the per-ROI readers and writers — is safe in both settings. The dask iterator surface (iter_as_dask, map_as_dask, reduce_as_dask) is deprecated and will be removed in ngio=1.2; for lazy whole-region access use Image.get_as_dask instead.

For per-ROI measurement without writing anything, use reduce — it returns one result per ROI, in ROI order, and takes the same mapper argument:

means = seg_iterator.reduce(lambda patch: float(patch.mean()))
print([round(mean, 1) for mean in means])
[20.7]

Batched inference

For neural-network inference the per-call overhead usually dominates, and the model wants one batched (B, ...) input rather than one patch at a time. BatchedMapper does exactly that: it stacks up to batch_size patches into a single array, calls the func once per batch, and writes each result back to its own region:

from ngio import ImageProcessingIterator
from ngio.iterators import BatchedMapper


def fake_model(batch: np.ndarray) -> np.ndarray:
    # Stands in for a neural network: one (B, y, x) input, one call per batch.
    return batch // 2


halved = demo.derive_image(store=MemoryStore())
nn_iterator = ImageProcessingIterator(demo_image, halved.get_image())

# size 24 over 64 clips the border tiles to 16 - the ragged shapes are padded
# to each batch's maximum before stacking, and sliced back before the write.
nn_iterator = nn_iterator.by_grid(size_x=24, size_y=24)
nn_iterator.map(fake_model, mapper=BatchedMapper(batch_size=4))
print(halved.get_image().get_as_numpy().mean().round(2))
10.11

This is the one mapper that changes the func's contract: it receives a stacked batch with a leading batch axis and must return an array with the same leading axis. Everything else composes as usual — tilings are ragged in general (border tiles clip, halos shrink at image edges, masked regions are arbitrary), so each batch is padded up to its per-axis maximum before stacking (pad_mode/pad_values choose how) and every output is sliced back to its region's true shape before the write; a halo is trimmed after that, exactly as under any other mapper; and for_job partitions batch their own share.

Within a batch the reads fan out on a thread pool (read_workers), while the writes run serially on the calling thread. Batches are cut over the same canonical (wave) order as every other mapper, so contested pixels land identically — and the serial writes make batched mapping write-safe on any tiling.

Prefer to drive the loop yourself? iter(batch_size=...) yields (patches, writers) — two aligned lists of up to batch_size items, in ROI order — and leaves the stacking (and any ragged-tile policy) to you. The run finalizes when the loop completes, exactly like the unbatched iter; on the read-only iterators it yields the payload lists alone ((image, label, roi) tuples for features, (patch, roi) for detection):

manual = demo.derive_image(store=MemoryStore())
loop_iterator = ImageProcessingIterator(demo_image, manual.get_image())
# Uniform 32px tiles, so stacking the batch is a plain np.stack.
loop_iterator = loop_iterator.by_grid(size_x=32, size_y=32)

for patches, writers in loop_iterator.iter(data_mode="numpy", batch_size=4):
    outs = fake_model(np.stack(patches))
    for writer, out in zip(writers, outs, strict=True):
        writer(out)
print(manual.get_image().get_as_numpy().mean().round(2))
10.11

Distributed runs

A mapper parallelizes within one machine. On a cluster — SLURM array tasks over a shared filesystem, where processes cannot coordinate — restrict each task to one partition of the work instead:

n_jobs = 4
job_index = 0  # e.g. int($SLURM_ARRAY_TASK_ID)

iterator = SegmentationIterator(image, label, ...).by_chunks()
iterator.for_job(job_index, n_jobs=n_jobs).segment(func)

and, in a dependent job once every array task has finished, the gather step:

iterator = SegmentationIterator(image, label, ...).by_chunks()  # same construction
iterator.finalize()

for_job is a builder call like by_grid or with_halo — it returns a new iterator restricted to that partition's regions, and everything else reads as usual, segment(func, mapper=...) included. It comes last in the chain (reshaping a restricted iterator refuses), and a slice's segment (or map) deliberately does not finalize: the pyramid resolve is the one global step, and it belongs to the single gather job — which, thanks to region-scoped consolidation, rebuilds only what the jobs wrote. Until the gather runs, only the iterated level is up to date. (That also makes for_job(0, 1) the sanctioned way to defer a finalize on purpose: one job that writes everything, gathered whenever you choose.)

Each job builds the identical iterator — construction is metadata-only and deterministic, so this is cheap — and derives the same partition on its own; there is nothing to hand from one job to another. Partitions never share a write unit, so the jobs need no locks and no coordination, in any order and any overlap in time; regions whose footprints conflict simply travel in the same partition, where the ordinary wave planning handles them.

Effective parallelism therefore equals the number of independent groups, which follows the output's chunking. Inspect it before submitting: [it.for_job(i, n_jobs=n).partition_indices for i in range(n)]. One fat list plus empties means the output chunking (or a tiling like by_zyx, which splits along t only) is the constraint, not the cluster — a single-chunk output is one group by construction, since a chunk is one atomic write object. Surplus partitions are harmless no-ops, and write_conflict_components makes the grouping auditable.

The requirements mirror the model: every job must use the same n_jobs and the same iterator construction; the store must not be in-memory (each process would write its own private copy). The read-only iterators partition too — their global join cannot be reproduced piecewise (greedy NMS is not hierarchical: suppressing per job and then merging can keep different boxes than one global pass), so on a slice measure/detect bank a partial instead of joining, and finalize() runs the one global join and returns the table. A declared join on a slice is inert — the slice banks regardless, and the gather runs it. finalize refuses a half-finished run (a missing job errors instead of producing a plausible-looking, silently incomplete table), refuses on a for_job slice (the gather is global), and refuses when nothing was prepared or banked.

Schedulers like Fractal run distributed work as init → parallel tasks → consolidate, and prepare_jobs is the init step: it performs any setup the run needs — always wiping stale scratch state from earlier runs first — and returns the parallelization list (one JSON-ready argument set per non-empty partition, splatting into for_job(**args)). It is optional for a plain writer — the two-step recipe above works on its own — and required for a stitching segmentation (the scratch band arrays must exist, race-free, before any job banks into them) and for the read-only iterators. The distributed processing tutorial makes the whole three-phase recipe executable — partition layouts, distributed stitching, and distributed measurement.

Halos: context without seams

Tiling an image and processing each tile independently leaves artifacts at the joins — a smoothing kernel at a tile edge has no neighbours to work with, and a segmentation cuts objects at the boundary. with_halo fixes that by reading a margin around each ROI and writing only the ROI back:

How a halo reads a grown region and writes only the core The region is grown by the halo margin on the read, the function sees the grown patch, and the margin is cropped off before the write — so the written region is exactly the region you asked for. Margins clip at the image border. 123 THE ROITHE READTHE WRITE base ROIwith_halo(x=12, y=12)halo stripped on write
from scipy.ndimage import uniform_filter


def smooth(patch: np.ndarray) -> np.ndarray:
    return uniform_filter(patch, size=5)


blurred = demo.derive_image(store=MemoryStore())
blur_iterator = ImageProcessingIterator(demo_image, blurred.get_image())

# 8 px of context per side; the function returns the grown region and the
# border is cropped off before the write - no seams, same write footprints.
blur_iterator = blur_iterator.by_write_units().with_halo(x=8, y=8)
blur_iterator.map(smooth, mapper=ThreadedMapper("auto"))
print(blurred.get_image().get_as_numpy().mean().round(2))
19.94

smooth receives the grown region and must return it grown too; the border is cropped off before the write, so it never lands on disk. Margins are in pixels and clip at the image borders, so an edge tile simply grows on the sides where there is room.

Two read-only iterators take a halo too, as a pure read margin — there is no write to crop it from, so the overlap must be reconciled after the fact. Detection reconciles it itself: NMS removes the duplicate boxes. Feature extraction delegates it to you: patches and the roi argument cover the grown region, a border object is measured by every region that sees it, and the resulting duplicate label rows are yours to reconcile in a declared join (with_join) — every row carries roi_index/roi_name for exactly that (the default join keeps the duplicates as-is).

The ROIs themselves do not move, which is the whole point of doing this on the read side: write footprints are unchanged, so a haloed iterator parallelizes exactly as far as it did without one. Overlapping writes would have to be serialized; overlapping reads cost nothing.

Read the same trick backwards and it is a "trim": if you want each tile's outer margin discarded rather than written, that is exactly a halo of that width.

Overlap and reconciliation

Every iterator has exactly one reconciliation declaration — the chain call that says how per-region results become one consistent answer — and each is backed by a swappable protocol, so a custom implementation drops in without touching internals:

Iterator Declaration Required? Behind it
ImageProcessing on_overlap(policy) optional — undeclared overlap is last-writer-wins in schedule order (write_order="roi" for reproducible seams) the write path's merge= policies ("max", "sum", a callable, …)
Segmentation with_stitch(config) or on_overlap(policy) required when write footprints overlap — undeclared overlapping label writes refuse loudly SeamMatcherProtocol (StitchConfig(seam_matcher=...), IouSeamMatcher default)
MaskedSegmentation with_stitch(config) (within a mask) never — mask-protected writes cannot contest, on_overlap is refused same
FeatureExtractor with_join(join) optional — the default join keeps duplicate rows, provenance columns attached JoinProtocol (ConcatJoin default)
ObjectDetection with_nms(nms) optional — GreedyNms() by default NmsProtocol (GreedyNms default)

The segmentation rule is the one hard requirement, and it is deliberate: last-writer-wins on two overlapping label writes produces torn objects — deterministic, but almost never the intent — so the writing verbs refuse until you say what should happen. The check is pixel-exact on the write footprints: a halo never triggers it (the margin is cropped before the write), and tiles that merely share a chunk without sharing pixels pass. on_overlap("last") declares exactly the old behavior; any merge rule combines with what is on disk instead.

Stitching a tiled segmentation

Segmenting tile by tile leaves an object that crosses a boundary as two objects with two ids — and, because every tile numbers its objects from 1, leaves ids that mean nothing outside their own tile. with_stitch() fixes both:

How stitching merges objects split across tile boundaries Each tile predicts its own objects, numbered from one within that tile. Where two tiles predicted the same pixels and their objects agree above the IoU threshold, a union-find joins the two ids into one object; objects that merely abut across a cut are left as two. 1 EACH TILE IS SEGMENTED ALONE, WITH ITS OWN IDS tile Atile B 2 UNION-FIND DECIDES WHICH PAIRS ARE ONE OBJECT IoU 0.78 ≥ 0.5 IoU 0 < 0.5

Any ROI list stitches: a regular grid with a halo, an overlapping-FOV microscope layout, a ragged ROI table. The criterion is that two tiles' predictions overlap, not that their objects merely touch across a cut — two distinct objects that abut at a boundary are adjacent but do not overlap, so an adjacency rule would merge them and the overlap rule does not. The shared opinion comes from a halo (each tile reads past its own edge), from the tiles genuinely overlapping (FOV layouts need no halo — the overlap is the evidence), or both; with neither, stitching refuses, since no two tiles ever predict the same pixel.

During the map each tile banks its grown prediction into transient per-tile scratch arrays, removed once the stitch resolves. Where two tiles wrote the same output pixels and their objects did not match, the schedule picks the owner (unification itself is exact and order-free); StitchConfig(write_order="roi") makes those seams deterministically owned by the later ROI instead.

MaskedSegmentationIterator takes with_stitch() too, for tiling within a mask: a huge masked object tiled with by_grid + with_halo gets its split sub-objects merged, each tile banks only what its own mask can write, and tiles of different masks are never compared — an object cannot span two masks. Ids come out unique and dense across every object, so no UniqueLabelsTransform is needed (combining it with stitch raises).

The stitching tutorial walks the whole flow on a real image — the naive per-tile route, with_stitch(), and the StitchConfig tuning knobs (iou_threshold, block_size, scratch_store).

Detecting objects into a ROI table

Not every model produces a mask. An object detector — a YOLO network, a spot finder — reports bounding boxes, and the natural home for those is a ROI table, not a label image. The ObjectDetectionIterator runs a detector tile by tile and returns one RoiTable of the objects it found:

How the object detection iterator turns per-tile boxes into one deduplicated ROI table Each tile reads a halo past its edge so a spot at the boundary is seen whole by at least one tile. Both neighbours then report it, so non-maximum suppression keeps the higher-scoring box, and the survivors are merged into one ROI table in world coordinates. 123 DETECT PER TILEDEDUPLICATEMERGE non-maximum suppression 0.910.68 iou ≥ 0.5 labelx, yconf 1234… worldworldworldworld… 0.940.910.880.82…

NMS is declared with with_nms(GreedyNms(iou_threshold=..., score_column=...)), exactly as stitching is with with_stitch(StitchConfig(...)) — and both defaults are swappable protocols: any object with score_column, max_detections_per_tile, and a deterministic suppress(detections) satisfies NmsProtocol (soft-NMS, class-aware suppression), and a StitchConfig(seam_matcher=...) replaces the IoU criterion with your own (patch_a, patch_b) -> [(id_a, id_b), ...] pair decision. A parallel mapper= on detect fans the tiles out like any reduce.

The detector sees one tile at a time and answers in the tile's own pixels; the iterator does the bookkeeping the detector should not — anchoring each tile's boxes into the reference image's world coordinates, and resolving the boundary problem.

Box contract

  • (patch) -> list[Roi], each box in the tile's own pixels: Roi.from_values(slices={"x": (x0, width), "y": (y0, height)}, name=None, space="pixel", confidence=0.9). space="pixel" is required; a world-space box is refused — patch-local numbers in a world-labelled Roi would land every box in the wrong place silently.
  • Boxes pin x and y (optionally z), and every tile must report the same dimensionality, scored or unscored — mixtures raise.
  • Extra fields (confidence, class) ride into the table unchanged; name and label are refused — the iterator assigns them when it renumbers the survivors (put a class in an extra field, e.g. class_id).
  • A 2D detector on a 3D or timelapse image is deduplicated only within each z-slab or time point; boxes keep their tile's extent along the axes they do not pin.

The boundary problem is the sliding-window one. An object cut by a tile edge is seen only partially by either tile, so each tile reads a halo past its edge — on a read-only iterator the halo is a pure read margin, there being no write to crop it from — and the object is seen whole by at least one of them. The cost is that both neighbours now report it, and the cure is standard non-maximum suppression: boxes overlapping at or above iou_threshold (default 0.5) are one object, and the one ranked higher by the score_column ("confidence" by default; box volume when the detector reports no score) survives. Per-tile NMS inside the detector composes cleanly with this cross-tile pass. The survivors are renumbered to a dense 1..N and returned; like measure, nothing is written — storing the table is your add_table call.

The object detection tutorial runs a spot finder through detect end to end, shows the suppression happening on the raw pre-NMS boxes, and covers Roi.anchor — the coordinate call you can reuse in a custom flow.

What each iterator supports

ImageProcessing Segmentation MaskedSegmentation FeatureExtractor ObjectDetection
Writes to image label label (masked) — —
Topic verb process [5] segment [5] segment [5] measure detect
reduce ✓ ✓ ✓ ⚠ [6] ⚠ [1]
with_halo ✓ ✓ ✓ ✓ [6] ✓ [1]
with_stitch — ✓ [2] ✓ within a mask [2] — —
Overlapping ROIs ✓ [3] ✓ [3] ✓ [3] ✓ ✓ [1]
ThreadedMapper / ProcessMapper ✓ ✓ ✓ ✓ ✓
BatchedMapper ✓ ✓ ✓ ✗ [4] ✗ [4]
Distributed gather finalize finalize finalize finalize → table finalize → table
  1. Detection reconciles overlap itself: the halo is a pure read margin and NMS removes the duplicate boxes — but only through detect. reduce returns raw per-tile boxes.
  2. Stitching takes any ROI list. On the masked iterator it merges sub-objects split by a tile boundary within one mask; tiles of different masks are never compared, and ids come out unique across every object — no UniqueLabelsTransform needed (combining it with stitch raises).
  3. Overlapping writes are safe under every mapper: they are wave-scheduled and run in the same order everywhere, so a run is deterministic per version. Which write wins a contested pixel is schedule-defined by default; write_order="roi" makes it the later ROI and map bit-identical to the manual iter loop. Masked writes never contest — each touches only its own object's pixels. A commutative merge= ("max", "sum") removes the order-dependence outright. One configuration to avoid: an in-place run (same array in and out) with an overlapping tiling — reads then race the neighbouring writes; use a separate output (a halo makes this refuse outright).
  4. BatchedMapper stacks plain arrays; these iterators hand func tuple payloads.
  5. On the writers the topic verb is the generic map under its domain name; both remain available.
  6. The feature halo is a pure read margin: patches and the roi argument grow, a border object can be measured by several regions, and the duplicate rows are yours to reconcile in a declared join (with_join) via the stamped roi_index/roi_name (the default join keeps them, silently). reduce/iter read the grown regions too. Without a halo, reduce is unrestricted.

Restrictions that hold everywhere:

  • with_stitch needs a halo or overlapping ROIs, stays on the numpy path, and cannot combine with on_overlap (the stitch owns the contested pixels).
  • A haloed writer refuses reduce and read-only iter; an in-place run (same array in and out) refuses a halo.
  • by_write_units on the read-only iterators falls back to the input's chunks.
  • The deprecated dask verbs run serially, and ProcessMapper refuses in-memory stores.

Next steps