# chronozarr > An open Zarr v3 layout for raster time series. ## chronozarr v0.2 **Status:** Draft **Date:** 2026-09-30 **Spec version string:** `0.2.0` **Supersedes:** chronozarr v0.1 (`0.1.0`); `CHRONOFABRIC.md` (ChronoFabric v1: Zarr v2, one array per cell, `manifest.json`) ### 0. Purpose chronozarr is a Zarr v3 layout convention for raster time series (satellite image stacks, gridded products). A browser opens a decade of analysis-ready data from a static bucket, scrubs through time like video, and returns the exact stored value on click. It is a **convention plus a reader**. It is not a codec and not a new container: every object in a chronozarr store is a plain Zarr v3 object, readable by any Zarr v3 library that supports the store's codecs. A store uses one of two temporal encodings (§4). With `none`, the data array holds the true values of every timestep and any Zarr v3 reader reads the store without this spec. With `star-delta`, non-anchor timesteps hold residuals against an anchor timestep, and a chronozarr-aware reader is required to reconstruct them. The writer chooses between the two per store and enables `star-delta` only where it reduces compressed size (§4.3). The star-delta encoding cannot be a per-chunk Zarr codec: decoding a delta timestep needs the anchor chunk, which is a different chunk. A plain Zarr reader (`xarray.open_zarr`, zarrita) of a `star-delta` store sees anchor timesteps correctly and non-anchor timesteps as residuals in the stored dtype. This is the documented cost of that encoding. MUST, MUST NOT, SHOULD, SHOULD NOT, MAY are used as in RFC 2119. ### 1. Invariants 1. **Lossless roundtrip.** `encode(data)` then a read of level 0 returns the exact input values in the stored dtype, under either temporal encoding. No quantization, no rounding, no lossy compression of the data; coarser levels are block means of it (§6). 2. **O(1) random timestep access.** Any timestep at any level decodes with at most two chunk reads: one anchor plus one delta (one read under `none`). No chain accumulation, no sequential dependency. In a sharded store the two chunks may live in different shards (§7.4). 3. **dtype preservation.** Values are stored and served in the store's dtype (§2.3). The format never converts between integer and float and never quantizes. Physical values are derived at read time from `scale` and `offset` (§3.6); they are not stored. 4. **Self-describing.** Group and array attributes fully describe layout, temporal encoding, CRS and chunk grid. No sidecar files, no external metadata. ### 2. Store layout A chronozarr store is a Zarr v3 hierarchy. Zarr v3 arrays are leaf nodes, so each LOD level is a **group** named `"0"`, `"1"`, ... holding one data array (default name `data`, declared in `chronozarr.variable`), optional `mask` and `coverage` arrays, and coordinate arrays. This is the shape ndpyramid writes and carbonplan/zarr-layer reads (`{store}/{level}/{variable}`). ``` {store}/ zarr.json root group; attrs: multiscales, chronozarr (§3.1) volatility/ zarr.json array (grid_rows_0, grid_cols_0) float32 (§5) c/0/0 0/ zarr.json level group; attrs: crs, transform, resolution (§3.3) data/ zarr.json array (time, band, y, x), chunks (1, n_band, cs, cs) (§2.1) c/{t}/0/{r}/{c} unsharded (default): one object per (timestep, cell) (§7) c/{ts}/0/{r}/{c} sharded: one object per (time shard, cell) mask/ optional: array (time, y, x) uint8 (§2.4) zarr.json c/{t}/{r}/{c} unsharded (default); sharded is c/{ts}/{r}/{c} coverage/ optional: array (time, y, x) uint8 (§2.5) zarr.json c/{ts}/{r}/{c} time/ zarr.json c/0 int64 ms since epoch, length n_time band/ zarr.json c/0 band names, length n_band y/ zarr.json c/0 float64 projected y of pixel centres, length H_0 x/ zarr.json c/0 float64 projected x of pixel centres, length W_0 1/ ... same members; H_1 = ceil(H_0 / 2), W_1 = ceil(W_0 / 2) 2/ ... consecutive levels until the grid is 1x1 (§6) ``` #### 2.1 Data array `{level}/{variable}` | Field | Value | | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | `shape` | `[n_time, n_band, H_k, W_k]` | | `data_type` | `uint8`, `uint16`, `int16` or `float32` (§2.3). Identical at every level. | | `chunk_grid` | `regular`. Unsharded (the writer default): `chunk_shape = [1, n_band, cs, cs]`. Sharded: §7. | | `chunk_key_encoding` | `default`, separator `/`, so keys are `c/...` | | `fill_value` | `chronozarr.nodata` when it is a number; otherwise `0` | | `codecs` | `bytes` (`endian: little`) then one compression codec (§8), inside `sharding_indexed` when sharded | | `dimension_names` | `["time", "band", "y", "x"]` | | `attributes` | `_ARRAY_DIMENSIONS: ["time","band","y","x"]`; `nodata` (same value as `chronozarr.nodata`) present only when that value is a number | * `cs` (chunk size) MUST be identical at every level and even. Writers SHOULD use 256 or 512 (default 512; 256 suits low-latency stores). Readers MUST take `cs` from the data array's spatial inner chunk size, `chunk_shape[2]` (of the `sharding_indexed` codec's `chunk_shape` when sharded; `chunk_shape[3]` is equal), MUST NOT assume either value, and MUST NOT reject another even value (the reference writer accepts smaller even sizes for test fixtures). Readers MUST NOT take `cs` from `multiscales`, which carries no tile size (§3.4). * The band chunk index is always 0: one chunk holds every band of one timestep of one cell. * Cell `(r, c)` at level `k` is array slice `[:, :, r*cs:(r+1)*cs, c*cs:(c+1)*cs]`; row 0 is the top (north) edge. `grid_rows_k = ceil(H_k / cs)`, `grid_cols_k = ceil(W_k / cs)`. * **Edge chunks** are ordinary Zarr chunks: stored at full `cs x cs`; elements outside `shape` equal `fill_value`. Readers MUST discard elements beyond `shape`. There is no edge-size formula and no per-chunk size metadata. Padding is `fill_value` in anchors and 0 in residuals, so it costs almost nothing after compression. #### 2.2 Coordinate arrays `{level}/{time,band,y,x}` | Array | `data_type` | Contents | | ------ | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `time` | `int64` | Milliseconds since Unix epoch. Attrs `units: "milliseconds since 1970-01-01T00:00:00"`, `calendar: "proleptic_gregorian"` (CF), so xarray decodes to `datetime64` natively. MUST be strictly increasing. Any cadence: monthly, weekly, irregular. | | `band` | `string` (codec `vlen-utf8`) or `int32` | Band names in array order where the writer supports Zarr v3 strings; otherwise the band index `0..n_band-1`. | | `y` | `float64` | Projected y of pixel centres in `crs`, decreasing (north up). | | `x` | `float64` | Projected x of pixel centres in `crs`, increasing. | * Every array in the store MUST set `dimension_names` and the attribute `_ARRAY_DIMENSIONS` to the same list (`dimension_names` is native Zarr v3; `_ARRAY_DIMENSIONS` is what xarray-v2-era tools read). * Each coordinate array SHOULD be a single chunk. `time` and `band` MUST be identical at every level. * The timestamps as ISO-8601 strings and the band names are duplicated in root attrs `chronozarr.times` and `chronozarr.band_names` (§3.2), so a JS reader never needs int64 or vlen strings. `chronozarr.times[i]` MUST be the ISO-8601 rendering of `time[i]`; `chronozarr.band_names[i]` MUST equal `band[i]` where `band` holds names. #### 2.3 dtype and nodata profiles | `data_type` | `nodata` | `star-delta` | Pyramid mean (§6) | | ----------- | ----------------------------- | -------------------- | ----------------- | | `uint8` | integer 0 to 255, or `null` | allowed | integer | | `uint16` | integer 0 to 65535, or `null` | allowed | integer | | `int16` | integer, or `null` | not allowed (`none`) | integer | | `float32` | finite number, or `null` | not allowed (`none`) | float | * `chronozarr.nodata` is one number or `null`. `null` means the store declares no nodata value. The writer default is `0` for `uint8` and `uint16` (as in v0.1 for `uint16`) and `null` for `int16` and `float32`. * When `nodata` is a number it MUST equal every data array's `fill_value` and `nodata` attribute. When it is `null`, `fill_value` is `0` and the attribute is absent. * **Validity.** A pixel is *valid* if `mask` exists and is 1 at that pixel; or, when no mask exists and `nodata` is a number, if its value differs from `nodata`; or, when neither applies, always. A reader that has a mask MUST use it and MUST NOT also compare against `nodata`. Invalid pixels are missing in band math, statistics, block means and charts. * Where `mask` is 0, the stored data value is not interpreted; writers SHOULD store `fill_value` there. * NaN is not interpreted by this spec and is not a valid `nodata`. Writers SHOULD NOT store NaN; they SHOULD mark gaps with `nodata` or `mask`. * `bytes` for a one-byte dtype has no byte order; its `endian` configuration MAY be omitted. #### 2.4 Mask `{level}/mask` (optional) `uint8`, dims `["time", "y", "x"]`, shape `[n_time, H_k, W_k]`, values 1 (valid) and 0 (invalid), `fill_value` 0, `_ARRAY_DIMENSIONS` set. Chunked like `data` without the band axis: unsharded `[1, cs, cs]`; sharded per §7 with shard shape `[shard_time, cs, cs]` and inner chunk `[1, cs, cs]`. Keys carry no band index: `c/{t}/{r}/{c}` unsharded, `c/{ts}/{r}/{c}` sharded. * A store has a mask at every level or at none. When present, root attr `chronozarr.mask_variable` is `"mask"`; when absent, that attribute is absent. * The mask is always stored as true values: it is never star-delta encoded. * Level `k` mask is the maximum over each 2 x 2 block of level `k-1` (1 if any source pixel is valid). #### 2.5 Coverage `{level}/coverage` (optional) `uint8`, dims `["time", "y", "x"]`, shape `[n_time, H_k, W_k]`. The value is the number of valid observations behind the pixel (saturating at 255); `0` means the pixel was gap-filled or never observed. `fill_value` 0. Layout and chunking are identical to `mask` (§2.4), and it is never star-delta encoded. * A store has coverage at every level or at none. When present, `chronozarr.coverage_variable` is `"coverage"`; when absent, that attribute is absent. * Level `k` coverage is the mean of each 2 x 2 block of level `k-1`, rounded to the nearest integer with halves rounded up: `(sum + 2) // 4`. * Coverage is independent of `mask`: a gap-filled pixel is valid data with coverage 0. How gaps were filled is recorded in `chronozarr.provenance.gap_fill` (§3.8). ### 3. Attributes #### 3.1 Root group | Key | Type | Rule | | ------------- | ------ | ------------------------------------------- | | `multiscales` | list | ndpyramid schema (§3.4). Exactly one entry. | | `chronozarr` | object | §3.2. | Writers SHOULD also write zarr-python 3 `consolidated_metadata` into the root `zarr.json` so a reader learns every array's shape and codecs in one GET. Readers MUST fall back to per-array `zarr.json` when it is absent. #### 3.2 `chronozarr` block Complete instance in the root example, §3.9. | Field | Type | Rule | | ------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `spec_version` | string | `0.2.x` for stores written to this spec. Readers MUST accept `0.1.x` and `0.2.x` and MUST reject every other version. | | `variable` | string | Name of the data array inside each level group. Default `"data"`. Readers MUST NOT hardcode it. | | `times` | string\[] | ISO-8601 timestamp per timestep, length `n_time`, equal to the `time` coordinate. | | `bands` | object\[] | Band objects (§3.6) in array order, length `n_band`. Readers MUST also accept the v0.1 form, a list of strings, each read as a band object with only `name`. | | `band_names` | string\[] | `name` of each entry of `bands`, in order. Writers MUST write it. Readers that find only `bands` derive it. | | `nodata` | number or null | §2.3. | | `crs` | string | `EPSG:`. MUST equal every level's `crs` and every `multiscales.datasets[].crs`. A projected per-AOI CRS (UTM) SHOULD be used. Web Mercator SHOULD NOT: its pixels are not equal-area, so block means and statistics are biased. | | `temporal` | object | §4.1. | | `volatility_path` | string | Path of the volatility array relative to the root. `"volatility"` in v0.2. | | `mask_variable` | string | Optional. `"mask"` when `{level}/mask` exists (§2.4). | | `coverage_variable` | string | Optional. `"coverage"` when `{level}/coverage` exists (§2.5). | | `provenance` | object | Optional. §3.8. | | `levels` | object\[] | §3.7. Writers MUST write it; readers MUST fall back to level metadata when it is absent. | | `shard_bytes` | object | Optional, sharded stores only. §7.3. | The temporal block applies to every level: anchor schedule and references are shared across the pyramid. #### 3.3 Level group attributes | Key | Rule | | ------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `crs` | Same string as `chronozarr.crs`. | | `transform` | 6 floats, rasterio/Affine order `[a, b, c, d, e, f]`: `x = a*col + b*row + c`, `y = d*col + e*row + f` for pixel corners. Level `k`: `[a*2^k, b, c, d, e*2^k, f]` (origin preserved). | | `resolution` | Ground sample distance in CRS units at this level: `abs(a) * 2^k`. | The data array SHOULD additionally carry `proj:code`, `spatial:dimensions` (`["y","x"]`), `spatial:shape`, `spatial:transform` and `spatial:bbox` (zarr-conventions `proj` and `spatial`, same values). zarr-layer reads these to detect a non-4326/3857 CRS instead of inferring it from bounds magnitude. When `crs` is an EPSG code the writer SHOULD also set `_CRS` on the array, `{ "url": "http://www.opengis.net/def/crs/EPSG/0/" }` optionally with a `wkt` member: GDAL's Zarr driver reads this attribute to assign a CRS (its documentation names it the only CRS convention supported before GDAL 3.13). The array MAY also repeat `crs` and `transform`. Readers MUST NOT require any of these. #### 3.4 `multiscales` (ndpyramid schema) One entry with `datasets[]`, `type` and `metadata` (instance in §3.9). * `datasets[i].path` is the level group name; `crs` equals `chronozarr.crs`. Writers emit only `path` and `crs` in an entry. * Writers MUST NOT write `pixels_per_tile`. Readers MUST ignore it if present: stores written before this rule carry it, and it is never the source of `cs` (§2.1). CarbonPlan zarr-layer 0.10.0 reads the key as the marker of a global Web Mercator slippy-map pyramid; with it present a native-CRS store opens as a blank map, and without it zarr-layer takes the extent from `proj:code`, `spatial:transform` and `spatial:bbox` (§3.3). The cell size is carried by the chunk shape of each level's data array (§2.1) and the cell grid by `chronozarr.levels` (§3.7), so nothing is lost. A validator MUST accept a store with or without the key. * Levels MUST be listed in order `"0"`, `"1"`, ... with no gaps. * `type` and `metadata.{method, version, args}` follow the ndpyramid schema page verbatim; `method` is nested under `metadata`, not top-level. `metadata.version` is `chronozarr 0.2.0` for new stores. * The zarr-conventions `multiscales` (object with `layout[]`) uses the same key with a different shape; a store cannot carry both. v0.2 keeps the ndpyramid list form, without `pixels_per_tile`, because zarr-layer reads it that way and the GeoZarr layout is still a proposal. Revisit when GeoZarr's multiscale layout is stable. #### 3.5 Mirrors Three root attributes duplicate information that also lives in arrays or level groups, so a reader needs no vlen strings, int64 or per-level requests: `chronozarr.times` (mirrors the `time` coordinate), `chronozarr.band_names` (mirrors the `band` coordinate and `chronozarr.bands`) and `chronozarr.levels` (mirrors level attributes and array shapes, §3.7). A mirror MUST agree with its source; a validator checks this. Readers use the mirror as the primary source and MUST NOT require the source array to be read. #### 3.6 Band objects Each entry of `chronozarr.bands` is an object: | Field | Type | Rule | | ------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------- | | `name` | string | Required. Unique within the store. Equals the `band` coordinate value where that holds names. | | `common_name` | string | Optional. SHOULD come from the STAC `eo:bands` common-name vocabulary where one applies (`red`, `green`, `blue`, `nir`, `swir16`, `swir22`, ...). | | `scale` | number | Optional, default 1. | | `offset` | number | Optional, default 0. | | `units` | string | Optional. Free text, for example `"reflectance"`. | * Physical value = `stored * scale + offset`. Invalid pixels (§2.3) are not scaled; they are missing. * The v0.1 string form carries no `scale`, so it reads as `scale` 1. A consumer that knows the source may supply one: the JavaScript reader assumes 0.0001, the Sentinel-2 L2A reflectance scale, for `0.1.x` stores. New stores SHOULD write `scale` and `offset` explicitly. * chronozarr does not set CF `scale_factor` or `add_offset` on the data array, so plain Zarr readers see stored values and scale-aware decoding is the consumer's choice. * Consumers that do band math or compute indices SHOULD select bands by `common_name`, falling back to `name`, and MUST apply `scale` and `offset` rather than assume a fixed reflectance scale. An RGB `uint8` store with bands `red`, `green`, `blue` and `scale` 1 therefore renders as stored. #### 3.7 `levels` mirror `chronozarr.levels` is a list with one object per level, in level order, mirroring level metadata so a reader opens the store with one GET: | Field | Type | Rule | | ------------ | ---------- | ------------------------------------------------------------------- | | `path` | string | Level group name; equals `multiscales[0].datasets[i].path`. | | `resolution` | number | Equals the level group's `resolution`. | | `transform` | number\[6] | Equals the level group's `transform`. | | `shape` | int\[4] | `[n_time, n_band, H_k, W_k]`; equals the data array's `shape`. | | `grid` | int\[2] | `[rows, cols]` of cells; equals `[ceil(H_k / cs), ceil(W_k / cs)]`. | The mirror MUST agree with the level group and array metadata; a validator checks it. #### 3.8 `provenance` (optional) ```json { "sources": ["sentinel-2-l2a"], "composite": "monthly median", "gap_fill": "carry-forward", "notes": "SCL classes 4, 5, 6, 7 and 11 kept" } ``` | Field | Type | Rule | | ----------- | --------- | ---------------------------------------------------------------------------------------------------------------- | | `sources` | string\[] | Required. Collection ids or URLs of the input data. | | `composite` | string | Required. Free text describing the temporal composite, for example `"monthly median"`. | | `gap_fill` | string | Required. `"carry-forward"` (a pixel with no valid observation takes the previous timestep's value) or `"none"`. | | `notes` | string | Optional. | #### 3.9 Complete example Store: 5 monthly timesteps, 2 bands, 700 x 600 pixels, `cs = 512`, EPSG:32631 at 10 m, `star-delta` with `anchor_interval = 2`, sharded with `shard_time = 4` (two time shards; the writer default is unsharded, §7.1), with `mask` and `coverage`. LOD 0 grid is 2 x 2; LOD 1 is 350 x 300, grid 1 x 1, so the pyramid stops there. Shard byte lengths are illustrative. The example uses `zstd` level 5, the writer default (§8). `zarr.json` (root): ```json { "zarr_format": 3, "node_type": "group", "attributes": { "multiscales": [{ "datasets": [ { "path": "0", "crs": "EPSG:32631" }, { "path": "1", "crs": "EPSG:32631" } ], "type": "reduce", "metadata": { "method": "block_mean", "version": "chronozarr 0.2.0", "args": [] } }], "chronozarr": { "spec_version": "0.2.0", "variable": "data", "times": ["2024-01-01T00:00:00Z", "2024-02-01T00:00:00Z", "2024-03-01T00:00:00Z", "2024-04-01T00:00:00Z", "2024-05-01T00:00:00Z"], "bands": [ { "name": "B04", "common_name": "red", "scale": 0.0001, "offset": 0.0, "units": "reflectance" }, { "name": "B08", "common_name": "nir", "scale": 0.0001, "offset": 0.0, "units": "reflectance" } ], "band_names": ["B04", "B08"], "nodata": 0, "crs": "EPSG:32631", "temporal": { "encoding": "star-delta", "anchor_interval": 2, "anchor_indices": [0, 2, 4], "delta_reference": { "1": 0, "3": 2 }, "selection": { "mode": "auto", "sampled_cells": 3, "ratio": 0.78 } }, "volatility_path": "volatility", "mask_variable": "mask", "coverage_variable": "coverage", "provenance": { "sources": ["sentinel-2-l2a"], "composite": "monthly median", "gap_fill": "carry-forward", "notes": "SCL classes 4, 5, 6, 7 and 11 kept" }, "levels": [ { "path": "0", "resolution": 10.0, "transform": [10.0, 0.0, 746090.0, 0.0, -10.0, 2540440.0], "shape": [5, 2, 700, 600], "grid": [2, 2] }, { "path": "1", "resolution": 20.0, "transform": [20.0, 0.0, 746090.0, 0.0, -20.0, 2540440.0], "shape": [5, 2, 350, 300], "grid": [1, 1] } ], "shard_bytes": { "0": { "0/0/0": 1412807, "0/0/1": 388120, "0/1/0": 1290455, "0/1/1": 351902, "1/0/0": 402211, "1/0/1": 101347, "1/1/0": 377689, "1/1/1": 96410 }, "1": { "0/0/0": 388554, "1/0/0": 98203 } } } } } ``` `0/zarr.json` (level group): `{ "zarr_format": 3, "node_type": "group", "attributes": { "crs": "EPSG:32631", "transform": [10.0, 0.0, 746090.0, 0.0, -10.0, 2540440.0], "resolution": 10.0 } }` `0/data/zarr.json` (sharded, two time shards of up to 4 timesteps): ```json { "zarr_format": 3, "node_type": "array", "shape": [5, 2, 700, 600], "data_type": "uint16", "chunk_grid": { "name": "regular", "configuration": { "chunk_shape": [4, 2, 512, 512] } }, "chunk_key_encoding": { "name": "default", "configuration": { "separator": "/" } }, "fill_value": 0, "codecs": [{ "name": "sharding_indexed", "configuration": { "chunk_shape": [1, 2, 512, 512], "codecs": [{ "name": "bytes", "configuration": { "endian": "little" } }, { "name": "zstd", "configuration": { "level": 5, "checksum": false } }], "index_codecs": [{ "name": "bytes", "configuration": { "endian": "little" } }, { "name": "crc32c" }], "index_location": "end" } }], "dimension_names": ["time", "band", "y", "x"], "attributes": { "_ARRAY_DIMENSIONS": ["time", "band", "y", "x"], "nodata": 0 } } ``` Shard objects: `0/data/c/{0,1}/0/{0,1}/{0,1}` (8 objects). Timesteps 0 to 3 are in time shard 0 (inner position `t`), timestep 4 is in time shard 1 at inner position 0. Shard `c/1/0/r/c` holds one timestep, but its index still has `shard_time = 4` entries (16 x 4 + 4 = 68 bytes), three of them empty. Shards at `r=1` hold rows 512..699 plus 324 rows of fill; shards at `c=1` hold columns 512..599 plus 424 columns of fill. `0/mask/zarr.json` (`coverage` differs only by name and `_ARRAY_DIMENSIONS` is identical): ```json { "zarr_format": 3, "node_type": "array", "shape": [5, 700, 600], "data_type": "uint8", "chunk_grid": { "name": "regular", "configuration": { "chunk_shape": [4, 512, 512] } }, "chunk_key_encoding": { "name": "default", "configuration": { "separator": "/" } }, "fill_value": 0, "codecs": [{ "name": "sharding_indexed", "configuration": { "chunk_shape": [1, 512, 512], "codecs": [{ "name": "bytes" }, { "name": "zstd", "configuration": { "level": 5, "checksum": false } }], "index_codecs": [{ "name": "bytes", "configuration": { "endian": "little" } }, { "name": "crc32c" }], "index_location": "end" } }], "dimension_names": ["time", "y", "x"], "attributes": { "_ARRAY_DIMENSIONS": ["time", "y", "x"] } } ``` Shard objects: `0/mask/c/{0,1}/{0,1}/{0,1}` (8 objects), `0/coverage/c/{0,1}/{0,1}/{0,1}` (8 objects). `1/zarr.json` has `"transform": [20.0, 0.0, 746090.0, 0.0, -20.0, 2540440.0]`, `"resolution": 20.0`. Level 1 arrays differ from level 0 only in `shape` (`[5, 2, 350, 300]`, `[5, 350, 300]`). Shard objects: `1/data/c/{0,1}/0/0/0`, `1/mask/c/{0,1}/0/0`, `1/coverage/c/{0,1}/0/0`. `0/time/zarr.json`: ```json { "zarr_format": 3, "node_type": "array", "shape": [5], "data_type": "int64", "fill_value": 0, "chunk_grid": { "name": "regular", "configuration": { "chunk_shape": [5] } }, "chunk_key_encoding": { "name": "default", "configuration": { "separator": "/" } }, "codecs": [{ "name": "bytes", "configuration": { "endian": "little" } }, { "name": "zstd", "configuration": { "level": 5, "checksum": false } }], "dimension_names": ["time"], "attributes": { "_ARRAY_DIMENSIONS": ["time"], "units": "milliseconds since 1970-01-01T00:00:00", "calendar": "proleptic_gregorian" } } ``` `0/time/c/0` decodes to `[1704067200000, 1706745600000, 1709251200000, 1711929600000, 1714521600000]`. `band`, `x` and `y` follow the same pattern with `data_type` `string` (codec `vlen-utf8`) or `int32`, `float64` and `float64` respectively. `volatility/zarr.json`: `shape [2, 2]`, `data_type "float32"`, `chunk_shape [2, 2]`, `fill_value 0.0`, `dimension_names ["row", "col"]`. A `none` store differs in the `temporal` block only (`{ "encoding": "none" }`, optionally with `selection`); the arrays, layout and every other attribute are the same. ### 4. Temporal encoding #### 4.1 `temporal` block | Field | Type | Rule | | ----------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `encoding` | string | `"none"` or `"star-delta"`. Readers MUST support both. | | `anchor_interval` | int >= 1 | `star-delta` only. Timesteps between anchors. `1` means every timestep is an anchor (no deltas). | | `anchor_indices` | int\[] | `star-delta` only. `[0, k, 2k, ...]` for `k = anchor_interval`, `< n_time`. | | `delta_reference` | object | `star-delta` only. Key: non-anchor timestep index as a decimal string (JSON object keys). Value: the anchor index it references (§4.2), at index distance `d` with `0 < d < anchor_interval`. Every non-anchor index MUST appear and no anchor may. | | `selection` | object | Optional. `{ "mode": "auto", "sampled_cells": int, "ratio": number }`, written only when the writer chose the encoding by measurement (§4.3). | With `"none"`, the data array holds true values and readers do nothing beyond §10. `anchor_interval`, `anchor_indices` and `delta_reference` MUST be absent. `star-delta` is valid only for `uint8` and `uint16` data (§2.3). #### 4.2 Star-delta Applies identically at every level. 1. **Anchor indices:** `[0, anchor_interval, 2*anchor_interval, ...]` while `< n_time`. 2. **Delta reference:** each non-anchor timestep references one anchor, recorded in `delta_reference`. A reference is valid when it names an anchor at index distance `d` with `0 < d < anchor_interval`; readers MUST decode through the recorded map and MUST NOT recompute it. A writer SHOULD reference the nearest anchor by absolute index distance, ties toward the earlier anchor, among the anchors that exist when it writes the timestep: a fresh encode considers every anchor `< n_time`, an append the anchors that exist after the append, including those in the appended batch. A writer MUST NOT change the reference of a timestep that is already written. Caches and open readers hold its bytes, and a changed reference would change what those bytes decode to (§14). The nearest-anchor rule is therefore the default for new timesteps (it gives the smallest residuals), not a validity condition. 3. **Anchor storage:** `data[t] = source[t]` (true values). 4. **Delta storage:** `data[t] = (source[t] - source[a]) mod 2^bits`, where `a` is the reference anchor and `bits` is 8 or 16: unsigned wraparound subtraction in the stored dtype. Anchors and deltas live in the same array. There is no clipping, no range check and no failure mode: modular arithmetic reconstructs every input exactly, so Invariant 1 holds for every input. When `|source[t] - source[a]| < 2^(bits-1)` the stored bits equal the two's-complement signed difference, which is what v0.1 stored for `uint16` (int16 residuals viewed as uint16); every v0.1 star-delta store is therefore also a valid v0.2 store. Decoding: ``` if encoding == "none" or t in anchor_indices: return data[t] # true values, use directly else: anchor_t = delta_reference[str(t)] return (data[anchor_t] + data[t]) mod 2^bits # unsigned wraparound in the stored dtype ``` Readers MUST NOT clamp. A residual is not special-cased for `nodata`: the stored residual of a nodata pixel is `(nodata - anchor) mod 2^bits`, and reconstruction returns `nodata` exactly. **Why star, not chain.** Chain-delta (each timestep references the previous) gives roughly 10-15% better compression but requires sequential decoding. Star-delta gives O(1) random access to any timestep, which is what scrubbing needs. #### 4.3 Writer selection The writer option `encoding` takes `auto` (default), `none` or `star-delta`. * `auto` applies to `uint8` and `uint16` only. For `int16` and `float32` the writer uses `none` and writes no `selection`. * `auto` encodes a sample of LOD 0 cells both ways with the codec the store will use, at least 3 cells or all cells when there are fewer, and keeps `star-delta` only if its total compressed bytes are at most 0.85 times the plain total. It records `selection = { "mode": "auto", "sampled_cells": n, "ratio": r }`, where `r` is star-delta compressed bytes divided by plain compressed bytes over the sampled cells. * `none` and `star-delta` MAY be forced; a forced choice writes no `selection`. * Basis for the threshold: star-delta reduced compressed size by about 25% on arid scenes and about 6% on vegetated scenes (2026-09-30). A plain store reads correctly in xarray and zarrita without an adapter, and in zarr-layer too when it carries no `pixels_per_tile` (§3.4, §12), so the extra decode path is worth carrying only where the saving is material. ### 5. Volatility A float32 array at `volatility_path` (`volatility`) in the root group, shape `(grid_rows_0, grid_cols_0)`, one value per LOD 0 cell, single chunk. ``` volatility[r, c] = clip( mean(|source[t] - source[ref(t)]| over all non-anchor t, all bands, all pixels of cell (r, c)) / 10000 , 0, 1) ``` * `ref(t)` is `delta_reference[t]` for `star-delta` stores. For `none` stores the writer evaluates the same expression against a nominal nearest-anchor schedule (anchor interval chosen by the writer, default 6); that schedule is not recorded and readers MUST NOT depend on it beyond ordering. * Computed on exact differences (int32 for integer dtypes, float64 for `float32`), over all pixels including invalid ones, at LOD 0 only. * `0.0` when the cell has no delta timesteps (`anchor_interval = 1` or `n_time = 1`). * The divisor 10000 is a fixed normalization constant for every dtype (it is the Sentinel-2 reflectance scale). It is not a physical unit; for other sources the value is still monotone in temporal change. Readers use it to order prefetch (volatile cells first) and to draw change overviews. A store without it is decodable but not conforming. ### 6. Multiscale pyramid Each level halves pixel dimensions and doubles ground sample distance by block averaging that excludes invalid pixels (§2.3). A block whose pixels are all invalid yields `nodata` (`0` when `nodata` is `null`). * Level 0: native resolution (10 m for Sentinel-2). Level `k`: resolution `* 2^k`. * Level `k` is derived from level `k-1` by a factor-2 block mean. Before averaging, the source is padded to an even height and width by edge replication, so `H_k = ceil(H_{k-1} / 2) = ceil(H_0 / 2^k)`, same for `W`. * The mean of integer dtypes (`uint8`, `uint16`, `int16`) is computed on the exact sum in a wider integer type (`uint32` or `int32`) with floor division by the valid-pixel count, and stored in the data dtype. The mean of `float32` accumulates in float64 and is stored as float32. * `mask` reduces by maximum and `coverage` by rounded mean (§2.4, §2.5). * Downsampling is applied to source timesteps; star-delta is then applied per level. Deltas are never downsampled. * **Chunk size stays constant.** Only the grid shrinks: `grid_rows_k = ceil(H_k / cs)`. * Levels MUST be consecutive from 0. The encoder default stops at the first level whose grid is 1 x 1 (that level is included). Fewer or more levels MAY be present. **Categorical and binary bands.** Coarser levels are block means. That is correct for continuous quantities and for fractions, and wrong for a class code or a 0/1 flag stored as the integers 0 and 1: an integer mean floors, so a block with 3 of 4 pixels set becomes 0, and the flagged share shrinks with every level (a water flag fell from 12.1 % at level 0 to 8.5 % at level 3 in one test). A binary band SHOULD be stored as a scaled fraction, for example 0 or 10000 with `scale` 1e-4 and `units` `"fraction"` (§3.6), so that each coarse value is the fraction of the block that is set and a reader can threshold it. A categorical band has no meaningful mean at all. A writer that needs true categorical overviews must downsample them itself, with a class rule such as majority or nearest, before calling the encoder; the encoder always builds its pyramid by block mean, so each such overview is a store of its own. ### 7. Sharding #### 7.1 Layout **Default: unsharded.** Each chunk `(1, n_band, cs, cs)` is its own object, key `c/{t}/0/{r}/{c}`: one object per timestep, cell and level (about 5,900 for a 117-month store of 36 level-0 cells). A timestep of a cell is one `GET` of one object, two for a `star-delta` delta timestep. There is no index to read, a CDN miss costs one chunk, and an append writes only new objects (§14). **Option: sharded.** The array uses the `sharding_indexed` codec with shard shape `(shard_time, n_band, cs, cs)` and inner chunk shape `(1, n_band, cs, cs)`. A shard object holds one spatial cell over `shard_time` consecutive timesteps (93 objects for the same store), and once its index is cached one timestep is one HTTP byte-range read. The index is a separate read at the end of an object that can be tens or hundreds of megabytes, and a CDN that fetches the whole object from the origin on a miss makes that read as slow as the object is large, so the cost of a cold read grows with the shard size (§13 gives the measurement). An append rewrites the shard that receives the new timestep (§14). Zarrita 0.7.5 read sharded stores over byte ranges with bytes transferred equal to the unsharded store and bit-exact reconstruction (spike, 2026-09-29). A sharded array: ```json "chunk_grid": { "name": "regular", "configuration": { "chunk_shape": [4, 2, 512, 512] } }, "codecs": [{ "name": "sharding_indexed", "configuration": { "chunk_shape": [1, 2, 512, 512], "codecs": [{ "name": "bytes", "configuration": { "endian": "little" } }, { "name": "zstd", "configuration": { "level": 5, "checksum": false } }], "index_codecs": [{ "name": "bytes", "configuration": { "endian": "little" } }, { "name": "crc32c" }], "index_location": "end" } }] ``` * In Zarr v3 terms the array's `chunk_grid.chunk_shape` is the **shard** shape and the codec's `chunk_shape` is the **inner chunk** shape. Inner chunks are the same bytes an unsharded store would hold as separate objects. * **`shard_time`** is an integer `>= 1`; the writer default, when sharding is asked for, is `n_time` (one shard holds a cell's whole time axis). It MAY exceed `n_time`: the array then holds one partial shard, which a later append (§14) fills. The time grid has `n_shards_t = ceil(n_time / shard_time)` shards, so the shard grid of the data array is `(n_shards_t, 1, rows, cols)`. Readers MUST handle more than one shard along time. * Timestep `t` lives in time shard `ts = floor(t / shard_time)` at inner position `t mod shard_time`. The shard key is `c/{ts}/0/{r}/{c}`. When `shard_time = n_time`, `ts` is always 0 and the keys are `c/0/0/{r}/{c}`, as in v0.1. `mask` and `coverage` have no band axis: shard shape `(shard_time, cs, cs)`, key `c/{ts}/{r}/{c}`. * Objects per level drop from `n_time * grid_rows * grid_cols` to `n_shards_t * grid_rows * grid_cols`. * The writer SHOULD choose `shard_time` so that no shard object exceeds the largest object the intended host or CDN will cache or range-serve, and so that a miss on one shard is tolerable (a CDN that pulls the whole object for a small range read charges the shard's size for its 2 KB index). For `star-delta` stores it SHOULD choose a multiple of `anchor_interval`. * **Both forms are valid.** A reader MUST handle both; it learns which from the `codecs` list. Every store written with the earlier sharded default stays valid. Unsharded (§2.1) also works on hosts without byte-range support. #### 7.2 Shard index * The index holds one `(offset, nbytes)` uint64 pair per inner chunk of the shard: `16 * shard_time` bytes plus 4 bytes crc32c, `N = 16 * shard_time + 4`. A partial last time shard is stored at full shard shape: its index still has `shard_time` entries, and entries for timesteps `>= n_time` are empty. * With `index_location: "end"` a reader fetches the index with the last `N` bytes of the shard; with `"start"`, the first `N` bytes. Readers MUST support both. Writers SHOULD use `"end"`: zarrita 0.7.5 ignores `index_location` and decodes start-indexed shards as garbage without error, which would break every stock zarrita reader including CarbonPlan zarr-layer. * Without `shard_bytes` (§7.3), a reader fetches the end-located index with a suffix range (`Range: bytes=-N`, which triggers a CORS preflight) or learns the object length with a `HEAD` first. Zarrita issues one `HEAD` before the index read, so a cold cell costs `HEAD`, index, chunk, then one range per timestep. * Readers MUST cache the index per shard and reuse one array handle per level (zarrita caches the index per array instance). * Empty inner chunks (both index values `2^64 - 1`) decode as all `fill_value`. A shard object that does not exist (`404`) is an entirely empty shard, and readers MUST decode it as all `fill_value`; writers MAY omit shards that are entirely fill. #### 7.3 `shard_bytes` Optional root attribute `chronozarr.shard_bytes`, sharded stores only, covering the data array (not `mask` or `coverage`): ```text { "": { "//": } } ``` * It lists exactly the shard objects of the data array that exist, for every level. * With the length of an end-located shard known, a reader issues the index read as the bounded range `bytes=(L-N)-(L-1)`, needing no `HEAD` and no preflight. Readers MUST use `shard_bytes` when present and MUST fall back to a `HEAD` for any shard it does not list. * Because shard objects are immutable (§9), the lengths cannot go stale, with one exception: an append (§14) replaces the shard that receives a new timestep, and its listed length changes with it. #### 7.4 Star-delta across shards A non-anchor timestep and its reference anchor can sit in different time shards (for example the last timesteps before a shard boundary whose nearest anchor is the first timestep of the next shard). Such a read still touches two chunks (Invariant 2) but two shards, so a cold read also pays the second shard's index. Choosing `shard_time` as a multiple of `anchor_interval` keeps almost every delta and its anchor in one shard. ### 8. Compression codec | Codec | Status | Configuration | | ------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `zstd` | Default | `{ "level": 5, "checksum": false }` | | `blosc` | Permitted | `cname` `zstd` or `lz4`; `shuffle` `noshuffle` or `shuffle`; `clevel` 0 to 9; `typesize` the element size in bytes of the data dtype; `blocksize` 0 | | `gzip` | Permitted | `{ "level": 1..9 }` | * Readers MUST support all three, including both `cname` values and both `shuffle` values of `blosc`. Writers MUST use exactly one, preceded by `bytes` little-endian. * Other codecs, other `blosc` `cname` values and `shuffle: "bitshuffle"` MUST NOT be used: the browser reader carries no WASM dependency for them. * The same codec chain applies to `data`, `mask` and `coverage`; the `time`, `band`, `x`, `y` and `volatility` arrays MAY use any of the three. Rationale: size and browser decode speed were measured, not assumed. On one real Sentinel-2 chunk (4 x 512 x 512 uint16, 2,097,152 bytes, median of 15 in Chromium, 2026-09-29) zstd via the numcodecs-js WASM codec that zarrita uses decoded in 4.5 ms at 1,284,781 bytes, native gzip via `DecompressionStream` in 6.7 ms at 1,387,409 bytes, and zstd via the pure-JS fzstd in 14.1 ms. On four real Ucayali LOD 0 chunks of 2,097,152 bytes (zarrita 0.7.5 with vendored codec modules served locally, headless Chrome, median of 20, 2026-09-30), zstd level 5 stored 1,376,469 bytes and `blosc` with zstd level 1 and byte shuffle stored 1,520,154 bytes, 10.4% more. Steady-state decode took 4.4 ms for zstd and 3.5 ms for `blosc` (4.4 and 3.3 ms inside a worker), and every reconstruction was bit-exact. Byte shuffle did not reduce the size of the zstd stream on this data. The 10.4% size penalty alone decides the default: zstd level 5 stores the fewest bytes, and the decode difference is about 1 ms per chunk. `blosc` stays permitted for writers that prefer encode speed or a faster steady-state decode, and `gzip` for writers without a zstd implementation. Encode cost is irrelevant to the default. ### 9. Static hosting requirements A chronozarr store is served by any HTTP server or object store that returns files by path. No server-side code. Host recipes and a header checklist are in `docs/hosting.md`. | Requirement | Level | | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------- | | `GET {store}/{key}` returns the object bytes; an absent key returns `404` (not a fallback page with `200`) | MUST | | `Range` requests honoured with `206 Partial Content` and `Content-Range` (needed for sharded stores; an unsharded store, the writer default, needs only GET) | MUST for sharded | | Chunk and shard objects served as stored: no `Content-Encoding` or other transformation, so byte offsets match the shard index | MUST | | `Access-Control-Allow-Origin: *` | MUST | | `Access-Control-Allow-Headers: Range` (or `*`) and a `200`/`204` answer to `OPTIONS` preflight | SHOULD | | `Access-Control-Expose-Headers: Content-Range, Content-Length` | SHOULD | | `Cache-Control: public, max-age=31536000, immutable` | SHOULD | | `Timing-Allow-Origin: *` | MAY | * Browsers do not preflight a bounded `Range: bytes=a-b`; only suffix ranges `bytes=-N` trigger one. A reader that uses `shard_bytes` (§7.3) issues only bounded ranges. * Without `Timing-Allow-Origin`, `PerformanceResourceTiming.transferSize` is 0 for cross-origin requests, so a reader can count bytes only from `Content-Length`. * Stores are treated as immutable. A re-encode MUST be written under a new prefix, never in place. The one in-place change is an append (§14), which leaves every existing chunk's meaning unchanged and replaces only the objects §14 lists. * Upload order SHOULD be: all chunk and shard objects, then group and array `zarr.json` below the root, then the root `zarr.json` last. A reader treats the root `zarr.json` as the marker that the store exists, so it never describes missing data. * Zarr v3 metadata is `zarr.json`, never a dotfile, so hosts that hide dotfiles (GitHub Pages, some CDNs) serve a store correctly. No `.zarray`, `.zattrs` or `.zmetadata` exist. * Directory listing is never required (a reader derives every key from metadata) and object content types are ignored. ### 10. Reader requirements A conforming reader MUST: 1. `GET {store}/zarr.json`. Reject the store if `attributes.chronozarr.spec_version` is not `0.1.x` or `0.2.x`, or if `temporal.encoding` is not `"none"` or `"star-delta"` (a v0.1 store is `star-delta`). Take timestamps from `chronozarr.times`, band objects from `chronozarr.bands` (strings read as `{ name }`), the data array name from `chronozarr.variable`, and the validity rule inputs from `chronozarr.nodata`, `mask_variable` and `coverage_variable`. 2. Enumerate levels from `chronozarr.levels` when present; otherwise from `multiscales[0].datasets[].path`, reading each level's `zarr.json` (`transform`, `resolution`). Read `{variable}/zarr.json` (`shape`, `data_type`, `chunk_grid`, `codecs`; `cs` is the spatial inner chunk size, §2.1) for each level, or take them from consolidated metadata when present. Fail with an error naming any unsupported `data_type` or codec. 3. Select a level: the largest `k` whose `resolution` does not exceed the requested output ground sample distance; `k = 0` if none. 4. For timestep `t` and cell `(r, c)`: under `none`, or when `t` is in `anchor_indices`, read one chunk. Otherwise read the chunk at `delta_reference[str(t)]` and the chunk at `t`. Never more than two reads (Invariant 2). 5. Reconstruct per §4.2: unsigned wraparound add in the stored dtype, no clamp. Discard elements beyond `shape` (§2.1). 6. For sharded arrays, compute the shard and inner position from `shard_time` (§7.1), read the shard index (§7.2, using `shard_bytes` when present), then one byte range per inner chunk. Decode a missing shard as `fill_value`. 7. Apply the validity rule (§2.3): invalid pixels are missing in band math and statistics. Read `{level}/mask` and `{level}/coverage` only when `mask_variable` and `coverage_variable` are present. 8. Apply `scale` and `offset` (§3.6) wherever physical values are shown or combined. A reader SHOULD cache decoded anchors and shard indices for the session, and SHOULD prefetch anchors before deltas: once every anchor in the viewport is resident, any timestep costs one delta read. Under `none` there are no deltas and prefetch order is free. Reference implementations: the Python package `chronozarr` (`open_store(path_or_url)` returns a reader with `read(t)`, `read_cell(t, row, col)` and `to_xarray()`; `validate(store)` checks a store against this spec) and the JavaScript reader `js/chronozarr/decoder.js` (zarrita store, reconstruction, cache, prefetch), DOM-free so it runs in a Worker, under MapLibre, or on a bare canvas. A plain Zarr reader (`xarray.open_zarr(store, group="0")`, zarrita) reads a `none` store without this spec. For a `star-delta` store it reads anchor timesteps correctly and non-anchor timesteps as residuals. ### 11. What this is not * Not a new binary format. It is Zarr v3 with a layout convention. * Not a codec. Star-delta spans chunks, so it cannot sit in a Zarr `codecs` list; it lives in group attributes and the reader. * Not a video codec. Deltas are block-compressed arrays, not I/P/B frames. * Not a spatial index. A regular grid of chunks with a power-of-2 pyramid. * Not a server or an API. There is nothing to run; a bucket is the deployment. * Not a viewer. chronozarr is one consumer; zarr-layer and xarray are others. ### 12. Relationship to prior art A row-by-row comparison with guidance on when to choose each tool is in `docs/format-comparison.md`. * **PMTiles** (Protomaps): single-file archive of z/x/y tiles for static hosting. The hosting model (one bucket, range reads, no server) is the same. Tiles are images; per-tile values are whatever the image encoding carries, and there is no native time axis. * **Mapbox raster-array (MRT)**: multi-band numeric raster tiles with a time-like band dimension, decoded client-side. The decoder code is published in mapbox-gl-js (`src/data/mrt`); the format is produced by the Mapbox Tiling Service and consumed by Mapbox's renderer, so using it means using that service and renderer. chronozarr is a Zarr-based alternative that a static bucket serves. * **carbonplan ndpyramid + zarr-layer**: Zarr pyramids rendered in MapLibre with a time selector. chronozarr writes the same `multiscales` attribute and level/variable shape, without `pixels_per_tile`: zarr-layer reads that key as the marker of a global Web Mercator pyramid (§3.4). A `none` store written without it opens in zarr-layer 0.10.0 unmodified (verified on the Ucayali store at level 1; point values equal to a direct Zarr read, `docs/comparisons.md`). A store written earlier carries the key and opens with zarr-layer's `crs` and `bounds` constructor options. zarr-layer reads the `proj` and `spatial` attributes and supports arbitrary CRS through proj4 reprojection. A `star-delta` store needs an adapter that reconstructs the residuals before zarr-layer renders it; stock zarr-layer does not. chronozarr stores stay in a projected per-AOI CRS and declare it via `proj`/`spatial` attributes. * **GeoZarr** and the zarr-conventions `proj`/`spatial`/`multiscales` drafts: CRS and affine conventions chronozarr aligns with (`crs`, `transform`, `proj:code`, `spatial:*`). GeoZarr's multiscale layout is still a proposal. chronozarr adds the temporal block, the band objects, the `times`/`band_names`/`levels` mirrors, `mask`, `coverage` and the volatility array on top. * **COG + TiTiler**: single-timestep GeoTIFFs rendered by a tile server. It needs a running server and one request path per timestep; chronozarr needs neither. ### 13. Changes from 0.1 * **Version.** `spec_version` is `0.2.0`. Readers accept `0.1.x` and `0.2.x`; every v0.1 store is a valid v0.2 store (its `bands` is a list of strings, it has no `levels`, `band_names`, `mask_variable` or `selection`, and its residual bytes are valid modular residuals). A v0.1-only reader rejects `0.2.x` stores at the version check; the `multiscales` attribute, `dimension_names` and plain arrays keep plain Zarr consumers working. * **Temporal encoding is optional** (§4). `temporal.encoding` is `none` or `star-delta`. The writer default `auto` measures a sample of LOD 0 cells and keeps star-delta only at 0.85 of the plain size or better, recording `temporal.selection`. * **Residuals are modular** (§4.2). The clip and the encoder's overflow failure are gone; decoders wrap instead of clamping. Bytes are identical to v0.1 wherever the difference fits the signed range. `uint8` joins `uint16` as a star-delta dtype. * **Codecs** (§8). Readers must support `zstd`, `gzip` and `blosc` (`cname` zstd or lz4, `shuffle` or `noshuffle`). The default stays `zstd` level 5. * **dtype and nodata profiles** (§2.3). `uint8`, `uint16`, `int16` and `float32`; `nodata` is a number or `null`; optional `mask` variable (§2.4). `int16` and `float32` are `none` only. * **Band objects** (§3.6). `bands` becomes a list of `{ name, common_name?, scale?, offset?, units? }` with a `band_names` string mirror; physical value is `stored * scale + offset`. * **Coverage and provenance** (§2.5, §3.8). Optional `coverage` variable and `provenance` object. * **Layout options and length hints** (§7, §3.7). `chunk_size` (256 or 512 recommended); `shard_time` and a multi-shard time grid; `shard_bytes`; `levels`. * **Volatility** (§5). Written for every store, `none` included, with a fixed divisor of 10000 for every dtype. * **Hosting** (§9). `404` for absent keys, no `Content-Encoding` on chunk and shard objects, `Timing-Allow-Origin` as a MAY, a missing shard object decodes as fill, upload order with the root `zarr.json` last. * **GDAL CRS** (§3.3). The data array carries `_CRS` for EPSG stores so GDAL assigns the CRS. * **Prior art** (§12). Corrected statements about Mapbox raster-array and zarr-layer; COG + TiTiler added. * **Unsharded is the writer default** (§7.1, §14). `encode()` writes one object per chunk unless sharding is asked for (`shard=True`, `--shard`, with `shard_time`); `--no-shard` remains accepted. `spec_version` stays `0.2.0`: both layouts were valid and readers handle both, so every store written with the sharded default is still valid and readable. Reason, from two measurements on the published 117-month Ucayali store (`docs/comparisons.md`, `docs/append.md`). First, a cold open spent 6.5 of 6.8 s on the nine shard-index reads: each is a 1,876-byte range read at the end of a shard of 83 to 174 MB, and a Cloudflare cache miss on it pulled the whole object from R2 (2 to 11 s; 0.36 s on a 41 MB shard). Second, an append to a sharded store rewrites the trailing shard (4389 MB against 686 MB unsharded over a 12-month cycle). Sharding had been chosen for object count (93 against about 5,900), which object storage does not charge for. Unsharded costs one object per cell, level and timestep, no index read, one chunk per CDN miss, and only new objects per append; sharded costs far fewer objects and one range read per timestep once the index is cached, with a miss proportional to the shard size. * **Delta references are the recorded map** (§4.2). A reference is valid when it names an anchor at distance `0 < d < anchor_interval`; the nearest-anchor schedule is the writer's default for new timesteps, not a validity rule, so an append can keep every published reference. `spec_version` stays `0.2.0`: every store written under the earlier rule (nearest anchor for `n_time`) is valid under this one, and readers already decoded through the recorded map. * **Appending** (§14). A store may grow at the end of its time axis. * **`pixels_per_tile` is no longer written** (§2.1, §3.4). Writers MUST NOT write `multiscales[0].datasets[].pixels_per_tile`; readers MUST ignore it if present; a validator accepts stores with or without it. Reason: CarbonPlan zarr-layer 0.10.0 reads the key as the marker of a global Web Mercator slippy-map pyramid, so a native-CRS store that carries it opens as a blank map; with the key removed from the root `zarr.json` the Ucayali store opens georegistered within 1 px (`docs/comparisons.md`). The cell size was already the chunk shape of the data arrays and the grid is in `chronozarr.levels`, so no information is lost. `spec_version` stays `0.2.0`: stores written earlier carry the key and remain valid, and for zarr-layer they need its `crs` and `bounds` constructor options or a root `zarr.json` rewritten without the key. An append (§14) leaves `multiscales` as it found it, key included. * **Unchanged.** Zarr v3 groups per level, `multiscales` in the ndpyramid form, `dimension_names` everywhere, consolidated metadata, native CRS, the index at the end of a sharded object, the anchor positions, the pyramid geometry. (The writer's layout default is the exception: the next entry.) ### 14. Appending A store MAY grow at the end of its time axis (`chronozarr append`). The new timesteps are strictly after the last one and have the store's grid, bands, dtype, CRS and nodata; a store with `mask` or `coverage` needs both planes for every new timestep, and a store without them cannot gain them. **What changes.** * Every level's `data`, `mask` and `coverage` grow to the new `n_time`. `time` is rewritten as one chunk `c/0` of the new length. * The objects that gain data are written: for an unsharded store one chunk per new timestep, cell and level; for a sharded store the time shard that holds each new timestep, for every cell and level. A shard that already holds earlier timesteps is replaced by one holding the same chunks followed by the new ones and a new index. zarr-python 3.1 keeps the old chunks at their old offsets, so an index a reader cached earlier still points at the same bytes. * Root `zarr.json`: `chronozarr.times`, `temporal` (`anchor_indices`, and `delta_reference` gains the entries of the new non-anchor timesteps), `levels[].shape`, `shard_bytes` (the listed lengths of the shards written), and the consolidated metadata. Each array's `zarr.json` carries the new shape. `volatility` is updated: the new deltas join each cell's mean, so it equals the §5 definition over the recorded references up to float32 rounding. * A new timestep at level `k` is the block mean of the new timestep at level `k-1`, exactly as a fresh encode writes it. A new non-anchor timestep of a `star-delta` store references the nearest anchor that exists after the append, including anchors in the appended batch (§4.2); an anchor from an earlier append is read back from the store. **What stays byte-identical.** Every chunk of an existing timestep, at every level; every time shard that receives no new timestep (all shards before the one holding the old last timestep, and that one too when the old axis ended exactly at its boundary); `band`, `x`, `y`; level and `multiscales` attributes; every existing `delta_reference` entry and anchor; and the `mask` and `coverage` planes of existing timesteps. Nothing a published chunk decodes to changes. **Choosing the layout.** A store meant for appends SHOULD be unsharded (§7.1), which is the writer default. An unsharded append writes only new chunk objects and the metadata: 55 MB per month in the Ucayali measurement (`docs/append.md`), nothing rewritten, no shard index for an open reader to hold stale, and the same chunk reads per timestep with no index read. A sharded append rewrites the shard that receives each new timestep, whole: with `shard_time` S the k-th append into a shard writes k chunks per cell, an average of `(S + 1) / 2`, measured as 4389 MB against 686 MB unsharded over a 12-month cycle at S = 12. A sharded store written with the whole-axis `shard_time` (`n_time` at creation) opens a second shard of the same length, which grows and is rewritten on every append until it is full. The price of unsharded is object count: the 117-month Ucayali store is about 5,900 objects against 93 sharded. Choose a finite `shard_time` (12 for monthly data) for an appendable store only when object count matters more than the rewrite cost, and the whole-axis `shard_time` for archives that are not appended to. `shard_time` MAY exceed `n_time` (§7.1), so the first batch of a sharded store can be a single timestep. **Publishing.** An append is not atomic and not reversible: write it to a working copy, run validation, then publish (`docs/hosting.md`). Objects change in place only in these classes, and a host SHOULD give them a short cache lifetime: the root and every other `zarr.json`; `time/c/0` of every level; `volatility/c/0/0`; and, for each cell, level and variable, the shard holding the last timestep (or, unsharded, nothing: new chunks are new keys). Every other chunk and shard object is immutable. Upload shards and chunks first, then `time/c/0` and `volatility`, then the `zarr.json` objects below the root, then the root `zarr.json` (§9). **Readers.** A reader holding the previous root `zarr.json` still decodes every timestep it lists, because their chunks, offsets and references are unchanged. It cannot see the new timesteps until it reloads the root. One hazard: the index of a shard is read with the shard length from `shard_bytes`. After an append the trailing shard is longer, so a reader that fetches such an index with the old length reads the wrong bytes and fails the index checksum; reloading the root `zarr.json` fixes it. Once the next shard has started, the earlier shard never changes again. An unsharded store has no shard index, so the hazard does not arise. ## Examples The [live chronozarr demo](https://chronozarr.org/demo/) explores a Sentinel-2 time series of the Ucayali River, Peru. The homepage previews are generated from that store's level-3 RGB bands. They use a fixed display stretch and are illustrative images, not level-0 measurements. | Workflow | Start here | | ----------------------------------------------------- | ------------------------------------------------------------------------------------------ | | Sentinel-2 monthly composites from Planetary Computer | [Ingest example](https://github.com/chronozarr/chronozarr/tree/main/examples/sentinel2_pc) | | Water masks, validity, and coverage | [Worked example](/guides/water-masks) | | Georeferenced PNG frames | [Conversion guide](/guides/png-frames) | | leafmap notebooks | [Notebook examples](https://github.com/chronozarr/chronozarr/tree/main/examples/leafmap) | | anywidget notebooks | [Widget examples](https://github.com/chronozarr/chronozarr/tree/main/examples/anywidget) | | SWOT rasters | [Raster example](https://github.com/chronozarr/chronozarr/tree/main/examples/swot_raster) | | NISAR | [Example directory](https://github.com/chronozarr/chronozarr/tree/main/examples/nisar) | Use the [synthetic quickstart](/getting-started) for a small local example without external data access. The [measured comparisons](/guides/comparisons) record specific stores, codecs, network conditions, and benchmark limitations. ## Getting started This example builds a small synthetic time series, writes it, and checks an exact level-0 roundtrip. You do not need a bucket or satellite imagery to try it. ### Install Python 3.11 or later: ```sh uv pip install chronozarr ``` ### Create a raster time series Save this as `quickstart.py`. The coordinates describe 10-metre pixels in UTM zone 31N. The data is synthetic; it demonstrates the layout rather than a measured phenomenon. ```python from pathlib import Path import tempfile import numpy as np import xarray as xr import chronozarr values = np.arange(3 * 1 * 64 * 64, dtype=np.uint16).reshape(3, 1, 64, 64) da = xr.DataArray( values, dims=("time", "band", "y", "x"), coords={ "time": np.array(["2024-01-01", "2024-02-01", "2024-03-01"], dtype="datetime64[ns]"), "band": ["example"], "y": 5000000 - (np.arange(64) + 0.5) * 10, "x": 500000 + (np.arange(64) + 0.5) * 10, }, ) # A new output directory each run; encode does not overwrite a populated store. path = Path(tempfile.mkdtemp(prefix="chronozarr-quickstart-")) / "my_store" chronozarr.encode(da, path, crs="EPSG:32631", encoding="none", nodata=None) store = chronozarr.open_store(path) np.testing.assert_array_equal(store.read(t=1), values[1]) print(path) print(store.to_xarray(lod=0)) ``` ```sh uv run python quickstart.py chronozarr validate /path/printed/by/the/script/my_store chronozarr info /path/printed/by/the/script/my_store ``` Here `encoding="none"` makes every timestep readable by ordinary compatible Zarr v3 clients. The normal writer default, `auto`, samples the data and enables star-delta only where it saves enough compressed bytes. `nodata=None` keeps zero as a valid value. ### Read with xarray ```python # Reconstruct temporal encoding and return physical values by default. ds = xr.open_dataset(path, engine="chronozarr") # This example has encoding none, so a plain Zarr reader also sees true values. plain = xr.open_zarr(path, group="0", zarr_format=3, chunks=None) ``` `store.to_xarray()` loads the requested level into memory. The xarray backend is lazy; install `chronozarr[dask]` if you want dask chunks. See [Python and xarray](/reference/python). ### Convert an existing stack ```sh chronozarr encode scenes.nc my_store chronozarr convert manifest.csv my_store ``` A COG manifest has `uri,datetime` columns and optional `bands`. GeoTIFF input needs `chronozarr[geo]`; NetCDF input may need `chronozarr[netcdf]`. For PNG input, follow [the georeferencing guide](/guides/png-frames). Next, [publish the store](/publishing), or [open it in a notebook](/integrate). ## How it works chronozarr is a layout convention plus readers. Every object is an ordinary Zarr v3 object; the convention specifies a time axis, a multiscale pyramid, coordinates, and optional temporal encoding. ### A pyramid of arrays Each numbered group is a resolution level. Level 0 preserves the stored input values; coarser levels contain block means. Data has dimensions `(time, band, y, x)`. Coordinates and metadata describe its native projection. ```text my_store/ zarr.json consolidated metadata and format attributes 0/ original resolution data/ time × band × y × x time/ band/ y/ x/ coordinate arrays mask/ coverage/ optional validity and observed coverage 1/ half resolution 2/ quarter resolution ``` ### Random access through time The default unsharded layout uses one object per spatial cell and timestep. Sharding is an option when reducing object count is useful, with byte-range reads and index caching as the tradeoff. Growing stores generally favor unsharded output: new timesteps add objects rather than rewriting a trailing shard. With temporal encoding `none`, every chunk holds true values. With `star-delta`, anchor timesteps hold true unsigned integer values and other timesteps hold residuals against a recorded anchor. Decoding needs at most two data chunks, without reading all preceding timesteps. The writer default `auto` enables it only when sampled compressed size is at most 85% of plain storage. Signed and float data use `none`. ### Values and validity The format supports uint8, uint16, int16, and float32. Scale and offset describe physical values; a validity mask can distinguish valid zero from missing data. Coverage records how much input was observed, separately from validity. Compatible ordinary Zarr clients read the layout. They see true values when encoding is `none`, and residuals at non-anchor timesteps when encoding is `star-delta`. Use a chronozarr reader or export reconstructed COGs for the latter. See the [draft specification](/specification) for the normative requirements and [format comparison](/guides/format-comparison) for alternatives. ## Integrate Choose the interface that fits your application. | Interface | Use it for | Guide | | ------------------- | ------------------------------------------------------------------------------- | ------------------------------- | | Python and xarray | Read and analyze stores, reconstruct temporal encoding, access physical values. | [Python](/reference/python) | | JavaScript reader | Fetch decoded cells without a DOM or renderer. | [Reader](/reference/javascript) | | MapLibre layer | Add a store to an existing map. | [MapLibre](/reference/maplibre) | | chronozarr iframe | Embed a ready-made viewer with playback and inspection. | [Embedding](/guides/embedding) | | COG export and STAC | Share reconstructed values and catalog metadata with other tools. | [CLI](/reference/cli) | ### Python ```python import chronozarr store = chronozarr.open_store("my_store") values = store.read(t=0) ``` For a local notebook viewer, install `chronozarr[notebook]` and call `chronozarr.view("my_store")`. It starts a local range-capable server for local stores. ### JavaScript ```sh npm install chronozarr ``` ```js import { openStore } from 'chronozarr' const store = await openStore('https://data.tileripper.com/ucayali_santa_maria/chronozarr-4') const lod = store.levels.length - 1 const cell = await store.getCell(lod, 0, 0, 0) console.log(cell.data) store.close() ``` The returned typed array contains reconstructed stored values. Scale and offset are applied separately when you want physical units. ### Embed a viewer ```html ``` The [embedding guide](/guides/embedding) covers store URLs, controls, themes, and the postMessage contract for host applications. ## Publish a store A store is a directory of static objects. Upload it to an HTTPS host, keep its relative paths intact, then share the root URL. No application server is required. ### Choose the prefix and cache policy For an immutable release, use a new versioned prefix and long-lived caching. Upload chunks first, child metadata second, and the root `zarr.json` last. The root then marks a complete store. Avoid requesting the final URL before the upload finishes. If you plan to append, use short cache lifetimes for metadata and objects that can change. A CDN rule must respect those origin headers. The [append guide](/guides/append) and [hosting recipes](/guides/hosting) explain the object classes and upload order. ### Check the live URL ```sh chronozarr doctor https://your-host/my_store ``` Browser access needs CORS. Sharded stores also need correct byte-range responses; missing objects must return 404, rather than an HTML fallback. Doctor checks the layout, HTTP behavior, and decoding and reports failures separately from advice. ### Open the viewer Use your URL as the `store` parameter: ```text https://chronozarr.org/demo/?store=https://your-host/my_store ``` The reader and renderer run in the browser. The host serves data objects. Follow the detailed recipes for [R2, S3 with CloudFront, GCS, or Source Cooperative](/guides/hosting). ## chronozarr Browser and Node reader for [chronozarr](https://github.com/chronozarr/chronozarr) stores, plus a MapLibre GL JS custom layer that draws one. A chronozarr store is a Zarr v3 time series of rasters with a multiscale pyramid, written one object per chunk by default and optionally sharded, laid out so a client reads one timestep of one map cell with one HTTP request (a plain `GET` of one chunk; for a sharded store one range read of a shard once its index is cached). The reader turns `(lod, row, col, t)` into a typed array: it caches shard indexes (sharded stores), decodes in a worker pool, reconstructs star-delta timesteps, and prefetches a window of the time axis around the current timestep. It reads spec 0.1 and 0.2 stores ([spec](https://github.com/chronozarr/chronozarr/blob/main/spec/CHRONOZARR.md)). The Python package `chronozarr` writes them. The package is plain ES modules. There is no build step and no runtime dependency: zarrita and numcodecs are vendored (see [Licenses](#licenses)). ### Install ```bash npm install chronozarr ``` ```js import { openStore } from 'chronozarr'; import { ChronozarrLayer } from 'chronozarr/maplibre'; ``` Without a bundler, map the names in an import map. Either point them into `node_modules`: ```html ``` or at a CDN: ```html ``` | Import | Provides | | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `chronozarr` | `openStore(url, options)`, the `ChronoStore` it resolves to, and helpers (`applyDelta`, `chunkKey`, `scrubCost`, `windowOrder`, `samplePixelFrom`, `FetchError`) | | `chronozarr/maplibre` | `ChronozarrLayer`, a MapLibre GL JS custom layer | | `chronozarr/decode-worker` | the module worker that `openStore` starts for decoding; `import.meta.resolve('chronozarr/decode-worker')` finds it, for example for the `spawnWorker` option | The decode worker is found relative to `decoder.js` (`new URL('./decode-worker.js', import.meta.url)`), so it loads wherever the package files are served from: `node_modules`, a static host, or a CDN. A cross-origin worker script, which is what a CDN is, is started through a same-origin `blob:` URL that imports it; the CDN has to send CORS headers (jsDelivr and unpkg do), and a page with a Content Security Policy needs `worker-src blob:`. In Node, and with `{ workers: 0 }`, chunks decode on the calling thread. ### Read a store ```js import { openStore } from 'chronozarr'; const store = await openStore('https://data.tileripper.com/ucayali_santa_maria/chronozarr-4'); console.log(store.times.length, store.bands, store.dtype, store.crs); // 117 [ 'B02', 'B03', 'B04', 'B08' ] 'uint16' 'EPSG:32718' const lod = store.levels.length - 1; // the coarsest pyramid level const { data, chunkWidth, chunkHeight } = await store.getCell(lod, 0, 0, 5); // row 0, col 0, timestep 5 // data: typed array of the store's dtype, laid out [band][y][x] over the padded chunk. Read-only. const stored = (band, y, x) => data[band * chunkHeight * chunkWidth + y * chunkWidth + x]; const { scale, offset } = store.attrs.bands[2]; // B04 console.log(stored(2, 100, 100) * scale + offset); // reflectance store.close(); // aborts in-flight requests and releases the decode workers ``` `getCell` returns exact stored values (anchors are true values; residuals of star-delta stores are added back). `store.levels[lod]` describes each pyramid level (`gridRows`, `gridCols`, `width`, `height`, `resolution`, `transform`), `store.prefetch({ lod, cells, t })` fills the caches around a timestep, and `store.stats()` reports requests, bytes and cache hits. `openStore` options include `fetch`, `workers`, `totalBytes` (the joint cap for the decoded and compressed tiers, 1.5 GiB on machines reporting 8 GB or more, else 768 MiB), `horizonSteps`, `idleBytes` and `idleMs` (how far and how much idle prefetch reaches: 12 timesteps either side and 64 MiB per view by default), `maxRequests` and `retryDelaysMs`; they are documented in `chronozarr/decoder.js`. `prefetch` takes `playing: true` to extend to the whole loop and `masks: true` to fetch masks alongside chunks. The host must serve the store with byte ranges and CORS; `chronozarr doctor ` from the Python package checks that. ### Draw a store on a MapLibre map ```js import * as maplibregl from 'maplibre-gl'; import { ChronozarrLayer } from 'chronozarr/maplibre'; const map = new maplibregl.Map({ container: 'map', style: 'https://demotiles.maplibre.org/style.json' }); const layer = new ChronozarrLayer({ url: 'https://data.tileripper.com/ucayali_santa_maria/chronozarr-4', // store root (holds zarr.json) product: 'true_color', // layer.products lists what the store's bands support t: 0, // timestep index prefetch: true, // fill the reader's caches around t in the background }); layer.on('open', () => map.fitBounds(layer.bounds, { padding: 40, duration: 0 })); layer.on('error', (event) => console.error(event.error)); map.on('load', () => map.addLayer(layer)); slider.oninput = () => layer.setTime(Number(slider.value)); // no refetch when the timestep is cached button.onclick = () => layer.setProduct('ndvi'); map.on('click', async (event) => console.log(await layer.getValueAt(event.lngLat))); // stored values of the pixel ``` The layer needs MapLibre GL JS 5 or later (tested with 6.10.0), which is the host page's to load; it is not a dependency of this package. It draws on the Web Mercator projection and supports stores in UTM, EPSG:3857 and EPSG:4326. Options, events, limits and the level-of-detail rule are in [js/maplibre/README.md](https://github.com/chronozarr/chronozarr/blob/main/js/maplibre/README.md). ### Licenses chronozarr is Apache-2.0 (`LICENSE`). Three MIT-licensed packages by Trevor Manz are vendored unmodified apart from a header comment, with their licenses beside them; each file's header records the package version and the SHA-256 of the published file. | Path | Package | Version | License | | ------------------------- | ------------------------------------------------------- | ------- | ------- | | `vendor/zarrita/` | [zarrita](https://github.com/manzt/zarrita.js) | 0.7.5 | MIT | | `vendor/zarrita-storage/` | [@zarrita/storage](https://github.com/manzt/zarrita.js) | 0.2.0 | MIT | | `vendor/numcodecs/` | [numcodecs](https://github.com/manzt/numcodecs.js) | 0.3.2 | MIT | `vendor/numcodecs/blosc.js`, `lz4.js` and `zstd.js` embed WebAssembly builds of Blosc (with its bundled zlib and snappy), LZ4 and Zstandard. Those C libraries carry their own permissive upstream licenses (BSD-style, zlib), and numcodecs ships no separate notice for them. ## chronozarr on MapLibre `ChronozarrLayer` draws a chronozarr store on a MapLibre GL JS map as a custom layer. It opens the store with `js/chronozarr/decoder.js`, picks the pyramid level from the map zoom, uploads the raw chunks of the visible cells as integer (or float) textures, and reconstructs star-delta timesteps and runs the product band math in the fragment shader. Each cell is a small mesh whose vertices were projected from the store's CRS to Web Mercator, so the raster lands on the basemap with no resampling step. No dependencies: MapLibre is the host page's. Demo: `js/maplibre/index.html` (the live Ucayali store over MapLibre's demo tiles, with a time slider, product buttons, opacity and a click readout). Serve the repository root and open `/js/maplibre/index.html`, for example `uv run python -m http.server 8000`. The page loads MapLibre GL JS **6.10.0** (published 2026-09-15) from cdn.jsdelivr.net through an import map, which is its only external script, plus the matching stylesheet. The layer reads `defaultProjectionData.mainMatrix`, a float64 mercator-to-clip matrix that MapLibre passes from version 5 on; it was tested with 6.10.0 only, and reports an error if the matrix is missing. ### Integration ```js import * as maplibregl from 'maplibre-gl'; import { ChronozarrLayer } from './js/maplibre/layer.js'; const map = new maplibregl.Map({ container: 'map', style: 'https://demotiles.maplibre.org/style.json' }); const layer = new ChronozarrLayer({ url: 'https://data.tileripper.com/ucayali_santa_maria/chronozarr-4', // store root (holds zarr.json) product: 'true_color', // see layer.products; 'band' shows one band t: 0, // timestep index prefetch: true, // fill the reader's caches around t in the background (default false) }); layer.on('open', () => map.fitBounds(layer.bounds, { padding: 40, duration: 0 })); layer.on('error', (event) => console.error(event.error)); map.on('load', () => map.addLayer(layer)); // optional second argument: the layer to draw below slider.oninput = () => layer.setTime(Number(slider.value)); // no refetch when the timestep is cached button.onclick = () => layer.setProduct('ndvi'); map.on('click', async (e) => console.log(await layer.getValueAt(e.lngLat))); // stored values of the pixel // layer.setOpacity(0.6); layer.remove() frees the GPU objects, workers and in-flight requests. ``` ### API `new ChronozarrLayer(options)`: `url` (required), `id` (`'chronozarr'`), `product` (`'true_color'`), `band` (index or name, for the `'band'` product), `t` (0), `opacity` (1), `prefetch` (false), `meshDivisions` (8), `gpuBudgetBytes` (256 MiB), `lodBias` (0.5), `stretchLo`, `range`, `storeOptions` (passed to `openStore`: `fetch`, `workers`, `decodedBytes`, ...). | | | | -------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `setTime(t)` | Show timestep `t` (integer index into `layer.times`). Throws `RangeError` outside the axis. | | `setProduct(id, band?)` | `true_color`, `false_color`, `ndvi`, `ndwi`, `water`, `band`; throws if the store lacks the bands. | | `setOpacity(o)` | 0 to 1. | | `getValueAt(lngLat)` | Promise of `{t, time, col, row, lngLat, x, y, valid, bands: [{name, units, reflectance, stored, value}], ndvi, ndwi, isWater}` from the finest level (exact stored values), or `null` outside the footprint. `lngLat` is `{lng, lat}` or `[lng, lat]`. Fetches the level-0 chunks if they are not cached. | | `on(type, fn)` / `off` | `open` (store metadata read; `layer.times`, `layer.bounds`, `layer.products` are valid), `loading` (the view needs data), `ready` (everything the view needs is on the GPU; fires again after each pan, zoom or time change that had to wait), `error` (`event.error`; the layer keeps running). `on` returns an unsubscribe function. | | `remove()` | Removes the layer from the map and releases GPU buffers, textures and the program, aborts in-flight requests and releases the decode workers (the reader ends an idle worker pool after 30 s). The layer cannot be re-added. | | `opened`, `store`, `times`, `bounds`, `bandNames`, `products`, `t`, `product`, `opacity`, `lod`, `stats` | State. `stats` counts GPU slots, uploads, evictions, meshes, pending loads. | A time change never refetches what the reader has cached: the GPU pool keeps recently shown chunks, and the reader cache (`layer.store`) holds decoded chunks. While a new timestep loads, the previous one stays on screen. With `prefetch: true` the reader also fills its caches in the background (`store.prefetch`: nearest timesteps first, within its cache and speculative-bandwidth budgets), which is what makes scrubbing a whole time series smooth; it can move hundreds of MB (the demo moved 260 MB in 40 s for a 4-cell view of the Ucayali store). Stores written to spec v0.2 work as they are: `uint8`, `uint16`, `int16` and `float32` data, star-delta (modular residuals) or no temporal encoding, per-band `scale` and `offset`, `nodata` or a validity `mask`. The product colors are `PRODUCT_GLSL` from `js/demo/products-glsl.js`, the same shader code as the viewer, and `displayMode` decides between the tone-mapped reflectance look and a linear stretch (an 8-bit RGB store is shown as stored; other single bands get the 2nd to 98th percentile of the valid values on screen, or `range: [min, max]` in physical units). ### Limits * **One projection per store.** The store's CRS is fixed; the layer supports WGS84 UTM (EPSG:326zz and 327zz), EPSG:3857 and EPSG:4326 and throws naming the supported set for anything else. A store that crosses the antimeridian or a UTM zone boundary is out of scope. The store must declare its level-0 `transform`. * **Mercator maps only.** In globe projection the layer draws nothing and emits an `error` once. * **LOD selection rule.** Level = `clamp(floor(log2(1 / p) + lodBias), 0, levels - 1)` with `p = mercatorPerTexel0 * 512 * 2^zoom`, the size of a level-0 texel in CSS pixels at the store's centre (UTM scale and mercator stretch included). `lodBias` 0.5 picks the level whose texels are closest to one CSS pixel; 0 is the spec's reader rule (largest level whose texels are no bigger than a pixel); negative values pick finer levels (crisper on high-density screens, more data). The level comes from the map centre's zoom, so a pitched view undersamples its far half and loads more cells. Zooming in past level 0 magnifies texels; nothing is interpolated (values are exact, edges are square). * **Footprint.** Coarse levels are padded to whole texels; the layer draws only the part inside the level-0 footprint, so the outline does not move when the level changes. * **Memory budget.** GPU: `gpuBudgetBytes` (default 256 MiB) sets the texture pool, in slots of one chunk (`n_band * chunk^2 * bytes`; 2 MiB for 4 bands of uint16 at 512 px, so 128 slots); a cell on screen needs two slots (anchor and delta, one for unencoded stores), plus one per chunk of a mask. A view that needs more cells than half the slots draws the cells nearest the centre and emits an `error`. Uploads are capped at 16 MiB per frame. CPU: the reader caches (`storeOptions.totalBytes`, 1.5 GiB on machines reporting 8 GB or more and 768 MiB below, shared by the decoded and compressed tiers) are separate and are what make a warm time change a zero-request operation; `prefetch` fills them within those budgets. * **Precision.** Vertex positions are float32 offsets from each mesh's centre, the matrix is composed in float64: the error is under 0.01 px at zoom 22. Adjacent cells share edge vertices to 1e-11 of the world. * **Mesh.** 8 x 8 quads per cell (`meshDivisions`): the warp error over a 512 px cell of a 10 m UTM store is below 0.01 px at zoom 22. A store with much larger cells (kilometres per pixel) needs more divisions. * Stand-ins: while a cell loads, the layer paints the last timestep shown there, else the same area from a coarser level that is already loaded. The first view fetches the coarsest level's cells first and the level the zoom asks for after them: on a 3 MB/s link with 120 ms latency (Ucayali, 4 cells at level 0, 12 MB) the first pixels appear after 2.5 s instead of 5.5 s, and the full-detail view is ready after 7.6 s instead of 5.8 s. On a fast link neither differs. ### Checking alignment `verify/` holds the check used to validate the placement against pyproj (not MapLibre's own tiles: the demo tiles have no linework near the Ucayali store). `truth.py` writes pyproj positions of the store outline and of texel boundaries; `align.js` runs in the demo page and compares them with what the GPU drew, from screenshots taken with the layer at opacity 1 and 0. Results for the Ucayali store, Chromium on an M-series Mac, 1280 x 800, bearing 0 unless noted: | Check | Result | | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Footprint outline, zoom 9.5 (level 3), 11 (2), 12.5 (0), 15.5 (0), plus bearing 30, pitch 45, bearing -50 with pitch 40 | pixels drawn outside the pyproj outline: 0 to 8 per view, none more than 0.001 px beyond it; of about 400 unlit pixels per view sampled evenly within 8 px inside the outline, every one is a nodata gap at level 0 | | Interior texel boundaries, levels 0 to 3 at about 4 px per texel | 60 to 84 boundaries per level (wherever neighbouring texels differ), all within 0.496 px of pyproj (0.5 px is the limit of pixel-centre sampling), mean offset below 0.03 px, no stray boundaries | | Colors vs `js/demo/viewer.js`, same store, timestep and texel | 20 of 20 texel and product pairs identical in 8 bits (NDVI, NDWI, water, true and false color) | Steps: `uv run --with pyproj python3 js/maplibre/verify/truth.py --epsg 32718 --x0 485650 --y0 9169880 --res 10 --width 2759 --height 2765 > truth.json`; open the demo page, hide its panel, add `verify/align.js` as a script, `jumpTo` a view, take a screenshot with `layer.setOpacity(1)` and another with `setOpacity(0)`, and pass them with `truth.outline` to `__align.mask`. ## Command line Install `chronozarr`, then run `chronozarr COMMAND --help` for all options. | Command | Purpose | | -------------------------- | -------------------------------------------------------------------- | | `encode INPUT OUT` | Write a store from a Zarr store, NetCDF file, or dated GeoTIFF glob. | | `convert SOURCE OUT` | Stream a manifest, Zarr variable, or NetCDF file into a store. | | `append STORE INPUT` | Add later timesteps to an existing store. | | `validate STORE` | Check layout conformance; failures exit with status 1. | | `info STORE` | Show times, bands, encoding, and levels. | | `doctor TARGET` | Check a local path or live URL, HTTP behavior, and decoding. | | `export-cog STORE OUT_DIR` | Export reconstructed true-value COGs. Needs the geo extra. | | `stac STORE --out DIR` | Write a static STAC Collection and Item. Needs the geo extra. | ### Common workflow ```sh chronozarr convert manifest.csv my_store chronozarr validate my_store chronozarr info my_store # After uploading: chronozarr doctor https://your-host/my_store ``` ### Encoding and storage `--encoding auto|none|star-delta` controls temporal encoding. The default is `auto`. `--shard` opts into sharding; `--shard-time` requires it. The default is unsharded. Use [the append guide](/guides/append) to choose a layout for growing stores. ### Interoperability ```sh chronozarr export-cog my_store cogs/ chronozarr stac my_store --out catalog/ ``` COG export reconstructs star-delta values. Static STAC metadata points clients to the store; it does not serve rendered tiles. ## Python and xarray ### Encode ```python import chronozarr chronozarr.encode(da, "my_store", crs="EPSG:32631") ``` `da` has dimensions `(time, band, y, x)`, supported stored dtype, and projected pixel-centre coordinates. The output must be new or empty. Encoding defaults to `auto`; sharding defaults off. Use `encoding="none"` for direct ordinary Zarr reads. ### Read stored values ```python store = chronozarr.open_store("my_store") # local path or HTTP URL frame = store.read(t=0, lod=0) # band × y × x cell = store.read_cell(t=0, row=0, col=0, lod=0) mask = store.read_mask(t=0, lod=0) # None when absent coverage = store.read_coverage(t=0, lod=0) da = store.to_xarray(lod=0, times=[0], physical=False) ``` Reads reconstruct temporal encoding. `to_xarray()` returns an in-memory DataArray; `physical=True` applies band scale and offset and represents invalid pixels as NaN. ### Lazy xarray backend ```python import xarray as xr ds = xr.open_dataset("my_store", engine="chronozarr") raw = xr.open_dataset("my_store", engine="chronozarr", physical=False) ``` The backend returns physical values by default. Install `chronozarr[dask]` to use `chunks=`. An ordinary `xr.open_zarr` reads stored arrays directly and does not reconstruct residual timesteps. ### View ```python chronozarr.view("my_store") ``` Install `chronozarr[notebook]` for the notebook viewer. The local server supports byte ranges, including sharded stores. Leafmap and geemap integrations have separate extras; their [examples](/examples) provide the complete notebook setup. See [the Python source](https://github.com/chronozarr/chronozarr/tree/main/src/chronozarr) for the full signatures and optional arguments. ## Appending: what one new date costs `chronozarr append STORE INPUT` adds timesteps to the end of a store in place (spec section 14, publishing in [hosting.md](/guides/hosting) section 7). Before it existed the only way to add a date was a re-encode under a new prefix: rewrite and re-upload every object. This page measures the append against that and against the layout choices that decide its cost. ### Which layout to append to The encoder default, unsharded, is the right layout for appends. Measured on the Ucayali mosaics: an unsharded append writes 55 MB per month and rewrites nothing but metadata. A store sharded at `shard_time` 12 writes 4389 MB over a 12-month cycle against 686 MB unsharded (6.4x). The whole-axis shard layout (`--shard`, `shard_time` = the timesteps at creation) rewrites a growing second shard every month (116.9 MB for the second append to a 115-month store, and up to 115 chunks per cell by the end). An unsharded store has no shard index for an open viewer to hold stale, and costs the same chunk reads per timestep with no index read. The price of unsharded is object count: about 5,900 objects for the 117 Ucayali months against 93 for the whole-axis shard, which is one object per cell and level for the whole time axis. Shard an appendable store (`--shard --shard-time 12` for monthly data) only when object count matters more than the rewrite cost, and use the whole-axis `--shard` for archives that are not appended to. The first batch of a sharded store may be a single timestep, because `shard_time` may exceed the timesteps present. (The tables below were measured when whole-axis sharding was still the encoder default; they are kept because they are the cost of the sharded layouts.) ### What an append writes * the objects that gain data: for each level and cell, the chunk of an unsharded store, or the time shard holding the new timestep (a shard that already has earlier timesteps is rewritten whole, its old chunks at their old offsets); * per level, the array `zarr.json` files, `time/zarr.json` and `time/c/0`, then `volatility/c/0/0` when it changes, then the root `zarr.json` with the consolidated metadata. Nothing else is opened. Star-delta references already recorded are never changed (section 4.2): an appended timestep references the nearest anchor that exists after the append. ### Method Data: `data/mosaics/ucayali_santa_maria/*.npz`, 117 monthly Sentinel-2 mosaics, 4 bands uint16, 2765 x 2759 px, chunk 512, four levels, 50 cells (36 at level 0). Stores are built from the first months, then months 13 and 14 are appended one at a time. Every operation runs in its own process (`scripts/bench_append.py`), so peak RSS is per operation. After every step `chronozarr validate` passes and every timestep decodes equal to its source mosaic, bit for bit (level 0, all timesteps). Objects written come from file size, mtime and sha256 before and after. The machine is a 16-core Mac with 128 GB shared with other jobs (load average about 10), so wall seconds are the median of three runs of the whole matrix and carry roughly 20 % noise. Object counts and bytes come from the last two runs, which were identical; the first run also rewrote the unchanged `volatility` chunk at month 13, which `append` now skips. Three layouts: whole-axis shard (`shard=True`, `shard_time` = the timesteps at creation; the encoder default when these were measured), `shard_time=12`, and unsharded (the encoder default now). At 12 months the first two are the same store, and the table shows identical numbers for them. Two encodings: `auto`, which chose `none` at 12 months (star-delta/plain 0.89 on the sample), and a forced `star-delta` with anchor interval 6. ### Results: 12-month stores, month 13 and 14 appended Store size after the 12 months: 653 MB `none`, 584 MB star-delta; 93 objects sharded, 643 / 614 unsharded. Building it takes 3.9 s with 1.3 GB peak RSS. "Objects written" counts files that are new or whose size or mtime changed (new + rewritten). `aws s3 sync` uploads exactly that set, because a rewritten file is newer than its remote copy; `sync --size-only` uploads fewer (last column). | layout | encoding | append | wall s | peak RSS MB | objects written (new + rewritten) | MB written = MB `sync` uploads | `sync --size-only`: objects / MB | | -------------------------- | ---------- | -------- | ------ | ----------- | --------------------------------- | ------------------------------ | -------------------------------- | | whole-axis = shard_time 12 | none | month 13 | 0.60 | 329 | 63 (50 + 13) | 55.3 | 55 / 55.3 | | | | month 14 | 0.70 | 381 | 64 (0 + 64) | 110.7 | 55 / 110.7 | | unsharded | none | month 13 | 0.62 | 325 | 63 (50 + 13) | 55.3 | 55 / 55.3 | | | | month 14 | 0.68 | 368 | 64 (50 + 14) | 55.4 | 55 / 55.4 | | whole-axis = shard_time 12 | star-delta | month 13 | 0.60 | 330 | 63 (50 + 13) | 55.3 | 55 / 55.3 | | | | month 14 | 0.67 | 381 | 64 (0 + 64) | 55.7 | 34 / 29.1 | | unsharded | star-delta | month 13 | 0.64 | 340 | 63 (50 + 13) | 55.3 | 55 / 55.3 | | | | month 14 | 0.60 | 338 | 43 (29 + 14) | 0.4 | 34 / 0.4 | Reading the table: * **Month 13 opens a new time shard** (sharded) or adds new chunk keys (unsharded): 50 new objects, one per cell and level, plus 13 metadata files. Peak RSS is a quarter of the build's and wall time a sixth. * **Month 14 rewrites the shard that month 13 created**, so a sharded store writes 2 chunks of data per cell (110.7 MB) while an unsharded store writes 1 (55.4 MB). The extra is old data written again. * **Star-delta, month 14, sharded: 64 objects and 55.7 MB, of which 21 shards are byte-identical to before.** Month 14 (2017-05) has 0.2 % clear coverage and its pixels are carried forward from month 13, so in 18 of the 36 level-0 cells it equals month 13, the delta chunk is all zeros and is not stored, but the shard is still rewritten, with the same bytes and a new mtime. The unsharded store writes nothing for those cells (29 new objects, 0.4 MB). A sync that compares size and mtime still uploads the 21 identical shards; one that compares checksums skips them. * **`sync --size-only` is not safe for appends.** It skips 8 to 9 changed objects whose size stays the same: each level's `data/zarr.json` and `time/zarr.json` (the shape goes from 12 to 13 in the same number of bytes) and, at times, `volatility/c/0/0`. Clients that read the per-array `zarr.json` instead of the consolidated metadata would see the old shape. Use size-and-mtime or checksum comparison, or `scripts/upload_stores.sh --newer-than`. #### The whole cycle of a shard Appending months 13 to 24 to the 12-month store, `none` encoding (star-delta in brackets): | layout | MB written, months 13-24 | objects per append | month 25 | | ------------- | ------------------------ | ------------------ | ------------------------------------------ | | shard_time 12 | 4389 (3978) | 63 to 64 | 60.8 MB, 50 new shards: the cycle restarts | | unsharded | 686 (672) | 63 to 64 | 60.8 MB | A sharded store writes the trailing shard again on each append, so month `k` of a shard writes `k` chunks per cell: the average is `(shard_time + 1) / 2` chunks, measured 6.4x (5.9x star-delta) the unsharded bytes at `shard_time` 12. The cost per append is bounded by the shard size, and `shard_time` trades object count against it: 12 gives up to 12 chunks per rewrite and 1 shard object per cell and level per year; 4 would give up to 4 and 3 per year; unsharded gives 1 and 12. ### Results: the production-sized store The published sharded Ucayali store (`chronozarr-3`) holds all its months in one shard per cell. Built the same way from the first 115 months (`--base-months 115`, `auto` chooses `none`), then months 116 and 117 appended: | operation | wall s | peak RSS MB | objects written | MB written | | ------------------------------------------------- | ------ | ----------- | -------------------------------- | ---------- | | build 115 months (today's only way to add a date) | 32.1 | 4724 | 93 (all, 6335 MB store) | 6335 | | append month 116 | 0.55 | 407 | 64 (50 new shards + 14 metadata) | 58.4 | | append month 117 | 0.57 | 392 | 64 (all rewritten) | 116.9 | Adding a month by re-encode writes and uploads 6.3 GB (93 whole-axis shards). Appending writes 58 MB (0.9 %). The first shard stays untouched, because its length was fixed at 115 timesteps when the store was created and the new timesteps start a second shard. The whole-axis layout therefore appends correctly, but the second shard is allowed to grow to 115 chunks per cell (the sharded 117-month store has shards of 161 MB at level 0) and is rewritten whole on every append until it fills. That is the cost of appending to a whole-axis shard: write an appendable store unsharded (the default), or with a finite `shard_time` when object count matters. A fresh encode of 13 months (`shard_time` 12) takes 4.8 s, 557 MB, 143 objects with 1.3 GB peak RSS; the append of month 13 takes 0.6 s and writes 55 MB. For reference, a fresh unsharded encode of all 117 months (`chronozarr-4`, `encoding` `auto` chose `none`) takes 41.9 s with 3.8 GB peak RSS and writes 6,451.8 MB in 5,893 objects; the sharded `chronozarr-3` is 6,451.9 MB in 93. ### Results: requests for a view A 3 x 3 block of level-0 cells, one timestep, read cold with the Python reader (it caches no shard index; the browser reader caches one index per shard, so after the first timestep it issues only the chunk reads). Each cell needs the shard index (one suffix range) and one range per chunk: a delta timestep reads its anchor and its delta. | store | before (t = 5 and 11) | after two appends (t = 5, 11 and the new t = 13) | | --------------------- | ----------------------- | -------------------------------------------------------------------------------------------------------- | | sharded, none | 9 index + 9 chunk = 18 | 18 for all three | | sharded, star-delta | 9 index + 18 chunk = 27 | 27 for t = 5 and 11; 9 index + 13 chunk for t = 13 (cells whose delta is all fill have no chunk to read) | | unsharded, none | 9 chunk | 9 for all three | | unsharded, star-delta | 18 chunk | 18 for all three | An append changes none of these: a shard index is `16 * shard_time + 4` bytes (196 at `shard_time` 12) whatever it holds. A fresh encode of the same 14 months with the nearest-anchor rule makes timesteps 10 and 11 reference anchor 12 in the next shard: their view costs 18 index reads instead of 9 (measured: t = 10: 18 index + 16 chunk; t = 11: 18 index + 12 chunk). Frozen references keep each delta in the same shard as its anchor. ### Frozen references: what they cost and what they avoid The nearest-anchor rule makes the next anchor, once it exists, the reference of the last timesteps of the interval before it. With interval 6, appending timestep 12 would change the references of timesteps 10 and 11 (6 to 12), and in general `anchor_interval - 1 - floor(anchor_interval / 2)` timesteps per new anchor. Those chunks are already published; changing them changes what cached copies decode to. The spec therefore freezes recorded references (section 4.2). **What re-referencing would have rewritten** at the month 13 append of the star-delta stores above: every time shard of every cell and level in a sharded store (50 objects, 584 MB, the whole store, because timesteps 10 and 11 sit in the only shard), or 100 chunk objects (110 MB) in an unsharded store. **What freezing costs in size.** A timestep whose reference stays on the preceding anchor sits 4 or 5 months from it instead of 2 or 1. Compressed bytes of level-0 chunks on the 117 Ucayali months (36 cells, zstd level 5): | anchor interval | timesteps affected | their size, nearest anchor | their size, preceding anchor | increase on them | increase of the level-0 star-delta store | | --------------- | ------------------ | -------------------------- | ---------------------------- | ---------------- | ---------------------------------------- | | 6 | 38 | 1496 MB | 1666 MB | +11.4 % | +3.6 % (of 4687 MB) | | 12 | 45 | 1903 MB | 2009 MB | +5.6 % | +2.2 % (of 4883 MB) | On this AOI star-delta barely beats plain storage (4687 MB against 4804 MB for the same level-0 chunks at interval 6; 4883 MB at interval 12, larger), and the affected timesteps stored against the preceding anchor (1666 MB) are larger than their plain chunks (1558 MB). That is why `auto` picks `none` here. Measured on an actual append, the first 12 months stored star-delta (584 MB) and then month 13 appended give 639.5 MB, against 556.8 MB for a fresh encode of the same 13 months: +82.7 MB (+14.9 %), all of it timesteps 10 and 11 (all levels). These early months are sparse (months 12 and 13 equal each other in 15 of 36 cells), so consecutive months are nearly equal and the distance to the anchor matters more than on dense data. For data that changes a lot between dates, use `encoding none`; for data that changes little, the nearest-anchor default still applies to every fresh encode. ### Limits found * An unsharded store is one object per timestep, cell and level: 5,893 for the 117 Ucayali months (17,088 for the water store, which has `mask` and `coverage` too). The first upload is slow with wrangler (hosting.md section 3.2), but each append uploads only the 63 objects it wrote. * Append is not transactional. Inputs are checked and spilled before the store is touched; a failure after that leaves it partly modified, and `append` refuses a store that no longer validates. Work on a copy. * `volatility` is updated incrementally. It equals the spec definition over the recorded references up to float32 rounding, except for a cell already clipped at 1.0. For a `none` store the nominal schedule is not recorded, so appended timesteps use interval 6. * In a sharded store, a viewer open during an append keeps the old metadata until reload and can fail on the index of a trailing shard it had not read (hosting.md section 7). An unsharded store has no shard index. ### Reproduce ```bash uv run python scripts/bench_append.py run --work-dir /tmp/append-bench uv run python scripts/bench_append.py run --work-dir /tmp/append-bench-115 --base-months 115 \ --layouts whole-axis --encodings auto --view-steps 5 100 uv run python scripts/bench_append.py run --work-dir /tmp/append-cycle --appends 13 \ --layouts shard-time-12 unsharded --view-steps 5 uv run python scripts/bench_append.py reference-cost ``` The tables above came from the first command (three runs), the second, the third, and the fourth. Each `--work-dir` must be new. ## Measured comparisons: zarr-layer, and one COG per date Two comparisons of a chronozarr store, measured on 2026-10-01 on one machine (Apple M3 Max, macOS, Node 24.16.0, Chromium 153 through Playwright 1.63.0 with the Metal GPU backend). All code, raw results and a lockfile are in `bench/`; the exact commands are in the last section. * **A. CarbonPlan zarr-layer** (`@carbonplan/zarr-layer` 0.10.0, `maplibre-gl` 6.11.2): does it open the published Ucayali store, and how does it compare with the chronozarr viewer on the same store, view and level. * **B. One COG per date**: the same imagery as 117 Cloud Optimized GeoTIFFs, against the chronozarr store, for the delivery cost of one session. Store used in both: `https://data.tileripper.com/ucayali_santa_maria/chronozarr-4` (spec 0.2.0, `temporal.encoding: none`, 117 monthly timesteps, 4 bands B02/B03/B04/B08 as uint16, 4 levels in UTM 18S, level 0 = 2759 x 2765 px at 10 m, shards of (117, 4, 512, 512) holding one zstd-5 chunk of (1, 4, 512, 512) per timestep, consolidated metadata, `shard_bytes` hints). Level 1 is 1380 x 1383 px at 20 m, a 3 x 3 grid of cells. A local copy is `data/stores/ucayali_santa_maria/chronozarr-3` (6.45 GB). ### Summary * **zarr-layer does not open the published store with its default options.** The map stays blank and zarr-layer says nothing: the `pixels_per_tile` key in `multiscales[0].datasets[*]` makes it treat the store as a global slippy-map pyramid, so it places a UTM raster on the whole Web Mercator world and finds no cell in view (section 1.1). It opens the store, drawn in the right place with the same values as the store, either with two constructor options (`crs`, `bounds`) or with `pixels_per_tile` removed from the store's attributes (section 1.2). The second was adopted as the format rule afterwards (note in section 1.2). * **Per step the two renderers move the same chunks** (9 chunks, 10.7 MB for the 3 x 3 view). chronozarr's speculative prefetch adds 10 to 64 % bytes to a 20-step scrub (14 to 17 % on a saturated local link). On a saturated link with a stable origin (local 50 and 10 Mbit/s, remote 10 Mbit/s) a chronozarr step is 15 to 19 % slower than zarr-layer's (10.2 s against 8.6 s at 10 Mbit/s); on an idle local link the same prefetch makes steps cost 0 ms against zarr-layer's 80 ms. chronozarr never shows a frame that mixes timesteps; zarr-layer does for 22 to 70 % of the time of a stepped scrub. * **Cold open on the published store is dominated by the CDN, not by either tool**: the nine shard-index reads at the end of 83 to 174 MB shard objects are cache misses that take 1.7 to 11 s, for both tools alike (section 1.5). * **One COG per date costs about the same bytes (+0 to +4 % per phase, +1.4 % over the 20-step scrub, +3.2 % on disk) and differs in requests.** Per new date a COG needs a header read before its tiles; chronozarr reads a shard index once and then only chunks. With geotiff.js 3.0.5's defaults a header is 7 chained requests; with 64 KB blocks it is 1. The modelled delivery time of the 20-step scrub is 1.75x (defaults) or 1.13x (64 KB blocks) of chronozarr on a 90 Mbit/s, 112 ms link, 1.18x / 1.05x at 50 Mbit/s + 40 ms and 1.10x / 1.04x at 10 Mbit/s + 100 ms. On slow links both are bandwidth-bound and the formats converge. * **Reading one pixel's complete history costs 164 MB in both layouts** (117 chunks or tiles of 512 x 512 x 4), because neither chunks along time; the two readers return identical values. * Local wall-clock is 3 to 5x longer for the COGs, almost all of it JavaScript decoding of DEFLATE in geotiff.js (257 ms for one 9-tile view against 37 ms for the chronozarr reader, 116 ms for a ZSTD COG), not the format. What these numbers do not show is in sections 1.7 and 2.4. ### 1. zarr-layer #### 1.1 Does it open the store? No, not as published `bench/zarr-layer/page.html` + `page.js` (bundled by `build.mjs`) put a MapLibre map at a fixed view (centre of the AOI, zoom 12, 1500 x 1500 px, black background, no interaction) and add a `ZarrLayer` with `source` = the store URL, `variable: 'data'`, selector `{ band: ['B04','B03','B02'], time: { selected: 60, type: 'index' } }` and a natural-colour fragment shader. `bench/zarr-layer/probe.mjs` runs it headless and reports what zarr-layer made of the store (`describe()`), whether every visible region loaded, and where the drawn pixels are against where the store's `spatial:bbox` puts the AOI (columns 21 to 1479, rows 15 to 1485 at that zoom). | configuration | result | what zarr-layer did | | ------------------------------------------------------------------------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | unmodified store, default options | **FAIL**, silently | level 3 chosen, 0 visible regions, 0 chunk requests (6 requests, 0.1 MB: `zarr.json`, three 404 probes for Zarr v2 files, two coordinate arrays). `describe()`: crs `EPSG:3857` with proj4 `EPSG:32718`, extent = the whole Web Mercator world (+-20,037,508 m), `latIsAscending: false`. No warning in the console. Blank map. | | unmodified store, options `crs: 'EPSG:32718'`, `bounds: [485650, 9142230, 513240, 9169880]` | PASS | level 1, 9 of 9 regions, drawn box (21,15)-(1478,1484), 33 requests, 10.3 MB | | store root attributes rewritten in the browser: `pixels_per_tile` removed | PASS | level 1, 9 of 9 regions, same box and the same 2,142,525 lit pixels as the options variant, 32 requests, 10.3 MB | Raw output: `bench/results/zarr-layer-probe-as-published.json`, `...-crs-and-bounds-options.json`, `...-without-pixels-per-tile.json`. **Cause.** In `@carbonplan/zarr-layer` 0.10.0 (`dist/index.js`), `_getPyramidMetadata` (line 1819 to 1822) reads `multiscales[0].datasets[0].pixels_per_tile` and sets `_usesSlippyMapDefaults = Boolean(pixelsPerTile)`, and takes the CRS from `datasets[0].crs` only to tell `EPSG:4326` from "everything else is `EPSG:3857`". `_loadSpatialMetadata` (line 1634) then returns early for a slippy-map store: extent is the whole world, rows run north to south, and the store's own `proj:code`, `spatial:transform` and `spatial:bbox` are never consulted. The `proj:code` is still applied as `proj4`, so the raster is stretched over a ±20,037,508 m square read as UTM metres, which projects nowhere near the view, and no region intersects the view. The same store with the key removed goes through the self-describing path (`proj:code` `EPSG:32718` is one of the UTM codes proj4 ships, `spatial:transform`/`spatial:bbox` give the extent), and the constructor options `crs` and `bounds` bypass the slippy-map branch the same way. Level order is not a problem: zarr-layer chooses the level by size, not by position in `datasets`, so the finest-first order of chronozarr works. #### 1.2 Change to the store attributes **Status (2026-10-01).** Adopted, in a follow-up to the measurement below (which tested it without touching the published store). `spec/CHRONOZARR.md` sections 2.1, 3.4 and 13 now say writers MUST NOT write `pixels_per_tile` and readers MUST ignore it; `spec_version` stays 0.2.0. The writer (`encode`, `append` leaves an existing `multiscales` alone), the Python reader (the cell size is the chunk shape of the data array) and the validator (accepts a store with or without the key) follow it, and `chronozarr doctor` prints one info line for a remote store that still carries the key. Stores written earlier stay valid and open in zarr-layer with the `crs` and `bounds` options. The root `zarr.json` of `ucayali_santa_maria/chronozarr-3` and `ucayali_santa_maria/water-1` was rewritten locally without the key (the only difference from the published roots is the four removed keys per store); the published copies carry it until those two files are uploaded. The JS reader (`js/chronozarr/decoder.js` line 408) still rejects a `pixels_per_tile` that disagrees with the chunk shape, which is stricter than "readers MUST ignore it"; not changed here. The text below is the proposal as measured. Smallest change that makes zarr-layer 0.10.0 open the store without constructor options: in the root `zarr.json`, `attributes.multiscales[0].datasets[i]` for i = 0..3, delete `pixels_per_tile`: ```diff "multiscales": [{ "datasets": [ - { "path": "0", "pixels_per_tile": 512, "crs": "EPSG:32718" }, + { "path": "0", "crs": "EPSG:32718" }, - { "path": "1", "pixels_per_tile": 512, "crs": "EPSG:32718" }, + { "path": "1", "crs": "EPSG:32718" }, ... ``` It occurs only there (the level groups and arrays carry no copy; the consolidated metadata does not repeat root attributes). Tested without touching the published store, by rewriting the root `zarr.json` in the browser (`--patch no-pixels-per-tile` of `probe.mjs`): the row above. What the change touches in this repository, as found when it was measured (all of it except the JS reader has since been edited, see the status note): * `spec/CHRONOZARR.md` section 2.1 (line 69: "Readers MUST take `cs` from `multiscales[0].datasets[].pixels_per_tile`") and section 3.4 (line 167: "`pixels_per_tile` equals `cs`"). The chunk size is also the inner `chunk_shape` of the `data` array's sharding codec, which every reader already parses. * Python: `chronozarr validate` rejects an edited store (`multiscales[0].datasets[0]: missing required key 'pixels_per_tile'`, `src/chronozarr/schema.py` around lines 697 to 705; `schema.py` 1030 and 1138 and `decode.py` 77 use the value). * The JS reader opens the edited store and reads a cell (`js/chronozarr/decoder.js:379` already treats the key as optional and only cross-checks it against the chunk shape). * Claims that a `none` store opens in zarr-layer unchanged: `spec/CHRONOZARR.md` line 170 and section 12 (line 554); `README.md` line 11 and `docs/format-comparison.md` line 12 say the same with "not yet tested". They hold only with the constructor options or after this change. Alternatives: the constructor options above (no store change; `zarrVersion: 3` additionally skips three sequential 404 probes for Zarr v2 files); an upstream change in zarr-layer so that `pixels_per_tile` only implies slippy-map defaults when the dataset CRS is `EPSG:4326` or `EPSG:3857` (not tested); the zarr-conventions `multiscales` layout form (not tested; the spec keeps the ndpyramid list form on purpose). #### 1.3 Values With the options variant, `layer.queryData` at level `finest` for pixel (row 1388, col 1380) at 11 timesteps (0, 2, 3, 40, 41, 60, 80, 100, 101, 102, 116) returns the same 44 band values as a read straight from the Zarr arrays with no chronozarr code (`bench/results/zarr-layer-values.json`); the four values of timestep 0 are nodata (0 in the store), which zarr-layer leaves out as NaN. #### 1.4 Method of the comparison * **View and level.** The whole AOI (3 x 3 cells of level 1) in a 1800 x 1700 px window at device pixel ratio 1, AOI 1457.6 px wide, centred. zarr-layer chooses level 1 by itself at map zoom 12 (every run reports level 1). chronozarr is pinned to level 1 (`viewer.loadStore(url, { lod: 1, viewSearch })`, `viewSearch` = `?t=40&z=<0.5283>&c=`), which also switches off its coarse-first staging; its canvas (1500 x 1592) shows the whole AOI. `chronozarr.interactionBench` was not used because it opens at the viewer's own fit camera with the adaptive level and takes neither; the page script replays its scrub with the same probe events and the same `analyzeLatency` from `js/demo/perf.js`. * **zarr-layer configuration.** Unmodified store; options `crs: 'EPSG:32718'`, `bounds`, `zarrVersion: 3`; one `setSelector` per step with `time` by index. A frame counts as complete when every visible region of the active level holds the current selector and has its textures uploaded (read from `layer.regionRenderer`'s region cache at each MapLibre render). * **Open.** From the call that opens the store (`loadStore` / `map.addLayer`, both include reading the metadata) to the first frame that shows all 9 cells at level 1. * **Scrub forward 20**, from timestep 40 (steps 41 to 60), in two modes. *Burst*: one input every 100 ms (chronozarr's `ArrowRight`, zarr-layer's `setSelector`), as in `interactionBench`; a step the tool skipped is satisfied by the later frame that shows a later step. *Paced*: the next input 100 ms after the frame of the previous step is complete, so every step is shown and its latency is the time to load and draw it. * **Network.** A fresh Chromium per run (empty caches; HTTP cache also disabled over CDP). CDP `Network.emulateNetworkConditions` is applied after the page has loaded, so only the data requests are throttled (checked once with 10 MB range reads: 1.66 s at 50 Mbit/s and 8.1 s at 10 Mbit/s). Requests and bytes are counted from CDP events of the page, identically for both tools: `encodedDataLength` (headers and body) of finished requests, and the body bytes received so far for aborted ones, which is a lower bound of what crossed the wire. The count was checked against chronozarr's own reader statistics (19 requests, 10.03 MB against 19 and 10.04 MB at open; the reader counts only completed requests). * **Two sources.** *remote*: the published store; `natural` is then the real link from this machine to Cloudflare (HTTP/2). *local*: the same files from this repository's range server (`js/support/static-server.js`, HTTP/1.1, so at most 6 connections per origin for either tool); `natural` is then unthrottled localhost, and CDP throttling defines the whole link. The local runs exist because the CDN adds stalls that are not the tools' (section 1.5). * **Repetitions.** remote: natural 3, 50 Mbit/s 2, 10 Mbit/s 2 (natural and 50 Mbit/s after one unrecorded warm-up per tool); local: 3, 2, 2. The two tools and two modes are interleaved, the tool order alternates by repetition. Tables show the median and the range. The machine was shared (load average 10 to 20) and the remote link varied by an order of magnitude between probes (27 to 964 Mbit/s). #### 1.5 Results: the published store ##### Cold open, level 1, 9 cells: published store (data.tileripper.com) Median over repetitions, range in parentheses. Burst and paced runs both open the store the same way and are pooled. | link | tool | runs | time to first complete frame | of which until the shard-index reads are done | requests | MB | | ------------ | ---------- | ---- | ---------------------------- | --------------------------------------------- | ---------- | ---------------- | | natural | chronozarr | 6 | 6.79 s (3.08 s-9.92 s) | 6.58 s (3.01 s-9.61 s) | 19 (19-19) | 10.0 (10.0-10.0) | | natural | zarr-layer | 6 | 6.72 s (4.84 s-7.91 s) | 6.51 s (4.46 s-7.54 s) | 30 (30-30) | 10.0 (10.0-10.0) | | 50Mbit-40ms | chronozarr | 4 | 7.69 s (3.76 s-10.1 s) | 7.42 s (3.38 s-9.85 s) | 19 (19-19) | 10.0 (10.0-10.0) | | 50Mbit-40ms | zarr-layer | 4 | 5.76 s (3.04 s-6.80 s) | 5.49 s (2.22 s-6.52 s) | 30 (30-30) | 10.0 (10.0-10.0) | | 10Mbit-100ms | chronozarr | 4 | 14.8 s (14.7 s-16.8 s) | 9.34 s (2.97 s-15.7 s) | 19 (19-19) | 10.0 (10.0-10.0) | | 10Mbit-100ms | zarr-layer | 4 | 16.9 s (9.36 s-21.3 s) | 15.8 s (3.88 s-20.1 s) | 30 (30-30) | 10.0 (10.0-10.0) | ##### Scrub forward 20 timesteps, paced: published store (data.tileripper.com) | link | tool | runs | requests | of which aborted | MB transferred | scrub duration | steps shown exactly | step latency median / p95 | time showing mixed-timestep frames | MB in the next 3 s | | ------------ | ---------- | ---- | ------------- | ---------------- | -------------- | ------------------------- | ------------------- | ------------------------- | ---------------------------------- | ------------------ | | natural | chronozarr | 3 | 227 (219-235) | 36 (28-50) | 235 (234-235) | 14.6 s (10.5 s-16.3 s) | 20 (20-20) of 20 | 424 ms / 1.45 s | 0 ms (0 ms-0 ms) | 49 (29-59) | | natural | zarr-layer | 3 | 180 (180-180) | 0 (0-0) | 214 (214-214) | 17.2 s (16.6 s-19.6 s) | 20 (20-20) of 20 | 451 ms / 1.77 s | 9.71 s (9.42 s-12.7 s) | 0 (0-0) | | 50Mbit-40ms | chronozarr | 2 | 257 (255-257) | 77 (75-77) | 297 (263-297) | 50.2 s (44.0 s-50.2 s) | 20 (20-20) of 20 | 2.12 s / 4.62 s | 0 ms (0 ms-0 ms) | 15 (15-15) | | 50Mbit-40ms | zarr-layer | 2 | 180 (180-180) | 0 (0-0) | 214 (214-214) | 65.1 s (38.9 s-65.1 s) | 20 (20-20) of 20 | 2.27 s / 7.70 s | 45.7 s (13.6 s-45.7 s) | 0 (0-0) | | 10Mbit-100ms | chronozarr | 2 | 275 (273-275) | 94 (93-94) | 270 (269-270) | 207.9 s (207.8 s-207.9 s) | 20 (20-20) of 20 | 10.3 s / 12.2 s | 0 ms (0 ms-0 ms) | 3 (2-3) | | 10Mbit-100ms | zarr-layer | 2 | 180 (180-180) | 0 (0-0) | 214 (214-214) | 175.8 s (175.7 s-175.8 s) | 20 (20-20) of 20 | 8.64 s / 9.73 s | 59.2 s (56.7 s-59.2 s) | 0 (0-0) | ##### Scrub forward 20 timesteps, burst (one step every 100 ms): published store (data.tileripper.com) | link | tool | runs | requests | of which aborted | MB transferred | scrub duration | steps shown exactly | step latency median / p95 | time showing mixed-timestep frames | MB in the next 3 s | | ------------ | ---------- | ---- | ------------- | ---------------- | -------------- | ---------------------- | ------------------- | ------------------------- | ---------------------------------- | ------------------ | | natural | chronozarr | 3 | 192 (192-192) | 177 (175-177) | 11 (11-21) | 2.91 s (2.33 s-3.11 s) | 1 (1-1) of 20 | 2.01 s / 2.91 s | 0 ms (0 ms-0 ms) | 29 (16-52) | | natural | zarr-layer | 3 | 180 (180-180) | 169 (169-171) | 28 (11-30) | 3.23 s (2.33 s-3.50 s) | 1 (1-1) of 20 | 2.33 s / 3.23 s | 661 ms (219 ms-1.11 s) | 0 (0-0) | | 50Mbit-40ms | chronozarr | 2 | 192 (192-192) | 177 (177-177) | 22 (12-22) | 4.14 s (3.60 s-4.14 s) | 1 (1-1) of 20 | 3.24 s / 4.14 s | 0 ms (0 ms-0 ms) | 20 (20-20) | | 50Mbit-40ms | zarr-layer | 2 | 180 (180-180) | 171 (171-171) | 19 (13-19) | 5.86 s (3.66 s-5.86 s) | 1 (1-1) of 20 | 4.96 s / 5.86 s | 2.40 s (377 ms-2.40 s) | 0 (0-0) | | 10Mbit-100ms | chronozarr | 2 | 192 (192-192) | 177 (177-177) | 10 (10-10) | 10.3 s (10.2 s-10.3 s) | 1 (1-1) of 20 | 9.35 s / 10.3 s | 0 ms (0 ms-0 ms) | 6 (5-6) | | 10Mbit-100ms | zarr-layer | 2 | 180 (180-180) | 171 (171-171) | 10 (10-10) | 10.2 s (10.2 s-10.2 s) | 1 (1-1) of 20 | 9.35 s / 10.2 s | 2.64 s (2.53 s-2.64 s) | 0 (0-0) | Paced scrub: both tools show every step. zarr-layer transfers exactly the 180 chunks it needs (214 MB, 10.7 MB a step) and aborts nothing; chronozarr transfers 235 to 297 MB: speculative requests for later steps (28 to 94 of its 219 to 275 requests are aborted when the user moves on) and, on the faster links, data that keeps arriving after the last step (up to 59 MB in the next 3 s). The median step takes 0.42 s (chronozarr) against 0.45 s (zarr-layer) on the natural link, 2.1 against 2.3 s at 50 Mbit/s and 10.3 against 8.6 s at 10 Mbit/s, where the link is the limit (10.7 MB need 8.6 s) and chronozarr's prefetch is on it while the next step is requested. The remote runs at 50 Mbit/s vary with the CDN (zarr-layer's scrub took 38.9 s in one run and 65.1 s in the other). zarr-layer replaces the nine regions of a step one by one as their chunks arrive, so the canvas showed a mixture of old and new timestep for 9.7 s of its 17 s scrub (natural) and 59 s of 176 s (10 Mbit/s); chronozarr draws a step only when all nine cells are there (no mixed frame). Burst scrub: neither tool keeps up with one step every 100 ms (a step is 9 chunk reads and at least one round trip). Both skip to the last step (1 of 20 shown exactly) and abort nearly every request (zarr-layer 169 to 171 of 180, chronozarr 175 to 177 of 192); the bytes that count are those of the final step plus what was already in flight (10 to 30 MB). From the first input to the last frame: 2.9 s (chronozarr) against 3.2 s (zarr-layer) on the natural link, 4.1 against 5.9 s at 50 Mbit/s, 10.3 against 10.2 s at 10 Mbit/s. How the cold open divides: of the 6.8 s (natural link) both tools need, 6.5 to 6.6 s is the wait for the nine shard-index reads, the last request before any chunk can be asked for. Each is a 1,876-byte range read at the end of a shard object of 83 to 174 MB; at the CDN they are `cf-cache-status: MISS` and the body arrives 1.7 to 11 s after the headers (`curl` for one of them: 2.4, 4.5 and 5.2 s on a miss, 0.34 s on a hit). Both tools make the same nine reads, so the open times on the published store do not separate them; the spread is the CDN's. zarr-layer reads the shard sizes with 9 `HEAD` requests first and chronozarr takes them from the `shard_bytes` attribute, which is why it needs 19 requests at open and zarr-layer 30 (33 with its defaults). This measurement is why the encoder default became unsharded (spec section 13): an unsharded store has no shard index and its largest object is one 1.8 MB chunk. The cold-open numbers for the unsharded layout will be measured on the live store after its upload; every number in this document is for the sharded `chronozarr-3` unless it says otherwise. #### 1.6 Results: the same files from the local range server ##### Cold open, level 1, 9 cells: same files, local range server Median over repetitions, range in parentheses. Burst and paced runs both open the store the same way and are pooled. | link | tool | runs | time to first complete frame | of which until the shard-index reads are done | requests | MB | | ------------ | ---------- | ---- | ---------------------------- | --------------------------------------------- | ---------- | ---------------- | | natural | chronozarr | 6 | 101 ms (95 ms-108 ms) | 9 ms (9 ms-12 ms) | 19 (19-19) | 10.0 (10.0-10.0) | | natural | zarr-layer | 6 | 127 ms (122 ms-130 ms) | 33 ms (30 ms-34 ms) | 30 (30-30) | 10.0 (10.0-10.0) | | 50Mbit-40ms | chronozarr | 4 | 2.12 s (2.11 s-2.12 s) | 191 ms (190 ms-193 ms) | 19 (19-19) | 10.0 (10.0-10.0) | | 50Mbit-40ms | zarr-layer | 4 | 1.89 s (1.89 s-1.89 s) | 287 ms (280 ms-288 ms) | 30 (30-30) | 10.0 (10.0-10.0) | | 10Mbit-100ms | chronozarr | 4 | 9.69 s (9.69 s-9.69 s) | 476 ms (475 ms-477 ms) | 19 (19-19) | 10.0 (10.0-10.0) | | 10Mbit-100ms | zarr-layer | 4 | 8.67 s (8.67 s-8.67 s) | 680 ms (678 ms-682 ms) | 30 (30-30) | 10.0 (10.0-10.0) | ##### Scrub forward 20 timesteps, paced: same files, local range server | link | tool | runs | requests | of which aborted | MB transferred | scrub duration | steps shown exactly | step latency median / p95 | time showing mixed-timestep frames | MB in the next 3 s | | ------------ | ---------- | ---- | ------------- | ---------------- | -------------- | ------------------------- | ------------------- | ------------------------- | ---------------------------------- | ------------------ | | natural | chronozarr | 3 | 305 (305-314) | 7 (3-11) | 351 (350-355) | 2.18 s (2.18 s-2.19 s) | 20 (20-20) of 20 | 0 ms / 4 ms | 0 ms (0 ms-0 ms) | 2 (1-2) | | natural | zarr-layer | 3 | 180 (180-180) | 0 (0-0) | 214 (214-214) | 3.79 s (3.78 s-3.85 s) | 20 (20-20) of 20 | 80 ms / 100 ms | 824 ms (805 ms-839 ms) | 0 (0-0) | | 50Mbit-40ms | chronozarr | 2 | 250 (249-250) | 70 (68-70) | 250 (249-250) | 43.0 s (42.9 s-43.0 s) | 20 (20-20) of 20 | 2.03 s / 2.82 s | 0 ms (0 ms-0 ms) | 16 (16-16) | | 50Mbit-40ms | zarr-layer | 2 | 180 (180-180) | 0 (0-0) | 214 (214-214) | 37.6 s (37.6 s-37.6 s) | 20 (20-20) of 20 | 1.76 s / 1.98 s | 14.6 s (14.6 s-14.6 s) | 0 (0-0) | | 10Mbit-100ms | chronozarr | 2 | 251 (251-251) | 71 (71-71) | 245 (244-245) | 204.7 s (204.2 s-204.7 s) | 20 (20-20) of 20 | 10.2 s / 14.1 s | 0 ms (0 ms-0 ms) | 4 (2-4) | | 10Mbit-100ms | zarr-layer | 2 | 180 (180-180) | 0 (0-0) | 214 (214-214) | 175.6 s (175.4 s-175.6 s) | 20 (20-20) of 20 | 8.63 s / 9.73 s | 72.6 s (72.5 s-72.6 s) | 0 (0-0) | ##### Scrub forward 20 timesteps, burst (one step every 100 ms): same files, local range server | link | tool | runs | requests | of which aborted | MB transferred | scrub duration | steps shown exactly | step latency median / p95 | time showing mixed-timestep frames | MB in the next 3 s | | ------------ | ---------- | ---- | ------------- | ---------------- | -------------- | ---------------------- | ------------------- | ------------------------- | ---------------------------------- | ------------------ | | natural | chronozarr | 3 | 306 (301-311) | 9 (7-10) | 346 (344-352) | 1.90 s (1.90 s-1.90 s) | 20 (20-20) of 20 | 0 ms / 53 ms | 0 ms (0 ms-0 ms) | 3 (2-3) | | natural | zarr-layer | 3 | 180 (180-180) | 0 (0-0) | 214 (214-214) | 1.98 s (1.98 s-1.98 s) | 20 (20-20) of 20 | 79 ms / 84 ms | 863 ms (789 ms-900 ms) | 0 (0-0) | | 50Mbit-40ms | chronozarr | 2 | 192 (192-192) | 177 (177-177) | 14 (14-14) | 3.62 s (3.61 s-3.62 s) | 1 (1-1) of 20 | 2.72 s / 3.62 s | 0 ms (0 ms-0 ms) | 17 (17-17) | | 50Mbit-40ms | zarr-layer | 2 | 180 (180-180) | 171 (171-171) | 16 (16-16) | 3.60 s (3.59 s-3.60 s) | 1 (1-1) of 20 | 2.70 s / 3.60 s | 681 ms (680 ms-681 ms) | 0 (0-0) | | 10Mbit-100ms | chronozarr | 2 | 192 (192-192) | 177 (177-177) | 10 (10-10) | 12.1 s (12.1 s-12.1 s) | 1 (1-1) of 20 | 11.2 s / 12.1 s | 0 ms (0 ms-0 ms) | 3 (3-3) | | 10Mbit-100ms | zarr-layer | 2 | 180 (180-180) | 171 (171-171) | 10 (10-10) | 10.2 s (10.2 s-10.2 s) | 1 (1-1) of 20 | 9.30 s / 10.2 s | 3.42 s (3.41 s-3.42 s) | 0 (0-0) | Paced scrub: zarr-layer transfers 214 MB exactly; chronozarr 245 MB at 10 Mbit/s (+14 %), 250 MB at 50 Mbit/s (+17 %) and 351 MB unthrottled (+64 %), where its prefetch fills what the link can carry. The median step is 2.03 s (chronozarr) against 1.76 s (zarr-layer) at 50 Mbit/s (+15 %), 10.2 against 8.63 s at 10 Mbit/s (+18 %), and 0 ms against 80 ms unthrottled, where chronozarr has already loaded the next steps. zarr-layer showed mixed-timestep frames for 0.8 s of 3.8 s unthrottled, 14.6 s of 37.6 s at 50 Mbit/s and 72.6 s of 175.6 s at 10 Mbit/s; chronozarr never did. Burst scrub: unthrottled both keep up (20 of 20 steps shown, 1.9 s against 2.0 s); at 50 and 10 Mbit/s both skip to the last step, as on the published store. With the CDN out of the picture the open is transfer-bound (10 MB: 1.6 s at 50 Mbit/s, 8.0 s at 10 Mbit/s plus round trips). zarr-layer reaches its first complete frame 0.23 s (11 %) and 1.0 s (10 %) sooner than chronozarr at 50 and 10 Mbit/s and 26 ms later unthrottled. chronozarr's last byte arrives 0.01 s (50 Mbit/s) and 0.41 s (10 Mbit/s) after zarr-layer's (1.89 s against 1.90 s, 8.66 s against 9.08 s); the rest of the difference is what happens afterwards: chronozarr takes 0.07 s (unthrottled), 0.23 s and 0.61 s after the last byte to paint, zarr-layer 0.01 to 0.08 s. The sequential round trips before the chunk reads are 3 for chronozarr (`zarr.json`, shard index, chunk) and 5 for zarr-layer with `zarrVersion: 3` (`zarr.json`, the `band` and `time` coordinate arrays, `HEAD`, shard index, chunk; 8 with its defaults). #### 1.7 What the zarr-layer numbers show and do not show They show the cost of a step at a fixed level for the two designs on this store: the same 9 chunk reads, and what each does around them (chronozarr: prefetch and whole frames; zarr-layer: nothing speculative, per-region replacement, abort and re-request on every new selector). They show that the viewer's speculation is a trade: it takes the idle link (steps cost nothing) and loses to a tool that fetches only what is on screen when the link is already full. They do not show: * anything about zarr-layer's design goals: arbitrary variables and dimensions, any CRS reprojected on the GPU to Web Mercator and globe projections, queries, other basemaps. chronozarr draws in the store's native projection in its own canvas and decodes a single dtype layout; the render cost of the two is not isolated and zarr-layer's reprojection is not charged to chronozarr. * cold open of the published store (CDN-bound, same for both) or the CDN at all beyond the stalls above; one machine, one Chromium, one store. * chronozarr's default behaviour: with the level pinned its coarse-first staging, adaptive movie level and playback buffering are off, as asked. The first frame a user sees with the viewer's defaults is earlier than the numbers here. * zarr-layer with a debounced selector (its README advises one; the burst runs deliberately do not), with `renderingMode '2d'`, or with its decoded-chunk cache warm (all runs are cold). * Bytes of aborted requests beyond what Chromium reported (data in flight when a stream is reset is not seen); the aborted counts are exact. ### 2. One COG per date #### 2.1 Method * **Representations.** *chronozarr*: the local store, read with the chronozarr JS reader (`js/chronozarr/decoder.js`, `openStore(url, { workers: 0 })`, `getCell`; explicit reads only, no `prefetch`). *COG*: `uv run chronozarr export-cog ... --level 0` writes one `L0_.tif` per timestep from the same decoded values: GDAL COG driver, **DEFLATE with predictor 2**, 512 x 512 tiles, pixel interleave, overviews at 2, 4 and 8 (AVERAGE) written by the same call, so no separate GDAL step was needed. 117 files, 6.66 GB (56.9 MB per date) against the store's 6.45 GB. Read with geotiff.js 3.0.5 in two configurations: `fromUrl(url)` (its defaults; no block cache) and `fromUrl(url, { blockSize: 65536, cacheSize: 400 })` (64 KB aligned blocks). * **Serving.** Both over the same local range server (`js/support/static-server.js`), reached by Node's `fetch`. A recorder wraps `fetch` and logs method, range, status, body bytes, start and end time of every request, tagged with the phase and step it belongs to. * **Session**, fixed level 1 with a 3 x 3 cell view (the whole level) except where noted, state kept from phase to phase (opened files, shard indexes, caches): | phase | what is read | | ---------------- | ----------------------------------------------------------------------------------- | | open | date 40: metadata/header, then the 9 cells | | scrub forward 20 | dates 41 to 60, one after the other (each step waits for the previous) | | jump | date 100 | | playback | dates 61 to 80, pipelined (no step waits for another) | | zoom | an uncached 2 x 2 cell area of **level 0**, rows 2 to 3, columns 2 to 3, at date 80 | | pixel history | one L0 pixel (row 1388, col 1380) at all 117 dates | * **Classification.** Requests are metadata (the store's root `zarr.json`), header (a COG read that ends before the file's first tile), index (a read of the last bytes of a shard, from `shard_bytes`) or pixel (the rest). * **Delivery time.** `bench/cog/model.mjs`: each request waits for its dependencies, then one round trip, then shares the link equally with the other requests that are transferring (processor sharing), with at most `parallel` requests in flight: **6 for COG** (HTTP/1.1 connections per origin in a browser) and **12 for the chronozarr reader** (its own cap). 500 bytes of headers are added to each request. Dependencies are inferred from the recorded local timing: a request of the same file or shard that had finished when this one started is a dependency, so a header read chain stays a chain, and a chunk read waits for its shard index; a shard index waits for the store metadata. Sequential phases (open, scrub, jump, zoom) start each step when the previous one is done; pipelined phases (playback, pixel history) are limited only by dependencies and `parallel`. The model has tests (`bench/cog/model.test.mjs`; breaking the link sharing makes two of them fail). Links: 50 Mbit/s + 40 ms and 10 Mbit/s + 100 ms as given, and `natural` = the medians of three probe rounds (start, middle, end of the run) of 1-byte and 8 MB range reads against the published store: 90 Mbit/s and 112 ms, observed 27 to 964 Mbit/s and 27 to 454 ms. The natural numbers move with that probe; compare their ratios, not the seconds. * **Wall-clock** is measured locally for each phase, median of 3 repetitions after one unrecorded warm-up (which warms the file cache); the three representations rotate in order. It includes decoding. * **Equality.** The pixel history (117 timesteps x 4 bands, level 0) of both readers is compared with a reference read straight from the Zarr arrays with no chronozarr code (`bench/cog/pixel_reference.py`); the run fails if any value differs. #### 2.2 Tables ##### Requests and bytes | phase | representation | header / index / metadata requests | header / index / metadata bytes | pixel requests | pixel bytes | total requests | total bytes | | ------------------------------------------ | ------------------------ | ---------------------------------- | ------------------------------- | -------------- | ----------- | -------------- | ----------- | | open | chronozarr | 10 | 55.9 kB | 9 | 10.0 MB | 19 | 10.0 MB | | open | COG, geotiff.js defaults | 7 | 3.2 kB | 9 | 10.4 MB | 16 | 10.4 MB | | open | COG, 64 KB blocks | 1 | 65.5 kB | 1 | 10.5 MB | 2 | 10.6 MB | | scrub forward 20 | chronozarr | 0 | 0.0 kB | 180 | 214 MB | 180 | 214 MB | | scrub forward 20 | COG, geotiff.js defaults | 140 | 64.8 kB | 180 | 217 MB | 320 | 217 MB | | scrub forward 20 | COG, 64 KB blocks | 20 | 1.3 MB | 20 | 218 MB | 40 | 220 MB | | jump to a distant date | chronozarr | 0 | 0.0 kB | 9 | 10.0 MB | 9 | 10.0 MB | | jump to a distant date | COG, geotiff.js defaults | 7 | 3.2 kB | 9 | 10.4 MB | 16 | 10.4 MB | | jump to a distant date | COG, 64 KB blocks | 1 | 65.5 kB | 1 | 10.5 MB | 2 | 10.6 MB | | playback, 20 consecutive dates | chronozarr | 0 | 0.0 kB | 180 | 204 MB | 180 | 204 MB | | playback, 20 consecutive dates | COG, geotiff.js defaults | 140 | 64.8 kB | 180 | 211 MB | 320 | 211 MB | | playback, 20 consecutive dates | COG, 64 KB blocks | 20 | 1.3 MB | 20 | 212 MB | 40 | 214 MB | | zoom to 4 uncached cells, next finer level | chronozarr | 4 | 7.5 kB | 4 | 5.6 MB | 8 | 5.6 MB | | zoom to 4 uncached cells, next finer level | COG, geotiff.js defaults | 8 | 0.0 kB | 4 | 5.7 MB | 12 | 5.7 MB | | zoom to 4 uncached cells, next finer level | COG, 64 KB blocks | 0 | 0.0 kB | 4 | 5.8 MB | 4 | 5.8 MB | | one pixel, complete history | chronozarr | 0 | 0.0 kB | 116 | 164 MB | 116 | 164 MB | | one pixel, complete history | COG, geotiff.js defaults | 682 | 0.2 MB | 117 | 164 MB | 799 | 164 MB | | one pixel, complete history | COG, 64 KB blocks | 75 | 4.9 MB | 116 | 170 MB | 191 | 175 MB | ##### Delivery time from the recorded requests Links: natural = 90 Mbit/s, 112 ms; 50Mbit-40ms = 50 Mbit/s, 40 ms; 10Mbit-100ms = 10 Mbit/s, 100 ms (natural: medians of 3 probe rounds to the published store, observed 27 to 964 Mbit/s and 27 to 454 ms). Parallel requests: COG 6, chronozarr reader 12. 500 bytes of headers per request. **natural** (90 Mbit/s, 112 ms). Time, and in parentheses the ratio to chronozarr: | phase | chronozarr | COG, geotiff.js defaults | COG, 64 KB blocks | | ------------------------------------------ | ---------- | ------------------------ | ----------------- | | open | 1.23 s | 1.82 s (1.48x) | 1.16 s (0.95x) | | scrub forward 20 | 21.2 s | 37.2 s (1.75x) | 24.0 s (1.13x) | | jump to a distant date | 999 ms | 1.82 s (1.83x) | 1.16 s (1.16x) | | playback, 20 consecutive dates | 18.2 s | 21.5 s (1.18x) | 19.3 s (1.06x) | | zoom to 4 uncached cells, next finer level | 719 ms | 842 ms (1.17x) | 631 ms (0.88x) | | one pixel, complete history | 14.7 s | 23.6 s (1.60x) | 16.1 s (1.09x) | **50Mbit-40ms** (50 Mbit/s, 40 ms). Time, and in parentheses the ratio to chronozarr: | phase | chronozarr | COG, geotiff.js defaults | COG, 64 KB blocks | | ------------------------------------------ | ---------- | ------------------------ | ----------------- | | open | 1.73 s | 1.99 s (1.15x) | 1.77 s (1.02x) | | scrub forward 20 | 35.0 s | 41.2 s (1.18x) | 36.8 s (1.05x) | | jump to a distant date | 1.64 s | 1.99 s (1.22x) | 1.77 s (1.08x) | | playback, 20 consecutive dates | 32.6 s | 34.1 s (1.05x) | 34.3 s (1.05x) | | zoom to 4 uncached cells, next finer level | 971 ms | 1.03 s (1.06x) | 974 ms (1.00x) | | one pixel, complete history | 26.3 s | 27.5 s (1.04x) | 28.1 s (1.07x) | **10Mbit-100ms** (10 Mbit/s, 100 ms). Time, and in parentheses the ratio to chronozarr: | phase | chronozarr | COG, geotiff.js defaults | COG, 64 KB blocks | | ------------------------------------------ | ---------- | ------------------------ | ----------------- | | open | 8.33 s | 9.14 s (1.10x) | 8.64 s (1.04x) | | scrub forward 20 | 173.0 s | 189.9 s (1.10x) | 179.8 s (1.04x) | | jump to a distant date | 8.08 s | 9.15 s (1.13x) | 8.64 s (1.07x) | | playback, 20 consecutive dates | 163.0 s | 169.8 s (1.04x) | 171.2 s (1.05x) | | zoom to 4 uncached cells, next finer level | 4.65 s | 4.86 s (1.04x) | 4.77 s (1.02x) | | one pixel, complete history | 131.6 s | 133.1 s (1.01x) | 140.3 s (1.07x) | ##### Per step (sequential phases): median step time | phase | natural: chronozarr | natural: COG, geotiff.js defaults | natural: COG, 64 KB blocks | 50Mbit-40ms: chronozarr | 50Mbit-40ms: COG, geotiff.js defaults | 50Mbit-40ms: COG, 64 KB blocks | 10Mbit-100ms: chronozarr | 10Mbit-100ms: COG, geotiff.js defaults | 10Mbit-100ms: COG, 64 KB blocks | | ---------------- | ------------------- | --------------------------------- | -------------------------- | ----------------------- | ------------------------------------- | ------------------------------ | ------------------------ | -------------------------------------- | ------------------------------- | | scrub forward 20 | 1.06 s | 1.87 s | 1.20 s | 1.74 s | 2.07 s | 1.84 s | 8.61 s | 9.53 s | 9.01 s | ##### Wall-clock on this machine, local range server (median over repetitions; includes decoding) | phase | chronozarr | COG, geotiff.js defaults | COG, 64 KB blocks | COG, geotiff.js defaults / chronozarr | COG, 64 KB blocks / chronozarr | | ------------------------------------------ | ---------- | ------------------------ | ----------------- | ------------------------------------- | ------------------------------ | | open | 69 ms | 284 ms | 291 ms | 4.12x | 4.22x | | scrub forward 20 | 1.13 s | 5.58 s | 5.87 s | 4.95x | 5.21x | | jump to a distant date | 58 ms | 275 ms | 285 ms | 4.74x | 4.91x | | playback, 20 consecutive dates | 1.15 s | 5.45 s | 5.41 s | 4.73x | 4.70x | | zoom to 4 uncached cells, next finer level | 33 ms | 144 ms | 147 ms | 4.32x | 4.39x | | one pixel, complete history | 881 ms | 2.97 s | 2.95 s | 3.38x | 3.35x | ##### Checks * Pixel history (117 timesteps x 4 bands, level 0): identical to `bench/results/pixel-reference-L0-r1388-c1380.json` for both readers: true. * Level 1 of 2019-09-01T00:00:00Z (COG overview 1382 x 1379, store level 1383 x 1380), 7623112 values compared: 3.43 % identical, 9.99 % within 1, mean absolute difference 36.6207, maximum 5935. * Bytes on disk: store 6.45 GB; 117 COGs 6.66 GB (56.9 MB per date). Per-phase requests are identical across repetitions for every representation (fingerprints in the JSON). The pixel-history phase of the 64 KB-block configuration reads 4.9 MB of header blocks (64 KB for each of 75 files not opened before) and a few percent extra pixel bytes from block alignment. #### 2.3 Codec and decode checks One date (2016-05-01, `bench/results/codec-sizes-2016-05.json`), bytes to deliver all four levels of it: | representation | bytes | | -------------------------------------------------------- | ------------------- | | chronozarr store, zstd 5 chunks (from the shard indexes) | 56,852,283 | | COG as exported (DEFLATE, predictor 2) | 57,876,521 (+1.8 %) | | COG, ZSTD 5, predictor 2 | 53,843,180 (-5.3 %) | | COG, ZSTD 5, no predictor | 62,225,767 (+9.5 %) | The store's chunks are band-planar; a pixel-interleaved COG compresses worse with zstd unless the predictor is on. Decoding the 9-cell view of level 1 with the bytes already in memory (`bench/results/decode-bench-2016-05.json`, median of 7): COG DEFLATE + predictor in geotiff.js (pako) 257 ms; COG ZSTD + predictor in geotiff.js (WASM) 116 ms; chronozarr reader (WASM zstd, one thread) 37 ms. That is the wall-clock gap in the table: 20 steps x 257 ms is 5.1 s of the 5.6 s measured. Overview equivalence (date 2019-09-01, level 1): the COG overview is 1382 x 1379 px, the store's level 1 is 1383 x 1380; the store's level 1 is exactly the floor of the 2 x 2 block mean of level 0 over the 1382 x 1379 complete blocks (checked), GDAL resamples by 2765/1382 = 2.0007 for an odd size, so the two grids drift apart by up to one source pixel at the far edge. 3.4 % of the 7.6 million values are identical, 10 % are within 1, the mean absolute difference is 36.6 DN on a mean of 1126 (mean difference 22 DN in the first 50 columns, 50 DN in the last), the maximum 5935. The overviews have the same ground footprint and resolution; they are not the same pixels. #### 2.4 What these numbers show and do not show They show, for imagery that is the same at level 0 (one pixel's history over all 117 dates identical in both readers and the Zarr reference; the whole COG level 0 equals the store's for the one date checked, 2019-09-01), at compression equal to within 2 % for the date checked, and with equivalent but not identical overviews: * bytes are not the difference (+0 to +4 % per phase for the exported COGs, +1.4 % over the 20-step scrub, +3.2 % on disk); requests and their dependencies are. A new COG costs a header first: 7 chained reads with geotiff.js defaults, 1 (64 KB) with blocks. chronozarr pays a shard index once per shard (9 at open, 4 at zoom), then only chunks. * the gap in delivery time is a latency effect and closes as the link becomes bandwidth-bound: scrub forward 20 is 1.75x / 1.13x (defaults / 64 KB blocks) at the natural link, 1.18x / 1.05x at 50 Mbit/s + 40 ms, 1.10x / 1.04x at 10 Mbit/s + 100 ms; pixel history 1.60x / 1.09x, 1.04x / 1.07x, 1.01x / 1.07x. Blocks cost bytes where reads are small: the pixel-history phase reads 175 MB instead of 164 MB. * reading a pixel's history needs the whole 512 x 512 x 4 tile of every date in both layouts (164 MB for 116 or 117 reads); neither is a time-series format. A layout chunked along time would make that one small read. * local wall-clock is decoder-bound (section 2.3). They do not show: * anything about a real CDN or the browser: both are served by a local HTTP/1.1 server to Node. Request count matters differently behind HTTP/2 or HTTP/3, TLS and slow start, and edge caches treat 117 files of 57 MB and a few shard objects of up to 174 MB differently; section 1.5 found 2 to 11 s for a shard-tail read on a cache miss, which one COG per date does not have (its header read is 3 to 64 KB at the start of a 57 MB file). This comparison does not measure that, and it would count against the shard layout. * a COG reader better than geotiff.js's two configurations. A COG client that reads the header in one request and then exactly the tiles would sit between the two columns on requests and below both on bytes; the floor per new date is two sequential round trips (header, tiles) against one for chronozarr with its index cached. * other COG layouts (ZSTD with predictor is 7 % smaller than the exported files; band-sequential; larger tiles), imagery with more or fewer dates (headers matter more for daily data), writing or appending, or the ecosystem: GDAL and QGIS read COGs natively and chronozarr needs its readers or `export-cog`. * the model's assumptions: one bottleneck link shared equally, no slow start, dependencies inferred from local timing, 500 bytes of headers per request, `natural` taken from a probe that varied by an order of magnitude. ### 3. Commands Everything from the repository root unless noted. `bench/.npmrc` keeps `min-release-age=7` (every locked package is at least 7.1 days old: the youngest, `maplibre-gl` 6.11.2, was 7.16 days old when installed). ```bash # setup cd bench && npm ci && node zarr-layer/build.mjs && cd .. # A. compatibility cd bench node zarr-layer/probe.mjs --out results/zarr-layer-probe-as-published.json node zarr-layer/probe.mjs --extra '{"crs":"EPSG:32718","bounds":[485650,9142230,513240,9169880]}' --out results/zarr-layer-probe-crs-and-bounds-options.json node zarr-layer/probe.mjs --patch no-pixels-per-tile --out results/zarr-layer-probe-without-pixels-per-tile.json node zarr-layer/values.mjs # A. comparison (tool order alternates; --warmup adds an unrecorded run of each tool first) node zarr-layer/run-all.mjs --source remote --profile natural --reps 3 --warmup node zarr-layer/run-all.mjs --source remote --profile 50Mbit-40ms --reps 2 --warmup node zarr-layer/run-all.mjs --source remote --profile 10Mbit-100ms --reps 2 node zarr-layer/run-all.mjs --source local --profile natural --reps 3 node zarr-layer/run-all.mjs --source local --profile 50Mbit-40ms --reps 2 node zarr-layer/run-all.mjs --source local --profile 10Mbit-100ms --reps 2 node zarr-layer/summarize.mjs # the tables of sections 1.5 and 1.6 cd .. # B. data uv run chronozarr export-cog data/stores/ucayali_santa_maria/chronozarr-3 data/cogs/ucayali_santa_maria --level 0 uv run python bench/cog/pixel_reference.py data/stores/ucayali_santa_maria/chronozarr-3 0 1388 1380 bench/results/pixel-reference-L0-r1388-c1380.json uv run python bench/cog/codec_sizes.py data/stores/ucayali_santa_maria/chronozarr-3 data/cogs/ucayali_santa_maria/L0_2016-05-01.tif 2 data/cogs-variants/2016-05 bench/results/codec-sizes-2016-05.json # B. session, tables, decode, model tests cd bench node cog/session.mjs --reps 3 # writes results/cog-vs-chronozarr.json node cog/report.mjs # the tables of section 2.2 node cog/decode-bench.mjs 2 7 > results/decode-bench-2016-05.json node --test cog/model.test.mjs ``` Files: `bench/package.json`, `bench/package-lock.json`, `bench/.npmrc`, `bench/.gitignore`; `bench/lib/harness.mjs` (server, browser, CDP throttling and accounting); `bench/zarr-layer/{page.html,page.js,build.mjs,probe.mjs,values.mjs,compare.mjs,run-all.mjs,summarize.mjs}`; `bench/cog/{session.mjs,recorder.mjs,model.mjs,model.test.mjs,report.mjs,decode-bench.mjs,pixel_reference.py,codec_sizes.py}`; raw results in `bench/results/*.json`. `data/cogs/` and `data/cogs-variants/` are under the ignored `data/`. ## Embedding the viewer `https://chronozarr.org/demo/?store=` opens any chronozarr store. Add `embed=1` and the same page is a compact viewer for an ` ``` Build the `src` with `URLSearchParams` or `encodeURIComponent`: the `store` and `origin` values are URLs inside a URL. `origin` must be your page's `location.origin`. Point the iframe at `/demo/` as above: `/demo` redirects there, preserving the query string. ### 2. URL parameters All parameters go in the iframe URL. Every parameter the full viewer understands works with `embed=1` too, so a permalink from the full viewer becomes an embed by adding `embed=1&origin=...`. | Parameter | Value | Default | Effect | | ---------- | ---------------------------------------------------------------------------------- | ------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `store` | base URL of a chronozarr store (the directory that holds `zarr.json`), URL-encoded | none: the viewer shows an error and sends `chronozarr:error` (`no_store`) | The store to open. An embed never falls back to another store and never fetches the catalog. | | `embed` | `1` | off | Compact layout, and the `postMessage` API. Any other value is not an embed. | | `controls` | `0` | on | Only the canvas and a thin readout of the time of the frame on screen. No header, wordmark, inspector or transport controls. Clicks are still reported. Needs `embed=1`. | | `theme` | `light` | `dark` | Light palette for the chrome (header, timeline, inspector). The canvas and the pixels it draws are the same in both themes. Needs `embed=1`. | | `origin` | the origin of the host page, URL-encoded | the origin of `document.referrer` | The only origin the viewer listens to and sends to; see section 3. | | `t` | timestep index | 0 | Initial timestep. | | `p` | `true_color`, `false_color`, `ndvi`, `ndwi`, `water`, `band` | the first product the store's bands allow | Initial product. | | `b` | band name | first band | Band of the single-band product (`p=band`). | | `z` | number above 0 | the fitted view | Initial zoom, in CSS pixels per full-resolution data pixel (so the same ground width shows in a window of any size). | | `c` | `x,y` | the store center | Initial view center, in the store's projected coordinates (level-0 pixels for a store with no georeferencing). | The embed layout drops the catalog selector and the export button, keeps the product buttons and the timeline, and shows the inspector as a drawer that opens on a click (Escape or the close button closes it), at every width. A small "chronozarr" wordmark in the header links to the full viewer on the same store and view (`target="_blank"`, without the embed parameters). Initial state belongs in the URL: the viewer also keeps the current time, product and camera in the iframe's own address, so a reload of the iframe restores them. ### 3. The origin rule Messages go only between the iframe and its parent, and only for one origin: * `origin=` in the iframe URL names it. It must be exactly an `http(s)` origin such as `https://example.com` or `http://localhost:3000`: no path, no `*`. * Without `origin`, the viewer uses the origin of `document.referrer`, which is your page when your `Referrer-Policy` sends at least the origin (the browser default does). With `Referrer-Policy: no-referrer` there is no referrer, so pass `origin`. * If neither gives a valid origin, or `origin` is present but invalid (the viewer does not then fall back to the referrer), the viewer still works but the API is off, and it says why with a `console.warn`. * The API is also off when the page is not framed. The viewer reads a message only if `event.origin` is that origin **and** `event.source` is its parent window. Anything else is dropped without a reply, an error, or any effect. It sends every message to that origin only, never to `*`; if the real parent is another origin, the browser refuses to deliver and warns "The target origin provided ... does not match the recipient window's origin" in the console of the viewer frame, which is the symptom of a wrong `origin`. A page opened from `file://` has the origin `null` and cannot be a host: serve it over `http(s)`, `localhost` included. ### 4. Messages All messages are plain JSON objects with a `type` that starts with `chronozarr:`. The contract version is `v: 1`. Every message the viewer sends carries `v: 1`. A message to the viewer may leave `v` out (meaning 1); any other value is answered with `chronozarr:error` (`bad_message`). The viewer ignores messages whose `type` is not its own (another protocol on the same window) and the types it sends itself (a host that echoes messages back gets no errors). It also ignores unknown fields in a `set`, so a later version can add some. #### Host to viewer ##### `chronozarr:set` Change the view. Every field is optional; the ones present are applied together, in this order: product and band, time, camera, speed, playback. | Field | Type | Meaning | | --------- | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `t` | integer, or ISO 8601 string | A timestep index (0 to `times.length - 1`), or a date such as `"2024-03-01"` or `"2024-03-01T12:00:00Z"`: the nearest timestep wins (the earlier one on a tie, the first or last one outside the series). A date-time without a zone is read as UTC. | | `product` | string | A product id from `ready.products` that is `available`. | | `band` | string | A band name from `ready.bands`. Chooses the band of the single-band product: send `product: "band"` with it to show it. | | `zoom` | number above 0 | CSS pixels per full-resolution data pixel, as `z=`. Limited like the mouse: out to half of the fitted view, in to 16 canvas pixels per data pixel. | | `center` | `{x, y}` or `{lon, lat}` | View center: `x`, `y` in the store's coordinates (as `c=`; level-0 pixels for a store with no georeferencing), or `lon`, `lat` in WGS84 degrees (needs a georeferenced store whose CRS is WGS84 UTM, EPSG:3857 or EPSG:4326). When `x` and `y` are both present they win, so the `center` of a `chronozarr:view` message can be sent back as it is. | | `playing` | boolean | Start or stop playback. | | `speed` | number above 0 | Playback speed in timesteps per second, snapped to the viewer's steps: 1 to 15, 20, 24, 30, 40, 48, 60. | A `set` is accepted or rejected as a whole. If a field is invalid, or the store cannot satisfy it (a `t` past the end, a product the store's bands do not allow, a band it does not have, `lon`/`lat` for a store without a transform), nothing is applied and the viewer answers with `chronozarr:error` (`bad_set`) naming the field. A `set` that arrives before the store has opened is kept (several are merged, later fields win) and applied right after `chronozarr:ready`. For the first frame, prefer the URL parameters: with them the viewer paints the requested view at once. ```js frame.postMessage({ v: 1, type: 'chronozarr:set', t: '2024-03-01' }, VIEWER); frame.postMessage({ type: 'chronozarr:set', product: 'ndvi', zoom: 2, center: { lon: -73.6, lat: -9.4 } }, VIEWER); frame.postMessage({ type: 'chronozarr:set', speed: 10, playing: true }, VIEWER); ``` ##### `chronozarr:get` ```js frame.postMessage({ type: 'chronozarr:get' }, VIEWER); ``` The viewer answers with `chronozarr:ready` again, whose `state` is current. Use it when your listener was attached after the first `ready` (a cached iframe, a framework that mounts late) and to read `state.playing`. Before the store has opened there is no answer; `ready` follows when it does. #### Viewer to host ##### `chronozarr:ready` Sent when the store's metadata has been read and the view is set up, before the first frame is painted. Commands can be sent from here on. ```json { "v": 1, "type": "chronozarr:ready", "store": { "url": "https://data.tileripper.com/ucayali_santa_maria/chronozarr-4", "name": "chronozarr-3", "crs": "EPSG:32618" }, "times": ["2015-11-01T00:00:00Z", "2015-12-01T00:00:00Z"], "bands": [{ "name": "B02", "common_name": "blue", "units": "reflectance", "scale": 0.0001, "offset": 0 }], "products": [{ "id": "true_color", "name": "True color", "available": true }], "levels": [{ "lod": 0, "width": 2759, "height": 2765, "resolution": 10 }], "state": { "t": 0, "time": "2015-11-01T00:00:00Z", "product": "true_color", "band": "B02", "zoom": 0.28, "center": { "x": 653100, "y": 9013200, "lon": -73.6, "lat": -9.4 }, "playing": false, "speed": 4 } } ``` `times` and `state.time` are the ISO strings of the store. `bands[].common_name` and `units` are `null` when the store does not say. `levels[].resolution` is the pixel size in the store's units (`null` without georeferencing). In `state.center`, `x` and `y` are in the store's coordinates and `lon`, `lat` are `null` when they cannot be computed. ##### `chronozarr:time` Sent on every change of the timestep: a `set`, a click on the timeline, the arrow keys, the chart, and each step of playback (up to 60 a second). `t` is the timestep asked for; the frame on screen may still be the previous one for a moment while chunks load. ```json { "v": 1, "type": "chronozarr:time", "t": 12, "time": "2016-11-01T00:00:00Z" } ``` ##### `chronozarr:click` Sent when the user clicks a pixel inside the store (not after a drag, not outside the raster), also with `controls=0`. The values are those of the frame on screen: `t` and `time` are its timestep, which is what the numbers belong to, and `level` is the pyramid level it is drawn from (0 is full resolution; a coarser level while a view is still loading, so the values are then those of the larger pixel). ```json { "v": 1, "type": "chronozarr:click", "pixel": { "x": 1204, "y": 311 }, "lon": -73.61204, "lat": -9.38117, "t": 12, "time": "2016-11-01T00:00:00Z", "level": 0, "valid": true, "values": { "B02": 0.0412, "B03": 0.0633, "B04": 0.0587, "B08": 0.2841 } } ``` `pixel` is the column and row of the clicked data pixel at full resolution. `lon` and `lat` are the WGS84 coordinates of that pixel's center, `null` when the store is not georeferenced or its CRS is not WGS84 UTM, EPSG:3857 or EPSG:4326. `values` holds the physical value of each band (stored value times the band's `scale`, plus its `offset`; reflectance for Sentinel-2 bands), keyed by band name. `valid` is `false`, and every value `null`, for a pixel at the store's nodata value or masked out. ##### `chronozarr:view` Sent when the camera has been still for 120 ms after a pan, zoom or resize, and only when it differs from the last view sent. Also follows a `set` that moved the camera. ```json { "v": 1, "type": "chronozarr:view", "zoom": 2, "center": { "x": 653100, "y": 9013200, "lon": -73.6, "lat": -9.4 } } ``` `zoom` and `center` mean what they mean in `set`, so a saved `view` can be sent back to restore the camera. ##### `chronozarr:error` ```json { "v": 1, "type": "chronozarr:error", "code": "bad_set", "message": "zoom must be a number above 0 (CSS pixels per level-0 data pixel), got \"big\"." } ``` | `code` | Cause | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------- | | `bad_message` | A message of yours has an unsupported `v` or an unknown `chronozarr:` type. | | `bad_set` | A `set` with an invalid field, or one the store cannot satisfy. Nothing was applied. | | `no_store` | The iframe URL has no `store`. | | `store_open_failed` | The store could not be opened (not found, no CORS, not a chronozarr store). `message` has the reason. | | `chunk_load_failed` | A chunk could not be read after retries. The cell keeps its previous data and is retried; this is the toast the viewer shows. | | `playback_failed` | Playback paused because the loop could not be buffered. | | `internal` | A command failed unexpectedly. The browser console has the stack. | A wrong or missing `origin` is not reported by message, since there is nobody the viewer may tell; its console says why (section 3). ### 5. Size The viewer reflows down to 320 px wide. Give it at least 320 px of height, and about 4:3 on a wide page. Below 700 px the product buttons give way to one select. The inspector is a drawer over the map at every width, up to 340 px wide (the whole width on a narrow iframe). ```css #chronozarr { width: 100%; aspect-ratio: 4 / 3; min-height: 320px; max-height: 80vh; border: 0; } ``` The map takes the mouse wheel (zoom) and one-finger drag on touch screens (pan) while the pointer is over the iframe, so the page does not scroll from there. Leave some page around the iframe, as a map embed does. ### 6. Stores and access The iframe runs on `chronozarr.org`, in its own origin. It does not inherit your page's login, cookies or headers, and its requests for the store carry no credentials. The store must therefore be readable by anyone who can open the page: publicly readable, `Access-Control-Allow-Origin: *` (or `https://chronozarr.org`), byte ranges and the other requirements of [hosting.md](/guides/hosting); `chronozarr doctor ` checks them. For a private store the host can put a prefix-wide signed URL in `store=`. The viewer appends the query string of `store` to every request it makes, so the token must authorize every object under the prefix: a CloudFront signed URL with a custom policy and a wildcard resource, an Azure Blob container or prefix SAS, or a token your own gateway accepts. A signature for one object (an S3 or GCS presigned URL) does not work, since a store is many files. The token is part of the iframe URL, so anyone who can see the page can read it; give it a short lifetime, and note that the wordmark link carries the same `store` value to the full viewer. `chronozarr.org` sends no `X-Frame-Options` and no `frame-ancestors` restriction: any page may embed it. If your page sets a Content-Security-Policy, allow `frame-src https://chronozarr.org`. A `sandbox` attribute on the iframe needs `allow-scripts allow-same-origin` (without the second the viewer's origin is `null` and your `event.origin` check fails); other sandbox settings are untested. ### 7. Stability The contract is versioned by `v`. Within `v: 1` the viewer may add optional fields to its messages and optional fields to `set`; hosts should ignore fields and message types they do not know. A change that removes or changes the meaning of something bumps `v`, and the viewer then answers a message with the old `v` with `bad_message`. ## chronozarr compared with other formats As of 2026-09-30. Statements about other projects come from their documentation and source and can go stale; check the project if a decision depends on one. Statements about chronozarr are measured or read from [spec/CHRONOZARR.md](/specification) and the tests in this repository. The question each format answers is different, so the table ends with the case for choosing the other tool. | Tool | What it stores and serves | Time axis | Values | CRS and serving | Choose it instead when | | -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **chronozarr** | A Zarr v3 pyramid of `(time, band, y, x)` arrays, one group per level, one object per chunk by default (sharding is an option); optional star-delta temporal encoding | Native: any timestep of any level is at most two chunk reads | Exact stored values in `uint8`, `uint16`, `int16` or `float32`; lossless; per-band `scale` and `offset` | Native projected CRS per store (UTM), never resampled to Web Mercator for storage; static bucket with byte ranges and CORS, no server | (reference row) | | **PMTiles** | One archive of z/x/y tiles (vector or image) addressed by byte range | None. A time series is one archive per timestep | Whatever the tile image encodes: rendered 8-bit pixels, or values packed into color channels at the packing's precision | Web Mercator tile scheme in practice; static bucket with byte ranges, no server | You want finished map tiles (basemaps, vector layers, rendered imagery) for one moment, and broad client support across MapLibre, Leaflet and OpenLayers. You do not need raw values or a time scrubber. | | **Mapbox raster-array** | Multi-band numeric raster tiles in the MRT format, produced by the Mapbox Tiling Service; bands can carry time steps or variables | As bands | Numeric values, decoded in the client | Web Mercator tiles. The format is tied to Mapbox's tiling service and renderer. The decoder code is published in mapbox-gl-js (`src/data/mrt`), so the dependency is on that service and renderer, not on a hidden decoder. | You already build on Mapbox GL JS and the Mapbox Tiling Service, want managed tiling and hosting, and can accept data resampled to Web Mercator. | | **ndpyramid + zarr-layer** | Zarr pyramids written by ndpyramid (reprojected to a tile grid, or coarsened in place) and rendered by CarbonPlan's zarr-layer in MapLibre or Mapbox GL | Through zarr-layer selectors on a time dimension; each timestep is normally its own chunk, so there is no compression across time | Any Zarr dtype, typically float32 | zarr-layer reads `proj:` and `spatial:` attributes and supports arbitrary CRS through proj4 reprojection (its README lists the WGS84 UTM zones as built in). Static Zarr over HTTP. | You want the established MapLibre Zarr layer, float data, selectors over arbitrary dimensions; or your data already comes out of ndpyramid. A chronozarr store with `encoding: none` written after `pixels_per_tile` was dropped from `multiscales` (spec section 13) opens in zarr-layer unmodified (verified on the Ucayali store at level 1, point values equal to a direct Zarr read); a store written earlier opens with zarr-layer's `crs` and `bounds` constructor options. A star-delta store needs an adapter that reconstructs the residuals; stock zarr-layer does not do it. | | **COG + TiTiler** | Single-timestep Cloud Optimized GeoTIFFs with internal tiles and overviews, rendered to image tiles by a tile server | One file per timestep, or STAC mosaics selected server-side | Exact values in any GDAL dtype, lossless compression available; point queries and band math run on the server | Any CRS; the server reprojects to a tile matrix. Needs a running service (container, serverless function) in front of the files. | Your data are already COGs or in a STAC catalog; you need server-side band math, many clients that read COG directly (GDAL, QGIS, rasterio), or single-date imagery; a server is acceptable. | | **GeoZarr** | A set of conventions for geospatial Zarr (CRS, affine transform, multiscales), not a product; the multiscale layout is still a proposal | Whatever the dataset's dimensions are; no temporal encoding | Whatever the array holds | Any CRS the conventions can describe; static Zarr | You are publishing general-purpose geospatial arrays for analysis tools and want the community convention as it settles. chronozarr follows the same `proj` and `spatial` attributes and adds the temporal profile, band objects, `mask`, `coverage` and a browser reader on top. | ### What chronozarr costs * **Modest temporal gain.** Star-delta reduced compressed size by about 25% on arid scenes and about 6% on vegetated ones, which is why the writer keeps it only when a sample of cells compresses to 0.85 of the plain size or better. The encoding's measured effect is on size; the scrub-latency gains in the README come from worker-pool decoding and time-window prefetch (before and after table), which a plain store also gets. * **Large downloads for wide views.** Values are lossless and unquantized. A 9-cell overview of the 117-month Ucayali store is about 11.7 MB per timestep; a cold loop over it is bound by the link. * **One reader implementation.** The viewer is the only consumer that reconstructs star-delta and that uses the time-window prefetch. Any Zarr v3 client reads a `none` store; other tools see residuals in a star-delta store. * **GDAL support depends on the version.** GDAL 3.12.4 opens an unsharded store (the writer default) as a raster of time-major bands with the geotransform, and fails on a sharded one (`Unsupported codec: sharding_indexed`, tested here). The Zarr driver documents sharding support from GDAL 3.13 ([GDAL Zarr driver](https://gdal.org/en/stable/drivers/raster/zarr.html)); that version is not tested here. The writer emits the `_CRS` array attribute GDAL reads, so GDAL assigns the CRS of an EPSG store (in a 3.12.4 test an added `_CRS` was picked up while `proj:code` alone was not; GDAL's documentation calls `_CRS` the only CRS convention supported before 3.13). In a star-delta store GDAL sees residuals for non-anchor timesteps. `chronozarr export-cog` writes true-value Cloud Optimized GeoTIFFs from any store for any GDAL or QGIS version. * **Not a tile server.** Products (true color, NDVI, water) are band math in the client's fragment shader. Anything that needs server-side rendering, mosaicking of many sources or arbitrary expressions belongs with COG + TiTiler. ### What the others cost that chronozarr avoids * A server to run and scale (TiTiler). * Resampling every timestep to Web Mercator before storage (PMTiles, Mapbox raster-array, ndpyramid's reprojected pyramids), which biases block statistics and discards native pixels. * A fresh chunk fetch per timestep with no shared state between timesteps (stores with one timestep per chunk, the common zarr-layer layout); the time-window prefetch and the optional temporal encoding address that cost. * A dependency on one vendor's tiling service and renderer (Mapbox raster-array). ## Hosting a chronozarr store A chronozarr store is a directory of static files. Any host that returns a file by path, answers byte-range requests and sends CORS headers serves it. There is nothing to run. This page turns the requirements in [spec/CHRONOZARR.md](/specification) section 9 into a checklist, an upload procedure, and recipes for four hosts: S3 with CloudFront, Cloudflare R2, Google Cloud Storage, and Source Cooperative. ### 1. Checklist `chronozarr doctor ` runs the HTTP and decode checks below against a live URL (and the decode checks against a local directory). It sends `Origin: https://chronozarr.org` by default. The last column is what doctor reports when the requirement is missing (`fail` violates a MUST, `warn` a SHOULD, `info` is reported without judgement), read from `src/chronozarr/doctor.py` on 2026-10-01; if that file changes, it is the authority. The `edge cache` and `timing-allow-origin` lines are advice: doctor reports them as `info`, and only `fail` lines make it exit with status 1 (warnings and info do not). | # | Requirement | doctor check | If missing | | -- | ------------------------------------------------------------------------------------------------------------- | --------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | | 1 | `GET {store}/zarr.json` returns 200 and a chronozarr root group | `root zarr.json` | fail | | 2 | `Access-Control-Allow-Origin: *` (or the viewer origin) on GET responses | `CORS on zarr.json`, `CORS on byte range ...` | fail | | 3 | Consolidated metadata in the root `zarr.json` | `consolidated metadata` | warn | | 4 | A bounded `Range: bytes=a-b` returns `206` with `Content-Range` | `byte range on ` | fail if sharded; warn if unsharded and the server answers 200 | | 5 | `Access-Control-Expose-Headers` lists `Content-Range` (or `*`) | `CORS expose Content-Range` | fail | | 6 | A suffix range `Range: bytes=-N` returns `206` with `Content-Range` (shard index reads) | `suffix range` | fail if sharded; info if unsharded (an unsharded store has no shard index) | | 7 | `OPTIONS` preflight answers 200 or 204 with `Access-Control-Allow-Headers` containing `Range` (or `*`) | `CORS preflight` | warn if sharded; info if unsharded | | 8 | `HEAD` returns 200 with `Content-Length` and CORS headers | `HEAD` | fail if sharded without `shard_bytes`; warn otherwise | | 9 | `Cache-Control` contains `immutable`, or `max-age` of at least one day | `cache-control` | warn when the prefix looks versioned (last path segment ends in a version or date, such as `chronozarr-2`); info otherwise | | 10 | `Timing-Allow-Origin` is sent (MAY) | `timing-allow-origin` | info; reports the value or its absence | | 11 | The host caches objects at the edge (MAY) | `edge cache` | info; reports `cf-cache-status` and says when it is `DYNAMIC` or `BYPASS` | | 12 | The layout validates, every level decodes, and one pixel read through plain Zarr equals the chronozarr decode | `validate`, `decode level N` | fail | Doctor does not check these; verify them with `curl` (section 4): | Requirement | Level in the spec | Why it matters | | ------------------------------------------------------------------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | An absent key returns `404`, not a fallback page with `200` | MUST | A missing shard, `mask` or `coverage` array is meaningful: readers decode a missing shard as fill. A host that answers 200 with HTML breaks that. | | Chunk and shard responses carry no `Content-Encoding` | MUST | The shard index stores byte offsets into the stored object. A CDN that gzips or re-encodes the body moves the bytes. | | Shards fit the host or CDN's cacheable object size (sharded stores) | SHOULD | Choose `shard_time` at encode time (section 5). An unsharded store's objects are single chunks, about 2 MB. | | Metadata uploaded last; the prefix never rewritten | SHOULD / MUST | Section 2. | A preflight is sent only for a suffix range. Browsers do not preflight a bounded `bytes=a-b`. A store written with `shard_bytes` lets a reader fetch every shard index as a bounded range, so checks 6 and 7 cost nothing for it in practice, but doctor still reports them. An unsharded store, the encoder default, is read with plain `GET`s: no `Range` header at all, so for it check 4 is a warning at most, check 6 is always `info`, and check 7 is `info` instead of `warn` when it fails. Without `Timing-Allow-Origin`, `PerformanceResourceTiming.transferSize` is 0 for cross-origin requests, so in-page benchmarks under-report bytes; the viewer then counts bytes only from `Content-Length`. ### 2. Immutable prefixes and upload order **Immutable prefixes.** A store is never edited in place. Every encode is written under a new prefix (for example `ucayali_santa_maria/chronozarr-3`), the catalog or URL is switched to it, and the old prefix is deleted afterwards. This is what makes `Cache-Control: public, max-age=31536000, immutable` safe: browsers and CDNs may hold any object for a year, and no object at that URL ever changes. **Metadata last.** The root `zarr.json` is the marker that a store exists. Upload in three phases: 1. Chunk and shard objects (everything except `zarr.json`), in parallel. 2. Group and array `zarr.json` below the root. 3. The root `zarr.json`. A reader that finds the root therefore finds a complete store. Do not request the final URL before phase 3 finishes: CDNs cache a `404` for seconds to minutes, and a cached `404` on the root `zarr.json` makes a finished store look absent. The same three phases work for every host. Save this as `upload.sh`, run it with `bash`, and define `put` from the recipe of your host. `put FILE` uploads `./FILE` (relative to the store directory) to `$PREFIX/FILE` with `Cache-Control: $CC`. ```bash #!/usr/bin/env bash set -euo pipefail STORE=data/stores/ucayali_santa_maria/chronozarr-3 # local store directory PREFIX=ucayali_santa_maria/chronozarr-3 # new for every encode CC="public, max-age=31536000, immutable" # put() goes here, from the host recipe below export -f put export STORE PREFIX CC BUCKET # plus anything else put() reads cd "$STORE" # 1. chunk and shard objects find . -type f ! -name zarr.json -print0 | xargs -0 -n 1 -P 8 bash -c 'put "$1"' _ # 2. group and array metadata below the root find . -type f -name zarr.json ! -path ./zarr.json -print0 | xargs -0 -n 1 -P 8 bash -c 'put "$1"' _ # 3. root metadata, last put ./zarr.json ``` `set -e` stops before phase 2 or 3 if any upload in the previous phase failed (`xargs` exits 123). The repository's `scripts/upload_stores.sh` uploads to R2 in the same three phases and stops before the next phase when an upload fails. It reads `data/stores//$STORE` (`STORE=chronozarr-3` for a versioned prefix) and sets `Cache-Control` per object: `immutable` for every chunk and shard except the objects an append rewrites, which get `max-age=300` (`--trailing-ttl SECONDS` changes it; `--dry-run` prints the class of every object without uploading). Section 7 lists those objects. ### 3. Recipes #### 3.1 Amazon S3 with CloudFront Keep the bucket private and let CloudFront read it through an origin access control (OAC). CloudFront gives HTTPS on your own domain, caching, and the only way to add `Timing-Allow-Origin` in front of S3. ```bash BUCKET=my-chronozarr-stores REGION=us-east-1 aws s3api create-bucket --bucket "$BUCKET" --region "$REGION" # outside us-east-1 add --create-bucket-configuration LocationConstraint="$REGION" aws s3api put-public-access-block --bucket "$BUCKET" --public-access-block-configuration \ BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true ``` Create a CloudFront distribution (console or `aws cloudfront create-distribution`) with: * Origin: the bucket's REST endpoint `BUCKET.s3.REGION.amazonaws.com`, origin access: origin access control with signing on. * Default cache behavior: viewer protocol policy redirect HTTP to HTTPS; allowed methods `GET, HEAD, OPTIONS`; cache policy `CachingOptimized` (no headers or query strings in the cache key). * Response headers policy: the custom policy below. Bucket policy allowing only that distribution (fill in the account id and distribution id): ```json { "Version": "2012-10-17", "Statement": [{ "Sid": "AllowCloudFrontRead", "Effect": "Allow", "Principal": { "Service": "cloudfront.amazonaws.com" }, "Action": "s3:GetObject", "Resource": "arn:aws:s3:::my-chronozarr-stores/*", "Condition": { "StringEquals": { "AWS:SourceArn": "arn:aws:cloudfront:::distribution/" } } }] } ``` Response headers policy (`aws cloudfront create-response-headers-policy --response-headers-policy-config file://chronozarr-headers.json`). It adds the CORS headers with `OriginOverride`, so the bucket needs no CORS configuration, and `Timing-Allow-Origin`: ```json { "Name": "chronozarr-store", "CorsConfig": { "AccessControlAllowOrigins": { "Quantity": 1, "Items": ["*"] }, "AccessControlAllowHeaders": { "Quantity": 1, "Items": ["Range"] }, "AccessControlAllowMethods": { "Quantity": 2, "Items": ["GET", "HEAD"] }, "AccessControlAllowCredentials": false, "AccessControlExposeHeaders": { "Quantity": 4, "Items": ["Content-Range", "Content-Length", "ETag", "Accept-Ranges"] }, "AccessControlMaxAgeSec": 86400, "OriginOverride": true }, "CustomHeadersConfig": { "Quantity": 1, "Items": [{ "Header": "Timing-Allow-Origin", "Value": "*", "Override": true }] } } ``` Upload with the section 2 script, using: ```bash put() { aws s3 cp "$1" "s3://$BUCKET/$PREFIX/${1#./}" --cache-control "$CC" --only-show-errors; } ``` The store URL is `https:///$PREFIX`. Notes: * `Cache-Control` is stored on the objects at upload and passed through by CloudFront. Nothing needs invalidating because prefixes are immutable. * `aws s3 cp` labels extensionless objects `binary/octet-stream`. CloudFront compresses only a fixed list of content types, between 1,000 and 10,000,000 bytes, and `binary/octet-stream` is not on it, so shards are served as stored. Confirm with section 4. * Serving a public bucket directly (no CloudFront) also works for `Range` and `Cache-Control`. It needs an S3 CORS configuration instead of the response headers policy, and it cannot send `Timing-Allow-Origin`: ```json { "CORSRules": [{ "AllowedOrigins": ["*"], "AllowedMethods": ["GET", "HEAD"], "AllowedHeaders": ["Range"], "ExposeHeaders": ["Content-Range", "Content-Length", "ETag", "Accept-Ranges"], "MaxAgeSeconds": 86400 }] } ``` Apply it with `aws s3api put-bucket-cors --bucket "$BUCKET" --cors-configuration file://s3-cors.json`. S3 sends CORS headers only on requests that carry an `Origin` header. #### 3.2 Cloudflare R2 This is how `data.tileripper.com` is served (see `deploy/README.md`). ```bash BUCKET=my-chronozarr-stores npx wrangler r2 bucket create "$BUCKET" --location enam npx wrangler r2 bucket cors set "$BUCKET" --file deploy/r2-cors.json --force npx wrangler r2 bucket domain add "$BUCKET" --domain data.example.com --zone-id ``` `deploy/r2-cors.json` uses wrangler's rule format, not the S3 array: ```json { "rules": [{ "allowed": { "origins": ["*"], "methods": ["GET", "HEAD"], "headers": ["Range", "If-Match", "If-None-Match", "Content-Type"] }, "exposeHeaders": ["Content-Range", "Content-Length", "ETag", "Accept-Ranges"], "maxAgeSeconds": 86400 }] } ``` Upload with the section 2 script, using: ```bash put() { npx --no-install wrangler r2 object put "$BUCKET/$PREFIX/${1#./}" --file "$1" --cache-control "$CC" --remote > /dev/null; } ``` Wrangler starts in about 2 seconds per object. An unsharded store (the encoder default) is thousands of objects: 5,893 for the 117-month Ucayali imagery store and 17,088 for its water store, which at 4 parallel is about 50 minutes and 2.4 hours. A sharded store of around 100 objects takes a minute or two. For the first upload of an unsharded store use an S3 client against R2's S3 endpoint instead; each later append uploads only the objects it wrote (section 7). Two Cloudflare rules on the hostname make the domain cache and time correctly. Both are created in the dashboard under Rules and need zone write access; the `wrangler login` token used for deploys does not have it. * **Cache Rule.** Cloudflare does not cache extensionless objects by default, so responses stay `cf-cache-status: DYNAMIC`. Create a Cache Rule: * Expression: `(http.host eq "data.example.com")`. In the builder: Field `Hostname`, Operator `equals`. * Cache eligibility: Eligible for cache. * Edge TTL: Ignore cache-control header and use this TTL, 1 year. * Browser TTL: Override origin and use this TTL, 1 year. Verified on `data.tileripper.com` on 2026-09-30: `zarr.json` returned MISS then HIT with an `age` header, and a `206` range request on a shard returned MISS then HIT. Overriding both TTLs is safe only because prefixes are immutable (section 2). Setting `Cache-Control` on each object at upload still matters for clients of the bare bucket and for any rule that respects the origin. * **`Timing-Allow-Origin`.** Add a Transform Rule, Modify Response Header, with the same hostname expression and the action Set static: `Timing-Allow-Origin` = `*`. * **Match on the hostname.** Put the expression in the expression editor, or use the `Hostname` field. Pasting it into a URI Full wildcard value silently matches nothing and every response stays `DYNAMIC`. The dashboard's warning that the rule "may not apply to your traffic" for the R2 hostname is a false alarm; ignore it. * **Object size and miss cost.** Cloudflare documents a 512 MB cacheable object limit on the Free, Pro and Business plans; check the current figure. A shard is about `n_time` times the compressed size of one cell-timestep (the sharded Ucayali store has 117 timesteps and shards up to 174 MB). A miss on a range read of an object that is not yet cached pulled the whole object from R2: the 1,876-byte shard-index read at the end of an 83 to 174 MB shard took 2 to 11 s, and 0.36 s on a 41 MB shard; nine of them were 6.5 of a 6.8 s cold open (`docs/comparisons.md`). That cost grows with shard size, and is why the encoder default is unsharded, where a miss costs one chunk of about 2 MB (the largest object of the unsharded Ucayali store is 1.8 MB). If you shard, a smaller `shard_time` shrinks the miss. * **`r2.dev`.** The public development URL (`wrangler r2 bucket dev-url enable`) is rate limited and not for production. Use it for a first check only. #### 3.3 Google Cloud Storage ```bash BUCKET=my-chronozarr-stores gcloud storage buckets create "gs://$BUCKET" --location=US --uniform-bucket-level-access gcloud storage buckets add-iam-policy-binding "gs://$BUCKET" --member=allUsers --role=roles/storage.objectViewer gcloud storage buckets update "gs://$BUCKET" --cors-file=gcs-cors.json ``` `gcs-cors.json`. GCS takes one `responseHeader` list and uses it for both `Access-Control-Allow-Headers` and `Access-Control-Expose-Headers`, so `Content-Range` goes in it alongside `Range`: ```json [{ "origin": ["*"], "method": ["GET", "HEAD"], "responseHeader": ["Content-Type", "Range", "Content-Range", "Content-Length", "ETag", "Accept-Ranges"], "maxAgeSeconds": 86400 }] ``` Upload with the section 2 script, using: ```bash put() { gcloud storage cp "$1" "gs://$BUCKET/$PREFIX/${1#./}" --cache-control="$CC" --quiet; } ``` The store URL is `https://storage.googleapis.com/$BUCKET/$PREFIX`. Notes: * CORS is honored on `storage.googleapis.com`, not on the cookie-authenticated `storage.cloud.google.com`. * A public object with no `Cache-Control` is served with `public, max-age=3600`; the upload command above overrides it. * `allUsers` is rejected when the organization enforces public access prevention. Either change that policy for this bucket or front the bucket with a backend bucket behind Cloud CDN. * `Timing-Allow-Origin` cannot be set on a bucket or object. It is available as a custom response header on a backend bucket behind an external HTTPS load balancer (`gcloud compute backend-buckets create ... --custom-response-header='Timing-Allow-Origin: *'`). It is a MAY; skip it unless you benchmark the viewer. * Do not upload objects with `Content-Encoding: gzip`; GCS applies transcoding that changes the bytes it serves. #### 3.4 Source Cooperative [Source Cooperative](https://source.coop) (Radiant Earth) hosts open datasets and serves them through a data proxy at `https://data.source.coop`. Objects are addressed as `https://data.source.coop///`. The proxy is documented as beta, and the storage behind it (S3, GCS, Azure, R2) is not your choice, so you do not configure CORS or cache headers. Observed on one public object on 2026-09-30 (`kerner-lab/fields-of-the-world`): a bounded range returns `206` with `Content-Range` and `Accept-Ranges: bytes`; `Access-Control-Allow-Origin: *`, `Access-Control-Allow-Headers: *`, `Access-Control-Expose-Headers: *`; `OPTIONS` returns `204`; there is no `Cache-Control` and no `Timing-Allow-Origin`; `cf-cache-status: DYNAMIC`. Expect `chronozarr doctor` to pass checks 1 to 8 and 12, warn on `cache-control` for a versioned prefix, and report `edge cache` and `timing-allow-origin` as info. Upload needs an account and upload access for the product (contact [hello@source.coop](mailto\:hello@source.coop) if the product page shows no upload option). Two routes: * **UI.** Product page, lock icon, Edit Mode, drag in files or directories. Uploads in the UI are not ordered: add the level directories (`0/`, `1/`, ...) and `volatility/` first, then the root `zarr.json` by itself. A sharded store (`--shard`) has roughly a hundred objects, which is workable here; an unsharded store, the default, has thousands and is not. * **S3 client through the proxy**, with temporary credentials from the Source CLI (`source-coop login`, then an AWS profile named `source-coop` as described in the Source docs). Use the section 2 script with: ```bash ACCOUNT=my-account PRODUCT=my-product put() { aws s3 cp "$1" "s3://$ACCOUNT/$PRODUCT/$PREFIX/${1#./}" \ --endpoint-url https://data.source.coop --profile source-coop \ --cache-control "$CC" --only-show-errors } ``` Whether the proxy stores and serves the `--cache-control` value is untested here; check the `cache-control` line of doctor. The store URL is `https://data.source.coop/$ACCOUNT/$PRODUCT/$PREFIX`. ### 4. Verifying by hand ```bash URL=https://data.example.com/aoi/chronozarr-2 KEY=0/data/c/0/0/0/0 # level 0, cell (0, 0): timestep 0 (unsharded) or time shard 0 (sharded) O='Origin: https://chronozarr.org' # 206 + Content-Range + CORS + exposed headers + caching curl -s -D - -o /dev/null -H "$O" -H 'Range: bytes=0-99' "$URL/$KEY" # suffix range (shard index; only sharded stores need it) curl -s -D - -o /dev/null -H "$O" -H 'Range: bytes=-16' "$URL/$KEY" # preflight curl -s -D - -o /dev/null -X OPTIONS -H "$O" -H 'Access-Control-Request-Method: GET' \ -H 'Access-Control-Request-Headers: range' "$URL/$KEY" # no Content-Encoding on a chunk or shard, even when the client offers compression curl -s -D - -o /dev/null -H "$O" -H 'Accept-Encoding: gzip, br' -H 'Range: bytes=0-99' "$URL/$KEY" | grep -i '^content-encoding' || echo "no content-encoding: ok" # absent keys are 404 curl -s -o /dev/null -w '%{http_code}\n' "$URL/no-such-key" ``` The key is the same in both layouts: for an unsharded store it is the chunk of timestep 0, band 0, cell (0, 0), for a sharded store the first shard. An all-fill chunk is not written, so a store that is empty at that cell answers `404` there; pick another cell. The expected answers are `206` with `Content-Range: bytes 0-99/`, `Access-Control-Allow-Origin: *`, an `Access-Control-Expose-Headers` that lists `Content-Range`, `204` or `200` for the preflight with `Range` allowed, and `404` for the missing key. Then run `chronozarr doctor "$URL"`. ### 5. Pitfalls * **Compression in front of the store.** A CDN, proxy or bucket setting that adds `Content-Encoding` to chunk or shard objects breaks reads (a chunk is decoded by the codec chain, not by the HTTP layer; for shards the index offsets are in stored bytes and a `Range` header applies to the encoded representation). Turn compression off for the store path, or rely on content types the host does not compress. * **Cloudflare rules that match nothing.** A Cache Rule or Transform Rule whose expression was pasted into a URI wildcard value instead of matching `Hostname` applies to no request, and R2 responses stay `DYNAMIC` with no `Timing-Allow-Origin` (section 3.2). * **Fallback pages.** Static hosts configured for single-page apps answer unknown paths with `200` and `index.html`. The store then appears to have every shard. Serve the store from a host or prefix with real `404` behaviour. * **Rewriting a prefix.** With `immutable` and a one-year `max-age`, a rewritten object stays stale in browsers and CDNs for up to a year. Always re-encode to a new prefix. The only in-place change is an append, and only the objects of section 7 change. * **Local testing.** `python -m http.server` ignores `Range`, so sharded stores read whole shards (doctor reports it as a failure); an unsharded store reads fine from it. Use `chronozarr.view(store)` from a notebook or any range-capable static server. * **Shard size.** Stores are unsharded unless you pass `--shard`. With it, `shard_time` defaults to `n_time`, so one shard holds a cell's whole time axis. Lower it (and keep it a multiple of `anchor_interval` for star-delta stores) when a shard would exceed what the host or CDN caches or serves in one object, or when a CDN miss on a shard is too slow (section 3.2). A store that will grow should stay unsharded (section 7). ### 6. The live store The published demo store is `ucayali_santa_maria/chronozarr-3` (plain encoding, sharded, 93 objects; see the README status). `chronozarr-4` is the same data written with the unsharded default (5,893 objects) and replaces it once uploaded. It replaces `chronozarr-2`, which was live on 2026-09-30 when `chronozarr doctor` was run against `https://data.tileripper.com/ucayali_santa_maria/chronozarr-4`: 15 ok, 2 info, 0 warnings (`edge cache` HIT, `timing-allow-origin` `*`, `cache-control` `max-age=31536000`). The host is R2 with `deploy/r2-cors.json` plus the Cache Rule and the Transform Rule of section 3.2; both rules match the hostname, so they apply to every prefix under it. ### 7. Appending to a live store `chronozarr append STORE INPUT` adds timesteps to the end of a store in place (spec section 14). It writes the shards or chunks that gain data and the metadata; every other object keeps its bytes and its cache entry. **The default layout is the right one for appends.** An unsharded store (the encoder default) appends by writing only new chunk objects, 55 MB per month on the Ucayali mosaics, with nothing rewritten but the metadata. It has no shard index for an open viewer to hold stale, and the same chunk reads per timestep with no index read. A sharded store (`--shard`) rewrites the shard that receives each new timestep, whole: `(shard_time + 1) / 2` chunks per cell per append on average, 4389 MB against 686 MB over a 12-month cycle at `shard_time` 12. A whole-axis shard (one shard for the initial axis) opens a second shard as long as the first and rewrites it on every append, up to 117 chunks per cell for a 117-month store. The price of unsharded is object count: about 5,900 objects for 117 months against 93 sharded. Choose a finite `--shard-time` (12 for monthly data, a multiple of `--anchor-interval` for star-delta) only when object count matters more than the rewrite cost, and the whole-axis `--shard-time` for archives that are not appended to. The first batch of a sharded store may be a single timestep at any `--shard-time`. An unsharded store of thousands of objects is slow to upload for the first time (section 3.2 says how), but an append uploads only the objects it wrote, 63 in the measurement. Measured costs: `docs/append.md`. **Procedure.** ```bash chronozarr convert new_month.csv work/new_month # one new timestep, or encode/convert several cp -R data/stores/aoi/chronozarr-3 work/aoi-working # append is not atomic: work on a copy touch work/stamp chronozarr append work/aoi-working work/new_month chronozarr validate work/aoi-working # publish: shards and chunks first, root zarr.json last; only objects the append touched STORE=chronozarr-3 ROOT=work scripts/upload_stores.sh --newer-than work/stamp --trailing-ttl 300 aoi-working ``` (`scripts/upload_stores.sh` reads `$ROOT//$STORE`; put the working copy at `work//chronozarr-3` for the layout it expects.) **Cache classes.** Only these objects change in place, so only these need a short lifetime: | Object | Why it changes | Cache-Control | | ------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | ------------------------------------- | | root `zarr.json` and every other `zarr.json` | new `times`, `temporal`, shapes, `shard_bytes`, consolidated metadata | `public, max-age=300` | | `{level}/time/c/0` | the time axis is one chunk | `public, max-age=300` | | `volatility/c/0/0` | new deltas join each cell's mean | `public, max-age=300` | | for each cell, level and sharded variable (`data`, `mask`, `coverage`), the shard with the highest time index | it gains the new chunks | `public, max-age=300` | | every other shard; every chunk of an unsharded store (new timesteps are new keys) | never rewritten | `public, max-age=31536000, immutable` | The trailing shard of a cell is the one holding its last timestep. When the next shard opens, the old one becomes immutable (nothing rewrites it again) but keeps the `max-age=300` it was uploaded with, because `--newer-than` does not touch it. That costs an origin revalidation every five minutes per shard; upload those shards once without `--newer-than` if that matters. The script classifies by the current state of the whole store, so a run without `--newer-than` sets every object's header correctly. **Cloudflare.** The Cache Rule in section 3.2 sets Edge TTL to *Ignore cache-control header and use this TTL, 1 year* and Browser TTL to a 1-year override, on every URL of the host. Behind that rule the live `data.tileripper.com` serves a rewritten `zarr.json` or trailing shard from the edge for a year, whatever the origin sent, so appends are invisible there. A provider needs a rule that respects the headers the upload script sets: change Edge TTL to *Use cache-control header if present, use default TTL if not* (default 1 year) and Browser TTL to *Respect origin TTL*, or keep the override rule for immutable paths and add a rule above it that matches the mutable objects. Trailing shards cannot be matched by path pattern (their time index grows), so the header-based rule is the simple one. Purging the changed URLs after the upload also works; it does not replace the TTL, because browsers keep their own copy. **Open viewers.** A viewer keeps the root `zarr.json` it loaded until the page is reloaded, so it keeps showing the old timesteps and does not see the new ones. Everything it already reads stays correct: the chunks, offsets and references of existing timesteps do not change. The one failure, for a sharded store, is a shard index it fetches after an append using the old shard length from `shard_bytes`: the trailing shard is longer, the range lands on the wrong bytes and the index checksum fails. Reloading fixes it. An unsharded store has no shard index, so a stale viewer has no such failure. The same mismatch can occur for up to the short lifetime when a CDN holds an old shard object next to a new `zarr.json` (or the reverse), which is why the upload order is shards, then metadata, and why both lifetimes are the same. **What `doctor` says.** `chronozarr doctor` on an appended store passes the same checks. Its `cache-control` line reads the header of one object, `0/data/c/0/0/0/0` (cell (0, 0)). In an unsharded store that is the chunk of timestep 0, immutable for good. In a sharded store it is time shard 0, immutable once a second time shard exists; while the store still has one time shard it is the trailing shard with `max-age=300`, and doctor warns on a versioned prefix; that warning is expected then. Check `zarr.json` and, for a sharded store, a trailing shard with `curl -I` (section 4). ## PNG frames to a store A sequence of georeferenced PNGs, one per date, becomes a scrubbable chronozarr store with `chronozarr convert`. There is no GeoTIFF step. The frames are what exporters hand over: Earth Engine thumbnails, QGIS "Save as image" (with its world file option), matplotlib figures, drone and orthomosaic pipelines. `convert` reads them with GDAL, so everything it does for COGs applies: manifest, time order, `--dry-run`, `--resume`, validity rules, the pyramid. The store holds exactly the values in the PNGs. A rendered image is display values, not measurements; see "Limits" before using one for anything but looking. ### Locating the frames A PNG carries no georeferencing of its own. GDAL finds it in a sidecar file next to the frame, and `convert` needs one of three things: | Frame has | Holds | `convert` also needs | | ------------------------------------ | ---------------------------- | ------------------------------------------------------- | | `2019-01.pgw` (or `.wld`) world file | the transform, no CRS | `--crs EPSG:xxxxx` | | `2019-01.png.aux.xml` | the CRS and the geotransform | nothing | | neither | nothing | `--crs EPSG:xxxxx` and `--bounds west,south,east,north` | ```bash # a world file per frame: the transform is there, the CRS is not uv run chronozarr convert frames/manifest.csv out --crs EPSG:32718 # an .aux.xml per frame (GDAL writes one when it copies a georeferenced raster to PNG) uv run chronozarr convert frames/manifest.csv out # no sidecar: say where the frames are uv run chronozarr convert frames/manifest.csv out --crs EPSG:32718 \ --bounds 498450,9151960,508690,9162200 ``` The manifest is the usual one: `uri,datetime` rows in CSV (or JSON), relative URIs resolved against the manifest's folder. Details: * With both sidecars on one frame, GDAL takes the transform from the world file and the CRS from the `.aux.xml`. * A world file gives the *centre* of the upper-left pixel. GDAL accounts for that; nothing for you to do. `--bounds` is the other way: it takes the outer edges of the image. * `--crs` has two jobs, as for COGs: it is the target CRS, and it is the CRS of any frame that declares none. If the frames carry a CRS and `--crs` names a different one, the frames are warped (`--resampling` required). * The sidecars of `http(s)` PNGs are found too, at the cost of a few extra requests per frame. * A frame with a geotransform but no CRS and no `--crs` fails and names the first such frame. A frame with no georeferencing at all fails and lists the three fixes above. #### `--bounds` `--bounds west,south,east,north` (units of `--crs`: metres for UTM, degrees for EPSG:4326) is the extent of every frame, edge to edge. `convert` derives the north-up transform from each frame's pixel size: `(east - west) / width` by `(north - south) / height`, so pixels need not be square. In a JSON manifest the same extent can sit at the top level: ```json {"bounds": [498450, 9151960, 508690, 9162200], "items": [{"uri": "2019-01.png", "datetime": "2019-01-01"}]} ``` Give it once, not in both places. It is refused when: * `--crs` is missing; * west is not less than east, or south not less than north; * a frame has a different size from the first (the same extent over a different pixel count would be a different pixel size, which would be a guess); * a frame carries its own geotransform (a world file or `.aux.xml`): drop `--bounds` or remove the sidecar; * a frame declares a CRS that is not `--crs`. Check that the extent you pass is the extent of the rendered image. An exporter that pads or crops the region you asked for has moved every pixel by that amount, and nothing in the PNG can tell `convert`. ### What the frames become | In the PNG | In the store | | ---------------------------------- | ----------------------------------------------------------------------------------- | | red, green, blue channels | bands `red`, `green`, `blue` with those `common_name`s, scale 1, offset 0, no units | | alpha channel (RGBA, gray + alpha) | the `mask` variable, 1 where alpha is nonzero; not a data band | | gray channel | band `1` | | 8-bit values | `uint8` data; the viewer shows them as they are | | palette (indexed colour) | refused, see below | The viewer picks products by band name, so the three colour bands enable True color. False color, NDVI, NDWI and Water need a near-infrared band and stay off; Single band is available. Alpha rule. Alpha wins over everything else, as in GDAL: a pixel is valid where alpha is nonzero. A partly transparent pixel (alpha 128) is valid and is drawn fully opaque, because the mask is 0 or 1. The RGB values under alpha 0 are kept in the store, hidden by the mask. Without an alpha channel the store has no mask and no nodata: every pixel is valid, a stored 0 is data. Palette rule. A palette PNG fails with the way to expand it: ``` cannot open source .../frame.png: it is a palette (indexed colour) PNG, so its values are palette indices, not colours. Expand it to RGB first, for example `gdal_translate -expand rgba in.png out.png`, or save the frames as RGB ``` Expanding silently was the alternative. It was not chosen because the indices of a classified map are the data, and the converter cannot tell a class map from a picture. This applies to PNG only: a GeoTIFF with a colour table is converted with its indices as the values, as before. Band names can still be given in the manifest (`bands` column or key); they replace `red`, `green`, `blue` and drop the common names. ### Worked example: Ucayali, 36 months `examples/png_frames/` turns the Sentinel-2 monthly mosaics already in `data/mosaics/` into frames, the way an exporter would, then converts them: ```bash uv run python examples/png_frames/render_frames.py # 36 PNGs + world files + manifest examples/png_frames/convert.sh # convert --crs EPSG:32718, then validate uv run python examples/png_frames/check_store.py # every frame bit-exact in the store cd js && node ../examples/png_frames/viewer_check.mjs # headless viewer, repo served on :8000 ``` The convert step is one command: ```bash uv run chronozarr convert data/png_frames/ucayali/manifest.csv \ data/stores/ucayali_santa_maria/png-1 --crs EPSG:32718 ``` What the frames are: * 2019-01 to 2021-12, a 1024 x 1024 pixel window (row 768, column 1280) of the 2765 x 2759 mosaic, 10 m pixels, EPSG:32718, 8-bit RGBA: red = B04, green = B03, blue = B02. * Alpha 0 where the mosaic's `coverage` is 0 (no valid scene that month), with RGB set to 0 there. The masked share of a frame runs from 0 to 76.8 %; 6 of the 36 frames are over 3 % masked. * One fixed linear stretch for all frames: DN 185 to 1758 mapped to 0 to 255, the 2nd and 98th percentile of the valid red, green and blue values of the whole series. It is not the viewer's tone mapping, which works on reflectance. Because it is the same for every month, brightness compares across months. Sizes and times (M3 Max laptop, other work running): | | | | -------------------------- | --------------------------------------------------------------------------------------------- | | Frames on disk | 74.7 MB for 36 PNGs (2.08 MB each; raw RGBA is 4.19 MB), GDAL's default PNG compression | | Raw in the store | 113.2 MB bands + 37.7 MB mask | | Store | 112.9 MB in 35 files: level 0 89.6 MB, level 1 23.2 MB (the masks, 0.23 MB, are inside those) | | `convert` | read 0.5 s, encode 1.0 s, 1.5 s in all | | Temporal encoding (`auto`) | `none`: star-delta / plain = 1.05 on the sampled cells, and it must be 0.85 or less | The store is 1.5 times the size of the PNGs: level 1 is a quarter of level 0, and PNG's row filters compress rendered imagery better than zstd on the raw bytes (level 0 alone is 1.2 times the PNGs). What the store adds is random access and a pyramid, not compression. Checks that were run on `png-1`: * `chronozarr validate`: conforms to chronozarr 0.2.0. `chronozarr doctor` on the local path: 3 ok, 0 info, 0 warning, 0 failure. * `check_store.py`: for all 36 frames the store's red, green, blue and mask equal the PNG's red, green, blue and alpha != 0, bit for bit. * The same frames without world files, converted with `--crs EPSG:32718 --bounds 498450,9151960,508690,9162200`, and with an `.aux.xml` per frame and no `--crs`, give stores identical to `png-1` in data, mask and transform. * Headless viewer (software WebGL): True color is the active product and the only colour product enabled (False color, NDVI, NDWI and Water are disabled); a masked pixel is drawn as the background colour (9, 12, 18) and the valid pixel six pixels to its right is drawn as the stored colour; a click reads the stored `red`, `green`, `blue` values (200, 238, 184 at pixel 154, 43 of 2019-10, which is what the PNG holds there); playback buffers, steps through six timesteps and stops. No console errors. Screenshots: `data/reports/viewer_ucayali_png_truecolor.png` (2019-10, 38 % masked) and `data/reports/viewer_ucayali_png_edge.png` (the masked edge at 8x with the inspector open). ### Limits * Display values. A rendered PNG is not reflectance: the store has no units and scale 1, so the index products stay off and no value in it is a measurement. If the data behind the picture matters, convert the data (COGs, Zarr), not its rendering. * Eight bits. Each channel has 256 levels. The exporter's stretch is baked in: whatever it clipped or compressed is gone, and the viewer does not stretch 8-bit colour, it shows it as it is. * Pyramid levels are block means of the stored values, an average of display values, which is fine for browsing and wrong for radiometry. * The mask is binary. Soft alpha (anti-aliased edges, a feathered mosaic seam) becomes opaque wherever alpha is above 0. * One grid for the series. `--bounds` needs every frame the same size. Frames with different world files or sizes are warped onto the first frame's grid with `--resampling`, as COGs are, and warped pixels with no source are masked. * PNG is read whole, top to bottom, with no tiles or overviews. A 1024 x 1024 frame reads in about 0.03 s; memory during staging is a few frames, so a very large PNG costs its raw size times `1 + --read-ahead`. * Sidecar probing is on for `.png` URIs only. Other image formats (JPEG with `.jgw`) are not covered, and 16-bit and 1 to 4-bit PNGs are not tested; only 8-bit gray, gray + alpha, RGB and RGBA are. * World files and `.aux.xml` next to `http(s)` PNGs were checked against a local range server, not a CDN. ## User zero: monthly water-mask stacks Two derived chronozarr stores, built from the Sentinel-2 monthly mosaics that already sit in `data/mosaics/`, to see what a fluvial geomorphologist can do with the viewer: scrub months, click a pixel, chart it. Nothing was downloaded. | | Ucayali (Santa María reach, Peru) | Lake Mead (NV/AZ) | | -------------------------- | ----------------------------------------- | ------------------------------------------------------ | | Store | `data/stores/ucayali_santa_maria/water-1` | `data/stores/lake_mead/water-1` | | Months in the store | 117 (2015-11 to 2026-02) | 94 (2015-08 to 2023-06) | | Mosaics skipped | none | 33 (2023-07 to 2026-03, see "Problems in the mosaics") | | Grid | 2765 x 2759 px, 10 m, EPSG:32718 | 2827 x 2316 px, 10 m, EPSG:32611 | | Size | 1730.8 MB, 201 files | 1260.6 MB, 183 files | | Raw int16 bands | 3570 MB (2.06x) | 2462 MB (1.95x) | | Levels (MB) | L0 1286.7, L1 332.8, L2 87.6, L3 23.7 | L0 933.0, L1 246.6, L2 64.0, L3 16.9 | | Derive + encode | 62.9 s (69 s wall with the scan) | 33.2 s (38 s wall) | | Temporal encoding (`auto`) | `none` | `none` | | Threshold floor | -0.15 | 0.0 | | Floor binding | 116 of 117 months | 9 of 94 months | `auto` did not measure anything: for `int16` data the writer always uses `none` (spec 4.3), so there is no star-delta ratio to report. Codec zstd level 5, 4 pyramid levels. Both `water-1` stores are sharded, (T, 2, 512, 512), the encoder default when they were built. `ucayali_santa_maria/water-2` is the Ucayali store rebuilt unsharded, the default now: 17,088 files, 1730.5 MB, and every value, mask plane and coverage plane identical to `water-1` at every level, timestep and cell. Lake Mead has no `water-2`. Both stores together are 2991 MB. With about 3.5 GB left in R2 they fit and leave about 0.5 GB. Publish Ucayali first (1.73 GB; it is the reach the project is for), Lake Mead second. Nothing has been uploaded; the script takes the store name from `STORE`: ```bash STORE=water-1 scripts/upload_stores.sh --dry-run ucayali_santa_maria # 201 objects STORE=water-1 scripts/upload_stores.sh ucayali_santa_maria ``` ### How they were built ```bash uv run python examples/water_masks/build_water_stack.py --aoi ucayali_santa_maria --floor -0.15 uv run python examples/water_masks/build_water_stack.py --aoi lake_mead --boa-offset-from 2022-02 ``` Those commands wrote `water-1` sharded because that was the encoder default. The default is now unsharded, so add `--shard` to reproduce `water-1`, or `--store-name water-2` for the unsharded layout. Per month, from `B03` (green) and `B08` (nir) of the monthly median mosaic: * `ndwi` = (green - nir) / (green + nir), computed in float64, stored x10000 as `int16` (`scale` 1e-4, `units` "index"). Masked pixels hold 0. * `water` = 1 where `ndwi` is above the month's threshold, else 0, stored as a fraction: 0 or 10000 (`int16`, `scale` 1e-4, `units` "fraction"), so the physical value is 0 or 1 at level 0 and the water fraction of the block at coarser levels (see "Categorical bands"). Masked pixels hold 0. * `mask` = 0 where the mosaic's `coverage` is 0, that is where no scene was valid that month and the mosaic holds a carry-forward copy of an earlier month (or nothing). Gaps are not filled (`provenance.gap_fill` is "none"). * `coverage` = 1 where at least one scene was valid, else 0. The mosaics keep the valid fraction of scenes (multiples of 1/n), not a count, and the number of scenes is not stored, so this is the same 0/1 flag as in the `sentinel2_pc` example. It equals `mask` here; the viewer shows it as "Observed by 1 scene". * Threshold = max(Otsu, floor). Otsu is computed in numpy on the histogram of the month's valid `ndwi` values, one bin per stored integer, so the threshold is a stored value and the check below can reproduce it exactly. Ties inside an empty gap between two modes take the middle. * Dark pixels: a valid pixel with green and nir both at or below 5 DN is water and is left out of the Otsu histogram. L2A clips reflectance at DN 1, and in many pre-2022 winter months the whole of Lake Mead is DN 1 in every band (2021-11, interior of the lake: 1, 1, 1, 1; 2021-09: 216, 240, 71, 25). `ndwi` of two floor values is 0 or noise, so NDWI cannot see that water. Without the rule Lake Mead water fell to 1.7 % of valid pixels in 2021-11 (15.1 % with it), 3.9 % to 16.7 % in 2015-12, 1.5 % to 16.4 % in 2016-02, 2.6 % to 16.9 % in 2020-01. 24 of 94 Lake Mead months have 1 % or more dark pixels. Ucayali has none. The stored `ndwi` at a dark pixel stays what the clipped DNs give, so at level 0 `water == 10000 * (ndwi > threshold)` holds everywhere except at dark pixels. * `--boa-offset-from 2022-02` (Lake Mead only): subtracts 1000 DN from every valid pixel of the months from 2022-02 on (clipped to 1), the Sentinel-2 processing-baseline 04.00 offset the Ucayali mosaics were corrected for at ingest and the Lake Mead mosaics were not (see below). * Months with no valid pixel are not written (Lake Mead, 33 of them). Months with few valid pixels are written; their row in the CSV says how few. #### Per-month record `data/stores//water-1.months.csv`, one row per mosaic, in store order: | Column | Meaning | | ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | | `valid_px`, `valid_frac` | pixels with `mask` 1, and their share of the grid | | `dark_frac` | share of valid pixels at the DN floor, counted as water | | `otsu_threshold`, `threshold`, `floored` | raw Otsu, applied threshold (NDWI units), 1 when the floor replaced Otsu | | `eta` | Otsu separability: between-class variance over total variance (0.5 to 0.97; low = the split cuts through one mode) | | `water_frac`, `water_frac_otsu` | water share of valid pixels at the applied threshold and at the raw Otsu threshold (floor off) | | `water_km2` | water pixels x 100 m² | | `blue_median_dn` | median B02 DN of valid pixels, a haze screen | `blue_median_dn` separates hazy months: clear Ucayali months sit at 190 to 500 DN, hazy ones at 550 to 2300; Lake Mead desert is 535 to 1500 and the screen does not apply there. #### Why the Ucayali floor is -0.15, not 0 Otsu on the Ucayali never finds water: the median threshold is -0.29 and it is below the floor in 116 of 117 months, because the AOI is dense forest (NDWI about -0.75) and Otsu separates forest from everything else (bare bars, pasture, wet soil), labelling about 14 % of valid pixels as water against 9.6 % with the floor. So the floor is the threshold on this reach, and its value matters. At 0.0 (the usual NDWI convention) the main channel drops out of turbid months: sediment-laden water has green close to nir. 2020-04 gave 4.4 % water against 8.1 % the month before and 8.6 % the month after while the quicklook shows the main stem at NDWI near 0. A sweep of the floor over the 83 Ucayali months with at least 90 % valid pixels (hazy ones included; 65 pairs of adjacent months) shows month-to-month noise in the water fraction at its minimum near -0.15 to -0.2: | NDWI floor | -0.30 | -0.20 | -0.15 | -0.10 | -0.05 | 0.00 | | -------------------------------------------- | ------ | ------ | ------ | ------ | ------ | ------ | | Mean water share of valid px | 10.5 % | 9.6 % | 9.2 % | 8.6 % | 8.0 % | 7.2 % | | Std / mean over months | 0.14 | 0.11 | 0.11 | 0.13 | 0.14 | 0.19 | | 90th percentile of the adjacent-month change | 1.8 pt | 1.2 pt | 1.1 pt | 1.3 pt | 1.7 pt | 2.1 pt | The sweep measures stability, not accuracy: a lower floor also takes in wet sand and mixed bank pixels (+28 % area from 0.0 to -0.15), and nothing here is ground truth. -0.15 is a choice that removes the turbidity artifact; compare with the SWIR ensemble before using areas as numbers. The Lake Mead default of 0.0 would be wrong for that desert land (land NDWI is about -0.2 there, so a floor of -0.15 would label bare ground as water in a month without water); its Otsu is bimodal (median eta 0.92) and the floor rarely binds. ### Categorical bands Block-mean pyramids are right for continuous data and for fractions, and a 0/1 band must be stored as a scaled fraction (0 and 10000 with scale 1e-4, units "fraction") or its coarse levels undercount, because the integer block mean floors to 0 unless every pixel of the block is water. Stored as 0/1, the Ucayali 2019-03 water share of valid pixels fell from 12.1 % at level 0 to 8.5 % at level 3; as a fraction the valid-pixel-weighted share stays within 0.04 points of level 0. ### Checks ```bash uv run chronozarr validate data/stores/ucayali_santa_maria/water-1 uv run chronozarr info data/stores/ucayali_santa_maria/water-1 uv run chronozarr doctor data/stores/ucayali_santa_maria/water-1 uv run python examples/water_masks/check_water_stack.py --aoi ucayali_santa_maria --month 2019-03 uv run python examples/water_masks/check_water_stack.py --aoi lake_mead --boa-offset-from 2022-02 --month 2021-11 ``` * `validate`: both stores conform to chronozarr 0.2.0. `doctor`: 5 ok, 0 info, 0 warnings, 0 failures on each (one cell per level decoded and compared with a plain Zarr read). * `xr.open_dataset(path, engine="chronozarr")`: variables `data` (time, band, y, x), `mask` and `coverage` (time, y, x, uint8). `mask` is a data variable here; it is a coordinate on `open_store(path).to_xarray()`. By default values are float32 physical values with NaN where the mask is 0 (NDWI in -1..1; checked: -0.92..0.99 and -1.00..1.00 over a month, NaN exactly where `mask` is 0). With `physical=False` the stored `int16` values appear. Scale, offset and units are not attributes of the xarray variable (spec 3.6: no CF `scale_factor`); they are in the store attributes, `open_store(path).attrs.bands` (ndwi 1e-4, 0, "index"; water 1e-4, 0, "fraction"), and are applied by `physical=True`, so `water` reads 0.0 or 1.0 at level 0 and a fraction at `lod` 1 to 3. * Pixels against a direct numpy computation from the npz (month 2019-03 for Ucayali, 2021-11 for Lake Mead, using the month's threshold from the CSV): a water, a land and a masked pixel agree in `water`, stored `ndwi` and `mask`, for example Ucayali (1415, 1471): water 10000, ndwi 760, mask 1; Lake Mead (1764, 1304), a dark-water pixel: water 10000, ndwi 0, mask 1. The whole month agrees at every pixel (7.6 and 6.5 million), masked pixels hold 0. * Level 1 `water` equals the floor mean of level 0 over valid pixels at every block (1.9 and 1.6 million blocks), and fractions do occur: 40,238 blocks of Ucayali 2019-03 hold a value strictly between 0 and 1 at level 1, and 7,452 of Lake Mead 2017-03. Water share of valid pixels by pyramid level, from the stored `water` at each `lod`. "Weighted" weights each block by its valid level-0 pixels; "valid blocks" is the plain mean over the level's valid blocks, which is what a click on an overview shows on average: | | Level 0 | Level 1 | Level 2 | Level 3 | | ------------------------------- | -------- | -------- | -------- | -------- | | Ucayali 2019-03, weighted | 12.135 % | 12.133 % | 12.150 % | 12.171 % | | Ucayali 2019-03, valid blocks | 12.135 % | 12.237 % | 12.341 % | 12.419 % | | Lake Mead 2017-03, weighted | 16.310 % | 16.316 % | 16.298 % | 16.254 % | | Lake Mead 2017-03, valid blocks | 16.310 % | 15.713 % | 15.461 % | 15.303 % | The weighted share is the level-0 share up to the floor division of the integer mean (0.04 and 0.06 points at most). The plain mean over valid blocks drifts, +0.28 points (Ucayali) and -1.0 points (Lake Mead) by level 3, and that is not rounding: a level's mask is the maximum over a block, so a block with one valid pixel counts as a whole block and the edges of masked areas get more weight at coarse levels. Weight by valid pixels, or measure areas at level 0 or from the CSV. #### Viewer ```bash uv run --with rangehttpserver python -m RangeHTTPServer 8000 # from the repo root cd js && node ../examples/water_masks/viewer_check.mjs --aoi ucayali_santa_maria --month 2019-03 cd js && node ../examples/water_masks/viewer_check.mjs --aoi lake_mead ``` Headless Chromium with software WebGL, `index.html?store=http://localhost:8000/data/stores//water-1`. All checks pass on both stores: * Opens with first paint between 55 and 700 ms over the runs; `int16`, mask present, 117 and 94 timesteps (as in the CSV). * Only "Single band" is enabled among the products (no red, green or nir band); the band selector offers `ndwi` and `water`. * A masked pixel at level 0 draws the background colour (9, 12, 18); a valid pixel six pixels to its right does not. * Click on a river pixel at level 0 (Ucayali 1035, 1035, Mar 2019): "ndwi 0.828 index, water 1 fraction", stored 8279 and 10000; the same values come out of xarray. Lake Mead (1275, 1200, Mar 2017): 0.890, stored 8903, "water 1 fraction". A land pixel (Ucayali 1069, 1027) reads "water 0 fraction", stored 0. * A click on level 2 (the viewer at zoom 0.25; the inspector shows "Level 2 (4× coarser)") reads a fraction: Ucayali Mar 2019 "water 0.500 fraction", stored 5000; Lake Mead Mar 2017 "water 0.555 fraction", stored 5555. Both equal the stored values at that level pixel read with xarray (`lod=2`) and with the reader. * The chart of the level-0 pixel with the `water` band on screen is "water (fraction) over time": 92 charted months, 90 of them at 1 (Ucayali), 25 gap months drawn as breaks. The chart of the level-2 pixel plots fractions: 97 charted months, 68 of them strictly between 0 and 1. * No console errors. Screenshots in `data/reports/`: `viewer_ucayali_santa_maria_ndwi.png` (ndwi, May 2024, 40 % masked), `viewer_ucayali_santa_maria_water_chart.png` (water, Mar 2019, inspector and chart), `viewer_ucayali_santa_maria_water_coarse.png` (water, level 2 click reading 0.500 fraction), `viewer_lake_mead_ndwi.png`, `viewer_lake_mead_water_chart.png`, `viewer_lake_mead_water_coarse.png`. Quicklooks of true colour, ndwi and water for 9 months are `data/reports/water__.png` (Ucayali 2016-04, 2017-05, 2019-11, 2019-12; Lake Mead 2016-01, 2017-02, 2017-03, 2018-04, 2022-01). ### Water fraction by year Share of valid pixels that are water, yearly minimum and maximum over the well-observed months (Ucayali: 90 % valid pixels and blue median at most 500 DN, 78 of 117 months; Lake Mead: 90 % valid pixels, minus the mixed-baseline months 2021-12 and 2022-01, 82 of 94). Area in km² is the same months' `water_km2`. | Year | Ucayali months | Ucayali min % (month) | Ucayali max % (month) | km² | | ---- | -------------- | --------------------- | --------------------- | ------------ | | 2015 | 1 | 8.27 (11) | 8.27 (11) | 57.2 | | 2016 | 5 | 6.74 (09) | 8.97 (06) | 50.9 to 68.0 | | 2017 | 5 | 8.40 (10) | 9.53 (12) | 63.0 to 71.6 | | 2018 | 7 | 8.16 (10) | 9.66 (07) | 62.1 to 73.3 | | 2019 | 7 | 8.22 (08) | 12.13 (03) | 62.6 to 90.8 | | 2020 | 10 | 7.96 (10) | 10.58 (03) | 60.7 to 80.6 | | 2021 | 10 | 8.37 (10) | 10.90 (04) | 63.6 to 83.0 | | 2022 | 9 | 8.49 (09) | 10.18 (05) | 63.5 to 77.7 | | 2023 | 9 | 7.96 (10) | 9.93 (01) | 60.6 to 75.6 | | 2024 | 6 | 8.01 (10) | 11.94 (03) | 61.1 to 83.1 | | 2025 | 9 | 8.74 (01) | 11.20 (05) | 65.7 to 82.1 | | Year | Lake Mead months | Lake Mead min % (month) | Lake Mead max % (month) | km² | | ---- | ---------------- | ----------------------- | ----------------------- | ------------- | | 2015 | 2 | 14.99 (08) | 15.68 (09) | 97.1 to 102.6 | | 2016 | 9 | 13.85 (06) | 15.06 (11) | 89.4 to 96.7 | | 2017 | 10 | 14.67 (07) | 16.31 (03) | 95.0 to 99.8 | | 2018 | 12 | 14.63 (07) | 16.16 (01) | 94.9 to 99.9 | | 2019 | 11 | 13.64 (03) | 16.16 (02) | 87.4 to 100.4 | | 2020 | 11 | 15.07 (03) | 16.19 (12) | 94.0 to 102.7 | | 2021 | 11 | 14.31 (09) | 15.70 (01) | 92.7 to 97.8 | | 2022 | 11 | 13.23 (08) | 14.29 (02) | 85.7 to 92.3 | | 2023 | 5 | 13.21 (04) | 13.83 (01) | 86.2 to 89.4 | The Ucayali series has a seasonal cycle (highest March to May, lowest August to October, 8 % to 12 %). Lake Mead's share is higher in winter than its area in km² is, because winter shadow pixels are masked and leave the lake a larger share of what is valid; prefer `water_km2` there. ### Candidate observations Starting points to look at in the viewer, not conclusions. **Ucayali** 1. Persistent bend migration. Compare the dry seasons (August to October) of 2016 and 2025 with the `water` band. Pixels that are water in more than half of the valid dry-season observations: 30.3 km² gained and 15.3 km² lost between the two (all pixels valid in both). The change is coherent, not speckle: crescents of gain on one bank and of loss on the other bank of the large bends (the lower-right bend and the central loop; a gain/loss map of 2016 to 2025 shows it). Gross change between consecutive dry seasons is 5 to 13 km², largest 2016 to 2017 (13.4 gained, 4.8 lost) and 2024 to 2025 (12.7 gained, 2.8 lost); 2018 to 2019 and 2019 to 2020 lose more than they gain (10.0 and 9.5 lost). 2. When it happens. For the 455,661 pixels whose dry-season state differs between 2016 and 2025, a step fit of the monthly `water` series (at least 12 valid months on each side, at most 5 % misfit) dates 200,685 of them. Land-to-water steps run at 1.2 to 3.3 km² per year in 2017 to 2024 with no single dominant event; the largest month is 2022-03 (1.5 km², and 2022-02 has no mosaic, so it spans 2022-01 to 2022-03), then 2024-11 (0.8), 2024-02 (0.7), 2021-02 (0.7), 2019-02 (0.6). Water-to-land steps cluster in June and July of 2018 to 2021 (0.2 km² in each of 2019-06 and 2019-07), which a falling dry-season stage would explain. 3. Largest month-to-month change in water area, well-observed adjacent months: 2023-01 to 2023-02 (-13.6 km²), 2019-04 to 2019-05 (-12.7), 2020-12 to 2021-01 (+12.3), 2022-12 to 2023-01 (+11.8), 2023-04 to 2023-05 (-11.4), 2016-08 to 2016-09 (-10.0). They are 14 to 18 % of the water area and probably stage (floodplain lakes and the tributary at the left filling and emptying), not planform; 2023-02 has 92 % valid pixels, so part of that drop may be mask holes on the channel. Area differences are the wrong statistic for cutoffs: an abandoned channel stays water as an oxbow. 4. Months that look like events and are not: 2019-11 (6.1 %) then 2019-12 (12.2 %): cloud and shadow holes over the channel in November, haze and cloud edges counted as water in December (blue medians 568 and 716 DN). 22 of 117 months have a blue median above 500 DN. **Lake Mead** 1. Drawdown. The lake area in the AOI is flat from 2016 to 2021 (yearly means 93.8 to 98.5 km²) and then falls: 97.8 km² in 2021-03, 92.3 in 2022-02, 88.2 in 2022-06, 86.5 in 2022-08, about 86 through 2023-04. The decline is smooth between 2022-02 and 2022-08, which uses the offset-corrected months. 2. Reversal. 86.2 km² in 2023-04 to 89.4 in 2023-05 (+3.2 km², 13.2 % to 13.7 %); 2023-06 has 72 % valid pixels and a share of 17.3 % that is not comparable, and the mosaics after it are dead. 3. Winter months with dark water (2015-12, 2016-01, 2016-02, 2020-01, 2021-11, 2021-12): the lake is DN 1 in every band and is carried by the dark rule. Look at their ndwi (about 0 over the lake) and `water` side by side. 4. Not events: 2022-01 (71.2 km², -20 km² from 2021-12) is a mix of processing baselines, with patches at NDWI near 0 and scene-boundary rectangles; 2019-03 (87.4 km², against 95.3 and 97.9 around it) has a pale haze or cloud patch over the lake that NDWI reads as land. ### What this stack cannot show * **NDWI only.** No SWIR band (the mosaics are B02, B03, B04, B08). NDWI is high for clear water and near 0 for turbid water, shallow water over sand, and dark pixels at the DN floor; it is also raised by haze and thin cloud. The Ucayali floor and the Lake Mead dark rule are patches over those cases, not fixes. * **Cloud gaps.** The mask marks only pixels where no scene was valid. SCL classes 4, 5, 6, 7 and 11 pass, and class 7 (unclassified) lets haze and thin cloud through, so hazy months carry false water and missed water (2019-12, 2019-03 Lake Mead). Ucayali has 14 months with fewer than half the pixels observed, 6 with fewer than 20 %. Water fractions are shares of valid pixels, so a cloud over the channel changes them. * **Composite blur.** Each month is a per-pixel median of the month's valid scenes. A pixel that changed during the month takes the majority state; the stack has no sub-monthly timing, and a bar or bank that moves between scenes can appear as a mixture. * **Mosaic problems you inherit** (below), and `coverage` as a flag, not a scene count. * **Overview levels of `water` are block fractions, not a mask.** A zoomed-out click reads the fraction of the block's valid pixels that are water, and averaging a level over its valid blocks differs from the level-0 share where blocks are partly valid (table under "Checks"). Measure areas at level 0 or from the CSV. * **Ground truth.** Nothing was compared with an independent water map. #### What the sword-water-masks ensemble adds `/Users/jakegearon/projects/sword-water-masks` (`water_ensemble.py`) votes six methods per pixel, `min_votes` 4 of 6: NDWI, MNDWI (green and SWIR1), AWEI_nsh and AWEI_sh (SWIR1 and SWIR2, the second with a shadow correction), ML4Floods and DeepWaterMap (which also read SWIR). The SWIR indices are what separate turbid river water, which is near NDWI 0, from bare sand and wet soil, and AWEI_sh keeps shadow from reading as water; the two networks need more bands than B02 to B08. Its SCL bad-pixel set also masks cloud shadow and thin cirrus, which the mosaics let through, and it has centerline extraction and stable/new/abandoned change classes, which is the bend-migration question above done properly. Its NDWI member uses a fixed threshold of 0, which is the floor used here for Lake Mead. To run it on these reaches the ingest would need B11 and B12 (20 m, resampled to 10 m) added to the monthly mosaics, about 1.5 times the stored bytes for six bands instead of four, and a rerun of the download. ### Problems in the mosaics Found while building; the mosaics were not modified. * **Lake Mead 2023-07 to 2026-03 are copies of 2023-06.** Each of the 33 files matches its predecessor (compared at every seventh pixel) and has coverage 0 everywhere. The ingest wrote them on 2026-04-19, before the signed-URL fix of a21e68f, which is the failure `.napkin.md` describes for Ucayali (later months carried forward after the token expired). 2023-06 itself is only 72 % valid and its mean coverage is 0.04 against 0.3 to 0.5 in other months, so it is probably cut short too. To recover: move `2023-06.npz` to `2026-03.npz` out of `data/mosaics/lake_mead/`, rerun `examples/sentinel2_pc/ingest.py --aoi lake_mead` (it skips months whose file exists), then rebuild. Until then the store ends in 2023-06. Appending the recovered months to this store would rewrite its single whole-axis shard per cell (spec 14); a store built to grow should be unsharded, which is the encoder default now. * **Lake Mead from 2022-02 still carries the +1000 DN processing-baseline offset.** Band medians jump by about 1100 DN between 2022-01 and 2022-02 (B02 1166 to 2122) in every band; Ucayali, corrected, does not. NDWI is not invariant to an additive offset: the 99th percentile of NDWI went from 0.89 to 0.99 before 2022 to 0.09 to 0.10 after. The build corrects it with `--boa-offset-from 2022-02`. 2021-12 and 2022-01 are mixed months (scenes from both baselines under one per-pixel median) and cannot be corrected; they are in the store, and 2022-01 has a water fraction of 11.2 % against 14 to 15 % around it. * **Lake Mead mosaics have DN 1 over dark winter water** (see dark pixels above). Not a mosaic fault; the L2A floor. * **Ucayali has 7 months without a mosaic inside its range** (2015-12 to 2016-03, 2017-04, 2017-06, 2022-02), so the time axis has gaps; the viewer labels each step by its month. ### Findings about the tools * By the viewer's code (`isReflectance` in `products.js`), a band with `scale` 1e-4 and no `units` is reflectance (tone mapped, four decimals), so these bands carry `units` "index" and "fraction". The sidebar prints them after the value ("0.828 index", "1 fraction"). * Only "Single band" is enabled for a store without red, green and nir bands, so there is no water-coloured product; `water` is drawn as a grey ramp, black and white at level 0, grey where an overview block is part water. * With a linear stretch, forest at the low end of the ndwi stretch draws near black (7, 7, 7) next to masked pixels at the background (9, 12, 18); in the screenshots masked areas are hard to tell from low values except at their outlines. * `coverage` of 0/1 is shown as "Observed by 1 scene", a count label for a flag. At coarse levels coverage is the rounded mean of the flag, so a valid block with fewer than half its pixels observed reads "no scene (gap-filled)" (Lake Mead level 2, 2017-03, with water 0.555). * `xr.open_dataset(engine="chronozarr")` returns both bands as one float32 `data` variable, so `water` is 0.0 or 1.0 at level 0 (a fraction at `lod` 1 to 3) and NaN where masked; `physical=False` gives the stored integers.