Changelog¶
0.9.5 (2026-07-21)¶
Two rounds from the 2026-07-20 rechunkit/cfdb dual-blind code review: the bug-fix round (correctness) and the performance round. Requires the matching rechunkit release (>= 0.6.0 — cfdb calls the new planner parameters and will TypeError on 0.5.1).
Performance round¶
- Groupby-style iteration is dramatically faster. Iterating a variable in small chunks along one coordinate with full extent along the others (
groupby('time'), dailygroupby({'time': 'D'})with non-divisor chunk sizes,iter_chunkswith flat wide shapes) previously re-read and re-decompressed every overlapping storage chunk per group — measured 27x read amplification and 58 MB/s on the benchmark case. With rechunkit's clipped read planning and batched reads this is now 2.4x planned reads and 552 MB/s (9.6x faster) at the same memory budget, and exactly 1x amplification when the budget accommodates the clipped ideal - Rechunking after coordinate prepends no longer re-decompresses neighbouring chunks. The rechunker now declares the source grid to rechunkit at the real storage-chunk alignment (coordinate origins are folded into rechunkit's phase machinery), so every read maps to exactly one storage chunk: decompression amplification on origin-shifted variables drops from ~3.4x to 1.0x, and the prepended-rechunk benchmark improves ~60%. Yields are byte-identical to the previous implementation on every consumer path (verified); the multi-chunk assembly remains as a tested defensive fallback
max_memis now an honest total (via rechunkit): it bounds the read buffer plus reorder and batch allocations, instead of silently exceeding the budget (previously up to hundreds of times over on wide arrays). The documented exceptions are the irreducible floors and the wide-array pending residual — see the rewritten Memory Model indocs/concepts/rechunking-internals.md. In rare tight-budget configurations honesty costs bounded extra reads (never more than brute force)- Changed:
Rechunker.rechunk()andRechunker.calc_n_reads_rechunker()defaultmax_memraised from 227 (128 MB) to 229 (512 MB), matchingiter_chunks/groupby/map/DatasetRechunker— the clipped ideal path is then reachable for typical full-field iteration patterns - Fixed:
Rechunker.calc_ideal_read_chunk_shape/memreported wildly inflated numbers for non-divisor targets (the unclipped LCM — e.g. a 193 GiB "ideal" on a 3 GB array); they now report the extent-clipped shape the rechunker actually uses, at every level includingDatasetRechunker.calc_ideal_read_chunk_mem Rechunker.calc_n_reads_rechunkerpredictions match instrumented read counts on origin-shifted variables and views (same storage-space transform as the rechunker)
Bug-fix round¶
- Fixed: missing chunks read through rechunker-based paths decoded to wrong values for encoded dtypes.
rechunk(),iter_chunks(chunk_shape=...),groupby(), and netCDF export filled missing chunks with encoded 0 regardless of the dtype's reserved fillvalue — with an explicit nonzero fillvalue a missing chunk decoded to a plausible real value (the offset) instead of NaN, and int-packed dtypes returnedmin_value - 1through these paths while.datareturned 0. Missing data now reads as the decoded blank (NaN/NaT/0) identically on every path - Changed: auto-packed integer dtypes now carry
fillvalue=0(encoded 0 was always structurally reserved:offset = min_value - 1puts every legitimate value at code ≥ 1).decode()maps the reserved code to decoded 0;encode()is unchanged, so a legitimate 0 (whenmin_value <= 0) still packs to1 - min_valueand can never collide with the missing marker (in files or netCDF exports). The fillvalue now round-trips through stored dtype dicts (previously the reconstruction branch dropped it on reopen). Legacy files (created before this release; no fillvalue in the stored dtype) keep their old read behavior — aUserWarningfires when rechunker/encoded paths read such int-packed variables, since missing chunks there still read asmin_value - 1(and a rechunk-copy of such a file materializes that value); recreate the variable to adopt the new semantics.combine()accepts the old/new fillvalue difference (output adopts the new-style dtype) - Fixed: partial chunk writes to int-packed variables with
min_value >= 2raised (integer value(s) outside the encodable range) — the blank cells of a partially-written chunk went throughencode(). Writes for packed dtypes now read-modify-write in encoded space: only the user's values are encoded, untouched cells keep their exact stored codes (no decode→re-encode round trip), and unwritten cells remain honestly missing on disk instead of becoming written values. The same fix covers int-packed coordinates with a partial last chunk (append/prependraised the same way) - Fixed: indexing bounds now follow numpy semantics. Previously
var[n]on a length-n variable andvar[-1]silently returned blank fill data (a phantom chunk), and over-long slices returned phantom-padded arrays (var[35:60]on a length-40 variable → 25 values). Now: negative ints wrap (var[-1]= last element), out-of-range ints raiseIndexError, slice bounds clamp to the variable extent exactly like numpy, and fully-out-of-range or descending slices raiseValueError. Consequences:set()through an over-long slice expects the clamped shape, and.locwith a scalar label beyond the coordinate end raisesIndexError(previously returned a phantom blank) - Fixed:
groupbywith month/year periods on a sliced variable view returned wrong groups — boundaries were computed from the full coordinate but applied view-relative (a Mar–Jun view got Jan–Apr month lengths). Groups are now computed from the view's own coordinate window. (Dataset-levelselect().groupby()was already correct) - Fixed: time-period groupby with an enforced
stepignored calendar alignment —{'time': 'D'}on hourly data starting 06:00 produced fixed 24-value windows each spanning two calendar days, while the same call on a step-less coordinate produced true calendar days. The fast path now additionally requires the first coordinate value to sit on the period boundary and otherwise falls back to calendar groups, so both paths always produce identical groups. Note the documented anchoring rules: single-unit periods are calendar-aligned; multi-count periods ('7D','6h') anchor at the first coordinate value;'W'follows numpy's Thursday-anchored week epoch (use'7D'for first-timestamp-anchored weeks) - Fixed: creating a data variable with
chunk_shape=Nonewhile a referenced coordinate was still empty baked a zero chunk shape into the file, making the variable permanently unwritable (ZeroDivisionError). This now raises aValueErrornaming the empty coordinate(s) — passchunk_shapeexplicitly or fill the coordinate first (see the new guidance in the data-variables guide). Explicitchunk_shapevalues are now also validated (> 0) and accept numpy integers - Fixed: netCDF export of int-packed variables crashed (
TypeErrorcomputingscale_factorfrom the integer packing's absent factor); such variables now export withscale_factor = 1.0,add_offset, and_FillValue = 0, so external CF readers see missing chunks as masked and unpack written values correctly - Fixed: netCDF export of packed datetime variables wrote offset-shifted timestamps (the packed values were written raw under
units: ... since 1970-01-01, shifting exported times by the packing offset — decades off). Datetimes now always export as true epoch ticks - Fixed:
compute_int_and_offsetwidth selection was off by one — at ranges exactly filling an integer width, the declaredmax_valuewas unencodable (packed ints raised; packed floats silently decoded the declared max to NaN). Auto-selected widths now reserve room for the fill code; dtypes at exact boundary ranges may select the next width up compared with earlier releases (new files only) - Fixed: auto-packed datetime dtypes (and integer/float packings given numpy-scalar
min_value/max_value) produced a numpy-typedoffsetthat crashed dataset close (msgspeccannot encodenumpy.int64) — packed-datetime variables created viamin_value/max_valuewere never persistable before this fix - Changed:
iter_chunks()yields a fresh blank array for each missing chunk (previously one shared blank was yielded for all missing chunks, so mutating a yielded array in place poisoned later yields). The general yield-lifetime contract is now documented on all chunk generators: consume or copy each yielded array before advancing the generator, and treat yields as read-only - Changed:
DatasetRechunker.calc_n_reads_rechunkernow sums write counts across variables (matching how reads were already counted), so both numbers measure physical I/O operations for the batch - Fixed:
Variable.items(decoded=...)ignored itsdecodedparameter
0.9.4 (2026-07-19)¶
- Fixed (follow-up, same guard): the float branch compares against bounds prepared as exactly-representable values in the data's own dtype (an unrepresentable bound shrinks one ULP toward the valid range) — a naive float32 comparison against
iinfo(uint32).maxpromoted the bound tofloat32(2**32), letting the top ~256-ULP band through to the undefined cast. Native-dtype comparison also keeps this hot per-chunk write path cheap: ~0.1 ms guard overhead per 500k-value chunk (vs ~1.6 ms for a float64-copy approach), with the fillvalue substitution pass skipped entirely for chunks containing no unencodable values; and the packed-Integer path now raises on values whose scaled representation exceeds the encoded width (previously wrapped modulo the width into plausible ints; ints have no reserved missing value, so mapping to fillvalue is not an option — fail loud). Values inside the width but outside the declared min/max remain the caller's policy, as for floats. - Fixed: packed-dtype
encode()fabricated data for unencodable values. Writing NaN (or inf, or a finite value outside themin_value/max_valueencoding range, or NaT for packed datetimes) through an int-packed dtype hit numpy's undefined float→integer cast: instead of the documented "out-of-range becomes NaN", values wrapped into plausible in-range numbers (e.g. NaN → 214747.36 on aprecision=4, 0–100000uint32 encoding; 60.0 → −5.536 on a−5–55uint16 one). Worse, the fabricated result differed between numpy's scalar and SIMD code paths, so small arrays could round-trip "correctly" while production-sized arrays corrupted — which is exactly how it evaded the test suite.encode()now maps every unencodable value (non-finite, NaT, out-of-range) to the reserved fillvalue before the cast, so they decode back as NaN/NaT as documented. Found by the envlib-ingest-base Phase-7 adversarial review (dual-blind: both reviewers hit the same function from different inputs). Any dataset previously written with NaN-holed arrays through a uint32+ packing on a SIMD path should be checked for the fabricated constant ((2147483648 / 10^precision) + offset); uint16 packings happened to cast NaN→0 (the fillvalue) on x86 and round-tripped correctly by luck — the live esa-sst dataset was verified to be in that category.
0.9.3 (2026-07-14)¶
- Fixed: chained
.locaccess on a temporary variable —ds[var].loc[...](read or write; on datasets, views, and EDatasets alike) crashed withReferenceError: weakly-referenced object no longer exists; only the bound form (v = ds[var]; v.loc[...]) worked. The loc indexer is now created per access via a property and holds a strong reference to its variable for the duration of the expression; because it is no longer stored on the variable, no reference cycle is introduced and finalizer/collection timing is unchanged. Found by the envlib esa-sst ingest — the chained form is the natural notebook idiom (cat.query(...)[0].open()[var].loc[...]) - Note for user code:
v.locnow returns a fresh indexer object on each access (v.loc is v.locisFalse), and assigning tov.locraisesAttributeError
0.9.2 (2026-07-13)¶
- Changed:
EDataset.push()now returns ebooklet'sPushResult(passthrough; requires ebooklet >= 0.10.0):result.updated,result.failures(failed keys → error strings; pending changes retained for retry), andbool(result)= fully-successful push. Previously it returned True/False/dict — note the old partial-failure dict was truthy, soif ds.push():misread partial failure as success; any failure is now falsy - Updated dependency pins:
ebooklet>=0.10.0(persistent journal, generational storage format 2, typed exceptions, PushResult, offline read mode — see ebooklet's changelog; format-1 remotes need the one-time re-push upgrade described there)
0.9.1 (2026-07-12)¶
- Fixed:
open_edatasetnow supportsdataset_type='ts_ortho'(it used to raise TypeError at create) and returns the class matching the dataset's STORED type for existing remotes (it always returnedEGrid, even for ts_ortho remotes) - Changed: in both
open_datasetandopen_edataset, thedataset_typeparameter now applies only when CREATING — existing datasets always open with their stored type and the matching class (previously the class silently followed the parameter, so an existing ts_ortho file opened without the parameter got a Grid-classed object) - Added public
dataset_typeproperty ('grid'or'ts_ortho') on all dataset classes and views — consumers no longer need to read the private sys-metadata - Fixed: the in-memory sys-metadata
dataset_typewas a plain string on freshly-created datasets but an enum on reopened ones (msgspec does not coerce on construction); it is now always the enum. On-disk format unchanged - Fixed: creating a coordinate on a REOPENED dataset crashed with
TypeError: argument of type 'Type' is not iterable(latent consequence of the string/enum inconsistency above) - Updated dependency pins:
booklet>=0.12.6(iterator-contract release; bump happened post-0.9.0 and was previously unchronicled),ebooklet>=0.9.4(5xx retry, 404-integrity re-check, and the flag='n' second-push data-loss fix — found by this release's live tier)
0.9.0¶
- Fixed:
EDataset.push()mid-session used to publish a stale/empty SysMeta (no variables) and no attributes —push()/changes()now flush all in-memory state to the local file first, so a push at any point publishes exactly the current structure. Pushing remains explicit: nothing is ever pushed automatically on close - Fixed:
open_edatasetwith'w'/'c'and a fresh local file used to treat an EXISTING remote dataset as create-new, which would overwrite the remote's structure metadata on the next push — the create decision now also checks the remote, so a fresh local file attaches to the existing remote dataset - Fixed: attributes split-brain — re-accessing a variable could create a second, independent in-memory attrs dict, and close() flushed the stale one last, silently losing newer edits; all attribute access now shares a single dataset-level source of truth
- Fixed: deleting a variable no longer lets its attributes resurrect at close;
attrs.clear()now persists - Added public
Dataset.sync(): flush variable definitions and attributes to the local file without closing (never touches a remote) - EDataset write sessions that change nothing no longer rewrite the metadata slot at open (previously every session pushed an unchanged metadata object)
- Updated dependency pins:
booklet>=0.12.4,ebooklet>=0.9.0
0.6.0¶
- Added
DatasetRechunkerclass for native multivariable rechunking - Added
rechunker()method toDatasetBaseto expose the new multivariableDatasetRechunker - Upgraded
iter_chunksto useDatasetRechunkerinternally, ensuring optimized synchronization and shared memory budget across variables - Fixed prepend and append bugs
- Fixed
copy()dataset issue - Updated dependencies
0.5.0¶
- Added Xarray backend for
cfdb(read-only mapping) - Added time aggregation capabilities (e.g.
groupby(period)) - Fixed Geometry insertion order
- Added comprehensive benchmarks
0.4.3¶
- Added
map()method toDataVariableViewfor parallel chunk processing using multiprocessing - Added pickling support (
__reduce__) toGeometry,String, andCompressorclasses - Fixed
Dataset.prune()to match updated booklet API (removedreindexparameter)
0.4.2¶
- Renamed
timeparameter toiter_dimin interpolation classes (GridInterp,PointInterp) andDataVariableView.interp()to generalize iteration over any non-spatial dimension - Fixed API reference docs rendering by correcting mkdocstrings parser from
googletonumpy
0.3.5¶
- Updated EDataset with new
num_groupsparameter - Uses cfdb-vars for variable definitions
- Added variable templates
Previous Releases¶
See the GitHub releases page for full version history.