Merging datasets and filling missing data#

Two related, but distinct, problems come up repeatedly when producing ancillary fields:

  1. Merging - combining a primary source with an alternate source, either to build a larger coverage than either source has on its own, or to embed a high quality local source within a dataset that has wider coverage.

  2. Filling - replacing missing data (masked or NaN values) with valid data taken from elsewhere in the field, so that every point that should have a value does have one.

Note that “merge” here is used differently to Iris, which uses the term for combining cubes along a new dimension - see the Iris Documentation for that meaning. See ants.analysis for the full set of routines covered here.

Merging two sources#

ants.analysis.merge() takes values from a primary source and an alternate source, optionally guided by a validity polygon that marks the region where the primary source should be trusted:

import numpy as np
import ants
import ants.tests.stock as stock

# data must already be shaped to match the cube's shape, so reshape
# the flat arrays of values first
primary = stock.geodetic((4, 4), data=np.arange(16,dtype=float).reshape(4, 4))
# Fill the first two rows with nans so there's something to merge into
primary.data[:2,:] = np.nan

alternate = stock.geodetic((4, 4), data=np.arange(16,dtype=float).reshape(4, 4))
# Fill the last two rows with nans so we know data hasn't come in from the alternate
alternate.data[2:,:] = np.nan

# everywhere valid data is present in the primary, it takes priority;
# elsewhere, the alternate is used.
result = ants.analysis.merge(primary, alternate)

Where you have a shapefile describing the region the primary source is valid for, pass it as a validity_polygon (see Worked examples: the ANTS command line tools for how to produce one with ancil_create_shapefile). A blending_distance can also be supplied to linearly blend between the sources over a number of grid cells either side of the polygon boundary, rather than having a hard edge.

Filling missing data#

Once you have a merged field (or even just a single source with gaps), ants.analysis provides several algorithms for filling missing points by searching for the nearest valid neighbour:

Note

Use KDTreeFill unless you have a specific reason not to. It is the preferred fill algorithm for consistency across both UM and LFRic ancillary generation pipelines. The other algorithms exist for specialised or legacy cases - for example, UMSpiralSearch reproduces the UM’s historical spiral search behaviour bit-for-bit, but depends on the optional um_spiral_search package (see Installing ANTS) and is primarily intended where matching existing UM output exactly is required.

In practice you will usually reach these fill algorithms indirectly, via ants.analysis.make_consistent_with_lsm(), which fills points so that a field is consistent with a target land sea mask (land points filled from land, sea points filled from sea):

target_lsm = ants.io.load.load_landsea_mask("/path/to/target_lsm.nc")

ants.analysis.make_consistent_with_lsm(
    result, target_lsm, invert_mask=True, method="kdtree"
)

The method argument currently accepts "kdtree" (using KDTreeFill, and recommended) or "spiral" (using UMSpiralSearch). The ancil_fill_n_merge and ancil_general_regrid command line tools expose the same choice via their --search-method argument (see Worked examples: the ANTS command line tools).

Key Points#

  • Merging and filling solve different problems: merging combines two sources, filling replaces missing points within a single field.

  • Prefer KDTreeFill ( method="kdtree" in make_consistent_with_lsm(), or --search-method kdtree on the command line tools) for consistency across UM and LFRic ancillary generation pipelines. Only reach for UMSpiralSearch where bit-comparable UM legacy behaviour is specifically required.

  • ants.analysis.merge() and the fill algorithms both check that the cubes involved are compatible (matching STASH code and units) before combining them.

See Also#