Skip to content

Tool · Python · MIT

FlowDivide

Version 1.0released

Version 1.0 is the current release: the Python package on GitHub under the MIT licence. The partitions of both FullBasin collections on this site were made with its C version; the Python package follows it step by step and writes the same tables.

Hydrographic attributes from flow-direction grids too large for a computer's memory. FlowDivide cuts a D8 grid into regions that fit a memory capacity you choose. The cuts follow the rivers: whole basins are grouped, and a basin that is too large is cut at its tributary outlets (Pfafstetter). Each region is then computed once, upstream to downstream, and one value crosses each cut. The results are exact: the same, pixel for pixel, as a run on the whole grid in one array.

Tested on

Tested on three large grids, all on one laptop with 64 GB of memory: HydroSHEDS v2 at 1 arc-second (30 m) over South America (259,200 × 219,600 pixels) and over North America (277,200 × 424,800), and MERIT Hydro at 3 arc-seconds (90 m) over the whole globe (174,000 × 432,000). On all three the results were compared with those of the C code.

How it works

The FlowDivide workflow: FD1 the partition, FD2 the views, FD3 the attributes region by region
The workflow. FD1 makes the partition, FD2 the coarse views and polygons, FD3 the attributes region by region.
South America: the basins, the groups by HydroBASINS Level-03 unit, and the groups along a Hilbert curve
South America, HydroSHEDS v2 at 1 arc-second: (a) the basins, (b) grouped by HydroBASINS Level-03 unit, (c) grouped automatically along a Hilbert curve; hatched, the groups over a capacity of 231 pixels.
The Amazon cut into Pfafstetter pieces and merged into 11 regions
The Amazon at a capacity of 160 deg² (231 pixels at 1 arc-second): cut into Pfafstetter pieces, the pieces too large cut again, then merged into 11 regions numbered so that every flow goes to a larger number.

Get it

Python 3.10 or later. A conda environment with GDAL is the easy way:

conda create -n flowdivide -c conda-forge python=3.11 numpy numba rasterio gdal pandas geopandas pyogrio pyarrow shapely scipy matplotlib
conda activate flowdivide
git clone https://github.com/hellolulujiang/flowdivide.git

The three grids above are built in; point the package at your downloaded copies with environment variables (see the user guide on GitHub), then

python flowdivide.py run south-america --out-root /data/flowdivide
python flowdivide.py run south-america --out-root /data/flowdivide --steps fd3 --capacity 2^31

Any other flow-direction grid, in any coding, projection or resolution:

python flowdivide.py run mygrid --dir /data/mygrid_dir.tif --convention esri --out-root /data/flowdivide

Out come the partition at three capacities (231, 230, 229 pixels), the basin and region tables, the coarse views as rasters and polygons (a GeoPackage opens in QGIS already coloured), and six attributes: Shreve magnitude, distance to the outlet, Hack order, upstream flow length, longest flow path and Strahler order. Every step checks its own output, and a stopped run resumes where it stopped.

Licence

MIT; see the licence. The grids FlowDivide runs on keep their producers' terms. Author: Lulu Jiang (lulujiang.me).

Site under development — pages may change daily. Comments and corrections are welcome: [email protected]