Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LivingEarth-in-a-Box (LE-Box) Pipeline — Installation & Usage Guide

The LE-Box pipeline provides an automated, globally applicable workflow for generating Living Earth land cover products from Earth Observation (EO) data. It transforms EO data into ready-to-use Environmental Descriptors (EDs) for the Living Earth Land Cover Classification System (LCCS), automatically generates a ready-to-run le_lccs configuration, and produces ready-to-use Living Earth land cover products for any geographic area worldwide. These products can be directly used for subsequent analyses without requiring additional classification processing.

The pipeline can be run using a default configuration, providing a standardized and reproducible workflow that can be applied across different geographic areas. It can also be customized to meet local requirements, allowing the selection and adaptation of input datasets, parameters, and processing settings to account for local data availability, environmental characteristics, and specific user/stakeholder needs.

The pipeline consists of four scripts that are normally executed sequentially, either individually or through the single orchestrator, run_le_pipeline.py.

Step Script Purpose
1 land_cover_classification_updated.py Trains (or applies) a Random Forest model per category, producing one ED_{model}_{region}_{year}_{category}.tif raster per category
2 Var_4LE_processing.py Combines the per‑category rasters into the final Living Earth environmental descriptors: VEG01, CULT01, ART01, AQU01, LIFEF, (LEAFTY)
3 create_config.py Writes the le_lccs classification config YAML pointing at the step 2 outputs
4 (optional) le_lccs_odc.py (external, livingearth_lccs package) Runs the Living Earth LCCS classification using the config from step 3

run_le_pipeline.py chains steps 1–4, with flags to skip any step and rerun downstream steps against existing outputs.


1. One‑command environment setup: install.sh

If you're setting up a fresh LE-Box environment from scratch (rather than adding these scripts to an existing JupyterHub and/or Open Data Cube), install.sh automates the full stack: it provisions the Open Data Cube deployment (cube-in-a-box), lays out the shared folders the pipeline scripts expect, and runs your own install_lib/ installers on top.

1.1 Prerequisites

Because it delegates to cube-in-a-box's make setup, you need, in addition to a working bash/git:

  • Docker and Docker Compose (v2, i.e. the docker compose CLI used by cube-in-a-box's Makefile)
  • make
  • Sufficient free disk space and RAM. By default, the version of cube-in-a-box deployed with LE-Box references EO datasets through their source href locations rather than downloading and storing complete copies locally. This significantly reduces local storage requirements, avoids unnecessary duplication of large EO datasets, and contributes to the lightweight and portable nature of LE-Box, allowing it to be deployed in different geographic and computational environments. Nevertheless, a few GB of local disk space/RAM are required for the ODC database, PostgreSQL database, JupyterHub, Docker images, and associated components of the local Docker stack.
  • Outbound internet access to clone GitHub and, during make setup, to pull Docker images and read href indexed starter data.
  • All requirements for installing cube-in-a-box — needed, as a consequence, for LE‑Box too — are available on the cube-in-a-box page: requirements.txt

1.2 Directory layout it expects

Run install.sh from the root of your pipeline repo, with this layout in place beforehand:

LE-Box/
├── install.sh
├── install_lib/
│   ├── install_lccs.sh
│   └── install_python.sh
└── scripts/
    ├── run_le_pipeline.py
    ├── land_cover_classification_updated.py
    ├── Var_4LE_processing.py
    └── create_config.py
    # (+ anything else you want available at /notebooks/shared/lebox)

external/ is created by the script itself and should generally be left out of version control (it's a full clone of cube-in-a-box plus the containers' shared state).

1.3 Install a fresh LE-Box environment (with ODC)

First, clone the LE-Box pipeline repository itself:

git clone <your-pipeline-repo-url>

Then run the installer from its root:

cd <your-pipeline-repo>
chmod +x install.sh
./install.sh

Or, to pass a custom bounding box and/or period of interest through to cube-in-a-box's make setup (Sentinel‑2 auto‑indexing region):

./install.sh BBOX=min_x,min_y,max_x,max_y DATETIME=YYYY-MM-DD/YYYY-MM-DD

For example, you can install a LE-Box environment for a study site for the period 2017-2025:

./install.sh BBOX=147.01,-19.55,147.25,-19.40 DATETIME=2017-01-01/2025-12-31

Because of set -e, the script stops on the first failing command — if it exits partway through, check the last printed step (clone/pull, the install_lib/*.sh scripts, or make setup) to see which stage failed before re‑running.

Re‑running install.sh later is safe for the cube-in-a-box clone step (it git pulls instead of re‑cloning) but will overwrite .env with .env.default again on every run — back up any local .env edits first.

1.5 After it finishes

make setup brings up the Docker stack, so once install.sh completes:

  • JupyterHub (from cube-in-a-box) should be reachable per that project's own README/.env port settings, with /notebooks/shared mounted inside
    • LE-Box JupyterHub environment is now available at http://localhost/jupyter. Only Authorized Users listed in JUPYTERHUB_ADMINS or JUPYTERHUB_USERS in the .env file, or manually added by an admin user can successfully sign up. Signup using admin or guest pre-registered usernames and a chosen password and then login. Additional users and admins can be added using the protocol decribe in the cube-in-a-box's User Management guidance: https://github.com/LivingEarthLab/cube-in-a-box
    • When installing a fresh LE-Box environment (with ODC), LE-Box scripts are available under /notebooks/shared/lebox. shared/ is read-only shared folder for all users. Please copy the lebox folder into your user environment /notebooks/ (i.e., writable environment).
  • Run the LE-Box classification pipeline (§2) either inside your user LE-Box Jupyter environment, or from any host that can reach the same ODC index (see §2 for what --mosaic_path odc:.../--ref_path odc:... need).

2. Running the LE-Box pipeline

2.1 Automated default version from EO to Living Earth product: run_le_pipeline.py

This command line automatically orchestrates all four steps.

python run_le_pipeline.py \
    --region AU \
    --year 2020 \
    --model_name rf_model \
    --output_dir AU_2020 \
    --mosaic_path odc:aef_annual \
    --ref_path odc:esa_worldcover \
    --train_model --save_model \
    --categories ARTIFICIAL BARE NTVherba CULT NTVwoody WATER NAV VEG \
    --odc_bbox "147.01,-19.55,147.25,-19.40" \
    --output_crs EPSG:32755 \
    --output_resolution 10 \
    --config_output config_files/AU_2020_config.yml \
    --run-lccs

Key flags:

  • --region, --year, --model_name, --output_dir — shared naming/output location across all steps.
  • --output_crs / --output_resolution — propagated to step 1's --odc_crs/--odc_resolution, step 2's --epsg, and step 3's --crs.
  • --odc_bbox (+ --odc_bbox_crs, default EPSG:4326) — used for the ODC query in step 1 and auto‑reprojected into --output_crs to fill in step 3's extent (--min_x/--max_x/--min_y/--max_y), unless those are given explicitly.
  • --skip-predict, --skip-combine, --skip-config — rerun only part of the pipeline, e.g. after fixing a downstream step without redoing the (slow) classification.
  • --lc-args "..." / --config-args "..." — shell‑quoted pass‑through for any flag of the underlying script not already exposed directly.
  • --run-lccs — runs the Living Earth LCCS classification step.

Training instead of predicting: add --train_model --ref_path <geotiff or odc:product> --save_model; --mosaic_path is also required in this mode.

2.2 Customizing the run with parameters

run_le_pipeline.py exposes most of the useful flags from all three underlying scripts directly, and forwards anything else via --lc-args (step 1) / --config-args (step 3). Parameters are grouped below by what they control.

Shared / naming

Flag Effect
--region (required) Region code used in output filenames and as the config's --region-code.
--year (required) Processing year (YYYY), used in output filenames.
--model_name Prefix used in all model filenames (default rf_model).
--output_dir (required) Where step 1/2 write rasters.
--output_crs Target CRS for everything: ODC query CRS, reprojection target, and config CRS (e.g. EPSG:32755).
--output_resolution Output resolution in --output_crs units.

Input data (step 1)

Flag Effect
--mosaic_path EO GeoTIFF path, or odc:<product> for direct ODC loading.
--tiles_dir Use pre‑tiled EO GeoTIFFs (mosaic_tile_*.tif) instead of --mosaic_path.
--categories Space‑separated list of categories to train and/or classify (e.g. ARTIFICIAL BARE CULT VEG WATER). Omit to use the script's
built‑in default set.
--categories_file YAML file overriding/extending category definitions (see §2.3, and example_categories_template.yaml).
--odc_product / --odc_ref_product ODC product names when using odc: mosaic/reference paths.
--odc_bbox "minx,miny,maxx,maxy" Query bounding box; also auto‑reprojected into --output_crs to fill step 3's extent if --min_x/etc. aren't given explicitly.
--odc_bbox_crs CRS of --odc_bbox (default EPSG:4326).
--veg_product, --veg_scl_classes, --veg_maxndvi_threshold Tune the NDVI‑based VEG category (source product, SCL cloud‑mask classes, NDVI threshold).

Training mode (step 1)

Flag Effect
--train_model Train in addition to predict. Requires --mosaic_path and --ref_path.
--ref_path Ground‑truth raster (GeoTIFF or odc:<product>) to train against.
--sample_size Number of training samples drawn per category (default: 20000 for presence and 20000 for abscence).
--test_size Held‑out fraction for evaluation (e.g. 0.3).
--n_estimators Random Forest tree count.
--save_model Persist the trained model to models/ for reuse in prediction runs.

Combination step (step 2)

Flag Effect
--roads_path Roads raster added into ART01; omit to build ART01 without roads.
--veg_input Override the VEG01 input path (default: derived from Sentinel-2 Level2A --year time series).

Config output (step 3)

Flag Effect
--min_x/--max_x/--min_y/--max_y Config extent, in --output_crs. Auto‑computed from --odc_bbox if omitted.
--res_x/--res_y Config resolution; default to --output_resolution / -1 × --output_resolution.
--classification_time Classification date written into the config (default {args.year}-12-31).
--product_name Product name written into the config.
--config_output (required) Output path for the YAML config for Living Earth LCCS software run.

Escape hatches and step control

Flag Effect
--run-lccs Also run the Living Earth classification step at the end.
--lc-args "..." Extra raw, shell‑quoted arguments appended to the land_cover_classification_updated.py call — for anything not already exposed above.
--config-args "..." Same, for create_config.py (e.g. --producer, --classification-scheme, --output-location).
--skip-predict Skip step 1 (reuse existing ED_*.tif files).
--skip-combine Skip step 2 (reuse existing ART01/VEG01/AQU01/...).
--skip-config Skip step 3 (classification + combination only).

3. Output naming conventions

All rasters share one naming scheme so steps 1–3 chain together without renaming:

ED_{model_name}_{region}_{year}_{CATEGORY}.tif   # step 1 output, per category
ART01_{model_name}_{region}_{year}.tif           # step 2
VEG01_{model_name}_{region}_{year}.tif           # step 2
AQU01_{model_name}_{region}_{year}.tif           # step 2
LIFEF_{model_name}_{region}_{year}.tif           # step 2
LEAFTY_{model_name}_{region}_{year}.tif          # step 2 (only if CONIFEROUS was classified)

CULT is used directly from step 1's ED_..._CULT.tif — it is not recombined by Var_4LE_processing.py.


4. Troubleshooting

  • "expected ART01/VEG01/... output not found" (step 3, via the orchestrator) — one of steps 1/2 failed or produced files under a different --model_name/--region/--year. Check the printed logs for that step.
  • ImportError: No module named 'osgeo' — install GDAL's Python bindings matching your system GDAL (conda install -c conda-forge gdal is the most reliable route).
  • ODC query returns no data — check --year, --odc_bbox, --odc_crs, --odc_resolution against what's actually indexed for that product in your datacube.
  • Out‑of‑memory / killed (exit code -9) during reprojection of a large reference raster — stream via rasterio.band() rather than loading the full array into memory before reprojecting (already applied in the current load_reference_from_geotiff/reprojection path; if you've customized this, watch for full‑array reads of large rasters).

Authors

  • Carole Planque — University of Geneva, Institute for Environmental Sciences (ORCID)
  • Bruno Chatenoux — UNEP/GRID-Geneva
  • Sebastien Chognard — University of Geneva, Institute for Environmental Sciences (ORCID)
  • Thomas Piller — UNEP/GRID-Geneva (ORCID)
  • Richard Lucas — Aberystwyth University, Department of Geography and Earth Sciences (ORCID)
  • Gregory Giuliani — University of Geneva, Institute for Environmental Sciences (ORCID)

Acknowledgements

All the developments have made possible thanks to the financial support of the European Union ‘Horizon Europe Program’ that funded the LandShift (Grant Agreement no. 101182007), Nostradamus (Grant Agreement no. 101134888), and NEMESIS (Grant Agreement no. 101219087) projects.

About

Living Earth in a Box (LE-Box): automated environmental descriptor retrieval and Living Earth land cover classification in Open Data Cube.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages