The LE-Box pipeline provides an automated, globally applicable workflow for generating
Living Earth land cover products from Earth Observation (EO) data. It transforms
EO data into ready-to-use Environmental Descriptors (EDs) for the Living Earth Land
Cover Classification System (LCCS), automatically generates a ready-to-run le_lccs
configuration, and produces ready-to-use Living Earth land cover products for any
geographic area worldwide. These products can be directly used for subsequent analyses
without requiring additional classification processing.
The pipeline can be run using a default configuration, providing a standardized and reproducible workflow that can be applied across different geographic areas. It can also be customized to meet local requirements, allowing the selection and adaptation of input datasets, parameters, and processing settings to account for local data availability, environmental characteristics, and specific user/stakeholder needs.
The pipeline consists of four scripts that are normally executed sequentially, either
individually or through the single orchestrator, run_le_pipeline.py.
| Step | Script | Purpose |
|---|---|---|
| 1 | land_cover_classification_updated.py |
Trains (or applies) a Random Forest model per category, producing one ED_{model}_{region}_{year}_{category}.tif raster per category |
| 2 | Var_4LE_processing.py |
Combines the per‑category rasters into the final Living Earth environmental descriptors: VEG01, CULT01, ART01, AQU01, LIFEF, (LEAFTY) |
| 3 | create_config.py |
Writes the le_lccs classification config YAML pointing at the step 2 outputs |
| 4 (optional) | le_lccs_odc.py (external, livingearth_lccs package) |
Runs the Living Earth LCCS classification using the config from step 3 |
run_le_pipeline.py chains steps 1–4, with flags to skip any step and rerun
downstream steps against existing outputs.
If you're setting up a fresh LE-Box environment from scratch (rather
than adding these scripts to an existing JupyterHub and/or Open Data Cube), install.sh automates
the full stack: it provisions the Open Data Cube deployment
(cube-in-a-box), lays out the shared folders the pipeline scripts expect,
and runs your own install_lib/ installers on top.
Because it delegates to cube-in-a-box's make setup, you need, in
addition to a working bash/git:
- Docker and Docker Compose (v2, i.e. the
docker composeCLI used bycube-in-a-box's Makefile) make- Sufficient free disk space and RAM. By default, the version of
cube-in-a-boxdeployed with LE-Box references EO datasets through their sourcehreflocations rather than downloading and storing complete copies locally. This significantly reduces local storage requirements, avoids unnecessary duplication of large EO datasets, and contributes to the lightweight and portable nature of LE-Box, allowing it to be deployed in different geographic and computational environments. Nevertheless, a few GB of local disk space/RAM are required for the ODC database, PostgreSQL database, JupyterHub, Docker images, and associated components of the local Docker stack. - Outbound internet access to clone GitHub and, during
make setup, to pull Docker images and readhrefindexed starter data. - All requirements for installing
cube-in-a-box— needed, as a consequence, for LE‑Box too — are available on thecube-in-a-boxpage:requirements.txt
Run install.sh from the root of your pipeline repo, with this layout in
place beforehand:
LE-Box/
├── install.sh
├── install_lib/
│ ├── install_lccs.sh
│ └── install_python.sh
└── scripts/
├── run_le_pipeline.py
├── land_cover_classification_updated.py
├── Var_4LE_processing.py
└── create_config.py
# (+ anything else you want available at /notebooks/shared/lebox)
external/ is created by the script itself and should generally be left
out of version control (it's a full clone of cube-in-a-box plus the
containers' shared state).
First, clone the LE-Box pipeline repository itself:
git clone <your-pipeline-repo-url>Then run the installer from its root:
cd <your-pipeline-repo>
chmod +x install.sh
./install.shOr, to pass a custom bounding box and/or period of interest through to cube-in-a-box's make setup (Sentinel‑2 auto‑indexing region):
./install.sh BBOX=min_x,min_y,max_x,max_y DATETIME=YYYY-MM-DD/YYYY-MM-DDFor example, you can install a LE-Box environment for a study site for the period 2017-2025:
./install.sh BBOX=147.01,-19.55,147.25,-19.40 DATETIME=2017-01-01/2025-12-31Because of set -e, the script stops on the first failing command — if it
exits partway through, check the last printed step (clone/pull, the
install_lib/*.sh scripts, or make setup) to see which stage failed
before re‑running.
Re‑running install.sh later is safe for the cube-in-a-box clone step
(it git pulls instead of re‑cloning) but will overwrite .env with
.env.default again on every run — back up any local .env edits first.
make setup brings up the Docker stack, so once install.sh completes:
- JupyterHub (from
cube-in-a-box) should be reachable per that project's ownREADME/.envport settings, with/notebooks/sharedmounted inside- LE-Box JupyterHub environment is now available at http://localhost/jupyter.
Only Authorized Users listed in JUPYTERHUB_ADMINS or JUPYTERHUB_USERS in the .env file,
or manually added by an admin user can successfully sign up. Signup using
adminorguestpre-registered usernames and a chosen password and then login. Additional users and admins can be added using the protocol decribe in thecube-in-a-box's User Management guidance: https://github.com/LivingEarthLab/cube-in-a-box - When installing a fresh LE-Box environment (with ODC), LE-Box scripts are available
under
/notebooks/shared/lebox.shared/is read-only shared folder for all users. Please copy theleboxfolder into your user environment/notebooks/(i.e., writable environment).
- LE-Box JupyterHub environment is now available at http://localhost/jupyter.
Only Authorized Users listed in JUPYTERHUB_ADMINS or JUPYTERHUB_USERS in the .env file,
or manually added by an admin user can successfully sign up. Signup using
- Run the LE-Box classification pipeline (§2) either inside your user LE-Box Jupyter
environment, or from any host that can reach the same ODC index (see
§2 for what
--mosaic_path odc:.../--ref_path odc:...need).
This command line automatically orchestrates all four steps.
python run_le_pipeline.py \
--region AU \
--year 2020 \
--model_name rf_model \
--output_dir AU_2020 \
--mosaic_path odc:aef_annual \
--ref_path odc:esa_worldcover \
--train_model --save_model \
--categories ARTIFICIAL BARE NTVherba CULT NTVwoody WATER NAV VEG \
--odc_bbox "147.01,-19.55,147.25,-19.40" \
--output_crs EPSG:32755 \
--output_resolution 10 \
--config_output config_files/AU_2020_config.yml \
--run-lccsKey flags:
--region,--year,--model_name,--output_dir— shared naming/output location across all steps.--output_crs/--output_resolution— propagated to step 1's--odc_crs/--odc_resolution, step 2's--epsg, and step 3's--crs.--odc_bbox(+--odc_bbox_crs, defaultEPSG:4326) — used for the ODC query in step 1 and auto‑reprojected into--output_crsto fill in step 3's extent (--min_x/--max_x/--min_y/--max_y), unless those are given explicitly.--skip-predict,--skip-combine,--skip-config— rerun only part of the pipeline, e.g. after fixing a downstream step without redoing the (slow) classification.--lc-args "..."/--config-args "..."— shell‑quoted pass‑through for any flag of the underlying script not already exposed directly.--run-lccs— runs the Living Earth LCCS classification step.
Training instead of predicting: add --train_model --ref_path <geotiff or odc:product> --save_model; --mosaic_path is also required in this mode.
run_le_pipeline.py exposes most of the useful flags from all three
underlying scripts directly, and forwards anything else via --lc-args
(step 1) / --config-args (step 3). Parameters are grouped below by what
they control.
Shared / naming
| Flag | Effect |
|---|---|
--region (required) |
Region code used in output filenames and as the config's --region-code. |
--year (required) |
Processing year (YYYY), used in output filenames. |
--model_name |
Prefix used in all model filenames (default rf_model). |
--output_dir (required) |
Where step 1/2 write rasters. |
--output_crs |
Target CRS for everything: ODC query CRS, reprojection target, and config CRS (e.g. EPSG:32755). |
--output_resolution |
Output resolution in --output_crs units. |
Input data (step 1)
| Flag | Effect |
|---|---|
--mosaic_path |
EO GeoTIFF path, or odc:<product> for direct ODC loading. |
--tiles_dir |
Use pre‑tiled EO GeoTIFFs (mosaic_tile_*.tif) instead of --mosaic_path. |
--categories |
Space‑separated list of categories to train and/or classify (e.g. ARTIFICIAL BARE CULT VEG WATER). Omit to use the script's |
| built‑in default set. | |
--categories_file |
YAML file overriding/extending category definitions (see §2.3, and example_categories_template.yaml). |
--odc_product / --odc_ref_product |
ODC product names when using odc: mosaic/reference paths. |
--odc_bbox "minx,miny,maxx,maxy" |
Query bounding box; also auto‑reprojected into --output_crs to fill step 3's extent if --min_x/etc. aren't given explicitly. |
--odc_bbox_crs |
CRS of --odc_bbox (default EPSG:4326). |
--veg_product, --veg_scl_classes, --veg_maxndvi_threshold |
Tune the NDVI‑based VEG category (source product, SCL cloud‑mask classes, NDVI threshold). |
Training mode (step 1)
| Flag | Effect |
|---|---|
--train_model |
Train in addition to predict. Requires --mosaic_path and --ref_path. |
--ref_path |
Ground‑truth raster (GeoTIFF or odc:<product>) to train against. |
--sample_size |
Number of training samples drawn per category (default: 20000 for presence and 20000 for abscence). |
--test_size |
Held‑out fraction for evaluation (e.g. 0.3). |
--n_estimators |
Random Forest tree count. |
--save_model |
Persist the trained model to models/ for reuse in prediction runs. |
Combination step (step 2)
| Flag | Effect |
|---|---|
--roads_path |
Roads raster added into ART01; omit to build ART01 without roads. |
--veg_input |
Override the VEG01 input path (default: derived from Sentinel-2 Level2A --year time series). |
Config output (step 3)
| Flag | Effect |
|---|---|
--min_x/--max_x/--min_y/--max_y |
Config extent, in --output_crs. Auto‑computed from --odc_bbox if omitted. |
--res_x/--res_y |
Config resolution; default to --output_resolution / -1 × --output_resolution. |
--classification_time |
Classification date written into the config (default {args.year}-12-31). |
--product_name |
Product name written into the config. |
--config_output (required) |
Output path for the YAML config for Living Earth LCCS software run. |
Escape hatches and step control
| Flag | Effect |
|---|---|
--run-lccs |
Also run the Living Earth classification step at the end. |
--lc-args "..." |
Extra raw, shell‑quoted arguments appended to the land_cover_classification_updated.py call — for anything not already exposed above. |
--config-args "..." |
Same, for create_config.py (e.g. --producer, --classification-scheme, --output-location). |
--skip-predict |
Skip step 1 (reuse existing ED_*.tif files). |
--skip-combine |
Skip step 2 (reuse existing ART01/VEG01/AQU01/...). |
--skip-config |
Skip step 3 (classification + combination only). |
All rasters share one naming scheme so steps 1–3 chain together without renaming:
ED_{model_name}_{region}_{year}_{CATEGORY}.tif # step 1 output, per category
ART01_{model_name}_{region}_{year}.tif # step 2
VEG01_{model_name}_{region}_{year}.tif # step 2
AQU01_{model_name}_{region}_{year}.tif # step 2
LIFEF_{model_name}_{region}_{year}.tif # step 2
LEAFTY_{model_name}_{region}_{year}.tif # step 2 (only if CONIFEROUS was classified)
CULT is used directly from step 1's ED_..._CULT.tif — it is not
recombined by Var_4LE_processing.py.
- "expected ART01/VEG01/... output not found" (step 3, via the
orchestrator) — one of steps 1/2 failed or produced files under a
different
--model_name/--region/--year. Check the printed logs for that step. ImportError: No module named 'osgeo'— install GDAL's Python bindings matching your system GDAL (conda install -c conda-forge gdalis the most reliable route).- ODC query returns no data — check
--year,--odc_bbox,--odc_crs,--odc_resolutionagainst what's actually indexed for that product in your datacube. - Out‑of‑memory / killed (exit code -9) during reprojection of a large
reference raster — stream via
rasterio.band()rather than loading the full array into memory before reprojecting (already applied in the currentload_reference_from_geotiff/reprojection path; if you've customized this, watch for full‑array reads of large rasters).
- Carole Planque — University of Geneva, Institute for Environmental Sciences (ORCID)
- Bruno Chatenoux — UNEP/GRID-Geneva
- Sebastien Chognard — University of Geneva, Institute for Environmental Sciences (ORCID)
- Thomas Piller — UNEP/GRID-Geneva (ORCID)
- Richard Lucas — Aberystwyth University, Department of Geography and Earth Sciences (ORCID)
- Gregory Giuliani — University of Geneva, Institute for Environmental Sciences (ORCID)
All the developments have made possible thanks to the financial support of the European Union ‘Horizon Europe Program’ that funded the LandShift (Grant Agreement no. 101182007), Nostradamus (Grant Agreement no. 101134888), and NEMESIS (Grant Agreement no. 101219087) projects.