Stages
Stages
ESPnet3 recipes run named stages, invoked with --stages <name> [<name> ...] (or --stages all) on run.py. A stage is a plain method on the recipe's System class (BaseSystem, ASRSystem, TTSSystem); the CLI only selects which stage methods to call, in a fixed canonical order.
Canonical stage order
Whatever order stage names are passed on the CLI, they always run in this order (from DEFAULT_STAGES in egs3/TEMPLATE/asr/run.py):
| # | Stage | System | Page |
|---|---|---|---|
| 1 | create_dataset | all | create-dataset.html |
| 2 | train_tokenizer | ASR | (see ASRSystem) |
| — | create_token_list, remove_long_short | TTS only, before train | (see TTSSystem) |
| 3 | collect_stats | all | collect-stats.html |
| 4 | train | all | train.html |
| 5 | infer | all | inference.html |
| 6 | measure | all | metrics.html |
| 7 | pack_model / upload_model | all | publish.html |
| 8 | pack_demo / upload_demo | all | demo.html |
TTS recipes additionally run create_token_list and remove_long_short ahead of train; these are TTSSystem-only stages with no shipped TTS recipe yet, so they are not documented as standalone pages here. --stages all expands to every stage a recipe's run.py declares (resolve_stages in espnet3/utils/stages_utils.py); passing an explicit subset (e.g. --stages infer measure) still executes in the table order above, not CLI order.
How to use this overview
Use the cards below to choose the stage you need. Each stage guide explains its purpose, required configuration, inputs and outputs, implementation API, safe re-run behavior, and the next stage to run. The execution implementation is available through BaseSystem, ASRSystem, TTSSystem, and run_stages.
For a normal ASR workflow, start with dataset preparation, then statistics, training, inference, measurement, and finally publication or demo packaging. See configuration files before running a stage and the generated Python API for implementation details.
No rank guard on multi-GPU local launches
run_stages() (espnet3/utils/stages_utils.py) only special-cases the train stage's log file naming per rank (rank0 vs. per_rank in stage_log_mode); it does not gate any stage on process rank. If a multi-GPU training.yaml (devices > 1 with Lightning's default local subprocess launcher, not torchrun/srun) requests multiple stages (e.g. the default --stages all), Lightning re-executes run.py once per local rank; after trainer.fit() returns in each rank's process, every rank continues on to run the remaining stages (infer, measure, pack_model, upload_model, ...) concurrently against the same output directories. Launch multi-GPU jobs with torchrun/srun (which does not re-exec run.py per rank), or run the post-train stages as a separate single-process run.py invocation after training finishes.
Stage reference
create_dataset
Download or build datasets for your recipe.
collect_stats
Compute feature shapes and global statistics.
train
Run Lightning training with training.yaml.
infer
Write hypothesis outputs under inference_dir.
measure
Compute metrics (WER, MOS, SI-SDR, …) from inference outputs.
pack_model / upload_model
Bundle and publish a trained model to HuggingFace Hub.
pack_demo / upload_demo
Generate and upload a Gradio demo UI.
