espnet3.components.data.data_organizer.DataOrganizer
espnet3.components.data.data_organizer.DataOrganizer
class espnet3.components.data.data_organizer.DataOrganizer(train: List[DatasetConfig | Dict[str, Any] | DictConfig] | None = None, valid: List[DatasetConfig | Dict[str, Any] | DictConfig] | None = None, test: List[DatasetConfig | Dict[str, Any] | DictConfig] | None = None, preprocessor: Callable[[dict], dict] | None = None, recipe_dir: str | None = None)
Bases: object
Organizes training, validation, and test datasets into a unified interface.
This class constructs combined datasets for training and validation, and individual named datasets for testing, optionally applying a transform and preprocessor per dataset.
- Parameters:
train (Optional *[*List *[*Union [DatasetConfig , Dict *[*str , Any ] , DictConfig ] ] ]) – A list of training dataset configuration objects.
valid (Optional *[*List *[*Union [DatasetConfig , Dict *[*str , Any ] , DictConfig ] ] ]) – A list of validation dataset configuration objects.
test (Optional *[*List *[*Union [DatasetConfig , Dict *[*str , Any ] , DictConfig ] ] ]) – A list of test dataset configurations, each with a name and corresponding data source and optional transform.
preprocessor (Optional *[*Callable ]) –
A global preprocessor function or Hydra config applied after each dataset’s transform. If it is an instance of
AbsPreprocessor, each sample is passed as(uid, sample). Dict-style configs support both shared and split-specific forms:- Shared config: the same preprocessor is used for train/valid/test.
- Split config:
train,valid, andtestcan each define their own_target_block. - Shared keys outside
train/valid/testare merged into each split-specific config. This is useful for common tokenizer settings such astoken_type,token_list, andbpemodel.
Split config fallback rules are:
- If
testis missing, test uses the valid preprocessor. - If
validis missing, valid uses the train preprocessor. - If
trainis missing, train uses no preprocessor.
recipe_dir (Optional *[*str ]) – Recipe root used to resolve local dataset modules when dataset entries omit
data_srcand rely onrecipe_dir/dataset.
train
Combined dataset built from training configurations, or None if not provided.
- Type:CombinedDataset
valid
Combined dataset built from validation configurations, or None if not provided.
- Type:CombinedDataset
test_sets
Dictionary mapping test set names to DatasetWithTransform instances.
Type: Dict[str, DatasetWithTransform]
Raises:
- RuntimeError – If only one of
trainorvalidis provided. - RuntimeError – If
trainandvalidare of mismatched types (e.g., one is CombinedDataset, the other is None). - AssertionError – If
preprocessoris not callable.
- RuntimeError – If only one of
Notes
The DataOrganizer is designed to support both training and testing workflows:
- For training: provide both
trainandvalid. - For testing only: provide
testand omittrain/valid. - All three (
train,valid,test) can also be provided simultaneously.
If any of the train, valid, or test are omitted, the corresponding : attributes will be set to None or empty.
Split-specific preprocessor configs are useful when train and valid should not share augmentation behavior.
Examples
Training and validation with one shared preprocessor:
organizer = DataOrganizer( … train=training_configs, … valid=valid_configs, … preprocessor=MyPreprocessor() … ) sample = organizer.train[0]
Testing only:
organizer = DataOrganizer( … test=test_configs, … preprocessor=MyPreprocessor() … ) test_sample = organizer.test[“test_clean”][0]
Shared Hydra config for all splits: : ```python
config = { ... "target": "my_project.preprocess.BuildPreprocessor", ... "token_type": "bpe", ... "token_list": "tokens.txt", ... } organizer = DataOrganizer( ... train=training_configs, ... valid=valid_configs, ... preprocessor=config, ... )
Split-specific Hydra config with shared tokenizer settings:
: ```python
>>> config = {
... "token_type": "bpe",
... "token_list": "tokens.txt",
... "bpemodel": "bpe.model",
... "train": {
... "_target_": "my_project.preprocess.TrainPreprocessor",
... "speed_perturb_prob": 0.1,
... },
... "valid": {
... "_target_": "my_project.preprocess.ValidPreprocessor",
... },
... }
>>> organizer = DataOrganizer(
... train=training_configs,
... valid=valid_configs,
... test=test_configs,
... preprocessor=config,
... )Split-specific config with test fallback to valid: : ```python
config = { ... "train": { ... "target": "my_project.preprocess.TrainPreprocessor", ... }, ... "valid": { ... "target": "my_project.preprocess.ValidPreprocessor", ... }, ... } organizer = DataOrganizer( ... train=training_configs, ... valid=valid_configs, ... test=test_configs, ... preprocessor=config, ... )
Initialize DataOrganizer object.
<div class='custom-h4'><p>log_summary<span class='small-bracket'>(log: Logger | None = None)</span> → None</p></div>
Log a concise dataset summary.
<div class='custom-h4'><p><em>property</em> test</p></div>
Get the dictionary of test datasets.