CLI Reference#
Run from the repository root:
python -m src.tabstruct.experiment.run_experiment --help
The parser in common/runtime/config/argument.py is the authoritative option
source. The CLI starts a complete experiment; there are no separate fit or
predict subcommands. See Tutorials for runnable sequences.
Required arguments#
Flag |
Values |
|---|---|
|
|
|
Identifier from Models Reference, chosen for the selected pipeline. |
|
TabCamel dataset name or a supported dataset source with metadata. |
The parser’s model choices combine both registries. A parser-accepted identifier
still needs a matching adapter in the chosen pipeline. unsupervision is
available only in generation.
Runtime and logging#
Flag |
Default |
Purpose |
|---|---|---|
|
|
|
|
CUDA if available, otherwise CPU |
Adapter execution device. |
|
|
Lightning accelerator selection. |
|
|
Model/runtime randomness. |
|
Empty |
One or more W&B tags for organizing runs. |
|
Disabled |
Disable logging; tag-based lookup still contacts W&B. |
|
Disabled |
Enable W&B model logging through the Lightning logger. |
|
Disabled |
Adapter/Lightning execution controls. |
Entity and project are configured in src/tabstruct/common/__init__.py.
There are no --wandb_entity or --wandb_project flags.
Dataset and splits#
Flag |
Default |
Purpose |
|---|---|---|
|
Task-dependent |
|
|
|
Fractions, or counts when greater than one. Validation splits the remaining pool. |
|
|
Split random states. |
|
Unset |
Classification filtering. |
|
|
Torch dataloader options. |
See Data and preprocessing for the preprocessing flags, schema requirements, and
split formulas. --model_specific_preprocessing enables adapter overrides;
--disable_preprocessing_tentative requests raw data.
Synthetic training data#
--curate_mode sharing replaces training rows using
--synthetic_data_path PATH or W&B lookup via --generator and
--generator_tags. --curate_ratio defaults to 1.0.
--disable_synthetic_data_validation bypasses generator-provenance lookup
for an explicit CSV; it does not disable data parsing.
Model persistence#
Flag |
Behavior (all disabled or unset by default) |
|---|---|
|
Save a pickle wrapper under the run directory. |
|
Evaluate an existing model or synthetic CSV. |
|
Load the selected checkpoint directly. |
|
Load a wrapper before the fitting workflow. |
|
Resolve a finished run’s |
--use_best_hyperparams is parsed, but its current post-processing branch is
a placeholder. It does not retrieve tuned parameters. Use the Optuna workflow
for implemented parameter search. See Workflows for persistence
constraints and methods reconstructed from reference data.
Generation controls#
Flag |
Default |
Purpose |
|---|---|---|
|
|
|
|
Unset |
Explicit sample count. |
|
|
Synthetic count relative to processed training rows. |
|
Disabled |
Generate and save CSV without metrics. |
|
Unset |
Evaluate an existing original-schema table. |
Evaluation controls#
--enable_eval_structure enables per-feature utility (default disabled).
--disable_eval_density and --disable_eval_privacy set the corresponding
flags to false. --enable_full_split_eval enables density/privacy flag
handling on every split. Structure still runs only on test. Without full split
evaluation, training density/privacy are always evaluated and validation metrics
are empty. See Evaluation before choosing these flags.
Training and tuning#
Flag |
Default |
Purpose |
|---|---|---|
|
|
Requested training budget; at least one epoch is enforced. |
|
|
Capped by processed training size. |
|
|
|
|
Unset |
|
|
|
|
|
|
Metric monitored by supporting Lightning models. |
|
Disabled |
Run the adapter’s search space. |
|
|
Maximum trial count. |
|
|
Split IDs evaluated within each tuning trial. |
|
|
Processes per tuning trial; reduced to one under multi-rank launch. |
|
|
|
|
|
Validation metric used by tuning. |
|
Disabled |
Disable the median pruner. |
Training flags affect only adapters that consume them. --help also lists
scheduler-specific values, gradient clipping, and Lightning logging/validation
cadence. Runtime-derived values are resolved after data preparation.
Distributed execution#
The runtime recognizes torchrun ranks and disables W&B logging on nonzero
ranks. The helper keeps inference collective and assigns serial metrics and
artifact writes to rank zero. Distributed support is adapter-specific; this
behavior does not make every registered generator or predictor support DDP.