Developer guide#
Read the implemented workflow#
Start with run_experiment and follow its calls to setup_runtime and
PipelineHelper.run_pipeline. The selected helper runs
benchmark_model in this order:
Resolve the model identifier to an adapter class.
Prepare the data and resolve tentative preprocessing/training arguments.
Create or restore the wrapper and fit it when appropriate.
Run inference, save generation output, and compute split metrics.
The Overview maps public modules to these responsibilities.
Add a predictor#
Use BaseSklearnPredictor for an estimator with fit, predict, and
predict_proba interfaces, or BaseLitPredictor for Lightning training.
A concrete adapter supplies its estimator, parameter definitions, and any
special preprocessing. Regression probability prediction returns None.
Use an existing wrapper such as
src/tabstruct/prediction/models/sklearn/Linear.py as the pattern. Then
register its identifier in predictior_list (the implementation spelling)
in src/tabstruct/common/__init__.py and add a matching branch to
PredictorHelper._model_handler.
Callers use public fit, predict, and predict_proba methods.
Underscored hooks implement adapter behavior. Validate probability shape
(n_rows, n_classes) and encoded class ordering before metric computation.
Add a generator#
Choose an existing base strategy in BaseGenerator.py: joint, conditional,
class-focused, TabEval-backed, or mixed/Lightning generation. Implement its
required fitting and generation hooks, plus parameter definitions. Use the
concrete bases to reuse class-allocation and conditional sampling behavior.
Register the identifier in generator_list and add the corresponding
GeneratorHelper._model_handler branch. generate() reads counts and
class distribution from runtime arguments; CSV restoration and evaluation
remain in the helper. Check that exported columns match the original dataset
and that repeated generation calls use the intended seed behavior.
Parameter and preprocessing contracts#
Define ordinary model parameters in _define_default_params and an optional
Optuna search space in _define_optuna_params.
get_model_specific_scaler_config() returns context,
feature_scaler, and target_scaler dictionaries. The helper applies
these when --model_specific_preprocessing is enabled.
AddOnlyNamespace retains parsed values and adds derived runtime fields.
Avoid mutating an existing argument through this interface. Pass new adapter
parameters explicitly and preserve feature/target transform state for restore.
Proportionate validation#
For a new adapter, check loading, fitting, output schema, inference, and save/restore where supported on a small realistic dataset. Compare metrics using the same split IDs and preprocessing. Run the public style checker from the repository root:
bash scripts/utils/style_check.sh
Build the documentation#
The site uses Sphinx with the Book Theme. Install the documentation dependencies in your working environment if necessary:
python -m pip install sphinx myst-parser sphinx-book-theme
python -m sphinx \
-b html \
-a \
-W \
--keep-going \
docs/wiki \
docs/wiki/_build/html
python -m http.server \
8081 \
--bind 127.0.0.1 \
--directory docs/wiki/_build/html
Open http://127.0.0.1:8081/. The public documentation deployment also uses
make html inside docs/wiki. Keep existing page routes and links working,
include new pages in the toctree, and verify mobile navigation, search, theme
switching, code copy controls, and figure/table readability.