Quickstart#

Requirements#

Use Python 3.10 or later and a Git checkout of the public repository. Run all commands from the repository root: the runtime discovers logs/ by walking upwards to a .git directory.

Dataset downloads and online W&B logging require network access. Classical baselines can run on CPU. Deep generators and pretrained models may require GPU resources, model downloads, and model-specific dependencies.

Installation#

git clone https://github.com/SilenceX12138/TabStruct.git
cd TabStruct
conda create -n tabstruct python=3.10.18
conda activate tabstruct
bash scripts/utils/install.sh

The script installs the package, Mostly AI, DGL and PyG extensions, and PyTorch 2.2.2 with CUDA 12.1 wheels. Use it in a compatible Linux environment. For a prepared environment, install the checkout directly:

python -m pip install -e .

This installs declared dependencies; it does not run the additional wheel installation steps in install.sh. A successful CPU baseline does not establish that every GPU generator dependency is installed.

Configure logging#

Set WANDB_ENTITY and WANDB_PROJECT in src/tabstruct/common/__init__.py to a workspace you can write to:

WANDB_ENTITY = "your-wandb-entity"
WANDB_PROJECT = "tabstruct"
wandb login

The checked-in values select a development workspace. Configure your own workspace before running examples. Logging is online by default, with local W&B files under logs/wandb. Use --disable_wandb for a local check; W&B run lookup by tags still requires access to the configured project.

Run a classifier#

python -m src.tabstruct.experiment.run_experiment \
  --pipeline prediction \
  --task classification \
  --model lr \
  --dataset credit-g \
  --device cpu \
  --save_model \
  --tags tutorial-prediction

--task, --model, and --dataset are required. TabCamel loads the named dataset and may download it on first use. You do not need to create a CSV for this example.

Expected outputs#

The terminal shows data preparation, training, and evaluation stages. W&B summaries contain train_metrics/, valid_metrics/, and test_metrics/ keys, including balanced_accuracy, F1_weighted, and AUROC_weighted. Exact scores depend on the dataset and installed model versions.

With --save_model, the runner saves logs/<configured-project>/<run-id>/lr.pkl and records best_model_path. A Python call to run_experiment also returns the metric dictionary.

Restore the model#

Use the saved path from the preceding run with the same dataset, task, model, split IDs, and preprocessing configuration:

bash docs/tutorial/example_scripts/prediction/eval.sh \
  logs/<configured-project>/<run-id>/lr.pkl

The script evaluates the restored model on all three splits. It does not expose a separate prediction-file export command.

Next steps#

Follow Tutorials for CSV generation and evaluation, Data and preprocessing for split semantics, and CLI Reference for the complete public option groups.