TabStruct Documentation#
Choose a workflow#
Generate a table
Create synthetic tables with the original feature and target schema.
Compare generators for a research data-sharing workflow.
Generator APIEvaluate synthetic data
Assess density, privacy, ML efficacy, and structural fidelity.
Check whether a synthetic table retains relationships across columns.
TabEval integrationBenchmark a predictor
Train classifiers or regressors and compare held-out performance.
Establish a credit classification baseline and restore its checkpoint.
Predictor APIFirst run#
After installation and W&B setup, run from the repository root:
python -m src.tabstruct.experiment.run_experiment \
--pipeline prediction \
--task classification \
--model lr \
--dataset credit-g \
--device cpu \
--save_model \
--tags tutorial-prediction
The runner returns split-specific metrics and records best_model_path
when saving a model. Continue with the tutorials or
browse the model registry.
Citations#
@inproceedings{jiang2026tabstruct,
title={TabStruct: Measuring Structural Fidelity of Tabular Data},
author={Jiang, Xiangjian and Simidjievski, Nikola and Jamnik, Mateja},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026}
}
@inproceedings{jiang2025well,
title={How Well Does Your Tabular Generator Learn the Structure of Tabular Data?},
author={Jiang, Xiangjian and Simidjievski, Nikola and Jamnik, Mateja},
booktitle={ICLR 2025 Workshop on Deep Generative Models in Machine Learning: Theory, Principle and Efficacy},
year={2025}
}