TabStruct Documentation

TabStruct Documentation#

Prepare data with TabCamel, then generate or predict; evaluate synthetic data with TabEval and measure ML efficacy through synthetic-training prediction.
TabCamel prepares the data; TabEval supplies synthetic-data evaluators.

Choose a workflow#

Generate a table

Create synthetic tables with the original feature and target schema.

Compare generators for a research data-sharing workflow.

Generator API
An input table becomes a synthetic table Generation tutorial

Evaluate synthetic data

Assess density, privacy, ML efficacy, and structural fidelity.

Check whether a synthetic table retains relationships across columns.

TabEval integration
Four complementary evaluation dimensions Evaluation guide

Benchmark a predictor

Train classifiers or regressors and compare held-out performance.

Establish a credit classification baseline and restore its checkpoint.

Predictor API
Tabular features map to target predictions Prediction tutorial

First run#

After installation and W&B setup, run from the repository root:

python -m src.tabstruct.experiment.run_experiment \
  --pipeline prediction \
  --task classification \
  --model lr \
  --dataset credit-g \
  --device cpu \
  --save_model \
  --tags tutorial-prediction

The runner returns split-specific metrics and records best_model_path when saving a model. Continue with the tutorials or browse the model registry.

Citations#

@inproceedings{jiang2026tabstruct,
  title={TabStruct: Measuring Structural Fidelity of Tabular Data},
  author={Jiang, Xiangjian and Simidjievski, Nikola and Jamnik, Mateja},
  booktitle={The Fourteenth International Conference on Learning Representations},
  year={2026}
}

@inproceedings{jiang2025well,
  title={How Well Does Your Tabular Generator Learn the Structure of Tabular Data?},
  author={Jiang, Xiangjian and Simidjievski, Nikola and Jamnik, Mateja},
  booktitle={ICLR 2025 Workshop on Deep Generative Models in Machine Learning: Theory, Principle and Efficacy},
  year={2025}
}