Evaluation#
TabStruct evaluates synthetic tables across four complementary dimensions. TabEval supplies synthetic-data evaluators, and the prediction pipeline measures ML efficacy on held-out real rows after training on synthetic data.
Dimension |
Question |
|---|---|
Density estimation |
Does the synthetic table preserve the real data distribution? |
Privacy preservation |
How much information might synthetic records reveal about training rows? |
ML efficacy |
Can a model trained on synthetic rows predict held-out real outcomes? |
Structural fidelity |
Does the synthetic table preserve relationships across columns? |
These dimensions answer different questions. Global utility adds a structural perspective that is independent of a single designated target. The following sections describe the current evaluator configuration and split policy for reproducible comparisons.
Structural fidelity and global utility#
The paper introduces global utility: use each column as a prediction target in turn and assess what the remaining columns predict about it. This probes relationships throughout a table without requiring a ground-truth causal graph.
The generation runner delegates structural evaluation to TabEval’s
UtilityPerFeature through compute_structure_metrics. It includes
features and the original supervised target in full_col_list_eval and uses
an ordinal-encoded view. Enable it with --enable_eval_structure.
Returned keys are
prefixed structure_ and incorporate the installed TabEval metric name and
its result keys. Inspect the per-feature results as well as any aggregates.
Evaluator configuration#
Dimension |
Implementation |
Input and summary keys |
|---|---|---|
Density |
|
Original columns plus SDMetrics metadata for low-order metrics;
one-hot columns for high-order metrics. Keys start with |
Privacy |
|
Distance to closest record on original columns with metadata;
at most 3,000 samples. Keys start with |
Structure |
|
Ordinal view of all columns; enabled explicitly.
Keys start with |
ML efficacy |
Prediction with |
Train a predictor on synthetic rows, evaluate on held-out real rows. See Data and preprocessing. |
A DCR measurement describes record proximity; it does not provide a formal privacy guarantee. Metric schemas come from the installed TabEval version.
Evaluation splits#
The ordinary generation policy is fixed by split:
Split |
Evaluators |
|---|---|
|
Density and privacy, regardless of the disable flags. |
|
No generation metrics (an empty dictionary). |
|
Structure if |
With --enable_full_split_eval, density and privacy follow their enable/
disable flags on all three splits. Structure still runs only on test:
compute_tabeval_metrics also checks split == "test". Tuning enables
full split evaluation automatically so supported validation metrics can be used
for model selection.
For a structure-only evaluation that respects disable flags, use:
bash docs/tutorial/example_scripts/generation/eval.sh \
logs/<configured-project>/<run-id>/synthetic_samples.csv \
--enable_full_split_eval \
--disable_eval_density \
--disable_eval_privacy
--generation_only bypasses the complete metric suite and saves the CSV.
--model real and --model real-test use real training or test tables as
reference baselines without fitting a generator.
Prediction metrics#
Classification reports balanced_accuracy, F1_weighted,
precision_weighted, recall_weighted, AUROC_weighted, ECE, and
cross_entropy_loss. Probability columns follow the encoded class order.
Regression reports rmse, mse, and r2 in the restored target space.
Prediction evaluates all three splits. W&B also records weighted class-recall statistics for classification and feature-selection output for supported predictors.
Read and compare outputs#
run_experiment returns a dictionary with train_metrics,
valid_metrics, and test_metrics for evaluated runs. W&B stores scalar
keys as <split>_metrics/<metric>. Generation-only runs return an empty
metric dictionary.
Keep dataset versions, split IDs, row counts, preprocessing, and evaluator configuration constant when comparing results. Record generator parameters and installed package versions. Use real-data references to interpret scores; do not infer structural fidelity from a single downstream classification score.