Evaluation#

TabStruct evaluates synthetic tables across four complementary dimensions. TabEval supplies synthetic-data evaluators, and the prediction pipeline measures ML efficacy on held-out real rows after training on synthetic data.

Four evaluation dimensions#

Dimension

Question

Density estimation

Does the synthetic table preserve the real data distribution?

Privacy preservation

How much information might synthetic records reveal about training rows?

ML efficacy

Can a model trained on synthetic rows predict held-out real outcomes?

Structural fidelity

Does the synthetic table preserve relationships across columns?

These dimensions answer different questions. Global utility adds a structural perspective that is independent of a single designated target. The following sections describe the current evaluator configuration and split policy for reproducible comparisons.

Structural fidelity and global utility#

The paper introduces global utility: use each column as a prediction target in turn and assess what the remaining columns predict about it. This probes relationships throughout a table without requiring a ground-truth causal graph.

The generation runner delegates structural evaluation to TabEval’s UtilityPerFeature through compute_structure_metrics. It includes features and the original supervised target in full_col_list_eval and uses an ordinal-encoded view. Enable it with --enable_eval_structure.

Returned keys are prefixed structure_ and incorporate the installed TabEval metric name and its result keys. Inspect the per-feature results as well as any aggregates.

Evaluator configuration#

Generation evaluators#

Dimension

Implementation

Input and summary keys

Density

LowOrderMetrics / HighOrderMetrics

Original columns plus SDMetrics metadata for low-order metrics; one-hot columns for high-order metrics. Keys start with density_.

Privacy

DCR

Distance to closest record on original columns with metadata; at most 3,000 samples. Keys start with privacy_.

Structure

UtilityPerFeature

Ordinal view of all columns; enabled explicitly. Keys start with structure_.

ML efficacy

Prediction with --curate_mode sharing

Train a predictor on synthetic rows, evaluate on held-out real rows. See Data and preprocessing.

A DCR measurement describes record proximity; it does not provide a formal privacy guarantee. Metric schemas come from the installed TabEval version.

Evaluation splits#

The ordinary generation policy is fixed by split:

Without --enable_full_split_eval#

Split

Evaluators

train

Density and privacy, regardless of the disable flags.

valid

No generation metrics (an empty dictionary).

test

Structure if --enable_eval_structure is supplied.

With --enable_full_split_eval, density and privacy follow their enable/ disable flags on all three splits. Structure still runs only on test: compute_tabeval_metrics also checks split == "test". Tuning enables full split evaluation automatically so supported validation metrics can be used for model selection.

For a structure-only evaluation that respects disable flags, use:

bash docs/tutorial/example_scripts/generation/eval.sh \
  logs/<configured-project>/<run-id>/synthetic_samples.csv \
  --enable_full_split_eval \
  --disable_eval_density \
  --disable_eval_privacy

--generation_only bypasses the complete metric suite and saves the CSV. --model real and --model real-test use real training or test tables as reference baselines without fitting a generator.

Prediction metrics#

Classification reports balanced_accuracy, F1_weighted, precision_weighted, recall_weighted, AUROC_weighted, ECE, and cross_entropy_loss. Probability columns follow the encoded class order. Regression reports rmse, mse, and r2 in the restored target space.

Prediction evaluates all three splits. W&B also records weighted class-recall statistics for classification and feature-selection output for supported predictors.

Read and compare outputs#

run_experiment returns a dictionary with train_metrics, valid_metrics, and test_metrics for evaluated runs. W&B stores scalar keys as <split>_metrics/<metric>. Generation-only runs return an empty metric dictionary.

Keep dataset versions, split IDs, row counts, preprocessing, and evaluator configuration constant when comparing results. Record generator parameters and installed package versions. Use real-data references to interpret scores; do not infer structural fidelity from a single downstream classification score.