Flow Matching for Structured Data Synthesis and Imputation
Recent advances in deep generative models have concentrated on media synthesis. Structured data of scientific and applied origin, such as measurement tables, spectra, and sensor time series, are of a different nature: variables obey hard domain constraints, state spaces are often discrete or mixed, and reproducing the data distribution, including its tails, matters as much as minimizing error. This qualification text investigates Flow Matching, a recent generative framework still little applied to such data, as a tool for synthesis and imputation, through three works. The first applies Discrete Flow Matching to generate chess configurations conditioned on strategic labels, where a post-processing step raises full-label accuracy. The second combines continuous Flow Matching with Kolmogorov-Arnold Networks to impute INMET meteorological series under geographic holdout: it outperforms deterministic reference imputers in distributional fidelity across all six evaluated variables and quantifies uncertainty, at the cost of higher pointwise error on high-variance variables. The third and most mature work introduces a stochastic classifier-based pruning step that aligns the distributions of synthetic soil data from the LUCAS dataset a posteriori: it reduces adversarial AUC across all four evaluated variants and outperforms TabSyn, a state-of-the-art reference in tabular synthesis, at one to two orders of magnitude lower computational cost, while preserving predictive utility. The first and third works further suggest a hypothesis to be explored: refining already generated samples can be more effective and cheaper than embedding the same signals into generator training. The text is completed by the theoretical background, the literature review, and the work plan up to the defense, which prioritizes the second work's experimental consolidation, including comparison with state-of-the-art generative imputers.