Résumé
Synthetic data generation has emerged as a promising avenue for preserving the privacy of sensitive information, such as medical records, while enabling meaningful data analysis. Yet, the landscape of available techniques is highly fragmented: numerous methods have been proposed, each excelling under specific conditions, and their evaluation is often inconsistent across studies. Fidelity to the original data is frequently prioritized, whereas privacy preservation receives comparatively less attention. In this work, we introduce a versatile ensemble framework; a super learner for synthetic data generation; that adaptively selects and combines methods based on the characteristics of the target dataset. We demonstrate this approach on medium-sized datasets (hundreds of samples), addressing challenges in sharing small biomedical datasets. We systematically assess the influence of different fidelity metrics, including the Energy distance and the Wasserstein distance, on the quality of generated data. Experimental results demonstrate that our approach consistently matches or outperforms the strongest individual method in terms of utility, while delivering enhanced privacy protection.