Résumé
Background/Objectives: Advancements in genomics have greatly increased the volume of available genomic data, offering significant potential for disease understanding and personalized medicine. However, this also raises privacy risks due to the sensitive nature of genomic information. Addressing this challenge, our study explores innovative methods to minimize privacy vulnerabilities in genomic data sharing.Methods: We propose an approach based on k-mers to estimate re-identification risks. K-mers, short DNA sequences, inherently reduce privacy risks by minimizing direct associations with identifiable genomic features such as Short Tandem Repeats or transposons. By analyzing Single Nucleotide Polymorphisms (SNPs) within these k-mers against public databases like dbSNP, we identify potential privacy vulnerabilities. The novelty lies in developing a risk score based on the SNP analysis within k-mers, aiming to assess re-identification risks. We aim to develop a risk score indicating the likelihood of re-identification from shared k-mer sets, offering a novel approach to mitigate privacy concerns.Results: We expect to produce a risk score that enables precise evaluation of re-identification risks in genomic data sharing. This tool will facilitate the selective sharing of genomic data by filtering out sensitive k-mers, reducing privacy risks and supporting the safe advancement of genomics research.Conclusion: By balancing genomic research benefits with privacy protection, our study contributes to safer data sharing practices. The development of a re-identification risk metric and data filtering tool represents a significant step towards secure, open science and research reproducibility.