Abstract
Measuring chromatin-RNA contacts is the first step to understand their role(s) in gene expression. Several technologies have been developed to detect genome-wide RNA-DNA interactions, each with different protocols and computational analysis methods. While these methods use various statistical tests, they all include RNA-focused corrections to account for confounding factors such as spurious trans-chromosomal mRNA-DNA interactions, RNA abundance and distance between the RNA-emitting gene and RNA-receiving DNA region. However, the RNA-DNA interaction counts can also be biased by the amount of DNA regions that can be sequenced and which varies along the chromosomes due notably not only to Copy Number Variations (CNVs) but also DNA replication timing (RT). DNA is replicated in a precise spatiotemporal program, which is local, cell-type specific and conserved in evolution, that causes some genomic regions to be duplicated earlier than others during the cell cycle. Consequently, depending on the number of dividing cells, early-replicating regions are more likely to be overrepresented in sequencing data, especially in proliferating cells, leading to uneven DNA coverage across the genome. Using RADICL-seq data1 from the FANTOM consortium across multiple cell types, we show that, despite RNA-centric normalization, chromatin-RNA contact counts remain correlated with both CNVs and replication timing. Since both CNVs and RT can be inferred from Whole Genome Sequencing (WGS) 2,3, we propose this data to estimate and correct for local DNA abundance. To evaluate our approach's validity, we use ATAC-seq data, which strongly correlates with RADICL-seq signals. We demonstrate that the newly predicted peaks are more biologically relevant.