Résumé
Branched glycerol dialkyl glycerol tetraethers (brGDGTs) are critical molecular biomarkers for the quantitative reconstruction of past environments, ambient temperature, and pH across various archives. However, numerous issues persist that limit their application. The distribution of brGDGTs varies significantly based on provenance, resulting in biases in environmental reconstructions that rely on fractional abundances and derived indices, such as MBT5ME′. This issue is especially significant in shallow lakes, wetlands, and peatlands, where ecosystems are sensitive to diverse environmental and climatic factors. Recent advancements, such as machine learning techniques, have been developed to identify changes in provenance; however, these techniques are insufficient for detecting mixed environments. The probability estimates derived from five machine learning algorithms are employed here to detect provenance changes in brGDGT downcore records and to identify periods of mixed provenance. A new global modern database (n=2031) was compiled to train, validate, test, and apply these algorithms to two sedimentary records. Our findings are corroborated by pollen, non-pollen palynomorphs, and X-ray fluorescence (XRF) obtained from the same sedimentary core sequence. These microfossil and geochemical proxies are utilized to discuss changes in provenance, hydrology, and ecology that influence brGDGT provenance. Probability estimates derived from random forest with a sigmoid calibration are most effective in detecting changes in brGDGT provenance. Minor changes in the relative contributions of brGDGT provenance can significantly influence the distribution of brGDGT, especially regarding the MBT5ME′ index.