Résumé
Contrastive learning requires creating distinct views of the same input while preserving important details. Standard audio augmentations, such as random resized crops, can remove species-specific characteristics. We propose mixing vocalizations as a domain-agnostic data augmentation, which preserves the unique features of the species of interest while forming a distinct view. This simple strategy allows contrastive learning to capture species-specific features in bird vocalizations from unlabeled data.