Abstract
The comparison of protein sequences has been for long a very effective tool in producing biological knowledge. It was initially based on the alignment of sequences, that is to say organizing the set of sequences in columns (of a spreadsheet) of sites which have evolved from a common site of the ancestral sequence. Alignments are generally obtained by minimizing an evolution or an edition cost. Sequence comparisons are now often performed without alignments by comparing the-mer compositions of the sequences. We present here the most popular methods used by biologists to compare sequences and place emphasis on an approach to augment the alphabet of a set of sequences in order to ease their comparison. The family of DNA topoisomerases, a set of ancient proteins whose history can be traced back 4 billion years, is used to illustrate this approach.