Abstract
Distributed infrastructures for computation and analytics are now evolving towards an interconnected ecosystem allowing complex applications to be executed on the Edge-to-Cloud continuum. Understanding and optimizing end-to-end performance in such a complex continuum is challenging. One crucial challenge is accurately reproducing relevant behaviors of a given application workflow and representative settings of the physical infrastructure underlying this complex continuum. This thesis is a conceptual and practical contribution to the Edge-to-Cloud continuum, proposing and applying methodologies in novel environments. Our methodologies aim at overcoming the complexity of understanding and optimizing Edge-to-Cloud workflows. As such, they enable reproducible experiment design, application optimization, efficient workflow provenance capture, and costeffective experiment reproducibility. We validated our proposal by first developing the E2Clab framework that supports the complete analysis cycle of an application on the Edge-to-Cloud Continuum and then using E2Clab to optimize Pl@ntNet, a global plant identification application. Large-scale experimental validation on Grid’5000 shows that our methodology has proven helpful for understanding and improving the performance of Pl@ntNet.