Résumé
Accurate egocentric body-pose estimation is essential for delivering immersive experiences in Virtual, Augmented, and Mixed Reality (VR/AR/MR) applications. A major challenge in this setting arises from the limited field of view provided by head-mounted cameras, which often leads to self-occlusion and poor visibility of certain body joints—particularly in the lower body. Current state-of-the-art methods typically perform pose estimation frame-by-frame in a view-aligned manner. While this is computationally efficient and performs well for visible joints, it struggles to estimate occluded joints due to the lack of contextual information. In this work, we explore the role of temporal information in improving joint estimation. We propose a lightweight model that refines pose predictions by leveraging a large temporal window. Our results show that incorporating such module improves the accuracy of joint estimation without incurring a large computation overhead.