Résumé
For hierarchical multivarariates outcomes, the FDA recommends the Win Ratio and Generalized Pairwise Comparisons approaches Pocock et al. [2011], Buyse [2010]. However, as far as we know, these empirical methods lack causal or statistical foundations to justify their broader use in recent studies. To address this gap, we establish causal foundations for hierarchical comparison methods. We define related causal effect measures, and highlight that depending on the methodology used to compute Win Ratio, the causal estimand targeted can be different, as proved by our consistency results, which may then lead to reversed and incorrect treatment recommendations in heterogeneous populations, as we illustrate through striking examples.
In order to compensate for this fallacy, we introduce a novel, individual-level yet identifiable causal effect measure that better approximates the ideal, non-identifiable individuallevel estimand. We prove that computing Win Ratio or Net Benefits using a Nearest Neighbor pairing approach between treated and controlled patients, an approach that can be seen as an extreme form of stratification, leads to estimating this new causal estimand measure. We extend our methods to observational settings via propensity weighting, distributional regression to address the curse of dimensionality, and a doubly robust framework. We prove the consistency of our methods, and the double robustness of our augmented estimator. These methods are straightforward to implement, making them accessible to practitioners. Finally, we validate our approach using synthetic data and on CRASH-3 [CRASH et al., 2019], a major clinical trial focused on assessing the effects of tranexamic acid in patients with traumatic brain injury.