Abstract
Causality is a fundamental concept in science and philosophy, and with the increasing complexity of data collection and structure, statistics plays a pivotal role in inferring causes and effects. This thesis delves into advanced causal inference methods, with a focus on policy learning, instrumental variables (IV), and difference-in-differences (DiD) approaches.The IV and DiD methods are critical tools widely used by researchers in fields like epidemiology, medicine, biostatistics, econometrics, and quantitative social sciences. However, these methods often face challenges due to restrictive assumptions, such as the IV's requirement to have no direct effect on the outcome other than through the treatment, and the parallel trends assumption in DiD, which may be violated in the presence of unmeasured confounding.In that context, this thesis introduces an innovative instrumented DiD approach to policy learning, which combines these two natural experiments to relax some of the key assumptions of conventional IV and DiD methods. To the best of our knowledge, the thesis presents the first comprehensive study of policy learning under the DiD setting. The direct policy search approach is proposed to learn optimal policies, based on the conditional average treatment effect estimators using instrumented DiD. Novel identification results for optimal policies under unmeasured confounding are established. Moreover, a range of estimators, including a Wald estimator, inverse probability weighting (IPW) estimators, and semiparametric efficient and multiply robust estimators, are introduced. Theoretical guarantees for these multiply robust policy learning approaches are provided, including the cubic rate of convergence for parametric policies and valid statistical inference with flexible machine learning algorithms for nuisance parameter estimation. These methods are further extended to the panel data setup.The majority of causal inference methods in the literature heavily depend on three standard causal assumptions to identify causal effects and optimal policies. While there has been progress in relaxing the consistency and unconfoundedness assumptions, addressing the violations of the positivity assumption has seen limited advancements.In that context, this thesis presents a novel policy learning framework that does not rely on the positivity assumption, instead focusing on dynamic and stochastic policies that are practical for real-world applications. Incremental propensity score policies, which adjust propensity scores by individualized parameters, are proposed, requiring only the consistency and unconfoundedness assumptions. This approach enhances the concept of incremental intervention effects, adapting it to individualized treatment policy contexts, and employs semiparametric theory to develop efficient influence functions and debiased machine learning estimators. Methods to optimize policy by maximizing the value function under specific constraints are also introduced.Additionally, the optimal individualized treatment regime (ITR) learned from a source population may not generalize well to a target population due to covariate shifts. A transfer learning framework is proposed for ITR estimation in heterogeneous populations with right-censored survival data, which is common in clinical studies and motivated by medical applications. This framework characterizes the efficient influence function (EIF) and proposes a doubly robust estimator for the targeted value function, accommodating a broad class of survival distribution functionals. For a pre-specified class of ITRs, a cubic rate of convergence for the estimated parameter indexing the optimal ITR is established. The use of cross-fitting procedures ensures the consistency and asymptotic normality of the proposed optimal value estimator, even with flexible machine learning methods for nuisance parameter estimation.