Abstract
In switchback experiments with unequal cluster sizes, outcome dispersion across randomization units inflates estimator variance and limits statistical power. Standard control-using-prediction-as-covariate (CUPAC) adjustment may be suboptimal for variance reduction in this setting, because in its basic form it targets overall predictive accuracy and does not distinguish between the components of variance that vary across the randomization units of switchback experiments and those that vary across individual observations, even though these components contribute unequally to estimator variance. We propose a power-optimal variance reduction methodology via CUPAC that balances prediction of the noise between and within randomization units to achieve maximum statistical power. The methodology utilizes the framework for decomposition of the variance of the treatment-effect estimator for switchback experiments, and adapts both the outcome prediction and the analysis-time residualization to minimize the treatment-effect variance. The study first develops the theoretical framework for the power-optimal CUPAC. We then validate the theoretical framework through an extensive Monte Carlo simulation study. Finally, we discuss the practical considerations of the proposed methodology, including its potential efficiency gains and limitations.
Abstract
Switchback experiments---in which treatment is assigned at the level of a cluster crossed with a time period---are widely used in marketplace and platform settings, yet no closed-form power formula exists for them. We fill this gap by deriving a closed-form, multi-level asymptotic variance approximation for the individual-level OLS estimator, facilitating power budgeting. Using this formula, we reveal a structural floor on statistical power: while idiosyncratic noise vanishes with observation density, macro-level shocks are multiplicatively penalized by cluster size imbalance. We confirm through analytical derivations and Monte Carlo simulations that the formula is exact across typical parameters and serves as a mathematically conservative upper bound in extreme boundary regimes. We study three methodological applications. First, we prove that advanced assignment designs like stratification only partially eliminate the penalty of cluster size imbalance on power. Second, we demonstrate that variance reduction techniques targeting macro-level shocks yield disproportionately greater efficiency gains than those targeting residual noise. Third, we formalize the finite-sample power trade-offs between individual-level and cell-level estimators.
Abstract
Switchback experiments and other clustered randomized designs are widely used on online platforms, but the clustered, time-dependent nature of these designs can make standard variance reduction methods behave differently than in standard A/B tests. We evaluate design-aware variance reduction methods for switchbacks---CUPED, CUPAC (ML-based covariate adjustment), and doubly robust (DR) estimators---relative to a baseline switchback analysis with cluster-robust standard errors. Through a hierarchical simulation framework that varies key regime parameters—number of clusters, cluster-size imbalance, within-cluster autocorrelation, carryover, and predictive signal strength—we evaluate validity (false positive rate and confidence interval coverage) and efficiency (standard error reduction, power, and minimum detectable effect as a function of run length). We also include a sensitivity analysis for cross-cluster spillovers to quantify bias and inference degradation under mild interference. The primary outcome is a practitioner-oriented regime map: when CUPED, CUPAC, or DR are most beneficial, and when time and cluster dependence and finite-cluster effects limit improvements.
Presented at:
American Causal Inference Conference 2026 (Society for Causal Inference)
Symposium on Data Science and Statistics 2026 (American Statistical Association)
Abstract
Controlled-experiment Using Pre-Experiment Data (CUPED) reduces the variance of online experiments by adjusting the outcome for a pre-experiment covariate, and is typically deployed with a single lagged metric. Recent practical frameworks extend this idea by predicting the in-experiment outcome with a gradient-boosted model over many pre-experiment attributes, raising the covariate's explanatory power and the resulting variance reduction (Brazhenko, 2025). We formalize and generalize this direction as multivariate covariate enrichment. Rather than modifying how a model learns---by re-engineering the machine-learning loss function, which is fragile and prone to degenerate solutions---we enrich what it learns from, expanding the covariate set fed to a control-variate adjustment. We establish that enrichment is monotone in population, since the standard error scales by (1-R2)^1/2 and the coefficient of determination $R^2$ is non-decreasing in the covariate set, and that cross-fitting renders the realized reduction unbiased. We then connect the method to the multi-level variance structure of clustered and switchback experiments, demonstrating that covariate enrichment yields disproportionately larger efficiency gains when the added covariates explain macro-level shocks---spatial, temporal, and cluster-temporal interaction components---because those components are amplified by cluster size imbalance. In particular, enriching the covariate set with multiple historical lags captures the recurring cluster-by-time interaction that single-lag CUPED leaves unaddressed and that no fixed assignment design can balance.
Abstract
Can more economically successful parents have stronger influence on their kids? This paper tests empirically whether economic achievement of parents is associated with higher similarity in norms and values between them and their children. I use German Socio-Economic Panel (SOEP) survey data and show that parental income strengthens transmission of trust and democratic preferences from parents to their children. Both cross-sectional differences and changes of parental income over time explain differences in transmission of values and norms. I also show that results are robust to various alternative interpretations and sample restrictions.
Presented at:
GLO-Berlin-2024
47th EBES Conference
Abstract
Does increasing the time allocated to a predictive task improve quality of individual predictions and accuracy of beliefs? In our online experiment, participants observe data from a bivariate data-generating process with asymmetrically correlated variables and make predictions for new data points. When participants choose the time they devote to the predictive task endogenously, it is positively correlated with the accuracy of their predictions; this finding remains robust after controlling for individual characteristics. Moreover, exogenous variables that have a relatively simple relationship with the outcome variable are identified more accurately when participants spend more time on the task, whereas this does not hold for variables with a more complex relationship. However, exogenously increasing the time required for a predictive task does not affect the accuracy of predictions. Taken together, the findings suggest that time is a significant predictor of the accuracy of individual predictions and beliefs; nonetheless, exogenous time constraints appear ineffective in online experiments.
Awarded: Orlando Bravo Center for Economic Research Award
Abstract
This paper studies historical origins and persistence of culture of morality in places where enforcement of formal institutions is restricted by geographical factors. I develop a measure of local isolation due to irregularities of land surface and show that local isolation is associated with morality-related ethnic folklore. Moreover, importance of morality from World Values Survey is robustly positively correlated with local isolation on subnational level. Finally, I show that differences in strength of individual morality values persist on population of second generation migrants surveyed by European Social Survey. I provide an extensive set of robustness checks to address alternative explanations and bring additional evidence on possible mechanisms driving these differences.
Presented at:
The 6th International Conference on European Economics and Politics
EconTR 2024
Accepted for:
3rd Bolzano Workshop on Historical Economics
Abstract
This paper studies the emergence of motivated beliefs in response to different information structures. I employ a simple urn-draw framework for belief elicitation, randomizing whether past ball draws provide information about future utility-relevant draws. An additional treatment arm increases the payoff per rewarded ball to explore the effect of stake size on belief distortions. The study finds no evidence of motivated beliefs, as participants' beliefs are, on average, the same across the treatment and control groups. Moreover, the distributions of participants' beliefs do not differ between the control and treatment groups. These results hold both conditional and unconditional on the ball color in the first round. Finally, increasing the payoff of a future ball draw does not induce belief distortions either.
Awarded: Orlando Bravo Center for Economic Research Award