How can restricted randomization to achieve covariate balance lead to imbalance in unobserved variables?

How can restricted randomization to achieve covariate balance lead to imbalance in unobserved variables?

Manage alerts

Loading saved threads...

retodomax · External communityPost link
External question — Cross Validated Stack Exchange Author: retodomax Original post: https://stats.stackexchange.com/questions/649609 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. In literature, designing an experiment is considered a trade-off between covariate balance and robustness. For example Harshaw et al. (2024) writes In an effort to make the estimators more precise, experimenters sometimes restrict the randomization to achieve covariate balance between treatment groups. A concern with this approach is that unobserved characteristics , including potential outcomes, may not be similar between the groups even if the observed characteristics are. [...] Experimenters must weigh the robustness granted by randomness against possible gains in precision granted by balancing prognostically important covariates. This implies that methods to reduce covariate imbalance (blocking, matched pair, rerandomization, ...) can accidentally lead to larger differences in unobserved variables then a complete randomized design. Question: I don't understand how balancing a covariate could make another unobserved variable non-similar. Can you make a toy example where balancing for one covariate makes unobserved variable non-similar? My thinking: If we balance for covariates with restricted randomization, any positively/negatively correlated unobserved variable would also become better balanced. For uncorrelated variables, they would not be better balanced but also not worse than with a completely randomized design.
Quote
Report
Demetri Pananos · External communityPost link
External answer — Cross Validated Stack Exchange Author: Demetri Pananos Original post: https://stats.stackexchange.com/a/649881 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. A few things. First, my read of the passage is not that restricted randomization leads to imbalance in unobserved covariates. Rather, despite the researcher's best intentions, restricted randomization does not protect against imbalances in unobserved confounders. Second, this is true and it is of little consequence. In his paper Seven myths of randomisation in clinical trials [1], Stephen Senn lists the second myth as "Balance of prognostic factors is necessary for valid inference". Senn begins this section by discussing a game of dice which he uses as an analogy to randomized experiments. While I won't go through that here, the relevant material is as follows It is not necessary for the groups to be balanced. In fact, the probability calculation applied to a clinical trial automatically makes an allowance for the fact that groups will almost certainly be unbalanced , and if one knew that they were balanced, then the calculation that is usually performed would not be correct. (Emphasis Senn) In short, the statistical calculations we perform are intended to account for possible imbalances between groups. Additionally, any observed imbalance is not a function of the design (were we to perform the experiment again, we might achieve balance or imbalance in the other direction), and so the two groups are balanced in expectation which is the more important point. In any case, the passage you quote is a non-sequitur. Even though randomization my lead to imbalance on unobservables, the calculations we perform explicitly allow for this and hence do not make imbalance -- actual or otherwise -- a problem. References Senn, Stephen. "Seven myths of randomisation in clinical trials." Statistics in medicine 32.9 (2013): 1439-1450.
Quote
Report
kjetil b halvorsen · External communityPost link
External answer — Cross Validated Stack Exchange Author: kjetil b halvorsen Original post: https://stats.stackexchange.com/a/649894 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. You say If we balance for covariates with restricted randomization, any positively/negatively correlated unobserved variable would also become better balanced. For uncorrelated variables, they would not be better balanced but also not worse than with a completely randomized design. This is a linear intuition! I will denote the observed covariate by $x$ and the unobserved by $z$ . If there is a linear correlation between $x$ and $z$ , if one is balanced the other will be. But the association between covariates need not be linear. For one example, where we cover $x$ uniformly but the relation to $z$ is very curved, so the corresponding $z$ sample is very unbalanced. For another example, suppose there is a "triangular correlation" (explained by the plot), also here $x$ is covered uniformly but $z$ is very unbalanced. Random sampling cannot be a complete solution to this, it will probably make $x$ with less balance, but then $z$ more balanced. The referenced paper (which I will not repeat here) makes a more mathematical argument. It is in the context of the potential-outcomes model and Horwitz-Thompson estimation, analyzing its mean squared error. It defines robustness as having good precision whatever is the potential outcomes, and shows that a measure of robustness is minimized by random sampling. Then, using restricted randomization to get better balance, minimizing the measure of robustness in the resulting restricted space, necessarily must give a higher minimum, for the same reason as in Why is sum of squared residuals non-increasing when adding explanatory variable? They also refer an old debate in statistics, between randomization on the one hand and systematic design on the other. Referring to an old paper by B Efron, Forcing a sequential experiment to be balanced where the same phenomenon is called accidental bias , randomization then leads to a freedom from accidental bias, while balancing can lead to accidental bias
Quote
Report
retodomax · External communityPost link
External answer — Cross Validated Stack Exchange Author: retodomax Original post: https://stats.stackexchange.com/a/651084 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. After some thought, I have a toy example where rerandomization (to achieve balance in the observed covariate $x$ ) leads to imbalance in another variable $z$ . Imagine we need to randomize a binary treatment to 20 experimental units with covariate $x$ We want to have similar average value of $x$ across the two treatments. There are ${{20}\choose{10}} = 184\ 756$ different ranomizations possible. We randomly sample 1000 randomizations and choose the one with the lowest covariate imbalance Notice the large outlier on the right. By choosing a randomization with low covariate imbalance, we will always implicitly 'compensate' for such an outlier by assigning the same treatment to many experimental units at the other end. This can lead to imbalance in any unobserved covariate $z$ , which has a non-linear relationship with $x$ (as similarly pointed out by @kjetil b halvorsen ). In the example below, we assume $z = 1/x$ as non-linear relationship. Instead of assigning the treatments with rerandomization, we could also use another famous restricted randomization scheme: Blocking . Pairs of experimental units with similar $x$ values are created and within each of those blocks one unit is randomized to Control and the other to the Treatment condition. Below, a potential treatment allocation This type of restricted randomization might lead to a less perfect balance in $x$ but does not suffer from the same 'outlier compensation' problem as rerandomization. Can someone can give a toy example where restricted randomization via blocking leads to unobserved covariate imbalance? I would happily accept it as the best solution to the question. Edit 2024-08-12 Below the code for the last plot (as requested in the comments). I used ggbeeswarm::geom_quasirandom() to avoid overplotting in the dot plots and cowplot::axis_canvas() to add margin plots (similarly as in this post ) library(tidyverse) library(ggbeeswarm) library(cowplot) theme_set(theme_bw()) ## Log-Normal distribution with outlier set.seed(18) x_orig <- rnorm(n = 20) dat <- tibble(x_orig = x_orig, x = exp(x_orig)) %>% mutate(z = 1/x) ## Treatment randomization set.seed(1) dat <- dat %>% mutate(x_rank = rank(x), block = as.factor(ceiling(x_rank/2))) %>% group_by(block) %>% mutate(treatment = sample(c("Control", "Treatment"))) %>% ungroup() ## Treatment group average agg_dat <- dat %>% group_by(treatment) %>% summarise(x_mean = mean(x), z_mean = mean(z)) ## Plot pmain <- dat %>% ggplot(aes(x = x, y = z, color = treatment)) + geom_function(fun = \(x) 1/x, inherit.aes = FALSE) + geom_point() xbox <- axis_canvas(pmain, axis = "x") + geom_quasirandom(data = dat, aes(x = x, y = treatment, color = treatment)) + geom_crossbar(aes(x = x_mean, xmin = x_mean, xmax = x_mean, y = treatment, color = treatment), data = agg_dat) + scale_y_discrete() ybox <- axis_canvas(pmain, axis = "y") + geom_quasirandom(data = dat, aes(x = treatment, y = z, color = treatment)) + geom_crossbar(aes(y = z_mean, ymin = z_mean, ymax = z_mean, x = treatment, color = treatment), data = agg_dat) + scale_x_discrete() p1 <- insert_xaxis_grob(pmain, xbox, grid::unit(1, "in"), position = "top") p2 <- insert_yaxis_grob(p1, ybox, grid::unit(1, "in"), position = "right") ggdraw(p2)
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: retodomax Source score (net votes, not local likes): 12 Original post: https://stats.stackexchange.com/questions/649609 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. In literature, designing an experiment is considered a trade-off between covariate balance and robustness. For example Harshaw et al. (2024) writes In an effort to make the estimators more precise, experimenters sometimes restrict the randomization to achieve covariate balance between treatment groups. A concern with this approach is that unobserved characteristics , including potential outcomes, may not be similar between the groups even if the observed characteristics are. [...] Experimenters must weigh the robustness granted by randomness against possible gains in precision granted by balancing prognostically important covariates. This implies that methods to reduce covariate imbalance (blocking, matched pair, rerandomization, ...) can accidentally lead to larger differences in unobserved variables then a complete randomized design. Question: I don't understand how balancing a covariate could make another unobserved variable non-similar. Can you make a toy example where balancing for one covariate makes unobserved variable non-similar? My thinking: If we balance for covariates with restricted randomization, any positively/negatively correlated unobserved variable would also become better balanced. For uncorrelated variables, they would not be better balanced but also not worse than with a completely randomized design.

Cancel quote

Checking account access…