In the Background

FRT & Interference: A General Method for Detection

Reference

Aronow, P. M. (2012). A general method for detecting interference between units in randomized experiments. Sociological Methods & Research, 41(1), 3–16.

Aronow’s paper proposes a specific Fisher Randomization Test (FRT) based on a known “distance” between units.

Two things I want to do:

  • Replicate some of Aronow’s Simulation Studies
  • Add a twist to the simulations by
    • Allowing for “unlucky” fixed sets by randomly making some units “immune” to interference
    • See if I can find a way to guard against these unlucky fixed sets

Replication

Aronow’s procedure

  1. Select an α\alpha for rejecting the null hypothesis of no interference.
  2. Choose a random “fixed set” among the control set, with size nfixedn_\text{fixed}. All statistics for the test are calculated within the fixed set.
  3. For each unit in the fixed set, calculate the distance to the nearest treated unit, DD.
  4. The baseline statistic, tobst_\text{obs}, is the absolute Spearman rank correlation ∣ρ∣|\rho| between outcome YY and DD (both restricted to the fixed set).
  5. Calculate the same statistic across nitern_\text{iter} randomizations
    • Randomize the treatment assignment ZZ for all units except the fixed set (shuffling ntreatn_\text{treat} 1’s and nctrl−nfixedn_{ctrl} - n_\text{fixed} 0’s)
    • For each unit of the fixed set, re-calculate the distance to the nearest treated unit, DkD_k
    • Calculate tk=∣ρ∣t_k = |\rho|, the absolute Spearman rank correlation between the (fixed) YY and (new) DkD_k
  6. Calculate p=∣tk:tk≥tobs∣/niterp = |{t_k: t_k \geq t_\text{obs}}| / n_\text{iter} and reject if p<αp < \alpha

Aronow explains the use of the fixed set:

without further steps, randomization inference is capable of testing only the joint significance of both direct and indirect effects, rather than either individually.

A solution lies in developing a conditional permutation test for which the treatment status for a “fixed” subset of units remains invariant, so that the randomization inference is testing only the significance of indirect effects on this fixed subset.

For step 6, you could add 1 to both numerator and denominator of the pp-value (I did), but it shouldn’t make much difference for the simulations here.

Replication results

For the replication, I used fewer trials but the same number of randomizations per trial.

The plots above show results similar to Aronow’s Figure 1. However, for the trials with interference I got 90% rejection vs Aronow’s 83%. This is a big enough difference that it’s likely not due to chance. I have not been able to find the source of the discrepancy, but I’m reassured by the alignment of the main conclusions.

The plots below are also very similar to the top of Aronow’s Figure 2.

Replication of Aronow's Figure 2
Left: Plot of how the strength of components within Y fall off with distance in Aronow's Gaussian model.
Right: Rejection rates for α = 0.05, 0.10, 0.20.

For my purposes, this is sufficient replication.

Lucky / unlucky fixed sets with the original DGP

Assuming distance-based interference, some fixed sets will produce stronger Spearman ρ\rho values (∣ρ∣|\rho| close to 1) and some will produce weaker values (∣ρ∣|\rho| close to zero). In other words, some random choices will produce fixed sets with stronger signals and some will produce weaker signals.

In Aronow’s Gaussian model, the outcome YY with interference is given by

Yi=Zi+Θi+Ψi+εiΘi′=∑j((1−δij)Zjexp⁡(−(1.5dij)2))(indirect effect, i.e., interference)Θi=Θi′/(∑iΘi′/n)(mean normalized to 1)Ψi′=∑j((1−δij)exp⁡(−(1.5dij)2))(spatial effect)Ψi=Ψi′/(∑iΨi′/n)(mean normalized to 1)δij={1ifi=j0ifi≠j(i.e., 1−δij evaluates to zero when i=j).\begin{aligned} Y_i &= Z_i + \Theta_i + \Psi_i + \varepsilon_i & \\ \Theta_i' &= \sum_j \left( (1 - \delta_{ij}) Z_j \exp( -(1.5 d_{ij})^2)\right) & \text{(indirect effect, i.e., interference)} \\ \Theta_i &= \Theta_i' / \left(\textstyle\sum_i \Theta_i' / n\right) & \text{(mean normalized to 1)} \\ \Psi_i' &= \sum_j \left( (1 - \delta_{ij}) \exp( -(1.5 d_{ij})^2) \right) & \text{(spatial effect)} \\ \Psi_i &= \Psi_i' / \left(\textstyle\sum_i \Psi_i' / n\right) & \text{(mean normalized to 1)} \\ \delta_{ij} &= \begin{cases} 1 & \text{if} & i = j \\ 0 & \text{if} & i \not = j \end{cases} & \text{(i.e., } 1 - \delta_{ij} \text{ evaluates to zero when } i = j \text{).} \end{aligned}

So, there is a random ε\varepsilon component to the outcome YY, but the interference component is deterministic based on ZZ and distance.

Recall that ρ\rho is the rank correlation (within the fixed set) between YY and DD, where DD is the distance to the closest treated unit. I took 200 samples from the DGP (Gaussian model, 400 units, 200 control), and for each of those 200 samples I took 200 random fixed sets of 100 units each from the control units. I calculated ∣ρ∣|\rho| for each of those 200×200=40,000200 \times 200 = 40{,}000 fixed sets, for both the null (no interference) and with-interference simulations, and plotted the histograms.

Histograms of ρ for the original DGP
Histogram of Spearman |ρ| for the null (no interference) data-generation process (DGP) and with-interference DGP. In each plot the other histogram appears in light gray to facilitate comparison.

There’s some overlap, but there were no iterations with ∣ρno interf.∣≥∣ρinterf.∣|\rho_\text{no interf.}| \geq |\rho_\text{interf.}|. Basically, this reflects what the QQ-plots above show: The null and with-interference cases are easily distinguishable. Put another way, the with-interference fixed sets will usually be “lucky”.

Adding “immunity” to the interference DGP

To degrade the inference signal, I removed some of the interference. I made some units “immune” to others’ influence while maintaining their influence on others, by setting some Θi\Theta_i to zero.

Using the same settings as earlier, I added a large number of immune units, chosen randomly:

nunits=400ntreat=200nfixed=100nimmune∈{100,200,300}(random across test and control).\begin{aligned} n_\text{units} &= 400 \\ n_\text{treat} &= 200 \\ n_\text{fixed} &= 100 \\ n_\text{immune} &\in \{100, 200, 300\} & \text{(random across test and control)}. \end{aligned}

So 25%, 50%, or 75% of all units were randomly selected to be immune.

As expected, the rejection rate decreases:

            Rejection Rates
            with α = 0.05
  Original: 0.899
100 Immune: 0.626
200 Immune: 0.346
300 Immune: 0.131

The QQ-plot also gets closer to uniform: QQ-plot for with-interference DGP with immunity

And the histogram of ∣ρ∣|\rho| values approaches the histogram of the null DGP. The plot below repeats the histograms for ∣ρ∣|\rho| above, but with 50% immunity (from receiving interference) in the with-interference DGP.

Histograms of ρ for the DGP with immunity
Histogram of Spearman |ρ| for the null (no interference) data-generation process (DGP) and with-interference DGP. The with-interference DGP now includes 50% immunity to interference.

So, as intended, the added immunity weakens the interference signal overall, and will create more unlucky fixed sets. The interference is still present, though, and we want to find a robust way to detect it.

Guarding against unlucky fixed sets

What if you choose a subset with low signal? If you have some way of guessing which subsets have the clearest signal, you should use that, but if you don’t, how can you guard against an unlucky choice of subset?

Using multiple fixed sets

My initial idea was to choose multiple fixed sets, and then somehow combine the information from all of them. The more subsets you have, the more likely at least one of them has a strong signal.

The modified procedure, for getting raw pp-values from multiple subsets:

  1. Select an α\alpha for rejecting the null hypothesis of no interference.
  2. Choose K random fixed sets among the control set, each with size nfixedn_\text{fixed}. All statistics for the test are calculated within the fixed sets.
  3. Do the following separately for each fixed set. Compute one conditional pp‑value per fixed set, treating each set as the fixed controls for its own conditional test.
    1. For each unit in the fixed set, calculate the distance to the nearest treated unit, DD.
    2. The baseline statistic, tobst_\text{obs}, is the absolute Spearman rank correlation ∣ρ∣|\rho| between outcome YY and DD (both restricted to the fixed set).
    3. Calculate the same statistic across nitern_\text{iter} randomizations
      • Randomize the treatment assignment ZZ for all units except the fixed set (shuffling ntreatn_\text{treat} 1’s and nctrl−nfixedn_{ctrl} - n_\text{fixed} 0’s)
      • For each unit of the fixed set, re-calculate the distance to the nearest treated unit, DkD_k
      • Calculate tk=∣ρ∣t_k = |\rho|, the absolute Spearman rank correlation between the (fixed) YY and (new) DkD_k
    4. Calculate p=∣tk:tk≥tobs∣/niterp = |{t_k: t_k \geq t_\text{obs}}| / n_\text{iter}

Combining pp-values

Rereading Aronow’s paper I saw that combining pp-values from multiple fixed sets was already partially covered:

the pp values for multiple subsets may be combined to obtain a single pp value for the entire sample (e.g., Kost and McDermott 2002) if a covariance matrix for the pp values is known

Unfortunately, I don’t know the covariance matrix for the pp values here, and I suspect there are many real-life situations where it would be hard to determine. What should we do?

MinP?

I initially tried the MinP procedure, with Aronow’s FRT in an inner loop and a permutation test in the outer loop. Essentially, MinP is a permutation test with Tippett’s method as the statistic (combine pp-values by just using the min). That turned out not to work; the rejection rate was worse than the single-fixed-set rejection rate and got worse as the number of fixed sets was increased. This might make sense because MinP is designed for sparse situations with one extremely strong pp-value among many nulls. In the current situation there’s likely some signal in all of the subsets.

MinP procedure but Fisher’s method?

Using the same FRT-within-permutation test, I replaced Tippett’s method with Fisher’s method: −2∑ln⁡p-2 \sum \ln p. In Fisher’s method, the pp-values should be from independent tests, and the combined pp-value has a χ2\chi^2 distribution. Neither of those are necessary here, because the combined pp-value is just used as a statistic within the permutation test. (A bit more technical: The null distribution is generated through permutation. This is true for Tippett’s method as well. The pp-values should be independent, but in MinP the method is just used as a statistic.) This worked pretty well, and works increasingly better with more fixed sets.

Stouffer’s method?

I then tried the same procedure with Stouffer’s method. Stouffer’s method produces a zz-value instead of a pp-value, but again, we’re just using it as a statistic in the permutation. It’s less sensitive to extremely small pp-values than either Tippett’s or Fisher’s methods. Stouffer’s method seems to work slightly better than Fisher’s method.

Results

For simulations, I used

nunits=400nctrl=200nimmune=200(random across test and control)nfixed∈{100,75,50}(number of units per fixed set)nsets∈{3,5,10}(number of fixed sets)ntrials=500niter., FRT=200(inner loop)niter., adj.=200(outer loop).\begin{aligned} n_\text{units} &= 400 \\ n_\text{ctrl} &= 200 \\ n_\text{immune} &= 200 & \text{(random across test and control)} \\ n_\text{fixed} &\in \{100, 75, 50\} & \text{(number of units per fixed set)} \\ n_\text{sets} &\in \{3, 5, 10\} & \text{(number of fixed sets)} \\ n_\text{trials} &= 500 \\ n_\text{iter., FRT} &= 200 & \text{(inner loop)} \\ n_\text{iter., adj.} &= 200 & \text{(outer loop)} &. \\ \end{aligned}

The values for ntrialsn_\text{trials}, niter., FRTn_\text{iter., FRT}, and niter., adj.n_\text{iter., adj.} are all low. Even with these values, the simulations are very time consuming, and I’m ok with rough results here.

Rejection rate for the DGPs with and without interference, for three adjustment methods.
All fixed sets had 100 units.
# Fixed Sets
Interference Method 3 5 10
No MinP 0.032 0.024 0.018
Fisher 0.046 0.042 0.052
Stouffer 0.046 0.042 0.050
Yes MinP 0.342 0.348 0.254
Fisher 0.480 0.546 0.592
Stouffer 0.492 0.552 0.592

The null DGP has roughly the correct rejection rate, but a little low for MinP with 5 and 10 subsets. For the DGP with interference (and 200 immune), recall the 1-fixed-set rejection rate was 0.346. MinP doesn’t improve the rejection rate with 3 or 5 fixed sets, and in fact does worse with 10 fixed sets. Fisher and Stouffer both improve the rejection rate, with Stouffer doing slightly better, and they get better with more fixed sets.

Rejection rate for the DGP with interference,
with varying size of fixed sets
# Fixed Sets
Units Per Fixed Set Method 3 5 10
100 MinP 0.342 0.348 0.254
Fisher 0.480 0.546 0.592
Stouffer 0.492 0.552 0.592
75 MinP 0.352 0.350 0.266
Fisher 0.484 0.556 0.608
Stouffer 0.490 0.558 0.630
50 MinP 0.286 0.274 0.192
Fisher 0.428 0.482 0.572
Stouffer 0.458 0.504 0.592

To see if smaller fixed sets might work better (by improving the chance of a stronger signal), I reduced the fixed set size to 75 and 50.

  • Having 75 units per fixed set seems to work slightly better, and MinP with 3 or 5 fixed sets gives slightly better than 1 fixed set without adjustment, but the differences are small and not statistically significant for typical values of α\alpha.
  • Having 50 units per fixed set seems to be worse for all three types of adjustment.

Something simpler?

The nested permutation test for the methods above is complicated and time consuming. It would be nice if there were a simpler, quicker method, but I haven’t found any. I tried multiple other metrics in place of Spearman’s ρ\rho, trying to find a metric that might be more robust to sparse signal in a single fixed set, but haven’t found anything that has worked yet.