Replication in R
R code that regenerates the numbers this paper reports, starting from the dataset published on this site. It reproduces 33 of 36. The 3 that do not match are listed below.
Run it
Rscript run.R
run.R is the whole package. Save it and run it: it reads the dataset from
this site if the CSV is not sitting beside it, and loads the shared conventions file the same way,
so it works on its own in an empty directory. It needs R with fixest, and depending on
the paper lme4, metafor, plm, BMS or
LowRankQP. It writes results.json, one value per number, each named for
where it appears in the paper.
How it compares with Stata
Most of these papers were estimated in Stata, and the two programs differ in places that change
printed digits. Those conventions are stated once, in
stata_compat.R, and shared by every replication on this site:
ivreg2's large-sample variance, SSC winsor's order statistics,
xtreg's treatment of singleton groups, and the restricted-ML default of
xtmixed, which is not the default of the mixed command that replaced it.
Numbers that do not match the printed paper
Cells are named as they are in results.json. The reason for each difference follows the table.
| cell | paper | this code |
|---|---|---|
| PanelA_Pub_N | 370 | 378 |
| PanelB_Pub_N | 241 | 249 |
| PanelC_Pub_N | 305 | 321 |
Table 1's published-only columns print N = 370, 241 and 305; this code computes 378, 249 and 321, which is what the authors' own log of the published run records. The coefficients and standard errors of these same columns reproduce. The printed values are left as published.
Files
- run.R, the replication
- stata_compat.R, the Stata conventions
- targets.json, the numbers as printed in the paper, taken from the paper rather than from the code
- results.json, the numbers this code produced
- REPLICATION_STATUS.md, the full comparison, generated from this package's own output