Transcript. Lecture 10, Meta-Research, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.
- Lecture 10. Meta-Research 00:00:00
- A typical meta-analysis dataset 00:02:53
- Coding what the authors prefer 00:04:08
- Meta-research 00:17:44
- Example 1: Attenuation bias, aka regression dilution 00:18:47
- Elasticity of skill substitution 00:24:38
- Many studies far from consensus 00:32:05
- Three estimation frameworks 00:33:54
- OLS: no effect (but attenuation and endogeneity biases) 00:35:35
- IV: strong effect (no biases beyond publication) 00:36:39
- Natural experiments: no effect (but attenuation bias) 00:37:40
- Publication bias > attenuation bias 00:38:51
- Two opposing biases 00:40:19
- Five correction techniques and random effects 00:42:48
- Assumptions of Andrews and Kasy not satisfied 00:50:00
- Model uncertainty; biases confirmed 00:54:17
- Best practice estimate around 4 00:56:03
- Example 2: Too much heterogeneity? 01:08:34
- Fat tails instead of normality 01:10:01
- Test of Infinite Variance 01:12:48
- Estimating the tail index and the cut-off 01:14:31
- Nudges 01:16:31
- Summary 01:18:30
Lecture 10. Meta-Research 00:00:00
A typical meta-analysis dataset 00:02:53
00:02:53 Tomas Havranek: Yesterday we talked about heterogeneity and why we do it. Let's have a look at how the final dataset can look when you collect all the data. As an example, I will open my class size meta-analysis site. It is about class size and student achievement and was published in the Journal of Labor Economics, and we have the data online. It is a big table where we always have the ID of the primary study, the study label, the publication year, and many other things that we collected. Of course, the key ones are the effect size and the standard error. Then you compute the statistics based on these, which is easy.
Coding what the authors prefer 00:04:08
One thing which I did not mention yesterday: quite often it might be useful to code a variable that captures what the actual authors think about the estimate. Sometimes you collect 10, 15, 20 estimates from one paper, but the authors perhaps highlight just one of these estimates, in the abstract, in the introduction or somewhere else. The rest might be provided just for robustness checks, or even to show what happens when you do not do it in the right way, let's say. In this case, we always try to take into account what the authors think about each estimate which we collect.
What do the authors say about it? We have three categories. "Preferred" means that the estimate is explicitly used as the main result, or they also say, "This is what we prefer, this is what we base our conclusions and our inference on." We also collect the page on which this is stated. "Discounted" is the opposite. If the authors say, "We provide this estimate, but it is probably not good for this other reason," we code it as discounted, and then we provide a justification. We include a note on where exactly in the paper the authors say that they prefer this estimate, and why.
This could be useful because, even though you collect these characteristics of data, methodology and so on, you will never be able to capture all of these differences. So it might be useful to take note of what the authors say, because this allows you in the end, when you do the meta-analysis, either as your main analysis or as a robustness check, to focus on the subset of estimates which are explicitly preferred by the authors. Or you can use a broader sample and just get rid of the estimates which are explicitly discounted, when the authors say, "Okay, I have this table, but it is not correct for this and that reason."
I tend to believe it is a good idea to do this when you code the data. Of course, you need to collect the information on what the justification is. If you classify an estimate as neutral, you do not need any justification. You need it only for preferred estimates and discounted ones. Then you can either put it in as additional variables in your analysis, BMA or whatever you use, or you can do some subsamples. And I think this is useful to show referees, editors and other people who will read your analysis that you really do your best in terms of how you capture quality.
00:08:12 Participant: [A participant asks, faintly, whether Havranek ever distinguishes between such estimates in his own analyses.]
00:08:42 Tomas Havranek: I think we have definitely done it a couple of times, at least in two or three meta-analyses. This could be especially relevant for publication bias or p-hacking. We tried to look into it, and we did not find anything. For a control, if it is just one additional control which you do not focus on, then you do not expect much. There could be some, but probably not so much publication bias or p-hacking.
In principle, I think it is a good idea. Well, you cannot collect everything all the time, and this is one of the variables where it is not always clear where you should draw the line between the central focus and just a control. So you need to make a judgment. That is why we also provide, in the dataset, a column with justifications. If anyone wants to check what we did and why, it is there. But in principle, this is more like an indirect control for quality. For example, in this meta-analysis we have one paper which is all about showing how this kind of literature can be wrong.
Essentially all of the estimates in the paper are explicitly labeled as wrong. And the editor pointed out that the paper was published in the AER and that its whole point is that these IV estimates are wrong. The paper was showing this in different ways, using class size as an example. You estimate it this way and it makes no sense; then you do it in a different way, and again it does not work. It was like five tables.
We had it in the data because the default approach is: it is an estimate, so let's collect it, and let's also collect the characteristics of methods, data and so on. But I think the editor had a good point that it is really useful information if the authors themselves tell you, "It is wrong. We believe it makes no sense to do it this way," for reasons which are hard to codify, hard to collect as an additional variable.
The best we can do is to have this classification: either an estimate is preferred, or neutral, or explicitly discounted. And I think that in future meta-analyses I will try to do it, because I believe it is the right thing to do. It reflects, maybe, some dimensions of quality which are otherwise hard to code, hard to observe. Also from the practical point of view, it makes it much easier for me to argue with referees. When they have these kinds of questions, I can always point to this classification. There are many things people in top journals do not like about the way meta-analysis is done.
One thing is that sometimes they do not like it when you include lower-ranked journals. This classification helps you. Plus, of course, you have the general classification. And sometimes they do not like the fact that your meta-dataset is really unbalanced, that you can have 50 estimates from one study and two estimates from another study, and then, of course, you try different weights.
This is a systematic theme in referee reports. They do not feel too comfortable about this. They want me to show what happens when we use one estimate for a study. And then, of course, if you want to use one estimate for a study, you can look at the median or the mean, but it might actually be much better to focus on something which is preferred by the authors.
00:13:50 Participant: I would be a little concerned that when they choose a preferred estimate, they are just implementing their own publication bias. So I am going to tout the result that is the big highlight of what I want to show. And so, in a sense, that is exacerbating the publication bias.
00:14:12 Tomas Havranek: I mean, that is a good comment. It shows how it looks when you include the interaction term or when you look at the subset: you see whether there is more bias, and if there is not more bias, then that is also interesting.
But I think it is good to think about it in advance, when you read the papers, and to try to collect it. In the worst case, it is going to be just two or three insignificant variables in BMA. So you can do it much more simply. You can just collect the information on whether it is explicitly preferred or not. And then you can use it for publication bias analysis, as an extra moderator for subsample analysis. So it allows you to do more, especially, I think, in the review process. I think we should maybe do it more often.
And this is how a typical meta-analysis dataset looks. You have plenty of dummy variables: do we look at math or reading or writing, and so on. It is either zero or one when you code for it, and that is it. So again, as I said yesterday, I think the datasets may be larger than would be necessary. This is a strange paper in a way, in that we find essentially no significant results anywhere. So you could say, why would you need a meta-analysis? Because we essentially just confirm what many people who know the literature say: there is really not much, if any, evidence of an effect of class size on learning.
But still, if you look at public discussions on the issue, if you look at the policies, it is very different. Then the main perception is that we have solid evidence for small classes being much, much better. In the paper, we just showed that if you collect all of this, then whether you look at all estimates or just at the preferred ones, it is zero; there is no publication bias, or not much at least. After correction, therefore, you still have zero and not much heterogeneity. So really, it is a null result paper. But still, it got published.
Meta-research 00:17:44
So let's move to a new topic, which will build on what we discussed yesterday. I will mostly talk about two big projects which are not exactly meta-analysis but meta-research, which is something broader. It builds on a meta-analysis or a couple of meta-analyses right now, and that is something which I believe is natural: when you want to specialize in meta-analysis, you will never stay just with it. That is what I have been doing for a large part of my life. But at some point we will move beyond and do something which builds on this meta-dataset.
Example 1: Attenuation bias, aka regression dilution 00:18:47
The first idea, or meta-research project, that I would like to talk about is a comparison between publication bias and attenuation bias. We know that publication bias is an issue because it makes published results look too big. Because of publication bias, and also p-hacking, which is more difficult, but which will have similar results, similar effects, what we see published in journals and also in working papers and reports is commonly bigger than it should be, because statistically significant results are more likely to be selected for publication, or just for reporting in the working paper.
We know publication bias is a bias in the upward direction. We have at least three papers, including the paper by Ioannidis, Stanley and Doucouliagos in the Economic Journal, and two more papers, which show that on average in economics and finance, the exaggeration due to publication bias or p-hacking is about two.
This means that on average, if you collect estimates for a meta-analysis, the true effect is about half. And of course it depends on the topic a lot, but on average you have 100% exaggeration in the reported evidence. That is a huge, huge difference. So when you want to base your policy on published literature, you need to make the adjustment for publication bias, otherwise you will be really badly off. But you can also ask, wait a minute, publication bias works in the upward direction. What about regression dilution?
Let's go back a little bit. Let's say you want to estimate a regression slope, a regression relationship, and the true data are the black ones [in the figure on the slide "Example 1: Attenuation bias, aka regression dilution"]. This is the regression line, the black one, and the points are the ones in the original. The red dots and the red line show what happens when you add noise to the x variable, the independent variable. So when you add noise to your regression variable, your estimate is diluted.
The slope you estimate is smaller than it was before. One way to understand it is to imagine what would happen if you regressed Y on something which is extremely noisy, so just noise. What will you get if you regress Y on noise? You will get zero, because it is random noise. So if your X variable is extremely noisy, you will get a flat line.
No relationship, because the underlying relationship is completely destroyed by the huge amount of random noise. That is called attenuation bias, or regression dilution. We know it exists in economics, or in any regression, because when you measure your X, you always have some error. There is some noise even in GDP. GDP is computed, but it is not a super precise process, and there is a lot of noise. GDP is a noisy variable. Inflation is a noisy variable. If you ask people how much they earn, there is going to be noise, because some people will lie or they will not pay attention. So in every single economics context you will have some measurement error.
The question is how important it is. I will step back a little. We know from previous research that publication bias creates an enormous upward exaggeration. We know that on average there is also attenuation bias, but we do not know how much of a bias it is, how important it is. So our idea is to compare attenuation bias with publication bias. I will show you how we do it in one paper, in one specific meta-analysis. But you can also do it overall, and that is our current project: we try to do it for economics as a discipline, many hundreds of times, with Chris Doucouliagos.
Elasticity of skill substitution 00:24:38
Let's move to the first example, where we have one meta-analysis. It is a meta-analysis of the elasticity of substitution between skilled and unskilled labor. Skilled labor is you, who will have college degrees, and unskilled labor is people who do not have a college education. The formal classification in the literature almost always uses skilled labor as a synonym for college-educated and non-skilled for everybody else.
[…] This is actually an important problem if you are concerned about inequality. This is the key parameter used in wage inequality models: how easy it is to substitute one skilled worker, like any of you, with a certain number of people who have a high school education. The bigger the elasticity, the easier it is to substitute skilled and non-skilled.
The problem here is that people in the primary studies never, or almost never, estimate the actual elasticity. They estimate the inverse of it: 1 over the elasticity, or even minus 1 over the elasticity. That is because they regress the wage premium, which is the difference between the earnings of college-educated people and non-college-educated people, on the relative supply of skilled and unskilled labor.
You have to do it in this way because the treatment is changes in the supply of educated people. For example, if you are lucky, you have a natural experiment: it could be that you build new colleges, like in Palestine in the 80s, where the first colleges were built. There is a nice paper by Joshua Angrist that uses this exogenous variation. The causality goes from the supply to the skill premium: you change the supply, and then something happens to the skill premium. So you need to regress the skill premium on relative labor supply, but that means that you estimate not the elasticity, but one over the elasticity.
I think this is very important, because in the original paper, when we first submitted to a journal, we just converted these inverse elasticities to elasticities and then did the analysis. But one of the referees wrote a nice report: it is a nice idea, it is all wrong, but I think you can fix it if you completely change the way you compute it.
So we did, because the editor followed the referees' advice. The problem is that if you first convert, if you invert your estimates, and then do the analysis, it is not easy to see directly, but when you write it down, you have to compute the standard error using the delta method. You automatically introduce a mechanical relation between estimates and standard errors. It is a bit like when you standardize coefficients, like when you compute the PCC. The logic is similar. When you translate something to PCC, partial correlation, then by the definition of the PCC you introduce a correlation between estimates and standard errors, or precision. When you invert, it is the same problem.
00:30:22 Participant: [A participant restates the point: the primary studies report the inverse elasticity, the authors inverted it and used the delta approximation to get the standard error, which brings the elasticity back into the standard error, and the referee caught this.]
00:30:59 Tomas Havranek: It was one sentence in the report: this is wrong, for this and this reason. Then there was another set of problems. So when you have a literature like this, what you should do, as I learned five years ago or so, is work with the original coefficients. They are comparable. The only issue is that to understand what they mean, you have to invert them in your head. That is why we do the analysis in the entire paper on the inverted elasticities. That is what people almost always get when they do this. There are a couple of studies that did it the opposite way, so we put them in a separate appendix and focused on the ones that use the standard identification.
Many studies far from consensus 00:32:05
When we collect all of these estimates and compute a simple mean, it is about minus 0.67, which translates to an elasticity of 1.5. So it is mildly elastic.
00:32:21 Participant: [A participant asks whether a confidence interval is computed around that mean, and whether it is done the way described earlier.]
00:32:28 Tomas Havranek: It is just a standard approximation. You can compute confidence intervals here and then also invert the confidence intervals. I think that is the way we do it in the paper. But now we need to go back to the elasticity, because that is what people are interested in. It is the same issue as in the original papers: they also estimate something and then need to comment on the elasticity.
So for the inversion, we just follow what they do in the primary equation. It is the same kind of problem. What we liked a lot is that this 1.5, the overall average of all the estimates we have, is exactly what is considered the consensus elasticity in the literature. [Note, 2026: in the published paper the 682 estimates imply a mean elasticity of 1.8, and 1.5 is the consensus value; see meta-analysis.cz/skill/.] Of course, they do not use 77 studies; it is based on three studies, published in the QJE and the Review of Economics and Statistics. So it is great: we collect a lot, and it is still the same. It is the same starting point. We started from the consensus, and then we try to see whether the consensus is built on strong ground.
00:33:52 Participant: Did you find publication bias?
Three estimation frameworks 00:33:54
00:33:54 Tomas Havranek: Yes, I will talk about it as well. We have different approaches to identification. First, you can just do some sort of time series regression and hope for the best.
Or you can do some sort of IV estimation if you have good data. In the best case, but it is just a couple of studies, you can do a natural experiment, like the paper on Palestine by Angrist, or the paper by David Card on the Mariel boatlift in Miami, where suddenly you have tens of thousands of unskilled workers from Cuba arriving. So it was an exogenous shock to the relative supply of unskilled labor, which allows you to have essentially a natural experiment in which you can identify the causal effect on the wage premium, comparing Miami with other cities in the US.
So we have three basic frameworks. It seems there is a lot of publication bias, or p-hacking, and it seems to be pretty strong.
OLS: no effect (but attenuation and endogeneity biases) 00:35:35
Then we do the analysis separately for OLS, IV and natural experiment studies, and I will tell you why. OLS is a naive technique, a time series regression where you hope for the best. You get some results, but this is just a correlation. You cannot claim a causal effect, so your elasticity estimates are based on shaky ground. You can have attenuation, of course, the problem I was talking about, and you can also have endogeneity. So we collect the OLS literature and correct for publication bias, and we get estimates which are really close to zero, which would imply infinite elasticity of substitution, because these are inverse elasticities. Let's keep it in mind.
IV: strong effect (no biases beyond publication) 00:36:39
Then we move to IV. We assume here that these IV estimates are done well, so we do not judge how valid the instruments are, since we do not have data on that. When we correct for publication bias in the IV literature, we get, not always but almost always, a significant negative effect, which implies elasticities of around three to four, depending on the specification. If the IV is done well, it should correct for both attenuation bias and any other endogeneity bias. That is important.
Natural experiments: no effect (but attenuation bias) 00:37:40
Finally, we have a subset of natural experiments. These are again OLS regressions, but based on data that have arguably exogenous variation in labor supply. So here we have no endogeneity problem, but we have the attenuation bias problem. Even in cases like when the Berlin Wall fell, when you suddenly had an influx of workers from East Germany into West Berlin and places like that, you can use it as a natural experiment, but you still do not measure it with perfect precision, so you will have some measurement error, some attenuation bias. When we correct for publication bias, we find no effect. So again, it is infinite elasticity of substitution. And now, let me put it all together.
Publication bias > attenuation bias 00:38:51
What did we find? We found that after correction for publication bias, the IV estimates were bigger than the OLS estimates. The OLS estimates were essentially zero. The IV estimates were negative, but bigger in absolute value. Again, after correction for publication bias, the OLS estimates were essentially the same as the natural experiment estimates, close to zero. This second bullet point tells us that probably these other potential endogeneity problems are not so large on average, because the OLS and natural experiment results are similar.
The first bullet point tells us that if IV estimates are bigger than OLS estimates in absolute value, then either attenuation bias or endogeneity bias is important. If we can rule out endogeneity, the difference has to be due to attenuation bias. So our conclusion is that the difference is a proxy for attenuation bias.
Two opposing biases 00:40:19
Recall that the mean reported inverted elasticity was minus two thirds. When we take what we want to take as the baseline estimate, which is corrected for publication bias and also for attenuation bias, which means the IV estimate, it is about minus one fourth. So the implied elasticity is four: we invert this and take the negative of it. You can also see that publication bias is still stronger than attenuation bias. So this is what we think is closest to the true value, minus 0.25. And this is the mean of the reported estimates.
From this, you have two opposite forces for the true value. One force is publication bias, which takes you farther away from zero. The opposite force is attenuation bias, which brings you closer to zero. So these two forces go against each other. If the corrected estimate of -0.25 were the same as the mean of the reported estimates of -0.67, we could say that they cancel each other on average, attenuation bias and publication bias. But because the mean of the reported estimates is farther away from zero, the implication is that publication bias is stronger than attenuation bias. So the tendency to exaggerate is bigger than the tendency to attenuate. That is the idea of the paper.
It was published in the Review of Economics and Statistics. Now we take the same idea and try to apply it to the economics discipline as a whole. Again, you need at least two comparison groups, IV and OLS. Hopefully we will also have natural experiments, but they are hard to collect if you do not have a well-identified, well-specified literature. We will essentially do a similar analysis.
Five correction techniques and random effects 00:42:48
00:42:48 Participant: [A participant asks which estimator the meta-analysis estimates on the slide are based on.]
00:43:10 Tomas Havranek: We take the median. This is also based on what the referee wanted us to do. We have five correction techniques: we have our MAIVE here, fixed effects, between effects, endogenous kink and AK. So all five, and we take the median. In spirit, we are approaching something like RoBMA, Bayesian meta-analysis, but in a very simple way: just take the median. That is the best we can do. Or we could just take the MAIVE, which I would prefer, but the paper is unpublished. [Note, 2026: MAIVE has since been published: Irsova, Bom, Havranek and Rachinger (2025), "Spurious Precision in Meta-Analysis of Observational Research," Nature Communications 16, 8454. See meta-analysis.cz/maive/.] So we just went for the median.
00:44:09 Participant: [A participant says they are surprised that random-effects estimates are not shown, since outside economics reviewers would want to see them.]
00:44:29 Tomas Havranek: The random effects estimator essentially does not go down well in economics.
00:44:47 Participant: [The participant adds that economics is unusual in this respect.]
00:44:55 Tomas Havranek: But I think it is an important discussion.
00:44:58 Participant: [The participant asks whether none of the reviewers gave the authors a hard time about this.]
00:45:04 Tomas Havranek: Never. No one has ever asked me for random effects. But I don't think we need random effects here. I'll try to explain. First of all, I am not sure, but AK is like a random effects estimator, isn't it?
00:45:25 Participant: Yes, it is a random effects estimator.
00:45:26 Tomas Havranek: So you have at least one here.
00:45:29 Participant: [A participant asks what the between-effects estimator is.]
00:45:32 Tomas Havranek: Between effects is like between effects in panel data. So you take the average for your study.
00:45:45 Participant: So that's the study effect?
00:45:48 Tomas Havranek: It is called between effects because it captures only variation between studies. So it is the complete opposite of fixed effects in the econometric sense, which means that we take just the within-study variation.
For a classical meta-analysis random effect, it depends on the structure, on how many levels I would have. But in panel data, when I say random effects, it means that I would estimate a parameter and assume that it is randomly distributed across studies. Here I am more flexible: I add dummies for every single study. So it is a non-parametric way to do it. It is more flexible. But you are right, we allow for a lot of heterogeneity across studies. In the fixed effect estimator, I mean fixed effects in the econometric sense, we add a control for each of the 77 studies. So we get rid of all study-level characteristics.
And then of course the comment you would have is that now it is something different: now we focus on the within-study variation. So we remove a lot of the variation, anything in between. That is why right next to it we have the complete opposite. These two specifications are completely independent, because it looks as if they have nothing in common in terms of the identification and variation. Each of them uses a completely different source of variation. And they still give you the same idea overall, which I think is nice here. Actually, I think the two together are better than if I just added my random effect. I can do it, of course. That should be a combination of the two, like a weighted average, if I mean a random effect in the panel data sense.
00:48:31 Participant: [A participant remarks that one can also do random effects in the estimated model, or use something different, like two-way effects.]
00:48:38 Tomas Havranek: Yes, that would be something different. But at least I can say that Andrews and Kasy have the classical method of this random effect as their base. So we cover a lot, and we do not completely ignore this heterogeneity.
00:49:00 Participant: [A participant asks which of the estimates is the median.]
00:49:07 Tomas Havranek: It would be the MAIVE. So this is the idea. And again, we want to do the same thing on a bigger scale, so that we can see how general it is.
Assumptions of Andrews and Kasy not satisfied 00:50:00
Because we focused a lot on this estimator in the paper, I think it is good to point to a test that you can use to test the assumptions of the Andrews and Kasy model. I have never seen anything like it done for selection models in meta-analysis. But Andrews himself, as an editor at the AER, handled a comment on a different paper by Abel Brodeur. Abel and colleagues have this paper on whether publication bias or p-hacking depends on which technique you use. And they say yes: if you use IV, for example, you are more likely to p-hack, which for me makes sense, because with IV you have many more degrees of freedom as a researcher, what to do, which instrument, and so on.
The comment is Kranz and Pütz (2022), and the test for Andrews and Kasy in it was recommended by Isaiah Andrews himself. Essentially, it is a test of the underlying assumption that, beyond publication bias, there should be no correlation between estimates and standard errors.
The test uses the weights estimated by the Andrews and Kasy model, I mean the weights for insignificant estimates: how much less likely they are to be published than significant estimates. These weights are then used to re-weight the original estimates, and then you compute the correlation. If all the assumptions of the model hold, then you should have no relation ex post between estimates and standard errors.
Again, it was suggested by Andrews himself, and it is not exactly a test of this one assumption, but it is the most important assumption of the model. So when it is broken, like here, you can see that when you test it on the re-weighted data, it is very far from zero for most subsamples. So this is a nice way to see, if you base your conclusions on a selection model, whether you are on solid ground.
If it is broken, then again, the most likely source of this violation of the assumption is the problem that, in fact, you have some relation between estimates and precision even beyond publication bias. In that case, I would tell you: let us use MAIVE. But in theory, it might also be because of the other assumptions, for example that you do not have clear categories. The selection model assumes that you have a step function in the probability of publication when you cross a significance threshold: your publication probability changes suddenly, which might not be the case. And there are a bunch of other assumptions. Again, it is easy to do. Kranz and Pütz have simple code for it.
Model uncertainty; biases confirmed 00:54:17
Going back to the previously discussed paper on skills, we do BMA, as we discussed yesterday. Even in BMA, we have evidence for publication bias, that is the variable on top. We have less publication bias in IV estimates, which could be surprising at first sight. I think the explanation is quite logical, because in our sample the OLS estimates are on average much closer to zero before correction for publication bias.
Because of attenuation bias, the effect estimated by OLS is typically closer to zero than for IV. With IV, because it does not have attenuation bias, it is easier to get negative estimates of inverse elasticity. You do not need to try that hard. If you do IV, you are more likely to get something which is negative, which makes sense even without publication bias. But with OLS, you need to select. You need to try harder. That is why we think there is more publication bias in OLS than in IV estimates. I will just move on, because there are a lot of variables.
Best practice estimate around 4 00:56:03
And then finally, as the bottom line, we do the best practice, also called the implied elasticity or implied estimate analysis. We take the results of BMA and compute the elasticity, or in the first step the negative inverse elasticity, conditional on a specific definition of data and method characteristics. That is the first column. We prefer IV, we prefer large data sets, new data and so on. And what we get is pretty close to the previous IV result of four. Then we also do it in the remaining columns. We define best practice based on prominent studies in the literature.
Autor is the most prominent OLS study. So when you do OLS, your inverse elasticity is close to zero, even in this kind of complex exercise. When you do IV, it is close to one quarter, so the implied elasticity is four. Card, again, is the most important example of an IV study, well executed. And then Carneiro et al. is a nice natural experiment, I think, from Norway, where they also built new colleges. You can have a natural experiment where you can exploit a plausibly exogenous change in the labor supply in regions, and you can estimate the effect. And again, it is close to zero, the negative inverse coefficient.
00:58:28 Participant: [A participant asks about attenuation bias in the natural experiments.]
00:58:41 Tomas Havranek: Yes. I change what I call best practice, the values of the variables I put into BMA, based on exactly what they do.
00:58:53 Participant: [A participant asks what is done when the chosen study has no values for some of the variables.]
00:59:04 Tomas Havranek: For the USA I change it separately, independently. I want it to be for the US. But for the US there is a lot of noise, and we are not really able to identify any strong results there.
00:59:32 Participant: [A participant asks whether, to define best practice like this, one could simply use the most cited paper.]
00:59:56 Tomas Havranek: Yes, that could be one way. The idea is that you want to do something like this, but if you use just the first column, it is really subjective. People can tell you it is not convincing enough. What I like to do is to find at least one really well-cited and well-respected study on which there is general agreement that it is well done. Instead of your view of what best practice is, you use what the study uses, and impose it on the entire literature. So in a way you ask: what would the average effect be if all studies did it the same way as Card 2009, which we believe is a good way here? So it is not always best practice. I have these three categories of identification approaches: OLS, IV, natural experiments. From each category I choose the leading study based on citations, but it is also kind of obvious. There is one study which people really like to use a lot, which is based on OLS. Another is the hugely influential paper by Card.
And then there is a new study, it does not really have many citations, but it is a new, well-executed study on Norway. So I think at least at this time the editor liked it.
01:01:45 Participant: [A participant asks whether the original estimates of these studies were reported, and notes that they rest on only a few observations.]
01:01:51 Tomas Havranek: No. We could also compare it.
01:02:03 Participant: [A participant suggests adding a row that reports what the original studies found, to compare with the predicted values, even where the studies do not cover all countries.]
01:02:20 Tomas Havranek: Yes, that might actually be a good thing to do. So here is the takeaway. Attenuation bias in this case seems to be significant, but smaller than publication bias. So when we put it all together, we have evidence for a mean implied elasticity around 4 compared to 1.5, which is now commonly used.
01:03:25 Participant: [A participant remarks that this was a clever thing to do.]
01:03:36 Tomas Havranek: I think the natural extension is to do it for another field, and then...
01:03:52 Participant: [A participant concludes that the exercise would then cover all of the roughly 400 meta-analyses.]
01:04:00 Tomas Havranek: Let us go big. By the way, I think that would be a nice value added for any future meta-analysis. Like your question from yesterday, when I was talking about how to measure beauty: you asked whether, if it were measured by AI, it might be more precise or more reliable. And I think there is something to it, because you always have this attenuation bias, measurement-error problem. Sometimes in the primary studies they find ways to reduce it. When you measure beauty, for example, people disagree on what it means. Maybe a way to decrease the noise around the measurement is to use a more objective measure, such as ratings by ChatGPT. It might have other issues. Let me go back to the figure here. Imagine this is beauty and this is earnings. What do we see? What we should see is a big effect, but we see a diluted effect, because there is noise in how we measure beauty.
But if we were able to measure it more precisely, so that, instead of the noise in the red dots, we go back to the actual black ones, then we would be able to estimate the beauty effect without the attenuation bias, and it could actually be bigger. That is why, in the beauty paper (I did not have time to talk about it at my seminar or yesterday), we control for so many things which are related to attenuation bias.
For example, we also have this in the paper: sometimes these beauty studies report how much raters disagree on the beauty ranking. So you have 10 people giving the rankings. And then some papers would report what was the disagreement rate on beauty. The disagreement rate is a measure of measurement error: how much uncertainty there is around beauty. So you always have the publication bias problem with meta-analysis.
And quite often you can also play around with attenuation bias. You can either compare OLS or IV, which we did in this paper. You can look at different ways to make your X more precisely measured. Some studies find a way to do it in the primary literature, or at least they try something which is related to it.
If you ask different people what beauty means, they will have different opinions. There is uncertainty: one rater will tell you something, another will tell you something else. And when I employ them and ask them to rank 100 pictures, they will give me different answers, and somehow I need to construct the beauty score. So the idea is that maybe if ChatGPT does it instead of human raters, it may be more precise.
Or the rater wants to be nice to the people evaluated, maybe. GPT will not get tired. It might be systematically off, but it will not have this disagreement between raters, or attention issues. So this is the idea. And really quickly about the other project I'm working on. Yesterday we talked about heterogeneity.
Example 2: Too much heterogeneity? 01:08:34
Heterogeneity is a key issue in meta-analysis. We learned about how we can handle it, what it means, why it's important and so on. But still, sometimes it seems you have too much heterogeneity, and people will ask you whether it makes sense to combine, to put all of these studies that you have into one meta-analysis, or whether you should maybe do two separate meta-analyses. In practice that's a common dilemma. Should I go for one, or should I do the analysis separately for men and women? In this project, which is still in its infancy and ongoing, we try to develop a test which would give us a clear answer.
Is there too much heterogeneity in your meta-analysis or not? Should you separate it into two meta-analyses or more, or can you essentially go ahead with what you have? We talked yesterday about the I-squared, and we discussed how imperfect a measure it is.
Fat tails instead of normality 01:10:01
So we do something else. I will probably go directly to the next slide. You can ignore the equations. They are just for people who are really interested in them. The typical assumption in meta-analysis is that the underlying effects have a normal distribution. We almost always use this assumption. If you do a random effects model, you use it. If you do a selection model, you use it. But if you look at the actual distributions in practice, they are not normal. It is obvious that you have these huge outliers. So the idea, which was not mine, is by my co-author, Chishio Furukawa from Yokohama.
Chishio has a PhD from MIT. He was looking at this meta-analysis of mine and others, and he said it is definitely not normal. The normality assumption is way off. It is not even approximately true.
We have fat tails in meta-analysis, like in finance. We really have these heavy outliers on both sides, quite commonly. And Chishio says, well, this seems like typical fat-tail behavior, as in many fields, for example when people want to model the distribution of earthquakes, or wealth distribution, or the size of cities, or the number of words used when you speak. The distribution which is used for these extreme events is called the Pareto distribution. You do not really have to pay attention to the math, but what it means is that when you are at the extreme tail, let's say 5 to 10% of the distribution, it has this functional form: it is a power-law distribution.
Test of Infinite Variance 01:12:48
If you estimate the alpha here, you can test directly whether the implied variance of the distribution is finite or not. Now, when you reject the hypothesis of finite variance, your meta-average would have infinite variance. That's really bad. It really tells you that it is not a good idea to do one meta-analysis with apples, oranges, crocodiles, everything put together. You should do it separately. That's, I think, a very low bar for a meta-analysis: the variance of your meta-estimate should not be infinite.
It turns out that all we need to do is to estimate the alpha parameter from the distribution. We can show that we can actually model meta-analysis data using something like a Pareto distribution. It does not have to be precisely Pareto, because Chishio finds a way to make it more flexible, and you can have some variation in detail, but I will not go into these details.
Estimating the tail index and the cut-off 01:14:31
Again, in the literature on earthquakes, wealth distribution, word frequency, there are established techniques for how to do this estimation. So you estimate the alpha using maximum likelihood, and you choose where the tail of the distribution starts, the cut-off, using some sort of goodness-of-fit test. It looks complicated, but actually it is not so difficult.
01:14:59 Participant: [A participant makes a brief, faint remark.]
01:15:14 Tomas Havranek: Sometimes you have outliers which you can legitimately get rid of. So this would be an additional thing you can try. And it's pretty clear: if the implication is that the variance is infinite, you should do something about it. It might be driven by just outliers, which maybe don't belong. But still, maybe first you should do the test, and if you don't pass the test, then you should do something about it: either split it, or get rid of outliers, or seriously think about what is going on there.
Nudges 01:16:31
So far we have just managed to apply it to one dataset. This is the meta-analysis on the effect of nudges: when you nudge people to do something, when you encourage them to, I don't know, save more for retirement, donate more blood or recycle more. It covers completely different kinds of life scenarios and completely different behaviors, from charity donations to pension saving, recycling, whatever.
It was published in Econometrica, and it's a nice paper. Because it's obvious there is a lot of heterogeneity, it's a good example where we can test it. It looks like this: we estimate it in logarithmic form, because then it's linear. You can see the tail of the distribution is essentially linear. So you have a nice way to estimate the alpha parameter, which is substantially below 2, which means that we can reject the hypothesis of finite variance.
Summary 01:18:30
I talked about two projects today, meta-research things, as two examples of what you can do beyond just applied meta-analysis, how you can extend it. One example was how you could incorporate attenuation bias when you do a meta-analysis. So you can have a nice comparison between publication and attenuation bias. And the second example is still in the works, but hopefully in Ottawa Chishio will present the final product, and we will have a package in R, which will be really simple and will allow you to run this test, and then you will decide what it means for your meta-analysis. But probably if you have infinite variance, you should do something about it.
01:19:18 Participant: [A participant raises the question of converting effect sizes, for example to partial correlation coefficients (PCC).]
01:19:39 Tomas Havranek: And then, by the way, when we have this paper ready, it would be easy to extend the analysis to different transformations. So this could be a clever thing to try to see what happens. Does it still give us similar results? So I don't know what it would mean for our paper, but I think it could be a nice extension. If our paper is published somewhere, then you can...
01:20:22 Participant: [A participant makes a faint remark about what publishing the paper would mean for people who transform effect sizes.]
[The participant says that transformations such as PCC make it easier to get a normal distribution.]
01:20:46 Tomas Havranek: Maybe. We want to publish it in an economics journal, so we are not in any rush; we just want to think it through. I have some examples, tests, there is going to be mathematics by Chishio, and some empirical applications on these hundreds of meta-analyses in economics.
First we will do it for economics. Then we want to maybe recruit more people like you guys and, if you are willing, do it for many disciplines. Let's have 20 people on the data, because then it's easy. When it's finished, and once we have the R package, we will just need these different datasets put together.
Essentially, we should show what share of published meta-analyses have infinite implied variance. That is something which is highly interesting and has big implications for evidence-based policy, if you care about that, and for how people use these meta-analytic results. So it's the medium-term plan, I think. But it's really slow. That's all I had for you today.
01:22:37 Participant: [A participant asks whether the test is applied before or after heterogeneity has been explained by covariates, as with I-squared.]
01:23:04 Tomas Havranek: Yes, right. In principle, you could also do it at the end, using some residuals.
01:23:23 Participant: [The participant adds that the fat tails may be due to covariates that are left out of the analysis.]
01:23:37 Tomas Havranek: Maybe. But I think first we want to keep it simple and just focus on the averages. The connection would be Simonsohn and his post on "Meaningless Means", which many people know. And I see some people sending me the link: "You see how problematic meta-analysis is, you see this post by Simonsohn." So that would be a response. Let's look at it. When the variance is infinite, the mean is probably quite meaningless. It doesn't mean that the entire meta-analysis is meaningless, if you focus on heterogeneity analysis.
But the truth is, even though in economics we do a lot of heterogeneity analysis, we often focus on the mean overall. In my papers, for example. Because that's what people want to know. They want to have one number and then maybe more numbers for different contexts. But they want to have a number. It's not enough to say, okay, it depends on what frequency you use and what kind of identification strategy you employ. They want to have concrete results that they can use in their models. It's a fair criticism, and this is one way to tackle it. Whether we succeed or not, I don't know, but it looks promising based on this one pilot study on the nudge dataset. […]
Corrections
Slips of the tongue corrected in the text:
- [00:46:10] said "67 studies"; the text has "77 studies".
- [00:33:05] said "the AER, the QJE and another journal"; the text has "the QJE and the Review of Economics and Statistics".