Transcript. Lecture 11, Limitations of Meta-Analysis, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.
- Lecture 11. Limitations of Meta-Analysis 00:00:00
- Scholar search replicability 00:05:45
- Universal coverage 00:08:37
- Measurement error 00:12:46
- Rounding 00:16:00
- Outliers 00:20:38
- Too much heterogeneity 00:28:56
- Identification 00:42:15
- Attenuation bias 00:52:12
- Partial correlations 00:56:06
- Inverse coefficients 01:05:20
- Delta method 01:09:33
- Tests disagree 01:13:01
- P-hacking 01:17:45
- Omitted variables 01:19:35
- Conclusion 01:24:46
Lecture 11. Limitations of Meta-Analysis 00:00:00
00:05:20 Tomas Havranek: Today we will mostly talk about some potential problems and pitfalls, as you can see from the title. You're welcome to bring up your own issues. I'm sure some of you will. We can talk about things that can go wrong in meta-analysis or meta-research.
Scholar search replicability 00:05:45
In the first week, we talked about how you search for papers and collect the data. If you use Google Scholar, which is what I recommend and what I mostly do in my own research because of the universal coverage, there is one issue. Google Scholar is constantly changing the underlying algorithm. When you want to replicate exactly what you did last year, it's not going to be possible. This is a caveat that some people find annoying: it makes it impossible for people to replicate exactly what you did.
But I think it's still worth using Google Scholar because of its many other advantages. Be aware that this could be a potential issue. One reason I don't think it's a great problem is that, when you search for papers for a meta-analysis, you eventually also want to do snowballing. This means that you look at the studies that are frequently cited by the other ones.
You don't rely just on this Google Scholar search. You use it as the first step, the starting point, and then you build on it. Even if the algorithm changes slightly, it shouldn't really affect your overall coverage of the papers that you find because of the snowballing, which is completely different from the first starting point.
00:07:46 Participant: I noticed that when you do the snowballing, you might say, "Maybe I should add this search term." I don't want to start again from scratch. [The participant describes finding extra papers by adding a search term.]
00:08:10 Tomas Havranek: I think it's perfectly fine to say that you had a starting point in your search. Then you build on it, which could be snowballing, or you could add some keywords, as you say. Be aware that some people ask about this when you submit a paper or a thesis. Now you know how to respond and why it's not a big problem.
Universal coverage 00:08:37
A related problem that many people tend to be nervous about is that, even if you do snowballing and sometimes use different databases, you can never be sure you have collected all of the papers on the topic. In fact, you are almost guaranteed not to. Maybe you will if it's a new topic, so the universe of papers you need to collect is not that large.
In this case, I believe you might be able to find all of them. But if you write a paper on capital-labor substitution, there have been hundreds and hundreds of papers. It's highly likely you will not have all of them. Is it a problem? Again, some people can hold this against you.
What can you do? First of all, you have a clear approach, and this paper didn't turn up in your approach. You can say, "It's impossible to do everything." But it's a technical excuse. What has more gravity, I think, is saying, "I have my search strategy, and we didn't omit any studies on the basis of the results."
If we omitted a paper, we did what we could. If we omitted some papers, the omission should be random in terms of the results. You will always have this problem in almost every meta-analysis. You don't have the full universe of the published results that you collect because sometimes there are so many of them.
Sometimes it's also quite difficult to construct these PRISMA diagrams. I talked to Bob about it the other day. It's always nice to save the information as you do the search. But sometimes you change your query a little bit.
00:11:28 Participant: But again, couldn't you just make your PRISMA graph for your initial search, your starting point? Then you say, "After that, I still added those five studies based on..."
00:11:42 Tomas Havranek: That would be ideal, but the Google Scholar algorithm can change, and you have a different number. You can also have new studies there.
It's best, as you say, to take the number you get from your initial search. We do this PRISMA diagram mainly to illustrate what the general strategy was.
Measurement error 00:12:46
Some people care a lot about measurement error. There are many ways in which you can have measurement error. My first example is from one of the papers I showed you a couple of weeks ago. If you collect data from graphical results, you have, for instance, studies where the results are not reported in tables but in figures. In this case, we look at the effect of monetary policy on house prices.
You have to measure the coordinates in the figures and try to recover the numerical value when you want to look at eight quarters, two years after the original shock. You measure it, and it's never exactly precise, of course, so you have some measurement error. In this case, it's fine because it should be random. If you measure it in the same way, it should be random.
You have a different type of measurement error when we think about mistakes we make when we code data. That could be the effect size, precision, or these measures of heterogeneity, which are sometimes difficult to collect because you need to think about them hard.
[For example, whether the study was experimental or not.] The best measure you can take against it is to have more than one person collecting data, which is not always possible. What is almost always feasible is to ask someone to check at least a small portion of your data. You get an independent opinion on how you collect your data for your meta-analysis. It could be just one study.
That's one thing you can do. Of course, if you really want to make it bulletproof, you can have two people doing all the coding separately. Then they compare the results, and you also record the disagreement rate. But it's really costly in terms of labor. We did it a couple of times, but we don't do it all the time for all the papers.
Rounding 00:16:00
You can also have another type of measurement error that arises from rounding. If you collect your estimates from regression tables in different papers, rather than from figures, these papers always round, so they do not report all of the decimal points. They typically report two or three decimal points, but sometimes it's just one, and sometimes it's four. In some cases, that could have an effect, for example, when you look at the distribution of t-statistics or z-statistics and want to do what are called caliper tests. You compare the number of estimates just below and just above a threshold, like 1.96.
Then you need to be careful. Even if the rounding is relatively small, it will give you a large number of t-statistics that are exactly one or two, because both the estimate and the standard error are rounded. It doesn't happen always, but it's going to be frequent enough to make a big impact on this caliper test.
You can see these spikes at one and two in t-curves. There is a way to de-round it, which you need to think about if your primary goal is to look at publication bias or p-hacking using these t-curves or z-curves, the way Abel Brodeur does in his papers. It's not exactly a meta-analysis; it's meta-research. But even in any classical meta-analysis, potentially, especially if people ask about this, it could be nice to have a robustness check when you de-round.
When it's rounded, something is cut off from the decimal points. You can add a small amount of random noise, so that you have a consistent way in which all of these papers are rounded, let's say to 5 decimal points. There are ways to do this. On the slide [Rounding], the figure on the right-hand side was produced essentially this way.
There are ways to deal with this rounding issue. For normal meta-analyses, rounding is typically not a big issue. But if you focus on caliper tests, it could have a large effect. There is a comment on one paper by Abel Brodeur. If you read the original paper in AEJ Applied, from 2016, he uses these de-rounding techniques.
It's important. It can have big effects on this type of test. Then he has a similar paper in the AER. There is a comment on the AER paper by, I think, Sebastian Kranz. I think I already mentioned this. They do the de-rounding. […]
Outliers 00:20:38
A big issue that is still, I think, unresolved is outliers. It's not clear what to do. […]
00:21:26 Tomas Havranek: There are outliers both in the size of the effect and in precision. You can have estimates that are otherwise normal but super precise, suspiciously precise. What can you do? There is no universal solution that everyone agrees on. When you see a problematic paper or estimate, you should check whether you coded it well. Check whether it's not a typo on your side or maybe an obvious typo in the original study, like a decimal point issue.
Sometimes, and I have seen this many times, the authors have a regression table with point estimates and something in parentheses. They say, "We have standard errors in parentheses," but sometimes it's actually t-statistics in parentheses, or the other way around. There is a mistake in the label. If you take it at face value, you have an obvious mistake. Sometimes you can see it clearly from the number of stars. It doesn't really match the description. That's something that will sometimes happen. But even if you do your best, sometimes your final plot will look like the one on the slide [Outliers].
It really tells you that you are putting a lot of different things together. You have super precise estimates close to zero, but also really large estimates in economic terms, if you talk about the elasticity of substitution between capital and labor in this case. This could be heterogeneity, but if you have numbers like five or sometimes ten, you know that it doesn't make much sense to put it all together and continue as if it were a normal estimate.
What I typically do in my papers is use some sort of winsorization. This means I select a percentile, typically one percent, and say that if the estimate is bigger than my 99th percentile, I will still include it, but squeeze it to make it equal to the number at the 99th percentile. I do the same from the other side if it is too small. You keep the information in the data that there is a large observation, but you squeeze it. What else can you do with it?
If you check the data and it still looks okay, I prefer to squeeze it and keep it in the data. Other people will tell you that you should delete it, maybe. The threshold is also not clear. I like to use 1% if possible. Other people would tell you that maybe you should look at estimates that are more than two or three standard deviations from the mean. When you do these PET-PEESE regressions, you can use outlier detection tools as in normal regression analysis. Again, there is no consensus in economics on what to do with outliers.
My take is that there is no universal way to handle outliers. Nowadays, if I look at papers published in really good journals, they mostly include all observations and are really careful to justify it clearly if they get rid of some data points. They try to include as much of the data as possible. What would you say?
00:26:07 Participant: [A participant agrees, saying that removing extreme estimates or standard errors often makes little difference in their experience, and recommends a robustness check. Part of the comment about defining outliers is hard to hear.]
00:26:36 Tomas Havranek: That is ideal. If it doesn't make much of a difference, you are fine. In my experience, sometimes it still makes a big difference. Sometimes, if you do all of these tests on the raw dataset, even after checking carefully for typos, it gives you results that switch completely when you get rid of, or do something to, the most extreme 1%. Then the question is what to do. I find it more reasonable to focus on results driven by 99% of the data than on ones driven by the 1%.
But sometimes the 1% contains the most precise estimates. I sometimes have small arguments with other people: maybe you should put it all in. It is not set in stone.
00:27:42 Participant: [A participant suggests another argument for using random effects and asks whether Havranek excludes an outlier if its data are unreasonable and squeezes it if they are reasonable.]
00:28:01 Tomas Havranek: If I go into the paper, sometimes I find something weird about it. Either you find an obvious mistake in a decimal point, for example, or you find that the paper is doing something slightly different from what you collected. This can happen afterwards when you carefully check these outliers, these influential points. But if you don't have a reason to exclude it apart from the fact that it doesn't look good, you should include it and do something about it, such as winsorization, the squeezing. That's what I would do.
Too much heterogeneity 00:28:56
I think last time I also talked about this meta-analysis of nudges by DellaVigna and Linos, published in Econometrica. That might be the most prominently published meta-analysis ever. It's a great meta-analysis, but if you look at the funnel plot, you sometimes have huge estimates.
00:30:35 Participant: To me, the first thing is that it looks like they have two different funnel plots.
00:30:40 Tomas Havranek: Maybe many more than just two. You have so much heterogeneity. They use random effects, so they do not hide it. But we should be really careful in how we interpret such a meta-analysis. It has widely different effects.
00:31:19 Participant: But you don't know in advance.
00:31:23 Tomas Havranek: You know a little bit when you design your topic. In this case, the effect is what happens to how people behave if you nudge them as a government. But it's a very broad definition of an effect: it can include blood donations. It can include sending you letters, and the effect is how much you save for a pension. The idea is called nudges, based on the book [Nudge, by Thaler and Sunstein]. […]
You have these nudge units. In the U.S., at least, they used to have them. In Britain as well, you have institutions that collect data on these things. That's why you have the data.
00:32:42 Participant: I think there's a trade-off here. In some way, did you say that you only want to do a meta-analysis on something for which you have many studies on exactly the same topic? That means that the number of times you can apply a meta-analysis becomes less and less.
00:32:57 Tomas Havranek: Certainly, you have a trade-off. You also need a topic that is general enough for people to hear about.
00:33:15 Participant: You see the effect of AI, of generative AI, on productivity. Even there, you can say some look at productivity in design, others look at productivity in coding, others look at productivity in search. It's also very different areas where you would expect very different effects. […]
How do you decide whether it's too much heterogeneity or not?
00:33:53 Tomas Havranek: We have this new project with Chishio Furukawa, which I presented at some point. We try to address precisely the question you just raised: when do you have too much heterogeneity? We have a test based on the Pareto tail. To simplify it, there is a way to show this if you assume that the extremes in the distribution of your metadata behave like the Pareto distribution, which is heavily used for modeling earthquakes and wealth distribution.
It is universally used for these tail behaviors. If you estimate this parameter and it is less than 2, it implies that the variance is infinite. That's a natural test, because if you have an infinite variance, it doesn't make much sense to compute a mean. You can still compute the mean, but it doesn't have much meaning.
00:35:14 Participant: [A participant questions the usefulness of a test that reveals too much heterogeneity only after all the work has been done, suggesting that the work would have been for nothing.]
00:35:23 Tomas Havranek: No, it doesn't mean that you have to scrap it. It tells you that maybe you should divide it: you should do two papers.
00:35:38 Participant: What do you have to say about doing this on the front end versus the back end? You're doing this at the front, but maybe there are some explanatory variables, like maybe the type of area we're talking about for productivity, that explain those things. It seems to me that what you want is the heterogeneity after you've used your observables to explain the variance.
00:36:03 Tomas Havranek: Right now, what we are working on is, as you said, the raw data. But you could use residuals from your meta-regression. I think the same thing would work on the residuals as well.
00:36:23 Participant: [A participant asks which approach would be right if tests were available for both the raw variables and the meta-regression.]
00:36:36 Tomas Havranek: These are different questions. If you want to focus on the mean, which many methodologies do in the end, then raw data analysis would be the right one. If you focus on being able to explain heterogeneity, then it's fair to use the test on the residuals. My impression is that in economics and finance, we always focus on heterogeneity in meta-analysis.
But quite often, when I look at how people quote my papers, in 95% of cases they quote the big number, which is typically the mean corrected for publication bias. I try not to overstress it; maybe I do sometimes. The fact is that this is the overwhelming usage of my research. If you want to calibrate your model, you find a couple of meta-analyses. It's natural. That's what people do.
00:38:02 Participant: When you're on the front end of the test, as you're doing now, do you treat all these points as equal, or do you weight them by the standard errors?
00:38:18 Tomas Havranek: So far, we haven't. We just look at the point estimates because you are interested in the tail of the distribution. Whether we should weight them higher is a good question.
00:38:40 Participant: I don't know the answer.
00:38:42 Tomas Havranek: No, that's a good question. In the end, you always apply some weights in meta-analysis. We will do it for a hundred meta-analyses in economics. Someone could say that if we use precision weights, then the importance of this outlier, of this issue, is smaller and it is okay. I think that's a good comment that we should anticipate.
In principle, it should be straightforward to design a weighted version of this test, because it's a simple regression. It's easy to do; there is no reason why not. We can do it in weighted form as well. I have to talk about it with Chishio, to see if it makes sense from a statistical point of view.
But I think we should definitely think about it. Hopefully, in Ottawa, Chishio and I will present a finished paper and R package that will be able to do this test for infinite variance for any meta-analysis. It is not the final answer, but it tells you that if, under some realistic assumptions, your distribution has an infinite variance, then maybe you should consider splitting the data, or looking hard at your outliers. Maybe the problem is not heterogeneity, but extreme outliers that you kept in your data.
00:40:30 Participant: It would also be interesting in your research to see if you can find a nice relationship between what you're doing here with the Pareto and I-squared. That's your motivation with I-squared. I wonder if, in general, there's a clear, monotonic relationship between I-squared and what you're going to find in your test statistic here for inference.
00:41:03 Tomas Havranek: That's what we should do as well. We should compute I-squared for each of these analyses and see at which point there is a systematic relation between rejecting the null hypothesis of infinite variance, or vice versa, and I-squared. I think that's also a good idea.
This is something to keep in mind. Another example is from our paper on beauty. Sometimes it's really clear that you have two populations, two effects. In the funnel plot, they have two different peaks. It's best to keep them separate. In this case, it's really obvious. It's almost never so obvious, but here it was clear.
Identification 00:42:15
Another issue that I think we should pay much more attention to is identification. We need to keep meta-analysis as a way to put results together. A nice example is different experiments on the same medicine, on the same drug. If you look at it this way, you have five different experiments on different populations. You can put them together and do whatever meta-analysis statistical test you want. But in economics, we mostly look at research that is observational rather than experimental.
This means we don't have a randomly selected treatment and control group, unfortunately. Even in medicine, most medical research doesn't. If there is a new medicine, you will need to do RCTs. You will have the data on RCTs. But if you want to say something about how eating meat impacts your health, or alcohol, there are essentially no RCTs. It's so hard to randomly select people and then force them to either eat meat or not eat meat, or drink or not drink.
You have to rely on observational correlations. That's all you have. It's tricky because if you don't specify your model well, you will have a correlation, which doesn't have to be the causal effect you are interested in at all. An example I gave you was alcohol. For many years, and even now when I talk to people, many people still believe that it's healthy to drink a little bit compared to not drinking at all.
You can hear, even now, that it's healthy for you to drink one, two or three glasses of wine per day. There was a long-term consensus that it's healthier to drink.
But it's completely wrong. The problem is that these observational studies didn't control for medical status, or didn't control well for the medical status of people who don't drink at all. If you don't drink at all, maybe you do it because you have really strong willpower, or you don't like the taste of alcohol. But sometimes you don't drink because the doctors told you that if you drink, you will die because you already have liver disease. Then, of course, even if you don't drink, you can live a shorter life than people who drink a little bit.
But it's not because of your alcohol consumption. It's because of the disease that you have, and that's the reason why you don't drink. It's very easy to mess it up in observational research. That's why I think it's really important not just to rely on the mean when we do meta-analyses, but to seriously take identification into account. It's difficult. Sometimes it's not clear whether, for example, IV studies are better than other studies. It can be the opposite. When you do a meta-analysis of economics or finance research, you should always be careful about identification. I think this applies to a meta-analysis of any observational research.
00:46:47 Participant: For those kinds of things, do they have a proxy that they typically use for willpower or drive? You're talking about people who don't drink. They might have more drive, and then they might have more drive to exercise or do other things.
00:47:03 Tomas Havranek: That's a good example of how difficult it is to measure it. In an ideal experiment, I would take 100 people, or even 30 people, and randomly assign them to drink or not to drink. The problem is that I cannot do it. They can promise me something, maybe, but then they go back home. I will not keep them in my hospital for a year and watch what they do and how much they drink.
00:47:35 Participant: [A participant makes a comment about the experiment that is hard to hear.]
00:47:37 Tomas Havranek: Maybe if you pay a lot, but then it will most likely not be a random sample. Now imagine you had to be under surveillance for a long period of time, and you were forced to eat and drink what the experimenters wanted you to. That's hard to do. That's why, by the way, research on nutrition is so weak in medicine.
We know very little about it, surprisingly. In economics, it is difficult, but you can always do something along these lines. You can look at whether the study used difference-in-differences, RDD or IV, and if it used IV, whether the instrument was strong or not. You can try to control for this identification. On a related front, sometimes you cannot really do much more than OLS or some versions of OLS.
For example, I keep going back to this paper on the beauty effect. There are no experiments. It's very hard to do any IV or difference-in-differences; it's really rare. What people do almost always is OLS, and they hope they have enough controls to make the estimate reasonably consistent, reasonably unbiased. This is from a different paper on monetary policy and house prices. Again, it's mostly about which variables people have in these vector autoregressions.
For example, here we found that if they control for credit and money supply, there is a systematically different effect. That was one of the main outcomes of our paper. You can always find a way to at least partially control for the identification employed in the primary study. That's what we always have to do, and I think that's what distinguishes economics and finance from medicine, not just in terms of meta-analysis, but in terms of the empirical procedure itself.
In medicine, they hope to have an RCT. If they do an RCT, they know how to do it well. But if they don't do an RCT, then often, I think, identification in medicine is really weak, surprisingly so. I think that's one more reason why you study econometrics.
We can almost never just compare two means, for the control and treatment groups, in economics. Of course, you have experimental economics, but it's a small part of the discipline. Even though it's growing, it's still tiny. In medicine, it's the default gold standard when you have an RCT. That's your benchmark, and you don't have to go much further.
Attenuation bias 00:52:12
Last time I also talked a little bit about this attenuation bias issue, and I want to stress it a little bit more. When we do meta-analysis, we are used to stressing exaggeration or overestimation because of publication bias. We like to stress that the estimates which are published are likely too large because the ones which are insignificant are often not reported. But we need to be aware that there is always also this opposite problem, called attenuation bias or regression dilution. Any time you do a regression, measurement error in your regressor will bias your coefficient toward zero a little bit.
I think it's fair to say that because you will always have some random measurement error. You can also have non-random measurement error, but you always have some mistakes in measurement, even if you look at macroeconomic data like GDP or micro surveys in labor economics, or any kind of regression analysis. You will have publication bias, which pushes your literature upwards, away from zero. But you also have attenuation bias, or regression dilution, which works the other way.
It's much more difficult to handle. In this course, and in meta-analysis in general, we have seen many ways to control and correct for publication bias and potentially also p-hacking. But with attenuation bias, we know it's there, but it's not discussed much. It's almost never quantified. In labor economics, they do it a little bit. Sometimes they can have a validation set of data, so they can measure how large attenuation bias is. But in general, it's difficult.
I talked last time about some ways to do it. You can compare IV to OLS under some assumptions. In the example of the beauty paper, you can use a rating of beauty done by people who disagree with each other, and then a ranking of beauty done by a machine, an AI like ChatGPT, which might be systematically off, maybe, but is completely precise once you define it.
This way, you could reduce the amount of attenuation bias and get results closer to the original causal effect, rather than the one diluted by random noise, which gives you attenuation bias in the process. […]
Partial correlations 00:56:06
Partial correlations are easy to compute, but you should really take them as the last resort because you lose a lot of information. Even if you have a smaller correlation, the underlying effect can be pretty large, even if you compute it correctly, as I'm sure you do. You have some guidelines on how you should interpret correlations, but these are really just rules of thumb. It's better to have some guidelines than no guidelines. If you look at the original paper by Chris Doucouliagos, there is a relation between the size of these correlations and the underlying elasticities, but it's really weak.
It gives you an indication that it's never a good idea to rely solely on PCCs. Again, one recommendation is that if you need to combine your datasets and translate your estimates to PCCs, always also use a subsample of effects for which you can use the actual economic effect.
00:57:35 Participant: [A participant asks what to do if there is no sufficiently large subset of comparable studies, with effects measured in different ways.]
00:58:06 Tomas Havranek: I would take the biggest subset. That's what I would do.
00:58:11 Participant: [A participant asks about putting effects into percentage changes.]
00:58:26 Tomas Havranek: Then you're able to use something like an elasticity or semi-elasticity. That's ideal, if you can do it.
00:58:37 Participant: No. […]
My other point is that, and I can't tell off the top of my head, for reasonable-sized partial correlation coefficients, they are effectively equivalent to standardized beta coefficients. That's something I hadn't appreciated before. Standardized beta coefficients make a little more sense to me because now you're talking about not just the size of the coefficient, but how much a standard deviation change in X affects a standard deviation change. It's not a perfect thing, but for PCCs that are, I'm going to say, 0.5 or lower, the correspondence is quite high.
00:59:36 Tomas Havranek: This is also new to me.
00:59:38 Participant: [Participants discuss a small correlation alongside a larger percentage effect, whether the percentage change is an elasticity, and whether the X variable is consistent across studies. They note that a consistent X variable is a better case for comparison.]
01:00:43 Tomas Havranek: This can help you if one of the variables is essentially 0 or 1, and then you find a way to translate the Y variable into percentage changes. I think that's much better. In your case, I would probably forget about PCCs altogether.
01:01:04 Participant: [A participant notes that percentage changes are not always easy to compute because the base category is not always known.] […]
01:01:23 Tomas Havranek: I wanted to briefly mention Bob's favorite topic, and I hope I will get it right. This is the partial correlation coefficient, and that's the way I used to compute the standard error for many years. Bob and others showed us that, from a statistical point of view, it's incorrect. The correct formula for the standard error is the one on the slide [Partial correlations]. The numerator is smaller than one, so when you square it, it's even smaller.
That's why you get more precision with this correct formula. That's another reason why PCCs are problematic. Tom Stanley, sometimes with me as well, has a couple of papers on what happens when you use different measures of standard error. My summary is that not really much happens in practice. In some cases, you would prefer the original, wrong estimator because, even though it's not correct for the PCC, in meta-analysis it gives you a smaller bias. This is precisely because squaring the numerator here gives you more precision, which makes the correlation between estimates and standard errors in PCCs even worse. But I think, from a practical point of view, it doesn't really change the recommendation I gave you before: PCCs should be the last resort.
In many papers, I use PCCs, and some papers are exclusively based on PCCs. But if I am about to write a new meta-analysis, I will try to avoid them. You heard how sometimes you can have real discrepancies: a small PCC, with a large economic effect behind it.
01:04:22 Participant: [A participant thinks Tom Stanley made a great point and discusses comparing results using the common and theoretically correct standard-error formulas.]
01:04:53 Tomas Havranek: I would call it "common".
01:04:58 Participant: [A participant says that the labels have changed and are now "theoretically correct" and "common".]
01:05:04 Tomas Havranek: It's more of a small digression.
Inverse coefficients 01:05:20
Another point I would like to emphasize, which I already mentioned last time, I think on Monday or Tuesday, is something I learned a couple of years ago from a referee report on our paper in the Review of Economics and Statistics.
We were doing an analysis of the substitution elasticity between skilled and unskilled labor. I think I showed this to you last time. We did it on the elasticity, so we converted it first because most people estimated the inverse. We inverted it, computed the standard error using the delta method, and then we would have the elasticity. I don't have it on the slides, but it's not difficult to show that the inversion itself introduces a relation between estimates and standard errors.
When you do publication-bias tests, you find a correlation in your FAT, but you don't know whether it is because of publication bias or because of your inversion. The referee told us that this was not how we should do it. We should do it on the original coefficients. We completely redid the analysis on the original reported coefficients, not the elasticity. We did the meta-analysis and then computed the implied elasticity, not the other way around.
01:07:03 Participant: When do you work with the inverse elasticity? What was the application? Which way was it? You're making a general comment, which is really great, but I'm wondering what the application is. In your paper, the key variable in the primary studies was the inverse elasticity. Where am I likely to see that?
01:07:27 Tomas Havranek: I think in many cases in economics, when you estimate an elasticity, you actually estimate its inverse, especially if it's an elasticity of substitution. Not always, but the reason is that you have two variables, x and y. You would like to have y on the left-hand side and x on the right-hand side, but sometimes it's really hard to find exogenous variation for x, while you can find it for y. You switch them and estimate one over the elasticity. That's what they do here.
01:08:10 Participant: Is it a reverse regression?
01:08:12 Tomas Havranek: It's a reverse regression because that's how causality works in the field. You can have quasi-experiments, like the one in Miami.
Suddenly there was an increase in unskilled labor in Miami, and you can use it as an experiment compared to other cities in the US. But you will not have such experiments with the skilled wage premium, which would suddenly change for exogenous reasons. That's why it would be easier to estimate it the other way around, but they don't have the exogenous variation in the variable they want to have on the right-hand side.
They have to switch it. I think this is quite common when people estimate substitution elasticities, and also in some other fields. I think Tom Stanley has one paper in environmental economics, but maybe it's also another elasticity of substitution.
01:09:30 Participant: [A participant begins a comment about being careful.]
Delta method 01:09:33
01:09:33 Tomas Havranek: When inverting your values, if most of the literature does it in the inverted way, you should also do the analysis on the original parameters and then invert them. Generally, the delta method is used when, for example, you use inversion and need to compute the standard error, or when you have different functional forms of your x variable. You can use the delta method, which is based on the Taylor series expansion, but it's really an approximation.
You will have some measurement error in your precision, which will probably work against finding publication bias in your method. It's tricky. You know it exists and that you need to use it when you have these complicated functional forms. But if you can avoid it, if you can work with the actual reported regression coefficient, then it's better and you avoid a bunch of problems.
01:11:04 Participant: [A participant asks whether publication bias should be analyzed in a subsample with economically meaningful transformed coefficients or in the full sample of original coefficients.]
01:11:30 Tomas Havranek: We always have it all.
01:11:32 Participant: [The participant clarifies that some coefficients require an approximation and others do not, and asks about using subsamples.]
01:11:42 Tomas Havranek: Yes, you're right. It depends on the fraction of the sample. In general, you will need to do this, especially when your dataset is small and you really need to include all that you possibly can. It's useful.
But if it's 50% of the data that you have to transform, be careful. On a similar issue, I'm still not completely 100% sure what really is going on when you do these transformations, which are common in meta-analysis, or standardization. People tell me it's okay to do one standardization, then do a meta-analysis, then go back or go to something else, or do two steps in between. That's why I like your project. I think it shows how tricky it could be. That's a similar issue along similar lines.
Tests disagree 01:13:01
Especially with Tom [Stanley], we always have this issue: "You show us so many methods for publication-bias correction. Which one should I choose when they disagree?" That's a good question. We don't have a clear answer. One answer could be robust Bayesian meta-analysis as the benchmark. Right now, we don't have any benchmark. I would say the benchmark is FAT-PET or PEESE: the OLS regression. That's the simplest way. But you would like some sophisticated weighted average, which is what robust Bayesian meta-analysis delivers. When it's ready for economics, I think that would be a nice benchmark. [Note, 2026: the current guidance on meta-analysis.cz uses RoBMA to correct for publication bias and MAIVE and Mathur's RTMA to correct for p-hacking. See meta-analysis.cz/maive/.] […]
01:13:52 Participant: [A participant makes a comment that is hard to hear.]
01:14:39 Tomas Havranek: There are two things. When you talk about economics and finance, I-squared is almost always super high, by default. That's not an answer to the question of what we should report. But is PET-PEESE in trouble? Almost always. Tom Stanley would tell you that, by default, we do multivariate PEESE, so we add these additional variables. The issue is not settled, but I think the counterargument is relatively clear. Maybe you should compute the residual I-squared and look at whether it's low enough for your PEESE to have sufficient probability.
01:15:37 Participant: I think that's an issue, though. I agree, obviously, with what you're saying about adding the additional variables. What you have up there is a standard presentation of addressing publication bias with a univariate regression. You've got a standard error and a constant term. It's precisely that situation where high I-squared impedes the performance of PET-PEESE.
01:16:04 Tomas Havranek: Right. But if you add study-level dummies, you capture a lot of this heterogeneity. It's not perfect, and you look at somewhat weird within-study variation, but you do take a lot of heterogeneity explicitly into account with these fixed effects. In this case, it does matter, but the corrected mean is not affected dramatically. It should give you some peace of mind when you want to focus on your OLS coefficient.
I think more work is needed on what the baseline should be. I would, of course, say MAIVE, the instrumental approach. Then, as Bob said, look at the assumptions of these techniques. PEESE works well when heterogeneity is small; other techniques work better when heterogeneity is large. In general, I think the consensus is that selection models like Andrews and Kasy or p-uniform* work better with heterogeneity, while meta-regression models work better when there is a lower degree of heterogeneity. I don't think it's fully true, but that's the general consensus.
P-hacking 01:17:45
I talked about p-hacking. If you have p-hacking on precision, that might, again, make your funnel plot biased, because even the top of the funnel can be off. It's not clear whether the bias is negative or positive; it goes both ways.
That's why, in any classical meta-analysis model built for publication bias but not for p-hacking, you can have substantial problems when there is a lot of p-hacking. This was the main presentation last week. Keep in mind that p-hacking, as I have said, I think is a great focus. If you want to work on meta-research in the future, I think p-hacking will be the most promising avenue because too little work has been done on it so far. It's obviously important because it happens in practice to some degree and can have big effects. This will be the final caveat, and you will have your own.
Omitted variables 01:19:35
01:19:35 Participant: I think you didn't discuss the fact that, in many cases, you have a limited number of studies. When you code variables, many of them are unique to only a small number of studies, so it's very hard to distinguish the effects of different study characteristics on the outcome. Every study is unique in many aspects. By definition, you then make an abstraction. As a consequence, you're always going to have omitted variable bias in your meta-regression. I think people now like to forget that.
01:20:19 Tomas Havranek: I think you're right. It's a trade-off. When you control for too few issues, it's not a good thing because you will always have some of the problems. But if you overdo it, you control for issues that are really idiosyncratic to these individual studies. Then it becomes like study-level fixed effects, or study-level dummies.
01:20:49 Participant: A combination of factors that actually points to one study.
01:20:53 Tomas Havranek: Exactly, it could be. My general recommendation would be not to overdo it with the number of characteristics we collect.
01:21:04 Participant: And let's not overinterpret the results either.
01:21:07 Tomas Havranek: Yes, because sometimes one of your variables is based on just a couple of studies. Even if you have statistical significance there, it could mean something specific to the study. It doesn't have to mean that the driver itself is significant in any conventional way in which we would understand it. I think that's a good takeaway.
In many of my papers, we have 40 or 50 controls because we were so anxious that we would miss an important data or method characteristic. I don't think it's a good approach anymore. I think it's better to use more judgment before you start. You really think harder: "Do I need this? Should I group it together?" I can never capture all of the differences, as you said. You don't have enough data points.
01:22:28 Participant: Yes, indeed. You have 50 studies, and by definition there are more than 50 characteristics that you can go through in each paper.
01:22:35 Tomas Havranek: Maybe it's not by definition, but in practice that's always the case.
01:22:45 Participant: Oftentimes, at the front of the paper, they'll list the variables, and virtually all the coding variables we're talking about are zeros and ones. Right away, you can tell how unique that particular variable is. If it turns out you have 100 observations and there are only four ones, that's telling you right away that it's something very idiosyncratic to those four studies. To some extent, you should be able to sum it up in the summary statistics.
01:23:15 Tomas Havranek: We have these guidelines on how to do meta-analyses in economics and finance. I think the level we recommend, the minimum, is 3% variation.
01:23:36 Participant: That's assuming that you have 100 studies. But in practice you have 50 studies, so 3% of 50 studies is one study.
01:23:47 Tomas Havranek: The way we give it is as a percentage of estimates. But of course, you can make the argument that estimates within one study can be highly related to each other. In life, you need to find some balance. I think your comment is really important.
Sometimes you have things like difference-in-differences identification that was used in just two or three studies out of 50. But you really want to have it there because these are the three new studies, and it's very important to use a proper identification technique in this field. I might want to make an exception and put it there, with a disclaimer that I should be careful in how I interpret it.
Conclusion 01:24:46
That's what I wanted to show you in this session: we have many problems, as in any empirical field. Many issues are the same, like outliers. It's not special to meta-analysis. It's the same for any regression or any empirical analysis. Many issues are the same in meta-analysis and medical research in general, identification problems and so on. But it doesn't mean that we cannot be hugely influential with what we do here.
On balance, if you do research, I think doing at least a couple of meta-analyses is very good for you individually. Even if you don't publish in a super high journal, sometimes it gets a lot of citations. […] I think it's valuable. First, it's useful to other people, but it's also personally rewarding.
01:26:06 Participant: I think it's useful because you see how other people do things and all the little decisions you have to make when you run a regression. It also shows that, when you look at one study, you shouldn't really put too much weight on it because there are so many small decisions you have to make that can affect the outcome of that one single study.
01:26:28 Tomas Havranek: Exactly. People often ask me when I give seminars: "It's all nice and beautiful, but why don't we just choose one study that you really trust?"
Why don't you base your policy implications on this one good primary study, the best one? The answer, of course, could be publication bias, but it could also be just coincidence. Even if there is no publication bias, you can have one study that has impeccable identification and just happens to have a large point estimate by chance.
01:27:21 Participant: That specific sector, or that specific country, or that specific whatever. Or let me rephrase it. You have random effects, which reflect the reality out there. The one study that was really incredibly well done happened to pick one true effect from that distribution of effects, and that's what you ended up doing.
01:27:39 Tomas Havranek: That's the same. I think that's why we need meta-analysis. In medicine, meta-analysis is the gold standard of evidence. Then there are RCTs, and below that is observational stuff.
In economics and finance, I think we still have plenty of room to do better meta-research, to be more useful to society in general, and also to get more recognition from our colleagues in the profession.