Transcript. Lecture 7, Publication Bias, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.

Lecture 7. Publication Bias 00:00:00

Student presentation 00:00:28

00:00:28 Participant: [A student presents a meta-analysis of income inequality and economic growth.]

Here it says the number of estimates, but the number of studies was only about 50. I don't know, maybe, in your experience, isn't it typical that a couple of studies have many estimates, but most studies have a fairly limited number? Shouldn't we always also indicate how many studies these estimates are coming from?

00:07:46 Tomas Havranek: That would probably be a good idea in many cases.

00:07:51 Participant: Do you have the same experience? Whenever I look at this, there are always one or two papers with a huge number of estimates, and that can skew everything enormously.

00:08:01 Tomas Havranek: That's right. It is commonly highly unbalanced.

00:08:06 Participant: That's a feeling I have with this whole meta-analysis approach: we hide away how few papers we have by always focusing on the number of estimates.

00:08:23 Tomas Havranek: It would be more informative to also provide the number of papers.

Did you try to do it [the replication]?

00:13:08 Participant: [The student answers.]

00:13:32 Tomas Havranek: You don't have to do it, but you can check whether the code provided can give you the same main results. There should be a data file and R code. You don't really have to know much R to run it. I think they have data and code online, if I'm not mistaken. I think you nicely summarized what the paper is about.

00:14:13 Participant: [Participants discuss whether an estimate of 0.03 is economically meaningful.]

00:15:02 Tomas Havranek: […] This might be a nice topic for a project to convince the readers. It seems it might be economically significant in some cases, but in statistical terms, it's really very close to zero. I'm also not sure: I think these are partial correlations, but I might be wrong.

Thank you again. It is a nice project. […]

Intuition 00:16:08

I will talk about publication bias, which will partly overlap with my seminar on Friday. Most of you were there on Friday, so you will hear it again. By the way, many of you already know quite a lot about this, so we can talk about some of the nuances. Still, for the benefit of the rest of the class, I want to go through it from scratch. Of course, because there are so many models of publication bias, we could choose one model and spend one or two hours just talking about it.

We don't have time for that. I will go through the intuition, what we want to do, and then briefly summarize the models that are most commonly used. […] What does it mean when we say publication bias?

It means that what we see published, either in journals or in working papers, is not a good reflection of the results people get when they run their regressions or do their estimations. The results are selectively chosen to be reported and selectively chosen to be highlighted. One way to imagine it is to think about a case where someone comes to you and says, "When you eat jelly beans, it causes acne or cancer or whatever."

You go and test it. You find your data; it would probably be an observational study. You run your regression and find no link between eating jelly beans and acne or whatever. Because there is no link, it's fine. But then the same people might tell you, "It's not about jelly beans in general, but about the specific color of jelly beans. Go and try each individual color. What happens when you eat green jelly beans, pink jelly beans, or brown jelly beans?" You test the effect of eating jelly beans on your health for different colors of jelly beans.

Again, you find no significance. You do not reject the null hypothesis of no effect at the 5% level. But if you test a large number of different colors, most likely at some point, as here with the green one, you will find a significant effect just by chance.

You have 20 or more different tests, almost all of which are non-significant: no results, no strong results. But at some point, you're going to have significance. If you want to make your study more interesting and easier to publish, these could be the results you focus on. Then it gets into the newspaper: "If you eat green jelly beans, it causes acne." This is an example I use when I try to give high school students a short intuition about meta-analysis, because it's easy to relate to.

Paxil scandal 00:21:00

But we can, of course, imagine examples that are much more sinister, such as the one I mentioned on Friday with the Paxil antidepressant. You might have heard about it. Essentially, with the Paxil case, publication bias first appeared as a major topic, not just in statistical journals, but in general discussion in the media and so on.

When you have a new drug, a new medicine, and you want it to go on the market, you need clinical trials that would show it works. It is better than the placebo, at least. By the way, some people point out that this procedure for registering new drugs is really suboptimal, because what you should compare the new drug to is not a placebo, but the existing best functioning alternative. But to my knowledge, this is still the way it is done: you compare it to a placebo.

These clinical trials are paid for by the pharmaceutical companies, in this case GSK. But at least in this case, it was later shown, when the drug was prescribed on a massive basis, that other studies, or even other parts of the same clinical trial in the same study, were hidden and were not submitted for publication.

These additional results would show much less efficacy and much more potential for negative side effects. It was a big scandal in the media as well. There was a big lawsuit.

Then there was an independent reanalysis of the key study [Study 329, on teenagers]. It showed that if you don't just focus on the significant results, on the ones which are published, Paxil has no effect beyond placebo [in teenagers]. That's not very good, but it's worse: it has side effects which make suicidal ideation more likely for teenagers. You would have an antidepressant drug which was supposed to make you feel better if you suffer from depression, but which wouldn't work. It could even make you more likely to think about suicide. This was terrible. Here you can imagine the worst-case implications of publication bias.

That was the main thing that started pre-registration in medicine. Nowadays, if you want to do an RCT and then publish it in a medical journal, you need to register first.

The results are not hidden. Even if you don't submit your paper for publication, we can still see the results somewhere. We now have the same thing in economics. If you want to do a field experiment and publish in AER, you have to register it.

Pre-registration has now expanded beyond experimental research to observational research too. The problem with observational research is that it's very hard to guarantee that you don't already know what the results could be at the time when you submit your pre-registration or pre-analysis plan. With experimental data, it's much easier because you haven't done the experiment yet. You have nothing to base your analysis plan on in terms of results. But if it's observational, it's much trickier.

What publication bias implies is that when you do a meta-analysis, it's not enough to just look at the mean of the reported results, even if there is no heterogeneity. You somehow need to take publication bias into account. In theory, you have two responses. Response number one, as we have discussed, is pre-registration, which could work for experiments. Response number two is meta-analysis, which means you use some statistical techniques to correct the literature for publication bias.

No bias, no correlation (?) 00:27:28

This is the topic we will discuss. All the techniques which I will talk about today are based on the assumption that if you have no publication bias, you have no correlation between reported effects and standard errors. That's the key assumption of almost all publication bias correction techniques in economics, but also elsewhere. I put a question mark in the slide title [No bias, no correlation (?)], because we will talk about this assumption tomorrow in more detail: whether it really is an assumption we should be making in economics. Nevertheless, it's important.

Imagine that you are doing a meta-analysis of regression coefficients, as in the beauty paper which I presented on Friday: the effect of beauty on your salary, your earnings. These are regression estimates. In almost all settings, when you estimate this, for example by OLS, to make it simple, the implication of your estimation is that the ratio of the point estimate, of the beauty premium in this case, to the standard error has a t-distribution.

Or it could be a different type of symmetrical distribution. When this is the case, your estimates and standard errors should be statistically independent. I'm not saying this always holds in reality, but the assumption makes sense in many ways. I will try to show either tomorrow or next week that, for instance, if you use instrumental variables, it doesn't work. I think Bob and I discussed the paper by Keane and Neal in the Journal of Econometrics, which shows that with instrumental variables you have a systematic dependence between estimates and standard errors, which few people are aware of.

I wasn't aware of it until I read the paper by Keane and Neal. But let's leave it aside for now. We will come to it later. This is the assumption we really have to rely on in meta-analysis, also if you want to use the simple fixed or random effects which we discussed last week, the inverse variance weights. You need to assume there is no relationship between the estimates and the weights.

00:30:30 Participant: [A participant asks whether independence requires all the OLS assumptions to be satisfied, and whether omitted variable bias or measurement error could create dependencies.]

00:30:45 Tomas Havranek: Sure. If this is not specified well, if you have a correlation between the beauty variable and the error term, essentially anything could happen. I feel bad about defending this assumption because I don't like it at all. The main point of my current research is to point out that it's not a good assumption to make in economics and finance especially, or in observational research in general. But it makes some sense. You can say that if my model is correctly specified, that's the implication of the properties of OLS.

I like to quote this paper by Card and Krueger from 1995 because I think it's probably the first one, or certainly among the first ones, which really stress this mechanism. There should be no correlation if there is no publication bias, and then they test for a statistical relationship. People in the meta-analysis literature would typically call it Egger regression, after the Egger et al. paper, but I think it's incorrect. I think Card and Krueger do the same thing, but they did it first.

We can have two basic types of publication bias, and of course we can also have both. People can prefer to report estimates which have the intuitive sign. For instance, if we go back to the effect of beauty on salary, people probably expect it to be positive. It makes more sense. Some people might say the opposite: it's easier to report completely surprising results. You could also have some of this effect, but it's a bit hard to see in meta-analysis. Typically, in meta-analysis, we see the opposite.

We see a preference for results which are intuitive. Then, of course, everyone could have a preference for statistical significance. If you have more stars for significance in your main regression tables, that's a bit better for you. That has always been my intuition. When I was younger, I liked these significance stars. Nowadays, some journals like AER, or essentially all AEJ journals, tell you not to put any asterisks in your tables, any significance stars. Just put standard errors and let the reader know. That's in response to concerns about publication bias. Explicitly, it's about a paper by Abel Brodeur from 2016. That's the reference that they use.

00:34:26 Participant: [A participant asks whether studying contexts where an effect is more likely, such as beauty among sex workers rather than blue-collar workers, is covered by Type I or Type II publication bias, or whether these types assume there are studies on the topic.]

00:34:51 Tomas Havranek: I think that's more like heterogeneity. I would not call it publication bias: if you focus on a context, or a subset of your data, for which it's more likely you will find something.

00:35:12 Participant: Would that be covered here or not?

00:35:17 Tomas Havranek: That would probably not be covered. My answer is no. For that, you would need to take a look at heterogeneity.

00:35:26 Participant: These techniques do not correct for that issue?

00:35:31 Tomas Havranek: No.

00:35:40 Participant: [A participant begins a follow-up question that is hard to hear.]

00:35:51 Tomas Havranek: I don't think we can correct for that. It belongs more to the heterogeneity part, which we will talk about in two weeks. For now, I completely ignore this systematic heterogeneity. If you ignore it in practice, you will have trouble. You will find maybe some sort of correlation which you assign to publication bias, but maybe it's not publication bias, maybe it's heterogeneity, or vice versa.

Simulation 00:36:57

For now, let's keep it simple. Let's assume there are no issues of the kind you talked about. I will show a very simple simulation where we have estimates on the vertical axis and standard errors on the horizontal axis. It's a fun example. It's not existing data. I decide that the true effect of some elasticity or beauty effect is 1. I set it to 1 in my simulation. Then I simulate studies or estimates which try to estimate this parameter, this elasticity. They all have the same functional form, just simple regressions. They only differ in the data set they have. I simulate random data sets.

Some studies have really large data sets. They have a lot of observations, a big sample size. They are able to estimate really precisely the underlying effect of one. Other studies have very small data sets. They have a lot of dispersion and large standard errors. There is nothing special about this graph. If you collect these estimates for your meta-analysis and then regress estimates on standard errors, you find no relation, no correlation. That's what you expect if there is no publication bias. There is no publication bias: all the studies are reported.

But if the researchers, the authors in the literature, don't like to publish negative estimates, maybe because they make no sense, what we see is a positive correlation between estimates and standard errors. You can use the slope of the regression as a test for publication bias. That's the idea of Card and Krueger, Egger et al. I will show how it is connected to funnel plots. It's a very simple test, but the one most commonly used in all the meta-analysis literature about publication bias. The slope measures publication bias.

I think it was Thomas Stanley who went one step further and said that you can use the intercept of the regression as an estimate of the true elasticity or true effect corrected for publication bias. You can see that it's not an exact match. When you just run a linear function, you will have an underestimation. What you need is some sort of polynomial. I will talk about how the quadratic function works reasonably well. When you look at the intercept, you come close to the underlying true effect.

That was the first type of publication bias, selection based on sign. Then you can have the second type, which doesn't care about sign, but is selection based on statistical significance. You will only publish estimates which are significant at the 5% level. You will achieve statistical significance when your estimate is big enough to offset the standard error. Depending on the threshold you use, if it's 5% significance, you need the statistic to be 2 or 1.96, or bigger. You need the ratio of estimates to standard errors to be at least 2. You will only see the significant ones. Again, if you run the regression, it will have a positive slope.

In this case [this simulation, with selection on statistical significance], it's okay to have just a linear approximation, and you will perfectly capture the underlying effect beyond publication bias. Do you have any questions at this point before we move to formulas?

00:42:01 Participant: If the true effect is zero, under the null assumption there's no effect, then it should be symmetric, right?

00:42:11 Tomas Havranek: In this case, if it were zero, we could still have plenty of publication bias, but of course this test would show you a flat regression line. There are two ways to look at it. The first interpretation would be that you need a different test because there is publication bias, but you don't find it.

What I mean is that if the true effect is zero and I publish just positive significant and just negative significant effects with the same likelihood, then there is no bias on average. Even though the published estimates can be heavily selected, there is no bias in the mean. On average, if I just look at the mean reported effect, it's fine. There are situations in which you would like to go further and also investigate a selection process. But if your goal is just to say, "This is what the literature seems to imply. Is it biased relative to what we should observe if people reported everything?", then the answer is that it's not an issue.

00:43:51 Participant: [A participant begins a question about the simulation; the rest is hard to hear.]

00:44:13 Tomas Havranek: The simulation is just my idea of how to show how it works. I would put it in a different way. I would say this specific test, the regression of estimates on standard errors, has low power against some types of publication bias. For instance, if the effect is very small, close to zero, and you just look at statistical significance, there could be cases in which you have essentially no power.

But these are also the cases in which there is selection, but it does not create bias in the overall mean.

Now I use the same simulated data and switch the axes. I have estimates on the horizontal instead of the vertical axis, and now I have precision, which is 1 over the standard error, on the vertical axis. It is very similar, but the axes are switched, and this is inverted. These are the same data. I do not really know why I do it, but this is how people have reported it in meta-analysis for decades, at least since the Egger et al. paper.

00:45:47 Participant: I often wonder why people do that, because I find that far less intuitive than the other graph.

00:46:04 Tomas Havranek: I will give you an example. In the MAIVE paper, which many of you have seen and which I will also talk about, a paper on meta-analysis methodology for p-hacking, we used a figure like this [estimates against standard errors], as an illustration, because it is immediately apparent what the regression is, and it is very simple. For us in economics and finance, this is much more intuitive.

But I do not know. That is what people use, so that is what they want to see. That is called a funnel plot, by the way, because it looks like an inverted funnel.

00:47:06 Participant: [A participant comments on whether the graph looks like an inverted funnel.]

00:47:29 Tomas Havranek: I think that is a useful discussion. I also do not know. By the way, that is a different thing. A funnel plot with estimates on the horizontal axis and some sort of standard error on the vertical axis is more common in meta-analysis in other fields. This particular funnel plot, where we have one over the standard error, so we have precision, I think is more common in economics. I think that is because, for many reasons, people want to give even more weight to the most precise estimates. When you do it this way, you stress the precise estimates even more.

I think that is connected to how reasonable it is to use inverse variance weights in many situations, and how reasonable it is to assume that there is no correlation between estimates and precision in the absence of publication bias. We will get back to it. To be honest, for me this would be the most intuitive way [estimates on the vertical axis, standard errors on the horizontal axis], which easily connects to the statistical tests. It is easy to grasp.

We have many different strange graphs in economics. The demand and supply diagram, depending on how you think about causality, has P, price, on the vertical axis, which is not intuitive for many people. They would like to switch it as well. Again, this is the same simulation, the same data. We switched the axes and inverted the new vertical axis. Now, if there is no publication bias and you take the average of all estimates, you get one, which is the true mean value, the true parameter that I set in my simulation. It is fine. If you collect your estimates and take the mean, it corresponds to the true effect, and you are fine.

But if you have publication bias, so people do not report negative estimates, and you take the mean of the reported estimates, the positive reported estimates, because these are the only ones you see, you have a much bigger mean. It can be 70% bigger than in the previous unbiased case [in this simulation]. Of course, depending on the type of publication bias, if you also get rid of insignificant results, you have an even bigger bias. This is the problem: the difference between the mean of the estimates you see published and the true value of the mean that you should observe if all of these estimates were reported. This is publication bias: the difference between one and the new average you get from the studies you actually observe.

Now we use the same intuition from these graphs and do the regression we talked about before. Again, as Bob said, it is better to go back to this graph, where it is more obvious: estimates on the vertical axis, standard errors on the horizontal axis. This is the regression we keep in mind when we think about Egger regression.

00:51:36 Participant: One thing I was thinking is that we would also avoid having unreasonably big or unreasonably small estimates: publication bias. If the coefficient is very big or very small, people say, "I have to do some additional stuff."

00:51:55 Tomas Havranek: You mean the ones you collect?

00:51:57 Participant: Yes. No, the ones you make. When you run regressions and want to publish something, if your coefficient is too big, you are also going to continue to work to produce it so it becomes feasible.

00:52:11 Tomas Havranek: You might have many different types of selection rules or p-hacking rules, which are really hard for me to model. Here we use first-order approximations, the most fundamental ways we can simplify. It is really reduced-form thinking here. You could make some sort of selection model in which you would have many different types of potential behavior, and then you would estimate the weights for each behavior, but that is really hard to do in practice. Maybe in some specific cases it could be a good idea, where you have strong reasons to believe there are four different types of behavior, and then you would try to assign probabilities to these different types of behavior.

Funnel asymmetry test 00:53:06

This is reduced-form analysis. We take any correlation between estimates and standard errors as a potential sign of publication bias, and I say potential. This is the regression we want to estimate. The slope captures publication bias. Under some assumptions, the intercept captures the mean corrected for publication bias. You saw that the regression was heteroskedastic.

What does heteroskedasticity mean? You know from your econometrics classes with Bob, I guess, that the variance of the response variable, the dependent variable, is systematically related, in this case, to the size of the independent variable. As a solution, you can use weighted least squares, in which you give more weight to the observations that are more precise [those with small standard errors].

This is the version of the regression we commonly use in practice, the one weighted by inverse variance. Again, as always in traditional meta-analysis, the inverse variance weight is used not to get rid of any bias, but to increase your meta-analysis precision. This was the case for the fixed effect meta-analysis estimator. It is also the case here. This is a version of the fixed effect model. Some people would take it further and say that the inverse variance weight per se can help you reduce publication bias by giving more weight to the precise estimates, which are less likely to be biased.

That is a completely different thought. The funnel asymmetry test is the test on the slope coefficient: we test whether it is zero or not. If it is zero, the funnel plot is going to be symmetrical. If it is not zero, we will have some sort of asymmetry, which again is probably better seen in the previous form in the figure here. But it is the same thing. Asymmetry of this scatter plot is the same as asymmetry of the funnel plot.

In FAT-PET, the precision effect test is a test for beta, for the intercept, whether it is systematically different from zero or not.

PEESE 00:56:30

Again, I will go back to this regression here. We mentioned this problem when we try to estimate the intercept. For some types of publication bias, you will have underestimation of the true effect. The intercept is below one, below the true effect. What you need is some sort of quadratic function, and Tom Stanley and Chris Doucouliagos show that the quadratic function works quite well in many situations.

They call it PEESE, the precision-effect estimate with standard error. In the weighted form, we do not include an intercept, which we had in the previous case. Instead of an intercept, we have a term with the standard error. That is called PEESE. That is the most commonly used correction for publication bias in economics and finance. This is the benchmark, which in many cases is relatively hard to beat because it is very simple. You do not need to estimate many parameters. It is just two parameters, a very simple function, a very simple model.

PEESE, again, is the benchmark. [Note, 2026: the current recommendation on meta-analysis.cz is RoBMA to correct for publication bias and MAIVE to correct for p-hacking; see meta-analysis.cz/maive/how-to/.] Then you have related techniques that are also based on the funnel plot, on the same idea that what you want in meta-analysis is to somehow estimate the mean of the most precise estimates, because you assume that when you have a lot of precision you will be close to the true effect. You will probably have statistical significance anyway if your precision is large enough, so you will not be affected by publication bias.

WAAP (Ioannidis et al., 2017) 00:58:47

Another commonly used technique is the weighted average of adequately powered estimates. What does it do? First, you need to somehow estimate roughly what the true effect could be. I think Ioannidis et al. recommend UWLS, but you can choose any of the basic meta-analysis means we discussed last week. Once you know what the true effect is, which you do not, but you take some proxy for it, you can compute power for the estimates you have collected.

You have the truth, the significance level, typically you would use 5%, and the sample size. That should be all for the computation of power. Again, you know power from your econometrics classes. WAAP is simply the inverse variance weighted mean of all estimates that have power of at least 80%, with 80% being a benchmark commonly recommended in the statistical literature. Very often it will give you results close to PEESE. The slight problem here, for me at least, is that you need this true effect estimate before you start.

01:00:20 Participant: Why do you need that? If you know the true effect, you know the true effect.

01:00:23 Tomas Havranek: But as you will see, many estimators work like this. It is dependent on that because, of course, if you start with the assumption that the estimate is large, most of your estimates will be powerful enough. You will use all of them, so in the limit you will approach UWLS. If you assume first that your estimate is 1000, then probably all estimates will have enough power, and you will take the weighted average of all of them. Then you will have UWLS.

01:01:16 Participant: You say the true effect is 5, and then you do your estimates and get 4. But then it already rejects your original true effect.

01:01:27 Tomas Havranek: That is it. Again, I think it is important to be aware of how it works here.

01:01:39 Participant: [A participant comments on the assumption of only one true effect. The rest of the comment is hard to hear.]

01:03:21 Tomas Havranek: That is true. But on the other hand, I think you can easily extend, tweak or change the estimator to take what you said into account. You can take an external benchmark for power, like Cohen's d, which is typically used. When you write a grant application for an experiment, you need to provide some ex ante estimates of power. You could do something like that, which would make the estimator less circular and even more conservative. I think this would be a small parametric change, but the logic behind it makes sense to me.

Stem-based correction technique (Furukawa, 2019) 01:04:14

In a way, all of these techniques are similar because they look at the most precise estimates. One of them is the stem-based technique by Chishio Furukawa, which is non-parametric. The model itself is complicated and technical, but what it does, in a sense, is this. When we go back to the funnel, you focus on the most precise estimates, if you believe the funnel plot is the best idea to work with. But you also know that if you have just a couple of the most precise estimates, you get rid of a lot of information.

You have a trade-off between bias and variance. If you look at the most precise ones, you have no bias, probably. You can include more studies and decrease variance, but you can increase bias because you have more potential for publication bias when you decrease precision. Chishio shows how you can reasonably use the trade-off between bias and variance to construct an objective function [the mean squared error], which you would then minimize.

Based on this, you will find a cutoff for how many studies you want to focus on, and you take an average of these, again, in various variants. Some of you maybe know that there is a paper by Tom Stanley, Chris Doucouliagos and Stephen Jarrell in The American Statistician. It was about a paradox: when you get rid of 90% of the data, you can actually get a more accurate estimate. This paper I just mentioned takes 10% of the most precise estimates, and you can think of the Furukawa stem-based technique as a way to endogenously estimate the ratio.

That is the intuition we should keep in mind. We take X percent of the most precise studies. What should X be? Either you say 10%, because you had this paper by Tom and Chris, or you can estimate it using non-parametric techniques, which is what Chishio does in his paper, which is still unpublished. Hopefully it is going to be published well, and I think it is a very clever idea.

Endogenous kink model (Bom & Rachinger, 2019) 01:07:24

A similar idea, but for a different estimator, is provided by Pedro Bom and Heiko Rachinger in their RSM paper from a couple of years back. They start with a PEESE model. Again, I will go back to the scatter plot here. The PEESE model would be a quadratic regression here. But as you can see, up to here in this case, we have no publication bias. Publication bias only kicks in once we have enough imprecision to get any negative estimates. For this segment here, we should have no publication bias at all. There is nothing going on. It only becomes a problem once you have any negative estimates that you do not want to report.

The true function is not quadratic. It is horizontal first, and then it bends upwards. There should be a kink somewhere. The endogenous kink model tries to estimate the location of the kink. I will go back to the regression. To be able to estimate the location of the kink, where the horizontal segment ends, you need a first estimate of the true effect.

In a way, it is similar to what WAAP does. First, you need a first estimate of the true effect. Then, because you observe the dispersion in your data, you know where the 5% significance threshold is. You will use it to estimate the location of the kink. That is definitely an improvement in terms of theory. But I have two observations. First, in practice, it tends to be really close to PEESE. Since PEESE is simpler and is just a quadratic regression, you may ask how useful the kink model is in practice. Second, before you start, you need a first estimate of the true effect to be able to estimate the kink.

01:10:32 Participant: To be able to find the kink, don't you need lots and lots of observations?

01:10:36 Tomas Havranek: There is an additional parameter.

01:10:41 Participant: [A participant asks about using all the estimates to estimate linearity with PEESE, compared with having estimates to locate a kink.]

01:10:52 Tomas Havranek: You will have fewer degrees of freedom.

01:10:56 Participant: [A participant asks whether the kink model could also include other variables.]

01:11:15 Tomas Havranek: Yes, I think so. But I have never done it.

You would have to divide your variables into two groups. One group, which you assume is related to publication bias, would go into the slope. You could have different slopes based on different characteristics of the data, maybe other techniques. Then you could have different intercepts, which is probably what you are mainly concerned about. I do not see why this would not be possible in this framework.

Essentially, you would probably have to estimate different locations of the kink for each subgroup, for each characteristic. I guess computation could be more difficult. It should be possible. But in PEESE, it is very easy. I think the endogenous kink model is more rigorous. To be honest, I like to report it because I think it is a very clever idea: take PEESE and make it fit the way the process actually works. So far, we have only talked about models based on the funnel plot or on regression. We were trying to find the top of the funnel, the most precise estimates.

Selection model (Andrews & Kasy, 2019) 01:13:37

You can think of it as trying to estimate a potential study that is infinitely precise. Now we completely abandon the funnel plot and look at something else: a selection model. I put a reference here to Andrews and Kasy because that is the most commonly known selection model in economics and finance. It is published in AER, and it is very rigorous and detailed.

As I understand it, it is similar to previous selection models by Larry Hedges and other people. I think it was mostly Hedges; he is the person who really developed these kinds of models. I do not think I have enough information to comment on it. Andrews and Kasy have a different procedure for using maximum likelihood to estimate the different weights.

It is a completely different approach. In the selection model, we try to estimate the probability with which each type of estimate is published or not. For example, positive, insignificant, negative, and so on. You can choose different brackets for which you want to estimate different publication probabilities. Then you need to make some assumptions about the distribution of the underlying effect. Typically, you would say it is probably a random effects model, and the effect is going to be normally distributed. Sometimes you can also choose a t-distribution, so you can put more weight in the heavy tails of the data and allow for them.

You use maximum likelihood to estimate the weights. Once you have the weights, you can use them to reweight the published estimates. If you compute that estimates which are, for example, negative have a 50% likelihood of being published, you will give double weight to the negative estimates in your data set. That is how selection models work. It is very simple.

But the procedure is not simple. You commonly need a large sample for maximum likelihood to work well. Sometimes it works perfectly, but there are also cases when it will not converge at all, even if you have a large sample. Nowadays, many people outside economics would say selection models are better, especially when we have heterogeneity. That is based on a number of simulation studies by Carter et al. and other people. We will talk about that later, but I do not think it is so simple.

Another reason why some people prefer selection models is that they can be more tied to theory. You can have different scenarios in which people either publish or hide the estimates, and you can incorporate them into your selection model. That is much more difficult to do when you look at the funnel plot. In general, we have two competing approaches: techniques based on the funnel plot and selection models. There are dozens of different selection models.

Example: labor supply elasticity 01:18:16

This is about economics and finance, so we probably want to use the Andrews and Kasy model because the referees will also know about it. If you do not use it, they will ask, "Why do you use something else?" Here is a brief example from our applied paper on labor supply elasticity. This is the funnel plot. You can see it is asymmetrical, which could mean some of these negative or insignificant estimates are not reported. Then the mean of the reported effects is too large; it is exaggerated. Even if you do not want to use techniques based on the funnel plot, it is always useful to show a funnel plot because people are used to it, and the asymmetry is typically quite apparent if you have publication bias.

Results vary: RoBMA? 01:19:21

You do not have to read the numbers. I want to show you one thing. You have these different techniques for publication bias correction. We discussed it the other day. What will happen to you quite often is that you have really different estimates. In this case, they are not so different, but still, an estimate can be three times bigger in some cases.

01:19:55 Participant: But still, 0.13 versus 0.25?

01:19:59 Tomas Havranek: It is a difference. What should we do? There is no consensus on which technique is the best one, which is the benchmark. I would say PEESE is the benchmark because of its simplicity.

01:20:18 Participant: [A participant comments that other covariates could give even more different estimates.]

01:20:30 Tomas Havranek: Of course, this is before any considerations of heterogeneity.

01:20:36 Participant: When you consider heterogeneity, do you include all the variables, or do you exclude bad controls? You end up with as much variation in your estimates as in your original study.

01:20:51 Tomas Havranek: You can choose one, of course. What I think will be really useful is robust Bayesian meta-analysis. I do not have it here in detail because the statistics are complicated. It is an estimator, or a set of estimators, by František Bartoš and colleagues from the Netherlands.

They estimate many models in a Bayesian framework, with different publication bias corrections: selection models, PET and PEESE. I think they have 36 different models, also including simple fixed effects and random effects. They give each model a weight which corresponds to how well the model fits the data. You can think about adjusted R squared, for example, from a regression framework, as if you gave more weight to models which have higher R squared or lower information criteria.

They do it in a Bayesian framework because it is much easier to compute, but you need some priors. Currently, we are still not able to use it for economics. [Note, 2026: RoBMA is now used for economic estimates, and meta-analysis.cz recommends it as the correction for publication bias; see meta-analysis.cz/maive/how-to/.] You can use it in economics if you work with Cohen's d, for example, because the priors need to be specifically set for a particular set of estimates. I think they have it for Cohen's d and for log-odds ratios, but it is not clear how to use it when you have an elasticity, for example. I know František Bartoš will work on a paper which will make it more feasible to use in economics, but so far, to my knowledge, it is not.

If you are looking for a benchmark, even though it might be wrong for many reasons, this is a good candidate because it is a sophisticated average over many different models.

01:23:25 Participant: [A participant asks why it cannot be used for elasticities.]

01:23:29 Tomas Havranek: You need priors tailored to elasticities for technical reasons related to these Bayesian procedures.

01:23:41 Participant: [A participant asks whether applying it in R to elasticities would still run.]

01:23:45 Tomas Havranek: It would run, but it would give you a nonsensical result.

01:23:51 Participant: [A participant asks whether one would recognize the result as nonsensical.]

01:24:05 Tomas Havranek: That is a disclaimer, but relatively soon it should be ready. When it is ready for use with elasticities and other parameters that we have in economics, such as correlations, I think it will be a very nice benchmark.

Maybe then, instead of running all these different estimators, you can run RoBMA, robust Bayesian meta-analysis. Then you can choose specific extensions. I will talk about our instrumental variable approach later. Then you can compare. I think it would be really useful to be able to use RoBMA easily in economics.

I did not talk about the p-uniform* paper by van Aert and van Assen. It is essentially similar to selection models, like the one by Hedges and the one by Andrews and Kasy, but it is a simplified version in the sense that it uses fewer parameters. When you have a smaller sample, it is more plausible to use it than a full-scale selection model a la Andrews and Kasy, for example.

We have seen all the rest. In practice, when you estimate a PEESE or PET regression, I like to use fixed effects in the econometric sense, which means I include dummies for different studies because we have many estimates from different studies. You can ask what that means, because essentially we estimate publication bias within studies, which is a bit difficult to explain. We have a paper suggesting that maybe it captures more p-hacking than publication bias if you look at within-study choices and within-study variation. When you compare across studies, it could be publication bias. But I will talk about p-hacking tomorrow.

01:27:04 Participant: [A participant comments that including a dummy for each study means having a separate constant for each paper rather than one constant.]

01:27:15 Tomas Havranek: You look at the mean.

01:27:17 Participant: For each paper. Then you take the mean of the means of these dummies.

01:27:24 Tomas Havranek: You run a normal fixed effect estimator, as you would with panel data.

01:27:35 Participant: I understand that, but in the end, the main thing of interest is your mean beyond bias. If you use panel data, you do not have this constant term anymore.

01:27:48 Tomas Havranek: You have it, but it is an average. You can look at it as a random effect estimator, but more flexible, because you do not estimate one parameter; you have a set of dummies. The big benefit is that you get rid of all concerns about differences in study-level quality. It comes with a big disclaimer, but I think it is useful to at least look at it.

The precision weight means the standard inverse variance weight, but it might also be useful to use different weights. In applied meta-analysis, we often use weights based on the number of estimates reported in each study. For the reasons you mentioned, some papers have many estimates, and their weight in your results would be quite substantial. It might make sense, at least as a robustness check, to give each study the same weight, irrespective of the number of estimates reported in the study. I will talk about the IV, the MAIVE estimator, later.

Bias caused by preference for positive estimates 01:29:19

It might also be nice to plot a histogram of t-statistics or z-statistics to see which type of publication bias is more likely to be prevalent in your meta-data. In this case [slide 'Bias caused by preference for positive estimates'], look at the threshold for 5% significance. You see a jump, but not a large one compared to the rest of the data, and you see a big jump at zero. Probably the first type of publication bias, selection for the correct sign, is more likely to be going on here. […]

Questions and discussion 01:30:45

[To a question about who selects the results:] You mean in the literature that you collect? That is a good question, because it is not clear. It could be the authors themselves. Even before you submit your paper, you can look at your results and say, "I do not think I will ever be able to publish this. It is not really intuitive." You will not submit it. Or you might slightly adjust it, or adjust it more; that is p-hacking, which will be tomorrow's topic. Or you submit your paper, and the referee takes a look at it and says, "What's the point? It's insignificant, so nothing is going on. Reject." Or the editor can say, "We don't learn much new from it. We want something surprising and strong." So the decision comes from a combination of the editor and the referee.

01:31:52 Participant: The author as well.

01:31:53 Tomas Havranek: The author decides as well, because it's your paper and you can make a choice. It doesn't have to be conscious. You might not trust your results as much if they don't align with your priors. Again, this is not true for all people, but for many. […]

01:32:35 Participant: [The participant jokingly objects to how meta-analysis was presented among the ways of dealing with publication bias.]

01:32:56 Tomas Havranek: I knew you would raise this issue later. But this is also something you can do ex post.

01:33:03 Participant: [A participant comments that meta-analysis should be mentioned in the same suite of approaches, that it often is not, and that they do not understand why. The rest of the question is hard to hear.]

01:34:40 Tomas Havranek: Sometimes, even in my mind, I don't think there is a clear line [between publication bias and p-hacking]. For me, the line would be this: if you can still use these publication bias correction models to correct for it, and nothing else would remain, I would still call it publication bias. If not, if you need something else, then I would call it p-hacking.

01:35:13 Participant: [A participant welcomes this functional definition based on a set of correction tools. The rest of the comment is hard to hear.]

01:35:50 Tomas Havranek: There is plenty of confusion about what publication bias means compared to p-hacking. I used to do this as well, as recently as a couple of years ago. Sometimes, when I said publication bias, I would mean anything related to bias, including p-hacking. I would think it was similar, or that we could adjust for it using these funnel-based methods. I wouldn't fully appreciate that this does not apply to all these p-hacking cases.

There are some cases for which PEESE works well, and the question is whether we should call it p-hacking or not. For example, if you change your point estimate in response to how much precision you get, PEESE works perfectly fine, even if you have p-hacking. But I don't think there is any reference.

01:36:54 Participant: That might be a topic for future research.

[A participant says the discussion convinced them that p-hacking is different from publication bias.]

01:38:20 Tomas Havranek: This was publication bias. Tomorrow we will cover p-hacking, and then even more on p-hacking in a week, if you are still not fed up with my presentations, which use similar pieces of evidence.

But I think it's important.

01:38:40 Participant: [A participant comments on covariate effects and robust BMA; the precise claim is hard to hear. They also say they find it frustrating that there is not much analysis of which publication bias correction methods work best.]

01:40:08 Tomas Havranek: I wish I could tell you, but I agree, I don't think we do. I have to admit, I like PEESE because it's so simple. It seems to work. In my experience, in the 40-50 meta-analyses I have done, it never seems to be a big outlier when I compare these five, six, seven different techniques. I like PEESE a lot. But I agree: it's not a rigorous analysis.

I don't have any proof that we should use PEESE and treat all the rest as robustness checks. I don't know. The best I can tell you is that I'm looking forward to robust Bayesian meta-analysis in its full form, when we can use it for elasticities, for all the coefficients we do use in economic meta-analysis.

We can discuss the shortcomings and the problems, but I still think RoBMA should probably be the benchmark, unless we have some strong evidence that we should prefer something else. Hopefully, that could be MAIVE. […]

Corrections

Slips of the tongue corrected in the text:

  • [00:25:13] said "an experiment"; the text has "a field experiment".