Transcript. Lecture 8, p-hacking, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.
- Lecture 8. p-hacking 00:00:00
- Recall publication bias 00:01:45
- Key difference 00:09:06
- Example: education premium 00:20:28
- Spurious precision 00:29:41
- Example: Changing controls changes precision 00:33:56
- Changing clustering changes precision 00:37:05
- Funnel plot 00:40:10
- Conventional p-hacking 00:41:24
- Spurious precision 00:43:08
- Key meta assumption broken 00:44:06
- Options for the meta-analyst 00:45:53
- Meta-analysis instrumental variable estimator 00:58:24
- Extended MAIVE 01:04:01
- MAIVE reduces PET-PEESE in 70% of the cases 01:08:19
- Bonus: Maya Mathur's p-hacking correction 01:09:20
- Questions and discussion 01:11:58
Lecture 8. p-hacking 00:00:00
00:00:33 Tomas Havranek: Thank you again for coming. For some of you, it is maybe the 15th seminar or lecture I have given in the last couple of weeks. I appreciate your endurance. I have to say I like the topic of today's session, because I think that if someone wants to specialize in meta-research, p-hacking could be the way. That is because there is a lot to do here. We have seen a lot of work on publication bias, but much less on p-hacking. I will try to talk about why it is important, why it is really different from publication bias, and about some of the approaches which already exist. There is a lot of potential to work on this issue, to publish well and to have a lot of impact.
Recall publication bias 00:01:45
Let us start with a repetition of the main mechanism which we focused on yesterday, in the last session, when we talked about publication bias. There are many ways to correct, or try to correct, for publication bias with these selection models, in which you estimate the likelihood that insignificant estimates are published, that negative estimates are published. But we are mostly focusing on meta-regression, which means you look at something like the funnel plot, which is a scatter plot of estimates and their standard errors, or precision.
I am going back to the version of the scatter plot which is easier to interpret as a regression. I have the estimates on the vertical axis, and I have on the horizontal axis the uncertainty around the estimates, the standard errors. If for some reason negative estimates are not published, maybe because authors do not like to submit these estimates, maybe referees do not like them, maybe editors do not like to publish them, we only observe positive estimates of, for example, the beauty premium or whatever we are looking at. When this happens, when there is publication bias, which means publication selection bias, as it is sometimes called, some estimates are selectively omitted, selectively not reported, hidden in the file drawer.
When you run a regression of the estimates on standard errors, you will have a regression slope which is non-zero. In this case it is positive, because the bias is against negative estimates. But it could be negative if, for example, you focus on the price elasticity of gasoline. Then you would assume that the price elasticity of gasoline is probably negative. […] So you can have positive or negative bias. If you run this regression and you have a slope of the regression line which is not zero, you can take it as evidence [of publication bias].
Then the idea of Tom Stanley was that we can go a bit further. We can also use the same framework, the same regression, as an estimator for the true effect corrected for publication bias. So what is the underlying true effect if there was no publication bias? To my knowledge, it was Tom Stanley who came up with this idea. Let's have a good look at this. Now you see that when you run the regression and look at the intercept, it is pretty close to the true effect, which in the simulation was one.
That is how we set it in there, but it is not precisely one, so it is a little bit lower. The reason is that the shape, the pattern of the estimates, is not really a linear relationship. It is a nonlinear one: it approaches a flat line here, and then it goes up. So that is why, in practice, we do not use only the linear approximation, but also the second-order quadratic approximation. One way to put it is that PET-PEESE, the main meta-regression estimator, is a reduced-form, second-order approximation of how publication bias can work. And then we talked about these extensions, like the endogenous kink model by Bom and Rachinger, which is very similar, but it explicitly models the flat line in the beginning, then the kink, and then the linear segment.
But essentially all models which somehow use the funnel plot to estimate the meta-average, the mean effect beyond publication bias, are based on some version of this idea that we try to look at estimates which are really precise. The stem of the funnel plot, used by Chishio Furukawa, also consists essentially of really precise estimates.
Precision, the reported precision which you see in the papers you collect, is really key in meta-analysis, especially in the bias correction. It is also, of course, used as a weight, as inverse variance. But here the role is magnified: you use both, you weight by inverse variance, and you also try to explicitly, or at least implicitly, estimate a study which would be really precise, or infinitely precise, really an intercept of the regression line, which means you condition your mean on a standard error being zero, so an infinitely precise study. And this is perfectly fine, as long as the selection process works like this. Or it can also work in this way, which means people do not really care about the sign, but they do care about statistical significance.
So only statistically significant estimates, which means those which have t-statistics above 1.96 in absolute value, are published. Again, if this is the selection rule, the selection process, you will still have the same kind of pattern. You will have a positive regression line in meta-regression. And the intercept in this case, even in the linear case, is good.
Key difference 00:09:06
So far, so good. But today we are going to talk about something slightly more complicated. The complication lies in the fact that yesterday's discussion was about publication bias, selective reporting. It means that you only observe some estimates, positive ones, significant ones, and the rest are not reported. But if you can find the selection process, which is to say if, for example, in a selection model, you can somehow estimate the likelihood that an insignificant estimate is published compared to a significant estimate, then you just put more weight on the insignificant estimates compared to the significant ones, and voila, you have the mean, the meta-estimate.
As for p-hacking, the topic of today's session, we discussed last time how to define it. One way to define it is to say that in publication bias the estimates which you observe can be selectively reported, but on their own they are unbiased. Or they could be biased in a way that they are not done well, but they were not selected for this reason. They were just selected because you got the estimate and you either report it or not. With p-hacking, the individual estimates can be biased because you actively try to get statistical significance, or you can try to get positive estimates.
It is still not a clear-cut definition, I realize that, and I don't think there will be a clear-cut definition in the future, but you get the idea of what is going on here. So in this p-hacking framework, it is not enough just to say, I observe this number of published insignificant and significant estimates, I compute the likelihood with which insignificant estimates are published, and then I just put more weight on the insignificant estimates. So it doesn't work anymore. Well, maybe it can work with some types of p-hacking, but in general it doesn't work. Let me give you an extreme example. Suppose someone p-hacks in a way that completely cheats: they get their results, just erase them, and write new numbers, which give them a lot of statistical significance.
That is a very extreme example. That is almost never how it works, I think. But you can imagine that you can have bias in the way observational research is done, how regression is done. After a lot of work, you can essentially get whatever results you want. It does not have to mean that they have the same degree of defensibility, that you can justify them. If you just play around with the controls, with the subsamples, with the methodology, whether you use IV, what kind of instrument you use and so on, you can really change your estimates a lot, almost to the point of getting whatever you want in some contexts, in some scenarios.
So the approach here needs to be completely different, at least for selection models. Selection models do not really work here unless you have a good idea about how specifically p-hacking works. That is very hard to do, because we talked about a paper which has 12 p-hacking scenarios, and again, you can come up with 120 scenarios. So it is very hard to put them into a selection model and to make the selection model feasible to estimate.
There is one approach which has some affinity to selection models, and that is the approach by Maya Mathur, the right-truncated meta-analysis. I will talk about it only very briefly. What Maya Mathur does is completely ignore all significant estimates, and the assumption is that significant estimates can be p-hacked. [Note, 2026: strictly, RTMA sets aside the affirmative estimates, those significant in the expected direction, and fits a right-truncated normal distribution to the remaining, nonaffirmative ones.] Because I don't know what the function of p-hacking is, what the mechanism is, I cannot do anything about it. So I just ignore all significant estimates, and I only focus on the ones which are not significant, because I believe that if they are insignificant, most likely they will not be hacked. Now you can question the assumption, but to me it is possible as a first approximation.
What you try to do is observe just the insignificant estimates. Then you have to assume the underlying distribution of the true effects in the analysis; in her case you assume it is normal. Because you observe the insignificant estimates, which are just part of the distribution, you would like to fit the rest of the normal distribution based on this. The problem is that in many cases you will have a relatively small number of insignificant results, so it is very hard to do it using frequentist techniques like maximum likelihood. That is why Maya Mathur uses Bayesian techniques. You need some priors, but it is much more computationally friendly. But still, you need a good number of insignificant results. So that is one approach. I will focus on something else, which we have been working on and which is more grounded in the classical meta-regression. I think it is also more intuitive for an economist who uses instrumental variables.
00:16:29 Participant: What is interesting is when things that work for publication bias don't work for p-hacking. A really good example here is that with p-hacking, you focus on the insignificant estimates, but there are these other approaches to publication bias that focus precisely on the significant ones: the exact opposite. [The participant adds that it is hard to see how the two sets of papers arrive at opposite solutions.]
00:17:08 Tomas Havranek: It is exactly the opposite. So I think what you mean is p-curve and the original version of p-uniform, before p-uniform star. But you are right, it is completely the opposite. Until recently, people would mention p-hacking, but it would all be lumped together under the same umbrella, like some sort of publication bias or selective reporting, and the same cures, the same techniques would be put forward for both issues.
But as Bob says, it is actually very different, to the point that you can have one solution for publication bias which is the exact opposite of a solution to p-hacking. That is really striking. […]
I still want to talk about models based on funnel plots and how we can address some specific forms of behavior. In the extreme case I mentioned to you, where you just scrap what you have and you write your own numbers, anything can happen. You have no way to correct for it if all people do it, if all people just scrap results and put a random signal. Or maybe if just a small fraction do it, then you can find these suspicious estimates. So p-hacking is really, really hard to tackle. Some people say, okay, we cannot do anything really about it: selection models do not work, we give up, we have to assume there is no p-hacking. I think that is not very useful, because I would say that the p-hacking process is at least as likely as publication bias.
Most of you have already done some empirical work, and you know that you feel better when you get statistically significant results. It is not just about whether or not you send your assignment or your paper. Sometimes there are many choices you can make along the way, many choices which can make your estimates more significant or change their magnitude. So let's try to see how we can improve, or how we can extend, funnel-based techniques to take into account the p-hacking problem.
Example: education premium 00:20:28
I am using here an example from labor economics. Suppose you do a meta-analysis of the literature which looks at the effect of education on your salary: how much money you make if you have one additional year. Now this is a very large literature. It is an important question. Now suppose the true model has not just education on the right-hand side, but also some ability. So if you are, for example, more clever, you are likely to have more education, and you are also likely to earn more money.
It does not have to be because of your education; it can be because you are more clever. So this is a classical example of endogeneity bias or an identification problem in economics. Of course, you can have many other issues which would play the same role, so you can have many more omitted variables. You can call it the omitted variable problem, because the issue is that in most cases, when you have data for different people, you do not have data on ability. Ability is hard to observe. You can have some proxies, but you do not really measure ability. So what do the primary studies do in the literature?
Well, some studies, especially the older ones, just ignore ability. They omit the variable because they have no data on it. They ignore it. So they just do a simple regression. […] Now, if you do it, you are almost guaranteed to have an education premium which is too large, because now it also includes all of the effect of ability, of your intelligence. So it is an upward bias.
It is not guaranteed, but it is quite likely and plausible that you will also have too small a standard error. And why? Because you reduce collinearity in the model: you remove one variable that was heavily correlated with the important variable on the right-hand side. So you are likely to increase the precision of your estimate. So your gamma is too large and too precise at the same time.
Now, if you are luckier with your data, you will have some sort of proxies for ability. For example, if you use data from Finland: in Finland they have compulsory military service, and when you enter the draft you have to take an IQ test. So they have IQ results for millions of people, as I was told. When you do this kind of wage regression in Finland and you have access to the data, which is anonymized, but you can still link the IDs of different people, you can add a control for IQ.
This is not a perfect proxy for ability, but it will definitely capture some of it. So when you do it, you get closer to the truth with gamma. It will be smaller than in the previous case. It will also most likely be less precise, because you increase the amount of collinearity in the model. So your gamma will be smaller and your standard error will be larger.
So less precision, a smaller effect. You are much less likely to have statistical significance than in the previous case. Finally, you can do some sort of quasi-experiment. So you can use difference-in-differences, maybe some sort of RDD. People have done instrumental variables quite a lot here, so you can have an instrument that is exogenous in the sense that it is related to education, but not to ability.
00:25:30 Participant: [A participant suggests an instrument: living closer to a school.]
00:25:39 Tomas Havranek: Those who live closer are more likely to attend a school, they are more likely to get education, but it does not have to be related to ability. Anyway, if you do the quasi-experiment, you are likely to get even closer to the true value of gamma, so gamma is going to be even smaller. And again, most likely, though not always, your precision is going to be smaller, especially when you use IV, because IV is very noisy.
As you get closer to the correct way to measure it, you get smaller estimates and bigger standard errors. And of course, you don't have just three options. There are many ways to do quasi-experiments. You can choose typically from difference-in-differences, IV, maybe RDD, maybe something else. You can choose different proxies, maybe. So there are many different choices from which you can choose. What is happening in terms of the funnel plot is that you can move in the funnel plot. Suppose that you estimate it right, using a good quasi-experiment.
Let's just suppose that, in my case, this is what I would get if I estimate it using the correct IV. So my estimate is not statistically significant. But if I do something else, for example, if I completely ignore ability, or if I use a proxy, or if I use a bad instrument, I can get an estimate that is bigger and more precise at the same time.
Again, it is not guaranteed. But I am very likely to find a combination of changes, either in control variables or in methodology, or maybe by using a different subset of data, to move in the funnel plot, ideally beyond the significance line here, to have statistical significance and to make it easier for me to publish the estimates. So that's the idea. It's one of the potential mechanisms of p-hacking. The ways and forms of what could be going on are essentially unlimited. But this is easy to understand. This could be realistic. This is probably what is going on.
Now, if we again just do the quadratic meta-regression, PEESE, or something similar, and focus on the most precise estimates, even among the most precise estimates we can have those which are simply p-hacked, not correct. And that's why Mathur says: in my model I will ignore all the significant estimates, because they can be p-hacked, and I just focus on the insignificant ones. So the original approach doesn't work anymore [under this kind of p-hacking]. It could work, but we have no guarantee. In general, it's incorrect.
Spurious precision 00:29:41
In our paper on this issue [Irsova, Bom, Havranek and Rachinger, Nature Communications 2025], we describe the mechanism I explained before: you can change your estimation so that you have more precision, but we don't feel it's really right. We debate whether there is such a thing as a correct precision, but when you p-hack you can get more precision than you would have if you did not p-hack. So we call it spurious precision. And we do some simulations in our paper. I will not go into the details of the simulations. I will just say that we try different scenarios, different degrees of the p-hacking problem.
You can imagine that, going back to the example of education, earnings and ability, we play with the correlation between education and ability. If ability and education are not correlated at all, then you have no problem and you have no p-hacking. If the correlation is high, then you have more potential for p-hacking when you omit ability, or you use different proxies. We simulate it. You can choose different proxies, and in this case all of these publication bias correction methods fail at some point. They don't have to fail when the degree of p-hacking is relatively small. In the paper we also show that, even though a correlation of 0.8 sounds like a big number, it doesn't translate to a big percentage of selection on spurious precision versus point estimates. So it doesn't have to mean that 0.8 is an extreme value in all cases.
00:31:50 Participant: [A participant asks what the correlation refers to.]
00:31:56 Tomas Havranek: It's the correlation between the independent variable that you're interested in and the control. So this is a simple way to simulate. The black dashed line is just a simple mean. You can see that when p-hacking is modest, these correction techniques correct almost all of the values, maybe not all of them, but they are better than just taking the simple mean of all estimates. At some point, p-hacking makes these correction techniques actually worse. So the p-hacking or spurious precision can make the cures to publication bias even worse than the original disease. They can make that even worse than if you just do a simple average. That's a big degree of p-hacking, but it could happen anyway.
Before I go to the actual estimator, let me also give you some more motivation on why people could perhaps play with precision and not just with the point estimates. In the funnel plots we cannot assume that precision is given, and maybe people can p-hack it. But if they p-hack by changing the point estimates, and the precision stays the same, then you can rely on inverse variance weighting, you can put more weight on precise results, and then you can actually estimate.
Example: Changing controls changes precision 00:33:56
So this is an example from an Alan Krueger paper from the QJE, the Quarterly Journal of Economics, on the effect of class size on student achievement in primary schools. It was a big experiment. I think I mentioned it a couple of weeks ago. The interesting thing is that you have the treatment effect, small class versus regular-size class. And then you can add controls, student background, gender, teacher controls and so on. Essentially your treatment effect stays the same or changes a little bit, but it is a very small change.
But what really changes a lot is your precision. If you add these controls, which in this case are not related to the treatment, you only increase the fit of your regression and you increase your precision. So this is the opposite of the collinearity case that I mentioned before. Typically, when we add variables, we think that we have more collinearity and less precision in the model. But it could also be the other way around. Suppose you do a meta-analysis on the effect of class size on student achievement, which we did and published in the Journal of Labor Economics. Now you collect these estimates of the treatment effect, you collect the standard errors, and then some of the estimates have twice the weight of the other ones.
In this case, it's one paper, so it doesn't really matter, but you could have one paper which only does this, and then another paper would do this, with a different experiment, different data. And then you do a meta-analysis, and using the normal inverse variance weights, you would give, well, not twice, because you use the square of it as a weight, but 4 times the weight to a paper along the lines of specification 4 compared to a paper along the lines of specification 1. And it's not completely clear to me that it's what we want to do. So it's another example of how people can play around with variance. I'm not implying that Alan Krueger was doing that. He was just adding controls and seeing what happens.
Changing clustering changes precision 00:37:05
But this is one possible mechanism: you can just change controls and your precision changes. Another example is, again, we use the same data. Now, for simplicity, I ignore the controls. I use the specification number one, where we have no controls. It's just the variable for treatment. And now, again with no control variables, we only changed the way that Alan Krueger computed standard errors.
Some people would just use the cluster estimator in Stata. You could also argue which clustered standard errors you could use. There are at least several ways to compute them, as I have learned from you. So we could add more. Other people would argue that it's too strict to look at the class level.
And it's less defensible, but some people would just ignore clustering altogether. They would just compute heteroskedasticity-robust standard errors. And you can also just do a regression with no adjustment at all. So what you can see is that your precision could be three times smaller if you do it correctly, compared to a plain vanilla variance estimator, which means nine times more weight in meta-analysis. So that shows you the strength of the choices we make, not just on the point estimate, but also on precision.
Funnel plot 00:40:10
Once again, let's say this would be the unbiased funnel plot for your meta-analysis. The blue dots are significant, the hollow circles are insignificant, but all are published, so you can see all of them. When we talk about publication bias, we have in mind that some estimates are not published, that they are hidden. So realistically, insignificant estimates are often not reported. So that's publication bias. Then what happens is, if you just do meta-analysis and you compute the average of the reported estimates, your mean would be too large, it would be biased, it would be exaggerated. So that's not good. And then you can use the selection models, or you can use PEESE, you can use whatever correction technique is there, which is well established.
Conventional p-hacking 00:41:24
Now you can have what I call conventional p-hacking, just because it's easier to handle methodologically. Your precision and your center are given. But you can play around with your model to push your estimates higher, to make them statistically significant. So instead of the hollow circles, you observe the black filled ones. So you move your estimates eastwards. What happens is, again, you have a bias, because now you can only see the black and the blue ones. So the mean is biased upwards. But you can still do PEESE. PEESE is fine, because PEESE will run a quadratic regression here, and it will focus on the top of the funnel plot. And you are okay, because the most precise estimates are not affected by this type of p-hacking.
So the conventional p-hacking, which I think Tom Stanley has in mind, is okay for the funnel plot, for PEESE, for the endogenous kink, probably. It's not okay for selection models. Selection models still do break down. That's one of the reasons I focus more on funnel plot techniques than on selection models.
Spurious precision 00:43:08
What I was talking about in the last couple of minutes is the more problematic p-hacking in which the precision can also be p-hacked. So in this extreme case, there is only p-hacking on precision, not estimates. But of course, in practice, you can have both. So that's the point which I'm making again.
Key meta assumption broken 00:44:06
What is really going on in meta-regression when you have p-hacking? When you have p-hacking, especially the p-hacking which can also work on precision, you have a problem with your main assumption in regression, which is the same assumption you always make when you run a regression. You have to assume that the variable you have there as your main independent variable is not correlated with the residual.
And we will come later to the question of why we don't just put more controls here to solve the problem. So in the simple univariate case, we will always have this potential problem that there could be something hidden here in the residuals which can at the same time affect your precision and also your point estimate.
And that's one reason why Bob says, correctly, that we should not rely on univariate PEESE meta-regression. Tom Stanley would also say that we should not rely on univariate meta-regression, we should go beyond, we should add controls. My point is that it's not always possible to specify the model fully and correctly and to add all potential omitted variables. I will talk more about this in just a few minutes. But for now, let's still focus on the univariate meta-regression.
Options for the meta-analyst 00:45:53
You have this classical identification problem in regression. So what can you do? The simplest solution to a violation of the orthogonality condition, the exogeneity condition, is about 100 years old: instrumental variables.
00:46:19 Participant: [A participant comments on the correlation that p-hacking creates, basing this on the figure where the points are moved.]
00:46:42 Tomas Havranek: Yes. But even if you had this, you could still in principle use PEESE. So the main problem is that you have something hidden in the residual of the univariate meta-regression, which influences both precision and estimates at the same time.
00:47:15 Participant: [A participant comments on the difference between publication bias and p-hacking and suggests that the key difference is endogeneity.]
00:47:33 Tomas Havranek: Yes. Instrumental variables are the most basic quasi-experimental approach we can imagine. If we can figure out a meta-analysis technique based on, for example, regression discontinuity, that would be much cooler.
But maybe you will find a way to tackle this endogeneity problem using a different quasi-experimental approach, difference-in-differences or RDD. I didn't find a way to do it. We focus on instrumental variable models.
Now, what do we do exactly? Before we move to the estimator, what are the other options you have? You collect a bunch of studies. People can tell you that you should get rid of the bad studies. The problem is, of course, that you don't know which ones are bad. But sometimes the editor, for example, will tell you: "You don't know which ones are bad, but certainly you know which ones are quasi-experimental. So I want you to just focus on the quasi-experimental studies."
The problem is that even if you have an older study, it can sometimes be done well. So you have situations in which even with OLS you can approach a quasi-experimental setup. By the way, if you have experimental data and you run OLS, you are perfectly fine. So sometimes it is tricky to say what is quasi-experimental and what is not. Or even if you have experimental data, like the experiment from Tennessee I mentioned: is this a real experiment? So it is not a zero-or-one decision, what is a good study and what is a bad study.
You can include controls for OLS, RDD and so on. Again, sometimes it is not clear. A study uses difference-in-differences, but it completely messes up the assumptions. Or it uses RDD, but not that well. So should I put one for RDD or not? It is very hard, in a multivariate meta-regression model, to properly control for all these ways in which a study could maybe have been done better.
As I write here, unfortunately, we don't know what the true model is. If you run a regression of salary on beauty, we don't know what the correct control variables are that you should include in your regression. You can come up with a big number of potential controls, but you are never completely sure. That is our point. If you are lucky, you include controls and you will be fine. In general, that doesn't hold. It doesn't have to hold all the time.
00:52:01 Participant: [A participant asks which example he means.]
00:52:10 Tomas Havranek: There could be many other things that are happening. On Friday, I had a seminar on this. The point of our paper on beauty is that intelligence could be related to beauty. Does that sound strange? There is a big biological literature on this.
00:52:35 Participant: [A participant suggests that beauty and intelligence may be correlated because people tend to marry within their own socioeconomic class, for example within the same elite universities, while noting that beautiful people are not always intelligent.]
00:53:14 Tomas Havranek: True. I think we are not in disagreement. But the point is, if you run a regression model, there are often controls you should include in your regression. Beauty was just an example. Maybe it is not the best example. You can have a look at the paper if you want. But maybe the example with education was a better one. If you regress your earnings on your education, you should also control for things like your family background, your intelligence and so on. And it is never completely clear whether you control for all of these things.
00:54:07 Participant: [A participant repeats the question about the beauty and salary example.]
00:54:20 Tomas Havranek: I mentioned ability. Then you can have things like cognitive ability, and then social ability, like your self-confidence.
Meta-analysis instrumental variable estimator 00:58:24
When we try to solve this endogeneity problem, we try to find an instrument that would be related to precision somehow, but not to the problem in the residual in the PET-PEESE regression. So it would not be related to p-hacking. In the paper we use sample size as such an instrument. Sample size by definition is related to precision. If you have a larger sample, you have more precision, and vice versa.
It is much more difficult to somehow manipulate your sample size to get statistical significance than to change your estimation technique or your control variables to get more statistical significance. In more technical words, the instrument should be strong. There should be a strong correlation between precision and sample size, just by the definition of the standard error. And we assume there is not much of a relation between sample size and p-hacking. There could be some. You could p-hack by changing the subset you are drawing from your data. But it is hard to artificially increase your sample size just to get statistical significance.
The reason is that typically you would use all the data you have. There is no reason you would just start with a small sample. One exception: in some fields of experimental research, when you plan your experiment, you can also plan how many people you will have in the experiment, how large a sample size you will have. And sometimes you do these power computations, in which you plan to have a bigger sample size, a bigger number of people in your experiment, if you expect the effect in your experiment to be small.
Sometimes you have some idea. If I am doing an experiment on class size and student performance, the effect is probably not so large, so I want to have a big number of students to be able to really isolate and identify the effect. Sometimes you know that the coefficient is going to be large, so you need just a normal size, because experiments are costly and increasing the sample size in experiments is going to be expensive for you. So this is the intuition for the instrumental variable estimator, which we call MAIVE, the meta-analysis instrumental variable estimator.
Since, as you know, IV is a two-stage regression approach, you first need to put your instrument on the right-hand side and regress the endogenous variable on the instrument. In this case, you regress variance on inverse sample size. And it gets rid of the p-hacking and misspecification things, assuming they are not related to the sample size. So you will use the fitted values from the first-stage regression instead of the reported ones.
In MAIVE, when you use PEESE with the MAIVE adjustment, you have a standard instrumental variable adjustment, but you can also use these adjusted variance measures in other estimators. That is also possible. And of course, when you estimate PEESE with the adjusted variance, you can also include other control variables for whatever you think should be included in a multivariate method. In the paper, we have the same simulations as before, and we just compare PET-PEESE with the instrumented version. You can see that it substantially helps to improve the performance of these estimators in the worst cases.
Extended MAIVE 01:04:01
That is essentially the main idea of the estimator. We are also working on an extended version, in which we put it in a different way. What I showed you is a normal IV estimator.
You have p-hacked studies and you have studies which are not p-hacked. For the p-hacked studies, we assume that the movement in the funnel plot is in this direction. So when you have more p-hacking, you are likely to have bigger effects, which are also more precise.
But we only adjust standard errors. That is the only thing that MAIVE does. We only correct the spurious precision, because we don't know which estimates are p-hacked and which are not p-hacked.
Maybe the first stage of the instrumental variable approach also tells you something about how biased the point estimates potentially are. Because if you have too much precision relative to sample size, your residual here, your pi, would be negative. If your pi is negative, it means the standard error that you report is too small in relation to the sample size that you have. So it doesn't have to mean that there is p-hacking in this estimate, but it could be more likely.
What we experiment with is this: we look at this first-stage equation from the MAIVE estimator, and we put less weight on the estimates which have a negative pi, which means estimates which are more likely, perhaps, to have a p-hacking problem, which show too small standard errors compared to the sample size.
So we put 0.9 weight, just a slight adjustment, on these estimates where we have these negative residuals, which are too precise somehow. And again, we do the simulations, and it turns out that the extended version of MAIVE, when we try to actively diminish the weight of spuriously precise estimates even further, has a smaller bias with these large values, large degrees of p-hacking. So hopefully it is going to be a follow-up paper. It will not always work, because you do not always have spurious precision and at the same time a spuriously large estimate. Sometimes it is going to be just one of these issues. But still, I think it is worth exploring.
MAIVE reduces PET-PEESE in 70% of the cases 01:08:19
These simulations were meant to add a bit more credibility to our results. We have theoretical motivation, then we have simulations, and we also try to see what happens in real meta-analysis data, when we compare MAIVE to PET-PEESE.
We use the data set of many meta-analyses collected by Chris Doucouliagos. And if you use MAIVE, not always, but very often, you have smaller estimates, which would be consistent with p-hacking on average, pushing the estimates northeast in the funnel plot, causing a positive bias in your estimates, and also causing a positive bias in meta-analysis models which correct for publication bias like this.
Bonus: Maya Mathur's p-hacking correction 01:09:20
That is actually it. And again, as a reminder, there is another approach by Maya Mathur, which is completely different. It is a different approach to p-hacking, in which you simply focus on the insignificant estimates, for example, and you ignore the significant ones, because you assume these are potentially p-hacked. So that is really completely different from what we do.
If you are concerned about p-hacking, you have two completely separate approaches. One, our approach, is based on the funnel plot and instrumental variables. The other approach is somewhat analogous to selection models, but in a way the opposite, in that you only look at the insignificant estimates. It is also Bayesian, but the Bayesian technique here is just used for computational feasibility.
That is all.
Questions and discussion 01:11:58
Do you have any questions on the lecture, on the substance? […]
01:14:08 Participant: [A participant says that publication bias and p-hacking might be distinguished in simulations by whether the estimates are correlated with their standard errors, and that a p-hacking simulation program could be used to check this. The participant says they will try it out and show the result afterwards.]
01:16:02 Tomas Havranek: Thanks a lot.
There is a lot of work to be done on this. I think there is a lot of potential for you guys, if you want to specialize in the field later. This is really where a lot of contribution can still be made.
Before, the message was more that, with p-hacking, you really can't use the publication bias correction tools anymore. So there is still a lot of potential, and this is just the beginning.
01:17:08 Participant: I think people like myself appreciate that the things you do to fix publication bias don't necessarily work for p-hacking. In fact, they can be counterproductive. That's a great summary.
01:17:24 Tomas Havranek: Thank you, guys. Next time we're going to talk about heterogeneity.
Corrections
Slips of the tongue corrected in the text:
- [00:24:51] said "a bigger effect"; the text has "a smaller effect".
- [00:26:14] said "instrumental variables"; the text has "quasi-experiments".
- [00:30:27] said "0.9 is an extreme value"; the text has "0.8 is an extreme value".
- [00:40:10] said "selection technique"; the text has "correction technique".