Transcript. The recorded lectures of Meta-Analysis in Economics: Practical Tools for Synthesizing Research, given by Tomas Havranek at the Institute of Social and Economic Research (ISER), the University of Osaka, on 16 March 2026. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. A few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the hosts where Tomas Havranek refers to them. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the recordings (recording 1 is the morning session, recording 2 the afternoon); the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.

Morning session (recording 1) 00:00:00

00:00:00 Tomas Havranek: First of all, as some of you know, Chishio, Taisuke, Martina, I love to go back to Japan, and I try to come as often as I can. I am sorry to say that I am unable to give the workshop in Japanese, because my Japanese is essentially limited to one sentence, something like nama biru o hitotsu onegaishimasu. That is all I can say. […] By the way, if you want to see a great example of how meta-analyses are done in economics, you should probably not look at my papers but at Taisuke's.

There are many reasons, which I have also explained, but one of the things I really like is the typesetting and the presentation of figures and tables. It is the best I have seen, not just in meta-analysis but in any empirical work in economics. It is really great. And I am happy to be here at the Institute, the place where the International Economic Review is edited. I know there are a couple of great places for economics in Japan, Tokyo University, Kyoto, Osaka, but for me the Institute really is one of the flagship places for economics research in Japan, especially for behavioral and experimental economics, and maybe one of the top places anywhere.

I am really honored to be here. We have plenty of time today, and my plan is to adjust the pace to what you would like to hear about. So please, if you want me to spend more or less time on something, or if something is not clear or I can make it clearer, just let me know. I know Japanese people are really polite, but in economics we are not so polite in general, and this is an economics workshop. So feel free to interrupt me at any point. I am used to it.

I have given lectures in Chicago, where sometimes I just say a couple of sentences and the rest of the seminar is people arguing with each other. We will not do that here, but I want to emphasize that it is completely OK to jump in at any point and ask me questions. If something is not clear, it is my fault. […] You are my customers, so let me know if something is not clear.

My plan is to show you how to do a meta-analysis from scratch in economics. I will also talk a little about the differences between economics and psychology and other fields, but mostly it will be about economics. I will not talk about some basic meta-analysis tools that are not used much in economics. I will just describe them briefly. We can skip some things. We will see how it goes and how tired we are, and we will adjust to your needs.

[…] Essentially, we have four blocks, but I have seven chapters, as you can see. First, let's talk about what meta-analysis is and why we need it in the first place, and why we do not just look at one best identified study, and so on. Second, a brief digression: even if you do not want to do any meta-analysis, it might be useful for you to understand how research synthesis and meta-analysis are done, to be able to judge for yourself and not just rely on AI. Is this meta-analysis good? What should I do in terms of my research or in terms of policy, and what should I rely on? I will give examples from policy, because I spent most of my professional life at a central bank, where I would use meta-analysis essentially every week for policy recommendations and not just for my own research.

We will talk about how to search for studies to be included in a meta-analysis, how to extract data points from these studies, and, very briefly, some conventional meta-analysis tools. But mostly I want to talk about the methods we use in economics as our baseline. Some of these were developed by people like Chishio, who developed a great correction for publication bias and for p-hacking. And of course we will also talk about heterogeneity, by which I mean this: why do different studies produce different estimates? That is one of the key issues we want to get into.

Motivation 00:05:52

Let's start with some introduction. Why do a meta-analysis? I will give you a research perspective and a policy perspective. In terms of policy, again, I worked for the central bank, the Czech National Bank, for many years, and I was an adviser to the board, which means I would be in touch with the governor, mostly the vice governor, and they would ask me any question they had. That was before AI, so now you would probably delegate most of it to AI. But I would read materials, and my job would often be to summarize research, so I would often turn to meta-analysis. Of course, if you have to produce an answer within one week, you do not have time to do your own, but you can do a little bit. I would use it every week, essentially, sometimes every day.

Then I served on the Czech National Economic Council. We would do taxation and fiscal policy. Again, meta-analysis would be important: how to set taxes, how to calibrate responses to fiscal policy, and so on. […]

But of course, even if you do not care about policy at all and you just do research, for example in macroeconomics, you need to calibrate your models. You need to take some parameters that guide risk preferences. For example, you can look at Taisuke's meta-analysis for calibration. You need to look somewhere. And you may really dislike meta-analysis. That is probably not your case, because you would not be here, but maybe a little bit. But some people do dislike it. In economics especially, you still find many people who say, "I do not want to do any kind of average over published papers. I just want to select the best, most powerful study and look at that one."

But that effectively means you do a meta-analysis in which you assign weight zero to all the other studies. It is a special case of a meta-analysis. So why not start with the entire universe? I will show you examples where we essentially arrive at something similar: we put a lot of weight on just the best studies and essentially ignore the rest. But why would you say ex ante, "I will not even look at them, because I know there are just one or two papers in the AER which are the best, and I should only look at them"? There are many problems with that. We will talk about it. One of them is publication bias. If you look at just one published study, you never know: was it just a lucky result, or what is going on? Even if it looks impeccable on paper in terms of identification, in terms of sample size, everything.

00:09:26 Participant: I would like to ask you a question about the historical perspective on meta-analysis and how special, or not, we are in econ. It seems to me that we picked this up a bit late relative to other social sciences and to medicine and public health, maybe, but I am not sure. I have a very vague understanding of this. Do you have an intuition for why that is the case, if it is true?

00:09:53 Tomas Havranek: In economics, I think, compared to other disciplines, we really value contribution, methodological contribution and innovation. If you want to publish a paper high, it often needs to have something new in terms of methodology. That is a big difference from, for example, psychology, where a well-executed meta-analysis without any methodological value added can be published in the very top journals, such as Nature Human Behaviour, or even Science or Nature. This focus on contribution, which is really very competitive in economics, leads some people, I think, to say, "This is just an average. We want to see something new." And there is something to it, by the way.

I am an economist, so I have to defend the profession as well. But I think we need to publish much more of this meta-research, and I am not talking just about my papers or Taisuke's papers, but in general. On publication bias and the p-hacking issue, if you do not have meta-analysis, you cannot really say anything about it. We will talk about how important it is. That was a great question. Thank you.

Here is one brief example from fiscal policy. You do not have to understand the model. If you want to see details, we have a paper that talks about it in more detail, but this is just the general intuition. Let's say you work for the Ministry of Finance and you are an analyst. The minister asks you, "We want to increase spending by this percentage point. Tell me what will happen to real wages, for example."

Most likely you will have some sort of structural model. You can also have a VAR, but let's say you have a structural model, which is a typical tool to evaluate policy changes. The answer you give will depend on how this model is calibrated. For fiscal policy, the key parameter is the intertemporal elasticity of labor supply. Here it just shows a standard model, like a New Keynesian model, in which you change one parameter.

You change how you calibrate the labor supply elasticity, and you will have completely different results, sometimes, not always, but for the real wage you will. So your answer to the minister, your boss, will depend heavily on how you calibrate it. And how you calibrate will depend on the studies, because I will show you how the results differ for different studies which estimate the labor supply elasticity. So you need some sort of literature review, synthesis, or meta-analysis, however you want to call it. Now let's turn to a central bank. Let's say the governor asks you, "We want to increase interest rates because we fear inflation, da-da-da. What will happen to investment?" Again, you can run your structural model, DSGE or whatever, and here it will again depend on many parameters, but the key parameter is the elasticity of intertemporal substitution in consumption.

That is essentially how people adjust consumption in response to changes in interest rates. Again, with different calibrations, you will have completely different answers, so this is important. And then the governor can ask you, "Why do you calibrate the elasticity at 0.5?" and you say, "Because I read this paper by Taisuke or Martina." This shows you how important even a change in one parameter can be. But of course, in practice, you have many parameters.

If you work in a central bank and you run these sophisticated DSGE models, you have dozens and dozens of parameters you have to calibrate. So the model uncertainty is huge. Even when you have a great meta-analysis, you will often be in trouble. But however you do it, you will need literature surveys for each and every one of these parameters. So again, meta-analysis is crucial. We should do it much more in economics. For many of these parameters, we do not even have a good survey. Now, back to the fiscal policy problem, labor supply elasticity. Let's suppose I want to calibrate my model in the Ministry of Finance, the Treasury.

What do you call it in Japan? Is it the Treasury or the Ministry of Finance? It is the Ministry of Finance in my country too, but in the US and in Australia they call it the Treasury. So you want to look at individual empirical studies that estimate labor supply elasticity. Now, what do you see? They agree that the elasticity is positive, which is more or less what would be intuitive. But in terms of the precise value, you can justify any value from zero to maybe two if you pick a specific study. So if you pick one study, you can do anything. That is not really useful. You have to somehow summarize it.

00:15:33 Participant: Is this the compensated elasticity?

00:15:36 Tomas Havranek: It is a Frisch elasticity. Again, we have the details in the paper. It is just this one paper in the Review of Economic Dynamics, if you are interested. But I think for the general intuition you do not need to worry about compensated or not, Hicks or Frisch.

What is commonly done, for example, by the Congressional Budget Office (CBO) in the United States is to take the average of the reported estimates, which seems like a good idea, or the median value of what is published in solid journals. In this case, it is often calibrated at 0.5. […] But the problem here, and we will talk about it in much more detail, is that the distribution of estimates is not really what you would expect.

You have these big jumps, especially at zero. The intuition is very strong that labor supply elasticity should be positive; otherwise it does not really make sense. But since it is estimated in a normal regression, you should not expect such a big jump. Estimates that are just positive would be much more likely to be reported than estimates that are just negative. So something is going on there, some selection. Again, we will talk about it much more. That is the reason why, if you just take the mean or the median, it is probably not going to be a good representation of the underlying research results. So that is one thing: publication bias or p-hacking.

Another is heterogeneity. If you estimate, for example, the labor supply elasticity and compare men and women, you will most likely get a different elasticity. You will also most likely get one if you compare young people, such as Taisuke, with people near retirement, such as me, or someone who is 65.

00:18:14 Participant: I think you are younger than me.

00:18:16 Tomas Havranek: I wanted you to say that, that is why. No, but you certainly do look younger. There are many different contexts in which you can estimate conceptually the same thing, but you have good reasons to expect that it will differ across people in a systematic way.

The difference between men and women is clear. For me, if you slash my salary in half, does it affect the way I work? No, because I like what I do. […] if you do labor economics, you will know that women tend to have a much higher labor supply elasticity.

Context is very important, even for the calibration of these models. What I will show you is the bottom line of the workshop and of any meta-analysis. You take the literature and try to correct it for different biases: publication bias, p-hacking, identification problems. You put more weight on better identified studies. Then you construct the best guess, with confidence intervals for different contexts. Here I show examples for the labor supply elasticity for men, for women and for older workers. If you are close to retirement, you are more likely to be responsive to the price of your free time and to how much you are paid.

That was a general outline of why we should do a meta-analysis. If I were to recommend one short paper, it would be the guidelines we put together with Zuzana, Chris and Tom Stanley. It is a very short, non-technical document of just a couple of pages, but it is really broad. You have it in the documents that I shared with Taisuke. You also have the data and code, which I will briefly go through, but do not feel forced to do it. If you just want to sit and listen, that is completely okay. There is going to be no test at the end, I guess. Whatever suits you. In these guidelines, we go through essentially all of the things I will talk about.

It is quite new: it was published online in November 2023. [Note, 2026: the guide is Irsova, Doucouliagos, Havranek and Stanley (2024), Journal of Economic Surveys 38(5): 1547-1566, free at meta-analysis.cz/guidelines/. Since 2026 the site pairs it with three principles: correct for publication bias (RoBMA), correct for p-hacking (MAIVE and RTMA), and cluster by study (CR2 standard errors).] I will tell you if there is something new that is not covered in the guidelines, especially AI. I will have quite a few slides about how to use AI at different stages of meta-analysis, and also for those who do not want to use AI. But this is not in the paper. It was done at a time when we were still learning how to do it. We are still learning, but at that time AI was not really useful for this kind of work. Only in the last year, when these new tools were developed, can we save quite a lot of time in a safe, productive way using AI in meta-analysis. As a guiding example in this workshop, I will use a paper on whether you earn more if you are more beautiful. That is a weird question, but a hypothetical one.

Does it pay off for me to go for plastic surgery to make myself more beautiful? Will I make more money? Will people pay me more? That would be a practical question. So what is the causal effect of beauty on salary? It is quite easy to understand. […]

[…] We will work on this data set on how beauty affects earnings. It is a difficult question because you cannot really run experiments. You could, but it would be unethical to randomly disfigure some people. So there are clever ways to look at it, but it is mostly observational regressions, which also makes it easy for meta-analysis. We will talk about it more.

Let's very briefly look at the data and code that we have. [Hands-on: opens the Stata do file in the code folder.] The code is fully annotated, so even if you do not want to open it now, you can get back to it later, and you should be able to understand it, hopefully all of it. [Hands-on: loads the beauty data set into Stata.] Here you can see a couple of examples of estimates. What does it mean when we have a beauty premium of nine?

It means that if people move from average looks by one standard deviation to better looks, which means really pretty or even beautiful, you tend to earn 9% more. That is the estimate in this paper. It is not a small effect, but a big one. We have the estimate, the standard error, the sample size, and much more information in the data set.

Now let's also open R and do the same thing. We will upload our data in R. [Hands-on: loads the data file in R.] It works and shows you the basic structure of the data. We have a lot of variables, which we will discuss later. One is how beauty is measured. Do people just look at pictures? Or do they do personal interviews? Or do they let AI, some algorithms, look at facial symmetry? By the way, beauty can be defined in many ways, in terms of physical attributes.

We look at facial features, just the face. You could also look at height. I should be careful not to say something stupid here, but for men there is of course a premium for height. A taller man tends to make more money. It is well established. I read on Greg Mankiw's website that we should tax height. Not income, but height. Because there is no distortion here. And there is a correlation between height and income. […] Here is the structure of the data. It is not height, by the way.

What we find in the paper (and I will show you essentially how we do the paper from scratch) is that, on average, before you do any corrections, if you just take the average of the estimates, a change of one standard deviation from plain to beautiful makes you earn 5% more, which is not really huge. It is not zero, so economically it is an effect, but it is not huge. So do not worry too much about it. That is the first thing. Now, if you correct it for these biases, you will get very close to zero. So it is actually a non-issue. What is interesting is that our results imply there could be a correlation between beauty and ability. […]

Again, very weak. I want to emphasize how small even the 5% is. Even before any correction for publication bias, which tends to bring it much smaller still, 5% is very small. That is because one standard deviation would be a major plastic surgery for me, for any of us, or at least for most of us, I guess. So it is probably not really worth it. Even though, in purely financial terms (I do not know what the average salary in Japan is), if you undergo the surgery at age 30 and count the additional income for the next 20 years or something, the 5% might actually tell you, in terms of cost and benefit, that you should go for the surgery. But I would correct the estimates, and then we will definitely know: just forget about it. Let's just do something else.

You have the paper in your folder, but we also have this website, meta-analysis.cz/beauty. And as I said, we are working, mostly Martina, on a big revision of this paper. […]

00:33:48 Participant: Did you guys look at the interaction with age? I'm guessing there is heterogeneity across the life cycle in the beauty premium.

00:33:54 Tomas Havranek: […] We do look at age, but we will look at it in more detail in the next revision. We have age coded somehow. The problem is that in meta-analysis you can only do as much as was already done in primary studies, and not too many primary studies look at age. A couple of them do. But then the question is: if you have three studies which look into it, do you have enough power in a meta-analysis to say something new about this? That is one of the issues we face. Another good question.

00:34:39 Participant: I was thinking about the correlation with ability. If you think about ability as being endogenous and reflecting human capital investments, you could think that early on, if you're beautiful and you attract attention, then, I don't know, teachers spend more time with you. That could actually also be important. I was trying to think a little bit about this. […]

00:35:04 Tomas Havranek: […] That is a good point. Yes, that might happen. It is one of the explanations we now have in the paper. There are a couple of papers which we look at, and they suggest, exactly as you say, that more beautiful children tend to get more attention even in kindergarten. It might then translate into human capital, even cognitive skills. We distinguish in the paper between cognitive things like IQ, ability, and non-cognitive skills, which, I would suppose, are probably much more related to beauty, like self-confidence. We will talk about it much more. […]

00:36:04 Participant: Sorry, I have another question. I was looking at point number one and at the number of estimates you have overall. Given that you have 67 studies, do you have huge heterogeneity in the number of estimates available per study? Are there specific studies that drive the...

00:36:20 Tomas Havranek: Yes, we do. There are a couple of studies that give you only one estimate, and then you might have a study that gives you 100 estimates. That is common in economics because sometimes you do many subsamples. The question is then what you should collect: should you collect all of them? I think it is better to collect everything, because you can always show what happens if you ignore it. If you think that one study should be one estimate, I will also show you estimations in which we do so. But it is safer to collect all of it, and then you can always decide: I only trust this one study, this one estimate from this one study. That is an extreme case, but you can do it.

But my recommendation would be not to do it ex ante. Now, literature search. Before you search for your studies, you need to decide on the topic. Most of you, if you want to do a meta-analysis, will already have a topic. You will want to do it for specific reasons, such as parameters of, let's say, risk preferences. If you want some inspiration on what kind of topics could be good for meta-analysis, you can see examples of very different topics from different fields on the meta-analysis.cz website, where I have all the papers I have ever done on meta-analysis, more than 50 papers. It is mostly economics, but there are topics even a little on the boundary between economics and psychology, and on climate science.

Here are a couple of examples where you have a good motivation for a meta-analysis. Of course, in almost any case you need one, but it is nice when you can sell it to the audience immediately in the introduction. One case: these graphs you will only get after you do a meta-analysis. But even before, if you know the field, you might have some idea about the main things, the way it looks.

For the beauty effect, what we see in the literature is increasing disagreement: the newer studies tend to disagree more than the older studies. This is apparent even before you do the meta-analysis. You just know the literature, and you know there is increasing disagreement. That is a great motivation. Why do you need a survey? Another is hedge fund alpha: how much do hedge funds perform better compared to the general stock market, the S&P 500? You have some vague knowledge from just looking at the literature that there seems to be a downward trend in the reported alphas, and you might want to explain why. That is another good topic for a meta-analysis. Or you might have a parameter like the elasticity of substitution, which seems to differ a lot across countries. Again, you might want to write a paper and explain why, which is what we did for the Journal of International Economics. […] Or you might have a vague suspicion that there might be publication bias. I will explain this graph in detail later, so don't worry about it now. That is another good motivation for a meta-analysis. Or, my favorite: different techniques in the studies give you different results. This is the capital-labor substitution elasticity.

It really matters a lot, when you estimate the substitution between capital and labor, whether you focus on the first-order condition of labor or of capital. It is striking. This is the histogram of published estimates. It is well known that it is great to be able to show it in such a way using a meta-analysis. I think there is a paper in REStat that is completely based on this result of ours, on this picture, and tries to explain it using a model. This might be a good motivation for why you would want something like that.

Now you have your topic, and you want to search for studies to be included in your meta-analysis. What I will talk about here is the way we do it in economics. Then there is the psychology way, which I will only briefly mention. Our economics process is a bit non-standard for meta-science in general.

Outside of economics, people would almost never use Google Scholar. They would hire a librarian and go through all of these databases, but not with full-text search, because Google Scholar is one of the few databases that give you full-text search. They would look at the keywords, the abstract and the title. That is very laborious, and it is what we are now doing for this paper: again using a librarian and these many databases, with no full-text search.

I think it is much better to use one database that gives you full-text search and is universal, with full coverage. But this approach has its own problems. The main problem is that you cannot replicate a Google Scholar search. It keeps changing. That is a real downside. Also, by the definition of how it works, it gives more weight in the results to more cited studies, although not always. So in principle you might get some citation bias. These are two downsides. But the big advantage is full coverage and full-text search. You are unlikely to miss many studies compared to a search with rigid keywords, where you search just by the keywords or just by the abstract. That is my take.

00:43:19 Participant: I also have the impression that Google Scholar has some priority keywords. If I add an extra keyword by combining with AND, it should reduce the number of hits, but sometimes it increases the number. I wonder how Google Scholar reads a combination of keywords.

00:43:38 Tomas Havranek: Let's wait a little bit, and I will tell you more about that. There is no perfect solution. My personal preference is that if you want to publish in economics, it is perfectly fine. I have published 50 papers using Google Scholar. But if you submit it elsewhere, you might get complaints.

For now, let's forget about AI for a while. Then we will also talk about AI. Let's say you don't want to use AI, which is completely fine. The way we do it is in the guidelines, but it is not set in stone. It is just one way to do it, and it works. What I like to do is find the most relevant studies that I know I need to have in the paper. I know the topic, so I know five studies that will be there.

Then I play a lot with the search in Google Scholar, with which keywords and which operators to include and so on, until I have a search that gives me these famous studies near the top, because I know they should be there. They are prominent. If my search doesn't show any of them, it is probably not the right one. Usually I get about one million hits, and I don't have time to go through one million papers without AI, and even with AI, no. So I go through the first 500 hits. It is arbitrary, but it is how we do it. I will show you why the precise number is probably not an issue. I read the abstracts of these hits.

For most of them, you will see that you cannot use them: it is not an empirical paper, it is something completely different. So you can see which hits you can delete immediately. But you will be left with maybe 100 to 200 hits where you are not sure. They could contain some empirical estimates. So you download these papers and then you skim or even read them, just to see if they have empirical estimates. Then you can do the same thing, just to be sure you haven't missed any really new studies, which are not so much cited and will normally not be in the top of Google Scholar. You can just restrict your search to the last couple of years.

On the website meta-analysis.cz, for different papers, we have different examples of the specific queries and how we do it, so you can take a look. I will not go through the examples, but you can see in general what we do.

I don't think you need to use the AND operator in Google Scholar. If you just add another one, it is the same thing. Anyway, you will end up with a list of papers. You will create a basic table like that, where you have the key information about the studies. What is really important now is to do something called snowballing. Snowballing makes your search robust to the initial definition of the search, to how many abstracts you screen and to how it works.

You try a completely different procedure to search for studies you might have missed. You take the studies you have identified in the previous step, just from Google Scholar, and then you can use a database like Scopus or Web of Science, or whatever, and download the reference list for each study you have found. You can do it easily; now you can do it with AI really quickly. Then you can construct a list of studies that are frequently cited by the papers you have already identified. You have to check them. And sometimes, well, every time, you will see that some papers are frequently cited by your studies, but you don't have them in your list.

Some of them might be empirical papers that actually report estimates you can use. That is a very important check, because it is not much dependent on the way you define your keywords and your search in the first step. It is complementary. If you do these two steps, you are unlikely to miss any important study. You will still probably miss some studies; it is hard to be totally sure you haven't missed anything. But with these two steps, you can be reasonably sure. Maybe once or twice there were some complaints that we don't cover all of the studies, so we had to expand.

00:49:59 Participant: When you have a previous meta-analysis or review papers, is there any change to this procedure?

00:50:12 Tomas Havranek: That is a good question. Essentially, I was talking here about a case in which you do the first meta-analysis on a given topic. But you are right: in many cases you will already have previous meta-analyses, which have their own sets of studies. Of course, you would like to include them as well. By definition, you probably start with that, but do it as a combination: take these previous studies and then do this. Nowadays, probably many of us will want to use AI in some way.

It is changing rapidly every week, and it is very hard for me to keep up. Maybe you are doing much better than I do, but I always have to learn new things every month. This is written for ChatGPT, but instead of ChatGPT you can use Claude or whatever. My main point here, which I would like to emphasize and will say a couple of times, is this: do not rely on one AI model.

You should have at least two subscriptions, such as ChatGPT and Claude. If you use them to help you collect the studies or collect data, you should use both, because the hallucinations and mistakes are not fully correlated within models, and even less so across models. If you ask the same model repeatedly, double check, triple check, it will get rid of most of the problems. And if you complement it with a completely different model, you will get something which is much less likely to be wrong than if you just rely on one model.

[An example of an AI tool inventing a plausible data set, down to a series number in a database, that does not exist.] […]

[…] That is the point even here: it is very useful to use more models.

00:53:56 Participant: I had three questions. The first one was, when you mentioned, much earlier, age being correlated with the outcome: is it actually a good practice, when I know that this particular variable is of interest but is not reported in a lot of primary studies, to contact the authors to see if they are okay to share the complete summary statistics?

00:54:24 Tomas Havranek: It is a very good idea, but in practice it is often not feasible. When I was a young PhD student, I did the same thing. When I was doing my first meta-analysis, later published in JIE, I would write to each and every author and ask them for additional data. It took me a couple of years to do my paper. Now I don't do it. I have this project, and I want to move on to something else. Still, it is probably a good thing to do.

But no one expects you to do it. It is not something that you would have to do, and in many ways it is not very feasible. People will not always reply, and you will get only a bit more data. In this case, what people usually do is look at the subset of studies that do report it. For age, because it is important, we might want to do something along the lines you recommend, but in general I don't have a clear answer for you. So, as you can see...

00:55:39 Participant: The second question is on the usage of Google Scholar. You mentioned that Google Scholar does create a citation bias: it pushes the studies that have been heavily cited. But when we do snowballing, there is a good chance that those studies will only keep showing up. I don't know if I can explain it properly.

00:56:05 Tomas Havranek: Snowballing even makes the citation bias worse, if there is one. And the question is: is there a citation bias? In general, more highly cited studies will definitely be the more prominent ones, the ones that people know about. So if you miss a highly cited study, it is definitely more of an issue for you than if you miss a study that was published 10 years ago and has one citation.

[…] But I agree that the principal problem is there, and the solution is difficult: to do something completely different, with a librarian and seven different databases, takes a couple of months of your life.

00:57:10 Participant: The third one is related. It's not really a question, but I think you can now modify Claude by making it generate a particular skill. It would keep running its own responses, and you can ask it to create a skill to check if its replies actually make sense. So you wouldn't essentially need two or three models to do the checking. I have not tried this with any particular meta-analysis data set, but in general it is now possible to make it check its own responses over and over, until it gives you the absolutely correct answer.

00:57:55 Tomas Havranek: That is a great comment. You should definitely do that. Within a model, you should double check and triple check. But my point is that this still does not necessarily correct for all of these problems. All of these models will overlap to a large extent. They are trained on data sets that are not the same, and the way they are trained, the methodology, is also different, but not completely different. Still, you will have additional information from these different models. That is what I mean by within and across model correlations of these, let's say, mistakes or hallucinations. So I think you should do both. Why not? It is very simple and does not cost you much. For example, I like to use Grok as an additional model. It is not my primary model, but it is very different from the other ones. We like meta-analysis, so we should apply the meta-analysis approach to AI as well: not just one study, not just one model, but more. But I'm not an AI expert.

[…] Here is one example of how you can use AI, and maybe it will come up with much better ways to do it. Of course, you should start with deep research anyway, just to help you: you know your topic, but you might still have missed a couple of nuances. You will want to define your search query in Google Scholar or in any database, but we use Google Scholar here.

For the search, it might be very useful to use a thinking model to make the search as powerful as possible. I will show you some examples. Then download the abstracts, maybe using an AI agent: do not do it manually, but ask an agent to do it. That is a good assignment. You can also use an agent to download the PDFs, so you do not have to do it yourself. Then you can upload them, not all of them, but in batches or maybe one by one, and ask one or two AI models what the chance is that the paper actually has an empirical estimate. This is to help you decide: should I drop it, or should I read it myself? You can also check out ASReview, which uses AI. I haven't really tried it much.

That is the general outline; now in more detail. You can read these examples of simple prompts, which we did not use for our paper but which might be useful in general for such a project. Which model to use is a different question. Again, I would definitely use more than one. I've been using all four main models: ChatGPT, Claude, Gemini and Grok. My impression right now is that Claude with thinking and ChatGPT with extended thinking are the most powerful models. Gemini, in my experience in research, usually makes a lot of mistakes, and so does Grok, but it is nice to use them as another source of potential information.

Here are other examples of prompts, which I will not read out, to help you decide whether it makes sense to download and read these papers. Otherwise it takes a really long time to scan all of these abstracts and to read all of these papers. AI can be used, as you know, for many questions with a different degree of efficiency. But it is a language model, and for these questions I think it is well suited to estimate the likelihood: if you read the abstract, if you read the paper, what is the likelihood that it has empirical estimates? Then you can ask it to generate simple summary statistics or information.

You do the same with the actual PDFs, not just abstracts: you upload them. Given the current state of the models, you should not upload many papers at the same time, because the models could get confused. Maybe it is still best to upload the papers one by one. Nowadays you would ask an agent to do it step by step, to look at each paper individually. If you upload 10 different papers in one session, it might get confused and make up results for individual papers.

01:04:14 Participant: Do you request quotes? I am always worried that the LLM might make stuff up, but if you request quotes from the paper, then you can verify them yourself on a few papers. I don't know if this is something you have done.

01:04:31 Tomas Havranek: You mean replicate a paper?

01:04:33 Participant: No, for instance, you ask it: does it contain new empirical estimates? It could say: yes, I am basing myself on page 24 of the paper, as a way of verifying.

01:04:48 Tomas Havranek: I have it on the slides later for data collection, but of course you can also ask already here, at this stage. Again, if you use a different AI model, the chance that you will make a mistake is smaller. Certainly, I see value in that, and you are right.

I already mentioned ASReview, which you might want to check out. It is a way to screen studies in meta-analysis with AI, in which you help to co-train your screening. Some people do use it, but not much in economics. Now, of course, there is Semantic Scholar, and Google Scholar also has this new AI search, as you probably know. I haven't tried it yet for a new meta-analysis, but it is something you can also incorporate.

As of today, you cannot just fully trust these models. Even if you verify within a model and across models, you still need someone to cross-check. But it is great because before AI, you would need more people to do this. Again, as with AI, you cannot rely on one person who will collect the studies and the data set by herself or by himself. You would need more people to cross-check at least portions of the search and the data set. Now, I think you can do with just one person and a couple of AI models. But you still need a human there, even though you might have heard about otto-SR, this new startup, which promises to do a meta-analysis from start to finish, fully automated.

You still need some human. They have a paper in which they compare meta-analyses done by them, fully automated, to human meta-analyses from the Cochrane database, and they show that it is better if you fully automate it. […]

01:07:26 Participant: I used otto-SR. They allow you to use it for one study, but you cannot invite anyone as a reviewer in otto-SR. You also mentioned something about a beta version. Was it ASReview or otto-SR? I checked it once, but there was another one that I found, Nested Knowledge, which had a very similar interface to otto-SR. They also have the same constraint: they also let you do only one project. Then you need to pay. And with almost all the other ones I tried, there is either a paywall at the beginning, or they will let you do only one, or they won't let you do data extraction.

01:08:12 Tomas Havranek: But definitely, something along these lines will be super useful pretty soon. I was really excited about it in the spring, but since then I haven't really seen much progress, at least in the publicly available sphere. So we are still waiting, but a lot of help should be on the way. For multiple-agent or multiple-model AI debate or cross-checking, we have a protocol on GitHub. You have the link in the slides in your folder, which you might want to check out. If you want to check anything or stress test your own ideas, here is what I found useful. We also talked about it with Taisuke quite a lot.

You have a paper or a grant proposal, and you upload it to these different models. You assign each model a different role: one model is supposed to be the editor of the journal, another is a referee on meta-analysis, another is a referee on econometrics, and another is a referee on labor economics. Their task is to find the weakest points in your paper. Then you let them discuss, and the models know that the other responses come from AI, which makes them more assertive. Otherwise, many times they will just agree with whatever you put there. So you tell them: this is my paper, I want a brutal stress test. Then you let it run for a couple of iterations.

You sometimes get really useful, out-of-the-box criticisms, and sometimes you can implement them to make your paper stronger, or implement them even before you start to work on your paper. You may go to a fancy conference, like AEA or ASSA. You fly to San Diego or Philadelphia, in your case. You present your paper at the best conference in economics. And what kind of feedback do you get? If the purpose is to get feedback on your paper, I would bet it is more cost efficient to just sit at home, open these AI models, wait for about one hour, and then read what they give you. Of course, in practice, we want to do both. We still want to talk to people, but if you look purely at the scientific feedback you get from many general conferences, where you don't have discussants and it is not specialized, then I don't know.

But I definitely want to emphasize the point made earlier by Taisuke. If you do a meta-analysis, if you have already done one, or if you think about doing a meta-analysis, please come to our colloquium. Every year we have a colloquium, mostly in a cool place, so we try to make it fun. Typically it is really a superb location, in a World Heritage site. […] You will be really treated well.

This year it is in Chemnitz in Germany. Next year it is going to be in Greenwich, near London, in the UK. […] You should try this with AI, I think just for fun. You will see that it might be really useful in many cases. The bottom line is "don't rely on one model." Regarding AI, we already have guidelines on how to use AI in economics meta-analysis.

You have it in your folder. It is still not public, but it will be published pretty soon in the Journal of Economic Surveys. The Journal of Economic Surveys, by the way, is our flagship journal. It is like a field journal for meta-analysis in economics, and we are striving to make it better and better. If you write a meta-analysis and you follow the basic things we will talk about today, any of you who does a good meta-analysis will be welcome to submit it to the Journal of Economic Surveys. It is a journal with a solid ranking and a solid impact factor, if you care about that. Some universities do care. The guidelines would not be restrictive, just something that would give an overview of what can be done and what should not be done. [Note, 2026: the AI guidance is now published: Cook et al. (2026) in the Journal of Economic Surveys, linked from meta-analysis.cz/guidelines/.]

Sometimes people ask me what the minimum requirements are, what minimum number of papers you need for a meta-analysis. There are no strong reasons for these numbers, but we had to come up with some. We say you need at least 30 estimates. As in normal statistical inference, 30 seems an intuitive number, and that is what I was taught as a student: 30 estimates or 30 observations at least, and about 10 studies. You can have more estimates per study, but if you have fewer than 10 studies, that is difficult to do in economics. [Note, 2026: 30 estimates from 10 studies is the practitioner's guide's floor for modern meta-regression; as a rule of thumb, our notes for theses (meta-analysis.cz/ai/thesis-notes.md) ask for at least 10 studies and 50 estimates, or 20 studies and 100 estimates if you want to publish.]

Many meta-analyses in medicine would be just a couple of clinical trials. Let's say you have a new medicine, a new drug, and these clinical trials are really costly. They are precious. So even if you have three of them, it is worth doing a meta-analysis. Then, of course, you cannot use all of these methods I will be talking about, and you mainly look at some averages, some weighted averages, or at the median.

The main things you need to collect are, of course, your effect size, the effect you are interested in, and also the uncertainty around it: the confidence interval or standard error, which you can compute somehow from the p-values if only p-values are reported. You also need the sample size for many of the techniques, so it is good to have the number of observations from each paper. [Note, 2026: MAIVE needs the total number of observations behind each estimate, so collect it for every estimate; see meta-analysis.cz/maive/how-to/.]

Another question is: should we collect studies where we know these studies are not good? The answer is yes. In the end, you don't have to put any weight on these studies, but it might be useful for you to show how they differ from the good studies. If they don't differ systematically from the good studies, that is also interesting for a paper, and you might want to emphasize it. So try to collect all of them. The differences between the good and the bad studies might be a good part of your contribution in the analysis.

01:16:23 Participant: So how do you think about quality? That is something we have been struggling with a lot. I guess there are more established procedures in other disciplines, but beyond sample size and pre-registration, have you thought about systematic ways of measuring what is good and bad?

01:16:42 Tomas Havranek: No, that has multiple dimensions. We code many aspects of the context in which studies are produced: identification, sample size, and many other ways. There is not one dimension on which I would judge quality, and that is one of the reasons. Ex ante, I don't know. My prior could be that, for example, in economics, I would want to place more weight on experimental or quasi-experimental results.

But in the beauty paper, it is very hard to do because you have no experiments. You have only a couple of quasi-experiments, like some difference in differences and IV. We have an IV estimator for meta-analysis, so I cannot say that IV is always bad. But by the way, there is a nice paper by Abel Brodeur in AER 2020, I think, where he shows that IV estimates are the ones most likely to be affected by publication bias, because there are so many choices: which instrument you choose, and then how you play with what you report. That is not a good start. But my point is that for some researchers, you just have observational data, and you still want to have an answer to the research question. So what do you do? It is not clear before you start your meta-analysis.

01:18:34 Participant: And so in practice, will you look at subgroups, or have you tried to compute some index?

01:18:43 Tomas Havranek: I will talk about it in the heterogeneity section, so don't worry, we will cover it. I don't try to judge too much, and I don't try to construct any index of quality. But in many cases you can make a good argument that different ways of identification are better. If I have a good RDD, for example, that is probably more defensible than IV in many literatures, not to speak of OLS regressions where you just hope for the best, you put in some controls and you control for observables. So sometimes you can make this claim, and then, what we like to finish with when we do a meta-analysis is to have some implied estimates conditional on this. What is the implied estimate for this specific subset, for this demographic group, conditional on the use of the best quasi-experimental or experimental techniques?

Experimental economics is great. The problem is that in economics, when you do macroeconomics, for example, you would have to do experiments with countries. If you ask the central bank to just increase interest rates to create some random variation, it is unlikely the governor will be happy to do it. So you might find clever ways to get around it with quasi-experiments. […] It is so hard in economics to have a good identification for this macro time series. […]

Anyway, that was a digression, sorry. Yes, you mentioned that some fields do have quality benchmarks. They call it risk of bias. I don't like that much. They tell you they will be excluding studies based on it. I like to have all the studies in my analysis, and then in the end I show you explicitly what is the weight I place on this. It could be zero or very small. Typically it is not zero, just very small. But you use the entirety of the literature to derive the best guess: now that we look at the entire literature and some of the studies are biased, what is the best guess we can make about the parameter in question? But again, what I tell you is my personal preference, and it doesn't have to mean that it's the right way or the best way to do it.

01:22:18 Participant: But if the studies are somehow homogeneous, up to some extent, the standard error should capture part of the quality. And when the studies are more heterogeneous, which could be independent of the standard error, then I think, conditional on the study characteristics, the standard error should capture some kind of quality.

01:22:52 Tomas Havranek: We will also talk about this specific issue, how precision is related to quality and bias. That is a big issue in meta-analysis because, of course, as you know, the traditional approach in meta-analysis is to use precision as weight. So more precise estimates get more weight, inverse variance weights. I will show you some examples where it works and some examples where it may actually make things worse, more problematic.

What about unpublished papers? Gray literature. Again, my recommendation would be, if feasible for you, to try to collect everything. But sometimes it's just not feasible. In some of my papers, you have 200 published papers, so many that you can justify just looking at the published ones. But it's better if you can compare, because then you can play much more with publication bias. But it's hard, because if you have a paper which is unpublished but the draft is from 2022, what does it mean? Does it mean that this guy is still revising it for AER, or that it will never be published? It's not on his website. So it's tricky if you want to identify publication bias just from comparing published and unpublished papers.

So you would have to say, let's say, I will only take unpublished papers that are many years old. They can still be published, but it's unlikely. And I take them as papers which didn't make it for some reason. Compare them to published papers and then you can maybe say something about the publication process. But it's hard to do: typically you don't have too many papers which are never published. I know, for example, that Taisuke has a great paper on this registry of trials in economics experiments, which shows that insignificant results still tend to be published, but not really well, not in top five journals. So you could look at something like that. But it's more like meta-research than meta-analysis, so it's broader.

Nowadays, since you can use AI on this data collection or literature search, it is, I think, much more plausible and feasible to collect all of these studies as well. And you need some sort of documentation, which can be really simple. It's called PRISMA; it's like reporting standards for meta-analysis. So you essentially just code which keywords you used, what your search procedure was, and what you did. This is okay for economics, you have this example here, so you have the data in any case as you go, and you should just make notes of what you did. If you want to do this in psychology, it's more complicated. PRISMA typically looks more complex, and you have to follow these guidelines, which are more detailed. But for economics, this should be completely okay. […]

Data Collection 01:27:17

So now we have our topic, we have our studies, and we really want to collect the data for our meta-analysis. The main point which needs to be made is that the estimates we focus on in the studies need to be comparable. It's obvious, but I just want to emphasize it just in case. If you take one estimate from study A, which is 2, and another estimate from study B, which is 1, the comparison between 2 and 1 must make sense. That is not the case, by the way, if you use t-statistics. T-statistics are not a good summary measure because they depend on sample size. So don't use t-statistics, please. A few people still do it. It's standardized, but it's not good. Of course, in meta-analysis, because of the infeasibility of large-scale experiments, we do have observational data, which means regressions.

The best thing is if we can work with log on log, so elasticities. You can't always do it, but often you can. Elasticities are great because they have a clear economic interpretation, and you can compare them. This is perfect. Please, if you can work with elasticities, do it. Sometimes you will need to convert some studies from something else. Even if you have part of your sample in elasticities and need to recompute the rest, it's often worth it. So logarithms are amazing if you can do it. Sometimes, if you focus on the value of statistical life, you will have it not in elasticities but in dollars or yen. That is always perfect and a great summary because you can clearly compare it.

One example is a paper we have on class size: how important is class size for study outcomes? You have a lot of kids in class, which is typically the case in Japan. Compared to, for example, my hometown, where I chair the school board and the average class size is about 18 kids in a class, so these are typically small classes. How does it matter for the education of children? Spoiler: it doesn't really matter much. It matters less than we believe in Europe. But how do you do a meta-analysis on that?

Because the test scores across studies differ, like in the US and Japan, you need to use standard deviation changes in test scores as a response to a change in the number of kids by 10, because that's the most commonly used change, say from 25 to 15. You can divide something like that. Standard deviations will often be useful. The beauty paper also probably uses standard deviations. That is not ideal, because a standard deviation is not immediately intuitive. But it's probably the best you can do if you have very different measurements of the same thing. But at least one side of the comparison should be something absolute, like the number of kids or percentages.

Then it is at least partly intuitive and interpretable. If you cannot have absolute values on one side, you cannot do elasticity or normal values, and you cannot even do something like what we do in the beauty paper, a standard deviation effect on percentages. You have to do something standardized. If you have experiments, you would often do standardized mean differences: the result for the treatment group minus the control group, divided by the standard deviation. In economics, because we do not have much experimental research, you cannot use them for most questions. But if you have an experimental field, this is great: they are standardized, and you can use them. But be aware that this is not an economic effect, only a statistical effect, so it is a bit difficult to interpret. But there are guidelines on how to say which is large and which is small.

That is one option. Mostly in economics we would just look at correlations, partial correlations: you compute partial correlation coefficients from the regression. The good thing about this is that it is very easy to do. You can do this PCC in any observational data context where you have continuous variables. The problem with any standardization is that you will have statistical issues when you do a meta-analysis because (I do not have it in the slides) it can be shown that if you compute the variance of this, it is actually a function of the PCC itself.

So you will have endogeneity in a meta-analysis, and I will show you later ways to tackle it. But you should know that if you use this, it is problematic. The mean differences have the same issue: there is a mechanical correlation between estimates and standard errors. It is a bit complicated to explain, so I just ask you to trust me, because it would take a lot of time to show it to you. But this is the last resort. If you cannot do anything else, you can at least always do partial correlations, and you can do a meta-analysis. But then the interpretation is a bit more difficult. What does it mean if you get a PCC of 0.1: is it a lot or not? That is why I much prefer elasticity or dollar values.

As I say repeatedly, it is great if you can collect all these estimates on the same topic from the literature. If you can do that, it helps if you can also code information from the authors themselves on what they think about the estimates. Sometimes they will tell you, "We do this just to show how wrong it is." This is an example from the class size paper. There are entire papers that do not have any new estimates or estimators, but just show how wrong the previous literature is. And the question is, should we collect these too? My answer is yes, but we should take this into account.

We should take into account what the author says about his or her results. Typically you would have preferred results, then some robustness checks. And you might also have results that the author says are incorrect: "When I use RDD, I get this, and I will show how wrong I would be if I relied on IV." It is probably useful to collect this information and see if it systematically matters for results, because then you will probably want to put more weight on the preferred estimates. Here I will just briefly mention that you need standard errors for any advanced meta-analysis technique. If you have just the estimates reported and no information on confidence intervals, you are unfortunately in trouble in meta-analysis.

Sometimes, when you really need to do it, you have many studies without confidence intervals, without precision. You can somehow approximate it, maybe using a bootstrap: you have many estimates per study, which tells you something about uncertainty. But I would say it is a last resort. We have also published papers that do this. Even our meta-analysis of discount rates, published in Experimental Economics, heavily relies on this approximation, because discount rates quite often are reported just as numbers, as estimates, without the confidence intervals.

How can you use AI to collect data? This will be similar to the previous points about literature search. Again, as of now, people unfortunately still need to be involved. We cannot just ask the AI agent, "Just collect data for me, and when I am back from lunch, I want the data set." Or, ideally, I want the entire meta-analysis after dinner. Sorry, but it is reasonably good as an additional coder. It can help you cross-check how you are collecting the data and how reliable the human who collects it is.

Prior to that, it was recommended to have at least two coders who would code at least some of the studies, not everything, independently and then compare how they agree or disagree before they go on and collect the rest of the data, to settle this agreement and make sure they are doing it in a correct way, consistent between the coders.

We have some experience with that because we collected data for this huge paper, not in economics but in health science, a part of sports medicine, on how exercise is important for cognition. There are so many papers in sports medicine on exercise and cognition, almost all of them experiments, that you even have dozens of different meta-analyses on the same topic. So there was a huge influential meta-meta-analysis, a meta-analysis of meta-analyses, whose conclusion was that exercise is great for cognition. You will be more powerful in terms of cognition if you exercise.

But the problem is that it completely ignored publication bias, because it just took the results of the meta-analyses. Some of the meta-analyses did correct for publication bias, our next topic, and some did not. So we used the same set of studies, but we collected the data from the individual meta-analyses to be able to work on publication bias. The result is that it is not as great. It is still good to exercise, but it will not really improve your cognition much. It has other benefits.

01:40:00 Participant: Can I ask something about this meta-meta-analysis? I never really understood the point of doing this, precisely because you already have access to the individual studies and there could be double counting across meta-analyses. I do not know. I always wondered why people do this.

01:40:17 Tomas Havranek: I think that is a reasonable objection. First of all, if you just use the results of this meta-analysis, it is really as you say: you have these problems, you cannot correct for some biases, you have double counting, you have overlap. By the way, sample overlap is a problem not just for meta-meta-analysis but for any meta-analysis in economics, because, especially for observational data in, let's say, macro, you have just one series for GDP and inflation.

Different studies will use essentially the same data but different techniques, and we will talk a little about how to take this into account. You can maybe bootstrap on some of these issues, but I just wanted to make this point: it is not a problem just for meta-meta-analysis, or even just for meta-analyses. The studies we have are not independent, because they either use similar data or techniques or have similar ways to selectively report results. But with such an important topic as exercise and cognition, you have 50 different meta-analyses with different results.

[…] And I think it is our role to give them some answer, whether they decide to use it or not. So you need somehow to synthesize all these things. But of course, much better than to take the meta-analysis results and then synthesize them again would be to go to the original results and synthesize them freshly, and take care of the double counting, which is what we do in this paper. […]

01:42:51 Participant: How about Elicit or other search engines that are AI-assisted?

01:42:59 Tomas Havranek: I have not tried that, so thank you. Again, the problem is that it is changing too fast, and I struggle to keep up. Probably in a couple of months these slides will be completely obsolete, but as of now they are at least something. My advice for AI here is very similar to my advice when we were talking about the collection of studies. Again, multiple models, if possible, are better, just to cross-check. I like to use NotebookLM. Do you know it from Google? You may correct me, but my impression is that it is less likely to hallucinate, because it is really based on what you upload there.

I did not mention it before because it was not so important, but especially here, when you want to extract data from PDFs, NotebookLM could be really important, and you could maybe even use it as your main tool.

01:44:11 Participant: Some people I know do that.

01:44:17 Tomas Havranek: We have some examples of how you can prompt it, but I wanted to stress NotebookLM. I will not read another example, but I will tell you a little about it. We have these AI guidelines, which I already mentioned, where we discuss it in a bit more detail, and we also have a detailed discussion on the website of MAER-Net (maer-net.org). If you go there and navigate to the blog section, you will see quite long discussions on the use of AI, where you can also get some interesting points. I also have a blog post there where I summarize, as of, I think, August last year, the ways you can use AI for data extraction. You have specific steps there, and you can look into it. But again, I will probably need to update it very soon, maybe next summer. So stay tuned.

Very soon we will have a series on AI by Bob Reed on our blog website. Do you know Bob Reed from Canterbury? The series will have a concrete example of a meta-analysis done in AI-assisted mode this month. It is going to use all the fresh models and all the fresh information, so you will have a concise, brief, up-to-date explanation of precisely how you can now use AI models to help you do a meta-analysis. This should be out very soon. […]

One example is that in a meta-analysis you do not have to collect just numerical estimates. Especially in macroeconomics, which is my field (I worked for the central bank and also for the government), you would often have impulse responses from vector autoregressions, for example. The message of the paper would then not be a table but a figure. But of course, you can extract numbers from the figure, and now you can do that again using AI. Before AI, we would just measure pixel coordinates.

We have two papers on this. One example is this paper from the IMF Economic Review, where we look at what happens to house prices, or home prices, when you change interest rates: the monetary policy transmission from interest rates to home prices. People would report this, and we recompute it to slice it into different horizons. We compute the response and the confidence interval. So it can be done, and then we proceed as in a normal analysis.

One objection would be that you have measurement error when you transform figures to numbers. You do not have the actual numbers from the authors, and they use a thick line here, so especially before AI, when I had to measure pixel coordinates, you would inevitably make small mistakes. Is it a big problem? Do you think it is a bigger problem than when you have numerical data?

I do not know the answer, but my intuition would be that this is actually better, because when you collect the data from tables, the numbers will be rounded. Rounding could introduce a systematic bias. There are papers on that by Abel Brodeur. For example, if they are t-statistics, they will often be rounded to full numbers like 2. But then if you run a publication bias analysis, it may show a big spike at 2, which looks like publication bias but could actually be rounding.

You do not have this here. The measurement error in meta-analysis from collecting numerical estimates could be non-classical. It could be systematic, but here it really should be random. It is just my thick fingers when I click on the line. So classical measurement error is relatively fine: the only thing it will do depends on which side of the equation it is on, but it is much less problematic than the non-classical one. Here, in principle, I can ask two people to measure it.

And because it is a classical random measurement error, I can use the second measurement as an instrument for the first measurement, and I am fine. The referee will be completely happy about that. I cannot do anything like that for numerical results, or I could, but it is much more complicated. This is just a small caveat. You might have referees who will ask about that. There are ways to tackle it, but it is a niche topic in meta-analysis. But it could arise. […]

Afternoon session (recording 2) 00:00:00

00:00:01 Tomas Havranek: […] If you don't mind, I will skip the code. It is in your materials and should be self-explanatory. I will not run the computation here because it would take additional time, and you can easily do it yourself. On the slides, there is always an indication of where you can switch to look at this code, so it should be pretty easy.

We are talking about how you can also collect, in a meta-analysis, data from studies that do not report numerical results but do report graphical results. For such studies, you can use either AI or measure pixel coordinates.

The next problem you will always face in meta-analysis is what to do with outliers. Unfortunately, there is no clear answer. I will give you mine, but other people have different answers. You will always have outliers, which means that sometimes you have estimates that are just strange. It could be a typo in the original study, or it could be a real estimate that is really far away from the mean, but only for some specific context, subset or demographic group, or for something really wild.

If you don't do anything about outliers, they will mostly be a mess for your meta-analysis procedure, so you should do something. The solution I recommend is winsorizing. Say we winsorize at the 1% level. That means we shrink both tails, positive and negative, to the first percentile and the 99th percentile: anything outside these percentiles, we shrink to them. We do not omit or drop these outliers. We keep the information in the data that it is a large estimate, but not an extremely crazy large one that would create problems. I do this for both estimates and precision, because these are the key things that you have to consider in meta-analysis.

00:02:43 Participant: What about the combination of those? Suppose I have an estimate like here. It is inside.

00:02:55 Tomas Havranek: Then you would have to consider both variables, that is, more dimensions. As I said, there are many ways to do it, and some people are more into complex solutions. They look at the funnel plot, which we will talk about shortly, and if something does not belong, they use some metrics, maybe regression metrics such as DFBETA, the standard treatment of outliers in regression when you have not just one dimension but more dimensions, to decide how to trim it.

But this is often difficult to justify in an economics paper. You can justify winsorization. All of my papers essentially use it. You should also show different levels of winsorization, and what will happen if you do not use any winsorization. That is feasible to publish, and people can accept it. If you run a sophisticated mix, in which you try to make it look better, I sometimes struggle with how to justify it.

Usually, if you have something like this, it is going to be an outlier in both directions, and there is going to be a deep reason why it is there. […] Or it could really be a mistake that needs to be deleted, something that is completely off. People might mess up the units in the original study, so this might be a reason to look at the codes of the primary study. But the simplest solution really usually is some sort of winsorizing. That is essentially what I have just explained: you choose a winsorization level, and if something is really super large, it will be kept in the data, but it will be brought a bit closer to the domain.

I have an example in the code, so you can look at how it works with the beauty data set. For the people who haven't been here before: in the first session we used a meta-analysis of the beauty premium as an example, that is, how being more beautiful will increase your salary. There have been dozens of papers on that, and we have it as a test data set for meta-analysis in the folder I provided. As I said, I will skip the demonstration in order to be able to finish, but it is all there, fully described and annotated, so it should be okay.

00:06:03 Participant: How do you justify this procedure?

00:06:12 Tomas Havranek: How do I justify it? I need to do something about outliers, because there are huge outliers and, despite the best of my efforts, sometimes I cannot explain why an estimate is 1000 when it is totally implausible. If you keep it there, it will really mess up your analysis. It is not a perfect solution, but I think the other solutions are worse: do nothing or delete it. The justification is that it is a symmetrical rule. If something is out of line on both sides, we shrink it in a symmetrical way.

If you just delete something, there might be a different number of these large outliers on one side than on the other. So the justification is that this is the least intrusive data cleaning solution there is, but again, that is just my personal preference. How to handle outliers is an unsettled issue, I think, in any kind of empirical work in general. But here it is very important.

00:07:29 Participant: [A participant asks whether winsorizing is still advisable when the distribution is asymmetric, for example when there are no observations on one side.]

00:07:47 Tomas Havranek: […] Still, that is the least bad solution I know of. That is what I do most often, with robustness checks.

That is all I can say about that. We are still in the data collection session, so a little about how you collect the context. For this beauty effect analysis, this is how the final data set looks in Excel. Of course, we do not collect just the estimates, their precision and the sample size. What we care about a lot is the context in which the estimates were derived, computed.

For example, how do you measure beauty? Do we have raters in the room? Do we invite people and rate them from 0 to 10? Do we give them photos, which is the most frequent way to do it in empirical work? Do we use some sort of AI rating, which is of course much more prominent now, but even before the current AI wave, people would use algorithms to measure facial symmetry? Because if you ask what beauty is, it is mostly facial symmetry.

Symmetry is heavily related to how people perceive beauty, and there are some other ways too. The other main thing we care about is beauty and success, or beauty and earnings. Sometimes it is earnings, but sometimes we have study outcomes, because the paper is broader than just salary or wages. We have research outcomes as well: some studies do look at how beauty makes you more productive as a researcher. Sports, elections, sales. But it is not just about the definition of the two most important variables in these primary studies, but also about the subjects. Do we look at men, women, sex workers?

What kind of culture do we look at? Is it Western, or maybe Japan? As for estimation, it is mostly OLS. You cannot really do good experiments or even quasi-experiments. You have a couple of diff-in-diff estimates from, I think, plastic surgeries, and you can do a little diff-in-diff on accidents, such as car accidents. If you do OLS, you can try to measure non-cognitive skills or cognitive skills, which will be a main focus of our paper. Non-cognitive skills are hard to measure. Sometimes it is self-confidence, so you either ask people to self-rate or let others evaluate them.

For a cognitive skill, you may want to try to measure IQ. Sometimes you have something related. As you can see, the goal is, in general, to identify the most important characteristics in which the estimates differ and in which the studies differ. You can never collect all of the reasons, as that would be essentially an infinite pool of potential moderators, but you can focus on the ones that should be important: a couple of dozen at most, maybe 20 variables, maybe fewer.

We have a dummy variable, which we coded as one for these top five journals, plus maybe REStat or something like that, because some referees would tell you, "I don't care about this. I just care about the top five, maybe REStat, and these top journals. Can you just show me what your results would be if we only focus on top journals?" One way to do it is just to control for this, to try to do ceteris paribus. But I think we also had a subsample on that in the paper. You can also control for things like whether the paper was published in a peer-reviewed journal. You can add citations, whatever. Essentially, these are the types of things you would like to code when you collect your data.

00:13:12 Participant: What about citations, actually? How do you think about this, especially when you are putting studies across different disciplines, or potentially fields, and then you have huge variation and outliers?

00:13:24 Tomas Havranek: That is true, even within economics, of course, in different fields. But when you do a meta-analysis, it is commonly just one field, so the citation norms would not be the same, but would be relatively related. What we most often do is just use per-year citations, because you need to take a look at how old the study is. Even if that is not perfect, it is good enough. You can use it: some referees will tell you that it is a good indicator of unobserved quality, because you have the methods and the data, but it does not fully capture how well the study is done.

You might want to use this: high quality top five journals, whether the paper is peer-reviewed, how many citations it gets, as a proxy for quality that is not directly observed. But you need to be careful, because the number of citations could also reflect the results themselves. If the results are intuitive and easy to use, people will probably quote the paper more, as with many of my papers. (A brief digression: I have written many dozen meta-analyses, and I am happy for any citation.)

Sometimes the citations do not take the main estimate, corrected for publication bias, but quote the mean. In the middle of the meta-analysis, I just say, "We collect these estimates, and the mean is such and such." That is, of course, not all: we continue, we do the publication bias and p-hacking correction, and then we give our bottom line result. But some people like to quote the mean because that is what they need for the calibration of the model. So I have tons of citations that do not really reflect what I find in the paper. But I do not complain, I have to admit. I have never written to anyone saying, "You quote my paper wrong."

So there could be a citation bias. If the results are such that they make a paper easy to use, people will quote it more, certainly. One needs to be careful. It is the same thing with high-quality peer review, as Taisuke's paper shows. The insignificant estimates are less likely to be published in top five journals. That is probably not about quality at all; it is about publication bias. Sometimes you are able to do in a meta-analysis something you would not be able to do in a primary study.

For example, we have this paper on how much daylight saving time actually saves energy. Do you have daylight saving time in Japan? [Someone in the room answers no.] That is probably good, because it does not save much. We do in Europe, as you know: we change the clocks, in most of the US as well, and people are used to it now. But the main reason why it was implemented first was savings in terms of energy, electricity, lighting and now other things as well.

We do an analysis of these, mostly quasi-experimental papers, where typically every paper is a different country. Then we try to see whether the country level matters. For example, if a country is really close to the equator, it probably does not make much sense to have daylight saving time, because there is not much variation in daylight through the year. Yet there are still subtropical countries which do have daylight saving time, for historical reasons: they used to be a colony, or they do what the developed countries do […].

So you can actually test it. You can use either the latitude of the country, essentially how far away it is from the equator, so how many daylight hours you have on the longest day of the year, in June or depending on whether it is north or south. I will explain this figure later, when we talk about heterogeneity. Just very briefly, what we find in this paper, and this cannot be done in primary studies because they are done for each individual country, is that the position of the country matters for savings from daylight saving time. If you are closer to the equator, you are likely to have negative savings, so you consume more energy, which is related to how you use air conditioning and so on.

It is a complex issue and you can read the paper if you want, but it is just an example of a contribution which meta-analysis can give you. On top of a sophisticated summary of research, you can also include original hypotheses which are unfeasible in primary studies. Now I will go through the classical meta-analysis techniques used outside of economics, really briefly so that we have more time.

Conventional Tools 00:19:22

One of the points I want to make is that if you do not know what kind of summary to use, and maybe you have just a couple of studies, you should use the median. The median, as we all know, is very simple and surprisingly robust to many of the problems we will discuss, like publication bias. Of course, it is not a full solution, but it is so simple and so robust that, if you compare it to these more complex techniques, it is often better. If you cannot do anything else, use the median, the simple, unweighted median.

00:20:05 Participant: [A participant objects that the median estimate could be very imprecise.]

00:20:14 Tomas Havranek: But it is still in the middle of the continuum of the estimates which you see in the literature.

00:20:34 Tomas Havranek: You can also compute confidence intervals around the median. Sure, that is not a bad idea, but it is not what people do in a classical meta-analysis.

In a classical meta-analysis, this is just a simple simulation where we have the estimates on the vertical axis and the standard errors on the horizontal axis. In an ideal situation, it looks something like this. If you have really precise estimates or studies, they will be close to the true value, which in this simulation is 1, so we know what it is. If you have small samples and a lot of imprecision, you will be all over the place. So it makes sense intuitively to give more weight to more precise estimates, and that is what people do. That is what has been done since meta-analysis was invented in the 70s. If you know nothing else, it makes more sense to give more weight to these precise estimates.

That is called the fixed effect estimator in meta-analysis. It is not the fixed effects we know from econometrics. It is something completely different. The fixed effect estimator means that you assume all of these studies estimate the same thing, which is highly unrealistic in economics, but you should still know how it works. You have two studies, one of them more precise than the other: one estimate, another estimate. […] You assume they are both drawn from a distribution with the same mean, and then you compute the weighted average. In practice, fixed effect estimation is simply weighted by inverse variance. That is all. There are some equations, but we do not have to spend time on them. You simply use the inverse of the variance as the weight and then compute it.

You have examples in Stata. Sometimes you can also see, especially in economics, something called UWLS, unrestricted weighted least squares. It is essentially the same as the fixed effect estimator, but the variance of the meta-analysis estimator is computed in the typical economics weighted least squares way, so it is more conservative. The meta-analysis fixed effect, for reasons we do not have time to talk about, is restricted, and typically it will be super precise. It is just not realistic. So the common weighted least squares, where you use inverse variance as weights, will usually give you more conservative confidence intervals. To repeat, it is nothing more than a weighted average, with inverse variance being the weight.

For example, if you have a couple of studies which estimate aptitude at the same college, using the same sample of people or the same pool, it makes sense to do this weighted average, because these are homogeneous: they estimate the same thing. Now, in practical meta-analysis, what is used much more often is random effects. Again, it is something different from random effects in panel data, but somehow related. You assume these studies can be drawn from different distributions, distributions which have different means, but that any differences across studies or estimates are random, a little bit like in panel data with random effects. So you estimate not just the final estimate, but also the heterogeneity across studies.

The way it differs from the previous weighted average is that in the weight you do not just have the variance, but you add this heterogeneity part, the between-study heterogeneity, that is, how much the underlying effect differs across studies, which you need to estimate. This is tau. It is important. If there is just a little bit of heterogeneity, if tau is very small, you will be close to the fixed effect setting we had previously. If it is super large, this will be essentially close to what?

00:26:16 Participant: [A participant answers from the room; unclear.]

00:26:20 Tomas Havranek: If there is a lot of heterogeneity, the weights will be close to equal weights for each study. So if you use no weights at all, which is the way I like to do it quite a lot, you are implicitly assuming, by this logic, that there could be a lot of heterogeneity. You need to estimate tau. There are some techniques to do it, but they are too complicated for today's lecture. It would take a lot of time.

Now, a simple example. First we looked at fixed effects when measuring aptitude in one college. If you want to do it across colleges, it is more plausible to assume random effects, because these are different colleges and the effects will be different. So it is better to use random effects. What will happen here: the weight of the less precise studies will be a bit bigger. The weight is diluted by the addition of the tau squared in the denominator of the formula.

This is a figure you often see in meta-analysis. It is called a forest plot: you have many studies which you include, for each study the estimate and the confidence interval, and then the overall meta-analysis estimate, typically from random effects. In economics, as we discussed on Saturday with Chishio, it is hard to produce these forest plots because typically you have many estimates per study. You would first need to do some sort of meta-analysis within the study, which is an assumption, so it is hard to do.

One way to do it is to use the median from each study and then somehow construct the confidence interval for this median. […] We did something like that. But in general it is difficult when you have multiple estimates per study.

What you can also sometimes see is a cumulative meta-analysis, which looks like the forest plot from the previous picture, but it starts from the first study, the first paper published in 1987. When you add more studies, it is not just that this is the result of study two, but it is already a meta-analysis of these two studies. And this one is a meta-analysis of the three studies. When you add more information, you get more precision, and you converge, in an ideal case, to a value which is more in line with what you are actually trying to estimate, the underlying value.

So this was real quick, and now we have more time to focus on what is really crucial, publication bias. Let's go into it.

Publication Bias 00:29:49

A little bit of background on why publication bias is bad. You might have heard about the scandal with antidepressants, with Paxil in the US: there was an antidepressant which was heavily prescribed […]. The problem was that the published clinical trials were just a subset of all clinical trials. There were many others which showed that it does not work well, that it has side effects and so on, which was discovered only afterwards, when many people would suffer from this […].

[…] It was a big scandal, and based on this the pre-registration movement started, first in medicine. For many years now we have it also in economics for experimental research.

Now it is much more common even for observational research. Pre-registration means that if you want to do an experiment or a clinical trial, you first need to pre-register, and then even if you decide not to publish the results, there is some sort of record that the experiment was conducted, and maybe a way to check the results. That is one way to combat this publication bias. Of course, if you do not do an experiment, but run a regression on inflation or GDP, or do what we often do in economics, it is very hard to rely just on pre-registration, because you can never know whether the people have already seen the data and are writing the pre-analysis plan when they already know what they will try to push. Still, it is good practice nowadays to try to have a registered protocol before you do your paper.

[…] Now, how does it work in meta-analysis? If the problem is not fully resolved by pre-registration, which it probably is not, how can you correct for it? The key thing here, which is common to almost all models, is that if you have no publication bias, there is no problem at all. You call these estimates of the beauty premium gamma, and these are just regression coefficients.

If you think about how regression works, then in this standard normal setting the implication is that the ratio between the gamma estimate and the standard error has a t-distribution. That is how people interpret it, and it can be shown that it follows from the regression techniques. So the distribution of this ratio is symmetrical. Put more simply, it implies that there should be no relationship, no systematic relationship between estimates and standard errors. They should be two statistically independent quantities. There are some disclaimers, but that is the basic idea. I will show you in a simple simulation how this correlation arises and how it is related to publication bias.

Let's assume there is a parameter whose true value is one, and different studies try to estimate it. We will have different studies here, with their estimates on the vertical axis, and we will have, as before, the standard errors, a kind of inverse precision, on the horizontal axis. So let's simulate a lot of studies. Again, the most precise will be close to the true value of one. When you have less precision, you will have more dispersion. Since there is no publication bias, if you just look at the correlation between estimates and standard errors, there is going to be no correlation. The regression will be a flat line. That is fine.

Suppose that people do not want to report negative results because it does not make sense that beauty would make you earn less on the labor market. It might, but a justification is difficult and we would have special cases. That is a plausible thing that could happen: the negative estimates are not reported as much as the positive ones. Even though we know here in the simulation that the true effect is positive, sometimes, if you have enough imprecision, you can get an estimate which is negative. If you get rid of such estimates and then, in a meta-analysis, do a regression of estimates on standard errors, you will get a slope. The slope is a simple way to test for publication bias. We will get to disclaimers and these different problems quite soon. Now, there could be a different type of publication bias.

For instance, people might want to publish negative estimates if they are statistically significant. The slope again allows you to test for publication bias under certain assumptions. What is beautiful about this figure is that the slope is a test for publication bias and the intercept is a solid estimate of the true value beyond publication bias. That is the background, or the justification, of the most basic but quite powerful bias correction techniques.

I use the same simulation, but I switch the axes: instead of the standard error here, I show the precision, which is one over the standard error. This is the standard way we show it in meta-analysis, following medical research. It is the same simulation with no publication bias, and if you compute the mean, the average is 1. If you have publication bias and you compute the mean, it is more than 1, because you get rid of the negative results. If you have a stronger bias and you only report positive estimates which are statistically significant, the mean is going to be even more biased. You can take the difference between the mean and the most precise estimates as a measure of publication bias.

That is another motivation for how these methods work. To put it in a very simple regression form, the most basic model you can have for publication bias correction is the regression of estimates on standard errors. The slope, again, is a test for publication bias. The intercept is one way to estimate the true effect. It is very simple, but it is not so bad. If you can do this, that is actually okay. In practice, there is obvious heteroskedasticity here: the distribution of estimates will be related to the standard error, because the standard error is a measure of variance, or dispersion. So you have heteroskedasticity by definition, and many people will tell you that you should do weighted least squares.

You should weight it by inverse variance and then run the regression. But even if you do this and then use some heteroskedasticity-robust estimator, it is usually okay. [Note, 2026: a funnel asymmetry test alone is not enough. FAT-PET is now an optional robustness check; the core is RoBMA, MAIVE and RTMA, with standard errors clustered by study (CR2).] We have examples in Stata if you want to see how it works, but it is relatively simple. I will go back to this picture here. What you can see here is that the intercept is not a perfect estimate of the true effect, so there is going to be downward bias. In many simulations, it is better to fit a quadratic line than a linear one.

Instead of a linear regression, it is usually better to run a quadratic one. It is going to be less biased. This is called PEESE (precision-effect estimate with standard error), which is the common name for it. This very simple regression is actually quite hard to beat. You can devise really complicated selection models for publication bias, Bayesian models with many components, but this very simple one usually works quite well. We will get to disclaimers quite soon. [Note, 2026: for publication bias we now recommend RoBMA, which averages selection models and funnel-based methods such as PET-PEESE, weighting each by fit.]

Most of these models try to estimate the top of the funnel here, because the top of the funnel should be close to the true effect and should not reflect publication bias, which will typically drive the mean upwards.

One way is the PEESE estimator, which fits a quadratic line here, but you can do it in a more sophisticated way. For example, there is the WAAP estimator by John Ioannidis and colleagues, published in the Economic Journal almost 10 years ago. I will not go into details again, but the estimator computes retrospective power and focuses only on the estimates which are likely to be powered enough, such as those with 80% power. In effect, it is very similar to just looking at the most precise estimates at the top of the funnel. […]

It works relatively well. Of course, the problem is that you do not know the power, because you need to compute the estimate first, so it is a circle. First you compute an estimate of the effect, then you compute power, then you refine your estimate. That is the drawback. What I like a lot is Chishio's stem-based technique. Chishio might tell you much more about it later on, or during dinner, but the main idea is that you exploit the variance-bias trade-off.

If I go back to the funnel plot, as you go down you get more bias, because the less precise estimates are likely to be more biased. You would like to use the less biased estimates, but as you go up, you lose data and information, so your variance will also increase. Chishio shows in a nice, sophisticated way that you can exploit this to construct a nicely founded meta-estimator. The paper was originally part of his MIT dissertation, and hopefully it will be published soon.

There is another way that you can do this. I will go back to the funnel plot again. What is really going on here with this publication bias process is not a quadratic relation: the relation is horizontal up to some point, as here, and then it is linear. Pedro Bom and Heiko Rachinger found a way to find the kink. The method is called the endogenous kink. You can try to estimate first the horizontal segment and then the linear segment. Because it is quite a complicated estimator, you have the full code in Stata in the files, so you can check for yourself how it works. If the technique does not find any kink, it will just reduce to this linear regression, which is what happens with our beauty data. It just does not find any kink.

In principle, you can also do a non-parametric fit. I have not seen any paper on that, but I think it would work too. The kink in the Bom and Rachinger paper is estimated analytically, but you could also do it non-parametrically.

That was the kink. Here we have the formula for the kink and the way it is done, and the code is in this data file. A completely different way to correct for publication bias is to do a selection model. That is not based on this funnel plot, which is the scatter plot of estimates and standard errors, at all. The idea is that first you assume some distribution which the underlying estimates probably have.

Usually you assume a normal distribution, or you might assume a t-distribution. Then you look at which estimates are really reported, which are significant and which are not. Then you compute, given that your assumption is a normal distribution and this is the distribution that you observe, how many more significant estimates you have than you would have assumed. You compute this likelihood and use it as a weight. Essentially, you then give more weight to insignificant estimates, because they seem to be penalized by original authors who do not like to report them or by journals which do not like to publish them so much.

That is the essence of a selection model. You try to compute the likelihood under some assumptions, typically a distribution assumption, that significant estimates are more likely to be reported than insignificant estimates. Many people in the meta-research field like selection models more than funnel-based techniques like this PEESE and all of these previous ones, because they are really based on assumptions about how people operate, so they are not so much data-driven.

The problem is that, in practice, these selection rules are not really strict. You need a large sample to be able to do this, and it will be difficult especially if you have only a few insignificant estimates. Sometimes the Andrews and Kasy model does not converge. It has all sorts of troubles, but I think in general it is good, when you do a meta-analysis, to have both: to show a selection model approach, but also some of these funnel-based techniques, which I showed before.

We will talk about it more. To emphasize again, if you want to correct for publication bias, you have two main strategies. One is based on a funnel plot and meta-regression, a simple regression of estimates; the other is a selection model, which is typically maximum likelihood estimation. Here are some estimates for the beauty effect, but you do not have to read them; they are just for this case. We run the selection model, the stem-based techniques by Chishio, the kink model, and different versions of the regression, where we add study dummies or do between effects. We do an IV estimate, which I will talk about soon, and we run different versions of weights. Essentially all of the corrected mean effects are much smaller than five. The unweighted mean was five. When we correct for this publication bias, we get something like three, and sometimes even much, much less if weighted by precision.

There is publication bias in this beauty effect literature. You can see the funnel plot here, and this is the mean. It looks very much like the simulation I was showing before. There are some negative estimates, but it seems that there should be many more, which are not reported because it is just awkward to report that being more beautiful is bad for you on the labor market.

In general, the most precise estimates are pretty small. The only exceptions are sex workers: their beauty is important for their professional success. If you see a funnel like that, a scatter plot like that, it tells you that you should probably treat this group completely separately and not mix it up with the rest of the sample, because that is really a special profession in this case. The rest of the funnel looks essentially okay, with no huge outliers. I think we do 1% winsorization, but we probably do not even need it that much in this paper.

The question is why: you have found some evidence for publication bias, but what is the behavior behind that? How do we get this publication bias? How do researchers behave? A simple way to look into this a little bit is to do a caliper test around an important threshold for the t-statistic. I mentioned that zero is an important psychological threshold: if you have an estimate which is just positive, even if it is not statistically significant, it looks much better on paper than something negative with the wrong sign. So you might expect people to be a bit more cautious. If you have a negative beauty effect, maybe you try to run the regression again, but differently.

Even if it is just a robustness check, you do not want a robustness check which shows a negative effect of beauty. Indeed, that is what you see in the data: just negative estimates are much less likely to be published than just positive estimates. You can look at the boundary and compare how many estimates there are. That is what the caliper test does: you just compare the number of estimates just below the threshold with the number just above it.

It has its own problems. I think Graham Elliott has a nice paper in Econometrica, a follow-up paper on why caliper tests can sometimes be misleading. The rounding problem could be an issue here, but around zero probably not so much. You can do the same for these significance thresholds, where you see a little action, but it is not out of line with the other changes here. Zero is what is really important there. […]

P-hacking 00:53:44

Now to p-hacking, which is different from publication bias, as I will explain soon. First, some self-promotion. We have a website, EasyMeta.org, which allows you to correct for p-hacking and publication bias really easily. As the title suggests, you need no coding skills, neither Stata nor R. You just upload your data, and it does it for you. So check it out; it should be very simple. In the end, it gives you the R code for the results, so you can replicate them. If you write a paper, it gives you a full replication package for what you have produced.

Publication bias is what we described: some estimates are estimated but then not reported. What can happen with p-hacking? We can actually move these points. It is not just that you decide to publish or not. You can say: the sign is not right, so I think my model is not correct, so I need to add more moderators, or maybe I don't need this moderator, or I should do something about outliers a bit more aggressively. So you tweak the model a little bit.

[…] I think that is how empirical research operates. Even if you are not consciously trying to gain statistical significance, it is better, even unconsciously, if you get something that looks stronger and more plausible, and that looks as if you might have a chance to publish this really well.

The main difference between publication bias and p-hacking is that publication bias is like zero or one: you either see the study or you do not see it. With p-hacking, it might change, and you do not know where it will go. Once again, publication bias treats all published estimates as individually unbiased. So the literature as a whole can be biased, but if it is something published, we can just reweight it, and it is going to be okay. With p-hacking, it is not, because the individual estimates can be biased by this process of multiple specification changes: you go back to the data, you get rid of more outliers, then you produce your results, you use a subset, and so on.

That is much more difficult to correct for, and much more difficult to model. Some people would tell you it is impossible, but that is not very useful. We know it is a big issue, and we want to do something about it. I will show you a couple of ways that can be used. Here is an example of how this could happen. Let's say you want to estimate the education premium, the effect of education on earnings.

But of course, you cannot just run a correlation. You will have unobserved ability, which is related to both. We all know it; it is a common economics problem. If you ignore ability, you will have a large education premium, because the education premium you estimate will also capture the ability premium. And it is likely to be quite precise, because you have omitted variables. It is not 100%, but in most cases you will have higher precision. Imagine now what will happen to your meta-analysis when you have studies like that.

Of course, you can never publish anything now on this topic if you just ignore ability. You can include some proxies like polygenic scores, or before that you would include IQ, which is better. You will have a smaller gamma, and it is going to be less precise. Now you have two estimates. One is more precise, but it is obviously worse. So in a normal meta-analysis framework, you would put more weight on something that is obviously less correct. So putting more weight on more precise estimates could sometimes backfire.

Not too frequently, but in observational research, when you do have to make choices like that, it is quite plausible. And then, of course, to be really able to publish this paper, you need some sort of quasi-experiment: you can use some policy reform as an instrument. Earlier, people would use distance to school as an instrument. That is not perfect either, but it is likely to get closer to what gamma really is. But by the way, it is going to be much less precise.

So the least precise estimates could be the ones that are most credible. That is the point here. What you will see in the funnel plot with p-hacking are these hollow circles: if you just use a different instrument or a different technique, you can push your estimate beyond the significance line if you are so inclined. There are so many choices you can make as a researcher. That is the problem of p-hacking. By the way, the p-hacking problem is amplified in meta-analysis when you put more weight on more precise estimates.

Because you can have p-hacking on both, not just on the t-statistic but on both estimate and precision. We have a paper in Nature Communications in which we simulate how p-hacking works in environments where people can p-hack. They have a true model, which we know, and they also know how they should do it, but they can change the control variable to something else that is correlated with the true variable, which should be the control, but not perfectly correlated.

So there is potential for p-hacking there. This measures the correlation here between the control variable and the variable of interest, but that is not so important. When you allow people to p-hack a lot, they will change both the estimate and the standard error, because when you use a moderator, a proxy, which is really irrelevant, it will not be related much to the other variable, and you will have less collinearity in the model.

These are different meta-analysis estimators, and the vertical axis shows the bias of these estimators as we increase the potential for p-hacking. You can see that all of these models eventually become as bad as a simple mean, which is the naive, completely stupid way: just take the mean. Even really sophisticated models, like the Andrews and Kasy AER paper, can be as bad as that when you have enough p-hacking. The point is that these models were developed for publication bias, where the individual estimates are unbiased. When you have p-hacking, the individual estimates can, and many times will, be biased, which makes these models really break with p-hacking, when there is enough p-hacking.

01:02:12 Participant: Is this the proportion of p-hacked results?

01:02:14 Tomas Havranek: It is a measure, but we should have a better interpretation. It would take me a long time to explain. It is just the way we design a simulation.

01:02:31 Participant: How strong is the 0.9? I'm trying to think of how realistic it would be.

01:02:38 Tomas Havranek: I will show you some comparison with real data. That is probably the best. I agree: we should maybe choose a different scale or explain it a bit more. One more example of how it could work.

This is taken from a paper by Krueger in the QJE, again on how the size of a class, how many children are in a class, can affect education, that is, test scores. It is from the STAR experiment from Tennessee, which you might already have known, but that is not so important. It shows that Alan Krueger ran a couple of regressions with different moderators. He is always interested in the treatment effect. You see that the treatment effect is essentially the same, around five. The standard error really depends on the moderators you include. These are characteristics of the children and teachers and so on.

The point is that you could easily just publish this one here or this one. The choice is yours, but then you will have completely different precision, and the weight in classical meta-analysis will be completely different. I am not saying that this is p-hacking, but it shows you the potential for p-hacking if you want to do it. […]

We tried to replicate his data. It is an old paper, so it is difficult to fully replicate. We did different computations of variance. He does, of course, bootstrap clusters, but we were not able to get such a large standard error. This is good: it survived all of our efforts. But of course, if you are not so strict and, for example, instead of the bootstrap you do just clusters, you get much more precision.

You get a much smaller standard error. If you just do plain vanilla regression, you get about one-third of the correct standard error. So in most observational work (and this is actually experimental work), it is really crucial how you compute the standard error. That is very different from medicine, because in medical clinical trials you usually have a clear way to compute the standard error: it is just a function of sample size. But in observational research, even in some experimental research, in the social sciences and economics, the computation of precision is a key part of the empirical exercise. So if we just take the more precise estimates to be more credible in meta-analysis, that does not really capture the way we do empirical research in the social sciences.

Again, there is no p-hacking here, but there could be. We could p-hack the t-statistic to three times the original. If you want to see more details, we have a paper with Milan Scasny in the Journal of Labor Economics on this specific issue.

One more brief summary. Here is the funnel plot again, because the shape should resemble an inverted funnel. That is how Taisuke started the lecture, by showing the funnel. If there is no p-hacking and no publication bias, that is fine, everything is published. If you have publication bias, the insignificant estimates just disappear, so you have an inflated mean of the published ones.

You can still have p-hacking when people cannot change the precision but can somehow change the point estimate. In theory. In reality it is probably not like that, because it happens at the same time. But in theory, if you can do that, I call it conventional p-hacking. With such p-hacking, these funnel plot-based corrections will still be okay, because they estimate the top of the funnel, and if you run a quadratic regression, like PEESE, it is going to be perfectly fine. This type of p-hacking is okay for funnel-type correction models, but not for selection models, because selection models assume that all of these published things are individually unbiased.

But funnel-based models will probably be okay. But what if you have p-hacking also on precision, for example by changing the way you cluster? There are many ways to compute the variance-covariance matrix, and if you just play around with it, that is exactly what you are doing. Given the same estimate, you can get more precision. Sometimes you can even cross the significance boundary and have something that looks stronger, if you can justify your clustering choice given the constraints that you need.

There, as you can see, if you just look at the top of the funnel, you are in trouble, because you will have this bias. In reality, as we show in the education premium slide, p-hacking works this way: it is not north or east, but northeast, at the same time in both directions. So you will move these estimates here. But you do not know which way the bias will go. In the simulations in the Nature Communications paper, we show that most likely the bias is going to be positive. So if you just look at the top of the funnel, it is going to drive your typical meta-analysis results upwards.

01:09:26 Participant: [Asks how we can tell publication bias apart from a failure of the two assumptions behind the independence of estimates and standard errors: normally distributed data and random sampling.]

01:09:47 Tomas Havranek: I do not think you need a normal distribution. That is a good question, and we can get to it later as well. I think all you need is actually the properties of regression analysis. Maybe Chishio will correct me or will have a better answer, but I think that in the end, if you just run OLS and the assumptions of OLS are met in some general form, then what should follow is this independence between estimates and standard errors. By the way, if you run instrumental variables, that does not hold anymore. There is a nice paper by Keane and Neal, in the Journal of Econometrics, where they show that, by definition, if you use IV, you will always have this correlation.

I will try to be brief so that we end on time. With p-hacking, we have a classical endogeneity problem, a classical identification problem in economics. Take the PEESE regression, the quadratic fit. For it to work, we need the variance to be exogenous, that is, not correlated with u. But p-hacking is something in u that influences both the variance and the estimate. So it is a classical problem, and I am sorry to say that we have a classical solution: we propose simply to use sample size as an instrument for variance. Why? Because variance, by definition, is related to sample size.

It is also much more difficult to manipulate sample size than precision. In observational research, sample size is typically just given to you, and you use as much data as possible. Of course, there are exceptions, and I am sure you can give me five different reasons why this could also be manipulated. But I hope that, if you read the paper, it will at least convince you that sample size is much more likely to be exogenous than the standard error. Other options are to use only the good studies, but which ones are good? Or you can control for all observables, but I can never control for all the choices that people make when they p-hack. So we try to remove the bad variation, the variation related to p-hacking, from what we see in the standard errors. That is the instrumental variable solution.

Again, IV is not really sexy nowadays, but here I think we have a strong case for it, because the instrument is almost always strong and likely to be valid in many cases. We use the definition of variance in the first stage: first we regress variance on inverse sample size and take the fitted values. Whatever p-hacking happens is not related to sample size. It is the choices you make after you have your sample, when you choose which variables to include and so on. These choices do not affect sample size, but they will affect variance. [Note, 2026: the settings recommended at meta-analysis.cz/maive/how-to/ use a log first stage.]

The residuals are key here: we get rid of them and just take the fitted values. […] In our simulations, it works. At high, extreme values of p-hacking, you still get something substantially better than the naive mean, which is not the case for PEESE, for the selection models, or for anything else.

We are also working on something called WAIVE, where we use the residuals. In MAIVE we just get rid of the residuals, but here we use them as a weight, because the residuals from the first stage are likely to be related to p-hacking. If the residual is negative, it means that the variance is too small given the sample size, which may or may not be the case, but it could be consistent with some p-hacking. And if there is p-hacking, it likely affects both the estimate and the variance.

So both are going to be biased, and we use this information in the second stage of IV to give less weight to estimates that have a negative residual. This is still a work in progress, so we are thinking about how to improve it and maybe to apply the weight only around significant thresholds, but you get the idea. The residual is an important part, which is why we call it the weighted adjustment instrumental variable estimator. In simulations, it seems to work and to improve the behavior of MAIVE. MAIVE means meta-analysis instrumental variable estimator, and that is from the Nature Communications paper.

These are all simulations. But the problem is that, when you do not use simulations, you do not know what the truth is. You never know in observational research what the true model is or what the true value is. There are a couple of things you can do. For experimental data, you can look at multi-lab replications, where you have many laboratories running the same experiment. That is probably as close to the truth as we can get in empirical research. You can compare previous meta-analysis data and different estimators to the multi-lab result and see which ones were closer. We do this as well in the Nature Comms paper. It works, but there are only about 15 pairs of a multi-lab replication and a meta-analysis that you can use, so the sample size is very small.

What I want to show you is this slide, where Chris Doucouliagos, our co-author, has collected about 500 meta-analyses from economics, finance and accounting, but mostly economics and finance. We do a simple thing. For each meta-analysis, we compute PEESE, the quadratic fit, and then we compute MAIVE, the instrumental version of the same simple quadratic regression. What you can see is quite striking, I think. It is definitely not a coincidence. Very often, MAIVE, which is the dark one, tends to be smaller than PEESE. So the instrumental version tends to be substantially different in a systematic way.

Now, it is not always smaller than the OLS version, but it is often smaller. That is not a proof of anything, but it is highly suggestive that the instrumental treatment matters and, by extension, that the p-hacking problem is serious. […] It is also consistent with the simulations, because in the simulations you saw an upward bias from p-hacking […]. You can run this instrumental estimator on our website, EasyMeta.org. Of course, you know how to run IV in R or Stata, but on EasyMeta it computes the Anderson-Rubin confidence interval automatically.

It allows for different clustering choices, so you can do a lot of things that are possible in R but not always in Stata. Even in R, there are many ways to do it, and it is not always straightforward. That is why we created EasyMeta.org, which I might very briefly show you.

Just a couple of seconds. This is how it works. You can upload your data. There is an AI component that helps you sort the columns, that is, which column is the estimate and which is the variance, but you can do it yourself. I will briefly run a demo. We have demo data there. You can choose whether you want to use the standard weighted least squares, the Nature Comms technique, or the experimental WAIVE technique. You can compute the Anderson-Rubin interval, which is, by the way, not straightforward for the intercept, so if you want to do it in Stata, it is going to be a pain. Then you can choose what kind of weights to use and what kind of winsorization, so I may try 1%. You can use a log first stage just to improve the fit, but I will not do it now. When I run it, it takes, hopefully, a couple of seconds. [Note, 2026: the settings now recommended at meta-analysis.cz/maive/how-to/ are PET-PEESE, a log first stage, equal weights and CR2 standard errors clustered by study.]

You have the estimate with the Anderson-Rubin confidence interval, because you need to take the strength of the instrument into account. We always have an interpretation. Of course, you can use AI for interpretation, but here you have a built-in, verified interpretation from the creators of the technique. You also have a test for publication bias and p-hacking. You have a funnel plot that you can just download and put into your paper. You see the original estimates and then the estimates where we adjust precision using the instrumental fitted value. You can also export the code for the entire operation. […] You also have Stata code for this. One more thing I want to mention: there is another correction for p-hacking.

Our MAIVE is one way to do it. A completely different way was developed by Maya Mathur of Stanford. It is called RTMA, Right-Truncated Meta-Analysis. Maya assumes that if something is statistically significant, it is likely to be p-hacked in some way. We do not know how, but it is likely to be, or can be, p-hacked. If it is not significant, then it is probably not p-hacked. You can question the second assumption, but it is quite clever and simple.

What she does is this: you need to assume a distribution for the true effects, again as in a selection model, say a normal or t-distribution or something like that. Then you focus only on the insignificant estimates and use Bayesian techniques to reconstruct the latent distribution of the true effects. It is just a very simple statistical process. You could use maximum likelihood, but in this case you would need a much bigger sample, so the Bayesian technique is more feasible here. In the folder you have the code for how to do it in R.

It is really straightforward to run this. I recommend that you use it, or you can use both, MAIVE and RTMA. For MAIVE you need a strong instrument, which will often be the case because, by definition, variance is a function of sample size. But sometimes it is not going to be strong. So even if you have this Anderson-Rubin confidence interval, it is worth cross-checking it with RTMA. But RTMA has different problems. One of them is the assumption that insignificant estimates are not p-hacked, which we do not need in MAIVE. [Note, 2026: always report the first-stage F. Below 10, report MAIVE's Anderson-Rubin interval instead of the point estimate and cross-check it with RTMA; see meta-analysis.cz/maive/how-to/.]

01:24:16 Participant: It is very interesting. I think it would be interesting to compare those, because they are going to fail under different environments. For instance, with your exogeneity assumption, if you think about experimental data, especially if people have not registered their sample size, you could think that they might have some stopping rule on the sample size, and that could create failures of the exogeneity assumption. I also find this very interesting, as you are mentioning. In a sense, all of these approaches assume that what we care about is statistical significance and creating positive results, or whatever, in absolute value.

But I think we have to think about what the prior in the literature is. It could be that, if you found a lot of positive results, now your incentive to undo the literature is to say, well, there is no result. […] One of the initial slides that you showed was very interesting, where you showed that there is more variance over time in the results that people find. It could be that this also reflects some further publication bias, where, for applications, for instance, you would rather show that, in fact, the effect is so much weaker than what the literature says.

01:25:35 Tomas Havranek: It is an exciting field. You can do a lot of things. You just mentioned three ideas for different papers.

01:25:43 Participant: We already have a ton of projects.

01:25:47 Tomas Havranek: You can also combine both techniques, the instrumental approach and RTMA. We are working on something like that with Maya.

01:25:56 Participant: Do you differentiate the main estimates within the study, or not? Nowadays, I think, they try to report more robustness checks, so it could be that wide variance is not the main result.

01:26:14 Tomas Havranek: I need to move on, but very briefly: I gave this one example with class size, where we have preferred estimates and the rest. They pick the preferred, or main, estimates, which could be more than one from each paper, maybe just one, maybe more. So you can do that, and I think you should. In R, we have an example of how RTMA works. Now, the final thing is heterogeneity. This is the slide you just mentioned, increasing variance.

Heterogeneity 01:26:49

You want to explain why studies report different estimates. One reason could be a different degree of publication bias, but it is definitely not the entire story. You will have objective differences across studies. I need to speed up, so I will just skip these motivation figures. But I told you about how people can measure beauty differently, and I told you about how people can measure success, professional success, differently. It does not have to be earnings; it can be sales, research outcomes and so on. Across studies, but also within studies, you could have different subjects and different contexts. And even if you have the same data, you can use a different technique.

You also want to take into account some publication characteristics, such as the number of citations. We have already covered this, which will save me a good amount of time. For the beauty paper, as you can see, the measurement of beauty does not really matter in a systematic way. All of these ways are essentially similar in terms of the resulting beauty premium.

It is the same, but there is more variation. Again, with the measurement of success, there is no strong difference. A little is going on for sports and politicians. […]

[…] But what matters is if the study includes a control for cognitive skill. It is quite strong. You can already see from the summary statistic here, from the histogram, that if there is a cognitive skill control, you are likely to have a smaller beauty premium. The logical interpretation is that there needs to be some correlation between beauty and cognitive skill, […]

[…] One mechanism, such as teachers or even parents maybe preferring more beautiful children, which might then be reflected in better cognitive results. That could well be true. We are still talking about only four or five percentage points here, which is very small.

These results, by the way, are used in litigation in the US. If there is a car accident and I get disfigured, part of the compensation I would get is for forgone earnings based on damage to my beauty or facial symmetry. It is really applied. It is not a huge market, but it is not tiny either. […]

01:31:15 Participant: Are the results symmetric around the status quo? If you have a beauty improvement versus if you become more ugly, do you see anything?

01:31:25 Tomas Havranek: These studies do not really look at being ugly, only a little, but it is not so fun.

01:31:30 Participant: I asked because you were mentioning car crashes and things.

01:31:34 Tomas Havranek: A few of these studies do, but we do not have a big enough sample to study it in our paper. There is some evidence that the adverse premium (it is not really a premium, it is a penalty) is more serious for some contexts. But that is just my interpretation of what I have read. As for heterogeneity, what you might want to do is run subsamples. Then you can just run it for sex workers separately, and so on. So what we like to do is something cool, some Bayesian thing. We do Bayesian model averaging. As for the idea of Bayesian model averaging, I have about three minutes, but let's see where I go.

You have collected these different variables, which reflect the context in which the estimate was obtained: how it was measured, how success was measured, what kind of subjects were included, and so on. If I put all these variables in one regression, that would be a mess. You know that these variables might explain it, but you do not know the true model. So what we do here is try to consider all combinations of all these variables.

There are going to be millions and billions of regressions. To do it in a feasible way, we use a Bayesian model where we can use Markov chain Monte Carlo to go through the models. We do not go through all of them, but again, there is no time to explain the details of Bayesian econometrics. You need some priors, which are quite standard, so you do not need to invent your own priors. We also use a dilution prior, so we put less weight on models that have plenty of collinearity.

It is one way to handle this collinearity problem. It is not a perfect way. We have an example in R of how to implement it, which you can, by the way, use in any regression analysis when you have many, many variables. The colors are the regression signs, and you can see how robust the signs are. Each column is a different regression with different variables. Some models have just a few variables, but we can also put all of them into the model, so that is some millions, or thousands, of models.

The weight on the horizontal axis is proportional to parsimony and model fit. It is something like information criteria, or adjusted R-squared, that is, the model fit given its complexity. The best models are on the left. You can see that standard error is important,. […]

This will give you the most important factors related to the reported beauty premium in the literature. I think it is a nice way to summarize the results. Then you can look at the individual variables. You have these millions of regressions, so you can look at the distribution of the variables. Taisuke has much better visualizations in his paper, much more beautiful than this. For example, this one is from a different paper. This is how it looks: what you report is the posterior mean, standard deviation, and inclusion probability, which is something close to statistical significance in a Bayesian setting, let's say, for each variable. Again, you have fully annotated code in the package.

As the bottom line of the analysis, we want to use this Bayesian model averaging regression result to compute the fitted values conditional on something that is good, such as good methodology or a data set of good size. We need to be subjective here and say what is good, and you need to justify it well. What is good would be to use diff-in-diff, maybe. In this case, what would be good is to control for cognitive skills, because we see that it matters, and it is an important omitted variable.

What would be good is to have less publication bias, so you set standard error to a small value, to zero. That is one way to correct for this publication bias: looking at the intercept here. You do that. Again, you have the precise code in the zip file. You get something that, for the beauty premium, is about 0.8, and the confidence interval includes zero. So you get from 5%, the mean reported beauty premium, to essentially one. […]

That is all in Stata. Here we have examples of the same thing, best practice implied estimates for the class size JOLE paper, where we show it for different contexts. What is the implied estimate if you focus on the experiment, the STAR experiment, or if you focus on RDD, on kindergarten or primary school or Scandinavian countries, and so on? This is the table you would like to show to policymakers, and then they can decide for their own context. In summary, there are seven steps on how to do meta-analysis.

Takeaways 01:38:24

The first step is to use Google Scholar and snowballing, and to use help from AI as an independent, or quasi-independent, actor. Focus on something that can be compared, not the statistics, ideally some economic effects, such as elasticities, mean differences in experimental data, or correlations in observational data. [Note, 2026: partial correlations are a last resort, for when effects cannot be translated into a common economic metric; with them, add a robustness check on the largest subset with a comparable economic effect.] What you need is always some sort of precision, sample size, and key differences in context. Then you probably want to do something about outliers. A relatively safe choice is to winsorize at the 1% level.

You want to correct for both publication bias and p-hacking. There are many models, so you can use some of those I showed. Of course, I want to promote my own. That is not surprising. If you have weak instruments, you should definitely also run Maya Mathur's RTMA estimator. [Note, 2026: with a first-stage F below 10, report MAIVE's Anderson-Rubin interval instead of the point estimate and cross-check it with RTMA.] Do look at heterogeneity, for instance using BMA with the dilution prior. Then you will have this fancy figure I showed you. And then, as the bottom line, there should always be some computation: what is the implied number for different contexts? [Note, 2026: the three principles as of 2026 are to correct for publication bias (RoBMA), correct for p-hacking (MAIVE and Mathur's RTMA), and cluster by study (CR2 standard errors); see meta-analysis.cz/guidelines/.]

I think I almost made it on time. I am five or ten seconds beyond. Thank you very much.

Corrections

Slips of the tongue corrected in the text:

  • [morning session, 01:49:13] said "non-classical measurement error is relatively fine"; the text has "classical measurement error is relatively fine".
  • [afternoon session, 01:14:21] said "If the pi is positive"; the text has "If the residual is negative".
  • [afternoon session, 01:15:06] said "weighted average instrumental variable estimator"; the text has "weighted adjustment instrumental variable estimator".