Transcript. Lecture 3, Literature Search, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.

00:01:41 Tomas Havranek: I guess I will start. Last week we had some introduction and motivation for why to do a meta-analysis in the first place. Then we had some examples of papers which you can do.

Today we will talk about how you can choose your topic, if you want to do a meta-analysis, and then how you search for the studies which you will include. I think it's really important before you start, because it's a lot of work, it's really labor intensive, to have a good idea what the bottom line would be, I think, if people ask you why you do it.

I can see this because I'm also an editor of the Journal of Economic Surveys, so I handle many finished papers. Quite often you get papers which are solid meta-analyses, but what they do is just update a previous one. They take a meta-analysis, let's say, published seven years ago, and then do an update. Sometimes it's not completely clear what the scientific value of it is. Obviously, it is quite often better, because you have more data. But you need to be able to show, in very strong terms, if you want to publish it well, what the contribution is, the principal one. The fact that you have seven or ten more years of data is not really a sufficient contribution per se for most readers.

Choosing your topic 00:03:22

So you need something stronger. Ideally, you need one figure in the introduction which would show why you need either a new meta-analysis or the first one on the topic. I will talk a little bit about the topic selection first, and then I will go to the search strategy for the literature. I think last week I mentioned that we had these guidelines, brief guidelines, on how we think meta-analysis should or could be done. Essentially, we also write there about many of the topics I will cover today.

One of the things we stress is that when you do a meta-analysis, you should really have some underlying knowledge of the field, of the topic on which you want to write. Again, some people just want to do a meta-analysis on something. But if you don't have a relatively deep knowledge of the field, it's very hard to do it well. You don't have to be a super expert, but you need to have read about it. It should somehow align with your expertise. Or you should have someone as your co-author who would be in that role, who would be an expert on the field. Typically, my papers have three, four, five co-authors, people who know more about the topic than I do. By the way, feel free to ask me at any point or stop me.

00:05:08 Participant: [A participant disagrees, suggesting that doing a meta-analysis may itself be a way to get knowledge of the field.]

00:05:38 Tomas Havranek: In terms of learning, yes, I agree. But if you want to publish it, typically you will find out that you omitted something which is just obvious to people in the field. So there is a trade-off here. I encourage you to do a meta-analysis in any case, but be careful if your goal is to publish it eventually. You need either previous knowledge of the topic or a co-author. Which, by the way, you can recruit along the way. If you have a draft of the paper, you can ask someone, but it's better to start already with the knowledge, because that's important, especially for data collection.

Inspiration 00:06:21

As I put on the slides, it's good to have some previous familiarity with the topic. And then you need a punchline, ideally a punchline figure in the introduction. That could be, if you do a new meta-analysis, some sort of policy importance. Last week I showed these figures, from models, of what happens in a policy model in a central bank if you change the parameter, for example. So you can use these examples. Maybe I can show it again, because it should be easy.

For instance, this figure: this is the result of a structural model of the type commonly used in central banks, ministries of finance, or treasuries. It shows what the model tells you about what happens to, in this case, investment when you, as a central bank, increase the interest rate. That's what central banks are interested in. They change the interest rates, and they want to affect the economy in a certain way.

So they need to know what happens to the economy, and for this purpose they have this model. The model gives them the answer: we increase interest rates, what happens to investment? But the answer of the model depends on one parameter, in this case the elasticity of intertemporal substitution, so how people respond in terms of consumption and savings to changes in the interest rate. If you want to do a meta-analysis of this parameter, the elasticity of intertemporal substitution, it's useful to show a figure like this, which makes it clear to everyone who reads the paper that this parameter is important for policy.

That's because if you work for a central bank and you need to calibrate this model well, you need good evidence on the size of the parameter. That's one example. In my papers, you can often see impulse responses like this in the introduction. We take some model, and we show what happens in the model if you change the parameter on which we do the meta-analysis.

But of course you can have a different motivation. If you do an updated meta-analysis, so there was a previous one, let's say, 10 years ago, and you want to explain why your paper deserves to be published, why you deserve to defend your thesis or whatever, a good way is to look at the methods used in the previous one.

Quite commonly, these older meta-analyses ignore publication bias. They have some sort of summary, like fixed effects or random effects. We will get to that. But sometimes they ignore publication bias, or they just mention there could be publication bias. They don't really correct for it. So that's a clear motivation which you can use in your paper. Now, if we talk about the recent meta-analyses, from the last, let's say, five years in economics, most of them would take publication bias into account.

But quite often they ignore p-hacking, which is similar but different. I will talk about it, I think, in two weeks, because it has different solutions. So if you cannot use the publication bias motivation, you can almost always use the p-hacking motivation. You can also look at heterogeneity. Again, some papers maybe do publication bias adjustments, but they don't properly do heterogeneity, by which I mean looking at the characteristics which make different studies report different results.

Or sometimes they do it in a really sketchy manner. In a couple of cases, in my own experience, we did a meta-analysis on a topic which was done previously but not in economics. For instance, now we have a paper under revision in JPE Micro, the Journal of Political Economy Microeconomics, and it's a meta-analysis of the effect of financial incentives on performance. [Note, 2026: the paper is now published in the Journal of Political Economy Microeconomics; see meta-analysis.cz/incentives/.] We do a meta-analysis of the economics evidence, the experimental economics evidence, but you already have a couple of meta-analyses in psychology. But the papers in psychology are quite different from economics in terms of how experiments are designed. Do you have a question?

00:11:51 Participant: [A participant asks a question about p-hacking.]

00:11:59 Tomas Havranek: For p-hacking, as far as I know, there are two major estimators. One estimator, or a set of estimators, by Maya Mathur, the RTMA, I think it's called. And then you have our estimator, MAIVE. It corrects for many forms of p-hacking. And then, finally, what you can use as a motivation for your meta-analysis is if, since the previous meta-analysis was published, there has been an improvement, a substantial improvement, in the methods in the primary studies.

For example, you can have a meta-analysis from 10 years ago that only looked at correlation studies or OLS regression studies, but since then you could have an expansion of quasi-experimental techniques, IV, difference-in-differences, and so on. So you would have a clear improvement in the quality of the new studies in your meta-analysis. Now, I will show some specific examples.

I will use our website. We have this website, meta-analysis.cz, where we have examples of our own work, many different meta-analyses. Here you have links to the papers on different topics, and we also provide data and codes for these meta-analyses.

Motivation example 1 00:13:57

Example number one, how you can motivate your paper. This is from a meta-analysis on the effect of beauty on productivity. I will present this paper in a couple of weeks at a departmental seminar. There has been, surprisingly, a lot of interest in how beauty affects success or productivity in many settings. And if you look at the studies, what we do here is plot the median estimates from each study and the year of the data used in the study. What you can see is how the dispersion increases when time passes. When you look at newer studies, they have much more disagreement regarding the effect of beauty on productivity.

By the way, a small spoiler for my presentation. In the paper we show that the beauty effect essentially is quite small for most professions. There are two professions where the beauty effect is really strong, whatever you do. Can you guess which ones? There are two. [A participant answers: politics.] Okay, I expected that. We don't really have a separate entertainment category. There is one obvious answer. It's sex workers. [Note, 2026: in the published meta-analysis, the premium stays large after controlling for cognitive ability only for sex workers; see meta-analysis.cz/beauty/.]

But the rest I will leave for my seminar. There is one paper, I forget which; I saw its presentation eight years ago. I think it also finds a beauty effect in citations and productivity, which is hard to explain. Come to my seminar. I will talk more about it, because we have an explanation, I think. But if you read the literature, you will have some idea, even after a couple of days. There seems to be increasing uncertainty. Then you can plot it in a nice way, and you have a great motivation for why you need a meta-analysis.

Motivation example 2 00:16:50

A second example: you don't have more disagreement in time, but you have a clear trend. In this case, it's the alpha of hedge funds. So this is an example from finance. I would say you have a continuum from purely passive investing, when you just buy ETFs, exchange-traded funds, on the S&P 500 or whatever. Then you have the other extreme, the completely active ones, which are hedge funds. And then you have something in between, which can be mutual funds, which are typically tied to some sort of benchmark, but not 100%.

Anyway, hedge funds really are aggressive, completely active management. The common wisdom is that they can earn you more on your investment because they are less regulated. They are typically smaller, or at least they have a small number of investors. You cannot just come and invest in a hedge fund. You need to have substantial wealth. So it's for investors who are qualified. And then when you have a hedge fund, you face fewer restrictions from the regulatory authorities.

So the idea is that it allows you to make more money and deliver an alpha, a positive alpha compared to other strategies. In the paper, we collect papers on hedge fund performance. We show that the alpha was there in the 90s and before, but it seems to be decreasing in time. And in the paper we ask: is it a true trend, or is it just driven by different techniques used in these newer studies, maybe? That's what we explore in the paper. But the trend itself is quite interesting, because it seems that now it doesn't really matter if you invest in a hedge fund or an S&P 500 ETF fund.

Motivation example 3 00:19:25

Example number three could be about publication bias, or maybe not about publication bias, but just about the mean effect. And this is the effect of class size, how many students you have in a class, in primary and secondary school, and how it affects the performance of students. I think I mentioned it in the first lecture. Most people would say it's much better to have smaller classes. If I talk to parents, even me, before I knew about this empirical evidence, I would, as a parent, always prefer a school which would provide smaller classes, like 15 students in first grade or something.

But if you look at the empirical evidence, if you just read the five or ten most recent studies, there is really not much there. The effect seems to be really small, the effect of changing the size of the class on performance. The number actually means by how many percent student performance decreases when you increase class size by one student. When I have 20 students in a class and I compare it to 21 students, what is the difference in performance in terms of standard deviations? The result is that, on average, performance decreases by 0.36% of one standard deviation, which is really zero. [Note, 2026: in the published paper the mean reported effect is 0.25% of a standard deviation per additional student and the implied effect after corrections is 0.21%; see meta-analysis.cz/class/.] It's nothing.

We have an entire paper in the Journal of Labor Economics, essentially, about this figure. There's no publication bias. [Note, 2026: the published paper finds little publication bias; see meta-analysis.cz/class/.] The mean effect is pretty close to zero. And that's it. But it comes as a surprise to many people, because there are two influential papers from the 90s, by Joshua Angrist and by Alan Krueger, which use really sophisticated techniques to show that there is an effect.

What you can read in the literature, in the recent papers: you have a class, as a teacher, which had 25 students. Then it gets to 20 students. Your life is easier, but you don't really change the way you teach in many cases, on average.

[In reply to a participant's remark about teachers having more time for each child:] If you still put the same amount of time in, then yes, it should be okay. So I don't have a clear answer. It's just what I observe in other people mentioning this explanation. Also, in many cases you could have benefits of larger classes when you have more interactions between kids, and sometimes kids can help each other. It really depends much more on the context. But if you have a very small class, then I think it's plausible that you can have these adverse effects. Then you don't have so many friends, and it's harder for you to ask, maybe, for help or for explanations.

So this is quite surprising. You're right, we don't have a clear, super clear explanation. After we had just started to read a little bit on the topic, what was really striking to me was that I read the five most recent papers, published well, and they all essentially agree there is no effect. But then you read policy: there's a recent change in class size policy in New York, and the policy recommendation is justified by empirical evidence. It says studies show that smaller classes help students. And it's not there. So for some reason there is this perception, because it's so intuitive, as you said: you have more time for the kids, so this should be better. But it's not so much better.

It's a little bit better, maybe, on average, but not so much. But the point is, some of these things you can really know in advance, for example, that it's around zero or that you have a decrease in time. You see it after reading a couple of papers or from the dispersion.

Motivation example 4 00:24:32

Another example is that quite often you have one parameter but really different results for different countries. That's what you mentioned, I think, last week.

That's really true. So again, this is the elasticity of intertemporal substitution. For different countries you have really different results, and it's hard to find any clear pattern, such as a larger elasticity for richer countries and a smaller one for poorer countries. [Note, 2026: after controlling for study design, the published paper finds that the elasticity is larger in richer countries and in countries with higher stock market participation; see meta-analysis.cz/substitution/.] So it's not obvious. And again, we have a paper in the Journal of International Economics essentially based on this figure. We try to explain why the parameter varies across countries.

Motivation example 5 00:25:35

A classical example: if you have publication bias, it's always an easy way to explain. We will have a separate lecture on publication bias, so I will talk about it in much more detail. But for instance, if your topic is some sort of price elasticity, it's very hard to publish price elasticities which would be positive, which would imply that if you increase the price of, for example, gasoline, you would have more demand for gasoline.

People expect it to be negative, even though sometimes it could be positive. And of course, we know that if you have a model which is imprecise and has some noise, you can easily have positive estimates even though the true effect is most likely negative. That's a great motivation to say: we have this previous meta-analysis, but they ignore publication bias, which means that, most likely, the results they have are exaggerated, and we correct for publication bias.

Motivation example 6 00:26:46

Final example before we move to literature search: in some cases you have a clear difference in the results of studies depending on methodology. For instance, this is from our paper on capital-labor substitution. Do you know the Cobb-Douglas production function? In the Cobb-Douglas production function you set this elasticity of substitution between capital and labor to one. But if it's not one, then you cannot use the Cobb-Douglas production function, and you need to use a more complex constant elasticity of substitution framework, which is a little bit more demanding computationally if you build a structural model. Anyway, you have hundreds of studies which estimate the elasticity, and it's quite commonly known that if you use the first order condition for capital as your main identification device, you get much smaller estimates.

But if you use labor, you get much larger ones. Why? That's another good motivation for a meta-analysis. You have a clear heterogeneity, which is really crisp, and you want to explain: does it still hold if I control for the other data characteristics and so on? What could be the reason for the difference? So that was motivation. I think before you start to collect the data, which really takes a lot of time, a lot of energy, you need to have a clear idea how you will sell the paper.

Because, of course, you don't know anything about the results you will get. But you should have this one figure in the introduction with motivation. You should have some idea about how it could look, how you would explain to people if they ask you, "So why do you do it? We have this old meta-analysis from 2015. Why do you do a new one?"

00:29:07 Participant: If you're the first, you're fine.

00:29:09 Tomas Havranek: If you are the first one, you just need to explain why, because if you have a first meta-analysis on a topic, typically it means the topic is not super prominent, otherwise someone else would have done it.

[A short exchange about the importance of a participant's topic.]

But of course, it could happen to you that someone else is also working on it. So what I do in such cases: we did this capital-labor substitution meta-analysis, which is a huge topic, and I was concerned whether someone else could be working on it. What we did, before pre-registration really started, was just to put it on my website: we are working on it. Which I think is...

00:31:16 Participant: You claim the field.

00:31:17 Tomas Havranek: Or you should pre-register officially.

00:31:32 Participant: But even if you pre-register, it's...

00:31:35 Tomas Havranek: No, it doesn't guarantee. You never have a guarantee. So yes, but you are the first, right? As far as you know. What I mean is: make it public that you are working on it, or, even before that, just put up a pre-registration. Because before people start collecting data on a meta-analysis, they probably Google, "Is there a meta-analysis on this topic?" And they will find your pre-registration.

So that's what I would do. I would do it soon. It doesn't guarantee that no one will do it. I think that's important. I should probably do it. I do it sometimes, but I should do it more, especially...

00:33:14 Participant:...if you try to be the first.

00:33:16 Tomas Havranek: Yes. Quite often I'm not the first; there is a meta-analysis published 15 years ago. Of course, if you already have some preliminary results, it's great to put them online as soon as possible.

Google Scholar 00:33:47

I use Google Scholar. Of course, you might have other databases which you like, but the advantage of Google Scholar is really that it is inclusive. All papers, published and unpublished, are contained in Google Scholar. Thing number one, which is really important, is inclusiveness. And second, it doesn't just use the title, keywords and abstract when you search. It goes through the full text of the studies. So most of the other databases used in meta-analysis, not all but most, will just search for the title and keywords.

00:34:31 Participant: So if this is true, then...

00:34:37 Tomas Havranek: I will show you specific examples of how this could be done. So these are advantages. Disadvantages: it changes every day, so it's hard to replicate. But what can you do? You have your search, and you can save it. You can save the keywords that you used, the hits that you got from Google Scholar, and then you can put it in your online appendix. And I will show you examples. When you search in Google Scholar, it's never perfectly replicable. But all the other alternatives are worse.

Search steps 00:35:30

When I start a new meta-analysis, I first try to find what I believe are the most important empirical studies. Typically 5 to 10, you will know about them. If you start reading, you will know what the most prominent ones are. And then, if you do the search, again I will show you examples, I try to calibrate it in a way that it really finds those prominent studies. Because if I have a search and I know there is this one AER paper which I really need to include, and the AER paper is not anywhere among the hits, I know I need to change the way I search for the papers.

That's step number one and two. Then, because Google Scholar will give you hundreds and hundreds of hits, it's infeasible to go through all of them. In most cases, you select a number where you stop. So, for example, 500 or 200, that's up to you. I use 500, that's my habit. Then I read the title and abstract of each hit. And for most hits, you can see that it doesn't make any sense, it doesn't belong here.

It's a theoretical study or it's a completely different topic, so you just throw them away. And you are left with a small subsample of these 500 or whatever hits, where you see from the abstract that there could be some empirical estimates of the effect of ChatGPT on productivity, for example. So then I download these studies, which out of 500 could be maybe 100 or slightly more, and I skim the PDFs. Again, if you see that this is something else, you just drop it, but then you will see which papers you can actually use on the first try.

And then what you can also do, I think it's a good idea, is to do the search again just for the last couple of years, because you can have really new studies which don't show up among the first couple of hundred hits in Google Scholar. They are not so much cited because they are new. The Google Scholar ranking is highly related to citations, so the most cited papers will be on the top. If you have a new paper which is perfectly useful for you and can be included, it could be quite low in the ranking. So it's good to also do the same search, but just limit it in Google Scholar, let's say, since 2022 or something. So that's the first round. Let me open the website to show you a more concrete example.

Search query 1 00:38:38

It's from the paper on intertemporal substitution. There's a link to the examined studies, and there's the search. What we do is put in quotation marks "elasticity of intertemporal substitution", because that's the name of the parameter. Any paper which estimates it will use the name of the parameter in this form, or, because sometimes people would switch the word, they could also call it "intertemporal elasticity of substitution". So we say this or this different version of the definition.

And then, just to increase the relevance, we don't use quotation marks anymore. We don't impose that it needs to be there. But we also put, again, elasticity, substitution, consumption, interest rates, Euler equation, estimate, regression, empirical. These help you to put the more relevant studies to the front line, but they don't exclude studies which don't have this wording explicitly. And then we go through the first couple of hundred of these hits the way I described. That's one example.

I think the search terms are implicitly combined with AND. The way I think Google Scholar operates, if you put quotation marks, it needs to be there in this form. That's why we have this definition or the second form of the definition. And when I put these words without quotation marks, it can also find synonyms or similar words. So it doesn't have to be the same thing. I want the paper to have something about consumption, something about interest rates, because these are the variables which you regress on each other, something about regression, something about the Euler equation.

Again, Google Scholar's ranking is a bit fuzzy. So you have citations, which are given some weight, and so on, the prominence of the journal. I will first show the other example, of capital-labor elasticity. Capital-labor is here.

Search query 2 00:42:16

Again, we have examined studies. It's the elasticity of substitution between capital and labor. But it's too long, and people use different permutations of the definition. But we definitely want to have capital in the paper. We want to have labor, in US or British spelling, or New Zealand as well. You use the British. So labor, labour. You want to have substitution there, and you need to have elasticity there. And then you also want to have either constant elasticity of substitution or CES, because that's the estimation framework that people use.

And then again, you don't have to, but it's nice if there is something about production function, first order, estimate, regression and so on. Cobb-Douglas, labor, capital, and so on. So this is the result of calibration we did. We tried to change the search in a way that the studies we knew were among the first hits, let's say.

00:43:55 Participant: [A participant asks about the search terms.]

00:43:55 Tomas Havranek: Again, I may be wrong, but the way I understand how Google Scholar works, if you put a term in quotation marks, it's very strict: it must appear in the same form which you have between the quotation marks. I need the paper to have labor, but I don't care about the British or the American spelling. If you don't put a term in quotation marks, Google will search for either the word, for example regression, or synonyms, or something similar. It could be regress, or it could be identification, it could be something which is not exactly in the same form.

Primary studies 00:45:31

And then we have a set of studies which we can include. So you have a list of studies which, after skimming through the PDFs, you think you can include. It's just a list of papers. But I will go to the next stage. This was the beginning.

Snowballing 00:46:25

But as we discussed, it's a bit fuzzy. So you are not sure if you really have all of the relevant studies, because maybe, even though you do your best, the search could not be perfect in terms of capturing these empirical studies.

In our newer meta-analyses, what we try to do is to have a completely different approach in the second stage. It is called snowballing. First, you try to identify your empirical papers using a Google Scholar search. You have this list of papers, maybe 50, maybe more, maybe fewer. But no search using keywords is perfect, because some studies, even with a full-text search, can use different terminology and so on. Or for some reason, you will just not have them there. So what we do, for each study that we have in the data set and previously identified, is collect the references, the list of papers quoted in this primary study.

Suppose you have 50 studies from the previous stage, from Google Scholar. For each study, you download the references. You can do it either from Google Scholar, or you can use something like Scopus or Web of Science. It is easy. You just click, download the references, and have an Excel file or whatever. Then we put it all together and try to see which are the most commonly quoted studies in this literature.

[A participant asks a question that is hard to hear.] Of course, if you have a narrative survey on the same topic, that could be a good starting point before you do all of these steps, when you locate the most important papers in the first stage. What we do here is something similar, but we do it for every single paper, automatically: we just download the citations and then see which papers are really commonly cited in the literature. If you have a paper that is cited about 20 times in your primary studies, it is probably important. Quite often it could be a theoretical study, but very often you find at least a couple of new ones that were not identified in the Google Scholar search.

[A participant asks about collecting the data.] If it is just me and I have about 120 studies, I cannot collect the data myself. But when we have more people, you can spread it. For the capital-labor paper, we had about 120 studies, and I did not collect the data myself; I did a small portion.

But when you collect it, I think that when you are sure that you understand what you are coding, you do not read it from title to conclusion. But you need to read at least a couple of papers. Sometimes it is so unclear that you have to read a lot of it to have a good idea of what you are actually collecting from this paper.

To summarize: first, we search in Google Scholar using some keywords, as we showed before, and you have some examples online on our website. Then, in the second stage, we take the studies, collect them, and look at the references to see if we missed an important study. This is called snowballing. It complements the previous stage in a completely independent manner. Or not completely, but it is a very different procedure.

Snowballing example 00:51:37

This is an example of how you can do it. Obviously, if you do it by hand, it would take you forever. But if you go to, for example, Scopus and you find a paper, you have a list of references which you can easily export to XLS, for example. If you have 50 studies, you will have 50 XLS files, but it is relatively easy to put them into one database. You can be done with a snowballing exercise in maybe one hour, and then you just go through the most prominent studies, the ones that are quoted most often.

I do not know about the other databases, but Scopus works well for snowballing because of this export function. Of course, what we are talking about here today is very important for meta-analysis. But in the actual paper, you will spend a couple of sentences on how you collected the data, so people do not really expect you to write a lot about it in the paper.

For example, if you want to do snowballing just with published papers, because the published ones will be in Scopus, you can use just the published ones for snowballing. I think that is completely fine, because if they are unpublished, you do not have this option to download references, and it could be more difficult. You can use Scholar, but I think for Google Scholar you would need to write your own code to scrape it from the web. Here, you can just point and click. Publish or Perish: is it automatic? [A participant answers.] Amazing, that is even better. Perfect. So now you have a solution.

00:53:56 Participant: [A participant asks a question that is hard to hear, apparently whether he uses other databases.]

00:54:22 Tomas Havranek: No, I do not, because this is cleaner for me. I have one search, even though it has the problem that we discussed: it is a bit fuzzy, it is not 100%. You need to stop somewhere, because it will give you thousands of hits. Publish or Perish, I think, is based on Google Scholar, isn't it? It is the same data.

00:54:43 Participant: [A participant replies, but it is hard to hear.]

00:54:55 Tomas Havranek: I do not have a definitive answer why you should just use Google Scholar. I think there is no way to justify it completely, apart from that it is cleaner. And if you do this and you do snowballing, you can be reasonably sure that your coverage is...

00:55:14 Participant: [A participant adds that one then does not have to explain why RePEc or EBSCO were not used.]

00:55:23 Tomas Havranek: No, it never happens to me, because I think RePEc is a subset of Google Scholar. Have you seen a paper that is indexed in RePEc but not in Google Scholar?

00:55:38 Participant: [A participant replies that there are differences.]

00:56:08 Tomas Havranek: Google Scholar is quite fast compared to RePEc. But I am surprised. Do you mean that there are papers that are in RePEc but not in Google Scholar? That is surprising. My experience is that even when I put an online appendix on my website, it shows up in Google Scholar after a couple of days as a new paper. Maybe there are some websites that are not explored by the bots from Google, so it could happen. You see, there are these caveats, but still, if you do the snowballing, if you use more databases, no one will complain.

It will never be perfect here. When you do a meta-analysis, there is an element of art involved. It is never 100%, because you don't have a closed solution. You can never say, "Now I have all of the studies." Or it could be just in a local journal somewhere, and it is not online. But I think this is not so difficult, it is quite inclusive, and you have these two steps, and it is quite elegant: just one database.

Minimum requirements for meta-analysis 00:58:44

Now, people often ask, especially when you do your MA thesis or PhD thesis, what the minimum requirements are for data collection. How many studies do I need? Do I need to collect 200 studies if I am just working on my bachelor's thesis? The guidelines we typically provide, also in our paper with Tom Stanley and others, are that if you want to use these techniques for publication bias and p-hacking correction that we were talking about, you need at least 30 estimates. And 30 is the common number.

00:59:45 Participant: [A participant asks whether this number is based on simulations and makes a remark that is hard to hear.]

00:59:54 Tomas Havranek: So let me ask you a question: how much data do you need to run a regression, a simple regression with one variable? 30. That is the common number. Some people would say 50 or 40, but 30 is the minimum. Definitely, you need 30. Sometimes, when the question is so important and you have just a few observations, it still gives you some idea, and at least you can take it quasi-seriously, a little bit. Of course, you want to have more, but at least you should have 30 observations to do some sort of regression. It is the same question as how much data you need for a regression.

But this course is about meta-analysis in economics and finance, and in economics and finance, we want to do regression, and we always want to correct for publication bias. If you have fewer than 30 observations, you cannot really do it. It is essentially impossible. If you want to use these more sophisticated models, you will need more. But if you have 30, you can do at least something. So that is the minimum. How many studies? I say at least 10. But again, if you have fewer than 10 studies, then you can do some weighted average.

It is not just about the number of estimates but about the number of studies, because the number of studies typically tells you how many independent estimates you have. You can have many estimates in one study, sometimes from different treatments or different data sets, but they will be to some extent correlated with each other within the study. So you cannot have 30 estimates from two studies. That is no good. I would say at least 10 studies and altogether at least 30 estimates. Now, at minimum, what you have to collect is the effects you are interested in.

For example, the effects of performance fees on performance, measured as partial correlations or as the actual effect, like in percentage terms, whatever the literature allows you to include. Then you need standard errors: how precise these estimates are. If you don't have standard errors, well, of course, you can compute them if you have t-statistics, because it is a simple formula. Or you can compute them from p-values if you have just p-values reported. But you need them. You need some measure of uncertainty around these estimates. If you do not have it, you cannot even do the weighted averages, because you don't have any weights in your model.

You also need the sample size or degrees of freedom, that is, how many observations are in the primary study. The reason is that when you have the sample size, you can adjust your weights, and you can apply the p-hacking correction techniques like MAIVE, so you can solve problems in your measure of standard errors. [Note, 2026: MAIVE uses the total sample size behind each estimate, not the degrees of freedom or another effective sample size. See meta-analysis.cz/maive/how-to/.] That is basically it.

01:04:04 Participant: [A participant asks whether the sample size is enough.]

01:04:12 Tomas Havranek: You are exactly right. It is a big mess. Sometimes it is almost impossible to compute, because, for example, people do a regression and they do not tell you. They have fixed effects or controls for industry dummies or whatever. Sometimes you have regressions in the primary studies, and it is not completely clear what the precise sample size used in this specification was, for the reasons you mentioned. For example, in one specification they remove some outliers. It should be there, but sometimes the number is not there. But still, you will have a pretty good idea about the sample size, because they always write roughly how large the data set is, and even if they do smaller specifications with small changes, it is not like one-half of the data set or one-tenth.

So in these cases, you can reasonably impute. Also, when you collect data on methodology, sometimes it is not 100% clear. Is this author using maximum likelihood or not? But in most cases, you can make a good guess, which is justified.

01:06:06 Participant: [A participant asks whether one typically collects the sample size and forgets about the degrees of freedom.]

01:06:13 Tomas Havranek: Yes, because it is impossible.

01:06:33 Participant: [A participant agrees and adds that it takes a good knowledge of economics.]

01:06:38 Tomas Havranek: Of course, Bob is right that it would be neater in theory. The problem is that it is quite often not reported. You need to do your imputation. It is a mess. It is better to use the sample size.

But you always find a way. For instance, the referee says, "You should use degrees of freedom, not sample size." I have about 5,000 estimates from 200 studies. So what I can do is collect it from 20 studies, and I will show you what the difference is. I can show you that there is a very tiny difference. It is not statistically significant. If you want, I will do it, of course.

This is the minimum that you need. So you can do a small meta-analysis, for example, for a project in this class, if you want. It doesn't take you months of your life, but it can be done relatively easily. Or you can take existing data and add a couple of new studies to them. If you do a thesis, for example, it does not have to be a full-fledged new meta-analysis.

What we are discussing here is mostly a research paper, so the best practices. But of course, in other settings, for example when you work for a central bank, your boss will ask you, "So what is the effect of interest rates on consumption? You are a scientist, what does the science say about this?" And you have one day. So what can you do? Well, you will look at some recent meta-analyses, and you can collect a little bit of your own to update it maybe and to give your boss a qualified guess.

But it will still be a meta-analysis even though it is small scale. It happened to me many times when I was working for a central bank. There is a question, and they know you are an advisor, you are a researcher: "Tomorrow we have a meeting. We want to know what the science says about this question, so give me an answer." And you need to do some sort of meta-analysis quickly.

Quality 01:10:02

Referees typically do not ask about degrees of freedom, and they do not care which database you use. What they do care about is study quality. Especially if I want to publish my paper in a prestigious economics journal, such as the Journal of Labor Economics, REStat or the AER, at least one referee will ask: why do you include these studies from local journals?

The referee says, "I just trust top-quality peer review, so only top field journals and above." Of course, I do not like that approach. What is the difference? How much does it matter? That is also an interesting question. But typically at least one referee will say, "I just care about quasi-experimental results, which means no OLS. Or it can be OLS but using experimental data. And I just want you to look at top journals, like the top five, maybe also REStat, the top ten journals in economics. That is all I care about."

I would not recommend focusing just on the best, because then you lose a lot of information: you lose plenty of other studies, which increase your power even if they are poor. It actually helps you to identify the effect of quality on results. So it is always better to include as many studies as you can, if you have the time and you are doing a research paper.

Then you can try to code for quality objectively. What you do is try to pin down characteristics of these studies: what methods are used. Is it quasi-experimental or not? Is it OLS? Is it IV? Is it difference-in-differences? Something else. How do they treat outliers? There are a number of things that you can code for. Then, if the referee insists, what you can do as a robustness check is, in your Table 23 or Table 5, whatever, just after the main results, get rid of these lower-quality journals.

And just try to see what happens if you focus on the top journals. That works perfectly. Sometimes there is a difference, so it might be useful to talk about it and put more weight on the top results. Sometimes it is relatively the same.

Sometimes you read about things like risk of bias. In some fields they have a standardized checklist of what a primary study needs to have to be considered high quality. You see it used in fields where, from an economist's point of view, it does not fit: you have these guidelines for medicine, and then they use it somewhere in social science.

So I think it is better to code for study quality in terms of the methods used in the study. It is essentially the same thing as the risk-of-bias classification, which you often see in medical journals, but the way we do it in economics is more detailed, I think. If you want to call it risk of bias, you can call it risk of bias. You can use the dummy variables that you collect, and we will talk about that. You collect a dummy for IV, OLS and so on, and you can use them to construct a risk-of-bias category, essentially.

Unpublished papers 01:14:57

Issue number one: quality. Referees ask about quality. Referees also ask about unpublished papers. In some of my studies, some of my meta-analyses, if the literature is too large, even if I have around four co-authors, we decide to ignore unpublished papers. That is not completely a good thing. You would like to collect everything. The problem is that you have some limitations in terms of how many person-hours you have on your team. What can you feasibly collect?

Second, in unpublished papers it is much more common to have typos. They are hard to read sometimes, and there are more mistakes on average. So that is our justification for why sometimes, not always, we ignore unpublished papers. Now, if referees ask about unpublished papers, which they quite often do, and I have ignored them, they almost always seem to think that you would solve publication bias if you just included unpublished papers.

Unfortunately, it does not work like this. When I get this comment from a referee, I say, okay, let's take a look. We collect data from, let's say, 10 working papers, and we try to run the publication bias models, the same models that we have in the main paper, also on the working papers, if possible, and we see there is no difference. Or maybe there is a difference, and then you have to explain why. But typically it is very similar. Why? I am asking you.

01:17:05 Participant: [A participant answers.]

01:17:05 Tomas Havranek: Well, almost. But if you have a working paper, you want to publish it, so it is not as if it were in your file drawer somewhere. There is a paper by Abel Brodeur on this. He has data from one journal, I think the Journal of Human Resources, with thousands of submissions, and he compares the first submission and the final version. I do not think he finds much difference, but it is a nice paper, published in the AER, by the way.

So if you have the data, the referees are probably happier. You can have a dummy for published and unpublished. What you can also do, and I think that is a nice exercise, is to do the publication bias analysis separately for published and unpublished studies. My point is actually yours: most working papers will be published anyway. It is just a matter of time.

01:18:21 Participant: It's just a matter of which journal ranking.

01:18:25 Tomas Havranek: I think a better way to do it, if you want to compare published and unpublished studies in economics and finance, is to look at working papers which are really old, like 10 years old. It is unlikely that a 10-year-old paper is still unpublished. So probably there is something about the paper which made it unpublishable. Maybe it was publication bias; maybe the authors were principled: no p-hacking, no publication bias, we will just report what we got.

01:19:26 Participant: [A participant asks a question about working at a treasury or central bank.]

[A participant comments on whether publication counts.]

01:19:54 Tomas Havranek: It is the same in central banks, typically. The working paper is the goal. There is the policy work, or at least a policy component, and that is it. But that is a good point. Quite often you will have these old working papers from treasuries, reserve banks and central banks, which you could then use to distinguish between published and unpublished studies.

But again, in economics, when we write a paper, typically, even if you do not want to publish it but want your boss to like it, you want to have significant results there. In academia, you sometimes can: we have this paper on class size, all zeros. It can be published. […]

01:22:51 Tomas Havranek: By the way, my favorite book on the causes of economic growth is by Deirdre McCloskey. Do you know Deirdre McCloskey, the economic historian? I think she is totally brilliant. She has a trilogy on the origins of the industrial revolution, which is really transformative. The main idea is that the main driver of growth is not savings, not investment, not even institutions, and not natural resources, but the way people treat innovations. For almost the entire history of the human race, innovations were not treated as a good thing.

You would have tradition, and you would not like to mess with tradition, so you discouraged innovation. We have liked innovation only since 1800 or something, on average. A little bit earlier in the Netherlands and England.

So anyway, I like these books by Deirdre McCloskey. She is also one of the best-writing economists of all time. I think there was a survey some 20 years ago about who is the best writer in economics, and she won first prize. She has a paper on how to write, Economical Writing, I think it is called. She is really good.

If you want to read a bit about writing in economics, about writing research, Deirdre McCloskey is the first reference. The second reference is John Cochrane. He has a nice brief booklet, Writing Tips for PhD Students, which is also really helpful, really good.

So, this was a brief digression. If you still want to do a meta-analysis but there are too many papers for you to collect, what can you do? In my case, I typically try to find co-authors, PhD students. We now have a paper on financial incentives and performance, for which we have a revision at JPE Micro, and we have around six co-authors. At revision they ask us to collect more data, and we added another co-author to help us, because it is a lot of work.

But if you do not have the luxury of extra co-authors, as a last resort you can select a random subset, a subsample of the data. That is the cleanest solution if you document it: this is my full sample, and this is how I randomly select from it. So that is also an option. But I still think that, if you want to publish the paper, that is not a good option. In that case it will be better to really focus on better journals, maybe. Just say, "Because of the breadth of the literature, I focus on the top 20 journals in economics or something," which would be easier for referees to understand and to actually appreciate.

This is the last slide, and then we can talk some more about the details. You probably know this figure. It is called the PRISMA diagram, in which you report how you collected your data. And it is very simple. It does not have to be complicated. You have a brief description of your search in, for instance, Google Scholar, the way I do it, and the snowballing. Then you say, okay, you got these, in this case 1500 hits, and in this paper we screened a lot of them.

I do not know why, but in this paper we did more, I think because it was a house price paper. We skimmed a lot of papers. Most of them we just excluded because it was obvious they are not empirical. Then we were left with more than 300, which we tried to download. But again, almost all of them we could not use. So we were left with just 37 papers in total. And that is it. That is all you need for a PRISMA diagram. Essentially, you have this information because you collect the data, so you have notes to be able to construct it at the end. We still have some time, if you want to ask any questions or anything related to essentially anything.

01:30:33 Participant: [A student demonstrates a keyword-based literature search workflow in Scopus.]

01:32:36 Tomas Havranek: I see. That's useful. Even if you want to do your primary search in Google Scholar, you can use this for identifying keywords. So that's pretty useful, I think.

01:32:54 Participant: We keep that and put it on Google Scholar.

01:32:56 Tomas Havranek: Or you can just ask ChatGPT what would be good keywords.

01:33:01 Participant: [The student continues the demonstration.]

01:34:29 Tomas Havranek: Thank you. By the way, people often ask me, when I present at conferences, "So you do all this labor-intensive data collection by hand? What about AI?"

To be honest, I think that in the medium to long term there is nothing really which would prevent a solid AI from doing it on its own. Right now it's not even close. But in a couple of years I'm pretty sure we will have it as a co-author. It will collect, and you will just check it, maybe randomly, if you see a problem there. In the medium and long run, I think, it should essentially be able to collect most of the data itself. Or what you could do: you would have 200 papers, so you code five papers to help the AI do the rest.

Now, when you look at the outputs of ChatGPT, in some ways they are better than, for example, what I could produce on some fronts, I believe. This is some sort of mechanical task which would really benefit a lot, and it's feasible that AI could do it, and then we could focus on the creative things around it, such as what topic to choose. Even there, ChatGPT will help you. But you would have exponentially more time to focus on these important issues which are not laborious, but where you really need to be creative. And then, when at some point ChatGPT or any type of language model can actually be more creative than you are in research, then we have a big problem. Then we can retire.

Corrections

Slips of the tongue corrected in the text:

  • [00:10:57] said "JPE"; the text has "JPE Micro".
  • [00:19:25] said "example number two"; the text has "example number three".
  • [00:42:16] said "intertemporal substitution"; the text has "elasticity of substitution".
  • [00:49:48] said "about 150 studies"; the text has "about 120 studies".