Transcript. Lecture 1, Introduction to Meta-Analysis, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.
- Lecture 1. Introduction to Meta-Analysis 00:00:00
- Meta important for research, policy, practice 00:02:20
- What happens if the government spends more? 00:20:14
- Meta in monetary policy making: DSGE models 00:25:31
- Dozens or hundreds of equations 00:30:15
- Psychohistory accomplished? Civ VII? 00:31:49
- Main output: fan charts 00:34:44
- Calibration uncertainty: meta needed! 00:36:59
- Estimates of the labor supply elasticity vary 00:54:21
- Values around 0.4 to 0.5 used for calibration (CBO) 01:00:23
- Bias: some estimates less likely to be published 01:01:10
- Context: different situations yield different estimates 01:13:30
- Course 01:21:48
- Guidelines 01:27:47
- Resources 01:29:21
Lecture 1. Introduction to Meta-Analysis 00:00:00
00:00:06 Tomas Havranek: Good afternoon. […] I am happy to be here. Can I just ask you, because I know some of you are not from economics or finance: is there anyone who does not do economics or finance? I already know of one student. So it is going to be mostly about economics and finance, but I will try not to go into many technical details from macroeconomics.
What is it going to be about? I will talk about ways to put together empirical research on the same topic or a pretty similar topic. That is what we call research synthesis. I will mostly say meta-analysis. Meta-analysis is the method to do it: the different statistical techniques that are used to do research synthesis. It is also easier for me to pronounce. Now, why do we care?
Meta important for research, policy, practice 00:02:20
Why do we do it? First of all, if you do any kind of research, typically you need to start somewhere. What is the current stock of knowledge? You either read a survey on the topic, or you read an influential paper, or you read a document. But you always need some sort of mechanism through which you take stock of the literature. For instance, take a study of the effect of performance fees on the performance of mutual funds in New Zealand and Australia. When you do something like this, specifically for New Zealand or Australia, it might be useful to know what the general picture in the literature is, what the typical finding is elsewhere, maybe for Europe or the US.
Today, as a motivation, I will mostly talk about policy, how meta-analysis can be useful for policy, and specifically in the central bank. We will try to make it accessible. For instance, you work in a central bank and the question is whether to raise interest rates or not. Your boss, the governor or the board, may ask you, if you are an analyst in a central bank: what kind of results can we expect from this policy? You will have to go into research. You will have to try to take stock of it. You might want to either look at the analysis or do some quick summary.
The course is quite long, 12 lectures, so we can go into quite a lot of detail in terms of how we can do it. But I will also stress along the way, if you are a policy analyst and have just one day, how you can make a reasonable survey in one day, for example. What are the shortcuts you can take if your boss tells you, "I need it in the morning"? I will try to also make it relevant for practical use. Or, for instance, if you work in a hedge fund and you think about how, as a team, you can change your strategy in your investment portfolio, you may be interested in what is called, in finance, pricing factors.
For instance, there are plenty of papers that show that if you buy stock in small companies, they tend to have a bigger return, typically, because they are more risky, and so on. But if you do a practical policy change based on a strategy, you want to know how precise it is, what the precise effect is, how it relates to risk, and so on. Central banks, hedge funds, any policy institutions: at some point you will need to do something like a meta-analysis, or you will need to read a meta-analysis or a survey that maybe could be better done as a meta-analysis.
It is important first to know how it is done and, second, how to know which survey is a reasonable one, because in many cases you have many on the same topic. It also matters in practical life. For instance, if you like to drink wine, is it good for your health to drink a little bit of wine, like one glass per day? Because when I was younger, you could read in the newspapers, and even prominent surgeons would tell you, "Oh, one or two glasses, it actually helps you." Do you know what the current state of knowledge is?
For many years you would have these studies, which would show that drinking a little bit is better than not drinking at all. In terms of life expectancy, you would have a little bit less life expectancy if you don't drink at all, it would be more if you drink one or two drinks per day, and it would go down if you drink a lot, which is obvious.
This was the basis for the recommendation that it is good to drink a little bit. But the problem was that these were not experiments, because it is hard to randomly select people. Typically, what we have in the data are what we call observational studies, which means they just ask people, "Do you drink? How much do you drink?" And then they observe your health outcomes, how healthy you are, and so on.
The issue is that if you just ask people, even if they tell you the truth, quite often it is not the choice of those who don't drink at all. They don't drink because they cannot, for some other health reasons. Quite often researchers are just looking at a correlation, just the graph of life expectancy. So you would have dozens of studies that would give you a very flawed policy recommendation, a real-life recommendation.
It turns out that if you take this into account, so you control in these studies also for other health issues, there is no hump-shaped life expectancy function. We are talking about how much alcohol is good for your health. It turns out that no amount of alcohol is good for you. If you look at the most recent meta-analysis on this issue, it is not just a summary of previous studies: it gives more weight to the studies that actually control for these omitted variables, as we call them in economics and finance.
By the way, if you have any questions at any point, just speak or just raise your hand whenever you want. We have plenty of time, so I think we can do it really slowly and interactively, because I am not a native speaker and may mispronounce or confuse things.
That is the basic motivation, and why you should be here. It is interesting for research, it is important for policy work, in case you want to be in a public policy or finance institution. And sometimes it is also useful to know for your own or personal purposes, like these little questions, and you can take a look at meta-analyses. Hopefully you will be able to tell which one is a good one, which one you can rely on.
Now, why do we need another one? If you look online, there are many courses on meta-analysis, but almost all of them are on medical research. There is a basic difference, because meta-analysis originally developed in medical research as a way to increase statistical power. What I mean is, when you do an experiment, a proper experiment in medicine means you have randomly assigned groups: you have a treatment group, and you have a control group.
For example, you have a new drug. The treatment group gets the new drug, and the control group gets a placebo, which means a sugar pill. People think they are actually taking a great new drug, but they are not. The placebo effect would be quite important, which is why the control group gets the placebo. And then you just compare the means: how much the treatment group improves compared to the control group. The difference is your experimental estimate. So when you can do it, it is great, because if the selection into these two groups is random, your estimate is purely causal. There is no confounding and no other omitted variable problems, typically, if you do it well.
Meta-analysis was originally used to take a couple of experiments like this, RCTs, randomized controlled trials, on the same topic, for the same drug, and put them together, just to increase the precision, so the statistical power, of the final outcome, to get tighter confidence intervals. So in medicine, people would use meta-analysis to get more precise estimates.
00:12:43 Participant: And in that situation, the number of studies you have doesn't really matter. It works from two studies or three studies.
00:12:49 Tomas Havranek: Yes, the point that you increase statistical power does work. Then, when you have techniques for publication bias correction (we will get to it, and I will explain what it is), you typically need many more observations.
But the primary objective is really to get more precise estimates. So that is medicine, experimental research in medicine, because, as we have shown in the example of alcohol consumption, not all research in medicine is experimental; actually, most of it is observational. So even in medicine, you do experiments with new drugs, but for most questions, like how much it hurts you to eat red meat, we have studies, but these are not experimental studies. We have maybe one or two experimental studies. But it is very hard to force people to eat red meat and then actually know that they are eating it, or to force people to eat just vegetables:
So it is very difficult to do RCTs. So you have, again, observational data, and you have the same problems we have in economics. In economics, again, sometimes we have experiments, but most commonly we work with regression results. So you try to estimate the effect of how many children you have in a class on the outcomes of the children. Many people would say, well, it is better to have a small class because the teacher can pay more attention, and then they have better education results, and so on. So what do you do in economics? It is very hard to do an experiment: you would have to randomly create small classes and randomly assign children to them. There was one experiment in Tennessee in the 80s.
The rest are regressions. You try to use different quasi-experimental techniques to get closer to an experimental setup. You are never really there. And then, when you do a meta-analysis, because not all these studies are of the same quality, you need to be much more judgmental, actually, than in a meta-analysis of medical experimental research, but in a way that is structured and can be replicated. So it is much more complicated to do a meta-analysis in economics and finance than in medicine in general. And the interesting complications will also result from the fact that we are not dealing with experimental results, experimental data, but with observations.
Do you have any questions on that so far? That is why this course actually exists. Why do we not just play YouTube videos by Ioannidis or others who are much more experienced in medical meta-analysis than I am? I have done about 50 meta-analyses in total, so I have some experience with how it should be done, how it should not be done, what mistakes I have made, and so on.
And I just want to tell you that I have really used it a lot in the different roles I have had. When I was working for the central bank in my country, the Czech Republic, as advisor to the board, I would use it every week. A question would come up, and I would either do a quick literature search, a quick literature survey, or look at an analysis, or even write a paper. I was also on the Czech National Economic Council, which is not about monetary policy. So what happens if you increase taxes?
What do you do if the prime minister asks you for your input? Again, quite often I used evidence from meta-analysis. I am also involved a little bit with the Centre for Economic Policy Research in London. For instance, in 2020 we did Europe-wide recommendations on COVID-related policies, which would also rely on it. And then I do research, academic work, in Prague, and also a little bit at Stanford, where, for example, I work with John Ioannidis. Again, we use meta-analysis a lot. It is my primary vehicle to do research. For a couple of years, I was also on an evaluation panel of the European Research Council. And again, in many of these proposals, meta-analysis was contained somewhere.
00:18:15 Participant: Don't you run into the problem that you estimate typically an average effect? For your practical application, you want to have the effect for that specific circumstance. Rather than doing research synthesis, we should do research selection.
00:18:35 Tomas Havranek: Yes. We will talk about it in lecture 9, on heterogeneity. In medicine, meta-analysis is primarily about one number. In economics, for the reason that was mentioned, we are much more interested in the context. For instance, what is the effect of class size, the number of children in a class, on their outcomes in different contexts, in different countries, and maybe also when you have changes from different initial status? If you decrease class size by 5 students, from 30 to 25, it is a different problem than if you decrease it from 20 to 15. So we will pay a lot of attention to context and how to take it into account.
Along with the observational versus experimental dimension, this is the other one, which is related to it: in economics we are really interested not just in one number, even though many times you will need one number, but in one conditional on context. I will touch on it a little bit in today's presentation as well. […]
What happens if the government spends more? 00:20:14
Now, one example: I apologize to the non-economists here, but I will explain what it means. Suppose you work for a central bank and your governor asks you, "I heard the government wants to spend more money, so there is going to be a fiscal stimulus. What should we expect in the economy? What will happen to the economy?" Typically, when you work for a central bank, you will answer, "Okay, I will run my structural model." So you have a model, which has some assumptions, some equations, and you change input data and you get your outcome.
Quite commonly, what people look at are so-called impulse responses. What does that mean? It is just a graph. The question is what happens if the government increases spending. So there is a shock, a positive shock to government spending: the government spends more, by 1% in this case. An impulse response shows you what will happen to different other quantities in the economy. On the horizontal axis you have time. At time zero we have the policy action, and then what happens in the future. On the vertical axis you have the scale of the change in these other variables.
For instance, if the government spends more money, if there is more government consumption, what you have is higher GDP, so more output, quite immediately, in this model. This is a structural model, which you would run as a central bank employee and then show the results to the governor. For now, ignore the different colors and just look at the black one. You have more output, higher GDP, higher wages. You also have more inflation, which is just intuition: you would expect it even without the structural model. Then you have some results for private capital, private investment. You have lower private consumption because the public consumption crowds out private consumption. This course is not about macroeconomics, so I will not go into the details.
Now, if your governor is clever, he can ask you what assumptions you made. He knows that, for the effects of fiscal policy, it is important to know what we call the labor supply elasticity, which means how people actually react if they are offered higher wages. Do they want to work more, or do they not care at all? This is very important because it changes the answer to the original question raised by the governor. So if he is clever, he asks: you are showing me these black impulse responses, but what kind of labor supply elasticity?
You would go into the data and say, "It is 0.25." Then he would ask you, "Based on what empirical evidence do you calibrate the model?" And then you should have a meta-analysis to show him: this is a survey which does this and that, and, given this evidence, the best elasticity for this country is 0.25. Because if you use a different parameter, you will have very different reactions. It will be much more muted for output, for real wage, and for inflation, if you have a higher elasticity.
Again, I will not go into details, but the intuition is this: if people change a lot, if they want to work much more when they are offered more money, higher wages, the labor supply will go up a lot. So in equilibrium the wages will not increase as much, because many people now want to work more, so the employers do not have to offer as high salaries as would be the case if people did not react to the change. But that is not really important. The important thing is that it changes your answer if you use a different calibration. The calibration is almost always, or should be, based on empirical evidence. If your calibration is wrong, if you assume that people behave in a way which is not consistent with the data, then your answer, your policy recommendation, will be wrong.
Meta in monetary policy making: DSGE models 00:25:31
That is one example. I will talk a little bit in more detail about how it works in the central banks. In the central banks, the structural model for which I showed the outcome, the impulse responses, is called a DSGE model. It is a dynamic stochastic general equilibrium model, but do not be too intimidated by the terminology. It just means that you make some assumptions on how people behave, how firms and companies behave, how the government behaves, and then you put it together in different equations.
To be a proper DSGE model, you need to have a dynamic component, which means you have some time series, some changes in time. It is stochastic, which means you can have some random shocks, which come in from outside, by God or by whoever: exogenous random shocks. And it is in general equilibrium, which means it is not just one industry or one part of the economy like the labor market.
These models are used in almost all central banks and ministries of finance, as we call them in Europe.
In Europe, we call central banks national banks: Swiss National Bank, Czech National Bank, and so on. Here, in the Anglo-Saxon tradition, it is the Reserve Bank and the Treasury. But that is the main customer for people who know how to build these models: these two institutions, in every single country. Now, how do they differ from simply looking at empirical evidence? Many people would say these models are better because they are immune to the so-called Lucas critique in macroeconomics.
The Lucas critique means this: you have a huge empirical model, like a regression model, which gives you what happens to one variable when you change another. It is based purely on regression results. For example, what happens to inflation when you change interest rates, if you just run a regression? It is useful, but Professor Lucas would tell you: be careful, because people might change their behavior, and the next time you run the regression, it will be different.
DSGE models are just a way to get around this problem by going to the fundamentals, which should not change at all. That is also questionable, but what people in central banks and ministries of finance usually do, or in macroeconomics in general, is assume that there are things like basic parameters of preference and of behavior which do not change. You learn them by aggregating microeconomic or experimental empirical evidence on these parameters in a meta-analysis. Then you can build a model using different equations on top of these deep parameters. For instance, the labor supply elasticity would be a deep parameter, which supposedly should not change, should be the same.
Dozens or hundreds of equations 00:30:15
So you are appointed to the central bank board, and you look at these models, which I just copied, a random page from one of the working papers. One reaction is that you feel intimidated. Reaction number two is: "I do not get it."
Ideally, it is something in between. Ideally, you know how it works. You know it is not a model which would perfectly fit the economy, and no model does. And you know that there are assumptions. You can question the assumptions, but it is still useful as a way of comparing what might happen if I raise interest rates and if I just hold. So you still need some sort of model. The model could be just intuition, which is really unstructured. It could be anything. Or it could be disciplined by some data and theory, even though the data, theory and the equations are not perfect.
Psychohistory accomplished? Civ VII? 00:31:49
How many of you have read Foundation? It is a science fiction book, set in the future, about a mathematician, originally, who developed a model which allows you to predict the future. It is a complicated model, called psychohistory, and if you know the correct approach, you can gently nudge the society to a desired outcome.
So these DSGE models are a little bit like that. For inflation, the DSGE model is like psychohistory. I guess it would seem like science fiction for inflation: we can nudge the economy based on what the model tells us, what the mathematics tells us, and we will get inflation to 2%. Again, that is a little bit of an extreme view, but I think the analogy is quite useful. Or have you ever played Civilization, the strategy game?
00:33:28 Participant: [A participant asks whether this is the book with the robot laws.]
00:33:35 Tomas Havranek: The entire universe created by Asimov is connected. When you read all of these books, you also get to the Foundation, just one part of the story, but the robots are at the beginning and at the end.
00:33:51 Participant: [A participant remarks that in those stories you learn not to hurt humans.]
00:33:57 Tomas Havranek: It is also there. Or if you play any type of strategy game, it also needs to have some sort of structural microeconomic model involved, if you build the economy and then you have consequences for how much money you get, how many taxes you collect, what your economy is worth. Civilization, a game which I enjoyed as a kid, also has essentially a DSGE model built in, even though they call it differently. So what do you get from these?
Main output: fan charts 00:34:44
When you work on these DSGE models, you are in the end most interested in something called fan charts. It is called a fan chart because you have uncertainty around your prediction which looks a little bit like a fan. For central banks, the key thing is inflation. You want to keep inflation on the target. On the horizontal axis you have the future, different time periods. On the vertical axis you have how large inflation is. The model wants to keep inflation at 2%. That is a number which, by the way, came from New Zealand.
Everybody keeps repeating 2%. It is like the ideal inflation. By the way, there are also many papers which estimate optimal inflation, and I am in the middle of a project which would do a meta-analysis on optimal inflation. But it is difficult because these papers are mostly not empirical estimates: they have their own structural models inside. So it is a bit hard to know how to tackle it.
The fan chart is the main output, which the central bank would then show to the public. Now you see, inflation is going to be 2%. Maybe not, but that is the center of the prediction interval. You can do the same fan charts for interest rates. The model will tell you what level of interest rates you need at each point to get inflation to 2%. At every monetary policy meeting, you will get some sort of recommendation from this model.
Calibration uncertainty: meta needed! 00:36:59
Why do we need meta-analysis? With these models, I have shown you how many equations you need, how many parameters you need to describe the deep behavior of people and companies. Again, it is not a full list, but I just took a screenshot from a paper by my colleagues. These are some of the parameters you need to choose numbers for in the model. You need to say, for example, that the Frisch elasticity of labor supply is 1.25. Of course, then you can ask, why do you calibrate it this way? And the answer should be: I have some empirical evidence, typically in a meta-analysis.
00:37:49 Participant: [A participant asks how many of these parameters are based on meta-analysis in practice.]
00:37:55 Tomas Havranek: As you probably suspect, very few [parameters are calibrated from meta-analysis]. I think they should be.
00:38:10 Participant: Also, all these numbers should come with a standard error, too.
00:38:14 Tomas Havranek: Of course. It is important to know, because if tomorrow you are appointed governor of the RBNZ, you can ask about these things.
00:38:32 Participant: No, but I mean, it is for any decision you make in any company. In a company, you have to make lots of decisions, too. And in some way, for any decision that you make, you would need a meta-analysis parameter.
00:38:44 Tomas Havranek: Yes. But sometimes some of the parameters are more important than others. Of course, you are not going to do 50 different meta-analyses. But I have shown you the labor supply elasticity, which is really important for fiscal policies: if you do these sensitivity analyses, it is typically the most important parameter for fiscal policy. For monetary policy, there are different ones, which I will show you. But one of the most important ones is habit formation. Habit formation means: does your utility, your happiness, let's say, depend just on what you currently have, or does it also take into account what you had last year?
For instance, if I earn a hundred thousand dollars per year and last year I earned $50,000, I am quite happy. But if I earn $100,000 and last year I earned $200,000, I would probably feel a bit different about it. That is the idea about habit formation: you get into the habit of getting a lot of money, of doing well, and then your fortunes change, and in your head you do not just think about what you currently have, but also what you had in the past. That is called habit formation in macroeconomics.
00:40:18 Participant: [A participant asks whether this concerns only salary or also money coming in in other ways.]
00:40:29 Tomas Havranek: Sure. It depends on the use. It depends on the model. It depends on the situation. But the concept would apply to other changes as well.
Economists, typically, historically, would ignore this kind of thing. They would just look at the present. How much money do I have? What are my sources of happiness today? Maybe also the future, but not the past, not what I had last year. That is, again, not the focus of this course, so I will not talk about details, but it is one of the many parameters which go into these models.
You need to assume something about how people behave on this front. You have the labor supply elasticity, and you have, for instance, the elasticity of substitution between domestic and foreign goods. So how much do people prefer apples grown in New Zealand compared to apples grown in Germany or Australia? There are so many things, and I think I will just skip this, because there are too many different parameters.
I said that for fiscal policy the labor supply elasticity was really important. […] I want to show you another example of a parameter that is important for monetary policy. It's called the elasticity of intertemporal substitution in consumption, and it means how much more you will want to save if interest rates increase by one percentage point. Typically, people will save more if interest rates are higher, because you can put your money in a bank and get a bigger return than if you just spend the money. That's very important for what happens if the central bank increases interest rates, because that's how central banks operate. They typically change the interest rate to control inflation.
They can also do other things, like buying bonds or foreign currency, but mostly they change interest rates, which influences mortgage rates and savings rates and then influences how the economy behaves in general. And when you have, again, an impulse response, first there is a shock in the interest rate, then time passes, and you see what happens to, in this case, investment, like private investment. And you can see again that the answer will differ a lot depending on the elasticity of substitution in consumption. So this is a key parameter. We have dozens of different parameters that we have to calibrate, but this one is really the most important for monetary policy.
The second parameter, very important for monetary policy, is habit formation, which we talked about before. So again, it's the same thing: you change the parameter, and the complicated model will give you very different...
00:44:29 Participant: But again, here your meta-analysis will give you a point estimate and then a confidence interval. So are you sure that your confidence interval isn't going to be as wide as your...
00:44:40 Tomas Havranek: It could be true, especially in this case, where after a couple of quarters the difference is not so huge.
00:44:51 Participant: I agree that we should do a meta-analysis, but in the end, even with a meta-analysis, we think we are still quite uncertain about the range.
00:44:59 Tomas Havranek: Yeah, but at least you would acknowledge the uncertainty. Normally these models take calibration as given. This is again technical. Are you familiar with Bayesian statistics? Ideally, this kind of model should, and actually can, be done using Bayesian statistics, Bayesian econometrics, where you would have these meta results, if you have them, a point estimate for the parameter and a confidence interval, which you would take as your prior. Then you would use the rest of the model and the data, which come from the economy, to estimate the posterior.
That would be the outcome of the model. So you would have an impulse response with a confidence interval, which is what you are most interested in. That's actually how some central banks are using this kind of approach I'm talking about. So let's focus on a couple of meta-analyses for the key parameters. Then for some parameters you don't have enough empirical evidence, so you still have to guess a little bit, because there are simply too many of them. But for the main ones, you want to do a survey and a meta-analysis, and then you use Bayesian statistics to combine it with the data, and you get the posterior estimate, either for the parameter or, especially, for the behavior of the economy.
00:46:43 Participant: [A participant asks how to calibrate a parameter for a specific country.]
00:46:58 Tomas Havranek: That's a good question. Typically, for small countries like Czechia, there are no studies at all. One way is to actually do a new study: not a meta-analysis, but a study estimated just for Czechia. The problem is that typically there are no studies, because the data period is too short.
The main problem with countries like mine is that we had communism until 1989. It was not an economy in a modern sense: everything was planned, so you would have no real prices. The price was what the planning committee decided. So you have no data, no GDP: there was no concept of GDP. Inflation was officially forbidden, because the government would just say that the price would be this and that. One famous example, which my mother told me about, was toilet paper.
There was not enough toilet paper, but the government, to keep inflation down, would say that the price of toilet paper would be this low. What happened was that there was not enough toilet paper to buy. So there was no toilet paper for half a year.
00:48:22 Participant: [A participant asks whether people could queue for the toilet paper.]
00:48:25 Tomas Havranek: You queue, but at some point there was nothing to queue for, until they changed the plan. And for the next five years there was too much toilet paper, because they overdid it. So it's hard to measure the economy, and you have no data. It's very difficult to run your model just based on this short data period. Now it's quite a bit better: it's actually 35 years, so you can do things, but it's still not much.
To answer your question, what you can do is look at these other studies for different countries, compute the average estimate, and then compute how the results change depending on the country's characteristics. Maybe in small open economies like Czechia, the Czech Republic, you have a bigger habit formation for some reason. You can take a look. Or if you find no variable that explains any differences, you can say, okay, it seems to be universal, similar for all countries. But certainly you can find reasons why parameters like these could be different for different countries.
Especially, we mentioned the substitution between apples from New Zealand and Australia. It will certainly differ across countries, because some countries are more patriotic, and you want to consume domestic goods, even if they are sometimes more expensive, and some people in other countries don't care so much, so they mostly look at the price. It's very plausible that you will have cross-country differences in the extent to which people prefer domestic goods. That has implications for the exchange rate, and it has implications for all sorts of things in the model. So that's an important question.
So far it was just motivation, for why it matters in the specific case of central banking, but it will be very similar for treasuries, ministries of finance, and for institutions, for example, which tackle environmental sustainability, things like how much should we tax carbon emissions?
If you want to compute the optimal carbon tax, you will need a model like this one, but much more complicated, because it's not just the economy: you will also have physics there, climate science. You will have environmental issues, like biodiversity, in the model. It's called the social cost of carbon, and the models for the social cost of carbon and the optimal carbon tax are called integrated assessment models. That's what my wife focused on a couple of years back. We went to Berkeley for a couple of months and we wrote a meta-analysis of social cost of carbon estimates. That's one example definitely outside of central banking and treasuries.
We in economics and finance are very lucky because we have these strong policy institutions like the central banks, treasuries, and other regulatory agencies, but especially the central bank.
If you're a psychologist, you don't have that, but you have something a little bit similar: the nudge units. In the US they are actually quite powerful in designing how websites are built and how, when you offer people options, for instance different pension plans, it should be done. We often call it just economics as well, but it's mostly psychology.
These nudge units are mostly in psychology, and they would also need meta-analyses. There is a famous meta-analysis of the effectiveness of nudges published in PNAS, and also another one in Econometrica, an economics journal. When you do meta-analyses of the effects of nudges, what does it mean?
It could be that you have a visual thing that makes you click on this side but not on the other side. It can also mean you save more money. It can also mean you recycle more. So when you put it all together, you probably have an apples-and-oranges problem, which we'll talk about in heterogeneity. And so I think there are good reasons why meta-analyses of this kind are criticized, because you put together too many things, and in the end it's just unclear if you can compare them. But we will get to that.
Estimates of the labor supply elasticity vary 00:54:21
Suppose you want to calibrate a parameter which is easier: labor supply elasticity. Again, this is the parameter important for fiscal policy, which means how much more people want to work if they are offered more. They should work more, so it should probably be more than zero. But the different studies disagree quite strongly on what the estimate should be. The different studies are here on the vertical axis. On the horizontal axis we have the size of the estimate.
And you can see that even within one study you can have pretty different estimates, for one country and essentially the same technique. Quite often in economics people do robustness checks, so for instance the baseline could be something like difference-in-differences, then they could do IV, simple regression, different techniques of regression analysis, and they would have different results. So in this case, and it's the case for many other parameters as well, studies would disagree both across and within. Sometimes you will have these differences. So it's not clear, when you just look at it at first sight, how you should calibrate it.
If you just want to do a summary for calibration and you don't care about meta-analyses, what many people have done before is quite often just the average.
00:56:12 Participant: [A participant asks whether any studies check how different the conclusions would be if one used just the average instead of the meta-analysis estimate.]
00:56:24 Tomas Havranek: Yes. We have a new paper in the Journal of Economic Surveys with Sebastian Gechert and other colleagues, which compares, for 15 prominent meta-analyses, the conventional knowledge, quite often the average, with the meta-corrected average. [Note, 2026: the published paper covers 24 influential meta-analyses; see meta-analysis.cz/conventional_wisdom/.]
00:56:51 Participant: [A participant suggests doing the same with the thousands of meta-analyses in medicine, comparing how different the conclusions are if one just uses the average, which would answer whether meta-analyses are actually needed.]
00:57:22 Tomas Havranek: It would be, I think, a nice thing, easy for you to do.
00:57:30 Participant: [A participant agrees that it would be fantastic: as with econometrics and its many fancy methods, one can ask whether they actually matter, and here this could be done with a huge data set.]
00:57:42 Tomas Havranek: And I would add a small spoiler. I have done dozens of meta-analyses. Quite often, though not always, what happens to me is that the average is way above [the final result]. But if I look at the median, and then I do all those complications, in many cases the final result is not far away from the median.
00:58:10 Participant: [A participant remarks that one could make a decent decision with simple rules instead of a meta-analysis.]
00:58:15 Tomas Havranek: At some point I wanted to write a paper on it, but you can go ahead.
00:58:20 Participant: [A participant notes that the data already exist, so this can be done.]
00:58:22 Tomas Havranek: So I think it's interesting. If you want to do something quickly for your boss in the central bank, my recommendation would be to just do the median. It's not perfect, but it takes care of a surprising number of problems with publication bias, p-hacking, outliers, and so on. So at first guess, the median is pretty good.
00:58:49 Participant: [A participant suggests that one could quantify how good the median is, for example the share of cases in which it gives the same result, and asks what the conclusion was in that paper.]
00:59:02 Tomas Havranek: We didn't have the median, I think. We just compared what people would typically use, I assume the conventional wisdom, which is quite commonly the average, to the corrected result. And so on average, we found that the corrected result is one half of the average. But we are not the first: using different data sets, Ioannidis et al. found the same thing. Also now Bartos et al.
00:59:30 Participant: [A participant asks whether the data sets were big or small.]
00:59:33 Tomas Havranek: Yeah, ours was small and they have big ones. But I think it's just in economics, so you could do it cross-study. There is a new paper by Bartos et al. in Research Synthesis Methods, so maybe I can point you to that as well. It's related to this idea as well. But not the median. I have never seen a meta-analysis of the median.
Values around 0.4 to 0.5 used for calibration (CBO) 01:00:23
This was just a digression. But, for instance, in the US the CBO, the Congressional Budget Office, is another kind of institution which uses structural models.
It is a little bit like a DSGE model, but anyway, it is a structural model, which the CBO runs to analyze fiscal policy. So they need to calibrate this model, including the labor supply elasticity. They collected these studies, and they take the average.
Bias: some estimates less likely to be published 01:01:10
It is a problem especially in this case, as you can see from the histogram. We have the t-statistics of the estimates on the horizontal axis and, on the vertical axis, the frequency, how often they appear in the literature. You can see a huge jump at zero. If the estimate is just positive, just a tiny bit positive, it is much more likely to appear in the literature.
Now, how could this be? Of course, the intuition here is extremely strong that when you offer more money to your people, they will not work less. We know from Microeconomics 101 that it could happen if I pay someone one million dollars per hour of coding and then offer him two million dollars: he will probably not work more, because he will already have so much money that he will either not care or actually work less, because he can afford it, and he will just want more leisure. But normally, when you offer people more money, you expect them to work a little bit more, to earn more. So it is difficult to publish estimates which show the opposite. The labor supply in total would react in the opposite way, and that is really weird. So we know that these negative estimates are probably not really realistic.
But the problem is that sometimes you get these negative estimates just because you have some estimation error, some noise in the data, some noise in your model, in the estimation technique. So when you run a regression like these people are doing, in these quasi-experiments and OLS regressions, there are no bounds that would tell you that it should not be negative. So they get many negative results.
The most plausible interpretation of the figure here is that people just do not report these negative numbers, because they would seem weird. So you can report a positive and statistically insignificant one. So what happens is that there should be a whole bunch of estimates here which you do not see published. You just see most of the positive ones. So in consequence, the average, and also the median, but especially the average, is too large. It is exaggerated compared to what should be the true mean. […]
So again, we will talk a lot about publication bias and p-hacking in much detail, if you want me to. But for now, just keep in mind that this is one of the primary reasons why we need a meta-analysis, why it is not okay that many of my friends in economics would say, "I do not care about these other studies, I just need to look at the one study published in the AER, the most prestigious economics journal. And I trust this one the most. Go for the best study."
So, it could be. But you have no idea about the publication pressures behind it, and it might be just a lucky estimate which ended up being published, also because it is statistically significant. So, if you do research on meta-analysis and you want to publish high in economics, you need to tackle these issues somehow, because many referees will tell you, "I do not care about 90% of the literature."
So in most of our papers, we somehow do some sort of subsample analysis just for the top journals, or top five journals plus, because in economics we have a strong obsession with top five journals. I do not think that is the case in any other discipline, but in economics, if you do not publish in the top five journals, you are essentially not regarded as a kind of all-star economist.
But other people will tell you, "I just care about the very top results." But sometimes it will be really misleading.
01:07:26 Participant: [A participant remarks that the incentives for p-hacking and similar practices are strongest at the top journals, so publication bias should be larger there.]
01:07:38 Tomas Havranek: And I think there are some studies which show that as well. But also a little bit on publishing in economics: it has changed a lot since I was a PhD student. When I was a PhD student, you would see no meta-papers: not just meta-analysis, but any kind of meta-research. "Meta" comes from Greek. It means beyond or above. That is why the Facebook parent company was renamed Meta. When I was young, it was really impossible to publish a meta-paper in a good journal. But now every year you see meta-analysis and meta-research in the top five journals. It is not so rare anymore.
Of course, we had a Nobel Prize for David Card. David Card is a labor economist, but he has written about six meta-analyses. So David Card got a Nobel Prize, partly for his research on the effects of the minimum wage, which included the 1995 meta-analysis. Isaiah Andrews got the John Bates Clark Medal, partly for his work on publication bias.
Meta-analysis is much more accepted in the profession. The top journals are not rejecting these papers anymore. I think it is a good field for a young postdoc or young professional who wants to do research, because many fields in economics which publish well are super intensive in terms of finances. You need to do an RCT in Uganda, or you need to have a few million observations.
But if you have a clever idea in meta-research, or you have already collected a good-sized data set, you can do plenty of interesting research on these issues, and all you need is to work hard. You do not need to be in the top one percentile, and you do not need a PhD from MIT, for example.
01:10:42 Participant: [A participant says that meta-analysis is now respectable, but that qualitative literature reviews, for example in the Journal of Economic Literature, are still even more respectable and get cited more.]
01:11:09 Tomas Havranek: I think that is still the case, which means there is still plenty of potential. That is not how it works in other fields. In psychology, people would look mainly at meta-analysis, I think, more so than at just narrative surveys.
01:11:26 Participant: [A participant adds that it goes the opposite way in some fields: surveys of the literature written in words have always been accepted.]
01:11:48 Tomas Havranek: Yeah, but now you have meta-analyses published. […] So, you are right, it has still not switched, and maybe it will never switch, and by the way, maybe it should not, I do not know. But there is much more acceptance, and you will be listened to much more if you do a meta-research paper in economics than was the case when I was starting in my PhD studies.
So I think a good meta-analysis is also by definition a good research survey. It is hard to do, by the way. I do know how to do a meta-analysis, but I always need someone who understands the literature. That is what I recommend to anyone who wants to do a meta-analysis, like a survey of a given field: you will know how to do a meta-analysis, and you need someone who understands it. Maybe you already understand the particular field.
But since I do meta-analysis for macro topics, micro labor, finance, environmental, and so on, I need to recruit people who understand it deeply. Otherwise, I would make mistakes in data collection and interpretation. They are the people who then actually write the narrative part.
Context: different situations yield different estimates 01:13:30
Now, context. Related to what was mentioned earlier, it is very important not just to look at one number, but at how it differs for different types of data, types of households, methodologies, and so on. Here is just an example of the different characteristics we took into account when we did a meta-analysis of the labor supply elasticity studies. For instance, there are too many of them, so I will just talk about these ones: prime age, near retirement, and women.
In terms of labor supply elasticity, it matters a lot what your age is. If you are 40, like me, you will definitely work anyway. And if I am offered more money per hour, my labor supply would not change much, because I still need to spend time with my family and so on. And I will work anyway, even if they decrease my salary. What else could I do?
So my labor supply elasticity would be pretty small. But when I am about 65, my kids are grown up, and there is a crisis and the university slashes my salary, I can just go into retirement. So I will be much more responsive to changes in my wages, which is the common economic intuition.
Historically, it was typically the case that the husband would make more money than the wife, and the wife would work maybe part-time to also be able to take care of the kids and the household. So any changes in the wage of the husband would have essentially no effect on his labor supply.
But for the wife, if you increase the wage, now it is much more economical to work, because maybe it could be better for you to hire a nanny for the kids and to work. So historically, women were more responsive to changes in wages, which means in economics terminology that they had a higher elasticity of labor supply.
So one of the things which are interesting to look at is whether this is really the case. Do we observe in the literature higher elasticities for older people, near retirement, after 60, and for women? We look at many other issues, but these are two which are intuitive to explain. And we find that, yes, it is true. I will not show the results. I will show the results in some other lecture, when we talk about heterogeneity and how it is done. But I will just talk a little bit about how, when you collect your data for your meta-analysis, you somehow correct it for publication bias and p-hacking.
Then you estimate how the results depend on context. And then you want somehow to summarize everything you have done into a number which would have confidence intervals and would be useful for, for example, calibration in central banks or congressional budget offices or whatever. So one thing you can do is to take your results and compute the conditional estimate, which from a regression is something like a fitted value based on certain characteristics of the research literature. So you can do it separately for overall elasticity, for older workers, and also for women and men.
For the rest of the variables, you can plug in the mean if you have no preference. But typically you have some preference for, for example, quasi-experimental results, so you want to give more weight to better identification. So if it is a difference-in-differences estimation using, for example, tax holidays in Switzerland, you give more weight to it. Again, later on I will explain more how it is done and how the details work, but just as a taste.
I will give you an example of what we get from a typical analysis. We get something like this table [Implied elasticity: taking bias and context into account], which shows you your best guess, based on the analysis, for the parameter, both in total, on average, and also for the main specific contexts, which in this case would be older workers and women. And you can see there is some difference.
Overall we find no evidence for substitution, not much evidence, but for older workers and women there is some evidence. If you wonder why we have different results in these different columns, it is because one column uses what, based on our reading of the literature, is the best practice. So, do we prefer IV or diff-in-diff or something? And the other column is based on the most prominent study in the literature, which most people would say is the best one. And we just take the characteristics of the methodology there and put them in our meta-regression model.
I will later explain the details, how it works. So, for example, CBO can then go and use these numbers for the calibration of their model, because, at least in some versions of their model, they can have different values for women, for workers in retirement, and so on.
So this could be used in practice. Of course, the confidence intervals are quite wide, but that is simply what we get. But you can notice that even the largest confidence interval here does not include 0.5, which is the upper end of the values they now use for the calibration in the CBO. [Note, 2026: in the published paper the upper 95% limits for the intensive margin are 0.52 to 0.59, just above 0.5; see meta-analysis.cz/frisch/.]
Course 01:21:48
So this was the labor supply elasticity. Now I guess I should talk a little bit about how the course would be structured, what we will do. What we have been discussing so far was motivation: why meta-analysis matters. Now let me turn to what we will actually do. Tomorrow I want to talk about some applied examples.
I think it's best to start by showing you what a meta-analysis looks like, including some problems and mistakes which I made in my previous papers. I will point these issues out, because some of these papers are a bit old, from different fields, such as environmental, labor and macro, so different areas of economics. And again, if I'm going into much detail, we will leave the details for the other lectures. After tomorrow, which will be just an overview of three papers, we will go into much detail in the literature search: what could be a good way to first search for the papers to include in the analysis, then how to actually collect these data, what works and what doesn't work.
Then we will talk about basic synthesis methods, which are typically taken from medical research: fixed effects, random effects. Some of you already know them. [Session six is for student presentations.] And then we have two sessions on publication bias and p-hacking, because these are different issues and they have different solutions.
I think that will be probably the core. If you really want to do hardcore meta-analysis for research purposes, these are the main issues there. And then, of course, heterogeneity right after, but if you fail to account properly for p-hacking, for example, it's difficult.
And then you have heterogeneity, so context, which we already hinted at a little bit. And meta-research. What meta-research means is not exactly meta-analysis but, for example, using some meta-analysis results to deliver some additional scientific value added, or it can be some practical value added. So I will talk about different ideas of what has been done and what could be done, maybe.
It's also important to talk about the problems we have. I think we should explicitly have one session on these limitations. Especially now, when, not just in economics but in general, meta-analysis has exploded in terms of the number of papers published, and you can see many metas which are not really well done. So one needs to be careful, but even at the frontier there are still problems which I don't know how to address, so I want to name them explicitly.
And again, the final session would be some presentations of some small projects, or, if you are not officially registered, you can just present whatever you wish. It would be optional. So that would be the structure. […]
Guidelines 01:27:47
I should also do some self-promotion. We have this guidelines paper on meta-analysis, which is short and briefly outlines how we think meta-analysis should be done. It's me, my wife Zuzana, Tom Stanley and Chris Doucouliagos. And it was very difficult to agree between the four of us on some issues.
Originally we wanted to do it on the MAER-Net basis. MAER-Net is the Meta-Analysis of Economics Research Network. But we started with Tom and Zuzana and Chris. So we have the guidelines, which I think are pretty judgmental, and I'm glad about it.
So we are able to put in some concrete recommendations, which doesn't mean that this is set in stone, but at least if you want some advice, you get it. It's up to you whether you use it or not. We may be wrong on many accounts, which we can also talk about. This one paper I would like you to read at some point.
Resources 01:29:21
And also, if you want to do some replication or have a look at what meta-research can look like, we have on my website meta-analysis.cz dozens of different examples of different applications from different fields. That would be my website, and I have reached the end of my talk today. […]
Corrections
Slips of the tongue corrected in the text:
- [00:12:08] said "randomized controlled experiments"; the text has "randomized controlled trials".
- [00:18:35] said "lecture 8"; the text has "lecture 9".
- [00:21:45] said "crowds out private profits"; the text has "crowds out private consumption".
- [00:39:29] said "what you have in the future"; the text has "what you had in the past".