Transcript. Lecture 5, Basic Meta-Analysis Tools, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.

Lecture 5. Basic Meta-Analysis Tools 00:00:00

Summary 00:00:33

00:00:33 Tomas Havranek: In the last two weeks, we talked about how you can or should choose your topic and what your motivation punchline should be. You need a rationale for doing a meta-analysis. Then we talked about how you search for studies and, finally, how you collect data. That was last time. Now we will talk a little bit about the brief first-stage meta-analysis that you do.

Mean, median, something else? 00:01:24

Essentially, in medicine, quite often they would have just a few studies. Also, if your boss, the minister or the governor of a central bank asks you to do a quick survey, typically that's what you will do. You will have a couple of studies, so you will not go for a full-fledged meta-analysis, but you will do some sort of weighted averages.

It's useful to know this, even though in economics and finance we usually go beyond it. In applied papers, we commonly skip this entire summary statistics stage. But it's important to know it nonetheless.

If you want to start with something really simple, without any weighting or anything complex or complicated, it's best to look at medians. You have a couple of studies, maybe a couple of dozen estimates, or you can have many more. But if you want to show your boss some first results that could be useful immediately, it's much better to look at the median, not the mean. The reason is, as we already discussed, I think at least twice at these sessions, the median is more robust to outliers. It can also correct a little bit for publication bias, sometimes quite a lot.

I mentioned an example from my own research: in many cases, when I do a paper, I do a big meta-analysis and start with the summary statistics. After all these complicated tests, I get results that, not always but quite often, are not far away from the median. Almost always, they are far away from the average because of publication bias, outliers and other issues. But the median is, I think, underrated in meta-analysis. We should use medians more. It's not perfect, but as a first guess, it's not so bad. It's probably the best you can find.

More precision, more weight 00:04:09

But as I have said, that's not what people typically do in meta-analysis. To most people who have done it, meta-analysis means that you take the estimates and weight them by inverse variance. What do I mean by inverse variance? You collect your estimates and their standard errors, which measure how certain you are about these estimates, how uncertain they are. Variance is the square of the standard error, as you know from our econometrics classes. Inverse variance is 1 over variance. You put more weight on results that have lower variance and are more precise.

In economics, we use the weighting mostly as a way to reduce publication bias, and we will talk about it next week. But originally, in medicine and education, where meta-analysis originated in the 70s, 80s and later, it was primarily meant as a way to decrease the uncertainty of your meta-analysis estimates. The primary goal was to increase precision in your meta-analysis. Still, many people in medicine think about inverse-variance weighting in this way: to get more precise estimates. This works if the assumptions are fulfilled.

Let's have a brief example of how it works and what we do. It's not complicated. Let's do a simple simulation. We have estimates, maybe of some elasticity or something, on the vertical axis. We have standard errors on the horizontal axis. The standard error tells you how imprecise the estimates are. Now let's suppose the true value of the elasticity is 1. We do a simple simulation, so I can control what the true value is.

I set it to 1. Then I simulate many studies that are the same, but have samples of different sizes. They have random samples from the population. One study has a small sample, and another has a large sample. If you have a small sample, you have a small number of observations, so you will have large standard errors, and vice versa. If you have a large sample, you will be precise.

I simulate hundreds of studies. These are the studies with small samples and therefore a lot of variance, a lot of dispersion. If you have a lot of information, a large sample, it will be really precise, and there is not going to be much dispersion because these studies will be really close to the true effect, which in my simulation I set to 1. When I do this, I control the simulation. In practice, of course, you don't know what the true value is. It may not be the same across studies. We'll talk about that in a couple of minutes.

These are the data points you collect from the literature, from the primary studies. You want to estimate the underlying effect, the true effect. You also need some degree of uncertainty, some standard error around it. We are used to thinking about regressions in economics and finance. Essentially, you estimate a regression, but you are interested in the intercept here. You can regress estimates on the standard errors, but you care about the intercept.

That is the infinitely precise estimate, and that's what you want to find. If you think about it this way, the regression has a problem. What is the problem with this regression of estimates on the standard error? It looks like there is a specific econometric problem in the regression, called heteroscedasticity. This means that the dispersion of the variable on the left-hand side, on the vertical axis, is related to the size of the variable on the horizontal axis, the independent variable.

The dispersion, the variance of the dependent variable, is related to the size of the independent variable. That's heteroscedasticity in regression. What do you do in regression when you have heteroscedasticity? You can either use heteroscedasticity-robust standard errors, or you can do what? You can do weighted least squares when you know the pattern of heteroscedasticity. Here you know it because, by definition, it depends on the size of the standard error. The standard error is a measure of dispersion of the estimate.

You have heteroscedasticity by definition. One solution is to use weighted least squares, in which you would use some sort of function of the variable on the horizontal axis as your weight. For many economists, this is an easy way to understand why people would want to use inverse-variance weights in meta-analysis: it removes heteroscedasticity.

If you do a regression and don't include a variable, so you just estimate the intercept without any controls, then it should be exactly the same as a weighted mean.

If you do this regression and don't weight or use heteroscedasticity-robust standard errors, your estimate will be okay, but your measure of variance will be off. This is just a motivation. […]

Fixed effect and UWLS 00:11:38

00:11:38 Participant: [A participant makes a comment that is hard to hear.]

00:11:43 Tomas Havranek: That's what I'm going to talk about next, exactly as you say. It's easier to start with a slightly less realistic scenario in which you assume that all the studies you collect, all the primary studies, really estimate the same thing. Not just the same thing in terms of the concept, such as the elasticity of substitution between capital and labor, but really the same thing: you assume it is the same population. There are no differences across countries, no systematic differences across techniques people can use.

That's called a fixed-effect model in meta-analysis. It's different from what we know from panel data econometrics and finance. When we say fixed effects, it means you run a panel data model with dummy variables for different cross-sectional units, typically. There is something completely different here, and in a while I will explain why. Fixed means there is just one effect, and it's fixed. In all studies, it's the same. If you observe two estimates from two studies, you have an estimate here and here, and a variance, a standard error for each estimate, V1 and V2.

If you assume a fixed effect in your meta-analysis, your model will, by definition, assume that there is just one true effect common to these two studies. Your fixed-effect meta-analysis will try to estimate it with some degree of precision. You assume there is just one effect that all the studies try to estimate. To decide which you should use, let me explain the other alternative, and then I will come back to your question.

How does it work? It's very easy, so you can listen to what I say. You don't have to read the formulas, but they are not so difficult. First, you have an estimate. Somewhere in that estimate is the true value, theta, and then some noise. If you work in the fixed-effect framework in meta-analysis, that's what you assume. We will also discuss the other framework. You assume that your estimates are the true effect plus some sampling error, some random noise in your data, which is given by the size of your data set and also by some random problems, maybe, in your data.

But there is no systematic difference across studies. You have the variances of these estimates, which are, of course, inversely proportional to sample size, as you know from statistics. But the important issue is the weight you use in meta-analysis. The weight is 1 over variance, 1 over the standard error squared, which you collect from these studies. Really precise studies will have a large weight because V is small. Conversely, if your estimate from a primary study is imprecise, V is going to be large and your weight W is going to be small.

For your meta-analysis estimator, you weight the estimates y by a weight proportional to inverse variance. You have a weighted average, and that's your estimate. That's very simple and not really so surprising or interesting.

What is interesting is how you compute the variance of the meta-analysis model. You completely ignore differences between the studies: you compute it as the inverse of the sum of the weights. There is no information about how each study differs from the mean value, the true value or the estimate you compute. It's just one over the sum of weights. That's not a common formula we would use in econometrics for a weighted average. For this to work, you need the assumption at the top that there is just one effect, the same for all studies.

I will show you the other alternative in a few seconds, where you will have a different formula. There is no role for any between-study heterogeneity. What you often see in meta-analysis in economics is so-called UWLS, which means unrestricted weighted least squares. That's essentially the same as the fixed-effect estimator. It's also a weighted average, weighted in the same way as for fixed effects. If you do meta-analysis using fixed effects and UWLS, you will have the same result in terms of your meta-analysis point estimate.

But you will have different variances, different measures of precision for each of these meta-analysis estimators.

This is the variance for the fixed-effect estimator, which we discussed before, which has 1 in the numerator. For UWLS, you have a standard sampling error for each study, how far away it is from the value you estimate in the meta-analysis. Typically, not always but often, you will have a larger variance using the second formula. This would be the default computation of variance if you do any kind of weighted least squares in economics, in econometrics. You will compute it this way. If you do it in R or Stata, that's what you will get.

It depends on the underlying assumptions. If you want a more robust measure of variance, I would definitely, always prefer the second one. In most cases it's larger, more conservative. It also takes into account the possible heterogeneity, which might be there, or maybe it's not there, but you typically don't know ex ante. The fixed-effect estimate is conditional on very specific assumptions, which are unlikely to hold, especially in the fields we are discussing, economics and finance. They might be more likely to hold in medicine in some cases.

If you estimate completely the same thing, it might make sense. I will give you an example soon. You can say it's more correct. In many real-world applications, UWLS is going to suit you better than fixed effect.

One example where fixed-effect estimation can be reasonable or useful is if you estimate the average score of children in the same school and take different subsamples of kids from the school. You have one, two, three, four, five different experiments or five different studies in which you take different samples from the student population and test them in math or something, I don't know. Then you see the scores. You know you are estimating the same thing, not just conceptually, but really the same number.

You have some noise based on the sample that you selected from the same population. Here, it makes sense to use fixed effects. It is a very specific case. You will not see many examples like this in meta-analysis, but it could happen.

In one school, you are estimating the same effect. It should be the same. You have five different studies, five different experiments, five different samples. The samples give you different point estimates. They give you some measure of precision, some measure of variance, depending, in this case, just on the sample size that you have. It is 200 in four experiments, but in one you have a really large study, number three, study C, which is more precise. It has not 200 students, but 800 students.

When you do fixed-effect estimation in meta-analysis, again, the weight is going to be proportional to precision, proportional to inverse variance. If you have more precision, there is going to be more weight. This large study, study C, will have 50% of the weight, and the other studies will have less because they are less precise and have smaller samples. In practical applications, especially in economics and finance, you will quite often have studies which are much more precise than the other studies.

You can have 50 studies, and one study is super precise, 100 times more precise than all the others. When you use the fixed-effect framework, you should be aware that you give much more weight to this one really precise study. It becomes an outlier or leverage point which can really drive your results a lot. This is a point which Bob Reed has raised a couple of times, and I think it is important to realize when you do fixed effects, even UWLS.

The point estimate is really going to be driven by these studies. They can be large, but they can also be super precise for some other reason. It could be a typo, a decimal point mistake. It could happen. Be careful about these super precise studies. You want to place more weight on them, but be careful not to place too much. If you have one super precise study, double or triple check that it is really correctly coded.

Sometimes, and we will get to that later, in economics it is much more difficult than in, for example, medicine to compute the variance of your estimate. In medicine, when you do an experiment, an RCT, your variance is essentially driven by your sample size, as in this example. There could be a little bit of something else as well. Of course, you have some sampling error, but what drives almost all of it is sample size. In economics, of course, sample size is also important.

If you have millions of observations, you will have huge statistical significance; you will have precise estimates. But it is much more important how you compute your standard error. We talked about heteroscedasticity. That is one issue. Heteroscedasticity is typically not a huge issue. It can change your variance by maybe 20%. It can still make your estimates insignificant. But what is more important is clustering. When you have more estimates from one study, which is typical in economics and finance, you can have 20 estimates from one study, five estimates from another study, just one estimate from another study, and so on.

That is a clustering problem in meta-analysis. But before that, you have different studies which, for instance, have panel data: data for different countries for different years. When you do panel data estimation, the right thing to do is to cluster your variance, your standard error, so that you take into account that the estimates within one country, for example, or within one year are likely to be somehow related and are not completely independent.

For instance, if you try to estimate what the determinants of growth are, why some countries grow faster than other countries, there have been hundreds of papers like that. When you want to do it, you collect data for as many countries as you can, and then typically you take a couple of points in time. You can do a panel each year, or, more commonly, you take 10-year periods. Then you regress the growth in these countries on some country characteristics. You can include dozens of different characteristics.

These include investment, education, institutions, and whatnot. If you do it properly, for example, when you are interested in the effect of institutions on growth, you should cluster your standard errors at the country level in your regression, maybe also at the decade level, but especially at the country level, because your observations within countries are most likely going to be dependent on each other.

For time series for Afghanistan, for example, or for New Zealand, there is going to be dependence because it is still the same country. But not everyone does it this way. Some people ignore it, especially in the older studies, but even now. Or they do it in the wrong way. Sometimes you have a small number of countries; you can have just 20 countries. If you have just 20 countries, you cannot really use clustering because for clustering you need a large sample.

These cluster-robust standard errors will not work, or they will work, but they will be heavily biased downwards when you have a small sample. With a small number of clusters, you cannot really use the cluster-robust correction, but you should do something like bootstrap, wild bootstrap, as people do a lot in labor economics, for example.

It will also be an issue in meta-analysis, but primarily it is a problem in primary studies. What I am trying to say is that sometimes you can have studies which seem to be really precise, but that is just because the authors did a sloppy job in terms of how they compute precision. Then you will give the sloppiest studies the most weight, which is not optimal. Be careful when you collect data about what precision actually means. I will talk about it much more in two weeks, I think.

I think it is important to stress this even here. That is the fixed-effect estimator. More commonly, if you want to work in this simple framework just to do a quick meta-analysis, you will want to do something different.

Random effects 00:31:05

You will want to do random-effects meta-analysis. Now you relax the assumption that there is just one common effect across your studies and allow for potential differences. You still estimate the elasticity of substitution between capital and labor, but because you do so for different countries, you can have small differences. The elasticity could be bigger, for instance, in more developed countries than in poorer countries, or there could be some other sources of heterogeneity.

00:31:45 Participant: How does it work?

00:31:47 Tomas Havranek: Again, you have two studies, two point estimates, plus some variance, which you see reported in the papers, in the primary studies. But when you do meta-analysis, you do not necessarily assume that there is just one common effect. It could be something else.

Then you compute your estimate, and I will show you how, but you also compute something which is commonly denoted as tau squared. It is a measure of heterogeneity, a measure of dispersion across studies. This tau squared was assumed to be zero in the previous fixed-effect case, where you would have just one estimate, one effect, the same for all studies. That would be your assumption. Here you allow for some heterogeneity, even in the true effect.

How do we interpret it? As we discussed, I think maybe two weeks ago, when you have some practical application, for instance, in a central bank, and you do it for New Zealand, you are really interested in the New Zealand value. This average would not be really informative for you. You would want to go further, which we will do anyway. But if the question is more general, what could I typically expect to be the most likely number? Of course, it can be on one side or the other in different countries. Then the random-effects estimate, I think, is a good start, because, as we will see, it does not put so much weight on these outlying, super precise studies. It is less sensitive to the problems which I described, which we call spurious precision. I will also talk about it.

But if you want to use it for practical applications, you need to go a couple of steps further, of course. It is a much more conservative, much safer estimate for the typical true average, even though you know there are going to be systematic differences across contexts.

Now we again have a couple of formulas. Do not be afraid. It is essentially the same as we had before. The only difference is that here we add this new guy, which captures some systematic dispersion across studies. Before, we had just the true effect and the sampling error. Now we also have some underlying differences across studies, which you can call structural differences. The variance is the same; we could put i here.

But now, when we compute the weight which we will use in the random-effects model, we do not have just the variance in the denominator, but we also have tau squared. Tau squared, as I have mentioned, is a measure of dispersion between studies, the heterogeneity across studies. If there is no heterogeneity, tau squared is zero, and we go back to the fixed-effect model.

This animal here, tau squared, will disappear. We would just have one over variance, standard inverse-variance weights. What happens here is that the heterogeneity term decreases the importance of variance. The bigger the heterogeneity you have here, which you need to estimate, the lower the importance of variance you will get. It is a natural way to decrease the super high importance of outliers, or you can think about it in this way.

The meta-analysis estimator M, again, is the weighted average. It is the same formula. The only difference is that, for the weight, you do not have just inverse variance, but you have the inverse variance adjusted for heterogeneity, which means, in effect, that you give less weight to precision than in the previous case.

[A participant asks about the difference between these estimators.] Let us talk about UWLS. It only adjusts variance for heterogeneity. The estimate M, or however you denote it, your meta-analysis estimate, is the same as it was for fixed effect. That is UWLS. You leave the point estimate as it is, and you typically increase your variance, which is fine. That is reasonable. It is better in most cases than fixed effect. When you have random effects, you change the point estimate. You change it in a way which takes this heterogeneity into account by accounting for this tau squared term in the definition of the weight.

But you compute the variance in the same way as for fixed effects. Of course, the weight is going to be different. This variance is almost always much bigger than in the fixed-effect case because you have tau squared here, and it will propagate here.

In total, it will increase this animal here [the variance]. Now you could also probably combine both approaches. You can take the random-effects weighted average here. Using this weight, you could then compute the variance of the random-effects estimate using the same logic which was applied by Tom Stanley and Chris Doucouliagos in their UWLS estimation here. The logic here is that you do not need to add anything in the denominator because the cross-country heterogeneity is already included here as tau squared in the denominator of the weight.

These are technical details, so do not be afraid if you get lost a little bit. The main point to remember is that the fixed-effect estimator in meta-analysis will typically have two basic properties. First, it will give a lot of weight to precise results. Second, your meta-analysis estimate will appear to be really precise by definition. You will have small variance, but it is not because you have little uncertainty; it is because of the assumptions that you impose for the fixed-effect model.

UWLS would give the same large weight to precise results. The point estimate is going to be the same, but your meta-estimator will not be so super precise. But you essentially say, "I am less sure it is true." That is UWLS. Now, when you do random effects, you have your estimate. You still give more weight to precise results, but the weight is not so large, depending on your estimate of heterogeneity across studies.

Your variance is typically quite large. You can imagine that your variance could be similar to UWLS, but your point estimate will often be bigger because, as we will see, I think next time or in a week, this inverse-variance weighting typically reduces the meta-analysis estimates [when there is publication bias]. The more you weight by precision, by inverse variance, the smaller the meta estimates you will get.

Estimating tau squared 00:42:19

Now, on a more technical note, there is one problem. How do you estimate tau squared, the measure of heterogeneity? You do not have a clean, nice analytical formula. You can have it here [the DerSimonian-Laird formula], but then you need to understand what Q is, and it is complicated. It is biased anyway. Most people would use a maximum likelihood approach, for example, this REML here, to jointly estimate the meta-analysis mean and also tau. I do not want to spend much time on it because it is quite technical.

It is not really central to understanding what meta-analysis is about or how to do it in practice. Just so you know, if you do, for example, random effects in R and use this metafor package, which many people do, there is going to be some sort of restricted maximum likelihood procedure which jointly estimates your point estimate and tau squared and then also gives you the variance.

It is not so simple, and you can have many different ways to compute it. That is one disadvantage of the random-effects model. It is not super clear how you should estimate tau squared. Fixed effect is completely easy. You can do it by hand easily or in Excel, just using simple formulas. On the technical side, random effects is much more demanding. You need to have maximum likelihood, so you can have trouble with convergence and so on.

You typically assume that the distribution of true effects across studies is normal. Why should it be normal? It could be a t-distribution; it could have fat tails. People almost always, or I would say always, when they talk about this, just assume it is normally distributed. This is again fine for RCTs in medicine, when you take a random sample and do your RCT, but in economics, maybe not so much.

To be able to compute tau squared, you need it [the normality assumption]. You could impose a different distribution here, but it is going to be much more complicated, of course.

You can say, for example, that the t-distribution is essentially almost the same, but it allows for fat tails. It looks very similar, depending on the parameters you use, but it could allow for some interesting cases. There are some selection models, like the Andrews and Kasy model, that we will hopefully talk about in the session on publication bias. You can choose whether you use the normal distribution assumption or the t-distribution assumption. The t-distribution is arguably more robust because it allows for outliers and fat tails. In this case, we are talking about the true effect being normally distributed across studies.

The parameter is degrees of freedom, but I'm not sure whether there are other complications that it would bring. There is definitely one more parameter. […] Even restricted maximum likelihood will have trouble converging in many cases when you don't have a huge data set. It can be done, with maybe just one more parameter, degrees of freedom.

The REML approach is typically the default in many packages because many people claim that, although it is relatively hard to compute, it is less biased than, for example, the DerSimonian-Laird procedure.

Random effects: an example 00:47:57

Before, when we talked about the fixed-effect model, we had an example in which we wanted to estimate the average score of students here at UC. You could use a fixed-effect model because people are estimating the same thing. If you want to estimate the average score at different colleges, you have UC, the University of Auckland, maybe Waikato and Otago. Each study has samples of different students from different universities. Then, of course, you know that you are estimating a similar thing, let's say the average score for New Zealand students. Different schools are going to be different, so you will assume that you have heterogeneity across schools.

It is reasonable. Again, you have the same five primary studies. Each study can have a different sample size. One study is really large and has 800 students included. It will have more weight, but now it is not going to be 50%, just 30%. In this example, the weight is going to be smaller than in the previous case because of the heterogeneity term, tau squared, which is now included in the denominator of the weight formula.

Overall, you will get an estimate that is slightly different because you weigh the studies differently. To be more explicit, you give less weight to this large study than in the previous case. Because you assume this additional heterogeneity, you will also have more uncertainty around your meta-analysis estimate here. Once again, random effects will make studies more equal. The ones that are more precise will get more weight, but not so drastically more. You will also be more conservative. I think these are quite attractive characteristics for the estimator in many situations in economics, if you don't do anything beyond these weighted averages.

Forest plot 00:50:42

Now let's talk briefly about some plots that are useful to summarize your results before you go to some more sophisticated analysis. What we were discussing here, these figures, is commonly called a forest plot. A forest plot means that you have different primary studies and their estimates, and then you show the uncertainty around these estimates, which is given by standard errors in the original studies. At the bottom, you have the overall weighted average from the meta-analysis and its confidence interval. In this case, I think it is fixed effect. It is quite useful, especially if you have just a few studies. […]

[A participant asks about studies with many estimates.] I will show you. I hope it is the next slide. This is a forest plot for a study that I think I mentioned, the effect of class size on student achievement. We have many studies, but more to the point of your question, we have many estimates per study.

It looks the same, but I will tell you how we did it. Of course, there is no procedure for what you should do when you have more estimates per study. We did the following. We focused just on the estimates that were preferred by the authors. But I think it makes sense. If you just want to visualize what the study says, you should focus on what the authors actually say: "These are our main results." When you do the meta-analysis itself, you want to include all of the other estimates too, but if you just want to visualize the main results, we focus on the preferred ones.

Still, for some studies, you have five preferred estimates, because the author can tell you, "This table contains my main results," but the table can have 10 different specifications with different control variables. The authors can be completely mute on which specific one they prefer. That makes sense. Sometimes you don't know. What we do is take the median. I told you I like medians. When I'm not sure what to do, I take the median.

Now, the median itself doesn't have a standard error. That is a problem. What should we do? One way would be to take the confidence interval around the median value. But when you take just the median of the point estimates, it could happen that the median is in the middle of the point estimates but happens to be super imprecise for some reason, or super precise.

What we do is take the median of the standard errors as well. It is hard to do it statistically 100% correctly.

This is an approximation. But in my view, it is better than just taking one particular standard error. Again, if the table has 10 columns, there could be big differences. When you take the median estimate and the median standard error, you have a good representation of what the study says is the point estimate, and what the study says about the uncertainty of this estimate. Then we can construct the forest plot.

[…] Here we did use random effects. [A participant suggests another approach.] You would first do a meta-analysis for each study separately. We also thought about this, but it is statistically even more problematic. First of all, when you do a simple meta-analysis like this, you assume that the estimates are independent. When you do it for one study, you know that the basic assumption is broken.

We have thought about it. It could be a way to summarize the study, but you would already put more weight on the more precise estimate. Right now, I think it is perhaps more problematic than just using the medians. But maybe it could be a way. We take the medians of point estimates for each study, and then we take the medians of standard errors for each study as well.

Finally, we compute the random-effects meta-analysis estimate. We use random effects here because there is obviously heterogeneity. These are studies for different countries. Some studies are for primary school; a couple of studies are for high schools. We want to be on the safer side. Even with random effects and the medians, we have quite a tight confidence interval.

It is a very small number, but the confidence interval doesn't include zero. There is a lot of precision, even though we are relatively conservative in how we measure it. You can see the weight: even the most precise studies have a weight of at most 4%. When you use random effects, you typically won't have any huge weights for individual studies. If we first used fixed effects for each study and then did an overall fixed-effect meta-analysis here, we would have some super precise study, with maybe 50% weight for this study, or definitely 30%.

I'm not sure how the mean would change, but it would probably be smaller and much more precise. This would be really tight. This is one practical example of how you can do a forest plot in the context of a meta-analysis of economics or finance data. We were specifically asked for the forest plot by the referee. Normally, we don't do it because it is hard to do when we have more estimates per study. But maybe this is not a bad way to summarize your results as well. Here, when you have 60 studies, it is hard to grasp. But when you have 20 or 30 studies, I think such a figure is a good idea. It nicely shows you what is going on, especially when you combine the preferred estimates.

Box plot 00:59:32

What is much easier to do in many meta-analyses in economics and finance is the so-called box plot. For a box plot, you take the estimates you have from each study and plot them. You plot them in a structured way. You show the median from each study. You show the 25th percentile and the 75th percentile. Then, if there are any extreme values, you show them as outliers. It shows you the dispersion of the estimates within studies, but also across studies, and you can compare them.

Maybe the forest plot, which I showed previously, is probably more informative. But you need additional assumptions to do it. You need to choose fixed or random effects, and how to do it for each study. Maybe it would be better if we did this instead of the box plot. The main difference between the forest plot and the box plot is that the box plot shows you just the estimates as they are from the studies.

It doesn't show confidence intervals for these estimates, just the point values. The forest plot shows you just one estimate for each study, which could be the median, and then the confidence interval. It also shows you some weighted average. That is what we have. You can see that it is also from the class size paper. We have both in this paper. In many journals, the editors would tell you, "Decide. Which one do you want? Pick one. There are too many figures." By the way, on the box plot, or on the forest plot, what referees commonly ask us to do is not just to list the studies in alphabetical order, but to sort them, for instance, by the age of the data.

You have the oldest studies at the top and the newest ones at the bottom. In one figure, you see dispersion within studies, across studies, and also some time trend, if there is one. You can have some milestones in terms of data age. In this case, the conclusion is that there is not much of a time trend in the data.

Maybe there is some room to improve or to combine these plots into one. We do report both, but there is a lot of information that is the same.

Some sort of combination would be nice. That is from another paper on daylight saving time, where you can see that many studies have relatively tight estimates, close to zero, but some are completely off. They have many estimates all the way to positive values, which would say that daylight saving time is actually bad for electricity consumption. There are many ways in which you can use these box plots or forest plots.

There is another one on the effect of working while you study. Again, in this case, the story would be that almost all countries show negative estimates. That is the message from the literature. These are not studies but countries. You can do the same forest or box plot for countries or different characteristics for which you would assume there could be some systematic differences.

Here, we thought it might be useful to do it for countries. We found that Germany is the only country for which studies say it actually helps you to have better grades when you work while you study.

As I was talking about it in lecture two or something, I mentioned that it is not really clear why that is. But one explanation could be that in Germany you have this long tradition of combining study and work in vocational schools. We don't focus on vocational schools; these are colleges. Even so, you could still have this tradition that students and companies are able to combine work and study in an efficient manner. Maybe that is not the case in many other countries.

For New Zealand, it is negative. It is not so large, but typically taking a part-time job will not improve your score, especially if it is a high-intensity job.

Histograms and the STAR experiment 01:06:08

In most meta-analyses, you also want to report some basic histograms to show people what is in your data before you do any sophisticated analysis. For instance, for this class-size paper, it is useful to show what happens when you have larger and smaller classes. What happens to scores in different subjects, like math, reading, writing and so on, when you move students between these classes? The message in this case was that it doesn't really matter much. It is all relatively close to zero.

You can obviously make very different points using histograms. This is the same paper, and instead of different subjects we look at different students: men, women, disadvantaged students. Again, it is a very similar story, pretty close to zero.

Sometimes it is very obvious that you have something going on there. Again, in the same paper, on the same topic, we have different ways in which people measure it. We have some RDD, regression discontinuity, instruments, and student fixed effects. They are all quite close to zero again, as in the previous cases. But you have one approach that really stands out, and that is the STAR experiment from the 80s. […] You want to compute the effect of class size, how many students you have in a class, on students' outcomes.

It's hard to do an experiment. It's very expensive. You need to call in schools and create artificially small classes and artificially large classes. In the 80s in Tennessee, the state government decided to do an experiment. "We have plenty of money, and it is going to be voluntary, but we will give plenty of money to schools that participate and create new small classes. Instead of 25 students, which was common at the time, these new classes would have 15 students. We will give you money to pay for the classes, for the new teachers, and something extra as well for your effort.

"But we want you to distribute teachers and students randomly across classes." [He explains that in practice the random assignment may not always have been kept, for example if parents asked for their children to be placed in the smaller classes.]

We had this one experiment, which shows us that children put in smaller classes perform better, substantially so. But it is just this one experiment. You have one experiment that was super expensive and huge. It involved about 330 classrooms and 6,500 students. Normally, in economics and finance, what is your gold standard evidence? RCT experiments are the gold standard. This was not perfectly controlled, but it is still an experiment.

Then you have plenty of data that you can look at. You use RDD, IV and different statistical techniques to get around omitted variable bias, selection bias and so on. If you do anything other than this one experiment in Tennessee, you will always find 0 or something very close to 0. The question is essentially: do you trust this one data set or do you trust everything else? We would say everything else.

In the paper, which was then published in the Journal of Labor Economics, the conclusion is zero. [Note, 2026: the published paper (Journal of Labor Economics; see meta-analysis.cz/class/) reports a negligible class size effect for all identification approaches except the Tennessee STAR project and for all contexts except classes of fewer than 15 students.]

Essentially, the experiment fails to replicate. We might be wrong. Maybe all of these studies really are completely off, and the only relevant piece of data was the Tennessee experiment. But it is quite unlikely, for all the reasons I have explained.

Cumulative meta-analysis 01:13:39

Something I do not do much, but maybe should do more often, is called a cumulative meta-analysis. What does it mean? You have studies published: the oldest study is study number one, then study two is published, and so on. You do a fixed-effect or random-effects meta-analysis. And you update it. First, you do not need any meta-analysis; you have just one study. Then you do a meta-analysis using two studies, then a meta-analysis using three studies.

This is not the result of the study, but the meta-analysis after three studies have been published. As you go, you typically have more and more precise results. Sometimes, maybe, you do not get more precise results if a completely different study comes in. But you gradually increase your precision over time, even though your estimates can change.

That is cumulative meta-analysis, which I would say is more relevant in medicine than in economics, especially when you have a new medication. It is promising, but you do not have enough evidence to show that it really helps the patients. You have one study, or it could be the second one also, but when you put them together, at some point your cumulative evidence is enough to convince the FDA (the Food and Drug Administration in the U.S.) to formally authorize your medicine for distribution. I think this would be the context in which you could do cumulative meta-analysis.

[A participant comments on cumulative meta-analysis.] I have heard about something like this from Mehmet Ugur from Greenwich. What I think would be really cool, now that it is easier and easier, is that if you publish your meta-analysis and have your data and code, you could easily make a simple web application, a Shiny app or something, where you could allow people to add studies. Then it would automatically run the code and give you the main results, not all of them. It would not be difficult to do. Thank you for the idea. Maybe we will do it.

We could do it with the class size paper, because it is a contentious idea. Many people feel strongly that smaller classes must be better. If it does not show in the data, it just means you do not measure it well, which could be true. It is hard to identify some effects in economics. But that would be a way to do it: "This is our paper. You have a new study. Plug it in. Let's see what happens."

Another example: in the Journal of Economic Surveys, where I am one of the editors, we could make it obligatory. We could ask people to provide the data and code in a certain way, because not everyone is familiar with Shiny apps. We could make it live and very simple, with just the main result, not all the results, because, as we will see later, you can do many things. Or we could use the main estimate corrected for publication bias. […]

01:19:08 Participant: It is very good.

01:19:11 Tomas Havranek: Cumulative meta-analysis: I think that is a very good idea. I will give some more thought to whether we could implement it. I think we could. […]

Publication bias (topic of the next lecture) 01:20:04

[…] I will talk about publication bias, which I think is the key to any good meta-analysis. If you ignore publication bias, it is essentially as if you did nothing. In the class size paper we find no publication bias, but in most cases in economics and finance, if you ignore publication bias, you are likely, not always, to really mislead the reader. [Note, 2026: the published class size paper finds little publication bias; see meta-analysis.cz/class/.] Publication bias tends to exaggerate reported results. I will explain in detail how it works, why it is a problem, and how we can correct for it, at least to some degree. […]

Discussion: multilevel models and the selection of contexts 01:21:02

[…] I did a little bit on multilevel models 15 years ago, when I was starting with meta-analysis. It might be good for these summary statistics, maybe for reasons that you know. When you go further and want to estimate regressions, the assumptions behind these multilevel models are relatively strong. It depends on which multilevel model we have in mind, but quite commonly you have random effects at different levels: random effects at the level of study, author and something like this. When you add these additional variables in your meta-regression, even when you add just standard error as a variable, you need to assume that these variables are not related to these random effects.

As an example, imagine you have panel data and want to run a regression with random effects. The main assumption is that the random effect is not correlated with the other variables, which is almost always broken. That is why, if you look at the panel data literature, random effects in econometrics are almost never used nowadays. Sometimes you see random effects in very special cases, but when you send it to a journal, referees will immediately ask, "How can you expect the assumption to hold?"

Sometimes it does, but it is hard. Here, for example, you can have different degrees of publication bias for different studies or different authors. When you do FAT-PET-PEESE using a multilevel model, this assumption could be broken because the standard error variable could be related to random effects at the different levels that you model. It is one of the reasons why Thomas Stanley does not like the application of random effects in meta-regression, such as PET-PEESE type models, which we will discuss. That is why, in most applications in economics and finance, we treat heterogeneity by adding extra control variables rather than random effects, because of this orthogonality assumption. You could maybe use something like a Hausman test to test whether it is fine.

The multilevel model solves one problem and maybe creates two more. It depends: if you do not have any additional controls, as in the models we discussed today, I agree that this could be better for simple summary statistics in many ways. You will have levels such as countries, studies, maybe authors. In more advanced applications, you run into these issues, and then it is really hard to justify.

But this is the problem with essentially all approaches. You have some advantages and some disadvantages, and you need to choose what you need in your specific context. When I want to do all the things that we commonly do in meta-analysis, it will be hard for me to justify a multilevel model in the Bayesian model averaging (BMA) case, for example. All the assumptions are really difficult for me to justify. That is why I do not do it. If you look at my older papers, we have a paper on FDI spillovers from 2011 in the Journal of International Economics. I use the multilevel estimator there for these simple PET-PEESE regressions. […]

01:29:34 Participant: [A participant asks whether publication bias also includes selecting contexts in which ChatGPT is expected to have an effect, so that studies in contexts with no expected effect are never done, rather than merely withholding specifications.]

01:30:55 Tomas Havranek: Now I get it. You have really heterogeneous effects.

01:31:00 Participant: [A participant refers to heterogeneous effects of ChatGPT in different scenarios.]

01:31:05 Tomas Havranek: People focus on the context, then maybe shift the entire distribution.

01:31:12 Participant: [A participant refers to shifting the distribution to the right.]

01:31:17 Tomas Havranek: There is no way you can do anything about this [selection of the contexts in which experiments are run].

01:31:18 Participant: [A participant says that this selection should still be discussed as publication bias, even if it cannot be fixed. Studies of Copilot and ChatGPT in the software industry may find a big effect, whereas nobody would experiment with ChatGPT for digging holes, where no effect is expected.]

01:31:44 Tomas Havranek: It is not a representative estimate for the entire economy.

Corrections

Slips of the tongue corrected in the text:

  • [00:09:09] said "cluster-robust standard errors"; the text has "heteroscedasticity-robust standard errors".
  • [00:14:54] said "proportional to sample size"; the text has "inversely proportional to sample size".
  • [00:19:49] said "1 in the denominator"; the text has "1 in the numerator".
  • [00:35:06] said "the true estimate"; the text has "the true effect".
  • [00:35:49] said "inverse variance in the denominator"; the text has "the variance in the denominator".
  • [01:14:48] said "Federal Drug Administration"; the text has "Food and Drug Administration".
  • [00:57:43] said "huge outliers in your data"; the text has "huge weights for individual studies".