Transcript. Lecture 2, Applied Meta-Analysis Examples, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.

Lecture 2. Applied Meta-Analysis Examples 00:00:00

00:02:53 Tomas Havranek: There are the guidelines I mentioned, a very short, non-technical paper on how to do a meta-analysis. Then we have a nice introduction to meta-analysis published in Nature.

It's a paper co-authored by Shinichi Nakagawa and others, and again it's a nice non-technical introduction to meta-analysis. And then we have the paper in Research Synthesis Methods, which compares the means reported in the literature and in the meta-analysis results. So again, it's nice reading as an introduction to meta-analysis. Today I want to talk about three applications of meta-analysis techniques in different fields of economics and finance. I will start with something from energy and environmental economics, maybe a little bit of lifestyle. Then we will have a finance and macroeconomics application, and we will end with a microeconomic paper on labor. […]

Let's go to lecture number two. The idea of today's session is to give you some sort of flavor of what it looks like when you do a meta-analysis and you have the final product. I will not go into many technical details, but just give you some idea of how it looks in slightly different fields. We will go through the entire papers and all their results, but we will leave the technical details for other lectures, because we will have plenty of time to talk about publication bias, p-hacking and other topics in much more detail.

Example 1: Daylight saving time 00:07:38

The first example I have is the effect of daylight saving time on energy consumption. You might have heard recently that some people want to abolish daylight saving time; it's a policy issue right now, maybe a political one as well. But this is an old paper [Havranek, Herman and Irsova, Energy Journal 2018], so it predates all of the discussion. And the question is: does daylight saving actually save energy? The original rationale behind the policy, about a hundred years ago, during the First World War, was to conserve electricity and coal, to conserve energy by expanding the amount of daylight that you actually use during the summer.

Let me put it in a different way. What should your sleep pattern be? When should you go to bed if you want to maximize the daylight hours when you are awake? If you typically sleep for eight hours, at what time should you go to sleep if you want to center your sleep on the darkest part of the day? [A participant answers: after nine.] Nine is not far from the answer. But let's suppose you sleep eight hours, as a typical person does. But at what specific time should you go to sleep if you want to center it?

[A participant answers that it depends on daylight saving time.] But let's suppose there was no daylight saving time and you would have true noon at 12 o'clock and true midnight at 12 at night. So if this was the arrangement, you would need to go to sleep at 8 p.m., because then you would sleep 4 hours before midnight and 4 hours after midnight. So your sleep would be centered on the darkest part of the day, because midnight is the darkest period. So when you are awake, you would always maximize daylight hours. But that's not how people behave.

People don't like to go to sleep at 8 p.m. I don't go to sleep at 8 p.m., and you probably don't either. Many people would say they go to sleep at 10 p.m. or maybe even after that. Anyway, to help us a little bit in the summer to make better use of daylight, we use daylight saving time. We move the clock a little bit, so instead of 8 a.m. we say it's already 9, or instead of 8 p.m. we say it's already 9 p.m. It's not a perfect correction, but in that way we go to sleep a bit sooner. It forces us to go to sleep sooner, because we think it's already 9 p.m., although by sunlight it's 8 p.m. of true time. So this is the idea behind it. You don't use as much electricity for lighting and heating and so on.

But the issue is not settled, and a lot has changed since the time the policy was first implemented in Europe in 1916 or something like that. There are many studies which estimate the effect of daylight saving: does it actually save electricity? We completely ignore all the other effects of daylight saving time, for example that we enjoy longer daylight hours in the evening.

There are also some studies which look at the effect of daylight saving on, for example, crime, on economic activity, on traffic accidents, because when it's 5 p.m. but in reality it's just 4 p.m., you have more sunlight and it's easier to drive safely. A little bit. So there are many, many different effects of daylight saving. We focus on the one which was the original reason why it was established, which is energy savings. And the way people usually measure it is they do a regression analysis.

They compare days when you have a daylight saving policy with days when you have no daylight saving policy. And especially, you look at the transition from normal time to daylight saving, and also the other way around, from daylight saving to normal time. So you typically take a couple of days before and after the change and compare them using, for example, difference in differences. You can use different simulation techniques as well, but this is the most common one.

So you have a regression of energy consumption or electricity consumption (we look just at electricity consumption) on a dummy variable which equals 1 for daylight saving time. You estimate the effect of this dummy variable. And then you have some control variables for things like weather. You can have different weather, which would also affect consumption of electricity. You can have holidays like Easter, which interfere with the changes in daylight saving time, and so on. So this is how people do it. In the literature we call them primary studies: when you collect studies for a meta-analysis, these individual studies are called primary studies. It's just a matter of terminology.

No consensus in the literature 00:14:43

To get some motivation, when I want to introduce readers to what I do, I like to make a graphical summary. In this case, we show on the horizontal axis the publication year of the study, and on the vertical axis we have the actual effect of DST. What you see is that more recent studies have more dispersion in the reported effects, which is disappointing from the position of a reader who would just like to look at the newest studies, like here. You have plenty of dispersion. It's hard to take stock of the literature without any additional work, some additional meta-analysis.

Here we have zero, and it should be negative if it works. A negative value means that on the days when you have the daylight saving policy, you have lower consumption of electricity than in other parts of the year. So if it's negative, it means you save energy. But you can see that you also have quite a few estimates which are actually positive, which means that you consume more energy when the daylight saving policy is in place.

Why could this be? Some people say it's because of more use of air conditioning in recent years, which sometimes interferes with the original heating savings, which are intended. And also, just after the change, it is suddenly really dark in the morning. So you use more electricity for lighting, heating and so on. So there could be unintended effects as well, which might partly explain the puzzling positive effects. And you can see there is no trend.

All you have in the data, if you just look at it, are estimates for different years. You can see there is increasing disagreement among researchers about what the actual effect of daylight saving policy is. So these are perfect conditions for a meta-analysis, because people in the field disagree. We would like to see first what the average effect is after correction for publication bias and all the other issues.

And second, why do people disagree? What are the underlying reasons for this heterogeneity? Is it because they look at different countries and different countries have different effects? It could be a good reason if you compare a daylight saving policy in a country like Jordan, which is in the subtropical zone, with one in Norway. You will have plenty of differences in how it affects society and energy consumption. And there is a good reason why you have no daylight saving policy in tropical countries: when you are on the equator, you have no changes in daylight hours during the year. So there is no need to adjust [the clock]. But still, in some subtropical countries, you have a daylight saving policy.

Journals report smaller savings 00:18:36

Now, we look at the distribution of empirical estimates among working papers and published journal articles. And it's quite similar. We can see a little bit bigger, more negative effects in working papers compared to journal estimates. But it's not large.

Publication bias? Probably not 00:19:13

Next, we do a funnel plot, which I will talk about in much more detail during further lectures, but this is just a brief introduction. The funnel plot is a scatter plot where we have the estimates on the horizontal axis and their precision on the vertical axis. Precision means one over the standard error of the estimate. It's called a funnel plot because it should resemble something like an inverted funnel, which in this case you can see a little bit there. So the top of the funnel is narrow when you have the most precise estimates. As you decrease precision, you get more dispersion at the bottom of the funnel. The stem of the funnel is pretty narrow. If there is no publication bias, the funnel should be symmetrical. Again, I will explain it later in much more detail.

For now, let's just say it should be symmetrical, because all imprecise estimates, both positive and negative, should have the same likelihood of being published. And essentially, in meta-analysis, we often look at the top of the funnel. So we put more weight on the more precise estimates, which are at the top and are actually quite close to the mean effect of around minus 0.3%, that is, savings of about 0.3% due to DST. In this paper we don't have so many estimates, as you can see, but it looks relatively symmetric. When we just eyeball it, it's at least not heavily asymmetrical.

Funnel asymmetry tests 00:21:10

Now we do the same thing as in the figure, but we just use a simple statistic to measure how asymmetrical it is. So we essentially do what is called either an Egger regression or a funnel asymmetry test, and we regress the reported effects, the estimates, on standard errors. So if I go back, essentially you take the horizontal axis, you switch it to the vertical axis, and then you invert the other axis to get, instead of precision, standard errors.

So the slope coefficient will measure the amount of asymmetry in the literature. That's one thing. The other thing is that the intercept will measure the top of the funnel. So under some assumptions, that's going to be the true effect beyond any publication bias. Again, I will discuss it in much more detail. This is just for you to have a flavor before we get into the details.

Now, this is the basic regression we want to run, and there are many ways to do it. […] You have the simple regression setup, and then you can estimate it in many ways. So it's a standard econometrics exercise. You can either just do OLS, a regression by ordinary least squares, or you can use fixed effects, and I mean fixed effects in the econometric panel data sense, by including dummies for different studies. Because we have more estimates per study, we can include dummies for studies to get rid of individual study-level effects like study quality, some study-level biases. We get rid of them.

Or we can use between effects, which means we use study-level means. This way we just take into account between-study heterogeneity, between-study variation, and we ignore variation within studies. Or we can give the same weight to each country, because in the data set we have data for different countries, and of course some countries, like the United States, have many more observations than others, for example New Zealand.

Then there are some complications, but I will not talk about them right now. And what you can see is that, no matter what you do, you always get a statistically insignificant estimate of publication bias, the slope coefficient. So the asymmetry is close to zero, or it's statistically insignificant. So we find little evidence. It's negative, so you might say that if there is any asymmetry, it's a bit to the negative side. The negative side is a bit heavier than the other side. But the asymmetry is slight, and that's also what you see in the empirical results. So you have negative estimates of publication bias, that is, selection towards negative estimates, but it's not statistically significant.

The true effect, the mean corrected for publication bias, doesn't really change much from the simple mean, because there is not much publication bias there in the first place. So you have something like 0.3% savings, which means that with daylight saving time you save 0.3% of your electricity consumption. That is not much, but there are some savings on average. So this was the mean, the average.

Country heterogeneity 00:26:11

But we also want to look at why the estimates differ so much. For example, we can look at cross-country heterogeneity. You can see that for different countries we have different estimates. Jordan, as I mentioned, is subtropical, and for Jordan we have no energy savings at all; it is more like the opposite. New Zealand is an outlier here: there are big savings for New Zealand, for some reason. That could either be a true result, since for New Zealand it is more efficient due to the location, the latitude and so on, or it could be just a coincidence, because we have just one estimate.

Method heterogeneity (just studies for the US) 00:27:02

Now, we also definitely have plenty of method heterogeneity, which means that even for one country, the same country, you have different results. This has to be explained by the way people measure it, because the data are essentially the same. This is an example for the United States. What we do is collect the characteristics of these estimates, the ways in which they can differ. For instance, what type of data they use and what type of data frequency: is it hourly data, or do people use a lower frequency? We also use data periods, which means the age of the data, how new they are.

We control for maximum daylight hours, which essentially means how much sunlight you have on the longest day during the year. It is something similar to the distance from the equator, to account for the local climate of the country. Then we have methodology controls: is it a regression or a simulation? Do people use, instead of OLS, difference-in-differences, and so on? These are the typical variables you would use in most meta-analyses on economic topics. We also control for publication characteristics, such as whether the paper was published in a journal, how many citations it has, and how prestigious the journal is.

Bayesian model averaging 00:29:10

We put it all together and try to see which of these characteristics are important for the reported effects of daylight saving time. I will leave the Bayesian model averaging, the details, for session number seven or eight or something, but I will just briefly tell you what it is about. The figure that you see has many different regression models in different columns: each column is a regression model. The colors denote the signs of the regression parameters. If it's red, it means a negative sign; if it's blue, it means a positive sign. The width of the column shows how good the model is: technically the posterior model probability, but you can think of it as something like R squared or adjusted R squared.

Results 00:30:13

Again, I will explain the details later on. The best model, the first model, has five variables: four of them have a negative coefficient, one has a positive coefficient. What does it mean? Estimates published in journals with a bigger impact factor are typically more positive or less negative, which means they find smaller savings from daylight saving policy, which you can interpret in two ways. One way is that there could be more publication bias in these better journals, but we don't find evidence for any publication bias.

Or these studies published in better journals are of higher quality, which could be quite plausible. Then we have a negative effect for maximum daylight hours, which means that in countries which are farther away from the equator, you have more savings. You have more negative coefficients, which means more savings. So in Norway, which is far away from the equator, you will have better effects, more savings from daylight saving time, than in Jordan, which is relatively close to the equator. We can also see that it matters what kind of methodology you use, whether simulation, regression or difference-in-differences, and what type of data frequency you use. This is an example of what a heterogeneity section in a meta-analysis can look like.

In summary, in this paper we find that if you take the literature and summarize it, on average there seem to be some savings of energy related to daylight saving time policy, but they are not large, around 0.3%. [Note, 2026: in the published paper the simple mean is 0.34% savings, but the best practice estimate, which also corrects for data and method choices, is 0.01%, essentially zero; see meta-analysis.cz/dst/.] We find little or no publication bias. We also find that the methodology can be important, whether you use simulation or regression, and it is especially important for which countries you measure it. If you have a country which is in a temperate zone or far away from the equator, you are more likely to get some significant, some substantial results.

As for all the papers that we do, we always have an appendix with data and codes, so in case you want to replicate it, you can find it there. I will just show you the website, meta-analysis.cz. It mainly has a list of different meta-analyses with data and codes. This is just for you to know that these data are there. If you want to look at some examples, we have featured papers on the website.

[He shows the featured papers on meta-analysis.cz, which he considers good examples, and the list of the newest papers.]

Example 2: Interest rates and house prices 00:34:18

Now let's go to the second example. The second example is from central banking, so it is related to my talk yesterday. It is macroeconomics and finance. We look at what happens if the central bank increases interest rates by one percentage point: what happens to house prices, home prices, real estate prices?

In many countries house prices are outside of the central bank's mandate. So again, central banks work mostly by changing the interest rates. That is how they affect the economy. So that is the key question for a central bank: what happens to the economy if I increase my interest rate? And we ask this question for house prices. Now, the literature on house prices is a bit different from the typical literature we collect for meta-analysis, because these studies don't report numbers, but they report figures.

They report something like this, which is an impulse response. I think I talked about it yesterday. What do these impulse responses mean? It means what happens to some variable, in this case house prices, if you increase interest rates by one percentage point, and you have time on the horizontal axis. So the central bank increases interest rates, and house prices go down. Why? Well, because mortgages are more expensive, so it is harder to buy a house, so typically prices go down, a little bit. But after some time the shock fades away and it goes back to the original state. That is how it should be. So this is a typical figure that these empirical studies would report.

Converting graphs to data 00:37:01

And now you want to do a meta-analysis because, for example, your boss in the central bank asks you, "Okay, so we have these different studies, so what is the true effect, the mean effect? Which study should I trust? I have 10 different studies." So you can do a meta-analysis, but you don't have the numerical estimates. You don't have regression tables. You have these figures. So what do you do? Well, you try to somehow measure it.

By the way, this figure shows the mean response we collected, but I use it for now as an example of what the individual studies report. What we do in the meta-analysis is download the papers and measure pixel coordinates, because all you have are these figures and you need to convert figures to numbers. We use a tool called WebPlotDigitizer, which allows us to measure the coordinates more precisely. So what is the number here, here and here? And the confidence intervals. It is pretty useful, and it makes the work relatively speedy. This is what is reported in the paper, and this is what we code for the meta-analysis.

So we convert figures to numbers. We did the paper a couple of years ago. I think nowadays you could probably use some AI tools to help you do it even faster. Again, I will talk in more detail about data collection next week. I will also talk about literature search next week, but again, this is just a brief taste of how it is done in practice.

At the very beginning of any meta-analysis, when you know you want to do a meta-analysis on monetary policy and house prices, first you start with a literature search. Typically you use something like Google Scholar, because Google Scholar is universal in the way that it includes almost all scientific studies. Another benefit is that it goes through not just the title but the actual full text of these studies. You have many databases, like EconLit, Scopus, Web of Science, many different ones, but typically they just look at keywords, which is often not enough.

You use some keywords, which we will discuss next time, and then you have a large number of studies which you could use. But you typically exclude the studies that your keywords found but that are obviously theoretical, with no empirical work. Then you are left with some studies for which you have good reason to believe they will be useful. You download them, you read them. Again, we will talk about it in detail, but this is called a PRISMA diagram. In any new meta-analysis, you will need to have something like a PRISMA diagram to show people how you collected the data, how you searched for studies. This is the way to do it.

In the PRISMA diagram we have in total 37 studies. Because of the nature of the data which I described, we need to decide not just how to convert these figures to numbers, but also which horizons to focus on, because the effect is different at a 5 quarter horizon than at a 10 quarter horizon.

Histograms for different horizons 00:41:29

What we do is select a couple of horizons as representative of short-run, medium-run and long-run effects. We don't have the entire function, but we have some points on the way. So we have 1 quarter, 2 quarters, 4 quarters, 8 quarters and 12 quarters, and we collect the estimates of the effect of interest rates on house prices for each of these horizons. And these are histograms. What you can see is that the histograms are relatively asymmetrical. These are not funnel plots yet, just histograms. But you can see that the extreme negative estimates are probably more likely to be published than positive ones, especially for the longest horizon, the long run.

Publication bias for longer horizons? 00:42:32

Now, this is quite interesting: you see more asymmetry as you go into the long run. And the reason is the price puzzle. There are a lot of papers on what is called the price puzzle, which means that after the central bank increases interest rates, in the short run, in a couple of quarters, prices tend to increase as well, which is not very intuitive, because we would like prices to go down. That is why the central bank increases interest rates: to cool prices, to cool inflation, to cool the economy. But in many empirical papers you see this counterintuitive, puzzling increase of prices.

That is why it is called the price puzzle. We also often have it in the case of house prices, not just prices in general. This is really well known: almost half of the studies have some price puzzle in the short run. This is quite typical, and many people do publish price puzzle results in the short run. That is universally known and accepted. But in the long run there should really be the intuitive response. You increase interest rates, you should have lower prices. That is very strong intuition. So you can have some price puzzle in the short run, you can have these positive estimates here in the short run, but not so much in the long run.

That is essentially what you see. People do publish a significant proportion of positive results for the short run, but a very small one for the long run. So this could be consistent with publication bias, consistent with the theory I just described to you. So why could there be a price puzzle in the short run? Let's go a little bit into macroeconomic theory. One potential explanation is that when you increase interest rates, for companies it also increases their costs.

That is because they have loans, and if they have loans with a flexible interest rate, they immediately have to pay more on repayments, on the interest costs. So they have bigger costs, and they may pass the costs to the clients, to customers. That is one of the reasons why we could have a price puzzle in the short run. So the bottom line is that in the short run, positive and negative estimates are acceptable. In the medium and long run, you should really just have negative estimates, and you should have the intuitive effect of monetary policy.

Funnel asymmetry test (Card & Krueger) 00:45:58

So then we do test for publication bias using the very same approach as in the previous paper. I will not repeat all of the motivation, because it is the same. Again, we have this regression of estimates on standard errors, but we commonly use a weighted least squares version of the estimator to get rid of heteroscedasticity. This equation here is heteroscedastic because on the left-hand side you have estimates, and on the right-hand side you have standard errors, which also measure the dispersion in the left-hand side. But don't worry about it; again, I will explain in much more detail later on.

Corroborated by funnel asymmetry tests 00:46:58

And here we have funnel plots. In the funnel plots, you can see the same story as the one with the histograms. In the short run, there is not much asymmetry. As you go to the medium and long run, you have more and more asymmetry. So that would be a nice story of publication bias: more publication bias when you have a strong theory against positive results in the long run.

But the problem is, when we test this publication bias for different horizons, statistically we do not really find much difference. So we find a similar amount of publication bias across all horizons: publication bias against positive results and against insignificant results. That is what the table says. I will not go through every single estimate, because it is essentially pretty much similar.

Nonlinear tests also imply bias 00:48:34

This publication bias correction is the one that we most commonly use in economics and finance, but it is very simple. It is a brutally simple regression: just one regressor, no non-linearity, nothing. It often works really well, but in practice we will also use non-linear estimators, which are sometimes better in some situations and less biased, and so on. I will not talk much about them today, but I will just mention that the results we get from them are pretty similar to the linear estimation that I showed you before. It is pretty much the same in this case.

Mean impulse response beyond publication bias 00:49:28

What we do next is take these results from the publication bias analysis, the corrected estimates, the intercepts from the meta-regression, and use them to construct an impulse response that is corrected for publication bias. When we started talking about the paper, we had the mean: we just collected the data and took the mean from the entire literature. That is the figure here. Then we did the publication bias analysis and we have the result corrected for publication bias. It looks a bit similar, but if you look at the vertical axis, here we have a maximum effect of around minus 0.2 to minus 0.25. In the uncorrected analysis, we had about minus 1.2.

So the corrected effect is much, much smaller. If you are a central banker, it tells you that you need a more aggressive increase in the interest rate to be able to bring down house prices. So that is quite policy relevant.

Cross-country heterogeneity 00:51:11

But that is not the end. It is just publication bias. We also want to take a look at heterogeneity, differences in context. So again, the natural thing to do is to look at cross-country differences. You have different people doing these kinds of regressions, these kinds of estimations, for different countries, so we can have a look at how it looks for different countries. It looks different, but you can also see the basic shape.

These are just averages, just means of what people report for the US, Italy and so on. They are different, of course, but the basic shape is surprisingly similar. You always have a little bit of effect in the short run, and then in the medium run, or maybe sometimes in the long run, like for Switzerland, it gets stronger. Then often at the end it dissipates; it gets weaker again. So it is nice to show that despite all of these differences, probably the underlying theory behind monetary policy is not so wrong, because for very different countries, from different studies, you get essentially a very similar story in qualitative terms. Of course, in terms of numbers, the responses are different. For the United Kingdom, you have the maximum effect of minus 2.5, while for Germany it is minus 0.6. [Note, 2026: in the published paper the maximum decrease in house prices is 2.2% for the United Kingdom and 0.6% for Germany; see meta-analysis.cz/house_prices/.]

Estimation context 00:52:53

What we do is collect plenty of information on the context in which these estimates are derived, that is, what people actually do when they estimate the effect of interest rates on house prices. There are many choices you can make, again, in terms of data, frequency, how much data you have, how new the data are, and what variables you include in your model. One small digression on the models behind these estimates.

If you want to estimate the effect of interest rates on house prices, you typically use something called vector autoregression. Vector autoregression is a set of time series variables that you put together in one system using lagged values, and you try to get rid of endogeneity in a way. Chris Sims was the one who introduced this in the 1980s, and then he got a Nobel Prize for it in 2011. For instance, when you estimate this vector autoregression, VAR, you can either include information on credit and money, or you can exclude it. Some people don't include it, some people do. When you estimate it, you can use different econometric techniques. You can use Bayesian vector autoregression, which has some benefits in terms of efficiency and so on. You can do sign restrictions.

Have you heard about sign restrictions in econometrics? Sign restrictions mean that, a priori, you say, "I know from theory that the effect of interest rates on house prices needs to be negative." And you essentially institute publication bias ex ante. The upside is that you essentially cut off one half of the model space. So when you do it for a couple of variables, it helps you identify the model more easily, because vector autoregression uses a lot of parameters. And for some countries you do not have a long data series, for some Eastern European countries, for example. So you need to increase your efficiency, which you can do, for example, by using sign restrictions.

But that is a small digression. We also control again for the number of citations, impact and publication status. And then we include country characteristics, which is important because it measures the actual context, the actual economic structural characteristics that can affect what people in central banks call the transmission of monetary policy, and what we should expect from interest rate changes. Especially in this case, we need to include the size of the mortgage market, so mortgage-to-GDP, because in some countries, even in developed countries like Germany, you still have mortgages, but a substantial number of deals on the real estate market are not financed by mortgages.

People can just buy houses from their own savings. When this is the case, when you do not need a mortgage to buy a house, then you should not expect an increase in the interest rate to decrease house prices, because it does not affect your ability to buy a house. So we include this cross-country measure to be able to say how important it is in explaining differences between, let's say, Germany and the United Kingdom in the strength of transmission. We also include the percentage of loans that are floating.

In the US, typically, when you buy a house using a mortgage, you have a 30-year fixation. So your interest rate stays the same for 30 years. In Czechia, in my country, the mortgage rate gets reset every five years, or even every two years, and you get a different rate. In some countries, it is completely floating, so every month you have a different interest rate, related to the rate that is set by the central bank. This could have huge implications for the transmission of monetary policy, because if you have the fixation for a long time period, then the effect is much smaller. It does not affect people who already own a house, and so on.

Bayesian model averaging 00:58:49

I think that is enough. Again, we do this Bayesian model averaging, which is the same as before, but now we have many, many more variables there. What is interesting to see here is that the publication bias variable, the standard error, is the most important one, even after you include so many of them. You do multivariate FAT-PET, multivariate meta-regression, with so many variables, and you still have the standard error as the most important one. So publication bias is important, even when you control for this context.

The size of the mortgage market is also important, which is intuitive. If you have no mortgages, then you should not expect monetary policy to have any effects on the prices. The effect is also stronger, more negative, if you have really high house prices for some time: that is the variable prolonged high house prices. It is also stronger when the price-to-income ratio is high, which means that houses are expensive relative to incomes. You also have stronger monetary policy effects if the yield curve is flat, but that is really technical, financial, so I will skip it.

But generally what these results tell us is that it is not just about the mean or publication bias: we can tell a story to a central banker on top of the primary studies which already exist. So I can tell a story about when monetary policy has bigger effects, not just if it does. When you have countries with bigger mortgage markets, like the UK, you can expect bigger effects than in Germany. [A participant asks a question about the figure.] Again, do not worry, I will talk about it later in much more detail, but it is a good question.

Technically, the width of a column is the posterior model probability, but in practice you can understand it this way: the thicker the column, the better the model. These are different models in the different columns. Each model has a different performance, which you can measure, for example, by R squared. This is Bayesian, so it is slightly different, but you can imagine it as R squared. If you have a model with a high R squared, it will be on the left here, and it will have a high score. This is the percentage of the total probability of all the models.

This is because you can run many different models with different combinations of these variables. That is what Bayesian model averaging actually does. And then it ranks the models according to something like R squared. The horizontal axis measures the goodness of these models. It shows you that this model is the best one, but it has only 0.3% of the entire goodness of all the models. In the previous example, the previous paper, we had many fewer variables, which means there are still thousands of combinations, but not millions. So the first model, the best one, has 25% of the entire possible goodness.

In the second case, because there are so many models, even the best model is just a small slice of the entire model space, the entire model probability. Again, do not worry, I will talk in much more detail about what it means later.

Best practice 01:04:13

This is a bit similar to what I showed you yesterday, when I talked about the implied estimates for labor supply elasticity. Here we do the same kind of thing for the effects of interest rates on house prices.

We try to take the Bayesian model averaging estimates, and we try to compute fitted values from this estimation, from this regression, conditional on some selected issues. For instance, we say that we do not want any publication bias there, so we plug in zero for the standard error. We plug in the sample maximum for the data year and for the size of the data set. So essentially we compute the mean value of the reported estimate, conditional on zero publication bias, and conditional on better approaches which are identified in the literature. Again, we will have this discussion much later; this is just a taste. And we do it for different countries, but I will just show you the figure, how it looks on average.

Impulse response implied by best practice 01:05:37

So again, we have this hump-shaped estimate. In this case, it is difficult to construct confidence intervals, so that is why we have just one line. But that would be our best guess regarding the mean effect of interest rates on house prices.

Results 01:06:00

And this is the summary. We have 37 studies and 237 impulse responses, figures like these that we took, and we put them together. The main result is that yes, monetary policy works, so it brings house prices down. Typically, the maximum effect is after two years and it is a little bit more than 1%. There is publication bias, but methodology is also important. And this transmission, this effect, is stronger for countries which have bigger mortgage markets. Typically also later in the economic cycle, when you have high house prices for a longer time, a flat yield curve and so on. Again, we have data and codes online.

[A participant asks a question about the results.]

Yes, you had the publication bias correction, but on the other side you have some methodology issues. So actually, the best-practice estimate is not so far away from the original mean.

Example 3: Working while in school 01:07:19

Now the final example I have, which could be interesting for some of you, is: what is the effect if you work while at school? You study, but you can also take a job. So what is the effect on your grades, on your performance at school, if you also work? There is a large literature on this. In some countries it is more common than in other countries. In the US, many students work at high school. In Europe, it is pretty common to work when you are in college. So you may ask, how does it affect my grades or my study performance?

This is interesting, and it is not easy to compute, because there is the endogeneity problem. And what do I mean by the endogeneity problem? Some people, probably all of you, are really clever and capable. With these characteristics, you could take a job on the side and still have good grades, compared to other people who do not work and have bad grades at the same time.

Do you know what I mean by endogenous? It is about your capabilities, essentially, which are hard to observe for other people when you just do a regression. If you just regress your grades on a dummy variable, which is one if you work, you can get a positive coefficient, but it would not be a causal effect that working makes you better at school. It could just capture that you are more capable than other people, for example.

It is complicated, but you want the causal estimate. And it is unlikely that it will be positive. If you work, it takes some of your time, so most likely you will be able to focus a little bit less on your studies. Maybe it is not a big effect, but probably it is going to be a negative effect. In theory, it is possible that if you work, it can make you better at school, for instance, if you learn to be more efficient. In practice, it is a little bit hard to imagine, but it could.

What we do is collect estimates from 69 studies, almost 900 estimates. And again, we do a meta-analysis.

The way we measure it is that we use partial correlation coefficients. I will talk more about effect sizes later on, but the reason is that these different studies in the literature use different units, because they measure performance in different ways. For example, some people use grades, some people use graduation rates, some people use the likelihood of dropping out of college or university, and so on. So you have different measures, and you cannot just take what people report in their papers. You have to translate it to a common metric. There are many common metrics you can use, but a very simple one used in economics is a partial correlation. I will show you later how it is computed.

Essentially, you take the t-statistic of each regression coefficient and the degrees of freedom and convert them into something which can be compared across all of these studies. It creates some statistical problems, which we will talk about as well in quite some detail later on. But for now, this is just to let you know what the PCC, the partial correlation coefficient, is. So again, if you plot the median estimates from each study, with time on the horizontal axis and the size of the estimates on the vertical axis, you see that there is no clear trend. And if anything, you again get more dispersion over time. When you get closer to the present, you get much more disagreement among the studies. So a meta-analysis is useful to try to bring some clarity.

Negative for all countries except Germany 01:13:01

We were talking about whether it should be positive or negative, and I was trying to say that negative is what you would expect in most scenarios, in most contexts. That is also what you see in the literature. For most countries, working while in school has a negative effect on your performance. The only exception is Germany. Germany is famous in Europe because of the vocational school system. Many of the schools that we would call high schools here are actually practical schools in which you spend one week in the classroom and the next week you have an internship in a company. That is vocational school in Germany.

We exclude these estimates because they would be for vocational schools, and that is a very different type of education. It is not like normal working while in school, but it is part of the curriculum, actually. It is part of the studies. We keep only the normal schools in Germany. But even for these normal schools, you still see a positive effect, and that is surprising. We do not really know why this could be the case, and you will also see it in the final results.

It survives correction for publication bias, it survives correction for heterogeneity, and so on. One explanation could be that these are not vocational schools, but you still have this great example of vocational schools in Germany, so maybe people are used to it. When you go to university, you already know how to combine work and study well, and it can even make you a better student in some ways. So Germany is interesting.

Work intensity clearly matters 01:15:18

Now we have a couple of histograms which I think are quite revealing. Here, on the horizontal axis, we again have the size of the effect: how much does it hurt or help if you work while in school, in terms of your study performance? And we have two comparisons. We have red if you work a lot (I forgot the actual threshold, but I think it was 20 hours a week). And the gray histogram is if you work a little bit, so fewer than 20 hours a week. [Note, 2026: the published paper defines low-intensity work as fewer than 15 hours a week and high-intensity work as more than 30; see meta-analysis.cz/students/.] If you work a little bit, the effect is really close to zero on average. If you work a lot, well, you can see it can have a relatively significant effect on your performance. That makes sense: if you work a lot, you will have much less time for school. So this is a good sanity check that this is what you need to have there, otherwise it would be really weird.

Endogeneity control is important 01:16:28

This is more interesting. Here we compare studies, or estimates to be more precise: estimates which ignore endogeneity, the thing I mentioned before about ability, and studies which take endogeneity into account. How can you take endogeneity into account? Well, for example, if you have data on IQ for the students, which typically you do not. But, for instance, in Finland, all men have to do compulsory military service. I am pretty sure it is still the case.

And when you enter military service, they force you to take an IQ test. So you have data for most Finnish males on IQ, and some people can use Finnish data sets to control for ability. So anyway, it is a nice data set for millions of people, which the Finnish government has, and sometimes people can use it. Or if you do not have data on IQ, you can use some sort of instrumental variables, but it is always difficult to find good instruments. Or you can use some twin studies.

You have data on twins: one of the twins is not working, the other is working. If they are identical twins, they will have the same IQ, essentially. So you have a nice quasi-experiment. So there are many clever ways in which you can do it. But the point is, if you ignore endogeneity, you typically have more positive estimates than when you take it into account. Why? Well, it makes perfect sense, because if you ignore endogeneity, you will have these strange things.

Clever students can work and have good grades at the same time. That is the endogeneity problem. If you ignore it, you will get these positive effects for students who are good at it. And the other way round for students of lower ability, who maybe do not work as much on average: they would have worse grades and worse performance. So it is not super strong, but you can see that the red one is shifted to the right and the blue one is shifted to the left. That is pretty interesting.

Little publication bias on average 01:20:22

Now, publication bias. When we look at publication bias on average, there seems to be not much of it. There is a little bit: maybe you can also say that the left-hand side is heavier, so there is some publication bias towards negative results, but definitely it is not strong.

No bias in studies addressing endogeneity 01:20:46

The key thing, going back to this distinction between the two histograms, is that we try to test for publication bias separately for studies which do take endogeneity into account and studies which ignore it. Why? Well, the reason is that if you ignore endogeneity, so these are the red estimates, you have a bigger chance of getting positive results. The positive results are the ones which are not very intuitive: if you work more, you also get better performance at school. That is not what you want if you are in this field. That is hard to explain and to justify.

So you do not want it. A researcher who gets these results might feel some pressure, maybe to try again, use a different control variable, use a different specification, do a little bit of p-hacking. For now, let's just call it all publication bias. Later we will distinguish between p-hacking and publication bias. Our intuition is that you would expect more selective reporting, more publication bias, if you ignore endogeneity.

Strong bias in studies ignoring endogeneity 01:22:16

So these estimates would be doubly problematic, and that is actually what we find. If we look at the studies which do take endogeneity into account, we find really little publication bias. But if you look at the other studies, we find quite strong publication bias. So there is an interesting interplay between publication bias on one side and identification, endogeneity, on the other side. It tells you something about how researchers can behave in many ways. It does not have to be just that you do publication bias, you do heterogeneity, and that is all. You can use different subsamples, you can separate them.

Model uncertainty; endogeneity control is key 01:23:20

Again, we look at the heterogeneity in the same way as in the two previous examples. Again, we have many variables, but you can see that at the top this time is Germany. So the positive effect in Germany survives all of these other adjustments. You also see the interaction between publication bias and endogeneity almost at the top. So that again survives the addition of these context variables, and it is important.

And then we have many other things. If you do just simple OLS, that means you ignore endogeneity, you tend to have bigger estimates in the sense of more positive, less negative, closer to zero or even positive. If you control for IQ, you tend to have more negative estimates, and so on. There are many, many avenues which you can explore. If you work a lot, in high-intensity employment, your grades will be worse. Again, that makes sense.

Closer look: work intensity, measurement 01:24:42

What we can also do in this Bayesian model averaging exercise is look at the distribution of the coefficients for each variable. What I mean is that when you do Bayesian model averaging, which we will discuss in technical terms later, you run many, many different regressions. These are the different columns, the vertical lines, different models. If it is blank, it means the variable is not included. So you run many regressions, and you get many regression parameters for each variable. So you have a distribution of these coefficients, and that is what we can also plot like this. So we can take, for example, the variable high-intensity employment and take the coefficients from the different regressions from this Bayesian model averaging, put them on the horizontal axis, and have a density on the vertical one.

And you can see that in almost any model we run, we have a negative coefficient on high-intensity employment. So that is a Bayesian analogy of statistical significance, or you can use it in a similar way. Low intensity is the opposite, which makes sense. We also have the variable choices. This means that the primary study looks not at grades or graduation rates, but at choices, whether to continue with studies. So if you are in high school and you work on the side, does it affect the likelihood that you will apply for college? That is the variable choices. It is negative again in almost all contexts, which means that your grades are not as much affected by your working as your decision to continue with your education is. So that is one important aspect. We can do the same kind of analysis for different variables, but I will just skip it.

Best practice estimates 01:27:11

Again, you can compute best practice implied estimates for different countries. So here what is interesting is Germany. We have a positive effect, but it is not statistically significant. So even in Germany, we do not find a strong positive effect. Overall, the best-practice estimate is almost statistically significant, but the issue is that it is -0.04 in terms of partial correlation.

Results 01:27:51

An effect of 0.04 is not much. It is very small. So what are the takeaways of the paper? First, what we find, and what we think is most interesting in terms of meta-analysis here, is that you have this nice interplay between publication bias and endogeneity. Second, you have a negative effect of working while in school on your study outcomes.

But in economic terms, it is really, really essentially nothing, except if you really work a lot and if what you look at are choices, like whether to continue with school. Quite often what happens is that people start to work, they find they can earn quite a lot of money, maybe they do not always need a degree, so they are like Steve Jobs or Bill Gates and they decide, "I will just work. I do not care about my college." That shows up in the data. It is not just a story; it actually shows up there. Again, we have data and code online. And that is it. That is the end of lecture two.

Now, do you have any questions at this point? Both days were essentially an extended introduction: why we need meta-analysis, how a meta-analysis could look. In the remainder of the course, we will look into how specifically we can do it, and what we should do if we want to do a good meta-analysis or a good policy report for a boss at work. If you did not fully catch some details, do not worry: we will talk about it, and we have plenty of time to address any questions you would have.

[…] And I am looking forward to the next section, which will be on literature search: how you search for studies, what kind of studies you can collect, how you do it. Any questions?

01:30:32 Participant: [A participant asks how to code the studies.]

01:30:50 Tomas Havranek: Sometimes it is hard to code. For instance, you can have a study which uses OLS, but it can actually be fine because the data are of such high quality that you do not have any sample selection issues or any endogeneity problem in this particular case. So sometimes it is very difficult to code it. Coding OLS versus IV is a crude way to make this distinction.

You can also have an IV study which is really bad, much worse than OLS, if you have a weak instrument. So this is difficult. Like all of the choices when we collect data in a meta-analysis, you need to make subjective decisions at some point, to some degree. Sometimes more, sometimes less. But it is an integral part, I would say, of any empirical work. You always have to make these subjective decisions at some point. Thank you. […]

Corrections

Slips of the tongue corrected in the text:

  • [00:07:38] said "after the First World War"; the text has "during the First World War".
  • [00:19:13] said "If there is publication bias"; the text has "If there is no publication bias".
  • [00:22:02] said "beta zero, the slope"; the text has "the slope coefficient".
  • [00:57:47] said "20-year fixation"; the text has "30-year fixation".
  • [01:27:11] said "0.04"; the text has "-0.04".