Transcript. Lecture 9, Heterogeneity, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.

Lecture 9. Heterogeneity 00:00:00

00:03:07 Tomas Havranek: We will talk about underlying differences in the primary studies when we do meta-analysis. We call it heterogeneity, and we care about it for at least three reasons. Reason number one: if you do just a simple, quick meta-analysis, you need to choose between the fixed and random effects, which we talked about two weeks ago. Suppose there is little difference among the primary studies and they all really look at the same thing, the same drug or something similar.

00:03:51 Participant: It should be the same drug in the same country.

00:03:57 Tomas Havranek: If you really have very little or no underlying heterogeneity, then you will have more precision. But you are unlikely to see such homogeneous literatures in practice. So if your goal is just to produce this weighted average, which you would typically do in maybe medical research, maybe also for some experiments in economics, you will choose the random effects version, which allows for heterogeneity, as we saw and as we will also remind ourselves today. So that is reason number one why you care about heterogeneity.

Reason number two: sometimes, especially when we have a lot of data, as we often do in a meta-analysis of an economics or finance topic, you want to go beyond just doing a weighted average, assuming heterogeneity. You want to model heterogeneity explicitly. You want to see if some specific studies, specific data or specific methodology produce results which are systematically different from other studies and other contexts.

It might be useful, especially for the calibration in macroeconomic models, like in central banks, that I talked a lot about during my first session. Sometimes you want to calibrate your model, for example, for a small country like New Zealand, but you do not have enough estimates of this parameter for New Zealand. So you want to rely on data for different countries. But because many of these estimates are for, let's say, big countries like the US or European ones, you would like to see if, for example, trade openness plays a role, or country size, level of development and so on.

If you have these systematic patterns, you can adjust your meta-analysis mean estimate to correspond better to countries like New Zealand, even though it is not a perfect match. So the second reason is that you want to model heterogeneity explicitly for many different purposes, or to see if it matters systematically whether you use a different methodology like OLS versus IV versus, for example, RDD. Now, the third and final reason to talk about this is testing for publication bias and p-hacking, which we talked about last week. I will go back to it very briefly.

Identification of publication bias 00:07:19

We were discussing a lot this simple meta-regression, where you regress the estimates you collect on variance, the standard error squared. Your primary goal is to recover the true effect beyond publication bias, corrected for publication bias: the beta zero. The beta coefficient should capture publication bias. Now, if you do not put any additional variables there, that means you essentially ignore heterogeneity, and you might have a problem here with identification. When you run a regression, you have to assume that the variable on the right-hand side is not related to the residual, the error term. If this assumption is broken, then your beta is going to be biased, and your underlying meta-analysis estimate of beta zero will also be biased.

Last time I mentioned one solution: we can use an instrument for variance. We can use a function of sample size in the first stage and then estimate the PEESE using the two-stage least squares estimator. But what you can also do, and what is in a way a different answer to this problem, is to add controls to the regression instead of IV. For example, you can add controls for whether OLS or IV is used in the primary study.

Then you leave the univariate meta-regression and you have what we typically call a multiple or multivariate meta-regression. So you have many variables there. If you choose them well, hopefully it might help you with p-hacking sometimes, if these choices between different methods are related to p-hacking. It may also help you even if there is no p-hacking, but simply some techniques influence both estimates and their precision at the same time. Again, a simple example is whether people use OLS or IV.

In my example from last time, you try to recover the effect of education on your income, your earnings. Then you should also control for ability, but you do not observe it. So if you omit ability and just regress earnings on education using OLS, you will get some estimate, and it will probably have some precision. The correct way, probably, is to use instrumental variables.

But if you use IV, you need to estimate two equations. First you need to find an instrument, and you regress education on the instrument, like distance to school, maybe in some developing countries. Then you estimate the actual model, where you use instrumented education, using the fitted values for education from the first regression.

When you have two stages, you will have much more imprecision there. So your IV estimate will almost always, I think always, be less precise than the OLS estimate. So if you just do the simple PEESE, the univariate PEESE, you have an omitted variable problem, where you should include controls for which technique the primary study used.

Do they use OLS or IV or something else? The IV versus OLS choice influences precision, and it can easily also influence how large the effect is. It can go both ways, but in any case you break this assumption. So certainly if you use OLS instead of IV, in the case of the education example, you have this endogeneity problem. So your IV, if specified well, will give you a systematically different answer, most likely.

And it will also give you a systematically different precision. With IV, less precision. The solution: you can use multivariate or multiple meta-regression with additional controls as another way to solve this problem, which I talked about last week. When you have omitted variables in your univariate PEESE, your estimate of the degree of publication bias and also of the corrected effect is therefore likely to be biased.

In almost all applications of meta-analysis in economics or finance, you will want to go beyond these publication bias tests. Typically in the paper, the way I do it, first I look at publication bias using different techniques, and then you also add the moderators. Some people would prefer to do it all at the same time. I like to try different publication bias techniques in the first stage, because some of them, especially some types of p-uniform, do not allow you to add these additional controls I was talking about. But you can still argue it is useful to have them as a robustness check.

00:14:24 Participant: [A participant asks whether, if a technique cannot control for covariates, one should simply not use it.]

00:14:41 Tomas Havranek: No. You might have techniques, like the one we have, MAIVE, where you do not have to include controls, because by using an instrument for variance here, you solve the omitted variable bias problem. Sometimes a model cannot include covariates but has another advantage: people who like selection models, for example, tell you that maybe in some cases it is better even though we cannot add these things.

There is no definite answer. It is just a convention for me, which I find useful: first, in my paper, to focus on publication bias, with a disclaimer that many of these techniques, because in the first stage we do not include these controls, can be off for this reason. But then we always move to this heterogeneity analysis. You can think about adding covariates as another way to make your publication bias analysis hopefully more robust, less biased.

Classical heterogeneity measure 00:16:23

How do we measure heterogeneity in meta-analysis? In economics we typically do not. The reason is that we always know, or at least are safe to assume, that it is huge. When you start doing a meta-analysis on an economics topic, you know there are going to be plenty of differences in the underlying results, because these are not experiments on a single thing. You have different countries, different samples of the population, different techniques. The differences are much more substantial than in medical research, for example, or I think in almost any other scientific field. There are also different identification approaches, which, by the way, I think is what defines economics and to some extent also finance: that we really focus on identification.

I was recently reading a lot about nutrition epidemiology: medical papers on what happens to you if you drink more alcohol or eat more meat. They are regressions on survey data: they ask you how you feel, and they ask you how much you drink and how much meat you eat. Then they run a regression, and then you see the headlines in the newspapers: meat is bad for you, or drinking can be good for you sometimes.

So they ignore identification. So what do people use? It is called I-squared, a measure of heterogeneity. It is a proportion. You take the total variation in the literature and try to compute how much of it is driven by real underlying heterogeneity in the effects: what is the percentage of the total variation which can be explained just by underlying heterogeneity across the true effects?

How is it computed? We go back to tau squared from the lecture on random effects versus fixed effects, if you remember. Tau squared, which you can compute, for instance, using some maximum likelihood techniques, is a measure of heterogeneity from the random effects model. So we know how to compute it, essentially. So we have it. And the other animal here is the average variance in reported studies. You can see immediately that if tau squared is really low, if it approaches zero, then your I-squared will also approach zero.

00:20:28 Participant: [A participant asks whether this measure assumes there are no covariates in the regression, and what happens if observable covariates can explain all the variation in the estimates.]

00:21:06 Tomas Havranek: The way I understand it, it is for simple averages in meta-analysis. It is not meant for the case when you add these moderators in that meta-regression. Then you can compute it, but I have not really seen it that often. So again, that is another reason why I think we do not really do it in economics, because it is meant for a univariate framework. Or you can look at R-squared in a regression. This would be a similar issue, maybe.

00:21:51 Participant: If you do a meta-regression in random effects, it is going to estimate that parameter, tau squared. Is that coming from that first stage, where it is by itself, or is it already accounting for the moderators?

00:22:04 Tomas Havranek: I think it must already be accounting for all of these moderators.

00:22:09 Participant: Then you want to see tau squared from a univariate regression.

00:22:16 Tomas Havranek: Right. Personally, I have an issue with including controls when you have a random effects model, because then you need to assume that these random effects are not related. And so I do not know. In my view, I would understand this multivariate regression as a way to explain these random effects. So when you put random effects, there is random heterogeneity, and when I do this, I want to explain it. I want to see what the pattern is. I want to model it explicitly.

Even though I do just this in the meta-regression, it is a fixed effect framework, and I add controls here. Meta-analysis people would still say you are in the fixed effect framework. But in fact, you do much more about heterogeneity than when you do random effects, because you do not have just one parameter of the random effects distribution there. Then you have a good number of explicit controls for these systematic differences; we can talk about how many. So you do more.

Sometimes I struggle to explain this to meta-analysis people who are not used to it. When you do [a meta-regression], yes, it is a fixed effect framework, but you allow for heterogeneity, and in a way which can actually explain it, give a pattern, a story to your estimation.

Back to I-squared. Again, you compute it as the fraction of the total variation which you can attribute to heterogeneity.

The total variation is tau squared plus the average variance in the published studies. The general recommendations are that if it is more than 75%, you have a lot of heterogeneity. What does that mean? Definitely, you should then use random effects, but many people would argue that you should use them even if you have a lower I-squared. In economics, it is almost always 90% or more. It is really rare to have a meta-analysis in economics or finance that would have an I-squared of 75%, for example.

Now, I personally don't like this I-squared too much because, first of all, I think it is not completely clear what tau squared actually means, and sometimes what it actually captures. And then there is this weird characteristic: if you have a lot of precise studies in the literature, you will have a smaller denominator and your I-squared will be larger. So when you have more and more precise studies, you will have higher implied heterogeneity, which, I think, is not always what we would like to have, or intuitive.

00:26:09 Participant: [A participant asks about higher precision and heterogeneity.]

00:26:14 Tomas Havranek: Yes, but not necessarily. You can have a literature with a lot of heterogeneity, and no matter how precisely you measure it, it will still be there. That is what distinguishes true heterogeneity from plain errors, variation which is random.

That is the way I understand it. Next time, which is tomorrow, I will briefly talk about some alternatives for how we can measure heterogeneity, or how we can say whether we have too much heterogeneity or not. That is I-squared. It is really important to know what it is, because you will see it a lot if you read a meta-analysis from any field.

Guiding example: beauty and success 00:27:18

I will use as an example the beauty effect paper, which I presented a couple of weeks ago, almost two weeks ago. Most of you have already seen something about it, so it will be familiar. I will go through, step by step, how we think about heterogeneity when we develop the analysis.

Heterogeneity both within and across 00:27:40

First of all, it is nice to have some motivation. For example, our motivation for the paper was that we saw a lot of heterogeneity, which increases over time. That is a good reason why you should focus on heterogeneity in a paper, because you see more and more of it as new studies are published. The figure shows the age of the study and the size of the beauty effect estimate.

In this paper we don't focus on it, but in many cases it will be interesting for you to see cross-country differences. Maybe, for the reason I mentioned in the beginning, you want to calibrate your model for a small country. You don't have estimates for your specific country, but you have estimates for many others, so you can compute the implied estimate that would be close, in terms of country characteristics, to something like New Zealand. It is often useful to take a look at how it is across countries. Here we don't see a clear pattern, so we don't focus on it in the paper that much, but quite often it might be useful.

Wide dispersion 00:29:18

Another thing that is obvious to show in an applied meta-analysis is just the distribution, a histogram. In this case it shows you that you have quite common estimates around zero, but also a relatively good number of estimates around 20. These are percentage increases in your salary when you become more beautiful. So, in the economic sense, there are a lot of differences, and you want to be able to explain them.

To me, probably the most intellectually difficult part of any meta-analysis is just before you start your data collection. You somehow choose a topic. Sometimes it is not obvious to everyone, but once you have this idea, you know it is a very good topic. But then you need to really capture what the main dimensions are in which these papers differ. You should not overdo it; maybe in this paper we do overdo it. We collect a lot of variables.

00:30:44 Participant: [A participant asks whether the minimum number of studies, six, is by accident or reflects a cutoff for including a characteristic.]

00:30:53 Tomas Havranek: Yes, that is by accident.

00:30:55 Participant: What do you recommend as a cutoff?

00:30:59 Tomas Havranek: It also depends on how many estimates you have in total. I have it for all of these characteristics, so you will see on the next slide that there might be much lower numbers.

Again, as I said, for me the most difficult part is to sit down with my co-authors and really try to isolate what we should collect. What should be the main drivers of the difference? Coding is, of course, when you actually collect the data. That is the time-consuming part, but if you prepare it well in the previous stage, it should be relatively straightforward. The question always is how much detail you should go into. Sometimes you might have, for example, method choices or data choices that are specific to one or two studies.

If you control for all of these things, it becomes almost like a fixed effect, or fixed effects at the level of studies. So you probably don't want to do it. But sometimes it might happen that only the two most recent studies do it, and it seems so promising, this new technique, that you want to control for it, with a disclaimer. Then, of course, most likely you will not have any statistical significance anyway.

00:32:58 Participant: [A participant says that, honestly, every study differs from every other study in many dimensions, so that picking only one characteristic attributes all the difference between two studies to it. They add that they are always annoyed by meta-analyses that make sweeping statements about differences based on very few studies.]

00:33:25 Tomas Havranek: Yes, you need to find a balance. Again, I chose this paper because it has many moderators. My view has recently been that less is more, and that this paper has slightly too many controls. Maybe if we had just three or four measures of beauty, that would be enough.

Measurement of beauty 00:33:48

Interview means that someone sits down with you and then marks the subject with a specific number of points. Photo-rated means you just get the photos, you give them to a group of maybe ten people, and then take the average ranking. Software means the same: it takes the pictures, but you let some artificial intelligence assign, for example, a facial symmetry score. Dummy beauty means that the study has dummies for average-looking and beautiful; it is not continuous.

You can have different settings. You can either compare pretty people to average people, or you can compare ugly people to average people. The latter is called the beauty penalty. They call it that, but it should not be a beauty penalty but an ugliness penalty.

00:35:44 Participant: [A participant asks whether the larger mean effect for some beauty measures means that those measures are more accurate.]

00:35:58 Tomas Havranek: I am not sure which one is more accurate, but you are right to observe that for software-rated beauty, the beauty effect on earnings seems to be bigger. And actually, we wanted to have it for, for example, this reason, which is related to attenuation bias.

Let me explain. Attenuation bias means that when you do a regression and you measure your independent variable with some noise, like beauty, then, if you have ten people and they cannot really agree on how beautiful the person in the picture is, you have some average beauty ranking, but it can be noisy. And then you have this attenuation bias in the estimated beauty effect. But maybe if you let ChatGPT or whatever rank these pictures, it could be more accurate, and therefore your independent variable in the regression, beauty, will not have so much noise in it, and you can have less attenuation bias in your beauty effect estimate.

That could be consistent with it, but later, when we do all the analysis together with other variables, it doesn't seem like too much of a difference. But you are right. That is actually a good observation, that the mean is the largest. But the means are also not so super far away from it.

Measurement of success 00:38:09

In this case, we are again looking at how beauty affects your success. We talked before about different measures of beauty. We also need to code different measures of success. Here you see that we have just two studies on athletic success, which means professionals such as footballers or rugby players. We have mostly data on earnings and a little bit on study outcomes and research outcomes. There is one paper which looks at economics and examines this question.

00:39:04 Participant: Are these numbers the mean estimated effect of beauty for those different categories? Are these subsamples?

00:39:19 Tomas Havranek: Yes, I should have mentioned that before. These are just summary statistics.

So these are different subsamples. These are the variables we then choose to include in the multivariate meta-regression. You have the number of estimates, the number of studies, the mean beauty premium, and then confidence intervals around this mean beauty premium for the specific subsample.

These numbers are not really super important now; it is just an example of how you think when you select these moderators before you start the data collection. Typically, you need to know, when you collect your data on any effect, that quite commonly it is going to come from a regression. So you will have the two most important issues: first, how Y is defined, and second, how X is defined. You always need to include some information on that.

Data characteristics 00:41:18

That would be Y, and beauty was the X. And then you want to say something about how the data are constructed in these primary studies. One thing you can always include is the age of the data, that is, to which data period they correspond. Of course, here we also have controls for different subsets of the population to which they correspond in the primary study, whether it is just men, women, or both genders together, or sex workers, and so on. So again, we might have included maybe too many.

00:42:09 Participant: [A participant comments on the number of studies with panel data, where beauty is observed over time.]

00:42:15 Tomas Havranek: Yes, for example, if I remember correctly, you would really track these people over time, or you would have access to photos, pictures over time. You would have pictures from high school, for example, and then you could look at grades.

00:42:38 Participant: [A participant remarks that most studies in this literature seem to use cross-sectional data.] [The participant then asks how occupation is coded as dressy or non-dressy, and whether it is the job the person ended up in or what they were trained in.]

00:43:07 Tomas Havranek: Oh, so I think we don't have it for students, because we don't know where they will end up. So it is just for people in a dataset who already work.

00:43:19 Participant: [A participant asks how the dressy or non-dressy occupation variable applies when the measure of success is study outcomes rather than an occupation.]

00:43:37 Tomas Havranek: I see. A small number of studies do not look at the job market, but look at how you do while at school, at college. There the outcome is what your grades are, for example. But I don't have to classify between dressy and non-dressy occupations, because you are a student. So this is a subset. And then we have a subset of estimates that are specific to lawyers, actors, for example, so we have fewer occupations, and we have a precise definition in the paper.

We could maybe just focus on what effect beauty has on your earnings when you work. But we also want to look at these other things, like what effect beauty has on your study outcomes, because there are some nice differences in studies on the COVID period. Then you have how politicians do in elections, and so on.

00:45:08 Participant: [A participant asks whether some studies use study outcomes as a control variable while others use them as the outcome.]

00:45:26 Tomas Havranek: That is a good question. To be honest, I don't remember. It might happen. Let me see.

00:45:35 Participant: [Another participant asks for the question to be clarified.]

00:45:40 Tomas Havranek: I think your question is, for example, that people might regress earnings on beauty, and they may also control for study outcomes in the primary studies. So I guess it is probably true that some studies do it: they would also control for education. Then you assume that beauty also doesn't affect education at the same time, so it gets more complicated. But I am pretty sure that some studies do.

Estimation technique 00:46:40

Right. You had something about the X variable, something about the Y variable, something about the subsamples, the different groups of people that the primary studies would focus on. And then you always have to code, in a meta-analysis, information on how it was actually identified. What was the estimation procedure? As I was talking about the nutrition example in medicine, it's very important to get identification, to be able to say something about causality. That is especially important if you want to publish your meta-analysis in a top economic journal, because in economics we are so obsessed with identification.

They also want you to be really careful about how you control for these differences. Typically you would have maybe a little bit of experimental evidence, but not much, a few RCTs. You would have some quasi-experimental approaches like IV, RDD, difference-in-differences. In this case, unfortunately, we don't have any experiments. It's hard to do experiments like that. But we do have a few IV studies. Difference-in-differences, well, we have two difference-in-differences studies on COVID. I think it's a good example of a control for which we really have just 12 observations from two studies.

So normally I would say don't bother, let's not collect it, there is too little variation in this dummy variable. By the way, these are all dummy variables. It means: does the primary study, does the estimate use OLS or not? These are binary dummy variables, which we collect in meta-analysis. The difference-in-differences studies, well, we wanted to include them, because these are the new studies on COVID, and you can claim they are really well identified, at least in principle. So you want to take them into account as a separate group for this reason.

Of course, you could also group them together with IV, and you could say: was it experimental, or is it purely observational? By observational, I mean when you do OLS, you try to include control variables. But even there, the boundary is not clear. When you do an RCT and you do OLS on your RCT data, well, it's still an experiment. Even though you use an observational technique, it's not a quasi-experiment, it's an experiment. And sometimes, when you have really good data, which probably isn't the case in the beauty literature, you can claim that even though it's not an experiment per se, your X variable really has exogenous variation. It is really almost like an experiment.

00:50:18 Participant: [A participant objects that the two difference-in-differences studies may capture something else by accident, for example if they were all done in the US, and that this could drive the difference.]

00:50:31 Tomas Havranek: So certainly you have to be careful about how you interpret it. But in this case, we don't find any difference in the final Bayesian model averaging, which I will show. But if we found that, we would have to be really careful about the interpretation.

00:50:58 Participant: [A participant adds that the fewer studies there are, the more careful one has to be.]

00:51:03 Tomas Havranek: Yes, I think that's a good principle.

00:51:05 Participant: [A participant asks whether Havranek has ever used, or seen used, a directed acyclic graph (DAG) to indicate which estimates capture total effects and which only direct effects, because controls can mix the two.]

[The participant explains that a DAG is a way of drawing the causal relations between variables, like the pictures used in structural equation modelling, and that it is associated with Judea Pearl. If beauty affects wages only through promotion, for example, holding promotion constant would hide much of the effect. A meta-regression does this only in a very crude way.]

00:52:56 Tomas Havranek: So that is a graphical summary of causality.

00:53:01 Participant: Yes, exactly. It's used to help identify.

00:53:03 Tomas Havranek: I will check it out. Thank you. Good.

Publication characteristics 00:53:28

I will try to speed it up a little bit. Another thing you can always control for is whether the paper was published or not, in a journal, or whether it's just a working paper. Here we have a control for high quality peer review, so that's something which might potentially capture quality, unobserved quality characteristics: was it published in a top five journal in economics with high quality peer review? You can also include the impact factor of the journal, or the number of citations of the study. Again, if you find that more citations are connected with, for example, larger effects, the question is how we should interpret it.

I can say, maybe, if a study is more cited, it might mean it's better conducted in ways which are difficult for me to capture. For example, is the instrument good or not? You can measure how strong it is, but not always how appropriate it is. So the number of citations could be a measure of quality, which is otherwise hard to codify. But also, the studies might be more cited because they have these bigger estimates which people want to cite. They need to cite. So when you find a correlation there, it's just an OLS regression, so you don't know what is causal there.

But it's a good idea to at least try to control for some of these publication characteristics. If you have no other reason, well, I think we can make a good case that some of these proxies, like where it was published, could in principle be related to quality, which is otherwise difficult to codify. You want to control for whether the publication rank matters. If you find that it doesn't matter, it's also an interesting story.

Beauty measurement doesn't matter 00:56:54

These are the same things, but in graphical form. Beauty measurement, well, doesn't seem to matter much, just by looking at the histograms. Success measurement, well, maybe a little bit, but it's hard to see any systematic pattern. Maybe it shifts a little bit, but it's really all close to the mean.

Ability control matters 00:57:26

Again, it's the same with the methodology: different methods give us a lot of similar answers. What we stress in the paper is whether the study controls for ability. Even before you do any sophisticated analysis, you can see that if you control for cognitive skills, that's the gray histogram, you are much less likely to find any substantial beauty premium. If you ignore ability, you are more likely to find large positive effects. So this might be an indication for you: I should focus on this, especially if you get the same or similar results just by looking at the histogram and then from a more detailed analysis.

We do have a control for non-cognitive skills. The problem is it's much more difficult to measure, and also much more difficult to distinguish from physical beauty.

I'm not a psychologist, but I would say some people have charisma, which is partly based on how they look, in a good way or a bad one. But we do have that control. Some studies try to do it. We don't find any difference, or not a significant one. But it's important to realize it's really hard to do, so there's going to be a huge measurement error in this charisma variable. For IQ, it's easier. You still have some noise, but it's much clearer how you can measure it.

Occupation matters 00:59:46

Obviously also occupation matters. For sex workers, beauty is much more important than for other occupations. And the difference is really huge. When you do a funnel plot for sex workers, you get a completely different shape, a completely different top of the funnel. So it might make sense. When you find a characteristic which gives you a completely different funnel plot, it's a good idea to completely separate it from your analysis, at least in some specifications. For some of the analyses, we take sex workers out of the sample.

01:00:38 Participant: How many studies did you have?

01:00:41 Tomas Havranek: Well, we can go back. I thought there weren't many. Where was it?

01:00:47 Participant: Four studies. So you cannot really do...

01:00:55 Tomas Havranek: Yes, so that's why we mostly take it out. We also do it separately, but you're right. Maybe here that would be the case where, if you wanted to do a proper meta-analysis of the beauty effect...

01:01:15 Participant: [A participant suggests that with so few studies one would not do a meta-analysis at all, or would just report averages.]

01:01:19 Tomas Havranek: Maybe. You're right. So that's a disclaimer we have in the paper, of course. But it's so obvious. So here we do it just for sex workers, but as you say, it's a very small sample.

01:02:25 Participant: [A participant remarks that tables should report both the number of estimates and the number of studies.]

01:02:31 Tomas Havranek: Yes, I think you have a good point. We should always also have the number of studies. This table, I don't think we even have it in the paper. Anyway, this is more interesting. You remove what is completely different, and you look at the sample without sex workers. And there, well, I will not talk about all these numbers, of course, but you still find some evidence for a beauty premium in many of these models. [Note, 2026: in the published meta-analysis, the premium stays large after controlling for cognitive ability only for sex workers; see meta-analysis.cz/beauty/.]

Option 1: subsamples 01:03:17

But still it's mixed. Maybe let me give you some more context on this. When you have substantial heterogeneity and, for example, it's because one part of your sample is really different, you have two options. Option number one, you split your sample. And this is a perfect idea. Again, in this case, we shouldn't really pay much attention to this small sample of sex workers. Or you can keep all the data together, and then you would do the multivariate meta-regression with all the controls on the right-hand side. So my inclination is to put it all together, if it makes sense conceptually, if it's the same topic, which it is here.

01:04:26 Participant: [A participant remarks that, a priori, one would not expect a different publication bias here.]

01:04:34 Tomas Havranek: Now, in my experience, top journals often force me to do a lot of subsamples, because they do not fully trust it. When you have all of these W variables in your model, you assume that there is a linear relation between these dummies and what the studies actually find. But of course, it doesn't always have to be linear. So we do a lot of univariate regressions. Then you have another problem, of course, because you have omitted variable problems there. So we have both in the paper.

When you put many controls in your model, it doesn't have to be a meta-regression, but any regression model, obviously at some point you will run into issues with collinearity. And second, which is related to it, you can overfit the model. I don't have time to explain all of it, but then crazy things can happen, because of all these different correlations. So there is a trade-off. I will show you a little bit how I try to treat it in Bayesian model averaging. But still, it's probably better not to overdo the number of regressors, again in any regression, not just in meta-regression. But you also want to control for the main drivers.

01:06:36 Participant: [A participant remarks that there are many possibilities for the variables.]

Option 2: Bayesian model averaging 01:06:46

01:06:46 Tomas Havranek: I showed different subsamples, but I will not go through them. What you will see often in meta-analysis, especially in my meta-analyses, are these Bayesian model averaging figures and tables. So what does it mean? Well, it's an answer to a problem. And the problem is, when you collect a good number of different variables which would explain why people report different estimates, why you have so much variation in your sample, you would like to run a regression.

You regress the estimates on these characteristics of data, methods, variables, identification, and so on. But the problem is that you should have good reasons why you include these parameters, but still, you are not sure a priori if all of them really belong to the true model. So when you run just an OLS regression, you are quite sure the variance will be overestimated. You will have too little precision in your meta-regression.

And by the way, in any regression in which you include many controls, the problem is called model uncertainty in statistics. You are not certain about the underlying structure of your model, be it a meta-regression or any other regression.

One solution: you do not know which combinations of your control variables you should include in the model, so let's try all of these combinations, which is going to be two to the power of the number of dummy variables or characteristics you have collected. Then you weigh these specifications according to something like how well they fit the data and how small they are. You would like to prefer smaller models; more parsimony is better. So you can imagine that you do many, many different regressions and weigh them by something like adjusted R-squared.

But if you want to do adjusted R-squared in the frequentist framework, it would take a lot of time. In a Bayesian setting, I will not go into the details of Bayesian Markov chain Monte Carlo and so on, but there is a way to make it much faster, much more efficient, when you use Bayesian techniques. That is why we have Bayesian model averaging. But Bayesian is really just for the sake of computational feasibility. It is not because you have to adopt a Bayesian way of looking at things. I like Bayesian approaches, even though in economics we are still mostly on the frequentist side. But you do not have to be a convinced Bayesian to use Bayesian model averaging.

01:10:13 Participant: Is it the same thing with lasso regression or something like that?

01:10:15 Tomas Havranek: Yes, the idea is the same. But lasso is frequentist; I don't think you can go through all of these combinations. You need the Markov chain. Well, sometimes, of course, if you have seven moderators, it does not matter.

01:10:37 Participant: [A participant says that something people sometimes do does not seem right: they run BMA and choose the variables with a high posterior inclusion probability.]

01:11:02 Tomas Havranek: I disagree. I also do it as a robustness check [selecting variables with BMA and then running OLS]. And the reason I disagree is, well, you can also use Bayesian model averaging as a model selection device. It is completely legitimate. Of course, you then move away completely from the Bayesian setting, and you combine essentially a Bayesian approach to model selection with a frequentist approach to estimation.

01:11:31 Participant: The problem is the standard errors.

01:11:35 Tomas Havranek: I also prefer the model averaging solution. Some people like to see something simple which they understand. If you just run a normal regression, it is not completely off. Some people would argue for the model selection approach. I agree with you that model averaging is better. I just think that as a robustness check it is not so bad to do a simple OLS based on the model averaging result, just to show people what happens if we just do it with OLS, which is not the best way to do it.

01:12:22 Participant: [A participant asks why the lecturer does not prefer lasso, which picks out a single specification. The participant notes that standard errors can be valid for a small subset of variables chosen this way and take model uncertainty into account, whereas a one-time regression on selected variables ignores the preceding search. The participant thinks there are better ways of model selection than what people do with model averaging, though no single best way exists, and that the standard errors remain a problem.]

01:13:35 Tomas Havranek: I think you are completely right on this. So I would prefer Bayesian model averaging. There is a nice paper in the Journal of Economic Literature by Steel on model averaging, which I think is really good. It explains what model uncertainty is. […] It is really a good paper, which is easy to read. You do not have to know much about Bayesian econometrics to appreciate what it does.

01:14:24 Participant: I don't think it was that easy to read. It may be easier to read for you, but I don't think it was that easy.

01:14:30 Tomas Havranek: Well, at least the introduction. Yes, so that is mostly the underlying idea. And then he has plenty of examples and some equations. So this is what I have actually just said on the slide, just a summary. There is not enough time to go through all of these technical aspects of BMA. You can read about it in this paper, which I love. […]

Just a couple of basic points. It is a Bayesian approach, so we have to choose priors. We choose priors in a way that should not really have a big effect on the results, just for computational convenience, feasibility and efficiency. Our prior is that all regression coefficients are zero. The weight of this prior is the same as one observation, one data point, so it is a very small weight. If you have a meta-analysis of typical size in economics, you could have a couple of hundred estimates.

It will slightly push your estimates to zero. That is what you should realize: it works against you if you want to find something, but in a really slight way. In practice you will not notice. This is called the UIP, unit information prior. A prior with the same weight as one data point is commonly used in Bayesian econometrics. Then you also need to use a prior for models.

You have the g-prior I was talking about previously, a prior on specific coefficients in your regression, your meta-regression. And then you need a prior on the different sizes of different models. What is the easiest way? I give all models the same weight. You can choose different priors. Maybe sometimes you want to give less weight to even smaller models, but I will leave that. What I think is important, and what has been a relatively recent development in Bayesian model averaging, is that we also add a dilution model prior.

And what does a dilution model prior mean? We compute the determinant of the correlation matrix for the variables included. Let me simplify it: if there is plenty of collinearity in the model, you will place less weight on this model. Again, what Bayesian model averaging does is run many regressions on different combinations of these variables, which you have collected. So, two to the power of the number of variables you have. So you have many, many different models, typically millions, sometimes billions. And they will have a different collinearity problem, depending on which variables are included in these models.

A clever way, I think, an idea by [Edward] George, is to use the determinant of the correlation matrix as the prior weight. If there is a lot of correlation, I think the determinant will be close to zero. And vice versa. So these more problematic models in terms of collinearity will be downweighted: you get not only a weighted average, but one that penalizes models which have these collinearity issues. Again, it is not a silver bullet. You could also do it in the frequentist framework, but here it is all built in; it is elegant.

Model inclusion in BMA 01:19:18

How does it look when we estimate it? What I would like to report are these figures. I think it is a great way when you can show your results in a graphical form. For many people it is easier to grasp, maybe not for all people. This picture shows you only a kind of statistical significance, but it shows you which variables are most important in statistical terms, in terms of explaining differences in the reported beauty premium.

You have these millions of different models, in columns, the variables in rows. If the color is blank, if it is white, it means the variable is not included in this model. The color means which sign: it is the reported or estimated sign in your meta-regression for this variable.

Blue is a positive sign, red is a negative sign. You can see that one, two, three, maybe four variables belong in all the best models. The horizontal axis measures how good these models are in terms of goodness of fit and parsimony combined. You can think about adjusted R-squared, or information criteria. As for the axes, the best models are on the left, the best variables are on the top. You can see publication bias, cognitive skills, sex workers being a separate thing, and also the impact factor of the journal.

Regression results 01:21:46

That is just statistical significance. For the actual size of these estimates, you need to go to the table where you have a posterior mean, which is the weighted mean of all of these regressions that were run by BMA. Then you have something like statistical significance, the posterior inclusion probability. Anyway, if you open any of my papers, you will have it explained in non-technical language as well. This is how we report it.

01:23:03 Participant: Why didn't you take your first regression from the BMA, the best regression from the BMA?

The first regression from the BMA is simple too.

01:23:18 Tomas Havranek: Yes, but then it is not independent of the BMA. We wanted something which is really simple and completely different. But in our case, it does not differ so much. So it shows the reader that it is not the Bayesian framework that drives the results. But of course, we trust the Bayesian BMA result more. I actually cut a portion of this table. It is a big table, so for you to be able to read it, I just cut it. But the two sets of results are not so far away from each other: 4.8 and 4.4.

We mentioned briefly, when we were talking about publication bias, the RoBMA, which is something really different, but also the same in a way. Robust Bayesian meta-analysis is a weighted mean of different publication bias correction techniques and fixed effect and random effect in simple models. The mechanism here is the same. You specify some priors, hopefully in a way which will not disturb what you do, just for computational purposes. Then you compute the weighted average, and the weight is something like model fit.

So the approach is similar. It is a different kind of thing, but the logic, or the mechanism, the Bayesian mechanism, is the same.

Posterior densities 01:26:36

Finally, you can also put your results in a graphical form, but this is from a different paper. We do not have it in the beauty paper; this is a paper on substitution between capital and labor. Anyway, don't read the labels, just look at the figures. You can take the variables from your BMA and plot the posterior distribution of the estimates. There you can see what the economic effect is and also what the significance is.

Here, for industry-level data, you can see it is about minus 0.2, and it is almost always negative. So it would be statistically significant. But the number is small. Or maybe it is big, because it is an elasticity, so it is not so small. If you do not want people to read through all of these numbers, you can put this table in the appendix and focus first on the inclusion probability, which is statistical significance, as an overview of the most important variables. Then for the most important variables you can do these posterior density plots, where you select maybe just the four most important variables based on the model inclusion from before. So these are the final two slides, the final points.

Best practice (implied estimate) 01:28:28

After all these steps, you have collected your data, you have run your regression, maybe Bayesian model averaging, maybe just a simple regression. So now you hope you know what the relationship is between results in the literature and data characteristics, estimation characteristics, publication characteristics, and so on.

But people will ask you, "So what? I need to calibrate my model at the Reserve Bank of New Zealand, and I need one number. What is the best number for me?" What makes sense to me is to use the BMA result and try to compute the implied, fitted value, which would be useful for your context. In many cases, you will compute several fitted values.

Best practice (class size effect) 01:29:33

Let me show what I mean. This is one such table from a paper on class size effect, I think I also mentioned it, which we published in the Journal of Labor Economics. This is the bottom line of the paper, this table, where we have the BMA table and we compute fitted values from the previous BMA estimation.

Then you have to select the value of each variable. For the standard error, you put zero. Then you have variables related to identification, so you would prefer ones which are quasi-experimental, for example, and so on. If you don't know, you just have no preference. So what do you do? You just use the sample mean for this variable.

Let's say some estimates are for Scandinavia. If you want to compute the overall mean, you don't put one or zero for Scandinavia. Or you can, and then you have a specific result for Scandinavia, but for the overall mean, you have no preference between Scandinavia and Wales or New Zealand.

So you put the simple mean for the Scandinavian countries, and so on. The overall mean would be based on no publication bias: you plug in zero for SE, you plug in one for quasi-experimental approaches, and so on. And for the rest you plug in sample means, because you have no preference.

01:31:36 Participant: Why do you call it best practice if you say you give some weight to OLS? OLS is unlikely to be the best practice.

01:31:44 Tomas Havranek: Yes, you're right. It should be called implied estimates or something. The thing is, OLS is not included here; it appears only in its own row. So we compute what the implied estimate would be if all studies used the STAR experiment. So you plug in one for the STAR experiment.

01:32:16 Participant: [A participant asks about the remaining variables.]

01:32:17 Tomas Havranek: Yes, except SE, the standard error. And then you do the same for regression discontinuity, for IV, for fixed effects, and for OLS. Then you forget about this. You define your best practice, which also means you make a choice on these values. So you could essentially put zero weight on OLS here, and you prefer something which is at least a little bit experimental.

Then you compute it for kindergarten, so zero for primary school. In the second row, you do the same for primary school, and so on, for different contexts and also for different sizes of the class.

01:33:12 Participant: [A participant makes a comment; the audio is too faint to follow.]

01:33:40 Tomas Havranek: We call it best practice, so it's the standard error as a proxy for publication bias at zero, OLS at zero, and a few other variables, I don't remember now which ones, also at zero.

01:33:59 Participant: [A participant adds a comment; the audio is too faint to follow.]

01:34:02 Tomas Havranek: Yes, because in some samples you have the omitted variable problem, so this will be the takeaway table. And here it shows it's never too large, because with a large STAR experiment it is still a very small fraction of a standard deviation effect when you increase or decrease your class size. [Note, 2026: in the published paper (Journal of Labor Economics), the implied class size effect is negligible for all identification approaches except Tennessee's STAR experiment and for all contexts except classes of fewer than 15 students. See meta-analysis.cz/class/.]

But something like this, it's not perfect, and if you look at the confidence intervals, they are quite wide. But you also see how, in most cases, we are able to reject these really huge effects of class size beyond 3, which are typically used to justify policy actions. The idea is to put it all together: to correct for publication bias, to correct for other biases in identification, and to produce estimates which would be suited for different contexts.

01:35:15 Participant: [A participant asks how publication bias is corrected in the best-practice estimate, and whether it is just PET.]

01:35:26 Tomas Havranek: I do just PET and PET-PEESE. That is of course a simplification, and that's why I think I need the previous part of the paper, where I discuss these different approaches to publication bias, and if my PET-PEESE is an outlier, then I shouldn't really rely much on these publication-based results here. So it gives credence and justification to the way I do it here. Like in the beauty effect paper, PET-PEESE was quite conservative, if you recall my presentation from the other day.

Most of the other techniques, the selection models for example, would push for more correction. So with the disclaimer, it's probably worse than what we have. But otherwise, it would be hard for me to justify why I should select this approach, because to do it in a selection model framework would probably be possible, but very difficult.

You would need a Bayesian version of Andrews and Kasy. So in principle it would be possible to do a BMA version of Andrews and Kasy, but I don't know how to do it. That's the only way I can show robustness in terms of publication bias results.

01:37:11 Participant: [A participant remarks that few people actually report a best-practice estimate.]

01:37:22 Tomas Havranek: Yes, one caveat: quite often you will really have wide confidence intervals. Here it's not so bad, but since you take all these regression coefficients, your final estimate reflects all the uncertainty in the BMA. That may be another reason not to overdo it with the number of variables we collect for meta-analysis. Anyway, I think I have reached my limits for today. […]

01:38:14 Participant: [A participant asks about the proxy for cognitive ability.]

01:38:35 Tomas Havranek: I don't think they use it as a proxy for cognitive ability. They would use it maybe as an initial control.

01:38:48 Participant: [A participant objects to the control; the audio is too faint to follow.]

01:38:57 Tomas Havranek: Let me think about it.

01:39:00 Participant: Because beauty impacts study abilities.

01:39:07 Tomas Havranek: Or not? Well, certainly, if you want to estimate how something affects your earnings, you also need to take education into account somehow.

01:39:18 Participant: [The participant adds that if beauty affects education outcomes, education in turn affects the outcome of interest.]

01:39:28 Tomas Havranek: Yes, exactly. I would need to look back at how precisely they do it, so I don't have an answer for it right now.

01:39:42 Participant: [The participant elaborates that in some studies, education itself is one of the outcomes affected by beauty.]

01:39:56 Tomas Havranek: There is one different study, which takes the COVID period and looks at how grades of students differ in that period. If I remember correctly, there was a substantial effect on the grades. So there could be an effect somehow. Maybe not on education, but on the grades.

See you tomorrow, guys. Thank you. […]

Corrections

Slips of the tongue corrected in the text:

  • [00:11:52] said "IV instead of OLS"; the text has "OLS instead of IV".
  • [01:28:56] said "Federal Reserve Bank of New Zealand"; the text has "Reserve Bank of New Zealand".