Transcript. Lecture 4, Data Collection, from Research Synthesis in Economics and Finance, given by Tomas Havranek at the University of Canterbury, Christchurch, in February and March 2025. An edited machine transcript. Tomas Havranek's words are edited like an authorized interview: in standard written English, without fillers, repetitions and unfinished sentences; numbers and negations are kept as spoken. Course administration and a few passages are left out; […] marks a cut, and editorial notes are in square brackets. Questions and comments from the audience are labelled Participant or summarised in brackets; participants are not named, except the host, Bob Reed, where Tomas Havranek refers to him. A few slips of the tongue are corrected in the text; they are listed at the end. Times are positions in the lecture recording; the slides follow the same order. What is said here is spoken and informal; the written guidelines take precedence.

Lecture 4. Data Collection 00:00:00

00:02:42 Tomas Havranek: Going back briefly to yesterday's lecture on literature search and on how to choose your topic: I gave you my approach to how I would do a meta-analysis, or a meta-research project, but other people have different ideas. I think it is nice to have a bit more balanced knowledge about how people approach it.

Today's topic is data collection. This is the part of meta-analysis which typically takes most of your time. It can be quite difficult and also laborious, not just in terms of time but in terms of how you have to think about it. First I will talk about what kind of effects we need. What do we need for them to be comparable in meta-analysis? And if not, what should we do? How can we make them comparable? And then, in a more practical way, how exactly should we do it? What is the most efficient way to collect the data? So let's go.

Effects must be quantitatively comparable 00:04:42

For meta-analysis to make any sense, we need the estimates from different studies to be comparable. They need to be comparable not just in a qualitative manner but directly numerically, so that we can say how much bigger this estimate from study A is than the estimate from study B. That is really important, because if you do not have it, you can be in big trouble, and then in the middle you find out that it does not really add up.

It seems trivial, but in many cases it is not, because, especially in economics, studies on the same topic can use really different measures: different units, different functional forms. I will talk about that in a while. Some people, especially in the past, would use t-statistics. The t-statistics are comparable, but the problem is that they do not measure how large the effect is; they just measure how strong it is in terms of statistical significance. You will need these statistics for something else, but do not use them as your effect size. That is not what you should focus on when you do a meta-analysis. Again, I will explain why.

In economics and finance, and I think generally in the social sciences, we will often have regression coefficients. We will have regressions in the primary studies. In the studies you are interested in, you will want to work with regression coefficients. But again, as I said, you cannot always take directly what people report, because some people would use New Zealand dollars, other people would use US dollars. Maybe sometimes you can translate it, but not always. Some people would use regressions in logarithms, other people would do it in levels, other people would do it in differences, and so on. So we will talk about it. It needs to be comparable.

Elasticities 00:07:02

Elasticity, which you know well from your previous studies, is very useful in meta-analysis because elasticity is directly comparable. If your topic is, for example, the price elasticity of gasoline, it is clear that you will focus on these elasticities. Elasticity is a percent change in one variable as a response to a one percent change in another variable.

How does it look in practice? If you have a regression of logarithm on logarithm, then the regression coefficient can be interpreted as an elasticity, which I am sure you learned a couple of years back, just as a reminder. That is what we want. It is perfect. If the studies in the literature regress logarithms on logarithms, then the coefficient is an elasticity, and you can directly collect it and directly compare it. It is great for meta-analysis.

In some fields we have this, and it is easy, but in others it is more complicated. Elasticity, I just want to stress once more, is probably the best effect size in economics, if you can use it, because it does not have all of these troubles that are connected to the use of standardized effect sizes, which we will discuss later. So elasticity is great. And sometimes, even if not all studies use elasticities, you can somehow recompute the results to elasticities. But it is a bit more difficult. Or you can at least take the subsample of the literature which focuses on elasticities.

Semi-elasticities, dollar values, std. dev. changes 00:09:04

In many cases we simply do not have elasticities, but sometimes you have something which is called a semi-elasticity. It is a bit similar, since you still have a logarithm on the left-hand side. I have some examples there: for example, the effect of euro adoption, the European common currency, on trade.

What happens to your trade with a different country, or with other countries in Europe, if you adopt the euro? My country, the Czech Republic, does not use the euro. We use our own Czech crown. But there are many studies which look at what would happen if you adopted the euro, and the studies use econometric specifications from previous adoptions of the euro. So you would have trade on the left-hand side, the logarithm of trade.

And on the right-hand side of the regression you would have whether the country has it or not, zero or one. The coefficient on the dummy variable on the right-hand side can be called a semi-elasticity, which means the response is in percent. So what happens to trade in percent? It is in logarithms. But the treatment, the initial thing which moves first, is zero or one. So it is kind of standardized as well, and you can also use it easily as your effect size in meta-analysis.

That is the euro adoption, but you can also have things like the effects of foreign ownership on domestic productivity. You have foreign investment, foreign ownership of domestic firms, going from 0% to 100%, from 0 to 1, and then what are the effects on the productivity of these other firms in the country?

Or borders and trade. There is a big literature on what happens to trade if you suddenly have borders. The most famous case is between the US and Canada. These are countries which otherwise, until recently, did not have so many tariff barriers, problems for export and import. But you had the border, and the studies show there is a big difference in trade just because of the existence of the border.

For some questions, you will be able to use just dollar values. For example, what is the value of a statistical life? You know the concept. For instance, if the government wants to build a new freeway, in many countries you can build it in a really expensive way, which is very safe, or you can build it in a way which is much cheaper, but not so safe.

Of course, you would like to have your freeways amazing, beautiful and safe, but you also have a limited budget. So you need to compute somehow how much you value human life, even though that is a strange thing. It is really used in policy evaluation, because you need it.

There are ways to estimate it. It is in dollars, so that is easy. Or the social cost of carbon. Do you know the concept of the social cost of carbon? Essentially, you compute the damage done by one additional ton of CO2 emitted into the atmosphere. You need it when you want to determine what the optimal carbon tax is. For example, if you wanted to establish a global tax on carbon emissions, that is what you should use as your scientific basis: the social cost of carbon. Again, dollar values.

These are simple cases; then it can become a bit more complicated. For example, last time I think I mentioned a paper we have on class size: what is the effect of class size in primary schools on how kids learn, on student performance? You can also do it at university level, but that is a bit less interesting. You have many studies on different class sizes and the performance of students, but these studies measure student performance in different ways. Almost always you look at grades, but the way grades are assigned is different in different countries, and even within one country you can look at different levels.

It would be great if you had similar studies, so that you would be able to say by how many points your grades improve if you are placed in a smaller class, but that is not the case, because you have very different ways in which people measure it. So the best you can do, in this case probably, is to look at standard deviation changes. What I mean is: what happens to your performance when you are moved to a smaller class, in terms of how you move in the distribution of students?

You can move, for example, by one tenth of one standard deviation. It is a bit more difficult, but it still makes sense, and many policy interventions, in education especially, are measured in terms of standard deviation changes. So you can also use it as an effect size in meta-analysis.

Standardized mean differences 00:15:55

Even if you try all of these things, looking at elasticities, semi-elasticities and dollar values, it is still not applicable in your field, in your meta-analysis, because people have really used different things. It is not possible to combine it, even using some sort of standard deviation changes. What is actually pretty common outside economics, in sociology, in psychology, in education research, is that people often look at mean differences, standardized mean differences.

That is designed for the case when you have something like an experiment, or ideally an actual randomized experiment. You have a treatment group and a control group, and you just compute the difference between the two outcomes. Then you divide it by a standard deviation. It can vary a little bit, but usually it is the pooled standard deviation of the two groups. In other words, what you have is how large the treatment effect is when you compute it in standard deviation units. It is a bit similar to the previous case with class size, but more general.

You can only use it well when you have two groups, when you can make a difference. Going back to my example with classes, because it is a simple example: if I was able to do an experiment and force one school to create smaller classes and bigger classes, and randomly assign kids and teachers to these classes, then I could simply measure the difference between test scores for the small classes and the big classes, and divide by the standard deviation, and I would have something similar. And then I could compare it across different studies. As you can probably see quite clearly, the problem is that in economics and finance we do not often have these experimental settings. We would have something which is continuous.

Sometimes you can be close, but quite often it is something really different. That is the problem we face here, because most of the research in economics is not experimental. We do regression analysis, which is mostly observational. You try to control for things and so on. So it is not binary; the treatment variable is continuous.

Partial correlation coefficients 00:18:45

What is much more commonly used in economics is partial correlations. It is like a correlation. You try to recompute the different effects that people report on the same thing, but using different approaches, into a correlation. And it is a partial correlation because it comes typically from the regressions that people report in primary studies. They have not just one variable but more. So "partial" means like partial derivative: it is conditional on the other variables included.

But it is just a correlation. That is the most important thing. Now you can compute it using this formula. You collect t-statistics, so that is where t-statistics are useful, and you divide it by the square root of the t-statistic squared plus degrees of freedom. As we discussed yesterday, sometimes it is difficult to collect degrees of freedom, so you will just use the sample size if you have trouble with degrees of freedom. That is all right. Then you have a correlation. This approach, like all approaches which use standardized coefficients, has some statistical problems. We will come back to it later. But for now, this is important, and it is used quite a lot in economics because very often that is the only thing you can do.

When you do meta-analysis, you can use partial correlations. I have one example there. We have a paper on tuition, so how much you pay for your schooling, your fees at university, and enrollment, so how tuition affects enrollment. And again, we would like to have an elasticity, but many studies, for many reasons, do not report elasticity. They are not logarithms on logarithms, but they have different specifications. So as the main thing in this paper, in this meta-analysis, we use partial correlation coefficients.

But then, as a robustness check, we also have a subsample of the papers which actually use elasticities, to convince people that the results would not be so different if we just focused on elasticities all the time. You can still use partial correlations. It is OK. It is doable. It is established. There are these small statistical problems which you can solve, so it is fine. But still, when you do it and look at correlations, you lose a lot of information. [Note, 2026: the thesis notes on meta-analysis.cz advise partial correlations only as a last resort, when definitions cannot be translated, with a robustness check on a comparable subset; see meta-analysis.cz/ai/thesis-notes.md.]

It is better than just the t-statistic. It is not only a measure of significance; it tells you something about the strength of the association. But it is still a bit hard to interpret, so if you can use something else, it is better. If you can use something else, at least as a robustness check, as I mentioned in the education paper, that is also good, because then you can at least show people: "Okay, I do my best, but if you really want an economic effect size, not a correlation, this is how it looks."

If you use partial correlations, which might happen if you write a couple of meta-analyses, it is very useful to use the guidelines by Chris Doucouliagos from 2011. I have done it many times. The guidelines are unpublished, but nevertheless very useful. He collects many different partial correlations from economics, and he shows you which number you can actually consider to be small, medium, large, or essentially zero. It is not a silver bullet for your problem with the interpretation of correlations, but it is better than just saying, "I don't know."

We use these guidelines quite a lot. I think they are really useful. As I mentioned, if you use standardized effect sizes, which means you recompute what people report into something else, you will have some statistical issues, which I will talk about later in a more detailed statistical lecture. So, not today. A robustness check is useful, as we discussed. Now, what should we do? You have already found the papers and collected the PDFs. You have them in your folder.

Collect all estimates 00:24:15

Now you want to start collecting them. Some papers are okay in the sense that they just report a couple of estimates. That is easy. But others report many of them. They could have many tables with robustness checks. So what to do? The general advice I try to follow is to collect all of them, if possible. Of course, that might not be feasible for a small-scale student project. But if you can do it, if you have co-authors, if you want to publish the paper, it is better to collect all of the estimates.

But it is also useful, when you collect them, to make a note of what the authors of the primary study actually think about the estimate. Sometimes they really put it forward, saying, "This is what we prefer." Sometimes they say, "No, no, no, this is just to show what happens when you do it wrong." The traditional approach in meta-analysis, what we have done for many years, is just to ignore the thoughts or the preferences of the author of the original study. Maybe we just control for the methodology. But recently it has happened to me many times that referees or editors forced me, and I think they were right, to also collect information on what the original authors think.

Because sometimes it is an IV estimate or a difference-in-differences estimate, which normally you would prefer, but they say, "Be careful, there is a mistake: the assumptions are not met. This is a wrong estimate, and this is what happens when you do it wrong." So I think it is useful, as we do in this class size meta-analysis, to distinguish between when the estimate is explicitly preferred by the authors, when they say nothing about it and it is just neutral, like a robustness check and nothing else, or when they say explicitly, "It is rubbish, and we just do it to show you how not to do it." You could also omit these estimates, but I think it is much more interesting to collect them, and then you can play around with them.

There could be differences in publication bias, for example, between these different groups. Sometimes it is a bit hard to classify, but when you have the neutral category, you have either an explicit preference, or the explicit opposite, an explicit dismissal, or something in between, which will be neutral. At a top journal, the editor said that we should just focus on the preferred estimates. But then, of course, what happens is that I do not have enough observations for my heterogeneity analysis.

There could be more than one preferred estimate per study, because, for instance, they could show you a table and say, "We prefer this table." It has ten different specifications, slightly different. Or it could be five different countries, and you would have five. But often it is one. So you cannot really do the second part of a typical meta-analysis, which is heterogeneity, in a structured manner. The first part of most meta-analyses is publication bias and p-hacking analysis. We will come to that, I think, in two weeks. The second part is typically some sort of analysis of heterogeneity, with Bayesian model averaging and the pictures you saw last week. We will get to that later.

So at least the publication bias analysis or some simple summary statistics you can easily do for the subsample of preferred estimates. And I think it is a powerful way to show people that your meta-analysis really has strong fundamental principles, that it does not rest on a couple of poorly identified studies. If you can do it, I think that is pretty good. It also adds flavor to your results if you can show that the preferred estimates are typically bigger, even after I control for methodology and data, and the discounted estimates are typically smaller. In this paper [the class size meta-analysis], we do not find much of a difference, maybe a little bit, but not much.

Standard errors 00:30:10

For the first part of a meta-analysis, which is publication bias analysis, you will definitely need standard errors. You need a measure of precision for each estimate. For the second part, heterogeneity, in some cases you do not need it. Again, I advise you to do your best to collect it, but in some cases it is simply really difficult, like if you have some simulation studies. For example, we have a paper on the social cost of carbon.

In the social cost of carbon literature, you do not really have regression estimates. You have simulations mostly. So it is hard; people do not report standard errors. But you still want to say something about what the literature says. So the publication bias analysis is difficult, but you can do the heterogeneity analysis at least. That is also possible, but in most cases you will want to have your standard errors. [Note, 2026: the published paper, Havranek, Irsova, Janda and Zilberman (2015, Energy Economics), does test for publication bias in estimates of the social cost of carbon and finds substantial selective reporting; see meta-analysis.cz/scc/.]

00:31:21 Participant: [A participant asks about synthetic control studies.]

00:31:26 Tomas Havranek: Right. It is like difference-in-differences in a way. But I think the more recent applications would give you a confidence interval, from which you should be able to approximate the standard error.

00:31:46 Participant: [A participant makes a short remark that is hard to make out.]

00:31:52 Tomas Havranek: Okay, that is a good remark. Sometimes you do not have precisely a standard error; sometimes you have a confidence interval. The same issue arises with vector autoregressions. When you have vector autoregressions, you have these impulse responses, which I talked about during the first lecture. You also have a graph, and then you have confidence bands. You do not have an explicit standard error, but you can make some reasonable assumptions, maybe that the distribution is close to normal, and then you can compute it.

00:32:31 Participant: [A participant remarks that any measure of uncertainty will do.]

00:32:34 Tomas Havranek: That will be fine, of course. The best thing is if you have directly reported standard errors.

00:32:43 Participant: [A participant points out that, for a binary outcome, studies report different effect sizes, such as marginal effects at the mean or average marginal effects, and asks what to do then.]

00:33:14 Tomas Havranek: I do not think it has a 100% clear solution. There is no correct solution. I would probably go for the marginal effect. There must be one way which is the most common in the literature, and there could be a way to recompute the other ones to the one which is the baseline. It should be possible if you have the data, if you have sample means. Pick one, because it is not obvious. There is no right or wrong.

Typically you have a regression table with your point estimate and, in parentheses, your standard error. Sometimes you have a t-statistic, but that is essentially the same thing, and you can easily compute it. If you have the t-statistic and the effect size, you can compute the standard error easily. You can also compute it from the p-value; that is relatively similar.

What we did in one paper was a robustness check [Matousek, Havranek and Irsova 2022, a meta-analysis of individual discount rates]. Suppose you have a literature on an effect, and sometimes you have regression estimates and sometimes you have simulations or something like the synthetic control method without any confidence intervals. You still want to do a publication bias analysis for the entire sample as well. As a last resort, you can take the dispersion of results within one study and use it as the uncertainty measure for all estimates in the study.

Or you can just maybe take the mean value from the study. It would probably make more sense to take the mean or the median and then use the dispersion within the study as a measure of uncertainty. That is not really what you want, but it is a last resort. The paper was published in Experimental Economics.

00:35:51 Participant: If you were willing to make bold assumptions.

00:35:54 Tomas Havranek: Maybe. Sometimes you need to. So, of course, that is true. Now, commonly in meta-analysis, you will encounter problems when you have interactions in your regressions or transformations. So instead of a simple linear function, you have a quadratic function.

Delta method 00:36:15

What to do now? The question is how to compute the standard error. Computing the effect size is easy: that is the partial derivative, you just compute a derivative. But if you want to compute precision, variance, standard errors, you have to do something like the delta method, which is based on a Taylor series expansion. We do not have to talk about the details, but in practice there is a formula.

[A discussion of how to compute the standard error of a transformed estimate with the delta method follows.] [Note, 2026: in a primary regression, the delta method is applied to the estimated coefficients. For a regression of y on x and x squared, with coefficients β on x and γ on x squared, the effect of x at the sample mean x̄ is β + 2γx̄, and its variance is Var(β̂) + 4x̄²Var(γ̂) + 4x̄Cov(β̂, γ̂), with the variances and the covariance taken from the estimated covariance matrix of the coefficients.]

So this is a bit tricky. A common case is that you have an inverse on the right-hand side. It is just slightly more complicated, but the same logic applies here.

00:42:46 Participant: But is the best practice to include interactions or squares?

00:42:51 Tomas Havranek: That is what I do. The best practice would be to include it, and then you can always have a robustness check where you throw it out. Or you can have an indicator variable. It also depends on how large your sample is, because sometimes you have a small sample and you really would like to have these interactions.

But if you have 2,000 observations, it is probably okay to say, "I do not include interactions, because you need additional assumptions to do it."

Graphical estimates, measurement error 00:43:38

I think I also mentioned briefly in one of the lectures that you can do meta-analysis of graphical data. And now we go back to vector autoregressions. Suppose you have results like this. It is an impulse response, which shows you in time what happens to, in this case, house prices, but it can be any variable, when the central bank increases interest rates by one percentage point.

These are the results you will have in papers: you do not have numbers. But you can use some software, as I mentioned, and I forget the name, but it is in lecture one or two, I think. WebPlotDigitizer. It can help you measure the pixel coordinates. When I presented this in Prague and in Zurich, a frequent objection was, "But it will never be precise, so you will have measurement error in your meta-analysis."

Well, you always have measurement error. Even if you take regression results from regression tables, they are rounded. When people report their regression tables, they round their numbers to a couple of decimal points. How many? Two or three, typically, and it varies. Here, at least, when I do it myself, it might depend on the quality of the figure, but essentially I am equally imprecise. If the same person does it, you have just random measurement error. But if you collect estimates that are published as numbers in regression tables, you will have a different amount of measurement error in different tables, based on the rounding.

By the way, I have never seen a paper that explores this in meta-analysis. That could be interesting, because it could have consequences for some of the publication bias tests, for example. So the answer here is that it is not a big issue. You always have measurement error, not just in meta-analysis: you have measurement error in any regression analysis, in any empirical analysis you do in any kind of scientific discipline. If you regress GDP on inflation, of course you will have measurement error, because GDP is measured with huge uncertainty.

There is plenty of measurement error in inflation as well. There are many ways to measure inflation. The point is that we always have measurement error, and there is a consequence. If the measurement error is really random, which quite often I think it is, then you do a regression on two variables that have random noise. What will happen? I am asking you. What will happen to the regression coefficient when you add noise to Y and X, or when you add noise just to X?

You have a regression here, and you add noise to X. What will happen to your estimate of beta? What do you think? Will it be bigger or smaller or the same or just more imprecise? [A participant answers.] Exactly. You will have attenuation bias. Imagine you have X and you add to X a lot of noise, so much noise that X is just noise. If you regress something on noise, you are likely to get a beta of 0.

If you have random measurement error in X, you have attenuation bias toward zero. So if you think about it, we should expect, in most cases in economics, in any regression, a bias to zero, just by the nature of how the world works, because you always have some kind of noise in your data. I think that is quite important to realize. Next week, in our journal discussion group, I will present a project I have with a colleague, where we compare attenuation bias and publication bias in economics using a large data set. It is tricky, but I think it is pretty interesting, because you have these two opposing effects, opposing phenomena.

Outliers 00:49:07

Another issue is what to do with outliers. You collect your data, and your funnel plot can look like this [slide: Outliers]. It is not very beautiful. It is not one funnel, but it looks like a couple of funnels put together, so there are plenty of different things.

What to do with outliers? 00:49:28

Rule number one is that there is no rule on outliers. There is no consensus. Even in the guidelines paper we have on how to do meta-analysis, we are four co-authors, and each of us has a different view on how you should handle outliers. This should be the disclaimer I give you first: this is just my personal preference. What some people, for example Tom Stanley, would tell you is that you should keep all the outliers as they are. But if it is really huge, then you should omit it, but only in extreme cases, for example when it is more than three standard deviations away from the mean.

Some other people would tell you that you can either run some robust regression, which nowadays is not really done very often, so I do not recommend it, or you can use some outlier detection tools. But there are so many different outlier detection tools that it is hard to choose what to do. What I personally do is winsorizing. I will explain what it means.

Winsorization 00:50:53

Suppose my data set looks like the histogram on the top [slide: Winsorization]. You have some extreme observations here, and some people would tell you, "This is impossible, let's completely ignore it, let's cut it." The same applies to the other extreme. What winsorizing does, if you say I winsorize at the 1% level, is this: you go to the 99th percentile of your data, and you do not drop the observations, but you reduce them, all the observations that are bigger than your 99th percentile.

You reduce them to the 99th percentile here. It is important to be symmetrical, so you do the same thing from the left as well. So your median estimate is the same. Your average, your mean, can change a little bit, but your median is the same, because you treat the big and the small estimates in the same way. But if you just remove this estimate here, your median will change a little bit.

What I like about winsorizing is that you keep in your data set the information that there is a large observation. You just decrease the impact, the weight, the importance it has on your results. It is still there, but it is not so extreme. What I think we really miss here is that there is no paper that would systematically compare different ways to treat outliers and what happens in your meta-analysis. [Note, 2026: a later paper, Havranek, Irsova, Luskova and Stanley (2026), examines whether decisions about outliers and influential effects matter in 358 behavioral science meta-analyses; see meta-analysis.cz/outliers/.]

With winsorizing, you really change the data subjectively. So you should be able to show that this is not what is driving your results. You should show different winsorizing levels. And of course, when it makes no difference for the results, that is perfect. If nothing really happens, then do not winsorize. You do not want your results to be driven just by a few outliers.

I think many meta-analyses can be driven by a couple of really super precise estimates, especially when they use inverse variance weights, and quite often these super precise estimates are a bit fishy. Sometimes you see it when you look at how they compute precision, so be careful about it. That is why I would winsorize estimates, but also standard errors, in the same way, at 1%. It is a very light way to do it, and it is relatively conservative as well.

But of course, always do some sort of robustness checks, especially because we do not have a clear way to handle outliers, as I have stressed repeatedly. We have my preference, you have Tom Stanley's preference, you have Chris Doucouliagos's preference. I think it is a big issue, especially, again, with inverse variance weights, and we will get to it. Now I will show you some examples of data sets. Typically you do it in Excel or in some sort of spreadsheet.

Reflecting context: the effect of beauty on success 00:55:17

Before you really start collecting this data, because you do it by hand, you need to spend a couple of days, if you do a big project, thinking about what should be the main drivers of the differences across studies and between different estimates. Of course, you will collect estimates, standard errors, sample sizes, the number of citations of the papers, the publication outlets, and so on. Maybe the impact factor. It is easy. You can always do it in meta-analysis. But then you should be able to choose a couple of really key characteristics of the studies, which are maybe discussed in the literature, in some literature surveys done previously, or which prominent studies say should matter, or which prominent studies use for robustness checks.

For example, this is from a paper which I will present in two weeks at the departmental seminar here, on beauty. Here there are a couple of characteristics that are crucial, which people before us have repeatedly stressed: if you do it this way, you might have different estimates.

Measurement 00:56:40

Again, the question is: how does beauty affect your success or productivity? Success, we say. How do you measure beauty? Either you can hire some people who look at photos and put numbers to the photos, or you can have some sort of AI, which nowadays would probably be quite common, but not so much in our data set. You would put the photos in ChatGPT and ask for a rating on a scale from 0 to 10, and so on. You can use self-reporting: what do I think about myself, and so on. And you need to distinguish whether the paper focuses on the beauty premium for people who are above average or looks at people who are below average. That is the variable on the right-hand side.

00:57:55 Participant: Shouldn't you only code things that happen often? If you only have three studies with self-rated beauty, should we actually code it? Because, if you think about it, we have, say, 50 studies in our data set, but we can easily find 50 variables describing the different studies, and then we cannot include them all anyway. Or if we include them, say there are three self-rated, and two dummies, and four photo-rated, and they are all so small, your estimate will then only be determined by these four studies.

00:58:31 Tomas Havranek: Yes, because then you would have no degrees of freedom. So you have to make assumptions. There is a threshold. And these will be dummy variables that you code: is it software-rated or not? One or zero, and so on. In these dummy variables, you will need some variance. You need some variance to be able to do something with them in a regression analysis.

00:59:04 Participant: For each variable, you have at least 10 such studies, or something like that.

00:59:11 Tomas Havranek: That would be hard to have. Sometimes you have just three studies. But still, it is probably imprecise. But again, it is not set in stone. There is no clear threshold.

00:59:45 Participant: There are fewer than 10 studies that have the characteristics.

00:59:48 Tomas Havranek: I think in the guidelines we say that the dummy should have a mean above 0.05, that is, 5% of all estimates. [Note, 2026: the published guide warns against dummies with a mean below 0.03 or above 0.97; see meta-analysis.cz/guidelines/guide/.]

01:00:04 Participant: Many of these things are at the study level.

01:00:11 Tomas Havranek: Of course. So the rule could be: if it is just one study, it is like a fixed effect for the study, so do not collect it. That is part of the art involved in meta-analysis: to select well. This was a big project. We are three people on the paper. Well, it is not so many, but still we could afford to collect many. But I think in a sense you could live with three characteristics here: is it AI, is it self-rated, or something else?

01:00:44 Participant: Cast the net as wide as possible?

01:00:51 Tomas Havranek: Historically, I would cast it as wide as possible, but as I get older, I think that sometimes referees actually appreciate parsimony, when it is not so exhaustive. I think it is better to invest time into thinking about what is really crucial for me to include than in the actual data collection of dozens of different variables. So I think I am going the way that you are suggesting. It is okay to have just 10 variables in total in your meta-analysis. In my papers, you sometimes see 50 variables.

In this paper we have quite a few. Another thing is how to measure success. And here I think you need to distinguish between earnings, grades, number of papers, sports, and elections, because you have a good number of studies for each of these characteristics.

Data, estimation, publication 01:02:21

Then you can distinguish between the kinds of people you are looking at. Are you looking at men or women, or is it all together? Do you consider just people who earn a lot of money? That will probably have different results. Do you look at people who work in what we call dressy occupations, where appearance is important, like lawyers, actors, and so on? Do you do it in Europe or in, for example, China? That could matter, and so on. And then, how do you estimate it? Unfortunately, it is hard to do an experiment, because you would need volunteers. It is essentially impossible to do a good experiment on beauty.

01:03:29 Participant: [A participant makes a remark about dressing up versus coming in casual clothes.]

01:03:40 Tomas Havranek: Some of these papers try to do it. They try to enforce homogeneous pictures, for example. But to be honest, I do not think it can be done 100%. I think the point is who is rating it. Who would rate the picture? So you cannot do a good experiment.

So essentially you are left with regression, OLS. There are a couple of nice papers with difference-in-differences using the COVID pandemic for students. There is some evidence that grades can be related to appearance. As a consequence of the pandemic, you would have no in-person schooling, lectures or exams; it was all online. The idea was that maybe then appearance would be less important, and indeed the paper shows a strong impact on the grades of good-looking students during the pandemic. So that is a nice way to measure it.

What is important about our paper is that we want to deliver something new, not just a nice summary, a nice meta-analysis, but something else. So we distinguish between papers that have some data on cognitive ability, and also non-cognitive, but especially cognitive, which means intelligence or some measures of intelligence. Because there is a big literature in biology which implies there should be a correlation between beauty and intelligence.

We put in IQ, and so we would like to see if the studies which control for cognitive ability find a smaller beauty premium. That's what you would expect based on the biological evidence. But no study has done that comparison before in the context of economics, because you wouldn't have the data. But now we have the data, because we have the meta-analysis. I will not show you the results; I will show them to you in my seminar. In a meta-analysis, you can always put in dummies for whether the study was published, whether it is a working paper, whether it was published in a top journal, and so on.

Artificial intelligence and context 01:09:42

Nowadays, I think it's very useful to ask your good friend AI for help here. Not, of course, to do it for me, but to help me, for instance, choose which characteristics in this literature are important. It will give you some ideas. I think you will probably want to use some of them, and some of them not. So don't be afraid to use it. Again, as we discussed last time, I believe that in a couple of years it will be much more useful for meta-analysis at many stages. When you collect data, you will have a lot of help from AI.

Value added 01:10:33

I hinted at it a little bit with the IQ and beauty issue, but if you want to publish your meta-analysis in a high-ranking journal, what really helps is when you are able to show some sort of value added beyond a summary of the literature, beyond correction for publication bias. These things are important, but you can also bring something completely new: even though you do a meta-analysis, it's like primary study research. What I mean, for instance: the other day I showed you our paper on the effect of daylight saving time on energy consumption.

DST savings and latitude 01:11:22

So we collected these different papers on the effect of daylight saving time from different countries, in total, I think, about 30-something countries. [Note, 2026: the published paper covers 21 countries; see meta-analysis.cz/dst/.] But each study was done just for one country, or maybe two countries, or maybe the European Union, but not individually for different countries.

Now we have this data set, and we can actually look at how the effects differ in terms of the climate of each country. If the country is in Scandinavia, like Norway, it really has a big variation in sunlight. So daylight saving might have different effects there than in Jordan, which has relatively modest variation because it's much closer to the equator. And indeed, we find that in countries which are closer to the poles, where you have more variation, the effects are better in terms of higher savings.

01:12:25 Participant: But doesn't that require a huge amount of studies? If you suppose you only have two studies for each country, then in fact the uncertainty that you get for the estimate for each country is huge. And the regression you estimate there is only at the level of the country, so you don't have that many observations. The amount of evidence that you have there is always very, very small. I understand it's a good selling point, but...

01:12:59 Tomas Havranek: Yes, that's a good point. I would respond that in many cases you would have, for example, one study which does the same thing for 15 countries, and then you have 15 observations. One study with 15 observations cannot really do the kind of thing that you would do. But with 15 from this study and three countries in another paper, you get to 30 relatively quickly. Not always, but...

01:13:33 Participant: But that goes back to: is 30 enough?

01:13:37 Tomas Havranek: Yes, so it will not be super powerful, but it's a combination: you have this, and this is the icing on the cake which you deliver. And I think it's useful to do it. So daylight saving time would be one example, and the IQ would be another example.

01:14:16 Participant: [A participant suggests doing the analysis just for China.]

01:14:19 Tomas Havranek: That's a good idea. Yes, sure. You get rid of all this noise, and you have much more focused estimates.

01:14:47 Participant: How you sell things is important.

01:14:48 Tomas Havranek: It's the way you sell it, but it's not just in research; it's in any professional career. You need to not just work hard and be correct, but be able to sell what you do in a way which is not really intrusive, but persuasive. And I'm still learning how to do it, but I think it's important to emphasize.

01:15:21 Participant: As scientists, we hate it.

01:15:22 Tomas Havranek: We hate it, but it's true.

01:15:25 Participant: The older you get, the more you see it.

01:15:28 Tomas Havranek: Yes, exactly. So this is the daylight saving paper [slide: DST savings and latitude]. You can see that where you have more daylight hours in the country, the effect is more negative, which means the savings are more pronounced.

House prices and monetary policy 01:15:42

Or we have the paper on the effects of monetary policy on house prices. For instance, if the mortgage sector in the country is more important relative to GDP, we have faster, bigger effects of monetary policy on house prices, so that's also intuitive.

Questions and discussion 01:16:16

That's all. We have time for some comments or additional discussion, if you want. […]

01:16:36 Participant: [A participant asks about a BMA in which, if the standard error is included, all the other variables appear unimportant.]

01:17:10 Tomas Havranek: If all the other variables appear unimportant once you include the standard error, this could happen when these other variables are somehow related to the way precision is computed, so either related to publication bias or just to the computation of the variance-covariance matrix.

It also depends: do you use weights, inverse-variance weights, in that BMA? We will talk about it later, about inverse-variance weights and so on. But if you use them in BMA, where you have so many variables, it can create a lot of mess, including some of the things that you describe. This could happen. But if it happens when you don't use weights in BMA, that's a bit strange. So you should check the correlations of the individual variables with precision. But it could be that some of them are really important for the way precision is calculated in primary studies. That must be the reason. I can't think of a different explanation.

01:18:42 Participant: [A participant asks whether two people always code the data.]

01:18:49 Tomas Havranek: No, no. Most of our papers nowadays are done by two people, so that they can compare at least some part of the data. Most of them, the new ones. To be honest, when two people code it and then we compare it, it's not 99% agreement. It's slightly less, sometimes more than slightly. So that convinces me that it's important to have this check.

Especially, again, if it is a long-term project that you want to publish well. Or I think what could work is if just one person collects it and the other co-author would just check: select 10%, and then you can sit together and say, okay, what do you think? And then the other person goes back, maybe, and corrects it. Or you can start by collecting the first 10%. It might be the most efficient way.

Collect the first 10% independently and then compare it. And then the one person would use the new consensus to collect the rest of it. I think that should eliminate most of the issues. And it should save some time.

01:20:24 Participant: Yes, because from my experience, the coding is the hardest and the most annoying, and it's so difficult. That's why I'm really looking for ways to do it in a more efficient way. And that's why I think we should start being more restrictive on the number of variables you select.

01:20:43 Tomas Havranek: I think it's a good way: a smaller number of variables. I still think there should be some sort of check; it cannot be just one person. It also relates to our discussion last time about degrees of freedom versus sample size. So I completely agree. It's good to simplify. Then it also makes the paper easier to understand and grasp for other people. Thank you very much. […]

Corrections

Slips of the tongue corrected in the text:

  • [00:10:12] said "percentage points"; the text has "percent".