<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Research notes</title>
    <link>https://meta-analysis.cz/notes/</link>
    <atom:link href="https://meta-analysis.cz/notes/feed.xml" rel="self" type="application/rss+xml" />
    <description>Short, self-contained notes on economics research: methods, tools, findings, and how the work gets done.</description>
    <language>en</language>
    <lastBuildDate>Mon, 27 Jul 2026 00:00:00 +0000</lastBuildDate>
    <item>
      <title>Two new pre-registered papers: outlier decisions in meta-analysis and AI feedback on meta-analyses</title>
      <link>https://meta-analysis.cz/notes/maer-two-pre-registered-papers/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-two-pre-registered-papers/</guid>
      <pubDate>Mon, 27 Jul 2026 00:00:00 +0000</pubDate>
      <description>Two pre-registered papers are announced: recomputing 358 meta-analyses under five outlier treatments changes conclusions in up to 15.9% of cases, and blinded authors of 44 meta-analyses rank a single AI pass above two multi-agent debate tools.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/two-new-pre-registered-papers-outlier-decisions-in-meta-analysis-and-ai-feedback-on-meta" rel="external">MAER-Net</a>, 27 July 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-two-pre-registered-papers/">/komentare/maer-two-pre-registered-papers/</a>.</p>

<p>
My colleagues and I have two new pre-registered papers that may interest MAER-Net members.
</p>

<h2>1. Do decisions about outliers and influential effects matter?</h2>

<p>
(<a href="https://meta-analysis.cz/outliers">https://meta-analysis.cz/outliers</a>, <a href="https://arxiv.org/abs/2607.23174">https://arxiv.org/abs/2607.23174</a>)
</p>

<p>
With Zuzana Irsova, Martina Luskova, and Tom Stanley, we recompute 358 behavioral science meta-analyses under five outlier treatments: do nothing, drop the most extreme estimate, remove studentized residuals above 3, winsorize at 5/95, and remove estimates with |DFBETAS| above 2/sqrt(k). Each runs under random effects and UWLS. All data, thresholds, and rules were registered before we saw any results.
</p>

<p>
The mean effect barely moves: the median absolute change in Cohen's d is at most 0.047. Interpretation moves more. In 11.5% of the meta-analyses at least one treatment changes statistical significance, and in 15.9% whether the effect reaches a smallest effect size of interest (|d| &gt;= 0.20). Nearly all flips are in results already close to the boundary; strongly significant results essentially never change. Winsorizing changes the fewest conclusions, DFBETAS the most, and DFBETAS computed with UWLS flags the most influential estimates. Takeaway: pre-register the outlier rule and report results with and without it.
</p>

<h2>2. Does multi-agent debate improve AI feedback on research papers?</h2>

<p>
(<a href="https://meta-analysis.cz/debate">https://meta-analysis.cz/debate</a>, <a href="https://arxiv.org/abs/2607.14713">https://arxiv.org/abs/2607.14713</a>)
</p>

<p>
Many of you took part in this experiment with Zuzana and me -- thank you! Authors of 44 economics meta-analyses ranked three blinded AI reports on their own paper: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. The single pass won, by 0.66 rank points over mad-research and 0.57 over paper-workshop, although paper-workshop spends about thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; the AI judges almost always put the same human report last. And an independent AI judge (Gemini) would have reversed the authors' verdict and picked the most expensive tool. Takeaway: an AI judge is not a substitute for the author, so be careful with LLM-as-a-judge designs.
</p>

<p>
Both tools are open source: <a href="https://github.com/tjhavranek/mad-research">https://github.com/tjhavranek/mad-research</a> and <a href="https://github.com/tjhavranek/paper-workshop">https://github.com/tjhavranek/paper-workshop</a>
</p>

<p>
Comments are welcome!!
</p>

<figure><img src="https://meta-analysis.cz/komentare/item-img/2026-07_outliers-table3.png" alt="Two-panel table. Panel A, statistical significance at two-sided p below 0.05, gives for random effects, UWLS and pooled each of four outlier treatments — drop-extreme, studentized residual above 3, winsorize 5/95, and DFBETAS above 2 over the square root of k — with the number of results turning significant, the number turning non-significant, and the combined percentage. Discordance runs from 2.23 percent for winsorizing under UWLS to 6.42 percent for DFBETAS under UWLS; at least one treatment changes significance in 7.7 percent of results, 55 of 715. Panel B repeats the layout for whether the effect reaches the smallest effect size of interest, absolute pooled d of at least 0.20: winsorizing changes the fewest, 1.68 to 1.96 percent, and DFBETAS the most, up to 7.54 percent; at least one treatment changes this in 10.3 percent of results, 74 of 715." /></figure>

<figure><img src="https://meta-analysis.cz/komentare/item-img/2026-07_debate-cost-vs-rank.png" alt="Scatter plot with error bars. Author mean rank, where 1 is most useful, runs down an inverted vertical axis; tokens per paper run along a logarithmic horizontal axis. The single pass is both best and cheapest, at a mean rank near 1.6 and roughly 25 thousand tokens, with an interval from about 1.4 to 1.8. mad-research sits near 2.25 at roughly 250 thousand tokens and paper-workshop near 2.15 at about 800 thousand tokens; the intervals of the two multi-agent tools overlap each other but not the single pass." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>Do outlier treatment decisions matter in meta-analysis?</title>
      <link>https://meta-analysis.cz/notes/do-outlier-decisions-matter/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/do-outlier-decisions-matter/</guid>
      <pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate>
      <description>Recomputing 358 behavioral science meta-analyses under five outlier treatments barely moves the mean effect, but changes statistical significance in 11.5% of cases and whether the effect clears a smallest-effect-size threshold in 15.9%, with winsorizing changing conclusions least and DFBETAS most.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/posts/zuzanairsova_so-how-much-do-different-outlier-treatments-activity-7486767472440827905-HD5D" rel="external">LinkedIn</a>, 25 July 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-07-25-do-outlier-decisions-matter/">/komentare/posts/2026-07-25-do-outlier-decisions-matter/</a>.</p>

<p>
So, how much do different outlier treatments matter for meta-analysis results?
</p>

<p>
We have just preprinted a study that recomputes 358 behavioral science meta-analyses under five different outlier treatments.
</p>

<p>
The mean effect barely moves: the median absolute change in Cohen’s d is at most 0.047 and usually much smaller. But the interpretation moves more. In 11.5% of the meta-analyses, at least one treatment changes statistical significance; in 15.9%, it changes whether the effect reaches a smallest effect size of interest.
</p>

<p>
Winsorizing changes conclusions least often; DFBETAS changes them most often.
</p>

<p>
Paper, online appendix, and data: <a href="https://meta-analysis.cz/outliers">https://meta-analysis.cz/outliers</a> Pre-registration: <a href="https://doi.org/10.17605/OSF.IO/97CMV">https://doi.org/10.17605/OSF.IO/97CMV</a> Replication package: <a href="https://doi.org/10.5281/zenodo.21216506">https://doi.org/10.5281/zenodo.21216506</a>
</p>

<p>
Joint work with Tomas Havranek, Martina Lušková, and T. D. Stanley
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-07-25_outliers_title.png" alt="Title page of the working paper Do decisions about outliers and influential effects matter? Evidence from 358 behavioral science meta-analyses by Tomas Havranek, Zuzana Irsova, Martina Luskova and T. D. Stanley, dated July 2026, with the full abstract below the title." /></figure>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-07-25_outliers_table.png" alt="Table 3 of the paper, Changes in statistical significance and smallest effect size of interest. Two panels compare four outlier treatments (drop-extreme, absolute studentized residual above 3, Winsorize 5/95, and absolute DFBETAS above 2 over the square root of k) against the do-nothing baseline, each under random-effects and unrestricted weighted least squares estimators. Winsorizing changes the fewest conclusions and DFBETAS the most." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>Does multi-agent AI debate improve feedback on research papers?</title>
      <link>https://meta-analysis.cz/notes/multi-agent-debate-ai-feedback/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/multi-agent-debate-ai-feedback/</guid>
      <pubDate>Thu, 16 Jul 2026 00:00:00 +0000</pubDate>
      <description>No, at least not for economics meta-analyses. Authors of 44 meta-analyses ranked three blinded AI reports on their own paper; a single prompt beat two multi-agent debate tools, one of which spent thirty times the tokens.</description>
      <content:encoded><![CDATA[<p>
Probably not. We built two multi-agent debate tools, expected them to win, and they lost to a single prompt. This note summarises what we did and what we found; the full paper, the pre-registration, and the replication package are linked in the sidebar.
</p>

<h2>The question</h2>

<p>
Feedback from large language models is now cheap enough that many researchers use it on their own drafts. A natural conjecture is that spending more computation at inference time should produce better feedback: let several agents argue, critique each other, and converge. We wanted to know whether that conjecture survives contact with the people best placed to judge, namely the authors of the papers being reviewed.
</p>

<h2>What we did</h2>

<p>
We ran a pre-registered, identity-masked, within-paper experiment. For each of 44 meta-analyses in economics, we generated three AI reports on that paper: one from a single pass by a frontier model, and one from each of two multi-agent debate tools we had written ourselves, <a href="https://github.com/tjhavranek/mad-research">mad-research</a> and <a href="https://github.com/tjhavranek/paper-workshop">paper-workshop</a>. All three were held to a common length and template, so authors could not tell them apart by format. We then asked the authors to rank the three reports by how useful each would be for improving their own paper. The study was registered before any report was generated.
</p>

<h2>What we found</h2>

<p>
Authors preferred the single pass. It beat <i>mad-research</i> by 0.66 rank points (95% CI 0.32 to 1.00) and <i>paper-workshop</i> by 0.57 (0.16 to 0.95). This is despite <i>paper-workshop</i> spending roughly thirty times the tokens: about 800,000 per paper against about 27,000 for the single pass.
</p>

<p>
Two further results struck us as more interesting than the headline.
</p>

<p>
First, authors who could recall the referee report they had received from a journal usually ranked it above every AI report, and never ranked it last. When we asked AI judges to rank the same material, they almost always put the human referee report last. Author and AI rankings agree only weakly, with a correlation of 0.14.
</p>

<p>
Second, the choice of AI judge can reverse the finding. Gemini, the only judge whose model family had written none of the reports, would have ranked <i>paper-workshop</i> first in the authors' place. That is the opposite of what the authors themselves said. The warning here is narrow but sharp: an AI judge is not a substitute for the author, and a judge drawn from the same family as one of the systems under test is not a neutral instrument.
</p>

<h2>What this does not show</h2>

<p>
We measured perceived usefulness, judged by authors, on finished papers. That is not the same as measuring whether a report would improve a paper if acted on, and it is not the same as asking whether AI should referee papers at all. Both are separate questions and we do not answer them here. Our papers are meta-analyses in economics, so the result may not carry to other designs or fields.
</p>

<h2>Why we are publishing the negative result</h2>

<p>
We built these tools, we expected them to win, and we pre-registered the comparison before finding out. Reporting the outcome either way was the point of registering it. Both tools remain open source and are linked above; the finding is about test-time compute in this setting, not about whether the tools are useful for anything.
</p>]]></content:encoded>
    </item>
    <item>
      <title>A new AI tool for reviewing ERC grant proposals</title>
      <link>https://meta-analysis.cz/notes/a-tool-for-your-erc-proposal/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/a-tool-for-your-erc-proposal/</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate>
      <description>A new tool checks ERC grant drafts against the official evaluation criteria and rules, flagging routine weak spots before human reviewers see the proposal. It is meant to clear easy problems, not replace human review.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7475160943409135616" rel="external">LinkedIn</a>, 23 June 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-06-23-a-tool-for-your-erc-proposal/">/komentare/posts/2026-06-23-a-tool-for-your-erc-proposal/</a>.</p>

<p>
If you plan to work on your ERC proposal this summer, the following tool might help you:
</p>

<p>
<a href="https://github.com/tjhavranek/erc-ai-feedback">https://github.com/tjhavranek/erc-ai-feedback</a>
</p>

<p>
It checks your draft at various stages against the official ERC evaluation criteria and rules, flags the routine weak spots, gives you tips, etc. The idea is to clear the easy problems before you ask humans, not to replace them.
</p>

<p>
Just read the privacy note first: use a paid model with training turned off, or don't paste anything sensitive.
</p>

<p>
We built this with Tomas Havranek, who was on the ERC Advanced economics panel in 2020 and is involved in the expert group supporting ERC applicants in Czechia.
</p>

<p>
You might also check our related tools: mad-research (https://github.com/tjhavranek/mad-research), which stress-tests your paper or proposal with a Claude/Codex debate, and paper-workshop (https://github.com/tjhavranek/paper-workshop), which simulates an expert workshop on your stuff.
</p>

<p>
If you have any comments or feedback, we'll be happy to incorporate them into the tools.
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-06-23_p5_1.jpeg" alt="Screenshot of a repository page headed erc-ai-feedback, describing a small package that gives ERC Starting and Consolidator Grant applicants a rubric-based pre-review of a draft proposal in one chat session with a frontier model, to clear routine structural problems before workshop time is spent on them. It notes that it does not replace human review." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>A simulated expert panel that stress-tests and rebuilds your paper</title>
      <link>https://meta-analysis.cz/notes/a-panel-of-experts-for-your-paper/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/a-panel-of-experts-for-your-paper/</guid>
      <pubDate>Wed, 10 Jun 2026 00:00:00 +0000</pubDate>
      <description>The paper-workshop Claude Code skill assembles a simulated panel of experts to stress-test a research paper, then rebuilds it in tracked changes with a replication package, re-running the author&#x27;s Stata and R code.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7470445147063803904" rel="external">LinkedIn</a>, 10 June 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-06-10-a-panel-of-experts-for-your-paper/">/komentare/posts/2026-06-10-a-panel-of-experts-for-your-paper/</a>.</p>

<p>
Imagine a panel of the world's leading experts, assembled for your research paper, arguing it out from rival schools and then revising it themselves.
</p>

<p>
Well, that's not possible. But we built an approximation with Claude Code, and it now uses the full power of Anthropic's Mythos-class model, Claude Fable 5.
</p>

<p>
Below is the skill, open and free. You give it your paper, ideally with the data and draft code. You get a stress test of your paper + a revision in track changes + a replication package. We ran it end to end on colleagues' papers and our own paper accepted at JPE Micro, Stata and R re-runs included. Enjoy:
</p>

<p>
<a href="https://github.com/tjhavranek/paper-workshop">https://github.com/tjhavranek/paper-workshop</a>
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-06-10_p6_1.jpeg" alt="Screenshot of a page headed CRUCIBLE, the paper-workshop skill: a panel of leading experts assembled for one specific paper, arguing it out from rival schools and then rebuilding it themselves, re-running the author's code. It promises a tracked redline, a clean draft and a replication package." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>Stress-testing research with AI, now super easy and fully automated</title>
      <link>https://meta-analysis.cz/notes/maer-stress-testing/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-stress-testing/</guid>
      <pubDate>Sun, 31 May 2026 00:00:00 +0000</pubDate>
      <description>The research-stress-testing protocol is now automated as a Claude Code skill: describe a task in one sentence, and Claude calls OpenAI&#x27;s Codex to run critique and synthesis rounds and returns a memo with the full debate trail.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/stress-testing-research-with-ai-now-super-easy-and-fully-automated" rel="external">MAER-Net</a>, 31 May 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-stress-testing/">/komentare/maer-stress-testing/</a>.</p>

<p>
Last December we shared a protocol for stress-testing (meta-)research by making AI models argue and keeping what survives. But running it by hand (opening several models, copying outputs back and forth) is a chore and our original automation via a GPT agent was not reliable.
</p>

<p>
So we automated the protocol using Claude Code. After a one-time setup it is a single sentence: you describe the task, and the skill (built with Zuzana Irsova) has Claude call OpenAI's Codex, runs the critique and synthesis rounds, and gives you a memo with the full trail of the debate.
</p>

<p>
<a href="https://github.com/tjhavranek/mad-research">https://github.com/tjhavranek/mad-research</a>
</p>

<p>
We built it with meta-analysis in mind, but it works on any paper, proposal, or task. More generally, it is a simple way to have one AI check another's work. A worked example (WAIVE vs. MAIVE) is included so you can see what it produces.
</p>

<p>
The manual protocol (including a four-model version with Gemini and Grok in addition to Claude and GPT) is still here if you prefer copy-paste:
</p>

<p>
<a href="https://github.com/tjhavranek/research-audit-duel-protocol">https://github.com/tjhavranek/research-audit-duel-protocol</a>
</p>

<p>
Try it on something you are working on and tell us how we could improve it!
</p>]]></content:encoded>
    </item>
    <item>
      <title>Claude Code can call Codex to stress-test its own work</title>
      <link>https://meta-analysis.cz/notes/claude-code-with-codex-stress-test/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/claude-code-with-codex-stress-test/</guid>
      <pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate>
      <description>A Claude Code skill can call OpenAI&#x27;s Codex to stress-test its own work, letting researchers switch between the two tools or fall back to one when the other hits its usage limit.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7466382253208645634" rel="external">LinkedIn</a>, 30 May 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-05-30-claude-code-with-codex-stress-test/">/komentare/posts/2026-05-30-claude-code-with-codex-stress-test/</a>.</p>

<p>
Do you know that Claude Code can use Codex to stress test its work?
</p>

<p>
I had both installed and used them interchangeably: for some tasks, one seems to be better, and vice versa. Also, sometimes I hit usage limit in one, so I continue in the other.
</p>

<p>
This can be automated easily: work from Claude Code and call Codex when needed. Here is the Claude skill, applied to stress-testing research:
</p>

<p>
<a href="https://github.com/tjhavranek/mad-research">https://github.com/tjhavranek/mad-research</a>
</p>

<p>
This builds on our previous manual protocols for using different AI models to test research ideas or papers:
</p>

<p>
<a href="https://github.com/tjhavranek/research-audit-duel-protocol">https://github.com/tjhavranek/research-audit-duel-protocol</a>
</p>

<p>
Does the skill work for you? What should we change? Should we add Gemini?
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-05-30_p8_1.jpeg" alt="Screenshot showing Claude Code invoking Codex to stress-test its own work, as described in the post." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>Reporting guidelines for meta-analysis, updated for AI</title>
      <link>https://meta-analysis.cz/notes/reporting-guidelines-updated-for-ai/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/reporting-guidelines-updated-for-ai/</guid>
      <pubDate>Wed, 13 May 2026 00:00:00 +0000</pubDate>
      <description>The Journal of Economic Surveys has published updated reporting guidelines for meta-analysis in economics, this time addressing AI use in search, screening, and coding, with a personal recommendation to combine RoBMA with MAIVE and RTMA.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7460202147213791232" rel="external">LinkedIn</a>, 13 May 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-05-13-reporting-guidelines-updated-for-ai/">/komentare/posts/2026-05-13-reporting-guidelines-updated-for-ai/</a>.</p>

<p>
Reporting Guidelines for Meta-Analysis in Economics, updated for AI, just published in the Journal of Economic Surveys: <a href="https://onlinelibrary.wiley.com/doi/10.1111/joes.70116">https://onlinelibrary.wiley.com/doi/10.1111/joes.70116</a>
</p>

<p>
Two practical points I would emphasize (my personal opinion), beyond the reporting checklist itself:
</p>

<p>
1️⃣ If you use AI for searching, screening, or coding, don't rely on a single model. Use meta-analysis thinking: each model is trained differently and on different data (think Claude vs. Grok). Even if one model strictly dominates, there will be useful information in the others, and you need to stress-test your favorite model brutally regardless. We have developed a simple Research Audit Protocol based on Multi-Agent Debate (MAD) for exactly this: <a href="https://github.com/tjhavranek/research-audit-duel-protocol/">https://github.com/tjhavranek/research-audit-duel-protocol/</a>
</p>

<p>
2️⃣ These guidelines intentionally do not recommend any particular methodology. We do so in our 2024 method guidelines (https://onlinelibrary.wiley.com/doi/full/10.1111/joes.12595). Brief update: I think the baseline meta-analysis technique is now Robust Bayesian Meta-Analysis (RoBMA) by František Bartoš, Maximilian Maier, and Eric-Jan Wagenmakers -- a principled way to average over various bias-correction methods. But these methods don't address p-hacking, so RoBMA should be complemented with MAIVE (easy to apply via <a href="https://easymeta.org">https://easymeta.org</a>) and RTMA (Maya Mathur).
</p>

<p>
The updated reporting guidelines were led by Nikolai Cook and co-authored with František Bartoš, Pedro Bom, Sebastian Gechert, Klára Kantová, Jerome Geyer-Klingeberg, Dr.-Ing., Tomas Havranek, Martina Lušková, Matej Opatrny, Franz Prante, Heiko Rachinger, and Tom Stanley.
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-05-13_p9_1.jpeg" alt="Journal of Economic Surveys article header, open access: Reporting Guidelines for Meta-Analysis in Economics, Updated for AI, by Nikolai Cook, Frantisek Bartos, Pedro R. D. Bom, Sebastian Gechert, Klara Kantova, Jerome Geyer-Klingeberg, Tomas Havranek, Zuzana Irsova, Martina Luskova and others. First published 12 May 2026." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>Pre-registering a full redo of the beauty premium meta-analysis</title>
      <link>https://meta-analysis.cz/notes/redoing-a-meta-analysis-from-scratch/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/redoing-a-meta-analysis-from-scratch/</guid>
      <pubDate>Tue, 28 Apr 2026 00:00:00 +0000</pubDate>
      <description>A pre-registered revision redoes a meta-analysis of beauty and professional success using multidisciplinary systematic-review methods: librarian-designed multi-database search, dual screening, coding checks, and PRISMA documentation, as a natural experiment on the team&#x27;s own earlier work.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7454914753266692097" rel="external">LinkedIn</a>, 28 April 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-04-28-redoing-a-meta-analysis-from-scratch/">/komentare/posts/2026-04-28-redoing-a-meta-analysis-from-scratch/</a>.</p>

<p>
Will our results change if we redo the meta-analysis from scratch?
</p>

<p>
We just pre-registered a big revision of our meta-analysis on beauty and professional success:
</p>

<p>
<a href="https://doi.org/10.17605/OSF.IO/B3D7W">https://doi.org/10.17605/OSF.IO/B3D7W</a>
</p>

<p>
The current paper follows standards commonly used in economics meta-analysis. For the revision we decided to do the search and data collection differently, following multidisciplinary systematic-review practice: a librarian-designed multi-database search, dual screening, coding reliability checks, PRISMA documentation, and a full audit trail.
</p>

<p>
A nice side effect is that this becomes a natural experiment on our own work. Honestly, I'm curious how much it will move the results.
</p>

<p>
The amazing Martina Lušková joined the team to help lead study selection and coding, and we are working with a librarian at the University of Amsterdam on the search strategy.
</p>

<p>
In the current version we find that the effect of beauty on earnings is smaller than commonly thought once you correct for publication bias and p-hacking, and smaller still when more weight is given to studies that control for cognitive ability. The one clear exception is sex workers. For politicians, the beauty premium mostly goes away after correction.
</p>

<p>
To make the comparison fully transparent, we also uploaded the current paper, data, and code to the registration.
</p>

<p>
We should do much more pre-registration in observational research. It doesn't fully prevent p-hacking, but it helps a lot and the cost is low.
</p>

<p>
Co-authored with Tomas Havranek, František Bartoš, Xenia Bortnikova, and Martina Lušková
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-04-28_p10_1.jpeg" alt="First page of a pre-registered protocol headed Systematic review, update protocol: Meta-Analysis of Field Studies on Beauty and Professional Success. A table gives the protocol type, describing substantial revisions to the literature search, screening, coding, analysis and reporting of the previous version, and records the registration on the Open Science Framework." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>New guidance on using AI in meta-analysis</title>
      <link>https://meta-analysis.cz/notes/note-in-journal-of-economic-surveys/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/note-in-journal-of-economic-surveys/</guid>
      <pubDate>Wed, 22 Apr 2026 00:00:00 +0000</pubDate>
      <description>A new Journal of Economic Surveys note sets a floor for AI use in meta-analysis: human leadership, human accountability, human auditing of at least 10 percent of screening and coding, and full disclosure of AI use in prompts and models.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7452590050799853568" rel="external">LinkedIn</a>, 22 April 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-04-22-note-in-journal-of-economic-surveys/">/komentare/posts/2026-04-22-note-in-journal-of-economic-surveys/</a>.</p>

<p>
Not an easy task, but in a new note just out in the Journal of Economic Surveys we try to set a basic floor on the use of AI in meta-analysis.
</p>

<p>
<a href="https://onlinelibrary.wiley.com/doi/10.1111/joes.70105">https://onlinelibrary.wiley.com/doi/10.1111/joes.70105</a>
</p>

<p>
Short version:
</p>

<p>
🔹 Human leadership. Humans direct the search, coding, and analysis, and record where they override the AI.
</p>

<p>
🔹 Human accountability. AI cannot be a co-author. If your name is on the paper, the errors are yours.
</p>

<p>
🔹 Human auditing. AI can serve as one of the coders, as long as humans audit at least 10% of screening records and 10% of coded studies (or 100 and 20, whichever is larger), and report a measure of agreement.
</p>

<p>
🔹 Human disclosure. Anything that shapes search, screening, coding, analysis, or conclusions should be disclosed, with prompts and model versions saved.
</p>

<p>
The effort was led by the amazing Nikolai Cook. It will be periodically updated at maer-net.org.
</p>

<p>
For the full discussion of how we agreed on these guidelines (scroll down to comments): <a href="https://www.maer-net.org/post/developing-guidelines-for-the-use-of-ai-in-meta-analysis-of-economics-research-guai-maer-and">https://www.maer-net.org/post/developing-guidelines-for-the-use-of-ai-in-meta-analysis-of-economics-research-guai-maer-and</a>
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-04-22_p11_1.jpeg" alt="Journal of Economic Surveys article header, open access: Guidance for the Use of AI in the Meta-Analysis of Economics Research, by Nikolai Cook, Frantisek Bartos, Pedro R. D. Bom, Sebastian Gechert, Klara Kantova, Jerome Geyer-Klingeberg, Tomas Havranek, Zuzana Irsova, Martina Luskova, Matej Opatrny, Heiko J. Rachinger and T. D. Stanley. First published 21 April 2026." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>Reproducibility numbers from the Brodeur team&#x27;s Nature study</title>
      <link>https://meta-analysis.cz/notes/brodeur-team-reproducibility-numbers/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/brodeur-team-reproducibility-numbers/</guid>
      <pubDate>Thu, 02 Apr 2026 00:00:00 +0000</pubDate>
      <description>A Nature study led by Abel Brodeur&#x27;s team reports reproducibility and robustness numbers for economics and political science research, based on papers from journals with mandatory data and code sharing.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7445392769788968960" rel="external">LinkedIn</a>, 2 April 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-04-02-brodeur-team-reproducibility-numbers/">/komentare/posts/2026-04-02-brodeur-team-reproducibility-numbers/</a>.</p>

<p>
It was fun to be (a very small) part of the amazing team led by Abel Brodeur. The numbers look great, but note that we replicate papers in journals that have mandatory data and code sharing. So: share your data and code! And while you're at it, pre-register your papers and upload pre-analysis plans. This is the way!
</p>]]></content:encoded>
    </item>
    <item>
      <title>Four AI models debating works better than two</title>
      <link>https://meta-analysis.cz/notes/four-models-debating-works-better/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/four-models-debating-works-better/</guid>
      <pubDate>Thu, 19 Mar 2026 00:00:00 +0000</pubDate>
      <description>An updated research audit protocol (MAD v2.0) runs ChatGPT, Claude, Gemini, and Grok through independent critique, cross-examination, and a final synthesis, using copy-paste prompts with no coding required, or full automation via API frameworks.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7440296636783960064" rel="external">LinkedIn</a>, 19 March 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-03-19-four-models-debating-works-better/">/komentare/posts/2026-03-19-four-models-debating-works-better/</a>.</p>

<p>
Two AI models dueling worked. Four models debating works better.
</p>

<p>
We updated our research audit protocol. The new version (MAD v2.0) uses ChatGPT, Claude, Gemini, and Grok in structured adversarial rounds:
</p>

<p>
1️⃣ Independent critique — no model sees the others, every claim grounded in the document. 2️⃣ Cross-examination — each model attacks the weakest peer arguments. 3️⃣ Final arbiter synthesizes what survived.
</p>

<p>
No code needed. Copy-paste prompts. Free model versions work (just register for each model).
</p>

<p>
Advanced users: the entire workflow can be automated via the models' APIs using frameworks like AutoGen, LangGraph, or CrewAI.
</p>

<p>
Use it for high-stakes documents: stress-testing your papers, grant proposals, referee reports.
</p>

<p>
Protocol (GitHub): 👉 <a href="https://github.com/tjhavranek/research-audit-duel-protocol">https://github.com/tjhavranek/research-audit-duel-protocol</a>
</p>]]></content:encoded>
    </item>
    <item>
      <title>Meta-analysis should correct for p-hacking too</title>
      <link>https://meta-analysis.cz/notes/correcting-for-p-hacking-too/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/correcting-for-p-hacking-too/</guid>
      <pubDate>Mon, 09 Mar 2026 00:00:00 +0000</pubDate>
      <description>Corrections for publication bias assume individually unbiased estimates, an assumption p-hacking violates. The Nature Communications MAIVE paper shows that under some forms of p-hacking, classical publication-bias corrections can be more biased than a simple average.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7436800791929151488" rel="external">LinkedIn</a>, 9 March 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-03-09-correcting-for-p-hacking-too/">/komentare/posts/2026-03-09-correcting-for-p-hacking-too/</a>.</p>

<p>
Meta-analyses should try to correct not just for publication bias, but also for p-hacking.
</p>

<p>
Some estimates are more likely to be reported than others, so every good summary of research should correct for this publication bias. In case you're wondering — yes, this can be done, there are dozens of methods and decades of research on this. There is even a great way to put these different correction techniques together: see RoBMA by František Bartoš and colleagues.
</p>

<p>
The problem is that these techniques assume that the reported estimates are individually unbiased. This is a strong assumption, as researchers can tweak models (consciously or unconsciously) to get more "sensible" results. This is called p-hacking. Most of us do it.
</p>

<p>
In our recent Nature Communications paper we show that under some forms of p-hacking, classical models correcting for publication bias can actually be more biased than a simple average of published estimates.
</p>

<p>
As far as I know, there are only two meta-analysis corrections for (some forms of) p-hacking: our MAIVE (from the Nature Comms paper) and RTMA by Maya Mathur. If you know of more, please let me know in the comments!
</p>

<p>
If you want to see MAIVE and RTMA applied, take a look at our meta-analysis of the beauty premium. Spoiler: apart from the sex industry, beauty doesn't matter much in the labor market.
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-03-09_p16_1.jpeg" alt="Funnel plot of estimates of the beauty effect on earnings. The bulk of estimates cluster near zero; a red line marks the mean of 4.3 per cent, and estimates for sex workers, shown separately in red, sit well to the right of it." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>A browser tool for correcting publication bias</title>
      <link>https://meta-analysis.cz/notes/browser-tool-for-publication-bias/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/browser-tool-for-publication-bias/</guid>
      <pubDate>Tue, 27 Jan 2026 00:00:00 +0000</pubDate>
      <description>EasyMeta.org lets researchers upload a dataset and run bias corrections, including MAIVE and PET-PEESE, directly in the browser, with clustering options and exportable R code, and no installation or coding required.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3AugcPost%3A7421813900372918272" rel="external">LinkedIn</a>, 27 January 2026. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2026-01-27-browser-tool-for-publication-bias/">/komentare/posts/2026-01-27-browser-tool-for-publication-bias/</a>.</p>

<p>
A browser-based tool for correcting publication bias and p-hacking.
</p>

<p>
We built EasyMeta.org to make bias-corrected meta-analysis easier to run. Upload your dataset and benchmark your conclusions against modern corrections -- directly in your browser.
</p>

<p>
Why: Selective reporting can distort what enters the literature and how results are reported, but applying corrections often requires specialized software or a complex workflow.
</p>

<p>
What it offers: • Bias corrections: MAIVE (our Nature Communications method) plus benchmarks like PET-PEESE. • Robust inference: options for clustering and different data structures. • Minimal setup: free, open, no installation, no coding. • Reproducibility: export R code for the results you generate.
</p>

<p>
Swipe through the four slides to see the workflow.
</p>

<p>
Try it with your own data: 👉 easymeta.org (https://www.easymeta.org/)
</p>

<p>
With Pedro Bom, Tomas Havranek, Heiko Rachinger, and Petr Čala
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2026-01-27_easymeta.png" alt="First slide of the EasyMeta walkthrough: Seamless Meta-Analysis with MAIVE, adjust your data for publication bias, p-hacking, and spurious precision, with buttons to upload data or run a demo." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>Stress-Testing Meta-Research with AI Duels</title>
      <link>https://meta-analysis.cz/notes/maer-ai-duels/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-ai-duels/</guid>
      <pubDate>Fri, 12 Dec 2025 00:00:00 +0000</pubDate>
      <description>The Research Audit Protocol coordinates ChatGPT and Gemini in a structured, human-in-the-loop duel of anchor assessment, adversarial probing, and synthesis, illustrated with a case study auditing the proposed WAIVE idea against the MAIVE framework.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/ai_duel" rel="external">MAER-Net</a>, 12 December 2025. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-ai-duels/">/komentare/maer-ai-duels/</a>.</p>

<p>
Many of us now use large language models for meta-analysis tasks like coding, short summaries, or quick checks. Where they become genuinely valuable for research, though, is not in producing a single clean answer. It is in creating <b>structured disagreement</b>: two models pushing on each other’s reasoning, with a researcher steering the process.
</p>

<p>
That is the idea behind the <a href="https://github.com/tjhavranek/research-audit-duel-protocol"><b>Research Audit Protocol (v1.7)</b></a>. It is a structured, human-in-the-loop workflow that coordinates ChatGPT and Gemini in a deliberate “duel.” The goal is not AI approval. The goal is to generate the kinds of counterexamples, boundary conditions, and missing assumptions that a single model (and often a single human pass) would not surface. (Of course, advanced users can substitute Claude for the auditor role, or run the same workflow via an API-based multi-agent setup.)
</p>

<p>
<b>Why duels work</b> A normal chat is optimized for flow. A duel is optimized for scrutiny:
</p>

<ul>
<li><b>Anchor (ChatGPT):</b> Start with a full, file-grounded assessment before seeing any critique.</li>
<li><b>Duel (Gemini):</b> Probe hard for identification problems, hidden assumptions, and failure modes—and force specificity.</li>
<li><b>Synthesis (You + ChatGPT):</b> Map the disagreement (or convergence) and record what changed and why, so the final view is auditable.</li>
</ul>

<p>
<b>A MAER-Net Case Study: </b><a href="https://github.com/tjhavranek/research-audit-duel-protocol/tree/main/examples"><b>WAIVE vs. MAIVE</b></a> As a proof of concept for meta-method work, we applied the protocol to an audit of the proposed <b>WAIVE</b> idea against the <b>MAIVE</b> framework. The duel did not just restate the two approaches. It forced us to pin down the key tension that matters for applied meta-analysis: when does downweighting “suspiciously precise” results reduce spurious precision, and when might it also penalize genuinely informative studies?
</p>

<p>
In other words, it pushed us to state the boundary conditions clearly—the kind of slow thinking that improves methods before they hit peer review.
</p>

<p>
<b>Resources</b>
</p>

<ul>
<li><b>Protocol (v1.7):</b> <a href="https://github.com/tjhavranek/research-audit-duel-protocol">GitHub Repository</a></li>
<li><b>Worked Example (WAIVE/MAIVE):</b> <a href="https://github.com/tjhavranek/research-audit-duel-protocol/tree/main/examples">Examples Folder</a></li>
<li><b>Permanent DOI:</b> <a href="https://doi.org/10.5281/zenodo.17898869">Zenodo</a></li>
<li><b>MAIVE Code:</b> <a href="https://cran.r-project.org/package=MAIVE">CRAN</a> | <a href="https://easymeta.org/">EasyMeta.org</a></li>
</ul>]]></content:encoded>
    </item>
    <item>
      <title>MAIVE Is Now on CRAN</title>
      <link>https://meta-analysis.cz/notes/maer-maive-cran/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-maive-cran/</guid>
      <pubDate>Wed, 10 Dec 2025 00:00:00 +0000</pubDate>
      <description>MAIVE, the bias-correction estimator for meta-analysis published in Nature Communications, is now installable directly from CRAN, alongside the existing EasyMeta.org web app that runs MAIVE, PET-PEESE, and the endogenous kink model with one click.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/maive-is-now-on-cran" rel="external">MAER-Net</a>, 10 December 2025. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-maive-cran/">/komentare/maer-maive-cran/</a>.</p>

<p>
We are happy to share that <b>MAIVE</b> (Meta-Analysis Instrumental Variable Estimator) is now available on <b>CRAN</b>:
</p>

<p>
<code>install.packages("MAIVE")</code>
</p>

<p>
MAIVE is a simple tool that helps detect when study results look <i>too precise</i> because of method choices or selective reporting, and it provides an IV-based correction. The method was published recently in <i>Nature Communications</i>.
</p>

<p>
MAIVE is also available in the <a href="https://www.easymeta.org">EasyMeta.org</a> web app, where you can run MAIVE, PET-PEESE, and the EK model with one click, no coding needed:
</p>

<p>
👉 <a href="https://www.easymeta.org">https://www.easymeta.org</a>
</p>

<p>
👉 CRAN package: <a href="https://CRAN.R-project.org/package=MAIVE">https://CRAN.R-project.org/package=MAIVE</a>
</p>

<p>
👉 Nature Communications article: <a href="https://www.nature.com/articles/s41467-025-63261-0">https://www.nature.com/articles/s41467-025-63261-0</a>
</p>

<p>
Many thanks to <b>Petr Čala</b> for preparing the CRAN release and building the EasyMeta interface.
</p>]]></content:encoded>
    </item>
    <item>
      <title>Highlights from the 2025 MAER-Net Colloquium in Ottawa</title>
      <link>https://meta-analysis.cz/notes/maer-ottawa-colloquium/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-ottawa-colloquium/</guid>
      <pubDate>Fri, 31 Oct 2025 00:00:00 +0000</pubDate>
      <description>A recap of the 2025 MAER-Net Colloquium in Ottawa, where Abel Brodeur received the Founders&#x27; Medal and Shinichi Nakagawa and Andrew Gelman gave keynotes; the 2026 colloquium moves to Chemnitz, Germany, hosted by Sebastian Gechert.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/highlights-from-the-2025-maer-net-colloquium-in-ottawa" rel="external">MAER-Net</a>, 31 October 2025. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-ottawa-colloquium/">/komentare/maer-ottawa-colloquium/</a>.</p>

<p>
The Ottawa Colloquium was a wonderful experience and a great success. The quality of presentations and the richness of discussions have never been better. AI featured prominently this year, sparking lively debates that continue on <a href="https://www.maer-net.org/post/developing-guidelines-for-the-use-of-ai-in-meta-analysis-of-economics-research-guai-maer-and">MAER-Net’s blog</a>. We also had many methodological contributions, each offering new ways to understand economics research or to avoid misinterpretation.
</p>

<p>
Of particular note were the excellent keynotes and the awarding of our Founder’s Medal. Shinichi Nakagawa adeptly presented five fascinating meta-meta studies from ecology and environmental science, and Andrew Gelman’s entertaining talk-and-chalk reminded us that “all of our default models are wrong.” Abel Brodeur received the <a href="https://www.maer-net.org/award">MAER-Net’s Founders’ Medal</a> for his outstanding contributions to economic science, including his Herculean efforts to make economics research reproducible and replicable.
</p>

<p>
Our sincere thanks go to Abel and the University of Ottawa for their gracious and generous hospitality. Abel and his team offered a master class in hosting research conferences: from meticulous planning and catering to expert tour guiding and insightful commentary on others’ research.
</p>

<p>
We also thank our core members for their continued support and willingness to share their work, and our brilliant young researchers for their eagerness to learn and to further develop meta-analytic methods. See the Ottawa Colloquium website for more details.
</p>

<p>
Lastly, we are pleased to announce that the 2026 MAER-Net Colloquium will be held at Chemnitz University of Technology, Germany, hosted by Sebastian Gechert and his team. Chemnitz is the 2025 European Capital of Culture, known for its heritage as a powerhouse of early industrialisation in Central Europe, Art Nouveau architecture, its changeful history before, during and after the German partition, and its recent reinvention as a city of subculture and indie-rock-music. The region around Chemnitz, including the Ore Mountains is characterized by a high density of medieval, Renaissance and Baroque castles as well as picturesque hiking and cycling routes. We are very much looking forward to seeing you in Chemnitz, September 23-25, 2026. See the <a href="https://www.maer-net.org/2026-chemnitz">Chemnitz Colloquium website</a> for more details and updates.
</p>]]></content:encoded>
    </item>
    <item>
      <title>Spurious precision in meta-analysis, published in Nature Communications</title>
      <link>https://meta-analysis.cz/notes/spurious-precision-nature-communications/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/spurious-precision-nature-communications/</guid>
      <pubDate>Sun, 26 Oct 2025 00:00:00 +0000</pubDate>
      <description>Nature Communications has published the MAIVE paper, showing that meta-analyses can be misled when a study&#x27;s reported precision reflects method choices rather than real evidence strength, and introducing a correction, MAIVE, for this bias.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7388172095639265282" rel="external">LinkedIn</a>, 26 October 2025. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/posts/2025-10-26-spurious-precision-nature-communications/">/komentare/posts/2025-10-26-spurious-precision-nature-communications/</a>.</p>

<p>
🔎 New in Nature Communications: “Spurious precision in meta-analysis of observational research.”
</p>

<p>
Sometimes studies appear too precise because their reported uncertainty reflects method choices rather than real evidence strength. This can mislead meta-analyses.
</p>

<p>
We introduce MAIVE, a simple way to detect and correct such bias, including publication bias and p-hacking.
</p>

<p>
🖥️ Try it in your browser (free, no coding): <a href="https://www.easymeta.org">https://www.easymeta.org</a>
</p>

<p>
📄 Paper: <a href="https://www.nature.com/articles/s41467-025-63261-0">https://www.nature.com/articles/s41467-025-63261-0</a>
</p>

<p>
📰 Blog: <a href="https://communities.springernature.com/posts/spurious-precision-in-meta-analysis-of-observational-research">https://communities.springernature.com/posts/spurious-precision-in-meta-analysis-of-observational-research</a>
</p>

<p>
— Kudos to my great co-authors Pedro Bom, Tomas Havranek, and Heiko Rachinger
</p>

<figure><img src="https://meta-analysis.cz/komentare/social-img/2025-10-26_spurious.png" alt="Two funnel plots side by side from the Nature Communications paper. In (a), selection on estimates, the surviving studies sit above the t=1.96 line and the effect size is inflated. In (b), selection on standard errors, the funnel is equally asymmetric but the effect size is not inflated. Both plot standard errors against estimates." /></figure>]]></content:encoded>
    </item>
    <item>
      <title>Spurious Precision in Meta-Analysis</title>
      <link>https://meta-analysis.cz/notes/springer-spurious-precision/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/springer-spurious-precision/</guid>
      <pubDate>Sat, 27 Sep 2025 00:00:00 +0000</pubDate>
      <description>Meta-analyses give more weight to precise studies. But what if the reported precision is spurious? We introduce MAIVE, a new estimator that tackles this problem.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://communities.springernature.com/posts/spurious-precision-in-meta-analysis" rel="external">Springer Nature Research Communities</a>, 27 September 2025. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/springer-spurious-precision/">/komentare/springer-spurious-precision/</a>.</p>

<p>
When I began working on meta-analyses two decades ago, I assumed the standard tools, developed mainly with experiments in mind, were safe to apply to observational research. Meta-analysis sounds straightforward: gather comparable estimates from many studies, give more weight to the precise ones, apply rigorous bias-correction techniques, and arrive at reliable averages across different contexts.
</p>

<p>
After dozens of meta-analyses in the social sciences, we kept noticing something odd. Reported effects were often too strongly correlated with their standard errors. And the correlation persisted even after we applied bias corrections using selection models. That worried us. If the link between estimated effects and precision arises beyond standard publication bias, the ratonale of inverse-variance weighting, the backbone of meta-analysis, starts to crumble.
</p>

<h2>Spurious precision</h2>

<p>
One mechanism that can cause this problem is <i>spurious precision</i>. In theory, standard errors should measure study uncertainty. They are given to the researcher, and meta-analysts can use them as one indicator of study reliability. That is the clean textbook story.
</p>

<p>
In practice, reported precision is rarely that clean. Researchers make many choices: which controls to include, how to cluster standard errors, how to treat outliers, which estimation method to use. Ignoring heteroskedasticity often leads to artificially low standard errors, and commonly used cluster-robust standard errors are downward biased in small- and medium-sized samples.
</p>

<p>
Standard errors may also be underestimated in experiments due to violations of randomization assumptions, finite sample issues, or model misspecifications. And because significant results are easier to publish, there are incentives to favor smaller standard errors. A small effect with an even smaller standard error often looks better than a large but insignificant one.
</p>

<h2>When existing methods fail</h2>

<p>
The funnel plot below illustrates what happens with spurious precision. Each dot represents a study's estimate against its standard error. Ideally, the funnel narrows symmetrically toward the true effect. Standard corrections account for asymmetry from publication bias or p-hacking on the effect size. But when precision is p-hacked as well, some estimates appear both larger and more precise than they really are: the hollow dots become solid black ones.
</p>

<p>
<i>Figure: Spurious precision makes some studies look more precise than reality, distorting meta-analysis results.</i>
</p>

<p>
These overly precise estimates (e.g. those that omit an important control variable) receive too much weight in meta-analysis. That is how spurious precision creeps in and distorts the overall result. In such cases, all standard inverse-variance weighting methods begin to break down.
</p>

<p>
Publication bias has been studied for decades, and correction methods like PET-PEESE or selection models are widely used. The catch is that they still rely heavily on reported precision and assume that the most precise estimates are unbiased. But if precision itself is p-hacked, these corrections do not necessarily solve the problem.
</p>

<p>
Consider the funnel plot above. Standard inverse-variance-weighted averages are biased upward. Funnel-plot corrections fail because they assume any p-hacking targets effect sizes rather than standard errors. And selection models break down as well, since they rely on the assumption that individual reported estimates are unbiased, an assumption violated under p-hacking.
</p>

<p>
Indeed, in some of our simulations with moderate spurious precision, simple unweighted averages ended up less biased than sophisticated bias-correction estimators.
</p>

<h2>The idea behind MAIVE</h2>

<p>
This led us to MAIVE, the <b>Meta-Analysis Instrumental Variable Estimator</b>. The idea is simple. Reported standard errors may be biased, but sample size is much harder to p-hack. It is strongly (if imperfectly) linked to true precision and, in much of empirical research, authors typically use the largest feasible sample from the start. You can easily change controls, clustering, or estimation techniques; you usually cannot easily increase <i>N</i>.
</p>

<p>
MAIVE uses sample size as an instrument for precision. We keep the familiar funnel-based framework but rebuild it on a more robust foundation: predicted precision based on <i>N</i>, not reported precision alone, with confidence intervals considering prediction uncertainty. In simulations, MAIVE substantially reduced the overall bias (publication and p-hacking) compared to existing estimators. And in datasets with replication benchmarks, MAIVE moved meta-analytic results closer to the replications.
</p>

<h2>Making it easy to use</h2>

<p>
To make MAIVE easy to apply, we built a simple web tool: <a href="https://spuriousprecision.com/">spuriousprecision.com</a>. Upload your dataset, click a button, and MAIVE runs in seconds, correcting your meta-analysis for publication bias, p-hacking, and spurious precision. If you just want to see MAIVE in action, the site offers a <a href="https://spuriousprecision.com/demo">demo dataset</a> you can run instantly.
</p>

<p>
The tool supports study-level clustering (CR1, CR2, wild bootstrap), allows for extreme heterogeneity and weak instruments, and enables fixed-intercept multilevel specifications that account for within-study dependence and between-study differences in methods and quality. No coding, no setup, no R or Stata needed.
</p>

<p>
In the web tool we also included existing funnel-based methods — <b>PET-PEESE</b> and the <b>Endogenous Kink</b> model. Users can compare approaches side by side. This fills a gap: there are web tools for producing forest plots and basic meta-analysis models, but none that let you run advanced funnel-based corrections straight from the browser.
</p>

<h2>What we learned</h2>

<p>
The biggest takeaway was how damaging spurious precision can be. Publication bias is well known, but the meta-analysis bias working through distorted standard errors can be just as harmful. Because almost every meta-analytic method relies on inverse-variance weights, the problem is potentially widespread.
</p>

<p>
We do not claim MAIVE is a cure-all. But it is a practical safeguard against a form of bias that existing methods largely ignore. With the web tool at <a href="https://spuriousprecision.com/">spuriousprecision.com</a> (also available at <a href="http://easymeta.org/">easymeta.org</a>), we hope MAIVE becomes a useful part of the toolbox for meta-analysis, whether in medicine, psychology, economics, or beyond.
</p>]]></content:encoded>
    </item>
    <item>
      <title>Bias Correction Made Easy: A Web App for Meta-Analysis at EasyMeta.org</title>
      <link>https://meta-analysis.cz/notes/maer-easymeta/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-easymeta/</guid>
      <pubDate>Sat, 27 Sep 2025 00:00:00 +0000</pubDate>
      <description>Run MAIVE, PET-PEESE, and EK with one click, no coding, no installation.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/bias-correction-made-easy-a-web-app-for-meta-analysis-at-easymeta-org" rel="external">MAER-Net</a>, 27 September 2025. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-easymeta/">/komentare/maer-easymeta/</a>.</p>

<p>
We are happy to share a new tool now available at <a href="https://www.easymeta.org/">EasyMeta.org</a>. The app makes the new MAIVE method (Meta-Analysis Instrumental Variable Estimator, published yesterday in <a href="https://www.nature.com/articles/s41467-025-63261-0"><i>Nature Communications</i></a>) easy to apply with a few clicks. The site offers a <a href="https://www.spuriousprecision.com/demo">demo dataset</a> you can run instantly.
</p>

<p>
At the same time, the app allows for seamless use of methods well known in the MAER-Net community: PET-PEESE and the Endogenous Kink model. Until now, these approaches were available only in R or Stata. Now they can be run with a single click, no software installation needed.
</p>

<p>
The app supports options that applied researchers often need, including different types of clustering (classical, CR2, wild bootstrap), weighting schemes (including one that accounts for extreme heterogeneity), weak-instrument-robust confidence intervals, and fixed-intercept multilevel specifications to account for between-study differences in methods and quality.
</p>

<p>
We hope this tool will lower entry barriers to modern meta-analysis for students, collaborators, and applied researchers, and make it easier for MAER-Net members to demonstrate, compare, and teach bias-correction methods.
</p>

<p>
We'd be very grateful if you could use your favorite social network to share this tool with colleagues, students, or collaborators who might find it useful. And please let us know about any errors, missing features, or ideas for improvement — we'll keep refining the app based on your feedback. Together we can make bias correction in meta-analysis more accessible.
</p>

<p>
Special thanks to <b>Petr Čala</b> for creating the app, and to <b>Heiko Rachinger</b> and <b>Pedro Bom</b> for developing the R package that powers it.
</p>

<p>
👉 Try a demo at <a href="https://www.spuriousprecision.com/demo">EasyMeta.org/demo</a>
</p>]]></content:encoded>
    </item>
    <item>
      <title>AI Tools for Meta-Analysis</title>
      <link>https://meta-analysis.cz/notes/maer-ai-tools/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-ai-tools/</guid>
      <pubDate>Wed, 30 Jul 2025 00:00:00 +0000</pubDate>
      <description>A practical rundown of how ChatGPT&#x27;s deep research, o3, and Agent modes speed up literature search and data collection for meta-analysis, with the reminder that competent humans still need to check every AI-assisted extraction.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/ai-tools-for-meta-analysis" rel="external">MAER-Net</a>, 30 July 2025. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-ai-tools/">/komentare/maer-ai-tools/</a>.</p>

<p>
I believe we’ve finally reached the point where AI tools can genuinely save a lot of time in meta-analysis during literature search and data collection. Below is a summary of what works well as of late July 2025. My experience is mostly with ChatGPT, so I’d be interested to hear what has worked for others. There must be better ways of doing this.
</p>

<h2>A) Literature search</h2>

<ul>
<li><b>Start with deep research. </b>The deep research function in ChatGPT helps you explore the literature before you begin the actual meta-analysis. It improves your understanding of the topic and also improves how useful ChatGPT replies become in later stages.</li>
<li><b>Refine search strategy using o3. </b>Upload a few key papers and ask the o3 reasoning model to help you create a good Google Scholar query, also based on the deep research results. Go back and forth with it until the first page of results gives you mostly relevant studies.</li>
<li><b>Use ChatGPT Agent to scan results. </b>The Agent can go through the first few hundred hits on Google Scholar and look at the abstracts. Ask it to flag papers that have at least a small chance of including the type of estimates you want. <i>Example prompt: You are an expert in meta-analysis. For each abstract and based on all information you can find about the paper, say whether the paper reports, with more than 20 percent probability, quantitative estimates on the effect of beauty on labor market outcomes. Justify briefly.</i></li>
<li><b>Download flagged papers and filter again. </b>The Agent can download some PDFs but usually not those behind paywalls, so you’ll have to do that part. Upload them to ChatGPT in batches of five. Ask whether the papers include effect size estimates and where in the paper they are located (table numbers or page references). <i>Example prompt: For each of these papers, does it contain new empirical estimates of the effect of beauty on wages that I can use in a meta-analysis? Point to specific tables or pages.</i></li>
<li><b>Cross-check with other tools. </b>Use ASReview or Semantic Scholar to double-check or supplement your results. ASReview is good if you want to train an active learning model on what counts as relevant in your case.</li>
<li><b>Use snowballing to expand the dataset. </b>Ask the Agent to identify studies that are closely connected to those already in your list. For example, you can request the 100 most frequently cited papers <i>by</i> the studies in your dataset, or the 100 most recent papers that <i>cite</i> at least five of them. (The latter helps surface newer work that may be underrepresented in standard Google Scholar results due to low citation counts.) Then return to step 4 and screen these new candidates for relevance.</li>
</ul>

<h2>B) Data collection</h2>

<ul>
<li><b>Upload papers and identify coding dimensions. </b>Upload your included papers in small batches. Ask ChatGPT to identify the main ways in which the studies differ, both in terms of methodology and data. These differences typically become your coding dimensions (e.g. IV vs. RDD vs. DID, different professions examined, different time periods).</li>
<li><b>Use o3 to review the coding structure. </b>Before you finalize the structure of your dataset, ask the o3 reasoning model to help you think through whether the coding dimensions you’ve selected make sense. Adjust if needed. Then fix the structure yourself. Keep all related chats in one project folder in ChatGPT.</li>
<li><b>Train the Agent on a few coded papers. </b>Code five papers manually. Upload the PDFs and your data entries for three of them. Then ask ChatGPT Agent to extract the same information from the other two using GPT-4.1. Iterate your prompt until it works well, then continue in batches. <i>Example prompt: Here are three PDFs and the corresponding rows from my Excel sheet. Learn the structure. These data are collected well and I want you to collect data from other papers. I will now send two new PDFs. Extract estimates of the beauty premium and the corresponding standard error. Collect information on the estimation method used, the definition of the beauty variable, definition of the earnings variable.</i></li>
<li><b>Use NotebookLM as a separate check. </b>Upload the same batch to NotebookLM. Let it try the same extraction task independently. Compare the two outputs. Manually resolve discrepancies and randomly spot-check the rest.</li>
</ul>

<p>
As of July 2025, you still need competent humans to collect data for meta-analysis. All of the above should be complemented with your usual expertise. You have to check carefully what the AI returns. There’s still a risk of hallucinations, although with good prompting the risk is quite small. But where you needed four co-authors a year ago, now you probably need just two. And they can focus on more intellectually demanding tasks.
</p>

<p>
Let me know if you’ve found better ways to do this. I’m especially curious about tools outside the OpenAI ecosystem. Some colleagues report success with Claude, which seems strong in reasoning, but I haven’t tested it systematically yet.
</p>

<p>
<b>Update (July 31, 2025):</b> Several MAER-Net colleagues wrote to me about a promising new platform called <a href="https://ottosr.com/"><i>otto-SR</i></a>. It offers an end-to-end AI pipeline for systematic reviews, using tools like GPT-4.1 and o3 for screening and data extraction: essentially a polished, integrated version of the workflow I described above. It’s currently in preview only and not publicly available. While the <a href="https://www.medrxiv.org/content/10.1101/2025.06.13.25329541v2">early results</a> sound more than impressive, otto-SR has been designed and benchmarked mainly for Cochrane-style reviews, which differ quite a bit from meta-analyses in economics. Still, it’s worth watching closely.
</p>

<p>
<b>Update (August 19, 2025):</b> ChatGPT now runs on GPT‑5. In the UI you pick Auto, Fast, or Thinking; I use GPT‑5 Thinking for screening and PDF extraction because it reasons longer and stays more consistent. Agent mode also feels steadier, but careful human checks are still essential. (If you work via API, OpenAI still documents separate “reasoning” models like o3, but that label isn’t shown in the ChatGPT menu.)
</p>]]></content:encoded>
    </item>
    <item>
      <title>New Challenges for Meta-Analysis: Attenuation Bias, P-Hacking, Preferred Estimates</title>
      <link>https://meta-analysis.cz/notes/maer-new-challenges/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-new-challenges/</guid>
      <pubDate>Mon, 07 Jul 2025 00:00:00 +0000</pubDate>
      <description>Our recent meta-analyses highlight three issues for the field: attenuation bias can rival publication bias in distorting results; new methods like MAIVE address p-hacking more effectively; and author preferred estimates may systematically differ from others.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/new-challenges-for-meta-analysis-attenuation-bias-p-hacking-preferred-estimates" rel="external">MAER-Net</a>, 7 July 2025. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-new-challenges/">/komentare/maer-new-challenges/</a>.</p>

<p>
My colleagues and I have recently published or revised three meta-analyses, each raising issues that may matter for how we do meta-analysis in economics. I’d be grateful for thoughts or feedback -- here, by e-mail, or in person at our colloquium in Ottawa.
</p>

<p>
<b>1. Attenuation bias (</b><a href="https://meta-analysis.cz/skill/"><b>Review of Economics and Statistics, 2024</b></a><b>)</b>
</p>

<p>
Attenuation bias, aka regression dilution, arises when explanatory variables are measured with (classical random) error, biasing regression coefficients toward zero. We show that attenuation bias can be quantitatively important in estimating the elasticity of substitution between skilled and unskilled labor, although publication bias is still the bigger problem. We’re now working on comparing the two biases more broadly across meta-analyses in economics. Is it possible that on average, the two wrongs make a right? On a technical note, we also argue it’s risky to meta-analyze inverted regression coefficients -- especially common when elasticities are estimated primarily in their inverse form. Working with these transformed estimates can violate key meta-analysis assumptions. We should meta-analyze the originally reported regression coefficients.
</p>

<p>
<b>2. P-hacking (</b><a href="https://meta-analysis.cz/incentives/"><b>JPE Microeconomics, revised &amp; resubmitted</b></a><b>)</b>
</p>

<p>
Publication bias and p-hacking are commonly treated as the same problem, and in many settings, this is defensible: both are often observationally equivalent and create a correlation between effect sizes and standard errors. But the distinction can matter for methods. Selection models don’t handle p-hacking, while meta-regression (such as PET-PEESE) is robust to some forms. In this paper on the effect of financial incentives on performance we apply MAIVE, a new extension of PET-PEESE <a href="https://meta-analysis.cz/maive/">forthcoming in Nature Communications</a> robust to more p-hacking strategies that work on precision (such as clustering choices or changing controls) as well as omitted-variable bias in meta-regression. Our results, in the Nature C and JPE Micro papers, suggest a need for more attention to the mechanisms behind selective reporting and for broader adoption of estimators robust to p-hacking, not just MAIVE but also RTMA developed by Maya Mathur.
</p>

<p>
<b>3. Preferred estimates (</b><a href="https://meta-analysis.cz/class/"><b>Journal of Labor Economics, forthcoming</b></a><b>)</b>
</p>

<p>
Most economics meta-analyses collect all reported estimates -- a good default. But some estimates are clearly marked by the original authors as less or more trustworthy. The entire point of some papers is that a particular estimation strategy is wrong. In our class size meta-analysis, we classified estimates as “preferred,” “neutral,” or “discounted” according to how study authors described them. Preferred estimates were systematically larger, and this could not be explained by method choices or publication bias (based on tests we didn’t include in the final version of the paper because they were not relevant for JOLE readership though the finding could be relevant for MAER-Net). Takeaway: It’s worth coding which results are preferred in primary studies, as these author judgments may capture information that’s otherwise hard to quantify.
</p>

<p>
Comments, questions, or counterexamples are welcome!
</p>

<p>
<i>Links to papers and methods are at </i><a href="https://meta-analysis.cz/"><i>meta-analysis.cz</i></a><i>.</i>
</p>]]></content:encoded>
    </item>
    <item>
      <title>Methods Guidelines for Meta-Analysis</title>
      <link>https://meta-analysis.cz/notes/maer-methods-guidelines/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-methods-guidelines/</guid>
      <pubDate>Mon, 27 Nov 2023 00:00:00 +0000</pubDate>
      <description>Zuzana Irsova highlights seven recommendations from the new Journal of Economic Surveys methods guidelines for meta-analysis, covering topic choice, comparability, study quality, correction techniques, Bayesian model averaging, and implied best-practice estimates.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/methods-guidelines-for-meta-analysis" rel="external">MAER-Net</a>, 27 November 2023. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-methods-guidelines/">/komentare/maer-methods-guidelines/</a>.</p>

<p>
As discussed at the Palma colloquium, together with Tom, Chris, and Tomas we prepared a brief practical guide on how to do modern meta-analysis – especially in social sciences but hopefully useful to aspiring meta-analysts in any field. The paper has just been published by the <i>Journal of Economic Surveys</i> and is available <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/joes.12595">here</a>. The purpose of this column is to provide context to the guidelines from my own personal perspective.
</p>

<p>
Following these guidelines is not obligatory for publication in <i>JoES</i> or elsewhere. Nevertheless, we believe there is value in summarizing in one place the comments that we often provide, as editors or referees, on meta-analysis manuscripts in economics, management, psychology, education, environmental sciences, and medical research. We try to be as specific as possible, but the tone of the guidelines is tempered by referees' comments and interactions among the four co-authors. The paper is mostly intended for researchers new to meta-analysis. In this role, the methods guidelines complement the <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/joes.12363">reporting guidelines</a> published in 2020.
</p>

<p>
We provide a step-by-step guide on how to do a meta-analysis from scratch. We cover the following issues: topic selection, literature search, data collection, correction for publication bias and p-hacking, heterogeneity, implied (best practice) estimate, and use of artificial intelligence. The guidelines draw, first, on the vast experience of Tom and Chris with both methods and applications, and second, on the data and codes for dozens of meta-analyses available at <a href="https://meta-analysis.cz/">meta-analysis.cz</a>. Each discussed step and method characteristic is accompanied by a practical example.
</p>

<p>
The guidelines are brief and non-technical, so please read them if you are interested in the content. Here I highlight seven issues that I personally find particularly important:
</p>

<p>
<b>Topic</b>. If possible, choose a topic that you or your co-authors understand well from your own primary research. It's often risky to write a meta-analysis on a topic you are not intimate with, even if you have a lot of previous experience with meta-analysis. Seek co-authors who have written primary studies in this field.
</p>

<p>
<b>Comparability</b>. Make sure it makes sense to quantitatively compare the estimates reported in different studies. While some systematic heterogeneity is inevitable and can be explored in meta-analysis (indeed, it can be the main reason why to do the meta-analysis in the first place), one should not mix estimates that cover scientifically different concepts. If you are unsure, divide the dataset and conduct separate meta-analyses.
</p>

<p>
<b>Quality</b>. Do not exclude any primary studies ex ante. If you can include a paper, include it. You can always show what happens when you give more weight to "good" studies or if you remove the "bad" studies entirely. Importantly, doing so forces you to specify carefully what it is exactly that makes bad studies bad. Any such differentiation of "good" and "bad" need to be determined by an explicitly coded variable (e.g., quasi-experimental or observational) or some objective measure (e.g., retrospective power). If "bad" studies yield results similar to those of "good" studies, that's also a useful finding as it makes your main findings more robust.
</p>

<p>
<b>Multiple estimates per study</b>. Collect all available estimates and use clustering/bootstrapping. Beware <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/jrsm.1441">sample overlap</a>. You can base a robustness check on the estimates that researchers prefer. You can also introduce a robustness check that focuses on one (random or median) estimate per study.
</p>

<p>
<b>Correction techniques</b>. Always correct for potential publication bias or p-hacking. Use at least one model based on the funnel plot (such as PET-PEESE) and one selection model (such as 3PSM). Both families of correction techniques have quite different assumptions. Recently, Bartoš et al. (2023) have offered a model average across both families of publication selection bias corrections and models without any selection bias. Bartoš et al. offer easy to use software, JASP: (<a href="https://fbartos.github.io/RoBMA/">https://fbartos.github.io/RoBMA/</a>) to calculate RoBMA-PSMA. Also, there are a number of video and <a href="https://journals.sagepub.com/doi/full/10.1177/25152459221109259">published tutorials</a> that provide step-by-step guidance. Additionally, you can also use our new <a href="https://meta-analysis.cz/maive/">MAIVE technique</a>, which allows for spurious precision.
</p>

<p>
<b>Bayesian approaches</b>. You don't have to be a convinced Bayesian to recognize the advantages of Bayesian model averaging techniques, both when it comes to <a href="https://fbartos.github.io/RoBMA/">correcting for publication bias</a> (as described above) and explaining <a href="https://www.maer-net.org/post/model-averaging">heterogeneity</a>. When using Bayesian model averaging to explain heterogeneity, it is useful to add a dilution prior to address collinearity.
</p>

<p>
<b>Implied estimates</b>. Report implied (best-practice) meta-analysis means for different contexts (for example, different demographic characteristics) conditional on correction for publication bias and potential misspecifications in primary studies.
</p>

<p>
We hope you will find the guidelines useful. Sometimes they might sound too demanding: do not despair if you cannot accommodate all of these issues. You can find good reasons to disagree with us for specific applications. A well-executed meta-analysis, even if it does not follow all of our recommendations, is immensely helpful to the scientific community and likely to attract a healthy number of citations. Doing a meta-analysis is worth the effort!
</p>

<p>
The guidelines provide many specific examples, but if I were to choose one general example from the meta-analyses that I co-authored, it would be the <a href="https://direct.mit.edu/rest/article/doi/10.1162/rest_a_01227/112420/Publication-and-Attenuation-Biases-in-Measuring">one on skill substitution</a> forthcoming in the <i>Review of Economics and Statistics</i> and available open-access. The paper incorporates almost all the principles discussed here and in the guidelines.
</p>

<p>
To be clear, we do not claim that our guidelines must always be followed nor are they the definitive steps in conducting a rigorous meta-analysis in economics or the social sciences. It is impossible to provide a definitive guide on methodology: methods change rapidly, and opinions on best practices sometimes differ even among the co-authors of our guidelines paper. It's OK, and on some issues expected, if you disagree with us.
</p>]]></content:encoded>
    </item>
    <item>
      <title>Spurious Precision in Meta-Analysis</title>
      <link>https://meta-analysis.cz/notes/maer-spurious-precision/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-spurious-precision/</guid>
      <pubDate>Thu, 16 Feb 2023 00:00:00 +0000</pubDate>
      <description>Meta-analysis upweights studies that report lower standard errors, but reported precision can be p-hacked rather than given. The authors show existing corrections then fail and introduce MAIVE, an instrumental-variable estimator that corrects for this spurious precision.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/maive" rel="external">MAER-Net</a>, 16 February 2023. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-spurious-precision/">/komentare/maer-spurious-precision/</a>.</p>

<p>
<i>Meta-analysis upweights studies reporting lower standard errors and hence more precision. But in empirical practice, notably in observational research, precision is not given to the researcher. Precision must be estimated, and thus can be p-hacked to achieve statistical significance. Simulations show that a modest dose of spurious precision creates a formidable problem for inverse-variance weighting and bias-correction methods based on the funnel plot. Selection models fail to solve the problem, and the simple mean can dominate sophisticated estimators. Cures to publication bias may become worse than the disease. We introduce an approach that surmounts spuriousness: the Meta-Analysis Instrumental Variable Estimator (MAIVE). </i>
</p>

<p>
<b>The paper is available at </b><a href="https://meta-analysis.cz/maive"><b>meta-analysis.cz/maive</b></a><b>. We provide a package for R (</b><i><b>maive</b></i><b>), which makes it easy to use the new method.</b>
</p>

<h2>What is already known</h2>

<ul>
<li>In meta-analysis it's optimal to give more weight to more precise studies.</li>
<li>Inverse-variance weighting maximizes efficiency and may attenuate publication bias.</li>
<li>Inverse-variance weighting is used by all common estimators.</li>
</ul>

<h2>What is new</h2>

<ul>
<li>If reported precision exaggerates real one, inverse-variance weighting creates a bias.</li>
<li>Bias in current methods due to spurious precision can exceed publication bias.</li>
<li>Spurious precision arises naturally in observational research via <i>p</i>-hacking.</li>
<li>Meta-Analysis Instrumental Variable Estimator (MAIVE) corrects for spuriousness.</li>
</ul>

<h2>Potential impact</h2>

<ul>
<li>Meta-analysts should use MAIVE if they suspect <i>p</i>-hacking.</li>
<li>The difference between MAIVE and unadjusted estimators can measure spuriousness.</li>
<li>MAIVE substantially improves the robustness of the current meta-analysis toolkit.</li>
</ul>

<p>
Inverse-variance weighting reigns in meta-analysis. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[1</a>] More precise studies, or rather those seemingly more precise based on lower reported standard errors, get a greater weight explicitly or implicitly. The weight is explicit in traditional summaries, such as the fixed-effect model (assuming a common effect) and the random-effects model (allowing for heterogeneity). <a href="https://www.maer-net.org/post/maive#viewer-2137l">[2,3</a>] These models work as weighted averages, the weight diluted in random effects by a heterogeneity term. The weight is also explicit in publication bias correction models based on the funnel plot. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[4–12</a>] <b>In funnel-based models, reported precision is particularly important</b> because the weighted average gets reinforced by assigning more importance to supposedly less biased (nominally more precise) studies. The weight is implicit in selection models estimated using the maximum likelihood approach, <a href="https://www.maer-net.org/post/maive#viewer-2137l">[13–18</a>] which often reduce to the random-effects model in the absence of publication bias.
</p>

<h2>Reported precision turned spurious</h2>

<p>
The tacit assumption behind all these techniques is that the reported, nominal precision represents the true, underlying precision. The standard error, inverse of precision, is given to the researcher by her data and methods. It's fixed and can't be manipulated, consciously or unconsciously. The assumption is plausible in experimental research, for which most meta-analysis methods were developed. But in observational research, where thousands of meta-analyses are produced each year, <b>the derivation of the standard error is often a key part of the empirical exercise</b>. Consider a regression analysis with longitudinal data: explaining the health outcomes of patients treated by different physicians and observed over several years. Individual observations aren't independent, and standard errors need to be clustered. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[19</a>] But how? At the level of physicians, patients, or years? Should one use double clustering <a href="https://www.maer-net.org/post/maive#viewer-2137l">[20</a>] or perhaps wild bootstrap <a href="https://www.maer-net.org/post/maive#viewer-2137l">[21</a>]? It's complicated, and with a different computation of confidence intervals the researcher will report different precision for the same estimated effect size.
</p>

<p>
Spurious precision can arise in many contexts other than longitudinal data analysis. Ordinary least squares, the workhorse of observational research, assume homoskedasticity of residuals. The assumption is often violated, and in these cases researchers should use heteroskedasticity-robust standard errors, <a href="https://www.maer-net.org/post/maive#viewer-2137l">[22</a>] typically larger than plain vanilla standard errors. <b>If researchers ignore heteroskedasticity, they report precise estimates, but the precision is spurious</b>. Similar problems may arise due to nonstationarity in time series <a href="https://www.maer-net.org/post/maive#viewer-2137l">[23</a>] and a myriad of other issues. When a study with exaggerated precision enters meta-analysis, it gets too much impact because of inverse-variance weighting. If a meta-analyst spots the methodological problem, she can exclude the study or add a corresponding control. Either way, the weighting problem isn't properly addressed. And spotting misspecification is hard, because the tweak lifting reported precision can be hidden within a complex model.
</p>

<h2>Various sources of spuriousness</h2>

<p>
Spurious precision can also arise due to cheating. For economics journals, quasi-experimental evidence shows that the introduction of obligatory data sharing substantially reduced the reported <i>t</i>-statistics. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[24</a>] Prior to the introduction of data sharing some authors had probably cheated by manipulating data or results. Pütz and Bruns find hundreds of reporting errors in top economics journals; when they ask authors to explain the errors, the authors are four times more likely to admit a mistake in the standard error than in the estimated effect size. [67] But cheating, mistakes, and other issues that can affect the standard error independently of the estimated effect size aren't necessary to produce spurious precision. <b>A realistic mechanism is </b><i><b>p</b></i><b>-hacking, in which the researcher adjusts the entire model to produce statistically significant results.</b> After adjusting the model, both the effect size and standard error change, and both can jointly contribute to statistical significance. We examine, by employing Monte Carlo simulations, the consequences of cheating and the more realistic <i>p</i>-hacking behavior, of which spurious precision is a natural result.
</p>

<p>
Figure 1 gives intuition on the cheating/clustering/heteroskedasticity/nonstationarity simulation. For brevity we call it a cheating scenario. Researchers crave statistically significant estimates and to that effect manipulate effect sizes or standard errors at will, but not both at the same time. The scenario is simplistic, and we start with it because it allows for a clean separation of selection on estimates (conventional in the literature) and selection on standard errors (our focus). The separation isn't so clean in the <i>p</i>-hacking scenario but can be mapped back to the cheating scenario. The mechanism of the left-hand panel of Figure 1 is analogous to the Lombard effect in psychoacoustic: <a href="https://www.maer-net.org/post/maive#viewer-2137l">[25,26</a>] speakers increase their vocal effort in response to noise. Here <b>researchers increase their selection effort in response to noise in data or methods</b>, noise that produces imprecision and insignificance. When researchers so cheat with effect sizes, the results are consistent with funnel-based models of publication bias: funnel asymmetry arises, the most precise estimates remain close to the true effect, and inverse-variance weighting helps mitigate the bias—aside from improving the efficiency of the aggregate estimate, the original rationale for using the weights. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[27,28</a>]
</p>

<h2>Bias in inverse-variance weighting</h2>

<p>
The right-hand panel of Figure 1 paints a different picture. Here the mechanism is analogous to Taylor’s law in ecology: <a href="https://www.maer-net.org/post/maive#viewer-2137l">[29</a>] the variance can decrease with a smaller mean (originally describing population density for various species). <b>When researchers achieve significance by lowering the standard error, we again observe funnel asymmetry. But this time no bias arises </b>in the reported effect sizes: the black-filled circles and the hollow circles denote the same effect size, only precision changes. The simple unweighted mean of reported estimates is unbiased, and inverse-variance weighting paradoxically creates a downward bias. The bias increases when we use a correction based on the funnel plot: effectively, when we estimate the size of a hypothetical infinitely precise study, the intercept of a regression curve.
</p>

<p>
In practice, as noted, selection on estimates and standard errors arises simultaneously. We generate this quality in simulations by allowing researchers to replace control variables in a regression context, a mechanism that also gives rise to sizable heterogeneity. Control variables are correlated with the main regressor of interest (for example, a treatment variable), and their <b>replacement affects both the estimated treatment effect and the corresponding precision</b>. Then <i>p</i>-hacked estimates move not strictly north or west, as in the figure, but northwest. Even spuriously large estimates can now be spuriously precise. The resulting bias direction due to inverse-variance weighting is unclear. Our simulations suggest that an upwards bias is plausible.
</p>

<h2>Current methods fail with spurious precision</h2>

<p>
Does any technique yield little bias and good coverage rates in the case of panel B of Figure 1, or at least with a small ratio of selection on standard errors relative to selection on estimates? <b>We examine 7 current estimators</b>: simple unweighted mean, fixed effects (weighted least squares, FE/WLS), <a href="https://www.maer-net.org/post/maive#viewer-2137l">[30</a>] precision-effect test and precision-effect estimate with standard errors (PET-PEESE), <a href="https://www.maer-net.org/post/maive#viewer-2137l">[9</a>] endogenous kink (EK), <a href="https://www.maer-net.org/post/maive#viewer-2137l">[11</a>] weighted average of adequately powered estimates (WAAP), <a href="https://www.maer-net.org/post/maive#viewer-2137l">[10</a>] the selection model by Andrews and Kasy, <a href="https://www.maer-net.org/post/maive#viewer-2137l">[17</a>] and <i>p</i>-uniform∗ <a href="https://www.maer-net.org/post/maive#viewer-2137l">[18</a>]. The first two are basic summary statistics, the next three are correction methods based on the funnel plot, and the last two are selection models. The choice of estimators is subjective, but the three funnel-based techniques are commonly used in observational research. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[31–38</a>] The two selection models are also used often <a href="https://www.maer-net.org/post/maive#viewer-2137l">[39–46</a>] and represent the latest incarnations of models in the tradition of Hedges <a href="https://www.maer-net.org/post/maive#viewer-2137l">[13–16</a>] and their simplifications <a href="https://www.maer-net.org/post/maive#viewer-2137l">[47–52</a>].
</p>

<p>
The importance of reported precision for these estimators is summarized in Table 1. In most of them precision has two roles: <b>weight and identification</b>. Identification can be achieved through meta-regression (where the standard error or a function thereof is included as a regressor), selection model, or a combination of both—such as the EK model.
</p>

<p>
<b>None of these 7 estimators work well with even a sprinkle of spurious precision</b>. The simple unweighted mean plagued by publication bias can be the best, but still no good. The reader might expect selection models to beat funnel-based models, because of the latter’s heavier reliance on precision. Alas, this is generally not the case, and even selection models are often defeated by the simple mean when selection on standard errors is modest (about 1:5 and more compared to selection on estimates). We propose a straightforward adjustment of funnel-based techniques, the meta-analysis instrumental variable estimator (MAIVE), which corrects most of the bias and restores valid coverage rates. <b>MAIVE replaces, in all meta-analysis contexts, reported variance with the portion of reported variance that can be explained by the inverse sample size</b> used in the primary study. We justify the idea by starting with a version of the Egger regression: <a href="https://www.maer-net.org/post/maive#viewer-2137l">[4</a>]
</p>

<p>
where <i>αi_hat</i> on the left-hand side denotes effects estimated in primary studies and <i>SE </i>their standard errors. This is the PEESE model due to Stanley and Doucouliagos, but for simplicity without additional inverse-variance weights—since the model searches for the effect conditional on maximum precision, it already features an implicit, built-in weight. In panel A of Figure 1, the quadratic regression would fit the data quite well, <a href="https://www.maer-net.org/post/maive#viewer-2137l">[9</a>] and estimated <i>α</i>0 would lie close to the mean underlying effect. <b>In panel B, however, the regression fails to recover the underlying coefficients.</b> The regression fails because it assumes a causal effect of the standard error on the estimate: a good description of panel A (Lombard effect), but not panel B (Taylor’s law). In panel B, the standard error sometimes depends on the estimated effect size and is thus correlated with the error term, <i>vi</i>. The resulting estimates of <i>α</i>0 (true effect) and <i>β </i>(intensity of selection) are biased.
</p>

<h2>Endogeneity problem in meta-regression</h2>

<p>
The problem is the correlation between <i>SE </i>and <i>vi</i>, which can arise for three reasons: First, selective reporting based on standard errors, which we simulate. Second, measurement error in <i>SE</i>. This issue was mentioned in 2005 by Tom Stanley, <a href="https://www.maer-net.org/post/maive#viewer-2137l">[53</a>] <b>who was the first to instrument the standard error in a meta-analysis context</b>. Nevertheless, Stanley didn't discuss the adjustment of weights nor did he pursue the idea further as a bias-correction estimator. We don't consider this source of correlation in simulations. Third, the correlation can be caused by unobserved heterogeneity: some method choices affect both estimates and standard errors, and some standardized meta-analysis effects feature a mechanical correlation between both quantities. [36] (A careless meta-analyst may also mix estimates measured in different units. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[54</a>]) Our <i>p</i>-hacking simulation only partly addresses this mechanism by allowing researchers to change control variables, which can affect both estimates and standard errors at the same time—a combination of panel A and panel B of Figure 1. In other words, we model only some of the mechanisms which give rise to spurious precision.
</p>

<p>
The statistical solution to the problem, often called endogeneity, is to find an instrument for the standard error. A valid instrument is correlated with the standard error, but not with the error term (and thus unrelated to the three sources of endogeneity mentioned above). While finding good instruments is often challenging, here the answer beckons. By definition, reported variance (<i>SE</i>2) is a linear function of the inverse of the sample size used in the primary study. <b>The sample size is plausibly robust to selection,</b> or at least it's more difficult to collect more data than to <i>p</i>-hack the standard error to achieve significance. The sample size isn't estimated, and so it doesn't suffer from measurement error. The sample size is typically not affected by changing methodology, certainly not by changing control variables. Some endogeneity may remain if researchers correctly expecting smaller effects design larger experiments. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[41</a>] But, at least in observational research, authors often use as much data as available from the start. Indeed, the sample size, unlike the standard error, is often given to the researcher: the very word <i>data </i>means things given.
</p>

<h2>Meta-Analysis Instrumental Variable Estimator (MAIVE)</h2>

<p>
<b>We regress the squared reported standard errors on the inverse sample size and plug the fitted values instead of the variance</b> to the right-hand side of the aforementioned equation. Thence we obtain the baseline MAIVE estimator. For the baseline MAIVE we choose the instrumented version of PEESE without additional inverse-variance weights because it works well in simulations. The version with additional adjusted weights (again, using fitted values instead of reported precision) often performs similarly but is more complex, so we prefer the former, parsimonious solution. In principle, any funnel-based technique (and the funnel plot itself) can be adjusted by the procedure described above: just replace the standard error with the square root of the fitted values. The adjustment helps the fixed-effect, WAAP, and endogenous kink model to typically defeat both the simple unweighted mean and selection models in the presence of spurious precision. MAIVE can be easily applied using our <i>maive </i>package in R.
</p>

<p>
Table 2 shows the variants of individual estimators we consider in simulations. We always start with the unadjusted, plain-vanilla variant. Where easily possible, we consider the adjustment of weights and identification devices separately. So, for PET-PEESE and EK we have 5 different flavors. Note that the <b>separation is not straightforward for selection models</b>, and we do not pursue it here. This is one of many low-hanging fruits (thesis topics?) that grow from the spurious precision project and await reaping in future research; we discuss more at the end of this column.
</p>

<p>
In Figure 2 below we report one set of simulation results: the case of the <i>p</i>-hacking scenario with a positive underlying effect size. In this scenario the authors of primary studies run regressions with two variables on the right-hand side and are interested in the slope coefficient on the first variable. Both variables belong to the correctly specified regression model. A meta-analyst collects the slope coefficients estimated for the first variable (e.g., treatment); no one is interested in the second variable (control). The vertical axis in Figure 2 measures the bias of meta-analysis estimators relative to the true value of the slope coefficient, the true treatment effect (<i>α1</i> = 1). The horizontal axis measures the correlation between the regression variable of interest and a control variable that should be included—but can be replaced by some researchers with another, less relevant control, <b>a practice that affects both reported estimates and their standard errors</b>.
</p>

<p>
The higher the correlation, the more potential for <i>p</i>-hacking via the replacement of the control variable. If there is no correlation, removing or replacing the control won't systematically affect the main estimated parameter. With a positive correlation and a positive underlying value of the second slope coefficient, replacing the control variable with a less relevant proxy creates an upwards omitted-variable bias. Importantly for our purposes, a <b>higher correlation increases selection on standard errors more than proportionally compared to selection on estimates.</b> With a higher correlation and thus more <i>p</i>-hacking and also more relative selection on standard errors, the bias of standard meta-analysis estimators increases. Note that even a large correlation still corresponds to a relatively small ratio of selection on standard errors (spurious precision, Taylor's law) relative to selection on estimates (Lombard effect). In <a href="https://meta-analysis.cz/maive/maive.pdf">the paper</a> we compute and tabulate this correspondence for different values of the true effect.
</p>

<h2>Simple mean beats complex models</h2>

<p>
Eventually, the bias of the classical, unadjusted techniques gets even larger than the bias of the simple unweighted mean. That is, when there is enough spuriousness, corrections for publication bias do more harm than good. MAIVE corrects most of the spuriousness bias (see panels B-G in Figure 2 and compare panel A, classical estimators, to panel H, MAIVE versions of these estimators), and the MAIVE versions with adjusted or omitted weights work similarly well. MAIVE performs comparably to conventional estimators if spurious precision is negligible, and dominates unadjusted estimators if spuriousness is non-negligible (as we show in the paper, about 1:10 of selection on standard errors to selection on estimates). In <a href="https://meta-analysis.cz/maive/maive.pdf">the paper</a> we report the results of many more simulation scenarios, both cheating and <i>p</i>-hacking, for bias, MSE, and coverage rates—with comparable results in qualitative terms. <b>Even a modest dose of spurious precision makes inverse-variance weighting (explicit or implicit) unreliable and warrants a MAIVE treatment.</b>
</p>

<p>
Why, instead of instrumenting, don't we simply<b> replace variance with inverse sample size</b>? <a href="https://www.maer-net.org/post/maive#viewer-2137l">[36,55–57</a>] While the replacement would also address spurious precision, the instrumental approach has many advantages, as discussed in Section 3 of <a href="https://meta-analysis.cz/maive/maive.pdf">the paper.</a> One advantage is flexibility: the instrumental approach can incorporate other aspects of study design, besides sample size, that affect standard errors. Sample size rarely forms a perfect proxy for precision, and MAIVE can be extended by adding instruments to improve the fit. Moreover, the instrumental approach remains statistically valid even if, for some reason, the correlation between the reported variance and the inverse sample size is small.
</p>

<h2>Current methods adjusted to spuriousness</h2>

<p>
We don't argue that spurious precision is common. We argue that it can plausibly arise in observational research. Even in experimental settings, randomization can fail, <a href="https://www.maer-net.org/post/maive#viewer-2137l">[58</a>] and authors often use regressions to control for pre-treatment covariates or make other adjustments <a href="https://www.maer-net.org/post/maive#viewer-2137l">[59</a>] that can yield spurious precision. When it arises, a small dose can render the simple mean more reliable than sophisticated correction techniques. The Meta-Analysis Instrumental Variable Estimator (MAIVE) solves the problem by using inverse sample size as an instrument for reported variance. That is, we regress the reported squared standard errors on the inverse of the number of observations used in the primary study. The fitted values from this regression are then used instead of reported variance in the PEESE meta-regression. Standard weighted means, funnel plots, and funnel-based methods can be adjusted similarly to make them robust to spurious precision. <b>The entire meta-analysis toolkit can be salvaged with this modification.</b>
</p>

<p>
The instrumental approach has seven benefits over using sample size as a proxy for precision, as noted, and we explain them in <a href="https://meta-analysis.cz/maive/maive.pdf">the paper</a>. There are at least two costs as well, both compared to the proxy approach and the classical one that relies on reported precision. First, MAIVE is more complex since it involves an additional regression and computation of fitted values and valid confidence intervals. But the instrumental approach is readily available in most statistical programs. <b>We create the </b><i><b>maive </b></i><b>package for R, which makes estimation easy for meta-analysts unfamiliar with instrumental variables.</b> Second, the additional regression makes MAIVE noisier compared to conventional techniques. When a meta-analyst is sure there can be absolutely no spurious precision in her data, using reported precision without instruments will yield unbiased and more efficient estimates. The lack of spurious precision can be tested approximately by employing the Hausman specification test: <a href="https://www.maer-net.org/post/maive#viewer-2137l">[60</a>] if the coefficients estimated in MAIVE are far from those of an unadjusted PEESE, spurious precision is likely an issue.
</p>

<h2>Using MAIVE in practice</h2>

<p>
A discussion is in order regarding the application of MAIVE—pronounced, by the way, as the Irish name Maeve. The instrument is the overall sample size, not degrees of freedom, because the latter depends on clustering units. <b>We prefer the MAIVE version of PEESE without weights</b> (after testing with unweighted MAIVE-PET whether the true effect is nonzero). This parsimonious specification intuitively fits both panels of Figure 1. The <i>maive </i>package allows for optional adjusted weights. Researchers may choose a MAIVE version of another estimator, such as endogenous kink. The package also runs the Hausman test, a rough indicator of spuriousness. Because PEESE is heteroskedastic by definition and we prefer not to use inverse-variance weights, the package produces heteroskedasticity-robust standard errors by default. When some studies report multiple estimates, standard errors in MAIVE—and any meta-analysis estimator—<b>should be clustered at the study level,</b> again a default option. With fewer than 30 studies we recommend wild bootstrap. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[21</a>] It's a good idea to include study-level dummies (econometric fixed effects) to filter out study-specific idiosyncrasies related to unobserved heterogeneity. The package also reports a robust F-statistic of the first-stage regression. If the F-statistic is below 10, the instrument is weak and MAIVE results should be treated with caution. Researchers may want to use confidence intervals robust to weak instruments.
</p>

<p>
The reader will object that our simulation is unfair to correction methods. <b>The methods were designed to counter publication bias; we simulate </b><i><b>p</b></i><b>-hacking</b>. Individual estimates and standard errors get biased, which is why selection models don't work well here—though they don't assume, as funnel methods assume, that selection works only on estimates (the Lombard effect discussed earlier). The distinction between publication bias and <i>p</i>-hacking is clear in theory, but in practice both are often observationally equivalent to the meta-analyst. (But <i>p</i>-hacking likely predominates. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[61</a>]) As long as we believe our <i>p</i>-hacking environment is broadly realistic, we need a technique that corrects the resulting bias. MAIVE is the only such technique. One can design <i>p</i>-hacking scenarios in which misspecifications make it almost impossible for meta-analysis methods to uncover the true mean. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[58,62</a>] If that is a realistic description of observational research, unconditional meta-analysis means are meaningless. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[63</a>] <b>MAIVE can be extended to allow for observed heterogeneity</b> and deliver context-specific means via incorporation into <a href="https://www.maer-net.org/post/model-averaging">Bayesian model averaging meta-regression</a> approaches addressing model uncertainty. <a href="https://www.maer-net.org/post/maive#viewer-2137l">[42–46</a>]
</p>

<h2>Low-hanging fruit for future research</h2>

<p>
We leave many questions open regarding spurious precision. How common is it in practice? How does measurement error influence the relative performance of MAIVE? What happens when method heterogeneity explicitly affects both estimates and their precision? Does spurious precision help explain why meta-analyses often exaggerate the true effect compared to multi-lab pre-registered replications? <a href="https://www.maer-net.org/post/maive#viewer-2137l">[64–66</a>] <b>How to correctly adjust selection models for spuriousness?</b> The last is perhaps the most important question for future research because many meta-analysts prefer selection models over funnel-based techniques. The adjustment of selection models isn't straightforward since here precision has two intertwined roles: identification and weighting. For identification, we need the reported, nominal precision, which determines statistical significance. But for weights we need the underlying, true precision. The maximum likelihood approach has to be modified to allow a different measure of precision for each role.
</p>

<h2>Bottom line</h2>

<p>
Spurious precision, while plausibly destructive, is surmounted by adjusting funnel-based methods.
</p>

<h2>References</h2>

<ul>
<li>Gurevitch J, Koricheva J, Nakagawa S, Stewart G. Meta-analysis and the science of research synthesis. <i>Nature </i>2018; 555: 175–182.</li>
<li>Borenstein M, Hedges L, Higgins J, Rothstein H. A basic introduction to fixed-effect and random-effects models for meta-analysis. <i>Research Synthesis Methods </i>2010; 1(2): 97–111.</li>
<li>Stanley TD, Doucouliagos H. Neither fixed nor random: weighted least squares meta-analysis. <i>Statistics in Medicine </i>2015; 34(13): 2116-2127.</li>
<li>Egger M, Smith GD, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. <i>British Medical Journal </i>1997; 315(7109): 629–634.</li>
<li>Duval S, Tweedie R. Trim and fill: A simple funnel-plot–based method of testing and adjusting for publication bias in meta-analysis. <i>Biometrics </i>2000; 56(2): 455–463.</li>
<li>Stanley TD. Meta-Regression Methods for Detecting and Estimating Empirical Effects in the Presence of Publication Selection. <i>Oxford Bulletin of Economics and Statistics </i>2008; 70(1): 103–127.</li>
<li>Stanley TD, Jarrell SB, Doucouliagos H. Could It Be Better to Discard 90% of the Data? A Statistical Paradox. <i>The American Statistician </i>2010; 64(1): 70–77.</li>
<li>Stanley TD, Doucouliagos H. <i>Meta-regression analysis in economics and business</i>. New York: Routledge. 2012.</li>
<li>Stanley TD, Doucouliagos H. Meta-regression approximations to reduce publication selection bias. <i>Research Synthesis Methods </i>2014; 5(1): 60–78.</li>
<li>Ioannidis JP, Stanley TD, Doucouliagos H. The Power of Bias in Economics Research. <i>The Economic Journal </i>2017; 127(605): F236–F265.</li>
<li>Bom PRD, Rachinger H. A kinked meta-regression model for publication bias correction. <i>Research Synthesis Methods </i>2019; 10(4): 497–514.</li>
<li>Furukawa C. Publication Bias under Aggregation Frictions: Theory, Evidence, and a New Correction Method. <i>MIT </i>2019; working paper.</li>
<li>Hedges L. Estimation of effect size under nonrandom sampling: The effect of censoring studies yielding statistically insignificant mean differences. <i>Journal of Educational Statistics </i>1984; 9: 61–85.</li>
<li>Iyengar S, Greenhouse JB. Selection Models and the File Drawer Problem. <i>Statistical Science </i>1988; 3(1): 109–117.</li>
<li>Hedges LV. Modeling Publication Selection Effects in Meta-Analysis. <i>Statistical Science</i>1992; 72(2): 246–255.</li>
<li>Vevea J, Hedges LV. A general linear model for estimating effect size in the presence of publication bias. <i>Psychometrika </i>1995; 60(3): 419–435.</li>
<li>Andrews I, Kasy M. Identification of and correction for publication bias. <i>American Economic Review </i>2019; 109(8): 2766–2794.</li>
<li>Aert vRC, Assen vM. Correcting for publication bias in a meta-analysis with the p-uniform<i> method. </i>Tilburg University &amp; Utrecht University *2021; working paper.</li>
<li>Abadie A, Athey S, Imbens GW, Wooldridge JM. When Should You Adjust Standard Errors for Clustering? <i>The Quarterly Journal of Economics </i>2022; 138(1): 1–35.</li>
<li>Cameron AC, Miller DL. A practitioner’s guide to cluster-robust inference. <i>Journal of Human Resources </i>2015; 50(2): 317–372.</li>
<li>Roodman D, Nielsen MØ, MacKinnon JG, Webb MD. Fast and Wild: Bootstrap Inference in Stata Using Boottest. <i>The Stata Journal </i>2019; 19(1): 4–60.</li>
<li>White H. A Heteroskedasticity-Consistent Covariance Matrix Estimator and a Direct Test for Heteroskedasticity. <i>Econometrica </i>1980; 48(4): 817–838.</li>
<li>Bom PRD, Ligthart JE. What Have We Learned from Three Decades of Research on the Productivity of Public Capital? <i>Journal of Economic Surveys </i>2014; 28(5): 889–916.</li>
<li>Askarov Z, Doucouliagos A, Doucouliagos H, Stanley TD. The Significance of Data-Sharing Policy. <i>Journal of the European Economic Association </i>2023; forthcoming.</li>
<li>Lane H, Tranel B. The Lombard Sign and the Role of Hearing in Speech. <i>Journal of Speech and Hearing Research </i>1971; 14(4): 677–709.</li>
<li>McCloskey DN, Ziliak ST. What quantitative methods should we teach to graduate students? A comment on Swann’s 'Is precise econometrics an illusion'? <i>The Journal of Economic Education </i>2019; 50(4): 356–361.</li>
<li>Hedges LV. A random effects model for effect sizes. <i>Psychological Bulletin </i>1983; 93(2): 388–395.</li>
<li>Hedges LV, Olkin I. <i>Statistical methods for meta-analysis</i>. Orlando, FL: Academic Press. 1985.</li>
<li>Taylor LR. Aggregation, variance and the mean. <i>Nature </i>1961;189(4766): 732–735.</li>
<li>Stanley TD, Doucouliagos H. Neither fixed nor random: weighted least squares meta-regression. <i>Research Synthesis Methods </i>2017; 8(1): 19–42.</li>
<li>Havranek T, Stanley TD, Doucouliagos H, et al. Reporting Guidelines for Meta-Analysis in Economics. <i>Journal of Economic Surveys </i>2020; 34(3): 469–475.</li>
<li>Ugur M, Awaworyi Churchill S, Luong H. What do we know about R&amp;D spillovers and productivity? Meta-analysis evidence on heterogeneity and statistical power. <i>Research Policy </i>2020; 49(1): 103866.</li>
<li>Xue X, Reed WR, Menclova A. Social capital and health: A meta-analysis. <i>Journal of Health Economics </i>2020; 72(C): 102317.</li>
<li>Neisser C. The Elasticity of Taxable Income: A Meta-Regression Analysis. <i>Economic Journal </i>2021; 131(640): 3365–3391.</li>
<li>Zigraiova D, Havranek T, Irsova Z, Novak J. How puzzling is the forward premium puzzle? A meta-analysis. <i>European Economic Review </i>2021; 134(C): 103714.</li>
<li>Nakagawa S, Lagisz M, Jennions MD, et al. Methods for testing publication bias in ecological and evolutionary meta-analyses. <i>Methods in Ecology and Evolution </i>2022; 13(1): 4–21.</li>
<li>Brown AL, Imai T, Vieider F, Camerer C. Meta-Analysis of Empirical Estimates of Loss-Aversion. <i>Journal of Economic Literature </i>2023; forthcoming.</li>
<li>Heimberger P. Do Higher Public Debt Levels Reduce Economic Growth? <i>Journal of Economic Surveys </i>2023; forthcoming.</li>
<li>Carter EC, Schonbrodt FD, Gervais WM, Hilgard J. Correcting for Bias in Psychology: A Comparison of Meta-Analytic Methods. <i>Advances in Methods and Practices in Psychological Science </i>2019; 2(2): 115-144.</li>
<li>Brodeur A, Cook N, Heyes A. Methods Matter: P-Hacking and Causal Inference in Economics. <i>American Economic Review </i>2020; 110(11): 3634–3660.</li>
<li>DellaVigna S, Linos E. RCTs to Scale: Comprehensive Evidence From Two Nudge Units. <i>Econometrica </i>2022; 90(1): 81–116.</li>
<li>Gechert S, Havranek T, Irsova Z, Kolcunova D. Measuring Capital-Labor Substitution: The Importance of Method Choices and Publication Bias. <i>Review of Economic Dynamics </i>2022; 45(C): 55–82.</li>
<li>Imai T, Rutter TA, Camerer CF. Meta-Analysis of Present-Bias Estimation Using Convex Time Budgets. <i>The Economic Journal </i>2021; 131(636): 1788–1814.</li>
<li>Gechert S, Heimberger P. Do corporate tax cuts boost economic growth? <i>European Economic Review </i>2022; 147(C): 104157.</li>
<li>Havranek T, Irsova Z, Laslopova L, Zeynalova O. Publication and Attenuation Biases in Measuring Skill Substitution. <i>The Review of Economics and Statistics </i>2023; forthcoming.</li>
<li>Matousek J, Havranek T, Irsova Z. Individual discount rates: A meta-analysis of experimental evidence. <i>Experimental Economics </i>2022; 25(1): 318–358.</li>
<li>Simonsohn U, Nelson LD, Simmons JP. P-curve: A key to the file-Drawer. <i>Journal of Experimental Psychology: General </i>2014; 143(2): 534–547.</li>
<li>Simonsohn U, Nelson LD, Simmons JP. p-Curve and Effect Size: Correcting for Publication Bias Using Only Significant Results. <i>Perspectives on Psychological Science </i>2014; 9(6): 666-681. PMID: 26186117.</li>
<li>Assen vM, Aert vRC, Wicherts JM. Meta-analysis using effect size distributions of only statistically significant studies. <i>Psychological Methods </i>2015; 20(3): 293–309.</li>
<li>Simonsohn U, Simmons JP, Nelson LD. Better p-curves: Making p-curve analysis more robust to errors, fraud, and ambitious p-hacking, a Reply to Ulrich and Miller (2015). <i>Journal of Experimental Psychology: General </i>2015; 144(6): 1146–1152.</li>
<li>Aert vRC, Assen vM. Bayesian evaluation of effect size after replicating an original study. <i>Plos ONE </i>2017; 12(4): e0175302.</li>
<li>Aert vRC, Assen vM. Examining reproducibility in psychology: A hybrid method for combining a statistically significant original study and a replication. <i>Behavior Research Methods </i>2018; 50: 1515–1539.</li>
<li>Stanley TD. Beyond Publication Bias. <i>Journal of Economic Surveys </i>2005; 19(3):309–345.</li>
<li>Kranz S, Putz P. Methods Matter: p-Hacking and Publication Bias in Causal Analysis in Economics: Comment. <i>American Economic Review </i>2022; 112(9): 3124–3136.</li>
<li>Sanchez-Meca J, Marın-Martınez F. Weighting by Inverse Variance or by Sample Size in Meta-Analysis: A Simulation Study. <i>Educational and Psychological Measurement </i>1998; 58(2): 211-220.</li>
<li>Deeks JJ, Macaskill P, Irwig L. The performance of tests of publication bias and other sample size effects in systematic reviews. <i>Journal of Clinical Epidemiology </i>2005; 58(9): 882–893.</li>
<li>Peters JL, Sutton AJ, Jones DR, Abrams KR, Rushton L. Comparison of Two Methods to Detect Publication Bias in Meta-analysis. <i>JAMA </i>2006; 295(6): 676-680.</li>
<li>Bruns SB, Ioannidis JP. p-Curve and p-Hacking in Observational Research. <i>PloS ONE </i>2016; 11(2): e0149144.</li>
<li>Freedman DA. On regression adjustments to experimental data. <i>Advances in Applied Mathematics </i>2008; 40(2): 180–193.</li>
<li>Hausman JA. Specification Tests in Econometrics. <i>Econometrica </i>1978;46(6): 1251–1271.</li>
<li>Brodeur A, Carrell S, Figlio D, Lusher L. Unpacking p-hacking and publication bias. <i>American Economic Review </i>2023; forthcoming.</li>
<li>Bruns SB. Meta-Regression Models and Observational Research. <i>Oxford Bulletin of Economics and Statistics </i>2017; 79(5): 637–653.</li>
<li>Simonsohn U, Simmons J, Nelson LD. Above averaging in literature reviews. <i>Nature Reviews Psychology </i>2022; 1: 551–552.</li>
<li>Kvarven A, Stromland E, Johannesson M. Comparing meta-analyses and preregistered multiple-laboratory replication projects. <i>Nature Human Behavior </i>2020; 4: 423–434.</li>
<li>Lewis M, Mathur MB, VanderWeele TJ, Frank MC. The puzzling relationship between multi-laboratory replications and meta-analyses of the published literature. <i>Royal Society Open Science </i>2022; 9(2): 211499.</li>
<li>Stanley TD, Doucouliagos H, Ioannidis JPA. Retrospective median power, false positive meta-analysis and large-scale replication. <i>Research Synthesis Methods </i>2022; 13(1): 88-108.</li>
<li>Putz P, Bruns SB. The (Non-)Significance of Reporting Errors in Economics: Evidence from Three Top Journals. <i>Journal of Economic Surveys</i> 2021; 35(1): 348-373.</li>
</ul>]]></content:encoded>
    </item>
    <item>
      <title>How financial incentives affect performance</title>
      <link>https://meta-analysis.cz/notes/voxeu-incentives/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/voxeu-incentives/</guid>
      <pubDate>Mon, 13 Feb 2023 00:00:00 +0000</pubDate>
      <description>A meta-analysis of 44 experimental economics studies finds a negligible effect of financial incentives on performance once publication bias and differences in experimental context are corrected, suggesting money-based nudges are less effective than commonly assumed.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://cepr.org/voxeu/columns/how-financial-incentives-affect-performance" rel="external">VoxEU / CEPR</a>, 13 February 2023. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/voxeu-incentives/">/komentare/voxeu-incentives/</a>.</p>

<p>
Economists tend to assume that monetary rewards make people try harder, but many psychologists argue that the opposite can be true. This column takes stock of 44 experimental studies in economics to test the relationship between financial incentives and performance. After correcting publication bias and differences in experimental contexts, the analysis suggests a negligible effect of financial incentives on performance across all contexts of field experiments. Incentives or nudges that rely mainly on financial motives, such as offering money to people for getting their Covid-19 shots, may be less effective than commonly thought.
</p>

<p>
Since at least the 1970s, psychologists have been pointing out that financial incentives can harm performance by crowding out the enjoyment we would otherwise earn while working on a task (Deci 1971). An enjoyable task morphs into one that we do for the money, which crowds out intrinsic motivation. If extrinsic motivation provided by the financial incentive is not strong enough, money rewards result in reduced performance. While not universally accepted, the motivation crowding theory is the default incentive model in psychology and related fields (Engström and Bengtsson 2014, Deserranno 2015, Bowles 2016, Pope and DellaVigna 2016).
</p>

<p>
A widely cited meta-analysis by Weibel et al. (2010) reports that financial incentives indeed hurt performance in the case of interesting tasks. While economists have long been aware of the psychology theory and evidence (Camerer and Hogarth 1999, Gneezy and Rustichini 2000, Frey and Jegen 2001, Gneezy et al. 2011, Esteves-Sorenson and Broce 2022), few models used in economics allow for motivation crowding. The following statement recently expressed prominently on the website of a leading management consultancy reflects the prior of many economists:
</p>

<p>
"Generous and specific financial incentives can help drive and sustain a rapid performance improvement" (McKinsey 2022)
</p>

<p>
In this column, based on a synthesis of empirical studies on the topic (Cala et al. 2022), we argue that empirical evidence, even in economics, does not support the prior. The finding has far-reaching policy consequences: incentives or nudges that rely mainly on financial motives, such as offering money to people for getting their Covid-19 shots, may be less effective than commonly thought.
</p>

<p>
The purpose of our meta-analysis is threefold (Cala et al. 2022). First, we correct the literature for publication bias, which can exaggerate the underlying effect multiplicatively (Ioannidis et al. 2017). Second, we allow for model uncertainty (Steel 2020), which is important given how individual experiments differ. Third, we focus on economics. Existing meta-analyses have focused exclusively or to a large extent on psychology. The economics literature is thus largely unexplored, although researchers have pointed out the vast differences in priors and methodological approaches between economics and psychology experiments when it comes to the effect of money on behaviour (Camerer and Hogarth 1999, Hertwig and Ortmann 2001, Esteves-Sorenson and Broce 2022).
</p>

<p>
Figure 1 presents a bird’s-eye view of the experimental economics literature on the topic. The median estimates from each study, recomputed to partial correlations for comparability, range commonly between 0 and 0.2, though some studies report correlations of −0.3 or 0.5. The estimates do not converge to a consensus value. Figure 2 shows that results vary across countries. Surprisingly to an economist, estimates are far from being robustly and consistently positive.
</p>

<p>
Figure 1 Estimated effects of incentives vary across studies, and do not converge to a consensus value
</p>

<p>
Figure 2 Effects of incentives vary across and within countries, and centre around small positive values
</p>

<p>
None of the previous meta-analyses in psychology (Jenkins et al. 1998, Condly et al. 2003, Weibel et al. 2010, Garbers and Konradt 2014, Kim et al. 2022) corrected the literature for publication bias. Publication bias arises when some results – typically those that are intuitive and statistically significant – are preferentially selected for publication. Selective reporting can work at the level of entire studies – for example, studies may end up unpublished, forever hidden in a file drawer, because of their insignificant results.
</p>

<p>
More plausibly, however, selective reporting works as self-censorship practised by the authors themselves (Brodeur et al. 2022). In the context of the incentive performance literature, researchers can, for example, alter the measure of performance they report (Esteves-Sorenson and Broce 2022) or choose a subset of the data until they get a desired outcome.
</p>

<p>
Selective reporting does not equal cheating and can be unintentional. McCloskey and Ziliak (2019) draw a useful analogy to the Lombard effect in psychoacoustics: speakers involuntarily increase their vocal effort in the presence of noise. In a similar way, researchers may increase their effort to find a plausible estimate when there is noise in their data.
</p>

<p>
Figure 3 shows two graphs used to detect publication bias. The left-hand panel plots estimates on the horizontal axis against their precision on the vertical axis. In the absence of bias, this ‘funnel plot’ should be symmetrical. But we can see that many negative estimates are missing – perhaps not reported because they are not intuitive. In a similar vein, the right-hand panel shows that estimates are just statistically insignificant at the 5% level, and thus with t-statistics just below 1.96 in absolute value, are under-reported.
</p>

<p>
We use a battery of statistical methods that formalise the ideas of both panels of Figure 3. All methods find evidence of publication bias, which pushes the mean reported estimate upwards. After correction for the bias, the mean experimental result suggests a negligible effect of financial incentives on performance – a striking result to an economist.
</p>

<p>
The experiments on this topic vary so much that a reader may ask how a mean estimate is informative. Indeed, researchers focus on different definitions of performance: work outcomes, school grades, games, blood donations, among others. The task itself can be appealing or unappealing, cognitive or manual. Outputs can be measured quantitatively or qualitatively.
</p>

<p>
Reward size and framing also differ across experiments – sometimes only individual people are paid, sometimes the rewards are group-specific. Some experiments are conducted in a lab, most are field studies. Subjects differ in terms of gender, occupation, age, and culture. Various statistical techniques are used to produce the results.
</p>

<p>
To account for the differences in estimation context, we employ Bayesian model averaging, the natural solution to model uncertainty in the Bayesian framework (Steel 2020). The results suggest that some method choices drive the results systematically, as depicted in Figure 4. The most important drivers of heterogeneity are shown on the top. Blue denotes a positive effect on the estimated incentive-performance nexus, red denotes a negative effect. Columns denote different models, and the horizontal axis measures each model’s importance.
</p>

<p>
The composition of the subject pool matters, as does the framing of rewards, individual versus group rewards, and qualitative versus quantitative measurement of output. Financial incentives are even less efficient in improving grades and pro-social behaviour than they are in improving performance at games and work. But, importantly, the differences are small.
</p>

<p>
The implied correlations for various experimental contexts after correction for bias and accounting for model uncertainty are always statistically insignificant and negligible according to the Doucouliagos (2011) guidelines for the interpretation of partial correlations. The only exception is laboratory experiments, but even here the implied effect is tiny. In Cala et al. (2022) we show the numerical results for different scenarios.
</p>

<p>
We conclude that, regarding the effect of financial incentives on performance, the experimental economics literature is inconsistent with most models commonly employed in economics.
</p>

<p>
The results do not fit neatly in the mainstream psychology framework either. The motivation crowding theory assumes that the crowding out of intrinsic motivation happens only in the case of interesting tasks, exactly as reported by Weibel et al. (2010). The problem is that the definition of an interesting task is subjective, and some people will enjoy tasks that others find unappealing.
</p>

<p>
Another potential explanation is that reward cues distract people from the task itself; a recent meta-analysis shows that this effect can be important (Rusz et al. 2020). The distraction effect can exist for both interesting and uninteresting tasks and is more likely in field settings, where the experimenter does not always have full control over the connection between reward cues and the task itself.
</p>]]></content:encoded>
    </item>
    <item>
      <title>Armington elasticity and international trade models: Fifty years on</title>
      <link>https://meta-analysis.cz/notes/voxeu-armington/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/voxeu-armington/</guid>
      <pubDate>Wed, 23 Sep 2020 00:00:00 +0000</pubDate>
      <description>A meta-analysis of 3,524 estimates of the Armington elasticity of substitution between domestic and foreign goods, corrected for publication bias, implies a range of 2.5-5.1 with a median of 3.8, equivalent to a trade cost elasticity of 2.8.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://cepr.org/voxeu/columns/armington-elasticity-and-international-trade-models-fifty-years" rel="external">VoxEU / CEPR</a>, 23 September 2020. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/voxeu-armington/">/komentare/voxeu-armington/</a>.</p>

<p>
A key parameter informing policy models in international economics is the elasticity of substitution between domestic and foreign goods, also known as the Armington elasticity. Yet elasticity estimates have varied widely since Armington’s seminal 1969 contribution. This column considers 3,524 previous estimates and discusses how these historical analyses can be corrected for various biases. The previous research implies that the elasticity lies in the range 2.5-5.1 with a median of 3.8. In a simple model this translates to a trade cost elasticity of 2.8.
</p>

<p>
How does the demand for domestic versus foreign goods react to a change in relative prices? The answer is central to a host of policy problems in international trade and macroeconomics. To name but a few, these include the welfare effects of globalisation (Costinot and Rodriguez-Clare 2014), trade balance adjustments (Imbs and Mejean 2015), and the exchange rate pass-through of monetary policy (Auer and Schoenle 2016). In particular, any attempt to evaluate the effect of tariffs depends crucially on the assumed reaction of relative demand to relative prices (Waugh 2019, Freund et al. 2020, Larch et al. 2019).
</p>

<p>
In most models, the demand reaction is governed by the constant elasticity of substitution between domestic and foreign goods. The size of the elasticity used for calibration often drives the conclusions of the model, as shown by Schurenberg-Frosch (2015), who recomputes the results of 50 previously published models using different values for the elasticity. She finds that, with plausible changes in the elasticity, the results change qualitatively in more than half of the cases. As Hillberry and Hummels (2013) put it, “it is no exaggeration to say that [the elasticity] is the most important parameter in modern trade theory”.
</p>

<p>
Elasticity is also crucial within discussions concerning unconventional monetary policy, a topic that will gain renewed urgency if the Covid-19 malaise lingers into 2021 and beyond. Consider, for example, two European central banks that, in the wake of the Great Recession, introduced exchange rate floors to limit their currencies' appreciation against the euro: the Swiss National Bank and the Czech National Bank. Currency depreciation (relative to the counterfactual without the currency floor) produces two effects relevant to the aggregate price level.
</p>

<p>
First, imported goods become more expensive, which directly increases inflation. With a large elasticity of substitution between domestic and foreign goods, however, this effect becomes muted and delayed because consumers shift toward relatively cheaper domestic goods. Second, currency depreciation stimulates the economy by encouraging exports and discouraging imports, which raises inflation in the medium term. With a larger elasticity of substitution, this effect strengthens. Because both the Swiss National Bank and the Czech National Bank use open-economy dynamic stochastic general equilibrium models for policy analysis, the assumed size of the ‘Armington elasticity’ played an important role in the decision on when and how to implement the exchange rate floor.
</p>

<p>
Despite this, no consensus on the magnitude of the elasticity exists. In different contexts, researchers tend to obtain substantially different estimates, as observed by Feenstra et al. (2018) and many other commentators before them. In new paper (Bajzik et al. 2020), we assign a pattern to these differences, a pattern that we hope will be useful for calibrating policy models in international trade and macroeconomics. As the Armington-style literature reaches its 50th anniversary, the time is ripe for taking stock. In one of the largest meta-analyses ever conducted in economics, we collect 3,524 estimates of the elasticity of substitution between domestic and foreign goods and construct 32 variables that reflect the context in which researchers produce their estimates.
</p>

<p>
A bird's-eye view of the literature (Figure 1 and Figure 2) shows three stylised facts. First, the estimates of the elasticities vary substantially, even within individual countries. A researcher wishing to calibrate her policy model has plenty of degrees of freedom; she can easily find empirical evidence for any value of the elasticity between zero and eight. Such plausible (that is, justifiable by an empirical study) changes in the elasticity can have decisive effects on the results of the model. For example, Engler and Tervala (2016) show that changing the elasticity from three to eight more than doubles the estimated welfare gains from the Transatlantic Trade and Investment Partnership.
</p>

<p>
Second, the reported elasticity seems to be increasing in time, but it is not clear whether the apparent trend reflects fundamental changes in preferences or improved data and techniques used by more recent studies. Finally, the third stylised fact is that newer studies show more disagreement on the value of the elasticity of substitution. That is, instead of converging to a consensus value, the literature diverges.
</p>

<p>
The increased variance in the estimated elasticities provides additional rationale for a systematic evaluation of the published results. For this evaluation we use the methods of meta-analysis, which were originally developed in, or inspired by, medical research. Recent applications of meta-analysis in economics include Imai et al. (2020) on the present bias, Card et al. (2018) on the effectiveness of active labor market programmes, and Havranek and Irsova (2017) on the border effect in international trade. For recent Vox columns on meta-analysis, see, for example, Ahlfeldt and Pietrostefani (2019), Bruno et al. (2017), and Rose (2016).
</p>

<p>
Meta-analysis allows us to correct for potential publication bias. Publication bias arises when, holding other aspects of study design constant, some results (for example, those that are statistically insignificant at standard levels or have the ‘wrong’ sign) have a lower probability of publication than other results (Stanley 2001). In the context of the elasticity of substitution, it is safe to assume that its sign is positive: a negative value is not compatible with any commonly applied model of preferences. Similarly, it is difficult to interpret a zero elasticity.
</p>

<p>
As a result, from the point of view of an individual study it makes sense not to report such unintuitive estimates, but instead to find a specification where the elasticity is positive (because non-positive elasticity suggests that something is wrong with the data or the model). Nevertheless, non-positive estimates will occur from time to time simply because of sampling error. For the same reason, researchers will sometimes obtain estimates much larger than the true value. If large estimates (which can be intuitive) are kept but non-positive ones are omitted, an upward bias arises.
</p>

<p>
Paradoxically, publication bias can improve inferences drawn from some individual studies since they avoid making central conclusions based on negative or zero elasticities. However, publication bias inevitably distorts inference drawn from the literature as a whole. Ioannidis et al. (2017) show that, in economics, the effects of publication selection are dramatic and exaggerate the mean reported estimate twofold.
</p>

<p>
To correct for publication bias, we use meta-regression techniques based on Egger et al. (1997) and their extensions, together with three new non-linear techniques developed specifically for meta-analysis in economics. The first one is due to Ioannidis et al. (2017) and relies on estimates that are adequately powered. The second technique was developed by Andrews and Kasy (2019) and employs a selection model that estimates the probability of publication for results with different ‘p-values’. The third non-linear technique is the so-called stem-based method by Furukawa (2019), a non-parametric estimator that exploits the variance-bias trade-off. All techniques suggest an upward bias in the mean reported elasticities due to publication selection.
</p>

<p>
While publication selection creates an upward bias, estimates of lower quality seem to yield a downward bias. We exploit our large dataset and the relationships unearthed by model averaging analysis of 32 variables (reflecting the context in which the estimates were obtained) to compute a mean effect corrected for publication bias but conditional on the design of the most reliable studies. In an alternative approach, we divide the estimates into groups based on the quality of the journal and the preferences of the authors themselves, while still controlling for publication bias.
</p>

<p>
The implied Armington elasticity corrected for the biases lies in the range 2.5-5.1, with a median estimate at 3.8. We interpret the number and the interval as our best guess on how to calibrate a model that allows for only one parameter to govern the aggregate elasticity of substitution between domestic and foreign goods (for example, an open economy dynamic stochastic general equilibrium model of the type used in many central banks). In our paper we also report implied mean elasticities for individual countries. The results are relevant to a broader discussion on the trade costs elasticity, because in a simple setting an Armington elasticity of 3.8 translates to a trade cost elasticity of 2.8 (see Costinot and Rodriguez-Clare 2014).
</p>]]></content:encoded>
    </item>
    <item>
      <title>Revision of Reporting Guidelines</title>
      <link>https://meta-analysis.cz/notes/maer-reporting-guidelines/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-reporting-guidelines/</guid>
      <pubDate>Mon, 11 Nov 2019 00:00:00 +0000</pubDate>
      <description>Tomáš Havránek proposes 12 recommendations to revise the reporting guidelines for meta-analysis in economics, covering weights, outliers, reconstructed standard errors, clustering, model averaging, robustness checks, and data sharing, and invites MAER-Net members to weigh in.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/revision-of-reporting-guidelines" rel="external">MAER-Net</a>, 11 November 2019. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-reporting-guidelines/">/komentare/maer-reporting-guidelines/</a>.</p>

<p>
Dear fellow meta-analysts:
</p>

<p>
It has been 7 years since we drafted the current <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/joes.12008">reporting guidelines</a>. I believe the time is ripe for updating them – and based on our debate at the MAER-Net Open Forum in Greenwich, I think at least some of you share that belief. Let me kick off a discussion on the shape of the possible revision.
</p>

<p>
The intention is not to create a list of rules and then cast them in stone. Rather, the guidelines should, apart from defining the minimum standard for conducting a meta-analysis in economics (well covered by the current guidelines), offer a set of practical recommendations. These recommendations may change again in a couple of year as our tools evolve.
</p>

<p>
My motivation for proposing the update is that since I started to serve as an associate editor at the Journal of Economic Surveys, I’ve handled quite a few meta-analyses with basic econometric and interpretation errors, studies that do not enhance the reputation of meta-analysis in economics. I find myself providing similar feedback all over again (and the same, I assume, goes for Tom and Chris at the JoES and many of you who write referee reports), so I believe a concrete set of recommendations published in the JoES will help.
</p>

<p>
Below I offer 12 subjective recommendations that I miss in the current guidelines. You may disagree about some points or may want to add others. We will prepare the revision of the guidelines (if any) based on the discussion that will follow. Thank you for contributing!
</p>

<ul>
<li><b>Weights</b>. If the meta-analyst doesn't use inverse-variance weights, she or he should explain why. The requirement is in line with the recent <a href="https://www.nature.com/articles/nature25753?sf184289833=1">review paper on meta-analysis in Nature</a>: <i>“Meta-analyses that are not weighted by inverse variances are common and often poorly justified.”</i></li>
<li><b>Outliers</b>. The meta-analyst is encouraged to specify how outliers, both in the estimated effects and standard errors, are treated: whether all observations are included, or which rule was used to omit outlying observations (for example, Hadi or winsorizing and the respective thresholds).</li>
<li><b>Reconstructed standard errors.</b> If standard errors are not directly reported in some primary studies, the meta-analyst should state how the standard errors were obtained (for example, using the delta method with the assumption of zero covariance). A robustness check is encouraged that excludes observations with reconstructed standard errors.</li>
<li><b>Study-level dummies</b>. When using meta-regression analysis to investigate the extent of publication bias, the meta-analyst is encouraged to include a robustness check with study-level dummies (fixed effects in the econometric sense) and thus control for unobserved characteristics of individual studies. Note that such specification only captures within-study bias (p-hacking).</li>
<li><b>Random effects</b>. Study-level random effects in economics meta-analyses can be correlated with publication bias or other aspects of studies. The meta-analyst should exercise caution when adding random effects to multiple MRA, because doing so likely violates the exogeneity condition.</li>
<li><b>Clustering or bootstrapping</b>. Standard errors in meta-regression analysis should be clustered or bootstrapped. <a href="https://ideas.repec.org/c/boc/bocode/s458121.html">Bootstrapping </a>is the only viable option when the number of studies is small.</li>
<li><b>General-to-specific</b>. Instead of sequential t-tests, in multiple MRA it is recommended to use the more holistic general-to-specific approach due to <a href="https://ideas.repec.org/p/fip/fedgif/838.html">Hendry and colleagues</a>. Alternatively, the meta-analyst may want to employ Bayesian or frequentist model averaging to address <a href="https://www.maer-net.org/post/model-averaging">model uncertainty</a>.</li>
<li><b>Sensitivity of model averaging</b>. If the meta-analyst uses Bayesian or frequentist model averaging, she or he should report robustness checks that show how the results depend on the selected priors (in the Bayesian case) or the selected weights (Mallows or other; in the frequentist case). The procedure employed to simplify model space (Markov Chain Monte Carlo or orthogonalization) should be mentioned.</li>
<li><b>Collinearity</b>. The meta-analyst is encouraged to report collinearity statistics for multiple MRA, for example the correlation matrix or variance-inflation factors. Note that collinearity increases when inverse-variance weights are used.</li>
<li><b>Robustness checks</b>. The meta-analyst should report robustness checks to the baseline test of publication bias and the underlying effect. Note that different estimators have different performance in different environments (as shown by <a href="https://journals.sagepub.com/doi/abs/10.1177/2515245919847196?journalCode=ampa">Carter et al.</a>). Choose several robustness checks, for example: <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/jrsm.1095">PET-PEESE</a>, <a href="https://onlinelibrary.wiley.com/doi/abs/10.1111/ecoj.12461">WAAP</a>, <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/jrsm.1352">Bom &amp; Rachinger</a>, <a href="https://ideas.repec.org/p/zbw/esprep/194798.html">Furukawa</a>, <a href="https://projecteuclid.org/euclid.ss/1177011364">Hedges </a>(and variants thereof), <a href="https://www.aeaweb.org/articles?id=10.1257/aer.20180310">Andrews &amp; Kasy</a>, <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2377290">p-curve</a>, <a href="https://osf.io/preprints/metaarxiv/zqjr9/">p-uniform</a>.</li>
<li><b>Data</b>. The meta-analyst is encouraged to provide the data to the editor and referees so that they can check their structure. This can be done either through the journal's submission system or (preferably) publicly through the author's website.</li>
<li><b>Economic significance</b>. The meta-analyst should discuss the economic significance of results. For example, publication bias or the underlying effect can be significant statistically, but not material in practice. If partial correlation coefficients are used, <a href="https://ideas.repec.org/p/dkn/econwp/eco_2011_5.html">Doucouliagos's guidelines</a> for the practical strength of the effect should be consulted.</li>
</ul>]]></content:encoded>
    </item>
    <item>
      <title>Death to the Cobb-Douglas Production Function!</title>
      <link>https://meta-analysis.cz/notes/maer-cobb-douglas/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-cobb-douglas/</guid>
      <pubDate>Wed, 25 Sep 2019 00:00:00 +0000</pubDate>
      <description>A meta-analysis of 3,186 estimates from 121 studies finds a mean capital-labor elasticity of substitution of 0.9, close to the Cobb-Douglas value of 1, but correcting for publication bias, data aggregation, and omitted first-order conditions lowers the recommended calibration to 0.3.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/death-to-the-cobb-douglas-production-function" rel="external">MAER-Net</a>, 25 September 2019. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-cobb-douglas/">/komentare/maer-cobb-douglas/</a>.</p>

<p>
<i>A key parameter in economics is the elasticity of substitution between capital and labor, but empirical estimates vary. This column takes stock of the literature. Among 3,186 estimates produced by 121 studies, the mean elasticity is 0.9, not far from the Cobb-Douglas assumption of 1. But the mean is biased by 3 factors: publication selection, use of aggregated data, and omission of the first-order condition for capital. The mean corrected for these biases is 0.3. The weight of evidence accumulated in the empirical literature thus emphatically rejects the Cobb-Douglas specification.</i>
</p>

<p>
The elasticity of substitution between capital and labor is central to a host of economic problems. Our understanding of long-run growth depends on the value of the elasticity (Solow, 1956). Klump and de La Grandeville (2000) suggest that a larger elasticity in a country results in higher per capita income at any stage of development. Turnovsky (2002) argues that a smaller elasticity leads to faster convergence. The explanation for the decline of the labor share in income during the recent decades that was put forward by Piketty (2014) and Karabarbounis and Neiman (2013) holds only when the elasticity surpasses one. Cantore et al. (2014) show how the effect of technology shocks on hours worked is sensitive to the elasticity. Nekarda and Ramey (2013) argue that the countercyclicality of the price markup over marginal cost also depends on the elasticity of substitution. In addition, the elasticity represents an important parameter in analyzing the effects of fiscal policies, including the effect of corporate taxation on capital formation, and in determining optimal taxation of capital (Chirinko, 2002).
</p>

<p>
<b>Figure 1</b>: The elasticity of substitution matters for monetary policy
</p>

<p>
<i>Note: The figure shows simulated impulse responses of inflation to a monetary policy shock. We use the SIGMA model of Erceg et al. (2008) developed for the Federal Reserve Board and vary the value of the capital-labor substitution elasticity while leaving other parameters at their original values. The model does not have a stable solution for the elasticity larger than one.</i>
</p>

<p>
The size of the elasticity has practical consequences for monetary policy, as Figure 1 illustrates. For example, in the SIGMA model used by the Federal Reserve Board (Erceg et al., 2008), the effectiveness of interest rate changes in steering inflation doubles when one assumes the elasticity to equal 0.9 instead of 0.5, yielding wildly different policy implications. We choose the SIGMA model for the illustration because, as one of very few models employed by central banks, it actually allows for different values of the elasticity of substitution. Almost all models use the convenient simplification of the Cobb-Douglas production function, which implicitly assumes that the elasticity equals one. If the true elasticity is smaller, these models overstate the strength of monetary policy and should imply a more aggressive campaign of interest rate cuts in response to a recession.
</p>

<p>
Empirical estimates of the elasticity vary widely both within and between studies (Figure 2). Solid references can be found for calibrating the elasticity anywhere between 0 and 1.5, which effectively means that the empirical literature does not discipline calibrations at all – despite decades of research on the value of the elasticity and the work of dozens of prominent economists.
</p>

<p>
<b>Figure 2</b>: No consensus on the value of the elasticity
</p>

<p>
<b>Publication bias</b>
</p>

<p>
To take stock of the voluminous literature and provide concrete guidelines for the calibration of the elasticity, we conduct a meta-analysis (Gechert et al., 2019).[1] We collect 3,186 coefficients from 121 studies, which produce a mean estimate of 0.9. But we show that the picture is seriously distorted by publication bias. After correcting for the bias, the mean reported elasticity shrinks to 0.5. This correction alone implies halving the effectiveness of monetary policy in a structural model, as shown by Figure 1. Moreover, some data and method choices bias the estimated elasticity systematically. If one agrees that sector-level data dominate more aggregated country- or state-level data and that including information from the first-order condition for capital dominates ignoring it, the implied mean estimate further decreases to 0.3. Thus we recommend 0.3 for the calibration of the elasticity, consistent with burying the Cobb-Douglas production function.
</p>

<p>
The finding of strong publication bias predominates in our results and is responsible for most of the reduction from 0.9 (the simple mean) to 0.3 (our recommended calibration). The bias arises when different estimates have a different probability of being reported depending on sign and statistical significance. The identification we use builds on the fact that almost all econometric techniques used to estimate the elasticity assume that the ratio of the estimate to its standard error has a symmetrical distribution, typically a <i>t</i>-distribution. So the estimates and standard errors should represent independent quantities. But if statistically significant positive estimates are preferentially selected for publication, large standard errors (given by noise in data or imprecision in estimation) become associated with large estimates.
</p>

<p>
Because empirical economists command plenty of degrees of freedom, a large estimate of the elasticity can always emerge if the researcher looks for it long enough, and an upward bias in the literature arises. A useful analogy appears in McCloskey and Ziliak (2019), who liken publication bias to the Lombard effect in biology: speakers increase their effort in the presence of noise. Apart from linear techniques based on the Lombard effect, we employ recently developed methods by Ioannidis et al. (2017), Andrews and Kasy (2019), Bom and Rachinger (2019), and Furukawa (2019), which account for the potential nonlinearity between the standard error and selection effort.
</p>

<p>
<b>Figure 3</b>: Publication bias in the literature
</p>

<p>
Figure 3 provides a graphical illustration of the mechanism outlined in the previous paragraph. In the scatter plot the horizontal axis measures the magnitude of the estimated elasticities, and the vertical axis measures their precision. In the absence of publication bias, the scatter plot will form an inverted funnel: the most precise estimates will lie close to the true mean elasticity, imprecise estimates will be more dispersed, and both small and large imprecise estimates will appear with the same frequency. The figure shows the predicted funnel shape, still with plenty of heterogeneity at the top (see the paper for a detailed analysis of heterogeneity using both Bayesian and frequentist model averaging) – but also shows asymmetry. For the funnel to be symmetrical, and hence consistent with the absence of publication bias, we should observe many more reported negative and zero estimates. The finding of strong publication selection and the corrected mean effect around 0.5 is confirmed by all techniques we apply.
</p>

<p>
<b>Concluding remarks</b>
</p>

<p>
We are not the first to highlight the disconnect between the Cobb-Douglas specification commonly used in macroeconomic models and the empirical literature estimating the elasticity of substitution. Chirinko (2008) and Knoblach et al. (2019) provide useful surveys of portions of the literature, and both studies suggest that the Cobb-Douglas production function is not backed by the available evidence. We argue that after controlling for publication bias the case against Cobb-Douglas strengthens to the point where one must warn against the continued use of this convenient simplification. As we show in Figure 1, a structural model built to aid monetary policy is biased from the beginning if it uses an elasticity of one for capital-labor substitution. Computational convenience should yield to the stylized fact established by half a century of meticulous research: capital and labor are gross complements.
</p>

<p>
<b>References</b>
</p>

<ul>
<li>Andrews, I. &amp; M. Kasy (2019): “Identification of and Correction for Publication Bias.” Americal Economic Review 109(8): pp. 2766-2794.</li>
<li>Bom, P. R. D. &amp; H. Rachinger (2019): “A Kinked Meta-Regression Model for Publication Bias Correction.” Research Synthesis Methods, forthcoming.</li>
<li>Cantore, C., M. Leon-Ledesma, P. McAdam, &amp; A. Willman (2014): “Shocking Stuff: Technology, Hours, and Factor Substitution.” Journal of the European Economic Association 12(1): pp. 108-128.</li>
<li>Chirinko, R. S. (2002): “Corporate Taxation, Capital Formation, and the Substitution Elasticity between Labor and Capital.” National Tax Journal 55(2): pp. 339-355.</li>
<li>Chirinko, R. S. (2008): “σ: The Long and Short of it.” Journal of Macroeconomics 30(2): pp. 671 686.</li>
<li>Erceg, C. J., L. Guerrieri, &amp; C. Gust (2008): “Trade Adjustment and the Composition of Trade.” Journal of Economic Dynamics and Control 32(8): pp. 2622-2650.</li>
<li>Furukawa, C. (2019): “Publication Bias under Aggregation Frictions: Theory, Evidence, and a New Correction Method.” <a href="https://ideas.repec.org/p/zbw/esprep/194798.html">Unpublished paper</a>, MIT.</li>
<li>Gechert, S., T. Havranek, Z. Irsova, &amp; D. Kolcunova (2019): “Death to the Cobb-Douglas Production Function.” <a href="https://www.boeckler.de/pdf/p_fmm_wp_imk_51_2019.pdf">FMM working paper 51</a>, Hans-Böckler-Stiftung.</li>
<li>Ioannidis, J., T. Stanley, &amp; H. Doucouliagos (2017): “The Power of Bias in Economics Research.” Economic Journal 127(605): F236-F265.</li>
<li>Karabarbounis, L. &amp; B. Neiman (2013): “The Global Decline of the Labor Share.” Quarterly Journal of Economics 129(1): pp. 61-103.</li>
<li>Klump, R. &amp; O. de La Grandville (2000): “Economic Growth and the Elasticity of Substitution: Two Theorems and Some Suggestions.” American Economic Review 90(1): pp. 282-291.</li>
<li>Knoblach, M., M. Rossler, &amp; P. Zwerschke (2019): “The Elasticity of Substitution Between Capital and Labour in the US Economy: A Meta-Regression Analysis.” Oxford Bulletin of Economic and Statistics, forthcoming.</li>
<li>McCloskey, D. N. &amp; S. T. Ziliak (2019): “What Quantitative Methods Should We Teach to Graduate Students? A Comment on Swann's ‘Is Precise Econometrics an Illusion?’” Journal of Economic Education, forthcoming.</li>
<li>Nekarda, C. J. &amp; V. A. Ramey (2013): “The Cyclical Behavior of the Price-Cost Markup.” <a href="https://www.nber.org/papers/w19099">NBER Working Paper</a> 19099.</li>
<li>Piketty, T. (2014): “Capital in the 21st Century.” Cambridge, MA: Harvard University Press.</li>
<li>Solow, R. M. (1956): “A Contribution to the Theory of Economic Growth.” Quarterly Journal of Economics 70(1): pp. 65-94.</li>
<li>Turnovsky, S. J. (2002): “Intertemporal and Intratemporal Substitution, and the Speed of Convergence in the Neoclassical Growth Model.” Journal of Economic Dynamics and Control 26(9-10): pp. 1765-1785.</li>
</ul>

<p>
<b>Endnotes</b>
</p>

<p>
[1] The full paper, together with data and codes, is available at <a href="https://meta-analysis.cz/sigma">https://meta-analysis.cz/sigma</a>.
</p>]]></content:encoded>
    </item>
    <item>
      <title>Why Model Averaging Is Useful in Meta-Analysis</title>
      <link>https://meta-analysis.cz/notes/maer-model-averaging/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/maer-model-averaging/</guid>
      <pubDate>Wed, 27 Feb 2019 00:00:00 +0000</pubDate>
      <description>An introduction to model averaging as a response to model uncertainty in regression and meta-regression, explaining why weighting many specifications by fit and parsimony beats picking a single best model.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://www.maer-net.org/post/model-averaging" rel="external">MAER-Net</a>, 27 February 2019. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/maer-model-averaging/">/komentare/maer-model-averaging/</a>.</p>

<p>
If you’ve ever run a regression with more than a handful of variables, you know the problem. It’s called <b>model uncertainty</b>: which variables should go into the baseline? One solution is to simply put all in one model, but doing so attenuates precision. Another approach is to get rid of the variables that are more difficult to interpret, but that’s not kosher either:
</p>

<blockquote><p>It is important to realize that this uncertainty is an inherent part of economic modelling, whether we acknowledge it or not. Putting on blinkers and narrowly focusing on a limited set of possible models implies that we may fail to capture important aspects of economic reality. (<a href="https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/steel/steel_homepage/techrep/modelaveraging_jel_final.pdf">Steel, 2019</a>)</p></blockquote>

<p>
A formal response to model uncertainty is <b>model averaging</b>. The idea is to run regression models with different combinations of variables, and then give these models weights based on how they fit the data and how parsimonious, and possibly well-specified, they are. I’ll write from an economist’s perspective (model averaging has its <a href="https://link.springer.com/article/10.1057/jors.1969.103">roots </a> in economics), but I hope this post will appeal to anyone who uses regression analysis: social scientists, medical scientists, ecologists, biologists, meta-analysts.
</p>

<h2>How it Works</h2>

<p>
Suppose we have 5 explanatory variables. To do model averaging, we first estimate 2^5 = 32 regressions. Then we assign each model a weight and compute a <b>weighted average</b> over the 32 regressions. The weight increases with data fit, but decreases with model complexity (given the same fit, a regression with 4 variables will get more weight than a regression with 5 variables). So, think of adjusted R-squared as an intuitive weight for model averaging. It’s not an optimal weight, but you get the idea.
</p>

<p>
Many methods of model averaging exist, and the technicalities may look intimidating. Luckily for us, the statistician Mark Steel has crafted a <a href="https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/steel/steel_homepage/techrep/modelaveraging_jel_final.pdf"><b>superb survey</b></a><b>,</b> which is now forthcoming in the Journal of Economic Literature. Mark’s paper has it all: it’s easily accessible to scientists from various fields, but also contains useful technical details. In this post I quote Mark’s paper liberally – I couldn’t put it better.
</p>

<h2>Key for Meta-Analysis</h2>

<p>
Mark makes a strong case that model averaging should form an essential part of an empirical economist’s toolkit. If that holds for economics, it must hold double for <b>meta-regression </b>analysis (in any field), where model uncertainty <a href="https://journals.sagepub.com/doi/10.1177/1740774508101279">runs rampant</a>. In most meta-analysis contexts we can think of <a href="https://meta-analysis.cz/excess_sensitivity/">dozens</a> of factors that may influence (or just be correlated with) the reported outcomes. Sometimes there’s the theory to guide us, but more often than not we are on our own.
</p>

<p>
Our traditional response to model uncertainty is <b>model selection</b>: let’s choose the best model, either via sequential t-tests or sophisticated <a href="https://onlinelibrary.wiley.com/doi/full/10.1002/jae.615">general-to-specific</a> modelling. But with 20 variables (a common number in economics meta-analyses), we already have more than a million, 2^20, models to choose from – so we’re virtually guaranteed to choose the wrong one. One can’t possibly run thorough specification checks for a million models. In addition, as Mark explains:
</p>

<blockquote><p>The most important common characteristic of model selection methods is that they choose a model and then conduct inference conditionally upon the assumption that this model actually generated the data. So these methods only deal with the uncertainty in a limited sense: they try to select the ‘best’ model, and their inference can only be relied upon if that model happens to be (a really good approximation to) the data generating process. In the much more likely case where the best model captures some aspects of reality, but there are other models that capture other aspects, model selection implies that our inference is almost always misleading, either in the sense of being systematically wrong or overly precise. (<a href="https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/steel/steel_homepage/techrep/modelaveraging_jel_final.pdf">Steel, 2019</a>)</p></blockquote>

<p>
Sure, that’s not to say we should bury the theory for good and just go all in for data mining. We always have some idea about the <b>structure of the underlying model</b>. We know that certain variables must play a role – so we can pin these to all regressions and only use model averaging to work over the ‘control’ variables (for which we have no theory, but still don’t want to ignore them). Roman Horvath, Zuzana Irsova, Marek Rusnak, and myself use this strategy in our 2015 <a href="https://meta-analysis.cz/substitution/">Journal of International Economics</a> paper, among other applications.
</p>

<h2>But It’s All Bayesian, Right?</h2>

<p>
The most common model averaging technique is <b>Bayesian model averaging</b> (BMA). Model averaging arises naturally as a response to model uncertainty in the Bayesian setting, but I suspect that most of us BMA-users are not convinced Bayesians. We are Bayesian opportunists. True, frequentist methods of model averaging <a href="https://arxiv.org/abs/1802.03511">do exist</a> (Mark discusses them thoroughly), but BMA is more flexible and easier to compute.
</p>

<p>
Why? With more than 30 variables, one would have to compute over a billion, 2^30, models. That would take you months, so you need to simplify the model space. Such simplification methods (and I won’t go to details here; read Mark’s paper if you’re interested) are readily available for BMA, but not for the frequentist alternatives – well, with some recent <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/jae.2288">exceptions</a>. Mark quotes a line by <a href="https://www.sciencedirect.com/science/article/pii/S0304407608001115">Jonathan Wright</a>, who is spot on:
</p>

<blockquote><p>One does not have to be a subjectivist Bayesian to believe in the usefulness of BMA, or of Bayesian shrinkage techniques more generally. A frequentist econometrician can interpret these methods as pragmatic devices that may be useful for out-of-sample forecasting in the face of model and parameter uncertainty. (<a href="https://www.sciencedirect.com/science/article/pii/S0304407608001115">Wright, 2008</a>)</p></blockquote>

<h2>Priors Matter</h2>

<p>
A disclaimer is in order. If you decide to use BMA, note that different priors will give you different results. The natural baseline is to use <b>agnostic priors</b>: for example, that all coefficients are zero and that the weight of this prior is the same as the weight of one observation of data (so, <a href="https://cran.r-project.org/web/packages/BMS/vignettes/bmsmanual.pdf">pretty small</a>). You must also specify the prior on model space. Again, the natural starting point is a prior in which all models have the same probability. In any case, robustness checks are crucial, and doing frequentist together with Bayesian model averaging in the same paper often yields the ultimate robustness check.
</p>

<p>
Yes, <a href="https://link.springer.com/article/10.1007/s11424-009-9198-y">frequentist techniques</a> do not require priors (not explicitly, anyway) – but you have to choose your assumptions concerning optimal weights. In my eyes, the challenge is essentially the same, but concealed. Mark stresses the importance of priors, but the same goes for assumptions in the frequentist setting:
</p>

<blockquote><p>For BMA, it is important to understand that the weights (based on posterior model probabilities) are typically quite sensitive to the prior assumptions, in contrast to the usually much more robust results for the model parameters given a specific model. In addition, this sensitivity does not vanish as the sample size grows. Thus, a good understanding of the effect of (seemingly arbitrary) prior choices is critical. (<a href="https://warwick.ac.uk/fac/sci/statistics/staff/academic-research/steel/steel_homepage/techrep/modelaveraging_jel_final.pdf">Steel, 2019</a>)</p></blockquote>

<h2>Model Averaging Is User-Friendly</h2>

<p>
How difficult is to do model averaging in <b>your next meta-analysis</b>? Not difficult at all. You can use the packages described by Mark or the R code from <a href="https://meta-analysis.cz/">our website</a> (you don’t have to know R syntax to apply these estimators; the code is self-contained). We have three recent papers in which we use both Bayesian and frequentist model averaging and for which we provide data and code: one in the <a href="https://meta-analysis.cz/habits/">European Economic Review</a> (as far as we know, the first application of frequentist model averaging in meta-analysis), another in the <a href="https://meta-analysis.cz/dst/">Energy Journal</a>, and the third in the <a href="https://meta-analysis.cz/education/">Oxford Bulletin of Economics and Statistics</a>.
</p>

<h2>Anything Else?</h2>

<p>
Not related to model averaging, but important news for the meta-analysis community: the <b>American Economic Review</b> has accepted the first regular <a href="https://maxkasy.github.io/home/files/papers/PublicationBias.pdf">paper</a> ever that focuses on meta-analysis and publication bias. Congratulations to Isaiah Andrews and Maximilian Kasy! They also have an <a href="https://maxkasy.github.io/home/metastudy/">app</a> for their method, so now you can do meta-analysis using your iPhone (or Xiaomi Mi, if you prefer). Another <a href="https://github.com/Chishio318/stem-based_method">superb new technique</a> was developed by <b>Chishio Furukawa</b> from the MIT. It’s a clever, intuitive non-parametric estimator that in my experience works great, and I recommend you give it a try. In our new paper conditionally accepted by the <a href="https://meta-analysis.cz/excess_sensitivity/">Review of Economic Dynamics</a> we employ these two techniques along with more traditional ones (including the ingenious and already classical <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/ecoj.12461">WAAP</a> by John Ioannidis, Tom Stanley, and Chris Doucouliagos), and all point to the same direction.
</p>

<p>
By now it won’t surprise you that we do model averaging in that paper, too.
</p>]]></content:encoded>
    </item>
    <item>
      <title>Natural resources and economic growth: the research evidence</title>
      <link>https://meta-analysis.cz/notes/globaldev-natural-resources/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/globaldev-natural-resources/</guid>
      <pubDate>Tue, 13 Feb 2018 00:00:00 +0000</pubDate>
      <description>A review of more than 40 studies on natural resources and economic growth finds only weak support for a resource curse once publication bias and method heterogeneity are accounted for, with the effect strongly dependent on the quality of a country&#x27;s institutions.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://globaldev.blog/natural-resources-and-economic-growth-research-evidence/" rel="external">GlobalDev Blog</a>, 13 February 2018. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/globaldev-natural-resources/">/komentare/globaldev-natural-resources/</a>.</p>

<p>
Can large reserves of oil, coal, diamonds, uranium or other natural resources be translated into national prosperity? This column provides an overview of the evidence from more than 40 studies. One common theme is that the better a country’s national institutions – including stronger rule of law, limited corruption and effective government – the more the economy will benefit from its natural resources. But more needs to be understood about the effect of discoveries of natural resources on inequality.
</p>

<p>
Numerous researchers with an interest in development studies have asked whether natural resources are beneficial for long-term economic growth. Yet despite nearly three decades of intensive empirical research, no consensus has emerged.
</p>

<p>
Observing this lack of consensus, we ask the following:
</p>

<ul>
<li>What is the typical effect of natural resources on economic growth?</li>
</ul>

<ul>
<li>Why do the reported results differ so much?</li>
</ul>

<ul>
<li>Are there any policy-relevant factors that have systematic importance with regard to how natural resources affect economic development?</li>
</ul>

<p>
We look at a large set of empirical studies of the natural resources-economic growth nexus using regression analysis. The results show a contradictory picture: roughly 40% of empirical papers find a negative effect (what is typically called the ‘natural resource curse’); 40% find no effect; and 20% find a positive effect of natural resources on economic growth.
</p>

<p>
We aim to summarize what we know about the quantitative effect of natural resources on economic growth based on previous empirical evidence, and provide guidance for other researchers. A somewhat tacit ambition of our research is to provide policy implications on how to manage natural resources so that societies benefit from them.
</p>

<p>
We collect 43 econometric studies, which report 605 regression estimates of the effect of natural resources on economic growth. Our definition of natural resources is point-source non-renewable resources – those extracted from a narrow geographical or economic base, such as oil, diamonds or metals.
</p>

<p>
We describe a long list of characteristics of these studies (all available in Havranek et al. 2016 ). For example, which explanatory variables are included in the models examining the effect of resources on growth? More specifically, do the empirical studies control for the country’s institutional quality (such as the World Bank’s measures of rule of law) and its potentially non-linear effect on economic growth via natural resources?
</p>

<p>
We also ask: What are the econometric methods employed? Do the studies address endogeneity issues? Are the studies published in prestigious peer-reviewed journals with many citations? Which measure of natural resources do the studies employ: measures of so-called natural resource abundance or dependence?
</p>

<p>
In our meta-analysis, we regress the previously reported estimates of the effect of natural resources on economic growth on the characteristics of data, estimation methods, and other aspects reflecting the context in which the estimates were produced.
</p>

<p>
Our research suggests only very weak support for the idea of the natural resource curse, once publication bias and method heterogeneity are taken into account. Therefore, the pessimistic prediction of some important studies in this field – that the natural resource curse is inevitable – is not warranted when all empirical evidence is considered.
</p>

<p>
Our statistical investigation of publication selection suggests that most researchers do not present results in a systematically biased manner – for example, publishing mostly negative estimates that support the notion of the natural resource curse while hiding more positive results in a ‘file drawer’. Such publication selection bias, resulting from the preference for intuitive or statistically significant results, has previously been shown to plague many fields of empirical economics.
</p>

<p>
We show that several factors systematically account for the heterogeneity in the results from studies of the effect of natural resources on economic growth:
</p>

<ul>
<li>Controlling for institutional quality.</li>
</ul>

<ul>
<li>Controlling for the level of investment activity.</li>
</ul>

<ul>
<li>Differentiating between resource dependence and abundance – that is, differentiating between flows and stocks of natural resources.</li>
</ul>

<ul>
<li>Distinguishing between different types of natural resources – in particular, oil seems to bring more benefits than other natural resources.</li>
</ul>

<p>
Especially important is the effect of institutional quality: Countries with better institutions tend to benefit much more from having a wealth of natural resources.
</p>

<p>
What should be the next steps? In our view, we know little about the effect that natural resources discoveries can have on inequality of income and wealth. In addition, more appropriate identification of the precise effect of natural resources on economic growth is vital (for example, using micro-level evidence or synthetic control methods).
</p>]]></content:encoded>
    </item>
    <item>
      <title>Daylight saving saves no energy</title>
      <link>https://meta-analysis.cz/notes/voxeu-daylight-saving/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/voxeu-daylight-saving/</guid>
      <pubDate>Sat, 02 Dec 2017 00:00:00 +0000</pubDate>
      <description>A meta-analysis of 162 estimates from 44 studies finds no publication bias and an essentially zero average effect of daylight saving time on energy consumption, with even the best case, Norway, saving only about 0.3% of annual energy use.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://cepr.org/voxeu/columns/daylight-saving-saves-no-energy" rel="external">VoxEU / CEPR</a>, 2 December 2017. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/voxeu-daylight-saving/">/komentare/voxeu-daylight-saving/</a>.</p>

<p>
The original rationale for daylight saving time was energy savings. This column reveals, however, that the modern empirical literature on the topic finds no savings on average. The extent of savings is related to latitude – regions at higher latitude enjoy slightly more savings, but subtropical regions consume more energy because of daylight saving time. Even in Scandinavia, the savings amount to just 0.3% of annual energy consumption. Policymakers must look at other effects of daylight saving time to justify the continued use of the policy.
</p>

<p>
Most Europeans and Americans are taught at school that daylight saving time (DST) reduces energy consumption. This common wisdom has been repeated by such august authors as Thaler and Sunstein, who praise DST in their book Nudge (Thaler and Sunstein 2008). Of course, daylight saving time was originally adopted by several countries during WWI to reduce energy use, but academic research on the extent of energy savings related to DST in a modern economy is surprisingly thin. We have gone through journal articles, unpublished working papers, energy company reports, government papers, and PhD dissertations and found 44 usable studies carried out since the pioneering Ebersole (1974) report. Unfortunately, a casual look at the literature does not help us. Estimates are all over the place, and are far from converging to a consensus value.
</p>

<p>
Note: Studies measure the effect of daylight saving time on energy use, so a negative estimate implies energy savings.
</p>

<p>
Two existing surveys of the literature – Reincke and van den Broek (1999) and Aries and Newsham (2008) – also show that different researchers obtain substantially different results. One can find evidence in support of energy savings resulting from DST, just as one can find evidence of increased energy demand associated with DST. For example, the most-cited empirical study, Kotchen and Grant (2011), concludes that, contrary to the policy’s objective, DST increases energy use. (The result might be the reason why the study receives so many citations, although it was also published in a prestigious journal, The Review of Economics and Statistics.) Aries and Newsham (2008: 1864) conclude that “the existing knowledge about how DST affects energy use is limited, incomplete, or contradictory”. Indeed, one can find contradictory results even within many individual studies, as Figure 2 demonstrates.
</p>

<p>
In a forthcoming paper, we conduct a quantitative synthesis of the literature: a meta-analysis (Irsova et al. 2018).1 Our intention is to trace the differences in results back to differences in data, methods, and potentially also general study quality. From the 44 studies on the effect of DST on energy consumption, we collect 162 usable estimates. First, we test for publication bias, which typically exagerrates results in empirical economics by a factor of 2 (Ioannidis et al. 2017). We find no publication bias whatsoever, which is a remarkable finding in itself. Yes, our dataset includes many unpublished studies, but in economics publication bias is typically found even in working papers, as many authors routinely discriminate against unintuitive and statistically insignificant results.
</p>

<p>
The lack of publication bias means that we can proceed to the main part of our analysis – an examination of why the reported estimates vary so much. To this end, we regress the estimates of the DST effect on factors associated with the context in which the estimates were obtained. For example, we include the number of maximum daylight hours in the region to which the estimate corresponds. Among other factors, we control for the frequency of data on energy use (hourly or daily), estimation technique (simulation, difference-in-differences, simple regression), definition of energy consumption (commercial, residential, lighting only), and aspects that may be related to quality (journal publication, impact factor of the outlet, number of citations).
</p>

<p>
Note: Factors are sorted from top to bottom by importance. Their combinations (models) are shown in columns sorted from left to right by usefulness. Relative usefulness is represented by column width. Blue color means that the factor contributes to finding less savings from DST.
</p>

<p>
The results show that studies published in more prestigious outlets typically report less energy savings from daylight saving time, and that countries with more summer daylight hours (that is, higher latitude) enjoy more energy savings. The frequency of data and estimation methodology are important as well.
</p>

<p>
Next, for each country in our sample we compute estimates of the DST effect conditional on best practice in the literature. Essentially, using the meta-analysis results we re-compute estimates as if they were all produced by studies using the difference-in-differences approach, hourly data, and were published in outlets with the maximum impact factor. The best-practice estimates are shown in Table 1.
</p>

<p>
Negative estimates mean that for these countries DST reduces energy consumption. For some countries with lower latitude, even in Europe, DST seems to increase the use of energy. But all estimates are statistically insignificant and very small. The mean for the entire dataset is almost precisely zero. Norway is the country with the highest benefits from DST, but even there the effect is only 0.5% during the days when DST applies, and therefore about 0.3% of annual consumption.
</p>

<p>
Daylight saving time affects 1.5 billion people around the world twice a year – sometimes fatally, as shown by studies documenting traffic accidents due to time shifts (Smith 2016). With the finding that daylight saving saves no energy to speak of, the original and still commonly used rationale for the policy falls. Perhaps DST really isn’t such a good nudge; dozens of countries have abandoned it in recent decades. Or perhaps the convenience benefits of longer evening daylight hours prevail over the inconvenience of sleep deprivation and even loss of lives resulting from traffic accidents (and possibly from an increased incidence of heart attacks and depression). Or perhaps a year-round DST would retain most of the benefits while removing the many problems associated with time shifts. We simply don’t know. We still await a study that would systematically compare all the different benefits and costs of daylight saving time.
</p>]]></content:encoded>
    </item>
    <item>
      <title>Headline inflation measures shouldn&#x27;t ignore costs of home ownership</title>
      <link>https://meta-analysis.cz/notes/voxeu-home-ownership/</link>
      <guid isPermaLink="true">https://meta-analysis.cz/notes/voxeu-home-ownership/</guid>
      <pubDate>Tue, 12 Sep 2017 00:00:00 +0000</pubDate>
      <description>Excluding owner-occupied housing costs from the EU&#x27;s harmonised index of consumer prices leaves out what most people experience as inflation; including imputed rents, as the US and Japan already do, would make Eurozone monetary policy more countercyclical.</description>
      <content:encoded><![CDATA[<p class="byline">First published on <a href="https://cepr.org/voxeu/columns/headline-inflation-measures-shouldnt-ignore-costs-home-ownership" rel="external">VoxEU / CEPR</a>, 12 September 2017. Archived with the rest of our writing at <a href="https://meta-analysis.cz/komentare/voxeu-home-ownership/">/komentare/voxeu-home-ownership/</a>.</p>

<p>
Seven out of every ten Europeans live in their own homes, yet Europe’s most important inflation measure excludes the costs associated with owner-occupied housing. This column argues that including the costs of home ownership would prove beneficial to the conduct of monetary and macroprudential policy. It would also bring the measure closer to what most people consider inflation to be.
</p>

<p>
Statistical offices of many countries measure the costs of home ownership by computing imputed rents, which are then included in headline inflation measures. This is the case for the US, Japan, and Switzerland, among others. In contrast, the harmonised index of consumer prices (HICP) – the EU’s most important inflation statistic – excludes owner-occupied housing, for the technical reason that imputed transactions are inconsistent with the definition of the HICP, and a more complex approach based on net acquisitions would be required (Eurostat 2012, 2013).
</p>

<p>
Eurostat has been considering adding the cost of owner-occupied housing to the HICP for many years. It does not, partly because the net acquisitions approach implies that house prices are directly included in the HICP (alongside property charges, and costs of repairs and maintenance). Because house purchases involve a substantial investment component, their inclusion in headline inflation makes many statisticians uneasy. Conceptually, however, homes are a special case of durable goods, because they provide a claim on a stream of future services. Cecchetti (2007), for example, showed the long-term capital gain from home ownership is very small.
</p>

<p>
House prices, of course, are important for financial stability in their own right, and in our recent work we argue that including them in official inflation measures can help integrate monetary and macroprudential policies (Hampl and Havranek (2017). Many economists have constructed early warning systems for financial crises in which house prices play a prominent role (e.g. Reimers 2012, Babecky et al. 2013, Antunes et al. 2014, Laina et al. 2015, Tölö 2015). The prominence of house prices in early warning indicators leads some to stress the interaction between their value and the monetary policy stance. As with many other issues in the recent discussion on macroprudential policy, however, there is no clear consensus.
</p>

<p>
One stream of thought, represented by Assenmacher-Wesche and Gerlach (2010) and Svensson (2014) among others, asserted that it is too costly and detrimental to the welfare of the country to use monetary policy to slow an increase in house prices. Williams (2015) conducted a meta-analysis of the empirical estimates reported in this body of research and found that typically, a 4% reduction in house prices delivered by monetary policy contraction is associated with a 1% loss of GDP.
</p>

<p>
But this discussion often omits the positive effects of this policy on GDP and employment during the downturn, when traditional CPI targeting, taking house prices into account, implies less easing than what would otherwise be optimal. In other words, it is important to highlight that inflation targeting is symmetrical, whatever the composition of the inflation series.
</p>

<p>
Several studies have demonstrated the usefulness of incorporating financial stability considerations (including, most prominently, house prices) into monetary policy rules under inflation targeting. For example, Aydin and Volkan (2011) provided evidence for this. They used a structural model of monetary policy calibrated for South Korea, and found that paying attention to house prices creates smoother business cycle fluctuations than conventional inflation targeting.
</p>

<p>
House prices are typically excluded from official inflation measures, although other goods that also provide a flow of future services (durables such as motor vehicles and washing machines) are included. There is no clear theoretical reason beyond intuition and convenience for this convention. The argument in favour is that for houses, the investment component relative to the consumption component is larger than for durables such as cars. Also, a portion of the value, such as land, does not depreciate, and is therefore often considered a good store of value.
</p>

<p>
Anecdotal evidence, however, suggests that many households treat at least their first home purchase more as consumption than investment. And theoretically, the prices of all assets, including houses, stocks, and bonds, should in principle be included in inflation if we are to measure the current cost of expected lifetime consumption, instead of merely current consumption (Alchian and Klein 1973).
</p>

<p>
Aside from the well-known studies by Alchian and Klein (1973) and Goodhart (2001), many other authors have argued for the inclusion of house prices in the consumer price index. For example, Bryan et al. (2002) showed that, in the US, the omission of house prices introduces an excluded goods bias, and results in underestimation of CPI by about 0.25 percentage points annually. Diewert and Nakamura (2009) also pointed to the need for a more direct measure of house price inflation in the official CPI index. They suggested that the recent period of low official inflation may be a mismeasurement of underlying consumer prices.
</p>

<p>
Figure 1 shows Eurozone quarterly year-on-year changes in the HICP, an index of owner-occupied housing consistent with the HICP, and a pure house price index. The growth in the latter two indices was below official inflation between 2011 and 2014, and has exceeded official inflation since 2015. It follows that including the cost of home ownership in the HICP would make the monetary policy of the ECB, if anything, more countercyclical. The extent of this effect depends on the weight attributed to the owner-occupied housing index or the house price index, but even a weight of 10% would mean a difference in the HICP of up to half a percentage point in some periods.
</p>

<p>
Figure 1 Giving non-zero weight to house prices would make monetary policy in the Eurozone more countercyclical
</p>

<p>
Notes: Aggregate index of owner-occupied housing for the Eurozone computed using the weights for each country in Eurostat’s construction of the harmonized index of consumer prices. Source: Eurostat.
</p>

<p>
The delay in data availability is a frequent argument against the inclusion of house prices. This is a problem, but one that has been overcome by several statistical offices (Hampl and Havranek 2017). For example, Czech headline monthly inflation includes house prices with a 1.4% weight for most regions, and with a 2.3% weight for Prague, the capital city. The Czech National Bank, unusually for a central bank, computes its own supplementary inflation index (the CPIH) in which house prices get a 15% weight, based on the share in consumers’ expenditure. The Bank makes this index available in its inflation report. In some countries, it is possible to take the data directly from the land registry, where all price information is available within a few days after property changes hands.
</p>

<p>
Among the many arguments for including the costs of home ownership in headline CPI, a prominent one is that it would bring the CPI index closer to what most people consider inflation to be. In a well-known, colourfully titled paper, “Measuring inflation: the core is rotten", James Bullard, the President of the Federal Reserve Bank of St. Louis, criticised the Federal Reserve’s focus on core inflation and argued that we should pay more attention to a broader gauge. To paraphrase Bullard’s (2011) provocative statement: an immediate benefit of moving away from the sole emphasis on an inflation measure that excludes the costs of home ownership would be to reconnect central banks and statistical bureaus with households and businesses who know price changes when they see them.
</p>]]></content:encoded>
    </item>
  </channel>
</rss>
