KomentářeTomáš Havránek, Zuzana Iršová Havránková

Results of the AI report ranking experiment

Correspondence — not a published text.

A letter to the participants, not a published text, which is why it is filed as correspondence. It reports the results of "Does Multi-Agent Debate Improve AI Feedback on Research Papers?" on the day the preprint appeared. Everything it points to is public: the preprint, CEPR Discussion Paper 21752, the OSF pre-registration, the Zenodo replication package and both tools on GitHub. The text is reproduced unchanged; only the recipient list and the attached PDF are omitted.

Dear colleague,

Thank you again for reading the AI reports on your paper and ranking them. Our experiment exists only because 47 authors were willing to do it! The paper is now out.

What we found:

  • A report produced by a single prompt using one frontier model beat both elaborate multi-agent tools, even though one of the tools spent about 30 times the tokens. In our experiment, multi-agent debate did not help.
  • Had an external AI model ranked the reports in your place, the most elaborate tool would have come first.
  • Authors who recalled their real journal referee report usually ranked it above all the AI reports. In contrast, the AI judges almost always ranked that same human report last.
  • Author rankings and the external AI model's rankings agree only weakly (correlation 0.14). Several of you told us the ranking took a lot of time, and we are grateful!

The two multi-agent tools from the experiment are open source: mad-research runs a cross-model adversarial audit, and paper-workshop runs a Claude-only expert workshop. They didn't beat a single prompt in our experiment, but both still produce detailed comments that can be useful as you revise your paper.

Links:

Havranek, T and Z Irsova (2026), "Does Multi-Agent Debate Improve AI Feedback on Research Papers?", CEPR Discussion Paper No. 21752. CEPR Press, Paris & London.

Stay well,

Tomas Havranek and Zuzana Irsova

Correspondence, not a published text. Sent 16 July 2026. Sent as a blind copy to the authors of 44 meta-analyses who had ranked three blinded AI reports on their own paper; the recipient list is not reproduced here. Project page and paper.