Does multi-agent debate improve AI feedback on research papers?
Does multi-agent debate improve AI feedback on research papers?
Not in our experiment.
We just posted a pre-registered study in which authors ranked three AI reports on their own paper. The reports were blinded and came from three setups: a single prompt and two multi-agent tools. We expected the multi-agent tools to win.
We find:
1) Multi-agent debate does not help here. The single prompt beat both multi-agent tools, even though one of them spent about 30x the tokens.
2) If an independent AI model ranked the reports in the authors' place, it would put the most expensive multi-agent tool first.
3) Authors who recalled their real journal referee feedback usually ranked it above all the AI reports. In contrast, the AI judges almost always ranked the human feedback last.
I was also surprised by how little the author and AI rankings agree (correlation 0.14). Authors seem to have actually read the AI reports (!) and thought about them, not just used a chatbot to rank them.
Thanks to all 47 authors who participated!
Paper: https://lnkd.in/d7YsNxWc
Both multi-agent tools are open source:
https://lnkd.in/dtGSi8zP
https://lnkd.in/dk5NQpcE
Links the author added in the comments: