Claude Code can call Codex to stress-test its own work

Do you know that Claude Code can use Codex to stress test its work?

I had both installed and used them interchangeably: for some tasks, one seems to be better, and vice versa. Also, sometimes I hit usage limit in one, so I continue in the other.

This can be automated easily: work from Claude Code and call Codex when needed. Here is the Claude skill, applied to stress-testing research:

https://github.com/tjhavranek/mad-research

This builds on our previous manual protocols for using different AI models to test research ideas or papers:

https://github.com/tjhavranek/research-audit-duel-protocol

Does the skill work for you? What should we change? Should we add Gemini?

Screenshot showing Claude Code invoking Codex to stress-test its own work, as described in the post.