Agentic AI

Post

When one model checks another: multi-model review becomes product

Source: Microsoft, DRACO benchmark for deep research quality

AI models reviewing each other is becoming a reality, and it deserves close attention. Microsoft has announced Critique, a multi-model deep research capability for Microsoft 365 Copilot Researcher. One model generates the initial research and draft. A second reviews it to challenge assumptions, identify gaps, improve accuracy and strengthen citations.

In many of my AI strategy sessions we have discussed exactly this: checking AI output with another model rather than relying on a single response. Seeing it operationalised directly into a product is a significant moment for agentic workflow design. On the deep-research benchmark Microsoft published alongside the announcement, the two-model configuration scored ahead of every single-model system it was compared against.

For anyone designing analytical or advisory workflows, the multi-model pattern is going to be central to producing output that is robust, accurate and trustworthy enough to put in front of a decision-maker.

Have a question these don’t answer?

For advisory, training and speaking enquiries, or a conversation about a decision you are weighing.