Can AI review its own code?

When your reviewer agent and your developer agent use the same model, when they are given the same context and operate with the same reasoning process, you haven’t added a second opinion. In fact, you’ve just run the first one twice.

A reviewer agent using the same model and context is likely to have the same blind spots the developer agent had, make similar assumptions and validate the same choices. Similar reasoning processes, when faced with similar problems, tend to produce unsurprisingly similar results. Going over the same ground with a slightly finer comb may improve the output, but it’s unlikely to catch real issues or spot the flawed logic behind decisions.

When a reviewer adds value

Here are three things that make the biggest difference to reviewer agents: 

  1. Different objectives. A developer agent’s main task is to build a working solution. A reviewer agent, on the other hand, should be tasked with assuming that solution is flawed. Then it will actively search for security vulnerabilities, performance bottlenecks, architectural issues and defects, not validating the logic that produced the code. Without the context of how decisions were made, the reviewer is not invested in the approach, which makes it far more likely to question it. 
  1. Different tools. A reviewer that only uses language model reasoning is still limited to reasoning. The more powerful version of a reviewer agent includes external verification: it’s able to run unit and integration tests, execute the compiler, inspect static analysis report, analyse code coverage and review security scan results. This objective data can surface things that no amount of reasoning, however good, can find on its own. 
  1. Different models. You can introduce diversity in reasoning and failure modes by using a different LLM. Different models have different strengths, training emphases and tendencies. And,importantly, they have different blind spots. Where a single model repeatedly misses something, combining two different ones meaningfully increases the chance of catching it. 

What we got wrong first 

We learned this by doing. When we first built the reviewer agent into Damilah’s Multi-Agent Platform (DMAP , the review artefact did not include acceptance criteria status. That meant we had no visibility into whether the criteria added by the Product Manager had been met, only whether the code had been written.

Essentially, we had built a reviewer that reviewed code, when what we needed was one that reviewed outcomes.

Once we spotted the gap, we added acceptance criteria to the artefact and included UI tests and screenshots of implemented features.

And we are still iterating: refining what tests we run, how we execute them and how to genuinely verify that the outcome matches the goal of the workflow against the quality standards set for each project.

“We are still improving how we review. Each project shows us something the previous one did not. That is just how building this kind of system works.”
– Andrea Stankovikj, Principal Product Owner at Damilah 

The bottom line

To become a real safety net, a reviewer agent needs to be built carefully and intentionally. With a clearly defined objective that differs from that of a developer, with access to the necessary external tools and, where it matters, running on a different model.

The strongest case for a multiple-agent system, along with parallel execution and cost routing, is task specialisation combined with access to different information, different goals and objective external verification.

The most practical architecture therefore seems to be a hybrid: a central orchestration agent coordinating specialised execution agents that use tooling, rather than relying on reasoning alone.These are the principles at the foundation of DMAP.

If you are thinking about how to build governance and quality review into your AI development workflows, book a meeting with us to discuss it further.