← Back to articlesGoverned agents

Why the agent that writes the work should not be its only evaluator

Freeze the behavioral contract from the Product Spec before the maker begins. Separate the maker from the evaluator, then keep final acceptance with an accountable human.

Separation of duties workflow with Maker coding agent, independent QA evaluator agent, and final accountable human sign-off

Creation and evaluation have different jobs

A maker agent is rewarded for producing a coherent result. It chooses an approach, makes tradeoffs, and works around obstacles. Those same choices can become blind spots when the agent grades itself. An evaluator should begin from the accepted specification and required proof, then look for missing behavior, unsupported claims, security or accessibility concerns, incomplete tests, and conditions the maker did not exercise.

Freeze the behavioral contract before implementation

Product intent becomes an executable constraint before implementation begins. For Agile, Kanban, and SAFe teams that means the Product Spec, business rules, acceptance criteria, and known failure conditions become tests first. The tests start as executable examples of expected behavior. They join the Living Working Spec on the card. They become the boundary the implementation agent works inside.

The testing workflow therefore runs Product Spec, then executable behavioral tests, then implementation. End-to-end scenarios are the primary behavioral contract for AI-assisted delivery. Treat the application, or an important subsystem, as a black box. Test what goes in and what observable behavior comes out.

  • Run a small set of critical end-to-end scenarios on every change.
  • Run broader scenarios less frequently.
  • Keep traces, screenshots, logs, API responses, and videos on the card.
  • Use unit tests when critical logic needs isolation, or when an edge case is too expensive to exercise end to end.

The implementation adapts. The spec stays.

Freeze the behavioral contract before implementation where practical. The agent or person deriving the tests is different from the agent implementing the solution. When the implementation fails a test, the implementation changes. The accepted spec stays until a human changes it.

A production defect first asks which product requirement, business rule, acceptance condition, or failure mode was missing upstream. Strengthen the specification. Encode the missing behavior. Then fix the implementation. The same principle applies to security, performance, concurrency, resilience, privacy, and resource usage.

Testing cannot prove the absence of every possible defect. Properties important enough to govern release should be explicit and repeatably verifiable. Encode the governance into AGENTS.md, repository instructions, agent prompts, CI policies, Product Spec templates, testing conventions, and review rules. The stack is simple: Product Spec defines intent, tests encode executable constraints, and agent instructions enforce the workflow.

What independent means in practice

  • A different agent role with an evaluator prompt and permission set.
  • Evaluation against the accepted spec and proof profile, not only the maker summary.
  • Direct inspection of the diff, artifacts, tests, logs, screenshots, or connector results where permitted.
  • Permission to return inconclusive, request changes, or identify an untested condition.
  • No authority to turn its own favorable review into human acceptance or merge approval.

Using a different model can add diversity, but model difference alone does not create independence. If both agents receive the same flawed assumption, hide the same missing evidence, or optimize for the same success signal, they can agree and still be wrong.

Use a proof profile that matches the risk

QA verification results showing passed, failed, and inconclusive evidence on a governed card
Work typeUseful automated proofHuman verification still needed
Documentation changeLink checks, examples, terminology, build, and rendered pageAccuracy, audience fit, confidentiality, and publication decision
Application changeUnit, integration, accessibility, security, and browser evidenceCustomer behavior, tradeoffs, usability, and merge decision
Infrastructure changePlan validation, policy checks, dry run, drift, and rollback evidenceBlast radius, timing, authorization, and production decision
Customer communicationRequired fields, policy checks, link validation, and approved templateMeaning, tone, recipient suitability, consent, and send decision

Record inconclusive proof honestly

A trustworthy evaluator must be allowed to say that it could not run a test, access an environment, reproduce an issue, or verify an outcome. The interface should not translate "expected to pass" into "passed." The human reviewer can request more proof, accept a disclosed risk if authorized, or stop the work. That decision and its reason should remain visible.

Keep repository review controls in force

When the output is code, agent QA complements rather than replaces pull request review, required status checks, code owners, and branch protection. The governed work record should link to the commit and pull request, while GitHub remains the authority for repository review and merge controls.

Key terms

Maker agent
The agent role that plans or produces the work.
Evaluator agent
A separate agent role that checks the result against accepted requirements and proof without accepting the work on behalf of a human.
Behavioral contract
Executable tests derived from the Product Spec, business rules, acceptance criteria, and known failure conditions before implementation begins.
Inconclusive
A verification result that cannot support pass or fail because required evidence or access was unavailable.

Sources and standards context