Creation and evaluation have different jobs
A maker agent is rewarded for producing a coherent result. It chooses an approach, makes tradeoffs, and works around obstacles. Those same choices can become blind spots when the agent grades itself. An evaluator should begin from the accepted specification and required proof, then look for missing behavior, unsupported claims, security or accessibility concerns, incomplete tests, and conditions the maker did not exercise.
Freeze the behavioral contract before implementation
Product intent becomes an executable constraint before implementation begins. For Agile, Kanban, and SAFe teams that means the Product Spec, business rules, acceptance criteria, and known failure conditions become tests first. The tests start as executable examples of expected behavior. They join the Living Working Spec on the card. They become the boundary the implementation agent works inside.
The testing workflow therefore runs Product Spec, then executable behavioral tests, then implementation. End-to-end scenarios are the primary behavioral contract for AI-assisted delivery. Treat the application, or an important subsystem, as a black box. Test what goes in and what observable behavior comes out.
- Run a small set of critical end-to-end scenarios on every change.
- Run broader scenarios less frequently.
- Keep traces, screenshots, logs, API responses, and videos on the card.
- Use unit tests when critical logic needs isolation, or when an edge case is too expensive to exercise end to end.
The implementation adapts. The spec stays.
Freeze the behavioral contract before implementation where practical. The agent or person deriving the tests is different from the agent implementing the solution. When the implementation fails a test, the implementation changes. The accepted spec stays until a human changes it.
A production defect first asks which product requirement, business rule, acceptance condition, or failure mode was missing upstream. Strengthen the specification. Encode the missing behavior. Then fix the implementation. The same principle applies to security, performance, concurrency, resilience, privacy, and resource usage.
Testing cannot prove the absence of every possible defect. Properties important enough to govern release should be explicit and repeatably verifiable. Encode the governance into AGENTS.md, repository instructions, agent prompts, CI policies, Product Spec templates, testing conventions, and review rules. The stack is simple: Product Spec defines intent, tests encode executable constraints, and agent instructions enforce the workflow.
What independent means in practice
- A different agent role with an evaluator prompt and permission set.
- Evaluation against the accepted spec and proof profile, not only the maker summary.
- Direct inspection of the diff, artifacts, tests, logs, screenshots, or connector results where permitted.
- Permission to return inconclusive, request changes, or identify an untested condition.
- No authority to turn its own favorable review into human acceptance or merge approval.
Using a different model can add diversity, but model difference alone does not create independence. If both agents receive the same flawed assumption, hide the same missing evidence, or optimize for the same success signal, they can agree and still be wrong.
Use a proof profile that matches the risk
| Work type | Useful automated proof | Human verification still needed |
|---|---|---|
| Documentation change | Link checks, examples, terminology, build, and rendered page | Accuracy, audience fit, confidentiality, and publication decision |
| Application change | Unit, integration, accessibility, security, and browser evidence | Customer behavior, tradeoffs, usability, and merge decision |
| Infrastructure change | Plan validation, policy checks, dry run, drift, and rollback evidence | Blast radius, timing, authorization, and production decision |
| Customer communication | Required fields, policy checks, link validation, and approved template | Meaning, tone, recipient suitability, consent, and send decision |
Record inconclusive proof honestly
A trustworthy evaluator must be allowed to say that it could not run a test, access an environment, reproduce an issue, or verify an outcome. The interface should not translate "expected to pass" into "passed." The human reviewer can request more proof, accept a disclosed risk if authorized, or stop the work. That decision and its reason should remain visible.
Keep repository review controls in force
When the output is code, agent QA complements rather than replaces pull request review, required status checks, code owners, and branch protection. The governed work record should link to the commit and pull request, while GitHub remains the authority for repository review and merge controls.
Key terms
- Maker agent
- The agent role that plans or produces the work.
- Evaluator agent
- A separate agent role that checks the result against accepted requirements and proof without accepting the work on behalf of a human.
- Behavioral contract
- Executable tests derived from the Product Spec, business rules, acceptance criteria, and known failure conditions before implementation begins.
- Inconclusive
- A verification result that cannot support pass or fail because required evidence or access was unavailable.
Sources and standards context
- NIST AI Risk Management Framework CoreNIST describes governance as a continuous, cross-cutting function and calls for documented human oversight, roles, testing, and accountability across the AI lifecycle.
- GitHub documentation on pull request reviewsGitHub documents comment, approval, request-changes, required-review, and traceable discussion patterns for proposed code changes.
