Skip to content

[Feature]: Reviewer agents lack an explicit verification discipline (verify-before-pass, pre-verdict checklist) #1134

Description

@vnlebaoduy

Description

The two review-only agents (aidlc-architecture-reviewer-agent, aidlc-product-lead-agent) carry a strong ## Adversarial Posture (refute, ground every finding in checkable evidence) and a ## Turn Budget, but no explicit verification discipline: how the reviewer must behave between "assume it is broken" and "write the verdict". A grep of core/agents/ finds no pre-verdict checklist in any of the 14 agent personas.

Concretely, the current prose leaves these behaviors implicit:

  • Assume-AI-authored posture. Stage artifacts are increasingly produced by the lead sub-agent itself. Polished tables and confident prose are not evidence that an ID resolves or that a criterion is testable, but nothing tells the reviewer to discount presentation.
  • Verify-before-pass, not only verify-before-flag. The posture text grounds findings in evidence; it does not say that an unread section cannot count toward READY. Under the turn budget this is the failure mode that silently degrades a review: "ran out of turns" becomes "no findings" becomes READY.
  • Competing readings before declaring a gap. "Missing" vs "lives under another ID" vs "already pinned by a passed contract" - the reviewer should rule the alternatives out with a lookup before charging the author with an omission (each false omission costs a revision round).
  • Severity ranking. Findings are not required to be ordered by what blocks a developer / fails in production vs what is merely unclear.
  • A pre-verdict checklist the reviewer confirms line by line (every NOT-READY finding has an anchor, every cross-reference resolved or flagged, validation-tool output cited, no sibling-unit read, unverified concerns recorded as questions).

Use Case

Reviewer sub-agents run unattended inside the adversarial loop (Construction) and as advisory decision support at human gates (ideation/inception). In both classes the human reads only the findings. A reviewer that passes unread material, or reports a "gap" that was a lookup it owed the author, wastes a review iteration or a revision round, and the audit trail records a READY that was never earned.

Proposal

Add a ## Verification Discipline section to both reviewer personas, placed between ## Advisory Dispatch and ## Key Principles, voiced in each agent's domain (systems/blast-radius for the architecture reviewer; customer/testability for the product lead). Prose only: no frontmatter change, no protocol change, the pinned ## Adversarial Posture / ## Output Contract / ## Turn Budget text (t234, t279) untouched, harness-neutral (t146).

If this shape is acceptable, I have a PR ready and would follow with the same treatment for the producing personas (developer, quality, devsecops, then product/design/delivery/operations) in small separate PRs.

Version

2.8.2 (main @ c0eb292)

Area

Agents (core/agents/)

Additional Context

Related but orthogonal: #1059 (weighted rubric / measured gate) adds a scoring layer on top of the verdict; this issue only sharpens how the reviewer arrives at the findings that feed any such layer.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions