Skip to main content
← Back to list
01Issue
BugShippedSwamp CLIPublic
Assigneesskunk-ape

Relationships

#2587 verify-reviews: code review fails on format, not substance; submit verdicts through a tool that enforces them

Opened by skunk-ape · 9/28/2026· Shipped 9/28/2026

The pre-PR agent reviews run claude -p with a prompt ending: Output your review. Start with a single line: VERDICT: pass or VERDICT: fail. scripts/check_review_verdict.ts passes a review only when its opening line is exactly that marker (deliberately: a marker further down is quoted, never trusted).

The code review often opens with a sentence first, so the step fails with verdict missing even when the review itself says VERDICT: pass with no blocking issues. Observed openings:

  • I've read the whole diff and I'm writing up the review now.
  • I couldn't run the tests here (the command needed approval), so this is a read-through only. Writing up the review now.

Frequency: 3 of 7 code-review runs across PR 327 (swamp-club #2576) and the GW-3 branch (swamp-club #2584). The adversarial and CI-security reviews, which end with the same instruction, never did it. Each false failure costs verification_failed, a new commit (the attestation is bound to the SHA) and a full re-run of both workflows.

Likely cause (inferred): the code-review prompt has a Testing Rules section that nudges the reviewer to run tests; reviewers only have Read, Glob and Grep, the command is denied, and the reviewer narrates that before the verdict. The format instruction is the last line of a long prompt.

Proposed fix: make the verdict structured and deterministic instead of parsed from prose. The reviewer submits its result through a tool whose input schema enforces the contract, for example a swamp model method or a small script (submit_review) taking a verdict enum (pass|fail), findings with severities, and the review text. It is the reviewer's only way to record a result, so a missing or malformed verdict is rejected at submission with an error the agent can act on, and the gate reads the stored record rather than scanning the first line of free text. The allowlist would add exactly that tool (no general shell), and the tool itself becomes a pinned trust-root file.

Cheaper interim options: move the no-commands and verdict-first instructions to the top of the code-review prompt; or retry once on a missing verdict. Relaxing the gate to accept a verdict after leading prose is not recommended.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 8 MOREREVIEW+ 7 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

9/28/2026, 6:54:25 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
skunk-ape assigned skunk-ape9/28/2026, 5:40:12 PM

Sign in to post a ripple.