---
title: "Verify"
description: "Grade work that already exists and decide whether it can merge. Runs the project's current unit, integration, and E2E suites plus security scanning and type checking, scores every dimension 0-10, and returns a merge verdict with a VERIFIED-vs-CLAIMED evidence manifest. Writes no test files and edits no source. Use when verifying changes are ready to merge. Use /ork:cover instead when the tests still have to be written."
canonical: "https://orchestkit.yonyon.ai/docs/reference/skills/verify"
---

# Verify

Grade work that already exists and decide whether it can merge. Runs the project's current unit, integration, and E2E suites plus security scanning and type checking, scores every dimension 0-10, and returns a merge verdict with a VERIFIED-vs-CLAIMED evidence manifest. Writes no test files and edits no source. Use when verifying changes are ready to merge. Use /ork:cover instead when the tests still have to be written.

<span className="badge badge-blue">Command</span> <span className="badge badge-orange">high</span>

```bash title="Invoke"
/ork:verify
```

<ContextualSkillSidebar slug="verify" />

> **Verify** Grade work that already exists and decide whether it can merge. Runs the project's current unit, integration, and E2E suites plus security scanning and type checking, scores every dimension 0-10, and returns a merge verdict with a VERIFIED-vs-CLAIMED evidence manifest. Writes no test files and edits no source. Use when verifying changes are ready to merge.


# Verify Feature

Comprehensive verification using parallel specialized agents with nuanced grading (0-10 scale) and improvement suggestions.

## Quick Start

```bash
/ork:verify authentication flow
/ork:verify --model=opus user profile feature
/ork:verify --scope=backend database migrations
```

## Argument Resolution

```python
SCOPE = "$ARGUMENTS"       # Full argument string, e.g., "authentication flow"
SCOPE_TOKEN = "$ARGUMENTS[0]"  # First token for flag detection (e.g., "--scope=backend")
# $ARGUMENTS[0], $ARGUMENTS[1] etc. for indexed access (CC 2.1.59)

# Model override detection (CC 2.1.72)
MODEL_OVERRIDE = None
for token in "$ARGUMENTS".split():
    if token.startswith("--model="):
        MODEL_OVERRIDE = token.split("=", 1)[1]  # "opus", "sonnet", "haiku", "fable"
        SCOPE = SCOPE.replace(token, "").strip()

# Streak gate detection (#2540) — consecutive-pass mode
STREAK_TARGET = None
for token in "$ARGUMENTS".split():
    if token.startswith("--streak="):
        STREAK_TARGET = int(token.split("=", 1)[1])  # N consecutive READY verdicts required (N >= 2)
        SCOPE = SCOPE.replace(token, "").strip()
# When set, apply the Streak Gate (see below). Full protocol: references/streak-gate.md
```

Pass `MODEL_OVERRIDE` to all Agent() calls via `model=MODEL_OVERRIDE` when set. Accepts symbolic names (`opus`, `sonnet`, `haiku`, `fable` on harnesses whose Agent tool lists it; note fable is premium API spend after 2026-07-12) or full IDs (`claude-opus-5`) per CC 2.1.74.

> **Opus 5**: Agents use native adaptive thinking (no MCP sequential-thinking needed); defaults to `high` effort (CC 2.1.154+). Extended 128K output supports comprehensive verification reports.

---

## STEP 0: Effort-Aware Verification Scaling (CC 2.1.76)

Scale verification depth based on `/effort` level:

| Effort Level | Phases Run | Agents | Output |
|-------------|------------|--------|--------|
| **low** | Run tests only → pass/fail | 0 agents | Quick check |
| **medium** | Tests + code quality + security | 3 agents | Score + top issues |
| **high** (default) | All 8 phases + visual capture | 6-7 agents | Full report + grades |
| **xhigh** (Opus 5, CC 2.1.111+) | All 8 phases + additional cross-file pattern sweep + self-verification pass | 6-7 agents | Full report with uncertainty annotations |

> **Override:** Explicit user selection (e.g., "Full verification") overrides `/effort` downscaling.

## STEP 0a: Verify User Intent with AskUserQuestion

**BEFORE creating tasks**, clarify verification scope:

```python
AskUserQuestion(
  questions=[{
    "question": "What scope for this verification?",
    "header": "Scope",
    "options": [
      # multiSelect questions do not render previews (single-select only) — kept text-only
      {"label": "Full verification (Recommended)", "description": "All tests + security + code quality + visual + grades"},
      {"label": "Tests only", "description": "Run unit + integration + e2e tests"},
      {"label": "Security & code quality", "description": "Security audit (OWASP/CVE/secrets) + lint/types/complexity"},
      {"label": "Quick check", "description": "Just run tests, skip detailed analysis"}
    ],
    "multiSelect": true
  }]
)
```

**Based on answer, adjust workflow:**
- **Full verification**: All 9 phases (8 + 2.5), 7 parallel agents including visual capture
- **Tests only**: Skip phases 2 (security), 5 (UI/UX analysis)
- **Security & code quality**: Run security-auditor + code-quality-reviewer agents
- **Quick check**: Run tests only, skip grading and suggestions

---

## STEP 0b: Select Orchestration Mode

Load details: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/orchestration-mode.md")` for env var check logic, Agent Teams vs Task Tool comparison, and mode selection rules.

Choose **Agent Teams** (mesh -- verifiers share findings) or **Task tool** (star -- all report to lead) based on the orchestration mode reference.

---

### MCP Probe + Resume

```python
# memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541) — probe below kept as fallback for older CC:
ToolSearch(query="select:mcp__memory__search_nodes")
Write(".claude/chain/capabilities.json", { memory, timestamp })

Read(".claude/chain/state.json")  # resume if exists
```

### Handoff File

After verification completes, write results:

```python
Write(".claude/chain/verify-results.json", JSON.stringify({
  "phase": "verify", "skill": "verify",
  "timestamp": now(), "status": "completed",
  "outputs": {
    "tests_passed": N, "tests_failed": N,
    "coverage": "87%", "security_scan": "clean"
  }
}))
```

### Regression Monitor (CC 2.1.71)

Optionally schedule post-verification monitoring:

```python
# Guard: Skip cron in headless/CI (CLAUDE_CODE_DISABLE_CRON)
# if env CLAUDE_CODE_DISABLE_CRON is set, run a single check instead
CronCreate(
  schedule="0 8 * * *",
  prompt="Daily regression check: npm test.
    If 7 consecutive passes → CronDelete.
    If failures → alert with details."
)
```

---

## Task Management (CC 2.1.16)

```python
# 1. Create main verification task
TaskCreate(
  subject="Verify [feature-name] implementation",
  description="Comprehensive verification with nuanced grading",
  activeForm="Verifying [feature-name] implementation"
)

# 2. Create subtasks for 8-phase process
TaskCreate(subject="Run code quality checks", activeForm="Running quality checks")    # id=2
TaskCreate(subject="Execute security audit", activeForm="Running security audit")     # id=3
TaskCreate(subject="Verify test coverage", activeForm="Verifying test coverage")      # id=4
TaskCreate(subject="Validate API", activeForm="Validating API")                       # id=5
TaskCreate(subject="Check UI/UX", activeForm="Checking UI/UX")                       # id=6
TaskCreate(subject="Calculate grades", activeForm="Calculating grades")               # id=7
TaskCreate(subject="Generate suggestions", activeForm="Generating suggestions")       # id=8
TaskCreate(subject="Compile report", activeForm="Compiling report")                   # id=9

# 3. Set dependencies — phases 2-6 run in parallel, 7-9 are sequential
TaskUpdate(taskId="7", addBlockedBy=["2", "3", "4", "5", "6"])  # Grading needs all checks
TaskUpdate(taskId="8", addBlockedBy=["7"])  # Suggestions need grades
TaskUpdate(taskId="9", addBlockedBy=["8"])  # Report needs suggestions

# 4. Update status as you progress
TaskUpdate(taskId="2", status="in_progress")  # When starting
TaskUpdate(taskId="2", status="completed")    # When done — repeat for each subtask
```

---

## 8-Phase Workflow

Load details: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/verification-phases.md")` for complete phase details, agent spawn definitions, Agent Teams alternative, and team teardown.

| Phase | Activities | Output |
|-------|------------|--------|
| **1. Context Gathering** | Git diff, commit history | Changes summary |
| **2. Parallel Agent Dispatch** | 6 agents evaluate | 0-10 scores |
| **2.5 Visual Capture** | Screenshot routes, AI vision eval | Gallery + visual score |
| **3. Test Execution** | Backend + frontend tests | Coverage data |
| **4. Nuanced Grading** | Composite score calculation | Grade (A-F) |
| **5. Improvement Suggestions** | Effort vs impact analysis | Prioritized list |
| **6. Alternative Comparison** | Compare approaches (optional) | Recommendation |
| **7. Metrics Tracking** | Trend analysis | Historical data |
| **8. Report Compilation** | Evidence artifacts + gallery.html | Final report |

### Phase 2 Agents (Quick Reference)

| Agent | Focus | Output |
|-------|-------|--------|
| code-quality-reviewer | Lint, types, patterns | Quality 0-10 |
| security-auditor | OWASP, secrets, CVEs | Security 0-10 |
| test-generator | Coverage, test quality | Coverage 0-10 |
| backend-system-architect | API design, async | API 0-10 |
| frontend-ui-developer | React 19, Zod, a11y | UI 0-10 |
| python-performance-engineer | Latency, resources, scaling | Performance 0-10 |

Launch ALL agents in ONE message with `run_in_background=True` and `max_turns=25`.

### Progressive Output (CC 2.1.76+)

Output each agent's score **as soon as it completes** — don't wait for all 6-7 agents.

> **Focus mode (CC 2.1.101):** In focus mode, include the full composite score, all dimension scores, and the verdict in your final message — the user didn't see the incremental outputs.

```
Security:     8.2/10 — No critical vulnerabilities found
Code Quality: 7.5/10 — 3 complexity hotspots identified
[...remaining agents still running...]
```

This gives users real-time visibility into multi-agent verification. If any dimension scores below the `security_minimum` threshold (default 5.0), flag it as a **blocker immediately** — the user can terminate early without waiting for remaining agents.

### Monitor + Partial Results (CC 2.1.98)

Use `Monitor` for streaming test output. **A `run_in_background` request may not be honoured, so never wait unconditionally on the result.**

```python
task = Bash(command="npm test 2>&1", run_in_background=true)
if not task.id:                 # request ignored → output already returned inline
    use_inline_output(task)     # do NOT wait; there is no task to wait for
else:
    Monitor(pid=task.id)        # bounded: see the contract reference below
    # Still empty AND no live process → the run never happened.
    # Report NO VERDICT as a FAILURE. Never emit a grade from an empty run.
```

Measured (#3263): three backgrounded suites wrote **0 bytes**, npm's own banner never appeared, and the skill waited ~40 min on a completion signal that could not fire. An empty run must be loud, not pending. Contract and the refuted hypotheses: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/background-task-contract.md")`.

Full pattern reference (when to use vs. `TaskOutput`, until-condition gates, anti-patterns): `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/chain-patterns/references/monitor-patterns.md")`.

**Partial results (CC 2.1.98):** If a verification agent fails mid-analysis, synthesize partial scores rather than re-spawning:

```python
for agent_result in verification_results:
    if "[PARTIAL RESULT]" in agent_result.output: A `maxTurns` stop is also partial since CC 2.1.246 (summary: "stopped at its N-turn limit (partial result; continue it with SendMessage to the task-id)"); continue that agent with `SendMessage` instead of re-spawning it.
        # Extract whatever scores the agent produced before crashing
        partial_score = parse_score(agent_result.output)  # May be incomplete
        scores[agent_result.dimension] = {
            "score": partial_score, "partial": True,
            "note": "Agent crashed — score based on partial analysis"
        }
        # A 4-dimension score is better than no score. Do NOT re-spawn.
```

### Phase 2.5: Visual Capture (NEW — runs in parallel with Phase 2)

Load details: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/visual-capture.md")` for auto-detection, route discovery, screenshot capture, and AI vision evaluation.

**Summary**: Auto-detects project framework, starts dev server, discovers routes, uses agent-browser to screenshot each route, evaluates with Claude vision, generates self-contained `gallery.html` with base64-embedded images.

**Output**: `verification-output/\{timestamp\}/gallery.html` — open in browser to see all screenshots with AI evaluations, scores, and annotation diffs.

**Graceful degradation**: If no frontend detected or server won't start, skips visual capture with a warning — never blocks verification.

---

## Grading & Scoring

Load `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/quality-gates/references/unified-scoring-framework.md")` for dimensions, weights, grade thresholds, and improvement prioritization. Load `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/quality-model.md")` for verify-specific extensions (Visual dimension). Load `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/grading-rubric.md")` for per-agent scoring criteria.

### Dimension-Level Blockers (ork-rubric/1.0)

Composite is necessary but not sufficient — a strong composite can average away a critical dimension. In Phase 4 (Nuanced Grading), read per-dimension thresholds from `$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/rubric.json` (schema: `$\{CLAUDE_PLUGIN_ROOT\}/shared/rubric.schema.json`): security `min_blocker` 4.0, compliance `min_pass` 6.0.

- **ANY dimension below its `min_blocker` → verdict is BLOCKED regardless of composite.** Report it explicitly: `Security 3.2/10 (CRITICAL BLOCKER — below min_blocker 4.0)`.
- A dimension below its `min_pass` (but at/above `min_blocker`) caps the verdict at IMPROVEMENTS RECOMMENDED — it cannot grade READY FOR MERGE.
- Blocked verdicts list every tripped dimension first, each with the fix needed to clear it.
- A project `.claude/policies/verification-policy.json` (see Policy-as-Code) may tighten these thresholds, never loosen them below the rubric defaults.

Threshold bands and reporting format: `references/grading-rubric.md` ("Dimension-Level Blockers" section).

---

## Streak Gate (consecutive-pass mode)

A single green is not proof — flaky and order-dependent suites pass once and fail the next run. With `--streak=N`, verify declares **READY FOR MERGE only after N consecutive passing runs**, resetting the count to 0 on any non-ready verdict. The count persists across independent runs in `.claude/chain/verify-streak.json`, keyed by scope.

- `--streak=N` (N ≥ 2; 3 is the sensible default). Absent ⇒ today's single pass/fail behavior, unchanged. Target may also come from `.claude/policies/verification-policy.json` (`"streak_target"`); the flag wins.
- The gate sits **above** the verdict — it never loosens a blocker, it only withholds "done" until the streak is met. Each run re-executes the *actual* tests (no cached passes — that independence is the whole point).
- Reset rule: **any** non-`READY FOR MERGE` verdict (tripped blocker, failing test, or IMPROVEMENTS RECOMMENDED) zeroes the count. No partial credit.
- The verdict surfaces the count: `STREAK 2/3 — one more green to merge`, or `streak reset to 0/3 (security 3.2 &lt; 4.0)`.
- This is the native mechanism the `prd-to-goal` quality-streak recipe (#2539) leans on. Pair it with a `/goal` loop, but **`rm` the ledger first** — `/goal` reads `until` before the turn's verify, so a stale `met:true` exits with zero runs (see streak-gate.md "Stale-ledger guard").

Full protocol — ledger schema, run loop, `/goal` wiring, and `/ork:cover` reuse: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/streak-gate.md")`.

---

## Evidence & Test Execution

Load details: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/rules/evidence-collection.md")` for git commands, test execution patterns, metrics tracking, and post-verification feedback.

---

## Policy-as-Code

Load details: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/policy-as-code.md")` for configuration.

Define verification rules in `.claude/policies/verification-policy.json`:

```json
{
  "thresholds": {
    "composite_minimum": 6.0,
    "security_minimum": 7.0,
    "coverage_minimum": 70
  },
  "blocking_rules": [
    {"dimension": "security", "below": 5.0, "action": "block"}
  ]
}
```

---

## Verification Manifest (VERIFIED vs CLAIMED)

Agent scores, tool summaries, and every "X is clean / passing / fixed" sentence are **claims** until the lead re-runs the proof. Before the verdict, build a **Verification Manifest** marking every *load-bearing* claim ✅ VERIFIED (lead ran it fresh — cites `command · exit · key line`), 🟡 CLAIMED (an agent/tool/doc asserted it, not re-run), ⬜ UNCHECKED, or ⚪ WAIVED (accepted non-blocking, with a reason). An agent's "PASS" copied into the report is still CLAIMED — VERIFIED means *the lead* ran it; a sub-agent's number (price, model-id, count) is CLAIMED until checked against source.

> **Verdict rule:** any load-bearing claim still 🟡 CLAIMED or ⬜ UNCHECKED **caps the verdict at IMPROVEMENTS RECOMMENDED** (never READY FOR MERGE) until it is ✅ VERIFIED or ⚪ WAIVED — this stacks with the dimension-level blockers (both must clear), and under `--streak=N` it resets the streak.

Protocol — claim sources, build step, template, and anti-patterns (laundering, optimism-marking, omission): `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/verification-manifest.md")`.

### Reachability: is the green load-bearing? (REACHED vs UNREACHED)

Provenance answers *who ran it*. It does not answer *whether the pass means anything*. A row reading `✅ VERIFIED · pytest · exit 0 · 214 passed` is honest and can still be worthless, because a suite passing does not prove the suite **reached** the change. A validator shipped 2026-07-19 was fully defined, fully tested, and never called at its call site: every test passed against the old path.

For every test the diff **adds or modifies**, the manifest carries a second mark:

| Mark | Meaning |
|---|---|
| 🟢 **REACHED** | The run showed the test **fail without the change and pass with it**, citing both commands. |
| 🟡 **UNREACHED** | The test is green but has never been seen to fail. Not evidence. |
| ⚪ **WAIVED** | Deliberately accepted with a one-line reason. |

> **Verdict rule:** a test added or modified by this diff that is 🟡 UNREACHED **caps the verdict at IMPROVEMENTS RECOMMENDED** until the proof is shown or the row is ⚪ WAIVED. Stacks with the provenance cap and the dimension blockers — all must clear. Under `--streak=N` it resets the streak.

Two ordering rules make the proof safe, and both come from real damage: **commit before mutating** (`git checkout --` restores to HEAD, so mutating uncommitted work destroys the change on restore), and **mutate the call site, not the new unit** (mutating the unit proves the unit's tests work, and leaves a dead call site undetected).

**This skill does not perform the mutation** — it writes no test files and edits no source. The proof is produced upstream by `/ork:implement` or `/ork:cover` and graded here; absent a proof, the row is 🟡 UNREACHED and the verdict is capped.

Protocol — scope, the 5-step proof, what makes a mutation load-bearing, template, and anti-patterns (coverage-as-proof, batch proof, cosmetic mutation): `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/reachability-proof.md")`.

---

## Report Format

Load details: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/report-template.md")` for full format. Summary:

```markdown
# Feature Verification Report

**Composite Score: [N.N]/10** (Grade: [LETTER])

## Verdict
**[READY FOR MERGE | IMPROVEMENTS RECOMMENDED | BLOCKED]**

[--streak=N mode only: **STREAK [current]/[target]** — READY FOR MERGE requires the full target; any non-ready run resets to 0.]

## Verification Manifest
[✅ VERIFIED · 🟡 CLAIMED · ⬜ UNCHECKED · ⚪ WAIVED — any load-bearing 🟡/⬜ caps the verdict below READY FOR MERGE]
[Reached: 🟢 REACHED · 🟡 UNREACHED · n/a — any 🟡 on a test this diff added/modified also caps the verdict]
| # | Load-bearing claim | Asserted by | Provenance | Reached | Evidence (cmd · exit · key line) |
```

> **Push notifications (CC 2.1.110+):** Verify runs for >5 min are common on complex changes. When the final verdict is ready, call `PushNotification` to alert the user — they likely walked away from the terminal. Requires Remote Control with "Push when Claude decides" config; fails silently for users without it.
>
> ```python
> PushNotification(
>   message=f"ork:verify complete — \{verdict\} · \{score\}/10 · \{blockers_count\} blockers",
>   status="proactive"
> )
> ```

---

## References

Load on demand with `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/references/&lt;file&gt;")`:

| File | Content |
|------|---------|
| `verification-phases.md` | 8-phase workflow, agent spawn definitions, Agent Teams mode |
| `visual-capture.md` | Phase 2.5: screenshot capture, AI vision, gallery generation |
| `quality-model.md` | Scoring dimensions and weights (8 unified) |
| `grading-rubric.md` | Per-agent scoring criteria |
| `report-template.md` | Full report format with visual evidence section |
| `verification-manifest.md` | VERIFIED‑vs‑CLAIMED provenance ledger: states, verdict rule, claim sources, template, anti‑patterns |
| `reachability-proof.md` | REACHED‑vs‑UNREACHED: the mutate→red→restore→green proof, commit-first and call-site rules, verdict cap, anti‑patterns |
| `alternative-comparison.md` | Approach comparison template |
| `orchestration-mode.md` | Agent Teams vs Task Tool |
| `policy-as-code.md` | Verification policy configuration |
| `verification-checklist.md` | Pre-flight checklist |
| `streak-gate.md` | `--streak=N` consecutive-pass gate: ledger schema, reset rule, `/goal` wiring, cover reuse |

## Rules

Load on demand with `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/verify/rules/&lt;file&gt;")`:

| File | Content |
|------|---------|
| `scoring-rubric.md` | Composite scoring, grades, verdicts |
| `evidence-collection.md` | Evidence gathering and test patterns |

### Verification Gate (Cross-Cutting)

Load `Read("$\{CLAUDE_PLUGIN_ROOT\}/shared/rules/verification-gate.md")` — the minimum 5-step gate that applies to ALL completion claims across all skills. This is non-negotiable: NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE.

Producer findings must also satisfy the evidence-replay gate (machine-checkable `\{file, line, quote\}` or `\{command, expected_output\}`, replayed before entering any verdict or score): `Read("$\{CLAUDE_PLUGIN_ROOT\}/shared/rules/evidence-replay.md")`.

### Anti-Sycophancy Protocol

Load `Read("$\{CLAUDE_PLUGIN_ROOT\}/shared/rules/anti-sycophancy.md")` — all verification agents report findings directly without performative agreement. "Should be fine" is not evidence. "Tests pass (exit 0, 47/47)" is.

### Agent Status Protocol

All verification agents MUST report using the standardized protocol: `Read("$\{CLAUDE_PLUGIN_ROOT\}/shared/status-protocol.md")`. Never report DONE if concerns exist. Never silently produce work you're unsure about.

---

## Agent Coordination

### SendMessage (Cross-Agent Findings)

When a security agent finds a critical issue, share it with other verification agents:

```python
SendMessage(to="test-generator", message="Security: SQL injection in user_service.py:88 — add parameterized query test")
SendMessage(to="code-quality-reviewer", message="Security finding at user_service.py:88 — flag in review")
```

### Skill Chain

After verification, chain to commit if all gates pass:

```python
TaskCreate(subject="Commit verified changes", activeForm="Committing")
TaskUpdate(taskId=commit_id, addBlockedBy=[verify_task_id])
# Then: /ork:commit
```

> **Session recovery (CC 2.1.108+):** After idle periods or interruptions, use `/recap` to restore conversational context alongside checkpoint-resume state. Enabled by default since CC 2.1.110 (even with telemetry disabled).

## Quality Bar

Done means all of these hold:
- verdict is exactly one of READY FOR MERGE / IMPROVEMENTS RECOMMENDED / BLOCKED, with the composite and every dimension score cited
- every load-bearing "passing/clean/fixed" claim sits in the Verification Manifest marked VERIFIED (lead re-ran, cites command · exit · key line), CLAIMED, UNCHECKED, or WAIVED
- test evidence is the actual runner summary line (command, exit code, pass count) — never paraphrase
- every test the diff added or modified carries a Reached mark: REACHED cites the failing run AND the passing run; a green-only row is UNREACHED, not evidence
- any dimension below its `min_blocker` is reported BLOCKED regardless of composite
- READY FOR MERGE only when no load-bearing claim is still CLAIMED/UNCHECKED, no diff-added test is still UNREACHED (and under `--streak=N`, the full streak is met)

## Before trusting a green check

A check that prints the same thing whether or not the fault is present has
measured nothing, yet it still returns an answer and that answer reads as
evidence. The `paired-probe` skill exists for this and is model-invocable: it applies
to any check whose PASS decides the verdict,
especially a sweep that reported zero findings or a gate that passed
unexpectedly. It refuses the verdict rather than the check.

## Related Skills

- `ork:implement` - Full implementation with verification
- `ork:review-pr` - PR-specific verification
- `testing-unit` / `testing-integration` / `testing-e2e` - Test execution patterns
- `ork:quality-gates` - Quality gate patterns
- `browser-tools` - Browser automation for visual capture

---

**Version:** 4.6.0 (July 2026) — Added the Reachability Proof (REACHED vs UNREACHED): the manifest's second axis. Provenance grades who ran a claim; reachability grades whether the green means anything. A test the diff added that has never been seen to fail caps the verdict below READY FOR MERGE
**Version:** 4.5.0 (July 2026) — Added the Verification Manifest (VERIFIED vs CLAIMED) — a load-bearing-claim provenance ledger that caps the verdict below READY FOR MERGE until unverified claims are re-run or waived
**Version:** 4.4.0 (June 2026) — Added `--streak=N` consecutive-pass gate (#2540)


---

## Rules (2)

### Evidence Collection Patterns — HIGH


# Evidence Collection Patterns

## Phase 1: Context Gathering

Run these commands in parallel in ONE message:

```bash
git diff main --stat
git log main..HEAD --oneline
git diff main --name-only | sort -u
```

**Incorrect:**
```bash
# Sequential — wastes time, no coverage data
cd backend && pytest tests/
cd frontend && npm test
```

**Correct:**
```bash
# Parallel with coverage — run both in ONE message
cd backend && poetry run pytest tests/ -v --cov=app --cov-report=json
cd frontend && npm run test -- --coverage
```

## Phase 3: Parallel Test Execution

Run backend and frontend tests in parallel:

```bash
# PARALLEL - Backend and frontend
cd backend && poetry run pytest tests/ -v --cov=app --cov-report=json
cd frontend && npm run test -- --coverage
```

## Phase 7: Metrics Tracking

Store verification metrics in memory for trend analysis:

```python
mcp__memory__create_entities(entities=[{
  "name": "verification-{date}-{feature}",
  "entityType": "VerificationMetrics",
  "observations": [f"composite_score: {score}", ...]
}])
```

Query trends: `mcp__memory__search_nodes(query="VerificationMetrics")`

## Phase 2.5: Visual Evidence Collection

Run in parallel with Phase 2 agents. Auto-detects frontend framework and captures screenshots.

**Incorrect:**
```bash
# Manual screenshots with no structure
open http://localhost:3000
# Take manual screenshot...
```

**Correct:**
```python
# Automated visual capture with AI evaluation
Agent(
  subagent_type="general-purpose",
  prompt="Visual capture: detect framework, start server, screenshot routes via agent-browser, evaluate with Claude vision, generate gallery.html",
  run_in_background=True
)
```

Output structure:
```
verification-output/{timestamp}/
├── screenshots/          (PNGs per route, base64 in gallery)
├── ai-evaluations/       (JSON per screenshot with score + issues)
└── gallery.html          (self-contained, open in browser)
```

## Phase 8.5: Post-Verification Feedback

After report compilation, store verification scores in the memory graph for KPI baseline tracking:

Query trends: `mcp__memory__search_nodes(query="VerificationScores")`


### Scoring Rubric — HIGH


# Scoring Rubric

## Composite Score

Each agent produces a 0-10 score with decimals for nuance. The composite score is a weighted sum using the weights from [Quality Model](../references/quality-model.md).

## Grade Thresholds

&lt;!-- Canonical source: ../references/quality-model.md — keep in sync --&gt;

| Grade | Score Range | Verdict |
|-------|-------------|---------|
| A+ | 9.0-10.0 | EXCELLENT |
| A | 8.0-8.9 | READY FOR MERGE |
| B | 7.0-7.9 | READY FOR MERGE |
| C | 6.0-6.9 | IMPROVEMENTS RECOMMENDED |
| D | 5.0-5.9 | IMPROVEMENTS RECOMMENDED |
| F | 0.0-4.9 | BLOCKED |

## Key Decisions

| Decision | Choice | Rationale |
|----------|--------|-----------|
| Scoring scale | 0-10 with decimals | Nuanced, not binary |
| Improvement priority | Impact / Effort ratio | Do high-value first |
| Alternative comparison | Optional phase | Only when multiple valid approaches |
| Metrics persistence | Memory MCP | Track trends over time |

**Incorrect:**
```
Security: "looks fine"  → 8/10    # No evidence, subjective
Performance: "fast enough" → 7/10  # No benchmarks
```

**Correct:**
```
Security: "11/11 injection tests pass, 13 deny patterns, 0 CVEs" → 9/10
Performance: "p99 latency 142ms (budget: 300ms), 0 N+1 queries" → 8.5/10
```

## Improvement Suggestions

Each suggestion includes effort (1-5) and impact (1-5) with priority = impact/effort. See [Quality Model](../references/quality-model.md) for scale definitions and quick wins formula.

## Blocking Rules

Verification can be blocked by policy-as-code rules. See [Policy-as-Code](../references/policy-as-code.md) for configuration of composite minimums, dimension minimums, and blocking rules.



---

## References (13)

### Alternative Comparison

# Alternative Comparison

Evaluate current implementation against alternative approaches.

## When to Compare

- Multiple valid architectures exist
- User asks "is this the best way?"
- Major patterns were chosen (ORM vs raw SQL, REST vs GraphQL)
- Performance/scalability concerns raised

---

## Comparison Criteria

### For Each Alternative

| Criterion | Weight | Description |
|-----------|--------|-------------|
| Effort | 30% | Implementation complexity (1-5 scale) |
| Risk | 25% | Technical and operational risk (1-5 scale) |
| Benefit | 45% | Value delivered, performance, maintainability (1-5 scale) |

### Migration Cost

| Factor | Estimate |
|--------|----------|
| Code changes | Files/lines affected |
| Data migration | Schema changes, backfill |
| Testing | New test coverage needed |
| Rollback risk | Reversibility |

---

## Decision Matrix Format

| Approach | Effort | Risk | Benefit | Score |
|----------|--------|------|---------|-------|
| Current | N | N | N | (E*0.3 + R*0.25 + B*0.45) |
| Alt A | N | N | N | calculated |
| Alt B | N | N | N | calculated |

**Note:** Higher effort and risk are bad (invert for scoring), higher benefit is good.

**Recommendation Formula:**
```
Score = (5 - Effort) * 0.3 + (5 - Risk) * 0.25 + Benefit * 0.45
```

---

## Output Template

```markdown
### Alternative Comparison: [Topic]

**Current Approach:** [description]
- Score: N/10
- Pros: [strengths]
- Cons: [weaknesses]

**Alternative A:** [description]
- Score: N/10
- Pros: [strengths]
- Cons: [weaknesses]
- Migration effort: [1-5]

**Recommendation:** [Keep current / Switch to Alt A]
**Justification:** [1-2 sentences]
```


### Background task contract: never wait unconditionally


# Background task contract

`/ork:verify` hung for ~40 minutes producing **no verdict** (#3263, 2026-08-03).
Three backgrounded test suites each wrote **0 bytes** and no process was alive
afterwards. This file records the contract that prevents a repeat, and the
evidence behind it, so nobody re-derives the refuted hypotheses.

## The rule

1. **Capture the task id.** If `run_in_background` returned no id, the request
   was not honoured: the command ran synchronously and its output came back
   inline. There is nothing to wait for. Use the inline output.
2. **Bound every wait.** A wait that can only end on a notification will never
   end if the notification cannot fire.
3. **An empty run is a FAILURE, not a pending one.** If the output is still
   empty and no matching process is alive, the suite never ran. Report that.
   Never emit a grade, a score, or a "verified" claim from a run that produced
   no bytes.

Rule 3 is the important one. A verification skill that can silently no-op
undermines every claim ever made through it: the operator reads "no verdict
yet" as "still working" rather than "never started".

## Diagnosis: what npm's banner proves

npm writes `> pkg@version script` **before** executing any script body. So any
hypothesis in which npm started predicts a non-empty file. All three failing
tasks were 0 bytes, which refutes every "the suite failed / hung" explanation
at once. Useful primitive: **the absence of a tool's own startup banner locates
the failure before the tool, not inside it.**

## Hypotheses, settled

| # | hypothesis | verdict |
|---|---|---|
| H1 | dies under the #3235 concurrency bug | REFUTED, requires npm to start |
| H2 | cwd/PATH differ so `npm` is unresolvable | REFUTED with positive evidence: npm resolves, node v22, exits 0 |
| H3 | `Monitor` attaches to an already-exited task | reclassified: a consequence, not a cause |
| H4 | child reaped when the subagent exits | REFUTED, child ran a 20s sleep to completion |
| H5 | backgrounding from inside a fork is ignored | LEADING, not proven for the skill-fork path |

**H5 is not proven.** The probe that demonstrated inline execution was an
*Agent* fork and received no task id at all; `/ork:verify` is a *skill* fork and
did receive ids (`b6tvprwm6`, `bdycijacv`, `brvhtn4h3`) with empty files. The two
paths only share a mechanism if the same plumbing serves both, which is an
assumption. Do not write "fixed by H5" anywhere until that is measured.

The contract above deliberately does not depend on H5 being right. Whatever
empties the file, an unbounded wait on an unconfirmed task is a defect on its
own, and bounding it converts a silent hang into a reported failure.

## Reproducing the original signature is still blocked

It needs a real `/ork:verify` run, which runs `npm test`.
`tests/unit/test-sync-versions.sh:76` sets `FAKE_VER="999.888.777"` and
propagates it through `scripts/stamp-counts.sh` into tracked files, restoring
them with a trap. That makes `npm test` unsafe while any other session is live.

`--quick` does **not** avoid it: `tests/run-all-tests.sh:63` clears only
`RUN_INTEGRATION`, `RUN_E2E` and `RUN_PERFORMANCE`, leaving `RUN_UNIT="true"`,
so the sentinel still runs. Verified at HEAD `b6e65248f` on 2026-08-06.

Settling H5 therefore requires either a quiet tree or making the sentinel test
operate on a copy instead of the live working tree. That is tracked separately
under #3235.


### Grading Rubric

# Verification Grading Rubric

0-10 scoring criteria for each verification dimension.

## Score Levels

| Range | Level | Description |
|-------|-------|-------------|
| 0-3 | Poor | Critical issues, blocks merge |
| 4-6 | Adequate | Functional but needs improvement |
| 7-9 | Good | Ready for merge, minor suggestions |
| 10 | Excellent | Exemplary, reference quality |

---

## Dimension Rubrics

&lt;!-- Weights from canonical source: ../references/quality-model.md — keep in sync --&gt;

### Correctness (Weight: 14%)

| Score | Criteria |
|-------|----------|
| 10 | All functional requirements met, edge cases handled, zero regressions |
| 8-9 | Core requirements met, most edge cases handled |
| 6-7 | Core paths work, some edge cases missing |
| 4-5 | Partial functionality, notable gaps |
| 1-3 | Broken core paths |
| 0 | Does not run |

### Maintainability (Weight: 14%)

| Score | Criteria |
|-------|----------|
| 10 | Zero lint errors/warnings, strict types, exemplary patterns, low complexity |
| 8-9 | Zero errors, &lt; 5 warnings, minimal `any`, good patterns |
| 6-7 | 1-3 errors, some warnings, acceptable patterns |
| 4-5 | 4-10 errors, pattern issues, needs refactoring |
| 1-3 | Many errors, poor patterns, high complexity |
| 0 | Lint/type check fails to run |

### Performance (Weight: 11%)

| Score | Criteria |
|-------|----------|
| 10 | p99 within budget, zero N+1, optimal caching, efficient resource usage |
| 8-9 | Good latency, no N+1, reasonable caching |
| 6-7 | Acceptable latency, minor inefficiencies |
| 4-5 | Notable bottlenecks, missing caching |
| 1-3 | Severe bottlenecks, resource leaks |
| 0 | Unresponsive or crashes under load |

### Security (Weight: 18%)

| Score | Criteria |
|-------|----------|
| 10 | No vulnerabilities, all OWASP compliant, secure by design |
| 8-9 | No critical/high, all OWASP, excellent practices |
| 6-7 | No critical, 1-2 high, most OWASP compliant |
| 4-5 | No critical, 3-5 high, some gaps |
| 1-3 | 1+ critical or many high vulnerabilities |
| 0 | Multiple critical, secrets exposed |

### Scalability (Weight: 9%)

| Score | Criteria |
|-------|----------|
| 10 | Horizontal scaling ready, stateless design, efficient data patterns |
| 8-9 | Good scaling patterns, minor bottlenecks |
| 6-7 | Scales for current needs, some concerns |
| 4-5 | Will hit limits soon, needs rework |
| 1-3 | Single-instance only, monolithic state |
| 0 | Cannot handle production load |

### Testability (Weight: 12%)

| Score | Criteria |
|-------|----------|
| 10 | >= 90% coverage, meaningful assertions, edge cases, no flaky tests |
| 8-9 | >= 80% coverage, good assertions, critical paths |
| 6-7 | >= 70% coverage (target), basic assertions |
| 4-5 | 50-69% coverage |
| 1-3 | 30-49% coverage |
| 0 | &lt; 30% coverage or tests fail to run |

### Compliance (Weight: 12%)

| Score | Criteria |
|-------|----------|
| 10 | Perfect REST/UI contracts, RFC 9457 errors, full Zod, WCAG AA |
| 8-9 | Good conventions, proper validation, accessibility |
| 6-7 | Acceptable patterns, minor inconsistencies |
| 4-5 | Several convention violations |
| 1-3 | Poor API/UI design, missing validation |
| 0 | Broken contracts or inaccessible |

### Visual (Weight: 10%)

| Score | Criteria |
|-------|----------|
| 10 | Pixel-perfect layout, full a11y, complete content, responsive |
| 8-9 | Good layout, minor visual issues, WCAG AA |
| 6-7 | Acceptable layout, some a11y gaps |
| 4-5 | Layout issues, missing content, a11y problems |
| 1-3 | Broken layout, major content missing |
| 0 | Page fails to render |

> **Note**: Visual weight is 0.00 for API-only projects — redistributed proportionally. See [Quality Model](quality-model.md).

---

## Grade Interpretation

&lt;!-- Canonical source: quality-model.md — keep in sync --&gt;

| Composite | Grade | Verdict |
|-----------|-------|---------|
| 9.0-10.0 | A+ | EXCELLENT |
| 8.0-8.9 | A | READY FOR MERGE |
| 7.0-7.9 | B | READY FOR MERGE |
| 6.0-6.9 | C | IMPROVEMENTS RECOMMENDED |
| 5.0-5.9 | D | IMPROVEMENTS RECOMMENDED |
| 0.0-4.9 | F | BLOCKED |

---

## Dimension-Level Blockers (ork-rubric/1.0)

The composite table above is overridden by per-dimension floors. Thresholds live in `../rubric.json` (`min_blocker` / `min_pass` fields; schema: `../../shared/rubric.schema.json`) — they map to the 0-3 "Poor, blocks merge" band of each dimension rubric.

| Dimension | Threshold | Effect when below |
|-----------|-----------|-------------------|
| Security | `min_blocker` 4.0 | Verdict BLOCKED regardless of composite |
| Compliance | `min_pass` 6.0 | Verdict capped at IMPROVEMENTS RECOMMENDED |

Reporting format — the tripped dimension leads the verdict, with the threshold named:

```
Security 3.2/10 (CRITICAL BLOCKER — below min_blocker 4.0)
```

A passing composite never clears a tripped `min_blocker`; list every tripped dimension with the fix required to clear it.


### Orchestration Mode

&lt;!-- SHARED: keep in sync with ../../../assess/references/orchestration-mode.md --&gt;
# Orchestration Mode Selection

Shared logic for choosing between Agent Teams and Task tool orchestration in assess/verify skills.

## Environment Check

```python
# Agent Teams is GA since CC 2.1.33
import os
force_task_tool = os.environ.get("ORCHESTKIT_FORCE_TASK_TOOL") == "1"

if force_task_tool:
    mode = "task_tool"
else:
    # Teams available by default — use for full multi-dimensional work
    mode = "agent_teams" if scope == "full" else "task_tool"
```

## Decision Rules

1. Full assessment/verification scope --> **Agent Teams mode** (GA since CC 2.1.33)
2. Quick/single-dimension scope --> **Task tool mode**
3. `ORCHESTKIT_FORCE_TASK_TOOL=1` --> **Task tool** (override)

## Agent Teams vs Task Tool

| Aspect | Task Tool (Star) | Agent Teams (Mesh) |
|--------|------------------|-------------------|
| Topology | All agents report to lead | Agents communicate with each other |
| Finding correlation | Lead cross-references after completion | Agents share findings in real-time |
| Cross-domain overlap | Independent scoring | Agents alert each other about overlapping concerns |
| Cost | ~200K tokens | ~500K tokens |
| Best for | Focused/single-dimension work | Full multi-dimensional assessment/verification |

## Fallback

If Agent Teams encounters issues mid-execution, fall back to Task tool for remaining work. This is safe because both modes produce the same output format (dimensional scores 0-10).

## Context Window Note

For full codebase work (>20 files), use the 1M context window to avoid agent context exhaustion. On 200K context, scope discovery should limit files to prevent overflow.


### Policy As Code

# Policy-as-Code

Define verification policies as machine-readable configuration.

## Policy Structure

```yaml
version: "1.0"
name: policy-name
description: What this policy enforces

thresholds:
  composite_minimum: 6.0
  coverage_minimum: 70

rules:
  blockers: []    # Fail verification
  warnings: []    # Note but continue
  info: []        # Informational only
```

---

## Rule Definition

### Blocker Rules (Must Pass)

```yaml
blockers:
  - dimension: security
    condition: below
    value: 5.0
    message: "Security score below minimum"

  - check: critical_vulnerabilities
    condition: above
    value: 0
    message: "Critical vulnerabilities found"

  - check: type_errors
    condition: above
    value: 0
    message: "TypeScript errors must be zero"
```

### Warning Rules (Should Fix)

```yaml
warnings:
  - dimension: code_quality
    condition: below
    value: 7.0
    message: "Code quality could be improved"

  - check: test_coverage
    condition: below
    value: 80
    message: "Coverage below recommended 80%"
```

### Info Rules (Awareness)

```yaml
info:
  - check: todo_count
    condition: above
    value: 5
    message: "Multiple TODOs found in code"
```

---

## Threshold Configuration

| Threshold | Type | Description |
|-----------|------|-------------|
| composite_minimum | float | Overall score minimum (0-10) |
| coverage_minimum | int | Test coverage percentage |
| critical_vulnerabilities | int | Max critical vulns (0) |
| high_vulnerabilities | int | Max high vulns |
| lint_errors | int | Max lint errors (0) |
| type_errors | int | Max type errors (0) |

---

## Custom Rules

```yaml
custom_rules:
  - name: no_console_log
    pattern: "console\\.log"
    file_glob: "**/*.ts"
    exclude: ["**/*.test.ts"]
    severity: warning
    message: "Remove console.log from production"
```

---

## Policy Location

Store at: `.claude/policies/verification-policy.yaml`

Multiple policies: `.claude/policies/\{name\}-policy.yaml`


### Quality Model

# Quality Model (verify)

Extends the unified scoring framework with Visual as the 8th dimension.

> **Canonical source**: `quality-gates/references/unified-scoring-framework.md`
> Load: `Read("$\{CLAUDE_PLUGIN_ROOT\}/skills/quality-gates/references/unified-scoring-framework.md")`

## verify-Specific Extensions

### Visual Dimension (8th)

| Dimension | Weight | What It Measures |
|-----------|--------|------------------|
| Visual | 0.10 | Layout correctness, a11y, content completeness, responsiveness |

When Visual is active, base dimensions scale: `adjusted = base_weight * (1.0 / 1.10)`.
When Visual is skipped (API-only), base weights stay at 1.00.

## Dimensions Used (with Visual)

| Dimension | Adjusted Weight |
|-----------|----------------|
| Correctness | 0.14 |
| Maintainability | 0.14 |
| Performance | 0.11 |
| Security | 0.18 |
| Scalability | 0.09 |
| Testability | 0.12 |
| Compliance | 0.12 |
| Visual | 0.10 |

See unified framework for grade thresholds, improvement prioritization, effort/impact scales, and blocking rules.


### Reachability Proof

# Reachability Proof (REACHED vs UNREACHED)

The [Verification Manifest](verification-manifest.md) closes one seam: *did the lead
actually re-run the claim, or relay someone else's assertion?* It does not close the
second seam, and the second seam is the one that ships bugs:

> `✅ VERIFIED · uv run pytest tests/unit -q · exit 0 · 214 passed`

That row is honest. The lead ran it, read the output, cited the key line. And it can
still be worthless, because **a suite passing does not mean the suite exercised the
change**. A green test that never reaches the new code is a proxy signal standing in for
correctness. The manifest grades *who ran it*. Reachability grades *whether the green
means anything*.

This is not hypothetical. A validator shipped 2026-07-19 was fully defined, fully
tested, and never called at its call site. Every test passed while exercising the old
path. The suite was green, the manifest would have read VERIFIED, and the feature was
inert.

## The two axes

| | Question | States |
|---|---|---|
| **Provenance** (manifest) | Did the *lead* run this, or relay it? | ✅ VERIFIED · 🟡 CLAIMED · ⬜ UNCHECKED · ⚪ WAIVED |
| **Reachability** (this doc) | Would this test *fail* if the change were absent? | 🟢 REACHED · 🟡 UNREACHED · ⚪ WAIVED |

A row can be VERIFIED and UNREACHED at the same time. That combination is the most
dangerous state in the report, because it looks like the strongest one.

## Scope: which claims need a reachability proof

Only tests that are **new or modified in the diff under verification**. An untouched
regression suite does not need one; it is not being offered as evidence that *this*
change works.

Concretely, a proof is required for each test file the diff adds or edits, and for any
claim of the form "the fix is covered" / "there is a test for this."

## The proof

A test is 🟢 REACHED when the run has shown it **fail without the change and pass with
it**. Nothing weaker counts. Not "the test looks like it covers this." Not "coverage
reports the line." Seeing it fail is the proof.

```
1. COMMIT the change first.          # non-negotiable, see below
2. Mutate the CALL SITE.             # not the new unit, see below
3. Run the test  ->  expect RED.     # this is the proof
4. git checkout -- <mutated file>    # restores to HEAD == the committed change
5. Run the test  ->  expect GREEN.
```

Record both runs in the manifest row. A proof with only step 5 is not a proof.

### Rule 1: commit before mutating

`git checkout -- &lt;file&gt;` restores the file to **HEAD**, not to the working state you had
before the mutation. If the change is still uncommitted when you mutate, step 4 reverts
the mutation *and the change together*, silently destroying the work. This has already
cost a real fix plus roughly forty lines of a new guard in one stroke.

Commit first, then mutate, then restore. The order is the safety property.

### Rule 2: mutate the call site, not the new unit

The failure this proof exists to catch is *new code that is never invoked*. Mutating the
new unit proves only that the unit's own tests exercise the unit, which was never in
doubt. It leaves the actual defect, a dead call site, completely undetected.

Mutate where the new code is **called from**:

| Change | Weak mutation (proves little) | Correct mutation |
|---|---|---|
| New validator function | break the validator body | delete the call to it in the handler |
| New middleware | break the middleware logic | remove it from the middleware chain |
| New guard clause | invert the guard's condition | remove the guard entirely |
| New config flag read | change the default | delete the branch that reads it |

If deleting the call site does **not** turn the test red, the test is not testing the
change.

### What a good mutation looks like

It must be **semantically load-bearing**, not cosmetic. Renaming a local variable,
touching a comment, or reformatting proves nothing. Delete the call, invert the
condition, or return the opposite value.

It must also be **exactly one thing**. Two simultaneous mutations produce a red you
cannot attribute.

## Verify does not perform the mutation

`ork:verify` writes no test files and edits no source. That contract is not relaxed
here. Verify **grades** the reachability proof; it does not produce one.

- The proof is produced by whoever wrote the code, typically `/ork:implement` or
  `/ork:cover`, and carried into the run as evidence.
- If no proof exists, verify marks the row 🟡 UNREACHED and the verdict is capped. It
  does not mutate the tree to manufacture one.

This keeps verify read-only and puts the burden where the knowledge is. A skill that
edited source to grade itself would be the same category error the proof exists to
catch.

## The rule (wired to the verdict)

> **A test added or modified by this diff that is 🟡 UNREACHED caps the verdict at
> IMPROVEMENTS RECOMMENDED. It cannot grade READY FOR MERGE until the proof is shown or
> the row is explicitly ⚪ WAIVED with a reason.**

This stacks with the existing gates rather than replacing them: dimension blockers, the
provenance cap, and this reachability cap must all clear. Under `--streak=N`, a run
carrying an UNREACHED new test is not READY and therefore resets the streak to 0.

## Template

Extend the manifest with a Reached column. Leave it blank (`n/a`) for rows that are not
test-coverage claims.

```markdown
## Verification Manifest

| # | Load-bearing claim | Asserted by | Provenance | Reached | Evidence (cmd · exit · key line) |
|---|---|---|---|---|---|
| 1 | Types check clean (web) | frontend-ui-dev | ✅ VERIFIED | n/a | `pnpm --filter web type-check` · exit 0 · 0 errors |
| 2 | New portal-auth guard is covered | test-generator | ✅ VERIFIED | 🟢 REACHED | deleted `get_portal_context` call in routes/portal.py:44 · `pytest tests/unit/routes/test_portal.py` · exit 1 · 3 failed · restored · exit 0 · 3 passed |
| 3 | Retry logic is covered | test-generator | ✅ VERIFIED | 🟡 UNREACHED | suite green, but no failing run shown → **caps verdict** |
| 4 | Legacy suite still green | ci | ✅ VERIFIED | n/a | untouched by this diff |

**Manifest verdict impact:** row 3 is a new test with no reachability proof →
verdict capped at IMPROVEMENTS RECOMMENDED until REACHED or WAIVED.
```

Row 3 is the entire point. It is VERIFIED, it is green, and it is still not evidence.

## Anti-patterns

- **Coverage as proof.** A line-coverage report says a line executed, not that an
  assertion would fail if the behaviour changed. Coverage is not reachability.
- **Mutating the new unit.** Passes the ritual, misses the dead call site. See Rule 2.
- **Mutating before committing.** Destroys the change on restore. See Rule 1.
- **Cosmetic mutation.** A renamed variable that still turns the test red means the test
  is asserting on the wrong thing, which is its own finding, not a pass.
- **Proof by assertion.** "I verified the test fails without the fix" with no command and
  no exit code is 🟡 CLAIMED provenance *and* 🟡 UNREACHED. Both caps apply.
- **Batch proof.** One mutation and a whole-suite red does not prove which test caught
  it. Prove per claim, and name the test that went red.

## Effort scaling

| `/effort` | Reachability |
|---|---|
| **low** | Skipped. Note in the report that reachability was not graded. |
| **medium** | Required for new tests covering the diff's primary behaviour change. |
| **high** / **xhigh** | Required for every test the diff adds or modifies. |

At every level, an UNREACHED row that is *present* still caps the verdict. The effort
level controls how many rows you are obliged to open, not whether the cap applies.


### Report Template

# Verification Report Template

Copy this template and fill in results from parallel agent verification.

## Quick Copy Template

```markdown
# Feature Verification Report

**Date**: [TODAY'S DATE]
**Branch**: [branch-name]
**Feature**: [feature description]
**Reviewer**: Claude Code with 5 parallel subagents
**Verification Duration**: [X minutes]

---

## Summary

**Status**: [READY FOR MERGE | NEEDS ATTENTION | BLOCKED]

[1-2 sentence summary of verification results]

---

## Verification Manifest

> Provenance of every load-bearing claim. Any 🟡 CLAIMED or ⬜ UNCHECKED row caps the
> status below READY FOR MERGE until it is ✅ VERIFIED or ⚪ WAIVED. See
> `references/verification-manifest.md`.

| # | Load-bearing claim | Asserted by | Provenance | Evidence (cmd · exit · key line) |
|---|--------------------|-------------|------------|----------------------------------|
| 1 | [e.g. Types check clean] | [agent/tool/doc] | ✅ VERIFIED / 🟡 CLAIMED / ⬜ UNCHECKED / ⚪ WAIVED | [`cmd` · exit 0 · key line, or reason] |

**Manifest impact**: [none — all load-bearing claims VERIFIED] / [verdict capped at IMPROVEMENTS RECOMMENDED — rows [N] unverified]

---

## Agent Results

### 1. Code Quality (code-quality-reviewer)

| Check | Tool | Exit Code | Errors | Warnings | Status |
|-------|------|-----------|--------|----------|--------|
| Backend Lint | Ruff | 0/1 | N | N | PASS/FAIL |
| Backend Types | ty | 0/1 | N | N | PASS/FAIL |
| Frontend Lint | Biome | 0/1 | N | N | PASS/FAIL |
| Frontend Types | tsc | 0/1 | N | N | PASS/FAIL |

**Pattern Compliance:**
- [ ] No `console.log` in production code
- [ ] No `any` types in TypeScript
- [ ] Exhaustive switches with `assertNever`
- [ ] SOLID principles followed
- [ ] Cyclomatic complexity < 10

**Findings:**
- [List any pattern violations]

---

### 2. Security Audit (security-auditor)

| Check | Tool | Critical | High | Medium | Low | Status |
|-------|------|----------|------|--------|-----|--------|
| JS Dependencies | npm audit | N | N | N | N | PASS/BLOCK |
| Python Dependencies | pip-audit | N | N | N | N | PASS/BLOCK |
| Secrets Scan | grep/gitleaks | N/A | N/A | N/A | N | PASS/BLOCK |

**OWASP Top 10 Compliance:**
- [ ] A01: Broken Access Control
- [ ] A02: Cryptographic Failures
- [ ] A03: Injection
- [ ] A04: Insecure Design
- [ ] A05: Security Misconfiguration
- [ ] A06: Vulnerable Components
- [ ] A07: Auth Failures
- [ ] A08: Data Integrity Failures
- [ ] A09: Logging Failures
- [ ] A10: SSRF

**Findings:**
- [List any security issues]

---

### 3. Test Coverage (test-generator)

| Suite | Total | Passed | Failed | Skipped | Coverage | Target | Status |
|-------|-------|--------|--------|---------|----------|--------|--------|
| Backend Unit | N | N | N | N | X% | 70% | PASS/FAIL |
| Backend Integration | N | N | N | N | X% | 70% | PASS/FAIL |
| Frontend Unit | N | N | N | N | X% | 70% | PASS/FAIL |
| E2E | N | N | N | N | N/A | N/A | PASS/FAIL |

**Test Quality:**
- [ ] Meaningful assertions (not just `assert result`)
- [ ] Edge cases covered (empty, error, timeout)
- [ ] No flaky tests (no sleep, no timing deps)
- [ ] MSW used for API mocking (not jest.mock)

**Coverage Gaps:**
- [List uncovered critical paths]

---

### 4. API Compliance (backend-system-architect)

| Check | Compliant | Issues |
|-------|-----------|--------|
| REST Conventions | Yes/No | [details] |
| Pydantic v2 Validation | Yes/No | [details] |
| RFC 9457 Error Handling | Yes/No | [details] |
| Async Timeout Protection | Yes/No | [details] |
| No N+1 Queries | Yes/No | [details] |

**Findings:**
- [List any API compliance issues]

---

### 5. UI Compliance (frontend-ui-developer)

| Check | Compliant | Issues |
|-------|-----------|--------|
| React 19 APIs (useOptimistic, useFormStatus, use()) | Yes/No | [details] |
| Zod Validation on API Responses | Yes/No | [details] |
| Exhaustive Type Checking | Yes/No | [details] |
| Skeleton Loading States | Yes/No | [details] |
| Prefetching on Navigation | Yes/No | [details] |
| WCAG 2.1 AA Accessibility | Yes/No | [details] |

**Findings:**
- [List any UI compliance issues]

---

## Quality Gates Summary

| Gate | Required | Actual | Status |
|------|----------|--------|--------|
| Test Coverage | >= 70% | X% | PASS/FAIL |
| Security Critical | 0 | N | PASS/FAIL |
| Security High | <= 5 | N | PASS/FAIL |
| Type Errors | 0 | N | PASS/FAIL |
| Lint Errors | 0 | N | PASS/FAIL |

**Overall Gate Status**: [ALL PASS | SOME FAIL]

---

## Blockers (Must Fix Before Merge)

1. [Blocker description with file:line reference]
2. [Blocker description with file:line reference]

---

## Suggestions (Non-Blocking)

1. [Suggestion for improvement]
2. [Suggestion for improvement]

---

## Visual Verification

**Visual Score: [N.N]/10**

| Route | Screenshot | AI Score | Issues | Status |
|-------|-----------|----------|--------|--------|
| / | [thumbnail] | N.N/10 | N | PASS/WARN/FAIL |
| /dashboard | [thumbnail] | N.N/10 | N | PASS/WARN/FAIL |
| /settings | [thumbnail] | N.N/10 | N | PASS/WARN/FAIL |

**Gallery**: Open `verification-output/{timestamp}/gallery.html` for full screenshots with AI evaluations.

---

## Evidence Artifacts

| Artifact | Location | Generated |
|----------|----------|-----------|
| Test Results | `/tmp/test_results.log` | [timestamp] |
| Coverage Report | `/tmp/coverage.json` | [timestamp] |
| Security Scan | `/tmp/security_audit.json` | [timestamp] |
| Lint Report | `/tmp/lint_results.log` | [timestamp] |
| Visual Gallery | `verification-output/{timestamp}/gallery.html` | [timestamp] |
| Screenshots | `verification-output/{timestamp}/screenshots/` | [timestamp] |
| AI Evaluations | `verification-output/{timestamp}/ai-evaluations/` | [timestamp] |

---

## Verification Metadata

- **Agents Used**: 7 (code-quality-reviewer, security-auditor, test-generator, backend-system-architect, frontend-ui-developer, python-performance-engineer, visual-capture)
- **Parallel Execution**: Yes
- **Total Tool Calls**: ~N
- **Context Usage**: ~N tokens
```

---

## Status Definitions

| Status | Emoji | Meaning | Action Required |
|--------|-------|---------|-----------------|
| READY FOR MERGE | Green | All checks pass, no blockers | Approve PR |
| NEEDS ATTENTION | Yellow | Minor issues found | Review suggestions, optionally fix |
| BLOCKED | Red | Critical issues found | Must fix before merge |

## Severity Levels

| Level | Threshold | Action | Blocks Merge |
|-------|-----------|--------|--------------|
| Critical | Any | Fix immediately | YES |
| High | > 5 | Fix before merge | YES |
| Medium | > 20 | Should fix | NO (with justification) |
| Low | > 50 | Nice to have | NO |
| Info | N/A | Informational | NO |

## Agent Output JSON Schemas

### code-quality-reviewer Output
```json
{
  "linting": {"tool": "ruff|biome", "exit_code": 0, "errors": 0, "warnings": 0},
  "type_check": {"tool": "ty|tsc", "exit_code": 0, "errors": 0},
  "patterns": {"violations": [], "compliance": "PASS|FAIL"},
  "approval": {"status": "APPROVED|NEEDS_FIXES", "blockers": []}
}
```

### security-auditor Output
```json
{
  "scan_summary": {"files_scanned": 100, "vulnerabilities_found": 0},
  "critical": [],
  "high": [],
  "secrets_detected": [],
  "recommendations": [],
  "approval": {"status": "PASS|BLOCK", "blockers": []}
}
```

### test-generator Output
```json
{
  "coverage": {"current": 85, "target": 70, "passed": true},
  "test_summary": {"total": 100, "passed": 98, "failed": 2, "skipped": 0},
  "gaps": ["file:line - reason"],
  "quality_issues": [],
  "approval": {"status": "PASS|FAIL", "blockers": []}
}
```

### backend-system-architect Output
```json
{
  "api_compliance": {"rest_conventions": true, "issues": []},
  "validation": {"pydantic_v2": true, "issues": []},
  "error_handling": {"rfc9457": true, "issues": []},
  "async_safety": {"timeouts": true, "issues": []},
  "approval": {"status": "PASS|FAIL", "blockers": []}
}
```

### frontend-ui-developer Output
```json
{
  "react_19": {"apis_used": ["useOptimistic"], "missing": [], "compliant": true},
  "zod_validation": {"validated_endpoints": 10, "unvalidated": []},
  "type_safety": {"exhaustive_switches": true, "any_types": 0},
  "ux_patterns": {"skeletons": true, "prefetching": true},
  "accessibility": {"wcag_issues": []},
  "approval": {"status": "PASS|FAIL", "blockers": []}
}
```

### Streak Gate

# Streak Gate — consecutive-pass verification

A single green is not proof. Flaky suites, race conditions, and order-dependent tests pass once and fail the next run. The streak gate makes `/ork:verify` declare a feature done **only after N consecutive passing runs**, resetting the count to zero on any failure. It is the flakiness defense that single-shot pass/fail can't give.

> Loop-Library theme: **Quality Streak** ("fixes product failures until a defined streak of realistic tests passes"). This is the native ork mechanism the `prd-to-goal` quality-streak recipe leans on.

## Invocation

```bash
/ork:verify --streak=3 authentication flow     # need 3 greens in a row
/ork:verify --streak=3                          # continues an existing streak for the same scope
```

`--streak=N` (N ≥ 2). When absent, verify behaves exactly as before (single pass/fail). The target may also come from the verify rubric's `streak_target` slot (`src/skills/verify/rubric.json`, an integer ≥ 2 validated by `src/shared/rubric.schema.json` — a configured value is schema-checked, not silently ignored); the explicit flag always wins.

Parse it alongside the other flags in Argument Resolution:

```python
STREAK_TARGET = None
for token in "$ARGUMENTS".split():
    if token.startswith("--streak="):
        STREAK_TARGET = int(token.split("=", 1)[1])   # explicit override
        SCOPE = SCOPE.replace(token, "").strip()
STREAK_TARGET = STREAK_TARGET or rubric.get("streak_target")  # validated slot; may stay None
```

## Ledger

The streak persists across independent verify runs in `.claude/chain/verify-streak.json` so each invocation extends (or breaks) the run before it:

```json
{
  "schema": "verify-streak/1.0",
  "scope": "authentication flow",
  "target": 3,
  "current": 2,
  "met": false,
  "reset_count": 1,
  "last_run_ts": "2026-06-20T10:31Z",
  "history": [
    { "ts": "2026-06-20T10:01Z", "verdict": "fail", "composite": 5.4, "blocker": "security 3.2" },
    { "ts": "2026-06-20T10:14Z", "verdict": "pass", "composite": 7.8 },
    { "ts": "2026-06-20T10:31Z", "verdict": "pass", "composite": 8.1 }
  ]
}
```

`scope` keys the streak — switching scope starts a fresh streak. `current` is the live consecutive-pass count; `met` is `current >= target`. `last_run_ts` is the timestamp of the run that last wrote the ledger — the freshness stamp a `/goal` loop checks so it never trusts a `met:true` it didn't just produce (see Stale-ledger guard).

**Scope keying is normalized.** The raw scope string is trimmed and its internal whitespace collapsed before it keys the streak, so `"auth flow"` and `"auth  flow"` (or a trailing space) extend the *same* streak instead of silently starting fresh ones:

```python
def streak_key(scope: str) -> str:
    return " ".join(scope.split())   # trim + collapse runs of whitespace
```

## Run protocol

Each `/ork:verify --streak=N` invocation:

```python
key = streak_key(SCOPE)                     # normalized: trim + collapse whitespace
ledger = read(".claude/chain/verify-streak.json") or new_ledger(key, N)
if ledger.scope != key or ledger.target != N:
    ledger = new_ledger(key, N)            # scope/target change → fresh streak

verdict = run_full_verification()          # the normal 8-phase verify, UNCHANGED

if verdict == "READY FOR MERGE":
    ledger.current += 1                    # extend the streak
else:
    if ledger.current > 0: ledger.reset_count += 1
    ledger.current = 0                     # ANY non-ready verdict breaks it
ledger.history.append({ts, verdict, composite, blocker?})
ledger.met = ledger.current >= ledger.target
ledger.last_run_ts = now_iso()             # stamp THIS run — the freshness proof
write_atomic(".claude/chain/verify-streak.json", ledger)   # tmp + rename, never in-place
```

**Atomic write (concurrency safety).** The ledger *is* the counter, so a torn or last-writer-wins write corrupts the whole feature. Two verify runs on the same scope (e.g. parallel worktrees) must not race: write to `verify-streak.json.tmp.&lt;pid&gt;` then `rename()` over the target (an atomic filesystem op on POSIX). Never mutate the file in place. If two runs still interleave, rename-last-wins loses at most one increment — it never leaves a half-written ledger.

**Reset rule:** a streak breaks on *any* non-`READY FOR MERGE` verdict — a tripped dimension blocker, a failing test, or an IMPROVEMENTS-RECOMMENDED. One red zeroes the count. There is no partial credit.

**Independence rule (the whole point):** each run must re-execute the *actual* tests — no cached results, no "already passed last turn." A streak over cached runs proves nothing. If the suite is fast, run it fresh each turn; if it is slow, that cost is the price of trusting the green.

## Verdict mapping

The streak gate sits *above* the normal verdict — it never loosens a blocker, it only withholds "done" until the streak is met:

| Streak state | Reported verdict |
|---|---|
| this run not READY (blocker/fail) | the normal verdict (BLOCKED / IMPROVEMENTS RECOMMENDED) + `streak reset to 0/N` |
| READY but `current &lt; target` | **STREAK PROGRESS — `current`/`target`** (not done; run again) |
| READY and `current >= target` | **READY FOR MERGE** (streak `target`/`target` met) |

Always surface the count: `STREAK 2/3 — one more green to merge` or `streak reset to 0/3 (security 3.2 &lt; 4.0)`. The user must see how close (or how broken) the streak is.

## Wiring into `/goal`

The streak gate is what makes the quality-streak recipe converge. The `until`-clause reads the ledger; each loop turn runs verify:

```
# Stamp the loop start; honor met only for a ledger written AFTER it.
LOOP_START="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
/goal until jq -e --arg t "$LOOP_START" '.met==true and .last_run_ts >= $t' .claude/chain/verify-streak.json, or stop after 15 turns
```

`no_progress_for_4_turns` is deliberately generous: a streak that keeps resetting *is* progress information (it's surfacing real flakiness), so give it room before aborting.

**Stale-ledger guard (the first-run race).** `/goal` evaluates the `until`-clause at the *top* of each turn — before that turn's verify runs. If a previous completed streak for the same scope left `met:true` in the ledger, a bare `.met==true` check exits on turn 1 having run **zero** fresh verifications. Defense (robust form): the `until`-clause compares `last_run_ts` against the loop-start timestamp, so `met:true` is honored **only when the ledger was written this loop** — you never trust a `met` you didn't just produce. This needs no `rm` and is safe even if a stale ledger exists (the old `rm -f` reset still works as a simpler fallback, but the timestamp guard is preferred because it also survives a mid-loop scope reuse). ISO-8601 UTC timestamps compare correctly as strings.

## Reuse in `/ork:cover`

`cover` already auto-heals up to 3 iterations. The same ledger + reset protocol applies: after generating tests, require the suite to pass N times consecutively before declaring coverage done — this catches flaky *generated* tests before they land. Same `verify-streak/1.0` ledger, keyed by the cover scope. (Cover wiring is a follow-up; the protocol here is the shared contract.)

## Anti-patterns

| Bad | Why it fails | Good |
|---|---|---|
| Count cached "passes" toward the streak | Proves the cache is green, not the code | Re-run the real suite each turn |
| Let an IMPROVEMENTS-RECOMMENDED extend the streak | Streak then means "good enough once", not "green N times" | Only `READY FOR MERGE` extends; everything else resets |
| Streak target of 1 | Identical to single-shot verify — no flakiness defense | N ≥ 2 (3 is the sensible default) |
| Raise the streak target to force a pass | Gaming in reverse — moving the goalposts | Target is set once, up front, from intent/policy |


### Verification Checklist

# Verification Checklist

Pre-flight checklist for comprehensive feature verification with parallel agents.

## Pre-Verification Setup

### Context Gathering
- [ ] Run `git diff main --stat` to understand change scope
- [ ] Run `git log main..HEAD --oneline` to see commit history
- [ ] Identify affected domains (backend/frontend/both)
- [ ] Check for any existing failing tests

### Task Creation (CC 2.1.16)
- [ ] Create parent verification task
- [ ] Create subtasks for each agent domain
- [ ] Set proper dependencies if needed

## Agent Dispatch Checklist

### Required Agents (Full-Stack)
| Agent | Launched | Completed | Status |
|-------|----------|-----------|--------|
| code-quality-reviewer | [ ] | [ ] | Pending |
| security-auditor | [ ] | [ ] | Pending |
| test-generator | [ ] | [ ] | Pending |
| backend-system-architect | [ ] | [ ] | Pending |
| frontend-ui-developer | [ ] | [ ] | Pending |

### Optional Agents (Add as Needed)
| Condition | Agent | Launched |
|-----------|-------|----------|
| AI/ML features | llm-integrator | [ ] |
| Performance-critical | frontend-performance-engineer | [ ] |
| Database changes | database-engineer | [ ] |

## Quality Gate Checklist

### Mandatory Gates
| Gate | Threshold | Actual | Pass |
|------|-----------|--------|------|
| Test Coverage | >= 70% | ___% | [ ] |
| Security Critical | 0 | ___ | [ ] |
| Security High | &lt;= 5 | ___ | [ ] |
| Type Errors | 0 | ___ | [ ] |
| Lint Errors | 0 | ___ | [ ] |

### Code Quality Gates
| Check | Status |
|-------|--------|
| No console.log in production | [ ] |
| No `any` types | [ ] |
| Exhaustive switches (assertNever) | [ ] |
| Proper error handling | [ ] |
| No hardcoded secrets | [ ] |

### Frontend-Specific Gates (if applicable)
| Check | Status |
|-------|--------|
| React 19 APIs used | [ ] |
| Zod validation on API responses | [ ] |
| Skeleton loading states | [ ] |
| Prefetching on links | [ ] |
| WCAG 2.1 AA compliance | [ ] |

### Backend-Specific Gates (if applicable)
| Check | Status |
|-------|--------|
| REST conventions followed | [ ] |
| Pydantic v2 validation | [ ] |
| RFC 9457 error handling | [ ] |
| Async timeout protection | [ ] |
| No N+1 queries | [ ] |

## Evidence Collection

### Required Evidence
- [ ] Test results with exit code
- [ ] Coverage report (JSON format)
- [ ] Linting results
- [ ] Type checking results
- [ ] Security scan results

### Optional Evidence
- [ ] E2E test screenshots
- [ ] Performance benchmarks
- [ ] Bundle size analysis
- [ ] Accessibility audit

## Report Generation

### Report Sections
- [ ] Summary (READY/NEEDS ATTENTION/BLOCKED)
- [ ] Agent Results (all 5 domains)
- [ ] Quality Gates table
- [ ] Blockers list (if any)
- [ ] Suggestions list
- [ ] Evidence links

### Final Steps
- [ ] Update all task statuses to completed
- [ ] Store verification evidence in context
- [ ] Generate final report markdown

## Quick Reference: Agent Prompts

### code-quality-reviewer
Focus: Lint, type check, anti-patterns, SOLID, complexity

### security-auditor
Focus: Dependency audit, secrets, OWASP Top 10, rate limiting

### test-generator
Focus: Coverage gaps, test quality, edge cases, flaky tests

### backend-system-architect
Focus: REST, Pydantic v2, RFC 9457, async safety, N+1

### frontend-ui-developer
Focus: React 19, Zod, exhaustive types, skeletons, prefetch, a11y

## Troubleshooting

### Agent Not Responding
1. Check if agent was launched with `run_in_background=True`
2. Verify agent name matches exactly
3. Check for context window limits

### Tests Failing
1. Run tests locally first
2. Check for missing dependencies
3. Verify test database state
4. Look for timing-dependent tests

### Coverage Below Threshold
1. Identify uncovered files
2. Check for excluded patterns
3. Focus on critical paths first


### Verification Manifest

# Verification Manifest (VERIFIED vs CLAIMED)

The report's agent scores, tool summaries, and every "X is fixed / clean / passing"
sentence are **claims** until the lead independently re-runs the proof. The single
most common verification failure is relaying an agent's assertion as a fact — a
sub-agent reports "tsc clean", the lead copies that into the report, it ships, and it
does **not** work.

The **Verification Manifest** is a provenance ledger that closes that seam. For every
*load-bearing* claim in the run, the lead records whether it was independently
**VERIFIED** (with the exact command + output), merely **CLAIMED** (asserted by an
agent / tool / doc, never re-run), or **UNCHECKED** (assumed). It is the operational
form of the [Verification Gate](../../shared/rules/verification-gate.md) and
[Anti-Sycophancy](../../shared/rules/anti-sycophancy.md) rules — turning "no completion
claims without fresh evidence" from a principle into a filled-in table.

## Definitions

**Load-bearing claim** — one where flipping it false would change the verdict, ship a
bug, or invalidate the merge. You do **not** manifest every trivial statement; you
manifest the claims the merge *rests on*. If in doubt, it's load-bearing.

**Provenance states**

| State | Meaning | Row must include |
|-------|---------|------------------|
| ✅ **VERIFIED** | The **lead** ran the proof **fresh this session** and read the output. | The command · exit code · the key output line |
| 🟡 **CLAIMED** | A sub-agent, tool summary, doc, or memory asserted it — **not** independently re-run by the lead. | Who asserted it |
| ⬜ **UNCHECKED** | Assumed true; nobody ran a proof (e.g. "no other caller depends on this" with no grep). | Why it was assumed |
| ⚪ **WAIVED** | A CLAIMED/UNCHECKED item deliberately accepted as non-blocking. | A one-line reason + issue ref if deferred |

> An agent reporting "PASS" and the lead **copying** that into the manifest is still
> 🟡 CLAIMED — not ✅ VERIFIED. VERIFIED means *the lead* ran it. The distinction is the
> whole point.

## The rule (wired to the verdict)

> **Any load-bearing claim that is 🟡 CLAIMED or ⬜ UNCHECKED caps the verdict at
> IMPROVEMENTS RECOMMENDED — it cannot grade READY FOR MERGE — until it is ✅ VERIFIED
> or explicitly ⚪ WAIVED with a reason.**

This sits *alongside* the dimension-level blockers (see `grading-rubric.md`): a strong
composite score never launders an unverified load-bearing claim into "done." Both gates
must clear. Under `--streak=N`, a run carrying any load-bearing CLAIMED/UNCHECKED item is
not READY and therefore **resets the streak to 0** (consistent with the existing reset
rule).

## How to build it (a synthesis step, after agents return, before the verdict)

1. **Enumerate claims** from three sources:
   - **a. Agent output** — every sub-agent `approval` / score / "PASS" / "clean" /
     "no X found" is a CLAIM by default.
   - **b. The working narrative** — every "X is fixed / passing / done" sentence.
   - **c. Premises the change rests on** — facts pulled from a doc, memory, or prior
     session ("the migration already ran", "the endpoint returns Y"). Docs and memory are
     Tier 2–4 (see the context-precedence rule): **CLAIMED until re-checked at HEAD.**
2. **For each, decide: can I cheaply run the proof now?**
   - Yes → run it **fresh**, capture `command · exit · key line` → ✅ VERIFIED.
   - No → 🟡 CLAIMED / ⬜ UNCHECKED; if load-bearing, it caps the verdict (or ⚪ WAIVE it
     with a reason).
3. **Spend verification on the load-bearing few.** You need not re-run *everything* — you
   need to never *present* a CLAIMED item *as* VERIFIED.

## Template

```markdown
## Verification Manifest

| # | Load-bearing claim | Asserted by | Provenance | Evidence (cmd · exit · key line) |
|---|--------------------|-------------|------------|----------------------------------|
| 1 | Types check clean (web) | frontend-ui-developer | ✅ VERIFIED | `pnpm --filter web type-check` · exit 0 · 0 errors |
| 2 | Unit suite green | test-generator | ✅ VERIFIED | `uv run pytest tests/unit -q` · exit 0 · 214 passed |
| 3 | No N+1 in order_service | backend-system-architect | 🟡 CLAIMED | agent asserted; not re-run → **caps verdict** |
| 4 | Migration a2f1b already applied | memory (2026-07-01) | ⬜ UNCHECKED | Tier-2 premise; re-check at HEAD before relying |
| 5 | Bundle within budget | — | ⚪ WAIVED | deferred, tracked #NNNN (non-blocking) |

**Manifest verdict impact:** rows 3 & 4 are load-bearing and unverified →
verdict capped at IMPROVEMENTS RECOMMENDED until VERIFIED or WAIVED.
```

## Anti-patterns

- **Laundering** — copying an agent's "PASS" into the table as ✅ VERIFIED without
  re-running. That's CLAIMED. VERIFIED is the *lead's* fresh run.
- **Optimism-marking** — flipping everything to ✅ to clear the gate. The manifest
  measures honesty, not confidence. "Should be fine" is ⬜, not ✅.
- **Convenient omission** — leaving a load-bearing claim off the table to dodge the
  verdict cap. The omission *is* the bug the manifest exists to catch.
- **Trusting agent numbers** — a price, model-id, cost, or count emitted by a sub-agent
  is CLAIMED until checked against source. (Sub-agents have fabricated off-by-1000×
  figures and retired model ids.) Central-verify before the row goes ✅.

## Effort scaling

| `/effort` | Manifest |
|-----------|----------|
| **low** (quick check) | Optional — skip unless a claim would block. |
| **medium** | Required for any claim that would block or pass the verdict. |
| **high** / **xhigh** | Full manifest, including doc/memory premises (source c). |


### Verification Phases

# Verification Phases — Detailed Workflow

## Phase Overview

| Phase | Activities | Output |
|-------|------------|--------|
| **1. Context Gathering** | Git diff, commit history | Changes summary |
| **2. Parallel Agent Dispatch** | 6 agents evaluate | 0-10 scores |
| **2.5 Visual Capture** | Screenshot routes, AI vision eval | Gallery + visual score |
| **3. Test Execution** | Backend + frontend tests | Coverage data |
| **4. Nuanced Grading** | Composite score calculation | Grade (A-F) |
| **5. Improvement Suggestions** | Effort vs impact analysis | Prioritized list |
| **6. Alternative Comparison** | Compare approaches (optional) | Recommendation |
| **7. Metrics Tracking** | Trend analysis | Historical data |
| **8. Report Compilation** | Evidence artifacts + gallery.html | Final report |

---

## Phase 2: Parallel Agent Dispatch (6 Agents)

Launch ALL agents in ONE message with `run_in_background=True` and `max_turns=25`. Pass `model=MODEL_OVERRIDE` when user specifies `--model=opus` (CC 2.1.72).

| Agent | Focus | Output |
|-------|-------|--------|
| code-quality-reviewer | Lint, types, patterns | Quality 0-10 |
| security-auditor | OWASP, secrets, CVEs | Security 0-10 |
| test-generator | Coverage, test quality | Coverage 0-10 |
| backend-system-architect | API design, async | API 0-10 |
| frontend-ui-developer | React 19, Zod, a11y | UI 0-10 |
| python-performance-engineer | Latency, resources, scaling | Performance 0-10 |

Use `python-performance-engineer` for backend-focused verification or `frontend-performance-engineer` for frontend-focused verification. See [Quality Model](quality-model.md) for Performance (0.11) and Scalability (0.09) weights.

Optionally add `monitoring-engineer` as a **conditional observability verifier** when the change touches services, handlers, background jobs, or infra (skip for pure UI/docs). It scores whether the new code is *operable in production* — structured logging on critical paths, metrics/SLIs, error/alert coverage — not just correct.

See [Grading Rubric](grading-rubric.md) for detailed scoring criteria.

### Task Tool Mode (Default)

```python
# PARALLEL — All 6 in ONE message
Agent(
  subagent_type="ork:code-quality-reviewer",
  model=MODEL_OVERRIDE,  # None inherits default; "opus" for thorough verification (CC 2.1.72)
  prompt="""# Cache-optimized: stable content first (CC 2.1.72)
  Verify code quality. Score 0-10.
  Check: lint errors, type coverage, cyclomatic complexity, DRY, SOLID.
  Budget: 15 tool calls max.
  Return: score (0-10), reasoning, evidence, 2-3 improvement suggestions.
  Feature: {feature}. Scope: ONLY review files in {scope_files}.""",
  run_in_background=True, max_turns=25
)
Agent(
  subagent_type="ork:security-auditor",
  model=MODEL_OVERRIDE,
  prompt="""# Cache-optimized: stable content first (CC 2.1.72)
  Security verification. Score 0-10.
  Check: OWASP Top 10, secrets in code, dependency CVEs, auth patterns.
  Budget: 15 tool calls max.
  Return: score (0-10), vulnerabilities found, severity ratings.
  Feature: {feature}. Scope: ONLY review files in {scope_files}.""",
  run_in_background=True, max_turns=25
)
Agent(
  subagent_type="ork:test-generator",
  model=MODEL_OVERRIDE,
  prompt="""# Cache-optimized: stable content first (CC 2.1.72)
  Verify test coverage. Score 0-10.
  Check: test existence, type matching, quality, edge cases, coverage %.
  Run existing tests and report results.
  Budget: 15 tool calls max.
  Return: score (0-10), coverage %, gaps identified.
  Feature: {feature}. Scope: ONLY review files in {scope_files}.""",
  run_in_background=True, max_turns=25
)
Agent(
  subagent_type="ork:backend-system-architect",
  model=MODEL_OVERRIDE,
  prompt="""# Cache-optimized: stable content first (CC 2.1.72)
  Verify API design and backend patterns. Score 0-10.
  Check: REST conventions, async patterns, transaction boundaries, error handling.
  Budget: 15 tool calls max.
  Return: score (0-10), pattern compliance, issues found.
  Feature: {feature}. Scope: ONLY review files in {scope_files}.""",
  run_in_background=True, max_turns=25
)
Agent(
  subagent_type="ork:frontend-ui-developer",
  model=MODEL_OVERRIDE,
  prompt="""# Cache-optimized: stable content first (CC 2.1.72)
  Verify frontend implementation. Score 0-10.
  Check: React 19 patterns, Zod schemas, accessibility (WCAG 2.1 AA), loading states.
  Budget: 15 tool calls max.
  Return: score (0-10), pattern compliance, a11y issues.
  Feature: {feature}. Scope: ONLY review files in {scope_files}.""",
  run_in_background=True, max_turns=25
)
Agent(
  subagent_type="ork:python-performance-engineer",
  model=MODEL_OVERRIDE,
  prompt="""# Cache-optimized: stable content first (CC 2.1.72)
  Verify performance and scalability. Score 0-10.
  Check: latency hotspots, N+1 queries, resource usage, caching, scaling patterns.
  Budget: 15 tool calls max.
  Return: score (0-10), bottlenecks found, optimization suggestions.
  Feature: {feature}. Scope: ONLY review files in {scope_files}.""",
  run_in_background=True, max_turns=25
)
# Conditional 7th agent — observability completeness. Spawn ONLY when the change
# touches services/handlers, background jobs, or infra (skip for pure UI/docs).
Agent(
  subagent_type="ork:monitoring-engineer",
  model=MODEL_OVERRIDE,
  prompt="""# Cache-optimized: stable content first (CC 2.1.72)
  Verify observability completeness. Score 0-10.
  Check: structured logging on critical paths, metrics/SLIs for the new surface,
  error and alert coverage, trace propagation. Skip pure UI/docs changes.
  Budget: 12 tool calls max.
  Return: score (0-10), observability gaps, instrumentation suggestions.
  Feature: {feature}. Scope: ONLY review files in {scope_files}.""",
  run_in_background=True, max_turns=20
)
```

### Agent Teams Alternative

In Agent Teams mode, form a verification team where agents share findings and coordinate scoring:

```python
# CC 2.1.178+: one implicit team per session — no TeamCreate.
# Spawn teammates directly via Agent(name=...). Requires
# CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 (set in ork.settings.json).

Agent(subagent_type="ork:code-quality-reviewer", name="quality-verifier",
     team_name="verify-{feature}", model=MODEL_OVERRIDE,
     prompt="""# Cache-optimized: stable content first (CC 2.1.72)
     Verify code quality. Score 0-10.
     When you find patterns that affect security, message security-verifier.
     When you find untested code paths, message test-verifier.
     Share your quality score with all teammates for composite calculation.
     Feature: {feature}.""")

Agent(subagent_type="ork:security-auditor", name="security-verifier",
     team_name="verify-{feature}", model=MODEL_OVERRIDE,
     prompt="""# Cache-optimized: stable content first (CC 2.1.72)
     Security verification. Score 0-10.
     When quality-verifier flags security-relevant patterns, investigate deeper.
     When you find vulnerabilities in API endpoints, message api-verifier.
     Share severity findings with test-verifier for test gap analysis.
     Feature: {feature}.""")

Agent(subagent_type="ork:test-generator", name="test-verifier",
     team_name="verify-{feature}", model=MODEL_OVERRIDE,
     prompt="""# Cache-optimized: stable content first (CC 2.1.72)
     Verify test coverage. Score 0-10.
     When quality-verifier or security-verifier flag untested paths, quantify the gap.
     Run existing tests and report coverage metrics.
     Message the lead with coverage data for composite scoring.
     Feature: {feature}.""")

Agent(subagent_type="ork:backend-system-architect", name="api-verifier",
     team_name="verify-{feature}", model=MODEL_OVERRIDE,
     prompt="""# Cache-optimized: stable content first (CC 2.1.72)
     Verify API design and backend patterns. Score 0-10.
     When security-verifier flags endpoint issues, validate and score.
     Share API compliance findings with ui-verifier for consistency check.
     Feature: {feature}.""")

Agent(subagent_type="ork:frontend-ui-developer", name="ui-verifier",
     team_name="verify-{feature}", model=MODEL_OVERRIDE,
     prompt="""# Cache-optimized: stable content first (CC 2.1.72)
     Verify frontend implementation. Score 0-10.
     When api-verifier shares API patterns, verify frontend matches.
     Check React 19 patterns, accessibility, and loading states.
     Share findings with quality-verifier for overall assessment.
     Feature: {feature}.""")

# Conditional 6th agent — use python-performance-engineer for backend,
# frontend-performance-engineer for frontend
Agent(subagent_type="ork:python-performance-engineer", name="perf-verifier",
     team_name="verify-{feature}", model=MODEL_OVERRIDE,
     prompt="""# Cache-optimized: stable content first (CC 2.1.72)
     Verify performance and scalability. Score 0-10.
     Assess latency, resource usage, caching, and scaling patterns.
     When security-verifier flags resource-intensive endpoints, profile them.
     Share performance findings with api-verifier and quality-verifier.
     Feature: {feature}.""")
```

**Team teardown** after report compilation:
```python
# After composite grading and report generation
# CC 2.1.178+: no TeamDelete — teammates wind down at turn end
# (press Ctrl+F twice to stop lingering background teammates).

# Worktree cleanup (CC 2.1.72)
ExitWorktree(action="keep")
```

> **Fallback:** If team formation fails, use standard Phase 2 Task spawns above.

> **Manual cleanup:** To stop lingering background teammates, press `Ctrl+F` twice (foreground/Stop).

---

## Phase 2.5: Visual Capture (Parallel with Phase 2)

**Runs as a 7th parallel agent** alongside the 6 verification agents. See [Visual Capture](visual-capture.md) for full details.

```python
# Launch IN THE SAME MESSAGE as Phase 2 agents
Agent(
  subagent_type="general-purpose",
  description="Visual capture and AI evaluation",
  prompt="""Visual verification capture for: {feature}
  1. Detect project type from package.json
  2. Start dev server (auto-detect framework)
  3. Discover routes (framework-aware scan)
  4. Use agent-browser to screenshot each route (max 20)
  5. Read each screenshot PNG for AI vision evaluation
  6. Score layout, accessibility, content completeness (0-10 per route)
  7. Read gallery template from ${CLAUDE_PLUGIN_ROOT}/skills/verify/assets/gallery-template.html
  8. Generate gallery.html with base64-embedded screenshots
  9. Write to verification-output/{timestamp}/gallery.html
  10. Kill dev server

  If no frontend detected, write skip notice and exit.
  If server fails to start, write warning and exit.
  Never block — graceful degradation only.""",
  run_in_background=True, max_turns=30
)
```

**Output**: `verification-output/\{timestamp\}/` folder with screenshots, AI evaluations (JSON), and `gallery.html`.

---

## Phase 4: Nuanced Grading

See [Quality Model](quality-model.md) for scoring dimensions, weights, and grade interpretation. See [Grading Rubric](grading-rubric.md) for detailed per-agent scoring criteria.

---

## Phase 5: Improvement Suggestions

Each suggestion includes effort (1-5) and impact (1-5) with priority = impact/effort. See [Quality Model](quality-model.md) for scale definitions and quick wins formula.

---

## Phase 6: Alternative Comparison (Optional)

See [Alternative Comparison](alternative-comparison.md) for template.

Use when:
- Multiple valid approaches exist
- User asked "is this the best way?"
- Major architectural decisions made

---

## Phase 8: Report Compilation

See [Report Template](report-template.md) for full format.

```markdown
# Feature Verification Report

**Composite Score: [N.N]/10** (Grade: [LETTER])

## Top Improvement Suggestions
| # | Suggestion | Effort | Impact | Priority |
|---|------------|--------|--------|----------|
| 1 | [highest] | [N] | [N] | [N.N] |

## Verdict
**[READY FOR MERGE | IMPROVEMENTS RECOMMENDED | BLOCKED]**
```


### Visual Capture

# Visual Capture — Phase 2.5

Visual verification that produces browsable screenshot evidence with AI evaluation.

## Architecture

```
Phase 2 agents (parallel)
         |
    Phase 2.5 (runs IN PARALLEL with Phase 2 agents)
         |
         v
┌─────────────────────────────────────────────────┐
│  1. Detect project type (package.json scan)      │
│  2. Start dev server (framework-aware)           │
│  3. Wait for server ready (poll localhost)        │
│  4. Discover routes (framework-aware)            │
│  5. agent-browser: navigate + screenshot each    │
│  6. Claude vision: evaluate each screenshot      │
│  7. Generate gallery.html (self-contained)       │
│  8. Stop dev server                              │
└─────────────────────────────────────────────────┘
```

## Step 1: Project Type Detection

Scan codebase to determine framework and dev server command:

```python
# PARALLEL — detect framework signals
Grep(pattern="\"next\":", glob="package.json", output_mode="content")
Grep(pattern="\"vite\":", glob="package.json", output_mode="content")
Grep(pattern="\"react-scripts\":", glob="package.json", output_mode="content")
Grep(pattern="\"vue\":", glob="package.json", output_mode="content")
Grep(pattern="\"nuxt\":", glob="package.json", output_mode="content")
Grep(pattern="\"@angular/core\":", glob="package.json", output_mode="content")
Glob(pattern="**/manage.py")
Glob(pattern="**/main.py")
Glob(pattern="**/app.py")
Glob(pattern="**/index.html")
```

### Detection Matrix

| Signal | Framework | Start Command | Default Port |
|--------|-----------|---------------|-------------|
| `"next":` in package.json | Next.js | `npm run dev` | 3000 |
| `"vite":` in package.json | Vite | `npm run dev` | 5173 |
| `"react-scripts":` | CRA | `npm start` | 3000 |
| `"vue":` + no vite | Vue CLI | `npm run serve` | 8080 |
| `"nuxt":` | Nuxt | `npm run dev` | 3000 |
| `"@angular/core":` | Angular | `npx ng serve` | 4200 |
| `manage.py` exists | Django | `python manage.py runserver` | 8000 |
| `main.py`/`app.py` + FastAPI | FastAPI | `uvicorn app:app` | 8000 |
| `index.html` only | Static | `npx serve .` | 3000 |
| None of the above | **Skip visual capture** | N/A | N/A |

### Override via Config

If `.claude/verification-config.yaml` exists with a `visual` section, use those settings instead of auto-detection.

## Step 2: Start Dev Server

```python
Bash(
  command=f"{start_command} &",
  description="Start dev server for visual capture",
  run_in_background=True
)
```

Wait for server readiness:
```python
Bash(command=f"for i in $(seq 1 30); do curl -s http://localhost:{port} > /dev/null && exit 0; sleep 1; done; exit 1",
     description="Wait for dev server to be ready (max 30s)")
```

**If server fails to start**: Skip visual capture with a warning in the report. Do NOT block verification.

## Step 3: Route Discovery

### Next.js App Router
```python
Glob(pattern="**/app/**/page.{tsx,jsx,ts,js}")
# Extract route from file path: app/dashboard/page.tsx → /dashboard
```

### Next.js Pages Router
```python
Glob(pattern="**/pages/**/*.{tsx,jsx,ts,js}")
# Exclude _app, _document, _error, api/
# Extract route: pages/about.tsx → /about
```

### React Router
```python
Grep(pattern="<Route.*path=[\"']([^\"']+)", glob="**/*.{tsx,jsx}", output_mode="content")
```

### FastAPI / Express
```python
Grep(pattern="@(app|router)\\.(get|post)\\([\"'](/[^\"']*)", glob="**/*.py", output_mode="content")
Grep(pattern="(app|router)\\.(get|post)\\([\"'](/[^\"']*)", glob="**/*.{ts,js}", output_mode="content")
```

### Fallback
If no routes discovered, screenshot just the root URL: `http://localhost:\{port\}/`

### Max Routes
Cap at **20 routes** to keep gallery manageable and generation fast. Prioritize:
1. Root `/`
2. Routes matching changed files (from Phase 1 git diff)
3. Routes with most sub-routes (likely important sections)

## Step 4: Screenshot Capture

Use `agent-browser` to navigate and screenshot each route:

```python
# For each route:
# 1. Navigate
agent-browser navigate http://localhost:{port}{route_path}
# 2. Wait for content
agent-browser wait --load networkidle
# 3. Capture (path is positional; --full = full page, not just viewport)
agent-browser screenshot --full verification-output/{timestamp}/screenshots/{idx}-{slug}.png
```

### Auth-Protected Routes

If `verification-config.yaml` specifies auth:

```python
# Login first
agent-browser navigate http://localhost:{port}/login
agent-browser fill "#email" "test@example.com"
agent-browser fill "#password" "test123"
agent-browser click "button[type=submit]"
agent-browser wait --load networkidle
# Then screenshot protected routes
```

### Viewport Options

Default: `1280x720`. If `mobile: true` in config, also capture at `375x812`.

## Step 5: AI Vision Evaluation

For each screenshot, use Claude's vision (Read tool on PNG) with a **structured evaluation prompt**:

```python
Read(file_path=f"verification-output/{timestamp}/screenshots/{filename}")
```

Then evaluate using this prompt template (include it in the visual capture agent's instructions):

```
Evaluate this screenshot of route "{route_path}" against these 6 criteria.
For EACH criterion, provide a severity (ok/warning/error) and specific observation.
Do NOT use generic "looks good" — cite what you actually see.

1. LAYOUT: Overflow, alignment, spacing, responsive grid. Check: content cut off? Overlapping elements? Scroll needed?
2. NAVIGATION: Is nav present and functional? Sidebar, breadcrumbs, TOC visible? Active state correct?
3. CONTENT: Text readable? Headings hierarchical? Data populated (not placeholder/loading)? Counts/numbers accurate?
4. ACCESSIBILITY: Contrast sufficient? Focus indicators visible? Text size adequate? Color-only information?
5. INTERACTIVITY: Buttons/links styled consistently? Hover/focus states? Forms labeled? CTAs discoverable?
6. BRANDING: Consistent with site theme? Dark/light mode correct? Typography matches design system?

Output as JSON array — exactly 6 items, one per criterion:
[{"severity": "ok|warning|error", "message": "CRITERION: specific observation with evidence"}]
Score 0-10 based on: 0 errors=9+, 1-2 warnings=7-8, errors=5-6, multiple errors=<5.
```

**Per-route evaluation output** (6+ items, never a single line):
```json
{
  "route": "/dashboard",
  "score": 7.5,
  "evaluation": [
    {"severity": "ok", "message": "LAYOUT: Content within viewport, no horizontal overflow, grid columns align properly"},
    {"severity": "ok", "message": "NAVIGATION: Sidebar present with 8 sections, 'Dashboard' correctly highlighted as active"},
    {"severity": "warning", "message": "CONTENT: Stats show '79 skills' but should be '89 skills' — stale count detected"},
    {"severity": "ok", "message": "ACCESSIBILITY: Body text ~16px on dark bg (#e6edf3 on #0d1117), contrast ratio ~13:1, passes WCAG AAA"},
    {"severity": "warning", "message": "INTERACTIVITY: Code block copy buttons present but no visible hover state change"},
    {"severity": "ok", "message": "BRANDING: Dark theme consistent, green accent (#3fb950) used for active states, monospace for code"}
  ]
}
```

### Cross-Route Summary

After evaluating all routes, synthesize a **summary** object for the gallery:

```python
# Build summary from all per-route evaluations
summary = {
  "total_routes": len(routes),
  "avg_score": round(sum(r.score for r in routes) / len(routes), 1),
  "pass_count": len([r for r in routes if r.score >= 7]),
  "warn_count": len([r for r in routes if 5 <= r.score < 7]),
  "fail_count": len([r for r in routes if r.score < 5]),
  "common_issues": [  # Issues appearing on 2+ routes
    {"count": 3, "message": "Stale skill count (79 instead of 89) on 3/5 pages"},
    {"count": 2, "message": "Code block copy buttons lack hover state feedback"}
  ],
  "strengths": [  # Positive patterns across routes
    "Consistent dark theme and typography across all pages",
    "Sidebar navigation present and correctly highlights active page"
  ]
}
```

Include this summary in `GALLERY_JSON` alongside `routes`.

## Step 6: Gallery Generation

Read the gallery template:
```python
Read(file_path="${CLAUDE_PLUGIN_ROOT}/skills/verify/assets/gallery-template.html")
```

Build the `GALLERY_JSON` data structure:
```json
{
  "branch": "feat/new-feature",
  "date": "2026-03-10",
  "timestamp": "2026-03-10T14:30:00Z",
  "compositeScore": 8.2,
  "visualScore": 7.8,
  "routes": [
    {
      "id": "homepage",
      "name": "Homepage",
      "path": "/",
      "screenshot": "data:image/png;base64,...",
      "score": 8.5,
      "evaluation": [
        {"severity": "ok", "message": "Layout consistent"},
        {"severity": "warning", "message": "Hero image loading slowly"}
      ],
      "apiResponse": null
    }
  ]
}
```

**Base64 encoding**: Convert each PNG to base64 for self-contained HTML:
```bash
base64 -i screenshots/01-homepage.png
```

**Size guard**: If total HTML > 10MB, use `maxDiffPixelRatio` compression or reduce to top 10 routes.

Write the final gallery:
```python
Write(file_path=f"verification-output/{timestamp}/gallery.html", content=rendered_html)
```

## Step 7: Cleanup

```python
# Kill dev server
Bash(command="kill $(lsof -ti :PORT) 2>/dev/null || true", description="Stop dev server")
```

## Graceful Degradation

| Failure | Behavior |
|---------|----------|
| No frontend detected | Skip visual capture, log info in report |
| Dev server won't start | Skip visual capture with warning |
| agent-browser unavailable | Skip screenshots, try `curl` for API-only |
| Screenshot fails on a route | Skip that route, continue with others |
| Base64 output too large | Compress or reduce route count |
| Auth flow fails | Skip protected routes, screenshot public only |



---

## Checklists (1)

### Verification Checklist

# Verification Checklist

Quick checklist for comprehensive feature verification.

## Grading Complete

- [ ] All 5 dimensions rated (0-10 scale)
- [ ] Weights applied correctly (20/25/20/20/15)
- [ ] Composite score calculated
- [ ] Grade letter assigned (A+ to F)

## Evidence Collected

- [ ] Test results with exit codes
- [ ] Coverage report (JSON)
- [ ] Security scan results
- [ ] Lint/type check output
- [ ] Evidence files linked in report

## Improvements Documented

- [ ] Each suggestion has effort estimate (1-5)
- [ ] Each suggestion has impact estimate (1-5)
- [ ] Priority calculated (Impact / Effort)
- [ ] Quick wins identified (low effort, high impact)

## Alternatives Considered

- [ ] Current approach scored
- [ ] At least one alternative evaluated
- [ ] Migration cost estimated
- [ ] Recommendation documented

## Policy Compliance

- [ ] No blocking rule violations
- [ ] Warning rules acknowledged
- [ ] Thresholds checked (composite, security, coverage)

## Report Generated

- [ ] All sections filled
- [ ] Verdict assigned (Ready/Recommended/Blocked)
- [ ] Tasks updated to completed
