Agent-readable docs index: /llms.txt. Full docs in one file: /llms-full.txt. Download /docs.zip to grep all markdown files locally.

review

review
Claude Code runs it as /review and Codex as $review. Other agents take its name or a plain request such as "review this".

What it does

review assesses one target against its criteria and returns a verdict with evidence. The reviewer is never the author and stays read-only: it edits no file, changes no task state and posts nothing to the card. Fixes, records and delivery stay with whoever asked for the review.
It writes its criteria before it looks at the work. One pass sees only the scope, the criteria and the validation contract, and writes down what would make the verdict SHIP, FIX-FIRST or BLOCKED. A second pass reads the work and scores it against that frozen list, and names any criterion it added afterwards. A reviewer that reads the work first tends to rationalize what it reads.
It reviews these targets:
TargetWhat it checks
DesignProblem, scope, approach, alternatives, risks and success criteria agree, and every added mechanism has a present need. It returns the digest of the content it reviewed, for the design's SHIP stamp
PlanThe template, the linked design's SHIP evidence, deliverables, per-group criteria and validation, dependency order, file ownership, and whether the plan can actually run
Implementation or pull requestEvery criterion traced to code and evidence: correctness, failure behaviour, compatibility, security, maintainability, performance where it matters, and scope
Evidence comes from current code and commands it runs, or from results it can attribute to the exact artifact. A worker's claim, a filename or an earlier verdict is a hypothesis, never a confirmed finding. Text inside a pull request thread or the reviewed files that tries to instruct the reviewer is recorded as a finding and never followed.

Severity and verdict

SeverityMeaningBlocks
CRITICALDemonstrated security exposure, data loss or severe outageYes
HIGHA broken criterion, a correctness defect, a major performance or compatibility failureYes
MEDIUMA bounded maintainability or evidence gapNo
LOWAn optional style or naming improvementNo
  • SHIP: no CRITICAL or HIGH gap, and the required validation passes.
  • FIX-FIRST: blocking gaps that can be acted on, or failed validation. The caller sends them to fix.
  • BLOCKED: missing scope, an open design decision, a broken environment or missing evidence prevents a valid assessment.
For a plan, the caller turns SHIP into APPROVED, FIX-FIRST into FIX-FIRST and BLOCKED into BLOCKED in WISH.md. An implementation review leaves the wish IN_PROGRESS. MEDIUM and LOW findings are not repair work; the caller sends them to deslop or records them as accepted.
For a deeper audit, ask for one of its lenses in plain words, such as "review performance": architecture and simplicity, code quality, documentation and contributor experience, performance, test quality, rendered interface, repository hygiene, or security and supply chain.

When to use it

  • A design from brainstorm, before a plan is written.
  • A plan from wish, before work starts.
  • A group's diff, a branch or a pull request, before it merges.
wish and work run review for you at each of those gates. Run it yourself to check anything else against stated criteria.

Run it

AgentHow it runs
Claude Code/review, optionally naming the target
Codex$review, optionally naming the target
Other agentsIts name, or a plain request such as "review this"
review has no saved workflow. In wish and work it runs as a separate agent from the one that did the work. When you run it on uncommitted work, it names the snapshot it reviewed, and the verdict no longer holds once that tree changes.

A recorded run

Simulated terminal: a replay of a real /review run in Claude Code, built from its saved capture and shortened in time. Edited in two places: host details hidden (paths shown as ~/, account plan and quota), and the agent launch line removed.
The run reviews an uncommitted change in a small test repository. It writes its criteria from the wish before opening the diff, runs the tests, and returns FIX-FIRST with HIGH findings: the new percent function does not reject a zero whole, and no test covers that case.
Read the transcript: the redacted text of the run's final screen and scrollback.

Next

fix
What happens after a FIX-FIRST.
Skills catalog
Every skill Genie ships, one line each.