Route by pipeline stage, not by a single favourite model: Claude Haiku 4.5 for high-volume PR triage, Claude Sonnet 5 as the default deep-review pass, and Claude Opus 5, escalating to Claude Fable 5.1 for the highest-stakes diffs, as the final gate on architecture, security and schema changes.
Last verified 7 September 2026 against the Claude 5 family: Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5. If a newer model has shipped since you are reading this, treat the roles below as still directionally correct and re-check the specifics before you wire them into CI.
Who this is for, and who it is not for
This is a routing guide for engineering teams putting an LLM into a code review pipeline: what to run at each stage, how the four current Claude models trade off against each other, and where a human still has to read the diff. It assumes you already have, or can add, a CI step that calls the Claude API. It does not assume any specific review-bot vendor; the routing logic holds whether you call the API directly or through a tool that wraps it.
Good fit
- Teams running an automated first-pass review before a human looks at a PR.
- Teams that already have, or are willing to write, a short style guide the model can read alongside the diff.
- Teams building or buying an agent loop that generates code and need a second, independent model checking its output.
- Teams that want a written, defensible routing policy instead of one engineer picking whichever model they used last.
Not a good fit
- Teams looking for a single model to replace human review entirely. None of the four models here should be the last set of eyes on a change that matters.
- Teams with no repository conventions written down anywhere. A model with no style guide produces generic best-practice noise regardless of tier, and the noise trains reviewers to stop reading the comments.
- Teams expecting the model choice alone to fix a review culture that already rubber-stamps PRs. That failure is organisational before it is technical, and swapping models will not touch it.
The four Claude 5 models, compared
Stated at the level a routing decision actually needs: relative cost and speed tier within the family, not a specific price or millisecond figure, since both change independently of this article and we are not going to guess at either.
| Model | Best role | Cost tier | Speed tier | Context to provide | Do not use for |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | First-pass triage on every push | Lowest | Fastest | The diff plus a short written style guide | Architecture review, security-sensitive diffs, or as the only reviewer before a merge |
| Claude Sonnet 5 | Default deep-review pass once triage has cleared a PR | Mid | Mid | The diff, the style guide and the files it touches | Final sign-off on schema migrations, auth changes or anything you would escalate to a senior engineer |
| Claude Opus 5 | Final gate on large diffs, architecture changes, security and infrastructure | Highest | Slowest | As much repository context as fits, plus related files beyond the diff | Routine style and lint-level comments on every PR |
| Claude Fable 5.1 | Escalation when Opus 5 flags low confidence or the change is unusually high-stakes (payments, data deletion, an auth rewrite) | Highest, at or above Opus 5 | Slowest | Everything Opus 5 gets, plus the design or incident history behind the change | Routine review of any kind |
Model ids, for the pipeline config: `claude-haiku-4-5-20251001`, `claude-sonnet-5`, `claude-opus-5` and `claude-fable-5-1`. Fable 5.1 is the most capable generally available model in the family and the one to reserve for cases the other three genuinely cannot resolve.
Which model for which review task
Model choice should follow the task, not the other way round. Five recurring review tasks and where each one routes:
| Review task | Model | Why |
|---|---|---|
| PR triage on every push | Haiku 4.5, with the style guide in the prompt | Highest volume, lowest stakes: catch the obvious and flag anything ambiguous for a heavier model or a human |
| Security review | Opus 5, escalating to Fable 5.1 for authentication, secrets handling or payment flows | Rewards a model that can hold more of the surrounding system in view, not just the changed lines |
| Architecture review of a large diff | Opus 5 by default | A large diff is where a cheaper tier starts missing cross-file patterns, the failure mode that matters most |
| Style and lint-level comments | Haiku 4.5, or better, a deterministic linter | Spending a frontier tier on comments a linter catches for free is the first routing mistake to fix |
| Reviewing generated code from an agent loop | Sonnet 5 as the second opinion, moved up to Opus 5 when the agent touched auth, billing, data deletion or schema | A model reviewing its own output is not an independent check; the review model should not be the model that wrote the code |
Why role, not benchmark, should drive the choice
It is tempting to ask which model 'scores best' at code review and stop there. Resist it. A single number collapses three different jobs, triage, deep review, and final gate, into one figure, and the job determines what good looks like far more than the model tier does. A model that is excellent at catching schema-migration risk is wasted, and expensive, running against a two-line copy change. Design the pipeline around the three jobs first, then assign a tier to each, the way our comparison pages set out the trade-offs for other tooling decisions, or our free guides set out an AI project brief.
The model choice is a routing decision, not a purchase decision: match Haiku's speed to triage volume, Sonnet's balance to the default pass, and Opus or Fable's depth to the diffs where a miss is expensive.
How much repository context to give each model
Context is the lever most teams under-use. Feeding a model an isolated diff, with no style guide, no related files and no history, produces the same generic best-practice commentary at every tier, which is why teams that skip this step conclude 'the model is not very good' when the model was never given enough to be good. As a rule of thumb: triage needs the diff and the style guide; a deep review benefits from the touched files, not just the changed lines; a final-gate review benefits from as much of the surrounding system as you can reasonably fit, because structural issues are, by definition, issues that only show up outside the diff.
Where a pipeline still needs a human
A model, at any tier, that always approves is not reviewing, it is rubber-stamping with extra steps. Keep a visible check on that: route a sample of already-approved PRs to a human on a schedule, track how often the model comments actually change a merge decision, and treat a near-zero comment rate as a signal to check the prompt, not as evidence the code is unusually clean. The four models here are a triage and drafting layer for review capacity. On anything that would need a senior engineer's sign-off today, keep the senior engineer in the loop; move their time to the diffs the models flagged as uncertain, not away from review entirely. Our write-ups on data-backed pipeline decisions cover this failure mode in more depth if you want the longer version.
Building this into your pipeline
None of this needs new tooling beyond a CI step that can call the Claude API and read a label or a path rule. If you want a second opinion on the routing policy before committing engineering time to it, built against your actual repository conventions rather than a generic template, that is the kind of spec work we do. Book time to talk it through, and bring the repository, not a slide deck.