Route by pipeline stage, not by a single favourite model: Claude Haiku 4.5 for high-volume PR triage, Claude Sonnet 5 as the default deep-review pass, and Claude Opus 5, escalating to Claude Fable 5.1 for the highest-stakes diffs, as the final gate on architecture, security and schema changes.

Last verified 7 September 2026 against the Claude 5 family: Claude Fable 5.1, Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5. If a newer model has shipped since you are reading this, treat the roles below as still directionally correct and re-check the specifics before you wire them into CI.

Who this is for, and who it is not for

This is a routing guide for engineering teams putting an LLM into a code review pipeline: what to run at each stage, how the four current Claude models trade off against each other, and where a human still has to read the diff. It assumes you already have, or can add, a CI step that calls the Claude API. It does not assume any specific review-bot vendor; the routing logic holds whether you call the API directly or through a tool that wraps it.

Good fit

  • Teams running an automated first-pass review before a human looks at a PR.
  • Teams that already have, or are willing to write, a short style guide the model can read alongside the diff.
  • Teams building or buying an agent loop that generates code and need a second, independent model checking its output.
  • Teams that want a written, defensible routing policy instead of one engineer picking whichever model they used last.

Not a good fit

  • Teams looking for a single model to replace human review entirely. None of the four models here should be the last set of eyes on a change that matters.
  • Teams with no repository conventions written down anywhere. A model with no style guide produces generic best-practice noise regardless of tier, and the noise trains reviewers to stop reading the comments.
  • Teams expecting the model choice alone to fix a review culture that already rubber-stamps PRs. That failure is organisational before it is technical, and swapping models will not touch it.

The four Claude 5 models, compared

Stated at the level a routing decision actually needs: relative cost and speed tier within the family, not a specific price or millisecond figure, since both change independently of this article and we are not going to guess at either.

The Claude 5 family in a code review pipeline, by role rather than benchmark
ModelBest roleCost tierSpeed tierContext to provideDo not use for
Claude Haiku 4.5First-pass triage on every pushLowestFastestThe diff plus a short written style guideArchitecture review, security-sensitive diffs, or as the only reviewer before a merge
Claude Sonnet 5Default deep-review pass once triage has cleared a PRMidMidThe diff, the style guide and the files it touchesFinal sign-off on schema migrations, auth changes or anything you would escalate to a senior engineer
Claude Opus 5Final gate on large diffs, architecture changes, security and infrastructureHighestSlowestAs much repository context as fits, plus related files beyond the diffRoutine style and lint-level comments on every PR
Claude Fable 5.1Escalation when Opus 5 flags low confidence or the change is unusually high-stakes (payments, data deletion, an auth rewrite)Highest, at or above Opus 5SlowestEverything Opus 5 gets, plus the design or incident history behind the changeRoutine review of any kind

Model ids, for the pipeline config: `claude-haiku-4-5-20251001`, `claude-sonnet-5`, `claude-opus-5` and `claude-fable-5-1`. Fable 5.1 is the most capable generally available model in the family and the one to reserve for cases the other three genuinely cannot resolve.

Which model for which review task

Model choice should follow the task, not the other way round. Five recurring review tasks and where each one routes:

Five recurring review tasks and where each one routes
Review taskModelWhy
PR triage on every pushHaiku 4.5, with the style guide in the promptHighest volume, lowest stakes: catch the obvious and flag anything ambiguous for a heavier model or a human
Security reviewOpus 5, escalating to Fable 5.1 for authentication, secrets handling or payment flowsRewards a model that can hold more of the surrounding system in view, not just the changed lines
Architecture review of a large diffOpus 5 by defaultA large diff is where a cheaper tier starts missing cross-file patterns, the failure mode that matters most
Style and lint-level commentsHaiku 4.5, or better, a deterministic linterSpending a frontier tier on comments a linter catches for free is the first routing mistake to fix
Reviewing generated code from an agent loopSonnet 5 as the second opinion, moved up to Opus 5 when the agent touched auth, billing, data deletion or schemaA model reviewing its own output is not an independent check; the review model should not be the model that wrote the code

Why role, not benchmark, should drive the choice

It is tempting to ask which model 'scores best' at code review and stop there. Resist it. A single number collapses three different jobs, triage, deep review, and final gate, into one figure, and the job determines what good looks like far more than the model tier does. A model that is excellent at catching schema-migration risk is wasted, and expensive, running against a two-line copy change. Design the pipeline around the three jobs first, then assign a tier to each, the way our comparison pages set out the trade-offs for other tooling decisions, or our free guides set out an AI project brief.

The model choice is a routing decision, not a purchase decision: match Haiku's speed to triage volume, Sonnet's balance to the default pass, and Opus or Fable's depth to the diffs where a miss is expensive.

How much repository context to give each model

Context is the lever most teams under-use. Feeding a model an isolated diff, with no style guide, no related files and no history, produces the same generic best-practice commentary at every tier, which is why teams that skip this step conclude 'the model is not very good' when the model was never given enough to be good. As a rule of thumb: triage needs the diff and the style guide; a deep review benefits from the touched files, not just the changed lines; a final-gate review benefits from as much of the surrounding system as you can reasonably fit, because structural issues are, by definition, issues that only show up outside the diff.

Where a pipeline still needs a human

A model, at any tier, that always approves is not reviewing, it is rubber-stamping with extra steps. Keep a visible check on that: route a sample of already-approved PRs to a human on a schedule, track how often the model comments actually change a merge decision, and treat a near-zero comment rate as a signal to check the prompt, not as evidence the code is unusually clean. The four models here are a triage and drafting layer for review capacity. On anything that would need a senior engineer's sign-off today, keep the senior engineer in the loop; move their time to the diffs the models flagged as uncertain, not away from review entirely. Our write-ups on data-backed pipeline decisions cover this failure mode in more depth if you want the longer version.

Building this into your pipeline

None of this needs new tooling beyond a CI step that can call the Claude API and read a label or a path rule. If you want a second opinion on the routing policy before committing engineering time to it, built against your actual repository conventions rather than a generic template, that is the kind of spec work we do. Book time to talk it through, and bring the repository, not a slide deck.

Or skip ahead and talk through it directly