Blog

Which Claude Model Is Best for Legal Work? What Lawyers Found Testing Fable 5.1, Opus 5.5, and Sonnet 5.5

Tue 06 Oct 2026

TL;DR

  • Default: Claude Sonnet 5.5 for well-scoped everyday legal work, such as extraction and analysis of a defined set of documents. It's fast and low-cost, and gives up little on most tasks
  • Hard calls: Claude Opus 5.5 for the hardest judgment calls and for drafting
  • Breadth: Claude Fable 5.1 when a task rewards wide legal knowledge or synthesis across several documents, and a little extra time is acceptable

The gaps between the three are narrow. For all of them, a precise instruction on scope and length is the most effective way to get a better answer. These findings come from the Litera Legal Knowledge Engineering (LKE) team. The team is made up of former practicing lawyers, and they tested each model on real legal work.

Three new frontier AI models arrived in roughly four weeks. Each launch came with benchmark charts, and each promised a step up on the last. None of those charts answers the question a practice leader actually has: which one should my lawyers use on Monday morning, and for what?

Picking wrong has a real cost. A model that's brilliant at synthesis but slow can stall a quick turnaround. A model that's fast but light on citations can create review work on a document-heavy matter. When the AI landscape shifts this fast, a 12-month technology roadmap is already out of date.

How Did Litera Test These Claude Models for Legal Work?

The LKE team evaluated Claude Fable 5.1, Claude Opus 5.5, and Claude Sonnet 5.5 on real legal tasks. The team is made up of former practicing lawyers across multiple jurisdictions and practice areas.

The goal wasn't to rank models on generic benchmarks. It was to describe what each model is like to work with on legal tasks, and when to reach for which. The work tested included:

  • Locating and extracting information
  • Summarization
  • Document analysis
  • Drafting
  • General legal-knowledge questions
  • Applied legal reasoning
  • Expert judgment

The team assessed each model's default behavior at a point in time. Model behavior evolves, and individual results will vary with prompts and documents.

Which Claude Model Should Lawyers Use?

All three sit alongside Claude Opus 5 as the strongest Claude models the LKE team has evaluated, and the gaps between them are narrow. Where they differ is in the shape of their strengths, and in speed and cost.

ModelReach for it whenStrongest atWatch forSpeed and cost
Claude Sonnet 5.5The work is well scoped and quick turnaround at low cost mattersExtraction and document analysis, with the highest overall result testedTrails on general legal-knowledge questions and drafting, and cites the document less oftenThe only top-tier model faster than the field average, at the lowest cost
Claude Opus 5.5The task turns on expert judgment or needs draftingThe most judgment-heavy tasks, plus drafting, analysis, and extractionTrails on general legal knowledge and summarization, and runs long after the direct answerFaster than Fable 5.1 and Opus 5, at lower cost than Fable 5.1
Claude Fable 5.1The task draws on wide legal knowledge or synthesis across several documentsKnowledge-based questions, and separating what a document says from what it infersDelivers a fully complete answer less often than Opus 5.5, and writes the longest responsesSlower than Opus 5.5, at the highest cost of the three

Is Claude Sonnet 5.5 Good Enough for Everyday Legal Work?

Yes, and often for more than everyday work. Sonnet 5.5 is the largest generational leap in this release.

Its predecessor, Sonnet 5, trailed well behind the leading Claude models. Sonnet 5.5 posted the highest overall result of any model the LKE team has evaluated. That said, it can't be meaningfully separated from Opus 5.5, Fable 5.1, or Opus 5.

Where it leads: Sonnet 5.5 performed best of any model on extraction and document analysis, and held up well on applied legal reasoning. Opus 5.5 kept its edge on the most judgment-heavy tasks.

Why it makes a strong default: Sonnet 5.5 was the only top-tier model to complete tasks faster than the field average. It also ran faster than Sonnet 5, at a fraction of the cost of Fable 5.1 and well below Opus 5.5.

Where it gives ground:

  • General legal knowledge: it trailed Fable 5.1, Opus 5.5, and Opus 5
  • Drafting: it was less reliable than Opus 5.5
  • Style: its responses now run about as long as those of Opus 5.5 and Fable 5.1, lean heavily on bullet points, and use tables less often
  • Citations: it answered document-based questions without citing the document more often than the other leading models

If you need every conclusion tied to the source, ask for citations explicitly.

Is Claude Opus 5.5 Good for Legal Drafting and Judgment Calls?

Yes. On the most judgment-heavy tasks, Opus 5.5 performed best of any model the LKE team has evaluated. It was also among the strongest on drafting, analysis, and extraction.

Where it leads: The margins at the top are narrow, best read as an edge rather than a clear lead, but the direction was consistent. The more a task leaned on expert judgment rather than locating information, the further Opus 5.5 pulled ahead of Fable 5.1. It also completed tasks somewhat faster than both Fable 5.1 and Opus 5, at a lower cost than Fable 5.1.

Where it gives ground: Its gains over Opus 5 were uneven. They concentrated in the hardest reasoning and in drafting. On general legal-knowledge questions and summarization, Opus 5.5 trailed both Opus 5 and Fable 5.1. For a question that turns on what the law is, rather than what a document says, Fable 5.1 is the stronger pick.

Style: Opus 5.5 reliably leads with a short, direct answer and explains legal concepts in plain language a lawyer can take in quickly. What follows often runs long. In legal testing, it spent as much text per required point as any model evaluated. A short, explicit instruction on length is the right first step.

When Should Lawyers Use Claude Fable 5.1?

Use Fable 5.1 when a task draws on wide legal knowledge rather than a single document, or when an answer has to be pulled together from several sources. Those are also where it improved most over Fable 5, with knowledge-based questions showing the clearest gain.

Where it leads: Fable 5.1 paired that breadth with care about where its answers come from. It consistently separated what a document actually says from what it was inferring.

Where it gives ground:

  • Completeness: on the most judgment-heavy tasks and on drafting, it covered the individual points a strong answer requires as reliably as any model evaluated. But it delivered a fully complete answer less often than Opus 5.5. The shortfall was rarely an error. More often, a sound answer was one point short. That makes Fable 5.1 an excellent first draft on complex work, worth checking against the original ask
  • Length and speed: Fable 5.1 writes more by default than any Claude model before it. It often reads a narrow question as an invitation to review the whole document, and its responses took somewhat longer than those of Opus 5.5

For complex questions, that careful answer is the trade Fable 5.1 is built for. For quick lookups, tell it to lead with the answer and keep it short.

Does a More Expensive AI Model Give Better Legal Answers?

Not reliably. In the LKE team's testing, paying more for a higher-tier model did not reliably buy a better result on legal work.

Sonnet 5.5, the fastest and lowest-cost model of the three, posted the highest overall result. It sits in a statistical tie with the more expensive models. The right model depends on the shape of the task, not the price tier.

How Should Lawyers Prompt AI Models for Legal Work?

Three habits improve results from every model:

  • Say what to leave out: all three models run long by default, and the extra material rarely makes an answer more likely to be right. For a narrow deliverable, say so directly, and name what to exclude, not only what to include
  • State your jurisdiction: when a task doesn't specify one, models tend to default to US law without flagging the assumption
  • Ground it in current sources: where a task turns on the current state of the law, provide source material rather than relying on the model's recall

How Should Lawyers Review AI Output on Legal Work?

Fable 5.1, Opus 5.5, and Sonnet 5.5 are among the strongest models evaluated for legal work. Even so, they carry the same limitations as any large language model (LLM). Three review habits still apply:

  • Check for the missing point: even the strongest current models satisfy nearly every individual point a legal answer needs, yet fall short of a fully complete answer far more often. The gap is usually missing content rather than wrong content, which makes it easy to overlook. Read the first response as a strong draft, check it against every part of the original ask, and use a follow-up prompt to fill any gap
  • Take more care as the task gets harder: every model performs strongly on locating, extracting, and summarizing information. Performance drops for all of them on applied legal reasoning and expert judgment. The more a task depends on judgment, the more closely its conclusions deserve review
  • Verify assertions on nuanced points of law: confirm that legal conclusions are grounded in the cited source rather than general knowledge. Take particular care where a statement about how the law works sits next to a citation to the document, which may only establish which law governs

What Is Lito?

Whichever model a firm uses, the LKE findings point to the same conclusion: even the strongest models need grounding, structure, and review on legal work. That's the problem Lito, Litera's award-winning Legal AI agent, is built to solve.

Lito is a Legal AI agent for the practice and business of law. It combines leading foundational AI models with rules-based engines Litera has built over 30 years, skills curated by Litera in-house lawyers, and the firm's own data. Lito runs inside Word, Outlook, on the web, and on iOS, and it's included in a Litera subscription with unlimited usage.

The design principle is simple: exact where the work can't be wrong, and flexible where lawyers need breadth. Litera is trusted Legal AI that best unifies the practice and business of law, and Lito is how lawyers put that platform to work.

How Does Lito Combine AI Models With Deterministic Engines?

Lito matches each part of a task to the technology best suited to it. Leading AI models handle the reasoning, and deterministic, rules-based engines establish the facts on high-stakes work, where an answer that varies from one run to the next presents too much risk. Lito brings four layers together:

LayerWhat it doesExample in Lito
Leading foundational modelsOpen-ended questions, drafting, analysis, and custom workflowsFoundational AI models available in Lito, with new models evaluated by the LKE team on real legal work
Deterministic enginesProduce exact, repeatable results on work that can't be wrongThe Litera comparison engine, which produces more than 10 million redlines a month
Lawyer-built skillsRun structured, repeatable legal tasks with prompts curated by Litera in-house lawyersA library of 30+ skills, including NDA Playbook Review, Analyze in Grid, and Mitigate Risk
Firm dataGround answers in the firm's own matters, relationships, and experienceFirm AI Search surfaces the right matter, relationship, and credentials

Here's how those layers work together on a counterparty draft:

  • The engine establishes the facts: the Litera comparison engine produces a deterministic redline of every change
  • The model reasons over those facts: the AI model explains what changed, and the Mitigate Risk skill flags the changes that carry risk and suggests mitigation strategies and rewrites
  • The lawyer decides: suggested language enters the document only when the lawyer chooses to insert it

Because the model reasons over an exact record rather than its own reading of the documents, lawyers spend their review time on judgment instead of re-checking what changed.

How Does Lito Build Good AI Habits Into Legal Work?

The LKE report shows that even the strongest models need grounding and review on legal work. Lito builds those habits into the workflow, so lawyers don't have to remember them on every prompt:

  • Grounded in exact results: on high-stakes steps like comparison, the model works from deterministic engine output, so conclusions can be verified against the source
  • Structured by lawyers: skills curated by Litera in-house lawyers define what a complete answer looks like for common tasks, which helps close the "missing point" gap
  • Encoded firm practice: Lito Studio lets lawyers turn their own best practices into skills and templates, built on the foundational models available in Lito
  • Tested by lawyers: the LKE team evaluates new foundational models on real legal work and publishes guidance like this report on what to expect from each

What Can Lawyers Do With Lito?

Lito covers both the practice and the business of law. Common workflows include:

  • Compare and mitigate risk: compare documents conversationally, then flag risky changes and get suggested rewrites
  • Manage counterparty changes in Word: get a recommendation and reasoning for every tracked change and comment
  • Review against a playbook: check NDAs against playbook standards with plain-language prompts
  • Analyze document sets: extract key information across many documents into a structured grid
  • Search firm intelligence: find the right matter, relationship, and credentials to win the next piece of work
  • Build custom workflows: automate firm best practices with Lito Studio

"Lito helped me review and compare over 40 leases… what would have taken five hours took 30–45 minutes." — Brian Corbett, Partner, Poyner Spruill LLP

How Do Lawyers Build Their Own Skills in Lito Studio?

Lito Studio is a no-code builder that lets lawyers, innovation teams, and knowledge managers create their own skills and multi-document review templates in a matter of minutes. Lawyers build skills on the foundational models available in Lito, with no prompt engineering required.

Building a skill takes six steps:

  1. Describe the task: explain what you need in plain language
  2. Get a structured prompt: Lito turns your description into a structured prompt you can edit
  3. Choose the model: select from the foundational models available in Lito
  4. Test it: preview and test the skill or template before you save it
  5. Save it: add a name, description, and category
  6. Share it: use it yourself, or request a firm-wide rollout that an admin approves and publishes with role-based permissions

Match the model to the task

The LKE findings carry straight into Studio. Models differ in the shape of their strengths, so the best skills pair each task with a model suited to it:

Skill or template you buildWhat to look for in the modelWhat LKE testing found
A grid template that extracts purchase price, completion date, and limitation of liability across a set of share purchase agreementsSpeed and low cost on well-scoped document workFaster, lower-cost models can now match top-tier results on extraction and document analysis
A negotiation skill that returns a clause-by-clause table of buyer-favorable changes and the rationale for eachStrength on expert judgment and draftingJudgment-heavy tasks and drafting are where models differ most, and where review matters most
A skill that answers questions drawing on wide legal knowledge or synthesizes several documents into one answerBreadth of legal knowledge and synthesisModels vary most on general legal-knowledge questions, so breadth is worth testing before rollout

Studio is also where the LKE prompting habits become permanent. A lawyer can build scope, length, jurisdiction, and citation instructions into a skill once, and every run after that follows them.

What Is the Return on Using the Right Model in Lito?

Choosing the right model for each task, and grounding it in deterministic engines where accuracy matters most, is how Litera delivers the full Return on AI (RoAI), not efficiency alone:

  • Efficiency Growth: the right model for each task means faster turnaround at lower cost, without paying for capability the task doesn't need
  • Relationship Growth: answers grounded in exact results, and reviewed with the right habits, give clients work they can trust
  • Business Growth: time returned to lawyers becomes capacity for more matters, and firm intelligence helps turn existing relationships into new work

See How Lito Puts the Right Model to Work

The best way to judge any Legal AI is on the work your lawyers do every day. Explore Lito, or let's talk about how your team can match the right model to each task.

Frequently Asked Questions About Claude Models and Lito

What is Lito?

Lito is Litera's award-winning Legal AI agent for the practice and business of law. It combines leading AI models with rules-based engines built over 30 years, skills curated by Litera in-house lawyers, and the firm's own data, and it runs inside Word, Outlook, on the web, and on iOS.

Which AI models does Lito use?

Lito uses leading foundational AI models alongside deterministic, rules-based engines. The Litera Legal Knowledge Engineering team evaluates new foundational models on real legal work and publishes guidance on what to expect from each.

Does Lito use deterministic engines or generative AI?

Both. Lito uses deterministic, rules-based engines, such as the Litera comparison engine, for high-stakes work where results must be exact and repeatable, and leading AI models for open-ended reasoning, drafting, and analysis.

What is Lito Studio?

Lito Studio is a no-code builder in Lito that lets lawyers, innovation teams, and knowledge managers create custom skills and multi-document review templates in minutes, built on the foundational models available in Lito, with no prompt engineering required.

Can I choose which AI model a Lito skill uses?

Yes. When you build a skill in Lito Studio, you choose from the foundational models available in Lito.

Do lawyers need prompt engineering experience to build a skill?

No. Lawyers describe the task in plain language, and Lito turns it into a structured prompt they can test, edit, and save.

Is Lito included in a Litera subscription?

Yes. Lito is included in a Litera subscription with unlimited usage.

Which Claude model is best for legal drafting?

Claude Opus 5.5. In LKE testing, it was among the strongest models on drafting and the most reliable of the three when a draft needs to be right the first time. Sonnet 5.5 was less reliable on drafting than Opus 5.5.

Which Claude model is best for contract extraction and document analysis?

Claude Sonnet 5.5. It performed best of any model the LKE team has evaluated on extraction and document analysis, and it does so faster and at lower cost than the other top-tier models.

Which Claude model is best for questions about what the law is?

Claude Fable 5.1. For general legal-knowledge questions, the ones that turn on what the law is rather than what a document says, Fable 5.1 is the stronger pick. Both Opus 5.5 and Sonnet 5.5 trailed it on these questions.

Do AI models assume US law?

Often, yes. When a task doesn't specify a jurisdiction, models tend to default to US law without flagging the assumption. State your jurisdiction in the prompt.

Why are AI answers to legal questions so long?

All three models run long by default, and the extra material rarely makes an answer more likely to be right. For a narrow deliverable, ask the model to lead with the answer, keep it short, and name what to leave out.

Can lawyers rely on AI output as final work product?

Treat the first response as a strong draft. Even the strongest models often fall one point short of a fully complete answer. Check the response against every part of the original ask, review conclusions more closely as tasks get harder, and verify assertions on nuanced points of law against the source.


Artificial Intelligence
Share on TwitterShare on FacebookShare on LinkedIn
On-Demand Webinar

From Back Office to Backbone: How Shell Modernized Records Governance

Retention, legal holds, and privacy requests too often run in separate lanes. Watch how Shell connected them across roughly 17 million...
Read More
Blog

How Do You Compare a PDF to a Word Document Without Converting It?

TL;DR The most reliable way to compare a PDF to a Word document is with a comparison engine that reads both formats directly and produces...
Read more
Blog

How Did Litera Rank in ILTA’s 2026 Technology Survey?

In ILTA's 2026 Technology Survey, Litera products ranked first or led among named vendor products across high-value workflows on both the...
Read more

Ready to get started?

Join over 4,000+ firms already growing with Litera.