Model highlight

Claude Sonnet 5.5 for Chemical Companies: First on Regulations, in a Crowd

Claude Sonnet 5.5 tops an independent test of regulatory research, but the test cannot separate it from 22 other models, some at a hundredth of the price. For a chemical distributor or producer, the question becomes which model is good enough for the price, not which one is first.

69.7%of US regulatory research questions answered fully right, first of 74 models
Anthropic · Released 28 Sep 2026Vailent model guide · Updated 5 min read

Claude Sonnet 5.5 at a glance

Made by
Anthropic5
Released
28 Sep 20265
Price
$2 per million tokens sent, $10 per million written6
Reads
Text and images, up to 1 million tokens3
Settings
Five, from low to max3
Index
56, 2nd on the Artificial Analysis Intelligence Index3
Regulatory
1st of 74, tied with 221

What Claude Sonnet 5.5 is.

Anthropic released Claude Sonnet 5.5 on 28 September 2026. It reads text and images, holds up to 1 million tokens, and costs $2 for every million tokens you send and $10 for every million it writes, the same as Sonnet 5.

Anthropic says it “runs 30% faster and costs up to 30% less for most work” than Sonnet 5. An independent test found otherwise on cost: at its max setting it used more tokens than any model the testers have measured, and a task cost about 50% more than on Sonnet 5.

First on regulatory research, level with 22 others.

vals.ai's test asks 66 US regulatory research questions, and an answer counts only if it meets every point lawyers listed. Sonnet 5.5 is first. Every model below it here scores within the test's margin of error. Sort by price to see what the same result costs.

22 results the test cannot separate from the leader, on 66 questions. The cheapest of them, GLM-5.3 Flash, costs 128 times less a question.

  1. Claude Sonnet 5.569.7%$12.64
  2. Gemini 4 Argon68.18%$4.81
  3. Claude Fable 5.168.18%$20.31
  4. Claude Opus 565.15%$6.58
  5. Claude Opus 5.565.15%$22.21
  6. Kimi K3Open63.64%$3.44
  7. Fireworks Ember 163.64%$3.17
  8. Grok 4.662.12%$1.52
  9. GLM-5.3Open62.12%$2.13
  10. Muse Spark 1.3 Max60.61%$0.54
  11. GPT-5.6 Sol60.61%$19.69
  12. MiMo V2.6 ProOpen59.09%$0.18
  13. Qwen3.8 Max59.09%$2.20
  14. Claude Fable 557.58%$9.38
  15. GLM-5.3 FlashOpen57.58%$0.099
  16. GPT-5.6 Terra57.58%$6.87
  17. Grok 4.754.55%$4.18
  18. Claude Sonnet 554.55%$2.51
  19. Tencent HY4 Preview54.55%$0.83
  20. DeepSeek V4.1 FlashOpen53.03%$0.25
  21. GPT-6 Astra53.03%$8.44
  22. GPT-6.1 Sol53.03%$2.76
  23. GPT-5.553.03%$7.95

The margin ends here. The next result, Muse Spark 1.3 and Grok 4.5 at 51.52%, is the first the test can tell apart from the leader. Scores are drawn from 50%; prices on a scale where each step is ten times dearer.

vals.ai Legal Research, regulatory questions6 Oct 2026

Where Sonnet 5.5 lands on the other jobs.

Two more of the guide's jobs have a test that ran it.

Preference rating on real professional tasks (GDPval-AA), at the max setting

  1. Claude Opus 5.51866Tie
  2. Claude Sonnet 5.51839Tie
  3. Claude Fable 5.11758
  4. Grok 4.71715
  5. MiMo V2.6 Pro, open weights1686

Sonnet 5.5 and Opus 5.5 are level: 27 points apart, inside the test's margin of error, on a rating where the next model is 81 points back.

Artificial Analysis: models and the Intelligence IndexNo date stated

Claude Sonnet 5.5 for chemical distributors and producers.

Where its results point. No public test used a chemical company's own documents, so treat each of these as a job to test, not a result.

Research a Regulation

Find what a rule requires before you ship a grade to a new market, with the source text quoted beside the answer.

What to check

Make it quote the dated official text, and check that text yourself. The test is US law, not REACH or GHS.

See this job in the guide

Write a Customer Report

Turn your sales data, account notes, and market news into a quarterly review a customer will read.

What to check

Check every figure against its source. Opus 5.5 rates level on this work and cost less a task in one test.

See this job in the guide

Read a Price List

Pull a supplier's multi-page price list into rows you can quote from.

What to check

It is 12th of 142 on tables. A free 1.2B model leads the independent document test, so try both on your longest lists.

See this job in the guide

Where Sonnet 5.5 costs more than it looks.

Its price per token is low. These are the places that price misleads.

Questions about Claude Sonnet 5.5.

Short answers, from the same sources as the rest of this page.

What is Claude Sonnet 5.5?

A general model from Anthropic, released on 28 September 2026. It reads text and images and costs $2 for every million tokens sent and $10 for every million written.

Is Claude Sonnet 5.5 the best model for regulatory research?

It is first on vals.ai's test of 66 US regulatory questions, at 69.70%. The test cannot separate it from 22 other models, including open models at under 20 cents a question.

Is Claude Sonnet 5.5 cheaper than Claude Opus 5.5?

Per token, yes: half the price. Per task it can cost more, because it uses more tokens. One index task cost $7.67 on Sonnet 5.5 and $5.98 on Opus 5.5. On regulatory questions Sonnet cost $12.64 a question against $22.21.

Can Claude Sonnet 5.5 read chemical price lists and supplier documents?

On ParseBench's table test it scores 91.04, 12th of 142, at 2.74 cents a page. A free 1.2B model, TeleOCR, leads the independent document test, so compare both on your own documents.

Does the regulatory result cover REACH or GHS?

No. The test is US law. For EU or GHS work, make it quote the dated official text, and check that text yourself.

Sources.

  1. vals.ai Legal Research, regulatory questions vals.ai, an independent testing company. Lawyers wrote and checked the questions; an answer counts only if it meets every point on the lawyers' list.Published 6 Oct 2026 · Vailent checked 6 Oct 2026
  2. Artificial Analysis: models and the Intelligence Index Artificial Analysis, an independent testing company. Each model's record gives its index score, its cost per index task, and its professional work rating.States no date · Vailent checked 6 Oct 2026
  3. Artificial Analysis: Claude Sonnet 5.5 reaches #2 on the Intelligence Index Artificial Analysis, an independent testing company, on its own tests of the model.Published 28 Sep 2026 · Vailent checked 30 Sep 2026
  4. ParseBench leaderboard LlamaIndex, which sells LlamaParse, a product it ranks.Published 29 Sep 2026 · Vailent checked 30 Sep 2026
  5. Anthropic: Claude Sonnet The maker's own page.States no date · Vailent checked 30 Sep 2026
  6. Anthropic: pricing The maker's own page.States no date · Vailent checked 30 Sep 2026