The model guide · Updated 6 Oct 2026

Which model to use where.

The best model for a job is often one you have never heard of. We read the public tests, task by task, and show who wins each job at a chemical company and what it costs.

Most recent highlightLast Week in AI for ChemicalsKimi K3, an open model you can run on your own servers, now leads the open models on regulatory research.63.64%Kimi K3 on regulatory research, the best open model on the test28 Sep to 4 Oct 2026Read the edition

Best in task.

One row per job, from public tests. Open a row to see every result, its price, and what to check before you use it.

Best resultMost Efficient OptionFree to run in-house
Read Supplier DocumentsPull the fields you need from a TDS, SDS, or CoA.TeleOCR Specialist · 1.2B · Free to download96.91No prices publishedSame as the best resultFree to download and run in-house

Every order, spec check, and compliance file starts with a supplier document. Reading it wrong means a wrong price, a wrong grade, or a missed hazard.

The five best readers are small models, free to download and run on your own computers.

What to check

Run it on 20 of your own supplier documents, including a scan and a multi-page SDS, and check every extracted value against the page.

Read the highlightTeleOCR: A Free Model That Tops One Document Test

Whole-document reading score: text, tables, and formulas, out of 100.

  1. TeleOCRSpecialist · 1.2B · Free to download96.91
  2. OvisOCR2Specialist · 0.8B · Free to download96.47
  3. PaddleOCR-VL-1.6Specialist · 0.9B · Free to download96.34
  4. MinerU2.5-ProSpecialist · 1.2B · Free to download · The test's own product95.75
  5. GLM-OCRSpecialist · 0.9B · Free to download95.22
  6. Gemini 3 ProGoogle92.91
  7. Gemini 3 FlashGoogle92.62
  8. GPT-5.2OpenAI86.59

OmniDocBench ↗Run by OpenDataLab (Shanghai AI Lab), which also makes MinerU, a model it ranks. Board updated 11 Sep 2026 · Vailent checked 30 Sep 2026

A second test puts its own maker's tool first. That tool is marked.

  1. LlamaParse Agentic PlusTool · The test's own product90.25.63 US cents a page
  2. LlamaParse AgenticTool · The test's own product87.011.25 US cents a page
  3. Pulse Ultra 2Tool81.61.50 US cents a page
  4. Claude Opus 5.5Anthropic79.856.13 US cents a page
  5. rakedoc-nanoSpecialist · 1.16B · Free to download77.23No price published
  6. Gemini 3 FlashGoogle75.052.41 US cents a page
  7. Claude Sonnet 5.5Anthropic70.182.74 US cents a page
  8. GPT-6 SolOpenAI69.693.47 US cents a page

ParseBench ↗Run by LlamaIndex, which sells LlamaParse, a product it ranks. Board updated 29 Sep 2026 · Vailent checked 30 Sep 2026

Pull Tables and Price ListsRead a supplier's multi-page price list into rows you can use.TeleOCR Specialist · 1.2B · Free to download96.82No prices publishedSame as the best resultFree to download and run in-house

Price lists and spec tables carry the numbers you quote from. One shifted column is a wrong price on a customer quote.

A free 1.2B model reads tables best on the independent test. On the priced test, a Gemini setting comes within 5 points of the leader at a ninth of the price.

What to check

Test your longest price lists. Check merged cells, footnotes, and where rows land across a page break.

Read the highlightClaude Sonnet 5.5: First on Regulations, in a CrowdRead the highlightClaude Opus 5.5: First Place, and What It CostsRead the highlightTeleOCR: A Free Model That Tops One Document Test

Table structure score, out of 100.

  1. TeleOCRSpecialist · 1.2B · Free to download96.82
  2. PaddleOCR-VL-1.6Specialist · 0.9B · Free to download94.76
  3. OvisOCR2Specialist · 0.8B · Free to download94.58
  4. MinerU2.5-ProSpecialist · 1.2B · Free to download · The test's own product93.42
  5. GLM-OCRSpecialist · 0.9B · Free to download92.83
  6. Gemini 3 FlashGoogle89.29
  7. Gemini 3 ProGoogle89.15

OmniDocBench ↗Run by OpenDataLab (Shanghai AI Lab), which also makes MinerU, a model it ranks. Board updated 11 Sep 2026 · Vailent checked 30 Sep 2026

This second test shows prices. It scores the small models above much lower: TeleOCR reaches 72.74 on tables. Its maker sells one of the tools, which is marked.

  1. Claude Opus 5.5Anthropic94.256.13 US cents a page
  2. LlamaParse Agentic PlusTool · The test's own product93.375.63 US cents a page
  3. GPT-6 AstraOpenAI93.179.87 US cents a page
  4. GPT-6 SolOpenAI92.633.47 US cents a page
  5. oi-parserTool92.62No price published
  6. Claude Fable 5.1Anthropic91.5216.05 US cents a page
  7. Gemini 3 FlashGoogle91.52.41 US cents a page
  8. Claude Sonnet 5.5Anthropic91.042.74 US cents a page
  9. Gemini 3 Flash (minimal thinking)Google89.850.65 US cents a page

ParseBench ↗Run by LlamaIndex, which sells LlamaParse, a product it ranks. Board updated 29 Sep 2026 · Vailent checked 30 Sep 2026

Read Scans and PhotosRead a photographed delivery note or a scanned certificate.Gemini 3 Pro Preview Google85.1No prices publisheddots.mocrSpecialist · 3B · Free to download77.2

Paper still arrives: delivery notes, stamped certificates, photos from the warehouse floor. A reader that fails on a photo sends a person back to typing.

Gemini 3 Pro Preview leads on photos, and free models under 3B come within 8 points. The newest Claude, GPT, and Gemini models have not been tested here.

What to check

Photograph ten real documents on the warehouse floor, in bad light, and check every value.

Reading score on photographed pages, out of 100.

  1. Gemini 3 Pro PreviewGoogle85.1
  2. MonkeyOCRv2-BSpecialist · 0.9B · Free to download · The test's own product81.7
  3. Kimi K3Moonshot AI · Free to download81.2
  4. dots.mocrSpecialist · 3B · Free to download77.2
  5. Chandra OCR 2Specialist · 5.3B77.1
  6. PaddleOCR-VL-1.6Specialist · 0.9B · Free to download76.3
  7. Claude Sonnet 4.6Anthropic69.3
  8. GPT-5.2OpenAI63

MDPBench ↗Run by the OCRBench authors' lab, which also makes MonkeyOCR, a model it ranks. Board updated 15 Aug 2026 · Vailent checked 30 Sep 2026

Sort Requests and RFQsSend each emailed RFQ, order, or complaint to the right team.Jev 1.13 Specialist95.9%No prices publishedNone tested

A shared inbox is where orders wait. Sorting it right the first time is hours back every week, and a request nobody sees is a lost order.

A model built to decide, not to write, sorted 95.9% right with no training. Keyword rules got 77.2%. It is one small study, so treat it as a lead.

What to check

Always give it a “none of these” option: without one, it labelled a cake recipe as a technical question. Let it act only when it is sure, and send the rest to a person.

Read the highlightJev 1.13: A Model That Decides Instead of Writing

Tickets sorted into the right category.

  1. Jev 1.13Specialist95.9%
  2. Hand-written keyword rulesNo AI needed77.2%
  3. Word-count classifierNo AI needed66%

PriorBench ↗A one-person study on 400 made-up support tickets in French, not run by the model's maker. Board updated 20 Sep 2026 · Vailent checked 30 Sep 2026

Draft Replies Without Made-Up FactsDraft a customer reply from the order, the price, and the lead time.Finix S1 32B Ant Group · 32B1.8%No prices publishedLlama 3.3 70BMeta · 70B · Free to download4.1%

A reply with an invented lead time or price is a promise your team has to break. The model that adds the fewest unsupported claims is the safest writer.

Small and open models add the fewest made-up claims, and the leader is a 32B model from Ant Group.

What to check

Give it only your approved facts, match every number, date, and promise against the order record, and have a person send.

Summaries with a claim the source does not support. Lower is better.

  1. Finix S1 32BAnt Group · 32B1.8%
  2. GPT-5.4 nanoOpenAI3.1%
  3. Gemini 2.5 Flash LiteGoogle3.3%
  4. Llama 3.3 70BMeta · 70B · Free to download4.1%
  5. Gemma 3 12BGoogle · 12B · Free to download4.4%
  6. Qwen3 8BAlibaba · 8B · Free to download4.8%
  7. Granite 4.0 H SmallIBM · Free to download5.2%
  8. GPT-6 SolOpenAI6.5%
  9. Claude Haiku 4.5Anthropic9.8%
  10. Gemini 3 FlashGoogle13.5%

Vectara hallucination leaderboard ↗Run by Vectara, which grades every model with its own judge and ranks no model of its own. Board updated 22 Sep 2026 · Vailent checked 30 Sep 2026

Research RegulationsFind what a rule requires before you ship a product to a new market.Claude Sonnet 5.5 Anthropic69.7%GLM-5.3 FlashZ.ai · Free to download57.58%128x cheaperKimi K3Moonshot AI · Free to download63.64%

A missed requirement stops a shipment at the border. And the answer is only as good as the dated official text behind it.

The test cannot separate the leader from open models that cost up to 128 times less.

Each score carries a margin of error of about 6 points on roughly 65 questions of US administrative law. GLM-5.3 Flash, MiMo V2.6 Pro, and Muse Spark 1.3 Max all sit inside the leader's margin, at under $0.60 a question. Kimi K3 is the best result you can run on your own servers, at $3.44 a question. GLM-5.3 is inside its margin, at $2.13.

What to check

Make it quote the dated official text, and check that text yourself. The test is US law, not REACH or GHS.

Read the highlightGPT-6.1 Sol: Near the Top for a Tenth of the BillRead the highlightClaude Sonnet 5.5: First on Regulations, in a CrowdRead the highlightClaude Opus 5.5: First Place, and What It Costs

Regulatory research questions answered correctly.

  1. Claude Sonnet 5.5Anthropic69.7%$12.64 a question
  2. Gemini 4 ArgonGoogle68.18%$4.81 a question
  3. Claude Fable 5.1Anthropic68.18%$20.31 a question
  4. Claude Opus 5Anthropic65.15%$6.58 a question
  5. Claude Opus 5.5Anthropic65.15%$22.21 a question
  6. Kimi K3Moonshot AI · Free to download63.64%$3.44 a question
  7. Grok 4.6xAI62.12%$1.52 a question
  8. GLM-5.3Z.ai · Free to download62.12%$2.13 a question
  9. Muse Spark 1.3 MaxMeta60.61%$0.54 a question
  10. GPT-5.6 SolOpenAI60.61%$19.69 a question
  11. MiMo V2.6 ProXiaomi · Free to download59.09%$0.18 a question
  12. GLM-5.3 FlashZ.ai · Free to download57.58%$0.099 a question
  13. GPT-6.1 SolOpenAI53.03%$2.76 a question
  14. DeepSeek V4.1 FlashDeepSeek · Free to download53.03%$0.25 a question

vals.ai Legal Research ↗Run by vals.ai, an independent testing company. Board updated 6 Oct 2026 · Vailent checked 6 Oct 2026

Find a Clause in a ContractFind the price-adjustment or force majeure clause in a supply agreement.Claude Opus 5 Anthropic73.19%Inkling SmallThinking Machines · Free to download69.62%34x cheaperKimi K3Moonshot AI · Free to download71.56%

The clause you cannot find is the one that costs you in a dispute. Page citations let a person check the answer in seconds.

Inkling Small, free to download, gives the most score for the money: 3.6 points below the leader, for a thirty-fourth of the price. The test has not run the newest models.

What to check

Require a page citation for every answer, open it, and treat “not found” as a result to check.

Questions on long credit agreements answered correctly.

  1. Claude Opus 5Anthropic73.19%$0.86 a question
  2. Claude Fable 5Anthropic71.83%$1.74 a question
  3. Kimi K3Moonshot AI · Free to download71.56%$0.064 a question
  4. Muse Spark 1.1Meta71.29%$0.11 a question
  5. Inkling SmallThinking Machines · Free to download69.62%$0.025 a question
  6. Grok 4.3xAI68.53%$0.13 a question
  7. GPT-5.5OpenAI68.42%$0.43 a question

vals.ai CorpFin v2 ↗Run by vals.ai, an independent testing company. Board updated 12 Aug 2026 · Vailent checked 30 Sep 2026

Search Years of EmailFind what was promised to a customer three years ago.Keyword search (BM25) No AI needed87.5%No prices publishedNone tested

Commitments hide in old threads. The fastest way to find them is not always the newest technology.

Plain keyword search found the right email far more often than a neural search model. Use AI to read what it finds.

What to check

The study's questions name specific people and things, which suits keyword search. Look for later messages that changed the commitment.

Right email found in the top five results.

  1. Keyword search (BM25)No AI needed87.5%
  2. ColBERTv2Specialist · Free to download54.1%

EnronQA ↗A published study on 103,638 real business emails. Board updated 1 May 2025 · Vailent checked 30 Sep 2026

Write a Report From Several SourcesTurn sales data, notes, and market news into a quarterly account review.Claude Opus 5.5 Anthropic1866No prices publishedMiMo V2.6 ProXiaomi · Free to download1686

Some work has to be right and look right for a customer or a board. Here a large model earns its price.

For finished professional work, large models lead, and Sonnet 5.5 matches Opus 5.5 at half the price.

Opus 5.5 and Sonnet 5.5 are level within the test's margin of error, and Sonnet costs half as much per token.

What to check

Check every figure against its source before it leaves the building.

Read the highlightGPT-6.1 Sol: Near the Top for a Tenth of the BillRead the highlightClaude Sonnet 5.5: First on Regulations, in a CrowdRead the highlightClaude Opus 5.5: First Place, and What It Costs

Preference rating on real professional tasks.

  1. Claude Opus 5.5Anthropic1866
  2. Claude Sonnet 5.5Anthropic1839
  3. Claude Fable 5.1Anthropic1758
  4. Grok 4.7xAI1715
  5. MiMo V2.6 ProXiaomi · Free to download1686
  6. GLM-5.3Z.ai · Free to download1653
  7. GPT-6.1 SolOpenAI1575

Artificial Analysis GDPval-AA ↗Run by Artificial Analysis, an independent testing company. The board states no update date · Vailent checked 5 Oct 2026

Match Products and GradesMatch a customer's product name to the grade in your catalogue.Kimi K2.6 Moonshot AI · Free to download87.2No prices publishedSame as the best resultFree to download and run in-house

The same grade has five names across suppliers and customers. A wrong match ships the wrong material.

An open model from Moonshot matched products best, ahead of GPT-5.2.

What to check

Match exact identifiers first. Send every unsure match, and any grade equivalence that changes a price, to an expert.

Product matches made correctly (F1), out of 100.

  1. Kimi K2.6Moonshot AI · Free to download87.2
  2. GPT-5.2OpenAI84.47
  3. DittoSpecialist · Free to download71.94

Steiner and Bizer, WDC Products ↗A published study by the University of Mannheim. Board updated 27 Jun 2026 · Vailent checked 23 Sep 2026

Forecast Demand for Slow-Moving StockPlan stock for a grade that sells a few drums a quarter.ADIDA No AI needed90.6%No prices publishedSame as the best resultFree to download and run in-house

Slow, lumpy demand is where stock-outs and dead stock both come from. The right method trades service against the cash tied up in stock.

Classic forecasting methods filled far more complete orders than AI forecasting models, at the cost of more stock on the shelf.

The classic methods fill more orders because they forecast high, so they hold more stock: Croston carried about four times the inventory of TimesFM. When each method's safety stock is tuned, the gap narrows.

What to check

Backtest on your own order history, and compare fill rate and stock held together, never one without the other.

Complete orders filled from stock.

  1. ADIDANo AI needed90.6%
  2. Croston's methodNo AI needed85.8%
  3. SBANo AI needed84.7%
  4. LightGBMNo AI needed77.1%
  5. MoiraiSpecialist · Free to download71.9%
  6. Chronos-Bolt smallSpecialist · Free to download56.7%
  7. TimesFMSpecialist · Free to download54.4%

SMU intermittent demand study ↗A published study by Singapore Management University and ST Logistics, on 20,330 real orders for industrial spare parts. Board updated 12 Sep 2026 · Vailent checked 30 Sep 2026

For products that sell steadily, AI forecasting models do better. Several are small and free. The best one, TimesFM-3, is not licensed for business use.

  1. TimesFM-3Specialist · 331M86.8%
  2. Chronos-2Specialist · 119M · Free to download · The test's own product81.3%
  3. t0-betaSpecialist · 256M · Free to download78.7%
  4. TiRex-2Specialist · 38M · Free to download77.8%
  5. Toto-2.0 2.5BSpecialist · 2.5B · Free to download77.5%
  6. AutoARIMANo AI needed29.3%

fev-bench ↗Run by the AutoGluon team at Amazon, which also makes Chronos, a model it ranks. Board updated 18 Sep 2026 · Vailent checked 30 Sep 2026

Screen Sanctions ListsCheck a new customer or its owners against sanctions lists.GPT-4o OpenAI98.95No prices publishedDeepSeek R1 Distill Qwen 14BDeepSeek · 14B · Free to download98.23

One missed match is a regulatory breach, and every false alarm is a person's time. The score below weighs both.

An open 14B model comes within a point of the best. The production matcher misses almost nothing, but one pair in seven it flags is a false alarm.

What to check

A person makes every decision. Check every hit, every name in another script, and a sample of non-hits.

Matches judged correctly, balancing misses and false alarms (F1), out of 100.

  1. GPT-4oOpenAI98.95
  2. GPT-5.2 ProOpenAI98.75
  3. DeepSeek R1 Distill Qwen 14BDeepSeek · 14B · Free to download98.23
  4. Claude 3.7 SonnetAnthropic97.5
  5. Claude Opus 4.5Anthropic95.45
  6. Llama 3.1 8BMeta · 8B · Free to download94.05
  7. Feature-based matcherNo AI needed91.33

OpenSanctions Pairs ↗A published study on OpenSanctions data. Board updated 25 Aug 2026 · Vailent checked 30 Sep 2026

Model highlights and updates.

New models and new test results, dated, and the tasks they change.

  1. Last Week in AI for Chemicals

    Kimi K3, an open model you can run on your own servers, now leads the open models on regulatory research.

    Read the edition
  2. GPT-6.1 Sol Arrives

    It ranks 11th of 223 in its price class on the general intelligence index, at a tenth of the leaders' cost a task. On regulatory research it scores 53% at $2.76 a question, just inside the leading group.

    Read the highlight
  3. Sonnet 5.5 Leads Regulatory Research and Joins the Document Test

    Top of 74 models on regulatory questions, though the test cannot separate it from open models at a hundredth of the price. On tables it lands 12th, at 2.74 cents a page.

    Read the highlight
  4. Claude Opus 5.5 Released

    First on the priced table test at its high setting, and first on finished professional work, level with Sonnet 5.5, whose tokens cost half as much. It is the dearest model in the regulatory leading group.

    Read the highlight
  5. Jev 1.13 Released

    A model that returns a decision and how sure it is, not text. In one small study it sorted 95.9% of tickets right with no training.

    Read the highlight
  6. A Free 1.2B Model Tops the Independent Document Test

    TeleOCR leads OmniDocBench on whole documents and on tables, ahead of Gemini 3 Pro.

    Read the highlight

Why one model is not enough.

One model was never the answer.

For years, using AI meant picking one model and asking it everything. The table above shows why that never held. A small specialist reads your documents best, classic forecasting fills more orders for your slow-moving stock, and the most expensive model is rarely the right buy.

Information is not context.

A model knows the public internet. It does not know your price lists, your allocation rules, your customers' history, or the exception your senior trader remembers. Give a smaller model that context, and it does work a larger one without it gets wrong.

The right size for every step.

Sorting a request, reading a certificate, and drafting a reply to a key account are different work. Small models should carry the volume, the largest should be saved for the work that needs them, and every answer should come back with its sources, ready for your review.

No two chemical companies run the same way, so the AI that serves them should not either.

See how Vailent works