The model guide · Updated 6 Oct 2026
Which model to use where.
The best model for a job is often one you have never heard of. We read the public tests, task by task, and show who wins each job at a chemical company and what it costs.
Best in task.
One row per job, from public tests. Open a row to see every result, its price, and what to check before you use it.
Every order, spec check, and compliance file starts with a supplier document. Reading it wrong means a wrong price, a wrong grade, or a missed hazard.
The five best readers are small models, free to download and run on your own computers.
Run it on 20 of your own supplier documents, including a scan and a multi-page SDS, and check every extracted value against the page.
Whole-document reading score: text, tables, and formulas, out of 100.
OmniDocBench ↗Run by OpenDataLab (Shanghai AI Lab), which also makes MinerU, a model it ranks. Board updated 11 Sep 2026 · Vailent checked 30 Sep 2026
A second test puts its own maker's tool first. That tool is marked.
ParseBench ↗Run by LlamaIndex, which sells LlamaParse, a product it ranks. Board updated 29 Sep 2026 · Vailent checked 30 Sep 2026
Price lists and spec tables carry the numbers you quote from. One shifted column is a wrong price on a customer quote.
A free 1.2B model reads tables best on the independent test. On the priced test, a Gemini setting comes within 5 points of the leader at a ninth of the price.
Test your longest price lists. Check merged cells, footnotes, and where rows land across a page break.
Table structure score, out of 100.
OmniDocBench ↗Run by OpenDataLab (Shanghai AI Lab), which also makes MinerU, a model it ranks. Board updated 11 Sep 2026 · Vailent checked 30 Sep 2026
This second test shows prices. It scores the small models above much lower: TeleOCR reaches 72.74 on tables. Its maker sells one of the tools, which is marked.
ParseBench ↗Run by LlamaIndex, which sells LlamaParse, a product it ranks. Board updated 29 Sep 2026 · Vailent checked 30 Sep 2026
Paper still arrives: delivery notes, stamped certificates, photos from the warehouse floor. A reader that fails on a photo sends a person back to typing.
Gemini 3 Pro Preview leads on photos, and free models under 3B come within 8 points. The newest Claude, GPT, and Gemini models have not been tested here.
Photograph ten real documents on the warehouse floor, in bad light, and check every value.
Reading score on photographed pages, out of 100.
MDPBench ↗Run by the OCRBench authors' lab, which also makes MonkeyOCR, a model it ranks. Board updated 15 Aug 2026 · Vailent checked 30 Sep 2026
A shared inbox is where orders wait. Sorting it right the first time is hours back every week, and a request nobody sees is a lost order.
A model built to decide, not to write, sorted 95.9% right with no training. Keyword rules got 77.2%. It is one small study, so treat it as a lead.
Always give it a “none of these” option: without one, it labelled a cake recipe as a technical question. Let it act only when it is sure, and send the rest to a person.
Tickets sorted into the right category.
PriorBench ↗A one-person study on 400 made-up support tickets in French, not run by the model's maker. Board updated 20 Sep 2026 · Vailent checked 30 Sep 2026
A reply with an invented lead time or price is a promise your team has to break. The model that adds the fewest unsupported claims is the safest writer.
Small and open models add the fewest made-up claims, and the leader is a 32B model from Ant Group.
Give it only your approved facts, match every number, date, and promise against the order record, and have a person send.
Summaries with a claim the source does not support. Lower is better.
Vectara hallucination leaderboard ↗Run by Vectara, which grades every model with its own judge and ranks no model of its own. Board updated 22 Sep 2026 · Vailent checked 30 Sep 2026
A missed requirement stops a shipment at the border. And the answer is only as good as the dated official text behind it.
The test cannot separate the leader from open models that cost up to 128 times less.
Each score carries a margin of error of about 6 points on roughly 65 questions of US administrative law. GLM-5.3 Flash, MiMo V2.6 Pro, and Muse Spark 1.3 Max all sit inside the leader's margin, at under $0.60 a question. Kimi K3 is the best result you can run on your own servers, at $3.44 a question. GLM-5.3 is inside its margin, at $2.13.
Make it quote the dated official text, and check that text yourself. The test is US law, not REACH or GHS.
Regulatory research questions answered correctly.
vals.ai Legal Research ↗Run by vals.ai, an independent testing company. Board updated 6 Oct 2026 · Vailent checked 6 Oct 2026
Prices and supply move on news. A search tool that finds the right source is the difference between a good answer and a confident guess.
TinyFish, a search tool few have heard of, gives the most score for the money: 8.8 points below the leader, for about an eighth of the price. Octen comes within 3 points for a third.
Open every source it cites and check the date. Confirm you can get access to the cheaper tool.
Answer quality with each search tool, out of 100.
Artificial Analysis Search Index ↗Run by Artificial Analysis, an independent testing company. Every tool is paired with the same answer model. Board updated 22 Sep 2026 · Vailent checked 30 Sep 2026
The clause you cannot find is the one that costs you in a dispute. Page citations let a person check the answer in seconds.
Inkling Small, free to download, gives the most score for the money: 3.6 points below the leader, for a thirty-fourth of the price. The test has not run the newest models.
Require a page citation for every answer, open it, and treat “not found” as a result to check.
Questions on long credit agreements answered correctly.
vals.ai CorpFin v2 ↗Run by vals.ai, an independent testing company. Board updated 12 Aug 2026 · Vailent checked 30 Sep 2026
Commitments hide in old threads. The fastest way to find them is not always the newest technology.
Plain keyword search found the right email far more often than a neural search model. Use AI to read what it finds.
The study's questions name specific people and things, which suits keyword search. Look for later messages that changed the commitment.
Right email found in the top five results.
EnronQA ↗A published study on 103,638 real business emails. Board updated 1 May 2025 · Vailent checked 30 Sep 2026
Some work has to be right and look right for a customer or a board. Here a large model earns its price.
For finished professional work, large models lead, and Sonnet 5.5 matches Opus 5.5 at half the price.
Opus 5.5 and Sonnet 5.5 are level within the test's margin of error, and Sonnet costs half as much per token.
Check every figure against its source before it leaves the building.
Preference rating on real professional tasks.
Artificial Analysis GDPval-AA ↗Run by Artificial Analysis, an independent testing company. The board states no update date · Vailent checked 5 Oct 2026
The same grade has five names across suppliers and customers. A wrong match ships the wrong material.
An open model from Moonshot matched products best, ahead of GPT-5.2.
Match exact identifiers first. Send every unsure match, and any grade equivalence that changes a price, to an expert.
Product matches made correctly (F1), out of 100.
Steiner and Bizer, WDC Products ↗A published study by the University of Mannheim. Board updated 27 Jun 2026 · Vailent checked 23 Sep 2026
Slow, lumpy demand is where stock-outs and dead stock both come from. The right method trades service against the cash tied up in stock.
Classic forecasting methods filled far more complete orders than AI forecasting models, at the cost of more stock on the shelf.
The classic methods fill more orders because they forecast high, so they hold more stock: Croston carried about four times the inventory of TimesFM. When each method's safety stock is tuned, the gap narrows.
Backtest on your own order history, and compare fill rate and stock held together, never one without the other.
Complete orders filled from stock.
SMU intermittent demand study ↗A published study by Singapore Management University and ST Logistics, on 20,330 real orders for industrial spare parts. Board updated 12 Sep 2026 · Vailent checked 30 Sep 2026
For products that sell steadily, AI forecasting models do better. Several are small and free. The best one, TimesFM-3, is not licensed for business use.
fev-bench ↗Run by the AutoGluon team at Amazon, which also makes Chronos, a model it ranks. Board updated 18 Sep 2026 · Vailent checked 30 Sep 2026
One missed match is a regulatory breach, and every false alarm is a person's time. The score below weighs both.
An open 14B model comes within a point of the best. The production matcher misses almost nothing, but one pair in seven it flags is a false alarm.
A person makes every decision. Check every hit, every name in another script, and a sample of non-hits.
Matches judged correctly, balancing misses and false alarms (F1), out of 100.
OpenSanctions Pairs ↗A published study on OpenSanctions data. Board updated 25 Aug 2026 · Vailent checked 30 Sep 2026
Model highlights and updates.
New models and new test results, dated, and the tasks they change.
Last Week in AI for Chemicals
Kimi K3, an open model you can run on your own servers, now leads the open models on regulatory research.
Read the editionGPT-6.1 Sol Arrives
It ranks 11th of 223 in its price class on the general intelligence index, at a tenth of the leaders' cost a task. On regulatory research it scores 53% at $2.76 a question, just inside the leading group.
Read the highlightSonnet 5.5 Leads Regulatory Research and Joins the Document Test
Top of 74 models on regulatory questions, though the test cannot separate it from open models at a hundredth of the price. On tables it lands 12th, at 2.74 cents a page.
Read the highlightClaude Opus 5.5 Released
First on the priced table test at its high setting, and first on finished professional work, level with Sonnet 5.5, whose tokens cost half as much. It is the dearest model in the regulatory leading group.
Read the highlightJev 1.13 Released
A model that returns a decision and how sure it is, not text. In one small study it sorted 95.9% of tickets right with no training.
Read the highlightA Free 1.2B Model Tops the Independent Document Test
TeleOCR leads OmniDocBench on whole documents and on tables, ahead of Gemini 3 Pro.
Read the highlight
Why one model is not enough.
One model was never the answer.
For years, using AI meant picking one model and asking it everything. The table above shows why that never held. A small specialist reads your documents best, classic forecasting fills more orders for your slow-moving stock, and the most expensive model is rarely the right buy.
Information is not context.
A model knows the public internet. It does not know your price lists, your allocation rules, your customers' history, or the exception your senior trader remembers. Give a smaller model that context, and it does work a larger one without it gets wrong.
The right size for every step.
Sorting a request, reading a certificate, and drafting a reply to a key account are different work. Small models should carry the volume, the largest should be saved for the work that needs them, and every answer should come back with its sources, ready for your review.
No two chemical companies run the same way, so the AI that serves them should not either.
See how Vailent works