Last Week in AI for Chemicals:
28 Sep to 4 Oct 2026
Kimi K3, an open model you can run on your own servers, now leads the open models on regulatory research.
Every Monday, Vailent reads the public tests of AI models and reports what changed for chemical companies. The model guide recommends models for 13 everyday jobs, from reading supplier documents to screening sanctions lists, and links each recommendation to the test behind it.
Recommendations that changed.
For each job, the guide recommends three models: the best result, the most efficient option, and the best model you can run on your own servers. Last week, one of those recommendations changed.
Research RegulationsFree to run in-house
The best model you can download and run on your own servers.
The test cannot tell Kimi K3 apart from GLM-5.3, and GLM-5.3 costs less a question, so both are reasonable choices to run yourself.
Models released last week.
What the tests showed.
Public tests score AI models on set tasks, such as answering regulatory questions. These are last week's results that matter for chemical companies.
Two More Readers for Photographed Documents
On MDPBench, U2-OCR enters 2nd at 84.1, behind Gemini 3 Pro Preview at 85.1. The board does not say who makes it. PepperOCR-VL, an open 4.5B model from Sionic AI, scores 80.7, above dots.mocr, the guide's in-house pick for scans, at 77.2. Its AGPL license limits how a business can use it.
Read Scans and Photos in the guideMDPBench leaderboard4 Oct 2026
Sionic AI: PepperOCR-VL4 Oct 2026
Ling 3.1 Flash Rates 31 Points Below GLM-5.3 at a Fifth of the Price
Ant Group's Ling 3.1 Flash scores 1622 on rated professional work, against 1653 for GLM-5.3, at $0.30 for every million tokens sent and $0.90 for every million written, against $1.40 and $4.40. Artificial Analysis lists it as open weights, but the weights are not on Hugging Face.
Write a Report From Several Sources in the guideArtificial Analysis: Ling 3.1 FlashNo date stated
Artificial Analysis: models and the Intelligence IndexNo date stated
Hugging Face: inclusionAINo date stated
Gemini 4 Argon Arrives Near the Top, at a Launch Price
Google's latest model scores 68.18% on regulatory research, level with Claude Fable 5.1 in 2nd of 74, and 76.23% on vals.ai's tax test, 2nd of 63. Its $1.99 a task on the intelligence index is a 50% launch discount that rises to $3.98. Google is rolling it out to selected users, so it is not generally available yet.
Research Regulations in the guidevals.ai Legal Research, regulatory questions6 Oct 2026
vals.ai Tax Agent Bench30 Sep 2026
Artificial Analysis: Gemini 4 Argon30 Sep 2026
vals.ai Adds a Tax Research Test
193 research-grade US corporate tax questions, from an independent tester. Claude Fable 5.1 leads at 77.64%, and GLM-5.3, an open model, scores 73.09% at $1.80 a question. It tests tax, not chemical rules, so read it as a sign of how models handle regulated research in general.
vals.ai Tax Agent Bench30 Sep 2026
A Public Test of Dangerous Goods Rules
A study revised on 28 September asked models 1,678 questions on the IMDG Code, the rules for shipping dangerous goods by sea. The best model, Gemini 3.1 Pro, was right on 95.8% of multiple choice questions and 89.6% of Dangerous Goods List lookups, but only 36.8% when asked to recall the code's text. The study puts trained people at 83.8% on the multiple choice part. Every model was weak on stowage and segregation, and none of the newest models was tested.
Released, not yet tested.
No independent test has run these, so the guide shows no score.
Xiaomi-OCR-0
No independent testA free 0.87B document reader, published by SeerRay-Lab on Hugging Face under the Apache 2.0 license. Its model card claims a result near the top of the independent document test, but the test's board does not list it, so the guide shows no score.
Published
SeerRay-Lab: Xiaomi-OCR-029 Sep 2026
Still no public test.
Jobs where nobody published a test last week. Knowing what has not been measured tells you what to check yourself.
- Read Supplier Documents
No public test last week measured reading a CoA, a TDS, or an SDS with current models.
- Research Regulations
No public test last week asked about GHS, CLP, REACH, or TSCA. The guide's regulatory test is US administrative law.
- Sort Requests and RFQs
No public test last week measured sorting RFQs or order emails.
- Screen Sanctions Lists
No public test last week measured sanctions screening with current models.
Sources.
- vals.ai Legal Research, regulatory questions vals.ai, an independent testing company.Published 6 Oct 2026 · Vailent checked 6 Oct 2026
- vals.ai Tax Agent Bench vals.ai, an independent testing company. 193 private questions.Published 30 Sep 2026 · Vailent checked 5 Oct 2026
- Artificial Analysis: Gemini 4 Argon Artificial Analysis, an independent testing company, on its own tests of the model.Published 30 Sep 2026 · Vailent checked 5 Oct 2026
- Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance Researchers at NCB Hazcheck, which sells dangerous goods software the study does not rank, and Durham University.Published 28 Sep 2026 · Vailent checked 5 Oct 2026
- MDPBench leaderboard The OCRBench authors' lab, which also makes MonkeyOCR, a model it ranks.Published 4 Oct 2026 · Vailent checked 5 Oct 2026
- Sionic AI: PepperOCR-VL The maker's own page.Published 4 Oct 2026 · Vailent checked 5 Oct 2026
- Artificial Analysis: Ling 3.1 Flash Artificial Analysis, an independent testing company. The model's record gives its professional work rating and price.States no date · Vailent checked 5 Oct 2026
- Artificial Analysis: models and the Intelligence Index Artificial Analysis, an independent testing company. Each model's record gives its index score, its cost per index task, and its professional work rating.States no date · Vailent checked 6 Oct 2026
- Hugging Face: inclusionAI Ant Group's own model page on Hugging Face.States no date · Vailent checked 5 Oct 2026
- SeerRay-Lab: Xiaomi-OCR-0 The maker's own page.Published 29 Sep 2026 · Vailent checked 5 Oct 2026