Weekly, every Monday

Last Week in AI for Chemicals:
28 Sep to 4 Oct 2026

Kimi K3, an open model you can run on your own servers, now leads the open models on regulatory research.

63.64%Kimi K3 on regulatory research, the best open model on the test
Vailent model guide · Updated 3 min read

Every Monday, Vailent reads the public tests of AI models and reports what changed for chemical companies. The model guide recommends models for 13 everyday jobs, from reading supplier documents to screening sanctions lists, and links each recommendation to the test behind it.

Recommendations that changed.

For each job, the guide recommends three models: the best result, the most efficient option, and the best model you can run on your own servers. Last week, one of those recommendations changed.

Research RegulationsFree to run in-house

The best model you can download and run on your own servers.

BeforeGLM-5.362.12%, $2.13 a question
NowKimi K363.64%, $3.44 a question

The test cannot tell Kimi K3 apart from GLM-5.3, and GLM-5.3 costs less a question, so both are reasonable choices to run yourself.

vals.ai Legal Research, regulatory questions6 Oct 2026

Models released last week.

What the tests showed.

Public tests score AI models on set tasks, such as answering regulatory questions. These are last week's results that matter for chemical companies.

  1. Two More Readers for Photographed Documents

    On MDPBench, U2-OCR enters 2nd at 84.1, behind Gemini 3 Pro Preview at 85.1. The board does not say who makes it. PepperOCR-VL, an open 4.5B model from Sionic AI, scores 80.7, above dots.mocr, the guide's in-house pick for scans, at 77.2. Its AGPL license limits how a business can use it.

    Read Scans and Photos in the guide
  2. Ling 3.1 Flash Rates 31 Points Below GLM-5.3 at a Fifth of the Price

    Ant Group's Ling 3.1 Flash scores 1622 on rated professional work, against 1653 for GLM-5.3, at $0.30 for every million tokens sent and $0.90 for every million written, against $1.40 and $4.40. Artificial Analysis lists it as open weights, but the weights are not on Hugging Face.

    Write a Report From Several Sources in the guide
  3. Gemini 4 Argon Arrives Near the Top, at a Launch Price

    Google's latest model scores 68.18% on regulatory research, level with Claude Fable 5.1 in 2nd of 74, and 76.23% on vals.ai's tax test, 2nd of 63. Its $1.99 a task on the intelligence index is a 50% launch discount that rises to $3.98. Google is rolling it out to selected users, so it is not generally available yet.

    Research Regulations in the guide
  4. vals.ai Adds a Tax Research Test

    193 research-grade US corporate tax questions, from an independent tester. Claude Fable 5.1 leads at 77.64%, and GLM-5.3, an open model, scores 73.09% at $1.80 a question. It tests tax, not chemical rules, so read it as a sign of how models handle regulated research in general.

  5. A Public Test of Dangerous Goods Rules

    A study revised on 28 September asked models 1,678 questions on the IMDG Code, the rules for shipping dangerous goods by sea. The best model, Gemini 3.1 Pro, was right on 95.8% of multiple choice questions and 89.6% of Dangerous Goods List lookups, but only 36.8% when asked to recall the code's text. The study puts trained people at 83.8% on the multiple choice part. Every model was weak on stowage and segregation, and none of the newest models was tested.

Released, not yet tested.

No independent test has run these, so the guide shows no score.

  • Xiaomi-OCR-0

    No independent test

    A free 0.87B document reader, published by SeerRay-Lab on Hugging Face under the Apache 2.0 license. Its model card claims a result near the top of the independent document test, but the test's board does not list it, so the guide shows no score.

    Published

    SeerRay-Lab: Xiaomi-OCR-029 Sep 2026

Still no public test.

Jobs where nobody published a test last week. Knowing what has not been measured tells you what to check yourself.

Sources.

  1. vals.ai Tax Agent Bench vals.ai, an independent testing company. 193 private questions.Published 30 Sep 2026 · Vailent checked 5 Oct 2026
  2. Artificial Analysis: Gemini 4 Argon Artificial Analysis, an independent testing company, on its own tests of the model.Published 30 Sep 2026 · Vailent checked 5 Oct 2026
  3. Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance Researchers at NCB Hazcheck, which sells dangerous goods software the study does not rank, and Durham University.Published 28 Sep 2026 · Vailent checked 5 Oct 2026
  4. MDPBench leaderboard The OCRBench authors' lab, which also makes MonkeyOCR, a model it ranks.Published 4 Oct 2026 · Vailent checked 5 Oct 2026
  5. Sionic AI: PepperOCR-VL The maker's own page.Published 4 Oct 2026 · Vailent checked 5 Oct 2026
  6. Artificial Analysis: Ling 3.1 Flash Artificial Analysis, an independent testing company. The model's record gives its professional work rating and price.States no date · Vailent checked 5 Oct 2026
  7. Artificial Analysis: models and the Intelligence Index Artificial Analysis, an independent testing company. Each model's record gives its index score, its cost per index task, and its professional work rating.States no date · Vailent checked 6 Oct 2026
  8. Hugging Face: inclusionAI Ant Group's own model page on Hugging Face.States no date · Vailent checked 5 Oct 2026
  9. SeerRay-Lab: Xiaomi-OCR-0 The maker's own page.Published 29 Sep 2026 · Vailent checked 5 Oct 2026