Jev 1.13 for Chemical Companies: A Model That Decides Instead of Writing
Most AI models write. Jev picks one of the options you give it and says how sure it is. For a chemical distributor or producer sorting RFQs, orders, and supplier documents, that is faster and far cheaper, with limits you need to know first.
Jev 1.13 at a glance
What Jev 1.13 is.
TypeSafe AI released Jev 1.13 on 15 September 2026. You send it a message and a short list of options, such as the teams that handle your requests. It returns one option and a probability, and nothing else.
That makes it a different tool from a chat model. It cannot draft a reply or summarize a thread. It can decide, quickly and cheaply, which of your options a message belongs to, and it can answer several such questions about one message in a single call.
It reads text only, so a scanned certificate or a photographed delivery note needs text recognition first. It handles up to 255 options in one question. The maker says English is its strongest language; the outside study ran in French and found French and English equally good on its hardest items.
How accurate Jev is.
Two outside studies have tested it so far. Both are small, and neither used a chemical company's own messages.
Tickets put in the right one of four categories, out of 400
Four in ten tickets were written to trip keyword rules, with a word from the wrong category in them. Jev had no examples; the classifier saw all 400 answers and still scored lower.
PriorBench: Jev, measured20 Sep 2026
When to trust its answer.
Jev says how sure it is on every answer. Set a bar: above it, Jev acts alone; below it, a person decides. Pick a bar to see what the study measured.
The bar most often recommended. The study found it no safer than 50%.
PriorBench: Jev, measured20 Sep 2026
Gate at 0.99, or do not gate.
PriorBench, recommendation 3
Ask it narrow questions.
How you ask matters more than which model you ask.
In the phishing study the same model rose 32 points when one broad question became five narrow ones. At a chemical company, that means asking “is this an RFQ?”, “does it name a grade?”, and “is it urgent?” as separate questions, in one call, rather than “where does this go?”
Jev phishing bench17 Sep 2026
Jev for chemical distributors and producers.
Anywhere a person reads a message only to decide where it goes. No public test has used a distributor's or a producer's inbox, so treat each of these as a job to test, not a result.
Sort Requests and RFQs
Send each emailed RFQ, order change, or complaint to the team that handles it, with a “none of these” option for the rest.
Count how many it gets right on 200 of your own requests before it routes anything alone.
Label Incoming Documents
Decide whether a supplier file is a certificate of analysis, a safety data sheet, an invoice, or a price change, so it lands in the right place.
It reads text only, so run text recognition on scans first, and check every label it is unsure of.
Clear Routine Exceptions
Clear the order holds and small spec deviations your team always decides the same way, and pass the rest to a person.
Let it act only when it is 99% sure. A named person owns every decision that changes a price, a credit limit, or a shipment.
Jev compared with a chat model and keyword rules.
The three things most teams would try for sorting requests, on the evidence there is.
| Jev 1.13 | Claude Haiku 4.5 | Keyword rules | |
|---|---|---|---|
| What it returns | One of your options, and how sure it is | Text, which you then read or parse | A match, or nothing |
| Needs examples first | No | No | Someone writes every rule |
| 400 support tickets | 95.9% | Not tested | 77.2% |
| 1,000 phishing emails, five narrow questions | 95.0% | 93.2%, a tie | 91.8%, a two-line rule |
| Cost per 1,000 emails, five narrow questions | $0.04 | $1.02 | No model fee |
| Speed on the same emails | 5 times faster | The baseline | No model call |
PriorBench: Jev, measured20 Sep 2026
Jev phishing bench17 Sep 2026
Where Jev goes wrong.
The outside study and the maker's own notes agree on most of these. Where they disagree, both are named.
It Always Answers
It put a cake recipe in the technical category at 94% sure. With no “none of these” option, it flagged none of 30 messages that fit no category.
PriorBench: Jev, measured20 Sep 2026
Your Descriptions Decide the Result
When the study swapped the descriptions of two categories, accuracy fell to 16.7%, below the 25% a random pick would get. Write each option plainly and check it.
PriorBench: Jev, measured20 Sep 2026
Text Can Steer It
The maker says instructions hidden in a message “can move the answer”. Keep a person on anything a customer or supplier could word to their advantage.
TypeSafe AI: where Jev 1.13 is weakNo date stated
Numbers and Dates
The maker says Jev is not a calculator and reads dates as text. The outside study found it compared numbers correctly 99.6% of the time. Do not use it to count or to do sums; test any comparison you rely on.
TypeSafe AI: where Jev 1.13 is weakNo date stated
What Jev costs.
TypeSafe AI charges $0.042 for every million tokens you send, and nothing for the answers. A token is about three quarters of a word.
| Messages a month | Short email100 words | Typical request300 words | Long document1,000 words |
|---|---|---|---|
| 1,000 | 1 cent | 2 cents | 6 cents |
| 10,000 | 6 cents | 17 cents | 56 cents |
| 50,000 | 28 cents | 84 cents | $2.80 |
| 100,000 | 56 cents | $1.68 | $5.60 |
| 500,000 | $2.80 | $8.40 | $28.00 |
| 1,000,000 | $5.60 | $16.80 | $56.00 |
A month's cost at the published price, counting the message only. Your options and their descriptions add a little to every call.
TypeSafe AI: models and pricingNo date stated
How to test Jev on your own requests.
An afternoon of labelling tells you more than either study can.
Pull 200 Real Requests
Take 200 recent RFQs, order changes, and complaints your team has already routed, and note where each one went.
You end with200 requests with the right answer
Write Your Options
Write one plain line for each team, and add “none of these” for everything else.
You end withA short list Jev chooses from
Run Them Through Jev
Send all 200 in the same form they arrive, and keep how sure Jev was about each answer.
You end with200 answers, each with how sure it was
Find Your 99% Line
Count how many it got right when it was at least 99% sure, and how many requests that was.
You end withYour own accuracy and share, not a study's
Route Only the Sure Ones
Let Jev route those on its own. Everything below the line goes to a person, as it does today.
You end withFewer requests for your team to sort
Questions about Jev.
Short answers, from the same sources as the rest of this page.
What is Jev 1.13?
A decision model from TypeSafe AI, released on 15 September 2026. It picks one of the options you give it and returns how sure it is. It does not write text.
How accurate is Jev?
In one small outside study it sorted 95.9% of 400 made-up French support tickets right with no training, against 77.2% for keyword rules. In a phishing study it tied with Claude Haiku 4.5 when the question was split into five narrow ones.
Can Jev sort RFQs and orders for a chemical distributor?
It is built for that kind of sorting, but no public test has used a chemical company's inbox. Test it on 200 of your own requests, give it a “none of these” option, and let it act alone only when it is 99% sure.
How much does Jev cost?
$0.042 for every million tokens you send, and the answers are free. At that price, 20,000 messages of 300 words cost about 34 cents a month, before your options and descriptions.
Is Jev better than Claude Haiku 4.5?
For sorting, one study found them tied on accuracy, with Jev about 27 times cheaper and 5 times faster. Haiku can write a reply and Jev cannot, so they do different jobs.
Does Jev work in languages other than English?
The maker says English is its strongest language and others work, but not equally well. The outside study found French and English equally good on 20 hard items.
Does TypeSafe AI train Jev on my messages?
The maker says Jev is not trained on customer requests or responses, and offers zero data retention to enterprise customers.
Sources.
- PriorBench: Jev, measured One outside author, not TypeSafe AI, in a single evening. The 400 test items are made-up French support tickets.Published 20 Sep 2026 · Vailent checked 30 Sep 2026
- Jev phishing bench One outside author, not TypeSafe AI. The 2,000 emails are made up by the dataset's authors around real web addresses.Published 17 Sep 2026 · Vailent checked 30 Sep 2026
- TypeSafe AI: models and pricing The maker's own page.States no date · Vailent checked 30 Sep 2026
- TypeSafe AI: where Jev 1.13 is weak The maker's own page.States no date · Vailent checked 30 Sep 2026