Jev and the case for AI that doesn't chat

A frontier model that can't write a sentence, can't code, and doesn't reason step by step. It classifies, routes and scores in a tenth of a second - and tells you how sure it is.

September 16, 2026 11 min read
What Jev is How it works How it differs from an LLM Where it could be useful Why this matters for industry Other approaches and AGI Our take References

For about four years, progress in AI has mostly meant better chat. Models write longer answers, reason for longer, and code better. So when a lab launches a frontier model that can't write a sentence, can't code and doesn't reason step by step, it's worth paying attention to why.

On 15 September 2026, TypeSafe AI came out of stealth with Jev, the first of what it calls System One Models. The San Francisco company was founded in 2024 by Diogo Almeida, who worked on the instruction-following research at OpenAI that fed into InstructGPT and ChatGPT, together with Erik Gafni and Sasha Sheng. Business Wire reported a seed round of roughly $40 million led by DCVC. Jev is in early access now, and developers are being let in from a waitlist.

A bright automated sorting facility where items on conveyor belts fan out and are routed into many bins

Almeida framed the launch with a question he says has bugged him for years: if chat models have been superhuman at conversation for a while, why hasn't that produced much real automation, or AGI?

“

If chat models have been superhuman at conversation for a while, why hasn't that produced much real automation?

What Jev is

Jev takes messy input - text describing the state of some program, a support ticket, an invoice, a security alert - and returns answers to questions you defined ahead of time. Each answer comes back as a typed value with a probability and a confidence score attached. TypeSafe describes it as a function call backed by frontier-level intelligence. You send unstructured state in and get typed, probabilistic decisions out.

In practice, a support app might ask Jev "which queue does this ticket belong to?" and "how likely is this customer to churn?", then apply its own normal code for routing and permissions. A separate LLM could write the reply to the customer if one is needed.

Where the names come from

System One borrows from Daniel Kahneman's Thinking, Fast and Slow, where System 1 is fast and intuitive and System 2 is slow and deliberate. Jev is a nod to William Stanley Jevons, the economist who noticed that more efficient steam engines made Britain burn more coal, not less. TypeSafe expects the same thing to happen with intelligence as its price falls.

How it works

TypeSafe says it rebuilt the stack around automation. There are three parts.

01 · Architecture
A parallel sampler

An LLM produces output one token at a time, each depending on the last. Jev produces all of its outputs in a single pass - in the launch demo you watch every probability fill in at once while the LLM is still typing.

02 · Output
Constrained, typed answers

You define the possible answers and their structure in advance; the model can only choose from those (up to 255 options per question). That is why TypeSafe says it "can't hallucinate" - meaning the output always matches the schema, not that it's always right.

03 · Training
RLCD, for honest odds

Reinforcement Learning for Calibrated Decisions. The goal is trustworthy probabilities: if Jev says 90%, it should be right about 90% of the time on that kind of question.

The second part is worth being precise about, and TypeSafe is fairly careful itself. Saying Jev "can't hallucinate" means the output always matches the schema. It does not mean the answer is always right. For choices bigger than its 255-option limit, it scores candidates independently and then makes a second, explicit pick.

The third part is the interesting bet. LLMs today are mostly shaped by reinforcement learning from human feedback (RLHF), which rewards answers people like, and reinforcement learning with verifiable rewards (RLVR), which rewards outputs a program can check. TypeSafe's RLCD instead optimises for honest probabilities.

The part to watch

A model that gets a task right 95% of the time is still hard to automate with if it can't tell you when it's in the other 5%. With a trustworthy confidence score, software can act on the confident answers and send the rest to a person or a bigger model. Calibration is the whole game.

How it differs from an LLM

The simplest summary is that Jev trades away strings to gain speed, cost and predictability.

 
Typical frontier LLM
Jev
Trained with
RLHF and RLVR
RLCD
Output
Free text that software has to parse and validate
Typed values chosen from a predefined set, each with a probability
Generation
One token at a time
All outputs in one parallel pass
Latency
Seconds to minutes for frontier models
70 to 500 ms, per TypeSafe
Price
~$0.20 to $10 per M input tokens; output costs more
$0.042 per M input tokens; output free
Best at
Chat, coding, open-ended work with a human checking
Classifying, routing, scoring and extracting inside code

Some numbers to put the pricing in context. The Rundown AI worked out that a million calls with 1,000 input tokens each would cost about $42 at Jev's listed rate. Claude Fable 5.1's base input price is $10 per million tokens, about 238 times higher, though that comparison ignores caching and batch discounts.

Claude Fable 5.1$10.00 / M input tokens
Jev$0.042 / M input · output free

About 238x cheaper on input, with output free - before any caching or batch discounts on the LLM side.

70-500 ms
latency on decision-shaped queries, per TypeSafe
$0.042/M
input tokens, and output is free
40-200x
faster than frontier LLMs on these queries, per TypeSafe

On speed, TypeSafe claims 40 to 200 times faster than frontier LLMs on decision-shaped queries. The headline figures on its homepage (193.6 times faster and 444.6 times cheaper) come from its own workflow evaluations, and the company itself says those are likely at the high end of what customers will see.

How much to trust the benchmarks

TypeSafe built a new kind of evaluation. Each model runs the same workflow, and its answers are compared with the average answer of GPT-6 Astra and Claude Fable 5.1. On that setup, Jev sits well ahead on the speed and cost trade-off. A few caveats, most of which TypeSafe states openly:

Read these with caution
  • Agreeing with two big models isn't the same as being correct. Per AI Wiki's summary of the evals, Jev matched the reference answers 67.8% of the time on average.
  • The workflows were written by people on TypeSafe's own model team.
  • Latency tests mostly ran near its West Coast servers.
  • There were no independent evaluations at launch.
  • TypeSafe notes it can't prove the pricing isn't subsidised yet.

None of this makes the results wrong. It means we should wait for outside testing, especially on whether the probabilities hold up on real customer data.

Where it could be useful

TypeSafe's own list centres on what it calls "smart if-statements": points in ordinary code where a rule is too brittle to write by hand. The examples it gives include customer service triage, invoice handling, security alert review and checking the work of AI agents after a run. A few use cases stand out to us.

Guardrails and verification get cheap

If a check costs fractions of a cent and returns in a tenth of a second, you can score every prompt, every agent step and every output for things like jailbreak attempts or policy problems - rather than sampling a few.

Large-scale data work becomes practical

Running an LLM over billions of records is expensive and slow. Jev is designed for map-reduce jobs that turn large piles of text into structured features.

Real-time products fit

At 70 to 500 ms, a model can sit inside a user-facing flow without making people wait - a place most frontier LLMs are simply too slow to go.

The routing layer itself

Cheap, calibrated decisions are exactly what you want deciding which request is simple enough for a small model and which should escalate to a frontier one.

The launch demos

The two demos make the speed point in a fun way. In one, Jev plays Doom from a text description of the game state - not from pixels - making about 10 decisions a second for roughly $7 an hour. In the other it plays Wikiracing, picking from hundreds or thousands of links at each step to get from one Wikipedia page to another. TypeSafe admits a hand-coded bot would play Doom better; the point was a bot that follows instructions and reacts to whatever state format you give it.

Why this matters for industry

Most businesses don't need a model to write an essay. They need thousands of small, dull, consistent judgements a day: is this invoice a duplicate, is this alert real, does this claim need a human. LLMs can do these, but they are slow, costly at volume, and their output needs parsing and checking. That friction is a big reason so much AI work in companies still looks like a copilot with a person in the loop.

TypeSafe's manifesto makes this argument directly. It says today's models are already smart enough to create huge economic value, and the real bottleneck is that their intelligence is hard to build on. Its comparison is databases before SQL: powerful, but every use was custom.

“

The company's tagline is "build prod, not God."

If that's right, the winners over the next few years won't only be the labs with the biggest models. They'll also include whoever makes judgement cheap and predictable enough to bury five layers deep in a system and forget about. That's a very different market from chat assistants, and it would change how companies budget for AI. You would pay for millions of tiny decisions instead of a few long conversations.

Other approaches and the road to AGI

Jev also belongs to a wider trend. Some of the most credible people in the field are betting that scaling chat-style LLMs isn't the whole answer.

The best-known example is Yann LeCun. After leaving Meta at the end of 2025, he launched AMI Labs in Paris, which raised a $1.03 billion seed round in March 2026. AMI is building world models based on LeCun's Joint Embedding Predictive Architecture (JEPA). Instead of predicting the next word, these models learn abstract representations of how the physical world behaves, so a system can predict what its actions will do before taking them. LeCun has said for years that scaling LLMs won't get us to human-level AI. As of mid-2026, AMI had not shipped a commercial product, and early published results show real gaps in robustness. Google DeepMind is also working on world models through its Genie project, alongside Gemini.

TypeSafe takes a different angle again. It explicitly says it isn't chasing AGI, calling the term a moving goalpost. Its footnotes point to the old idea of neuro-symbolic AI, where neural networks handle perception and judgement and symbolic code handles exact logic. The company's preferred measure of success is economic: global productivity growth reaching 3% a year within five years and staying there for a decade.

We think these efforts matter for anyone interested in general intelligence, for a few reasons.

Human thinking isn't one mode

Kahneman's fast/slow split is a simplification, but it fits what we see: a lot of intelligent behaviour is quick, intuitive and cheap, and only some needs slow deliberation. A system built only from slow, expensive reasoning will struggle to act at the speed the world moves.

Reliability may count as much as capability

A model that knows when it doesn't know is far easier to build on, and far easier to make safe, than one that is right most of the time but can't say when it's wrong. RLCD is a direct bet on that idea.

Monocultures are risky

When almost all money and talent flows into one architecture and one recipe, the field shares the same blind spots. World models, calibrated decision models and neuro-symbolic hybrids each test a different assumption - and the ones that work may end up as parts of a larger system alongside LLMs.

Our take

Jev is early, and most of the evidence so far comes from TypeSafe. Independent tests of calibration, accuracy on real business data and pricing over time will decide whether it lives up to the launch.

Still, the idea behind it strikes us as sound. A lot of useful AI work is small, repetitive and hidden inside software, and it doesn't need a chatbot. If a model can make those calls in a tenth of a second, for almost nothing, and tell you how sure it is, a lot of automation that isn't worth building today suddenly would be. We'll be testing it as soon as we're off the waitlist.

“

A model that knows when it doesn't know is far easier to build on - and far easier to make safe - than one that is right most of the time but can't say when it's wrong.

Route the easy calls cheap. Escalate the hard ones.

Whether the decision model is Jev, a fine-tuned 7B, or a frontier LLM, the pattern is the same: send each request to the right model by cost and confidence, with guardrails on every path. That routing layer is the Strongly.AI AI Gateway.

Explore the AI Gateway

References

TypeSafe and Jev
Coverage and analysis
  • The Rundown AI. (2026, September). TypeSafe launches Jev for AI decisions inside software. therundown.ai
  • Latent Space. (2026, September). [AINews] Jev: a "System One Model" that only decides/classifies/routes/scores. latent.space
  • StartupHub.ai. (2026, September). TypeSafe Jev: The model built to kill chat. startuphub.ai
  • AI Wiki. Jev (AI model). aiwiki.ai/wiki/jev
  • LLM Reference. Jev. llmreference.com/model/jev
World models and other approaches
  • MIT Technology Review. (2026, January 22). Yann LeCun's new venture is a contrarian bet against large language models. technologyreview.com
  • HPCwire / AIwire. (2026, March 11). Yann LeCun's AMI Secures $1B Seed to Develop AI World Models. hpcwire.com
  • Built In. Yann LeCun Launches AMI Labs to Build AI World Models. builtin.com
  • KAD8. (2026, June 22). Yann LeCun's World Model Vision: Inside AMI Labs and JEPA. kad8.com
Background