Back to articles

Jev by TypeSafe AI: the hype, reactions, and two-week clone war

Oct 4, 2026Dishant Sharma5 min read
Jev by TypeSafe AI: the hype, reactions, and two-week clone war

Somewhere in a roleplay community, people are feeding their own chat replies into a model built for industrial automation. The model scores their writing and tells them which reply sounds more in character. That model is Jev, and it is not what TypeSafe AI put in its press release.

i found this at 1 in the morning, scrolling thread titles i did not expect. Jev came out of stealth on September 15, 2026. TypeSafe AI is founded by Diogo Almeida, who helped build the instruction-following work inside OpenAI that became ChatGPT.

The pitch is simple. Jev never writes text. You send program state, ask typed questions, and get back probabilities in 70 to 500 milliseconds. Output tokens cost nothing.

The reaction split in a familiar way. r/LocalLLaMA ran 330 comments calling it old tech in a new costume. Someone in r/BetterOffline put it more simply:

Jev is Jian Yang's Hotdog / Not Hotdog app.

It is a Silicon Valley joke, and it landed.

And then the odd part. The busiest threads were in r/SillyTavernAI, a roleplay tool, and r/SideProject, people scoring their startup ideas. Two weeks later, OpenAI shipped a Decisions API built on Luna, and AWS shipped a local model called Strands Decider 2B.

TechCrunch ran the headline "Amazon releases its own Jev clone as decision models flood the web."

Here's why you should care. If the next phase of AI is models that stop talking, the way we build agents changes. And i don't think anyone has fully processed that yet.

What Jev actually is

Here's a question people keep asking: is Jev another chatbot? No. Jev is a discriminative model. It classifies and scores. It does not generate.

You give it a state, usually text or JSON, plus typed questions. It answers all of them in one call, in parallel. The whole vocabulary fits in three primitives:

  1. noul: a yes or no question, answered as a 0 to 1 probability.
  2. choice: pick from a fixed list, each option gets a probability.
  3. score: rate the input against a numerical scale or rubric.

That is the entire API surface.

There is no free text anywhere.

i kept seeing Jev mentioned and assumed it was a chat model. A thread on r/PromptEngineering said the same, and the top reply accused the poster of advertising. The confusion is fair. We are not used to a model that outputs nothing but verdicts.

The endpoint is a single line: api.typesafe.ai/v1/systemone.

Why 200 milliseconds matters

What actually happens is that Jev skips token generation. Normal models emit one token at a time, each conditioned on the last. Jev samples every answer at once, in parallel, hardware aware. That is where the speed comes from.

The training is different too. RLHF optimizes for what human raters prefer. RLCD, short for Reinforcement Learning for Calibrated Decisions, optimizes for honest probabilities. High confidence means high accuracy, and the model is built so you can trust the number.

That matters more than speed. If a model can do a task 95 percent of the time but cannot say when it is in the failing 5 percent, you cannot wire it into anything.

Jev Chat LLM
Output typed decisions free text
Latency 70 to 500 ms 3 to 329 s
Confidence calibrated often overconfident

The skeptics have a point

The loudest take on r/LocalLLaMA is that Jev is not new. A classifier with calibrated probabilities is a solved problem. TypeSafe has not published enough architecture detail to verify the interesting claims.

i could be wrong here, but the architecture silence bothers me more than the price. The company itself admits it cannot prove the pricing is not subsidized.

Input tokens cost $0.042 per million. Output tokens are free. That price is a bet, not a fact.

Jev also has a documented gap. It does not have deep knowledge of niche domains. You supply the context, which means you are doing some of the work. That is fine for routing. It is not fine for expert judgment.

Nobody expected the response

i used to think a model needed years to matter. Jev got cloned in two weeks.

OpenAI announced its Decisions API ahead of DevDay, built on the small Luna model, about 150 milliseconds per call. AWS shipped Strands Decider 2B, small enough to run locally. Two giants, two weeks, two takes on the same idea.

The strangest adoption is the roleplay one. People wired Jev into SillyTavern to score replies. Average response around 200 milliseconds, at roughly $0.0005 per call. The most human use of a machine-native model i have seen.

And builders went further. There is Claude Code memory tooling built on Jev. Idea scoring side projects with 130 comments. Cloudflare Workers running it. Everyone is bolting a verdict machine onto whatever they already made.


A quick detour about naming

All of this traces back to one book. Jev is named after Daniel Kahneman's System 1, the fast intuitive thinking from Thinking, Fast and Slow.

Every AI cycle, someone rediscovers that book and names a product after it. We survived years of System 2 reasoning marketing, and now the pendulum swings to System 1.

i do the same thing. There is a script in my home directory called "yesterday" because i was reading Murakami when i wrote it.

It has nothing to do with time travel. It parses CSV files.

Naming is a mood ring.

The funny part: Kahneman thought fast thinking was the source of bias. The confident automatic judgments were the ones to distrust. Half of the Jev hype is celebrating exactly that, which means we are misusing the book again.

But that argument can wait.


The honest version

Most people do not need Jev. If you make under a few hundred decisions a day, a chat model with a careful prompt works fine.

A normal classifier works fine too, as long as your labels do not change often. The calibrated probabilities only matter at volume, where a silent wrong answer costs real money.

And this is a bet. The pricing looks subsidized, and the company says so. The architecture is underpublished.

The clones are already here, and a decision model is easier to copy than a frontier model. If OpenAI's and AWS's versions are good, the moat is thin.

"cannot hallucinate" is also narrower than it sounds. Jev cannot type nonsense because it cannot type. It can still decide wrong, confidently, and your pipeline will trust it.

Last thought

i keep thinking about the roleplay people. A model built to route transactions and guardrail agents.

The most boring enterprise tool imaginable, and the community that adopted it fastest was people grading each other's fiction.

That is the part nobody predicted. The enterprise will take years to trust calibrated probabilities. The weird corners of the internet plugged it in and shipped.

Usually it works the other way around.

So here is the question i keep circling. When your model stops talking, what changes about how you build?

And the sharper one: what were you using the chat for in the first place?

Recent posts

View all posts