I went through how people actually wire Jev into their agents. I skipped the demos and looked at the plumbing: where it sits in the loop, what they feed it, and when they stop trusting it.
This is what the working setups have in common, and what happened to the Jev demo with 3.6M views once someone measured it.
What Jev is
Jev is the model TypeSafe AI launched on September 15. It does not generate text. You send it a state and a list of typed questions, and it answers all of them in one parallel pass: pick one of these options, place this on a scale, or how likely is this to be true. Each answer is a set of probabilities your code can branch on. A call takes a few hundred milliseconds and costs a fraction of a cent.
That is the whole interface. It does not chat, call tools or remember the previous call, so anything beyond a single decision has to live in your code.
The one sentence to keep:
Jev is a fast judge with no memory, no hands and no voice. It looks at the state you give it, answers the questions you wrote, from the options you listed. Everything else is your code.
So a Jev harness is three things: what you put in the state, how you word the questions and build the options, and what your code does with the answers.
Six seats people are giving Jev in a harness
I read the top hundred posts of the wave. Under the demos, the harness uses collapse into six seats.
1. Router
Before the expensive model runs, Jev picks which model runs. @dani_avila7 shipped a Claude Code mod that classifies the subagent model, the effort level and the main model. The detail that matters: the main model is routed "only at session start to avoid breaking the cache". LangChain ships the same idea as ModelRouterMiddleware.
2. Gate
A safety check on every command before it executes. @rauchg reports the fx auto-mode reviewer running "5-18x faster and more accurate" on Jev than on the small LLM they used before. LangChain's AutoModeMiddleware is the open version.
A gate that answers in a few hundred milliseconds can run on every call. A slow gate gets turned off.
3. Selector
Computer use and browser use. Jev never sees a screen. The harness observes, builds a list of candidates, Jev picks one ID, the harness executes and looks again.
@trycua has the cleanest statement of the loop: driver observes, perception parses if needed, client builds candidates, Jev selects one ID, driver executes and verifies fresh state. @awlevin did it with OCR boxes as the option list, about $0.003 and 1.5 seconds a step by his numbers, and hands off to a small LLM when something has to be typed.
4. Verifier
@omarsar0 put Jev behind the /goal feature of his harness: a check on whether the goal is actually complete, cheap enough to run far more often than the reasoning model he used before.
A reply under his post is worth more than most demos: a check at the moment the agent says "done" caught the fake finishes, but a check on every turn "started overruling steps the agent needed to try".
5. Advisor
@kunchenguid's compact-adviser does not compact anything. It answers one question: are you at a task boundary where compacting is safe.
He labelled checkpoints from 40 real sessions by hand, tuned the prompt against that set, and made the threshold move: favour precision while the window is mostly empty, shift toward recall as it fills, because the cost of a wrong "no" grows.
6. Triage
@redp314 reviews PRs with one call. 14 typed checks come back as probabilities: hardcoded secret, touches auth, deletes tests, does the description match the diff. Code turns them into BLOCK / security review / nits / merge, and anything between 0.35 and 0.65 on a critical check goes to a human or a bigger model. His cost: $0.00007 a PR.
In all six, Jev never acts on its own. The code builds the list of options, Jev picks one, and the code carries it out. All six also have a rule for what happens when Jev is unsure.

The seat that got audited
The biggest post of the wave, 3.6M views, was a seventh seat: compaction. Score every tool call, drop the irrelevant ones, no summarization prompt. One repost reported a session going from nearly 1M tokens to 86K in a second.
Then people measured it.
@Teknium ran it on the public Nous compaction eval. His finding: what the scoring learned was a rule you can write in one line of code, remove the tool calls. And it decays. The first pass frees half the window, the next pass a quarter, then less, because the only thing it can remove is the thing it already removed.
His numbers against a plain summary: 115K tokens retained versus 55K, recall 76 percent versus 79, and at the same budget the Jev ranking tied with plain recency, 77.8 to 77.8.
An independent tester, @servasyy_ai, reported the same pattern with a Claude Code compaction plugin: it behaved identically when he swapped Jev for a fake model that always answers 0. Several replies to the original pointed at a third problem: every filtered history is a prompt cache miss.
Teknium was clear that this is a problem with that compaction strategy, not with Jev. Nobody had measured the harness before it went viral, and from the outside the demo looks identical whether the scoring inside it works or not.
Nine rules the working harnesses share
1. Code builds the menu, Jev points
Never let it invent options, it cannot. Every selector that works builds the candidate list in code: from the DOM, from OCR boxes, from the accessibility tree, from the retriever, from the tool registry. And they rebuild it after every action, because a click changes what is on the screen. Otherwise, as @0xCodila put it, the model is choosing from yesterday's menu.
The cap is 255 options per question. For bigger lists the pattern is: filter the obvious mismatches in code, score what is left, then choose among the shortlist.
2. Put the exits in the menu
Jev will always pick something. If "none of these" is not an option, you will get a confident wrong pick. @CodingGarden built a tool-calling chatbot with no LLM at all and handles this inside the option list: one option says "no device or room is mentioned at all", another says "more than one tool call can satisfy this request", and either one triggers a follow-up question instead of an action.
3. Give it the facts it cannot know
It has no memory and no clock. @ctatedev found that Jev does not know today's date, and adding the date to the state moved one of his answers from 0.35 to 0.94.
It is also text only, with roughly 32K tokens of context. That limit is where the compaction and disk-cleanup experiments of @servasyy_ai died: to fit the limit he had to cut the content, and once the content was cut the model was judging blind. If the evidence does not fit in the state, the answer is a guess with a probability attached.
4. Ask everything in one request
Questions in the same request are answered in parallel, so twenty questions take about as long as one, and output tokens are free. Rosson's PR review is 14 checks in a single call. The catch: questions cannot read each other's answers. If a decision depends on a search result, run the search first.
5. Decide what "unsure" means before you ship
Every builder quoted above had to define this:
- Rosson sends anything between 0.35 and 0.65 on a critical check to a human or a bigger model.
- @DeRonin_: under 0.5 escalate, 0.85 or more before anything irreversible.
- Kun Chen moves the threshold as the context window fills.
- @servasyy_ai labelled 300 judgments by hand: all 255 answers at 90 percent confidence or higher were correct, and the middle band had too few samples to say anything.
Confidence is not an accuracy percentage. Treat it as a map of where to spend your expensive model.
6. Keep a small LLM on the bench
Jev cannot write. The moment a task needs text (type a URL, fill a form field, explain a finding) the harness hands off. @awlevin calls a small model for type_text. Browser Use does the same for input fields. Rosson runs a second LLM pass only on the PRs that Jev flags. The result is that the expensive model gets called far less often.
7. Respect the prompt cache
A router that switches the main model mid-session throws away the cached prompt, and the saving goes with it. That is why the Claude Code router only routes the main model at session start, and why the replies under the compaction demo were full of "you busted your cache". Any harness that edits history or swaps models has to price the cache miss in.
8. Check at the moment it matters, not on every turn
Cheap checks tempt you to run them constantly. The verifier reports say don't: a check when the agent claims "done" catches fake finishes, a check on every turn starts vetoing steps the agent needed to try. Put the judge at decision points: before an irreversible tool call, at a claimed finish, at a task boundary.
9. Measure it against a dumb baseline before you believe it
The compaction demo tied with "keep the most recent messages". A plugin behaved the same with a model that always answers 0. The projects that held up were measured first: Kun Chen hand-labelled 40 sessions before tuning a prompt, and LangChain ran every judge 100 times per case before writing a word.
A cheap test: replace Jev with a constant answer or a coin flip and rerun. If the result looks the same, Jev was not contributing.
And check the rest of the loop before blaming the model. In the Browser Use numbers @0xCodila collected, cutting browser protocol calls from 1,092 to 101 moved task time more than the model swap did.
A short note from my own testing
I ran a few of these ideas on a task with a free referee, chess, where an engine can grade every single decision. Two results I have not seen posted elsewhere:
When Jev hesitates between its own top options, its probabilities are close to a coin flip, but asking "is A better than B" in both orders picks the better one about two times out of three. And with the same answers in hand, weighting a yes/no veto much harder than a plain multiply cut the loss almost in half, with zero extra calls. It is a small sample from one domain, so treat it as a lead. It did convince me that what the code does with the answers deserves as much attention as the prompt.
The model costs almost nothing now, so the work has moved into the harness.
Which seat did you put Jev in, and what broke? Reply with it. The failures are teaching more than the demos right now.





