Model Architecture · 2026

Jev's Architecture, Unmasked

What 10,000 black-box API calls reveal about a model that returns probabilities instead of sentences - and why that one change rewrites the whole design.

0 / 6 sections
Read this first. Nothing here comes from TypeSafe's source code. Every claim is an inference from black-box API behaviour - 1,029 initial probes plus 446 follow-up trials against jev-1.13.0 on one account and one region, with latencies read off the x-envoy-upstream-service-time header. Published claims, observed behaviour, and architectural guesses are three different things. This page keeps them labelled.

▶ One request, end to end

0 ms
Press Play to watch one support ticket and three questions travel the whole pipeline. The clock on the right is the real latency budget - watch which stage eats it.
1
Numbers, not sentences
the probability readout
▶

Ask a normal chat model how sure it is and you get the string "I'm about 91% confident". That number was generated - sampled token by token, the same way it would write a poem. It is a plausible-looking sequence of characters, not a measurement of anything inside the model.

Jev takes a different exit. The final hidden state goes through a prediction head - a matrix W, then a softmax - and the resulting distribution is the answer. No decode loop, no parsing, no "please respond in JSON". The number never passes through the text channel, so it never has to be re-read back out of one.

headsoftmax(W · h)
decode steps0
outputfloat[n_options]
The catch. A softmax will hand you numbers whether or not they mean anything. Direct readout only buys you calibration if the head was trained against real outcomes. That is section 5.
New to W and softmax? Appendix A and B at the bottom of this page take both apart with live sliders. Nothing else here depends on reading them first.
Demo · two ways to produce "91%"
Press run to generate an answer.
Decode steps
-
Wall time
-
Parseable
-
A chat model writes "I am 91% confident". Why is that not a probability?
2
One shared state, many isolated questions
O(Q×S) → O(S)
▶

You have one support ticket and forty questions about it. The naive approach re-sends the ticket forty times - forty prefills of the same tokens. Jev encodes the shared state once, then runs each question as its own branch that attends back into that state.

Attention masks do the fencing. Every question branch can read the state; no branch can read another branch. So the answers stay independent while the expensive part gets paid for once. Compute drops from O(Q × S) to O(S), where Q is questions and S is state tokens.

The visibility test

Hume planted a secret phrase and then asked unrelated questions whether they could see it. Put it inside Question 2's option descriptions and the other questions score 0.00. Move the same phrase into the shared state and detection jumps to 0.90-0.92. Two facts in one experiment: branches are isolated, and all of them reach the state.

Demo · plant the secret, then probe
Pick a placement, then probe the other questions.
Demo · what the shared state saves
Questions asked about one 6,000-token ticket: 40
Re-send per question · O(Q×S)
240,000
state tokens encoded
Shared state · O(S)
6,000
state tokens encoded
40× less state encoding.
The secret phrase scores 0.00 from Q1 when planted in Q2's options, but 0.91 when planted in the shared state. What does that pair of results establish?
3
Causal, not bidirectional
position leaks the mask
▶

A classifier could plausibly use a bidirectional encoder - every token seeing every other token. Jev does not. It behaves like a standard left-to-right decoder: a token can only attend to what came before it.

The tell is positional. Hume gave the model three options - [Alpha, Beta, Gamma] - plus a "reference card" stating which one was correct, and moved the card around. If attention were bidirectional, position would not matter. It mattered a lot.

Why last beats first. A card placed before the option descriptions is visible to them, but the descriptions then talk past it. A card placed after all descriptions sits right where the decision is read - and causal attention lets the decision position look back over everything, card included. Put it in the shared state and every branch gets it cleanly.
Demo · move the reference card
Token order (left to right) · green = can see the card
Accuracy
-
Pick a placement for the reference card.
Why does moving the reference card from first to last raise accuracy from ~50% to ~88%?
4
Options compete with each other
listwise, not pointwise
▶

There are two ways to score a list of options. Pointwise: each option gets a score on its own, then you softmax the pile. Listwise: the model reads the whole list as context, so options influence each other before any score exists.

These make a clean, falsifiable split. Under pointwise scoring, adding an irrelevant fifth option just renormalises - the ratio between two existing options is mathematically untouched. Under listwise scoring, the ratio is free to move.

Hume added "weather caused it" to a four-option payout-failure classification, across 10 blocks. The log-odds between the surviving options shifted by -0.28 (95% CI -0.36 to -0.19). Pointwise predicts exactly 0. It is not 0.

Demo · add an irrelevant option, watch the ratio
Pointwise prediction
Observed from Jev
ln ratio · pointwise
1.082
ln ratio · observed
1.082
Shift
0.00
Ratio measured is p(customer) / p(unknown). Toggle the fifth option.
Adding an irrelevant option changed the log-odds between two existing options by -0.28. What does that rule out?
5
Calibration comes from outcomes
ECE 0.031 over 1,200 items
▶

TypeSafe describes a post-training stage they call RLCD - Reinforcement Learning for Calibrated Decisions. The methodology is unpublished, but the fingerprint is measurable: when Jev says 0.8, it is right about 80% of the time.

ECE0.031
items1,200 MMLU
in 0.9-1.0 bin990 / 1,200
MMLU-Pro84.6%

The freshly-generated check

Benchmark calibration can be contamination. So Hume generated new maths problems and compared claimed probability against realised accuracy: 3-digit multiplication drew 83% confidence against 86.7% accuracy; word problems drew 30% against 32%. The model does not just know the answers - it tracks how likely it is to be wrong, and the tracking moves with task difficulty.

Demo · reliability diagram, and what ECE measures
predicted confidence observed accuracy
low confidencehigh confidence
Drag to inject overconfidence: 0.00
Expected calibration error
0.031
Verdict
calibrated
Click a bar to inspect that bin. A perfectly calibrated model has both bars equal in every bin.
Why did Hume test calibration on freshly generated maths problems as well as MMLU?
6
Where the speed comes from
~30k tokens in ~160ms
▶

The latency numbers are the loudest evidence in the whole investigation. A 255-option response was billed at 2,714 output tokens and came back as fast as a 2-option question. That only makes sense if output tokens are a billing calculation applied after the fact, not a count of generation steps.

What actually scales is the state. One question over a 29.8k-token state: 214ms. Fifteen hundred questions over a short state: ~600ms. Prefill dominates; repeated state gets cheaper, which is the signature of prefix caching. And ~30k tokens in ~160ms is hard to hit with dense parameters - conditional computation, most likely a sparse mixture-of-experts, fits the budget.

Demo · the request pipeline (hover a stage)
📥Prefill shared state O(S)
Ticket text encoded once. Repeat requests with the same prefix hit the cache, which is why a 29.8k-token state still returns in ~214ms.
↓
🔀Fan out question branches batch units
Each question is scheduled as an independent work unit, not a turn in a conversation. 1,500 of them fit in ~600ms.
↓
🧠Causal transformer · sparse MoE inferred
Left-to-right decoder. Sparse experts are an inference from the latency budget, not something the API exposes directly.
↓
📋Read full option list listwise
All options enter as context before any score exists, which is why a fifth irrelevant option moves the ratio between the other four.
↓
📊Probability head 1 step
softmax(W · h) per question. One decision, no decode loop - which is why 255 options cost no more wall time than 2.
Demo · latency race
2 options
255 options
29.8k-token state
1,500 questions
if it decoded 2,714 tok
Scale is 0-700ms. The last bar is hypothetical: what the same response would cost if those 2,714 billed tokens were actually generated one at a time.
Still unknown. The tokenizer matches none of 192 public ones tested - Qwen is closest at 348/415 probe matches, but digit splitting and merge patterns differ. The base model is not identified. The MoE is inferred from latency, never observed. And the RLCD loss function - log loss, Brier, some other proper scoring rule - is anyone's guess.
Why it matters in practice. The design separates three jobs: the transformer builds a useful representation, the outcome-trained head turns it into a number you can trust, and your policy decides what to do - escalate if p(urgent) > 0.1. The threshold lives in your code, in the open, where you can tune it. That is the opposite of reading a self-reported confidence out of a sentence.
A 255-option response was billed 2,714 output tokens but returned as fast as a 2-option one. What follows?

🎉 All six sections opened

One design choice - return a probability instead of a sentence - cascades into everything else: no decode loop, so speed stops scaling with option count; a state/question split, so compute collapses to O(S); listwise options, so the list is context rather than a scoring queue; and an outcome-trained head, so the number is worth putting a threshold on.

Appendix

the two pieces of maths the whole design rests on · optional, not counted in progress
A
What "matrix W" actually is
the prediction head
▶

When the transformer finishes reading, it leaves behind a hidden state - written h. It is just a list of numbers, typically a few thousand of them. Nobody assigns those numbers meanings; training does, and the meanings end up smeared across many dimensions at once. But the useful mental model is: h is everything the model noticed, compressed into coordinates.

W is a matrix of learned weights that turns those coordinates into one score per option. If h has d numbers and there are n options, then W is d × n - and each column is a direction in h-space that means "this option". The score is a dot product: multiply matching entries, add them up.

logiti  =  h · W[:, i]  =  h₁W₁ᵢ + h₂W₂ᵢ + … + h_dW_dᵢ one number per option · can be negative · not yet a probability

That is the entire "prediction head". A matrix multiply. No loop, no sampling, no text - which is why it costs one step instead of one step per token.

Demo · move the hidden state, watch the logits
A toy h with 4 readable dimensions instead of 4,096. Drag them.
W · each column is one option's direction (fixed by training)
Resulting logits
One honest wrinkle. A fixed d × n matrix only works when n is fixed. Jev takes a different option list on every call - up to 255 of them. So the columns almost certainly are not stored weights; they are built from each option's own text as the model reads it. Same arithmetic, same dot product, but the "directions" arrive with the request. That is also why options can influence each other (section 4).
What does one column of W represent?
B
What softmax does to those logits
and the property section 4 tests
▶

Logits are unbounded - 1.55, -0.81, 0.22. You cannot hand those to a downstream rule. Softmax exponentiates each one (killing the negatives, since exp is always positive) and divides by the total, so the results are positive and sum to exactly 1.

pi  =  exp(zi) / Σj exp(zj) z = logits · output is a probability distribution over the options

Three properties worth internalising

1. Shift-invariant. Add the same constant to every logit and the output does not move at all - the constant cancels top and bottom. Only differences between logits carry information.

2. Ratios are logit differences. Divide two probabilities and the shared denominator cancels:

ln( pi / pj )  =  zi − zj this is the "log-odds" quantity measured in section 4

3. Temperature sharpens or flattens. Divide the logits by T before exponentiating. Small T pushes everything toward a single winner; large T flattens toward a uniform guess.

Demo · logits in, probabilities out
temperature T
1.00
1 · raw logits z
2 · exp(z)
3 · ÷ total
4 · probabilities
Sum
1.000
ln(p₁/p₂)
2.360
z₁ − z₂
2.360
Now section 4 reads differently. Press "Add a 4th option" above and watch ln(p₁/p₂) sit perfectly still while every probability moves. That immovability is not a modelling choice - it is arithmetic, true of any pointwise scorer. Jev moved it by -0.28. So the logits themselves must have changed when the new option was added, meaning the options were read together rather than scored apart.
You add a fifth option to a pointwise scorer. What happens to ln(p₁/p₂) for two existing options?
Learning Reference · Jev's Architecture Unmasked - Archer Hume

Get new posts in your inbox

No spam — just new posts and learning pages when they ship.

Share X LinkedIn Reddit Hacker News