What 10,000 black-box API calls reveal about a model that returns probabilities instead of sentences - and why that one change rewrites the whole design.
jev-1.13.0 on one account and one region, with latencies read off the
x-envoy-upstream-service-time header. Published claims, observed behaviour, and
architectural guesses are three different things. This page keeps them labelled.
Ask a normal chat model how sure it is and you get the string "I'm about 91% confident".
That number was generated - sampled token by token, the same way it would write a poem.
It is a plausible-looking sequence of characters, not a measurement of anything inside the model.
Jev takes a different exit. The final hidden state goes through a prediction head - a matrix
W, then a softmax - and the resulting distribution is the answer. No decode
loop, no parsing, no "please respond in JSON". The number never passes through the text channel,
so it never has to be re-read back out of one.
W and softmax? Appendix A and B at the bottom of this page
take both apart with live sliders. Nothing else here depends on reading them first.
You have one support ticket and forty questions about it. The naive approach re-sends the ticket forty times - forty prefills of the same tokens. Jev encodes the shared state once, then runs each question as its own branch that attends back into that state.
Attention masks do the fencing. Every question branch can read the state; no branch can read
another branch. So the answers stay independent while the expensive part gets paid for once.
Compute drops from O(Q × S) to O(S), where Q is questions
and S is state tokens.
Hume planted a secret phrase and then asked unrelated questions whether they could see it. Put it inside Question 2's option descriptions and the other questions score 0.00. Move the same phrase into the shared state and detection jumps to 0.90-0.92. Two facts in one experiment: branches are isolated, and all of them reach the state.
A classifier could plausibly use a bidirectional encoder - every token seeing every other token. Jev does not. It behaves like a standard left-to-right decoder: a token can only attend to what came before it.
The tell is positional. Hume gave the model three options -
[Alpha, Beta, Gamma] - plus a "reference card" stating which one was correct, and
moved the card around. If attention were bidirectional, position would not matter. It mattered a lot.
There are two ways to score a list of options. Pointwise: each option gets a score on its own, then you softmax the pile. Listwise: the model reads the whole list as context, so options influence each other before any score exists.
These make a clean, falsifiable split. Under pointwise scoring, adding an irrelevant fifth option just renormalises - the ratio between two existing options is mathematically untouched. Under listwise scoring, the ratio is free to move.
Hume added "weather caused it" to a four-option payout-failure classification, across 10 blocks. The log-odds between the surviving options shifted by -0.28 (95% CI -0.36 to -0.19). Pointwise predicts exactly 0. It is not 0.
TypeSafe describes a post-training stage they call RLCD - Reinforcement Learning for Calibrated Decisions. The methodology is unpublished, but the fingerprint is measurable: when Jev says 0.8, it is right about 80% of the time.
Benchmark calibration can be contamination. So Hume generated new maths problems and compared claimed probability against realised accuracy: 3-digit multiplication drew 83% confidence against 86.7% accuracy; word problems drew 30% against 32%. The model does not just know the answers - it tracks how likely it is to be wrong, and the tracking moves with task difficulty.
The latency numbers are the loudest evidence in the whole investigation. A 255-option response was billed at 2,714 output tokens and came back as fast as a 2-option question. That only makes sense if output tokens are a billing calculation applied after the fact, not a count of generation steps.
What actually scales is the state. One question over a 29.8k-token state: 214ms. Fifteen hundred questions over a short state: ~600ms. Prefill dominates; repeated state gets cheaper, which is the signature of prefix caching. And ~30k tokens in ~160ms is hard to hit with dense parameters - conditional computation, most likely a sparse mixture-of-experts, fits the budget.
escalate if p(urgent) > 0.1. The threshold
lives in your code, in the open, where you can tune it. That is the opposite of reading a
self-reported confidence out of a sentence.
One design choice - return a probability instead of a sentence - cascades into everything else: no decode loop, so speed stops scaling with option count; a state/question split, so compute collapses to O(S); listwise options, so the list is context rather than a scoring queue; and an outcome-trained head, so the number is worth putting a threshold on.
When the transformer finishes reading, it leaves behind a hidden state -
written h. It is just a list of numbers, typically a few thousand of them. Nobody
assigns those numbers meanings; training does, and the meanings end up smeared across many
dimensions at once. But the useful mental model is: h is everything the model noticed,
compressed into coordinates.
W is a matrix of learned weights that turns those coordinates into one score per
option. If h has d numbers and there are n options, then
W is d × n - and each column is a direction in
h-space that means "this option". The score is a dot product: multiply matching entries, add
them up.
That is the entire "prediction head". A matrix multiply. No loop, no sampling, no text - which is why it costs one step instead of one step per token.
h with 4 readable dimensions instead of 4,096. Drag them.
d × n matrix only works when
n is fixed. Jev takes a different option list on every call - up to 255 of them. So
the columns almost certainly are not stored weights; they are built from each option's own text
as the model reads it. Same arithmetic, same dot product, but the "directions" arrive with the
request. That is also why options can influence each other (section 4).
Logits are unbounded - 1.55, -0.81, 0.22. You cannot hand
those to a downstream rule. Softmax exponentiates each one (killing the negatives, since
exp is always positive) and divides by the total, so the results are positive and
sum to exactly 1.
1. Shift-invariant. Add the same constant to every logit and the output does not move at all - the constant cancels top and bottom. Only differences between logits carry information.
2. Ratios are logit differences. Divide two probabilities and the shared denominator cancels:
3. Temperature sharpens or flattens. Divide the logits by T before exponentiating. Small T pushes everything toward a single winner; large T flattens toward a uniform guess.
ln(p₁/p₂) sit perfectly still while every probability moves. That immovability is
not a modelling choice - it is arithmetic, true of any pointwise scorer. Jev moved it by
-0.28. So the logits themselves must have changed when the new option was
added, meaning the options were read together rather than scored apart.
No spam — just new posts and learning pages when they ship.
Drop your email and I will send new posts and learning pages your way.
Thanks — you are subscribed.
You can close this dialog now.
We will only email you when there is something new. Unsubscribe any time.