Each preset puts a document in the prompt that contradicts something the model is certain of. The conflict therefore exists whatever the model was trained on, which a real-world update cannot guarantee. Press Measure to see which one comes out.
The two distributions, side by side
Probability of each candidate without the document (grey) and with it (blue). The correct answer is outlined in green. If its bar is invisible, that is the finding.
50%4.6
1%9.2
0.01%
inert0.5
reshaping
Measurement details
| P(answer | query alone) | |
| P(answer | context + query) | |
| Answer tokenization | |
| Dsync definition | −ln P(answer given context and query): the negative natural logarithm of the correct answer's probability |
| Ictx definition | Kullback-Leibler divergence between the with-document and without-document distributions, over all 50,257 tokens |
The five stages
| Acquire | Pre-training compresses a corpus into weights. | Facts stored with no record of when or where they were learned. |
| Store | Facts live in weights, external indices, or both. | The two copies of a fact age independently. |
| Retrieve | Attention recalls; RAG fetches documents into context. | What was retrieved may not control the output. |
| Update | Weights edited, indices refreshed, conflicts resolved. | Edits damage neighbors; index updates leave weights stale. |
| Forget | Deliberate unlearning; accidental forgetting and eviction. | Removed facts recoverable; retained facts damaged. |
The diagnosis grid
| Ictx low | Ictx high | |
|---|---|---|
| Dsync low | fact never conflicted | the answer is within reach of the context |
| Dsync high | the answer is far down the distribution | context influential, answer still not first |
The three presets
| conflict | Counterfactual capital. The document contradicts a fact the model is certain of, so the conflict does not depend on when the model was trained. |
| rank | Watch the rank, not the bar. Across the paper's 432 document conditions the in-context answer carries more than 0.10 probability 332 times, and in 69 of those it is still not the token the model would emit. |
| control | When the fact is not held. GPT-2 base holds the plain capital only 25% of the time. Where it never held the fact there is no conflict to observe, which is why the paper reports its headline numbers only on the cases where it did. |
Terms used on this page
| Token | The unit a model reads and writes. Roughly a word or word fragment; " withdrawn" is one token, and the leading space is part of it. |
| Parametric memory | What the model absorbed into its weights during training. Fixed after training, and the model cannot tell you when it learned any of it. |
| Context | What you put in the prompt right now, including any document retrieved for the model to read. |
| Nat | A unit of information, measured with natural logarithms. Here it converts directly to probability: a value of n nats means the correct answer holds probability e−n. One nat is roughly 37%, nine nats is roughly one in ten thousand. |
| Surprisal | How surprised the model is by an answer, written as −ln P. Low when the model expected it, high when it did not. Dsync is the surprisal of the in-context answer. It is not the answer: probability spreads over tens of thousands of tokens, so a modest value can still be the token the model emits, and a large one need not mean the document was ignored. |
| KL divergence | A measure of how far one probability distribution sits from another. Ictx uses it to ask whether the document changed the model's mind about anything at all. |
| Temperature | How randomly a model picks among candidate answers. Low values make it repeat its favorite; high values spread the choice out. |
| Quantized | Weights stored at reduced precision so the model downloads and runs faster. This page uses 8-bit weights, which shift individual probabilities slightly without changing any conclusion. |
Build your own probe
Any fact change with a single-word answer works. Three rules make a clean probe: the query ends mid-sentence so the next token is the answer; the document states the new fact plainly; the answer is one word, because the estimator measures the first token. Leadership changes, product renames, and policy reversals all fit.
The poster
The whole argument on one page: the problem, the five-stage framework, the finding, the metric, and the proposed architecture. Open the full-resolution A0 PDF.
