Similar charts tell us more about range than direction
The finding
Tighter neighborhoods are genuinely more similar in what followed: the dispersion ratio falls monotonically as k shrinks, at every horizon. The tails narrow; the interquartile range barely moves. The state definition carries real information, but about how much, not which way.
The research question
Does ranking by shape within a situation carry information, or would k random members of the same situation have the same spread of outcomes?
Decision criteria
Against k random members of the same situation (identity bucket, same as-of cutoff, same 2-per-symbol cap), the k nearest by shape must have narrower 5-day outcome dispersion, with the ratio's interval clear of 1 and falling as k shrinks.
Result document
Repository source: research/results/study100_similarity_dispersion_2026_09_04.md
Reading (2026-09-04 12:30Z, written after the numbers; the bar was fixed before them). PASS. 478 subjects. Against k random members of the SAME situation (identity bucket, same as-of cutoff, same 2-per-symbol cap), the k nearest by shape have narrower 5 d outcome dispersion at every k, and the ratio falls monotonically as the neighborhood tightens: R(std) = 0.916 at k=200, 0.881 at 50, 0.865 at 20, 0.791 at 5; R(200) − R(5) = +0.125, CI90 [0.084, 0.193]; R(5)'s CI90 [0.721, 0.826] is clear of 1. Same picture at 1 d (0.921 → 0.826) and 10 d (0.918 → 0.811). So the state definition carries real information: tighter neighborhoods are genuinely more similar in what followed.
Where the narrowing lives: in the tails. The 80 % width narrows (R 0.975 → 0.935 from k=200 to 20) and the std narrows, but the interquartile range does NOT narrow at k ≥ 20 (R IQR 1.001 / 0.997 / 0.989) — only at k=5 (0.781). Shape similarity within a situation removes extreme outcomes among the nearest analogs more than it compresses the central half.
The construction confound was large: the RAW real std falls 37 % from k=200 to k=5 (4.73 → 2.96 pp), but the random-k control falls 30 % over the same range (5.44 → 3.81) from small-sample bias alone. Without the control the whole 37 % would have been credited to similarity; the honest number is the ratio.
The floor: the ratio is still falling at the tightest threshold tested (k=5, mean distance 0.307), so the irreducible-noise floor is NOT reached within k ∈ [5, 200]. The practical bound at this resolution: at the tightest servable neighborhood about 79 % of the within-situation outcome dispersion remains; at the served k (200–300) about 92 %. The situation itself (identity) is the larger lever (studies 90/91 measured the analog SET vs the base rate); the shape ranking adds ~8 % at served k and ~21 % at k=5.
Not licensed by this: any direction claim (none was measured). Candidate follow-up (its own registration): distance-weighted bands at k=200 (weight nearer analogs more) vs the unweighted band — measure band width at held coverage; the 21 % at k=5 is the ceiling of what such weighting could recover. Runtime 3.5 min at low priority with the clustered rank table.
Run 2026-09-04T12:21:15.896410+00:00 (smoke=False). Subjects 478 kept of 500 sampled ({'sampled': 500, 'no_embedding': 0, 'retrieve_not_ok': 0, 'short_set': 17, 'short_bucket': 5, 'kept': 478}). Pre-registration: research/specs/study100_similarity_dispersion_prereg_2026_09_04.md (ledger 100).
1d
| k | tightness | real std | random std | R std (CI90) | real IQR | random IQR | R IQR | real W80 | random W80 | R W80 |
|---|---|---|---|---|---|---|---|---|---|---|
| 5 | 0.3069 | 1.382 | 1.702 | 0.826 [0.773,0.902] | 1.171 | 1.374 | 0.848 [0.770,0.916] | None | None | - |
| 20 | 0.3419 | 1.699 | 2.063 | 0.882 [0.835,0.916] | 1.659 | 1.808 | 0.954 [0.922,0.997] | 3.53 | 3.912 | 0.921 [0.900,0.959] |
| 50 | 0.3711 | 1.898 | 2.215 | 0.878 [0.851,0.916] | 1.729 | 1.847 | 0.980 [0.949,1.020] | 3.759 | 4.225 | 0.944 [0.908,0.971] |
| 200 | 0.4314 | 2.083 | 2.396 | 0.921 [0.893,0.953] | 1.804 | 1.938 | 0.998 [0.981,1.022] | 3.987 | 4.432 | 0.965 [0.941,0.985] |
primary (std): {'R5': 0.826, 'R200': 0.9207, 'diff_R200_minus_R5': 0.0947, 'diff_ci90': [0.0332, 0.147], 'n': 478} · floor: {'k': 5, 'R_at_floor': 0.826, 'R_by_k': {5: 0.826, 20: 0.8819, 50: 0.8778, 200: 0.9207}, 'monotone_ratio': False, 'monotone_raw_std': True}
5d
| k | tightness | real std | random std | R std (CI90) | real IQR | random IQR | R IQR | real W80 | random W80 | R W80 |
|---|---|---|---|---|---|---|---|---|---|---|
| 5 | 0.3069 | 2.959 | 3.814 | 0.791 [0.721,0.826] | 2.782 | 3.23 | 0.781 [0.709,0.861] | None | None | - |
| 20 | 0.3419 | 3.898 | 4.781 | 0.865 [0.833,0.916] | 3.846 | 4.193 | 0.989 [0.931,1.032] | 8.023 | 9.094 | 0.935 [0.893,0.976] |
| 50 | 0.3711 | 4.355 | 5.135 | 0.880 [0.860,0.934] | 4.198 | 4.319 | 0.997 [0.960,1.021] | 8.831 | 9.603 | 0.940 [0.906,0.989] |
| 200 | 0.4314 | 4.731 | 5.442 | 0.916 [0.897,0.938] | 4.258 | 4.54 | 1.001 [0.979,1.019] | 9.397 | 10.291 | 0.975 [0.954,0.995] |
primary (std): {'R5': 0.7912, 'R200': 0.9163, 'diff_R200_minus_R5': 0.1251, 'diff_ci90': [0.0844, 0.1929], 'n': 478} · floor: {'k': 5, 'R_at_floor': 0.7912, 'R_by_k': {5: 0.7912, 20: 0.8645, 50: 0.8805, 200: 0.9163}, 'monotone_ratio': True, 'monotone_raw_std': True}
10d
| k | tightness | real std | random std | R std (CI90) | real IQR | random IQR | R IQR | real W80 | random W80 | R W80 |
|---|---|---|---|---|---|---|---|---|---|---|
| 5 | 0.3069 | 4.552 | 5.742 | 0.811 [0.780,0.856] | 4.352 | 4.882 | 0.932 [0.830,1.025] | None | None | - |
| 20 | 0.3419 | 5.702 | 6.886 | 0.876 [0.828,0.919] | 5.679 | 6.118 | 0.973 [0.926,1.006] | 11.959 | 13.479 | 0.959 [0.917,1.004] |
| 50 | 0.3711 | 6.115 | 7.337 | 0.909 [0.873,0.937] | 6.038 | 6.451 | 1.022 [0.988,1.060] | 13.338 | 14.352 | 0.964 [0.943,0.993] |
| 200 | 0.4314 | 6.605 | 7.893 | 0.918 [0.885,0.940] | 6.31 | 6.737 | 1.026 [1.000,1.046] | 13.793 | 15.046 | 0.999 [0.971,1.017] |
primary (std): {'R5': 0.8114, 'R200': 0.9182, 'diff_R200_minus_R5': 0.1068, 'diff_ci90': [0.0596, 0.1433], 'n': 478} · floor: {'k': 5, 'R_at_floor': 0.8114, 'R_by_k': {5: 0.8114, 20: 0.8764, 50: 0.9093, 200: 0.9182}, 'monotone_ratio': True, 'monotone_raw_std': True}
Study specification
Repository source: research/specs/study100_similarity_dispersion_prereg_2026_09_04.md
Registered 2026-09-04 ~12:20 UTC, before any number was pulled. Graham: "Test whether cohort similarity actually narrows forward outcome dispersion. For a set of subject symbol-date pairs, pull comps at several similarity thresholds — top five, top twenty, top fifty, top two hundred nearest analogs. At each threshold compute the dispersion of forward returns across cohort members: standard deviation, interquartile range, the width of the eighty percent interval. Then plot dispersion against cohort tightness. The hypothesis is that dispersion falls monotonically as similarity increases. If it does, the state definition carries real information and tighter neighborhoods mean genuinely more similar states. If dispersion flattens out past a certain tightness, that floor is the irreducible noise in the state representation, and it tells us the practical limit of forward predictability."
This is a measurement of the memory's RESOLUTION (how much, never which way). It ships nothing and trades nothing.
The confound, and the control that removes it
The sample dispersion of five members is smaller and noisier than that of two hundred by construction, so
"dispersion falls with k" is true even for random members. The question is whether the top-k by SHAPE within the
same SITUATION are tighter than k random members of that same situation. Control: for each subject and each k, k
random rows from the subject's own identity bucket (situation_context_1d, same happening_ident, dates ≤ the
same as-of cutoff, the same 2-per-symbol cap, the subject excluded), three disjoint draws averaged. The statistic
is the per-subject ratio real / random.
Frozen design
- Subjects: 500 random liquid (dollar_vol ≥ 5e7, close ≥ 5)
situation_contextrows with a happening identity and a 1d V5 embedding, dates 2022-01-01 … 2026-08-15 (so 10 d forward exists), TABLESAMPLE draw, seed 100. - Comps: the production happening-then-shape retrieve (
retrieve_happening_shape, scale 1d, k = 200, as-of cutoff = date − 15 d, 2-per-symbol cap), members sorted by distance; thresholds k ∈ {5, 20, 50, 200} are prefixes. Subjects whose retrieve returns < 200 members are skipped and counted. - Outcomes: date-matched excess forward return at 1 / 5 / 10 sessions (
forward_returns_cacheminusliquid_date_medians), in percentage points. Primary horizon 5 d. - Dispersion per subject per k: sample standard deviation, interquartile range, 80 % width (p90 − p10; reported only for k ≥ 20). Tightness = mean embedding distance of the prefix.
- Aggregation: median across subjects of each measure (real and random) per k; per-subject ratio real / random; median ratio with a 2,000-resample bootstrap CI90 over subjects.
PRIMARY (one bar, pre-stated)
R(k) = median per-subject ratio of 5 d standard deviations, real / random.
PASS ("similarity narrows outcomes") iff R(5) < R(200) with the bootstrap CI90 of the difference R(200) − R(5)
clear of zero AND R(5) < 1 with CI90 clear of 1.
FAIL otherwise: within a situation, the shape ranking does not narrow width; the situation IS the state.
Reported either way:
- the monotonicity of R(k) across 200 → 50 → 20 → 5 (and of the raw real dispersion, which is expected to fall regardless — the ratio is the claim);
- the floor: the largest k at which R(k) is within 5 % (relative) of R(5) — beyond that tightness, no further narrowing; R at the floor is the share of dispersion the representation cannot remove (irreducible noise);
- the same for IQR and, for k ≥ 20, the 80 % width; and for 1 d and 10 d.
Pre-stated readings:
- PASS with R(5) well below 1 (e.g. ≤ 0.85): the embedding's distance is informative about width within a situation; the calibration layer could use tightness (it already uses cohort tightness — this measures how much).
- PASS but R(5) near 1 (≥ 0.95): statistically real, practically nil — the situation carries the information.
- FAIL: the shape ranking is not where width lives; retrieval can be situation-only for width purposes.
No constant is tuned on this sample. Honest prior: studies 90/91 found width information in the analog set as a whole (IC 0.34–0.42 vs base rate); whether the WITHIN-situation shape ranking adds to that is genuinely open.
Machinery
scripts/research/study100_similarity_dispersion.py run [--smoke] → research/results/study100_*.{md,json,png}.
study_ledger row 100. Runs in the day at low priority (retrievals hit the clustered rank table, ~0.2 s each).