Skip to content
Research / Study 101Open access · Freely available
← Research ledger / Calibration

Closer analogs do not automatically deserve more weight

FailChart Library · Study 101

The finding

At k = 200 the member distances span too narrow a range for a kernel to separate anything: the median change in width is 0.2%. Inverse-rank weighting narrows the band 11% but loses 5.2 points of coverage, failing the guard. The unweighted band stays.

subjects574kernel width ratio0.998 [0.997, 1.000]kernel Winkler ratio0.999 [0.997, 1.000]inverse-rankwidth 0.887, coverage 0.777 vs 0.829

The research question

At the served k of 200, does weighting members by shape distance (Gaussian kernel, inverse rank, or a hard top-50 cut) give a narrower band without giving up coverage?

Decision criteria

Primary: Gaussian kernel scaled to the subject's own median distance, 5-day horizon, width ratio and Winkler ratio below 1 with the interval clear of 1, and no coverage loss beyond the guard. 574 fresh subjects, none from study 100.

Result document

Repository source: research/results/study101_distance_weighted_bands_2026_09_04.md

Reading (2026-09-04 12:50Z, written after the numbers; the bar was fixed before them). FAIL — the pre-stated third reading, with the second's mechanism visible in the secondaries. 574 fresh subjects (study 100's excluded).

Primary (Gaussian kernel, s = own median distance, vs unweighted, 5 d): width ratio 0.998 [0.997, 1.000], Winkler ratio 0.999 [0.997, 1.000] (touches 1), coverage 0.824 vs 0.829. At k = 200 the member distances span too narrow a range (≈ 0.3–0.45) for a kernel scaled to the subject's own median to separate anything: 54 % of subjects get a (slightly) narrower band, 46 % a wider one, and the median change is 0.2 %.

Secondaries: inverse-rank weighting narrows the band 11 % (0.887 [0.864, 0.902]) and scores better on Winkler (0.952 [0.935, 0.973]) — but loses 5.2 points of coverage (0.777 vs 0.829; loss CI90 [+0.030, +0.075]), failing the guard. The hard top-50 cut: width 0.983, coverage loss 2.6 points, Winkler 0.996 (CI spans 1) — nil.

The question that matters for the calibration layer — width at EQUAL coverage (each scheme rescaled by one scalar on its half-width until it covers exactly 0.80, i.e. what the conformal layer does; study101_equal_coverage_secondary_2026_09_04.json): unweighted 9.26 pp (multiplier 0.95 — the raw band over-covers on this draw, as the production conformal multipliers ≈ 0.95 say), Gaussian 0.999 [0.986, 1.020], inverse-rank 1.031 [0.974, 1.093] — WIDER, top-50 1.037 [0.992, 1.085]. Once coverage is restored, the narrower raw bands cost more width than they saved.

What this means with study 100: the nearer analogs are alike in their TAILS (study 100's narrowing was in std and the 80 % width of the members' outcomes), but they do not bracket the SUBJECT's own outcome any better than the whole served set does. Within a situation, the shape ranking narrows the members' dispersion without improving the subject's band at held coverage. The served unweighted band plus the conformal layer is already the efficient point for these weighting families at k = 200. No calibration-layer change is licensed. Direction was not measured. Runtime 2.5 min at low priority.

Run 2026-09-04T12:39:46.297173+00:00 (smoke=False). Subjects 574 kept ({'sampled': 600, 'excluded_overlap': 478, 'no_embedding': 0, 'retrieve_not_ok': 1, 'short_set': 25, 'no_own_outcome': 0, 'kept': 574}). Pre-registration: research/specs/study101_distance_weighted_bands_prereg_2026_09_04.md (ledger 101).

1d (n=574)

schemecoverage (CI90)median widthwidth ratio vs a (CI90)share narrowerWinkler ratio vs a (CI90)coverage loss vs a (CI90)
a_unweighted0.791 [0.763,0.817]4.091-1-
c_gauss0.794 [0.767,0.821]4.0940.999 [0.999,1.000]0.5260.999 [0.998,1.000]-0.004 [-0.009,+0.000]
b_invrank0.772 [0.744,0.800]3.7190.898 [0.881,0.916]0.6530.959 [0.931,0.989]+0.019 [-0.004,+0.042]
d_top500.801 [0.775,0.828]4.0941.001 [0.983,1.016]0.4980.999 [0.983,1.019]-0.011 [-0.030,+0.009]

5d (n=574)

schemecoverage (CI90)median widthwidth ratio vs a (CI90)share narrowerWinkler ratio vs a (CI90)coverage loss vs a (CI90)
a_unweighted0.829 [0.803,0.855]9.7431-1-
c_gauss0.824 [0.798,0.848]9.7320.998 [0.997,1.000]0.5440.999 [0.997,1.000]+0.005 [+0.000,+0.011]
b_invrank0.777 [0.747,0.805]8.8780.887 [0.864,0.902]0.6810.952 [0.935,0.973]+0.052 [+0.030,+0.075]
d_top500.803 [0.775,0.828]9.5930.983 [0.968,0.994]0.5470.996 [0.984,1.012]+0.026 [+0.007,+0.045]

10d (n=573)

schemecoverage (CI90)median widthwidth ratio vs a (CI90)share narrowerWinkler ratio vs a (CI90)coverage loss vs a (CI90)
a_unweighted0.834 [0.810,0.860]14.3891-1-
c_gauss0.833 [0.808,0.859]14.3870.999 [0.998,1.000]0.5430.999 [0.998,1.000]+0.002 [-0.005,+0.009]
b_invrank0.789 [0.761,0.817]12.8970.886 [0.871,0.905]0.6840.941 [0.918,0.954]+0.045 [+0.026,+0.065]
d_top500.820 [0.794,0.848]13.9910.991 [0.973,1.002]0.5251.003 [0.991,1.012]+0.014 [-0.002,+0.030]

PRIMARY (c_gauss vs a at 5d): {'coverage': 0.824, 'coverage_ci90': [0.7979, 0.8484], 'median_width': 9.732, 'median_winkler': 10.481, 'width_ratio_median': 0.9982, 'width_ratio_ci90': [0.9969, 0.9995], 'winkler_ratio_median': 0.999, 'winkler_ratio_ci90': [0.9974, 1.0003], 'coverage_loss_vs_a': 0.0052, 'coverage_loss_ci90': [0.0, 0.0105], 'share_narrower': 0.544}

Study specification

Repository source: research/specs/study101_distance_weighted_bands_prereg_2026_09_04.md

Registered 2026-09-04 ~12:40 UTC, before any number was pulled. Follow-up licensed by study 100 (PASS: within a situation, nearer analogs by shape have narrower forward dispersion, monotone in tightness, no floor within k 5..200; the narrowing lives in the tails). Question: can a band built from the SERVED analog set (k = 200) that weights nearer analogs more be narrower than the unweighted band without losing coverage on the subject's own realized outcome? Graham 2026-09-04: "register the distance-weighted bands study and run it."

Ships nothing. No direction is measured. A PASS licenses a calibration-layer change under its own gate; it does not change what is served.

Frozen design

  • Subjects: 600 random liquid (dollar_vol ≥ 5e7, close ≥ 5) situation_context rows with a happening identity and a 1d V5 embedding, dates 2022-01-01 … 2026-08-15, TABLESAMPLE draw, seed 101, excluding study 100's 478 subjects. Subjects whose retrieve returns < 200 members, or whose own realized outcome is missing, are skipped and counted.
  • Analog set: the production happening-then-shape retrieve (scale 1d, k = 200, as-of cutoff = date − 15 d, 2-per-symbol cap), members with their embedding distances d_i (ascending).
  • Outcomes: date-matched excess forward return (forward_returns_cacheliquid_date_medians) at 1 / 5 / 10 sessions for the members AND for the subject itself (the subject's outcome is realized after every member's, by the cutoff). Primary horizon 5 d.
  • Bands: nominal 80 %: [q10, q90] of the members' excess returns under each weighting scheme, using weighted quantiles with linear interpolation on the cumulative weight:
    • (a) unweighted — the baseline, what the served band is before the conformal layer.
    • (c) Gaussian kernel on distance — PRIMARY: w_i = exp(−(d_i / s)²), s = the subject's own median member distance. Self-normalising; no global constant.
    • (b) inverse rank: w_i = 1 / rank_i.
    • (d) hard cut: the nearest 50, unweighted.
  • Per subject, per scheme: width = q90 − q10; covered = subject's realized excess ∈ [q10, q90]; Winkler interval score at α = 0.2: S = width + (2/α) × distance outside the band (0 if covered). Lower is better; it is the proper scoring rule for an interval forecast and cannot be gamed by narrowing.
  • Aggregation over subjects: coverage rate (target 0.80) with a bootstrap CI90; median width; median per-subject width ratio scheme / (a); median per-subject Winkler ratio scheme / (a); paired bootstrap CI90 (2,000 resamples) over subjects for every ratio and for the coverage difference.

PRIMARY (one bar, pre-stated)

Scheme (c) versus (a) at 5 d: PASS iff the median per-subject Winkler ratio (c)/(a) has a bootstrap CI90 entirely below 1 and the paired coverage difference cov(a) − cov(c) has a CI90 upper bound ≤ 0.03 (the weighting may not buy its score by giving up coverage). FAIL otherwise.

Reported either way: coverage and median width per scheme; width ratios; Winkler ratios for (b) and (d); the same at 1 d and 10 d; the per-subject width-ratio distribution (how often the weighted band is narrower, and by how much); and whether coverage of (a) itself sits at its nominal 0.80 on this fresh draw (a check on the served band).

Pre-stated readings:

  • PASS: nearer-weighted bands are a real efficiency gain at held coverage; next is a registered gate on the calibration layer (the conformal multiplier is refit on the weighted band; nothing changes until that passes).
  • FAIL with narrower width but lost coverage: study 100's tail narrowing is real but the subject's own outcome is not better bracketed by the nearer analogs — the nearer analogs are alike in their tails, not in the subject's.
  • FAIL with no width gain: the weighting is too soft at k = 200 to matter; the 21 % at k = 5 is not recoverable by weighting alone.

No constant (s, α, k, the schemes) is tuned on this sample.

Machinery

scripts/research/study101_distance_weighted_bands.py run [--smoke] [--exclude study100.json]research/results/study101_*.{md,json,png}. study_ledger row 101. Runs in the day at low priority (~4 min).

Suggested citation: Chart Library (2026). Distance-weighting the cohort does not beat the unweighted band at held coverage. Study 101. chartlibrary.io/research/101-distance-weighted-bands.