Skip to content
Research / Study 102Open access · Freely available
← Research ledger / Calibration

Same event does not always mean a better comparison

FailChart Library · Study 102

The finding

The event-banded set is wider, not narrower, and no better at equal coverage. The finding underneath mattered more: the served analog set for an event day is the state's twins, not the event's, and the served band under-covers event days at 5 days. That under-coverage is what study 103 repaired.

subjects399 event-day, 304 with ≥ 30 same-event memberssame-event share of the served setmedian 4.5%, mean 8.2%served 5d coverage on event days0.694 vs 0.80 nominal

The research question

For a subject on an event day (a gap of 3% or more), is a band built only from analogs that shared the event narrower at equal coverage than the served band, which ranks on the state's shape?

Decision criteria

Primary: at equal coverage the event-banded set must be narrower than the served set, with the interval clear of 1. 399 event-day subjects, fixed before the numbers.

Result document

Repository source: research/results/study102_event_banded_retrieval_2026_09_04.md

Reading (2026-09-04 13:30Z, written after the numbers; the bar was fixed before them). FAIL on the primary — and the finding underneath is the one that matters. 399 event-day subjects (|gap| ≥ 3 %); 304 have ≥ 30 same-event members within the shape-ranked top 2,000 (24 % are too thin).

1. The day-1 read generalises. The served set's same-situation share on random event-day anchors: median 4.5 %, mean 8.2 %. HPE's 1-in-50 was typical. The production analog set for an event day is the STATE's twins, not the event's.

2. The served band UNDER-COVERS event-day subjects. Raw p10–p90 of the production set covers the subject's own 5 d excess 0.694 [0.651, 0.737] against a nominal 0.80 (1 d 0.694, 10 d 0.720; up-gaps 0.656, down-gaps 0.735). Event days are followed by wider outcomes than the state's twins suggest (the event-banded set's IQR is 1.98× the base IQR vs 1.47× for the served set).

3. The event-banded set is more honest about width but is not the better band. It is 35 % wider (19.2 vs 14.0 pp) and covers 0.766 — closer to nominal — but the Winkler ratio is 1.19 [1.14, 1.26] (worse: the extra width costs more than the recovered misses save), and at EQUAL 0.80 coverage the production set rescaled ×1.385 (19.4 pp) is narrower than the banded set rescaled ×1.145 (22.0 pp): ratio 1.137 at 5 d, 1.077 at 10 d, 0.991 at 1 d. The state's twins, once widened, bracket event-day subjects with less width than the event's own twins do.

What this licenses and what it does not. Not the banded set as the v2 packet's analog set (that was the question; the answer is no on width grounds). What it points at instead is a calibration multiplier for gap days: the served band needs ×1.385 at 5 d (×1.54 at 1 d, ×1.27 at 10 d) to hold 80 % on |gap| ≥ 3 % anchors. The existing earnings / catalyst conditioners widen by ~1.10–1.29 in the same direction and fall short of that. Candidate follow-up (own registration, held-out check): a gap-day conditioner in the calibration layer. For the desk: the state-set width the agent sizes stops from is too tight on event candidates by roughly that factor — a v2 INPUT item (hand the agent the gap-conditioned width), registered after v1's read. No direction measured. Runtime 10 min at low priority.

Run 2026-09-04T13:20:03.449547+00:00 (smoke=False). Subjects 399 kept ({'sampled': 400, 'no_embedding': 0, 'retrieve_not_ok': 0, 'short_a': 1, 'thin_banded': 95, 'no_own_outcome': 0, 'kept': 399}). Base 5d IQR 4.215 pp. Same-situation share of the production set: median 0.045, mean 0.0819. Banded set size median 67, thin (<30) share 0.238. Pre-registration: research/specs/study102_event_banded_retrieval_prereg_2026_09_04.md (ledger 102).

1d (n=304)

  • coverage a 0.694 [0.648, 0.7368] vs b 0.776 [0.7368, 0.8158]; coverage loss -0.082 [-0.1086, -0.0559]
  • median width a 6.324 vs b 9.019; width ratio 1.4261 [1.3501, 1.4724]; share narrower 0.079
  • Winkler ratio b/a 1.2403 [1.1948, 1.2877]
  • equal-coverage: {'mult_a': 1.54, 'width_a': 9.738, 'mult_b': 1.07, 'width_b': 9.65, 'ratio_b_over_a': 0.991}
  • up_gaps: {'n': 157, 'width_ratio': 1.5231, 'winkler_ratio': 1.3052, 'coverage_a': 0.675, 'coverage_b': 0.771}
  • down_gaps: {'n': 147, 'width_ratio': 1.3037, 'winkler_ratio': 1.2193, 'coverage_a': 0.714, 'coverage_b': 0.782}

5d (n=304)

  • coverage a 0.694 [0.6513, 0.7368] vs b 0.766 [0.7268, 0.8059]; coverage loss -0.072 [-0.0987, -0.0461]
  • median width a 14.001 vs b 19.247; width ratio 1.3541 [1.3078, 1.4073]; share narrower 0.128
  • Winkler ratio b/a 1.1946 [1.1417, 1.2594]
  • equal-coverage: {'mult_a': 1.385, 'width_a': 19.391, 'mult_b': 1.145, 'width_b': 22.038, 'ratio_b_over_a': 1.1365}
  • informative ratio (set IQR / base IQR): a 1.468, b 1.976
  • up_gaps: {'n': 157, 'width_ratio': 1.3715, 'winkler_ratio': 1.198, 'coverage_a': 0.656, 'coverage_b': 0.713}
  • down_gaps: {'n': 147, 'width_ratio': 1.3444, 'winkler_ratio': 1.1941, 'coverage_a': 0.735, 'coverage_b': 0.823}

10d (n=304)

  • coverage a 0.720 [0.6776, 0.7632] vs b 0.796 [0.7597, 0.8322]; coverage loss -0.076 [-0.1086, -0.0461]
  • median width a 19.353 vs b 25.978; width ratio 1.2938 [1.2686, 1.3279]; share narrower 0.158
  • Winkler ratio b/a 1.1831 [1.1427, 1.227]
  • equal-coverage: {'mult_a': 1.265, 'width_a': 24.482, 'mult_b': 1.015, 'width_b': 26.367, 'ratio_b_over_a': 1.077}
  • up_gaps: {'n': 157, 'width_ratio': 1.314, 'winkler_ratio': 1.1708, 'coverage_a': 0.694, 'coverage_b': 0.771}
  • down_gaps: {'n': 147, 'width_ratio': 1.2847, 'winkler_ratio': 1.1842, 'coverage_a': 0.748, 'coverage_b': 0.823}

Study specification

Repository source: research/specs/study102_event_banded_retrieval_prereg_2026_09_04.md

Registered 2026-09-04 ~13:05 UTC, before any number was pulled. Motivation: the desk's day-1 playbook read (2026-09-03, journaled as desk_review/playbook_read): for event-driven candidates the production analog set is the STATE's twins, not the EVENT's — the happening identity carries neither the gap's sign nor the move's magnitude (HPE's reclaimed gap-down had 1 same-situation analog in 50; FIVE's sold gap-up had 3). The proposed v2 packet input is situation-first bands on the event before shape. This study measures whether that band buys anything for the one thing a band is for: bracketing the subject's own outcome at held coverage.

Ships nothing; no direction is measured. A PASS licenses (i) the Desk v2 packet input change, registered after v1's read, and (ii) a gated event_bands retrieval option for the product. A FAIL keeps both off.

Frozen design

  • Subjects: 400 random liquid (dollar_vol ≥ 5e7, close ≥ 5) situation_context rows whose anchor day is an event: |gap_pct| ≥ 3, with a happening identity and a 1d V5 embedding, dates 2022-01-01 … 2026-08-15, TABLESAMPLE draw, seed 102.
  • Set (a), production: retrieve_happening_shape (scale 1d, k = 200, as-of cutoff = date − 15 d, 2-per-symbol cap) — identity-first, shape-ranked.
  • Set (b), event-banded: from ONE wide retrieve (k = 2,000, the API's own cohort_size ceiling, same identity-first shape ranking), the nearest 200 members whose own anchor day had the SAME gap sign and |gap_pct| ≥ 3 (looked up in one batch). Set (a) is the first 200 of the same wide retrieve, i.e. exactly the production set. Subjects with fewer than 30 banded members in the wide set are skipped from the primary and counted (the bucket is too thin for a band). Amended 13:15Z before any number: a correlated EXISTS inside the rank query was the first form and cost tens of seconds per subject on big buckets; the wide-retrieve form asks the same question within the shape-ranked top 2,000.
  • Outcomes: date-matched excess forward return at 1 / 5 / 10 sessions for members and for the subject. Primary horizon 5 d.
  • Bands: unweighted p10 … p90 of the members' excess (nominal 80 %), for (a) and (b). Per subject: width, covered (subject's realized excess inside), Winkler interval score at α = 0.2.
  • Aggregation: coverage with bootstrap CI90; median width; per-subject width ratio (b)/(a) and Winkler ratio (b)/(a) with paired bootstrap CI90 (2,000 resamples); coverage difference (a) − (b) with CI90; the equal-coverage width ratio (each set rescaled by one scalar on its half-width to exactly 0.80 coverage).

PRIMARY (one bar, pre-stated)

Set (b) versus (a) at 5 d on the subjects where both sets exist (≥ 30 banded members): PASS iff the median per-subject Winkler ratio (b)/(a) has a bootstrap CI90 entirely below 1 and the paired coverage difference cov(a) − cov(b) has a CI90 upper bound ≤ 0.03. FAIL otherwise.

Reported either way: the same-situation share of set (a) (members with the subject's gap sign and |gap| ≥ 3 — the day-1 read's number, now on 400 subjects); the banded set's typical size and the share of subjects with a thin bucket; the informative ratio (IQR of the set's 5 d excess over the base IQR) for (a) and (b); the equal-coverage width ratio; 1 d and 10 d; the same split by gap sign (up-gaps vs down-gaps).

Pre-stated readings:

  • PASS: the event's own history brackets event-day outcomes better than the state's twins → register the v2 packet input (bands before shape) and build the gated event_bands option; the desk's side still comes from the tape.
  • FAIL, bands not narrower: the state already carries what matters for width on event days; v2's input change is not licensed on width grounds (it may still be worth it for the AGENT's reading — that is a v2 forward question, not a width one).
  • FAIL, bands narrower but coverage lost: same reading as study 101 — the event's twins are alike among themselves, not around the subject.

No constant (3 %, k, the 30-member floor, α) is tuned on this sample.

Machinery

scripts/research/study102_event_banded_retrieval.py run [--smoke]research/results/study102_*.{md,json,png}. study_ledger row 102. Runs in the day at low priority (two retrieves per subject; the banded one carries a correlated EXISTS, ~1–3 s).

Suggested citation: Chart Library (2026). Restricting analogs to same-event members does not improve the band on event days. Study 102. chartlibrary.io/research/102-event-banded-retrieval.