Skip to content
Research / Study 103Open access · Freely available
← Research ledger / Calibration

Making uncertainty bands more honest on gap days

Pass at 5d, held out by timeChart Library · Study 103

The finding

One multiplier, ×1.48, fitted on the earlier half lifts held-out gap-day coverage from 0.70 to 0.83. The under-coverage lives in non-earnings gaps and in up-gaps; earnings gaps were already covered by the earnings conditioner, so the new one belongs on gap days that are not earnings sessions.

fit379 anchors, 2022-01 → 2024-06holdout385 anchors, 2024-07 → 2026-08raw 5d coverage0.701 [0.660, 0.738]with ×1.480.826 [0.795, 0.857]

The research question

The served 5-day band under-covers on gap days (study 102). Is that under-coverage stable across time, and does one multiplier fitted on the earlier half repair it on a later half it never saw?

Decision criteria

Fit on 2022-01 to 2024-06, freeze the multiplier, then score 2024-07 to 2026-08 once. Pass if the holdout's 90% interval contains 0.80 and the point sits inside [0.76, 0.84]. Written before any holdout number.

Result document

Repository source: research/results/study103_gap_day_conditioner_2026_09_04.md

Reading (2026-09-04 14:20Z, written after the numbers; the bar was fixed before them). PASS at 5 d. Fit half 379 anchors (2022-01..2024-06), holdout 385 (2024-07..2026-08), the fit frozen before any holdout number.

The under-coverage is stable across time, and one multiplier fixes it. Fitted m_5d = 1.48 (fit-half raw coverage 0.662). On the holdout the raw served band covers 0.701 [0.660, 0.738] — under-covering, as required — and with ×1.48 it covers 0.826 [0.795, 0.857]: the interval contains 0.80 and the point sits inside [0.76, 0.84]. The holdout on its own would have needed ×1.35, so the fitted value over-widens by ~9 % — inside the pre-stated window, and on the safe side of the stated 80 %.

Where the under-coverage lives: non-earnings gaps, and up-gaps. Earnings gaps (n = 56 on the holdout) are ALREADY covered by the raw band (0.821) and the single multiplier over-widens them to 0.911; non-earnings gaps go 0.681 → 0.812. Up-gaps 0.644 → 0.777, down-gaps 0.756 → 0.873. So the conditioner belongs on gap days that are not earnings sessions (earnings sessions already carry the earnings conditioner, ×1.10–1.29, which on these numbers is if anything generous), and a sign split is a reasonable refinement for the implementation gate — not tuned here.

1 d and 10 d are less stable than 5 d. Fitted ×1.615 / ×1.365 lift holdout coverage from 0.605 / 0.686 to 0.764 / 0.771, short of 0.80 (the holdout needed ×1.81 / ×1.505). Better than raw, not on target; the implementation gate should refit per horizon on the full sample and carry its own served-coverage receipt.

The honest cost: the Winkler interval score is WORSE with the conditioner (ratio 1.48): the raw band's misses are mostly small, so a proper scoring rule prefers the narrow band that misses 30 % of the time. The product promises a nominal 80 % band; that promise, not the score, is what the conditioner serves. Both numbers are on the record.

What this licenses: a flagged gap_day conditioner in the calibration layer (non-earnings gap sessions with |gap| ≥ 3 %, per-horizon multipliers refit on the full 764, default OFF, gated on its own served-coverage receipt over ≥ 30 sessions before default-on), and the desk v2 input (the agent gets the gap-conditioned width). No direction measured. Runtime 3 min.

Run 2026-09-04T14:12:59.073543+00:00 (smoke=False). Fit 379 / holdout 385 anchors ({'fit_sampled': 400, 'hold_sampled': 400, 'no_embedding': 0, 'retrieve_not_ok': 0, 'short_set': 35, 'no_own_outcome': 1, 'kept': 764}). Pre-registration: research/specs/study103_gap_day_conditioner_prereg_2026_09_04.md (ledger 103).

Fitted multipliers (fit half 2022-01..2024-06)

  • 1d: {'n': 379, 'm': 1.615, 'raw_coverage_fit': 0.628}
  • 5d: {'n': 379, 'm': 1.48, 'raw_coverage_fit': 0.6623}
  • 10d: {'n': 379, 'm': 1.365, 'raw_coverage_fit': 0.686}

Holdout (2024-07..2026-08) with the frozen multipliers

  • 1d: {'n': 385, 'm_applied': 1.615, 'raw_coverage': 0.6052, 'raw_ci90': [0.5636, 0.6468], 'conditioned_coverage': 0.7636, 'conditioned_ci90': [0.7273, 0.7974], 'm_holdout_would_need': 1.81, 'winkler_raw_median': 7.757, 'winkler_conditioned_median': 10.787, 'winkler_ratio_median': 1.615}
  • 5d: {'n': 385, 'm_applied': 1.48, 'raw_coverage': 0.7013, 'raw_ci90': [0.6597, 0.7377], 'conditioned_coverage': 0.826, 'conditioned_ci90': [0.7948, 0.8571], 'm_holdout_would_need': 1.35, 'winkler_raw_median': 15.181, 'winkler_conditioned_median': 21.59, 'winkler_ratio_median': 1.48, 'conditioned_coverage_earnings_gap': {'n': 56, 'raw': 0.8214, 'conditioned': 0.9107}, 'conditioned_coverage_non_earnings_gap': {'n': 329, 'raw': 0.6809, 'conditioned': 0.8116}, 'conditioned_coverage_up': {'n': 188, 'raw': 0.6436, 'conditioned': 0.7766}, 'conditioned_coverage_down': {'n': 197, 'raw': 0.7563, 'conditioned': 0.8731}}
  • 10d: {'n': 385, 'm_applied': 1.365, 'raw_coverage': 0.6857, 'raw_ci90': [0.6468, 0.7247], 'conditioned_coverage': 0.7714, 'conditioned_ci90': [0.7351, 0.8053], 'm_holdout_would_need': 1.505, 'winkler_raw_median': 22.088, 'winkler_conditioned_median': 28.65, 'winkler_ratio_median': 1.365}

Descriptive multipliers on the fit half (5d; none of these is the conditioner)

  • fit_m5_up: {'n': 181, 'm': 1.5, 'raw_coverage': 0.6409}
  • fit_m5_down: {'n': 198, 'm': 1.45, 'raw_coverage': 0.6818}
  • fit_m5_gap_3_5: {'n': 247, 'm': 1.46, 'raw_coverage': 0.6518}
  • fit_m5_gap_5_10: {'n': 95, 'm': 1.4, 'raw_coverage': 0.6737}
  • fit_m5_gap_gt_10: {'n': 37, 'm': 1.705, 'raw_coverage': 0.7027}
  • fit_m5_earnings_gap: {'n': 19, 'm': None, 'raw_coverage': 0.7895}
  • fit_m5_non_earnings_gap: {'n': 360, 'm': 1.51, 'raw_coverage': 0.6556}

PRIMARY: PASS

Study specification

Repository source: research/specs/study103_gap_day_conditioner_prereg_2026_09_04.md

Registered 2026-09-04 ~13:45 UTC, before any number was pulled. Licensed by study 102: on anchors whose session gapped ≥ 3 %, the served p10–p90 band of the production analog set covered the subject's own 5 d excess 0.694 against a nominal 0.80 (up-gaps 0.656), and the state's twins WIDENED were a better band than the event's own twins. Graham: "register study 103 and run it." Question: is the under-coverage stable enough across time that ONE multiplier per horizon, fitted on earlier gap days, restores 80 % coverage on later gap days it never saw?

Ships nothing by itself. A PASS licenses a gated calibration-layer change (a gap_day conditioner, implemented like the earnings conditioner, behind a flag, with its own served receipt). No direction is measured.

Frozen design

  • Anchors: liquid (dollar_vol ≥ 5e7, close ≥ 5) situation_context rows with a happening identity, a 1d V5 embedding, and |gap_pct| ≥ 3 on the anchor session. Two INDEPENDENT random draws, seed 103:
    • FIT half: 400 anchors dated 2022-01-01 … 2024-06-30.
    • HOLDOUT half: 400 anchors dated 2024-07-01 … 2026-08-15 (untouched until the multipliers are fixed).
  • Served band: the production happening-then-shape analog set at k = 300 (what cohort_analyze serves), as-of cutoff = date − 15 d; the raw p10 … p90 of the members' date-matched excess at 1 / 5 / 10 sessions. Subjects with < 300 members are skipped and counted.
  • Conditioner: a scalar m_h on the band's half-width about its midpoint: [mid − m·half, mid + m·half]. Fit: on the fit half, m_h = the smallest m on a 0.3 … 3.0 grid (step 0.005) at which coverage of the subject's own realized excess reaches 0.80. Fitted once per horizon; the 5 d value is the primary.
  • Test: apply the fitted m_5d to the holdout half. Coverage with a 2,000-resample bootstrap CI90 over holdout subjects; the raw (unconditioned) holdout coverage the same way.

PRIMARY (one bar, pre-stated), 5 d

PASS iff all three hold on the holdout half:

  1. the raw band under-covers there: raw coverage CI90 upper bound < 0.76 (the conditioner is needed);
  2. the conditioned coverage's CI90 contains 0.80;
  3. the conditioned coverage point estimate lies in [0.76, 0.84]. FAIL otherwise.

Reported either way: m_h at 1 / 5 / 10 d; the multiplier the holdout itself would have needed (drift); the Winkler interval score (α = 0.2) raw vs conditioned on the holdout; descriptive multipliers by gap sign, by |gap| bucket (3–5 %, 5–10 %, > 10 %) and by earnings vs non-earnings gaps (earnings_session.gap_date) — descriptive only, none of them is the conditioner; and the holdout coverage of the earnings-gap and non-earnings-gap subsets under the single multiplier.

Pre-stated readings:

  • PASS: the under-coverage is a stable property of gap days → build the gap_day conditioner (flagged, gated on its own served-coverage receipt over ≥ 30 sessions before default-on); the desk v2 input inherits the width.
  • FAIL by (1): the holdout no longer under-covers — study 102's finding was period-specific; nothing to fix.
  • FAIL by (2)/(3): the under-coverage is real but its size drifts — a fixed multiplier is the wrong shape; the online θ layer (which already adapts per slice) is where it belongs, as a gap-day slice. That is a different registration.

No constant (3 %, k = 300, the grid, the 0.76/0.84 window) is tuned on this sample. The holdout half is not touched until m is fixed from the fit half; both halves are drawn in the same run but the fit is computed first and frozen in the JSON before any holdout statistic is computed.

Machinery

scripts/research/study103_gap_day_conditioner.py run [--smoke]research/results/study103_*.{md,json,png}. study_ledger row 103. ~6 min at low priority.

Suggested citation: Chart Library (2026). A gap-day conditioner fitted on 2022–24 repairs 5-day coverage on 2024–26 gap days. Study 103. chartlibrary.io/research/103-gap-day-conditioner.