Status: proposed research program, published September 9, 2026. No agent benchmark has been run and no improvement has been demonstrated.
The question is practical: can an agent give a better-supported answer with Chart Library? An answer can be useful because it finds the right precedent, notices a mismatch, or correctly says the evidence is insufficient. We will measure those behaviors before claiming an improvement in decisions.
Six questions worth testing
| Priority | Question | A concrete task | What we would measure |
|---|---|---|---|
| 1. Keep the evidence intact | Does a structured research record reduce misreadings? | Explain the IONQ comparison while retaining the original sample, clock, missing observations and untested outcomes. | Fully supported answers; unsupported numerical claims; correctly stated unknowns. |
| 2. Retrieve the right precedent | Can the agent find a relevant study and a counterexample? | A user asks whether tighter similarity implies better directional predictions. Retrieve the dispersion result and its limits. | Relevant source recall; irrelevant citations; number of calls and time. |
| 3. Respect the clock | Does the agent distinguish a scheduled event from information released later? | Separate IONQ's noon cutoff from its scheduled 3 p.m. appearance. | Correct availability judgments; use of later information; missing timestamps left unknown. |
| 4. Preserve contradictory evidence | Does a summary retain differences that change the comparison? | Explain why IONQ and PENN's late declines follow opposite earlier paths. | Essential differences retained; unsupported pooling; correct reference prices. |
| 5. Use less context without losing the limits | How much can we shorten a record safely? | Compare full articles with concise records at fixed context budgets. | Evidence retention against tokens, latency and cost; omitted caveats. |
| 6. Remember a revision accurately | Can an agent update an answer when a publication changes or is withdrawn? | Retrieve a version, receive a correction, and cite the current source without erasing what the earlier version said. | Stale citations; correct version selection; withdrawal respected. |
These are evaluations of research use. A better score is not evidence of trading returns. The studies about market outcomes keep their own protocols and samples.
Start with a reading experiment
Agent Evidence 01: proposed protocol specifies the first comparison. Existing IONQ examples are development material: they can show the workflow and help write a rubric, but cannot count as unseen confirmation.
First freeze the eligible source families, tasks, exact model versions, prompts, budgets and grading rubric. Then run both conditions on the same tasks. Publish all attempts and inconclusive results, including failures to retrieve or abstain. Keep related articles together when splitting and calculating uncertainty.
A publication rhythm we can sustain
- Day 1: an inspectable example showing one agent failure mode and the exact evidence needed to avoid it.
- Day 2: a counterexample from the same frozen case, clearly labeled as a follow-up rather than independent evidence.
- Day 3: a source or clock audit, including unresolved availability.
- Day 4: a protocol or progress note with status, missing work and no invented result.
- Day 5: a completed result only when the registered test has actually finished; otherwise publish the limitation or the next concrete question.
Every public research article is available to agents and people. Agent-focused work adds a task, expected evidence, error categories and reproducible inputs. The first three Casebook articles use existing frozen data; subsequent market retrievals and outcome studies need their own admission and registration.
Why these evaluation dimensions
LongMemEval separates information extraction, reasoning across sessions, knowledge updates, temporal reasoning and abstention. Those categories suggest useful tests for a dated market-research library, although its chat-history benchmark does not establish that Chart Library works.
FinBen evaluates different financial tasks, including extraction, question answering and forecasting. That separation supports measuring research-reading quality separately from predictive performance. Our proposed tasks and thresholds are design choices, not results from either paper.