Skip to content
Research / For agentsOpen access · Freely available
Research for agents

Better evidence.
Better questions.

Help an agent find the right comparison, keep its limitations, and recognize what remains unknown. This is our research plan; the first experiment has not run.

Status: proposed research program, published September 9, 2026. No agent benchmark has been run and no improvement has been demonstrated.

The question is practical: can an agent give a better-supported answer with Chart Library? An answer can be useful because it finds the right precedent, notices a mismatch, or correctly says the evidence is insufficient. We will measure those behaviors before claiming an improvement in decisions.

Six questions worth testing

PriorityQuestionA concrete taskWhat we would measure
1. Keep the evidence intactDoes a structured research record reduce misreadings?Explain the IONQ comparison while retaining the original sample, clock, missing observations and untested outcomes.Fully supported answers; unsupported numerical claims; correctly stated unknowns.
2. Retrieve the right precedentCan the agent find a relevant study and a counterexample?A user asks whether tighter similarity implies better directional predictions. Retrieve the dispersion result and its limits.Relevant source recall; irrelevant citations; number of calls and time.
3. Respect the clockDoes the agent distinguish a scheduled event from information released later?Separate IONQ's noon cutoff from its scheduled 3 p.m. appearance.Correct availability judgments; use of later information; missing timestamps left unknown.
4. Preserve contradictory evidenceDoes a summary retain differences that change the comparison?Explain why IONQ and PENN's late declines follow opposite earlier paths.Essential differences retained; unsupported pooling; correct reference prices.
5. Use less context without losing the limitsHow much can we shorten a record safely?Compare full articles with concise records at fixed context budgets.Evidence retention against tokens, latency and cost; omitted caveats.
6. Remember a revision accuratelyCan an agent update an answer when a publication changes or is withdrawn?Retrieve a version, receive a correction, and cite the current source without erasing what the earlier version said.Stale citations; correct version selection; withdrawal respected.

These are evaluations of research use. A better score is not evidence of trading returns. The studies about market outcomes keep their own protocols and samples.

Start with a reading experiment

Agent Evidence 01: proposed protocol specifies the first comparison. Existing IONQ examples are development material: they can show the workflow and help write a rubric, but cannot count as unseen confirmation.

First freeze the eligible source families, tasks, exact model versions, prompts, budgets and grading rubric. Then run both conditions on the same tasks. Publish all attempts and inconclusive results, including failures to retrieve or abstain. Keep related articles together when splitting and calculating uncertainty.

A publication rhythm we can sustain

  • Day 1: an inspectable example showing one agent failure mode and the exact evidence needed to avoid it.
  • Day 2: a counterexample from the same frozen case, clearly labeled as a follow-up rather than independent evidence.
  • Day 3: a source or clock audit, including unresolved availability.
  • Day 4: a protocol or progress note with status, missing work and no invented result.
  • Day 5: a completed result only when the registered test has actually finished; otherwise publish the limitation or the next concrete question.

Every public research article is available to agents and people. Agent-focused work adds a task, expected evidence, error categories and reproducible inputs. The first three Casebook articles use existing frozen data; subsequent market retrievals and outcome studies need their own admission and registration.

Why these evaluation dimensions

LongMemEval separates information extraction, reasoning across sessions, knowledge updates, temporal reasoning and abstention. Those categories suggest useful tests for a dated market-research library, although its chat-history benchmark does not establish that Chart Library works.

FinBen evaluates different financial tasks, including extraction, question answering and forecasting. That separation supports measuring research-reading quality separately from predictive performance. Our proposed tasks and thresholds are design choices, not results from either paper.

Agent Evidence 01: proposed protocol

Status: proposed protocol. Task manifest and model configuration are not frozen; no evaluation has run. Publishing this design is not a completed preregistration.

Question and scope

Does presenting the same published evidence as a structured record plus its full source documents improve an agent's ability to give a correctly scoped, cited answer compared with plain source documents? This first experiment tests reading, not search quality or investment performance.

Development material

Use CB-001 and its IONQ/PENN and HYFM/KOD follow-ups to draft instructions and a grading rubric. All three belong to one source family and are excluded from confirmation. The IONQ end-of-day state and noon Casebook are different observations: never merge their clocks or analog samples.

Example task: "Do HYFM and KOD show the same ten-minute decline as IONQ, and does the sample establish that the setup is profitable?" Required facts: HYFM and KOD's ten stored bars cover 13 and 22 minutes; final clock-window returns are missing rather than zero; gaps do not establish halts; the original 50-member set is retained; outcomes were not tested. The answer must cite the original case or its directly sourced follow-up. A claim about later profitability fails.

Freeze before running

  1. Select 12 eligible, distinct public source families excluding the development family. Related horizon extensions, follow-ups and reused samples form one family. Do not manufacture independence by counting their articles separately.
  2. Write five tasks per family: extraction, comparison, time or scope, a numerical interpretation, and an unsupported-premise question. Total 60 paired tasks. A task must be answerable or explicitly unanswerable from its frozen sources. Record the required facts, permissible uncertainty and disqualifying claims before model calls.
  3. Record source IDs, full document hashes, publication versions, source-family mapping, task hashes, exact model identifiers, system prompts, temperatures, output budgets, randomization seed, provider settings, retry rules and software commit in a signed-off manifest. Record the freeze timestamp. If 12 qualified families are unavailable, report a feasibility pilot; do not quietly shrink the confirmatory design.
  4. Use two explicitly named model configurations fixed in that manifest. Keep each model's comparisons paired. No model or prompt selection after looking at evaluation outputs.

Conditions

A: the full frozen source documents and citation URLs in plain text.

B: the same full source documents and URLs, preceded by the corresponding structured overview: status, dates, sample receipt, limitations, document identities and hashes. Audit that B adds no substantive fact absent from A. Include those facts in A's source text if needed before freezing. This estimates the effect of that presentation package, not compression alone.

Both receive the same task, instruction, output budget and available evidence. Randomize condition order; use a fresh conversation for every attempt. No market lookup, outcome join, browsing or extra tool access. One attempt per condition and task; permit at most one retry only for a documented transport failure before a response begins. Report refusals, truncations and failures. Do not replace difficult tasks.

Primary score and grading

A task passes only if every required material fact is correct, relevant factual claims cite a source that supports them, units and clocks are preserved, and no disqualifying claim appears. An unknown answer can pass when the rubric requires abstention. A citation to a document containing the topic is insufficient if it does not support the claim.

Two reviewers grade shuffled outputs without model or condition labels using the frozen rubric. Disagreements are adjudicated and both initial judgments are retained. Publish reviewer agreement. Automated checks may verify exact numbers, IDs and URLs, but do not substitute keyword presence for meaning.

Primary: per-model paired difference in task pass rate, B minus A. Use 10,000 paired bootstrap resamples of source families with a fixed seed; keep every task and both conditions from a family together. Report a 95% interval for each model and disclose that two models are evaluated. A supported improvement requires both models' intervals above zero and both point differences at least 10 percentage points. Otherwise report mixed, no demonstrated improvement, or inconclusive according to the results; do not pool models to rescue a failed criterion.

Secondary: unsupported numerical claims, incorrect availability judgments, missing-value errors, misapplied calibration claims, abstention accuracy, input/output tokens, latency, cost and transport failures. These cannot rescue the primary. Any privacy leak or fabricated market outcome is a critical failure to investigate, regardless of average score.

Reporting and stopping

Publish the frozen manifest, tasks, rubric, raw outputs where licensing permits, scores, uncertainty, failures and costs. Keep provider credentials and participant information out of the public record. Stop and invalidate the affected run if source hashes drift, evidence differs across arms, the holdout leaks into prompt tuning, or graders see labels before scoring. A revised design needs a new unexamined evaluation set.

This test opens no new market outcomes and changes no historical analog membership. A pass would support a claim about reading this evidence under these conditions, not predictive accuracy, profitability, autonomous learning or all agents.