Separate intelligence & search

ONE RUN · THREE CONTRIBUTIONS

Inside one Exa Ultra research run

The assignment, the findings, and the presentation—separated so you can judge them.

$11.66 reported Exa cost92 reported searches54 + 30 returned records
Provider-generated findings; no independent source audit. One completed run, no comparison run. Search counts and output length do not measure quality.

PAGE NARRATIVE: ASSISTANT-AUTHORED · EXA SECTIONS LABELLED BELOW

Explore the guide +

A worked example of delegated research—and the judgments around it.

An assistant designed the assignment, Exa Ultra conducted the research, and the assistant organized the result. This page lets you inspect each contribution separately. You can read the explanation in a few minutes or explore the complete output below.

What happened

On September 26, 2026, one Exa Agent Ultra run investigated public benchmarks and evaluation frameworks for research agents. In plain language: how do people test whether an AI can find difficult answers, build complete lists, or produce a well-supported research report?

The run was restricted to public web research, with no Connect datasets requested and no follow-up run. It used Exa's documented default $20 run cap. The connected interface did not expose a custom budget field; the spending instruction in the prompt was not itself a technical limit.

This is an unaudited research output, not an authoritative benchmark directory or a comparison of research systems. No benchmark code was executed. The catalogue's underlying claims were not independently checked by the presenting assistant.

The service returned a completed status, reported a cost of $11.6589 and 92 searches, and supplied 54 records it classified as qualifying plus 30 adjacent, uncertain, or excluded records. These are reported activity and output counts—not measures of correctness. The returned response did not expose a stop reason or timestamps sufficient to establish run duration. It also does not contain a complete log of all 92 search queries.

Who contributed what

  1. Assistant: design the assignment. The assistant translated the research objective into a question, additional instructions, and a structured output schema. It selected a January 2025–September 2026 eligibility window and emphasized public artifacts, reproducibility, evidence, and results provenance.
  2. Exa: investigate and generate findings. Ultra retrieved sources, selected and classified candidates, assessed artifacts, and wrote the summary and records. Those decisions and claims belong to the provider-generated output.
  3. Assistant: package and assess. The assistant created this page, selected the three teaching examples below, and wrote the interpretation and limitations outside the labelled Exa sections. That editorial layer influences what readers notice.

The assistant that helped frame the assignment also prepared this presentation. Its assessment is not an independent review. The original Exa response was already shaped by the assignment, so inspecting the prompt matters as much as inspecting the answer. Neither the raw response nor this presentation is free of judgment.

Three examples of different research abilities

The examples below are selected and paraphrased by the assistant from Exa's records to illustrate different abilities. They are not a ranking, independently verified findings, or a representative sample of the whole catalogue.

  • Finding a difficult answer — BrowseComp. Exa describes questions requiring persistent browsing, with short answers scored for correctness. That tests a different ability from compiling a complete list. Open Exa's full record.
  • Finding and enriching a set — WideSearch. Exa describes table-based tasks with entity, attribute, row, and table-level evaluation. Missing entities and incomplete fields matter here. Open Exa's full record.
  • Producing a supported report — DeepResearch Bench (RACE + FACT). Exa describes separate assessment of report quality and citation support. A readable, comprehensive report is not automatically well evidenced. Open Exa's full record.

These examples suggest a useful question when reading a research-agent score: what ability did the evaluation actually test? The catalogue supplies leads for answering that question; its citations still need inspection before relying on a specific claim.

Assistant assessment

The useful deliverable is a starting reference: named evaluations, descriptions of what they measure, links to artifacts, and claimed barriers to reproduction. Its structure makes follow-up investigation easier than a bare list of search results.

The assignment also creates blind spots. It emphasizes public availability and reproducibility rather than giving equal weight to beginner accessibility, evaluation cost, multilingual representation, or agreement with real users' judgments. The chosen date window excludes older work unless substantively updated. These are framing choices, not neutral properties of the field.

Exa's scope statement includes controlled web-derived corpora and some literature or report-evaluation tasks. Readers may reasonably draw the boundary of “research-agent evaluation” differently. Preserve that scope statement when interpreting the count of 54; the count is Exa's classification, not a settled total for the field.

The result does not establish that every entry is accurate, that coverage is exhaustive, or that Ultra found more useful material than a cheaper Agent setting or assistant-directed search would have found. No comparison run was performed. More searches, more records, and polished explanations are not substitutes for evidence quality.

The review for this publication checked record counts, faithful rendering, links within this site, and removal of the operational run identifier. It did not audit the external citations or reproduce benchmark results. Exa's own descriptions of what it inspected remain provider claims.

Read Exa's output

The following summary, scope statement, limitations, and catalogue fields come from the saved Exa response. Field wording is preserved; labels, layout, search, and navigation were added by the assistant. Links are evidence supplied by Exa, not an endorsement or a claim that the presenting assistant checked them.

EXA OUTPUT · WORDING PRESERVED

Exa's summary

The catalogue contains 54 qualifying benchmarks/frameworks and 30 adjacent, uncertain or excluded entries. Three distinct capabilities recur: finding a difficult answer, discovering and enriching an entire set, and producing a supported research report. Their scores are not interchangeable: WideSearch (new tab) evaluates tables and set coverage, while DeepResearch Bench (new tab) separates report quality from citation support. Public availability ranges from released tasks and evaluators to encrypted, partially public or paper-only protocols. BrowseComp-Plus (new tab) controls web-state variation with a fixed corpus; live-web benchmarks still depend on retrieval, provider and judge versions. Most inspected results are benchmark-author runs, often also vendor runs. Reka’s third-party extension is identified separately without assuming organizational independence. Renames, forks and run reports are not counted as new benchmarks. No benchmark code was executed and no reported result was reproduced.

Exa's scope interpretation

Public benchmarks and evaluation frameworks with a documented release or substantive update from January 1, 2025 through September 26, 2026. Includes difficult web discovery, comprehensive entity/list retrieval, evidence-supported enrichment, literature discovery and research-report evaluation. Controlled web-derived corpora are included when they test research-agent workflows. A public paper or protocol can qualify even when tasks or scoring code are withheld; availability is assessed separately. Browser-action tasks, generic QA, preference-only search chat, internal-enterprise tasks and vendor run reports are separated. This is broad documented coverage, not a claim of exhaustiveness.

Exa's stated coverage limitations
  • Coverage is broad but not proven exhaustive. Newly released, poorly indexed, non-English or privately distributed projects may be missing; the catalogue does not claim every qualifying project exists here.
  • Artifact inspection establishes what was publicly documented or accessible, not that an end-to-end evaluation succeeds. “Not verified” means unresolved in the inspected sources, not proof that an artifact does not exist.
  • Release dates and benchmark versions matter. Paper revisions, corrected answers, expanded datasets and changing repositories can yield different task counts; later vendor runs alone do not establish a new benchmark release.
  • Complete-list scoring has different denominators: fixed expert gold, pooled discovered matches or a requested output quota. Recall and F1 therefore do not by themselves demonstrate exhaustive real-world coverage.
  • Encrypted answers, locked test sets, missing corpora, unavailable proprietary components, paid API dependencies and model/judge drift can prevent exact independent reproduction even when some code is public.
  • Public human preference or citation-quality evaluation is not automatically a test of autonomous discovery. Search Arena and SciArena are separated from agentic discovery benchmarks; broad science suites are included only for their relevant research components.
  • Independence was not inferred from a third-party name, leaderboard entry or open repository. Commercial comparisons remain attributed to their authors; no cross-benchmark vendor ranking is warranted.

Explore Exa's 54 catalogue records

These are the entries Exa classified as qualifying. Search covers all record fields; it does not score or validate them.

54 of 54 catalogue records

1. BrowseCompdifficult web discovery / short-answer research

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
difficult web discovery / short-answer research
Primary link
https://openai.com/index/browsecomp/ (new tab)
Release or update evidence
2025-04-10: OpenAI publicly announced and open-sourced the benchmark.
What it measures
Tests persistent browsing and reasoning for obscure, entangled facts across 1,266 questions with short, verifiable answers; it is not an exhaustive-list or report-quality test.
Task data
Public test CSV is referenced by the evaluator; questions and answers are obfuscated using canary-based decoding.
Evaluation code
Public scripts in openai/simple-evals load and decode tasks, then invoke an LLM grader.
Judging criteria
An LLM judges reference-answer equivalence, producing aggregate correctness or accuracy rather than source-quality or list-completeness scores.
Reproducibility and barriers
Requires an agent/search implementation and grader access; live-web changes, model settings, and browsing effort affect results. Public answers create leakage risk, and authors request that decoded examples not be republished.
Who reported the results
Launch results are benchmark-author and vendor-reported by OpenAI, not independently reproduced.

Evidence supplied by Exa

2. WideSearch: Benchmarking Agentic Broad Info-Seekingcomprehensive list enumeration + evidence-supported entity enrichment

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
comprehensive list enumeration + evidence-supported entity enrichment
Primary link
https://arxiv.org/abs/2508.07999 (new tab)
Release or update evidence
Released August 11, 2025 (arXiv 2508.07999; paper version dated August 28/September 5 in rendered materials); the GitHub README announces release on 2025/08/11.
What it measures
Measures item, column, and row precision/recall; table/task success requires complete, accurate atomic information. Max@N and human comparison are also reported.
Task data
Public 200-task dataset: 100 English and 100 Chinese tasks across 18 domains. Agents fill predefined tables using large-scale atomic facts about multiple entities; curation included exhaustive human gold research and automated-versus-human validation.
Evaluation code
Public MIT repository contains evaluation scripts, documentation, and agent/tool code; the public Hugging Face dataset and project page link to data and code.
Judging criteria
Entity/item precision-recall/F1, attribute correctness and exact complete-row/table correctness; the automated evaluator was compared with expert ratings.
Reproducibility and barriers
Code and dataset are public, but live-web execution requires API credentials and results may drift with web, model, and API changes. Commercial-system tests used web interfaces; this is an assessment of barriers.
Who reported the results
Original authors report benchmark runs across 10+ agentic systems and human tests; results are author-reported, not independently validated.

Evidence supplied by Exa

3. Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agentscomprehensive list enumeration + evidence-supported entity enrichment

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
comprehensive list enumeration + evidence-supported entity enrichment
Primary link
https://arxiv.org/abs/2606.27595 (new tab)
Release or update evidence
2026-06: arXiv release 2606.27595; official project, repository and dataset are public.
What it measures
Measures membership Item-F1, attribute-cell Column-F1, complete-row Row-F1, table success, and parse rate using normalization-aware comparison.
Task data
Public schema/metadata with gated/encrypted answer fields: 228 Korean tables cover 190 parent entities and 16 categories, including 4,262 gold rows and 14,560 attribute cells. Easy/Medium/Hard tiers vary table width and composite-key membership; 201 tables require cross-source attribute lookups.
Evaluation code
Public MIT-licensed pipeline and scorer are available in the official repository; answer fields remain encrypted or gated by request to reduce leakage.
Judging criteria
Scores membership precision/recall, per-column cell correctness, strict full-row correctness, and whole-table success; the comparator normalizes name variants, date granularity, and numeric formatting.
Reproducibility and barriers
Scorer and pipeline are public, but evaluation answers are gated/encrypted and live-web tasks are dynamic; API/model access and request-based data release remain practical barriers (assessment).
Who reported the results
Original authors report a 20-agent evaluation; these are benchmark-author runs, not independently validated results.

Evidence supplied by Exa

4. EnterList (introduced in WebLists)Structured list extraction from interactive websites; paper-defined benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Structured list extraction from interactive websites; paper-defined benchmark
Primary link
https://arxiv.org/html/2504.12682 (new tab)
Release or update evidence
ArXiv record dated 2025-04-17, within the eligibility window.
What it measures
Precision and recall for rows extracted from live interactive websites, with cost per output row; narrower than unconstrained open-web enumeration.
Task data
Availability not established as a public release: the paper defines 200 live tasks across four use cases and 50 websites, with annotated reference URLs and extraction scripts, capped at five pages per task.
Evaluation code
The paper describes reference extraction scripts and methodology, but a public repository or downloadable evaluator/data package was not established; treat it as paper-only.
Judging criteria
Exact field matching against refreshed reference extraction defines precision as retrieved gold rows and recall as retrieved gold coverage; live-site ground truth changes over time.
Reproducibility and barriers
Live websites and changing ground truth hinder reproduction; no verified public task, data, or evaluator release was established.
Who reported the results
Original authors report experiments including BardeenAgent results; no independent reproduction was established.

Evidence supplied by Exa

5. DeepResearch Bench (Ayanami0730 / Mingxuan Du; RACE + FACT)long-form deep-research report benchmark; citation/retrieval evaluation

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
long-form deep-research report benchmark; citation/retrieval evaluation
Primary link
https://github.com/Ayanami0730/deep_research_bench (new tab)
Release or update evidence
2025-06: paper arXiv:2506.11763 and public repository (created June 13) introduce RACE/FACT evaluation. Distinct from FutureSearch’s similarly named benchmark.
What it measures
RACE measures report quality through comprehensiveness, depth, instruction following, and readability; FACT measures citation accuracy and effective citation count.
Task data
Public repository materials include 100 PhD-level tasks, 50 Chinese and 50 English, across 22 fields, authored or refined by 100+ domain experts. It includes task, reference, report, and result materials, with raw articles and scores linked on the leaderboard.
Evaluation code
Public repository contains RACE and FACT evaluation code, prompts, configurations, and result directories; execution requires model/API and web-content retrieval services.
Judging criteria
RACE uses LLM-generated task criteria and weights to compare reports with high-quality references, with human-consistency validation. FACT extracts statement-URL pairs, retrieves cited pages, and LLM-judges support.
Reproducibility and barriers
Code and data are publicly downloadable under Apache-2.0, but reproducing scores requires paid or credentialed judge models and web retrieval; live URLs may change. These are assessed barriers.
Who reported the results
Authors report evaluations of commercial deep-research agents and search-enabled LLMs; linked raw articles and scores remain author-run, not independent results.

Evidence supplied by Exa

6. DeepResearch Bench II (imlrz / Ruizhe Li et al.)long-form deep-research report benchmark; expert-rubric diagnosis

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
long-form deep-research report benchmark; expert-rubric diagnosis
Primary link
https://github.com/imlrz/DeepResearch-Bench-II (new tab)
Release or update evidence
Official README records a November 2025 pipeline release; arXiv:2601.08536 was submitted 2026-01-13. It is explicitly a follow-up to DeepResearch Bench.
What it measures
Scores binary satisfaction of information-recall, analysis and presentation criteria across 9,430 fine-grained rubrics and 132 tasks.
Task data
Public tasks_and_rubrics.jsonl contains 132 tasks and 9,430 criteria across 22 domains. A February 24, 2026 Hugging Face release adds selected model-generated reports. May 2026 metadata assigns each task its source article’s license.
Evaluation code
Official Apache-2.0 repository provides rubrics, scripts and a batched evaluator for PDF, DOCX, image and text reports. Running requires Gemini or other model credentials.
Judging criteria
LLM judges assess binary satisfaction of atomic information-recall, analysis and presentation criteria. Criteria were extracted from expert reports and human-reviewed.
Reproducibility and barriers
Benchmark, rubrics and code are public; assessment: judge-model access, processing cost, evaluator versions and report formatting affect reproducibility.
Who reported the results
The paper reports author evaluations of several state-of-the-art agents; no independent result provenance is established here.

Evidence supplied by Exa

7. DeepResearch-ReportEval (HKUDS/DeepResearch-Eval)long-form research-report evaluation framework/dataset

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
long-form research-report evaluation framework/dataset
Primary link
https://github.com/HKUDS/DeepResearch-Eval (new tab)
Release or update evidence
Paper arXiv:2510.07861 was published 2025-10-09, and the official repository identifies the framework and dataset; it qualifies within the window.
What it measures
Scores comprehensiveness, coherence, clarity, insightfulness and overall quality, plus paragraph redundancy and citation-supported factuality.
Task data
100 queries cover 12 real-world categories and include 100 Qwen-DeepResearch reports collected in early September 2025; repository data and examples are public.
Evaluation code
Repository includes judge_score.py, judge_fact.py, utilities, prompts and examples. Fact checking uses Firecrawl or Jina Reader with -1/0/1 support labels.
Judging criteria
LLM judges score report quality and paragraph redundancy; a citation checker evaluates each claim as unsupported, uncertain or supported using retrieved source pages.
Reproducibility and barriers
MIT code and data are public; assessment: judge-model and Firecrawl/Jina credentials, web access and model selection are practical barriers. The dataset is a fixed snapshot.
Who reported the results
Paper comparisons cover four commercial systems and author-generated Qwen reports; no independent reproduction is identified here.

Evidence supplied by Exa

8. ResearcherBenchScientific web research and evidence-grounded report generation

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Scientific web research and evidence-grounded report generation
Primary link
https://github.com/GAIR-NLP/ResearcherBench (new tab)
Release or update evidence
Paper arXiv:2507.16280 was published 2025-07-22; the official site and repository are public within the window. It evaluates frontier-AI scientific questions rather than broad PhD-level web research.
What it measures
Scores weighted expert-insight coverage, citation-support accuracy (Faithfulness) and citation coverage (Groundedness).
Task data
Public repository provides 65 expert-curated questions across 35 AI subjects, with weighted reference insights; tasks include technical questions, literature reviews and research consulting.
Evaluation code
Official repository provides the benchmark platform, formats and eval.sh; users submit model responses and receive rubric and factuality results. OpenAI and Jina credentials are required.
Judging criteria
Researchers create weighted 1–3 insight criteria, with Claude-3.7-Sonnet assisting extraction; a factuality pipeline extracts claims and URLs, then judges source support.
Reproducibility and barriers
Questions and framework are public; assessment: API credentials, web retrieval and expert rubric interpretation are barriers. Baselines are author-run and dated March–June 2025.
Who reported the results
Official paper and site report author evaluations of OpenAI, Gemini, Grok, Perplexity and other baselines; no independent runs are established.

Evidence supplied by Exa

9. ResearchRubricsDeep-research report evaluation / human-authored rubrics

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Deep-research report evaluation / human-authored rubrics
Primary link
https://github.com/scaleapi/researchrubrics (new tab)
Release or update evidence
Paper/preprint arXiv:2511.07685 was published 2025-11-10 and the repository was created 2025-11-08; both fall within the window. Repository and Hugging Face data are public.
What it measures
Measures weighted criterion compliance across 101 prompts and 2,593 expert-written criteria, plus agreement between human and model judgments.
Task data
Public ScaleAI/researchrubrics download is documented in the repository: 101 human-written prompts and 2,593 reviewed criteria across nine domains, including penalty criteria.
Evaluation code
Actual MIT-licensed code ingests Markdown reports, chunks them, performs batch evaluation and calculates compliance. Default Gemini 2.5 Pro use through LiteLLM requires an API key and ScaleAI Hugging Face data.
Judging criteria
Human criteria cover requirements, reasoning, synthesis, references and communication. The released evaluator assigns Satisfied/Not Satisfied scores and computes weighted compliance; negative-weight rubrics are excluded from the denominator.
Reproducibility and barriers
Prompts, rubrics and code are public; external model access, cost and latency remain barriers. Commercial reports from the study may not be reproducible from the repository.
Who reported the results
Scale AI authors evaluated OpenAI, Gemini and Perplexity Deep Research; no independent leaderboard or result is found in the checked primary artifacts.

Evidence supplied by Exa

10. LiveResearchBench + DeepEvalDynamic web research benchmark and long-form report evaluator

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Dynamic web research benchmark and long-form report evaluator
Primary link
https://github.com/SalesforceAIResearch/LiveResearchBench (new tab)
Release or update evidence
Paper arXiv:2510.14240 was published in October 2025; the GitHub repository is dated 2025-10-17, and an ICLR 2026 artifact is public.
What it measures
Scores coverage, presentation, citation accuracy, citation traceability, fact-and-logic consistency and analysis depth using checklist, pointwise, pairwise and ensemble protocols.
Task data
Public static and realtime Hugging Face datasets: 100 expert-curated tasks across seven domains and ten categories, with checklists and dynamic date placeholders.
Evaluation code
Actual Salesforce code provides benchmark loading, DeepEval protocols, report handling and result output; the Hugging Face dataset is public. Enterprise-Deep-Research documents generation and invocation.
Judging criteria
Human checklists score coverage and presentation; rubric trees assess citation correctness and association; consistency and pairwise protocols assess factual logic and comparative depth.
Reproducibility and barriers
Tasks, code and static data are public; assessment: realtime evaluation depends on live web content, while generation and LLM judging require provider APIs.
Who reported the results
Salesforce-led evaluation covers 17 frontier systems; public leaderboard values are author/vendor results, not independent evaluations.

Evidence supplied by Exa

11. ReportBenchAcademic survey-grounded report quality and citation evaluation

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Academic survey-grounded report quality and citation evaluation
Primary link
https://github.com/ByteDance-BandAI/ReportBench (new tab)
Release or update evidence
arXiv:2508.15804 posted August 2025; GitHub repository public by the 2025 paper release, within window.
What it measures
Measures citation precision, recall, citation-match rate, reference counts, cited-statement coverage, and non-cited factual accuracy using statement-level verification.
Task data
Public ReportBench_v1.1.jsonl supplies 100 survey-grounded research tasks across ten domains and three prompt granularities; survey references furnish citation gold.
Evaluation code
Public repository contains processing, retrieval, citation and statement evaluators, metric calculators, benchmark data directories, and JSON-output support. Commercial web-product collection requires browser captures and dedicated processors.
Judging criteria
Published survey references are gold standards; cited claims are matched to retrieved passages, while non-cited claims use web-connected Gemini majority voting.
Reproducibility and barriers
Code and data are public, but web retrieval, browser capture, paid LLMs, and paid search services are required. Survey overlap may penalize valid divergent research; this is a methodological limitation assessment.
Who reported the results
ByteDance BandAI author-run results report precision, recall, citation matching, and non-cited accuracy for OpenAI and Gemini Deep Research. No independent result was located in checked primary artifacts.

Evidence supplied by Exa

12. FACTS Search (FACTS Benchmark Suite)Adjacent web-search factuality benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Adjacent web-search factuality benchmark
Primary link
https://www.kaggle.com/benchmarks/google/facts-search/leaderboard (new tab)
Release or update evidence
FACTS Benchmark Suite and Search benchmark launched December 2025 through the official announcement and arXiv:2512.10791; leaderboard remained public through 2026.
What it measures
Measures search-enabled answer F1, overall and attempted accuracy, hedging rate, and search count. A prompted auto-rater scores answer correctness.
Task data
1,884 Search questions comprise 890 public and 994 private items, including human-written hard-tail questions and three synthetic multi-hop subsets. The task evaluates factual search answers, not reports or list-building.
Evaluation code
Public benchmark examples and hosted Kaggle evaluation are available, but exact leaderboard reproduction requires the standardized Brave API, private set, hosted judge, and LLM setup.
Judging criteria
A prompted auto-rater compares answers with gold answers; F1 balances accuracy and attempted accuracy, while hedging and search count are diagnostics.
Reproducibility and barriers
Public examples and a common search API improve comparability, but the private set, API costs, availability, and auto-rater introduce barriers and possible bias. This is adjacent because it evaluates factual search answers.
Who reported the results
Google DeepMind, Google Research, and Kaggle author-hosted evaluation reports results for Gemini 3 Pro, GPT-5, and Claude 4.5 Opus. No independent result was counted.

Evidence supplied by Exa

13. BrowseComp-Plusfixed-corpus deep-research benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
fixed-corpus deep-research benchmark
Primary link
https://github.com/texttron/BrowseComp-Plus (new tab)
Release or update evidence
Public paper and repository released 2025-08-08 (arXiv:2508.06600; GitHub created/updated in 2025), within window.
What it measures
Measures answer accuracy, evidence-document recall, search-call count, calibration error, retrieval Recall@k, and nDCG@10.
Task data
Public fixed corpus, queries and qrels: 830 BrowseComp-derived questions over approximately 100,000 web documents, with human-verified evidence and hard negatives. This is a controlled web-derived corpus, not live search.
Evaluation code
Public official repository provides agent and retriever scripts, evaluation scripts, TREC-style qrels, reproduction documentation, and leaderboard instructions. Main evaluation uses top-five retrieval with a 512-token context limit.
Judging criteria
GPT-4.1 judges final-answer correctness; labeled evidence and gold documents determine retrieval Recall and nDCG. Citation accuracy is analyzed separately.
Reproducibility and barriers
The static corpus and released qrels improve reproducibility, while model/API access and substantial compute remain barriers. Dataset and retrieval artifacts are public through repository and Hugging Face links.
Who reported the results
Paper and official project results are author-reported comparisons of retrieval and model systems, including Search-R1, GPT-5, and GPT-5 with Qwen3 embeddings; no independent results are reported.

Evidence supplied by Exa

14. BrowseComp-ZHChinese live-web multi-hop browsing benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Chinese live-web multi-hop browsing benchmark
Primary link
https://github.com/PALIN2018/BrowseComp-ZH (new tab)
Release or update evidence
2025-04-24: original BrowseComp-ZH paper and repository. A separate AGI-Eval correction fork appeared in January 2026.
What it measures
Measures answer accuracy and calibration error for standalone LLMs and browsing agents, including model, reasoning, and browsing comparisons.
Task data
Public encrypted dataset of 289 Chinese multi-hop questions across 11 domains; repository decoding is required. The January 2026 AGI-Eval correction fork claims 24 answer corrections, so versions are not interchangeable.
Evaluation code
Public official repository includes decryption, model-evaluation, prediction/result, and calibration workflows. OpenCompass integration loads the dataset and uses an LLM judge.
Judging criteria
An LLM scorer judges answer correctness; the OpenCompass implementation explicitly uses LLMJudgeScorer. Official results report accuracy and calibration error.
Reproducibility and barriers
Encrypted questions and answers require the repository decryption procedure; live Chinese web access and proprietary model APIs are practical barriers. Encryption limits direct dataset inspection.
Who reported the results
Official author-run comparisons cover more than 20 systems, including OpenAI DeepResearch, O1, and Gemini-2.5-Pro. The reported figures are benchmark-author results.

Evidence supplied by Exa

15. WebWalkerQAweb traversal / information-seeking QA benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
web traversal / information-seeking QA benchmark
Primary link
https://github.com/Alibaba-NLP/WebAgent (new tab)
Release or update evidence
arXiv preprint 2501.07572 published 2025-01-13; official project announcement says WebWalker released 2025-01-14 and accepted ACL 2025.
What it measures
Measures question-answer accuracy and action count, with analyses by traversal depth, source count, domain, and language. Action count measures efficiency.
Task data
Public Hugging Face dataset: 680 Chinese/English questions over 1,373 webpages, with root URLs, hop counts and golden paths; tasks require single-site traversal or multisource navigation.
Evaluation code
Public official WebAgent repository provides the WebWalker framework, demo, and benchmark materials; the dataset is publicly distributed through Hugging Face. The paper specifies click-only interaction and a 15-step limit.
Judging criteria
GPT-4 evaluates answer correctness because generated answers vary, making exact match unsuitable. Correct-run action count measures efficiency.
Reproducibility and barriers
The public dataset and implementation support reproduction, but crawling, rendering, live-site changes, and web traversal create barriers. The benchmark requires access to current webpages.
Who reported the results
Original WebWalker results are benchmark-author runs. WebDancer reports separate WebWalkerQA runs; these are evaluations of the existing benchmark, not new benchmarks.

Evidence supplied by Exa

16. xbench-DeepSearchlive-web deep-search/tool-use benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
live-web deep-search/tool-use benchmark
Primary link
https://github.com/xbench-ai/xbench-evals (new tab)
Release or update evidence
Official README published 2025-05-28; DeepSearch-2505 and DeepSearch-2510 releases fall within the eligibility window.
What it measures
Measures task accuracy and supports cost/task and time/task reporting; the public runner enables repeated evaluation.
Task data
Public repository provides encrypted CSV datasets for DeepSearch-2505 and DeepSearch-2510, covering planning, search, reasoning and summarization; updates are quarterly.
Evaluation code
Public repository includes xbench_evals.py, data artifacts and decryption/evaluation workflows; versions 2505/2510 and an LLM-judge scorer are inspectable.
Judging criteria
An LLM judge scores answers; questions underwent manual collection, human validation and quality filtering.
Reproducibility and barriers
Encrypted files and proprietary agent interfaces limit full reproduction; public runner and data support model-side reproduction where APIs exist. Assessment: many product runs were manual.
Who reported the results
Official leaderboard reports author-run provider/product evaluations, including May and August 2025 tables; no independent reproductions are established.

Evidence supplied by Exa

17. SealQAsearch-augmented factual reasoning benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
search-augmented factual reasoning benchmark
Primary link
https://huggingface.co/datasets/vtllms/sealqa (new tab)
Release or update evidence
arXiv:2506.01062 was published 2025-06-01; the official dataset is publicly released within the eligibility window.
What it measures
Measures factual-answer accuracy with and without search across Seal-0 and Seal-Hard; LongSeal measures evidence selection among distractor documents.
Task data
Public data include Seal-0, Seal-Hard and LongSeal, covering conflicting, noisy or unhelpful search results across domains; LongSeal supplies many documents with one relevant answer source.
Evaluation code
Public Hugging Face data and paper artifacts are available; the benchmark is dynamic and periodically updated. No standalone official evaluator repository was verified.
Judging criteria
Scores factual answer accuracy under no-search and search/tool conditions. Exact automated judge implementation is not established from inspected sources.
Reproducibility and barriers
Public data aid access, but version updates, search-provider behavior and proprietary APIs limit frozen-test reproduction. A canary string addresses contamination.
Who reported the results
Reported results are author-run evaluations using specified models and search systems; no independent reproductions are established.

Evidence supplied by Exa

18. Mind2Web 2Core research benchmark: agentic search / long-horizon web research and citation-backed synthesis

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Core research benchmark: agentic search / long-horizon web research and citation-backed synthesis
Primary link
https://osu-nlp-group.github.io/Mind2Web-2/ (new tab)
Release or update evidence
Initial public release was 2025-06-26; the 2025-10-23 update released evaluation scripts for both public dev and test sets, within the eligibility window.
What it measures
Measures task completion through Partial Completion, Success Rate and Pass@3, alongside completion time, answer length, correctness and source attribution.
Task data
Public artifact provides 130 long-horizon web-search tasks involving synthesis, list retrieval, time-varying information and citations; dev and test data are available.
Evaluation code
Public MIT repository includes run_eval.py and task-specific Extractor/Verifier workflows; the 2025-10-23 release covers both public dev and test sets.
Judging criteria
An Extractor parses claims and citations, while a Verifier checks them against webpages or screenshots; leaf scores aggregate into task and root metrics.
Reproducibility and barriers
Public code and tasks improve reproduction, but live webpages, crawling, LLM/API judges and changing information introduce cost and nondeterminism. Assessment: results may drift.
Who reported the results
Original authors report 2025 agent and human evaluations, including private-test-era results; no independent reproduction is established.

Evidence supplied by Exa

19. Deep Research Bench (DRB; FutureSearch)Web research, dataset/list compilation and evidence discovery; frozen/live variants

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Web research, dataset/list compilation and evidence discovery; frozen/live variants
Primary link
https://arxiv.org/abs/2506.06287 (new tab)
Release or update evidence
2025-06: paper arXiv:2506.06287 introduces the benchmark; the official FutureSearch post is dated June 25, 2025.
What it measures
Eight research task families include dataset compilation, reference-class enumeration, evidence gathering, source tracing, numeric derivation and claim validation. The paper has 89 task instances; later leaderboard snapshots differ.
Task data
Full tasks are explicitly withheld to limit contamination; the paper publishes eight examples. Human-worked references and a frozen RetroSearch corpus support controlled evaluation, but a complete public bundle was not verified.
Evaluation code
The paper describes agent tooling and automated trace evaluation, and a public leaderboard exists. No complete public task/evaluator repository was verified; full reproduction code availability is partial/unclear.
Judging criteria
Task-specific scoring: row precision/recall/F1; URL or evidence recall; binary numeric/source correctness; and normalized distance from human probability judgments. LLMs assist entity matching and source-reliability checks; trace diagnostics are separate.
Reproducibility and barriers
RetroSearch controls web drift, but withheld tasks, proprietary product runs and incomplete evaluator artifacts limit independent reproduction. Assessment: repeatability is partial.
Who reported the results
FutureSearch authors report evaluations of LLMs, agents and commercial research products; leaderboard results are vendor/author-run, not independent validation.

Evidence supplied by Exa

20. DeepSearchQA (Google DeepMind)comprehensive multi-entity/list retrieval and deep-search benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
comprehensive multi-entity/list retrieval and deep-search benchmark
Primary link
https://arxiv.org/abs/2601.20975 (new tab)
Release or update evidence
Technical report posted 2026-01-28, within the cutoff; the public dataset is available on Hugging Face and the official leaderboard is hosted by Kaggle.
What it measures
Measures exhaustive set-answer generation through fully correct and incorrect set rates, extraneous-answer rate and F1, covering collation, deduplication, entity resolution and stopping.
Task data
Public dataset contains 900 expert-annotated prompts across 17 fields with objectively verifiable gold sets, answer types and time or source anchors; about 65% are set-answer tasks.
Evaluation code
Public dataset is downloadable, but no official standalone evaluator repository was verified. Third-party runners exist and are not Google code.
Judging criteria
A fully correct response exactly matches the gold set; F1 balances missing and extraneous entities. Evaluation uses an automated judging methodology.
Reproducibility and barriers
Static or time-anchored tasks and public data aid reproduction; web access, APIs and judge-model choice remain barriers. Official evaluator-code availability is unknown.
Who reported the results
Google DeepMind results are author-run; Kaggle results are platform-evaluated submissions and do not establish independent scientific replication.

Evidence supplied by Exa

21. DRACO Benchmark (Perplexity Research)long-form research report quality/citation benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
long-form research report quality/citation benchmark
Primary link
https://research.perplexity.ai/articles/evaluating-deep-research-performance-in-the-wild-with-the-draco-benchmark (new tab)
Release or update evidence
Eligible: public release announced in 2026; technical report arXiv:2602.11685 is dated 2026-02-12, before the 2026-09-26 cutoff. The announcement says the benchmark, rubrics and judge prompt are open sourced.
What it measures
Evaluates factual accuracy, breadth/depth, presentation quality, objectivity and citation quality across 100 tasks using weighted binary rubric criteria, including penalties for unsupported claims.
Task data
Public dataset availability: production-derived Perplexity Deep Research requests were de-identified, reformulated and filtered through five stages, with expert-reviewed rubrics.
Evaluation code
Public benchmark, rubrics, judge prompt and linked dataset are released; the exact harness boundary is unclear, and proprietary production sampling and research tools remain unreproducible.
Judging criteria
An LLM judge assigns weighted binary verdicts to rubric criteria; rubrics assess accuracy, completeness/depth, presentation and primary-source citation quality, with reliability checked across three judge models.
Reproducibility and barriers
Public tasks, rubrics and judge prompt support reproduction, but production-query sampling, proprietary tools and English single-turn scope remain barriers or limitations (assessment).
Who reported the results
Perplexity's own comparison of four deep-research systems; reported scores and leadership claims are vendor-run, not independent.

Evidence supplied by Exa

22. WANDR: A Benchmark for Wide and Deep Researchcomprehensive multi-entity discovery plus evidence-supported enrichment (wide-and-deep research)

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
comprehensive multi-entity discovery plus evidence-supported enrichment (wide-and-deep research)
Primary link
https://arxiv.org/abs/2608.14747 (new tab)
Release or update evidence
Submitted to arXiv 2026-08-14, within the 2026-09-26 cutoff. The official repository README was available by 2026-07-14 and contains the released task/evaluator tree.
What it measures
Scores record-level and hierarchical precision, recall and F1, including hard complete-subtree scores, soft partial-credit scores, retrieval-only versus full-record results, and task rollups.
Task data
Public 500-task packages encode qualification hierarchies and requested record volumes. Packages include instructions, schemas, fixtures, labels and evaluator artifacts, but not a static exhaustive gold universe.
Evaluation code
Public repository contains source tasks, Harbor adapter, generated packages, task-local evaluator, scripts, configurations and reports, with validation and full-run instructions.
Judging criteria
The evaluator fetches and canonicalizes cited pages, deduplicates entities, verifies claims against pages and excerpts, then aggregates task-specific record verdicts hierarchically. Precision measures submitted quality; recall measures coverage against requested volume.
Reproducibility and barriers
Public tasks, Docker packages, evaluator, configs and manifests support reproduction. End-to-end runs require live retrieval and paid APIs; web drift, bot walls, provider settings and judge nondeterminism remain barriers (assessment).
Who reported the results
Original authors' pinned runs over six production systems; these are author-reported benchmark results, not independent reproductions.

Evidence supplied by Exa

23. DeepWide (Table-as-Search business-development benchmark)constrained entity discovery and attribute enrichment; paper-defined evaluation

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
constrained entity discovery and attribute enrichment; paper-defined evaluation
Primary link
https://arxiv.org/abs/2602.06724 (new tab)
Release or update evidence
2026-02: Table-as-Search (arXiv:2602.06724) introduces the 20-query DeepWide evaluation.
What it measures
Column-F1 measures identification of entities satisfying complex constraints; Item-Precision measures correctness of retrieved attributes. Tasks use fixed retrieval quantities rather than an exhaustive-universe claim.
Task data
20 real-world business-development and e-commerce queries require constrained discovery and attribute enrichment. Expert-verified references combine pooled system matches; a complete downloadable bundle was not separately verified.
Evaluation code
Table-as-Search agent code and prompts are public in Marco-Search-Agent; an independently packaged scorer and reference bundle for these 20 tasks were not verified. The repository's DeepWideSearch data/eval directory is separate.
Judging criteria
Experts verify candidate validity and requested information against all stated constraints. Column-F1 scores entity identification, while Item-Precision scores attribute correctness; fixed quantities, exclusions and dynamic reference unions address open-ended completeness.
Reproducibility and barriers
Agent code is public, but a complete bundle of these 20 tasks, reference sets and scoring code was not verified. Live retrieval, commercial interfaces, dynamic reference unions and expert judgments complicate reproduction.
Who reported the results
Original authors' Table-as-Search runs are author-reported; no independent reproduction was established here.

Evidence supplied by Exa

24. DeepWeb-BenchDeep research benchmark; comprehensive multi-entity evidence and derivation

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Deep research benchmark; comprehensive multi-entity evidence and derivation
Primary link
https://arxiv.org/abs/2605.21482 (new tab)
Release or update evidence
Submitted 2026-05-20, within cutoff. The primary paper says data, rubrics and evaluation code are publicly released.
What it measures
Scores retrieval, derivation, reasoning and calibration across 100 cases using entity-by-dimension cells, family scores and source-provenance disclosure levels.
Task data
Public Hugging Face release: 100 cases, 900 model results/answers/score records, summaries and provenance. Tasks require 6–10 entities across 6–10 dimensions with cross-source evidence.
Evaluation code
Actually present in HF code/: validation, leaderboard/report rebuilding, rule-prompt grader reruns and an OpenAI-compatible model runner. Aggregation needs no keys; live reruns do.
Judging criteria
Explicit per-cell rules, provenance levels and cross-source checks determine scores; evaluation is not solely free-form judging.
Reproducibility and barriers
Data and code are downloadable, but model, grader, search and scrape APIs require keys. Raw tool traces, third-party source snapshots and local MCP/API state are excluded.
Who reported the results
Author-reported evaluation of nine frontier models; reported findings concern overall scores and retrieval, derivation and calibration error patterns.

Evidence supplied by Exa

25. From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task EvolutionAutomatically constructed verifiable deep-research benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Automatically constructed verifiable deep-research benchmark
Primary link
https://arxiv.org/abs/2608.02163 (new tab)
Release or update evidence
Submitted 2026-08-03, within cutoff. The primary paper says code and data are publicly available.
What it measures
Evaluates 500 tasks across 31 topics and 10 categories using DAG atomic steps, checkpoints, fact-grounded pointwise rubrics, and model discrimination and stability analyses.
Task data
Constructed from Wikipedia QA through Explorer–Formalizer–Challenger evolution; each task includes a query, DAG, checkpoints and aligned rubrics. Public availability is claimed by the paper.
Evaluation code
Paper links https://github.com/chr6192/TaskEvolving.git (new tab) and states that implementation, data and results are public. Repository contents were not retrievable in this check, so evaluator files remain unverified.
Judging criteria
Source-grounded checkpoint and DAG rubrics assign pointwise fact-based scores; the paper describes the method as human-aligned and stable.
Reproducibility and barriers
Public-release claim is explicit, but construction requires relatively high frontier-model cost. Exact repository data layout and license remain unverified.
Who reported the results
Author-reported experiments demonstrate discrimination across models and query types; no independent results were established.

Evidence supplied by Exa

26. DeepResearch-9Kchallenging multi-hop web research benchmark with agent trajectories

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
challenging multi-hop web research benchmark with agent trajectories
Primary link
https://arxiv.org/abs/2603.01152 (new tab)
Release or update evidence
ArXiv paper 2603.01152 was available in March 2026 and identifies SIGIR 2026 publication (July 20–24, 2026), within cutoff. Official repository was created 2026-02-06.
What it measures
Measures final-answer accuracy, search-tool usage, and trajectory correctness across three difficulty levels.
Task data
Public 9,000-question Hugging Face release includes difficulty labels, questions, answers, trajectories, and a 3,974-sample hard subset. Data are largely synthetic and teacher-generated.
Evaluation code
Public official repository includes training, inference, SFT/RL, environment, and evaluation scripts. The paper’s construction pipeline is released, with DeepSeek-V3 LLM judging.
Judging criteria
LLM judging scores final-answer correctness; difficulty reflects search-chain complexity and entity obfuscation. Hard items were filtered by incorrect teacher verdicts.
Reproducibility and barriers
Data and code are public, but training requires substantial resources and model/API dependencies. Teacher-generated trajectories and LLM judging create assessment dependence.
Who reported the results
Authors’ experiments; DeepResearch-R1 is the paired training/agent framework, not an additional benchmark. No independent replication established.

Evidence supplied by Exa

27. Mr.LHDRMultimodal real-world long-horizon deep-research benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Multimodal real-world long-horizon deep-research benchmark
Primary link
https://arxiv.org/abs/2609.11318 (new tab)
Release or update evidence
Submitted 2026-09-10, within 2026-09-26 cutoff.
What it measures
Measures final-answer and dependency-aware conclusion quality using OA, Strict Accuracy, Checklist Score, and Dependency-Aware Checklist Score.
Task data
Public dataset uses multimodal evidence including images, maps, PDFs, logos, charts, tables, and video frames. Metadata, checklists, sources, and images are released, but some fields are encrypted or canary protected.
Evaluation code
Public Apache-2.0 repository includes requirements, decryption, run, and evaluation scripts. Actual execution requires the dataset and model/web-search access.
Judging criteria
Scores final answers and intermediate conclusions against annotated dependencies using OA, SA, CS, and DACS.
Reproducibility and barriers
Code and dataset are public, but encrypted metadata and decrypt.py create a material reproduction barrier; raw data are not fully transparent.
Who reported the results
Author-reported evaluation; strongest system achieved 43.1% OA and 34.3% SA. Removing images reduced DACS by 12.6 points.

Evidence supplied by Exa

28. FinSearchCompfinancial open-web search and analyst-style reasoning

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
financial open-web search and analyst-style reasoning
Primary link
https://arxiv.org/abs/2509.13160 (new tab)
Release or update evidence
Submitted 2025-09-16, within cutoff. Official project and repository are public.
What it measures
Measures binary answer correctness across time-sensitive retrieval, historical lookup, and complex historical investigation.
Task data
Public 635-question expert-crafted dataset covers Global and Greater China subsets and includes questions, tool templates, answers, and traces.
Evaluation code
Actual public evaluator code is available at https://github.com/randomtutu/FinSearchComp (new tab), including data/finsearchcomp_data.json, eval/eval.py, and runnable commands.
Judging criteria
Rubric-guided LLM judging assigns 0/1 correctness, applying numerical tolerance and checking financial conventions and supporting evidence.
Reproducibility and barriers
Public data and evaluator support reproduction, but live APIs, web changes, and model/tool access remain practical assessment barriers.
Who reported the results
Author-run model comparison; no independent reproduction established.

Evidence supplied by Exa

29. Finance Agent Benchmarkfinancial SEC-filing research-agent benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
financial SEC-filing research-agent benchmark
Primary link
https://arxiv.org/abs/2508.00828 (new tab)
Release or update evidence
ArXiv release August 2025, within cutoff.
What it measures
Measures naive and class-balanced accuracy across nine finance categories, plus execution time and cost.
Task data
Partially public: 537 expert-authored SEC/EDGAR questions comprise 50 public validation, 150 private validation, and 337 private test items.
Evaluation code
Actual MIT-licensed harness and public validation data are available at https://github.com/vals-ai/finance-agent (new tab); Zenodo also provides the harness.
Judging criteria
Rubric-based component grading checks calculations and detects contradictions before assigning correctness.
Reproducibility and barriers
Public 50-item validation and code permit partial reproduction; private splits, live filings/search, and paid APIs limit full reproduction (assessment).
Who reported the results
Benchmark-author runs; no independent reproduction established.

Evidence supplied by Exa

30. FinRetrievalfinancial structured-data retrieval by AI agents

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
financial structured-data retrieval by AI agents
Primary link
https://arxiv.org/abs/2603.04403 (new tab)
Release or update evidence
January 2026 release; within cutoff.
What it measures
Measures numeric-answer accuracy across 500 questions and 14 model/tool configurations, with tool-call trace analysis.
Task data
Public release includes 500 questions, 7,000 responses, ground truth, scores, and complete tool-call traces comparing web-only and structured MCP/API access.
Evaluation code
Actual public evaluator is available at https://github.com/daloopa/finretrieval (new tab); README provides setup and execution commands.
Judging criteria
Automatic numeric matching compares answers with ground truth; edge cases and fiscal-period conventions receive manual review.
Reproducibility and barriers
Dataset, evaluator, scores, and traces are public; Daloopa MCP and commercial model/API access remain practical rerun barriers (assessment).
Who reported the results
Daloopa-author evaluation; no independent reproduction established.

Evidence supplied by Exa

31. MedBrowseCompmedical live-web/deep-research and biomedical evidence retrieval

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
medical live-web/deep-research and biomedical evidence retrieval
Primary link
https://arxiv.org/abs/2505.14963 (new tab)
Release or update evidence
Submitted 2025-05-20, within cutoff; repository and dataset match arXiv 2505.14963.
What it measures
Accuracy on MedBrowseComp-50 and MedBrowseComp-605, covering structured extraction and deep research, judged by GPT-4.1-mini with human checking.
Task data
Public 50-sample and 605-sample benchmarks cover hematology/oncology, PubMed, ClinicalTrials.gov, FDA Orange Book and market data.
Evaluation code
Public official repository https://github.com/shan23chen/MedBrowseComp (new tab) contains final50.csv, final121.csv and processing/evaluation scripts, including process_NCT_predictions.py. The 50/605 counts are verified; the paper describes 121 trials expanded into 605 tasks.
Judging criteria
GPT-4.1-mini judges answer correctness, with human checking of judge agreement.
Reproducibility and barriers
Partial reproduction is possible with public files and scripts, but live medical sources and proprietary systems create access and drift barriers (assessment).
Who reported the results
Original authors' commercial-system and Claude computer-use runs; no independent benchmark runner established.

Evidence supplied by Exa

32. PaSa / AutoScholarQuery and RealScholarQueryacademic literature discovery and comprehensive paper search

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
academic literature discovery and comprehensive paper search
Primary link
https://aclanthology.org/2025.acl-long.572/ (new tab)
Release or update evidence
ACL 2025 primary paper release, within cutoff; official ByteDance repository is linked by the paper and search result.
What it measures
Recall@20, Recall@50, Recall@100 and precision measure relevant-paper retrieval, document selection and crawling.
Task data
AutoScholarQuery has public 33,551 training, 1,000 development and 1,000 test synthetic queries. RealScholarQuery has 50 manually gathered researcher queries and an annotated 200 query-paper selector set.
Evaluation code
Public code and datasets are stated at https://github.com/bytedance/pasa (new tab). Repository search results show runnable agent scripts and benchmark support using scholarly retrieval, citation crawling and ar5iv parsing.
Judging criteria
Ranked relevant-paper retrieval is scored against annotated sets using recall and precision; separate human checks assess query and query–paper relevance.
Reproducibility and barriers
AutoScholarQuery is comparatively reproducible, while RealScholarQuery may drift with manually collected queries and changing indexes; search/API access and model compute are required (assessment).
Who reported the results
Authors' benchmark runs, not independent validation. Reported comparisons are benchmark-author results.

Evidence supplied by Exa

33. OpenBenchmarks Multi-turn Company Search Benchmarkmulti-turn web-search evaluation for comprehensive company-set discovery

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
multi-turn web-search evaluation for comprehensive company-set discovery
Primary link
https://openbenchmarks.com/multi-turn-company-search (new tab)
Release or update evidence
Official page records first complete benchmark publication on 2026-08-22, additions on 2026-08-26 and updates on 2026-09-15; all fall within the cutoff.
What it measures
Precision, recall, F1 and exact-set accuracy score company-set discovery; median latency and cost measure efficiency across search-only and search-plus-fetch settings.
Task data
Full board gold is locked: 45 hand-labelled questions with 375 canonical memberships. A separate public 10-question search-only sample with frozen reference companies is available via Hugging Face.
Evaluation code
Public MIT-licensed GitHub runner and deterministic offline judge validate schemas, receipts, canonical matching and aggregation. Running agents requires provider credentials; saved-artifact judging is offline.
Judging criteria
Canonical set comparison counts true positives, false positives and false negatives for precision, recall, F1 and exact-set accuracy. Three trials are aggregated; SD measures trial variability, not confidence.
Reproducibility and barriers
Partial reproduction is possible with the public runner, judge and 10-question sample, but locked gold, live providers, costs and web drift remain barriers (assessment).
Who reported the results
OpenBenchmarks editorial-board/vendor runs; no independent results established. Board scores use the locked 45-question set, while the public sample is not score-comparable.

Evidence supplied by Exa

34. MM-BrowseCompmultimodal deep web browsing benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
multimodal deep web browsing benchmark
Primary link
https://github.com/MMBrowseComp/MM-BrowseComp (new tab)
Release or update evidence
ArXiv v1 2025-08-14; official repository says full codebase released 2025-08-20 and dataset expanded to 400 questions 2026-01-02. Version drift remains: v1 describes 224/244 questions, current release 400.
What it measures
Final-answer accuracy measures correctness; checklist analysis measures multimodal dependencies and reasoning paths beyond aggregate scores.
Task data
Public encrypted JSONL includes the current 400-question update at data/MMBrowseComp_400.jsonl. Prompts and evidence may include images and videos; canary/decryption protects questions and answers.
Evaluation code
Public src/decrypt.py, src/gen_answer.py and src/eval.py implement decryption, answer generation and LLM judging. The repository includes released data and evaluation scripts.
Judging criteria
An LLM judges reference-answer correctness, while verified per-question checklists diagnose dependency and reasoning paths.
Reproducibility and barriers
Partial reproduction is possible with public code/data, but encrypted contents, canary/decryption, API-backed models, live web access and configuration create barriers (assessment).
Who reported the results
Paper/author-reported model evaluations; no independent result established here.

Evidence supplied by Exa

35. MMSearch-Plusprovenance-aware multimodal browsing/search benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
provenance-aware multimodal browsing/search benchmark
Primary link
https://github.com/mmsearch-plus/MMSearch-Plus (new tab)
Release or update evidence
ArXiv released 2025-08-29; official README records all data samples released to Hugging Face 2025-09-26, within cutoff.
What it measures
Answer accuracy measures correctness on 311 tasks; bounding-box, cropping and provenance analyses measure localized visual search and evidence-chain behavior.
Task data
Public encrypted dataset includes questions/images, ground truth, alternatives, metadata and Set-of-Mark annotations. Tasks require localized visual cues, spatial-temporal reasoning, iterative retrieval and cross-validation.
Evaluation code
Public official repository provides the agentic rollout framework, evaluation script and Set-of-Mark annotations; Hugging Face data are linked. External search infrastructure is not claimed to be bundled.
Judging criteria
Answers are scored for accuracy, while visual localization/cropping and provenance-aware retrieval analyses assess multimodal evidence chains rather than text-only shortcuts.
Reproducibility and barriers
Partial reproduction requires public data/code plus decryption, model/API and live-search dependencies; changing image sources and retrieval noise may affect results (assessment).
Who reported the results
Author-run paper/official leaderboard results; independent reproduction not verified.

Evidence supplied by Exa

36. BrowseComp-V³ (BrowseComp-V3)visual, vertical, verifiable multimodal deep-search benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
visual, vertical, verifiable multimodal deep-search benchmark
Primary link
https://github.com/Halcyon-Zhang/BrowseComp-V3 (new tab)
Release or update evidence
2026-02: arXiv:2602.12876 and official project/repository. Unicode V³ and ASCII V3 designate the same benchmark.
What it measures
Measures final-answer success and process-level adherence using expert-validated intermediate subgoals and search trajectories.
Task data
Partial/public: encrypted downloadable data contain 300 questions across 24 subdomains, visual assets, evidence, gold trajectories and subgoals; repository scripts decrypt JSON and images.
Evaluation code
Actual public repository code includes dataset download/decryption, rollout evaluation, score summarization, an OmniSeeker runner, documentation and smoke tests.
Judging criteria
Scores final-answer correctness and success alongside process and subgoal adherence for cross-modal, multi-hop evidence integration.
Reproducibility and barriers
Public GitHub/Hugging Face artifacts and CC BY 4.0 claim; encrypted samples, key, live search, judge model and APIs remain dependencies. Assessment: web volatility and closed baselines limit exact reproduction.
Who reported the results
Author-reported paper/project evaluations include OmniSeeker and MLLM comparisons; independent results are not verified.

Evidence supplied by Exa

37. ScholarQuestAcademic literature/paper-search benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Academic literature/paper-search benchmark
Primary link
https://arxiv.org/abs/2606.20235 (new tab)
Release or update evidence
Submitted 2026-06-18 (within window); public code/data repository linked by paper.
What it measures
Measures retrieval recall at 25, 100 and all results, plus search efficiency, tool use and robustness to intent and answer-set size.
Task data
Public: 1,111 queries from 1,000+ CS topic seeds use four intents, arXiv answer sets of 5–200 papers and a million-scale ScholarBase corpus with metadata and citations.
Evaluation code
Actual public repository includes the benchmark dataset, construction pipeline, analysis scripts and ScholarBase/Lewen search backend.
Judging criteria
Scores retrieval recall against answer_arxiv_ids after arXiv-ID normalization; process statistics cover rounds, calls, candidates and recall efficiency.
Reproducibility and barriers
Public code/data; reproduction requires deploying ScholarBase/Lewen and potentially using external search APIs or models (assessment).
Who reported the results
Paper authors report PaperScout and hybrid baselines; results are author-reported, with no independent validation established here.

Evidence supplied by Exa

38. Sage: Benchmarking and Improving Retrieval for Deep Research AgentsScientific literature retrieval benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Scientific literature retrieval benchmark
Primary link
https://arxiv.org/abs/2602.05975 (new tab)
Release or update evidence
Submitted 2026-02-06 (within window); public GitHub repository linked from paper.
What it measures
Reasoning-intensive scientific paper discovery: exact-match target-paper retrieval and relevance-weighted coverage of open-ended literature queries.
Task data
Public repository contains 600 short-form and 600 open-ended queries, with paper IDs/titles and relevance-tier references. The paper describes four domain-specific 50,000-paper corpora; a complete corpus bundle was not verified.
Evaluation code
Public question/reference JSON files and metric definitions are verified. The inspected repository README does not establish a turnkey evaluator or full agent/retriever implementation.
Judging criteria
Short-form exact match checks whether the target paper appears in answer text or citations. Open-ended weighted recall gives seed papers relevance 2, shared-reference papers 1 and other papers 0.
Reproducibility and barriers
Query/reference data are public; reproducing agent experiments requires the corresponding corpus, retrieval indexes and model/API setup. Corpus augmentation adds processing cost.
Who reported the results
Authors evaluate six deep-research agents and DR Tulu-backed retrievers; reported findings are author-generated, not independently validated here.

Evidence supplied by Exa

39. AstaBench (literature-search and research-synthesis components)Core subset — scientific literature discovery/synthesis; broader science suite is adjacent

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Core subset — scientific literature discovery/synthesis; broader science suite is adjacent
Primary link
https://allenai.org/asta/bench (new tab)
Release or update evidence
Public Ai2 announcement dated 2025-08-26 and benchmark paper submitted 2025-10-24; within window.
What it measures
Measures paper finding, literature retrieval and QA, long-form review answers, literature-review tables, quality, cost and tool-controlled agent behavior.
Task data
Public: 11 benchmarks and 2,400+ problems cover literature, coding, data analysis and discovery; literature components include PaperFindingBench, ScholarQA-CS2, LitQA2 and ArxivDIGESTables-Clean.
Evaluation code
Actual public GitHub suite provides task datasets, interfaces and standardized tools; Asta Scientific Corpus supports date- and corpus-restricted search and retrieval.
Judging criteria
Scores task-specific answers or retrieval; ScholarQA-CS2 adds coverage and citation precision, while logs record cost, tools and traces.
Reproducibility and barriers
Open benchmark/framework and baseline agents; reproduction depends on Asta Environment or corpus snapshots and compatible agent-evaluation tooling (assessment).
Who reported the results
Ai2 authors report 57-agent evaluations; announcement reports Asta Scholar QA/Elicit/SciSpace and Asta Paper Finder results. Author/vendor results, not independently validated.

Evidence supplied by Exa

40. LiveDRBench (Microsoft), from “Characterizing Deep Research: A Benchmark and Formal Definition”open-web deep-research claim discovery; includes entity/list retrieval and evidence-grounded synthesis

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
open-web deep-research claim discovery; includes entity/list retrieval and evidence-grounded synthesis
Primary link
https://github.com/microsoft/LiveDRBench (new tab)
Release or update evidence
Public repository created 2025-07-25; paper arXiv:2508.04183 is 2025. Data were collected May–June 2025, within the requested window.
What it measures
Measures claim-level precision, recall and F1, including recursive claim/subclaim correctness; also records sources, branching and backtracking.
Task data
Public: 100 live open-web tasks cover science and world events, with prompts, output formats, ground-truth JSON claims and references; Hugging Face provides the dataset.
Evaluation code
Actual public MIT repository includes src/evaluate.py and instructions; GPT-4o via OpenAI API judges claim agreement before computing information-retrieval metrics.
Judging criteria
Scores correctness and completeness of substantive claims; GPT-4o maps predictions to ground-truth claims, and unsupported incorrect subclaims receive no credit.
Reproducibility and barriers
Public code/data and a stated refresh plan support reproduction, but live-web changes and GPT/API dependence affect repeatability. It is explicitly open-web, not enterprise search.
Who reported the results
Microsoft authors report runs involving OpenAI, Perplexity, Google/Gemini and an open-source agent; independent reproduction is not established here.

Evidence supplied by Exa

41. DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation (LG AI Research)long-form expert report quality, evidence grounding and report-wide fact verification

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
long-form expert report quality, evidence grounding and report-wide fact verification
Primary link
https://github.com/hanjanghoon/DEER (new tab)
Release or update evidence
Paper arXiv:2512.17776 first released 2025-12-19; official repository created 2026-02-03. Both fall within the requested window.
What it measures
Measures report fulfillment, analytical soundness, coherence, style and ethics through 101 rubric items, plus claim factuality, citation support, evidence quality and sufficiency.
Task data
Public benchmark artifact, but access is gated by a password-protected archive; it contains 50 expert-report tasks across 13 domains derived from Humanity’s Last Exam and internal queries.
Evaluation code
Official MIT-licensed evaluation code is public; dataset access is gated under a custom license permitting non-commercial research but prohibiting redistribution, mirroring and public posting.
Judging criteria
LLM judges apply fixed rubric items and expert guidance; a fact-verification module extracts claims and citations, then checks external evidence through web search.
Reproducibility and barriers
Gated data, web-search verification, LLM judging, licensing restrictions and anti-contamination requirements materially limit fully frictionless reproduction.
Who reported the results
Reported human-correlation and system-comparison results are LG AI Research authors’ experiments; independent validation was not established here.

Evidence supplied by Exa

42. PaSaMaster-Benchmultidisciplinary scientific literature retrieval / comprehensive paper-set discovery

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
multidisciplinary scientific literature retrieval / comprehensive paper-set discovery
Primary link
https://arxiv.org/abs/2605.14306 (new tab)
Release or update evidence
2026-05: arXiv:2605.14306 introduces PaSaMaster-Bench, distinct from PaSa’s earlier AutoScholarQuery/RealScholarQuery datasets.
What it measures
Measures top-20 recall, precision, F1 and NDCG, plus source-hallucination rate and token cost; the paper reports 244 expert-curated tasks across 38 disciplines.
Task data
Availability not established: tasks contain multi-constraint literature intents, target paper sets and expert checklist annotations, but a public task/answer bundle was not confirmed.
Evaluation code
Public PaSaMaster system and retrieval code are available; benchmark-specific scorer and data files were not verified, so evaluation release status remains uncertain.
Judging criteria
Expert checklists determine whether retrieved papers satisfy the full intent; top-20 set and ranking metrics compare results with target papers, while source checks support hallucination scoring.
Reproducibility and barriers
Assessment: live scholarly sources, proprietary model/search comparisons and a potentially unreleased benchmark bundle impede exact reproduction, although system code is public.
Who reported the results
Reported results are the paper authors’ runs; no independent reproduction was established.

Evidence supplied by Exa

43. AutoResearchBenchscientific literature deep identification and comprehensive set discovery

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
scientific literature deep identification and comprehensive set discovery
Primary link
https://arxiv.org/abs/2604.25256 (new tab)
Release or update evidence
arXiv 2604.25256 (2026) falls within the window; official project resources publicly release code and benchmark data.
What it measures
Deep research uses exact-answer accuracy; wide research uses set IoU for coverage and precision. The benchmark contains 1,000 problems: 600 deep and 400 wide.
Task data
Public obfuscated JSONL bundle on Hugging Face, decrypted locally; it covers eight CS areas and contains human-verified deep and wide literature-search tasks.
Evaluation code
Public GitHub code includes inference, academic/web search tools, prompts, utilities, decryption scripts and separate deep- and wide-search evaluators.
Judging criteria
Deep answers receive exact-match scoring; wide answers are scored by IoU against gold paper sets using a standardized ReAct agent and DeepXiv search setup.
Reproducibility and barriers
Public code, data and evaluator support reproduction; decryption, model/API credentials, search access, token cost and live-search drift remain practical barriers.
Who reported the results
Results are the original authors’ baseline runs across more than 10 models and agents; independent reproduction was not established.

Evidence supplied by Exa

44. DRBENCHER (Deep Research Benchmarker)entity identification + property retrieval + quantitative computation for web research agents

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
entity identification + property retrieval + quantitative computation for web research agents
Primary link
https://arxiv.org/abs/2604.09251 (new tab)
Release or update evidence
arXiv first posted 2026-04-10 (2604.09251; v3 rendered 2026-08-09), within cutoff; IBM Research and the paper link the public implementation.
What it measures
Measures answer accuracy, entity identification, validity, human quality and semantic diversity for multi-hop entity, property-retrieval and calculation tasks; difficulty is summarized by CCI.
Task data
Public JSONL tasks are available. The paper describes 268 human-validated questions, while repository documentation lists a 255-question main evaluation set; pin versions rather than treating these counts as interchangeable.
Evaluation code
Public MIT-licensed IBM/DrBencher repository includes pipeline, schemas, released JSONL data, decryption/evaluation utilities and computation-based gold-answer checking.
Judging criteria
Programmatic checks require reproducible calculations, supported clues, no entity leakage and unambiguous questions; answers use domain-specific tolerances, while humans assess validity.
Reproducibility and barriers
Public code, schema and data support reproduction; live Wikidata/Wikipedia and domain APIs, changing values, model/API access and stale data remain barriers.
Who reported the results
Results are IBM authors’ human annotations and six-model evaluations, not independent reproductions; the paper identifies property retrieval as the dominant failure mode.

Evidence supplied by Exa

45. Reka Research-Eval (including the ndurner evaluator extension)search-augmented web research / grounded multi-hop question answering (static QA-style task suite, not comprehensive list-building)

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
search-augmented web research / grounded multi-hop question answering (static QA-style task suite, not comprehensive list-building)
Primary link
https://github.com/reka-ai/research-eval (new tab)
Release or update evidence
Public release announced by Reka on 2025-08-28; repository created 2025-08-27, within the 2025-01-01–2026-09-26 window.
What it measures
Measures checklist-based answer accuracy for grounded multi-hop web questions, with aggregate mean accuracy and cost per 1,000 requests; it does not measure exhaustive recall or report quality.
Task data
Public encrypted dataset of 374 questions with correctness checklists, constructed through generation, annotation, consensus refinement and filtering. Answering may use live web sources.
Evaluation code
Public Reka dataset and generation, scoring, and analysis scripts; ndurner/web-research-eval is a fork extending provider/model support on the same task suite, not a new benchmark.
Judging criteria
An LLM judge checks each answer against its question-specific checklist, and aggregate accuracy is the resulting score; source quality and citation entailment are not scored.
Reproducibility and barriers
Requires provider credentials, live search and an LLM judge. The Modified MIT license prohibits redistributing decrypted data. Pin fork, model, provider and run settings; author and third-party runs are not automatically comparable.
Who reported the results
Reka launch scores are benchmark-author/vendor results. Nils Durner reports separate third-party extension runs; organizational or financial independence was not established.

Evidence supplied by Exa

46. VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wildinteractive web research/discovery with evolving intent, multi-turn proactive search, and structured evidence enrichment

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
interactive web research/discovery with evolving intent, multi-turn proactive search, and structured evidence enrichment
Primary link
https://github.com/VibeBench/VibeSearchBench (new tab)
Release or update evidence
arXiv paper published/submitted May 2026 (arXiv:2605.27882); GitHub repository public May 20, 2026, within the window.
What it measures
Measures node and knowledge-graph triplet precision, recall, and F1 under average-at-N and best-at-N aggregation, evaluating discovery and structured enrichment.
Task data
Public dataset of 200 manually curated bilingual tasks across 20 domains: 100 professional and 100 daily scenarios, evenly split between Chinese and English, with expert-annotated ground-truth graphs.
Evaluation code
Official repository includes public task JSON, agent implementations, search/visit/Python toolkits, user simulation, evaluation and grading modules, compatible judges, and inference/evaluation scripts.
Judging criteria
Two-phase LLM graph matching handles aliases and translations, then semantic relations; recall allows direct, subsuming, collective, or compositional coverage, while precision counts covered predicted triples.
Reproducibility and barriers
Tasks, code, and evaluator are public, but execution requires LLM/search credentials and substantial live multi-turn inference; provider drift and public ground truth create assessment barriers.
Who reported the results
Paper authors evaluate seven frontier models with ReAct and OpenClaw and report F1 and ablations; no independent reproduction was established.

Evidence supplied by Exa

47. DeepWideSearchdeep-and-wide agentic information-seeking benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
deep-and-wide agentic information-seeking benchmark
Primary link
https://arxiv.org/abs/2510.20168 (new tab)
Release or update evidence
Submitted/published October 2025, within cutoff; this is the canonical DeepWideSearch paper, not the later Table-as-Search paper.
What it measures
Measures structured-table retrieval using exact task success, row/item F1, and core-entity accuracy across repeated runs.
Task data
Public dataset of 220 bilingual English/Chinese questions across 15 domains: 85 Deep2Wide and 135 Wide2Deep, averaging 414.10 information units and 4.21 reasoning depth.
Evaluation code
Public data and evaluation code are available in https://github.com/AIDC-AI/Marco-Search-Agent (new tab) under Marco-DeepResearch-Family/DeepWideSearch/data, eval, and scripts; a public HF artifact also exists.
Judging criteria
Human-verified table ground truth is used; exact success requires all rows, columns, and values, while row/item F1 and core-entity accuracy provide partial scores.
Reproducibility and barriers
Public data and evaluation scripts support reproduction, but live retrieval, multi-run cost, and human annotation remain assessment barriers.
Who reported the results
Authors report four-run system comparisons; no independent reproduction was established.

Evidence supplied by Exa

48. BrowseComp-VLmultimodal web discovery (introduced with WebWatcher)

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
multimodal web discovery (introduced with WebWatcher)
Primary link
https://arxiv.org/abs/2508.05748 (new tab)
Release or update evidence
2025-08: WebWatcher paper introduces BrowseComp-VL; official repository documents benchmark-specific inference/evaluation.
What it measures
Measures multimodal, multi-hop web information seeking through final-answer accuracy/Pass@1, including identification of obfuscated entities from images and textual clues.
Task data
Paper describes 199 level-1 and 200 level-2 image/question pairs. Task-bundle availability is partial: the repository expects JSONL files and separately downloaded images.
Evaluation code
Public WebWatcher repository documents benchmark inference and evaluation scripts; it provides an agent harness rather than a separately packaged benchmark evaluator.
Judging criteria
Scores reference-answer correctness separately for the two difficulty levels; the exact judging configuration was not verified.
Reproducibility and barriers
Assessment barriers include image or OSS download failures, required local evaluation data, live search/image retrieval, model access, and agent dependencies.
Who reported the results
Benchmark-author/Alibaba WebWatcher experiments; distinct from the separately authored MM-BrowseComp and BrowseComp-V³. No independent rerun was established.

Evidence supplied by Exa

49. Deep Research Comparatorhuman evaluation framework for research reports and intermediate steps

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
human evaluation framework for research reports and intermediate steps
Primary link
https://github.com/cxcscmu/Deep-Research-Comparator (new tab)
Release or update evidence
2025-07: paper arXiv:2507.05495; official MIT repository created 2025-07-05.
What it measures
Measures side-by-side report preferences, intermediate-step quality, and text-span feedback; it is not a fixed answer-key benchmark.
Task data
Availability is partial: the paper reports 176 user queries, 17 annotators, and three agents, but public release of the collected annotation corpus was not established and is promised.
Evaluation code
Public frontend, backend, agent-service integration, and configuration instructions establish platform-code availability, not release of the collected study data.
Judging criteria
Human pairwise preferences produce outcome rankings, while up/down votes on intermediate steps and report spans provide process-level feedback.
Reproducibility and barriers
Requires human annotators, Python 3.12, Node.js 18+, PostgreSQL, and agent API/search credentials; live sources and rater variation prevent deterministic replay.
Who reported the results
Authors’ proof-of-concept user study; no independent replication was established. Simple Deepresearch is the accompanying agent scaffold, not a separate benchmark.

Evidence supplied by Exa

50. Cross-Lingual BrowseComp-Plus (XBCP)controlled deep-research retrieval and evidence-grounded answering; multilingual extension distinct from BrowseComp-Plus

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
controlled deep-research retrieval and evidence-grounded answering; multilingual extension distinct from BrowseComp-Plus
Primary link
https://arxiv.org/abs/2606.15345 (new tab)
Release or update evidence
Submitted June 13, 2026 (v2 June 17, 2026), within cutoff; official repository created June 13, 2026.
What it measures
Measures end-to-end answer accuracy, gold-evidence recall, search-call cost, calibration, citation coverage/precision/recall, oracle-retrieval accuracy, and language/retrieval gaps.
Task data
Public availability is established through the repository and linked HF artifacts. It preserves English questions and answers while translating evidence documents into 12 languages, with cross-lingual and multilingual configurations.
Evaluation code
Public MIT repository includes preparation, translation, indexing, agent-running, oracle, LLM-judge, and per-language evaluation scripts; this is actual released code, using GPT-5.4 through OpenRouter.
Judging criteria
LLM judges final-answer correctness; evidence recall is computed against gold documents, citation metrics assess attribution, and oracle settings isolate retrieval from language-mismatch integration.
Reproducibility and barriers
Public code claims a complete pipeline, but reproduction requires decrypted original BrowseComp-Plus material, HF downloads, large indexes, APIs, and live models. The HF viewer reports a schema error.
Who reported the results
Authors report runs across four agents and several retrievers, including translated-evidence accuracy drops and weaker evidence and citation performance; no independent validation identified.

Evidence supplied by Exa

51. K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean ContextsKorean difficult web-discovery and short-answer browsing benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Korean difficult web-discovery and short-answer browsing benchmark
Primary link
https://arxiv.org/abs/2606.02404 (new tab)
Release or update evidence
ArXiv v1 dated June 1, 2026, within cutoff; repository published May 31, 2026 and dataset card is public.
What it measures
Measures Pass@1 answer accuracy and calibration on 300 verified items, separate synthetic-item accuracy, and trajectory/search-call behavior for multi-hop or branching questions.
Task data
Public dataset availability is established. It contains 400 items: 300 Korean-speaker-validated handcrafted problems and 100 synthetic diagnostic items, with answers, URLs, trajectories, checklists and metadata.
Evaluation code
Public official repository includes generation and runtime evaluation code, search-evals harness, dataset loading and fallback JSONL; it uses Perplexity Search API and GPT-5.4-mini extraction.
Judging criteria
Scores single-run Pass@1 against short gold answers and reports calibration; synthetic results remain separate, while trajectories and checklists support diagnostic analysis.
Reproducibility and barriers
Public MIT code and data are documented, but API/model access and changing Korean web results are required. Gitignored seed material limits exact reconstruction of synthetic generation.
Who reported the results
Authors report single-run evaluations using a common Perplexity pipeline and diagnose termination, trajectory, candidate-management and constraint-tracking failures; no independent results identified.

Evidence supplied by Exa

52. DR-Arena: an Automated Evaluation Framework for Deep Research AgentsAutomated dynamic deep-research evaluation of depth and breadth

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Automated dynamic deep-research evaluation of depth and breadth
Primary link
https://aclanthology.org/2026.acl-long.1249/ (new tab)
Release or update evidence
Public arXiv release January 15, 2026; ACL publication July 2–7, 2026, within cutoff.
What it measures
Measures reasoning depth through tree deduction, coverage breadth through aggregation, and pairwise win/Elo; the paper also reports correlation with LMSYS Search Arena.
Task data
Public retained dataset availability is established: 30 evaluation trees are included in the repository. New trees can be generated by crawling current web trends.
Evaluation code
Public GitHub repository includes arena logic, tree generation/crawling, Examiner question generation and judging, scoring/Elo, and the retained 30-tree dataset.
Judging criteria
An automated Examiner builds source-grounded rubrics and judges answers and reports for evidence-based correctness; adaptive evaluation escalates depth or breadth, with reported human validation.
Reproducibility and barriers
The retained trees support fixed comparisons, whereas live-tree generation is time-sensitive. Reproduction requires model, search and API access, and results may change with web or Examiner updates.
Who reported the results
Six-model results and LMSYS correlation are original author and benchmark runs; LMSYS Search Arena supplies the external human-comparison reference.

Evidence supplied by Exa

53. Personalized Deep Research Bench (PDR-Bench), in Towards Personalized Deep Research: Benchmarks and EvaluationsPersonalized deep-research report benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Personalized deep-research report benchmark
Primary link
https://arxiv.org/abs/2509.25106 (new tab)
Release or update evidence
ArXiv v1 submitted September 29, 2025 (revised March 4, 2026), within cutoff; repository and HF dataset are public.
What it measures
Measures personalization alignment across goal, content, presentation and actionability; content quality through depth, insight, coherence and clarity; and factual reliability through accuracy and citation coverage.
Task data
Public data availability is established. It contains 50 tasks, 25 structured or dynamic user profiles and 250 bilingual task–user queries; authors evaluate a 150-query subset.
Evaluation code
Public GitHub repository provides run and evaluation scripts and result directories, while the HF dataset is public; execution requires model and search API keys.
Judging criteria
GPT-5 judges personalization and quality; GPT-5-mini judges reliability. Criteria are dynamically generated, while reliability extracts claims and verifies retrieval and citation support.
Reproducibility and barriers
Queries, data and scripts are public, but proprietary judges, retrieval configuration, API access, dynamic context and live verification create practical reproducibility barriers.
Who reported the results
Comparisons cover commercial and open-source agents, search-augmented models and memory systems, but results are author-run; no independent benchmark results verified.

Evidence supplied by Exa

54. Dr. Bench (formerly Rigorous Bench; Yao et al.)Expert-curated long-form report benchmark

EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED

Link to this record
Category
Expert-curated long-form report benchmark
Primary link
https://arxiv.org/abs/2510.02190 (new tab)
Release or update evidence
Released October 2, 2025 through arXiv and repository; the October 2025 OpenReview submission used the earlier Rigorous Bench title.
What it measures
Measures semantic quality, topical drift, retrieval trustworthiness, contribution per token and retrieval index across 214 expert-curated queries in 10 domains.
Task data
Availability is partial: a public repository exists, but complete downloadable task and reference coverage was not verified. The paper describes manually constructed reference bundles.
Evaluation code
Public repository availability is established, but inspected materials provide an abstract and clone instructions only; a turnkey evaluator was not established, and the paper’s framework is not released scoring code.
Judging criteria
Semantic quality uses query-specific and general rubrics; topical focus penalizes missing or deviating anchor terms; retrieval trustworthiness checks exact and hostname matches against curated links.
Reproducibility and barriers
Explicit formulas and curated references support offline scoring, but report parsing, link extraction and LLM rubric judgments introduce implementation and model sensitivity.
Who reported the results
Thirteen-model comparisons and reported human agreement are authors' experiments; no independent reproduction verified.

Evidence supplied by Exa

30 adjacent, uncertain, or excluded entries — Exa's classifications

The catalogue search above does not filter this separate group.

Tavily search-evals

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/tavily-ai/tavily-search-evals (new tab)

Public provider-comparison evaluation framework (SimpleQA/document relevance) released in the requested period, but not primarily comprehensive list-building or entity enrichment.

WebChoreArena

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://doi.org/10.48550/arxiv.2506.01952 (new tab)

June 2, 2025 benchmark with 532 human-curated tasks in simulated WebArena sites, emphasizing memory, calculation and tedious browser operations; functional task success, not web research or citation-backed synthesis.

AgentRewardBench

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://doi.org/10.48550/arxiv.2504.08942 (new tab)

April 11, 2025 benchmark of 1,302 web-agent trajectories for comparing automatic judges across five existing benchmarks; valuable evaluation-method artifact, but it evaluates judges/trajectories rather than research browsing itself.

WebVoyager updates / Surfer-H evaluation

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://doi.org/10.48550/arxiv.2506.02865 (new tab)

June 3, 2025 paper reports a 92.2% WebVoyager run and introduces WebVoyagerExtended (15,000 synthetic tasks/330 sites), but this is primarily an agent/model paper and browser-action benchmark update, not a new research-centric benchmark.

BrowserArena

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://doi.org/10.48550/arxiv.2510.02418 (new tab)

Adjacent browser-action evaluation. October 2, 2025 paper introduces live user-submitted web-navigation tasks, pairwise comparisons and step-level human feedback. The date is eligible, but navigation success and failure analysis—not research completeness or evidence-grounded enrichment—are its main target.

Perplexity search_evals

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/perplexityai/search_evals (new tab)

Adjacent search-provider evaluation framework, released with the September 25, 2025 Search API report. Public runner, graders and traces reuse SimpleQA, FRAMES, BrowseComp and HLE. The associated provider comparisons are Perplexity/vendor runs, not new benchmarks or independent reproductions.

Parallel Task API DeepSearchQA evaluation

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/parallel-web/parallel-llms-txt/blob/f6b31ffe/public/blog/deepsearch-qa.md (new tab)

2026 vendor report of running Google DeepSearchQA; not a new benchmark. Treat the reported results as Parallel’s runs, not independent reproduction.

FACTS Grounding (v1; later Grounding v2)

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://www.kaggle.com/benchmarks/google/facts-grounding (new tab)

Static long-context grounded-answer evaluation rather than agentic web research; Grounding v2 should not be conflated with FACTS Search.

Online-Mind2Web

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/OSU-NLP-Group/Online-Mind2Web (new tab)

Browser task/action completion on live websites, not primarily multi-source research, exhaustive discovery or evidence-supported enrichment.

FRAMES (v3 / 2025 update)

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://arxiv.org/abs/2409.12941 (new tab)

The original 824-question multi-hop Wikipedia benchmark predates 2025. A January 24, 2025 paper revision alone does not establish a substantive benchmark release. A 2026 community evaluator exists, but its provenance should not be conflated with an official new benchmark.

R2MED

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://arxiv.org/abs/2505.14558 (new tab)

Medical retrieval benchmark with public query/corpus/qrels and strong reasoning-centric design, but primarily a static closed-corpus retrieval benchmark rather than live web/literature research; excluded per scope.

Talc-AI SearchBench

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/Talc-AI/search-bench (new tab)

Substantive public benchmark repository with 900 manually filtered Q&A items, four realistic categories and LLM-as-judge methodology, but its release/results are 2024 (scores as of 2024-08-30; launch 2024-09-19), outside the requested 2025-01-01–2026-09-26 eligibility window. It should not be counted as a qualifying record absent a qualifying 2025–26 update.

WebWatcher

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/Alibaba-NLP/DeepResearch/tree/main/WebAgent/WebWatcher (new tab)

Agent/model project, not a second benchmark record. Its August 2025 paper introduces BrowseComp-VL, catalogued separately; other evaluation runs reuse existing datasets.

WebQuest

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/google-deepmind/webquest (new tab)

Multimodal web-UI/page-sequence QA benchmark with public repository, but primary paper/repository are 2024 (outside eligibility window); static/browser-UI QA adjacent rather than core research/list-building.

DRBench (ServiceNow)

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/ServiceNow/drbench (new tab)

Retained as an adjacent enterprise/internal-search benchmark: it does include public web sources, but evaluates cross-application private synthetic enterprise evidence as a defining requirement.

ARC-Bench

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/aiming-lab/AutoResearchClaw/tree/main/experiments/arc_bench (new tab)

Scientific autonomous experimentation benchmark; research is explicit but web research is not the benchmark’s target.

Entity Enricher platform model benchmarks and benchmark scoring

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://entityenricher.ai/docs/platform/benchmarks (new tab)

Vendor evaluation documentation for organization-specific saved entity/schema scenarios and model comparisons, not an established public dated benchmark release. Reference-based completeness/correctness scoring is described, but public tasks, results and evaluator artifacts were not verified.

Exa Websets benchmark

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://exa.ai/blog/websets-evals (new tab)

Keep adjacent/vendor-only. The official February 19, 2025 post reports Exa's own comparison over 200 generated queries and GPT-4o grading, but no public raw dataset, evaluator or code was established; no independent validation should be implied.

Parallel FindAll 40-query benchmark

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://parallel.ai/products/findall (new tab)

Vendor-only evaluation report: 40 discovery/enrichment queries, with recall measured against pooled correct matches from compared systems. Public task, gold and scorer artifacts and a firm release date were not established. Results are Parallel-created and reported, not independent.

Reka Vibe-Eval

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/reka-ai/reka-vibe-eval (new tab)

Not eligible: multimodal chat benchmark repository was created in 2024 and evaluates image/multimodal generations, not web research; it is a false-positive despite the shared word 'Vibe'.

Vals Web Search Index

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://www.vals.ai/benchmarks/web_search (new tab)

July 16, 2026 controlled search-tool comparison on legal research and finance tasks. It reports 208 legal tasks and 450 finance questions, with rubric-based grading. Keep as a public comparative report: a standalone reusable task/evaluator package and organizational independence were not established; linked orchestration repositories alone do not establish open task data.

EntiWeave

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://huggingface.co/datasets/jxg25/EntiWeave (new tab)

Public graph-grounded web-search training data and a 100-question held-out evaluation split were found, but a qualifying release/update date and benchmark evaluator were not established. Retained as uncertain, not counted as dated core coverage.

X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://aclanthology.org/2025.findings-acl.988/ (new tab)

ACL Findings 2025 multilingual interactive-web benchmark. Evaluates instruction following and product/website interaction using WebShop-style task success, not research discovery or citation-backed synthesis. Public paper verified; data/evaluator release not established here.

Search Arena

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://arxiv.org/abs/2506.05334 (new tab)

2025 public search-LLM preference platform with conversation/vote data and analysis code. Kept adjacent because its primary outcome is pairwise user preference on search chats, rather than objective discovery correctness, evidence coverage or set completeness. Votes were collected by the authors; they are not independent reproduction.

DR³-Eval

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://arxiv.org/abs/2604.14683 (new tab)

April 2026 public code/data benchmark for multimodal, multi-file research reports in static task sandboxes. Scores information recall, factual accuracy, citations, instruction following and depth. Adjacent because supplied files/controlled workspaces, rather than open-web discovery, define its tasks.

Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://arxiv.org/abs/2602.15019 (new tab)

Uncertain standalone benchmark: February 16, 2026 drug-asset scouting paper describes 48 seed queries and 22 held-out query–asset pairs, with asset precision/recall/F1 and expert-calibrated LLM grading. Results are author-run; public task, gold and evaluator releases were not verified. Retained as a paper-defined evaluation study, not a confirmed reusable public benchmark.

GAIA

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://ai.meta.com/research/publications/gaia-a-benchmark-for-general-ai-assistants/ (new tab)

Relevant general-assistant benchmark with web browsing, but the original paper/release is 2023–2024. Later agent runs or leaderboard submissions do not themselves demonstrate a substantive 2025–September 2026 benchmark update; none was established here.

Humanity’s Last Exam (HLE)

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://arxiv.org/abs/2501.14249 (new tab)

Eligible 2025 release, but primarily closed-ended expert academic questions testing broad knowledge/reasoning, not open-web research, complete entity discovery or evidence-supported enrichment. Agent papers using search on HLE are benchmark runs, not new research benchmarks.

SimpleQA Verified

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://arxiv.org/html/2509.07968v1 (new tab)

September 9, 2025 factuality benchmark designed to measure parametric knowledge without tools. Kept separate from agentic web research despite use of related SimpleQA datasets in search-provider evaluations.

SciArena and SciArena-Eval

EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED

https://github.com/yale-nlp/SciArena (new tab)

Public July 1, 2025 release of platform, preference data and analysis code; later paper includes a meta-evaluation benchmark. Adjacent: primarily compares foundation-model literature-grounded responses and judge agreement with human votes, rather than autonomous discovery or complete-set retrieval. Citation-attribution analysis is relevant, and paper-bank data are available, but preferences are not evidence-completeness scores.

Inspect the original materials

The exact request includes the research question, additional instructions, requested output schema, effort: "ultra", and an empty Connect data-source list. The downloadable response preserves Exa's output, grounding, usage, and cost fields. Only the operational top-level run identifier was removed for publication; JSON indentation may differ. No research field was edited.

The compact provenance file records this redaction and SHA-256 hashes of the public files. These hashes help check file identity; they do not establish factual correctness. Future editorial corrections should be distinguished from the preserved provider response.

Exact research question sent to Exa

ASSISTANT-AUTHORED ASSIGNMENT

Conduct a standalone evidence-backed catalogue of public benchmarks and evaluation frameworks testing AI agents on web research, comprehensive list-building, or evidence-supported entity enrichment. Include projects with a public release or substantive update between January 1, 2025 and September 26, 2026. Seek broad coverage of this defined scope, not an arbitrary top ten. Distinguish research/discovery evaluations from adjacent browser-action or generic QA benchmarks; put adjacent or uncertain cases separately. For each qualifying project identify canonical name, primary project/paper/repository URLs, what ability it measures, dated evidence of eligibility, public availability of task data, evaluation code and judging criteria, and practical reproducibility barriers (dependencies, unavailable components, changing web state). Distinguish benchmark-author, vendor, and independent reported results without inventing independence. Support material field claims with primary-source URLs and short evidence notes. Unknown is acceptable; avoid inferring released code from a paper's promise. Deduplicate renamed projects and distinguish benchmarks from reports of running them. Return concise structured records and a short synthesis of coverage, limitations, and major exclusions. Do not run benchmark code or claim reproduced evaluations. Use only public web research; do not enable or use Exa Connect, premium connected datasets, contact enrichment, or purchases. This is one Ultra run under the documented default US$20 run cap; do not create follow-up runs. Keep the assessment neutral across vendors. No personal or private project information is relevant.
Additional instructions sent to Exa

ASSISTANT-AUTHORED INSTRUCTIONS

Research using public web sources only. No connected datasets or contact enrichment. Prioritize primary papers, repositories, official benchmark sites and documented artifacts. Be concise and evidence-specific. Distinguish observed source facts from your judgments. Do not equate broad search with proven exhaustive coverage. Provide factual findings without promotional framing.
Requested output structure

ASSISTANT-DESIGNED SCHEMA

{
  "type": "object",
  "properties": {
    "scope": {
      "type": "string"
    },
    "benchmarks": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "name": {
            "type": "string"
          },
          "canonical_url": {
            "type": "string"
          },
          "category": {
            "type": "string"
          },
          "eligibility_date_and_evidence": {
            "type": "string"
          },
          "measures": {
            "type": "string"
          },
          "task_data": {
            "type": "string"
          },
          "evaluation_code": {
            "type": "string"
          },
          "judging_criteria": {
            "type": "string"
          },
          "reproducibility_and_barriers": {
            "type": "string"
          },
          "reported_results_provenance": {
            "type": "string"
          },
          "evidence": {
            "type": "array",
            "items": {
              "type": "object",
              "properties": {
                "url": {
                  "type": "string"
                },
                "supports": {
                  "type": "string"
                }
              },
              "required": [
                "url",
                "supports"
              ]
            }
          }
        },
        "required": [
          "name",
          "canonical_url",
          "category",
          "eligibility_date_and_evidence",
          "measures",
          "task_data",
          "evaluation_code",
          "judging_criteria",
          "reproducibility_and_barriers",
          "reported_results_provenance",
          "evidence"
        ]
      }
    },
    "adjacent_or_uncertain": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "name": {
            "type": "string"
          },
          "url": {
            "type": "string"
          },
          "reason": {
            "type": "string"
          }
        },
        "required": [
          "name",
          "url",
          "reason"
        ]
      }
    },
    "summary": {
      "type": "string"
    },
    "coverage_gaps": {
      "type": "array",
      "items": {
        "type": "string"
      }
    }
  },
  "required": [
    "scope",
    "benchmarks",
    "adjacent_or_uncertain",
    "summary",
    "coverage_gaps"
  ]
}

What to take away

Use this example to inspect how delegated research works, not to choose a “winning” model. The assignment shapes the investigation; the provider makes research judgments; the assistant makes further presentation choices. Keeping those layers visible gives readers a basis for their own assessment.