A worked example of delegated research—and the judgments around it.
An assistant designed the assignment, Exa Ultra conducted the research, and the assistant organized the result. This page lets you inspect each contribution separately. You can read the explanation in a few minutes or explore the complete output below.
What happened
On September 26, 2026, one Exa Agent Ultra run investigated public benchmarks and evaluation frameworks for research agents. In plain language: how do people test whether an AI can find difficult answers, build complete lists, or produce a well-supported research report?
The run was restricted to public web research, with no Connect datasets requested and no follow-up run. It used Exa's documented default $20 run cap. The connected interface did not expose a custom budget field; the spending instruction in the prompt was not itself a technical limit.
This is an unaudited research output, not an authoritative benchmark directory or a comparison of research systems. No benchmark code was executed. The catalogue's underlying claims were not independently checked by the presenting assistant.
The service returned a completed status, reported a cost of $11.6589 and 92 searches, and supplied 54 records it classified as qualifying plus 30 adjacent, uncertain, or excluded records. These are reported activity and output counts—not measures of correctness. The returned response did not expose a stop reason or timestamps sufficient to establish run duration. It also does not contain a complete log of all 92 search queries.
Who contributed what
- Assistant: design the assignment. The assistant translated the research objective into a question, additional instructions, and a structured output schema. It selected a January 2025–September 2026 eligibility window and emphasized public artifacts, reproducibility, evidence, and results provenance.
- Exa: investigate and generate findings. Ultra retrieved sources, selected and classified candidates, assessed artifacts, and wrote the summary and records. Those decisions and claims belong to the provider-generated output.
- Assistant: package and assess. The assistant created this page, selected the three teaching examples below, and wrote the interpretation and limitations outside the labelled Exa sections. That editorial layer influences what readers notice.
The assistant that helped frame the assignment also prepared this presentation. Its assessment is not an independent review. The original Exa response was already shaped by the assignment, so inspecting the prompt matters as much as inspecting the answer. Neither the raw response nor this presentation is free of judgment.
Three examples of different research abilities
The examples below are selected and paraphrased by the assistant from Exa's records to illustrate different abilities. They are not a ranking, independently verified findings, or a representative sample of the whole catalogue.
- Finding a difficult answer — BrowseComp. Exa describes questions requiring persistent browsing, with short answers scored for correctness. That tests a different ability from compiling a complete list. Open Exa's full record.
- Finding and enriching a set — WideSearch. Exa describes table-based tasks with entity, attribute, row, and table-level evaluation. Missing entities and incomplete fields matter here. Open Exa's full record.
- Producing a supported report — DeepResearch Bench (RACE + FACT). Exa describes separate assessment of report quality and citation support. A readable, comprehensive report is not automatically well evidenced. Open Exa's full record.
These examples suggest a useful question when reading a research-agent score: what ability did the evaluation actually test? The catalogue supplies leads for answering that question; its citations still need inspection before relying on a specific claim.
Assistant assessment
The useful deliverable is a starting reference: named evaluations, descriptions of what they measure, links to artifacts, and claimed barriers to reproduction. Its structure makes follow-up investigation easier than a bare list of search results.
The assignment also creates blind spots. It emphasizes public availability and reproducibility rather than giving equal weight to beginner accessibility, evaluation cost, multilingual representation, or agreement with real users' judgments. The chosen date window excludes older work unless substantively updated. These are framing choices, not neutral properties of the field.
Exa's scope statement includes controlled web-derived corpora and some literature or report-evaluation tasks. Readers may reasonably draw the boundary of “research-agent evaluation” differently. Preserve that scope statement when interpreting the count of 54; the count is Exa's classification, not a settled total for the field.
The result does not establish that every entry is accurate, that coverage is exhaustive, or that Ultra found more useful material than a cheaper Agent setting or assistant-directed search would have found. No comparison run was performed. More searches, more records, and polished explanations are not substitutes for evidence quality.
The review for this publication checked record counts, faithful rendering, links within this site, and removal of the operational run identifier. It did not audit the external citations or reproduce benchmark results. Exa's own descriptions of what it inspected remain provider claims.
Read Exa's output
The following summary, scope statement, limitations, and catalogue fields come from the saved Exa response. Field wording is preserved; labels, layout, search, and navigation were added by the assistant. Links are evidence supplied by Exa, not an endorsement or a claim that the presenting assistant checked them.
EXA OUTPUT · WORDING PRESERVED
Exa's summary
The catalogue contains 54 qualifying benchmarks/frameworks and 30 adjacent, uncertain or excluded entries. Three distinct capabilities recur: finding a difficult answer, discovering and enriching an entire set, and producing a supported research report. Their scores are not interchangeable: WideSearch (new tab) evaluates tables and set coverage, while DeepResearch Bench (new tab) separates report quality from citation support. Public availability ranges from released tasks and evaluators to encrypted, partially public or paper-only protocols. BrowseComp-Plus (new tab) controls web-state variation with a fixed corpus; live-web benchmarks still depend on retrieval, provider and judge versions. Most inspected results are benchmark-author runs, often also vendor runs. Reka’s third-party extension is identified separately without assuming organizational independence. Renames, forks and run reports are not counted as new benchmarks. No benchmark code was executed and no reported result was reproduced.
Exa's scope interpretation
Public benchmarks and evaluation frameworks with a documented release or substantive update from January 1, 2025 through September 26, 2026. Includes difficult web discovery, comprehensive entity/list retrieval, evidence-supported enrichment, literature discovery and research-report evaluation. Controlled web-derived corpora are included when they test research-agent workflows. A public paper or protocol can qualify even when tasks or scoring code are withheld; availability is assessed separately. Browser-action tasks, generic QA, preference-only search chat, internal-enterprise tasks and vendor run reports are separated. This is broad documented coverage, not a claim of exhaustiveness.
Exa's stated coverage limitations
- Coverage is broad but not proven exhaustive. Newly released, poorly indexed, non-English or privately distributed projects may be missing; the catalogue does not claim every qualifying project exists here.
- Artifact inspection establishes what was publicly documented or accessible, not that an end-to-end evaluation succeeds. “Not verified” means unresolved in the inspected sources, not proof that an artifact does not exist.
- Release dates and benchmark versions matter. Paper revisions, corrected answers, expanded datasets and changing repositories can yield different task counts; later vendor runs alone do not establish a new benchmark release.
- Complete-list scoring has different denominators: fixed expert gold, pooled discovered matches or a requested output quota. Recall and F1 therefore do not by themselves demonstrate exhaustive real-world coverage.
- Encrypted answers, locked test sets, missing corpora, unavailable proprietary components, paid API dependencies and model/judge drift can prevent exact independent reproduction even when some code is public.
- Public human preference or citation-quality evaluation is not automatically a test of autonomous discovery. Search Arena and SciArena are separated from agentic discovery benchmarks; broad science suites are included only for their relevant research components.
- Independence was not inferred from a third-party name, leaderboard entry or open repository. Commercial comparisons remain attributed to their authors; no cross-benchmark vendor ranking is warranted.
Explore Exa's 54 catalogue records
These are the entries Exa classified as qualifying. Search covers all record fields; it does not score or validate them.
54 of 54 catalogue records
1. BrowseCompdifficult web discovery / short-answer research
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- difficult web discovery / short-answer research
- Primary link
- https://openai.com/index/browsecomp/ (new tab)
- Release or update evidence
- 2025-04-10: OpenAI publicly announced and open-sourced the benchmark.
- What it measures
- Tests persistent browsing and reasoning for obscure, entangled facts across 1,266 questions with short, verifiable answers; it is not an exhaustive-list or report-quality test.
- Task data
- Public test CSV is referenced by the evaluator; questions and answers are obfuscated using canary-based decoding.
- Evaluation code
- Public scripts in openai/simple-evals load and decode tasks, then invoke an LLM grader.
- Judging criteria
- An LLM judges reference-answer equivalence, producing aggregate correctness or accuracy rather than source-quality or list-completeness scores.
- Reproducibility and barriers
- Requires an agent/search implementation and grader access; live-web changes, model settings, and browsing effort affect results. Public answers create leakage risk, and authors request that decoded examples not be republished.
- Who reported the results
- Launch results are benchmark-author and vendor-reported by OpenAI, not independently reproduced.
Evidence supplied by Exa
- https://openai.com/index/browsecomp/ (new tab)
Dated release, task construction, size, original model comparisons and leakage warning.
- https://github.com/openai/simple-evals/blob/main/browsecomp_eval.py (new tab)
Observed task-loading, decryption and model-grading implementation.
- https://arxiv.org/abs/2504.12516 (new tab)
Primary paper linked from the official announcement.
2. WideSearch: Benchmarking Agentic Broad Info-Seekingcomprehensive list enumeration + evidence-supported entity enrichment
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- comprehensive list enumeration + evidence-supported entity enrichment
- Primary link
- https://arxiv.org/abs/2508.07999 (new tab)
- Release or update evidence
- Released August 11, 2025 (arXiv 2508.07999; paper version dated August 28/September 5 in rendered materials); the GitHub README announces release on 2025/08/11.
- What it measures
- Measures item, column, and row precision/recall; table/task success requires complete, accurate atomic information. Max@N and human comparison are also reported.
- Task data
- Public 200-task dataset: 100 English and 100 Chinese tasks across 18 domains. Agents fill predefined tables using large-scale atomic facts about multiple entities; curation included exhaustive human gold research and automated-versus-human validation.
- Evaluation code
- Public MIT repository contains evaluation scripts, documentation, and agent/tool code; the public Hugging Face dataset and project page link to data and code.
- Judging criteria
- Entity/item precision-recall/F1, attribute correctness and exact complete-row/table correctness; the automated evaluator was compared with expert ratings.
- Reproducibility and barriers
- Code and dataset are public, but live-web execution requires API credentials and results may drift with web, model, and API changes. Commercial-system tests used web interfaces; this is an assessment of barriers.
- Who reported the results
- Original authors report benchmark runs across 10+ agentic systems and human tests; results are author-reported, not independently validated.
Evidence supplied by Exa
- https://arxiv.org/html/2508.07999 (new tab)
Primary paper: benchmark purpose, 200 bilingual tasks, curation, metrics, tools, judging and reported runs.
- https://github.com/ByteDance-Seed/WideSearch (new tab)
Primary repository: 2025/08/11 release notice, evaluation code, instructions and dataset/project links.
- https://widesearch-seed.github.io/ (new tab)
Official project page: full benchmark dataset/ground-truth tables download and evaluation-code access.
- https://huggingface.co/datasets/ByteDance-Seed/WideSearch (new tab)
Official dataset artifact linked by the project/repository.
3. Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agentscomprehensive list enumeration + evidence-supported entity enrichment
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- comprehensive list enumeration + evidence-supported entity enrichment
- Primary link
- https://arxiv.org/abs/2606.27595 (new tab)
- Release or update evidence
- 2026-06: arXiv release 2606.27595; official project, repository and dataset are public.
- What it measures
- Measures membership Item-F1, attribute-cell Column-F1, complete-row Row-F1, table success, and parse rate using normalization-aware comparison.
- Task data
- Public schema/metadata with gated/encrypted answer fields: 228 Korean tables cover 190 parent entities and 16 categories, including 4,262 gold rows and 14,560 attribute cells. Easy/Medium/Hard tiers vary table width and composite-key membership; 201 tables require cross-source attribute lookups.
- Evaluation code
- Public MIT-licensed pipeline and scorer are available in the official repository; answer fields remain encrypted or gated by request to reduce leakage.
- Judging criteria
- Scores membership precision/recall, per-column cell correctness, strict full-row correctness, and whole-table success; the comparator normalizes name variants, date granularity, and numeric formatting.
- Reproducibility and barriers
- Scorer and pipeline are public, but evaluation answers are gated/encrypted and live-web tasks are dynamic; API/model access and request-based data release remain practical barriers (assessment).
- Who reported the results
- Original authors report a 20-agent evaluation; these are benchmark-author runs, not independently validated results.
Evidence supplied by Exa
- https://arxiv.org/html/2606.27595 (new tab)
Primary paper: task construction, 228-table composition, tiers, cross-source enrichment, metrics, comparator, gating and author results.
- https://github.com/minstar/Ko-widesearch (new tab)
Official code repository: open pipeline/scorer, project citation and dataset/release information.
- https://minstar.github.io/Ko-widesearch/ (new tab)
Official project page: benchmark scope, metrics, table counts, categories and results.
- https://huggingface.co/datasets/Minbyul/Ko-widesearch (new tab)
Official dataset card: 228-task schema, encrypted question/answer fields, tiers and scoring format.
4. EnterList (introduced in WebLists)Structured list extraction from interactive websites; paper-defined benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Structured list extraction from interactive websites; paper-defined benchmark
- Primary link
- https://arxiv.org/html/2504.12682 (new tab)
- Release or update evidence
- ArXiv record dated 2025-04-17, within the eligibility window.
- What it measures
- Precision and recall for rows extracted from live interactive websites, with cost per output row; narrower than unconstrained open-web enumeration.
- Task data
- Availability not established as a public release: the paper defines 200 live tasks across four use cases and 50 websites, with annotated reference URLs and extraction scripts, capped at five pages per task.
- Evaluation code
- The paper describes reference extraction scripts and methodology, but a public repository or downloadable evaluator/data package was not established; treat it as paper-only.
- Judging criteria
- Exact field matching against refreshed reference extraction defines precision as retrieved gold rows and recall as retrieved gold coverage; live-site ground truth changes over time.
- Reproducibility and barriers
- Live websites and changing ground truth hinder reproduction; no verified public task, data, or evaluator release was established.
- Who reported the results
- Original authors report experiments including BardeenAgent results; no independent reproduction was established.
Evidence supplied by Exa
- https://arxiv.org/html/2504.12682 (new tab)
Primary paper: date, 200 tasks/50 sites, reference-script design, live-ground-truth caveat, metrics and results.
- https://doi.org/10.48550/arxiv.2504.12682 (new tab)
Primary DOI/arXiv record confirming 2025-04-17 publication and benchmark definition.
5. DeepResearch Bench (Ayanami0730 / Mingxuan Du; RACE + FACT)long-form deep-research report benchmark; citation/retrieval evaluation
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- long-form deep-research report benchmark; citation/retrieval evaluation
- Primary link
- https://github.com/Ayanami0730/deep_research_bench (new tab)
- Release or update evidence
- 2025-06: paper arXiv:2506.11763 and public repository (created June 13) introduce RACE/FACT evaluation. Distinct from FutureSearch’s similarly named benchmark.
- What it measures
- RACE measures report quality through comprehensiveness, depth, instruction following, and readability; FACT measures citation accuracy and effective citation count.
- Task data
- Public repository materials include 100 PhD-level tasks, 50 Chinese and 50 English, across 22 fields, authored or refined by 100+ domain experts. It includes task, reference, report, and result materials, with raw articles and scores linked on the leaderboard.
- Evaluation code
- Public repository contains RACE and FACT evaluation code, prompts, configurations, and result directories; execution requires model/API and web-content retrieval services.
- Judging criteria
- RACE uses LLM-generated task criteria and weights to compare reports with high-quality references, with human-consistency validation. FACT extracts statement-URL pairs, retrieves cited pages, and LLM-judges support.
- Reproducibility and barriers
- Code and data are publicly downloadable under Apache-2.0, but reproducing scores requires paid or credentialed judge models and web retrieval; live URLs may change. These are assessed barriers.
- Who reported the results
- Authors report evaluations of commercial deep-research agents and search-enabled LLMs; linked raw articles and scores remain author-run, not independent results.
Evidence supplied by Exa
- https://arxiv.org/abs/2506.11763 (new tab)
Paper identity, 100 tasks, RACE/FACT methodology, release link and submission date.
- https://github.com/Ayanami0730/deep_research_bench (new tab)
Official repository, code/data/result layout, update history and evaluation instructions.
- https://deepresearch-bench.github.io/ (new tab)
Official project description, task composition and author-reported leaderboard/results.
6. DeepResearch Bench II (imlrz / Ruizhe Li et al.)long-form deep-research report benchmark; expert-rubric diagnosis
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- long-form deep-research report benchmark; expert-rubric diagnosis
- Primary link
- https://github.com/imlrz/DeepResearch-Bench-II (new tab)
- Release or update evidence
- Official README records a November 2025 pipeline release; arXiv:2601.08536 was submitted 2026-01-13. It is explicitly a follow-up to DeepResearch Bench.
- What it measures
- Scores binary satisfaction of information-recall, analysis and presentation criteria across 9,430 fine-grained rubrics and 132 tasks.
- Task data
- Public tasks_and_rubrics.jsonl contains 132 tasks and 9,430 criteria across 22 domains. A February 24, 2026 Hugging Face release adds selected model-generated reports. May 2026 metadata assigns each task its source article’s license.
- Evaluation code
- Official Apache-2.0 repository provides rubrics, scripts and a batched evaluator for PDF, DOCX, image and text reports. Running requires Gemini or other model credentials.
- Judging criteria
- LLM judges assess binary satisfaction of atomic information-recall, analysis and presentation criteria. Criteria were extracted from expert reports and human-reviewed.
- Reproducibility and barriers
- Benchmark, rubrics and code are public; assessment: judge-model access, processing cost, evaluator versions and report formatting affect reproducibility.
- Who reported the results
- The paper reports author evaluations of several state-of-the-art agents; no independent result provenance is established here.
Evidence supplied by Exa
- https://arxiv.org/abs/2601.08536 (new tab)
Paper identity/date, 132 tasks, 9,430 rubrics, dimensions, human review and release claims.
- https://github.com/imlrz/DeepResearch-Bench-II (new tab)
Official repo, November 2025 release, Apache-2.0 status and evaluator artifacts.
- https://arxiv.org/html/2601.08536 (new tab)
Benchmark abstract and rubric construction/evaluation description.
7. DeepResearch-ReportEval (HKUDS/DeepResearch-Eval)long-form research-report evaluation framework/dataset
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- long-form research-report evaluation framework/dataset
- Primary link
- https://github.com/HKUDS/DeepResearch-Eval (new tab)
- Release or update evidence
- Paper arXiv:2510.07861 was published 2025-10-09, and the official repository identifies the framework and dataset; it qualifies within the window.
- What it measures
- Scores comprehensiveness, coherence, clarity, insightfulness and overall quality, plus paragraph redundancy and citation-supported factuality.
- Task data
- 100 queries cover 12 real-world categories and include 100 Qwen-DeepResearch reports collected in early September 2025; repository data and examples are public.
- Evaluation code
- Repository includes judge_score.py, judge_fact.py, utilities, prompts and examples. Fact checking uses Firecrawl or Jina Reader with -1/0/1 support labels.
- Judging criteria
- LLM judges score report quality and paragraph redundancy; a citation checker evaluates each claim as unsupported, uncertain or supported using retrieved source pages.
- Reproducibility and barriers
- MIT code and data are public; assessment: judge-model and Firecrawl/Jina credentials, web access and model selection are practical barriers. The dataset is a fixed snapshot.
- Who reported the results
- Paper comparisons cover four commercial systems and author-generated Qwen reports; no independent reproduction is identified here.
Evidence supplied by Exa
- https://github.com/HKUDS/DeepResearch-Eval (new tab)
Official scripts, data structure, dimensions, fact labels and dataset description.
- https://doi.org/10.48550/arxiv.2510.07861 (new tab)
Paper date, framework identity, 100-query dataset, methodology and author-reported results.
- https://arxiv.org/pdf/2510.07861 (new tab)
Primary source for quality, redundancy, factuality and expert-alignment claims.
8. ResearcherBenchScientific web research and evidence-grounded report generation
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Scientific web research and evidence-grounded report generation
- Primary link
- https://github.com/GAIR-NLP/ResearcherBench (new tab)
- Release or update evidence
- Paper arXiv:2507.16280 was published 2025-07-22; the official site and repository are public within the window. It evaluates frontier-AI scientific questions rather than broad PhD-level web research.
- What it measures
- Scores weighted expert-insight coverage, citation-support accuracy (Faithfulness) and citation coverage (Groundedness).
- Task data
- Public repository provides 65 expert-curated questions across 35 AI subjects, with weighted reference insights; tasks include technical questions, literature reviews and research consulting.
- Evaluation code
- Official repository provides the benchmark platform, formats and eval.sh; users submit model responses and receive rubric and factuality results. OpenAI and Jina credentials are required.
- Judging criteria
- Researchers create weighted 1–3 insight criteria, with Claude-3.7-Sonnet assisting extraction; a factuality pipeline extracts claims and URLs, then judges source support.
- Reproducibility and barriers
- Questions and framework are public; assessment: API credentials, web retrieval and expert rubric interpretation are barriers. Baselines are author-run and dated March–June 2025.
- Who reported the results
- Official paper and site report author evaluations of OpenAI, Gemini, Grok, Perplexity and other baselines; no independent runs are established.
Evidence supplied by Exa
- https://arxiv.org/abs/2507.16280 (new tab)
Paper identity/date, 65 questions, 35 subjects, metrics and release.
- https://github.com/GAIR-NLP/ResearcherBench (new tab)
Official code, data format, evaluation workflow and API requirements.
- https://researcherbench.github.io/ (new tab)
Task construction, methodology, leaderboard provenance and evaluation timing.
9. ResearchRubricsDeep-research report evaluation / human-authored rubrics
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Deep-research report evaluation / human-authored rubrics
- Primary link
- https://github.com/scaleapi/researchrubrics (new tab)
- Release or update evidence
- Paper/preprint arXiv:2511.07685 was published 2025-11-10 and the repository was created 2025-11-08; both fall within the window. Repository and Hugging Face data are public.
- What it measures
- Measures weighted criterion compliance across 101 prompts and 2,593 expert-written criteria, plus agreement between human and model judgments.
- Task data
- Public ScaleAI/researchrubrics download is documented in the repository: 101 human-written prompts and 2,593 reviewed criteria across nine domains, including penalty criteria.
- Evaluation code
- Actual MIT-licensed code ingests Markdown reports, chunks them, performs batch evaluation and calculates compliance. Default Gemini 2.5 Pro use through LiteLLM requires an API key and ScaleAI Hugging Face data.
- Judging criteria
- Human criteria cover requirements, reasoning, synthesis, references and communication. The released evaluator assigns Satisfied/Not Satisfied scores and computes weighted compliance; negative-weight rubrics are excluded from the denominator.
- Reproducibility and barriers
- Prompts, rubrics and code are public; external model access, cost and latency remain barriers. Commercial reports from the study may not be reproducible from the repository.
- Who reported the results
- Scale AI authors evaluated OpenAI, Gemini and Perplexity Deep Research; no independent leaderboard or result is found in the checked primary artifacts.
Evidence supplied by Exa
- https://arxiv.org/html/2511.07685 (new tab)
Benchmark size, human-authored rubrics, dimensions, systems and reported compliance.
- https://github.com/scaleapi/researchrubrics (new tab)
Public MIT code, data expectations and executable evaluation pipeline.
10. LiveResearchBench + DeepEvalDynamic web research benchmark and long-form report evaluator
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Dynamic web research benchmark and long-form report evaluator
- Primary link
- https://github.com/SalesforceAIResearch/LiveResearchBench (new tab)
- Release or update evidence
- Paper arXiv:2510.14240 was published in October 2025; the GitHub repository is dated 2025-10-17, and an ICLR 2026 artifact is public.
- What it measures
- Scores coverage, presentation, citation accuracy, citation traceability, fact-and-logic consistency and analysis depth using checklist, pointwise, pairwise and ensemble protocols.
- Task data
- Public static and realtime Hugging Face datasets: 100 expert-curated tasks across seven domains and ten categories, with checklists and dynamic date placeholders.
- Evaluation code
- Actual Salesforce code provides benchmark loading, DeepEval protocols, report handling and result output; the Hugging Face dataset is public. Enterprise-Deep-Research documents generation and invocation.
- Judging criteria
- Human checklists score coverage and presentation; rubric trees assess citation correctness and association; consistency and pairwise protocols assess factual logic and comparative depth.
- Reproducibility and barriers
- Tasks, code and static data are public; assessment: realtime evaluation depends on live web content, while generation and LLM judging require provider APIs.
- Who reported the results
- Salesforce-led evaluation covers 17 frontier systems; public leaderboard values are author/vendor results, not independent evaluations.
Evidence supplied by Exa
- https://github.com/SalesforceAIResearch/LiveResearchBench (new tab)
Public benchmark and DeepEval code, protocols, tasks and outputs.
- https://arxiv.org/abs/2510.14240 (new tab)
Release date, design, 100 tasks, research effort and 17-system evaluation.
- https://huggingface.co/datasets/Salesforce/LiveResearchBench (new tab)
Public static/realtime datasets and date-placeholder behavior.
- https://livedeepresearch.github.io/ (new tab)
Public leaderboard and reported system-level scores.
11. ReportBenchAcademic survey-grounded report quality and citation evaluation
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Academic survey-grounded report quality and citation evaluation
- Primary link
- https://github.com/ByteDance-BandAI/ReportBench (new tab)
- Release or update evidence
- arXiv:2508.15804 posted August 2025; GitHub repository public by the 2025 paper release, within window.
- What it measures
- Measures citation precision, recall, citation-match rate, reference counts, cited-statement coverage, and non-cited factual accuracy using statement-level verification.
- Task data
- Public ReportBench_v1.1.jsonl supplies 100 survey-grounded research tasks across ten domains and three prompt granularities; survey references furnish citation gold.
- Evaluation code
- Public repository contains processing, retrieval, citation and statement evaluators, metric calculators, benchmark data directories, and JSON-output support. Commercial web-product collection requires browser captures and dedicated processors.
- Judging criteria
- Published survey references are gold standards; cited claims are matched to retrieved passages, while non-cited claims use web-connected Gemini majority voting.
- Reproducibility and barriers
- Code and data are public, but web retrieval, browser capture, paid LLMs, and paid search services are required. Survey overlap may penalize valid divergent research; this is a methodological limitation assessment.
- Who reported the results
- ByteDance BandAI author-run results report precision, recall, citation matching, and non-cited accuracy for OpenAI and Gemini Deep Research. No independent result was located in checked primary artifacts.
Evidence supplied by Exa
- https://github.com/ByteDance-BandAI/ReportBench (new tab)
Actual public code/data structure, evaluation workflow and full reported result table.
- https://openreview.net/forum?id=zvL42fmtbG (new tab)
Public conference record and abstract confirming benchmark purpose and public-release status.
12. FACTS Search (FACTS Benchmark Suite)Adjacent web-search factuality benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Adjacent web-search factuality benchmark
- Primary link
- https://www.kaggle.com/benchmarks/google/facts-search/leaderboard (new tab)
- Release or update evidence
- FACTS Benchmark Suite and Search benchmark launched December 2025 through the official announcement and arXiv:2512.10791; leaderboard remained public through 2026.
- What it measures
- Measures search-enabled answer F1, overall and attempted accuracy, hedging rate, and search count. A prompted auto-rater scores answer correctness.
- Task data
- 1,884 Search questions comprise 890 public and 994 private items, including human-written hard-tail questions and three synthetic multi-hop subsets. The task evaluates factual search answers, not reports or list-building.
- Evaluation code
- Public benchmark examples and hosted Kaggle evaluation are available, but exact leaderboard reproduction requires the standardized Brave API, private set, hosted judge, and LLM setup.
- Judging criteria
- A prompted auto-rater compares answers with gold answers; F1 balances accuracy and attempted accuracy, while hedging and search count are diagnostics.
- Reproducibility and barriers
- Public examples and a common search API improve comparability, but the private set, API costs, availability, and auto-rater introduce barriers and possible bias. This is adjacent because it evaluates factual search answers.
- Who reported the results
- Google DeepMind, Google Research, and Kaggle author-hosted evaluation reports results for Gemini 3 Pro, GPT-5, and Claude 4.5 Opus. No independent result was counted.
Evidence supplied by Exa
- https://deepmind.google/blog/facts-benchmark-suite-systematically-evaluating-the-factuality-of-large-language-models/ (new tab)
Official December 2025 release, Search purpose, public/private sizes and suite leaderboard.
- https://arxiv.org/abs/2512.10791 (new tab)
Search dataset composition, Brave API, metrics and reported results.
- https://www.kaggle.com/benchmarks/google/facts-search/leaderboard (new tab)
Public live Search leaderboard and operational evaluation description.
13. BrowseComp-Plusfixed-corpus deep-research benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- fixed-corpus deep-research benchmark
- Primary link
- https://github.com/texttron/BrowseComp-Plus (new tab)
- Release or update evidence
- Public paper and repository released 2025-08-08 (arXiv:2508.06600; GitHub created/updated in 2025), within window.
- What it measures
- Measures answer accuracy, evidence-document recall, search-call count, calibration error, retrieval Recall@k, and nDCG@10.
- Task data
- Public fixed corpus, queries and qrels: 830 BrowseComp-derived questions over approximately 100,000 web documents, with human-verified evidence and hard negatives. This is a controlled web-derived corpus, not live search.
- Evaluation code
- Public official repository provides agent and retriever scripts, evaluation scripts, TREC-style qrels, reproduction documentation, and leaderboard instructions. Main evaluation uses top-five retrieval with a 512-token context limit.
- Judging criteria
- GPT-4.1 judges final-answer correctness; labeled evidence and gold documents determine retrieval Recall and nDCG. Citation accuracy is analyzed separately.
- Reproducibility and barriers
- The static corpus and released qrels improve reproducibility, while model/API access and substantial compute remain barriers. Dataset and retrieval artifacts are public through repository and Hugging Face links.
- Who reported the results
- Paper and official project results are author-reported comparisons of retrieval and model systems, including Search-R1, GPT-5, and GPT-5 with Qwen3 embeddings; no independent results are reported.
Evidence supplied by Exa
- https://arxiv.org/html/2508.06600v1 (new tab)
Paper date, fixed curated corpus, 830 queries, metrics, and reported results.
- https://github.com/texttron/BrowseComp-Plus (new tab)
Official code, reproduction scripts, qrels/evaluation instructions, and static-corpus description.
- https://texttron.github.io/BrowseComp-Plus/ (new tab)
Official project description and metric definitions.
14. BrowseComp-ZHChinese live-web multi-hop browsing benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Chinese live-web multi-hop browsing benchmark
- Primary link
- https://github.com/PALIN2018/BrowseComp-ZH (new tab)
- Release or update evidence
- 2025-04-24: original BrowseComp-ZH paper and repository. A separate AGI-Eval correction fork appeared in January 2026.
- What it measures
- Measures answer accuracy and calibration error for standalone LLMs and browsing agents, including model, reasoning, and browsing comparisons.
- Task data
- Public encrypted dataset of 289 Chinese multi-hop questions across 11 domains; repository decoding is required. The January 2026 AGI-Eval correction fork claims 24 answer corrections, so versions are not interchangeable.
- Evaluation code
- Public official repository includes decryption, model-evaluation, prediction/result, and calibration workflows. OpenCompass integration loads the dataset and uses an LLM judge.
- Judging criteria
- An LLM scorer judges answer correctness; the OpenCompass implementation explicitly uses LLMJudgeScorer. Official results report accuracy and calibration error.
- Reproducibility and barriers
- Encrypted questions and answers require the repository decryption procedure; live Chinese web access and proprietary model APIs are practical barriers. Encryption limits direct dataset inspection.
- Who reported the results
- Official author-run comparisons cover more than 20 systems, including OpenAI DeepResearch, O1, and Gemini-2.5-Pro. The reported figures are benchmark-author results.
Evidence supplied by Exa
- https://arxiv.org/html/2504.19314 (new tab)
Paper date, 289-question design, domains, construction, live-web task and reported evaluations.
- https://github.com/PALIN2018/BrowseComp-ZH (new tab)
Official dataset/code, encrypted-data procedure, validation and evaluation README.
- https://github.com/open-compass/AgentCompass/blob/a7c30989/src/agentcompass/benchmarks/browsecomp_zh.py (new tab)
Public evaluator integration showing dataset loading and LLM judge scoring.
- https://github.com/AGI-Eval-Official/BrowseComp-ZH-revised (new tab)
Third-party correction fork, January 2026 creation, 24 corrections and decoding instructions; not official benchmark-author validation.
15. WebWalkerQAweb traversal / information-seeking QA benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- web traversal / information-seeking QA benchmark
- Primary link
- https://github.com/Alibaba-NLP/WebAgent (new tab)
- Release or update evidence
- arXiv preprint 2501.07572 published 2025-01-13; official project announcement says WebWalker released 2025-01-14 and accepted ACL 2025.
- What it measures
- Measures question-answer accuracy and action count, with analyses by traversal depth, source count, domain, and language. Action count measures efficiency.
- Task data
- Public Hugging Face dataset: 680 Chinese/English questions over 1,373 webpages, with root URLs, hop counts and golden paths; tasks require single-site traversal or multisource navigation.
- Evaluation code
- Public official WebAgent repository provides the WebWalker framework, demo, and benchmark materials; the dataset is publicly distributed through Hugging Face. The paper specifies click-only interaction and a 15-step limit.
- Judging criteria
- GPT-4 evaluates answer correctness because generated answers vary, making exact match unsuitable. Correct-run action count measures efficiency.
- Reproducibility and barriers
- The public dataset and implementation support reproduction, but crawling, rendering, live-site changes, and web traversal create barriers. The benchmark requires access to current webpages.
- Who reported the results
- Original WebWalker results are benchmark-author runs. WebDancer reports separate WebWalkerQA runs; these are evaluations of the existing benchmark, not new benchmarks.
Evidence supplied by Exa
- https://arxiv.org/abs/2501.07572 (new tab)
Official paper, benchmark size/task formulation, traversal setup, GPT-4 judging and metrics.
- https://aclanthology.org/2025.acl-long.508.pdf (new tab)
ACL paper details, dataset construction, 680 QA pairs, 1,373 pages, bilingual/domain statistics and evaluation.
- https://github.com/Alibaba-NLP/WebAgent (new tab)
Official implementation/project and benchmark release links.
- https://huggingface.co/datasets/callanwu/WebWalkerQA (new tab)
Public dataset card describing 680 QA records and fields including root URL, hop and golden path.
16. xbench-DeepSearchlive-web deep-search/tool-use benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- live-web deep-search/tool-use benchmark
- Primary link
- https://github.com/xbench-ai/xbench-evals (new tab)
- Release or update evidence
- Official README published 2025-05-28; DeepSearch-2505 and DeepSearch-2510 releases fall within the eligibility window.
- What it measures
- Measures task accuracy and supports cost/task and time/task reporting; the public runner enables repeated evaluation.
- Task data
- Public repository provides encrypted CSV datasets for DeepSearch-2505 and DeepSearch-2510, covering planning, search, reasoning and summarization; updates are quarterly.
- Evaluation code
- Public repository includes xbench_evals.py, data artifacts and decryption/evaluation workflows; versions 2505/2510 and an LLM-judge scorer are inspectable.
- Judging criteria
- An LLM judge scores answers; questions underwent manual collection, human validation and quality filtering.
- Reproducibility and barriers
- Encrypted files and proprietary agent interfaces limit full reproduction; public runner and data support model-side reproduction where APIs exist. Assessment: many product runs were manual.
- Who reported the results
- Official leaderboard reports author-run provider/product evaluations, including May and August 2025 tables; no independent reproductions are established.
Evidence supplied by Exa
- https://github.com/xbench-ai/xbench-evals (new tab)
Official datasets, runner, releases, metrics and leaderboard tables.
- https://xbench.org/agi/aisearch (new tab)
Official scope, open-source links, update information and Eval Card link.
- https://xbench.org/files/Eval%20Card%20xbench-DeepSearch.pdf (new tab)
Evaluation design, validation, manual execution and LLM judging.
- https://github.com/xbench-ai/xbench-evals/blob/main/data/DeepSearch-2510.csv (new tab)
Released dataset schema and encrypted task artifact.
17. SealQAsearch-augmented factual reasoning benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- search-augmented factual reasoning benchmark
- Primary link
- https://huggingface.co/datasets/vtllms/sealqa (new tab)
- Release or update evidence
- arXiv:2506.01062 was published 2025-06-01; the official dataset is publicly released within the eligibility window.
- What it measures
- Measures factual-answer accuracy with and without search across Seal-0 and Seal-Hard; LongSeal measures evidence selection among distractor documents.
- Task data
- Public data include Seal-0, Seal-Hard and LongSeal, covering conflicting, noisy or unhelpful search results across domains; LongSeal supplies many documents with one relevant answer source.
- Evaluation code
- Public Hugging Face data and paper artifacts are available; the benchmark is dynamic and periodically updated. No standalone official evaluator repository was verified.
- Judging criteria
- Scores factual answer accuracy under no-search and search/tool conditions. Exact automated judge implementation is not established from inspected sources.
- Reproducibility and barriers
- Public data aid access, but version updates, search-provider behavior and proprietary APIs limit frozen-test reproduction. A canary string addresses contamination.
- Who reported the results
- Reported results are author-run evaluations using specified models and search systems; no independent reproductions are established.
Evidence supplied by Exa
- https://doi.org/10.48550/arxiv.2506.01062 (new tab)
Publication date, benchmark variants, task design, versioning and reported results.
- https://huggingface.co/datasets/vtllms/sealqa (new tab)
Official public dataset, splits, results and contamination canary.
18. Mind2Web 2Core research benchmark: agentic search / long-horizon web research and citation-backed synthesis
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Core research benchmark: agentic search / long-horizon web research and citation-backed synthesis
- Primary link
- https://osu-nlp-group.github.io/Mind2Web-2/ (new tab)
- Release or update evidence
- Initial public release was 2025-06-26; the 2025-10-23 update released evaluation scripts for both public dev and test sets, within the eligibility window.
- What it measures
- Measures task completion through Partial Completion, Success Rate and Pass@3, alongside completion time, answer length, correctness and source attribution.
- Task data
- Public artifact provides 130 long-horizon web-search tasks involving synthesis, list retrieval, time-varying information and citations; dev and test data are available.
- Evaluation code
- Public MIT repository includes run_eval.py and task-specific Extractor/Verifier workflows; the 2025-10-23 release covers both public dev and test sets.
- Judging criteria
- An Extractor parses claims and citations, while a Verifier checks them against webpages or screenshots; leaf scores aggregate into task and root metrics.
- Reproducibility and barriers
- Public code and tasks improve reproduction, but live webpages, crawling, LLM/API judges and changing information introduce cost and nondeterminism. Assessment: results may drift.
- Who reported the results
- Original authors report 2025 agent and human evaluations, including private-test-era results; no independent reproduction is established.
Evidence supplied by Exa
- https://github.com/OSU-NLP-Group/Mind2Web-2 (new tab)
Release dates, evaluation-script updates, repository and instructions.
- https://github.com/OSU-NLP-Group/Mind2Web-2/blob/main/run_eval.py (new tab)
Evaluation runner, version options and local execution.
- https://huggingface.co/datasets/osunlp/Mind2Web-2 (new tab)
Dataset artifact and all-scripts release corroboration.
- https://arxiv.org/html/2506.21506v2 (new tab)
Task design, metrics, rubric workflow and reported evaluations.
19. Deep Research Bench (DRB; FutureSearch)Web research, dataset/list compilation and evidence discovery; frozen/live variants
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Web research, dataset/list compilation and evidence discovery; frozen/live variants
- Primary link
- https://arxiv.org/abs/2506.06287 (new tab)
- Release or update evidence
- 2025-06: paper arXiv:2506.06287 introduces the benchmark; the official FutureSearch post is dated June 25, 2025.
- What it measures
- Eight research task families include dataset compilation, reference-class enumeration, evidence gathering, source tracing, numeric derivation and claim validation. The paper has 89 task instances; later leaderboard snapshots differ.
- Task data
- Full tasks are explicitly withheld to limit contamination; the paper publishes eight examples. Human-worked references and a frozen RetroSearch corpus support controlled evaluation, but a complete public bundle was not verified.
- Evaluation code
- The paper describes agent tooling and automated trace evaluation, and a public leaderboard exists. No complete public task/evaluator repository was verified; full reproduction code availability is partial/unclear.
- Judging criteria
- Task-specific scoring: row precision/recall/F1; URL or evidence recall; binary numeric/source correctness; and normalized distance from human probability judgments. LLMs assist entity matching and source-reliability checks; trace diagnostics are separate.
- Reproducibility and barriers
- RetroSearch controls web drift, but withheld tasks, proprietary product runs and incomplete evaluator artifacts limit independent reproduction. Assessment: repeatability is partial.
- Who reported the results
- FutureSearch authors report evaluations of LLMs, agents and commercial research products; leaderboard results are vendor/author-run, not independent validation.
Evidence supplied by Exa
- https://arxiv.org/html/2506.06287v1 (new tab)
Paper date, task scope, RetroSearch, withheld data and methodology.
- https://futuresearch.ai/deep-research-bench/ (new tab)
Official benchmark page, later snapshot, frozen pages and update policy.
- https://drb.futuresearch.ai/ (new tab)
Official public leaderboard and evaluation endpoint.
20. DeepSearchQA (Google DeepMind)comprehensive multi-entity/list retrieval and deep-search benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- comprehensive multi-entity/list retrieval and deep-search benchmark
- Primary link
- https://arxiv.org/abs/2601.20975 (new tab)
- Release or update evidence
- Technical report posted 2026-01-28, within the cutoff; the public dataset is available on Hugging Face and the official leaderboard is hosted by Kaggle.
- What it measures
- Measures exhaustive set-answer generation through fully correct and incorrect set rates, extraneous-answer rate and F1, covering collation, deduplication, entity resolution and stopping.
- Task data
- Public dataset contains 900 expert-annotated prompts across 17 fields with objectively verifiable gold sets, answer types and time or source anchors; about 65% are set-answer tasks.
- Evaluation code
- Public dataset is downloadable, but no official standalone evaluator repository was verified. Third-party runners exist and are not Google code.
- Judging criteria
- A fully correct response exactly matches the gold set; F1 balances missing and extraneous entities. Evaluation uses an automated judging methodology.
- Reproducibility and barriers
- Static or time-anchored tasks and public data aid reproduction; web access, APIs and judge-model choice remain barriers. Official evaluator-code availability is unknown.
- Who reported the results
- Google DeepMind results are author-run; Kaggle results are platform-evaluated submissions and do not establish independent scientific replication.
Evidence supplied by Exa
- https://arxiv.org/html/2601.20975v1 (new tab)
Report scope, task construction, metrics, judging and results.
- https://huggingface.co/datasets/google/deepsearchqa (new tab)
Public dataset artifact, schema, answer types and evaluator guidance.
- https://www.kaggle.com/benchmarks/google/dsqa/leaderboard (new tab)
Official Google/Kaggle leaderboard endpoint.
21. DRACO Benchmark (Perplexity Research)long-form research report quality/citation benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- long-form research report quality/citation benchmark
- Primary link
- https://research.perplexity.ai/articles/evaluating-deep-research-performance-in-the-wild-with-the-draco-benchmark (new tab)
- Release or update evidence
- Eligible: public release announced in 2026; technical report arXiv:2602.11685 is dated 2026-02-12, before the 2026-09-26 cutoff. The announcement says the benchmark, rubrics and judge prompt are open sourced.
- What it measures
- Evaluates factual accuracy, breadth/depth, presentation quality, objectivity and citation quality across 100 tasks using weighted binary rubric criteria, including penalties for unsupported claims.
- Task data
- Public dataset availability: production-derived Perplexity Deep Research requests were de-identified, reformulated and filtered through five stages, with expert-reviewed rubrics.
- Evaluation code
- Public benchmark, rubrics, judge prompt and linked dataset are released; the exact harness boundary is unclear, and proprietary production sampling and research tools remain unreproducible.
- Judging criteria
- An LLM judge assigns weighted binary verdicts to rubric criteria; rubrics assess accuracy, completeness/depth, presentation and primary-source citation quality, with reliability checked across three judge models.
- Reproducibility and barriers
- Public tasks, rubrics and judge prompt support reproduction, but production-query sampling, proprietary tools and English single-turn scope remain barriers or limitations (assessment).
- Who reported the results
- Perplexity's own comparison of four deep-research systems; reported scores and leadership claims are vendor-run, not independent.
Evidence supplied by Exa
- https://research.perplexity.ai/articles/evaluating-deep-research-performance-in-the-wild-with-the-draco-benchmark (new tab)
Official release: tasks, construction, rubrics, judge prompt, dimensions, results and limitations.
- https://arxiv.org/html/2602.11685v1 (new tab)
Primary technical report dated 2026-02-12 and detailed DRACO methodology/results.
- https://hf.co/datasets/perplexity-ai/draco (new tab)
Public dataset location linked by the official release.
22. WANDR: A Benchmark for Wide and Deep Researchcomprehensive multi-entity discovery plus evidence-supported enrichment (wide-and-deep research)
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- comprehensive multi-entity discovery plus evidence-supported enrichment (wide-and-deep research)
- Primary link
- https://arxiv.org/abs/2608.14747 (new tab)
- Release or update evidence
- Submitted to arXiv 2026-08-14, within the 2026-09-26 cutoff. The official repository README was available by 2026-07-14 and contains the released task/evaluator tree.
- What it measures
- Scores record-level and hierarchical precision, recall and F1, including hard complete-subtree scores, soft partial-credit scores, retrieval-only versus full-record results, and task rollups.
- Task data
- Public 500-task packages encode qualification hierarchies and requested record volumes. Packages include instructions, schemas, fixtures, labels and evaluator artifacts, but not a static exhaustive gold universe.
- Evaluation code
- Public repository contains source tasks, Harbor adapter, generated packages, task-local evaluator, scripts, configurations and reports, with validation and full-run instructions.
- Judging criteria
- The evaluator fetches and canonicalizes cited pages, deduplicates entities, verifies claims against pages and excerpts, then aggregates task-specific record verdicts hierarchically. Precision measures submitted quality; recall measures coverage against requested volume.
- Reproducibility and barriers
- Public tasks, Docker packages, evaluator, configs and manifests support reproduction. End-to-end runs require live retrieval and paid APIs; web drift, bot walls, provider settings and judge nondeterminism remain barriers (assessment).
- Who reported the results
- Original authors' pinned runs over six production systems; these are author-reported benchmark results, not independent reproductions.
Evidence supplied by Exa
- https://arxiv.org/html/2608.14747 (new tab)
Primary paper: release, tasks, hierarchy, construction, scoring pipeline, metrics and system results.
- https://github.com/perplexityai/wandr (new tab)
Observed repository: tasks, Harbor packages, evaluator, configs, instructions and paid-API requirements.
- https://research.perplexity.ai/articles/wandr-benchmark-evaluating-research-agents-that-must-search-wide-and-deep (new tab)
Official description of evidence, dynamic judging, rollups and benchmark setup.
23. DeepWide (Table-as-Search business-development benchmark)constrained entity discovery and attribute enrichment; paper-defined evaluation
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- constrained entity discovery and attribute enrichment; paper-defined evaluation
- Primary link
- https://arxiv.org/abs/2602.06724 (new tab)
- Release or update evidence
- 2026-02: Table-as-Search (arXiv:2602.06724) introduces the 20-query DeepWide evaluation.
- What it measures
- Column-F1 measures identification of entities satisfying complex constraints; Item-Precision measures correctness of retrieved attributes. Tasks use fixed retrieval quantities rather than an exhaustive-universe claim.
- Task data
- 20 real-world business-development and e-commerce queries require constrained discovery and attribute enrichment. Expert-verified references combine pooled system matches; a complete downloadable bundle was not separately verified.
- Evaluation code
- Table-as-Search agent code and prompts are public in Marco-Search-Agent; an independently packaged scorer and reference bundle for these 20 tasks were not verified. The repository's DeepWideSearch data/eval directory is separate.
- Judging criteria
- Experts verify candidate validity and requested information against all stated constraints. Column-F1 scores entity identification, while Item-Precision scores attribute correctness; fixed quantities, exclusions and dynamic reference unions address open-ended completeness.
- Reproducibility and barriers
- Agent code is public, but a complete bundle of these 20 tasks, reference sets and scoring code was not verified. Live retrieval, commercial interfaces, dynamic reference unions and expert judgments complicate reproduction.
- Who reported the results
- Original authors' Table-as-Search runs are author-reported; no independent reproduction was established here.
Evidence supplied by Exa
- https://arxiv.org/html/2602.06724 (new tab)
Primary paper: task definitions, 20-query data, references, judging, metrics, limitations and results.
- https://github.com/AIDC-AI/Marco-Search-Agent (new tab)
Observed repository: DeepWideSearch artifacts and Table-as-Search implementation, prompts and tools.
- https://arxiv.org/abs/2510.20168 (new tab)
Distinguishes earlier 220-question DeepWideSearch from the 20-query evaluation.
24. DeepWeb-BenchDeep research benchmark; comprehensive multi-entity evidence and derivation
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Deep research benchmark; comprehensive multi-entity evidence and derivation
- Primary link
- https://arxiv.org/abs/2605.21482 (new tab)
- Release or update evidence
- Submitted 2026-05-20, within cutoff. The primary paper says data, rubrics and evaluation code are publicly released.
- What it measures
- Scores retrieval, derivation, reasoning and calibration across 100 cases using entity-by-dimension cells, family scores and source-provenance disclosure levels.
- Task data
- Public Hugging Face release: 100 cases, 900 model results/answers/score records, summaries and provenance. Tasks require 6–10 entities across 6–10 dimensions with cross-source evidence.
- Evaluation code
- Actually present in HF
code/: validation, leaderboard/report rebuilding, rule-prompt grader reruns and an OpenAI-compatible model runner. Aggregation needs no keys; live reruns do. - Judging criteria
- Explicit per-cell rules, provenance levels and cross-source checks determine scores; evaluation is not solely free-form judging.
- Reproducibility and barriers
- Data and code are downloadable, but model, grader, search and scrape APIs require keys. Raw tool traces, third-party source snapshots and local MCP/API state are excluded.
- Who reported the results
- Author-reported evaluation of nine frontier models; reported findings concern overall scores and retrieval, derivation and calibration error patterns.
Evidence supplied by Exa
- https://arxiv.org/abs/2605.21482 (new tab)
Primary paper, date, task design, metrics, findings and release claim.
- https://huggingface.co/datasets/deepweb-bench-anon/deepweb-bench (new tab)
Released files, cases, result records, executable code and exclusions.
- https://sixiongxie1001-dot.github.io/deep-research-benchmark2.0 (new tab)
Official project page.
25. From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task EvolutionAutomatically constructed verifiable deep-research benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Automatically constructed verifiable deep-research benchmark
- Primary link
- https://arxiv.org/abs/2608.02163 (new tab)
- Release or update evidence
- Submitted 2026-08-03, within cutoff. The primary paper says code and data are publicly available.
- What it measures
- Evaluates 500 tasks across 31 topics and 10 categories using DAG atomic steps, checkpoints, fact-grounded pointwise rubrics, and model discrimination and stability analyses.
- Task data
- Constructed from Wikipedia QA through Explorer–Formalizer–Challenger evolution; each task includes a query, DAG, checkpoints and aligned rubrics. Public availability is claimed by the paper.
- Evaluation code
- Paper links https://github.com/chr6192/TaskEvolving.git (new tab) and states that implementation, data and results are public. Repository contents were not retrievable in this check, so evaluator files remain unverified.
- Judging criteria
- Source-grounded checkpoint and DAG rubrics assign pointwise fact-based scores; the paper describes the method as human-aligned and stable.
- Reproducibility and barriers
- Public-release claim is explicit, but construction requires relatively high frontier-model cost. Exact repository data layout and license remain unverified.
- Who reported the results
- Author-reported experiments demonstrate discrimination across models and query types; no independent results were established.
Evidence supplied by Exa
- https://arxiv.org/abs/2608.02163 (new tab)
Primary date, scope, 500-task design and public code/data claim.
- https://arxiv.org/html/2608.02163v1 (new tab)
Official HTML identifies GitHub TaskEvolving repository and method details.
- https://github.com/chr6192/TaskEvolving.git (new tab)
Primary artifact link cited by paper; file-level availability not verified here.
26. DeepResearch-9Kchallenging multi-hop web research benchmark with agent trajectories
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- challenging multi-hop web research benchmark with agent trajectories
- Primary link
- https://arxiv.org/abs/2603.01152 (new tab)
- Release or update evidence
- ArXiv paper 2603.01152 was available in March 2026 and identifies SIGIR 2026 publication (July 20–24, 2026), within cutoff. Official repository was created 2026-02-06.
- What it measures
- Measures final-answer accuracy, search-tool usage, and trajectory correctness across three difficulty levels.
- Task data
- Public 9,000-question Hugging Face release includes difficulty labels, questions, answers, trajectories, and a 3,974-sample hard subset. Data are largely synthetic and teacher-generated.
- Evaluation code
- Public official repository includes training, inference, SFT/RL, environment, and evaluation scripts. The paper’s construction pipeline is released, with DeepSeek-V3 LLM judging.
- Judging criteria
- LLM judging scores final-answer correctness; difficulty reflects search-chain complexity and entity obfuscation. Hard items were filtered by incorrect teacher verdicts.
- Reproducibility and barriers
- Data and code are public, but training requires substantial resources and model/API dependencies. Teacher-generated trajectories and LLM judging create assessment dependence.
- Who reported the results
- Authors’ experiments; DeepResearch-R1 is the paired training/agent framework, not an additional benchmark. No independent replication established.
Evidence supplied by Exa
- https://arxiv.org/html/2603.01152v2 (new tab)
Primary paper: benchmark construction, 9K/3,974 sizes, difficulty levels, metrics, judge protocol, results, and release links.
- https://github.com/Applied-Machine-Learning-Lab/SIGIR2026_DeepResearch-R1 (new tab)
Official code repository: dataset links, hard subset, data format, SFT/RL and inference/evaluation scripts.
- https://huggingface.co/datasets/artillerywu/DeepResearch-9K (new tab)
Official dataset artifact: 9,000/3,974 splits and train/test composition.
27. Mr.LHDRMultimodal real-world long-horizon deep-research benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Multimodal real-world long-horizon deep-research benchmark
- Primary link
- https://arxiv.org/abs/2609.11318 (new tab)
- Release or update evidence
- Submitted 2026-09-10, within 2026-09-26 cutoff.
- What it measures
- Measures final-answer and dependency-aware conclusion quality using OA, Strict Accuracy, Checklist Score, and Dependency-Aware Checklist Score.
- Task data
- Public dataset uses multimodal evidence including images, maps, PDFs, logos, charts, tables, and video frames. Metadata, checklists, sources, and images are released, but some fields are encrypted or canary protected.
- Evaluation code
- Public Apache-2.0 repository includes requirements, decryption, run, and evaluation scripts. Actual execution requires the dataset and model/web-search access.
- Judging criteria
- Scores final answers and intermediate conclusions against annotated dependencies using OA, SA, CS, and DACS.
- Reproducibility and barriers
- Code and dataset are public, but encrypted metadata and decrypt.py create a material reproduction barrier; raw data are not fully transparent.
- Who reported the results
- Author-reported evaluation; strongest system achieved 43.1% OA and 34.3% SA. Removing images reduced DACS by 12.6 points.
Evidence supplied by Exa
- https://arxiv.org/abs/2609.11318 (new tab)
Primary paper/date, benchmark design, metrics and reported results.
- https://arxiv.org/html/2609.11318v1 (new tab)
Official HTML links code repository.
- https://github.com/minghaoguo20/Mr-LHDR-eval (new tab)
Actual public Apache-2.0 repository, requirements, data download/decrypt and run scripts.
- https://huggingface.co/datasets/Henryeahhh/Mr-LHDR (new tab)
Actual dataset schema, source/checklist/image fields and encrypted-looking released records.
28. FinSearchCompfinancial open-web search and analyst-style reasoning
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- financial open-web search and analyst-style reasoning
- Primary link
- https://arxiv.org/abs/2509.13160 (new tab)
- Release or update evidence
- Submitted 2025-09-16, within cutoff. Official project and repository are public.
- What it measures
- Measures binary answer correctness across time-sensitive retrieval, historical lookup, and complex historical investigation.
- Task data
- Public 635-question expert-crafted dataset covers Global and Greater China subsets and includes questions, tool templates, answers, and traces.
- Evaluation code
- Actual public evaluator code is available at https://github.com/randomtutu/FinSearchComp (new tab), including data/finsearchcomp_data.json, eval/eval.py, and runnable commands.
- Judging criteria
- Rubric-guided LLM judging assigns 0/1 correctness, applying numerical tolerance and checking financial conventions and supporting evidence.
- Reproducibility and barriers
- Public data and evaluator support reproduction, but live APIs, web changes, and model/tool access remain practical assessment barriers.
- Who reported the results
- Author-run model comparison; no independent reproduction established.
Evidence supplied by Exa
- https://arxiv.org/html/2509.13160 (new tab)
Paper date, 635 questions, task families, metrics, judge, and release links.
- https://github.com/randomtutu/FinSearchComp (new tab)
Observed official repository with data files and eval/eval.py.
- https://randomtutu.github.io/FinSearchComp/ (new tab)
Official project page and public-release description.
29. Finance Agent Benchmarkfinancial SEC-filing research-agent benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- financial SEC-filing research-agent benchmark
- Primary link
- https://arxiv.org/abs/2508.00828 (new tab)
- Release or update evidence
- ArXiv release August 2025, within cutoff.
- What it measures
- Measures naive and class-balanced accuracy across nine finance categories, plus execution time and cost.
- Task data
- Partially public: 537 expert-authored SEC/EDGAR questions comprise 50 public validation, 150 private validation, and 337 private test items.
- Evaluation code
- Actual MIT-licensed harness and public validation data are available at https://github.com/vals-ai/finance-agent (new tab); Zenodo also provides the harness.
- Judging criteria
- Rubric-based component grading checks calculations and detects contradictions before assigning correctness.
- Reproducibility and barriers
- Public 50-item validation and code permit partial reproduction; private splits, live filings/search, and paid APIs limit full reproduction (assessment).
- Who reported the results
- Benchmark-author runs; no independent reproduction established.
Evidence supplied by Exa
- https://arxiv.org/html/2508.00828 (new tab)
Primary paper, data splits, judging, metrics, and results.
- https://github.com/vals-ai/finance-agent (new tab)
Observed public harness repository.
- https://doi.org/10.5281/zenodo.15428823 (new tab)
Paper-linked harness artifact.
30. FinRetrievalfinancial structured-data retrieval by AI agents
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- financial structured-data retrieval by AI agents
- Primary link
- https://arxiv.org/abs/2603.04403 (new tab)
- Release or update evidence
- January 2026 release; within cutoff.
- What it measures
- Measures numeric-answer accuracy across 500 questions and 14 model/tool configurations, with tool-call trace analysis.
- Task data
- Public release includes 500 questions, 7,000 responses, ground truth, scores, and complete tool-call traces comparing web-only and structured MCP/API access.
- Evaluation code
- Actual public evaluator is available at https://github.com/daloopa/finretrieval (new tab); README provides setup and execution commands.
- Judging criteria
- Automatic numeric matching compares answers with ground truth; edge cases and fiscal-period conventions receive manual review.
- Reproducibility and barriers
- Dataset, evaluator, scores, and traces are public; Daloopa MCP and commercial model/API access remain practical rerun barriers (assessment).
- Who reported the results
- Daloopa-author evaluation; no independent reproduction established.
Evidence supplied by Exa
- https://arxiv.org/abs/2603.04403 (new tab)
Primary paper and release claims, task count, configurations, metrics, and results.
- https://github.com/daloopa/finretrieval (new tab)
Verified public evaluation-code repository and README.
- https://huggingface.co/datasets/daloopa/finretrieval (new tab)
January 2026 dataset card and evaluator link.
31. MedBrowseCompmedical live-web/deep-research and biomedical evidence retrieval
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- medical live-web/deep-research and biomedical evidence retrieval
- Primary link
- https://arxiv.org/abs/2505.14963 (new tab)
- Release or update evidence
- Submitted 2025-05-20, within cutoff; repository and dataset match arXiv 2505.14963.
- What it measures
- Accuracy on MedBrowseComp-50 and MedBrowseComp-605, covering structured extraction and deep research, judged by GPT-4.1-mini with human checking.
- Task data
- Public 50-sample and 605-sample benchmarks cover hematology/oncology, PubMed, ClinicalTrials.gov, FDA Orange Book and market data.
- Evaluation code
- Public official repository https://github.com/shan23chen/MedBrowseComp (new tab) contains final50.csv, final121.csv and processing/evaluation scripts, including process_NCT_predictions.py. The 50/605 counts are verified; the paper describes 121 trials expanded into 605 tasks.
- Judging criteria
- GPT-4.1-mini judges answer correctness, with human checking of judge agreement.
- Reproducibility and barriers
- Partial reproduction is possible with public files and scripts, but live medical sources and proprietary systems create access and drift barriers (assessment).
- Who reported the results
- Original authors' commercial-system and Claude computer-use runs; no independent benchmark runner established.
Evidence supplied by Exa
- https://arxiv.org/html/2505.14963 (new tab)
Canonical paper identity, date, construction, 50/605 evaluation variants, judging and scripts.
- https://github.com/shan23chen/MedBrowseComp (new tab)
Official project repository and actual final50/final121 files/scripts.
- https://huggingface.co/datasets/AIM-Harvard/MedBrowseComp (new tab)
Official dataset artifact with 50 and 605 harmonized datasets.
32. PaSa / AutoScholarQuery and RealScholarQueryacademic literature discovery and comprehensive paper search
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- academic literature discovery and comprehensive paper search
- Primary link
- https://aclanthology.org/2025.acl-long.572/ (new tab)
- Release or update evidence
- ACL 2025 primary paper release, within cutoff; official ByteDance repository is linked by the paper and search result.
- What it measures
- Recall@20, Recall@50, Recall@100 and precision measure relevant-paper retrieval, document selection and crawling.
- Task data
- AutoScholarQuery has public 33,551 training, 1,000 development and 1,000 test synthetic queries. RealScholarQuery has 50 manually gathered researcher queries and an annotated 200 query-paper selector set.
- Evaluation code
- Public code and datasets are stated at https://github.com/bytedance/pasa (new tab). Repository search results show runnable agent scripts and benchmark support using scholarly retrieval, citation crawling and ar5iv parsing.
- Judging criteria
- Ranked relevant-paper retrieval is scored against annotated sets using recall and precision; separate human checks assess query and query–paper relevance.
- Reproducibility and barriers
- AutoScholarQuery is comparatively reproducible, while RealScholarQuery may drift with manually collected queries and changing indexes; search/API access and model compute are required (assessment).
- Who reported the results
- Authors' benchmark runs, not independent validation. Reported comparisons are benchmark-author results.
Evidence supplied by Exa
- https://aclanthology.org/2025.acl-long.572/ (new tab)
Primary ACL paper: datasets, date-constrained scholarly retrieval task, metrics and reported results.
- https://github.com/bytedance/pasa (new tab)
Official code/data repository linked by the primary paper; search result identifies runnable PaSa agent and benchmark support.
33. OpenBenchmarks Multi-turn Company Search Benchmarkmulti-turn web-search evaluation for comprehensive company-set discovery
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- multi-turn web-search evaluation for comprehensive company-set discovery
- Primary link
- https://openbenchmarks.com/multi-turn-company-search (new tab)
- Release or update evidence
- Official page records first complete benchmark publication on 2026-08-22, additions on 2026-08-26 and updates on 2026-09-15; all fall within the cutoff.
- What it measures
- Precision, recall, F1 and exact-set accuracy score company-set discovery; median latency and cost measure efficiency across search-only and search-plus-fetch settings.
- Task data
- Full board gold is locked: 45 hand-labelled questions with 375 canonical memberships. A separate public 10-question search-only sample with frozen reference companies is available via Hugging Face.
- Evaluation code
- Public MIT-licensed GitHub runner and deterministic offline judge validate schemas, receipts, canonical matching and aggregation. Running agents requires provider credentials; saved-artifact judging is offline.
- Judging criteria
- Canonical set comparison counts true positives, false positives and false negatives for precision, recall, F1 and exact-set accuracy. Three trials are aggregated; SD measures trial variability, not confidence.
- Reproducibility and barriers
- Partial reproduction is possible with the public runner, judge and 10-question sample, but locked gold, live providers, costs and web drift remain barriers (assessment).
- Who reported the results
- OpenBenchmarks editorial-board/vendor runs; no independent results established. Board scores use the locked 45-question set, while the public sample is not score-comparable.
Evidence supplied by Exa
- https://openbenchmarks.com/multi-turn-company-search (new tab)
Primary benchmark page: 2026-08-22 publication, subsequent changelog dates, 45-question/375-membership board, public 10-question sample, metrics and locked-vs-public distinction.
- https://github.com/openbenchmarks-labs/multi-turn-company-search (new tab)
Official public runner/judge repository, MIT license, fixed agent, deterministic scorer, artifact contract, credentials and dataset/board caveats.
- https://huggingface.co/datasets/openbenchmarks/OB-Company-Websearch (new tab)
Official advertised endpoint for the separate 10-question public sample; availability should be checked at publication time.
34. MM-BrowseCompmultimodal deep web browsing benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- multimodal deep web browsing benchmark
- Primary link
- https://github.com/MMBrowseComp/MM-BrowseComp (new tab)
- Release or update evidence
- ArXiv v1 2025-08-14; official repository says full codebase released 2025-08-20 and dataset expanded to 400 questions 2026-01-02. Version drift remains: v1 describes 224/244 questions, current release 400.
- What it measures
- Final-answer accuracy measures correctness; checklist analysis measures multimodal dependencies and reasoning paths beyond aggregate scores.
- Task data
- Public encrypted JSONL includes the current 400-question update at data/MMBrowseComp_400.jsonl. Prompts and evidence may include images and videos; canary/decryption protects questions and answers.
- Evaluation code
- Public src/decrypt.py, src/gen_answer.py and src/eval.py implement decryption, answer generation and LLM judging. The repository includes released data and evaluation scripts.
- Judging criteria
- An LLM judges reference-answer correctness, while verified per-question checklists diagnose dependency and reasoning paths.
- Reproducibility and barriers
- Partial reproduction is possible with public code/data, but encrypted contents, canary/decryption, API-backed models, live web access and configuration create barriers (assessment).
- Who reported the results
- Paper/author-reported model evaluations; no independent result established here.
Evidence supplied by Exa
- https://arxiv.org/abs/2508.13186 (new tab)
Primary paper/date and multimodal browsing design; v1 task count and checklist concept.
- https://github.com/MMBrowseComp/MM-BrowseComp (new tab)
Official release history, current 400-question update, encrypted data and concrete decrypt/generate/evaluate commands.
35. MMSearch-Plusprovenance-aware multimodal browsing/search benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- provenance-aware multimodal browsing/search benchmark
- Primary link
- https://github.com/mmsearch-plus/MMSearch-Plus (new tab)
- Release or update evidence
- ArXiv released 2025-08-29; official README records all data samples released to Hugging Face 2025-09-26, within cutoff.
- What it measures
- Answer accuracy measures correctness on 311 tasks; bounding-box, cropping and provenance analyses measure localized visual search and evidence-chain behavior.
- Task data
- Public encrypted dataset includes questions/images, ground truth, alternatives, metadata and Set-of-Mark annotations. Tasks require localized visual cues, spatial-temporal reasoning, iterative retrieval and cross-validation.
- Evaluation code
- Public official repository provides the agentic rollout framework, evaluation script and Set-of-Mark annotations; Hugging Face data are linked. External search infrastructure is not claimed to be bundled.
- Judging criteria
- Answers are scored for accuracy, while visual localization/cropping and provenance-aware retrieval analyses assess multimodal evidence chains rather than text-only shortcuts.
- Reproducibility and barriers
- Partial reproduction requires public data/code plus decryption, model/API and live-search dependencies; changing image sources and retrieval noise may affect results (assessment).
- Who reported the results
- Author-run paper/official leaderboard results; independent reproduction not verified.
Evidence supplied by Exa
- https://arxiv.org/abs/2508.21475 (new tab)
Primary paper date, 311 tasks, curation/evaluation design and reported results.
- https://github.com/mmsearch-plus/MMSearch-Plus (new tab)
Official release dates, dataset link, decryption usage, framework/evaluation and annotation availability.
- https://huggingface.co/datasets/Cie1/MMSearch-Plus (new tab)
Public dataset artifact linked by the project.
36. BrowseComp-V³ (BrowseComp-V3)visual, vertical, verifiable multimodal deep-search benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- visual, vertical, verifiable multimodal deep-search benchmark
- Primary link
- https://github.com/Halcyon-Zhang/BrowseComp-V3 (new tab)
- Release or update evidence
- 2026-02: arXiv:2602.12876 and official project/repository. Unicode V³ and ASCII V3 designate the same benchmark.
- What it measures
- Measures final-answer success and process-level adherence using expert-validated intermediate subgoals and search trajectories.
- Task data
- Partial/public: encrypted downloadable data contain 300 questions across 24 subdomains, visual assets, evidence, gold trajectories and subgoals; repository scripts decrypt JSON and images.
- Evaluation code
- Actual public repository code includes dataset download/decryption, rollout evaluation, score summarization, an OmniSeeker runner, documentation and smoke tests.
- Judging criteria
- Scores final-answer correctness and success alongside process and subgoal adherence for cross-modal, multi-hop evidence integration.
- Reproducibility and barriers
- Public GitHub/Hugging Face artifacts and CC BY 4.0 claim; encrypted samples, key, live search, judge model and APIs remain dependencies. Assessment: web volatility and closed baselines limit exact reproduction.
- Who reported the results
- Author-reported paper/project evaluations include OmniSeeker and MLLM comparisons; independent results are not verified.
Evidence supplied by Exa
- https://arxiv.org/abs/2602.12876 (new tab)
Primary benchmark description, 300 tasks, process evaluation and reported results.
- https://github.com/Halcyon-Zhang/BrowseComp-V3 (new tab)
Official repository contents: dataset download/decryption, assets, evaluator, baseline runner and license documentation.
- https://halcyon-zhang.github.io/BrowseComp-V3/ (new tab)
Official project page and artifact links, including dataset/repository.
37. ScholarQuestAcademic literature/paper-search benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Academic literature/paper-search benchmark
- Primary link
- https://arxiv.org/abs/2606.20235 (new tab)
- Release or update evidence
- Submitted 2026-06-18 (within window); public code/data repository linked by paper.
- What it measures
- Measures retrieval recall at 25, 100 and all results, plus search efficiency, tool use and robustness to intent and answer-set size.
- Task data
- Public: 1,111 queries from 1,000+ CS topic seeds use four intents, arXiv answer sets of 5–200 papers and a million-scale ScholarBase corpus with metadata and citations.
- Evaluation code
- Actual public repository includes the benchmark dataset, construction pipeline, analysis scripts and ScholarBase/Lewen search backend.
- Judging criteria
- Scores retrieval recall against answer_arxiv_ids after arXiv-ID normalization; process statistics cover rounds, calls, candidates and recall efficiency.
- Reproducibility and barriers
- Public code/data; reproduction requires deploying ScholarBase/Lewen and potentially using external search APIs or models (assessment).
- Who reported the results
- Paper authors report PaperScout and hybrid baselines; results are author-reported, with no independent validation established here.
Evidence supplied by Exa
- https://arxiv.org/abs/2606.20235 (new tab)
Paper identity, 2026-06-18 submission, benchmark scope, 1,111 queries, four intents, metrics and headline results.
- https://github.com/pty12345/ScholarQuest (new tab)
Public benchmark dataset, code, construction pipeline, ScholarBase/Lewen backend and dataset structure.
- https://arxiv.org/html/2606.20235 (new tab)
Detailed benchmark construction, corpus/retrieval design and evaluation methodology.
38. Sage: Benchmarking and Improving Retrieval for Deep Research AgentsScientific literature retrieval benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Scientific literature retrieval benchmark
- Primary link
- https://arxiv.org/abs/2602.05975 (new tab)
- Release or update evidence
- Submitted 2026-02-06 (within window); public GitHub repository linked from paper.
- What it measures
- Reasoning-intensive scientific paper discovery: exact-match target-paper retrieval and relevance-weighted coverage of open-ended literature queries.
- Task data
- Public repository contains 600 short-form and 600 open-ended queries, with paper IDs/titles and relevance-tier references. The paper describes four domain-specific 50,000-paper corpora; a complete corpus bundle was not verified.
- Evaluation code
- Public question/reference JSON files and metric definitions are verified. The inspected repository README does not establish a turnkey evaluator or full agent/retriever implementation.
- Judging criteria
- Short-form exact match checks whether the target paper appears in answer text or citations. Open-ended weighted recall gives seed papers relevance 2, shared-reference papers 1 and other papers 0.
- Reproducibility and barriers
- Query/reference data are public; reproducing agent experiments requires the corresponding corpus, retrieval indexes and model/API setup. Corpus augmentation adds processing cost.
- Who reported the results
- Authors evaluate six deep-research agents and DR Tulu-backed retrievers; reported findings are author-generated, not independently validated here.
Evidence supplied by Exa
- https://arxiv.org/abs/2602.05975 (new tab)
Benchmark definition, 1,200 queries, four domains, 200,000-paper corpus and reported findings.
- https://arxiv.org/html/2602.05975v2 (new tab)
Dataset composition, corpus-search setup, metrics/results and corpus-scaling method.
- https://github.com/HughieHu/Sage (new tab)
Public implementation/artifact linked as the SAGE repository.
39. AstaBench (literature-search and research-synthesis components)Core subset — scientific literature discovery/synthesis; broader science suite is adjacent
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Core subset — scientific literature discovery/synthesis; broader science suite is adjacent
- Primary link
- https://allenai.org/asta/bench (new tab)
- Release or update evidence
- Public Ai2 announcement dated 2025-08-26 and benchmark paper submitted 2025-10-24; within window.
- What it measures
- Measures paper finding, literature retrieval and QA, long-form review answers, literature-review tables, quality, cost and tool-controlled agent behavior.
- Task data
- Public: 11 benchmarks and 2,400+ problems cover literature, coding, data analysis and discovery; literature components include PaperFindingBench, ScholarQA-CS2, LitQA2 and ArxivDIGESTables-Clean.
- Evaluation code
- Actual public GitHub suite provides task datasets, interfaces and standardized tools; Asta Scientific Corpus supports date- and corpus-restricted search and retrieval.
- Judging criteria
- Scores task-specific answers or retrieval; ScholarQA-CS2 adds coverage and citation precision, while logs record cost, tools and traces.
- Reproducibility and barriers
- Open benchmark/framework and baseline agents; reproduction depends on Asta Environment or corpus snapshots and compatible agent-evaluation tooling (assessment).
- Who reported the results
- Ai2 authors report 57-agent evaluations; announcement reports Asta Scholar QA/Elicit/SciSpace and Asta Paper Finder results. Author/vendor results, not independently validated.
Evidence supplied by Exa
- https://allenai.org/asta/bench (new tab)
Official benchmark page listing literature components, metrics and evaluation framework.
- https://github.com/allenai/asta-bench (new tab)
Public code, task names/datasets, standardized tools and date/corpus restrictions.
- https://arxiv.org/abs/2510.21652 (new tab)
Benchmark paper date, suite scope (2,400+ problems), environment and agent evaluation design.
- https://allenai.org/blog/astabench (new tab)
Official release date and reported literature-component results.
40. LiveDRBench (Microsoft), from “Characterizing Deep Research: A Benchmark and Formal Definition”open-web deep-research claim discovery; includes entity/list retrieval and evidence-grounded synthesis
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- open-web deep-research claim discovery; includes entity/list retrieval and evidence-grounded synthesis
- Primary link
- https://github.com/microsoft/LiveDRBench (new tab)
- Release or update evidence
- Public repository created 2025-07-25; paper arXiv:2508.04183 is 2025. Data were collected May–June 2025, within the requested window.
- What it measures
- Measures claim-level precision, recall and F1, including recursive claim/subclaim correctness; also records sources, branching and backtracking.
- Task data
- Public: 100 live open-web tasks cover science and world events, with prompts, output formats, ground-truth JSON claims and references; Hugging Face provides the dataset.
- Evaluation code
- Actual public MIT repository includes src/evaluate.py and instructions; GPT-4o via OpenAI API judges claim agreement before computing information-retrieval metrics.
- Judging criteria
- Scores correctness and completeness of substantive claims; GPT-4o maps predictions to ground-truth claims, and unsupported incorrect subclaims receive no credit.
- Reproducibility and barriers
- Public code/data and a stated refresh plan support reproduction, but live-web changes and GPT/API dependence affect repeatability. It is explicitly open-web, not enterprise search.
- Who reported the results
- Microsoft authors report runs involving OpenAI, Perplexity, Google/Gemini and an open-source agent; independent reproduction is not established here.
Evidence supplied by Exa
- https://arxiv.org/html/2508.04183 (new tab)
Formal definition, open-web scope, 100-task domains, claim representation, metrics, categories and author results.
- https://github.com/microsoft/LiveDRBench (new tab)
Public repository date, 100-task dataset description, Hugging Face loading, evaluation script/instructions and licenses.
41. DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation (LG AI Research)long-form expert report quality, evidence grounding and report-wide fact verification
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- long-form expert report quality, evidence grounding and report-wide fact verification
- Primary link
- https://github.com/hanjanghoon/DEER (new tab)
- Release or update evidence
- Paper arXiv:2512.17776 first released 2025-12-19; official repository created 2026-02-03. Both fall within the requested window.
- What it measures
- Measures report fulfillment, analytical soundness, coherence, style and ethics through 101 rubric items, plus claim factuality, citation support, evidence quality and sufficiency.
- Task data
- Public benchmark artifact, but access is gated by a password-protected archive; it contains 50 expert-report tasks across 13 domains derived from Humanity’s Last Exam and internal queries.
- Evaluation code
- Official MIT-licensed evaluation code is public; dataset access is gated under a custom license permitting non-commercial research but prohibiting redistribution, mirroring and public posting.
- Judging criteria
- LLM judges apply fixed rubric items and expert guidance; a fact-verification module extracts claims and citations, then checks external evidence through web search.
- Reproducibility and barriers
- Gated data, web-search verification, LLM judging, licensing restrictions and anti-contamination requirements materially limit fully frictionless reproduction.
- Who reported the results
- Reported human-correlation and system-comparison results are LG AI Research authors’ experiments; independent validation was not established here.
Evidence supplied by Exa
- https://arxiv.org/html/2512.17776v2 (new tab)
Task count/domains, taxonomy, 101 rubrics, expert guidance, fact verification and judging architecture.
- https://github.com/hanjanghoon/DEER (new tab)
Official code, repository date, dataset location, access procedure, licenses and redistribution restrictions.
- https://huggingface.co/datasets/LG-AI-Research/DEER-Deep-Research-Benchmark (new tab)
Official dataset artifact and custom-license/access status.
42. PaSaMaster-Benchmultidisciplinary scientific literature retrieval / comprehensive paper-set discovery
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- multidisciplinary scientific literature retrieval / comprehensive paper-set discovery
- Primary link
- https://arxiv.org/abs/2605.14306 (new tab)
- Release or update evidence
- 2026-05: arXiv:2605.14306 introduces PaSaMaster-Bench, distinct from PaSa’s earlier AutoScholarQuery/RealScholarQuery datasets.
- What it measures
- Measures top-20 recall, precision, F1 and NDCG, plus source-hallucination rate and token cost; the paper reports 244 expert-curated tasks across 38 disciplines.
- Task data
- Availability not established: tasks contain multi-constraint literature intents, target paper sets and expert checklist annotations, but a public task/answer bundle was not confirmed.
- Evaluation code
- Public PaSaMaster system and retrieval code are available; benchmark-specific scorer and data files were not verified, so evaluation release status remains uncertain.
- Judging criteria
- Expert checklists determine whether retrieved papers satisfy the full intent; top-20 set and ranking metrics compare results with target papers, while source checks support hallucination scoring.
- Reproducibility and barriers
- Assessment: live scholarly sources, proprietary model/search comparisons and a potentially unreleased benchmark bundle impede exact reproduction, although system code is public.
- Who reported the results
- Reported results are the paper authors’ runs; no independent reproduction was established.
Evidence supplied by Exa
- https://arxiv.org/abs/2605.14306 (new tab)
2026 paper identity, benchmark definition, 244 tasks, 38 disciplines, construction, metrics and author results.
- https://arxiv.org/abs/2605.14306 (new tab)
Benchmark protocol, checklist-based ground truth, top-20 metrics and reliability measures.
- https://github.com/sjtu-sai-agents/PaSaMaster (new tab)
Official system repository and release link; benchmark data/evaluator availability was not verified.
43. AutoResearchBenchscientific literature deep identification and comprehensive set discovery
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- scientific literature deep identification and comprehensive set discovery
- Primary link
- https://arxiv.org/abs/2604.25256 (new tab)
- Release or update evidence
- arXiv 2604.25256 (2026) falls within the window; official project resources publicly release code and benchmark data.
- What it measures
- Deep research uses exact-answer accuracy; wide research uses set IoU for coverage and precision. The benchmark contains 1,000 problems: 600 deep and 400 wide.
- Task data
- Public obfuscated JSONL bundle on Hugging Face, decrypted locally; it covers eight CS areas and contains human-verified deep and wide literature-search tasks.
- Evaluation code
- Public GitHub code includes inference, academic/web search tools, prompts, utilities, decryption scripts and separate deep- and wide-search evaluators.
- Judging criteria
- Deep answers receive exact-match scoring; wide answers are scored by IoU against gold paper sets using a standardized ReAct agent and DeepXiv search setup.
- Reproducibility and barriers
- Public code, data and evaluator support reproduction; decryption, model/API credentials, search access, token cost and live-search drift remain practical barriers.
- Who reported the results
- Results are the original authors’ baseline runs across more than 10 models and agents; independent reproduction was not established.
Evidence supplied by Exa
- https://arxiv.org/abs/2604.25256 (new tab)
Paper date, task paradigms, 1,000-query composition, construction, metrics and reported evaluation.
- https://github.com/CherYou/AutoResearchBench (new tab)
Official inference/evaluation code, decryptor, task mapping and Hugging Face data instructions.
- https://cheryou.github.io/autoresearchbench.github.io/ (new tab)
Official project page, split, schema, answers, license and public evaluation workflow.
- https://huggingface.co/datasets/Lk123/AutoResearchBench (new tab)
Officially referenced host for the obfuscated benchmark bundle.
44. DRBENCHER (Deep Research Benchmarker)entity identification + property retrieval + quantitative computation for web research agents
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- entity identification + property retrieval + quantitative computation for web research agents
- Primary link
- https://arxiv.org/abs/2604.09251 (new tab)
- Release or update evidence
- arXiv first posted 2026-04-10 (2604.09251; v3 rendered 2026-08-09), within cutoff; IBM Research and the paper link the public implementation.
- What it measures
- Measures answer accuracy, entity identification, validity, human quality and semantic diversity for multi-hop entity, property-retrieval and calculation tasks; difficulty is summarized by CCI.
- Task data
- Public JSONL tasks are available. The paper describes 268 human-validated questions, while repository documentation lists a 255-question main evaluation set; pin versions rather than treating these counts as interchangeable.
- Evaluation code
- Public MIT-licensed IBM/DrBencher repository includes pipeline, schemas, released JSONL data, decryption/evaluation utilities and computation-based gold-answer checking.
- Judging criteria
- Programmatic checks require reproducible calculations, supported clues, no entity leakage and unambiguous questions; answers use domain-specific tolerances, while humans assess validity.
- Reproducibility and barriers
- Public code, schema and data support reproduction; live Wikidata/Wikipedia and domain APIs, changing values, model/API access and stale data remain barriers.
- Who reported the results
- Results are IBM authors’ human annotations and six-model evaluations, not independent reproductions; the paper identifies property retrieval as the dominant failure mode.
Evidence supplied by Exa
- https://research.ibm.com/publications/drbencher-can-your-agent-identify-the-entity-retrieve-its-properties-and-do-the-math (new tab)
IBM Research publication page, benchmark purpose, domains, validity, accuracy and release links.
- https://arxiv.org/html/2604.09251v3 (new tab)
Answer-first generation, CCI, sources, validation, metrics, human/model results and error analysis.
- https://github.com/IBM/DrBencher (new tab)
Primary code/data repository, released QA files, schema, license and reproducible output fields.
45. Reka Research-Eval (including the ndurner evaluator extension)search-augmented web research / grounded multi-hop question answering (static QA-style task suite, not comprehensive list-building)
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- search-augmented web research / grounded multi-hop question answering (static QA-style task suite, not comprehensive list-building)
- Primary link
- https://github.com/reka-ai/research-eval (new tab)
- Release or update evidence
- Public release announced by Reka on 2025-08-28; repository created 2025-08-27, within the 2025-01-01–2026-09-26 window.
- What it measures
- Measures checklist-based answer accuracy for grounded multi-hop web questions, with aggregate mean accuracy and cost per 1,000 requests; it does not measure exhaustive recall or report quality.
- Task data
- Public encrypted dataset of 374 questions with correctness checklists, constructed through generation, annotation, consensus refinement and filtering. Answering may use live web sources.
- Evaluation code
- Public Reka dataset and generation, scoring, and analysis scripts; ndurner/web-research-eval is a fork extending provider/model support on the same task suite, not a new benchmark.
- Judging criteria
- An LLM judge checks each answer against its question-specific checklist, and aggregate accuracy is the resulting score; source quality and citation entailment are not scored.
- Reproducibility and barriers
- Requires provider credentials, live search and an LLM judge. The Modified MIT license prohibits redistributing decrypted data. Pin fork, model, provider and run settings; author and third-party runs are not automatically comparable.
- Who reported the results
- Reka launch scores are benchmark-author/vendor results. Nils Durner reports separate third-party extension runs; organizational or financial independence was not established.
Evidence supplied by Exa
- https://reka.ai/news/introducing-research-eval-a-benchmark-for-search-augmented-llms (new tab)
Dated release, 374 questions, checklist judging, six-stage annotation process, and author-reported results.
- https://github.com/reka-ai/research-eval (new tab)
Upstream repository establishes the original task/evaluator provenance and distinguishes the fork from the source suite.
- https://github.com/ndurner/web-research-eval (new tab)
Fork README identifies the inherited benchmark, added APIs, workflow, five-run leaderboard, Durner-marked rows, and reproduction script.
- https://ndurner.github.io/reka-websearch-benchmark (new tab)
Author's extension page explains that Reka released the 374-question benchmark and that the fork adds OpenAI Responses API and Exa Answers comparisons.
46. VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wildinteractive web research/discovery with evolving intent, multi-turn proactive search, and structured evidence enrichment
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- interactive web research/discovery with evolving intent, multi-turn proactive search, and structured evidence enrichment
- Primary link
- https://github.com/VibeBench/VibeSearchBench (new tab)
- Release or update evidence
- arXiv paper published/submitted May 2026 (arXiv:2605.27882); GitHub repository public May 20, 2026, within the window.
- What it measures
- Measures node and knowledge-graph triplet precision, recall, and F1 under average-at-N and best-at-N aggregation, evaluating discovery and structured enrichment.
- Task data
- Public dataset of 200 manually curated bilingual tasks across 20 domains: 100 professional and 100 daily scenarios, evenly split between Chinese and English, with expert-annotated ground-truth graphs.
- Evaluation code
- Official repository includes public task JSON, agent implementations, search/visit/Python toolkits, user simulation, evaluation and grading modules, compatible judges, and inference/evaluation scripts.
- Judging criteria
- Two-phase LLM graph matching handles aliases and translations, then semantic relations; recall allows direct, subsuming, collective, or compositional coverage, while precision counts covered predicted triples.
- Reproducibility and barriers
- Tasks, code, and evaluator are public, but execution requires LLM/search credentials and substantial live multi-turn inference; provider drift and public ground truth create assessment barriers.
- Who reported the results
- Paper authors evaluate seven frontier models with ReAct and OpenClaw and report F1 and ablations; no independent reproduction was established.
Evidence supplied by Exa
- https://arxiv.org/html/2605.27882 (new tab)
Primary paper: date, 200 bilingual tasks/20 domains, annotation and dual review, simulator, graph-matching evaluator, metrics, experimental setup, and reported results.
- https://github.com/VibeBench/VibeSearchBench (new tab)
Official implementation: task files, agent/tool code, eval/grader/evaluator modules, run scripts, dataset fields, and metric definitions.
- https://vibebench.github.io/VibeSearchBench.github.io/index.html (new tab)
Official project page: proactive-search framing, professional/daily task subsets, multi-turn tools, persona simulator, and graph-F1 evaluation.
47. DeepWideSearchdeep-and-wide agentic information-seeking benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- deep-and-wide agentic information-seeking benchmark
- Primary link
- https://arxiv.org/abs/2510.20168 (new tab)
- Release or update evidence
- Submitted/published October 2025, within cutoff; this is the canonical DeepWideSearch paper, not the later Table-as-Search paper.
- What it measures
- Measures structured-table retrieval using exact task success, row/item F1, and core-entity accuracy across repeated runs.
- Task data
- Public dataset of 220 bilingual English/Chinese questions across 15 domains: 85 Deep2Wide and 135 Wide2Deep, averaging 414.10 information units and 4.21 reasoning depth.
- Evaluation code
- Public data and evaluation code are available in https://github.com/AIDC-AI/Marco-Search-Agent (new tab) under Marco-DeepResearch-Family/DeepWideSearch/data, eval, and scripts; a public HF artifact also exists.
- Judging criteria
- Human-verified table ground truth is used; exact success requires all rows, columns, and values, while row/item F1 and core-entity accuracy provide partial scores.
- Reproducibility and barriers
- Public data and evaluation scripts support reproduction, but live retrieval, multi-run cost, and human annotation remain assessment barriers.
- Who reported the results
- Authors report four-run system comparisons; no independent reproduction was established.
Evidence supplied by Exa
- https://arxiv.org/html/2510.20168 (new tab)
Canonical paper, 220 count, 15 domains, split composition, metrics, and results.
- https://github.com/AIDC-AI/Marco-Search-Agent (new tab)
Repository tree showing DeepWideSearch data/eval/scripts.
- https://huggingface.co/datasets/AIDC-AI/DeepWideSearch (new tab)
Public 220-question dataset and reproduction instructions.
48. BrowseComp-VLmultimodal web discovery (introduced with WebWatcher)
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- multimodal web discovery (introduced with WebWatcher)
- Primary link
- https://arxiv.org/abs/2508.05748 (new tab)
- Release or update evidence
- 2025-08: WebWatcher paper introduces BrowseComp-VL; official repository documents benchmark-specific inference/evaluation.
- What it measures
- Measures multimodal, multi-hop web information seeking through final-answer accuracy/Pass@1, including identification of obfuscated entities from images and textual clues.
- Task data
- Paper describes 199 level-1 and 200 level-2 image/question pairs. Task-bundle availability is partial: the repository expects JSONL files and separately downloaded images.
- Evaluation code
- Public WebWatcher repository documents benchmark inference and evaluation scripts; it provides an agent harness rather than a separately packaged benchmark evaluator.
- Judging criteria
- Scores reference-answer correctness separately for the two difficulty levels; the exact judging configuration was not verified.
- Reproducibility and barriers
- Assessment barriers include image or OSS download failures, required local evaluation data, live search/image retrieval, model access, and agent dependencies.
- Who reported the results
- Benchmark-author/Alibaba WebWatcher experiments; distinct from the separately authored MM-BrowseComp and BrowseComp-V³. No independent rerun was established.
Evidence supplied by Exa
- https://arxiv.org/pdf/2508.05748 (new tab)
Primary paper introduces BrowseComp-VL, task construction, level counts and original comparisons.
- https://github.com/alibaba-nlp/deepresearch/blob/main/WebAgent/WebWatcher/README.md (new tab)
Official benchmark options, image-download caveat and inference/evaluation instructions.
49. Deep Research Comparatorhuman evaluation framework for research reports and intermediate steps
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- human evaluation framework for research reports and intermediate steps
- Primary link
- https://github.com/cxcscmu/Deep-Research-Comparator (new tab)
- Release or update evidence
- 2025-07: paper arXiv:2507.05495; official MIT repository created 2025-07-05.
- What it measures
- Measures side-by-side report preferences, intermediate-step quality, and text-span feedback; it is not a fixed answer-key benchmark.
- Task data
- Availability is partial: the paper reports 176 user queries, 17 annotators, and three agents, but public release of the collected annotation corpus was not established and is promised.
- Evaluation code
- Public frontend, backend, agent-service integration, and configuration instructions establish platform-code availability, not release of the collected study data.
- Judging criteria
- Human pairwise preferences produce outcome rankings, while up/down votes on intermediate steps and report spans provide process-level feedback.
- Reproducibility and barriers
- Requires human annotators, Python 3.12, Node.js 18+, PostgreSQL, and agent API/search credentials; live sources and rater variation prevent deterministic replay.
- Who reported the results
- Authors’ proof-of-concept user study; no independent replication was established. Simple Deepresearch is the accompanying agent scaffold, not a separate benchmark.
Evidence supplied by Exa
- https://arxiv.org/html/2507.05495v1 (new tab)
Evaluation design, dated paper, study size and data-release promise.
- https://github.com/cxcscmu/Deep-Research-Comparator (new tab)
Actual public platform repository, creation date, dependencies and setup.
50. Cross-Lingual BrowseComp-Plus (XBCP)controlled deep-research retrieval and evidence-grounded answering; multilingual extension distinct from BrowseComp-Plus
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- controlled deep-research retrieval and evidence-grounded answering; multilingual extension distinct from BrowseComp-Plus
- Primary link
- https://arxiv.org/abs/2606.15345 (new tab)
- Release or update evidence
- Submitted June 13, 2026 (v2 June 17, 2026), within cutoff; official repository created June 13, 2026.
- What it measures
- Measures end-to-end answer accuracy, gold-evidence recall, search-call cost, calibration, citation coverage/precision/recall, oracle-retrieval accuracy, and language/retrieval gaps.
- Task data
- Public availability is established through the repository and linked HF artifacts. It preserves English questions and answers while translating evidence documents into 12 languages, with cross-lingual and multilingual configurations.
- Evaluation code
- Public MIT repository includes preparation, translation, indexing, agent-running, oracle, LLM-judge, and per-language evaluation scripts; this is actual released code, using GPT-5.4 through OpenRouter.
- Judging criteria
- LLM judges final-answer correctness; evidence recall is computed against gold documents, citation metrics assess attribution, and oracle settings isolate retrieval from language-mismatch integration.
- Reproducibility and barriers
- Public code claims a complete pipeline, but reproduction requires decrypted original BrowseComp-Plus material, HF downloads, large indexes, APIs, and live models. The HF viewer reports a schema error.
- Who reported the results
- Authors report runs across four agents and several retrievers, including translated-evidence accuracy drops and weaker evidence and citation performance; no independent validation identified.
Evidence supplied by Exa
- https://arxiv.org/html/2606.15345 (new tab)
Paper: date, benchmark distinction, evidence construction, languages, task counts, metrics, oracle analysis and results.
- https://github.com/paddler2022/XBCP (new tab)
Repository: downloads, indexing, agent, oracle, evaluation scripts and artifact layout.
- https://huggingface.co/datasets/UTokyo-Yokoya-Lab/XBCP (new tab)
Dataset location; viewer reports schema-generation error, not absent files.
51. K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean ContextsKorean difficult web-discovery and short-answer browsing benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Korean difficult web-discovery and short-answer browsing benchmark
- Primary link
- https://arxiv.org/abs/2606.02404 (new tab)
- Release or update evidence
- ArXiv v1 dated June 1, 2026, within cutoff; repository published May 31, 2026 and dataset card is public.
- What it measures
- Measures Pass@1 answer accuracy and calibration on 300 verified items, separate synthetic-item accuracy, and trajectory/search-call behavior for multi-hop or branching questions.
- Task data
- Public dataset availability is established. It contains 400 items: 300 Korean-speaker-validated handcrafted problems and 100 synthetic diagnostic items, with answers, URLs, trajectories, checklists and metadata.
- Evaluation code
- Public official repository includes generation and runtime evaluation code, search-evals harness, dataset loading and fallback JSONL; it uses Perplexity Search API and GPT-5.4-mini extraction.
- Judging criteria
- Scores single-run Pass@1 against short gold answers and reports calibration; synthetic results remain separate, while trajectories and checklists support diagnostic analysis.
- Reproducibility and barriers
- Public MIT code and data are documented, but API/model access and changing Korean web results are required. Gitignored seed material limits exact reconstruction of synthetic generation.
- Who reported the results
- Authors report single-run evaluations using a common Perplexity pipeline and diagnose termination, trajectory, candidate-management and constraint-tracking failures; no independent results identified.
Evidence supplied by Exa
- https://arxiv.org/html/2606.02404 (new tab)
Paper: date, 400-item design, validation, metrics, protocol and author results.
- https://github.com/prometheus-eval/K-BrowseComp (new tab)
Official code: evaluation, generation, loading, fallback data and trajectory fields.
- https://huggingface.co/datasets/prometheus-eval/k-browsecomp (new tab)
Public dataset: splits, sizes, source metadata, trajectories and license.
52. DR-Arena: an Automated Evaluation Framework for Deep Research AgentsAutomated dynamic deep-research evaluation of depth and breadth
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Automated dynamic deep-research evaluation of depth and breadth
- Primary link
- https://aclanthology.org/2026.acl-long.1249/ (new tab)
- Release or update evidence
- Public arXiv release January 15, 2026; ACL publication July 2–7, 2026, within cutoff.
- What it measures
- Measures reasoning depth through tree deduction, coverage breadth through aggregation, and pairwise win/Elo; the paper also reports correlation with LMSYS Search Arena.
- Task data
- Public retained dataset availability is established: 30 evaluation trees are included in the repository. New trees can be generated by crawling current web trends.
- Evaluation code
- Public GitHub repository includes arena logic, tree generation/crawling, Examiner question generation and judging, scoring/Elo, and the retained 30-tree dataset.
- Judging criteria
- An automated Examiner builds source-grounded rubrics and judges answers and reports for evidence-based correctness; adaptive evaluation escalates depth or breadth, with reported human validation.
- Reproducibility and barriers
- The retained trees support fixed comparisons, whereas live-tree generation is time-sensitive. Reproduction requires model, search and API access, and results may change with web or Examiner updates.
- Who reported the results
- Six-model results and LMSYS correlation are original author and benchmark runs; LMSYS Search Arena supplies the external human-comparison reference.
Evidence supplied by Exa
- https://aclanthology.org/2026.acl-long.1249/ (new tab)
ACL publication date, framework, measures and reported correlation.
- https://arxiv.org/html/2601.10504v1 (new tab)
Release date, pipeline, Examiner judging and evaluation details.
- https://github.com/iNLP-Lab/DR-Arena (new tab)
Public implementation and retained 30-tree dataset location.
53. Personalized Deep Research Bench (PDR-Bench), in Towards Personalized Deep Research: Benchmarks and EvaluationsPersonalized deep-research report benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Personalized deep-research report benchmark
- Primary link
- https://arxiv.org/abs/2509.25106 (new tab)
- Release or update evidence
- ArXiv v1 submitted September 29, 2025 (revised March 4, 2026), within cutoff; repository and HF dataset are public.
- What it measures
- Measures personalization alignment across goal, content, presentation and actionability; content quality through depth, insight, coherence and clarity; and factual reliability through accuracy and citation coverage.
- Task data
- Public data availability is established. It contains 50 tasks, 25 structured or dynamic user profiles and 250 bilingual task–user queries; authors evaluate a 150-query subset.
- Evaluation code
- Public GitHub repository provides run and evaluation scripts and result directories, while the HF dataset is public; execution requires model and search API keys.
- Judging criteria
- GPT-5 judges personalization and quality; GPT-5-mini judges reliability. Criteria are dynamically generated, while reliability extracts claims and verifies retrieval and citation support.
- Reproducibility and barriers
- Queries, data and scripts are public, but proprietary judges, retrieval configuration, API access, dynamic context and live verification create practical reproducibility barriers.
- Who reported the results
- Comparisons cover commercial and open-source agents, search-augmented models and memory systems, but results are author-run; no independent benchmark results verified.
Evidence supplied by Exa
- https://arxiv.org/abs/2509.25106 (new tab)
Submission and revision dates and benchmark scope.
- https://github.com/OPPO-PersonalAI/PersonalizedDeepResearchBench (new tab)
Task/profile counts, data paths, scripts, evaluators and API requirements.
- https://huggingface.co/datasets/PersonalAILab/PersonalizedDeepResearchBench (new tab)
Public dataset release artifact.
54. Dr. Bench (formerly Rigorous Bench; Yao et al.)Expert-curated long-form report benchmark
EXA-GENERATED RECORD · NOT INDEPENDENTLY VERIFIED
Link to this record- Category
- Expert-curated long-form report benchmark
- Primary link
- https://arxiv.org/abs/2510.02190 (new tab)
- Release or update evidence
- Released October 2, 2025 through arXiv and repository; the October 2025 OpenReview submission used the earlier Rigorous Bench title.
- What it measures
- Measures semantic quality, topical drift, retrieval trustworthiness, contribution per token and retrieval index across 214 expert-curated queries in 10 domains.
- Task data
- Availability is partial: a public repository exists, but complete downloadable task and reference coverage was not verified. The paper describes manually constructed reference bundles.
- Evaluation code
- Public repository availability is established, but inspected materials provide an abstract and clone instructions only; a turnkey evaluator was not established, and the paper’s framework is not released scoring code.
- Judging criteria
- Semantic quality uses query-specific and general rubrics; topical focus penalizes missing or deviating anchor terms; retrieval trustworthiness checks exact and hostname matches against curated links.
- Reproducibility and barriers
- Explicit formulas and curated references support offline scoring, but report parsing, link extraction and LLM rubric judgments introduce implementation and model sensitivity.
- Who reported the results
- Thirteen-model comparisons and reported human agreement are authors' experiments; no independent reproduction verified.
Evidence supplied by Exa
- https://openreview.net/forum?id=EYUG4Su6ZU (new tab)
Title, October 8, 2025 public record, ICLR submission and 214-query abstract.
- https://arxiv.org/abs/2510.02190 (new tab)
Paper identity, date and canonical arXiv record.
- https://ar5iv.labs.arxiv.org/html/2510.02190 (new tab)
Reference bundles, formulas, metrics and repository link.
- https://github.com/evigbyen/rigorousbench/ (new tab)
Repository now uses Dr. Bench name; README does not establish turnkey evaluator availability.
No matching catalogue records. Clear the search to see all entries.
30 adjacent, uncertain, or excluded entries — Exa's classifications
The catalogue search above does not filter this separate group.
Tavily search-evals
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/tavily-ai/tavily-search-evals (new tab)Public provider-comparison evaluation framework (SimpleQA/document relevance) released in the requested period, but not primarily comprehensive list-building or entity enrichment.
WebChoreArena
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://doi.org/10.48550/arxiv.2506.01952 (new tab)June 2, 2025 benchmark with 532 human-curated tasks in simulated WebArena sites, emphasizing memory, calculation and tedious browser operations; functional task success, not web research or citation-backed synthesis.
AgentRewardBench
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://doi.org/10.48550/arxiv.2504.08942 (new tab)April 11, 2025 benchmark of 1,302 web-agent trajectories for comparing automatic judges across five existing benchmarks; valuable evaluation-method artifact, but it evaluates judges/trajectories rather than research browsing itself.
WebVoyager updates / Surfer-H evaluation
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://doi.org/10.48550/arxiv.2506.02865 (new tab)June 3, 2025 paper reports a 92.2% WebVoyager run and introduces WebVoyagerExtended (15,000 synthetic tasks/330 sites), but this is primarily an agent/model paper and browser-action benchmark update, not a new research-centric benchmark.
BrowserArena
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://doi.org/10.48550/arxiv.2510.02418 (new tab)Adjacent browser-action evaluation. October 2, 2025 paper introduces live user-submitted web-navigation tasks, pairwise comparisons and step-level human feedback. The date is eligible, but navigation success and failure analysis—not research completeness or evidence-grounded enrichment—are its main target.
Perplexity search_evals
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/perplexityai/search_evals (new tab)Adjacent search-provider evaluation framework, released with the September 25, 2025 Search API report. Public runner, graders and traces reuse SimpleQA, FRAMES, BrowseComp and HLE. The associated provider comparisons are Perplexity/vendor runs, not new benchmarks or independent reproductions.
Parallel Task API DeepSearchQA evaluation
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/parallel-web/parallel-llms-txt/blob/f6b31ffe/public/blog/deepsearch-qa.md (new tab)2026 vendor report of running Google DeepSearchQA; not a new benchmark. Treat the reported results as Parallel’s runs, not independent reproduction.
FACTS Grounding (v1; later Grounding v2)
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://www.kaggle.com/benchmarks/google/facts-grounding (new tab)Static long-context grounded-answer evaluation rather than agentic web research; Grounding v2 should not be conflated with FACTS Search.
Online-Mind2Web
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/OSU-NLP-Group/Online-Mind2Web (new tab)Browser task/action completion on live websites, not primarily multi-source research, exhaustive discovery or evidence-supported enrichment.
FRAMES (v3 / 2025 update)
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://arxiv.org/abs/2409.12941 (new tab)The original 824-question multi-hop Wikipedia benchmark predates 2025. A January 24, 2025 paper revision alone does not establish a substantive benchmark release. A 2026 community evaluator exists, but its provenance should not be conflated with an official new benchmark.
R2MED
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://arxiv.org/abs/2505.14558 (new tab)Medical retrieval benchmark with public query/corpus/qrels and strong reasoning-centric design, but primarily a static closed-corpus retrieval benchmark rather than live web/literature research; excluded per scope.
Talc-AI SearchBench
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/Talc-AI/search-bench (new tab)Substantive public benchmark repository with 900 manually filtered Q&A items, four realistic categories and LLM-as-judge methodology, but its release/results are 2024 (scores as of 2024-08-30; launch 2024-09-19), outside the requested 2025-01-01–2026-09-26 eligibility window. It should not be counted as a qualifying record absent a qualifying 2025–26 update.
WebWatcher
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/Alibaba-NLP/DeepResearch/tree/main/WebAgent/WebWatcher (new tab)Agent/model project, not a second benchmark record. Its August 2025 paper introduces BrowseComp-VL, catalogued separately; other evaluation runs reuse existing datasets.
WebQuest
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/google-deepmind/webquest (new tab)Multimodal web-UI/page-sequence QA benchmark with public repository, but primary paper/repository are 2024 (outside eligibility window); static/browser-UI QA adjacent rather than core research/list-building.
DRBench (ServiceNow)
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/ServiceNow/drbench (new tab)Retained as an adjacent enterprise/internal-search benchmark: it does include public web sources, but evaluates cross-application private synthetic enterprise evidence as a defining requirement.
ARC-Bench
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/aiming-lab/AutoResearchClaw/tree/main/experiments/arc_bench (new tab)Scientific autonomous experimentation benchmark; research is explicit but web research is not the benchmark’s target.
Entity Enricher platform model benchmarks and benchmark scoring
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://entityenricher.ai/docs/platform/benchmarks (new tab)Vendor evaluation documentation for organization-specific saved entity/schema scenarios and model comparisons, not an established public dated benchmark release. Reference-based completeness/correctness scoring is described, but public tasks, results and evaluator artifacts were not verified.
Exa Websets benchmark
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://exa.ai/blog/websets-evals (new tab)Keep adjacent/vendor-only. The official February 19, 2025 post reports Exa's own comparison over 200 generated queries and GPT-4o grading, but no public raw dataset, evaluator or code was established; no independent validation should be implied.
Parallel FindAll 40-query benchmark
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://parallel.ai/products/findall (new tab)Vendor-only evaluation report: 40 discovery/enrichment queries, with recall measured against pooled correct matches from compared systems. Public task, gold and scorer artifacts and a firm release date were not established. Results are Parallel-created and reported, not independent.
Reka Vibe-Eval
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/reka-ai/reka-vibe-eval (new tab)Not eligible: multimodal chat benchmark repository was created in 2024 and evaluates image/multimodal generations, not web research; it is a false-positive despite the shared word 'Vibe'.
Vals Web Search Index
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://www.vals.ai/benchmarks/web_search (new tab)July 16, 2026 controlled search-tool comparison on legal research and finance tasks. It reports 208 legal tasks and 450 finance questions, with rubric-based grading. Keep as a public comparative report: a standalone reusable task/evaluator package and organizational independence were not established; linked orchestration repositories alone do not establish open task data.
EntiWeave
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://huggingface.co/datasets/jxg25/EntiWeave (new tab)Public graph-grounded web-search training data and a 100-question held-out evaluation split were found, but a qualifying release/update date and benchmark evaluator were not established. Retained as uncertain, not counted as dated core coverage.
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://aclanthology.org/2025.findings-acl.988/ (new tab)ACL Findings 2025 multilingual interactive-web benchmark. Evaluates instruction following and product/website interaction using WebShop-style task success, not research discovery or citation-backed synthesis. Public paper verified; data/evaluator release not established here.
Search Arena
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://arxiv.org/abs/2506.05334 (new tab)2025 public search-LLM preference platform with conversation/vote data and analysis code. Kept adjacent because its primary outcome is pairwise user preference on search chats, rather than objective discovery correctness, evidence coverage or set completeness. Votes were collected by the authors; they are not independent reproduction.
DR³-Eval
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://arxiv.org/abs/2604.14683 (new tab)April 2026 public code/data benchmark for multimodal, multi-file research reports in static task sandboxes. Scores information recall, factual accuracy, citations, instruction following and depth. Adjacent because supplied files/controlled workspaces, rather than open-web discovery, define its tasks.
Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://arxiv.org/abs/2602.15019 (new tab)Uncertain standalone benchmark: February 16, 2026 drug-asset scouting paper describes 48 seed queries and 22 held-out query–asset pairs, with asset precision/recall/F1 and expert-calibrated LLM grading. Results are author-run; public task, gold and evaluator releases were not verified. Retained as a paper-defined evaluation study, not a confirmed reusable public benchmark.
GAIA
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://ai.meta.com/research/publications/gaia-a-benchmark-for-general-ai-assistants/ (new tab)Relevant general-assistant benchmark with web browsing, but the original paper/release is 2023–2024. Later agent runs or leaderboard submissions do not themselves demonstrate a substantive 2025–September 2026 benchmark update; none was established here.
Humanity’s Last Exam (HLE)
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://arxiv.org/abs/2501.14249 (new tab)Eligible 2025 release, but primarily closed-ended expert academic questions testing broad knowledge/reasoning, not open-web research, complete entity discovery or evidence-supported enrichment. Agent papers using search on HLE are benchmark runs, not new research benchmarks.
SimpleQA Verified
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://arxiv.org/html/2509.07968v1 (new tab)September 9, 2025 factuality benchmark designed to measure parametric knowledge without tools. Kept separate from agentic web research despite use of related SimpleQA datasets in search-provider evaluations.
SciArena and SciArena-Eval
EXA-GENERATED CLASSIFICATION · NOT INDEPENDENTLY VERIFIED
https://github.com/yale-nlp/SciArena (new tab)Public July 1, 2025 release of platform, preference data and analysis code; later paper includes a meta-evaluation benchmark. Adjacent: primarily compares foundation-model literature-grounded responses and judge agreement with human votes, rather than autonomous discovery or complete-set retrieval. Citation-attribution analysis is relevant, and paper-bank data are available, but preferences are not evidence-completeness scores.
Inspect the original materials
The exact request includes the research question, additional instructions, requested output schema, effort: "ultra", and an empty Connect data-source list. The downloadable response preserves Exa's output, grounding, usage, and cost fields. Only the operational top-level run identifier was removed for publication; JSON indentation may differ. No research field was edited.
The compact provenance file records this redaction and SHA-256 hashes of the public files. These hashes help check file identity; they do not establish factual correctness. Future editorial corrections should be distinguished from the preserved provider response.
Exact research question sent to Exa
ASSISTANT-AUTHORED ASSIGNMENT
Conduct a standalone evidence-backed catalogue of public benchmarks and evaluation frameworks testing AI agents on web research, comprehensive list-building, or evidence-supported entity enrichment. Include projects with a public release or substantive update between January 1, 2025 and September 26, 2026. Seek broad coverage of this defined scope, not an arbitrary top ten. Distinguish research/discovery evaluations from adjacent browser-action or generic QA benchmarks; put adjacent or uncertain cases separately. For each qualifying project identify canonical name, primary project/paper/repository URLs, what ability it measures, dated evidence of eligibility, public availability of task data, evaluation code and judging criteria, and practical reproducibility barriers (dependencies, unavailable components, changing web state). Distinguish benchmark-author, vendor, and independent reported results without inventing independence. Support material field claims with primary-source URLs and short evidence notes. Unknown is acceptable; avoid inferring released code from a paper's promise. Deduplicate renamed projects and distinguish benchmarks from reports of running them. Return concise structured records and a short synthesis of coverage, limitations, and major exclusions. Do not run benchmark code or claim reproduced evaluations. Use only public web research; do not enable or use Exa Connect, premium connected datasets, contact enrichment, or purchases. This is one Ultra run under the documented default US$20 run cap; do not create follow-up runs. Keep the assessment neutral across vendors. No personal or private project information is relevant.
Additional instructions sent to Exa
ASSISTANT-AUTHORED INSTRUCTIONS
Research using public web sources only. No connected datasets or contact enrichment. Prioritize primary papers, repositories, official benchmark sites and documented artifacts. Be concise and evidence-specific. Distinguish observed source facts from your judgments. Do not equate broad search with proven exhaustive coverage. Provide factual findings without promotional framing.
Requested output structure
ASSISTANT-DESIGNED SCHEMA
{
"type": "object",
"properties": {
"scope": {
"type": "string"
},
"benchmarks": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {
"type": "string"
},
"canonical_url": {
"type": "string"
},
"category": {
"type": "string"
},
"eligibility_date_and_evidence": {
"type": "string"
},
"measures": {
"type": "string"
},
"task_data": {
"type": "string"
},
"evaluation_code": {
"type": "string"
},
"judging_criteria": {
"type": "string"
},
"reproducibility_and_barriers": {
"type": "string"
},
"reported_results_provenance": {
"type": "string"
},
"evidence": {
"type": "array",
"items": {
"type": "object",
"properties": {
"url": {
"type": "string"
},
"supports": {
"type": "string"
}
},
"required": [
"url",
"supports"
]
}
}
},
"required": [
"name",
"canonical_url",
"category",
"eligibility_date_and_evidence",
"measures",
"task_data",
"evaluation_code",
"judging_criteria",
"reproducibility_and_barriers",
"reported_results_provenance",
"evidence"
]
}
},
"adjacent_or_uncertain": {
"type": "array",
"items": {
"type": "object",
"properties": {
"name": {
"type": "string"
},
"url": {
"type": "string"
},
"reason": {
"type": "string"
}
},
"required": [
"name",
"url",
"reason"
]
}
},
"summary": {
"type": "string"
},
"coverage_gaps": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"scope",
"benchmarks",
"adjacent_or_uncertain",
"summary",
"coverage_gaps"
]
}- Download exact request (JSON)
- Download Exa response (JSON; run identifier removed)
- Download provenance and file hashes (JSON)
What to take away
Use this example to inspect how delegated research works, not to choose a “winning” model. The assignment shapes the investigation; the provider makes research judgments; the assistant makes further presentation choices. Keeping those layers visible gives readers a basis for their own assessment.