THE CASE STUDY / SEPARATING SEARCH FROM INTELLIGENCE
Multiple providers filled gaps.
Their contributions also overlapped.
Two assistants researched Glamsterdam. A third combined their findings into the guide you can read here.
This is a worked example of Choose your AI. Choose how it searches. The idea is simple: choose the assistant that interprets information separately from the tools that retrieve its evidence. You can understand this case study without reading the companion guide first.
The research brief asked Codex and Grok Build to investigate the practical impact of Ethereum’s Glamsterdam upgrade using Octen, Exa, Perplexity Search and authenticated X data. Claude Opus 5.5 in Claude Code then synthesized the two reports without new web research. Start with the Glamsterdam takeaways or read the full synthesis.
Combining the reports retained detail and made disagreements visible.
The synthesis did more than shorten two documents. It retained useful material from each and explained how differences were resolved. Two examples show the kind of judgment involved:
- ETH transfer logs: narrow the claim. Grok Build used the broad wording “Every ETH transfer emits a log.” Codex specified nonzero transfers between different accounts, with exclusions. Claude adopted the narrower scope and retained Grok Build’s additional detail about how the cost is accounted for.
- Gas constants: keep the qualification. Codex presented the current state-access gas table. Grok Build highlighted a statement in the companion specification that some constants were not final. Claude retained that caveat so the combined account did not imply more certainty than the inputs supported.
These are decisions documented in the synthesis’s comparison section. They illustrate editorial value from comparing the reports; the synthesis did not independently re-check the underlying sources.
Providers retrieved evidence; assistants interpreted it in two stages.
A harness is the application coordinating the work. A model reasons and writes. A provider supplies search or retrieval. Keeping those roles separate makes the process easier to inspect.
- Retrieve and research. Each research harness used the four providers under substantially the same written brief. Search results and source content informed its own report.
- Synthesize within each harness. Sol 6.1 in Codex and Grok 4.7 in Grok Build interpreted the retrieved material and wrote separate accounts.
- Synthesize across harnesses. Claude Opus 5.5 in Claude Code compared those accounts, combined useful coverage and recorded disagreements.
The provider roles in these runs concerned evidence retrieval. No Exa Agent or Perplexity Ask, Reason or Research narrative was adopted as a finding, according to the reports. Search and extraction may still use models internally. Different viewpoints came from sources and their interpretation, rather than four providers writing four competing opinions.
The ledgers record complementary contributions, with shared ground.
This assessment uses the saved Codex ledger and Grok Build ledger. It describes what those reports record; raw tool responses were not supplied for a fresh execution audit.
| Provider | Recorded contribution | Value and limits |
|---|---|---|
| Octen | Broad discovery, targeted searches and substantial full-text retrieval of specifications, scope records and releases. | A major retrieval role. In Codex, raw-source retrieval completed two EIP reads that Exa Fetch returned only as headings. |
| Exa | Semantic discovery of authoritative sources and historical records; Codex also used Fetch for full specifications. | An alternative route for particular gaps. Codex pursued an exclusion/scope gap through Exa after an Octen search returned many irrelevant results. Headline sources overlapped. |
| Perplexity Search | Targeted discovery of repricing guidance, devnet updates and removal or deferral discussions. | Useful gap-filling, but its unique contribution is harder to isolate: some documents also appeared through other providers, and fuller reads often used Octen. |
| Authenticated X API | Original developer statements, implementation concerns and design discussions. | The clearest complementary evidence category alongside specifications. Posts establish what people said, not upgrade inclusion or measured performance by themselves. |
Another retrieval route helped when the first was incomplete.
The Codex ledger records that Exa Fetch returned only headings for EIP-8045 and EIP-8061; Octen retrieval of raw sources completed those reads. It also records the reverse kind of assistance at discovery time: an Exa follow-up pursued a scope gap after an Octen search produced considerable noise. These are concrete examples of tools helping at different stages, rather than a general ranking of either provider.
Different search choices changed the evidence even within one provider.
Codex’s X searches reached earlier design debates; Grok Build’s emphasized recent announcements. The synthesis reports little overlap between the two samples. Different date windows, queries, result limits and follow-up choices broadened their combined coverage. Neither run exhausted X archive pagination. Compare the recorded search approaches →
Different indexes or rankings could also affect discovery, but these runs do not isolate those causes. We cannot attribute a result difference to an index alone.
Overlap can support checking—or become unnecessary work.
Finding the same primary document through two services can help researchers locate and compare the relevant version. It does not create two independent sources for the same claim. Both research reports often relied on the same underlying Ethereum records.
Some repetition was avoidable: Codex’s ledger explicitly records one unnecessary repeat read of the meta-EIP. Both runs also encountered stale or irrelevant results. Additional providers introduce more material to reconcile and can add service charges, latency and setup work. The records here do not quantify that overhead.
Were all four necessary? This case does not answer that. There was no comparable one-provider baseline, no complete accounting of which final claims depended uniquely on each provider, and no complete cost or timing record. The reports document useful contributions, but do not establish that four providers beat one—or that two would have been insufficient.
The model settings describe the runs, not a controlled comparison.
The reported model and reasoning settings are listed below. Execution logs were not available to independently verify them.
| Step / harness | Model | Reasoning setting |
|---|---|---|
| Research / Codex | Sol 6.1 | Medium |
| Research / Grok Build | Grok 4.7 | Medium |
| Synthesis / Claude Code | Opus 5.5 | Medium |
“Medium” is the selected setting, not evidence of equivalent reasoning effort across models. Harness behavior, available tools, retrieval results and execution choices also varied. The outputs cannot establish a model ranking.
Inspect the reports and the decisions behind the synthesis.
The Grok Build report and its download have one privacy edit: the authenticated X account identifier is anonymized. Research findings and cited researchers’ public statements are unchanged.
The prompts show how retrieval and synthesis were assigned.
The two research prompts differ in three places: two references to “Codex” become “this harness,” and the destination filename changes. The question, provider roles, evidence rules and requested coverage are otherwise identical. Their publication copies replace local output paths with filenames and omit directory-creation wording.
The Claude prompt is edited for grammar, structure and clarity, not presented as the literal wording used in the run. It preserves the request for full-depth synthesis without new web research, a one-page TL;DR, harness attribution confined to the differences section and anonymization of the requester’s X handle.
This case demonstrates a workflow and preserves its evidence limits.
The reports and synthesis reflect a 30 September 2026 snapshot. Agreement does not establish independent verification. Benchmarks, chain replays and client tests were not reproduced; source conflicts and incomplete retrievals remain documented in the research limits. Raw tool responses, full execution transcripts, complete model settings and cost records are absent from this collection.
This website adds an introductory reading route and an assessment of the recorded workflow. It does not refresh the Ethereum research or certify that the original tool calls executed as described.
Start with one provider. Add another to fill a specific gap.
The multi-provider-research skill can work with one suitable provider. Add a second when you need a different source category, an unresolved search angle or a way around incomplete retrieval. Judge it by the usable evidence it adds. If it repeatedly returns the same material without resolving a gap, more calls may not help.
The companion guide explains provider setup and skill installation. Assistant subscriptions and provider access are separate choices; you do not need to reproduce this four-provider setup to try the approach.