serp.fast

Scores how web search tools lift agent accuracy on finance and legal research

Nathan Kessler
By Nathan KesslerUpdated

Each tool is evaluated against our methodology using public docs, vendor demos, and hands-on testing.

What is Vals AI Web Search?

Independent benchmark measuring how much web search tools like Exa lift AI agent accuracy over native provider search, across 208 legal-research and 450 finance tasks.

Our verdict

One of the few benchmarks that isolates the search layer instead of the model. Vals AI keeps the model fixed and swaps only the retrieval tool, a provider's native search versus an independent API like Exa, across 208 expert-written legal-research tasks and 450 equity-research finance questions. It then fits mixed-effects models to separate each tool's contribution from raw task difficulty. The datasets are proprietary and non-public, which blunts test-set leakage but also means you can't reproduce the runs yourself. The July 2026 snapshot is a useful reality check. The best pairing, Claude Fable 5 with Exa, clears just 48.5% overall, and the search tool only moves the numbers where the task is data-heavy. On finance the gap is real: Exa beats native search by roughly 6.5 points (p < 0.001), concentrated in market and earnings analysis. On legal research the two are statistically indistinguishable (p = 0.86). If you're weighing whether to add a dedicated search API to an agent, this is a rare source that quantifies when it actually pays off.

Categories:

Benchmarks are how you separate marketing claims from measured reality. Instead of trusting vendor-reported numbers, benchmarks run the same tasks against every system under a shared methodology and publish the results. For AI product builders picking an agentic extraction or search stack, a trustworthy benchmark is a strong input to the build-vs-buy decision – and a fast way to spot when a category is still too immature to rely on.

Share:

Similar tools

See Benchmarks

Embeddings-based neural search that finds semantically related pages

FreemiumMar 2026AI-Native Search APIs

Real-time search API built for AI agents and RAG pipelines

FreemiumMar 2026AI-Native Search APIs

Benchmark testing browser agents on 153 tasks across 144 live sites

FreeApr 2026Benchmarks

How Vals AI Web Search compares

ClawBench

ClawBench is the directory's other live benchmark, but it scores browser agents on task completion rather than search-answer accuracy.

Frequently asked questions

What does the Vals AI Web Search benchmark measure?

It measures how much a web search tool improves an AI agent's answers, holding the model constant and changing only the retrieval layer. Vals AI runs the same tasks with a provider's native search and with an independent API like Exa, then compares accuracy. Coverage spans two domains: 208 expert-authored legal-research tasks and 450 finance questions modeled on equity-research work. The result is a per-tool accuracy score, not a latency or cost figure.

Is the Vals AI Web Search benchmark free and open source?

The leaderboard is free to view at vals.ai, but the benchmark itself is not open source. Vals AI builds its datasets in-house and keeps them private, partly to keep them out of model training data. You can read the published results and methodology, but you can't download the tasks or reproduce the runs yourself the way you can with an open benchmark like WebArena.

Which search tool wins the Vals AI Web Search benchmark?

In the July 2026 snapshot the strongest result is Claude Fable 5 paired with Exa, at 48.5% overall accuracy, the highest of any model-and-tool combination tested. The advantage is domain-specific: on finance tasks Exa beats native provider search by about 6.5 points with high statistical confidence, while on legal research the tools score about the same. Standings shift as Vals adds models, so check the live leaderboard for current numbers.

Does an independent search API actually beat a model's native search?

By this benchmark, it depends on the task. On data-heavy finance work an independent tool like Exa shows a clear edge, about 6.5 points, concentrated in market analysis and earnings analysis. On legal research the difference isn't statistically significant. The practical read is that a dedicated search API helps most when answers depend on fresh, specific data, and matters less when the model's built-in search already covers the ground.

Who runs Vals AI?

Vals AI is an independent evaluation company that benchmarks frontier models on real-world, domain-specific tasks across finance, law, software, and other fields. It runs its own evaluations and builds many of its benchmarks in-house, which is how it positions itself as a neutral third party rather than a model vendor. The web search benchmark is one entry in a broader suite that also covers legal, finance, and coding.

How is the Vals AI Web Search benchmark different from agent benchmarks like WebArena?

WebArena, WebVoyager, and ClawBench score whether an agent can finish multi-step tasks on websites. The Vals AI Web Search benchmark asks a narrower question: given a fixed model, how much does the search tool it calls improve the accuracy of its answers. It isolates retrieval quality rather than end-to-end task completion, which makes it the more useful reference when you're choosing a search API instead of an agent framework.

Visit

Vals AI Web Search

Visit