Skip to content
WonderSearch
ComparePricingBenchmarksBlogTeam
Create an Account ↗
WonderSearch
01Compare↗02Pricing↗03Benchmarks↗04Blog↗05Team↗06Book a call↗
Create an Account ↗

Search so good it feels like magic.

↖ Back to WonderSearch

WONDERSEARCH / BENCHMARKS

We Measured
the Magic

See how WonderSearch compares with published sparse and dense retrieval results.

nDCG@108 datasetsUpdated October 4, 2026

Small, Medium, and Large: internal benchmarks using the same harness as production, run on an NVIDIA H100. Ultra is an internal research configuration and is not currently available for public use.

THE RETRIEVAL LANDSCAPE

See where we land.

SciFact Public dataset

300 questions from the BEIR test set.

WonderSearch Ultra (research) Published reference
Compare against
nDCG@10 · higher is better
050100
Score
  • WonderSearch 1.1 UltraResearch · Not currently available for public use
    80.3995% interval
    76.75–83.99
  • WonderSearch 1.1 LargeInternal · production harness · H100
    80.0695% interval
    76.43–83.72
  • Qwen/Qwen3-Embedding-8B ↗Dense · Published MTEB results
    78.46
  • Qwen/Qwen3-Embedding-4B ↗Dense · Published MTEB results
    78.33
  • WonderSearch 1.1 MediumInternal · production harness · H100
    77.7995% interval
    73.98–81.71
  • openai/text-embedding-3-large ↗Dense · Published MTEB results
    77.77
  • WonderSearch 1.1 SmallInternal · production harness · H100
    74.6895% interval
    70.53–78.91
  • BAAI/bge-large-en-v1.5 ↗Dense · Published MTEB results
    74.64
  • thenlper/gte-large ↗Dense · Published MTEB results
    74.27
  • openai/text-embedding-3-small ↗Dense · Published MTEB results
    73.37
  • intfloat/e5-large-v2 ↗Dense · Published MTEB results
    72.24
  • ELSER v2 ↗Sparse · Elastic, October 17, 2023
    72.00
  • intfloat/multilingual-e5-large ↗Dense · Published MTEB results
    70.20
  • DeepRetrieval-3B + BM25 ↗Sparse · DeepRetrieval, Table 2 (April 2025)
    64.60
  • BAAI/bge-m3 ↗Dense · Published MTEB results
    64.37

Small, Medium, and Large use internal product-mode evaluations. Benchmarks were recorded using a copy of our production harness, and on a H100. Ultra uses research evaluation settings and is not currently available for public use. Other scores come from their linked publications. Lines on WonderSearch bars show 95% bootstrap intervals.

The benchmark that matters next?
Your data.

Book a call ↗
WonderSearch

Search so good it feels like magic.

A search product by Evokoa ↗

Product

  • Pricing
  • Compare
  • Benchmarks
  • Create an account ↗

Company

  • Team
  • Manifesto
  • Blog
  • Book a call

Legal

  • Terms of Service
  • Privacy Policy
  • Privacy choices
  • Security Disclosure
© 2026 Evokoa. All rights reserved.team@evokoa.comBack to top ↑