---
title: "How to Really Benchmark Search Engines"
author: "Max Irwin"
date: 2026-10-05
canonical_url: https://bonsai.io/blog/how-to-really-benchmark-search-engines/
license: CC BY-SA 4.0
license_url: https://creativecommons.org/licenses/by-sa/4.0/
copyright: Bonsai.io, 2026
---

Benchmarks need to measure how the technology is actually used. For search engines like OpenSearch and Elasticsearch, this means we need to be ingesting documents/vectors at the same time as querying them, and the setup should be built around a use case.

I've grown tired of benchmarks that don't mean anything. I see this all the time. They measure latency and throughput on an imaginary dataset in a clean room with narrow parameters. Benchmarks like this don't help anyone make decisions because the teams who are evaluating technology won't ever be in that scenario.

At Bonsai we manage multiple search engine technologies and versions across thousands of customers and use cases. Benchmarks came up in the news recently because two unnamed vendors were having a bit of a scuffle. I'm not picking sides here, but I'm going to trash their methodology. They tested variously optimized indices of 1M documents at search time only.

Let the trashing commence:

1.  Nobody is just searching a static/readonly index
2.  1M documents is too tiny for a meaningful assessment
3.  Only testing the optimized index is cheating
4.  A lone vector field without metadata filtering+aggs is unrealistic

If you're just trying to see how fast your search engine can query with the above, then you've lost sight of the real problems the engine should solve.

**Goodhart's Law**

When a measure becomes a target, it ceases to be a good measure.

Now, I definitely understand pushing the limit to finding the maximum capability in a test. It's fun and thrilling. The salt flat races for search engines, if you will. But most of us are just driving around in regular conditions and need to understand the practicality.

## Our goal and process

All benchmarks stem from a use case. Ours is generalized hosting of OpenSearch on EC2 instances for customers.

Last summer, we were looking into whether we could host our vector/hybrid clusters on ARM instead of x86. In AWS, the Graviton ARM instances are much less expensive (saving about 40%). We don't want to give up performance to save money and ultimately we needed to know what AWS instance types we should be using when recommending hybrid search setups. We _also_ needed to understand how OpenSearch fared against a variety of query variants. So our tests involved both.

This meant we _had_ to use a real-world scenario, otherwise our findings would be useless for a recommendation. This was an internal test - but I'm removing the embargo and opening up our results :)

We started with 30M hybrid documents. We tested a hybrid search with aggregations and filters while pummeling the index with new data across three processor types.

These were the steps:

1.  Choose several candidate instance types
2.  Standup an OpenSearch (2.19\*) cluster on an instance of type X
3.  Ingest most of the dataset into the cluster (30M documents)
4.  Start a sustained high queries-per-second (qps) benchmark experiment while avoiding cache
5.  Simultaneously write as much as possible (using the rest of the dataset) during the query load
6.  Capture the stats and notes
7.  Perform the test for several query variants
8.  Loop back to step 1 for another instance type Y
9.  Compare results after we test several instance types
10.  Choose ideal candidate for instance type, extrapolate and adjust cost expectations

### Additional notes from our process:

I was looking for noisy "worst case" things. I don't want a perfect outcome!

-   We skipped an optimization after step 3 - optimizing can speed things up after a large ingest, but unoptimized gives us a better picture of practicality.
-   I ran these tests _from my office_ to the cluster. This is because we don't know where our customers run their application! I have a fast 1GB connection (ethernet+fiber).

That being said, I also didnt delete any documents - which is an important thing that happens in the real world. If/when we rerun these tests for newer Grav5 and OS3 we'll consider including deletes.

_\*Yes, OpenSearch 2.19 - this was last summer and most of our clients still run this version_

## Findings

Instead of burying these at the bottom, here's the important things we discovered:

-   Newer ARM Gravitons (gen 4 and up) are better than x86 from a cost/value perspective. They are fast and efficient. Use them.
-   Some aggregation operations (particularly ranges) are very slow with Hybrid search on OS 2.19. Fixes for this have been rolled out in v3.2
-   Running your own fit-for-purpose benchmark is fun and worthwhile and you should do it!

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/in_benchmark_we_trust_emblem_latin_fixed.svg)

## The dataset and index

We used the now defunct wikipedia-22-12-simple-embeddings dataset from Cohere. It's a shame the dataset was taken down, but a fork/update exists hosted by [timescale](https://huggingface.co/datasets/timescale/wikipedia-22-12-simple-embeddings/blob/main/README.md).

What you really need to know is that we have a solid mix of lexical and vector features in the dataset, in the following fields:

-   id (unique)
-   title (title of the article)
-   paragraph\_id (integer between 1 through 8 - articles were chunked)
-   text (body of a chunked paragraph)
-   url (wikipedia URL)
-   wiki\_id (keyword - duplicate for article chunks)
-   views (integer - number of times the article was viewed)
-   langs (keyword - two-letter language code)
-   emb (768 dimension f32 vector)

We use these to design a real-world schema with real analyzers and optimizations, as well as a completions search-as-you-type field for the title and the phrase shingles from my [autocomplete](https://bonsai.io/blog/how-to-really-do-autocomplete/) experiments.

```json
{
  "mappings": {
    "properties": {
      "id": { "type": "keyword" },
      "title": {
        "type": "text",
        "analyzer": "analyze_english",
        "copy_to": [
          "completions",
          "suggest_word",
          "suggest_word_all",
          "suggest_phrase"
        ]
      },
      "text": {
        "type": "text",
        "analyzer": "analyze_english",
        "copy_to": ["suggest_word_all"]
      },
      "url": { "type": "keyword" },
      "wiki_id": { "type": "keyword" },
      "views": { "type": "float" },
      "paragraph_id": { "type": "integer" },
      "langs": { "type": "keyword" },
      "emb": {
        "type": "knn_vector",
        "dimension": 768,
        "method": {
          "name": "hnsw",
          "space_type": "cosinesimil",
          "parameters": {
            "ef_construction": 512,
            "m": 16
          }
        },
        "mode": "on_disk",
        "compression_level": "32x"
      },
      "completions": { "type": "search_as_you_type", "max_shingle_size": 3 },
      "suggest_word": {
        "type": "text",
        "analyzer": "analyze_suggest_word",
        "search_analyzer": "analyze_suggest_search",
        "fielddata": true
      },
      "suggest_word_all": {
        "type": "text",
        "analyzer": "analyze_suggest_word",
        "search_analyzer": "analyze_suggest_search",
        "fielddata": true
      },
      "suggest_phrase": {
        "type": "text",
        "analyzer": "analyze_suggest_phrase",
        "search_analyzer": "analyze_suggest_search",
        "fielddata": true
      }
    }
  },
  "settings": {
    "index": {
      "number_of_shards": 3,
      "number_of_replicas": 1,
      "replication.type": "SEGMENT",
      "knn": true,
      "knn.derived_source.enabled": true
    },
    "analysis": {
      "char_filter": {
        "strip_html": { "type": "html_strip" }
      },
      "filter": {
        "english_stop": { "type": "stop", "stopwords": "_english_" },
        "english_light_stem": {
          "type": "stemmer",
          "language": "light_english"
        },
        "english_possessive_stem": {
          "type": "stemmer",
          "language": "possessive_english"
        },
        "shingles": {
          "type": "shingle",
          "min_shingle_size": 2,
          "max_shingle_size": 4,
          "output_unigrams": false
        },
        "edge_ngram_filter": {
          "type": "edge_ngram",
          "min_gram": 2,
          "max_gram": 20
        }
      },
      "analyzer": {
        "analyze_english": {
          "tokenizer": "standard",
          "char_filter": ["html_strip"],
          "filter": [
            "lowercase",
            "english_possessive_stem",
            "english_stop",
            "english_light_stem"
          ]
        },
        "analyze_suggest_word": {
          "tokenizer": "standard",
          "char_filter": ["html_strip"],
          "filter": ["lowercase", "english_stop"]
        },
        "analyze_suggest_phrase": {
          "tokenizer": "standard",
          "char_filter": ["strip_html"],
          "filter": ["lowercase", "english_stop", "shingles", "unique"]
        },
        "analyze_suggest_search": {
          "tokenizer": "standard",
          "char_filter": ["strip_html"],
          "filter": ["lowercase", "english_stop"]
        }
      }
    }
  }
}
```

## The indexer

Before we test anything we need index all our starting documents, and also maximize index throughput as much as possible. We need to find the maximum here so that when we index while searching we can be confident we're doing all we can. Normally a team would use something like Kafka - but that would be overkill for our needs and take too long to setup and get right. So I rolled my own indexer. After fighting Python for speed I gave up and wrote something in Rust.

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/python-v-rust.jpg)

On average I was able to index between 5k and 6k docs per second using the Rust indexer. This gives us a very helpful datapoint as a bonus: we now know how fast we can load a 1B size dataset into an index (about 50 hours before optimization).

## The querier

When load testing for queries, we used locust.io (no affiliation). We simulated various query variants all with the same parameters of 12 'users'. Each user is a separate client to the cluster, running as many queries as possibe in sequence. In other words, one user will send a query, wait for a response, and then send the next query and there are 12 users doing this at the same time.

When a test is started, the number of users will ramp up to maximum for about 20 seconds. We ran all tests for at least 1 minute after rampup.

BTW, Locust is a Python framework and I really like it.

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/rust-v-python.jpg)

## The clusters

We tested on a trio of the following instance types:

-   i7i.2xlarge (Intel x86\_64)
-   m7gd.2xlarge (Graviton3 ARM64)
-   m8gd.2xlarge (Graviton4 ARM64)

# Results

Here are the results for the test. No need to scroll to the bottom :)

| Family | Query | Requests | Failures | qps | 50%\* | 99%\* | 99.90%\* | 100%\* | $/1M Queries |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| grav4 | hybrid | 3344 | 0 | 56.550 | 190 | 370 | 590 | 740 | $4.53 |
| grav3 | hybrid | 1921 | 0 | 32.450 | 320 | 670 | 5700 | 6000 | $7.31 |
| grav4 | hybrid\_aggs | 3265 | 0 | 55.172 | 200 | 360 | 510 | 600 | $4.65 |
| grav3 | hybrid\_aggs | 1317 | 0 | 22.257 | 480 | 1000 | 1100 | 1400 | $10.66 |
| grav4 | hybrid\_filter | 3407 | 0 | 57.594 | 180 | 340 | 460 | 500 | $4.45 |
| grav3 | hybrid\_filter | 2299 | 0 | 38.819 | 270 | 530 | 630 | 700 | $6.11 |
| grav4 | knn | 4283 | 0 | 72.395 | 150 | 230 | 260 | 280 | $3.54 |
| grav3 | knn | 2833 | 0 | 47.822 | 220 | 360 | 450 | 500 | $4.96 |
| grav4 | lexical\_aggs | 22384 | 0 | 378.633 | 26 | 60 | 100 | 230 | $0.68 |
| grav3 | lexical\_aggs | 18419 | 0 | 311.342 | 29 | 99 | 180 | 1300 | $0.76 |

_\*Response time percentiles in milliseconds_

Note that these stats are for queries _while hammering the index with new data_. Remember, I queried from my office - so the qps is a "worst case" due to bonus network latency.

# Science Notes

What follows are my raw notes as I was running the tests. Yes, I actually used those emoji as I was taking my run notes.

## i7i.2xlarge

This was the instance type we're currently recommending and want to move off of it to Graviton. These numbers are just for compare and to warm up while I get used to testing.

### Indexing

I managed to index 36M hybrid docs at 4k docs/sec from a single small client using a fast rust indexer.

Cluster stayed healthy, RAM peaked and CPU was 100% for the duration.  We stayed up and maxed it out! 

Grafana view:

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-md.png)

Extrapolating these numbers we can index 1B docs in a couple days - which helps answer our first question “is it feasible to index lots of vectors on OpenSearch all at once?“.  I expect that number to be better on a bigger cluster and a bigger ingestion client.

### Querying a static index

All queries simulate concurrent use with no rampup (Tests are from RochesterNY to us-east-1).

Concurrency can be seen in “number of users” plots. We measure successful requests per second, failed requests, p50 latency, and p95 latency. In all tests aside from #1, we randomize the query keyword and vector to avoid cache hits. Lexical randomization is a single random term from a 100k word dictionary, and Vector randomization is a random numpy 768 dim ndarray.

1.  Null hypothesis base load test (simple match\_all query for 5 concurrent users):

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-1-md.png)

2.  knn query with np.rand for the vector on each request (no caching)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-3-md.png)

15 concurrent connections, 150qps, 100ms latency (p50), 140ms latency (p95)

3.  Hybrid query (methodology: random keyword, random vector, 50/50 weighted norm, no filters, no aggs)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-4-md.png)

4.  Hybrid query with aggs (methodology: same as #3 plus aggressive aggregation on paragaph\_id (int count), terms (word cound), and terms (phrases count)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-5-md.png)

5.  Hybrid query with aggs and filter (methology: same as #4, plus filtering on paragraph\_id with random value)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-6-md.png)

Grafana view for the most aggressive queries. Shows high CPU and disk I/O due to hybrid lexical+vector+aggs:

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-7-md.png)

### Querying AND Indexing

The above were for querying on a static index. We perform the same query operations as above while heavy writes are ongoing.

Important indexing notes: all the data being indexed is always new (we’re not doing updates). replication (segment) is ON and refresh is ON (1s).

1.  Query: Null hypothesis (match all query size 1) - 195QPS, p95@22ms Index: 140k docs in 1m25s

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-8-md.png)

2.  Query: Knn (np.rand so no caching) for 15 concurrent “users”: 81QPS, p95@230ms Index: 140k docs in 2m16s

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-9-md.png)

3.  Query: Hybrid query (methodology: random keyword, random vector, 50/50 weighted norm, no filters, no aggs): 69QPS, p95@210ms Index: 140k docs in 2m16s

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-10-md.png)

4.  ‼️ Query: Hybrid query with aggs (methodology: same as #3 plus aggressive aggregation on paragaph\_id (int count), terms (word cound), and terms (phrases count): 1.8QPS, p95@9000ms Index: 140k docs in 1m37s Note: **CPU can’t cope with complex aggs during hybrid while also indexing**

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-11-md.png)

5.  ‼️Query: Hybrid query with aggs and filter (methology: same as #4, plus filtering on paragraph\_id with random value): 1.6QPS, p95@8500ms Index: 140k docs in 1m37s Note: **CPU can’t cope with complex aggs and filters during hybrid while also indexing**

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-12-md.png)

6.  ‼️ Query: Lexical only with aggs and filter (methodolgy: same as #5 but without KNN): 3.6QPS, p95@4000ms Index: 140k docs in 1m29s Note: **CPU can’t cope with indexing AND heavy term aggs, even during lexical only**

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-13-md.png)

7.  Query: Hybrid with less aggressive aggs (methodology: same as #3 plus light aggregation on paragaph\_id (int count)): 87.7QPS, p95@190ms Index: 140k docs in 1m48s Note: **MUCH BETTER**

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-14-md.png) ![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-15-md.png)

8.  Query: Hybrid with filter and no aggs: 97QPS, p95@180ms Index: 140k docs in 1m47s

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-16-md.png)

### Grafana view for all 8 runs

There are 8 peaks, each corresponding to a run.

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-17-md.png)

## m7gd.2xlarge (Graviton3)

Indexing: (5,359 docs per second)

```bash
cd rustmax@ip-172-31-78-74:/mnt/nvme/omc-vector-tests/rust$ ./fast-index.sh
Index 'wikipedia-22-12-en' does not exist
{
  "acknowledged": true,
  "shards_acknowledged": true,
  "index": "wikipedia-22-12-en"
}
Replication off for wikipedia-22-12-en
2025-07-23 17:21:07  indexer 0 started
2025-07-23 17:21:07  indexer 2 started
2025-07-23 17:21:07  indexer 3 started
2025-07-23 17:21:07  indexer 4 started
2025-07-23 17:21:07  indexer 6 started
2025-07-23 17:21:07  indexer 5 started
2025-07-23 17:21:07  indexer 1 started
⠁   [01:37:07] Ingest complete
2025-07-23 18:58:14 Done!
Replication on for wikipedia-22-12-en
Refreshed wikipedia-22-12-en
[
  {
    "health": "yellow",
    "status": "open",
    "index": "wikipedia-22-12-en",
    "uuid": "PC2_J1iWR3-2dUhN18yL5A",
    "pri": "3",
    "rep": "1",
    "docs.count": "31191813",
    "docs.deleted": "0",
    "store.size": "161015816346",
    "pri.store.size": "161015816346"
  }
]
```

### Querying a static index

1.  Null hypothesis base load test (simple match\_all query for 12 concurrent users):

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-18-md.png)

2.  knn query with np.rand for the vector on each request (no caching)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-19-md.png)

3.  Hybrid query (methodology: random keyword, random vector, 50/50 weighted norm, no filters, no aggs)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-20-md.png)

4.  Hybrid query with aggs (methodology: same as #3 plus light aggregation on paragaph\_id (int count)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-21-md.png)

5.  Hybrid query with filter (methodology: same as #3, plus filtering on paragraph\_id with random value)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-22-md.png)

## m8gd.2xlarge

Indexing: (6,823 docs per second)

```bash
max@ip-172-31-78-74:/mnt/nvme/omc-vector-tests/rust$ ./fast-index.sh
Index 'wikipedia-22-12-en' does not exist
{
  "acknowledged": true,
  "shards_acknowledged": true,
  "index": "wikipedia-22-12-en"
}
Replication off for wikipedia-22-12-en
2025-07-23 19:02:40  indexer 0 started
2025-07-23 19:02:40  indexer 2 started
2025-07-23 19:02:40  indexer 1 started
2025-07-23 19:02:40  indexer 3 started
2025-07-23 19:02:40  indexer 4 started
2025-07-23 19:02:40  indexer 5 started
2025-07-23 19:02:40  indexer 6 started                                                                                                                                                    ⠁ [00:00:00] Loading data1/train-00001-of-00253-2840fd802467fbe7.parquet                                                                                                                                                     ⠁ [00:00:00] Loading data2/train-00002-of-00253-0ecc6c7ff8c4fa3c.parquet                                                                                                                                                     ⠁ [00:00:00] Loading data0/train-00000-of-00253-8d3dffb4e6ef0304.parquet                                                                                                                                                     ⠁ [00:00:00] Loading data4/train-00004-of-00253-8c33c9247a95d7f8.parquet                                                                                                                                                     ⠁ [00:00:00] Loading data3/train-00003-of-00253-32f0ed655d4213a4.parquet                                                                                                                                                     ⠁ [00:03:30] Posting rows 15000–20000 to wikipedia-22-12-en of data6/train-00022-of-00253-f9c5d70b7960cda7.parquet
  [01:16:26] Ingest complete
2025-07-23 20:19:06 Done!
Replication on for wikipedia-22-12-en
Refreshed wikipedia-22-12-en
[
  {
    "health": "yellow",
    "status": "open",
    "index": "wikipedia-22-12-en",
    "uuid": "FMfZAfYxSNeaKbVfD3Sj-A",
    "pri": "3",
    "rep": "1",
    "docs.count": "31113807",
    "docs.deleted": "0",
    "store.size": "154817937620",
    "pri.store.size": "154817937620"
  }
]
```

### Querying a static index

1.  Null hypothesis base load test (simple match\_all query for 12 concurrent users):

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-23-md.png)

2.  knn query with np.rand for the vector on each request (no caching)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-24-md.png)

3.  Hybrid query (methodology: random keyword, random vector, 50/50 weighted norm, no filters, no aggs)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-25-md.png)

4.  Hybrid query with aggs (methodology: same as #3 plus light aggregation on paragaph\_id (int count)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-26-md.png)

5.  Hybrid query with filter (methodology: same as #3, plus filtering on paragraph\_id with random value)

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/image-27-md.png)

## Graviton Hybrid Search w/ Aggs comparison

Grav4 is the clear winner

### Requests per second

| Type | \# reqs | \# fails | Avg | Min | Max | Med | req/s |
| --- | --- | --- | --- | --- | --- | --- | --- |
| GRAV3 | 3366 | 0 (0.00%) | 206 | 76 | 447 | 210 | 52.75 |
| GRAV4 | 4368 | 0 (0.00%) | 144 | 64 | 298 | 150 | 74.76 |

### Response time percentiles (milliseconds)

| Type | \# reqs | 50% | 66% | 75% | 80% | 90% | 95% | 98% | 99% | 99.9% | 99.99% | 100% |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| GRAV3 | 3366 | 210 | 240 | 260 | 270 | 290 | 310 | 340 | 360 | 420 | 450 | 450 |
| GRAV4 | 4368 | 150 | 160 | 170 | 180 | 190 | 210 | 230 | 240 | 270 | 300 | 300 |

## In Closing

The most important things about benchmarks is the use case and methodology. We needed to measure practical outcomes with real skin in the game. We used this data to not only test viability of Graviton 4, but also have a solid foundation for recommendations to our customers.

Pushing the needle is fun, and as you see we did measure static index performance as well, but only as a baseline to see how read+write performance compared. We haven't run this test against the newer Graviton 5's yet, but I'm looking forward to it!

See you next time.

![](https://bonsai.io/blog/how-to-really-benchmark-search-engines/in_benchmark_we_trust_emblem_latin_fixed.svg)

_Copyright ©️ Bonsai.io, 2026 · By Max Irwin · Originally published at https://bonsai.io/blog/how-to-really-benchmark-search-engines/ · CC BY-SA 4.0_
