ScreenshotNeo

BlogEngineering

How to Build a Fast Web Search API

Design a fast web search API with inverted indexes, bounded queries, measured ranking, caching, and production troubleshooting.

By the ScreenshotNeo team1 October 20269 min read

Short answer: build the first version on an inverted-index search engine, analyze text consistently at index and query time, keep each request bounded, return only the fields clients need, and benchmark realistic traffic before tuning storage, shards, caches, or semantic ranking. Fast search is a measured retrieval system, not simply a low-latency HTTP handler.

This guide shows a practical architecture, a lexical BM25 baseline, runnable API examples, production controls, performance work, reliability practices, cost decisions, and troubleshooting.

1. Start with an inverted index

Full-text search begins by analyzing text into tokens and building an inverted index that maps terms to the documents containing them. Lowercasing, stemming, stop-word handling, and token positions determine what matches. Query text should use a compatible analyzer so that indexing and searching agree. OpenSearch describes this inverted-index model and token-position support in its introduction to search.

Use separate field types for separate jobs:

  • text fields for analyzed full-text matching.
  • keyword, numeric, date, and boolean fields for exact filters, grouping, and sorting.
  • A combined analyzed field when users commonly search several text fields together.

Model documents around the queries you need to serve. Denormalizing safe, frequently-read data can avoid expensive joins, but it creates update and consistency work.

Elastic’s search guidance recommends benchmarking a realistic workload before committing to storage architecture or tuning parameters.

2. Choose a sensible first ranking model

Start with lexical retrieval and BM25. OpenSearch documents BM25 as its default lexical scoring algorithm. It combines term frequency, inverse document frequency, and field-length normalization. Treat the default as a baseline: evaluate it against judged queries from your own corpus.

Use semantic or hybrid retrieval only when evaluation shows that lexical matching misses meaning or intent. A common design retrieves a bounded candidate set cheaply, then reranks those candidates with a more expensive model. This can improve relevance while limiting model cost and tail latency.

3. Reference architecture

  1. Ingest: validate documents, normalize fields, assign stable IDs, and version mappings and analyzers.
  2. Index: write analyzed text plus exact filter and sort fields. Decide whether updates become visible synchronously or after a refresh.
  3. Query API: authenticate, validate and bound query text, filters, page size, sort choices, and timeouts.
  4. Retrieve: execute a lexical query over only the required fields.
  5. Rank: apply field weights, function scores, or a limited reranking stage when justified by evaluation.
  6. Serve: return a small response with stable pagination and useful metadata.
  7. Measure: record client-visible p50, p95, and p99 latency, engine time, queueing, errors, throughput, cache state, and freshness.

4. Define the API contract

A predictable contract prevents accidental expensive searches.

GET /v1/search?q=wireless+headphones&page=1&page_size=20&category=audio
  • q: required or explicitly allowed to be empty for browse requests; enforce a maximum length.
  • page_size: impose a hard upper bound and reject unreasonable values.
  • Filters: expose named, typed filters instead of accepting arbitrary engine JSON from untrusted clients.
  • Sort: allow only indexed keyword, numeric, or date fields; do not sort on analyzed text. See Elastic’s sorting guidance.
  • Fields: request and return only the properties the caller needs.
  • Timeouts and cancellation: propagate request cancellation to the search engine.
  • Pagination: use a bounded page number for simple interfaces; use a cursor such as search_after for deep scrolling.

5. Minimal OpenSearch index mapping

curl -X PUT 'http://localhost:9200/products-v1' \
  -H 'Content-Type: application/json' \
  -d '{
    "settings": {
      "number_of_shards": 1,
      "number_of_replicas": 1
    },
    "mappings": {
      "properties": {
        "title": {"type": "text"},
        "description": {"type": "text"},
        "category": {"type": "keyword"},
        "price_cents": {"type": "integer"},
        "published_at": {"type": "date"}
      }
    }
  }'

Choose shard counts from measured document size, query concurrency, recovery time, and growth. A copied shard layout can be slower or more expensive on a different corpus.

6. Runnable search API in Python

The example below uses FastAPI and the OpenSearch Python client. It searches only the title and description, applies a typed category filter, caps page size, and returns selected fields.

from fastapi import FastAPI, HTTPException, Query
from opensearchpy import OpenSearch

app = FastAPI()
client = OpenSearch("http://localhost:9200")
INDEX = "products-v1"

@app.get("/v1/search")
def search(
    q: str = Query("", max_length=200),
    category: str | None = Query(None, max_length=80),
    page: int = Query(1, ge=1, le=1000),
    page_size: int = Query(20, ge=1, le=100),
):
    must = []
    if q.strip():
        must.append({
            "multi_match": {
                "query": q,
                "fields": ["title^3", "description"],
                "operator": "and"
            }
        })
    filters = []
    if category:
        filters.append({"term": {"category": category}})

    body = {
        "from": (page - 1) * page_size,
        "size": page_size,
        "track_total_hits": False,
        "_source": ["title", "description", "category", "price_cents", "published_at"],
        "query": {"bool": {"must": must or [{"match_all": {}}], "filter": filters}},
    }
    try:
        result = client.search(index=INDEX, body=body, request_timeout=2)
    except Exception as exc:
        raise HTTPException(status_code=502, detail="search backend unavailable") from exc

    return {
        "items": [
            {"score": hit.get("_score"), "id": hit["_id"], **hit.get("_source", {})}
            for hit in result["hits"]["hits"]
        ],
        "timed_out": result.get("timed_out", False),
    }

Run it with pip install fastapi uvicorn opensearch-py, then uvicorn app:app --reload. In production, configure authentication, TLS, connection pooling, retries, and a client timeout appropriate for your workload.

7. Equivalent requests in cURL, Python, and Node.js

cURL

curl -G 'http://localhost:8000/v1/search' \
  --data-urlencode 'q=wireless headphones' \
  --data-urlencode 'category=audio' \
  --data-urlencode 'page_size=20'

Python

import requests

response = requests.get(
    "http://localhost:8000/v1/search",
    params={"q": "wireless headphones", "category": "audio", "page_size": 20},
    timeout=3,
)
response.raise_for_status()
print(response.json())

Node.js

const params = new URLSearchParams({
  q: 'wireless headphones',
  category: 'audio',
  page_size: '20'
});
const response = await fetch(`http://localhost:8000/v1/search?${params}`, {
  signal: AbortSignal.timeout(3000)
});
if (!response.ok) throw new Error(`search failed: ${response.status}`);
console.log(await response.json());

8. Make the common query cheap

  • Search only necessary fields and use field boosts deliberately.
  • Use filters for exact constraints; filter clauses can be cached by the engine and do not need scoring.
  • Return a bounded result set and disable total-hit counting when the UI does not need an exact count.
  • Use keyword or numeric fields for sorting.
  • Avoid wildcard, regexp, script, and unrestricted fuzzy queries on public endpoints.
  • Prefer denormalized documents when they remove a join from a hot path.
  • Bundle independent requests with OpenSearch Multi-Search when it reduces client orchestration, then measure server resource use.

Explain APIs are diagnostic tools. OpenSearch warns that explanations consume resources and time; use them on representative failures, not every production request.

9. Caches, memory, and shard layout

Search engines rely heavily on the operating-system filesystem cache. Elastic’s self-managed guidance says that, in general, at least half of available memory should go to filesystem cache so hot index regions can remain in physical memory. Treat that as vendor guidance to validate, not a universal optimum.

Repeated requests can lose cache locality when routed to different shard copies. Review routing, replica selection, shard size, query cost, and concurrency together. Keep shard counts and index layouts under version control and benchmark changes with warm and cold behavior.

10. Add semantic retrieval only when it earns its cost

Hybrid search combines lexical and vector retrieval; reranking then applies a more expensive model to a smaller candidate set. Measure relevance lift, p95 and p99 latency, memory, model-serving cost, and fallback behavior. OpenSearch’s vector performance guidance also calls out segment count, shard parallelism, and warming native indexes because first queries can be slower.

Keep a lexical fallback for model timeouts, missing embeddings, or index migration. Do not assume semantic search is faster.

11. Benchmark the whole API

Define “fast” at the API boundary, including authentication, queueing, serialization, and network time. Build a workload that includes frequent and rare queries, filters, pagination, concurrent users, empty-result searches, and cold and warm cache runs.

  1. Record a fixed query corpus and expected relevance judgments.
  2. Replay the same mix at several concurrency levels.
  3. Report p50, p95, and p99 latency, throughput, errors, timeouts, and freshness.
  4. Separate client-visible latency from engine time to find queueing and network overhead.
  5. Change one variable at a time: mapping, analyzer, query fields, shard count, replicas, refresh behavior, or hardware.
  6. Repeat after every production-like change.

Elastic’s tuning documentation states: “Before committing to a particular storage architecture, benchmark your system with a realistic workload to determine the effects of any tuning parameters.”

12. Freshness, reliability, and operations

  • Version mappings and analyzers; incompatible changes usually require a new index and an alias cutover.
  • Choose refresh behavior from freshness requirements and write load. More frequent refreshes can increase indexing overhead.
  • Use replicas and tested snapshots for recovery.
  • Set bounded engine and API timeouts, propagate cancellation, and return a clear degraded response when the backend is unavailable.
  • Apply authentication, authorization, rate limits, request-size limits, and query allowlists.
  • Monitor rejected tasks, heap and filesystem-cache pressure, segment counts, refresh time, indexing lag, error rate, and tail latency.

13. Troubleshooting

Symptom Likely cause Fix
Relevant documents are missing Index and query analyzers differ, or the field is mapped as keyword Inspect analyzed tokens, align analyzers, and use a text field for full-text matching.
Sorting is slow or rejected Sorting on an analyzed text field Add a keyword, numeric, or date field and sort on that field.
p99 spikes under load Unbounded queries, queueing, cache misses, or too much shard fan-out Cap query work, inspect slow requests, test routing and shard layout, and measure concurrency.
First vector query is slow Native vector index or segment warm-up Warm indexes where supported and benchmark cold-start behavior.
Results are stale Refresh and ingestion visibility settings Document the freshness contract and tune refresh or read-after-write behavior accordingly.
High memory pressure Too many shards, large aggregations, broad field loads, or insufficient filesystem cache Reduce shard overhead, bound aggregations, return fewer fields, and validate heap/cache allocation.
Explain requests hurt production Explain API used on normal traffic Use explanations only for sampled diagnosis.
Deep pages time out Large from/size offsets Use cursor pagination such as search_after and cap page depth.
Option Useful when Trade-offs
Self-managed Elasticsearch/OpenSearch You need direct control over mappings, shards, and cluster settings. Your team owns upgrades, capacity, recovery, and operations.
Amazon OpenSearch Service You want AWS to provide a managed deployment, operation, and scaling path. Compare regional pricing, limits, integration, control, and latency for your configuration. Use the current AWS pricing calculator.
Lexical BM25 Queries are primarily term based and explainability matters. Evaluate relevance and latency on your own judged queries.
Hybrid or semantic retrieval Evaluation shows lexical search misses intent. Budget model serving, memory, tail latency, and fallback complexity.

No source-backed benchmark establishes a universally fastest engine or a universal p95 target. Select by workload, operational capacity, relevance, availability, freshness, and measured cost.

15. Further reading

Elasticsearch in Action, Second Edition (Manning, 2023) covers architecture, APIs, indexing, and tuning for Elasticsearch. It is useful if that is your chosen stack, but it is not required for every implementation.

16. Or skip the browser setup

If your search product needs screenshots for result previews, documentation, or visual regression pages, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Cookie banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. The MCP server lets AI agents call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Is BM25 enough for a new search API?

Often yes for a term-oriented corpus. Establish a relevance baseline before adding vectors or reranking.

What latency target should I use?

Set it from your product’s interaction budget and measure p50, p95, and p99 at the API boundary. The research does not establish a universal target.

Should I use one index or many?

Use the layout that matches retention, freshness, tenant isolation, and query patterns. Benchmark shard count and routing with realistic data.

When should I add caching?

After measuring repeated-query frequency and cache locality. Cache only responses whose freshness and authorization rules permit it.

How do I debug a relevance complaint?

Capture the query, filters, mapping version, and returned scores; inspect analyzed tokens and use an explain request on representative cases, outside the normal hot path.