How to Collect 1,000 Web Search Results in Under 5 Minutes
Learn how to collect 1,000 authorized search results efficiently, page through APIs, deduplicate URLs, measure five-minute performance, and avoid scraping violations.
Short answer: Use an authorized search API, request its largest documented page size, advance the pagination offset until you reach 1,000 unique URLs or the provider runs out of results, and measure the complete run. With Bing Web Search API’s documented maximum of 50 requested results per page, 1,000 rows require at least 20 requests if every page is full and has no overlap. That is arithmetic, not a five-minute performance guarantee.
The five-minute target depends on your query, market, API eligibility, quotas, latency, retries, duplicate results, and storage time. Record those variables and report the measured wall-clock time instead of promising that every query will produce 1,000 distinct results in five minutes.
Choose an authorized route
For general commercial collection, select a search API whose terms permit your use. Microsoft’s Bing Web Search API documents pagination with count and offset; count can be requested up to 50, although a response can contain fewer items and pages can overlap. See the Bing query-parameter reference.
Google’s Search Researcher Result API is limited to approved researcher projects and non-commercial use. Its documented program quota of 1,000 queries per day per approved project is a query quota, not a promise of 1,000 results from one query. Do not use it as a general commercial collection route.
Do not automate Google result-page scraping without express permission. Google Search Central states that automated queries and scraping search results without permission violate its spam policies and Terms of Service; review the Google spam policies before building any workflow.
Plan the collection before writing code
- Define the target. Decide whether “1,000” means raw rows or 1,000 unique normalized URLs. Use unique URLs for a useful dataset.
- Fix the search context. Record the exact query, market or geography, language, safe-search setting, device setting, and timestamp.
- Set a stopping rule. Stop at 1,000 unique URLs, when the API returns no more results, when the offset reaches the provider’s documented limit, or when your time budget expires.
- Plan persistence. Save the query, offset, rank, URL, title, snippet, retrieval timestamp, and raw response or response identifier.
- Budget retries. Use bounded exponential backoff for transient errors and avoid sending requests faster than the provider permits.
Reference implementation with Bing Web Search API
The following Python program requests 50 results at a time, follows offset, normalizes URLs for deduplication, preserves ranks, retries transient failures, and reports elapsed time. Set BING_SUBSCRIPTION_KEY and adjust the endpoint or market to match your account and documentation.
import json
import os
import sys
import time
from datetime import datetime, timezone
from urllib.parse import urlsplit, urlunsplit
import requests
ENDPOINT = "https://api.bing.microsoft.com/v7.0/search"
QUERY = sys.argv[1] if len(sys.argv) > 1 else "renewable energy storage"
TARGET = 1000
PAGE_SIZE = 50
MARKET = "en-US"
MAX_RETRIES = 4
key = os.environ.get("BING_SUBSCRIPTION_KEY")
if not key:
raise SystemExit("Set BING_SUBSCRIPTION_KEY before running")
def normalize_url(value):
parts = urlsplit(value)
scheme = parts.scheme.lower()
host = (parts.hostname or "").lower()
if not host:
return value
port = parts.port
netloc = host
if port and not ((scheme == "http" and port == 80) or (scheme == "https" and port == 443)):
netloc += f":{port}"
path = parts.path or "/"
return urlunsplit((scheme, netloc, path, parts.query, ""))
def get_page(offset):
params = {
"q": QUERY,
"count": PAGE_SIZE,
"offset": offset,
"mkt": MARKET,
"responseFilter": "Webpages",
}
headers = {"Ocp-Apim-Subscription-Key": key}
for attempt in range(MAX_RETRIES):
response = requests.get(ENDPOINT, params=params, headers=headers, timeout=30)
if response.status_code == 200:
return response.json()
if response.status_code not in (408, 429, 500, 502, 503, 504):
response.raise_for_status()
if attempt == MAX_RETRIES - 1:
response.raise_for_status()
time.sleep(2 ** attempt)
raise RuntimeError("unreachable")
started = time.monotonic()
seen = set()
rows = []
offset = 0
requests_made = 0
raw_rows = 0
while len(seen) < TARGET:
data = get_page(offset)
requests_made += 1
items = data.get("webPages", {}).get("value", [])
if not items:
break
raw_rows += len(items)
for item in items:
url = item.get("url")
if not url:
continue
normalized = normalize_url(url)
if normalized in seen:
continue
seen.add(normalized)
rows.append({
"query": QUERY,
"rank": len(rows) + 1,
"offset": offset,
"url": url,
"normalized_url": normalized,
"name": item.get("name"),
"snippet": item.get("snippet"),
"retrieved_at": datetime.now(timezone.utc).isoformat(),
})
if len(rows) >= TARGET:
break
offset += len(items)
if len(items) < PAGE_SIZE:
# A short page may mean fewer available results; continue once to
# let the API confirm exhaustion, then stop on an empty page.
continue
elapsed = time.monotonic() - started
with open("search-results.jsonl", "w", encoding="utf-8") as output:
for row in rows:
output.write(json.dumps(row, ensure_ascii=False) + "\n")
print(json.dumps({
"query": QUERY,
"market": MARKET,
"raw_rows": raw_rows,
"unique_urls": len(rows),
"requests": requests_made,
"elapsed_seconds": round(elapsed, 3),
"under_five_minutes": elapsed < 300 and len(rows) >= TARGET,
}, indent=2))
This implementation treats the returned URL as the source record and removes URL fragments. Depending on your data policy, you may also remove tracking parameters, canonicalize trailing slashes, or keep query strings because they can identify different resources.
cURL: page through the API
export BING_SUBSCRIPTION_KEY='YOUR_KEY'
query='renewable energy storage'
offset=0
count=50
curl --fail-with-body -G 'https://api.bing.microsoft.com/v7.0/search' \
-H "Ocp-Apim-Subscription-Key: $BING_SUBSCRIPTION_KEY" \
--data-urlencode "q=$query" \
--data "count=$count" \
--data "offset=$offset" \
--data "mkt=en-US" \
--data "responseFilter=Webpages" \
-o page-000.json
Repeat the request with offset=50, offset=100, and so on. Use a script for production collection so you can deduplicate, retry, persist metadata, and stop at a unique-result count.
Node.js implementation
const fs = require('node:fs/promises');
const endpoint = 'https://api.bing.microsoft.com/v7.0/search';
const key = process.env.BING_SUBSCRIPTION_KEY;
const query = process.argv.slice(2).join(' ') || 'renewable energy storage';
const target = 1000;
const pageSize = 50;
if (!key) throw new Error('Set BING_SUBSCRIPTION_KEY before running');
function normalizeUrl(value) {
const u = new URL(value);
u.hash = '';
return u.toString();
}
async function fetchPage(offset) {
const url = new URL(endpoint);
url.searchParams.set('q', query);
url.searchParams.set('count', String(pageSize));
url.searchParams.set('offset', String(offset));
url.searchParams.set('mkt', 'en-US');
url.searchParams.set('responseFilter', 'Webpages');
for (let attempt = 0; attempt < 4; attempt++) {
const res = await fetch(url, {
headers: { 'Ocp-Apim-Subscription-Key': key }
});
if (res.ok) return res.json();
if (![408, 429, 500, 502, 503, 504].includes(res.status) || attempt === 3) {
throw new Error(`${res.status}: ${await res.text()}`);
}
await new Promise(resolve => setTimeout(resolve, 2 ** attempt * 1000));
}
}
(async () => {
const started = Date.now();
const seen = new Set();
const rows = [];
let offset = 0;
let requests = 0;
let rawRows = 0;
while (rows.length < target) {
const data = await fetchPage(offset);
requests++;
const items = data.webPages?.value || [];
if (!items.length) break;
rawRows += items.length;
for (const item of items) {
if (!item.url) continue;
const normalized = normalizeUrl(item.url);
if (seen.has(normalized)) continue;
seen.add(normalized);
rows.push({ query, rank: rows.length + 1, offset, url: item.url,
normalized_url: normalized, name: item.name, snippet: item.snippet,
retrieved_at: new Date().toISOString() });
if (rows.length >= target) break;
}
offset += items.length;
}
await fs.writeFile('search-results.json', JSON.stringify(rows, null, 2));
console.log({ query, rawRows, uniqueUrls: rows.length, requests,
elapsedSeconds: (Date.now() - started) / 1000,
underFiveMinutes: rows.length >= target && Date.now() - started < 300000 });
})();
Concurrency, speed, and the five-minute measurement
Sequential paging is easiest to reason about because each offset follows the previous response. Parallel requests can reduce waiting time only when the provider permits them and your offsets are independent. Before adding concurrency, confirm rate limits and quota rules. Use a small worker pool, bounded retries, and a queue so a burst does not trigger throttling.
| Measure | Why it matters |
|---|---|
| Raw rows | Shows how many items the provider returned, including duplicates. |
| Unique URLs | Shows whether you actually reached 1,000 distinct results. |
| Requests and retries | Explains quota consumption and latency. |
| Elapsed wall-clock time | Includes network, backoff, parsing, deduplication, and storage. |
| Query, market, and settings | Makes the result reproducible and comparable. |
Start the timer before the first request and stop it after the final record is durably written. Run several trials for the same query if you need a meaningful estimate. A single run cannot establish a repeatable service guarantee.
Duplicates, short pages, and result limits
- Overlapping pages: deduplicate by normalized URL while retaining the original rank and offset.
- Short pages: never assume that requesting 50 returns 50. Continue until an empty page or another documented exhaustion condition.
- Fewer than 1,000 available results: report the unique total; do not manufacture rows by counting duplicate responses.
- Changing results: search indexes can change between requests. Save retrieval timestamps and the original response.
- Canonicalization: preserve the original URL for auditability and store a separate normalized key for deduplication.
- Quota exhaustion: stop safely, persist the checkpoint, and resume only when the provider permits it.
Reliability and cost checklist
- Keep API keys in environment variables or a secret manager.
- Use request timeouts and bounded retries for 408, 429, and server errors.
- Persist each page or checkpoint before requesting the next page.
- Log status codes, offsets, retry counts, and elapsed time without logging secrets.
- Estimate cost from the provider’s current pricing and the actual number of requests; the 20-request figure is only a theoretical minimum at 50 rows per page.
- Review commercial terms, retention rules, geography controls, and permitted downstream use before collecting or redistributing results.
Common errors and fixes
| Error | Likely cause | Fix |
|---|---|---|
| 401 or 403 | Missing, invalid, expired, or unauthorized key. | Check the environment variable, subscription, endpoint, and account permissions. |
| 400 | Invalid query parameter or unsupported market/value. | Compare every parameter with the provider’s current reference and URL-encode the query. |
| 429 | Rate limit or quota exceeded. | Honor retry-after guidance, reduce concurrency, and checkpoint progress. |
| 5xx or timeout | Transient provider or network failure. | Use a timeout, exponential backoff, and a maximum retry count. |
| Only a few unique URLs | Short pages, overlap, restrictive query, or exhausted result depth. | Inspect raw rows and offsets; report the measured unique total. |
| Run exceeds five minutes | High latency, retries, storage overhead, or insufficient page size. | Measure each phase, verify the documented maximum page size, and use permitted bounded concurrency. |
| Results cannot be reproduced | Query context or timestamp was not saved. | Persist query, market, settings, offsets, timestamps, and raw responses. |
Or skip the browser setup
If your workflow also needs visual snapshots of result pages or other URLs, ScreenshotNeo provides a single screenshot API call instead of maintaining browser automation. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients take screenshots. The free plan includes 1,000 screenshots each month with no card, and paid plans start at $5 for 3,000 shots.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Create a free ScreenshotNeo account with 1,000 screenshots per month and no card.
FAQ
Can I get 1,000 results from one query?
Only if the selected provider exposes that much result depth for the query and your account is allowed to request it. A query quota does not guarantee a result count.
Is 20 requests enough?
Twenty is the theoretical minimum when every request returns 50 distinct results. Short pages, overlap, and retries can require more requests or make 1,000 unique results unavailable.
Should I count rows or URLs?
Count normalized unique URLs for the final total, while retaining raw rows for auditing and diagnosing overlap.
Can I scrape Google automatically?
Do not do so without express permission. Google’s published policy says automated queries and scraping search results without permission violate its spam policies and Terms of Service.
How do I prove the run was under five minutes?
Save start and end timestamps, request and retry counts, raw rows, unique URLs, query context, and the final persisted file. Publish those conditions with the measured result.


