ScreenshotNeo

BlogEngineering

How to Use a Rust SDK for Web Scraping APIs

Call any web scraping API from Rust with reqwest, choose a Rust SDK, handle JavaScript and proxies, and build reliable production workflows.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: You do not need a dedicated Rust SDK to call a web scraping API. A scraping provider exposes HTTP, so a reusable asynchronous reqwest::Client is enough for authentication, JSON requests, timeouts, retries and response parsing. Use a provider crate such as webscrapingapi when its builder matches your provider and you want less boilerplate. Use raw reqwest when you need portability, custom middleware or parameters added after a wrapper release.

This guide builds a complete Rust client, then covers JavaScript rendering, proxies, synchronous and asynchronous jobs, retries, response contracts, cost and production troubleshooting. The endpoint, authentication method and payload fields are provider-specific; copy them from the provider documentation.

1. Choose the Rust integration

Approach Use it when Trade-off
reqwest You want provider portability, custom retries, tracing or new parameters. You write request and response handling yourself.
webscrapingapi crate Your account matches its API and a typed builder is useful. Its documented version is 0.1.0; verify maintenance and compatibility before production.
Managed API such as Oxylabs Web Scraper API You need proxy rotation, JavaScript rendering, access handling, parsing or job delivery. Provider pricing and response contracts become part of your system.

The reqwest documentation covers asynchronous and blocking clients, JSON and form bodies, proxies, TLS, cookies, redirects and connection reuse. Reuse one client for repeated calls so its connection pool can keep connections alive.

2. Create a provider-neutral Rust client with reqwest

Project setup

[package]
name = "scrape-client"
version = "0.1.0"
edition = "2021"

[dependencies]
anyhow = "1"
reqwest = { version = "0.12", default-features = false, features = ["json", "rustls-tls"] }
serde_json = "1"
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }

Complete asynchronous example

use anyhow::{Context, Result};
use reqwest::{Client, StatusCode};
use serde_json::{json, Value};
use std::{env, time::Duration};

#[tokio::main]
async fn main() -> Result<()> {
    let api_key = env::var("API_KEY").context("API_KEY is not set")?;
    let target_url = env::args()
        .nth(1)
        .context("usage: cargo run -- https://example.com")?;

    // Reuse this client for every request in your process.
    let client = Client::builder()
        .connect_timeout(Duration::from_secs(10))
        .timeout(Duration::from_secs(90))
        .redirect(reqwest::redirect::Policy::limited(10))
        .build()?;

    // Replace this URL, authentication header and fields with your provider's docs.
    let response = client
        .post("https://provider.example/v1/query")
        .bearer_auth(api_key)
        .json(&json!({ "url": target_url }))
        .send()
        .await
        .context("request failed")?;

    let status = response.status();
    if status == StatusCode::TOO_MANY_REQUESTS || status.is_server_error() {
        anyhow::bail!("retryable provider response: {status}");
    }
    let response = response.error_for_status()?;
    let body: Value = response.json().await.context("invalid JSON response")?;
    println!("{}", serde_json::to_string_pretty(&body)?);
    Ok(())
}

Run it with API_KEY=... cargo run -- https://example.com. Some providers use an API-key query parameter, a custom header or basic authentication instead of Bearer auth. Do not guess: use the provider’s documented mechanism.

Reading HTML or text instead of JSON

let response = client
    .get("https://provider.example/v1/query")
    .query(&[("url", target_url)])
    .send()
    .await?
    .error_for_status()?;

let content_type = response
    .headers()
    .get(reqwest::header::CONTENT_TYPE)
    .and_then(|v| v.to_str().ok())
    .unwrap_or("")
    .to_owned();
let text = response.text().await?;
if content_type.contains("application/json") {
    let value: serde_json::Value = serde_json::from_str(&text)?;
    println!("{value}");
} else {
    println!("{text}");
}

3. Use the webscrapingapi Rust crate

The documented webscrapingapi crate exposes WebScrapingAPI and QueryBuilder. Its example sets a URL, enables JavaScript rendering, adds headers and reads response text.

[dependencies]
webscrapingapi = "0.1.0"
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }
anyhow = "1"
use anyhow::Result;
use std::collections::HashMap;
use webscrapingapi::{QueryBuilder, WebScrapingAPI};

#[tokio::main]
async fn main() -> Result<()> {
    let client = WebScrapingAPI::new(std::env::var("API_KEY")?);
    let mut query = QueryBuilder::new();
    query.url("https://example.com");
    query.render_js("1");

    let mut headers = HashMap::new();
    headers.insert("Accept-Language".to_string(), "en-US".to_string());
    query.headers(headers);

    let html = client.get(query).await?.text().await?;
    println!("{html}");
    Ok(())
}

When the wrapper does not expose a newly added provider parameter, use its raw_get or raw_post methods. POST requests can include a body. Check the crate source and your provider account before depending on this API in a long-lived service.

4. JavaScript pages, headers, cookies and proxies

  • JavaScript: request browser rendering through the provider’s documented option (the crate example uses render_js("1")). A plain HTTP client cannot execute page JavaScript.
  • Headers and cookies: send only the headers and cookies required by the target and provider. Keep credentials in environment variables or a secret manager.
  • Proxies: a provider may rotate proxies for you, or you may configure a proxy in reqwest with Client::builder().proxy(reqwest::Proxy::all("http://user:pass@host:port")?). Proxy responsibility, geography and acceptable use differ by provider.
  • TLS and redirects: enable a maintained TLS backend, set a redirect limit and validate the final URL if SSRF is a concern.

5. Managed workflows: realtime, push-pull and proxy endpoint

Oxylabs documents three Web Scraper API integration methods: Realtime keeps the connection open and returns one result; Push-Pull creates asynchronous jobs that you retrieve later; Proxy Endpoint lets an application use the service as an HTTPS proxy. See the integration-method documentation and the official repository.

// Realtime shape: adapt URL, credentials and JSON fields to the provider docs.
let result = client
    .post("https://realtime.oxylabs.io/v1/queries")
    .basic_auth(env::var("OXYLABS_USER")?, Some(env::var("OXYLABS_PASSWORD")?))
    .json(&json!({ "source": "universal", "url": target_url }))
    .send()
    .await?
    .error_for_status()?;
let json: serde_json::Value = result.json().await?;

Push-Pull fits large or long-running workloads because submission and retrieval are separate operations. The official repository documents batch submission of up to 5,000 query or url values in one POST and delivery to S3-compatible storage. Realtime fits a request that must wait for one result. Proxy Endpoint fits URL-based scraping when your application already expects an HTTPS proxy.

6. Reliability checklist

  1. Reuse one reqwest::Client per process or service.
  2. Set connect, request and total-operation timeouts.
  3. Call error_for_status() before deserializing a success schema.
  4. Retry only transient transport failures, rate limits and provider 5xx responses. Use bounded exponential backoff with jitter.
  5. Record provider request IDs and job IDs, but never log API keys or sensitive page contents.
  6. Make asynchronous submissions idempotent when the provider supports an idempotency key.
  7. Validate required fields separately for HTML, parsed JSON and Markdown responses.
  8. Respect target-site terms, robots directives where applicable, privacy duties and provider acceptable-use rules.

7. Performance, cost and output decisions

There is no neutral benchmark in the available sources that establishes a universally fastest or cheapest provider. Measure with your target domains, geography, concurrency, rendering mode and output format. Browser rendering, proxy rotation and parsing generally add work compared with a direct HTTP fetch, so enable only what the page requires.

Estimate cost from successful-result volume and provider pricing, then include retries, JavaScript rendering, proxy geography, storage and callback delivery. Compare the stability of the returned schema as well as the headline request price. Batch and cloud-storage delivery can reduce application overhead for large jobs, while realtime responses simplify small synchronous tasks.

8. Troubleshooting common Rust errors

Symptom Likely cause Fix
401 Unauthorized or 403 Forbidden Wrong credential, auth scheme or account permission. Check the provider’s required header, basic-auth pair or query parameter; rotate the secret and verify target access.
429 Too Many Requests Rate or concurrency limit. Honor Retry-After when present, reduce concurrency and apply bounded backoff.
Timeouts Slow target, browser rendering or overloaded provider. Set explicit timeouts, reduce page work, use an asynchronous workflow and retry only when safe.
HTML is empty or contains a challenge The page requires JavaScript, cookies or access handling. Enable the provider’s rendering/access option or choose a managed scraper with those capabilities.
JSON parse error Error HTML or a different output contract was returned. Inspect status and content type first; save a redacted response sample and validate the documented schema.
Too many redirects Redirect loop or a redirect limit that is too low. Set a finite redirect policy, inspect the chain and verify the final host.
Crate method or field missing The wrapper version predates a provider parameter. Use raw_get/raw_post or switch to raw reqwest; verify crate maintenance.

9. Test safely before production

  • Use a provider sandbox or fixture where available.
  • Test plain HTML, JavaScript-rendered pages, redirects, rate limits, malformed responses and provider errors.
  • Run concurrency tests with the same geography and rendering settings you will use in production.
  • Assert the fields your pipeline needs rather than assuming every result is HTML.
  • Keep target URLs and captured content out of logs when they contain personal or confidential data.

10. Or skip the browser setup

If your goal is a clean screenshot rather than extracted HTML, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP or PDF. Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response reports the verdict in X-Page-Verdict and billing in X-Billed.

See the ScreenshotNeo API documentation for all options, including full-page and element capture, JavaScript, custom CSS, waits, blocking rules, headers, cookies, user agents, timezone, geolocation, resizing, caching, signed links, async jobs, webhooks, bulk capture and PDF output.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
use reqwest::Client;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let client = Client::new();
    let bytes = client
        .get("https://api.screenshotneo.com/v1/shot")
        .query(&[("access_key", "YOUR_API_KEY"), ("url", "https://stripe.com")])
        .send()
        .await?
        .error_for_status()?
        .bytes()
        .await?;
    tokio::fs::write("shot.webp", bytes).await?;
    Ok(())
}
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

An MCP server lets AI agents such as Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account and start with the 1,000 included screenshots.

11. FAQ

Is a dedicated Rust SDK required?

No. The API is HTTP, so reqwest is sufficient. A crate is an optional convenience layer.

Can reqwest scrape a JavaScript application by itself?

No. It fetches HTTP responses; use a provider’s browser-rendering feature for pages whose content is created in JavaScript.

When should I choose Push-Pull?

Choose it when jobs are numerous or long-running and your application can poll, receive callbacks or read results from storage.

How do I compare providers fairly?

Use the same targets, geography, concurrency, rendering mode and output contract, and measure successful results and total cost. The supplied sources do not provide a neutral benchmark.