How to Use a TypeScript SDK for Web Scraping APIs
Connect TypeScript to a web scraping API safely: install an SDK, render JavaScript when needed, parse results, retry failures, and scale jobs.
A TypeScript SDK is a provider-specific client for a web scraping API. Install the provider’s package, keep its credential on your server, make one basic request, inspect both API and target-page status, then add JavaScript rendering, parsing, retries, and asynchronous jobs only when your workload needs them.
SDK methods, option names, response formats, token types, limits, and billing differ by provider. Use the provider’s current reference beside your code. The examples below show documented integration shapes from Scrapfly and Crawlbase; they were not independently executed here.
1. Choose the API before choosing its SDK
Decide what the application actually needs:
- One static HTML or JSON response
- JavaScript rendering for client-side content
- Structured extraction or a provider scraper
- Cookies, headers, geography, scrolling, or clicks
- Screenshots or PDFs instead of parsed text
- Async jobs, callbacks, crawling, or batch submission
Compare runtime support, package maintenance, authentication, rendering controls, output formats, error visibility, concurrency, batch features, pricing, privacy terms, and permitted use. An SDK simplifies HTTP calls; it does not decide whether collecting data from a particular site is allowed. Check the target site’s terms and applicable law.
2. Install a provider package
Follow the package manager and runtime version in the provider’s current documentation.
# Crawlbase (documented package)
npm install crawlbase
# A typical TypeScript project
npm install -D typescript tsx @types/node
npx tsc --init
Crawlbase documents Node.js 16 or later, ESM and CommonJS imports, and a package named crawlbase. Scrapfly distributes its JavaScript/TypeScript SDK through npm, JSR, and Deno. Confirm the package version and current method names before copying an example.
3. Store credentials safely
Read keys from server-side environment configuration or a secrets manager. Never put a paid scraping key in browser-delivered JavaScript, source control, issue reports, or request logs.
# .env (keep this file out of version control)
SCRAPFLY_KEY=replace_me
CRAWLBASE_TOKEN=replace_me
// config.ts
const key = process.env.SCRAPFLY_KEY;
if (!key) throw new Error('SCRAPFLY_KEY is not configured');
In production, grant the process only the secret it needs, rotate credentials through your secret manager, and redact authorization headers in logs.
4. Make the smallest possible request
Scrapfly TypeScript example
The official repository initializes ScrapflyClient with a key and passes a ScrapeConfig to client.scrape. It demonstrates JavaScript rendering, a country option, and anti-bot handling. The repository says unblocker is the current name and asp is a deprecated alias. Verify names against the installed version.
import { ScrapflyClient, ScrapeConfig } from 'scrapfly-sdk';
const key = process.env.SCRAPFLY_KEY;
if (!key) throw new Error('SCRAPFLY_KEY is required');
const client = new ScrapflyClient({ key });
const response = await client.scrape(
new ScrapeConfig({
url: 'https://example.com',
render_js: true,
// Add only options documented by your installed version.
// country: 'us',
// unblocker: true,
}),
);
console.log(response.result.content);
The sample response exposes HTML through response.result.content and includes a selector helper in the repository’s example. Treat this as documentation shape, not a tested fixture.
Crawlbase TypeScript example
import { CrawlingAPI } from 'crawlbase';
const token = process.env.CRAWLBASE_TOKEN;
if (!token) throw new Error('CRAWLBASE_TOKEN is required');
const api = new CrawlingAPI({ token });
const response = await api.get('https://example.com');
console.log('API status:', response.statusCode);
console.log('Target status:', response.headers?.cb_status);
if (response.statusCode !== 200) {
throw new Error(`Crawlbase request failed: ${response.statusCode}`);
}
if (response.headers?.cb_status && response.headers.cb_status !== '200') {
throw new Error(`Target retrieval failed: ${response.headers.cb_status}`);
}
console.log(response.body);
Crawlbase describes its Node SDK as “a thin wrapper around the same HTTP API documented in API Reference.” Its response includes a body, an API statusCode, and headers. A 200 API response can still accompany an empty body and a non-200 cb_status, so check both.
5. Render JavaScript only when the page needs it
Fetch the page without rendering first. If the returned HTML contains the required data, keep the cheaper and usually simpler path. Enable the provider’s browser mode when the content is inserted by client-side JavaScript, loaded after an XHR/fetch call, revealed after scrolling, or hidden behind an interaction.
Crawlbase separates a Normal Token for static HTML and JSON endpoints from a JavaScript Token for SPAs and client-rendered or lazy-loaded content. Its documented page_wait, ajax_wait, scroll, and css_click_selector options require the JavaScript Token. Scrapfly’s example uses render_js: true. These controls are provider-specific; do not assume the same names or pricing elsewhere.
Prefer a condition that proves the page is ready, such as a selector or completed network request, when the provider supports it. Fixed sleeps can waste time and still miss slow content. Keep waits bounded.
6. Parse and validate the result
Choose structured extraction only when the provider supports the exact target and fields you need. Otherwise parse the returned HTML or text with a parser such as Cheerio.
import * as cheerio from 'cheerio';
export function parseTitle(html: string): string {
const $ = cheerio.load(html);
const title = $('title').first().text().trim();
if (!title) throw new Error('Required title field is missing');
return title;
}
const title = parseTitle(response.body);
console.log({ title });
Validate required fields before sending data downstream. Record the URL, retrieval time, provider request ID (if supplied), parser version, and a short failure reason. Page markup changes; treat missing selectors as an expected operational error.
7. Handle authentication, status, and retries
- Authenticate through the SDK’s constructor or documented request option.
- Check the SDK request’s HTTP status.
- Check the provider’s target verdict or status field.
- Validate that the body is non-empty and has the expected content type.
- Retry only transient failures with bounded exponential backoff and jitter.
function sleep(ms: number) {
return new Promise(resolve => setTimeout(resolve, ms));
}
async function withRetry<T>(operation: () => Promise<T>, attempts = 3): Promise<T> {
let lastError: unknown;
for (let attempt = 0; attempt < attempts; attempt++) {
try {
return await operation();
} catch (error) {
lastError = error;
if (attempt === attempts - 1) break;
const delay = Math.min(8000, 500 * 2 ** attempt) + Math.floor(Math.random() * 250);
await sleep(delay);
}
}
throw lastError;
}
Do not blindly retry authentication errors, invalid URLs, unsupported options, quota exhaustion, or other documented client errors. Honor provider retry-after headers and account limits. Idempotent single-page fetches are safer to retry than operations that trigger side effects.
8. Move recurring work to async jobs
For large batches, slow targets, or scheduled collection, investigate the provider’s async request, crawl job, callback, or webhook features. Crawlbase documents an async request that returns a request ID and callback delivery, and recommends async processing for sustained high-volume submission. Verify callback signing, retry behavior, retention, concurrency, and plan limits in the current documentation.
Use one long-lived client instance when the provider recommends it instead of constructing a client for every URL. Bound concurrency with a queue, observe quota and rate-limit headers, and persist job state so a process restart does not lose work.
9. Complete command-line and language examples
cURL
curl -G "https://api.example-provider.test/v1/fetch" \
-H "Authorization: Bearer $SCRAPING_API_KEY" \
--data-urlencode "url=https://example.com"
Replace the endpoint, authentication header, and options with the provider’s documented values. cURL is useful for confirming credentials and response fields before debugging TypeScript.
Python
import os
import requests
response = requests.get(
"https://api.example-provider.test/v1/fetch",
params={"url": "https://example.com"},
headers={"Authorization": f"Bearer {os.environ['SCRAPING_API_KEY']}"},
timeout=90,
)
response.raise_for_status()
print(response.text)
Node.js without an SDK
const target = encodeURIComponent('https://example.com');
const res = await fetch(
`https://api.example-provider.test/v1/fetch?url=${target}`,
{ headers: { Authorization: `Bearer ${process.env.SCRAPING_API_KEY}` } },
);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
console.log(await res.text());
These generic examples make the HTTP contract visible. An SDK should add typed configuration, convenience methods, and provider-specific response handling; it does not remove the need to read the API reference.
10. Or skip the browser setup
If your goal is a clean screenshot or PDF rather than parsed HTML, ScreenshotNeo provides a single GET request and an MCP server for AI agents. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP tools include take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
You can add full-page capture with lazy images, CSS-element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and margin settings, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agent, Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTL, signed image links, async signed webhooks, bulk capture for up to 100 URLs per call, and usage reporting. Every feature is on every plan. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
11. Performance, reliability, and cost
- Start static: browser rendering can add latency and resource use. Enable it for pages that require it.
- Bound work: set request, page-wait, queue, and overall job timeouts.
- Reuse clients: avoid repeated setup and let the SDK manage connections when supported.
- Control concurrency: match provider quotas and the target site’s capacity.
- Cache deliberately: cache stable pages, but include the relevant URL, rendering mode, headers, cookies, and parser version in your cache key.
- Measure the right result: track API errors, target-status failures, empty bodies, parse failures, latency, retries, and cost separately.
- Estimate rendering cost: provider billing models differ; confirm whether JavaScript, bandwidth, proxy, browser time, retries, or async jobs consume additional units.
12. Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 | Missing, expired, or wrong credential | Read the key from the server environment, check the account, and avoid logging it. |
| HTTP 200 but empty content | Target failed while the API request succeeded | Inspect provider target-status fields such as Crawlbase cb_status; check the body and request ID. |
| HTML lacks visible data | Content is client-rendered or lazy-loaded | Enable the provider’s documented browser mode, wait, scroll, or click option. |
| Timeouts | Slow target, excessive waits, or blocked resources | Set bounded waits, block unnecessary resource types where supported, and retry transient failures with backoff. |
| Parser returns no fields | Markup changed or selector is wrong | Save a sanitized response sample, inspect selectors, validate required fields, and version the parser. |
| 429 or quota errors | Rate or plan limit reached | Honor retry-after, reduce concurrency, queue jobs, or change the plan after checking current limits. |
| TypeScript import error | ESM/CommonJS mismatch or unsupported runtime | Follow the package’s module example, check package.json, and use the documented Node version. |
| Retries multiply charges | Provider bills attempts or rendered work | Confirm billing rules, retry only transient errors, use idempotency where offered, and prefer caching. |
13. Short FAQ
Do I need TypeScript to use a scraping API SDK?
No. Most providers expose HTTP endpoints, so cURL, Python, and plain Node.js can call them. TypeScript adds static types and editor assistance when the provider supplies useful declarations.
Should I always enable JavaScript rendering?
No. Try a static request first. Render only when required content is absent or loaded by browser code, and confirm the provider’s cost and latency behavior.
Does an SDK make scraping permitted?
No. Review the target site’s terms, robots guidance where relevant, data rules, and applicable legal advice for your use case.
When should I use async jobs?
Use them for large or slow workloads, scheduled collection, and callback-based delivery. Keep single interactive requests synchronous when the user needs an immediate result.
Can I use ScreenshotNeo for extracted HTML?
ScreenshotNeo is designed for screenshots, PDFs, page information, and AI-agent capture workflows. Use a scraping API when you need a provider’s HTML or structured extraction response.
Sources and provider notes
Implementation details in this guide come from the Scrapfly TypeScript/JavaScript SDK repository, Crawlbase Node.js SDK documentation, and the Scrapeless SDK overview. Recheck each provider’s versioned documentation for current option names, limits, pricing, and runtime support.


