How to Capture Screenshots of Indian News Websites at Scale with Browserless
Use Browserless to capture Indian news pages in batches, handle lazy-loaded content, and detect blocked or incomplete screenshots before scaling up.
To capture Indian news websites at scale with Browserless, send one authenticated POST /screenshot request per page, configure the viewport and capture options, and save each image with a record of its URL and outcome. For long articles, enable scrolling before a full-page capture so lazy-loaded content has a chance to appear. Start with a small representative batch: Browserless documents the endpoint and options, but its reviewed documentation does not establish a universal safe request rate or permission to automate any particular publisher.
This guide covers a repeatable REST workflow, a runnable Python batch example, cURL and Node.js requests, capture settings, validation, failure handling, and the point at which a scripted browser session may be a better fit.
1. Choose the right Browserless capture path
For independent snapshots where each input URL needs one image, use Browserless’s stateless REST POST /screenshot endpoint. It accepts a URL or raw HTML and can return PNG, JPEG, or WebP. A REST request does not require you to run Puppeteer or Playwright locally. See the Browserless Screenshot API and REST API overview.
| Workflow | Use it when | Output or behavior |
|---|---|---|
| Screenshot REST endpoint | Each URL is an independent capture | Image response with screenshot options |
| Smart Scrape with screenshot format | You want screenshot output alongside other requested page outputs | Screenshot is returned as a base64 PNG field; requesting it forces a headless browser |
| Scripted browser workflow | You need multiple actions, branches, or conditional logic before capture | A managed browser session with interaction control |
Smart Scrape’s screenshot behavior is documented in the Smart Scrape API documentation. Browserless describes one-shot REST calls and scripted browser paths in its website scraping guide. A simple screenshot request cannot branch interactively mid-execution; use a scripted workflow when the page must be inspected and acted on before deciding what to capture.
2. Set up access and a small input manifest
- Get a Browserless API token using its account and access process. Keep it in an environment variable such as
BROWSERLESS_TOKEN; do not commit it to source control. - Prepare a manifest of target URLs and stable record IDs. Keep the original URL alongside every result.
- Choose consistent capture options for pages you plan to compare, including viewport, scale, and format.
- Run a few representative pages first: an article, a homepage, a long article, and a page with dynamic content.
- Review whether each publisher’s current terms, access controls, and intended-use requirements allow the planned automated access, retention, and sharing.
The research reviewed for this guide does not establish terms or automation permissions for any specific Indian publisher. A successful image response does not itself grant permission to reuse or redistribute the page or its contents.
3. Make one screenshot request with cURL
Browserless documents POST /screenshot with API token authentication. The following shell example sends a URL and common capture settings as JSON. The exact options are configurable; consult the endpoint reference when adding options such as selectors, clipping, waits, or device scale.
export BROWSERLESS_TOKEN='YOUR_API_TOKEN'
curl --fail-with-body --silent --show-error \
-X POST 'https://production-sfo.browserless.io/screenshot?token='"$BROWSERLESS_TOKEN" \
-H 'Content-Type: application/json' \
--data '{
"url": "https://example.com/news/article",
"options": {
"fullPage": true,
"type": "webp",
"quality": 80,
"viewport": { "width": 1365, "height": 900 },
"deviceScaleFactor": 1,
"scrollPage": true
}
}' \
-o article.webp
Replace the example URL with a page you are authorized to capture. If your Browserless account uses a different endpoint host, use the endpoint shown for your account. A successful HTTP response should be checked and recorded; do not assume every returned image is a valid article capture.
4. Capture a batch with Python
This runnable example processes a modest number of URLs sequentially, saves the raw image response, and writes a JSON Lines record for each attempt. Sequential processing is a conservative starting point, not a documented Browserless rate recommendation. Expand concurrency only after checking your account limits, publisher conditions, and results.
import json
import os
import time
from datetime import datetime, timezone
from pathlib import Path
import requests
TOKEN = os.environ["BROWSERLESS_TOKEN"]
ENDPOINT = "https://production-sfo.browserless.io/screenshot"
OUT = Path("captures")
OUT.mkdir(exist_ok=True)
urls = [
"https://example.com/news/article-one",
"https://example.com/news/article-two",
]
options = {
"fullPage": True,
"type": "webp",
"quality": 80,
"viewport": {"width": 1365, "height": 900},
"deviceScaleFactor": 1,
"scrollPage": True,
}
with (OUT / "manifest.jsonl").open("a", encoding="utf-8") as manifest:
for index, url in enumerate(urls, start=1):
started = datetime.now(timezone.utc).isoformat()
record = {
"id": f"page-{index:04d}",
"url": url,
"started_at": started,
"options": options,
}
try:
response = requests.post(
ENDPOINT,
params={"token": TOKEN},
json={"url": url, "options": options},
timeout=90,
)
record["http_status"] = response.status_code
response.raise_for_status()
path = OUT / f"page-{index:04d}.webp"
path.write_bytes(response.content)
record["file"] = str(path)
record["bytes"] = len(response.content)
record["outcome"] = "response_received_review_required"
except requests.RequestException as exc:
record["outcome"] = "request_error"
record["error"] = str(exc)
record["finished_at"] = datetime.now(timezone.utc).isoformat()
manifest.write(json.dumps(record, ensure_ascii=False) + "\n")
manifest.flush()
time.sleep(1)
Replace the sample URLs with your manifest. The deliberate one-second pause is an example of pacing under your control, not a claim that this is the right or permitted rate for a publisher or a Browserless plan. The script labels successful HTTP responses for review because status alone cannot distinguish a genuine article from a CAPTCHA, access-denied page, or other unexpected result.
5. Send the same kind of request with Python or Node.js
For a single capture in Python, the request can be kept small. Refer to the endpoint documentation for the exact request shape and available options for your account.
import os
import requests
endpoint = "https://production-sfo.browserless.io/screenshot"
response = requests.post(
endpoint,
params={"token": os.environ["BROWSERLESS_TOKEN"]},
json={
"url": "https://example.com/news/article",
"options": {
"fullPage": True,
"type": "png",
"viewport": {"width": 1365, "height": 900},
"scrollPage": True,
},
},
timeout=90,
)
response.raise_for_status()
with open("article.png", "wb") as image:
image.write(response.content)
In Node.js, use fetch and write the response body to a file. This example uses built-in modules available in current Node.js releases.
import { writeFile } from "node:fs/promises";
const token = process.env.BROWSERLESS_TOKEN;
if (!token) throw new Error("Set BROWSERLESS_TOKEN first");
const endpoint = new URL("https://production-sfo.browserless.io/screenshot");
endpoint.searchParams.set("token", token);
const response = await fetch(endpoint, {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({
url: "https://example.com/news/article",
options: {
fullPage: true,
type: "webp",
quality: 80,
viewport: { width: 1365, height: 900 },
deviceScaleFactor: 1,
scrollPage: true,
},
}),
signal: AbortSignal.timeout(90_000),
});
if (!response.ok) {
const detail = await response.text();
throw new Error(`Browserless returned HTTP ${response.status}: ${detail}`);
}
await writeFile("article.webp", Buffer.from(await response.arrayBuffer()));
6. Configure the capture for news pages
| Need | Setting or approach | Notes |
|---|---|---|
| Entire article | Full-page capture | Use for a whole page, then verify it is not clipped or excessively tall. |
| Lazy-loaded images or sections | scrollPage: true with full-page mode |
Scrolling helps trigger lazy loading before capture; inspect the output for sections that still did not load. |
| Comparable captures | Consistent viewport and device scale | Changing either can alter line wrapping, image dimensions, and total page height. |
| Lossless detail | PNG | Useful when exact visual detail matters; files are often larger than compressed formats. |
| Smaller output | JPEG or WebP | Choose a format supported by the downstream system; use quality configuration where applicable. |
| Specific page region | CSS selector or clipping rectangle | Useful for a headline, article body, or fixed region; the selector must exist after the page is ready. |
| Dynamic rendering | Wait for a meaningful selector or page condition | Prefer an observable page state to an arbitrary short pause when possible. |
The Screenshot API documents output type, quality, clipping, viewport, device scale, selectors, waits, and full-page capture. It also documents scrollPage: true for pages where scrolling is needed to trigger lazy content. Browserless does not prescribe a single viewport, wait, or quality setting for Indian news sites; determine these from representative pages.
7. Make batch results observable
For every attempt, store the stable record ID, URL, capture timestamp, options, HTTP status, output path, and outcome. Add a content review status if a person or downstream validator checks whether the image contains the expected article. Keep failures and successful responses in the same manifest so omissions are visible.
- Retry only clearly transient failures, with a bounded attempt count and recorded delays.
- Do not retry a repeated CAPTCHA or access-denied result as if it were a network glitch; mark it blocked for review.
- Keep screenshot capture separate from text extraction. Browserless documents screenshots, rendered content, and structured scraping as different tasks.
- Use a stable filename derived from your record ID rather than a URL, which may contain characters unsuitable for filenames.
- Record the exact options so a later recapture can be compared against the original configuration.
The reviewed documentation gives no universal concurrency or throughput target for this workload. There is therefore no defensible number of simultaneous requests to recommend here. Increase request volume only in line with account limits and the access conditions for each target site.
8. Validate a representative set before scaling
- Open samples from each page type and check that the intended article, date, and images are present.
- Confirm the image dimensions, file type, and whether the full page was captured without unwanted clipping.
- Look for blank white output, CAPTCHA screens, access-denied or 403 content, and missing elements.
- Compare the saved image with its manifest URL and timestamp to catch mismatched filenames or stale outputs.
- Review failures before retrying, changing waits, or raising concurrency.
Browserless identifies blank screenshots, CAPTCHA pages, access-denied/403 screens, and missing elements as signs that automation may be blocked. Its documentation says: “If screenshots return blank images, CAPTCHA pages, or content that differs from what a real browser would show, the site is likely blocking automation.” See its troubleshooting guidance.
9. Troubleshooting common problems
| Symptom | Likely cause | What to do |
|---|---|---|
| Authentication error | Missing, invalid, or incorrectly supplied token | Check the token environment variable and the authentication format shown in your Browserless account documentation. Do not print the token into logs. |
| Blank or white image | Capture occurred before rendering completed, or the site blocked automation | Check the response and page state, use a meaningful wait if supported, and inspect logs. Treat repeated blank output as a failed or blocked capture. |
| CAPTCHA or 403/access denied | Publisher or intermediary blocked the automated visit | Record it as blocked and review the site’s access conditions. Do not count it as a successful article screenshot. |
| Article body or images are missing | Lazy content was not loaded, or the chosen wait condition was too early | Try scrolling before full-page capture, wait for a page-specific element, and inspect a sample. Browserless documents scrollPage: true for lazy loading. |
| Selector capture is empty | Selector is wrong or the element was absent when capture began | Check the selector against the rendered page and wait for the element before capture. |
| Output appears clipped | Viewport, full-page setting, selector, or clip rectangle does not match the intended target | Review capture mode and dimensions; compare a viewport screenshot with a full-page sample. |
| Request times out | Slow navigation, heavy assets, or a wait that never completes | Use a bounded timeout, check whether the page is reachable, and choose a wait tied to useful content rather than an unbounded condition. |
| Batch has missing records | An exception interrupted processing or results were not persisted | Write each outcome as it completes, keep a manifest, and reconcile attempted IDs against output records. |
10. Performance, reliability, and cost considerations
Full-page images and high device scale can increase image dimensions and transfer size. PNG preserves detail but may create larger files; JPEG or WebP can reduce file size when supported by the receiving system. Scrolling, selector waits, and page readiness conditions add work to each capture, so configure them only when the page needs them.
For reliability, make each URL an independently recorded job, save outcomes incrementally, bound retries, and review blocked or malformed images. A timeout should be an explicit failure record, not a missing row. Because the research sources do not establish a sustainable rate, throughput limit, or comparative price for this workload, estimate capacity from your actual account terms and a small authorized sample rather than extrapolating an unsupported benchmark.
11. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers, made by Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF, and its documented options include full-page capture with lazy images loaded, CSS selector capture, device presets, viewport and retina scale, waits, custom headers and cookies, request blocking, caching, async jobs, and bulk capture. See ScreenshotNeo and the ScreenshotNeo API documentation.
For a one-off capture, make this request and save the response body as an image:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/news/article -o shot.webp
Equivalent Python and Node.js requests are available if those fit your batch workflow:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/news/article"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/news/article' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers identify the page verdict and whether the shot was billed.
- An MCP server exposes
take_screenshot,get_page_info, andcapture_pdfto Claude, Cursor, and other MCP clients. - The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
12. Frequently asked questions
Can I capture a page that requires a login?
The researched Browserless screenshot documentation establishes URL or HTML input and capture options, but does not establish a general login-session recipe. If access requires authentication, consult the current Browserless endpoint documentation for supported request configuration and ensure your access is authorized.
Does a successful screenshot prove the page is authentic or complete?
No. A response can contain a block page, a blank page, or incomplete content. Review the image and record its validation outcome.
Can I use the screenshot as a republication of the article?
A screenshot response does not grant reuse rights. Check the publisher’s applicable terms and rights for your intended retention, publication, or redistribution.
Should I use screenshots to build a searchable archive?
Screenshots preserve appearance. If you also need searchable text, handle extraction as a separate workflow and keep its output associated with the screenshot record.


