How to Automate Website Screenshots with Python and Apify
Use Python and Playwright for automated screenshots, then package the workflow as an Apify Actor for cloud runs, storage, APIs, and schedules.
Direct answer: use Python with Playwright to launch a headless browser, navigate to the target URL, wait for the page to reach a useful state, and call Playwright’s screenshot method. Run the script locally while developing. When the job needs cloud execution, structured JSON input, persistent output, API invocation, integrations, or schedules, package the same workflow as an Apify Actor. Apify’s Python SDK is the official library for creating Python Actors, and Apify supports browser automation with Playwright.
This guide shows a complete local implementation, an Apify Actor shape, remote invocation, scheduling considerations, reliability practices, troubleshooting, and an API alternative with ScreenshotNeo.
1. Choose local Playwright or an Apify Actor
| Need | Best fit |
|---|---|
| Experiment, debugging, or a one-off image | Local Python and Playwright |
| Structured inputs and platform-managed output | Apify Actor |
| API-triggered runs, schedules, or integrations | Apify Actor |
| Maximum control over the host and files | Local script or your own server |
An Apify Actor takes structured JSON input, performs a job such as browser automation, and stores results on the platform. Apify’s supported Actor environment includes Playwright and browser binaries; local development still requires Playwright setup. See the Apify SDK for Python documentation and Actor platform documentation.
2. Install Python and Playwright locally
- Create and activate a virtual environment.
- Install Playwright.
- Install the browser binaries required by your installed Playwright version.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
# .venv\\Scripts\\Activate.ps1
python -m pip install --upgrade pip
pip install playwright
playwright install chromium
The browser installation step is required for a normal local machine. The Apify image used by the supported browser Actor template already includes Playwright and browsers.
3. Capture a website with Python and Playwright
The following script accepts a URL and common capture options from the command line. It uses a deterministic viewport, waits for network idle, and writes a PNG. Replace the readiness strategy for pages whose content appears after a specific request or selector.
import argparse
import asyncio
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import TimeoutError as PlaywrightTimeoutError
from playwright.async_api import async_playwright
def validate_url(value: str) -> str:
parsed = urlparse(value)
if parsed.scheme not in {"http", "https"} or not parsed.netloc:
raise ValueError("url must be an absolute http or https URL")
return value
async def capture(
url: str,
output: str = "page.png",
full_page: bool = True,
image_type: str = "png",
width: int = 1440,
height: int = 900,
ready_selector: str | None = None,
timeout_ms: int = 45_000,
) -> None:
validate_url(url)
if image_type not in {"png", "jpeg"}:
raise ValueError("image_type must be png or jpeg")
output_path = Path(output)
output_path.parent.mkdir(parents=True, exist_ok=True)
async with async_playwright() as playwright:
browser = await playwright.chromium.launch(headless=True)
context = await browser.new_context(
viewport={"width": width, "height": height},
device_scale_factor=1,
)
page = await context.new_page()
page.set_default_timeout(timeout_ms)
try:
await page.goto(url, wait_until="domcontentloaded", timeout=timeout_ms)
if ready_selector:
await page.locator(ready_selector).wait_for(state="visible")
else:
# Useful for pages that finish loading resources shortly after HTML.
await page.wait_for_load_state("networkidle", timeout=timeout_ms)
await page.screenshot(
path=str(output_path),
full_page=full_page,
type=image_type,
quality=85 if image_type == "jpeg" else None,
animations="disabled",
)
except PlaywrightTimeoutError as exc:
raise RuntimeError(f"timed out while loading or capturing {url}") from exc
finally:
await browser.close()
if __name__ == "__main__":
parser = argparse.ArgumentParser()
parser.add_argument("url")
parser.add_argument("--output", default="page.png")
parser.add_argument("--viewport-width", type=int, default=1440)
parser.add_argument("--viewport-height", type=int, default=900)
parser.add_argument("--viewport-only", action="store_true")
parser.add_argument("--type", choices=["png", "jpeg"], default="png")
parser.add_argument("--ready-selector")
args = parser.parse_args()
asyncio.run(
capture(
args.url,
output=args.output,
full_page=not args.viewport_only,
image_type=args.type,
width=args.viewport_width,
height=args.viewport_height,
ready_selector=args.ready_selector,
)
)
Run it with:
python screenshot.py https://example.com --output example.png
python screenshot.py https://example.com/dashboard --ready-selector "main[data-loaded='true']" --viewport-only
Playwright’s Python screenshot API supports image format, clipping, quality, and full-page capture. The official screenshot reference documents those options.
4. Full-page, viewport, element, and clipped screenshots
Full-page capture
Set full_page=True when the entire document is the subject, such as documentation, an article, or an archive. Very tall pages can produce large images and longer browser work.
Viewport capture
Set full_page=False for visual checks where the visible fold matters. Keep width, height, device scale factor, and color scheme fixed so comparisons remain meaningful.
Element capture
card = page.locator("article.product-card").first
await card.screenshot(path="card.png", type="png")
Clip a region
await page.screenshot(
path="header.png",
clip={"x": 0, "y": 0, "width": 1440, "height": 220},
type="png",
)
Element screenshots depend on the selector being present and visible. A selector that changes between deployments should be treated as an input, not hard-coded as a permanent assumption.
5. Wait for JavaScript-rendered pages correctly
wait_until="networkidle" can be useful, but it is not a universal definition of “ready.” Analytics, advertisements, and long-lived connections can prevent network idle. Prefer a meaningful readiness condition when one exists:
await page.goto(url, wait_until="domcontentloaded")
await page.locator("main[data-hydrated='true']").wait_for(state="visible")
await page.wait_for_function("document.fonts.status === 'loaded'")
Use a short, bounded delay only for a known animation or delayed widget. Playwright provides auto-waiting for browser interactions; combine that with a page-specific selector or state for dynamic applications.
6. Turn the script into an Apify Actor
Define a JSON input contract. A practical input includes the URL, full_page, image type, viewport dimensions, output name, and an optional readiness selector.
{
"url": "https://example.com",
"full_page": true,
"image_type": "png",
"viewport_width": 1440,
"viewport_height": 900,
"ready_selector": "main"
}
In an Actor, read the input, run the same Playwright capture, and push a result containing metadata. Store the image in the Actor’s platform storage according to the storage API exposed by the selected Actor template.
import asyncio
from datetime import datetime, timezone
from pathlib import Path
from apify import Actor
from playwright.async_api import async_playwright
async def main() -> None:
async with Actor:
actor_input = await Actor.get_input() or {}
url = actor_input.get("url")
if not url:
raise ValueError("input.url is required")
full_page = bool(actor_input.get("full_page", True))
image_type = actor_input.get("image_type", "png")
width = int(actor_input.get("viewport_width", 1440))
height = int(actor_input.get("viewport_height", 900))
ready_selector = actor_input.get("ready_selector")
output_name = actor_input.get("output_name", "screenshot.png")
output_path = Path("/tmp") / output_name
async with async_playwright() as playwright:
browser = await playwright.chromium.launch(headless=True)
page = await browser.new_page(viewport={"width": width, "height": height})
try:
await page.goto(url, wait_until="domcontentloaded", timeout=45_000)
if ready_selector:
await page.locator(ready_selector).wait_for(state="visible")
else:
await page.wait_for_load_state("networkidle", timeout=45_000)
await page.screenshot(
path=str(output_path),
full_page=full_page,
type=image_type,
quality=85 if image_type == "jpeg" else None,
)
finally:
await browser.close()
await Actor.push_data({
"url": url,
"screenshot_path": str(output_path),
"captured_at": datetime.now(timezone.utc).isoformat(),
"viewport": {"width": width, "height": height},
"full_page": full_page,
"image_type": image_type,
})
if __name__ == "__main__":
asyncio.run(main())
The exact file-upload call depends on the Actor template and storage choice. Keep the metadata record even when the binary is stored separately so downstream systems can identify the URL, viewport, timestamp, and capture mode.
Apify’s documentation describes the Actor lifecycle, structured input, platform storage, and browser automation capabilities. Adapt imports and storage calls to the SDK version selected for the project.
7. Run the Actor remotely
Use the Apify API or the official Python client to start a run with JSON input. The client workflow is: create an ApifyClient, invoke the Actor, inspect the run, and read its dataset or storage output.
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("YOUR_USERNAME/YOUR_ACTOR").call(
run_input={
"url": "https://example.com",
"full_page": True,
"image_type": "png",
"viewport_width": 1440,
"viewport_height": 900,
}
)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)
For a production integration, persist the run ID, check the run status, and handle failed runs separately from successful runs that produce an image. The official Python client example documents invocation and dataset iteration: Apify API client for Python.
8. Schedule recurring screenshots
- Make the Actor input deterministic and store the target URL and viewport in the schedule configuration.
- Choose a frequency appropriate to the page and the purpose of the capture.
- Write stable output names or include a timestamp in metadata.
- Send the run result to the downstream system that compares, archives, or reviews the image.
- Alert on navigation errors, timeouts, missing readiness selectors, and empty output.
Apify supports manual starts, API calls, schedules, and integrations. Scheduling does not remove the need to respect the target site’s terms, robots directives, authentication boundaries, and privacy requirements.
9. Reliability checklist
- Use a fixed viewport and record it with every result.
- Prefer a meaningful selector or application state over an arbitrary long sleep.
- Set bounded navigation and selector timeouts.
- Retry transient navigation failures with a small maximum retry count.
- Keep output names stable enough for downstream systems to locate them.
- Disable or account for animations when pixel consistency matters.
- Decide how cookie banners, ads, lazy images, and chat widgets should appear before capture.
- Use full-page mode only when page length is relevant.
- Record the final URL after redirects and the capture timestamp.
- Do not capture private or authenticated content unless you have permission and an appropriate data-handling policy.
10. Performance, file size, and cost considerations
Viewport screenshots are generally faster and smaller than full-page images. PNG preserves lossless detail but can be large; JPEG reduces size with lossy compression and requires a quality value. A larger viewport, retina scale, heavy page, lazy-loaded images, and additional readiness checks all increase browser work.
For Apify, account for Actor runtime, browser memory, storage, and the frequency of scheduled runs. Keep the browser context short-lived, avoid unnecessary pages, and reject invalid URLs before launching Chromium. The research sources do not establish a universal Apify screenshot price; check the current Apify pricing for the selected resources.
11. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
Executable doesn't exist |
Local browser binaries were not installed. | Run playwright install chromium in the active environment. |
| Blank or incomplete page | Capture happened before client-side rendering finished. | Wait for a meaningful selector or application state after navigation. |
| Network-idle timeout | Analytics, ads, or persistent connections never become idle. | Use domcontentloaded plus a readiness selector and a bounded timeout. |
| Missing lazy images | Images load only after scrolling or intersection. | Scroll the page before capture or use a site-specific readiness condition. |
| Cookie banner obscures content | The page requires consent interaction. | Locate and accept or hide the banner when permitted, then wait for it to disappear. |
| Element selector fails | The selector is absent, hidden, or changed. | Verify it in the target deployment and wait for visibility. |
| Actor input is empty | The run was started without valid JSON input. | Require url and validate every option before launching the browser. |
| Remote run succeeds but image is unavailable | Only metadata was pushed, or the binary was not uploaded to storage. | Use the selected Apify storage API for the image and return its key or URL in metadata. |
| Huge output file | Very tall page, large viewport, or lossless PNG. | Capture the viewport, reduce dimensions, or use JPEG where visual requirements allow. |
12. Or skip the browser setup
ScreenshotNeo provides a website screenshot API: one GET request returns PNG, JPEG, WebP, or PDF. The API accepts the URL and capture options without maintaining Playwright, Chromium, browser binaries, or an Apify Actor.
See the ScreenshotNeo API documentation for the complete option list.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo can accept cookie and consent banners before capture and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also provides an MCP server for AI agents, including Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools.
Options include full-page capture with lazy images loaded, CSS element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and margin settings, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request and resource blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification.
Free usage includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan. Create a free ScreenshotNeo account.
13. FAQ
Can Apify run a Playwright screenshot Actor?
Yes. Apify supports browser automation with Playwright, and its Python SDK is the official library for creating and running Python Actors.
Should a visual monitor use full-page screenshots?
Use full-page mode when changes anywhere in the document matter. Use a fixed viewport for fold-level checks or faster, smaller artifacts.
Why is a readiness selector better than a long sleep?
A selector expresses the page state the capture actually needs. A fixed sleep can be too short for a slow run and unnecessarily long for a fast one.
How should screenshots be stored?
Store the binary in the platform or object storage and keep metadata for URL, timestamp, viewport, capture mode, and final URL. This makes later comparison and debugging possible.
Is an Apify Actor required for a single screenshot?
No. A local Playwright script is sufficient for one-off work. Use an Actor when managed cloud execution, API calls, storage, integrations, or schedules are useful.


