ScreenshotNeo

BlogHow-to

How to Capture WordPress Websites with an API

Learn how to capture WordPress content with the REST API, authenticate safely, paginate results, handle media, and choose screenshots when you need rendered pages.

By the ScreenshotNeo team1 October 202610 min read

Direct answer: To capture a WordPress website as structured data, discover its site-specific REST API at /wp-json/, choose a resource route such as /wp-json/wp/v2/posts, /wp/v2/pages or /wp/v2/media, then make HTTP requests and parse the JSON response. Public content usually works without credentials. Private content and write operations require the permissions and authentication configured by that site.

This guide uses “capture” to mean collecting WordPress content and metadata through the REST API. The REST API returns JSON; it does not by itself create a rendered screenshot or a complete backup. If you need a visual image of a page, use a browser capture workflow such as ScreenshotNeo instead.

1. Understand the WordPress REST API

WordPress describes its REST API as an interface for applications to interact with a site by sending and receiving JSON objects. The API is distributed per site, so there is no single universal root for every installation. Start with the site’s own API index and inspect the routes it exposes.

Primary references: REST API Handbook and the REST API Reference.

Discover the API root

curl -i https://example.com/wp-json/

A successful response is JSON describing namespaces, routes and available links. The common built-in namespace is wp/v2, but plugins and custom code can add others. A site may also use a different front-controller or host configuration, so use the discovered links and route descriptions rather than assuming every route exists.

Check a route before collecting data

curl -i "https://example.com/wp-json/wp/v2/posts?per_page=5"

Inspect the HTTP status, response headers and JSON fields. A route being listed does not guarantee that your caller can read every record or field.

2. Capture public posts

The posts collection is normally available anonymously when posts are public. The endpoint reference documents pagination, search and date filters; confirm the exact arguments for any custom route.

cURL

curl --fail-with-body -sS \
  "https://example.com/wp-json/wp/v2/posts?per_page=10&page=1&_fields=id,date,slug,title,content,link"

Python

import requests

base = "https://example.com/wp-json/wp/v2/posts"
params = {
    "per_page": 10,
    "page": 1,
    "_fields": "id,date,slug,title,content,link",
}
response = requests.get(base, params=params, timeout=30)
response.raise_for_status()
posts = response.json()

for post in posts:
    print(post["id"], post["slug"], post["link"])

Node.js

const endpoint = new URL("https://example.com/wp-json/wp/v2/posts");
endpoint.searchParams.set("per_page", "10");
endpoint.searchParams.set("page", "1");
endpoint.searchParams.set("_fields", "id,date,slug,title,content,link");

const response = await fetch(endpoint);
if (!response.ok) {
  throw new Error(`WordPress returned ${response.status}: ${await response.text()}`);
}
const posts = await response.json();
for (const post of posts) {
  console.log(post.id, post.slug, post.link);
}

Important response fields

Field Use
id Stable WordPress post identifier.
date and modified Publication and last-modified timestamps.
slug URL-friendly identifier.
title.rendered Rendered title HTML.
content.rendered Rendered post HTML.
content.protected Whether the content is protected.
link Canonical public URL.
_links Related resources, including author and featured media links when exposed.

3. Paginate and filter large collections

Do not assume one request returns every record. Collection responses include pagination headers such as X-WP-Total and X-WP-TotalPages when the server provides them. Request a bounded per_page value (the endpoint documents the allowed range), then advance page until the final page.

import requests

url = "https://example.com/wp-json/wp/v2/pages"
page = 1
all_pages = []

while True:
    response = requests.get(
        url,
        params={"page": page, "per_page": 100, "orderby": "modified", "order": "desc"},
        timeout=30,
    )
    if response.status_code == 400 and page > 1:
        break  # WordPress reports an invalid page after the last page
    response.raise_for_status()
    batch = response.json()
    all_pages.extend(batch)
    total_pages = int(response.headers.get("X-WP-TotalPages", page))
    if page >= total_pages or not batch:
        break
    page += 1

print(f"Collected {len(all_pages)} pages")

Useful documented filters for built-in collections include search, after, before, author, include, exclude, orderby and order. The exact filter set varies by endpoint and plugin. Use _fields to reduce response size when you only need selected properties.

4. Capture pages and media

Pages

Pages use the parallel route /wp-json/wp/v2/pages. The pages reference documents collection parameters such as page, per_page, search and date filters.

curl -G "https://example.com/wp-json/wp/v2/pages" \
  --data-urlencode "slug=about" \
  --data-urlencode "_fields=id,slug,title,content,link"

Media

Media has its own endpoint at /wp-json/wp/v2/media. You can list public attachment metadata and follow fields such as source_url, mime_type, media_details and alt_text when the site exposes them. See the media endpoint reference.

curl -G "https://example.com/wp-json/wp/v2/media" \
  --data-urlencode "media_type=image" \
  --data-urlencode "per_page=20" \
  --data-urlencode "_fields=id,date,slug,source_url,mime_type,alt_text,media_details"

Upload requests are host- and configuration-dependent. Verify the endpoint’s required headers, file format, authentication and size limits against the target installation before automating uploads. WordPress.com documents a separate media upload flow in its media API documentation.

5. Authenticate for private data and writes

Anonymous requests generally expose published public content. Drafts, private or password-protected content, user-specific metadata and write operations require authentication and the relevant WordPress capability.

Self-hosted WordPress: Application Passwords

For self-hosted sites, WordPress documents Application Passwords as revocable, per-application credentials. Create one for the integration and send it over HTTPS with HTTP Basic Authentication. Do not store or distribute a user’s normal interactive login password. Read the Application Passwords security documentation.

curl --user "API_USER:APPLICATION_PASSWORD" \
  "https://example.com/wp-json/wp/v2/posts?context=edit&per_page=5"
import os
import requests

response = requests.get(
    "https://example.com/wp-json/wp/v2/posts",
    params={"context": "edit", "per_page": 5},
    auth=(os.environ["WP_USER"], os.environ["WP_APPLICATION_PASSWORD"]),
    timeout=30,
)
response.raise_for_status()
print(response.json())

Use the smallest account capability that meets the job, keep credentials in a secret store or environment variables, rotate or revoke them when an integration changes, and never log the Authorization header.

Creating a post

The posts endpoint supports creation with POST /wp-json/wp/v2/posts when the authenticated user has permission.

curl --user "API_USER:APPLICATION_PASSWORD" \
  -H "Content-Type: application/json" \
  -d '{"title":"API draft","content":"Created through the REST API.","status":"draft"}' \
  "https://example.com/wp-json/wp/v2/posts"

Validate the returned ID and status, and make write operations idempotent in your application where possible. A successful HTTP request does not mean a plugin workflow, editorial rule or cache has completed.

WordPress.com

WordPress.com uses its own URL patterns and access-token setup. Its documented service covers WordPress.com sites and Jetpack-connected self-hosted sites. Follow the WordPress.com API getting-started guide instead of copying self-hosted /wp-json/ assumptions.

6. Custom post types, taxonomies and metadata

A custom post type is available through the REST API only when it is registered for REST exposure (commonly with show_in_rest) and the caller has permission. The route and namespace may differ from wp/v2. Custom fields likewise require explicit exposure and may be hidden by permissions or a plugin.

  • Inspect the API index for the namespace and route.
  • Check the route’s schema and collection arguments.
  • Confirm the post type is publicly queryable or authenticated for your use case.
  • Do not infer that a field exists because it appears in the WordPress admin.

7. Data capture versus screenshots and backups

The REST API gives you structured records: titles, rendered HTML, dates, links, taxonomy references and media metadata. It does not prove that a page visually rendered correctly, include every asset loaded by a browser, or constitute a complete database and uploads backup. For a backup, use a backup workflow that covers the database and files. For a rendered visual, use a browser screenshot service.

8. Or skip the browser setup

If your goal is a rendered screenshot of a WordPress URL, ScreenshotNeo provides a single GET request that returns PNG, JPEG, WebP or PDF. Its capture steps can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets before the shot; each step can be turned off.

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

See the ScreenshotNeo API documentation for all options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture with lazy images loaded, CSS element capture, dark mode, device presets or custom viewports, retina scale, PDF controls, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, selectable cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work.

The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is on every plan. Start with 1,000 free screenshots.

9. Troubleshooting

Symptom Likely cause Fix
404 on /wp-json/ Wrong host, rewrite configuration or a nonstandard API setup. Open the site in a browser, verify HTTPS and inspect the site’s API documentation or hosting configuration.
401 or 403 Missing, invalid or insufficient credentials; application passwords may be disabled. Use the site’s documented authentication method, check account capability and send credentials only over HTTPS.
200 response but missing private fields The request is anonymous or lacks the required context and permission. Authenticate and request the documented context; confirm the field is explicitly exposed.
400 “invalid page” The requested page exceeds the collection’s available pages. Read X-WP-TotalPages or stop when a page returns no records.
429 or intermittent 5xx Host rate limits, overloaded PHP workers, a proxy or a plugin. Use bounded pages, exponential backoff with jitter, caching and a lower request rate; inspect server logs.
HTML appears escaped or incomplete You are reading the JSON representation incorrectly or relying on content that a plugin renders only in the browser. Parse content.rendered as HTML and determine whether the missing data comes from a separate endpoint or client-side script.
Media URL fails Hotlink protection, private media, CDN rules or a deleted attachment. Check the media record, permissions and the returned URL directly; do not assume every attachment is public.
Screenshot includes a popup A visual capture tool did not dismiss the site’s consent or widget elements. Use a browser workflow that supports waits, clicks and selector hiding, or use ScreenshotNeo’s cleanup steps.

10. Performance, reliability and cost

  • Reduce payloads: request only needed fields with _fields, use reasonable per_page values and avoid repeatedly downloading rendered HTML.
  • Cache safely: cache public responses according to your freshness requirement and use the post’s modified value to detect changes.
  • Handle pagination explicitly: persist progress so a failed page does not restart a large collection.
  • Retry carefully: retry transient 429 and 5xx responses with bounded exponential backoff; do not retry authentication failures indefinitely.
  • Respect the site: coordinate with the owner for high-volume jobs and account for host CPU, database load, CDN rules and plugin behavior.
  • Separate data and visual jobs: REST requests are usually cheaper and more deterministic for content extraction; browser rendering is needed when layout, JavaScript or the final pixels matter.
  • ScreenshotNeo billing: only clean shots are billed; failed loads, blank pages, bot checks, timeouts and cache hits are not billed, and the response headers report the verdict and billing result.

11. Practical checklist

  1. Confirm whether you need JSON content, a backup or a rendered screenshot.
  2. For self-hosted WordPress, open https://your-site.example/wp-json/ and inspect routes.
  3. Choose the resource endpoint: posts, pages, media or an exposed custom type.
  4. Start anonymously for public data; authenticate only when the resource and operation require it.
  5. Use pagination, field selection, timeouts, retry limits and a progress checkpoint.
  6. Verify permissions and exposure for private records, custom fields and writes.
  7. For visual output, use a browser capture service and verify how it handles consent banners, failed pages and billing.

FAQ

Does the WordPress REST API return a screenshot?

No. It returns structured JSON. Use a browser screenshot workflow for rendered pixels.

Can I read every post without logging in?

Only public, exposed content is generally anonymous. Drafts, private posts and restricted fields require permission and authentication.

Is /wp-json/wp/v2/ universal?

It is common for self-hosted WordPress, but routes are site-specific and can be changed or extended by plugins and custom code.

Should I use my normal WordPress password in a script?

No. Use a revocable Application Password for self-hosted integrations, or the token flow documented by WordPress.com.

How do I capture a custom post type?

Find its exposed route in the API index and confirm its registration, schema and permissions. A post type visible in wp-admin may not be available through REST.

What is the quickest way to screenshot many WordPress URLs?

Use a screenshot API with bulk and asynchronous options. ScreenshotNeo supports bulk capture for up to 100 URLs per call and provides signed webhooks for async jobs.