How to Capture a Screenshot of a Canonical URL After Redirects
Follow redirects with Playwright, record the destination, and capture the page when it is ready. Learn how that differs from a page’s declared SEO canonical URL.
To screenshot the page reached after redirects, navigate to the starting URL in a real browser, record the browser’s final URL, check the response status, wait for the page condition you need, and then capture it. In Playwright, page.goto() follows navigation redirects and returns the main resource response for the first non-redirect response in the chain. Use page.url() to record the address the browser ended on.
“Canonical URL” can mean two different things: the final address reached by navigation, or a URL the page declares in an SEO canonical link. A redirect to a destination does not prove that destination declares itself canonical. This guide captures the visited destination and shows how to inspect the declared canonical separately.
1. Set up Playwright
The examples use Node.js and Playwright. Install the package and a browser in a project directory:
npm init -y
npm install playwright
npx playwright install chromium
Save the following as screenshot-redirect.js. Pass the starting URL as the first command-line argument. The URL must include a scheme such as https://.
2. Navigate, record redirects, and capture
const { chromium } = require('playwright');
async function main() {
const startUrl = process.argv[2];
if (!startUrl) {
throw new Error('Usage: node screenshot-redirect.js https://example.com');
}
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
const response = await page.goto(startUrl, {
waitUntil: 'domcontentloaded',
timeout: 45_000
});
const finalUrl = page.url();
const status = response?.status() ?? null;
const redirectChain = [];
let request = response?.request();
while (request) {
redirectChain.unshift(request.url());
request = request.redirectedFrom();
}
console.log(JSON.stringify({ startUrl, finalUrl, status, redirectChain }, null, 2));
if (status === null || status >= 400) {
throw new Error(`Navigation did not return a successful HTTP status: ${status}`);
}
// Replace this with an application-specific readiness condition when needed.
await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
await browser.close();
}
}
main().catch(error => {
console.error(error);
process.exitCode = 1;
});
Run it with:
node screenshot-redirect.js https://example.com/old-path
The log includes the requested URL, final browser URL, response status, and the server redirect chain represented by the response request and its redirectedFrom() links. Playwright documents that a completed navigation may still have an HTTP error status such as 404 or 500, so inspect the response instead of treating navigation completion as proof of success. See the official Playwright Page API, Request API, and network documentation.
3. Choose what “canonical” means for your capture
Capture the destination reached by redirects
This is the usual interpretation when documenting where a URL sends a browser. Capture the current page after page.goto(), and store page.url() alongside the image. This records the destination that was actually visited.
Inspect the page’s declared SEO canonical URL
A page may declare a canonical link in its HTML. Read it independently from the final browser URL:
const declaredCanonical = await page.locator('link[rel="canonical"]').first().getAttribute('href').catch(() => null);
console.log({ finalUrl: page.url(), declaredCanonical });
The canonical value may be relative, absent, or one of several matching elements in malformed markup. Resolve a relative value against the document URL and decide how your workflow handles missing or multiple declarations. Reading the declaration does not navigate the browser to that address. If your requirement is to screenshot the page at the declared canonical target, validate that target and navigate to it explicitly; record both the original destination and the canonical target so the artifact’s provenance is clear.
4. Configure readiness and screenshot scope
Navigation wait conditions
page.goto() supports readiness points such as commit, domcontentloaded, and load. Choose one that fits the site, then wait for the specific application state needed for a meaningful image. For example:
await page.goto(startUrl, { waitUntil: 'domcontentloaded' });
await page.locator('main article').waitFor({ state: 'visible', timeout: 15_000 });
await page.screenshot({ path: 'article.png', fullPage: true });
For a simple static page, waiting for the load event can be sufficient. For a client-rendered page, wait for a visible content selector or another app-specific signal. A fixed delay can help with a known animation or delayed widget, but it is less reliable than waiting for the condition itself. Playwright’s Page API marks networkidle as discouraged as a general testing readiness rule; analytics, polling, and long-lived connections can prevent it from becoming idle.
Viewport or full page
By default, a screenshot covers the viewport. Set fullPage: true to capture the full scrollable page. Full-page images can be very tall and use more memory; choose viewport capture if the question is what a visitor sees above the fold. Set the viewport when reproducibility matters. Use page.screenshot({ path: 'page.jpg', type: 'jpeg', quality: 80 }) for JPEG, or type: 'png' for PNG. Screenshot type and path are documented by Playwright’s Page API.
5. cURL, Python, and Node.js alternatives
cURL can follow HTTP redirects and save the final response body, but it does not render a browser page or create a screenshot image. It is useful for inspecting the redirect destination and response headers:
curl -L -sS -D redirect-headers.txt -o /dev/null -w 'final_url=%{url_effective}\nstatus=%{http_code}\n' 'https://example.com/old-path'
For HTML inspection, save the response body instead of discarding it:
curl -L -sS 'https://example.com/old-path' -o page.html
Python’s requests library follows redirects by default for GET requests and exposes the redirect history, but it also does not render the page. Use Playwright’s Python bindings when you need a browser screenshot:
from playwright.sync_api import sync_playwright
start_url = 'https://example.com/old-path'
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page(viewport={"width": 1440, "height": 900})
response = page.goto(start_url, wait_until='domcontentloaded', timeout=45_000)
print({
"start_url": start_url,
"final_url": page.url,
"status": response.status if response else None,
})
if response is None or response.status >= 400:
raise RuntimeError(f"Unsuccessful navigation: {response.status if response else 'no response'}")
page.screenshot(path='page.png', full_page=True)
browser.close()
Install and provision it with pip install playwright and playwright install chromium. For Node.js Playwright, the runnable code above uses Chromium and can be adapted to Firefox or WebKit if those engines are installed.
6. Or skip the browser setup
For a one-call screenshot, use ScreenshotNeo, a website screenshot API and MCP server for developers. Its API accepts a URL and returns an image or PDF. See the ScreenshotNeo API documentation for parameters and options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/old-path -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/old-path"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/old-path' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.
7. Reliability, performance, and cost
- Record provenance: Save the starting URL, final browser URL, response status, timestamp, browser version, viewport, and readiness rule with the image when captures need to be repeatable.
- Use bounded timeouts: Set a navigation timeout and a separate timeout for app readiness. A slow site should fail clearly rather than leave a job hanging indefinitely.
- Keep browser resources bounded: Reuse a browser for a batch of pages, create isolated pages or contexts as appropriate, and close them in a
finallyblock. Full-page screenshots of long pages require more memory than viewport captures. - Retry selectively: Retry transient network failures with a limit and backoff. Do not blindly retry 4xx responses or a deterministic bad URL. Recheck the final URL and status on each attempt.
- Account for local compute: Self-hosted Playwright costs depend on the machine, browser runtime, concurrency, and frequency of captures. Browser screenshots are not just an HTTP download; pages execute scripts and load resources.
- Account for API usage: ScreenshotNeo plans are Free 1,000/month, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan. See current details at ScreenshotNeo.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
page.goto() says the URL is invalid |
The input has no URL scheme or is malformed. | Pass a complete address such as https://example.com. |
| The script captured a 404 or 500 page | Playwright treats a valid HTTP response as a completed navigation; HTTP error statuses do not necessarily throw. | Check response.status() and decide whether to retain error-page screenshots or fail the capture. |
| No response object was returned | The navigation may have been interrupted or may not have produced a main-resource response. | Handle the null response, inspect the navigation error, and check whether the page is a download or another special case. |
| The screenshot shows a loading shell | The document event occurred before the application rendered its content. | Wait for a stable, visible content selector or application-specific readiness signal before taking the image. |
| Navigation times out on a page that keeps making requests | Waiting for network idle may be inappropriate for analytics, polling, or persistent connections. | Use a more suitable navigation event and wait for the element that proves the required content is ready. |
| The final URL differs from the SEO canonical | Redirect destination and declared canonical are separate signals. | Read link[rel="canonical"] from the loaded document. Navigate to that address separately only if the task requires a screenshot of the declared target. |
| The redirect chain appears incomplete | The code may be inspecting the wrong request, or the navigation involved client-side routing rather than HTTP redirects. | Start from response.request() and walk redirectedFrom(); record the final URL too. Client-side navigation may not appear as a server redirect chain. |
| Full-page capture is unexpectedly large or slow | The page is very tall, contains many images, or loads content as the page scrolls. | Use viewport capture when sufficient, wait for lazy-loaded content if required, and avoid unnecessary concurrent full-page captures. |
| Certificate or connection errors | The host is unreachable, TLS is invalid, or the environment cannot access the destination. | Check the URL from the same machine, DNS and network access, and the site’s certificate. Avoid disabling certificate checks except in a controlled environment where that is explicitly required. |
9. Frequently asked questions
Does a redirect tell me the SEO canonical URL?
No. It tells you where navigation ended. Inspect the document’s canonical link separately if you need its declared SEO target.
Should I screenshot the original URL or the final URL?
Capture after navigation when you want the page a visitor reaches, and record both addresses. If you need the original response itself, capture it through an HTTP inspection workflow; a browser normally proceeds through its redirects.
Can I use a screenshot to verify a redirect?
A screenshot shows rendered appearance, not the full redirect evidence. Save the final URL, status, and redirect request history with it.
Does Playwright follow redirects automatically?
For ordinary browser navigation through page.goto(), it follows navigation redirects and resolves with the main resource response for the first non-redirect response. Check the returned status and inspect the final page URL.


