ScreenshotNeo

BlogHow-to

Render a Long HTML Page as a Tiled JPEG Image

Use Playwright to render HTML as a full-page JPEG, then crop it into fixed-size tiles. Includes runnable code, settings, and troubleshooting.

By the ScreenshotNeo team4 October 202611 min read

To render a long HTML page as JPEG tiles, first let a browser render the page, capture it as a full-page JPEG buffer, then crop that buffer into tiles with an image-processing library. A screenshot library does not automatically split a full-page capture into fixed-size tiles. This guide uses Playwright with Node.js and Sharp; it also shows Playwright in Python, Puppeteer, and a direct screenshot API option.

1. Choose full-page capture or viewport tiles

There are two useful workflows:

  • Capture once, crop afterward: simplest when a single tall image fits your renderer and memory budget. Capture with fullPage: true, then crop the JPEG buffer into rows and columns.
  • Capture page rectangles: useful when you need fixed-size outputs without first creating one very tall bitmap. Scroll and capture defined viewport regions, or use clipping where supported. You must handle sticky elements, lazy loading, overlap, and coordinate scaling.

Playwright documents full-page capture, JPEG quality, clipping, and screenshot output as a buffer. Its scale option controls whether output uses CSS-pixel or device-pixel dimensions. A full-page capture produces one tall image; cropping or multiple captures are extra processing steps. There is no universal tile size or browser maximum established by the cited documentation, so validate dimensions and memory use with the browser, page, and image library you deploy. Playwright Page API · Playwright Screenshots guide.

2. Runnable Node.js example: render, then crop into JPEG tiles

This example navigates to a URL, waits for the document load event, takes a full-page JPEG buffer, and writes tiles of at most 1200 by 1600 pixels. The final row or column can be smaller. It does not add overlap.

npm install playwright sharp
npx playwright install chromium
// save as tile-page.mjs
import { chromium } from 'playwright';
import sharp from 'sharp';
import { mkdir } from 'node:fs/promises';

const url = process.argv[2] ?? 'https://example.com';
const outputDir = process.argv[3] ?? 'tiles';
const tileWidth = 1200;
const tileHeight = 1600;

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
try {
  const page = await browser.newPage({ viewport: { width: tileWidth, height: 900 } });
  await page.goto(url, { waitUntil: 'load', timeout: 60_000 });
  // If the page has lazy-loaded content, see the notes below before capture.
  const jpeg = await page.screenshot({
    type: 'jpeg',
    quality: 80,
    fullPage: true,
    scale: 'css'
  });
  const metadata = await sharp(jpeg).metadata();
  if (!metadata.width || !metadata.height) throw new Error('Could not read screenshot dimensions');

  let index = 0;
  for (let top = 0; top < metadata.height; top += tileHeight) {
    for (let left = 0; left < metadata.width; left += tileWidth) {
      const width = Math.min(tileWidth, metadata.width - left);
      const height = Math.min(tileHeight, metadata.height - top);
      await sharp(jpeg)
        .extract({ left, top, width, height })
        .jpeg({ quality: 80 })
        .toFile(`${outputDir}/tile-${String(index).padStart(4, '0')}.jpg`);
      index += 1;
    }
  }
  console.log(`Wrote ${index} tiles from ${metadata.width}x${metadata.height} screenshot to ${outputDir}`);
} finally {
  await browser.close();
}
node tile-page.mjs https://example.com output-tiles

Tile order is row-major: left to right, then top to bottom. The input screenshot is decoded as a single image for each crop operation by Sharp; for very large pages this can demand substantial memory. If that is unsuitable, use viewport-region captures and write each region directly, or reduce the viewport width and scale.

Use HTML instead of navigating to a URL

For supplied markup, replace page.goto with page.setContent. External scripts, stylesheets, and images still need to be reachable and loaded if they are part of the intended rendering.

await page.setContent(`<!doctype html>
<html><head><style>body{font:16px sans-serif}</style></head>
<body><h1>Rendered HTML</h1><p>Page content</p></body></html>`, {
  waitUntil: 'load',
  timeout: 60_000
});

Playwright’s Page API documents setContent and screenshot settings. Its screenshots guide describes getting image data as a buffer for further processing.

3. Make tiles by capturing page regions

If a huge full-page bitmap is impractical, capture viewport-height strips and crop the last strip as needed. This approach still creates image buffers, but avoids constructing one full-page screenshot. The example scrolls to CSS-pixel offsets and captures the viewport. Sticky headers may appear in every tile, and content that changes during scrolling can cause seams or inconsistent states.

const viewport = { width: 1200, height: 1600 };
const page = await browser.newPage({ viewport, deviceScaleFactor: 1 });
await page.goto(url, { waitUntil: 'load' });
const fullHeight = await page.evaluate(() => document.documentElement.scrollHeight);
let tile = 0;
for (let y = 0; y < fullHeight; y += viewport.height) {
  await page.evaluate((offset) => window.scrollTo(0, offset), y);
  await page.waitForTimeout(150); // allow scroll-triggered layout to settle; tune for the site
  const buffer = await page.screenshot({ type: 'jpeg', quality: 80 });
  const height = Math.min(viewport.height, fullHeight - y);
  await sharp(buffer)
    .extract({ left: 0, top: 0, width: viewport.width, height })
    .jpeg({ quality: 80 })
    .toFile(`tiles/strip-${String(tile++).padStart(4, '0')}.jpg`);
}

This basic strip method is not equivalent to a robust archival capture for every site. It does not force lazy images to load, freeze animations, or account for a sticky header consuming part of each viewport. For exact tiling, decide whether tile coordinates mean CSS pixels or output pixels, use scale: 'css' or device scale deliberately, and test the last partial strip and any overlap in the target browser.

4. Python example: Playwright capture and Pillow cropping

Install Playwright and Pillow, then install a browser. This captures the page as JPEG bytes and crops those bytes into fixed-size tiles.

python -m pip install playwright Pillow
python -m playwright install chromium
# save as tile_page.py
import asyncio
from io import BytesIO
from pathlib import Path
import sys
from PIL import Image
from playwright.async_api import async_playwright

async def main():
    url = sys.argv[1] if len(sys.argv) > 1 else 'https://example.com'
    out = Path(sys.argv[2] if len(sys.argv) > 2 else 'tiles')
    out.mkdir(parents=True, exist_ok=True)
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        try:
            page = await browser.new_page(viewport={"width": 1200, "height": 900})
            await page.goto(url, wait_until='load', timeout=60_000)
            data = await page.screenshot(type='jpeg', quality=80, full_page=True, scale='css')
            image = Image.open(BytesIO(data))
            tile_w, tile_h = 1200, 1600
            count = 0
            for top in range(0, image.height, tile_h):
                for left in range(0, image.width, tile_w):
                    right = min(left + tile_w, image.width)
                    bottom = min(top + tile_h, image.height)
                    image.crop((left, top, right, bottom)).save(
                        out / f'tile-{count:04d}.jpg', 'JPEG', quality=80
                    )
                    count += 1
            print(f'Wrote {count} tiles from {image.width}x{image.height} screenshot to {out}')
        finally:
            await browser.close()

asyncio.run(main())
python tile_page.py https://example.com output-tiles

For HTML strings, use await page.set_content(html, wait_until='load') instead of goto. The image is held in memory as both encoded bytes and a decoded bitmap; for exceptionally tall pages, measure actual memory use and prefer viewport strips if needed.

5. Puppeteer alternative

If your project already uses Puppeteer, its screenshot options include full-page capture, clipping, capture beyond the viewport, output type, and JPEG quality. Check the documentation for the version installed in your project before depending on option defaults or combinations.

npm install puppeteer sharp
// save as puppeteer-tiles.mjs
import puppeteer from 'puppeteer';
import sharp from 'sharp';
import { mkdir } from 'node:fs/promises';

const url = process.argv[2] ?? 'https://example.com';
const out = process.argv[3] ?? 'tiles';
const tileW = 1200, tileH = 1600;
await mkdir(out, { recursive: true });
const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.setViewport({ width: tileW, height: 900, deviceScaleFactor: 1 });
  await page.goto(url, { waitUntil: 'load', timeout: 60_000 });
  const buffer = await page.screenshot({ type: 'jpeg', quality: 80, fullPage: true });
  const meta = await sharp(buffer).metadata();
  let n = 0;
  for (let top = 0; top < meta.height; top += tileH) {
    for (let left = 0; left < meta.width; left += tileW) {
      const width = Math.min(tileW, meta.width - left);
      const height = Math.min(tileH, meta.height - top);
      await sharp(buffer).extract({ left, top, width, height })
        .jpeg({ quality: 80 }).toFile(`${out}/tile-${String(n++).padStart(4, '0')}.jpg`);
    }
  }
  console.log(`Wrote ${n} tiles`);
} finally {
  await browser.close();
}

Puppeteer’s ScreenshotOptions interface documents its capture controls. Playwright and Puppeteer both provide browser-based rendering; choose based on the runtime and browser setup already supported by your application, not on an undocumented speed ranking.

6. JPEG, scale, dimensions, and image quality

Setting Effect Practical guidance
type: 'jpeg' Produces a lossy JPEG. JPEG does not preserve transparency; use PNG if transparent output is required.
quality Controls JPEG quality from 0 to 100 in Playwright; documented default is 80. Start at 80 and inspect small text, gradients, and file size for your content.
fullPage: true Captures the scrollable page as one image. Crop the result afterward if you need tiles.
scale: 'css' Uses CSS-pixel dimensions. Often makes tile dimensions easier to reason about.
scale: 'device' Uses device-pixel dimensions and can produce a larger image. Account for the scale factor when mapping CSS coordinates to pixel crops.
Clip rectangle Captures a defined region in APIs that support clipping. Useful for controlled regions; verify whether coordinates are relative to the page or viewport in the installed tool version.

These Playwright settings are described in the official Page API. Exact supported behavior can depend on library version and browser, so consult the matching docs when upgrading.

7. Handle lazy content, fonts, and page state

  1. Wait for the right readiness condition. load means page resources have reached that event; a single-page application may still be rendering data. Wait for a known selector when the page has a stable completion marker.
  2. Load lazy images. Scroll through the page in increments before the final capture, allowing images and scroll-triggered sections to load. A full-page screenshot option does not promise to trigger every site’s lazy-loading logic in the way its application expects.
  3. Wait for fonts and layout. If web fonts or late content shift the page, wait for the relevant font or element before measuring height and capturing.
  4. Stabilize animated content. Disable animations or hide blinking content with page-specific CSS if deterministic tiles matter. This is especially important for strips captured at different times.
  5. Set a fixed viewport. Responsive layout changes with viewport width. Set viewport dimensions and device scale explicitly so runs are comparable.

A simple lazy-load pre-pass can scroll down the page, then return to the top. Tune the delay to the page rather than assuming one delay works universally:

await page.evaluate(async () => {
  const step = Math.max(300, window.innerHeight);
  for (let y = 0; y < document.documentElement.scrollHeight; y += step) {
    window.scrollTo(0, y);
    await new Promise(resolve => setTimeout(resolve, 100));
  }
  window.scrollTo(0, 0);
});
await page.waitForTimeout(300);

8. Troubleshooting

Symptom Likely cause Fix
Tiles are blank or content is missing Capture happened before client rendering, images, or fonts finished. Wait for a stable selector or app-specific ready state; scroll through lazy content; verify external resources load.
The last tile is the wrong size The page dimensions are not an exact multiple of tile dimensions. Use min(tileSize, remainingPixels) for each crop, as in the examples.
Seams or duplicated content between strips Scroll timing, sticky elements, viewport overlap, or dynamic layout changes. Capture a small overlap and crop a consistent interior region; hide sticky UI if appropriate; freeze page state.
Text looks soft JPEG compression, device scale, or repeated JPEG encoding. Use quality 80 or higher after visual inspection, use device scale if larger output is needed, and avoid repeated recompression.
Unexpected huge output or memory failure Full-page dimensions or device scale create a large bitmap, and buffers plus decoded pixels consume memory. Use CSS scale, reduce viewport width, capture strips, or process one region at a time. Test with representative pages; no universal safe dimension applies.
Transparent areas become a solid color JPEG has no transparency channel. Choose PNG for transparency, or set an intentional background before JPEG capture.
Navigation times out The page continues network activity or is slow to load. Use an appropriate readiness condition and timeout; for dynamic pages wait for the content selector you need rather than all network activity.
Different tile count between runs Responsive width, content height, or late-loaded modules changed. Fix viewport and page state, wait for layout to settle, and record the resulting image dimensions.

9. Performance, reliability, and cost

Browser rendering cost grows with page complexity and captured pixel area. Full-page capture keeps a large encoded buffer, and cropping requires decoding image data; high device scale can multiply pixel count. Strip capture bounds each screenshot’s height but introduces more browser work and can expose page changes between captures. Choose the method by measuring memory and elapsed time on representative pages; the cited docs do not provide a universal performance benchmark.

For repeatable output, pin the browser and automation-library versions in your deployment, set viewport and scale explicitly, wait for application-specific readiness, and keep a timeout around navigation and capture. Treat externally hosted pages as variable inputs: content, ads, and layout may change. Store dimensions and capture metadata alongside the tiles if later reconstruction matters. JPEG tiles are lossy, and each new encode can add artifacts.

Local automation has no per-screenshot API fee, but it uses compute, memory, browser installation, and maintenance. For repeated or production capture, compare those operating costs with a hosted screenshot service and include retries, failed navigations, and image storage in the estimate.

10. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A single GET returns an image or PDF; request JPEG and choose full-page capture, then crop the returned image into tiles with the library of your choice. The service provides options including resizing, custom CSS and JavaScript, waiting, and caching; see the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.jpg -d format=jpeg -d full_page=true
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={
    "access_key": "YOUR_API_KEY", "url": "https://stripe.com",
    "format": "jpeg", "full_page": "true"
}, timeout=90)
r.raise_for_status()
open("shot.jpg", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY', url: 'https://stripe.com',
  format: 'jpeg', full_page: 'true'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.jpg', Buffer.from(await res.arrayBuffer()));

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Crop the resulting full-page JPEG into tiles using the code above. Sign up for 1,000 free screenshots a month, no card required.

11. Frequently asked questions

Can I create one JPEG tile per screenful?

Yes. Capture successive viewport regions and save each as a separate JPEG. Decide how to handle sticky headers and overlap so joining the tiles later does not duplicate or omit content.

Does full-page capture mean the browser gives me tiles?

No. It produces one full-page image. Split it afterward with an image library, or capture page regions separately.

Can JPEG tiles have transparent backgrounds?

No. JPEG does not support transparency. Use PNG when transparent pixels must be preserved.

Which tile size should I use?

Choose dimensions that fit your downstream viewer, upload limit, and memory budget. There is no universal tile size specified by the browser screenshot documentation.

Should I use Playwright or Puppeteer?

Use the tool your project already supports unless a required browser or configuration detail points elsewhere. Both document full-page screenshot options; the supplied sources establish no comparative speed ranking.