ScreenshotNeo

BlogHow-to

Generate Screenshots of Indian News Article Pages with Java Playwright

Capture an Indian news article’s viewport, full scrollable page, or a selected element with Java Playwright, and handle common capture issues.

By the ScreenshotNeo team4 October 20266 min read

Use Java Playwright to open the article URL, navigate to the page, and call page.screenshot(...). The default capture is the visible viewport. Set fullPage to true for the full scrollable page, or use a locator screenshot to capture one element. The examples below use https://example.com/article as a placeholder; choose an article URL you are permitted to access and inspect its markup before relying on a selector.

1. Set up Java Playwright

Add the Playwright Java dependency using the installation instructions for your project, then install the browser binaries required by the chosen engine. The official guide shows the Java screenshot API and the Page API demonstrates navigation with WebKit. See Playwright Java screenshots and the Java Page API. Browser output can differ; the available documentation does not establish which engine best renders any particular Indian news publisher.

2. Capture an article page

This runnable class saves a viewport screenshot as article.png. Replace the placeholder URL with the article you want to capture.

import com.microsoft.playwright.*;
import java.nio.file.Paths;

public class ArticleScreenshot {
  public static void main(String[] args) {
    String articleUrl = "https://example.com/article";

    try (Playwright playwright = Playwright.create()) {
      Browser browser = playwright.webkit().launch();
      try {
        BrowserContext context = browser.newContext();
        Page page = context.newPage();
        page.navigate(articleUrl);
        page.screenshot(new Page.ScreenshotOptions()
            .setPath(Paths.get("article.png")));
      } finally {
        browser.close();
      }
    }
  }
}

The sequence is straightforward: create Playwright, launch a browser, create a context and page, navigate, then save a screenshot. The example uses WebKit as an available browser choice, not as a recommendation for all publisher pages. The official Java documentation also demonstrates screenshot bytes when you need to process or send the image in memory rather than write it directly to disk.

3. Choose what to capture

Need Java API What it captures
Visible viewport page.screenshot(options) The currently visible page area. Full-page capture is false by default.
Entire scrollable page page.screenshot(options.setFullPage(true)) The page as if displayed on one very tall screen; it is not a sequence of viewport screenshots.
One element locator.screenshot(options) The matched element’s bounds. The selector must match the actual page markup.
Image bytes page.screenshot() Returns a byte[] for downstream processing or storage.

These behaviors are documented by Playwright’s Java screenshot guide and its screenshot options reference.

Full scrollable page

For an article archive or review where the whole page should appear in one image, enable full-page capture:

page.screenshot(new Page.ScreenshotOptions()
    .setPath(Paths.get("article-full.png"))
    .setFullPage(true));

A very long page creates a correspondingly tall image. If the recipient needs manageable dimensions or page-by-page review, consider a different output workflow rather than assuming full-page mode produces separate viewport files.

One article element

Use a locator when you only need the article body or another specific region. Inspect the selected page and choose a selector that exists there; .header below is illustrative only, not a known selector for any publisher.

Locator article = page.locator("article").first();
article.screenshot(new Locator.ScreenshotOptions()
    .setPath(Paths.get("article-element.png")));

If the selector does not match, inspect the DOM and update it for that page. Publisher markup is not established by Playwright documentation, and article structures can differ between sites and pages.

Capture to memory

When another part of your Java program will store, transform, or transmit the image, use the byte array returned by the screenshot call:

byte[] image = page.screenshot();
// Pass image to your image-processing or storage code.

The guide documents both saving to a path and obtaining screenshot bytes. Pick one based on how the rest of your application consumes the result.

4. Practical considerations for news pages

  • Choose a scope deliberately: viewport, full page, and element screenshots have different dimensions and uses.
  • Use a real target URL: the examples use a placeholder and make no claim about access or rendering on a live publisher.
  • Inspect selectors: do not assume a shared article selector across publishers.
  • Choose an engine for your use case: the API example uses WebKit; no source here compares output for Indian news pages.
  • Save with an explicit path and extension: the documented example uses a PNG path. Keep naming and storage conventions consistent in your own application.

Playwright’s documented screenshot API does not establish a page’s access rules, whether a consent dialog appears, how a particular site loads content, or whether a publisher permits automated access. Check the relevant site’s terms and access requirements, and verify the actual page before building a repeatable capture job.

5. Troubleshooting

Symptom Likely cause What to check
No screenshot file appears The path is not where you expect, or the capture failed before writing. Use an explicit output path, check that its parent directory exists, and confirm the program reaches the screenshot call.
Only the visible portion appears The default is viewport capture. Set .setFullPage(true) for the full scrollable page.
Element capture fails or is empty The locator may not match the page. Inspect the actual DOM and replace the illustrative selector with one that matches the target element.
The page image differs from another browser Rendering can vary by engine and page. Use a browser engine appropriate to the application and verify the output on the chosen target; the cited sources do not identify a best engine for specific publishers.
The captured page does not show the expected content The target page may not have rendered as expected or access may be restricted. Check the destination URL and page state in the browser workflow. Playwright’s screenshot API documentation does not establish publisher-specific access behavior.

6. Performance, reliability, and cost

Screenshot size and capture work depend on the selected scope: full-page output can be much taller than a viewport image, while an element screenshot limits the capture area. The cited API documentation describes the available capture forms but provides no performance benchmark, reliability guarantee, or cost figures for running Playwright. For recurring work, choose the smallest scope that meets the requirement, keep output paths predictable, and verify saved output as part of your application’s own workflow.

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A GET request with a URL returns an image or PDF; its documented API options include full-page capture, CSS selector capture, and format selection. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/article \
  -o article.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/article"},
    timeout=90,
)
r.raise_for_status()
open("article.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/article'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = new Uint8Array(await res.arrayBuffer());
await Bun.write('article.webp', image);

Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000 screenshots. Sign up free for 1,000 screenshots a month, with no card required.

8. FAQ

Does full-page mode stitch together viewport screenshots?

The API describes it as capturing the full scrollable page as if shown on a very tall screen, rather than a sequence of viewport screenshots.

Does the example selector work on every Indian news site?

No selector is guaranteed by the cited documentation. Inspect each target page and choose a locator that matches its markup.

Can I use the screenshot without creating a file?

Yes. The Java screenshot method can return image bytes as a byte[] for your application to handle.

Which browser engine should I use?

The official Java Page API example uses WebKit, but the available sources do not compare engines on Indian news publishers. Select and verify an engine for your specific page and workflow.