ScreenshotNeo

BlogAI agents

How to Give a LlamaIndex Agent Website Screenshots

Capture a webpage with Playwright, pass it to a LlamaIndex agent as an ImageBlock, and handle multimodal tools, failures, and production trade-offs.

By the ScreenshotNeo team29 September 20269 min read

How to Give a LlamaIndex Agent Website Screenshots

Direct answer: capture the page with Playwright, save the PNG (or keep the returned bytes), and send it to your LlamaIndex agent in a ChatMessage containing an ImageBlock. The agent must use a multimodal model and provider path. LlamaIndex’s current agent documentation demonstrates this pattern with FunctionAgent and ImageBlock(path="./screenshot.png"); Playwright’s Page API provides navigation and screenshot capture. See the LlamaIndex agents documentation and Playwright Page API.

What the integration looks like

There are two useful designs:

Design How it works Best fit Main trade-off
Application-side capture Your code opens the URL, captures the image, then sends a user message with an ImageBlock. Fixed or externally triggered captures. Browser state and timing remain outside the agent.
Custom screenshot tool The agent calls a tool that uses Playwright and returns image content for its next reasoning step. Workflows where the agent decides when visual inspection is needed. You must verify image-bearing tool results with your LlamaIndex version, model provider, and agent class.

The reviewed LlamaIndex Playwright tool reference documents navigation, text and link extraction, inspection, clicking, and filling. It does not document a screenshot operation. Plan to call page.screenshot in your own function or capture in application code.

Prerequisites

  • Python 3.9 or newer.
  • A LlamaIndex installation and a multimodal model integration supported by your deployment.
  • Playwright and its browser binaries.
  • An API key for the model provider you choose.
python -m pip install llama-index playwright
playwright install chromium

Package names and model adapters vary. Keep the model configuration used by your application, and confirm that the selected model accepts image input. LlamaIndex’s agent guide notes that some LLMs support multiple modalities; an ordinary text-only model cannot inspect pixels just because an image file is attached.

The integration pipeline: navigate, capture pixels, and send an ImageBlock to the agent.
The integration pipeline: navigate, capture pixels, and send an ImageBlock to the agent.

Complete Python example: capture, then send an ImageBlock

This example separates browser capture from agent execution. Replace the model initialization with the provider integration used by your project.

import asyncio
from pathlib import Path

from playwright.async_api import async_playwright
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.llms import ChatMessage, ImageBlock, TextBlock

TARGET_URL = "https://example.com"
SCREENSHOT_PATH = Path("screenshot.png")


async def capture_page(url: str, output_path: Path) -> None:
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        try:
            page = await browser.new_page(
                viewport={"width": 1440, "height": 900},
                device_scale_factor=1,
            )
            await page.goto(url, wait_until="networkidle", timeout=60_000)
            await page.screenshot(path=str(output_path), full_page=True, type="png")
        finally:
            await browser.close()


async def main() -> None:
    await capture_page(TARGET_URL, SCREENSHOT_PATH)

    # Configure this agent with your multimodal LlamaIndex LLM.
    agent = FunctionAgent(
        tools=[],
        llm=YOUR_MULTIMODAL_LLAMA_INDEX_LLM,
        system_prompt=(
            "You inspect website screenshots. Describe only what is visible "
            "and call out uncertainty when text is unreadable."
        ),
    )

    message = ChatMessage(
        role="user",
        blocks=[
            TextBlock(
                text=(
                    "Describe the visible layout, identify the primary call to "
                    "action, and report any sign-in form."
                )
            ),
            ImageBlock(path=str(SCREENSHOT_PATH)),
        ],
    )
    response = await agent.run(message)
    print(response)


if __name__ == "__main__":
    asyncio.run(main())

The important detail is the message structure: the instruction is a TextBlock, and the actual pixels are an ImageBlock. Do not replace the image with a textual description if the task depends on layout, typography, visual state, or content that is not exposed in the DOM.

Use screenshot bytes instead of a file

Playwright can return image bytes. This is useful in a service that uploads directly to object storage or passes encoded content through an internal message layer. The exact image block constructor can differ between LlamaIndex releases, so check the versioned API reference before replacing the documented path form.

async with async_playwright() as p:
    browser = await p.chromium.launch()
    page = await browser.new_page()
    await page.goto("https://example.com", wait_until="domcontentloaded")
    png_bytes = await page.screenshot(type="png", full_page=True)
    await browser.close()

# Persisting the bytes keeps the documented ImageBlock(path=...) shape simple.
with open("screenshot.png", "wb") as image_file:
    image_file.write(png_bytes)

Capture settings that affect what the agent sees

Viewport, device scale, and full-page mode

viewport controls responsive breakpoints. Capture at the width your question refers to: a 390-pixel mobile layout and a 1440-pixel desktop layout can contain different navigation and content. device_scale_factor changes pixel density; a higher value produces a larger image and may increase model processing cost. full_page=True stitches the page’s scrollable height into one image. Use a viewport screenshot when you need the initial fold or when very long pages would exceed model image limits.

page = await browser.new_page(
    viewport={"width": 390, "height": 844},
    device_scale_factor=2,
)
await page.screenshot(path="mobile.png", full_page=False)

Wait for the state you want

wait_until="networkidle" can be unsuitable for pages with analytics, chat, or long polling. Use domcontentloaded plus an explicit selector, or a bounded delay, when the page never becomes idle.

await page.goto(url, wait_until="domcontentloaded", timeout=60_000)
await page.locator("main").wait_for(state="visible", timeout=20_000)
await page.wait_for_timeout(1_000)
await page.screenshot(path="ready.png", full_page=True)

For lazy-loaded images, scroll before capture so the browser loads content below the fold:

await page.goto(url, wait_until="domcontentloaded")
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
await page.wait_for_timeout(1_000)
await page.evaluate("window.scrollTo(0, 0)")
await page.screenshot(path="lazy-loaded.png", full_page=True)

Stable browser state

Authentication, cookies, locale, timezone, geolocation, and user-agent all change the rendered page. Create a browser context with the state required by the task. For repeatable analysis, fix the viewport and locale, and avoid relying on a user’s existing profile.

context = await browser.new_context(
    viewport={"width": 1440, "height": 900},
    locale="en-US",
    timezone_id="UTC",
    color_scheme="light",
)
page = await context.new_page()

Giving the agent a screenshot through a custom tool

A custom tool is appropriate when the agent should choose which URL or state to inspect. The tool should navigate, capture, and return image content in a form your selected LlamaIndex agent and provider actually preserve as an image for the next model call. The official documentation clearly demonstrates image input in a user message, but it does not establish universal support for image-bearing tool outputs across every agent class and provider.

Start with application-side capture. Then add the tool only after a small integration test confirms that the model receives pixels rather than a file path or an empty tool result. Pin LlamaIndex and provider versions in production and recheck this behavior during upgrades.

Common failure modes and fixes

Symptom Likely cause Fix
Executable doesn't exist Playwright package is installed but browser binaries are missing. Run playwright install chromium in the same environment used by the service.
Navigation timeout The site is slow, blocked, or never reaches the selected load state. Set a bounded timeout, try domcontentloaded, then wait for a required selector.
Blank or partial screenshot Capture occurred before client rendering or lazy loading completed. Wait for a visible application selector, scroll to trigger lazy content, and use a short bounded delay.
Agent describes text but misses layout The image was not included as an ImageBlock, or the model is text-only. Inspect the outgoing ChatMessage; verify the model and provider support image input.
Image works in a message but not from a tool Tool-result image handling differs by agent class, provider, or package version. Capture in application code, or implement provider-specific image content and test the exact deployment.
Cookie banner covers the page The site requires consent before showing content. Click the consent control with Playwright, inject a known consent state, or use a capture service that handles consent before capture.
Screenshot is too large for the model Full-page height or device scale creates excessive pixels. Capture the relevant element or viewport, reduce scale, or split a long page into sections.
Different results across runs Ads, animations, time, locale, or responsive dimensions vary. Fix context settings, disable or wait for animations, and record the URL, viewport, and capture timestamp.

Reliability and performance checklist

  1. Reuse a browser process when handling many jobs, but create a fresh context per tenant or authentication state.
  2. Set navigation and selector timeouts; never allow an unbounded page load to occupy a worker.
  3. Close pages, contexts, and browsers in finally blocks.
  4. Record the target URL, viewport, wait strategy, screenshot dimensions, and model response ID for debugging.
  5. Retry transient navigation failures with a small limit and backoff. Do not blindly retry authentication failures or bot challenges.
  6. Prefer PNG for text and UI inspection. Use JPEG only when a smaller file is more valuable than crisp edges.
  7. Crop to the relevant element when the agent needs one form, chart, or navigation area; fewer pixels generally mean faster image transfer and easier visual grounding.
  8. Remove secrets from logs. A screenshot can contain account data, tokens rendered in a page, or personal information.

Cost and operational choices

Self-hosted Playwright gives you control over browser versions, timing, cookies, and network policy. You pay for browser CPU, memory, storage, and model image input. Full-page captures and high device scale increase both transfer size and model work. A remote capture API trades browser maintenance for a per-capture price and a simpler request path. Cache screenshots when the page and state are unchanged, but include viewport, locale, authentication state, and any relevant query parameters in the cache key.

Or skip the browser setup

ScreenshotNeo is a website screenshot API with a single GET request for a PNG, JPEG, WebP, or PDF. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the result with X-Page-Verdict and X-Billed headers. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Use the API result as the image file you pass to ImageBlock:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

See the ScreenshotNeo API documentation for the complete option set. Relevant controls include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, custom CSS and JavaScript, clicks, selector or network-idle waits, request and resource blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage data, and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, which can simplify migration.

ScreenshotNeo has 1,000 free shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account and use the returned image as the input to your LlamaIndex workflow.

FAQ

Does the LlamaIndex Playwright tool take screenshots?

The reviewed tool reference lists browser interaction and extraction operations but does not document screenshot capture. Use Playwright’s Page.screenshot in application code or expose your own screenshot function.

A clean capture removes obstructing overlays before visual analysis.
A clean capture removes obstructing overlays before visual analysis.

Can I send a screenshot as plain text?

No. A textual description loses visual information. Attach the actual image with ImageBlock and use a multimodal model.

Should I use a full-page screenshot?

Use it when the agent must inspect content below the fold. Use a viewport or element capture when the task concerns one visible state or when image size is constrained.

Can every LlamaIndex agent consume an image returned by a tool?

Do not assume that. Verify the exact LlamaIndex agent class, provider adapter, and package versions. Application-side capture followed by a user message is the clearest documented path.

How do I make captures reproducible?

Fix viewport, device scale, locale, timezone, color scheme, authentication state, wait conditions, and animation behavior. Store those settings with the capture metadata.