ScreenshotNeo

BlogAI agents

How to Debug Garbage Output from an AutoGen Screenshot Tool

If AutoGen returns garbage from a screenshot tool, first check whether PNG bytes became text. Here’s how to trace the payload and deliver a real image to a vision model.

By the ScreenshotNeo team29 September 202611 min read

How to Debug Garbage Output from an AutoGen Screenshot Tool

If an AutoGen screenshot tool returns garbage, the first thing to check is whether the screenshot’s binary bytes were converted into ordinary text before they reached the model. A tool call can complete successfully while the model receives a string such as b'\\x89PNG...' instead of image pixels. A fluent answer does not prove the model saw the page.

The reliable repair is to fetch the screenshot as bytes, decode it into an image, and put that image in multimodal message content. This guide focuses on Microsoft’s autogen-agentchat, autogen-core, and autogen-ext family. The older autogen package and the ag2 project have different APIs and may need a different implementation.

1. Diagnose the payload before changing your agent

Start at the boundary between your screenshot function and AutoGen. Record the returned Python type, the byte count, and the first eight bytes. A PNG normally starts with the byte signature 89 50 4e 47 0d 0a 1a 0a. If your value is a string starting with b'\\x89PNG, bytes have already been rendered as their Python representation.

from pathlib import Path

result = capture_screenshot("https://example.com")
print("type:", type(result).__name__)

if isinstance(result, bytes):
    print("length:", len(result))
    print("first 8 bytes:", result[:8].hex(" "))
    Path("debug-shot.png").write_bytes(result)
elif isinstance(result, str):
    print("length:", len(result))
    print("prefix:", repr(result[:40]))

Next inspect the actual message content passed to the model client. It should contain an image object, not a textual byte representation, a base64 blob pasted into prose, or just a screenshot URL. Keep a local copy of the original response bytes so you can distinguish a capture failure from a later encoding or message-construction failure.

Observed value What it suggests Next check
bytes beginning with a PNG or JPEG signature The HTTP fetch probably returned image bytes. Decode the bytes and inspect the image before passing it to the model.
String like b'\\x89PNG...' Bytes were stringified. Find the tool-result or logging conversion and keep the image out of that text path.
Base64 text The payload may have been encoded, truncated, or treated as ordinary text. Decode it strictly, then create an image object and place that in multimodal content.
HTML, JSON, or an error page The endpoint response may not be an image at all. Check HTTP status, content type, endpoint parameters, and response body.

This separates the two common failure classes: the capture did not produce an image, or a valid image was flattened into text while moving through the agent framework.

2. Why a normal tool result can corrupt an image

In Microsoft AutoGen’s standard tool path, a function result is represented as text. The documented tool flow produces a FunctionExecutionResult with a string content field; the tool helper’s return_value_as_string converts a value with str(value). Returning raw PNG bytes through that path can therefore yield a Python bytes representation rather than an image part. The model gets tokens describing bytes, not the visual content encoded by those bytes. See the [AutoGen tool documentation](https://microsoft.github.io/autogen/dev/user-guide/core-user-guide/components/tools.html) and [tool API reference](https://microsoft.github.io/autogen/dev/reference/python/autogen_core.tools.html).

Trace the screenshot through each representation boundary: bytes, decoded image, and multimodal message content.
Trace the screenshot through each representation boundary: bytes, decoded image, and multimodal message content.

This is why an agent may confidently describe a plausible page after the tool call. It might be guessing from the URL, prior context, or its language model priors. The useful debugging question is not “Did the tool run?” but “Did the model request contain image content?”

3. Repair: fetch bytes and pass a multimodal message

For a screenshot selected by your application, use a binary-capable HTTP client, decode the response with Pillow, wrap it as an AutoGen image, and send a MultiModalMessage. Install the Microsoft packages and Pillow in the same Python environment that runs the script:

pip install autogen-agentchat autogen-core httpx pillow

The example below uses Site-Shot as the screenshot endpoint named in the research dossier. Set its API key in the environment; do not commit credentials in source. The capture result stays binary until Pillow decodes it.

import io
import os

import httpx
from PIL import Image as PILImage
from autogen_core import Image as AGImage
from autogen_agentchat.messages import MultiModalMessage


def capture(page_url: str) -> AGImage:
    response = httpx.get(
        "https://api.site-shot.com/",
        params={
            "url": page_url,
            "userkey": os.environ["SITESHOT_API_KEY"],
            "full_size": 1,
            "no_ads": 1,
            "no_cookie_popup": 1,
        },
        timeout=60.0,
    )
    response.raise_for_status()

    # Decode response bytes as an image before creating multimodal content.
    with PILImage.open(io.BytesIO(response.content)) as image:
        image.load()  # Force decode while the response buffer is available.
        return AGImage(image.copy())


async def ask_about_screenshot(agent) -> None:
    shot = capture("https://example.com")
    task = MultiModalMessage(
        content=[
            "Does this pricing page show a free tier above the fold?",
            shot,
        ],
        source="user",
    )
    result = await agent.run(task=task)
    print(result.messages[-1].content)

The path is HTTP response bytes → BytesIO → Pillow image → autogen_core.Image → multimodal message. The Microsoft documentation shows the same general shape for multimodal input: text and an image object can coexist in message content. See [AutoGen’s multimodal agent guide](https://microsoft.github.io/autogen/dev/user-guide/agentchat-user-guide/tutorial/agents.html).

This example defines the image transport and message construction, but it intentionally accepts an already configured agent. Your agent’s model client must accept images and support any function or tool calling you use. AutoGen represents those as separate model capabilities: vision and function calling are distinct. The [model capability FAQ](https://microsoft.github.io/autogen/dev/user-guide/core-user-guide/faqs.html) explains those flags.

Validate the response before decoding

Production code should fail clearly if the capture endpoint returns an error page or a non-image response. Add checks before Pillow decoding, and log metadata rather than dumping image bytes into application logs:

def fetch_image_bytes(url: str) -> bytes:
    response = httpx.get(url, timeout=60.0, follow_redirects=True)
    response.raise_for_status()

    content_type = response.headers.get("content-type", "").lower()
    if not content_type.startswith("image/"):
        preview = response.text[:200] if response.text else "<empty body>"
        raise ValueError(
            f"Expected image response, got {content_type!r}; body starts {preview!r}"
        )

    if not response.content:
        raise ValueError("Screenshot endpoint returned an empty body")
    return response.content

If the provider returns a PDF or JSON metadata under a particular option, handle that response type explicitly instead of trying to open it as PNG. Check the endpoint’s response contract and headers for the configuration you selected.

4. Why MCP, HttpTool, and image URLs may still fail

  • MCP image result: an MCP server can return image content, but that alone does not ensure the AutoGen model request preserves it. The standard AssistantAgent tool-result path may turn tool results into text via to_text(), which can render image data as base64 text. Inspect the final model message, not merely the MCP tool response.
  • HttpTool: the documented HTTP tool route is for text or JSON, and its GET path returns response.text. That is not a safe way to transport arbitrary PNG bytes. Use an HTTP client that exposes response.content for binary data.
  • Image.from_uri() with HTTPS: despite the name, the AutoGen image URI helper described in the dossier matches base64 data URIs for PNG or JPEG, not an ordinary hosted URL. Fetch the URL first, then decode the downloaded bytes.
  • Base64 embedded in prose: a long string is not automatically vision input. Decode it and construct the framework’s image type, then pass that object as multimodal content.

These pitfalls all cross a representation boundary. At every handoff, ask whether the value is still binary data, valid encoded image data, a decoded image object, or merely text.

5. Choose who controls screenshot timing

The repair above is best when your application knows which page to capture and when. Your code fetches one screenshot, builds a multimodal user message, and asks a question about that frame. The tradeoff is that the application, rather than a standard AssistantAgent, decides when to capture.

Choose application-controlled capture for a known page state, or a browser agent for repeated navigation.
Choose application-controlled capture for a known page state, or a browser agent for repeated navigation.

For repeated browser interaction where the agent must navigate, inspect, and act across screenshots, Microsoft’s MultimodalWebSurfer is the built-in alternative. It launches Chromium through Playwright, takes and scales screenshots, wraps them with AGImage.from_pil, and includes them in multimodal messages. Its official documentation requires a multimodal model client that supports function/tool calling; see the [MultimodalWebSurfer reference](https://microsoft.github.io/autogen/dev/reference/python/autogen_ext.agents.web_surfer.html). This component manages browser turns rather than simply returning a screenshot from a normal function tool.

For a custom agent that owns its browser loop, follow the same architecture: make the agent a chat agent that can emit multimodal message types and include screenshots as image objects. Don’t try to make a text-only function return channel carry the image implicitly.

6. Troubleshooting common errors

Symptom Likely cause Fix
The answer is plausible but unrelated to page details. Image bytes were stringified or image data was omitted from the model request. Print the message content types and verify an image object is present. Ask a visual question with a verifiable detail.
Invalid base64-encoded string Base64 was malformed, truncated, had characters removed, or was decoded from the wrong text field. Prefer raw response bytes. If base64 is required, decode strictly and do not treat a Python bytes repr as base64.
Pillow raises UnidentifiedImageError. The endpoint returned HTML, JSON, an error body, or a zero-length response. Check status and content type first; log a short text preview only for non-image responses.
Image.from_uri() rejects an HTTPS URL. The helper expects an image data URI form rather than a regular URL. Download with an HTTP client and construct the image from decoded bytes.
Request times out during a full-page capture. Rendering, loading, or the endpoint took longer than the client’s timeout. Set an explicit timeout appropriate to the capture; the dossier’s working example uses 60 seconds. Avoid assuming the documented five-second HttpTool default is enough for full-page rendering.
MultimodalWebSurfer rejects the model client. The client lacks vision or function/tool calling capability, or its capability metadata is wrong. Choose a compatible model and set capability metadata to reflect actual provider support; flags cannot add unsupported model features.
Code works in one AutoGen install but not another. Different package families or versions are installed. Check pip show autogen-agentchat autogen-core autogen-ext ag2 autogen, pin the intended family, and consult that family’s docs.

An invalid base64 warning has appeared in a Microsoft AutoGen GitHub issue involving a vision workflow, but the report itself is not proof of a universal package defect. Treat it as a clue to inspect encoding and payload handling: [issue #2204](https://github.com/microsoft/autogen/issues/2204).

7. Performance, reliability, and cost considerations

Images consume more bandwidth and model input capacity than a short text result. Avoid base64-in-prose because it adds a large text payload without guaranteeing visual interpretation. For a stable application-selected capture, make one capture for the question, reuse it if the page has not changed, and resize or crop only when doing so preserves the details the question depends on. Microsoft’s MultimodalWebSurfer source contains specific scaled-image constants, but those are implementation details, not a model-quality benchmark.

Set an explicit HTTP timeout and use raise_for_status(). For transient network errors, a bounded retry with backoff can help, but do not blindly retry malformed-image or authorization errors. If the page itself changes while you capture and ask, include the URL and capture time in your own trace metadata so you can reproduce which state was inspected.

Log status code, content type, byte length, image dimensions, capture duration, and whether the message contained an image object. Keep API keys and page content out of logs unless your data policy permits them. Browser-driven agents add browser startup and page interaction to the request path; the AutoGen documentation notes that the browser is initialized lazily and reused until closed. Reusing that browser can avoid repeated startup work, while application-managed screenshots can be simpler for one-off captures.

Costs vary by the model provider, image detail, and number of model turns; the dossier provides no controlled cost or latency benchmark. Count capture requests and model calls separately in your own application. A browser-agent loop may need multiple screenshot-and-reason turns, while an application-selected screenshot can use a single captured frame. Neither path removes the need to verify successful image transport.

8. Or skip the browser setup

If you want the screenshot delivered by an API instead of setting up a browser capture stack, ScreenshotNeo is a website screenshot API and MCP server. It returns a screenshot or PDF from one GET request. It accepts cookie or consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing result.

This produces an image response that your Python code can fetch as bytes, decode to a PIL image, and insert into the same multimodal message pattern above. The call below is a runnable endpoint example; set your access key and choose a target URL. See the ScreenshotNeo API documentation for request options.

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Then preserve the image as an object when you pass it to AutoGen:

from io import BytesIO
from PIL import Image as PILImage
from autogen_core import Image as AGImage

with PILImage.open(BytesIO(r.content)) as pil_image:
    screenshot = AGImage(pil_image.copy())

# Use screenshot in MultiModalMessage content, alongside your question.

ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. If you use it through AutoGen, still inspect how your particular agent path transports MCP image results; MCP support by itself does not guarantee the model receives image pixels.

There are 1,000 screenshots a month on the free plan with no card required. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Sign up free and get 1,000 screenshots a month with no card.

9. FAQ

Can a model infer the screenshot from the URL?

It may generate a plausible answer, but a URL is not a screenshot. If the answer depends on visual state, confirm an image object was part of the model request.

Should I return a base64 string from my tool?

Only if the receiving layer explicitly decodes that string into image content. Base64 returned as ordinary tool text is still text to the model.

When should I use MultimodalWebSurfer?

Use it when the agent should control ongoing browser navigation and actions. Use a multimodal message with an application-captured image when your application already knows what page and state to inspect.

Does an MCP screenshot tool solve this automatically?

No. Confirm the final model request contains image content. A tool response can be converted to text between the MCP server and the model.

Final checklist

  1. Confirm the installed AutoGen package family.
  2. Log the result type, byte length, and image signature before any conversion.
  3. Check HTTP status and content type; decode the bytes as an image.
  4. Put an AutoGen image object in multimodal message content.
  5. Verify that the selected model supports vision and any required tool calling.
  6. Inspect the final request structure when answers remain ungrounded.