How to Capture Competitor Website Pages for a Market Research Report
Capture public competitor pages with clear timestamps, documented methods, and honest notes about what screenshots and archives can miss.
For a market research report, capture each relevant public competitor page with a screenshot when visual layout or an offer is the evidence you need. Record the exact URL, page title, capture date and time in UTC, and method alongside it. Use an archive snapshot or a local web harvest when links, structure, or replay matter. Inspect every capture and describe what is missing or uncertain: a screenshot is a static view, while an archive or harvest can still omit resources and dynamic content.
A capture records an observation at a particular time. It does not prove that the page was complete, continuously available, or displayed identically to every visitor. Keep that distinction clear in your report.
1. Define the question and scope
Start with the report question. Select only pages that can answer it, such as a competitor’s pricing, product, feature, or landing page. Avoid collecting whole sites by default; focused scope makes comparisons easier to explain and review.
For each page, record these details before or during capture:
- Competitor or organization name
- Exact page URL, including relevant path and query string
- Page title as observed
- Capture date and time in UTC
- Capture method and tool
- Research question or comparison dimension the page supports
Use a common time window and comparable page types where possible. If one competitor’s pricing page is compared with another’s product home page, note the difference rather than implying they are equivalent evidence.
2. Choose a capture method
| Method | Useful for | What it preserves | Limitations to record |
|---|---|---|---|
| Browser screenshot | Visible layout, hierarchy, imagery, and offer | A static visual record of the rendered view | Does not preserve linked structure or interactions; content below the fold or behind controls may require additional captures |
| Public web archive | Dated snapshots and shareable replay | A captured page and, depending on the archive and capture, some linked resources for replay | Capture timing and completeness are outside your control; pages and resources may be absent, blocked, or incomplete |
| Local web harvest | Researcher-controlled capture of pages and linked resources | Potentially more structure and navigation for later inspection | Requires more setup and review; scripts, interactive content, and externally hosted resources may not work |
The Library of Congress explains that a crawler starts from a seed URL and follows links, and notes that images, CSS, and JavaScript can matter to reproducing a page’s look and function. Its guidance says, “A crawler can only capture websites that it knows about.” A sitemap can expose pages that a crawler might not discover by following links alone. These points help assess archive coverage; they do not mean an independent researcher should ignore a site’s access controls or preferences. See the Library of Congress guidance on web archiving for site owners and the Library’s web archiving program.
For permanent agency web records, the National Archives and Records Administration advises capture methods that retain hypertext functionality, such as harvesting. That is agency records guidance, not a universal requirement for market research. Consider structure and replay needs for your own report rather than treating one method as suitable for every purpose. See NARA Guidance on Managing Web Records.
3. Capture a visual record with a browser
- Open the exact public URL in a browser and note the time in UTC.
- Wait for the content relevant to your question to render. If a page loads content as you scroll, inspect the full page before deciding what needs capture.
- Capture the initial viewport when it establishes the page’s first impression. Capture additional views for content below the fold or important sections.
- Save the screenshot without editing the original. Use a filename that identifies the competitor, page, and UTC date, and keep a separate record with the full URL and method.
- Inspect the saved image. Note any consent dialog, popup, missing image, clipped content, or other condition that affects interpretation.
A screenshot can establish what was visible in that image, but it does not preserve links, page behavior, or content that was not captured. If you need a record that can be navigated or replayed, supplement it with an archive snapshot or web harvest.
4. Record archive snapshots and web harvests
For a public archive, save the archive URL or capture identifier in your evidence record, along with the original URL, timestamp shown by the archive, and the archive or method used. Review the replay instead of assuming every element was captured. An archive date is evidence of a dated capture; it does not by itself prove that every page element was available then or that the content first appeared on that date.
For a local harvest, document the tool and its settings, the starting URL, scope or link-following rules, and where the resulting files are stored. Check whether stylesheets, images, scripts, and embedded resources were retained. WARC and ARC are formats used for web archiving; when evaluating a tool, consider whether it supports useful indexing, original URLs, chronology, and link relationships. Do not assume every harvest has these properties.
Archive evidence can be useful for showing when content appeared online, but access restrictions, blocking, partial captures, removal requests, and sporadic crawl schedules can limit what is available. The EUIPO’s Common Communication CP12 on evidence in trade mark appeal proceedings discusses both the value and limits of web archive evidence. Treat it as context for careful evidence handling, not as a guarantee that an archived page is complete.
5. Keep evidence understandable and reviewable
Store the capture with a record that preserves context. A useful report entry can include:
Competitor: Example Co.
Page title: Example Product Pricing
Original URL: https://example.com/pricing
Observed at (UTC): 2026-10-04T14:30:00Z
Method: Browser screenshot, full-page capture
Capture file: example-co-pricing-2026-10-04T143000Z.png
Archive URL or capture ID: Not used
Scope: Public pricing page visible without signing in
Limitations: Consent dialog dismissed; interactive calculator not captured
Observation: Three plan columns were visible.
Interpretation: The page presents plan comparison as the primary purchase path.
Keep observation separate from interpretation. Preserve an unmodified copy of the capture where possible, and make any cropping or annotation on a clearly labeled working copy. In the report, state the capture date and method near the evidence. NARA’s recordkeeping guidance describes reliability, authenticity, integrity, and usability as useful qualities for trustworthy web records; context and site structure can support later interpretation.
Do not include private or access-controlled pages in a workflow intended for public competitor research unless you have authorization. The sources cited here do not settle copyright, privacy, terms-of-use, or republication questions across jurisdictions. For a report distributed outside your organization, check the rules and permissions that apply to the material and location.
6. Compare competitor pages fairly
Choose dimensions tied to the research question, then capture comparable pages in a shared time window. Examples include stated plan names and prices, visible feature claims, calls to action, navigation structure, or the order in which benefits appear. Record what you observed before drawing conclusions.
- Use the same page type where possible, such as pricing against pricing.
- Record the capture time for every competitor; pages can change between observations.
- Note differences in locale, currency, personalization, consent state, or viewport if they affect the visible content.
- Do not infer that an uncaptured element was absent from the live site.
- Label interpretation as analysis, not as a directly visible fact.
Or skip the browser setup
For repeatable visual captures, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns an image or PDF. This call saves a screenshot of a public page; see the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
ScreenshotNeo accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
For research records, retain the original URL and capture timestamp with the returned file, and describe any limitations just as you would for a browser capture. Sign up for 1,000 free screenshots a month, with no card required.
Options for repeatable API captures
ScreenshotNeo supports PNG, JPEG, WebP, and PDF output. For a market research workflow, relevant options include full-page capture with lazy images loaded, a CSS selector for capturing one element, custom viewport or one of 12 device presets, retina scale, dark mode, and waiting for a selector, a delay, or network idle. PDF captures support paper size, margins, landscape orientation, and page ranges.
It also supports custom CSS and JavaScript, clicking an element before capture, hiding selectors, blocking ads, trackers, requests, or resource types, and setting headers, cookies, user agent, authorization, timezone, and geolocation. Transparent backgrounds, image resizing, configurable cache TTL, signed links for public image tags, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI spec are available. Parameter names used by other screenshot APIs also work, which can ease switching. All features are on every plan.
Use options that make captures comparable and document them. A changed viewport, cookie state, user agent, or custom script can change what is visible. If a custom wait condition is used, include it in the method record. A cached result may represent an earlier capture; record the cache behavior and TTL relevant to your evidence. For asynchronous or bulk captures, associate each result with its requested URL and timestamp before comparing pages.
Performance, reliability, and cost
Page complexity, network requests, lazy content, and wait conditions affect capture time. Waiting for network idle can be slow on sites with continuous background requests; a targeted selector or bounded delay may be more appropriate when it matches the research question. Capture only the pages and page regions you need, and use bulk requests when collecting a defined set of URLs.
No capture method guarantees a complete rendering. Websites may block automated access, show bot checks, fail to load, or vary by location, account, cookies, or time. Preserve failure outcomes in your notes rather than treating a failed capture as evidence about the page’s content. For a serious comparison, recheck unexpected differences and keep original files and records together.
ScreenshotNeo’s listed pricing is Free: 1,000 shots per month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free. Clean shots are billed; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not. Each response includes X-Page-Verdict and X-Billed headers to show the outcome. Consult the documentation for request details and current plan information.
Troubleshooting capture problems
| Problem | Likely cause | What to do |
|---|---|---|
| Screenshot is blank or incomplete | The page did not finish loading, content is lazy-loaded, or a bot check blocked access | Open the URL normally to inspect it, wait for a specific element or content state, scroll where needed, and note any challenge or failure. Do not present a blank capture as a complete page record. |
| Images or styling are missing in an archive replay | Resources were not captured, were hosted elsewhere, or were unavailable to the crawler | Record the missing resources, check whether another dated snapshot exists, and capture a screenshot of the visible live page if appropriate. |
| Interactive area does not work in replay | Scripts, embeds, or dynamic state did not carry over | Capture the visible state separately and document the interaction you could not reproduce. |
| Competitor pages appear different | Captures used different times, viewports, locales, cookie states, or personalization | Align those conditions where possible, recapture, and explicitly document remaining differences. |
| ScreenshotNeo request returns an error | Invalid or missing key, malformed URL, timeout, blocked page, or request option mismatch | Check the API key and URL encoding, inspect the HTTP response and X-Page-Verdict/X-Billed headers, then adjust the wait or capture settings. Keep failure details in the research record. |
| Capture is stale | A cached result was returned | Check the response outcome and configured cache TTL; use a suitable TTL or cache setting for time-sensitive captures and record it. |
FAQ
Does a screenshot prove what every visitor saw?
No. It records one rendered view under particular conditions. Note the time, viewport, and any relevant locale, cookies, or personalization.
Should I use an archive or a screenshot?
Use a screenshot for visible appearance and an archive or harvest when replay, navigation, or structure matters. You can preserve both when the report needs both kinds of evidence.
Can I treat an archive timestamp as the date content first appeared?
No. It dates the archived capture, not necessarily the content’s first publication or the availability of every element.
Can I republish competitor screenshots in a report?
That depends on applicable copyright, privacy, contractual, and jurisdictional rules. The cited web archiving guidance does not resolve those questions for every report.


