ScreenshotNeo

BlogHow-to

How to Crawl Websites with Python

Use Python for one-off fetches, then build a responsible Scrapy crawler that follows links, extracts structured data, and exports results.

By the ScreenshotNeo team4 October 20268 min read

For a single page, Python’s standard-library urllib.request can fetch and read the response. To crawl multiple pages, follow links, extract structured data, and export results, use Scrapy: a spider parses each response, yields items, and schedules requests for relevant links. Start with a limited scope, identify your crawler, and check the target site’s robots.txt, terms, and applicable requirements before sending requests.

1. Fetch one page with Python

This small script retrieves one URL and prints its response body. It is useful for a one-off fetch; it does not discover links, control a crawl queue, or provide structured exports.

from urllib.request import urlopen

url = "https://example.com/"
with urlopen(url, timeout=20) as response:
    body = response.read()
    charset = response.headers.get_content_charset() or "utf-8"

print(body.decode(charset, errors="replace"))

Replace the example URL with a page you are permitted to retrieve. For a multi-page project, a crawler framework avoids having to build request scheduling, retry behavior, link traversal, and output handling from scratch.

2. Create a Scrapy project

Install Scrapy in a virtual environment, create a project, and generate a spider. These commands follow Scrapy’s documented tutorial workflow. Consult the current installation documentation for platform-specific requirements.

python -m venv .venv

# macOS or Linux
. .venv/bin/activate

# Windows PowerShell
# .venv\Scripts\Activate.ps1

python -m pip install scrapy
scrapy startproject site_crawl
cd site_crawl
scrapy genspider pages example.com

Set a descriptive project-specific user agent in site_crawl/settings.py. Use contact details that you control; do not copy a placeholder contact address into a real crawl.

BOT_NAME = "site_crawl"
USER_AGENT = "site_crawl (contact: your-team@example.org)"
ROBOTSTXT_OBEY = True
CONCURRENT_REQUESTS = 4
DOWNLOAD_DELAY = 1.0

Replace the sample email with your actual contact address. Scrapy’s tutorial recommends an identifiable user agent so site owners can contact the crawler operator. Scrapy supports robots.txt handling; set behavior deliberately and verify the target’s instructions before crawling. Robots rules do not replace checking site terms or applicable law.

A Scrapy spider defines where to start and how to handle each response. The example below collects each page’s URL, title, and description, then follows links only within the example.com domain. Save it as site_crawl/spiders/pages.py.

import scrapy


class PagesSpider(scrapy.Spider):
    name = "pages"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/"]

    custom_settings = {
        # Keep the crawl modest; adjust based on the target's guidance.
        "CONCURRENT_REQUESTS": 4,
        "DOWNLOAD_DELAY": 1.0,
        "ROBOTSTXT_OBEY": True,
    }

    def parse(self, response):
        title = response.css("title::text").get()
        description = response.css('meta[name="description"]::attr(content)').get()

        yield {
            "url": response.url,
            "title": title.strip() if title else None,
            "description": description.strip() if description else None,
        }

        for href in response.css("a::attr(href)").getall():
            next_url = response.urljoin(href)
            if next_url.startswith(("http://", "https://")):
                yield scrapy.Request(next_url, callback=self.parse)

allowed_domains keeps off-domain links from being scheduled by Scrapy. The URL check retains HTTP and HTTPS links; Scrapy’s domain filtering enforces the declared domain boundary. For a narrow crawl, also filter paths, query parameters, file extensions, and pagination links before yielding requests. This helps avoid calendars, search pages, tracking URLs, and effectively unbounded URL patterns.

CSS selectors depend on each page’s markup. If the target uses different elements, inspect representative HTML and change the selectors. Missing title or description values are returned as null in the exported JSON rather than causing the spider to fail.

4. Run the crawl and export results

From the project directory, run the spider and write extracted items as JSON Lines:

scrapy crawl pages -O pages.jsonl

Each item is written as one JSON object per line. Use -O to overwrite the output file; use -o to append to an existing feed where that format supports appending. Scrapy also supports other feed formats and destinations. For larger workflows, item pipelines can validate, clean, and store records.

Approach Use it when Trade-off
Plain Scrapy Spider You need custom traversal, parsing, or filtering. Most flexible; you write the crawl logic.
CrawlSpider The site is regular and links can be described with rules. Convenient rule-based following, but it does not suit every site and custom callbacks need care.
SitemapSpider The target provides useful sitemap URLs. Discovers URLs through sitemap structure instead of relying only on page links.

Choose based on the site’s structure, whether a usable sitemap exists, the fields you need, and the request pace the target permits. Scrapy describes its spiders as producing requests and processing responses; callbacks can yield both data items and more requests.

6. Set crawl scope and request behavior

  • Start narrowly: use a small set of start URLs and a domain allowlist. Add path or page-type rules where needed.
  • Respect site instructions: inspect the top-level /robots.txt for each host and review site terms and applicable requirements. RFC 9309 specifies the Robots Exclusion Protocol and the top-level robots file location.
  • Identify the crawler: use a descriptive user agent and a contact method you control.
  • Limit load: tune concurrency and download delay for the target. Maximum speed is not the goal; use a pace appropriate to the site and stop if it signals a problem.
  • Watch crawl growth: query strings, faceted navigation, calendars, and repeated pagination can create very large or unbounded URL spaces. Filter them intentionally.

A robots file is a set of crawler instructions, not a grant of permission. The technical sources cited here do not decide whether a particular crawl is lawful or allowed by a particular site.

7. Handle common extraction and crawl edge cases

Use response.urljoin(href) to resolve links against the current page. A raw relative path is not a complete request URL.

Duplicate URLs and query parameters

Equivalent pages may have different tracking parameters or URL forms. Normalize or reject irrelevant parameters and fragments, and define which URL variants belong in your dataset. Avoid discarding parameters that change page content.

Redirects and errors

Redirects may lead outside your intended scope. Check the final response URL and keep domain and path boundaries explicit. Review Scrapy’s crawl statistics and logs for repeated HTTP failures rather than silently treating failed pages as extracted records.

JavaScript-rendered pages

A normal HTTP response may not contain content assembled by client-side JavaScript. First determine whether the data is present in the HTML response or available through a documented public endpoint. Browser rendering is a separate requirement; do not assume a plain Scrapy download executes page scripts.

Changing markup and missing fields

Selectors can stop matching after a site redesign. Treat extracted fields as optional, monitor missing-value rates, and validate records in a pipeline or after export.

Large crawls

For a large crawl, write outputs incrementally, keep a clear crawl boundary, and use pipelines or feed destinations suited to the volume. Revisit scheduling, storage, and retention before expanding the URL set.

Or skip the browser setup

If you need screenshots of pages as part of a review or visual archive, ScreenshotNeo is a website screenshot API and MCP server. A GET request captures a URL as PNG, JPEG, WebP, or PDF. It is not a replacement for crawling and extracting a whole site; it handles page captures.

One call with cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

See the ScreenshotNeo API documentation for request parameters and response details. Cookie and consent banners are accepted or removed before the capture, along with known newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Performance, reliability, and cost

For a crawl, total work depends on the number of URLs you schedule, response sizes, server behavior, and your configured concurrency and delays. There is no universal safe or appropriate crawl rate. Increase scope gradually, observe responses and logs, and stop or slow down when the target indicates trouble. Retries can improve resilience to transient failures, but repeated requests to a failing host can make the problem worse.

Scrapy provides scheduling, concurrent requests, item pipelines, and feed exports, so you do not need to implement those basics around a one-page fetching function. Storage, bandwidth, compute, and any third-party services still have costs that depend on your setup; this research does not establish a universal cost or speed benchmark. A one-off urllib request has little framework overhead, while the engineering and operating cost of a crawler grows with scope and data handling needs.

Troubleshooting

Symptom Likely cause Fix
scrapy: command not found Scrapy is not installed in the active environment, or the virtual environment is not activated. Activate .venv, then run python -m pip install scrapy.
The spider finds no pages Start URL, domain, selector, or robots behavior is wrong for the target. Check the spider name and start URL, inspect logs and response HTML, confirm allowed_domains, and review robots instructions.
Requests leave the intended site Links point to another host or a redirect changes the destination. Keep allowed_domains narrow and verify final response URLs before accepting data.
Fields are empty Selectors do not match the returned markup, or content is rendered after the initial response. Inspect the response source, adjust selectors, and determine whether a rendering-capable approach is required.
Output is unexpectedly appended or overwritten The feed option differs: -o appends where supported; -O overwrites. Choose the intended option and use a new output path for a separate crawl run.
Crawl grows without bound Calendars, query variants, faceted filters, or pagination create endless unique URLs. Constrain paths and query parameters, define pagination rules, and monitor scheduled URL counts.

FAQ

Is web crawling the same as web scraping?

Crawling discovers and requests pages; scraping extracts information from responses. Many projects do both, as the Scrapy spider above does.

Can Scrapy crawl pages that require a login?

Only if you have authorization and implement the target’s permitted authentication flow. The example is for public pages and does not include credentials.

Does robots.txt settle whether I may crawl a site?

No. It communicates crawler rules. Review site terms and applicable law separately.

When should I use urllib instead of Scrapy?

Use urllib for a simple single fetch. Use Scrapy when you need link traversal, structured item handling, and crawl controls.

Sources