How to Extract Data From a Website: A Practical Guide
Choose the right source, extract fields from HTML with Python, and handle pagination, dynamic pages, validation, and crawler rules.

To extract data from a website, first look for an official API, feed, downloadable dataset, or other supported source. If the fields you need are already in the page’s HTML, request the page and select the relevant elements with CSS or XPath. For many pages, use a crawler such as Scrapy to follow links and produce structured records. If the data appears only after JavaScript runs, inspect the browser’s network requests and try the underlying data request; use a headless browser when that is impractical or when you need the rendered page itself.
The key decision is where the data lives. A page that looks complete in a browser may return only a shell to a basic HTTP client. Inspect a representative response before committing to a scraping method.
1. Define the data and scope
Write down the fields you need, the pages that contain them, the number of pages, and whether the collection will run once or recur. This short inventory keeps the extraction focused and gives you a checklist for validating results.
For example, a product catalog task might require product name, price, product URL, and availability. Decide how to represent missing prices or unavailable items before collecting records. Keep each source URL with its extracted record if you may need to revisit or audit it.
2. Check for an API or structured source first
Look for a documented API, RSS or Atom feed, downloadable file, or public structured data. When a supported API exists, use its documentation and access requirements instead of depending on page markup that may change. APIs often return fields in a structured format directly, while HTML extraction requires you to identify and maintain selectors.
Scrapy can retrieve data from APIs as well as HTML, so using an API does not rule out using a crawler for link discovery, scheduling, or record output. See the Scrapy overview for its crawl and item workflow.
3. Inspect a page response before parsing
Fetch one representative page and inspect the returned HTML. Search it for a distinctive value that you can see in the browser. If the value is present, a normal HTTP request plus an HTML parser may be enough. If it is absent, the browser could be loading it through JavaScript or a separate data request.

Here is a compact Python example using Requests and Beautiful Soup. Install the dependencies with python -m pip install requests beautifulsoup4. Replace the example URL and CSS selectors with those from a page you are permitted to access.
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
url = "https://example.com/catalog/item-1"
response = requests.get(
url,
headers={"User-Agent": "ExampleResearchBot/1.0 (contact: you@example.com)"},
timeout=30,
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
name = soup.select_one("h1.product-title")
price = soup.select_one(".product-price")
link = soup.select_one("a.canonical-link")
record = {
"source_url": response.url,
"name": name.get_text(" ", strip=True) if name else None,
"price": price.get_text(" ", strip=True) if price else None,
"canonical_url": urljoin(response.url, link["href"]) if link and link.has_attr("href") else None,
}
print(record)
The selectors above are examples, not selectors for a real site. Inspect the document and choose selectors based on stable structure. CSS selectors such as .product-price target classes; XPath is useful for relationships that are awkward to express in CSS. Scrapy documents both CSS and XPath selectors, along with alternatives such as Beautiful Soup and lxml.
Extract attributes as well as text
Some useful values are stored in attributes rather than visible text: links use href, images use src or data-src, and metadata can use content. Always account for relative links with urljoin. Consider lazy-loaded images: the actual address may be in a data attribute, not src.
image = soup.select_one(".product-image")
image_url = None
if image:
image_url = image.get("src") or image.get("data-src")
if image_url:
image_url = urljoin(response.url, image_url)
4. Extract data from multiple pages with Scrapy
For a set of pages connected by pagination or category links, a crawler framework helps separate requesting, parsing, link following, and output. Scrapy spiders define start URLs and callbacks; callbacks can yield dictionaries or items, and pipelines can process output. The following minimal spider follows links matching an example pagination selector and yields records from pages.
import scrapy
class CatalogSpider(scrapy.Spider):
name = "catalog"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/catalog"]
custom_settings = {"ROBOTSTXT_OBEY": True}
def parse(self, response):
for card in response.css("article.product"):
name = card.css("h2 a::text").get()
href = card.css("h2 a::attr(href)").get()
price = card.css(".price::text").get()
yield {
"name": name.strip() if name else None,
"url": response.urljoin(href) if href else None,
"price": price.strip() if price else None,
"source_page": response.url,
}
next_page = response.css("a.next::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
Install Scrapy with python -m pip install scrapy, save the code as catalog_spider.py in a Scrapy project’s spiders directory, and run scrapy crawl catalog -O items.json from the project directory. The domain, selectors, and next-page link are placeholders; inspect the target site and adjust them. Scrapy’s overview shows callbacks, link following, structured items, and pipelines.
Pagination and duplicate records
Pagination may use a next link, numbered pages, a cursor, or an API parameter. Follow the site’s actual mechanism rather than guessing page numbers. If records can appear on several listing pages, deduplicate using a stable key such as a canonical URL or source identifier. Preserve the source URL and retrieval time when freshness or traceability matters.
5. Handle JavaScript-loaded data
If the desired value is missing from the HTTP response, open the page in a browser and inspect the Network panel while it loads or while you perform the relevant interaction. Look for a request that returns JSON or another structured response. If it is an ordinary data request and permitted for your use, reproduce it with the necessary documented parameters and headers. This is usually simpler and transfers less page content than parsing a rendered DOM.
Data can also be embedded in a script element as JSON. Inspect the script payload and parse the relevant object rather than executing arbitrary page scripts. If request reproduction is impractical, or the browser-rendered result itself is what you need, use browser automation. Scrapy’s dynamic content guide explains these options and discusses Playwright integration; it notes that direct Playwright use can bypass Scrapy components, while scrapy-playwright provides tighter integration.
When a screenshot is useful
A screenshot is a visual record of a page, not a structured dataset. It can help document what a rendered page looked like or produce an image or PDF for review, but you still need an API, parser, or browser DOM extraction for named fields such as price or title.
6. Respect crawler rules and access restrictions
Read the target site’s robots.txt and applicable terms, respect stated restrictions, and obtain permission when needed. The robots protocol communicates crawler rules requested by site operators; it does not grant access to restricted material. RFC 9309 says, “These rules are not a form of access authorization.” A path not disallowed by robots.txt is not automatically permitted for every purpose.
Scrapy offers robots middleware. Set ROBOTSTXT_OBEY to True so the crawler observes robots.txt rules. Do not bypass authentication, technical access controls, or explicit restrictions. Keep request rates restrained, and stop if the site indicates that automated requests are unwanted. No universal request-rate number applies to every site.
7. Validate and store the output
Before relying on a dataset, inspect representative records and check:
- Required fields exist, and missing values are represented consistently.
- Text is decoded correctly and whitespace is normalized.
- URLs are resolved and point to the expected pages.
- Pagination did not skip pages or create excessive duplicates.
- Prices, dates, or other values have the expected format and units.
Store records as JSON, CSV, or another format suited to the next step. Keep source URLs and retrieval times when you may need to verify a value later. For recurring collection, compare a fresh sample with prior output so a changed page structure does not silently produce empty fields.
Choosing the right extraction method
| Page/data situation | Good starting method | Tradeoff |
|---|---|---|
| Official API or downloadable data | Use the documented source | Follow its terms, authentication, and limits |
| Fields present in initial HTML, one or a few pages | HTTP client plus parser | Selectors need maintenance when markup changes |
| Many pages with links and structured output | Scrapy crawler | More project setup; crawl rules and output need configuration |
| Data returned by a browser network request | Reproduce the request when permitted | Request parameters or behavior may change |
| Rendered DOM is necessary or request route is impractical | Headless browser automation | More browser setup and resource use |
Or skip the browser setup
If your task is to save the rendered page as an image or PDF, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns a screenshot or PDF; it is not a replacement for extracting structured fields. The request below saves an image of a page. See the ScreenshotNeo API documentation for the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and responses identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, no card required.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Selector returns no value | Selector does not match, markup differs, or data is dynamic | Inspect the response HTML and test the selector against that document; then check the browser Network panel if the value is absent. |
| HTTP 403 or 429 | The site refuses the request or is limiting traffic | Check permission and site rules, reduce request frequency, and stop if automated access is unwanted. Do not try to evade access controls. |
| Some pages are missing | Pagination or link discovery is incomplete | Inspect the actual next-page mechanism, follow permitted links, and compare discovered pages with the expected scope. |
| Fields are intermittently empty | Markup variants, delayed content, or a changed page template | Handle optional elements, sample several pages, and validate required fields before writing records. |
| Duplicate rows | Overlapping listings or repeated links | Deduplicate on a stable page ID or canonical URL and retain the source URL for diagnosis. |
| Broken image URLs | Relative or lazy-loaded address was treated as absolute or read from the wrong attribute | Resolve paths with urljoin and inspect attributes such as data-src. |
Performance, reliability, and cost
Prefer a documented data endpoint when it serves the exact fields you need: it can avoid downloading and parsing a full page. For static HTML, use a parser before introducing a full browser. A crawler becomes useful when link following and many structured records are central to the task. Browser automation should be reserved for cases where the rendered result matters or simpler request-level methods are unsuitable.
Reliability depends on validating assumptions: page templates change, network requests can fail, and selectors can stop matching. Set request timeouts, check HTTP status, make fields optional where the source is inconsistent, and validate output before treating a run as complete. Keep the collection limited to the scope you need and operate within the site’s access rules.
Cost is mainly engineering and operating time: parser maintenance for markup changes, crawler infrastructure for recurring multi-page jobs, or browser resources when rendering is required. There is no universal cost or speed figure for these methods; page complexity, volume, and infrastructure vary. If you need visual screenshots rather than field extraction, ScreenshotNeo’s free allowance and published paid tiers provide a separate per-shot option.
FAQ
Is website data extraction the same as web scraping?
Website data extraction is the goal: collecting selected information in a useful structure. Web scraping is one way to do it by retrieving and parsing pages. An official API or dataset may be a better source when available.
Should I use CSS or XPath?
Use whichever expresses the target clearly and remains understandable to maintain. CSS works well for common class, attribute, and descendant selection; XPath can express more complex relationships. Both are supported by Scrapy selectors.
Can a screenshot API extract prices into a spreadsheet?
A screenshot API returns a visual image or PDF, not a table of extracted fields. Use a documented data source, HTML parser, or browser DOM extraction for structured records.
Does robots.txt give permission to scrape a site?
No. It communicates crawler rules and is not access authorization. Check applicable terms and restrictions, and get permission where needed.


