ScreenshotNeo

BlogHow-to

How to Extract Data from a Website for Free

Extract website data for free with spreadsheet imports, Python, or Scrapy. Learn how to handle JavaScript pages, validate results, and troubleshoot common issues.

By the ScreenshotNeo team4 October 20269 min read

To extract data from a website for free, first check for an official CSV, XML, RSS feed, API, or sitemap. For a small table, try a spreadsheet import formula. If the data is present in the page HTML, use Python with an HTML parser; for repeatable crawls and pagination, use Scrapy. If the browser shows data that your request does not, inspect the browser’s network requests for the source before reaching for a headless browser.

Before collecting anything, check the site’s terms and crawler guidance, keep requests restrained, and confirm that you have permission for your intended use. A publicly viewable page is not automatically permission to collect or republish its contents.

1. Choose the smallest method that fits

Method Use it when Limits
Spreadsheet import You need a small table or feed for a one-off task. The page must expose data in a supported format. Google Sheets documents traffic-related limits and throttling for import functions.
Python parser The values are in the fetched HTML and you need a modest script. You need code, and selectors can break when the page structure changes.
Scrapy You need repeatable crawling, pagination, structured exports, or crawl controls. It takes a Python project and maintenance.
Inspect network requests The page fills in data after JavaScript runs. You may need to reproduce request parameters, headers, or browser behavior. A headless browser is an escalation if the data request cannot reasonably be reproduced.

Google documents the spreadsheet functions IMPORTHTML, IMPORTXML, IMPORTDATA, and IMPORTFEED. [Google Sheets function list] Scrapy supports CSS and XPath selectors, pagination, download-delay and concurrency controls, and exports such as CSV and JSON Lines. Its selector documentation also identifies Beautiful Soup and lxml as parsing options. [Scrapy documentation] [Scrapy selectors]

2. Check for a structured source first

  1. Write down the fields you need, how many pages you need, how often they change, and the output format.
  2. Look for a downloadable CSV, XML or RSS feed, API, or sitemap. Prefer a source the site publishes for reuse when available.
  3. Review the site’s terms and robots.txt before collecting. robots.txt communicates crawler guidance; it is not access control or a privacy mechanism.
  4. Start with one representative page and the smallest collection scope that answers your question.

Google Search Central describes robots.txt as instructions about which URLs crawlers can access, mainly for managing crawler traffic. It is not a security measure, and blocked URLs can still appear in search results. Use authentication or suitable indexing controls to protect private content. [Google Search Central: robots.txt] Google’s terms also give a service-specific example of restrictions on automated access: they prohibit automated access to Google services in violation of machine-readable instructions. That example does not establish the terms for other websites. [Google Terms of Service]

3. Try a spreadsheet import for a small table

In Google Sheets, enter a formula in an empty cell and replace the URL and table index with the page and table you need:

=IMPORTHTML("https://example.com/catalog", "table", 1)

The final argument is the table index on the page. If the page publishes a CSV, try IMPORTDATA; for an XML source, use IMPORTXML with an XPath expression; for a feed, try IMPORTFEED. Consult the function help for syntax and supported inputs. [Google Sheets function list]

Use this route when you need a quick import, not as a reliable crawler. Import functions can be throttled, and pages may not expose a compatible table or feed. If the formula fails, check that the URL is publicly reachable and that the desired values are present in the page source. For a larger or recurring task, move to a script with explicit validation and request pacing.

4. Extract HTML data with Python

When the requested values appear in the server-returned HTML, a small script using Requests and Beautiful Soup can fetch and parse one page. Install the dependencies with python -m pip install requests beautifulsoup4.

import csv
import requests
from bs4 import BeautifulSoup

url = "https://example.com/catalog"
response = requests.get(
    url,
    headers={"User-Agent": "ExampleResearchScript/1.0 (contact: you@example.com)"},
    timeout=30,
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
rows = []

for item in soup.select(".product"):
    name = item.select_one(".product-name")
    price = item.select_one(".price")
    if not name:
        continue
    rows.append({
        "name": name.get_text(" ", strip=True),
        "price": price.get_text(" ", strip=True) if price else "",
    })

with open("products.csv", "w", newline="", encoding="utf-8") as output:
    writer = csv.DictWriter(output, fieldnames=["name", "price"])
    writer.writeheader()
    writer.writerows(rows)

print(f"Saved {len(rows)} rows to products.csv")

Replace .product, .product-name, and .price with selectors that match the target page. Inspect the HTML first: if the selector matches nothing, the page may use a different structure or only add the content after JavaScript executes. Scrapy’s selector guide covers CSS and XPath extraction patterns. [Scrapy selectors]

Request the page with cURL

Use cURL to inspect the server response before writing selectors. Save the response body and examine the HTML locally:

curl -L --fail --max-time 30 \
  -A 'ExampleResearchScript/1.0 (contact: you@example.com)' \
  'https://example.com/catalog' \
  -o page.html

-L follows redirects, --fail returns an error for HTTP failures, and --max-time bounds the request duration. A successful response does not guarantee that the page contains the data you see in a browser.

5. Use Scrapy for pagination and repeatable crawls

When a task spans pages or needs to run repeatedly, Scrapy provides request scheduling, pagination, crawl controls, and structured feed exports. Create a project with python -m pip install scrapy and scrapy startproject catalog_crawl. Add a spider such as this under the project’s spiders directory:

import scrapy

class CatalogSpider(scrapy.Spider):
    name = "catalog"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/catalog"]

    custom_settings = {
        "DOWNLOAD_DELAY": 1.0,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 2,
        "FEEDS": {
            "products.jsonl": {"format": "jsonlines", "encoding": "utf8"},
        },
    }

    def parse(self, response):
        for item in response.css(".product"):
            yield {
                "name": item.css(".product-name::text").get(default="").strip(),
                "price": item.css(".price::text").get(default="").strip(),
            }

        next_page = response.css("a.next::attr(href)").get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

Run it from the project directory with scrapy crawl catalog. Adjust the domain, start URL, selectors, pagination link, and pacing for the target site. Scrapy documents download delay, per-domain concurrency, and feed exports; set controls conservatively and follow the target site’s instructions. [Scrapy settings] [Scrapy feed exports]

6. Handle JavaScript-rendered pages

If the browser displays values that the Python or cURL response does not contain, treat it as a source-discovery problem first:

  1. Open browser developer tools and inspect the Network activity while loading the page.
  2. Look for a request that returns the missing data, often as JSON or another structured response.
  3. Check whether the request is an official, permitted source and whether its terms allow your intended use.
  4. If suitable, reproduce that request and validate its parameters and response. Do not assume that copying a browser request grants permission.
  5. Use a headless browser only when reproducing the source request is impractical and you are permitted to automate the page.

Scrapy recommends inspecting browser network activity to discover requests that provide dynamically loaded data; reproducing the underlying request can be simpler than rendering the whole page. [Scrapy: dynamic content]

7. Validate and save the results

Before using an export, review a sample and check:

  • Required fields are present and have the expected type and format.
  • Missing values are represented consistently rather than silently shifted into the wrong columns.
  • Pagination reached the expected end and did not revisit the same page indefinitely.
  • Duplicate records are identified using a stable key where one exists.
  • Text encoding and CSV quoting preserve punctuation and non-ASCII characters.
  • A sample of extracted records agrees with the source page.

For recurring jobs, log request failures and record counts. Recheck selectors when the page changes. Keep only the fields needed for the task, and define a retention period appropriate to your use.

8. Troubleshoot common failures

Symptom Likely cause What to do
Spreadsheet formula returns an error or no rows The source is unsupported, throttled, inaccessible, or not exposed as the specified table/feed. Check the URL and function syntax, try the published CSV/XML/feed if available, and use a script for a modest task. Google documents traffic-related import limits. [Google Sheets function list]
Python gets zero matches The selectors do not match the response HTML, or the content is rendered later in JavaScript. Save and inspect the response body; verify selectors against that HTML; inspect network requests if browser content differs.
HTTP 403 or 429 response The server denied or rate-limited the request. Stop rapid retries, review the site’s terms and instructions, reduce request frequency, and use an official source or ask for access where appropriate.
HTTP 404 or redirect loop The URL is stale, requires a different canonical path, or the request is being redirected unexpectedly. Check the page URL in a browser, inspect the redirect chain, and update the start URL.
Timeout or intermittent failures The site is slow, the connection is unstable, or requests are too frequent. Use a bounded timeout, reduce concurrency, add restrained retries with backoff, and log failures for later review.
Repeated or missing pages Pagination selectors are wrong, next links repeat, or the site uses a different paging model. Inspect links and page parameters, keep a visited-page set, and stop when there is no new page.
CSV opens with broken characters or columns Encoding or quoting was handled inconsistently. Write UTF-8 output and use a CSV library rather than joining values with commas manually.

9. Keep collection reliable, efficient, and proportionate

  • Use the least work that gets the data. A published feed or one spreadsheet formula is cheaper to maintain than a crawler; a parser is simpler than a browser for static HTML.
  • Control request load. Limit concurrency, use delays, and avoid repeatedly fetching pages that have not changed. Scrapy documents download delay and concurrency settings. [Scrapy settings]
  • Bound failures. Set timeouts, cap retries, log errors, and make repeated runs safe to resume where practical.
  • Expect page changes. Selectors and pagination can break; validate counts and sample records on every run.
  • Keep costs visible. The tools above are free software or built-in spreadsheet functions, but compute, storage, and maintenance still take time. A headless browser generally requires more resources than fetching static HTML, so reserve it for pages that need rendering.
  • Use data responsibly. Terms, copyright, privacy, database rights, authentication, and local law can affect collection and reuse. This guide cannot determine whether a specific use is permitted; get qualified advice for consequential projects.

10. Or skip the browser setup

If the next step is capturing a clean visual record of a page, ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request returns an image or PDF. For a screenshot, use the API key from your account and replace the target URL as needed:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and responses identify page verdict and billing status. Its MCP server exposes screenshot, page-information, and PDF capture tools to AI agents. One thousand shots per month are free without a card; paid plans begin at $5 for 3,000 shots. Learn about ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.

Frequently asked questions

Can I extract a website into Excel or Google Sheets without writing code?

For a small compatible table or feed, try a Sheets import function and export or copy the result to your spreadsheet. Import functions can be throttled, so they are not a dependable large-scale crawler.

Is robots.txt permission to scrape a site?

No. It provides crawler guidance, not access authorization. Check the site’s terms and applicable rules separately.

What should I do when the data is not in the HTML?

Inspect browser network requests for the source that supplies it. Reproduce an appropriate request if permitted; use a headless browser only if needed.

Which output format should I choose?

CSV is convenient for spreadsheets and simple tables; JSON Lines works well for record-by-record processing and larger exports. Choose based on the next tool in your workflow.