ScreenshotNeo

BlogHow-to

How to Scrape Cass Art Product Pages

Extract Cass Art products, prices, stock and variants into reliable CSV records with a sitemap-first crawler, validation and change detection.

By the ScreenshotNeo team1 October 202610 min read

Direct answer: Scrape Cass Art in layers. Start with its published category sitemap and category pages to discover product URLs, request each page with a stable user agent and conservative rate limits, parse server-rendered HTML and JSON-LD first, and use a browser only when a required field appears after JavaScript runs. Store one record per SKU or variant, preserve raw responses and timestamps, validate that price, stock, colour, size and image belong to the same variant, and obey Cass Art’s terms and the live rules in /robots.txt.

Cass Art says stock shown on its site is only a guide, and that displayed colour can differ by screen or device. Treat every price and availability value as time-sensitive evidence, not a permanent fact. Its terms also state that site content is copyrighted and may not be reproduced without written permission beyond limited individual use. Review the Cass Art site and terms before operating a production crawler.

1. Decide what a product record means

A product page can represent several sellable choices. A scraper that emits one row per page can mix a selected colour’s price with another colour’s SKU. Define the output before writing selectors.

Field Purpose
url Requested page URL.
canonical_url Canonical link, when supplied by the page.
title, brand Product identity.
product_code, sku Stable identifiers used for deduplication.
variant_name, colour, size The selected sellable option.
price, sale_price, currency Prices as displayed at retrieval time.
availability Exact visible wording, such as an in-stock or unavailable message.
description, specifications, breadcrumbs Descriptive and classification data.
image_urls Images associated with that product or variant.
source_retrieved_at, http_status Evidence about when and how the record was obtained.
content_hash, parser_version Change detection and reproducibility.
raw_snapshot_reference Pointer to retained HTML/JSON, stored under your own retention policy.

Make one record per SKU or variant where possible. If a selector changes but the extracted SKU or price does not change, reject the record for review instead of guessing.

2. Discover pages with the category sitemap

Use Cass Art’s category sitemap as the first seed. It covers paths for paints, brushes, paper, canvas, studio equipment, art books, craft and brands. The representative watercolour paint sets category is a category page with a heading and descriptive copy, not a single product record.

  1. Fetch https://www.cassart.co.uk/robots.txt before production crawling.
  2. Parse the user-agent group that applies to your crawler, including Disallow and any crawl-delay signal.
  3. Follow sitemap declarations from that file and identify category URLs.
  4. Request category pages and collect product links, canonicalising absolute URLs and removing fragments.
  5. Keep a URL ledger with discovery source, first-seen time and status.

Robots and sitemaps are operational controls, not permission to copy content. Cass Art’s own terms and any written permission remain controlling.

3. Request pages conservatively

import time
import requests
from urllib.parse import urljoin, urldefrag

SESSION = requests.Session()
SESSION.headers.update({
    "User-Agent": "ProductResearchBot/1.0 (+https://example.com/contact)",
    "Accept": "text/html,application/xhtml+xml",
})

def fetch(url, attempts=3):
    for attempt in range(attempts):
        response = SESSION.get(url, timeout=30)
        if response.status_code not in (403, 429) and response.status_code < 500:
            response.raise_for_status()
            return response
        time.sleep(min(30, 2 ** attempt))
    response.raise_for_status()

html = fetch("https://www.cassart.co.uk/watercolours/watercolour-paint-sets").text
print(len(html))

Use a low request rate, cache successful responses, and back off on 403, 429 and 5xx responses. Do not retry a permanent 404 indefinitely. A stable, honest user agent makes your traffic identifiable; replace the example contact URL with a real address before deployment.

4. Parse HTML and JSON-LD before using a browser

Server HTML is cheaper and easier to reproduce. Extract canonical links, headings, visible price and availability, breadcrumbs, images, specifications, product code/SKU and every JSON-LD block. JSON-LD often gives a coherent product object, but still validate it against visible page content.

# pip install requests beautifulsoup4 lxml
import csv, hashlib, json, re
from datetime import datetime, timezone
from bs4 import BeautifulSoup

PRODUCT_URL = "https://www.cassart.co.uk/REPLACE-WITH-A-PRODUCT-PATH"
PARSER_VERSION = "cassart-v1"

html = fetch(PRODUCT_URL).text
soup = BeautifulSoup(html, "lxml")
def text(css):
    node = soup.select_one(css)
    return " ".join(node.stripped_strings) if node else None

def first_meta(*selectors):
    for selector in selectors:
        node = soup.select_one(selector)
        if node and node.get("content"):
            return node["content"].strip()
    return None

json_ld = []
for node in soup.select('script[type="application/ld+json"]'):
    try:
        value = json.loads(node.string or node.get_text())
        json_ld.extend(value if isinstance(value, list) else [value])
    except (TypeError, json.JSONDecodeError):
        continue
product_ld = next((x for x in json_ld if isinstance(x, dict) and x.get("@type") in ("Product", ["Product"])), {})
offers = product_ld.get("offers", {})
if isinstance(offers, list):
    offers = offers[0] if offers else {}

record = {
    "url": PRODUCT_URL,
    "canonical_url": first_meta('link[rel="canonical"]'),
    "title": product_ld.get("name") or text("h1"),
    "brand": (product_ld.get("brand") or {}).get("name") if isinstance(product_ld.get("brand"), dict) else product_ld.get("brand"),
    "product_code": product_ld.get("mpn") or product_ld.get("productID"),
    "sku": product_ld.get("sku"),
    "variant_name": None,
    "colour": None,
    "size": None,
    "price": offers.get("price"),
    "currency": offers.get("priceCurrency"),
    "sale_price": None,
    "availability": offers.get("availability") or text("[class*=availability], [class*=stock]"),
    "description": product_ld.get("description") or text("[class*=description]"),
    "specifications": {},
    "breadcrumbs": [x.get_text(" ", strip=True) for x in soup.select('[aria-label="breadcrumb"] a, nav.breadcrumb a')],
    "image_urls": product_ld.get("image", []),
    "source_retrieved_at": datetime.now(timezone.utc).isoformat(),
    "http_status": 200,
    "content_hash": hashlib.sha256(html.encode("utf-8")).hexdigest(),
    "parser_version": PARSER_VERSION,
    "raw_snapshot_reference": None,
}
if isinstance(record["image_urls"], str):
    record["image_urls"] = [record["image_urls"]]
print(json.dumps(record, indent=2, ensure_ascii=False))

The CSS selectors for visible price, stock and specifications are deliberately fallbacks rather than claims about a permanent Cass Art DOM. Inspect a current page, keep selectors in version control, and alert when they return no value. Never silently replace a missing value with a neighbouring variant’s value.

5. Build a sitemap-first crawler in Python

# pip install requests beautifulsoup4 lxml
import csv, hashlib, json, time
from datetime import datetime, timezone
from urllib.parse import urljoin, urldefrag
from bs4 import BeautifulSoup

CATEGORY_URLS = ["https://www.cassart.co.uk/watercolours/watercolour-paint-sets"]
OUT = "cassart_products.csv"

def clean_url(href, base):
    if not href: return None
    absolute, _ = urldefrag(urljoin(base, href))
    return absolute if absolute.startswith("https://www.cassart.co.uk/") else None

def product_links(category_url):
    response = fetch(category_url)
    soup = BeautifulSoup(response.text, "lxml")
    links = set()
    for a in soup.select("a[href]"):
        link = clean_url(a.get("href"), category_url)
        if link and link.rstrip("/") != category_url.rstrip("/"):
            links.add(link)
    return sorted(links)

rows = []
for category in CATEGORY_URLS:
    for url in product_links(category):
        try:
            response = fetch(url)
            soup = BeautifulSoup(response.text, "lxml")
            canonical = soup.select_one('link[rel="canonical"]')
            title = soup.select_one("h1")
            rows.append({
                "url": url,
                "canonical_url": canonical.get("href") if canonical else None,
                "title": title.get_text(" ", strip=True) if title else None,
                "source_retrieved_at": datetime.now(timezone.utc).isoformat(),
                "http_status": response.status_code,
                "content_hash": hashlib.sha256(response.content).hexdigest(),
                "parser_version": "cassart-v1",
            })
        except Exception as exc:
            rows.append({"url": url, "http_status": None, "error": str(exc)})
        time.sleep(1.0)

fields = sorted({key for row in rows for key in row})
with open(OUT, "w", newline="", encoding="utf-8") as handle:
    writer = csv.DictWriter(handle, fieldnames=fields)
    writer.writeheader(); writer.writerows(rows)

Replace the discovery selector with selectors reviewed against the current category markup. A category page can contain editorial links, filters and pagination, so apply product URL rules and deduplicate before requesting detail pages.

6. Handle variants without mixing data

  1. Locate the page’s variant controls and enumerate their option values.
  2. For each option, perform the smallest supported interaction that changes the product object.
  3. After the change, read the selected option, SKU, price, availability and image from the same DOM state.
  4. Require a changed SKU or another explicit product identifier when the site exposes one.
  5. Reject a row when the option changed but the SKU and price stayed unchanged, unless the page explicitly confirms they are shared.
  6. Store the raw response or browser snapshot used for each variant.

Do not infer colour, size or availability from a title, URL slug or adjacent option. If a variant is unavailable, preserve the exact unavailable wording and use a null price when no price is shown.

7. Browser-rendering fallback

Use a browser only for fields missing from server HTML, such as a selector populated after JavaScript runs. Keep the same URL ledger and parser output so browser and non-browser records are comparable.

# pip install playwright
# playwright install chromium
from playwright.sync_api import sync_playwright

url = "https://www.cassart.co.uk/REPLACE-WITH-A-PRODUCT-PATH"
with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto(url, wait_until="networkidle", timeout=60000)
    page.locator("h1").wait_for(timeout=15000)
    html = page.content()
    selected = page.locator("[aria-selected='true'], option:checked").all_text_contents()
    print(selected)
    browser.close()

Do not wait for network idle forever: analytics or chat requests can keep a page busy. Prefer a meaningful selector, a bounded delay and a maximum navigation timeout. Block nonessential resources only after confirming that images or API responses are not required for the fields you extract.

8. Export a trustworthy CSV

import csv

columns = [
    "url", "canonical_url", "title", "brand", "product_code", "sku",
    "variant_name", "colour", "size", "price", "currency", "sale_price",
    "availability", "description", "specifications", "breadcrumbs",
    "image_urls", "source_retrieved_at", "http_status", "content_hash",
    "parser_version", "raw_snapshot_reference"
]
with open("cassart_products.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=columns, extrasaction="ignore")
    writer.writeheader()
    for row in records:
        writer.writerow({k: json.dumps(row[k], ensure_ascii=False) if isinstance(row.get(k), (dict, list)) else row.get(k) for k in columns})

Keep currency and promotion text alongside numeric prices. Store retrieval time in UTC. A later consumer can then distinguish a sale price from a regular price and decide when a record is stale.

9. Node.js and cURL request examples

// Node.js 18+; npm install cheerio
import * as cheerio from "cheerio";

const url = "https://www.cassart.co.uk/REPLACE-WITH-A-PRODUCT-PATH";
const response = await fetch(url, { headers: { "User-Agent": "ProductResearchBot/1.0" } });
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);
const html = await response.text();
const $ = cheerio.load(html);
const product = {
  url,
  canonical_url: $('link[rel="canonical"]').attr("href") ?? null,
  title: $("h1").first().text().trim() || null,
  availability: $('[class*=availability], [class*=stock]').first().text().trim() || null
};
console.log(product);
curl --fail --location --user-agent 'ProductResearchBot/1.0' \
  'https://www.cassart.co.uk/REPLACE-WITH-A-PRODUCT-PATH' \
  --output product.html

10. Reliability, performance and cost

  • Coverage: Sitemap-first discovery usually sends fewer requests than blindly walking every category link.
  • Resilience: Retain raw HTML/JSON, content hashes and parser versions so you can reprocess records after selector changes.
  • Freshness: Refresh stock and price immediately before displaying a buying recommendation.
  • Throughput: Cache unchanged pages, limit concurrency and back off on 403, 429 and 5xx responses. Browser rendering consumes more CPU and bandwidth than HTML requests.
  • Failure handling: Record nulls and errors explicitly. A missing field is different from a zero price or an unavailable product.
  • Compliance: Robots rules do not replace Cass Art’s terms or written permission. Keep request volume and retained content within the permission you have.

11. Troubleshooting

Symptom Likely cause Fix
403 or 429 Requests are too frequent, disallowed or look automated. Stop, review robots.txt and terms, reduce concurrency, add backoff and seek permission where needed.
Empty price or stock Field is rendered by JavaScript or selector changed. Inspect raw HTML and JSON-LD, then use a bounded browser fallback and version the selector.
Category rows are not products Category pages contain editorial and filter links. Filter to canonical product URLs and verify that a detail page has a product identity.
Colour and price disagree Variant state was not captured atomically. Read SKU, option, price, stock and image after each selection and reject inconsistent rows.
Duplicate products Tracking parameters, fragments or alternate paths. Remove fragments, prefer canonical URLs and deduplicate by canonical URL plus SKU.
Browser never finishes Analytics or chat keeps network activity open. Use a selector wait and maximum timeout instead of an unlimited network-idle wait.
Prices change between runs Promotions and stock are time-sensitive. Store timestamps, currency, sale text and exact availability; refresh before use.

12. Or skip the browser setup

If you need a clean image of a Cass Art page for a catalog, review workflow or agent, ScreenshotNeo provides a single screenshot request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.cassart.co.uk/watercolours/watercolour-paint-sets -o cassart.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.cassart.co.uk/watercolours/watercolour-paint-sets"}, timeout=90)
r.raise_for_status()
open("cassart.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.cassart.co.uk/watercolours/watercolour-paint-sets' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`); if (!res.ok) throw new Error(`${res.status}`); const image = Buffer.from(await res.arrayBuffer());

ScreenshotNeo also has an MCP server so Claude, Cursor and other MCP clients can call take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

13. FAQ

Does Cass Art provide a product feed?

The research establishes a published category sitemap. Treat it as a discovery source and confirm the current sitemap declarations in /robots.txt; do not assume it is a complete, permission-granting feed.

How often should stock be scraped?

There is no universally correct interval. Choose a rate that fits your use case and permission, then refresh immediately before showing a buying recommendation because Cass Art describes stock as a guide.

Should I save the full HTML?

Retaining a raw snapshot reference and the parser version makes corrections and audits possible. Retain only what your permission and retention policy allow.

Can I use displayed colour as an exact colour value?

No. Cass Art warns that colour can vary by display and that swatches are guides. Preserve the product’s own colour name and treat rendered pixels as illustrative.