ScreenshotNeo

BlogHow-to

How to Scrape AutomationDirect Product Pages

Use AutomationDirect’s product API first, then reconcile HTML, documents, and PDFs into a reliable parts dataset.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: Start with AutomationDirect’s official Product Data API, if its access terms fit your project. Use the Products taxonomy and selectors to discover canonical product URLs and part numbers, fetch HTML only for fields the API does not expose, and attach manuals, CAD, compliance files, and certificates as linked records. Keep the manufacturer part number as your reconciliation key, store raw values with retrieval timestamps, and use PDF catalogs for bulk discovery and archival snapshots rather than as the only source of current price or stock.

1. Choose the acquisition path

Source Best use Limits and checks
Product Data API Structured, current product fields Confirm authentication, quotas, pagination, field names, and permitted uses with AutomationDirect before production.
Product pages (HTML) Page-specific text, visible price/stock, links and labels Markup can change; rendered content may require a browser.
Manuals, CAD, compliance files Technical and regulatory detail Store as linked child records with their own hashes and timestamps.
PDF catalogs Bulk discovery and historical snapshots Values can lag revisions; reconcile important fields against current API or item pages.

2. Design a durable record

Use the displayed manufacturer part number as the stable identity across API responses, pages, and documents. A practical record keeps:

  • part_number, title, category and product-family fields
  • normalized fields plus the original text for every specification and unit
  • canonical URL and source type (api, html, pdf, manual, cad, compliance)
  • price and stock strings exactly as observed, with retrieved_at
  • HTTP status, content hash, revision/status fields when exposed
  • document children containing URL, file hash, document type and retrieval time

Do not silently convert voltage ranges, dimensions, temperature ratings or other units. Preserve the source wording beside any normalized value so a later schema change is auditable.

3. Discover products and reconcile identities

  1. Start at the Products taxonomy and selectors; enqueue category results and canonical product URLs.
  2. Record the displayed part number before following document links.
  3. Deduplicate by part number and canonical URL, retaining aliases and product-family context.
  4. Queue manuals, CAD, compliance and certificate links as child resources instead of flattening them into the product row.

The same part can appear in a selector, a product page and a catalog. Treat the part number as the join key and retain every source URL.

4. Use the Product Data API first

AutomationDirect publishes a discovery page describing a Product Data API for accurate product information retrieval by AI assistants and agents. The public discovery material does not define authentication, quotas, pagination or a final schema, so verify those details with AutomationDirect and implement the client behind a small adapter.

API client pattern (Python)

import os
import time
import requests

BASE_URL = os.environ["AUTOMATIONDIRECT_API_URL"]
TOKEN = os.environ.get("AUTOMATIONDIRECT_API_TOKEN")

session = requests.Session()
if TOKEN:
    session.headers["Authorization"] = f"Bearer {TOKEN}"

params = {"page": 1, "page_size": 100}
while True:
    response = session.get(BASE_URL, params=params, timeout=30)
    response.raise_for_status()
    payload = response.json()
    items = payload.get("items", payload if isinstance(payload, list) else [])
    for item in items:
        part = item.get("part_number") or item.get("partNumber")
        if not part:
            continue
        print(part, item)
    next_page = payload.get("next_page") if isinstance(payload, dict) else None
    if not next_page or not items:
        break
    params["page"] = next_page
    time.sleep(0.2)

Map the actual field names and pagination contract after you receive API documentation or credentials. Keep the unmodified response for replay and change detection.

5. HTML fallback for missing fields

Fetch HTML when a field is absent from the API or when you need the exact page label, visible commercial text, or document links. Identify elements by stable attributes and labels where possible; avoid relying on positional selectors.

Runnable HTML collector (Python)

from datetime import datetime, timezone
from hashlib import sha256
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup

url = "https://example.invalid/product"
r = requests.get(url, headers={"User-Agent": "product-research/1.0 (contact: you@example.com)"}, timeout=30)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")

def text(selector):
    node = soup.select_one(selector)
    return node.get_text(" ", strip=True) if node else None

record = {
    "url": r.url,
    "retrieved_at": datetime.now(timezone.utc).isoformat(),
    "http_status": r.status_code,
    "content_sha256": sha256(r.content).hexdigest(),
    "title": text("h1"),
    "part_number": text("[data-part-number], .part-number"),
    "price_text": text("[data-price], .price"),
    "stock_text": text("[data-stock], .stock"),
    "documents": [],
}
for a in soup.select("a[href]"):
    href = urljoin(r.url, a["href"])
    label = a.get_text(" ", strip=True).lower()
    if any(word in label for word in ("manual", "cad", "compliance", "certificate", "pdf")):
        record["documents"].append({"label": label, "url": href})
print(record)

The selectors above are generic. Inspect a current page and adapt them to the labels and attributes you observe; keep a fixture page in your tests so selector changes are visible.

When JavaScript rendering is required

If a direct request returns an empty shell, use a browser only for that URL class. Wait for a product identifier or specification container, capture the rendered HTML, then parse it with the same normalization code. Set a bounded navigation timeout and avoid infinite scrolling unless the page requires it.

6. Download manuals, CAD and compliance files

For each document link, stream the file, hash the bytes, and store the retrieval timestamp. Keep the document URL and part-number association even when the filename is generic. Manuals and compliance documents are authoritative for technical or regulatory details; they do not replace current commercial fields such as stock.

from pathlib import Path
from hashlib import sha256
import requests

def download(url, out_dir="documents"):
    Path(out_dir).mkdir(exist_ok=True)
    with requests.get(url, stream=True, timeout=60) as r:
        r.raise_for_status()
        digest = sha256()
        path = Path(out_dir) / "download.bin"
        with path.open("wb") as f:
            for chunk in r.iter_content(1024 * 64):
                if chunk:
                    digest.update(chunk)
                    f.write(chunk)
    return {"url": url, "path": str(path), "sha256": digest.hexdigest()}

7. Use PDF catalogs as a bulk and archival fallback

AutomationDirect describes its catalogs as searchable PDFs whose part numbers link to online pricing, specifications and stocking information. Extract part numbers and candidate URLs from a catalog, then reconcile important values against the current API or product page. Keep the catalog edition and retrieval date because product information and revisions can change; the current catalog index includes a price-change notice effective September 2, 2026.

PDF workflow

  1. Save the original PDF and compute a file hash.
  2. Extract text and links with a PDF parser.
  3. Normalize part-number formatting without discarding the original token.
  4. Enqueue linked product pages for current values.
  5. Mark catalog-only values as snapshots, never as silently current stock or price.

8. A complete crawl loop with freshness controls

from datetime import datetime, timezone
import time

queue = load_discovered_urls()
while queue:
    url = queue.pop()
    if recently_fetched(url, hours=24):
        continue
    try:
        record = fetch_and_parse(url)
        record["retrieved_at"] = datetime.now(timezone.utc).isoformat()
        save_product(record)
        for child in record.get("documents", []):
            save_document_link(record["part_number"], child)
    except TemporaryError:
        requeue_with_backoff(url)
    except PermanentError as exc:
        save_error(url, str(exc))
    time.sleep(rate_limit_delay())

Use conditional requests when supported, content hashes when they are not, and a queue that can resume after interruption. Confirm crawl permissions, terms and rate limits before scaling; never bypass authentication, CAPTCHAs, access controls or published limits.

9. Validation checklist

  • Sample API records against the corresponding product page.
  • Flag missing part numbers, duplicate canonical URLs and changed specification labels.
  • Compare document links and file hashes between runs.
  • Check that price and stock values carry retrieval timestamps.
  • Reconcile PDF snapshots after catalog revisions or price notices.
  • Review a manual sample for unit and range normalization errors.

10. cURL and Node.js request examples

curl -L --max-time 30 "https://example.invalid/product" -o product.html
const res = await fetch('https://example.invalid/product');
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const html = await res.text();
console.log(html.length);

Replace the placeholder with a URL discovered from AutomationDirect’s taxonomy or API. Add retry logic only for transient failures and respect the site’s published limits and terms.

11. Troubleshooting

Symptom Likely cause Fix
401/403 from API Missing or invalid credentials, or an unapproved use Confirm authentication and permitted uses with AutomationDirect; do not rotate around access controls.
429 responses Quota or rate limit exceeded Read the published quota, slow the queue, honor Retry-After and request a higher limit if available.
HTML has no product data Content is rendered by JavaScript Use a bounded browser render for that page class, wait for a product selector, then parse.
Duplicate products Multiple URLs or selector paths for one item Canonicalize URLs and deduplicate on manufacturer part number plus canonical URL.
PDF price differs from page Catalog snapshot is older Keep the dated PDF value and reconcile against the current API or page.
Missing CAD or compliance links Resources live in separate tabs or lookup tools Follow the document vault and part-number lookup paths; store links as child records.
Parser breaks after a redesign Selector or label changed Retain raw HTML, run fixture checks, and alert on missing required fields.

12. Performance, reliability and cost

  • Performance: Prefer API pages and conditional requests; parallelize only within the confirmed quota. Stream large PDFs and cache hashes.
  • Reliability: Use retries with exponential backoff for transient network errors, checkpoint the queue, and preserve raw responses for replay.
  • Freshness: Refresh price and stock more often than slow-changing specifications, and always expose retrieved_at.
  • Cost: API, bandwidth and browser execution costs depend on your account and infrastructure. Measure request volume and avoid recrawling unchanged resources.
  • Compliance: Verify terms and crawl permission before production collection; never defeat CAPTCHAs or access controls.

Or skip the browser setup

ScreenshotNeo can capture a product page with one request when you need a visual record alongside structured data. It accepts consent banners like a visitor, removes 60+ known consent platforms, newsletter popups and chat widgets before the shot, and lets you turn each step off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and each response reports the page verdict and billing status.

See the ScreenshotNeo API docs. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Its MCP server gives Claude, Cursor and other MCP clients take_screenshot, get_page_info and capture_pdf tools. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Is there an AutomationDirect API?

AutomationDirect publishes Product Data API discovery information. Confirm credentials, quotas, schema and permitted uses directly with AutomationDirect before relying on it.

What should be the primary key?

The manufacturer part number, paired with the canonical URL and any exposed revision or status, is the practical reconciliation key.

Should I scrape PDFs instead of pages?

Use PDFs for bulk discovery and archives; reconcile current commercial and specification values against the API or live page.

How do I keep a dataset auditable?

Store raw responses, hashes, source URLs, HTTP status and retrieval timestamps, and link every manual, CAD or compliance file to its part number.