How to Scrape a Shopify Store with BrowserQL
Use BrowserQL to navigate and extract Shopify storefront data, with runnable examples, authorization guidance, troubleshooting, and API alternatives.
BrowserQL lets you describe browser actions in GraphQL: navigate to an authorized Shopify storefront page, wait for its content, and extract the fields you need. Use a browser when the information is rendered in the page or requires browser interaction. If you own the store or have merchant authorization and Shopify’s supported API exposes the fields you need, prefer the versioned Storefront API.
A successful browser request does not establish permission to collect, store, republish, or commercially use a store’s data. Confirm authorization and data-use rights first, collect only what you need, and use the minimum permissions required.
1. Choose BrowserQL or Shopify’s Storefront API
| Need | Use | Considerations |
|---|---|---|
| Page-rendered content or browser interactions | BrowserQL | Page structure and selectors can vary by theme. BrowserQL provides navigation, waits, interaction, and extraction. |
| Buyer-facing fields exposed by Shopify for an authorized integration | Storefront API | It is versioned and permissioned. Choose a supported version and account for query complexity and automated-traffic limits. |
| Backend merchant data or store administration | Shopify Admin API | This is a separate API. It requires merchant-granted scopes and is not a substitute for public storefront browsing. |
These are implementation tradeoffs, not benchmark results. Browserless documents BrowserQL as a declarative GraphQL browser automation interface. Its BAP TypeScript and Python SDKs construct the same mutations; the SDK is the documented route for those languages, while direct GraphQL is useful from other languages or the hosted IDE. Browserless BrowserQL documentation.
2. Check authorization and scope
Shopify’s API Terms restrict systematic or automated collection through the Shopify API, including scraping and extraction, unless authorized by Shopify or where applicable law expressly prevents that restriction. The terms also call for limiting requests to the data needed for the app’s intended functionality and to granted permissions. This is not a complete legal analysis of public webpage access in every jurisdiction. Shopify API Terms of Use.
For authorized crawling of a public Shopify online store, Shopify documents Web Bot Auth as a way to securely authorize crawlers, scripts, or tools. Its examples include audits, testing, and data analysis; it is not blanket permission for any collection or reuse. Shopify guidance on crawling storefronts.
- Identify the store owner or other source of authorization and the permitted purpose.
- Confirm whether storage, redistribution, and commercial reuse are allowed.
- Request only the pages and fields needed; avoid collecting unrelated customer or personal data.
- Do not treat a successful request, stealth feature, proxy, or CAPTCHA handling as evidence of permission.
3. Get a Browserless token and select an endpoint
BrowserQL requests need a Browserless API token, passed as a token query parameter. The documented HTTP endpoints are /chromium/bql for open-source Chromium, /chrome/bql for a genuine Google Chrome build, and /stealth/bql for a privacy-hardened browser configuration. Use the endpoint that matches the documented behavior and configuration of your account; no endpoint is universally best. See the current BrowserQL endpoint documentation before deployment.
Keep the token in an environment variable or secret manager. Do not commit it to source control, put it in client-side code, or log full endpoint URLs containing the token. The examples below use placeholders; set the endpoint and token for your account.
4. Run a minimal BrowserQL extraction
Start with one product or collection page you are allowed to access. The mutation below navigates, waits for the initial DOM to be ready, and returns page text. Replace the example URL with an authorized page. The dossier does not verify any Shopify-specific selector or live store structure.
mutation ScrapePage {
goto(url: "https://your-authorized-store.example/products/example", waitUntil: domContentLoaded) {
status
}
text {
text
}
}
Send the GraphQL document as the request body to your chosen BrowserQL endpoint. For example, using cURL (set BROWSERLESS_ENDPOINT to the endpoint URL without a token query parameter):
export BROWSERLESS_ENDPOINT="https://production-sfo.browserless.io/chromium/bql"
export BROWSERLESS_TOKEN="YOUR_BROWSERLESS_TOKEN"
curl --fail-with-body --silent --show-error \
--get "${BROWSERLESS_ENDPOINT}" \
--data-urlencode "token=${BROWSERLESS_TOKEN}" \
--header "Content-Type: application/json" \
--data-urlencode 'query=mutation ScrapePage { goto(url: "https://your-authorized-store.example/products/example", waitUntil: domContentLoaded) { status } text { text } }'
BrowserQL is GraphQL over HTTP, but account endpoint details and accepted request encoding should be checked against Browserless’s current docs. If your account expects a JSON POST body, use this equivalent form instead:
curl --fail-with-body --silent --show-error \
"${BROWSERLESS_ENDPOINT}?token=${BROWSERLESS_TOKEN}" \
-H "Content-Type: application/json" \
--data '{"query":"mutation ScrapePage { goto(url: \"https://your-authorized-store.example/products/example\", waitUntil: domContentLoaded) { status } text { text } }"}'
For actual field extraction, identify the page’s rendered structure in an authorized session and use BrowserQL’s documented text, attribute, or structured extraction operations. Do not assume a theme-independent product selector: themes and custom storefronts differ. Keep the result limited to necessary fields.
5. Wait for JavaScript-rendered content
domContentLoaded means the initial document has been parsed; it does not guarantee that a storefront’s asynchronously rendered product details are ready. If needed, wait for a selector that you observed on the authorized page, or for a relevant browser event. BrowserQL documents waitForSelector and waitForEvent. Adapt the selector and operation syntax to the current BrowserQL schema and the actual page; no store-specific selector is asserted here.
mutation ScrapeRenderedPage {
goto(url: "https://your-authorized-store.example/products/example", waitUntil: domContentLoaded) {
status
}
waitForSelector(selector: "YOUR_OBSERVED_PRODUCT_SELECTOR") {
time
}
text {
text
}
}
Prefer a condition tied to the content you need over an arbitrary long delay. A delay can waste session time on fast pages and still fail on slow ones. If the page requires a click, use the documented interaction operation before extraction, and verify that the interaction is within your authorization.
6. Use the Storefront API when it fits
For a store you own or where the merchant has granted the required authorization, Shopify’s Storefront API provides a versioned GraphQL interface for buyer-facing storefront functionality, including products, collections, search, pages, blogs and articles, and carts. The versioned reference surfaced for this dossier is 2026-04; Shopify advises selecting a currently supported version. Its endpoint pattern is https://{store_name}.myshopify.com/api/2026-04/graphql.json, with GraphQL POST requests. Confirm the supported version and access requirements in the Shopify Storefront API documentation.
Storefront API access can be tokenless, public-token, or private-token depending on use. Public tokens are intended for browser/mobile use and private tokens for server-side use. Tokenless requests have a query complexity limit of 1,000. Some features—including product tags, metafields/metaobjects, online-store menus, and customers—require token-based access. Shopify also limits automated Storefront API traffic and crawlers, with the strictest treatment applying to unsigned traffic; consult Shopify’s current documentation, including its Web Bot Auth guidance, before designing automated collection.
Do not confuse the Storefront API with the Admin API. The Admin API is for backend merchant data and requires scopes granted by the merchant. If your authorized use case needs fields unavailable through Storefront API, confirm the appropriate API and scopes with the store owner.
7. Plan a crawl beyond one page
A one-page extraction and a site-wide crawl are different jobs. First validate one permitted product or collection page, confirm the fields and waits, then determine whether broader collection is necessary and authorized. A larger crawl also increases request volume, session use, operational complexity, and the amount of data that must be governed.
- Write down the purpose, permitted domain/path scope, fields, and retention needs.
- Inspect a representative page and record which content is present initially versus rendered later.
- Build extraction around the observed page structure and handle missing fields as normal cases.
- Set conservative concurrency and retries consistent with your authorization and account limits.
- Record source URL, retrieval time, API/browser version, and extraction version where useful for debugging and provenance.
- Stop or reduce collection when access is denied, the authorization scope is unclear, or the page presents a challenge you are not authorized to bypass.
BrowserQL session ceilings are plan-dependent in Browserless’s documentation: Free 2 minutes, Prototyping (20k) 15 minutes, Starter (180k) 30 minutes, Scale (500k) 60 minutes, and Enterprise/self-hosted custom. These figures can change; verify current limits for your account before sizing work. Avoid designing a job that depends on keeping one browser session alive for a whole catalog.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Unauthorized or token error | Missing, invalid, expired, or incorrectly encoded Browserless token. | Check account credentials and pass the token in the documented query parameter. Keep it out of logs and client code. |
| GraphQL validation error | Operation or argument does not match the current BrowserQL schema. | Check the current schema/docs or use BAP for TypeScript/Python. Verify operation names and argument types. |
| Navigation succeeds but product text is absent | Content is rendered after initial DOM readiness, or the page uses a different structure. | Wait for an observed selector or event, then inspect the rendered page. Do not guess a universal Shopify selector. |
| Selector wait times out | Selector is wrong, content is absent, or the page did not reach the expected state. | Confirm the selector in an authorized browser session, check navigation status, and handle unavailable content explicitly. |
| Session ends before extraction | Task exceeded the plan’s maximum session duration or spent too long waiting. | Reduce unnecessary waits and work per session; check the current plan ceiling. Split authorized work into smaller jobs. |
| Storefront API rejects a field or request | Wrong API version, schema, token type, permission, or query complexity. | Use a supported version, confirm token requirements/scopes, and reduce the query. Check Shopify’s current API reference. |
| Automated requests are limited or denied | Traffic policy, authorization, bot controls, or API limits apply. | Confirm authorization and the applicable limits with the merchant/platform. Do not infer permission from an anti-detection capability. |
| Extraction returns less data than expected | Fields are not in the rendered page, are conditional, or are outside API access permissions. | Check the data source and authorization; use the supported API or request appropriate merchant access when available. |
9. Performance, reliability, and cost
No comparative performance measurements or live-store tests are available for this guide. In general, browser work includes navigation, rendering, waits, and extraction; unnecessary waits and repeated page loads add latency and consume session capacity. An API request can avoid browser rendering when its authorized schema contains the required fields, but query limits and automated-traffic rules still apply.
- Reliability: Browser selectors depend on theme markup and may need maintenance after storefront changes. Detect missing or malformed fields rather than silently treating them as valid.
- Retries: Retry transient transport failures cautiously, with a bounded policy and backoff. Do not retry authorization failures or access denials as if they were transient.
- Session use: Browserless publishes plan-specific maximum session durations; check current limits and keep individual jobs comfortably within them.
- Cost: Browserless plan and usage pricing are account-specific and were not verified here. Check current Browserless pricing and your account usage before estimating a crawl. Shopify API limits are operational constraints, not a cost estimate.
- Data handling: Minimize fields, access, retention, and onward sharing according to the authorization and applicable requirements.
10. Or skip the browser setup
If your goal is a screenshot rather than structured extraction, ScreenshotNeo is a website screenshot API and MCP server. It does not replace BrowserQL for extracting structured product fields. Its one-request API returns a PNG, JPEG, WebP, or PDF; see the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://your-authorized-store.example/products/example -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://your-authorized-store.example/products/example"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://your-authorized-store.example/products/example' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners are accepted like a visitor and removed before the shot, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free plan.
11. Frequently asked questions
Does BrowserQL require a browser installed on my machine?
BrowserQL is a hosted Browserless browser automation interface. You send a request to a Browserless endpoint, so the browser runs in its managed environment rather than as a locally launched browser.
Can BrowserQL extract every field shown in Shopify Admin?
No. A public storefront page and the Storefront API expose different data from the merchant backend. Admin data requires the appropriate merchant authorization and API scopes.
Does using stealth or CAPTCHA features make scraping permitted?
No. Those are technical product capabilities. Authorization and downstream data-use rights must be established separately.
When is a screenshot API useful here?
Use one when the deliverable is a visual capture for review, documentation, or monitoring. For product fields, use an authorized API or browser extraction workflow that returns the data you need.


