Web Scraping with Client-Side Vanilla JavaScript
Learn how vanilla JavaScript fetches and parses pages the browser is allowed to read, and what to do when CORS blocks access.

Yes, you can scrape a webpage with client-side vanilla JavaScript when the browser is allowed to read its response. That includes pages and endpoints on your own origin, plus cross-origin resources whose servers permit your site through CORS. If another site does not grant that access, a browser script cannot override its decision: parsing HTML only works after you have obtained the HTML.
The usual workflow is fetch(), check the HTTP status, read the response body, and—if it is HTML—parse the text with DOMParser. This guide shows runnable examples, explains the browser boundary, and covers common errors and practical alternatives.
1. Understand what browser-side scraping can access
A web origin is the combination of scheme, host, and port. For example, https://example.com and https://example.com/products share an origin because the path does not change it. Changing the scheme, host, or port creates a different origin. The browser’s same-origin policy restricts how a page interacts with resources from other origins. MDN’s same-origin policy guide describes that boundary.

| Source | Can browser JavaScript read it? | Typical approach |
|---|---|---|
| Your page or endpoint on the same origin | Generally yes, subject to normal authentication and application rules | Fetch a page or endpoint directly |
| A cross-origin API or page with suitable CORS headers | Yes, if the server permits the requesting origin | Fetch and inspect the response |
| A cross-origin site that does not grant CORS access | No, not from ordinary page JavaScript | Use a source intended for browser access, or consider a server-mediated design where permitted |
Fetch uses CORS mode by default for cross-origin requests. The target server controls whether the browser exposes the response to your script. Some requests trigger a preflight request before the browser sends the main request. Setting a different client option does not grant permission. See the MDN Fetch API guide.
2. Fetch and parse HTML with vanilla JavaScript
The following example fetches an HTML page from the same origin, checks for an HTTP error, parses the HTML string, and extracts article titles and links. Save it as an HTML file next to a page you control, or adapt the URL to an accessible same-origin page. Opening a local file directly can behave differently from serving the page over HTTP; use your normal local development server if requests fail.
<!doctype html>
<meta charset="utf-8">
<title>Read article titles</title>
<pre id="output">Loading…</pre>
<script>
async function scrapeArticles() {
const output = document.querySelector("#output");
try {
const response = await fetch("/articles");
if (!response.ok) {
throw new Error(`HTTP ${response.status} ${response.statusText}`);
}
const html = await response.text();
const doc = new DOMParser().parseFromString(html, "text/html");
const articles = [...doc.querySelectorAll("article")].map((article) => ({
title: article.querySelector("h2")?.textContent.trim() ?? "",
href: article.querySelector("a[href]")?.getAttribute("href") ?? ""
}));
output.textContent = JSON.stringify(articles, null, 2);
} catch (error) {
output.textContent = `Could not read articles: ${error.message}`;
}
}
scrapeArticles();
</script>
Replace /articles and the selectors with the path and markup you actually control. The example uses textContent to display the result rather than injecting scraped strings as HTML.
Step by step
- Choose an accessible source. Start with a same-origin page or an API/page that documents browser access.
- Make the request.
fetch()returns a promise that resolves to aResponse; it does not synchronously return the page. - Check the status. A 404 or 500 response usually still fulfills the fetch promise. Check
response.okorresponse.status. - Read the body. Use
response.text()for HTML andresponse.json()for JSON. These are asynchronous operations. - Parse and select. For accessible HTML, pass the string to
DOMParser, then select just the fields you need. - Handle failures visibly. Catch network-level errors and show a useful message; do not treat a failed request as an empty result.
3. Prefer JSON when a suitable endpoint exists
If the source offers a documented JSON endpoint, consume that data instead of scraping presentation markup. JSON avoids selectors tied to page layout and is already structured. The request still needs to be same-origin or permitted by CORS.
async function loadProducts() {
const response = await fetch("/api/products");
if (!response.ok) {
throw new Error(`Product request failed: HTTP ${response.status}`);
}
const products = await response.json();
return products.map(({ id, name, price }) => ({ id, name, price }));
}
loadProducts()
.then((products) => console.table(products))
.catch((error) => console.error(error));
For a cross-origin endpoint, the API’s response must permit your page’s origin. An endpoint working when opened in a browser tab does not by itself prove that JavaScript on your origin may read it.
4. Configure requests without expecting a CORS bypass
Fetch accepts a URL and an options object. Common options include method, headers, body, mode, credentials, and signal. Use only what the endpoint requires.

const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 10_000);
try {
const response = await fetch("/api/catalog", {
method: "GET",
headers: { Accept: "application/json" },
credentials: "same-origin",
signal: controller.signal
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const data = await response.json();
console.log(data);
} finally {
clearTimeout(timeout);
}
mode: "cors" is the normal default for Fetch and uses the CORS mechanism for cross-origin requests. mode: "same-origin" expressly disallows cross-origin requests. mode: "no-cors" is not a scraping workaround: it produces an opaque response, so JavaScript cannot inspect its status, headers, or body.
Credentials are a separate setting. Fetch defaults to sending credentials only for same-origin requests. Cross-origin credentialed access requires agreement from the server, including an explicit allowed origin rather than *. Credentials and cross-origin requests also have security implications, including CSRF risk. Do not add cookies or authorization headers casually; follow the API’s documented authentication and security model.
5. Parse HTML carefully
DOMParser turns HTML text you already have into a document that can be queried. It does not fetch a URL and does not change browser permissions. Treat scraped content as untrusted input: prefer text extraction, validate URLs before using them, and avoid copying untrusted markup into your live page.
function extractLinks(html, baseUrl) {
const doc = new DOMParser().parseFromString(html, "text/html");
return [...doc.querySelectorAll("a[href]")].map((anchor) => {
const rawHref = anchor.getAttribute("href");
let href = "";
try {
href = new URL(rawHref, baseUrl).href;
} catch {
// Ignore invalid URL values.
}
return { text: anchor.textContent.trim(), href };
}).filter((link) => link.href);
}
Selectors should match the actual markup. Optional chaining and fallback values help when a field is missing. If the page changes its HTML structure, selectors may need maintenance. If content is rendered later by scripts, the fetched HTML may contain only the initial document rather than the final visible page; browser fetch does not run the target page as a full browsing session.
6. When the target blocks browser access
If a cross-origin fetch fails due to CORS, the useful fix is architectural: choose an endpoint that supports browser access, or make a request from a server you control when the target’s access rules and applicable terms permit it. A server relay moves the request out of the browser’s same-origin boundary; it does not automatically make every request authorized or guarantee that the target will provide the desired content. Protect any relay from becoming an open proxy, and keep secrets on the server rather than shipping them in browser JavaScript.
Browser extensions and proxies also change the architecture and may raise security, privacy, terms-of-service, or legal questions. Check the rules that apply to the particular source. This guide does not establish the terms or permissions for any specific website.
7. Troubleshooting common failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Console reports a CORS error | The responding server did not grant readable access to your page’s origin, or a preflight did not succeed. | Use an endpoint that supports your origin, ask the API owner to configure CORS, or move the request to an appropriate server-side design. |
| Fetch returns a response but the page reports failure | HTTP errors such as 404 do not necessarily reject the promise. | Check response.ok and inspect the status before reading the result as success. |
Response body is unavailable with no-cors |
The response is opaque by design. | Remove the attempted workaround and use a source that grants access. |
| JSON parsing throws | The endpoint may have returned HTML, an empty body, or malformed JSON. | Check the status and response content type; inspect response.text() while debugging before assuming JSON. |
| Selectors return no elements | The markup differs, the selector is wrong, or desired content is inserted after the fetched HTML is generated. | Inspect the returned HTML and verify selectors against it. Look for a documented data endpoint when available. |
| Request appears to hang | The request is slow, stalled, or waiting on the network. | Use an AbortController timeout, show a loading state, and provide a retry path for transient failures. |
| Credentials are missing or rejected | Fetch defaults to same-origin credentials; cross-origin credentials need server agreement and may be restricted. | Follow the endpoint’s authentication instructions and avoid exposing sensitive credentials in public client code. |
8. Performance, reliability, and cost
For a browser-only workflow, the main costs are request time, response size, parsing work, and the user’s network and device resources. Fetch only the endpoint or page you need, extract only the fields you need, and avoid repeatedly downloading a large document when a structured endpoint is available. There is no universal performance figure: it depends on the source, network, response, and client device.
Make the flow reliable by checking status codes, distinguishing network failures from HTTP failures, timing out requests that take too long, and showing whether no matching records were found. Retry only when appropriate, with a limit; repeating a request indefinitely can waste bandwidth and burden the source. Respect the source’s documented usage expectations. Client-side code runs on the visitor’s device, so it is a poor place to keep secrets or centralize a dependable scheduled collection job.
A browser request does not have an API bill simply because it uses Fetch, but the source may impose its own limits or terms, and requests consume user bandwidth and device resources. A server-mediated design adds server operation and maintenance costs. Assess those tradeoffs for your application rather than assuming that moving a request server-side removes all constraints.
9. Or skip the browser setup
If what you need is a rendered screenshot rather than structured text, ScreenshotNeo provides a website screenshot API and MCP server for developers. Its one-call API returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Those points make it useful when the deliverable is a clean page capture, though a screenshot is not a substitute for structured page data.
Sign up for 1,000 free screenshots a month with no card.
10. FAQ
Can I scrape any website from a browser script?
No. The browser can read same-origin responses and cross-origin responses the server permits through CORS. A page being publicly viewable does not mean another origin’s JavaScript can read it.
Does DOMParser fetch a remote page?
No. It parses a string of HTML that your code has already obtained.
Can I use vanilla JavaScript without a library?
Yes. Fetch, promises, and DOMParser are browser APIs; a library is not required for the basic request and parsing flow.
Should I scrape HTML or use JSON?
Use a documented JSON endpoint when it provides the fields you need. Parse HTML when you have legitimate access to the document and need information represented in its markup.


