Web Scraping with JavaScript and Selenium: A Practical Guide
Learn when Selenium is worth using, how to scrape JavaScript-rendered pages with Node.js, and how to wait for the exact content your script needs.
Selenium lets a JavaScript program control a real browser, so it can read content after page scripts render it and interact with controls when needed. Install the selenium-webdriver package, open the target page, wait for the specific element or state your extraction depends on, collect only the fields you need, and close the browser in a finally block.
A completed navigation does not guarantee that a JavaScript application has finished rendering. Use Selenium when browser rendering or interaction is necessary; if the data is already available through a permitted HTTP response or documented interface, a direct request may be simpler and lighter. [Selenium WebDriver](https://www.selenium.dev/documentation/webdriver/) [Waiting Strategies](https://www.selenium.dev/documentation/webdriver/waits/)
1. What Selenium does, and when to use it
Selenium WebDriver is a browser automation interface with language bindings and browser-specific implementations. A JavaScript script sends WebDriver commands; the browser loads and renders the site, and Selenium can inspect the resulting DOM and interact with it. It can run locally or connect to a remote Selenium server. [Selenium WebDriver](https://www.selenium.dev/documentation/webdriver/)
That browser fidelity has operating costs: starting and keeping browser processes consumes more resources than making direct HTTP requests, and browser automation brings more setup and maintenance. The research sources do not establish a numerical speed comparison, so evaluate the trade-off for your own target and workload.
| Choose | When it fits | Trade-off |
|---|---|---|
| Direct HTTP request | The needed data is present in a response you are allowed to access, or exposed through a documented interface. | Less browser machinery; does not itself run page JavaScript or reproduce interactive browser behavior. |
| Selenium | The information appears only after client-side rendering, or reaching it requires browser interactions such as opening a panel or changing a selection. | Can observe browser-rendered behavior, with browser startup, resource use, and synchronization to manage. |
Start with the smallest approach that reliably provides the required data. A screenshot or visible-page check may need a browser even when structured data collection does not.
2. Install Selenium and run a JavaScript scraper
The Selenium JavaScript API currently documents Node.js 22 or newer and installation with npm install selenium-webdriver. Runtime requirements can change, so check the [current JavaScript API reference](https://www.selenium.dev/selenium/docs/api/javascript/) when setting up a new project. Selenium Manager handles browser driver installation in the documented quick start. [JavaScript API](https://www.selenium.dev/selenium/docs/api/javascript/)
Set up a small project
mkdir selenium-scraper
cd selenium-scraper
npm init -y
npm install selenium-webdriver
Save the following as scrape.js. This example uses the public Selenium demonstration page, waits for the page heading, extracts its title and visible text, and always closes the driver. The locator is specific to that example page; replace it with a selector for the data you are authorized to collect.
const { Builder, Browser, By, until } = require('selenium-webdriver');
async function main() {
const driver = await new Builder().forBrowser(Browser.CHROME).build();
try {
await driver.get('https://www.selenium.dev/');
// Navigation completion and application readiness are different states.
const heading = await driver.wait(
until.elementLocated(By.css('h1')),
10000,
'The page heading did not appear within 10 seconds'
);
await driver.wait(until.elementIsVisible(heading), 5000);
const result = {
title: await driver.getTitle(),
heading: (await heading.getText()).trim(),
url: await driver.getCurrentUrl()
};
console.log(JSON.stringify(result, null, 2));
} finally {
await driver.quit();
}
}
main().catch((error) => {
console.error(error);
process.exitCode = 1;
});
Run it with node scrape.js. For a JavaScript-rendered target, replace the URL and locator with the target page and the actual result container or field. Confirm the selector in the current DOM, and wait for the state the next operation requires.
Extract multiple records
Once the result container is ready, read the relevant elements and normalize the values. For example, after adapting .result-card, .result-title, and .result-link to the target page:
const cards = await driver.findElements(By.css('.result-card'));
const results = [];
for (const card of cards) {
const title = await card.findElement(By.css('.result-title')).getText();
const href = await card.findElement(By.css('.result-link')).getAttribute('href');
results.push({ title: title.trim(), href });
}
console.log(JSON.stringify(results, null, 2));
Use stable selectors where the site provides them, and verify that the expected number of records or fields was actually collected. An empty array can mean the page has not rendered, the locator no longer matches, or the page legitimately has no results.
3. Wait for the page state you need
WebDriver navigation waits for a page-load state determined by its page-load strategy; the default is complete. That state concerns assets defined in the HTML. Page scripts may then add elements or change visibility, so the next Selenium command can run before the application is ready. [Waiting Strategies](https://www.selenium.dev/documentation/webdriver/waits/)
Use explicit waits for actual preconditions
Wait for the specific condition needed by the next action: an element located, an element visible, a result count reached, or a loading indicator gone. Selenium’s JavaScript API provides driver.wait and conditions such as until.elementLocated and until.elementIsVisible. [JavaScript WebDriver API](https://www.selenium.dev/selenium/docs/api/javascript/WebDriver.html)
// Wait for a result element to exist in the DOM.
const result = await driver.wait(
until.elementLocated(By.css('[data-testid="search-result"]')),
15000
);
// If the next operation needs to interact with it, also wait until visible.
await driver.wait(until.elementIsVisible(result), 5000);
// A custom condition can wait until a result list has enough items.
await driver.wait(async () => {
const items = await driver.findElements(By.css('.result-item'));
return items.length >= 5 ? items.length : false;
}, 15000, 'Fewer than five results appeared');
Set timeouts according to the operation and the page behavior you expect, then handle a timeout as an observable failure. Do not simply raise every timeout when a wait fails: first determine which condition was false and whether the selector is still correct.
Avoid fixed sleeps as the normal synchronization strategy
A fixed delay may be too short on a slow response and waste time on a fast one. Selenium recommends condition-based synchronization. It also warns against mixing implicit and explicit waits because the resulting total wait can be unpredictable. Prefer explicit waits for the relevant condition and leave the implicit wait at its default of zero unless you have a deliberate, consistent strategy. [Waiting Strategies](https://www.selenium.dev/documentation/webdriver/waits/)
4. Interact with dynamic pages
Some pages need an action before the data appears. Locate and interact with the control, then wait for the post-action state before extracting:
const search = await driver.wait(
until.elementLocated(By.css('input[name="q"]')),
10000
);
await search.sendKeys('Selenium WebDriver');
const submit = await driver.findElement(By.css('button[type="submit"]'));
await submit.click();
await driver.wait(
until.elementLocated(By.css('.search-results .result-item')),
15000,
'Search results did not appear'
);
For pagination, dropdowns, or “load more” controls, treat each action as a transition: perform the action, wait for a concrete indication that the page changed, then collect the next batch. If the page updates the same element in place, wait for its text, attribute, or record count to change instead of waiting for an element that was already present.
For authenticated content, use an authorized account and handle credentials as secrets. Do not place credentials in source control or log them. Whether a particular collection is permitted depends on the target’s rules and the circumstances; this guide does not determine that for a specific site.
5. Configure browser startup, timeouts, and remote execution
The Selenium JavaScript Builder selects a browser and accepts browser-specific options. The current API documents Chrome and Firefox options, browser selection, Selenium Manager, and remote server configuration. Browser-specific options vary, so consult the official [browser options documentation](https://www.selenium.dev/documentation/webdriver/drivers/options/) and [JavaScript API](https://www.selenium.dev/selenium/docs/api/javascript/).
For example, use an explicit Chrome option object when configuring browser startup, and keep options for the selected browser:
const { Builder, Browser } = require('selenium-webdriver');
const chrome = require('selenium-webdriver/chrome');
const options = new chrome.Options();
// Add only options appropriate for your environment here.
const driver = await new Builder()
.forBrowser(Browser.CHROME)
.setChromeOptions(options)
.build();
Set navigation, script, and element wait timeouts for their separate purposes. Navigation timeout limits page-load events; script timeout applies to scripts evaluated through WebDriver; explicit waits define application conditions. Avoid treating a longer navigation timeout as a solution to a missing client-rendered element. [Timeouts API](https://www.selenium.dev/selenium/docs/api/javascript/Timeouts.html)
When local browser execution becomes difficult to operate or you need remote browser sessions, Selenium Grid routes WebDriver commands to remote browser instances. It is an operational option, not a guarantee of a particular provider, price, or performance level. [Selenium Grid](https://www.selenium.dev/documentation/grid/)
6. Responsible access and collection boundaries
Before collecting data, check the target site’s published instructions, terms, and any permissions or obligations applicable to your use. MDN describes robots.txt as crawler instructions, typically at the site’s root; it is publicly accessible and does not secure the site. A robots file is not, by itself, a grant of permission or a complete answer about whether collection is allowed. [MDN: robots.txt](https://developer.mozilla.org/en-US/docs/Web/Security/Practical_implementation_guides/Robots_txt)
Keep request volume proportionate, avoid collecting information you do not need, and stop if the site signals that access is not allowed. The available sources do not establish jurisdiction-specific legal rules or determine whether any particular dataset may be collected.
7. Troubleshooting common Selenium scraping failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
NoSuchElementError |
The element has not rendered yet, the selector is wrong, or the page is different than expected. | Inspect the current DOM and URL; wait for the element’s presence; verify the locator against the loaded page. |
TimeoutError from a wait |
The awaited condition never became true within the timeout. | Check that the selector matches, the action succeeded, and the target page actually contains results. Increase the timeout only if the condition is right but legitimately takes longer. |
| Element is present but interaction fails | It is hidden, disabled, covered, or not yet ready for interaction. | Wait for visibility or the needed enabled state; verify that the correct frame or page section is active. |
| Navigation returns but extracted data is empty | Client-side rendering continued after navigation, or the page’s data path changed. | Wait on the result container or a meaningful state change; inspect the rendered DOM and locator. |
| Works sometimes, fails sometimes | A timing race, unstable selector, or varying page state. | Replace arbitrary sleeps with condition-based waits; verify a stable locator and the expected state before extraction. |
| Browser or driver fails to start | Browser installation, version, environment, or configuration issue. | Check that the selected browser is installed and supported in the environment, review Selenium’s setup output, and consult the official [troubleshooting guide](https://www.selenium.dev/documentation/webdriver/troubleshooting/). |
| Wait takes unexpectedly long | Implicit and explicit waits may be combined, or the wrong state is being awaited. | Use one deliberate synchronization strategy, identify the exact unmet condition, and avoid mixing implicit and explicit waits. |
| Script exits but browser remains open | The code did not reach driver cleanup after an error. | Put await driver.quit() in a finally block so cleanup runs on success and failure. |
8. Performance, reliability, and operating cost
- Use the browser selectively. A real browser is justified when page rendering or interaction is required. For data already present in an allowed response, direct HTTP may avoid browser startup and rendering work.
- Wait narrowly. Waiting for the exact result condition reduces needless delay while making timing failures easier to diagnose. Avoid fixed sleeps throughout a workflow.
- Extract only needed fields. Limit DOM reads and stored data to the task’s requirements.
- Always close sessions. A
finallycleanup path releases browser resources even if navigation or extraction throws. - Scale deliberately. Remote execution through Selenium Grid is available for routing scripts to remote browsers. Assess setup, capacity, failure handling, and provider cost for your own workload; the cited sources do not provide benchmarks or pricing.
- Make failures visible. Log the page URL, failed condition, and concise error context, while keeping credentials and sensitive page data out of logs.
9. Or skip the browser setup
If your task is to capture a website screenshot rather than extract arbitrary records, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns an image or PDF, with options for full-page or element capture, viewport and device settings, waits, custom CSS or JavaScript, caching, and other capture controls. See the ScreenshotNeo API documentation for parameters and examples.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));
- Cookie banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
- Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Responses include page-verdict and billing headers.
- An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
- The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.
10. Frequently asked questions
Can Selenium scrape a single-page application?
Yes. Selenium can control the browser while the application adds or updates page content. Wait for the relevant rendered state before reading it.
Does a successful driver.get() mean the page data is ready?
No. Navigation completion follows a page-load state; client-side scripts may continue changing the page afterward.
Should I use implicit or explicit waits?
For dynamic application content, explicit waits make the needed condition clear. Selenium warns not to combine implicit and explicit waits in the same session.
Can Selenium run in a remote browser?
Yes. Selenium supports remote execution through a Selenium server, and Grid routes WebDriver commands to remote browser instances.
Is robots.txt permission to scrape?
No. It is crawler guidance, not access control or blanket authorization. Check the site’s rules and applicable obligations for your situation.


