ScreenshotNeo

BlogHow-to

How to Use js-crawler to Crawl Websites

Install js-crawler, crawl links from a starting URL, control scope and request rate, and handle results and failures in Node.js.

By the ScreenshotNeo team4 October 20268 min read

js-crawler is a Node.js package for crawling links over HTTP and HTTPS. Install it with npm install js-crawler, create a crawler, and call crawl with a starting URL. Set depth to control how many links outward it follows, and use shouldCrawl to keep requests within your intended scope. Its callbacks expose the URL, response content, and status for each successful page, plus callbacks for failures and crawl completion. The package documentation is in the js-crawler repository README.

1. Install js-crawler

Use a Node.js project and install the package from npm:

mkdir js-crawler-demo
cd js-crawler-demo
npm init -y
npm install js-crawler

The documented API uses CommonJS and imports the package’s default export. Save the following as crawl.js:

const Crawler = require('js-crawler').default;

const crawler = new Crawler().configure({ depth: 2 });

crawler.crawl('https://example.com', function onSuccess(page) {
  console.log(`${page.status}\t${page.url}`);
});

Run it with node crawl.js. Replace https://example.com with a site you are allowed to crawl. With depth 2, the crawl starts at the supplied page and follows links outward up to the configured depth. The configure call is optional; omitting it uses the package defaults.

2. Collect successful pages, failures, and completion

The success callback runs once for each page that was accessed. The page object includes url, content (usually the HTML body), and HTTP status. The README also describes response-related fields such as error, response, and body, along with referer, which identifies the page that linked to the current URL.

For a crawl you want to process later, collect results and handle failures separately. The options-based form names all three callbacks:

const Crawler = require('js-crawler').default;

const pages = [];
const failures = [];

const crawler = new Crawler().configure({
  depth: 2,
  maxRequestsPerSecond: 2,
  maxConcurrentRequests: 2
});

crawler.crawl({
  url: 'https://example.com',
  success(page) {
    pages.push({
      url: page.url,
      status: page.status,
      html: page.content
    });
    console.log(`Fetched ${page.status}: ${page.url}`);
  },
  failure(page) {
    failures.push({
      url: page.url,
      referer: page.referer,
      status: page.status,
      error: page.error
    });
    console.error(`Could not access ${page.url}; status=${page.status}`);
  },
  finished(crawledUrls) {
    console.log(`Finished. Crawled URL count: ${crawledUrls.length}`);
    console.log(`Successful pages: ${pages.length}; failures: ${failures.length}`);
  }
});

A failed request may not have an HTTP status at all, so treat page.status as potentially undefined. A network error can occur before the server sends an HTTP response. The completion callback receives the crawled URL collection; keep success and failure records separately if you need outcome details.

The package also supports positional callbacks: crawl(url, success, failure, finished). Use the options object when named callbacks make control flow clearer.

3. Control crawl depth, URL scope, and request load

The documented options let you choose what enters the crawl, which pages contribute links, and how quickly requests may be issued.

Option Documented default What it controls
depth 2 How many link levels outward from the starting page are crawled.
ignoreRelative false Whether relative URLs are skipped. When false, relative links are not ignored.
userAgent crawler/js-crawler The User-Agent string sent in requests.
maxRequestsPerSecond 100 Upper bound on requests issued per second; can also be fractional.
maxConcurrentRequests 10 Maximum number of active requests at once.
shouldCrawl(url) Always true Whether a candidate URL should be requested.
shouldCrawlLinksFrom(url) Always true Whether links found on a fetched page should be added to the queue.

Restrict requests to the starting host and selected paths

Use an explicit URL parser instead of loose substring checks. This example allows only the starting hostname and excludes paths outside /docs/. It also restricts link harvesting to that host:

const Crawler = require('js-crawler').default;

const startUrl = 'https://example.com/docs/';
const start = new URL(startUrl);

function inScope(value) {
  try {
    const candidate = new URL(value, startUrl);
    return candidate.protocol.startsWith('http')
      && candidate.hostname === start.hostname
      && (candidate.pathname === '/docs' || candidate.pathname.startsWith('/docs/'));
  } catch {
    return false;
  }
}

const crawler = new Crawler().configure({
  depth: 3,
  ignoreRelative: false,
  userAgent: 'ExampleResearchCrawler/1.0',
  maxRequestsPerSecond: 2,
  maxConcurrentRequests: 2,
  shouldCrawl: inScope,
  shouldCrawlLinksFrom: inScope
});

crawler.crawl({
  url: startUrl,
  success(page) {
    console.log(page.status, page.url);
  },
  failure(page) {
    console.error('Failed:', page.url, page.status, page.error);
  },
  finished(urls) {
    console.log(`Visited ${urls.length} URLs`);
  }
});

URL scope needs deliberate choices. Hostname comparison excludes sibling subdomains; add them explicitly if they belong in the crawl. Query strings can create many URL variants, so decide whether parameters should be allowed or normalized. The documented options do not describe a URL canonicalization setting; do not assume the crawler merges URLs with different query strings, fragments, or trailing slashes.

Separate request rate from concurrency

maxRequestsPerSecond caps how many requests may be issued per second. maxConcurrentRequests caps how many requests are active simultaneously. They solve different problems and can be set together. A limit of two requests per second is an upper bound, not a guarantee that the crawler will reach that rate; network and server response times affect actual throughput. Start conservatively, especially for a site you do not operate.

Technical rate limits do not establish permission to crawl. Follow the site’s applicable terms, access rules, and operational guidance. Stop or reduce traffic if the site signals that requests should cease.

4. Understand what js-crawler fetches

js-crawler is documented as an HTTP/HTTPS crawler that returns page response content, usually HTML. The documentation does not establish that it runs a browser or executes page JavaScript. If a site inserts links or content only after client-side scripts run, do not assume those links or rendered values will be present in page.content. Confirm that the information you need exists in the HTTP response, or use a browser-based capture approach for rendered output.

This distinction matters for single-page applications, consent overlays, and pages that require browser interaction. The package’s documented controls cover request scope, depth, callbacks, and request load; they do not describe browser viewport, clicking, or DOM-rendered screenshot options.

5. Reuse a crawler instance safely

A crawler instance remembers URLs it has already crawled and does not crawl them again by default. If you intentionally want to crawl the same URLs again with the same instance, call forgetCrawled() before starting another crawl, or create a fresh instance:

const Crawler = require('js-crawler').default;

const crawler = new Crawler().configure({ depth: 1 });

function runCrawl() {
  return new Promise((resolve) => {
    crawler.crawl({
      url: 'https://example.com',
      success(page) {
        console.log(page.url);
      },
      finished(urls) {
        resolve(urls);
      }
    });
  });
}

async function main() {
  await runCrawl();
  crawler.forgetCrawled();
  await runCrawl();
}

main().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

The promise wrapper waits for the documented finished callback. For repeated independent jobs, a new crawler instance can make the URL memory boundary easier to reason about.

6. Troubleshooting

Symptom Likely cause What to do
No pages beyond the start URL The page has no links in the fetched HTML, depth is too low, or link harvesting is filtered. Inspect page.content, raise depth carefully, and check shouldCrawlLinksFrom.
Expected links are missing Links may be relative while ignoreRelative is enabled, or they may be added by browser JavaScript. Keep ignoreRelative: false when relative links should count. If links are client-rendered, use a browser-based method; js-crawler’s documentation does not claim script execution.
Unexpected off-site requests The default scope predicates allow all URLs. Set shouldCrawl to an explicit allowlist and constrain link harvesting with shouldCrawlLinksFrom.
Failure callback has no status The request failed before an HTTP response supplied a status code. Log URL, referer, and available error information; handle status as optional.
Crawl seems slow The configured request rate or concurrency is low, or the network and servers respond slowly. Measure the crawl on an authorized target, then adjust rate and concurrency incrementally. The configured rate is only a ceiling.
Second run skips pages The same instance remembers URLs it already crawled. Call forgetCrawled() or instantiate a new crawler.
Import is undefined The documented example expects the package’s default export under CommonJS. Use const Crawler = require('js-crawler').default as shown in the README. Check the installed package and module format if your project uses a different loader.

7. Performance, reliability, and cost

The package README documents defaults of up to 100 requests per second and 10 concurrent requests. Those are configuration defaults, not a measured performance benchmark or a promise of achieved throughput. A large crawl can still be slow because of network latency, server response time, and the number of URLs reached within the selected depth.

For reliability, keep the crawl bounded with a deliberate depth and URL predicate, record failures separately, and expect some failures to lack an HTTP status. The supplied documentation does not specify retry policies, durable queues, or a timeout configuration; do not assume those behaviors. For a long-running or repeatable production workflow, define how the surrounding application will persist progress and recover from interruption.

js-crawler is installed as software from npm; its README does not establish a package usage fee. Your operational costs may include the machine and network used to run it. Keep request volume appropriate to the target site.

Or skip the browser setup

If the goal is a screenshot of a rendered page rather than crawling links from raw HTTP responses, ScreenshotNeo offers a one-request website screenshot API and MCP server. Its API accepts a URL and returns PNG, JPEG, WebP, or PDF, with options for full-page capture, selected elements, waiting, and other capture settings. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server includes screenshot, page information, and PDF capture tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

FAQ

Does js-crawler execute JavaScript in a page?

The README documents HTTP/HTTPS fetching and response content, but does not establish browser JavaScript execution. Treat client-rendered content as unavailable unless it appears in the fetched response.

What does depth 2 mean?

It limits traversal to two link levels outward from the starting page, as described by the package README.

Can I crawl the same URL twice with one instance?

Yes. Clear the instance’s remembered URLs with forgetCrawled(), or create a new crawler instance.

Is the configured request rate the achieved speed?

No. It is an upper limit; actual throughput also depends on network and server response times.