PageCrawl.io API Setup in Node.js for Indian Developers
Create a PageCrawl monitor with Node.js, protect your token, and choose reliable polling, webhooks, or both.
To set up the PageCrawl.io API in Node.js, create an API token in PageCrawl, store it in a server-side environment variable, and send it as a Bearer token in the Authorization header. The shortest documented route to create a monitor is POST https://pagecrawl.io/api/track-simple. After creation, choose polling, webhooks, or a hybrid approach to receive changes.
PageCrawl is a website monitoring service: it checks pages and reports changes. If your task is to capture a page as an image or PDF, see the ScreenshotNeo option below; a screenshot API does not replace a monitoring service.
1. Create and protect an API token
- In PageCrawl, open Settings > API > API Tokens and create a token.
- Copy it when it is shown. The help material says the token is not shown again.
- Store it in a server-side secret store or environment variable. Do not put it in browser JavaScript, a URL, source control, or logs.
PageCrawl documents Bearer authentication in the Authorization header. OAuth access tokens are also supported. A query-string api_token is mentioned for quick browser tests, but the documented supported form is the Bearer header; do not use a query parameter in a production integration because URLs can be recorded in logs and telemetry.
export PAGECRAWL_API_TOKEN='replace-with-your-token'
For deployment, configure the variable through your hosting provider’s secret settings. Avoid committing a real .env file. If a token is exposed, revoke or rotate it in PageCrawl and update the secret wherever the application runs.
2. Create your first monitor with Node.js
This example uses the built-in fetch available in current Node.js releases. Save it as create-monitor.mjs and run it in an environment where PAGECRAWL_API_TOKEN is set.
const token = process.env.PAGECRAWL_API_TOKEN;
if (!token) {
throw new Error('Set PAGECRAWL_API_TOKEN before running this script');
}
const response = await fetch('https://pagecrawl.io/api/track-simple', {
method: 'POST',
headers: {
Authorization: `Bearer ${token}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
url: 'https://example.com/pricing',
tracking_mode: 'fullpage',
}),
});
if (!response.ok) {
const detail = await response.text();
throw new Error(`PageCrawl returned HTTP ${response.status}: ${detail}`);
}
const page = await response.json();
console.log(`Monitoring: ${page.name} (${page.id})`);
The documented quick start returns JSON containing the created monitor’s name and ID. A new monitor is documented as returning HTTP 201; checking response.ok accepts any successful 2xx response and avoids assuming that every successful response has exactly the same status.
Request shape and tracking modes
The example sends a URL and tracking_mode. The official guide describes these modes:
| Mode | Use it for |
|---|---|
fullpage |
All visible text; documented as the default. |
content_only |
Page content with navigation, header, and footer content stripped. |
reader |
Reader-mode content extraction. |
price |
Price detection. |
specific_text |
Text selected using a selector. |
specific_number |
A number selected using a selector. |
feed |
Repeated listings. |
seo |
SEO fields such as title, meta, canonical, robots, and Open Graph data. |
Exact accepted values and selector request fields can depend on the current API schema. Check PageCrawl’s API reference/OpenAPI specification before relying on optional fields; its developer guide describes the reference as generated from that specification. Avoid guessing field names from a mode’s label.
cURL equivalent
curl --request POST 'https://pagecrawl.io/api/track-simple' \
--header "Authorization: Bearer ${PAGECRAWL_API_TOKEN}" \
--header 'Content-Type: application/json' \
--data '{"url":"https://example.com/pricing","tracking_mode":"fullpage"}'
Python equivalent
import os
import requests
token = os.environ["PAGECRAWL_API_TOKEN"]
response = requests.post(
"https://pagecrawl.io/api/track-simple",
headers={
"Authorization": f"Bearer {token}",
"Content-Type": "application/json",
},
json={"url": "https://example.com/pricing", "tracking_mode": "fullpage"},
timeout=30,
)
response.raise_for_status()
page = response.json()
print(f"Monitoring: {page['name']} ({page['id']})")
3. Choose how your application receives changes
| Pattern | Choose it when | Operational tradeoff |
|---|---|---|
| Polling | A dashboard or report can refresh periodically. | Simple to operate, but request volume grows with polling frequency and pagination. |
| Webhooks | Changes should trigger near-real-time work. | Requires a reachable receiver and signature verification. |
| Hybrid | Fast updates matter and missed events must be recovered. | Use webhooks for prompt updates and slower polling to reconcile stored state. |
Polling PageCrawl pages
The documented Node.js polling pattern requests GET /api/pages?simple=1, follows links.next for pagination, reads latest.contents, and maps individual element values using stable element_id values. The following illustrates a guarded page fetch; adapt pagination and response-field handling to the current API reference.
async function getPages() {
const response = await fetch('https://pagecrawl.io/api/pages?simple=1', {
headers: { Authorization: `Bearer ${process.env.PAGECRAWL_API_TOKEN}` },
});
if (response.status === 429) {
const retryAfter = response.headers.get('retry-after');
throw new Error(`Rate limited; retry after ${retryAfter ?? 'a backoff interval'}`);
}
if (!response.ok) {
throw new Error(`PageCrawl returned HTTP ${response.status}`);
}
return response.json();
}
const result = await getPages();
console.log(result);
For a production poller, follow every next-page link until none remains, persist a checkpoint, and schedule the next poll with enough spacing for the total number of pages. Do not treat a single page of results as the complete monitor list.
Rate limits and backoff
PageCrawl’s reference lists 60 requests per minute for Free accounts and 300 requests per minute for paid accounts. These are service limits, not performance benchmarks, and can change. If a request receives HTTP 429, honor Retry-After when supplied. A simple delay helper is:
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
async function fetchWith429Retry(url, options, maxRetries = 4) {
for (let attempt = 0; ; attempt += 1) {
const response = await fetch(url, options);
if (response.status !== 429 || attempt >= maxRetries) return response;
const retryAfter = Number(response.headers.get('retry-after'));
const delayMs = Number.isFinite(retryAfter) && retryAfter > 0
? retryAfter * 1000
: Math.min(30_000, 1_000 * 2 ** attempt);
await sleep(delayMs);
}
}
For multiple workers, coordinate request pacing centrally rather than letting each worker independently consume the full account limit. Retrying every error immediately can worsen an outage or rate limit.
4. Receive and verify webhooks
PageCrawl’s Node.js reference verifies X-PageCrawl-Signature and X-PageCrawl-Timestamp with HMAC-SHA256 over the timestamp, a period, and the exact raw request body. Verify before trusting or processing the JSON. In Express, capture the raw bytes for this route before JSON parsing; parsing and re-serializing JSON can change whitespace or byte representation.
import express from 'express';
import crypto from 'node:crypto';
const app = express();
const webhookSecret = process.env.PAGECRAWL_WEBHOOK_SECRET;
if (!webhookSecret) throw new Error('Set PAGECRAWL_WEBHOOK_SECRET');
app.post('/webhooks/pagecrawl', express.raw({ type: 'application/json' }), (req, res) => {
const timestamp = req.header('X-PageCrawl-Timestamp');
const suppliedHex = req.header('X-PageCrawl-Signature');
if (!timestamp || !suppliedHex || !Buffer.isBuffer(req.body)) {
return res.status(400).send('Missing signature headers or raw body');
}
const timestampSeconds = Number(timestamp);
const nowSeconds = Math.floor(Date.now() / 1000);
const allowedAgeSeconds = 5 * 60;
if (!Number.isFinite(timestampSeconds) || Math.abs(nowSeconds - timestampSeconds) > allowedAgeSeconds) {
return res.status(400).send('Stale or invalid timestamp');
}
const expectedHex = crypto
.createHmac('sha256', webhookSecret)
.update(`${timestamp}.`)
.update(req.body)
.digest('hex');
const supplied = Buffer.from(suppliedHex, 'hex');
const expected = Buffer.from(expectedHex, 'hex');
if (supplied.length !== expected.length || !crypto.timingSafeEqual(supplied, expected)) {
return res.status(401).send('Invalid signature');
}
const event = JSON.parse(req.body.toString('utf8'));
// Enqueue event for processing; acknowledge promptly after durable enqueue.
console.log('Verified PageCrawl webhook event', event);
return res.sendStatus(200);
});
app.listen(3000);
Use the exact signature encoding and timestamp units specified by PageCrawl’s current webhook guide. The example follows the described hex digest and timestamp-seconds pattern; confirm the current reference before deployment. Configure this route so global JSON middleware does not consume the body first. Return a 2xx acknowledgment promptly after validation and durable queueing; PageCrawl documents retries with backoff when delivery fails, and treats 2xx as acknowledgment. Make event handling idempotent because retries can deliver an event again.
5. Keep the integration reliable
- Separate creation from routine reads. Create a monitor once and persist its returned ID. Do not accidentally create a new monitor every time a dashboard loads.
- Paginate completely. Follow
links.nextand store stable IDs rather than relying on list position. - Make webhook handlers quick. Validate, enqueue, and acknowledge; perform slow downstream work asynchronously.
- Reconcile state. If missing an update would matter, combine webhooks with a slower poll to recover from receiver downtime.
- Watch plan capacity. PageCrawl says monitoring checks pause when plan limits are exceeded, so a successful API setup alone does not guarantee ongoing checks.
- Handle configuration drift. Keep token rotation and webhook-secret changes in your deployment runbook.
6. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| HTTP 401 or 403 | Missing, invalid, revoked, or incorrectly formatted token. | Check the server-side secret and send Authorization: Bearer TOKEN. Rotate an exposed token. |
| HTTP 422 | Invalid request data or a missing/incorrect field. | Read the field-level validation response and compare the payload to the current API schema. |
| HTTP 429 | Request rate exceeded. | Wait for Retry-After, reduce polling or concurrency, and apply bounded backoff. |
| Webhook signature always fails | JSON middleware changed or consumed the body, wrong secret, or incorrect signed message construction. | Capture raw bytes first, verify timestamp + period + raw body, and check the configured secret and signature encoding. |
| Webhook rejected as stale | Server clock skew or delayed delivery. | Synchronize the host clock and use the timestamp tolerance documented by PageCrawl. |
| Monitor exists but no continuing checks | Plan limits may have been reached, pausing checks. | Check current account usage and plan capacity in PageCrawl. |
| Only some pages appear in a poll | The client did not follow pagination. | Continue through links.next until there is no next link. |
Node reports fetch is not defined |
The runtime lacks the built-in fetch used by this example. | Use a current Node.js runtime with built-in fetch, or use an HTTP client already approved for your project. |
7. Plans, geography, and cost considerations
PageCrawl says API access and webhooks are available on every plan, including Free. Its published Free plan lists up to six pages, 220 checks, and a 60-minute check frequency; paid tiers have higher limits and more frequent checks. The pricing page says prices exclude VAT. The reviewed material does not establish India-specific GST treatment, INR billing, or acceptance of every Indian-issued card, so verify the current checkout and pricing details before budgeting. Limits, prices, and plan terms can change.
Estimate usage from the number of monitored pages, check frequency, polling traffic, and any reconciliation requests. A frequent poll can consume API request capacity even when the monitoring checks themselves are within plan limits. Keep the monitor cadence and the application polling cadence as separate settings.
Or skip the browser setup
PageCrawl monitors pages for changes. If what you need is a screenshot or PDF capture, ScreenshotNeo is a website screenshot API and MCP server for developers. It can accept cookie banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify page verdict and billing status in headers. AI agents can use its MCP server tools for screenshots, page information, and PDFs.
Here is the one-call Node.js request. See the ScreenshotNeo API documentation for the available options.
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo HTTP ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', res));
The response body is the image; for a production integration, check the response status and save the returned bytes. ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Can I call PageCrawl directly from a browser app?
Keep the API token on a server. A browser bundle exposes credentials to visitors, so route requests through your backend.
Should I poll or use webhooks?
Poll when periodic refresh is sufficient. Use webhooks for event-driven updates, and add a slower reconciliation poll when recovering missed events matters.
Does a Free PageCrawl account include API access?
PageCrawl says API and webhook access are available on Free; page, check, and frequency limits still apply and can change.
Is ScreenshotNeo a PageCrawl replacement?
No. PageCrawl is for recurring website monitoring. ScreenshotNeo captures images or PDFs on request.


