How to Wait for a Web Page to Finish Loading Before PageCrawl.io Captures It
Configure PageCrawl.io to wait for the content you need before capture, with practical steps for dynamic pages, lazy loading, custom JavaScript, and troubleshooting.
To make PageCrawl.io wait before capturing a page, edit the monitored page, add pre-capture actions, and choose a readiness condition tied to the content you need. Use Wait for text or Wait for element when dynamic content has a stable marker. For lazy-loaded content, use Scroll to bottom followed by a short Wait. Use custom JavaScript when the page needs multiple interactions or a condition that the built-in actions cannot express.
PageCrawl.io runs configured actions in order before taking the snapshot. “Finished loading” is not one universal browser state: the document can finish loading or network traffic can become quiet before the application renders the results you are tracking. Choose a signal that represents those results.
Configure a wait in PageCrawl.io
- Open the monitored page and click Edit.
- Find Actions, choose Add Action, and select the action to run before capture. Actions are configured per tracked element and run from top to bottom.
- For content that appears asynchronously, add Wait for text and enter a stable phrase that appears when the content is ready. Alternatively, add Wait for element and provide a CSS or XPath selector for the content.
- For content that loads as a visitor scrolls, add Scroll to bottom, then add Wait. PageCrawl.io’s guide suggests 2–3 seconds as an example; adjust for the page and verify the content appears.
- Save the action sequence and inspect a capture to confirm the marker and action order match the page’s behavior.
Wait for text and Wait for text to disappear can wait up to 15 seconds, according to PageCrawl.io’s documentation. A wait limit is a ceiling, not a guarantee that the page will finish loading.
Choose the readiness condition that matches the page
| Page behavior | Condition to try | Why |
|---|---|---|
| Results arrive after an API request | Wait for a unique result heading, status phrase, or result element | It targets the content being monitored instead of guessing how long the request will take. |
| A loading message is replaced | Wait for the loading text to disappear, then wait for the result marker if available | The disappearance confirms one transition; the result marker confirms the desired content. |
| More content loads while scrolling | Scroll to bottom, then Wait briefly | This gives scroll-triggered lazy loading a chance to run. |
| A button must be clicked or several steps must run | Custom JavaScript with bounded interactions and waits | It can sequence actions that a single built-in wait cannot express. |
| A browser script is navigating to a page | Wait for a navigation milestone, then assert the expected content | Navigation completion and application readiness are separate conditions. |
Prefer a specific marker that is unlikely to appear in a loading shell, stale content, or unrelated page region. If the page can show an empty state as a valid result, choose a marker that distinguishes “results finished” from “results still loading.”
Wait for dynamic content with a built-in action
For a page that eventually displays a known phrase, configure Wait for text with that phrase. For a result container or other stable node, use Wait for element with its selector. A selector should identify the content whose readiness matters, not merely a general page wrapper that exists before its contents render.
For a loading indicator, Wait for text to disappear can help, but disappearance alone may be ambiguous: the indicator could be removed because of an error or because the application stopped rendering. When possible, follow it with a wait for the successful result or empty-state marker.
Handle lazy-loaded content
- Add Scroll to bottom so the page triggers content that loads on scroll.
- Add Wait after the scroll action. A few seconds is a starting point, not a universal timing value.
- Inspect a capture and confirm the expected lower-page content is present. If it is not, check whether the site uses a “Load more” button or requires multiple scrolls.
Some pages load content incrementally only when the visitor scrolls in steps. In that case, one scroll-to-bottom action may not trigger every section. Use custom JavaScript for a bounded sequence of scrolls or clicks if the built-in actions are insufficient, and keep the total time inside the action and check limits.
Use custom JavaScript for multi-step waits
PageCrawl.io runs a custom JavaScript action in the page’s browser context as part of the pre-capture action sequence. The return value is ignored, so use it for side effects such as clicking a load-more button, changing page state, or sequencing interactions. Its documentation gives this example for clicking a load-more control twice and allowing rows to render:
(async () => {
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
document.querySelector('#load-more')?.click();
await sleep(800);
document.querySelector('#load-more')?.click();
await sleep(800);
})();
The 800 ms pauses are illustrative, not a universal delay recommendation. When the page exposes a reliable condition, checking for that condition is generally more robust than relying only on a fixed pause. Keep custom action logic bounded: PageCrawl.io documents a 30-second safety timeout for custom JavaScript actions and JavaScript tracked elements.
For a page with a known result selector, a bounded polling pattern can make the condition explicit:
(async () => {
const selector = '.results-ready';
const deadline = Date.now() + 10000;
while (Date.now() < deadline) {
if (document.querySelector(selector)) return;
await new Promise(resolve => setTimeout(resolve, 250));
}
})();
This example only waits for an element to exist. Adapt the selector and condition to the page, and remember that the action’s return value is ignored. If the selector exists before the results are populated, test for a more specific condition, such as non-empty text or a child result node. The PageCrawl.io action timeout still applies.
Browser navigation waits are not application readiness
In browser automation, navigation lifecycle options describe document events: Playwright distinguishes commit, domcontentloaded, load, and networkidle. Its documentation defines networkidle as no network connections for at least 500 ms and discourages using it as a test readiness condition; it recommends assertions about the page instead. For capture or extraction, wait for the content the task actually needs.
Puppeteer provides waitForFunction for waiting until a page-context function becomes truthy, as well as waitForNetworkIdle. A predicate that checks the desired content is usually more meaningful than treating network quiet as proof that application work is complete.
Timeouts and limits
| Limit | Documented value | How to use it |
|---|---|---|
| Wait for text / wait for text to disappear | Up to 15 seconds | Choose a marker likely to appear within the action window; a longer overall check limit does not necessarily extend this action. |
| Custom JavaScript safety timeout | 30 seconds | Bound loops and delays so the script can complete within the action limit. |
| Overall check timeout: Free | 45 seconds | The full check has a separate time ceiling. |
| Overall check timeout: Standard | 90 seconds | The full check has a separate time ceiling. |
| Overall check timeout: Enterprise and Ultimate | 180 seconds | The full check has a separate time ceiling. |
These are documented limits, not performance guarantees. Increasing a wait does not fix an inaccessible page, a broken selector, authentication requirements, or a site that never renders the expected state.
Troubleshooting PageCrawl.io capture waits
| Symptom | Likely cause | What to check |
|---|---|---|
| The capture times out | The site responded slowly, temporarily failed, or the configured condition never occurred before a time limit. | Inspect the captured page and error. Confirm the marker is correct and appears on a normal visit; reduce unnecessary actions and delays. Do not assume a longer wait will fix access or rendering problems. |
| “Selector not found” | The page changed, the selector is invalid, or the target is inside a different page context. | Inspect the current page structure, update the CSS or XPath selector, and make sure it identifies the content rather than a transient wrapper. |
| The wait completes but the capture is still incomplete | The condition is too broad, appears before the data is populated, or only one lazy-load stage was triggered. | Use a result-specific marker or condition. For scroll-driven pages, verify the scroll action triggers the content; add a bounded interaction sequence if needed. |
| 401 response | The page requires authentication. | Check whether the monitor has an authorized way to access the page. Waiting does not provide a login or grant account permissions. |
| 403 response | The site refused access. | Inspect the access error and site requirements. A longer wait does not resolve a refusal. |
| Custom JavaScript reaches its limit | The action exceeds its 30-second safety timeout or gets stuck polling. | Set an explicit deadline, use shorter bounded steps, and remove waits that do not contribute to the target condition. |
Start diagnosis by inspecting the captured page and the reported error. A timeout, a missing selector, an authentication response, and an access refusal are different problems and need different fixes.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a screenshot or PDF; see the API documentation for options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
Performance and reliability notes
- Prefer conditions over guessed delays. A content-specific wait can finish as soon as the marker appears; a fixed delay always consumes its configured time and can still be too short.
- Keep action sequences small. Every interaction and wait uses part of the action and overall check time budget.
- Use stable markers. A selector tied to generated classes or a phrase that changes often can make monitoring fragile. Review it when the page changes.
- Do not equate network quiet with complete rendering. Pages can render after requests finish, and some keep background requests open.
- Expect temporary site failures. A longer wait may help a slow response, but it cannot repair a failed request, permission problem, or access block.
PageCrawl.io’s documented timing limits describe maximums, not measured capture speed. No independent performance measurements are established here.
Frequently asked questions
Does PageCrawl.io wait for all JavaScript to finish?
There is no single “all JavaScript finished” condition for a modern page. Configure a wait for the text, element, or state that matters to the tracked content.
Should I always use a 2–3 second wait after scrolling?
No. That is an example in PageCrawl.io’s guidance for lazy-loaded content. Confirm the page’s behavior and use the shortest delay that reliably allows the content to appear.
Can I use a custom JavaScript return value to tell PageCrawl.io the page is ready?
The custom action’s return value is ignored. Use it to perform bounded page interactions and waits before extraction.
Will increasing the timeout fix a 401 or 403?
No. Those responses point to authentication or access refusal, not merely slow rendering. Inspect the error and resolve the access requirement separately.


