Waiting for page load
Wait for a rendered page to reach a load state, from its HTML being parsed to its network going idle, before it is read or before actions run.
A rendered page is read once it has loaded and settled, which suits most pages. When a page fills itself in later,
after its load event or after a run of API calls, tell the request what to wait for:
waitUntilnames the state a rendered page must reach before it is read, or before the first action runs.- The
waitForNetworkIdleaction waits for the network to go idle at any point in a browser actions sequence, such as after a click.
A wait that runs out does not fail the request by default. The page is read as it is, and the response says so.
waitUntil
curl https://request.usestring.ai/v1/fetch \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://www.example.com/dashboard",
"executeJS": true,
"waitUntil": "networkidle0"
}'waitUntil | The page is read once |
|---|---|
domcontentloaded | Its HTML has been parsed (the DOMContentLoaded event). |
load | Its load event has fired, which follows its images, stylesheets and iframes. |
networkidle0 | No request has been in flight for 500 ms. networkidle means the same. |
networkidle2 | At most 2 requests have been in flight for 500 ms. For pages that hold a long-poll or analytics request open. |
waitUntil requires a browser render: executeJS: true, actions or screenshot. Without one of them the request
returns a 400.
On a rendered fetch
With executeJS: true, the wait can make a page take longer to read but never makes it come back sooner. It gives up
after 10 seconds, and the page is read as it is.
With actions or screenshot
With actions or screenshot, waitUntil applies to loading url, before the first action runs:
- It replaces the usual settle after loading
url(see Timing). Withdomcontentloaded, the first action starts as soon as the HTML is parsed. - A wait for
loador for the network gives up 10 seconds after the HTML is parsed, and the actions then run on the page as it is. - Later navigations are not affected. A
navigateaction, or a click that loads a new page, gets the usual settle. To wait for the network after one of them, add awaitForNetworkIdleaction.
What counts as in flight
The network-idle states and waitForNetworkIdle count only document, XHR, fetch() and script requests, from the page
and from every iframe in it, cross-origin iframes included. Images, fonts, media, stylesheets,
navigator.sendBeacon() pings, EventSource streams and WebSockets never count.
The waitForNetworkIdle action
{ "type": "waitForNetworkIdle", "idleTime": 500, "concurrency": 0, "timeout": 10000, "onTimeout": "continue" }| Field | Type | Default | Description |
|---|---|---|---|
idleTime | integer, 0–10000 | 500 | Milliseconds the network must stay idle. |
concurrency | integer, 0–10 | 0 | Requests that may stay in flight while the network counts as idle. 1 or 2 suits a page that holds a long-poll or analytics request open. |
timeout | integer, 1–30000 | 30000 | Milliseconds to wait for the network to go idle. |
onTimeout | continue fail | continue | continue: the sequence carries on with the page as it is, and the action is listed in timedOutActions. fail: the sequence fails at this action. |
The defaults wait like networkidle0, and concurrency: 2 waits like networkidle2.
Wait for the results a click loads before you take a screenshot:
{
"url": "https://www.example.com/flights?from=LHR&to=JFK",
"actions": [
{ "type": "click", "selector": "button.search" },
{ "type": "waitForNetworkIdle", "idleTime": 1000, "timeout": 15000 },
{ "type": "screenshot", "full_page": true }
]
}Pages that never go idle
Some pages keep requests going for as long as they are open: asking for new data every few seconds, holding a
long-poll open, refreshing ads, or sending analytics as fetch() calls. On such a page networkidle0, and
waitForNetworkIdle with its defaults, wait out their full timeout every time. To read these pages faster:
- Allow a request or two in flight. Use
networkidle2, orwaitForNetworkIdlewithconcurrency: 1or2. - Wait for what you need instead. A
waitwith aselectorwaits for an element, and awaitForResponsewaits for one API response. - Shorten the wait. Give
waitForNetworkIdleatimeoutyou can afford to spend. - Use a load event. When the HTML is all you need,
domcontentloadedorloaddoes not depend on the network going idle.
When a wait runs out
| Response | Where it shows |
|---|---|
waitUntilTimedOut: true | In the JSON envelope, when the page was read, or the actions ran, before the page reached the waitUntil state. |
x-wait-until-timed-out: true header | On the same responses, in every format, so a raw, markdown or extracted jsonSchema response reports it too. |
timedOutActions | With actions: the 0-based indexes into your actions of the waitForNetworkIdle actions that ran out under onTimeout: "continue". |
Each is absent when no wait ran out.
{
"statusCode": 200,
"headers": { "content-type": ["text/html; charset=utf-8"] },
"timedOutActions": [1],
"waitUntilTimedOut": true,
"finalUrl": "https://www.example.com/live-scores",
"data": "<!doctype html>…"
}A waitForNetworkIdle with onTimeout: "fail" that runs out stops the sequence there. The response is still a 200,
with error: "Browser action sequence failed", a failedActionIndex that points at the action, and the page as it
stood in data.
Errors
Validation failures return a 400 with the standard error shape before anything runs:
message | Fix |
|---|---|
waitUntil requires executeJS: true, actions or screenshot | Add executeJS: true, or use actions. |
| Status | error | What to do |
|---|---|---|
503 | waitUntil and waitForNetworkIdle are temporarily unavailable | Retry later, with backoff. |
Billing
Waiting adds no fee of its own. A rendered fetch is billed per request, however long it waits. A request with actions
or screenshot is billed as a browser session, by bandwidth and duration, so time spent
waiting counts toward its duration.