XHR capture
Return the XHR and fetch() requests a rendered page makes, filter them down to the API calls you want, and wait for a specific response.
Many pages load their data from a JSON API after the HTML arrives. With captureXHR, a rendered /fetch also returns
the XHR and fetch() requests the page made: each request's URL, method, headers and body, and the response's status,
headers and body. You can read the API's own response instead of parsing the HTML it was rendered into.
There are three building blocks:
captureXHRreturns the page's requests asxhr, either all of them or only the ones that match filters.- Browser actions extend the capture to a whole actions session, so the requests your clicks, typing and scrolling trigger are included.
- The
waitForResponseaction waits for one matching response, and withasResultreturns it as the response body in place of the page.
Quick start
curl https://request.usestring.ai/v1/fetch \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://shop.example.com/search?q=headphones",
"executeJS": true,
"captureXHR": true
}'The usual JSON envelope comes back with an xhr array next to the page:
{
"statusCode": 200,
"headers": { "content-type": ["text/html; charset=utf-8"] },
"data": "<!doctype html>…",
"finalUrl": "https://shop.example.com/search?q=headphones",
"xhr": [
{
"resourceType": "fetch",
"request": {
"url": "https://shop.example.com/api/search?q=headphones&page=1",
"method": "GET",
"headers": { "accept": "application/json" }
},
"response": {
"statusCode": 200,
"headers": { "content-type": "application/json" },
"body": "{\"total\":412,\"items\":[{\"sku\":\"WH-1000\",\"price\":349.99}]}",
"bodyEncoding": "utf8"
},
"timing": { "startedAt": 1791400000123, "durationMs": 184 }
}
]
}Bodies are strings. A JSON response arrives as JSON text in response.body, so parse it yourself.
How long the capture runs
The capture ends when the page is read, and a request the page sends after that is not in xhr. When that happens
depends on whether the request runs browser actions.
A plain fetch waits for the network to go quiet. With captureXHR on, the page loads and then the request waits
until the page's network has been quiet for 500 ms, for up to 10 seconds. A page that never goes quiet, because it
polls, streams or sends analytics heartbeats, is read when the 10 seconds are up. The request doesn't fail. Bodies
still downloading at that point get up to 3 more seconds to finish.
With actions, you decide how long to wait. The page load and each navigate, click, press, scroll and
selectOption get a settle of about 3 seconds at most, and other actions get none. The whole sequence has 180
seconds. After the last action, bodies still downloading get up to 3 seconds, and then the session ends. See
Timing for the details.
The settles and the 10-second wait are best effort. To make sure a particular request is in the response, wait for it:
| You need | Send |
|---|---|
| The requests a page makes while it loads | A plain fetch with captureXHR. |
| A request that fires late, or one you can't afford to miss | actions that start with a waitForResponse. It waits up to 30 seconds. |
| The requests a click, search or scroll triggers | actions with that interaction, then a waitForResponse for the request. |
| More time on a page that keeps loading, with no single request to wait for | actions with a { "type": "wait", "milliseconds": 5000 }, up to 30000 per wait. |
Turning it on
captureXHR | What xhr contains |
|---|---|
true or {} | Every XHR and fetch() request the page made. |
{ "include": [ … ] } | Only the requests that match at least one filter. |
false or omitted | Nothing. The response has no xhr field. |
When captureXHR is on, xhr is always present, and it is an empty array when nothing qualified.
Requirements
executeJS: trueis required, including on requests that useactionsorscreenshot. Without it the request returns a400.- A plain fetch must use
format: "json"(the default). A request withactionsorscreenshotcan also useformat: "markdown", since its response is a JSON envelope either way. jsonSchemacannot be combined withcaptureXHR.- The
executeJSrules still apply: the page is loaded with aGET, with nobodyand no customheaders.
What gets captured
- Requests the page sends with
XMLHttpRequestorfetch()tohttpandhttpsURLs, in the order the page sent them. The page's own navigation, scripts, stylesheets, images, fonts, media,navigator.sendBeacon()pings, EventSource streams and WebSockets are not captured. Neither is a request whose URL is longer than 8 KiB. - From the first navigation until the page is read, across the whole session when the request uses
actionsorscreenshot. See How long the capture runs. - A request still loading when the page is read comes back incomplete. A response whose body was still downloading
has its status and headers but no
body, and a request with neitherresponsenorerrorhad not received a response yet. - Your organization's access rules apply to every XHR and
fetch()request the page makes. A request they deny is blocked in the browser, so it never reaches its destination: it has anerrorand noresponse. - Traffic that belongs to a site's bot-protection challenge is left out.
- Proxy headers, such as
Proxy-*, are removed from captured request and response headers.
The xhr entry
| Field | Type | Description |
|---|---|---|
resourceType | xhr | fetch | Whether the page used XMLHttpRequest or fetch(). |
request.url | string | The full request URL, query string included. |
request.method | string | The request method, such as GET or POST. |
request.headers | object | The request headers, each value a single string. |
request.body | string | The body the page sent, encoded as bodyEncoding says. Absent when there was no body or it was omitted. |
request.bodyEncoding | utf8 | base64 | utf8 for valid UTF-8 text, base64 for anything else. |
request.bodyOmitted | too_large | budget_exhausted | unavailable | Why request.body is absent. See Omitted bodies. |
response | object | Absent when no response arrived. |
response.statusCode | number | The response status. |
response.headers | object | The response headers, each value a single string. |
response.body | string | The response body, decoded from any content-encoding and encoded as bodyEncoding says. |
response.bodyEncoding | utf8 | base64 | utf8 for valid UTF-8 text, base64 for anything else. |
response.bodyOmitted | too_large | budget_exhausted | unavailable | Why response.body is absent. See Omitted bodies. |
error | string | The network error, when the request failed. |
matchedFilter | number | Present when captureXHR set include: the 0-based index of the first filter this request matched. |
timing.startedAt | number | When the page sent the request, in milliseconds since the Unix epoch. |
timing.durationMs | number | Milliseconds from sending the request until its response finished loading or it failed. Absent while in flight. |
Captured headers hold one string per name. The envelope's top-level headers, which describe the page, can hold
arrays.
Omitted bodies
bodyOmitted | Why the body is absent |
|---|---|
too_large | A response body over 2 MiB once decoded, or a request body over 64 KiB. |
budget_exhausted | The capture's 10 MiB budget was used up, so there was no room left for this body. |
unavailable | The body could not be captured, for example on a redirect or an aborted request. |
Limits and truncation
| Limit | Value |
|---|---|
| Captured requests per response | 200 |
Total size of xhr, as JSON | 10 MiB |
| Response body | 2 MiB, decoded |
| Request body | 64 KiB |
| Request URL | 8 KiB. Longer URLs are not captured. |
A limit never fails the request. The page is returned and billed as usual. When a limit left requests or bodies
out, the response carries xhrTruncated: true; otherwise the field is absent. A very busy page can also make the
capture stop early, and xhrTruncated reports that too.
On a busy page, analytics and polling calls can use up the limits before the request you want arrives. Filters prevent
that: a request that matches no filter takes none of the 200 entries or the 10 MiB, and never sets xhrTruncated.
Filter with include
include is a list of 1 to 10 filters. A request is captured when it matches at least one of them, and it matches a
filter when it meets every condition that filter sets. Each captured request carries matchedFilter, the index of the
first filter it matched.
{
"url": "https://shop.example.com/search?q=headphones",
"executeJS": true,
"captureXHR": {
"include": [
{ "url": "/api/search", "methods": ["GET"], "statusCodes": [200] },
{ "url": "^https://shop\\.example\\.com/graphql\\?op=Products", "match": "regex" }
]
}
}| Field | Type | Default | Description |
|---|---|---|---|
url | string, 1–2048 characters | — | Required. The pattern, matched against the full request URL with its query string. A request URL over 8192 characters matches no filter. |
match | contains glob regex | contains | How url is matched. See Match modes. |
methods | array of methods | any method | Matches only these methods, in any case: GET, POST, PUT, PATCH, DELETE, HEAD, OPTIONS. |
statusCodes | array of 1–20 integers, 100–599 | any status | Matches only responses with one of these statuses. A request that failed, or had no response yet, never matches. |
Unknown fields in a filter return a 400. A request that matched a filter's url but came back with a status none of
its matching filters list is dropped when its response arrives, and gives back its share of the limits.
Match modes
match | url is | Example url |
|---|---|---|
contains | A case-sensitive substring of the request URL. | /api/search |
glob | A case-sensitive pattern for the whole URL. * matches any run of characters, / included. ? matches exactly one character. Every other character matches itself. | https://*.example.com/api/v?/products* |
regex | An RE2 regular expression that may match anywhere in the URL. Anchor it with ^ and $ to match the whole URL. | ^https://api\.example\.com/v[0-9]+/items\? |
Notes on the patterns:
globmatches the whole URL, so a pattern without a trailing*must end exactly where the URL ends, query string included. Because?stands for any one character, it also matches a literal?in the URL.- A glob part between two
*that contains?can be at most 64 characters. Longer parts return a400; shorten the part or useregex. - RE2 has no lookaheads, lookbehinds or backreferences. A pattern that uses them, or is otherwise invalid, returns
a
400that names the problem. - Size limits apply to regular expressions. A pattern that compiles too large, usually from counted repetitions
such as
{n,m}, returns a400. The regular expressions in one request also share a combined size limit. - Escape backslashes for JSON. The regex
\.is written"\\."inside a JSON string.
Capture during browser actions
Add captureXHR to a request with actions and the capture covers the whole session, including the requests each
action triggers. executeJS: true is still required.
curl https://request.usestring.ai/v1/fetch \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://shop.example.com/category/headphones",
"executeJS": true,
"captureXHR": { "include": [{ "url": "/api/products", "statusCodes": [200] }] },
"actions": [
{ "type": "click", "selector": "button.load-more" },
{ "type": "waitForResponse", "url": "/api/products", "statusCodes": [200], "timeout": 10000 }
]
}'The response is the browser actions envelope with xhr and, when a limit
was hit, xhrTruncated added. The waitForResponse action here holds the session open until the response the click
triggered has arrived, so it is in xhr when the session ends.
Wait for a response
waitForResponse is a browser action. It waits for an XHR or fetch() response
that matches its conditions and has finished loading. It does not need captureXHR or executeJS.
{ "type": "waitForResponse", "url": "/api/search", "methods": ["GET"], "statusCodes": [200], "timeout": 10000 }| Field | Type | Default | Description |
|---|---|---|---|
url | string | — | Required. Matched like a filter's url. |
match | contains glob regex | contains | As for filters. |
methods | array of methods | any method | As for filters. |
statusCodes | array of integers | any status | As for filters. |
timeout | integer, 0–30000 | 30000 | Milliseconds to wait for a matching response. |
onTimeout | fail continue | fail | fail: no match within timeout fails the sequence at this action. continue: carry on with the next action. |
asResult | boolean | false | Return the matched response as the response body. |
Which responses count
You don't have to start waiting before the request is sent. A wait looks back as well as forward:
- A response counts if it finished after the action before the wait started. "The action before" is the most
recent action that can start traffic: anything except
wait,screenshotandwaitForResponse. A response that aclick,pressorscrolltriggered is therefore not missed, even when it finished before the wait began, and even with awaitfor a selector in between. - Before the first such action, the window starts with the page load. A
waitForResponseat the start of the sequence can match a request the page makes while it loads. - Consecutive waits share one window and each takes a different response. Two waits after one click can match two responses that the click triggered, and two waits with the same filter wait for two different matching responses.
- The earliest matching response wins. When no unclaimed match has finished yet, the action waits up to
timeoutfor the next one. - Redirects are passed over. A
3xxresponse that redirects never matches; the response it leads to can.
When a wait with onTimeout: "fail" times out, the sequence stops there. The response carries error and a
failedActionIndex that points at the wait, and data holds the page as it stood.
Return a response as the result
Set asResult: true on a waitForResponse action to get the matched response back as data, instead of the page.
curl https://request.usestring.ai/v1/fetch \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://shop.example.com/search?q=headphones",
"actions": [
{ "type": "click", "selector": "button.load-more" },
{ "type": "waitForResponse", "url": "/api/search", "statusCodes": [200], "timeout": 10000, "asResult": true }
]
}'{
"statusCode": 200,
"headers": { "content-type": ["text/html; charset=utf-8"] },
"finalUrl": "https://shop.example.com/search?q=headphones",
"data": "{\"total\":412,\"page\":2,\"items\":[…]}",
"matchedResponse": {
"url": "https://shop.example.com/api/search?q=headphones&page=2",
"method": "GET",
"statusCode": 200,
"headers": { "content-type": "application/json" },
"bodyEncoding": "utf8"
}
}datais the matched response's body, as a string. Parse it when the API returns JSON. WhenmatchedResponse.bodyEncodingisbase64, decodedatafirst. Those bytes are not valid UTF-8, so read them with the charset the response'scontent-typedeclares.matchedResponsedescribes the matched response: itsurl,method,statusCode,headersandbodyEncoding.statusCode,headersandfinalUrlstill describe the page, not the matched response.- It applies only when the whole sequence succeeds and that action matched a response. If any action fails, or the
wait used
onTimeout: "continue"and nothing matched,datais the page andmatchedResponseis absent. Check formatchedResponsebefore you parsedata. - The body can be up to 10 MiB. A larger or unreadable body fails the action, whatever
onTimeoutsays. A match that cannot be returned fails it with the errorThe response waitForResponse matched could not be returned. - Only one action per request can set
asResult. It requiresformat: "json"(the default) and cannot be combined withjsonSchema.
asResult combines with captureXHR, so one response can carry the matched body in data and the captured requests in
xhr.
Recipes
The examples below use these two helpers. The client retries a 503, honoring Retry-After when it is present. Its
timeout leaves room for a browser actions session, whose last action can start just before the 180-second limit. The
decoder turns a captured body back into text. A base64 body is not valid UTF-8, so the decoder reads its bytes with
the charset you pass, which should be the one the body's content-type declares. Without one it uses Latin-1.
import base64
import json
import os
import time
import requests
API = "https://request.usestring.ai/v1/fetch"
HEADERS = {"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"}
def string_fetch(payload, attempts=3):
for attempt in range(1, attempts + 1):
res = requests.post(API, headers=HEADERS, json=payload, timeout=300)
if res.status_code == 503 and attempt < attempts:
time.sleep(int(res.headers.get("Retry-After", 2 * attempt)))
continue
res.raise_for_status()
return res.json()
def body_text(body, encoding, charset="latin-1"):
return base64.b64decode(body).decode(charset) if encoding == "base64" else bodyconst API = "https://request.usestring.ai/v1/fetch";
async function stringFetch(payload, attempts = 3) {
for (let attempt = 1; ; attempt++) {
const res = await fetch(API, {
method: "POST",
headers: { Authorization: `Bearer ${process.env.STRING_API_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify(payload)
});
if (res.status === 503 && attempt < attempts) {
const seconds = Number(res.headers.get("retry-after") ?? 2 * attempt);
await new Promise((resolve) => setTimeout(resolve, seconds * 1000));
continue;
}
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
return res.json();
}
}
const bodyText = (body, encoding, charset = "latin1") =>
encoding === "base64" ? new TextDecoder(charset).decode(Buffer.from(body, "base64")) : body;Read the API behind a single-page app
Load the page once and keep only the search API's successful responses.
result = string_fetch({
"url": "https://shop.example.com/search?q=headphones",
"executeJS": True,
"captureXHR": {"include": [{"url": "/api/search", "methods": ["GET"], "statusCodes": [200]}]},
})
if result.get("xhrTruncated"):
print("A capture limit left something out; narrow the filters.")
pages = [
json.loads(body_text(entry["response"]["body"], entry["response"]["bodyEncoding"]))
for entry in result["xhr"]
if "body" in entry.get("response", {})
]
if not pages:
raise RuntimeError("The page did not call /api/search before it was read")
items = [item for page in pages for item in page["items"]]const result = await stringFetch({
url: "https://shop.example.com/search?q=headphones",
executeJS: true,
captureXHR: { include: [{ url: "/api/search", methods: ["GET"], statusCodes: [200] }] }
});
if (result.xhrTruncated) console.warn("A capture limit left something out; narrow the filters.");
const pages = result.xhr
.filter((entry) => entry.response?.body !== undefined)
.map((entry) => JSON.parse(bodyText(entry.response.body, entry.response.bodyEncoding)));
if (pages.length === 0) throw new Error("The page did not call /api/search before it was read");
const items = pages.flatMap((page) => page.items);If the list comes back empty, the page may call the API later than the page is read. Use a wait instead.
Pick one GraphQL operation out of a shared endpoint
GraphQL apps often send every query to one URL, so the URL alone can't tell the operations apart. Capture the endpoint
and pick the operation by its request body. A second filter catches the same operation when it is sent as a GET with
the operation name in the query string; matchedFilter tells you which filter caught each entry.
result = string_fetch({
"url": "https://www.example.com/products/123",
"executeJS": True,
"captureXHR": {
"include": [
{"url": "/graphql", "methods": ["POST"], "statusCodes": [200]},
{
"url": "^https://www\\.example\\.com/graphql\\?.*operationName=ProductDetail(&|$)",
"match": "regex",
"statusCodes": [200],
},
]
},
})
def operation_name(entry):
if entry["matchedFilter"] == 1:
return "ProductDetail"
request = entry["request"]
if "body" not in request:
return None
payload = json.loads(body_text(request["body"], request["bodyEncoding"]))
return payload.get("operationName") if isinstance(payload, dict) else None
detail = next((entry for entry in result["xhr"] if operation_name(entry) == "ProductDetail"), None)
if detail is None or "body" not in detail["response"]:
raise RuntimeError("No ProductDetail response body was captured")
product = json.loads(body_text(detail["response"]["body"], detail["response"]["bodyEncoding"]))["data"]["product"]const result = await stringFetch({
url: "https://www.example.com/products/123",
executeJS: true,
captureXHR: {
include: [
{ url: "/graphql", methods: ["POST"], statusCodes: [200] },
{ url: "^https://www\\.example\\.com/graphql\\?.*operationName=ProductDetail(&|$)", match: "regex", statusCodes: [200] }
]
}
});
function operationName(entry) {
if (entry.matchedFilter === 1) return "ProductDetail";
const { body, bodyEncoding } = entry.request;
if (body === undefined) return undefined;
const payload = JSON.parse(bodyText(body, bodyEncoding));
return Array.isArray(payload) ? undefined : payload.operationName;
}
const detail = result.xhr.find((entry) => operationName(entry) === "ProductDetail");
if (detail?.response.body === undefined) throw new Error("No ProductDetail response body was captured");
const { product } = JSON.parse(bodyText(detail.response.body, detail.response.bodyEncoding)).data;A request body over 64 KiB is omitted (bodyOmitted: "too_large"), so an operation sent with a very large body can't
be identified this way.
Click "Load more" and return the next page's data
asResult makes the API response the response body, so there is no page to parse.
result = string_fetch({
"url": "https://shop.example.com/category/headphones",
"actions": [
{"type": "click", "selector": "button.load-more"},
{"type": "waitForResponse", "url": "/api/products", "statusCodes": [200], "timeout": 15000, "asResult": True},
],
})
if "matchedResponse" not in result:
raise RuntimeError(f"{result.get('error')} at action {result.get('failedActionIndex')}")
next_page = json.loads(body_text(result["data"], result["matchedResponse"]["bodyEncoding"]))const result = await stringFetch({
url: "https://shop.example.com/category/headphones",
actions: [
{ type: "click", selector: "button.load-more" },
{ type: "waitForResponse", url: "/api/products", statusCodes: [200], timeout: 15000, asResult: true }
]
});
if (!result.matchedResponse) throw new Error(`${result.error} at action ${result.failedActionIndex}`);
const nextPage = JSON.parse(bodyText(result.data, result.matchedResponse.bodyEncoding));Submit a search form and return the results API
Click the search box to focus it, type, and submit. The wait counts the response the press triggered, even if it
finished before the wait started.
{
"url": "https://www.example.com/",
"actions": [
{ "type": "click", "selector": "input[name=q]" },
{ "type": "write", "text": "noise cancelling headphones" },
{ "type": "press", "key": "Enter" },
{
"type": "waitForResponse",
"url": "https://www.example.com/api/search?*",
"match": "glob",
"methods": ["GET"],
"statusCodes": [200],
"asResult": true
}
]
}Wait for a call the page makes on load
A plain fetch reads the page once its network has been quiet for 500 ms, so a request the page sends on a timer after
that is missed. A waitForResponse as the first action matches responses from the page load onward, so it catches the
call whenever it lands within timeout.
{
"url": "https://www.example.com/flights?from=LHR&to=JFK&date=2026-11-02",
"actions": [{ "type": "waitForResponse", "url": "/api/fares", "statusCodes": [200], "timeout": 25000, "asResult": true }]
}Collect several pages from one session
Only one action can set asResult, so collect several responses through captureXHR and let the waits hold the
session open until each one arrives. Each wait follows its own click, so it matches that click's response.
actions = []
for page in (2, 3, 4):
actions += [
{"type": "click", "selector": "button.load-more"},
{"type": "waitForResponse", "url": f"/api/products?page={page}", "statusCodes": [200], "timeout": 15000},
]
result = string_fetch({
"url": "https://shop.example.com/category/headphones",
"executeJS": True,
"captureXHR": {"include": [{"url": "/api/products", "statusCodes": [200]}]},
"actions": actions,
})
if "failedActionIndex" in result:
print(f"Stopped at action {result['failedActionIndex']}; keeping the pages that arrived.")
pages = [
json.loads(body_text(entry["response"]["body"], entry["response"]["bodyEncoding"]))
for entry in result["xhr"]
if "body" in entry["response"]
]const actions = [2, 3, 4].flatMap((page) => [
{ type: "click", selector: "button.load-more" },
{ type: "waitForResponse", url: `/api/products?page=${page}`, statusCodes: [200], timeout: 15000 }
]);
const result = await stringFetch({
url: "https://shop.example.com/category/headphones",
executeJS: true,
captureXHR: { include: [{ url: "/api/products", statusCodes: [200] }] },
actions
});
if (result.failedActionIndex !== undefined) {
console.warn(`Stopped at action ${result.failedActionIndex}; keeping the pages that arrived.`);
}
const pages = result.xhr
.filter((entry) => entry.response.body !== undefined)
.map((entry) => JSON.parse(bodyText(entry.response.body, entry.response.bodyEncoding)));The first page, which the page requested as it loaded, is in xhr too. Keep the session within the
limits: around 200 responses and 10 MiB in total.
Scroll an infinite feed without failing on the last page
On the last page the feed stops loading, so a wait that can't find a match should let the sequence carry on.
onTimeout: "continue" does that.
{
"url": "https://www.example.com/feed",
"executeJS": true,
"captureXHR": { "include": [{ "url": "/api/feed", "statusCodes": [200] }] },
"actions": [
{ "type": "scroll", "direction": "down", "amount": 4000 },
{ "type": "waitForResponse", "url": "/api/feed", "timeout": 5000, "onTimeout": "continue" },
{ "type": "scroll", "direction": "down", "amount": 4000 },
{ "type": "waitForResponse", "url": "/api/feed", "timeout": 5000, "onTimeout": "continue" },
{ "type": "scroll", "direction": "down", "amount": 4000 },
{ "type": "waitForResponse", "url": "/api/feed", "timeout": 5000, "onTimeout": "continue" }
]
}Wait for two calls one click triggers
Consecutive waits share the click's window and each claims a different response, so the order in which the two calls finish doesn't matter.
{
"url": "https://www.example.com/products/123",
"executeJS": true,
"captureXHR": {
"include": [
{ "url": "/api/inventory", "statusCodes": [200] },
{ "url": "/api/shipping-quote", "statusCodes": [200] }
]
},
"actions": [
{ "type": "click", "selector": "button.check-availability" },
{ "type": "waitForResponse", "url": "/api/inventory", "statusCodes": [200] },
{ "type": "waitForResponse", "url": "/api/shipping-quote", "statusCodes": [200] }
]
}Errors
Validation failures return a 400 with the standard error shape before anything runs.
The messages that belong to this feature, as they appear in issues[].message and, for the first problem found, in
message:
message | Fix |
|---|---|
captureXHR requires executeJS: true | Add executeJS: true. |
format must be "json" when captureXHR is set | Use format: "json" on a plain fetch. |
captureXHR cannot be combined with jsonSchema | Drop one of the two. |
url is not a valid RE2 regular expression: … | Fix the pattern. RE2 has no lookarounds or backreferences. |
url is too large a regular expression; … | Use fewer or smaller counted repetitions such as {n,m}. |
url takes this request's regular expressions over their combined size limit; … | Use fewer or simpler patterns. |
url has a glob part between two * that contains ? and is over 64 characters; … | Shorten that part, or use match: "regex". |
at most one waitForResponse action can set asResult | Collect the others through captureXHR. |
format must be "json" when a waitForResponse action sets asResult | Remove format: "markdown". |
jsonSchema cannot be combined with a waitForResponse action that sets asResult | Drop one of the two. |
A 503 means the feature can't take the request right now. Requests that don't use it are unaffected.
| Status | error | What to do |
|---|---|---|
503 | XHR capture is at capacity; retry shortly | Retry after the number of seconds in the Retry-After header. |
503 | XHR capture is temporarily unavailable | Retry with backoff. |
503 | waitForResponse is temporarily unavailable | Retry with backoff. |
500 | XHR capture is temporarily unavailable | The capture could not start. Retry the request. |
A browser actions sequence that fails still returns 200, with error, the page in data, and a failedActionIndex
unless loading url itself failed:
error | Cause |
|---|---|
Browser action sequence failed | An action failed. For a waitForResponse, nothing matched within timeout under onTimeout: "fail", or the asResult body was over 10 MiB or unreadable. |
The response waitForResponse matched could not be returned | The asResult action matched a response that could not be returned. |
Billing
Capture has no fee of its own. A /fetch with captureXHR is billed per request, like any other fetch, at the
rate of the fetch path that served it. Because it requires executeJS, that is normally
the browser-based rate. With actions or screenshot, the request is billed as a
browser session, with or without capture. A capture limit never fails the request, so a
response with xhrTruncated: true is billed like any other successful request.