Iframes
Inline the content of a page's iframes, list them with their HTML, and run browser actions inside them.
Reviews, maps, booking and payment forms, video players and comment widgets are often embedded in iframes. A rendered
page's HTML holds only the <iframe> element, not the document inside it. Three request features reach that content,
cross-origin iframes included:
includeIframesinlines each iframe's content intodata, so Markdown,jsonSchemaextraction and your own parser see it as part of the page.captureIframeslists the page's iframes asiframes, each with its URL, a selector and its HTML.frameruns a browser action inside an iframe, such as a click on a button in an embedded form.
To also get the XHR and fetch() requests iframes make, see
Requests from iframes in XHR capture.
Quick start
curl https://request.usestring.ai/v1/fetch \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://shop.example.com/products/123",
"executeJS": true,
"includeIframes": true
}'data is the page with each iframe's content right after its <iframe> element:
<iframe src="https://reviews.example.net/widget?product=123" name="reviews"></iframe><div data-iframe-id="4F1D2C9AB0E37A65C1D8F2B49E0A7C13" data-iframe-url="https://reviews.example.net/widget?product=123" data-iframe-name="reviews"><link rel="stylesheet" href="https://reviews.example.net/widget.css"><ul class="reviews"><li>Great sound, comfortable fit.</li>…</ul></div>Requirements
- A browser render.
includeIframesandcaptureIframesneedexecuteJS: true,actionsorscreenshot. Without one of them the request returns a400.frameis a field of browser actions. captureIframeson a plain render needsformat: "json"(the default). Withactionsorscreenshotany supported format works.captureIframescannot be combined withjsonSchema.includeIframesworks withjsonSchema.
Iframes are read from the final page: after the page has loaded on a plain render, and after the last action ran with
actions.
Inline iframes with includeIframes
Set includeIframes: true and every iframe's content is placed into the page you get back:
- Each iframe's content follows its
<iframe>element, in adivwithdata-iframe-id,data-iframe-url(the URL of the iframe's document) and, when the iframe has aname,data-iframe-name. - The
divholds what is inside the iframe's<head>and<body>, not a second document. Relative URLs are made absolute. - An iframe's own iframes are inlined inside its
divthe same way.
data-iframe-id is the iframe's id in iframes and the frameId of the
requests it sent in xhr.
Formats and extraction
includeIframes applies to every format and to jsonSchema:
| Request | What you get |
|---|---|
format: "json" | The page with its iframes inlined as data. |
format: "raw" | The page with its iframes inlined as the response body. |
format: "markdown" | The page converted with its iframes. An iframe's text appears where the iframe was; the div markers do not survive conversion. |
jsonSchema | Extraction reads the page with its iframes inlined. |
Inlined iframes make the page larger, and Markdown conversion and extraction both have a size limit. See
Large pages and
body_too_large.
List iframes with captureIframes
captureIframes returns the final page's iframes as iframes in the JSON envelope.
captureIframes | What iframes contains |
|---|---|
true or {} | Every iframe, each with its HTML. |
{ "include": [ … ] } | Every iframe. Only the iframes whose URL matches a filter carry HTML. |
false or omitted | Nothing. The response has no iframes field. |
When captureIframes is set, iframes is always present, and it is an empty array when the page has none.
curl https://request.usestring.ai/v1/fetch \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://shop.example.com/products/123",
"executeJS": true,
"captureIframes": { "include": [{ "url": "https://reviews.example.net/", "match": "contains" }] }
}'{
"statusCode": 200,
"headers": { "content-type": ["text/html; charset=utf-8"] },
"data": "<!doctype html>…",
"finalUrl": "https://shop.example.com/products/123",
"iframes": [
{
"id": "4F1D2C9AB0E37A65C1D8F2B49E0A7C13",
"depth": 1,
"url": "https://reviews.example.net/widget?product=123",
"src": "https://reviews.example.net/widget?product=123",
"name": "reviews",
"selector": "iframe[name=\"reviews\"]",
"crossOrigin": true,
"statusCode": 200,
"html": "<!DOCTYPE html><html><head>…</head><body>…</body></html>",
"matchedFilter": 0
},
{
"id": "9B27E0C4D18F3A56B2E7C90D4F1A8E36",
"parentId": "4F1D2C9AB0E37A65C1D8F2B49E0A7C13",
"depth": 2,
"url": "https://media.example.org/player/42",
"src": "https://media.example.org/player/42",
"selector": "#player",
"crossOrigin": true,
"statusCode": 200,
"htmlOmitted": "not_matched"
}
]
}Each iframe is listed before the iframes nested in it.
The iframes entry
| Field | Type | Description |
|---|---|---|
id | string | An opaque id for the iframe. The same value is its data-iframe-id with includeIframes and the frameId of its requests in xhr. |
parentId | string | The id of the iframe this one is nested in. Absent for an iframe of the page itself. |
depth | integer | 1 for an iframe of the page itself, 2 for one nested inside it, and so on. |
url | string | The URL of the iframe's document when the page was read. Empty when the iframe had no document yet. |
src | string | The <iframe> element's src attribute, as written. Absent when empty. |
name | string | The element's name attribute. Absent when empty. |
selector | string | A CSS selector for the <iframe> element in its parent's document, usable in an action's frame. Absent when none could be built. |
crossOrigin | boolean | Whether the iframe's document is from a different origin than the page. |
statusCode | integer | The HTTP status of the iframe's latest document response. Absent when its document was not loaded over HTTP, as for about:blank. |
html | string | The iframe's whole document, as it stood when the page was read. |
htmlOmitted | not_matched | too_large | budget_exhausted | unavailable | Why html is absent. See Omitted HTML. |
matchedFilter | integer | Present when captureIframes set include and the iframe's URL matched: the 0-based index of the first filter it matched. |
A selector can be a positional path such as body:nth-of-type(1) > iframe:nth-of-type(2). On a page whose iframes
vary between loads, such as one with ads, it can point at a different iframe next time.
Omitted HTML
htmlOmitted | Why the HTML is absent |
|---|---|
not_matched | No include filter matched the iframe's URL. The iframe is still listed. |
too_large | The iframe's HTML is over 2 MiB. |
budget_exhausted | Too little of the request's 4 MiB of iframe HTML was left for this iframe. |
unavailable | The iframe could not be read, for example because it took too long or went away. |
Filter with include
include is a list of 1 to 10 filters, matched against the full URL of each iframe's document. Each filter takes the
same url and match fields as an XHR capture filter, with the same
match modes: contains (the default), glob and regex. Other fields return a
400.
{ "captureIframes": { "include": [{ "url": "https://*.example.net/widget?*", "match": "glob" }] } }Act inside an iframe
The wait, click, scroll, hover and selectOption browser actions take a
frame field. The action then runs inside that iframe, cross-origin ones included, and its own selector is looked up
in the iframe's document.
{
"url": "https://shop.example.com/products/123",
"includeIframes": true,
"actions": [
{ "type": "wait", "frame": "iframe[name=\"reviews\"]", "selector": ".review" },
{ "type": "click", "frame": "iframe[name=\"reviews\"]", "selector": "button.show-more" },
{ "type": "wait", "frame": "iframe[name=\"reviews\"]", "selector": ".review:nth-of-type(20)" }
]
}frameis a selector for the<iframe>element, such asiframe[name="reviews"]oriframe[src*="checkout"].- For an iframe inside another iframe,
frameis a list of selectors from the outermost iframe in, at most 5:["#store-locator", "iframe.map"]. - Each selector uses its first match, as an action's
selectordoes. - Actions wait for the iframe as they wait for an element. An iframe or element that doesn't appear within the action's timeout fails the sequence at that action. See Timeouts.
- A
waitwithframeneeds aselector. writeandpresstype into whatever has focus, soclickthe field inside the iframe first, thenwrite.scrollwithframeand noselectorscrolls the iframe's own document bydirectionandamount.
A form inside an iframe:
{
"url": "https://www.example.com/book",
"actions": [
{ "type": "click", "frame": "#booking-widget", "selector": "input[name=date]" },
{ "type": "write", "text": "2026-11-02" },
{ "type": "selectOption", "frame": "#booking-widget", "selector": "select[name=party]", "value": "4" },
{ "type": "click", "frame": "#booking-widget", "selector": "button[type=submit]" },
{ "type": "wait", "frame": "#booking-widget", "selector": ".confirmation" }
]
}Build frame from iframes
captureIframes gives each iframe a selector and the parentId of the iframe around it, which together make its
frame. The helper below walks from an iframe up to the page.
def frame_for(iframes, iframe_id):
by_id = {iframe["id"]: iframe for iframe in iframes}
path = []
iframe = by_id.get(iframe_id)
while iframe is not None:
if "selector" not in iframe:
return None
path.insert(0, iframe["selector"])
iframe = by_id.get(iframe.get("parentId"))
return path or Nonefunction frameFor(iframes, iframeId) {
const byId = new Map(iframes.map((iframe) => [iframe.id, iframe]));
const path = [];
for (let iframe = byId.get(iframeId); iframe; iframe = byId.get(iframe.parentId)) {
if (iframe.selector === undefined) return undefined;
path.unshift(iframe.selector);
}
return path.length > 0 ? path : undefined;
}Limits and truncation
| Limit | Value |
|---|---|
| Iframes per page, nested ones included | 50 |
| Nesting depth | 5 levels |
| HTML per listed iframe | 2 MiB |
| HTML read per request | 4 MiB, counting each iframe once for includeIframes and once for captureIframes, and with includeIframes the page's own HTML |
The page's own HTML, with includeIframes | 4 MiB. A larger page comes back as it is, without its iframes. |
A limit never fails the request. When a limit, or an iframe that could not be read, left an iframe or its content
out, the response carries iframesTruncated: true and the x-iframes-truncated: true header; otherwise both are
absent. The header is sent on every format, so a raw, markdown or extracted jsonSchema response reports it too.
Errors
Validation failures return a 400 with the standard error shape before anything runs:
message | Fix |
|---|---|
includeIframes requires executeJS: true, actions or screenshot | Add executeJS: true, or use actions. |
captureIframes requires executeJS: true, actions or screenshot | Add executeJS: true, or use actions. |
format must be "json" when captureIframes is set | Use format: "json" on a plain render. |
captureIframes cannot be combined with jsonSchema | Drop one of the two, or use includeIframes with jsonSchema. |
frame requires selector | Give the wait a selector, or drop frame. |
Too big: expected array to have <=5 items (field actions.N.frame) | Reach the iframe through at most 5 selectors. |
| Status | error | What to do |
|---|---|---|
503 | iframe support is temporarily unavailable | Retry later, with backoff. |
Billing
Iframes add no fee of their own.
- A plain render is billed per request, like any
executeJSfetch. The iframe content indataandiframescounts toward the response'sx-data-transfer-bytes. - A request with
actionsorscreenshotis billed as a browser session, by bandwidth and duration. The time spent reading iframes counts toward the session's duration.
Waiting for page load
Wait for a rendered page to reach a load state, from its HTML being parsed to its network going idle, before it is read or before actions run.
XHR capture
Return the XHR and fetch() requests a rendered page makes, filter them down to the API calls you want, and wait for a specific response.