MCP tools
Search the web, fetch any URL with GET or POST and map sites as clean Markdown or structured data — parameters and return shapes for every Web Access MCP tool.
The remote server exposes four tools. web_access_search is read-only. web_access_fetch can run
browser actions that modify destination data, web_access_request sends POST, PUT and PATCH requests, and
web_access_sitemap creates and manages billed crawl jobs.
The self-hosted source includes five tools; the published package may lag these docs. It has no
web_access_request, and its
web_access_fetch still takes a method, so it sends POST, PUT and PATCH directly while advertising a read-only
hint. Prompt on it the way you would on web_access_request. It also has web_access_product_help and
web_access_report, which the remote server does not.
web_access_fetch
Fetch any webpage and get clean, LLM-ready Markdown back. The common case is passing only a url.
The direct HTTP path issues a GET. To send a direct POST, PUT, or PATCH with a body, use
web_access_request. Browser actions can still modify destination data or submit forms.
| Parameter | Type | Default | Description |
|---|---|---|---|
url | string | — | Required. The full http/https URL to fetch. |
format | markdown raw json | markdown | markdown for clean text, raw for the verbatim body, json for a { statusCode, headers, data } envelope. |
headers | object | — | Custom request headers (max 50), sent in the order you give them. Not combinable with executeJS or actions. |
countryCode | string(2) | — | ISO 3166-1 alpha-2 country for geolocated proxy routing. |
executeJS | boolean | false | Render JavaScript for SPAs. Cannot be combined with headers. |
actions | array | — | Browser actions to run before the page is captured. See below. |
solveCaptcha | boolean | true | Set false to fail fast instead of solving challenges. Not combinable with actions. |
A header name given twice is an error; send one value per header name. headers together with executeJS is refused
before anything is sent, since a rendered fetch cannot carry them.
Returns: Markdown by default; the verbatim body or a JSON envelope when format is set accordingly.
{ "url": "https://example.com/article" }Browser actions
Pass actions when the content you want does not exist in the document until something happens to the page — a click
past a consent gate, a search form submitted, a "load more" button, or rows that render only once scrolled into view.
The sequence runs in a real browser session, and the page is captured after the last step.
Treat the remote tool as mutating and potentially destructive: browser actions can change destination data or submit a form, even though a URL-only call just reads a page.
Escalate in this order, cheapest first:
- Plain
url— always try this first. executeJS: true— the page renders client-side but needs no interaction.actions— the content requires interaction. Takes tens of seconds and holds a browser session.
If the data is missing at every rung, it is probably not in the HTML at all. Look for the JSON API the page itself calls and fetch that endpoint directly — it is faster and returns exact values.
{
"url": "https://example.com/listings",
"actions": [
{ "type": "wait", "selector": ".card", "timeout": 20000 },
{ "type": "scroll", "direction": "down" },
{ "type": "wait", "milliseconds": 1500 },
{ "type": "scroll", "direction": "down" },
{ "type": "wait", "milliseconds": 1500 }
]
}Waiting on a selector rather than a fixed duration continues as soon as the content is there, so prefer it wherever
you can name one. Up to 50 actions per request, one screenshot at most, and 30s caps any single wait. The full step
vocabulary is in Browser actions.
Returns with actions: a JSON object — data (the final page, Markdown by default), finalUrl, statusCode, and
screenshot when the sequence took one. A failed step still returns successfully, carrying error and
failedActionIndex (0-based, into your actions array) with data holding the page as it stood at that point.
actions cannot be combined with headers, format: "raw" or solveCaptcha: false; the call fails before anything is
sent, with an error naming the field. A browser session is always a GET.
Tool surface vs. the full API
The MCP web_access_fetch tool exposes the most common fetch options. Advanced /fetch features — structured
extraction (jsonSchema), requireWSS, and
ignoreCertificateErrors — are available on the HTTP
API directly.
web_access_request
Send a POST, PUT, or PATCH with a body to a URL you supply, over the same path
web_access_fetch uses. Plenty of those calls only read: a GraphQL query, a search backend that takes a body, a JSON
API that refuses GET. It is a separate tool because method and body do not exist on web_access_fetch at all:
requests with a body get their own schema and their own warning. The split is about request shape, not safety — browser
actions make web_access_fetch mutating too.
| Parameter | Type | Default | Description |
|---|---|---|---|
url | string | — | Required. The full http/https URL to send the request to. |
method | POST PUT PATCH | — | Required. GET is refused here — read a page with web_access_fetch. |
body | string | object | — | Request body. A string is sent as-is; an object is sent as the exact JSON text you wrote. |
format | markdown raw json | json | Defaults to the envelope: an API caller usually needs the status code, not prose. |
headers | object | — | Custom request headers (max 50), sent in the order you give them. A name given twice is an error. |
countryCode | string(2) | — | ISO 3166-1 alpha-2 country for geolocated proxy routing. |
solveCaptcha | boolean | true | Set false to fail fast instead of solving challenges. |
executeJS and actions are not available here — both run a browser session, which is always a GET.
Returns: a { statusCode, headers, data } envelope by default.
{
"url": "https://example.com/api/items",
"method": "POST",
"body": { "name": "widget" }
}web_access_search
Search the web from a query and get ranked organic results back, plus every surface Google rendered around them. Set
format to raw for the Google results page as HTML, one page per call.
| Parameter | Type | Description |
|---|---|---|
query | string | Required. The search query. Be specific for best results. |
searchCount | integer | Optional. How many organic results you want, 1 to 300; above 300 is rejected. Google is paged, up to 36 pages, until that many are in hand or it has no more, and each page is billed as one search. Omit it for one page. Rejected with format: "raw". |
aiMode | boolean | Optional. true asks Google AI Mode instead of the results page: the response carries aiMode, the answer with its cited sources and any products, places and videos it shows, and results is empty. Billed as one search. Not combinable with searchCount. |
aiOverview | boolean | Optional, off by default. true makes a second, sequential fetch to fill the AI Overview Google streams in after the page loads. Filling a streamed overview adds a few seconds to the request. Costs 2x a normal search. Without it, an overview already present in the page is still returned, and a streamed one comes back as declined: true. Not combinable with aiMode. |
country | string | Optional. ISO 3166-1 alpha-2 code the search runs from, such as GB. Default US. |
language | string | Optional. Results language tag such as en or pt-br. |
location | string | Optional. A place name such as London or Austin,Texas,United States, 1 to 200 characters; the search is sent from its country. Google results are biased toward the place, most strongly for queries like plumbers near me, and may still cover a wider area; with aiMode the answer is given for that place. A name that cannot be placed is rejected; send coordinates instead. |
coordinates | object | Optional. A point to search from, as { latitude, longitude, radius } with radius in meters (default 5000), sent from the country it lies in. Google results are biased toward it; with aiMode the answer is given for it. If location is also sent, coordinates win. |
page | integer | Optional. Google results page to start from, 1 to 30, where page N is the page Google shows as N; default 1. Without searchCount the response is that one page, and with format: "raw" it always is; with searchCount, results are collected starting from that page. page and searchCount together stay within the first 300 results. position stays 1-based within the response; rank is the result's Google rank. A page past the last result returns zeroResults: true. A page holds about eight to ten results, so separate page calls can repeat or skip a result; send one call with searchCount for a list without repeats. |
dateRange | string or object | Optional. Limit results to a publication window: hour, day, week, month or year for the past hour through the past year, or { "from": "2024-01-01", "to": "2024-06-30" } with ISO dates, inclusive. A custom range needs at least one end, and from must not be after to. |
sortBy | string | Optional. relevance (the default) or date for the newest results first. |
format | string | Optional. structured (the default) for results as JSON, or raw for the Google results page as HTML, one page per call. We recommend structured. Raw supports page only: searchCount is rejected with raw. Raw is web only. dateRange and sortBy work with raw. See below for the size limit. |
searchType | string | Optional. The Google tab to search: web (the default), images, videos, shopping, books, places or forums. images answers images (3 pages of about 100), shopping answers products (one page of about 55, so page must be 1), places answers places (20 a page), and the others answer results. See Search type. |
safeSearch | boolean | Optional. true removes explicit results. |
includeOmittedResults | boolean | Optional. true includes the results Google hides as very similar to ones already shown. |
autocorrect | boolean | Optional. false searches the query exactly as typed rather than Google's corrected spelling. |
restrictCountry | string | Optional. A two-letter code; returns only pages from that country. Unlike country, it does not change where the search runs from. |
verbatim | boolean | Optional. true matches the query's words exactly, without synonyms. web, videos and forums only. |
page, dateRange, sortBy, format, searchType and the filters apply to Google results only and are rejected with
aiMode. dateRange, sortBy and verbatim apply to the web, videos and forums tabs; format: "raw" and
aiOverview to web only.
With format: "raw", the response is html, the Google results page as HTML, in place of results and the surfaces,
with htmlBytes and htmlTruncated. Ask for page N with page, one call per page, each billed as one search; a raw
call with searchCount is rejected. See Raw HTML.
The tool sends at most 60,000 bytes of HTML per call. A page cut short has htmlTruncated: true, and htmlBytes is the
size of the whole page. For the whole page, call POST /search with
"format": "raw". Prefer the default structured results unless you need HTML.
Returns: results, the ranked organic documents, each with position (its place in this response), rank
(its Google rank; Google only), title, url (the link to fetch; absent when the destination is not known, and the result is still returned),
snippet, displayUrl (Google's displayed URL line, empty when Google shows none, as on Reddit and YouTube results),
displayText (the source line Google shows under the title, verbatim: a URL, engagement counts such as
20+ comments · 3 months ago, or other text), source when Google names the site (such as Reddit · r/buildapc),
and a timestamp when the engine dated the result; and zeroResults, true when the engine itself reported that
nothing matched. Every other field is a surface Google rendered around the documents. Each is present only when the
page carried it, is never merged into results, and only Google returns it — the same shape as
POST /search:
| Field | What it carries |
|---|---|
entity | The knowledge panel for the one business or person the query named: title, subtitle, description and its descriptionSource, rating, reviews, website, attributes, profiles, unread. |
places | Local-pack business listings: name, category, rating, reviews, address, phone, hours, url, mapsUrl. With searchType: "places", the whole answer. |
products | Only with searchType: "shopping": product listings with title, productId, price, originalPrice, merchant, moreMerchants, delivery, returns, rating, reviews; no url. |
overviews | Google's AI overviews: the query's own first (no topic), then the "Things to know" tabs (topic, question); declined: true marks a frame Google did not fill, including a streamed overview when aiOverview was not set. Each has text and cited sources as { title, url }. |
peopleAlsoAsk | { question } entries; answers are not on the page. |
relatedSearches | Query strings Google suggests. |
answers | One single-purpose widget under localTime, currency, unitConversion, weather, translation, sports or flights. |
spelling | { kind, query, asked } — substituted means the results are for the corrected query, suggested means they are for the query as typed. |
ads | Sponsored results, with advertiser, kept out of results. |
videos, shortVideos | platform, channel, date, duration where shown. |
discussions | site, forum, comments, age, excerpt where shown. |
images | Source pages of the Images block, with source. With searchType: "images", the whole answer, each also with imageUrl, imageWidth, imageHeight and thumbnail. |
sitelinks | Sub-pages under a result, with resultPosition. |
aiMode | Only with aiMode: true: Google's AI Mode answer as plain text, as markdown, and as blocks in reading order (paragraph, heading, list, table with header and rows, code); the cited sources as { title, url, snippet, source }, url absent when a citation could not be resolved; and, when the answer shows them, products (title, price, oldPrice, merchant, rating, reviews, url…), places (name, category, rating, reviews, status, address, url…) and videos (title, url, channel, duration). See Google AI Mode. |
paging | { pages, complete, stoppedBy }, only when searchCount was sent. pages is how many results pages answered, each billed as one search. stoppedBy is why paging stopped: search_count (the count was reached), end_of_results (Google had no more, or the first page carried no organic results), page_cap (the 36-page limit; Google may have more), page_failed (a later page could not be fetched) or deadline (the time budget ran out). complete is false only for page_failed and deadline, and results then holds what was collected. Surfaces describe the first page only; positions run on across pages. |
A query naming a single business often comes back with results empty and the answer in entity or places, so read
those before treating an empty results as no answer. A snippet is the engine's description of the page, not the page
itself, and an overview is a summary, not a source — fetch the url with web_access_fetch to
read it.
{ "query": "latest developments in AI agents 2026" }{ "query": "construction consulting firms Ohio", "searchCount": 30 }{ "query": "heat pump grants", "format": "raw" }{ "query": "plumbers near me", "location": "Austin,Texas,United States" }{ "query": "how does a heat pump work", "aiMode": true }{ "query": "what is the weather today", "aiMode": true, "location": "London" }{ "query": "best bakeries near me", "aiMode": true, "coordinates": { "latitude": 48.8566, "longitude": 2.3522 } }web_access_sitemap
Crawl a whole site and map its URLs. Starting from one URL the crawler follows same-domain links breadth-first
(optionally seeded from the site's /sitemap.xml) and records every URL it reaches with fetch status, depth, and
parent. Crawls run asynchronously as sitemap jobs: one tool drives the whole
quote → approve → poll → read lifecycle through its action parameter, and nothing is crawled or billed until the
quote is explicitly approved.
action | Requires | Description |
|---|---|---|
submit | url | Quote a crawl. Returns jobId, estimatedPages, and estimatedCostUsd with status awaiting_approval. Quotes expire after one hour. |
approve | jobId | Start the quoted crawl — the billing-consent step. Fails with 402 on insufficient balance; retry on a 409 partial_state. |
status | jobId | Poll progress: awaiting_approval → running (pending / processed counts) → completed | failed | canceled | token_cap_exceeded. |
results | jobId | Page through discovered URLs. total tells you when to stop paging. |
cancel | jobId | Stop a non-terminal job. Pages already fetched stay billed and readable. |
list | — | Your organization's recent crawl jobs, most recent first. |
| Parameter | Type | Used by | Description |
|---|---|---|---|
action | enum | — | Required. One of the six actions above. |
url | string (URL) | submit | Required for submit. Starting URL. The crawl stays on this hostname. |
maxPages | integer | submit | Maximum pages to crawl (1–10,000, default 10 when omitted). Each fetched page is billed. |
maxDepth | integer | submit | Maximum link depth from the starting URL (1–100, default 2 when omitted). |
pathPrefix | string | submit | Only crawl URLs whose path starts with this prefix. |
budgetUsd | number | submit | Spend cap in USD (min 0.0001). Defaults to the quote; the crawl stops with token_cap_exceeded before exceeding it. |
useSitemap | boolean | submit | Also seed the crawl from the site's root /sitemap.xml (one extra billed page; default false). |
jobId | string | job-scoped actions | The id returned by submit. |
limit | integer | results / list | Page size — results 1–5,000, default 1000; list 1–100, default 20. |
offset | integer | results / list | Number of entries to skip, 0 or more. |
A bound outside its range, 0 included, is an error rather than being replaced by the default. Omit the field to get
the default.
Returns: the JSON envelope for the chosen action (quote, status, URL page, or job list) alongside a one-line summary.
{ "action": "submit", "url": "https://example.com", "maxPages": 200, "maxDepth": 3 }Same jobs as the HTTP API
The tool fronts the sitemap endpoints directly — response shapes, statuses, billing,
and result retention (per-URL discoveredUrls is only available while results are live) all match. See the Sitemap
guide for a walkthrough.
Recommended workflow
- Call
web_access_searchto find relevant pages for a query. - Call
web_access_fetchon the most relevant URLs to pull their full content.
To cover a whole site instead of single pages, run web_access_sitemap first (submit → approve → poll → results) to
build the URL list, then web_access_fetch the pages you need.