String API
MCP Server

MCP tools

Search the web, fetch any URL with GET or POST and map sites as clean Markdown or structured data — parameters and return shapes for every Web Access MCP tool.

The remote server exposes four tools. web_access_search is read-only. web_access_fetch can run browser actions that modify destination data, web_access_request sends POST, PUT and PATCH requests, and web_access_sitemap creates and manages billed crawl jobs.

The self-hosted source includes five tools; the published package may lag these docs. It has no web_access_request, and its web_access_fetch still takes a method, so it sends POST, PUT and PATCH directly while advertising a read-only hint. Prompt on it the way you would on web_access_request. It also has web_access_product_help and web_access_report, which the remote server does not.

web_access_fetch

Fetch any webpage and get clean, LLM-ready Markdown back. The common case is passing only a url.

The direct HTTP path issues a GET. To send a direct POST, PUT, or PATCH with a body, use web_access_request. Browser actions can still modify destination data or submit forms.

ParameterTypeDefaultDescription
urlstring—Required. The full http/https URL to fetch.
formatmarkdown raw jsonmarkdownmarkdown for clean text, raw for the verbatim body, json for a { statusCode, headers, data } envelope.
headersobject—Custom request headers (max 50), sent in the order you give them. Not combinable with executeJS or actions.
countryCodestring(2)—ISO 3166-1 alpha-2 country for geolocated proxy routing.
executeJSbooleanfalseRender JavaScript for SPAs. Cannot be combined with headers.
actionsarray—Browser actions to run before the page is captured. See below.
solveCaptchabooleantrueSet false to fail fast instead of solving challenges. Not combinable with actions.

A header name given twice is an error; send one value per header name. headers together with executeJS is refused before anything is sent, since a rendered fetch cannot carry them.

Returns: Markdown by default; the verbatim body or a JSON envelope when format is set accordingly.

Example call
{ "url": "https://example.com/article" }

Browser actions

Pass actions when the content you want does not exist in the document until something happens to the page — a click past a consent gate, a search form submitted, a "load more" button, or rows that render only once scrolled into view. The sequence runs in a real browser session, and the page is captured after the last step.

Treat the remote tool as mutating and potentially destructive: browser actions can change destination data or submit a form, even though a URL-only call just reads a page.

Escalate in this order, cheapest first:

  1. Plain url — always try this first.
  2. executeJS: true — the page renders client-side but needs no interaction.
  3. actions — the content requires interaction. Takes tens of seconds and holds a browser session.

If the data is missing at every rung, it is probably not in the HTML at all. Look for the JSON API the page itself calls and fetch that endpoint directly — it is faster and returns exact values.

Load rows that render as they scroll into view
{
  "url": "https://example.com/listings",
  "actions": [
    { "type": "wait", "selector": ".card", "timeout": 20000 },
    { "type": "scroll", "direction": "down" },
    { "type": "wait", "milliseconds": 1500 },
    { "type": "scroll", "direction": "down" },
    { "type": "wait", "milliseconds": 1500 }
  ]
}

Waiting on a selector rather than a fixed duration continues as soon as the content is there, so prefer it wherever you can name one. Up to 50 actions per request, one screenshot at most, and 30s caps any single wait. The full step vocabulary is in Browser actions.

Returns with actions: a JSON object — data (the final page, Markdown by default), finalUrl, statusCode, and screenshot when the sequence took one. A failed step still returns successfully, carrying error and failedActionIndex (0-based, into your actions array) with data holding the page as it stood at that point.

actions cannot be combined with headers, format: "raw" or solveCaptcha: false; the call fails before anything is sent, with an error naming the field. A browser session is always a GET.

Tool surface vs. the full API

The MCP web_access_fetch tool exposes the most common fetch options. Advanced /fetch features — structured extraction (jsonSchema), requireWSS, and ignoreCertificateErrors — are available on the HTTP API directly.

web_access_request

Send a POST, PUT, or PATCH with a body to a URL you supply, over the same path web_access_fetch uses. Plenty of those calls only read: a GraphQL query, a search backend that takes a body, a JSON API that refuses GET. It is a separate tool because method and body do not exist on web_access_fetch at all: requests with a body get their own schema and their own warning. The split is about request shape, not safety — browser actions make web_access_fetch mutating too.

ParameterTypeDefaultDescription
urlstring—Required. The full http/https URL to send the request to.
methodPOST PUT PATCH—Required. GET is refused here — read a page with web_access_fetch.
bodystring | object—Request body. A string is sent as-is; an object is sent as the exact JSON text you wrote.
formatmarkdown raw jsonjsonDefaults to the envelope: an API caller usually needs the status code, not prose.
headersobject—Custom request headers (max 50), sent in the order you give them. A name given twice is an error.
countryCodestring(2)—ISO 3166-1 alpha-2 country for geolocated proxy routing.
solveCaptchabooleantrueSet false to fail fast instead of solving challenges.

executeJS and actions are not available here — both run a browser session, which is always a GET.

Returns: a { statusCode, headers, data } envelope by default.

Example call
{
  "url": "https://example.com/api/items",
  "method": "POST",
  "body": { "name": "widget" }
}

Search the web from a query and get ranked organic results back, plus every surface Google rendered around them. Set format to raw for the Google results page as HTML, one page per call.

ParameterTypeDescription
querystringRequired. The search query. Be specific for best results.
searchCountintegerOptional. How many organic results you want, 1 to 300; above 300 is rejected. Google is paged, up to 36 pages, until that many are in hand or it has no more, and each page is billed as one search. Omit it for one page. Rejected with format: "raw".
aiModebooleanOptional. true asks Google AI Mode instead of the results page: the response carries aiMode, the answer with its cited sources and any products, places and videos it shows, and results is empty. Billed as one search. Not combinable with searchCount.
aiOverviewbooleanOptional, off by default. true makes a second, sequential fetch to fill the AI Overview Google streams in after the page loads. Filling a streamed overview adds a few seconds to the request. Costs 2x a normal search. Without it, an overview already present in the page is still returned, and a streamed one comes back as declined: true. Not combinable with aiMode.
countrystringOptional. ISO 3166-1 alpha-2 code the search runs from, such as GB. Default US.
languagestringOptional. Results language tag such as en or pt-br.
locationstringOptional. A place name such as London or Austin,Texas,United States, 1 to 200 characters; the search is sent from its country. Google results are biased toward the place, most strongly for queries like plumbers near me, and may still cover a wider area; with aiMode the answer is given for that place. A name that cannot be placed is rejected; send coordinates instead.
coordinatesobjectOptional. A point to search from, as { latitude, longitude, radius } with radius in meters (default 5000), sent from the country it lies in. Google results are biased toward it; with aiMode the answer is given for it. If location is also sent, coordinates win.
pageintegerOptional. Google results page to start from, 1 to 30, where page N is the page Google shows as N; default 1. Without searchCount the response is that one page, and with format: "raw" it always is; with searchCount, results are collected starting from that page. page and searchCount together stay within the first 300 results. position stays 1-based within the response; rank is the result's Google rank. A page past the last result returns zeroResults: true. A page holds about eight to ten results, so separate page calls can repeat or skip a result; send one call with searchCount for a list without repeats.
dateRangestring or objectOptional. Limit results to a publication window: hour, day, week, month or year for the past hour through the past year, or { "from": "2024-01-01", "to": "2024-06-30" } with ISO dates, inclusive. A custom range needs at least one end, and from must not be after to.
sortBystringOptional. relevance (the default) or date for the newest results first.
formatstringOptional. structured (the default) for results as JSON, or raw for the Google results page as HTML, one page per call. We recommend structured. Raw supports page only: searchCount is rejected with raw. Raw is web only. dateRange and sortBy work with raw. See below for the size limit.
searchTypestringOptional. The Google tab to search: web (the default), images, videos, shopping, books, places or forums. images answers images (3 pages of about 100), shopping answers products (one page of about 55, so page must be 1), places answers places (20 a page), and the others answer results. See Search type.
safeSearchbooleanOptional. true removes explicit results.
includeOmittedResultsbooleanOptional. true includes the results Google hides as very similar to ones already shown.
autocorrectbooleanOptional. false searches the query exactly as typed rather than Google's corrected spelling.
restrictCountrystringOptional. A two-letter code; returns only pages from that country. Unlike country, it does not change where the search runs from.
verbatimbooleanOptional. true matches the query's words exactly, without synonyms. web, videos and forums only.

page, dateRange, sortBy, format, searchType and the filters apply to Google results only and are rejected with aiMode. dateRange, sortBy and verbatim apply to the web, videos and forums tabs; format: "raw" and aiOverview to web only.

With format: "raw", the response is html, the Google results page as HTML, in place of results and the surfaces, with htmlBytes and htmlTruncated. Ask for page N with page, one call per page, each billed as one search; a raw call with searchCount is rejected. See Raw HTML.

The tool sends at most 60,000 bytes of HTML per call. A page cut short has htmlTruncated: true, and htmlBytes is the size of the whole page. For the whole page, call POST /search with "format": "raw". Prefer the default structured results unless you need HTML.

Returns: results, the ranked organic documents, each with position (its place in this response), rank (its Google rank; Google only), title, url (the link to fetch; absent when the destination is not known, and the result is still returned), snippet, displayUrl (Google's displayed URL line, empty when Google shows none, as on Reddit and YouTube results), displayText (the source line Google shows under the title, verbatim: a URL, engagement counts such as 20+ comments · 3 months ago, or other text), source when Google names the site (such as Reddit · r/buildapc), and a timestamp when the engine dated the result; and zeroResults, true when the engine itself reported that nothing matched. Every other field is a surface Google rendered around the documents. Each is present only when the page carried it, is never merged into results, and only Google returns it — the same shape as POST /search:

FieldWhat it carries
entityThe knowledge panel for the one business or person the query named: title, subtitle, description and its descriptionSource, rating, reviews, website, attributes, profiles, unread.
placesLocal-pack business listings: name, category, rating, reviews, address, phone, hours, url, mapsUrl. With searchType: "places", the whole answer.
productsOnly with searchType: "shopping": product listings with title, productId, price, originalPrice, merchant, moreMerchants, delivery, returns, rating, reviews; no url.
overviewsGoogle's AI overviews: the query's own first (no topic), then the "Things to know" tabs (topic, question); declined: true marks a frame Google did not fill, including a streamed overview when aiOverview was not set. Each has text and cited sources as { title, url }.
peopleAlsoAsk{ question } entries; answers are not on the page.
relatedSearchesQuery strings Google suggests.
answersOne single-purpose widget under localTime, currency, unitConversion, weather, translation, sports or flights.
spelling{ kind, query, asked } — substituted means the results are for the corrected query, suggested means they are for the query as typed.
adsSponsored results, with advertiser, kept out of results.
videos, shortVideosplatform, channel, date, duration where shown.
discussionssite, forum, comments, age, excerpt where shown.
imagesSource pages of the Images block, with source. With searchType: "images", the whole answer, each also with imageUrl, imageWidth, imageHeight and thumbnail.
sitelinksSub-pages under a result, with resultPosition.
aiModeOnly with aiMode: true: Google's AI Mode answer as plain text, as markdown, and as blocks in reading order (paragraph, heading, list, table with header and rows, code); the cited sources as { title, url, snippet, source }, url absent when a citation could not be resolved; and, when the answer shows them, products (title, price, oldPrice, merchant, rating, reviews, url…), places (name, category, rating, reviews, status, address, url…) and videos (title, url, channel, duration). See Google AI Mode.
paging{ pages, complete, stoppedBy }, only when searchCount was sent. pages is how many results pages answered, each billed as one search. stoppedBy is why paging stopped: search_count (the count was reached), end_of_results (Google had no more, or the first page carried no organic results), page_cap (the 36-page limit; Google may have more), page_failed (a later page could not be fetched) or deadline (the time budget ran out). complete is false only for page_failed and deadline, and results then holds what was collected. Surfaces describe the first page only; positions run on across pages.

A query naming a single business often comes back with results empty and the answer in entity or places, so read those before treating an empty results as no answer. A snippet is the engine's description of the page, not the page itself, and an overview is a summary, not a source — fetch the url with web_access_fetch to read it.

Example call
{ "query": "latest developments in AI agents 2026" }
Thirty results, paged
{ "query": "construction consulting firms Ohio", "searchCount": 30 }
Google's results page as HTML
{ "query": "heat pump grants", "format": "raw" }
Results biased toward Austin
{ "query": "plumbers near me", "location": "Austin,Texas,United States" }
Google AI Mode answer
{ "query": "how does a heat pump work", "aiMode": true }
Google AI Mode answer for London
{ "query": "what is the weather today", "aiMode": true, "location": "London" }
Google AI Mode answer for a point in Paris
{ "query": "best bakeries near me", "aiMode": true, "coordinates": { "latitude": 48.8566, "longitude": 2.3522 } }

web_access_sitemap

Crawl a whole site and map its URLs. Starting from one URL the crawler follows same-domain links breadth-first (optionally seeded from the site's /sitemap.xml) and records every URL it reaches with fetch status, depth, and parent. Crawls run asynchronously as sitemap jobs: one tool drives the whole quote → approve → poll → read lifecycle through its action parameter, and nothing is crawled or billed until the quote is explicitly approved.

actionRequiresDescription
submiturlQuote a crawl. Returns jobId, estimatedPages, and estimatedCostUsd with status awaiting_approval. Quotes expire after one hour.
approvejobIdStart the quoted crawl — the billing-consent step. Fails with 402 on insufficient balance; retry on a 409 partial_state.
statusjobIdPoll progress: awaiting_approval → running (pending / processed counts) → completed | failed | canceled | token_cap_exceeded.
resultsjobIdPage through discovered URLs. total tells you when to stop paging.
canceljobIdStop a non-terminal job. Pages already fetched stay billed and readable.
list—Your organization's recent crawl jobs, most recent first.
ParameterTypeUsed byDescription
actionenum—Required. One of the six actions above.
urlstring (URL)submitRequired for submit. Starting URL. The crawl stays on this hostname.
maxPagesintegersubmitMaximum pages to crawl (1–10,000, default 10 when omitted). Each fetched page is billed.
maxDepthintegersubmitMaximum link depth from the starting URL (1–100, default 2 when omitted).
pathPrefixstringsubmitOnly crawl URLs whose path starts with this prefix.
budgetUsdnumbersubmitSpend cap in USD (min 0.0001). Defaults to the quote; the crawl stops with token_cap_exceeded before exceeding it.
useSitemapbooleansubmitAlso seed the crawl from the site's root /sitemap.xml (one extra billed page; default false).
jobIdstringjob-scoped actionsThe id returned by submit.
limitintegerresults / listPage size — results 1–5,000, default 1000; list 1–100, default 20.
offsetintegerresults / listNumber of entries to skip, 0 or more.

A bound outside its range, 0 included, is an error rather than being replaced by the default. Omit the field to get the default.

Returns: the JSON envelope for the chosen action (quote, status, URL page, or job list) alongside a one-line summary.

Example call
{ "action": "submit", "url": "https://example.com", "maxPages": 200, "maxDepth": 3 }

Same jobs as the HTTP API

The tool fronts the sitemap endpoints directly — response shapes, statuses, billing, and result retention (per-URL discoveredUrls is only available while results are live) all match. See the Sitemap guide for a walkthrough.

  1. Call web_access_search to find relevant pages for a query.
  2. Call web_access_fetch on the most relevant URLs to pull their full content.

To cover a whole site instead of single pages, run web_access_sitemap first (submit → approve → poll → results) to build the URL list, then web_access_fetch the pages you need.