LangChain
Give a LangChain agent web search and clean Markdown page fetches through the Web Access API.
The langchain-string package wraps the Web Access API as two LangChain tools. StringSearch runs a search and
returns the results as JSON; StringFetch fetches a URL and returns the page as Markdown, with proxy rotation,
anti-bot handling, CAPTCHA solving and JavaScript rendering happening server-side.
Both are ordinary BaseTool implementations with sync and async paths, so they work in any LangChain agent.
Install
pip install -U langchain-stringSet your API key, or pass api_key=SecretStr("...") to either tool:
export STRING_API_KEY=...Search
from langchain_string import StringSearch
search = StringSearch(max_results=5)
print(search.invoke({"query": "who makes the best web scraping api"}))The tool returns the search response as JSON: results holds the ranked pages, and a
Google search can also carry the knowledge panel, local listings, AI overviews and other surfaces the results page
showed around them. A query answered by a panel alone therefore still has an answer, even though results is empty.
The agent can set query, engine, country and language per call.
| Field | Default | Description |
|---|---|---|
query | — | Required. The search query. Quoted phrases and site: pass through to the engine. |
engine | google | google, duckduckgo, brave or mojeek. Only Google returns the surfaces above. |
country | US | ISO 3166-1 alpha-2 country code used to localize the results. |
language | — | Language tag such as en or pt-br. |
max_results | — | Constructor only. Organic results wanted, 1 to 50. See below. |
max_results is a billing decision, which is why the agent cannot set it
Without it, a search returns one results page, about ten results, billed as one search. With it, Google is paged until it has that many, and each page is billed as one search.
Fetch
from langchain_string import StringFetch
fetch = StringFetch(main_content_only=True)
print(fetch.invoke({"url": "https://example.com"}))The agent can set url, and execute_js when a JavaScript-rendered page comes back empty, and country_code to
route the request through a specific country. Everything else is set when you construct the tool.
| Field | Default | Description |
|---|---|---|
url | — | Required. The http/https URL to fetch. |
execute_js | false | Render the page in a browser before capturing it. |
country_code | — | ISO 3166-1 alpha-2 country to route the request through. |
markdown_mode | full | Constructor only. full or readable. See Markdown for LLMs. |
main_content_only | false | Constructor only. Strip navigation, footers and other page chrome. |
max_content_length | — | Constructor only. Truncate the page to this many characters. |
When the site answers with an error status, the output says so above the page it returned, so the agent can tell a 404 page from an article.
Both tools in an agent
StringWebAccessToolkit returns both tools and passes its configuration to each.
from langchain.agents import create_agent
from langchain_string import StringWebAccessToolkit
agent = create_agent(
model="anthropic:claude-sonnet-5",
tools=StringWebAccessToolkit(max_results=5).get_tools(),
system_prompt=(
"You research questions using the live web. Search first, then read the pages "
"worth reading, and cite every URL you used."
),
)
result = agent.invoke(
{"messages": [{"role": "user", "content": "What does String charge for a CAPTCHA solve?"}]}
)
print(result["messages"][-1].content)Failures reach the model
API and network failures are returned to the model as the tool's output, carrying the reason the API gave: the
rejected field, the status the destination answered with, or an account problem such as an
insufficient balance. The model can then change its call instead of the run ending. Pass
handle_tool_error=False to raise instead.
A missing API key is not treated this way. It raises before any request, because no retry fixes it.
Shared options
| Field | Default | Description |
|---|---|---|
api_key | STRING_API_KEY | Held as a SecretStr, so it stays out of reprs and traces. |
base_url | https://request.usestring.ai/v1 | Change only to point at a different String deployment. |
timeout | 600.0 | Seconds. A browser-rendered fetch of a slow page can take minutes. |
http_client / http_async_client | — | Your own httpx clients, for proxies or custom transports. |
See also
- POST /search and POST /fetch for the underlying API.
- MCP server for the same capabilities over MCP, in any MCP client.