String API
Guides

Markdown for LLMs

Turn web pages into clean, token-efficient Markdown for AI pipelines.

When feeding web pages into an LLM, raw HTML wastes tokens and confuses the model. Set format: markdown and the Web Access API converts the page to clean Markdown. The default full mode preserves the complete cleaned page; choose readable when a smaller response matters more than completeness.

curl https://request.usestring.ai/v1/fetch \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "url": "https://example.com/article", "format": "markdown" }'

What you get

  • Full cleaned page content in document order, including navigation, sidebars, and footers.
  • Scripts, styles, tracking pixels, and other non-content markup removed.
  • Headings, lists, links, tables, code blocks, and inline formatting preserved.
  • Relative URLs rewritten against the page origin.
  • Embedded structured data preserved as fenced JSON code blocks.

Choose how much to preserve

markdownMode: "full" is the default and favors completeness over response size. Set markdownMode: "readable" to distill the page with Readability, which removes boilerplate and uses fewer tokens at the cost of content that falls outside the detected article.

curl https://request.usestring.ai/v1/fetch \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/article",
    "format": "markdown",
    "markdownMode": "readable"
  }'

Set mainContentOnly: true for the opposite tradeoff: remove navigation, headers, sidebars, and footers. You can use it with either mode; with full, it converts the complete remaining main content without Readability distillation.

If you need specific fields rather than the whole page, use structured extraction instead.