The call
GET /api/fetch reads one page live and returns it as Markdown: its text, headings, links and tables, without scripts, styling, navigation, headers or footers. 1 credit.
| Name | Type | What it does |
|---|---|---|
url | string | Required. The page; https:// is assumed. Private and internal addresses are refused. |
all | 1 | Optional. Keep navigation, headers and footers. |
max_age | integer | Optional. Accept a copy we fetched up to this many seconds ago (default 3600; 0 = always live). |
max_chars | integer | Optional. Cut the Markdown at this many characters (default 100,000). |
Shell
curl -G https://stackgrep.com/api/fetch -H "Authorization: Bearer $STACKGREP_KEY" \
--data-urlencode 'url=https://example.com/pricing'Response
{"url": "https://example.com/pricing", "final_url": "https://www.example.com/pricing",
"status": 200, "title": "Pricing - Example", "description": "Plans for teams of every size.",
"markdown": "# Pricing\n\n| Plan | Price |\n|---|---|\n| Team | $29/mo |\n...",
"truncated": false, "total_chars": 4210, "rendered": false, "blocked": "",
"cached": false, "age_s": 0, "ms": 448}| Name | Type | What it does |
|---|---|---|
status | integer | The page's HTTP status (0 when it couldn't be reached; description then says why). |
rendered | boolean | We ran the page in a browser because its content needs JavaScript. Most pages don't. |
blocked | string | The bot protection that answered instead of the page, if one did. We don't try to get around it. |
truncated | boolean | The Markdown was cut at max_chars; total_chars says how long it was. Never cut silently. |
cached | boolean | It came from a copy fetched age_s seconds ago. |
ms | integer | Time taken, in milliseconds. |
We respect robots.txt and follow redirects, and remember both per site, so repeat fetches from a site are faster. A fetch gives up after 15 seconds.
Try it on your own documentsExact and regex search for your agents, from one API. Or see the engine on npm's source first.