# Scrape

The fastest way to get several extractions from the same URL: one request, one page visit. Every operation shares a single page load — the slowest part of any analysis — instead of each paying for its own, and the AI extractions run in parallel on top of it.

**Speed is the whole point of this endpoint.** The results themselves are identical to the individual endpoints', and billing is unchanged too: the request spends the sum of what its operations [cost](https://urlpipe.dev/docs/credits), and operations served from cache stay free. You pay the same — you just get everything much sooner.

There is one exception, and it is in your favour. A [residential exit](https://urlpipe.dev/docs/residential) is charged per page fetch, and a scrape fetches once, so `residential: true` costs one exit here where five separate calls would cost five. (`lighthouse` runs its own audit on its own engine, so a scrape including it fetches twice, and so does a screenshot with an option that changes how the page loads — see [How it works](#how).)

POST/scrape

## Run multiple operations at once

### Body parameters

- **Name**
  : `url`
  **Type**
  : string
  **Required**
  : Required
  **Description**
  : The absolute URL of the page to process. Rendered with headless Chrome, so JavaScript runs and redirects are followed. It must not include a username or password (`https://user:pass@example.com`).
- **Name**
  : `operations`
  **Type**
  : array
  **Required**
  : Required
  **Description**
  : The operations to run — any non-empty subset of `html`, `markdown`, `meta`, `summarize`, `keywords`, `screenshot`, `console` and `lighthouse`. A JSON array (`["markdown","meta"]`) or a comma-separated string (`markdown,meta`). An unknown operation returns `422`.
- **Name**
  : `page_options`
  **Type**
  : object
  **Description**
  : Wait for the page, and remove ads, cookie banners or your own elements from it before it is read — gone from this result, not merely hidden. See [Page options](https://urlpipe.dev/docs/page-options).
- **Name**
  : `residential`
  **Type**
  : boolean
  **Description**
  : Fetch the page from a [residential exit](https://urlpipe.dev/docs/residential) — an address on a home broadband line rather than one in a datacentre. Reach for it when a site serves you less than it serves a browser, or nothing at all. Defaults to `false`. Adds **25 credits** per page fetch on top of what the operation costs, and its results are kept separate from the ordinary ones.
- **Name**
  : `report_to`
  **Type**
  : string
  **Description**
  : Webhook URL — an `http` or `https` address URLpipe POSTs the result to when it's ready. **Optional**: without it we deliver to your project's [default endpoint](https://urlpipe.dev/docs/async#where-results-go) if it has one, and otherwise send no webhook at all — the result still waits for you at [GET /result/:token](https://urlpipe.dev/docs/results). A value we cannot deliver to returns `422`. Ignored on a `sync=true` request. Deliveries can be [signed](https://urlpipe.dev/docs/async#signature) so your endpoint can verify they came from us.
- **Name**
  : `sync`
  **Type**
  : boolean
  **Description**
  : Process the request synchronously, returning the result inline in the response. Defaults to `false` (async: return a token now, and either receive the result at a [webhook](https://urlpipe.dev/docs/async#where-results-go) or fetch it with [GET /result/:token](https://urlpipe.dev/docs/results)). See [Async & sync modes](https://urlpipe.dev/docs/async) for the full contract.
- **Name**
  : `max_age`
  **Type**
  : string | integer
  **Description**
  : How fresh a cached result must be to be accepted. Either an **integer** number of seconds (`3600`) or a **duration string** of the form `"<number> <unit>"` — units `s`/`min`/`h`/`d`/`w` (e.g. `"2 hours"`, `"3 days"`, `"30m"`). Defaults to `7 days`, clamped to a max of `30 days`; `0` always bypasses the cache. See [Caching](https://urlpipe.dev/docs/caching) for all accepted units.
- **Name**
  : `labels`
  **Type**
  : object
  **Description**
  : Your own keys to find and account for this request by — a client, a project, a campaign: `{"client": "acme"}`. Returned with the result, in the webhook and in the `X-Labels` header, and your dashboard filters history and totals credits by them. Up to 16 keys; string values. See [Labels](https://urlpipe.dev/docs/labels).
- **Name**
  : `include_audits`
  **Type**
  : boolean
  **Description**
  : Only used when `operations` includes `lighthouse`: include the detailed audits section in its result. Defaults to `false`.
- **Name**
  : `device`
  **Type**
  : string
  **Description**
  : Only used when `operations` includes `lighthouse`: the device profile to audit with, `mobile` (default) or `desktop`.
- **Name**
  : `screenshot_options`
  **Type**
  : object
  **Description**
  : Only used when `operations` includes `screenshot`: how to take it, with the same [keys](https://urlpipe.dev/docs/screenshot#options) `/screenshot` takes.

### Response

Content type `application/json` — one entry per requested operation, each with its own `success` flag. A failing operation (for example a page too large for an AI extraction) never affects its siblings; each entry carries either its `result` in that operation's usual format or its `error`. `cached` tells you when an operation was served from cache (free).

```
curl -X POST https://urlpipe.dev/scrape \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "operations": ["markdown","meta","screenshot"], "sync": "true"}'
```

Response

```
{
  "url": "https://example.com",
  "final_url": "https://example.com/",
  "operations": {
    "markdown": {
      "success": true,
      "result": "# Example Domain\n\nThis domain is for use in…",
      "cached": false
    },
    "meta": {
      "success": true,
      "result": { "title": "Example Domain", "description": "…", "language": "en" },
      "cached": true
    },
    "screenshot": {
      "success": true,
      "result": "iVBORw0KGgoAAAANSUhEUg…",
      "cached": false
    }
  }
}
```

## How it works

URLpipe loads the page once in a headless browser and captures, from that single visit, everything the requested operations need: the rendered HTML (which feeds `html`, `markdown`, `meta`, `summarize` and `keywords`), the `screenshot` and the `console` messages. The AI extractions then run in parallel. Every result is identical to what the individual endpoint would return, and is cached under the same key — a scrape can be served by earlier individual calls, and later individual calls can be served by a scrape.

The one exception is `lighthouse`: a performance audit needs its own instrumented page load, so it runs alongside the unified visit and its result is merged into the combined response when it finishes.

A screenshot shares the visit too, with most of its `screenshot_options`: a format, a selector and hidden elements all apply to the image alone. Four change how the page itself loads — `viewport_width`, `viewport_height`, `device_scale_factor` and `dark_mode` — so a screenshot with any of them loads the page again on its own, and the `html` and `markdown` beside it stay the page as it normally renders.

[page\_options](https://urlpipe.dev/docs/page-options) apply to the shared visit, and so to every operation of the scrape except `lighthouse`, whose audit loads the page for itself.

Including `lighthouse`? Prefer **async mode** — audits regularly outlast the 60-second synchronous window, in which case the sync response is a `504` with a token to fetch later. Also note that a combined response embeds the screenshot as Base64, so webhook payloads can be several megabytes.

## Partial failures

Operations fail independently: one failed extraction leaves the others intact, and the response is still `200 OK`. Only when every operation fails — typically because the page itself was unreachable — is the whole request a `422`. Failed operations never spend credits. If a lighthouse audit is still running when the rest of the scrape finishes, its entry reports `processing_timeout`. It keeps running — fetch the scrape again from [GET /result/:token](https://urlpipe.dev/docs/results) with this request's token and the entry will be filled in once it lands. The individual operations have no tokens of their own; the scrape is addressed as a whole.

## Responses

Whatever the status, the response carries [metadata headers](https://urlpipe.dev/docs/response-headers): the result token, whether it was served from cache and how old that result is, how long we took, what it cost in credits, and the allowance you have left.

Status

When

Body

`200 OK`

The request succeeded.

Sync: the result, in this endpoint's format (see Response above). Async: a job token and the request's labels — { "token": "…", "status": "accepted", "labels": {} }.

`422 Unprocessable Entity`

The analysis failed, or a parameter was invalid (a bad max\_age, a screenshot\_options value out of range, labels that break the rules, or a report\_to we will not deliver to).

{ "error": "<message>" }

`429 Too Many Requests`

Three causes, told apart by the error field: concurrency\_limit (too many of your requests already running), rate\_limited (sending too fast), or quota\_exceeded (Free plan only — out of credits with no card on file to bill the extra to, and checked only on a cache miss; a paid plan keeps serving at the overage rate and is never refused for credits).

All three carry "error" and "message". Extra fields: concurrency\_limit → "limit", "running" · rate\_limited → "retry\_after" · quota\_exceeded → "limit", "used", "needed", "resets\_at".

`504 Gateway Timeout`

Sync only: the analysis didn't finish within 60s. It keeps running — fetch it via GET /result/:token.

{ "error": "processing\_timeout", "token": "…" }

`401 Unauthorized`

Missing or invalid API key.

{ "error": "invalid\_api\_key", "message": "…" }

Try it live — no API key needed

Run this endpoint against any URL right in your browser.

[Open tool](https://urlpipe.dev/tools/analyze-url)

Source: https://urlpipe.dev/docs/scrape
