# Metadata

Get everything a page declares about itself as one JSON object: title, description, site name, authors, publication dates, share image, favicon and icons, feeds, language, canonical URL and more. URLpipe reads Open Graph and Twitter card tags, standard meta tags, JSON-LD structured data, microdata and link tags, and reconciles them for you.

The object has the same keys on every page. A value the page does not declare is `null` — or an empty list — never a guess, so your code handles one shape. The page is rendered with headless Chrome first, so values set by JavaScript are included.

It uses no AI: the page's own declarations are read out of the rendered DOM, so the same page always gives the same answer, and it costs **1 credit** per call, the same as [/html](https://urlpipe.dev/docs/html). See [Credits](https://urlpipe.dev/docs/credits).

POST/meta

## Extract metadata

### Body parameters

- **Name**
  : `url`
  **Type**
  : string
  **Required**
  : Required
  **Description**
  : The absolute URL of the page to process. Rendered with headless Chrome, so JavaScript runs and redirects are followed. It must not include a username or password (`https://user:pass@example.com`).
- **Name**
  : `page_options`
  **Type**
  : object
  **Description**
  : Wait for the page, and remove ads, cookie banners or your own elements from it before it is read — gone from this result, not merely hidden. See [Page options](https://urlpipe.dev/docs/page-options).
- **Name**
  : `residential`
  **Type**
  : boolean
  **Description**
  : Fetch the page from a [residential exit](https://urlpipe.dev/docs/residential) — an address on a home broadband line rather than one in a datacentre. Reach for it when a site serves you less than it serves a browser, or nothing at all. Defaults to `false`. Adds **25 credits** per page fetch on top of what the operation costs, and its results are kept separate from the ordinary ones.
- **Name**
  : `report_to`
  **Type**
  : string
  **Description**
  : Webhook URL — an `http` or `https` address URLpipe POSTs the result to when it's ready. **Optional**: without it we deliver to your project's [default endpoint](https://urlpipe.dev/docs/async#where-results-go) if it has one, and otherwise send no webhook at all — the result still waits for you at [GET /result/:token](https://urlpipe.dev/docs/results). A value we cannot deliver to returns `422`. Ignored on a `sync=true` request. Deliveries can be [signed](https://urlpipe.dev/docs/async#signature) so your endpoint can verify they came from us.
- **Name**
  : `sync`
  **Type**
  : boolean
  **Description**
  : Process the request synchronously, returning the result inline in the response. Defaults to `false` (async: return a token now, and either receive the result at a [webhook](https://urlpipe.dev/docs/async#where-results-go) or fetch it with [GET /result/:token](https://urlpipe.dev/docs/results)). See [Async & sync modes](https://urlpipe.dev/docs/async) for the full contract.
- **Name**
  : `max_age`
  **Type**
  : string | integer
  **Description**
  : How fresh a cached result must be to be accepted. Either an **integer** number of seconds (`3600`) or a **duration string** of the form `"<number> <unit>"` — units `s`/`min`/`h`/`d`/`w` (e.g. `"2 hours"`, `"3 days"`, `"30m"`). Defaults to `7 days`, clamped to a max of `30 days`; `0` always bypasses the cache. See [Caching](https://urlpipe.dev/docs/caching) for all accepted units.
- **Name**
  : `labels`
  **Type**
  : object
  **Description**
  : Your own keys to find and account for this request by — a client, a project, a campaign: `{"client": "acme"}`. Returned with the result, in the webhook and in the `X-Labels` header, and your dashboard filters history and totals credits by them. Up to 16 keys; string values. See [Labels](https://urlpipe.dev/docs/labels).

```
curl -X POST https://urlpipe.dev/meta \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com"}'
```

Response

```
{
  "final_url": "https://example.com/blog/launch-week",
  "canonical_url": "https://example.com/blog/launch-week",
  "title": "Everything we shipped in launch week",
  "description": "Five days, five releases: what's new and why it matters.",
  "site_name": "Example Blog",
  "type": "article",
  "language": "en",
  "locale": "en-US",
  "authors": [
    {
      "name": "Jane Doe",
      "url": "https://example.com/authors/jane",
      "profiles": [
        "https://x.com/janedoe"
      ]
    }
  ],
  "published_at": "2026-09-01T08:00:00+00:00",
  "modified_at": "2026-09-03T14:30:00+00:00",
  "image": {
    "url": "https://example.com/images/launch-week.jpg",
    "width": 1200,
    "height": 630,
    "alt": "The team on stage",
    "type": "image/jpeg"
  },
  "video": null,
  "favicon_url": "https://example.com/favicon.ico",
  "icons": [
    {
      "url": "https://example.com/favicon.ico",
      "rel": "icon",
      "sizes": "32x32",
      "type": null
    },
    {
      "url": "https://example.com/apple-touch-icon.png",
      "rel": "apple-touch-icon",
      "sizes": "180x180",
      "type": null
    }
  ],
  "logo_url": "https://example.com/logo.svg",
  "feeds": [
    {
      "url": "https://example.com/feed.xml",
      "type": "application/rss+xml",
      "title": "Example Blog"
    }
  ],
  "keywords": [
    "launch",
    "product"
  ],
  "theme_color": "#0f172a",
  "generator": "WordPress 6.6",
  "manifest_url": "https://example.com/site.webmanifest",
  "robots": {
    "index": true,
    "follow": true
  },
  "open_graph": {
    "title": "Everything we shipped in launch week",
    "description": "Five days, five releases.",
    "image": "https://example.com/images/launch-week.jpg",
    "url": "https://example.com/blog/launch-week",
    "type": "article",
    "site_name": "Example Blog",
    "locale": "en_US"
  },
  "twitter": {
    "card": "summary_large_image",
    "site": "@example",
    "creator": "@janedoe",
    "title": null,
    "description": null,
    "image": null
  },
  "alternates": [
    {
      "hreflang": "es",
      "url": "https://example.com/es/blog/launch-week"
    }
  ],
  "oembed_url": "https://example.com/wp-json/oembed/1.0/embed?url=https%3A%2F%2Fexample.com%2Fblog%2Flaunch-week",
  "structured_data_types": [
    "BlogPosting",
    "Organization"
  ]
}
```

## Response fields

Every URL comes back absolute, resolved against the page, and dates are ISO 8601 — a date alone (`2026-09-01`) when that is all the page gives, a moment with its offset otherwise.

### The page

- **Name**
  : `final_url`
  **Type**
  : string | null
  **Description**
  : The address the page ended on after its redirects — the same one the `X-Final-Url` header carries. Relative URLs in the page are resolved against it.
- **Name**
  : `title`
  **Type**
  : string | null
  **Description**
  : The page's own title, without the site's name around it.
- **Name**
  : `description`
  **Type**
  : string | null
  **Description**
  : The page description.
- **Name**
  : `site_name`
  **Type**
  : string | null
  **Description**
  : The name of the site the page belongs to.
- **Name**
  : `type`
  **Type**
  : string | null
  **Description**
  : The kind of page, as the page declares it: `article`, `website`, `profile`, `book`, `product`, `video`, `music`, `place`, `event`, `recipe`. A site's homepage is always `website`.
- **Name**
  : `canonical_url`
  **Type**
  : string | null
  **Description**
  : The canonical URL the page declares.
- **Name**
  : `language`
  **Type**
  : string | null
  **Description**
  : ISO 639-1 code of the language the content is written in, e.g. `en`. When the page declares none, or declares one its text plainly contradicts, it is read from the text.
- **Name**
  : `locale`
  **Type**
  : string | null
  **Description**
  : The most specific language tag the page declares for that language, e.g. `en-US`.
- **Name**
  : `keywords`
  **Type**
  : string\[\]
  **Description**
  : Keywords and tags, from meta tags, article tags and structured data.

### Who and when

- **Name**
  : `authors`
  **Type**
  : object\[\]
  **Description**
  : Each with `name`, `url` (their page, or null) and `profiles` (other profiles the page links them to).
- **Name**
  : `published_at`
  **Type**
  : string | null
  **Description**
  : When the page was first published.
- **Name**
  : `modified_at`
  **Type**
  : string | null
  **Description**
  : When it was last updated.

### Images and icons

- **Name**
  : `image`
  **Type**
  : object | null
  **Description**
  : The share image: `url`, and the `width`, `height`, `alt` and `type` the page declares for it.
- **Name**
  : `video`
  **Type**
  : object | null
  **Description**
  : `url`, `type`, `width` and `height` of the page's video, when it has one.
- **Name**
  : `favicon_url`
  **Type**
  : string | null
  **Description**
  : The icon a browser tab shows.
- **Name**
  : `icons`
  **Type**
  : object\[\]
  **Description**
  : Every icon the page declares, with its `rel`, `sizes` and `type`.
- **Name**
  : `logo_url`
  **Type**
  : string | null
  **Description**
  : The site's or publisher's logo.
- **Name**
  : `theme_color`
  **Type**
  : string | null
  **Description**
  : The browser theme colour.

### Discovery

- **Name**
  : `feeds`
  **Type**
  : object\[\]
  **Description**
  : RSS, Atom and JSON feeds, with their `type` and `title`.
- **Name**
  : `alternates`
  **Type**
  : object\[\]
  **Description**
  : The page in other languages: `hreflang` and `url`.
- **Name**
  : `oembed_url`
  **Type**
  : string | null
  **Description**
  : The page's oEmbed endpoint, for embedding it.
- **Name**
  : `manifest_url`
  **Type**
  : string | null
  **Description**
  : The web app manifest.
- **Name**
  : `robots`
  **Type**
  : object
  **Description**
  : `index` and `follow`: whether the page lets search engines index it and follow its links.
- **Name**
  : `generator`
  **Type**
  : string | null
  **Description**
  : The software that built the page, as it declares it.
- **Name**
  : `structured_data_types`
  **Type**
  : string\[\]
  **Description**
  : The schema.org types the page's JSON-LD and microdata describe, e.g. `Article`.

### Raw tags

- **Name**
  : `open_graph`
  **Type**
  : object
  **Description**
  : The page's Open Graph tags exactly as declared: `title`, `description`, `image`, `url`, `type`, `site_name`, `locale`. Useful for checking what a share card will show.
- **Name**
  : `twitter`
  **Type**
  : object
  **Description**
  : The page's Twitter card tags exactly as declared: `card`, `site`, `creator`, `title`, `description`, `image`.

## Responses

Whatever the status, the response carries [metadata headers](https://urlpipe.dev/docs/response-headers): the result token, whether it was served from cache and how old that result is, how long we took, what it cost in credits, and the allowance you have left.

Status

When

Body

`200 OK`

The request succeeded.

Sync: the result, in this endpoint's format (see Response above). Async: a job token and the request's labels — { "token": "…", "status": "accepted", "labels": {} }.

`422 Unprocessable Entity`

The analysis failed, or a parameter was invalid (a bad max\_age, a screenshot\_options value out of range, labels that break the rules, or a report\_to we will not deliver to).

{ "error": "<message>" }

`429 Too Many Requests`

Three causes, told apart by the error field: concurrency\_limit (too many of your requests already running), rate\_limited (sending too fast), or quota\_exceeded (Free plan only — out of credits with no card on file to bill the extra to, and checked only on a cache miss; a paid plan keeps serving at the overage rate and is never refused for credits).

All three carry "error" and "message". Extra fields: concurrency\_limit → "limit", "running" · rate\_limited → "retry\_after" · quota\_exceeded → "limit", "used", "needed", "resets\_at".

`504 Gateway Timeout`

Sync only: the analysis didn't finish within 60s. It keeps running — fetch it via GET /result/:token.

{ "error": "processing\_timeout", "token": "…" }

`401 Unauthorized`

Missing or invalid API key.

{ "error": "invalid\_api\_key", "message": "…" }

Try it live — no API key needed

Run this endpoint against any URL right in your browser.

[Open tool](https://urlpipe.dev/tools/metadata-extractor)

Source: https://urlpipe.dev/docs/meta
