Python · Code recipe
Get a page's metadata and Open Graph tags in Python
Read the title, description and share image of any page, in Python 3.9+ with requests. Every program on this page runs as it stands — each one was run against a stub of the API before it was published.
By Roger Campos · Last updated: September 2026
TL;DR
To read a page's title, description, Open Graph image and other metadata in Python, POST its URL to https://urlpipe.dev/meta with sync set to true and parse the JSON with res.json(). It answers with nine fields, any of which can be null, for 5 credits: it is one of the three endpoints that call a language model.
Free plan, no credit card. 1,000 credits a month.
One POST to /meta renders the page and returns what it says about itself: title, description, language, main image, favicon, author, feed, first publication date and extra author details. URLs come back absolute, resolved against the page.
A language model reads the page's metadata declarations — Open Graph, Twitter cards, JSON-LD, plain tags — and settles conflicts by a fixed order: og:title, then twitter:title, then <title>, then the <h1>. A field the page never declares is null, never a guess. That is why it costs 5 credits where a page fetch costs 1 credit. In Python, res.json() gives you the fields; the Open Graph guide covers which tags each platform reads.
Setup
Before you start
One dependency, requests. Put your API key in the environment so it never lands in the file.
python3 -m pip install requests
export URLPIPE_API_KEY="your_api_key"
The request
Print the title, description and main image
"sync": True holds the connection open until the result is ready and answers with it. Pass timeout= every time: requests has no default, and a page that never finishes loading would otherwise hang your process. res.json() gives a dict; any field can be None, hence the or "none".
import os
import requests
res = requests.post(
"https://urlpipe.dev/meta",
headers={"Authorization": f"Bearer {os.environ['URLPIPE_API_KEY']}"},
json={"url": "https://example.com", "sync": True},
# requests waits forever unless told otherwise; a sync call can take up to 60 s.
timeout=90,
)
res.raise_for_status()
meta = res.json()
print("Title:", meta["title"] or "none")
print("Description:", meta["description"] or "none")
print("Image:", meta["main_image_url"] or "none")
Run it: python3 page_metadata.py
Given the example response on the docs page, it prints:
Title: Example Domain
Description: Illustrative examples in documents.
Image: https://example.com/cover.jpgAsync
The async variant: a token, a webhook and a poll
Leave out sync and the answer is a token, straight away. for … else is the idiom for "ran out of attempts": the else runs only when the loop never hit break.
import os
import time
import requests
API = "https://urlpipe.dev"
HEADERS = {"Authorization": f"Bearer {os.environ['URLPIPE_API_KEY']}"}
# No "sync": the request is accepted at once and the work carries on without you.
res = requests.post(
f"{API}/meta",
headers=HEADERS,
json={
"url": "https://example.com",
"report_to": "https://your-app.com/webhooks/urlpipe",
"labels": {"customer": "acme"},
},
timeout=30,
)
res.raise_for_status()
token = res.json()["token"]
print(f"Accepted {token}")
# The result is POSTed to report_to when it is ready. Polling by token is the
# other way to collect it: no endpoint needed, and a backup for the webhook.
for _ in range(60):
res = requests.get(f"{API}/result/{token}", headers=HEADERS, timeout=30)
if res.status_code != 202: # 202 means still processing
break
time.sleep(2)
else:
raise SystemExit("Still processing after two minutes; try the token again later.")
if res.status_code == 422:
raise SystemExit(f"The analysis failed: {res.json()['error']}")
if res.status_code == 410:
raise SystemExit("The result is past the 30-day window; send the request again.")
res.raise_for_status()
meta = res.json()
print("Title:", meta["title"] or "none")
print("Description:", meta["description"] or "none")
print("Image:", meta["main_image_url"] or "none")
Run it: python3 page_metadata_async.py
Errors
Handle errors and retries
raise_for_status() is fine for a script; a service wants to tell the failures apart. Check the status before calling res.json() — a 401 body is not JSON. sys.exit(message) prints to stderr and exits 1.
import os
import sys
import time
import requests
API_KEY = os.environ["URLPIPE_API_KEY"]
class URLpipeError(Exception):
pass
def urlpipe(path, payload, attempts=5):
"""POST a sync request and return the response, or raise URLpipeError."""
for attempt in range(attempts):
res = requests.post(
f"https://urlpipe.dev{path}",
headers={"Authorization": f"Bearer {API_KEY}"},
json={**payload, "sync": True},
timeout=90,
)
if res.status_code == 200:
return res
if res.status_code == 401:
raise URLpipeError("401: the API key is missing or wrong. Check URLPIPE_API_KEY.")
try:
body = res.json()
except ValueError:
body = {}
code = body.get("error", "")
detail = f"{code}: {body['message']}" if body.get("message") else code
if res.status_code == 429 and code == "rate_limited":
# Sending too fast: Retry-After says how long the window has left.
time.sleep(int(res.headers.get("Retry-After", 1)))
elif res.status_code == 429 and code == "concurrency_limit":
# Every parallel slot on your plan is busy with your own requests.
time.sleep(2**attempt)
elif res.status_code == 504:
# Still running on our side; the token collects it from GET /result/:token.
raise URLpipeError(f"504 processing_timeout: collect it later with token {body['token']}")
else:
# 403 email_unverified, 422 (a bad parameter, or a page that would not load),
# 429 quota_exceeded: sending the same request again gets the same answer.
raise URLpipeError(f"{res.status_code}: {detail}")
raise URLpipeError(f"429: still refused after {attempts} attempts")
try:
res = urlpipe("/meta", {"url": "https://example.com"})
except URLpipeError as error:
sys.exit(f"URLpipe: {error}")
except requests.RequestException as error:
sys.exit(f"Network error: {error}")
meta = res.json()
print("Title:", meta["title"] or "none")
print("Description:", meta["description"] or "none")
print("Image:", meta["main_image_url"] or "none")
Run it: python3 page_metadata_errors.py
Details
What to know about /meta
- Any field can be
nullwhen the page does not have it — code for that, as the program does. - There is no
canonicalfield; the nine fields are the whole response. - Image and favicon URLs that are data URIs come back as
nullrather than as a blob. - Pages over 10 MB of HTML are refused before the model sees them.
FAQ
Frequently asked questions
Which fields does /meta return?
Why does metadata cost more than fetching the HTML?
Is AI processing done in the EU?
Do I need an SDK to call URLpipe from Python?
Get a key and run it.
Free plan, no card. Paste your key into URLPIPE_API_KEY and every program on this page runs as it is.