Skip to main content

Confirm

Are you sure?

Python · Code recipe

Get the rendered HTML of a JavaScript page in Python

Fetch the HTML a browser ends up with, JavaScript and all, in Python 3.9+ with requests. Every program on this page runs as it stands — each one was run against a stub of the API before it was published.

By · Last updated: September 2026

TL;DR

To get the HTML of a JavaScript-rendered page in Python, POST its URL to https://urlpipe.dev/html with sync set to true. The body is the DOM after the page's scripts ran in real Chrome, serialized as HTML — what a visitor's browser holds, not what the server first sent — read with res.text, for 1 credit.

Free plan, no credit card. 1,000 credits a month.

One POST to /html loads the page in real Chrome, lets its JavaScript run, and serializes the DOM it ended up with. For a single-page app that is the difference between an empty <div id="root"> and the content; the rendered vs raw HTML guide shows how far apart the two get.

The body is the HTML as text, read with res.text. The program saves it and reports its size in bytes — compare that with curl -s https://example.com | wc -c and you see how much of the page only exists after the scripts ran.

Setup

Before you start

One dependency, requests. Put your API key in the environment so it never lands in the file.

Terminal
python3 -m pip install requests
export URLPIPE_API_KEY="your_api_key"

The request

Save the rendered HTML and report its size

"sync": True holds the connection open until the result is ready and answers with it. Pass timeout= every time: requests has no default, and a page that never finishes loading would otherwise hang your process. Write res.content (bytes) rather than res.text, so the file is byte-for-byte what came back and its size is the real size.

rendered_html.py
import os

import requests

res = requests.post(
    "https://urlpipe.dev/html",
    headers={"Authorization": f"Bearer {os.environ['URLPIPE_API_KEY']}"},
    json={"url": "https://example.com", "sync": True},
    # requests waits forever unless told otherwise; a sync call can take up to 60 s.
    timeout=90,
)
res.raise_for_status()

with open("page.html", "wb") as f:
    f.write(res.content)
print(f"Saved page.html ({len(res.content)} bytes)")

Run it: python3 rendered_html.py

Async

The async variant: a token, a webhook and a poll

Leave out sync and the answer is a token, straight away. for … else is the idiom for "ran out of attempts": the else runs only when the loop never hit break.

rendered_html_async.py
import os
import time

import requests

API = "https://urlpipe.dev"
HEADERS = {"Authorization": f"Bearer {os.environ['URLPIPE_API_KEY']}"}

# No "sync": the request is accepted at once and the work carries on without you.
res = requests.post(
    f"{API}/html",
    headers=HEADERS,
    json={
        "url": "https://example.com",
        "report_to": "https://your-app.com/webhooks/urlpipe",
        "labels": {"customer": "acme"},
    },
    timeout=30,
)
res.raise_for_status()
token = res.json()["token"]
print(f"Accepted {token}")

# The result is POSTed to report_to when it is ready. Polling by token is the
# other way to collect it: no endpoint needed, and a backup for the webhook.
for _ in range(60):
    res = requests.get(f"{API}/result/{token}", headers=HEADERS, timeout=30)
    if res.status_code != 202:  # 202 means still processing
        break
    time.sleep(2)
else:
    raise SystemExit("Still processing after two minutes; try the token again later.")

if res.status_code == 422:
    raise SystemExit(f"The analysis failed: {res.json()['error']}")
if res.status_code == 410:
    raise SystemExit("The result is past the 30-day window; send the request again.")
res.raise_for_status()

with open("page.html", "wb") as f:
    f.write(res.content)
print(f"Saved page.html ({len(res.content)} bytes)")

Run it: python3 rendered_html_async.py

Errors

Handle errors and retries

raise_for_status() is fine for a script; a service wants to tell the failures apart. Check the status before calling res.json() — a 401 body is not JSON. sys.exit(message) prints to stderr and exits 1.

rendered_html_errors.py
import os
import sys
import time

import requests

API_KEY = os.environ["URLPIPE_API_KEY"]


class URLpipeError(Exception):
    pass


def urlpipe(path, payload, attempts=5):
    """POST a sync request and return the response, or raise URLpipeError."""
    for attempt in range(attempts):
        res = requests.post(
            f"https://urlpipe.dev{path}",
            headers={"Authorization": f"Bearer {API_KEY}"},
            json={**payload, "sync": True},
            timeout=90,
        )
        if res.status_code == 200:
            return res
        if res.status_code == 401:
            raise URLpipeError("401: the API key is missing or wrong. Check URLPIPE_API_KEY.")

        try:
            body = res.json()
        except ValueError:
            body = {}
        code = body.get("error", "")
        detail = f"{code}: {body['message']}" if body.get("message") else code

        if res.status_code == 429 and code == "rate_limited":
            # Sending too fast: Retry-After says how long the window has left.
            time.sleep(int(res.headers.get("Retry-After", 1)))
        elif res.status_code == 429 and code == "concurrency_limit":
            # Every parallel slot on your plan is busy with your own requests.
            time.sleep(2**attempt)
        elif res.status_code == 504:
            # Still running on our side; the token collects it from GET /result/:token.
            raise URLpipeError(f"504 processing_timeout: collect it later with token {body['token']}")
        else:
            # 403 email_unverified, 422 (a bad parameter, or a page that would not load),
            # 429 quota_exceeded: sending the same request again gets the same answer.
            raise URLpipeError(f"{res.status_code}: {detail}")
    raise URLpipeError(f"429: still refused after {attempts} attempts")


try:
    res = urlpipe("/html", {"url": "https://example.com"})
except URLpipeError as error:
    sys.exit(f"URLpipe: {error}")
except requests.RequestException as error:
    sys.exit(f"Network error: {error}")

with open("page.html", "wb") as f:
    f.write(res.content)
print(f"Saved page.html ({len(res.content)} bytes)")

Run it: python3 rendered_html_errors.py

Details

What to know about /html

  • The HTML is serialized from the live DOM, so it is well-formed but not byte-identical to any file on the server.
  • Pages over 10 MB of HTML are refused with The page is too big to be processed.
  • page_options.wait_for_selector waits for an element that loads late; delay waits a fixed time on top.
  • It is the whole document, scripts and styles included, not a cleaned-up version of it. For clean text, use Markdown instead.

Other languages

Get the rendered HTML of a JavaScript page in another language

More Python: every Python recipe

FAQ

Frequently asked questions

How is this different from fetching the URL myself?
A plain GET returns what the server sent before any JavaScript ran. /html returns the DOM after the page's scripts ran in real Chrome — the HTML a visitor's browser actually holds.
Can I wait for content that loads late?
Yes: page_options.wait_for_selector waits up to 10 seconds for an element to appear, and delay adds a fixed wait. If the element never appears the request is a 422 and costs nothing.
Does URLpipe follow robots.txt?
Yes, by default, for every project. You can turn it off per project, and then you are responsible for having the right to fetch those pages.
Do I need an SDK to call URLpipe from Python?
requests is the only dependency; there is no SDK to install, because the API is one POST per job.

Get a key and run it.

Free plan, no card. Paste your key into URLPIPE_API_KEY and every program on this page runs as it is.