Guides
Turn the web into clean data
Practical guides to getting clean data out of web pages — for LLMs, agents, screenshots and audits — each with a free tool to try it on a real page.
Each guide pairs the concept with a free, no-signup tool so you can try it on a real page.
Web pages as LLM input
- HTML to Markdown for LLMs
Why Markdown is the right format for language models, what a clean conversion keeps and drops, and how to handle client-rendered pages.
Read the guide - Get keywords from a URL
TF-IDF, RAKE and YAKE vs model-based extraction, a competitor's page as a keyword brief, clustering, and the limits.
Read the guide - Chunking web pages for RAG
Markdown first, heading-aware splitting, chunk size and overlap, metadata, and keeping chunks fresh.
Read the guide - What llms.txt is
The format section by section, a real example, and an honest look at which tools read it.
Read the guide
Agents that read the web
- Why AI agents read empty pages
The reference fetch server does a plain GET, so client-rendered pages reach your agent empty — measured on three real pages.
Read the guide - Web access for Claude and Cursor
Exact configuration for Claude Code, Claude Desktop and Cursor, token scopes, first prompts and costs.
Read the guide
Capturing a page
- Extracting page metadata
The three metadata standards, why they conflict in practice, and how to reconcile them into one clean object.
Read the guide - Open Graph image sizes
The og:image size and file limits each network documents, checked against its own docs, and the one size that works everywhere.
Read the guide - Screenshot a website programmatically
Puppeteer vs Playwright vs an API, and the full-page gotchas: lazy loading, sticky headers, banners and size limits.
Read the guide
Page health, from outside
- Reading a Lighthouse audit
What the four categories and the Performance score mean, how Core Web Vitals are measured, and how to prioritize fixes.
Read the guide - Core Web Vitals: lab vs field
LCP, INP and CLS, their thresholds, and why a Lighthouse run and real-user data tell different stories.
Read the guide - Find JavaScript errors on any site
DevTools, headless Chrome and load-time captures — and what each can and cannot see on a site you don't own.
Read the guide
Fetching the modern web properly
- Rendered vs. raw HTML
Why modern sites return an empty shell to curl, how client-side rendering works, and how to get the post-JavaScript DOM.
Read the guide - Verify webhook signatures (HMAC)
The signed string, the raw body, timestamp tolerance, constant-time compares and rotation — with code in five languages.
Read the guide - Polite scraping: robots.txt and pacing
What RFC 9309 says about robots.txt, how to pace requests per host, and how to back off when a site struggles.
Read the guide - Idempotency keys explained
What the header does, how servers implement it, and the mistakes that make a retry charge twice.
Read the guide - Cookie banners and ads in scraped pages
Why automated visits get every banner and ad, and how to remove them without consenting on anyone's behalf.
Read the guide
From reading to shipping
The URLpipe API powers every tool in these guides. Grab a free key and call any endpoint from your code — 1,000 credits a month, no card.