Web Scraping with Python and AI: Turn Pages Into Clean Data
Fetch pages politely, strip them to the text that matters, and have Claude return validated JSON. Includes robots.txt checks, caching and cost control.
In this article
- Read this before you scrape anything
- When AI extraction is worth it (and when it isn't)
- Set up
- Step 1: Fetch politely
- Step 2: Shrink the page to the text that matters
- Step 3: Ask Claude for structured data
- Step 4: Validate before you trust
- Step 5: Put it together
- Keeping the bill small
- Troubleshooting
- Run it on a schedule
- Where to go from here
Classic web scraping has one big weakness: it breaks the moment a site changes its HTML. You write a selector like div.product > span.price, it works for a month, then a redesign renames a class and your pipeline quietly returns nothing.
AI extraction flips the trade-off. Instead of telling the program where the data lives, you tell it what you want and let a language model find it in the page text. It is slower and costs a little per page, but it shrugs off most layout changes and handles messy, inconsistent pages that selectors struggle with.
In this guide you will build a small, polite scraper that checks robots.txt, caches what it downloads, asks Claude for structured JSON, validates the result and writes a CSV. It is deliberately short so you can read all of it, and every part is something you can swap out.
#Read this before you scrape anything
Good scraper manners also keep you out of trouble technically:
- Identify yourself with a descriptive
User-Agentthat includes a way to contact you. - Slow down. One request every couple of seconds is a polite default.
- Cache everything so you never download the same page twice while you develop.
- Stop on errors. A
403or429is a request to back off, not a puzzle to defeat.
#When AI extraction is worth it (and when it isn't)
Use plain selectors when the page structure is stable and simple. They're fast, free and deterministic. Reach for AI when:
- The data is scattered across free text, such as a job posting or an event description.
- Many different sites carry the same kind of information in different layouts.
- The markup changes often and you are tired of fixing selectors.
A hybrid is often best: try the selectors first and fall back to AI only for pages where they return nothing.
#Set up
You need Python 3.9 or newer and four packages:
pip install requests beautifulsoup4 anthropic pydantic
export ANTHROPIC_API_KEY="your-key-here"New to the API? The Claude API beginner's guide covers keys, requests and errors in detail.
#Step 1: Fetch politely
This part has nothing to do with AI, and it is where most scrapers go wrong. We check robots.txt before every new domain, wait between requests, set a clear User-Agent, and store each page on disk so a re-run costs nothing.
import hashlib
import pathlib
import time
from urllib import robotparser
from urllib.parse import urlparse
import requests
USER_AGENT = "MyResearchBot/1.0 (contact: you@example.com)"
DELAY_SECONDS = 2
CACHE = pathlib.Path(".cache")
CACHE.mkdir(exist_ok=True)
_robots = {}
def allowed(url: str) -> bool:
parts = urlparse(url)
base = f"{parts.scheme}://{parts.netloc}"
if base not in _robots:
rp = robotparser.RobotFileParser()
rp.set_url(f"{base}/robots.txt")
try:
rp.read()
except Exception:
return False # can't confirm permission, so don't fetch
_robots[base] = rp
return _robots[base].can_fetch(USER_AGENT, url)
def fetch(url: str) -> str:
cached = CACHE / f"{hashlib.sha256(url.encode()).hexdigest()[:16]}.html"
if cached.exists():
return cached.read_text(encoding="utf-8")
if not allowed(url):
raise PermissionError(f"robots.txt does not allow fetching {url}")
time.sleep(DELAY_SECONDS)
response = requests.get(url, headers={"User-Agent": USER_AGENT}, timeout=20)
response.raise_for_status()
cached.write_text(response.text, encoding="utf-8")
return response.textSet USER_AGENT to something that really identifies you. Pretending to be a normal browser is exactly the behaviour site owners object to.
#Step 2: Shrink the page to the text that matters
Raw HTML is mostly noise: scripts, styles, menus, tracking pixels. You pay for every token you send, so strip the page down first. BeautifulSoup does this in a few lines:
from bs4 import BeautifulSoup
NOISE = ["script", "style", "nav", "footer", "header", "aside", "form", "noscript", "svg"]
def page_text(html: str, limit: int = 12000) -> str:
soup = BeautifulSoup(html, "html.parser")
for tag in soup(NOISE):
tag.decompose()
text = soup.get_text("\n", strip=True)
return text[:limit] # hard cap keeps cost predictableThe limit is a blunt but effective cost control. If the data you want sits at the bottom of long pages, raise it, or select the main content container first with a selector and pass only that.
#Step 3: Ask Claude for structured data
You could ask for JSON in the prompt and parse whatever comes back, but there is a more reliable way: define a tool whose input schema is the shape you want, and force Claude to call it. The reply then arrives as a Python dictionary that already matches your schema.
import anthropic
MODEL = "claude-haiku-5-5" # a small, fast model is usually enough for extraction
client = anthropic.Anthropic()
TOOL = {
"name": "save_product",
"description": "Save the product details found on the page.",
"input_schema": {
"type": "object",
"properties": {
"name": {"type": "string"},
"price": {"type": ["number", "null"], "description": "Numeric price, or null if not shown"},
"currency": {"type": ["string", "null"], "description": "ISO code such as USD, or null"},
"in_stock": {"type": ["boolean", "null"]},
"summary": {"type": "string", "description": "One factual sentence, no marketing language"},
},
"required": ["name", "summary"],
},
}
def extract(text: str, url: str):
response = client.messages.create(
model=MODEL,
max_tokens=800,
system=(
"Extract product details from the page text. "
"Use null for anything the page does not clearly state. Never guess."
),
tools=[TOOL],
tool_choice={"type": "tool", "name": "save_product"},
messages=[{"role": "user", "content": f"URL: {url}\n\nPAGE TEXT:\n{text}"}],
)
call = next(block for block in response.content if block.type == "tool_use")
return call.input, response.usageTwo things make this robust. The system prompt gives the model permission to say "null" instead of inventing a value, which cuts down on made-up data. And tool_choice forces the structured reply, so you don't have to coax or repair free text.
#Step 4: Validate before you trust
A model can follow your schema and still be wrong. Validate types, then check the values against the source text where you can. Pydantic handles the first part:
from typing import Optional
from pydantic import BaseModel, ValidationError
class Product(BaseModel):
name: str
price: Optional[float] = None
currency: Optional[str] = None
in_stock: Optional[bool] = None
summary: str
def validate(data: dict, source_text: str):
try:
product = Product.model_validate(data)
except ValidationError as err:
raise ValueError(f"Schema check failed: {err}") from err
# Cheap sanity check: a price the page never mentions is a red flag.
if product.price is not None and str(int(product.price)) not in source_text:
raise ValueError(f"Price {product.price} not found in page text; check by hand")
return productThat last check is crude on purpose. It catches the classic failure where the model returns a plausible but absent number. Rows that fail go into a "review" list instead of your dataset.
#Step 5: Put it together
Now a short script ties the stages into a loop, writes a CSV, and prints what the run cost in tokens so there are no surprises:
import csv
from clean import page_text
from extract import extract
from fetch import fetch
from validate import validate
URLS = [
"https://example.com/products/1",
"https://example.com/products/2",
]
good, review = [], []
tokens_in = tokens_out = 0
for url in URLS:
try:
text = page_text(fetch(url))
data, usage = extract(text, url)
tokens_in += usage.input_tokens
tokens_out += usage.output_tokens
product = validate(data, text)
good.append({"url": url, **product.model_dump()})
except Exception as err: # keep going; one bad page shouldn't stop the batch
review.append({"url": url, "problem": str(err)})
if good:
with open("products.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=good[0].keys())
writer.writeheader()
writer.writerows(good)
print(f"{len(good)} saved, {len(review)} need review")
print(f"Tokens used: {tokens_in} in, {tokens_out} out")
for item in review:
print("REVIEW:", item["url"], "-", item["problem"])Replace the example URLs with pages you are allowed to fetch. Because the fetcher caches to disk, you can tweak the prompt and re-run as often as you like while only paying for the model calls.
#Keeping the bill small
Extraction is the cheapest kind of AI work to optimise, because the output is small and the task is narrow:
- Use a small model. Start with a Haiku-class model and move up only if accuracy needs it.
- Trim the input. Pass the main content container, not the entire page.
- Cache both stages. You already cache the HTML; also cache the extraction result keyed by a hash of the cleaned text.
- Skip what hasn't changed. Compare a hash of the page text with the last run and only re-extract when it differs.
- Batch when you can wait. For large overnight jobs, a batch interface is cheaper than real-time requests.
#Troubleshooting
| Problem | What it usually means | What to do |
|---|---|---|
403 or 429 errors |
The site is refusing or rate-limiting you | Slow down, check the terms and look for an official API. Don't try to evade blocks |
| Page text is nearly empty | The content is built by JavaScript after load | Look for an official feed or API. Only if the terms allow it, use a headless browser |
| Wrong or invented values | The model guessed | Keep the "use null, never guess" rule and the source-text check; tighten the schema descriptions |
| Costs creeping up | Pages are long or you re-run everything | Lower limit, narrow the container, cache results |
| Works on one site, fails on another | Different layouts hide the data differently | Add one or two example pages to the prompt as guidance |
#Run it on a schedule
A scraper is most useful when it runs without you. You can schedule run.py with cron on a server or with a scheduled GitHub Actions workflow. If you'd rather click than code, a visual workflow tool such as n8n can run a Python script or call the Claude API on a timer, and send the results to a sheet or an email. Whichever you choose, send yourself a short report after each run listing how many rows were saved and how many need review.
#Where to go from here
You now have a reusable pattern: fetch, clean, extract against a schema, validate, store. The same skeleton works for job listings, event pages, recipes, documentation changelogs and price tracking; you only change the schema and the prompt. Each new use case is a good candidate for the weekly workflow we send out, so subscribe if you want the next one.
Get the next workflow by email
One tested workflow every week. No filler, unsubscribe in one click.