Async web scraper that collects the full catalogue of books.toscrape.com (a sandbox site built for scraping practice) into clean CSV and JSON.
It crawls every listing page, visits all 1,000 book detail pages concurrently, and exports structured data ready for Excel, a database, or further analysis.
Measured run: 1,000 / 1,000 books scraped, 0 failures, 89 seconds (--concurrency 8).
- Headless browser scraping with Playwright (Chromium), so the same approach works on JavaScript-rendered sites
- Bounded concurrency: fetches many pages in parallel without hammering the server
- Retries with exponential backoff + jitter for timeouts,
429, and5xxresponses; permanent errors such as404fail fast - Fault tolerant: a broken page is logged and skipped, the rest of the crawl continues
- Faster page loads by blocking images, fonts, and CSS the scraper does not need
- Clean data: numeric prices, ratings as
1–5, stock as an integer, absolute image URLs, truncation markers removed - Excel-friendly CSV (UTF-8 with BOM, so
£and accented titles display correctly) and pretty-printed JSON - Unit-tested parser that runs offline against saved pages
Each book record:
{
"upc": "a897fe39b1053632",
"title": "A Light in the Attic",
"category": "Poetry",
"price": 51.77,
"currency": "GBP",
"rating": 3,
"stock": 22,
"reviews": 0,
"description": "It's hard to imagine a world without A Light in the Attic. ...",
"image_url": "https://books.toscrape.com/media/cache/fe/72/fe72f0532301ec28892ae79a629a293c.jpg",
"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
}Files are written to data/books.csv and data/books.json.
Requires Python 3.10+ (tested on 3.12).
python -m venv .venv
# Windows: .venv\Scripts\activate macOS/Linux: source .venv/bin/activate
pip install -r requirements.txt
playwright install chromium
python -m scraper # full catalogue
python -m scraper --max-pages 2 # quick test: first 40 books| Option | Default | Description |
|---|---|---|
--max-pages N |
all | Limit listing pages (20 books each) |
--concurrency N |
5 | Pages fetched in parallel |
--retries N |
3 | Attempts per URL before giving up |
--out DIR |
data |
Output directory |
--formats |
csv json |
Output formats, one or both |
--headed |
off | Show the browser window |
-v, --verbose |
off | Log every request and status code |
Exit codes: 0 success, 1 crawl aborted, 2 invalid arguments, 3 finished but some pages failed, 130 interrupted.
listing page 1 ──► total page count
│
▼
listing pages 2..N (parallel) ──► 1,000 book URLs
│
▼
book detail pages (parallel, retried) ──► parse ──► CSV / JSON
Fetching and parsing are deliberately separate: fetcher.py only returns HTML, and parser.py only turns HTML into data. This keeps the parser testable without a browser or network, and lets the fetch layer change (proxies, a plain HTTP client, a different site) without touching extraction logic.
scraper/
cli.py command-line options, logging, exit codes
crawler.py crawl flow and concurrency control
fetcher.py Playwright browser, resource blocking, retries
parser.py HTML -> Book, pure functions
exporters.py CSV and JSON writers
models.py Book data model
tests/
fixtures/ saved listing and detail pages
test_parser.py
pip install -r requirements-dev.txt
pytest