Skip to content

About

Async Playwright web scraper: full books.toscrape.com catalogue to clean CSV/JSON, with retries, bounded concurrency and tests

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Book Scraper

Async web scraper that collects the full catalogue of books.toscrape.com (a sandbox site built for scraping practice) into clean CSV and JSON.

It crawls every listing page, visits all 1,000 book detail pages concurrently, and exports structured data ready for Excel, a database, or further analysis.

Measured run: 1,000 / 1,000 books scraped, 0 failures, 89 seconds (--concurrency 8).

Features

  • Headless browser scraping with Playwright (Chromium), so the same approach works on JavaScript-rendered sites
  • Bounded concurrency: fetches many pages in parallel without hammering the server
  • Retries with exponential backoff + jitter for timeouts, 429, and 5xx responses; permanent errors such as 404 fail fast
  • Fault tolerant: a broken page is logged and skipped, the rest of the crawl continues
  • Faster page loads by blocking images, fonts, and CSS the scraper does not need
  • Clean data: numeric prices, ratings as 1–5, stock as an integer, absolute image URLs, truncation markers removed
  • Excel-friendly CSV (UTF-8 with BOM, so £ and accented titles display correctly) and pretty-printed JSON
  • Unit-tested parser that runs offline against saved pages

Output

Each book record:

{
  "upc": "a897fe39b1053632",
  "title": "A Light in the Attic",
  "category": "Poetry",
  "price": 51.77,
  "currency": "GBP",
  "rating": 3,
  "stock": 22,
  "reviews": 0,
  "description": "It's hard to imagine a world without A Light in the Attic. ...",
  "image_url": "https://books.toscrape.com/media/cache/fe/72/fe72f0532301ec28892ae79a629a293c.jpg",
  "url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html"
}

Files are written to data/books.csv and data/books.json.

Quick start

Requires Python 3.10+ (tested on 3.12).

python -m venv .venv
# Windows: .venv\Scripts\activate    macOS/Linux: source .venv/bin/activate
pip install -r requirements.txt
playwright install chromium

python -m scraper                  # full catalogue
python -m scraper --max-pages 2    # quick test: first 40 books

Options

Option Default Description
--max-pages N all Limit listing pages (20 books each)
--concurrency N 5 Pages fetched in parallel
--retries N 3 Attempts per URL before giving up
--out DIR data Output directory
--formats csv json Output formats, one or both
--headed off Show the browser window
-v, --verbose off Log every request and status code

Exit codes: 0 success, 1 crawl aborted, 2 invalid arguments, 3 finished but some pages failed, 130 interrupted.

How it works

listing page 1 ──► total page count
      │
      ▼
listing pages 2..N (parallel) ──► 1,000 book URLs
      │
      ▼
book detail pages (parallel, retried) ──► parse ──► CSV / JSON

Fetching and parsing are deliberately separate: fetcher.py only returns HTML, and parser.py only turns HTML into data. This keeps the parser testable without a browser or network, and lets the fetch layer change (proxies, a plain HTTP client, a different site) without touching extraction logic.

Project structure

scraper/
  cli.py        command-line options, logging, exit codes
  crawler.py    crawl flow and concurrency control
  fetcher.py    Playwright browser, resource blocking, retries
  parser.py     HTML -> Book, pure functions
  exporters.py  CSV and JSON writers
  models.py     Book data model
tests/
  fixtures/     saved listing and detail pages
  test_parser.py

Tests

pip install -r requirements-dev.txt
pytest

About

Async Playwright web scraper: full books.toscrape.com catalogue to clean CSV/JSON, with retries, bounded concurrency and tests

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages