Skip to content

Repository files navigation

ScrapeForge

A web scraper that learns page layouts and remembers them.
One LLM call per page template, not per page.

Python FastAPI Crawl4AI Docker Tests


What is this?

ScrapeForge is a self-learning, LLM-powered web scraper that doesn't just extract data , it understands page templates and caches them for reuse.

Traditional scrapers break when a site redesigns. ScrapeForge breaks this cycle.

The problem with every other "AI scraper"

Most of them do the same thing: send the page to an LLM, get structured data back, charge you for the token, repeat for every single URL. Scrape 1,000 products, pay for 1,000 calls. Site redesigns? Start over.

That's not intelligence. That's an API invoice with extra steps.

ScrapeForge works differently. It looks at a page once, figures out the underlying template (product/iphone and product/samsung are the same layout), generates a CSS extraction schema, and caches it. From then on, extraction is pure CSS selectors no LLM, no cost, no waiting.

1,000 product pages on the same template = 1 LLM call. That's the whole pitch.


Origin & Design

I needed a solid scraping system for AI agents and RAG pipelines something reliable, fast, and cheap enough to use at scale. Sending every page to an LLM wasn’t sustainable, and maintaining brittle CSS selectors wasn’t appealing. The idea of caching page templates rather than page data stayed in my mind for months, but I never had the right foundation to build it cleanly.

That changed when I discovered Crawl4AI. Its flexible extraction strategies and anti‑bot support gave me exactly the building blocks I needed. I started researching how to separate layout understanding from data extraction, and iterated on the design until the system could:

  • Learn a template from a single page using an LLM (most time taken part).
  • Cache the schema with a matching URL pattern.
  • Test itself empirically before trusting a cached schema.
  • Merge templates automatically when they share the same field structure.

The result is ScrapeForge: learn once with an LLM, extract thousands of times with pure CSS cutting cost and latency by over 90% compared to per‑page AI scrapers.


How it works

First visit to a URL:
  crawl HTML  ->  LLM generates extraction schema  ->  validate it
  ->  test it on the live page  ->  cache it in SQLite  ->  return data

Every visit after that:
  match URL against cached patterns  ->  run CSS extraction  ->  return data

And because websites love ruining your week, the cache isn't trusted blindly:

  • Empirical verification before a cached schema is used for a new URL, the extractor actually runs and the fill-rate is measured. Too many empty fields? The schema gets flagged, not silently returned.
  • Rot detection when a cached schema keeps failing, hit /schemas/refresh and a fresh one is generated and swapped in.
  • Template merging new schemas are compared against existing ones using Jaccard similarity on field signatures. Same layout detected? The URL patterns get generalized (product/iphone + product/samsung -> product/*) instead of duplicating schemas.

Architecture

flowchart TD
    A["Client<br/>curl / Claude / Cursor / any MCP client"] --> B["FastAPI<br/>+ fastapi-mcp"]
    B --> C["SecurityRateLimitMiddleware<br/>X-API-Key check + per-IP rate limit"]
    C --> D["/scrape request"]
    D --> D1{"crawl_type?"}

    %% Markdown / HTML path – skips schema service entirely
    D1 -->|"markdown / html"| R1["crawl_with_filter<br/>Crawl4AI + Playwright (stealth mode)"]
    R1 --> P1["Markdown or HTML response"]
    R1 -.->|"pdf + screenshot<br/>if requested"| Q[("media volume")]

    %% Structured path – goes through schema service
    D1 -->|"structured"| E["Schema service<br/>(per-domain asyncio lock)"]
    E --> F{"Cache lookup<br/>SQLite via SQLAlchemy"}
    F -->|"exact URL match"| J
    F -->|"pattern candidates"| G["Empirical test:<br/>run extraction, measure fill-rate"]
    G -->|"passes (>= 50% fields filled)"| J
    G -->|"fails"| H

    H["Crawl raw HTML<br/>Crawl4AI + Playwright (stealth mode)"] --> I["LLM generates<br/>JsonCssExtractionStrategy schema"]
    %% H -.->|"pdf + screenshot<br/>if requested"| Q

    I --> K["Validate schema shape"]
    K --> L{"Jaccard similarity >= 0.8<br/>vs existing domain schemas?"}
    L -->|"same template"| M["Merge: generalize URL pattern,<br/>append example URL"]
    L -->|"new template"| N["Insert new schema row"]

    M --> J["Structured extraction<br/>CSS-based, zero LLM cost"]
    N --> J
    J --> O["Verify output<br/>fill-rate check"]
    J -.->|"pdf + screenshot<br/>if requested"| Q
    O --> P2["JSON response"]

    E -.->|schemas + patterns| R[("SQLite schema.db")]

    style E fill:#1f2937,color:#fff
    style J fill:#065f46,color:#fff
    style R fill:#78350f,color:#fff
Loading

Features

  • LLM schema generation: Feed raw HTML to Gemini, OpenAI, Groq, Ollama, anything supported by crawl4ai.LLMConfig. Provider (uses LiteLLM as proxy) is a runtime parameter, not a hardcoded import.
  • Pattern learning: URLs sharing a layout are automatically grouped into wildcard patterns. One schema ends up covering thousands of URLs.
  • Empirical cache validation: cached schemas are tested, not assumed. A schema that extracts nothing useful gets rejected before it wastes your time.
  • Self-healing: stale schema after a redesign? One call to /schemas/refresh regenerates and replaces it.
  • Three output modes: markdown, raw HTML, or structured JSON from a single endpoint.
  • Anti-bot stealth: Crawl4AI's magic mode (stealth, user simulation) on by default, with robots.txt support when you want to be polite.
  • MCP server built in: every endpoint is automatically exposed as a Model Context Protocol tool. Codex, Claude, Cursor, and any MCP client can call the scraper natively.
  • Media capture: full-page PDF and screenshots, saved per-request and served via /download/{id}.
  • Actually tested: 25 unit tests covering the URL generalizer, schema similarity, validators, repository layer, retry logic, and rate limiter.

Quickstart

With Docker (recommended)

git clone https://github.com/ajazhussainsiddiqui/ScrapeForge.git
cd ScrapeForge

docker build -t scrapeforge .
docker run -d --name scrapeforge \
  -p 8000:8000 \
  -e ENDPOINT_API_KEY="pick-something-strong" \
  -v scrapeforge-data:/data \
  --shm-size=1g \
  scrapeforge

Local

python -m venv venv && source venv/bin/activate
pip install -r requirements.txt
playwright install

echo "ENDPOINT_API_KEY=pick-something-strong" > .env
echo "LLM_API_KEY=your_key_here" >> .env        # only needed for structured mode

uvicorn api:app --reload

Markdown/HTML scraping needs no LLM key at all. Structured extraction needs one, passed via .env.


API

All endpoints except /health and / require the X-API-Key header.

Method Endpoint What it does
POST /scrape Scrape a URL as markdown, HTML, or structured JSON
GET /schemas List cached schemas (optional ?domain= filter)
DELETE /schemas/{id} Delete a cached schema
POST /schemas/refresh Force schema regeneration for a URL (post-redesign rescue)
GET /download/{request_id}/{pdf\|screenshot} Fetch saved media
GET /health Health check

Scrape a page

you can test this via Swagger UI (http://localhost:8000/docs)

curl -X POST http://localhost:8000/scrape \
  -H "X-API-Key: pick-something-strong" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://www.bbc.com/news/articles/cvgy5k4n07ko",
    "crawl_type": "structured",
    "model_provider": "gemini/gemini-2.0-flash",
    "api_key": "your_llm_key"
  }'

First call: slow (LLM is generating the schema). Second call to any URL on the same template: fast, cheap, pure CSS.

{
  "success": true,
  "data": "{\"content\": \"[{ \\\"title\\\": \\\"...\\\", \\\"author\\\": \\\"...\\\" }]\", \"type\": \"structured\"}"
}

Interactive docs at /docs. MCP tools at the mounted MCP endpoint for Claude/Cursor.


Configuration

Everything lives in config.py, overridable via environment variables:

Variable Default Notes
ENDPOINT_API_KEY - Required. No key, no server startup.
DATABASE_URL sqlite:///./schema.db Any SQLAlchemy URL. Postgres works too.
RATE_LIMIT_REQUESTS 10 Max requests per IP per window
RATE_LIMIT_WINDOW 3600 Window in seconds
SIMILARITY_THRESHOLD 0.8 Jaccard score for merging templates
SCHEMA_FAILURE_SCORE_THRESHOLD 0.5 Max empty-field ratio before a schema is rejected
MAX_RETRIES / BACKOFF_BASE_SECONDS 3 / 2.0 Exponential backoff for crawls and LLM calls
CRAWL_TIMEOUT_SECONDS 70 Hard timeout per crawl
MEDIA_DIR / LOG_DIR ./media / ./logs Overridden to /data/... in Docker

Project Structure (high‑level)

ScrapeForge/
├── core/
│   ├── crawler.py          # Crawl4AI wrapper: markdown / HTML / structured + media
│   ├── middleware.py       # API key auth + per-IP rate limiting
│   ├── schema_gen.py       # LLM schema generation (any provider)
│   └── validators.py       # schema shape validation + fill-rate verification
├── services/
│   └── schema_service.py   # the brain: cache lookup, empirical tests, pattern merging
├── db/
│   ├── connection.py       # SQLAlchemy engine
│   ├── models.py           # Schema table
│   └── repository.py       # CRUD + pattern queries
├── utils/
│   ├── url.py              # domain extraction, wildcard matching, URL generalization
│   ├── schema.py           # field signatures + Jaccard similarity
│   └── async_helpers.py    # retry decorator + RateLimiter
├── tests/                  # 25 tests, no network required
├── api.py                  # FastAPI app + MCP server
├── main.py                 # orchestration entrypoint
└── Dockerfile

Deployed on AWS

The included Dockerfile is production-shaped: non-root user, Chromium + system deps baked in, healthcheck on /health, /data volume for the SQLite DB and media, shm-size handled.

docker build -t scrapeforge .
docker run -d -p 8000:8000 \
  -e ENDPOINT_API_KEY=... \
  -v scrapeforge-data:/data \
  --shm-size=1g scrapeforge

Runs happily on a single EC2 instance, Lightsail container, or any host that can hold a Docker container. Two honest caveats for multi-instance deployments: the rate limiter and per-domain locks are in-memory (pin to one replica, or move the limiter to Redis), and SQLite is a single-writer database (swap DATABASE_URL to Postgres when you scale out, the models are already portable).


Honest limitations

A README that only shows strengths is a marketing page. Here's where ScrapeForge struggles:

  • Detectable on some anti‑bot pages. Heavily restricted pages (like Bloomberg) aren't always bypassed (in current version).
  • LLM output is non-deterministic. Schema generation mostly works, occasionally produces a weird selector. The validators catch most of it; the retry wrapper catches the rest. Still, expect the occasional failed first crawl on ugly HTML.
  • Heavily personalized / infinite-scroll pages can defeat static schema caching. scan_full_page and stealth mode help, but some SPAs will need update_schema=true more often.
  • LLM cost is per-template, not zero. If every URL on a domain has a unique layout (some forums, some listing sites), you pay per page anyway. ScrapeForge shines when templates repeat, on the web, they almost always do.

Testing

pytest tests/ -v

25 tests across the URL generalizer, schema similarity, validators, repository layer, retry logic, and rate limiter. All run in-memory no network, no API keys, no Playwright.


Why I built this

Every scraping tools ends the same way: "now maintain your selectors forever." (or use LLM tokens for each scraping) I wanted to see what happens if you make the maintenance the interesting part schemas that test themselves, patterns that generalize themselves, a cache that knows when it's wrong.


License

MIT. Scrape responsibly, respect robots.txt, don't be a jerk with other people's servers.

About

A self-learning web scraper that uses LLMs to generate CSS extraction schemas once per page template, then caches and auto-generalizes URL patterns. Zero manual selectors. Zero maintenance. FastAPI + MCP server included.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages