A web scraper that learns page layouts and remembers them.
One LLM call per page template, not per page.
ScrapeForge is a self-learning, LLM-powered web scraper that doesn't just extract data , it understands page templates and caches them for reuse.
Traditional scrapers break when a site redesigns. ScrapeForge breaks this cycle.
Most of them do the same thing: send the page to an LLM, get structured data back, charge you for the token, repeat for every single URL. Scrape 1,000 products, pay for 1,000 calls. Site redesigns? Start over.
That's not intelligence. That's an API invoice with extra steps.
ScrapeForge works differently. It looks at a page once, figures out the underlying template (product/iphone and product/samsung are the same layout), generates a CSS extraction schema, and caches it. From then on, extraction is pure CSS selectors no LLM, no cost, no waiting.
1,000 product pages on the same template = 1 LLM call. That's the whole pitch.
I needed a solid scraping system for AI agents and RAG pipelines something reliable, fast, and cheap enough to use at scale. Sending every page to an LLM wasn’t sustainable, and maintaining brittle CSS selectors wasn’t appealing. The idea of caching page templates rather than page data stayed in my mind for months, but I never had the right foundation to build it cleanly.
That changed when I discovered Crawl4AI. Its flexible extraction strategies and anti‑bot support gave me exactly the building blocks I needed. I started researching how to separate layout understanding from data extraction, and iterated on the design until the system could:
- Learn a template from a single page using an LLM (most time taken part).
- Cache the schema with a matching URL pattern.
- Test itself empirically before trusting a cached schema.
- Merge templates automatically when they share the same field structure.
The result is ScrapeForge: learn once with an LLM, extract thousands of times with pure CSS cutting cost and latency by over 90% compared to per‑page AI scrapers.
First visit to a URL:
crawl HTML -> LLM generates extraction schema -> validate it
-> test it on the live page -> cache it in SQLite -> return data
Every visit after that:
match URL against cached patterns -> run CSS extraction -> return dataAnd because websites love ruining your week, the cache isn't trusted blindly:
- Empirical verification before a cached schema is used for a new URL, the extractor actually runs and the fill-rate is measured. Too many empty fields? The schema gets flagged, not silently returned.
- Rot detection when a cached schema keeps failing, hit
/schemas/refreshand a fresh one is generated and swapped in. - Template merging new schemas are compared against existing ones using Jaccard similarity on field signatures. Same layout detected? The URL patterns get generalized (
product/iphone+product/samsung->product/*) instead of duplicating schemas.
flowchart TD
A["Client<br/>curl / Claude / Cursor / any MCP client"] --> B["FastAPI<br/>+ fastapi-mcp"]
B --> C["SecurityRateLimitMiddleware<br/>X-API-Key check + per-IP rate limit"]
C --> D["/scrape request"]
D --> D1{"crawl_type?"}
%% Markdown / HTML path – skips schema service entirely
D1 -->|"markdown / html"| R1["crawl_with_filter<br/>Crawl4AI + Playwright (stealth mode)"]
R1 --> P1["Markdown or HTML response"]
R1 -.->|"pdf + screenshot<br/>if requested"| Q[("media volume")]
%% Structured path – goes through schema service
D1 -->|"structured"| E["Schema service<br/>(per-domain asyncio lock)"]
E --> F{"Cache lookup<br/>SQLite via SQLAlchemy"}
F -->|"exact URL match"| J
F -->|"pattern candidates"| G["Empirical test:<br/>run extraction, measure fill-rate"]
G -->|"passes (>= 50% fields filled)"| J
G -->|"fails"| H
H["Crawl raw HTML<br/>Crawl4AI + Playwright (stealth mode)"] --> I["LLM generates<br/>JsonCssExtractionStrategy schema"]
%% H -.->|"pdf + screenshot<br/>if requested"| Q
I --> K["Validate schema shape"]
K --> L{"Jaccard similarity >= 0.8<br/>vs existing domain schemas?"}
L -->|"same template"| M["Merge: generalize URL pattern,<br/>append example URL"]
L -->|"new template"| N["Insert new schema row"]
M --> J["Structured extraction<br/>CSS-based, zero LLM cost"]
N --> J
J --> O["Verify output<br/>fill-rate check"]
J -.->|"pdf + screenshot<br/>if requested"| Q
O --> P2["JSON response"]
E -.->|schemas + patterns| R[("SQLite schema.db")]
style E fill:#1f2937,color:#fff
style J fill:#065f46,color:#fff
style R fill:#78350f,color:#fff
- LLM schema generation: Feed raw HTML to Gemini, OpenAI, Groq, Ollama, anything supported by
crawl4ai.LLMConfig. Provider (usesLiteLLMas proxy) is a runtime parameter, not a hardcoded import. - Pattern learning: URLs sharing a layout are automatically grouped into wildcard patterns. One schema ends up covering thousands of URLs.
- Empirical cache validation: cached schemas are tested, not assumed. A schema that extracts nothing useful gets rejected before it wastes your time.
- Self-healing: stale schema after a redesign? One call to
/schemas/refreshregenerates and replaces it. - Three output modes: markdown, raw HTML, or structured JSON from a single endpoint.
- Anti-bot stealth: Crawl4AI's
magicmode (stealth, user simulation) on by default, with robots.txt support when you want to be polite. - MCP server built in: every endpoint is automatically exposed as a Model Context Protocol tool. Codex, Claude, Cursor, and any MCP client can call the scraper natively.
- Media capture: full-page PDF and screenshots, saved per-request and served via
/download/{id}. - Actually tested: 25 unit tests covering the URL generalizer, schema similarity, validators, repository layer, retry logic, and rate limiter.
git clone https://github.com/ajazhussainsiddiqui/ScrapeForge.git
cd ScrapeForge
docker build -t scrapeforge .
docker run -d --name scrapeforge \
-p 8000:8000 \
-e ENDPOINT_API_KEY="pick-something-strong" \
-v scrapeforge-data:/data \
--shm-size=1g \
scrapeforgepython -m venv venv && source venv/bin/activate
pip install -r requirements.txt
playwright install
echo "ENDPOINT_API_KEY=pick-something-strong" > .env
echo "LLM_API_KEY=your_key_here" >> .env # only needed for structured mode
uvicorn api:app --reloadMarkdown/HTML scraping needs no LLM key at all. Structured extraction needs one, passed via
.env.
All endpoints except /health and / require the X-API-Key header.
| Method | Endpoint | What it does |
|---|---|---|
POST |
/scrape |
Scrape a URL as markdown, HTML, or structured JSON |
GET |
/schemas |
List cached schemas (optional ?domain= filter) |
DELETE |
/schemas/{id} |
Delete a cached schema |
POST |
/schemas/refresh |
Force schema regeneration for a URL (post-redesign rescue) |
GET |
/download/{request_id}/{pdf\|screenshot} |
Fetch saved media |
GET |
/health |
Health check |
you can test this via Swagger UI (
http://localhost:8000/docs)
curl -X POST http://localhost:8000/scrape \
-H "X-API-Key: pick-something-strong" \
-H "Content-Type: application/json" \
-d '{
"url": "https://www.bbc.com/news/articles/cvgy5k4n07ko",
"crawl_type": "structured",
"model_provider": "gemini/gemini-2.0-flash",
"api_key": "your_llm_key"
}'First call: slow (LLM is generating the schema). Second call to any URL on the same template: fast, cheap, pure CSS.
{
"success": true,
"data": "{\"content\": \"[{ \\\"title\\\": \\\"...\\\", \\\"author\\\": \\\"...\\\" }]\", \"type\": \"structured\"}"
}Interactive docs at /docs. MCP tools at the mounted MCP endpoint for Claude/Cursor.
Everything lives in config.py, overridable via environment variables:
| Variable | Default | Notes |
|---|---|---|
ENDPOINT_API_KEY |
- | Required. No key, no server startup. |
DATABASE_URL |
sqlite:///./schema.db |
Any SQLAlchemy URL. Postgres works too. |
RATE_LIMIT_REQUESTS |
10 |
Max requests per IP per window |
RATE_LIMIT_WINDOW |
3600 |
Window in seconds |
SIMILARITY_THRESHOLD |
0.8 |
Jaccard score for merging templates |
SCHEMA_FAILURE_SCORE_THRESHOLD |
0.5 |
Max empty-field ratio before a schema is rejected |
MAX_RETRIES / BACKOFF_BASE_SECONDS |
3 / 2.0 |
Exponential backoff for crawls and LLM calls |
CRAWL_TIMEOUT_SECONDS |
70 |
Hard timeout per crawl |
MEDIA_DIR / LOG_DIR |
./media / ./logs |
Overridden to /data/... in Docker |
ScrapeForge/
├── core/
│ ├── crawler.py # Crawl4AI wrapper: markdown / HTML / structured + media
│ ├── middleware.py # API key auth + per-IP rate limiting
│ ├── schema_gen.py # LLM schema generation (any provider)
│ └── validators.py # schema shape validation + fill-rate verification
├── services/
│ └── schema_service.py # the brain: cache lookup, empirical tests, pattern merging
├── db/
│ ├── connection.py # SQLAlchemy engine
│ ├── models.py # Schema table
│ └── repository.py # CRUD + pattern queries
├── utils/
│ ├── url.py # domain extraction, wildcard matching, URL generalization
│ ├── schema.py # field signatures + Jaccard similarity
│ └── async_helpers.py # retry decorator + RateLimiter
├── tests/ # 25 tests, no network required
├── api.py # FastAPI app + MCP server
├── main.py # orchestration entrypoint
└── DockerfileThe included Dockerfile is production-shaped: non-root user, Chromium + system deps baked in, healthcheck on /health, /data volume for the SQLite DB and media, shm-size handled.
docker build -t scrapeforge .
docker run -d -p 8000:8000 \
-e ENDPOINT_API_KEY=... \
-v scrapeforge-data:/data \
--shm-size=1g scrapeforgeRuns happily on a single EC2 instance, Lightsail container, or any host that can hold a Docker container. Two honest caveats for multi-instance deployments: the rate limiter and per-domain locks are in-memory (pin to one replica, or move the limiter to Redis), and SQLite is a single-writer database (swap DATABASE_URL to Postgres when you scale out, the models are already portable).
A README that only shows strengths is a marketing page. Here's where ScrapeForge struggles:
- Detectable on some anti‑bot pages. Heavily restricted pages (like Bloomberg) aren't always bypassed (in current version).
- LLM output is non-deterministic. Schema generation mostly works, occasionally produces a weird selector. The validators catch most of it; the retry wrapper catches the rest. Still, expect the occasional failed first crawl on ugly HTML.
- Heavily personalized / infinite-scroll pages can defeat static schema caching.
scan_full_pageand stealth mode help, but some SPAs will needupdate_schema=truemore often. - LLM cost is per-template, not zero. If every URL on a domain has a unique layout (some forums, some listing sites), you pay per page anyway. ScrapeForge shines when templates repeat, on the web, they almost always do.
pytest tests/ -v25 tests across the URL generalizer, schema similarity, validators, repository layer, retry logic, and rate limiter. All run in-memory no network, no API keys, no Playwright.
Every scraping tools ends the same way: "now maintain your selectors forever." (or use LLM tokens for each scraping) I wanted to see what happens if you make the maintenance the interesting part schemas that test themselves, patterns that generalize themselves, a cache that knows when it's wrong.
MIT. Scrape responsibly, respect robots.txt, don't be a jerk with other people's servers.