Work in progress — This project is an experiment in "vibe coding" with AI assistance. It comes with no guarantee of correctness or completeness. Use at your own risk.
Stability promise (since v1.0.0) — only the stations referential export is stable: its format and its JSON Schema follow SemVer (see Versioning). Everything else (SDK, generated clients, sync tools) stays experimental and may change in any release.
Export a stations referential of Île-de-France, auto-sync PRIM (Île-de-France Mobilités) OpenAPI/Swagger specs, generate Python clients, and sync/validate Opendatasoft datasets.
This repository maintains up-to-date interface contracts from PRIM APIs and dataset exports from the IDFM Opendatasoft portal. Everything is manifest-driven and idempotent.
- Exports a stations referential of Île-de-France (logical stations with their lines, one JSON file, strict numbered format, closed JSON Schema) — see Stations referential export
- Syncs OpenAPI/Swagger specs from PRIM APIs (supports direct URLs and PRIM page scraping)
- Generates Python clients from specs using OpenAPI Generator
- Downloads dataset exports from Opendatasoft (JSONL format, full exports without pagination limits)
- Validates datasets against JSON Schema
- Runs nightly via GitHub Actions and opens PRs when updates are detected
The sync pipeline runs in 4 steps:
- sync_specs — downloads OpenAPI/Swagger specs from
manifests/apis.yml, resolves PRIM page URLs, caches with ETag/Last-Modified/sha256 - generate_clients — regenerates Python clients in
generated/clients/when specs change - sync_datasets — downloads dataset exports from Opendatasoft portal as defined in
manifests/datasets.yml - validate_datasets — retrieves JSON Schema for each dataset, validates records, generates reports
Each step is conditional: resources are only re-fetched or regenerated when changes are detected.
prim_api/ # Python SDK (IdFMPrimAPI, dataset sync, background updater)
referential/ # Stations referential exporter + JSON Schemas (stable contract)
samples/ # Runnable usage examples (update when adding endpoints/data)
manifests/ # YAML manifests (apis.yml, datasets.yml, urls_of_interest.yml)
specs/ # Downloaded OpenAPI/Swagger specs (committed)
generated/clients/ # Generated Python clients (committed)
data/schema/ # JSON Schemas for datasets (committed)
data/raw/ # Dataset exports in JSONL (gitignored)
data/reports/ # Validation reports (gitignored)
tools/ # CLI scripts
docs/site/ # Generated API docs (gitignored)
.github/workflows/ # CI and nightly sync workflows
Committed:
- Manifests (
manifests/*.yml) - Tools (
tools/*.py) - Tests (
tests/) - CI workflows (
.github/workflows/) - Project config (
pyproject.toml,.gitignore) - OpenAPI specs + metadata (
specs/) — updated by nightly sync - Generated Python clients (
generated/clients/) — regenerated when specs change - Dataset schemas (
data/schema/) — kept in sync with portal metadata - Referential schemas (
prim_api/referential/schemas/) — the contract, changed only with a new format - Test fixtures (
tests/fixtures/referential/) — small real extracts of the five source datasets
Gitignored (downloaded on demand by devs):
data/raw/— dataset exports (JSONL, can be large)data/reports/— validation reports
- Python 3.12+
- uv (recommended) or pip
- Docker (for client generation)
uv syncexport-referential builds, from five IDFM open datasets, a single JSON file listing the logical
stations of Île-de-France with their lines — meant to be embedded offline in apps. It knows nothing
about cities or consumers: it always covers the whole network, and its only options are the
format and the modes to keep.
# From a tag, without cloning (what consumers should pin):
uvx --from git+https://github.com/MaxLeb/idfm-prim-api@v1.1.0 \
export-referential --format 2 --modes METRO,RER,TRAIN,TRAM > stations.json
# From a checkout (always pass --format: the default stays 1 for compatibility):
uv run export-referential --format 2 --modes METRO,RER -o stations.json
uv run export-referential --format 2 --no-download -o stations.json # reuse downloaded datasetsDatasets are downloaded to data/raw/ in a checkout, or to ~/.cache/idfm-prim-api/raw
($XDG_CACHE_HOME) when installed. The file is validated against its schema before being written.
Each release also publishes, for every produced
format, the export of all modes and its schema.
{
"format": 1,
"version": "2026-09-27T07:05:39Z",
"generatedAt": "2026-09-27T09:00:00Z",
"generator": "idfm-prim-api@1.0.0",
"source": "idfm-opendata",
"license": "ODbL-1.0",
"attribution": "Île-de-France Mobilités",
"modes": ["METRO", "RER", "TRAIN", "TRAM"],
"stops": [
{ "id": "IDFM:474151", "name": "Châtelet - Les Halles", "lat": 48.86174, "lon": 2.34697,
"areas": ["IDFM:monomodalStopPlace:45102"],
"lines": [
{ "id": "IDFM:C01742", "label": "A", "mode": "RER", "color": "EB2132", "textColor": "FFFFFF" }
] }
]
}| Field | Meaning |
|---|---|
format |
Version of the data contract (integer). |
version |
Most recent data_processed date among the five source datasets (UTC): changes if and only if the content may have changed. |
generatedAt |
Time of the export (UTC). |
generator |
idfm-prim-api@<version> that produced the file. With modes, enough to reproduce it. |
source, license, attribution |
idfm-opendata, ODbL-1.0, « Île-de-France Mobilités ». |
modes |
Modes kept at export time: a station of another mode was not exported, not absent. |
stops[].id, name, lat, lon |
The IDFM zone de correspondance (multimodal hub): its id, its name, its official point converted from Lambert 93 to WGS 84 (six decimals). |
stops[].areas |
Stop areas of the hub served by a kept line (what SIRI Stop Monitoring queries). |
stops[].lines[] |
IDFM line id, label, mode, colours (six uppercase hex digits, or null). |
Construction rules (all part of the format): a station is a zone de correspondance, never a
grouping computed by distance or name; only active lines are kept; a hub without a kept line is
not exported; stations are sorted by id, lines by mode then natural label order then id. Modes:
| Value | IDFM modes |
|---|---|
METRO |
metro |
RER |
RER A to E (rail / local) |
TRAIN |
Transilien H to V, TER |
TRAM |
tramway |
BUS |
bus (all sub-modes) |
SHUTTLE |
CDG VAL, Orlyval |
FUNICULAR |
Montmartre funicular |
CABLEWAY |
Câble C1 |
Sources: zones-de-correspondance, zones-d-arrets, arrets (Licence Ouverte 2.0) and
arrets-lignes, referentiel-des-lignes (ODbL). Any inconsistency in them (unresolvable stop,
line, area or hub, unknown mode) fails the export rather than dropping data silently.
Format 2 is format 1 plus, for every station, the accesses (entrances and exits) of its stop
areas, from the IDFM datasets acces and relations-acces (Licence Ouverte 2.0). Consumers that do
not need them keep asking for --format 1, which stays produced.
{ "id": "IDFM:474151", "name": "Châtelet - Les Halles", "…": "…",
"accesses": [
{ "id": "IDFM:50148652", "name": "Porte Marguerite de Navarre", "number": 1,
"lat": 48.860433, "lon": 2.346346, "entry": true, "exit": true }
] }| Field | Meaning |
|---|---|
accesses[].id |
IDFM:<access id>: the id that platform positioning (positionnement-dans-la-rame) points to. |
name, number |
Name as published; number shown on the signage, or null. |
lat, lon |
Published WGS 84 point, six decimals. |
entry, exit |
Whether one can enter, exit, or both. |
Rules: every access linked to one of the station's areas (the stop areas of the kept modes), each
listed once, sorted by numeric id; [] when none is published (tram stops, a few TER stations
outside Île-de-France). The version is the most recent data_processed of the seven datasets.
- The format is strict: every field is always present (
nullwhen the source has no value), each schema is closed (prim_api/referential/schemas/stations.format-<n>.schema.json), and any change of shape — a field, a mode value, a construction rule — is a new format number. Consumers read exactly the format they know and refuse the others. - The repository follows SemVer on that contract: producing a new format while still producing the previous ones is a minor release; dropping a format is a major release; a fix that keeps format and output unchanged is a patch.
- Format 1 is frozen by
v1.0.0, format 2 byv1.1.0. Tags are pushed by a maintainer, never by automation. - The format is owned by its first consumer (the Everyday France app); this repository is its first producer and hosts its schema.
The referential is a derived database of Île-de-France Mobilités open data, two of its sources being
under ODbL: every export is published under the ODbL 1.0, with the attribution
« Île-de-France Mobilités », carried by the license and attribution fields of the file itself.
Releases make the full export available; a filtered export is reproducible from generator + modes.
- New: format 2 (
--format 2,prim_api.referential.format2, closed schemastations.format-2.schema.json): format 1 plus the accesses of each station. Format 1 is still produced unchanged. - New datasets in
manifests/datasets.yml:acces,relations-acces. - The release publishes every format (all modes) and every schema.
- CLI: one registry of formats (sources and builder); the summary counts accesses for format 2.
- New:
export-referentialandprim_api.referential— stations referential, format 1, closed JSON Schema, release workflow publishing the export (all modes) and the schema. - New datasets in
manifests/datasets.yml:arrets,zones-de-correspondance. prim_api.datasets: explicitdata_dir,load_metadata,fetch_dataset_metadata(portaldata_processed, the export API gives no ETag/Last-Modified),record_metadata_fields.import prim_apino longer loads the generated PRIM client eagerly (IdFMPrimAPIis imported on first access), so the package works when installed outside a checkout.- Packaging: build system declared (
uvx --from git+…@<tag>works),pyprojdependency, stability promise limited to the export.
The prim_api package provides a high-level Python interface to PRIM APIs and datasets.
from prim_api import IdFMPrimAPI
api = IdFMPrimAPI(api_key="your-prim-key")
# Query real-time next passages at a stop
passages = api.get_passages("IDFM:473921")
# Filter by line
passages = api.get_passages("IDFM:473921", line_id="IDFM:C01742")
# Access downloaded datasets
zones = api.get_zones_darrets()
lignes = api.get_referentiel_lignes()
# Cleanup (stops background dataset updater)
api.stop()IdFMPrimAPI(
api_key="...", # Required. PRIM API key.
auto_sync=True, # Download missing datasets on init.
sync_interval=3600, # Background refresh interval in seconds.
)| Method | Description |
|---|---|
get_passages(stop_id, *, line_id=None) |
Real-time next passages at a stop/area |
get_zones_darrets() |
Load zones-d-arrets dataset as list of dicts |
get_referentiel_lignes() |
Load referentiel-des-lignes dataset as list of dicts |
get_arrets_lignes() |
Load arrets-lignes (stop-line associations) as list of dicts |
ensure_datasets() |
Download datasets if missing or stale |
refresh_datasets() |
Force re-check all datasets |
stop() |
Stop the background updater thread |
The prim_api.refs module provides helpers to convert between IDFM and STIF identifier formats:
| Helper | Description |
|---|---|
parse_stop_ref(idfm_id) |
Auto-detect StopPointRef or StopAreaRef from an IDFM ID |
parse_line_ref(idfm_id) |
Parse an IDFM line ID into a LineRef |
StopPointRef / StopAreaRef / LineRef |
Dataclasses with .to_stif() and .from_idfm() |
from prim_api.refs import parse_stop_ref, parse_line_ref
stop = parse_stop_ref("IDFM:473921")
print(stop.to_stif()) # "STIF:StopPoint:Q:473921:"
line = parse_line_ref("IDFM:C01742")
print(line.to_stif()) # "STIF:Line::C01742:"Datasets are open data and don't require authentication:
from prim_api.datasets import ensure_all_datasets, load_dataset
ensure_all_datasets()
zones = load_dataset("zones-d-arrets")
lignes = load_dataset("referentiel-des-lignes")
arrets_lignes = load_dataset("arrets-lignes")See samples/ for runnable examples.
The sync tools live in tools/, which is not part of the installed package: run them as modules
from a checkout.
# Sync OpenAPI/Swagger specs
uv run python -m tools.sync_specs
# Generate Python clients
uv run python -m tools.generate_clients
# Download datasets
uv run python -m tools.sync_datasets
# Validate datasets
uv run python -m tools.validate_datasetsuv run python -m tools.sync_allMost tools support --dry-run to preview changes without modifying files:
uv run python -m tools.sync_all --dry-run# Run tests
uv run pytest
# Lint
uv run ruff check .
# Format
uv run ruff format .PRIM_TOKEN— Bearer token for authenticated PRIM spec exports (optional, required only if the API enforces auth)
Set in GitHub repository secrets for CI.
Runs on every PR and push to main:
- Install dependencies
- Lint with ruff
- Run tests
- Dry-run the sync pipeline
Runs nightly at 01:00 UTC (≈ 02:00 Europe/Paris):
- Syncs OpenAPI/Swagger specs from PRIM
- Regenerates Python clients if specs changed
- Opens a PR automatically if anything changed
Dataset sync is not part of the nightly — devs download data locally on demand via uv run python -m tools.sync_datasets.
Builds API documentation with pdoc and deploys to GitHub Pages on push to main.
CI runs pytest --cov and pushes a dynamic badge to a GitHub gist. See setup instructions below.
Defines APIs to sync. Supports two types:
type: direct— URL returns OpenAPI/Swagger JSON directlytype: prim_page— PRIM page URL; script scrapes HTML to find the spec export link
Example:
apis:
idfm_ivtr_requete_unitaire:
type: prim_page
page_url: "https://prim.iledefrance-mobilites.fr/fr/apis/idfm-ivtr-requete_unitaire"Defines Opendatasoft datasets to download and validate.
Example:
datasets:
- dataset_id: "zones-d-arrets"
portal_base: "https://data.iledefrance-mobilites.fr"
export_format: "jsonl"
validate: trueCurated list of useful URLs (docs, consoles, examples).
Example:
urls:
prim_api_example: "https://prim.iledefrance-mobilites.fr/fr/apis/idfm-ivtr-requete_unitaire"
dataset_zones_d_arrets: "https://data.iledefrance-mobilites.fr/explore/dataset/zones-d-arrets/"
explore_api_docs: "https://help.opendatasoft.com/apis/ods-explore-v2/"API reference is auto-generated from docstrings and published to GitHub Pages:
https://MaxLeb.github.io/idfm-prim-api/
This project's own code is licensed under the MIT License.
- IDFM data — datasets and API responses from Île-de-France Mobilités are published under the Open Database License (ODbL 1.0) or the Licence Ouverte 2.0 (Etalab), depending on the dataset.
- Stations referential export — a derived database, published under the ODbL 1.0 with the attribution « Île-de-France Mobilités » (see Licence of the export). Test fixtures in
tests/fixtures/referential/are small extracts of the source datasets, under their own licences. - Generated clients — Python clients in
generated/clients/are produced by OpenAPI Generator, licensed under Apache 2.0.