We use docker containers to run the application locally.
Make sure the DCAT-US 3.0 schema submodule is checked out first — see
DCAT-US 3.0 schemas. If you cloned with
git clone --recurse-submodules, you already have it.
Build the static assets (requires npm):
% make install-static
Build and bring up docker containers:
% make build
% make up
% make load-test-data # optional: load fixture orgs, sources, and jobs
make up starts one Compose stack with a single database. The Flask app,
OpenSearch, transformer, and harvest source nginx all share that network, so
the app can reach the database at the db hostname and harvest jobs started
via LocalTaskHandler use the same database.
This is the default local workflow: run make up, use the app at
http://localhost:8080, and trigger harvests from the UI. You do not need a
separate harvest-runner container or a second database.
To reset to a clean database with fixtures loaded:
% make re-up
Refer to the Makefile for additional commands.
Note that you do not need to set the CF_SERVICE_USER and CF_SERVICE_AUTH variables. They are needed only in the Cloud.gov environment.
In deployed (Cloud.gov) environments, harvest jobs run as Cloud Foundry tasks via CFHandler. Locally there is no CF task API, so the app falls back to LocalTaskHandler, which runs the same python harvester/harvest.py <job_id> <job_type> command as a child subprocess of the running app. The handler is selected automatically by create_task_handler() (harvester/lib/task_handler.py): it uses CFHandler only when running on Cloud Foundry or when all three CF_* credentials are configured, and otherwise uses LocalTaskHandler. This means you can register a harvest source and trigger a harvest locally without any Cloud Foundry credentials.
The normal path is to trigger harvests from the app (UI or API) and let
LocalTaskHandler run harvester/harvest.py inside the app container.
For debugging harvester code directly — for example to use a debugger in your
IDE on harvester modules — you can still run the harvest runner on the host
after make up:
poetry install
poetry run python harvester/harvest.py <job_id> <job_type>Host-side runs use DATABASE_URI from .env (localhost:5432), which is the
same database exposed by make up. No second database or separate make
target is required.
Point your web browser to http://localhost:8080
To use local login, set ENABLE_LOCAL_DEV_LOGIN=true in your .env, then visit http://localhost:8080/login and sign in with:
- Username:
admin - Password:
admin
This bypasses Login.gov and the harvest_user allow list for local development only. It is disabled by default and remains disabled in deployed environments. Login.gov sandbox remains available via the link on the login page.
Alternatively, you can log in with Login.gov. For local development you must have an account at the login.gov sandbox https://idp.int.identitysandbox.gov. (Click "Create an account" if you don't already have one.)
Add your user account to the local app, using an email address that matches your login.gov sandbox account (see also "user management" below):
% docker compose exec app flask user add your.i.name@gsa.gov --name yourName
User added successfully!
Now you should be able to log in at http://localhost:8080/login, add an organization, and add a feed to it.
This is primarily a python project.
We use Ruff to format and lint our Python files. If you use VS Code, you can install the formatter here.
- This repo contains pre-commit actions. Learn how to configure your IDE to run those here.
- Create a branch from
main. We prefer short descriptive branch names. - To test changes in the
developmentspace in Cloud.gov, merge changes into thedevelopbranch. Coordinate with other developers by announcing your plans in #datagov-devsecops.
The DCAT-US 3.0 JSON Schemas are not in this repo. They come from
GSA/dcat-us, tracked as a git submodule at
_external/dcat-us. Paths are defined once, in
harvester/utils/schema_paths.py — import
from there rather than rebuilding paths from __file__.
Don't edit anything under _external/dcat-us. It belongs to GSA/dcat-us; open
a PR against that repo instead. A commit here records only which dcat-us
commit to check out, never file contents.
DCAT-US 1.1 is different — it has no GSA/dcat-us equivalent and stays
vendored in this repo under schemas/dcatus1.1/.
Cloning fresh:
% git clone --recurse-submodules https://github.com/GSA/datagov-harvester.git
If you already cloned without that flag:
% git submodule update --init _external/dcat-us
To stop having to remember the flag on every clone and pull:
% git config --global submodule.recurse true
An uninitialized submodule leaves _external/dcat-us as an empty directory,
not a missing one, so the failure is easy to misread. Two things to know:
make buildwill not fix it.docker-compose.ymlbind-mounts.:/app, so the container sees your host working tree, empty directory included. Rebuilding the image accomplishes nothing.- The symptom is a
FileNotFoundErrorfrombuild_dcatus3_validatornaming the directory and the fix. Anything touching DCAT-US 3.0 validation raises it, including at test-collection time.
Confirm with git submodule status. A leading - means uninitialized:
-24f6f1e... _external/dcat-us # not initialized — run the update above
24f6f1e... _external/dcat-us (...) # good (note the leading space)
+abc1234... _external/dcat-us (...) # checked out at a different commit than pinned
Don't use sparse-checkout inside the submodule. Sparse patterns live in
.git/modules/_external/dcat-us/info/sparse-checkout, which cannot be
committed, so CI and fresh clones get the full tree regardless. All it does is
make your machine disagree with CI about which schema files exist — in the one
code path whose entire job is schema validation.
The submodule is pinned to a specific dcat-us commit on purpose, so upgrades
are deliberate. Dependabot opens a monthly PR that bumps it; you can also do it
by hand.
Note that --remote follows the branch named in .gitmodules, which is
dcat-us's main — not this repo's main.
% git -C _external/dcat-us rev-parse HEAD # record the old SHA first
% git submodule update --remote _external/dcat-us
Review what actually moved, then test:
% git diff --submodule=log # commit range
% git -C _external/dcat-us diff "$old" HEAD -- jsonschema/definitions
% poetry run pytest tests/unit
Commit only the gitlink (_external/dcat-us) — there are no file changes to
stage.
A schema change can legitimately alter which harvest records validate, so failing tests here are the point of pinning, not an obstacle. Fix the code or reject the bump. Never skip the tests to land the SHA.
Local configuration should be stored in .env, which is ignored by git.
Use .env.sample as the template for required local variables.
Do not commit real credentials, environment-specific secrets, or generated .env files.
Production and deployed environment variables are provided by the deployment platform.
If you absolutely need to hit a breakpoint in your Flask app, you can setup local Flask debugging in your IDE.
NOTE: To use the VS-Code debugger, you will first need to sacrifice the reloading support for flask
-
Build new containers with development requirements by running
make build-dev -
Launch containers by running
make up-debug -
In VS-Code, launch debug process
Python: Remote Attach -
Set breakpoints
-
Visit the site at
http://localhost:8080and invoke the route which contains the code you've set the breakpoint on.
We use poetry to manage this project, and to run the tests. Install poetry here. (Poetry is also installed and run automatically within the app container, which is why you didn't need it to get the app up and running.)
Once poetry is installed, poetry install installs dependencies into a local virtual environment.
To update poetry itself locally (matching CI, which will always use the latest version), run poetry self update (or make poetry-update).
A number of "test" and "test-*" targets are defined in the Makefile.
For tests to pass, you may have to pull the latest MDTranslator. Use docker compose pull to get the latest versions of the docker images.
make test and make test-integration run against the database in .env
(localhost:5432). Integration tests reset schema/data during the run, so use
make re-up or make load-test-data afterward if you need fixture data back in
the dev app.
If you've added, updated, or removed any python dependencies, be sure to export requirements.txt:
poetry export -f requirements.txt --output requirements.txt --without-hashesWhen altering the db during development, you first want to stamp the db before making any changes to the model.
make clean up
docker compose exec app bashOnce inside the container, you run:
flask db stamp headApply your changes to the model file, then run:
flask db migrate -m "your migration message here"Then, finally, to apply your changes in place to the local db, run:
flask db upgradeGithub workflows automatically deploy:
- to the
developmentspace when thedevelopbranch is updated - to
stagingandprodwhen themainbranch is updated
Data.gov team members can deploy to development from the command line. The remainder of this document provides background on the Cloud.gov configuration.
Warning: this documentation has not been tested recently!
A database service is required for use on cloud.gov.
In a given Cloud Foundry space, a db can be created with
cf create-service <service offering> <plan> <service instance>.
In dev, for example, the db was created with
cf create-service aws-rds micro-psql datagov-harvest-db.
Creating databases for the other spaces should follow the same pattern, though the size may need to be adjusted (see available AWS RDS service offerings with cf marketplace -e aws-rds).
Any created service needs to be bound to an app with cf bind-service <app> <service>. With the above example, the db can be bound with
cf bind-service harvesting-logic datagov-harvest-db.
Alternately, you can just push the app up and it will bind with the services so long as they are named following the expected pattern in manifest.yml.
The harvester also expects an OpenSearch service named datagov-catalog-opensearch. The provisioning script creates it with Cloud.gov's aws-elasticsearch broker and requests OpenSearch_2.11, using es-medium in development and es-medium-ha in staging and es-large in production.
A user provided service by the name of datagov-harvest-secrets is also expected to be in place and populated with the following secrets:
- CF_SERVICE_AUTH
- CF_SERVICE_USER
- FLASK_APP_SECRET_KEY
- HARVEST_API_TOKEN
- OPENID_PRIVATE_KEY
CF_SERVICE_* variables can be extracted from from service-keys by running cf service-key ci-deployer datagov-harvest-deployer in the appropriate space.
datagov-harvest-validator binds its own datagov-harvest-validator-secrets and datagov-harvest-validator-db instead, so a compromised validator (it takes public, server-side URL-fetch requests) doesn't inherit the admin app's full secret set (GSA/data.gov#6293). create_cloudgov_services.sh creates both empty/placeholder if missing, but only the app code enforces that FLASK_APP_SECRET_KEY/HARVEST_API_TOKEN are non-empty - it never checks their value, and nothing the validator serves is @login_required, so these two can be any throwaway string, independent of the admin app's real ones. CF_SERVICE_USER/CF_SERVICE_AUTH/NEW_RELIC_LICENSE_KEY must be the same real values as datagov-harvest-secrets - LoadManager authenticates to the CF API with them at import time regardless of which app is running, so fake values crash the boot rather than sitting unused. After create_cloudgov_services.sh runs (or by hand the first time), populate it once per space:
cf cups datagov-harvest-validator-secrets -p '{
"FLASK_APP_SECRET_KEY": "<any random string, independent of the admin app'"'"'s>",
"HARVEST_API_TOKEN": "<any random string, independent of the admin app'"'"'s>",
"CF_SERVICE_AUTH": "<same value as datagov-harvest-secrets>",
"CF_SERVICE_USER": "<same value as datagov-harvest-secrets>",
"NEW_RELIC_LICENSE_KEY": "<same value as datagov-harvest-secrets>"
}'datagov-harvest-validator-db needs no manual step - create_cloudgov_services.sh populates it with a fake, non-resolving URI, which is all harvester/__init__.py's create_engine() call needs to boot (it never connects at construction, and nothing the validator serves ever queries it).
Note: we prefer that you deploy to the development environment by pushing to the develop branch, which triggers deployment. That approach provides better team visibility. However, there are circumstances where deploying from the command line is necessary; for example if a failing action is preventing deployment.
-
Ensure you have a
manifest.ymlandvars.development.ymlfile configured for your Flask application. The vars file may include variables:app_name: datagov-harvest database_name: datagov-harvest-db route_external: harvest-dev.data.gov route_internal: datagov-harvest-dev.apps.internal proxy_instances: 1 basic_auth_enabled: on
-
Deploy the application using Cloud Foundry's
cf pushcommand with the variable file:poetry export -f requirements.txt --output requirements.txt --without-hashes cf push --vars-file vars.development.yml
This is an nginx app which owns the public route and proxies traffic to the internal Flask app route. It also splits off validator traffic: /validate, /validate/, /api/validate, and /api/v1/validate are proxied to datagov-harvest-validator instead of datagov-harvest (see proxy/nginx-common.conf and proxy/nginx-maps.conf). Everything else goes to datagov-harvest as before.
This is a Flask app which manages the configuration of harvest sources, organizations, and the creation of harvest jobs.
Runs the identical codebase as datagov-harvest, deployed as a separate Cloud Foundry app so public validator traffic (schema validation, including the server-side URL-fetch option) can't consume the admin app's capacity or be scaled/restarted independently of it (GSA/data.gov#6293). Its manifest entry:
- Binds its own
datagov-harvest-validator-dbanddatagov-harvest-validator-secretsrather than the admin app's — see Services > User provided for what goes in each and why. It does not bind OpenSearch or SMTP — nothing on the validator's request path uses either. - Still sets
CLIENT_ID/ISSUER/REDIRECT_URIeven though the validator never serves/loginor/callback:app/main/auth.pyandharvester/utils/general_utils.py'sSMTP_CONFIGboth dereferenceISSUER/REDIRECT_URIunconditionally at import time, and those modules load duringcreate_app()regardless of which routes a given app actually serves. - Sets
SKIP_DB_MIGRATIONS=trueso its own instance 0 doesn't also runflask db upgradeagainst the shared database (app-start.shchecks this before migrating). - Has its own
validator_instances/validator_memory_quota/validator_disk_quotavars per space, separate fromadmin_*.
The /validate* routes still exist on datagov-harvest itself (unchanged) - the split is only in how the proxy routes traffic, so reverting the two proxy/nginx-*.conf location/map additions is a full rollback to routing everything through the admin app.
The Data.gov team are the only intended users of the harvester admin app. Since users are rarely added or deleted, no user management UI has been built and adding/deleting users must be done via the command line.
To add a user:
cf run-task datagov-harvest --name "add new user" --command "flask user add xxx@gsa.gov --name xxx"
Or, if doing for local development:
docker compose exec app flask user add your.i.name@gsa.gov --name yourName
To delete a user:
cf run-task datagov-harvest --command "flask user remove xxx@gsa.gov"
You can add organizations using the harvester UI. Alternatively, you can run this command:
cf run-task datagov-harvest-admin --name "add new org" --command "flask org add 'Name of Org' --log https://some-url.png --id 1234"