This is triggered by a recent Slack discussion on deciding what field names we should use on one of date column in harvest_job table. The discussion itself maybe can easily be resolved by adding/renaming certain fields on the db/ui, but in the future if we have other similar discussion on the current way of how harvest jobs are created, we can come back to this ticket to see my thought on it.
Three components in harvest job creation:
Harvest Source table:
| id,... |
schedule |
date_next_run |
| ... |
weekly |
2025-09-15 15:12:44 |
| ... |
daily |
2025-08-01 01:47:28 |
| ... |
daily |
|
| ... |
manual |
|
| ... |
|
|
date_next_run can be empty, a future date, or a past date.
date_next_run will be filled in by Job scheduler, or it gets emptied when schedule field is modified.
- An empty
date_next_run means either the schedule is manual, or it is waiting for the next Job scheduler run to figure out a value. Once Job scheduler is run, there should not be any empty date_next_run in the table, except manual ones.
Harvest Job table:
| id,... |
date_created |
date_started |
date_completed |
status |
| ... |
2025-07-15 15:12:44 |
2025-07-15 15:20:11 |
2025-07-15 16:06:38 |
complete |
| ... |
2025-08-15 15:12:44 |
2025-07-15 15:20:11 |
|
in_progress |
| ... |
2025-08-15 16:12:44 |
|
|
new |
| ... |
|
|
|
|
- all dates should be past dates
- Jobs are created by Job scheduler, or by developer manually. No job can be created if the a
new or in_progress job already exists for that source.
- A job with a
new status will be deleted when its harvest source schedule field is changed.
Job Scheduler
- This runs on a cron, say every 15 minutes.
- It reads Harvest Source table, and writes Harvest Job table to create a job for each non-manual sources if the
date_next_run is past or empty.
- After a harvest job is created, it writes to Harvest Source table to update
date_next_run according to its schedule.
- If there is already a harvest job in
new or in_progress status, no new job will be created for that source.
Other thoughts
- Harvest job created outside of Job Scheduler, i.e., manual by developer, has no impact on
date_next_run.
- We will have the same schedule shifting issue as CKAN. After a few months, an source that originally gets harvested every morning might end up in the afternoon.
This is triggered by a recent Slack discussion on deciding what field names we should use on one of date column in harvest_job table. The discussion itself maybe can easily be resolved by adding/renaming certain fields on the db/ui, but in the future if we have other similar discussion on the current way of how harvest jobs are created, we can come back to this ticket to see my thought on it.
Three components in harvest job creation:
Harvest Source table:
date_next_runcan be empty, a future date, or a past date.date_next_runwill be filled in by Job scheduler, or it gets emptied when schedule field is modified.date_next_runmeans either the schedule is manual, or it is waiting for the next Job scheduler run to figure out a value. Once Job scheduler is run, there should not be any emptydate_next_runin the table, except manual ones.Harvest Job table:
neworin_progressjob already exists for that source.newstatus will be deleted when its harvest sourceschedulefield is changed.Job Scheduler
date_next_runis past or empty.date_next_runaccording to its schedule.neworin_progressstatus, no new job will be created for that source.Other thoughts
date_next_run.