Description
An interrupted or crashed search leaves its evaluations (and, with #359, its best candidate) but not the search itself: the documentation says a new directory is used for the next attempt, the runner refuses a directory that holds records (runner.py#L488-L495), and the next attempt starts from a fresh initial population. Passing <config>-calibrated.json back as -c restarts from the best point alone, with a new population around it, which is not the population the search had built.
Everything a resume needs is at hand. The callback receives, every generation, intermediate_result.population already scaled to parameter values, and population_energies (SciPy 1.18, DifferentialEvolutionSolver._result); CalibrationSettings.init accepts an explicit (S, N) array as the initial population; and the records of the first part are on disk.
Proposed solution
- The callback writes
population.json atomically after every generation (temporary file, renamed): the generation number, the free parameter names, the bounds, the fixed parameters, the members, their energies, and the settings that describe the search (variable, spinup_steps, stations, popsize, strategy, mutation, recombination, maxiter, seed).
rubem calibrate --resume (CalibrationSettings(resume=True)) with the same -o: refused without a population file; refused when the settings of the command differ from the saved ones, --maxiter and --workers excepted. The saved members are passed as init, x0 is omitted (SciPy overwrites population[0] with x0), maxiter becomes what is left of the saved one. The refusal of a directory with records is lifted for a resume, the records of both parts are consolidated into one table, and result.json records resumed_from (the generation) and counts the generations and evaluations of both parts.
- SciPy re-evaluates the initial population it is given, so a resume repeats one generation of model runs; the documentation says so. The random draws of the resumed search are not the ones the interrupted search would have made: a resumed search starts from where the other stood, it is not the other continued candidate by candidate.
Tests that prove it: the population file is written every generation, atomically, with the members SciPy holds; a search interrupted after generation g and resumed ends with the generations of both parts counted and a table holding the records of both; a resume is refused without the file and with different settings; a directory with records is still refused without --resume.
Documentation: "An interrupted run", the run-directory warning, the result.json field list; changelog.
Alternative solutions
- Save only the best member and restart around it. That is what
-c <config>-calibrated.json does today; it throws the population away.
- Save the population as
.npy. A JSON file is readable without NumPy and carries the header the settings check needs in the same file.
Additional context
Part of the calibrator robustness work assessed on 2026-09-24. Touches _Progress like #360 (rebase, not a dependency); builds on #359 for termination.
Description
An interrupted or crashed search leaves its evaluations (and, with #359, its best candidate) but not the search itself: the documentation says a new directory is used for the next attempt, the runner refuses a directory that holds records (runner.py#L488-L495), and the next attempt starts from a fresh initial population. Passing
<config>-calibrated.jsonback as-crestarts from the best point alone, with a new population around it, which is not the population the search had built.Everything a resume needs is at hand. The callback receives, every generation,
intermediate_result.populationalready scaled to parameter values, andpopulation_energies(SciPy 1.18,DifferentialEvolutionSolver._result);CalibrationSettings.initaccepts an explicit(S, N)array as the initial population; and the records of the first part are on disk.Proposed solution
population.jsonatomically after every generation (temporary file, renamed): the generation number, the free parameter names, the bounds, the fixed parameters, the members, their energies, and the settings that describe the search (variable,spinup_steps,stations,popsize,strategy,mutation,recombination,maxiter,seed).rubem calibrate --resume(CalibrationSettings(resume=True)) with the same-o: refused without a population file; refused when the settings of the command differ from the saved ones,--maxiterand--workersexcepted. The saved members are passed asinit,x0is omitted (SciPy overwritespopulation[0]withx0),maxiterbecomes what is left of the saved one. The refusal of a directory with records is lifted for a resume, the records of both parts are consolidated into one table, andresult.jsonrecordsresumed_from(the generation) and counts the generations and evaluations of both parts.Tests that prove it: the population file is written every generation, atomically, with the members SciPy holds; a search interrupted after generation g and resumed ends with the generations of both parts counted and a table holding the records of both; a resume is refused without the file and with different settings; a directory with records is still refused without
--resume.Documentation: "An interrupted run", the run-directory warning, the
result.jsonfield list; changelog.Alternative solutions
-c <config>-calibrated.jsondoes today; it throws the population away..npy. A JSON file is readable without NumPy and carries the header the settings check needs in the same file.Additional context
Part of the calibrator robustness work assessed on 2026-09-24. Touches
_Progresslike #360 (rebase, not a dependency); builds on #359 fortermination.