Skip to content

Resume a calibration from the population of its last generation #361

Description

@soaressgabriel

Description

An interrupted or crashed search leaves its evaluations (and, with #359, its best candidate) but not the search itself: the documentation says a new directory is used for the next attempt, the runner refuses a directory that holds records (runner.py#L488-L495), and the next attempt starts from a fresh initial population. Passing <config>-calibrated.json back as -c restarts from the best point alone, with a new population around it, which is not the population the search had built.

Everything a resume needs is at hand. The callback receives, every generation, intermediate_result.population already scaled to parameter values, and population_energies (SciPy 1.18, DifferentialEvolutionSolver._result); CalibrationSettings.init accepts an explicit (S, N) array as the initial population; and the records of the first part are on disk.

Proposed solution

  • The callback writes population.json atomically after every generation (temporary file, renamed): the generation number, the free parameter names, the bounds, the fixed parameters, the members, their energies, and the settings that describe the search (variable, spinup_steps, stations, popsize, strategy, mutation, recombination, maxiter, seed).
  • rubem calibrate --resume (CalibrationSettings(resume=True)) with the same -o: refused without a population file; refused when the settings of the command differ from the saved ones, --maxiter and --workers excepted. The saved members are passed as init, x0 is omitted (SciPy overwrites population[0] with x0), maxiter becomes what is left of the saved one. The refusal of a directory with records is lifted for a resume, the records of both parts are consolidated into one table, and result.json records resumed_from (the generation) and counts the generations and evaluations of both parts.
  • SciPy re-evaluates the initial population it is given, so a resume repeats one generation of model runs; the documentation says so. The random draws of the resumed search are not the ones the interrupted search would have made: a resumed search starts from where the other stood, it is not the other continued candidate by candidate.

Tests that prove it: the population file is written every generation, atomically, with the members SciPy holds; a search interrupted after generation g and resumed ends with the generations of both parts counted and a table holding the records of both; a resume is refused without the file and with different settings; a directory with records is still refused without --resume.

Documentation: "An interrupted run", the run-directory warning, the result.json field list; changelog.

Alternative solutions

  • Save only the best member and restart around it. That is what -c <config>-calibrated.json does today; it throws the population away.
  • Save the population as .npy. A JSON file is readable without NumPy and carries the header the settings check needs in the same file.

Additional context

Part of the calibrator robustness work assessed on 2026-09-24. Touches _Progress like #360 (rebase, not a dependency); builds on #359 for termination.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions