You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The search stops only when SciPy's convergence test fires or --maxiter is reached. Three things a person watching a calibration needs are missing, and the reason the search ended is not recorded in its own right:
tol and atol of differential_evolution are not exposed; CalibrationSettings has neither (runner.py#L150-L237) and every search runs with the 1 % default the documentation describes ("The evaluation budget").
The per-generation callback logs and never stops (runner.py#L346-L376). A search whose best NSE has not moved for twenty generations keeps spending a population of model runs per generation until --maxiter.
A configuration whose every evaluation fails (a station mismatch the pre-checks cannot see under the zones aggregation, a MODFLOW setup that never converges) spends the whole budget before the runner reports "Every evaluation of the calibration failed" (runner.py#L586-L590). With the defaults that is 12 928 runs to learn nothing.
result.json records only SciPy's message and success; a stop asked by the callback gives success=False and the generic "callback function requested stop early".
Two SciPy facts shape the design. INADMISSIBLE_OBJECTIVE is finite, 1e30 (objective.py#L43-L52), so one failed member blows up the standard deviation of the population energies and the tol test cannot fire while any evaluation of the generation failed: whole-generation failure needs a rule of its own. And the callback first fires after generation 1: the initial population goes through the workers map before it (runner.py#L552), so a failure of the whole initial population is seen by the map, not by the callback.
Proposed solution
--tol (default 0.01) and --atol (default 0), passed through as tol and atol and recorded in settings.
--stall-generations K (default 0, off) and --stall-min-delta-nse D (default 0.001, read only with K): the callback keeps the best NSE of every generation, from intermediate_result.fun (NSE = 1 - sqrt(f / 1000) / 100, no records re-read), and stops the search when it improved by less than D over the last K generations.
A generation whose every evaluation failed (errors other than inadmissible) stops the search: a wrapper around executor.map reads the records of the batch it just ran and raises a dedicated exception; calibrate consolidates and raises the existing CalibrationError with the first error, after one population instead of the whole budget. Because the wrapper sees the initial population, a hopeless configuration costs one population.
The generation line gains the number of failed evaluations of the generation and the time since the previous line.
Tests that prove it: tol and atol reach the search arguments (the "documented arguments" test); a stall stop on an objective that stops improving ends with termination == "stalled" after K generations without a gain; a configuration whose every evaluation fails ends after the initial population (nfev equal to the population size) with the existing message; the generation line.
Documentation: "Settings", the convergence paragraph of "The evaluation budget", the result.json field list, "Progress on the terminal"; changelog.
Alternative solutions
Stop on stalled objective instead of NSE. The objective is a monotone function of the mean NSE, so the two are the same rule; NSE is the unit the user reads.
Detect the failed generation in the callback. Rejected: it fires only after generation 1, so the initial population would be spent twice over before the first check.
Additional context
Part of the calibrator robustness work assessed on 2026-09-24. Depends on #359 for the termination field; touches _Progress, which #361 touches too (rebase, not a dependency).
Description
The search stops only when SciPy's convergence test fires or
--maxiteris reached. Three things a person watching a calibration needs are missing, and the reason the search ended is not recorded in its own right:tolandatolofdifferential_evolutionare not exposed;CalibrationSettingshas neither (runner.py#L150-L237) and every search runs with the 1 % default the documentation describes ("The evaluation budget").--maxiter.zonesaggregation, a MODFLOW setup that never converges) spends the whole budget before the runner reports "Every evaluation of the calibration failed" (runner.py#L586-L590). With the defaults that is 12 928 runs to learn nothing.result.jsonrecords only SciPy'smessageandsuccess; a stop asked by the callback givessuccess=Falseand the generic "callback function requested stop early".Two SciPy facts shape the design.
INADMISSIBLE_OBJECTIVEis finite,1e30(objective.py#L43-L52), so one failed member blows up the standard deviation of the population energies and thetoltest cannot fire while any evaluation of the generation failed: whole-generation failure needs a rule of its own. And the callback first fires after generation 1: the initial population goes through theworkersmap before it (runner.py#L552), so a failure of the whole initial population is seen by the map, not by the callback.Proposed solution
--tol(default 0.01) and--atol(default 0), passed through astolandatoland recorded insettings.--stall-generations K(default 0, off) and--stall-min-delta-nse D(default 0.001, read only with K): the callback keeps the best NSE of every generation, fromintermediate_result.fun(NSE = 1 - sqrt(f / 1000) / 100, no records re-read), and stops the search when it improved by less than D over the last K generations.inadmissible) stops the search: a wrapper aroundexecutor.mapreads the records of the batch it just ran and raises a dedicated exception;calibrateconsolidates and raises the existingCalibrationErrorwith the first error, after one population instead of the whole budget. Because the wrapper sees the initial population, a hopeless configuration costs one population.termination(from Write the result of a calibration that ends before the search finishes #359) gainsconverged(SciPy's tolerance),maxiterandstalled(with the K and D that fired).successkeeps the optimizer's flag; the documentation saysterminationis the field to read.Tests that prove it:
tolandatolreach the search arguments (the "documented arguments" test); a stall stop on an objective that stops improving ends withtermination == "stalled"after K generations without a gain; a configuration whose every evaluation fails ends after the initial population (nfevequal to the population size) with the existing message; the generation line.Documentation: "Settings", the convergence paragraph of "The evaluation budget", the
result.jsonfield list, "Progress on the terminal"; changelog.Alternative solutions
Additional context
Part of the calibrator robustness work assessed on 2026-09-24. Depends on #359 for the
terminationfield; touches_Progress, which #361 touches too (rebase, not a dependency).