You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Calibrating RUBEM today means editing the nine CALIBRATION parameters by hand, running the model, comparing the station time series with the observed streamflow and iterating. The correction of the soil moisture coefficient (#318, #321) invalidated the published parameter sets of the three basins, so a reproducible, automatic calibration is now the way to recover them, and the same tool is needed by anyone applying the model to a new basin.
This issue is the single place for the calibration command: it absorbs #345 (documentation of the method and the dataset smoke test) and #348 (the capabilities found missing when the command was compared with the calibration procedure the model was calibrated with before).
Proposed solution
A rubem calibrate command, built on the Python API (#341), with an independent implementation (SciPy API and the Nash-Sutcliffe definition; no external calibration script is ported).
Objective and observations
FO = 1000 (100 (1 - NSE))^2, minimized with SciPy's differential_evolution. The NSE compares the simulated series of one output variable at the sample stations (arn, the accumulated runoff in m3/s, by default; --variable overrides it) with the observed series per station id, averaged over the stations that have a value, with a guard for a station of zero variance, excluding the first --spinup-steps steps.
Observed series in the table layout of the model's own time series (; separated, header 0;<id>;..., one row per step) or as a PCRaster time series with its header (Write the PCRaster header in the time series outputs #347); a file without a header is refused naming the two accepted layouts. Gaps: NaN, negative values, -9999 and the PCRaster missing value 1e31.
--stations ID[,ID...] chooses the stations that enter the objective; the others are still measured, which gives a calibration/validation split.
Per-station goodness-of-fit for every evaluation: valid pairs, mean and standard deviation of both series, Pearson correlation, RMSE and NSE.
An observed series that shares no time step (after the spin-up) or no station with what the configuration samples is refused before the search starts.
Decision space
Free parameters alpha, beta, w1, w2, rcd, f, alpha_gw, x within the ranges of appsettings.json; w3 = 1 - w1 - w2 is derived, with the linear constraint w1 + w2 <= 1.
--bound NAME=MIN:MAX (repeatable) narrows the range of a parameter for a run; --fix NAME=VALUE (repeatable) takes a parameter out of the decision vector and keeps its value in every candidate. Both are recorded in result.json.
strategy="best1exp", init="sobol" (population max(5, popsize * N) rounded up to a power of two, N the number of free parameters), x0 from the configuration, mutation=(0.5, 1.0), recombination=0.7, no polishing, updating="deferred"; --seed, --maxiter, --popsize, --workers, --temp-dir, and --init, --strategy, --polish/--no-polish for other settings.
The ProcessPoolExecutor handed to differential_evolution as workers=executor.map is the isolation layer (spawn, one task per worker, silenced workers); each worker builds the derived configuration (format 1.0 document: every raster series disabled, only the target time series enabled, a per-evaluation temporary output directory, the candidate parameters) and calls Model.run() in-process, never run_isolated(), which would nest pools. One JSON record per evaluation, written atomically; failures are recorded and ranked last.
Progress on the terminal: the budget, one line per generation (best objective and NSE) and the closing summary. An interrupted or failed search still consolidates the evaluations; a worker killed by the system is reported with the advice to lower --workers.
Artifacts, in the run directory
observed.csv (per station: selection, pairs in the compared window, dropped values, summary statistics), evaluations.csv (one row per evaluation: start time, parameters, NSE, objective, error), stations.csv and best_<variable>.csv (the per-station metrics and the simulated series of the best candidate beside the observed one), result.json (best parameters including the derived w3, NSE, evaluations, generations, settings) and <config>-calibrated.json, the input configuration with the best parameters in its own format.
Packaging, tests and documentation
scipy as the optional extra rubem[calibration], installed by the CI test job; a clean error naming the extra when it is missing; ci/dev-requirements.in committed and the audit lock recompiled.
Unit tests of the objective, the metrics, the readers, the decision space, the worker and the runner (a deterministic small differential evolution on the synthetic dataset whose optimum is the configuration itself), and of the command.
A dataset pytest marker with the RUBEM_DATASET_DIR convention (tests that need the published datasets skip elsewhere) and a smoke test of the calibration on the Ipojuca basin over a short window with a fixed drainage network.
A calibration page in the documentation: the objective, the parameters and their bounds, the search settings and the evaluation budget, the observed series layouts, the fixed LDD raster a calibration on arn needs (lddcreate does not reproduce the network between calls on real basins), the artifacts, the process model and hints for large machines; the command in the user guide; the Python entry point named in the API page.
Alternative solutions
Running Model.run_isolated() inside the differential evolution workers doubles the process layers; a single pool whose workers run the model in-process is enough because each worker handles one evaluation and exits. Accepting a head-less observed file by numbering its columns would silently mislabel stations; refusing it with the expected layouts is safer and the conversion is one header line.
Additional context
Stack K of the modernization plan, brought forward before the architecture work; scientific motivation in #217 (the effect of #321 on the three basins). Delivered by #344.
Description
Calibrating RUBEM today means editing the nine
CALIBRATIONparameters by hand, running the model, comparing the station time series with the observed streamflow and iterating. The correction of the soil moisture coefficient (#318, #321) invalidated the published parameter sets of the three basins, so a reproducible, automatic calibration is now the way to recover them, and the same tool is needed by anyone applying the model to a new basin.This issue is the single place for the calibration command: it absorbs #345 (documentation of the method and the dataset smoke test) and #348 (the capabilities found missing when the command was compared with the calibration procedure the model was calibrated with before).
Proposed solution
A
rubem calibratecommand, built on the Python API (#341), with an independent implementation (SciPy API and the Nash-Sutcliffe definition; no external calibration script is ported).Objective and observations
FO = 1000 (100 (1 - NSE))^2, minimized with SciPy'sdifferential_evolution. The NSE compares the simulated series of one output variable at the sample stations (arn, the accumulated runoff in m3/s, by default;--variableoverrides it) with the observed series per station id, averaged over the stations that have a value, with a guard for a station of zero variance, excluding the first--spinup-stepssteps.;separated, header0;<id>;..., one row per step) or as a PCRaster time series with its header (Write the PCRaster header in the time series outputs #347); a file without a header is refused naming the two accepted layouts. Gaps:NaN, negative values,-9999and the PCRaster missing value1e31.--stations ID[,ID...]chooses the stations that enter the objective; the others are still measured, which gives a calibration/validation split.Decision space
alpha, beta, w1, w2, rcd, f, alpha_gw, xwithin the ranges ofappsettings.json;w3 = 1 - w1 - w2is derived, with the linear constraintw1 + w2 <= 1.--bound NAME=MIN:MAX(repeatable) narrows the range of a parameter for a run;--fix NAME=VALUE(repeatable) takes a parameter out of the decision vector and keeps its value in every candidate. Both are recorded inresult.json.C_wp <= 1, Validate the domain of the weighted runoff coefficient (C_wp <= 1) from the tables and the calibration weights #328) is rejected without running the model. Pending: it needs the check merged by Validate the domain of the weighted runoff coefficient from the tables and the weights #336, which the branch does not contain until it is rebased ontomain.Search and process
strategy="best1exp",init="sobol"(populationmax(5, popsize * N)rounded up to a power of two,Nthe number of free parameters),x0from the configuration,mutation=(0.5, 1.0),recombination=0.7, no polishing,updating="deferred";--seed,--maxiter,--popsize,--workers,--temp-dir, and--init,--strategy,--polish/--no-polishfor other settings.ProcessPoolExecutorhanded todifferential_evolutionasworkers=executor.mapis the isolation layer (spawn, one task per worker, silenced workers); each worker builds the derived configuration (format 1.0 document: every raster series disabled, only the target time series enabled, a per-evaluation temporary output directory, the candidate parameters) and callsModel.run()in-process, neverrun_isolated(), which would nest pools. One JSON record per evaluation, written atomically; failures are recorded and ranked last.--workers.Artifacts, in the run directory
observed.csv(per station: selection, pairs in the compared window, dropped values, summary statistics),evaluations.csv(one row per evaluation: start time, parameters, NSE, objective, error),stations.csvandbest_<variable>.csv(the per-station metrics and the simulated series of the best candidate beside the observed one),result.json(best parameters including the derivedw3, NSE, evaluations, generations, settings) and<config>-calibrated.json, the input configuration with the best parameters in its own format.Packaging, tests and documentation
scipyas the optional extrarubem[calibration], installed by the CI test job; a clean error naming the extra when it is missing;ci/dev-requirements.incommitted and the audit lock recompiled.datasetpytest marker with theRUBEM_DATASET_DIRconvention (tests that need the published datasets skip elsewhere) and a smoke test of the calibration on the Ipojuca basin over a short window with a fixed drainage network.arnneeds (lddcreatedoes not reproduce the network between calls on real basins), the artifacts, the process model and hints for large machines; the command in the user guide; the Python entry point named in the API page.Alternative solutions
Running
Model.run_isolated()inside the differential evolution workers doubles the process layers; a single pool whose workers run the model in-process is enough because each worker handles one evaluation and exits. Accepting a head-less observed file by numbering its columns would silently mislabel stations; refusing it with the expected layouts is safer and the conversion is one header line.Additional context
Stack K of the modernization plan, brought forward before the architecture work; scientific motivation in #217 (the effect of #321 on the three basins). Delivered by #344.