Description
rubem run offers two ways to handle the input validation, and neither one lets a user run past a blocking problem and still see it:
- Default: every check runs, the problems are logged as warnings, and any blocking problem raises
ConfigurationError, so the run stops with exit code 1 (model_configuration.py#L502-L507).
-s/--skip-inputs-validation (cli.py#L117-L124): the content checks do not run at all. InputRasterFiles, InputRasterSeries and InputTableFiles return early. The lookup table checks and the runoff coefficient domain are skipped (model_configuration.py#L111-L121), and so are the series resolver checks and the grid cell size check. The only output is a "... validation is disabled." warning, so the user never learns which problems the inputs have.
When exploring a dataset, or reproducing a published run whose inputs fail a check, the user has to choose between not running and running blind.
Proposed solution
Add an opt-in rubem run --allow-blocking-problems option (no short form):
- It runs the same checks as the default. Every problem is still logged: non-blocking ones at
WRN as today, blocking ones at ERR. A final ERR line states that the simulation continues despite N blocking problems. The run then goes on instead of raising ConfigurationError.
- It cannot be combined with
-s. That is a usage error (exit code 2), because -s skips the very checks this option exists to report.
- The library gets a matching keyword,
ModelConfiguration(..., allow_blocking_problems=False), so the command line only passes it through. ModelConfiguration.problems keeps the full list so that a caller can inspect it.
- The default does not change. Blocking problems still stop the run unless the option is given, including the zero-denominator domains (Tw >= 1, C_wp >= 1).
- Failures that are not
Problems stay fatal because the model cannot be built without them: schema validation (pydantic), missing lookup table files (FileNotFoundError), no raster format enabled (ValueError), and the DEM/clone/georeference mismatch that OutputRasterBase refuses.
- Exit code: 0 when the run completes, 1 when it fails.
- The option is not added to the deprecated
rubem -c <config> spelling (_LEGACY_OPTIONS).
- Tests: a configuration with one blocking problem exits 1 without the option. With the option it logs the problem at
ERR and starts the simulation. Combining the option with -s exits 2.
- Docs: the
rubem run -h block in the user guide and the changelog.
Alternative solutions
- Let the option relax only the data domain problems (raster content ranges, Tw >= 1, C_wp >= 1, grid cell size) and keep the structural ones blocking (missing series steps, clone geometry or coordinate reference system mismatch, duplicate step matches, unreadable tables). Forcing a structural problem only moves the failure into the run. However,
Problem has no attribute for that distinction today, and every place that creates a blocking problem would have to be classified (input rasters, raster series, series resolver, lookup tables, raster content, grid cell size).
- Change
-s into "check and report, never block". This is rejected: it changes the documented meaning of an existing option, and users who rely on -s to skip the content checks would lose that.
Additional context
Open point: should a forced run also record the blocking problems in metadata.json, so that the outputs themselves say they were produced from inputs that failed validation, not only the log? metadata.json exists only for configuration format 1.0.
The open PRs #342 (Python API) and #344 (rubem calibrate) pass validate_input through to ModelConfiguration. Whichever of the two sides lands later (this change, or #342 and #344) must expose the new keyword on the same path, or state why not (calibration evaluations already run with validate_input=False after one validated load).
Description
rubem runoffers two ways to handle the input validation, and neither one lets a user run past a blocking problem and still see it:ConfigurationError, so the run stops with exit code 1 (model_configuration.py#L502-L507).-s/--skip-inputs-validation(cli.py#L117-L124): the content checks do not run at all.InputRasterFiles,InputRasterSeriesandInputTableFilesreturn early. The lookup table checks and the runoff coefficient domain are skipped (model_configuration.py#L111-L121), and so are the series resolver checks and the grid cell size check. The only output is a "... validation is disabled." warning, so the user never learns which problems the inputs have.When exploring a dataset, or reproducing a published run whose inputs fail a check, the user has to choose between not running and running blind.
Proposed solution
Add an opt-in
rubem run --allow-blocking-problemsoption (no short form):WRNas today, blocking ones atERR. A finalERRline states that the simulation continues despite N blocking problems. The run then goes on instead of raisingConfigurationError.-s. That is a usage error (exit code 2), because-sskips the very checks this option exists to report.ModelConfiguration(..., allow_blocking_problems=False), so the command line only passes it through.ModelConfiguration.problemskeeps the full list so that a caller can inspect it.Problems stay fatal because the model cannot be built without them: schema validation (pydantic), missing lookup table files (FileNotFoundError), no raster format enabled (ValueError), and the DEM/clone/georeference mismatch thatOutputRasterBaserefuses.rubem -c <config>spelling (_LEGACY_OPTIONS).ERRand starts the simulation. Combining the option with-sexits 2.rubem run -hblock in the user guide and the changelog.Alternative solutions
Problemhas no attribute for that distinction today, and every place that creates a blocking problem would have to be classified (input rasters, raster series, series resolver, lookup tables, raster content, grid cell size).-sinto "check and report, never block". This is rejected: it changes the documented meaning of an existing option, and users who rely on-sto skip the content checks would lose that.Additional context
Open point: should a forced run also record the blocking problems in
metadata.json, so that the outputs themselves say they were produced from inputs that failed validation, not only the log?metadata.jsonexists only for configuration format 1.0.The open PRs #342 (Python API) and #344 (
rubem calibrate) passvalidate_inputthrough toModelConfiguration. Whichever of the two sides lands later (this change, or #342 and #344) must expose the new keyword on the same path, or state why not (calibration evaluations already run withvalidate_input=Falseafter one validated load).