0.6: new aircraft tasks, honest baselines, release-ready docs and repo - #34
Merged
Merged
Conversation
… the answer The environments claim a learned policy has something real to beat, and no learned policy's numbers are published. This is the recording side of that comparison; the training belongs elsewhere. Results go in data/rl_results.json through target_gym.rl_results.record_result, which stamps each entry with a fingerprint of the environment it was measured against. That is the mechanism the MPC baselines already use, reused rather than reinvented: training costs GPU-hours, the answer only moves when the environment or the agent moves, and a number that outlived its environment should be refused rather than quoted. The fingerprint for a learned result is deliberately narrower than the one the shipped baselines use. An agent never calls a PID, so re-tuning a controller must not throw away a training run; a change to the dynamics, the reward, the parameters or the integrator must. environment_fingerprint is composed separately rather than by reusing the baseline digest, so adding it could not change any of the sixteen fingerprints already recorded in data/baseline_returns.json -- checked, and none went stale. The agent's own configuration is stored verbatim and pointedly *not* fingerprinted. Two runs of the same agent at different learning rates are different results, not stale ones, and both are worth keeping; a tag separates them. Four tests, none of which asserts that RL beats anything -- whether it does is the question the library exists to ask, and a test that presumed the answer would be worthless. They check that a record names a real environment, that it is well formed, that it still describes that environment, and that it was scored over the environment's own episode so it sits in the same column as the PID and MPC numbers rather than a different one. That last check earned its place immediately: a round-trip record claiming 200 steps on the cstr was rejected, because that environment's episode is 100. The dependency runs one way, from the RL library to here. Ajax already depends on this package -- pinned, as it happens, to the branch merged earlier today -- and already handles the two things that would otherwise bias every number on these tasks: termination against truncation, and preserving the terminal observation at a time-limit truncation so the value bootstrap is correct. These episodes truncate at max_steps constantly. Installing TargetGym still drags in no RL framework. docs/rl-baselines.md records the design and two things to hold onto when there are numbers to read. If RL loses, the result is on trial rather than the environments, and one cross-check against stable-baselines3 -- already a dev dependency with a working PPO smoke test -- settles "your agent was under-trained". If RL wins, check what it is beating: this library has shipped a PID pinned to the edge of its search grid and an expert that could not fly a third of its own task's radius range, worth 3.5% and 31% once fixed. A win against a defective baseline measures the defect. Fast suite 1260 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC39GNmL25Z61NXKGoZmoK
Almost every choice in an RL-versus-classical-control comparison can be made in a way that decides the result before training starts. This writes them down first, so the design cannot drift toward whichever answer arrives. Two questions, run separately and never averaged. Tabula rasa asks whether a learned policy beats a tuned PID from nothing, which is the benchmark framing. Expert-based asks whether, given a tuned PID, learning improves on it -- the industrial framing, since nobody replaces a working loop with a random network. A no to the first and a yes to the second is coherent, and is the outcome a plant engineer would care about most. Three things measured in this repository shaped the design rather than being assumed. The MPC is not a peer. It calls _extract_x0(state) and reads the true state directly, and it plans against the exact dynamics; the PID and any learned policy see observations only, and several environments are deliberately partially observed -- the furnace hides 6 of 9 states, the reactor 7 of 11. So the PID is the peer and the MPC is an upper bound *with more information*. Beating it would be a claim needing explanation, and the first hypothesis should be a simulator exploit. Raw returns are not comparable across environments: episodes run from 100 steps to 1200 and the reward is bounded per step, so return scales with horizon. Cross -environment reporting uses mean reward per step, which is in [0, 1] and means the same thing everywhere. Within an environment the statistic is the paired per-seed difference against the PID on identical episodes. Observations span 9e-3 to 8.4e3 across the suite and about four orders of magnitude inside single environments, so normalisation is mandatory and is stated rather than left as a tuned flag; without it the experiment would partly measure a network's tolerance of unscaled inputs. The rest follows from evidence already in the repository. Hyperparameters are searched per environment because the PID and MPC are tuned per environment and a comparison against an untuned opponent measures the tuning -- re-tuning the aircraft PIDs was worth +26% to +139%, and the furnace's gains were pinned at a grid edge. The search runs on seeds disjoint from those reported, because the winner of a search is the maximum of noisy draws and quoting it on its own seeds publishes that bias; the aircraft PID tuner already follows exactly this rule. Results are reported as an interquartile mean with bootstrap intervals and a win rate rather than a mean, because the battery MPC reads +14.0 on the mean and -4.1 on the median while winning 1 seed in 10. Thirty agent seeds, because two seeds have misled this project three separate times and Ajax's scaling makes a hundred nearly free. Two choices are less obvious. The discount factor is set from physics -- gamma = exp(-delta_t / tau) -- because delta_t spans 0.05 s to 900 s here, so a fixed gamma would mean a five-second horizon in one environment and twenty-five hours in another; expressing it as a number of settling times makes it the same claim everywhere. And an average-reward agent is run as a secondary study, because these are continuing maintenance tasks ended by a time limit rather than episodic goal tasks, so discounting imposes a horizon the task does not have -- which distinguishes "RL cannot do this" from "the discounted formulation was the wrong tool". The expert arm is a bounded residual on the shipped PID, with alpha swept and always reported. That sweep is the falsifiability: at alpha = 1 the residual can overwrite the expert entirely and the method degenerates toward tabula rasa with an odd prior, so without the sweep "learning improved the PID" cannot be distinguished from "learning ignored the PID". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC39GNmL25Z61NXKGoZmoK
…rch space Two placeholders in the protocol are now numbers, and measuring them overturned one of the rules I had written. The draft set the discount from physics as gamma = exp(-delta_t / tau), a fixed number of the plant's own time constants. Measuring the time constants killed it. Linearising one step and reading the slowest stable eigenvalue gives the wrong quantity -- for the 2D aircraft that mode is fuel burn at 29 694 s and for the reactor it is xenon at 61 908 s, neither of which is what a controller acts through. The right measurement is the open-loop actuator-to-output step response, the same quantity relay tuning assumes: time to 63.2% of the total change, taken differentially between two step levels so drift cancels. That gives sensible, checkable numbers -- 23 s for the aircraft altitude loop, 66 minutes for the furnace crown, 15.5 hours for building thermal mass -- and it breaks the rule. The reactor's flux answers the rod in *one step*, because prompt neutron response really is that fast, so five time constants is a five-step horizon while the agent must hold that flux for 1200 steps against xenon it cannot see resolve. The turbine and battery are the same shape. A horizon set from the actuator response would have made three environments myopic by construction. So the discount is the episode: gamma = 1 - 1/N. Evaluation scores the undiscounted return over exactly N steps, so an effective horizon of N aligns what the agent optimises with what it is measured on, and anything shorter makes it deliberately blind to part of its own score. It also removes gamma as a free parameter, which matters in a comparison whose result is the point -- it cannot be tuned to flatter either side, and it is excluded from the search for that reason. The measurement keeps a job: checking the horizon covers the plant's response. On three environments it does not. The glass furnace would need 2.8x its episode for five open-loop time constants, HVAC 1.6x and the cement kiln 1.2x, so on those no controller can demonstrate steady-state holding and every score is partly a measure of the approach rather than of maintenance. That is a real limitation of the suite. It is recorded rather than fixed, because lengthening those episodes would invalidate every baseline number in the repository and is a decision to take on its own. The hyperparameter search space is published in full, identical across environments, so that "tuned" means one mechanical procedure rather than the experimenter's taste. Conventional continuous-control ranges; the point is that they were fixed before any result existed. Two exclusions are argued rather than assumed: gamma, above, and observation normalisation, which is held on because observations span four orders of magnitude inside single environments and turning it off would not produce a worse agent but a different experiment. scripts/measure_time_constants.py keeps the measurement runnable -- `make time-constants` -- since the discounts in the protocol are derived from it and a discount quoted against dynamics that have moved is the same stale claim the recorded baselines already guard against. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC39GNmL25Z61NXKGoZmoK
The content was strong and the packaging was not. A visitor met a 623-line README with no badges, whose first visual sat at line 131, whose install instructions were at line 353, and 22% of which was internal planning. The top of the README is now what a newcomer needs in order: what it is, that it is alive, what it looks like, and code that runs. Badges for version, Python support, CI and licence; three of the eighteen environment gifs -- a figure-8, a glass furnace and a boiler drum -- since working control systems are the most compelling thing this project has and they were buried; then pip install and a quickstart. That quickstart imported a name the package does not export. Nothing would have caught it: tests/test_docs.py globbed docs/ only, so the single most-read code sample in the project was the only one not executed. It now collects README blocks too, which is how the bug surfaced. The SB3 training block is marked `# doc: skip` -- tests/plane/test_agent.py already covers that path and 10 000 training steps do not belong in the fast job. Added a "where this fits" table against gymnax, brax and pc-gym, saying plainly what each is better at. A reader's first question is why not one of those, and answering it honestly sends the right person further in and saves the wrong one the time. Being the wrong tool for locomotion is not a weakness worth hiding. docs/index.md was stale -- it advertised eleven model review checks when there are thirteen, and omitted four of the twelve pages, including the testing and RL protocol pages written this week. It is now grouped by what a reader is trying to do (start here / beat the baselines / trust the numbers / contribute) rather than listed flat, and a check asserts every page is linked and every link resolves. The roadmap and known gaps moved to docs/roadmap.md, leaving three lines and a link. They are for contributors, and the README is the shop window. What is deliberately unchanged is the candour. "Where a baseline is weak, the docs say how weak" stays on the front page, and so does the table of failure modes -- irrecoverable states, non-minimum phase, transport delay. That is the differentiator, and sanding it down to look more polished would remove the reason the audience this wants would choose it. README 623 -> 551 lines. Doc tests 14 -> 16. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC39GNmL25Z61NXKGoZmoK
…k the docs The sixteen process and industrial plants render through render_kit.frame -- schematic panel, instrument stack, strip chart, one dark palette. The aircraft did not: it drew a sky-blue pygame scene with pure #00FF00 ground and a text box. Beside a furnace or a reactor in the same gallery that read as a different project rather than a different environment, and no amount of recolouring fixes it, because they are different genres. plane/rendering_console.py puts the flight inside that chrome. The scene stays, since an aircraft visibly climbing toward a commanded altitude is the most legible thing this project can show, but it becomes the schematic panel, with real gauges beside it and the altitude-versus-commanded strip underneath. It is drawn in matplotlib rather than blitted from pygame, so the anti-aliasing, typography and colours are the ones every other environment already uses. Four things needed measuring rather than guessing. The silhouette was first drawn by mirroring a wing about the fuselage, which is a plan-view convention and renders as a cross in side elevation; it is now a side profile with the wing edge-on, a vertical fin and a row of cabin windows, which is what makes a shape read as an airliner rather than a dart. The rotation matrix had the sign that pitches the nose down while climbing. The status bands were absolute metres, so a commanded four-kilometre climb showed ALARM for being mid-task; they are now relative to the altitude envelope. And the trail was drawn across a fixed span regardless of elapsed time, so two steps of history drew the same streak as two hundred. The old pygame renderer keeps its palette fix in the meantime -- sky lifted off near-black so the aircraft has a field to read against, clouds flattened to barely above the sky after they were more visible than the aircraft, hull brightened to be the brightest thing in frame. Documentation, alongside: Eighteen generated per-environment pages, Gymnasium-style -- picture, action space, observation space, rewards, starting state, episode end, arguments, in that order. Generated from the registry by scripts/generate_env_pages.py rather than written, because eighteen hand-maintained pages drift the first time a parameter moves. Action meanings come only from environments whose own docstrings state them, ten of eighteen; the rest show bounds with a blank meaning rather than a label invented by the generator. The homepage is a mosaic, install, and a working example. The six-gif grid it replaced was ragged for reasons HTML cannot fix: aspect ratios from 3.0:1 to 1.5:1 and frame counts from 90 to 160, so nothing lined up and the tiles drifted out of phase. One pre-rendered mosaic fixes both. It is animated WebP at the sources' own resolution -- GIF caps at 256 colours and compressed these instrument panels so badly that fitting one under a few megabytes meant downscaling until the gauge text was mush. WebP is 1962x710 at 1.8 MB, against 1224x450 at 2.8 MB for the GIF, which is kept as a fallback. Also: the right-hand table of contents is hidden on the landing page so the mosaic gets the width, and the comparison table is dropped. Two follow-ups are recorded in docs/roadmap.md rather than left implicit: action labels for the eight environments whose docstrings do not state them, and wiring both page generators' --check into CI so the pages cannot go stale. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC39GNmL25Z61NXKGoZmoK
The hand-drawn silhouette was a second aircraft: fourteen faces of A320 geometry already exist in plane3d/rendering.py, sourced from that environment's PHYSICS.md at 37.57 m, and drawing a different shape beside it meant the two environments showed different aircraft. The console now projects that mesh -- orthographic, camera along -y, faces drawn back to front by mean y, which is all the depth ordering a side elevation needs. The mesh's greyscale shading is kept and re-tinted onto the console ramp, so it belongs to the same drawing as the gauges without losing the information about which way each face points. Clouds, at three parallax depths keyed to the distance actually flown. The settled scene was motionless -- the aircraft is pinned mid-panel holding a constant altitude -- so nothing in it said it was moving at 230 m/s. The first attempt at alpha 0.16 filled the panel with grey blobs that competed with the aircraft, which is the opposite of the point; they are now 0.045 and below. The trail is a recent window rather than the whole flight. Over a 10 000-step episode the history holds a thousand samples, and drawing all of them compressed the climb into a vertical spike detached from the aircraft. The strip chart underneath is what shows the whole episode. Recorded while measuring the "it looks too steady" observation, which was correct and is not a rendering artefact: turbulence_sigma is 0.0 in both the default and the benchmark parameters of both aircraft, so the machinery is present -- Ornstein-Uhlenbeck gusts, gust_x/gust_z, wind shear -- and switched off. Settled pitch has a standard deviation of 0.027 degrees and altitude of 0.00 m. Turning it on works and costs return: sigma 0.6 gives 0.25 deg and 1.25 m for -6%, sigma 1.5 gives 0.63 deg and 3.12 m for -11%. Changing the default would invalidate every recorded baseline, so it is a decision to take deliberately rather than a side effect of making a video look better. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC39GNmL25Z61NXKGoZmoK
…wrong env Holding one altitude is a task the tuned PID finishes with 0.0 m of settled error, so it can no longer tell two controllers apart. plane_steps and plane_sine are the same plant and the same controllers with a commanded altitude that moves, which is what makes them informative. Measured with the shipped PID: hold 1595 at 0.0 m, steps 1472 at 57.6 m, sinusoid 870 at 120.3 m -- monotone in difficulty, and each exposes something a constant setpoint cannot. Steps gives the transient response repeatedly; the sinusoid gives the closed loop's bandwidth, since amplitude ratio and phase lag against frequency *are* its frequency response. A ramp and a chirp are implemented too and not registered: the ramp leaves 7.4 m of steady-state error because tracking one needs an extra integrator, and the chirp walks the loop off the end of its bandwidth until it traces triangles against a sine. The pattern is a parameter, defaulting to hold, so no recorded baseline, tuned gain or published number moves. Episode length follows the criterion already in use: 800 steps is 3.3 periods of the 240 s sinusoid, clearing the three periods a periodic task needs. Both inherit the aircraft's shock-stall seam at 308 m/s and are allowlisted for it in the conformance suite -- checked against the reported seam rather than assumed, since it is the same plant. Then a real defect, found because plane_steps rendered a clip whose total reward was byte-identical to plane's: 9783.4140625, the same trajectory. The media runner built its parameters from spec.params_cls() -- the bare dataclass defaults -- at four call sites covering both figures and videos, so EnvSpec.test_params was ignored by everything that draws. Wherever a spec overrides a parameter, the pictures disagreed with the benchmark. Media now takes the registered parameters, with episode length bounded to 600-1200 steps: the aircraft's own default is 10 000, which is 41 cycles of the sinusoid and unwatchable, and its benchmark 280 ends before the climb does. videos/plane3d/*_short.gif were hand-placed leftovers of an older layout that no generator refreshes, and they went stale silently -- the gallery showed pre-re-skin aircraft while the environment had moved on. Removed, and the README, the environment pages and the mosaic now use videos/<env>/pid_output.gif, which is what the runner actually writes. Two mosaics instead of one, because a vehicle moving through a scene and an instrument panel of a process are different pictures and neither read well side by side. Plants is 4x3 over twelve environments; aircraft is 3x2 and now shows the difficulty ladder rather than the same flat hold three times. Finishing the plane3d re-skin: render_side_scene had its own ground rectangle at (100, 160, 80) that the earlier pass missed, so the side view kept bright green terrain while the top-down panel was already dark, and its clouds were near-white at (200, 220, 240) -- more visible than the aircraft. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC39GNmL25Z61NXKGoZmoK
`aero_coefficients` wrote its separation sigmoid as `CL_linear / (1 + exp(u))`. Past about 77 degrees of incidence the exponent overflows float32, `exp` returns inf, and while the forward value is still a perfectly good 0, the reverse-mode derivative of `x / (1 + inf)` is NaN. Any rollout long enough for the aircraft to depart reached that incidence, so every gradient taken through the aircraft came back NaN. That is what had `tune_plane3d_heading_pid` and `tune_plane3d_circle_pid` pinned as strict xfails, with a comment saying the mechanism was known but "the specific operation has not been localised". It is this one. `jax.nn.sigmoid` is the same function evaluated stably, and all seven tuners now pass. Localised by bisecting the horizon at which the gradient first went NaN, then taking the single-step Jacobian at that state: 3421 steps in, at 72 degrees of incidence and 9.7 m/s of airspeed, with 160 of its entries NaN at once. NAN_TUNERS is kept as an empty set rather than deleted, because it is where the next tuner of this kind would be pinned. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Check 7 asserts that no state accelerates under zero input, comparing the mean increment over the last fifth of an unforced run against the first. It ignored fields that do not move, with an absolute floor of 1e-12. That floor is far below float32 resolution, so a stationary field still cleared it on round-off. `plane3d_figure8` flies wings level under zero input and its heading is constant at 1.4296085 rad; the increments are 1e-9 early against 1.5e-8 late, both an order of magnitude below the 1.7e-7 that is one float32 ULP at that magnitude. The check read that as 15x acceleration and failed. The floor now scales with the field's own magnitude, at 32 ULP. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
`patrol` was importing eleven private names out of `plane3d.rendering`, so a palette change in one aircraft environment silently restyled another and nothing said so. The shared half now lives in `render_aircraft`: the palette, the console chrome (header strip, gauge stack, world grid, scale bar), the A320 mesh and its three projections, and the side-elevation and trace panels. All four aircraft renderers draw with it. `plane/rendering.py` is deleted. It had no importer left once `plane3d` and `patrol` moved off it, and a comment kept it "available for anyone rendering long episodes interactively" while it drew in a style matching nothing shipped. The 3D panels move from a floating text HUD to the same labelled bars every other environment reads its values off, over a daylight scene rather than a night one, with altitudes in feet and airspeed in knots. The 2D console was matplotlib with its own palette, its own projection of the same mesh, its own parallax clouds and its own gauges; it is now the shared side elevation plus a trace of altitude against its command over the episode. Clip timing was wrong twice over. `utils.save_video` builds at 60 fps and writes the gif at 30, so moviepy dropped every other frame: a 100-step clip came out 49 frames and played at double the intended speed. `runners.video` now passes FPS=30, and with the halving gone TARGET_FRAMES has to say 100 rather than 200 or every process clip doubles in length and in bytes. Aircraft clips show the opening of an episode at a fixed ten times real time instead of a whole episode time-lapsed into twenty seconds, which read as an aerobatic display rather than an airliner on 8 km lobes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
It had thermal inertia but no dead time, and a first-order plant with no transport delay has no bandwidth limit, so a PID could be tuned arbitrarily tight against it: one held the crown to 0.078 K mean against a 10 K open-loop drift, roughly thirty times better than a real furnace is held, and the task had no headroom left for anything to beat. Its only load variation was an AR(1) drift on pull at a 50 min correlation time, slower than the plant, and a slow smooth load is precisely what integral action cancels perfectly. Added: a 120 s crown thermocouple lag on the observation, with the reward still scoring the true crown temperature, because what the glass sees is what matters and the instrument is the controller's problem; 60 s of fuel transport delay covering the gas train, the air adjustment behind it and flame development; discrete batch charging on a 300 s charger cycle with dose-to-dose mass jitter, averaged over the step so the mass is conserved exactly whatever delta_t is; and the 40 s firing interruption at each reversal that deviation D2 had recorded as missing. Burners re-sized 2.6% to deliver the same time-averaged heat despite losing 40 s in every 1500, which is what real burners are sized to do. Setpoints are a trim walk rather than five independent draws from a 45 K band. Independent draws span 30 K on average and can step 40 C between slots, against the 10-20 C trim the band's own comment describes. Reset now places the crown within 4 K of its first setpoint: a running furnace is found at its setpoint, not 18 K off it. The controller had drifted from the plant in three places. Its objective normalised error by `params.tracking_scale`, 40 K, a field left over from a reward the environment stopped using and dead everywhere else, so with the loop operating at 1 K the tracking term was 6e-4 against an O(1) fuel penalty and the problem was nearly flat in the direction being scored. That is both why this MPC trailed its own PID by 16% on 10 of 10 seeds and why IPOPT needed 349 iterations a step on the worst one. It reads `params.tracking_band` now, the name the four-tank, the column and the pH loop already use. Its regenerator runs at two nodes a chamber against the plant's four, halving the state vector that IPOPT's cost is superlinear in, with `_extract_x0` averaging the plant's nodes down in pairs. And it now models the firing gate and the committed fuel pipeline, both deterministic in time and so legitimately known to a controller. Worst seed: 5.5 s a step at 349 iterations, to 0.12 s at 17. Open-loop crown swing 10 K to 31 K, the PID from 99% of the reward ceiling to 90%, and the MPC ahead of it on every seed measured. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Seven rewards carried a running cost: fuel on the boiler drum and the cement kiln, energy on the building, reboiler duty on the column, reagent on the pH loop, import cost on the battery, and fuel on the glass furnace in the previous commit. All seven weights are now zero. Running cost is real; nobody operates a furnace without caring what the gas costs. The problem is the weight. Against a tracking term already normalised into [0, 1], a cost weight silently picks a point on a Pareto front, and none of the seven had an argument behind its number. The glass furnace showed what that costs: a 0.1 fuel weight made a 3.3 K standing error the optimum of what its MPC was asked to minimise, so the controller sat 6 K cold with fuel at minimum 80% of the time and lost to its own PID. Removing it made that seed the MPC's best and collapsed its solve cost by a factor of 40. The fields and the terms stay wired, so restoring a weight is a one-line change once there is a defensible way to set one. What a proper framing needs is on the roadmap: the two terms in commensurable units rather than one normalised and one priced, so the exchange rate is a physical statement instead of a tuning constant. Actuator-activity terms are deliberately kept. Pitch activity on the turbine and rod motion on the reactor are regularisers that keep the control problem well-posed rather than pricing it, and without them the optimal policy is a bang-bang chatterer, which is its own implausibility. The two tests that asserted a cost reduces reward now assert it both ways, so the wiring stays covered whichever the weight is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
…pisodes nobody needs An audit of every tracked output against what its plant can be asked found two aircraft tasks specifying the impossible. `plane_sine` commanded a peak climb rate of 20.9 m/s against an aircraft that sustains 14.6, so the target was unreachable by construction and the task measured actuator saturation rather than closed-loop bandwidth, which is what the sinusoid is documented to be for. Amplitude 800 m to 300, peaking at 7.9 m/s. The 3D tasks drew their initial altitude independently of the commanded one over the same 5 km band, so the median episode began 1590 m from its assigned level and opened with minutes of open-loop climb before the pattern could be flown at all. An aircraft handed a heading, a circle or a hold is at the level it was given. Both the 2D and 3D resets now draw the start around the target; median gap 337 m. The staircase was a square wave between two altitudes 2.4 km apart. Sizing it by tracking quality optimised the wrong thing: the aircraft can fly it, but it is not what an altitude-hold task looks like. It is a ladder of eight levels now, adjacent changes 0.2 to 0.8 of the amplitude, neither level nor direction repeating on a two-tread cycle. `plane_steps` is absorbed into `plane_energy`. They were the same environment: same plant, same schedule, same disturbances, same episode, differing only in `speed_weight` being 0.0 rather than 0.5. Two registered environments for one reward coefficient is not two tasks, and the pair cost 9 h of the 12.9 h the aircraft took to record. `PlaneParams(speed_weight=0.0)` still gives the pure altitude staircase. Gains stay tuned on the plain hold task and are deliberately not re-tuned for the airspeed trade, since a PID fitted to that trade is no longer the honest reference for whether the trade is worth making. Five episodes were far above what the protocol's own criterion asks. It says N >= max(10 tau, one period) and the suite clusters at 10 to 15 tau; `plane_energy` ran 104, and the path tasks up to 3.6 laps where one shows whether the path can be flown. Cut to 52 tau and about one lap. Recording the aircraft goes from 12.9 h to roughly 4.5. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Carries the pieces the recorded baselines depend on: `provenance` fingerprints, `rl_results` stamping so a learned result cannot outlive the environment it describes, the recording script, and the aircraft gains from the coordinate search in `scripts/tune_pid.py`. The racetrack hold gets its bank-rate term. Its cascaded PID had no `Kd_bank`, so the tuner scored every candidate at -inf because `Plane3DRacetrack` was missing `obs_target_index` and `rollout` raised before any of them ran. With both fixed, cross-track error went from 3.02 km to 0.31 km. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Adds the GitHub Pages workflow, a code of conduct, and the contribution terms the README links to, and moves the trove classifier from Alpha to 4 - Beta. The quickstart notebook is under test rather than decorative: `test_docs.py` executes it, so an example that stops working is a failure rather than something a reader discovers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
The study separates two things a single MPC number conflates: what accounting for uncertainty is worth, and what is left for learning. Deterministic against scenario against oracle, on the four environments whose disturbances are large enough for the distinction to show. The roadmap gains what this session's audits turned up and did not fix: deriving the MPC error bands rather than choosing them one at a time, a defensible framing for running cost before restoring any weight, measurement noise anywhere at all, cullet ratio on the furnace as the first gain disturbance in the suite, whether the aircraft warrants transport delay, and a conformance check to stop the do-mpc controller models drifting from the plants they mirror. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
The counts were stale everywhere: twenty-two environments became twenty-one and twenty MPCs nineteen once `plane_steps` was absorbed. `rl-protocol.md` was the worst of it, since gamma is derived from the episode length by rule, so every discount for a changed environment was wrong and three environments were missing from the table entirely. It now also records the shortening pass rather than only the lengthening one that preceded it. Three PHYSICS files documented a cost term their reward no longer applies, two of them still describing the pre-log clipped-square shape. The furnace's state and parameter tables were missing everything added with the dead time. `generate_env_pages.py` preferred a hand-placed `pid_output_short.gif` and fell back to the generated clip. Its own comment recorded that going wrong once, "the gallery showed pre-re-skin aircraft for a week". It went wrong the same way again: thirteen pages were still on shorts that no generator refreshes, so the gallery was split between two visual styles and the furnace's played 1.8 s against 9.8. A fallback chain whose first entry nothing maintains is a trap, so there is no chain. A page with no picture now means no clip was rendered, which is worth seeing. The furnace's `mpc_degraded` note asserted the MPC is 16% behind its PID on 10 of 10 seeds, which is no longer true. Rewritten rather than deleted: it says what was wrong, that all three causes are fixed, and that the flag stays until that is confirmed at ten seeds, since it drives an xfail and a page warning. Removed: 25 MB of stale short clips, 4 MB of mosaic gifs nothing references against the webp the docs actually use, an orphaned `videos/plane_steps`, and the OS junk. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
A pass over every environment's code and physics before committing hours to a re-record. It found one defect that would have gone into the baselines, five that change behaviour, and a systematic documentation problem. The defect: `battery.cost_weight` had been zeroed with the running costs. It is not a consumption cost. It gates degradation and state-of-charge comfort, which are what keep that control problem well posed, and without them the optimal policy follows dispatch until the pack hits a limit and the episode ends. Restored, and the field now says what it is. Behaviour changed. Patrol moves to the suite's log-scaled reward, in both the single-agent and MARL variants: it was the last Gaussian, and with a 60 m sigma against a 1500 m terminal bound it was flat at 1.4e-6 from roughly 250 m out, so most of the reachable range carried no gradient. That flatness is a candidate explanation for the chaotic tuning objective its own D1 documents, and migrating separates reward shape from guidance law as the cause. The reactor loses a setpoint schedule it stopped using long ago, starts its xenon at the equilibrium for an independently sampled recent power level so the poison is a live bias rather than exactly zero, and runs 8640 physics steps rather than 1200 -- which is what `effectiveness_overrides` already used, on the grounds that anything shorter cannot distinguish controllers on the environment's headline physics. That override is gone, so the conformance check and the baseline finally measure the same task. Its rod penalty moves to the wind turbine's form, charging for demanding motion the rods cannot deliver rather than for holding them where the physics requires. The cement kiln loses a trip that fired on a refractory temperature its own contract calls unreadable. The boiler drum starts with feedwater matched to the steam actually leaving, which is what its comment always claimed. The systematic problem is that documentation had drifted from code, and it is the repository's most common defect. Three of this review's own findings were wrong because of it, including one where a duplicate implementation was written and reverted after a deviation claimed post-stall lift decays to zero when the fix had been in place for months. Eleven parameters described a reward that does not read them; four made claims about it that are checkably false. Five `precision_floor` declarations had swallowed the comment belonging to the bound below them, so five bounds silently lost their explanation. Two registry comments quoted episode lengths from before an audit changed them. Two hedges, both wired into `make ci-docs`, CI and the test suite. `scripts/check_doc_drift.py` catches four mechanical cases: a contract naming a symbol its package no longer defines, a parameter whose comment claims the reward uses it, an episode comment disagreeing with its value, and the doubled `#` that is the signature of the formatter merge. On its first run, after the manual sweep, it found three more. `scripts/generate_physics_facts.py` writes the computable facts into all fifteen contracts between markers, so they cannot be wrong. What is left hand-written is what cannot be derived, which is the point of the documents. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
It said every environment "ships with MPC, PID and RL baseline controllers". There are no learned baselines, and publishing them is an open roadmap item. The front page now says what is actually recorded, PID everywhere and MPC on nineteen, and points at the protocol for what the learned arm will be measured against. `docs/rl-protocol.md` also drops the random-policy and best-constant-action reference rows it specified. The constant-action bar is not lost: the conformance suite already asserts every PID beats it, which is the claim that row existed to support. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Benchmarking all twenty-one environments to check they are still fit for RL: at the protocol's largest budget of 1e7 environment steps, the slowest of them costs about 16 seconds of environment time and the fastest under a millisecond. Network passes dominate any learning loop by two to three orders of magnitude, so the environment is never the bottleneck and none of the twenty-one is unfit on speed. Nothing has regressed. Measured at the batch the contracts quote, every environment is at or above its documented figure except the building, and with the rest uniformly faster that reads as a different machine rather than a code change. What the measurement did expose is that these numbers do not belong in the contracts. They move with the machine and with the batch size, while every other number in a PHYSICS.md is a fact about the process being modelled. Two contracts both claimed to be "the slowest environment in the suite", and at batch 4096 neither is -- patrol is, and at batch 256 it is the column, so the superlative is not even well defined. The README's range of "0.5-17 M steps/s" understated the top end by a factor of forty. Throughput now lives in docs/performance.md with the machine, the batch and the command to reproduce it, and the contracts point at it. Deliberately not checked by CI: a check that fails on somebody else's laptop is worse than no check. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
The recorder merges its rows onto whatever the file already holds, which is what makes a partial re-record safe and is exactly why nothing ever gets cleaned up. `plane_steps` sat in `baseline_returns.json` after being absorbed into `plane_energy`, publishing returns for a task that can no longer be constructed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
IPOPT was left at its default of 3000 iterations with no time limit. In a receding-horizon loop that is a hang, not a safety net: nine glass-furnace seeds finished in about three and a half minutes each while the tenth was still going after seventy, and it blocked a whole re-record. CasadiMPC now caps iterations at 150 and CPU time as a backstop. The iteration cap is the one meant to bind because it is deterministic, so a baseline recorded on one machine still reproduces on another. Capping alone would not have been enough, because do-mpc neither raises nor warns when IPOPT gives up. It stores the failed iterate, hands it back as the action, and warm-starts the next step from it, so nothing downstream could tell a failure from a converged solve and an MPC baseline could quietly stop being the upper bound it is presented as. Every record now carries solver_calls, solver_failures, solver_capped and solver_mean_iters, and a solve that failed for any reason other than the cap makes the controller hold its previous action and restore its previous warm start. The cap has never actually fired. What fixed the furnace was variable scaling, which no CasADi MPC declared. IPOPT auto-scales the objective and the constraints but not the decision variables, and the reactor was handing it a vector spanning rho_ext around 0.0016 up to a precursor concentration around 377, a factor of 605 000, with hard bounds on the smallest entry of it. Each subclass now declares a SCALING table of typical magnitudes measured over a PID episode. The furnace averages 16.8 iterations over 16 000 solves and the reactor 5.2 over 86 400; a pathological furnace seed used to take 349. The glass furnace MPC consequently stops losing: 1513.4 against its PID's 1443.5, winning 10 of 10 seeds where it was 16.0% behind and lost 10 of 10. Its mpc_degraded flag is removed. The battery takes its place, and for a reason worth stating precisely: it loses on 9 of 10 seeds and its mean leads only because seed 0 scores 350.4 against 154.7 for the other nine. Two explanations are ruled out in the registry entry. The four-tank MPC bounded h_min hard and had no h_max bound at all, so it was blind to half of a termination condition it is scored on. Both bounds are now present and soft: the optimiser owns its inputs and can always satisfy their bounds, but a state bound the plant can walk the initial state onto makes the NLP infeasible at x0, and IPOPT answers that with a restoration phase. Recording is now parallel across environments, not just across seeds, and checkpoints after each one. It used to write once after the loop, so an interrupted run threw away everything finished; that cost two full re-records. A monitor prints every thirty seconds and names any seed running past four times the median for its own environment, which is the line that would have caught the furnace at minute fifteen instead of minute seventy. A signal handler reaps the pool, because ProcessPoolExecutor orphans its children when the parent is signalled and fifty of them survived two kills and ate the machine for two and a half hours. The baseline table's ceiling was wrong for the reactor. It divided by control_period on the belief that max_steps_in_episode counted physics steps there; it does not, and a reactor rollout pays 8640 rewards for 8640 env steps. That put its share at 1.251 and tripped the guard against publishing. Full suite green: 1573 passed, 3 xfailed. Recording the whole suite is now 30 minutes against 53. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
…oise Two environments were not measuring control skill, and both were found the same way: by asking what a number could possibly be a function of. The battery tracked a dispatch signal whose one-step innovation had a standard deviation of 63.6 kW against a 150 kW tracking band, which puts the best attainable tracking reward at 0.429. The shipped PID scored 0.447 and the MPC 0.430: both sat on an irreducible noise floor and nothing that could be plugged into the environment would have scored meaningfully better. The dispatch is now a schedule of twelve 300 s market blocks in +/-0.8 MW with 2 kW of regulation jitter, which is also what a grid battery is actually handed. PID 111 -> 262, MPC 174 -> 266. The patrol follower chased a lead that held one constant turn rate for a whole episode. Its settled error was exactly linear in that rate and exactly symmetric in its sign -- 25.9 m per 0.001 rad/step -- which is proportional control against a rotating reference, not a mistuning. No gain could close it because the term was missing: close to the slot the law blends onto "fly parallel", which sets heading and never commands the rate a turn needs. Feeding the lead's turn rate forward, recovered by differencing psi + rel_heading, takes the hardest case from 77.8 m to 2.4 m against a 60 m tolerance and a 3 m reward precision floor. The strict xfail is removed and now runs at both +/-0.003. The lead flies a routed circuit of eight legs, so there is a turn onset to anticipate roughly every minute instead of once per episode. Patrol gets its first MPC. The obstacle recorded against one -- a manoeuvring lead needing its trajectory as a time-varying parameter -- is real for CasADi and irrelevant for a gradient planner, since the lead is scripted and step_env propagates it for free. It took 300 iterations rather than the suite's usual 50: at 100 it merely matches the PID, at 300 it reaches 0.84 of ceiling and leads by 51% over ten seeds. The planner was never stuck, it was stopping early on 90 decision variables, and every objective variant compared before that was being compared at a non-converged optimum. Planners no longer plan against an invented disturbance. GradientMPC and SamplingMPC roll the true environment forward under a hardcoded PRNGKey(0) while rollout drives the plant with PRNGKey(seed), so on seed 0 the planner's simulated disturbance *was* the plant's -- perfect foresight, worth 350.4 against an honest 151.8 on the battery. experts.mpc.plan_params now zeroes the parameters named in EnvSpec.noise_fields for the planner's copy, which is certainty equivalence and which docs/baselines.md had already flagged as absent. It also controls better: the wind turbine went 343.9 -> 348.3. Seed parallelism is now chosen by backend. vmap is ~4x more efficient per iteration but uses ~1.4 cores on CPU, so ten processes win the wall-clock there; on a GPU the reasoning inverts and batching is near-free. The worker pool is sized by *available memory* rather than cores: one planner per core put 14 processes and ~10 GB on this machine, exhausted RAM and sent it swapping, and the load average of 158 was threads blocked on memory while every worker sat at a well-behaved 95% CPU. Also: PlanePatrol.get_obs now takes the optional params every other environment takes, which the shared MPC contract needs and which only surfaced once patrol had an MPC to exercise it. Six stale roadmap entries corrected, two of them invalidated by this commit and two being the same duplicated item. README and docs/baselines.md now say plainly that a PID losing here is a claim about these tasks and not about PID control, with both measurements above as the evidence. data/baseline_returns.json is stale by design: experts/mpc.py changed, and it is a _SHARED_SOURCE, so all 20 fingerprints moved. The 20 failing test_recorded_baseline_still_describes_this_tree cases are that machinery working. One re-record clears them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Both are things a first reader hits before anything else, and neither is tracked anywhere. All 40 gifs date from 2026-09-09 and 25 environment modules have changed since. Two of those changed the *task*: the battery clips show an OU random-walk dispatch that no longer exists, and the patrol clips show a lead holding one constant turn rate rather than flying a routed circuit. Showing behaviour the library no longer has is worse than showing it late, so this is a regeneration after the final record, not a nice-to-have. The README is 410 lines and 3199 words over fourteen top-level sections, so a reader deciding whether the library is for them scrolls past physics validation, performance benchmarks and related projects before reaching anything they can act on. Most of that already lives in docs/, so the fix is deletion and linking rather than rewriting, and it gets cheaper once the docs site is live. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
A clean 39-minute run on an idle machine. Every environment has the MPC leading, no environment terminates early on any seed, and all seven CasADi plants report 100% solver convergence -- 119,600 solves with zero failures and zero hitting the iteration cap added this morning. The two environments this work was about both landed where the earlier partial runs said they would. The glass furnace records 1443.5 against 1513.4, winning 10 of 10, where it began the day 16.0% behind its own PID and losing 10 of 10 with one seed that ran seventy minutes against a three-and-a-half-minute median. Patrol records 111.2 against 168.1, a 51% lead at 0.84 of ceiling, on an environment that had no MPC at all. ``EnvSpec.mpc_degraded`` is now empty for the first time. Both entries were retired by fixing the cause rather than the wording: the furnace by variable scaling, and the battery -- which lost on 9 of 10 seeds with its mean carried entirely by seed 0, where the planner shared the plant's PRNG key and knew the future noise exactly -- by closing that leak and reshaping its dispatch signal. It now wins 8 of 10 on merit. docs/baselines.md says what the field is for and what was retired from it, since an empty field with no explanation invites the assumption that nothing was ever wrong. The recorder also survives a pool that dies under it. BrokenProcessPool is not recoverable in place -- the executor refuses further work once a child dies abruptly -- and it took down two runs in one evening, once under genuine memory exhaustion and once at eight workers with 17 GB free, which is not understood. Outstanding jobs are now tracked explicitly and a fresh pool is built for whatever is left, halving the worker count each time. Seed parallelism stays batched on both backends. The per-seed route has never been measured to completion here -- two attempts, killed at 26 and 37 minutes -- and the "26 minutes" that briefly justified switching to it was a progress line misread as a result. The docstring now says that rather than the swap-pressure story it replaced, and names the experiment that would settle it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
…edia
The 40 clips shipped in videos/ were rendered on 2026-09-09, and 25 environment
modules have changed since. Two of those changed the *task*: the battery clip
showed an Ornstein-Uhlenbeck dispatch that no longer exists, and the patrol clip
a lead holding one constant turn rate rather than flying a routed circuit. All
21 are re-rendered from the shipped PID, the console clips re-quantised through
scripts/make_gallery_clips.py, and the five gallery mosaics rebuilt.
Nineteen mpc_output*.gif files are deleted. Nothing referenced them -- not the
env-page generator, not the mosaics, not the Makefile, not a single document --
and they were stale on both counts, predating today's MPC changes as well as the
task reshapes. They were 114 MB of the 246 MB videos/ directory, which is now
125 MB.
The README goes from 410 lines and 3199 words to 353 and 2053, a 36% cut, by
moving reference material to the documentation rather than deleting it:
- The per-environment validation findings -- what each model is checked
against and something that check caught -- existed *only* in the README.
They now live in docs/PHYSICS_METHODOLOGY.md, beside the method they
illustrate.
- The complexity ladder becomes docs/complexity.md and joins the nav, with
the README keeping a three-line summary and a link.
- The performance table is replaced by a headline range and a link. It was a
second measurement of a quantity docs/performance.md already measures, at a
different batch size and depth, which is how two tables drift apart.
- The stable-baselines3 example is folded into a disclosure, and the
deliberate-modelling bullets are compressed to a paragraph, both already
covered in docs/index.md.
The quickstart, the environment families, the failure-mode table and the
baselines claim are untouched: they are what a reader needs to decide whether
this library is for them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
docs/videos is a symlink to videos/, so mkdocs copies the clips in from the working tree. A local build therefore passes on a tree that happens to have them, and a build from a clean checkout would have published twenty-one environment pages with broken images. The docs-deploy workflow now renders the clips before building, headless via SDL_VIDEODRIVER=dummy and MPLBACKEND=Agg, with the job timeout raised from 10 to 45 minutes because rendering is about eight minutes on a developer machine and a runner has fewer, slower cores. The tracked set is now exactly the five gallery mosaics, which the README and the environment index embed. .gitignore had excluded videos/**/*.gif with an exception for *_short.gif, left from when the pages embedded the shorts. They do not: scripts/generate_env_pages.py deliberately references pid_output.gif, and the comment there records the fallback chain going stale twice. So the repository was carrying 55 MB of shorts nothing published referenced while missing every file the pages pointed at. docs/baselines.md now states the convention and why it changed, rather than claiming the shorts are what is committed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
…started There was no guide: the API appeared as a six-line example in getting-started.md, the two toolkits were documented only in their own module docstrings, and nothing told a reader how to regenerate the shipped media or what a headless machine needs. docs/rendering.md now covers rendering an episode (including why FPS has to be set deliberately -- passing the default 60 against a clip built at 30 makes moviepy drop every other frame, so a 100-step episode comes out 49 frames at twice the intended speed), what the dashboard regions mean, the hidden-state marking that tells a viewer which quantities the agent is flying blind on, the split between the matplotlib render_kit and the pygame render_aircraft and why it exists, the regeneration commands and which of them actually shrinks a clip, and how to add a renderer. The README's Rendering subsection and getting-started's both keep a short description and link out, so the detail has one home. The first example is executed by tests/test_docs.py, which caught a placeholder in the second: the render_kit snippet is a call shape rather than a runnable program and is marked `doc: skip` with that reason. Verified while writing this that `mkdocs build --strict` exits 0 with the clips absent -- mkdocs does not validate image sources -- so the per-PR docs job stays fast and green, and only docs-deploy, which runs on merge to main, pays the render. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
…g numbers
This is the version that gets advertised, so the page has to make its case in
the first ten seconds and every number on it has to be true.
The opening led with inventory -- "21 environments: 9 aircraft, 5 process
control..." -- before saying what the library is or why the setting is
different. It now leads with the distinction: most benchmarks ask you to reach
a goal once and stop, industrial control asks you to reach a setpoint and hold
it indefinitely, and that fails in its own ways. Then the plants, then three
scannable claims in a table: you can tell immediately whether your agent is any
good, the physics is a tested contract rather than an assertion, and the
environment is not the bottleneck.
Four factual errors, one of them introduced by the previous trim:
- The Performance section claimed "0.5 M to 1.7 G steps/s". The measured
range is 0.6 M (patrol) to 700 M (first_order). The header had it right and
the section did not.
- The environment table said 9 aircraft in the header and 10 in the family
table, which summed to 22 environments rather than 21.
- It also described "four target patterns" for the 2D aircraft, where there
are three: hold, sine, and altitude-and-airspeed.
- The documentation table offered "All 22" environments.
The Performance section is gone, since the intro table now carries the number
and docs/performance.md carries the detail; its link moves into the
documentation table alongside the new rendering and complexity pages. The
physics section linked the methodology twice and now links it once. The
baselines section keeps the PID-framing paragraph, as a pull quote, at about
half the length.
354 lines and 2047 words to 338 and 2071: barely shorter, considerably clearer,
which is the trade worth making on a page whose job is to be read.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
A pass over the wording rather than the structure, on the page that gets
advertised.
The gallery caption read "held on setpoint by its shipped PID baseline", which
implies a PID is all that ships. It now says the clips are one example task per
family under PID control, which is what they are.
The failure-mode table led with control-theory labels. "Non-minimum phase" says
nothing to most readers, so the rows now say what actually happens: open the
steam valve and the drum level rises as water leaves; you cannot see most of it;
it answers late; tracking now costs you later. The phenomenon is the
interesting part and the jargon was hiding it.
Removed throughout: every "X, not Y" and "rather than a" construction, and all
five em-dashes. The Colab line was written like a landing page ("Prefer a
browser? ... on a free CPU runtime in about a minute") for an audience that
knows what Colab is, and is now one clause. "NREL 5 MW wind turbine" is just a
wind turbine; the model is cited in its PHYSICS.md where that matters.
The Rendering section explained the dashboard layout at length directly under a
gallery that shows it. Cut to three lines and a link.
tests/test_docs.py caught the one substantive risk in the pass: rewording
"covered by fifteen contracts" broke the check that keeps that number true.
The wording the check depends on is restored rather than the check loosened.
331 lines to 329, 2030 words to 2012.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
"20 of them an MPC" invited the obvious question and gave the reader no way to answer it without grepping the registry. The README now names patrol_bearing_only and says why in the same breath: it hides the slot error a planner would read, which is the point of the variant. The registry note gave the wrong reason. It ended "No MPC: the follower's plant is the full 3D aircraft and the reference is a manoeuvring lead", which was the stated obstacle for both patrol variants until yesterday, when patrol got a GradientMPC that differentiates step_env and propagates the scripted lead for free. The manoeuvring lead blocks a CasADi model and nothing else. What actually blocks a planner here is that it reads the slot error out of the state, and this variant withholds exactly that; handing it the true state anyway would make it an oracle on a task defined by what is hidden. The note now says so, and says what it used to claim, because a reason that was disproved is worth recording rather than quietly replacing. docs/baselines.md already had this right. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
The previous version addressed the reader directly ("you have something real to
beat", "you can tell straight away"), explained things an RL researcher already
knows, and read as one continuous argument. A README is skimmed, so sections now
stand alone and state what the library does.
Concrete changes:
- Second-person address removed entirely: 0 occurrences, from 14.
- The three-claim table became four feature bullets, which is the convention
and scans faster.
- Vague subjects named. "These are meant to be plants you can believe"
became "Each environment carries a PHYSICS.md stating what it models".
- The baselines paragraph that buried three controller-structure examples in
one sentence is now a list of three, one line each.
- Installation lost the poetry and uv variants; anyone using them can
translate `pip install`.
- Related projects and Contributing lost their explanatory asides.
- Citation lost its preamble.
- "Why these environments" became "Why setpoint tracking", which says what
the section answers, and its row labels use the standard terms again
(inverse response, transport delay) since the audience is control-literate.
329 lines and 2012 words to 312 and 1597, a 21% cut with nothing factual
dropped. The tagline no longer says every environment has an MPC, since
patrol_bearing_only does not.
Every internal link still resolves and the guarded "covered by fifteen
contracts" claim is intact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
The one-line description of what the library is sat at body-text size under a slogan. It is now an <h3> directly under the title, with "Reach the target. Then hold it." demoted to an italic line beneath it. The failure-mode table named control-theory properties and left the reader to work out why they matter for RL. "Non-minimum phase" became "inverse response" in an earlier pass, which was no more informative. It is now "wrong-way-first response", and the table gained a third column saying what each property costs a learner: exploration that ends an episode permanently, a policy that has to infer what it cannot measure, credit assignment spanning hundreds of steps, one control interval that cannot serve both ends of a timescale split. The middle column also explains the mechanism where it was merely asserted: drum level rises before it falls because the steam bubbles expand. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
The quickstart opened with `REGISTRY["plane"]` and `spec.make_env()`, which is
not how any other environment library is used and is not required here. It now
uses what already existed:
from target_gym import Plane, PlaneParams
env = Plane()
params = PlaneParams() # or env.default_params
pid = env.make_pid()
obs, state = env.reset(key, params)
`reset` and `step` already follow the gymnax API with optional params, and every
environment already exposes `make_pid`, `make_mpc` and `save_video`. The registry
is now introduced only where it is actually needed, in a separate subsection on
reproducing the published numbers.
That subsection exists because of a trap worth stating rather than hiding: the
recorded returns are measured at each environment's benchmark settings, and those
differ from the defaults. `plane` runs 10 000 steps by default and is scored over
280, so a default-params rollout compared against a recorded MPC number is
meaningless. The README now says so and shows `spec.make_test_params()`.
The flagship mosaic captioned the four-tank "Process - non-minimum phase", which
conveys nothing. It now reads "the obvious valve pairing is unstable", which is
concrete and is also why the shipped four-tank PID crosses its loops. The other
three captions lost their shorthand too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Custom parameters belong in the documentation, not in the first example. The quickstart now constructs nothing but the environment and relies on reset and step falling back to default_params, with a sentence noting that each environment exports its parameter class for custom configurations. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
… stale
docs/rl-protocol.md states N >= max(10 * tau_actuator, 3 * T_period). Six
episodes were lengthened to satisfy it during the episode-length audit and
nothing has enforced it since, so the rule held only for as long as nobody
changed a plant.
Somebody did. Yesterday's furnace work added a regenerator, a reversal cycle, a
120 s thermocouple lag and a two-step fuel dead time, which took its crown
response from 132 steps to 822. Its 1600-step episode went from a documented
12.1 tau to 1.9, and no number in the repository moved. The protocol's table
still said 132, still said the reactor runs 1200 steps when it runs 8640, and
still gave every aircraft a 14 or 23-step response where they measure 36 to 365.
Three pieces:
- target_gym/episode_length.py holds the measurement, so the script and the
test cannot disagree about what a time constant is. Integrating plants have
no tau_63 and are reported as such rather than fitted.
- tests/test_episode_length.py asserts the actuator clause for every
environment, marked slow because it drives two open-loop episodes each.
Nine pass, nine skip as integrating, and three are strict xfails carrying
their measurement: glass_furnace at 1.9 tau, plane_energy at 3.3,
plane_sine at 2.2. Strict, so lengthening one breaks the test and forces
the entry out rather than letting it linger.
- scripts/measure_time_constants.py now writes the table into
docs/rl-protocol.md between markers, flagging any episode below the rule.
The period clause is not checked: it needs each task's reference period, which
the registry does not expose. plane_sine fails it too (240 s period, 480-step
episode against the three periods required), which is recorded in its xfail.
Fixing the three costs a re-record of each, and lengthening the furnace to 8220
steps would make it five times more expensive to record, so they are reported
rather than fixed in passing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
A clean `uv sync --frozen --group dev` resolved CPython 3.14.6 and then failed with "Distribution jaxlib==0.6.2 ... doesn't have a source distribution or wheel for the current platform", which is an obscure way of saying the interpreter is too new. jaxlib publishes cp311 through cp313. requires-python said <3.15 and the test matrix listed 3.14, so that job could never have passed. Both now stop at 3.13, with a note to raise them when jaxlib ships cp314. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Withdraws the test and the generated table added in 22c303e, and fixes the documentation instead. The test was wrong in a way worth recording. It measured tau_actuator inside the episode it was judging, which makes the rule circular: a longer episode sees more of the response, reports a larger tau, and demands a longer episode. The same aircraft read "integrating" at N=280, 217 steps at N=480 and 365 from N=800 on, and plane, plane_sine and plane_energy returned identical numbers at equal N because they are one plant. Three apparent violations were three episode lengths. Decoupling the probe did not rescue it. With a converged window `plane` failed too, which is implausible for a task documented at 12.2 tau, and adding the monotonicity precondition a first-order fit actually requires left almost nothing measurable, `first_order` included. Three attempts gave three answers about the same plants. The finding is that tau_actuator does not exist for most of this suite. Nine environments integrate, so there is no steady state to settle to. The aircraft oscillate, because a held elevator excites the phugoid. The glass furnace does not settle inside eight thousand steps, and its response reads 822, 1520, 2383 and 3415 steps at successive window doublings since its physics gained a regenerator, a reversal cycle, a thermocouple lag and a fuel dead time. So docs/rl-protocol.md now says the actuator clause binds only where a settling time exists, names what sets the episode otherwise (laps, disturbance timescale, setpoint schedule), and records why enforcing it as a test was withdrawn. The period clause is also corrected from three periods to one, in all five places it was stated. The relaxation was decided in the second audit pass and its reasoning recorded there, but the rule itself was never updated, so the document argued with itself. plane_sine's two periods keep their frequency-probe justification. No baseline is re-recorded. The evidence for lengthening any episode was an artifact of the measurement. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
…ded at
Eleven of twenty-one environments now have default_params identical to what
their baselines were measured with, so a plain rollout is directly comparable to
the published number. test_params is gone for those, since a field that exists
only to restate the default is one that drifts.
Every recorded value is unchanged. The nine re-recorded environments returned
the same numbers to the decimal (battery 261.97/265.76, glass_furnace
1443.47/1513.44, and so on); only the fingerprints moved, because the env
sources did. The episode lengths that became defaults are the reasoned ones:
laps for the path-following aircraft, the disturbance timescale for the turbine
and battery, the setpoint schedule for the furnace, building and kiln. The old
defaults were not. All nine aircraft said exactly 10 000, which is one number
standing in for nine judgements, and hvac and cement_kiln defaulted *shorter*
than the lengths the episode-length audit justified.
Ten environments still override, for two reasons that are not going away:
- Seven share a params class with their siblings. PlaneParams serves plane,
plane_sine and plane_energy at 280, 480 and 1200 steps; PlaneParams3D serves
four tasks at 200 to 650; PatrolParams inherits PlaneParams3D. One class
cannot default to several variants, so the episode length belongs to the
variant and test_params is the right place for it.
- The reactor counts max_steps_in_episode in *physics* steps, 86 400 for 24 h,
while its benchmark 8640 is the same 24 h in env steps, since step_env runs
ten physics sub-steps. Aligning those numerals would make the episode ten
times too long. Left alone deliberately.
Four duration comments were wrong once the values moved and are corrected:
boiler_drum is 800 s rather than an hour, glass_furnace 13.3 h rather than 48,
cement_kiln 5.8 h rather than 4, hvac 7.5 days rather than 7.
The README's caveat about defaults differing from the scored settings is
narrowed to the variant-sharing plants, which is all it still describes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Roadmap. Three 0.6 items are complete and were still open: the media regeneration, the README trim, and the baseline re-record, which finished in one 39-minute run with every MPC leading, no early terminations and all seven CasADi plants at 100% solver convergence. A fourth is added and marked done rather than silently closed, because it was broken rather than missing: publishable media never reached the site, since docs/videos is a symlink into a gitignored directory, so a local strict build passed on a tree that happened to have the clips while CI would have published broken images. "Release before advertising" said 22 environments. There are 21. New item: the glass furnace's episode length is justified by a crown response of 132 steps that the physics no longer has. Since it gained a regenerator, a reversal cycle, a thermocouple lag and a fuel dead time, its open-loop response does not settle at all on this timescale, reading 822, 1520, 2383 and 3415 steps at successive probe windows. The recorded return is unaffected, the *reason* for the episode length is not, and something measurable has to replace it. CHANGELOG gains the second half of the work: patrol's first MPC and why the recorded obstacle against one was wrong, the patrol PID's missing feedforward term, the routed lead, the aligned default parameters, the requires-python cap, and the withdrawn episode-length rule. Also cleaned 1.9 GB of disposable clones out of /tmp, keeping the pre-rewrite backup bundle. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
The comment said the bound tracked jaxlib and to raise it when jaxlib ships a cp314 wheel. jaxlib has shipped cp314 since 0.10.1 and is now at 0.11.1, so that trigger has already fired and following it would not work. The real constraint is gymnax. Its latest release, 1.0.0, declares requires-python <3.14 and caps jax<0.7, which is what holds jaxlib at 0.6.x, the version without a cp314 wheel. Every environment here is a gymnax environment, so upgrading jax is not a route around it. Same correction in the CI matrix comment. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
Two CI failures, both mine. scripts/generate_env_pages.py tested whether a clip file existed before referencing it. The clips are a build artifact: .gitignore excludes videos/**/*.gif and docs-deploy renders them before mkdocs runs. So pages committed from a working tree that had them disagreed with pages generated in CI from a clean checkout, and all twenty-one came back stale. The generator now emits the path unconditionally, which is correct for a file the site build produces, and --check is stable with the clips present or absent. That is the last consequence of moving media out of the repository. mypy does not model attributes on function objects, so `policy.controller = mpc` in runners.mpc_policy failed [attr-defined]. Silenced at the assignment with a note, since the alternative is a wrapper class for one attribute. `make ci` is green locally: 1468 passed, 85 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012WdRJEyGAZwKDD1EFwwH3U
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Environments
plane_sine(a frequency response)and
plane_energy(altitude and airspeed, removing the spare actuator). Atuned PID finishes an altitude hold with 0.0 m of error and can no longer
separate two controllers.
lag, fuel transport delay and pulsed charging. Also 43x faster.
signal capped the best attainable reward at 0.429 while the PID scored 0.447;
the patrol lead held one turn rate all episode, which a single feedforward
term cancels. Both reshaped.
Baselines
All 20 environments with an MPC recorded in one 39-minute run. Every MPC leads
its PID, zero early terminations, and 100% solver convergence across 119 600
CasADi solves.
EnvSpec.mpc_degradedis empty for the first time. The glass furnace went from16% behind its own PID, losing 10 of 10 seeds, to leading all ten; the fix was
variable scaling, which no CasADi MPC declared (the reactor was handing IPOPT a
vector spanning a factor of 605 000). Patrol gets its first MPC, leading by 51%.
Correctness
gradient or sampling planner: they simulate under a hardcoded
PRNGKey(0)while
rolloutdrives the plant withPRNGKey(seed). Worth 350.4 against anhonest 151.8 on the battery. Planners now predict the mean disturbance.
raises nor warns when a solve fails. Now capped, with convergence published.
h_minand noth_max, while the plantterminates on either.
requires-pythonallowed 3.14, which jaxlib has no wheel for.Docs and repository
from the registry to the direct
Plane()API.docs/videossymlinks into a gitignored directory, so local strict builds passed on a tree
that happened to have the clips. CI now renders them.
the working tree byte-identical.
environments that do not share a params class.
Known gaps
longer has; it does not settle inside 8000 steps. Recorded returns unaffected.
does not exist for two thirds of the suite. Now scoped; an attempt to enforce
it as a test was withdrawn and the reasoning recorded.
docs/rl-protocol.md.docs-deployrender has never run on an Ubuntu runner. It fires on merge.