Repository navigation
Count MLX's retained cache in the harness watchdogs - #42
Merged
Merged
Conversation
The showcase, benchmark, latent-capture and A/B workers compared active memory alone against the ceiling; freed buffers sit in MLX's retained cache, resident but not active, and the default cache limit is near device memory. Both watchdogs now sample active plus cache, bound the pool at install time (2 GiB generation workers, 4 GiB decode reps so the timed decode keeps a warm pool, A/B per-condition bounds passed through), poll every 50 ms, record the policy in the abort artifact and the worker result, and abort with reason sample_error if a sample raises instead of letting the thread die silently.
The cache bound does not close arithmetic the memory limit already closes (mlx 0.32.2 drains the retained cache before an allocation that would cross it), so the comments, CHANGELOG and COMPARISON now say what the change is: an honest active-plus-cache reading with a stated policy. Each watchdog keeps high-water marks of the sum it compares and of the cache term, recorded as watchdog_observed in the live result, the bench sentinel and its report block, and the A/B unit result. The orchestrators print every term of an abort plus the error text of a failed sample. The live bound rises to 4 GiB after an interleaved timing check of live_preview at 4 and 20 GiB (10.76 s vs 11.07 s median, three reps each); the capture bound follows it. The A/B condition-to-cap mapping is tested for real again.
…ad tests The operator line for a watchdog abort now lists every term the artifact carries, including the wall budget, and skips terms that are absent instead of printing None. The thread tests wait for the sample count they need rather than sleeping a fixed interval. The CHANGELOG no longer calls the two allocator counters a resident figure.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The showcase, benchmark, latent-capture and A/B harness workers each run under a watchdog thread that aborts the process before it can exhaust unified memory. Until now that thread compared MLX's active memory against the ceiling. MLX keeps freed buffers in a retained cache that is still resident, and it reports that cache separately, so the number the watchdog compared was not the number MLX considers in use. Under the harness caps the memory limit already drains that cache before any allocation that would cross it (checked on mlx 0.32.2 with synthetic buffers), so this was never an open door to a kernel panic; it was a reading that could not say how close a run actually came, and a policy the report could not state.
Both watchdogs (
scripts/run_showcase.py,scripts/_capture_latent.py) now sample active plus cached bytes and compare the sum against the ceiling. Each one bounds the cache pool withmx.set_cache_limitwhen it starts, so every model-loading path gets the bound from the same place and the abort artifact can state it: 4 GiB for the generation workers, 4 GiB for the decode reps (their heaviest peak is 3.7 GiB, so the timed decode still runs from a warm pool), and the A/B script passes its existing per-condition bounds through. Polling moves from 0.5 s to 50 ms; both getters are plain reads and the poll runs whilemx.evalholds no GIL. The watchdog keeps high-water marks of the sum it compares and of the cache term alone, and each worker records those next to the policy (ceiling, cache bound, cadence, wall budget): in the showcase live result, in the bench sentinel and the per-condition report block, and in the A/B unit result. A memory sample that raises used to kill the daemon thread and leave the run with no backstop; it now aborts with reasonsample_errorand the error text, and the orchestrators print every term the watchdog compared.Nothing in the library changes. The showcase report and COMPARISON numbers were not re-measured; COMPARISON says which accounting its numbers came from. The 4 GiB live bound was checked against the committed live wall clocks:
live_previewat 512x512 ran three interleaved reps at 4 GiB and at 20 GiB, median 10.76 s against 11.07 s, with the retained pool topping out at about 4.1 GiB against 17 GiB, so the bound is neutral for the timings and only limits how much the process holds at teardown.Evidence: 19 new offline tests drive the real watchdog thread against a fake MLX memory API, and each was checked against the one-line bug it names (dropping the cache from the sum, skipping the cache bound, installing it after the first sample, a swallowed sample error, a cadence regression, a rep worker inheriting the live bound, a result or report block that omits the policy or the observed peak, an orchestrator message that hides the cache term or the error text, the A/B condition-to-cap mapping collapsing). On the M1 Max, a throwaway check with a 1.5 GiB synthetic ceiling aborted a worker whose active memory was 0.9 GiB and retained cache 0.8 GiB (exit 70, artifact showing both terms), and let a 1.0 GiB cached-only run finish. Full suite: 558 passed, coverage 98.6%.