Repository navigation
Decode the normalized latent with TAEF2 by default (v0.8.2) - #41
Merged
Merged
Conversation
The FLUX.2 test checked four hand-picked elements; two of three deliberate axis-order breaks passed it. It now composes mflux's two static unpack steps as the oracle on a non-square arange tensor and asserts exact equality. The Qwen oracle moves from a random tensor with a tolerance to the same form.
Adds a resumable one-model-per-process A/B that decodes a captured FLUX.2 latent with the full VAE and with TAEF2 on the normalized and the batch-norm-denormalized latent, then scores both against the full VAE (SSIM, LPIPS, MAE, mean RGB shift). Committed results: on the 512x512 fixture the normalized latent scores SSIM 0.920 / LPIPS 0.058 against 0.588 / 0.241 with the inverse applied; a fully denoised 768x512 control gives 0.945 / 0.022 against 0.716 / 0.150.
LivePreviewCallback used to extract the Flux2VAE batch-norm statistics from flux=model and apply the inverse before TAEF2. Measured against the full VAE that path is the worse one (SSIM 0.59 vs 0.92 on the committed latent, 0.72 vs 0.94 on a fully denoised control): TAEF2 was distilled on the normalized latent mflux hands the callback. auto_bn now defaults to False. Passing auto_bn=True or explicit bn_mean/bn_var still applies the inverse and logs a warning saying what it costs. The warning that fired on the identity path is gone, since that path is now the correct one.
The TAEF2 bench condition and the FLUX.2 showcase and example recipes no longer apply the batch-norm inverse, so they time and score the image the library produces by default. Both harness scripts now forward their auto_bn recipe flag to the callback instead of relying on flux= alone.
--update-report refreshes a single --scenario, keeps every other scenario and its artifacts, moves only the refreshed scenario's old artifacts to the Trash, and records the refresh's time, hardware and versions under scenario_updates. A failed refresh leaves the report untouched.
TAEF2 against the full FLUX.2 VAE on the benchmark latent is now SSIM 0.960 and LPIPS 0.057 (was 0.616 and 0.216) at 30 ms against 0.27 s. live_preview and combined were re-captured under mflux 0.19.1 and MLX 0.32.2. The FLUX.1 and Z-Image scenarios are unchanged.
A failed --update-report moved the scenario's artifacts to the Trash before running and left the unchanged report pointing at them; it now puts them back. Report provenance records a dirty tree. The A/B script keys its resume on the latent's sha256, bounds the MLX cache pool in every worker, keeps one arm's LPIPS when the other arm's scoring fails, and labels its wall-clock and process-peak fields as what they are. The fidelity warning no longer fires for variants whose unpack never reads batch-norm statistics. New tests cover each of these, the provenance history across two refreshes, a truncated result file, and the opt-in inverse against mflux's own denormalize and unpatchify order.
… control Both measurements reproduce to the last digit. The control latent is not committed; its capture recipe and sha256 now sit next to its images, its files carry a distinct name, and the docs say so. The headline 0.960 is labelled as webp-scored wherever it appears.
…rupt-safe The A/B script wrote each unit straight into --out-dir, so re-running it on a different latent and failing midway left the old report beside images decoded from the new latent. Units now go to a .staging-<sha> directory and move over in one step after every decode and the scoring succeeded; a retry resumes from the staged units, and a complete result is not re-run. --update-report now restores the previous artifacts on Ctrl-C as well as on errors, and refuses --no-trash-prior, which would leave nothing to restore. The report's source version reads dirty only when src, scripts, pyproject.toml or uv.lock have uncommitted changes: the harnesses rewrite tracked artifacts while they run, so a repo-wide check said dirty every time. Worker memory limits move into a helper so a test can see that the cache limit is really installed. New tests cover each of these, the orchestrator stopping at the first failed unit, both auto_bn branches of the live scenario wiring, and the sign of the image statistics.
The docs stated as fact that TAEF2 was distilled on the normalized latent. What is known is the measurement, that upstream's reference code and ComfyUI feed it the same way, and that the inverse widens the latent about 1.8 times per channel. The control recipe now captures into a private temp directory and describes reproducibility in terms of the latent's sha256.
Publishing moved a unit's result file before its image, and completeness looked at result files only, so Ctrl-C between two moves could leave --out-dir reading as already measured with an image still in staging. A unit now counts only when its image exists, images move first and the report last, and the next run copies back whatever an interrupted publish already moved instead of decoding again. Tested at every interruption point. The real-git dirtiness test also switches off commit and tag signing so a signing git config cannot break it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TAEF2 live previews of FLUX.2 Klein have been darker and more saturated than the image mflux finally returns. This PR fixes the cause and re-measures everything the fix touches.
Since v0.2.0, passing
flux=modeltoLivePreviewCallbackmade it read the Flux2VAE batch-norm running statistics and apply the VAE's inverse normalization to the latent before handing it to TAEF2. The reasoning was that the full VAE does the same. TAEF2, though, wants the normalized latent the diffusion model emits, which is also how its upstream reference code and ComfyUI feed it; the inverse widens the latent by about 1.8× per channel. A new script,scripts/ab_taef2_bn_domain.py, decodes one captured latent with the full VAE and with TAEF2 both ways, one model per process, with lossless PNGs between decoder and scorer:The statistics used in the old path were checked against the VAE weights and are exact, and a test now pins that path to mflux's own denormalize-then-unpatchify order, so the old arm was not handicapped. The images and reports for both latents are committed under
_artifacts/ab_taef2_bn_domain/. The 512×512 latent is the committed fixture; the control latent is not committed, and its capture recipe and sha256 sit next to its images.What users see:
auto_bnnow defaults toFalse, so previews get closer to the final image with no code change.auto_bn=Truewithflux=model, or explicitbn_mean/bn_var, still apply the inverse and now log a warning that says what it costs.callback.resolved_bnis"none"by default; code that asserts"auto"needs to drop the assertion or opt in. Nothing else in the public API changes, and the FLUX.1, Z-Image, Qwen-Image and Krea 2 paths are untouched.Re-measured under the showcase protocol (M1 Max, mflux 0.19.1, MLX 0.32.2): TAEF2 against the full FLUX.2 VAE is SSIM 0.960 / LPIPS 0.057, up from 0.616 / 0.216, at 30 ms against 0.27 s and 0.59 GB against 2.80 GB. The FLUX.2 live-preview and TeaCache scenarios and the FLUX.2 Klein example frames were re-captured; the whole-generation TeaCache ratio now reads 1.27× (it was 1.41× under mflux 0.18.1, from a single run per scenario, and the page says so). The FLUX.1 and Z-Image rows keep their earlier run, which the report now records per scenario: the showcase gained
--update-reportso one scenario can be re-measured without discarding the rest, and a refresh that fails or is interrupted leaves both the report and that scenario's previous artifacts as they were. The A/B script stages its work the same way, so re-running it into a directory that already holds results never leaves an old report next to new images. The 0.960 is scored on the showcase's webp files; the A/B's 0.920 is the same decode scored on lossless PNGs.Tests: the FLUX.2 unpack is now compared element for element with mflux's own inverse on a non-square
arangetensor. The test it replaces pinned four values and let two of three deliberate axis-order breaks through. The Qwen-Image unpack oracle moves to the same form. 540 offline tests pass, coverage 98.6%.The version entry in CHANGELOG, README and ROADMAP is 0.8.2.