Skip to content

Decode the normalized latent with TAEF2 by default (v0.8.2) - #41

Merged
IonDen merged 12 commits into
mainfrom
fix/flux2-unpack-oracles-and-taef2-domain
Sep 17, 2026
Merged

IonDen merged 12 commits into
mainfrom
fix/flux2-unpack-oracles-and-taef2-domain

Conversation

@IonDen

@IonDen IonDen commented Sep 17, 2026 •

Copy link
Copy Markdown
Owner

TAEF2 live previews of FLUX.2 Klein have been darker and more saturated than the image mflux finally returns. This PR fixes the cause and re-measures everything the fix touches.

Since v0.2.0, passing flux=model to LivePreviewCallback made it read the Flux2VAE batch-norm running statistics and apply the VAE's inverse normalization to the latent before handing it to TAEF2. The reasoning was that the full VAE does the same. TAEF2, though, wants the normalized latent the diffusion model emits, which is also how its upstream reference code and ComfyUI feed it; the inverse widens the latent by about 1.8× per channel. A new script, scripts/ab_taef2_bn_domain.py, decodes one captured latent with the full VAE and with TAEF2 both ways, one model per process, with lossless PNGs between decoder and scorer:

Latent TAEF2 input SSIM vs full VAE LPIPS
committed 512×512 fixture latent as mflux produces it 0.920 0.058
same batch-norm inverse applied (old default) 0.588 0.241
fully denoised 768×512 control latent as mflux produces it 0.945 0.022
same batch-norm inverse applied 0.716 0.150

The statistics used in the old path were checked against the VAE weights and are exact, and a test now pins that path to mflux's own denormalize-then-unpatchify order, so the old arm was not handicapped. The images and reports for both latents are committed under _artifacts/ab_taef2_bn_domain/. The 512×512 latent is the committed fixture; the control latent is not committed, and its capture recipe and sha256 sit next to its images.

What users see: auto_bn now defaults to False, so previews get closer to the final image with no code change. auto_bn=True with flux=model, or explicit bn_mean / bn_var, still apply the inverse and now log a warning that says what it costs. callback.resolved_bn is "none" by default; code that asserts "auto" needs to drop the assertion or opt in. Nothing else in the public API changes, and the FLUX.1, Z-Image, Qwen-Image and Krea 2 paths are untouched.

Re-measured under the showcase protocol (M1 Max, mflux 0.19.1, MLX 0.32.2): TAEF2 against the full FLUX.2 VAE is SSIM 0.960 / LPIPS 0.057, up from 0.616 / 0.216, at 30 ms against 0.27 s and 0.59 GB against 2.80 GB. The FLUX.2 live-preview and TeaCache scenarios and the FLUX.2 Klein example frames were re-captured; the whole-generation TeaCache ratio now reads 1.27× (it was 1.41× under mflux 0.18.1, from a single run per scenario, and the page says so). The FLUX.1 and Z-Image rows keep their earlier run, which the report now records per scenario: the showcase gained --update-report so one scenario can be re-measured without discarding the rest, and a refresh that fails or is interrupted leaves both the report and that scenario's previous artifacts as they were. The A/B script stages its work the same way, so re-running it into a directory that already holds results never leaves an old report next to new images. The 0.960 is scored on the showcase's webp files; the A/B's 0.920 is the same decode scored on lossless PNGs.

Tests: the FLUX.2 unpack is now compared element for element with mflux's own inverse on a non-square arange tensor. The test it replaces pinned four values and let two of three deliberate axis-order breaks through. The Qwen-Image unpack oracle moves to the same form. 540 offline tests pass, coverage 98.6%.

The version entry in CHANGELOG, README and ROADMAP is 0.8.2.

The FLUX.2 test checked four hand-picked elements; two of three deliberate
axis-order breaks passed it. It now composes mflux's two static unpack steps
as the oracle on a non-square arange tensor and asserts exact equality. The
Qwen oracle moves from a random tensor with a tolerance to the same form.
Adds a resumable one-model-per-process A/B that decodes a captured FLUX.2
latent with the full VAE and with TAEF2 on the normalized and the
batch-norm-denormalized latent, then scores both against the full VAE
(SSIM, LPIPS, MAE, mean RGB shift). Committed results: on the 512x512
fixture the normalized latent scores SSIM 0.920 / LPIPS 0.058 against
0.588 / 0.241 with the inverse applied; a fully denoised 768x512 control
gives 0.945 / 0.022 against 0.716 / 0.150.
LivePreviewCallback used to extract the Flux2VAE batch-norm statistics from
flux=model and apply the inverse before TAEF2. Measured against the full VAE
that path is the worse one (SSIM 0.59 vs 0.92 on the committed latent, 0.72
vs 0.94 on a fully denoised control): TAEF2 was distilled on the normalized
latent mflux hands the callback. auto_bn now defaults to False. Passing
auto_bn=True or explicit bn_mean/bn_var still applies the inverse and logs a
warning saying what it costs. The warning that fired on the identity path is
gone, since that path is now the correct one.
The TAEF2 bench condition and the FLUX.2 showcase and example recipes no
longer apply the batch-norm inverse, so they time and score the image the
library produces by default. Both harness scripts now forward their auto_bn
recipe flag to the callback instead of relying on flux= alone.
--update-report refreshes a single --scenario, keeps every other scenario
and its artifacts, moves only the refreshed scenario's old artifacts to the
Trash, and records the refresh's time, hardware and versions under
scenario_updates. A failed refresh leaves the report untouched.
TAEF2 against the full FLUX.2 VAE on the benchmark latent is now SSIM 0.960
and LPIPS 0.057 (was 0.616 and 0.216) at 30 ms against 0.27 s. live_preview
and combined were re-captured under mflux 0.19.1 and MLX 0.32.2. The FLUX.1
and Z-Image scenarios are unchanged.
@IonDen IonDen self-assigned this Sep 17, 2026
@IonDen IonDen added bug Something isn't working documentation Improvements or additions to documentation labels Sep 17, 2026
A failed --update-report moved the scenario's artifacts to the Trash before
running and left the unchanged report pointing at them; it now puts them
back. Report provenance records a dirty tree. The A/B script keys its resume
on the latent's sha256, bounds the MLX cache pool in every worker, keeps one
arm's LPIPS when the other arm's scoring fails, and labels its wall-clock and
process-peak fields as what they are. The fidelity warning no longer fires
for variants whose unpack never reads batch-norm statistics. New tests cover
each of these, the provenance history across two refreshes, a truncated
result file, and the opt-in inverse against mflux's own denormalize and
unpatchify order.
… control

Both measurements reproduce to the last digit. The control latent is not
committed; its capture recipe and sha256 now sit next to its images, its
files carry a distinct name, and the docs say so. The headline 0.960 is
labelled as webp-scored wherever it appears.
…rupt-safe

The A/B script wrote each unit straight into --out-dir, so re-running it on a
different latent and failing midway left the old report beside images decoded
from the new latent. Units now go to a .staging-<sha> directory and move over
in one step after every decode and the scoring succeeded; a retry resumes from
the staged units, and a complete result is not re-run. --update-report now
restores the previous artifacts on Ctrl-C as well as on errors, and refuses
--no-trash-prior, which would leave nothing to restore. The report's source
version reads dirty only when src, scripts, pyproject.toml or uv.lock have
uncommitted changes: the harnesses rewrite tracked artifacts while they run,
so a repo-wide check said dirty every time. Worker memory limits move into a
helper so a test can see that the cache limit is really installed. New tests
cover each of these, the orchestrator stopping at the first failed unit, both
auto_bn branches of the live scenario wiring, and the sign of the image
statistics.
The docs stated as fact that TAEF2 was distilled on the normalized latent.
What is known is the measurement, that upstream's reference code and ComfyUI
feed it the same way, and that the inverse widens the latent about 1.8 times
per channel. The control recipe now captures into a private temp directory
and describes reproducibility in terms of the latent's sha256.
Publishing moved a unit's result file before its image, and completeness
looked at result files only, so Ctrl-C between two moves could leave
--out-dir reading as already measured with an image still in staging. A unit
now counts only when its image exists, images move first and the report
last, and the next run copies back whatever an interrupted publish already
moved instead of decoding again. Tested at every interruption point. The
real-git dirtiness test also switches off commit and tag signing so a signing
git config cannot break it.
@IonDen
IonDen merged commit a16ec15 into main Sep 17, 2026
7 checks passed
@IonDen
IonDen deleted the fix/flux2-unpack-oracles-and-taef2-domain branch September 17, 2026 20:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant