Repository navigation
Add the latent-contracts research memo - #40
Merged
Merged
Conversation
Documents what mflux's in-loop callback hands a preview decoder for nine image models across three VAE families: the packing, the sub-pixel fold order, the spatial divisor and the normalization all live in the generator, so sharing a VAE, even byte-identical weights, does not mean sharing a latent contract. Includes a synthetic-latent figure of the FLUX.2 and Ideogram 4 fold orders with its generator script, a hash table for the shared VAE files, and a README link.
Sorted imports, type annotations and a docstring on the two helpers, one statement per line. The rendered figure is byte-identical.
The memo no longer asserts that TAEF2 was distilled on the raw latent: mlx-taef applies the batch-norm inverse before it, the reference wrapper and ComfyUI do not, and the repository has not measured which is right. The identity-BN degradation is described as documented rather than observed, the canonical taew2.1 link is pinned to a commit, the Klein row and one line anchor are corrected, and the README summary matches the memo.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds the first research memo to
docs/papers/: The latent in the callback is not the latent the decoder wants.The memo reads nine mflux image models at the public
v.0.19.1tag and records what each one hands its in-loop callback: the tensor layout, the spatial divisor, where the 2×2 sub-pixel unpack happens, and which component applies the normalization. Four of those models decode through the same VAE class, and the Hub confirms two of them ship byte-identical weights, yet they arrive in three layouts and two normalization schemes. FLUX.2 Klein and Ideogram 4 fold the sub-pixels into the packed channel axis in different orders, and a synthetic latent pushed through both orders shows the plausible-looking scramble that produces. The memo also records that a Hugging Face mirror of the taew2.1 tiny decoder differs from the canonical file in 87% of its values while matching it in size, tensor names, shapes and dtype, which is why every weight source in this library is pinned by digest.Nothing was generated or benchmarked for it. Every per-model claim is read from pinned source; two file-level checks (Hub digests, the weight comparison) are dated in the text. The README gains a short "Research notes" section linking the memo. No package code changes.
A follow-up commit, after review, corrects one claim: the memo now states that mlx-taef applies the VAE's batch-norm inverse before TAEF2 while the reference wrapper and ComfyUI do not, and that which input domain the decoder expects has not been measured. That measurement is tracked separately.