The Flash MXFP4 GGUF (ds4f-mxfp4) cannot be used as a donor or base with gguf-tools/mixed/splice_mixed_expert_layers_gguf.py, because its GGML_QUANT_SIZES table has no entry for type 39.
error: unsupported GGML tensor type 39; add it to GGML_QUANT_SIZES
Everything else in the splicer already handles it correctly — it deliberately does not require base and donor tensor types to match, and it writes each tensor's own ggml_type through. Only the block-size table is missing an entry.
One-line fix, using the values from ds4's own table at ds4.c:2052 ([39] = {"mxfp4", 32, 17}):
GGML_QUANT_SIZES = {
...
26: (1, 4, "I32"),
+ 39: (32, 17, "MXFP4"),
}
Verified locally with --dry-run, splicing MXFP4 routed experts for layers 25-42 onto the IQ2XXS/Q2_K hybrid base:
selected donor tensors: 54
selected donor types: MXFP4:54
base tensor payload: 90.88 GiB
mixed tensor payload: 107.76 GiB
Motivation, in case it is useful: this produces a quant sized for a single 128 GB Mac. The published Flash quants are 80.8 / 90.9 GiB (fit, leaving ~21 GiB of the 112 GiB wired limit unused) then jump to 145.3 / 153.3 GiB (do not fit). Splicing MXFP4 experts into the upper layers fills that gap at 107.76 GiB.
Happy to open a PR if useful.
The Flash MXFP4 GGUF (
ds4f-mxfp4) cannot be used as a donor or base withgguf-tools/mixed/splice_mixed_expert_layers_gguf.py, because itsGGML_QUANT_SIZEStable has no entry for type 39.Everything else in the splicer already handles it correctly — it deliberately does not require base and donor tensor types to match, and it writes each tensor's own
ggml_typethrough. Only the block-size table is missing an entry.One-line fix, using the values from ds4's own table at
ds4.c:2052([39] = {"mxfp4", 32, 17}):Verified locally with
--dry-run, splicing MXFP4 routed experts for layers 25-42 onto the IQ2XXS/Q2_K hybrid base:Motivation, in case it is useful: this produces a quant sized for a single 128 GB Mac. The published Flash quants are 80.8 / 90.9 GiB (fit, leaving ~21 GiB of the 112 GiB wired limit unused) then jump to 145.3 / 153.3 GiB (do not fit). Splicing MXFP4 experts into the upper layers fills that gap at 107.76 GiB.
Happy to open a PR if useful.