build: generate header prerequisites with -MMD -MP (#1741) - #1758
Conversation
Every engine and test rule listed its headers by hand on one line, the most edited line in the Makefile and a merge conflict for every open PR touching the same rule. The lists also drifted: many built targets on dev read headers their rule does not name (measured in the PR), so editing one of those leaves a stale binary while make reports success. -MMD -MP in CFLAGS now writes <target>.d beside each binary and object, and the Makefile includes those files. Header lists are removed from 166 rules (9 engines, 154 tests, 3 objects). Three cases the compiler cannot cover on its own: - A command compiling several .c files writes the last unit's deps only. Eight such tests keep their lists (AUTODEP_MULTI_TU). With CUDA/HIP, qwen36_tier.c joined 14 command lines the same way; it is now its own object, qwen36_tier.o, and QWEN36_TIER_SRC is renamed QWEN36_TIER_OBJ. - Only 6 of the 166 rules depend on .build-config, so the CFLAGS change would not rebuild the rest and they would have no .d. Each target therefore depends on its own .d through an empty rule: a missing .d makes it out of date, and the rebuild writes one. - Rules built through Makefile.deepseek-v4 or the adapter targets (AUTODEP_ELSEWHERE), and the nvcc, MSVC-hosted CUDA and Metal objects, keep their lists; nvcc never receives CFLAGS. test_makefile_deps.py now checks that every rule in AUTODEP_BINS compiles one unit with -MMD in the default and the GPU configuration, that no such rule lists headers by hand, that every built target's .d covers its source's unconditional includes, and that make -q -W on a header from a .d answers "rebuild". clean.py removes *.d; c/.gitignore ignores them. test_glm53_metal_source.py no longer expects backend_metal.h in the glm53 rule.
test_registry_engine_agreement treats an engine as accelerated when its rule names a backend object. QWEN36_TIER_SRC became QWEN36_TIER_OBJ in the previous commit, so the old entry matched nothing; qwen36 and qwen38 still passed through CUDA_OBJ, which hid it.
- test_makefile_deps: every make call keeps .build-config's contents and timestamp. Parsing with CUDA=1 or CUDA_DLL=1 rewrote it, so running the suite switched the tree's recorded configuration and the next make relinked everything that depends on it. A new test pins this. - clean.py: remove the .d files under tools/ and build/segment/ too, and the backend_vulkan.o and backend_xdna.o objects. After a clean that removed their .d, those objects stayed in the tree with no record of the headers they read, and the coverage test rightly flagged them. - One sed pass over the rule names now serves both TEST_RULES and ENGINE_RULES; both variables are unchanged. - A stale comment still said the qwen36 tests compile qwen36_tier.c; the deepseek_v41 rule loses a leftover line continuation.
Eight tests compiled a second .c on their own command line (segment_runtime.c, edge_runtime.c, backend_xdna.c, deepseek_v4.c and a fixtures file). GCC writes a .d for the last unit only, so they kept hand-written header lists as exceptions, and those lists had the same kind of gaps as the ones -MMD replaced elsewhere. Each helper is now its own object in tests/, compiled with exactly the flags that test gave it before: -DCOLI_XDNA for the XDNA lane object, -DCOLI_V4_UNIT_NATIVE_QUANT for deepseek_v4.c, and NOCUDA_CFLAGS for the two qwen38 tests, which get their own copies of segment_runtime.o and edge_runtime.o since those flags differ from CFLAGS in GPU builds. Each unit was already compiled on its own, so the result is the same. AUTODEP_MULTI_TU is gone, and with it the last hand-written header lists on gcc-built tests; the one-unit check in test_makefile_deps.py now points at objects instead of an exception list. clean.py removes tests/*.o.
Headers under #if reach a .d only through the compile that reads them: backend_xdna.h with XDNA=1, backend_vulkan.h with VK=1, the qwen36 tier only with CUDA or HIP. The coverage test cannot require them from the source text, so what stands behind them is the one-unit check, and that only ran in the default and one GPU configuration. It now runs in every flavour the host's make accepts: CUDA=1, CUDA_DLL=1, HIP=1 (with a fixed HIP_ARCH, as the CI syntax job uses), HIP_DLL=1, VK=1, XDNA=1 and METAL=1, each recognised by the define colibri then gets. On Windows that is CUDA_DLL, HIP_DLL, VK and XDNA; on Linux CUDA, HIP, VK and XDNA. Putting backend_xdna.c on colibri's command line under XDNA=1 passes the default check and fails this one, in XDNA=1 only.
The comment said all six are built outside this Makefile's own $(CC) recipes. Only the two DeepSeek V4 tests are; the four segment/edge adapter tests are compiled here, three units in one command.
The first slice covered the engines, the tests/test_* rules and the objects they link. The other rules gcc compiles with $(CFLAGS) still listed their headers by hand, and drifted the same way: the benchmarks under tests/, the XDNA probe, the objects of the segment library and of the V4 ownership test, and the rANS ctypes library. - AUTODEP_BINS now takes every rule under tests/, not only tests/test_*, plus $(SEGMENT_ALL_OBJS), $(SEGMENT_RUNTIME_OBJS), $(V4_OWNERSHIP_OBJS) and $(RANSLIB). Their header lists are gone. - AUTODEP_OWN_FLAGS names the three tests/ rules that do not compile with $(CFLAGS) (a GPU compiler, and two benchmarks pinned to their own flags whose sources include no local header), so they write no .d. - xdna_physical_probe compiled backend_xdna.c beside its own source; it now links tests/backend_xdna_lane.o, which has exactly its flags. - clean.py removes build/segment/ as a whole, as it does build/ownership/. It used to leave the segment objects, and one left without its .d is a built target whose headers nothing tracks. - test_makefile_deps.py resolves rules spelled through variables ($(SEGMENT_BUILD_DIR)/glm.o:, $(RANSLIB):, the V4 static pattern rule) by asking make for them, so the hand-list and coverage checks see those rules instead of skipping them. The run targets fuzz-rans, dsv4-cuda-test and dsv4-cuda-loader-test keep their lists: their names are not files, so they rebuild every time.
|
@JustVugg macOS is green, so this is out of draft. What the macOS logs do and don't show is under CI on this PR in the description. |
|
@huppiflupp a small, optional ask, since you run a Strix Halo on Linux: if ROCm is already installed there, could you build one engine from this branch with git clone --depth 1 -b build/makefile-autodeps https://github.com/Kenneth-Javier/colibri colibri-1758
cd colibri-1758/c
/opt/rocm/bin/hipcc --version | head -1
make colibri HIP=1 # HIP_ARCH defaults to native
grep -c backend_cuda.h colibri.d # expect 1: a header only a CUDA/HIP build reads
make -q -W backend_cuda.h colibri HIP=1; echo $? # expect 1: make would rebuild after it changes
python3 -m unittest tests.test_makefile_deps # expect OK (a skip or two is fine)The last line or two of each is plenty. No rush, and if ROCm isn't on that machine, never mind: it isn't worth installing for this. |
|
@huppiflupp one correction for Nobara, going by your HARDWARE.md and setup.sh: Fedora's ROCm packages install under hipcc --version | head -1
rpm -q rocwmma-devel # gfx1151 needs its headers; if it isn't installed, skip the whole thing
make colibri HIP=1 ROCM_HOME=/usr HIP_ARCH=gfx1151
grep -c backend_cuda.h colibri.d
make -q -W backend_cuda.h colibri HIP=1 ROCM_HOME=/usr HIP_ARCH=gfx1151; echo $?
python3 -m unittest tests.test_makefile_deps( |
|
Built on Nobara 44 / gfx1151 (hipcc 7.1.52802, rocwmma-devel 7.1.0) with |
|
Thanks, that closes the last gap outside macOS, on the Fedora packaging I couldn't test. You're right about the count: my "expect 1" should have said 2, and the |
The VK_OBJ test errored in setUpClass on the macOS runner: it read $(CC) and $(EXE) through `make --eval`, and /usr/bin/make on macOS is GNU Make 3.81, which predates --eval (3.82). Its stdout was empty; the rest of the suite passed (1479 run, 156 skipped, this one error). $(EXE) now comes from make's own database (`make -pn`), which 3.81 prints the same way, and the compiler is simply the first word of each link line. Two holes in reading `make -n` go with it: - recipe continuations are printed as written, so the five link commands that span lines (test_deepseek_v4 and the four segment/edge adapter tests) were never checked. They are folded first now. - a rule that writes a differently named file (fuzz-rans, bench-omp-grain, the dsv4 CUDA tests, glm53-metal-check) was skipped, because its output was not a rule name. Every printed command that compiles a .c into an output without -c is checked now, and such a rule is reported by its output name if it ever needs $(VK_OBJ). The scan covers 204 link lines on dev plus this branch (198 before). None of the added ones calls the backend, so the results against the linker are unchanged: - Linux (WSL Ubuntu 24.04, make 4.3): dev fails with the thirteen rules; this branch passes, and the 60 lines that call the backend are the 60 that link backend_vulkan.o; with JustVugg#1338 merged only test_qwen36_slot_int8 is reported. About 22 s. - Windows (make 4.4.1): the same, with twelve on dev, since bench_idot cannot preprocess without ARCH=native. About 20 s. - dev + JustVugg#1758 + this branch, conflicts resolved: passes together with JustVugg#1758's test_makefile_deps and writes no .d file. Not verified locally: GNU Make 3.81 itself. The macOS runner is the check.
dev added tok_unicode_deepseek.h to 84 hand-written prerequisite lists (JustVugg#1775). All 84 rules take their headers from -MMD here, so each conflict resolves to the line without headers, and the new tests/test_tok_deepseek rule keeps its recipe without a list. dev touched none of the rules that still carry a list.
dev added qwen38_vision.h to 19 hand-written prerequisite lists (JustVugg#1777). All 19 rules take their headers from -MMD here, so each conflict resolves to the line without headers; dev touched none of the rules that still carry a list. JustVugg#1762's Windows Vulkan link merges cleanly.
The first version of test_makefile_vk_obj.py read preprocessor conditionals as text: a rule needed $(VK_OBJ) when a .c it compiled reached a file with `#if ... COLI_VULKAN`. That is exact on dev, but not in general. With #1338 (the qwen36 Vulkan expert tier) merged on top, it flagged 25 rules where the linker fails one: - the qwen36 tier tests pass -DCOLI_CUDA, and qwen36_tier.c picks CUDA over Vulkan with #elif; - qwen36.c's Vulkan hook only reads COLI_VULKAN from the environment; - test_qwen36_tier_vk_fake defines the coli_vk_* functions itself. Skipping lines that pass -DCOLI_CUDA is not a fix either: kimi_k3.c keeps independent #ifdef COLI_VULKAN blocks, and test_kimi_cuda_expert does need backend_vulkan.o. The test now asks the tools. `make -Bn VK=1` gives each expanded link line, and every .c on it is preprocessed with that line's own flags (link-only and -MMD/-MF arguments dropped). A line needs $(VK_OBJ) when the surviving code of this tree, outside backend_vulkan.h, calls a function backend_vulkan.h declares and does not define it. String literals are ignored. A line this host cannot preprocess is left to the platforms that can build it. .build-config is restored afterwards, so the dry run does not make the next real build relink. The engines whose own source calls the backend anchor the non-vacuity check. Measured against the linker (make -k test-c VK=1 plus the on-demand targets): - dev 4e28e39: flags exactly the thirteen rules the previous commit fixed on Linux; on Windows twelve, since bench_idot cannot even preprocess there without ARCH=native (its #error). - dev + the previous commit: passes. On Linux the 60 lines it says call the backend are exactly the 60 that link backend_vulkan.o. - dev + #1338 + this branch: flags only test_qwen36_slot_int8, which is also the only link the linker fails there. - dev + #1758 + this branch, conflicts resolved: passes together with #1758's test_makefile_deps; with -MMD -MP in CFLAGS no .d file is written. Runtime: 15 s on Windows (gcc 16.1.0, 32 threads), 20 s on WSL Ubuntu 24.04 (gcc 13.3.0). The first version took 1.5 s. Not verified: macOS / Apple clang, where the linemarker format is the same but the test was not run.
|
This is still wanted: dev has had many missing header dependencies. Two things before it can go in, and it should be the last of the current batch.
-qwenimage$(EXE): qwenimage.c qi_gemm.h qwenimage_vae.h qwenimage_preview.h st.h json.h tok.h tok_unicode.h tok_unicode_o200k.h tok_unicode_deepseek.h qwen38_nfc.h qwen38_nfc_tables.h serve_poll.h compat.h .build-config
+qwenimage$(EXE): qwenimage.c .build-config
-tests/test_qi_gemm$(EXE): tests/test_qi_gemm.c qi_gemm.h
+tests/test_qi_gemm$(EXE): tests/test_qi_gemm.c
-tests/test_qwenimage_vae$(EXE): tests/test_qwenimage_vae.c qwenimage_vae.h qi_gemm.h st.h json.h compat.h
+tests/test_qwenimage_vae$(EXE): tests/test_qwenimage_vae.cWith those, the guards pass and touching |
dev brought new rules with hand-written header lists and new multi-source compiles (JustVugg#1323, JustVugg#1343, JustVugg#1780); here they take their headers from -MMD: - qwenimage, test_qi_gemm, test_affine_quant, test_affine_dequant, test_qwenimage_vae and test_qwen36_dense_affine list no headers. - qwen36's qpack sources (METAL=1 only) are objects, qwen36_qpack.o and qpack.o, compiled with $(QWEN36_CFLAGS) as qwen36's command line gave them; test_qwen36_qpack and test_qpack link helper objects compiled with their own flags. clean.py removes the two root objects. - QWEN36_TIER_SRC in dev's new lines is QWEN36_TIER_OBJ here. qpack-inspect keeps its two-unit compile and its list: the '-' in its name keeps it out of the generated set. The Metal rules keep theirs.
JustVugg#1763 added $(VK_OBJ) to the prerequisites and link line of 13 test and benchmark rules; each conflict keeps this branch's rule and adds $(VK_OBJ) where JustVugg#1763 put it. JustVugg#1790's bench_matmul_f32 and test_matmul_f32 list no headers here. TEST_BINS is the same 178 targets as on dev, and test_tok_deepseek is among them.
|
Done, merged up to |
|
The batch before this one is now in (#1739, #1740, #1761, #1798, #1799, #1812 and #1813 among others), so this is the moment for the final rebase. Two things on top of what we wrote before:
-mimo$(EXE): mimo.c mimo_vision.h sse41_kernels.h cli_args.h compat.h json.h st.h quant.h idot.h fp8_format.h \
- ... .build-config
+mimo$(EXE): mimo.c .build-configWith it |
JustVugg#1798 added matmul_f32.h to 10 hand-written prerequisite lists; all 10 rules take their headers from -MMD here, so each conflict resolves to the line without headers. JustVugg#1813's mimo and the new test_dsv41_untrusted_load list no headers either: mimo.d names the 19 headers mimo's list had plus matmul_f32.h, which it did not. TEST_BINS is the same 179 targets as on dev, test_tok_deepseek among them.
|
Merged |
…ustVugg#1758) JustVugg#1758 removed hand-written header lists from 202 rules and renamed QWEN36_TIER_SRC to QWEN36_TIER_OBJ. The rule this PR added predates it and named five headers plus the old variable, so after the rebase onto dev it no longer resolved. Now spelled exactly like its neighbour tests/test_qwen36_slot_int8: one translation unit, $(QWEN36_CFLAGS) so -MMD -MP writes the .d, no header list. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Implements #1741 as specified there: every engine, test, benchmark and object that gcc or clang compiles with
$(CFLAGS)now takes its header prerequisites from the compiler, and those rules no longer list headers by hand. What still keeps a list, and why, is under point 3.What changes
CFLAGS += -MMD -MP, after the platform blocks, so it is part ofBUILD_CONFIG.build/segment/, 5 objects underbuild/ownership/, 3 backend objects and the ctypes library.colibri$(EXE):is nowcolibri.c $(CUDA_OBJ) $(METAL_OBJ) $(VK_OBJ) $(VK_SPV) $(XDNA_OBJ) .build-config..con their own command line; each helper is now its own object intests/, compiled with the flags that test gave it.qwen36_tier.clikewise becomesqwen36_tier.o, andQWEN36_TIER_SRCis renamedQWEN36_TIER_OBJ..dfiles are-included at the end of the Makefile, next toAUTODEP_BINS, the 239 targets they cover. The engines and every rule undertests/are derived from onesedpass, the wayTEST_RULESwas, so a new one joins on its own; the objects come from the variables that already list them.clean.pyremoves the.dfiles (c/,tests/,tools/), the newtests/*.o, andbackend_vulkan.oandbackend_xdna.o, which it never removed. It now removesbuild/segment/as a whole, as it already didbuild/ownership/: it used to leave the segment objects behind, and one left without its.dis a built target whose headers nothing tracks.c/.gitignoreignores*.d.test_makefile_deps.pyis rewritten, see below; it reads rules spelled through variables ($(SEGMENT_BUILD_DIR)/glm.o:,$(RANSLIB):, the static pattern rule for the V4 units) by asking make for those variables.test_registry_engine_agreement.pyfollows the renamed tier variable, andtest_glm53_metal_source.pyno longer expectsbackend_metal.hin the glm53 rule.Against the five points in #1741
1. Multi-source recipes. Confirmed: a command compiling two
.cfiles writes a.dfor the last one only. Of the four you named,test_segment_runtime,test_segment_conformanceandtest_edge_runtimenow link their helpers as objects;dsv4_decode_loader_testis built by thedsv4-cuda-loader-testrun target, which rebuilds every time (point 3), so it keeps its list untouched. Asking make for every command it would run found five more of the$< other.cshape (test_xdna_qt_state,test_xdna_failure,test_qwen38_tokenizer,test_qwen38_native_weights,test_native_quant), handled the same way, and one in a tool of mine from #1261:xdna_physical_probecompiledbackend_xdna.cbeside its own source and now linkstests/backend_xdna_lane.o, which has exactly its flags. One more only appears withCUDA,CUDA_DLL,HIPorHIP_DLL:qwen36_tier.con the command line of qwen36, qwen38, eleven tests and a benchmark. The helper objects:A test asks make for the command of every rule in
AUTODEP_BINSand fails on any that compiles more than one unit or lacks-MMD. It does so in the default build and again in every build flavour the host accepts, because a header under#ifreaches a.donly through the compile that reads it: hereCUDA_DLL,HIP_DLL,VKandXDNAon Windows,CUDA,HIP(with a fixedHIP_ARCH),VKandXDNAon Linux, andMETALon macOS.2.
-MMD -MPin CFLAGS. Done. One thing the forced rebuild does not reach, though: only 6 of those 202 rules depend on.build-config(colibri, deepseek_v41, qwen36 and three objects). The other 196, six of the nine engines among them, would stay "up to date" after this change with no.d, and then miss every header edit. So each target also depends on its own.dthrough an empty rule: a missing.dmakes the target out of date, and the rebuild writes it. Measured by building ondev, then checking this branch out over the same tree:3. First slice. gcc and clang only. The nvcc recipes, the MSVC-hosted CUDA build, Metal and
Makefile.deepseek-v4.unitskeep their lists, as agreed; nvcc and hipcc receiveNVCCFLAGSandHIPCCFLAGS, neverCFLAGS, so nothing there writes a.d. Also keeping a list, each named in the Makefile:AUTODEP_ELSEWHERE: six gcc tests with an explicit list, the other option point 1 allows.test_deepseek_v4andtest_v4_serve_framingare built byMakefile.deepseek-v4, and the four segment/edge adapter tests compile three units in one command.AUTODEP_OWN_FLAGS: three rules undertests/that do not compile with$(CFLAGS), so they write no.d:bench_cuda_resident_batch(the GPU compiler) and two benchmarks pinned to their own flags, whose sources include no local header.fuzz-rans,dsv4-cuda-testanddsv4-cuda-loader-testare run targets whose name is not a file, so they rebuild every time and a list changes nothing there.A new rule under
tests/joins the generated set on its own; one that compiles more than one unit, or without$(CFLAGS), fails the one-unit test until it is listed above.4. macOS. Shown by this PR's CI with Apple clang and the make Xcode ships, with one thing the log does not print; see CI on this PR.
5.
clean.pyandtest_makefile_deps.py. Rule names are unchanged.test_glm53_metal_source.pyassertedbackend_metal.hinside the glm53 rule; that line now documents that the header comes fromglm53.d.What the hand lists were missing
Method. In fresh trees on both platforms, the tree was built with this branch, so every target in
AUTODEP_BINSthat was built has a.dnaming the headers its compile actually read: 224 targets on Linux, 226 on Windows. Besidesmake checkand every engine, the benchmarks, the XDNA probe and the segment and ownership objects were built too: 52 of 53 on Windows, wherebench_idotstops at its own#errorunless built withARCH=native, and 52 of 53 on Linux, where the XDNA probe compiles but links on Windows only. Both are the same ondev, and both compiles still wrote their.d, so both are compared. The seven new helper objects have no rule ondevto compare with. Not built in that configuration, so not compared: on Linuxtest_xdna_execution,test_xdna_failure,test_qwen38_tier_engine,test_e8x4g64_loader,backend_vulkan.o,backend_loader.o,backend_xdna.o,qwen36_tier.o; on Windowstest_qwen38_tier_engine,test_uring,test_e8x4g64_loader,backend_vulkan.o,backend_xdna.o,qwen36_tier.o. Each.dwas compared with that target's prerequisite list indev's Makefile at499757d: headers only, paths normalised sotests/../st.handst.hcount as the same file.Result.
devomitsThe most frequent, identical on both:
decode_batch.h(102),pin_pool.h(94),cli_args.h(91),tok_unicode_o200k.h(82),omp_tune.h(77) targets. The platforms differ only where the builds do:test_uringis built on Linux only andtest_xdna_failureis built on Windows only, anduring.his read on Linux only. The previoustest_makefile_deps.pyalready named some of these in a docstring and left them for a separate fix. Two of them are still there at499757d(tok_unicode_o200k.hin colibri and deepseek_v41,decode_batch.hin five engines); the other two (sse41_kernels.h,edge_adapter_internal.h) have since been added by hand.One gap is outside both default runs and is mine:
backend_xdna.o, built only withXDNA=1, includescompat.hunconditionally. Its rule from #1261 ondevisbackend_xdna.o: backend_xdna.c backend_xdna.h .build-config; its.dfrom theXDNA=1build on Windows isbackend_xdna.o: backend_xdna.c backend_xdna.h compat.h.What that means in practice, shown on
devitself with the exit status ofmake -q(0: "up to date", nothing rebuilt; 1: rebuild needed), the same on both platforms:In the other build flavours. The same comparison after building everything this host builds with each flavour, against a baseline built the same way in the default configuration (231 targets compared on Windows, 153 with gaps; 228 and 152 on Linux). The last two columns are what the flavour adds: targets whose list on
devmisses a header only that flavour reads.XDNA=1backend_xdna.h(48)CUDA_DLL=1backend_cuda.h(51),backend_cuda_ink.h(4)HIP_DLL=1backend_cuda.h(51),backend_cuda_ink.h(4)VK=1backend_vulkan.h(57)CUDA=1backend_cuda.h(50),backend_cuda_ink.h(4)HIP=1In each GPU or NPU flavour, about fifty targets read that backend's own header without their list on
devnaming it, so editing it rebuilds nothing there. The flavour adds entries to targets that already had gaps, which is why the count with gaps does not grow.Limits of this measurement. Linux and Windows, the default configuration and the flavours above; not macOS, so not
METAL=1. A missing entry meansmakewill not rebuild after that header changes; whether the stale binary then behaves differently depends on the edit.All 150 targets, with the headers each reads but does not list on
devThe comparison script
Run on a tree built with this branch, with
dev's Makefile saved alongside:python3 gaps.py c/ dev.Makefile. It asks make forAUTODEP_BINS, so it compares exactly the targets this PR moves to generated dependencies, and for the variables rule names are spelled with.The new test_makefile_deps.py
AUTODEP_BINS; the Makefile's derivation matches the test's-MMD, default build-MMD, every build flavour#if: the flavours this host accepts.build-configalone.dcovering its includes#include "..."of its source is in its.dmake -q -W <header>, spelled as the.dspells itEach check was seen to fail before being trusted: the change on the left, made on its own, turns the test on the right red, on both platforms.
Testing
The CI jobs that build differently from
make check, run locally the way the workflow runs them:Windows HIP, built and run for real. No hosted runner has a Windows HIP toolchain or an AMD GPU, so no CI job covers this path (
docs/windows.mdsays so); it was run here, on the machine #788 was validated on, followingdocs/windows.md: TheRock HIP7.14.60850-2b22ab01, MSVC14.44.35207pinned,HIP_ARCH=gfx1151, in a fresh worktree of this branch.The two mismatches are in
devas well, on the CPU and on the GPU alike, so they do not come from this change;test_glm_oracle.pyallows two floating-point near ties in teacher forcing.Every number in this description comes from two scripted runs in fresh trees of this branch at
fbeef6aondev499757d, one per platform (the Windows one includes the HIP build), the same HIP script run once ondevfor comparison, the WindowsVK=1, LinuxCUDA=1and LinuxHIP=1builds, the every-header and per-flavour runs after CI, and reading the two Makefiles straight from git; the real-hardware HIP build is huppiflupp's. I can attach the scripts if that helps.CI on this PR
All 29 checks passed on the merge of
fbeef6aintodev4e28e39. The ones that close what this description listed as missing when it was opened:On macOS
$(MAKE)is/Applications/Xcode_26.6.app/Contents/Developer/usr/bin/make(check job, engines job). The Python suite runs quietly there, so the log does not nametest_makefile_deps, but its wiring and coverage tests cannot have been skipped: they skip only when make is missing or nothing inAUTODEP_BINSis built, and a failure to read the Makefile through-f -is an error, not a skip. So under that make the new Makefile lines and the-f -read work, and every target the job built had a.dcovering its source's includes, which is the question whether Apple clang writes<output>.dfor a one-command compile and link.backend_metal.ois built by clang++ without$(CFLAGS)and keeps its list, as planned.Built here after CI
Three build flavours no CI job builds or links through this Makefile, run on the branch at
fbeef6aafter the checks above:The HIP CI job's comment says a container cannot link the engine; with the 7.2.4 image it does, on
devas well.On real hardware. @huppiflupp built it on a Strix Halo under Linux (comment): Nobara 44, gfx1151, Fedora's ROCm 7.1 (hipcc 7.1.52802, rocwmma-devel 7.1.0),
make colibri HIP=1 ROCM_HOME=/usr HIP_ARCH=gfx1151atfbeef6a. It builds and linkslibamdhip64.so.7,colibri.dnamesbackend_cuda.h,make -q -W backend_cuda.h colibri …answers 1, andtest_makefile_depsis OK. Thank you.Every target, every header.
make -q -W <header> <target>for every built, up-to-date target and every header its.dnames:With
CUDA=1on Linux, 72 targets do not build, and the same 72 do not build ondev: most are tests and benchmarks that compilecolibri.cwith-DCOLI_CUDAbut do not link the CUDA object (undefined reference to coli_cuda_pipe_scratch); olmoe's link keeps-lcublaswithout the CUDA library path;backend_vulkan.oneeds Vulkan headers that environment does not have.One thing found on the way, on
devas well and not touched here: under MSYS2 the Makefile's-lvulkandoes not link (cannot find -lvulkan). The loader shipslibvulkan-1.dll.a, and the Windows link is-static, which makes ld skip import libraries for-l. The build above links that import library by path.Still not verified
All on macOS, which I do not have:
make --versionline in the macOS job, if you want it pinnedMETAL=1itself is built for real byAll engines (macos)METAL=1'sNot listed, because this change cannot affect it: CUDA or HIP kernels at run time on NVIDIA or Linux AMD hardware. The change decides which files make rebuilds, not what they compile to; the kernels that did run (HIP on gfx1151, Vulkan on the Radeon, the CPU path of the CUDA build) match
dev.On timing: as you said, this will conflict once with open PRs that edit the Makefile. Where it does, the conflict is on a rule's prerequisite line and the resolution is the same each time: take the incoming line without its headers. A new rule that arrives with a header list fails
test_no_generated_rule_lists_headers_by_hand, which names it.