codegen: support grid-constant kernel parameters - #1225
Conversation
|
The generic-kernel path needs to be part of the design before this lands:
#[kernel]
pub fn k<T: Copy>(#[grid_constant] d: &Desc, out: *mut u32, _tag: T) {
unsafe { *out = d.values[3] }
} fails with kernel contains duplicate grid-constant marker for source
other(d: &Desc, out: *mut u32) {
k::<u8>(d, out, 0)
} with no attribute comes out with
And we need tests in the PR exercise a generic kernel with To land this: the strip in both generic paths, a generic |
nihalpasham
left a comment
There was a problem hiding this comment.
Changes requested, details in #1225 (comment)
|
Addressed the generic-helper ABI concern in f2b6b15. The macro now captures the grid-constant marker for the generated entry and strips it from the re-emitted callable helper in both generic expansion paths. The contract example now covers both sides:
Measured validation:
|
nihalpasham
left a comment
There was a problem hiding this comment.
Reviewed f2b6b150: changes requested.
- The generic marker repair is present. The legacy NVVM path still loses the descriptor's pointee type:
host payload: 128 bytes -> i8* byval -> PTX parameter: 1 byte
- Preserve the aggregate type through the legacy signature and body references, and add a real legacy NVVM control. The modern path passing does not cover this.
- Define the read-only/lifetime contract across entry routes. Give the generic GPU test one writer; it currently has 32 threads writing one element.
- Sign the revised stack and restore its missing DCO trailers.
Independent review and libNVVM output confirm the ABI blocker. #1223 remains open.
Let #[kernel] declarations mark immutable reference parameters as grid constants, derive the generated host launch ABI from the same declaration, and carry pointee layout through MIR and LLVM lowering. Emit LLVM byval alignment and NVVM grid_constant metadata, with compiler and SM120 contract coverage. Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
Convert generic destinations explicitly with cvta.to.shared::cluster in inline PTX instead of constructing AS7, which legacy NVVM does not support. Retain the LLVM intrinsic path and test all G2S ranks and multicast forms. Signed-off-by: nihalp <nihalp@nvidia.com>
Retain complete typed byval pointees in legacy NVVM signatures, symbol references and entry adapters. Share physical parameter mapping with reference validity and reject malformed metadata and target-dependent storage layouts. Check monomorphized Freeze and launch-scoped reference lifetimes, keep ABI markers out of callable helpers, and require unsafe use of the hidden ABI marker. Exercise mixed parameter shapes, preserved helper calls, real TMA descriptors and rejection paths across the native backends. Signed-off-by: nihalp <nihalp@nvidia.com>
Reject an ordinary device-extern declaration that would suppress a grid-constant kernel declaration after erased-pointer shape matching. Cover both modern and legacy public exporter paths. Signed-off-by: nihalp <nihalp@nvidia.com>
f2b6b15 to
7853230
Compare
Copy and immutable parameter storage do not prove device validity, lifetime or synchronization of nested references. Keep that obligation explicit at every grid host launch boundary, including prepared sync, borrowed async and owned async methods. Preserve device function signatures and scalar marshalling. Add compile-fail and positive controls, generated safety documentation, and explicit contracts at the example call sites. Signed-off-by: nihalp <nihalp@nvidia.com>
|
I added the needed follow-ups to the descriptor-by-value support:
At This addresses #1223 for the native backend. CUTLASS integration is separate; no speedup has been measured yet. |
What this adds
TMA descriptors now travel with the kernel launch, removing their separate GPU allocation/upload in
tma_copy.The host passes a
Copydescriptor by value; device threads borrow one read-only copy shared by the grid for that launch.What we fixed
Launch contract
unsafe, including prepared/async calls, numeric payloads and kernels declared safe.Implementation
byvaland NVVMgrid_constantmetadata. Legacy declarations, symbol references and entry adapters retain the complete pointee type.cvta.to.shared::cluster, avoiding unsupported LLVM address-space handling while preserving cluster semantics.Verification
Checked at
f27fd082:just check: 5,393 tests/doctests, formatting, strict Clippy, guards and docs. Book passes with warnings as errors.tma_copymemcheck/synccheck on RTX 5090 through LLVM NVPTX, modern libNVVM and legacy libNVVM: 4,096 copied values and all 256 pipeline threads verified.Caveats
&Tor&'_ Tfor the parameter; named or'staticouter lifetimes are rejected.KernelScalar: Copylaunches retain an existing nested-reference safety gap; this PR does not redesign that API.Closes #1223.