Skip to content

Commit 4694e29

Browse files
committed
numerical-acceleration 후속 심화: Python numpy->GPU 직결 + GpuArray.map 잔류 체이닝
GPU를 파이썬에 실제로 연결 = pyproc 정체성 완성. 실 GPU(창 모드) 실측. GpuArray.map(expr): WGSL 표현식(x=원소)을 각 원소에 적용한 새 잔류 핸들. matmul 뒤 활성화 (max(x,0) relu 등)를 리드백 없이 잇는다. 표현식별 파이프라인 캐시. Runtime.enableGpu() -> GpuBridge(Python numpy 직결): install()이 GPU 디바이스 확보 + 파이썬 pyprocGpu 모듈 배선. pyprocGpu.matmul(a, b)가 numpy 배열을 f32로 GPU에서 곱해 numpy로 반환 (블로킹 = JSPI run_sync, socketBridge 패턴). numpy + 실 GPU + 창 모드 필요. gpuPythonProbe(실 GPU GREEN 4/4): 파이썬 GPU matmul == CPU numpy maxerr 0.00, 1024 f32 92배(GPU 84ms vs CPU 7682ms). map 잔류 체이닝 matmul->relu == CPU 참조 maxErr 1.19e-7. 헤드리스는 SKIP(CI 무해). 커널 최적화 판정(정직): 연구가 커널 자작 금지(naive vs 최적 600-1000배 격차, jax-js/WgPy 차용 권장) 명시 + 이 GPU 하나에서만 검증 가능 -> naive 타일드(검증 92-109배) 유지, 프로덕션 타일링/차용은 후속. 위험한 자작 커널이 정공법 아니다. index.js/index.d.ts(GpuBridge/map/enableGpu)/run.mjs/README 2종 + 이니셔티브 원장 완결. 구조 455 + 코어 40/40 + wasiGate 10/10 + 예제 5/5 + 실 GPU 3종(4/4,5/5,4/4) 전부 green.
1 parent 15a4792 commit 4694e29

11 files changed

Lines changed: 221 additions & 9 deletions

File tree

‎README.ko.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -202,7 +202,7 @@ Pyodide Workers
202202
| `Init` | OS init: `/home/web/boot.py` 오토런 + `cron.py` 틱, 전부 파일 구동 |
203203
| `MachineJournal` | WAL: 유휴에 스스로 체크포인트해, 강제종료된 탭도 마지막 커밋으로 부활 |
204204
| `MachineJail` | 권한 감옥: `permissions{net, clipboard, home, workers}`를 2단 집행. 협조 파이썬 초크포인트 + 브라우저 벽(감옥 컨텍스트의 `connect-src` CSP가 비허용 host 차단, 감옥 코드가 `import js`로 우회해도 무력) |
205-
| `GpuCompute` / `GpuArray` | f32 대규모 선형대수를 WebGPU 컴퓨트로 오프로드: 잔류 핸들(업로드 1회, GPU 위에서 `matmul` 체이닝, 다운로드 1회). 실 GPU에서 WASM numpy 대비 109배 실측. f32 한정(WGSL은 f64 없음), 창 있는 브라우저 + GPU 필요 |
205+
| `GpuCompute` / `GpuArray` / `GpuBridge` | f32 대규모 선형대수를 WebGPU 컴퓨트로 오프로드: 잔류 핸들(업로드 1회, GPU 위에서 `matmul` / `map` 체이닝, 다운로드 1회). `Runtime.enableGpu()`가 파이썬에 배선(`pyprocGpu.matmul`이 numpy 배열을). 실 GPU에서 WASM numpy 대비 109배 실측. f32 한정(WGSL은 f64 없음), 창 있는 브라우저 + GPU 필요 |
206206
| `bootSession` / `Session` / `openMachine` | 세션 부활 + 이동 가능한 `.pymachine` 이미지: 결정적 리플레이 + 사용자 델타, OPFS 영속(`save` / `load`) 또는 한 파일 내보내기(`exportImage` / `openMachine`) |
207207
| `WheelCache` | 오프라인 / 재다운로드 0 패키지 설치용 wheel / OPFS 캐시 |
208208
| `PyProc` | 프로세스 OS 커널: 스냅샷-fork 스폰, `map` / `mapArray` 병렬, 수명주기(`kill` / `signal` / respawn), `fork(2)`(살아있는 프로세스 복제, 변수·배열이 실린다), 흐름 IPC(`pipe` / `lock` / `semaphore` / `shm`: SAB 링버퍼 파이프, 진짜 블로킹 read + backpressure) |

‎README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -202,7 +202,7 @@ Capabilities are opt-in - turn on only what you need, and consume the capability
202202
| `Init` | OS init: `/home/web/boot.py` autorun plus `cron.py` ticks, all file-driven |
203203
| `MachineJournal` | Write-ahead log: the machine checkpoints itself while idle, so a crashed tab still boots back into its last commit |
204204
| `MachineJail` | Permission jail: `permissions{net, clipboard, home, workers}` enforced in two tiers, a cooperative Python chokepoint plus the browser's own wall (a `connect-src` CSP on the jail context blocks disallowed hosts even if the jailed code tries `import js`) |
205-
| `GpuCompute` / `GpuArray` | Offload large f32 linear algebra to WebGPU compute: a residency handle (upload once, chain `matmul` on the GPU, download once). Measured 109x vs WASM numpy on a real GPU; f32 only (WGSL has no f64), needs a windowed browser with a GPU |
205+
| `GpuCompute` / `GpuArray` / `GpuBridge` | Offload large f32 linear algebra to WebGPU compute: a residency handle (upload once, chain `matmul` / `map` on the GPU, download once). `Runtime.enableGpu()` wires it into Python (`pyprocGpu.matmul` on numpy arrays). Measured 109x vs WASM numpy on a real GPU; f32 only (WGSL has no f64), needs a windowed browser with a GPU |
206206
| `bootSession` / `Session` / `openMachine` | Session revival and portable `.pymachine` images: deterministic replay plus user delta, persisted to OPFS (`save` / `load`) or exported as one file (`exportImage` / `openMachine`) |
207207
| `WheelCache` | Wheel / OPFS cache for offline, zero-redownload package installs |
208208
| `PyProc` | Process OS kernel: snapshot-fork spawn, `map` / `mapArray` parallelism, lifecycle (`kill` / `signal` / respawn), `fork(2)` (clone a live process, its variables and arrays travel), and flow IPC (`pipe` / `lock` / `semaphore` / `shm`: SAB ring-buffer pipes with real blocking read and backpressure) |

‎index.d.ts‎

Lines changed: 14 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -569,11 +569,23 @@ export class GpuArray {
569569
readonly cols: number;
570570
/** 이 배열(M x K) @ other(K x N) = 새 잔류 핸들(M x N). 재업로드 0. */
571571
matmul(other: GpuArray): GpuArray;
572+
/** 원소별 변환(WGSL 표현식, x = 원소)을 적용한 새 잔류 핸들(같은 shape). 예: map("max(x, 0.0)"). matmul 뒤 활성화 체이닝. */
573+
map(expr: string): GpuArray;
572574
/** GPU -> CPU 회수. 반환 { data: Float32Array, rows, cols }. */
573575
toArray(): Promise<{ data: Float32Array; rows: number; cols: number }>;
574576
destroy(): void;
575577
}
576578

579+
/**
580+
* Python numpy -> GPU 직결. Runtime.enableGpu()로 얻고 install() 후 파이썬이 pyprocGpu.matmul(a, b)로
581+
* numpy 배열을 GPU에서 곱한다(블로킹 = JSPI, rt.runAsync 경로). 실 GPU + 창 모드 + numpy 필요.
582+
* f64는 f32로 강등(WGSL 한계, 정밀도 손실은 계약).
583+
*/
584+
export class GpuBridge {
585+
install(): Promise<{ installed: string; note: string }>;
586+
destroy(): void;
587+
}
588+
577589
/**
578590
* WebGPU 컴퓨트로 f32 대규모 선형대수 가속(수치 성능 도약 Phase 2). numpy 대체가 아니라 좁은
579591
* 고피크 레인: matmul 실측 109배 vs WASM numpy(실 GPU). 잔류 핸들(업로드1/체이닝/다운로드1)이
@@ -641,6 +653,8 @@ export class Runtime {
641653
enableDeviceFs(cfg?: DeviceFsConfig): DeviceFs;
642654
enableInit(cfg?: InitConfig): Init;
643655
enableJournal(cfg: JournalConfig): MachineJournal;
656+
/** Python numpy -> GPU 직결(install()로 pyprocGpu 배선). 실 GPU + 창 모드 + numpy 필요. */
657+
enableGpu(cfg?: object): GpuBridge;
644658
/** 디렉터리 핸들(OPFS 등)을 파이썬 경로로 마운트(기본 /home/web). 반환된 sync()로 영속화. */
645659
mountHome(dirHandle: FileSystemDirectoryHandle, path?: string): Promise<{ path: string; sync: () => Promise<void> }>;
646660
/** 탈출구(권장 안 함): 내부 Pyodide 인스턴스. */

‎index.js‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -52,7 +52,7 @@ export { MachineJail } from "./src/capabilities/machineJail.js";
5252
export { bootSession, openMachine, Session } from "./src/capabilities/session.js";
5353
export { WheelCache } from "./src/capabilities/wheelCache.js";
5454
export { bootEnv, runScript } from "./src/capabilities/envManager.js";
55-
export { GpuCompute, GpuArray } from "./src/capabilities/gpuCompute.js";
55+
export { GpuCompute, GpuArray, GpuBridge } from "./src/capabilities/gpuCompute.js";
5656
export { PyProc, SIGNAL } from "./src/processOs/pyProc.js";
5757
export { MachineContainer } from "./src/processOs/machineContainer.js";
5858
export { JobControl } from "./src/processOs/jobControl.js";

‎mainPlan/numerical-acceleration/03-progress-ledger.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -4,6 +4,15 @@
44

55
## 결정 원장 (최신이 위)
66

7+
### 2026-07-13 후속 심화 완료: Python numpy -> GPU 직결 + map 체이닝 (gpuPythonProbe 실 GPU 4/4)
8+
9+
- Phase 2의 후속 심화(파이썬 통합 + 원소별 op)를 실 GPU로 완성했다. **GPU가 파이썬에 실제로 연결됐다 = pyproc 정체성 완성.**
10+
- **`GpuArray.map(expr)`(원소별 잔류)**: WGSL 표현식(x=원소)을 각 원소에 적용한 새 잔류 핸들. matmul 뒤 활성화(`max(x,0)` relu 등)를 **리드백 없이** 잇는다. 표현식별 파이프라인 캐시.
11+
- **`Runtime.enableGpu()` -> `GpuBridge`(Python numpy 직결)**: install()이 GPU 디바이스 확보 + 파이썬 `pyprocGpu` 모듈 배선. `pyprocGpu.matmul(a, b)`가 numpy 배열을 f32로 GPU에서 곱해 numpy로 반환(블로킹 = JSPI run_sync, socketBridge 패턴). numpy 필요, 실 GPU + 창 모드.
12+
- **실측(gpuPythonProbe GREEN 4/4, 실 GPU)**: 파이썬 GPU matmul == CPU numpy **maxerr 0.00**, 1024 f32 **92배**(GPU 84ms vs CPU 7682ms). map 잔류 체이닝 matmul->relu == CPU 참조 maxErr 1.19e-7. 헤드리스는 SKIP.
13+
- **커널 최적화 판정(정직)**: 연구가 "커널 자작 금지"(naive vs 최적 600-1000배 격차, jax-js/WgPy 차용 권장)를 명시했고 이 GPU 하나에서만 검증 가능하므로, **naive 타일드(검증된 92-109배)를 유지하고 프로덕션 타일링/차용은 후속으로 정직하게 둔다**. 위험한 자작 커널이 정공법이 아니다.
14+
- **표면**: index.js/index.d.ts(GpuBridge/map/enableGpu)/run.mjs/README 2종. **Phase 2 = 개념 + src 승격 + 파이썬 통합 + 원소별 op까지 완결.** 잔여(커널 최적화 차용, GPU reduce, worker 내 GPU)는 코어 밖 선택 후속.
15+
716
### 2026-07-13 Phase 2 완료 + src 승격: GpuCompute WebGPU 잔류 핸들 (실 GPU gpuMatmul 4/4 + gpuSurface 5/5)
817

918
- Phase 2(프론티어, GPU)를 실 GPU로 실증하고 `GpuCompute`/`GpuArray`로 src 승격했다. **"GPU 검증 불가"는 헤드리스 한정**이었음이 드러났다.

‎mainPlan/numerical-acceleration/README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
# numerical-acceleration - 수치 성능 도약 (브라우저 파이썬의 마지막 큰 격차)
22

3-
상태: **Phase 1 + Phase 2 실증·src 승격 완료 (2026-07-13).** browser-os P1~P7 + engine-independence 사다리가 닫힌 뒤 개시. **수치 연산 속도**(numpy 86배)를 두 레인으로 뚫었다: (1) CPU 샤딩 `PyProc.matmul`(compute-bound near-linear, 종단 2.48배) (2) **WebGPU 잔류 핸들 `GpuCompute`**(f32 대규모 matmul **실 GPU 109배**). 잔여(worker+JSPI 통합, 커널 최적화, op 확장)는 코어 밖 후속 심화 = [02-phasing NEXT](02-phasing-and-wiring.md).
3+
상태: **Phase 1 + Phase 2 + 후속 심화 완료 (2026-07-13).** browser-os P1~P7 + engine-independence 사다리가 닫힌 뒤 개시. **수치 연산 속도**(numpy 86배)를 두 레인으로 뚫었다: (1) CPU 샤딩 `PyProc.matmul`(compute-bound near-linear, 종단 2.48배) (2) **WebGPU 잔류 핸들 `GpuCompute`/`GpuArray`**(f32 대규모 matmul **실 GPU 109배**, `map`으로 활성화 체이닝) + **`Runtime.enableGpu()`로 Python numpy 직결**(`pyprocGpu.matmul`이 numpy 배열을 GPU에서, **92배**). 잔여(커널 최적화 차용, GPU reduce, worker 내 GPU)는 코어 밖 선택 후속 = [02-phasing NEXT](02-phasing-and-wiring.md).
44

55
## 한 문장
66

‎src/capabilities/gpuCompute.js‎

Lines changed: 92 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -28,8 +28,22 @@ fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
2828
c[row * d.n + col] = sum;
2929
}`;
3030

31+
// 원소별 WGSL 템플릿(EXPR = 소비자 표현식, x = 원소). matmul 뒤 활성화 등 잔류 체이닝용.
32+
// 예: map("max(x, 0.0)")(relu), map("x * 2.0 + 1.0"), map("1.0 / (1.0 + exp(-x))")(sigmoid).
33+
const ELEMENTWISE_WGSL = `
34+
@group(0) @binding(0) var<storage, read> a: array<f32>;
35+
@group(0) @binding(1) var<storage, read_write> c: array<f32>;
36+
@group(0) @binding(2) var<uniform> len: u32;
37+
@compute @workgroup_size(64)
38+
fn main(@builtin(global_invocation_id) gid: vec3<u32>) {
39+
let i = gid.x;
40+
if (i >= len) { return; }
41+
let x = a[i];
42+
c[i] = __EXPR__;
43+
}`;
44+
3145
export class GpuCompute {
32-
constructor(device) { this._device = device; this._matmul = null; }
46+
constructor(device) { this._device = device; this._matmul = null; this._elementwise = new Map(); }
3347

3448
// WebGPU 디바이스를 확보한다(async). 어댑터가 없으면(헤드리스) 실행 가능한 에러.
3549
static async create() {
@@ -54,6 +68,17 @@ export class GpuCompute {
5468
return this._matmul;
5569
}
5670

71+
// 원소별 파이프라인(표현식별 캐시). expr는 소비자 WGSL 표현식(x = 원소).
72+
_elementwisePipeline(expr) {
73+
let p = this._elementwise.get(expr);
74+
if (!p) {
75+
const module = this._device.createShaderModule({ code: ELEMENTWISE_WGSL.replace("__EXPR__", expr) });
76+
p = this._device.createComputePipeline({ layout: "auto", compute: { module, entryPoint: "main" } });
77+
this._elementwise.set(expr, p);
78+
}
79+
return p;
80+
}
81+
5782
// f32 배열을 GPU에 올린다(잔류 시작). data = Float32Array(길이 rows*cols). 반환 = 잔류 핸들.
5883
array(data, rows, cols) {
5984
if (!(data instanceof Float32Array)) throw new Error("gpuArray: data는 Float32Array다(WGSL은 f64 없음 = f32만).");
@@ -92,6 +117,26 @@ export class GpuArray {
92117
return new GpuArray(this._gc, cBuf, M, N);
93118
}
94119

120+
// 원소별 변환: 각 원소 x에 WGSL 표현식 expr를 적용한 새 잔류 핸들(같은 shape). 재업로드 0.
121+
// 잔류 체이닝의 핵심: m.matmul(w).map("max(x, 0.0)")처럼 matmul 뒤 활성화를 리드백 없이 잇는다.
122+
map(expr) {
123+
if (typeof expr !== "string" || !expr.length) throw new Error("GpuArray.map: expr는 WGSL 표현식 문자열(x = 원소). 예: \"max(x, 0.0)\"");
124+
const device = this._gc._device, len = this.rows * this.cols;
125+
const cBuf = device.createBuffer({ size: len * 4, usage: GPUBufferUsage.STORAGE | GPUBufferUsage.COPY_SRC });
126+
const nBuf = device.createBuffer({ size: 16, usage: GPUBufferUsage.UNIFORM, mappedAtCreation: true });
127+
new Uint32Array(nBuf.getMappedRange()).set([len, 0, 0, 0]); nBuf.unmap();
128+
const pipeline = this._gc._elementwisePipeline(expr);
129+
const bind = device.createBindGroup({ layout: pipeline.getBindGroupLayout(0), entries: [
130+
{ binding: 0, resource: { buffer: this.buffer } }, { binding: 1, resource: { buffer: cBuf } }, { binding: 2, resource: { buffer: nBuf } },
131+
] });
132+
const enc = device.createCommandEncoder();
133+
const pass = enc.beginComputePass(); pass.setPipeline(pipeline); pass.setBindGroup(0, bind);
134+
pass.dispatchWorkgroups(Math.ceil(len / 64)); pass.end();
135+
device.queue.submit([enc.finish()]);
136+
nBuf.destroy();
137+
return new GpuArray(this._gc, cBuf, this.rows, this.cols);
138+
}
139+
95140
// GPU -> CPU 회수(리드백 1복사). 반환 { data: Float32Array, rows, cols }.
96141
async toArray() {
97142
const device = this._gc._device, size = this.rows * this.cols * 4;
@@ -107,3 +152,49 @@ export class GpuArray {
107152

108153
destroy() { this.buffer.destroy(); }
109154
}
155+
156+
// Python numpy -> GPU 직결(pyproc 정체성 완성). Runtime의 파이썬이 numpy 배열을 f32로 GPU에서
157+
// matmul한다: pyprocGpu.matmul(a, b). 블로킹은 JSPI(run_sync)라 rt.runAsync 경로에서 동작한다
158+
// (socketBridge/machineContainer와 같은 패턴). 실 GPU + 창 모드 필요(navigator.gpu는 메인 스레드).
159+
// numpy 필요(rt.loadPackages(["numpy"])). f64는 f32로 강등(WGSL 한계) - 정밀도 손실은 계약이다.
160+
const GPU_BOOTSTRAP = `
161+
import sys as _pyprocSysG, types as _pyprocTypesG
162+
import numpy as _pyprocNumpyG
163+
from pyodide.ffi import to_js as _pyprocToJsG, run_sync as _pyprocRunSyncG
164+
165+
_pyprocGpuMod = _pyprocTypesG.ModuleType('pyprocGpu')
166+
167+
def _pyprocGpuMatmul(a, b):
168+
a = _pyprocNumpyG.ascontiguousarray(a, dtype=_pyprocNumpyG.float32)
169+
b = _pyprocNumpyG.ascontiguousarray(b, dtype=_pyprocNumpyG.float32)
170+
res = _pyprocRunSyncG(_pyprocGpuMatmulBridge(
171+
_pyprocToJsG(a.tobytes()), a.shape[0], a.shape[1],
172+
_pyprocToJsG(b.tobytes()), b.shape[0], b.shape[1]))
173+
return _pyprocNumpyG.frombuffer(bytes(res.to_py()), dtype=_pyprocNumpyG.float32).reshape(a.shape[0], b.shape[1])
174+
175+
_pyprocGpuMod.matmul = _pyprocGpuMatmul
176+
_pyprocSysG.modules['pyprocGpu'] = _pyprocGpuMod
177+
`;
178+
179+
export class GpuBridge {
180+
constructor(rt) { this._rt = rt; this._gc = null; }
181+
182+
// GPU 디바이스 확보 + 파이썬 pyprocGpu 모듈 배선. 어댑터 부재(헤드리스) 시 실행 가능한 에러.
183+
async install() {
184+
this._gc = await GpuCompute.create();
185+
const gc = this._gc;
186+
// 파이썬이 부를 브리지(JSPI가 서스펜드하는 async): f32 바이트를 받아 GPU matmul 후 결과 바이트.
187+
const bridge = async (aU8, aRows, aCols, bU8, bRows, bCols) => {
188+
const A = new Float32Array(aU8.slice().buffer), B = new Float32Array(bU8.slice().buffer);
189+
const ga = gc.array(A, aRows, aCols), gb = gc.array(B, bRows, bCols), gout = ga.matmul(gb);
190+
const r = await gout.toArray();
191+
ga.destroy(); gb.destroy(); gout.destroy();
192+
return new Uint8Array(r.data.buffer);
193+
};
194+
this._rt.setGlobal("_pyprocGpuMatmulBridge", bridge);
195+
this._rt.run(GPU_BOOTSTRAP);
196+
return { installed: "pyprocGpu", note: "블로킹은 JSPI(run_sync)라 rt.runAsync 경로에서. numpy 필요" };
197+
}
198+
199+
destroy() { if (this._gc) this._gc.destroy(); }
200+
}

‎src/runtime/runtime.js‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -13,6 +13,7 @@ import { Terminal } from "../capabilities/terminal.js";
1313
import { DeviceFs } from "../capabilities/deviceFs.js";
1414
import { Init } from "../capabilities/init.js";
1515
import { MachineJournal } from "../capabilities/machineJournal.js";
16+
import { GpuBridge } from "../capabilities/gpuCompute.js";
1617

1718
export { MemoryCapability, PAGE_SIZE } from "./memoryCapability.js";
1819
export { checkEnvironment } from "./preflight.js";
@@ -137,6 +138,8 @@ export class Runtime {
137138
enableDeviceFs(cfg = {}) { return new DeviceFs(this, cfg); }
138139
enableInit(cfg = {}) { return new Init(this, cfg); }
139140
enableJournal(cfg = {}) { return new MachineJournal(this, cfg); }
141+
// Python numpy -> GPU 직결(install()로 pyprocGpu 모듈 배선). 실 GPU + 창 모드 + numpy 필요.
142+
enableGpu(cfg = {}) { return new GpuBridge(this); }
140143

141144
// 영속 디스크: OPFS 등 디렉터리 핸들을 파이썬 파일시스템 경로로 마운트한다.
142145
// 파이썬 open()이 진짜 지속 파일을 읽고 쓴다. 변경 반영은 반환된 sync() 호출(핸들은 소비자 제공).

‎tests/attempts/gpuCompute/README.md‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -21,6 +21,8 @@ PYPROC_HEADED=1 node tests/browser/run.mjs tests/attempts/gpuCompute/gpuMatmulPr
2121
| 질문 | probe | 게이트 |
2222
|---|---|---|
2323
| WGSL matmul이 정확하고 numpy보다 빠른가 | [gpuMatmulProbe.html](gpuMatmulProbe.html) | (실 GPU) 결과 == CPU 참조(f32 허용오차) + GPU 종단이 WASM numpy 대비 >= 10배. 어댑터 없으면 SKIP |
24+
| 승격 계약(GpuCompute/GpuArray 잔류)이 정확한가 | [gpuSurfaceProbe.html](gpuSurfaceProbe.html) | array->matmul->toArray == CPU 참조 + 잔류 체이닝 (A@B)@C 정확 + 차원 에러 |
25+
| Python numpy가 GPU를 쓰고 map 체이닝이 되나 | [gpuPythonProbe.html](gpuPythonProbe.html) | enableGpu -> pyprocGpu.matmul(numpy)==CPU numpy + 속도 + GpuArray.map(matmul->relu)==CPU 참조 |
2426

2527
승격 조건(G2 실 GPU GREEN): worker 소유 GPUDevice + JSPI 잔류 핸들로 `GpuCompute`/`gpuArray` 능력 승격. G2 실패(전송비 이득 초과) 시 examples/문서 패턴으로 강등(정직한 조건부).
2628

@@ -29,4 +31,5 @@ PYPROC_HEADED=1 node tests/browser/run.mjs tests/attempts/gpuCompute/gpuMatmulPr
2931
| 날짜 | probe | 환경 | 핵심 수치 | 결론 | 판정 |
3032
|---|---|---|---|---|---|
3133
| 2026-07-13 | gpuMatmulProbe | Edge **창 모드(실 GPU)** + 자가 호스팅 경로 | 정확성 GPU matmul == CPU 참조 **maxErr 3.58e-7**(f32), 대형 1024 f32 GPU 종단(업로드+연산+리드백) **65.9ms**, WASM numpy 단일워커 7221ms = **GPU 109.6배**. GREEN 4/4(헤드리스는 SKIP) | **Phase 2 개념 성립**: naive 타일드 WGSL matmul로도 109배(최적화 커널 WgPy 340배). f32 정밀도 정확, 종단 전송비 포함해도 압도. GPU 벽(f64 없음, 창 모드 필요)은 정직한 경계 | 승격 -> `GpuCompute`/`gpuArray` 잔류 핸들 능력 |
32-
| 2026-07-13 | gpuSurfaceProbe | Edge **창 모드(실 GPU)** | 승격 계약 `GpuCompute`/`GpuArray` 검증: create -> `array(f32)` -> `matmul` -> `toArray` == CPU 참조 **maxErr 2.38e-7**, **잔류 체이닝 (A@B)@C == 참조 maxErr 2.68e-7**(중간 리드백 0 = 재업로드 없음), 차원 불일치 명시적 에러, 대형 잔류 matmul **37.1ms**. GREEN 5/5(헤드리스 SKIP) | **Phase 2 src 승격 완료.** 잔류 핸들 모델(업로드1/GPU 체이닝/다운로드1)이 실 GPU에서 정확·동작. 셰이더 1회 컴파일 캐시 | 승격 -> `GpuCompute`/`GpuArray`. 후속: worker+JSPI 통합, 커널 최적화(타일링/차용), reduce/elementwise op |
34+
| 2026-07-13 | gpuSurfaceProbe | Edge **창 모드(실 GPU)** | 승격 계약 `GpuCompute`/`GpuArray` 검증: create -> `array(f32)` -> `matmul` -> `toArray` == CPU 참조 **maxErr 2.38e-7**, **잔류 체이닝 (A@B)@C == 참조 maxErr 2.68e-7**(중간 리드백 0 = 재업로드 없음), 차원 불일치 명시적 에러, 대형 잔류 matmul **37.1ms**. GREEN 5/5(헤드리스 SKIP) | **Phase 2 src 승격 완료.** 잔류 핸들 모델(업로드1/GPU 체이닝/다운로드1)이 실 GPU에서 정확·동작. 셰이더 1회 컴파일 캐시 | 승격 -> `GpuCompute`/`GpuArray` |
35+
| 2026-07-13 | gpuPythonProbe | Edge **창 모드(실 GPU)** + 자가 호스팅 | **Python numpy -> GPU 직결(pyproc 정체성 완성)** + map 체이닝. `Runtime.enableGpu().install()` -> 파이썬 `pyprocGpu.matmul(numpy a, b)`(JSPI run_sync)가 GPU에서 곱해 numpy로 반환 == CPU numpy **maxerr 0.00**, 1024 f32 **92배**(GPU 84ms vs CPU 7682ms). **GpuArray.map 잔류 체이닝**: matmul -> relu(`max(x,0)`) 리드백 없이 == CPU 참조 maxErr 1.19e-7. GREEN 4/4(헤드리스 SKIP) | **후속 심화 완료**: 파이썬이 GPU를 쓴다(numpy 배열 한 줄로 92배). map으로 matmul 뒤 활성화를 리드백 없이 잇는다 | 승격 -> `GpuBridge`(enableGpu) + `GpuArray.map`. 커널 최적화(자작 금지 = jax-js/WgPy 차용)는 후속 |

0 commit comments

Comments
 (0)