Repository navigation
feat: add Core ML, Core AI and ExecuTorch Core ML runtimes for the Mac - #19
Open
leeclemnet wants to merge 2 commits into
Open
leeclemnet wants to merge 2 commits into
leeclemnet wants to merge 2 commits into
Conversation
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
7 tasks done
leeclemnet
force-pushed
the
feat/executorch-runtime
branch
from
September 30, 2026 16:34
82d725a to
068f8a6
Compare
leeclemnet
force-pushed
the
feat/mac-runtimes
branch
from
September 30, 2026 16:34
5ae2510 to
b35c431
Compare
leeclemnet
force-pushed
the
feat/executorch-runtime
branch
from
September 30, 2026 16:53
068f8a6 to
56ecd8f
Compare
leeclemnet
force-pushed
the
feat/mac-runtimes
branch
from
September 30, 2026 16:53
b35c431 to
c1b22fb
Compare
- CoreMLRuntime and CoreAIRuntime run on cpu, gpu and npu. - ExecuTorchRuntime gets the npu device for Core ML programs. - ThermalStateMonitor reads the macOS thermal state during each run. - New coreml, coreai and mac extras. The scripts get no new rows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Remove coreai from the mac extra: coreai-core has wheels only for macOS 26+. - Pin coreai-core to 1.0.0b2, and close the Core AI event loop. - Return False when the private ExecuTorch backend API is missing. - Add busy() to the macOS thermal monitor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
leeclemnet
force-pushed
the
feat/executorch-runtime
branch
from
September 30, 2026 19:35
56ecd8f to
e0c7228
Compare
leeclemnet
force-pushed
the
feat/mac-runtimes
branch
from
September 30, 2026 19:35
c1b22fb to
03c0492
Compare
leeclemnet
marked this pull request as ready for review
September 30, 2026 19:37
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Mac runtimes
What it does: Adds
CoreMLRuntimeandCoreAIRuntime(devicescpu,gpu,npu), thenpudevice for Core ML programs inExecuTorchRuntime, and a macOS throttle monitor. No script rows use the new runtimes.Why: SAB can now benchmark the Apple formats that rfdetr and Ultralytics both export. The benchmark scripts keep rows only for the
.onnxartifacts in the bucket. A throwaway script outside the repo exports and benchmarks the other formats, so the bucket holds no confusing artifacts. Stacked on #18.Change map
sab/runtimes/coreml.py) — loads.mlpackage/.mlmodelwithCPU_ONLY,CPU_AND_GPUorCPU_AND_NE. The timed call isMLModel.predict(), which includes the conversion toMLMultiArray. Image-type inputs (Ultralytics exports) take 0-255 pixels, so their rows setnormalized_in_graph.sab/runtimes/coreai.py) — loads an.aimodelwithcoreai.runtime(macOS 27).cpuallows only the CPU.gpuandnpuset the preferred compute unit. The timed call is one call of the inference function.sab/runtimes/executorch.py) — devicenpuruns.pteprograms that delegate to Core ML. Each device needs its backend in the installed ExecuTorch:XnnpackBackendforcpu,CoreMLBackendfornpu.sab/monitors/macos.py) —ThermalStateMonitorpollsProcessInfo.thermalStatethrough ctypes.fairor higher counts as throttling.select_monitoruses it for every device on macOS.tests/data/coreai, 36 KB), because the converter (coreai-torch) is not a SAB dependency.pyproject.toml,uv.lock,.github/workflows/tests.yml,README.md) — extrascoreml,coreaiandmac.macleaves outcoreai, because thecoreai-corewheels need macOS 26+ and the runtime needs macOS 27. On macOS 27, install--extra mac --extra coreai. CI does not installcoreai, so the Core AI tests skip there.Behavior changes
ExecuTorchRuntime.is_available("cpu")now also needsXnnpackBackendin the installed ExecuTorch.Watch out for
sab/runtimes/executorch.py:32— the backend check usesportable_lib._get_registered_backend_names(), a private ExecuTorch function. ExecuTorch has no public one. If the function is missing,is_availablereturns False.sab/runtimes/coreai.py:51— the Core AI Python API is async. One event loop serves every call, so the timed call does not start a loop.pyproject.toml:36—coreai-coreis a pre-release, pinned to 1.0.0b2 until its API is stable.npulabel is not reliable for fp32 artifacts. Core ML fp32 onnpuruns on the CPU, because the Neural Engine runs only fp16. Core AI fp32 onnpuis as fast as thegpurow. The Python API does not show which unit ran the model.Related Issue(s): n/a
Type of Change
Testing
--extra mac --extra coreai.Apple M4 Max, macOS 27 (on battery). The artifacts are the official exports in
gs://rfdetr/export-2026-09-30. The.ptefp16 programs delegate to Core ML. ORT reads the.onnxof the same export. YOLO26 has no official exports, so it has no rows.rfdetr-nanorfdetr-smallrfdetr-mediumrfdetr-largerfdetr-xlargerfdetr-2xlargenpuloses 1.9–4.5 mAP on RF-DETR. The same models in ExecuTorch Core ML onnpustay within 0.3 mAP of ORT.npuruns on the CPU: the Neural Engine runs only fp16. Core AI fp32 onnpuis as fast as thegpurow.cpuloses 0.5–1.2 mAP on xlarge and 2xlarge. The same artifacts keep their mAP ongpuandnpu.fairor higher).Checklist
Additional Context
🤖 Generated with Claude Code