Skip to content

feat: add Core ML, Core AI and ExecuTorch Core ML runtimes for the Mac - #19

Open
leeclemnet wants to merge 2 commits into
feat/executorch-runtimefrom
feat/mac-runtimes
Open

leeclemnet wants to merge 2 commits into
feat/executorch-runtimefrom
feat/mac-runtimes

Conversation

@leeclemnet

@leeclemnet leeclemnet commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

Mac runtimes

What it does: Adds CoreMLRuntime and CoreAIRuntime (devices cpu, gpu, npu), the npu device for Core ML programs in ExecuTorchRuntime, and a macOS throttle monitor. No script rows use the new runtimes.
Why: SAB can now benchmark the Apple formats that rfdetr and Ultralytics both export. The benchmark scripts keep rows only for the .onnx artifacts in the bucket. A throwaway script outside the repo exports and benchmarks the other formats, so the bucket holds no confusing artifacts. Stacked on #18.

Change map

  • Core ML (sab/runtimes/coreml.py) — loads .mlpackage / .mlmodel with CPU_ONLY, CPU_AND_GPU or CPU_AND_NE. The timed call is MLModel.predict(), which includes the conversion to MLMultiArray. Image-type inputs (Ultralytics exports) take 0-255 pixels, so their rows set normalized_in_graph.
  • Core AI (sab/runtimes/coreai.py) — loads an .aimodel with coreai.runtime (macOS 27). cpu allows only the CPU. gpu and npu set the preferred compute unit. The timed call is one call of the inference function.
  • ExecuTorch (sab/runtimes/executorch.py) — device npu runs .pte programs that delegate to Core ML. Each device needs its backend in the installed ExecuTorch: XnnpackBackend for cpu, CoreMLBackend for npu.
  • Monitor (sab/monitors/macos.py) — ThermalStateMonitor polls ProcessInfo.thermalState through ctypes. fair or higher counts as throttling. select_monitor uses it for every device on macOS.
  • Tests — Core ML models are built in the tests with coremltools. Core AI uses three small committed fixtures (tests/data/coreai, 36 KB), because the converter (coreai-torch) is not a SAB dependency.
  • Packaging, CI and docs (pyproject.toml, uv.lock, .github/workflows/tests.yml, README.md) — extras coreml, coreai and mac. mac leaves out coreai, because the coreai-core wheels need macOS 26+ and the runtime needs macOS 27. On macOS 27, install --extra mac --extra coreai. CI does not install coreai, so the Core AI tests skip there.

Behavior changes

  • On macOS, rows record a throttle verdict. Before, the verdict was always unknown.
  • ExecuTorchRuntime.is_available("cpu") now also needs XnnpackBackend in the installed ExecuTorch.

Watch out for

  • sab/runtimes/executorch.py:32 — the backend check uses portable_lib._get_registered_backend_names(), a private ExecuTorch function. ExecuTorch has no public one. If the function is missing, is_available returns False.
  • sab/runtimes/coreai.py:51 — the Core AI Python API is async. One event loop serves every call, so the timed call does not start a loop.
  • pyproject.toml:36 — coreai-core is a pre-release, pinned to 1.0.0b2 until its API is stable.
  • The npu label is not reliable for fp32 artifacts. Core ML fp32 on npu runs on the CPU, because the Neural Engine runs only fp16. Core AI fp32 on npu is as fast as the gpu row. The Python API does not show which unit ran the model.

Related Issue(s): n/a

Type of Change

  • New feature (non-breaking change that adds functionality)

Testing

  • Unit tests. The Core AI tests ran on an Apple M4 Max with macOS 27, with --extra mac --extra coreai.
  • Core ML, Core AI and ExecuTorch Core ML vs ORT CPU on an Apple M4 Max (50 coco val2017 images)

Apple M4 Max, macOS 27 (on battery). The artifacts are the official exports in gs://rfdetr/export-2026-09-30. The .pte fp16 programs delegate to Core ML. ORT reads the .onnx of the same export. YOLO26 has no official exports, so it has no rows.

Model Format Precision mAP50:95 cpu mAP50:95 gpu mAP50:95 npu Median ms cpu Median ms gpu Median ms npu
rfdetr-nano ONNX Runtime fp32 53.9 – – 34.5 – –
Core ML fp32 53.9 53.9 53.9 23.7† 6.8† 23.9†
Core ML fp16 54.0 53.9 49.4 14.3† 6.3† 8.6†
Core AI fp32 53.9 53.9 53.9 26.0† 7.4† 6.9†
Core AI fp16 54.2 53.9 53.3 21.3† 6.9† 14.6†
ExecuTorch Core ML fp16 – – 54.0 – – 10.3
rfdetr-small ONNX Runtime fp32 58.3 – – 58.3 – –
Core ML fp32 58.3 58.3 58.3 40.1 9.5 40.1
Core ML fp16 58.3 57.5 55.0 25.4 8.6 17.8
Core AI fp32 58.3 58.3 58.3 43.6 9.6 9.8
Core AI fp16 58.0 57.9 57.7 35.6 8.8 25.7
ExecuTorch Core ML fp16 – – 58.0 – – 22.3
rfdetr-medium ONNX Runtime fp32 57.1 – – 68.3 – –
Core ML fp32 57.1 57.1 57.1 50.6 11.8 50.6
Core ML fp16 57.3 57.1 55.2 32.3 10.8 25.8
Core AI fp32 57.1 57.1 57.1 55.7 12.0 12.3
Core AI fp16 57.3 57.0 57.3 45.3 10.8 35.5
ExecuTorch Core ML fp16 – – 57.0 – – 31.5
rfdetr-large ONNX Runtime fp32 60.2 – – 123.4 – –
Core ML fp32 60.2 60.2 60.2 80.4 16.1 80.2
Core ML fp16 59.0 59.9 58.1 46.7 14.1 48.9
Core AI fp32 60.2 60.2 60.2 87.1 16.3 16.5
Core AI fp16 60.1 59.9 58.9 62.1 14.5 51.2
ExecuTorch Core ML fp16 – – 60.1 – – 54.8
rfdetr-xlarge ONNX Runtime fp32 61.1 – – 234.9 – –
Core ML fp32 61.1 61.1 61.1 157.1 34.8 157.5
Core ML fp16 61.1 60.8 58.5 102.5 29.7 101.2
Core AI fp32 60.6 61.1 61.1 183.3 34.0 35.2
Core AI fp16 60.5 60.9 60.6 125.9 29.8 110.2
ExecuTorch Core ML fp16 – – 60.8 – – 92.1
rfdetr-2xlarge ONNX Runtime fp32 62.3 – – 322.3 – –
Core ML fp32 62.3 62.3 62.3 215.9 47.2 216.5
Core ML fp16 62.3 62.6 59.6 114.3 39.3 131.8
Core AI fp32 61.1 62.3 62.3 244.4 44.6 45.9
Core AI fp16 60.8 62.3 61.9 148.4 39.0 151.3
ExecuTorch Core ML fp16 – – 62.0 – – 126.6
  • Core ML fp16 on npu loses 1.9–4.5 mAP on RF-DETR. The same models in ExecuTorch Core ML on npu stay within 0.3 mAP of ORT.
  • Core ML fp32 on npu runs on the CPU: the Neural Engine runs only fp16. Core AI fp32 on npu is as fast as the gpu row.
  • Core AI fp32 on cpu loses 0.5–1.2 mAP on xlarge and 2xlarge. The same artifacts keep their mAP on gpu and npu.
  • † SAB marked the row as throttled (macOS thermal state fair or higher).

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code where necessary, particularly in hard-to-understand areas
  • My changes generate no new warnings or errors
  • I have updated the documentation accordingly (if applicable)

Additional Context

🤖 Generated with Claude Code

@socket-security

socket-security Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Addedcoreai-core@​1.0.0b299100100100100

View full report

leeclemnet and others added 2 commits September 30, 2026 15:34
- CoreMLRuntime and CoreAIRuntime run on cpu, gpu and npu.
- ExecuTorchRuntime gets the npu device for Core ML programs.
- ThermalStateMonitor reads the macOS thermal state during each run.
- New coreml, coreai and mac extras. The scripts get no new rows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Remove coreai from the mac extra: coreai-core has wheels only for macOS 26+.
- Pin coreai-core to 1.0.0b2, and close the Core AI event loop.
- Return False when the private ExecuTorch backend API is missing.
- Add busy() to the macOS thermal monitor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@leeclemnet
leeclemnet force-pushed the feat/executorch-runtime branch from 56ecd8f to e0c7228 Compare September 30, 2026 19:35
@leeclemnet
leeclemnet marked this pull request as ready for review September 30, 2026 19:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant