Skip to content

feat: add OpenVINO runtime and all RF-DETR detection sizes - #16

Open
leeclemnet wants to merge 3 commits into
refactor/runtime-processor-compositionfrom
feat/cpu-runtimes
Open

leeclemnet wants to merge 3 commits into
refactor/runtime-processor-compositionfrom
feat/cpu-runtimes

Conversation

@leeclemnet

@leeclemnet leeclemnet commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

Native OpenVINO rows, and all RF-DETR detection sizes

What it does: Adds OpenVINORuntime on the Runtime contract from #15. OpenVINO compiles the existing .onnx files on the host, as TensorRT does, so rfdetr, yolov8, yolov11 and yolo26 get OpenVINO fp32 and fp16 rows with no new artifacts. benchmark_rfdetr.py also gets the large, xlarge and xxlarge sizes.
Why: Part of the work toward benchmarks of every export format that rfdetr and Ultralytics share. Stacked on #15. LiteRT and ExecuTorch are in their own PRs, because they need exported artifacts.

Change map

  • OpenVINO runtime (sab/runtimes/openvino.py) — times only InferRequest.infer(). It reads .onnx, .xml, or a directory with one .xml, and sets INFERENCE_PRECISION_HINT from the row precision. Inputs are cast to the model's input dtypes before the timed call, and outputs are copied after it. input_shape reshapes a dynamic image input to a static shape before compilation. The OpenVINO IR from the rf-detr export has the input [?,?,?,?]. A dynamic image input without input_shape fails at load.
  • Host limits (sab/runtimes/base.py, sab/runner.py) — UnavailableOnHost: a runtime raises it at load when the host cannot run the row as requested, and the runner skips the row. The first case is OpenVINO fp16 on a CPU with no native fp16, which compiles in f32.
  • Rows (sab/models/benchmark_{rfdetr,yolov8,yolov11,yolo26}.py) — OpenVINO fp32 and fp16 for each existing .onnx file, after the old rows. RF-DETR large, xlarge and xxlarge get the same rows as the other sizes.
  • Packaging and CI (pyproject.toml, uv.lock, .github/workflows/tests.yml, README.md) — extra openvino. nvidia now includes it.

Behavior changes

  • A script run on a host with the openvino extra (including nvidia) → also runs 2 OpenVINO rows for each ONNX file. Use --runtimes to leave them out.
  • On a CPU with no native fp16 → the OpenVINO fp16 rows are skipped with a reason, not run in f32 under an fp16 label.
  • benchmark_rfdetr.py → 30 rows (was 9): 6 sizes × (TensorRT fp32, TensorRT fp16, ORT CPU, OpenVINO fp32, OpenVINO fp16).

Watch out for

  • sab/models/benchmark_rfdetr.py:56 — large uses rf-detr-large-new.onnx. rf-detr-large.onnx in the bucket is the deprecated RFDETRLargeDeprecated.
  • sab/runtimes/openvino.py:89 — the fp16 check compares the compiled precision with the request. On x86 CPUs with no native fp16, every OpenVINO fp16 row skips.
  • sab/runtimes/openvino.py:126 — outputs keep the dtype of the model (for example int64 labels), as the ONNX Runtime rows do.
  • pyproject.toml:35 — nvidia now installs OpenVINO, so an NVIDIA host runs the OpenVINO CPU rows by default.

Related Issue(s): n/a

Type of Change

  • New feature (non-breaking change that adds functionality)

Testing

  • Unit tests
  • OpenVINO vs ORT CPU on a T4 VM and an Apple M4 Max (50 coco val2017 images)

RF-DETR rows use the official exports in gs://rfdetr/export-2026-09-30: ORT reads the .onnx, and OpenVINO reads the OpenVINO IR (.xml). YOLO26 rows use the bucket .onnx for both runtimes.

CPU: T4 VM (Haswell, 8 cores × 2 HT, fp32)

Model ORT CPU mAP50:95 OpenVINO IR mAP50:95 ORT CPU median ms OpenVINO IR median ms OpenVINO speedup
rfdetr-nano 53.9 53.9 104.9 85.2 1.23×
rfdetr-small 58.3 58.3 170.3 145.2 1.17×
rfdetr-medium 57.1 57.1 223.8 184.4 1.21×
rfdetr-large 60.2 60.2 352.1 289.4 1.22×
rfdetr-xlarge 61.1 61.1 693.7 686.1 1.01×
rfdetr-2xlarge 62.3 62.3 985.7 963.6 1.02×
Model ORT CPU mAP50:95 OpenVINO mAP50:95 ORT CPU median ms OpenVINO median ms OpenVINO speedup
yolo26n.onnx 47.3 47.3 37.5 29.9 1.26×
yolo26s.onnx 54.2 54.2 65.8 58.2 1.13×
yolo26m.onnx 59.1 59.1 150.5 156.7 0.96×
yolo26l.onnx 60.5 60.5 185.2 190.5 0.97×
yolo26x.onnx 61.8 61.8 364.3 420.6 0.87×

Both runtimes use all 16 logical CPUs. By default, the OpenVINO LATENCY hint uses only the 8 physical cores. ENABLE_HYPER_THREADING=True removes this limit.

CPU: Apple M4 Max (12 performance + 4 efficiency cores, on battery)

Model ORT CPU fp32 mAP50:95 OpenVINO IR fp32 mAP50:95 OpenVINO IR fp16 mAP50:95 ORT CPU fp32 median ms OpenVINO IR fp32 median ms OpenVINO IR fp16 median ms
rfdetr-nano 53.9 53.9 4.9 34.5 60.7 39.3
rfdetr-small 58.3 58.3 9.5 58.3 100.2 68.3
rfdetr-medium 57.1 57.1 10.5 68.3 124.1 84.0
rfdetr-large 60.2 60.2 10.5 123.4 196.6 121.9
rfdetr-xlarge 61.1 61.1 29.5 234.9 397.7 240.2
rfdetr-2xlarge 62.3 62.3 29.2 322.3 556.9† 324.9†
Model ORT CPU fp32 mAP50:95 OpenVINO fp32 mAP50:95 OpenVINO fp16 mAP50:95 ORT CPU fp32 median ms OpenVINO fp32 median ms OpenVINO fp16 median ms
yolo26n.onnx 47.3 47.3 47.3 14.5 20.9 14.7
yolo26s.onnx 54.2 54.2 54.3 33.7 43.1 29.2
yolo26m.onnx 59.1 59.1 59.2 82.9 98.4 60.6
yolo26l.onnx 60.5 60.5 60.7 100.3 122.7 77.3
yolo26x.onnx 61.8 61.8 61.7 169.0 236.5 138.6
  • ORT fp32 is faster than OpenVINO fp32 here, because ORT uses KleidiAI SME2 kernels on the M4. With KleidiAI off, ORT is as slow as OpenVINO fp32 or slower.
  • OpenVINO fp16 loses most of the RF-DETR mAP, with the IR as with the ONNX. The CPU plugin runs the LayerNorm (MVN) and Softmax layers of RF-DETR in f16. YOLO26 fp16 keeps its mAP and is the fastest CPU path for YOLO26 s–x.
  • † SAB marked the row as throttled (macOS thermal state fair or higher).

Checklist

  • My code follows the style guidelines of this project
  • I have performed a self-review of my own code
  • I have commented my code where necessary, particularly in hard-to-understand areas
  • My changes generate no new warnings or errors
  • I have updated the documentation accordingly (if applicable)

Additional Context

🤖 Generated with Claude Code

@socket-security

socket-security Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Addedopenvino@​2026.4.07710010010070

View full report

@leeclemnet leeclemnet changed the title feat: add OpenVINO, LiteRT, ExecuTorch runtimes and ORT execution providers feat: add OpenVINO, LiteRT and ExecuTorch runtimes Sep 29, 2026
@leeclemnet
leeclemnet force-pushed the feat/cpu-runtimes branch 2 times, most recently from ea209a5 to 29dfcf5 Compare September 29, 2026 20:34
@leeclemnet leeclemnet changed the title feat: add OpenVINO, LiteRT and ExecuTorch runtimes feat: add OpenVINO runtime and all RF-DETR detection sizes Sep 29, 2026
@leeclemnet
leeclemnet marked this pull request as ready for review September 29, 2026 20:55
@leeclemnet
leeclemnet force-pushed the feat/cpu-runtimes branch 2 times, most recently from 65ed1a7 to b40aaf6 Compare September 30, 2026 16:53
leeclemnet and others added 3 commits September 30, 2026 15:33
Add OpenVINORuntime on the Runtime contract from the previous PR.

- OpenVINORuntime times only InferRequest.infer(). It reads .onnx, .xml,
  or a directory with one .xml, and sets INFERENCE_PRECISION_HINT from
  the row precision.
- OpenVINO compiles the existing ONNX files on the host, as TensorRT
  does. rfdetr, yolov8, yolov11 and yolo26 get OpenVINO fp32 and fp16
  rows after the old rows.
- UnavailableOnHost: a runtime raises it at load when this host cannot
  run the row as requested (OpenVINO fp16 on a CPU without native fp16),
  and the runner skips the row.
- Extra openvino; nvidia now includes it. CI installs it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add large, xlarge and xxlarge to benchmark_rfdetr.py, with the same rows
as nano, small and medium: TensorRT fp32 and fp16, ONNX Runtime CPU, and
OpenVINO fp32 and fp16. The old rows keep their order.

rf-detr-large.onnx in the bucket is the deprecated large model
(RFDETRLargeDeprecated). The rows use rf-detr-large-new.onnx, the
current RFDETRLarge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant