Skip to content

feat(models): add RFDETRAtto, RFDETRFemto and RFDETRPico (PE-Core-T backbone) - #68

Open
Matvezy wants to merge 10 commits into
mainfrom
feat/pe-core-t-atto-femto-pico
Open

Matvezy wants to merge 10 commits into
mainfrom
feat/pe-core-t-atto-femto-pico

Conversation

@Matvezy

@Matvezy Matvezy commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

Adds RFDETRAtto, RFDETRFemto and RFDETRPico: real-time detectors on a PE-Core-T backbone, built like RFDETRXLarge (a ModelConfig and RFDETR subclass per size, weights in ModelWeights).

Atto Femto Pico
resolution / patch 380 / 20 384 / 16 560 / 20
attention windows 1 2 2
decoder layers / queries 0 / 300 2 / 200 3 / 200
params 7.4M 7.9M 8.4M
COCO AP50 / AP50:95 49.4 / 30.5 55.9 / 37.8 60.2 / 41.6
RF100-VL AP50 / AP50:95 78.2 / 48.3 82.5 / 53.8 84.1 / 56.0
T4 TensorRT FP16 latency 1.0 ms 1.4 ms 1.7 ms

The architectures come from a neural architecture search over a COCO-trained PE-Core-T RF-DETR supernet; the weights are hosted in gs://rfdetr/platform-licensed/ in the rf-detr-xlarge.pth format.

Changes

  • models/pe_core.py: the PE-Core-T encoder, registered with rfdetr's register_backbone as pe_core_t.
    • timm trunk vit_pe_core_tiny_patch16_384, without PE-CLIP's unused attention-pool head.
    • Windowed attention, merged to full attention at blocks [2, 5, 8, 11], with RoPE.
    • Layer-wise learning-rate and weight-decay rules for the trunk.
    • Inputs at another resolution resample the position embedding and RoPE in the forward pass, without replacing any parameter, so multi-scale training keeps training pos_embed.
    • A checkpoint pos_embed saved at another grid is resampled on load, so a custom resolution works.
    • Export bakes position embeddings for the export shape, square or not, so ONNX graphs contain no runtime resample.
  • models/detection.py: RFDETRPECoreTConfig, per-size configs and the three model classes.
  • assets/model_weights.py: URLs and MD5s of the three checkpoints.
  • Dependencies: rfdetr>=1.12.0 (companion rf-detr PR: backbone registry, dim_feedforward, export-shape hook) and timm>=1.0.27,<2.
  • Docs: README model table, CHANGELOG, and the dependency and model lists in the contributor docs.

Testing

  • tests/test_pe_core.py, tests/test_pe_models.py (offline, fake release-format checkpoints):
    • windowing and RoPE;
    • off-resolution forward;
    • load-time resampling;
    • export graphs at native and non-square shapes;
    • ONNX export;
    • a multi-scale training epoch;
    • LoRA;
    • inference().
  • tests/test_config.py config checks. tests/test_inference.py COCO benchmarks against the hosted weights, with GPU thresholds measured on a T4.

Before merging

🤖 Generated with Claude Code

Matvezy and others added 4 commits September 29, 2026 06:18
…ackbone)

Three real-time detection models built like RFDETRXLarge: a ModelConfig subclass per size and an RFDETR subclass,
with release weights registered in ModelWeights. They use Meta's Perception Encoder PE-Core-T trunk (timm
vit_pe_core_tiny_patch16_384) with windowed attention and full attention at out_feature_indexes, in the
NAS-selected architectures (Atto 380px/1 window/0 decoder layers/300 queries, Femto 384px/patch 16/2 windows/
2 layers/200 queries, Pico 560px/2 windows/3 layers/200 queries) and the COCO weights of the PE-Core-T supernet.

rfdetr_plus.models.pe_core ports rf-detr-internal's fixed PE backbone and registers it with rfdetr's backbone
registry as encoder "pe_core_t", with rf-detr-internal's layer-decay rules. Differences from rf-detr-internal:
- Inputs at another resolution (multi-scale training, custom resolution) resample pos_embed and rebuild RoPE inside
  the forward pass. rf-detr-internal called timm's set_input_size there, which swapped pos_embed for a new
  Parameter outside the optimizer and resampled it down and back up on every such batch.
- A checkpoint pos_embed from another position grid is resampled on load; export() bakes position embeddings for
  the export shape (also via rfdetr's set_export_shape hook).
- PE-CLIP's attention-pool head, which the detector never runs, is dropped (0.54M parameters).
- timm builds the trunk directly (no open_clip), identical to the open_clip trunk.

Parity (private harness in rf-detr-internal tests/pe_parity/plus_port): weights, eval/export/off-resolution
outputs, train-mode outputs, every gradient, criterion losses, per-parameter lr/weight decay and a 3-step AdamW
trajectory are bit-identical to rf-detr-internal once the shared code's known numerics differences are aligned.
The encoder alone is bit-identical with no alignment.

Requires rfdetr>=1.12.0 (backbone registry, ModelConfig.dim_feedforward, export-shape hook) and timm>=1.0.27.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the PE-Core-T models

- backbone_lora=True wraps PE-Core-T's attention qkv projections, freezes the rest of the encoder and gets
  gradients into every adapter.
- RFDETR.inference() (export graph, TorchScript trace, dtype cast) reproduces eager predict() in fp32 and stays
  within bf16 rounding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… encoder

- The forward adds the class token and position embedding itself and resamples only when the input grid differs
  from the stored one. timm's _pos_embed also resampled non-square grids of the stored size, which left
  antialiased bicubic interpolation (no ONNX symbolic) in graphs exported at a non-square shape. Numerics are
  unchanged: the parity suite is still bit-identical.
- Position-embedding resampling and RoPE rebuilds run outside torch.compile graphs, as rfdetr's DINOv2
  interpolation does, and resampling skips antialiasing on MPS.
- set_export_shape refuses a second, different shape instead of silently keeping the first one.
- Loading drops PE-CLIP attention-pool/head weights (still in rf-detr-internal checkpoints) and rejects a
  checkpoint of another patch size with a clear error instead of a size-mismatch dump.
- RFDETRPECoreTConfig keeps core's dim_feedforward >= 1 bound.
- Tests: non-square export (encoder and full graph), repeated export shapes, internal checkpoint keys, patch-size
  mismatch, gradient checkpointing actually engaging, and a training check that pos_embed moves but is not resampled
  in place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Use repeat_interleave, as rf-detr-internal's develop does, so each window's class-token copy belongs to the image
of that window. The token is identical across images when it is inserted, so outputs are unchanged; the plus
encoder is now bit-identical to rf-detr-internal's develop and rfml-image backbones in forward and backward.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@socket-security

socket-security Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Addedpypi/​timm@​1.0.3098100100100100

View full report

Matvezy and others added 3 commits September 29, 2026 19:58
…e hosted weights

Measured on a T4 over 500 COCO val2017 images: Femto mAP@50 0.595 / F1 0.589, Pico 0.626 / 0.611. Thresholds
sit ~0.02 below, as for Atto.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Core-T comments

COCO AP50 49.4 / 55.9 / 60.2 comes from the same COCO val2017 evaluation as the AP50:95 column. Comments, docstrings
and the changelog now describe the behavior without referring to other code bases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… so format="tflite" converts

timm's AttentionRope applies RoPE as cat([q[:, :, :npt], rope(q[:, :, npt:])], dim=2). The encoder sets npt to 0 (the
class token has a no-op RoPE row), so the first operand is an empty slice, which onnx2tf reads as the whole tensor:
the TFLite conversion failed in the first attention block. export() now binds an attention forward that applies RoPE to
all tokens directly. Only the exported copy changes; its values are identical, and TFLite then matches PyTorch to
~1e-6 on COCO val2017 for Atto, Femto and Pico.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Matvezy
Matvezy marked this pull request as ready for review September 30, 2026 05:26
@Borda
Borda requested a balanced review from Copilot September 30, 2026 05:38

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The required rfdetr>=1.12.0 release is not yet published, so dependency resolution currently fails.

Review effort: Balanced
Findings: None

What changed in this PR

Adds Atto, Femto, and Pico detector variants using a registered PE-Core-T backbone.

Changes:

  • Implements PE-Core-T windowed attention, RoPE, checkpoint loading, training, and export support.
  • Adds model configurations, public exports, weights, benchmarks, and extensive tests.
  • Raises dependency requirements and updates project documentation.
File Description
src/​rfdetr_plus/​models/​pe_core.py Implements the PE-Core-T backbone.
src/​rfdetr_plus/​models/​detection.py Defines the three model variants.
src/​rfdetr_plus/​models/​__init__.py Exports model classes and configurations.
src/​rfdetr_plus/​assets/​model_weights.py Registers pretrained checkpoints.
src/​rfdetr_plus/​__init__.py Exposes models publicly.
tests/​test_pe_core.py Tests backbone behavior and export.
tests/​test_pe_models.py Tests public model workflows.
tests/​test_inference.py Adds accuracy and inference coverage.
tests/​test_config.py Validates new configurations.
pyproject.toml Adds required dependency versions.
README.md Documents models and benchmarks.
CHANGELOG.md Records the feature.
AGENTS.md Updates agent guidance.
.github/​copilot-instructions.md Updates Copilot project context.
.github/​CONTRIBUTING.md Updates contributor documentation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

rfdetr_plus needs the backbone registry and export hooks from roboflow/rf-detr#1568, which is not on PyPI yet.
[tool.uv.sources] builds rfdetr from that PR's branch for every uv install in CI, and the floor is relaxed to 1.11.0
because the branch still carries that version. Revert this commit (restoring rfdetr>=1.12.0) once 1.12.0 is released.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov-commenter

codecov-commenter commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.27273% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 99%. Comparing base (05e551a) to head (767a8e9).

Additional details and impacted files
@@         Coverage Diff         @@
##           main   #68    +/-   ##
===================================
+ Coverage    97%   99%    +2%     
===================================
  Files         6     7     +1     
  Lines        69   342   +273     
===================================
+ Hits         67   338   +271     
- Misses        2     4     +2     
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Matvezy and others added 2 commits September 30, 2026 15:04
…ts scores with fp32's

bf16 scores drift from fp32's by up to ~0.05 (core models' too) and vary by CPU, so the 1e-2 check failed on some
CI runners. fp32 still matches eager predict within 1e-5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…point round-trip test

Restores `pretrain_weights=f"{model_cls.size}.pth"` for every Plus model, so the XLarge and 2XLarge cases test exactly
what they did before; the PE-Core-T models resolve from those names too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants