Conversation
…ackbone) Three real-time detection models built like RFDETRXLarge: a ModelConfig subclass per size and an RFDETR subclass, with release weights registered in ModelWeights. They use Meta's Perception Encoder PE-Core-T trunk (timm vit_pe_core_tiny_patch16_384) with windowed attention and full attention at out_feature_indexes, in the NAS-selected architectures (Atto 380px/1 window/0 decoder layers/300 queries, Femto 384px/patch 16/2 windows/ 2 layers/200 queries, Pico 560px/2 windows/3 layers/200 queries) and the COCO weights of the PE-Core-T supernet. rfdetr_plus.models.pe_core ports rf-detr-internal's fixed PE backbone and registers it with rfdetr's backbone registry as encoder "pe_core_t", with rf-detr-internal's layer-decay rules. Differences from rf-detr-internal: - Inputs at another resolution (multi-scale training, custom resolution) resample pos_embed and rebuild RoPE inside the forward pass. rf-detr-internal called timm's set_input_size there, which swapped pos_embed for a new Parameter outside the optimizer and resampled it down and back up on every such batch. - A checkpoint pos_embed from another position grid is resampled on load; export() bakes position embeddings for the export shape (also via rfdetr's set_export_shape hook). - PE-CLIP's attention-pool head, which the detector never runs, is dropped (0.54M parameters). - timm builds the trunk directly (no open_clip), identical to the open_clip trunk. Parity (private harness in rf-detr-internal tests/pe_parity/plus_port): weights, eval/export/off-resolution outputs, train-mode outputs, every gradient, criterion losses, per-parameter lr/weight decay and a 3-step AdamW trajectory are bit-identical to rf-detr-internal once the shared code's known numerics differences are aligned. The encoder alone is bit-identical with no alignment. Requires rfdetr>=1.12.0 (backbone registry, ModelConfig.dim_feedforward, export-shape hook) and timm>=1.0.27. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the PE-Core-T models - backbone_lora=True wraps PE-Core-T's attention qkv projections, freezes the rest of the encoder and gets gradients into every adapter. - RFDETR.inference() (export graph, TorchScript trace, dtype cast) reproduces eager predict() in fp32 and stays within bf16 rounding. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… encoder - The forward adds the class token and position embedding itself and resamples only when the input grid differs from the stored one. timm's _pos_embed also resampled non-square grids of the stored size, which left antialiased bicubic interpolation (no ONNX symbolic) in graphs exported at a non-square shape. Numerics are unchanged: the parity suite is still bit-identical. - Position-embedding resampling and RoPE rebuilds run outside torch.compile graphs, as rfdetr's DINOv2 interpolation does, and resampling skips antialiasing on MPS. - set_export_shape refuses a second, different shape instead of silently keeping the first one. - Loading drops PE-CLIP attention-pool/head weights (still in rf-detr-internal checkpoints) and rejects a checkpoint of another patch size with a clear error instead of a size-mismatch dump. - RFDETRPECoreTConfig keeps core's dim_feedforward >= 1 bound. - Tests: non-square export (encoder and full graph), repeated export shapes, internal checkpoint keys, patch-size mismatch, gradient checkpointing actually engaging, and a training check that pos_embed moves but is not resampled in place. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Use repeat_interleave, as rf-detr-internal's develop does, so each window's class-token copy belongs to the image of that window. The token is identical across images when it is inserted, so outputs are unchanged; the plus encoder is now bit-identical to rf-detr-internal's develop and rfml-image backbones in forward and backward. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
…e hosted weights Measured on a T4 over 500 COCO val2017 images: Femto mAP@50 0.595 / F1 0.589, Pico 0.626 / 0.611. Thresholds sit ~0.02 below, as for Atto. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Core-T comments COCO AP50 49.4 / 55.9 / 60.2 comes from the same COCO val2017 evaluation as the AP50:95 column. Comments, docstrings and the changelog now describe the behavior without referring to other code bases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… so format="tflite" converts timm's AttentionRope applies RoPE as cat([q[:, :, :npt], rope(q[:, :, npt:])], dim=2). The encoder sets npt to 0 (the class token has a no-op RoPE row), so the first operand is an empty slice, which onnx2tf reads as the whole tensor: the TFLite conversion failed in the first attention block. export() now binds an attention forward that applies RoPE to all tokens directly. Only the exported copy changes; its values are identical, and TFLite then matches PyTorch to ~1e-6 on COCO val2017 for Atto, Femto and Pico. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Matvezy
marked this pull request as ready for review
September 30, 2026 05:26
Matvezy
requested review from
Borda,
SkalskiP,
isaacrob and
probicheaux
as code owners
September 30, 2026 05:26
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The required rfdetr>=1.12.0 release is not yet published, so dependency resolution currently fails.
Review effort: Balanced
Findings: None
What changed in this PR
Adds Atto, Femto, and Pico detector variants using a registered PE-Core-T backbone.
Changes:
- Implements PE-Core-T windowed attention, RoPE, checkpoint loading, training, and export support.
- Adds model configurations, public exports, weights, benchmarks, and extensive tests.
- Raises dependency requirements and updates project documentation.
| File | Description |
|---|---|
src/rfdetr_plus/models/pe_core.py |
Implements the PE-Core-T backbone. |
src/rfdetr_plus/models/detection.py |
Defines the three model variants. |
src/rfdetr_plus/models/__init__.py |
Exports model classes and configurations. |
src/rfdetr_plus/assets/model_weights.py |
Registers pretrained checkpoints. |
src/rfdetr_plus/__init__.py |
Exposes models publicly. |
tests/test_pe_core.py |
Tests backbone behavior and export. |
tests/test_pe_models.py |
Tests public model workflows. |
tests/test_inference.py |
Adds accuracy and inference coverage. |
tests/test_config.py |
Validates new configurations. |
pyproject.toml |
Adds required dependency versions. |
README.md |
Documents models and benchmarks. |
CHANGELOG.md |
Records the feature. |
AGENTS.md |
Updates agent guidance. |
.github/copilot-instructions.md |
Updates Copilot project context. |
.github/CONTRIBUTING.md |
Updates contributor documentation. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
rfdetr_plus needs the backbone registry and export hooks from roboflow/rf-detr#1568, which is not on PyPI yet. [tool.uv.sources] builds rfdetr from that PR's branch for every uv install in CI, and the floor is relaxed to 1.11.0 because the branch still carries that version. Revert this commit (restoring rfdetr>=1.12.0) once 1.12.0 is released. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #68 +/- ##
===================================
+ Coverage 97% 99% +2%
===================================
Files 6 7 +1
Lines 69 342 +273
===================================
+ Hits 67 338 +271
- Misses 2 4 +2 🚀 New features to boost your workflow:
|
…ts scores with fp32's bf16 scores drift from fp32's by up to ~0.05 (core models' too) and vary by CPU, so the 1e-2 check failed on some CI runners. fp32 still matches eager predict within 1e-5. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…point round-trip test
Restores `pretrain_weights=f"{model_cls.size}.pth"` for every Plus model, so the XLarge and 2XLarge cases test exactly
what they did before; the PE-Core-T models resolve from those names too.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
RFDETRAtto,RFDETRFemtoandRFDETRPico: real-time detectors on a PE-Core-T backbone, built likeRFDETRXLarge(aModelConfigandRFDETRsubclass per size, weights inModelWeights).The architectures come from a neural architecture search over a COCO-trained PE-Core-T RF-DETR supernet; the weights are hosted in
gs://rfdetr/platform-licensed/in therf-detr-xlarge.pthformat.Changes
models/pe_core.py: the PE-Core-T encoder, registered with rfdetr'sregister_backboneaspe_core_t.vit_pe_core_tiny_patch16_384, without PE-CLIP's unused attention-pool head.[2, 5, 8, 11], with RoPE.pos_embed.pos_embedsaved at another grid is resampled on load, so a customresolutionworks.models/detection.py:RFDETRPECoreTConfig, per-size configs and the three model classes.assets/model_weights.py: URLs and MD5s of the three checkpoints.rfdetr>=1.12.0(companion rf-detr PR: backbone registry,dim_feedforward, export-shape hook) andtimm>=1.0.27,<2.Testing
tests/test_pe_core.py,tests/test_pe_models.py(offline, fake release-format checkpoints):inference().tests/test_config.pyconfig checks.tests/test_inference.pyCOCO benchmarks against the hosted weights, with GPU thresholds measured on a T4.Before merging
[tool.uv.sources]inpyproject.toml, floor temporarily>=1.11.0).rfdetr>=1.12.0from PyPI.🤖 Generated with Claude Code