Skip to content

fix(ae1): keep reference-resolution caveats from flagging SKILL.md self-references - #661

Merged
rng1995 merged 2 commits into
mainfrom
naren/fix-ae1-missing-reference-source
Sep 29, 2026
Merged

rng1995 merged 2 commits into
mainfrom
naren/fix-ae1-missing-reference-source

Conversation

@rng1995

@rng1995 rng1995 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A skill author reported that SkillEvaluator's Tier 1 check (SkillSpector static scan) blocks ordinary skills with a HIGH AE1 "Referenced artifact was not completely inspected" finding on SKILL.md itself.

Two things in SKILL.md combine to cause it:

  1. A backticked path that isn't in the bundle. For example, "Other skills keep templates under `templates/config.yaml`". Reference resolution records a reference_missing ledger event, and records it against SKILL.md, the file that mentions the path. An ambiguous mention (one that matches more than one bundled file) is recorded the same way, as reference_unresolved.
  2. A backticked `SKILL.md`. For example, "Keep `SKILL.md` under 500 lines." This resolves as a reference to SKILL.md itself. A markdown link such as [this file](SKILL.md) does the same.

_reference_coverage_findings then sees a partial event on the target SKILL.md and emits a HIGH AE1 for every such mention. SkillEvaluator blocks on any HIGH finding.

A missing reference names a path the bundle does not carry, and an ambiguous one leaves open which bundled file a mention means. Neither leaves bytes of the mentioning file uninspected. The NVCARPS-154 acceptance tests already expect each half separately:

  • a self-reference stays complete with no AE1;
  • a missing or ambiguous primary reference makes the scan incomplete, also with no AE1.

Only the combination produced AE1.

Change

src/skillspector/nodes/finalize_inspection_ledger.py: when indexing ledger events by path for AE1, skip reference-resolution caveats. The skip matches exactly the event shape build_context emits: outcome=partial, phase="reference_resolution", and reason_code of reference_missing or reference_unresolved. These events are only recorded on the primary file, which is SKILL.md or, when that is the only primary, lowercase skill.md. So in practice this affects only references that resolve to the primary file.

The filter runs once, while building the per-path index; that index is only used to find a reference target's events. The other checks (fatal paths, status evidence, truncation, canonicalization) still read the unfiltered ledger.

What stays the same:

  • The missing or ambiguous path is still a reference_missing / reference_unresolved ledger exception.
  • is_complete stays false, the recommendation stays CAUTION, and --fail-on-incomplete still exits 1.
  • The MCP safe_to_install handling is unchanged. Missing-reference-only caveats already pass. Ambiguous references still return safe_to_install=false, because that gate depends only on the ledger exception (mcp_server.py), not on AE1. That is where the fail-closed signal for ambiguity comes from.
  • Real content limitations on SKILL.md, such as static_parse_limit, still produce AE1 for a self-reference. Its evidence now lists only the real limitation.
  • Any other event shape still counts toward AE1: a failed outcome, a different phase, or reference_extraction_limit (which leaves part of SKILL.md unexamined).

Before and after

skillspector scan <skill> --no-llm --format json, with SKILL.md containing `templates/config.yaml` (missing) and `SKILL.md`:

main @ 8831219 this branch
Issues AE1 HIGH, SKILL.md (partial) none
Ledger exceptions SKILL.md: reference_missing SKILL.md: reference_missing
Recommendation CAUTION, max severity HIGH CAUTION, max severity NONE

The ambiguous case (`config.yaml` with both a/config.yaml and b/config.yaml present, plus `SKILL.md`) behaves the same way: no AE1 on this branch, SKILL.md: reference_unresolved in the ledger exceptions, CAUTION, and MCP run_scan still returns safe_to_install=false.

The markdown-link self-reference, the markdown-link missing target and the lowercase skill.md primary also produce no AE1, with the caveat kept.

Controls:

Tests

  • tests/nodes/test_finalize_inspection_ledger.py
    • test_reference_caveat_does_not_make_its_source_an_ae1_target[missing|ambiguous]
    • test_self_reference_ae1_still_reports_source_content_limitations[missing|ambiguous]: a real static_parse_limit on SKILL.md still yields AE1, and its evidence excludes the reference caveat.
    • test_other_reference_event_shapes_still_yield_self_reference_ae1: a failed outcome or another phase with either reason, and reference_extraction_limit, all still yield AE1.
  • tests/nodes/test_security_remediation.py::test_reference_caveat_does_not_turn_self_reference_into_ae1: full graph scan, parametrized over a backticked self-reference, a markdown-link self-reference, a markdown-link missing target, a lowercase skill.md primary and an ambiguous reference. Each checks no AE1, the caveat retained, is_complete=false and CAUTION.

On main, the four no-AE1 unit cases and all five end-to-end cases fail; they pass here. The out-of-shape cases pass on main too, because they pin the AE1 that the filter must keep.

Local runs:

  • ruff check src/ tests/ and ruff format --check src/ tests/ pass.
  • The two touched test files: 569 passed.
  • -m integration tests/integration/test_opaque_reference_reporting.py: 47 passed.
  • The full unit suite (-m "not integration and not provider"): 8403 passed, 14 skipped, 4 xfailed, 0 failed.

Related

🤖 Generated with Claude Code

A path in SKILL.md that is not in the bundle (e.g. `templates/config.yaml`)
records a reference_missing ledger event on SKILL.md. AE1 then treated
SKILL.md as partially inspected, so every backticked `SKILL.md` in the same
file became a HIGH analysis-evasion finding. SkillEvaluator blocks Tier 1 on
any HIGH finding, so ordinary skills failed.

A missing reference names a path the bundle does not carry; it leaves no
bytes of the mentioning file uninspected. Ignore reference_missing events
when judging whether a resolved reference target was completely inspected.
The completeness ledger, CAUTION verdict and MCP gate still report the
missing path. Content limitations on SKILL.md, such as static_parse_limit,
still produce AE1, and its evidence no longer lists the missing reference.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[SkillSpector Review]

Hi @rng1995, thank you for this fix. It unblocks a real customer workflow, and the tests and before/after evidence made it easy to check. I found no blocking issues; the comments below are follow-ups.

What the change does. When a resolved reference points at SKILL.md, _reference_coverage_findings now ignores reference_missing ledger events for that target. It does this before deciding the disposition, building the AE1 evidence and running the format-only check. The missing path still shows up as a completeness caveat.

How I checked it

  • Traced every producer and consumer of these events.
  • Ran the 3 new tests on origin/main in a scratch worktree: all 3 fail. On this branch they pass.
  • Scanned 9 fixtures on main and the branch through JSON, SARIF, markdown, the CLI --fail-on-* flags and MCP run_scan.

Verified OK

  • Only one producer. reference_missing is emitted in one place (build_context.py ~2826-2843), always with path=primary_path. Other scan modes don't reach this code:
    • Nested archives don't resolve references.
    • Recursive scans run a separate graph per skill.
    • Transitive results are merged after the graph runs.
      So the filter can only affect references whose target is the primary file.
  • SKILL.md inventory disposition stays analyzed when references are missing. Only reference_extraction_limit / source_partial set it to partial, so the fix is complete. Leaving reference_extraction_limit alone is correct, because the inventory still triggers AE1 for it.
  • Other checks still read unfiltered data. fatal_paths, invalid_status_paths, duplicate_inventory_paths and ledger_evidence_complete all use the unfiltered ledger or state. _has_only_format_limitations needs a verified passive PNG, so it can't be affected.
  • Real limitations still produce AE1. With a real static_parse_limit on SKILL.md, AE1 is still HIGH and its evidence now lists only static_parse_limit.
  • Odd reason_code values keep AE1. None, an absent key or a list all still produce it.
  • The caveat is still reported everywhere:
    • JSON ledger_exceptions
    • SARIF toolExecutionNotifications
    • the markdown table
    • is_complete=false / CAUTION
    • --fail-on-incomplete still exits 1
    • --fail-on-findings now exits 0 (it was 1 on main)
    • MCP safe_to_install is unchanged
  • Lint and tests pass. ruff check and ruff format --check pass. The touched test files pass (558), as do the integration tests in test_opaque_reference_reporting.py (47).

Non-blocking follow-ups (details inline)

  1. Ambiguous references (reference_unresolved) still turn a SKILL.md self-reference into a false AE1 HIGH. Either filter them too, or pin the current behavior with a test and explain why.
  2. Filter once, when building events_by_path, instead of for every reference.
  3. Match the exact event shape build_context emits, so a future producer can't erase AE1.
  4. The test helper's type annotation doesn't match what ledger_event returns.
  5. Parametrize the end-to-end test over the other self-reference forms.

PR description. Two small fixes:

  • "Only SKILL.md (the primary file) can carry these" should say the primary file can be lowercase skill.md (build_context.py:2708-2710). The fix already covers that case; I checked it with a scan.
  • "Ambiguous references … deliberately left unchanged" gives no reason; see follow-up 1.

Decision: Comment (no blocking issues). Reviewed head 8d1c5a97ca6d4153fb951267854db463df0a9b21. CI is green: changes, lint, test-unit, OpenCode TypeScript Tests, DCO Check, docker-smoke.

Comment thread src/skillspector/nodes/finalize_inspection_ledger.py Outdated
Comment thread src/skillspector/nodes/finalize_inspection_ledger.py Outdated
Comment thread src/skillspector/nodes/finalize_inspection_ledger.py Outdated
Comment thread tests/nodes/test_finalize_inspection_ledger.py Outdated
Comment thread tests/nodes/test_security_remediation.py Outdated
Reference resolution records an ambiguous mention (reference_unresolved)
on SKILL.md the same way it records a missing one. It is about which
bundled file a mention means, not SKILL.md bytes left uninspected, so a
SKILL.md self-reference no longer turns it into AE1 either. The caveat
still reaches the completeness ledger, and MCP safe_to_install stays
False for it.

Skip only the exact event shape build_context emits (partial outcome,
reference_resolution phase, one of the two reasons), once while indexing
events by path. Any other shape still counts toward AE1, as does
reference_extraction_limit.

Parametrize the end-to-end test over markdown-link, lowercase skill.md
and ambiguous variants, pin the out-of-shape events in unit tests, and
type the test helper with InspectionLedgerEvent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
@rng1995 rng1995 changed the title fix(ae1): keep missing references from flagging SKILL.md self-references fix(ae1): keep reference-resolution caveats from flagging SKILL.md self-references Sep 29, 2026
@rng1995
rng1995 merged commit ec6633e into main Sep 29, 2026
6 checks passed
@rng1995
rng1995 deleted the naren/fix-ae1-missing-reference-source branch September 29, 2026 17:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant