Skip to content

Creates false positives on pure text-based content #37

Description

@francoisdp

I ran Opus 4.8 on all the skills in my Github repo (https://github.com/francoisdp/skills) it came up with the following report:

All scans are complete. Here are the results before I change any file or push anything. None of the eight skills ships executable code, so every finding is a
pattern match against prose, a LICENSE file, or template comments.

SkillSpector results (the 8 deployed skills)

Skill Score Tool verdict Findings
deep-research 5 SAFE 1 LOW (LICENSE boilerplate)
study-note 5 SAFE 1 LOW (LICENSE)
youtube2article 5 SAFE 1 LOW (LICENSE)
branded-document-toolkit 15 SAFE 1 LOW (LICENSE), 1 MEDIUM (README)
model-router 35 CAUTION 1 HIGH, 1 MEDIUM
anti-ai-writing 65 DO_NOT_INSTALL 2 HIGH, 1 MEDIUM, 1 LOW
deploy-skill 65 DO_NOT_INSTALL 1 CRITICAL, 1 MEDIUM, 1 LOW
marp-deck 80 DO_NOT_INSTALL 3 HIGH, 1 LOW

My assessment of the flagged items

I read every HIGH and CRITICAL line in the source. They are all false positives produced by the scanner pattern-matching documentation and template text:

  • deploy-skill — CRITICAL "Harmful Content Injection" at README.md:33. The line is ordinary prose describing the /deploy-skill trigger and that the skill
    refuses to publish another person's work. Nothing harmful.
  • marp-deck — 3× HIGH "Hidden Instructions" in templates. These are the Marp HTML comment blocks (<!-- Template: default … Replace [TITLE] -->). Legitimate
    template authoring comments, not hidden directives.
  • anti-ai-writing — HIGH "Direct Prompt Extraction" at SKILL.md:916 is a section header about output types. HIGH "Tool Parameter Abuse" was matched inside the
    binary bytes of a PDF example.
  • model-router — HIGH "Memory Manipulation" at SKILL.md:81 is the instruction describing what context to pass to a worker subagent, including reading CLAUDE.md.
    Normal dispatch description.
  • The LOW at LICENSE:16 appears on all eight. It is the MIT licence text.

In substance, all eight skills are documentation-only, with no executable scripts, no network calls, no credential access, and no persistence mechanism. The tool
over-flags instruction-heavy skills. SkillSpector has no suppression or baseline flag, so I cannot make the raw verdict turn green for the four flagged skills
without editing the skills themselves.

Activity

  1. HetCreep commented on Jun 13, 2026

    @HetCreep

    Another concrete repro — and an ironic one: a security-audit skill suite (CoalMine, code-quality "canary" skills) scores 58/100 HIGH / DO NOT INSTALL on the static pass, entirely on benign text:

    Static finding What it actually is
    HIGH · P2 Hidden Instructions An HTML-comment metadata stamp <!-- coalmine: verified <date> · revalidate Nd --> (machine-readable rule-freshness marker — reviewable, no invisible/bidi chars). The static pass appears to flag every HTML comment as a possible hidden instruction.
    MED · EA2 Autonomous Decision The literal line "ADOPT and every CONFORM fix are gated through ask_question — never assume approval" — i.e. the consent gate itself.
    MED · RA2 Session Persistence A user-facing message string that names a documented kill-switch file path.

    The LLM semantic stage returned findings=0 on the files it evaluated, so it reads these correctly — but on a free hosted tier the semantic pass gets rate-limited (#10) and aborts to static-only (#9), so the false-positive score is what ships.

    Possible mitigations: don't treat every HTML comment as P2 (gate on invisible/bidi chars + agent-directing content, cf. #39); recognise an explicit consent-gate clause before raising EA2.

  2. Spectorian commented on Jun 16, 2026

    @Spectorian
    Collaborator

    Thanks for the detailed examples. This is useful calibration feedback.

    Static analysis is intentionally conservative, but documentation-only or template-heavy skills should not receive high-risk recommendations solely from benign prose, template comments, license text, or example artifacts. The examples here are helpful because they point to specific places where static findings may need better context or confidence handling.

    We’ll add this to our false-positive review and calibration.

  3. HetCreep commented on Jun 20, 2026

    @HetCreep

    Update (SkillSpector v2.2.3, 2026-06-20) — a 4th context-blind class, and the score got worse. After we shipped a consent-gated self-update feature to the same suite, all three skills now score 100/100 CRITICAL / DO_NOT_INSTALL on the static pass — CoalMine rose 58→100, CoalTipple 0→100, CoalBoard 100. The new finding is RA1 self-modification (6–8 hits/repo), firing on every self-update string: the /<skill>:update command, the conductor hook that only schedules a throttled version-check (no network, no file write), and // self-update comments. None of it rewrites the skill's own files — the agent only ever offers the host's native claude plugin update. A benign consent-gated update-check reading as malicious self-modification — the same token-not-intent blindness as the EA2/P2 cases above. (Static-only again — the LLM stage 429'd, per #10.)

  4. rng1995 commented on Aug 14, 2026

    @rng1995
    Collaborator

    Active implementation: PR #49 addresses the structural Markdown-comment/P2 false-positive subset reported here. The PR references this issue; keeping the issue open until that work merges and the broader false-positive classes are evaluated.

  5. rng1995 commented on Sep 28, 2026

    @rng1995
    Collaborator

    Updated implementation links: PR #622 is a new proposed approach for the structural-comment/P2 subset, presented as a replacement for the still-open PR #49. Neither has merged. This report spans additional context-sensitive false-positive classes, so it remains open beyond that single slice.

  6. rng1995 commented on Sep 28, 2026

    @rng1995
    Collaborator

    Related subset now merged: PR #513 resolves the report-formatting-heading P6 false positive tracked specifically in #512. Current-main verification passed 302 focused heading/context, reporting, prepared-runner, and CLI tests, retaining malicious extraction controls.

    Keeping this broader issue open: fixing that heading class does not establish resolution of all documentation/template false positives reported here, including the structural-comment/P2 and other context-sensitive classes. The pending structural-comment alternatives remain PR #49 and PR #622.

  7. rng1995 commented on Oct 4, 2026

    @rng1995
    Collaborator

    Implementation-link update: PR #49 is now closed without merging. PR #622 remains open for the narrowly proven benign header-comment carve-out tracked in #677; it is not a general comment exemption or a fix for every spanning-comment case. The P2 comment-boundary work in PR #452 also remains open.

    Keeping this broader report open for its remaining documentation/template false-positive classes.

    PR-state snapshot checked on 2026-10-04:

    • PR #622 — open, non-draft; GitHub review decision: changes requested; latest reported check rollup: success.
    • PR #452 — open, non-draft; GitHub review decision: changes requested; merge conflicts reported; latest reported check rollup: success.

    The header-comment carve-out and comment-boundary work are separate partial fixes; the broader documentation/template false-positive scope remains open.

    Keeping this issue open: the relevant implementation is not merged and the remaining scope still needs verification. Check/review status is a point-in-time snapshot, not a claim of merge readiness.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions