Repository navigation
Creates false positives on pure text-based content #37
Description
Activity
Another concrete repro — and an ironic one: a security-audit skill suite (CoalMine, code-quality "canary" skills) scores 58/100 HIGH / DO NOT INSTALL on the static pass, entirely on benign text:
Static finding What it actually is HIGH · P2 Hidden Instructions An HTML-comment metadata stamp <!-- coalmine: verified <date> · revalidate Nd -->(machine-readable rule-freshness marker — reviewable, no invisible/bidi chars). The static pass appears to flag every HTML comment as a possible hidden instruction.MED · EA2 Autonomous Decision The literal line "ADOPT and every CONFORM fix are gated through ask_question— never assume approval" — i.e. the consent gate itself.MED · RA2 Session Persistence A user-facing message string that names a documented kill-switch file path. The LLM semantic stage returned
findings=0on the files it evaluated, so it reads these correctly — but on a free hosted tier the semantic pass gets rate-limited (#10) and aborts to static-only (#9), so the false-positive score is what ships.Possible mitigations: don't treat every HTML comment as P2 (gate on invisible/bidi chars + agent-directing content, cf. #39); recognise an explicit consent-gate clause before raising EA2.
Thanks for the detailed examples. This is useful calibration feedback.
Static analysis is intentionally conservative, but documentation-only or template-heavy skills should not receive high-risk recommendations solely from benign prose, template comments, license text, or example artifacts. The examples here are helpful because they point to specific places where static findings may need better context or confidence handling.
We’ll add this to our false-positive review and calibration.
Update (SkillSpector v2.2.3, 2026-06-20) — a 4th context-blind class, and the score got worse. After we shipped a consent-gated self-update feature to the same suite, all three skills now score 100/100 CRITICAL / DO_NOT_INSTALL on the static pass — CoalMine rose 58→100, CoalTipple 0→100, CoalBoard 100. The new finding is RA1 self-modification (6–8 hits/repo), firing on every
self-updatestring: the/<skill>:updatecommand, the conductor hook that only schedules a throttled version-check (no network, no file write), and// self-updatecomments. None of it rewrites the skill's own files — the agent only ever offers the host's nativeclaude plugin update. A benign consent-gated update-check reading as malicious self-modification — the same token-not-intent blindness as the EA2/P2 cases above. (Static-only again — the LLM stage 429'd, per #10.)Active implementation: PR #49 addresses the structural Markdown-comment/P2 false-positive subset reported here. The PR references this issue; keeping the issue open until that work merges and the broader false-positive classes are evaluated.
Related subset now merged: PR #513 resolves the report-formatting-heading P6 false positive tracked specifically in #512. Current-main verification passed 302 focused heading/context, reporting, prepared-runner, and CLI tests, retaining malicious extraction controls.
Keeping this broader issue open: fixing that heading class does not establish resolution of all documentation/template false positives reported here, including the structural-comment/P2 and other context-sensitive classes. The pending structural-comment alternatives remain PR #49 and PR #622.
Implementation-link update: PR #49 is now closed without merging. PR #622 remains open for the narrowly proven benign header-comment carve-out tracked in #677; it is not a general comment exemption or a fix for every spanning-comment case. The P2 comment-boundary work in PR #452 also remains open.
Keeping this broader report open for its remaining documentation/template false-positive classes.
PR-state snapshot checked on 2026-10-04:
- PR #622 — open, non-draft; GitHub review decision: changes requested; latest reported check rollup: success.
- PR #452 — open, non-draft; GitHub review decision: changes requested; merge conflicts reported; latest reported check rollup: success.
The header-comment carve-out and comment-boundary work are separate partial fixes; the broader documentation/template false-positive scope remains open.
Keeping this issue open: the relevant implementation is not merged and the remaining scope still needs verification. Check/review status is a point-in-time snapshot, not a claim of merge readiness.
I ran Opus 4.8 on all the skills in my Github repo (https://github.com/francoisdp/skills) it came up with the following report:
All scans are complete. Here are the results before I change any file or push anything. None of the eight skills ships executable code, so every finding is a
pattern match against prose, a LICENSE file, or template comments.
SkillSpector results (the 8 deployed skills)
My assessment of the flagged items
I read every HIGH and CRITICAL line in the source. They are all false positives produced by the scanner pattern-matching documentation and template text:
refuses to publish another person's work. Nothing harmful.
<!-- Template: default … Replace [TITLE] -->). Legitimatetemplate authoring comments, not hidden directives.
binary bytes of a PDF example.
Normal dispatch description.
In substance, all eight skills are documentation-only, with no executable scripts, no network calls, no credential access, and no persistence mechanism. The tool
over-flags instruction-heavy skills. SkillSpector has no suppression or baseline flag, so I cannot make the raw verdict turn green for the four flagged skills
without editing the skills themselves.