Two heuristic bugs in the section and trigger checks, surfaced in
addyosmani#387 and triaged by @addyosmani and @nucliweb:
Section check: content.includes('## Overview') matches headings inside
fenced code blocks and sub-headings (### Overview) because substring
matching can't distinguish prose from examples. Replace with a
line-anchored regex match against content that has had fenced code
blocks stripped, so only real top-level headings satisfy the check.
Trigger check: 'Do not use when testing' satisfied the DESCRIPTION_TRIGGER
regex because it contains 'use … when'. Add a negation pattern and reject
descriptions where the only trigger match is negated.
Verified: all 24 skills pass (output unchanged — no current skill has
either bug pattern). Crafted edge cases confirm the fixes catch both
issues. Eval suite green, rank-1 rate unchanged at 86%.
Co-authored-by: Cursor <cursoragent@cursor.com>
Per review: the policy collections (REQUIRED_SECTIONS, SECTION_EXEMPT_SKILLS,
SKILL_REF_PATTERNS, regexes) were exported by reference, so a consumer could
mutate shared state and change lint results process-wide. Export only the
linting functions; keep the policy collections private. No behavior change
(validator output byte-identical).
Split the skill validation rules out of the validate-skills.js CLI into a
shared, importable scripts/lib/skill-lint.js: a pure lintSkillContent() with
no I/O plus a thin lintSkill() file wrapper. validate-skills.js becomes a thin
CLI over the lib. No behavior change to the validator's output or exit codes;
this only makes the rules unit-testable for a follow-up test battery.
Test Plugin Installation / Validate skill content (push) Has been cancelled
Test Plugin Installation / Validate command parity and description sync (push) Has been cancelled
Test Plugin Installation / Validate plugin structure (push) Has been cancelled
Test Plugin Installation / Test plugin installation (push) Has been cancelled
Two latent Tier-3 bugs caught in review before the path was ever exercised:
- The grader prompt embeds the full stream-json trace (up to megabytes) and
was passed as an argv entry, which would fail with E2BIG on any real run.
It now goes to `claude -p` over stdin; the executor prompt moves to stdin
for the same reason.
- The executor ran headless with no permission mode, so file edits and
command runs could be denied, forcing the narrate-instead-of-perform
failure mode that trace grading exists to catch. It now runs with
--permission-mode acceptEdits and a pre-approved tool list
(Read,Glob,Grep,Edit,Write,Bash), documented in the README.
Per review from @federicobartoli and @nucliweb on #342:
Tier 3 (behavioral):
- Grade the execution trace, not the final output: executor runs with
--output-format stream-json --verbose so the grader judges tool calls
and file edits rather than the model's self-reporting.
- Run each eval in a throwaway workspace; files[] fixtures materialize
from evals/fixtures/ so evals can operate on real code.
- Node-level timeouts on executor and grader calls; grader output parsed
and shape-validated before writing (raw saved on failure); the trace is
fenced as untrusted data in the grader prompt.
- All 24 behavioral evals flagged trust_level: "provisional" until they
gain fixtures; the runner surfaces this and exits nonzero on failed
expectations.
Tier 2 (deterministic):
- Negative triggers accept an "owner" skill that must outrank this one,
turning them into pairwise routing tests that cannot pass vacuously;
37 of 48 negatives now declare owners (the rest are tracked in #351).
- Warn when a case file is below the documented minimums (3 positive /
2 negative / 1 behavioral); promotion to error tracked in #352.
- Stemmer: cluster trailing y/i ("simplify"/"simplifies").
Baseline holds: 120 checks, 0 errors, 85% trigger rank-1 rate.
There was no way to measure whether skills trigger correctly, stay
distinct, or change agent behavior. This adds evals, aligned with what
the community has converged on, with a deterministic CI tier on top:
- evals/cases/<skill>.json for all 24 skills. The evals[] block uses
Anthropic skill-creator's evals.json schema verbatim (id, prompt,
expected_output, expectations[]) so its runner, benchmarks, and eval
viewer work against our files unmodified. A trigger block (this
repo's extension) adds positive/negative routing prompts per skill.
- scripts/run-evals.js, zero-dependency runner:
Tier 2 (CI): trigger evals via stemmed TF-IDF ranking over skill
descriptions (positive prompts must rank top-k, negative prompts
must not rank first), catalog collision detection between skill
descriptions, schema and coverage checks.
Tier 3 (opt-in): --behavioral <skill> executes each eval through
headless claude -p and grades the transcript against expectations[]
(superpowers-style); --dry-run previews without spending tokens.
- CI: run the deterministic tier in the validate-skills job.
- Docs: evals/README.md defines the framework and prior art;
CONTRIBUTING requires an eval file for new skills (warning-level in
the runner until in-flight skill PRs clear); CLAUDE.md pointers.
Current baseline: 120 checks pass, 85% trigger rank-1 rate across 72
positive prompts, zero catalog collisions.
The validator claims to check skills "against the rules in
docs/skill-anatomy.md", but two rules that doc marks as Required were
never enforced:
- Directory names must be lowercase-hyphen-separated (Naming Conventions).
Previously only `name === dirName` was checked, so `My_Skill/` passed.
- Descriptions must say *when* to use the skill, not just what it does
(Required vs Recommended). Formalized in #167 but the validator was
never updated to match.
Adds both as blocking checks. All 24 existing skills pass, so this is a
non-breaking guardrail. Closes the remaining gap in #233 (frontmatter,
name-match, and section checks already shipped in the original validator).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Support single-quoted TOML description strings (previously silently returned null)
- Wrap file reads in try/catch to produce actionable CI errors instead of raw stack traces
- Remove dead tomlStems variable; declare allTomlStems once and compute allCanonicalStems for an accurate command count in the summary
- Simplify description mismatch output to always show all three values, removing a redundant conditional branch
- scripts/validate-skills.js: zero-dependency Node.js validator that
checks every skill for valid frontmatter, description length (≤1024),
required sections (Overview, When to Use, Common Rationalizations,
Red Flags, Verification), and dead cross-skill references
- Skills with type:meta or exempt:sections in frontmatter skip section
checks; applied to using-agent-skills (meta) and idea-refine (legacy
structure predating the anatomy spec)
- CI: validate-skills job runs before plugin-manifest validation and
blocks merge on any error; uses Node 20, no npm install required