- Expanded SKILL.md to 480 lines covering all three evaluation layers,
composite scoring formula with blend weights, dimension grade interpretation,
all five anti-pattern flags with fix guidance, Elo ranking mechanics, CLI
reference with code examples, and a troubleshooting section
- Added references/rubrics.md (510 lines) with full anchored rubrics for all
four judge dimensions (triggering accuracy, orchestration fitness, output
quality, scope calibration) including per-level examples and calibration norms
- Updated description to include specific trigger contexts for autonomous invocation
Add plugin.json, three commands (eval/compare/certify), eval-orchestrator agent,
and evaluation-methodology skill to wire plugin-eval into the claude-agents ecosystem.