* feat(dgx-spark-ops): scaffold plugin and register in marketplace Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add spark-environment-setup skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(dgx-spark-ops): tighten spark-environment-setup per review Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add spark-training-gotchas skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(dgx-spark-ops): align preflight.sh output contract with docs Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add spark-memory-thermal-ops skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(dgx-spark-ops): trim spark-memory-thermal-ops and source 68GB anchor Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add dgx-spark-ops-engineer agent Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(dgx-spark-ops): defer agent facts to skills Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add /spark-preflight command Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): scaffold plugin and register in marketplace Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add finetuning-method-selection skill with dated model catalog Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add lora-qlora-recipes skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add preference-optimization skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add grpo-rlvr-training skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): add isolation warning to execution reward Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add vision-sft skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add dataset-curation skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): correct loss-masking example in dataset-curation Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add eval-harness-first skill (Phase 0 gate) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): bring eval-harness-first into line band Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): reclaim byte headroom in eval-harness-first Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add trace-to-training-data skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add checkpoint-promotion skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add quantized-export skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add llm-finetuning-architect agent Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add llm-finetuning-training-engineer agent Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add llm-finetuning-eval-engineer agent Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add /finetune phase-gated lifecycle command Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): thread checkpoint path into /finetune Phase 6 Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add /promote-checkpoint command Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): thread checkpoint path via phase5.output in /promote-checkpoint Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): robust checkpoint discovery and goldens fingerprint in re-gate Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * docs: register llm-finetuning and dgx-spark-ops (94 plugins, 203 agents, 175 skills, 109 commands) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * docs(agents): add fine-tuning and spark-ops agent entries (203 agents) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * chore: generate per-harness artifacts for llm-finetuning and dgx-spark-ops Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: final-review cleanups for fine-tuning plugins Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: address PR #624 review feedback (G1 NGC detection, Phase 3 fallback dispatch, registry metadata) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: enforce sandbox boundary in execution grader and reward examples Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: address CodeRabbit review findings on PR #624 Verified and fixed 37 of 43 outstanding CodeRabbit findings across the llm-finetuning and dgx-spark-ops plugins (skipping 6 confirmed false positives/already-fixed, with reasons in the disposition report). Highlights: cross-file contracts (goldens fingerprint persistence, paired-arena stage numbering, canonical golden-ID field, RERUN resolution before promotion) now match between finetune.md, promote-checkpoint.md, and checkpoint-promotion's templates. TRL API usages (SFTConfig.max_length, trl.experimental ORPO/CPO imports) verified against live TRL docs rather than blindly renamed. Several runnable examples hardened against real failure modes: malformed judge output, empty arena results, non-distinct DPO pairs, unbounded rejection-sampling fan-out, non-deterministic smoke-test comparisons, silent FP8-to-bf16 fallback, and orphaned background thermal-sampler processes. Container detection (G9) and LoRA adapter-size math fixed in both dgx-spark-ops and llm-finetuning where the same bugs were independently present in each plugin. Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: apply dogfood friction-log remediations (F1-F32) from DGX Spark run Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: address CodeRabbit round-2 findings (digest pins, suite-size math, monkeypatch scoping) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf
5.3 KiB
claude-agents — multi-harness agentic plugin marketplace
Production-ready agentic-workflow building blocks: 94 plugins (90 local + 4 external), 203 agents, 175 skills, 109 commands. Native source-of-truth for Claude Code; also consumed by OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source.
This file is the canonical context file. Codex / Cursor / OpenCode read it directly. Claude Code reads it via CLAUDE.md, a symlink to this file. Gemini CLI reads it via gemini-extension.json (contextFileName) / .gemini/settings.json.
Read this file like a table of contents. Detail lives in
docs/. Authoring conventions live indocs/authoring.md. Per-harness setup and capability deltas live indocs/harnesses.md. Gemini-specific setup is inGEMINI.md(also auto-loaded by Gemini CLI). This file should never grow beyond ~150 lines (per OpenAI's harness-engineering practice).
Map
- ARCHITECTURE.md — top-level architectural overview (adapter framework, source-of-truth invariant, capability matrix summary)
- docs/architecture.md — detailed design principles
- docs/plugins.md — full plugin catalog (94 plugins by category)
- docs/agents.md — agent reference (203 agents, model tiers)
- docs/agent-skills.md — skill reference (progressive disclosure model)
- docs/usage.md — commands, workflows, examples
- docs/authoring.md — portable-content style guide (read before adding plugins)
- docs/harnesses.md — per-harness capability matrix
- docs/plugin-eval.md — three-layer quality evaluation framework
- docs/round-trip-results.md — real-CLI verification recipes
- CONTRIBUTING.md — how to contribute
Working in this repo
- Python tooling: uv (package manager), ruff (lint/format), ty (type check). Do not use pip / mypy / black.
- Plugins live under
plugins/<name>/with auto-discovery — seedocs/authoring.mdfor frontmatter shapes. - Plugin names: lowercase, hyphen-separated. Never use
__(it's the adapter namespace separator). - Never commit secrets. Never run destructive git (force-push,
reset --hard, branch -D) without explicit ask.
Quality gates (run these before pushing)
make validate STRICT=1 # structural validation across all harness outputs
make garden # drift detection (dead links, stale artifacts, oversize skills)
make test # full pytest suite (plugin-eval + tools/tests/)
make smoke-test # real-CLI subprocess tests against generated artifacts
CI (.github/workflows/validate.yml) runs all four on every PR plus installs OpenCode + Gemini CLI for live verification.
Regenerating per-harness artifacts
make generate HARNESS=codex # .codex/skills, .codex/agents, .codex/plugins/<p>/, .agents/plugins/marketplace.json
make generate HARNESS=cursor # .cursor-plugin/{marketplace,plugin}.json, .cursor/rules/
make generate HARNESS=opencode # .opencode/{skills,agents,commands,plugins}/, opencode.json
make generate HARNESS=gemini # skills/, agents/, commands/ at extension root
make generate-all # all four
Generated artifacts are committed so each harness installs natively from a clone / GitHub URL (native-install commands in docs/harnesses.md). Run make generate-all before committing source changes — CI fails on drift. Source-of-truth lives only under plugins/; never hand-edit generated files.
Skills (cross-harness)
175 skills under plugins/*/skills/<n>/SKILL.md — discoverable by every harness:
- Claude Code: auto-discovery via Anthropic's SKILL.md spec
- Codex CLI: mirrored to
.codex/skills/<plugin>__<skill>/(8 KB body cap; detail inreferences/details.md) - OpenCode: mirrored to
.opencode/skills/<plugin>-<skill>/using hyphenated names for global install - Cursor: reads
.claude/skills/directly (no re-emit) - Gemini CLI: native skills at
skills/<plugin>__<skill>/SKILL.md
Top-level skills/ is Gemini output; do not use it for OpenCode installs.
Subagents (cross-harness)
203 subagents under plugins/*/agents/<name>.md. Per-harness transpilation:
- Codex:
.codex/agents/<plugin>__<agent>.toml(droptools:, map model alias to the GPT-5.x family, infersandbox_mode) - OpenCode:
.opencode/agents/<plugin>__<agent>.mdwithmode: subagent+permission:block (locked agents — those with sourcetools: []— get deny-everything except baseskill/task) - Gemini:
agents/<plugin>__<agent>.md(April 2026 subagent spec) - Cursor: reads
.claude/agents/directly
Why this file is short
Per OpenAI's harness-engineering practice: this file is a map, not an encyclopedia. Procedural detail lives in skills (loaded on demand by agents). Reference material lives in docs/ (loaded when an agent navigates). A single bloated AGENTS.md crowds out the task, rots quickly, and is hard to verify mechanically. Keep it lean; push detail elsewhere.