wshobson--agents
baa5bd7997
* feat(dgx-spark-ops): scaffold plugin and register in marketplace Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add spark-environment-setup skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(dgx-spark-ops): tighten spark-environment-setup per review Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add spark-training-gotchas skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(dgx-spark-ops): align preflight.sh output contract with docs Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add spark-memory-thermal-ops skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(dgx-spark-ops): trim spark-memory-thermal-ops and source 68GB anchor Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add dgx-spark-ops-engineer agent Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(dgx-spark-ops): defer agent facts to skills Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(dgx-spark-ops): add /spark-preflight command Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): scaffold plugin and register in marketplace Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add finetuning-method-selection skill with dated model catalog Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add lora-qlora-recipes skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add preference-optimization skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add grpo-rlvr-training skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): add isolation warning to execution reward Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add vision-sft skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add dataset-curation skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): correct loss-masking example in dataset-curation Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add eval-harness-first skill (Phase 0 gate) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): bring eval-harness-first into line band Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): reclaim byte headroom in eval-harness-first Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add trace-to-training-data skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add checkpoint-promotion skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add quantized-export skill Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add llm-finetuning-architect agent Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add llm-finetuning-training-engineer agent Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add llm-finetuning-eval-engineer agent Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add /finetune phase-gated lifecycle command Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): thread checkpoint path into /finetune Phase 6 Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * feat(llm-finetuning): add /promote-checkpoint command Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): thread checkpoint path via phase5.output in /promote-checkpoint Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix(llm-finetuning): robust checkpoint discovery and goldens fingerprint in re-gate Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * docs: register llm-finetuning and dgx-spark-ops (94 plugins, 203 agents, 175 skills, 109 commands) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * docs(agents): add fine-tuning and spark-ops agent entries (203 agents) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * chore: generate per-harness artifacts for llm-finetuning and dgx-spark-ops Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: final-review cleanups for fine-tuning plugins Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: address PR #624 review feedback (G1 NGC detection, Phase 3 fallback dispatch, registry metadata) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: enforce sandbox boundary in execution grader and reward examples Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: address CodeRabbit review findings on PR #624 Verified and fixed 37 of 43 outstanding CodeRabbit findings across the llm-finetuning and dgx-spark-ops plugins (skipping 6 confirmed false positives/already-fixed, with reasons in the disposition report). Highlights: cross-file contracts (goldens fingerprint persistence, paired-arena stage numbering, canonical golden-ID field, RERUN resolution before promotion) now match between finetune.md, promote-checkpoint.md, and checkpoint-promotion's templates. TRL API usages (SFTConfig.max_length, trl.experimental ORPO/CPO imports) verified against live TRL docs rather than blindly renamed. Several runnable examples hardened against real failure modes: malformed judge output, empty arena results, non-distinct DPO pairs, unbounded rejection-sampling fan-out, non-deterministic smoke-test comparisons, silent FP8-to-bf16 fallback, and orphaned background thermal-sampler processes. Container detection (G9) and LoRA adapter-size math fixed in both dgx-spark-ops and llm-finetuning where the same bugs were independently present in each plugin. Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: apply dogfood friction-log remediations (F1-F32) from DGX Spark run Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf * fix: address CodeRabbit round-2 findings (digest pins, suite-size math, monkeypatch scoping) Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf
77 行
5.3 KiB
Markdown
77 行
5.3 KiB
Markdown
# claude-agents — multi-harness agentic plugin marketplace
|
|
|
|
Production-ready agentic-workflow building blocks: **94 plugins** (90 local + 4 external), **203 agents**, **175 skills**, **109 commands**. Native source-of-truth for Claude Code; also consumed by OpenAI Codex CLI, Cursor, OpenCode, and Gemini CLI from a single Markdown source.
|
|
|
|
This file is the canonical context file. Codex / Cursor / OpenCode read it directly. Claude Code reads it via `CLAUDE.md`, a symlink to this file. Gemini CLI reads it via `gemini-extension.json` (`contextFileName`) / `.gemini/settings.json`.
|
|
|
|
> **Read this file like a table of contents.** Detail lives in `docs/`. Authoring conventions live in `docs/authoring.md`. Per-harness setup and capability deltas live in [`docs/harnesses.md`](docs/harnesses.md). Gemini-specific setup is in `GEMINI.md` (also auto-loaded by Gemini CLI). This file should never grow beyond ~150 lines (per OpenAI's [harness-engineering](https://openai.com/index/harness-engineering/) practice).
|
|
|
|
## Map
|
|
|
|
- **[ARCHITECTURE.md](ARCHITECTURE.md)** — top-level architectural overview (adapter framework, source-of-truth invariant, capability matrix summary)
|
|
- **[docs/architecture.md](docs/architecture.md)** — detailed design principles
|
|
- **[docs/plugins.md](docs/plugins.md)** — full plugin catalog (94 plugins by category)
|
|
- **[docs/agents.md](docs/agents.md)** — agent reference (203 agents, model tiers)
|
|
- **[docs/agent-skills.md](docs/agent-skills.md)** — skill reference (progressive disclosure model)
|
|
- **[docs/usage.md](docs/usage.md)** — commands, workflows, examples
|
|
- **[docs/authoring.md](docs/authoring.md)** — portable-content style guide (read before adding plugins)
|
|
- **[docs/harnesses.md](docs/harnesses.md)** — per-harness capability matrix
|
|
- **[docs/plugin-eval.md](docs/plugin-eval.md)** — three-layer quality evaluation framework
|
|
- **[docs/round-trip-results.md](docs/round-trip-results.md)** — real-CLI verification recipes
|
|
- **[CONTRIBUTING.md](CONTRIBUTING.md)** — how to contribute
|
|
|
|
## Working in this repo
|
|
|
|
- Python tooling: **uv** (package manager), **ruff** (lint/format), **ty** (type check). Do not use pip / mypy / black.
|
|
- Plugins live under `plugins/<name>/` with auto-discovery — see `docs/authoring.md` for frontmatter shapes.
|
|
- Plugin names: lowercase, hyphen-separated. Never use `__` (it's the adapter namespace separator).
|
|
- Never commit secrets. Never run destructive git (force-push, `reset --hard`, branch -D) without explicit ask.
|
|
|
|
## Quality gates (run these before pushing)
|
|
|
|
```bash
|
|
make validate STRICT=1 # structural validation across all harness outputs
|
|
make garden # drift detection (dead links, stale artifacts, oversize skills)
|
|
make test # full pytest suite (plugin-eval + tools/tests/)
|
|
make smoke-test # real-CLI subprocess tests against generated artifacts
|
|
```
|
|
|
|
CI (`.github/workflows/validate.yml`) runs all four on every PR plus installs OpenCode + Gemini CLI for live verification.
|
|
|
|
## Regenerating per-harness artifacts
|
|
|
|
```bash
|
|
make generate HARNESS=codex # .codex/skills, .codex/agents, .codex/plugins/<p>/, .agents/plugins/marketplace.json
|
|
make generate HARNESS=cursor # .cursor-plugin/{marketplace,plugin}.json, .cursor/rules/
|
|
make generate HARNESS=opencode # .opencode/{skills,agents,commands,plugins}/, opencode.json
|
|
make generate HARNESS=gemini # skills/, agents/, commands/ at extension root
|
|
make generate-all # all four
|
|
```
|
|
|
|
Generated artifacts are **committed** so each harness installs natively from a clone / GitHub URL (native-install commands in [`docs/harnesses.md`](docs/harnesses.md)). Run `make generate-all` before committing source changes — CI fails on drift. Source-of-truth lives only under `plugins/`; never hand-edit generated files.
|
|
|
|
## Skills (cross-harness)
|
|
|
|
175 skills under `plugins/*/skills/<n>/SKILL.md` — discoverable by every harness:
|
|
|
|
- **Claude Code**: auto-discovery via Anthropic's SKILL.md spec
|
|
- **Codex CLI**: mirrored to `.codex/skills/<plugin>__<skill>/` (8 KB body cap; detail in `references/details.md`)
|
|
- **OpenCode**: mirrored to `.opencode/skills/<plugin>-<skill>/` using hyphenated names for global install
|
|
- **Cursor**: reads `.claude/skills/` directly (no re-emit)
|
|
- **Gemini CLI**: native skills at `skills/<plugin>__<skill>/SKILL.md`
|
|
|
|
Top-level `skills/` is Gemini output; do not use it for OpenCode installs.
|
|
|
|
## Subagents (cross-harness)
|
|
|
|
203 subagents under `plugins/*/agents/<name>.md`. Per-harness transpilation:
|
|
|
|
- **Codex**: `.codex/agents/<plugin>__<agent>.toml` (drop `tools:`, map model alias to the GPT-5.x family, infer `sandbox_mode`)
|
|
- **OpenCode**: `.opencode/agents/<plugin>__<agent>.md` with `mode: subagent` + `permission:` block (locked agents — those with source `tools: []` — get deny-everything except base `skill`/`task`)
|
|
- **Gemini**: `agents/<plugin>__<agent>.md` (April 2026 subagent spec)
|
|
- **Cursor**: reads `.claude/agents/` directly
|
|
|
|
## Why this file is short
|
|
|
|
Per OpenAI's harness-engineering practice: this file is a **map**, not an encyclopedia. Procedural detail lives in skills (loaded on demand by agents). Reference material lives in `docs/` (loaded when an agent navigates). A single bloated AGENTS.md crowds out the task, rots quickly, and is hard to verify mechanically. Keep it lean; push detail elsewhere.
|