项目文件夹

文件
Seth Hobson baa5bd7997 feat: add llm-finetuning and dgx-spark-ops plugins (eval-gated fine-tuning lifecycle) (#624)
* feat(dgx-spark-ops): scaffold plugin and register in marketplace

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(dgx-spark-ops): add spark-environment-setup skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(dgx-spark-ops): tighten spark-environment-setup per review

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(dgx-spark-ops): add spark-training-gotchas skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(dgx-spark-ops): align preflight.sh output contract with docs

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(dgx-spark-ops): add spark-memory-thermal-ops skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(dgx-spark-ops): trim spark-memory-thermal-ops and source 68GB anchor

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(dgx-spark-ops): add dgx-spark-ops-engineer agent

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(dgx-spark-ops): defer agent facts to skills

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(dgx-spark-ops): add /spark-preflight command

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): scaffold plugin and register in marketplace

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add finetuning-method-selection skill with dated model catalog

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add lora-qlora-recipes skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add preference-optimization skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add grpo-rlvr-training skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(llm-finetuning): add isolation warning to execution reward

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add vision-sft skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add dataset-curation skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(llm-finetuning): correct loss-masking example in dataset-curation

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add eval-harness-first skill (Phase 0 gate)

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(llm-finetuning): bring eval-harness-first into line band

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(llm-finetuning): reclaim byte headroom in eval-harness-first

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add trace-to-training-data skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add checkpoint-promotion skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add quantized-export skill

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add llm-finetuning-architect agent

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add llm-finetuning-training-engineer agent

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add llm-finetuning-eval-engineer agent

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add /finetune phase-gated lifecycle command

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(llm-finetuning): thread checkpoint path into /finetune Phase 6

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* feat(llm-finetuning): add /promote-checkpoint command

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(llm-finetuning): thread checkpoint path via phase5.output in /promote-checkpoint

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix(llm-finetuning): robust checkpoint discovery and goldens fingerprint in re-gate

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* docs: register llm-finetuning and dgx-spark-ops (94 plugins, 203 agents, 175 skills, 109 commands)

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* docs(agents): add fine-tuning and spark-ops agent entries (203 agents)

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* chore: generate per-harness artifacts for llm-finetuning and dgx-spark-ops

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix: final-review cleanups for fine-tuning plugins

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix: address PR #624 review feedback (G1 NGC detection, Phase 3 fallback dispatch, registry metadata)

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix: enforce sandbox boundary in execution grader and reward examples

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix: address CodeRabbit review findings on PR #624

Verified and fixed 37 of 43 outstanding CodeRabbit findings across the
llm-finetuning and dgx-spark-ops plugins (skipping 6 confirmed false
positives/already-fixed, with reasons in the disposition report).

Highlights: cross-file contracts (goldens fingerprint persistence,
paired-arena stage numbering, canonical golden-ID field, RERUN
resolution before promotion) now match between finetune.md,
promote-checkpoint.md, and checkpoint-promotion's templates. TRL API
usages (SFTConfig.max_length, trl.experimental ORPO/CPO imports)
verified against live TRL docs rather than blindly renamed. Several
runnable examples hardened against real failure modes: malformed
judge output, empty arena results, non-distinct DPO pairs, unbounded
rejection-sampling fan-out, non-deterministic smoke-test comparisons,
silent FP8-to-bf16 fallback, and orphaned background thermal-sampler
processes. Container detection (G9) and LoRA adapter-size math fixed
in both dgx-spark-ops and llm-finetuning where the same bugs were
independently present in each plugin.

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix: apply dogfood friction-log remediations (F1-F32) from DGX Spark run

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf

* fix: address CodeRabbit round-2 findings (digest pins, suite-size math, monkeypatch scoping)

Claude-Session: https://claude.ai/code/session_01RsN3Vz5fZRTdVMkNtmxhSf
2026-07-14 13:18:02 -04:00
..