Backstop for the gate: previously loop-triage only fired on Harness (E2E) failures, so a red lint or test on master (e.g. the misspell that slipped past because golangci-lint isn't a required check) produced no fix issue. Now triage watches all the gate workflows. - micro loop: `--ci-workflow` accepts a comma-separated list of workflow names, rendered into the triage workflow_run trigger as a YAML array; the issue names the actual failed workflow via github.event.workflow_run.name. (generic CLI) - go-micro: regenerate loop-triage.yml to watch "Harness (E2E)", "Lint", "Run Tests"; generalize the triage prompt beyond the harness (a lint/test failure on master is a real regression to fix, not a flake to ignore). - Docs: update CONTINUOUS_IMPROVEMENT.md triage description. Note: this is defense-in-depth. The primary fix is making golangci-lint a required status check so red lint can't merge in the first place — that stays with the human (branch protection). Claude-Session: https://claude.ai/code/session_01CmdEY7pYmV5zzwCjNJ4ykL Co-authored-by: Claude <noreply@anthropic.com>
1.4 KiB
Triage the failed CI run at RUNURL. It may be the linter (Lint), the unit/integration tests (Run Tests), or the provider-conformance harness (Harness (E2E)).
Read the logs and root-cause each distinct failure. DEDUPE against open issues — if a failure matches an existing issue, comment "recurred" there instead of filing a duplicate.
For each genuine, self-contained defect, file a scoped issue (gh issue create --label codex --label enhancement --title "<scoped fix>" --body "<root cause, where, acceptance criteria>") so the increment loop builds it and the next CI/harness run verifies it. A lint or test failure on master is a real regression — file it so it is fixed promptly; do NOT ignore it.
IGNORE only genuine transient flakes — live-model latency, provider outages, rate limits, network timeouts with no code cause (mostly relevant to the harness). Anything needing a breaking or architectural change: file it as needs-human and describe it, rather than auto-queuing it as a routine fix.
Close this issue (gh issue close __ISSUE__) when triage is done. Open any PR yourself from the shell with gh; do not use the make_pr tool.