项目文件夹

文件
Gelei Deng ab5fbb4d90 feat: ship the durable multi-model autonomous PentestGPT runtime (#493)
* first refactor

* feat: dockerized tool with persistent Claude+Codex login + multi-model benchmark

Run the autonomous CTF/pentest tool in Docker with a one-time, persistent login for
BOTH Claude Code and Codex, and add a multi-model benchmark harness.

Backend (multi-model):
- Add `--backend {claude,codex}` to the CTF pipeline. CodexBackend (pentestgpt/core/
  backend.py) wraps unified_agent's Codex backend and translates its events into
  AgentMessages, so the same pipeline runs on Claude (opus/sonnet) or Codex
  (gpt-5.5/gpt-5.4-mini). Wired through config.backend, pipeline stage construction,
  and the CLI (+ PENTESTGPT_CODEX_EFFORT; greppable [CODEX_USAGE] under PENTESTGPT_BENCH=1).

Docker tool (tool-only image; the benchmark stays OUTSIDE the image):
- Extend Dockerfile: Codex CLI (@openai/codex) + openai_codex SDK + unified_agent/
  pentestgpt_agent/pentestgpt_legacy packages + gobuster/dirb + socat. Add .dockerignore
  (keeps creds/benchmark/workspace out of the build context).
- Persistent dual login (the hard part) — asymmetric by token model:
  * Claude: `setup-token` -> token stored in the pentestgpt-claude volume; entrypoint
    exports CLAUDE_CODE_OAUTH_TOKEN (setup-token does not write .credentials.json; macOS
    host creds live in the Keychain and can't be copied).
  * Codex: the container does its OWN `codex login` (NOT seeding -- ChatGPT refresh tokens
    are single-use, so a shared/copied login 401s on first refresh). The 127.0.0.1:1455
    OAuth callback is forwarded into the container via a socat hop (-p 1455:8455).
  * scripts/docker-login.sh is idempotent: checks logins live, logs in only the missing one(s).
- docker-compose codex-config volume (+ pinned names); entrypoint token-export + non-blocking
  preflight; scripts/docker-auth-status.sh; Make targets (docker-build/login/auth-status/
  run/shell/down/nuke).
- Verified end-to-end: one `make docker-login` -> a fresh container reports claude+codex
  logged in with live round-trips; the CTF pipeline (Codex) captured a flag against an
  isolated fixture and the pentest pipeline ran cleanly; persists across recreation, no re-login.

Benchmark (multi-model, host-side):
- benchmark/pilot/ harness (run_pilot.py + report.py): builds each xbow challenge, discovers
  the loopback port, runs the pipeline across the 4 model combos, judges by the baked
  FLAG{sha256(UPPER-dir)}, and renders REPORT.md (infra failures excluded from solve rates).
  Includes the partial pilot's results (results.jsonl + REPORT.md).

Docs: docs/docker-dev-plan.md (full plan + implementation status); CLAUDE.md and README
docker quickstart; benchmark/pilot/README.md; design-doc roadmap (docs/redesign).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: fail controller on backend error messages

* fix: allow listing sessions without target

* docs: add docker xbow benchmark report

* fix: infer concrete backend constructor type

* docs: refresh docker benchmark documentation

* feat(benchmark): add pure single-agent baseline + pipeline comparison

Add a "pure single agent" benchmark variant -- one bare `claude -p` /
`codex exec` call per target (no pipeline) -- to quantify what the 3-stage
PentestGPT pipeline buys over an un-orchestrated agent on the xbow targets.

- pentestgpt/prompts/stages.py: ctf_single_agent_{system,task}_prompt -- the
  pipeline's shared fragments collapsed into ONE turn, so prompt content is
  held constant and the only variable is the multi-stage decomposition.
- benchmark/pilot/run_docker_bench.py: docker-network runner
  (--variant single|pipeline). Brings the target up, discovers the container's
  internal IP+network (skips DB side-cars/ports), docker-runs the tool image on
  that network, and scores the ground-truth flag against the agent's *assistant
  text* only (parity with the pipeline's raw streaming). Reads stdout in chunks
  to handle >64KB JSON lines. Resumable; --dry-run supported.
- benchmark/pilot/report_comparison.py -> DOCKER_COMPARISON.md: head-to-head
  pipeline-vs-single per model on the common non-infra set.
- tests/unit/test_single_agent_prompt.py: prompt-builder coverage.
- docs: README, CLAUDE.md, benchmark README, DOCKER_REPORT updated.

Recorded result (10 medium/hard targets x 4 models, container-to-container,
same baseline image digest 0c4c0f3e..., commit dca0019 image):

  Model               Pipeline   Single
  Claude Opus           5/10      7/10   (single +2)
  Claude Sonnet         6/10      4/10   (pipeline +2)
  Codex gpt-5.5         7/10      7/10   (tie)
  Codex gpt-5.4-mini    3/10      4/10   (single +1)
  TOTAL                21/40     22/40

Single agent matches the pipeline on solve rate (55% vs 52%) while using
~40% fewer Codex tokens (13.0M vs 21.8M) and solving faster. The pipeline
only clearly helps Claude Sonnet (which times out solo); Opus is better solo.
Full per-challenge grid in DOCKER_COMPARISON.md; raw records in
docker_single_results.jsonl.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(benchmark): add pentestgpt_agent docker harness

* bench: refresh pentestgpt_agent smoke result

* fix(benchmark): make repeat rows variant-aware

* fix(agent): fall back for semantic executor labels

* fix(agent): tolerate executor prose evidence

* fix(benchmark): score accepted framework findings

* bench: append partial framework repeat results

* bench: complete framework repeat sweep

* bench: expose framework executor concurrency

* bench: add extended parallel framework sweep

* checkpoint: preserve working agent and benchmark state

* feat: harden durable agent loop and xbow qualification

* fix: reserve an exploit result turn

* docs: record clean xbow qualification

* build: consume unified-agent from the git wrapper repo

Repoint pentestgpt_agent_new's unified-agent dependency from the local
editable path (../../UnifiedAgentPoC, now renamed and gone) to the pinned
git source PentestGPT-Project/UnifedAgentWrapper@d05d21f. Regenerate uv.lock
and update test_dependency.py to assert the external package is installed
from that VCS URL (not the repo-root vendored copy) at version 0.2.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: make pentestgpt_agent_new the sole framework

Remove the retired ledger-based pentestgpt_agent package (instructor/executor/
judge) and its orphaned unit + smoke tests. The nested pentestgpt_agent_new
project (Supervisor/Executor over a durable SQLite loop, consuming unified-agent
from the git wrapper) is now the single maintained framework.

Repoint the top-level tooling to it:
- pyproject: drop the pentestgpt-agent console script and pentestgpt_agent from
  the wheel packages.
- Makefile: lint/format target parent code only; typecheck/check/ci now run the
  nested framework's own gate (ruff, format, mypy, pytest) via test-agent-new /
  check-agent-new, so `make check` finally covers it; `make run` delegates to the
  pentestgpt-agent-new CLI.
- Dockerfile: stop copying the removed package (kept the build working); note the
  framework is not baked into the image yet.
- docker container-health test: import the substrate packages that actually ship.
- CLAUDE.md / AGENT.md: describe the new framework, the git-sourced wrapper, and
  the deprioritized benchmark/Docker rewire.

The XBOW `--variant framework` path and docker-bench Makefile targets still point
at the old in-image framework and are left as a pending rewire (benchmarks
deprioritized); the naive `--variant single` path is unaffected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: rename pentestgpt_agent_new -> pentestgpt_agent

The framework reclaims the clean name now that the old ledger-based package is
gone. Rename the nested project folder, its src package, the distribution
(pentestgpt-agent-new -> pentestgpt-agent) and CLI, and every import/reference in
the package, the umbrella Makefile, the Dockerfile, the docker health test, and
CLAUDE.md / AGENT.md. Regenerate uv.lock. The audit CLI stays pentestgpt-agent-audit;
the git-sourced unified-agent dependency is unchanged. `make check` is green
(108 nested tests). The two historical *_REPORT.md files keep the old name as
dated records.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: extract benchmark harness to sibling xbow-benchmark repo

Move PentestGPT/benchmark/ out to ../xbow-benchmark (its own repo) to keep this
project clean. The harness was decoupled from the framework code (it scores
container output, never imports pentestgpt_agent/unified_agent), so only
operational ties remain and they now live in the sibling repo.

- Remove benchmark/ and the 4 harness unit tests (relocated + repointed there).
- Strip the docker-bench-*/bench-* targets and their config vars from the
  Makefile; keep the tool-image lifecycle (docker-build/login/run/...) and add a
  help pointer to `make -C ../xbow-benchmark help`.

The sibling repo mounts this checkout read-only (--source-root ../PentestGPT) and
runs the pentestgpt:latest image built here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: harden autonomous framework and runtime integration

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 16:49:08 +08:00

350 行
12 KiB
Python

import json
from collections.abc import AsyncIterator
from pathlib import Path
import pytest
from unified_agent import AgentEvent, RunOptions, SandboxPolicy, TurnCompleted, UnifiedAgent
from pentestgpt_agent.agents import (
EXECUTOR_SCHEMA,
SUPERVISOR_INSTRUCTIONS,
SUPERVISOR_SCHEMA,
AgentContractError,
Supervisor,
_supervisor_prompt,
parse_supervisor_decision,
)
from pentestgpt_agent.memory import (
AttemptRecord,
AttemptStatus,
MemoryKernel,
ObservationRecord,
RunSnapshot,
RunSpec,
RunStatus,
)
from pentestgpt_agent.plan import TaskKind, TaskRecord, TaskStatus
from pentestgpt_agent.trace import EpisodeRunner, TraceStore
class SupervisorBackend:
name = "scripted"
async def stream(self, prompt: str, opts: RunOptions) -> AsyncIterator[AgentEvent]:
assert opts.sandbox is SandboxPolicy.FULL_ACCESS
assert opts.output_schema is not None
task_schema = opts.output_schema["properties"]["new_tasks"]["items"]
assert task_schema["properties"]["kind"]["enum"] == [
"discover",
"enumerate",
"test",
"exploit",
"verify",
"recover",
]
assert "maxItems" not in opts.output_schema["properties"]["new_tasks"]
assert "speculative backlog" in opts.output_schema["properties"]["new_tasks"]["description"]
assert "uniqueItems" not in task_schema["properties"]["basis_ids"]
assert "newest same-target TEST" in task_schema["properties"]["basis_ids"]["description"]
assert "uniqueItems" not in task_schema["properties"]["depends_on"]
assert "basis-producing task" in task_schema["properties"]["depends_on"]["description"]
assert "canonical evidence" in opts.output_schema["properties"]["finish"]["description"]
assert "finish_basis_ids" in opts.output_schema["required"]
finish_basis_schema = opts.output_schema["properties"]["finish_basis_ids"]
assert "uniqueItems" not in finish_basis_schema
assert "supplied canonical observation IDs" in finish_basis_schema["description"]
yield TurnCompleted(
success=True,
final_text="structured decision",
structured_output={
"base_revision": 0,
"new_tasks": [
{
"id": "discover-http",
"kind": "discover",
"target": "http://127.0.0.1:8080",
"objective": "Inspect the HTTP service.",
"done_when": "The reachable HTTP surface is recorded.",
"basis_ids": [],
"depends_on": [],
}
],
"next_task_id": "discover-http",
"finish": False,
"finish_basis_ids": [],
"summary": "Begin with HTTP discovery.",
},
)
def test_provider_output_schemas_use_the_codex_supported_subset() -> None:
unsupported = {
"format",
"maxItems",
"maxLength",
"maximum",
"minItems",
"minLength",
"minimum",
"multipleOf",
"pattern",
"uniqueItems",
}
def walk(value: object) -> None:
if isinstance(value, dict):
assert not unsupported.intersection(value)
for child in value.values():
walk(child)
elif isinstance(value, list):
for child in value:
walk(child)
walk(SUPERVISOR_SCHEMA)
walk(EXECUTOR_SCHEMA)
@pytest.mark.asyncio
async def test_supervisor_turns_authoritative_state_into_a_typed_decision(
tmp_path: Path,
) -> None:
assert "An exact anomaly is a lead" in SUPERVISOR_INSTRUCTIONS
assert "one syntax-preserving, goal-directed derivative" in SUPERVISOR_INSTRUCTIONS
assert "Use only supplied observation IDs" in SUPERVISOR_INSTRUCTIONS
assert "newest completed TEST observation" in SUPERVISOR_INSTRUCTIONS
assert "Finish only when canonical evidence" in SUPERVISOR_INSTRUCTIONS
assert "one bounded TEST task" in SUPERVISOR_INSTRUCTIONS
assert "Do not create one task per nearby payload" in SUPERVISOR_INSTRUCTIONS
assert "whitespace or argument-shape control" in SUPERVISOR_INSTRUCTIONS
assert "single-token redirection or IFS-style payload" in SUPERVISOR_INSTRUCTIONS
assert "Recent diagnostics are noncanonical" in SUPERVISOR_INSTRUCTIONS
assert "finish_basis_ids is empty unless finish is true" in SUPERVISOR_INSTRUCTIONS
assert "one or more supplied canonical observation IDs" in SUPERVISOR_INSTRUCTIONS
assert "every tool exposed by the provider" in SUPERVISOR_INSTRUCTIONS
assert "file read/write tools" in SUPERVISOR_INSTRUCTIONS
assert "byte-for-byte copy one supplied allowed target" in SUPERVISOR_INSTRUCTIONS
assert "Ports, schemes, vhosts, URLs, and paths belong only in the objective" in (
SUPERVISOR_INSTRUCTIONS
)
snapshot = MemoryKernel(tmp_path / "state.sqlite3").create_run(
RunSpec(
run_id="run-1",
goal="Assess the authorized target.",
allowed_targets=("http://127.0.0.1:8080",),
)
)
agent = UnifiedAgent(
SupervisorBackend(),
workspace=tmp_path / "workspace",
sandbox=SandboxPolicy.FULL_ACCESS,
instructions=SUPERVISOR_INSTRUCTIONS,
)
supervisor = Supervisor(EpisodeRunner(agent, TraceStore(tmp_path / "runs")))
decision = await supervisor.decide(snapshot, episode_id="supervisor-1")
assert decision.base_revision == snapshot.revision
assert decision.next_task_id == "discover-http"
assert decision.finish is False
assert decision.finish_basis_ids == ()
assert len(decision.new_tasks) == 1
assert decision.new_tasks[0].kind is TaskKind.DISCOVER
assert decision.new_tasks[0].target == "http://127.0.0.1:8080"
def test_supervisor_retrieval_keeps_a_bounded_working_set_and_compact_history() -> None:
tasks = tuple(
TaskRecord(
id=f"task-{index}",
kind=TaskKind.TEST,
target="http://target.test/input",
objective=f"Full objective {index}",
done_when=f"Completion condition {index}",
basis_ids=(),
depends_on=(),
status=TaskStatus.DONE,
created_revision=index,
)
for index in range(1, 11)
)
snapshot = RunSnapshot(
run_id="run-1",
goal="Capture the flag.",
allowed_targets=("http://target.test",),
status=RunStatus.RUNNING,
revision=12,
max_attempts_per_task=2,
tasks=tasks,
)
state = json.loads(_supervisor_prompt(snapshot).split("\n\n", 1)[1])
assert [task["id"] for task in state["tasks"]] == [
"task-7",
"task-8",
"task-9",
"task-10",
]
assert state["task_history"]["total_closed"] == 10
assert state["task_history"]["counts_by_status"] == {"done": 10}
assert [task["id"] for task in state["task_history"]["recent"]] == [
"task-3",
"task-4",
"task-5",
"task-6",
]
assert len(state["task_history"]["recent"]) == 4
assert all(task.get("objective") != "Full objective 1" for task in state["tasks"])
def test_supervisor_bounds_observations_diagnostics_and_required_context() -> None:
closed_tasks = tuple(
TaskRecord(
id=task_id,
kind=TaskKind.TEST,
target="http://target.test/input",
objective=f"Objective for {task_id}",
done_when=f"Done condition for {task_id}",
basis_ids=(),
depends_on=(),
status=TaskStatus.DONE,
created_revision=index,
)
for index, task_id in enumerate(
(
"basis-task",
"dependency-task",
"closed-3",
"closed-4",
"closed-5",
"closed-6",
"closed-7",
"closed-8",
),
start=1,
)
)
open_task = TaskRecord(
id="open-task",
kind=TaskKind.EXPLOIT,
target="http://target.test/input",
objective="Use the confirmed primitive.",
done_when="The goal artifact is captured.",
basis_ids=("obs-required",),
depends_on=("basis-task", "dependency-task"),
status=TaskStatus.READY,
created_revision=9,
)
observations = tuple(
ObservationRecord(
id=observation_id,
task_id="basis-task" if observation_id == "obs-required" else "closed-8",
attempt_id=f"attempt-{index}",
statement=f"Evidence {observation_id}",
trace_episode_id=f"executor-{index}",
evidence_sequences=(1,),
created_revision=index,
)
for index, observation_id in enumerate(
("obs-required", *(f"obs-{number}" for number in range(2, 10))),
start=1,
)
)
attempts = tuple(
AttemptRecord(
id=f"attempt-{index}",
task_id="closed-8",
status=status,
started_revision=index,
finished_revision=index + 1,
trace_episode_id=f"executor-{index}",
summary=summary,
failure_kind="provider" if status is AttemptStatus.ERROR else None,
failure_message="provider failed" if status is AttemptStatus.ERROR else None,
)
for index, (status, summary) in enumerate(
(
(AttemptStatus.FAILED, "old failure"),
(AttemptStatus.PROGRESS, "more work remains"),
(AttemptStatus.BLOCKED, "missing prerequisite"),
(AttemptStatus.DONE, "successful DONE summary must stay out"),
(AttemptStatus.ERROR, "provider error"),
(AttemptStatus.FAILED, "latest failure"),
),
start=1,
)
)
snapshot = RunSnapshot(
run_id="run-1",
goal="Capture the flag.",
allowed_targets=("http://target.test",),
status=RunStatus.RUNNING,
revision=20,
max_attempts_per_task=2,
tasks=(*closed_tasks, open_task),
observations=observations,
attempts=attempts,
)
state = json.loads(_supervisor_prompt(snapshot).split("\n\n", 1)[1])
assert [task["id"] for task in state["tasks"]] == [
"closed-5",
"closed-6",
"closed-7",
"closed-8",
"open-task",
]
assert [task["id"] for task in state["required_task_context"]] == [
"basis-task",
"dependency-task",
]
assert [observation["id"] for observation in state["observations"]] == [
"obs-required",
"obs-4",
"obs-5",
"obs-6",
"obs-7",
"obs-8",
"obs-9",
]
assert [diagnostic["status"] for diagnostic in state["recent_diagnostics"]] == [
"progress",
"blocked",
"error",
"failed",
]
assert "successful DONE summary must stay out" not in json.dumps(state)
def test_supervisor_prompt_bounds_optional_validation_feedback() -> None:
snapshot = RunSnapshot(
run_id="run-1",
goal="Capture the flag.",
allowed_targets=("http://target.test",),
status=RunStatus.RUNNING,
revision=2,
max_attempts_per_task=2,
)
state_without_feedback = json.loads(_supervisor_prompt(snapshot).split("\n\n", 1)[1])
state_with_feedback = json.loads(
_supervisor_prompt(snapshot, feedback="x" * 2_000).split("\n\n", 1)[1]
)
assert "validation_feedback" not in state_without_feedback
assert state_with_feedback["validation_feedback"] == "x" * 1_000
def test_live_supervisor_output_requires_finish_basis_ids() -> None:
raw = {
"base_revision": 2,
"new_tasks": [],
"next_task_id": "ready-task",
"finish": False,
"summary": "Continue the ready task.",
}
with pytest.raises(AgentContractError, match="unexpected or missing fields"):
parse_supervisor_decision(raw)