项目文件夹

文件
Gelei Deng ab5fbb4d90 feat: ship the durable multi-model autonomous PentestGPT runtime (#493)
* first refactor

* feat: dockerized tool with persistent Claude+Codex login + multi-model benchmark

Run the autonomous CTF/pentest tool in Docker with a one-time, persistent login for
BOTH Claude Code and Codex, and add a multi-model benchmark harness.

Backend (multi-model):
- Add `--backend {claude,codex}` to the CTF pipeline. CodexBackend (pentestgpt/core/
  backend.py) wraps unified_agent's Codex backend and translates its events into
  AgentMessages, so the same pipeline runs on Claude (opus/sonnet) or Codex
  (gpt-5.5/gpt-5.4-mini). Wired through config.backend, pipeline stage construction,
  and the CLI (+ PENTESTGPT_CODEX_EFFORT; greppable [CODEX_USAGE] under PENTESTGPT_BENCH=1).

Docker tool (tool-only image; the benchmark stays OUTSIDE the image):
- Extend Dockerfile: Codex CLI (@openai/codex) + openai_codex SDK + unified_agent/
  pentestgpt_agent/pentestgpt_legacy packages + gobuster/dirb + socat. Add .dockerignore
  (keeps creds/benchmark/workspace out of the build context).
- Persistent dual login (the hard part) — asymmetric by token model:
  * Claude: `setup-token` -> token stored in the pentestgpt-claude volume; entrypoint
    exports CLAUDE_CODE_OAUTH_TOKEN (setup-token does not write .credentials.json; macOS
    host creds live in the Keychain and can't be copied).
  * Codex: the container does its OWN `codex login` (NOT seeding -- ChatGPT refresh tokens
    are single-use, so a shared/copied login 401s on first refresh). The 127.0.0.1:1455
    OAuth callback is forwarded into the container via a socat hop (-p 1455:8455).
  * scripts/docker-login.sh is idempotent: checks logins live, logs in only the missing one(s).
- docker-compose codex-config volume (+ pinned names); entrypoint token-export + non-blocking
  preflight; scripts/docker-auth-status.sh; Make targets (docker-build/login/auth-status/
  run/shell/down/nuke).
- Verified end-to-end: one `make docker-login` -> a fresh container reports claude+codex
  logged in with live round-trips; the CTF pipeline (Codex) captured a flag against an
  isolated fixture and the pentest pipeline ran cleanly; persists across recreation, no re-login.

Benchmark (multi-model, host-side):
- benchmark/pilot/ harness (run_pilot.py + report.py): builds each xbow challenge, discovers
  the loopback port, runs the pipeline across the 4 model combos, judges by the baked
  FLAG{sha256(UPPER-dir)}, and renders REPORT.md (infra failures excluded from solve rates).
  Includes the partial pilot's results (results.jsonl + REPORT.md).

Docs: docs/docker-dev-plan.md (full plan + implementation status); CLAUDE.md and README
docker quickstart; benchmark/pilot/README.md; design-doc roadmap (docs/redesign).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: fail controller on backend error messages

* fix: allow listing sessions without target

* docs: add docker xbow benchmark report

* fix: infer concrete backend constructor type

* docs: refresh docker benchmark documentation

* feat(benchmark): add pure single-agent baseline + pipeline comparison

Add a "pure single agent" benchmark variant -- one bare `claude -p` /
`codex exec` call per target (no pipeline) -- to quantify what the 3-stage
PentestGPT pipeline buys over an un-orchestrated agent on the xbow targets.

- pentestgpt/prompts/stages.py: ctf_single_agent_{system,task}_prompt -- the
  pipeline's shared fragments collapsed into ONE turn, so prompt content is
  held constant and the only variable is the multi-stage decomposition.
- benchmark/pilot/run_docker_bench.py: docker-network runner
  (--variant single|pipeline). Brings the target up, discovers the container's
  internal IP+network (skips DB side-cars/ports), docker-runs the tool image on
  that network, and scores the ground-truth flag against the agent's *assistant
  text* only (parity with the pipeline's raw streaming). Reads stdout in chunks
  to handle >64KB JSON lines. Resumable; --dry-run supported.
- benchmark/pilot/report_comparison.py -> DOCKER_COMPARISON.md: head-to-head
  pipeline-vs-single per model on the common non-infra set.
- tests/unit/test_single_agent_prompt.py: prompt-builder coverage.
- docs: README, CLAUDE.md, benchmark README, DOCKER_REPORT updated.

Recorded result (10 medium/hard targets x 4 models, container-to-container,
same baseline image digest 0c4c0f3e..., commit dca0019 image):

  Model               Pipeline   Single
  Claude Opus           5/10      7/10   (single +2)
  Claude Sonnet         6/10      4/10   (pipeline +2)
  Codex gpt-5.5         7/10      7/10   (tie)
  Codex gpt-5.4-mini    3/10      4/10   (single +1)
  TOTAL                21/40     22/40

Single agent matches the pipeline on solve rate (55% vs 52%) while using
~40% fewer Codex tokens (13.0M vs 21.8M) and solving faster. The pipeline
only clearly helps Claude Sonnet (which times out solo); Opus is better solo.
Full per-challenge grid in DOCKER_COMPARISON.md; raw records in
docker_single_results.jsonl.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(benchmark): add pentestgpt_agent docker harness

* bench: refresh pentestgpt_agent smoke result

* fix(benchmark): make repeat rows variant-aware

* fix(agent): fall back for semantic executor labels

* fix(agent): tolerate executor prose evidence

* fix(benchmark): score accepted framework findings

* bench: append partial framework repeat results

* bench: complete framework repeat sweep

* bench: expose framework executor concurrency

* bench: add extended parallel framework sweep

* checkpoint: preserve working agent and benchmark state

* feat: harden durable agent loop and xbow qualification

* fix: reserve an exploit result turn

* docs: record clean xbow qualification

* build: consume unified-agent from the git wrapper repo

Repoint pentestgpt_agent_new's unified-agent dependency from the local
editable path (../../UnifiedAgentPoC, now renamed and gone) to the pinned
git source PentestGPT-Project/UnifedAgentWrapper@d05d21f. Regenerate uv.lock
and update test_dependency.py to assert the external package is installed
from that VCS URL (not the repo-root vendored copy) at version 0.2.0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: make pentestgpt_agent_new the sole framework

Remove the retired ledger-based pentestgpt_agent package (instructor/executor/
judge) and its orphaned unit + smoke tests. The nested pentestgpt_agent_new
project (Supervisor/Executor over a durable SQLite loop, consuming unified-agent
from the git wrapper) is now the single maintained framework.

Repoint the top-level tooling to it:
- pyproject: drop the pentestgpt-agent console script and pentestgpt_agent from
  the wheel packages.
- Makefile: lint/format target parent code only; typecheck/check/ci now run the
  nested framework's own gate (ruff, format, mypy, pytest) via test-agent-new /
  check-agent-new, so `make check` finally covers it; `make run` delegates to the
  pentestgpt-agent-new CLI.
- Dockerfile: stop copying the removed package (kept the build working); note the
  framework is not baked into the image yet.
- docker container-health test: import the substrate packages that actually ship.
- CLAUDE.md / AGENT.md: describe the new framework, the git-sourced wrapper, and
  the deprioritized benchmark/Docker rewire.

The XBOW `--variant framework` path and docker-bench Makefile targets still point
at the old in-image framework and are left as a pending rewire (benchmarks
deprioritized); the naive `--variant single` path is unaffected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: rename pentestgpt_agent_new -> pentestgpt_agent

The framework reclaims the clean name now that the old ledger-based package is
gone. Rename the nested project folder, its src package, the distribution
(pentestgpt-agent-new -> pentestgpt-agent) and CLI, and every import/reference in
the package, the umbrella Makefile, the Dockerfile, the docker health test, and
CLAUDE.md / AGENT.md. Regenerate uv.lock. The audit CLI stays pentestgpt-agent-audit;
the git-sourced unified-agent dependency is unchanged. `make check` is green
(108 nested tests). The two historical *_REPORT.md files keep the old name as
dated records.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: extract benchmark harness to sibling xbow-benchmark repo

Move PentestGPT/benchmark/ out to ../xbow-benchmark (its own repo) to keep this
project clean. The harness was decoupled from the framework code (it scores
container output, never imports pentestgpt_agent/unified_agent), so only
operational ties remain and they now live in the sibling repo.

- Remove benchmark/ and the 4 harness unit tests (relocated + repointed there).
- Strip the docker-bench-*/bench-* targets and their config vars from the
  Makefile; keep the tool-image lifecycle (docker-build/login/run/...) and add a
  help pointer to `make -C ../xbow-benchmark help`.

The sibling repo mounts this checkout read-only (--source-root ../PentestGPT) and
runs the pentestgpt:latest image built here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: harden autonomous framework and runtime integration

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 16:49:08 +08:00

168 行
6.7 KiB
Python

"""One tool definition, two agents.
Tools are plain Python functions registered on a :class:`ToolRegistry`. Both
backends launch the same stdio MCP server subprocess
(``python -m unified_agent.tool_server pkg.mod:REGISTRY``), so the model-visible
tool names are byte-identical in Claude Code and Codex: ``mcp__<server>__<tool>``.
Constraints enforced here (lowest common denominator of both hosts):
- names match ``[a-z][a-z0-9_]*`` (Codex rewrites ``-`` to ``_`` which would
desync Claude permission strings, so hyphens are banned),
- the fully-qualified name fits in 64 bytes (Codex hard cap),
- the registry lives in an importable module (the server subprocess imports it).
"""
from __future__ import annotations
import importlib
import inspect
import os
import re
import sys
from collections.abc import Callable
from dataclasses import dataclass
from pathlib import Path
from typing import Any
from .types import ToolRegistryError, ToolServerSpec
NAME_RE = re.compile(r"^[a-z][a-z0-9_]*$")
MAX_FQ_NAME_BYTES = 64
@dataclass(frozen=True)
class ToolEntry:
fn: Callable[..., Any]
name: str
description: str
class ToolRegistry:
"""Collects tool functions; served to both agents via the MCP tool server."""
def __init__(self, server_name: str = "unified"):
if not NAME_RE.match(server_name):
raise ToolRegistryError(
f"invalid server name {server_name!r}: must match {NAME_RE.pattern}"
)
self.server_name = server_name
self._entries: dict[str, ToolEntry] = {}
frame = sys._getframe(1)
self._defining_module: str | None = frame.f_globals.get("__name__")
@property
def tools(self) -> list[ToolEntry]:
return list(self._entries.values())
def tool(
self, name: str | None = None, description: str | None = None
) -> Callable[[Callable[..., Any]], Callable[..., Any]]:
"""Register a (sync or async) function as a tool on both agents.
The input schema is derived from type hints by the MCP server; the
description (here or the docstring's first line) is what both models
use to decide when to call the tool — make it count.
"""
def decorator(fn: Callable[..., Any]) -> Callable[..., Any]:
tool_name = name if name is not None else fn.__name__
if not NAME_RE.match(tool_name or ""):
raise ToolRegistryError(
f"invalid tool name {tool_name!r}: must match {NAME_RE.pattern} "
"(lowercase, digits, underscores; no hyphens)"
)
fq = f"mcp__{self.server_name}__{tool_name}"
if len(fq.encode()) > MAX_FQ_NAME_BYTES:
raise ToolRegistryError(
f"fully-qualified tool name {fq!r} exceeds {MAX_FQ_NAME_BYTES} bytes "
"(Codex hard limit); shorten the tool or server name"
)
if tool_name in self._entries:
raise ToolRegistryError(f"duplicate tool name {tool_name!r}")
desc = description
if desc is None:
doc = inspect.getdoc(fn) or ""
desc = doc.strip().splitlines()[0].strip() if doc.strip() else ""
if not desc:
raise ToolRegistryError(
f"tool {tool_name!r} needs a description (docstring or description=)"
)
self._entries[tool_name] = ToolEntry(fn=fn, name=tool_name, description=desc)
return fn
return decorator
def resolve_registry(spec: str) -> ToolRegistry:
"""Import a registry from a ``package.module:ATTR`` spec string."""
module_name, sep, attr = spec.partition(":")
if not sep or not module_name or not attr:
raise ToolRegistryError(f"invalid registry spec {spec!r}: expected 'pkg.mod:ATTR'")
try:
module = importlib.import_module(module_name)
except ImportError as e:
raise ToolRegistryError(f"cannot import module {module_name!r}: {e}") from e
try:
registry = getattr(module, attr)
except AttributeError as e:
raise ToolRegistryError(f"module {module_name!r} has no attribute {attr!r}") from e
if not isinstance(registry, ToolRegistry):
raise ToolRegistryError(f"{spec!r} is not a ToolRegistry (got {type(registry).__name__})")
return registry
def registry_spec(registry: ToolRegistry) -> str:
"""Derive the ``pkg.mod:ATTR`` spec for a registry instance."""
module_name = registry._defining_module
if not module_name or module_name == "__main__":
raise ToolRegistryError(
"the tool registry must be defined in an importable module (not __main__) "
"so the MCP server subprocess can import it; move it into a module or pass "
"the 'pkg.mod:ATTR' spec string explicitly"
)
module = importlib.import_module(module_name)
for attr, value in vars(module).items():
if value is registry:
return f"{module_name}:{attr}"
raise ToolRegistryError(
f"registry not found as a top-level attribute of module {module_name!r}"
)
def _module_root_dir(module_name: str) -> Path:
"""Directory that must be on sys.path so ``module_name`` is importable."""
module = importlib.import_module(module_name)
module_file = getattr(module, "__file__", None)
if module_file is None: # namespace pkg / builtin: nothing to add
raise ToolRegistryError(f"module {module_name!r} has no __file__")
path = Path(module_file).resolve()
parts = module_name.count(".") + 1
levels = parts if path.name == "__init__.py" else parts - 1
return path.parents[levels] if levels > 0 else path.parent
def build_tool_server_spec(
tools: ToolRegistry | str, extra_env: dict[str, str] | None = None
) -> ToolServerSpec:
"""Build the launch spec both backends use for the shared tool server.
The env is passed explicitly because Codex does NOT inherit the parent
environment into MCP server processes (Claude Code does, but we pass the
same env there too so both servers run identically).
"""
spec = tools if isinstance(tools, str) else registry_spec(tools)
registry = resolve_registry(spec)
if not registry.tools:
raise ToolRegistryError(f"registry {spec!r} has no tools registered")
module_name = spec.partition(":")[0]
unified_agent_root = Path(__file__).resolve().parent.parent
paths = [str(_module_root_dir(module_name)), str(unified_agent_root)]
env = dict(extra_env or {})
if env.get("PYTHONPATH"):
paths.append(env["PYTHONPATH"])
env["PYTHONPATH"] = os.pathsep.join(dict.fromkeys(paths))
command = [sys.executable, "-m", "unified_agent.tool_server", spec]
return ToolServerSpec(server_name=registry.server_name, command=command, env=env)