greydgl--pentestgpt
ab5fbb4d90
* first refactor
* feat: dockerized tool with persistent Claude+Codex login + multi-model benchmark
Run the autonomous CTF/pentest tool in Docker with a one-time, persistent login for
BOTH Claude Code and Codex, and add a multi-model benchmark harness.
Backend (multi-model):
- Add `--backend {claude,codex}` to the CTF pipeline. CodexBackend (pentestgpt/core/
backend.py) wraps unified_agent's Codex backend and translates its events into
AgentMessages, so the same pipeline runs on Claude (opus/sonnet) or Codex
(gpt-5.5/gpt-5.4-mini). Wired through config.backend, pipeline stage construction,
and the CLI (+ PENTESTGPT_CODEX_EFFORT; greppable [CODEX_USAGE] under PENTESTGPT_BENCH=1).
Docker tool (tool-only image; the benchmark stays OUTSIDE the image):
- Extend Dockerfile: Codex CLI (@openai/codex) + openai_codex SDK + unified_agent/
pentestgpt_agent/pentestgpt_legacy packages + gobuster/dirb + socat. Add .dockerignore
(keeps creds/benchmark/workspace out of the build context).
- Persistent dual login (the hard part) — asymmetric by token model:
* Claude: `setup-token` -> token stored in the pentestgpt-claude volume; entrypoint
exports CLAUDE_CODE_OAUTH_TOKEN (setup-token does not write .credentials.json; macOS
host creds live in the Keychain and can't be copied).
* Codex: the container does its OWN `codex login` (NOT seeding -- ChatGPT refresh tokens
are single-use, so a shared/copied login 401s on first refresh). The 127.0.0.1:1455
OAuth callback is forwarded into the container via a socat hop (-p 1455:8455).
* scripts/docker-login.sh is idempotent: checks logins live, logs in only the missing one(s).
- docker-compose codex-config volume (+ pinned names); entrypoint token-export + non-blocking
preflight; scripts/docker-auth-status.sh; Make targets (docker-build/login/auth-status/
run/shell/down/nuke).
- Verified end-to-end: one `make docker-login` -> a fresh container reports claude+codex
logged in with live round-trips; the CTF pipeline (Codex) captured a flag against an
isolated fixture and the pentest pipeline ran cleanly; persists across recreation, no re-login.
Benchmark (multi-model, host-side):
- benchmark/pilot/ harness (run_pilot.py + report.py): builds each xbow challenge, discovers
the loopback port, runs the pipeline across the 4 model combos, judges by the baked
FLAG{sha256(UPPER-dir)}, and renders REPORT.md (infra failures excluded from solve rates).
Includes the partial pilot's results (results.jsonl + REPORT.md).
Docs: docs/docker-dev-plan.md (full plan + implementation status); CLAUDE.md and README
docker quickstart; benchmark/pilot/README.md; design-doc roadmap (docs/redesign).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: fail controller on backend error messages
* fix: allow listing sessions without target
* docs: add docker xbow benchmark report
* fix: infer concrete backend constructor type
* docs: refresh docker benchmark documentation
* feat(benchmark): add pure single-agent baseline + pipeline comparison
Add a "pure single agent" benchmark variant -- one bare `claude -p` /
`codex exec` call per target (no pipeline) -- to quantify what the 3-stage
PentestGPT pipeline buys over an un-orchestrated agent on the xbow targets.
- pentestgpt/prompts/stages.py: ctf_single_agent_{system,task}_prompt -- the
pipeline's shared fragments collapsed into ONE turn, so prompt content is
held constant and the only variable is the multi-stage decomposition.
- benchmark/pilot/run_docker_bench.py: docker-network runner
(--variant single|pipeline). Brings the target up, discovers the container's
internal IP+network (skips DB side-cars/ports), docker-runs the tool image on
that network, and scores the ground-truth flag against the agent's *assistant
text* only (parity with the pipeline's raw streaming). Reads stdout in chunks
to handle >64KB JSON lines. Resumable; --dry-run supported.
- benchmark/pilot/report_comparison.py -> DOCKER_COMPARISON.md: head-to-head
pipeline-vs-single per model on the common non-infra set.
- tests/unit/test_single_agent_prompt.py: prompt-builder coverage.
- docs: README, CLAUDE.md, benchmark README, DOCKER_REPORT updated.
Recorded result (10 medium/hard targets x 4 models, container-to-container,
same baseline image digest 0c4c0f3e..., commit dca0019 image):
Model Pipeline Single
Claude Opus 5/10 7/10 (single +2)
Claude Sonnet 6/10 4/10 (pipeline +2)
Codex gpt-5.5 7/10 7/10 (tie)
Codex gpt-5.4-mini 3/10 4/10 (single +1)
TOTAL 21/40 22/40
Single agent matches the pipeline on solve rate (55% vs 52%) while using
~40% fewer Codex tokens (13.0M vs 21.8M) and solving faster. The pipeline
only clearly helps Claude Sonnet (which times out solo); Opus is better solo.
Full per-challenge grid in DOCKER_COMPARISON.md; raw records in
docker_single_results.jsonl.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(benchmark): add pentestgpt_agent docker harness
* bench: refresh pentestgpt_agent smoke result
* fix(benchmark): make repeat rows variant-aware
* fix(agent): fall back for semantic executor labels
* fix(agent): tolerate executor prose evidence
* fix(benchmark): score accepted framework findings
* bench: append partial framework repeat results
* bench: complete framework repeat sweep
* bench: expose framework executor concurrency
* bench: add extended parallel framework sweep
* checkpoint: preserve working agent and benchmark state
* feat: harden durable agent loop and xbow qualification
* fix: reserve an exploit result turn
* docs: record clean xbow qualification
* build: consume unified-agent from the git wrapper repo
Repoint pentestgpt_agent_new's unified-agent dependency from the local
editable path (../../UnifiedAgentPoC, now renamed and gone) to the pinned
git source PentestGPT-Project/UnifedAgentWrapper@d05d21f. Regenerate uv.lock
and update test_dependency.py to assert the external package is installed
from that VCS URL (not the repo-root vendored copy) at version 0.2.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor: make pentestgpt_agent_new the sole framework
Remove the retired ledger-based pentestgpt_agent package (instructor/executor/
judge) and its orphaned unit + smoke tests. The nested pentestgpt_agent_new
project (Supervisor/Executor over a durable SQLite loop, consuming unified-agent
from the git wrapper) is now the single maintained framework.
Repoint the top-level tooling to it:
- pyproject: drop the pentestgpt-agent console script and pentestgpt_agent from
the wheel packages.
- Makefile: lint/format target parent code only; typecheck/check/ci now run the
nested framework's own gate (ruff, format, mypy, pytest) via test-agent-new /
check-agent-new, so `make check` finally covers it; `make run` delegates to the
pentestgpt-agent-new CLI.
- Dockerfile: stop copying the removed package (kept the build working); note the
framework is not baked into the image yet.
- docker container-health test: import the substrate packages that actually ship.
- CLAUDE.md / AGENT.md: describe the new framework, the git-sourced wrapper, and
the deprioritized benchmark/Docker rewire.
The XBOW `--variant framework` path and docker-bench Makefile targets still point
at the old in-image framework and are left as a pending rewire (benchmarks
deprioritized); the naive `--variant single` path is unaffected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor: rename pentestgpt_agent_new -> pentestgpt_agent
The framework reclaims the clean name now that the old ledger-based package is
gone. Rename the nested project folder, its src package, the distribution
(pentestgpt-agent-new -> pentestgpt-agent) and CLI, and every import/reference in
the package, the umbrella Makefile, the Dockerfile, the docker health test, and
CLAUDE.md / AGENT.md. Regenerate uv.lock. The audit CLI stays pentestgpt-agent-audit;
the git-sourced unified-agent dependency is unchanged. `make check` is green
(108 nested tests). The two historical *_REPORT.md files keep the old name as
dated records.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: extract benchmark harness to sibling xbow-benchmark repo
Move PentestGPT/benchmark/ out to ../xbow-benchmark (its own repo) to keep this
project clean. The harness was decoupled from the framework code (it scores
container output, never imports pentestgpt_agent/unified_agent), so only
operational ties remain and they now live in the sibling repo.
- Remove benchmark/ and the 4 harness unit tests (relocated + repointed there).
- Strip the docker-bench-*/bench-* targets and their config vars from the
Makefile; keep the tool-image lifecycle (docker-build/login/run/...) and add a
help pointer to `make -C ../xbow-benchmark help`.
The sibling repo mounts this checkout read-only (--source-root ../PentestGPT) and
runs the pentestgpt:latest image built here.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: harden autonomous framework and runtime integration
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
168 行
6.7 KiB
Python
168 行
6.7 KiB
Python
"""One tool definition, two agents.
|
|
|
|
Tools are plain Python functions registered on a :class:`ToolRegistry`. Both
|
|
backends launch the same stdio MCP server subprocess
|
|
(``python -m unified_agent.tool_server pkg.mod:REGISTRY``), so the model-visible
|
|
tool names are byte-identical in Claude Code and Codex: ``mcp__<server>__<tool>``.
|
|
|
|
Constraints enforced here (lowest common denominator of both hosts):
|
|
- names match ``[a-z][a-z0-9_]*`` (Codex rewrites ``-`` to ``_`` which would
|
|
desync Claude permission strings, so hyphens are banned),
|
|
- the fully-qualified name fits in 64 bytes (Codex hard cap),
|
|
- the registry lives in an importable module (the server subprocess imports it).
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import importlib
|
|
import inspect
|
|
import os
|
|
import re
|
|
import sys
|
|
from collections.abc import Callable
|
|
from dataclasses import dataclass
|
|
from pathlib import Path
|
|
from typing import Any
|
|
|
|
from .types import ToolRegistryError, ToolServerSpec
|
|
|
|
NAME_RE = re.compile(r"^[a-z][a-z0-9_]*$")
|
|
MAX_FQ_NAME_BYTES = 64
|
|
|
|
|
|
@dataclass(frozen=True)
|
|
class ToolEntry:
|
|
fn: Callable[..., Any]
|
|
name: str
|
|
description: str
|
|
|
|
|
|
class ToolRegistry:
|
|
"""Collects tool functions; served to both agents via the MCP tool server."""
|
|
|
|
def __init__(self, server_name: str = "unified"):
|
|
if not NAME_RE.match(server_name):
|
|
raise ToolRegistryError(
|
|
f"invalid server name {server_name!r}: must match {NAME_RE.pattern}"
|
|
)
|
|
self.server_name = server_name
|
|
self._entries: dict[str, ToolEntry] = {}
|
|
frame = sys._getframe(1)
|
|
self._defining_module: str | None = frame.f_globals.get("__name__")
|
|
|
|
@property
|
|
def tools(self) -> list[ToolEntry]:
|
|
return list(self._entries.values())
|
|
|
|
def tool(
|
|
self, name: str | None = None, description: str | None = None
|
|
) -> Callable[[Callable[..., Any]], Callable[..., Any]]:
|
|
"""Register a (sync or async) function as a tool on both agents.
|
|
|
|
The input schema is derived from type hints by the MCP server; the
|
|
description (here or the docstring's first line) is what both models
|
|
use to decide when to call the tool — make it count.
|
|
"""
|
|
|
|
def decorator(fn: Callable[..., Any]) -> Callable[..., Any]:
|
|
tool_name = name if name is not None else fn.__name__
|
|
if not NAME_RE.match(tool_name or ""):
|
|
raise ToolRegistryError(
|
|
f"invalid tool name {tool_name!r}: must match {NAME_RE.pattern} "
|
|
"(lowercase, digits, underscores; no hyphens)"
|
|
)
|
|
fq = f"mcp__{self.server_name}__{tool_name}"
|
|
if len(fq.encode()) > MAX_FQ_NAME_BYTES:
|
|
raise ToolRegistryError(
|
|
f"fully-qualified tool name {fq!r} exceeds {MAX_FQ_NAME_BYTES} bytes "
|
|
"(Codex hard limit); shorten the tool or server name"
|
|
)
|
|
if tool_name in self._entries:
|
|
raise ToolRegistryError(f"duplicate tool name {tool_name!r}")
|
|
desc = description
|
|
if desc is None:
|
|
doc = inspect.getdoc(fn) or ""
|
|
desc = doc.strip().splitlines()[0].strip() if doc.strip() else ""
|
|
if not desc:
|
|
raise ToolRegistryError(
|
|
f"tool {tool_name!r} needs a description (docstring or description=)"
|
|
)
|
|
self._entries[tool_name] = ToolEntry(fn=fn, name=tool_name, description=desc)
|
|
return fn
|
|
|
|
return decorator
|
|
|
|
|
|
def resolve_registry(spec: str) -> ToolRegistry:
|
|
"""Import a registry from a ``package.module:ATTR`` spec string."""
|
|
module_name, sep, attr = spec.partition(":")
|
|
if not sep or not module_name or not attr:
|
|
raise ToolRegistryError(f"invalid registry spec {spec!r}: expected 'pkg.mod:ATTR'")
|
|
try:
|
|
module = importlib.import_module(module_name)
|
|
except ImportError as e:
|
|
raise ToolRegistryError(f"cannot import module {module_name!r}: {e}") from e
|
|
try:
|
|
registry = getattr(module, attr)
|
|
except AttributeError as e:
|
|
raise ToolRegistryError(f"module {module_name!r} has no attribute {attr!r}") from e
|
|
if not isinstance(registry, ToolRegistry):
|
|
raise ToolRegistryError(f"{spec!r} is not a ToolRegistry (got {type(registry).__name__})")
|
|
return registry
|
|
|
|
|
|
def registry_spec(registry: ToolRegistry) -> str:
|
|
"""Derive the ``pkg.mod:ATTR`` spec for a registry instance."""
|
|
module_name = registry._defining_module
|
|
if not module_name or module_name == "__main__":
|
|
raise ToolRegistryError(
|
|
"the tool registry must be defined in an importable module (not __main__) "
|
|
"so the MCP server subprocess can import it; move it into a module or pass "
|
|
"the 'pkg.mod:ATTR' spec string explicitly"
|
|
)
|
|
module = importlib.import_module(module_name)
|
|
for attr, value in vars(module).items():
|
|
if value is registry:
|
|
return f"{module_name}:{attr}"
|
|
raise ToolRegistryError(
|
|
f"registry not found as a top-level attribute of module {module_name!r}"
|
|
)
|
|
|
|
|
|
def _module_root_dir(module_name: str) -> Path:
|
|
"""Directory that must be on sys.path so ``module_name`` is importable."""
|
|
module = importlib.import_module(module_name)
|
|
module_file = getattr(module, "__file__", None)
|
|
if module_file is None: # namespace pkg / builtin: nothing to add
|
|
raise ToolRegistryError(f"module {module_name!r} has no __file__")
|
|
path = Path(module_file).resolve()
|
|
parts = module_name.count(".") + 1
|
|
levels = parts if path.name == "__init__.py" else parts - 1
|
|
return path.parents[levels] if levels > 0 else path.parent
|
|
|
|
|
|
def build_tool_server_spec(
|
|
tools: ToolRegistry | str, extra_env: dict[str, str] | None = None
|
|
) -> ToolServerSpec:
|
|
"""Build the launch spec both backends use for the shared tool server.
|
|
|
|
The env is passed explicitly because Codex does NOT inherit the parent
|
|
environment into MCP server processes (Claude Code does, but we pass the
|
|
same env there too so both servers run identically).
|
|
"""
|
|
spec = tools if isinstance(tools, str) else registry_spec(tools)
|
|
registry = resolve_registry(spec)
|
|
if not registry.tools:
|
|
raise ToolRegistryError(f"registry {spec!r} has no tools registered")
|
|
|
|
module_name = spec.partition(":")[0]
|
|
unified_agent_root = Path(__file__).resolve().parent.parent
|
|
paths = [str(_module_root_dir(module_name)), str(unified_agent_root)]
|
|
env = dict(extra_env or {})
|
|
if env.get("PYTHONPATH"):
|
|
paths.append(env["PYTHONPATH"])
|
|
env["PYTHONPATH"] = os.pathsep.join(dict.fromkeys(paths))
|
|
|
|
command = [sys.executable, "-m", "unified_agent.tool_server", spec]
|
|
return ToolServerSpec(server_name=registry.server_name, command=command, env=env)
|