* first refactor
* feat: dockerized tool with persistent Claude+Codex login + multi-model benchmark
Run the autonomous CTF/pentest tool in Docker with a one-time, persistent login for
BOTH Claude Code and Codex, and add a multi-model benchmark harness.
Backend (multi-model):
- Add `--backend {claude,codex}` to the CTF pipeline. CodexBackend (pentestgpt/core/
backend.py) wraps unified_agent's Codex backend and translates its events into
AgentMessages, so the same pipeline runs on Claude (opus/sonnet) or Codex
(gpt-5.5/gpt-5.4-mini). Wired through config.backend, pipeline stage construction,
and the CLI (+ PENTESTGPT_CODEX_EFFORT; greppable [CODEX_USAGE] under PENTESTGPT_BENCH=1).
Docker tool (tool-only image; the benchmark stays OUTSIDE the image):
- Extend Dockerfile: Codex CLI (@openai/codex) + openai_codex SDK + unified_agent/
pentestgpt_agent/pentestgpt_legacy packages + gobuster/dirb + socat. Add .dockerignore
(keeps creds/benchmark/workspace out of the build context).
- Persistent dual login (the hard part) — asymmetric by token model:
* Claude: `setup-token` -> token stored in the pentestgpt-claude volume; entrypoint
exports CLAUDE_CODE_OAUTH_TOKEN (setup-token does not write .credentials.json; macOS
host creds live in the Keychain and can't be copied).
* Codex: the container does its OWN `codex login` (NOT seeding -- ChatGPT refresh tokens
are single-use, so a shared/copied login 401s on first refresh). The 127.0.0.1:1455
OAuth callback is forwarded into the container via a socat hop (-p 1455:8455).
* scripts/docker-login.sh is idempotent: checks logins live, logs in only the missing one(s).
- docker-compose codex-config volume (+ pinned names); entrypoint token-export + non-blocking
preflight; scripts/docker-auth-status.sh; Make targets (docker-build/login/auth-status/
run/shell/down/nuke).
- Verified end-to-end: one `make docker-login` -> a fresh container reports claude+codex
logged in with live round-trips; the CTF pipeline (Codex) captured a flag against an
isolated fixture and the pentest pipeline ran cleanly; persists across recreation, no re-login.
Benchmark (multi-model, host-side):
- benchmark/pilot/ harness (run_pilot.py + report.py): builds each xbow challenge, discovers
the loopback port, runs the pipeline across the 4 model combos, judges by the baked
FLAG{sha256(UPPER-dir)}, and renders REPORT.md (infra failures excluded from solve rates).
Includes the partial pilot's results (results.jsonl + REPORT.md).
Docs: docs/docker-dev-plan.md (full plan + implementation status); CLAUDE.md and README
docker quickstart; benchmark/pilot/README.md; design-doc roadmap (docs/redesign).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: fail controller on backend error messages
* fix: allow listing sessions without target
* docs: add docker xbow benchmark report
* fix: infer concrete backend constructor type
* docs: refresh docker benchmark documentation
* feat(benchmark): add pure single-agent baseline + pipeline comparison
Add a "pure single agent" benchmark variant -- one bare `claude -p` /
`codex exec` call per target (no pipeline) -- to quantify what the 3-stage
PentestGPT pipeline buys over an un-orchestrated agent on the xbow targets.
- pentestgpt/prompts/stages.py: ctf_single_agent_{system,task}_prompt -- the
pipeline's shared fragments collapsed into ONE turn, so prompt content is
held constant and the only variable is the multi-stage decomposition.
- benchmark/pilot/run_docker_bench.py: docker-network runner
(--variant single|pipeline). Brings the target up, discovers the container's
internal IP+network (skips DB side-cars/ports), docker-runs the tool image on
that network, and scores the ground-truth flag against the agent's *assistant
text* only (parity with the pipeline's raw streaming). Reads stdout in chunks
to handle >64KB JSON lines. Resumable; --dry-run supported.
- benchmark/pilot/report_comparison.py -> DOCKER_COMPARISON.md: head-to-head
pipeline-vs-single per model on the common non-infra set.
- tests/unit/test_single_agent_prompt.py: prompt-builder coverage.
- docs: README, CLAUDE.md, benchmark README, DOCKER_REPORT updated.
Recorded result (10 medium/hard targets x 4 models, container-to-container,
same baseline image digest 0c4c0f3e..., commit dca0019 image):
Model Pipeline Single
Claude Opus 5/10 7/10 (single +2)
Claude Sonnet 6/10 4/10 (pipeline +2)
Codex gpt-5.5 7/10 7/10 (tie)
Codex gpt-5.4-mini 3/10 4/10 (single +1)
TOTAL 21/40 22/40
Single agent matches the pipeline on solve rate (55% vs 52%) while using
~40% fewer Codex tokens (13.0M vs 21.8M) and solving faster. The pipeline
only clearly helps Claude Sonnet (which times out solo); Opus is better solo.
Full per-challenge grid in DOCKER_COMPARISON.md; raw records in
docker_single_results.jsonl.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(benchmark): add pentestgpt_agent docker harness
* bench: refresh pentestgpt_agent smoke result
* fix(benchmark): make repeat rows variant-aware
* fix(agent): fall back for semantic executor labels
* fix(agent): tolerate executor prose evidence
* fix(benchmark): score accepted framework findings
* bench: append partial framework repeat results
* bench: complete framework repeat sweep
* bench: expose framework executor concurrency
* bench: add extended parallel framework sweep
* checkpoint: preserve working agent and benchmark state
* feat: harden durable agent loop and xbow qualification
* fix: reserve an exploit result turn
* docs: record clean xbow qualification
* build: consume unified-agent from the git wrapper repo
Repoint pentestgpt_agent_new's unified-agent dependency from the local
editable path (../../UnifiedAgentPoC, now renamed and gone) to the pinned
git source PentestGPT-Project/UnifedAgentWrapper@d05d21f. Regenerate uv.lock
and update test_dependency.py to assert the external package is installed
from that VCS URL (not the repo-root vendored copy) at version 0.2.0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor: make pentestgpt_agent_new the sole framework
Remove the retired ledger-based pentestgpt_agent package (instructor/executor/
judge) and its orphaned unit + smoke tests. The nested pentestgpt_agent_new
project (Supervisor/Executor over a durable SQLite loop, consuming unified-agent
from the git wrapper) is now the single maintained framework.
Repoint the top-level tooling to it:
- pyproject: drop the pentestgpt-agent console script and pentestgpt_agent from
the wheel packages.
- Makefile: lint/format target parent code only; typecheck/check/ci now run the
nested framework's own gate (ruff, format, mypy, pytest) via test-agent-new /
check-agent-new, so `make check` finally covers it; `make run` delegates to the
pentestgpt-agent-new CLI.
- Dockerfile: stop copying the removed package (kept the build working); note the
framework is not baked into the image yet.
- docker container-health test: import the substrate packages that actually ship.
- CLAUDE.md / AGENT.md: describe the new framework, the git-sourced wrapper, and
the deprioritized benchmark/Docker rewire.
The XBOW `--variant framework` path and docker-bench Makefile targets still point
at the old in-image framework and are left as a pending rewire (benchmarks
deprioritized); the naive `--variant single` path is unaffected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* refactor: rename pentestgpt_agent_new -> pentestgpt_agent
The framework reclaims the clean name now that the old ledger-based package is
gone. Rename the nested project folder, its src package, the distribution
(pentestgpt-agent-new -> pentestgpt-agent) and CLI, and every import/reference in
the package, the umbrella Makefile, the Dockerfile, the docker health test, and
CLAUDE.md / AGENT.md. Regenerate uv.lock. The audit CLI stays pentestgpt-agent-audit;
the git-sourced unified-agent dependency is unchanged. `make check` is green
(108 nested tests). The two historical *_REPORT.md files keep the old name as
dated records.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: extract benchmark harness to sibling xbow-benchmark repo
Move PentestGPT/benchmark/ out to ../xbow-benchmark (its own repo) to keep this
project clean. The harness was decoupled from the framework code (it scores
container output, never imports pentestgpt_agent/unified_agent), so only
operational ties remain and they now live in the sibling repo.
- Remove benchmark/ and the 4 harness unit tests (relocated + repointed there).
- Strip the docker-bench-*/bench-* targets and their config vars from the
Makefile; keep the tool-image lifecycle (docker-build/login/run/...) and add a
help pointer to `make -C ../xbow-benchmark help`.
The sibling repo mounts this checkout read-only (--source-root ../PentestGPT) and
runs the pentestgpt:latest image built here.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: harden autonomous framework and runtime integration
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
4.9 KiB
Docker runtime status
Status: 2026-07-12
This document describes the repository as it exists now. It replaces the original implementation plan, whose phase matrix and in-repo benchmark paths are obsolete.
Current image responsibility
pentestgpt:latest is a disposable pentest-tool and provider-CLI environment. It contains:
- Ubuntu 24.04, Python 3.12, Node 20,
uv, Claude Code, and Codex; - common network/pentest tools such as nmap, gobuster, dirb, netcat, curl, DNS utilities, jq, and ripgrep;
- the root
pentestgpt_legacypackage; - the old root
unified_agentcompatibility copy; - persistent Claude and Codex authentication helpers.
It deliberately excludes benchmark fixtures/results, credentials, workspaces, and run artifacts.
The maintained pentestgpt_agent nested project is not baked into this image. Consequently,
make docker-run deliberately fails fast with a wiring diagnostic instead of invoking an absent
CLI. The temporary qualification runner in pentestgpt_agent/scripts/run_xbow.py works around this
by building the framework and external wrapper wheels on the host, then installing those wheels into
an isolated container environment.
Repository ownership
PentestGPT/ image, auth helpers, framework source, legacy client
UnifedAgentWrapper/ canonical provider-wrapper package
xbow-benchmark/ Docker-network targets, scheduling, scoring, reports
The benchmark harness no longer belongs under this repository. Run benchmark commands from
../xbow-benchmark; do not add result JSONL or target orchestration here.
Persistent provider login
Authentication state lives in named volumes and is never copied into image layers:
pentestgpt-claude -> /home/pentester/.claude
pentestgpt-codex -> /home/pentester/.codex
The providers require different setup paths:
- Claude uses a long-lived
setup-token, stored as.claude/oauth_tokenand exported asCLAUDE_CODE_OAUTH_TOKENby the entrypoint. - Codex performs its own in-container OAuth login. Its callback is forwarded through a
socathop; hostauth.jsonmust not be copied because ChatGPT refresh tokens rotate.
Useful commands:
make docker-build
make docker-login
make docker-auth-status
ROUNDTRIP=1 make docker-auth-status # spends a minimal provider call
make docker-shell
make docker-down # keeps auth volumes
make docker-nuke # deliberately removes auth volumes
The auth-status check is advisory by default. Named volumes are credentials and must be protected like a logged-in workstation.
Isolation contract
Both PentestGPT roles use provider FULL_ACCESS. The container or dedicated attack box is therefore
the blast radius and security boundary. A deployment must:
- contain only authorized target routes;
- avoid mounting unrelated source, home directories, tokens, or host sockets;
- mount run state only when persistence is required;
- treat traces and SQLite state as sensitive;
- tear down the environment after the assessment.
The tool image runs as pentester, which has passwordless sudo. It is isolation from the developer
host only when mounts, capabilities, devices, and networking are deliberately constrained.
Framework-image decision still open
There are two reasonable future shapes:
- Build framework and
unified-agentwheels outside Docker, then copy them into a dedicated runtime image. This matches the proven qualification runner and preserves exact package hashes. - Publish both packages and install pinned releases during the Docker build.
Do not copy the root unified_agent/ directory into the maintained framework. It is version 0.1-era
compatibility code; the agent requires the pinned external 0.2 package.
Whichever shape is selected must make these checks true in a fresh container:
pentestgpt-agent --help
python -c "import pentestgpt_agent, unified_agent; print(unified_agent.__version__)"
Only after that should make docker-run be advertised as supported.
Build cleanup opportunities
These are independent of framework design:
- remove
apt-get upgradefor faster, more reproducible builds; - install
socatin the main apt layer; - combine global npm installs;
- use BuildKit cache mounts for apt, npm, and uv;
- install from lockfiles/wheels before copying frequently changing source;
- remove the root
unified_agentcopy and its SDK dependencies when the public-package cleanup is performed.
Benchmark handoff
The sibling benchmark repository is host-driven: it starts each authorized XBOW fixture, attaches a fresh tool container to the target network, records one result, and tears the target down. Claude jobs may share their token volume; Codex jobs should remain serialized unless refresh-token safety is demonstrated.
Its framework variant is still pending rewire to the renamed pentestgpt_agent. The working wheel
build/install sequence in pentestgpt_agent/scripts/run_xbow.py is the implementation reference for
that transfer.