项目文件夹

文件
wehub-resource-sync 23f7624596
ADR-166 MCP Bridge Security Lock / Static-source security lock (push) Failing after 0s
ADR-166 MCP Bridge Security Lock / Compose default binds loopback + Mongo has auth (push) Failing after 2s
CodeQL Advanced / Analyze (rust) (push) Failing after 0s
ADR-166 MCP Bridge Security Lock / plugin-agent-federation bindHost default (push) Failing after 1s
ADR-166 MCP Bridge Security Lock / Runtime behavior — 401 + terminal gate + fail-closed (push) Failing after 4s
business-pods-smoke / smoke (push) Failing after 1s
all-plugins-smoke / smoke-all (push) Failing after 2s
CI/CD Pipeline / Security & Code Quality (push) Failing after 1s
CI/CD Pipeline / Test Suite (ubuntu-latest) (push) Failing after 1s
CI/CD Pipeline / Build & Package (macos-latest) (push) Has been skipped
CI/CD Pipeline / Build & Package (ubuntu-latest) (push) Has been skipped
CI/CD Pipeline / Build & Package (windows-latest) (push) Has been skipped
CI/CD Pipeline / Documentation & Examples (push) Failing after 1s
Clone Tracker (14-day rolling) / Snapshot clones for ruflo ecosystem (push) Failing after 1s
CodeQL Advanced / Analyze (actions) (push) Failing after 1s
CodeQL Advanced / Analyze (javascript-typescript) (push) Failing after 1s
federation-peer-rust / stable-noop (push) Failing after 1s
metaharness-ci / score (push) Failing after 1s
metaharness-ci / router-compat (push) Failing after 0s
metaharness-ci / similarity-tests (push) Failing after 0s
no-agentbbs-smoke / smoke-without-agentbbs (push) Failing after 1s
V3 CI/CD Pipeline / Build V3 (windows-latest) (push) Has been skipped
codex-integration-audit / Codex integration audit (push) Failing after 1s
helpers-manifest-guard / guard (push) Failing after 1s
🔗 Cross-Agent Integration Tests / 🤝 Agent Coordination Tests (push) Has been skipped
🔗 Cross-Agent Integration Tests / 🧠 Memory Sharing Integration (push) Has been skipped
🔗 Cross-Agent Integration Tests / 🛡️ Fault Tolerance Tests (push) Has been skipped
🔗 Cross-Agent Integration Tests / ⚡ Performance Integration Tests (push) Has been skipped
metaharness-ci / mcp-scan (push) Failing after 1s
metaharness-ci / eject-dryrun (push) Failing after 1s
metaharness-ci / metaharness-real-data (push) Failing after 0s
no-cli-optdep-bloat-2561 / guard (push) Failing after 1s
no-metaharness-smoke / smoke-without-metaharness (push) Failing after 1s
no-phantom-agentic-flow-subpath / guard (push) Failing after 1s
🔄 Automated Rollback Manager / 🚨 Failure Detection (push) Failing after 1s
V3 CI/CD Pipeline / Plugin hooks smoke / ubuntu-latest / Node 22 (push) Failing after 1s
V3 CI/CD Pipeline / ruflo-graph-intelligence build + test smoke (#2044, ADR-123) (push) Failing after 1s
CVE Audit Gate / Audit root (critical-blocking) (push) Failing after 2s
cost-tracker-smoke / smoke (push) Failing after 3s
oia-audit-weekly / audit (push) Failing after 2s
ruflo-agent-smoke / ruflo-agent structural smoke (push) Failing after 1s
📊 Status Badges Update / 📊 Update Status Badges (push) Failing after 1s
V3 CI/CD Pipeline / Static regression guards (#2267 YAML + (push) Failing after 1s
V3 CI/CD Pipeline / Test V3 Packages (push) Failing after 0s
V3 CI/CD Pipeline / agent_execute provider routing smoke (#2042) (push) Failing after 0s
CVE Audit Gate / Audit v3 (critical-blocking) (push) Failing after 1s
federation-peer-rust / stable-native (push) Failing after 2s
🔗 Cross-Agent Integration Tests / 🚀 Integration Test Setup (push) Failing after 2s
neural-trader-smoke / runtime-smoke (push) Failing after 1s
V3 CI/CD Pipeline / Build V3 (macos-latest) (push) Has been skipped
V3 CI/CD Pipeline / Build V3 (ubuntu-latest) (push) Has been skipped
V3 CI/CD Pipeline / Type Check V3 (push) Failing after 1s
V3 CI/CD Pipeline / Smoke (no better-sqlite3) / ubuntu-latest / Node 24 (push) Failing after 1s
V3 CI/CD Pipeline / Smoke (no better-sqlite3) / ubuntu-latest / Node 22 (push) Failing after 2s
V3 CI/CD Pipeline / browser rvf create flag smoke (#2015) (push) Failing after 0s
V3 CI/CD Pipeline / Dependency review (#2046) (push) Has been skipped
V3 CI/CD Pipeline / Supply-chain audit (#2046) (push) Failing after 0s
V3 CI/CD Pipeline / witness marker drift smoke (#2021) (push) Failing after 1s
V3 CI/CD Pipeline / neural-trader portfolio CG smoke (#2068, ADR-126 Phase 3) (push) Failing after 1s
V3 CI/CD Pipeline / neural-trader backtest signing smoke (#2068, ADR-126 Phase 4) (push) Failing after 1s
V3 CI/CD Pipeline / kg-extract type-import classification smoke (#2049) (push) Failing after 0s
V3 CI/CD Pipeline / witness verify precondition smoke (#1880) (push) Failing after 2s
V3 CI/CD Pipeline / neural-trader pipeline risk-gate smoke (#2068, ADR-126 Phase 5) (push) Failing after 0s
V3 CI/CD Pipeline / neural-trader feature attribution smoke (#2068, ADR-126 Phase 6) (push) Failing after 0s
V3 CI/CD Pipeline / plugin-registry signature verification smoke (#1922, CWE-347) (push) Failing after 4s
V3 CI/CD Pipeline / memory stats legacy-DB smoke (#2120) (push) Failing after 4s
V3 CI/CD Pipeline / github deprecated actions smoke (#2089, ADR-127 Phase 3) (push) Failing after 1s
V3 CI/CD Pipeline / graph query + pathfinder smoke (ADR-130 P2+P5) (push) Has been skipped
V3 CI/CD Pipeline / graph trajectory hooks smoke (ADR-130 P3) (push) Has been skipped
V3 CI/CD Pipeline / graph plugin adapter smoke (ADR-130 P4) (push) Has been skipped
V3 CI/CD Pipeline / graph benchmark (ADR-130 P6) (push) Has been skipped
V3 CI/CD Pipeline / statusline generator delegation smoke (#2195) (push) Failing after 1s
V3 CI/CD Pipeline / wizard init regression guard (#2206 (push) Failing after 1s
V3 CI/CD Pipeline / memory no-stray-db smoke (ADR-125 P7) (push) Failing after 1s
V3 CI/CD Pipeline / github-safe injection smoke (#2089, ADR-127 Phase 1) (push) Failing after 1s
V3 CI/CD Pipeline / github actions pin smoke (#2089, ADR-127 Phase 1) (push) Failing after 1s
V3 CI/CD Pipeline / github attribution opt-in smoke (#2089, ADR-127 Phase 4) (push) Failing after 1s
V3 CI/CD Pipeline / pre-bash hook safety smoke (#2017) (push) Failing after 1s
V3 CI/CD Pipeline / Memory import smoke / ubuntu-latest (push) Failing after 0s
V3 CI/CD Pipeline / MCP protocol smoke / ubuntu-latest (push) Failing after 2s
V3 CI/CD Pipeline / ruvllm WASM auto-init smoke (#2086) (push) Failing after 4s
V3 CI/CD Pipeline / MCP paired-tool round-trip smoke (#1889) (push) Failing after 1s
V3 CI/CD Pipeline / Plugin package install-safety (#1902/#1903/#1904) (push) Failing after 1s
V3 CI/CD Pipeline / Tool description discoverability (ADR-112) (push) Failing after 3s
V3 CI/CD Pipeline / CLI npx-install smoke (#1147 / (22) (push) Failing after 1s
V3 CI/CD Pipeline / CLI npx-install smoke (#1147 / (24) (push) Failing after 1s
V3 CI/CD Pipeline / Windows hook shim smoke (#2132) / ubuntu-latest (push) Failing after 2s
V3 CI/CD Pipeline / Windows hook execution smoke (#2132) / ubuntu-latest (push) Failing after 1s
V3 CI/CD Pipeline / Windows init hooks smoke (#2132) / ubuntu-latest (push) Failing after 1s
V3 CI/CD Pipeline / Vector-index dimension audit (#1947) (push) Failing after 0s
V3 CI/CD Pipeline / Hook-command install safety (#1921) (push) Failing after 1s
V3 CI/CD Pipeline / ToolOutputGuardrail smoke (ADR-131, (push) Failing after 1s
V3 CI/CD Pipeline / init-bundle invariants smoke (#2095, ADR-128 Phase 5) (push) Failing after 1s
V3 CI/CD Pipeline / wasm provider bridge smoke (ADR-129 P1) (push) Failing after 2s
V3 CI/CD Pipeline / wasm gallery CRUD smoke (ADR-129 P3) (push) Failing after 1s
V3 CI/CD Pipeline / wasm plugin bridge smoke (ADR-129 P4) (push) Failing after 0s
V3 CI/CD Pipeline / wasm compose smoke (ADR-129 P2) (push) Failing after 4s
V3 CI/CD Pipeline / graph schema smoke (ADR-130 P1) (push) Failing after 0s
Validate Marketplace / validate (push) Failing after 1s
🔍 Verification Pipeline / 🚀 Setup Verification (push) Failing after 1s
🔍 Verification Pipeline / 🛡️ Security Verification (push) Has been skipped
🔍 Verification Pipeline / 📝 Code Quality (push) Has been skipped
🔍 Verification Pipeline / 🧪 Test Verification (${{ matrix.os }}, Node ${{ matrix.node }}) (push) Has been skipped
🔍 Verification Pipeline / 🏗️ Build Verification (push) Has been skipped
🔍 Verification Pipeline / 📚 Documentation Verification (push) Has been skipped
CVE Audit Gate / High-severity report (warn only) (push) Has been cancelled
🔄 Automated Rollback Manager / 🔄 Execute Rollback (push) Has been cancelled
🔄 Automated Rollback Manager / ✅ Post-Rollback Verification (push) Has been cancelled
🔄 Automated Rollback Manager / 📊 Rollback Monitoring (push) Has been cancelled
V3 CI/CD Pipeline / Windows init hooks smoke (#2132) / windows-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows hook execution smoke (#2132) / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows hook execution smoke (#2132) / windows-latest (push) Has been cancelled
🔄 Automated Rollback Manager / ⏳ Manual Rollback Approval (push) Has been cancelled
V3 CI/CD Pipeline / MCP protocol smoke / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Memory import smoke / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows hook shim smoke (#2132) / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows hook shim smoke (#2132) / windows-latest (push) Has been cancelled
V3 CI/CD Pipeline / Windows init hooks smoke (#2132) / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Witness verify (signed manifest) / macos-latest (push) Has been cancelled
V3 CI/CD Pipeline / Witness verify (signed manifest) / ubuntu-latest (push) Has been cancelled
V3 CI/CD Pipeline / Witness verify (signed manifest) / windows-latest (push) Has been cancelled
V3 CI/CD Pipeline / Publish to npm (alpha) (push) Has been cancelled
V3 CI/CD Pipeline / Smoke (no better-sqlite3) / macos-latest / Node 22 (push) Has been cancelled
V3 CI/CD Pipeline / Plugin hooks smoke / macos-latest / Node 22 (push) Has been cancelled
CI/CD Pipeline / Deploy & Release (push) Has been cancelled
CI/CD Pipeline / CI Status (push) Has been cancelled
🔗 Cross-Agent Integration Tests / 📊 Integration Test Report (push) Has been cancelled
🔄 Automated Rollback Manager / 🔍 Pre-Rollback Validation (push) Has been cancelled
🔍 Verification Pipeline / ⚡ Performance Verification (push) Has been cancelled
🔍 Verification Pipeline / 📊 Verification Report (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:02:19 +08:00

313 行
14 KiB
JavaScript

#!/usr/bin/env node
/**
* ADR-129 rvagent benchmark suite.
*
* Measures the four performance wins introduced by ADR-129 Phases 1–4:
*
* P1 — Provider routing latency (JsModelProvider wiring).
* Compares echo-stub round-trip (keyless baseline) against provider
* attachment overhead (measured with a fake key that will 401 — we
* capture the latency to the WASM call boundary, not the network hop).
*
* P2 — wasm_agent_compose throughput with N=10/50/100 MCP tools.
* Shows cost of allowlist validation + descriptor building at scale.
*
* P3 — Gallery CRUD throughput: add → list → import → export → remove.
* Validates that the 10 new gallery operations land in <1 ms each.
*
* P4 — Plugin enumeration overhead: wasm_agent_compose with/without
* includePlugins. Measures manifest-lookup cost when the plugin
* directory does not exist (graceful no-op path).
*
* Output: docs/benchmarks/rvagent-baseline.json
*
* Usage:
* node v3/@claude-flow/cli/scripts/bench-rvagent.mjs [--tag=baseline] [--trials=5]
*
* Pattern: mirrors v3/@claude-flow/guidance/scripts/bench-phase-1.mjs.
* Standalone — does NOT require the test runner, only Node ≥20 + the WASM
* package (@ruvector/rvagent-wasm must be installed for realistic numbers;
* the bench gracefully degrades to timing the WASM-unavailable error path).
*/
import { performance } from 'node:perf_hooks';
import { mkdirSync, writeFileSync } from 'node:fs';
import { fileURLToPath } from 'node:url';
import { dirname, resolve } from 'node:path';
const __dirname = dirname(fileURLToPath(import.meta.url));
const REPO_ROOT = resolve(__dirname, '../../../..');
const OUT_DIR = resolve(REPO_ROOT, 'docs', 'benchmarks');
const args = Object.fromEntries(
process.argv.slice(2).map(a => {
const [k, v] = a.replace(/^--/, '').split('=');
return [k, v ?? true];
}),
);
const TAG = args.tag || 'baseline';
const TRIALS = Math.max(3, parseInt(args.trials || '5', 10));
// ─────────────────────────────────────────────────────────────────────────────
// Timing harness — 5-trial median, same as bench-phase-1.mjs
// ─────────────────────────────────────────────────────────────────────────────
/**
* Time N async calls (or sync if fn returns non-Promise), return median ms.
*/
async function bench(name, fn, reps = 20) {
// Warmup — trigger any lazy init so it doesn't inflate the first trial.
for (let i = 0; i < Math.min(3, reps); i++) {
try { await fn(); } catch { /* ignore warmup errors */ }
}
const latencies = [];
for (let t = 0; t < TRIALS; t++) {
const start = performance.now();
for (let i = 0; i < reps; i++) {
try { await fn(); } catch { /* measure call overhead, not success */ }
}
latencies.push((performance.now() - start) / reps);
}
latencies.sort((a, b) => a - b);
const median = latencies[Math.floor(TRIALS / 2)];
const min = latencies[0];
const max = latencies[TRIALS - 1];
return {
name,
trials: TRIALS,
reps,
medianMs: Math.round(median * 1000) / 1000,
minMs: Math.round(min * 1000) / 1000,
maxMs: Math.round(max * 1000) / 1000,
variance: Math.round(((max - min) / (median || 1)) * 1000) / 1000,
};
}
// ─────────────────────────────────────────────────────────────────────────────
// Load the live agent-wasm module (dist preferred, src fallback)
// ─────────────────────────────────────────────────────────────────────────────
// dist layout: dist/src/**/*.js (tsc rootDir=.)
const DIST_SRC = resolve(__dirname, '../dist/src');
let wasmMod;
try {
wasmMod = await import(resolve(DIST_SRC, 'ruvector/agent-wasm.js'));
} catch {
console.warn('[bench] Could not load agent-wasm dist — using stub metrics only.');
wasmMod = null;
}
// Load the wasm tools module to bench wasm_agent_compose handler directly
let composeHandler;
let galleryHandlers = {};
try {
const toolsMod = await import(resolve(DIST_SRC, 'mcp-tools/wasm-agent-tools.js'));
const tools = toolsMod.wasmAgentTools ?? toolsMod.default ?? [];
const findHandler = (name) => tools.find(t => t.name === name)?.handler;
composeHandler = findHandler('wasm_agent_compose');
galleryHandlers = {
add: findHandler('wasm_gallery_add_custom'),
list: findHandler('wasm_gallery_list'),
importFn: findHandler('wasm_gallery_import'),
exportFn: findHandler('wasm_gallery_export'),
remove: findHandler('wasm_gallery_remove_custom'),
categories: findHandler('wasm_gallery_categories'),
listByCategory: findHandler('wasm_gallery_list_by_category'),
loadRvf: findHandler('wasm_gallery_load_rvf'),
active: findHandler('wasm_gallery_active'),
config: findHandler('wasm_gallery_config'),
};
} catch {
// Handlers unavailable — bench will time the missing-module error path.
composeHandler = null;
}
// ─────────────────────────────────────────────────────────────────────────────
// P1 — Provider routing latency
//
// Baseline: echo-stub path (no keys set) — measures raw WASM call overhead.
// Provider: createWasmAgent with a fake ANTHROPIC_API_KEY. The key is 401'd
// by the Anthropic API before any tokens are billed, but the routing
// logic (attachJsModelProvider, JsModelProvider instantiation, provider
// lookup) runs synchronously before the network call. We measure the
// in-process portion only by timing createWasmAgent + promptWasmAgent
// with a deliberately short input and catching the 401 network error.
// ─────────────────────────────────────────────────────────────────────────────
const originalKey = process.env.ANTHROPIC_API_KEY;
console.log('\nADR-129 rvagent benchmark suite');
console.log('================================');
console.log(`tag=${TAG} trials=${TRIALS} node=${process.version}`);
console.log('');
// P1a — echo-stub baseline (no key)
console.log('P1a: provider routing — echo-stub baseline (keyless)...');
delete process.env.ANTHROPIC_API_KEY;
delete process.env.OPENROUTER_API_KEY;
delete process.env.OLLAMA_API_KEY;
const r_p1_echo = await bench('P1-echo-stub: createWasmAgent (keyless)', async () => {
if (!wasmMod) return; // module unavailable — timing the absence
try {
const info = await wasmMod.createWasmAgent({ maxTurns: 1 });
wasmMod.terminateWasmAgent(info.id);
} catch { /* WASM unavailable is expected in keyless CI */ }
}, 5);
// P1b — provider-path: fake key triggers attachJsModelProvider logic
console.log('P1b: provider routing — with fake key (measures attach overhead)...');
process.env.ANTHROPIC_API_KEY = 'sk-ant-bench-fake-key-00000000000000000000000000';
const r_p1_provider = await bench('P1-provider-path: createWasmAgent (fake key)', async () => {
if (!wasmMod) return;
try {
const info = await wasmMod.createWasmAgent({ maxTurns: 1 });
wasmMod.terminateWasmAgent(info.id);
} catch { /* 401 expected */ }
}, 5);
// Restore original key state
if (originalKey !== undefined) {
process.env.ANTHROPIC_API_KEY = originalKey;
} else {
delete process.env.ANTHROPIC_API_KEY;
}
// ─────────────────────────────────────────────────────────────────────────────
// P2 — wasm_agent_compose throughput with N=10/50/100 MCP tools
// ─────────────────────────────────────────────────────────────────────────────
console.log('P2: compose throughput — N=10/50/100 MCP tools...');
function makeMcpToolList(n) {
return Array.from({ length: n }, (_, i) => `memory_search_${i % 30 === 0 ? i : 'search'}`);
}
async function timeCompose(toolCount) {
if (!composeHandler) return null;
const tools = makeMcpToolList(toolCount);
return bench(`P2-compose: ${toolCount} MCP tools`, async () => {
await composeHandler({ mcpTools: tools, skills: [], prompts: [], tools: [] });
}, 10);
}
const r_p2_10 = await timeCompose(10);
const r_p2_50 = await timeCompose(50);
const r_p2_100 = await timeCompose(100);
// ─────────────────────────────────────────────────────────────────────────────
// P3 — Gallery CRUD throughput: full add → list → import → export → remove cycle
// ─────────────────────────────────────────────────────────────────────────────
console.log('P3: gallery CRUD — add/list/import/export/remove cycle...');
const CUSTOM_TEMPLATE = JSON.stringify({
id: 'bench-custom-1',
name: 'bench-template',
description: 'Benchmark custom template',
category: 'testing',
tags: ['bench'],
version: '0.1.0',
author: 'bench',
builtin: false,
tools: [],
prompts: [],
skills: [],
mcp_tools: [],
capabilities: [],
});
const r_p3_cycle = await bench('P3-gallery-crud: add→list→import→export→remove', async () => {
const fns = galleryHandlers;
if (fns.add) await fns.add({ template: JSON.parse(CUSTOM_TEMPLATE) }).catch(() => {});
if (fns.list) await fns.list({}).catch(() => {});
if (fns.importFn) await fns.importFn({ templatesJson: `[${CUSTOM_TEMPLATE}]` }).catch(() => {});
if (fns.exportFn) await fns.exportFn({}).catch(() => {});
if (fns.remove) await fns.remove({ id: 'bench-custom-1' }).catch(() => {});
}, 5);
const r_p3_categories = await bench('P3-gallery-categories: getCategories', async () => {
if (galleryHandlers.categories) await galleryHandlers.categories({}).catch(() => {});
}, 10);
// ─────────────────────────────────────────────────────────────────────────────
// P4 — Plugin enumeration: compose with/without includePlugins
//
// Tests the manifest-lookup path for plugins that don't exist on disk
// (graceful no-op / warning). This measures that skipping absent plugins
// doesn't add significant overhead.
// ─────────────────────────────────────────────────────────────────────────────
console.log('P4: plugin enumeration — compose with/without includePlugins...');
const r_p4_without = await bench('P4-plugin-enum: compose WITHOUT includePlugins', async () => {
if (!composeHandler) return;
await composeHandler({ mcpTools: ['memory_search'], skills: [], prompts: [], tools: [] });
}, 10);
const r_p4_with = await bench('P4-plugin-enum: compose WITH includePlugins (absent plugins)', async () => {
if (!composeHandler) return;
await composeHandler({
mcpTools: ['memory_search'],
includePlugins: ['nonexistent-plugin-a', 'nonexistent-plugin-b'],
skills: [],
prompts: [],
tools: [],
});
}, 10);
// ─────────────────────────────────────────────────────────────────────────────
// Emit results
// ─────────────────────────────────────────────────────────────────────────────
const results = [
r_p1_echo,
r_p1_provider,
r_p2_10,
r_p2_50,
r_p2_100,
r_p3_cycle,
r_p3_categories,
r_p4_without,
r_p4_with,
].filter(Boolean);
const out = {
tag: TAG,
trials: TRIALS,
node: process.version,
platform: `${process.platform}-${process.arch}`,
capturedAt: new Date().toISOString(),
adr: 'ADR-129',
phases: ['P1-provider-routing', 'P2-compose-throughput', 'P3-gallery-crud', 'P4-plugin-enum'],
results,
};
mkdirSync(OUT_DIR, { recursive: true });
const outPath = resolve(OUT_DIR, `rvagent-${TAG}.json`);
writeFileSync(outPath, JSON.stringify(out, null, 2));
console.log(`\nWrote ${outPath}\n`);
const COL_NAME = 52;
const COL_MS = 10;
const COL_VAR = 8;
console.log(`| ${'Benchmark'.padEnd(COL_NAME)} | ${'Median ms'.padStart(COL_MS)} | ${'Variance'.padStart(COL_VAR)} |`);
console.log(`|${'-'.repeat(COL_NAME + 2)}|${'-'.repeat(COL_MS + 2)}|${'-'.repeat(COL_VAR + 2)}|`);
for (const r of results) {
const name = r.name.padEnd(COL_NAME).slice(0, COL_NAME);
const ms = String(r.medianMs).padStart(COL_MS);
const vari = String(r.variance).padStart(COL_VAR);
console.log(`| ${name} | ${ms} | ${vari} |`);
}
console.log('');
console.log('Note: ms values measure in-process overhead including WASM unavailable error path.');
console.log('When @ruvector/rvagent-wasm is not installed, timings reflect the import error cost only.');