项目文件夹

文件
wehub-resource-sync 94057c3d3e
PR Test (NPU) / check-changes (push) Has been cancelled
PR Test (NPU) / pr-gate (push) Has been cancelled
PR Test (NPU) / set-image-config (push) Has been cancelled
PR Test (NPU) / stage-b-test-1-npu-a2 (0) (push) Has been cancelled
PR Test (NPU) / stage-b-test-1-npu-a2 (1) (push) Has been cancelled
PR Test (NPU) / stage-b-test-2-npu-a2 (0) (push) Has been cancelled
PR Test (NPU) / stage-b-test-2-npu-a2 (1) (push) Has been cancelled
PR Test (NPU) / stage-b-test-4-npu-a3 (push) Has been cancelled
PR Test (NPU) / stage-b-test-16-npu-a3 (push) Has been cancelled
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (push) Has been cancelled
PR Test (NPU) / multimodal-gen-test-2-npu-a3 (push) Has been cancelled
PR Test (Arm64) / pr-gate (push) Has been cancelled
PR Test (Arm64) / check-changes (push) Has been cancelled
PR Test (Arm64) / build-test (push) Has been cancelled
PR Test (sgl-router) / gate (push) Has been cancelled
PR Test (sgl-router) / tier-1 — lint (push) Has been cancelled
PR Test (sgl-router) / tier-2 — build + test (push) Has been cancelled
PR Test (sgl-router) / tier-3 — docker (placeholder) (push) Has been cancelled
PR Test (sgl-router) / tier-3 — k8s integration (push) Has been cancelled
PR Test (sgl-router) / tier-3 — e2e (push) Has been cancelled
PR Test (sgl-router) / finish (push) Has been cancelled
PR Test (NPU) / single-node-poc (map[name:qwen3_6_27b_w8a8_1p_in64k_out1k_50ms runner:linux-aarch64-a3-2 test_case:test/registered/ascend/performance/qwen3_6_27b/test_npu_qwen3_6_27b_w8a8_1p_in64k_out1k_50ms.py test_type:perf]) (push) Has been cancelled
PR Test (NPU) / pr-test-npu-finish (push) Has been cancelled
PR Test (Xeon) / pr-gate (push) Has been cancelled
PR Test (Xeon) / check-changes (push) Has been cancelled
PR Test (Xeon) / build-test (, xeon-gnr, base-b-test-cpu) (push) Has been cancelled
PR Test (XPU) / check-changes (push) Has been cancelled
PR Test (XPU) / pr-gate (push) Has been cancelled
PR Test (XPU) / stage-a-test-1-gpu-xpu (push) Has been cancelled
PR Test (XPU) / wait-for-stage-a (push) Has been cancelled
PR Test (XPU) / stage-b-test-1-gpu-xpu (push) Has been cancelled
PR Test (XPU) / finish (push) Has been cancelled
CI Model Inventory / build-inventory (push) Has been cancelled
Lint / lint (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark Compilation Check (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark - Manual Policy (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark - Request Processing (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark Summary (push) Has been cancelled
PR Test (SMG) / build-wheel (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on windows (x86_64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on macos (x86_64 - auto) (push) Has been cancelled
PR Test (SMG) / python-unit-tests (push) Has been cancelled
PR Test (SMG) / unit-tests (push) Has been cancelled
PR Test (SMG) / benchmarks (push) Has been cancelled
PR Test (SMG) / chat-completions (push) Has been cancelled
PR Test (SMG) / chat-completions-4gpu (push) Has been cancelled
PR Test (SMG) / e2e (push) Has been cancelled
PR Test (SMG) / docker-build-test (push) Has been cancelled
PR Test (SMG) / k8s-integration (push) Has been cancelled
PR Test (SMG) / finish (push) Has been cancelled
PR Test (SMG) / summarize-benchmarks (push) Has been cancelled
Release SGLang Model Gateway Docker Image / publish (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on macos (aarch64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (aarch64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (x86_64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (aarch64 - musllinux_1_1) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (x86_64 - musllinux_1_1) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / Build SDist (push) Has been cancelled
Release SGLang Model Gateway to PyPI / Upload to PyPI (push) Has been cancelled
Release SGLang Kernels / build-cu129-matrix (aarch64, 12.9, 3.10, arm-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / build-cu129-matrix (x86_64, 12.9, 3.10, x64-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / release-cu129 (push) Has been cancelled
Release SGLang Kernels / build-cu130-matrix (aarch64, 13.0, 3.10, arm-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / build-cu130-matrix (x86_64, 13.0, 3.10, x64-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / release-cu130 (push) Has been cancelled
Release SGLang Kernels / build-rocm-matrix (3.10, 700) (push) Has been cancelled
Release SGLang Kernels / build-rocm-matrix (3.10, 720) (push) Has been cancelled
Release SGLang Kernels / release-rocm700 (push) Has been cancelled
Release SGLang Kernels / release-rocm720 (push) Has been cancelled
Release SGLang Kernels / build-musa43 (43, 3.10) (push) Has been cancelled
Release SGLang Kernels / release-musa43 (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:38:16 +08:00

185 行
5.8 KiB
Python

"""Unit tests for /v1/loads load snapshot response behavior."""
import asyncio
import os
import tempfile
import unittest
from types import SimpleNamespace
import msgspec.msgpack
from sglang.srt.entrypoints.v1_loads import get_loads
from sglang.srt.managers.load_snapshot import (
HEADER_STRUCT,
MAGIC,
SLOT_LEN_STRUCT,
SLOT_SIZE,
VERSION,
DisaggregationMetrics,
LoadSnapshot,
QueueMetrics,
ShmLoadSnapshotReader,
ShmLoadSnapshotWriter,
slot_offset,
)
from sglang.srt.managers.tokenizer_control_mixin import TokenizerControlMixin
from sglang.test.ci.ci_register import register_cpu_ci
from sglang.test.test_utils import CustomTestCase, maybe_stub_sgl_kernel
maybe_stub_sgl_kernel()
register_cpu_ci(est_time=10, suite="base-a-test-cpu")
def _temp_path() -> str:
fd, path = tempfile.mkstemp()
os.close(fd)
os.unlink(path)
return path
class _FakeTokenizerManager(TokenizerControlMixin):
def __init__(self, reader, dp_size: int):
self.load_snapshot_reader = reader
self.server_args = SimpleNamespace(
dp_size=dp_size,
enable_dp_attention=False,
nnodes=1,
)
def auto_create_handle_loop(self):
pass
class _FakeHttpTokenizerManager:
metrics_collector = None
def __init__(self, loads):
self.loads = loads
async def get_loads(self, include=None, dp_rank=None):
results = []
for load in self.loads:
if dp_rank is not None and load.dp_rank != dp_rank:
continue
results.append(load)
return results
class TestLoadsResponse(CustomTestCase):
def test_response_omits_server_side_aggregate_and_redundant_fields(self):
manager = _FakeHttpTokenizerManager(
[
LoadSnapshot(
dp_rank=0,
num_running_reqs=3,
num_waiting_reqs=2,
num_total_tokens=256,
)
]
)
response = asyncio.run(get_loads(tokenizer_manager=manager))
self.assertNotIn("dp_rank_count", response)
self.assertNotIn("aggregate", response)
self.assertEqual(len(response["loads"]), 1)
self.assertNotIn("num_total_reqs", response["loads"][0])
self.assertEqual(response["loads"][0]["num_running_reqs"], 3)
self.assertEqual(response["loads"][0]["num_waiting_reqs"], 2)
class TestGetLoads(CustomTestCase):
def test_load_snapshot_wire_format_is_msgpack_slots(self):
path = _temp_path()
writer = ShmLoadSnapshotWriter(path, dp_size=2, dp_rank=1)
try:
writer.write(
LoadSnapshot(
dp_rank=1,
num_running_reqs=3,
num_waiting_reqs=2,
token_usage=0.25,
)
)
with open(path, "rb") as f:
data = f.read()
self.assertEqual(len(data), HEADER_STRUCT.size + 2 * SLOT_SIZE)
magic, version, dp_size, slot_size = HEADER_STRUCT.unpack_from(data, 0)
self.assertEqual(magic, MAGIC)
self.assertEqual(version, VERSION)
self.assertEqual(dp_size, 2)
self.assertEqual(slot_size, SLOT_SIZE)
offset = slot_offset(1, slot_size)
(payload_len,) = SLOT_LEN_STRUCT.unpack_from(data, offset)
payload_start = offset + SLOT_LEN_STRUCT.size
payload = data[payload_start : payload_start + payload_len]
decoded = msgspec.msgpack.decode(payload)
self.assertEqual(decoded["dp_rank"], 1)
self.assertEqual(decoded["num_running_reqs"], 3)
self.assertEqual(decoded["num_waiting_reqs"], 2)
self.assertEqual(decoded["token_usage"], 0.25)
finally:
writer.close()
if os.path.exists(path):
os.unlink(path)
def test_reads_snapshot_and_filters_sections(self):
path = _temp_path()
writer = ShmLoadSnapshotWriter(path, dp_size=1, dp_rank=0)
reader = ShmLoadSnapshotReader(path, dp_size=1)
try:
initial_load = reader.read(0)
self.assertIsNotNone(initial_load)
self.assertEqual(initial_load.num_total_tokens, 0)
writer.write(
LoadSnapshot(
dp_rank=0,
timestamp=1.25,
num_running_reqs=3,
num_waiting_reqs=2,
num_used_tokens=128,
num_total_tokens=256,
max_total_num_tokens=4096,
token_usage=0.125,
gen_throughput=99.5,
cache_hit_rate=0.75,
utilization=0.5,
max_running_requests=128,
disaggregation=DisaggregationMetrics(
mode="decode", decode_transfer_queue_reqs=4
),
queues=QueueMetrics(waiting=2, grammar=1, paused=0, retracted=3),
)
)
manager = _FakeTokenizerManager(reader, dp_size=1)
loads = asyncio.run(manager.get_loads(include=["core"], dp_rank=0))
self.assertEqual(len(loads), 1)
self.assertEqual(loads[0].num_total_tokens, 256)
d = loads[0].to_dict({"core"})
self.assertNotIn("disaggregation", d)
self.assertNotIn("queues", d)
loads_all = asyncio.run(manager.get_loads(include=["all"], dp_rank=0))
d_all = loads_all[0].to_dict()
self.assertIn("disaggregation", d_all)
self.assertIn("queues", d_all)
finally:
reader.close()
writer.close()
if os.path.exists(path):
os.unlink(path)
if __name__ == "__main__":
unittest.main()