项目文件夹

文件
wehub-resource-sync 94057c3d3e
PR Test (NPU) / check-changes (push) Has been cancelled
PR Test (NPU) / pr-gate (push) Has been cancelled
PR Test (NPU) / set-image-config (push) Has been cancelled
PR Test (NPU) / stage-b-test-1-npu-a2 (0) (push) Has been cancelled
PR Test (NPU) / stage-b-test-1-npu-a2 (1) (push) Has been cancelled
PR Test (NPU) / stage-b-test-2-npu-a2 (0) (push) Has been cancelled
PR Test (NPU) / stage-b-test-2-npu-a2 (1) (push) Has been cancelled
PR Test (NPU) / stage-b-test-4-npu-a3 (push) Has been cancelled
PR Test (NPU) / stage-b-test-16-npu-a3 (push) Has been cancelled
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (push) Has been cancelled
PR Test (NPU) / multimodal-gen-test-2-npu-a3 (push) Has been cancelled
PR Test (Arm64) / pr-gate (push) Has been cancelled
PR Test (Arm64) / check-changes (push) Has been cancelled
PR Test (Arm64) / build-test (push) Has been cancelled
PR Test (sgl-router) / gate (push) Has been cancelled
PR Test (sgl-router) / tier-1 — lint (push) Has been cancelled
PR Test (sgl-router) / tier-2 — build + test (push) Has been cancelled
PR Test (sgl-router) / tier-3 — docker (placeholder) (push) Has been cancelled
PR Test (sgl-router) / tier-3 — k8s integration (push) Has been cancelled
PR Test (sgl-router) / tier-3 — e2e (push) Has been cancelled
PR Test (sgl-router) / finish (push) Has been cancelled
PR Test (NPU) / single-node-poc (map[name:qwen3_6_27b_w8a8_1p_in64k_out1k_50ms runner:linux-aarch64-a3-2 test_case:test/registered/ascend/performance/qwen3_6_27b/test_npu_qwen3_6_27b_w8a8_1p_in64k_out1k_50ms.py test_type:perf]) (push) Has been cancelled
PR Test (NPU) / pr-test-npu-finish (push) Has been cancelled
PR Test (Xeon) / pr-gate (push) Has been cancelled
PR Test (Xeon) / check-changes (push) Has been cancelled
PR Test (Xeon) / build-test (, xeon-gnr, base-b-test-cpu) (push) Has been cancelled
PR Test (XPU) / check-changes (push) Has been cancelled
PR Test (XPU) / pr-gate (push) Has been cancelled
PR Test (XPU) / stage-a-test-1-gpu-xpu (push) Has been cancelled
PR Test (XPU) / wait-for-stage-a (push) Has been cancelled
PR Test (XPU) / stage-b-test-1-gpu-xpu (push) Has been cancelled
PR Test (XPU) / finish (push) Has been cancelled
CI Model Inventory / build-inventory (push) Has been cancelled
Lint / lint (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark Compilation Check (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark - Manual Policy (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark - Request Processing (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark Summary (push) Has been cancelled
PR Test (SMG) / build-wheel (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on windows (x86_64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on macos (x86_64 - auto) (push) Has been cancelled
PR Test (SMG) / python-unit-tests (push) Has been cancelled
PR Test (SMG) / unit-tests (push) Has been cancelled
PR Test (SMG) / benchmarks (push) Has been cancelled
PR Test (SMG) / chat-completions (push) Has been cancelled
PR Test (SMG) / chat-completions-4gpu (push) Has been cancelled
PR Test (SMG) / e2e (push) Has been cancelled
PR Test (SMG) / docker-build-test (push) Has been cancelled
PR Test (SMG) / k8s-integration (push) Has been cancelled
PR Test (SMG) / finish (push) Has been cancelled
PR Test (SMG) / summarize-benchmarks (push) Has been cancelled
Release SGLang Model Gateway Docker Image / publish (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on macos (aarch64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (aarch64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (x86_64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (aarch64 - musllinux_1_1) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (x86_64 - musllinux_1_1) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / Build SDist (push) Has been cancelled
Release SGLang Model Gateway to PyPI / Upload to PyPI (push) Has been cancelled
Release SGLang Kernels / build-cu129-matrix (aarch64, 12.9, 3.10, arm-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / build-cu129-matrix (x86_64, 12.9, 3.10, x64-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / release-cu129 (push) Has been cancelled
Release SGLang Kernels / build-cu130-matrix (aarch64, 13.0, 3.10, arm-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / build-cu130-matrix (x86_64, 13.0, 3.10, x64-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / release-cu130 (push) Has been cancelled
Release SGLang Kernels / build-rocm-matrix (3.10, 700) (push) Has been cancelled
Release SGLang Kernels / build-rocm-matrix (3.10, 720) (push) Has been cancelled
Release SGLang Kernels / release-rocm700 (push) Has been cancelled
Release SGLang Kernels / release-rocm720 (push) Has been cancelled
Release SGLang Kernels / build-musa43 (43, 3.10) (push) Has been cancelled
Release SGLang Kernels / release-musa43 (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:38:16 +08:00

121 行
3.8 KiB
Python

import pytest
import torch
import sglang.srt.layers.mhc as mhc
from sglang.srt.layers.mhc import mhc_fused_post_pre, mhc_post, mhc_pre
from sglang.test.ci.ci_register import register_cuda_ci
register_cuda_ci(est_time=30, stage="base-b", runner_config="1-gpu-large")
@pytest.mark.parametrize("hidden_size", [4096, 7168])
@pytest.mark.parametrize("num_tokens", [0, 1, 8, 17, 32, 64])
@pytest.mark.parametrize("use_norm", [False, True])
def test_mhc_fused_post_pre_matches_unfused(
monkeypatch, hidden_size, num_tokens, use_norm
):
if not torch.cuda.is_available():
pytest.skip("CUDA is required for TileLang mHC kernels")
monkeypatch.setattr(mhc, "is_dsa_prefill_cp_round_robin_split", lambda: False)
torch.manual_seed(0)
device = torch.device("cuda")
hc_mult = 4
hc_mult3 = hc_mult * 2 + hc_mult * hc_mult
hc_hidden_size = hc_mult * hidden_size
x = torch.randn(num_tokens, hidden_size, device=device, dtype=torch.bfloat16) * 0.1
residual = (
torch.randn(
num_tokens, hc_mult, hidden_size, device=device, dtype=torch.bfloat16
)
* 0.1
)
post_prev = torch.rand(num_tokens, hc_mult, 1, device=device, dtype=torch.float32)
comb_prev = (
torch.rand(num_tokens, hc_mult, hc_mult, device=device, dtype=torch.float32)
* 0.25
)
fn = (
torch.randn(hc_mult3, hc_hidden_size, device=device, dtype=torch.float32) * 0.01
)
hc_scale = torch.tensor([0.5, 0.25, 0.25], device=device, dtype=torch.float32)
hc_base = torch.zeros(hc_mult3, device=device, dtype=torch.float32)
norm_weight = (
torch.ones(hidden_size, device=device, dtype=torch.bfloat16)
if use_norm
else None
)
norm_eps = 1e-6 if use_norm else None
rms_eps = 1e-6
hc_eps = 1e-6
sinkhorn_repeat = 2
residual_ref = post_ref = comb_ref = layer_ref = None
if num_tokens > 0:
residual_ref = mhc_post(x, residual, post_prev, comb_prev)
post_ref, comb_ref, layer_ref = mhc_pre(
residual_ref,
fn,
hc_scale,
hc_base,
rms_eps,
hc_eps,
hc_eps,
2.0,
sinkhorn_repeat,
norm_weight=norm_weight,
norm_eps=norm_eps,
)
residual_out, post_out, comb_out, layer_out = mhc_fused_post_pre(
x,
residual,
post_prev,
comb_prev,
fn,
hc_scale,
hc_base,
rms_eps,
hc_eps,
hc_eps,
2.0,
sinkhorn_repeat,
norm_weight=norm_weight,
norm_eps=norm_eps,
)
torch.cuda.synchronize()
if num_tokens == 0:
assert residual_out.shape == residual.shape
assert post_out.shape == (0, hc_mult, 1)
assert comb_out.shape == (0, hc_mult, hc_mult)
assert layer_out.shape == (0, hidden_size)
assert residual_out.dtype == torch.bfloat16
assert post_out.dtype == torch.float32
assert comb_out.dtype == torch.float32
assert layer_out.dtype == torch.bfloat16
return
assert residual_ref is not None
assert post_ref is not None
assert comb_ref is not None
assert layer_ref is not None
assert residual_out.shape == residual_ref.shape
assert post_out.shape == post_ref.shape
assert comb_out.shape == comb_ref.shape
assert layer_out.shape == layer_ref.shape
torch.testing.assert_close(residual_out, residual_ref, atol=0, rtol=0)
torch.testing.assert_close(post_out, post_ref, atol=1e-3, rtol=1e-3)
torch.testing.assert_close(comb_out, comb_ref, atol=1e-3, rtol=1e-3)
layer_atol = 2e-2 if use_norm else 2e-3
layer_rtol = 2e-2 if use_norm else 2e-3
torch.testing.assert_close(layer_out, layer_ref, atol=layer_atol, rtol=layer_rtol)
if __name__ == "__main__":
import sys
sys.exit(pytest.main([__file__]))