项目文件夹

文件
wehub-resource-sync 94057c3d3e
PR Test (NPU) / check-changes (push) Has been cancelled
PR Test (NPU) / pr-gate (push) Has been cancelled
PR Test (NPU) / set-image-config (push) Has been cancelled
PR Test (NPU) / stage-b-test-1-npu-a2 (0) (push) Has been cancelled
PR Test (NPU) / stage-b-test-1-npu-a2 (1) (push) Has been cancelled
PR Test (NPU) / stage-b-test-2-npu-a2 (0) (push) Has been cancelled
PR Test (NPU) / stage-b-test-2-npu-a2 (1) (push) Has been cancelled
PR Test (NPU) / stage-b-test-4-npu-a3 (push) Has been cancelled
PR Test (NPU) / stage-b-test-16-npu-a3 (push) Has been cancelled
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (push) Has been cancelled
PR Test (NPU) / multimodal-gen-test-2-npu-a3 (push) Has been cancelled
PR Test (Arm64) / pr-gate (push) Has been cancelled
PR Test (Arm64) / check-changes (push) Has been cancelled
PR Test (Arm64) / build-test (push) Has been cancelled
PR Test (sgl-router) / gate (push) Has been cancelled
PR Test (sgl-router) / tier-1 — lint (push) Has been cancelled
PR Test (sgl-router) / tier-2 — build + test (push) Has been cancelled
PR Test (sgl-router) / tier-3 — docker (placeholder) (push) Has been cancelled
PR Test (sgl-router) / tier-3 — k8s integration (push) Has been cancelled
PR Test (sgl-router) / tier-3 — e2e (push) Has been cancelled
PR Test (sgl-router) / finish (push) Has been cancelled
PR Test (NPU) / single-node-poc (map[name:qwen3_6_27b_w8a8_1p_in64k_out1k_50ms runner:linux-aarch64-a3-2 test_case:test/registered/ascend/performance/qwen3_6_27b/test_npu_qwen3_6_27b_w8a8_1p_in64k_out1k_50ms.py test_type:perf]) (push) Has been cancelled
PR Test (NPU) / pr-test-npu-finish (push) Has been cancelled
PR Test (Xeon) / pr-gate (push) Has been cancelled
PR Test (Xeon) / check-changes (push) Has been cancelled
PR Test (Xeon) / build-test (, xeon-gnr, base-b-test-cpu) (push) Has been cancelled
PR Test (XPU) / check-changes (push) Has been cancelled
PR Test (XPU) / pr-gate (push) Has been cancelled
PR Test (XPU) / stage-a-test-1-gpu-xpu (push) Has been cancelled
PR Test (XPU) / wait-for-stage-a (push) Has been cancelled
PR Test (XPU) / stage-b-test-1-gpu-xpu (push) Has been cancelled
PR Test (XPU) / finish (push) Has been cancelled
CI Model Inventory / build-inventory (push) Has been cancelled
Lint / lint (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark Compilation Check (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark - Manual Policy (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark - Request Processing (push) Has been cancelled
PR Benchmark (SMG Components) / Benchmark Summary (push) Has been cancelled
PR Test (SMG) / build-wheel (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on windows (x86_64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on macos (x86_64 - auto) (push) Has been cancelled
PR Test (SMG) / python-unit-tests (push) Has been cancelled
PR Test (SMG) / unit-tests (push) Has been cancelled
PR Test (SMG) / benchmarks (push) Has been cancelled
PR Test (SMG) / chat-completions (push) Has been cancelled
PR Test (SMG) / chat-completions-4gpu (push) Has been cancelled
PR Test (SMG) / e2e (push) Has been cancelled
PR Test (SMG) / docker-build-test (push) Has been cancelled
PR Test (SMG) / k8s-integration (push) Has been cancelled
PR Test (SMG) / finish (push) Has been cancelled
PR Test (SMG) / summarize-benchmarks (push) Has been cancelled
Release SGLang Model Gateway Docker Image / publish (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on macos (aarch64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (aarch64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (x86_64 - auto) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (aarch64 - musllinux_1_1) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / build on linux (x86_64 - musllinux_1_1) (push) Has been cancelled
Release SGLang Model Gateway to PyPI / Build SDist (push) Has been cancelled
Release SGLang Model Gateway to PyPI / Upload to PyPI (push) Has been cancelled
Release SGLang Kernels / build-cu129-matrix (aarch64, 12.9, 3.10, arm-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / build-cu129-matrix (x86_64, 12.9, 3.10, x64-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / release-cu129 (push) Has been cancelled
Release SGLang Kernels / build-cu130-matrix (aarch64, 13.0, 3.10, arm-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / build-cu130-matrix (x86_64, 13.0, 3.10, x64-kernel-build-node) (push) Has been cancelled
Release SGLang Kernels / release-cu130 (push) Has been cancelled
Release SGLang Kernels / build-rocm-matrix (3.10, 700) (push) Has been cancelled
Release SGLang Kernels / build-rocm-matrix (3.10, 720) (push) Has been cancelled
Release SGLang Kernels / release-rocm700 (push) Has been cancelled
Release SGLang Kernels / release-rocm720 (push) Has been cancelled
Release SGLang Kernels / build-musa43 (43, 3.10) (push) Has been cancelled
Release SGLang Kernels / release-musa43 (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:38:16 +08:00

165 行
4.3 KiB
Python

import itertools
import torch
import triton
import triton.testing
from sgl_kernel import concat_mla_absorb_q as aot_absorb_q
from sgl_kernel import concat_mla_k as aot_k
from sglang.jit_kernel.benchmark.utils import run_benchmark
from sglang.jit_kernel.concat_mla import concat_mla_absorb_q as jit_absorb_q
from sglang.jit_kernel.concat_mla import concat_mla_k as jit_k
from sglang.test.ci.ci_register import register_cuda_ci
from sglang.utils import is_in_ci
register_cuda_ci(
est_time=6, stage="base-b-kernel-benchmark", runner_config="1-gpu-large"
)
IS_CI = is_in_ci()
NUM_LOCAL_HEADS = 128
QK_NOPE_HEAD_DIM = 128
QK_ROPE_HEAD_DIM = 64
K_HEAD_DIM = QK_NOPE_HEAD_DIM + QK_ROPE_HEAD_DIM
A_LAST_DIM = 512
B_LAST_DIM = 64
DTYPE = torch.bfloat16
DEVICE = "cuda"
def aot_concat_mla_k(k, k_nope, k_rope):
aot_k(k, k_nope, k_rope)
def jit_concat_mla_k(k, k_nope, k_rope):
jit_k(k, k_nope, k_rope)
def torch_concat_mla_k(k, k_nope, k_rope):
nope_head_dim = k_nope.shape[-1]
k[:, :, :nope_head_dim] = k_nope
k[:, :, nope_head_dim:] = k_rope.expand(-1, k.shape[1], -1)
def aot_concat_mla_absorb_q(a, b):
return aot_absorb_q(a, b)
def jit_concat_mla_absorb_q(a, b):
return jit_absorb_q(a, b)
def torch_concat_mla_absorb_q(a, b, out):
a_last_dim = a.shape[-1]
out[:, :, :a_last_dim] = a
out[:, :, a_last_dim:] = b
if IS_CI:
NUM_TOKENS_VALS = [256, 1024]
else:
NUM_TOKENS_VALS = [256, 512, 1024, 2048, 4096, 8192, 16384, 32768]
K_LINE_VALS = ["aot", "jit", "torch"]
K_LINE_NAMES = ["SGL AOT Kernel", "SGL JIT Kernel", "PyTorch"]
K_STYLES = [("orange", "-"), ("blue", "--"), ("green", "-.")]
def _create_concat_mla_k_data(num_tokens):
"""Allocate oversized containers and slice to produce non-contiguous tensors."""
k_nope_container = torch.randn(
(num_tokens, NUM_LOCAL_HEADS, QK_NOPE_HEAD_DIM + 128),
dtype=DTYPE,
device=DEVICE,
)
k_nope = k_nope_container[:, :, :QK_NOPE_HEAD_DIM]
k_rope_container = torch.randn(
(num_tokens, 1, 128 + QK_ROPE_HEAD_DIM),
dtype=DTYPE,
device=DEVICE,
)
k_rope = k_rope_container[:, :, -QK_ROPE_HEAD_DIM:]
k = torch.empty(
(num_tokens, NUM_LOCAL_HEADS, K_HEAD_DIM),
dtype=DTYPE,
device=DEVICE,
)
return k, k_nope, k_rope
@triton.testing.perf_report(
triton.testing.Benchmark(
x_names=["num_tokens"],
x_vals=NUM_TOKENS_VALS,
line_arg="provider",
line_vals=K_LINE_VALS,
line_names=K_LINE_NAMES,
styles=K_STYLES,
ylabel="us",
plot_name="concat-mla-k-performance",
args={},
)
)
def bench_concat_mla_k(num_tokens: int, provider: str):
k, k_nope, k_rope = _create_concat_mla_k_data(num_tokens)
FN_MAP = {
"aot": aot_concat_mla_k,
"jit": jit_concat_mla_k,
"torch": torch_concat_mla_k,
}
fn = lambda: FN_MAP[provider](k, k_nope, k_rope)
return run_benchmark(fn)
if IS_CI:
ABSORB_Q_VALS = list(itertools.product([4, 16], [16]))
else:
ABSORB_Q_VALS = list(itertools.product([1, 4, 8, 16, 32], [1, 8, 32, 128]))
Q_LINE_VALS = ["aot", "jit", "torch"]
Q_LINE_NAMES = ["SGL AOT Kernel", "SGL JIT Kernel", "PyTorch"]
Q_STYLES = [("orange", "-"), ("blue", "--"), ("green", "-.")]
@triton.testing.perf_report(
triton.testing.Benchmark(
x_names=["dim_0", "dim_1"],
x_vals=ABSORB_Q_VALS,
line_arg="provider",
line_vals=Q_LINE_VALS,
line_names=Q_LINE_NAMES,
styles=Q_STYLES,
ylabel="us",
plot_name="concat-mla-absorb-q-performance",
args={},
)
)
def bench_concat_mla_absorb_q(dim_0: int, dim_1: int, provider: str):
a = torch.randn(dim_0, dim_1, A_LAST_DIM, dtype=DTYPE, device=DEVICE)
b = torch.randn(dim_0, dim_1, B_LAST_DIM, dtype=DTYPE, device=DEVICE)
if provider == "torch":
out = torch.empty(
dim_0, dim_1, A_LAST_DIM + B_LAST_DIM, dtype=DTYPE, device=DEVICE
)
fn = lambda: torch_concat_mla_absorb_q(a, b, out)
else:
FN_MAP = {
"aot": aot_concat_mla_absorb_q,
"jit": jit_concat_mla_absorb_q,
}
fn = lambda: FN_MAP[provider](a, b)
return run_benchmark(fn)
if __name__ == "__main__":
bench_concat_mla_k.run(print_data=True)
bench_concat_mla_absorb_q.run(print_data=True)