docs(package): include knowledge guide in defense bundle
参赛者:LindseyMei 目标仓库:MetaX-MACA/vLLM-metax(GitLink 主仓:metax-maca/vLLM-metax) 贡献内容:为 MetaX C500(MXC500)补充 3 组缺失的 fused-MoE Triton tuned 配置 PR:MetaX-MACA/vLLM-metax#313(已关闭,未合并) 重新提交 PR:MetaX-MACA/vLLM-metax#323(目标分支 v0.13.0-dev,源分支 feat/moe-c500-configs) 后续主仓 PR:MetaX-MACA/vLLM-metax#319(将 3 个 tuned config 补齐到 releases/v0.13.0) 包更新日期:2026-07-13
v0.13.0-dev
feat/moe-c500-configs
releases/v0.13.0
本文件包记录为 vLLM-MetaX(沐曦 MACA 后端) 做的 MoE(Mixture-of-Experts)Triton 内核配置调优 贡献:在真机 MetaX C500 上定位到 6 个热门 MoE 模型在 MXC500 上无 tuned 配置、回退默认 tile 的性能缺口,通过直接调优 MACA 实际运行的 vllm_metax fused-MoE 内核,给出针对性的两段式 Triton 配置,实现 kernel 级 1.17–3.38x 提速,并通过正确性与加载验证。
vllm_metax
vllm_metax 的 fused-MoE 走 Triton 内核,其 tile 配置按 (H, E, N, device_name[, dtype]) 从 vllm_metax/model_executor/layers/fused_moe/configs/ 加载 tuned JSON;未命中则回退到通用 get_default_config,并打印 Using default MoE config. Performance might be sub-optimal!。
(H, E, N, device_name[, dtype])
vllm_metax/model_executor/layers/fused_moe/configs/
get_default_config
Using default MoE config. Performance might be sub-optimal!
实测缺口:以下常见 MoE shape 在 MXC500 上没有对应的 tuned 配置文件:
Qwen/Qwen1.5-MoE-A2.7B
Qwen2MoeForCausalLM
deepseek-ai/DeepSeek-V2-Lite
DeepseekV2ForCausalLM
Qwen/Qwen3-30B-A3B
Qwen3MoeForCausalLM
deepseek-ai/DeepSeek-V2
mistralai/Mixtral-8x22B
MixtralForCausalLM
deepseek-ai/DeepSeek-V3
DeepSeek-R1
DeepseekV3ForCausalLM
关键点:上游 benchmarks/kernels/benchmark_moe.py 导入的是上游 fused_experts,而 MACA 运行时用的是 vllm_metax 自己的 fused_experts(经 OOT 注册,二者是不同对象,已核实)。因此上游 tuner 调的是错误的内核。
benchmarks/kernels/benchmark_moe.py
fused_experts
本方案自建微基准(moe_tuning/tools/moe_tune.py)直接调优 MACA 内核:
moe_tuning/tools/moe_tune.py
vllm.model_executor.layers.fused_moe.override_config
vllm_metax ... fused_experts()
torch.cuda.synchronize()
BLOCK_SIZE_M/N/K、GROUP_SIZE_M、num_warps、num_stages
stage1
stage2
ACCF32
SPLIT_K
pipeline
scenario
正确性验证:moe_tuning/tools/moe_verify.py 使用 torch.allclose(rtol=2e-2, atol=2e-2) 比较「默认 tile」与「调优 tile」的 MoE 输出,并调用 get_moe_configs() 确认配置文件会被运行时正确加载。
moe_tuning/tools/moe_verify.py
torch.allclose(rtol=2e-2, atol=2e-2)
get_moe_configs()
小 M(decode)保留默认 tile 或近似配置,避免延迟回退;大 M(prefill / 大 batch)获得显著加速。
图:6 个已覆盖 shape 在 MXC500 上的 fused-MoE kernel 加速比(默认 / 调优)。
原始数据:moe_tuning/results/moe_tune_result.json。
moe_tuning/results/moe_tune_result.json
max|Δ|≈1e-3
torch.allclose(rtol/atol=2e-2)=True
get_moe_configs(E, N, None, 0, 0, H)
device_name=MXC500
验证命令:
source /data/workspace/vLLM-metax-Project/env_vllm.sh cp moe_tuning/tools/moe_verify.py /tmp/ && cd /tmp && python3 /tmp/moe_verify.py
benchmark_moe.py
vLLM-metax-Contributions/ ├── README.md # 本文件 ├── LINKS.md # PR / issue 链接汇总 ├── docs/ │ └── PROGRESS_vLLM_metax.md # 路线 A 全栈安装 + 恢复文档 └── moe_tuning/ # 主贡献:MoE Triton 配置调优 ├── README.md ├── configs/ │ ├── H=2048,E=60,N=1408,device_name=MXC500.json # Qwen1.5-MoE │ ├── H=2048,E=64,N=1408,device_name=MXC500.json # DeepSeek-V2-Lite │ ├── H=2048,E=128,N=768,device_name=MXC500.json # Qwen3-30B-A3B │ ├── H=5120,E=64,N=1536,device_name=MXC500.json # DeepSeek-V2 │ ├── H=6144,E=8,N=2048,device_name=MXC500.json # Mixtral-8x22B TP8 │ └── H=7168,E=256,N=256,device_name=MXC500.json # DeepSeek-V3 / R1 TP8 ├── tools/ │ ├── moe_tune.py # MACA 内核微基准调优器 │ └── moe_verify.py # 正确性 + 配置加载验证 └── results/ └── moe_tune_result.json # 原始调优数据
source /data/workspace/vLLM-metax-Project/env_vllm.sh # 调优(无需下载模型,随机权重);注意从不含 vllm/ 子目录的 cwd 运行 cp moe_tuning/tools/moe_tune.py /tmp/ && cd /tmp python3 /tmp/moe_tune.py --H 2048 --E 60 --N 1408 --top-k 4 --Ms 1 8 16 32 64 128 256 512 1024 2048 # 正确性 + 加载验证 cp moe_tuning/tools/moe_verify.py /tmp/ && python3 /tmp/moe_verify.py # 部署到插件:把 configs/*.json 复制到 # vllm_metax/model_executor/layers/fused_moe/configs/ # 或者用 VLLM_TUNED_CONFIG_FOLDER 覆盖测试
FINAL_STAGE_PLAN.md
vllm_metax/attention/ops/triton_decode_attention.py
_fwd_kernel_stage1
_fwd_grouped_kernel_stage1
@triton.autotune
BLOCK_N ∈ {8,16}
BLOCK_N ∈ {16}
num_warps ∈ {1,2}
num_stages ∈ {1,2}
VLLM_MACA_DECODE_ATTN_AUTOTUNE
1
attention_tuning/tools/attn_decode_sweep.py
attention_tuning/README.md
feat(attn): add autotuned triton decode attention tiles for MXC500
vllm_metax/envs.py
USE_PRECOMPILED_KERNEL
MACA_DP_OPT
bool(os.environ.get(...))
bool(int(...))
vllm_metax/v1/attention/backends/triton_attn.py
MIN_LAUNCH_GRID_SIZE_2D
NUM_PAR_SOFTMAX_SEGMENTS
MACA_DP_OPT=0
False
USE_PRECOMPILED_KERNEL=0
moe_tuning/tools/mctlass_benchmark.py
fix(envs): correct bool parsing for MACA env vars and make attn thresholds overridable
torch.zeros(10, device='cuda')
attention_tuning/results/attn_decode_sweep.json
attention_tuning/results/attn_decode_baseline.json
moe_tuning/results/mctlass_ab_summary.json
WORKLOG_2026-08-07.md
MetaX-MACA:v0.13.0-dev
LindseyMei:feat/moe-more-c500-configs
config_list.txt
LindseyMei:feat/moe-c500-configs
you could merge this pr to the dev branch of v0.13.0 (v0.13.0-dev)
MetaX-MACA:releases/v0.13.0
LindseyMei:feat/moe-mxc500-missing-configs-v0.13.0
pyproject.toml
注:PR #313 为原始调优贡献(GitLink 留痕 #221),PR #319 为将配置正式合并到主仓库的后续补齐(GitLink 留痕 #223)。
版权所有:中国计算机学会技术支持:开源发展技术委员会 京ICP备13000930号-9 京公网安备 11010802047560号
vLLM-MetaX MoE Triton 配置调优 —— 参赛文件包
一句话概述
本文件包记录为 vLLM-MetaX(沐曦 MACA 后端) 做的 MoE(Mixture-of-Experts)Triton 内核配置调优 贡献:在真机 MetaX C500 上定位到 6 个热门 MoE 模型在 MXC500 上无 tuned 配置、回退默认 tile 的性能缺口,通过直接调优 MACA 实际运行的
vllm_metaxfused-MoE 内核,给出针对性的两段式 Triton 配置,实现 kernel 级 1.17–3.38x 提速,并通过正确性与加载验证。提交材料
v0.13.0-dev)v0.13.0-devv0.13.0-dev,等待 review / merge)v0.13.0-dev)1. 问题(真实缺口)
vllm_metax的 fused-MoE 走 Triton 内核,其 tile 配置按(H, E, N, device_name[, dtype])从vllm_metax/model_executor/layers/fused_moe/configs/加载 tuned JSON;未命中则回退到通用get_default_config,并打印Using default MoE config. Performance might be sub-optimal!。实测缺口:以下常见 MoE shape 在 MXC500 上没有对应的 tuned 配置文件:
Qwen/Qwen1.5-MoE-A2.7BQwen2MoeForCausalLMdeepseek-ai/DeepSeek-V2-LiteDeepseekV2ForCausalLMQwen/Qwen3-30B-A3BQwen3MoeForCausalLMdeepseek-ai/DeepSeek-V2DeepseekV2ForCausalLMmistralai/Mixtral-8x22BTP8MixtralForCausalLMdeepseek-ai/DeepSeek-V3/DeepSeek-R1TP8DeepseekV3ForCausalLM2. 方法(调优的是真正的 MACA 内核)
关键点:上游
benchmarks/kernels/benchmark_moe.py导入的是上游fused_experts,而 MACA 运行时用的是vllm_metax自己的fused_experts(经 OOT 注册,二者是不同对象,已核实)。因此上游 tuner 调的是错误的内核。本方案自建微基准(
moe_tuning/tools/moe_tune.py)直接调优 MACA 内核:vllm.model_executor.layers.fused_moe.override_config强制指定 tile,计时单次vllm_metax ... fused_experts()(随机权重,因此无需下载模型);torch.cuda.synchronize(),warmup + 多次取中位数;BLOCK_SIZE_M/N/K、GROUP_SIZE_M、num_warps、num_stages;stage1/stage2+ACCF32/SPLIT_K/pipeline/scenario),对称填充(低风险起点)。正确性验证:
moe_tuning/tools/moe_verify.py使用torch.allclose(rtol=2e-2, atol=2e-2)比较「默认 tile」与「调优 tile」的 MoE 输出,并调用get_moe_configs()确认配置文件会被运行时正确加载。3. 结果(kernel 级,真机 sGPU 切片)
Qwen1.5-MoE-A2.7B(H=2048, E=60, N=1408, top_k=4)
DeepSeek-V2-Lite(H=2048, E=64, N=1408, top_k=6)
Qwen3-30B-A3B(H=2048, E=128, N=768, top_k=8)
DeepSeek-V2(H=5120, E=64, N=1536, top_k=6)
Mixtral-8x22B TP8(H=6144, E=8, N=2048, top_k=2)
DeepSeek-V3 / DeepSeek-R1 TP8(H=7168, E=256, N=256, top_k=8)
小 M(decode)保留默认 tile 或近似配置,避免延迟回退;大 M(prefill / 大 batch)获得显著加速。
图:6 个已覆盖 shape 在 MXC500 上的 fused-MoE kernel 加速比(默认 / 调优)。
原始数据:
moe_tuning/results/moe_tune_result.json。4. 正确性与加载验证
max|Δ|≈1e-3,torch.allclose(rtol/atol=2e-2)=True—— 调优只改 tile 选择,不改数学。get_moe_configs(E, N, None, 0, 0, H)均能命中本配置(device_name=MXC500),返回全部 10 个 M-条目 → 运行时会被正确采用。验证命令:
5. 重要说明(诚实边界)
moe_tuning/tools/moe_tune.py复核。benchmark_moe.py同类证据,是 MoE 配置 PR 的标准依据)。6. 目录结构
7. 复现
8. 后续可扩展方向
moe_tuning/tools/moe_tune.py复核 tile。FINAL_STAGE_PLAN.md**(含 attention tile 调优、mctlass 量化路径验证、DeepSeek FlashMLA 调参等)。9. 决赛增量(2026-08-03 起)
9.1 方向 A:Triton decode attention tile 自动调优
vllm_metax/attention/ops/triton_decode_attention.py:为_fwd_kernel_stage1/_fwd_grouped_kernel_stage1增加@triton.autotuneBLOCK_N ∈ {8,16}、grouped kernelBLOCK_N ∈ {16};num_warps ∈ {1,2}、num_stages ∈ {1,2}VLLM_MACA_DECODE_ATTN_AUTOTUNE(默认1)attention_tuning/tools/attn_decode_sweep.pyattention_tuning/README.mdfeat(attn): add autotuned triton decode attention tiles for MXC500。9.2 方向 B:MACA 环境变量解析 bug 修复 + mctlass 路径基准
vllm_metax/envs.py:修复USE_PRECOMPILED_KERNEL/MACA_DP_OPT的bool(os.environ.get(...))解析 bug,统一为bool(int(...))vllm_metax/v1/attention/backends/triton_attn.py:MIN_LAUNCH_GRID_SIZE_2D/NUM_PAR_SOFTMAX_SEGMENTS改为环境变量可覆盖MACA_DP_OPT=0返回False,USE_PRECOMPILED_KERNEL=0返回False,默认值符合预期。moe_tuning/tools/mctlass_benchmark.pyfix(envs): correct bool parsing for MACA env vars and make attn thresholds overridable。9.3 当前状态
torch.zeros(10, device='cuda')可正常执行。attention_tuning/results/attn_decode_sweep.jsonattention_tuning/results/attn_decode_baseline.jsonmoe_tuning/results/mctlass_ab_summary.jsonWORKLOG_2026-08-07.md。10. 后续提交记录
2026-08-02:扩展 MXC500 fused-MoE tuned 配置到 DeepSeek-V2 / Mixtral-8x22B / DeepSeek-V3
MetaX-MACA:v0.13.0-devLindseyMei:feat/moe-more-c500-configsconfig_list.txt2026-07-13:按管理员建议将 MoE 调优重新提交到
v0.13.0-devMetaX-MACA:v0.13.0-devLindseyMei:feat/moe-c500-configsv0.13.0-dev重新整理并提交相同的 3 个 MXC500 tuned configyou could merge this pr to the dev branch of v0.13.0 (v0.13.0-dev)2026-07-06:将 tuned config 补齐到 vLLM-MetaX 主仓库
MetaX-MACA:releases/v0.13.0LindseyMei:feat/moe-mxc500-missing-configs-v0.13.0vllm_metax/model_executor/layers/fused_moe/configs/config_list.txtpyproject.tomllicense 字段以兼容新版 setuptools