refactor(skills): split inference upgrades and hardware adaptation by framework (#35)
What
Split inference upgrade and hardware adaptation into explicit vLLM and SGLang skills, then align their acceptance contracts with the latest plugin upgrade and vendor-validation work.
The branch is merged with current
upstream/main; the previous modify/delete and skill-whitelist conflicts are resolved while keeping the four framework-specific skills.Skills
infer-vllm-plugin-upgradeinfer-vllm-hw-adaptinfer-sglang-plugin-upgradeinfer-sglang-hw-adaptThe legacy
infer-plugin-upgradeandinfer-hw-adaptskills are removed. Related skill references, the English/Chinese catalogs, and the expected-skill whitelist are updated.Common full-stack gate
Final runtime acceptance must use FlagGems, FlagTree, and FlagCX simultaneously in the same pinned environment. The report must prove:
- FlagTree supplies the active Triton compiler/runtime without a shadowing standalone Triton installation;
- FlagGems actually dispatches the intended operators;
- FlagCX executes a real collective or TP/PP path without silently falling back to NCCL or a vendor communication backend.
Imports, isolated component tests, partial-stack runs, silent fallbacks, skipped cases, and missing assets do not pass the gate.
vLLM acceptance
- execute every case in
tools/adaptation-gate-cases;- run Qwen3.6-27B and Qwen3.6-35B-A3B in eager and graph modes;
- run text single/concurrent-8, image single/concurrent-8, and mixed concurrent-8 for every model/mode pair;
- the current canonical inventory is 20 pytest scenarios and 104 requests;
- retain every generated JSON result and validate request-level quality, not only exit codes;
- update CI path triggers, platform YAML, matrices, reusable workflows, full-stack setup checks, artifacts, and aggregate status.
This reflects the recent vLLM 0.28 work, where the complete gate covered both models, both execution modes, all five scenario types, and 20/20 scenarios.
SGLang acceptance
- inventory and run every applicable non-helper script under
examples/unchanged, including offline, concurrent, MTP, and multinode examples when the backend claims those capabilities;- execute every concurrent E2E model/case and every configured text, VL, and mixed mode through
tests/run.py ... --scope e2e --task concurrent;- verify request counts, configured input/output lengths, responses, worker health, and available latency/throughput evidence;
- treat warmup failure, zero served requests, early-phase aborts, corrupt or repeated output, cleanup tracebacks, ignored child exit codes, and missing assets as non-passing;
- update
examples/**triggers, platform concurrency matrices, runner/image/model configuration, reusable E2E jobs, artifacts, and aggregate CI status.These checks incorporate the recent SGLang 0.5.18 validations where text/VL/mixed concurrency phases, zero-request warmup failures, output corruption, and wrapper exit-code propagation materially changed the acceptance result.
Validation
- all four framework-specific skill directories pass
skill-creator/scripts/quick_validate.py;- direct full-stack and framework-specific content assertions pass;
tests/test_skill_manager.pyandtests/test_skill_validation.py:131 passed (run on Windows with an in-memory
fcntlcompatibility stub because currentupstream/mainimports the Unix-only module during collection);
git diff --check: passed.No FlagScale-Agent runtime behavior is changed by the PR diff.
版权所有:中国计算机学会技术支持:开源发展技术委员会
京ICP备13000930号-9
京公网安备 11010802047560号
FlagScale-Agent
English | 简体中文
面向大规模训练、推理和服务的自主 AI Agent
🌟 概述
FlagScale-Agent 是一个专注于大规模分布式训练、推理和服务基础设施的自主 AI Agent。基于 ReAct (推理 + 行动) 范式,它将 LLM 推理能力与领域专用工具、知识和安全约束相结合,自动化复杂工作流 — 从环境配置、数据准备到训练启动、监控、调试和模型移植。
为什么选择 FlagScale-Agent?
📋 快速开始
前置要求
安装
配置
设置 API 密钥:
可选:在
~/.flagscale/agent.yaml创建配置文件:第一个命令
📚 核心概念
架构
FlagScale-Agent 遵循 ReAct 循环模式:
Guard 系统有三种工作模式:
技能(Skills)
技能是领域专用的工作流指南,教会 Agent 如何处理特定任务。每个技能包含任务描述、工具推荐、安全约束和示例。
训练技能:
train-env-setup— 安装 FlagScale、依赖和 conda 环境train-data-prep— 准备训练数据(文本分词、多模态 WebDataset)train-config— 生成带并行策略的 Hydra 配置train-run— 启动、停止和管理分布式训练任务train-monitor— 监控日志、检测异常(NaN、OOM、NCCL 超时)train-parallel-strategy— 设计 TP/PP/DP/EP/CP/SP 策略train-precision-alignment— 调试跨迁移的精度不匹配train-model-porter— 从 HuggingFace 移植模型到 Megatron-LMtrain-reproduce— 从论文/代码库复现训练结果train-moe-perf— 分析和优化 MoE 训练性能推理技能:
infer-env-setup— 配置 vllm-plugin-FL 推理环境infer-model-adapt— 适配新模型到 vllm-plugin-FLinfer-vllm-hw-adapt— 移植 vllm-plugin-FL 到新硬件后端infer-sglang-hw-adapt— 移植 sglang-plugin-FL 到新硬件后端infer-vllm-plugin-upgrade— 升级 vllm-plugin-FL 到新 vLLM 版本infer-sglang-plugin-upgrade— 升级 sglang-plugin-FL 到新 SGLang 版本infer-precision-check— 验证推理输出正确性基础设施技能:
topo-detect— 检测硬件拓扑(NVLink、NUMA、RDMA)workspace-layout— 标准化工作空间目录管理debug-strategy— 系统化调试方法论ops-discipline— 通用运维最佳实践te-upstream-sync— 将 TransformerEngine-FL 与上游版本同步mg-fl-upstream-sync— 将 Megatron-LM-FL 与上游版本同步技能会根据任务自动加载。使用
load_skill(name)手动加载。知识(Knowledge)
知识模块为基础设施领域提供深度技术文档。Agent 在行动之前加载相关知识,避免试错式错误。
可用知识域:
know-megatron-parallel— TP/PP/DP/CP/EP 进程组和通信模式know-megatron-training— 训练循环、前向/反向传播、优化器步骤、检查点know-megatron-model— Transformer 层、MLA/MTP、RoPE、混合精度、MoEknow-te-fp8— TransformerEngine FP8 量化系统know-te-attention— TE 注意力后端(DotProductAttention、Context Parallel)know-te-comm— TE 通信优化(Userbuffers、comm-gemm overlap)know-nccl-core— NCCL 拓扑检测和 channel/ring/tree 算法know-nccl-runtime— NCCL 集合通信算法、传输层、调优know-flash-attn— FlashAttention tiling、TMA/WGMMA kernel、KV cacheknow-torch-distributed— PyTorch ProcessGroup、DDP、FSDP、DeviceMeshknow-cuda-kernel— CUDA 算子开发(CUTLASS、CuTe、TMA)know-profiling— Nsys、NCU、PyTorch Profiler 集成know-flagscale— FlagScale 代码结构、Hydra 配置、Runner 执行使用
load_knowledge(name)访问文档。工具(Tools)
Agent 可以使用以下工具:
文件操作:
read_file— 读取文件内容(带行号)write_file— 创建或覆盖文件(支持追加模式)edit_file— 通过精确字符串替换编辑文件Shell:
shell— 执行 shell 命令(支持超时和后台执行)训练基础设施:
flagscale_train_monitor— 监控 FlagScale 训练(check/watch 模式)inspect_checkpoint— 深度检查 PyTorch 检查点记忆系统:
memory_write— 保存 fact/pitfall/insight 用于跨会话memory_read— 读取特定记忆条目memory_list— 列出和搜索记忆条目计划系统:
plan_create— 创建带验收标准的结构化任务计划plan_update— 更新计划步骤(doing/done/skip)、添加笔记、添加步骤plan_status— 显示当前计划和进度上下文管理:
evict— 交换消息以释放上下文空间recall— 检索之前被驱逐的消息知识与技能:
load_skill— 加载领域专用工作流指南load_knowledge— 加载技术文档Web:
web_fetch— 获取并提取 URL 文本内容web_search— 搜索当前信息Guards(守卫)
Guards 通过生命周期钩子强制执行安全和质量。当 Guard 触发时:
_override_reason继续活跃的 Guards:
VerificationGuard— 标记计划步骤完成时强制验证SafetyGuard— 阻止破坏性操作(数据删除、基础设施变更)KnowledgeSkillGuard— 提醒为专业任务加载知识/技能MemoryDisciplineGuard— 提示在发现工作后保存发现PlanGuard— 建议为多步骤任务创建计划ContextPressureGuard— 上下文接近限制时强制驱逐TrainingMonitorGuard— 提醒为训练任务使用 flagscale_train_monitorPackageSearchGuard— 防止盲目的包位置搜索UnitTestGuard— 修改 agent 源代码时要求测试PlanUpdateGuard— 验证 plan_update 调用的正确使用ArgTypeGuard— 验证工具参数类型覆盖机制:
当 Guard 阻止时,用
_override_reason重新发起相同的工具调用:记忆系统(Memory System)
Agent 使用三种记忆类型跨会话持久化发现:
fact — 可验证的环境状态(路径、配置、值)
pitfall — 调试教训(现象 → 原因 → 解决)
insight — 待消化的可复用模式
使用
memory_list()、memory_read(key)或memory_list(keyword='nccl')查询。计划系统(Plan System)
计划为多步骤任务提供带验收标准和验证关卡的结构化跟踪。
基础计划:
结构化计划(推荐用于复杂任务):
带验证完成步骤:
跟踪进度:
上下文管理(Context Management)
长会话通过驱逐和召回自动管理上下文:
驱逐(Eviction) — 交换旧消息以释放空间:
召回(Recall) — 检索被驱逐的内容:
上下文压力自动监控。当上下文达到 80% 时,
ContextPressureGuard在允许进一步工具调用前强制驱逐。🎯 使用场景
1. 环境配置
Agent 将:
2. 训练配置
Agent 生成验证过的 Hydra YAML,包含:
3. 训练启动与监控
Agent 将:
4. 调试训练失败
Agent 将:
5. 多节点训练
Agent 将:
6. 模型移植
Agent 将:
🛠️ 高级用法
斜杠命令
在 agent 内部:
/quit— 退出 agent/reload— 热重载(重启进程,恢复会话并加载新代码)/reload config— 仅重载配置(无进程重启)/resume— 列出可恢复的会话/resume <number|session_id>— 恢复特定会话/session— 显示当前会话信息(ID、目录、turn 计数)自定义技能
在
~/.flagscale/skills/my-skill/SKILL.md创建自己的技能:在对话中使用
load_skill('my-skill')加载。配置文件选项
~/.flagscale/agent.yaml的完整配置选项:完整的
AgentConfig数据类参见flagscale_agent/react/config.py。提供商特定模型
Anthropic:
claude-sonnet-4-20250514(200K 上下文,推荐)claude-opus-4-20250514(200K 上下文,最强能力)claude-3-7-sonnet-20250219(200K 上下文)OpenAI:
gpt-4o(128K 上下文)o1(200K 上下文,推理专注)o3-mini(200K 上下文)DeepSeek:
deepseek-chat(200K 上下文)deepseek-reasoner(200K 上下文,R1 推理)🧪 开发
运行测试
代码质量
测试 Agent 变更
修改 agent 源代码(
flagscale_agent/**)时:pytest tests/验证 0 失败在交互模式下使用
/reload测试变更而无需重启:项目结构
🤝 贡献
我们欢迎贡献!以下是入门方法:
git checkout -b feature/my-featurepytest tests/ -vruff format .贡献指南:
添加新技能:
skills/<skill-name>/SKILL.md,包含前置和内容添加新知识:
knowledge/docs/<domain>/目录knowledge/knowledge_config.yaml📄 许可证
本项目采用 Apache License 2.0 许可。
🙏 致谢
基于以下项目构建:
📬 联系
为 AI 基础设施社区用 ❤️ 构建