目录
nate.river

build: vendor-named wheel versions and declared FlagTree/FlagGems/FlagCX (#471)

  • build: name the wheel’s local version segment after the vendor

The local segment was SDK-shaped and inconsistent: +dtk for DCU, +ppu for PPU, +metax for MetaX, and nothing at all for CUDA, Ascend, GCU, MUSA, TsingMicro and BPU – so a CUDA wheel and a MUSA wheel were both plain 2.10.0, and the MetaX hooks overrode the segment per build image (metax3.8.1.3 on one, metax3.8.0 on the other).

It is now the vendor name everywhere it exists, taken from the wheel_local column of cmake/flagos_platforms.json, which is also the Nexus pypi lane the artifact belongs in (flagos-pypi-nvidia, flagos-pypi-hygon, …):

cuda       2.10.0+nvidia       dcu    2.10.0+hygon
gcu        2.10.0+enflame      musa   2.10.0+mthreads
ppu        2.10.0+thead        ascend 2.10.0+ascend
metax      2.10.0+metax        tsingmicro 2.10.0+tsingmicro
bpu        2.10.0              (unchanged: no lane exists for it)

Two hook edits follow from that: set_env_metax.sh no longer exports FLAGOS_WHEEL_LOCAL per image, and set_env_ppu.sh no longer defaults it to ppu. Both names are also dropped from the hooks’ export_ci_env lists, which matter because export_ci_env reads each name with ${!name} under set -u and would fail on an unset one. Leaving the overrides in place would make the artifact name a property of the build script rather than of the platform table. The table is the single source, and it is where a new platform is added.

What the segment no longer carries is the SDK version, which the two MetaX images did distinguish and which the original comment justified the segment by. That belongs in FLAGOS_SDK_VERSION, which setup.py already records in torch_fl/compatibility.json and which –release requires; nothing sets it today, so both MetaX images now produce the same 2.10.0+metax and their SDK is recorded as unknown. Wiring it into the hooks is a separate change – docs/reference/environment-variables.md and the capability-matrix comment in the platform table both say so.

Tested:

  • FLAGOS_ACCELERATOR=<platform> python setup.py --version for all nine: 2.10.0+nvidia / +hygon / +enflame / +mthreads / +thead / +ascend / +metax / +tsingmicro / 2.10.0
  • bash -n on both edited hooks; python -c "json.load(...)" on the table
  • pytest tests/unit/test_setup_build.py tests/unit/test_env_registry.py -> 57 passed (two guards fired on this change and are updated with it: test_setup_build asserted the musa manifest’s wheel_version, which is now 2.10.0+mthreads, and test_env_registry compares the documented default against torch_fl/_env.py’s registry, which is updated to match)
  • ruff check / ruff format --check -> clean
  • full pytest tests/unit -> 951 passed, the same 27 environment-dependent failures as before this branch

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com

  • build: declare FlagTree, FlagGems and FlagCX at the versions this wheel pins

pip install torch_fl installed almost none of what the wheel needs. Measured per platform before this change: flagtree and flagcx appeared nowhere in setup.py on any platform – they were installed only by the CI hooks – and flag_gems was declared on five of nine platforms, excluded from DCU, Ascend and PPU by _vendor_supplies_triton(), whose stated reason was only about Triton. So a user’s environment got a flagos device with no operator source and a distributed path staged through the host; the DCU CI log shows that warning even in CI.

Declaring the three comes with two facts that rule out a range:

  • flagtree and flagcx are not on PyPI (both 404 on the pypi-proxy repo), and flag_gems there stops at 5.0.2 while the per-op routing tables in torch_fl/configs/backends_*.conf are generated against the 5.4.0 cohort – so flag_gems>=5.0.2 was satisfied by a build this wheel was never measured on;
  • a FlagTree build is per-platform: the package name carries the vendor’s own Triton backend, so flagtree==0.7.0rc2+hcu3.6 for DCU and ==0.7.0rc3+metax3.6 for MetaX. FlagTree installs the triton distribution, which is why the separate triton>=3.5.1 requirement is now declared only where no FlagTree build is pinned (tsingmicro, bpu) – declaring both is what produces Ascend’s “0 active drivers” failure.

The versions come from .github/version-pins.env, the file the CI scripts source, read as data. A build without it raises rather than falling back to a floor, because there is no floor that could stand in for the pins.

The resulting declaration matches the pins exactly: flagtree on the seven platforms whose hook calls install_flagtree, flagcx on the five whose hook installs it (CUDA, DCU, GCU, MetaX, MUSA), flag_gems everywhere. A new test in tests/unit/test_ci_version_pins.py holds the two halves together – it reads setup.py’s own function and asserts each platform declares its pins and that flagtree/triton never appear together. Verified by breaking it both ways (dropping the flagcx line and adding a spurious triton line each fail it).

Two incidental findings from building this, both fixed here because they are the same declaration:

  • pyproject.toml’s [project.optional-dependencies] already declared a flagcx extra, unpinned, for an artifact that is not on PyPI. It is removed: the package is now a pinned requirement wherever a build exists, and the platforms without one have no FlagCX at all.
  • setup.py’s extras_require={"cuda": ...} was dead config. setuptools reports “extras_require overwritten in pyproject.toml (optional-dependencies)” on every build, so no cuda extra was ever in an artifact; the CUDA runtime deps are hard requirements on that platform through _install_requires() regardless, and _cuda_runtime_requires() becomes unused with it.

Resolution was measured rather than assumed, because the index layout decides it. For DCU as a cp310 target, the full stack resolves – and each piece comes from a different place:

flag_gems 5.4.0                   flagos-pypi-hygon
flagcx 0.14.0rc2.post2+dtk2604    flagos-pypi-hygon
flagtree 0.7.0rc2+hcu3.6          flagos-pypi-hosted
PyYAML 6.0.1, packaging 26.3,     pypi-proxy
SQLAlchemy 2.0.52, numpy 2.2.6

Two things that measurement showed and the docs now state: the vendor lane alone is not enough even where it holds all three FlagOS packages, because flag_gems itself requires packaging>=26.0 and PyYAML==6.0.1; and flagtree is not in most lanes, so flagos-pypi-hosted is required. That split is already how CI is configured (FLAGTREE_INDEX_URL versus FLAGGEMS_INDEX_URL in the hooks). docs/getting-started/installation.md gains the table and a worked command.

One thing deliberately not changed: the resolver picks numpy 2.2.6, while several CI hooks and .github/constraints.txt pin numpy<2 with the reason “2.x breaks the stock +cpu torch C extensions at import”. That did not reproduce here – torch 2.10.0+cpu imports with numpy 2.2.6 installed – so no cap is declared on a constraint that cannot be demonstrated. The discrepancy between the hooks and the pinned torch is worth its own look.

Tested:

  • every declared requirement parses with packaging and matches its own specifier, for all nine platforms
  • python setup.py egg_info now emits flagtree==0.7.0rc2+hcu3.6, flag_gems==5.4.0, flagcx==0.14.0rc2.post2+dtk2604 and Version: 2.10.0+nvidia, with no setuptools “overwritten” warning, and no flagcx or cuda extra
  • pytest tests/unit/test_ci_version_pins.py tests/unit/test_setup_build.py tests/unit/test_wheel_preflight.py -> 70 passed
  • ruff check / ruff format --check -> clean
  • full pytest tests/unit -> 951 passed, the same 27 environment-dependent failures as before this branch

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com

  • ci: give the DCU manifest 90 minutes, because it is the one that runs tests/unit

The DCU pipeline was cancelled by its own limit on the last two runs: 58m18s for a green one (2026-09-28T16:26Z) and 62m26s for the next, which gh pr checks reports as a failure. Neither run had a failing test – the second had already printed all 72 of its per-file unit results PASS when the job was cut, and every step in the job is completed/success. The wall clock was the only thing red.

The reason DCU sits at the edge is that its manifest is the only one that also runs the whole tests/unit directory, one process per file, on top of a wheel build and thirteen integration groups. Trimming that group is the wrong fix: it is what caught the cp310 tomllib collection error this branch series is fixing (#470), and no other platform runs it. Ascend, the other long manifest, is already at 120.

Tested:

  • python -c "yaml.safe_load(open('.github/configs/dcu.yml'))" -> OK
  • no test or script asserts this value (grep -rn timeout_minutes tests/) – it is read by all-tests-common.yml and passed to the job
  • ruff check / ruff format --check -> clean

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com


Co-authored-by: Claude Opus 5 (1M context) noreply@anthropic.com

8小时前268次提交

torch-fl

License Python PyTorch CI

文档 · 安装 · 快速开始 · 兼容性 · English

FlagOS 软件栈的 PyTorch 设备插件。torch-fl 对外提供统一的 flagos 设备,并在可复用原生内核、可移植编译器内核、厂商原生实现和显式 CPU 回退之间路由算子。

概览

不同加速器厂商提供的运行时、编译器栈和 PyTorch 集成方式各不相同。torch-fl 通过统一的运行时和算子路由层屏蔽这些差异,提供符合 PyTorch 使用习惯的设备接口。

用户只需使用标准 PyTorch API 和 flagos 设备。插件根据平台能力和配置为每个算子选择内核实现,无需将不同厂商暴露为不同设备名称,也无需在迁移加速器时修改工作负载。

设计理念

torch-fl 遵循五项原则:

  1. PyTorch 原生接口 — 标准 PyTorch API 无需修改;用户面向 flagos 设备编程,而不是使用厂商专用扩展。

  2. 统一逻辑设备 — 单一设备名称(flagos)抽象厂商差异,平台相关路由在算子层透明完成。

  3. 分层算子后端 — 每个操作可以分发到不同实现。路由决策以算子为粒度,而不是以设备或模型为粒度。

  4. 优先复用而非重写 — 在 dispatch 和 ABI 边界允许的情况下集成成熟内核与编译器栈,避免重复实现已有能力。

  5. 明确能力边界 — 对不支持的操作和 CPU 回退路径进行明确说明,不将其描述为完整原生覆盖;通过状态等级区分已验证支持与实验性集成。

执行路径

torch-fl 主要支持三类算子执行策略:

  • 厂商原生内核:直接调用厂商运行时和算子库(ACLNN、mudnn、topsaten),由插件生成对应厂商 C/C++ API 的绑定代码。

  • 兼容性 boxing:当厂商栈提供可与 PrivateUse1 共存的独立 PyTorch dispatch key 时,以零拷贝方式转换张量元数据。CUDA boxing 通过外部 libtorch_cuda.so 复用 NVIDIA 内核。

  • 可移植编译器内核:使用 Triton 或兼容编译器后端生成的 FlagGems 内核,在多个加速器系列之间复用,无需逐平台重写。

同一平台可以组合多种执行路径。这些路径属于内部实现策略,而不是用户选择的产品等级。

架构

PyTorch API
    |
flagos 设备(PrivateUse1)
    |
设备运行时 + 按算子路由
    |
FlagGems/编译器内核 | 兼容性 boxing | 厂商原生内核 | CPU 回退
    |
加速器运行时

上图为概念架构。组件设计、dispatch 内部机制、分布式集合通信、编译集成和 profiler 设计详见 docs/architecture/。

能力

torch-fl 在项目层面提供以下能力:

  • PyTorch eager 张量操作和设备管理
  • Autograd 与模型训练
  • torch.compile 集成
  • 通过 ProcessGroupFlagOS 支持分布式集合通信和 DDP
  • torch.profiler 集成
  • FlagGems/Triton 算子集成
  • 在 PyTorch 语义允许时,为未覆盖操作提供显式 CPU 回退

各平台的能力可用性和验证状态并不相同。某项功能存在于 torch-fl 代码库中,并不代表每个平台都已实现或验证该功能。平台详情请参阅兼容性矩阵。

硬件支持

平台 执行路径 已验证能力 状态 指南
NVIDIA CUDA 基于外部 libtorch_cuda.so 的 CUDA boxing Eager、autograd、分布式(FlagCX/NCCL)、profiler(CUPTI)、FlagGems(Python + C++) 稳定 CUDA
MetaX 通过 cu-bridge 进行 CUDA boxing,或使用 MetaX 原生内核 Eager、autograd 稳定 MetaX
Ascend 原生 ACLNN 后端,通过 FlagTree(Triton 3.5)使用 FlagGems Eager、autograd、RNG 套件 Beta Ascend
PPU 针对 PPU CUDA 13 兼容 SDK 的 CUDA boxing Eager、autograd 实验性 PPU
海光 DCU 基于 hipify DTK torch 的 CUDA boxing Eager、autograd、profiler Beta DCU
燧原 GCU 原生 topsaten 后端,未路由及 int64 算子使用 CPU 回退 Eager 实验性 GCU
摩尔线程 MUSA FlagGems 优先的 Triton 内核,原生 mudnn 回退,未路由算子使用 CPU 回退 Eager、FP16/BF16 AMP 实验性 MUSA
地平线 BPU 无 eager 内核;通过 hbdk4 使用 torch.compile 图执行路径 仅图编译 仅运行时 BPU
清微智能 已提供运行时构建选择器 尚无安装文档 仅运行时 —

状态定义:

  • 稳定:关键路径持续接受测试,并已记录受支持的版本组合。
  • Beta:主要路径已经验证,但覆盖范围、打包或发布流程尚未稳定。
  • 实验性:已在特定配置、模型或硬件环境中完成验证;接口或构建流程仍可能变化。
  • 仅运行时:已提供设备运行时支持,但该平台不是通用 eager 算子后端。

Eager 执行、训练、编译、分布式、profiler、FlagGems 等能力的详细拆分以及真实硬件验证结果,请参阅兼容性矩阵。

兼容性

组件 支持范围 说明
Python 3.8 或更高版本 平台 SDK 和可用 wheel 可能要求更窄的版本范围。
PyTorch 2.10.x(>=2.10,<2.11) 生成的 ATen 绑定与该次版本线绑定。
FlagGems 取决于平台 仅在平台路由使用 FlagGems 时,从 PyPI 或厂商兼容构建安装。
Triton/编译器 取决于平台 使用所选加速器要求的编译器发行版。

ATen 次版本绑定

torch-fl 根据 PyTorch 内部 ATen 算子注册表生成原生绑定。这些绑定对 C++ ABI 和算子 schema 变化敏感,因此项目固定使用一个 PyTorch 次版本线。当前固定版本线为 2.10.x。

使用其他 PyTorch 次版本(例如 2.11.x)会导致构建或运行时失败。同一次版本线内的补丁版本(例如 2.10.0 到 2.10.1)兼容。

厂商 SDK、CUDA toolkit 及其他平台特定版本要求记录在各平台安装指南中。

快速开始

先在安装指南中选择平台,然后运行:

import torch
import torch_fl

# 在 flagos 设备上创建张量
x = torch.randn(4, 4, device="flagos:0")

# 操作会路由到适合当前平台的内核
y = torch.relu(x @ x)

# 将结果移回 CPU
print(y.cpu())

算子具体路由到 FlagGems、厂商内核、兼容性 boxing 或 CPU 回退,由平台检测和运行时配置决定。以上代码在所有受支持加速器上保持不变。

设备查询、同步和多设备用法请参阅快速开始指南。

文档

入门

参考

架构

平台指南

参与贡献

欢迎参与以下方向的贡献:

  • 算子:补充缺失算子,或针对特定后端优化现有实现。
  • 运行时和平台集成:将 torch-fl 移植到新加速器,或改进现有后端支持。
  • 编译器集成:扩展 torch.compile 支持、改进 Triton 代码生成或增加编译后端。
  • 分布式与 profiler:增强集合通信后端或 profiling 集成。
  • 测试:增加算子正确性测试、模型集成测试或性能基准。
  • 文档:改进安装指南、故障排查文档或厂商集成说明。

开发流程、代码生成、测试要求和 PR 规范详见 CONTRIBUTING.md。

所有 GitHub 对外文本必须使用英文,包括 PR 标题、PR 描述、commit message、issue 内容和 code review 评论。本仓库的代码、注释和历史记录均使用英文,PR 也由不阅读其他语言的贡献者审阅。

致谢

torch-fl 基于多个上游项目构建:

  • PyTorch — PrivateUse1 扩展机制和 ATen 算子接口
  • FlagGems — 基于 Triton 的可移植算子内核
  • Triton(通过 FlagTree 和厂商发行版)— GPU 内核编译基础设施
  • FlagCX — 异构集合通信库
  • 厂商运行时和算子库 — ACLNN(Ascend)、mudnn(MUSA)、topsaten(GCU)、cu-bridge(MetaX)、DTK(DCU)等

以上致谢不代表所列项目或厂商对 torch-fl 的认可或背书。

许可证

torch-fl 使用 Apache License 2.0 许可证。

邀请码
    Gitlink(确实开源)
  • 加入我们
  • 官网邮箱:gitlink@ccf.org.cn
  • QQ群
  • QQ群
  • 公众号
  • 公众号

版权所有:中国计算机学会技术支持:开源发展技术委员会
京ICP备13000930号-9 京公网安备 11010802047560号