build: vendor-named wheel versions and declared FlagTree/FlagGems/FlagCX (#471)
- build: name the wheel’s local version segment after the vendor
The local segment was SDK-shaped and inconsistent:
+dtkfor DCU,+ppufor PPU,+metaxfor MetaX, and nothing at all for CUDA, Ascend, GCU, MUSA, TsingMicro and BPU – so a CUDA wheel and a MUSA wheel were both plain2.10.0, and the MetaX hooks overrode the segment per build image (metax3.8.1.3on one,metax3.8.0on the other).It is now the vendor name everywhere it exists, taken from the
wheel_localcolumn of cmake/flagos_platforms.json, which is also the Nexus pypi lane the artifact belongs in (flagos-pypi-nvidia, flagos-pypi-hygon, …):cuda 2.10.0+nvidia dcu 2.10.0+hygon gcu 2.10.0+enflame musa 2.10.0+mthreads ppu 2.10.0+thead ascend 2.10.0+ascend metax 2.10.0+metax tsingmicro 2.10.0+tsingmicro bpu 2.10.0 (unchanged: no lane exists for it)Two hook edits follow from that: set_env_metax.sh no longer exports FLAGOS_WHEEL_LOCAL per image, and set_env_ppu.sh no longer defaults it to
ppu. Both names are also dropped from the hooks’ export_ci_env lists, which matter because export_ci_env reads each name with${!name}underset -uand would fail on an unset one. Leaving the overrides in place would make the artifact name a property of the build script rather than of the platform table. The table is the single source, and it is where a new platform is added.What the segment no longer carries is the SDK version, which the two MetaX images did distinguish and which the original comment justified the segment by. That belongs in FLAGOS_SDK_VERSION, which setup.py already records in torch_fl/compatibility.json and which –release requires; nothing sets it today, so both MetaX images now produce the same
2.10.0+metaxand their SDK is recorded as unknown. Wiring it into the hooks is a separate change – docs/reference/environment-variables.md and the capability-matrix comment in the platform table both say so.Tested:
FLAGOS_ACCELERATOR=<platform> python setup.py --versionfor all nine: 2.10.0+nvidia / +hygon / +enflame / +mthreads / +thead / +ascend / +metax / +tsingmicro / 2.10.0bash -non both edited hooks;python -c "json.load(...)"on the tablepytest tests/unit/test_setup_build.py tests/unit/test_env_registry.py-> 57 passed (two guards fired on this change and are updated with it: test_setup_build asserted the musa manifest’s wheel_version, which is now2.10.0+mthreads, and test_env_registry compares the documented default against torch_fl/_env.py’s registry, which is updated to match)ruff check/ruff format --check-> clean- full
pytest tests/unit-> 951 passed, the same 27 environment-dependent failures as before this branchCo-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
- build: declare FlagTree, FlagGems and FlagCX at the versions this wheel pins
pip install torch_flinstalled almost none of what the wheel needs. Measured per platform before this change:flagtreeandflagcxappeared nowhere in setup.py on any platform – they were installed only by the CI hooks – andflag_gemswas declared on five of nine platforms, excluded from DCU, Ascend and PPU by_vendor_supplies_triton(), whose stated reason was only about Triton. So a user’s environment got a flagos device with no operator source and a distributed path staged through the host; the DCU CI log shows that warning even in CI.Declaring the three comes with two facts that rule out a range:
flagtreeandflagcxare not on PyPI (both 404 on the pypi-proxy repo), andflag_gemsthere stops at 5.0.2 while the per-op routing tables in torch_fl/configs/backends_*.conf are generated against the 5.4.0 cohort – soflag_gems>=5.0.2was satisfied by a build this wheel was never measured on;- a FlagTree build is per-platform: the package name carries the vendor’s own Triton backend, so
flagtree==0.7.0rc2+hcu3.6for DCU and==0.7.0rc3+metax3.6for MetaX. FlagTree installs thetritondistribution, which is why the separatetriton>=3.5.1requirement is now declared only where no FlagTree build is pinned (tsingmicro, bpu) – declaring both is what produces Ascend’s “0 active drivers” failure.The versions come from .github/version-pins.env, the file the CI scripts source, read as data. A build without it raises rather than falling back to a floor, because there is no floor that could stand in for the pins.
The resulting declaration matches the pins exactly: flagtree on the seven platforms whose hook calls install_flagtree, flagcx on the five whose hook installs it (CUDA, DCU, GCU, MetaX, MUSA), flag_gems everywhere. A new test in tests/unit/test_ci_version_pins.py holds the two halves together – it reads setup.py’s own function and asserts each platform declares its pins and that flagtree/triton never appear together. Verified by breaking it both ways (dropping the flagcx line and adding a spurious triton line each fail it).
Two incidental findings from building this, both fixed here because they are the same declaration:
- pyproject.toml’s
[project.optional-dependencies]already declared aflagcxextra, unpinned, for an artifact that is not on PyPI. It is removed: the package is now a pinned requirement wherever a build exists, and the platforms without one have no FlagCX at all.- setup.py’s
extras_require={"cuda": ...}was dead config. setuptools reports “extras_requireoverwritten inpyproject.toml(optional-dependencies)” on every build, so nocudaextra was ever in an artifact; the CUDA runtime deps are hard requirements on that platform through _install_requires() regardless, and_cuda_runtime_requires()becomes unused with it.Resolution was measured rather than assumed, because the index layout decides it. For DCU as a cp310 target, the full stack resolves – and each piece comes from a different place:
flag_gems 5.4.0 flagos-pypi-hygon flagcx 0.14.0rc2.post2+dtk2604 flagos-pypi-hygon flagtree 0.7.0rc2+hcu3.6 flagos-pypi-hosted PyYAML 6.0.1, packaging 26.3, pypi-proxy SQLAlchemy 2.0.52, numpy 2.2.6Two things that measurement showed and the docs now state: the vendor lane alone is not enough even where it holds all three FlagOS packages, because flag_gems itself requires
packaging>=26.0andPyYAML==6.0.1; andflagtreeis not in most lanes, so flagos-pypi-hosted is required. That split is already how CI is configured (FLAGTREE_INDEX_URL versus FLAGGEMS_INDEX_URL in the hooks). docs/getting-started/installation.md gains the table and a worked command.One thing deliberately not changed: the resolver picks numpy 2.2.6, while several CI hooks and .github/constraints.txt pin
numpy<2with the reason “2.x breaks the stock +cpu torch C extensions at import”. That did not reproduce here – torch 2.10.0+cpu imports with numpy 2.2.6 installed – so no cap is declared on a constraint that cannot be demonstrated. The discrepancy between the hooks and the pinned torch is worth its own look.Tested:
- every declared requirement parses with packaging and matches its own specifier, for all nine platforms
python setup.py egg_infonow emitsflagtree==0.7.0rc2+hcu3.6,flag_gems==5.4.0,flagcx==0.14.0rc2.post2+dtk2604andVersion: 2.10.0+nvidia, with no setuptools “overwritten” warning, and noflagcxorcudaextrapytest tests/unit/test_ci_version_pins.py tests/unit/test_setup_build.py tests/unit/test_wheel_preflight.py-> 70 passedruff check/ruff format --check-> clean- full
pytest tests/unit-> 951 passed, the same 27 environment-dependent failures as before this branchCo-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
- ci: give the DCU manifest 90 minutes, because it is the one that runs tests/unit
The DCU pipeline was cancelled by its own limit on the last two runs: 58m18s for a green one (2026-09-28T16:26Z) and 62m26s for the next, which
gh pr checksreports as a failure. Neither run had a failing test – the second had already printed all 72 of its per-file unit results PASS when the job was cut, and every step in the job iscompleted/success. The wall clock was the only thing red.The reason DCU sits at the edge is that its manifest is the only one that also runs the whole tests/unit directory, one process per file, on top of a wheel build and thirteen integration groups. Trimming that group is the wrong fix: it is what caught the cp310
tomllibcollection error this branch series is fixing (#470), and no other platform runs it. Ascend, the other long manifest, is already at 120.Tested:
python -c "yaml.safe_load(open('.github/configs/dcu.yml'))"-> OK- no test or script asserts this value (
grep -rn timeout_minutes tests/) – it is read by all-tests-common.yml and passed to the jobruff check/ruff format --check-> cleanCo-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Co-authored-by: Claude Opus 5 (1M context) noreply@anthropic.com
版权所有:中国计算机学会技术支持:开源发展技术委员会
京ICP备13000930号-9
京公网安备 11010802047560号
torch-fl
文档 · 安装 · 快速开始 · 兼容性 · English
FlagOS 软件栈的 PyTorch 设备插件。torch-fl 对外提供统一的
flagos设备,并在可复用原生内核、可移植编译器内核、厂商原生实现和显式 CPU 回退之间路由算子。概览
不同加速器厂商提供的运行时、编译器栈和 PyTorch 集成方式各不相同。torch-fl 通过统一的运行时和算子路由层屏蔽这些差异,提供符合 PyTorch 使用习惯的设备接口。
用户只需使用标准 PyTorch API 和
flagos设备。插件根据平台能力和配置为每个算子选择内核实现,无需将不同厂商暴露为不同设备名称,也无需在迁移加速器时修改工作负载。设计理念
torch-fl 遵循五项原则:
PyTorch 原生接口 — 标准 PyTorch API 无需修改;用户面向
flagos设备编程,而不是使用厂商专用扩展。统一逻辑设备 — 单一设备名称(
flagos)抽象厂商差异,平台相关路由在算子层透明完成。分层算子后端 — 每个操作可以分发到不同实现。路由决策以算子为粒度,而不是以设备或模型为粒度。
优先复用而非重写 — 在 dispatch 和 ABI 边界允许的情况下集成成熟内核与编译器栈,避免重复实现已有能力。
明确能力边界 — 对不支持的操作和 CPU 回退路径进行明确说明,不将其描述为完整原生覆盖;通过状态等级区分已验证支持与实验性集成。
执行路径
torch-fl 主要支持三类算子执行策略:
厂商原生内核:直接调用厂商运行时和算子库(ACLNN、mudnn、topsaten),由插件生成对应厂商 C/C++ API 的绑定代码。
兼容性 boxing:当厂商栈提供可与 PrivateUse1 共存的独立 PyTorch dispatch key 时,以零拷贝方式转换张量元数据。CUDA boxing 通过外部
libtorch_cuda.so复用 NVIDIA 内核。可移植编译器内核:使用 Triton 或兼容编译器后端生成的 FlagGems 内核,在多个加速器系列之间复用,无需逐平台重写。
同一平台可以组合多种执行路径。这些路径属于内部实现策略,而不是用户选择的产品等级。
架构
上图为概念架构。组件设计、dispatch 内部机制、分布式集合通信、编译集成和 profiler 设计详见 docs/architecture/。
能力
torch-fl 在项目层面提供以下能力:
torch.compile集成ProcessGroupFlagOS支持分布式集合通信和 DDPtorch.profiler集成各平台的能力可用性和验证状态并不相同。某项功能存在于 torch-fl 代码库中,并不代表每个平台都已实现或验证该功能。平台详情请参阅兼容性矩阵。
硬件支持
libtorch_cuda.so的 CUDA boxingtorch.compile图执行路径状态定义:
Eager 执行、训练、编译、分布式、profiler、FlagGems 等能力的详细拆分以及真实硬件验证结果,请参阅兼容性矩阵。
兼容性
>=2.10,<2.11)ATen 次版本绑定
torch-fl 根据 PyTorch 内部 ATen 算子注册表生成原生绑定。这些绑定对 C++ ABI 和算子 schema 变化敏感,因此项目固定使用一个 PyTorch 次版本线。当前固定版本线为 2.10.x。
使用其他 PyTorch 次版本(例如 2.11.x)会导致构建或运行时失败。同一次版本线内的补丁版本(例如 2.10.0 到 2.10.1)兼容。
厂商 SDK、CUDA toolkit 及其他平台特定版本要求记录在各平台安装指南中。
快速开始
先在安装指南中选择平台,然后运行:
算子具体路由到 FlagGems、厂商内核、兼容性 boxing 或 CPU 回退,由平台检测和运行时配置决定。以上代码在所有受支持加速器上保持不变。
设备查询、同步和多设备用法请参阅快速开始指南。
文档
入门
参考
架构
平台指南
参与贡献
欢迎参与以下方向的贡献:
开发流程、代码生成、测试要求和 PR 规范详见 CONTRIBUTING.md。
所有 GitHub 对外文本必须使用英文,包括 PR 标题、PR 描述、commit message、issue 内容和 code review 评论。本仓库的代码、注释和历史记录均使用英文,PR 也由不阅读其他语言的贡献者审阅。
致谢
torch-fl 基于多个上游项目构建:
以上致谢不代表所列项目或厂商对 torch-fl 的认可或背书。
许可证
torch-fl 使用 Apache License 2.0 许可证。