FlagOS-Compressor converts and quantizes HuggingFace safetensors checkpoints.
Its INT4/INT8 command supports module-level selection: selected weights are
quantized, while other low-precision weights are converted to BF16.
Install
pip install -e .
The official AutoRound package is not required for native quantization. Install
pip install -e '.[official-autoround]' only when running the optional official
reference implementation or parity checks.
Device backends
Five device types have built-in backend names: cpu, cuda, npu, mlu,
and musa. CPU and CUDA use native PyTorch directly. NPU, MLU, and MUSA load
torch_npu, torch_mlu, and torch_musa respectively only when selected.
Other PyTorch device extensions can be used by passing their registered device
type and --device; they are accepted through the generic backend and are not
counted among the five built-ins.
Inspect
flagos-compressor inspect --input /path/to/model
This reports detected weight formats and selectable groups such as moe,
moe.routed, moe.shared, attention, mlp, and linear.
Per-selector quantization in one command
Assign a complete weight/activation scheme to each part of a checkpoint in one
execution. For example, quantize DSV4 attention to W8A8 and MoE weights to
W4A16:
Each mixed --select clause starts with SELECTOR=WEIGHT_FORMAT, followed by
selector-local settings that reuse the existing CLI option names without the
leading --: activation-bits, strategy, group-size, scale-dtype, and
chunk-size. activation-bits and strategy are required, so groupwise and
per-channel weights cannot be confused. Groupwise INT4/INT8 defaults to group
sizes 32/128 when group-size is omitted; channel strategy rejects a group
size. The weight format is explicitly int4 or int8, leaving distinct
fp4/fp8 extension points for future implementations.
Selectors can be a built-in group (attention, moe, moe.routed,
moe.shared, mlp, or linear). Regular expressions continue to use
--select-name, for example --select-name 'model\.layers\.0\..*=int8' activation-bits=16 strategy=group group-size=64. Name rules are applied after
built-in target rules, and the last matching rule wins. Global --exclude and
--exclude-name rules are applied after the local rules.
The legacy homogeneous form remains valid: --select moe --bits 4.
Selector-local and legacy selectors cannot be mixed in one command, so no
selection can silently inherit the wrong format.
The mapping is an artifact contract, not a runtime-compatibility hint. If a
DSV4 source FP8 attention weight matches the local attention=int8 W8A8 rule,
it is converted to INT8 W8A8 storage, including attn.wo_a; the planner does not silently
preserve FP8 for a DeepGEMM implementation. Runtime support for consuming that
layout is a separate concern.
The command writes one compressed-tensors checkpoint with a config group for
every requested scheme. A W4A16/W8A8 combination uses the official
mixed-precision top-level format and declares pack-quantized or
int-quantized on each group. DSV4 fused attention and shared-expert runtime
aliases are included. Source-quantized weights that match no rule follow the
existing unselected policy and are converted to BF16 by default.
Per-selector mode currently supports MSE W4A16, W8A16, and W8A8. The local
selectors own weight_format, activation_bits, scale_dtype, strategy,
group_size, and chunk_size, so do not combine them with those global
settings.
--n-candidates, exclusions, the unselected policy, and backend/device flags
remain configurable. Use --dry-run to inspect the resolved plan without
writing the output checkpoint.
quantize directly writes an inference-ready compressed-tensors
pack-quantized W4A16 checkpoint and updates config.json. There is no
post-quantization conversion step. Unselected source-quantized weights use
BF16 by default.
INT8 uses symmetric groupwise MSE quantization, BF16 scales, and
compressed-tensors pack-quantized int32 storage. Its default group size is
128; override it with --group-size when the model shape or runtime requires
a different value. INT4 keeps its existing default group size of 32.
Do not pass --group-size with --strategy channel. Channelwise W8A16 export
stores scales as [out_features, 1] and declares strategy: channel in the
compressed-tensors config. vLLM’s WNA16 routed-MoE path requires group
quantization, so W8A16 channel strategy is supported for ordinary Linear and
shared-expert Linear weights, but rejected for routed experts.
Dynamic per-token W8A8, including routed MoE experts:
W8A8 uses compressed-tensors int-quantized storage: raw signed INT8 weights,
FP32 (default) or BF16 per-output-channel scales, and dynamic symmetric
per-token INT8 activations. Fused routed-expert banks are expanded to the standard
experts.<id>.<projection>.weight and weight_scale names consumed by vLLM.
W8A8 requires --bits 8 --strategy channel; --group-size is not accepted.
Use --scale-dtype bf16 to emit BF16 weight_scale tensors.
For fused MoE model types without a registered layout adapter, the CLI can
infer the 3D bank order from a consistent gate_up_proj / down_proj pair:
[E, 2I, H] plus [E, H, I] is treated as [E, out, in], while
[E, H, 2I] plus [E, I, H] is treated as [E, in, out]. Every discovered
pair must be complete, valid, and agree on the same order; otherwise
quantization stops instead of guessing.
AWQ uses AutoAWQ’s activation/weight grid-search scaling, output-MSE clipping,
asymmetric zero points, and native GEMM packing order. The result has
qweight/qzeros/scales tensors and an AWQ quantization config. Native AWQ is
currently W4A16 GEMM with zero points.
AutoRound is implemented natively with PyTorch and does not depend on the
official auto-round package. It supports symmetric group-wise W4A16 and
W8A16, learnable rounding offsets, optional min/max tuning, quantized-input
cascading, and single-device execution. Device extensions such as torch_npu,
torch_mlu, or torch_musa are imported only when their backend is selected.
The output uses the established GPTQ tensor ABI for broad loader compatibility,
while config provenance records algorithm: autoround; algorithm and packing
are separate internally.
An official AutoRound config.json can be imported without installing the
official package:
The bridge recognizes the current public fields such as scheme, bits,
group_size, iters, nsamples, seqlen, and the official historical
spelling enable_quanted_input. CLI and recipe values take precedence over
imported values. Exported GPTQ metadata retains official AutoRound-compatible
field names while identifying FlagOS-Compressor as the provider. The optional
official Python entry point is lazy-loaded only for explicit reference/parity
work; normal installation and native execution do not import it.
All calibrated methods execute the original Transformers model definition and
discover decoder blocks through Transformers’ no-split contract; they do not
maintain a per-model forward adapter. Transformers-v5 fused expert modules are
temporarily exposed as ordinary per-expert nn.Linear modules, so the same
hooks handle dense and routed-MoE models. Routed experts must be selected as a
complete gate/up/down set. Source FP4/FP8 checkpoints are staged as BF16 before
calibration.
Calibration data can use any of these formats (text below can be changed with
--calibration-text-column):
.txt/.text: one sample per non-empty line;
.jsonl: one JSON string or {"text": "sample"} object per line;
.json: a top-level list of strings/objects, {"data": [...]}, or
{"text": [...]};
a Hugging Face dataset name whose selected split contains a string text
column; this form requires the optional datasets package.
For example:
{"text": "The first calibration sample."}
{"text": "The second calibration sample."}
The default is 128 examples packed into 512-token blocks; use
--calibration-samples and --calibration-seq-length to change it. Empty
records and examples longer than the configured sequence length are skipped.
Ready-made recipes are in examples/recipes/.
FlagOS-Compressor
FlagOS-Compressor converts and quantizes HuggingFace
safetensorscheckpoints. Its INT4/INT8 command supports module-level selection: selected weights are quantized, while other low-precision weights are converted to BF16.Install
The official AutoRound package is not required for native quantization. Install
pip install -e '.[official-autoround]'only when running the optional official reference implementation or parity checks.Device backends
Five device types have built-in backend names:
cpu,cuda,npu,mlu, andmusa. CPU and CUDA use native PyTorch directly. NPU, MLU, and MUSA loadtorch_npu,torch_mlu, andtorch_musarespectively only when selected. Other PyTorch device extensions can be used by passing their registered device type and--device; they are accepted through the generic backend and are not counted among the five built-ins.Inspect
This reports detected weight formats and selectable groups such as
moe,moe.routed,moe.shared,attention,mlp, andlinear.Per-selector quantization in one command
Assign a complete weight/activation scheme to each part of a checkpoint in one execution. For example, quantize DSV4 attention to W8A8 and MoE weights to W4A16:
Each mixed
--selectclause starts withSELECTOR=WEIGHT_FORMAT, followed by selector-local settings that reuse the existing CLI option names without the leading--:activation-bits,strategy,group-size,scale-dtype, andchunk-size.activation-bitsandstrategyare required, so groupwise and per-channel weights cannot be confused. Groupwise INT4/INT8 defaults to group sizes 32/128 whengroup-sizeis omitted; channel strategy rejects a group size. The weight format is explicitlyint4orint8, leaving distinctfp4/fp8extension points for future implementations.Selectors can be a built-in group (
attention,moe,moe.routed,moe.shared,mlp, orlinear). Regular expressions continue to use--select-name, for example--select-name 'model\.layers\.0\..*=int8' activation-bits=16 strategy=group group-size=64. Name rules are applied after built-in target rules, and the last matching rule wins. Global--excludeand--exclude-namerules are applied after the local rules.The legacy homogeneous form remains valid:
--select moe --bits 4. Selector-local and legacy selectors cannot be mixed in one command, so no selection can silently inherit the wrong format.The mapping is an artifact contract, not a runtime-compatibility hint. If a DSV4 source FP8 attention weight matches the local
attention=int8W8A8 rule, it is converted to INT8 W8A8 storage, includingattn.wo_a; the planner does not silently preserve FP8 for a DeepGEMM implementation. Runtime support for consuming that layout is a separate concern.The command writes one compressed-tensors checkpoint with a config group for every requested scheme. A W4A16/W8A8 combination uses the official
mixed-precisiontop-level format and declarespack-quantizedorint-quantizedon each group. DSV4 fused attention and shared-expert runtime aliases are included. Source-quantized weights that match no rule follow the existingunselectedpolicy and are converted to BF16 by default.Per-selector mode currently supports MSE W4A16, W8A16, and W8A8. The local selectors own
weight_format,activation_bits,scale_dtype,strategy,group_size, andchunk_size, so do not combine them with those global settings.--n-candidates, exclusions, the unselected policy, and backend/device flags remain configurable. Use--dry-runto inspect the resolved plan without writing the output checkpoint.Convert to BF16
Quantize selected weights
INT4 (the backwards-compatible default):
quantizedirectly writes an inference-ready compressed-tensorspack-quantizedW4A16 checkpoint and updatesconfig.json. There is no post-quantization conversion step. Unselected source-quantized weights use BF16 by default.INT8 weight-only W8A16:
INT8 uses symmetric groupwise MSE quantization, BF16 scales, and compressed-tensors
pack-quantizedint32 storage. Its default group size is 128; override it with--group-sizewhen the model shape or runtime requires a different value. INT4 keeps its existing default group size of 32.INT8 also supports one scale per output channel:
Do not pass
--group-sizewith--strategy channel. Channelwise W8A16 export stores scales as[out_features, 1]and declaresstrategy: channelin the compressed-tensors config. vLLM’s WNA16 routed-MoE path requires group quantization, so W8A16 channel strategy is supported for ordinary Linear and shared-expert Linear weights, but rejected for routed experts.Dynamic per-token W8A8, including routed MoE experts:
W8A8 uses compressed-tensors
int-quantizedstorage: raw signed INT8 weights, FP32 (default) or BF16 per-output-channel scales, and dynamic symmetric per-token INT8 activations. Fused routed-expert banks are expanded to the standardexperts.<id>.<projection>.weightandweight_scalenames consumed by vLLM. W8A8 requires--bits 8 --strategy channel;--group-sizeis not accepted. Use--scale-dtype bf16to emit BF16weight_scaletensors.For fused MoE model types without a registered layout adapter, the CLI can infer the 3D bank order from a consistent
gate_up_proj/down_projpair:[E, 2I, H]plus[E, H, I]is treated as[E, out, in], while[E, H, 2I]plus[E, I, H]is treated as[E, in, out]. Every discovered pair must be complete, valid, and agree on the same order; otherwise quantization stops instead of guessing.GPTQ (AutoGPTQ-compatible)
GPTQ uses AutoGPTQ’s running Hessian, Cholesky error feedback, activation ordering, true-sequential projection groups, and native
qweight/qzeros/scales/g_idxpacking. Defaults aredesc_act: true,static_groups: false,true_sequential: true, and 1% dampening. The output contains the standard GPTQ quantization config and canonical GPTQ safetensors filenames.AWQ (AutoAWQ-compatible)
AWQ uses AutoAWQ’s activation/weight grid-search scaling, output-MSE clipping, asymmetric zero points, and native GEMM packing order. The result has
qweight/qzeros/scalestensors and an AWQ quantization config. Native AWQ is currently W4A16 GEMM with zero points.AutoRound (native PyTorch)
AutoRound is implemented natively with PyTorch and does not depend on the official
auto-roundpackage. It supports symmetric group-wise W4A16 and W8A16, learnable rounding offsets, optional min/max tuning, quantized-input cascading, and single-device execution. Device extensions such astorch_npu,torch_mlu, ortorch_musaare imported only when their backend is selected. The output uses the established GPTQ tensor ABI for broad loader compatibility, while config provenance recordsalgorithm: autoround; algorithm and packing are separate internally.An official AutoRound
config.jsoncan be imported without installing the official package:The bridge recognizes the current public fields such as
scheme,bits,group_size,iters,nsamples,seqlen, and the official historical spellingenable_quanted_input. CLI and recipe values take precedence over imported values. Exported GPTQ metadata retains official AutoRound-compatible field names while identifying FlagOS-Compressor as the provider. The optional official Python entry point is lazy-loaded only for explicit reference/parity work; normal installation and native execution do not import it.All calibrated methods execute the original Transformers model definition and discover decoder blocks through Transformers’ no-split contract; they do not maintain a per-model forward adapter. Transformers-v5 fused expert modules are temporarily exposed as ordinary per-expert
nn.Linearmodules, so the same hooks handle dense and routed-MoE models. Routed experts must be selected as a complete gate/up/down set. Source FP4/FP8 checkpoints are staged as BF16 before calibration.Calibration data can use any of these formats (
textbelow can be changed with--calibration-text-column):.txt/.text: one sample per non-empty line;.jsonl: one JSON string or{"text": "sample"}object per line;.json: a top-level list of strings/objects,{"data": [...]}, or{"text": [...]};textcolumn; this form requires the optionaldatasetspackage.For example:
The default is 128 examples packed into 512-token blocks; use
--calibration-samplesand--calibration-seq-lengthto change it. Empty records and examples longer than the configured sequence length are skipped. Ready-made recipes are inexamples/recipes/.Selections can be combined: