Skip to main content
To load already quantized models, simply load the model weights and config. Again, if the model has been quantized offline, there’s no need to add --quantization argument when starting the engine. The quantization method will be automatically parsed from the downloaded quant_model_description.json or config.json config. SGLang supports mix-bits quantization (independently defines and loads each layer depending on the type of quantification specified in the quant_model_description.json). Advanced mix-bits for MoE in progress, will add independent quantization determination for the w13 (up-gate) and w2 (down) layers. ModelSlim on Ascend support
Quantization schemeLayer typeAscend A2 Series Products SupportedAscend A3 Series Products SupportedAscend 950PR/DT Series Products SupportedDiffusion models
W4A4 dynamicLinear√√TBD√
W8A8 staticLinear√√TBD√
W8A8 dynamicLinear√√TBD√
MXFP8 (Diffusion, LLM dense)Linearxx√√
MXFP4Linearxx√√
MXFP4 W4A8Linearxx√x
MXFP4 W4A8 (ModelSlim)MoExx√x
MXFP4 W4A4LinearxxWIPx
MXFP4 W4A4 (ModelSlim)MoExx√x
W4A4 dynamicMoE√√TBDx
W4A8 dynamicMoE√√TBDx
W8A8 dynamicMoE√√TBDx
MXFP8 (LLM MoE)MoExx√x
AWQ on Ascend support:
Quantization schemeLayer typeAscend A2 Series Products SupportedAscend A3 Series Products SupportedAscend 950PR/DT Series Products Supported
W4A16Linear√√TBD
W8A16Linear√√TBD
W4A16MoE√√TBD
GPTQ on Ascend support
Quantization schemeLayer typeAscend A2 Series Products SupportedAscend A3 Series Products SupportedAscend 950PR/DT Series Products Supported
W4A16Linear√√TBD
W8A16Linear√√TBD
W4A16 MOEMoE√√TBD
W8A16 MOEMoE√√TBD
Auto-round on Ascend support
Quantization schemeLayer typeAscend A2 Series Products SupportedAscend A3 Series Products SupportedAscend 950PR/DT Series Products Supported
W4A16Linear√√TBD
W8A16Linear√√TBD
W4A16MoE√√TBD
W8A16MoE√√TBD
Compressed-tensors (LLM Compressor) on Ascend support:
Quantization schemeLayer typeAscend A2 Series Products SupportedAscend A3 Series Products SupportedAscend 950PR/DT Series Products Supported
W8A8 dynamicLinear√√TBD
W4A8 dynamic with/without activation clipMoE√√TBD
W4A16 MOEMoE√√TBD
W8A8 dynamicMoE√√TBD
GGUF on Ascend support
Quantization typeLayer typeAscend A2 Series Products SupportedAscend A3 Series Products SupportedAscend 950PR/DT Series Products Supported
All GGUF types (standard, K-quant)Linear√√TBD
All GGUF types (standard, K-quant)MoE√√TBD
Usage Examples:
  • Dense model (e.g., Qwen3-14B-Q4_K_M.gguf):
Command
  • MoE model (e.g., Qwen3-30B-A3B-Q4_K_M.gguf):
Command
Implementation Notes:
  • GGUF weights are pre-dequantized to FP16/BF16 during model loading on CPU, then transferred to NPU for inference. This trades higher memory usage for faster runtime performance (no per-forward-pass dequantization overhead).
  • MoE layers use npu_grouped_matmul and npu_moe_init_routing / npu_moe_finalize_routing for high-performance expert computation.
  • TP (tensor parallelism) sharding is supported for both dense and MoE GGUF models.
MXFP8 for LLM dense models (e.g., Qwen3 / Qwen3.5): LLM dense W8A8 MXFP8 Linear support on Ascend was added in PR #22352. Requires Ascend 950PR/DT Series or newer (npu_dynamic_mx_quant is not available on A2/A3 Series).
  • Online MXFP8 quantization (BF16/FP16 weights β†’ MXFP8 at load time):
Command
  • Offline MXFP8 quantization (msmodelslim pre-quantized weights, W8A8_MXFP8 scheme; no --quantization flag needed β€” auto-detected from quant_model_description.json):
Command
Implementation Notes:
  • Online path: Fp8Config.get_quant_method() dispatches to NPUMXFP8LinearMethod. Weights are quantized once at load via npu_dynamic_mx_quant(weight, dst_type=torch_npu.float8_e4m3fn) and pre-transposed to [in, out]; activations are per-token quantized at inference and matmul runs via npu_quant_matmul(..., group_sizes=[1, 1, 32]) (block_size = 32).
  • Offline path: ModelSlimMXFP8Scheme loads float8_e4m3fn weights + float8_e8m0fnu block scales pre-exported by msmodelslim. Transpose is kept as a non-contiguous view (.data assignment) β€” calling .contiguous() would physically reorder the pre-quantized layout and break the block-scale mapping.
  • MoE MXFP8 (FusedMoE) for LLMs is documented in MXFP8 for LLM MoE models below.
MXFP8 for LLM MoE models (e.g. Qwen3-30B-A3B / Qwen3.5 MoE): LLM MoE W8A8 MXFP8 (FusedMoE) support builds on the dense MXFP8 path. Requires Ascend 950PR/DT Series or newer β€” the fused MoE MX kernels (npu_grouped_matmul_swiglu_quant_v2, npu_dynamic_mx_quant) are only available on the 950PR/DT Series.
  • Online MXFP8 quantization (BF16/FP16 expert weights β†’ MXFP8 at load time):
Command
  • Offline MXFP8 quantization (msmodelslim pre-quantized weights, W8A8_MXFP8 scheme). No --quantization flag is needed: the quant_model_description.json shipped with the checkpoint selects both the ModelSlim path and the scheme automatically.
Command
Implementation Notes:
  • Both paths share the per-gmm kernel NPUMXFP8MoEMethod (hardware_backend/npu/quantization/moe_methods.py), which tells online from offline by weight dtype. Expert weights and their e8m0 block scales are kept as non-contiguous transpose views β€” calling .contiguous() would tank HBM bandwidth.
  • Online path: Fp8Config.get_quant_method() dispatches FusedMoE layers to NPUMXFP8OnlineMoEMethod, which subclasses UnquantizedFusedMoEMethod and overrides only create_moe_runner to swap in the MXFP8 kernels β€” weight creation, weight post-processing and the forward pass are the unquantized Ascend ones. BF16 expert weights w13/w2 are quantized once at load via npu_dynamic_mx_quant(dst_type=torch.float8_e4m3fn) (a 3D [E, N, K] input is accepted directly).
  • Offline path: ModelSlimMXFP8MoEScheme (one instance per weight group) loads float8_e4m3fn expert weights + uint8 (e8m0, exponent + 127) block scales. The scale is reshaped [E, N, K/32] β†’ [E, N, K/64, 2] (contiguous pairing, matching npu_dynamic_mx_quant) then transposed.
  • Forward: AscendTPDispatcher runs npu_moe_init_routing_v2(quant_mode=3), which fuses the per-token MX activation quant into routing (e4m3 payload + e8m0 block scale, reshaped to the pair-split layout). AscendRunnerCore then runs gmm1 npu_grouped_matmul_swiglu_quant_v2 (cumulative group_list; fuses gate/up + swiglu + requant, so no separate activation step) β†’ gmm2 npu_grouped_matmul (count group_list). The UE8M0 (float8_e8m0fnu) scale dtypes are passed explicitly; the e4m3 x/weight dtypes are left implicit.
  • Router gate: msmodelslim may also quantize mlp.gate (W8A8_MXFP8). The gate is a ReplicatedLinear, so its quantization must be description-driven: for the offline modelslim path the gate is passed the quant config and dequantized correctly; the online path keeps it in BF16. Loading a quantized gate as BF16 without its block scale scrambles routing and produces garbage output.
  • Where the activation quant happens depends on the dispatcher. On ascend_tp it is fused into routing as described above. DeepEP has no MXFP8 dispatch dtype, so it keeps dispatching BF16 and gmm1 quantizes the hidden states itself via npu_dynamic_mx_quant before the fused kernel β€” the two paths reach the same gmm1 input. Only the ascend_tp path has been validated end-to-end on the Ascend 950PR/DT Series.
MXFP4 W4A8 for LLM dense models (e.g., Qwen3 / Qwen3.5): LLM dense W4A8 (MXFP4 4-bit weights + MXFP8 8-bit activations) Linear support was added in PR #23650. Requires Ascend 950PR/DT Series or newer.
  • Online W4A8 quantization (BF16/FP16 weights β†’ MXFP4 at load time):
Command
  • Offline W4A8 quantization (msmodelslim pre-quantized weights, W4A8_MXFP scheme; no --quantization flag needed β€” auto-detected from quant_model_description.json).
Implementation Notes:
  • Weights are packed FP4 (float4_e2m1fn_x2, two nibbles per byte) with a UE8M0 per-block shared exponent (block_size = 32); activations are per-token MXFP8. Matmul runs via npu_quant_matmul(..., x2_dtype=torch_npu.float4_e2m1fn_x2, group_sizes=[0, 0, 32]).
  • The packed-FP4 dtype passed to the NPU ops (dst_type / x2_dtype / input_dtype) must be resolved from torch_npu.float4_e2m1fn_x2 (an int enum), not the torch.float4_e2m1fn_x2 dtype object, which recent op-plugin builds reject.
  • Online and offline share the same kernel path and layout; they differ only in the weight source (RTN at load vs msmodelslim calibration).
ModelSlim W4A8 MXFP4 for LLM MoE models: SGLang auto-detects offline ModelSlim W4A8_MXFP MoE checkpoints from quant_model_description.json; do not pass --quantization. This path requires Ascend 950PR/DT Series or newer.
Command
Implementation Notes:
  • ModelSlim supplies packed MXFP4 w13 and w2 expert weights with UE8M0 block scales (block size 32).
  • Ascend TP and DeepEP dispatch activations as BF16; this path does not request MXFP8 dispatch. SGLang dynamically quantizes each expert input to MXFP8 immediately before grouped matmul.
MXFP4 W4A4 for LLM dense models (e.g. Qwen3 / Qwen3.5): LLM dense W4A4 (MXFP4 4-bit weights + 4-bit activations) Linear support was added in PR #23795. Requires Ascend 950PR/DT Series or newer β€” the dual-level online path uses the DualLevelQuantBatchMatmul op, which A2/A3 Series lack. On the Ascend NPU backend --quantization mxfp4 selects this W4A4 path (on GPU the same flag selects the upstream OCP MXFP4 MoE config instead).
  • Online W4A4 quantization (BF16/FP16 weights β†’ dual-level MXFP4 at load time):
Command
  • Offline W4A4 quantization (msmodelslim pre-quantized weights, W4A4_MXFP4 scheme; no --quantization flag needed β€” auto-detected from quant_model_description.json).
Implementation Notes:
  • Online (NPUDualLevelMXFP4LinearMethod) uses dual-level MXFP4: both weights and activations are quantized with a fine FP8 (E4M3) L0 block scale plus a coarser L1 scale via npu_dynamic_dual_level_mx_quant, and the matmul runs via npu_dual_level_quant_matmul (weight in FRACTAL_NZ). Dual-level captures per-block dynamic range far better than a single UE8M0 (power-of-2) scale, which is what made an earlier single-level RTN online path degenerate (greedy decoding could loop without emitting EOS).
  • Offline (ModelSlimMXFP4Scheme β†’ NPUSingleLevelMXFP4OfflineLinearMethod) is single-level: msmodelslim’s W4A4_MXFP4 checkpoint ships single-level UE8M0 block scales (block_size = 32), so the matmul runs via npu_quant_matmul(..., x1_dtype=x2_dtype=torch_npu.float4_e2m1fn_x2, group_sizes=[1, 1, 32]). The online and offline paths therefore use different matmul kernels β€” they no longer share the matmul path.
  • As with W4A8, the packed-FP4 dtype passed to the NPU ops (dst_type / x2_dtype) must be resolved from torch_npu.float4_e2m1fn_x2 (an int enum), not the torch.float4_e2m1fn_x2 dtype object, which recent op-plugin builds reject.
  • Validated end-to-end on Ascend 950PR/DT Series hardware.
ModelSlim W4A4 MXFP4 for LLM MoE models: SGLang auto-detects offline ModelSlim W4A4_MXFP4 MoE checkpoints from quant_model_description.json; do not pass --quantization. This path requires Ascend 950PR/DT Series or newer.
Command
Implementation Notes:
  • ModelSlim supplies packed MXFP4 w13 and w2 expert weights with UE8M0 block scales (block size 32).
  • Ascend TP and DeepEP dispatch activations as BF16; this path does not request MXFP4 dispatch. SGLang dynamically quantizes each expert input to MXFP4 immediately before grouped matmul.

Diffusion Model Quantization on Ascend NPU

SGLang-Diffusion supports MXFP8 online and offline quantization for diffusion models (such as Wan2.2) on Ascend NPUs. MXFP8 requires Ascend 950PR/DT Series; the ModelSlim W8A8/W4A4 schemes work on A2/A3 Series. Requirements for MXFP8: CANN β‰₯ 8.0.RC3, Ascend 950PR/DT Series
Quantization methodquant_type in JSONScheme classModeAscend A2/A3 Series Products SupportedAscend 950PR/DT Series Products SupportedTrigger
MXFP8 (W8A8)β€”MXFP8ConfigOnlinexβˆšβ€”quantization mxfp8
MXFP8 (W8A8)W8A8_MXFP8ModelSlimMXFP8SchemeOfflinex√auto-detected from quant_model_description.json
W8A8 staticW8A8ModelSlimW8A8Int8Offline√TBDauto-detected from quant_model_description.json
W8A8 dynamicW8A8_DYNAMICModelSlimW8A8Int8Offline√TBDauto-detected from quant_model_description.json
W4A4 dynamicW4A4_DYNAMICModelSlimW4A4Int4Offline√TBDauto-detected from quant_model_description.json

Online MXFP8 Quantization

Online quantization dynamically quantizes FP16/BF16 weights to MXFP8 at load time using npu_dynamic_mx_quant + npu_quant_matmul CANN kernels. Pass --quantization mxfp8 to override auto-detection.
Command
Command

Offline MXFP8 Quantization (ModelSlim)

For offline quantization, pre-quantize the model with msModelSlim and load the resulting checkpoint. The quantization scheme is auto-detected from quant_model_description.json, so no extra --quantization flag is needed. Step 1: Quantize with msModelSlim
Command
Note: SGLang does not support quantized embeddings; disable embedding quantization when using msmodelslim.
Step 2: Convert to Diffusers format msModelSlim saves quantized Wan2.2 weights in the original Wan format. Convert to Diffusers format using the provided repack script:
Command
Then copy all files from the original Diffusers checkpoint (except the transformer/transformer_2 folders) into the output directory. Step 3: Run inference
Command
For pre-quantized checkpoints available on ModelScope, see modelscope/Eco-Tech.