Composed Quantization (Rotation + AutoSmoothQuant + GPTQ)#
At low bit widths, no single pre-quantization algorithm is enough. Rotation, AutoSmoothQuant (ASQ) and GPTQ each attack a different source of quantization error, and AMD Quark lets you apply all three to the same model in one pass.
What each algorithm actually fixes#
The three algorithms are not interchangeable, and they are not redundant. Each targets a distinct failure mode of low-bit quantization.
Rotation — smooth activation outliers. Quantization error is driven by the dynamic range of a tensor. A single outlier channel forces a coarse scale on every other channel sharing that scale. Rotation multiplies the hidden state by an orthogonal matrix \(R\) and its inverse into the next weight, which is mathematically a no-op in floating point but spreads a concentrated outlier across all channels. The canonical illustration: the vector \((1, 10)\) has a 10x range; rotate it 45 degrees and you get \((7.78, 6.36)\), a 1.2x range. See Rotation pre-processing optimization for more details, and the runnable example at amd/Quark. Rotation is applied offline, fused into the weights, so it costs nothing at inference time.
ASQ — rebalances difficulty between activations and weights. Rotation makes outliers less extreme but does not equalize activations against weights. In W4A4 both tensors are quantized, and activations are typically the harder of the two. ASQ searches for a per-channel scale \(s\) and rewrites \(y = (x / s)(s \cdot W)\): the activation gets easier to quantize, the weight gets harder, and the product is unchanged. The scale is folded into the preceding LayerNorm or Linear, so this is also free at inference. See the runnable Auto SmoothQuant tutorial, and Activation/weight smoothing (SmoothQuant) for the smoothing family ASQ belongs to. Rotation cannot do this — an orthogonal matrix preserves norms and therefore cannot move difficulty from one tensor to the other.
GPTQ — minimizes the rounding error that remains. After rotation and ASQ have reshaped the tensors, the weights still have to land on the fp4 grid. Round-to-nearest treats each weight independently; GPTQ instead uses second-order (Hessian) information from calibration data to quantize column by column, compensating each rounding decision against the ones still to come. It fixes error the first two algorithms cannot touch, because it operates on the rounding step itself rather than on the distributions being rounded.
They compose because they act at different stages:
Rotation ASQ GPTQ
(reshape the -> (rebalance -> (choose the
distribution) act. vs. wt.) best rounding)
Order matters#
What works the best for Qwen3.5-397B-A17B model is: Rotation → ASQ → GPTQ
Rotation first, because it is a global change of basis. Any smoothing scale or Hessian computed before rotation would describe a tensor that no longer exists afterwards.
ASQ before GPTQ, because ASQ folds a scale into the weights. Running ASQ after GPTQ would rescale weights that GPTQ had already snapped onto the fp4 grid, knocking them back off it and discarding the entire benefit of the Hessian correction.
Quark applies algo_config entries in list order, so the order you write is the order you
get.
Target: what gets which treatment#
The recipe is deliberately not uniform across the model:
Module group |
Treatment |
|---|---|
|
Rotation + ASQ + GPTQ, then MXFP4 |
|
Rotation + ASQ + GPTQ, then MXFP4 |
|
Rotation, then MXFP4 only |
|
Rotation, then MXFP4 only |
|
Rotation only — stays bf16 |
|
excluded (Conv1d, not a Linear) |
|
excluded (scalar gate, stays bf16) |
|
excluded |
Not every projection gets every algorithm#
The three algorithms do not cover the same module set:
conv1dis not a Linear, so no algorithm applies.out_projgets no ASQ because the smoothing scale has nowhere to fold: the gated RMSNorm in front of it is sizedhead_v_dim, butout_projtakeshead_v_dim * num_v_headsinputs.The MoE rows are rotated because they read the rotated residual stream and would break otherwise. They skip ASQ and GPTQ because neither is needed for correctness.
Module |
Rotation |
ASQ |
GPTQ |
|---|---|---|---|
|
yes |
yes |
yes |
|
yes |
no |
yes |
|
n/a |
no |
no |
|
yes |
yes |
yes |
|
yes |
no |
yes |
|
yes |
no |
no |
|
yes |
no |
no |
|
yes |
no |
no |
Rotation config#
Rotation needs to know, for each RMSNorm, which modules produce its input and which consume its
output, so that \(R\) and \(R^{-1}\) can be fused into the right weights. Each group
names the union of both attention shapes; per-layer filtering keeps whichever exists. The JSON
below is deserialized into RotationConfig, whose reference documents every field:
{
"name": "rotation",
"backbone": "model.language_model",
"model_decoder_layers": "model.language_model.layers",
"rotation_size": 4096,
"r1": true, "r2": false, "r3": false, "r4": false,
"online_r1_rotation": false,
"scaling_layers": {
"first_layer": [
{"prev_modules": ["model.language_model.embed_tokens"],
"norm_module": "model.language_model.layers.layer_id.input_layernorm",
"next_modules": ["model.language_model.layers.layer_id.linear_attn.in_proj_qkv",
"model.language_model.layers.layer_id.linear_attn.in_proj_z",
"model.language_model.layers.layer_id.linear_attn.in_proj_a",
"model.language_model.layers.layer_id.linear_attn.in_proj_b",
"model.language_model.layers.layer_id.self_attn.q_proj",
"model.language_model.layers.layer_id.self_attn.k_proj",
"model.language_model.layers.layer_id.self_attn.v_proj"]},
{"prev_modules": ["model.language_model.layers.layer_id.linear_attn.out_proj",
"model.language_model.layers.layer_id.self_attn.o_proj"],
"norm_module": "model.language_model.layers.layer_id.post_attention_layernorm",
"next_modules": ["model.language_model.layers.layer_id.mlp.gate",
"model.language_model.layers.layer_id.mlp.experts.*.gate_proj",
"model.language_model.layers.layer_id.mlp.experts.*.up_proj",
"model.language_model.layers.layer_id.mlp.shared_expert.gate_proj",
"model.language_model.layers.layer_id.mlp.shared_expert.up_proj",
"model.language_model.layers.layer_id.mlp.shared_expert_gate"]}
],
"middle_layers": [
{"prev_modules": ["model.language_model.layers.pre_layer_id.mlp.experts.*.down_proj",
"model.language_model.layers.pre_layer_id.mlp.shared_expert.down_proj"],
"norm_module": "model.language_model.layers.layer_id.input_layernorm",
"next_modules": ["model.language_model.layers.layer_id.linear_attn.in_proj_qkv",
"model.language_model.layers.layer_id.linear_attn.in_proj_z",
"model.language_model.layers.layer_id.linear_attn.in_proj_a",
"model.language_model.layers.layer_id.linear_attn.in_proj_b",
"model.language_model.layers.layer_id.self_attn.q_proj",
"model.language_model.layers.layer_id.self_attn.k_proj",
"model.language_model.layers.layer_id.self_attn.v_proj"]},
{"prev_modules": ["model.language_model.layers.layer_id.linear_attn.out_proj",
"model.language_model.layers.layer_id.self_attn.o_proj"],
"norm_module": "model.language_model.layers.layer_id.post_attention_layernorm",
"next_modules": ["model.language_model.layers.layer_id.mlp.gate",
"model.language_model.layers.layer_id.mlp.experts.*.gate_proj",
"model.language_model.layers.layer_id.mlp.experts.*.up_proj",
"model.language_model.layers.layer_id.mlp.shared_expert.gate_proj",
"model.language_model.layers.layer_id.mlp.shared_expert.up_proj",
"model.language_model.layers.layer_id.mlp.shared_expert_gate"]}
],
"last_layer": [
{"prev_modules": ["model.language_model.layers.layer_id.mlp.experts.*.down_proj",
"model.language_model.layers.layer_id.mlp.shared_expert.down_proj"],
"norm_module": "model.language_model.norm",
"next_modules": ["lm_head"]}
]
}
}
Note that mlp.gate, mlp.experts.*, mlp.shared_expert.* and mlp.shared_expert_gate
all appear here even though two of them are excluded from quantization and none of them receives
an algorithm for offline merging.
ASQ config#
Each group names a prev_op (the module the smoothing scale is folded into) and the
layers that consume its output. Only the attention blocks are listed. The JSON below is
deserialized into AutoSmoothQuantConfig, whose reference documents every field:
{
"name": "autosmoothquant",
"model_decoder_layers": "model.language_model.layers",
"compute_scale_loss": "MAE",
"scaling_layers": [
{"prev_op": "input_layernorm",
"layers": ["linear_attn.in_proj_qkv", "linear_attn.in_proj_z",
"linear_attn.in_proj_a", "linear_attn.in_proj_b"],
"inp": "linear_attn.in_proj_qkv", "module2inspect": "linear_attn"},
{"prev_op": "input_layernorm",
"layers": ["self_attn.q_proj", "self_attn.k_proj", "self_attn.v_proj"],
"inp": "self_attn.q_proj", "module2inspect": "self_attn"}
]
}
Unlike rotation, ASQ resolves module names with fnmatch, so a group whose modules do not
exist in a given layer is genuinely skipped without any code support. Listing both attention
variants is safe: each layer matches exactly one of the two.
Two omissions are deliberate: linear_attn.out_proj has no foldable prev_op, and the
v_proj -> o_proj group for self_attn.o_proj is left out pending verification. Both are
explained in Not every projection gets every algorithm.
GPTQ config#
GPTQ just needs the list of Linear modules to correct, covering both attention shapes (full- and
linear-attention). The JSON below is deserialized into GPTQConfig, whose reference
documents every field; see Configuring PyTorch Quantization for how algorithm configs are
built in Python and for the constraints GPTQ places on the quantization scheme:
{
"name": "gptq",
"model_decoder_layers": "model.language_model.layers",
"damp_percent": 0.01, "desc_act": true, "static_groups": true, "block_size": 128,
"inside_layer_modules": [
"self_attn.q_proj", "self_attn.k_proj", "self_attn.v_proj", "self_attn.o_proj",
"linear_attn.in_proj_qkv", "linear_attn.in_proj_z",
"linear_attn.in_proj_a", "linear_attn.in_proj_b", "linear_attn.out_proj"
]
}
Important
R2 must be disabled for this model family ("r2": false). R2 rotates the v_proj output and relies on o_proj
applying the exact inverse. Qwen3.5 attention always multiplies an elementwise sigmoid gate into the attention output
before o_proj (attn_output = attn_output * torch.sigmoid(gate)), and an elementwise gate does not commute with a
dense rotation — so the inverse no longer cancels and the model is silently corrupted. R1 alone is safe.
Running Quantization#
The three configuration files are passed to the stock quantize_quark.py entry point with
--quant_algo selecting the composition and --quant_algo_config_file supplying each
config:
cd examples/torch/language_modeling/llm_ptq
export HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export QUARK_MXFP4_IMPL=triton
python3 quantize_quark.py \
--model_dir Qwen/Qwen3.5-397B-A17B \
--data_type bfloat16 --device cuda --multi_gpu balanced \
--model_attn_implementation sdpa \
--quant_scheme mxfp4 \
--quant_algo rotation,autosmoothquant,gptq \
--quant_algo_config_file rotation ./rotation.json \
--quant_algo_config_file autosmoothquant ./autosmoothquant.json \
--quant_algo_config_file gptq ./gptq.json \
--exclude_layers "lm_head" "*visual*" "*mtp*" "*conv1d*" \
"*mlp.gate" "*shared_expert_gate" \
--dataset pileval --num_calib_data 128 --seq_len 2048 --batch_size 1 \
--skip_evaluation \
--model_export hf_format \
--output_dir Qwen3.5-397B-A17B-MXFP4-Rot-ASQ-GPTQ
Serving the export#
SGLang is not part of this repository. Pull a ROCm build from lmsysorg/sglang-rocm and launch the server from the image. The MXFP4 export is ~213 GB, so two MI355X GPUs are enough:
MODEL=/path/to/Qwen3.5-397B-A17B-MXFP4-Rot-ASQ-GPTQ
IMAGE=lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260903
docker run -d --rm --name serve-mxfp4 \
--device=/dev/kfd --device=/dev/dri --group-add video \
--ipc=host --shm-size 64g --network host \
-e HIP_VISIBLE_DEVICES=0,1 \
-v "$MODEL":"$MODEL" \
--entrypoint python3 "$IMAGE" \
-m sglang.launch_server \
--model-path "$MODEL" \
--served-model-name Qwen3.5-397B-A17B-MXFP4-Rot-ASQ-GPTQ \
--tp 2 --context-length 131072 \
--reasoning-parser qwen3-thinking \
--host 0.0.0.0 --port 8000 --trust-remote-code
Results#
MMLU-Pro measures whether quality survives into multi-step reasoning, which is far more sensitive to quantization damage than perplexity. All quantized arms are MXFP4 W4A4 with identical calibration (128 pileval samples, sequence length 2048).
MMLU-Pro (0-shot, chat, full split)#
Configuration |
Accuracy |
|---|---|
No quantization |
88.33 % |
Rotation only |
62.72% |
Rotation + ASQ |
82.65% |
Rotation + ASQ + GPTQ |
87.45% |
Reproduce the MMLU-Pro numbers by serving each export and running lm_eval against it.
Install the API extra first (pip install 'lm_eval[api]') — the plain package cannot talk to
an OpenAI-compatible endpoint and fails at start-up with a missing tenacity:
lm_eval --model local-chat-completions \
--tasks mmlu_pro_chat \
--model_args "model=Qwen3.5-397B-A17B-MXFP4-Rot-ASQ-GPTQ,max_length=96000,base_url=http://0.0.0.0:8000/v1/chat/completions,num_concurrent=128,max_retries=3,tokenized_requests=False,tokenizer_backend=None,timeout=3600" \
--num_fewshot 0 \
--apply_chat_template \
--output_path results.json \
--seed 42 \
--gen_kwargs "do_sample=true,temperature=0.6,top_p=0.95,top_k=20,min_p=0.0,max_gen_toks=64000,presence_penalty=0.0,repetition_penalty=1.0,seed=42"
See also#
Rotation pre-processing optimization — rotation and QuaRot in depth
Activation/weight smoothing (SmoothQuant) — the smoothing family that ASQ belongs to