Troubleshooting#
Note
In this documentation, AMD Quark is sometimes referred to simply as “Quark” for ease of reference. When you encounter the term “Quark” without the “AMD” prefix, it specifically refers to the AMD Quark quantizer unless otherwise stated. Please do not confuse it with other products or technologies that share the name “Quark.”
AMD Quark for PyTorch#
Environment Issues#
Known Issue: Windows CPU mode does not support fp16.
Because of an existing PyTorch issue, Windows CPU mode cannot perfectly support fp16.
C++ Compilation Issues#
Known Issue: Stuck in the compilation phase for a long time (over ten minutes), and terminal shows:
[QUARK-INFO]: Configuration checking start.
[QUARK-INFO]: C++ kernel build directory [cache folder path]/torch_extensions/py39...
Solution:
Delete the cache folder [cache folder path]/torch_extensions and run AMD Quark again.
FlyDSL / native inference issues (gfx950)#
These apply to enable_native_inference with native_linear_mode set to
flydsl_a8w4, flydsl_svdquant, or mxfp4.
Environment requirements
Component |
Requirement |
Why |
|---|---|---|
GPU |
gfx950 (MI350) |
the A8W4 / MXFP4 kernels target this arch only |
|
exactly 0.2.4 |
the pin every recorded accuracy number was produced on; installed by
|
|
>= v0.1.20, installed from git |
the A8W4 GEMM imports its MFMA epilogue and pipeline helpers from
|
|
not used |
the GEMM is vendored in |
Known Issue: native_linear_mode="mxfp4" (w4a4) silently does nothing, while
flydsl_a8w4 works.
mxfp4 runs on aiter’s ASM gemm_a4w4. That kernel has no FlyDSL dependency, but
aiter/ops/flydsl/__init__ sits in aiter’s top-level import cascade
(aiter/__init__ -> gemm_op_a8w8 -> ops/flydsl/__init__) and older releases
raised there when the installed flydsl did not match their pin. That exception
aborts the entire aiter import and takes gemm_a4w4 with it – so an unrelated version
check disables w4a4, with no error that names the cause.
Solution: upgrade aiter to a release whose own minimum is flydsl >= 0.2.4. That
gate was raised in v0.1.18 (2026-07-19); v0.1.20 is the earliest verified
end-to-end here. ROCm’s aiter is not on PyPI under that name, so install from git:
pip install git+https://github.com/ROCm/aiter.git@v0.1.21.post1
python3 -c "from aiter import gemm_a4w4; print('gemm_a4w4 OK')"
Do not downgrade flydsl to satisfy an older aiter: that breaks Quark’s own A8W4
kernels, which need 0.2.x.
Warning
Do not pick an aiter tag by date. Tags are cut from a release branch rather than the
main tip, so a newer tag can lack a newer commit – and the tag dates are not even
monotonic (v0.1.16.post6 postdates v0.1.17). Verify by content – the version
gate is the only thing that moved, so it is the only check that discriminates:
git show <tag>:aiter/ops/flydsl/__init__.py | grep FLYDSL_VERSION # want 0.2.4
The releases still import aiter’s FlyDSL kernels eagerly, so a future flydsl API
break could take gemm_a4w4 down as collateral again. That is why
setup_flydsl.sh reports whether gemm_a4w4 is importable rather than assuming
it.
Known Issue: which A8W4 GEMM is Quark running?
Always the vendored snapshot in quark/torch/kernel/flydsl/kernels/. ROCm/FlyDSL#957
(the fused SVD epilogue) has not landed upstream, so the snapshot is the only
implementation that supports flydsl_svdquant, and Quark imports it directly rather
than searching a FlyDSL checkout. Only the GEMM is vendored – its unmodified MFMA
epilogue and pipeline helpers come from aiter.ops.flydsl.kernels. Do not try to
substitute upstream kernels by putting a checkout on PYTHONPATH – mixing the pinned
snapshot with a different helper revision is what previously broke native inference (a
compile-time TypeError, then a GPU memory fault).
Known Issue: fewer native linears than expected after enable_native_inference.
The FlyDSL A8W4 GEMM requires in_features to be >= 256 and a multiple of 256, and
out_features to be >= 128 and a multiple of 128. Layers that violate this are
silently left on the eager path, so a “native” model is often a mix.
enable_native_inference returns the number converted – count it rather than assuming.
(On Wan2.2-A14B this is 400/400 per expert, because the one offending layer, proj_out
with out_features=64, is excluded from quantization.)
Known Issue: any Quark import hangs for many minutes with no output.
Usually a stale C++ extension build lock, left behind when a process was killed mid-build
(for example with kill -9). Distinguish it from slow compilation with time: a
blocked process shows tiny user time against huge real time (for example 18 s of
CPU across 15 minutes).
Solution: with no Python process running, delete
[cache folder path]/torch_extensions/py3xx_cpu/kernel_ext/lock. The compiled .so
next to it stays valid, so nothing is rebuilt.
vLLM Integration Issues#
Known Issue: vLLM fails with AttributeError: 'CustomOp' has no attribute 'op_registry'.
Typical Error:
AttributeError: 'CustomOp' has no attribute 'op_registry'
Root Cause:
Some vLLM builds check for (or rely on) the
amd-quarkPython package for emulation MXFP4 kernels.If you intend to use native MXFP4 kernels,
amd-quarkis not required.
Solution:
Install
amd-quarkif your vLLM runtime requires emulation kernels.Otherwise, configure vLLM to use its native MXFP4 kernel path (when available) so that
amd-quarkis not needed.
Known Issue: vLLM weight loading fails with shape mismatch for some models/checkpoints.
Typical Error:
AssertionError: param_data.shape == loaded_weight.shape
Root Cause:
The checkpoint stores packed weights (e.g., packed QKV), but the corresponding vLLM model implementation does not provide the required mapping.
Solution:
Define
packed_modules_mappingin the corresponding vLLM model executor file. For example, to support Qwen3, the following mapping is required in vllm/vllm/model_executor/models/qwen3.py:class Qwen3ForCausalLM(nn.Module, SupportsLoRA, SupportsPP, SupportsEagle3): packed_modules_mapping = { "qkv_proj": ["q_proj", "k_proj", "v_proj"], "gate_up_proj": ["gate_proj", "up_proj"], }
Quantization Performance Issues#
Known Issue: Quantization can be extremely slow when running on CPU for very large LLM checkpoints.
Solution:
Use a shard-by-shard (file-by-file) loading/quantization workflow to reduce peak memory and improve throughput.
See Language Model PTQ for the recommended workflow and scripts in this repository.