Evaluating a Quark-Quantized Model with MINCE

Evaluating a Quark-Quantized Model with MINCE#

The key idea is that you do not re-run MINCE on the quantized model. You calibrate n* once on the bf16 model, freeze it, and then reuse that frozen subset for every downstream quantized variant of that model for efficient evaluation. Below is an example of how to use MINCE to evaluate Quark-quantized models.

Prerequisites#

Steps 1-3 of MINCE: Compact Evaluation via Monte-Carlo Subset Sizing completed for your benchmark, so you already have:

  • the bf16 full-benchmark results.json (the reference score),

  • the bf16 subset score and its validated drift, and

  • mince_frozen/<benchmark>/<model-label>/subset_samples.json.

Workflow#

(from example_quark_torch_mince: BF16 EVAL -> SIZING -> FREEZE -> validated n*)
        |
QUANTIZE         (quantize_quark.py)    -> Quark INT4/AWQ checkpoint (hf_format)
        |
SUBSET SCORING   (lm_eval --samples)    -> quantized model scored on the frozen subset
        |
COMPARE          (validate.py)          -> quantization impact on the same items

The steps below span two directories — quantize_quark.py lives under examples/torch/language_modeling/llm_ptq and the MINCE CLIs under examples/contrib/mince — so set the repo root once and use it throughout:

export QUARK_REPO=/path/to/Quark

1. Quantize the model — Example of INT4 weight-only, group size 128, with AWQ. Export in hf_format so the result is a Hugging Face-loadable checkpoint:

cd $QUARK_REPO/examples/torch/language_modeling/llm_ptq

python3 quantize_quark.py \
  --model_dir meta-llama/Llama-3.1-8B-Instruct \
  --output_dir $QUARK_REPO/quark_models/Llama-3.1-8B-Instruct-awq-int4-g128 \
  --quant_scheme int4_wo_128 \
  --quant_algo awq \
  --dataset pileval_for_awq_benchmark \
  --num_calib_data 128 \
  --seq_len 512 \
  --model_export hf_format

The --model_export hf_format flag is what matters for MINCE: it writes a checkpoint that lm-eval-harness can load directly with pretrained=, so the subset-scoring step below is identical to the bf16 flow apart from the model path.

2. Score the quantized model on the frozen subset — the same lm-eval command you used in step 4 of the bf16 walkthrough, pointing pretrained= at the quantized checkpoint. The --samples map is unchanged, because it is the frozen bf16-calibrated subset:

cd $QUARK_REPO/examples/contrib/mince

lm_eval --model hf \
  --model_args pretrained=$QUARK_REPO/quark_models/Llama-3.1-8B-Instruct-awq-int4-g128 \
  --tasks gsm8k --num_fewshot 0 --apply_chat_template \
  --samples "$(cat mince_frozen/gsm8k/meta-llama__Llama-3.1-8B-Instruct/subset_samples.json)" \
  --log_samples \
  --output_path eval_test_logs/Llama-3.1-8B-Instruct-awq-int4-g128-GSM8K \
  --batch_size 1

Reuse the same generation settings as the bf16 runs (--num_fewshot, --apply_chat_template, and any other options). Otherwise the difference you measure will include unwanted variables, not just quantization effects.

3. Compare — run validate.py on the quantized subset run against the bf16 subset run. Both scored the identical frozen items, so the difference is attributable to quantization:

python validate.py --benchmark gsm8k \
  --subset-results eval_test_logs/Llama-3.1-8B-Instruct-awq-int4-g128-GSM8K/*/results_*.json \
  --baseline-results eval_test_logs/Llama-3.1-8B-Instruct-GSM8K/*/results_*.json

Note

Because the frozen subset was already validated as faithful to the full benchmark within the P95 drift budget from sizing, the quantized model’s score on n* estimates its full-benchmark score to within roughly that same budget — without ever running the full benchmark on the quantized model. This cuts the evaluation time across multiple different quantized models that need evaluation and serves as one use case of MINCE for efficient evaluation. See the MINCE paper: arxiv.org/abs/2606.22826.