MINCE: Monte-Carlo Informed N-sizing for Compact Evaluation#
Note
MINCE is a community contribution in Quark’s contrib area. It is maintained by its author, @devledas.
For questions, bugs, or feedback, please open a GitHub issue and tag
@devledas. See Quark’s contrib Area for the contrib policies.
MINCE (Monte-Carlo Informed N-sizing for Compact Evaluation) cuts LLM benchmark evaluation time by evaluating a small, frozen subset of items instead of the full benchmark, while keeping the subset score close to the full-benchmark score.
Given a bf16 model’s per-item evaluation logs, MINCE:
sizes a representative subset (
n*) via a Monte-Carlo drift sweep,freezes that subset into a reproducible, ID-based artifact plus an lm-eval
--samplesmap, andlets you reuse the frozen subset to evaluate downstream model variants within a bounded, quantified accuracy drift.
For the method and experiments, see the MINCE paper: arxiv.org/abs/2606.22826.
Package and examples#
The importable core lives in the quark.contrib.mince package
(config, data_loader, mince_metrics, montecarlo, selection,
subset), with unit tests under quark/contrib/mince/test. The runnable
CLIs (size.py, freeze.py, validate.py, extract_inputs.py) live
under examples/contrib/mince.
The end-to-end walkthrough — generate bf16 logs, size n*, freeze the subset,
score it with lm-eval, and validate the drift — is below. We also share an
example that reuses a frozen subset to evaluate a Quark-quantized model. We provide a guide
to adding benchmarks beyond those that ship with MINCE. We also provide a guide for calibrating
n* on several bf16 models instead of one.