MINCE: Monte-Carlo Informed N-sizing for Compact Evaluation

MINCE: Monte-Carlo Informed N-sizing for Compact Evaluation#

Note

MINCE is a community contribution in Quark’s contrib area. It is maintained by its author, @devledas. For questions, bugs, or feedback, please open a GitHub issue and tag @devledas. See Quark’s contrib Area for the contrib policies.

MINCE (Monte-Carlo Informed N-sizing for Compact Evaluation) cuts LLM benchmark evaluation time by evaluating a small, frozen subset of items instead of the full benchmark, while keeping the subset score close to the full-benchmark score.

Given a bf16 model’s per-item evaluation logs, MINCE:

  1. sizes a representative subset (n*) via a Monte-Carlo drift sweep,

  2. freezes that subset into a reproducible, ID-based artifact plus an lm-eval --samples map, and

  3. lets you reuse the frozen subset to evaluate downstream model variants within a bounded, quantified accuracy drift.

For the method and experiments, see the MINCE paper: arxiv.org/abs/2606.22826.

Package and examples#

The importable core lives in the quark.contrib.mince package (config, data_loader, mince_metrics, montecarlo, selection, subset), with unit tests under quark/contrib/mince/test. The runnable CLIs (size.py, freeze.py, validate.py, extract_inputs.py) live under examples/contrib/mince.

The end-to-end walkthrough — generate bf16 logs, size n*, freeze the subset, score it with lm-eval, and validate the drift — is below. We also share an example that reuses a frozen subset to evaluate a Quark-quantized model. We provide a guide to adding benchmarks beyond those that ship with MINCE. We also provide a guide for calibrating n* on several bf16 models instead of one.