Adding a New Benchmark to MINCE#
Adding a benchmark takes three code changes — a config entry, a loader, and a
--samples mapping — plus a test fixture. Everything else in the pipeline
reads the config rather than hard-coding benchmark names.
This page uses GSM8K as its worked example because GSM8K ships in the repo, so
every snippet below can be checked against the real BENCHMARKS["gsm8k"] entry
in config.py and the real _load_gsm8k in data_loader.py.
Step 1 — Read the task definition#
Most of what you need is already declared in the task’s YAML in lm-eval-harness.
Read the file on the web under lm_eval/tasks,
pinning the URL to the tag matching your installed version rather than browsing
main — task definitions change between releases:
python -c "import lm_eval; print(lm_eval.__version__)"
# 0.4.12 -> .../lm-evaluation-harness/tree/v0.4.12/lm_eval/tasks
The YAML is not always named after the task. CommonsenseQA’s lives at
tasks/commonsense_qa/default.yaml, and task YAMLs also inherit: ARC’s
arc_challenge.yaml is three lines that include: arc_easy.yaml, with
metric_list in the parent. Follow include: before concluding a field is
missing.
Three things come out of that file:
metric_list[].metricThe metric names, which become
metric_names. GSM8K declaresexact_match; ARC declaresaccandacc_norm.filter_list[].nameIf present, the task writes one row per filter and you pick one with
sample_filter— GSM8K declaresstrict-matchandflexible-extract. If absent, it writes one row per item and you need no filter.aggregationon each metricmeanon every metric means MINCE’s existing scorer already handles the benchmark and you write no metrics code. An!functionaggregation means you do; see Metrics that are not a per-item mean.
Step 2 — Look at one sample row#
The YAML names the metrics but not the JSONL keys your loader has to read, and
those differ per task: the question text is under question for GSM8K and
CommonsenseQA, but goal for PIQA. Generate three rows with a tiny model to
see the real shape — no results are shipped in the repo, and the schema does not
depend on model quality:
TASK=<your_task> # set once, used by both commands
lm_eval --model hf --model_args pretrained=facebook/opt-125m \
--tasks "$TASK" --limit 3 --log_samples \
--output_path "schema_check/$TASK"
python -c "
import json, glob, os, sys
task = sys.argv[1]
path = max(glob.glob(f'schema_check/{task}/**/samples_{task}_*.jsonl', recursive=True),
key=os.path.getmtime)
print('reading:', os.path.basename(path))
row = json.loads(open(path).readline())
print('top level:', sorted(row)); print('doc:', sorted(row['doc']))" "$TASK"
The task name belongs in both the output path and the glob, and the snippet
prints the file it read. A bare samples_*.jsonl glob over a directory shared
with earlier runs will hand you a different benchmark’s rows, and nothing in the
output looks wrong — you just write your loader against the wrong schema.
Three things to take from the output:
doc_idis the per-task item index and is the usual item ID.docholds the raw dataset row. Yourtext=field comes from here.Each metric named in the YAML appears as a top-level key, not inside
docormetrics.
filter is always present, with the literal value "none" for tasks that
declare no filter_list, so a missing filter_list does not mean a missing
filter key.
Step 3 — Get the item count#
total_items is the one field neither the YAML nor a sample log gives you. It
needs a real count from lm-eval-harness, not a dataset card. Ask lm-eval:
from lm_eval.tasks import TaskManager
td = TaskManager().load(["gsm8k"])
print(sum(len(t.eval_docs) for t in td["tasks"].values()))
# 1319
td["tasks"] is keyed by task name and holds only leaf tasks, so the same sum
works for a task that expands into subtasks: load(["mmlu"]) returns all 57
subject tasks and sums to 14042.
This materializes the dataset, so expect a minute or two and a download on first
call. It also resolves the split for you, which a dataset card will not: PIQA
ships 1838 validation and 3084 test rows, but its YAML sets test_split: null,
so 1838 is the correct total_items.
Step 4 — Register the benchmark#
Add an entry to BENCHMARKS in quark/contrib/mince/config.py:
"gsm8k": BenchmarkConfig(
name="gsm8k",
total_items=1319,
metric_names=["exact_match"],
sample_glob="samples_gsm8k_*.jsonl",
candidate_ns=[100, 200, 300, 400, 500, 600, 700, 800],
sample_filter="flexible-extract",
),
The fields, and where each value comes from:
nameMust equal the dict key. Used to look up the loader.
total_itemsThe count from Step 3 — Get the item count. Acts as a completeness guard:
load_benchmark_itemsraises if the number of loaded items differs, which catches a truncated or partial log before it silently skews a sizing run.metric_namesDrawn from
metric_list[].metric, but these must match exactly the keys your loader writes intomodel_results— sizing and drift reporting iterate over this list and willKeyErroron a name the loader never populates. Listing a subset of the YAML’s metrics is fine: ARC declaresaccandacc_norm, and a config may size onaccalone provided the loader writes onlyacc.sample_filterThe
filter_list[].namevalue you picked from the task YAML. Omit it when the task declares nofilter_list; it defaults to"", meaning every row is loaded. Your loader must honor it — see step 5.sample_globGlob matching lm-eval’s output filenames, which are
samples_<task>_<timestamp>.jsonl.candidate_nsThe subset sizes swept during Monte-Carlo sizing, and the only values
n*can be chosen from. Follow the shipped convention: 8 to 14 values spanning roughly 10% to 60% oftotal_items. Do not try to cover the whole dataset — drift atn == total_itemsis zero by definition and tells you nothing.run_sizingraises outright if any candidate exceedstotal_items.
Step 5 — Write a loader#
Add a loader to quark/contrib/mince/data_loader.py and register it in
_LOADERS at the bottom of that file. Every loader takes the same two inputs
and returns the same type:
def _load_<name>(
benchmark: BenchmarkConfig, # the entry you just registered
model_paths: dict[str, str], # model_name -> dir holding sample JSONLs
) -> list[BenchmarkItem]: ...
Its job is to turn lm-eval’s per-sample JSONL logs into one BenchmarkItem per
benchmark item, with every model’s metric values attached to that item.
model_paths is not configured anywhere — size.py builds it per run from
--model-dirs, mapping a label (the directory’s basename) to a directory of
sample JSONLs. Those labels become the model_results keys. The CLI passes one
bf16 model today, but the signature is plural because sizing takes the worst P95
across every model it is given.
Four rules it must honor:
Return the items in a fixed canonical order. The frozen subset is a list of positions into this list, so the order must be reproducible across runs — and across models, which is why the first model fixes the order below and the rest attach by ID lookup rather than by file order.
Populate
model_results[model_name]for every model inmodel_paths, keyed by the names inmetric_names.Skip rows whose
filterdoes not matchbenchmark.sample_filter. Forget this on a filtered task and you load one item per filter variant, which trips thetotal_itemsguard with an error that does not name the cause.Raise on a duplicate item ID rather than overwriting.
This is the real _load_gsm8k, abridged. It needs json, BenchmarkItem
and _find_sample_file — all already imported or defined in
data_loader.py. _find_sample_file resolves a directory plus a glob to one
path, preferring the newest file when several match:
def _load_gsm8k(
benchmark: BenchmarkConfig,
model_paths: dict[str, str],
) -> list[BenchmarkItem]:
"""Load GSM8K items from one sample JSONL per model.
lm-eval logs one row per filter, so GSM8K writes two rows per item; rows are
kept only when they match ``benchmark.sample_filter``. Items are keyed by
``doc_id`` with a single ``exact_match`` metric.
Args:
benchmark: The GSM8K benchmark configuration, including ``sample_filter``.
model_paths: Maps model_name -> directory containing the sample JSONL.
Returns:
List of BenchmarkItem in the first model's file order, each carrying
per-model metrics in ``item.model_results[model_name]``.
"""
first_model = next(iter(model_paths))
first_path = _find_sample_file(model_paths[first_model], benchmark.sample_glob)
sample_filter = benchmark.sample_filter
canonical_ids: list[int] = []
items_by_id: dict[int, BenchmarkItem] = {}
# First model fixes the canonical order.
with open(first_path) as f:
for line in f:
doc = json.loads(line)
if sample_filter and doc.get("filter") != sample_filter:
continue
doc_id = doc["doc_id"]
if doc_id in items_by_id:
raise ValueError(f"{benchmark.name}: duplicate canonical id {doc_id!r} in {first_path}")
canonical_ids.append(doc_id)
items_by_id[doc_id] = BenchmarkItem(
item_id=doc_id,
text=doc["doc"]["question"],
)
for model_name, model_dir in model_paths.items():
path = _find_sample_file(model_dir, benchmark.sample_glob)
with open(path) as f:
for line in f:
doc = json.loads(line)
if sample_filter and doc.get("filter") != sample_filter:
continue
items_by_id[doc["doc_id"]].model_results[model_name] = {
"exact_match": doc.get("exact_match", doc.get("acc", 0.0)),
}
return [items_by_id[did] for did in canonical_ids]
Then add one entry to the existing _LOADERS dict — leave the other entries
alone:
_LOADERS = {
# ... existing entries ...
"gsm8k": _load_gsm8k,
}
If the task writes one file per subject instead, model the loader on
_load_mmlu: glob all files, derive the subject from each filename with
_extract_subject(basename, "samples_mmlu_"), and use a (subject, doc_id)
tuple as the item ID, since doc_id restarts at zero in each subtask’s file.
Note
BenchmarkItem.stratum is an optional label logged in the frozen artifact
for benchmarks with natural categories. It is never used for sampling, and
leaving it unset is fine.
Step 6 — Add the --samples mapping#
This step is required and is easy to miss, because nothing in steps 1-5 fails
without it. freeze.py translates a frozen artifact into lm-eval’s
--samples map via build_samples in quark/contrib/mince/subset.py,
which dispatches on a hardcoded set. Add your benchmark to whichever of the two
fits — these grow with every benchmark, so expect more entries than shown:
# Frozen indices map 1:1 onto a single lm-eval task's doc order.
_SINGLE_FILE = {"gsm8k", "ifeval", ...}
# Group tasks: (subtask name prefix, artifact key holding the grouping field).
_GROUP_SUBTASK = {
"mmlu": ("mmlu_", "subject"),
"mmlu_pro": ("mmlu_pro_", "category"),
}
Both build_samples and the module docstring at the top of data_loader.py
carry a per-benchmark inventory in prose. Update those too, or you will leave
them stale.
Skip this and freezing appears to succeed, then fails when the subset is written:
ValueError: no --samples mapping for benchmark 'my_benchmark': add it to
_SINGLE_FILE or _GROUP_SUBTASK in quark/contrib/mince/subset.py
Then add a branch to build_item_id in the same file, so the artifact records
readable identifiers:
if benchmark_name in ("gsm8k", "my_benchmark"):
return {"doc_id": item.item_id, "question": item.text[:80]}
These keys are only labels in the frozen artifact, so join the branch above only
if question actually describes your item text. A task whose text is not a
question — PIQA’s is a goal — is better off with its own branch and its own
key, or the artifact reads as if it holds something it does not.
Warning
build_item_id has a catch-all fallback that returns
{"item_id": str(item.item_id)}, so an unregistered benchmark does not
raise here. It silently writes a frozen artifact with the ID stringified and
the item text dropped. Unlike step 6’s first half, nothing will tell you.
To check your work after a real freeze, freeze.py writes
frozen_subset_seed<seed>.json and subset_samples.json into
mince_frozen/<benchmark>/<model-label>/ unless --out-dir says otherwise.
Every entry under items should carry your keys, not item_id.
Step 7 — Add a fixture and tests#
Fixtures build tiny synthetic logs, so the tests need no model and no network.
Fixture (an alias for tuple[BenchmarkConfig, dict[str, str]]) and
write_jsonl both come from quark/contrib/mince/test/utils.py. Give the
fixture and both tests docstrings, as the rest of the suite does.
Add a fixture to quark/contrib/mince/test/conftest.py. Note that it builds
its own small BenchmarkConfig rather than importing the real one from
BENCHMARKS — the real total_items would fail the completeness guard
against six synthetic rows:
@pytest.fixture
def gsm8k_data(tmp_path: Path) -> Fixture:
"""Provide a 6-item synthetic GSM8K benchmark with two filter rows per item.
Args:
tmp_path: pytest's per-test temporary directory.
Returns:
A ``(config, model_paths)`` pair ready for ``load_benchmark_items``.
"""
config = BenchmarkConfig(
name="gsm8k",
total_items=6,
metric_names=["exact_match"],
sample_glob="samples_gsm8k_*.jsonl",
candidate_ns=[2, 3, 4],
sample_filter="flexible-extract",
)
model_paths = {}
for mi, model in enumerate(["modelA"]):
d = os.path.join(tmp_path, "gsm8k", model)
rows = []
for i in range(6):
base = {"doc_id": i, "doc": {"question": f"gsm q{i}"}}
# A second filter row that the flexible-extract filter must drop.
rows.append({**base, "filter": "strict-match", "exact_match": 0.0})
rows.append({**base, "filter": "flexible-extract", "exact_match": float((i + mi) % 2)})
write_jsonl(os.path.join(d, "samples_gsm8k_2026-01-01T00-00-00.jsonl"), rows)
model_paths[model] = d
return config, model_paths
Then add two tests. The first, in test_data_loader.py, asserts the three
things you are responsible for — item count, canonical order, and metric keys:
def test_gsm8k_applies_flexible_extract_filter(gsm8k_data: Fixture) -> None:
"""GSM8K yields one item per doc, dropping rows from the non-matching filter."""
config, model_paths = gsm8k_data
items = load_benchmark_items(config, model_paths)
# 6 docs, two filter rows each -> only flexible-extract kept.
assert len(items) == 6
assert [it.item_id for it in items] == list(range(6))
assert all("exact_match" in it.model_results["modelA"] for it in items)
The second, in test_build_samples.py, covers step 6, which no other test
reaches. That file imports build_samples as _build_samples:
def test_my_benchmark_single_file(my_benchmark_data: Fixture) -> None:
"""My benchmark maps to flat indices with ``doc_id``-enriched frozen items."""
config, model_paths = my_benchmark_data
artifact, _ = build_frozen_subset(config, model_paths, n=3, seed=42)
samples = _build_samples(artifact)
assert set(samples) == {config.name}
assert samples[config.name] == artifact["indices"]
# Not the stringified fallback from build_item_id.
assert all("doc_id" in it for it in artifact["items"])
Finally, add your fixture to test_sizing_runs_for_all_simple_benchmarks in
test_smoke_core.py. The test takes its fixtures explicitly
rather than enumerating the registry, so it does not cover a new benchmark until
you add one. It exercises the sizing sweep end to end: that every value in
candidate_ns is in range, and that metric_names survives the vectorized
aggregation in run_model.
Metrics that are not a per-item mean#
A metric whose YAML declares an !function aggregation instead of mean is
not a per-item average, so _compute_simple_metrics does not apply.
IFEval is the shipped example. Please follow how it is wired.
Warning
It has to be wired in two places: compute_accuracy in mince_metrics.py
dispatches the exact scorer used at freeze time, and run_model in
montecarlo.py dispatches a vectorized twin used by the sizing sweep, which
runs thousands of iterations. They are separate implementations of the same
aggregation. Add one without the other and sizing and freezing will silently
report different accuracies for the same subset.
Workflow summary#
1. read the task YAML -> metric_names, sample_filter, aggregation
2. log 3 samples w/ opt-125m -> JSONL keys for item_id, text, metrics
3. count eval_docs -> total_items
4. config.py -> BenchmarkConfig entry in BENCHMARKS
5. data_loader.py -> loader + _LOADERS entry
6. subset.py -> _SINGLE_FILE/_GROUP_SUBTASK + build_item_id
7. test/ -> fixture + loader test + build_samples test
Steps 4 through 6 are all required for a working benchmark. With those done, the
benchmark is selectable everywhere: size.py and freeze.py derive their
--benchmark choices from BENCHMARKS.keys(), and validate.py takes the
task name as a free-form string. Sizing, freezing, and validating then follow the
standard flow in MINCE: Compact Evaluation via Monte-Carlo Subset Sizing.
If sizing reports drift that never settles below your budget across the whole
candidate_ns range, widen the range upward before concluding the benchmark is
unsuitable — n* can only be chosen from the values you swept.