EAGLE-3 Large-Model Best Recipe#

This page describes a portable baseline for adapting the EAGLE-3 pipeline to large, quantized target models. The accompanying recipes/eagle3_mxfp4_moe.yaml file uses amd/MiniMax-M3-MXFP4 as an example target and provides a starting point for adapting the workflow to other compatible models.

Warning

This recipe is a general starting point rather than a configuration tuned for a particular LLM. Run its small smoke manifest before preparing a full manifest and reserving a full node.

Configuration layout#

Keep reusable orchestration separate from per-model settings:

  • The shared runner owns data generation, hidden-state extraction, draft training, export, serving, and reporting.

  • A model adapter identifies the target, its chat-template family, nested-module keys required for loading, and supported quantization.

  • A domain manifest identifies data sources and their license information.

  • Run-local paths, endpoints, credentials, host details, and resolved manifests belong outside source control.

Domain manifests#

A domain manifest describes sources without embedding a fixed data mix. Useful starting domains include general_instruction, code, math_reasoning, question_answering, and multilingual. Optional domains can cover long-context or structured tool-use prompts.

The committed smoke manifest uses small repository-authored JSONL files. A full manifest can use Hugging Face sources or local JSONL files:

version: 1
seed: 0
max_samples_per_domain: 1000
splits:
  train: 0.98
  eval: 0.02
domains:
  general_instruction:
    sources:
      - type: jsonl
        path: REPLACE_WITH_SOURCE.jsonl
        license: REPLACE_WITH_LICENSE
  code:
    sources:
      - type: hf
        dataset: REPLACE_WITH_DATASET
        split: train
        license: REPLACE_WITH_LICENSE

The manifest processor applies deterministic normalization, global SHA256 deduplication, per-domain caps, domain-aware held-out splitting, and source and license provenance. Domain caps are user inputs; this recipe does not prescribe a fixed data mix. Assistant responses used for training should be regenerated by the selected target so that the resulting data is on-policy.

The Python runner materializes manifest prompts before launching Docker. The target then regenerates assistant responses on-policy; filtered train/eval JSONL and provenance summaries remain in the content-addressed run cache.

Because that step runs outside the container, type: hf sources require the optional datasets package in the same interpreter that runs the CLI, and network access to the dataset host. Install it into the environment you launch from, or export the prompts yourself and reference them as type: jsonl sources instead. Hugging Face sources are read in streaming mode, so no full dataset copy is written to disk.

Model adapter#

The packaged adapter template is quark/experimental/speculative_decoding/recipes/eagle3_mxfp4_moe.yaml. It identifies the MXFP4 target model and provides the embedding key needed for its nested text tower. Draft width, attention geometry, vocabulary size, auxiliary-layer selection, and training controls are derived from target metadata or runner settings. Because an MoE target’s intermediate_size describes one expert rather than a dense draft MLP, values below twice the hidden width use a transparent 3 * hidden_size dense SwiGLU default; the exported Llama EAGLE-3 draft uses silu to match its implementation.

The profile supplies a small smoke manifest and the target’s chat delimiters. A full run must override data.domain_manifest with a full manifest.

Eight-GPU topology and cost#

Plan for one node with eight compatible AMD Instinct GPUs and substantial local storage. The MiniMax adapter uses two TP4 target replicas for data generation, then one TP4 extraction engine plus four FSDP draft ranks for training. Baseline and speculative TP4 services share the node during matched evaluation. Other adapters derive replica counts from world_size and target_tp_size; if two target groups do not fit, evaluation runs baseline and speculative services sequentially on the same GPUs.

Each wall-clock hour on the full node consumes eight GPU-hours. The setup preflight recommends at least 600 GB free for this large-model profile. See Results for what one 300,000-prompt run actually consumed.

Command skeleton#

First run the packaged smoke profile:

python3 -m quark.experimental.speculative_decoding.setup \
    --base_model amd/MiniMax-M3-MXFP4

python3 -m quark.experimental.speculative_decoding.run \
    --config recipes/eagle3_mxfp4_moe.yaml \
    --base_model amd/MiniMax-M3-MXFP4

For a full run, provide a manifest and a fresh output directory:

python3 -m quark.experimental.speculative_decoding.run \
    --config recipes/eagle3_mxfp4_moe.yaml \
    --base_model amd/MiniMax-M3-MXFP4 \
    execution.profile=full \
    data.domain_manifest=/absolute/path/to/domain_manifest.yaml \
    training.output_dir=ckpts/minimax-m3-eagle3-full

Start with a small representative dataset and a fresh output directory. Preserve the resolved recipe, environment versions, and source manifest with the run report.

Report-only evaluation#

Large-model evaluation reports observed serving metrics without enforcing a Qwen-specific pass/fail threshold. The report includes the resolved-configuration hash, software and hardware identifiers, served acceptance length, matched throughput and latency, wall time, and GPU-hours. Report sanitization should exclude local paths, endpoints, credentials, hostnames, and unmaterialized source records.

Qwen retains its validated gates; the large-model profile sets them to zero and records the observed metrics for review.

Limitations#

  • The recorded result comes from one hardware and software combination, one manifest, and one evaluation suite. Treat it as a reference point when planning a run.

  • The Quark-native trainer is single-GPU, and its streaming extractor is an interface stub. Large-model execution therefore requires compatible external distributed training and transport support.

  • Compatibility depends on the selected ROCm, vLLM, remote model code, and target checkpoint versions.

  • Quality, latency, throughput, and capacity depend on the selected model, data, and serving setup.

Results#

The following reference run summarizes the resource use and serving metrics of this configuration on one node.

Environment and inputs#

Target

amd/MiniMax-M3-MXFP4

Resolved configuration

cac90317fc1ea5e3d19f6ab65e29dfc85102d96367a356fa44e64416a6c4658c

Hardware

8 x AMD Instinct MI350X (gfx950), one node

Platform

Linux 6.8.0 x86_64, glibc 2.39

Trainer revision

d230a3f13212e8a3eb35e04f99fbdd8fa45241f0

Manifest prompts

300,000 across five generic domains, 60,000 each

Manifest digest

98691720ac7a2f980cfc84ab58518747878898ebd02d08d293a33e7637da74aa

Licenses

Apache-2.0 120,000, MIT 120,000, ODC-BY-1.0 60,000

On-policy generation returned 299,887 of 300,000 prompts. Discarding generations that hit the token limit left 260,729 training and 600 held-out rows, trained for one epoch over 8,148 optimizer steps.

Cost#

Wall time

22 h 29 min end to end

GPU-hours

approximately 180

Peak run-directory storage

111 GB, plus 3.8 GB of cached on-policy data

Phase split

7 h 19 min generation, 13 h 50 min training, the remainder export and evaluation

Most run-directory storage is the training checkpoint and the held-out evaluation cache, which stores target hidden states and costs roughly 76 MB per held-out row. Size the held-out split against available disk rather than as a fixed fraction of a large prompt count.

Measured serving metrics#

Three matched rounds, 40 prompts each, 512 maximum new tokens, temperature 0, concurrency 1. Baseline and speculative services use the same target, tensor parallelism, and prompts.

Round

Baseline (tok/s)

Speculative (tok/s)

Speedup

1

124.8

206.2

1.652x

2

125.8

211.7

1.683x

3

125.7

212.5

1.690x

Median

125.7

211.7

1.683x

Served acceptance length was 2.716 with three speculative tokens per step.

Reading these numbers#

Acceptance length depends on the evaluation prompts, so a figure measured on a different suite is not comparable with this one. Speedup additionally depends on concurrency, sequence length, and the serving stack; this run measured a single stream, where speculative decoding has the most headroom. Both figures should be re-measured on the workload that matters to you.