EAGLE-3 Large-Model Best Recipe#
This page describes a portable baseline for adapting the EAGLE-3 pipeline to
large, quantized target models. The accompanying
recipes/eagle3_mxfp4_moe.yaml file uses
amd/MiniMax-M3-MXFP4 as an example target and provides a starting point for
adapting the workflow to other compatible models.
Warning
This recipe is a general starting point rather than a configuration tuned for a particular LLM. Run its small smoke manifest before preparing a full manifest and reserving a full node.
Configuration layout#
Keep reusable orchestration separate from per-model settings:
The shared runner owns data generation, hidden-state extraction, draft training, export, serving, and reporting.
A model adapter identifies the target, its chat-template family, nested-module keys required for loading, and supported quantization.
A domain manifest identifies data sources and their license information.
Run-local paths, endpoints, credentials, host details, and resolved manifests belong outside source control.
Domain manifests#
A domain manifest describes sources without embedding a fixed data mix.
Useful starting domains include general_instruction, code,
math_reasoning, question_answering, and multilingual. Optional
domains can cover long-context or structured tool-use prompts.
The committed smoke manifest uses small repository-authored JSONL files. A full manifest can use Hugging Face sources or local JSONL files:
version: 1
seed: 0
max_samples_per_domain: 1000
splits:
train: 0.98
eval: 0.02
domains:
general_instruction:
sources:
- type: jsonl
path: REPLACE_WITH_SOURCE.jsonl
license: REPLACE_WITH_LICENSE
code:
sources:
- type: hf
dataset: REPLACE_WITH_DATASET
split: train
license: REPLACE_WITH_LICENSE
The manifest processor applies deterministic normalization, global SHA256 deduplication, per-domain caps, domain-aware held-out splitting, and source and license provenance. Domain caps are user inputs; this recipe does not prescribe a fixed data mix. Assistant responses used for training should be regenerated by the selected target so that the resulting data is on-policy.
The Python runner materializes manifest prompts before launching Docker. The target then regenerates assistant responses on-policy; filtered train/eval JSONL and provenance summaries remain in the content-addressed run cache.
Because that step runs outside the container, type: hf sources require the
optional datasets package in the same interpreter that runs the CLI, and
network access to the dataset host. Install it into the environment you launch
from, or export the prompts yourself and reference them as type: jsonl
sources instead. Hugging Face sources are read in streaming mode, so no full
dataset copy is written to disk.
Model adapter#
The packaged adapter template is
quark/experimental/speculative_decoding/recipes/eagle3_mxfp4_moe.yaml. It
identifies the MXFP4 target model and provides the embedding key needed for its
nested text tower. Draft width, attention
geometry, vocabulary size, auxiliary-layer selection, and training controls are
derived from target metadata or runner settings. Because an MoE target’s
intermediate_size describes one expert rather than a dense draft MLP, values
below twice the hidden width use a transparent 3 * hidden_size dense SwiGLU
default; the exported Llama EAGLE-3 draft uses silu to match its
implementation.
The profile supplies a small smoke manifest and the target’s chat delimiters. A
full run must override data.domain_manifest with a full manifest.
Eight-GPU topology and cost#
Plan for one node with eight compatible AMD Instinct GPUs and substantial local
storage. The MiniMax adapter uses two TP4 target replicas for data
generation, then one TP4 extraction engine plus four FSDP draft ranks for
training. Baseline and speculative TP4 services share the node during matched
evaluation. Other adapters derive replica counts from world_size and
target_tp_size; if two target groups do not fit, evaluation runs baseline
and speculative services sequentially on the same GPUs.
Each wall-clock hour on the full node consumes eight GPU-hours. The setup preflight recommends at least 600 GB free for this large-model profile. See Results for what one 300,000-prompt run actually consumed.
Command skeleton#
First run the packaged smoke profile:
python3 -m quark.experimental.speculative_decoding.setup \
--base_model amd/MiniMax-M3-MXFP4
python3 -m quark.experimental.speculative_decoding.run \
--config recipes/eagle3_mxfp4_moe.yaml \
--base_model amd/MiniMax-M3-MXFP4
For a full run, provide a manifest and a fresh output directory:
python3 -m quark.experimental.speculative_decoding.run \
--config recipes/eagle3_mxfp4_moe.yaml \
--base_model amd/MiniMax-M3-MXFP4 \
execution.profile=full \
data.domain_manifest=/absolute/path/to/domain_manifest.yaml \
training.output_dir=ckpts/minimax-m3-eagle3-full
Start with a small representative dataset and a fresh output directory. Preserve the resolved recipe, environment versions, and source manifest with the run report.
Report-only evaluation#
Large-model evaluation reports observed serving metrics without enforcing a Qwen-specific pass/fail threshold. The report includes the resolved-configuration hash, software and hardware identifiers, served acceptance length, matched throughput and latency, wall time, and GPU-hours. Report sanitization should exclude local paths, endpoints, credentials, hostnames, and unmaterialized source records.
Qwen retains its validated gates; the large-model profile sets them to zero and records the observed metrics for review.
Limitations#
The recorded result comes from one hardware and software combination, one manifest, and one evaluation suite. Treat it as a reference point when planning a run.
The Quark-native trainer is single-GPU, and its streaming extractor is an interface stub. Large-model execution therefore requires compatible external distributed training and transport support.
Compatibility depends on the selected ROCm, vLLM, remote model code, and target checkpoint versions.
Quality, latency, throughput, and capacity depend on the selected model, data, and serving setup.
Results#
The following reference run summarizes the resource use and serving metrics of this configuration on one node.
Environment and inputs#
Target |
|
Resolved configuration |
|
Hardware |
8 x AMD Instinct MI350X ( |
Platform |
Linux 6.8.0 x86_64, glibc 2.39 |
Trainer revision |
|
Manifest prompts |
300,000 across five generic domains, 60,000 each |
Manifest digest |
|
Licenses |
Apache-2.0 120,000, MIT 120,000, ODC-BY-1.0 60,000 |
On-policy generation returned 299,887 of 300,000 prompts. Discarding generations that hit the token limit left 260,729 training and 600 held-out rows, trained for one epoch over 8,148 optimizer steps.
Cost#
Wall time |
22 h 29 min end to end |
GPU-hours |
approximately 180 |
Peak run-directory storage |
111 GB, plus 3.8 GB of cached on-policy data |
Phase split |
7 h 19 min generation, 13 h 50 min training, the remainder export and evaluation |
Most run-directory storage is the training checkpoint and the held-out evaluation cache, which stores target hidden states and costs roughly 76 MB per held-out row. Size the held-out split against available disk rather than as a fixed fraction of a large prompt count.
Measured serving metrics#
Three matched rounds, 40 prompts each, 512 maximum new tokens, temperature 0, concurrency 1. Baseline and speculative services use the same target, tensor parallelism, and prompts.
Round |
Baseline (tok/s) |
Speculative (tok/s) |
Speedup |
|---|---|---|---|
1 |
124.8 |
206.2 |
1.652x |
2 |
125.8 |
211.7 |
1.683x |
3 |
125.7 |
212.5 |
1.690x |
Median |
125.7 |
211.7 |
1.683x |
Served acceptance length was 2.716 with three speculative tokens per step.
Reading these numbers#
Acceptance length depends on the evaluation prompts, so a figure measured on a different suite is not comparable with this one. Speedup additionally depends on concurrency, sequence length, and the serving stack; this run measured a single stream, where speculative decoding has the most headroom. Both figures should be re-measured on the workload that matters to you.