EAGLE-3#
AMD Quark provides an AMD/ROCm-native, Quark-integrated pipeline to train EAGLE-3 draft models from scratch, synthesize on-policy training data, export a vLLM-loadable draft, deploy it on vLLM, and evaluate acceptance length and throughput. It trains the draft, aligned to a Quark (optionally quantized) target, using EAGLE-3’s low/mid/high-level hidden-state features and a training-time-test (TTT) objective.
Note
Choose an EAGLE-3 workflow:
Validated Qwen3-8B path. The end-to-end example on AMD Instinct is the fully GPU-validated, reproducible recipe (data, training, serving, and a measured speedup). The default YAML/CLI on this page packages and drives that same TorchSpec streaming path: about 150k target-generated samples, 8-GPU training, vLLM export, and a matched baseline/speculative benchmark.
Portable large-model baseline. The EAGLE-3 Large-Model Best Recipe provides a model-adapter configuration and domain-manifest workflow for adapting the pipeline to a large, quantized target. It is a general starting point: validate its smoke profile first, then provide a full manifest before starting a full run.
TorchSpec is an MIT-licensed, open-source distributed training framework maintained by the LightSeek Foundation; Quark uses its streaming EAGLE-3 trainer for the validated multi-GPU execution path.
The Python API (qsd.convert / qsd.train) remains the Quark-native
single-GPU reference implementation. Its export path (export_hf /
convert_to_vllm) emits a
draft in vLLM’s LlamaForCausalLMEagle3 format – architecture, weight key
names, and shapes are validated byte-for-byte against a known-good draft – so
exports load directly into a vLLM speculative serve. Its native training path
is format-validated only. Select it explicitly with
execution.backend=native.
See Extraction regimes and Status and limitations.
Background: EAGLE-3 and EAGLE-3.1#
EAGLE is a family of feature-level draft models. Instead of bolting on an unrelated small language model, an EAGLE draft is trained against the target model’s own internal representations, which is what makes its proposals so acceptable to the verifier. Across generations it has improved draft quality and acceptance rate (EAGLE, EAGLE2, EAGLE3).
EAGLE-3 combines low-, mid-, and high-level hidden states from the target with a training-time-test (TTT) objective: the draft is unrolled several steps during training so it learns to propose multiple tokens in a row, matching how it is used at inference.
This package trains an EAGLE-3.1 draft, a small refinement of EAGLE-3 that
adds extra normalization inside the draft head to stabilize cold-start
(from-scratch) training. The draft is a single decoder layer and exports to vLLM’s
LlamaForCausalLMEagle3 serving format.
Pipeline#
quant/ extraction/
│ │
▼ ▼
data/ ──▶ convert() ──▶ training/ ──▶ export/ ──▶ eval/
quant/ (a Quark MXFP4/FP8 target used as the verifier and hidden-state source)
feeds convert(); extraction/ (target hidden states, online / offline /
streaming) feeds training/. Each stage is described in
Module layout below.
Requirements#
Python 3.11+, PyTorch (ROCm build),
transformers,safetensors,pyyaml,tqdm.The default CLI requires one node with 8 AMD Instinct GPUs, Docker with access to
/dev/kfdand/dev/dri, internet access, and about 45 GB of free disk. Heavy dependencies run in the prepared ROCm image.For
online/streamingextraction: a vLLM-ROCm build whose worker exposes the hidden-state extraction hook.For quantized targets:
quark.torch(MXFP4/FP8 verifier loading).Data synthesis (
data/synth.py) depends only on the standard library andtqdm.
Quickstart: Python API#
The API surface is intentionally small:
import quark.experimental.speculative_decoding as qsd
# 1) target -> trainable EAGLE-3 spec model (frozen target verifier + fresh draft)
spec = qsd.convert(
"Qwen/Qwen3-8B",
spec_cfg={
"method": "eagle3",
"aux_hidden_layers": [2, -3, -1], # low / mid / high target layers
"ttt_length": 7,
"eagle_architecture_config": {"num_hidden_layers": 1, "fc_norm": True, "norm_output": True},
},
)
# 2) cold-start training (online extraction by default)
qsd.train(spec, data_cfg=qsd.DataConfig(train="data/onpolicy.jsonl", eval="data/eval.jsonl"),
train_cfg=qsd.TrainConfig(num_epochs=1, output_dir="ckpts/eagle3"))
# 3) export a vLLM-loadable EAGLE-3 draft
qsd.export_hf(spec, "release/draft_hf")
Data utilities and evaluation:
qsd.data.synthesize(prompts, target_endpoint, served_model_name, out) # on-policy data
qsd.data.calibrate_draft_vocab(tokenizer, data, draft_vocab_size=32000) # optional lm_head compression
qsd.eval.acceptance(draft_hf, target, dataset, target_endpoint=...) # served AL@NST
qsd.eval.throughput_sweep(endpoint, served_model_name, conc=[1,8,32], isl_osl=[(1024,1024)])
Serve the exported draft with vLLM:
vllm serve <target> --speculative-config \
'{"method":"eagle3","model":"release/draft_hf","num_speculative_tokens":3}'
Quickstart: YAML recipe and CLI#
A working Docker environment is
required before running the CLI; verify that docker info succeeds.
Prepare the validated environment once, then launch the packaged Qwen3-8B recipe:
python3 -m quark.experimental.speculative_decoding.setup \
--base_model Qwen/Qwen3-8B
python3 -m quark.experimental.speculative_decoding.run --base_model Qwen/Qwen3-8B
The CLI performs on-policy data generation, 8-GPU training, export, and measured serving validation. See the complete EAGLE-3 Quick Start for the explicit YAML form, quick/full profiles, hardware requirements, runtime cost, output layout, benchmark interpretation, and troubleshooting.
Module layout#
Path |
Purpose |
|---|---|
|
Typed configs ( |
|
|
|
Config-driven entry point. Defaults to the validated TorchSpec runner;
|
|
On-policy |
|
Target hidden-state extraction: |
|
EAGLE-3 draft |
|
Cold-start TTT |
|
|
|
|
|
Quark quantization integration (quantized target as verifier; optional draft PTQ/QAT hook). |
|
Ready-to-run YAML recipes. |
|
Logging, checkpointing, ROCm environment, and extraction watchdog. |
Extraction regimes#
All three regimes deliver the same training signal (the target’s auxiliary hidden states per position). They differ in where the target runs and how features reach the trainer.
Regime |
When to use it |
Status |
|---|---|---|
|
Target co-located with the trainer (for example Qwen3-8B). |
Implemented; the recommended path for the library. |
|
Dump hidden states to disk, then train the draft alone. |
Provides |
|
A live vLLM serve streams hidden states (large MoE targets). |
Interface and environment checks only; no concrete transport is implemented. Use the TorchSpec-based example for scale. |
Quark quantization integration#
A supported setup is a quantized target (MXFP4/FP8) plus a
speculative draft. quant/integrate.py loads a Quark-quantized target as the
verifier and hidden-state source (via
quark.torch.import_model_from_safetensors, falling back to plain Hugging Face
loading), and exposes a hook to optionally quantize the draft itself. This is the
Quark-specific angle of the package. See recipes/eagle3_mxfp4_moe.yaml and
the large-model recipe template.
Evaluation#
Acceptance length (AL) is the quality and diagnostic metric.
eval/acceptance.pydrives a running vLLM speculative serve and reads the real AL from the engine’s spec-decode Prometheus/metricsendpoint. Served AL is the metric to trust; in-training proxies overestimate it.Throughput is the deployment and SLA metric.
eval/throughput_sweep.pysweeps concurrency against (ISL, OSL) pairs and reports tokens/s per GPU.
Recipes#
Recipe |
Target |
Extraction |
|---|---|---|
|
Validated Qwen3-8B full-vocabulary recipe. |
|
|
Portable large-model baseline for |
|
Relationship to the example and TorchSpec#
The end-to-end example contains the canonical runner assets used by the default Python CLI. Those assets are included in wheels, so the CLI works from any directory after installation.
This package is the Quark-native reference implementation of the same pipeline. Its export format matches vLLM, and its distinguishing feature is Quark quantization integration.
The large-model recipe template documents the model-adapter configuration and large-model workflow.
Status and limitations#
Experimental. APIs may change.
The draft is a single EAGLE-3 layer:
num_hidden_layersmust be1, because the vLLMLlamaForCausalLMEagle3serving format is single-layer. Other values raise an error.onlineextraction is the implemented training path.offlineis partial andstreamingis an interface stub, as described in Extraction regimes.Export is byte-validated against vLLM, but the native trainer is not yet scale-validated here. Use the example for a reproduced speedup.