File size: 13,660 Bytes
337c554 0ac7575 337c554 0ac7575 337c554 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 | ---
library_name: speculators
base_model:
- qwen3/qwen3-8b
license: apache-2.0
tags:
- speculative-decoding
- dflash
- speculators
---
# RedHatAI/Qwen3-8B-speculator.dflash
This is a DFlash speculator model for [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).
## Training Details
This model was trained using the [Speculators](https://github.com/vllm-project/speculators) library on a subset of [Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered](https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered) and the `train_sft` split of [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k). Responses were regenerated by Qwen3-8B (with reasoning). Training compute for this model was sponsored by [Modal](https://modal.com).
<details>
<summary>Commands</summary>
Using the [Speculators](https://github.com/vllm-project/speculators) library and the helper scripts provided in the repo.
### Prepare data
```bash
# In virtual environment with speculators installed
python scripts/prepare_data.py \
--model Qwen/Qwen3-8B
--data ./regenerated_data.jsonl \
--output ./output \
--seq-length 8192
```
### Launch vLLM
```bash
# In (separate) virutal environment with vllm installed
CUDA_VISIBLE_DEVICES=0,1 vllm_venv/bin/python scripts/launch_vllm.py \
Qwen/Qwen3-8B \
--target-layer-ids 2 10 18 26 34 \
-- --port 8000 \
--gpu-memory-utilization 0.9 \
--disable-uvicorn-access-log \
--tensor-parallel-size 1 \
--data-parallel-size 2
```
### Launch training
Must be run once vLLM has finished launching and is running in the background.
```bash
# In virtual environment with speculators installed
CUDA_VISIBLE_DEVICES=2,3 torchrun \
--standalone \
--nproc_per_node 2 \
scripts/train.py \
--verifier-name-or-path Qwen/Qwen3-8B \
--speculator-type dflash \
--num-layers 5 \
--data-path ./output \
--vllm-endpoint http://localhost:8000/v1 \
--save-path ./output/checkpoints \
--epochs 3 \
--lr 0.0006 \
--total-seq-len 8192 \
--on-missing generate \
--on-generate delete \
--seed 42 \
--log-freq 100 \
--draft-vocab-size 32000 \
--draft-arch qwen3 \
--target-layer-ids 2 10 18 26 34 \
--draft-hidden-act silu \
--scheduler-type cosine \
--max-anchors 3072 \
--prefetch-factor 2 \
--num-workers 8
```
</details>
## Model Specifications
| | |
|---|---|
| **Base Model** | Qwen/Qwen3-8B |
| **Chat Template** | Qwen/Qwen3-8B (use `/chat/completions` endpoint) |
| **Format** | Safetensors |
| **License** | Apache 2.0 |
| **Validation Hardware** | Nvidia H100 |
## Deployment
```bash
# Install vLLM from the required PR
pip install git+https://github.com/vllm-project/vllm.git@refs/pull/41880/head
# Deploy with speculative decoding
vllm serve Qwen/Qwen3-8B \
--tensor-parallel-size 1 \
--max-model-len 16384 \
--speculative-config '{
"model": "RedHatAI/Qwen3-8B-speculator.dflash",
"num_speculative_tokens": 7,
"method": "dflash"
}'
```
## Preliminary Evaluations
Per-position token acceptance rates across datasets:
(with reasoning enabled)
| Dataset | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Pos 7 | Avg Length |
|---------|-------|-------|-------|-------|-------|-------|-------|------------|
| HumanEval | 79.9% | 58.0% | 40.3% | 27.0% | 17.8% | 11.3% | 6.8% | 3.410 |
| math_reasoning | 82.2% | 62.7% | 46.2% | 33.5% | 23.4% | 15.8% | 9.9% | 3.740 |
| qa | 68.9% | 42.6% | 25.0% | 14.4% | 8.1% | 4.4% | 2.3% | 2.660 |
| question | 73.0% | 47.6% | 30.1% | 18.9% | 11.7% | 7.1% | 4.1% | 2.930 |
| rag | 71.1% | 44.8% | 27.0% | 15.7% | 8.9% | 4.9% | 2.5% | 2.750 |
| summarization | 65.5% | 36.1% | 19.0% | 9.5% | 4.7% | 2.3% | 1.1% | 2.380 |
| tool_call | 71.3% | 44.6% | 25.8% | 14.4% | 7.8% | 4.1% | 2.1% | 2.700 |
| translation | 63.8% | 38.4% | 22.1% | 11.8% | 6.1% | 3.2% | 1.5% | 2.470 |
| writing | 73.2% | 47.7% | 30.1% | 18.9% | 11.8% | 7.2% | 4.2% | 2.930 |
## References
**Paper**: [DFlash: Block Diffusion for Flash Speculative Decoding](https://arxiv.org/abs/2602.06036) |