File size: 13,660 Bytes
337c554
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0ac7575
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
337c554
 
 
 
 
 
 
 
 
0ac7575
337c554
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
---
library_name: speculators
base_model:
- qwen3/qwen3-8b
license: apache-2.0
tags:
- speculative-decoding
- dflash
- speculators
---
                                                                                                                                                                                                                                                                                                        
# RedHatAI/Qwen3-8B-speculator.dflash                                                                                                                                                                                                                                                                     
   
This is a DFlash speculator model for [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B).                                                                                                                                                                                                              
                       
## Training Details                                                                                                                                                                                                                                                                                       
                       
This model was trained using the [Speculators](https://github.com/vllm-project/speculators) library on a subset of [Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered](https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered) and the `train_sft` split of [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k). Responses were regenerated by Qwen3-8B (with reasoning). Training compute for this model was sponsored by [Modal](https://modal.com). 

<details>
  <summary>Commands</summary>

  Using the [Speculators](https://github.com/vllm-project/speculators) library and the helper scripts provided in the repo.

  ### Prepare data
  ```bash
  # In virtual environment with speculators installed
  python scripts/prepare_data.py \
    --model Qwen/Qwen3-8B
    --data ./regenerated_data.jsonl \
    --output ./output \
    --seq-length 8192
  ```

  ### Launch vLLM
  ```bash
  # In (separate) virutal environment with vllm installed
  CUDA_VISIBLE_DEVICES=0,1 vllm_venv/bin/python scripts/launch_vllm.py \
    Qwen/Qwen3-8B \
    --target-layer-ids 2 10 18 26 34 \
    -- --port 8000 \
    --gpu-memory-utilization 0.9 \
    --disable-uvicorn-access-log \
    --tensor-parallel-size 1 \
    --data-parallel-size 2
  ```

  ### Launch training
  Must be run once vLLM has finished launching and is running in the background.
  ```bash
  # In virtual environment with speculators installed
  CUDA_VISIBLE_DEVICES=2,3 torchrun \
    --standalone \
    --nproc_per_node 2 \
    scripts/train.py \
    --verifier-name-or-path Qwen/Qwen3-8B \
    --speculator-type dflash \
    --num-layers 5 \
    --data-path ./output \
    --vllm-endpoint http://localhost:8000/v1 \
    --save-path ./output/checkpoints \
    --epochs 3 \
    --lr 0.0006 \
    --total-seq-len 8192 \
    --on-missing generate \
    --on-generate delete \
    --seed 42 \
    --log-freq 100 \
    --draft-vocab-size 32000 \
    --draft-arch qwen3 \
    --target-layer-ids 2 10 18 26 34 \
    --draft-hidden-act silu \
    --scheduler-type cosine \
    --max-anchors 3072 \
    --prefetch-factor 2 \
    --num-workers 8
  ```
  
</details>

                                                                                                                                                                                                                                                                                                            
## Model Specifications                                   

| | |
|---|---|
| **Base Model** | Qwen/Qwen3-8B |
| **Chat Template** | Qwen/Qwen3-8B (use `/chat/completions` endpoint) |                                                                                                                                                                                                                                  
| **Format** | Safetensors |                                                                                                                                                                                                                                                                              
| **License** | Apache 2.0 |                                                                                                                                                                                                                                                                              
| **Validation Hardware** | Nvidia H100 |                                                                                                                                                                                                                                                                 
                                                                                                                                                                                                                                                                                                            
## Deployment                                             
                                                                                                                                                                                                                                                                                                            
  ```bash                                                                                                                                                                                                                                                                                                   
  # Install vLLM from the required PR
  pip install git+https://github.com/vllm-project/vllm.git@refs/pull/41880/head                                                                                                                                                                                                                             
                                                                                                                                                                                                                                                                                                            
  # Deploy with speculative decoding                                                                                                                                                                                                                                                                        
  vllm serve Qwen/Qwen3-8B \                                                                                                                                                                                                                                                                                
      --tensor-parallel-size 1 \                                                                                                                                                                                                                                                                            
      --max-model-len 16384 \                               
      --speculative-config '{                                                                                                                                                                                                                                                                               
          "model": "RedHatAI/Qwen3-8B-speculator.dflash",                                                                                                                                                                                                                                                   
          "num_speculative_tokens": 7,                                                                                                                                                                                                                                                                      
          "method": "dflash"                                                                                                                                                                                                                                                                                
      }'
```                                                                                                                                                                                                                                                                                                 
                                                                                                                                                                                                                                                                                                            
## Preliminary Evaluations                                                                                                                                                                                                                                                                                   
                                                            
Per-position token acceptance rates across datasets:                                                                                                                                                                                                                                                      
(with reasoning enabled)

| Dataset | Pos 1 | Pos 2 | Pos 3 | Pos 4 | Pos 5 | Pos 6 | Pos 7 | Avg Length |                                                                                                                                                                                                                          
  |---------|-------|-------|-------|-------|-------|-------|-------|------------|
  | HumanEval | 79.9% | 58.0% | 40.3% | 27.0% | 17.8% | 11.3% | 6.8% | 3.410 |                                                                                                                                                                                                                              
  | math_reasoning | 82.2% | 62.7% | 46.2% | 33.5% | 23.4% | 15.8% | 9.9% | 3.740 |                                                                                                                                                                                                                         
  | qa | 68.9% | 42.6% | 25.0% | 14.4% | 8.1% | 4.4% | 2.3% | 2.660 |                                                                                                                                                                                                                                       
  | question | 73.0% | 47.6% | 30.1% | 18.9% | 11.7% | 7.1% | 4.1% | 2.930 |                                                                                                                                                                                                                                
  | rag | 71.1% | 44.8% | 27.0% | 15.7% | 8.9% | 4.9% | 2.5% | 2.750 |                                                                                                                                                                                                                                      
  | summarization | 65.5% | 36.1% | 19.0% | 9.5% | 4.7% | 2.3% | 1.1% | 2.380 |                                                                                                                                                                                                                             
  | tool_call | 71.3% | 44.6% | 25.8% | 14.4% | 7.8% | 4.1% | 2.1% | 2.700 |                                                                                                                                                                                                                                
  | translation | 63.8% | 38.4% | 22.1% | 11.8% | 6.1% | 3.2% | 1.5% | 2.470 |                                                                                                                                                                                                                              
  | writing | 73.2% | 47.7% | 30.1% | 18.9% | 11.8% | 7.2% | 4.2% | 2.930 |  
                                                                                                                                                                                                                                                                                                                                                                                                                  
                                                            
## References

**Paper**: [DFlash: Block Diffusion for Flash Speculative Decoding](https://arxiv.org/abs/2602.06036)