MTP Supported?

#3
by alexanderacuna - opened

Was the MTP feature tested with these quants? The base model has MTP layers but MTP is actually reducing throughput. Wondering if these specific quants created my unsloth were tested with the spec decoding.

Ref:
https://github.com/ggml-org/llama.cpp/issues/23924

Any info?

Sign up or log in to comment