Some evals on coding and results!

#4
by eidolon08 - opened

pi-bench results on single strix halo: https://pi-local-coding-bench.dev/

MODELS: IQ1_M MTP
SUCCESS RATE: 58.0%
TASKS: 29 / 50
AVG DURATION: 20m 30s

The results were quite lower than https://huggingface.co/YanissAmz/Hy3-295B-A21B-GGUF 's Hy3-UD128-*.gguf
Which scored
SUCCESS RATE: 68.0%
TASKS: 34 / 50
AVG DURATION: 14m 20s

I'm gonna try the Q2_K! Thx for the quantization!

full results on: https://github.com/kyuz0/pi-bench/pull/5

AngelSlim org
edited 4 days ago

pi-bench results on single strix halo: https://pi-local-coding-bench.dev/

MODELS: IQ1_M MTP
SUCCESS RATE: 58.0%
TASKS: 29 / 50
AVG DURATION: 20m 30s

The results were quite lower than https://huggingface.co/YanissAmz/Hy3-295B-A21B-GGUF 's Hy3-UD128-*.gguf
Which scored
SUCCESS RATE: 68.0%
TASKS: 34 / 50
AVG DURATION: 14m 20s

I'm gonna try the Q2_K! Thx for the quantization!

full results on: https://github.com/kyuz0/pi-bench/pull/5

Thanks for checking it out and sharing the benchmark results. The Hy3-UD128 model is an IQ3 quant (116.7 GB), so it might have higher precision than our IQ1 version (89.4 GB). We’ve now also uploaded a Q2_K_XL version (about 100 GB), which sits between the two in size while leaving more memory for the context/KV cache.

I'll also share results for Q2_K_XL too!

Tried Q2_K_XL MTP as well — thx for the quantization!

MODELS: Q2_K_XL MTP
SUCCESS RATE: 68.0%
TASKS: 34 / 50
AVG DURATION: 17m 53s

Q2_K_XL matched UD128 on pass rate (34/50) and sits between IQ1_M and UD128 on speed.

Full results:
• Q2_K_XL MTP: https://github.com/kyuz0/pi-bench/pull/6

Sign up or log in to comment