How to use from the
Use from the
llama-cpp-python library
# !pip install llama-cpp-python

from llama_cpp import Llama

llm = Llama.from_pretrained(
	repo_id="AtomicChat/ornith-35b-GGUF",
	filename="",
)
llm.create_chat_completion(
	messages = [
		{
			"role": "user",
			"content": "What is the capital of France?"
		}
	]
)
Atomic Chat Join Discord GitHub

Ornith 1.0 35B

Ornith 1.0 35B, self-quantized to GGUF by Atomic Chat. Built straight from DeepReinforce's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.

Highlights

  • 0.0B parameters: the weights this repo quantizes.
  • Context length: 262,144 tokens (256K), as published by DeepReinforce.
  • 40 layers: Mixture-of-Experts.
  • Modalities: the base model handles Text, Image; this repo ships text-only quants, it carries no vision projector.
  • Full imatrix ladder: every quant is calibrated with an importance matrix.
  • State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
  • Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scallfold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.

These GGUFs are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.

Always pass --jinja so the Ornith 1.0 35B chat template is applied. Without it the model can emit malformed turns.

Model Overview

Property Value
Base model deepreinforce-ai/Ornith-1.0-35B
Parameters 0.0B
Layers 40
Experts 256 routed (top-8)
Context length 262,144 tokens (256K)
Vocabulary 248,320
Modalities Text, Image in the base model; text only in this repo, it ships no vision projector
Architecture Mixture-of-Experts, 256 experts (top-8), 16 attention heads over 2 KV heads, Qwen3_5MoeForConditionalGeneration
This repo GGUF quants (imatrix). Quants: Q4_K_M, UD-Q4_K_XL, Q5_K_M, Q6_K, Q8_0
Ornith 1.0 35B benchmark scores

Scores are DeepReinforce's published results for the base deepreinforce-ai/Ornith-1.0-35B, not our own measurements. Quantization preserves the large majority of this; Q4_K_M and up stay close to full precision.

Choosing a quant

Quant Size Notes
Q4_K_M 21.2 GB Recommended default. Best balance of size, speed and quality.
UD-Q4_K_XL 21.5 GB Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint.
Q5_K_M 24.7 GB Higher quality, low loss.
Q6_K 28.5 GB Near lossless, noticeably lighter than Q8_0.
Q8_0 36.9 GB Effectively lossless, reference quality.

Pick the largest file that fits your (V)RAM with room for context. Q4_K_M or UD-Q4_K_XL is the sweet spot for most setups; Q6_K or Q8_0 for maximum fidelity.

Get started

Run Ornith 1.0 35B locally with:

  • Atomic Chat: the easiest path. Open the app, search AtomicChat/ornith-35b-GGUF, pick a quant, hit Use this model.
  • llama.cpp: llama-server -hf AtomicChat/ornith-35b-GGUF:Q4_K_M --jinja -c 8192
  • Ollama: ollama run hf.co/AtomicChat/ornith-35b-GGUF:Q4_K_M
  • LM Studio / Jan: search the repo id, download any quant.

Best practices

Parameter Value
temperature 1.0
top_p 1.0
top_k 20

DeepReinforce's recommended sampling configuration for deepreinforce-ai/Ornith-1.0-35B.

Run in llama.cpp

git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
./llama.cpp/build/bin/llama-server \
    -hf AtomicChat/ornith-35b-GGUF:Q4_K_M \
    --jinja -ngl 99 -c 8192 -fa on

How these were made

  1. Download deepreinforce-ai/Ornith-1.0-35B (original weights).
  2. Convert to f16 GGUF with llama.cpp.
  3. Build an importance matrix over our calibration corpus.
  4. Quantize the ladder with --imatrix.
  5. UD-Q4_K_XL additionally pins the token-embedding and output tensors to Q8_0.

License

Original model by DeepReinforce, released under the MIT license. Full terms: MIT. Quantized by Atomic Chat.

Downloads last month
3,951
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AtomicChat/ornith-35b-GGUF

Quantized
(151)
this model

Collection including AtomicChat/ornith-35b-GGUF