How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "AtomicChat/ornith-35b-MLX-4bit"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "AtomicChat/ornith-35b-MLX-4bit" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links
Atomic Chat Join Discord GitHub

Ornith 1.0 35B

Ornith 1.0 35B, self-quantized to MLX by Atomic Chat. Built straight from DeepReinforce's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.

Highlights

  • 0.0B parameters: the weights this repo quantizes.
  • Context length: 262,144 tokens (256K), as published by DeepReinforce.
  • 40 layers: Mixture-of-Experts.
  • Modalities: Text, Image.
  • Full imatrix ladder: every quant is calibrated with an importance matrix.
  • State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
  • Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scallfold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.

These MLXs are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.

Model Overview

Property Value
Base model deepreinforce-ai/Ornith-1.0-35B
Parameters 0.0B
Layers 40
Experts 256 routed (top-8)
Context length 262,144 tokens (256K)
Vocabulary 248,320
Modalities Text, Image
Architecture Mixture-of-Experts, 256 experts (top-8), 16 attention heads over 2 KV heads, Qwen3_5MoeForConditionalGeneration
This repo MLX weights

Get started

  • Atomic Chat: search AtomicChat/ornith-35b-MLX-4bit and hit Use this model.
  • mlx-lm: mlx_lm.generate --model AtomicChat/ornith-35b-MLX-4bit --prompt "Hello" --max-tokens 512
  • Server: mlx_lm.server --model AtomicChat/ornith-35b-MLX-4bit --port 8080

Best practices

Parameter Value
temperature 1.0
top_p 1.0
top_k 20

DeepReinforce's recommended sampling configuration for deepreinforce-ai/Ornith-1.0-35B.

How these were made

  1. Download deepreinforce-ai/Ornith-1.0-35B (original weights).
  2. Convert and quantize with mlx_lm.convert on our pipeline.

License

Original model by DeepReinforce, released under the MIT license. Full terms: MIT. Quantized by Atomic Chat.

Downloads last month
1,071
Safetensors
Model size
35B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AtomicChat/ornith-35b-MLX-4bit

Quantized
(152)
this model

Collection including AtomicChat/ornith-35b-MLX-4bit