Ornith-1.0-35B-MTP (GGUF)

Ornith-1.0-35B (DeepReinforce) with an embedded MTP (Multi-Token Prediction / nextn) head grafted in, enabling self-speculative decoding in llama.cpp at identical output quality.

Which file?

file size decode t/s acceptance notes
ornith-1.0-35b-MTP-Q4_K_M.gguf 20.6 GiB ~80 0.847 recommended
ornith-1.0-35b-MTP-Q8_0.gguf 35.2 GiB ~63–66 0.859 near-lossless weights

Sizes are weights only — budget additional headroom for the KV cache and compute buffers, which grow with context length. At long context (128K) plan well above the file size.

Both carry the MTP head at Q8_0 precision. That matters: the nextn.eh_proj tensor is only ~9 MB but it largely determines draft acceptance — quantizing it down costs a substantial slice of the speedup, so it is kept at Q8 even in the Q4_K_M build.

Why

Ornith-1.0-35B is an agentic coder fine-tuned from Qwen3.6-35B-A3B (same qwen35moe architecture, same tokenizer). The base Qwen ships with an embedded MTP head; Ornith's release does not — so it decodes without self-speculation (~50 t/s at Q8_0). Because the fine-tune barely shifts the relevant hidden states, the base model's MTP head transfers almost perfectly when grafted in.

Notably the head also transfers down the quant ladder: moving from a Q8_0 body to a Q4_K_M body costs only ~1.2 points of acceptance (0.859 → 0.847), so the smaller build keeps essentially all of the benefit.

Measured

Radeon 8060S / Strix Halo (gfx1151), llama.cpp build 387 (571d0d5), -c 8192, temp 0.

build backend decode t/s draft acceptance mean accepted len
Q4_K_M + MTP Vulkan 80.2 0.847 3.39
Q4_K_M + MTP ROCm/HIP 61.0 0.870 3.28
Q8_0 + MTP Vulkan 63–66 0.859 3.05
Q8_0, no MTP (reference) Vulkan ~50

Prefer Vulkan for this model — it is ~31% faster than ROCm/HIP on the same build, while acceptance is statistically identical across backends (acceptance is a property of the math, not the kernel), so the gap is pure kernel throughput.

⚠️ ROCm/HIP on gfx1151: throughput measured, output not validated at long context. The ROCm row above is a speed measurement from short prompts only. A different K-quant model on this hardware produced corrupted output on ROCm at long context while being clean on Vulkan, and the ROCm figure here was never checked for correctness at depth. Treat ROCm/HIP as unverified for this model and use Vulkan.

The weights are unchanged → output quality is identical; the speedup is pure self-speculation.

Usage (llama.cpp)

llama-server -m ornith-1.0-35b-MTP-Q4_K_M.gguf \
  --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.6 \
  -fa on -ngl 99 -c 131072 --jinja --alias ornith

The MTP head is embedded — no separate draft model (-md) required.

--spec-draft-n-max is worth a quick sweep on your hardware: this MoE peaks around 4, and drafting deeper eventually costs more than it wins because acceptance falls with draft depth.

How it was made (reproducible)

The donor MTP block is blk.40 — a full nextn layer: attention + MoE experts + nextn.{eh_proj, enorm, hnorm, shared_head_norm}, 20 tensors — originating from Qwen3.6-35B-A3B-Q8_0.gguf. Splice with gguf-py:

  1. Copy all of the target's tensors + metadata (raw quantized round-trip — no dequant/requant). Skip the reader's virtual GGUF.* fields, which GGUFWriter emits itself.
  2. Append the donor's 20 blk.40.* tensors.
  3. Set qwen35moe.block_count = 41 and add qwen35moe.nextn_predict_layers = 1.

The Q4_K_M build was produced the same way, using the Q8_0 build above as the donor so the head keeps its Q8 precision on top of a Q4_K_M body.

Both bases share the architecture + tokenizer, so the head plugs in directly with no retraining.

Licensing & attribution

A derivative of two permissively-licensed models; both are credited and their licenses apply to their respective parts:

  • Ornith-1.0-35B — © DeepReinforce — MIT — the base weights (733 of 753 tensors).
  • Qwen3.6-35B-A3B — © Alibaba Cloud / Qwen — Apache-2.0 — the grafted MTP (blk.40) head.

Quantizations: Q4_K_M (Q8_0 MTP head) and Q8_0. Not affiliated with or endorsed by DeepReinforce or the Qwen team.

Downloads last month
658
GGUF
Model size
36B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for singulared/Ornith-1.0-35B-MTP-GGUF