Loading Complexity
Loading Complexity
[Model releases]
130B pretrain → 32.07B refinement → full SFT
Official preprint
Deterministic multi-hash routing supports long-horizon training in a compact language model
DOI 10.21203/rs.3.rs-10788774/v1
201.2M
parameters
≈162B
source-token exposure
69.10%
PIQA acc_norm (SFT epoch 3)
The base pretraining run completed its 130B-token replay schedule. A fresh-optimizer full-parameter refinement then reached step 8,156 / 17,802, adding approximately 32.07B unique-token exposures before it was intentionally stopped; the refinement release is therefore an evaluated intermediate checkpoint, not a completed 70B pass.
The released assistant is the 32,000-token full-parameter SFT trained directly from the refinement checkpoint for three epochs on the audited 300K mixture. Epoch 3 is published at the repository root and reaches 68.01% PIQA accuracy and 69.10% normalized accuracy.
The cited checkpoint reaches 57.24% ARC-Easy, 27.13% ARC-Challenge and 47.29% Combined ARC on the complete public zero-shot splits. These results, the architecture and the complete training lineage are documented in the public Research Square preprint.
125B pretraining lineage → refinement checkpoint → full SFT
100.4M
parameters
500K
SFT training examples
25K
held-out examples
1.317482
matched eval loss
3.73
matched eval perplexity
The released assistant contains exactly 100,366,720 parameters and starts from checkpoint 9,791 in the public 100M Agentic refinement archive. It preserves deterministic top-2 token-ID routing across four experts, a shared SwiGLU path, 10 transformer layers and causal GQA.
Full-parameter instruction tuning ran for three epochs on 500,000 training examples selected from the 2M SFT corpus, with 25,000 held-out examples and protected n-gram exclusion for PIQA, ARC Easy, ARC Challenge, GSM8K and HellaSwag. The run uses the model's native 32K Agentic tokenizer and chat boundaries with a 2,048-token sequence length.
The final epoch-3 checkpoint at step 9,093 reached matched validation loss 1.317482 and perplexity 3.73. Its mixture covers general instructions, verified reasoning and arithmetic, calculator and general tool use, execution-checked code, and constraint following. Public zero-shot benchmark results for this new 500K checkpoint will be reported separately.
Compact object detection
2.53M
parameters
20.05
AP50-95 (SFT)
62.3%
of 32.2 AP YOLO26 reference
A hash-routed detector: a dense shared SwiGLU branch plus 8 narrow experts, top-2 routed by a deterministic hash of spatial-grid position — no learned gating, no auxiliary load-balancing loss.
Two-stage recipe — from-scratch pretraining (Mosaic + MixUp, MuSGD) followed by a full-parameter, clean-image supervised fine-tuning stage. SFT alone moved AP50-95 from 16.59 to 20.05 (+20.9% relative), reaching 62.3% of the 32.2 AP YOLO26 reference used on our model cards. Single random-init run, no hyperparameter search, neither stage had plateaued.
Independently reproduced — thanks to the community: O2M+NMS mAP50 0.325 / mAP50-95 0.200 / AR100 0.379, NMS-free mAP50 0.140 / mAP50-95 0.096 — matching our reported numbers.
First end-to-end pretrain — the guinea pig
492.1M
parameters
20B
tokens
4
routed experts
The first model that proved the architecture could learn end-to-end: deterministic token-identity routing across four residual experts, with one always-on shared SwiGLU path carrying shared context.
An SFT follow-up on this base (LoRA, one full-shard epoch) wasn't pretrained long enough to pass our behavioral promotion gate — it isn't published as an assistant. But the metrics show real learning, not a dead run: full PIQA accuracy was retained (0.6953 → 0.6964) and matched eval loss dropped (3.68 → 2.98). The base simply needed more pretraining before instruction-following behavior could reliably stick — an instruction-coverage limit, not a capability regression.