ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
QUESTION — How can a compact foundation model overcome parametric capacity limits by coupling internal thinking with external tool use?
The authors present ZGCM-1, a fully open 7B dense foundation model trained from scratch with high efficiency for math and agentic search. The model features interleaved gated sliding-window and full attention, a stable FP8 Muon optimizer, context scaling across 16K, 64K, and 256K, and Markov Decision Processes formatting. Evaluations show ZGCM-1-7B is competitive on general benchmarks and challenging mathematical reasoning suites, while delivering a ~4.2x efficiency improvement in 16K pre-training time-to-loss.
The pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss.
yshenaw · 11 Sept 2026
read the original ↗