CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 10 upvotes

Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

QUESTION — How does synthetic pre-pretraining perform at scale and under realistic data mixtures?

This paper evaluates synthetic pre-pretraining (PPT) across five tasks, four data mixtures, four parameter scales from 500M to 7B, and budgets up to 100B tokens. Results show that PPT downstream performance and token efficiency gains persist at scale, such as saving at least 21B tokens at the 3B scale. However, these gains stem from tasks improving long-range retrieval rather than a grammatical prior as previously hypothesized.

Saving at least 21B PT tokens at the 3B scale.

Spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens.

atsuki-yamaguchi · 30 Sept 2026 read the original ↗
↑