The repo for the first in the Biggerbrain 3 lineup, a 150m parameter model.
This model uses a Recurrent transformer(looping the central 5 layers), and soft averaged MoE architecture to achieve new levels of reasoning(reletive to the Biggerbrain lineup). This model used an improved training pipeline*, and a substantial parameter increase over V2, and a great boost to intelligence, focus, and overall usefullness. It used a custom trained 32k BPE tokenizer, with pre-tokenized .bin files saved in Uint16 format for random sampling of individual topics. The 5 topics all datasets were compiled into were as follows:
Text(Fineweb-edu), Books(cleaned project gutenberg), Math(a tiny collection of math datasets), reasoning(some reasoning, some 3d visualization questions, etc.), and Code(Some competition code & python-edu-cleaned).
For training I used 3 stages, slowly increasing the amount of Code, Math, and Code, to help prepare the model for SFT, and deployement.
Pretraining has finished, and the model is now going through SFT in order to achieve higher competence in user-chat scenarios, aswell as recovering the ability to code**. For SFT, we are using a ~1 epoch on a combined dataset including "Jackrong/Competitive-Programming-python-blend"(Which was also included in pretraining), "allenai/Dolci-Think-SFT-Python", "open-r1/OpenR1-Math-220k", "Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B", and "KingNish/reasoning-base-20k". All of these datasets had ~5k examples downloaded Except for OpenR1 math, which had 10k downloaded(it was the only math dataset; the others were 5k + another 5k for both reasoning and code)
*The training pipeline was modified to use percentage based sampling of datasets, warmup-stable-decay LR, and an MTP head(+1 extra token) to increase model peak intelligence. It integrated Muon aswell as adamW8bit.
**The model's tokenizer had python indentation stripped during pretraining, and thus we are using heavily code-weighted sampling during Supervised Fine Tuning to hopefully recover its ability to indent python correctly. Either way though, V4 will have fixed this issue with tokenization, and biggerbrain V3 is still a strong model at math, reasoning and creative tasks.