Parallax 8B lambda: training checkpoints
These are research checkpoints from lambda, a training run of Parallax. Training was stopped at about 900B tokens, before the planned end of the run. The designated final checkpoint is step 73973 (900.8B tokens), marked in the table below. Exports saved later than the final checkpoint were saved after training stopped and are not the final checkpoint. Parallax is a mixture-of-experts language model trained on consumer GPUs spread over several countries. The hosts have no direct connections to each other and synchronize over the public internet.
- Base model. It is pretrained on a mixture of web, document, math, code, knowledge-focused and book text. It is not instruction-tuned, chat-tuned or safety-tuned, and it will continue text rather than follow instructions.
- Training snapshots. Every export, the final one included, is a snapshot taken before the planned ~973B tokens. Quality between exports is not monotonic.
- Research artifact. It is published so the training can be followed and inspected. It is not intended for production use.
Live training dashboard: parallax.chutes.ai.
Model
| Parameters | ~7.77B total, ~1.20B active per token |
| Layers | 64: 32 mixers (18 EDA recurrent, 8 sliding-window attention, 6 global attention) interleaved with 32 mixture-of-experts layers |
| Recurrent mixers | EDA recurrence (a gated delta rule) with per-key-channel decay bounded below by e^-5 per token; erase rank 16; output normalization |
| Sliding-window attention | First and last mixers use 2048 tokens; the six interior layers share a per-step draw of 128 or 2048 tokens with equal probability throughout training; evaluation and inference use 2048 |
| Global attention | sparse-capable global attention (MSA) with 256-dimensional latent KV, 16 query heads, 4 KV heads and RoPE; dense causal attention throughout this run, with sparse top-k selection inactive |
| Experts | 4096 routed experts (128 per MoE layer), 12 routed + 1 shared expert per token |
| Expert weights | ternary (-1, 0, +1) with per-row scales; at most two nonzero pairs in every group of eight |
| Width | Model width 1152; expert intermediate width 2304, latent width 384; ReLUยฒ expert activation |
| Tokenizer | Llama 3 tokenizer (vocabulary padded to 128,384); tied input/output embeddings |
| Training context | 4096 tokens |
| Document masking | In training, attention and recurrent state are reset at document boundaries (up to 64 documents per 4096-token sequence; further boundaries are merged into the 64th document) |
| Output logit scale | learnable, capped at 3.0 |
Training
- Data: seven stationary sources totalling about 1.09T tokens: two general pretraining mixtures of web, document, math and code text, an additional web set, two math sets, a knowledge-focused set and a books set. At about 899.6B tokens the data switched to two further sources while the learning-rate anneal that began at 877.28B continued. Training stopped at about 900B tokens, shortly after this switch. The planned training total was about 972.9B tokens.
- Optimization: the dense trunk is synchronized with a decoupled DiLoCo scheme over the internet. Routed experts are trained through low-rank adapters on the GPUs that use them, and the adapter updates are folded into full-precision masters that publish new ternary versions. The tech report describes the details.
- Learning rate: linear warm-up from 0 to 6e-4 over the first 19.66B tokens, then 6e-4 until 400B tokens, 3e-4 until 700B and 1.5e-4 until 877.28B. A live anneal started at 877.28B tokens on 2026-10-09, with multiplier m = 1 - 0.9 sqrt((T - 877.28B) / (972.90B - 877.28B)) and learning rate m times 1.5e-4. When training stopped, the multiplier had reached 0.5489 (learning rate 8.234e-5). The anneal was partial.
- Batch: 10 gradient-accumulation micro-batches per step at launch (about 8.5M tokens per fleet step with 26 machines), switched to 20 at about 100B tokens on 2026-10-04 at 22:08 UTC. This gives about 17M tokens per step with all 26 machines; tokens per step vary with fleet size.
- Fleet: 26 machines and 208 GPUs at launch. Seven machines were lost over 5-8 October, with roles reassigned automatically while training continued. At the end of the run, 19 machines were assigned, totalling 152 GPUs.
- Document masking: In training, attention and recurrent state are reset at document boundaries (up to 64 documents per 4096-token sequence; further boundaries are merged into the 64th document).
- Output logits: the learnable output logit scale is capped at 3.0 in training and inference (see Files and format).
Exports
One folder per export under exports/, named by training tokens (rounded down to whole billions) and
fleet step. Exports were added about every 30 minutes while training ran. Training has stopped; the final checkpoint is marked in the table. Published exports are
never modified or removed.
| Export | Tokens | Step | Size | Exported (UTC) |
|---|---|---|---|---|
901B-tokens_step73973_fc88abba |
901.5B | 73973 | 2.43 GB | 2026-10-10 01:44 |
901B-tokens_step73973_6a749265 |
901.5B | 73973 | 2.43 GB | 2026-10-10 02:45 |
901B-tokens_step73973_b08bb757 |
901.5B | 73973 | 2.43 GB | 2026-10-10 02:18 |
901B-tokens_step73973 |
901.5B | 73973 | 2.43 GB | 2026-10-10 01:17 |
900B-tokens_step73973 final |
900.8B | 73973 | 2.43 GB | 2026-10-10 00:45 |
898B-tokens_step73878 |
898.7B | 73878 | 2.43 GB | 2026-10-10 00:18 |
895B-tokens_step73637 |
896.0B | 73637 | 2.43 GB | 2026-10-09 23:49 |
892B-tokens_step73364 |
892.8B | 73364 | 2.43 GB | 2026-10-09 23:18 |
890B-tokens_step73127 |
890.0B | 73127 | 2.43 GB | 2026-10-09 22:48 |
886B-tokens_step72845 |
886.9B | 72845 | 2.43 GB | 2026-10-09 22:19 |
884B-tokens_step72604 |
884.2B | 72604 | 2.43 GB | 2026-10-09 21:46 |
881B-tokens_step72330 |
881.0B | 72330 | 2.43 GB | 2026-10-09 21:17 |
878B-tokens_step72097 |
878.3B | 72097 | 2.43 GB | 2026-10-09 20:48 |
875B-tokens_step71820 |
875.2B | 71820 | 2.43 GB | 2026-10-09 20:20 |
872B-tokens_step71588 |
872.4B | 71588 | 2.43 GB | 2026-10-09 19:48 |
869B-tokens_step71310 |
869.3B | 71310 | 2.43 GB | 2026-10-09 19:19 |
866B-tokens_step71076 |
866.5B | 71076 | 2.43 GB | 2026-10-09 18:47 |
863B-tokens_step70798 |
863.4B | 70798 | 2.43 GB | 2026-10-09 18:16 |
860B-tokens_step70559 |
860.6B | 70559 | 2.43 GB | 2026-10-09 17:46 |
857B-tokens_step70296 |
857.5B | 70296 | 2.43 GB | 2026-10-09 17:17 |
854B-tokens_step70065 |
854.7B | 70065 | 2.43 GB | 2026-10-09 16:50 |
851B-tokens_step69789 |
851.7B | 69789 | 2.43 GB | 2026-10-09 16:16 |
848B-tokens_step69548 |
848.9B | 69548 | 2.43 GB | 2026-10-09 15:46 |
845B-tokens_step69283 |
845.7B | 69283 | 2.43 GB | 2026-10-09 15:15 |
843B-tokens_step69045 |
843.0B | 69045 | 2.43 GB | 2026-10-09 14:48 |
839B-tokens_step68773 |
839.9B | 68773 | 2.43 GB | 2026-10-09 14:15 |
837B-tokens_step68544 |
837.1B | 68544 | 2.43 GB | 2026-10-09 13:48 |
834B-tokens_step68273 |
834.0B | 68273 | 2.43 GB | 2026-10-09 13:15 |
831B-tokens_step68038 |
831.3B | 68038 | 2.43 GB | 2026-10-09 12:47 |
828B-tokens_step67772 |
828.2B | 67772 | 2.43 GB | 2026-10-09 12:16 |
825B-tokens_step67534 |
825.5B | 67534 | 2.43 GB | 2026-10-09 11:49 |
822B-tokens_step67258 |
822.3B | 67258 | 2.43 GB | 2026-10-09 11:16 |
819B-tokens_step67027 |
819.6B | 67027 | 2.43 GB | 2026-10-09 10:47 |
816B-tokens_step66758 |
816.4B | 66758 | 2.43 GB | 2026-10-09 10:15 |
813B-tokens_step66509 |
813.7B | 66509 | 2.43 GB | 2026-10-09 09:47 |
810B-tokens_step66235 |
810.5B | 66235 | 2.43 GB | 2026-10-09 09:16 |
807B-tokens_step66001 |
807.8B | 66001 | 2.43 GB | 2026-10-09 08:46 |
804B-tokens_step65728 |
804.7B | 65728 | 2.43 GB | 2026-10-09 08:14 |
801B-tokens_step65490 |
801.9B | 65490 | 2.43 GB | 2026-10-09 07:46 |
798B-tokens_step65214 |
798.8B | 65214 | 2.43 GB | 2026-10-09 07:16 |
796B-tokens_step64968 |
796.0B | 64968 | 2.43 GB | 2026-10-09 06:47 |
792B-tokens_step64700 |
792.9B | 64700 | 2.43 GB | 2026-10-09 06:14 |
790B-tokens_step64456 |
790.2B | 64456 | 2.43 GB | 2026-10-09 05:46 |
787B-tokens_step64188 |
787.0B | 64188 | 2.43 GB | 2026-10-09 05:13 |
784B-tokens_step63948 |
784.3B | 63948 | 2.43 GB | 2026-10-09 04:49 |
781B-tokens_step63679 |
781.1B | 63679 | 2.43 GB | 2026-10-09 04:15 |
778B-tokens_step63434 |
778.4B | 63434 | 2.43 GB | 2026-10-09 03:45 |
775B-tokens_step63160 |
775.3B | 63160 | 2.43 GB | 2026-10-09 03:15 |
772B-tokens_step62921 |
772.6B | 62921 | 2.43 GB | 2026-10-09 02:47 |
769B-tokens_step62646 |
769.4B | 62646 | 2.43 GB | 2026-10-09 02:16 |
766B-tokens_step62413 |
766.6B | 62413 | 2.43 GB | 2026-10-09 01:46 |
763B-tokens_step62131 |
763.6B | 62131 | 2.43 GB | 2026-10-09 01:16 |
760B-tokens_step61889 |
760.7B | 61889 | 2.43 GB | 2026-10-09 00:48 |
757B-tokens_step61611 |
757.6B | 61611 | 2.43 GB | 2026-10-09 00:18 |
754B-tokens_step61370 |
754.9B | 61370 | 2.43 GB | 2026-10-08 23:45 |
751B-tokens_step61097 |
751.8B | 61097 | 2.43 GB | 2026-10-08 23:15 |
748B-tokens_step60864 |
749.0B | 60864 | 2.43 GB | 2026-10-08 22:46 |
745B-tokens_step60600 |
745.9B | 60600 | 2.43 GB | 2026-10-08 22:20 |
743B-tokens_step60341 |
743.1B | 60341 | 2.43 GB | 2026-10-08 21:45 |
740B-tokens_step60075 |
740.0B | 60075 | 2.43 GB | 2026-10-08 21:16 |
737B-tokens_step59833 |
737.2B | 59833 | 2.43 GB | 2026-10-08 20:45 |
734B-tokens_step59575 |
734.2B | 59575 | 2.43 GB | 2026-10-08 20:16 |
731B-tokens_step59330 |
731.4B | 59330 | 2.43 GB | 2026-10-08 19:46 |
728B-tokens_step59064 |
728.1B | 59064 | 2.43 GB | 2026-10-08 19:16 |
725B-tokens_step58814 |
725.2B | 58814 | 2.43 GB | 2026-10-08 18:46 |
721B-tokens_step58547 |
722.0B | 58547 | 2.43 GB | 2026-10-08 18:18 |
719B-tokens_step58295 |
719.1B | 58295 | 2.43 GB | 2026-10-08 17:48 |
715B-tokens_step58037 |
715.9B | 58037 | 2.43 GB | 2026-10-08 17:16 |
712B-tokens_step57790 |
712.9B | 57790 | 2.43 GB | 2026-10-08 16:46 |
709B-tokens_step57529 |
709.7B | 57529 | 2.43 GB | 2026-10-08 16:17 |
706B-tokens_step57284 |
706.8B | 57284 | 2.43 GB | 2026-10-08 15:47 |
703B-tokens_step57012 |
703.7B | 57012 | 2.43 GB | 2026-10-08 15:16 |
700B-tokens_step56772 |
700.7B | 56772 | 2.43 GB | 2026-10-08 14:45 |
697B-tokens_step56506 |
697.5B | 56506 | 2.43 GB | 2026-10-08 14:16 |
694B-tokens_step56272 |
694.6B | 56272 | 2.43 GB | 2026-10-08 13:46 |
691B-tokens_step55992 |
691.4B | 55992 | 2.43 GB | 2026-10-08 13:19 |
688B-tokens_step55745 |
688.4B | 55745 | 2.43 GB | 2026-10-08 12:45 |
685B-tokens_step55474 |
685.2B | 55474 | 2.43 GB | 2026-10-08 12:15 |
682B-tokens_step55234 |
682.3B | 55234 | 2.43 GB | 2026-10-08 11:45 |
679B-tokens_step54971 |
679.1B | 54971 | 2.43 GB | 2026-10-08 11:14 |
676B-tokens_step54731 |
676.1B | 54731 | 2.43 GB | 2026-10-08 10:48 |
672B-tokens_step54458 |
672.9B | 54458 | 2.43 GB | 2026-10-08 10:16 |
670B-tokens_step54209 |
670.0B | 54209 | 2.43 GB | 2026-10-08 09:46 |
666B-tokens_step53941 |
666.8B | 53941 | 2.43 GB | 2026-10-08 09:14 |
663B-tokens_step53702 |
663.9B | 53702 | 2.43 GB | 2026-10-08 08:45 |
660B-tokens_step53438 |
660.7B | 53438 | 2.43 GB | 2026-10-08 08:17 |
657B-tokens_step53196 |
657.8B | 53196 | 2.43 GB | 2026-10-08 07:48 |
654B-tokens_step52696 |
654.5B | 52696 | 2.43 GB | 2026-10-08 07:17 |
651B-tokens_step52456 |
651.5B | 52456 | 2.43 GB | 2026-10-08 06:48 |
648B-tokens_step52190 |
648.2B | 52190 | 2.43 GB | 2026-10-08 06:16 |
645B-tokens_step51945 |
645.2B | 51945 | 2.43 GB | 2026-10-08 05:50 |
641B-tokens_step51674 |
641.9B | 51674 | 2.43 GB | 2026-10-08 05:15 |
638B-tokens_step51438 |
638.9B | 51438 | 2.43 GB | 2026-10-08 04:47 |
635B-tokens_step51165 |
635.5B | 51165 | 2.43 GB | 2026-10-08 04:16 |
632B-tokens_step50914 |
632.5B | 50914 | 2.43 GB | 2026-10-08 03:47 |
629B-tokens_step50646 |
629.1B | 50646 | 2.43 GB | 2026-10-08 03:17 |
626B-tokens_step50402 |
626.1B | 50402 | 2.43 GB | 2026-10-08 02:47 |
622B-tokens_step50135 |
622.7B | 50135 | 2.43 GB | 2026-10-08 02:18 |
619B-tokens_step49880 |
619.8B | 49880 | 2.43 GB | 2026-10-08 01:47 |
616B-tokens_step49610 |
616.4B | 49610 | 2.43 GB | 2026-10-08 01:22 |
613B-tokens_step49361 |
613.3B | 49361 | 2.43 GB | 2026-10-08 00:47 |
610B-tokens_step49091 |
610.0B | 49091 | 2.43 GB | 2026-10-08 00:17 |
606B-tokens_step48845 |
607.0B | 48845 | 2.43 GB | 2026-10-07 23:47 |
603B-tokens_step48573 |
603.6B | 48573 | 2.43 GB | 2026-10-07 23:17 |
600B-tokens_step48324 |
600.7B | 48324 | 2.43 GB | 2026-10-07 22:47 |
597B-tokens_step48055 |
597.3B | 48055 | 2.43 GB | 2026-10-07 22:14 |
594B-tokens_step47817 |
594.4B | 47817 | 2.43 GB | 2026-10-07 21:46 |
590B-tokens_step47547 |
591.0B | 47547 | 2.43 GB | 2026-10-07 21:18 |
587B-tokens_step47301 |
588.0B | 47301 | 2.43 GB | 2026-10-07 20:46 |
584B-tokens_step47029 |
584.6B | 47029 | 2.43 GB | 2026-10-07 20:16 |
581B-tokens_step46782 |
581.6B | 46782 | 2.43 GB | 2026-10-07 19:46 |
578B-tokens_step46521 |
578.3B | 46521 | 2.43 GB | 2026-10-07 19:16 |
575B-tokens_step45978 |
575.3B | 45978 | 2.43 GB | 2026-10-07 18:46 |
571B-tokens_step45785 |
571.9B | 45785 | 2.43 GB | 2026-10-07 18:16 |
568B-tokens_step45541 |
568.7B | 45541 | 2.43 GB | 2026-10-07 17:47 |
565B-tokens_step45273 |
565.2B | 45273 | 2.43 GB | 2026-10-07 17:14 |
562B-tokens_step45031 |
562.1B | 45031 | 2.43 GB | 2026-10-07 16:46 |
558B-tokens_step44774 |
558.5B | 44774 | 2.43 GB | 2026-10-07 16:18 |
555B-tokens_step44528 |
555.4B | 44528 | 2.43 GB | 2026-10-07 15:46 |
551B-tokens_step44262 |
551.9B | 44262 | 2.43 GB | 2026-10-07 15:17 |
548B-tokens_step44023 |
548.7B | 44023 | 2.43 GB | 2026-10-07 14:47 |
545B-tokens_step43760 |
545.2B | 43760 | 2.43 GB | 2026-10-07 14:17 |
542B-tokens_step43518 |
542.1B | 43518 | 2.43 GB | 2026-10-07 13:46 |
538B-tokens_step43262 |
538.6B | 43262 | 2.43 GB | 2026-10-07 13:17 |
535B-tokens_step43016 |
535.4B | 43016 | 2.43 GB | 2026-10-07 12:47 |
531B-tokens_step42753 |
531.9B | 42753 | 2.43 GB | 2026-10-07 12:20 |
528B-tokens_step42504 |
528.8B | 42504 | 2.43 GB | 2026-10-07 11:46 |
525B-tokens_step42242 |
525.2B | 42242 | 2.43 GB | 2026-10-07 11:21 |
522B-tokens_step42001 |
522.2B | 42001 | 2.43 GB | 2026-10-07 10:48 |
518B-tokens_step41738 |
518.6B | 41738 | 2.43 GB | 2026-10-07 10:18 |
515B-tokens_step41487 |
515.5B | 41487 | 2.43 GB | 2026-10-07 09:45 |
511B-tokens_step41224 |
511.9B | 41224 | 2.43 GB | 2026-10-07 09:15 |
508B-tokens_step40992 |
508.8B | 40992 | 2.43 GB | 2026-10-07 08:48 |
505B-tokens_step40722 |
505.3B | 40722 | 2.43 GB | 2026-10-07 08:20 |
502B-tokens_step40480 |
502.1B | 40480 | 2.43 GB | 2026-10-07 07:47 |
498B-tokens_step40216 |
498.6B | 40216 | 2.43 GB | 2026-10-07 07:17 |
495B-tokens_step39965 |
495.4B | 39965 | 2.43 GB | 2026-10-07 06:46 |
491B-tokens_step39706 |
492.0B | 39706 | 2.43 GB | 2026-10-07 06:17 |
488B-tokens_step39456 |
488.8B | 39456 | 2.43 GB | 2026-10-07 05:45 |
485B-tokens_step39196 |
485.2B | 39196 | 2.43 GB | 2026-10-07 05:19 |
482B-tokens_step38951 |
482.1B | 38951 | 2.43 GB | 2026-10-07 04:44 |
478B-tokens_step38687 |
478.6B | 38687 | 2.43 GB | 2026-10-07 04:18 |
475B-tokens_step38450 |
475.4B | 38450 | 2.43 GB | 2026-10-07 03:46 |
471B-tokens_step38181 |
471.9B | 38181 | 2.43 GB | 2026-10-07 03:16 |
468B-tokens_step37937 |
468.7B | 37937 | 2.43 GB | 2026-10-07 02:48 |
465B-tokens_step37674 |
465.2B | 37674 | 2.43 GB | 2026-10-07 02:20 |
462B-tokens_step37429 |
462.1B | 37429 | 2.43 GB | 2026-10-07 01:45 |
458B-tokens_step37164 |
458.5B | 37164 | 2.43 GB | 2026-10-07 01:17 |
455B-tokens_step36914 |
455.4B | 36914 | 2.43 GB | 2026-10-07 00:45 |
451B-tokens_step36659 |
451.9B | 36659 | 2.43 GB | 2026-10-07 00:19 |
448B-tokens_step36416 |
448.7B | 36416 | 2.43 GB | 2026-10-06 23:48 |
445B-tokens_step36150 |
445.2B | 36150 | 2.43 GB | 2026-10-06 23:17 |
442B-tokens_step35905 |
442.0B | 35905 | 2.43 GB | 2026-10-06 22:45 |
438B-tokens_step35646 |
438.5B | 35646 | 2.43 GB | 2026-10-06 22:17 |
435B-tokens_step35402 |
435.3B | 35402 | 2.43 GB | 2026-10-06 21:46 |
431B-tokens_step35140 |
431.9B | 35140 | 2.43 GB | 2026-10-06 21:18 |
428B-tokens_step34890 |
428.7B | 34890 | 2.43 GB | 2026-10-06 20:46 |
425B-tokens_step34634 |
425.2B | 34634 | 2.43 GB | 2026-10-06 20:21 |
422B-tokens_step34390 |
422.0B | 34390 | 2.43 GB | 2026-10-06 19:46 |
418B-tokens_step34134 |
418.5B | 34134 | 2.43 GB | 2026-10-06 19:17 |
415B-tokens_step33886 |
415.2B | 33886 | 2.43 GB | 2026-10-06 18:45 |
411B-tokens_step33618 |
411.8B | 33618 | 2.43 GB | 2026-10-06 18:20 |
408B-tokens_step33381 |
408.7B | 33381 | 2.43 GB | 2026-10-06 17:47 |
405B-tokens_step33116 |
405.1B | 33116 | 2.43 GB | 2026-10-06 17:17 |
402B-tokens_step32861 |
402.0B | 32861 | 2.43 GB | 2026-10-06 16:45 |
398B-tokens_step32598 |
398.6B | 32598 | 2.43 GB | 2026-10-06 16:16 |
395B-tokens_step32353 |
395.3B | 32353 | 2.43 GB | 2026-10-06 15:44 |
391B-tokens_step32098 |
391.8B | 32098 | 2.43 GB | 2026-10-06 15:18 |
388B-tokens_step31854 |
388.6B | 31854 | 2.43 GB | 2026-10-06 14:45 |
385B-tokens_step31191 |
385.1B | 31191 | 2.43 GB | 2026-10-06 14:18 |
381B-tokens_step30942 |
381.8B | 30942 | 2.43 GB | 2026-10-06 13:45 |
378B-tokens_step30696 |
378.2B | 30696 | 2.43 GB | 2026-10-06 13:18 |
374B-tokens_step30455 |
374.8B | 30455 | 2.43 GB | 2026-10-06 12:44 |
371B-tokens_step30200 |
371.2B | 30200 | 2.43 GB | 2026-10-06 12:14 |
367B-tokens_step29959 |
367.8B | 29959 | 2.43 GB | 2026-10-06 11:45 |
364B-tokens_step29699 |
364.3B | 29699 | 2.43 GB | 2026-10-06 11:17 |
360B-tokens_step29461 |
360.9B | 29461 | 2.43 GB | 2026-10-06 10:47 |
357B-tokens_step29209 |
357.3B | 29209 | 2.43 GB | 2026-10-06 10:17 |
353B-tokens_step28971 |
354.0B | 28971 | 2.43 GB | 2026-10-06 09:46 |
350B-tokens_step28721 |
350.4B | 28721 | 2.43 GB | 2026-10-06 09:19 |
347B-tokens_step28477 |
347.0B | 28477 | 2.43 GB | 2026-10-06 08:46 |
343B-tokens_step28219 |
343.4B | 28219 | 2.43 GB | 2026-10-06 08:17 |
340B-tokens_step27983 |
340.1B | 27983 | 2.43 GB | 2026-10-06 07:47 |
336B-tokens_step27720 |
336.4B | 27720 | 2.43 GB | 2026-10-06 07:18 |
333B-tokens_step27484 |
333.0B | 27484 | 2.43 GB | 2026-10-06 06:46 |
329B-tokens_step27229 |
329.5B | 27229 | 2.43 GB | 2026-10-06 06:20 |
326B-tokens_step26986 |
326.1B | 26986 | 2.43 GB | 2026-10-06 05:43 |
322B-tokens_step26728 |
322.5B | 26728 | 2.43 GB | 2026-10-06 05:17 |
319B-tokens_step26493 |
319.2B | 26493 | 2.43 GB | 2026-10-06 04:44 |
315B-tokens_step26240 |
315.6B | 26240 | 2.43 GB | 2026-10-06 04:17 |
312B-tokens_step25999 |
312.2B | 25999 | 2.43 GB | 2026-10-06 03:44 |
308B-tokens_step25748 |
308.6B | 25748 | 2.43 GB | 2026-10-06 03:18 |
305B-tokens_step25502 |
305.2B | 25502 | 2.43 GB | 2026-10-06 02:45 |
301B-tokens_step25253 |
301.7B | 25253 | 2.43 GB | 2026-10-06 02:19 |
298B-tokens_step25009 |
298.3B | 25009 | 2.43 GB | 2026-10-06 01:43 |
294B-tokens_step24762 |
294.7B | 24762 | 2.43 GB | 2026-10-06 01:18 |
291B-tokens_step24519 |
291.3B | 24519 | 2.43 GB | 2026-10-06 00:46 |
287B-tokens_step24258 |
287.7B | 24258 | 2.43 GB | 2026-10-06 00:15 |
284B-tokens_step24014 |
284.4B | 24014 | 2.43 GB | 2026-10-05 23:44 |
280B-tokens_step23768 |
280.7B | 23768 | 2.43 GB | 2026-10-05 23:15 |
277B-tokens_step23530 |
277.4B | 23530 | 2.43 GB | 2026-10-05 22:46 |
273B-tokens_step23276 |
273.8B | 23276 | 2.44 GB | 2026-10-05 22:16 |
270B-tokens_step23035 |
270.4B | 23035 | 2.44 GB | 2026-10-05 21:44 |
266B-tokens_step22784 |
266.9B | 22784 | 2.44 GB | 2026-10-05 21:19 |
263B-tokens_step22539 |
263.5B | 22539 | 2.44 GB | 2026-10-05 20:47 |
259B-tokens_step22282 |
259.9B | 22282 | 2.44 GB | 2026-10-05 20:17 |
256B-tokens_step22052 |
256.6B | 22052 | 2.44 GB | 2026-10-05 19:47 |
252B-tokens_step21791 |
252.9B | 21791 | 2.44 GB | 2026-10-05 19:17 |
249B-tokens_step21550 |
249.6B | 21550 | 2.44 GB | 2026-10-05 18:45 |
245B-tokens_step21250 |
245.3B | 21250 | 2.45 GB | 2026-10-05 18:09 |
242B-tokens_step21067 |
242.4B | 21067 | 2.45 GB | 2026-10-05 17:45 |
238B-tokens_step20776 |
238.0B | 20776 | 2.45 GB | 2026-10-05 17:14 |
235B-tokens_step20575 |
235.2B | 20575 | 2.46 GB | 2026-10-05 16:46 |
230B-tokens_step20280 |
230.7B | 20280 | 2.46 GB | 2026-10-05 16:12 |
227B-tokens_step20081 |
227.9B | 20081 | 2.46 GB | 2026-10-05 15:45 |
223B-tokens_step19786 |
223.4B | 19786 | 2.45 GB | 2026-10-05 15:11 |
220B-tokens_step19602 |
220.6B | 19602 | 2.45 GB | 2026-10-05 14:47 |
216B-tokens_step19301 |
216.2B | 19301 | 2.45 GB | 2026-10-05 14:11 |
213B-tokens_step19100 |
213.3B | 19100 | 2.45 GB | 2026-10-05 13:45 |
208B-tokens_step18807 |
208.9B | 18807 | 2.45 GB | 2026-10-05 13:15 |
206B-tokens_step18619 |
206.0B | 18619 | 2.45 GB | 2026-10-05 12:45 |
201B-tokens_step18313 |
201.6B | 18313 | 2.45 GB | 2026-10-05 12:12 |
198B-tokens_step18118 |
198.7B | 18118 | 2.45 GB | 2026-10-05 11:44 |
194B-tokens_step17828 |
194.3B | 17828 | 2.45 GB | 2026-10-05 11:11 |
191B-tokens_step17631 |
191.5B | 17631 | 2.45 GB | 2026-10-05 10:44 |
187B-tokens_step17334 |
187.0B | 17334 | 2.45 GB | 2026-10-05 10:12 |
184B-tokens_step17145 |
184.2B | 17145 | 2.45 GB | 2026-10-05 09:45 |
179B-tokens_step16851 |
179.8B | 16851 | 2.45 GB | 2026-10-05 09:11 |
176B-tokens_step16645 |
176.9B | 16645 | 2.45 GB | 2026-10-05 08:45 |
172B-tokens_step16359 |
172.6B | 16359 | 2.45 GB | 2026-10-05 08:12 |
169B-tokens_step16164 |
169.7B | 16164 | 2.45 GB | 2026-10-05 07:45 |
165B-tokens_step15861 |
165.2B | 15861 | 2.45 GB | 2026-10-05 07:10 |
162B-tokens_step15667 |
162.4B | 15667 | 2.46 GB | 2026-10-05 06:45 |
157B-tokens_step15384 |
157.9B | 15384 | 2.46 GB | 2026-10-05 06:13 |
155B-tokens_step15179 |
155.1B | 15179 | 2.46 GB | 2026-10-05 05:46 |
150B-tokens_step14880 |
150.6B | 14880 | 2.46 GB | 2026-10-05 05:12 |
147B-tokens_step14693 |
147.8B | 14693 | 2.46 GB | 2026-10-05 04:45 |
143B-tokens_step14118 |
143.4B | 14118 | 2.47 GB | 2026-10-05 04:13 |
140B-tokens_step13940 |
140.5B | 13940 | 2.47 GB | 2026-10-05 03:44 |
135B-tokens_step13714 |
135.8B | 13714 | 2.47 GB | 2026-10-05 03:14 |
132B-tokens_step13524 |
132.5B | 13524 | 2.47 GB | 2026-10-05 02:48 |
127B-tokens_step13242 |
127.7B | 13242 | 2.48 GB | 2026-10-05 02:12 |
123B-tokens_step13016 |
123.6B | 13016 | 2.48 GB | 2026-10-05 01:43 |
119B-tokens_step12790 |
119.5B | 12790 | 2.48 GB | 2026-10-05 01:13 |
115B-tokens_step12532 |
115.4B | 12532 | 2.48 GB | 2026-10-05 00:38 |
111B-tokens_step12302 |
111.4B | 12302 | 2.48 GB | 2026-10-05 00:10 |
107B-tokens_step12081 |
107.2B | 12081 | 2.48 GB | 2026-10-04 23:41 |
103B-tokens_step11842 |
103.2B | 11842 | 2.48 GB | 2026-10-04 23:14 |
99B-tokens_step11533 |
99.2B | 11533 | 2.48 GB | 2026-10-04 22:39 |
95B-tokens_step11160 |
96.0B | 11160 | 2.47 GB | 2026-10-04 22:12 |
92B-tokens_step10766 |
92.6B | 10766 | 2.47 GB | 2026-10-04 21:43 |
89B-tokens_step10388 |
89.3B | 10388 | 2.47 GB | 2026-10-04 21:12 |
85B-tokens_step9986 |
86.0B | 9986 | 2.47 GB | 2026-10-04 20:41 |
82B-tokens_step9591 |
82.7B | 9591 | 2.47 GB | 2026-10-04 20:09 |
79B-tokens_step9195 |
79.3B | 9195 | 2.47 GB | 2026-10-04 19:39 |
76B-tokens_step8835 |
76.1B | 8835 | 2.48 GB | 2026-10-04 19:13 |
72B-tokens_step8443 |
72.8B | 8443 | 2.48 GB | 2026-10-04 18:40 |
69B-tokens_step8039 |
69.5B | 8039 | 2.48 GB | 2026-10-04 18:11 |
66B-tokens_step7652 |
66.1B | 7652 | 2.48 GB | 2026-10-04 17:42 |
62B-tokens_step7275 |
62.9B | 7275 | 2.48 GB | 2026-10-04 17:10 |
59B-tokens_step6711 |
59.6B | 6711 | 2.48 GB | 2026-10-04 16:40 |
56B-tokens_step6444 |
56.3B | 6444 | 2.48 GB | 2026-10-04 16:09 |
53B-tokens_step6096 |
53.0B | 6096 | 2.48 GB | 2026-10-04 15:40 |
49B-tokens_step5723 |
49.8B | 5723 | 2.48 GB | 2026-10-04 15:12 |
46B-tokens_step5320 |
46.4B | 5320 | 2.48 GB | 2026-10-04 14:40 |
43B-tokens_step4946 |
43.1B | 4946 | 2.48 GB | 2026-10-04 14:10 |
39B-tokens_step4549 |
39.8B | 4549 | 2.48 GB | 2026-10-04 13:38 |
36B-tokens_step4166 |
36.5B | 4166 | 2.48 GB | 2026-10-04 13:15 |
33B-tokens_step3768 |
33.2B | 3768 | 2.48 GB | 2026-10-04 12:40 |
29B-tokens_step3391 |
29.8B | 3391 | 2.49 GB | 2026-10-04 12:16 |
26B-tokens_step3002 |
26.5B | 3002 | 2.49 GB | 2026-10-04 11:43 |
23B-tokens_step2604 |
23.1B | 2604 | 2.50 GB | 2026-10-04 11:15 |
19B-tokens_step2198 |
19.9B | 2198 | 2.50 GB | 2026-10-04 10:40 |
16B-tokens_step1817 |
16.7B | 1817 | 2.50 GB | 2026-10-04 10:11 |
13B-tokens_step1419 |
13.3B | 1419 | 2.50 GB | 2026-10-04 09:42 |
6B-tokens_step635 |
6.6B | 635 | 2.50 GB | 2026-10-04 08:41 |
- This table lists the exports only. Final benchmarks of the final checkpoint, in percent, from the same lm-eval harness used for every model in the comparison: 35.4 on MMLU (cloze, 5-shot), 46.5 on ARC-Challenge (5-shot), 59.8 on HellaSwag (0-shot), and, on the full generative test splits, 14.3 on GSM8K (5-shot), 15.4 on MBPP (3-shot) and 12.2 on HumanEval (pass@1). No statistically significant aggregate difference was detected among the three late exports (paired 95% confidence intervals). The tech report has the full comparison, including other public base models.
Final checkpoint
exports/900B-tokens_step73973: 900.8B training tokens, step 73973.
hf download chutesai/parallax-8b-lambda --include "exports/900B-tokens_step73973/*" --local-dir parallax-8b-lambda
cd parallax-8b-lambda/exports/900B-tokens_step73973
tar -xf packed_experts.tar
sha256sum -c --quiet SHA256SUMS # every file of the original export, byte for byte
Files and format
The files are in Parallax's native compact export format (inference only, no optimizer state). They are byte-identical to the export the training system produced:
| File | Contents |
|---|---|
manifest.json |
export manifest: tensor inventory, per-file sha256 digests, token clock |
model_config.json |
model configuration |
coverage.json, layouts.json |
tensor coverage and expert frame layouts |
indexer.bundle |
sparse-attention indexer weights |
relay_pack/ |
trunk (non-expert) weights in bf16, with their own manifest |
packed_experts.tar |
the 4096 routed experts (packed_experts/*.t24p, packed ternary codes and scales) |
SHA256SUMS |
sha256 of every file of the original export |
export_info.json |
step, tokens, time, sizes and digests of this export |
The only change from the original export is packaging. The 4096 expert files are stored in one
uncompressed tar to keep the repository's file count manageable. Extract it and check SHA256SUMS
as shown above. Every upload was checked against the training system's own digests before and
after it was published.
Logit-scale bound. The model's output logit scale is bounded: the forward pass uses
exp(min(logit_scale_log, log 3.0)). The stored trunk tensor is the raw training parameter (it can sit
slightly above the bound, e.g. from bf16 rounding), and each export records the bound in manifest.json
(logit_scale: max, raw, effective) and in coverage.json (inference_policy.logit_scale_max).
A loader must apply the recorded bound; the files themselves are left byte-identical.
Running it
This is a base model only. It is not chat- or instruction-tuned, so it does plain text completion: give it the start of a text and it continues it. It will not follow instructions or hold a conversation.
Standard transformers cannot load this format. Use the parallax-lambda branch of our llama.cpp fork,
https://github.com/chutesai/llama.cpp. It adds a
dedicated runtime, llama-parallax, for CPU (x86-64, ARM64) and Apple GPUs (Metal). Lambda's recurrent
mixers (EDA) are implemented only on this branch. The fork's master branch runs the earlier kappa
checkpoints and cannot run lambda.
A vLLM serving path for GPU servers is in preparation.
Ready-made GGUF
parallax-8b-lambda-878B-tokens.gguf at the repository root (2,133,006,592 bytes, 2.13 GB, sha256
ad38864d341fe0ab70fab02d91c3320e6d36424171e20076ebba022737091a98, also in the .sha256 file next to it) was
converted from exports/878B-tokens_step72097 (878.3B training tokens, step 72097).
It keeps the export's bf16 trunk and stores the routed experts' exact ternary codes and per-row scales;
no weight is requantized. Each GGUF is a snapshot of one export and is not updated as new exports are added.
To use a later export, convert it yourself (see tools/parallax/README.md on the branch).
The earlier snapshots remain at the repository root:
| File | Export | sha256 |
|---|---|---|
parallax-8b-lambda-860B-tokens.gguf |
exports/860B-tokens_step70559 |
a65ecadcb38cb8ed34bd73ecc8fedf564393609ac2bb87f1f9f48b6149a70ba7 |
parallax-8b-lambda-790B-tokens.gguf |
exports/790B-tokens_step64456 |
489cd1d23f6c36ac394dc558fc266b26b14f5fb29b68ef9cbea0bb93b0354ef4 |
parallax-8b-lambda-725B-tokens.gguf |
exports/725B-tokens_step58814 |
eca13c0bf0bb4a5fbdf33b1616c7f33823ee84eb40fe7edb8852c566c7f6275f |
parallax-8b-lambda-584B-tokens.gguf |
exports/584B-tokens_step47029 |
0026f760b29d4098da94d3ec88e6239a5216cc09df6db046d5474ab70d54f200 |
parallax-8b-lambda-525B-tokens.gguf |
exports/525B-tokens_step42242 |
7a20cb014a3b125e14979871e9adb6bc57a9b194f46d6638325ee9daf8cb5aec |
parallax-8b-lambda-431B-tokens.gguf |
exports/431B-tokens_step35140 |
687296380f21e2c133b6d00668e09a70e1fc5488186b27a4e4ca20cdf4ddd694 |
parallax-8b-lambda-381B-tokens.gguf |
exports/381B-tokens_step30942 |
86312194ffc72b7b8d87f37389fe0fdfe7c510e2f6066f0c424cc0eaf592aedf |
hf download chutesai/parallax-8b-lambda parallax-8b-lambda-878B-tokens.gguf --local-dir .
hf download NousResearch/Meta-Llama-3-8B tokenizer.json --local-dir . # Llama 3 tokenizer
Build and run
Requires CMake, a C++17 compiler, and Python 3 with tokenizers for the text front end.
git clone -b parallax-lambda https://github.com/chutesai/llama.cpp && cd llama.cpp
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON # -DGGML_METAL=OFF for CPU only
cmake --build build --target llama-parallax -j 8
build/bin/llama-parallax --backend metal --self-test # or --backend cpu
pip install tokenizers
# Apple GPU
python tools/parallax/run.py --binary build/bin/llama-parallax \
--model ../parallax-8b-lambda-878B-tokens.gguf --tokenizer ../tokenizer.json \
--backend metal --experts lut9 --prompt 'The capital of France is' --predict 64
# CPU (x86-64 or ARM64)
python tools/parallax/run.py --binary build/bin/llama-parallax \
--model ../parallax-8b-lambda-878B-tokens.gguf --tokenizer ../tokenizer.json \
--backend cpu --experts lut8 --threads 8 --prompt 'The capital of France is' --predict 64
Generation is greedy by default. build/bin/llama-parallax --help lists sampling, context size (-c, default 4096),
prompt scoring and benchmark options.
Speed and memory
Measured on an Apple M5 Max on AC power (batch 1, 128 generated tokens) with the bf16-trunk GGUF of the earlier
270B-tokens_step23035 export, which has the same architecture, using the same runtime code; speed does not
depend on the weight values. Prompt times used --prefill-batch 128:
| Backend | Experts | Tokens/s, short context | Tokens/s after 4,096 tokens | 2,048-token prompt |
|---|---|---|---|---|
| Metal | lut9 |
192 | 172 | 2.2 s |
| CPU, 12 threads | lut8 |
157 | 112 | 4.5 s |
| CPU, 12 threads | lut (exact) |
90 | 74 |
Peak resident memory was about 2.8 GB on Metal and 3.4 GB on CPU at the default 4,096-token context. Other Apple and ARM hardware will differ. ARM64 CPUs use NEON kernels. x86-64 CPUs use portable C++ that gives the same results but is not tuned for speed: on a shared, loaded 64-core AMD EPYC 9534 server it generated about 22 tokens/s with 16 threads.
Fidelity
The 878B GGUF was checked on CPU with lut against the training code's own model loaded from the same
export, with FP32 activations in both, on the same 2,555 held-out positions from public text and code
passages as before. Overall mean KL was 4.1e-7 nats/token; the top-1 next token agreed at 2,554 of 2,555 positions,
and three 48-token greedy continuations matched exactly.
The larger differences are in two passages (segments 3 and 4, maximum logit error 0.347 and 1.085; 0.00067, 0.00075 and 0.00080 in segments 0, 1 and 2). In each affected passage they begin at a token where the router's 12th- and 13th-ranked expert-selection scores were near-tied, so the two implementations picked different experts. In segment 3 the first flip occurs in layer 7 at position 465; in segment 4 it occurs in layer 17 at position 187. The differences carry forward to later positions in both passages. Top-1 predictions agree throughout segment 3; the single top-1 disagreement is at position 319 in segment 4, after its first flip. In segment 3, the score margins were 1.4e-6 in the training code and 9.5e-7 in the runtime near -4.70 (FP32 spacing 4.8e-7, respectively 3 and 2 units of that spacing); in segment 4, they were 1.9e-6 in both near -3.92 (FP32 spacing 2.4e-7, 8 units of that spacing in both). Forcing the runtime's selection at just the first flip token/layer in the training code reduced mean KL to 1.9e-9 and 2.2e-9 in those passages, with top-1 agreement at every position. Conservatively excluding every position from those tokens onward in those passages (2,185 positions remain), mean KL was 2.1e-9 nats/token with top-1 agreement at every position.
The training code itself runs in bf16; against that bf16 forward the top-1 agreement was 94.9% (mean
KL 0.0027, measured on the earlier 270B-tokens_step23035 export), which reflects the bf16 rounding
of the reference, not an error of the runtime.
On the 878B GGUF, the largest per-layer relative RMS difference between Metal and CPU was about 2.2e-5
at the final position of a 512-token text, and the two gave identical 48-token greedy continuations on three prompts. lut8 (CPU)
quantizes expert lookup tables to 8 bits and is slightly approximate; lut and lut9 are exact.
Limitations of the runtime
- One sequence per process, greedy or simple sampled generation. No chat template, batching or server.
- Not integrated into
llama-cli,llama-serveror the standard llama.cpp model loader. - Context is limited by
-c. Inference uses a 2,048-token window for all sliding-window layers and dense global attention, the settings the run is evaluated with. Lambda was trained on 4,096-token sequences; longer inputs run but were not qualified.
Tech report
The Parallax tech report: https://parallax.chutes.ai/tech-report.pdf. It is AI-generated from the project's measurements and logs, and it is a living document that changes as the run progresses. It describes lambda, including the end of training and the final benchmark comparison.
Limitations
This is an early base model. It can produce incorrect, biased or nonsensical output. It has no alignment or safety tuning, and it is small and far from converged.
License
MIT.
Table updated 2026-10-10 11:19 UTC.
- Downloads last month
- 53
We're not able to determine the quantization variants.