Parallax 8B lambda: training checkpoints

These are research checkpoints from lambda, a training run of Parallax. Training was stopped at about 900B tokens, before the planned end of the run. The designated final checkpoint is step 73973 (900.8B tokens), marked in the table below. Exports saved later than the final checkpoint were saved after training stopped and are not the final checkpoint. Parallax is a mixture-of-experts language model trained on consumer GPUs spread over several countries. The hosts have no direct connections to each other and synchronize over the public internet.

  • Base model. It is pretrained on a mixture of web, document, math, code, knowledge-focused and book text. It is not instruction-tuned, chat-tuned or safety-tuned, and it will continue text rather than follow instructions.
  • Training snapshots. Every export, the final one included, is a snapshot taken before the planned ~973B tokens. Quality between exports is not monotonic.
  • Research artifact. It is published so the training can be followed and inspected. It is not intended for production use.

Live training dashboard: parallax.chutes.ai.

Model

Parameters ~7.77B total, ~1.20B active per token
Layers 64: 32 mixers (18 EDA recurrent, 8 sliding-window attention, 6 global attention) interleaved with 32 mixture-of-experts layers
Recurrent mixers EDA recurrence (a gated delta rule) with per-key-channel decay bounded below by e^-5 per token; erase rank 16; output normalization
Sliding-window attention First and last mixers use 2048 tokens; the six interior layers share a per-step draw of 128 or 2048 tokens with equal probability throughout training; evaluation and inference use 2048
Global attention sparse-capable global attention (MSA) with 256-dimensional latent KV, 16 query heads, 4 KV heads and RoPE; dense causal attention throughout this run, with sparse top-k selection inactive
Experts 4096 routed experts (128 per MoE layer), 12 routed + 1 shared expert per token
Expert weights ternary (-1, 0, +1) with per-row scales; at most two nonzero pairs in every group of eight
Width Model width 1152; expert intermediate width 2304, latent width 384; ReLUยฒ expert activation
Tokenizer Llama 3 tokenizer (vocabulary padded to 128,384); tied input/output embeddings
Training context 4096 tokens
Document masking In training, attention and recurrent state are reset at document boundaries (up to 64 documents per 4096-token sequence; further boundaries are merged into the 64th document)
Output logit scale learnable, capped at 3.0

Training

  • Data: seven stationary sources totalling about 1.09T tokens: two general pretraining mixtures of web, document, math and code text, an additional web set, two math sets, a knowledge-focused set and a books set. At about 899.6B tokens the data switched to two further sources while the learning-rate anneal that began at 877.28B continued. Training stopped at about 900B tokens, shortly after this switch. The planned training total was about 972.9B tokens.
  • Optimization: the dense trunk is synchronized with a decoupled DiLoCo scheme over the internet. Routed experts are trained through low-rank adapters on the GPUs that use them, and the adapter updates are folded into full-precision masters that publish new ternary versions. The tech report describes the details.
  • Learning rate: linear warm-up from 0 to 6e-4 over the first 19.66B tokens, then 6e-4 until 400B tokens, 3e-4 until 700B and 1.5e-4 until 877.28B. A live anneal started at 877.28B tokens on 2026-10-09, with multiplier m = 1 - 0.9 sqrt((T - 877.28B) / (972.90B - 877.28B)) and learning rate m times 1.5e-4. When training stopped, the multiplier had reached 0.5489 (learning rate 8.234e-5). The anneal was partial.
  • Batch: 10 gradient-accumulation micro-batches per step at launch (about 8.5M tokens per fleet step with 26 machines), switched to 20 at about 100B tokens on 2026-10-04 at 22:08 UTC. This gives about 17M tokens per step with all 26 machines; tokens per step vary with fleet size.
  • Fleet: 26 machines and 208 GPUs at launch. Seven machines were lost over 5-8 October, with roles reassigned automatically while training continued. At the end of the run, 19 machines were assigned, totalling 152 GPUs.
  • Document masking: In training, attention and recurrent state are reset at document boundaries (up to 64 documents per 4096-token sequence; further boundaries are merged into the 64th document).
  • Output logits: the learnable output logit scale is capped at 3.0 in training and inference (see Files and format).

Exports

One folder per export under exports/, named by training tokens (rounded down to whole billions) and fleet step. Exports were added about every 30 minutes while training ran. Training has stopped; the final checkpoint is marked in the table. Published exports are never modified or removed.

Export Tokens Step Size Exported (UTC)
901B-tokens_step73973_fc88abba 901.5B 73973 2.43 GB 2026-10-10 01:44
901B-tokens_step73973_6a749265 901.5B 73973 2.43 GB 2026-10-10 02:45
901B-tokens_step73973_b08bb757 901.5B 73973 2.43 GB 2026-10-10 02:18
901B-tokens_step73973 901.5B 73973 2.43 GB 2026-10-10 01:17
900B-tokens_step73973 final 900.8B 73973 2.43 GB 2026-10-10 00:45
898B-tokens_step73878 898.7B 73878 2.43 GB 2026-10-10 00:18
895B-tokens_step73637 896.0B 73637 2.43 GB 2026-10-09 23:49
892B-tokens_step73364 892.8B 73364 2.43 GB 2026-10-09 23:18
890B-tokens_step73127 890.0B 73127 2.43 GB 2026-10-09 22:48
886B-tokens_step72845 886.9B 72845 2.43 GB 2026-10-09 22:19
884B-tokens_step72604 884.2B 72604 2.43 GB 2026-10-09 21:46
881B-tokens_step72330 881.0B 72330 2.43 GB 2026-10-09 21:17
878B-tokens_step72097 878.3B 72097 2.43 GB 2026-10-09 20:48
875B-tokens_step71820 875.2B 71820 2.43 GB 2026-10-09 20:20
872B-tokens_step71588 872.4B 71588 2.43 GB 2026-10-09 19:48
869B-tokens_step71310 869.3B 71310 2.43 GB 2026-10-09 19:19
866B-tokens_step71076 866.5B 71076 2.43 GB 2026-10-09 18:47
863B-tokens_step70798 863.4B 70798 2.43 GB 2026-10-09 18:16
860B-tokens_step70559 860.6B 70559 2.43 GB 2026-10-09 17:46
857B-tokens_step70296 857.5B 70296 2.43 GB 2026-10-09 17:17
854B-tokens_step70065 854.7B 70065 2.43 GB 2026-10-09 16:50
851B-tokens_step69789 851.7B 69789 2.43 GB 2026-10-09 16:16
848B-tokens_step69548 848.9B 69548 2.43 GB 2026-10-09 15:46
845B-tokens_step69283 845.7B 69283 2.43 GB 2026-10-09 15:15
843B-tokens_step69045 843.0B 69045 2.43 GB 2026-10-09 14:48
839B-tokens_step68773 839.9B 68773 2.43 GB 2026-10-09 14:15
837B-tokens_step68544 837.1B 68544 2.43 GB 2026-10-09 13:48
834B-tokens_step68273 834.0B 68273 2.43 GB 2026-10-09 13:15
831B-tokens_step68038 831.3B 68038 2.43 GB 2026-10-09 12:47
828B-tokens_step67772 828.2B 67772 2.43 GB 2026-10-09 12:16
825B-tokens_step67534 825.5B 67534 2.43 GB 2026-10-09 11:49
822B-tokens_step67258 822.3B 67258 2.43 GB 2026-10-09 11:16
819B-tokens_step67027 819.6B 67027 2.43 GB 2026-10-09 10:47
816B-tokens_step66758 816.4B 66758 2.43 GB 2026-10-09 10:15
813B-tokens_step66509 813.7B 66509 2.43 GB 2026-10-09 09:47
810B-tokens_step66235 810.5B 66235 2.43 GB 2026-10-09 09:16
807B-tokens_step66001 807.8B 66001 2.43 GB 2026-10-09 08:46
804B-tokens_step65728 804.7B 65728 2.43 GB 2026-10-09 08:14
801B-tokens_step65490 801.9B 65490 2.43 GB 2026-10-09 07:46
798B-tokens_step65214 798.8B 65214 2.43 GB 2026-10-09 07:16
796B-tokens_step64968 796.0B 64968 2.43 GB 2026-10-09 06:47
792B-tokens_step64700 792.9B 64700 2.43 GB 2026-10-09 06:14
790B-tokens_step64456 790.2B 64456 2.43 GB 2026-10-09 05:46
787B-tokens_step64188 787.0B 64188 2.43 GB 2026-10-09 05:13
784B-tokens_step63948 784.3B 63948 2.43 GB 2026-10-09 04:49
781B-tokens_step63679 781.1B 63679 2.43 GB 2026-10-09 04:15
778B-tokens_step63434 778.4B 63434 2.43 GB 2026-10-09 03:45
775B-tokens_step63160 775.3B 63160 2.43 GB 2026-10-09 03:15
772B-tokens_step62921 772.6B 62921 2.43 GB 2026-10-09 02:47
769B-tokens_step62646 769.4B 62646 2.43 GB 2026-10-09 02:16
766B-tokens_step62413 766.6B 62413 2.43 GB 2026-10-09 01:46
763B-tokens_step62131 763.6B 62131 2.43 GB 2026-10-09 01:16
760B-tokens_step61889 760.7B 61889 2.43 GB 2026-10-09 00:48
757B-tokens_step61611 757.6B 61611 2.43 GB 2026-10-09 00:18
754B-tokens_step61370 754.9B 61370 2.43 GB 2026-10-08 23:45
751B-tokens_step61097 751.8B 61097 2.43 GB 2026-10-08 23:15
748B-tokens_step60864 749.0B 60864 2.43 GB 2026-10-08 22:46
745B-tokens_step60600 745.9B 60600 2.43 GB 2026-10-08 22:20
743B-tokens_step60341 743.1B 60341 2.43 GB 2026-10-08 21:45
740B-tokens_step60075 740.0B 60075 2.43 GB 2026-10-08 21:16
737B-tokens_step59833 737.2B 59833 2.43 GB 2026-10-08 20:45
734B-tokens_step59575 734.2B 59575 2.43 GB 2026-10-08 20:16
731B-tokens_step59330 731.4B 59330 2.43 GB 2026-10-08 19:46
728B-tokens_step59064 728.1B 59064 2.43 GB 2026-10-08 19:16
725B-tokens_step58814 725.2B 58814 2.43 GB 2026-10-08 18:46
721B-tokens_step58547 722.0B 58547 2.43 GB 2026-10-08 18:18
719B-tokens_step58295 719.1B 58295 2.43 GB 2026-10-08 17:48
715B-tokens_step58037 715.9B 58037 2.43 GB 2026-10-08 17:16
712B-tokens_step57790 712.9B 57790 2.43 GB 2026-10-08 16:46
709B-tokens_step57529 709.7B 57529 2.43 GB 2026-10-08 16:17
706B-tokens_step57284 706.8B 57284 2.43 GB 2026-10-08 15:47
703B-tokens_step57012 703.7B 57012 2.43 GB 2026-10-08 15:16
700B-tokens_step56772 700.7B 56772 2.43 GB 2026-10-08 14:45
697B-tokens_step56506 697.5B 56506 2.43 GB 2026-10-08 14:16
694B-tokens_step56272 694.6B 56272 2.43 GB 2026-10-08 13:46
691B-tokens_step55992 691.4B 55992 2.43 GB 2026-10-08 13:19
688B-tokens_step55745 688.4B 55745 2.43 GB 2026-10-08 12:45
685B-tokens_step55474 685.2B 55474 2.43 GB 2026-10-08 12:15
682B-tokens_step55234 682.3B 55234 2.43 GB 2026-10-08 11:45
679B-tokens_step54971 679.1B 54971 2.43 GB 2026-10-08 11:14
676B-tokens_step54731 676.1B 54731 2.43 GB 2026-10-08 10:48
672B-tokens_step54458 672.9B 54458 2.43 GB 2026-10-08 10:16
670B-tokens_step54209 670.0B 54209 2.43 GB 2026-10-08 09:46
666B-tokens_step53941 666.8B 53941 2.43 GB 2026-10-08 09:14
663B-tokens_step53702 663.9B 53702 2.43 GB 2026-10-08 08:45
660B-tokens_step53438 660.7B 53438 2.43 GB 2026-10-08 08:17
657B-tokens_step53196 657.8B 53196 2.43 GB 2026-10-08 07:48
654B-tokens_step52696 654.5B 52696 2.43 GB 2026-10-08 07:17
651B-tokens_step52456 651.5B 52456 2.43 GB 2026-10-08 06:48
648B-tokens_step52190 648.2B 52190 2.43 GB 2026-10-08 06:16
645B-tokens_step51945 645.2B 51945 2.43 GB 2026-10-08 05:50
641B-tokens_step51674 641.9B 51674 2.43 GB 2026-10-08 05:15
638B-tokens_step51438 638.9B 51438 2.43 GB 2026-10-08 04:47
635B-tokens_step51165 635.5B 51165 2.43 GB 2026-10-08 04:16
632B-tokens_step50914 632.5B 50914 2.43 GB 2026-10-08 03:47
629B-tokens_step50646 629.1B 50646 2.43 GB 2026-10-08 03:17
626B-tokens_step50402 626.1B 50402 2.43 GB 2026-10-08 02:47
622B-tokens_step50135 622.7B 50135 2.43 GB 2026-10-08 02:18
619B-tokens_step49880 619.8B 49880 2.43 GB 2026-10-08 01:47
616B-tokens_step49610 616.4B 49610 2.43 GB 2026-10-08 01:22
613B-tokens_step49361 613.3B 49361 2.43 GB 2026-10-08 00:47
610B-tokens_step49091 610.0B 49091 2.43 GB 2026-10-08 00:17
606B-tokens_step48845 607.0B 48845 2.43 GB 2026-10-07 23:47
603B-tokens_step48573 603.6B 48573 2.43 GB 2026-10-07 23:17
600B-tokens_step48324 600.7B 48324 2.43 GB 2026-10-07 22:47
597B-tokens_step48055 597.3B 48055 2.43 GB 2026-10-07 22:14
594B-tokens_step47817 594.4B 47817 2.43 GB 2026-10-07 21:46
590B-tokens_step47547 591.0B 47547 2.43 GB 2026-10-07 21:18
587B-tokens_step47301 588.0B 47301 2.43 GB 2026-10-07 20:46
584B-tokens_step47029 584.6B 47029 2.43 GB 2026-10-07 20:16
581B-tokens_step46782 581.6B 46782 2.43 GB 2026-10-07 19:46
578B-tokens_step46521 578.3B 46521 2.43 GB 2026-10-07 19:16
575B-tokens_step45978 575.3B 45978 2.43 GB 2026-10-07 18:46
571B-tokens_step45785 571.9B 45785 2.43 GB 2026-10-07 18:16
568B-tokens_step45541 568.7B 45541 2.43 GB 2026-10-07 17:47
565B-tokens_step45273 565.2B 45273 2.43 GB 2026-10-07 17:14
562B-tokens_step45031 562.1B 45031 2.43 GB 2026-10-07 16:46
558B-tokens_step44774 558.5B 44774 2.43 GB 2026-10-07 16:18
555B-tokens_step44528 555.4B 44528 2.43 GB 2026-10-07 15:46
551B-tokens_step44262 551.9B 44262 2.43 GB 2026-10-07 15:17
548B-tokens_step44023 548.7B 44023 2.43 GB 2026-10-07 14:47
545B-tokens_step43760 545.2B 43760 2.43 GB 2026-10-07 14:17
542B-tokens_step43518 542.1B 43518 2.43 GB 2026-10-07 13:46
538B-tokens_step43262 538.6B 43262 2.43 GB 2026-10-07 13:17
535B-tokens_step43016 535.4B 43016 2.43 GB 2026-10-07 12:47
531B-tokens_step42753 531.9B 42753 2.43 GB 2026-10-07 12:20
528B-tokens_step42504 528.8B 42504 2.43 GB 2026-10-07 11:46
525B-tokens_step42242 525.2B 42242 2.43 GB 2026-10-07 11:21
522B-tokens_step42001 522.2B 42001 2.43 GB 2026-10-07 10:48
518B-tokens_step41738 518.6B 41738 2.43 GB 2026-10-07 10:18
515B-tokens_step41487 515.5B 41487 2.43 GB 2026-10-07 09:45
511B-tokens_step41224 511.9B 41224 2.43 GB 2026-10-07 09:15
508B-tokens_step40992 508.8B 40992 2.43 GB 2026-10-07 08:48
505B-tokens_step40722 505.3B 40722 2.43 GB 2026-10-07 08:20
502B-tokens_step40480 502.1B 40480 2.43 GB 2026-10-07 07:47
498B-tokens_step40216 498.6B 40216 2.43 GB 2026-10-07 07:17
495B-tokens_step39965 495.4B 39965 2.43 GB 2026-10-07 06:46
491B-tokens_step39706 492.0B 39706 2.43 GB 2026-10-07 06:17
488B-tokens_step39456 488.8B 39456 2.43 GB 2026-10-07 05:45
485B-tokens_step39196 485.2B 39196 2.43 GB 2026-10-07 05:19
482B-tokens_step38951 482.1B 38951 2.43 GB 2026-10-07 04:44
478B-tokens_step38687 478.6B 38687 2.43 GB 2026-10-07 04:18
475B-tokens_step38450 475.4B 38450 2.43 GB 2026-10-07 03:46
471B-tokens_step38181 471.9B 38181 2.43 GB 2026-10-07 03:16
468B-tokens_step37937 468.7B 37937 2.43 GB 2026-10-07 02:48
465B-tokens_step37674 465.2B 37674 2.43 GB 2026-10-07 02:20
462B-tokens_step37429 462.1B 37429 2.43 GB 2026-10-07 01:45
458B-tokens_step37164 458.5B 37164 2.43 GB 2026-10-07 01:17
455B-tokens_step36914 455.4B 36914 2.43 GB 2026-10-07 00:45
451B-tokens_step36659 451.9B 36659 2.43 GB 2026-10-07 00:19
448B-tokens_step36416 448.7B 36416 2.43 GB 2026-10-06 23:48
445B-tokens_step36150 445.2B 36150 2.43 GB 2026-10-06 23:17
442B-tokens_step35905 442.0B 35905 2.43 GB 2026-10-06 22:45
438B-tokens_step35646 438.5B 35646 2.43 GB 2026-10-06 22:17
435B-tokens_step35402 435.3B 35402 2.43 GB 2026-10-06 21:46
431B-tokens_step35140 431.9B 35140 2.43 GB 2026-10-06 21:18
428B-tokens_step34890 428.7B 34890 2.43 GB 2026-10-06 20:46
425B-tokens_step34634 425.2B 34634 2.43 GB 2026-10-06 20:21
422B-tokens_step34390 422.0B 34390 2.43 GB 2026-10-06 19:46
418B-tokens_step34134 418.5B 34134 2.43 GB 2026-10-06 19:17
415B-tokens_step33886 415.2B 33886 2.43 GB 2026-10-06 18:45
411B-tokens_step33618 411.8B 33618 2.43 GB 2026-10-06 18:20
408B-tokens_step33381 408.7B 33381 2.43 GB 2026-10-06 17:47
405B-tokens_step33116 405.1B 33116 2.43 GB 2026-10-06 17:17
402B-tokens_step32861 402.0B 32861 2.43 GB 2026-10-06 16:45
398B-tokens_step32598 398.6B 32598 2.43 GB 2026-10-06 16:16
395B-tokens_step32353 395.3B 32353 2.43 GB 2026-10-06 15:44
391B-tokens_step32098 391.8B 32098 2.43 GB 2026-10-06 15:18
388B-tokens_step31854 388.6B 31854 2.43 GB 2026-10-06 14:45
385B-tokens_step31191 385.1B 31191 2.43 GB 2026-10-06 14:18
381B-tokens_step30942 381.8B 30942 2.43 GB 2026-10-06 13:45
378B-tokens_step30696 378.2B 30696 2.43 GB 2026-10-06 13:18
374B-tokens_step30455 374.8B 30455 2.43 GB 2026-10-06 12:44
371B-tokens_step30200 371.2B 30200 2.43 GB 2026-10-06 12:14
367B-tokens_step29959 367.8B 29959 2.43 GB 2026-10-06 11:45
364B-tokens_step29699 364.3B 29699 2.43 GB 2026-10-06 11:17
360B-tokens_step29461 360.9B 29461 2.43 GB 2026-10-06 10:47
357B-tokens_step29209 357.3B 29209 2.43 GB 2026-10-06 10:17
353B-tokens_step28971 354.0B 28971 2.43 GB 2026-10-06 09:46
350B-tokens_step28721 350.4B 28721 2.43 GB 2026-10-06 09:19
347B-tokens_step28477 347.0B 28477 2.43 GB 2026-10-06 08:46
343B-tokens_step28219 343.4B 28219 2.43 GB 2026-10-06 08:17
340B-tokens_step27983 340.1B 27983 2.43 GB 2026-10-06 07:47
336B-tokens_step27720 336.4B 27720 2.43 GB 2026-10-06 07:18
333B-tokens_step27484 333.0B 27484 2.43 GB 2026-10-06 06:46
329B-tokens_step27229 329.5B 27229 2.43 GB 2026-10-06 06:20
326B-tokens_step26986 326.1B 26986 2.43 GB 2026-10-06 05:43
322B-tokens_step26728 322.5B 26728 2.43 GB 2026-10-06 05:17
319B-tokens_step26493 319.2B 26493 2.43 GB 2026-10-06 04:44
315B-tokens_step26240 315.6B 26240 2.43 GB 2026-10-06 04:17
312B-tokens_step25999 312.2B 25999 2.43 GB 2026-10-06 03:44
308B-tokens_step25748 308.6B 25748 2.43 GB 2026-10-06 03:18
305B-tokens_step25502 305.2B 25502 2.43 GB 2026-10-06 02:45
301B-tokens_step25253 301.7B 25253 2.43 GB 2026-10-06 02:19
298B-tokens_step25009 298.3B 25009 2.43 GB 2026-10-06 01:43
294B-tokens_step24762 294.7B 24762 2.43 GB 2026-10-06 01:18
291B-tokens_step24519 291.3B 24519 2.43 GB 2026-10-06 00:46
287B-tokens_step24258 287.7B 24258 2.43 GB 2026-10-06 00:15
284B-tokens_step24014 284.4B 24014 2.43 GB 2026-10-05 23:44
280B-tokens_step23768 280.7B 23768 2.43 GB 2026-10-05 23:15
277B-tokens_step23530 277.4B 23530 2.43 GB 2026-10-05 22:46
273B-tokens_step23276 273.8B 23276 2.44 GB 2026-10-05 22:16
270B-tokens_step23035 270.4B 23035 2.44 GB 2026-10-05 21:44
266B-tokens_step22784 266.9B 22784 2.44 GB 2026-10-05 21:19
263B-tokens_step22539 263.5B 22539 2.44 GB 2026-10-05 20:47
259B-tokens_step22282 259.9B 22282 2.44 GB 2026-10-05 20:17
256B-tokens_step22052 256.6B 22052 2.44 GB 2026-10-05 19:47
252B-tokens_step21791 252.9B 21791 2.44 GB 2026-10-05 19:17
249B-tokens_step21550 249.6B 21550 2.44 GB 2026-10-05 18:45
245B-tokens_step21250 245.3B 21250 2.45 GB 2026-10-05 18:09
242B-tokens_step21067 242.4B 21067 2.45 GB 2026-10-05 17:45
238B-tokens_step20776 238.0B 20776 2.45 GB 2026-10-05 17:14
235B-tokens_step20575 235.2B 20575 2.46 GB 2026-10-05 16:46
230B-tokens_step20280 230.7B 20280 2.46 GB 2026-10-05 16:12
227B-tokens_step20081 227.9B 20081 2.46 GB 2026-10-05 15:45
223B-tokens_step19786 223.4B 19786 2.45 GB 2026-10-05 15:11
220B-tokens_step19602 220.6B 19602 2.45 GB 2026-10-05 14:47
216B-tokens_step19301 216.2B 19301 2.45 GB 2026-10-05 14:11
213B-tokens_step19100 213.3B 19100 2.45 GB 2026-10-05 13:45
208B-tokens_step18807 208.9B 18807 2.45 GB 2026-10-05 13:15
206B-tokens_step18619 206.0B 18619 2.45 GB 2026-10-05 12:45
201B-tokens_step18313 201.6B 18313 2.45 GB 2026-10-05 12:12
198B-tokens_step18118 198.7B 18118 2.45 GB 2026-10-05 11:44
194B-tokens_step17828 194.3B 17828 2.45 GB 2026-10-05 11:11
191B-tokens_step17631 191.5B 17631 2.45 GB 2026-10-05 10:44
187B-tokens_step17334 187.0B 17334 2.45 GB 2026-10-05 10:12
184B-tokens_step17145 184.2B 17145 2.45 GB 2026-10-05 09:45
179B-tokens_step16851 179.8B 16851 2.45 GB 2026-10-05 09:11
176B-tokens_step16645 176.9B 16645 2.45 GB 2026-10-05 08:45
172B-tokens_step16359 172.6B 16359 2.45 GB 2026-10-05 08:12
169B-tokens_step16164 169.7B 16164 2.45 GB 2026-10-05 07:45
165B-tokens_step15861 165.2B 15861 2.45 GB 2026-10-05 07:10
162B-tokens_step15667 162.4B 15667 2.46 GB 2026-10-05 06:45
157B-tokens_step15384 157.9B 15384 2.46 GB 2026-10-05 06:13
155B-tokens_step15179 155.1B 15179 2.46 GB 2026-10-05 05:46
150B-tokens_step14880 150.6B 14880 2.46 GB 2026-10-05 05:12
147B-tokens_step14693 147.8B 14693 2.46 GB 2026-10-05 04:45
143B-tokens_step14118 143.4B 14118 2.47 GB 2026-10-05 04:13
140B-tokens_step13940 140.5B 13940 2.47 GB 2026-10-05 03:44
135B-tokens_step13714 135.8B 13714 2.47 GB 2026-10-05 03:14
132B-tokens_step13524 132.5B 13524 2.47 GB 2026-10-05 02:48
127B-tokens_step13242 127.7B 13242 2.48 GB 2026-10-05 02:12
123B-tokens_step13016 123.6B 13016 2.48 GB 2026-10-05 01:43
119B-tokens_step12790 119.5B 12790 2.48 GB 2026-10-05 01:13
115B-tokens_step12532 115.4B 12532 2.48 GB 2026-10-05 00:38
111B-tokens_step12302 111.4B 12302 2.48 GB 2026-10-05 00:10
107B-tokens_step12081 107.2B 12081 2.48 GB 2026-10-04 23:41
103B-tokens_step11842 103.2B 11842 2.48 GB 2026-10-04 23:14
99B-tokens_step11533 99.2B 11533 2.48 GB 2026-10-04 22:39
95B-tokens_step11160 96.0B 11160 2.47 GB 2026-10-04 22:12
92B-tokens_step10766 92.6B 10766 2.47 GB 2026-10-04 21:43
89B-tokens_step10388 89.3B 10388 2.47 GB 2026-10-04 21:12
85B-tokens_step9986 86.0B 9986 2.47 GB 2026-10-04 20:41
82B-tokens_step9591 82.7B 9591 2.47 GB 2026-10-04 20:09
79B-tokens_step9195 79.3B 9195 2.47 GB 2026-10-04 19:39
76B-tokens_step8835 76.1B 8835 2.48 GB 2026-10-04 19:13
72B-tokens_step8443 72.8B 8443 2.48 GB 2026-10-04 18:40
69B-tokens_step8039 69.5B 8039 2.48 GB 2026-10-04 18:11
66B-tokens_step7652 66.1B 7652 2.48 GB 2026-10-04 17:42
62B-tokens_step7275 62.9B 7275 2.48 GB 2026-10-04 17:10
59B-tokens_step6711 59.6B 6711 2.48 GB 2026-10-04 16:40
56B-tokens_step6444 56.3B 6444 2.48 GB 2026-10-04 16:09
53B-tokens_step6096 53.0B 6096 2.48 GB 2026-10-04 15:40
49B-tokens_step5723 49.8B 5723 2.48 GB 2026-10-04 15:12
46B-tokens_step5320 46.4B 5320 2.48 GB 2026-10-04 14:40
43B-tokens_step4946 43.1B 4946 2.48 GB 2026-10-04 14:10
39B-tokens_step4549 39.8B 4549 2.48 GB 2026-10-04 13:38
36B-tokens_step4166 36.5B 4166 2.48 GB 2026-10-04 13:15
33B-tokens_step3768 33.2B 3768 2.48 GB 2026-10-04 12:40
29B-tokens_step3391 29.8B 3391 2.49 GB 2026-10-04 12:16
26B-tokens_step3002 26.5B 3002 2.49 GB 2026-10-04 11:43
23B-tokens_step2604 23.1B 2604 2.50 GB 2026-10-04 11:15
19B-tokens_step2198 19.9B 2198 2.50 GB 2026-10-04 10:40
16B-tokens_step1817 16.7B 1817 2.50 GB 2026-10-04 10:11
13B-tokens_step1419 13.3B 1419 2.50 GB 2026-10-04 09:42
6B-tokens_step635 6.6B 635 2.50 GB 2026-10-04 08:41
  • This table lists the exports only. Final benchmarks of the final checkpoint, in percent, from the same lm-eval harness used for every model in the comparison: 35.4 on MMLU (cloze, 5-shot), 46.5 on ARC-Challenge (5-shot), 59.8 on HellaSwag (0-shot), and, on the full generative test splits, 14.3 on GSM8K (5-shot), 15.4 on MBPP (3-shot) and 12.2 on HumanEval (pass@1). No statistically significant aggregate difference was detected among the three late exports (paired 95% confidence intervals). The tech report has the full comparison, including other public base models.

Final checkpoint

exports/900B-tokens_step73973: 900.8B training tokens, step 73973.

hf download chutesai/parallax-8b-lambda --include "exports/900B-tokens_step73973/*" --local-dir parallax-8b-lambda
cd parallax-8b-lambda/exports/900B-tokens_step73973
tar -xf packed_experts.tar
sha256sum -c --quiet SHA256SUMS   # every file of the original export, byte for byte

Files and format

The files are in Parallax's native compact export format (inference only, no optimizer state). They are byte-identical to the export the training system produced:

File Contents
manifest.json export manifest: tensor inventory, per-file sha256 digests, token clock
model_config.json model configuration
coverage.json, layouts.json tensor coverage and expert frame layouts
indexer.bundle sparse-attention indexer weights
relay_pack/ trunk (non-expert) weights in bf16, with their own manifest
packed_experts.tar the 4096 routed experts (packed_experts/*.t24p, packed ternary codes and scales)
SHA256SUMS sha256 of every file of the original export
export_info.json step, tokens, time, sizes and digests of this export

The only change from the original export is packaging. The 4096 expert files are stored in one uncompressed tar to keep the repository's file count manageable. Extract it and check SHA256SUMS as shown above. Every upload was checked against the training system's own digests before and after it was published.

Logit-scale bound. The model's output logit scale is bounded: the forward pass uses exp(min(logit_scale_log, log 3.0)). The stored trunk tensor is the raw training parameter (it can sit slightly above the bound, e.g. from bf16 rounding), and each export records the bound in manifest.json (logit_scale: max, raw, effective) and in coverage.json (inference_policy.logit_scale_max). A loader must apply the recorded bound; the files themselves are left byte-identical.

Running it

This is a base model only. It is not chat- or instruction-tuned, so it does plain text completion: give it the start of a text and it continues it. It will not follow instructions or hold a conversation.

Standard transformers cannot load this format. Use the parallax-lambda branch of our llama.cpp fork, https://github.com/chutesai/llama.cpp. It adds a dedicated runtime, llama-parallax, for CPU (x86-64, ARM64) and Apple GPUs (Metal). Lambda's recurrent mixers (EDA) are implemented only on this branch. The fork's master branch runs the earlier kappa checkpoints and cannot run lambda.

A vLLM serving path for GPU servers is in preparation.

Ready-made GGUF

parallax-8b-lambda-878B-tokens.gguf at the repository root (2,133,006,592 bytes, 2.13 GB, sha256 ad38864d341fe0ab70fab02d91c3320e6d36424171e20076ebba022737091a98, also in the .sha256 file next to it) was converted from exports/878B-tokens_step72097 (878.3B training tokens, step 72097). It keeps the export's bf16 trunk and stores the routed experts' exact ternary codes and per-row scales; no weight is requantized. Each GGUF is a snapshot of one export and is not updated as new exports are added. To use a later export, convert it yourself (see tools/parallax/README.md on the branch).

The earlier snapshots remain at the repository root:

File Export sha256
parallax-8b-lambda-860B-tokens.gguf exports/860B-tokens_step70559 a65ecadcb38cb8ed34bd73ecc8fedf564393609ac2bb87f1f9f48b6149a70ba7
parallax-8b-lambda-790B-tokens.gguf exports/790B-tokens_step64456 489cd1d23f6c36ac394dc558fc266b26b14f5fb29b68ef9cbea0bb93b0354ef4
parallax-8b-lambda-725B-tokens.gguf exports/725B-tokens_step58814 eca13c0bf0bb4a5fbdf33b1616c7f33823ee84eb40fe7edb8852c566c7f6275f
parallax-8b-lambda-584B-tokens.gguf exports/584B-tokens_step47029 0026f760b29d4098da94d3ec88e6239a5216cc09df6db046d5474ab70d54f200
parallax-8b-lambda-525B-tokens.gguf exports/525B-tokens_step42242 7a20cb014a3b125e14979871e9adb6bc57a9b194f46d6638325ee9daf8cb5aec
parallax-8b-lambda-431B-tokens.gguf exports/431B-tokens_step35140 687296380f21e2c133b6d00668e09a70e1fc5488186b27a4e4ca20cdf4ddd694
parallax-8b-lambda-381B-tokens.gguf exports/381B-tokens_step30942 86312194ffc72b7b8d87f37389fe0fdfe7c510e2f6066f0c424cc0eaf592aedf
hf download chutesai/parallax-8b-lambda parallax-8b-lambda-878B-tokens.gguf --local-dir .
hf download NousResearch/Meta-Llama-3-8B tokenizer.json --local-dir .   # Llama 3 tokenizer

Build and run

Requires CMake, a C++17 compiler, and Python 3 with tokenizers for the text front end.

git clone -b parallax-lambda https://github.com/chutesai/llama.cpp && cd llama.cpp
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON   # -DGGML_METAL=OFF for CPU only
cmake --build build --target llama-parallax -j 8
build/bin/llama-parallax --backend metal --self-test               # or --backend cpu
pip install tokenizers

# Apple GPU
python tools/parallax/run.py --binary build/bin/llama-parallax \
  --model ../parallax-8b-lambda-878B-tokens.gguf --tokenizer ../tokenizer.json \
  --backend metal --experts lut9 --prompt 'The capital of France is' --predict 64

# CPU (x86-64 or ARM64)
python tools/parallax/run.py --binary build/bin/llama-parallax \
  --model ../parallax-8b-lambda-878B-tokens.gguf --tokenizer ../tokenizer.json \
  --backend cpu --experts lut8 --threads 8 --prompt 'The capital of France is' --predict 64

Generation is greedy by default. build/bin/llama-parallax --help lists sampling, context size (-c, default 4096), prompt scoring and benchmark options.

Speed and memory

Measured on an Apple M5 Max on AC power (batch 1, 128 generated tokens) with the bf16-trunk GGUF of the earlier 270B-tokens_step23035 export, which has the same architecture, using the same runtime code; speed does not depend on the weight values. Prompt times used --prefill-batch 128:

Backend Experts Tokens/s, short context Tokens/s after 4,096 tokens 2,048-token prompt
Metal lut9 192 172 2.2 s
CPU, 12 threads lut8 157 112 4.5 s
CPU, 12 threads lut (exact) 90 74

Peak resident memory was about 2.8 GB on Metal and 3.4 GB on CPU at the default 4,096-token context. Other Apple and ARM hardware will differ. ARM64 CPUs use NEON kernels. x86-64 CPUs use portable C++ that gives the same results but is not tuned for speed: on a shared, loaded 64-core AMD EPYC 9534 server it generated about 22 tokens/s with 16 threads.

Fidelity

The 878B GGUF was checked on CPU with lut against the training code's own model loaded from the same export, with FP32 activations in both, on the same 2,555 held-out positions from public text and code passages as before. Overall mean KL was 4.1e-7 nats/token; the top-1 next token agreed at 2,554 of 2,555 positions, and three 48-token greedy continuations matched exactly.

The larger differences are in two passages (segments 3 and 4, maximum logit error 0.347 and 1.085; 0.00067, 0.00075 and 0.00080 in segments 0, 1 and 2). In each affected passage they begin at a token where the router's 12th- and 13th-ranked expert-selection scores were near-tied, so the two implementations picked different experts. In segment 3 the first flip occurs in layer 7 at position 465; in segment 4 it occurs in layer 17 at position 187. The differences carry forward to later positions in both passages. Top-1 predictions agree throughout segment 3; the single top-1 disagreement is at position 319 in segment 4, after its first flip. In segment 3, the score margins were 1.4e-6 in the training code and 9.5e-7 in the runtime near -4.70 (FP32 spacing 4.8e-7, respectively 3 and 2 units of that spacing); in segment 4, they were 1.9e-6 in both near -3.92 (FP32 spacing 2.4e-7, 8 units of that spacing in both). Forcing the runtime's selection at just the first flip token/layer in the training code reduced mean KL to 1.9e-9 and 2.2e-9 in those passages, with top-1 agreement at every position. Conservatively excluding every position from those tokens onward in those passages (2,185 positions remain), mean KL was 2.1e-9 nats/token with top-1 agreement at every position.

The training code itself runs in bf16; against that bf16 forward the top-1 agreement was 94.9% (mean KL 0.0027, measured on the earlier 270B-tokens_step23035 export), which reflects the bf16 rounding of the reference, not an error of the runtime.

On the 878B GGUF, the largest per-layer relative RMS difference between Metal and CPU was about 2.2e-5 at the final position of a 512-token text, and the two gave identical 48-token greedy continuations on three prompts. lut8 (CPU) quantizes expert lookup tables to 8 bits and is slightly approximate; lut and lut9 are exact.

Limitations of the runtime

  • One sequence per process, greedy or simple sampled generation. No chat template, batching or server.
  • Not integrated into llama-cli, llama-server or the standard llama.cpp model loader.
  • Context is limited by -c. Inference uses a 2,048-token window for all sliding-window layers and dense global attention, the settings the run is evaluated with. Lambda was trained on 4,096-token sequences; longer inputs run but were not qualified.

Tech report

The Parallax tech report: https://parallax.chutes.ai/tech-report.pdf. It is AI-generated from the project's measurements and logs, and it is a living document that changes as the run progresses. It describes lambda, including the end of training and the final benchmark comparison.

Limitations

This is an early base model. It can produce incorrect, biased or nonsensical output. It has no alignment or safety tuning, and it is small and far from converged.

License

MIT.

Table updated 2026-10-10 11:19 UTC.

Downloads last month
53
GGUF
Model size
2B params
Architecture
parallax
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support