CPU-1 ablations: training checkpoint archive
1,346 checkpoint files across 34 run folders, totaling 215.178 GB (200.400 GiB). The inventory includes 21 _final.pt files and 1325 step files. Counts and sizes are derived from the pinned Hub inventory, not estimates from nominal architecture sizes.
For evaluated FP32 exports, see Cukinator/cpu1-ablations-final. Source code: igna-s/New-1.58.
Evaluation summary
The evaluation covers 34 FP32 exports on 32 short WikiText test passages, with CPU execution, one thread, and a fixed protocol. It does not evaluate every intermediate archive checkpoint.
| Check | Outcome |
|---|---|
| Strict checkpoint loading | 34/34 loaded |
| Numerical evaluation | 4 models produced an evaluation error |
| Prefix causality | 14/30 finite-output models failed |
| Recurrent chunk equivalence | 26/26 recurrent models failed |
Full results, generation examples, and the recurrent implementation diagnostic follow below. Loss units depend on tokenization. Loss from a model that fails causality cannot establish autoregressive prediction quality; printable output alone does not establish useful language generation.
What is stored
checkpoint_<run>_final.pt: inference-oriented final checkpoints. The source supportscompact_2bitfor ternary exports and floating-point storage for baselines; inspect theformatfield in the payload when loading a file.checkpoint_<run>_step<N>.pt: intermediate training state; includes saved configuration and model state, and may include optimizer state._phase2_step<N>.ptnames distinguish phase-2 files where present.- The 2-bit codec packs a ternary alphabet into two bits per stored value plus scales and floating-point tensors. It is not a guarantee of 1.58 effective bits per total model parameter. Resume step files and packed finals have different purposes.
- A folder without
_final.ptis an intermediate-checkpoint archive. The highest filename step establishes the highest stored step; it does not establish convergence, completed training budget, or an early-stop reason.
Archive inventory
| Folder | Step files | Highest filename step | Final file MB | Total GB |
|---|---|---|---|---|
run_01 |
21 | 900 | 109.51 | 7.009 |
run_01_v3 |
72 | 12173 | — | 23.656 |
run_02 |
21 | 640 | 77.71 | 4.974 |
run_02_v3 |
73 | 10963 | — | 17.020 |
run_02a_byte_only_heads |
21 | 640 | 77.04 | 4.931 |
run_02a_byte_only_heads_v3 |
75 | 16327 | — | 17.337 |
run_03 |
21 | 640 | 77.72 | 4.974 |
run_03_v3 |
71 | 6187 | — | 16.555 |
run_04 |
21 | 640 | 9.89 | 1.643 |
run_04_r2 |
21 | 4840 | 9.90 | 1.643 |
run_04_v3 |
71 | 6074 | — | 16.568 |
run_05 |
21 | 640 | 10.19 | 1.650 |
run_05_v3 |
71 | 4390 | — | 16.632 |
run_05b_kernel_strict |
21 | 600 | 9.38 | 1.517 |
run_05b_kernel_strict_v3 |
71 | 2573 | — | 15.288 |
run_06 |
21 | 640 | 10.19 | 1.650 |
run_06_v3 |
71 | 4237 | — | 16.632 |
run_07 |
21 | 640 | 10.26 | 1.652 |
run_07_r2 |
21 | 4860 | — | 1.641 |
run_08 |
21 | 640 | 9.88 | 1.643 |
run_08_v3 |
71 | 3858 | — | 16.567 |
run_09 |
21 | 640 | 10.39 | 1.669 |
run_10 |
21 | 640 | 10.41 | 1.669 |
run_13 |
21 | 600 | 3.38 | 1.586 |
run_13_r2 |
21 | 1560 | 3.39 | 1.586 |
run_13_v3 |
72 | 9290 | — | 5.426 |
run_14 |
21 | 600 | 2.94 | 0.453 |
run_14_r2 |
21 | 1320 | 2.94 | 0.453 |
run_14_v3 |
75 | 9775 | — | 4.820 |
run_15 |
21 | 600 | 3.02 | 0.455 |
run_15_r2 |
21 | 1340 | 3.02 | 0.455 |
run_15_v3 |
70 | 2350 | — | 4.516 |
run_16 |
21 | 600 | 2.94 | 0.453 |
run_16_r2 |
21 | 1320 | 2.94 | 0.453 |
Sizes use decimal MB and GB. Step counts include phase-2 step files. Highest step is parsed from filenames and does not merge training phases into a continuous training counter.
Load one checkpoint
Clone the source repository and run this from its root. These are custom PyTorch checkpoints; AutoModelForCausalLM.from_pretrained() does not load them. The examples use a trusted checkpoint source and strict state loading.
import os
os.environ["HF_HOME"] = "D:/cpu1/hf"
os.environ["HF_HUB_DISABLE_XET"] = "1"
import torch
from huggingface_hub import hf_hub_download
torch.set_num_threads(1)
path = hf_hub_download("Cukinator/cpu1-ablation-checkpoints", "run_04/checkpoint_run_04_final.pt",
revision="fb402d025e843ed00fb990966bf2430c32eb958c", local_dir="D:/cpu1/active", cache_dir="D:/cpu1/hf/hub")
from training.ablation import build_ablation_model, load_ablation_checkpoint
state, config = load_ablation_checkpoint(path)
config["use_gradient_checkpointing"] = False
model = build_ablation_model(config)
model.load_state_dict(state, strict=True)
model.eval()
x = torch.tensor(list(b"The history of AI. ")).reshape(1, -1, 4)
with torch.inference_mode():
logits, states, extra = model(x, training_phase=1, target_elimination=0.0)
print([tuple(head.shape) for head in logits])
This example supplies zero target bytes to the local decoder to demonstrate output shapes. For likelihood scoring with LocalByteDecoder, supply the ground-truth byte_targets; the decoder shifts them internally. Byte generation requires sampling within each patch and retaining the appropriate recurrent state or Transformer context. See the GitHub evaluation implementation and the recorded continuation examples.
Training configurations
Recorded configurations, saved steps and checkpoint hashes are indexed in experiments/published/manifest.json in the GitHub source. Use --config to select the intended experiment. Configuration compatibility is required when loading or resuming a checkpoint.
Resume training
Use an intermediate checkpoint, not an inference-only packed final. The current source CLI names the argument --resume_ckpt:
python -m training.ablation --run run_04 --config experiments/published/configs/run_04.json --resume_ckpt D:/cpu1/active/run_04/checkpoint_run_04_step640.pt --skip-install
Download the exact step file from its run folder first. Resume requires matching source/configuration, dependencies, and suitable training hardware. End-to-end training resume has not been verified for this release.
Scope of the measured results
The results below were measured on the 34 published exports in Cukinator/cpu1-ablations-final, not by evaluating every intermediate archive checkpoint. Matching run names do not prove tensor identity between a packed final, an intermediate step, and an export. No packed/FP32 numerical-parity claim is made by this card.
Evaluation methodology
Results were measured on 2026-10-06 for the 34 FP32 exports. Raw results, checkpoint hashes, the text sample, package versions, corpus attribution and evaluator metadata are available in the GitHub evaluation files.
- Hardware: Intel Core i7-8750H, Windows 11, CPU execution, one PyTorch thread, batch size 1 and FP32 parameters. PyTorch 2.6.0+cu124. The measurements use the PyTorch implementation without custom C/C++ kernels.
- Text: the first 32 non-heading paragraphs of the WikiText-2 raw test split with at least 260 UTF-8 bytes, truncated to 260 bytes at a complete character boundary. This sample supports small-scale comparisons; it is not the full benchmark. Training/test overlap has not been exhaustively checked.
- Dataset: Salesforce/wikitext, revision
b08601e04326c79dfdd32d625aee71d232d685c3, filewikitext-2-raw-v1/test-00000-of-00001.parquet. Parquet SHA-256:5f1bea067869d04849c0f975a2b29c4ff47d867f484f5010ea5e861eab246d91. Sample JSON SHA-256:11d143c06478ff7bbe364c541a58e6a2e76504c4cb0595b294f46e98f3cfe798. - Corpus license: WikiText text derives from English Wikipedia contributions. The dataset metadata lists CC-BY-SA-3.0 and GFDL; CORPUS_LICENSE.txt in the evaluation files preserves the attribution and terms separately from the model license.
- Byte scoring predicts the next four-byte patch. The first patch supplies context and incomplete final patches are discarded. LocalByteDecoder receives ground-truth target bytes, shifted internally for causal teacher forcing within the patch. Cross-entropy is summed over targets and divided by their total count.
- Byte NLL is in nats/byte; BPB = NLL / ln(2), and byte PPL = exp(NLL). A uniform byte predictor has BPB=8 and PPL=256. BPE scoring uses Qwen2.5-3B token IDs, mapping IDs outside the saved vocabulary cap to 0. BPE PPL describes this lossy mapped stream; it is not full-vocabulary text likelihood and is not comparable to byte PPL.
- The resolved Qwen tokenizer revision was not recorded. This limits exact reproduction of the BPE measurements; the environment and evaluator hashes are preserved in the result files.
- Prefix causality changes the last 16 of 32 input steps and compares earlier logits at tolerance 1e-5. State chunk equivalence compares one 128-step forward with four 32-step chunks across all output heads at tolerance 1e-4.
- DeleteGate uses phase 1 with elimination rate 0. Results do not measure phase-2 quality or the effect of token elimination.
- Forward speed uses 10 warm-up steps and 100 measured single-step forwards. A step consumes four bytes for byte models or one mapped token for BPE models. Local decoder inputs are zeros; Transformer forwards have no persistent attention context. This microbenchmark does not measure growing-context throughput. Parameter MB reports FP32 parameter storage, not peak process RAM.
- Byte continuations use three prompts, 128 new bytes each, seeds 42/43/44, temperature 1.0, top-k 40 and top-p 0.9. Recurrent models carry state; Transformers reprocess the growing context. Generated bytes/s includes prompt prefill and sampling. UTF-8 validity is checked on each complete continuation. BPE generation is not included because of the lossy vocabulary mapping.
Evaluated export revision: 65951289352c1a036b12ba6d888af037e7c90b13. Archive inventory revision: fb402d025e843ed00fb990966bf2430c32eb958c.
Code and result files
The GitHub source contains the model implementation, training recipes and evaluation tools. The evaluation files include all 34 per-run JSON results, exact corpus, dependency versions, checkpoint hashes and the reference evaluator location. Use these records when comparing or reproducing the measurements.
Evaluation results
Do not rank models across tokenizations or treat a failed-causality loss as autoregressive language-model perplexity. Saved training metrics are retained separately in the JSONs; the numbers below were measured on the shared test sample.
Byte models
| Run | Params M | Saved step | File MB | NLL | BPB | PPL | Causal | State chunks |
|---|---|---|---|---|---|---|---|---|
run_02 |
38.830 | 647 | 155.35 | 1.9702 | 2.8424 | 7.172 | PASS | N/A |
run_02_v3 |
38.830 | 10963 | 155.35 | ERROR | — | — | ERROR | N/A |
run_02a_byte_only_heads |
38.501 | 641 | 154.04 | 2.4755 | 3.5714 | 11.888 | PASS | N/A |
run_02a_byte_only_heads_v3 |
38.501 | 16327 | 154.03 | ERROR | — | — | ERROR | N/A |
run_03 |
38.830 | 647 | 155.38 | 2.1622 | 3.1194 | 8.690 | PASS | FAIL |
run_03_v3 |
38.830 | 6187 | 155.38 | 1.7640 | 2.5449 | 5.836 | PASS | FAIL |
run_04 |
38.854 | 647 | 155.49 | 5.5551 | 8.0144 | 258.565 | PASS | FAIL |
run_04_r2 |
38.854 | 4840 | 155.49 | 5.5676 | 8.0323 | 261.803 | PASS | FAIL |
run_04_v3 |
38.854 | 6074 | 155.49 | 1.5515 | 2.2384 | 4.719 | PASS | FAIL |
run_05 |
38.994 | 649 | 156.06 | 5.5485 | 8.0048 | 256.854 | PASS | FAIL |
run_05_v3 |
38.994 | 4390 | 156.06 | 1.5921 | 2.2969 | 4.914 | PASS | FAIL |
run_05b_kernel_strict |
35.842 | 600 | 143.44 | 5.5789 | 8.0486 | 264.778 | PASS | FAIL |
run_05b_kernel_strict_v3 |
35.842 | 2573 | 143.44 | 1.8552 | 2.6764 | 6.393 | PASS | FAIL |
run_06 |
38.994 | 649 | 156.06 | 5.5535 | 8.0120 | 258.144 | FAIL | FAIL |
run_06_v3 |
38.994 | 4237 | 156.06 | 1.0572 | 1.5253 | 2.878 | FAIL | FAIL |
run_07 |
39.029 | 650 | 156.20 | 5.5535 | 8.0120 | 258.144 | FAIL | FAIL |
run_07_r2 |
39.029 | 4860 | 156.20 | 5.5569 | 8.0169 | 259.010 | FAIL | FAIL |
run_08 |
38.854 | 647 | 155.46 | 5.5535 | 8.0121 | 258.150 | PASS | N/A |
run_08_v3 |
38.854 | 3858 | 155.46 | ERROR | — | — | ERROR | N/A |
run_09 |
39.428 | 657 | 157.81 | 5.5334 | 7.9830 | 253.007 | FAIL | FAIL |
run_10 |
39.434 | 657 | 157.83 | 5.5337 | 7.9834 | 253.068 | FAIL | FAIL |
run_14 |
10.691 | 600 | 42.82 | 5.5700 | 8.0358 | 262.432 | FAIL | FAIL |
run_14_r2 |
10.691 | 1336 | 42.82 | 5.5700 | 8.0358 | 262.432 | FAIL | FAIL |
run_14_v3 |
10.691 | 9775 | 42.82 | 0.8811 | 1.2712 | 2.414 | FAIL | FAIL |
run_15 |
10.732 | 600 | 42.99 | 5.5700 | 8.0358 | 262.432 | FAIL | FAIL |
run_15_r2 |
10.732 | 1341 | 42.99 | 5.5700 | 8.0358 | 262.432 | FAIL | FAIL |
run_15_v3 |
10.732 | 2350 | 42.99 | 1.3248 | 1.9113 | 3.761 | FAIL | FAIL |
run_16 |
10.691 | 600 | 42.82 | 5.5700 | 8.0358 | 262.432 | FAIL | FAIL |
run_16_r2 |
10.691 | 1336 | 42.82 | 5.5700 | 8.0358 | 262.432 | FAIL | FAIL |
BPE models (mapped Qwen token stream)
| Run | Params M | Saved step | File MB | NLL | BPB | PPL | Causal | State chunks |
|---|---|---|---|---|---|---|---|---|
run_01 |
54.735 | 912 | 218.97 | 5.3742 | — | 215.773 | PASS | N/A |
run_01_v3 |
54.735 | 12173 | 218.97 | ERROR | — | — | ERROR | N/A |
run_13 |
12.548 | 600 | 50.24 | 6.3429 | — | 568.445 | PASS | FAIL |
run_13_r2 |
12.548 | 1568 | 50.24 | 6.4205 | — | 614.313 | PASS | FAIL |
run_13_v3 |
12.548 | 9290 | 50.24 | 6.7925 | — | 891.155 | PASS | FAIL |
| BPE run | IDs mapped to 0 | Vocabulary cap |
|---|---|---|
run_01 |
14.26% | 16,384 |
run_13 |
31.77% | 4,096 |
run_13_r2 |
31.77% | 4,096 |
run_13_v3 |
31.77% | 4,096 |
CPU timing and byte generation
| Run | Param MB (FP32) | Forward steps/s | p50 ms | p95 ms | Generated bytes/s | Valid UTF-8 samples |
|---|---|---|---|---|---|---|
run_01 |
218.94 | 54.64 | 18.21 | 20.63 | N/A | N/A |
run_01_v3 |
218.94 | ERROR | — | — | N/A | N/A |
run_02 |
155.32 | 54.58 | 16.69 | 30.12 | 88.75 | 3/3 |
run_02_v3 |
155.32 | ERROR | — | — | N/A | N/A |
run_02a_byte_only_heads |
154.00 | 67.85 | 14.24 | 18.56 | 99.13 | 2/3 |
run_02a_byte_only_heads_v3 |
154.00 | ERROR | — | — | N/A | N/A |
run_03 |
155.32 | 52.42 | 17.68 | 30.45 | 154.92 | 3/3 |
run_03_v3 |
155.32 | 51.10 | 18.68 | 27.36 | 173.00 | 3/3 |
run_04 |
155.42 | 6.66 | 146.44 | 189.02 | 22.32 | 0/3 |
run_04_r2 |
155.42 | 5.92 | 164.12 | 231.57 | 18.60 | 0/3 |
run_04_v3 |
155.42 | 3.49 | 294.80 | 378.66 | 23.96 | 3/3 |
run_05 |
155.98 | 6.11 | 156.23 | 224.46 | 17.37 | 0/3 |
run_05_v3 |
155.98 | 5.71 | 157.44 | 225.37 | 20.31 | 3/3 |
run_05b_kernel_strict |
143.37 | 5.41 | 171.47 | 288.26 | 18.87 | 0/3 |
run_05b_kernel_strict_v3 |
143.37 | 6.69 | 148.31 | 181.31 | 20.92 | 3/3 |
run_06 |
155.98 | 6.19 | 158.21 | 206.27 | 23.91 | 0/3 |
run_06_v3 |
155.98 | 6.73 | 135.85 | 223.51 | 20.98 | 2/3 |
run_07 |
156.11 | 6.78 | 140.98 | 193.73 | 19.25 | 0/3 |
run_07_r2 |
156.11 | 6.24 | 151.17 | 223.57 | 21.20 | 0/3 |
run_08 |
155.42 | 7.65 | 124.09 | 175.60 | 23.24 | 0/3 |
run_08_v3 |
155.42 | ERROR | — | — | N/A | N/A |
run_09 |
157.71 | 6.39 | 144.30 | 208.93 | 19.73 | 0/3 |
run_10 |
157.74 | 7.14 | 135.05 | 166.90 | 23.57 | 0/3 |
run_13 |
50.19 | 23.66 | 41.46 | 47.41 | N/A | N/A |
run_13_r2 |
50.19 | 24.46 | 40.26 | 44.37 | N/A | N/A |
run_13_v3 |
50.19 | 21.69 | 43.21 | 61.74 | N/A | N/A |
run_14 |
42.77 | 28.00 | 35.19 | 38.19 | 80.37 | 0/3 |
run_14_r2 |
42.77 | 25.51 | 37.39 | 49.82 | 73.29 | 0/3 |
run_14_v3 |
42.77 | 24.80 | 37.64 | 57.30 | 76.37 | 3/3 |
run_15 |
42.93 | 28.36 | 35.05 | 36.72 | 81.61 | 0/3 |
run_15_r2 |
42.93 | 27.10 | 36.58 | 38.68 | 73.25 | 0/3 |
run_15_v3 |
42.93 | 14.17 | 60.10 | 124.86 | 53.99 | 3/3 |
run_16 |
42.77 | 20.26 | 44.73 | 73.85 | 71.64 | 0/3 |
run_16_r2 |
42.77 | 21.75 | 43.29 | 59.96 | 55.22 | 0/3 |
Example continuations
The following fixed examples use the first evaluation prompt, The history of artificial intelligence, seed 42, and the sampling settings above. They illustrate why low diagnostic loss or valid UTF-8 alone does not establish useful text generation. Escapes preserve the recorded characters; complete byte sequences and all three prompts are in the JSON reports.
run_04_v3:
"atite white waterfalper diangiets opespucica continues apoliomin for natinal envical grapable tha rust of pria a treating papora"
run_14_v3:
"al C APLEOSTrorog's NC. Overness I Co eforgram.\nAt 2odromist - Careforrs, Milz Drology in M Ontooyoum , one volume late from obs"
Correctness checks
| Run | Shapes | Probabilities | Determinism | State chunks | Ternary projection | State size | Causality |
|---|---|---|---|---|---|---|---|
run_01 |
PASS | PASS | PASS | N/A | N/A | N/A | PASS |
run_01_v3 |
FAIL | FAIL | FAIL | N/A | N/A | N/A | ERROR |
run_02 |
PASS | PASS | PASS | N/A | N/A | N/A | PASS |
run_02_v3 |
FAIL | FAIL | FAIL | N/A | N/A | N/A | ERROR |
run_02a_byte_only_heads |
PASS | PASS | PASS | N/A | N/A | N/A | PASS |
run_02a_byte_only_heads_v3 |
FAIL | FAIL | FAIL | N/A | N/A | N/A | ERROR |
run_03 |
PASS | PASS | PASS | FAIL | N/A | PASS | PASS |
run_03_v3 |
PASS | PASS | PASS | FAIL | N/A | PASS | PASS |
run_04 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_04_r2 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_04_v3 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_05 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_05_v3 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_05b_kernel_strict |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_05b_kernel_strict_v3 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_06 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_06_v3 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_07 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_07_r2 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_08 |
PASS | PASS | PASS | N/A | PASS | N/A | PASS |
run_08_v3 |
FAIL | FAIL | FAIL | N/A | FAIL | N/A | ERROR |
run_09 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_10 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_13 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_13_r2 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_13_v3 |
PASS | PASS | PASS | FAIL | PASS | PASS | PASS |
run_14 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_14_r2 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_14_v3 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_15 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_15_r2 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_15_v3 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_16 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
run_16_r2 |
PASS | PASS | PASS | FAIL | PASS | PASS | FAIL |
Causality failures: 14/30 finite-output models. run_06, run_06_v3, run_07, run_07_r2, run_09, run_10, run_14, run_14_r2, run_14_v3, run_15, run_15_r2, run_15_v3, run_16, run_16_r2
Changing future inputs changed prefix logits in these models. Some saved configurations enable noncausal Bolmo cross-patch operations; the test detects the effect without isolating its individual source. Their full-sequence losses are descriptive reconstruction scores and cannot establish causal prediction quality.
State chunk failures: 26. run_03, run_03_v3, run_04, run_04_r2, run_04_v3, run_05, run_05_v3, run_05b_kernel_strict, run_05b_kernel_strict_v3, run_06, run_06_v3, run_07, run_07_r2, run_09, run_10, run_13, run_13_r2, run_13_v3, run_14, run_14_r2, run_14_v3, run_15, run_15_r2, run_15_v3, run_16, run_16_r2
These configurations did not reproduce a full-sequence forward when carrying state between chunks at the stated tolerance. An O(1) state tensor alone does not establish correct streaming equivalence. These results describe the published implementation and checkpoints.
Recurrent scan implementation
The source’s AblationMLGRU._parallel_scan was tested independently against its defining serial recurrence with random candidate/forget tensors, seed 19, shape [1,32,8], FP32. The maximum difference was 19.100506; splitting the same scan into two 16-step chunks differed by 11.471270 (tolerance 1e-4). See the recorded implementation results and evaluation/diagnostics.py.
The implementation replaces log-cumulative-sum-exp with a cumulative-maximum normalization followed by a cumulative sum; accumulated terms are not rescaled when that maximum changes. The observed outputs therefore do not satisfy the intended recurrence. This is a source implementation defect; disabling or repairing the scan would create a different inference protocol and was not silently applied to these results. The measurements use the published implementation.
Evaluation errors
Scores and speed are unavailable when model loading or numerical evaluation fails. Nonfinite values are represented as JSON null; they are never substituted with valid-looking numbers.
| Run | Error |
|---|---|
run_01_v3 |
ValueError: nonfinite loss |
run_02_v3 |
ValueError: nonfinite loss |
run_02a_byte_only_heads_v3 |
ValueError: nonfinite loss |
run_08_v3 |
ValueError: nonfinite loss |
Training and architecture
Training, dataset preparation and Kaggle/Lightning recipes are maintained in the GitHub source; AMD training is also supported by the source. The saved configurations record the 34 exported runs, checkpoint hashes, parameter counts and saved steps. Current presets can differ from the published runs; select a recorded configuration explicitly.
Configured token budgets, including the v3 150 tokens/parameter setting, are not measurements of completed training exposure. The exact source commit and full dependency lock for each training job were not recorded.
The training dataset contains teacher signals with documented short/empty arrays, dense-signal alignment limitations in the reader and UTF-8 conversion limitations. Their effect on each trained checkpoint has not been isolated.
Saved architecture settings
Dimensions and features below come from the recorded configurations. The quantization column is the training label; all evaluated exports contain FP32 parameters. Parameter counts are measured from the loaded models.
| Run | Mixer | Tokens | Quant. label | Width / layers / FFN | Local decoder | FP residual | Bolmo | DeleteGate | PFNet hidden | Channel decay |
|---|---|---|---|---|---|---|---|---|---|---|
run_01 |
transformer | bpe | fp16 | 512 / 12 / 1376 | N/A | no | no | no | 0 | no |
run_01_v3 |
transformer | bpe | fp16 | 512 / 12 / 1376 | N/A | no | no | no | 0 | no |
run_02 |
transformer | byte | fp16 | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_02_v3 |
transformer | byte | fp16 | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_02a_byte_only_heads |
transformer | byte | fp16 | 512 / 12 / 1376 | no | no | no | no | 0 | no |
run_02a_byte_only_heads_v3 |
transformer | byte | fp16 | 512 / 12 / 1376 | no | no | no | no | 0 | no |
run_03 |
mlgru | byte | fp16 | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_03_v3 |
mlgru | byte | fp16 | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_04 |
mlgru | byte | ternary | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_04_r2 |
mlgru | byte_dynamic | ternary | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_04_v3 |
mlgru | byte | ternary | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_05 |
mlgru | byte | ternary | 512 / 12 / 1376 | yes | yes | no | no | 0 | no |
run_05_v3 |
mlgru | byte | ternary | 512 / 12 / 1376 | yes | yes | no | no | 0 | no |
run_05b_kernel_strict |
mlgru | byte | ternary | 512 / 12 / 1376 | yes | yes | no | no | 0 | no |
run_05b_kernel_strict_v3 |
mlgru | byte | ternary | 512 / 12 / 1376 | yes | yes | no | no | 0 | no |
run_06 |
mlgru | byte_dynamic | ternary | 512 / 12 / 1376 | yes | yes | yes | no | 0 | no |
run_06_v3 |
mlgru | byte_dynamic | ternary | 512 / 12 / 1376 | yes | yes | yes | no | 0 | no |
run_07 |
mlgru | byte_dynamic | ternary | 512 / 12 / 1376 | yes | yes | yes | yes | 0 | no |
run_07_r2 |
mlgru | byte_dynamic | ternary | 512 / 12 / 1376 | yes | yes | yes | yes | 0 | no |
run_08 |
transformer | byte | ternary | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_08_v3 |
transformer | byte | ternary | 512 / 12 / 1376 | yes | no | no | no | 0 | no |
run_09 |
mlgru | byte_dynamic | ternary | 512 / 12 / 1376 | yes | yes | yes | yes | 32 | no |
run_10 |
mlgru | byte_dynamic | ternary | 512 / 12 / 1376 | yes | yes | yes | yes | 32 | yes |
run_13 |
mlgru | bpe | ternary | 320 / 8 / 853 | N/A | yes | no | yes | 0 | no |
run_13_r2 |
mlgru | bpe | ternary | 320 / 8 / 853 | N/A | yes | no | yes | 0 | no |
run_13_v3 |
mlgru | bpe | ternary | 320 / 8 / 853 | N/A | yes | no | yes | 0 | no |
run_14 |
mlgru | byte_dynamic | ternary | 320 / 8 / 853 | yes | yes | yes | yes | 0 | no |
run_14_r2 |
mlgru | byte_dynamic | ternary | 320 / 8 / 853 | yes | yes | yes | yes | 0 | no |
run_14_v3 |
mlgru | byte_dynamic | ternary | 320 / 8 / 853 | yes | yes | yes | yes | 0 | no |
run_15 |
mlgru | byte_dynamic | ternary | 320 / 8 / 853 | yes | yes | yes | yes | 0 | no |
run_15_r2 |
mlgru | byte_dynamic | ternary | 320 / 8 / 853 | yes | yes | yes | yes | 0 | no |
run_15_v3 |
mlgru | byte_dynamic | ternary | 320 / 8 / 853 | yes | yes | yes | yes | 0 | no |
run_16 |
mlgru | byte_dynamic | ternary | 320 / 8 / 853 | yes | yes | yes | yes | 0 | no |
run_16_r2 |
mlgru | byte_dynamic | ternary | 320 / 8 / 853 | yes | yes | yes | yes | 0 | no |
These experiments change parameter counts, representations, training signal, budgets, saved steps, and sometimes multiple architectural settings. These differences prevent isolating the benefit of an individual component from these results alone. chinchilla_tokens_per_param is a configured budget; it does not prove that a run completed that budget. Actual saved steps are reported separately.
License
The model repository declares Apache-2.0. Dataset, corpus and tokenizer licenses are separate.
Related resources
- GitHub: igna-s/New-1.58: extraction pipeline, training and evaluation code.
- Hugging Face: Cukinator/cpu1-ablations-final: FP32 model exports.
- Hugging Face: Cukinator/cpu1-ablation-dataset: byte and BPE training data.