Instructions to use logits/sft_robodojo_mot_eefabs_sana_pixel_320x480_sanavideo_aligned_f33fps8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Sana
How to use logits/sft_robodojo_mot_eefabs_sana_pixel_320x480_sanavideo_aligned_f33fps8 with Sana:
# Load the model and infer image from text import torch from app.sana_pipeline import SanaPipeline from torchvision.utils import save_image sana = SanaPipeline("configs/sana_config/1024ms/Sana_1600M_img1024.yaml") sana.from_pretrained("hf://logits/sft_robodojo_mot_eefabs_sana_pixel_320x480_sanavideo_aligned_f33fps8") image = sana( prompt='a cyberpunk cat with a neon sign that says "Sana"', height=1024, width=1024, guidance_scale=5.0, pag_guidance_scale=2.0, num_inference_steps=18, ) - Notebooks
- Google Colab
- Kaggle
sft_robodojo_mot_eefabs_sana_pixel_320x480_sanavideo_aligned_f33fps8
Final RoboDojo ARX-X5 MoT SFT checkpoint: epoch 10, optimizer step 36,180.
Model and initialization
The video expert was initialized from Sana-Video EMA (sana_v2_5b_sft_2076k_ema.pth); the action expert was initialized fresh.
- Model:
SanaRWMMoTAttnResPolicy_5B_P1_D36. - Layout: Sana Pixel 2x2 canvas, height 320 x width 480.
- RoPE: aligned;
action_state_as_context=false. - One canvas prompt, with a separate caption embedder for each expert.
- Absolute robot-base EEF targets only, action mode ratio
[0, 1, 0]; joint slots are masked. - Rotation: column 6D representation. XYZ and gripper use q01/q99 statistics; rotation uses fixed min=-1/max=1 normalization. The included normalization artifact applies to initial state and action targets.
Training
All 3,500 episodes; 33 source rows at 25 Hz. video_fps=8, video frame stride 4: 9 sampled RGB frames and 32 dense action rows per window. Video latents were read from the precomputed cache.
Global batch 512, 10 epochs, AdamW learning rate 1e-4, cosine decay, 2,000 warmup steps, weight decay 1e-4, gradient norm clip 1, action loss weight 1.
This run used the frozen training revision b010eb9a502f0c5fe610a997290cec84ae2a0901 plus the local overlay identified by source-manifest SHA256 82f8f42a7c194aa8e22cb554b8d8445f07a22c4e9af1ee62a8c3bebb5daabdc4. Its original scheduler parameters are recorded in config.yaml and training_manifest.json; subsequent development-branch changes were not applied to this run.
W&B training run. Slurm job 47826 completed with exit code 0.
Files
model/pytorch_model_fsdp.bin: full native MoT state dictionary, including video and action experts.metadata.pth: saved epoch/step and checkpoint metadata.config.yaml: exact frozen source configuration used for SFT.normalization/robodojo_arx_x5_model_fps_25_f33_normalization.json: normalization artifact used for training.training_manifest.json: initialization, sampling, layout and training provenance.upload_metadata.json: file hashes and upload provenance.
This is a model-weight export; optimizer checkpoint files are not included. To train again on a development revision that renamed the SFT options block, rename data.extra.robotwin_sft to data.extra.robot_sft in a new configuration. The archived config here is preserved unchanged.
- Downloads last month
- 4