Instructions to use MC7ever/pe-single-L14 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MC7ever/pe-single-L14 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="MC7ever/pe-single-L14")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModel processor = AutoProcessor.from_pretrained("MC7ever/pe-single-L14") model = AutoModel.from_pretrained("MC7ever/pe-single-L14", device_map="auto") - PerceptionEncoder
How to use MC7ever/pe-single-L14 with PerceptionEncoder:
# Use any PE model as a vision encoder import core.vision_encoder.pe as pe model = pe.VisionTransformer.from_config("MC7ever/pe-single-L14", pretrained=True) - Notebooks
- Google Colab
- Kaggle
pe-single-L14 โ one fused Meta Perception Encoder, no router
Single-file joint embedding model (audio + image/video + text, 1024-d shared space, 1.03B params) fusing four Meta PE family members:
| source | HF id | what was fused |
|---|---|---|
| PE-AV base (anchor) | facebook/pe-av-base |
full joint model kept as scaffold |
| PE A-Frame base | facebook/pe-a-frame-base |
audio tower, weight 0.25 |
| PE-Core-L14-336 | facebook/PE-Core-L14-336 |
vision trunk (near-identical to anchor video tower โ init copy, ~no-op) |
| PE-Spatial-L14-448 | facebook/PE-Spatial-L14-448 |
vision trunk, weight 0.0 in this tag (hurt COCO alignment; see below) |
Why L-scale: PE-Core/Spatial-G14 are width-1536 ร 50 layers while PE-AV/A-Frame are width-1024 โ cross-width tensors cannot be weight-averaged. L/14 (1024-wide, 24 layers, patch 14) is the scale where Core + Spatial + AV-video share width/depth/patch, so true weight fusion is possible.
Status (2026-10-04): v3 TUNED โ local MPS fine-tune, new SOTA
Heads-only contrastive tune (COCO-train 4000 pairs x 2 epochs, frozen towers, local Apple-Silicon MPS, fp32). Full-slice eval (COCO n=150, ESC-50 n=200):
| model | COCO t2i R@1 / R@5 | COCO i2t R@1 / R@5 | ESC-50 acc |
|---|---|---|---|
facebook/pe-av-base (anchor) |
0.720 / 0.953 | 0.013 / 0.073 | 0.885 |
| frozen fusion (this repo, previous tag) | 0.713 / 0.953 | 0.013 / 0.047 | 0.895 |
| v3 tuned (this revision) | 0.827 / 0.980 | 0.727 / 0.973 | 0.895 |
The tune unlocked the text direction (i2t 0.013 -> 0.727) that no weight blend touched, and lifted t2i +0.11. ESC unchanged (v1 trained video<->text only; audio<->text tuning is next). Previous certification notes below remain the record for the frozen init.
Full-slice validation (COCO n=150, ESC-50 n=200) vs anchor:
| model | COCO t2i R@1 / R@5 | COCO i2t R@1 | ESC-50 acc |
|---|---|---|---|
facebook/pe-av-base (anchor) |
0.720 / 0.953 | 0.013 | 0.885 |
| this tag | 0.713 / 0.953 | 0.013 | 0.895 |
Verdict: statistical tie (diffs within SE). Weight fusion preserves anchor quality while carrying A-Frame/Core/Spatial weights in one file โ it does not beat the anchor on these benches by itself. Measured gains are deferred to alignment fine-tuning, for which this tag is the frozen init. Certification: 9-build grid + 18-genome evolution (fixed machinery) + full-slice validation, all agreeing (audio light, vision ~zero).
15-metric suite (COCO n=300, ESC-50 5x120, MPS fp32) vs anchor: vision tied (V1 t2i R@1 0.613 vs 0.617, V5 0.637 vs 0.643, V4 min 0.527 vs 0.547; V2/V3 at chance for both โ text embeds are dummy-video dominated in this harness, a comparative-only handicap); audio 4/5 folds +0.017..+0.025 (fold3 -0.017), prompt-mean +0.008; joint diagnostics mixed tiny deltas (J5 +0.033). No regression anywhere; consistent small audio edge. Full table: suite.json in the GitHub repo.
- Evolution rerun in progress after two voided attempts (documented below) โ if it certifies a better genome, that becomes the next revision.
- Layer-depth probe: naive mid-network readout does NOT beat the joint output (t2i R@1 โค 0.06 vs 0.79) โ mid-network features need trained per-layer pooling (fine-tune phase), not a free readout change.
- Voided runs: (1) blend-of-blend contamination via shared-storage state_dict;
(2) silent no-op
load_state_dicton this model's dual-prefix key layout; (3) joint-path audio mismatch that faked ESC collapse at high audio weight. All fixed (pristine clones,copy_by param name, standalone-tower-only audio mapping) before trusting any evolution number.
How it was built
- Audio tower: exact-shape suffix-mapped average of AV-audio and A-Frame-audio
(both
pe_audio_encoder, 1024/16L/8H โ 422/422 tensors matched). - Video trunk: explicit native-PE โ timm key map (patch_embed, cls_token, pos_embed@336px, norms, QKV, MLP ร24 blocks). Pool/head/proj kept from anchor.
- Blend weights chosen by measured 3ร3 grid search over COCO-retrieval +
ESC-50 slices (script
search_pe_weights.py), not by vibes. Vision donor weight degraded COCO t2i monotonically (0.766 โ 0.64), so the winning tag keeps anchor vision and blends audio at 0.25.
Usage
from transformers import AutoModel, AutoProcessor
m = AutoModel.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
p = AutoProcessor.from_pretrained("MC7ever/pe-single-L14", trust_remote_code=True)
inputs = p(videos=[frames], audio=[waveform], text=["a cat playing in rain"],
return_tensors="pt")
out = m(**inputs) # out.video_embeds, out.audio_embeds, out.text_video_embeds, ...
Retrieval: dot-product video_embeds @ text_video_embeds.T.
Zero-shot audio: argmax over audio_embeds @ text_audio_embeds.T with
"the sound of {label}" prompts.
Honest limits
- Numbers above are CPU-run slices (COCO n=64, ESC n=100), comparative anchor-vs-variant only โ not official MMEB/PE-paper figures.
- i2t (textโimage caption ranking over 150 similar COCO captions) sits near floor for anchor and variants alike; t2i and audio carry the signal.
- Encoder only: no generation, detection, or segmentation heads.
- Next: evolution-based per-tower weights, intermediate-layer readout (PE paper: best features are mid-network), GPU alignment fine-tune.
Reproduce
fuse_pe_single.py (fusion) ยท eval_pe_bench.py (COCO+ESC-50 bench) ยท
search_pe_weights.py (grid) ยท evo_fuse.py (evolution) ยท layer_sweep.py
(depth probe). Benchmark slice metadata: MMEB MSCOCO_i2t rows + COCO val2014
images + ashraq/esc50 streaming clips.
- Downloads last month
- 48
Evaluation results
- t2i R@1 (search slice n=64)self-reported0.766
- t2i R@5 (approx, search slice)self-reported0.950
- ESC-50 accuracy (search slice)self-reported0.960