mlboydaisuke commited on
Commit
a5f5672
·
verified ·
1 Parent(s): 3728fd8

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +23 -21
README.md CHANGED
@@ -33,13 +33,16 @@ Getting it wrong does not throw; it returns vectors that look fine and rank wron
33
 
34
  | build | file | size (MB) | Mac ms* | backend takes | worst cosine vs eager | retrieval budget |
35
  |---|---|---|---|---|---|---|
36
- | fp32 | `embed_paraphrase_multilingual_mpnet_xnnpack_fp32.pte` | 1110.0 | 39.6 | 77.8% | 1.000000 | 0% |
37
- | fp16 | `embed_paraphrase_multilingual_mpnet_xnnpack_fp16.pte` | 555.2 | 67.4 | 67.3% | 1.000000 | 1% |
38
- | Core ML (fp16, iOS) | `embed_paraphrase_multilingual_mpnet_coreml_all.pte` | 555.7 | 7.0 | 100.0% | 0.999997 | 2% |
39
 
40
- \*Mac arm64, median of 10, one 256-token sequence — a reference point for relative
41
- cost, not a device number. Torch eager fp32 on the same machine is
42
- 48.4 ms.
 
 
 
43
 
44
  Cosine is measured against the model run in eager through its own pooling, over eight
45
  sentences. The last column is the one that decides: rank those eight against each
@@ -50,21 +53,20 @@ top-1 results.
50
  ## The attention is eager, and that is the faster export
51
 
52
  `F.scaled_dot_product_attention` does not survive export as one operation. The edge
53
- dialect lowers it through `_safe_softmax`, whose guard against a row with no unmasked
54
- key at all leaves **eleven operations XNNPACK cannot take, in every attention block** —
55
- `scalar_tensor`, `where`, `mul.Scalar`, `logical_not`, `eq`, `full_like`, `any.dim`.
56
- Each one cuts the subgraph in two. The count is exact and does not vary by family:
57
- measured across this shelf, from a 4-layer cross-encoder to a 28-layer causal reranker,
58
- it is 11 per block every time.
59
-
60
- The guard is emitted whether or not it can ever fire, and here it cannot. It triggers
61
- only on `-inf`, which reaches the graph only because the sdpa path hands `F.sdpa` a
62
- **boolean** mask for PyTorch to fill; `attn_implementation="eager"` masks with
63
- `torch.finfo(dtype).min`, a large finite number, and never produces one. So the two
64
- arms differ only on rows that have no unmasked key — sdpa zeroes them, eager gives them
65
- a uniform row — and those are padding rows, which the pooling discards and every real
66
- query row masks out. Measured with all but eight positions masked, as adversarial as
67
- this shape gets, the two graphs agree to 1.4e-07.
68
 
69
  XNNPACK fp32 goes from **62.1% to 77.8%** delegated.
70
 
 
33
 
34
  | build | file | size (MB) | Mac ms* | backend takes | worst cosine vs eager | retrieval budget |
35
  |---|---|---|---|---|---|---|
36
+ | fp32 | `embed_paraphrase_multilingual_mpnet_xnnpack_fp32.pte` | 1110.0 | 32.1 | 77.8% | 1.000000 | 0% |
37
+ | fp16 | `embed_paraphrase_multilingual_mpnet_xnnpack_fp16.pte` | 555.2 | 53.2 | 67.3% | 1.000000 | 1% |
38
+ | Core ML (fp16, iOS) | `embed_paraphrase_multilingual_mpnet_coreml_all.pte` | 555.7 | 6.7 | 100.0% | 0.999997 | 2% |
39
 
40
+ \*Mac arm64, one 256-token sequence, **fastest of five medians of ten** — a reference
41
+ point for relative cost, not a device number. The host shares its cores with other work,
42
+ and a single median does not survive that: the same eager model here measured 19.6 ms and
43
+ 182.8 ms twenty minutes apart. Contention only ever adds time, so the fastest repetition is
44
+ the one that means something. Torch eager fp32, measured the same way, is
45
+ 34.5 ms.
46
 
47
  Cosine is measured against the model run in eager through its own pooling, over eight
48
  sentences. The last column is the one that decides: rank those eight against each
 
53
  ## The attention is eager, and that is the faster export
54
 
55
  `F.scaled_dot_product_attention` does not survive export as one operation. The edge
56
+ dialect lowers it through `_safe_softmax`, whose guard against a row with no unmasked key
57
+ at all leaves **11 operations XNNPACK cannot take, in every attention
58
+ block** — `scalar_tensor`, `where`, `mul.Scalar`, `logical_not`, `eq`, `full_like`, `any.dim`. Each one cuts the subgraph in two.
59
+
60
+ The switch is `attn_implementation="eager"`: transformers then builds the mask
61
+ itself, as `torch.finfo(dtype).min`, instead of handing `F.sdpa` a **boolean** mask
62
+ for PyTorch to fill with `-inf`.
63
+
64
+ The guard is emitted whether or not it can ever fire, and here it cannot: it triggers only
65
+ on `-inf`, and this arm never produces one. So the two differ only about rows that have no
66
+ unmasked key at all — sdpa zeroes them, this one gives them a uniform row — and those are
67
+ padding rows, which the pooling discards and which every real query row masks out anyway.
68
+ Measured with all but eight positions masked, as adversarial as this shape gets, the two
69
+ graphs agree to 1.4e-07.
 
70
 
71
  XNNPACK fp32 goes from **62.1% to 77.8%** delegated.
72