You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

PP-OCRv5-mobile-rec-vi

A fine-tune of latin_PP-OCRv5_mobile_rec whose classifier head was widened from 838 to 908 classes, so the model has an output class for the Vietnamese precomposed letters the base model cannot express at all.

Read this before you use it

On a held-out test set of Vietnamese exam papers, against stock latin_PP-OCRv5_mobile_rec:

this checkpoint stock base model
word error rate 32.27% 55.73%
character error rate 12.4% 20.98%
characters emitted from U+1EA0–U+1EF9 1,270 0
characters deleted 356 1,010
characters inserted 510 392

Use it if you care about words, not characters. WER is 12.46 points better. CER is 1.95 points worse, and paired by document the difference is +1.12% with a 95% interval of [−1.92%, +4.13%] — it includes zero, so on characters this is a tie, not a win.

The two metrics disagree because the two models fail differently and CER cannot tell the two failures apart. The base model drops the letters it has no class for (một → mt, định → đnh): 1,010 deletions. This model attempts them: 356 deletions but more substitutions. Deleting a character and mis-predicting one both cost exactly 1 in CER, so they trade off almost evenly. They do not trade evenly in WER: a diacritic the base model drops destroys the whole word, while a diacritic this model gets right saves it.

Where it is worse: base letters, not diacritics. Substitutions of unaccented letters number 518 against the base model's 79. It has learned to read tone marks while reading the underlying letter less reliably — đjnh for định is the characteristic error: the dot-below is in the right place, the ị became j. That is the next thing to fix.

Why the base model needed surgery

Measured on 250 hand-verified regions of Vietnamese exam papers: of the characters predicted by stock PaddleOCR recognition models, 0 fall in the Unicode block U+1EA0–U+1EF9, while 1,672 of 22,880 characters in the ground truth do — 7.31%. That block holds every Vietnamese letter with a hook-above or dot-below tone, and every tone on ă â ê ô ơ ư. The missing letters are not misread, they are unrepresentable: the network has no output class for them, so it deletes them.

Widening the table is therefore a precondition for any further work on Vietnamese, and it is the only thing this repository changes about the architecture.

How the head was widened

CTC class order is blank, then the character table in file order, then the space class. The 836 entries of the base table were kept in their original order and 70 characters appended, so the trained rows could be copied across instead of being discarded:

new[0]        = old[0]          blank
new[1..836]   = old[1..836]     the whole base table, unchanged
new[837..906] = fresh           the 70 appended characters
new[907]      = old[837]        space, moved from 837 to 907

Fresh weight columns are drawn from a normal distribution with the standard deviation of the existing matrix; fresh biases take the mean of the existing biases. head.ctc_head.fc.weight is [120, C], so the patch runs along axis 1; the bias along axis 0. 966 of the model's 968 parameters are untouched.

The two class-sized parameters of the auxiliary NRTR branch (head.gtc_head.embedding.embedding.weight, head.gtc_head.tgt_word_prj.weight) are not patched and are re-initialised by PaddleOCR. Its special-token layout could not be derived, and guessing it would teach the model to map characters to the wrong classes. Re-initialising that branch is not the same as switching it off — see Do not switch off the GTC branch below.

A quirk of the base table, measured and harmless

dict.txt has 906 lines but only 842 distinct characters: 64 accented Latin letters (À Á Â … ÿ ƒ) appear twice in upstream's latin_dict.txt, so 64 classes are decode-equivalent duplicates. In principle that is a second way to double a character, because CTC merges only identical adjacent labels and two different class indices are not identical. In practice it does not happen: across all 250 test regions, adjacent identical pairs drawn from those 64 letters number 0 for this checkpoint and 1 for the previous one. Left as-is, because de-duplicating the table would change the class count and invalidate the patched weights.

The 70 added characters

Ạ ạ Ả ả Ấ ấ Ầ ầ ẩ ẫ Ậ ậ Ắ ắ ằ ẳ ặ ẻ ẽ Ế ế Ề ề Ể ể ễ Ệ ệ ỉ Ị ị Ọ ọ ỏ Ố ố Ồ ồ Ổ ổ ỗ Ộ ộ ớ Ờ ờ ở ỡ ợ
Ụ ụ Ủ ủ Ứ ứ ừ ử Ữ ữ Ự ự Ỳ ỳ ỷ Ỹ ỹ – ₁ ₂ ₃

Training data

945 line crops cut from photographed and scanned Vietnamese university exam papers — 829 for training, 116 for validation. Labels are NFC-normalised. The corpus itself is not published.

Lines were filtered by four rules: the detector's own reading disagreeing with the label beyond a threshold, vertically overlapping detected rows, row polygons hanging outside the region crop, and — new in this round — labels that need more CTC output frames than the training geometry can give them (18 lines). Label files carry four tab-separated columns, image, label, width, height, because aspect-ratio-aware batching reads the width and height from the label file rather than opening the image.

What the previous release got wrong, and what fixed it

The previous checkpoint emitted the old class and the new class for the same character — một → môộxt, của → cựũủa, plus a spurious . on every line. It inserted 1,616 characters against this checkpoint's 510, and its CER was 32.59% against the base model's 20.16%.

ground truth : Thời gian làm bài 110 phút
stock base   : Thi gian làm bài 110 phút        <- 'ờ' dropped, unrepresentable
previous      : ThTời gian làạm bài 110 phúết.  <- doubled, plus spurious '.'
this checkpoint: Thời gian làm bài 110 phút     <- exact

Exact-match accuracy on validation went from 0.008 to 0.235: 27 of 115 validation lines now match character for character, against 1 before. A line carrying a doubled character cannot match exactly, so that number is the defect's death certificate, 27 times over.

Three changes did it, and all three are training-side:

  1. Aspect-ratio-aware batching (ds_width: true, MultiScaleSampler with max_w: 960). The frame count now tracks the actual amount of text instead of padding every line out to a fixed 960 px. At a fixed 960, the median line filled 63% of its frame and the 90th-percentile line filled 21%; the rest was black padding, and that is where an undecided model had room to emit both classes.
  2. Learning rate 1e-4 → 2.5e-5. The effective batch is 5.6, not the configured 8, because fix_bs: false holds total batch pixels constant rather than the sample count. Upstream trains this architecture at 5e-4 with batch 128, which scales linearly to about 2.2e-5 at batch 5.6 — the previous round ran about 4.5× too hot, and validation fell inside the warm-up.
  3. Metric.main_indicator: acc → norm_edit_dis. With acc pinned near zero and the comparison >=, best_accuracy was written the single time acc was non-zero and never again. The previous release was therefore epoch 4, not epoch 36 as its model card claimed. This is now a real selection: these weights are epoch 62 of 68, chosen on validation edit distance.

Do not switch off the GTC branch

With use_guide: True, EncoderWithSVTR.forward sets z.stop_gradient = True (ppocr/modeling/necks/rnn.py, whose docstring says "neck gradients do not propagate back to the backbone"). The CTC branch's gradient therefore never reaches the backbone: the auxiliary GTC/NRTR branch is the backbone's only gradient path. Setting Loss.weight_2: 0.0 does not remove a noisy signal, it freezes the backbone. Two 10-epoch probes differing in exactly that one value:

validation norm_edit_dis ep 1 2 4 6 8 12 15 18
weight_2 = 0.0 0.595 0.665 0.713 0.651 0.614 0.291 0.202 0.177
weight_2 = 1.0 0.646 0.713 0.741 0.791 0.809 0.825 0.844 0.850

The first diverges — training loss also rose, 36.9 to 63.8, while the learning rate was decaying — rather than overfitting. Leave weight_2 at its default.

Numbers

Test set: 250 regions from 15 exams, every label read and confirmed by hand, held out from training and from every parameter choice, and measured once per training round. The text slice excludes regions that are mostly mathematics. CER no-diacritics strips tone and vowel marks from both sides, so the gap between the two CER columns is what diacritics cost.

model CER (text) 95% CI CER no-diac. WER ins del subs
PP-OCRv6_medium_rec (stock) 20.98% [16.1, 27.5] 14.85% 55.97% 408 683 750
latin_PP-OCRv5_mobile_rec (stock) 20.16% [15.5, 26.5] 17.84% 55.73% 392 1010 462
previous release (epoch 4) 32.59% [25.2, 40.1] 28.17% 62.18% 1616 464 933
this checkpoint (epoch 62) 22.11% [15.1, 29.9] 17.64% 43.27% 510 356 1178

Characters emitted, counted raw across all 250 regions:

in U+1EA0–U+1EF9 total characters against ground truth
ground truth 1,672 22,880 —
latin_PP-OCRv5_mobile_rec (stock) 0 18,023 79%, short by 4,857 — it drops them
previous release 1,505 22,621 99%, but by way of 1,616 inserted characters
this checkpoint 1,270 20,320 89%, with 510 inserted

The base model emits fewer characters not because it is terse but because it deletes what it cannot represent. This checkpoint recovers 76% of the ground truth's characters from the missing block.

Validation split, 116 lines, norm_edit_dis (higher is better) and exact-match accuracy:

model norm_edit_dis accuracy
widened table, before training 0.8377 0.0435
previous release (epoch 4) 0.749 0.008
this checkpoint (epoch 62) 0.9072 0.2348

The previous round never beat the untrained widened table at any of its 53 evaluations. This one passes it at epoch 6 and climbs monotonically for 40 epochs.

Inference geometry

RecResizeImg treats the configured width as a floor, not a fixed frame: it computes imgW = imgH * max(imgW_cfg / imgH, actual_ratio), so a line narrower than the configured width is zero-padded out to it, and a wider one keeps its true ratio. That number travels: tools/export_model.py copies the Eval transforms verbatim into inference.yml, and PaddleX reads RecResizeImg.image_shape back from that file at inference time.

This build ships a floor of 160, and that matters much more than it did for the previous release. Measured on the validation split with these weights:

width floor 64 96 160 240 320 480 960
norm_edit_dis 0.9017 0.9017 0.9072 0.8967 0.8707 0.8084 0.5948

At a 960 floor the model collapses. The previous release shipped 960 and was trained at a fixed 960, so it was internally consistent; sweeping its width was worth only 0.57 points on a healthy model. The lesson is not that width matters but that a mismatch between training and inference geometry does. This model was trained on true aspect ratios and must be run on them. 64 and 96 are identical because below about 160 the floor stops binding for every line in the set.

If you re-export this model yourself, keep Eval…RecResizeImg.image_shape at [3, 48, 160].

Usage

from paddleocr import PaddleOCR

ocr = PaddleOCR(
    text_recognition_model_name="latin_PP-OCRv5_mobile_rec",  # the architecture
    text_recognition_model_dir="path/to/this/repo",           # these weights
)

The model name still has to be the base model's: PaddleOCR reads the name to decide which architecture to build, and the directory to decide which weights to load.

Files

file what it is
inference.json, inference.pdiparams, inference.yml the inference build, what PaddleOCR loads
best_accuracy.pdparams the training checkpoint (epoch 62), for continuing training
dict.txt the 906-entry character table, base order plus the 70 appended
vi_rec_v2.yaml the training recipe these weights were trained with
vi_rec_v1.yaml the previous release's recipe, kept only so that round can be reproduced

vi_rec_v1.yaml does not describe the weights in this repository. It is the recipe of the superseded epoch-4 release — fixed width 960, learning rate 1e-4, main_indicator: acc — and it is kept because the comparisons in this card are only checkable if that run can be rebuilt. Read vi_rec_v2.yaml for these weights.

Provenance

These weights are epoch 62 of a planned 75. The run was stopped by the host at epoch 68 when the machine ran out of RAM. The remaining epochs were not re-run: validation had been flat since epoch 43, oscillating in 0.893–0.907 without trend, and the cosine schedule had decayed the learning rate to 1e-6 by epoch 68 — smaller steps than the measurement's own noise, which on a 116-line split is 0.0087 of accuracy per line.

Limitations

  • 945 lines is very little data. Four of the added characters (Ậ ễ Ồ Ổ) appear exactly once, so they cannot be learned from this set at all; they need rendered data.
  • Almost no handwriting. The corpus is printed exam papers, photographed or scanned.
  • Mathematics is read as running text. A separate formula model handles that.
  • Unaccented letters are read less reliably than by the base model (518 substitutions against 79). If your text has no Vietnamese diacritics, the base model is the better choice.
  • Character error rate is a tie with the base model, not an improvement. The gain is in word error rate.

License

Apache 2.0, inherited from the base model latin_PP-OCRv5_mobile_rec by PaddlePaddle.

Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vphuc/PP-OCRv5-mobile-rec-vi

Finetuned
(2)
this model