Instructions to use Vphuc/PP-OCRv5-mobile-rec-vi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PaddleOCR
How to use Vphuc/PP-OCRv5-mobile-rec-vi with PaddleOCR:
# Please refer to the document for information on how to use the model. # https://paddlepaddle.github.io/PaddleOCR/latest/en/version3.x/module_usage/module_overview.html
- Notebooks
- Google Colab
- Kaggle
PP-OCRv5-mobile-rec-vi
A fine-tune of latin_PP-OCRv5_mobile_rec whose classifier head was widened from 838 to 908
classes, so the model has an output class for the Vietnamese precomposed letters the base model
cannot express at all.
Read this before you use it
On a held-out test set of Vietnamese exam papers, against stock latin_PP-OCRv5_mobile_rec:
| this checkpoint | stock base model | |
|---|---|---|
| word error rate | 32.27% | 55.73% |
| character error rate | 12.4% | 20.98% |
characters emitted from U+1EA0–U+1EF9 |
1,270 | 0 |
| characters deleted | 356 | 1,010 |
| characters inserted | 510 | 392 |
Use it if you care about words, not characters. WER is 12.46 points better. CER is 1.95 points worse, and paired by document the difference is +1.12% with a 95% interval of [−1.92%, +4.13%] — it includes zero, so on characters this is a tie, not a win.
The two metrics disagree because the two models fail differently and CER cannot tell the two failures
apart. The base model drops the letters it has no class for (một → mt, định → đnh): 1,010
deletions. This model attempts them: 356 deletions but more substitutions. Deleting a character and
mis-predicting one both cost exactly 1 in CER, so they trade off almost evenly. They do not trade evenly
in WER: a diacritic the base model drops destroys the whole word, while a diacritic this model gets right
saves it.
Where it is worse: base letters, not diacritics. Substitutions of unaccented letters number 518
against the base model's 79. It has learned to read tone marks while reading the underlying letter less
reliably — đjnh for định is the characteristic error: the dot-below is in the right place, the ị
became j. That is the next thing to fix.
Why the base model needed surgery
Measured on 250 hand-verified regions of Vietnamese exam papers: of the characters predicted by stock
PaddleOCR recognition models, 0 fall in the Unicode block U+1EA0–U+1EF9, while 1,672 of 22,880
characters in the ground truth do — 7.31%. That block holds every Vietnamese letter with a
hook-above or dot-below tone, and every tone on ă â ê ô ơ ư. The missing letters are not misread, they
are unrepresentable: the network has no output class for them, so it deletes them.
Widening the table is therefore a precondition for any further work on Vietnamese, and it is the only thing this repository changes about the architecture.
How the head was widened
CTC class order is blank, then the character table in file order, then the space class. The 836
entries of the base table were kept in their original order and 70 characters appended, so the
trained rows could be copied across instead of being discarded:
new[0] = old[0] blank
new[1..836] = old[1..836] the whole base table, unchanged
new[837..906] = fresh the 70 appended characters
new[907] = old[837] space, moved from 837 to 907
Fresh weight columns are drawn from a normal distribution with the standard deviation of the existing
matrix; fresh biases take the mean of the existing biases. head.ctc_head.fc.weight is [120, C], so
the patch runs along axis 1; the bias along axis 0. 966 of the model's 968 parameters are untouched.
The two class-sized parameters of the auxiliary NRTR branch
(head.gtc_head.embedding.embedding.weight, head.gtc_head.tgt_word_prj.weight) are not patched
and are re-initialised by PaddleOCR. Its special-token layout could not be derived, and guessing it
would teach the model to map characters to the wrong classes. Re-initialising that branch is not the
same as switching it off — see Do not switch off the GTC branch below.
A quirk of the base table, measured and harmless
dict.txt has 906 lines but only 842 distinct characters: 64 accented Latin letters
(À Á Â … ÿ ƒ) appear twice in upstream's latin_dict.txt, so 64 classes are decode-equivalent
duplicates. In principle that is a second way to double a character, because CTC merges only identical
adjacent labels and two different class indices are not identical. In practice it does not happen:
across all 250 test regions, adjacent identical pairs drawn from those 64 letters number 0 for this
checkpoint and 1 for the previous one. Left as-is, because de-duplicating the table would change the
class count and invalidate the patched weights.
The 70 added characters
Ạ ạ Ả ả Ấ ấ Ầ ầ ẩ ẫ Ậ ậ Ắ ắ ằ ẳ ặ ẻ ẽ Ế ế Ề ề Ể ể ễ Ệ ệ ỉ Ị ị Ọ ọ ỏ Ố ố Ồ ồ Ổ ổ ỗ Ộ ộ ớ Ờ ờ ở ỡ ợ
Ụ ụ Ủ ủ Ứ ứ ừ ử Ữ ữ Ự ự Ỳ ỳ ỷ Ỹ ỹ – ₁ ₂ ₃
Training data
945 line crops cut from photographed and scanned Vietnamese university exam papers — 829 for training, 116 for validation. Labels are NFC-normalised. The corpus itself is not published.
Lines were filtered by four rules: the detector's own reading disagreeing with the label beyond a
threshold, vertically overlapping detected rows, row polygons hanging outside the region crop, and — new
in this round — labels that need more CTC output frames than the training geometry can give them
(18 lines). Label files carry four tab-separated columns, image, label, width, height, because
aspect-ratio-aware batching reads the width and height from the label file rather than opening the image.
What the previous release got wrong, and what fixed it
The previous checkpoint emitted the old class and the new class for the same character — một →
môộxt, của → cựũủa, plus a spurious . on every line. It inserted 1,616 characters against
this checkpoint's 510, and its CER was 32.59% against the base model's 20.16%.
ground truth : Thời gian làm bài 110 phút
stock base : Thi gian làm bài 110 phút <- 'ờ' dropped, unrepresentable
previous : ThTời gian làạm bài 110 phúết. <- doubled, plus spurious '.'
this checkpoint: Thời gian làm bài 110 phút <- exact
Exact-match accuracy on validation went from 0.008 to 0.235: 27 of 115 validation lines now match character for character, against 1 before. A line carrying a doubled character cannot match exactly, so that number is the defect's death certificate, 27 times over.
Three changes did it, and all three are training-side:
- Aspect-ratio-aware batching (
ds_width: true,MultiScaleSamplerwithmax_w: 960). The frame count now tracks the actual amount of text instead of padding every line out to a fixed 960 px. At a fixed 960, the median line filled 63% of its frame and the 90th-percentile line filled 21%; the rest was black padding, and that is where an undecided model had room to emit both classes. - Learning rate 1e-4 → 2.5e-5. The effective batch is 5.6, not the configured 8, because
fix_bs: falseholds total batch pixels constant rather than the sample count. Upstream trains this architecture at 5e-4 with batch 128, which scales linearly to about 2.2e-5 at batch 5.6 — the previous round ran about 4.5× too hot, and validation fell inside the warm-up. Metric.main_indicator: acc→norm_edit_dis. Withaccpinned near zero and the comparison>=,best_accuracywas written the single timeaccwas non-zero and never again. The previous release was therefore epoch 4, not epoch 36 as its model card claimed. This is now a real selection: these weights are epoch 62 of 68, chosen on validation edit distance.
Do not switch off the GTC branch
With use_guide: True, EncoderWithSVTR.forward sets z.stop_gradient = True
(ppocr/modeling/necks/rnn.py, whose docstring says "neck gradients do not propagate back to the
backbone"). The CTC branch's gradient therefore never reaches the backbone: the auxiliary GTC/NRTR
branch is the backbone's only gradient path. Setting Loss.weight_2: 0.0 does not remove a noisy
signal, it freezes the backbone. Two 10-epoch probes differing in exactly that one value:
validation norm_edit_dis |
ep 1 | 2 | 4 | 6 | 8 | 12 | 15 | 18 |
|---|---|---|---|---|---|---|---|---|
weight_2 = 0.0 |
0.595 | 0.665 | 0.713 | 0.651 | 0.614 | 0.291 | 0.202 | 0.177 |
weight_2 = 1.0 |
0.646 | 0.713 | 0.741 | 0.791 | 0.809 | 0.825 | 0.844 | 0.850 |
The first diverges — training loss also rose, 36.9 to 63.8, while the learning rate was decaying —
rather than overfitting. Leave weight_2 at its default.
Numbers
Test set: 250 regions from 15 exams, every label read and confirmed by hand, held out from training and
from every parameter choice, and measured once per training round. The text slice excludes regions
that are mostly mathematics. CER no-diacritics strips tone and vowel marks from both sides, so the gap
between the two CER columns is what diacritics cost.
| model | CER (text) |
95% CI | CER no-diac. | WER | ins | del | subs |
|---|---|---|---|---|---|---|---|
PP-OCRv6_medium_rec (stock) |
20.98% | [16.1, 27.5] | 14.85% | 55.97% | 408 | 683 | 750 |
latin_PP-OCRv5_mobile_rec (stock) |
20.16% | [15.5, 26.5] | 17.84% | 55.73% | 392 | 1010 | 462 |
| previous release (epoch 4) | 32.59% | [25.2, 40.1] | 28.17% | 62.18% | 1616 | 464 | 933 |
| this checkpoint (epoch 62) | 22.11% | [15.1, 29.9] | 17.64% | 43.27% | 510 | 356 | 1178 |
Characters emitted, counted raw across all 250 regions:
in U+1EA0–U+1EF9 |
total characters | against ground truth | |
|---|---|---|---|
| ground truth | 1,672 | 22,880 | — |
latin_PP-OCRv5_mobile_rec (stock) |
0 | 18,023 | 79%, short by 4,857 — it drops them |
| previous release | 1,505 | 22,621 | 99%, but by way of 1,616 inserted characters |
| this checkpoint | 1,270 | 20,320 | 89%, with 510 inserted |
The base model emits fewer characters not because it is terse but because it deletes what it cannot represent. This checkpoint recovers 76% of the ground truth's characters from the missing block.
Validation split, 116 lines, norm_edit_dis (higher is better) and exact-match accuracy:
| model | norm_edit_dis | accuracy |
|---|---|---|
| widened table, before training | 0.8377 | 0.0435 |
| previous release (epoch 4) | 0.749 | 0.008 |
| this checkpoint (epoch 62) | 0.9072 | 0.2348 |
The previous round never beat the untrained widened table at any of its 53 evaluations. This one passes it at epoch 6 and climbs monotonically for 40 epochs.
Inference geometry
RecResizeImg treats the configured width as a floor, not a fixed frame: it computes
imgW = imgH * max(imgW_cfg / imgH, actual_ratio), so a line narrower than the configured width is
zero-padded out to it, and a wider one keeps its true ratio. That number travels:
tools/export_model.py copies the Eval transforms verbatim into inference.yml, and PaddleX reads
RecResizeImg.image_shape back from that file at inference time.
This build ships a floor of 160, and that matters much more than it did for the previous release. Measured on the validation split with these weights:
| width floor | 64 | 96 | 160 | 240 | 320 | 480 | 960 |
|---|---|---|---|---|---|---|---|
norm_edit_dis |
0.9017 | 0.9017 | 0.9072 | 0.8967 | 0.8707 | 0.8084 | 0.5948 |
At a 960 floor the model collapses. The previous release shipped 960 and was trained at a fixed 960, so it was internally consistent; sweeping its width was worth only 0.57 points on a healthy model. The lesson is not that width matters but that a mismatch between training and inference geometry does. This model was trained on true aspect ratios and must be run on them. 64 and 96 are identical because below about 160 the floor stops binding for every line in the set.
If you re-export this model yourself, keep Eval…RecResizeImg.image_shape at [3, 48, 160].
Usage
from paddleocr import PaddleOCR
ocr = PaddleOCR(
text_recognition_model_name="latin_PP-OCRv5_mobile_rec", # the architecture
text_recognition_model_dir="path/to/this/repo", # these weights
)
The model name still has to be the base model's: PaddleOCR reads the name to decide which architecture to build, and the directory to decide which weights to load.
Files
| file | what it is |
|---|---|
inference.json, inference.pdiparams, inference.yml |
the inference build, what PaddleOCR loads |
best_accuracy.pdparams |
the training checkpoint (epoch 62), for continuing training |
dict.txt |
the 906-entry character table, base order plus the 70 appended |
vi_rec_v2.yaml |
the training recipe these weights were trained with |
vi_rec_v1.yaml |
the previous release's recipe, kept only so that round can be reproduced |
vi_rec_v1.yaml does not describe the weights in this repository. It is the recipe of the superseded
epoch-4 release — fixed width 960, learning rate 1e-4, main_indicator: acc — and it is kept because the
comparisons in this card are only checkable if that run can be rebuilt. Read vi_rec_v2.yaml for these
weights.
Provenance
These weights are epoch 62 of a planned 75. The run was stopped by the host at epoch 68 when the machine ran out of RAM. The remaining epochs were not re-run: validation had been flat since epoch 43, oscillating in 0.893–0.907 without trend, and the cosine schedule had decayed the learning rate to 1e-6 by epoch 68 — smaller steps than the measurement's own noise, which on a 116-line split is 0.0087 of accuracy per line.
Limitations
- 945 lines is very little data. Four of the added characters (
Ậ ễ Ồ Ổ) appear exactly once, so they cannot be learned from this set at all; they need rendered data. - Almost no handwriting. The corpus is printed exam papers, photographed or scanned.
- Mathematics is read as running text. A separate formula model handles that.
- Unaccented letters are read less reliably than by the base model (518 substitutions against 79). If your text has no Vietnamese diacritics, the base model is the better choice.
- Character error rate is a tie with the base model, not an improvement. The gain is in word error rate.
License
Apache 2.0, inherited from the base model latin_PP-OCRv5_mobile_rec by PaddlePaddle.
- Downloads last month
- 2
Model tree for Vphuc/PP-OCRv5-mobile-rec-vi
Base model
PaddlePaddle/latin_PP-OCRv5_mobile_rec