Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis,...
Main Authors: | , , |
---|---|
Format: | Article |
Language: | English |
Published: |
MDPI AG
2022-11-01
|
Series: | Electronics |
Subjects: | |
Online Access: | https://www.mdpi.com/2079-9292/11/23/3919 |
_version_ | 1797463447316725760 |
---|---|
author | Ce Zhang Weilan Wang Guowei Zhang |
author_facet | Ce Zhang Weilan Wang Guowei Zhang |
author_sort | Ce Zhang |
collection | DOAJ |
description | The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets. |
first_indexed | 2024-03-09T17:50:49Z |
format | Article |
id | doaj.art-e11b18c912e94a9fa2e15f46f3f8556c |
institution | Directory Open Access Journal |
issn | 2079-9292 |
language | English |
last_indexed | 2024-03-09T17:50:49Z |
publishDate | 2022-11-01 |
publisher | MDPI AG |
record_format | Article |
series | Electronics |
spelling | doaj.art-e11b18c912e94a9fa2e15f46f3f8556c2023-11-24T10:47:43ZengMDPI AGElectronics2079-92922022-11-011123391910.3390/electronics11233919Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource ConditionsCe Zhang0Weilan Wang1Guowei Zhang2Key Laboratory of China’s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, ChinaKey Laboratory of China’s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, ChinaLinkDoc Technology, Beijing 100089, ChinaThe construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.https://www.mdpi.com/2079-9292/11/23/3919historical Tibetan documentscharacter annotationcharacter extractiondata augmentationcharacter recognition |
spellingShingle | Ce Zhang Weilan Wang Guowei Zhang Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions Electronics historical Tibetan documents character annotation character extraction data augmentation character recognition |
title | Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions |
title_full | Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions |
title_fullStr | Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions |
title_full_unstemmed | Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions |
title_short | Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions |
title_sort | construction of a character dataset for historical uchen tibetan documents under low resource conditions |
topic | historical Tibetan documents character annotation character extraction data augmentation character recognition |
url | https://www.mdpi.com/2079-9292/11/23/3919 |
work_keys_str_mv | AT cezhang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions AT weilanwang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions AT guoweizhang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions |