Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions

The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis,...

Full description

Bibliographic Details
Main Authors: Ce Zhang, Weilan Wang, Guowei Zhang
Format: Article
Language:English
Published: MDPI AG 2022-11-01
Series:Electronics
Subjects:
Online Access:https://www.mdpi.com/2079-9292/11/23/3919
_version_ 1797463447316725760
author Ce Zhang
Weilan Wang
Guowei Zhang
author_facet Ce Zhang
Weilan Wang
Guowei Zhang
author_sort Ce Zhang
collection DOAJ
description The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.
first_indexed 2024-03-09T17:50:49Z
format Article
id doaj.art-e11b18c912e94a9fa2e15f46f3f8556c
institution Directory Open Access Journal
issn 2079-9292
language English
last_indexed 2024-03-09T17:50:49Z
publishDate 2022-11-01
publisher MDPI AG
record_format Article
series Electronics
spelling doaj.art-e11b18c912e94a9fa2e15f46f3f8556c2023-11-24T10:47:43ZengMDPI AGElectronics2079-92922022-11-011123391910.3390/electronics11233919Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource ConditionsCe Zhang0Weilan Wang1Guowei Zhang2Key Laboratory of China’s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, ChinaKey Laboratory of China’s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, ChinaLinkDoc Technology, Beijing 100089, ChinaThe construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.https://www.mdpi.com/2079-9292/11/23/3919historical Tibetan documentscharacter annotationcharacter extractiondata augmentationcharacter recognition
spellingShingle Ce Zhang
Weilan Wang
Guowei Zhang
Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
Electronics
historical Tibetan documents
character annotation
character extraction
data augmentation
character recognition
title Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_full Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_fullStr Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_full_unstemmed Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_short Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_sort construction of a character dataset for historical uchen tibetan documents under low resource conditions
topic historical Tibetan documents
character annotation
character extraction
data augmentation
character recognition
url https://www.mdpi.com/2079-9292/11/23/3919
work_keys_str_mv AT cezhang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions
AT weilanwang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions
AT guoweizhang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions