Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions

The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis,...

Full description

Bibliographic Details
Main Authors:	Ce Zhang, Weilan Wang, Guowei Zhang
Format:	Article
Language:	English
Published:	MDPI AG 2022-11-01
Series:	Electronics
Subjects:	historical Tibetan documents character annotation character extraction data augmentation character recognition
Online Access:	https://www.mdpi.com/2079-9292/11/23/3919

_version_	1797463447316725760
author	Ce Zhang Weilan Wang Guowei Zhang
author_facet	Ce Zhang Weilan Wang Guowei Zhang
author_sort	Ce Zhang
collection	DOAJ
description	The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.
first_indexed	2024-03-09T17:50:49Z
format	Article
id	doaj.art-e11b18c912e94a9fa2e15f46f3f8556c
institution	Directory Open Access Journal
issn	2079-9292
language	English
last_indexed	2024-03-09T17:50:49Z
publishDate	2022-11-01
publisher	MDPI AG
record_format	Article
series	Electronics
spelling	doaj.art-e11b18c912e94a9fa2e15f46f3f8556c2023-11-24T10:47:43ZengMDPI AGElectronics2079-92922022-11-011123391910.3390/electronics11233919Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource ConditionsCe Zhang0Weilan Wang1Guowei Zhang2Key Laboratory of China’s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, ChinaKey Laboratory of China’s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, ChinaLinkDoc Technology, Beijing 100089, ChinaThe construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.https://www.mdpi.com/2079-9292/11/23/3919historical Tibetan documentscharacter annotationcharacter extractiondata augmentationcharacter recognition
spellingShingle	Ce Zhang Weilan Wang Guowei Zhang Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions Electronics historical Tibetan documents character annotation character extraction data augmentation character recognition
title	Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_full	Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_fullStr	Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_full_unstemmed	Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_short	Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions
title_sort	construction of a character dataset for historical uchen tibetan documents under low resource conditions
topic	historical Tibetan documents character annotation character extraction data augmentation character recognition
url	https://www.mdpi.com/2079-9292/11/23/3919
work_keys_str_mv	AT cezhang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions AT weilanwang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions AT guoweizhang constructionofacharacterdatasetforhistoricaluchentibetandocumentsunderlowresourceconditions

Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions

Similar Items