NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization

This paper describes the design and creation of a monolingual parallel corpus for the Malay language written in Jawi. This paper proposes a new corpus called the National University of Malaysia Word Tokenization (NUWT) corpora To the best of our knowledge, currently, there is no sufficiently compreh...

Full description

Bibliographic Details
Main Authors:	Abu Bakar, Juhaida, Omar, Khairuddin, Nasrudin, Mohammad Faidzul, Murah, Mohd Zamri
Format:	Article
Language:	English
Published:	Universiti Utara Malaysia Press 2016
Subjects:	T Technology (General)
Online Access:	https://repo.uum.edu.my/id/eprint/30364/1/JICT%2015%2001%202016%20107-131.pdf

_version_	1825806169747226624
author	Abu Bakar, Juhaida Omar, Khairuddin Nasrudin, Mohammad Faidzul Murah, Mohd Zamri
author_facet	Abu Bakar, Juhaida Omar, Khairuddin Nasrudin, Mohammad Faidzul Murah, Mohd Zamri
author_sort	Abu Bakar, Juhaida
collection	UUM
description	This paper describes the design and creation of a monolingual parallel corpus for the Malay language written in Jawi. This paper proposes a new corpus called the National University of Malaysia Word Tokenization (NUWT) corpora To the best of our knowledge, currently, there is no sufficiently comprehensive, well-designed standard corpus that is annotated and made available for the public for the Jawi script corpora. This corpus contains the Jawi-specific Buckwalter character code and can be used to evaluate the performance of word tokenization tasks, as well as further language processing. The objective of this work is to conform and standardize the corpora between similar characters in Jawi. It consists of three subcorporas with documents from different genres. The gathering and processing steps, as well as the definition of several evaluation tasks regarding the use of these corpora, are included in this paper. One of the important roles and fundamental tasks of the corpus, which is the tokenization, is also presented in this paper. The development of the Malay language tokenizer is based on the syntactic data compatibility of Malay words written in Jawi. A series of experiments were performed to validate the corpus and to fulfill the requirement of the Jawi script tokenizer with an average error rate of 0.020255. Based on this promising result, the token will be used for the disambiguation and unknown word resolution, such as out-of-vocabulary (OOV) problem in the tagging process.
first_indexed	2024-07-04T06:45:07Z
format	Article
id	uum-30364
institution	Universiti Utara Malaysia
language	English
last_indexed	2024-07-04T06:45:07Z
publishDate	2016
publisher	Universiti Utara Malaysia Press
record_format	eprints
spelling	uum-303642024-02-05T08:16:32Z https://repo.uum.edu.my/id/eprint/30364/ NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization Abu Bakar, Juhaida Omar, Khairuddin Nasrudin, Mohammad Faidzul Murah, Mohd Zamri T Technology (General) This paper describes the design and creation of a monolingual parallel corpus for the Malay language written in Jawi. This paper proposes a new corpus called the National University of Malaysia Word Tokenization (NUWT) corpora To the best of our knowledge, currently, there is no sufficiently comprehensive, well-designed standard corpus that is annotated and made available for the public for the Jawi script corpora. This corpus contains the Jawi-specific Buckwalter character code and can be used to evaluate the performance of word tokenization tasks, as well as further language processing. The objective of this work is to conform and standardize the corpora between similar characters in Jawi. It consists of three subcorporas with documents from different genres. The gathering and processing steps, as well as the definition of several evaluation tasks regarding the use of these corpora, are included in this paper. One of the important roles and fundamental tasks of the corpus, which is the tokenization, is also presented in this paper. The development of the Malay language tokenizer is based on the syntactic data compatibility of Malay words written in Jawi. A series of experiments were performed to validate the corpus and to fulfill the requirement of the Jawi script tokenizer with an average error rate of 0.020255. Based on this promising result, the token will be used for the disambiguation and unknown word resolution, such as out-of-vocabulary (OOV) problem in the tagging process. Universiti Utara Malaysia Press 2016 Article PeerReviewed application/pdf en cc4_by https://repo.uum.edu.my/id/eprint/30364/1/JICT%2015%2001%202016%20107-131.pdf Abu Bakar, Juhaida and Omar, Khairuddin and Nasrudin, Mohammad Faidzul and Murah, Mohd Zamri (2016) NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization. Journal of Information and Communication Technology, 15 (1). pp. 107-131. ISSN 2180-3862 https://e-journal.uum.edu.my/index.php/jict/article/view/8172
spellingShingle	T Technology (General) Abu Bakar, Juhaida Omar, Khairuddin Nasrudin, Mohammad Faidzul Murah, Mohd Zamri NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization
title	NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization
title_full	NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization
title_fullStr	NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization
title_full_unstemmed	NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization
title_short	NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization
title_sort	nuwt jawi specific buckwalter corpus for malay word tokenization
topic	T Technology (General)
url	https://repo.uum.edu.my/id/eprint/30364/1/JICT%2015%2001%202016%20107-131.pdf
work_keys_str_mv	AT abubakarjuhaida nuwtjawispecificbuckwaltercorpusformalaywordtokenization AT omarkhairuddin nuwtjawispecificbuckwaltercorpusformalaywordtokenization AT nasrudinmohammadfaidzul nuwtjawispecificbuckwaltercorpusformalaywordtokenization AT murahmohdzamri nuwtjawispecificbuckwaltercorpusformalaywordtokenization

NUWT: Jawi-Specific Buckwalter Corpus for Malay Word Tokenization

Similar Items