DT-PICS: An Efficient and Cost-Effective SNP Selection Method for the Germplasm Identification of Arabidopsis

Germplasm identification is essential for plant breeding and conservation. In this study, we developed a new method, DT-PICS, for efficient and cost-effective SNP selection in germplasm identification. The method, based on the decision tree concept, could efficiently select the most informative SNPs...

Full description

Bibliographic Details
Main Authors: Liwen Xiong, Zirong Li, Weihua Li, Lanzhi Li
Format: Article
Language:English
Published: MDPI AG 2023-05-01
Series:International Journal of Molecular Sciences
Subjects:
Online Access:https://www.mdpi.com/1422-0067/24/10/8742
_version_ 1797599833075220480
author Liwen Xiong
Zirong Li
Weihua Li
Lanzhi Li
author_facet Liwen Xiong
Zirong Li
Weihua Li
Lanzhi Li
author_sort Liwen Xiong
collection DOAJ
description Germplasm identification is essential for plant breeding and conservation. In this study, we developed a new method, DT-PICS, for efficient and cost-effective SNP selection in germplasm identification. The method, based on the decision tree concept, could efficiently select the most informative SNPs for germplasm identification by recursively partitioning the dataset based on their overall high PIC values, instead of considering individual SNP features. This method reduces redundancy in SNP selection and enhances the efficiency and automation of the selection process. DT-PICS demonstrated significant advantages in both the training and testing datasets and exhibited good performance on independent prediction, which validates its effectiveness. Thirteen simplified SNP sets were extracted from 749,636 SNPs in 1135 Arabidopsis varieties resequencing datasets, including a total of 769 DT-PICS SNPs, with an average of 59 SNPs per set. Each simplified SNP set could distinguish between the 1135 Arabidopsis varieties. Simulations demonstrated that using a combination of two simplified SNP sets for identification can effectively increase the fault tolerance in independent validation. In the testing dataset, two potentially mislabeled varieties (ICE169 and Star-8) were identified. For 68 same-named varieties, the identification process achieved 94.97% accuracy and only 30 shared markers on average; for 12 different-named varieties, the germplasm to be tested could be effectively distinguished from 1,134 other varieties while grouping extremely similar varieties (Col-0) together, reflecting their actual genetic relatedness. The results suggest that the DT-PICS provides an efficient and accurate approach to SNP selection in germplasm identification and management, offering strong support for future plant breeding and conservation efforts.
first_indexed 2024-03-11T03:39:56Z
format Article
id doaj.art-c01d4e68052343d38d9ad8a8226891fe
institution Directory Open Access Journal
issn 1661-6596
1422-0067
language English
last_indexed 2024-03-11T03:39:56Z
publishDate 2023-05-01
publisher MDPI AG
record_format Article
series International Journal of Molecular Sciences
spelling doaj.art-c01d4e68052343d38d9ad8a8226891fe2023-11-18T01:41:06ZengMDPI AGInternational Journal of Molecular Sciences1661-65961422-00672023-05-012410874210.3390/ijms24108742DT-PICS: An Efficient and Cost-Effective SNP Selection Method for the Germplasm Identification of ArabidopsisLiwen Xiong0Zirong Li1Weihua Li2Lanzhi Li3Hunan Engineering & Technology Research Center for Agricultural Big Data Analysis & Decision-Making, College of Plant Protection, Hunan Agricultural University, Changsha 410128, ChinaHunan Engineering & Technology Research Center for Agricultural Big Data Analysis & Decision-Making, College of Plant Protection, Hunan Agricultural University, Changsha 410128, ChinaHunan Engineering & Technology Research Center for Agricultural Big Data Analysis & Decision-Making, College of Plant Protection, Hunan Agricultural University, Changsha 410128, ChinaHunan Engineering & Technology Research Center for Agricultural Big Data Analysis & Decision-Making, College of Plant Protection, Hunan Agricultural University, Changsha 410128, ChinaGermplasm identification is essential for plant breeding and conservation. In this study, we developed a new method, DT-PICS, for efficient and cost-effective SNP selection in germplasm identification. The method, based on the decision tree concept, could efficiently select the most informative SNPs for germplasm identification by recursively partitioning the dataset based on their overall high PIC values, instead of considering individual SNP features. This method reduces redundancy in SNP selection and enhances the efficiency and automation of the selection process. DT-PICS demonstrated significant advantages in both the training and testing datasets and exhibited good performance on independent prediction, which validates its effectiveness. Thirteen simplified SNP sets were extracted from 749,636 SNPs in 1135 Arabidopsis varieties resequencing datasets, including a total of 769 DT-PICS SNPs, with an average of 59 SNPs per set. Each simplified SNP set could distinguish between the 1135 Arabidopsis varieties. Simulations demonstrated that using a combination of two simplified SNP sets for identification can effectively increase the fault tolerance in independent validation. In the testing dataset, two potentially mislabeled varieties (ICE169 and Star-8) were identified. For 68 same-named varieties, the identification process achieved 94.97% accuracy and only 30 shared markers on average; for 12 different-named varieties, the germplasm to be tested could be effectively distinguished from 1,134 other varieties while grouping extremely similar varieties (Col-0) together, reflecting their actual genetic relatedness. The results suggest that the DT-PICS provides an efficient and accurate approach to SNP selection in germplasm identification and management, offering strong support for future plant breeding and conservation efforts.https://www.mdpi.com/1422-0067/24/10/8742germplasm identificationArabidopsis thalianaSNPDNA fingerprinting
spellingShingle Liwen Xiong
Zirong Li
Weihua Li
Lanzhi Li
DT-PICS: An Efficient and Cost-Effective SNP Selection Method for the Germplasm Identification of Arabidopsis
International Journal of Molecular Sciences
germplasm identification
Arabidopsis thaliana
SNP
DNA fingerprinting
title DT-PICS: An Efficient and Cost-Effective SNP Selection Method for the Germplasm Identification of Arabidopsis
title_full DT-PICS: An Efficient and Cost-Effective SNP Selection Method for the Germplasm Identification of Arabidopsis
title_fullStr DT-PICS: An Efficient and Cost-Effective SNP Selection Method for the Germplasm Identification of Arabidopsis
title_full_unstemmed DT-PICS: An Efficient and Cost-Effective SNP Selection Method for the Germplasm Identification of Arabidopsis
title_short DT-PICS: An Efficient and Cost-Effective SNP Selection Method for the Germplasm Identification of Arabidopsis
title_sort dt pics an efficient and cost effective snp selection method for the germplasm identification of arabidopsis
topic germplasm identification
Arabidopsis thaliana
SNP
DNA fingerprinting
url https://www.mdpi.com/1422-0067/24/10/8742
work_keys_str_mv AT liwenxiong dtpicsanefficientandcosteffectivesnpselectionmethodforthegermplasmidentificationofarabidopsis
AT zirongli dtpicsanefficientandcosteffectivesnpselectionmethodforthegermplasmidentificationofarabidopsis
AT weihuali dtpicsanefficientandcosteffectivesnpselectionmethodforthegermplasmidentificationofarabidopsis
AT lanzhili dtpicsanefficientandcosteffectivesnpselectionmethodforthegermplasmidentificationofarabidopsis