This dataset is based on raw sequence data, which has been double compared with GenBank and local databases, and further analyzed using genetic distance analysis (such as Barcoding Gap and sequence similarity comparison) and phylogenetic tree analysis to systematically eliminate low-quality and redundant sequences. This processing flow significantly improves the accuracy of species identification and the overall reliability of the database, ultimately constructing a high-quality DNA barcode reference database for seed plants in arid regions of China. We have also updated the v2 version of the database, which contains sequence data that has been synchronized and made public in GenBank, in order to provide a more solid, continuously updated, and publicly accessible data foundation for regional biodiversity monitoring and protection.
This dataset focuses on the semi-arid desert in central and eastern Inner Mongolia as the core survey area, and investigates the main desert plants in the four major sandy areas in northern China and surrounding areas from 2017 to 2021, and collects DNA molecular materials. According to project requirements, sequence 5 DNA barcode fragments (rbcL, matK, psbA trnH, trnL-F, ITS) from each plant material. This data set contains 352 copies and 1760 pieces of DNA barcodes, including 56 copies and 280 pieces of plant DNA barcodes in Hulunbeier Sandy Land, 73 copies and 365 pieces of plant DNA barcodes in Horqin Sandy Land, 133 copies and 665 pieces of plant DNA barcodes in Hunshandake Sandy Land, and 90 copies and 450 pieces of plant DNA barcodes in Maowusu Sandy Land.
| collect time | 2017/01/01 - 2021/12/31 |
|---|---|
| collect place | Semiarid sandy land in central and eastern Inner Mongolia |
| data size | 79.7 KiB |
| data format | Excel |
Autonomously generated
Field collection, experimental testing, and digital processing.
1. Collection and production of plant specimens and acquisition of DNA material
Plant specimens were collected in the field with collection number plates bolted on, and the coordinates of the collection site and habitat were recorded in the field. In order to ensure the consistency between the voucher specimens and DNA materials, healthy leaves without obvious diseases or insect pests were collected directly from the voucher specimens in the process of pressing the specimens, and placed in small paper bags with very thin paper, and the numbers on the paper bags were consistent with the specimen collection numbers. After screening and removing decayed specimens, a total of 1,300 complete specimens were obtained, which were pressed dry and then put on the counterpane. the DNA material was permanently preserved in the refrigerator at -80℃.
2. Image data acquisition of plant specimens
The specimen images were captured with a high-resolution camera, and the image brightness was adjusted and sharpened to form the image data of the plant specimens.
3. DNA barcode generation
Plant DNA is extracted by magnetic bead method or CTAB method, and after extraction, a small amount of sample is taken to detect the quality and integrity of DNA by agarose gel electrophoresis. PCR was used to amplify the 5 DNA barcodes required by the project. The amplified products were handed over to a professional sequencing company for sequencing, and in order to ensure that the correct sequence was obtained, they were sequenced through from both ends respectively. Sequencing was performed using an ABI series automated sequencer (Applied Biosystems, Foster City, California, USA). Sequencing primers used PCR amplification primers.
After obtaining the peak maps returned by the sequencing company, the DNA sequences were manually aligned using Geneious or Chromas software to remove low-quality base sequences at both ends, and the forward and reverse sequences were spliced and edited to obtain the DNA sequences and saved as FASTA format files. To verify the accuracy of sequencing or whether the material was contaminated, the spliced DNA barcode sequences of each sample were checked through the BLASTn function on the NCBI website. If a sample's barcode sequence (query sequence) matches the highest score value (subject sequence) for the same family, genus or species, it is determined that the sample in the material collection and experimental process did not make mistakes, its barcode sequence can be initially determined as correct.
4. Generation of DNA barcode information table
Sampling location, geographic coordinates, elevation and habitat information is compiled, and at the same time, the sample number is connected with the image data of the plant specimen, and the sample collection number is connected with the physical specimen, which ultimately forms the DNA barcode description information table.
| # | number | name | type |
| 1 | 2017FY100200 | Survey of main plant communities in Chinese desert | Basic Resource Survey Project |
This work is licensed under
CC BY 4.0 (Creative Commons Attribution 4.0 International License).
| # | title | file size |
|---|---|---|
| 1 | 内蒙古中东部半干旱荒漠植物DNA条形码专题数据库(2017-2021).docx | 30.8 KiB |
| 2 | 内蒙古中东部半干旱荒漠植物DNA条形码数据集.xlsx | 48.9 KiB |
-
-
©Copyright 2005-. Northwest Institute of Eco-Environment and Resources, CAS.
Donggang West Road 320, Lanzhou, Gansu, China (730000)

