WanJuan 2.0 Dataset
WanJuan2.0 (WanJuan-CC) is a high-quality English web text dataset of 1T tokens obtained from CommonCrawl. The results show that, compared to various open-source English CC corpora, WanJuan2.0 exhibits higher safety in different dimensions of the Perspective API evaluation. Additionally, its practicality is demonstrated through perplexity (PPL) on 4 validation sets and accuracy on 6 downstream tasks. WanJuan2.0 shows competitive PPL performance across various validation sets, especially on datasets like tiny-storys that require higher language fluency. By comparing with similar datasets in 1B model training, using perplexity of the validation dataset and accuracy of downstream tasks as evaluation metrics, experiments prove that WanJuan2.0 significantly enhances the performance of English text completion and general English capability tasks.