OmniCorpus-CW Dataset
The OmniCorpus dataset is a large-scale image-text interleaved dataset that breaks the boundaries of scale and diversity by containing 8.6 billion images interleaved with 1,696 text labels from different sources, greatly surpassing previous datasets.