Published July 12, 2023
| Version 0.1.2
Dataset
Open
A collection of text embeddings of the arXiv corpus by title and abstract
Description
A popular online repository of arXiv is home to numerous preprints in many scientific domains. Other than playing a role of disseminating up-to-date knowledge in pertaining domains, arXiv is an interesting complex system by itself from text analytics point of view. In this repository, we provide a collection of text embedding outputs for (almost) all papers' from the arXiv corpus by their titles and abstracts in order to provide multi-faceted characteristics of scientific knowledge.
Files
Files
(28.6 GB)
| Name | Size | |
|---|---|---|
|
md5:97bf660f4e5f72045c7798b21e275d8e
|
3.6 GB | Download |
|
md5:bece0f637eaece168452be4d81d013f2
|
3.6 GB | Download |
|
md5:ff651a05e4ef1a146d5b8a9126b6c866
|
7.2 GB | Download |
|
md5:7303dbc26997936c85e629ad4f76feee
|
3.6 GB | Download |
|
md5:892e37a45ca6a092a43cfdc5e7e3429b
|
3.6 GB | Download |
|
md5:46736b067fe193015adc2d510f0ccf38
|
7.2 GB | Download |