Description
Datasets to reproduce the experiments associated with the paper: https://doi.org/10.48550/arXiv.2307.06440 The readme contains instructions for how to use them: https://github.com/JeanKaddour/NoTrainNoGain/blob/main/bert/README.md c4-subset-random.tar.bz2 is a subset of the C4 dataset (https://arxiv.org/abs/1910.10683), licensed under ODC-BY 1.0.
Data Citation
Kaddour, J., Key, O., Nawrot, P., Minervini, P., & Kusner, M. J. (2023). Data for "No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models" (Version 1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.8279728
| Date made available | 25 Jul 2023 |
|---|---|
| Publisher | Zenodo |
Cite this
- DataSetCite