Motif-Technologies/Motif-3
Text Generation • 315B • Updated • 50
@loubnabnl Thank you for your response, it’s been extremely helpful.
While reviewing the dataset sources, I came across a few questions and would like to ask for clarification:
1. Regarding the pes2o dataset, I couldn’t find it in the shared pretraining collection. Is it referring to the allenai/peS2o dataset?
2. Based on the config, I noticed three datasets: pull-requests, jupyter-scripts, and github-issues. Could you please clarify the sources for each of these?
3. For the kaggle dataset, did you use the HuggingFaceTB/issues-kaggle-notebooks source? I saw that there are two subsets, issues and kaggle — did you use only the kaggle subset?
Thanks again!