It depends on:
- the exact repository
- the quality of the metadata
Since it is not trivial at all to ensure that all technical requirements for successful harvesting are fulfilled, we advise to deposit your data with a CLARIN centre. This guarantees a successful harvesting process.
If this is not an option, then it is good to realise that repositories that combine an OAI-PMH endpoint (e.g. Zenodo, Invenio, DataVerse, FigShare, or any repository acquiring DOIs via DataCite) with rich metadata (depends on the repository) are more likely to be harvested in good order.
For an example, see Metadata university repository harvested in VLO: Radboud University PoC
Details about candidates and assessments of the metadata richness can be found here.
General-purpose systems such as Github and Gitlab and machine learning platforms such as Hugging Face and Kaggle in general do not have an OAI-PMH endpoint and therefore are not obvious candidates for harvesting. This does not mean it is not possible, but doing so would require a substantial amount of work and maintenance.
In some specific cases, CLARIN established a collaboration to harvest some larger external data sets: