# Can you harvest my data from Zenodo, DataVerse, FigShare, Github, Hugging Face?

**URL:** <https://forum.clarin.eu/t/can-you-harvest-my-data-from-zenodo-dataverse-figshare-github-hugging-face/1490>\
**Category:** FAQ\
**Tags:** harvesting\
**Created:** [23 July 2026 09:50 UTC](https://forum.clarin.eu/t/can-you-harvest-my-data-from-zenodo-dataverse-figshare-github-hugging-face/1490 "2026-07-23T09:50:44Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![dieter](https://avatars.discourse-cdn.com/v4/letter/d/977dab/32.png) [@dieter](https://forum.clarin.eu/u/dieter)\
**Post date:** [23 July 2026 09:50 UTC](https://forum.clarin.eu/t/can-you-harvest-my-data-from-zenodo-dataverse-figshare-github-hugging-face/1490/1 "2026-07-23T09:50:44Z")

</div>

It depends on:

- the exact repository
- the quality of the metadata

Since it is not trivial at all to ensure that all technical requirements for successful harvesting are fulfilled, we [advise](https://forum.clarin.eu/t/why-should-i-deposit-my-data-in-a-clarin-repository-and-not-somewhere-else/1489) to deposit your data [with a CLARIN centre](https://www.clarin.eu/content/depositing-services). This guarantees a successful harvesting process.

If this is not an option, then it is good to realise that repositories that combine an **OAI-PMH endpoint** (e.g. Zenodo, Invenio, DataVerse, FigShare, or any repository acquiring DOIs via DataCite) with **rich metadata** (depends on the repository) are more likely to be harvested in good order.

For an example, see [Metadata university repository harvested in VLO: Radboud University PoC](https://forum.clarin.eu/t/metadata-university-repository-harvested-in-vlo-radboud-university-poc/706)

Details about candidates and assessments of the metadata richness [can be found here](https://github.com/clarin-eric/oai-harvest-config/issues?q=is%3Aissue%20state%3Aopen%20label%3A%22harvest%20candidate%22).

General-purpose systems such as Github and Gitlab and machine learning platforms such as Hugging Face and Kaggle in general do not have an OAI-PMH endpoint and therefore are not obvious candidates for harvesting. This does not mean it is not possible, but doing so would require a substantial amount of work and maintenance.

In some specific cases, CLARIN established a collaboration to harvest some larger external data sets:

- [Mozilla Data Collective datasets can now be explored through CLARIN’s Virtual Language Observatory (VLO)](https://forum.clarin.eu/t/mozilla-data-collective-datasets-can-now-be-explored-through-clarins-virtual-language-observatory-vlo/1476)
- [Exploring New Europeana Resources in CLARIN’s Virtual Language Observatory](https://www.clarin.eu/blog/exploring-new-europeana-resources-clarins-virtual-language-observatory)
