Open Science Framework scaling down and alternatives via CLARIN

As our TROLLing colleagues have put it in their article:

Yesterday, the Center for Open Science announced that OSF Projects will be phased out starting 16 November 2026. If you’re a linguist who has been using OSF to store and share research data, code, and other materials, now is a good time to consider where those materials should live in the long term.

More details are available at the OSF website.

CLARIN repositories

While it is sad to see nice open science supporting services such as OSF Projects scale down so drastically, this is also an interesting wake-up call about the importance of long-term organisational commitment to data repositories. That is where the CLARIN centres come in: a very important prerequisite for establishing a B-centre repository is to have a long-term infrastructural horizon, including the existence of a national consortium.

While this is not a panacea for sustainable funding, together with the broader context of the integration of national infrastructures into the long-term ESFRI roadmap(s) it definitely helps towards achieving long-term data preservation and related services.

Generalist vs. specialised repositories

Both have their own pros and cons, but to mention one important difference: with the specialised ones, e.g. the CLARIN ones focused on language resources, there is more detailed metadata available, leading e.g. to guaranteed inclusion into subject-specific catalogues such as the Virtual Language Observatory.

Supporting collaborative data workflows

Two particular aspects in which OSF Projects stands out is the support for collaborative workflows and the integration with third-party service providers. Personally, I think this is something that CLARIN repositories can be inspired by.

Offering the same functionality as OSF is probably not realistic, but some basic features which have proven popular with other repositories, such as importing data from github or gitlab repositories (as Zenodo/InvenioRDM and DataVerse support) might help to attract extra data deposits.

Do you think CLARIN should do something in this area?

One idea that recently popped up is that CLARIN should help developing a CLARIN-DSpace plugin to import datasets from git repositories (such as github and gitlab). Would this be a good idea (more deposits) or not worth the effort (rubbish deposits)?