# FAIR Metadata for Specialised Corpora - How can CLARIN support the communities?

**URL:** <https://forum.clarin.eu/t/fair-metadata-for-specialised-corpora-how-can-clarin-support-the-communities/1025>\
**Category:** General\
**Tags:** metadata, cmdi, vlo\
**Created:** [7 October 2025 11:59 UTC](https://forum.clarin.eu/t/fair-metadata-for-specialised-corpora-how-can-clarin-support-the-communities/1025 "2025-10-07T11:59:00Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![egon](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.clarin.eu/egon/32/17_2.png) [@egon](https://forum.clarin.eu/u/egon)\
**Post date:** [7 October 2025 11:59 UTC](https://forum.clarin.eu/t/fair-metadata-for-specialised-corpora-how-can-clarin-support-the-communities/1025/1 "2025-10-07T11:59:00Z")

</div>

Returning to our presentation at the Annual Conference 2025: “Towards FAIR Metadata for Specialised Corpora: A Community-Informed Empirical Study of Schema Development in Two Communities” ([slides](https://www.clarin.eu/media/9180)), I wanted to link relevant Topics (from this forum):

- [What metadata scheme is used within CLARIN?](https://forum.clarin.eu/t/what-metadata-scheme-is-used-within-clarin/454)
- [If there is no single metadata scheme, how should I describe my resources in order for them to be compatible with the CLARIN infrastructure?](https://forum.clarin.eu/t/if-there-is-no-single-metadata-scheme-how-should-i-describe-my-resources-in-order-for-them-to-be-compatible-with-the-clarin-infrastructure/453)
- [What are the guidelines for creating good quality metadata?](https://forum.clarin.eu/t/what-are-the-guidelines-for-creating-good-quality-metadata/477)
- [How do I create a new CMDI metadata file?](https://forum.clarin.eu/t/how-do-i-create-a-new-cmdi-metadata-file/725)
- [What software is available for authoring CMDI metadata?](https://forum.clarin.eu/t/what-software-is-available-for-authoring-cmdi-metadata/476)

but also emphasise that the usability of the annotated data is unclear - community-specific metadata, with its domain-specific complexity:

- should it be integrated into existing infrastructures like the VLO?
- (how) can it be integrated into/connected to the Resource Families?
- how can it be made accessible (displayed/searched/browsed/…) to end users without reducing it to a common denominator?

and since this will (potentially) become relevant for a few more communities (~= K-centres, and likely C-Centres that would host the metadata - or B-centres) the question still remains:

**How can CLARIN (ERIC / B,C,K-centres) support each other and the community efforts [for community-specific metadata]?**

---

<div class="post-metadata">

**Author:** ![matthies](https://dub1.discourse-cdn.com/flex013/user_avatar/forum.clarin.eu/matthies/32/421_2.png) [@matthies](https://forum.clarin.eu/u/matthies)\
**Post date:** [24 October 2025 07:35 UTC](https://forum.clarin.eu/t/fair-metadata-for-specialised-corpora-how-can-clarin-support-the-communities/1025/2 "2025-10-24T07:35:44Z")

</div>

Dear Egon, to my knowledge there is indeed “only” CMDI as commonly agreed metastandard, but profiles (essentially XML schemas) vary. In Finland we use Profiles derived from the META-SHARE Schema in COMEDI. COMEDI currently exports them as [clarin.eu:cr1:p\_1361876010571](https://clarino.uib.no/oai?verb=ListMetadataFormats). So for CLARIN you should use CMDI. But the world is bigger than CLARIN. I understood your talk as touching on the questions: If we have metadata in different granularity, how to we make sure that we are dealing with variants of the same thing? Example: (PID-1 is here the placeholder for a Handle used in CLARIN)

- Very detailed metadata of dataset with “PID-1”: Contains names of subjects, etc. Sensitive.
- Pseudomymized version of dataset with “PID-1” above. Less sensitive.
- CMDI of dataset with PID-1: Describes the dataset, where and when and how it was created. Public, shown in VLO.
- EOSC compatible HTML metadata of PID-1 in VLO landing page. Subset of CMDI. Increases FAIR score of fair tools
- Subset of dataset with PID-1 exported to national service, like [etsin.fairdata.fi](http://etsin.fairdata.fi). (The Language Bank data is available there for search)
- Reference to dataset with PID in article. Has Author, Year, Name, repository, PID-1 (very small subset of the metadata)

My suggestion would now the following: PID-1 points to the CMDI descriptive metadata at the repository’s Metadata service (COMEDI in our case). This is the “master metadata”, subsets and supersets must be in sync with the data provided there. So if the superset of very detailed metadata mentions the Name of the dataset and it is not identical to CMDI, CMDI is the authoritative source.

Supersets should therefore not copy too much of the authoritative metadata, since it can be always found behind the PID.

The same holds true for subsets, like reference instructions. All sub and super sets need to contain the dataset PID (“PID-1”) as clear link between them. The CMDI metadata points to the data (via resource proxy). Also these pointers are authoritative.

If we can agree on this principle we can think of how to implement it. Descriptive Metadata does not change extremely often, but it does change, like due to incorrect creation which is detected later, etc. So mechanisms should be in place to deal with such changes.

What do you think?
