šŸ“œ Release of LexFCS 1.0 Specification

I’m happy and proud to announce that the German Text+ consortium has now released the 1.0 version of the LexFCS specification. This version marks the first stable release after ~ 5 years of development and testing, with lots of user feedback and discussions. Previous versions have been presented in various formats at different events over the years.

LexFCS 1.0 is now ready for implementation and we are excited to see many new lexical resources added to the FCS. Various CLARIN Java libraries used for SRU and FCS have already been updated, the reference implementation for LexFCS endpoints, the endpoint validator, and the integration in the CLARIN Content Search will follow shortly.

What is the LexFCS?

The Federated Content Search for Lexical Resources (LexFCS) specification is an extension of the CLARIN Federated Content Search (CLARIN-FCS) - Core 2 specification that allows search and retrieval of lexical resources including dictionaries, encyclopedias, normative data, terminological databases, ontologies etc. The LexFCS specification adds the following to the FCS:

  • a new query language, LexCQL, based on the Contextual Query Language (CQL) to allow querying lexical resources,
  • a new Lexical Data View to better represent lexical results,
  • extends the CLARIN-FCS specification with mechanisms to describe LexFCS resources, queryable fields and other functionality.

Lexical resources differ from corpora in the FCS in that they are structured differently and therefore require different approaches to querying and result presentation.

Currently, LexFCS supports the following fields: lemma (mandatory, default search field); entryId, phonetic, transcription, translation; definition, etymology; baseform, case, degree, frequency, gender, grammar, mood, number, pos, segmentation, sentiment, tense; antonym, holonym, hypernym, hyponym, meronym, synonym, related; ref, senseRef; citation. They can be used with LexCQL to search through records with the addition of lang (entry language) and any (search on any lexical field). Search terms can be matched in a tokenized (=) or untokenized (==, <>) manner, or use entity references (is). Complex queries can be constructed with boolean AND and OR and parentheses. More details and various examples are provided in the specification.

See LexFCS in action

LexFCS was proposed and developed in Text+ / Germany. There are already a few FCS endpoints implementing LexFCS and offering lexical resources. The majority of resources is still in German (for now).

Integration of version 1.0 in the CLARIN Content Search will follow shortly.

Future Plans

A number of exciting features are already in the pipeline, including support for and integration of multimedia content. These and other ideas will be optional and stay compatible with any major version.

Resources about LexFCS

1 Like