The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.1

Name: The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.1
License: https://creativecommons.org/licenses/by-sa/4.0/

Ljubešić, Nikola

The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Slovenian 1.1

CLARIN.SI data & tools

Authors: Ljubešić, Nikola

Item identifier: http://hdl.handle.net/11356/1312

Project URL: https://github.com/clarinsi/classla-stanfordnlp

Referenced by: https://www.aclweb.org/anthology/W19-3704/

Date issued: 2020-04-29

Type: toolService

Language(s): Slovenian

Description: This model for morphosyntactic annotation of standard Slovenian was built with the CLASSLA-StanfordNLP tool (https://github.com/clarinsi/classla-stanfordnlp) by training on the ssj500k training corpus (http://hdl.handle.net/11356/1210) and using the CLARIN.SI-embed.sl word embeddings (http://hdl.handle.net/11356/1204). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~97.06. The difference to the previous version of the model is that now the whole XPOS tag is predicted and not specific characters, as was the case in stanfordnlp, which resulted in illegal XPOS tags (and slightly decreased performance).