dc.contributor.author | Terčon, Luka |
dc.contributor.author | Ljubešić, Nikola |
dc.date.accessioned | 2023-05-16T06:50:00Z |
dc.date.available | 2023-05-16T06:50:00Z |
dc.date.issued | 2023-05-10 |
dc.identifier.uri | http://hdl.handle.net/11356/1832 |
dc.description | The model for morphosyntactic annotation of standard Croatian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the hr500k training corpus (http://hdl.handle.net/11356/1792) and using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1790). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~94.87. The difference to the previous version of the model is that this version was trained using the new version of the hr500k corpus and the new version of the Croatian word embeddings. |
dc.language.iso | hrv |
dc.publisher | Jožef Stefan Institute |
dc.relation.isreferencedby | http://dx.doi.org/10.18653/v1/W19-3704 |
dc.relation.replaces | http://hdl.handle.net/11356/1348 |
dc.rights | Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) |
dc.rights.uri | https://creativecommons.org/licenses/by-sa/4.0/ |
dc.rights.label | PUB |
dc.source.uri | https://github.com/clarinsi/classla |
dc.subject | language model |
dc.subject | part-of-speech tagging |
dc.title | The CLASSLA-Stanza model for morphosyntactic annotation of standard Croatian 2.1 |
dc.type | toolService |
metashare.ResourceInfo#ContentInfo.detailedType | tool |
metashare.ResourceInfo#ResourceComponentType#ToolServiceInfo.languageDependent | true |
has.files | yes |
branding | CLARIN.SI data & tools |
contact.person | Nikola Ljubešić nikola.ljubesic@ijs.si Jožef Stefan Institute |
contact.person | Luka Terčon luka.tercon@gmail.com Faculty of Computer and Information Science, University of Ljubljana |
sponsor | ARRS (Slovenian Research Agency) P6-0411 Language Resources and Technologies for Slovene nationalFunds |
sponsor | Jožef Stefan Institute CLARIN CLARIN.SI nationalFunds |
sponsor | ARRS (Slovenian Research Agency) J7-4642 MEZZANINE nationalFunds |
sponsor | Connecting Europe Facility (CEF) Telecom INEA/CEF/ICT/A2020/2278341 MaCoCu - Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages Other |
files.count | 2 |
files.size | 185620047 |
Datoteke v tem vnosu
Prenesi vse datoteke v vnosu (177.02 MB)To je vnos
Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
Publicly Available
z licenco:Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)




- Ime
- baseline_pos.zip
- Velikost
- 71.68 MB
- Format
- application/zip
- Opis
- Language model
- MD5
- 0dbefd95b263c269a5e01215f19112c9

- Ime
- hr_set.pretrain.zip
- Velikost
- 105.34 MB
- Format
- application/zip
- Opis
- Pretrained word embeddings
- MD5
- 49951c88b928436128dad24d68168f88