Prikaži enostavni zapis vnosa

 
dc.contributor.author Krek, Simon
dc.contributor.author Erjavec, Tomaž
dc.contributor.author Repar, Andraž
dc.contributor.author Čibej, Jaka
dc.contributor.author Arhar Holdt, Špela
dc.contributor.author Gantar, Polona
dc.contributor.author Kosem, Iztok
dc.contributor.author Robnik-Šikonja, Marko
dc.contributor.author Ljubešić, Nikola
dc.contributor.author Dobrovoljc, Kaja
dc.contributor.author Laskowski, Cyprian
dc.contributor.author Grčar, Miha
dc.contributor.author Holozan, Peter
dc.contributor.author Šuster, Simon
dc.contributor.author Gorjanc, Vojko
dc.contributor.author Stabej, Marko
dc.contributor.author Logar, Nataša
dc.date.accessioned 2026-04-02T14:15:31Z
dc.date.available 2026-04-02T14:15:31Z
dc.date.issued 2023-08-03
dc.identifier.uri http://hdl.handle.net/11356/2055
dc.description Gigafida 2.1 is a reference corpus of written Slovene texts published in the period 1990-2018. It is comprised of daily news, magazines, a selection of web texts (a certain portion of which covers news texts as well), and different types of publications (fiction, school books, and non-fiction). The texts have been selected and automatically processed with the aim of creating a corpus that represents a sample of modern standard Slovene and can be used for research in linguistics and other branches of the humanities, for compiling modern dictionaries, grammars, and learning materials, as well as for developing language technologies for Slovene. Version 2.1 contains the same texts as version 2.0, but includes four additional annotation layers: (1) syntactic dependency annotations based on the Universal Dependencies system (https://universaldependencies.org/); (2) syntactic dependency annotations based on the JOS system; (3) semantic role labelling annotations; (4) named entity annotations. Semantic role labels were assigned with "bilateral-srl" (https://github.com/clarinsi/bilateral-srl), named entities with "The CLASSLA-StanfordNLP model for named entity recognition of standard Slovenian 1.0" (http://hdl.handle.net/11356/1321"), and syntactic dependency annotations (both UD and JOS) with "Parser-V3" - a predecessor of Stanza (https://pypi.org/project/stanza) and CLASSLA (https://pypi.org/project/classla). References: Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem and Kaja Dobrovoljc. Gigafida 2.0: The Reference Corpus of Written Standard Slovene. Proceedings of The 12th Language Resources and Evaluation Conference. Marseille, May 2020. https://www.aclweb.org/anthology/2020.lrec-1.409/ LOGAR BERGINC, Nataša, GRČAR, Miha, BRAKUS, Marko, ERJAVEC, Tomaž, ARHAR HOLDT, Špela and KREK, Simon. Korpusi slovenskega jezika Gigafida, KRES, ccGigafida in ccKRES: gradnja, vsebina, uporaba. Ljubljana: Trojina, zavod za uporabno slovenistiko; Fakulteta za družbene vede, 2012. https://doi.org/10.4312/9789610603542
dc.language.iso slv
dc.publisher Centre for Language Resources and Technologies, University of Ljubljana
dc.relation.isreferencedby https://www.aclweb.org/anthology/2020.lrec-1.409/
dc.relation.isreferencedby https://doi.org/10.5281/zenodo.14165131
dc.relation.isreferencedby https://doi.org/10.4312/9789610603542
dc.relation.replaces http://hdl.handle.net/11356/1320
dc.relation.isreplacedby http://hdl.handle.net/11356/2106
dc.source.uri https://www.cjvt.si/en/research/cjvt-projects/gigafida-corpus/
dc.subject reference corpus
dc.subject representative corpus
dc.subject standard language
dc.subject lemmatisation
dc.subject morphosyntactic tags
dc.title Corpus of written standard Slovene Gigafida 2.1
dc.type corpus
metashare.ResourceInfo#ContentInfo.mediaType text
has.files no
branding CLARIN.SI data & tools
contact.person Tomaž Erjavec tomaz.erjavec@ijs.si Jožef Stefan Institute
contact.person Simon Krek simon.krek@guest.arnes.si Centre for Language Resources and Technologies, University of Ljubljana
sponsor ARRS (Slovenian Research Agency) P6-0411 Language Resources and Technologies for Slovene nationalFunds
sponsor University of Ljubljana I0-0022 Network of Research Infrastructure Centres (MRIC) nationalFunds
sponsor Ministry of Culture C3340-20-278001 Development of Slovene in a Digital Environment Other
size.info 59861870 sentences
size.info 1109441592 words
size.info 38310 texts
files.count 0
files.size 0
featuredService.noske search|https://www.clarin.si/ske/#dashboard?corpname=gfida21


Prikaži enostavni zapis vnosa