• Repozitorij
  • O repozitoriju
  • Kontakt
  • CLARIN
  •  Prijava
  • English Slovenščina
  • Repozitorij CLARIN.SI
  • Prikaz vnosa
  •  
  • CLARIN logo
  •   Brskanje  
    •    Celoten repozitorij  
      •   Datum izdaje
      •   Avtor
      •   Naslov
      •   Ključne besede
      •   Izdajatelj
      •   Jezik
      •   Vrsta
      •   Oznaka pravic
  •   Moj račun  
    •    Prijava
  •   Statistika  
    •    Statistika PiwikBETA
  •   Splošne informacije  
    •    O vnosu v repozitorij
    •    Citiranje
    •    Življenjski ciklus vnosa
    •    Pogosta vprašanja
    •    O repozitoriju
    •    Pomoč uporabnikom
 
 

Corpus of written standard Slovene Gigafida 2.1

 
CLARIN.SI data & tools
  Avtorji
Krek, Simon ; et al.prikaži vse Krek, Simon ; Erjavec, Tomaž ; Repar, Andraž ; Čibej, Jaka ; Arhar Holdt, Špela ; Gantar, Polona ; Kosem, Iztok ; Robnik-Šikonja, Marko ; Ljubešić, Nikola ; Dobrovoljc, Kaja ; Laskowski, Cyprian ; Grčar, Miha ; Holozan, Peter ; Šuster, Simon ; Gorjanc, Vojko ; Stabej, Marko ; Logar, Nataša
  Identifikator vnosa
http://hdl.handle.net/11356/2055
 URL projekta
https://www.cjvt.si/en/research/cjvt-projects/gigafida-corpus/
 Dokumentirano v
https://www.aclweb.org/anthology/2020.lrec-1.409/
https://doi.org/10.5281/zenodo.14165131
https://doi.org/10.4312/9789610603542
 Datum objave
2023-08-03
 Vrsta
corpus, text
 Velikost
59861870 sentences, 1109441592 words, 38310 texts
 Jezik(i)
Slovenian
 Opis
Gigafida 2.1 is a reference corpus of written Slovene texts published in the period 1990-2018. It is comprised of daily news, magazines, a selection of web texts (a certain portion of which covers news texts as well), and different types of publications (fiction, school books, and non-fiction). The texts have been selected and automatically processed with the aim of creating a corpus that represents a sample of modern standard Slovene and can be used for research in linguistics and other branches of the humanities, for compiling modern dictionaries, grammars, and learning materials, as well as for developing language technologies for Slovene. Version 2.1 contains the same texts as version 2.0, but includes four additional annotation layers: (1) syntactic dependency annotations based on the Universal Dependencies system (https://universaldependencies.org/); (2) syntactic dependency annotations based on the JOS system; (3) semantic role labelling annotations; (4) named entity annotations. Semantic role labels were assigned with "bilateral-srl" (https://github.com/clarinsi/bilateral-srl), named entities with "The CLASSLA-StanfordNLP model for named entity recognition of standard Slovenian 1.0" (http://hdl.handle.net/11356/1321"), and syntactic dependency annotations (both UD and JOS) with "Parser-V3" - a predecessor of Stanza (https://pypi.org/project/stanza) and CLASSLA (https://pypi.org/project/classla). References: Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem and Kaja Dobrovoljc. Gigafida 2.0: The Reference Corpus of Written Standard Slovene. Proceedings of The 12th Language Resources and Evaluation Conference. Marseille, May 2020. https://www.aclweb.org/anthology/2020.lrec-1.409/ LOGAR BERGINC, Nataša, GRČAR, Miha, BRAKUS, Marko, ERJAVEC, Tomaž, ARHAR HOLDT, Špela and KREK, Simon. Korpusi slovenskega jezika Gigafida, KRES, ccGigafida in ccKRES: gradnja, vsebina, uporaba. Ljubljana: Trojina, zavod za uporabno slovenistiko; Fakulteta za družbene vede, 2012. https://doi.org/10.4312/9789610603542
 Izdajatelj
Centre for Language Resources and Technologies, University of Ljubljana
 Zahvala
ARRS (Slovenian Research Agency) P6-0411 "Language Resources and Technologies for Slovene"
University of Ljubljana I0-0022 "Network of Research Infrastructure Centres (MRIC)"
Ministry of Culture C3340-20-278001 "Development of Slovene in a Digital Environment"
 Ključne besede
reference corpus representative corpus standard language lemmatisation morphosyntactic tags
 Zbirke
CLARIN.SI data & tools
 
Ta vnos je bil nadomeščen z novejšim.
http://hdl.handle.net/11356/2106
Prikaži polni zapis vnosa
 
 

Partnerji

  • Alpineon, d.o.o.
  • Amebis, d.o.o.
  • Inštitut za novejšo zgodovino
  • Institut "Jožef Stefan"
  • Narodna in univerzitetna knjižnica Slovenije
  • Slovensko društvo za jezikovne tehnologije

Partnerji

  • Univerza v Ljubljani
  • Univerza v Mariboru
  • Univerza v Novi Gorici
  • Univerza na Primorskem
  • ZRC SAZU
  • ZRS Koper

Repozitorij

  • Domača stran
  • Kontakt
  • Življenjski ciklus vnosa
  • Pogosta vprašanja
  • O repozitoriju in pravilih uporabe

Repozitorij uporablja programsko opremo, ki je bila razvita za LINDAT/CLARIAH-CZ jezikoslovni repozitorij in je dostopna na GitHubu.

CLARIN.SI podpira Ministrstvo za izobraževanje, znanost in šport
v okviru programa "Raziskovalne infrastrukture".