• Repository
  • About
  • Contact
  • CLARIN
  •  Login
  • English Slovenščina
  • CLARIN.SI repository
  • View Item
  •  
  • CLARIN logo
  •   Browse  
    •    All of the Repository  
      •   Issue Date
      •   Authors
      •   Titles
      •   Subjects
      •   Publisher
      •   Language
      •   Type
      •   Rights Label
  •   My Account  
    •    Login
  •   Statistics  
    •    Piwik StatisticsBETA
  •   General Information  
    •    Deposit
    •    Cite
    •    Submission Lifecycle
    •    FAQ
    •    About
    •    Help Desk
 
 

Corpus of written standard Slovene Gigafida 2.2

 
CLARIN.SI data & tools
  Authors
Krek, Simon ; et al.show everyone Krek, Simon ; Erjavec, Tomaž ; Repar, Andraž ; Čibej, Jaka ; Arhar Holdt, Špela ; Gantar, Polona ; Kosem, Iztok ; Robnik-Šikonja, Marko ; Ljubešić, Nikola ; Dobrovoljc, Kaja ; Laskowski, Cyprian ; Grčar, Miha ; Holozan, Peter ; Šuster, Simon ; Gorjanc, Vojko ; Stabej, Marko ; Logar, Nataša ; Terčon, Luka ; Škvorc, Tadej
  Item identifier
http://hdl.handle.net/11356/2106
 Project URL
https://viri.cjvt.si/gigafida/about/kolofon
 Demo URL
https://viri.cjvt.si/gigafida/
 Referenced by
https://www.aclweb.org/anthology/2020.lrec-1.409/
https://doi.org/10.5281/zenodo.14165131
https://doi.org/10.4312/9789610603542
 Date issued
2025-12-08
 Type
corpus, text
 Size
736267 texts, 61907096 sentences, 1133558970 words, 1363860705 tokens
 Language(s)
Slovenian
 Description
Gigafida 2.2 is a reference corpus of written Slovene texts published in the period 1990-2018. It is comprised of daily news, magazines, a selection of web texts (a certain portion of which covers news texts as well), and different types of publications (fiction, school books, and non-fiction). The texts have been selected and automatically processed with the aim of creating a corpus that represents a sample of modern standard Slovene and can be used for research in linguistics and other branches of the humanities, for compiling modern dictionaries, grammars, and learning materials, as well as for developing language technologies for Slovene. The main novelty of version 2.2 is the segmentation of texts from the newspapers Delo and Dnevnik, which represent the largest share of newspaper texts in the corpus. In version 2.1, these texts contained entire daily editions of newspapers with articles on various topics. Using a combination of automatic and manual methods, we segmented the editions into individual articles. In this way, the Gigafida corpus is also better prepared for future upgrades, as newer journalistic texts (e.g., the ones collected for the monitor corpus Trendi) are already being collected in the form of individual articles. A few other improvements have been made, e.g. invalid tags in several files which caused for the texts to be excluded from the corpus when uploaded to the concordances. On the other hand, Gigafida 2.2 does not contain Semantic role labels and Named Entity annotations. References: Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem and Kaja Dobrovoljc. Gigafida 2.0: The Reference Corpus of Written Standard Slovene. Proceedings of The 12th Language Resources and Evaluation Conference. Marseille, May 2020. https://www.aclweb.org/anthology/2020.lrec-1.409/ LOGAR BERGINC, Nataša, GRČAR, Miha, BRAKUS, Marko, ERJAVEC, Tomaž, ARHAR HOLDT, Špela and KREK, Simon. Korpusi slovenskega jezika Gigafida, KRES, ccGigafida in ccKRES: gradnja, vsebina, uporaba. Ljubljana: Trojina, zavod za uporabno slovenistiko; Fakulteta za družbene vede, 2012. https://doi.org/10.4312/9789610603542
 Publisher
Centre for Language Resources and Technologies, University of Ljubljana
 Acknowledgement
ARRS (Slovenian Research Agency) P6-0411 "Language Resources and Technologies for Slovene"
University of Ljubljana I0-0022 "Network of Research Infrastructure Centres (MRIC)"
 Subject(s)
reference corpus representative corpus standard language lemmatisation morphosyntactic tags
 Collection(s)
CLARIN.SI data & tools
 Other versions
Show full item record
 
 

Partners

  • Alpineon, d.o.o.
  • Amebis, d.o.o.
  • Institute of Contemporary History
  • Jožef Stefan Institute
  • National and University Library of Slovenia
  • Slovenian Language Technologies Society

Partners

  • University of Ljubljana
  • University of Maribor
  • University of Nova Gorica
  • University of Primorska
  • ZRC SAZU
  • ZRS Koper

Repository

  • Main page
  • Contact
  • Submission Lifecycle
  • FAQ
  • About and Policies

This platform runs under the software developed for the LINDAT/CLARIAH-CZ repository for linguistics, available on GitHub

CLARIN.SI is supported by the Ministry of Education, Science and Sport of the Republic of Slovenia
under the Programme of "Research Infrastructures".