• Repository
  • About
  • Contact
  • CLARIN
  •  Login
  • English Slovenščina
  • CLARIN.SI repository
  • View Item
  •  
  • CLARIN logo
  •   Browse  
    •    All of the Repository  
      •   Issue Date
      •   Authors
      •   Titles
      •   Subjects
      •   Publisher
      •   Language
      •   Type
      •   Rights Label
  •   My Account  
    •    Login
  •   Statistics  
    •    Piwik StatisticsBETA
  •   General Information  
    •    Deposit
    •    Cite
    •    Submission Lifecycle
    •    FAQ
    •    About
    •    Help Desk
 
 

Monitor corpus of Slovene Trendi 2022-05

 
CLARIN.SI data & tools
  Authors
Kosem, Iztok ; et al.show everyone Kosem, Iztok ; Čibej, Jaka ; Dobrovoljc, Kaja ; Erjavec, Tomaž ; Ljubešić, Nikola ; Ponikvar, Primož ; Šinkec, Mihael ; Krek, Simon
  Item identifier
http://hdl.handle.net/11356/1590
 Project URL
https://sled.ijs.si/
 Date issued
2022-06-23
 Type
corpus, text
 Size
565308991 tokens, 473161579 words, 25186942 sentences, 1436548 articles
 Language(s)
Slovenian
 Description
The Trendi corpus is a monitor corpus of Slovene. It contains news from 107 different media websites, published by 48 different publishers. Trendi 2022-05 covers the period from January 2019 to May 2022, complementing the Gigafida 2.0 reference corpus of written Slovene. All the contents of the Trendi corpus are at the moment obtained using the Jožef Stefan Institute Newsfeed service (http://newsfeed.ijs.si/). The texts have been annotated using the classla-stanza pipeline (https://github.com/clarinsi/classla), including syntactic parsing according to the Universal Dependencies (https://universaldependencies.org/sl/) and Named Entities (https://nl.ijs.si/janes/wp-content/uploads/2017/09/SlovenianNER-eng-v1.1.pdf). At the moment, the corpus is not available as a dataset due to copyright restrictions, we hope to make at least some of it available in the near future. The corpus is accessible through CLARIN.SI concordancers.
 Publisher
Jožef Stefan Institute
 Acknowledgement
Ministry of Culture of the Republic of Slovenia JR-infrastruktura-SJ-2021-2022 "SLED - Monitor corpus of Slovene and related resources"
 Subject(s)
monitor corpus news corpus universal dependencies temporal trends
 Collection(s)
CLARIN.SI data & tools
 
This item is replaced by a newer submission:
http://hdl.handle.net/11356/1681
Show full item record
 
 

Partners

  • Alpineon, d.o.o.
  • Amebis, d.o.o.
  • Institute of Contemporary History
  • Jožef Stefan Institute
  • National and University Library of Slovenia
  • Slovenian Language Technologies Society

Partners

  • University of Ljubljana
  • University of Maribor
  • University of Nova Gorica
  • University of Primorska
  • ZRC SAZU
  • ZRS Koper

Repository

  • Main page
  • Contact
  • Submission Lifecycle
  • FAQ
  • About and Policies

This platform runs under the software developed for the LINDAT/CLARIAH-CZ repository for linguistics, available on GitHub

CLARIN.SI is supported by the Ministry of Education, Science and Sport of the Republic of Slovenia
under the Programme of "Research Infrastructures".