What's New
corpus
Description:
The Trendi corpus is a monitor corpus of Slovenian. It contains news articles from 106 media websites, published by 64 publishers. Trendi 2026-07 covers the period from January 2019 to August 2026, complementing the Gigafida ...
Ta vnos ne vsebuje datotek.
corpus
Description:
AspectBench 1.0 (HBS Subset) is an HBS language dataset (hbs, ISO 639-3) containing online news articles, mainly in Serbian and Croatian, for studying sentiment towards target entities like companies and brand names. It ...
Ta vnos vsebuje 1 datoteko (190.74
MB).
Restricted Use
corpus
Description:
This entry contains the first part of the audiobook "Pobožal sem polje na desni" (I caressed the field on the right) by author Vili Telassani and Barbara Cerar (COBISS ID: 279985667, ISBN: 978-961-272-884-7).
Vili is ...
Ta vnos vsebuje 3 datotek(e) (128.23
MB).
Publicly Available
Največ ogledov
V preteklem tednu
corpus
Description:
Artur 1.0 is a speech database designed for the needs of automatic speech recognition for the Slovenian language. The database includes 1,067 hours of speech. 884 hours are transcribed, while the remaining 183 hours are ...
Ta vnos vsebuje 39 datotek(e) (324.53
GB).
Publicly Available
corpus
Description:
ParlaMint 5.0 is a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions, mostly starting in 2015 and extending to mid-2022. The individual corpora ...
Ta vnos vsebuje 31 datotek(e) (5.94
GB).
Publicly Available
lexicalConceptualResource
Description:
A lexicon of 751 emoji characters with automatically assigned sentiment.
The sentiment is computed from 70,000 tweets, labeled by 83 human annotators
in 13 European languages.
The process and analysis of emoji sentiment ...
Ta vnos vsebuje 3 datotek(e) (93.95
KB).
Publicly Available