What's New
corpus
Description:
ELEXIS-WSD is a parallel sense-annotated corpus in which content words (nouns, adjectives, verbs, and adverbs) have been assigned senses. Version 2.0 contains subcorpora with sentences for 17 languages: Bulgarian, Danish, ...
This item contains 1 file (14.08
MB).
Publicly Available
corpus
Description:
This entry contains the first part of the audiobook "Pramatija ali Bučman" (Pramatija, or the Bogeyman) by author Leopold Suhodolčan (COBISS ID: 264527107, ISBN: 978-961-7194-44-9).
Television spotlights were shining ...
This item contains 9 files (81.11
MB).
Publicly Available
corpus
Description:
This entry contains the first part of the audiobook "Rdeči lev" (The red lion) by author Leopold Suhodolčan (COBISS ID: 264850179, ISBN: 978-961-7194-48-7).
Blaž was faster and soon managed to escape them, but they still ...
This item contains 6 files (92.75
MB).
Publicly Available
Most Viewed Items
Top Last Week
corpus
Description:
ParlaMint 5.0 is a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions, mostly starting in 2015 and extending to mid-2022. The individual corpora ...
This item contains 31 files (5.94
GB).
Publicly Available
corpus
Description:
The dataset represents the Twitter production in Slovenian in the period from 2018 until 2020. It consists of tweet IDs, retweet IDs, pseudo-anonymized user IDs, publication dates, and automatically assigned hate labels ...
This item contains 1 file (182.04
MB).
Publicly Available
corpus
Description:
ParlaMint 2.1 is a multilingual set of 17 comparable corpora containing parliamentary debates mostly starting in 2015 and extending to mid-2020, with each corpus being about 20 million words in size. The sessions in the ...
This item contains 18 files (2.17
GB).
Publicly Available