What's New
corpus
Description:
This entry contains the Ilustrirani Slovenec Multimodal Document Understanding Dataset, a benchmark dataset designed for training and evaluating vision-language models on historical Slovenian newspaper pages. The dataset ...
This item contains 1 file (7.57
MB).
Publicly Available
corpus
Description:
This entry contains the first part of the e-book "Norost" (Madness) by Krištof Dovjak and published by Literary society "Hiša poezije" in Ljubljana in 2026 (COBISS.SI-ID: 285395203, ISBN: 978-961-7289-04-6).
The book ...
This item contains 1 file (385.76
KB).
Publicly Available
corpus
Description:
This entry contains the first part of the e-book "Lepoto zrl sem: Zbrane pesmi 1" (I gazed at beauty: Collected poems 1) by Konstandinos Kavafis, translated into Slovenian by Lara Unuk and published by Literary society ...
This item contains 1 file (2.42
MB).
Publicly Available
Most Viewed Items
Top Last Week
corpus
Description:
ParlaMint 5.0 is a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions, mostly starting in 2015 and extending to mid-2022. The individual corpora ...
This item contains 31 files (5.94
GB).
Publicly Available
lexicalConceptualResource
Description:
A lexicon of 751 emoji characters with automatically assigned sentiment.
The sentiment is computed from 70,000 tweets, labeled by 83 human annotators
in 13 European languages.
The process and analysis of emoji sentiment ...
This item contains 3 files (93.95
KB).
Publicly Available
corpus
Description:
goo300k is a manually annotated reference corpus of historical Slovene. It contains 1,100 pages (about 300,000 tokens) sampled from 89 texts from the period 1584-1899.
Each text contains extensive meta-data and per-page ...
This item contains 2 files (8.9
MB).
Publicly Available