What's New
corpus
Description:
This entry contains the Ilustrirani Slovenec Multimodal Document Understanding Dataset, a benchmark dataset designed for training and evaluating vision-language models on historical Slovenian newspaper pages. The dataset ...
This item contains 1 file (7.57
MB).
Publicly Available
corpus
Description:
This entry contains the first part of the e-book "Norost" (Madness) by Krištof Dovjak and published by Literary society "Hiša poezije" in Ljubljana in 2026 (COBISS.SI-ID: 285395203, ISBN: 978-961-7289-04-6).
The book ...
This item contains 1 file (385.76
KB).
Publicly Available
corpus
Description:
This entry contains the first part of the e-book "Lepoto zrl sem: Zbrane pesmi 1" (I gazed at beauty: Collected poems 1) by Konstandinos Kavafis, translated into Slovenian by Lara Unuk and published by Literary society ...
This item contains 1 file (2.42
MB).
Publicly Available
Most Viewed Items
Top Last Week
corpus
Description:
ParlaMint 5.0 is a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions, mostly starting in 2015 and extending to mid-2022. The individual corpora ...
This item contains 31 files (5.94
GB).
Publicly Available
lexicalConceptualResource
Description:
Frazeološki rječnik hrvatskoga jezika is an open-access dictionary of Croatian idioms based on data from a large electronic corpus. The resulting dictionary will serve as a gateway for a large number of users and researchers ...
This item contains no files.
corpus
Description:
goo300k is a manually annotated reference corpus of historical Slovene. It contains 1,100 pages (about 300,000 tokens) sampled from 89 texts from the period 1584-1899.
Each text contains extensive meta-data and per-page ...
This item contains 2 files (8.9
MB).
Publicly Available