What's New
lexicalConceptualResource
Description:
MEZZANINE-NstdLex is a dataset containing 4,281 potentially non-standard vocabulary candidates from the Sloleks Morphological Lexicon of Slovene (collected from among the manually inspected entries of version 3.0; ...
This item contains 1 file (83.42
KB).
Publicly Available
corpus
Description:
This entry contains the first part of the audiobook "Abonma" (Season ticket) by author Nataša Velikonja (COBISS ID: 287114755, ISBN: 978-961-7267-07-5).
The poetry collection "Abonma" is the first openly lesbian poetry ...
This item contains 1 file (10.07
MB).
Publicly Available
corpus
Description:
This entry contains the first part of the audiobook "Ampak, kdo?" (But who?) by author Nina Dragičević (COBISS ID: 287116803, ISBN: 978-961-7267-08-2).
In her poetry collection "Ampak, kdo?", Nina Dragičević uses her ...
This item contains 1 file (10.07
MB).
Publicly Available
Most Viewed Items
Top Last Week
corpus
Description:
ParlaMint 5.0 is a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions, mostly starting in 2015 and extending to mid-2022. The individual corpora ...
This item contains 31 files (5.94
GB).
Publicly Available
corpus
Description:
ReLDI-NormTagNER-sr 2.1 is a manually annotated corpus of Serbian tweets. It is meant as a gold-standard training and testing dataset for tokenisation, sentence segmentation, word normalisation, morphosyntactic tagging, ...
This item contains 4 files (4.51
MB).
Publicly Available
corpus
Description:
goo300k is a manually annotated reference corpus of historical Slovene. It contains 1,100 pages (about 300,000 tokens) sampled from 89 texts from the period 1584-1899.
Each text contains extensive meta-data and per-page ...
This item contains 2 files (8.9
MB).
Publicly Available