Najnovejše

 corpus 
corpus
Opis:
Gigafida 2.0, with about 1.1 billion words, is a reference corpus of written Slovene text published in the period 1990-2018. It is comprised of daily news, magazines, a selection of web texts (a certain portion of which ...
 Ta vnos ne vsebuje datotek.
 
Publicly Available
 lexicalConceptualResource 
lexicalConceptualResource
Opis:
The "School Dictionary of the Slovenian Language" ("Šolski slovar slovenskega jezika") includes 2045 dictionary entries and is aimed at pupils aged 6 to 10. It provides the most relevant language information for the target ...
 Ta vnos vsebuje 1 datoteko (480.06 KB).
 
Publicly Available Distributed under Creative Commons Attribution Required
 corpus 
corpus
Opis:
We present an Italian YouTube dataset manually annotated for hate speech types and targets. The comments to be annotated were sampled from the Italian YouTube comments on videos about the Covid-19 pandemic in the period ...
 Ta vnos vsebuje 3 datotek(e) (117.62 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike

Največ ogledov

V preteklem tednu
 corpus 
corpus
Opis:
The Croatian web corpus hrWaC was built by crawling the .hr top-level domain in 2011 and again in 2014. The corpus was near-deduplicated on paragraph level, normalised via diacritic restoration, morphosyntactically annotated ...
 Ta vnos vsebuje 15 datotek(e) (9.21 GB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike
 corpus 
corpus
Opis:
The dataset contains over 1.6 million tweets (tweet IDs), labeled with sentiment by human annotators. There are 15 Twitter corpora for the corresponding 15 European languages. The data can be used to train and evaluate ...
 Ta vnos vsebuje 16 datotek(e) (49.38 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike
 corpus 
corpus
Opis:
FRENK-STYRIA-24sata is a dataset of moderated newspaper comments from the website 24sata.hr with metadata on the time of publishing, user identifier, thread identifier and whether the comment was deleted by the moderators ...
 Ta vnos vsebuje 2 datotek(e) (7.62 GB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike