Najnovejše

 lexicalConceptualResource 
lexicalConceptualResource
Opis:
The lists contain consonant-vowel structures of all lemmas, word forms, and normalized word forms in the GOS 1.0 Corpus of Spoken Slovene (http://hdl.handle.net/11356/1040). In each unit, its characters were converted as ...
 Ta vnos vsebuje 7 datotek(e) (3.6 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike
 lexicalConceptualResource 
lexicalConceptualResource
Opis:
The lists contain consonant-vowel structures of all lemmas and word forms in the Gigafida 2.0 corpus. In each unit, its characters were converted as follows: C - consonant (in lists with finegrained character categorizations, ...
 Ta vnos vsebuje 5 datotek(e) (141.75 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike
 corpus 
corpus
Avtor(ji):
Opis:
KOMET 1.0 is a hand-annotated corpus for metaphorical expressions which contains about 200,000 words from Slovene journalistic, fiction and on-line texts. To annotate metaphors in the corpus an adapted and modified ...
 Ta vnos vsebuje 1 datoteko (6.97 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike

Največ ogledov

V preteklem tednu
 corpus 
corpus
Avtor(ji):
Opis:
SentiCoref 1.0 corpus consists of 837 documents selected from SentiNews 1.0 corpus (http://hdl.handle.net/11356/1110). The documents were selected based on the number of automatically detected named entities (using Polyglot, ...
 Ta vnos vsebuje 2 datotek(e) (6.64 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required
 corpus 
corpus
Avtor(ji):
Opis:
The corpus contains 256,567 documents from the Slovenian news portals 24ur, Dnevnik, Finance, Rtvslo, and Žurnal24. These portals contain political, business, economic and financial content. The submission contains 7 files: ...
 Ta vnos vsebuje 8 datotek(e) (616.88 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike
 corpus 
corpus
Opis:
The ssj500k training corpus contains about 500,000 tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, and lemmatisation. About half of the corpus is also manually ...
 Ta vnos vsebuje 4 datotek(e) (40.95 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike