Najnovejše

 lexicalConceptualResource 
lexicalConceptualResource
Opis:
The lists contain consonant-vowel structures of all lemmas and word forms in the Gigafida 2.0 corpus. In each unit, its characters were converted as follows: C - consonant (in lists with finegrained character categorizations, ...
 Ta vnos vsebuje 5 datotek(e) (141.75 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike
 corpus 
corpus
Avtor(ji):
Opis:
KOMET 1.0 is a hand-annotated corpus for metaphorical expressions which contains about 200,000 words from Slovene journalistic, fiction and on-line texts. To annotate metaphors in the corpus an adapted and modified ...
 Ta vnos vsebuje 1 datoteko (6.97 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike
 corpus 
corpus
Opis:
The GORDAN 1.0 corpus contains authentic data of spoken communication, annotated for dialogue acts. This entry contains the complete audio files of the corpus (seven wav files, 1 hour of recording), and video files (four ...
 Ta vnos vsebuje 1 datoteko (1.87 GB).
 
Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike

Največ ogledov

V preteklem tednu
 lexicalConceptualResource 
lexicalConceptualResource
Opis:
A lexicon of 751 emoji characters with automatically assigned sentiment. The sentiment is computed from 70,000 tweets, labeled by 83 human annotators in 13 European languages. The process and analysis of emoji sentiment ...
 Ta vnos vsebuje 3 datotek(e) (93.95 KB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike
 corpus 
corpus
Avtor(ji):
Opis:
The corpus contains 256,567 documents from the Slovenian news portals 24ur, Dnevnik, Finance, Rtvslo, and Žurnal24. These portals contain political, business, economic and financial content. The submission contains 7 files: ...
 Ta vnos vsebuje 8 datotek(e) (616.88 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Share Alike
 corpus 
corpus
Opis:
The novel "1984" by George Orwell is the central component of the MULTEXT-East corpus. This parallel and sentence aligned corpus contains the novel in the English original (about 100,000 words in length), and its translations ...
 Ta vnos vsebuje 1 datoteko (14.12 MB).
 
Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike