CLARIN.SI repository

What's New

corpus

CLARIN.SI data & tools

"Choice of plausible alternatives" datasets in South Slavic dialects DIALECT-COPA

Author(s):

Ljubešić, Nikola ; et al.show everyone

Ljubešić, Nikola ; Kuzman, Taja ; Rupnik, Peter ; Milosavljević, Stefan ; Galant, Nada ; Benčina, Sonja ; Čibej, Jaka

Description:

The DIALECT-COPA datasets comprise Choice of Plausible Alternatives (COPA) datasets for three South Slavic dialects: (1) COPA-SL-CER for the Cerkno dialect of Slovenian, spoken in the Slovenian Littoral region, specifically ...

This item contains 6 files (279.69 KB).

Publicly Available Distributed under Creative Commons

languageDescription

CLARIN.SI data & tools

Overview of inflectional paradigms in Slovenian

Author(s):

Štarkl, Ema ; Mišmaš, Petra and Simonović, Marko

Description:

The purpose of the overview is to provide a comprehensive overview of the inflectional features associated with specific endings. Each ending has a dedicated row in the table and is exemplified by a word in the relevant ...

This item contains 2 files (208.07 KB).

Publicly Available Distributed under Creative Commons

corpus

CLARIN.SI data & tools

The Sarajevo Corpus of SMS Messages in Bosnian

Author(s):

Wasserscheidt, Philipp ; et al.show everyone

Wasserscheidt, Philipp ; Bulić, Halid ; Durmišević, Elma ; Hodžić-Čavkić, Azra ; Bajraktarević, Enisa ; Ahmetspahić-Peljto, Azra ; Šabić, Belmin

Description:

This corpus is specialized, static (i.e., no future growth is planned), diachronic and covers the period from 2002 to 2022. All messages included in this Corpus were obtained from voluntary donors (informants). Both senders ...

This item contains 1 file (1.73 MB).

Publicly Available Distributed under Creative Commons

Most Viewed Items

Top Last Week

corpus

CLARIN.SI data & tools

Croatian corpus of non-professional written language by typical speakers and speakers with language disorders RAPUT 1.0

Author(s):

Kuvač Kraljević, Jelena ; Hržica, Gordana ; Štefanec, Vanja ; Kologranić Belić, Lana and Ljubešić, Nikola

Description:

The corpus consists of texts produced by nonprofessional typical speakers and speakers with different language disorders (developmental language disorder, dyslexia, traumatic brain injury, aphasia, other). Roughly half of ...

This item contains 2 files (8.11 MB).

Publicly Available Distributed under Creative Commons

languageDescription

CLARIN.SI data & tools

Overview of inflectional paradigms in Slovenian

Author(s):

Štarkl, Ema ; Mišmaš, Petra and Simonović, Marko

Description:

This item contains 2 files (208.07 KB).

Publicly Available Distributed under Creative Commons

corpus

CLARIN.SI data & tools

Multilingual comparable corpora of parliamentary debates ParlaMint 4.0

Author(s):

Erjavec, Tomaž ; et al.show everyone

Erjavec, Tomaž ; Kopp, Matyáš ; Ogrodniczuk, Maciej ; Osenova, Petya ; Agirrezabal, Manex ; Agnoloni, Tommaso ; Aires, José ; Albini, Monica ; Alkorta, Jon ; Antiba-Cartazo, Iván ; Arrieta, Ekain ; Barcala, Mario ; Bardanca, Daniel ; Barkarson, Starkaður ; Bartolini, Roberto ; Battistoni, Roberto ; Bel, Nuria ; Bonet Ramos, Maria del Mar ; Calzada Pérez, María ; Cardoso, Aida ; Çöltekin, Çağrı ; Coole, Matthew ; Darģis, Roberts ; de Libano, Ruben ; Depoorter, Griet ; Diwersy, Sascha ; Dodé, Réka ; Fernandez, Kike ; Fernández Rei, Elisa ; Frontini, Francesca ; Garcia, Marcos ; García Díaz, Noelia ; García Louzao, Pedro ; Gavriilidou, Maria ; Gkoumas, Dimitris ; Grigorov, Ilko ; Grigorova, Vladislava ; Haltrup Hansen, Dorte ; Iruskieta, Mikel ; Jarlbrink, Johan ; Jelencsik-Mátyus, Kinga ; Jongejan, Bart ; Kahusk, Neeme ; Kirnbauer, Martin ; Kryvenko, Anna ; Ligeti-Nagy, Noémi ; Ljubešić, Nikola ; Luxardo, Giancarlo ; Magariños, Carmen ; Magnusson, Måns ; Marchetti, Carlo ; Marx, Maarten ; Meden, Katja ; Mendes, Amália ; Mochtak, Michal ; Mölder, Martin ; Montemagni, Simonetta ; Navarretta, Costanza ; Nitoń, Bartłomiej ; Norén, Fredrik Mohammadi ; Nwadukwe, Amanda ; Ojsteršek, Mihael ; Pančur, Andrej ; Papavassiliou, Vassilis ; Pereira, Rui ; Pérez Lago, María ; Piperidis, Stelios ; Pirker, Hannes ; Pisani, Marilina ; Pol, Henk van der ; Prokopidis, Prokopis ; Quochi, Valeria ; Rayson, Paul ; Regueira, Xosé Luís ; Rudolf, Michał ; Ruisi, Manuela ; Rupnik, Peter ; Schopper, Daniel ; Simov, Kiril ; Sinikallio, Laura ; Skubic, Jure ; Tungland, Lars Magne ; Tuominen, Jouni ; van Heusden, Ruben ; Varga, Zsófia ; Vázquez Abuín, Marta ; Venturi, Giulia ; Vidal Miguéns, Adrián ; Vider, Kadri ; Vivel Couso, Ainhoa ; Vladu, Adina Ioana ; Wissik, Tanja ; Yrjänäinen, Väinö ; Zevallos, Rodolfo ; Fišer, Darja

Description:

ParlaMint 4.0 is a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions, mostly starting in 2015 and extending to mid-2022. The individual corpora ...

This item contains 30 files (5.67 GB).

Publicly Available Distributed under Creative Commons

Linguistic Data and NLP Tools

Find

Citation Support (with Persistent IDs)

Deposit Free and Safe

License of your Choice (Open licenses encouraged)

Easy to Find

Easy to Cite

What's New

Most Viewed Items

Partners

Partners

Repository