• Repository
  • About
  • Contact
  • CLARIN
  •  Login
  • English Slovenščina
  • CLARIN.SI repository
  • Search
  • CLARIN logo
  •   Browse  
    •    All of the Repository  
      •   Issue Date
      •   Authors
      •   Titles
      •   Subjects
      •   Publisher
      •   Language
      •   Type
      •   Rights Label
  •   My Account  
    •    Login
  •   General Information  
    •    Deposit
    •    Cite
    •    Submission Lifecycle
    •    FAQ
    •    About
    •    Help Desk
 

 
Selected Filters
 Author : Ljubešić, Nikola     Clear All
Advanced Search

Filters

Use filters to refine the search results.

Current Filters:
New Filters:

Limit your search

Author  
    • Erjavec, Tomaž (41)
    • Fišer, Darja (24)
    • Toral, Antonio (22)
    • Esplà-Gomis, Miquel (21)
    • Rupnik, Peter (21)
    • Kuzman, Taja (17)
    • Bañón, Marta (16)
    • Forcada, Mikel L. (16)
    • García-Romero, Cristian (16)
    • Pla Sempere, Leopoldo (16)
    • Ramírez-Sánchez, Gema (16)
    • Suchomel, Vít (16)
    • van der Werff, Tobias (16)
    • van Noord, Rik (16)
    • Zaragoza, Jaume (16)
    • Borovič, Mladen (9)
    • Boškovič, Borko (9)
    • Dobrovoljc, Kaja (9)
    • Ferme, Marko (9)
    • ... View More
Subject  
    • language model (31)
    • web corpus (27)
    • computer-mediated communication (20)
    • multilingual (17)
    • part-of-speech tagging (17)
    • lemmatisation (16)
    • parallel corpus (16)
    • TEI (16)
    • manual annotation (11)
    • academic writing (10)
    • named entities (10)
    • word normalisation (9)
    • BSc/BA theses (8)
    • MSc/MA theses (8)
    • PhD theses (8)
    • collocations (7)
    • named entity recognition (7)
    • parsing (7)
    • terminology (7)
    • word embeddings (6)
    • ... View More
Rights  
    • PUB (104)
    • ACA (16)
Language (ISO)  
    • Slovenian (51)
    • Croatian (37)
    • English (31)
    • Serbian (22)
    • Bulgarian (10)
    • Bosnian (7)
    • Macedonian (7)
    • Dutch (6)
    • Finnish (5)
    • Icelandic (5)
    • Montenegrin (5)
    • Spanish (5)
    • Turkish (5)
    • Czech (4)
    • Danish (4)
    • French (4)
    • Hungarian (4)
    • Italian (4)
    • Latvian (4)
    • Lithuanian (4)
    • ... View More
Type  
    • text (88)
    • corpus (73)
    • toolService (33)
    • lexicalConceptualResource (16)
    • audio (1)
Contain Files  
    • yes (120)
    • no (2)

Showing 1 through 80 out of 122 results

  • 1
  • 2
  •  
  •    
    • Sort items by
    •  Relevance
    • Title Asc
    • Title Desc
    • Issue Date Asc
    • Issue Date Desc
    •  
    • Results/page
    • 5
    • 10
    • 20
    • 40
    • 60
    •  80
    • 100

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    SimLex-999 Slovenian translation SimLex-999-sl 1.0
    (University of Ljubljana / 2020-05-15)
    
    Author(s):
    Pollak, Senja ; et al.show everyone Pollak, Senja ; Vulić, Ivan ; Pelicon, Andraž ; Repar, Andraž ; Armendariz, Carlos ; Matthew, Purver ; Ljubešić, Nikola
     This item contains 3 files (37.3 KB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    A Resource for Evaluating Graded Word Similarity in Context: CoSimLex
    (Queen Mary University / 2020)
    
    Author(s):
    Armendariz, Carlos ; et al.show everyone Armendariz, Carlos ; Matthew, Purver ; Ulčar, Matej ; Pollak, Senja ; Ljubešić, Nikola ; Robnik-Šikonja, Marko ; Granroth-Wilding, Mark ; Vaik, Kristiina
     This item contains 5 files (486.73 KB).
     
    Publicly Available

  • corpus
    CLARIN.SI data & tools
    corpus
    Machine Translation datasets from the KAS corpus KAS-MT 1.0
    (Faculty of Electrical Engineering and Computer Science, University of Maribor; Faculty of Computer and Information Science, University of Ljubljana / 2022-02-04)
    
    Author(s):
    Žagar, Aleš ; et al.show everyone Žagar, Aleš ; Kavaš, Matic ; Robnik-Šikonja, Marko ; Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola ; Ferme, Marko ; Borovič, Mladen ; Boškovič, Borko ; Ojsteršek, Milan ; Hrovat, Goran
     This item contains 1 file (182.14 MB).
     
    Academic Use Inform Before Use Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    Summarization datasets from the KAS corpus KAS-Sum 1.0
    (Faculty of Electrical Engineering and Computer Science, University of Maribor; Faculty of Computer and Information Science, University of Ljubljana / 2022-02-04)
    
    Author(s):
    Žagar, Aleš ; et al.show everyone Žagar, Aleš ; Kavaš, Matic ; Robnik-Šikonja, Marko ; Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola ; Ferme, Marko ; Borovič, Mladen ; Boškovič, Borko ; Ojsteršek, Milan ; Hrovat, Goran
     This item contains 1 file (4.11 GB).
     
    Academic Use Inform Before Use Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    Slovenian Twitter hate speech dataset IMSyPP-sl
    (Jožef Stefan Institute / 2021-02-17)
    
    Author(s):
    Kralj Novak, Petra ; Mozetič, Igor and Ljubešić, Nikola
     This item contains 4 files (5.19 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Abstracts from the KAS corpus KAS-Abs 1.0
    (Jožef Stefan Institute; Faculty of Electrical Engineering and Computer Science, University of Maribor / 2021-03-31)
    
    Author(s):
    Erjavec, Tomaž ; et al.show everyone Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola ; Ferme, Marko ; Borovič, Mladen ; Boškovič, Borko ; Ojsteršek, Milan ; Hrovat, Goran
     This item contains 1 file (178.99 MB).
     
    Academic Use Inform Before Use Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    Finnish web corpus fiWaC 1.0
    (Jožef Stefan Institute / 2016-09-20)
    
    Author(s):
    Ljubešić, Nikola ; Pirinen, Tommi and Toral, Antonio
     This item contains 38 files (15.28 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Serbian-English parallel corpus srenWaC 1.0
    (Jožef Stefan Institute / 2016-03-09)
    
    Author(s):
    Ljubešić, Nikola ; Esplà-Gomis, Miquel ; Ortiz Rojas, Sergio ; Klubička, Filip and Toral, Antonio
     This item contains 1 file (70.94 MB).
     
    Academic Use Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    Dataset and baseline model of moderated content FRENK-STYRIA-24sata 1.0
    (Jožef Stefan Institute / 2018-10-27)
    
    Author(s):
    Ljubešić, Nikola ; Erjavec, Tomaž and Fišer, Darja
     This item contains 2 files (7.62 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for UD dependency parsing of standard Bulgarian 1.0
    (Jožef Stefan Institute; IICT-BAS / 2020-06-24)
    
    Author(s):
    Ljubešić, Nikola ; Osenova, Petya and Simov, Kiril
     This item contains 2 files (476.56 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    The LiLaH Emotion Lexicon of Croatian, Dutch and Slovene
    (Jožef Stefan Institute; Centre for Computational Linguistics and Psycholinguistics (CLiPS) / 2020-06-04)
    
    Author(s):
    Daelemans, Walter ; et al.show everyone Daelemans, Walter ; Fišer, Darja ; Franza, Jasmin ; Kranjčić, Denis ; Lemmens, Jens ; Ljubešić, Nikola ; Markov, Ilia ; Popič, Damjan
     This item contains 1 file (199.85 KB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Finnish-English parallel corpus fienWaC 1.0
    (Jožef Stefan Institute / 2016-03-09)
    
    Author(s):
    Ljubešić, Nikola ; Esplà-Gomis, Miquel ; Ortiz Rojas, Sergio ; Klubička, Filip and Toral, Antonio
     This item contains 1 file (283.67 MB).
     
    Academic Use Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    Corpus of Academic Slovene (PhD theses) KAS-dr 1.0
    (Jožef Stefan Institute; Faculty of Electrical Engineering and Computer Science, University of Maribor / 2019-11-28)
    
    Author(s):
    Erjavec, Tomaž ; et al.show everyone Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola ; Ferme, Marko ; Borovič, Mladen ; Boškovič, Borko ; Ojsteršek, Milan ; Hrovat, Goran
     This item contains 3 files (2.52 GB).
     
    Academic Use Inform Before Use Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    English-Montenegrin parallel corpus of subtitles Opus-MontenegrinSubs 1.0
    (Jožef Stefan Institute / 2018-03-20)
    
    Author(s):
    Božović, Petar ; Erjavec, Tomaž ; Tiedemann, Jörg ; Ljubešić, Nikola and Gorjanc, Vojko
     This item contains 2 files (12.86 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Croatian language corpus Riznica 0.1
    (Institute of Croatian Language and Linguistics / 2018-03-07)
    
    Author(s):
    Brozović Rončević, Dunja ; et al.show everyone Brozović Rončević, Dunja ; Ćavar, Damir ; Ćavar, Małgorzata ; Stojanov, Tomislav ; Štrkalj Despot, Kristina ; Ljubešić, Nikola ; Erjavec, Tomaž
     This item contains 1 file (457.73 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for named entity recognition of standard Bulgarian 1.0
    (Jožef Stefan Institute; IICT-BAS / 2020-07-07)
    
    Author(s):
    Ljubešić, Nikola ; Osenova, Petya and Simov, Kiril
     This item contains 2 files (107.32 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Corpus of Academic Slovene (BSc/BA theses) KAS-dipl 1.0
    (Jožef Stefan Institute; Faculty of Electrical Engineering and Computer Science, University of Maribor / 2019-11-28)
    
    Author(s):
    Erjavec, Tomaž ; et al.show everyone Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola ; Ferme, Marko ; Borovič, Mladen ; Boškovič, Borko ; Ojsteršek, Milan ; Hrovat, Goran
     This item contains 5 files (27.63 GB).
     
    Academic Use Inform Before Use Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    Corpus of Academic Slovene (MSc/MA theses) KAS-mag 1.0
    (Jožef Stefan Institute; Faculty of Electrical Engineering and Computer Science, University of Maribor / 2019-11-28)
    
    Author(s):
    Erjavec, Tomaž ; et al.show everyone Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola ; Ferme, Marko ; Borovič, Mladen ; Boškovič, Borko ; Ojsteršek, Milan ; Hrovat, Goran
     This item contains 3 files (11.97 GB).
     
    Academic Use Inform Before Use Attribution Required Noncommercial

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Automatically constructed multiword lexicon slMWELex v0.5
    (Jožef Stefan Institute / 2015)
    
    Author(s):
    Ljubešić, Nikola ; Krek, Simon ; Dobrovoljc, Kaja and Erjavec, Tomaž
     This item contains 1 file (73.96 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Automatically constructed multiword lexicon hrMWELex v0.5
    (Jožef Stefan Institute / 2015)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 1 file (152.39 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Automatically constructed multiword lexicon srMWELex v0.5
    (Jožef Stefan Institute / 2015)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 1 file (40.26 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for JOS dependency parsing of standard Slovenian 1.0
    (Jožef Stefan Institute / 2020-06-24)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (1.53 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for named entity recognition of standard Slovenian 1.0
    (Jožef Stefan Institute / 2020-06-19)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (106.12 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Corpus of Croatian news portals ENGRI (2014-2018)
    (University of Rijeka, Faculty of Maritime Studies / 2021-03-14)
    
    Author(s):
    Bogunović, Irena ; Kučić, Mario ; Ljubešić, Nikola and Erjavec, Tomaž
     This item contains 12 files (8.48 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Tourism English-Croatian Parallel Corpus 2.0
    (Abu-MaTran project / 2016-01-28)
    
    Author(s):
    Toral, Antonio ; et al.show everyone Toral, Antonio ; Esplà-Gomis, Miquel ; Klubička, Filip ; Ljubešić, Nikola ; Papavassiliou, Vassilis ; Prokopidis, Prokopis ; Rubino, Raphael ; Way, Andy
     This item contains 1 file (69.36 MB).
     
    Academic Use Attribution Required Noncommercial

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Slovene ontology of semantic types for nouns SLONEST-noun 1.0
    (Centre for Language Resources and Technologies, University of Ljubljana / 2020-10-26)
    
    Author(s):
    Kosem, Iztok ; et al.show everyone Kosem, Iztok ; Pori, Eva ; Gantar, Polona ; Logar, Nataša ; Krek, Simon ; Laskowski, Cyprian ; Arhar Holdt, Špela ; Čibej, Jaka ; Dobrovoljc, Kaja ; Gorjanc, Vojko ; Klemenc, Bojan ; Ljubešić, Nikola
     This item contains 1 file (58.7 KB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Dataset of normalised Slovene text KonvNormSl 1.0
    (Jožef Stefan Institute / 2016-09-19)
    
    Author(s):
    Ljubešić, Nikola ; Zupan, Katja ; Fišer, Darja and Erjavec, Tomaž
     This item contains 1 file (4.57 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for named entity recognition of non-standard Croatian 1.0
    (Jožef Stefan Institute / 2020-08-07)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 1 file (46.14 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for named entity recognition of standard Croatian 1.0
    (Jožef Stefan Institute / 2020-06-19)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (106.34 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Concreteness and imageability lexicon MEGA.HR-Crossling
    (Jožef Stefan Institute; Faculty of Humanities and Social Sciences, University of Zagreb / 2018-05-28)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 1 file (164.76 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Corpus of Written Standard Slovene Gigafida 2.0
    (Centre for Language Resources and Technologies, University of Ljubljana / 2019-06-13)
    
    Author(s):
    Krek, Simon ; et al.show everyone Krek, Simon ; Erjavec, Tomaž ; Repar, Andraž ; Čibej, Jaka ; Arhar Holdt, Špela ; Gantar, Polona ; Kosem, Iztok ; Robnik-Šikonja, Marko ; Ljubešić, Nikola ; Dobrovoljc, Kaja ; Laskowski, Cyprian ; Grčar, Miha ; Holozan, Peter ; Šuster, Simon ; Gorjanc, Vojko ; Stabej, Marko ; Logar, Nataša
     This item contains no files.

  • toolService
    CLARIN.SI data & tools
    toolService
    The Orange workflow for observing collocation trends ColTrend 1.0
    (Centre for Language Resources and Technologies, University of Ljubljana / 2020-10-26)
    
    Author(s):
    Kosem, Iztok ; et al.show everyone Kosem, Iztok ; Krek, Simon ; Čibej, Jaka ; Gantar, Polona ; Arhar Holdt, Špela ; Logar, Nataša ; Laskowski, Cyprian ; Klemenc, Bojan ; Ljubešić, Nikola ; Dobrovoljc, Kaja ; Gorjanc, Vojko ; Pori, Eva
     This item contains 1 file (70.03 MB).
     
    Publicly Available

  • corpus
    CLARIN.SI data & tools
    corpus
    Slovene-English parallel corpus slenWaC 1.0
    (Jožef Stefan Institute / 2016-03-10)
    
    Author(s):
    Ljubešić, Nikola ; Esplà-Gomis, Miquel ; Ortiz Rojas, Sergio ; Klubička, Filip and Toral, Antonio
     This item contains 1 file (94.44 MB).
     
    Academic Use Attribution Required Noncommercial

  • toolService
    CLARIN.SI data & tools
    toolService
    The Orange workflow for observing collocation clusters ColEmbed 1.0
    (Centre for Language Resources and Technologies, University of Ljubljana / 2020-10-26)
    
    Author(s):
    Kosem, Iztok ; et al.show everyone Kosem, Iztok ; Čibej, Jaka ; Ljubešić, Nikola ; Krek, Simon ; Gantar, Polona ; Arhar Holdt, Špela ; Logar, Nataša ; Laskowski, Cyprian ; Klemenc, Bojan ; Dobrovoljc, Kaja ; Gorjanc, Vojko ; Pori, Eva
     This item contains 1 file (86.32 MB).
     
    Publicly Available

  • corpus
    CLARIN.SI data & tools
    corpus
    English YouTube Hate Speech Corpus
    (Jožef Stefan Institute / 2021-10-14)
    
    Author(s):
    Ljubešić, Nikola ; Mozetič, Igor ; Cinelli, Matteo and Kralj Novak, Petra
     This item contains 3 files (30.59 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    The Twitter user dataset for discriminating between Bosnian, Croatian, Montenegrin and Serbian Twitter-HBS 1.0
    (Jožef Stefan Institute / 2022-01-26)
    
    Author(s):
    Ljubešić, Nikola and Rupnik, Peter
     This item contains 1 file (12.98 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for UD dependency parsing of standard Croatian
    (Jožef Stefan Institute / 2019-10-11)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (1.13 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Bosnian web corpus bsWaC 1.1
    (Jožef Stefan Institute / 2016-05-12)
    
    Author(s):
    Ljubešić, Nikola and Klubička, Filip
     This item contains 3 files (1.85 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    Word embeddings CLARIN.SI-embed.mk 0.1
    (Jožef Stefan Institute / 2020-10-13)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (1.23 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Corpus of academic Slovene KAS 1.0
    (Jožef Stefan Institute; Faculty of Electrical Engineering and Computer Science, University of Maribor / 2019-11-28)
    
    Author(s):
    Erjavec, Tomaž ; et al.show everyone Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola ; Ferme, Marko ; Borovič, Mladen ; Boškovič, Borko ; Ojsteršek, Milan ; Hrovat, Goran
     This item contains 6 files (42.11 GB).
     
    Academic Use Inform Before Use Attribution Required Noncommercial

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Serbian 1.0
    (Jožef Stefan Institute / 2020-07-17)
    
    Author(s):
    Ljubešić, Nikola and Štefanec, Vanja
     This item contains 2 files (620.27 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Montenegrin web corpus meWaC 1.0
    (Jožef Stefan Institute / 2021-05-13)
    
    Author(s):
    Ljubešić, Nikola and Erjavec, Tomaž
     This item contains 2 files (2.47 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Word embeddings CLARIN.SI-embed.sl 1.0
    (Jožef Stefan Institute / 2018-11-26)
    
    Author(s):
    Ljubešić, Nikola and Erjavec, Tomaž
     This item contains 4 files (6.41 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Word embeddings CLARIN.SI-embed.sr 1.0
    (Jožef Stefan Institute / 2018-12-10)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 4 files (3.36 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • corpus
    CLARIN.SI data & tools
    corpus
    Training corpus hr500k 1.0
    (Jožef Stefan Institute / 2018-04-13)
    
    Author(s):
    Ljubešić, Nikola ; Agić, Željko ; Klubička, Filip ; Batanović, Vuk and Erjavec, Tomaž
     This item contains 3 files (91.53 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Croatian corpus of non-professional written language by typical speakers and speakers with language disorders RAPUT 1.0
    (Jožef Stefan Institute; Faculty of Education and Rehabilitation, University of Zagreb / 2021-06-15)
    
    Author(s):
    Kuvač Kraljević, Jelena ; Hržica, Gordana ; Štefanec, Vanja ; Kologranić Belić, Lana and Ljubešić, Nikola
     This item contains 2 files (8.11 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Twitter corpus Janes-Tweet 1.0
    (Jožef Stefan Institute / 2017-09-05)
    
    Author(s):
    Ljubešić, Nikola ; Erjavec, Tomaž and Fišer, Darja
     This item contains 2 files (1.17 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    News comment corpus Janes-News 1.0
    (Jožef Stefan Institute / 2017-08-17)
    
    Author(s):
    Erjavec, Tomaž ; Ljubešić, Nikola and Fišer, Darja
     This item contains 2 files (186.48 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • corpus
    CLARIN.SI data & tools
    corpus
    Croatian-English parallel corpus hrenWaC 2.0
    (Jožef Stefan Institute / 2016-03-09)
    
    Author(s):
    Ljubešić, Nikola ; Esplà-Gomis, Miquel ; Ortiz Rojas, Sergio ; Klubička, Filip and Toral, Antonio
     This item contains 1 file (186.46 MB).
     
    Academic Use Attribution Required Noncommercial

  • corpus
    CLARIN.SI data & tools
    corpus
    Semantic hypergraph corpus SemCRO 1.0
    (University of Mostar; University of Split; Jožef Stefan Institute / 2020-11-20)
    
    Author(s):
    Vasić, Daniel ; et al.show everyone Vasić, Daniel ; Žitko, Branko ; Gašpar, Angelina ; Ljubešić, Nikola ; Štrkalj Despot, Kristina ; Merkler, Danijela
     This item contains 1 file (21.66 KB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    English-Slovene term candidates KAS-biterm 1.0
    (Jožef Stefan Institute / 2020-05-05)
    
    Author(s):
    Erjavec, Tomaž ; Ljubešić, Nikola and Fišer, Darja
     This item contains 1 file (50.74 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • corpus
    CLARIN.SI data & tools
    corpus
    Slovenian Twitter dataset 2018-2020 1.0
    (Jožef Stefan Institute / 2021-07-20)
    
    Author(s):
    Evkoski, Bojan ; Pelicon, Andraž ; Mozetič, Igor ; Ljubešić, Nikola and Kralj Novak, Petra
     This item contains 1 file (182.04 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Dataset and baseline model of moderated content FRENK-MMC-RTV 1.0
    (Jožef Stefan Institute / 2018-10-27)
    
    Author(s):
    Ljubešić, Nikola ; Erjavec, Tomaž and Fišer, Darja
     This item contains 2 files (4.65 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Serbian web corpus srWaC 1.1
    (Jožef Stefan Institute / 2016-05-12)
    
    Author(s):
    Ljubešić, Nikola and Klubička, Filip
     This item contains 6 files (3.51 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Corpus of Serbian Forms of Address 1.0
    (Slavic Seminary, University of Zurich / 2021-04-06)
    
    Author(s):
    Lemmenmeier-Batinić, Dolores ; Ljubešić, Nikola and Samardžić, Tanja
     This item contains 1 file (2.39 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Noncommercial Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    The news dataset for discriminating between Bosnian, Croatian and Serbian SETimes.HBS 1.0
    (Jožef Stefan Institute / 2022-01-26)
    
    Author(s):
    Ljubešić, Nikola and Rupnik, Peter
     This item contains 1 file (20.15 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Forum corpus Janes-Forum 1.0
    (Jožef Stefan Institute / 2017-08-17)
    
    Author(s):
    Erjavec, Tomaž ; Ljubešić, Nikola and Fišer, Darja
     This item contains 2 files (573.23 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • corpus
    CLARIN.SI data & tools
    corpus
    Blog post and comment corpus Janes-Blog 1.0
    (Jožef Stefan Institute / 2017-08-17)
    
    Author(s):
    Erjavec, Tomaž ; Ljubešić, Nikola and Fišer, Darja
     This item contains 2 files (411.31 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • corpus
    CLARIN.SI data & tools
    corpus
    Wikipedia talk corpus Janes-Wiki 1.0
    (Jožef Stefan Institute / 2017-08-28)
    
    Author(s):
    Ljubešić, Nikola ; Erjavec, Tomaž and Fišer, Darja
     This item contains 2 files (55.35 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Choice of plausible alternatives dataset in Croatian COPA-HR
    (Jožef Stefan Institute / 2021-02-24)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 3 files (194.2 KB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.0
    (Jožef Stefan Institute / 2021-05-28)
    
    Author(s):
    Ljubešić, Nikola ; Fišer, Darja and Erjavec, Tomaž
     This item contains 1 file (4.17 MB).
     
    Academic Use Inform Before Use Attribution Required Noncommercial

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for morphosyntactic annotation of standard Serbian
    (Jožef Stefan Institute / 2019-10-10)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (549.14 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Word embeddings CLARIN.SI-embed.hr 1.0
    (Jožef Stefan Institute / 2018-12-10)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 4 files (4.88 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • corpus
    CLARIN.SI data & tools
    corpus
    Training corpus SETimes.SR 1.0
    (Regional Linguistic Data Initiative Centre ReLDI / 2018-08-20)
    
    Author(s):
    Batanović, Vuk ; Ljubešić, Nikola ; Samardžić, Tanja and Erjavec, Tomaž
     This item contains 3 files (10.91 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Croatian 1.0
    (Jožef Stefan Institute / 2020-07-17)
    
    Author(s):
    Ljubešić, Nikola and Štefanec, Vanja
     This item contains 2 files (1.12 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Terminology identification dataset KAS-term 1.0
    (Jožef Stefan Institute / 2018-08-18)
    
    Author(s):
    Erjavec, Tomaž ; et al.show everyone Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola ; Arhar Holdt, Špela ; Bren, Urban ; Robnik-Šikonja, Marko ; Udovič, Boštjan
     This item contains 4 files (17.26 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for UD dependency parsing of standard Slovenian
    (Jožef Stefan Institute / 2019-10-11)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (1.54 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for lemmatisation of standard Macedonian 1.0
    (Jožef Stefan Institute / 2020-11-05)
    
    Author(s):
    Ljubešić, Nikola ; Zdravkova, Katerina and Erjavec, Tomaž
     This item contains 1 file (2.9 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for named entity recognition of non-standard Slovenian 1.0
    (Jožef Stefan Institute / 2020-08-07)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 1 file (46.12 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Croatian web corpus hrWaC 2.1
    (Jožef Stefan Institute / 2016-05-12)
    
    Author(s):
    Ljubešić, Nikola and Klubička, Filip
     This item contains 15 files (9.21 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • lexicalConceptualResource
    CLARIN.SI data & tools
    lexicalConceptualResource
    Collocations Dictionary of Modern Slovene KSSS 1.0
    (Centre for Language Resources and Technologies, University of Ljubljana / 2019-09-20)
    
    Author(s):
    Kosem, Iztok ; et al.show everyone Kosem, Iztok ; Gantar, Polona ; Krek, Simon ; Arhar Holdt, Špela ; Čibej, Jaka ; Laskowski, Cyprian ; Pori, Eva ; Klemenc, Bojan ; Dobrovoljc, Kaja ; Gorjanc, Vojko ; Ljubešić, Nikola
     This item contains 1 file (311.31 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    JRC EU DGT Translation Memory Parsebank DGT-UD 1.0
    (Jožef Stefan Institute / 2018-08-15)
    
    Author(s):
    Ljubešić, Nikola and Erjavec, Tomaž
     This item contains 24 files (24.42 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for named entity recognition of non-standard Serbian 1.0
    (Jožef Stefan Institute / 2020-08-07)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 1 file (46.15 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for named entity recognition of standard Serbian 1.0
    (Jožef Stefan Institute / 2020-06-19)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (106.08 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Slovene Web genre identification corpus GINCO 1.0
    (Jožef Stefan Institute / 2021-12-02)
    
    Author(s):
    Kuzman, Taja ; Brglez, Mojca ; Rupnik, Peter and Ljubešić, Nikola
     This item contains 2 files (1.77 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Text collection for training the BERTić transformer model BERTić-data
    (Jožef Stefan Institute / 2021-05-05)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 10 files (21.14 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for UD dependency parsing of standard Serbian
    (Jožef Stefan Institute / 2019-10-11)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (624 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • toolService
    CLARIN.SI data & tools
    toolService
    The CLASSLA-StanfordNLP model for morphosyntactic annotation of non-standard Slovenian 1.0
    (Jožef Stefan Institute / 2020-08-06)
    
    Author(s):
    Ljubešić, Nikola
     This item contains 2 files (1.57 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    Bilingual terminology extraction dataset KAS-biterm 1.0
    (Jožef Stefan Institute / 2018-08-18)
    
    Author(s):
    Erjavec, Tomaž ; Fišer, Darja ; Ljubešić, Nikola and Bitenc, Maja
     This item contains 2 files (1.83 MB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • corpus
    CLARIN.SI data & tools
    corpus
    ASR training dataset for Croatian ParlaSpeech-HR v1.0
    (Jožef Stefan Institute / 2022-04-04)
    
    Author(s):
    Ljubešić, Nikola ; et al.show everyone Ljubešić, Nikola ; Koržinek, Danijel ; Rupnik, Peter ; Jazbec, Ivo-Pavao ; Batanović, Vuk ; Bajčetić, Lenka ; Evkoski, Bojan
     This item contains 5 files (117.25 GB).
     
    Publicly Available Distributed under Creative Commons Attribution Required Share Alike

  • 1
  • 2
  •  
  •    
    • Sort items by
    •  Relevance
    • Title Asc
    • Title Desc
    • Issue Date Asc
    • Issue Date Desc
    •  
    • Results/page
    • 5
    • 10
    • 20
    • 40
    • 60
    •  80
    • 100
 

Partners

  • Alpineon, d.o.o.
  • Amebis, d.o.o.
  • Institute of Contemporary History
  • Jožef Stefan Institute
  • Slovenian Language Technologies Society
  • Trojina, Institute for Applied Slovene Studies

Partners

  • University of Ljubljana
  • University of Maribor
  • University of Nova Gorica
  • University of Primorska
  • ZRC SAZU
  • ZRS Koper

Repository

  • Main page
  • Contact
  • Submission Lifecycle
  • FAQ
  • About and Policies

This platform runs under the software developed for the LINDAT/CLARIAH-CZ repository for linguistics, available on GitHub

CLARIN.SI is supported by the Ministry of Education, Science and Sport of the Republic of Slovenia
under the Programme of "Research Infrastructures".