<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
<title>CLARIN.SI data &amp; tools</title>
<link href="http://hdl.handle.net/11356/1024" rel="alternate"/>
<subtitle>CLARIN.SI repository language resources and tools</subtitle>
<id>http://hdl.handle.net/11356/1024</id>
<updated>2026-08-15T02:22:52Z</updated>
<dc:date>2026-08-15T02:22:52Z</dc:date>
<entry>
<title>List of headword candidates for the Dutch-Slovene Dictionary nl-sl-Lex 1.0</title>
<link href="http://hdl.handle.net/11356/2220" rel="alternate"/>
<author>
<name>Grahek, Egidij</name>
</author>
<author>
<name>Srebnik, Anita</name>
</author>
<author>
<name>Čibej, Jaka</name>
</author>
<id>http://hdl.handle.net/11356/2220</id>
<updated>2026-08-14T11:22:02Z</updated>
<published>2026-06-01T00:00:00Z</published>
<summary type="text">List of headword candidates for the Dutch-Slovene Dictionary nl-sl-Lex 1.0
Grahek, Egidij; Srebnik, Anita; Čibej, Jaka
nl-sl-Lex is a list of headword candidates that can be used as a basis to update the Dutch-Slovene Dictionary (Srebnik 2007). The list was extracted from the nlTenTen20 corpus (https://www.sketchengine.eu/nltenten-dutch-corpus/) using Sketch Engine (https://www.sketchengine.eu/) and custom Python scripts. For each part-of-speech (NOUN, VERB, ADJ, ADV; according to the Universal Dependencies POS-tagset: https://universaldependencies.org/u/pos/), lists of up to 1,000 lexeme candidates beginning with each letter of the Dutch alphabet (only lower-case candidates were considered) were extracted from nlTenTen20 (along with their absolute and relative frequencies), then merged, sorted by frequency in descending order, and cross-compared with several existing resources:&#13;
&#13;
• the Dutch-Slovene Dictionary (Srebnik 2007)&#13;
• the list of headwords from the Dictionary of Contemporary Dutch (Algemeen Nederlands Woordenboek - ANW; https://anw.ivdnt.org/lemmalist)&#13;
• the list of 5,000 most frequent Dutch words (Tiberius &amp; Schoonheim 2014)&#13;
• two frequency lists for Dutch - 0-2000 &amp; 2000-5000 by Hazenberg &amp; Hulstijn (de Boer et al. 2012)&#13;
&#13;
Version 1.0 contains 60,777 candidates and was compiled as a starting point for an analysis of the Dutch-Slovene Dictionary in order to determine priorities for future dictionary updates and facilitate linking with other resources such as WordNets. For more information on the structure of the list, please consult 00README.txt.
</summary>
<dc:date>2026-06-01T00:00:00Z</dc:date>
</entry>
<entry>
<title>English-Slovene sample of the ETHICS dataset ETHICS-EN-SL 1.0</title>
<link href="http://hdl.handle.net/11356/2337" rel="alternate"/>
<author>
<name>Novak, Ela</name>
</author>
<id>http://hdl.handle.net/11356/2337</id>
<updated>2026-08-13T13:43:12Z</updated>
<published>2026-08-13T00:00:00Z</published>
<summary type="text">English-Slovene sample of the ETHICS dataset ETHICS-EN-SL 1.0
Novak, Ela
English-Slovene sample of the ETHICS dataset ETHICS-EN-SL 1.0 contains an English-Slovene sample derived from the ETHICS dataset (Hendrycks et al. 2021), prepared for a master’s thesis on language as a factor in the moral evaluation of large language models. The resource contains selected examples from five ETHICS subsets: commonsense moral judgment, deontology, justice, virtue ethics, and utilitarianism.&#13;
&#13;
The dataset contains 100 examples or scenario pairs per subset. The English files contain the selected examples from the original ETHICS dataset, while the Slovene files contain translated and partially localized versions of the same examples. The bilingual files align the English and Slovene versions row by row. For the Slovene part of the study, a human reference baseline was established on the basis of majority judgments by Slovene annotators.&#13;
&#13;
The resource is intended for research on multilingual moral evaluation, language effects in large language models, value alignment, dataset translation and localization, and the creation of localized evaluation resources for less-resourced languages. The data are distributed as UTF-8 encoded TSV files with header rows. The package also includes a README file with the directory structure and column descriptions.&#13;
&#13;
The resource is based on the ETHICS dataset introduced by Hendrycks et al. (2021): &#13;
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. “Aligning AI With Shared Human Values.” ICLR 2021. https://arxiv.org/abs/2008.02275
</summary>
<dc:date>2026-08-13T00:00:00Z</dc:date>
</entry>
<entry>
<title>Monitor corpus of Slovene Trendi 2026-07</title>
<link href="http://hdl.handle.net/11356/2331" rel="alternate"/>
<author>
<name>Kosem, Iztok</name>
</author>
<author>
<name>Čibej, Jaka</name>
</author>
<author>
<name>Dobrovoljc, Kaja</name>
</author>
<author>
<name>Erjavec, Tomaž</name>
</author>
<author>
<name>Ljubešić, Nikola</name>
</author>
<author>
<name>Ponikvar, Primož</name>
</author>
<author>
<name>Šinkec, Mihael</name>
</author>
<author>
<name>Krek, Simon</name>
</author>
<id>http://hdl.handle.net/11356/2331</id>
<updated>2026-08-11T12:13:45Z</updated>
<published>2026-08-11T00:00:00Z</published>
<summary type="text">Monitor corpus of Slovene Trendi 2026-07
Kosem, Iztok; Čibej, Jaka; Dobrovoljc, Kaja; Erjavec, Tomaž; Ljubešić, Nikola; Ponikvar, Primož; Šinkec, Mihael; Krek, Simon
The Trendi corpus is a monitor corpus of Slovenian. It contains news articles from 106 media websites, published by 64 publishers. Trendi 2026-07 covers the period from January 2019 to July 2026, complementing the Gigafida 2.2 reference corpus of written Slovene (http://hdl.handle.net/11356/2106).&#13;
&#13;
The contents of the Trendi corpus are obtained using the Jožef Stefan Institute Newsfeed service (http://newsfeed.ijs.si/). The texts have been annotated using the CLASSLA-Stanza pipeline (https://github.com/clarinsi/classla), including syntactic parsing according to the Universal Dependencies (https://universaldependencies.org/sl/) and Named Entities (https://nl.ijs.si/janes/wp-content/uploads/2017/09/SlovenianNER-eng-v1.1.pdf).&#13;
&#13;
An important addition are topics or thematical categories, which have been automatically assigned to each text. There are 13 categories altogether: Arts and culture, Crime and accidents, Economy, Environment, Health, Leisure, Politics and Law, Science and Technology, Society, Sports, Weather, Entertainment, and Education. The text classification uses the following models: Text classification model SloBERTa-Trendi-Topics 1.0 (http://hdl.handle.net/11356/1709), Text classification model fastText-Trendi-Topics 1.0 (http://hdl.handle.net/11356/1710), and the SloBERTa model (https://huggingface.co/cjvt/sloberta-trendi-topics).&#13;
&#13;
The corpus is currently not available as a downloadable dataset due to copyright restrictions but we hope to make at least some of it available in the near future. The corpus is accessible through CLARIN.SI concordancers. If you would like to use the dataset for research purposes, please contact Iztok Kosem (iztok.kosem@ijs.si).&#13;
&#13;
This version adds texts from July 2026.
</summary>
<dc:date>2026-08-11T00:00:00Z</dc:date>
</entry>
<entry>
<title>List of Slovenian pseudowords</title>
<link href="http://hdl.handle.net/11356/2200" rel="alternate"/>
<author>
<name>Pavlič, Matic</name>
</author>
<author>
<name>Perdih, Andrej</name>
</author>
<author>
<name>Stepanov, Artur</name>
</author>
<id>http://hdl.handle.net/11356/2200</id>
<updated>2026-08-10T07:58:25Z</updated>
<published>2026-08-07T00:00:00Z</published>
<summary type="text">List of Slovenian pseudowords
Pavlič, Matic; Perdih, Andrej; Stepanov, Artur
The dataset contains 17,164 computer-generated Slovenian pseudowords empirically normed through an online lexical-decision task conducted from November 2024 to September 2025. The pseudowords were generated using Wuggy (Keuleers &amp; Brysbaert, 2010; https://github.com/WuggyCode/wuggy), based on source words drawn from the Dictionary of the Slovenian Standard Language, 2nd Edition (SSKJ2), the Dictionary of the Slovenian Standard Language, 3rd Edition (eSSKJ), and the Growing Dictionary of the Slovenian Language, as described by Perdih et al. (2026).&#13;
&#13;
A total of 39,064 unique participants completed 48,661 sessions, each comprising 84 words and 36 pseudowords. Although 19,742 pseudowords were included in the task, the final dataset contains 17,164 items following further data cleaning (see Pavlič &amp; Perdih, 2026).&#13;
&#13;
For each pseudoword, the dataset provides mean response time and false-alarm rate, along with length, syllable structure, number of syllables, proportion of consonants, source word, and two measures of orthographic neighbourhood density (OLD1 and OLD20). Syllabified forms of both the pseudoword and its source word are also provided.&#13;
&#13;
LITERATURE&#13;
Keuleers, E., &amp; Brysbaert, M. (2010). Wuggy: A multilingual pseudoword generator. Behavior Research Methods, 42(3), 627–633. https://doi.org/10.3758/BRM.42.3.627&#13;
Pavlič, M., &amp; Perdih, A. (2026). 17.164 slovenskih psevdobesed: empirično normiranje in psiholingvistični opis. Jezikovne tehnologije in digitalna humanistika : zbornik konference. [Accepted for publication].&#13;
Perdih, A., Gabrovšek, D., &amp; Pavlič, M. (2025). Izdelava seznama besed za množično raziskavo razširjenosti slovenskih besed. Slavistična revija, 73(1), 121–138. https://doi.org/10.57589/srl.v73i1.4231
</summary>
<dc:date>2026-08-07T00:00:00Z</dc:date>
</entry>
<entry>
<title>Corpus of the ZRC SAZU Terminology Consulting Service TermSvet 1.0</title>
<link href="http://hdl.handle.net/11356/2262" rel="alternate"/>
<author>
<name>Atelšek, Simon</name>
</author>
<author>
<name>Fajfar, Tanja</name>
</author>
<author>
<name>Jemec Tomazin, Mateja</name>
</author>
<author>
<name>Oman, Jera</name>
</author>
<author>
<name>Trojar, Mitja</name>
</author>
<author>
<name>Žagar Karer, Mojca</name>
</author>
<author>
<name>Zupan, Anja</name>
</author>
<author>
<name>Erjavec, Tomaž</name>
</author>
<id>http://hdl.handle.net/11356/2262</id>
<updated>2026-07-30T07:58:59Z</updated>
<published>2026-08-29T00:00:00Z</published>
<summary type="text">Corpus of the ZRC SAZU Terminology Consulting Service TermSvet 1.0
Atelšek, Simon; Fajfar, Tanja; Jemec Tomazin, Mateja; Oman, Jera; Trojar, Mitja; Žagar Karer, Mojca; Zupan, Anja; Erjavec, Tomaž
The Terminology Consulting Service of the Fran Ramovš Institute of the Slovenian Lanuage at ZRC SAZU has been active since 2013. It is intended especially for subject field experts, but also for other users who encounter terminological problems, which can be either the search for a term for a new concept that has not yet been named in Slovenian language, or the selection of the most suitable term among several designations for a concept. The compilation of each terminological answer is a joint work of all terminologists in the Department of Terminology. Each terminologist prepares their own opinion opinion on a terminological question, based on which a joint opinion is formed. Terminological answers are prepared taking into account terminological principles, i.e. from the point of view of terminological theory.&#13;
&#13;
The questions and answers of Terminology Consulting Service are available from its Web page and are internally stored in a database. This entry contains all the questions and answers of the Service from 2011-04-11 to 2026-02-26 formatted as a TEI XML document, which has been automatically converted from the database dump in TSV and XML formats.&#13;
&#13;
The corpus is available in three variants:&#13;
&#13;
- The base TEI corpus file structured into divisions each containing a question and answer. The corpus preserves basic formatting and hyperlinks in the texts. A portion of the divisions are classified into fields according to the supplied taxonomy and the authorship of those that prepared the answer is given.&#13;
&#13;
- The linguistically annotated TEI corpus, which discards all mark-up below the level of paragraph, and which has been marked up with Universal Dependencies morphological features and syntactic dependencies (and lemmatised) with the CLASSLA toolchain (https://github.com/clarinsi/classla).&#13;
&#13;
- The corpus as a vertical file, automatically converted from the linguistically annotated TEI, which is appropriate for use in CQP-type concordancers, such as noSketch Engine.
</summary>
<dc:date>2026-08-29T00:00:00Z</dc:date>
</entry>
<entry>
<title>Pragmatics understanding benchmark for Czech, Slovenian and Croatian PragMega CzeSloCro</title>
<link href="http://hdl.handle.net/11356/2261" rel="alternate"/>
<author>
<name>Vintar, Špela</name>
</author>
<author>
<name>Brglez, Mojca</name>
</author>
<author>
<name>Potočnjak, Mirna</name>
</author>
<author>
<name>Žižkova, Hana</name>
</author>
<author>
<name>Sangawa Hmeljak, Nina</name>
</author>
<id>http://hdl.handle.net/11356/2261</id>
<updated>2026-07-23T07:16:25Z</updated>
<published>2026-07-21T00:00:00Z</published>
<summary type="text">Pragmatics understanding benchmark for Czech, Slovenian and Croatian PragMega CzeSloCro
Vintar, Špela; Brglez, Mojca; Potočnjak, Mirna; Žižkova, Hana; Sangawa Hmeljak, Nina
PragMega CzeSloCro is a translation and adaptation of a section of the PragMega dataset (Floyd et al., 2026) into Czech, Slovenian and Croatian. The original dataset was manually crafted by psychologists and aimed at discovering whether "pragmatic inferencing" is a result of a single cognitive skill or, on the contrary, of different dissociable skills depending on the type of phenomena encountered. &#13;
A Slovenian version of the dataset was created first, for which we selected three tasks: Irony, Metaphor, and Humour. These consist of 50, 30, and 25 examples, respectively, or 105 examples in total. The Slovenian benchmark is described in Brglez &amp; Vintar (2026) and is implemented as SloPragMega on the SloBench (https://slobench.cjvt.si) platform. &#13;
&#13;
The dataset was then translated into Croatian and Czech by students of Digital Linguistics within a student project, then thoroughly revised by two professional linguists. Due to the highly nuanced and culturally specific nature of the dataset, some tasks were completely rewritten or replaced by more naturally sounding examples in the respective language.&#13;
&#13;
For each language, the dataset is divided into 3 subfolders (Metaphor, Irony, Humor), which contain the following files:&#13;
- stim.csv: The actual localised tasks with possible answers,&#13;
- stim_en.csv: The original English tasks with possible answers,&#13;
- stimOrder.csv: Template to create a randomized test with the order of questions and answers reshuffled,&#13;
- keys.csv: Solutions for the shuffled tasks. &#13;
&#13;
References:&#13;
Brglez, M., &amp; Vintar, S. (2026). From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene. In The Fourth Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL 2026) (pp. 44–54). European Language Resources Association (ELRA). https://doi.org/10.63317/4bpncy453r9k&#13;
&#13;
Floyd, S., Gibson, E., Fedorenko, E., &amp; Poliak, M. (2026, January 14). PragMega. https://doi.org/10.17605/OSF.IO/DPGE6
</summary>
<dc:date>2026-07-21T00:00:00Z</dc:date>
</entry>
<entry>
<title>Speech-level sentiment dataset of Slovenian parliamentary debates ParlaSent-SI 1.0</title>
<link href="http://hdl.handle.net/11356/2256" rel="alternate"/>
<author>
<name>Meden, Katja</name>
</author>
<author>
<name>Logar, Tamara</name>
</author>
<id>http://hdl.handle.net/11356/2256</id>
<updated>2026-07-09T10:04:34Z</updated>
<published>2026-06-25T00:00:00Z</published>
<summary type="text">Speech-level sentiment dataset of Slovenian parliamentary debates ParlaSent-SI 1.0
Meden, Katja; Logar, Tamara
The dataset comprises 1,000 manually annotated full utterances (i.e., speeches) from the parliamentary proceedings of Slovenia, extracted from the ParlaMint-SI 4.1 corpus (http://hdl.handle.net/11356/1912). The manual annotation campaign closely follows the setup used for the ParlaSent 1.0 multilingual sentiment dataset of parliamentary debates (http://hdl.handle.net/11356/1868), which provides sentiment annotations at sentence level.&#13;
&#13;
The ParlaSent-SI instances were randomly sampled and each speech was independently annotated by two trained annotators. The annotators underwent extensive training and also participated in the sentence-level sentiment annotation for the ParlaSent 1.0 dataset. The six-level annotation schema, originally based on the framework proposed by Batanović et al. (2020, DOI: https://doi.org/10.1371/journal.pone.0242050), was retained from the sentence-level annotation campaign and only minimally adapted in wording to suit full-utterance annotation:&#13;
&#13;
• Positive for utterances that are predominantly positive &#13;
• Negative for utterances that are predominantly negative &#13;
• M_Positive for utterances that convey an ambiguous sentiment or a mixture of sentiments, but lean more towards the positive sentiment &#13;
• M_Negative for utterances that convey an ambiguous sentiment or a mixture of sentiments, but lean more towards the negative sentiment &#13;
• P_Neutral for utterances that only contain non-sentiment-related statements, but still lean more towards the positive sentiment &#13;
• N_Neutral for utterances that only contain non-sentiment-related statements, but still lean more towards the negative sentiment.&#13;
&#13;
The final annotation for each utterance was determined in a separate reconciliation session, where the annotators reviewed their disagreements and agreed on the final tag. The 3-class labels (Positive, Negative, Neutral) are also provided.&#13;
&#13;
The dataset includes both procedural (i.e., those spoken by the session chair) and non-procedural parliamentary utterances. Procedural utterances are indicated in the "chair" column. Inter-annotator agreement (Krippendorff’s α) is reported for the full dataset and the non-procedural subset:&#13;
&#13;
6-class schema: 0.724 (full dataset), 0.570 (non-procedural subset)  &#13;
3-class schema: 0.852 (full dataset), 0.744 (non-procedural subset)&#13;
&#13;
The datasets are provided in both TSV and JSON formats and contain the initial annotations, annotator comments, procedural/non-procedural flag, flag for hard cases and the reconciled final label for 6- and 3-class sentiment annotation.
</summary>
<dc:date>2026-06-25T00:00:00Z</dc:date>
</entry>
<entry>
<title>Genus (proximum) in the SSKJ2 dictionary senses</title>
<link href="http://hdl.handle.net/11356/2254" rel="alternate"/>
<author>
<name>Perdih, Andrej</name>
</author>
<author>
<name>Bizjak Končar, Aleksandra</name>
</author>
<author>
<name>Divjak Race, Duša</name>
</author>
<author>
<name>Gabrovšek, Dejan</name>
</author>
<author>
<name>Ježovnik, Janoš</name>
</author>
<author>
<name>Krvina, Domen</name>
</author>
<author>
<name>Ledinek, Nina</name>
</author>
<author>
<name>Michelizza, Mija</name>
</author>
<author>
<name>Mirtič, Tanja</name>
</author>
<author>
<name>Petric Žižić, Špela</name>
</author>
<author>
<name>Sušnik, Miha</name>
</author>
<author>
<name>Trojar, Mitja</name>
</author>
<id>http://hdl.handle.net/11356/2254</id>
<updated>2026-06-30T08:46:36Z</updated>
<published>2026-06-22T00:00:00Z</published>
<summary type="text">Genus (proximum) in the SSKJ2 dictionary senses
Perdih, Andrej; Bizjak Končar, Aleksandra; Divjak Race, Duša; Gabrovšek, Dejan; Ježovnik, Janoš; Krvina, Domen; Ledinek, Nina; Michelizza, Mija; Mirtič, Tanja; Petric Žižić, Špela; Sušnik, Miha; Trojar, Mitja
The datasets contain sense–genus combinations from the Dictionary of the Slovenian Standard Language, 2nd Edition (Slovar slovenskega knjižnega jezika, druga, dopolnjena in deloma prenovljena izdaja; https://www.fran.si/133/sskj2-slovar-slovenskega-knjiznega-jezika-2). Genus is defined as a word denoting a broad, general category or superordinate class to which a defined word belongs. In the current version, 48,028 noun senses with 3,985 genera are included. Genera were attributed automatically and manually curated.&#13;
The first dataset (SSKJ2_headword_genus.xml) is focused on senses. Each dictionary sense contains the following information: headword or subheadword, entry ID, sense ID and one or more genera.&#13;
The second dataset (SSKJ2_genusGroups.xml) is focused on genera. One or more dictionary senses are attributed to each genus; for each dictionary sense, the following information are provided: headword or subheadword, entry ID and sense ID. No distinction between genera has been made with regard to homographs and homonyms.&#13;
For both XML files, the corresponding XML schemas are provided.&#13;
In rare cases, adjectival headwords are included, when the sense pertains to a multi-word unit containing an adjective and a substantive. Similarly, some noun senses are excluded, if they pertain to non-nominal phrases or are defined only by synonyms.&#13;
In the current version, words such as vsak, vsaka, vsako, and del, which form syntactic heads, are treated as genera. All genera are single-word units, even in cases where multi-word units would be expected.
</summary>
<dc:date>2026-06-22T00:00:00Z</dc:date>
</entry>
<entry>
<title>AI-generated text corpus AI-GenT 1.0</title>
<link href="http://hdl.handle.net/11356/2210" rel="alternate"/>
<author>
<name>Terčon, Luka</name>
</author>
<author>
<name>Dobrovoljc Zor, Kaja</name>
</author>
<id>http://hdl.handle.net/11356/2210</id>
<updated>2026-06-24T12:47:40Z</updated>
<published>2026-06-24T00:00:00Z</published>
<summary type="text">AI-generated text corpus AI-GenT 1.0
Terčon, Luka; Dobrovoljc Zor, Kaja
The AI-Generated Text (AI-GenT) corpus is a collection of English and Slovenian texts generated by several large language models. The corpus has been used in comparisons to collections of human-written texts in order to investigate the linguistic characteristics of the language generated by LLMs. &#13;
&#13;
The current version of the corpus contains texts that were constructed based on two preexisting human-written text corpora: the Šolar 3.0 corpus of Slovenian student essays (http://hdl.handle.net/11356/1589) and the LOCNESS corpus of English native speaker student essays (provided by the Centre for English Corpus Linguistics (CECL) at Université catholique de Louvain in Belgium - https://www.learnercorpusassociation.org/resources/tools/locness-corpus/). Three different LLMs—GPT-5 (https://developers.openai.com/api/docs/models/gpt-5), GaMS-27B (https://huggingface.co/cjvt/GaMS-27B-Instruct), and gemma-2-27b (https://huggingface.co/google/gemma-2-27b-it)—were instructed to produce corresponding texts to the texts in the human-written corpora using prompts containing information about the topic and length of the desired output. The AI-generated texts were produced by taking various subsets of the original human-written corpora as the basis for constructing the input prompts. For a full overview of the data, model, and prompt type combinations used to generate the AI-generated texts, please refer to the included AI-GenT_structure.png file which includes a full visual representation of the corpus structure.&#13;
&#13;
The corpus contains the AI-generated texts both in the form of raw text files as well as in the CoNLL-U file format containing grammatical annotations following the UD system of annotation (https://universaldependencies.org/). UD annotations were generated using the Trankit NLP pipeline (https://aclanthology.org/2021.eacl-demos.10/) with the default model used for English and a custom model used for Slovenian that is retrained on UD v2.15 data (http://hdl.handle.net/11356/1997). In the future, the corpus is planned to be extended with additional AI-generated news articles and Wikipedia articles.
</summary>
<dc:date>2026-06-24T00:00:00Z</dc:date>
</entry>
<entry>
<title>Slovene instruction-following safety dataset for large language models GaMS-Instruct-SAFE 0.5</title>
<link href="http://hdl.handle.net/11356/2218" rel="alternate"/>
<author>
<name>Čibej, Jaka</name>
</author>
<author>
<name>Kos, Sara</name>
</author>
<author>
<name>Kastelic, Maja</name>
</author>
<author>
<name>Gabrovšek, Dejan</name>
</author>
<author>
<name>Trojar, Mitja</name>
</author>
<author>
<name>Ježovnik, Janoš</name>
</author>
<author>
<name>Bizjak Končar, Aleksandra</name>
</author>
<author>
<name>Krvina, Domen</name>
</author>
<author>
<name>Petric Žižić, Špela</name>
</author>
<author>
<name>Divjak Race, Duša</name>
</author>
<author>
<name>Vreš, Domen</name>
</author>
<id>http://hdl.handle.net/11356/2218</id>
<updated>2026-06-01T11:01:50Z</updated>
<published>2026-06-01T00:00:00Z</published>
<summary type="text">Slovene instruction-following safety dataset for large language models GaMS-Instruct-SAFE 0.5
Čibej, Jaka; Kos, Sara; Kastelic, Maja; Gabrovšek, Dejan; Trojar, Mitja; Ježovnik, Janoš; Bizjak Končar, Aleksandra; Krvina, Domen; Petric Žižić, Špela; Divjak Race, Duša; Vreš, Domen
GaMS-Instruct-SAFE is a an instruction-following safety dataset designed to fine-tune Slovene large language models to provide safe responses (i.e. to train them to refuse responding to prompts that could lead to physical, economic or psychological harm). It consists of pairs of prompts and responses with various safety topics (e.g. sexual harassment, terrorism, violent crime, drugs). The prompts were written by human annotators using LabelStudio (Tkachenko et al. 2025) based on provided set of criteria (such as topic, expected prompt length, language standardness, different jailbreak strategies) to make the dataset as varied as possible (see Čibej 2024 for more details).&#13;
&#13;
In version 0.5, the responses to the prompts were generated using GaMS-27B-Instruct-Nemotron (https://huggingface.co/cjvt/GaMS-27B-Instruct-Nemotron). Only prompt-response pairs in which the model refused to cooperate were included. More responses will be added in future versions.&#13;
&#13;
The annotations for this dataset were created using Label Studio, open-source data labeling software developed by Heartex (Tkachenko et al. 2025).&#13;
&#13;
References:&#13;
Čibej, Jaka, 2024: First steps toward the compilation of a safety dataset for Slovene large language models. Jezikovne tehnologije in digitalna humanistika. https://repozitorij.uni-lj.si/IzpisGradiva.php?lang=slv&amp;id=164271 &#13;
&#13;
Tkachenko, Maxim, Mikhail Malyuk, Andrey Holmanyuk, Nikolai Liubimov, 2025: Label Studio: Data labeling software. https://github.com/HumanSignal/label-studio
</summary>
<dc:date>2026-06-01T00:00:00Z</dc:date>
</entry>
</feed>
