Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.1

Name: Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.1
License: https://clarin.si/repository/xmlui/page/licence-aca-id-by-nc-inf-nored-1.0

Ljubešić, Nikola; Fišer, Darja; Erjavec, Tomaž; Šulc, Ajda

Prikaži enostavni zapis vnosa

dc.contributor.author	Ljubešić, Nikola
dc.contributor.author	Fišer, Darja
dc.contributor.author	Erjavec, Tomaž
dc.contributor.author	Šulc, Ajda
dc.date.accessioned	2021-11-18T09:54:10Z
dc.date.available	2021-11-18T09:54:10Z
dc.date.issued	2021-11-17
dc.identifier.uri	http://hdl.handle.net/11356/1462
dc.description	The FRENK dataset consists of comments to Facebook posts (news articles) of mainstream media outlets from Croatia, Great Britain, and Slovenia, on the topics of migrants and LGBT. The dataset contains whole discussion threads. Each comment is annotated by the type of socially unacceptable discourse (e.g., inappropriate, offensive, violent speech) and its target (e.g., migrants/LGBT, commenters, media). The annotation schema in its details is described in https://arxiv.org/pdf/1906.02045.pdf. Usernames in the metadata are pseudo-anonymised and removed from the comments. The data in each language (Croatian (hr), English (en), Slovenian (sl), and topic (migrants, LGBT) is divided into a training and a testing portion. The training and testing data consist of separate discussion threads, i.e., there is no cross-discussion-thread contamination between training and testing data. The sizes of the splits are the following: Croatian, migrants: 4356 training comments, 978 testing comments; Croatian LGBT: 4494 training comments, 1142 comments; English, migrants: 4540 training comments, 1285 testing comments; English, LGBT: 4819 training comments, 1017 testing comments; Slovenian, migrants: 5145 training comments, 1277 testing comments; Slovenian, LGBT: 2842 training comments, 900 testing comments. The difference to the first version of the dataset are the additions of 1. the annotation guidelines in English and 2. the link to the huggingface dataset.
dc.language.iso	hrv
dc.language.iso	eng
dc.language.iso	slv
dc.publisher	Jožef Stefan Institute
dc.relation.isreferencedby	https://arxiv.org/pdf/1906.02045.pdf
dc.relation.replaces	http://hdl.handle.net/11356/1433
dc.rights	CLARIN.SI Licence ACA ID-BY-NC-INF-NORED 1.0
dc.rights.uri	https://clarin.si/repository/xmlui/page/licence-aca-id-by-nc-inf-nored-1.0
dc.rights.label	ACA
dc.source.uri	http://nl.ijs.si/frenk/
dc.subject	offensive language
dc.subject	hate speech
dc.subject	news comments
dc.title	Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.1
dc.type	corpus
metashare.ResourceInfo#ContentInfo.mediaType	text
has.files	yes
branding	CLARIN.SI data & tools
demo.uri	https://huggingface.co/datasets/classla/FRENK-hate-hr
contact.person	Nikola Ljubešić nikola.ljubesic@ijs.si Jožef Stefan Institute
sponsor	ARRS (Slovenian Research Agency) J7-8280 FRENK: Resources, methods, and tools for the understanding, identification, and classification of various forms of socially unacceptable discourse in the information society nationalFunds
sponsor	ARRS (Slovenian Research Agency) N6-0099 LiLaH: Linguistic Landscape of Hate Speech nationalFunds
sponsor	ARRS (Slovenian Research Agency) P6-0411 Language Resources and Technologies for Slovene nationalFunds
size.info	32795 texts
files.count	2
files.size	4701352