Show simple item record

 
dc.contributor.author Kosem, Iztok
dc.contributor.author Krek, Simon
dc.contributor.author Čibej, Jaka
dc.contributor.author Gantar, Polona
dc.contributor.author Arhar Holdt, Špela
dc.contributor.author Logar, Nataša
dc.contributor.author Laskowski, Cyprian
dc.contributor.author Klemenc, Bojan
dc.contributor.author Ljubešić, Nikola
dc.contributor.author Dobrovoljc, Kaja
dc.contributor.author Gorjanc, Vojko
dc.contributor.author Pori, Eva
dc.date.accessioned 2021-05-04T06:15:02Z
dc.date.available 2021-05-04T06:15:02Z
dc.date.issued 2020-10-26
dc.identifier.uri http://hdl.handle.net/11356/1424
dc.description The Orange workflow for observing collocation trends ColTrend 1.0 ColTrend is a workflow (.OWS file) for Orange Data Mining (an open-source machine learning and data visualization software: https://orangedatamining.com/) that allows the user to observe temporal collocation trends in corpora. The workflow consists of a series of Python scripts, data filters, and visualizers. As input, the workflow takes a .CSV file with data on collocations and their relative frequencies by year of publication extracted from a corpus. As output, it provides a .TSV file containing the same data (or a filtered selection thereof) enriched with four measures that indicate the collocation’s temporal trend in the corpus: (1) the slope (k) of a linear regression model fitted to the frequency data, which indicates whether the frequency of use of the collocation is increasing or declining; (2) the coefficient of determination (R2) of the linear regression model, indicating how linear the change in the collocation’s use is; (3) the ratio (m) of maximum relative frequency and average relative frequency, which indicates peaks in collocation usage; and (4) the coefficient of recent growth (t), which indicates an increased usage of the collocation in the last three years of the observed corpus data. The entry also contains three .CSV files that can be used to test the workflow. The files contain collocation candidates (along with their relative frequencies per year of publication) extracted from the Gigafida 2.0 Corpus of Written Slovene (https://viri.cjvt.si/gigafida/) with three different syntactic structures (as defined in http://hdl.handle.net/11356/1415): 1) p0-s0 (adjective + noun, e.g. rezervni sklad), 2) s0-s2 (noun + noun in the genitive case, e.g. ukinitev lastnine), and 3) gg-s4 (verb + noun in the accusative case, e.g. pripraviti besedilo). It should be noted that only collocation candidates with absolute frequency of 15 and above were extracted. Please note that the ColTrend workflow requires the installation of the Text Mining add-on for Orange. For installation instructions as well as a more detailed description of the different phases of the workflow and the measures used to observe the collocation trends, please consult the README file.
dc.publisher Centre for Language Resources and Technologies, University of Ljubljana
dc.rights Apache License 2.0
dc.rights.uri https://opensource.org/licenses/Apache-2.0
dc.rights.label PUB
dc.source.uri https://www.cjvt.si/kolos/
dc.subject collocations
dc.subject temporal trends
dc.subject relative frequency
dc.subject linear regression
dc.title The Orange workflow for observing collocation trends ColTrend 1.0
dc.type toolService
metashare.ResourceInfo#ContentInfo.detailedType tool
metashare.ResourceInfo#ResourceComponentType#ToolServiceInfo.languageDependent false
has.files yes
branding CLARIN.SI data & tools
contact.person Iztok Kosem iztok.kosem@cjvt.si Centre for Language Resources and Technologies, University of Ljubljana
sponsor ARRS (Slovenian Research Agency) J6-8255 Collocations as a basis for language description: semantic and temporal perspectives nationalFunds
sponsor ARRS (Slovenian Research Agency) P6-0411 Language Resources and Technologies for Slovene nationalFunds
sponsor ARRS (Slovenian Research Agency) J6-8256 New grammar of contemporary standard Slovene: sources and methods nationalFunds
files.count 1
files.size 73436084


 Files in this item

This item is
Publicly Available
and licensed under:
Apache License 2.0
Icon
Name
coltrend_1.0.zip
Size
70.03 MB
Format
application/zip
MD5
a5663ec7ed62547099a276ebc6e4ed58
 Download file  Preview
 File Preview  
  • coltrend_1.0
    • coltrend_gf2_csv
      • coltrend_gf2_34_p0-s0.csv172 MB
      • coltrend_gf2_23_gg-s4.csv65 MB
      • coltrend_gf2_53_s0-s2.csv117 MB
    • coltrend_1.0_workflow.ows50 kB
    • 00README.txt3 kB

Show simple item record