Vytautas Magnus University Research Management System (VDU CRIS)





Use this url to cite department: https://hdl.handle.net/20.500.12259/148370
Now showing 1 - 10 of 39
  • Item type:Dataset,
    Colloc – a tool for automatic identification of multiword expressions
    [Colloc - įrankis automatiniam pastoviųjų žodžių junginių nustatymui]
    dataset[2019][H004,N009]; ; ; ; ; ; ;
    Vilkaitė-Lozdienė, Laura
    Baltijos pažangių technologijų institutas / Baltic Institute of Advanced Technology, 2019

    Colloc -- a tool for automatic identification of multiword expressions (MWE) is freely available for online use at http://resursai.mwe.lt/atpazintuvas. As material for training DELFI.lt corpus (http://tekstynas.mwe.lt/) was used. For identification combination of 2 trained models (RNN bi-LSTM and CRF) is used. Automatically identified MWE can be retrieved in 2 formats -- list of MWE or / and text with annotated MWE.

      78
  • Item type:Dataset,
    Corpus of discourse on crime
    [Nusikalstamumo diskurso tekstynas]
    dataset[2020][H004,N009]
    Vytauto Didžiojo universitetas / Vytautas Magnus University, 2020-07-09

    Specialised "Corpus of Discourse on Crime" is synchronic, monolingual, unannotated, consists of two subcorpora. Subcorpus 1: all texts on crime, published in criminal columns on the most popular Lithuanian web portals (15min.lt, delfi.lt, lrytas.lt) and other sources (police websites, specialized newspaper Akistata). Period: from 6 September 2015 to 5 October 2015. Size: 329,227 tokens. Subcorpus 2: texts on two crimes that have caused widespread public resonance: (a) Suspect AB (15min.lt, delfi.lt, lrytas.lt, tv3.lt, individual authors, police websites, Akistata). Period: from 2 January 2016 to 28 January 2016. Size: 32,915 tokens. (b) Suspect GK (15min.lt, delfi.lt, lrytas.lt, tv3.lt, police websites). Period: from 26 January 2017 to 14 March 2017. Size: 48,849 tokens. The selection of the texts meets the criteria of readability and accessibility. The principle of random selection was less relevant, as the corpus included all the texts on crimes published in the scheduled sources. The 2nd subcorpus consists of texts on crimes against children and provides the opportunity to obtain more emotional evaluation data. The analysis of this corpus is presented in the dissertation by S. Jakimovienė “Evaluation in the Discourse on Crime: Identification and Interpretation”.

      83  2
  • Item type:Dataset,
    Corpus of the contemporary Lithuanian language
    [Dabartinės lietuvių kalbos tekstynas]

    Corpus of the Contemporary Lithuanian Language, which comprises 208 million words, is a collection of texts designed to represent the current Lithuanian. The corpus has been compiled since 1990. The corpus is designed to represent as wide a range of contemporary written Lithuanian as possible. The largest part of the corpus is comprised of General Press (texts from regional and national newspapers), Popular Press, and Special Press (specialized newspapers and magazines). The rest of the corpus consists of Fiction, Nonfiction, Administrative documents, and Spoken language. The corpus is morphologically annotated and freely accessible for online search at http://corpus.vdu.lt.

      211
  • Item type:Dataset,
    DELFI.lt corpus
    [DELFI.lt tekstynas]
    dataset[2019][H004,N009]; ; ; ; ; ; ;
    Vilkaitė-Lozdienė, Laura
    Baltijos pažangių technologijų institutas / Baltic Institute of Advanced Technology, 2019

    DELFI.lt is corpus made of articles published by news portal DELFI.lt since March 2014 till November 2016. Metadata was collected with articles as well: author, title, date, source, link, category, number of words. This corpus is made of 190 000 news articles from 12 thematic categories: DELFI Faces (DELFI Veidai), Projects (Projektai), DELFI Science (DELFI Mokslas), DELFI Auto, Unidentified category, Sport, DELFI Life (DELFI Gyvenimas), DELFI People (DELFI Žmonės), DELFI CItizen (DELFI Pilietis), Business (Verslas), DELFI FIT, DELFI News (DELFI Žinios). All in all DELFI.lt corpus consists of 70 million words. The corpus is morphologically annotated with Universal Dependencies tags and is freely accessible for online search at http://tekstynas.mwe.lt/.

      97
  • Item type:Dataset,
    DIGIRES COVID-19 Corpus v.1
    [DIGIRES COVID-19 tekstynas v.1]
    dataset[2023][H004,N009]; ;
    Meidutė, Aistė
    ;
    Vytauto Didžiojo universitetas / Vytautas Magnus University, 2023-02-20

    DIGIRES COVID-19 Corpus v.1 consists of 351 Lithuanian media articles about COVID-19 pandemics. The corpus was compiled from various internet public Lithuanian media sources. Corpus contains 351 files in plain text format (TXT) with UTF-8 encoding. Each article consists of a title (in the 1st line) and an article body. Files are classified into two subcorpora: 1) "unrealiable" that contains articles, which were identified by professional fact checkers as fake news; 2) "reliable" that contains trustworthy articles. Subcorpus Files Word tokens Reliable: 175 67902 Unreliable: 176 118747 Total 351 186649

      127  2
  • Item type:Dataset,
    DIGIRES COVID-19 ML Dataset v.1
    [DIGIRES COVID-19 ML duomenų rinkinys v.1]
    dataset[2023][H004,N009]; ;
    Meidutė, Aistė
    ;
    Vytauto Didžiojo universitetas / Vytautas Magnus University, 2023-02-20

    DIGIRES COVID-19 ML dataset v.1 is a tab-separated (.tsv) file prepared for training machine learning algorithms. The training dataset was compiled from various internet public Lithuanian media sources. It contains 351 records and has the following attributes: "Title": the title of a news article "Text": the text of the article "Label": a label that marks the article as 1: unreliable; 0: reliable 1) "unrealiable" marks articles, which were identified by professional fact checkers as fake news; 2) "reliable" marks trustworthy articles. Classes Labels Word tokens Reliable: 175 67902 Unreliable: 176 118747 Total 351 186649

      131  2
  • survey data[2024][H004,N009]
    Dviskaitos vertimo ypatumai : lyginamasis tekstynais grįstas tyrimas, 2024-11-04

    The resource offers two data sets: concordances of dual pronoun translations from Lithuanian into English (942 concordance lines) and translations of English pronouns into Lithuanian dual forms (1590 concordance lines). Parallel concordances were retrieved from: 1) The Lithuanian-English Corpus of Prose (LECOP). This corpus contains Lithuanian fiction and its English translations from 1990 to 2009. It includes 95 texts by 43 authors, translated by 39 translators, with a total word count 682,936. For more detailed information about the corpus, see Vaičenonienė, J. (2011). Lithuanian Literature in English: A Corpus-Based Approach to the Translation of Author-Specific Neologisms (Doctoral dissertation). Kaunas: VDU. 2) The English-Lithuanian Parallel Corpus available online at: https://sitti.vdu.lt/lygiagretus-tekstynas/. This corpus comprises 31 prose texts by 16 authors, with a total word count 1,199,730. For more detailed information about the resource, refer to: Vaičenonienė, Jurgita. (2024). “Dviskaitos vertimo ypatumai : lyginamasis tekstynais grįstas tyrimas.” Darnioji daugiakalbystė : periodinis mokslo žurnalas = Sustainable multilingualism: Biannual scientific journal 24. https://hdl.handle.net/20.500.12259/266695. This resource is valuable for generating activities for trainee translators or language editors and provides useful material for research on dual pronoun translation.

      49
  • Item type:Dataset,
    English-Lithuanian comparable cybersecurity corpus - DVITAS
    [Anglų–lietuvių kalbų palyginamasis tekstynas - DVITAS]
    dataset[2022][H004,N009];
    Rackevičienė, Sigita
    ;
    ; ;
    Mockienė, Liudmila
    ;
    Laurinaitis, Marius
    Vytauto Didžiojo universitetas, 2022-02-05

    The English-Lithuanian comparable corpus (DVITAS COMPARABLE) is morphologically annotated. It includes English and Lithuanian original texts on cybersecurity from the time period of 2010-2021. The corpus was compiled for the bilingual terminology extraction project together with English-Lithuanian parallel corpus. There are 1,708 files in English and 2,567 for Lithuanian. The total size of the corpus is 4m words (EN-2m; LT-2m) The corpus is composed of texts representing 4 text types: academic (EN-19%; LT-30%), administrative-informative (EN-8%; LT-11%), legal (EN-18%; LT-4%), media (EN-55%; LT-55%).

      123
  • Item type:Dataset,
    English-Lithuanian parallel cybersecurity corpus - DVITAS
    [Anglų-lietuvių lygiagretusis kibernetinio saugumo tekstynas DVITAS]
    dataset[2022][H004,N009];
    Rackevičienė, Sigita
    ;
    ; ;
    Mockienė, Liudmila
    ;
    Laurinaitis, Marius
    Vytauto Didžiojo universitetas, 2022-02-05

    Anglų-lietuvių lygiagretusis tekstynas DVITAS susideda iš originalių anglų kalba parašytų tekstų apie kibernetinį saugumą ir jų vertimų į lietuvių kalbą. Tekstai sukurti laikotarpyje nuo 2006 m. iki 2021 m. Šis tekstynas kartu su anglų-lietuvių palyginamuoju tekstynu buvo sukurti projekto „Dvikalbis automatinis terminų atpažinimas“ reikmėms. Duomenų rinkinys susideda iš 80 sulygiuotų failų TMX formate, kurie yra pusiau automatiškai sulygiuoti sakinio lygmenyje ir 160 nelygiuotų pradinių failų (80-anglų kalbos ir 80 lietuvių kalbos). Tekstyno dydis yra 1,4 mln. žodžių (EN-0,77 mln. ir LT-0,63 mln), 35415 sulygiuoti sakiniai.

      121  6
  • survey data[2024][H004,N009]
    Mickevič, Jolanta
    ;
    ;
    Rackevičienė, Sigita
    ;
    ; ;
    Mockienė, Liudmila
    ;
    Laurinaitis, Marius
    Vytauto Didžiojo universitetas / Vytautas Magnus University, 2024-12-31

    English-Lithuanian parallel corpus DVITAS v2 includes original English texts on cybersecurity and their Lithuanian translations aligned on the sentence level. Version 1 of the corpus was compiled for the bilingual terminology extraction project DVITAS together with English-Lithuanian comparable corpus. The current 2nd version of the corpus features expansion of the 1st version containing additional 27 files and metadata information. The parallel corpus includes the EU legal acts and other documents from the time period of 2006-2022. The documents have been extracted from the EUR-Lex database and other EU institutional repositories. There are 107 aligned files in TMX format in English and Lithuanian, as well as 214 raw files (107 in English, and 107 in Lithuanian) within the dataset. The total size of the corpus is 1.97m words (EN-1.08m; LT-0.88m). The corpus contains 53,792 aligned segments.

      64