Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry

The digitization of large archival collections of oral folk poetry in Finland and Estonia has opened possibilities for large-scale quantitative studies of intertextuality. As an initial methodological step in this direction, I present a method for pairwise line-by-line comparison of poems using the...

Full description

Bibliographic Details
Main Author: Maciej Janicki
Format: Article
Language:English
Published: Nicolas Turenne 2023-08-01
Series:Journal of Data Mining and Digital Humanities
Online Access:https://jdmdh.episciences.org/11390/pdf
_version_ 1797269963778555904
author Maciej Janicki
author_facet Maciej Janicki
author_sort Maciej Janicki
collection DOAJ
description The digitization of large archival collections of oral folk poetry in Finland and Estonia has opened possibilities for large-scale quantitative studies of intertextuality. As an initial methodological step in this direction, I present a method for pairwise line-by-line comparison of poems using the weighted sequence alignment algorithm (a.k.a. ‘weighted edit distance’). The main contribution of the paper is a novel description of the algorithm in terms of matrix operations, which allows for much faster alignment of a poem against the entire corpus by utilizing modern numeric libraries and GPU capabilities. This way we are able to compute pairwise alignment scores between all pairs from among a corpus of over 280,000 poems. The resulting table of over 40 million pairwise poem similarities can be used in various ways to study the oral tradition. Some starting points for such research are sketched in the latter part of the article.
first_indexed 2024-03-11T21:06:05Z
format Article
id doaj.art-c25e65692d3846cf8e38aa28e1ebb0da
institution Directory Open Access Journal
issn 2416-5999
language English
last_indexed 2024-04-25T01:56:44Z
publishDate 2023-08-01
publisher Nicolas Turenne
record_format Article
series Journal of Data Mining and Digital Humanities
spelling doaj.art-c25e65692d3846cf8e38aa28e1ebb0da2024-03-07T16:17:03ZengNicolas TurenneJournal of Data Mining and Digital Humanities2416-59992023-08-01NLP4DH10.46298/jdmdh.1139011390Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetryMaciej Janicki0https://orcid.org/0000-0003-3981-8021University of HelsinkiThe digitization of large archival collections of oral folk poetry in Finland and Estonia has opened possibilities for large-scale quantitative studies of intertextuality. As an initial methodological step in this direction, I present a method for pairwise line-by-line comparison of poems using the weighted sequence alignment algorithm (a.k.a. ‘weighted edit distance’). The main contribution of the paper is a novel description of the algorithm in terms of matrix operations, which allows for much faster alignment of a poem against the entire corpus by utilizing modern numeric libraries and GPU capabilities. This way we are able to compute pairwise alignment scores between all pairs from among a corpus of over 280,000 poems. The resulting table of over 40 million pairwise poem similarities can be used in various ways to study the oral tradition. Some starting points for such research are sketched in the latter part of the article.https://jdmdh.episciences.org/11390/pdf
spellingShingle Maciej Janicki
Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry
Journal of Data Mining and Digital Humanities
title Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry
title_full Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry
title_fullStr Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry
title_full_unstemmed Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry
title_short Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry
title_sort large scale weighted sequence alignment for the study of intertextuality in finnic oral folk poetry
url https://jdmdh.episciences.org/11390/pdf
work_keys_str_mv AT maciejjanicki largescaleweightedsequencealignmentforthestudyofintertextualityinfinnicoralfolkpoetry