Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry
The digitization of large archival collections of oral folk poetry in Finland and Estonia has opened possibilities for large-scale quantitative studies of intertextuality. As an initial methodological step in this direction, I present a method for pairwise line-by-line comparison of poems using the...
Main Author: | |
---|---|
Format: | Article |
Language: | English |
Published: |
Nicolas Turenne
2023-08-01
|
Series: | Journal of Data Mining and Digital Humanities |
Online Access: | https://jdmdh.episciences.org/11390/pdf |
_version_ | 1797269963778555904 |
---|---|
author | Maciej Janicki |
author_facet | Maciej Janicki |
author_sort | Maciej Janicki |
collection | DOAJ |
description | The digitization of large archival collections of oral folk poetry in Finland and Estonia has opened possibilities for large-scale quantitative studies of intertextuality. As an initial methodological step in this direction, I present a method for pairwise line-by-line comparison of poems using the weighted sequence alignment algorithm (a.k.a. ‘weighted edit distance’). The main contribution of the paper is a novel description of the algorithm in terms of matrix operations, which allows for much faster alignment of a poem against the entire corpus by utilizing modern numeric libraries and GPU capabilities. This way we are able to compute pairwise alignment scores between all pairs from among a corpus of over 280,000 poems. The resulting table of over 40 million pairwise poem similarities can be used in various
ways to study the oral tradition. Some starting points for such research are sketched in the latter part of the article. |
first_indexed | 2024-03-11T21:06:05Z |
format | Article |
id | doaj.art-c25e65692d3846cf8e38aa28e1ebb0da |
institution | Directory Open Access Journal |
issn | 2416-5999 |
language | English |
last_indexed | 2024-04-25T01:56:44Z |
publishDate | 2023-08-01 |
publisher | Nicolas Turenne |
record_format | Article |
series | Journal of Data Mining and Digital Humanities |
spelling | doaj.art-c25e65692d3846cf8e38aa28e1ebb0da2024-03-07T16:17:03ZengNicolas TurenneJournal of Data Mining and Digital Humanities2416-59992023-08-01NLP4DH10.46298/jdmdh.1139011390Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetryMaciej Janicki0https://orcid.org/0000-0003-3981-8021University of HelsinkiThe digitization of large archival collections of oral folk poetry in Finland and Estonia has opened possibilities for large-scale quantitative studies of intertextuality. As an initial methodological step in this direction, I present a method for pairwise line-by-line comparison of poems using the weighted sequence alignment algorithm (a.k.a. ‘weighted edit distance’). The main contribution of the paper is a novel description of the algorithm in terms of matrix operations, which allows for much faster alignment of a poem against the entire corpus by utilizing modern numeric libraries and GPU capabilities. This way we are able to compute pairwise alignment scores between all pairs from among a corpus of over 280,000 poems. The resulting table of over 40 million pairwise poem similarities can be used in various ways to study the oral tradition. Some starting points for such research are sketched in the latter part of the article.https://jdmdh.episciences.org/11390/pdf |
spellingShingle | Maciej Janicki Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry Journal of Data Mining and Digital Humanities |
title | Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry |
title_full | Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry |
title_fullStr | Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry |
title_full_unstemmed | Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry |
title_short | Large-scale weighted sequence alignment for the study of intertextuality in Finnic oral folk poetry |
title_sort | large scale weighted sequence alignment for the study of intertextuality in finnic oral folk poetry |
url | https://jdmdh.episciences.org/11390/pdf |
work_keys_str_mv | AT maciejjanicki largescaleweightedsequencealignmentforthestudyofintertextualityinfinnicoralfolkpoetry |