Identification of Parallel Passages Across a Large Hebrew/Aramaic Corpus

We propose a method for efficiently finding all parallel passages in a large corpus, even if the passages are not quite identical due to rephrasing and orthographic variation. The key ideas are the representation of each word in the corpus by its two most infrequent letters, finding matched pairs of...

Full description

Bibliographic Details
Main Authors: Avi Shmidman, Moshe Koppel, Ely Porat
Format: Article
Language:English
Published: Nicolas Turenne 2018-03-01
Series:Journal of Data Mining and Digital Humanities
Subjects:
Online Access:https://jdmdh.episciences.org/1388/pdf