Text Mining from Free Unstructured Text: An Experiment of Time Series Retrieval for Volcano Monitoring

Volcanic activity may influence climate parameters and impact people safety, and hence monitoring its characteristic indicators and their temporal evolution is crucial. Several databases, communications and literature providing data, information and updates on active volcanoes worldwide are availabl...

Full description

Bibliographic Details
Main Authors: Margherita Berardi, Luigi Santamaria Amato, Francesca Cigna, Deodato Tapete, Mario Siciliani de Cumis
Format: Article
Language:English
Published: MDPI AG 2022-03-01
Series:Applied Sciences
Subjects:
Online Access:https://www.mdpi.com/2076-3417/12/7/3503
_version_ 1797440439445356544
author Margherita Berardi
Luigi Santamaria Amato
Francesca Cigna
Deodato Tapete
Mario Siciliani de Cumis
author_facet Margherita Berardi
Luigi Santamaria Amato
Francesca Cigna
Deodato Tapete
Mario Siciliani de Cumis
author_sort Margherita Berardi
collection DOAJ
description Volcanic activity may influence climate parameters and impact people safety, and hence monitoring its characteristic indicators and their temporal evolution is crucial. Several databases, communications and literature providing data, information and updates on active volcanoes worldwide are available, and will likely increase in the future. Consequently, information extraction and text mining techniques aiming to efficiently analyze such databases and gather data and parameters of interest on a specific volcano can play an important role in this applied science field. This work presents a natural language processing (NLP) system that we developed to extract geochemical and geophysical data from free unstructured text included in monitoring reports and operational bulletins issued by volcanological observatories in HTML, PDF and MS Word formats. The NLP system enables the extraction of relevant gas parameters (e.g., SO<sub>2</sub> and CO<sub>2</sub> flux) from the text, and was tested on a series of 2839 daily and weekly bulletins published online between 2015 and 2021 for the Stromboli volcano (Italy). The experiment shows that the system proves capable in the extraction of the time series of a set of user-defined parameters that can be later analyzed and interpreted by specialists in relation with other monitoring and geospatial data. The text mining system can potentially be tuned to extract other target parameters from this and other databases.
first_indexed 2024-03-09T12:07:11Z
format Article
id doaj.art-ef7c4890d6094d9398c5252a4e60cae4
institution Directory Open Access Journal
issn 2076-3417
language English
last_indexed 2024-03-09T12:07:11Z
publishDate 2022-03-01
publisher MDPI AG
record_format Article
series Applied Sciences
spelling doaj.art-ef7c4890d6094d9398c5252a4e60cae42023-11-30T22:56:36ZengMDPI AGApplied Sciences2076-34172022-03-01127350310.3390/app12073503Text Mining from Free Unstructured Text: An Experiment of Time Series Retrieval for Volcano MonitoringMargherita Berardi0Luigi Santamaria Amato1Francesca Cigna2Deodato Tapete3Mario Siciliani de Cumis4Italian Space Agency (ASI), Space Center “G. Colombo”, Località Terlecchia s.n.c., 75100 Matera, ItalyItalian Space Agency (ASI), Space Center “G. Colombo”, Località Terlecchia s.n.c., 75100 Matera, ItalyItalian Space Agency (ASI), Via del Politecnico s.n.c., 00133 Roma, ItalyItalian Space Agency (ASI), Via del Politecnico s.n.c., 00133 Roma, ItalyItalian Space Agency (ASI), Space Center “G. Colombo”, Località Terlecchia s.n.c., 75100 Matera, ItalyVolcanic activity may influence climate parameters and impact people safety, and hence monitoring its characteristic indicators and their temporal evolution is crucial. Several databases, communications and literature providing data, information and updates on active volcanoes worldwide are available, and will likely increase in the future. Consequently, information extraction and text mining techniques aiming to efficiently analyze such databases and gather data and parameters of interest on a specific volcano can play an important role in this applied science field. This work presents a natural language processing (NLP) system that we developed to extract geochemical and geophysical data from free unstructured text included in monitoring reports and operational bulletins issued by volcanological observatories in HTML, PDF and MS Word formats. The NLP system enables the extraction of relevant gas parameters (e.g., SO<sub>2</sub> and CO<sub>2</sub> flux) from the text, and was tested on a series of 2839 daily and weekly bulletins published online between 2015 and 2021 for the Stromboli volcano (Italy). The experiment shows that the system proves capable in the extraction of the time series of a set of user-defined parameters that can be later analyzed and interpreted by specialists in relation with other monitoring and geospatial data. The text mining system can potentially be tuned to extract other target parameters from this and other databases.https://www.mdpi.com/2076-3417/12/7/3503text mininginformation extractionenvironmental monitoringvolcanic activitynatural language processing
spellingShingle Margherita Berardi
Luigi Santamaria Amato
Francesca Cigna
Deodato Tapete
Mario Siciliani de Cumis
Text Mining from Free Unstructured Text: An Experiment of Time Series Retrieval for Volcano Monitoring
Applied Sciences
text mining
information extraction
environmental monitoring
volcanic activity
natural language processing
title Text Mining from Free Unstructured Text: An Experiment of Time Series Retrieval for Volcano Monitoring
title_full Text Mining from Free Unstructured Text: An Experiment of Time Series Retrieval for Volcano Monitoring
title_fullStr Text Mining from Free Unstructured Text: An Experiment of Time Series Retrieval for Volcano Monitoring
title_full_unstemmed Text Mining from Free Unstructured Text: An Experiment of Time Series Retrieval for Volcano Monitoring
title_short Text Mining from Free Unstructured Text: An Experiment of Time Series Retrieval for Volcano Monitoring
title_sort text mining from free unstructured text an experiment of time series retrieval for volcano monitoring
topic text mining
information extraction
environmental monitoring
volcanic activity
natural language processing
url https://www.mdpi.com/2076-3417/12/7/3503
work_keys_str_mv AT margheritaberardi textminingfromfreeunstructuredtextanexperimentoftimeseriesretrievalforvolcanomonitoring
AT luigisantamariaamato textminingfromfreeunstructuredtextanexperimentoftimeseriesretrievalforvolcanomonitoring
AT francescacigna textminingfromfreeunstructuredtextanexperimentoftimeseriesretrievalforvolcanomonitoring
AT deodatotapete textminingfromfreeunstructuredtextanexperimentoftimeseriesretrievalforvolcanomonitoring
AT mariosicilianidecumis textminingfromfreeunstructuredtextanexperimentoftimeseriesretrievalforvolcanomonitoring