DOAJ Open Access 2023

Schema matching based on energy domain pre-trained language model

Zhiyu Pan Muchen Yang Antonello Monti

Abstrak

Abstract Data integration in the energy sector, which refers to the process of combining and harmonizing data from multiple heterogeneous sources, is becoming increasingly difficult due to the growing volume of heterogeneous data. Schema matching plays a crucial role in this process by giving each representation a unique identity by matching raw energy data to a generic data model. This study uses an energy domain language model to automate schema matching, reducing manual effort in integrating heterogeneous data. We developed two energy domain language models, Energy BERT and Energy Sentence Bert, and trained them using an open-source scientific corpus. The comparison of the developed models with the baseline model using real-life energy domain data shows that Energy BERT and Energy Sentence Bert models significantly improve the accuracy of schema matching.

Topik & Kata Kunci

Energy industries. Energy policy. Fuel trade

Penulis (3)

Zhiyu Pan

Muchen Yang

Antonello Monti

Format Sitasi

APA MLA BibTeX

Pan, Z., Yang, M., Monti, A. (2023). Schema matching based on energy domain pre-trained language model. https://doi.org/10.1186/s42162-023-00277-0

Akses Cepat

PDF tidak tersedia langsung

Cek di sumber asli →

Lihat di Sumber doi.org/10.1186/s42162-023-00277-0

Informasi Jurnal

Tahun Terbit: 2023
Sumber Database: DOAJ
DOI: 10.1186/s42162-023-00277-0
Akses: Open Access ✓