Schema Matching Using Federated Learning on a Hybrid Feature Set

Сопоставление схем с использованием федеративного обучения на гибридном наборе признаков
A. E. Shepelev, Sergey A. Stupnikov
2025-06-01

LSTM with attention and MLPfederated learninghybrid feature setname similarity metricsschema matching
To increase representativeness in data analysis, there is a need to extract and integrate data from various sources. This paper examines the application of machine learning to one of the main stages of data integration that is schema matching. A schema is the structure of a specific database, including the definition of relations, their attributes data types, and the definition of primary and foreign keys. The result of this stage is a matching of the elements of the source schemas and the elements of the target schema. A neural network model based on a combination of long short-term memory networks, attention mechanisms, and a multilayer perceptron is proposed. The model is trained using a hybrid set of features, including name similarity metrics, data types, tags, descriptive statistics, and correlation coefficients for numerical data. Experiments have been conducted showing that the proposed model outperforms the basic neural network model and classical schema matching methods. It is also shown that with federated learning, which preserves data privacy, the quality of the model is almost the same comparing to centralized learning.
1
A neural network combining LSTM, attention mechanisms, and a multilayer perceptron is proposed for schema matching.
2
Experiments show the proposed model outperforms a basic neural network model and classical schema matching methods.
3
Federated learning preserves data privacy while achieving model quality almost the same as centralized learning.
4
The model is trained on a hybrid feature set: name similarity metrics, data types, tags, descriptive statistics, and numerical correlation coefficients.

Schema matching for database schemas (matching elements of source and target schemas)

Performance of a neural-network-based schema matching approach using a hybrid feature set (name similarity, data types, tags, descriptive statistics, correlation coefficients) and its behaviour under federated learning versus centralized learning with respect to matching quality

Publication Details
Publication Date
2025-06-01
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
A. E. Shepelev
Sergey A. Stupnikov
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%