LogShrink: Effective Log Compression by Leveraging Commonality and Variability of Log Data

LogShrink: эффективное сжатие журналов за счёт использования их общих и вариативных характеристик
Xiaoyun Li, Hongyu Zhang, Van-Hoang Le, Pengfei Chen
2023-09-18

clustering-based sequence samplingcommonality and variabilityentropy analysislog compressionlongest common subsequence
Log data is a crucial resource for recording system events and states during system execution. However, as systems grow in scale, log data generation has become increasingly explosive, leading to an expensive overhead on log storage, such as several petabytes per day in production. To address this issue, log compression has become a crucial task in reducing disk storage while allowing for further log analysis. Unfortunately, existing general-purpose and log-specific compression methods have been limited in their ability to utilize log data characteristics. To overcome these limitations, we conduct an empirical study and obtain three major observations on the characteristics of log data that can facilitate the log compression task. Based on these observations, we propose LogShrink, a novel and effective log compression method by leveraging commonality and variability of log data. An analyzer based on longest common subsequence and entropy techniques is proposed to identify the latent commonality and variability in log messages. The key idea behind this is that the commonality and variability can be exploited to shrink log data with a shorter representation. Besides, a clustering-based sequence sampler is introduced to accelerate the commonality and variability analyzer. The extensive experimental results demonstrate that LogShrink can exceed baselines in compression ratio by 16% to 356% on average while preserving a reasonable compression speed.
1
A clustering-based sequence sampler accelerates the commonality and variability analysis process.
2
A longest-common-subsequence and entropy-based analyzer identifies latent commonality and variability in log messages.
3
An empirical study identifies three log-data characteristics that can be exploited to improve compression effectiveness.
4
Experiments show LogShrink exceeds baseline compression ratios by 16% to 356% on average while maintaining reasonable compression speed.
5
LogShrink leverages commonality and variability in log messages to produce shorter compressed representations than existing methods.

log data generated during system execution

the commonality and variability of log data and their exploitation for compression, including compression ratio and speed

Publication Details
Publication Date
2023-09-18
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Xiaoyun Li
Hongyu Zhang
Van-Hoang Le
Pengfei Chen
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%