LogShrink: Effective Log Compression by Leveraging Commonality and Variability of Log Data

LogShrink: эффективное сжатие журналов за счёт использования общих и вариативных характеристик данных журналов
Xiaoyun Li, Hongyu Zhang, Pengfei Chen, Van-Hoang Le
2024-02-06

clustering-based sequence samplercommonality and variabilityentropy analysislog compressionlongest common subsequence
Log data is a crucial resource for recording system events and states during system execution. However, as systems grow in scale, log data generation has become increasingly explosive, leading to an expensive overhead on log storage, such as several petabytes per day in production. To address this issue, log compression has become a crucial task in reducing disk storage while allowing for further log analysis. Unfortunately, existing general-purpose and log-specific compression methods have been limited in their ability to utilize log data characteristics. To overcome these limitations, we conduct an empirical study and obtain three major observations on the characteristics of log data that can facilitate the log compression task. Based on these observations, we propose LogShrink, a novel and effective log compression method by leveraging commonality and variability of log data. An analyzer based on longest common subsequence and entropy techniques is proposed to identify the latent commonality and variability in log messages. The key idea behind this is that the commonality and variability can be exploited to shrink log data with a shorter representation. Besides, a clustering-based sequence sampler is introduced to accelerate the commonality and variability analyzer. The extensive experimental results demonstrate that LogShrink can exceed baselines in compression ratio by 16% to 356% on average while preserving a reasonable compression speed.
1
Across extensive experiments, LogShrink improves compression ratios over baseline methods by 16% to 356% on average while maintaining reasonable compression speed.
2
An empirical study identifies three log-data characteristics that can be exploited to improve compression effectiveness.
3
LogShrink analyzes latent commonality and variability in log messages using longest common subsequence and entropy techniques.
4
The method compresses logs through shorter representations of shared and variable content, with clustering-based sequence sampling accelerating analysis.

log data, particularly log messages generated during system execution

the commonality and variability of log data and their exploitation to improve compression ratio and compression efficiency

Publication Details
Publication Date
2024-02-06
Journal
Publisher
ISSN
Cited by
29
Access Type
Author Information
Authors
Xiaoyun Li
Hongyu Zhang
Pengfei Chen
Van-Hoang Le
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%