No More Labelled Examples? An Unsupervised Log Parser with LLMs

Больше никаких размеченных примеров? Неразмеченный парсер логов на основе больших языковых моделей
Junjie Huang, Zhihan Jiang, Zhuangbin Chen, Michael R. Lyu
2025-06-19

Log Contrastive Unitshybrid rankinglarge language modelslog structure extractionunsupervised log parsing
Log parsing serves as an essential prerequisite for various log analysis tasks. Recent advancements in this field have improved parsing accuracy by leveraging the semantics in logs through fine-tuning large language models (LLMs) or learning from in-context demonstrations. However, these methods heavily depend on high-quality labeled examples to achieve optimal performance. In practice, continuously collecting high-quality labeled data is challenging since logs are huge in volume and under frequent evolution, leading to performance degradation or heavy maintenance efforts for existing log parsers after deployment. To address this issue, we propose LUNAR, an unsupervised LLM-based method for efficient and off-the-shelf log parsing. Our key insight is that while LLMs may struggle with direct log parsing, their performance can be significantly enhanced through comparative analysis across multiple logs that differ only in their parameter parts. We refer to such groups of logs as Log Contrastive Units (LCUs) . Given the vast volume of logs, obtaining LCUs is difficult. Therefore, LUNAR introduces a hybrid ranking scheme to effectively search for LCUs by jointly considering the commonality and variability among logs. Additionally, LUNAR crafts a novel parsing prompt for LLMs to identify contrastive patterns and extract meaningful log structures from LCUs. Experiments on large-scale public and industrial log datasets demonstrate that LUNAR significantly outperforms state-of-the-art log parsers in terms of accuracy and efficiency, providing an effective and practical solution for real-world deployment.
1
A contrastive parsing prompt enables LLMs to identify parameter differences and extract meaningful log structures from retrieved units.
2
Experiments on large-scale public and industrial datasets show that LUNAR significantly outperforms state-of-the-art parsers in both accuracy and efficiency.
3
LUNAR is an unsupervised, off-the-shelf LLM-based log parser that eliminates reliance on labeled examples for deployment.
4
LUNAR uses a hybrid ranking scheme combining log commonality and variability to efficiently retrieve Log Contrastive Units from large log collections.
5
The method improves LLM log parsing by comparatively analyzing logs that share templates but differ in parameter values, called Log Contrastive Units.

logs and log parsing systems in large-scale public and industrial environments

unsupervised extraction of meaningful log structures through contrastive analysis of logs differing in parameter parts, with emphasis on parsing accuracy and efficiency

Publication Details
Publication Date
2025-06-19
Journal
Publisher
ISSN
Cited by
19
Access Type
Author Information
Authors
Junjie Huang
Zhihan Jiang
Zhuangbin Chen
Michael R. Lyu
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%