Hybrid LexRank-LDA-MMR for Indonesian Text Summarization
Home Research Details
Nasrul Amin Muis, Yoga Pristyanto, Ika Nur Fajri

Hybrid LexRank-LDA-MMR for Indonesian Text Summarization

0.0 (0 ratings)

Introduction

Hybrid lexrank-lda-mmr for indonesian text summarization. Enhance Indonesian text summarization using a hybrid LexRank-LDA-MMR model. This study achieves higher ROUGE scores, producing more informative, relevant, and diverse summaries.

0
4 views

Abstract

The rapid growth of digital text information makes it crystal clear that there is a need for automated tools that summarize text for rapid retrieval. Extractive methods employed include LexRank, Latent Dirichlet Allocation (LDA), and Maximal Marginal Relevance (MMR), and the study aimed to enhance the quality of Indonesian text summaries beyond regular LexRank. In this study, the role of LexRank was to assist in selecting meaningful sentences that were centric to the center of the graphs, while the role of LDA was to ensure that the sentences were topically relevant. The strength of MMR lies in maintaining the document's relevance and diversity, thereby reducing redundancy in the summaries. Summaries from two publicly available datasets, IndoSum and Liputan6, containing texts in Bahasa Indonesia, were analyzed at 30% and 50% compression levels and graded using ROUGE (ROUGE-1, ROUGE-2, ROUGE-L F1) scores. Analysis of 5000 articles per dataset showed that implementing LexRank and LDA together with MMR resulted in a higher average ROUGE score than standard LexRank, irrespective of the set compression levels and across both datasets, demonstrating the approach's effectiveness in enhancing summary quality. The improvements recorded are most significant in ROUGE-1 and ROUGE-2, indicating that these combination approaches can produce more informative and relevant summaries while preserving sentence-level diversity, thereby deepening understanding of the information presented in the summary.


Review

This paper presents a timely and relevant study on enhancing automated text summarization, specifically for Indonesian language texts, a domain with growing informational needs. Addressing the demand for rapid information retrieval from vast digital text, the authors propose a novel hybrid extractive summarization method combining LexRank, Latent Dirichlet Allocation (LDA), and Maximal Marginal Relevance (MMR). The core objective is to move beyond the limitations of single-method approaches like standard LexRank, aiming to produce summaries that are not only relevant but also diverse and non-redundant. This innovative fusion of techniques promises a significant step forward in generating higher-quality summaries for Bahasa Indonesia. The methodological approach is well-articulated, detailing the synergistic roles of each component within the hybrid framework. LexRank is employed to identify salient sentences central to the document's structure, while LDA ensures the selected sentences maintain strong topical coherence, thereby grounding the summary in the main themes. The integration of MMR is crucial for mitigating redundancy and enhancing diversity, ensuring that the final summary provides a comprehensive yet concise overview without unnecessary repetition. The study rigorously evaluates this approach using two established Indonesian datasets, IndoSum and Liputan6, across 30% and 50% compression levels, with performance gauged by ROUGE-1, ROUGE-2, and ROUGE-L F1 scores. The findings strongly support the effectiveness of the proposed hybrid model, demonstrating a consistent and superior performance over standard LexRank across both datasets and compression levels. The reported improvements are particularly pronounced in ROUGE-1 and ROUGE-2 scores, which highlights the method's ability to generate summaries with higher precision in content and better sentence-level similarity to human-generated references. This suggests that the combined LexRank-LDA-MMR approach successfully produces more informative, topically relevant, and diverse summaries, thereby deepening the user's understanding of the source text. The study contributes significantly to the field of natural language processing, particularly in low-resource language summarization, and offers a robust model for future research in hybrid extractive methods.


Full Text

You need to be logged in to view the full text and Download file of this article - Hybrid LexRank-LDA-MMR for Indonesian Text Summarization from Jurnal Nasional Teknologi dan Sistem Informasi .

Login to View Full Text And Download

Comments


You need to be logged in to post a comment.