DESIGN OF AN AUTOMATED CLASSIFICATION SYSTEM USING INDOBERT TRANSFORMERS AND LOCAL NEWS TEXT SUMMARIZATION WITH LLAMA 3 ON RADARTEGAL.COM

Authors

  • Keisya Anazwa Octa Reviandy Universitas Islam Sultan Agung
  • Sam Farisa Chaerul Haviana Universitas Islam Sultan Agung

Abstract

In the digital news production process, editorial teams face challenges in manually categorizing news articles and generating summaries, which are time-consuming, inefficient, and prone to inconsistencies that affect content management quality and Search Engine Optimization (SEO). This study aims to design and develop an automated system for news classification and summarization on the radartegal.com platform using the Transformer-based IndoBERT model for automatic news category classification and the Large Language Model (LLM) LLaMA 3 for abstractive text summarization. The research methodology consisted of problem identification, literature review, dataset collection, text preprocessing, IndoBERT fine-tuning, LLaMA 3 implementation, pipeline integration, and system evaluation. Classification performance was evaluated using accuracy, precision, recall, and F1-score, while summarization quality was evaluated using ROUGE metrics. Experimental results showed that the IndoBERT model achieved an accuracy of 84.00%, precision of 84.58%, recall of 84.00%, and F1-score of 83.95%. Meanwhile, the LLaMA 3 summarization module achieved ROUGE-1, ROUGE-2, and ROUGE-L scores of 0.4672, 0.2732, and 0.4122, respectively. The integrated system successfully automated editorial workflows, improved categorization consistency, generated informative summaries, and supported SEO optimization. These findings demonstrate that the proposed system can improve editorial efficiency while maintaining content quality in local digital news publishing.

References

M. Adita Widianti and M. Budi Srikandi, “Netizen journalism and the democratization of local news: A case study of @abouttng on Instagram in Tangerang, Indonesia,” Channel, vol. 13, no. 1, pp. 62–75, 2025, doi: 10.12928/channel.

G. Airlangga, “Comparative analysis of machine learning models for chronic disease indicator classification using U.S. chronic disease indicators dataset,” MALCOM: Indonesian Journal of Machine Learning and Computer Science, vol. 4, no. 3, pp. 1034–1042, 2024, doi: 10.57152/malcom.v4i3.1403.

D. Al Akhdaan, T. E. Sutanto, and M. Liebenlito, “Confident Learning pada IndoBERT: Peningkatan kinerja klasifikasi sentimen,” The Indonesian Journal of Computer Science, vol. 13, no. 5, 2024, doi: 10.33022/ijcs.v13i5.4401.

A. Auriemma Citarella, M. Barbella, M. G. Ciobanu, F. De Marco, L. Di Biasi, and G. Tortora, “Assessing the effectiveness of ROUGE as unbiased metric in extractive vs. abstractive summarization techniques,” Journal of Computational Science, vol. 87, 2025, doi: 10.1016/j.jocs.2025.102571.

A. Bailly, A. Saubin, G. Kocevar, and J. Bodin, “Divide and summarize: Improve SLM text summarization,” Frontiers in Artificial Intelligence, vol. 8, 2025, doi: 10.3389/frai.2025.1604034.

T. B. Brown et al., “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165, 2020.

J. Federico Tantoro and I. D. M. B. A. Darmawan, “Klasifikasi berita berdasarkan kategori menggunakan convolutional neural network dengan IndoBERT,” JNATIA, vol. 3, no. 4, 2025.

L. Gao, X. Ma, J. Lin, and J. Callan, “Precise zero-shot dense retrieval without relevance labels,” arXiv preprint arXiv:2212.10496, 2022.

N. M. Gardazi, A. Daud, M. K. Malik, A. Bukhari, T. Alsahfi, and B. Alshemaimri, “BERT applications in natural language processing: A review,” Artificial Intelligence Review, vol. 58, no. 6, 2025, doi: 10.1007/s10462-025-11162-5.

A. Grattafiori et al., “The Llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024.

N. H. Masruri and D. A. Nihayati, “Optimasi SEO on-page pada long-tail keyword untuk meningkatkan visibilitas website di dalam SERP,” JRST (Jurnal Riset Sains dan Teknologi), vol. 6, no. 2, pp. 141–150, 2022, doi: 10.30595/jrst.v6i2.13071.

Y. Maulana, “Analisis dampak penggunaan content management system (CMS) open source terhadap industri portal berita media online,” Jurnal Algoritma, vol. 21, no. 1, pp. 1–8, 2024, doi: 10.33364/algoritma/v.21-1.1141.

M. Mazeika, B. Li, and D. Forsyth, “How to steer your adversary: Targeted and efficient model stealing defenses with gradient redirection,” arXiv preprint arXiv:2206.14157, 2022.

G. Michel, E. V. Epure, R. Hennequin, and C. Cerisara, “Evaluating LLMs for quotation attribution in literary texts: A case study of LLaMA3,” arXiv preprint arXiv:2406.11380, 2025.

T. Mshvidobadze, “Python for automating machine learning tasks,” JINAV: Journal of Information and Visualization, vol. 2, no. 2, pp. 77–82, 2021, doi: 10.35877/454ri.jinav373.

F. N. A. Ionendri, F. Candra, and A. Rizal, “News classification using natural language processing with TF-IDF and multinomial naïve Bayes,” Journal of Applied Computer Science and Technology, vol. 6, no. 1, pp. 37–45, 2025, doi: 10.52158/jacost.v6i1.1099.

A. F. Sonni, H. Hafied, I. Irwanto, and R. Latuheru, “Digital newsroom transformation: A systematic review of the impact of artificial intelligence on journalistic practices, news narratives, and ethical challenges,” Journalism and Media, vol. 5, no. 4, pp. 1554–1570, 2024, doi: 10.3390/journalmedia5040097.

M. Tahmid, R. Laskar, X.-Y. Fu, C. Chen, and S. Bhushan, “Building real-world meeting summarization systems using large language models: A practical perspective.”

T. Y. Tandi, T. F. Abidin, and H. Riza, “Incorporation of IndoBERT and machine learning features to improve the performance of Indonesian textual entailment recognition,” Journal of Information Systems Engineering and Business Intelligence, vol. 11, no. 2, pp. 173–186, 2025, doi: 10.20473/jisebi.11.2.173-186.

H. Touvron et al., “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.

R. Uddin, A. Basit, Y. Khan, S. J. Shazib, and S. Hossain, “BERT-based fake news detection: A transformer-driven approach for misinformation classification on Twitter,” International Journal on Science and Technology.

R. F. Wijaya, Z. Syahputra, Khairul, I. Purnama, and R. S. Hardinata, “Strategi pengelolaan konten dengan search engine optimization pada website,” Jurnal Komputer Teknologi Informasi Sistem Informasi (JUKTISI), vol. 4, no. 2, pp. 1417–1422, 2025, doi: 10.62712/juktisi.v4i2.680.

B. Wilie et al., “IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding,” arXiv preprint arXiv:2009.05387, 2020.

H. Zhang and M. O. Shafiq, “Survey of transformers and towards ensemble learning using transformers for natural language processing,” Journal of Big Data, vol. 11, no. 1, 2024, doi: 10.1186/s40537-023-00842-0.

J. Zhang, Y. Zhao, M. Saleh, and P. J. Liu, “PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization,” arXiv preprint arXiv:1912.08777, 2020.

X. Zhang, X. Xie, L. Ma, X. Du, Q. Hu, Y. Liu, J. Zhao, and M. Sun, “Towards characterizing adversarial defects of deep learning software from the lens of uncertainty,” in Proceedings of the International Conference on Software Engineering, 2020, pp. 739–751, doi: 10.1145/3377811.3380368.

Downloads

Published

2026-07-16

How to Cite

Reviandy, K. A. O., & Sam Farisa Chaerul Haviana. (2026). DESIGN OF AN AUTOMATED CLASSIFICATION SYSTEM USING INDOBERT TRANSFORMERS AND LOCAL NEWS TEXT SUMMARIZATION WITH LLAMA 3 ON RADARTEGAL.COM. Journal of Data Analytics, Information, and Computer Science, 3(3), 154–168. Retrieved from https://journal.ppmi.web.id/index.php/jdaics/article/view/3659

Issue

Section

Articles