Integrating Semantic Clustering-Based Recommendation Models and RAG Generative AI Chatbots for Regional Innovation Analytics
DOI:
https://doi.org/10.70609/gtech.v10i4.11030Keywords:
Agglomerative Clustering, Innovation Clustering, Regional Innovation, Retrieval-Augmented GenerationAbstract
Regional research and innovation agencies are responsible for monitoring large volumes of innovation data submitted by local government units. Non-standardized submissions, manually matched collaboration opportunities, and a growing data request burden have left this data underutilized for policymaking. This study aimed to design, implement, and evaluate an integrated platform for regional innovation monitoring at the Regional Research and Innovation Agency (BRIDA) of East Java Province, using a design and development research approach. The platform combines a clustering based recommendation model that groups semantically similar innovations using multilingual sentence embeddings, an AI chatbot built on Retrieval-Augmented Generation (RAG) for querying the innovation database, and descriptive analytical reports on innovation maturity. Tested on 643 innovation records, the platform achieved a 100 percent functional pass rate across 35 test scenarios, while the chatbot achieved mean faithfulness and relevance scores of 4.0 out of 5.0 under an LLM as a judge evaluation. These results suggest that embedding based clustering can reveal previously hidden collaboration opportunities, while a generative chatbot can meaningfully reduce manual search workload. To the authors' knowledge, this is the first platform integrating clustering based recommendation with RAG based querying for regional innovation monitoring, offering a replicable design for similar settings.
References
Alagarsamy, S., Tantithamthavorn, C., & Aleti, A. (2024). A3Test: Assertion-augmented automated test case generation. Information and Software Technology, 176, 107565. https://doi.org/10.1016/j.infsof.2024.107565 DOI: https://doi.org/10.1016/j.infsof.2024.107565
Aminah, S., & Wardani, D. K. (2018). Readiness analysis of regional innovation implementation. Jurnal Bina Praja, 10(1), 13–26. https://doi.org/10.21787/jbp.10.2018.13-26 DOI: https://doi.org/10.21787/jbp.10.2018.13-26
Artetxe, M., & Schwenk, H. (2019). Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. Transactions of the Association for Computational Linguistics, 7, 597–610. https://doi.org/10.1162/tacl_a_00288 DOI: https://doi.org/10.1162/tacl_a_00288
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., Ye, W., Zhang, Y., Chang, Y., Yu, P. S., Yang, Q., & Xie, X. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3), Article 39. https://doi.org/10.1145/3641289 DOI: https://doi.org/10.1145/3641289
Directorate General of Fiscal Balance. (2020). Kebijakan inovasi daerah [PDF]. Kementerian Keuangan Republik Indonesia & Kementerian Dalam Negeri Republik Indonesia. https://djpk.kemenkeu.go.id/wp-content/uploads/2020/11/Salinan-kebijakan-inovasi-daerah-DJPK.pdf
Hambarde, K. A., & Proença, H. (2023). Information retrieval: Recent advances and beyond. IEEE Access, 11, 76581–76604. https://doi.org/10.1109/ACCESS.2023.3295776 DOI: https://doi.org/10.1109/ACCESS.2023.3295776
Ikotun, A. M., Habyarimana, F., & Ezugwu, A. E. (2025). Cluster validity indices for automatic clustering: A comprehensive review. Heliyon, 11(2), e41953. https://doi.org/10.1016/j.heliyon.2025.e41953 DOI: https://doi.org/10.1016/j.heliyon.2025.e41953
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., & Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274 DOI: https://doi.org/10.1016/j.lindif.2023.102274
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474
Li, Z., Wang, Z., Wang, W., Hung, K., Xie, H., & Wang, F. L. (2025). Retrieval-augmented generation for educational application: A systematic survey. Computers and Education: Artificial Intelligence, 8, 100417. https://doi.org/10.1016/j.caeai.2025.100417 DOI: https://doi.org/10.1016/j.caeai.2025.100417
Papageorgiou, G., Sarlis, V., Maragoudakis, M., Magnisalis, I., & Tjortjis, C. (2025). Evaluating faithfulness in agentic RAG systems for e-governance applications using LLM-based judging frameworks. Big Data and Cognitive Computing, 9(12), 309. https://doi.org/10.3390/bdcc9120309 DOI: https://doi.org/10.3390/bdcc9120309
Pujiono, I., Agtyaputra, I. M., & Ruldeviyani, Y. (2024). Implementing retrieval-augmented generation and vector databases for chatbots in public services agencies context. JITK (Jurnal Ilmu Pengetahuan dan Teknologi Komputer), 10(1), 216–223. https://doi.org/10.33480/jitk.v10i1.5572 DOI: https://doi.org/10.33480/jitk.v10i1.5572
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 3982–3992. https://doi.org/10.18653/v1/D19-1410 DOI: https://doi.org/10.18653/v1/D19-1410
Reimers, N., & Gurevych, I. (2020). Making monolingual sentence embeddings multilingual using knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 4512–4525). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.365 DOI: https://doi.org/10.18653/v1/2020.emnlp-main.365
Riski, F., Auliasari, K., & Orisa, M. (2026). Job vacancy recommendation system based on text description analysis using word embedding and cosine similarity. bit-Tech, 8(3), 3718–3729. https://doi.org/10.32877/bt.v8i3.3690 DOI: https://doi.org/10.32877/bt.v8i3.3690
Saksono, H. (2020). Center for innovation: Collaborative media towards innovative local government. Nahkoda: Journal of Governance Science, 19(1), 1–16
Suherwin, Zainuddin, M., & Rachmat. (2026). Comparing FAQ-based and retrieval-augmented generation chatbots for academic service question answering. G-Tech: Jurnal Teknologi Terapan, 10(3), 1093–1105. https://doi.org/10.70609/g-tech.v10i3.10019 DOI: https://doi.org/10.70609/g-tech.v10i3.10019
Sun, X., Meng, Y., Ao, X., Wu, F., Zhang, T., Li, J., & Fan, C. (2022). Sentence similarity based on contexts. Transactions of the Association for Computational Linguistics, 10, 573–588. https://doi.org/10.1162/tacl_a_00477 DOI: https://doi.org/10.1162/tacl_a_00477
Swacha, J., & Gracel, M. (2025). Retrieval-Augmented Generation (RAG) chatbots for education: A survey of applications. Applied Sciences, 15(8), 4234. https://doi.org/10.3390/app15084234 DOI: https://doi.org/10.3390/app15084234
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., & Wen, J.-R. (2026). A survey of large language models. Frontiers of Computer Science, 20(12), 2012627. https://doi.org/10.1007/s11704-026-60308-3 DOI: https://doi.org/10.1007/s11704-026-60308-3
Zimazanya, B., Rizky Supriyatna, M. A., & Anwar, C. (2025). Functionality testing of web-based library information systems using the blackbox testing method using the ISO/IEC 29119 standard. Journal of Information Systems and Business Technology, 1(4), 85–93
Petukhova, A., Matos-Carvalho, J. P., & Fachada, N. (2025). Text clustering with large language model embeddings. International Journal of Cognitive Computing in Engineering, 6, 100–108. https://doi.org/10.1016/j.ijcce.2024.11.004 DOI: https://doi.org/10.1016/j.ijcce.2024.11.004
Sutrakar, V. K., & Mogre, N. (2025). An improved deep learning model for word embeddings based clustering for large text datasets. arXiv:2502.16139 DOI: https://doi.org/10.11648/j.mlr.20251001.14
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Alfi Fadliana, Dea Kayla Putri Darusman, Dinda Ayu Permatasari

This work is licensed under a Creative Commons Attribution 4.0 International License.









