Analysis of the Accuracy and Completeness of SINTA Author Data Extraction

Authors

  • Muhammad Arfah Asis Universitas Muslim Indonesia, Indonesia
  • Nia Kurniati Universitas Muslim Indonesia, Indonesia
  • Muhammad Alfarid Jufda Universitas Muslim Indonesia, Indonesia

DOI:

https://doi.org/10.70609/g-tech.v10i1.8832

Keywords:

Data completeness, Dynamic data ordering, Python, SINTA, Web scraping

Abstract

The advancement of information technology has increased the use of web scraping for scientific data collection, including from the SINTA (Science and Technology Index) platform, which provides researcher profiles, affiliations, publications, and citation data. However, scraping SINTA poses challenges, particularly when multiple authors share identical scores that trigger changes in display order. This instability can lead to duplicated or missing entries when using a single-pass scraping approach. This study evaluates the accuracy and completeness of SINTA author data collection by implementing repeated scraping as a strategy to handle dynamic data ordering. Experiments were conducted on the Universitas Muslim Indonesia (UMI) affiliation, targeting 915 active authors. The methodology involved page-structure analysis, spider development using Python and Scrapy, sequential scraping through pagination, and validation of data completeness and uniqueness. A three-second delay between requests was applied to maintain responsible scraping practices. The results show that a single scraping attempt failed to retrieve all authors, capturing an average of only 877.2 authors (95.86%). Due to unstable ordering, repeated iterations were required. Through 4–8 scraping cycles per trial, all 915 authors were successfully collected without duplication. These findings indicate that for platforms with dynamic data structures such as SINTA, repeated scraping provides a more reliable method for ensuring data completeness and accuracy, supporting the development of stable and responsible publication-data automation systems.

References

Adila, N. (2022). Implementation of Web Scraping for Journal Data Collection on the SINTA Website. Sinkron, 7(4), 2478–2485. https://doi.org/10.33395/sinkron.v7i4.11576 DOI: https://doi.org/10.33395/sinkron.v7i4.11576

Aidillah, K. H., Dodi Vionanda, & Dony Permana. (2025). Analysis on Scopus Articles Padang State University Based on SINTA Website. UNP Journal of Statistics and Data Science, 3(1), 79–85. https://doi.org/10.24036/ujsds/vol3-iss1/346 DOI: https://doi.org/10.24036/ujsds/vol3-iss1/346

Asis, M. A., Purnawansyah, P., & Salim, Y. (2024). Rancang Bangun Sistem Manajemen Data Akreditasi berbasis Web. Journal Cerita, 10(1), 32–38. https://doi.org/10.33050/cerita.v10i1.2989 DOI: https://doi.org/10.33050/cerita.v10i1.2989

Darmawan, I., Maulana, M., Gunawan, R., & Widiyasono, N. (2022). Evaluating Web Scraping Performance Using XPath, CSS Selector, Regular Expression, and HTML DOM With Multiprocessing Technical Applications. JOIV : International Journal on Informatics Visualization, 6(4), 904. https://doi.org/10.30630/joiv.6.4.1525 DOI: https://doi.org/10.30630/joiv.6.4.1525

Fahrudin, T. M., Funabiki, N., Brata, K. C., Naing, I., Aung, S. T., Muhaimin, A., & Prasetya, D. A. (2025). An Improved Reference Paper Collection System Using Web Scraping with Three Enhancements. Future Internet, 17(5), 195. https://doi.org/10.3390/fi17050195 DOI: https://doi.org/10.3390/fi17050195

Giuliana, B., Mariangela, L., & Marianna, L. (2024). Bibliometric Insights into Web Scraping and Advanced AI-Based Models for Valuable Business Data. Proceedings of the 26th International Conference on Enterprise Information Systems, 321–328. https://doi.org/10.5220/0012686900003690 DOI: https://doi.org/10.5220/0012686900003690

Goel, A., Zhu, J., Netravali, R., & Madhyastha, H. V. (2024). Sprinter: Speeding Up High-Fidelity Crawling of the Modern Web. NSDI’24: Proceedings of the 21st USENIX Symposium on Networked Systems Design and Implementation.

Guyt, J. Y., Datta, H., & Boegershausen, J. (2024). Unlocking the Potential of Web Data for Retailing Research. Journal of Retailing, 100(1), 130–147. https://doi.org/10.1016/j.jretai.2024.02.002 DOI: https://doi.org/10.1016/j.jretai.2024.02.002

Hidayat, M. K., Sugiarto, D., Fitriana, R., & Liang, Y.-C. (2025). Business intelligence system model to measure the performance of lecturers’ scientific publications. TELKOMNIKA (Telecommunication Computing Electronics and Control), 23(4), 954. https://doi.org/10.12928/telkomnika.v23i4.26221 DOI: https://doi.org/10.12928/telkomnika.v23i4.26221

Kazmali, A. S., & Sayar, A. (2025). Web Scraping: Legal and Ethical Considerations in General and Local Context - A Review. Procedia Computer Science, 259, 1563–1572. https://doi.org/10.1016/j.procs.2025.04.111 DOI: https://doi.org/10.1016/j.procs.2025.04.111

Kurniatuhadi, R., Hadary, F., Sulistyarini, S., Anzani, Y. M., Fahruna, Y., & Muthahhari, M. (2024). The Chronology of Development Tracer Study System at Tanjungpura University. Jurnal Teknologi Informasi Dan Pendidikan, 17(1), 218–228. https://doi.org/10.24036/jtip.v17i1.746 DOI: https://doi.org/10.24036/jtip.v17i1.746

Lengkong, J. S. J., Jacobus, S. N. H., Dondokambey, R., Ratumbuisang, K. F., Paath, D., & Liow, E. S. (2023). Web-Based Academic Information Systems in Vocational School. International Journal of Information Technology and Education, 2(4), 12–25. https://doi.org/10.62711/ijite.v2i4.153 DOI: https://doi.org/10.62711/ijite.v2i4.153

Logos, K., Brewer, R., Langos, C., & Westlake, B. (2023). Establishing a framework for the ethical and legal use of web scrapers by cybercrime and cybersecurity researchers: learnings from a systematic review of Australian research. International Journal of Law and Information Technology, 31(3), 186–212. https://doi.org/10.1093/ijlit/eaad023 DOI: https://doi.org/10.1093/ijlit/eaad023

Lotfi, C., Srinivasan, S., Ertz, M., & Latrous, I. (2021). Web Scraping Techniques and Applications: A Literature Review. In SCRS CONFERENCE PROCEEDINGS ON INTELLIGENT SYSTEMS (pp. 381–394). Soft Computing Research Society. https://doi.org/10.52458/978-93-91842-08-6-38 DOI: https://doi.org/10.52458/978-93-91842-08-6-38

Madjido, M., Espressivo, A., Maula, A. W., Fuad, A., & Hasanbasri, M. (2019). Health Information System Research Situation in Indonesia: A Bibliometric Analysis. Procedia Computer Science, 161, 781–787. https://doi.org/10.1016/j.procs.2019.11.183 DOI: https://doi.org/10.1016/j.procs.2019.11.183

Molina, I., Morales, J., & Keith, B. (2025). Web Scraping Chilean News Media: A Dataset for Analyzing Social Unrest Coverage (2019–2023). Data, 10(11), 174. https://doi.org/10.3390/data10110174 DOI: https://doi.org/10.3390/data10110174

Purnomo, L. M., & Ayub, M. (2021). Analisis Data Hasil Web Scraping untuk Menentukan Kualitas Jurnal Ilmiah. Jurnal Strategi, 3(1), 122–132.

Rahmatulloh, A., & Gunawan, R. (2020). Web Scraping with HTML DOM Method for Data Collection of Scientific Articles from Google Scholar. Indonesian Journal of Information Systems, 2(2), 95–104. https://doi.org/10.24002/ijis.v2i2.3029 DOI: https://doi.org/10.24002/ijis.v2i2.3029

Rochim, A. F., Nugraha, T., Widodo, A. P., Eridani, D., & Martono, K. T. (2021). Design data collection tool and weighting classification of authors in their scholar outputs based on percent-contribution-indicated (PCI) method. Journal of Physics: Conference Series, 1943(1), 012111. https://doi.org/10.1088/1742-6596/1943/1/012111 DOI: https://doi.org/10.1088/1742-6596/1943/1/012111

Sahria, Y. (2020). Implementasi Teknik Web Scraping pada Jurnal SINTA Untuk Analisis Topik Penelitian Kesehatan Indonesia. URECOL (Unversity Research Colloqium), 297–306.

Saniyah, A., & Ghozali, M. L. (2023). Implementation of the DSN-MUI Fatwa on Product Testimonials from Istihsan Perspective. EKSYAR : Jurnal Ekonomi Syari’ah & Bisnis Islam, 10(2), 223–232. https://doi.org/10.54956/eksyar.v10i2.474 DOI: https://doi.org/10.54956/eksyar.v10i2.474

Shah, A., Shah, H., Bafna, V., Khandor, C., & Nair, S. (2025). Validation and Extraction of Reliable Information Through Automated Scraping and Natural Language Inference. Engineering Applications of Artificial Intelligence, 147, 110284. https://doi.org/10.1016/j.engappai.2025.110284 DOI: https://doi.org/10.1016/j.engappai.2025.110284

Ulfah, A., & Najiah, I. (2023). Implementasi Web Scraping pada Situs Jurnal Sinta Menggunakan Framework Selenium Webdriver Python. JIKA (Jurnal Informatika), 7(1), 29. https://doi.org/10.31000/jika.v7i1.7037 DOI: https://doi.org/10.31000/jika.v7i1.7037

Yuda, M. A. D. (2025). Implementasi Web Scraping Untuk Ekstraksi Data Penjual dan Produk Panel Surya Di E-Commerce. KONSTELASI: Konvergensi Teknologi Dan Sistem Informasi, 5(1). https://doi.org/10.24002/konstelasi.v5i1.11751 DOI: https://doi.org/10.24002/konstelasi.v5i1.11751

Zbarcea, A., & Tudose, C. (2024). Migrating from Developing Asynchronous Multi-Threading Programs to Reactive Programs in Java. Applied Sciences, 14(24), 12062. https://doi.org/10.3390/app142412062 DOI: https://doi.org/10.3390/app142412062

Downloads

Published

2026-01-16

How to Cite

Analysis of the Accuracy and Completeness of SINTA Author Data Extraction. (2026). G-Tech: Jurnal Teknologi Terapan, 10(1), 381-389. https://doi.org/10.70609/g-tech.v10i1.8832

Most read articles by the same author(s)