Privacy-Focused AIoT: Implementing an Offline Voice Assistant for Smart Building Management Using Local LLMs
DOI:
https://doi.org/10.70609/g-tech.v10i2.9342Kata Kunci:
AIoT, Edge computing, Local large language model, Offline voice assistant, Smart buildingAbstrak
Voice assistants are increasingly used for smart building control, yet cloud-based architectures raise privacy risks and become unavailable during internet outages. This study designs and evaluates a fully offline AIoT voice assistant for smart building management using local speech and language models. The system employs an edge audio node (Raspberry Pi Zero 2W with ReSpeaker 2-Mics Pi HAT) and a local GPU server running containerized microservices for speech-to-text (Whisper), intent understanding and action planning (Ollama-hosted LLMs), and text-to-speech (Piper). Building devices and sensors are integrated through Home Assistant, enabling voice-driven control and monitoring without sending audio or interaction logs to external services. Experiments in a laboratory smart-building testbed evaluate speech recognition robustness under varying noise levels, LLM command understanding accuracy and memory footprint, and end-to-end IoT task execution. The speech subsystem achieves a Word Error Rate of 5–20% depending on background noise. Across 33 IoT entities, the assistant reaches a 96.67% execution success rate with an average response time of 5.5 s. Among the evaluated local models, Qwen3 8B achieves the highest intent-to-action accuracy (Acc_I2A=100% on an oracle-text command test set with N=43) with 6.8 GB memory use. The results demonstrate that privacy-preserving and resilient voice interaction for smart building management is feasible using current local LLM stacks.
Referensi
Andreyev, A. (2025). Quantization for OpenAI’s Whisper Models: A Comparative Analysis (arXiv:2503.09905). arXiv. https://doi.org/10.48550/arXiv.2503.09905
Arora, S., Kachari, K. K., Gupta, S., Saarthi, P., Yadav, A. K., & Patnaik, S. S. (2025). Bringing Llama-3 to the Edge: End-to-End Quantized Conversational AI on Raspberry Pi 5. 2025 IEEE International Conference on Computer Vision and Machine Intelligence (CVMI), 1–5. https://doi.org/10.1109/CVMI66673.2025.11337918
Aydin, O., Karaarslan, E., Erenay, F. S., & Bacanin, N. (2025). Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma (arXiv:2503.04765). arXiv. https://doi.org/10.48550/arXiv.2503.04765
Beňo, L., Kučera, E., Drahoš, P., & Pribiš, R. (2024). Transforming industrial automation: Voice recognition control via containerized PLC device. Scientific Reports, 14(1), 29387. https://doi.org/10.1038/s41598-024-81172-w
Birkmose, R., Reece, N. M., Norvin, E. H., Bjerva, J., & Zhang, M. (2025). On-Device LLMs for Home Assistant: Dual Role in Intent Detection and Response Generation (arXiv:2502.12923). arXiv. https://doi.org/10.48550/arXiv.2502.12923
Bolton, T., Dargahi, T., Belguith, S., Al-Rakhami, M. S., & Sodhro, A. H. (2021). On the Security and Privacy Challenges of Virtual Assistants. Sensors, 21(7), 2312. https://doi.org/10.3390/s21072312
Chai, Y., Kwen, M., Brooks, D., & Wei, G.-Y. (2025). FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices (arXiv:2501.07139). arXiv. https://doi.org/10.48550/arXiv.2501.07139
Dwarampudi, A., & Yogi, M. K. (2024). Novel Perspectives in Artificial IoT: Current Trends, Research Directions. Journal of Cyber Security, Privacy Issues and Challenges, 3(1), 22–31. https://doi.org/10.46610/JCSPIC.2024.v03i01.004
Gondi, S., & Pratap, V. (2021). Performance Evaluation of Offline Speech Recognition on Edge Devices. Electronics, 10(21), 2697. https://doi.org/10.3390/electronics10212697
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. de las, Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., & Sayed, W. E. (2023). Mistral 7B (arXiv:2310.06825). arXiv. https://doi.org/10.48550/arXiv.2310.06825
Li, J., chen, C., Pan, L., Azghadi, M. R., Ghodosi, H., & Zhang, J. (2023). Security and Privacy Problems in Voice Assistant Applications: A Survey (arXiv:2304.09486). arXiv. https://doi.org/10.48550/arXiv.2304.09486
Li, Q., Ren, X., Wang, Z., Yao, H., He, Y., & Liu, Y. (2025). MetaPipe: Incremental Deployment of Containerized AI Microservices for Edge Clouds. 2025 IEEE/ACM 33rd International Symposium on Quality of Service (IWQoS), 1–6. https://doi.org/10.1109/IWQoS65803.2025.11143291
Li, Y., Kim, S., & Sy, E. (2021). A Survey on Amazon Alexa Attack Surfaces (arXiv:2102.11442). arXiv. https://doi.org/10.48550/arXiv.2102.11442
Li, Z., Feng, W., Guizani, M., & Yu, H. (2024). TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices (arXiv:2410.00531). arXiv. https://doi.org/10.48550/arXiv.2410.00531
Maier, E., Doerk, M., Reimer, U., & Baldauf, M. (2023). Digital natives aren’t concerned much about privacy, or are they? I-Com, 22(1), 83–98. https://doi.org/10.1515/icom-2022-0041
Ollama. (2026). Ollama API Documentation (Chat Completions, Structured Outputs, and Tool Calling). https://github.com/ollama/ollama/blob/main/docs/api.md
Pelikan, M., Azam, S. S., Feldman, V., Silovsky, J. “Honza,” Talwar, K., Brinton, C. G., & Likhomanenko, T. (2025). Enabling Differentially Private Federated Learning for Speech Recognition: Benchmarks, Adaptive Optimizers and Gradient Clipping (arXiv:2310.00098). arXiv. https://doi.org/10.48550/arXiv.2310.00098
Shen, Y., Shao, J., Zhang, X., Lin, Z., Pan, H., Li, D., Zhang, J., & Letaief, K. B. (2023). Large Language Models Empowered Autonomous Edge AI for Connected Intelligence (arXiv:2307.02779). arXiv. https://doi.org/10.48550/arXiv.2307.02779
Vecino, B. T., Gabrys, A., Matwicki, D., Pomirski, A., Iddon, T., Cotescu, M., & Lorenzo-Trueba, J. (2023). Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications. 12th ISCA Speech Synthesis Workshop (SSW2023), 225–229. https://doi.org/10.21437/SSW.2023-35
Xu, S., & Li, W. (2024). A tool or a social being? A dynamic longitudinal investigation of functional use and relational use of AI voice assistants. New Media & Society, 26(7), 3912–3930. https://doi.org/10.1177/14614448221108112
Yan, C., Ji, X., Wang, K., Jiang, Q., Jin, Z., & Xu, W. (2023). A Survey on Voice Assistant Security: Attacks and Countermeasures. ACM Computing Surveys, 55(4), 1–36. https://doi.org/10.1145/3527153
Ye, S., Du, J., Zeng, L., Ou, W., Chu, X., Lu, Y., & Chen, X. (2024). Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference. IEEE INFOCOM 2024 - IEEE Conference on Computer Communications, 1001–1010. https://doi.org/10.1109/INFOCOM52122.2024.10621342
Yu, Z., Wang, Z., Li, Y., Gao, R., Zhou, X., Bommu, S. R., Zhao, Y. (Katie), & Lin, Y. (Celine). (2024). EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Unified Compression and Adaptive Layer Voting. Proceedings of the 61st ACM/IEEE Design Automation Conference, 1–6. https://doi.org/10.1145/3649329.3658473
Zhang, L., Wu, S., & Wang, Z. (2025). LoRA-INT8 Whisper: A Low-Cost Cantonese Speech Recognition Framework for Edge Devices. Sensors, 25(17), 5404. https://doi.org/10.3390/s25175404
Zheng, Y., Chen, Y., Qian, B., Shi, X., Shu, Y., & Chen, J. (2025). A Review on Edge Large Language Models: Design, Execution, and Applications. ACM Computing Surveys, 57(8), 1–35. https://doi.org/10.1145/3719664
Zhou, Z., Chen, X., Li, E., Zeng, L., Luo, K., & Zhang, J. (2019). Edge intelligence: Paving the last mile of artificial intelligence with edge computing. Proceedings of the IEEE, 107(8), 1738–1762.
Unduhan
Diterbitkan
Terbitan
Bagian
Lisensi
Hak Cipta (c) 2026 Fitri Wibowo, Suheri Suheri, Ferry Faisal, Freska Rolansa

Artikel ini berlisensi Creative Commons Attribution 4.0 International License.








