Exploración de las preferencias de usuarios expertos sobre las regiones de importancia como explicaciones en clasificación de actividades en vídeo

Autores/as

DOI:

https://doi.org/10.65234/interaccion.143

Palabras clave:

inteligencia artificial explicable, evaluación, XAI centrado en el ser humano, métodos XAI basados en vídeo

Resumen

Aunque existen numerosos métodos de inteligencia artificial explicable (XAI), todavía hay una falta de estudios que analicen cómo los usuarios perciben la explicabilidad y la confiabilidad que estos ofrecen. Consecuentemente, es difícil determinar cuáles son los métodos de XAI más adecuados en función de las preferencias y necesidades de los usuarios. En este trabajo, usuarios expertos en IA evaluaron seis métodos XAI basados en perturbación, aplicados a través de tres redes y dos conjuntos de datos para el reconocimiento de actividades en vídeo. Para ello, se pidió a los expertos puntuar cómo de razonables fueron las explicaciones, en base a las regiones del vídeo señaladas como importantes. Los resultados muestran la preferencia por el método RISE adaptado a vídeo, mientras que identifican el método de predictores univariados adaptado a vídeo como el menos razonable. Estos hallazgos ofrecen a investigadores y profesionales una visión sobre los métodos de XAI preferidos en vídeo, al tiempo que amplían la comprensión de la explicabilidad de la IA desde una perspectiva centrada en el usuario.

Referencias

Aechtner, J., Cabrera, L., Katwal, D., Onghena, P., Valenzuela, D. P., & Wilbik, A. (2022). Comparing User Perception of Explanations Developed with XAI Methods. Proceedings of the IEEE International Conference on Fuzzy Systems – FUZZ-IEEE ’22, 1-7. https://doi.org/10.1109/FUZZ-IEEE55066.2022.9882743 DOI: https://doi.org/10.1109/FUZZ-IEEE55066.2022.9882743

Alqaraawi, A., Schuessler, M., Weiss, P., Costanza, E., & Berthouze, N. (2020). Evaluating saliency map explanations for convolutional neural networks: A user study. Proceedings of the 25th International Conference on Intelligent User Interfaces – IUI ’20, 275-285. https://doi.org/10.1145/3377325.3377519 DOI: https://doi.org/10.1145/3377325.3377519

Arya, V., Bellamy, R. K. E., Chen, P.-Y., Dhurandhar, A., Hind, M., Hoffman, S. C., Houde, S., Liao, Q. V., Luss, R., Mojsilović, A., Mourad, S., Pedemonte, P., Raghavendra, R., Richards, J. T., Sattigeri, P., Shanmugam, K., Singh, M., Varshney, K. R., Wei, D., & Zhang, Y. (2020). AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models. Journal of Machine Learning Research, 21(130), 1-6. http://jmlr.org/papers/v21/19-1035.html

Barda, A. J., Horvat, C. M., & Hochheiser, H. (2020). A qualitative research framework for the design of user-centered displays of explanations for machine learning model predictions in healthcare. BMC Medical Informatics and Decision Making, 20(1). https://doi.org/10.1186/s12911-020-01276-x DOI: https://doi.org/10.1186/s12911-020-01276-x

Bertasius, G., Wang, H., & Torresani, L. (2021). Is Space-Time Attention All You Need for Video Understanding? Proceedings of the 38th International Conference on Machine Learning Research – PMLR ’21, 139, 813-824. https://proceedings.mlr.press/v139/bertasius21a/bertasius21a-supp.pdf

Dodge, J., Liao, Q. V., Zhang, Y., Bellamy, R. K. E., & Dugan, C. (2019). Explaining models: An empirical study of how explanations impact fairness judgment. Proceedings of the 24th International Conference on Intelligent User Interfaces – IUI ’19, 275-285. https://doi.org/10.1145/3301275.3302310 DOI: https://doi.org/10.1145/3301275.3302310

Ehsan, U., Passi, S., Liao, Q. V., Chan, L., Lee, I.-H., Muller, M., & Riedl, M. O. (2024). The Who in XAI: How AI Background Shapes Perceptions of AI Explanations. Proceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems – CHI ’24, 316.1-316.32. https://doi.org/10.1145/3613904.3642474 DOI: https://doi.org/10.1145/3613904.3642474

Ehsan, U., Wintersberger, P., Liao, Q. V., Watkins, E. A., Manger, C., Daumé III, H., Riener, A., & Riedl, M. O. (2022). Human-Centered Explainable AI (HCXAI): Beyond Opening the Black-Box of AI. Extended Abstracts of the ACM SIGCHI Conference on Human Factors in Computing Systems – CHI EA ’22. https://doi.org/10.1145/3491101.3503727 DOI: https://doi.org/10.1145/3491101.3503727

Gaya-Morey, F. X., Buades-Rubio, J. M., MacKenzie, I. S., & Manresa-Yee, C. (2024). REVEX: A Unified Framework for Removal-Based Explainable Artificial Intelligence in Video. https://doi.org/10.48550/arXiv.2401.11796

Guyon, I., & Elisseeff, A. (2003). An introduction to variable and feature selection. Journal of Machine Learning Research, 3(null), 1157-1182. https://www.jmlr.org/papers/volume3/guyon03a/guyon03a.pdf

Heimerl, A., Weitz, K., Baur, T., & Andre, E. (2020). Unraveling ML Models of Emotion with NOVA: Multi-Level Explainable AI for Non-Experts. IEEE Transactions on Affective Computing, 1(1), 1-13. https://doi.org/10.1109/TAFFC.2020.3043603 DOI: https://doi.org/10.1109/TAFFC.2020.3043603

Jang, J., Kim, D., Park, C., Jang, M., Lee, J., & Kim, J. (2020). ETRI-Activity3D: A Large-Scale RGB-D Dataset for Robots to Recognize Daily Activities of the Elderly. Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems – IROS ’20, 10990-10997. https://doi.org/10.1109/IROS45743.2020.9341160 DOI: https://doi.org/10.1109/IROS45743.2020.9341160

Kaplan, S., Uusitalo, H., & Lensu, L. (2024). A unified and practical user-centric framework for explainable artificial intelligence. Knowledge-Based Systems, 283, 111107. https://doi.org/10.1016/j.knosys.2023.111107 DOI: https://doi.org/10.1016/j.knosys.2023.111107

Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., & Zisserman, A. (2017). The Kinetics Human Action Video Dataset. https://doi.org/10.48550/arXiv.1705.06950

Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J., & Wasserman, L. (2018). Distribution-Free Predictive Inference for Regression. Journal of the American Statistical Association, 113(523), 1094-1111. https://doi.org/10.1080/01621459.2017.1307116 DOI: https://doi.org/10.1080/01621459.2017.1307116

Liao, Q. V., Gruen, D., & Miller, S. (2020). Questioning the AI: Informing Design Practices for Explainable AI User Experiences. Proceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems – CHI ’20, 1-15. https://doi.org/10.1145/3313831.3376590 DOI: https://doi.org/10.1145/3313831.3376590

Liao, Q. V., & Varshney, K. R. (2022). Human-centered explainable AI (XAI): From algorithms to user experiences. https://doi.org/10.48550/arXiv.2110.10790

Liu, Z., Wang, L., Wu, W., Qian, C., & Lu, T. (2021). TAM: Temporal Adaptive Module for video recognition. Proceedings of the IEEE/CVF International Conference on Computer Vision – ICCV ’21, 13688-13698. https://doi.org/10.1109/ICCV48922.2021.01345 DOI: https://doi.org/10.1109/ICCV48922.2021.01345

Lopes, P., Silva, E., Braga, C., Oliveira, T., & Rosado, L. (2022). XAI Systems Evaluation: A Review of Human and Computer-Centred Methods. Applied Sciences, 12(19). https://doi.org/10.3390/app12199423 DOI: https://doi.org/10.3390/app12199423

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Proceedings of the 31st International Conference on Neural Information Processing Systems – NIPE ’17, 4768-4777. https://proceedings.neurips.cc/paper/2017/file/8a20a8621978632d76c43dfd28b67767-Paper.pdf

Manresa-Yee, C., Ramis, S., Gaya-Morey, F. X., & Buades, J. M. (2024). Impact of Explanations for Trustworthy and Transparent Artificial Intelligence. Proceedings of the XXIII International Conference on Human Computer Interaction– Interacción ’23. https://doi.org/10.1145/3612783.3612798 DOI: https://doi.org/10.1145/3612783.3612798

Miller, T. (2019). Explanation in Artificial Intelligence: Insights From the Social Sciences. Artificial Intelligence, 267(C), 1-38. https://doi.org/10.1016/j.artint.2018.07.007 DOI: https://doi.org/10.1016/j.artint.2018.07.007

Mohseni, S., Zarei, N., & Ragan, E. D. (2018). A Multidisciplinary Survey and Framework for Design and Evaluation of Explainable AI Systems. ACM Transactions on Interactive Intelligent Systems, 11, 24:1–24:45. https://doi.org/10.1145/3387166 DOI: https://doi.org/10.1145/3387166

OpenMMLab. (2020). OpenMMLab’s Next Generation Video Understanding Toolbox and Benchmark.

Petsiuk, V., Das, A., & Saenko, K. (2018). RISE: Randomized Input Sampling for Explanation of Black-box Models. Proceedings of the British Machine Vision Conference – BMVC ’18, Newcastle, UK, 1-151. https://doi.org/10.48550/arXiv.1806.07421

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). «Why Should I Trust You?»: Explaining the Predictions of Any Classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining – KDD ’16, 1135-1144. https://doi.org/10.1145/2939672.2939778 DOI: https://doi.org/10.1145/2939672.2939778

Ridley, M. (2025). Human-centered explainable artificial intelligence: An Annual Review of Information Science and Technology (ARIST) paper. Journal of the Association for Information Science and Technology, 76(1), 98-120. https://doi.org/https://doi.org/10.1002/asi.24889 DOI: https://doi.org/10.1002/asi.24889

Rong, Y., Leemann, T., Nguyen, T.-T., Fiedler, L., Qian, P., Unhelkar, V., Seidel, T., Kasneci, G., & Kasneci, E. (2024). Towards human-centered explainable AI: A survey of user studies for model explanations . IEEE Transactions on Pattern Analysis & Machine Intelligence, 46(04), 2104-2122. https://doi.org/10.1109/TPAMI.2023.3331846 DOI: https://doi.org/10.1109/TPAMI.2023.3331846

Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. Proceedings of the International Conference on Computer Vision – ICCV ’17, 618-626. https://doi.org/10.1109/ICCV.2017.74 DOI: https://doi.org/10.1109/ICCV.2017.74

Shneiderman, B. (2020). Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy. International Journal of Human–Computer Interaction, 36(6), 495-504. https://doi.org/10.1080/10447318.2020.1741118 DOI: https://doi.org/10.1080/10447318.2020.1741118

Szymanski, M., Millecamp, M., & Verbert, K. (2021). Visual, textual or hybrid: The effect of user expertise on different explanations. Proceedings of the 26th International Conference on Intelligent User Interfaces – IUI ’21, 109-119. https://doi.org/10.1145/3397481.3450662 DOI: https://doi.org/10.1145/3397481.3450662

Wells, L., & Bednarz, T. (2021). Explainable AI and Reinforcement Learning: A systematic review of current approaches and trends. Frontiers in Artificial Intelligance, 4, 1-15. https://doi.org/10.3389/frai.2021.550030 DOI: https://doi.org/10.3389/frai.2021.550030

Yang, C., Xu, Y., Shi, J., Dai, B., & Zhou, B. (2020). Temporal Pyramid Network for Action Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition – CVPR ’20, 591-600. https://doi.org/10.1109/CVPR42600.2020.00067 DOI: https://doi.org/10.1109/CVPR42600.2020.00067

Zeiler, M. D., & Fergus, R. (2014). Visualizing and Understanding Convolutional Networks. Proceedings of the 13th European Conference on Computer Vision – ECCV ’14 (LNCS 8689), 818-833. https://doi.org/10.1007/978-3-319-10590-1_53 DOI: https://doi.org/10.1007/978-3-319-10590-1_53

Descargas

Publicado

2025-12-23