Exploración de las preferencias de usuarios expertos sobre las regiones de importancia como explicaciones en clasificación de actividades en vídeo
DOI:
https://doi.org/10.65234/interaccion.143Palabras clave:
inteligencia artificial explicable, evaluación, XAI centrado en el ser humano, métodos XAI basados en vídeoResumen
Aunque existen numerosos métodos de inteligencia artificial explicable (XAI), todavía hay una falta de estudios que analicen cómo los usuarios perciben la explicabilidad y la confiabilidad que estos ofrecen. Consecuentemente, es difícil determinar cuáles son los métodos de XAI más adecuados en función de las preferencias y necesidades de los usuarios. En este trabajo, usuarios expertos en IA evaluaron seis métodos XAI basados en perturbación, aplicados a través de tres redes y dos conjuntos de datos para el reconocimiento de actividades en vídeo. Para ello, se pidió a los expertos puntuar cómo de razonables fueron las explicaciones, en base a las regiones del vídeo señaladas como importantes. Los resultados muestran la preferencia por el método RISE adaptado a vídeo, mientras que identifican el método de predictores univariados adaptado a vídeo como el menos razonable. Estos hallazgos ofrecen a investigadores y profesionales una visión sobre los métodos de XAI preferidos en vídeo, al tiempo que amplían la comprensión de la explicabilidad de la IA desde una perspectiva centrada en el usuario.
Referencias
Aechtner, J., Cabrera, L., Katwal, D., Onghena, P., Valenzuela, D. P., & Wilbik, A. (2022). Comparing User Perception of Explanations Developed with XAI Methods. Proceedings of the IEEE International Conference on Fuzzy Systems – FUZZ-IEEE ’22, 1-7. https://doi.org/10.1109/FUZZ-IEEE55066.2022.9882743 DOI: https://doi.org/10.1109/FUZZ-IEEE55066.2022.9882743
Alqaraawi, A., Schuessler, M., Weiss, P., Costanza, E., & Berthouze, N. (2020). Evaluating saliency map explanations for convolutional neural networks: A user study. Proceedings of the 25th International Conference on Intelligent User Interfaces – IUI ’20, 275-285. https://doi.org/10.1145/3377325.3377519 DOI: https://doi.org/10.1145/3377325.3377519
Arya, V., Bellamy, R. K. E., Chen, P.-Y., Dhurandhar, A., Hind, M., Hoffman, S. C., Houde, S., Liao, Q. V., Luss, R., Mojsilović, A., Mourad, S., Pedemonte, P., Raghavendra, R., Richards, J. T., Sattigeri, P., Shanmugam, K., Singh, M., Varshney, K. R., Wei, D., & Zhang, Y. (2020). AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models. Journal of Machine Learning Research, 21(130), 1-6. http://jmlr.org/papers/v21/19-1035.html
Barda, A. J., Horvat, C. M., & Hochheiser, H. (2020). A qualitative research framework for the design of user-centered displays of explanations for machine learning model predictions in healthcare. BMC Medical Informatics and Decision Making, 20(1). https://doi.org/10.1186/s12911-020-01276-x DOI: https://doi.org/10.1186/s12911-020-01276-x
Bertasius, G., Wang, H., & Torresani, L. (2021). Is Space-Time Attention All You Need for Video Understanding? Proceedings of the 38th International Conference on Machine Learning Research – PMLR ’21, 139, 813-824. https://proceedings.mlr.press/v139/bertasius21a/bertasius21a-supp.pdf
Dodge, J., Liao, Q. V., Zhang, Y., Bellamy, R. K. E., & Dugan, C. (2019). Explaining models: An empirical study of how explanations impact fairness judgment. Proceedings of the 24th International Conference on Intelligent User Interfaces – IUI ’19, 275-285. https://doi.org/10.1145/3301275.3302310 DOI: https://doi.org/10.1145/3301275.3302310
Ehsan, U., Passi, S., Liao, Q. V., Chan, L., Lee, I.-H., Muller, M., & Riedl, M. O. (2024). The Who in XAI: How AI Background Shapes Perceptions of AI Explanations. Proceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems – CHI ’24, 316.1-316.32. https://doi.org/10.1145/3613904.3642474 DOI: https://doi.org/10.1145/3613904.3642474
Ehsan, U., Wintersberger, P., Liao, Q. V., Watkins, E. A., Manger, C., Daumé III, H., Riener, A., & Riedl, M. O. (2022). Human-Centered Explainable AI (HCXAI): Beyond Opening the Black-Box of AI. Extended Abstracts of the ACM SIGCHI Conference on Human Factors in Computing Systems – CHI EA ’22. https://doi.org/10.1145/3491101.3503727 DOI: https://doi.org/10.1145/3491101.3503727
Gaya-Morey, F. X., Buades-Rubio, J. M., MacKenzie, I. S., & Manresa-Yee, C. (2024). REVEX: A Unified Framework for Removal-Based Explainable Artificial Intelligence in Video. https://doi.org/10.48550/arXiv.2401.11796
Guyon, I., & Elisseeff, A. (2003). An introduction to variable and feature selection. Journal of Machine Learning Research, 3(null), 1157-1182. https://www.jmlr.org/papers/volume3/guyon03a/guyon03a.pdf
Heimerl, A., Weitz, K., Baur, T., & Andre, E. (2020). Unraveling ML Models of Emotion with NOVA: Multi-Level Explainable AI for Non-Experts. IEEE Transactions on Affective Computing, 1(1), 1-13. https://doi.org/10.1109/TAFFC.2020.3043603 DOI: https://doi.org/10.1109/TAFFC.2020.3043603
Jang, J., Kim, D., Park, C., Jang, M., Lee, J., & Kim, J. (2020). ETRI-Activity3D: A Large-Scale RGB-D Dataset for Robots to Recognize Daily Activities of the Elderly. Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems – IROS ’20, 10990-10997. https://doi.org/10.1109/IROS45743.2020.9341160 DOI: https://doi.org/10.1109/IROS45743.2020.9341160
Kaplan, S., Uusitalo, H., & Lensu, L. (2024). A unified and practical user-centric framework for explainable artificial intelligence. Knowledge-Based Systems, 283, 111107. https://doi.org/10.1016/j.knosys.2023.111107 DOI: https://doi.org/10.1016/j.knosys.2023.111107
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., & Zisserman, A. (2017). The Kinetics Human Action Video Dataset. https://doi.org/10.48550/arXiv.1705.06950
Lei, J., G’Sell, M., Rinaldo, A., Tibshirani, R. J., & Wasserman, L. (2018). Distribution-Free Predictive Inference for Regression. Journal of the American Statistical Association, 113(523), 1094-1111. https://doi.org/10.1080/01621459.2017.1307116 DOI: https://doi.org/10.1080/01621459.2017.1307116
Liao, Q. V., Gruen, D., & Miller, S. (2020). Questioning the AI: Informing Design Practices for Explainable AI User Experiences. Proceedings of the ACM SIGCHI Conference on Human Factors in Computing Systems – CHI ’20, 1-15. https://doi.org/10.1145/3313831.3376590 DOI: https://doi.org/10.1145/3313831.3376590
Liao, Q. V., & Varshney, K. R. (2022). Human-centered explainable AI (XAI): From algorithms to user experiences. https://doi.org/10.48550/arXiv.2110.10790
Liu, Z., Wang, L., Wu, W., Qian, C., & Lu, T. (2021). TAM: Temporal Adaptive Module for video recognition. Proceedings of the IEEE/CVF International Conference on Computer Vision – ICCV ’21, 13688-13698. https://doi.org/10.1109/ICCV48922.2021.01345 DOI: https://doi.org/10.1109/ICCV48922.2021.01345
Lopes, P., Silva, E., Braga, C., Oliveira, T., & Rosado, L. (2022). XAI Systems Evaluation: A Review of Human and Computer-Centred Methods. Applied Sciences, 12(19). https://doi.org/10.3390/app12199423 DOI: https://doi.org/10.3390/app12199423
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Proceedings of the 31st International Conference on Neural Information Processing Systems – NIPE ’17, 4768-4777. https://proceedings.neurips.cc/paper/2017/file/8a20a8621978632d76c43dfd28b67767-Paper.pdf
Manresa-Yee, C., Ramis, S., Gaya-Morey, F. X., & Buades, J. M. (2024). Impact of Explanations for Trustworthy and Transparent Artificial Intelligence. Proceedings of the XXIII International Conference on Human Computer Interaction– Interacción ’23. https://doi.org/10.1145/3612783.3612798 DOI: https://doi.org/10.1145/3612783.3612798
Miller, T. (2019). Explanation in Artificial Intelligence: Insights From the Social Sciences. Artificial Intelligence, 267(C), 1-38. https://doi.org/10.1016/j.artint.2018.07.007 DOI: https://doi.org/10.1016/j.artint.2018.07.007
Mohseni, S., Zarei, N., & Ragan, E. D. (2018). A Multidisciplinary Survey and Framework for Design and Evaluation of Explainable AI Systems. ACM Transactions on Interactive Intelligent Systems, 11, 24:1–24:45. https://doi.org/10.1145/3387166 DOI: https://doi.org/10.1145/3387166
OpenMMLab. (2020). OpenMMLab’s Next Generation Video Understanding Toolbox and Benchmark.
Petsiuk, V., Das, A., & Saenko, K. (2018). RISE: Randomized Input Sampling for Explanation of Black-box Models. Proceedings of the British Machine Vision Conference – BMVC ’18, Newcastle, UK, 1-151. https://doi.org/10.48550/arXiv.1806.07421
Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). «Why Should I Trust You?»: Explaining the Predictions of Any Classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining – KDD ’16, 1135-1144. https://doi.org/10.1145/2939672.2939778 DOI: https://doi.org/10.1145/2939672.2939778
Ridley, M. (2025). Human-centered explainable artificial intelligence: An Annual Review of Information Science and Technology (ARIST) paper. Journal of the Association for Information Science and Technology, 76(1), 98-120. https://doi.org/https://doi.org/10.1002/asi.24889 DOI: https://doi.org/10.1002/asi.24889
Rong, Y., Leemann, T., Nguyen, T.-T., Fiedler, L., Qian, P., Unhelkar, V., Seidel, T., Kasneci, G., & Kasneci, E. (2024). Towards human-centered explainable AI: A survey of user studies for model explanations . IEEE Transactions on Pattern Analysis & Machine Intelligence, 46(04), 2104-2122. https://doi.org/10.1109/TPAMI.2023.3331846 DOI: https://doi.org/10.1109/TPAMI.2023.3331846
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. Proceedings of the International Conference on Computer Vision – ICCV ’17, 618-626. https://doi.org/10.1109/ICCV.2017.74 DOI: https://doi.org/10.1109/ICCV.2017.74
Shneiderman, B. (2020). Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy. International Journal of Human–Computer Interaction, 36(6), 495-504. https://doi.org/10.1080/10447318.2020.1741118 DOI: https://doi.org/10.1080/10447318.2020.1741118
Szymanski, M., Millecamp, M., & Verbert, K. (2021). Visual, textual or hybrid: The effect of user expertise on different explanations. Proceedings of the 26th International Conference on Intelligent User Interfaces – IUI ’21, 109-119. https://doi.org/10.1145/3397481.3450662 DOI: https://doi.org/10.1145/3397481.3450662
Wells, L., & Bednarz, T. (2021). Explainable AI and Reinforcement Learning: A systematic review of current approaches and trends. Frontiers in Artificial Intelligance, 4, 1-15. https://doi.org/10.3389/frai.2021.550030 DOI: https://doi.org/10.3389/frai.2021.550030
Yang, C., Xu, Y., Shi, J., Dai, B., & Zhou, B. (2020). Temporal Pyramid Network for Action Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition – CVPR ’20, 591-600. https://doi.org/10.1109/CVPR42600.2020.00067 DOI: https://doi.org/10.1109/CVPR42600.2020.00067
Zeiler, M. D., & Fergus, R. (2014). Visualizing and Understanding Convolutional Networks. Proceedings of the 13th European Conference on Computer Vision – ECCV ’14 (LNCS 8689), 818-833. https://doi.org/10.1007/978-3-319-10590-1_53 DOI: https://doi.org/10.1007/978-3-319-10590-1_53
Descargas
Publicado
Número
Sección
Licencia
Derechos de autor 2025 F. Xavier Gaya-Morey, Jose M. Buades-Rubio, Scott MacKenzie, Raquel Lacuesta, Cristina Manresa-Yee

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial 4.0.
