83 “Al-Farg‘oniy avlodlari” elektron ilmiy jurnali ISSN 2181-4252. Tom: 1 | Son: 4 | 2025-yil "Descendants of Al-Farghani" electronic scientific journal. ISSN 2181-4252. Vol: 1 | Iss: 4 | 2025 year Электронный научный журнал "Потомки АльФаргани" ISSN 2181-4252. Том: 1 | Выпуск: 4 | 2025 год https://al-fargoniy.uz/ HUMAN PERCEPTION VERSUS ARTIFICIAL INTELLIGENCE IN DEEPFAKE VIDEO DETECTION: AN EMPIRICAL STUDY Maxmudjanov Sarvar, Associate professor of Tashkent University of Information Technologies named after Muhammad Al-Khwarizmi
[email protected] Primbetov Abbaz, Phd student, Tashkent University of Information Technologies named after Muhammad Al-Khwarizmi. Senior lecturer, University of Tashkent for applied sciences. [email protected] Alimbaeva Asalbanu, Phd student, Tashkent University of Information Technologies named after Muhammad Al-Khwarizmi alymbaevaaj9[email protected] Abstract. Deepfake technology has become a growing threat to information integrity and digital trust. This study examines how accurately humans can identify deepfake videos through an empirical survey conducted among 492 students of Tashkent University of Information Technologies named after Muhammad al-Khwarizmi. Participants watched five short videos and classified each as real or fake. The results show that respondents achieved an average accuracy of 57%, while 43% misclassified at least one video. Moreover, 46% of participants were unfamiliar with the term “deepfake.” These findings demonstrate that human perception alone is unreliable for detecting synthetic content and highlight the need for greater media awareness and AI-assisted verification tools in combating digital misinformation. Keywords: Deepfake detection, human perception, survey, misinformation, media literacy. 1. Indroduction In recent years, deepfake technology has rapidly evolved into one of the most pressing challenges in the digital era. Leveraging the power of deep learning and generative adversarial networks (GANs), deepfakes enable the realistic synthesis of human faces, voices, and actions in videos that are indistinguishable from authentic content. While initially developed for creative and entertainment purposes, deepfake applications have increasingly raised ethical, political, and security concerns worldwide [3]. Manipulated videos have been used to spread misinformation, fabricate public statements, and damage personal reputations—ultimately undermining public trust in visual media [2]. A growing body of research has examined how humans perceive and evaluate the authenticity of such manipulated media. Despite intuitive confidence in our visual judgment, empirical findings consistently show that human perception is highly unreliable in detecting AI-generated synthetic content. In a comprehensive meta-analysis, Diel et al. (2024) analyzed 56 scientific studies comprising responses from more than 86,000 participants across different modalities. Their results demonstrated that the average human accuracy in identifying deepfakes was 55.54%, barely above random guessing, while recognition of real content averaged 68.08%. Table 1 summarizes these findings, indicating that human observers struggle equally across audio, image, text, and video formats, with only minor variations in performance.
84 “Al-Farg‘oniy avlodlari” elektron ilmiy jurnali ISSN 2181-4252. Tom: 1 | Son: 4 | 2025-yil "Descendants of Al-Farghani" electronic scientific journal. ISSN 2181-4252. Vol: 1 | Iss: 4 | 2025 year Электронный научный журнал "Потомки АльФаргани" ISSN 2181-4252. Том: 1 | Выпуск: 4 | 2025 год https://al-fargoniy.uz/ Modality Deepfake Accuracy (%) Real Accuracy (%) Audio 62.08 70.67 Image 53.16 68.46 Text 52 66.81 Video 57.31 68 Table 1. Average accuracy of people in detecting deepfakes across modalities. These findings highlight a fundamental cognitive vulnerability: even educated or digitally experienced individuals are often unable to reliably distinguish authentic media from synthetic manipulations. As a result, research attention has increasingly turned toward artificial intelligence (AI)- based detection systems that can analyze subtle inconsistencies beyond the threshold of human perception—such as mismatches in eye blinking, lip synchronization, lighting, or motion dynamics. Figure 1. Average Human Accuracy in Detecting Deepfake and Real Content (Diel et al., 2024) The chart clearly shows that human accuracy is consistently higher for real content (≈68–71%) than for deepfake content (≈52–62%), confirming that humans struggle more with identifying manipulated media, especially in image and text modalities [1]. However, while numerous studies have examined AI models for deepfake detection, far fewer have focused on human performance within specific cultural or educational contexts. Most prior research has been conducted in Western or East Asian countries, leaving a notable gap in regional understanding. Addressing this gap, the present study explores human perception accuracy in deepfake video detection within the context of Uzbekistan, using an empirical survey conducted among 492 students at the Tashkent University of Information Technologies named after Muhammad al-Khwarizmi. Participants were asked to view short video clips and determine whether each was real or fake, enabling a quantitative assessment of perceptual accuracy and awareness levels [4]. By situating this investigation within a unique cultural and academic setting, this paper contributes new data to the global discourse on human deepfake perception. It aims to compare local findings with international benchmarks and to highlight the urgent need for media literacy education and AI-assisted verification systems to strengthen digital resilience against synthetic misinformation. 2. Methodology This study employed an empirical, surveybased research design to evaluate the accuracy of human perception in detecting deepfake videos. The main objective was to measure how effectively individuals could differentiate between authentic and manipulated video content based solely on visual observation. The study was conducted using an online questionnaire created through Google Forms, allowing participants to watch a series of short videos and record their judgments about whether each was real or fake. A total of 492 respondents participated in the survey. All participants were students of the Tashkent University of Information Technologies named after Muhammad al-Khwarizmi, enrolled in the external (part-time) education program. The sample included both undergraduate and postgraduate students across different age groups, academic years, and majors. Participation was voluntary, and all responses were recorded anonymously to ensure confidentiality [5]. The questionnaire consisted of five main videos presented sequentially, along with five corresponding questions that required participants to determine whether each video was real (authentic) or fake (deepfake). In addition, several supporting questions were included to assess: The participants’ prior awareness of deepfake technology, Their confidence level in visual judgment, and Their ability to identify deepfake cues across multiple examples.
85 “Al-Farg‘oniy avlodlari” elektron ilmiy jurnali ISSN 2181-4252. Tom: 1 | Son: 4 | 2025-yil "Descendants of Al-Farghani" electronic scientific journal. ISSN 2181-4252. Vol: 1 | Iss: 4 | 2025 year Электронный научный журнал "Потомки АльФаргани" ISSN 2181-4252. Том: 1 | Выпуск: 4 | 2025 год https://al-fargoniy.uz/ All videos were short clips (10–15 seconds) selected from open-source datasets and controlled to ensure consistent quality, lighting, and facial framing. The presentation order was randomized to minimize pattern recognition or learning effects. The survey was distributed online through institutional communication channels. Respondents were instructed to view each video completely before selecting an answer. The Google Forms platform automatically collected all responses in a structured spreadsheet format. Data were exported to Microsoft Excel for processing and basic descriptive analysis. Participation in the study was voluntary, and all respondents were informed about the academic purpose of the survey before participation. Although the questionnaire collected limited personal information such as the participants’ names and group numbers, these data were used exclusively for verification and statistical grouping purposes. No personal identifiers were disclosed or used in any public report or publication. All collected information was stored securely and analyzed anonymously to ensure privacy protection. The videos presented in the survey did not contain harmful, violent, or ethically sensitive material. All materials were used strictly for research and educational purposes within the framework of media literacy and artificial intelligence studies. 4. Results and Discussion The results revealed that participants achieved an average accuracy of 57% in detecting deepfake content, while 43% of responses represented misclassifications—either falsely identifying real videos as fake or accepting fake videos as real. These results are consistent with global averages reported by Diel et al. (2024), who found a mean deepfake detection accuracy of 55.54% across more than 86,000 participants. 3.1 Awareness of Deepfake Technology Before evaluating the videos, participants were asked about their familiarity with deepfake technology. As shown in Figure 2, 46.1% of respondents stated that the concept was entirely new to them, 38.4% had heard of it but lacked detailed knowledge, and only 15.4% claimed to have in-depth understanding. Figure 2. Distribution of Participants’ Prior Awareness of Deepfake Technology This distribution suggests that a large proportion of university students—despite being digital natives—remain insufficiently informed about emerging AI-based manipulation technologies. The relatively low awareness level likely contributed to the moderate accuracy rate observed in video classification tasks. 3.3 Accuracy and Error Patterns In the second question, participants were asked to classify a short video clip as either real or deepfake. Of the 492 responses, 271 (55.1%) identified the video as real, and 221 (44.9%) labeled it as deepfake. Assuming the correct classification was deepfake, the results show that approximately 55% correctly identified the video, while 45% misclassified it. Figure 2. Classification of Video as Real or Deepfake 3.4 Identification of Multiple Deepfakes Participants were shown five video clips and asked to identify which ones were deepfakes. The actual manipulated clips were Videos 1, 2, and 3, while Videos 4 and 5 were authentic. As illustrated in Figure 3, the responses revealed considerable variation, with participants frequently misclassifying both real and
86 “Al-Farg‘oniy avlodlari” elektron ilmiy jurnali ISSN 2181-4252. Tom: 1 | Son: 4 | 2025-yil "Descendants of Al-Farghani" electronic scientific journal. ISSN 2181-4252. Vol: 1 | Iss: 4 | 2025 year Электронный научный журнал "Потомки АльФаргани" ISSN 2181-4252. Том: 1 | Выпуск: 4 | 2025 год https://al-fargoniy.uz/ fake videos. Although the majority correctly recognized the presence of manipulation in some clips, the overall detection accuracy for genuine deepfakes remained below 50%, while many authentic videos were incorrectly labeled as fake. Figure 3. Frequency of Deepfake Identification Across Five Videos This inconsistency highlights two common perceptual errors: under-detection of true deepfakes and over-detection of real videos. Such tendencies suggest that human observers rely more on intuition and superficial visual cues—such as unnatural lighting or facial motion—than on objective indicators of manipulation. 3.5 Perception of Deepfake Presence Participants were asked whether any of the previously shown videos contained deepfake content. As shown in Figure 4, 56.3% of respondents answered “Yo‘q” (No), which was the correct response, while 43.7% incorrectly believed that at least one video contained a deepfake. This result indicates a notable level of false suspicion among participants. Nearly half of the respondents perceived manipulation where none existed, demonstrating the psychological bias toward over-detection often observed in digital media interpretation. Such overestimation suggests that human observers, when uncertain, tend to assume falsification rather than authenticity. Figure 4. Participants’ Responses Regarding the Presence of Deepfakes 3.6 Identification of Real (Authentic) Videos In the final question, participants were asked to identify which of the five clips were authentic (i.e., not deepfakes). The correct answers were Videos 4 and 5, both genuine. As illustrated in Figure 5, participants’ responses were widely distributed, with no strong consensus across the videos. The majority of respondents incorrectly labeled Videos 1–3—which were in fact deepfakes—as real, while less than half correctly identified Videos 4 and 5 as authentic. This demonstrates a persistent difficulty in recognizing genuine content, even when the proportion of real and fake videos was balanced. These results emphasize the asymmetry in human perception—people are prone to doubt authenticity but less capable of confirming it accurately. Figure 5. Accuracy Distribution in Recognizing Real Videos Conclusion This study examined human capability in detecting deepfake video content through an empirical survey conducted among 492 students at the Tashkent University of Information Technologies named after Muhammad al-Khwarizmi.
87 “Al-Farg‘oniy avlodlari” elektron ilmiy jurnali ISSN 2181-4252. Tom: 1 | Son: 4 | 2025-yil "Descendants of Al-Farghani" electronic scientific journal. ISSN 2181-4252. Vol: 1 | Iss: 4 | 2025 year Электронный научный журнал "Потомки АльФаргани" ISSN 2181-4252. Том: 1 | Выпуск: 4 | 2025 год https://al-fargoniy.uz/ The results revealed that despite technological awareness among participants, human accuracy in identifying manipulated media remains low and inconsistent. Most respondents lacked deep understanding of deepfake technology, and even those familiar with the concept frequently misclassified both real and fake content. The average detection accuracy across all tasks was approximately 57%, which aligns with international benchmarks reported in prior metaanalyses. The findings highlight two dominant perceptual errors—under-detection of actual deepfakes and overdetection of authentic videos—reflecting a broader cognitive uncertainty when evaluating visual authenticity. Participants tended to rely on intuitive cues rather than analytical observation, resulting in frequent false judgments. Such limitations underscore the inadequacy of unaided human perception as a reliable detection mechanism. Therefore, the study concludes that humanbased deepfake detection is inherently unreliable and must be supplemented by AI-driven verification systems capable of analyzing subtle spatial and temporal inconsistencies beyond human perception. Integrating automated detection tools with digital literacy education may significantly enhance users’ ability to critically evaluate visual content, contributing to greater resilience against misinformation and synthetic media manipulation in the digital age. References 1. Diel, A., Lalgi, T., Schröter, I. C., MacDorman, K. F., Teufel, M., & Bäuerle, A. (2024). Human performance in detecting deepfakes: A systematic review and meta-analysis of 56 papers. Computers in Human Behavior Reports, 16, 100538. 2. Abbaz, P., & Navruz, A. (2025). A comparative study of activation functions in deep learning models. Al-Farg’oniy avlodlari, 1(3), 80-84. 3. Primbetov A (2025). Deepfake detection using a hybrid resnext and lstm architecture. Al-farg’oniy avlodlari, 1(2), 87-94. 4. Primbetov (2025). DEEPFAKE TECHNOLOGY: THREAT LANDSCAPE, DETECTION TECHNIQUES, AND ETHICAL GOVERNANCE. Al-Farg’oniy avlodlari, 1(2), 106112. 5. Primbetov A (2025). Students survey list. Google Forms. 6. M. Tolosana, R. Vera-Rodriguez, J. Fierrez, A. Morales, and J. Ortega-Garcia, “Deepfakes and beyond: A survey of face manipulation and fake detection,” *Information Fusion*, vol. 64, pp. 131– 148, 2020. 7. A. Diel et al., “Human performance in detecting deepfakes: A systematic review and metaanalysis of 56 papers,” *Computers in Human Behavior Reports*, vol. 16, p. 100538, 2024. 8. F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in *Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR)*, 2017, pp. 1251–1258. 9. M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in *Proc. Int. Conf. Machine Learning (ICML)*, 2019, pp. 6105–6114. 10. S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in *Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR)*, 2017, pp. 1492–1500. 11. W. van Gansbeke, A. Dhoedt, and J. Van de Weijer, “Temporal awareness in deepfake detection,” *IEEE Trans. Biometrics, Behavior, and Identity Science*, vol. 3, no. 2, pp. 176–187, 2021.