scieee AI-readable full text Open interactive document viewer

Memorandum on Multimedia Applications & Virtual Reality

Mebarki, Abdelkrim

Abstract

Discover immersive applications with this new session, which focuses on the visual immersion in virtual reality. This series of summary notes provides a concise overview of my courses on multimedia applications and virtual reality, originally designed for master's students in computer science (Artificial Intelligence and Applications) at the University of Science and Technology of Oran.

Full text

Available on zenodo – DOI : 10.5281/zenodo.17918661 Memorandum on Multimedia Applications & Virtual Reality Visual Immersion Author: Abdelkrim Mebarki Master’s Program in Computer Science – Artificial Intelligence & Applications Department of Computer Science – Faculty of Mathematics and Computer Science University of Science and Technology of Oran December 12, 2025 1. Introduction The real world is perceived in two ways: 1. Natural direct perception –We perceive our environment without any interface or device. 2. Mediated (diffused / delayed) perception – It occurs through peripherals such as screens, projectors, or headsets. Mediated perception can represent both physical and virtual entities. The main goal of visual immersion is to create virtual objects and environments convincingly stimulating the user’s visual system, making them naturally integrated into the user’s perception. 1.1. Immersion of the Gaze Immersion of the gaze means that the user’s visual field is occupied only by the artificial environment. When the visual system is entirely surrounded by artificial imagery, the brain interprets it as a real and continuous space. This phenomenon reduces external visual references from the real world, strengthening the feeling of presence in virtual world. The quality of gaze immersion depends on: Field of view (FOV) ; Image resolution ; Refresh rate and response latency. The more natural and responsive the visual feedback, the stronger the sense of immersion. 2. Total and Partial Immersion through Image Framing Visual immersion is classified according to how completely the visual space is filled by the virtual scene: 2.1. Partial Immersion: Occurs when the user perceives the virtual content through a limited visual window (smartphone screen, computer monitor, or low-cost headset). The display frame remains visible, and the physical environment continues to influence user perception. 2.2. Total Immersion: Is achieved when the virtual environment occupies the entire visual field, leaving no visible boundary between the image and the external world. This is typically obtained through Head-Mounted Displays or large-scale projection systems. Here, the user perceives only the virtual scene, resulting in a powerful illusion of presence and interaction within the digital created world. 3. Conditions for Total Visual Immersion in Virtual Reality Visual immersion is reached when the user’s entire visual perception is dominated by the virtual scene, to the point that real-world visual cues are either excluded or forgotten. To satisfy this state, several technical and perceptual conditions have to be fulfilled: 3.1. Full Coverage of the Visual Field (Field of View – FOV) The display occupy almost all the user’s natural field of vision — horizontally & vertically. A wide FOV (>100°) eliminates peripheral references to the real world enhancing the sense of presence. 3.2. High Visual Quality (Resolution and Refresh Rate) High resolution and high refresh rate maintain stable and convincing visual experience: ➔ High resolution ensures sharp, detailed images and prevents pixelation effect. ➔ High refresh rate (>90 Hz) guarantees smooth motion and reduces flicker and motion sickness. Memorandum on Multimedia Applications & Virtual Reality (A. Mebarki, 2025) Page 2 of 6 3.3. Low Visual Latency and Real-Time Feedback Any delay between head/eye movement and the corresponding update of images must be imperceptible (<20 ms). Low latency preserves spatial coherence —world moving with user’s gaze, reinforcing realism. 3.4. Stereoscopic Vision and Depth Cues Each eye receives a slightly different image producing binocular disparity, the natural mechanism of depth perception. Combined with motion parallax (view changes as the user moves), this creates 3D illusion. 3.5. Correct Perspective & Natural Scaling Objects in the virtual scene must respect real-world proportions and perspective rules. Anomalies in scale or distance perception can instantly break immersion level and reduce scenario credibility. 3.6. Precise Head and Eye Tracking The visual system must dynamically react to user’s gaze and head orientation. The accurate tracking allows adjustment of natural viewpoint and maintaining alignment between the vision and the behavior. 3.7. Isolation from External Visual Stimuli The visual system must be shielded from all real-world distractions (light, room edges, visible devices). Head-Mounted Displays and immersive rooms achieve this by enclosing user in fully visual environment. 3.8. Consistency and Coherence of the Virtual Scene The artificial world has to remain visually stable, obeying physical and spatial rules as the real world (such as lighting and perspective). Visual incoherence or graphical artifacts instantly disrupt the illusion. ✅ In summary – Total visual immersion → User’s field of vision entirely filled with coherent, dynamic, and realistic virtual imagery that reacts naturally to their movements, excluding real-world visual entities. 4. Human Depth Perception – How It Works The ability to perceive distances and 3D structures is multi-factorial, not based only on a single mechanism. It comes from stereoscopic vision, but monocular cues contribute too. 4.1. Binocular (Stereoscopic) Vision – The Primary Mechanism ➔ Each eye sees the world from a slightly different angle → Binocular Disparity. ➔ The brain merges these two images to reconstruct the depth and the volume. ➔ This mechanism is very strong at both short and medium distances (~20 meters). ➔ It’s the foundation of 3D perception in VR — Headsets display two slightly offset images (one per eye). 4.2. Monocular Cues – Complementary Depth Information Even with one eye, the brain can still infer depth using several cues: Type Description Example Perspective Parallel lines converge in the distance A road narrowing toward the horizon Size Familiar objects appear smaller when farther A distant person looks smaller Occlusion One object hides part of another A tree blocking a building Shadows & Light Lighting gives clues about shape and position A sphere lit from one side looks 3D Motion Parallax Near objects move faster across the field of view Seen while moving or turning the head Texture Gradient Distant textures appear smoother Grass looks detailed nearby, blurry far away These cues simulate depth even on a flat screen or photograph. 4.3. Image Resolution – Only an Indirect Factor ➔ High resolution improves the clarity and precision of depth cues (texture gradients, edges, contrast). ➔ It does not generate depth by itself — Only helps the brain interpret depth information more accurately. ➔ We can have high-resolution flat image that still looks 2D if it lacks binocular or monocular depth cues. Mechanism Type Role in Depth Perception Stereoscopic Vision Binocular Main physiological source of real 3D depth Monocular Cues Cognitive/Visual Complementary, especially at long distances High Resolution Technical Enhances clarity of depth cues, but not a depth mechanism itself Memorandum on Multimedia Applications & Virtual Reality (A. Mebarki, 2025) Page 3 of 6 4.4. Can a person with only one eye perceive depth? Yes — but differently and less precisely. A person with only one eye vision (monocular vision) can still perceive depth and estimate distances relying entirely on monocular depth cues. 1. What is lost: ➔ The brain no longer receives two different images to compare (no binocular disparity). ➔ As a result, true stereoscopic depth at short range disappears. ➔ Tasks requiring fine distance judgment (e.g., catching a ball, threading a needle) become harder. 2. What remains: The brain uses monocular depth cues to reconstruct a sense of depth and distance: ➔ Perspective – Parallel lines converge in the distance. ➔ Relative size – Familiar objects appear smaller when farther away. ➔ Occlusion – One object covering another indicates it’s closer. ➔ Shading & light – Shadows reveal shape and depth. ➔ Motion parallax – Moving your head, nearer objects shift faster across your FoV than distant ones. Over time, people with one eye become skilled at using these cues. Their brain adapts and learns to estimate distances cognitively rather than physiologically. 3. Practical impact: ➔ Static depth (when not moving) is less accurate. ➔ Dynamic depth (while moving) can be quite effective thanks to motion parallax. ➔ Everyday activities are manageable after adaptation, but fine depth precision remains limited. ✅ In summary – A person with one eye can perceive depth, but through visual interpretation (monocular cues and experience), not through true stereoscopic fusion. 4.5. What Is Monocular Motion Parallax? Motion parallax is a depth perception mechanism that uses the relative motion between the objects and the observer to estimate the distances. When someone moves his head or body, objects at different distances appear to move at different speeds and directions across his retina: ➔ Nearby objects seem to move quickly and in the opposite direction of his motion. ➔ Distant objects move slowly and often in the same direction as him. His brain analyzes these differences to compute spatial depth, even when using only one eye. 👁️ Example – Imagine you’re sitting in a moving car: ➔ Trees near the road seem to race past you 🌳💨 ➔ Mountains far away barely seem to move at all ⛰️ Your brain interprets this as: “Trees are close, mountains are far”. ➔ Even with one eye closed, this depth sense works — because it’s based on motion, not stereoscopy. ⚙️ How It Works (Perceptual Mechanism) 1. Visual perspective changes slightly over time – each moment provides a new image of the scene. 2. The brain compares these temporal changes (like frames in a video). 3. From each object’s apparent motion Speed and Direction, the brain deduces their relative distances. This process is called temporal disparity, as opposed to the spatial disparity used in binocular vision. 🎮 In Virtual Reality and 3D Applications Motion parallax reinforces depth perception, even when stereoscopy is limited or absent. For example: ➔ Head tracking in VR (via sensors) allows the image to shift dynamically as the user moves. ➔ This creates realistic motion-based depth, similar to what the human brain expects in real life. ➔ Monoscopic VR (One image VR mobile) – parallax motion creates convincing illusion of depth. Mechanism Type Requires 2 Eyes? Works in VR? Depth Range Key Principle Stereoscopic Vision Binocular ✅ Yes ✅ Yes Short–medium Spatial Disparity (2 views) Motion Parallax Monocular ❌ No ✅ Yes Medium–long Temporal Disparity (motion-based) Memorandum on Multimedia Applications & Virtual Reality (A. Mebarki, 2025) Page 4 of 6 ✅ In summary – Motion parallax → Perception of depth → Relative motion between Observer and objects → Dynamic monocular cue → Brain understands 3D from movements, not from separate images. 4.6. 🧠 Is Motion Parallax a Simulation of Binocular Vision? Motion parallax can be seen as a temporal substitute for spatial binocular vision. But it’s not identical. 👁️ In Binocular Vision ➔ Two eyes separated horizontally (~6 cm). ➔ Each eye captures the world from a slightly different spatial viewpoint at the same time. ➔ The brain instantly compares two simultaneous perspectives (Spatial Disparity) to compute depth. 🌀 In Motion Parallax (Monocular Vision) ➔ Only one eye, but moving slightly as the head is shifted. ➔ At time t ₁ , the eye sees the scene from one position. ➔ At time t ₂ , the same eye sees it from a slightly shifted position. ➔ The brain compares these two successive images (Temporal Disparity) to estimate depth. 🔍 Simulated Binocular Vision The moving eye in motion parallax acts like two eyes, one before and the other after movement. The brain compares the difference between two “snapshots” like it compares images from left and right. However: ➔ In stereoscopy, images are captured simultaneously → Perfect for fast depth calculation. ➔ In motion parallax, the brain must rely on memory and temporal integration of what it just saw — it’s slightly slower and less precise. Motion parallax is not true stereoscopy, but a functional analog of it. Feature Binocular Vision Motion Parallax Type Spatial stereopsis Temporal stereopsis Source Two eyes (left/right) One eye (moving over time) Timing Simultaneous Sequential Precision Very high (short range) Medium (dynamic scenes) In VR True 3D Simulated 3D via head tracking 💡 Simplified Analogy Binocular vision = two eyes at once Motion parallax = one moving eye over time Both give the brain two perspectives — one in space, one in time — allowing it to reconstruct depth. 5. Visual Immersion Interfaces in Virtual Reality Visual immersion is the primary key of virtual reality (VR). It refers to the user’s ability to feel visually surrounded by the synthetic environment, perceiving objects as if they were truly present in his native space. This illusion is produced by a combination of optical, electronic, and perceptual mechanisms, implemented through a variety of peripherals and display technologies. These immersive interfaces are grouped into three main categories: fixed systems, wearable systems, and emerging technologies. 5.1. Fixed (Stationary) Immersive Interfaces Immersive Rooms (CAVEs) A CAVE (Cave Automatic Virtual Environment) is a room-sized cube where images are projected on several walls, and sometimes on the floor and ceiling. The user stands or moves inside the room space wearing stereo glasses tracked by motion sensors. Projection of stereoscopic synchronized images on each surface creates a powerful sense of spatial presence. CAVEs are used for scientific visualization, design review, and collaborative VR because they allow several users to share the same immersive scene. Memorandum on Multimedia Applications & Virtual Reality (A. Mebarki, 2025) Page 5 of 6 Immersive Desks and Workbenches Smaller projection environment — Horizontal or inclined surface where 3D objects appear in depth. The user views the virtual objects through stereoscopic glasses, sometimes combined with head tracking and haptic devices offering semi-immersive visualization suitable for medical imaging, industrial design, and education. Large Displays and Projection Screens This system uses one or more large high-resolution screens or curved projectors to surround the main frontal part of the user’s visual field. Although not fully immersive, they provide a wide field of view and are used in flight simulators, training systems, and panoramic installations. Stereoscopic Displays Stereoscopic displays present two slightly different images to the left and right eyes. The brain fusions them to form a 3D perception. These displays can be flat monitors, projectors, or even television screens relying on optical separation methods to deliver different images to each eye. ➔ Parallax Barrier Displays (Obsolete) A parallax barrier display uses a layer of slits placed in front of the screen to direct different pixels toward each eye. This allows depth perception without glasses (called autostereoscopic). However, the constrained narrow viewing angle and brightness loss make it unsuitable for immersive VR → Mainly used in small portable 3D displays or research prototypes. ➔ Lenticular Displays Similar to the parallax barrier, lenticular lenses refract light in different directions, sending one distinct image to each eye. This device is autostereoscopic too creates a 3D impression without eyewear but with limited accuracy. Lenticular technology is still used in advertising panels and handheld devices, but rarely in high-end VR systems. ➔ Glasses-Level Separation Techniques These interfaces rely on optical or electronic filtering to send different images from the screen to each eye, producing stereoscopic depth. ➔ Colorimetric Differentiation (Anaglyph Glasses) – Obsolete An anaglyph device encodes left and right images in complementary colors (usually red and cyan). Each glasses lens filters one color, allowing each eye to see its corresponding image. Although it offers a basic 3D experience, it distorts colors and causes eye fatigue. ➔ Polarizing Glasses (Passive) These systems use light waves oscillating in different planes. The projectors emit light polarized differently, and each lens passes only one orientation. This spatial multiplexing (two projectors side-byside) or temporal multiplexing (alternating frames) induces stereoscopic vision. Polarized projections are used in cinema and large-screen VR installations, as it allows several users to share the same content. ➔ Active Shutter Glasses (Electronic) Active glasses are based on liquid crystal shutters that alternately block each eye vision in synchronization with the screen refresh rate. Each eye receives their images in sequence, producing full-color stereoscopic view. This technique was dominant in early 3D computer graphics and simulations. 5.2. Mobile and Wearable Interfaces Head-Mounted Displays (HMDs) It is the most common immersive interface in modern VR integrating stereoscopic displays, motion tracking, and head orientation sensors into a headset worn by the user. Rendering images according to head movement allows the HMD providing a real-time egocentric experience. Memorandum on Multimedia Applications & Virtual Reality (A. Mebarki, 2025) Page 6 of 6 Smart Glasses and Augmented Reality Headsets Smart or AR glasses mix virtual content with real environment using semi-transparent lenses or waveguides to overlay information onto the physical world. These devices enable mixed reality, where users can interact with both real and virtual elements simultaneously. Smartphone-Based Viewers (VR Boxes, Cardboard) – Obsolete Early mobile VR emulators used smartphones inserted into plastic/cardboard headsets displaying two side-by-side images forming a simple stereoscopic view. These systems provided low-cost access to VR lacking positional tracking and interactivity. 5.3. Emerging and Hybrid Interfaces Light Field Displays Light field technology reproduces the direction and intensity of light rays in space. This approach enables perception of natural depth and changes of focus without stereoscopy technics. Holographic Displays Holography reconstructs the complete light wavefront of 3D models, allowing true volumetric perception without glasses. This is currently used in research labs and prototype installations. Panoramic Dome Projections Panoramic Dome Projection Systems (Fulldome or Immersive Dome Systems) are large hemispherical projection environments that surround the viewer with imagery on a curved surface, (180° or even 360°). Projection principle: 5 to 20 projectors are precisely calibrated to cover the entire dome. The projected images are blended together to form one continuous panoramic scene. A single fisheye projector can be used at the center to project the entire dome view. Immersion mechanism: The visual field is completely filled (from the floor to the ceiling) → The viewer feels inside the environment rather than watching from outside stimulating peripheral vision, enhancing immersion and spatial awareness. Applications: ➔ Museums and planetariums → Astronomy, natural history, and educational 360° experiences. ➔ Simulation rooms → Training for aviation, driving, military, or urban planning. ➔ Collaborative VR installations → Multiple users sharing the same environment without wearing HMDs. Interaction: Some domes integrate motion tracking or gesture interfaces, enabling limited interactivity. Status: Widely used in scientific visualization and cultural installations, though not personal VR: It’s a shared, large-scale immersive installation rather than an individual headset. 6. Conclusion From early anaglyph glasses to modern light field headsets, visual immersion interfaces have evolved from optical illusions to fully interactive installations. Technologies represent a compromise between cost, comfort, resolution, and immersion. HMDs dominate personal VR experiences, while CAVEs and domes persist in shared and professional contexts. Future systems will increasingly aim for removing devices while maintaining natural 3D vision, combining the best of optical science and perceptual psychology.