scieee AI-readable full text Open interactive document viewer

Pan-preferential stereo IBVS on ARMOS TurtleBot: Design and experimental evaluation

Martínez, Fredy Hernán

Abstract

Image-based visual servoing (IBVS) on nonholonomic platforms often suffers from transient conflicts between camera alignment and base motion when the viewpoint is fixed. This work addresses that difficulty by mounting a rectified stereo pair on a controllable pan axis and commanding a two-channel pan-preferential IBVS that regulates image features while the base heading follows the camera yaw. Depth at feature locations is recovered from disparity to build a well-conditioned visual Jacobian, and a damped least-squares inversion with feature weighting produces bounded linear speed and pan-rate commands. The methodology is implemented on the ARMOS TurtleBot differential base equipped with a low-cost ESP32+OV7670 stereo rig, and underpinned by fast encoder-based motor loops and disturbance-observer augmented PD control. Two experimental scenarios (static target and lateral target motion at roughly constant depth) demonstrate consistent feature tracking, smooth camera-base alignment, and tolerance to transient disparity dropouts through outlier rejection and command projection. Across repeated runs the proposed scheme reduces steady error and settling time compared with a fixed-camera baseline while respecting modest computing and driver limits. All components, rates, and bounds are reported to enable replication, and the results indicate that stereo-depth IBVS with pan alignment is feasible on affordable hardware for indoor mobile manipulation and navigation.

Full text

 Corresponding author: Fredy Martínez Copyright © 2025 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution Liscense 4.0. Pan-preferential stereo IBVS on ARMOS TurtleBot: Design and experimental evaluation Fredy Hernán Martínez * Facultad Tecnológica, Universidad Distrital Francisco José de Caldas, Bogotá, Colombia. Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 Publication history: Received on 10 July 2025; revised on 17 August 2025; accepted on 19 August 2025 Article DOI: https://doi.org/10.30574/gjeta.2025.24.2.0250 Abstract Image-based visual servoing (IBVS) on nonholonomic platforms often suffers from transient conflicts between camera alignment and base motion when the viewpoint is fixed. This work addresses that difficulty by mounting a rectified stereo pair on a controllable pan axis and commanding a two-channel pan-preferential IBVS that regulates image features while the base heading follows the camera yaw. Depth at feature locations is recovered from disparity to build a well-conditioned visual Jacobian, and a damped least-squares inversion with feature weighting produces bounded linear speed and pan-rate commands. The methodology is implemented on the ARMOS TurtleBot differential base equipped with a low-cost ESP32+OV7670 stereo rig, and underpinned by fast encoder-based motor loops and disturbance-observer augmented PD control. Two experimental scenarios (static target and lateral target motion at roughly constant depth) demonstrate consistent feature tracking, smooth camera-base alignment, and tolerance to transient disparity dropouts through outlier rejection and command projection. Across repeated runs the proposed scheme reduces steady error and settling time compared with a fixed-camera baseline while respecting modest computing and driver limits. All components, rates, and bounds are reported to enable replication, and the results indicate that stereo-depth IBVS with pan alignment is feasible on affordable hardware for indoor mobile manipulation and navigation. Keywords: Differential-drive robots; Image-based visual servoing; Low-cost stereo vision; Pan-preferential control; Stereo disparity 1. Introduction Image-based visual servoing (IBVS) remains a compelling strategy for mobile robots because it regulates quantities that are sensed directly, avoids brittle reconstruction pipelines, and can run with modest onboard resources [1–3]. When the camera is carried by a differential base, however, the camera-base mapping introduces configuration-dependent singularities that degrade control authority at certain poses [4]. Depth ambiguity in monocular setups further complicates gain selection and can slow convergence, particularly when features lie at moderate range [5]. Practical deployments add timing constraints and packet jitter that must be reflected explicitly in the control design rather than treated as afterthoughts. These factors together motivate a solution that addresses geometry, sensing, and real-time execution as a single, coherent problem [6]. Prior work has mitigated singularities by switching between control modes or by constraining motion near problematic configurations, which can produce discontinuities and transient spikes [7, 8]. Other approaches have relied on monocular depth surrogates or scene heuristics; while serviceable in narrow settings, they often yield poor conditioning of the interaction matrix and require heavy damping to remain stable [9, 10]. Continuous reparameterizations exist, yet many assume camera placements or actuation sets that do not match common mobile platforms. Low-cost stereo has Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 182 seen growing interest, but published integrations rarely document end-to-end timing, baseline stability under motion, and the interaction with inner motor loops [11]. A gap therefore remains for a reproducible, continuous IBVS policy that avoids mode switches, uses measured depth at the feature locations, and runs on accessible hardware. This paper addresses that gap with a pan-preferential scheme that couples a rigid stereo pair on a single-axis pan mount to a differential base. The image-space inversion commands the pair [𝑣,𝜙󰇗𝑐] through a minimum-norm solution, while the base yaw follows the camera with a proportional law 𝜔=𝐾𝑜𝜙𝑐, thereby keeping the inversion free from the directions that cause rank loss. Depth is obtained from disparity at the same pixels that drive the controller, with rectified intrinsics and a fixed baseline ensuring 𝑍=𝑓𝐵/𝑑 is well defined during pan motion. The system is realized on the ARMOS TurtleBot using two ESP32+OV7670 modules that stream frames to a host for rectification, disparity, and feature tracking at low resolution and moderate rate [12]. Safety interlocks (infrared stop, watchdog, and an e-stop) and a disturbance-observer layer on the wheel and pan loops complete the stack needed for a stable real-robot evaluation. The study is organized around operating points and metrics that reflect both control theory and field constraints. Vision runs primarily at 160×120 pixels and 10–20 Hz per camera; the IBVS loop updates at 50–100 Hz, and the motor layer executes at 1–2 ms periods. Trials comprise a static target and a laterally moving target (≈0.2 m/s) at indoor ranges of 1–3 m, with identical gains and caps across ablations. Key performance indicators include image-plane RMS error and settling time, the decay of |𝜙𝑐|, control energies for linear and angular channels, and the condition number of the composite visual Jacobian together with damping usage. End-to-end latency and stereo quality (inlier ratio and depth outliers) are reported to link closed-loop behavior to the timing and sensing budget. The contributions are fourfold and are backed by a complete real-robot implementation [13]. First, the paper formulates a continuous stereo-IBVS policy that preserves pan-preferential avoidance of singularities while injecting measured depth at the feature level. Second, it delivers an integration on an affordable platform with documented parameters, rates, and safety mechanisms, including PD plus disturbance observers on wheel and pan axes. Third, experiments against fixed-camera and monocular surrogates show improved feature regulation, smoother base-camera alignment, and fewer damped steps within the measured latency envelope, without inflating actuation energy. Fourth, the work provides a reproducible path by disclosing calibration assets, logging schema, and tuning guidelines. The remainder of the paper presents the system overview, the visual-servoing formulation with stereo depth, the low-level implementation, the experimental setup, the results with ablations, and a discussion of implications and limitations. 2. System Overview The platform is a differential-drive mobile base (ARMOS TurtleBot) equipped with wheel encoders, a single-axis pan joint for the vision payload, and a ring of proximity sensors for reflexive stopping in close-range situations. The locomotion module uses two independently driven wheels separated by a wheelbase 𝑙=0.25 m and wheel radius 𝑟= 0.024 m; these values are retained in both modeling and experiments to match the mechanical design. Let 𝜔𝑅 and 𝜔𝐿 denote the right and left wheel angular speeds, respectively. The body-frame linear and angular velocities (𝑣,𝜔) satisfy 𝑣=𝑟2(𝜔𝑅+𝜔𝐿), 𝜔=𝑟𝑙(𝜔𝑅−𝜔𝐿) …….. (1) and the inverse map used by the low-level controller is 𝜔𝑅=𝑣𝑟+𝑙2𝑟 𝜔, 𝜔𝐿=𝑣𝑟−𝑙2𝑟 𝜔 ……… (2) Nominal commanded speed is capped to 𝑣≤0.2 m/s on indoor flat floors to accommodate vision sampling and actuation limits, whereas the safety ring of nine infrared sensors issues an immediate halt when the estimated clearance falls below a fixed threshold. The vision payload is a rigid stereo pair mounted on a yaw (pan) axis with measured angle 𝜙𝑐, so that both cameras rotate together while their left–right extrinsics remain constant. Each camera is an OV7670 sensor interfaced to an ESP32 module; frames are sent over WebSocket to a host computer that runs rectification and disparity. The capture stage supports 160×120 pixels at 10–20 fps per stream and 320×240 pixels at 5–10 fps when bandwidth permits; we operate primarily at the lower resolution to reduce end-to-end latency. Stereo calibration yields intrinsics (𝑓𝑥,𝑓𝑦,𝑐𝑥,𝑐𝑦) and a fixed baseline 𝐵; after rectification, disparity 𝑑(𝑢,𝑣) provides per-feature depth via Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 183 𝑍(𝑢,𝑣)=𝑓 𝐵 𝑑(𝑢,𝑣), 𝑓=𝑓𝑥+𝑓𝑦 2 …… (3) Feature detection and tracking are performed on the rectified left image (fiducials or salient points), producing the image vector 𝑓=[𝑢𝑖,𝑣𝑖]𝑖=1..𝑘 and associated depths 𝑍=[𝑍𝑖]𝑖=1..𝑘, which feed the visual servo controller. Figure 1 summarizes the sensing-to-actuation data flow, including time alignment and control layers. The camera twist in the spatial frame, ⬚𝑠𝑤𝑐=[ ⬚𝑠𝑝󰇗𝑐⊤, ⬚𝑠𝜔𝑐⊤]⊤∈ℝ6, is related to the commanded generalized velocities 𝜈=[𝑣 𝜔 𝜙󰇗𝑐]⊤ ………. (4) through a configuration-dependent Jacobian, ⬚𝑠𝑤𝑐= ⬚𝑐𝐽(𝜙𝑐) 𝜈 ………. (5) which captures how base translation, base yaw, and pan rate induce camera translational and rotational velocities. This mapping partitions naturally as ⬚𝑐𝐽(𝜙𝑐)=[ ⬚𝑐𝐽13(𝜙𝑐) ⬚𝑐𝐽2(𝜙𝑐) ] with ⬚𝑐𝐽13∈ℝ6×2 selecting the {𝑣,𝜙󰇗𝑐} channels used by the outer visual loop, while ⬚𝑐𝐽2 corresponds to the base yaw channel 𝜔. By retaining the pan axis and selecting {𝑣,𝜙󰇗𝑐} in the image-space inversion, the proposed architecture avoids the loss of rank that occurs in camera-base configurations where the full three-channel inversion would be ill-conditioned. Figure 1 System overview. Dual ESP32+OV7670 streams are time-aligned by a GPIO trigger and timestamps, rectified on the host, and processed to obtain disparity and depth. Low-level PD+DOB loops track wheel and pan setpoints; the IR ring enforces a reflex stop The perception pipeline enforces deterministic pairing of left-right frames using a shared trigger pulse and host-side timestamps, followed by rectification with precomputed maps and disparity via block matching or semi-global methods, depending on compute budget. Depth values are sampled at feature locations to improve numerical conditioning of the interaction matrix used later by the IBVS layer. The outer loop produces 𝜈13=[𝑣,𝜙󰇗𝑐]⊤ through a pseudoinverse built on the composite visual Jacobian 𝐽vis(𝜙𝑐,𝑍)=𝐿‾(𝑓,𝑍) ⬚𝑐𝐽13(𝜙𝑐), whereas base yaw aligns to the pan angle through a simple proportional law 𝜔=𝐾𝑜𝜙𝑐. Commands are rate-limited and low-pass filtered prior to conversion into wheel and pan setpoints, which are closed by PD controllers augmented with disturbance observers to mitigate friction and load variations. Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 184 Table 1 lists the principal parameters and operating rates used in the experimental section. Vision runs at 10-20 Hz per stream at 160×120 pixels, with the IBVS node cycling at 50-100 Hz, and the motor-layer loops at 500-1000 Hz. Communications rely on WebSocket for image transport and ROS topics/services for control and logging; a watchdog supervises message freshness. The operating envelope is an indoor planar area; experiments limit peak linear speed, enforce braking when proximity sensors detect short clearance, and log all telemetry for reproducibility. These design choices balance sensing throughput, controller bandwidth, and actuation limits so that the subsequent analysis is grounded on a feasible real-time configuration. Table 1 Key System Parameters and Operating Rates Quantity Symbol Value / Range Wheel radius r 0.024 m Wheelbase l 0.25 m Pan angle φc measured (encoder) Baseline B fixed (rigid mount) Focal (effective) f (fx + fy)/2 Depth model Z Z = (f B)/ d Vision resolution — 160 × 120 (mono) Vision rate (per cam) — 10–20 Hz IBVS control rate — 50–100 Hz Motor control rate — 500–1000 Hz Speed cap (linear) v ≤ 0.2 m/s Yaw follow gain Ko tuned experimentally Safety stop — IR thresholded 3. Visual Servoing with Stereo Depth This section formulates the image-based control law that drives a differential mobile base equipped with a pan-mounted rigid stereo rig toward a desired visual configuration. We consider a rectified stereo pair with fixed extrinsics between the left and right cameras; the pan joint rotates the entire rig about a vertical axis by an angle 𝜙𝑐 without altering the baseline geometry. The control objective is to regulate a set of image features 𝑓=[𝑢𝑖,𝑣𝑖]𝑖=1..𝑘 to a target 𝑓𝑑, while the base yaw 𝜔 aligns to the pan angle so that the camera ultimately faces the robot heading. Pixels are mapped to normalized coordinates through the calibrated intrinsics, namely 𝑥=𝑢−𝑐𝑥𝑓𝑥 ⁄ and 𝑦=𝑣−𝑐𝑦𝑓𝑦 ⁄, which renders the interaction equations well scaled across lenses and resolutions. We assume local scene rigidity and modest camera motions over one sampling interval; these common conditions hold in the indoor, planar settings used for our experiments and allow a concise derivation of the stacked interaction model. Depth at each feature is obtained from stereo disparity after rectification. Let 𝐵 denote the rigid baseline and 𝑓= (𝑓𝑥+𝑓𝑦)2 ⁄ the effective focal length in pixels; for a measured disparity 𝑑(𝑢,𝑣) we compute: 𝑍(𝑢,𝑣)=𝑓 𝐵 𝑑(𝑢,𝑣) ……… (6) and sample 𝑍𝑖=𝑍(𝑢𝑖,𝑣𝑖) at the feature locations. A first-order approximation quantifies how disparity noise propagates to depth: if 𝜎𝑑2 denotes the local disparity variance, then: 𝜎𝑍2 ≈ (𝜕𝑍 𝜕𝑑)2𝜎𝑑2 = (𝑓𝐵 𝑑2)2𝜎𝑑2 …….. (7) Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 185 which later informs feature weighting. We operate the capture at 160×120 pixels with 10–20 Hz per stream to reduce latency and stabilize timing, and track either fiducials or salient corners on the rectified left image. For a feature with normalized coordinates (𝑥,𝑦) and depth 𝑍, the standard interaction matrix that links the camera twist to image-velocity reads: 𝐿𝑖(𝑥,𝑦,𝑍)=[−1𝑍 ⁄0 𝑥𝑍 ⁄𝑥𝑦 −(1+𝑥2)𝑦 0 −1𝑍 ⁄𝑦𝑍 ⁄1+𝑦2−𝑥𝑦 −𝑥]∈ℝ2×6 ……… (8) Stacking 𝑘 features produces 𝐿‾(𝑓,𝑍)=𝑏𝑙𝑘𝑑𝑖𝑎𝑔(𝐿1,…,𝐿𝑘)∈ℝ2𝑘×6. Let the camera spatial twist be ⬚𝑠𝑤𝑐= [ ⬚𝑠𝑝󰇗𝑐⊤, ⬚𝑠𝜔𝑐⊤]⊤∈ℝ6. The commanded generalized velocities of the mobile platform with pan are: 𝜈=[𝑣 𝜔 𝜙󰇗𝑐]⊤∈ℝ3 ………… (9) and the kinematic mapping from commands to camera twist is: ⬚𝑠𝑤𝑐 = ⬚𝑐𝐽(𝜙𝑐) 𝜈 , ⬚𝑐𝐽(𝜙𝑐)∈ℝ6×3 …………… (10) whose explicit form depends on the camera placement relative to the base and the pan axis. We partition ⬚𝑐𝐽(𝜙𝑐)= [ ⬚𝑐𝐽13(𝜙𝑐) ⬚𝑐𝐽2(𝜙𝑐) ], where ⬚𝑐𝐽13∈ℝ6×2 selects the {𝑣,𝜙󰇗𝑐} channels used in the image-space inversion, and ⬚𝑐𝐽2∈ ℝ6×1 corresponds to the base yaw 𝜔. The composite visual Jacobian that enters the inversion is defined as: 𝐽vis(𝜙𝑐,𝑓,𝑍) = 𝐿‾(𝑓,𝑍) ⬚𝑐𝐽13(𝜙𝑐) ∈ ℝ2𝑘×2 …………. (11) Selecting the pair {𝑣,𝜙󰇗𝑐} yields a two-column matrix whose full-column-rank condition is met for non-degenerate feature configurations and camera placements that avoid the pathological alignments reported for full three-channel inversions. The use of measured depths 𝑍𝑖 from (6) improves conditioning relative to monocular approximations, particularly when the target lies at moderate range or when the image spans a narrow field of view. Let 𝑒=𝑓𝑑−𝑓∈ ℝ2𝑘 denote the feature error; a damped, weighted least-squares provides a minimum-norm command for the two controlled channels: 𝜈13=(𝐽vis ⊤𝑊𝐽vis +𝜆2𝐼2)−1𝐽vis ⊤𝑊 𝐾img 𝑒, 𝜈13=[𝑣 𝜙󰇗𝑐]⊤ …. (12) where 𝜆≥0 is a Tikhonov parameter, 𝑊=𝑑𝑖𝑎𝑔(𝑤1,𝑤1,𝑤2,𝑤2,…) weights per-feature channels (e.g., from 𝜎𝑍, tracker confidence), and 𝐾img ≻0 is diagonal. The base yaw follows the pan angle through a proportional law that closes the orientation gap without switching modes: 𝜔 = 𝐾𝑜 𝜙𝑐, 𝐾𝑜>0 …………. (13) This separation keeps the image-space inversion free from the singularities that appear when 𝜔 is treated as a coequal control input in geometries that reduce the rank of ⬚𝑐𝐽(𝜙𝑐). The discrete-time implementation respects the heterogeneous sampling of perception and actuation. At each vision cycle, we compute 𝜈13 raw from (12), then apply a first-order filter to temper latency and quantization: 𝜈13[𝑛]=𝛼 𝜈13 raw[𝑛]+(1−𝛼) 𝜈13[𝑛−1], 𝛼= 𝑇𝑠 𝑇𝑠+𝜏LPF …………. (14) where 𝑇𝑠 is the vision sampling period and 𝜏LPF a small time constant. Command bounds are enforced by projection; a simple quadratic program clips the solution without discontinuities: 𝑚𝑖𝑛 𝜈13 ∥𝜈13−𝜈13[𝑛]∥22 s.t. 𝑣𝑚𝑖𝑛≤𝑣≤𝑣𝑚𝑎𝑥, |𝜙󰇗𝑐|≤𝜙󰇗𝑐,𝑚𝑎𝑥 ………….. (15) Outliers and dropouts are handled by rejecting features with inconsistent motion or disparity (e.g., |𝑑|<𝑑𝑚𝑖𝑛 or large photometric residuals), temporarily reducing their weights 𝑤𝑖, and, if needed, falling back to a monocular surrogate 𝑍∗ Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 186 with conservative gains until valid stereo recovers. Under full-column-rank 𝐽vis , sufficiently small 𝑇𝑠, and bounded delays, the linearized error dynamics around the operating point are exponentially stable for diagonal 𝐾img chosen within standard IBVS ranges; the separation with (13) yields smooth base alignment and avoids mode switching. Table 2 Controller Parameters and Rates for the Stereo-IBVS Layer Quantity Symbol Nominal Value / Range Image gains Kimg diag(ku, kv), 0.5–2.0 Damping (DLS) λ 10⁻3–10⁻¹ Feature weights W from σz, confidence Yaw follow gain Ko 0.6–1.2 s⁻¹ LPF time constant τLPF 60–120 ms Vision rate (per cam) — 10–20 Hz @ 160 × 120 IBVS update rate — 50–100 Hz Speed cap (linear) vmax 0.2 m/s Pan rate cap φ c,max 0.6–1.0 rad/s Disparity threshold dmin dataset dependent 4. Low-Level Control & Implementation This section describes the real-time actuator layer that tracks the high-level commands 𝜈=[𝑣, 𝜔, 𝜙󰇗𝑐]⊤ generated by the visual servoing loop. The ARMOS TurtleBot is a differential-drive platform with wheel radius 𝑟=0.024 m and wheelbase 𝑙=0.25 m; the pan joint carries the rigid stereo rig and provides the measured angle 𝜙𝑐. Wheel angular speeds 𝜔𝑅 and 𝜔𝐿 follow the standard inverse kinematics: 𝜔𝑅=𝑣𝑟+𝑙2𝑟 𝜔, 𝜔𝐿=𝑣𝑟−𝑙2𝑟 𝜔 ………. (16) while the pan-rate command is 𝜙󰇗𝑐 itself. We collect the actuator references as 𝜙󰇗 𝑑=[𝜔𝑅𝑑, 𝜔𝐿𝑑, 𝜙󰇗𝑐𝑑]⊤, subject to speed caps |𝑣|≤𝑣𝑚𝑎𝑥 and |𝜙󰇗𝑐|≤𝜙󰇗𝑐,𝑚𝑎𝑥 motivated by sensing and power limits; in experiments we adopt 𝑣𝑚𝑎𝑥=0.2 m/s and 𝜙󰇗𝑐,𝑚𝑎𝑥∈[0.6,1.0] rad/s. The closed loops for the two wheels and the pan axis share an identical per-axis structure that combines a proportionalderivative (PD) servo with a disturbance observer (DOB) to compensate friction and load changes without integral action. Each axis ⋆∈{𝑅,𝐿,𝑐} is modeled nominally by a first-order plant 𝑃𝑛,⋆(𝑠)=𝐾⋆ 𝜏⋆𝑠+1 from input 𝑢⋆ (voltage or duty cycle) to measured velocity 𝑦⋆≡𝜔⋆. The DOB uses a causal low-pass 𝑄⋆(𝑠)=𝜔𝑐,⋆ 𝑠+𝜔𝑐,⋆ to estimate and reject lumped disturbances, leading to the control law: 𝑢⋆ = 𝐾𝑝,⋆ (𝜔⋆𝑑−𝜔⋆) + 𝐾𝑑,⋆ (𝜔󰇗⋆𝑑−𝜔󰇗⋆) − 𝑑󰆹⋆, 𝑑󰆹⋆ = 𝑄⋆(𝑠)(𝑦⋆−𝑃𝑛,⋆(𝑠) 𝑢⋆)…….(17) The resulting inner loops track 𝜙󰇗 𝑑 with short settling times and high disturbance rejection, while avoiding windup effects and reducing sensitivity to parameter drift. On the electrical side, L298N-class drivers supply each motor channel with current limiting and thermal protection; duty-cycle saturation and slew-rate bounds are enforced in software to respect these constraints. A fast motor loop runs at period 𝑇𝑚∈[1,2] ms, where encoder counts are latched and converted to velocities. Let 𝑁⋆ be the counts per revolution (after gearing) and 𝛥counts⋆[𝑘] the increment over one period; the discrete-time velocity estimate is: Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 187 𝜔⋆[𝑘] = 2𝜋 𝑁⋆ 𝛥counts⋆[𝑘] 𝑇𝑚, ⋆∈{𝑅,𝐿,𝑐} ………. (18) The derivative term uses a noise-aware differentiator with prefilter: 𝜔󰇗⋆[𝑘] = 𝛽 𝜔󰇗⋆[𝑘−1] + (1−𝛽) 𝜔⋆[𝑘]−𝜔⋆[𝑘−1] 𝑇𝑚, 𝛽∈[0,1) ………. (19) which damps quantization effects at high rates. DOB and PD are discretized by Tustin or matched-𝑧 methods, yielding stable realizations under the chosen sampling. Command preprocessing ensures that actuator demands remain consistent with plant and driver limits as well as with perception throughput. Before applying (16), the high-level vector 𝜈 is smoothed by a first-order low-pass filter with time constant matched to the vision-update interval, reducing jitter due to asynchronous sensing. After mapping, a projection step clips the two controlled channels 𝑣 and 𝜙󰇗𝑐 without introducing discontinuities: 𝑚𝑖𝑛 𝜈13 ∥𝜈13−𝜈13 raw ∥22 s.t. 𝑣𝑚𝑖𝑛≤𝑣≤𝑣𝑚𝑎𝑥, |𝜙󰇗𝑐|≤𝜙󰇗𝑐,𝑚𝑎𝑥, 𝜈13=[𝑣 𝜙󰇗𝑐]⊤ ……….(20) with 𝜈13 raw the undamped output of the visual loop. At the motor level, per-axis PWM commands are limited to |𝑢⋆|≤ 𝑢𝑚𝑎𝑥 and slewed by |𝑢⋆[𝑘]−𝑢⋆[𝑘−1]|≤𝛥𝑢𝑚𝑎𝑥 to avoid current spikes. A reflex brake is triggered by the infrared ring when the measured clearance falls below a threshold, and a watchdog drops the platform into a safe idle if command freshness exceeds a fixed timeout. Odometry and calibration support reliable feedback and help detect slip conditions that may degrade servoing. Pan zeroing is performed at start-up using an index mark; wheel directions follow a right-handed convention so that positive 𝜔 corresponds to counterclockwise body yaw. At a slower rate, the encoder-based body velocities: 𝑣enc =𝑟2(𝜔𝑅+𝜔𝐿), 𝜔enc =𝑟𝑙(𝜔𝑅−𝜔𝐿) ………. (21) are compared with their commanded counterparts to build residual tests. If |𝑣enc −𝑣|>𝜀𝑣 or |𝜔enc −𝜔|>𝜀𝜔 persist beyond a dwell time, gains are temporarily reduced and the outer loop is notified to downscale the visual commands, which improves robustness on low-traction surfaces. Table 3 reports the parameters, limits, and rates used across the three axes so that the results in Section 6 can be reproduced. Values are given as nominal ranges reflecting actuator diversity and minor mechanical variations across builds. The chosen gains lead to 10–20% overshoot at step tests with brisk settling, while DOB cutoffs 𝜔𝑐,⋆ sit above the dominant disturbance bandwidth yet below the encoder noise corner. The motor loop maintains headroom above 30% CPU at the stated rates, and logging is buffered in a lower-priority thread to prevent jitter. These implementation details ensure that the low-level layer meets its deadlines and exposes well-behaved interfaces to the visual servoing module. Table 3 Low-Level Parameters, Limits, and Rates (per axis ⋆∈{𝑅,𝐿,𝑐}) Quantity Symbol Nominal Value / Range Wheel radius / wheelbase r, l 0.024 m, 0.25 m Motor nominal model Pn, ★ (s) K ★ / (τr, ★ s + 1) PD gains Kp, ★ , Kd, ★ tuned (step tests) DOB filter Q ★ (s) ωc, ★ / (s + ωc, ★ ) DOB cutoff ωc, ★ 20–80 rad/s Motor-loop period Tm 1–2 ms Derivative prefilter β 0.6–0.9 Linear speed cap vmax 0.2 m/s Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 188 Pan rate cap φ c,max 0.6–1.0 rad/s PWM magnitude / slew vmax; Δvmax driver-limited Safety threshold (IR) — lab-calibrated Watchdog timeout — 100–200 ms 5. Experimental Setup The experiments were conducted in an indoor laboratory bay with a planar floor and controlled lighting to reduce shadows and glare. The arena provided straight segments for line-of-sight tracking and lateral traverses, with marked start and goal zones as well as a safety buffer around the perimeter. To remain within sensing and actuation limits, the linear speed was capped at 𝑣𝑚𝑎𝑥=0.2 m/s, and the pan-rate was limited to |𝜙󰇗𝑐|≤1.0 rad/s. A hardware e-stop and a watchdog on command freshness were active throughout; the infrared proximity ring enforced a reflex stop when the estimated clearance dropped below a calibrated threshold. Power was delivered from an onboard battery pack sized for the full sensor load; drivers and regulators were dimensioned to avoid brownouts under peak demand. Figure 2 illustrates the arena and the robot configuration used for all trials, including the stereo mount and fiducial placement. Figure 2 Experimental arena and platform. The ARMOS TurtleBot carries a rigid stereo pair on a single-axis pan joint and an infrared safety ring. Start/goal areas and a lateral travel path for moving-target trials are marked on the floor. The mobile base is the ARMOS TurtleBot, modeled as a differential drive with wheel radius 𝑟=0.024 m and wheelbase 𝑙=0.25 m. Wheel encoders provide odometry and velocity feedback; a single-axis pan joint supports the camera rig and supplies the angle 𝜙𝑐 via an incremental encoder with index. Drivers of the L298N class interface the motor channels and enforce current and thermal limits; duty-cycle saturation and slew-rate constraints are applied in software to remain inside the electrical envelope. At a slower supervisory rate, encoder-based body velocities 𝑣enc =𝑟2(𝜔𝑅+𝜔𝐿), 𝜔enc =𝑟𝑙(𝜔𝑅−𝜔𝐿) …….. (22) are compared with commands to detect slip or traction loss, which gates the visual controller when residuals exceed thresholds. Initial conditions were set with |𝜙𝑐|<30∘ and a starting distance in the 1–3 m range to align with the camera field of view and the disparity operating regime. The vision payload comprises two OV7670 sensors rigidly mounted with baseline 𝐵 and actuated together by the pan joint, so that left–right extrinsics remain constant during yaw motion. Each camera feeds an ESP32 module that streams frames over WebSocket to a host machine, where rectification and disparity are computed. Primary operation used 160×120 pixels at 10–20 Hz per stream to reduce end-to-end delay; an optional 320×240 mode at 5–10 Hz was reserved for ablations. Stereo calibration produced intrinsics (𝑓𝑥,𝑓𝑦,𝑐𝑥,𝑐𝑦) and rectification maps; disparity 𝑑(𝑢,𝑣) was converted to depth by: Global Journal of Engineering and Technology Advances, 2025, 24(02), 181-195 189 𝑍(𝑢,𝑣)=𝑓 𝐵 𝑑(𝑢,𝑣), 𝑓=𝑓𝑥+𝑓𝑦 2 ………. (23) with per-feature sampling at tracked points on the rectified left image. Targets were realized with high-contrast fiducials (e.g., AprilTag) sized to be detectable across the stated range; natural features were enabled as fallback when fiducials were occluded. Figure 3 summarizes timing and synchronization across sensing and control layers. Figure 3 Timing diagram. Dual camera capture (Tcam), host-side rectification and disparity (Tstereo), IBVS update (Tvs), and motor loop (Tm) are shown with timestamps and a GPIO trigger for frame pairing. End-to-end latency is the sum Le2e = Tcam + Ttx + Tstereo + Ttrack + Tq. Frame pairing relied on a shared GPIO trigger and host-side timestamps; the IBVS layer executed at 50–100 Hz, while the motor loop ran at 𝑇𝑚∈[1,2] ms. The moving-target scenario used a lateral trajectory with nominal speed 𝑣𝑡≈ 0.2 m/s along a straight path at approximately constant depth. Two scenarios were tested: (A) a static target within 1–3 m, and (B) a target moving laterally for ∼5 s at the stated speed. Baselines included a fixed-camera condition (pan disabled) and a monocular surrogate using a nominal depth 𝑍∗, with all other parameters held equal for fair comparison. All runs were logged with synchronized timestamps for image features 𝑓,𝑓𝑑, depths 𝑍𝑖, commands 𝜈13=[𝑣,𝜙󰇗𝑐]⊤, yaw command 𝜔=𝐾𝑜𝜙𝑐, wheel and pan velocities, PWM, and proximity distances. Table 4 collects the principal experimental parameters and rates to support replication. Quality control required valid calibration residuals, consistent frame pairing over each run, and encoder health checks before motion. Trials were rejected when feature tracks were lost for more than a short dwell time or when network dropouts exceeded a pre-set percentage; at least five valid trials were recorded per scenario. Experiments were conducted in controlled indoor spaces without bystanders; procedures followed laboratory safety guidelines and institutional policies. These conditions provide a well-specified operating envelope that matches the modeling assumptions used by the control design. Table 4 Experimental Parameters and Operating Rates Quantity Symbol Value / Range Wheel radius / wheelbase r, l 0.024 m, 0.25 m Stereo baseline / focal (eff.) B, f fixed, (fx + fy)/2 Vision resolution / rate — 160×120 @ 10–20 Hz Optional vision mode — 320×240 @ 5–10 Hz IBVS update rate — 50–100 Hz