Full text
The project is carried out within the framework of the National Recovery and Resilience Plan Greece 2.0, funded by the European Union - NextGenerationEU (Implementation Body: HFRI) CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Abstract: This deliverable reports on the multimodal skill encoding methods employed, enhanced and developed in the CARPOS project, highlighting their generalization capabilities with respect to variations of the fruit’s position, orientation and size. The developed encoding mechanism is able to fuse data from multiple visual demonstrations towards the skill encoding and, in turn, allow the refinement of the learned skill through kinesthetic teaching. In the core of the developed method is a Dynamic Movements Primitives (DMP) model that encodes the pose of the robot (oriented trajectory), while the data fusion is achieved utilizing a variant of the Extended Kalman Filter (EKF). Furthermore, the forces applied by the fingertip are recorded and encoded using an Radial Base Function (RBF) Network. Lastly, the research activities related to this deliverable also include the investigation of employing LLMs (such as the ChatGPT) for segmenting a fruit picking motion pattern and classifying the segments into primitive actions, e.g. pull, slide, tilt etc; the related work will be presented as a poster in the 34th IEEE International Conference on Robot and Human Interactive Communication (IEEE RO-MAN 2025), attached at the Appendix of this deliverable. The encoding mechanism is evaluated on the CARPOS robotic system (Husky clear-path platform with a UR10e manipulator), demonstrating its adaptability and generalization capabilities in different fruit sizes and poses, as well as its ability to accumulate and encode knowledge from the human. Dissemination Level: PU
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Περίληψη στα Ελληνικά Αυτό το παραδοτέο καταγράφει τις μεθόδους πολυτροπικής κωδικοποίησης δεξιοτήτων που αναπτύχθηκαν στο έργο CARPOS, επισημαίνοντας τις δυνατότητες γενίκευσης σε σχέση με τις μεταβολές της θέσης, του προσανατολισμού και του μεγέθους των φρούτων. Ο μηχανισμός κωδικοποίησης που αναπτύχθηκε είναι σε θέση να συγχωνεύσει δεδομένα από πολλαπλές οπτικές επιδείξεις για την κωδικοποίηση δεξιοτήτων και, στην συνέχεια, να επιτρέψει τη βελτίωση της δεξιότητας μέσω της κιναισθητικής διδασκαλίας. Στον πυρήνα της αναπτυγμένης μεθόδου βρίσκεται ένα μοντέλο Dynamic Movement Primitives (DMP) που κωδικοποιεί την γενικευμένη θέση του ρομπότ (προσανατολισμένη τροχιά), ενώ η συγχώνευση επιτυγχάνεται χρησιμοποιώντας μια παραλλαγή του Εκτεταμένου Φίλτρου Kalman (EKF). Επιπλέον, οι δυνάμεις που ασκούνται από το άκρο του κύριου δακτύλου της αρπάγης καταγράφονται και κωδικοποιούνται χρησιμοποιώντας ένα δίκτυο Συναρτήσεων Ακτινικής Βάσης (RBFs). Τέλος, οι ερευνητικές δραστηριότητες που σχετίζονται με αυτό το παραδοτέο περιλαμβάνουν επίσης τη διερεύνηση της χρήσης μοντέλων LLM (όπως το ChatGPT) για την κατάτμηση ενός μοτίβου κίνησης συλλογής φρούτων και την ταξινόμηση των τμημάτων σε πρότυπες ενέργειες, π.χ. τράβηγμα, ολίσθηση, στρίψιμο κ.λπ. Η σχετική εργασία θα παρουσιαστεί ως poster στο 34th IEEE International Conference on Robot and Human Interactive Communication (IEEE RO-MAN 2025), το οποίο επισυνάπτεται στο Παράρτημα αυτού του παραδοτέου. Ο μηχανισμός κωδικοποίησης αξιολογείται στο ρομποτικό σύστημα CARPOS (πλατφόρμα Clearpath Husky με ρομποτικό βραχίονα UR10e), επιδεικνύοντας την προσαρμοστικότητα και τις δυνατότητες γενίκευσης σε διαφορετικά μεγέθη και πόζες φρούτων, καθώς και την ικανότητά του να συσσωρεύει και να κωδικοποιεί γνώση από τον άνθρωπο. Document Status Document Title Multimodal skill encoding and generalization in fruit geometry variations Version 1.0 Work Package 4 - Skill encoding Deliverable D7 Contributors CARPOS research team Due Date 15/07/2024 Completion Date 15/07/2024 Confidentiality PU
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Document Change Log Each change or set of changes made to this document will result in an increment to the version number of the document. This change log records the process and identifies for each version number of the document the modification(s) which caused the version number to be incremented. Change Log Version Date Template Creation 0.1 Apr 11th 2024 Draft document sent to the whole team for revision 0.5 Jul 4th 2025 Final version 1.0 Jul 15th 2025 INTRODUCTION 1.1 Introduction and problem description CARPOS project aims at developing a coachable robotic system, able to learn from human demonstrations, through multiple modalities, such as visual observations and kinesthetic teaching. It is therefore evident that the “encoding” of the skill that is captured through this process is of imperative importance. On the one hand, in order for the skill to be generalized in position, orientation and size variations1, the skill encoding mechanism has to possess generalization capabilities, which generally could be solved by using a relatively large training dataset. On the other hand, given the involvement of the human in the training loop and the limited resources (e.g. plants) during the training procedure, the minimization of the number of demonstrations is also critical, as each demonstration involves the detachment of one fruit and requires a non-negligible time from the human-teacher. To tackle these conflicting challenges, the research team of CARPOS employed and enhanced the Dynamic Movement Primitives (DMP) model, which is able to encode a skill even based only on a single demonstration, by providing it with multimodal learning capabilities, after taking into account the uncertainty involved in each modality. The proposed encoding mechanism fuses the visual data using a variant of the Extended Kalman Filter for yielding the “training dataset” of the DMP model in each iteration of demonstration. As detailed in D4 (“Control schemes for occlusion-free active perception”), we consider a visual sensor (i.e. camera), attached to the robot’s end-effector, which is actively moving in order to maximize the information gain about the human demonstration. This change of perspective, from iteration to iteration, means that the uncertainty of the estimated human hand changes accordingly. The developed data fusion mechanism, reported in this deliverable, takes into account the uncertainty of the visual data at each iteration, and complementary uses the data in 1 During the execution of the autonomous harvesting task the fruits are located in different positions and orientations, and come in different sizes.
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS order to optimally extract the spatial and temporal characteristics of the skill. During the last step of the training procedure, this of the inspection and kinesthetic correction, the visually trained skill is used as the basis of the motion and the human can complementary do final small modifications to the task along the directions with the higher uncertainty. This part essentially fuses the visually perceived data with some complementary ones from the kinesthetic corrections. The human-user can intervene whenever she/he deems it necessary through the application of forces (more details can be found in D5 - “Controller for unintentional and intentional pHRI during the coaching and autonomous operation”). For enabling the encoding of the forces that should be exerted by the main finger (thumb) of the gripper for a successful fruit detachment, the gripper is equipped with a force sensor that registers the forces applied towards its contact axis in real-time. To this aim, the force measurements are recorded and encoded through a Radial Base Function (RBF) interpolation, during the kinesthetic teaching. Lastly, for providing high level semantic understanding capabilities to the system during the training procedure, we perform an initial investigation on the capabilities of LLMs, such as ChatGPT, for segmenting a complete fruit picking motion sequence into primitive patterns, such as “pushing”, “twisting” etc. Our aim is to automatically extract the sequence of primitive actions that this specific fruit picking operation requires. This last investigation was not included in the initial research plan (i.e. the Technical Document), but it constitutes the first step towards the extension of the methods developed in CARPOS project for encoding a motion pattern. This investigation and its results will be presented as a Late Breaking Results (LBR) poster at the 34th IEEE International Conference on Robot and Human Interactive Communication (IEEE RO-MAN 2025); the related poster, as well as the corresponding technical report are provided in the Appendix of this deliverable. 1.2 Relation to Other Deliverables The purpose of this deliverable is to present the methods developed within the CARPOS project for tackling the problem of encoding a fruit picking skill. The deliverable is closely related to D3 - “Fruit Picking Action Identification and Skill Extraction from Visual Observation” and D4 - “Control schemes for occlusion-free active perception”, as the data fusion mechanism (presented here), the perception module and the control scheme for occlusion avoidance are interconnected. In particular, the purpose of the control schemes developed in D4 is to minimize the uncertainty of the pose of the human hand, which is essentially provided by the data fusion mechanism (detailed in this Deliverable), while the fusion mechanism, in turn, depends on the pose of the camera which is dictated by the aforementioned control methods. Furthermore, the deliverable is related to D5 - “Controller for unintentional and intentional pHRI during the coaching and autonomous operation”, as after the multiple iterations of visual observation of the human activity, the skill is kinesthetically refined and the encoding mechanism accounts for the physically merged visual and kinesthetic data. Lastly, as the encoding mechanism detailed in this deliverable constitutes the core of the CARPOS framework, this deliverable is also related to D8 – “System integration and Testing”.
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS METHODOLOGY As opposed to the explicit robot programming, Learning by Demonstration (LbD) was recently proposed to provide methods that make the system able of learning intuitively and seamlessly from human actions [1, 2]. Although relieving humans from the burden of possessing technical knowledge, LbD still requires a number of human demonstrations, in the form of time-series of the pose (position and orientation) of the human hand. Such demonstrated motions are captured, in the literature, either through external means, such as cameras or magnetic trackers [3-5], or through kinesthetic teaching, i.e. by directly interacting with the robot through the exertion of contact forces [6-8]. To encode the skill, all encoding mechanisms involve a number of parameters that has to be “trained”. Towards this direction, the “learning” procedure consists essentially of finding the optimal parameters that make the system yield a behavior which is closest to the demonstrated motion, as shown in Fig. 1. However, one of the key features of such mechanism concerns the generalization of the motion in different conditions, such as new initial states, new targets or even different task orientations. Figure 1 – Learning a demonstrated motion and generalizing in space Due to their real-time generation and generalization (spatial and temporal) capabilities, the most dominant encoding mechanisms, in robotics literature, are the ones that employ Dynamical Systems (DS). According to DSs, the motion reference is generated in real-time based on the state of the system and the environmental condition. The most popular DSs for encoding the spatial and temporal characteristics of a motion are the Gaussian Mixture Models (GMMs) [9, 10] and the Dynamic Movement Primitives (DMPs) [11, 12]; one can find more details about the current DSs for encoding motion patterns in [13]. Some of the advantages of the DMPs over the GMMs is that they do not require a large training dataset, which means that the system can “learn” based only on some demonstrations (even based on a single demonstration), as well as their low computation complexity. On the other hand, although GMMs possess better generalization capabilities, their performance, in terms of encoding and generalization, depends proportionally on the number of the motion demonstrations provided as a training dataset. In our Demonstration (one or more) Training Generalization Demonstration motion Robot’s kinematic behavior Target Initial pose New target Figure 2 – The considered {F} and {G} frames
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS case, in which the minimization of the required demonstrations is a crucial aspect, as the human has to perform the task in each iteration and the detachment of a fruit is involved, we selected the utilization of the DMPs for encoding the motion of the human hand during the task. As detailed in Deliverable 3 - “Fruit Picking Action Identification and Skill Extraction from Visual Observation”, in the initial step of the visual observation of the human action, the human hand grasp pose is captured with respect to the pose of the fruit, as depicted in Fig. 2. Let {𝐹} and {𝐺} be the coordinate frame of the fruit and human-hand respectively; which are defined based on six points of interest on the fruit and human-hand, as detailed in D3. This grasp information constitutes the basis of the learned skill, and it is stored as a part of the acquired knowledge. Let 𝒈∈𝑆𝐸(3) be the grasp pose of the human hand with respect to the fruit, in a homogeneous transformation form. Given that the grasp pose is stored as a part of the skill knowledge, the motion of the human hand is encoded with respect to its initial pose. In this way, during the autonomous execution of the motion pattern, after ensuring that the grasping of the fruit is identical to the one performed by the human, the motion is executed with respect to the initial pose of the robot’s gripper, which in turn depends on the exact position and orientation of the fruit. This ensures that the motion will be generalized to different positions and orientations of the fruits, as depicted in Fig. 3. Let 𝒈(𝑡)=𝑹(𝑡) 𝒑(𝑡) 𝟎× 1 ∈𝑆𝐸(3) be the pose of the human hand during the skill demonstration with respect to its initial pose, where 𝑹(𝑡)∈𝑆𝑂(3),𝒑(𝑡)∈ℝ are the orientation (in a rotation matrix form) and position (in Cartesian coordinates) of the hand with respect to its initial pose respectively, and 𝑡∈[0,𝑇] the time, with 𝑇∈ℝ being the total duration of the motion. The goal of the encoding mechanism is, based on this demonstrated (or data fused) pattern, to find the optimal parameters of the DMP model, in order its behavior to Case 1 Case 2 {F} {F} Reference motion based on the learned pattern Reference motion based on the learned pattern Figure 3 – Two different cases involving different poses of fruit
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS optimally mimic the given pattern. In the following section preliminaries about the specific DMP model variant utilized are provided. 2.1 Preliminaries on the Dynamic Movement Primitives model in SE(3) For the sake of simplicity, let us first explain the encoding of a scalar variable (e.g. a single axis of the motion of the end-effector) utilizing a Dynamic Movement Primitives (DMP) model. Let 𝑦(𝑡)∈ ℝ be the variable, whose behavior we want to encode and 𝑦 (𝑡)∈ℝ a demonstrated trajectory or a trajectory estimate after the fusion of multiple demonstrations, where 𝑡∈[0,𝑇] with 𝑇∈ ℝ being the total duration of the motion. The DMP model is a second order dynamical system of the following form: 𝜏 𝑦 = 𝑎 ( 𝑏 ( 𝑦 − 𝑦 ) − 𝜏 𝑦 ) + 𝑠 𝑓 ( 𝑥 ; 𝒘 ) , 𝜏 𝑥 = 𝑔 ( 𝑥 ) , (1) where 𝑦≜ ,𝑦≜ , 𝑦∈ℝ the final/goal value of 𝑦, i.e. 𝑦(𝑡)→𝑦, 𝜏, 𝑠∈ℝ the time scaling and spatial scaling parameters respectively, 𝑎,𝑏∈ℝ the constant parameters of the linear part of the DMP, and 𝑥∈ ℝ an auxiliary variable called the “phase variable” that replaces time, in order for the system to be time-independent. Notice that the dynamics of 𝑥 are governed by the function 𝑔(𝑥):ℝ→ℝ (the so-called “canonical system”), for which there are multiple different selections in the literature; the most common ones are 𝑔(𝑥)≜−𝑎𝑥 (the so-called “exponential” dynamics) and 𝑔(𝑥)≜1 with saturation such that 𝑥∈[0,1] (the so-called “linear” dynamics). However, the part which essentially encodes the motion pattern is the non-linear term of the DMP, which is represented by the function 𝑓(𝑥 ;𝒘):ℝ→ℝ, which is called the “forcing” term, where 𝒘∈ℝ are the parameters that are adjusted in order for the behavior of the dynamic system to optimally reflect the demonstrated motion, i.e. 𝑦 (𝑡),∀𝑡∈[0,𝑇], with 𝑁∈ℕ being the pre-defined number of the paramters. More specifically, this non-linear term is selected to be a weighted summation of Radial Base Function (RBFs), with the most common selection for RBF being the Gaussian function. Therefore, function 𝑓(𝑥 ;𝒘) has the following form: 𝑓 ( 𝑥 ; 𝒘 ) ≜ ∑ 𝑤 𝜓 ( 𝑥 ) ∑ 𝜓 ( 𝑥 ) = 𝒘 ⊺ 𝜳 ( 𝑥 ) , (2) where 𝒘≜[𝑤…𝑤]⊺∈ℝ the weights of the kernels (i.e. the RBFs), 𝜓(𝑥):ℝ→ℝ the RBF and 𝜳(𝑥)≜[𝜓(𝑥)…𝜓(𝑥)]⊺∈ℝ the kernel base (RBF base). Notice that the kernel base is constructed in the one dimensional space of the phase variable, i.e. is a function of 𝑥. The Gaussian RBF, utilized in CARPOS project, is the following: 𝜓 ( 𝑥 ) ≜ 𝑒 ( ) , (3)
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS where 𝑐,ℎ∈ℝ the centers and width of the 𝑖-th kernel. For more details on how to select these variables, one can refer to [11-13]. An example of the RBF base is provided visually in Fig. 4 for 𝑁=10 and considering a linear canonical system. Figure 4 – The kernel base for N=10. The RBFs appear in different colors. The phase variable is selected to be governed by linear dynamics. For ensuring that the behavior of the DMP, as a motion generation mechanism, will optimally reflect the demonstrated motion, i.e. 𝑦 (𝑡), one has to solve the following optimization problem: 𝒘 ∗ = arg min 𝒘 𝑓 ( 𝑡 ) − 𝑓 ( 𝑥 ( 𝑡 ) ; 𝒘 ) 𝑑𝑡 , (4) where 𝑓(𝑡)= 𝑦(𝑡)−𝑎(𝑏(𝑦(𝑇)−𝑦(𝑡))−𝑦(𝑡)) is the calculated “ideal” forcing term based on the demonstration, with 𝑦(𝑇)∈ℝ the final value of 𝑦(𝑡) of the demonstration, considering (for the training) that the demonstrated motion is neither scaled in time (i.e. 𝜏=1), nor in space (i.e. 𝑠=1). As the motion is captured in discrete time, e.g. with 30 fps if it is captured through a camera, the optimization problem given in (4) is written in discrete time as follows: 𝒘 ∗ = arg min 𝒘 𝑓 , − 𝑓 ( 𝑥 ; 𝒘 ) . (5) One can yield the analytic solution of (5), using the Least Square (LS) solver, provided by: 𝒘 ∗ = 𝜳 𝓛 𝐟 , (6)
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS where 𝜳 ≜ ⎣ ⎢ ⎢ ⎢ ⎢ ⎡ 𝜳 ( 𝑥 ) ∑ 𝜓 ( 𝑥 ) ⋮ 𝜳(𝑥) ∑ 𝜓 ( 𝑥 ) ⎦ ⎥ ⎥ ⎥ ⎥ ⎤ ∈ℝ×, (7) with 𝑀∈ℕ being the number of samples within the time-series of the demonstrated motion, 𝐟≜ 𝑓,… 𝑓,⊺∈ℝ is the array of the values of 𝑓 at each time instance of the demonstration and 𝜳 𝓛≜𝜳 ⊺𝜳 𝜳 ⊺∈ℝ× is the left MoorePenrose pseudo-inverse of 𝜳 . An example of the encoding achieved by the DMP model for a single axis is shown in Fig. 5, in which one can compare the initial demonstrated motion, with the behavior of the DMP. Notice that the encoding accuracy is improved with the increase of the number of kernel functions used (i.e. 𝑁). As Fig. 5 presents the exact value (not the derivative) of 𝑦, one can see the function approximation of 𝑓(𝑡) achieved by 𝑓(𝑥 ;𝒘) after the training of the DMP in Fig. 6, which is the goal of the optimization problem of (4). Figure 6 – Function approximation of the forcing term achieved with N=20 kernels. Orange line: 𝑓𝑥(𝑡), blue line: 𝑓(𝑡). Figure 5 – Encoding of a demonstrated trajectory with a DMP, using 10 and 20 kernels
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Given that the DMP model is trained based on the state estimate provided at the end of the previous repetition, it incorporates the uncertainty of this estimate. Therefore, assuming a sufficiently accurate endcoding, 𝝎(𝑡) is considered to be a random variable, having a normal distribution with covariance 𝑸 (𝑡)∈𝑆 , with 𝑚∈ℕ being the number of states of the DMP, which is equal to the uncertainty of the total estimate after at end of the 𝑖− 1 repetition (i.e. the previous total estimate), or in other words, we set: 𝑸 ( 𝑡 ) = 𝑷 ( 𝑡 ) , ∀ 𝑡 ∈ [ 0 , 𝑇 ] , (14) where 𝑷 (𝑡)∈𝑆 the covariance of the uncertainty of the total estimate during the (𝑖−1)-th repetition, the calculation of which follows. Remark 1. Notice that for 𝑖=1, i.e. during the first repetition, there is no previous demonstrations and therefore one can consider only the linear part of the DMP with a relatively high uncertainty, e.g. 𝑸 (𝑡)=𝜎𝜤, with 𝜎∈ℝ being a relatively large value. ∎ For fusing the position measurement provided by the sensor, i.e. 𝐩 (𝑡), with the position prediction provided by the DMP model, which is denoted in this section by 𝐩(𝑡), we propose the following covariance weighted mean, as a data fusion mechanism: 𝐩 ∗ ( 𝑡 ) = 𝑸 ( 𝑡 ) + 𝚺 ( 𝑡 ) 𝑸 ( 𝑡 ) 𝐩 ( 𝑡 ) + 𝚺 ( 𝑡 ) 𝐩 ( 𝑡 ) . (15) Notice that (15) is essentially the weighted left pseudoinverse solution (weighted least squares) of the following forward relationship: 𝒑 ( 𝑡 ) 𝒑 ( 𝑡 ) = 𝑰 𝑰 𝐩 ∗ ( 𝑡 ) , (16) considering the following weight matrix 𝑾=diag 𝑸 (𝑡),𝚺(𝑡), in order to account more for the directions with less uncertainty. The variance-covariance matrix of the total estimate, i.e. 𝑷 (𝑡)∈𝑆 , can then be calculated by: 𝑷 ( 𝑡 ) = 𝑸 ( 𝑡 ) + 𝚺 ( 𝑡 ) . (17)
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Proof of (17): Given that 𝐩(𝑡) is characterized by the modelling uncertainty, having a covariance of 𝑸 (𝑡), and 𝐩 ⊺(𝑡) by the measurement covariance, i.e. 𝚺(𝑡) and by taking into account (15), one gets: 𝑷 ( 𝑡 ) = var 𝐩 ∗ = 𝑸 +𝚺 𝑸 𝟏 𝑸 𝑸 ⊺𝑸 +𝚺⊺+ 𝑸 + 𝚺 𝚺 𝟏 𝚺 𝚺 ⊺ 𝑸 + 𝚺 ⊺ . (18) Due to the fact that both 𝑸 and 𝚺 are symmetric positive definite matrices, we have 𝑸 ⊺= 𝑸 , 𝚺⊺=𝚺, 𝑸 +𝚺⊺=𝑸 +𝚺and therefore, (17) yields from (18). ∎ Lemma 1. An ellipsoid 𝒮(𝚨) is enclosed (or radially surrounded) by another ellipsoid 𝒮(𝚩), where 𝚨,𝚩∈𝑆 , if 𝒚⊺𝚨𝐲>𝒚⊺𝐁𝐲 for all 𝒚∈ℝ. Proof of Lemma 1: Let 𝒌∈ℝ be a unit vector pointing towards an arbitrary direction. We can show that 𝒮(𝚨) is enclosed by 𝒮(𝚩), by showing that 𝑎<𝑏, where 𝑎∈ℝ:(𝑎𝒌)⊺𝑨(𝑎𝒌)= 𝑎𝒌⊺𝑨𝒌=1 and 𝑏∈ℝ:(𝑏𝒌)⊺𝑩(𝑏𝒌)=𝑏𝒌⊺𝑩𝒌=1, for any direction 𝒌∈ℝ:‖𝒌‖=1. Notice that 𝑎 and 𝑏 represent the length of the linear segments within 𝒮(𝚨) and 𝒮(𝐁) respectively, along the direction of 𝒌. To this aim, we have 𝒌⊺𝑨𝒌= and 𝒌⊺𝑩𝒌= . Then 𝒌⊺𝑨𝒌>𝒌⊺𝑩𝒌 implies > and consequently 𝑎<𝑏. ∎ Theorem 1. The covariance ellipsoid of the total estimate at the 𝑖 -th repetition, i.e. 𝒮 𝑷 (𝑡) , is enclosed (i.e. radially surrounded): 1. By the covariance ellipsoid of the (𝑖−1)-th repetition, i.e. 𝒮𝑷 . In other words, the uncertainty of the estimate at each repetition will always be less than the uncertainty at the previous repetition, along any direction. 2. By both the covariance ellipsoid of the modelling error, i.e. 𝒮𝑸 , and the covariance ellipsoid of the measurement, i.e. 𝒮(𝜮). In other words, the uncertainty of the total estimate will always be less than both the measurement and the modelling uncertainty, along any direction. Proof of Theorem 1: Given that 𝑷 (𝑡)= 𝑸 (𝑡) from (14), 𝒮𝑷 will constitute the surface for which: 𝒚 ⊺ 𝑷 ( ) 𝐲 = 𝒚 ⊺ 𝑸 𝐲 = 1 , 𝒚 ∈ ℝ , (19) while 𝒮𝑷 will constitute the surface for which:
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS 𝒚 ⊺ 𝑷 𝐲 = 𝒚 ⊺ 𝑸 + 𝜮 𝐲 = 1 , 𝒚 ∈ ℝ . (20) Given that both 𝑸 and 𝜮 are positive definite matrices, the following will hold: 𝒚 ⊺ 𝑸 + 𝜮 𝐲 > 𝒚 ⊺ 𝑸 𝐲 . (21) After substituting (19) and (20) into (21), we have: 𝒚 ⊺ 𝑷 𝐲 > 𝒚 ⊺ 𝑷 ( ) 𝐲 , 𝒚⊺𝑷 𝐲> 𝒚⊺𝜮𝐲, 𝒚 ⊺ 𝑷 𝐲 > 𝒚 ⊺ 𝑸 𝐲 , (22) for all 𝒚∈ℝ. Therefore, according to Lemma 1, the first inequality of (22) proves the first part of Theorem 1, while the second and third inequalities prove the second part of it. ∎ Remark 2. Assuming that one can rotate the covariance ellipsoid of the measurement, i.e. 𝒮(𝜮), by rotating the sensor, the second part of Theorem 1 implies that the uncertainty can be optimally minimized by ensuring that 𝒮(𝜮) is orthogonal to 𝒮𝑸 , i.e., the major axis of 𝜮 is aligned to the minor axis of 𝑸 and vice versa (given that they belong to 𝑆 ). This optimization rationale is the core of the active perception controllers developed in the project, provided in Deliverable 4 – “Control schemes for occlusion-free active perception”. ∎ 2.3 Generalization to different fruit sizes The generalization to different fruit geometries essentially involves the scaling of the motion given the different a) position, b) orientation and c) size of the fruit. As mentioned above, the DMP is encoded considering the initial pose at the exact time of grasping as the reference frame. This ensures that, when the grasping pose is exactly the same with that demonstrated, the whole motion will be performed with respect to this orientation, as depicted in Fig. 3. Therefore, this selection tackles the problem of generalizing the motion, given the different position and orientation of each fruit during the execution of the pattern. For addressing the scaling of the motion based on the size of the fruit, we exploit a metric for the size provided by the perception module, developed in CARPOS project, as detailed in Deliverable 3 – “Fruit Picking Action Identification and Skill Extraction from Visual Observation”, e.g. the radius of the fruit. Let 𝑟∈ℝ be this size measure during the human demonstration, and 𝑟∈ ℝ be the one during the autonomous execution of the fruit picking activity. Then, the scaling term and the goal of the position part of the DMP, i.e. equation (8) above, is selected to be the following respectively:
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS 𝑠 = 𝑟 𝑟 , 𝒑 = 𝑟 𝑟 𝒑 , (23) where 𝒑∈ℝ is the final position of the wrist with respect to the initial frame of the wrist, during the demonstration. Notice that the orientation part of the DMP is not scaled, as we assume that required rotation for the detachment of the fruit from the branch does not depend on the size of it, but mostly on the task itself. 2.4 Combining kinesthetic teaching with visual demonstrations As explained in the Introduction, after a number of visual demonstrations, the robot attempts to execute a fruit picking procedure autonomously. However, this first attempt constitutes a preliminary phase before the actual autonomous execution, in which the inspection and correction take place, with the human being able to intervene by applying forces for correcting the parts of the motion that needs small modifications. To this aim, in CARPOS project we developed a control scheme that exploits the uncertainty of the visual observation, i.e. 𝑷 (𝑡)∈ 𝑆 , with 𝜇∈ℕ being the total (final) number of iterations of visual demonstrations. In particular, the proposed LbD method, which is presented in detail in Deliverable 5 – “Controller for unintentional and intentional pHRI during the coaching and autonomous operation”, combines the advantages of both visual and kinesthetic teaching, exploiting the intuitiveness and effortlessness of learning through observation (detailed above) while benefiting from the accuracy of kinesthetic teaching. This deliverable focuses only on the physical fusion part of the visual data with the ones gathered through the kinesthetic teaching. The CARPOS approach leverages visual data from the visual demonstrations, captured by the in-hand camera, to assist the human-teacher in kinesthetic task instruction. Let us highlight that the proposed method does not rely on knowledge of the camera’s location relative to an inertial frame. As also detailed in Deliverable 5, in this stage, to provide the system with compliant characteristics, we employ the notion of admittance control. Therefore, a target impedance model in the task-space of the robot is defined for the robot’s end-effector, as follows: 𝑴 𝜹 + 𝑫 𝜹 + 𝑲𝒆 = 𝑭 , (24) where 𝑴,𝑫,𝑲∈𝑆 are the target 6D task-space positive definite inertia, damping and stiffness matrices respectively, 𝜹∈ℝ represents the desired deviation from the reference velocity induced by the admittance model, while 𝑭=[𝒇 ⊺ 𝝉 ⊺]⊺∈ℝ is the generalized force measurement at the end-effector with 𝒇,𝝉∈ℝ being the force and torque applied by the user to the end-effector. Moreover,
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS 𝐞 ≜ 𝒑 − 𝒑 log ( 𝑸 ∗ 𝑸 ) ∈ ℝ , (25) is the error between the robot’s pose, denoted with 𝒑,𝒑∈ℝ are the actual position of endeffector of the robot and the one generated in real time by the DMP model respectively, while 𝑸,𝑸∈𝕊 are the actual orientation of the end-effector of the robot and the one generated by the DMP model respectively. Here, log(.):𝕊→ℝ is the quaternion logarithmic mapping, “∗” represents the quaternion product, and 𝑸∈𝕊 the quaternion inverse of 𝑸 (more details on quaternion algebra can be found in [22]). For rendering these specific compliant characteristics, the commanded joint velocities are computed by: 𝒒 = 𝑱 𝒑 𝝎 − 𝜹 ∈ ℝ , (26) where 𝒒∈ℝ are the joint variables and 𝑱(𝒒)∈ℝ× the Jacobian matrix of the robot, with 𝜉 being the number of joints and “†” symbolizing any pseudo inverse matrix; for the CARPOS project 𝜉=8, involving the degrees of freedom of the platform, as well as those of the robotic manipulator. Based on (24), the selection of the impedance model parameters plays a crucial role on the directions in which the user is guided during the interaction with the robot. Specifically, the reactive force experienced by the user along a specific direction is proportional to the values of the corresponding elements of matrices 𝑴,𝑫 and 𝑲. Our aim is to append low impedance along the directions with high uncertainty. As described above, the hand pose motion (encoded by the DMP) includes a measurement error, which is considered to be a random variable having an uncertainty characterized by 𝑷 . Given this uncertainty, we propose the utilization of the following anisotropic stiffness matrix: 𝑲 ≜ diag 𝑰 , 𝑰 𝑷 , (27) where 𝜆,𝜆∈ℝ denote the maximum eigenvalues of the position and orientation part of 𝑷 respectively, while 𝑘,𝑘∈ℝ represent the selected stiffness value on the major axis of 𝑲 for position and orientation, respectively. At the end of the kinesthetic correction stage, the recorded refined motion of the end-effector, i.e. 𝒑(𝑡),𝑸(𝒕),∀ 𝑡∈[0,𝑇], is used to retrain the DMP model.
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Figure 12 – Example of the kinesthetics-enhanced skill encoding By taking into account the non-homogeneous total uncertainty of the visually perceived data in 3D-space and guiding the user along the directions with the lowest uncertainty, the developed method essentially fuses the visually perceived data with the ones from kinesthetic teaching in the physical (force) layer. In particular, as the human is facilitated to do modifications towards the directions with the higher (current) uncertainty, those correction (green-shaded area in Figure 12) are incremental and complements the already known skill. Notice that if the skill is already learned with relatively low uncertainty already from the visual observations when this stage is enabled, then the human will not intentionally apply any forces and therefore the skill knowledge will remain unaltered. Extensive experiments of this method can be found in Deliverable 5, in which the related paper at IEEE RO-MAN is also attached. The experimental evaluation of the current deliverable, which can be found below, involves the application of this method in the case of fruit picking. 2.5 Encoding of forces applied by the fingertip For the encoding of the forces, a Force-Sensing Resistor (FSR) is attached to the main finger of the CARPOS gripper, as shown in Figure 13, enabling measurement of contact force during grasping and harvesting. The measurement from the sensor is captured by the microcontroller of the gripper, which in turn is sent as a ROS topic to the whole CARPOS system. let 𝑓(𝑡),∀𝑡∈[0,𝑇] be this measurement during the kinesthetic demonstration of the harvesting skill. For encoding 𝑓(𝑡), a Radial Base Function (RBF) network is utilized as a regressor, which encodes the time-series through the weighted sum of RBFs. Selecting the Gaussian RBFs, the utilized regressor has the following form: 𝑓 ( 𝑥 ) = 𝑤 , 𝜓 ( 𝑥 ) = 𝑾 ⊺ 𝜳 ( 𝑥 ) , (27) 𝒑 ( 𝑡 ) , 𝑸 ( 𝑡 ) Actual 𝒑 ( 𝑡 ) , 𝑸 ( 𝑡 ) Generated by the DMP Allowable deviation based on 𝑷 ⬚ (uncertainty of visual obs.) Side view Top view 𝑲
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Figure 13 – Force sensor in the thumb of the robotic gripper where 𝑓∈ℝ is the encoded (regressed) force, 𝑥∈ℝ is the phase variable of the DMP, 𝑤, ∈ ℝ,𝑖=1,…,𝐿 are the weights that essentially encode the characteristics of the force time-series, 𝜓(𝑥),𝑖=1,…,𝐿 are the Gaussian kernels provided in (3), 𝑾≜𝑤, 𝑤, … 𝑤, ⊺∈ℝ is the vector of the weights and 𝜳(𝑥)≜[𝜓(𝑥) … 𝜓(𝑥)]⊺∈ℝ is the Gaussian Radial Base. i.e. the vector with the values of all the base functions, with 𝐿∈ℕ being the number of RBFs utilized, which is pre-defined. For finding the weights that optimally encode the demonstrated force pattern, one has to solve the following optimization problem: 𝑾 ∗ = arg min 𝑾 𝑾 ⊺ 𝜳 ( 𝑥 ( 𝑡 ) ) − 𝑓 ( 𝑡 ) 𝑑𝑡 , (28) which, given the discrete form of the time-series of 𝑓(𝑡), is solved using the Least-Squares formula, similarly to the one employed in (6) above. An example of this RBF-based encoding is presented in Fig. 14 for an arbitrary selected signal. Notice the difference in accuracy achieved with 10 and 40 kernels respectively. However, let us highlight that a more accurate approximation, which is achieved by using more kernels, possibly will also encode the high-frequency noise of the demonstrated time-series, therefore there is trade-off between high accuracy and noise rejection, which can be achieved by selecting a moderate number of kernels. Notice that the encoding of this contact force is done based on the phase variable of the DMP, in order for this value to be synchronized with the motion of the end-effector. This feature augments the original DMP in 𝑆𝐸(3) by one more dimension. In particular, in order to guide the motion of the robot towards achieving these learned forces, we propose the utilization of an additional term in the first equation of the position DMP, i.e. eq. (8), which proportionally moves the robot towards the direction of the contact force based on the contact force error. The complete “force-aware” DMP model, involving (8) and (10), is the following: Force sensor 𝒏
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS 𝜏 𝒑 = 𝑎 ( 𝑏 ( 𝒑 − 𝒑 ) − 𝜏 𝒑 ) + 𝑠 𝒇 ( 𝑥 ; 𝒘 ) + 𝑘 𝒏 ( 𝑓 𝑥 ; 𝑾 − 𝑓 ) , 𝜏𝒛=−𝑎𝑏𝒆+𝒛+𝑠 𝒇(𝑥 ;𝑾), 𝜏𝒆=𝒛, 𝜏 𝑥 = 𝑔 ( 𝑥 ) , (29) where 𝒏∈ℝ is the unit vector of the contact force measurement direction (shown in Fig. 13) and 𝑘∈ℝ a pre-defined tunable gain for this force-control action. Figure 14 – RBF-based encoding with 10 and 40 kernels. a) Encoding with 𝐿 = 10 RBFs b) Encoding with 𝐿 = 40 RBFs
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS EXPERIMENTAL VALIDATION Motion encoding Multiple full training procedures are executed and recorded, utilizing the complete CARPOS robotic system, i.e. the UR10e robotic manipulator attached on a Clearpath Husky platform (depicted in Figure 15), with a RealSense camera attached to the robot’s end-effector as depicted in Fig. 13. For enabling extensive testing of our methodologies horizontally within the CARPOS project, we designed and fabricated (using a 3D printer) an “artificial stem”, emulating the detachment of the fruit from the branch, employing a combination of magnets, as depicted in Figure 16. This artificial stem emulates the real stem abruption which firstly requires a rotation, followed by a translation, as a single translation without any rotation would require significantly larger forces. The experimental setup involved a real tomato plant with the mockup tomato fruit, as depicted in Fig. 15. For demonstrating the encoding functionality and performance of the CARPOS framework, described above, the results of a representative full training procedure are provided. More specifically, in Figure 17, the paths after the data fusion in each of the following steps are depicted, with respect to the frame of the fruit (red area in the figure): - Step1: First visual observation (Fig. 15a). o Sampling rate: 30 Hz (due to the camera fps). o The camera was observing the scene from its initial pose which was towards facing the x-z plane of the initial hand configuration. - Step 2: Data fusion based on the LAIF method, using the data from the second corrected visual observation, from another point of view (Fig. 15b). o Sampling rate: 30 Hz (also due to the camera fps) o The camera was firstly automatically raised above the scene, i.e. approximately facing the y-z plane of the initial hand’s pose, due to the action of the active perception of the LAIF method. - Step 3: Fusion in the physical layer based on the kinesthetic inspection and correction phase (Fig. 15c). o Sampling rate: 500 Hz (due to the robot’s high proprioceptive sampling rate) o The human was guided towards corrections based on the uncertainty from the visual observations, utilizing the proposed method.
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Figure 15 – The three phases of the training procedure for encoding the skill Figure 16 – The artificial stem emulating the detachment of the real fruit a) Initial visual demonstration b) Second visual demonstration c) Kinesthetic inspection and correction z x y a) The mockup tomato utilized b) Initial rotation of the art. stem c) Translation of the top part
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Custom GPT framework3. A key aspect of the configuration was the emphasis on rule-based classification, where the identification and segmentation of motions were grounded in the predefined kinematic characteristics listed in Table II, focusing on the dominant directional patterns observed in the translational and angular velocity components. Table II - Kinematic characteristics of primitive actions. Primitive Characteristic Behavior Pull Translation along 𝑥 -axis, without significant rotation Slide Translation on the 𝑦 − 𝑧 plane, without significant rotation Swing Translation along the 𝑦 -axis and rotation around the 𝑧 -axis Tilt Minimal translation 𝑧 − axis and rotation around 𝑦 -axis Twist No translation, but rotation around 𝑥 -axis Three fine-tuning (learning / explanation) approaches are considered for our case study. The first approach (Approach A) involves only linguistic explanation of the primitive actions. In particular, in this approach the GPT model is only provided with descriptions of the characteristics of the primitive actions with respect to the expected major axes of motion for each one of them; the characteristics defined in Table II. In the second approach (Approach B), a small dataset of representative example motions for each primitive action is provided to the GPT model without however providing any linguistic explanation regarding the characteristics of motion; namely five examples for each primitive action. Finally, Approach C combines both means of Approach A and B, i.e. both linguistic explanation and example datasets are provided to the GPT model. In all three approaches, the model was primed to recognize the motion primitives by examining the temporal evolution of the velocity signals and recognizing changes consistent with the prevalent axes defined for each motion type. Seg mentation indices were derived by recognizing the earliest significant change in either translational or angular velocity, depending on the defining characteristics of the motion. This ensured that each primitive began at the point where its characteristic movement pattern first emerged. The resulting sequence of motion primitives followed a strict chronological order, determined by their respective start indices. In order to allow consistent visualization of velocity profiles and to improve readability of motion transitions, a Python helper function was also provided in the GPT configurations to compute the supremum (maximum absolute value) of the translational and angular velocity components; the script is publicly available4. This method was then used to normalize the values of all plots so that an equivalent visual scale is shown regardless of motion magnitude. These resulting plots helped to visually inspect the segmentation boundary validity, making it easier to assess the accuracy and coherence of the identified primitives. In line with the recognition and segmentation rules described above, the custom GPT produced an easy readable, structured list of motion primitives with a start index and end index marking a segment border within the time series. An example output format for a complex sequence 3 https://help.openai.com/en/articles/8554397-creating-a-gpt 4https://github.com/CSRLHMU/VerbalDMPcorrections/blob/526440db1fe64ebccb89fe27721f18fc069b9526/motion_classificatio n_segmentation_time_series.ipynb
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS consisting of three dis tinct motion primitives, would appear as follows: twist (Index 0–62), tilt (Index 63–112), pull (Index 113–170). This clear and consistent representation supported both interpretability and smooth integration with tasks such as evaluation and visualization. Data acquisition and Pre-processing In order to create a smooth and cooperative human-robot interaction and capture the realistic motion behavior, a task-space admittance control scheme was employed, providing the robot with homogeneous compliance characteristics. For comparison purposes, a push-button was attached to the handles of the end-effector of the CARPOS robot, serving as a tool for precisely recording the exact moment of motion transitions while capturing the complex motion sequences which was used only as ground-truth data for the evaluation. When pressed, it generates a timestamp that effectively marks key transitions. To further emulate a real fruit-picking setup, a mock-up fruit was placed in the robot’s gripper, with the fruit being, at the start of each motion, anchored to a vertical stand that mimicked the presence of a fruit stem. The experimental setup used is depicted in Fig. 24. Figure 24 – Image from the demonstration of the pull primitive All experimental data employed in this work were organized into three datasets: training (provided to the model in Approaches B and C as examples), validation (used for assessing the performance of Approach C with feedback) and testing, each of which consisted of time-series motion recordings in the same format described above. The training dataset was made up of five demonstrations of each motion primitive separately, i.e. 25 time-series in total. As for the testing dataset, its purpose was to assess the model’s performance in all the considered approaches. It was made up of complex motion sequences made up of a sequence of the considered primitive actions. The experimental data captured the full 𝑆𝐸(3) pose in formation, i.e. 𝒑(𝑡),𝑸(𝑡),∀ 𝑡∈[0,𝑇], provided by the forward kinematics of the robot at 500Hz. However, one of the important aspects regarding the communication with most of the current GPT models, is the data size provided to the model, with most of the manufacturers imposing data limits or introducing additional costs. Therefore, to resolve the problem of the high sampling density of the initial time-series (i.e. 500
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Hz), the data were interpolated using the Nadaraya Watson kernel regression method [18], [19], significantly reducing, in this way, the size of the dataset, using Gaussian kernels [20]. This nonparametric approach provides an estimate, i.e. 𝑦(𝑡)∈ℝ, of an arbitrary timeseries 𝑦(𝑡), at a given time instance 𝑡, by computing a kernel-weighted average of observed values 𝑦 corresponding to time instances 𝑡 (i.e. 𝑦≜𝑦(𝑡)), as defined by: 𝑦 ( 𝑡 ) = ∑ 𝐾 ( 𝑡 − 𝑡 ) 𝑦 ∑ 𝐾 ( 𝑡 − 𝑡 ) (30) where 𝐾≜exp− is the Gaussian kernel, which assigns weights to each observation at 𝑡based on its proximity to t, with 𝜎∈ℝ reflecting its variance. For the interpolation, 𝑣=20 kernels per second were considered. Following the interpolation, to thoroughly analyze the dynamic behavior of the system, both translational and angular velocity vectors, i.e. 𝒑(𝑡),𝝎(𝑡) respectively, have been obtained from the non-causal (centered) numerical differentiation of the position and orientation data respectively. The latter ensures numerical consistency and preserves accuracy across all regions, yielding a discrete approximation of the continuous velocity profile from the estimated derivatives [21]. To compute the angular velocity vector from the unit quaternion derivative the following formula is used: 𝝎 ( 𝑡 ) = 𝟐 𝑱 ⊺ ( 𝑸 ) 𝑸 ( 𝑡 ) , (31) where 𝑱(𝑸)∈ℝ× is the matrix mapping the angular velocities to unit quaternion rates, defined as follows: 𝑱 ( 𝑸 ) ≜ − 𝝐 ⊺ 𝜂 𝑰 + 𝑺 ( 𝝐 ) , (32) with 𝑺(.):ℝ→ℝ× denoting the skew-symmetric matrix of an ℝ vector and 𝜂∈ℝ, 𝝐∈ℝ representing the scalar and vector part of the unit quaternion respectively, i.e. 𝑸≜[𝜂 𝝐⊺]⊺. More details about the algebra of unit quaternions and its use for describing the orientation of a frame can be found in [22]. Approach A: Explanation through language In this approach, the custom GPT model was tasked with identifying and segmenting a series of complex motion sequences using only the kinematic characteristics defined in advance in the model configuration. Importantly, in this approach, the model was not provided with any examples, training data, or data retrieved from motion primitives. Instead, is relied entirely on language-based reasoning, using descriptive rules about each motion behavior (see Table II). The prompt was aimed at looking at velocity profiles for the purpose of classifying motion boundaries, order of occurrence, and segmentation indices.
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS Approach B: Providing example measurements Unlike the language-based reasoning of Approach A, this method, introduces an intermediate stage of learning where the GPT is provided with representative motion examples, corresponding to the individual motion primitives: pull, swing, slide, tilt, twist. In this approach, the specified kinematic characteristics of each motion are not provided to the GPT. Instead, the model has to learn and understand these characteristics implicitly. Approach C: Combining explanations with examples Approach C integrates the language-based reasoning of Approach A with the example-based reasoning of the Approach B, aiming to leverage the strengths that emerge from the combination of these two approaches. In more details, the model here is provided with both the specific kinematic characteristics and the representative motion examples. Conclusions and the importance of providing feedback to the system Building upon these findings, the need to provide feedback to the system is evident. The performance variations observed among the three approaches mentioned before indicate the potential for improvement through a targeted feedback process. Towards this direction, five out of the 20 complex motion sequences were used as a “validation” dataset for supplying corrective feedback to the system. The results are detailed in there LBR report provided in the Appendix of the deliverable. In the context of this experiment, the feedback approach provides some performance improvement by allowing the system to adjust its internal representations and decision processes based on indicated past mistakes and inconsistencies. REFERENCES [1] A. Billard, S. Calinon, R. Dillmann, and S. Schaal, Robot Programming by Demonstration. Berlin, Heidelberg: Springer Berlin Heidelberg, 2008, pp. 1371–1394. [2] R. Dillmann, “Teaching and learning of robot tasks via observation of human performance,” Robot. Auton. Syst., vol. 47, no. 2, pp. 109–116,2004. [3] A. Sidiropoulos and Z. Doulgeri, “From RGB images to Dynamic Movement Primitives for planar tasks,” in IEEE/RAS Int. Conf. Humanoid Robots (Humanoids), 2023, pp. 1–8.
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS [4] R. Rahmatizadeh, P. Abolghasemi, L. Boloni, and S. Levine, “Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration,” in IEEE Int. Conf. Robot. Autom. (ICRA), 2018, pp. 3758–3765. [5] Y. Liu, A. Gupta, P. Abbeel, and S. Levine, “Imitation from observation: Learning to imitate behaviors from raw video via context translation,” in IEEE Int. Conf. Robot. Autom. (ICRA), 2018, pp. 1118–1125. [6] D. Papageorgiou and Z. Doulgeri, “A control scheme for haptic inspection and partial modification of kinematic behaviors,” in IEEE/RSJ Int. Conf. Intell. Rob. Syst. (IROS), 2020, pp. 9752–9758. [7] D. Papageorgiou, F. Dimeas, T. Kastritsi, and Z. Doulgeri, “Kinesthetic guidance utilizing dmp synchronization and assistive virtual fixtures for progressive automation,” Robotica, vol. 38, no. 10, p. 1824–1841, 2020. [8] D. Lee and C. Ott, “Incremental motion primitive learning by physical coaching using impedance control,” in IEEE/RSJ Int. Conf. Intell. Rob. Syst. (IROS), 2010, pp. 4133–4140. [9] Sylvain Calinon et al. “Handling of multiple constraints and motion alternatives in a robot programming by demonstration framework”. In: 2009 9th IEEE-RAS International Conference on Humanoid Robots. 2009, pp. 582–588. [10] E. Gribovskaya, S.M. Khansari-Zadeh, and A. Billard. “Learning Non-linear Multivariate Dynamics of Motion in Robotic Manipulators”. In: The International Journal of Robotics Research 30.1 (2011), pp. 80–117. url: https://doi.org/10.1177/0278364910376251 [11] A.J. Ijspeert, J. Nakanishi, and S. Schaal. “Movement imitation with nonlinear dynamical systems in humanoid robots”. In: Proceedings 2002 IEEE International Conference on Robotics and Automation (Cat. No.02CH37292). Vol. 2. 2002, 1398–1403 vol.2. [12] Auke Jan Ijspeert et al. “Dynamical Movement Primitives: Learning Attractor Models for Motor Behaviors”. In: Neural Comput. 25.2 (Feb. 2013), pp. 328–373. [13] A. Sidiropoulos, “Enhancing the generalization and adaptation capabilities of Dynamic Movement Primitives during robot learning from human demonstrations and execution in dynamic environments”, PhD Thesis, 2023, DOI: 10.26262/heal.auth.ir.354908, url: https://ikee.lib.auth.gr/record/354908/?ln=en [14] A. Ude, B. Nemec, T. Petric, and J. Morimoto. Orientation in cartesian space dynamic movement primitives. In 2014 IEEE International Conference on Robotics and Automation (ICRA),pages 2997–3004, May 2014. doi:10.1109/ICRA.2014.6907291. [15] Koutras, L., Doulgeri, Z.. (2020). A correct formulation for the Orientation Dynamic Movement Primitives for robot control in the Cartesian space. Proceedings of the Conference on Robot Learning, in Proceedings of Machine Learning Research. 100:293-302 Available from https://proceedings.mlr.press/v100/koutras20a.html.
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS [16] Bruno Siciliano, Oussama Khatib. Springer Handbook of Robotics. DOI: https://doi.org/10.1007/978-3-540-30301-5 [17] D. Papageorgiou, A. Sidiropoulos and Z. Doulgeri, "Sinc-Based Dynamic Movement Primitives for Encoding Point-to-point Kinematic Behaviors," 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 2018, pp. 8339-8345, doi: 10.1109/IROS.2018.8594479. [18] E. A. Nadaraya, “On estimating regression”, Theory of Probability & Its Applications, vol.9, no.1, pp.141–142, 1964. [19] G. S. Watson, “Smooth regression analysis,” Sankhya: The Indian Journal of Statistics, Series A (1961–2002), vol. 26, no. 4, pp. 359 372,1964. [20] C. E. Rasmussen and C. K. I. Williams, Gaussian Processes for Machine Learning. Cambridge, Massachusetts: The MIT Press, 2006. [21] R. L. Burden and J. D. Faires, Numerical Analysis, 9th ed. Boston: Brooks / Cole, Cengage Learning, 2010. [22] L. Koutras and Z. Doulgeri, “Exponential stability of an attitude trajectory tracking controller utilizing unit quaternions,” 06 2021, pp. 1126–1131.
CARPOS Coachable Robot for Fruit Picking Operations D7 - Multimodal skill encoding and generalization in fruit geometry variations Version: 1.0 CARPOS APPENDIX A LATE BREAKING RESULTS POSTER AND REPORT AT IEEE RO-MAN 2025
On the capabilities of LLMs for classifying and segmenting time series of fruit picking motions into primitive actions Eleni Konstantinidou1, Nikolaos Kounalakis2, Nikolaos Efstathopoulos1and Dimitrios Papageorgiou1,3 This paper is a Late Breaking Results report and it will be presented through a poster at the 34th IEEE International Conference on Robot and Human Interactive Communication (ROMAN), 2025 at Eindhoven, the Netherlands. Abstract— Despite their recent introduction to human society, Large Language Models (LLMs) have significantly affected the way we tackle mental challenges in our everyday lives. From optimizing our linguistic communication to assisting us in making important decisions, LLMs, such as ChatGPT, are notably reducing our cognitive load by gradually taking on an increasing share of our mental activities. In the context of Learning by Demonstration (LbD), classifying and segmenting complex motions into primitive actions, such as pushing, pulling, twisting etc, is considered to be a key-step towards encoding a task. In this work, we investigate the capabilities of LLMs to undertake this task, considering a finite set of predefined primitive actions found in fruit picking operations. By utilizing LLMs instead of simple supervised learning or analytic methods, we aim at making the method easily applicable and deployable in a real-life scenario. Three different fine-tuning approaches are investigated, compared on datasets captured kinesthetically, using a UR10e robot, during a fruit-picking scenario. I. INTRODUCTION Artificial Intelligence (AI) has become a powerful force that is revolutionizing various fields by developing systems that mimic human intelligence, encompassing this way areas like reasoning, learning and problem-solving [1]. This transformation is primarly due to Large Language Models (LLMs), deep learning models capable of processing and generating natural language data and other content [2]. The evolution of LLMs was driven by the growth of the internet and increasing computing power, which fuelled the vision of Natural Language Processing (NLP) [3] through deep learning methods such as Recurrent Neural Networks (RNNs) [4], and the Transformer model [5]. This nourished models like BERT [6] and Generative Pre-Trained Transformer (GPT) [7], revolutionizing fields like text generation and language translation. From a ”robotics” perspective, segmenting a complex motion into primitive actions is a crucial step in learning human kinematic behaviors through demonstrations. While many works employ Deep Neural Networks (DNN) [8][9] for segmentation, and others propose analytic solutions [10][11] This research is funded by CARPOS project, carried out within the framework of the National Recovery and Resilience Plan Greece 2.0, funded by the European Union - NextGenerationEU (Implementation Body: HFRI). Project Number: 16523. 1E. Konstantinidou, N. Efstathopoulos and D. Papageorgiou are with the Dept. of Electrical and Computer Engineering, Hellenic Mediterranean University, 714 10, Heraklion, Crete, Greece. [email protected] 2N. Kounalakis is with the Dept. of Mechanical Engineering, Hellenic Mediterranean University, 714 10, Heraklion, Crete, Greece. 3D. Papageorgiou is also with the Institute of Computer Science, Foundation for Research and Technology - Hellas, Greece. or even AI methods like Temporal Convolutional Networks (TCN) [12] [13] for explicitly segmenting motion timeseries, analytic methods require solid technical skills for action detection and classification, making them hard to deploy in a real-world scenario. This work investigates LLM capabilities for classifying primitive actions and segmenting complex fruit-picking motions. In particular, we explore and compare three LLM fine-tuning approaches: a) using only linguistic explanation of the primitive action, b) using a small dataset of representative example motions (e.g., five samples per action) and c) combining both a) and b) approaches. The utilized datasets for both the fine-tuning and experimental evaluation/comparison of the approaches, consists of kinesthetically captured data, i.e. motion time-series captured by having the human demonstrating the motion directly by moving the robot’s endeffector through physical contact. II. MOTIVATION AND PROBLEM DESCRIPTION The advantages of picking a fruit without utilizing cutting tools span from ensuring the safety of the plant and fruit, to being effective, as in many cases there are no cutting affordances for the insertion of a cutting tool (e.g. in the case of peaches). The complex motion required for detaching a fruit from the branch depends on the type of the plant. However, in most of the cases, this complex motion is composed by a sequence of simpler primitive actions. As we only account for the detachment of the fruit, these primitive actions are considered to be the following: Pull,Slide,Swing, Tilt,Twist, as presented in Fig. 1, with either one of the first two (Pull or Slide) always appearing at the end of the complex detachment motion, as they involve a displacement of the fruit from the branch. Fig. 1: The considered primitive actions To teach a fruit picking skill to a novice farmer, humans firstly utilize verbal communication, i.e. linguistic explanation. An example sentence provided by an experienced farmer to the novice one, to this aim, could be the following: ”To detach the fruit, you firstly have to twist it and then pull it from the branch.”. In most of the cases, the experienced farmer also provides an example of the fruit arXiv:2507.07745v1 [cs.RO] 10 Jul 2025
picking procedure, by demonstrating it in front of the learner. Inspired by the intuition characterizing the human-to-human skill transfer, in this work we aim at investigating the use of Large Language Models (LLMs) for segmenting a complex sequence of primitive actions, in the context of the fruit picking task. Our aim is to exploit these capabilities as a part of Learning by Demonstration (LbD) for a robotic system which will be capable to learn from humans through intuitive means. Let us assume the availability of a demonstrated fruit picking motion. Therefore, let {F}be the frame at the initial fruit’s pose and x(t)≜[p⊺(t)Q⊺(t)]⊺∈R3×S3be the pose of the robot’s end-effector with respect to {F}, where p(t)∈R3and Q(t)∈S3are the end-effector’s position and orientation in unit quaternion form respectively, with S3denoting the unit sphere in R4. Moreover, let v(t)≜ [˙ p⊺(t)ω⊺(t)]⊺∈R6be the generalized velocity of the endeffector with respect to {F}, with ˙ p(t),ω(t)∈R3being its linear and angular velocity respectively. Given that x(t)is provided by the forward kinematics of the robot, one can calculate a numerical approximate of v(t), based on the numerical differentiation of the pose, as described in the following section. Given a time series of the generalized velocity, v(t), t ∈ [0, T ], of a complex fruit-picking motion, with respect to the initial pose of the fruit, with T∈R>0being the duration of motion, our aim is to test the effectiveness of an LLM to simultaneously classify and segment the motion into the aforementioned five predefined primitive actions, shown in Fig. 1. Notice that the velocity of the end-effector is considered, instead of its pose, in order to facilitate the segmentation by rejecting bias. In other words, by selecting the velocity instead of the pose, each action maintains its characteristics even if it follows after a displacement that could possibly have occurred due to the sequential nature of the complex motion considered. III. CONSIDERED APPROACHES AND COMPARISON For classifying and segmenting the time series into primitive actions, a custom Generative Pre-trained Transformer (GPT) was created, based on OpenAI’s GPT-4-turbo model via the ChatGPT Custom GPT framework1. A key aspect of the configuration was the emphasis on rule-based classification, where the identification and segmentation of motions were grounded in the predefined kinematic characteristics listed in Table I, focusing on the dominant directional patterns observed in the translational and angular velocity components. Three fine-tuning (learning / explanation) approaches are considered for our case study. The first approach (Approach A) involves only linguistic explanation of the primitive actions. In particular, in this approach the GPT model is 1https://help.openai.com/en/articles/ 8554397-creating-a-gpt only provided with descriptions of the characteristics of the primitive actions with respect to the expected major axes of motion for each one of them; the characteristics defined in Table I. In the second approach (Approach B), a small dataset of representative example motions for each primitive action is provided to the GPT model without however providing any linguistic explanation regarding the characteristics of motion; namely five examples for each primitive action. Finally, Approach C combines both means of Approach A and B, i.e. both linguistic explanation and example datasets are provided to the GPT model. TABLE I: Kinematic characteristics of primitive actions. Primitive Characteristic Behavior Pull Translation along x-axis, without significant rotation Slide Translation on the y−zplane, without significant rotation Swing Translation along the y-axis and rotation around the z-axis Tilt Minimal translation z-axis and rotation around y-axis Twist No translation, but rotation around x-axis In all three approaches, the model was primed to recognize the motion primitives by examining the temporal evolution of the velocity signals and recognizing changes consistent with the prevalent axes defined for each motion type. Segmentation indices were derived by recognizing the earliest significant change in either translational or angular velocity, depending on the defining characteristics of the motion. This ensured that each primitive began at the point where its characteristic movement pattern first emerged. The resulting sequence of motion primitives followed a strict chronological order, determined by their respective start indices. In order to allow consistent visualization of velocity profiles and to improve readability of motion transitions, a Python helper function was also provided in the GPT configurations to compute the supremum (maximum absolute value) of the translational and angular velocity components; the script is publicly available2. This method was then used to normalize the values of all plots so that an equivalent visual scale is shown regardless of motion magnitude. These resulting plots helped to visually inspect the segmentation boundary validity, making it easier to assess the accuracy and coherence of the identified primitives. In line with the recognition and segmentation rules described above, the custom GPT produced an easy readable, structured list of motion primitives with a start index and end index marking a segment border within the time series. An example output format for a complex sequence consisting of three distinct motion primitives, would appear as follows: twist (Index 0–62),tilt (Index 63–112),pull (Index 113–170). This clear and consistent representation supported both interpretability and smooth integration with tasks such as evaluation and visualization. 2GitHub link to script