scieee AI-readable full text Open interactive document viewer

Chewing Detection from Commercial Smart-glasses

Papapanagiotou, Vasileios; Liapi, Anastasia; Delopoulos, Anastasios

Abstract

Abstract - Automatic dietary monitoring has progressed significantly during the last years, offering a variety of solutions, both in terms of sensors and algorithms as well as in terms of what aspect or parameters of eating behavior are measured and monitored. Automatic detection of eating based on chewing sounds has been studied extensively, however, it requires a microphone to be mounted on the subject's head for capturing the relevant sounds. In this work, we evaluate the feasibility of using an off-the-shelf commercial device, the Razer Anzu smart-glasses, for automatic chewing detection. The smart-glasses are equipped with stereo speakers and microphones that communicate with smart-phones via Bluetooth. The microphone placement is not optimal for capturing chewing sounds, however, we find that it does not significantly affect the detection effectiveness. We apply an algorithm from literature with some adjustments on a challenging dataset that we have collected in house. Leave-one-subject-out experiments yield promising results, with an F1-score of 0.96 for the best case of duration-based evaluation of eating time.

Full text

Chewing Detection from Commercial Smart-glasses Vasileios Papapanagiotou [email protected] Multimedia Understanding Group Electrical and Computer Engineering Dpt. Aristotle University of Thessaloniki Thessaloniki, Greece Anastasia Liapi [email protected] Electrical and Computer Engineering Dpt. Aristotle University of Thessaloniki Thessaloniki, Greece Anastasios Delopoulos [email protected] Multimedia Understanding Group Electrical and Computer Engineering Dpt. Aristotle University of Thessaloniki Thessaloniki, Greece ABSTRACT Automatic dietary monitoring has progressed significantly during the last years, offering a variety of solutions, both in terms of sensors and algorithms as well as in terms of what aspect or parameters of eating behavior are measured and monitored. Automatic detection of eating based on chewing sounds has been studied extensively, however, it requires a microphone to be mounted on the subject’s head for capturing the relevant sounds. In this work, we evaluate the feasibility of using an off-the-shelf commercial device, the Razer Anzu smart-glasses, for automatic chewing detection. The smart-glasses are equipped with stereo speakers and microphones that communicate with smart-phones via Bluetooth. The microphone placement is not optimal for capturing chewing sounds, however, we find that it does not significantly affect the detection effectiveness. We apply an algorithm from literature with some adjustments on a challenging dataset that we have collected in house. Leave-one-subject-out experiments yield promising results, with an F1-score of 0 . 96 for the best case of duration-based evaluation of eating time. CCS CONCEPTS •Human-centered computing →Ubiquitous and mobile computing systems and tools. KEYWORDS automatic dietary management, wearables, chewing, smart-glasses ACM Reference Format: Vasileios Papapanagiotou, Anastasia Liapi, and Anastasios Delopoulos. 2022. Chewing Detection from Commercial Smart-glasses. In Proceedings of the 7th International Workshop on Multimedia Assisted Dietary Management (MADiMa ’22), October 10, 2022, Lisboa, Portugal. ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3552484.3555746 1 INTRODUCTION Detection of chewing sounds is one of the first approaches that have been studied in the field of automatic dietary monitoring [3]. The idea is to capture (by audio) the distinct sound that occurs Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. MADiMa ’22, October 10, 2022, Lisboa, Portugal. ©2022 Association for Computing Machinery. ACM ISBN 978-1-4503-9502-1/22/10...$15.00 https://doi.org/10.1145/3552484.3555746 when food is crashed between the teeth. Different placements of the microphone have been considered, but the one that seems naturally advantageous is the outer in-ear canal, as it captures the chewing sounds naturally amplified by the skull, while external sounds are attenuated. Typically, microphones have been mounted on custom housing to support this placement, which creates the need for custom-made hardware. In recent years, various approaches have also emerged, focusing on different sensor solutions such as detection of swallowing sounds [4, 25, 1]. The naturally optimal placement of the microphone for this task is close to the neck, where swallowing sounds originate from. Other approaches attempt to leverage the power of modern wearables, such as smart-watches, by using the inertial sensors (i.e. accelerometer and optionally gyroscope) that are commonly found on such devices, in order to detect and identify eating gestures [15, 14, 12], Other approaches rely on the availability of cameras to identify food type [6, 5] from single plate photographs, or to directly segment food images into food components [7, 9]. It is also possible to estimate food volume using a depth camera [16] and the caloric content based on a photograph using a reference object [11]. In [10], authors reconstruct a 3D model of the food that is placed on a plate in order to estimate food volume; however, their approach requires two different views (photographs) of the plate. While such approaches can clearly extract very detailed information, they require the active participation of the user by taking a photograph and triggering the analysis, or in some cases specialized cameras (such as depth camera) or multiple views. However, combining with monocular depth-estimation methods [18] can help reduce user and hardware input requirements and simplify the estimation process. Besides smart-watches, another wearable that has received attention for dietary monitoring is the smart-glasses. Smart-glasses are mostly in experimental state currently, however, many alternative sensors have been mounted and evaluated in the literature. In [29], authors use 3D-printed glasses that incorporate an electromyography (EMG) sensor and evaluate its potential on detecting chewing and identifying food types, on a dataset of eight participants and five food types. The best reported chewing detection effectiveness is 0 . 8for both recall and precision, and classification accuracy is reported in the range of 0 . 64 to 0 . 84. In [27], they combine it with a vibration sensor. In [28], authors use 3D printed smart-glasses with a bilateral EMG sensor to detect chewing. They evaluate both in lab and free-living conditions. Results for free-living conditions in chewing detection achieve 0.79 recall and 0.77 precision. In [17], authors opt for a 3D accelerometer sensor (from Shimmer), mounted with plastic straps on regular glasses, as a prototype arXiv:2208.05735v1 [eess.AS] 11 Aug 2022 MADiMa ’22, October 10, 2022, Lisboa, Portugal. Papapanagiotou et al. device. They propose an algorithm based on feature extraction and classification, using either support vector machines (SVMs) or random forests, targeting three classes: chewing, non-chewing, and walking (as a challenging counterpart for chewing). They evaluate on a dataset from five volunteers, and achieve a classification accuracy of 0 . 74, however, the prior probability of each class in the dataset is not given. A pair of glasses with mounted EMG and additional electronics to support communication over Bluetooth with a smart-phone is used in [13]. The proposed algorithm is made of two parts: one that runs on the electronics mounted on the glasses and detects eating periods (vs. non-eating periods) and one that runs on a more powerful processing unit (after the data transfer via Bluetooth) that detects individual chews. Chewing detection is evaluated on dataset of four individuals that eat, drink, and talk, while sited for 40 minutes each, and achieves 0.96 accuracy. The authors of [8] propose a different approach by using 3D printed glasses that incorporate load cells. The principle of operation is that during chewing, the temples of the glasses (the part that is closer to the ear) is slightly pressed outward, and this increases the load on the cell that is placed at the hinge. The algorithm includes extraction of both time-domain and frequency-domain features and training of an SVM classifier with an radial-basis function (RBF) kernel. Authors evaluate on a dataset of 10 subjects on 6classes that include left/right side chewing, left/right side winking, talking, and head moving, and report an average F1-score of 0.94. In this current work, we try to combine the advantages of multiple approaches into a single solution. We employ audio-based chewing detection, as it can provide very accurate results, detailed (per chew) detection, can be used to identify food type [19], food texture [21], and even bite weight [2, 24], and also does not require active user input. We employ a commercial, off-the-shelf device to remove the requirement for specialized/custom-made hardware. We opt for a pair of Bluetooth-enabled smart-glasses by Razer. Finally, we employ a chewing-detection algorithm from literature that has been previously employed on a sensor that combines audio, photoplethysmography, and acceleration signals [22]; the algorithm is resilient to the challenges that this device introduces, such as the non-optimal placement of microphones. To evaluate our work, we conducted a data-collection trial. The rest of this work is organized as follows. Section 2 describes the sensor, data-collection process, and the algorithm for chewing detection. Section 3 presents the evaluation framework and results and discusses them. Finally, section 4 concludes the paper. 2 MATERIALS AND METHODS 2.1 Hardware and data-collection The hardware we use in this work for recording audio is the commercially available off-the-shelf smart-glasses by Razer, the Razer Anzu (Figure 1). They are regular glasses (which can also be used as sunglasses) that include stereo speakers and microphones, and communicate with smart-phones and laptops wirelessly via Bluetooth. The microphones are placed on the glasses temples, facing inside, and very close to the glasses end-pieces. They are essentially facing the subject’s eyes from the outer sides. Figure 1 marks the microphone placement with the red symbols: the microphone on Figure 1: The commercial smart-glasses that are used for collecting audio, Razer Anzu. (Photo from https://www.razer.co m/mobile-wearables/razer-anzu-smart-glasses). There are two microphones marked by the red circle (left side) and arrow (right side, microphone not visible due to perspective). the left side is at the two, very small, dark dots inside the red circle; the red arrow points to the right microphone which is not visible due to perspective. The audio recorded from the smart-glasses is stereo (2 channels). We sample at 16 kHz and we use the Advanced Audio Coding (AAC) standard for compression. The compression yields files with size less than 6MBs per hour which permits uploading the audio files to a server for processing and detection. Based on existing literature [3], chewing sounds travel well through the skull and are best captured with microphones placed inside the ear’s outer canal, facing inwards (this placement also naturally attenuates external sounds). However, continuous use of in-ear microphones can create discomfort to some users [26]. On the other hand, the Razer Anzu smart-glasses are worn as common glasses and are thus more comfortable. They are also available to buy commercially, which eliminates the need for specialized hardware. The only drawback to opting for the smart-glasses is that the placement of microphones is not optimal for chewing detection. Indeed, microphone placement in the Razer Anzu is oriented towards capturing speech, as their target use-case is to be used as a headset. This creates a bigger challenge for the chewing detection algorithm (compared to using in-ear microphones). 2.2 Dataset To record audio from the smart-glasses we have developed an application for Android smart-phones that can connect to the smartglasses via Bluetooth and record audio in real-time. The audio files are stored on the phone and we manually gather them to a computer for analysis and processing. However, it is technically feasible to upload them automatically to a server, given their relatively small size. The application includes an interface (Figure 2) for the subject to annotate the start and stop timestamps of eating sessions (meals). This helps us to manually annotate individual chews, which is required for training. Chewing Detection from Commercial Smart-glasses MADiMa ’22, October 10, 2022, Lisboa, Portugal. Figure 2: Main interface of the Android application that was developed for recording audio from the smart-glasses. It also enables the user to provide start and stop timestamps for eating sessions (meals). To analyze and evaluate our method we collected a dataset of audio data. A total of 5subjects (1male and 4female, age range 23 to 28) participated in the collection process. Subjects had no reported or diagnosed medical issues relevant to eating and digestion. Each recording lasted approximately an hour, and we collected 6such recordings (one subject contributed twice). Subjects were instructed to perform the following activities: eating, talking, walking, and resting. Subjects were free to perform these activities in any order, and in any way they wanted. The only instructions that were given were to perform each activity for at least 10 minutes, and have at least 2eating sessions (so at least 20 minutes of eating). The recorded audio signals were inspected in Audacity 1 . We used the eating session annotations that were provided by the subjects via the Android application interface as guides; based on the session annotations we manually annotated each individual 1https://www.audacityteam.org/ chew with start and stop timestamps. In total, the dataset contains 6hours and 14 minutes of audio, 12 eating sessions (meals), and 9 , 562 individual chews. For the chew duration, the mean is 0 . 348 sec and the standard deviation is 0 . 046 sec. Consumed food types include bread, cucumber, ice-cream, snack bars, and biscuits. The study was approved by the Ethics committee of the Aristotle University of Thessaloniki (151279/2022). The Ethics committee has also approved the public publishing of derivatives of the dataset (such as mel-frequency cepstral coefficients) along with the manual ground-truth annotations and demographics. All participants were informed about the study, the use of their data, and potential public publishing, and have signed a written consent form. 2.3 Detection algorithm To detect chewing activity, we employ an algorithm from literature [22]. The algorithm includes several steps: a pre-processing filter, feature extraction from short, overlapping windows, training of a binary SVM classifier (chewing vs. non-chewing), and SVMscore post-processing smoothing. The detector yields a sequence of binary labels which correspond to chewing activity. Based on that, individual chews, chewing bouts, and then eating sessions can be obtained, using a set of heuristic rules. The code for aggregating the classification labels to meals is available online 2 from [23]. The biggest challenge for detecting chewing sounds in this work is the different placement of the microphones, compared to the “traditional” in-ear placement commonly found in literature. The natural amplification of chewing sounds through the skull is not present here, while talking sounds are better amplified (since this is the target use-case of the smart-glasses). Based on that, we choose the algorithm of [22] as it is resilient to both challenges. First, each audio window is normalized before extracting the features by dividing each audio sample with the standard deviation of the audio samples within said window. This step, in combination with the high-pass filtering (with a very low cut-off point) during the pre-processing stage of the algorithm, essentially result in the standardization of each audio window (i.e. forces a mean of 0and a standard deviation of 1). It should be noted that based on our observations, the captured audio is already zero-mean, so the application of the high-pass filter is redundant. Second, the algorithm uses a small set of carefully selected features, some of which are exceptionally good at differentiating talking from other sounds, such as the fractal dimension [20], the condition number of the auto-correlation matrix, and skewness [23]). Finally, the sensor we use in this work includes two audio channels, left and right. Based on visual inspection of the signals we have identified that the left and right channels are practically equivalent, and can be used interchangeably. However, we also perform some aggregation tests to confirm our inspection conclusions. We examine the following cases: (1) Early fusion of features: for each window, the full feature set is extracted from each channel and the two feature vectors are then concatenated into a single feature vector 2https://github.com/mug-auth/chewing-detection-challenge MADiMa ’22, October 10, 2022, Lisboa, Portugal. Papapanagiotou et al. (2) Late fusion of SVM scores: we train one SVM model per channel, and obtain two SVM scores per window (one from the left and one from the right channel classifier); we then aggregate them using (a) the max operator, or (b) by training a third SVM on the 2D “feature vector” of scores (3) We only use one channel, either the left or the right one It is important to note that the algorithm of [23] is originally applied on audio signals with 2kHz sampling rate. In this work, we adjust the thresholds, filters, and feature extraction parameters accordingly, to accommodate the higher sampling rate of the smartglasses. 3 EVALUATION FRAMEWORK AND RESULTS 3.1 Classifier training Our dataset includes 5subjects. To train the binary SVM classifier we perform leave-one-subject-out experiments. We use the radial-basis function (RBF) kernel. For each training, we select the hyperparameters 𝐶 of the SVM classifier and 𝛾 of the RBF kernel by performing cross-validation on the training data. The hyperparameters space is traversed using Bayesian optimization (instead of a plain grid-search). Finally, to reduce the computational time we do not use all the available training data each time, but we randomly sample only 500 positive and 500 negative windows. Experimental results with greater samples (1 , 000 and 2 , 000 per class) yield similar results but significantly increase the computational time. 3.2 Evaluation framework Applying the algorithm on an audio recording yields a sequence of SVM scores. Thresholding these scores (typically at 0) yields binary labels that correspond to chewing vs. non-chewing. We first evaluate the effectiveness of the trained classifiers directly by computing a binary confusion matrix based on the window labels. We compute precision, recall, and F1-score for each subject and also the mean across subjects. We also extract eating events using the aggregation from chews to chewing bouts, and then to eating events, as described in [22]. The aggregation is applied both on the ground truth and the detected chews. This yields a sequence of ground truth eating sessions (i.e., start and stop timestamps) and a sequence of detected eating sessions. We evaluate by partitioning each recording duration into true positive (TP), true negative (TN), false positive (FP), and false negative (FN) time intervals. We then compute the same metrics as earlier. For this, we count the time that is marked as eating from both the ground truth and detector as TP positive time, the time that is marked as eating from ground truth and non-eating from the detector as FN, and similarly for FP and TN time. We also construct precision vs. recall plots by varying the decision threshold of the SVM scores (after the filtering). 3.3 Results and discussion We first compare the two channels using the early and late fusion methods described in section 2.3. Early fusion yields an average (across subjects) F1-score (on window based evaluation) of 0 . 657 ± 0 . 078. Late fusion yields 0 . 647 ± 0 . 063 using the max operator and 0 . 64 ± 0 . 076 using a third SVM. Finally, using only the left channel 0.05 0.1 0.2 0.3 wstep (sec) 0.68 0.7 0.72 0.74 0.76 F1-score 0.2 sec 0.4 sec 0.6 sec Figure 3: Evaluation results for different window parameters (size and step): F1-score for window-based classification, for three different window sizes (the three colored lines) across different window steps. Table 1: Evaluation results for selected points of the precision-recall curve (marked red in Figure 4). precision recall F1-score 0.9331 1.0000 0.9654 0.9382 0.9041 0.9208 0.9925 0.9032 0.9457 we obtain an F1-score of 0 . 657 ± 0 . 079 and using only the right channel 0 . 642 ± 0 . 08. Based on these results, we may pick any of the two channels. All following analysis is based on the right channel. We then examine the effect of the window size and step. The work of [22] uses two different window sizes: 0 . 2sec for some of the features and 0 . 1sec for the rest. In this we follow the simpler approach and use the same window size for all features; we test values 0 . 2,0 . 4, and 0 . 6. We also test the following window step values: 0 . 05,0 . 1,0 . 2,0 . 3. Figure 3 presents the evaluation results in terms of the F1-score for the window-level classification. The effectiveness benefits from smaller window sizes. The value for 0 . 2 sec window size and 0 . 3window step is expected to be worse since the window step is larger than its size, creating “gaps” of unused signal between successive windows. Based on these results, we select the values of 0.2and 0.05 for size and step. Finally, we formulate chews as “pulses” of the binary SVM score, and aggregate to chewing bouts and then eating events and evaluate based on duration. We then plot the precision-recall curve by varying the threshold for the SVM score (default is 0), to examine the limits of our approach in terms of precision and recall. Figure 4 shows the result. Results are very encouraging, as the area under curve (AUC) is 0 . 9874. We have chosen three points, one with very high precision, one with very high recall, and one balanced, and show the exact values in Table 1. F1-score is equal or greater than 0.92. Chewing Detection from Commercial Smart-glasses MADiMa ’22, October 10, 2022, Lisboa, Portugal. 0 0.25 0.5 0.75 1 1 - precision 0 0.25 0.5 0.75 1 recall Figure 4: Precision-recall curve for duration-based evaluation of eating time. Each point is the average (across subjects) for a different threshold of SVM scores (default is 0 ). Exact values for the three red points are shown in Table 1. Area under curve (AUC) is 0.9874. 4 CONCLUSIONS In this work we examine the feasibility of using off-the-shelf, commerciallyavailable hardware for automatic eating detection. We perform chewing detection based on audio captured by Bluetooth-enabled smart-glasses, the Razer Anzu. We combine this choice with a robust, chewing-detection algorithm from literature, and evaluate on an in-house yet challenging dataset of over 6hours. Detection effectiveness for eating vs. non-eating time is very promising, yielding F1-score above 0 . 92 and as high as 0 . 965 for the best case. Future work includes making a variant of our dataset public to enable direct comparison of similar approaches on this hardware, better leveraging the two microphones and examine the possibility of identifying the side of chewing (left, right, or middle), and finally evaluating on larger and more challenging datasets. ACKNOWLEDGMENTS The work leading to these results has received funding from the European Community’s Health, demographic change and wellbeing Program under Grant Agreement No. 965231, 01/04/2021 - 31/03/2025 https://rebeccaproject.eu/. REFERENCES [1] Mohammad Aboofazeli and Zahra Moussavi. 2009. Swallowing sound detection using hidden markov modeling of recurrence plot features. Chaos, Solitons & Fractals, 39, 2, 778–783. doi: https://doi.org/10.1016/j.chaos.2007.01.071. [2] Oliver Amft, Martin Kusserow, and Gerhard Troster. 2009. Bite weight prediction from acoustic recognition of chewing. IEEE Transactions on Biomedical Engineering, 56, 6, 1663–1672. doi: 10.1109/TBME.2009.2015873. [3] Oliver Amft, Mathias Stäger, Paul Lukowicz, and Gerhard Tröster. 2005. Analysis of chewing sounds for dietary monitoring. In UbiComp 2005: Ubiquitous Computing. Michael Beigl, Stephen Intille, Jun Rekimoto, and Hideyuki Tokuda, (Eds.) Springer Berlin Heidelberg, Berlin, Heidelberg, 56–72. isbn: 978-3-54031941-2. [4] Oliver Amft and Gerhard Troster. 2006. Methods for detection and classification of normal swallowing from muscle activation and sound. In 2006 Pervasive Health Conference and Workshops, 1–10. doi: 10.1109/PCTHEALTH.2006.36162 4. [5] Marios Anthimopoulos, Joachim Dehais, Sergey Shevchik, Botwey H. Ransford, David Duke, Peter Diem, and Stavroula Mougiakakou. 2015. Computer vision-based carbohydrate estimation for type 1 patients with diabetes using smartphones. Journal of Diabetes Science and Technology, 9, 3, 507–515. PMID: 25883163. eprint: https://doi.org/10.1177/1932296815580159. doi: 10.1177/1932296815580159. [6] Marios M. Anthimopoulos, Lauro Gianola, Luca Scarnato, Peter Diem, and Stavroula G. Mougiakakou. 2014. A food recognition system for diabetic patients based on an optimized bag-of-features model. IEEE Journal of Biomedical and Health Informatics, 18, 4, 1261–1271. doi: 10.1109/JBHI.2014.2308928. [7] Sinem Aslan, Gianluigi Ciocca, and Raimondo Schettini. 2018. Semantic food segmentation for automatic dietary monitoring. In 2018 IEEE 8th International Conference on Consumer Electronics - Berlin (ICCE-Berlin), 1–6. doi: 10.1109 /ICCE-Berlin.2018.8576231. [8] Jungman Chung, Jungmin Chung, Wonjun Oh, Yongkyu Yoo, Won Gu Lee, and Hyunwoo Bang. 2017. A glasses-type wearable device for monitoring the patterns of food intake and facial activity. Scientific Reports, 7, 1, (Jan. 2017), 41690. doi: 10.1038/srep41690. [9] Gianluigi Ciocca, Davide Mazzini, and Raimondo Schettini. 2019. Evaluating cnn-based semantic food segmentation across illuminants. In Computational Color Imaging. Shoji Tominaga, Raimondo Schettini, Alain Trémeau, and Takahiko Horiuchi, (Eds.) Springer International Publishing, Cham, 247–259. isbn: 978-3-030-13940-7. [10] Joachim Dehais, Marios Anthimopoulos, Sergey Shevchik, and Stavroula Mougiakakou. 2017. Two-view 3d reconstruction for food volume estimation. IEEE Transactions on Multimedia, 19, 5, 1090–1099. doi: 10.1109/TMM.2016.2642792. [11] Takumi Ege, Wataru Shimoda, and Keiji Yanai. 2019. A new large-scale food image segmentation dataset and its application to food calorie estimation based on grains of rice. In Proceedings of the 5th International Workshop on Multimedia Assisted Dietary Management (MADiMa ’19). Association for Computing Machinery, Nice, France, 82–87. isbn: 9781450369169. doi: 10.1145/3347448.33 57162. [12] Hamid Heydarian, Philipp V. Rouast, Marc T. P. Adam, Tracy Burrows, Clare E. Collins, and Megan E. Rollo. 2020. Deep learning for intake gesture detection from wrist-worn inertial sensors: the effects of data preprocessing, sensor modalities, and sensor positions. IEEE Access, 8, 164936–164949. doi: 10.1109 /ACCESS.2020.3022042. [13] Qianyi Huang, Wei Wang, and Qian Zhang. 2017. Your glasses know your diet: dietary monitoring using electromyography sensors. IEEE Internet of Things Journal, 4, 3, 705–712. doi: 10.1109/JIOT.2017.2656151. [14] Konstantinos Kyritsis, Christos Diou, and Anastasios Delopoulos. 2021. A data driven end-to-end approach for in-the-wild monitoring of eating behavior using smartwatches. IEEE Journal of Biomedical and Health Informatics, 25, 1, 22–34. doi: 10.1109/JBHI.2020.2984907. [15] Konstantinos Kyritsis, Petter Fagerberg, Ioannis Ioakimidis, K. Ray Chaudhuri, Heinz Reichmann, Lisa Klingelhoefer, and Anastasios Delopoulos. 2021. Assessment of real life eating difficulties in parkinson’s disease patients by measuring plate to mouth movement elongation with inertial sensors. Scientific Reports, 11, 1, (Jan. 2021), 1632. doi: 10.1038/s41598-020-80394-y. [16] Frank P. -W. Lo, Yingnan Sun, Jianing Qiu, and Benny Lo. 2018. Food volume estimation based on deep learning view synthesis from a single depth map. Nutrients, 10, 12. doi: 10.3390/nu10122005. [17] Gert Mertes, Hans Hallez, Tom Croonenborghs, and Bart Vanrumste. 2019. Detection of chewing motion using a glasses mounted accelerometer towards monitoring of food intake events in the elderly. In International Conference on Biomedical and Health Informatics. Yuan-Ting Zhang, Paulo Carvalho, and Ratko Magjarevic, (Eds.) Springer Singapore, Singapore, 73–77. isbn: 978-98110-4505-9. [18] Yue Ming, Xuyang Meng, Chunxiao Fan, and Hui Yu. 2021. Deep learning for monocular depth estimation: a review. Neurocomputing, 438, 14–33. doi: https://doi.org/10.1016/j.neucom.2020.12.089. [19] Mark Mirtchouk, Christopher Merck, and Samantha Kleinberg. 2016. Automated estimation of food type and amount consumed from body-worn audio and motion sensors. In Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp ’16). Association for Computing Machinery, Heidelberg, Germany, 451–462. isbn: 9781450344616. https://doi.org/10.1145/2971648.2971677. [20] Vasileios Papapanagiotou, Christos Diou, Zhou Lingchuan, Janet van den Boer, Monica Mars, and Anastasios Delopoulos. 2015. Fractal nature of chewing sounds. In New Trends in Image Analysis and Processing – ICIAP 2015 Workshops: ICIAP 2015 International Workshops, BioFor, CTMR, RHEUMA, ISCA, MADiMa, SBMI, and QoEM, Genoa, Italy, September 7-8, 2015, Proceedings. Vittorio Murino, Enrico Puppo, Diego Sona, Marco Cristani, and Carlo Sansone, (Eds.) Springer International Publishing, Cham, 401–408. isbn: 978-3-319-23222-5. doi: 10.100 7/978-3-319-23222-5_49. [21] Vasileios Papapanagiotou, Christos Diou, Janet van den Boer, Monica Mars, and Anastasios Delopoulos. 2021. Recognition of food-texture attributes using an in-ear microphone. In Pattern Recognition. ICPR International Workshops and Challenges. Alberto Del Bimbo, Rita Cucchiara, Stan Sclaroff, Giovanni Maria Farinella, Tao Mei, Marco Bertini, Hugo Jair Escalante, and Roberto Vezzani, MADiMa ’22, October 10, 2022, Lisboa, Portugal. Papapanagiotou et al. (Eds.) Springer International Publishing, Cham, 558–570. isbn: 978-3-030-688219. [22] Vasileios Papapanagiotou, Christos Diou, Lingchuan Zhou, Janet van den Boer, Monica Mars, and Anastasios Delopoulos. 2017. A novel chewing detection system based on ppg, audio, and accelerometry. IEEE Journal of Biomedical and Health Informatics, 21, 3, 607–618. doi: 10.1109/JBHI.2016.2625271. [23] Vasileios Papapanagiotou, Christos Diou, Lingchuan Zhou, Janet van den Boer, Monica Mars, and Anastasios Delopoulos. 2017. The splendid chewing detection challenge. In 2017 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). (July 2017), 817–820. doi: 10.1109/EMBC.2017.8036949. [24] Vasileios Papapanagiotou, Stefanos Ganotakis, and Anastasios Delopoulos. 2021. Bite-weight estimation using commercial ear buds. In 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), 7182–7185. doi: 10.1109/EMBC46164.2021.9630500. [25] P. Rayneau, R. Bouteloup, C. Rouf, P. Makris, and S. Moriniere. 2021. Automatic detection and analysis of swallowing sounds in healthy subjects and in patients with pharyngolaryngeal cancer. Dysphagia, 36, 6, (Dec. 2021), 984–992. doi: 10.1007/s00455-020-10225-9. [26] Janet van den Boer, Annemiek van der Lee, Lingchuan Zhou, Vasileios Papapanagiotou, Christos Diou, Anastasios Delopoulos, and Monica Mars. 2018. The splendid eating detection sensor: development and feasibility study. JMIR Mhealth Uhealth, 6, 9, (Sept. 2018), e170. doi: 10.2196/mhealth.9781. [27] Rui Zhang and Oliver Amft. 2016. Bite glasses: measuring chewing using emg and bone vibration in smart eyeglasses. In Proceedings of the 2016 ACM International Symposium on Wearable Computers (ISWC ’16). Association for Computing Machinery, Heidelberg, Germany, 50–52. isbn: 9781450344609. doi: 10.1145/2971763.2971799. [28] Rui Zhang and Oliver Amft. 2018. Monitoring chewing and eating in free-living using smart eyeglasses. IEEE Journal of Biomedical and Health Informatics, 22, 1, 23–32. doi: 10.1109/JBHI.2017.2698523. [29] Rui Zhang, Severin Bernhart, and Oliver Amft. 2016. Diet eyeglasses: recognising food chewing using emg and smart eyeglasses. In 2016 IEEE 13th International Conference on Wearable and Implantable Body Sensor Networks (BSN), 7–12. doi: 10.1109/BSN.2016.7516224.