Full text
Frontiers in Nutrition 01 frontiersin.org NutriBase– management system for the integration and interoperability of foodand nutrition-related data and knowledge EvaValenčič 1,2,3,4*, EmmaBeckett 4,5,6, TamaraBucher 4,5, ClareE.Collins 4,5 and BarbaraKoroušićSeljak 1,2 1 Computer Systems Department, Jožef Stefan Institute, Ljubljana, Slovenia, 2 Jožef Stefan International Postgraduate School, Ljubljana, Slovenia, 3 School of Health Sciences, College of Health, Medicine and Wellbeing, University of Newcastle, Newcastle, NSW, Australia, 4 Food and Nutrition Research Program, Hunter Medical Research Institute, Newcastle, NSW, Australia, 5 School of Environmental and Life Sciences, College of Engineering, Science and Environment, University of Newcastle, Newcastle, NSW, Australia, 6 Department of Science, Nutrition Research Australia, Sydney, NSW, Australia Introduction: Contemporary data and knowledge management and exploration are challenging due to regular releases, updates, and different types and formats. In the food and nutrition domain, solutions for integrating such data and knowledge with respect to the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles are still lacking. Methods: To address this issue, we have developed a data and knowledge management system called NutriBase, which supports the compilation of a food composition database and its integration with evidence-based knowledge. This research is a novel contribution because it allows for the interconnection and complementation of food composition data with knowledge and takes what has been done in the past a step further by enabling the integration of knowledge. NutriBase focuses on two important challenges; data (semantic) harmonization by using the existing ontologies, and reducing missing data by semi-automatic data imputation made from conflating with existing databases. Results and discussion: The developed web-based tool is highly modifiable and can befurther customized to meet national or international requirements. It can help create and maintain the quality management system needed to assure data quality. Newly generated data and knowledge can continuously beadded, as interoperability with other systems is enabled. The tool is intended for use by domain experts, food compilers, and researchers who can add and edit food-relevant data and knowledge. However, the tool is also accessible to food manufacturers, who can regularly update information about their products and thus give consumers access to current data. Moreover, the traceability of the data and knowledge provenance allows the compilation of a trustworthy management system. The system is designed to allow easy integration of data from different sources, which enables data borrowing and reduction of missing data. In this paper, the feasibility of NutriBase is demonstrated on Slovenian foodrelated data and knowledge, which is further linked with international resources. Outputs such as matched food components and food classifications have been integrated into semantic resources that are currently under development in various international projects. OPEN ACCESS EDITED BY Massimo Lucarini, Council for Agricultural Research and Economics, Italy REVIEWED BY Sercan Karav, Çanakkale Onsekiz Mart University, Türkiye Didier G. Leibovici, The University of Sheffield, UnitedKingdom Lindung Parningotan Manik, National Research and Innovation Agency (BRIN), Indonesia *CORRESPONDENCE Eva Valenčič [email protected] RECEIVED 28 September 2024 ACCEPTED 13 December 2024 PUBLISHED 06 January 2025 CITATION Valenčič E, Beckett E, Bucher T, Collins CE and Koroušić Seljak B (2025) NutriBase– management system for the integration and interoperability of foodand nutrition-related data and knowledge. Front. Nutr. 11:1503389. doi: 10.3389/fnut.2024.1503389 COPYRIGHT © 2025 Valenčič, Beckett, Bucher, Collins and Koroušić Seljak. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms. TYPE Original Research PUBLISHED 06 January 2025 DOI 10.3389/fnut.2024.1503389
Valenčič et al. 10.3389/fnut.2024.1503389 Frontiers in Nutrition 02 frontiersin.org KEYWORDS database management system, food data compilation, food composition data, food composition database, knowledge base 1 Introduction Food and nutrition-related data and knowledge (D&K) are essential for many research domains, including public health surveillance and promotion, dietary and health assessments, disease prevention, nutrition education, consumer protection, agriculture, food policy, and food labeling (1, 2). D&K, such as food composition data or dietary guidelines, are also necessary for stakeholders in the food industry, retail sector, non-government organisations, policymakers, and ultimately consumers. Consumers rely on D&K when making food and nutrition decisions, while policymakers use food and nutrition-related D&K to obtain accurate scientific evidence needed to design and promote strategies required to improve public health and overall well-being (3, 4). However, D&K are complex, covering diverse areas such as food composition, food safety, food authenticity, and consumption. This paper focuses on food composition data (FCD) and knowledge for dietary assessment and advising. This is highly important for domain experts and policymakers, as well as consumers, including patients. While FCD contains detailed compositional, biochemical, and physiological data of foods (e.g., how much vitamin C apples contain), knowledge provides additional food-related information (e.g., what is the recommended intake of vitamin C). FCD and knowledge are compiled in various databases; however, their integration and interoperability are lacking (5). Improved integration would enable easier access the latest evidencebased D&K from different research areas within a single system. Nowadays, FCD is compiled online in the form of a food composition database (FCDB). FCDBs are usually compiled at the national level but are often used internationally to conduct public health studies (2). Examples include multiple European FCDBs [available through the FoodEXplorer tool (6)], USDA’s FoodData Central (7), FAO/INFOODS databases (8), Canadian FooDB (9), and others. In general, FCDBs contain data on traditional, ethnic, and local foods and dishes, with some combining generic and branded foods [e.g., Serbian (10)] and others maintaining separate databases for different food types [e.g., Dutch branded food database (11)]. In addition to institutional databases, numerous company-owned FCDBs also exist, such as the Edamam’s food, grocery, and (restaurant) database composed using Natural Language Processing (NLP) techniques (12) and GS1 branded foods, and barcode databases maintained through the Global Data Synchronization Network (GDSN) (13). There are two main challenges with existing FCDBs. Namely, data harmonization and missing data. First, FCBDs may contain data of different quality due to differences in data production methods (food sampling, analyses or estimation, (re)calculation, borrowing), data compilation (collection, aggregation, compilation, and dissemination), and data management. The challenge of data harmonization has been addressed by several networks of excellence. For example, the Food CEN standard (14), which defines requirements on the structure and semantics of food datasets and of interchange of food data. Another initiative, the ESFRI research infrastructure Metrofood (15), contributes to the development of aligned metrology services in the food domain. Moreover, when compiling a FCDB, guidelines and frameworks to assess the quality of data, datasets, and databases (16, 17) need to beacknowledged. Several frameworks also enable unified data classification and description, which need to beconsidered when harmonizing various FCDB (2, 18, 19). While these standards and frameworks facilitate the harmonization of foodand nutrition-related data, the problem of linking it with other data types (e.g., medical, environmental, and consumption-related) remains unresolved. The second challenge is related to missing data in FCDBs, which distorts data integrity. Analyzing all components of specific foods poses a significant financial burden for institutions; thus, no FCDB is complete, and updates are not done continuously. The challenge of missing FCD is being addressed in various ways, including borrowing data from other databases, performing tedious manual work, or using computersupported methods for (semi-) automated data imputation (20, 21). On the other hand, together with databases, knowledge bases (KBs) are also very important resources. By definition, a KB is an easily accessible online library of collected and organized information and documentation about certain topics (22). The important knowledge that should beincluded in food and nutrition KB should include, but not be limited to: standardized classification and description of coding systems [e.g., LanguaL (23), FoodEx2 (24), INFOODS (8)]; standardized value documentation (e.g., acquisition type, method type) (18); a chemical databases of molecular entities– ChEBI (25); retention and yield factors used to calculate the nutrient content of composite dishes or recipes (26); standardized household measurement units; national dietary reference values and dietary guidelines; physical activity standards; food components’ bioavailability; food-drug interactions, and others. As knowledge accumulates quickly, the creation and maintenance of a KB is tedious work, usually done manually by domain experts. However, semantic resources have complemented KBs and allowed interoperability of D&K from various research domains. Semantic resources like the ontologies [e.g., FoodOn (27), ISO-FOOD (28), FNS-Harmony (29), COMFOCUS (30)] or knowledge graphs [e.g., describing complex relationships between food and biomedical factors (31)] are being developed to formally describe knowledge as a set of concepts and the relationships between those concepts within a domain. To link FCD with semantic resources, FCD needs to be annotated with standardized metadata in machine-readable formats to enable connectivity of terms across different data sources. Regardless of all research efforts, applicable KBs providing integrated knowledge on food and nutrition are still lacking. There are few KBs that focus on specific subdomains, such as FoodKG (32) for food recommendation based on diet-related knowledge or TasteAtlas (33), a world atlas of traditional dishes, local ingredients, and authentic restaurants. Abbreviations: API, Application Programming Interface; D&K, Data and knowledge; DKBMS, Dataand knowledge base management system; FCD, Food composition data; FCDB, Food composition database; KB, Knowledge base; NLP, Natural Language Processing; KPI, Key performance indicators; MTBF, Mean time between failures.
Valenčič et al. 10.3389/fnut.2024.1503389 Frontiers in Nutrition 03 frontiersin.org The food and nutrition community has created many FCDBs as well as few KBs, but their integration and interoperability are currently missing. Even when limited just to the integration within FCDB, information is not harmonized because different coding systems, documentation or standards are used. Some examples of best practice using harmonized FCD are FoodEXplorer (6), FoodCASE (34), FoodData Central, Glycemic Index Research and GI News (35). Some of these tools even enable comparison of FCD from multiple countries. This is important as, with increasing globalization, the availability of international foods and dishes is increasing, and obtaining datasets of non-local foods is necessary. Having databases composed on a national level is important; however, for applied science, it would beuseful if compilers could link and integrate not only FCD with each other but also FCDBs with KBs. This is something that webelieve does not yet exist or is not publicly available in the food and nutrition domain. The importance of integration and interoperability was also highlighted in the recent paper by Durazzo etal. (36), which further emphasized the necessity of cooperation and D&K sharing between compilers. However, the connectivity among computer systems and/ or online platforms is equally necessary. In the current paper, weintroduce a new database management system, called NutriBase, for integrating FCD from different databases with foodand nutrition-related knowledge. The integration is performed in a transparent way and enables, together with harmonization, a reduction in missing data. In Section 2, weexplain how publicly available D&K resources, which (currently) represent the baseline of the NutriBase, were identified and collected. Next, weintroduce NutriBase and describe its functionality. In Section 3, wedescribe the compilation process of the Slovenian FCDB and KB, identify issues, discuss possible solutions the system offers, and provide plans for future work. Weconclude the paper in Section 4. 2 Materials and methods 2.1 Data and knowledge collection To demonstrate the feasibility of NutriBase, Slovenian FCD and both, national and international semantic resources were collected. Firstly, the analytical compositional data on generic foods from the Slovenian FCDB composed in 2006 and updated in 2012 (37) were imported. The recipes included in the Slovenian FCDB were imported separately, as they require different data handling, such as consideration of yield and retention factors, as well as standards for calculating recipes (38, 39). In addition, branded foods that can currently bepurchased in Slovenia, are being uploaded through an application programming interface (API) from the Composition and Labeling Information System (CLAS) (40). To complement the Slovenian FCDB for generic foods, six publicly available FCDBs (Table1), together with associated metadata and documentation, were either downloaded or linked through an API in late 2020 or 2021. The imported FCDBs consisted of datasets in different formats, and not all of them adhered to the Food CEN standard (14). The imported metadata and documentation include various background information, such as explanations of data sources, procedures for data quality assurance, descriptions of foods and food group classification levels, and explanations of specific component descriptions, calculations and units used. Multiple foreign FCDBs needed to beimported because they contain different data. For example, FoodData Central (US in Table2) in addition to FCD, provides also the data for household measurement units (e.g., tablespoon, cup, dash) which can belinked to generic foods. Moreover, different components are collected or analyzed across different FCDBs. For instance, some datasets contain data for total carbohydrates (digestible and indigestible, including dietary fiber), whereas others contain only data for available carbohydrates. From the currently imported FCDBs only three provide data for total carbohydrate, however all of them contain data for available carbohydrates and total dietary fiber, thus the total carbohydrates could becalculated. Lastly, relevant evidence-based food and nutrition knowledge was systematically reviewed and collected from publicly available national and international resources, and was further compiled into the NutriBase KB (Table2). The approaches and tools applied and described in the current paper can beused for D&K from any country. The Slovenian D&K are used as an example only. Unlimited publicly available FCDBs and/or KBs can beuploaded or linked via an API to create a new database, as long as they comply with the NutriBase requirements. 2.2 NutriBase- dataand knowledge base management system NutriBase is designed to enable easy integration with other KBs and semantic resources conceptualizing the health, environmental, consumer behaviors, and food and nutrition domains in particular. This dataand knowledge base management system (DKBMS) has TABLE1 FCDBs currently included in the NutriBase. Currently Imported FCDBs Country code No. of components No. of food group levels‡ No. of foods / dishes Source file format SI 773*15 993 .CSV/.XSL 48 149 FR 60 10 2,807 .CSV/.XSL 58 83 NL 133 27 2,152 .CSV/.XSL DK 197 18 1,186 .CSV/.XSL 127 UK 178 14 2,910 .CSV/.XSL 71 54 AU 249 22 1,534 .CSV/.XSL 97 US 235 28 7,793†, 210¨ API *651 from EuroFIR Thesauri document and 122 subsequently added (own); ‡ = the top number is the highest level, the bottom number is the lowest (the most detailed) level (sublevel); † = SR Legacy Foods; ¨ = Foundation Foods.
Valenčič et al. 10.3389/fnut.2024.1503389 Frontiers in Nutrition 04 frontiersin.org TABLE2 Resources included in the NutriBase KB. Scemantic resources Resource name/type and reference Knowledge type Description Number of entities Standardized classifications and description coding system FoodEx2 classification A food classification and description system developed by EFSA- includes different hierarchies and facets for different food safety domains. (e.g., A00KR#F27.A00KV$F27.A00LN $F27. A00LB$F27.A00LG; mixed leafy vegetables) 4,445 Standardized value documentation (11) Component type Component identifiers and descriptors (e.g., CHO; carbohydrate; use for total of those carbohydrates digested and absorbed in the intestine; total accessible carbohydrates include free sugars, polyols and dextrins, starch, and glycogen). 660 (9 of these are for backward compatibility only) 56 classification identifiers (not used for new indexing) Unit E.g., grams, millimoles, alphatocopherol equivalent, per cent. 19 Additional 20 added (IU, g/kg body mass, etc.) Matrix unit E.g., per 100 g of total food, per 100 mL food volume, per unit, per 100 g edible portion. 20 matrixes Value type E.g., arithmetic mean, best estimate, average, below limit of detection, trace. 20 types Method type Reporting if the value was analyzed, calculated or imputed (e.g., calculated as recipes, calculated from related food, analytical result). 20 types Method indicator Providing details for the analytical method or formulas used for calculation (e.g., chromatography, difference, ash calculated as sum of minerals). 214 indicators Acquisition type Describes the origin of the value (e.g., laboratory, food composition table, authoritative document). 12 types Reference type E.g., article in journal, file or database, product label, software. 14 types LanguaL thesaurus (10) Cooking methods E.g., griddled, cooked by microwave, deep fried. 47 methods FoodData Central at US Department of Agriculture (USDA) Measurement and household units E.g., tea spoon, slice, filet, cup, could beused for volume to weight conversions. 115 (currently in use) out of 1923 ChEBI- a chemical database and ontology of molecular entities, which is part of the Open biomedical ontologies at the EBI, and European ELIXIR infrastructure Dictionary of molecular entities Providing detailed data of chemical entities of biological interest (e.g., definitions, formulas, ontologies, chemical reactions, IUPAC names and identifiers) 210 linked to added components SciName Finder (26) Search tool for scientific and common names of plants and animals Providing precise identify plants and animals Allows precise identification of plants and animals, and searching the information on scientific and common names provided by authoritative resources (and not from secondary sources) More than 1,000,000 scientific and common names Culinary groups [adapted from (18, 23)] Culinary groups / subgroups related to retention and yield factors. Providing the basics for obtaining nutrient content of foods by calculation methods (as recipe calculation), based on the amount of ingredients given in a recipe, nutrient composition of ingredients and factors that consider changes in nutrient content (retention factors), and weight (yield factors) during preparation. 31 groups and subgroups related to yield factors, and 38 related to retention factors Slovenian dietary reference values (DRVs) (27) based on the D-A-CH reference values adopted by the Ministry of Health of the Republic of Slovenia DRVs Reference values for energy and nutrient intake for children (at least 1-year old), adolescents, adults, elderly, pregnant women and nursing mothers. 34 references for energy, macroand micronutrients, for men and women (10 different age groups) Latest dietary guidelines and recommendations National and international dietary guidelines and recommendations Relevant evidence-based guidelines and recommendations for different consumers (athletes, pregnant women and nursing mothers, healthy individuals from different age groups). Currently defined for biomarkers (blood cholesterol and glucose) and endurance sports. (Continued)
Valenčič et al. 10.3389/fnut.2024.1503389 Frontiers in Nutrition 05 frontiersin.org been implemented as a web-based tool (Figure1) for food compilers to easily explore, compile and most importantly, link data from different FCDBs and KBs. The main goal of this process is achieving an optimal linking of D&K, which enables borrowing data respecting the FAIR (Findability, Accessibility, Interoperability, and Reusability) principles for data management (41), and reducing missing D&K. 2.2.1 FCDB compilation To ensure a semi-automatic connectivity among different sources (FCDBs), standardized components and food groups matching had to bemanually performed (Figure2; Step1 and 2). Since the composition of food depends on its geographical origin, it is important to also consider the data source and the data most closely related to local foods. Therefore, a pre-set priority list of data sources is integrated within the system and can beadapted if needed. For Slovenian example this means that European datasets are prioritized before non-EU datasets. This allows experts to semi-automatically compile datasets that are as complete as possible, while also transparently providing the source of specific data (e.g., component value). The pre-set priority list can easily beamended or set for different countries. Moreover, a comparison of a national dataset (in our case, Slovenian), with other, foreign datasets is also enabled. This feature allows borrowing specific data from other FCDBs. Together with food composition data, compilers can also check additional value information, such as value type and method type (if provided). Being able to check additional value information and standards, allows compilers to assess the quality of the data and select the most appropriate or accurate one. Additionally, during the FCDB compilation process, basic food information and metadata, such as generic and/or commercial names, allergens, ingredients, food origin, and food images, are also addressed and can beborrowed. NutriBase presents an infrastructure that can beadapted for FCD from any country. However, to achieve an optimal linking of D&K and to ease and expedite FCDB compilation, various knowledge resources had to beconsidered. 2.2.2 Knowledge base compilation In the NutriBase underlying thesaurus, knowledge about relevant foodand nutrition topics is collected and maintained. The KB, implemented within the DKBMS, is connected with all three steps of the workflow seen on Figure2. Thus, all updates of the KB content will have an immediate impact on linked data in FCDB. That means whenever a new data or knowledge is published, it can easily beimported and linked to existing D&K or substituted for the latest findings. An important part of the implemented KB is food naming by using tags. It provides functionality for unique food naming and metadata annotation. While much work has already been conducted on unifying food description and classification, food naming is still an open issue. Therefore, wehave implemented a new food-tagging approach to unify and standardize food naming within the FCDB. This is especially useful when different users are working on a FCDB, as it enables unambiguous communication between all users involved in the working process. In addition, together with using tags, setting rules for food naming has been proposed as another solution. 2.2.3 Usability of NutriBase Lastly, the usability of the newly developed system was evaluated. Wedistributed the System Usability Scale (SUS) questionnaire among regular NutriBase users with different profile roles. The SUS tool is a reliable and validated tool for measuring the usability, which is frequently used by evaluators of mHealth services (42). It consists of a 10-item questionnaire with five response options for respondents (strongly agree to strongly disagree). The survey was completely anonymous and after collecting the responses, the participant’s scores were carefully interpreted to produce a percentile ranking. 3 Results and discussion 3.1 The compilation of the Slovenian FCDB and KB Throughout the entire compilation process (Figure2), D&K were maintained in accordance with the FAIR principles. Managing D&K to ensure that the format of foreign FCDBs and KBs remains unchanged from the original sources has been a key requirement in NutriBase’s development (Figure3). 3.1.1 Components matching To create and link the Slovenian database, the compilation process was initiated by components matching (Figure 2, Step 1). The Slovenian FCDB complies with the CEN Food standard (14), therefore the components specified with respect to the EuroFIR thesaurus for components (18) were manually matched with components from the foreign FCDBs (Figure4 presents the user interface of this process). Although most of the foreign selected FCDBs also comply with the CEN Food standard, mismatched components (i.e., different names for the same components among different countries) were still present (examples are shown in Table3). Components were matched manually by domain experts to ensure a correct and unambiguous matching. Moreover, the result can beprovided as an input to the FNS-Harmony ontology (43), which has been developed within the FNS-Cloud project to support interoperability of foodand nutrition-related data in the European Open Science Cloud (EOSC) and is available through the NCBO Bioportal. NutriBase could beintegrated with FNS-Harmony, which reuses or incorporates several ontologies, including FoodOn (27). In TABLE2 (Continued) Scemantic resources Resource name/type and reference Knowledge type Description Number of entities Physical activity related standards Metabolic equivalent of task (METS) E.g., basketball, swimming, mopping, walking, sitting. 541 tasks Physical activity level (PAL) E.g., sedentary or light activity lifestyle. 5 levels per sex
Valenčič et al. 10.3389/fnut.2024.1503389 Frontiers in Nutrition 06 frontiersin.org FIGURE1 User interface of NutriBase. this case, food compilers would not only beable to provide but also use new knowledge about semantic integration with other systems, such as GS1 GDSN (44). 3.1.2 Food groups matching Firstly, food groups were designed based on the classification of foods used by relevant information systems in Slovenia, as well as the EuroFIR standard (18), which is intended for generic foods. Since Slovenian FCDB also includes branded foods, classification systems for these had to be considered as well. However, we found that different Slovenian institutions use different classification systems. This suggests that even within a single country, it might benecessary to follow and comply with several standards. For example, the Slovenian classification system, which is based on public procurement and is determined by law, or the Dunford classification system (45), specifically developed for branded foods. Currently, the Slovenian FCDB includes three hierarchical classification levels: 15 groups on the first, 48 groups on the second, and 160 on the third (and most detailed) level. In addition to manually matching national food classification systems with one another, the food groups used in Slovenian FCDB were also matched with those used in the foreign FCDBs (Figure2, Step2). An example of a matched food group- Fresh vegetables, among FCDBs is presented in Table 4. The task of food groups matching was especially challenging, as different countries use different numbers of classification levels. For example, foods in France and the UK are classified into up to three levels, in Australia and Denmark into two levels, and in the Netherlands and USA into just one level. Moreover, the level of detail within food groups varies. As shown in Table4, some countries group all vegetables together, while the others sub-classify them further (e.g., root vegetables, fruiting vegetables). To ensure accurate food classification and assist users in using NutriBase, a feature was implemented allowing compilers to add examples of foods allocated to specific food group. This feature was found to bevery useful, as it enables users to unambiguously select the correct food group. Additionally, manually matched food groups can also beprovided as inputs into FNS-Harmony. 3.1.3 FCDB compilation FCDB compilation process (Step 3 in Figure 2) began with manually checking and correcting a dataset of 14,064 entries for 443 generic foods analyzed by the Biotechnical Faculty of the University of Ljubljana in 2006 and 2012 (37). Together with the composition data, annotated metadata (e.g., value information) were also reviewed. Certain components were specifically checked to ensure compliance with the standards. For example, the differences between total available carbohydrates and total carbohydrates. This entire process aligns with the first 12 steps of the generic compilation process described by Westenbrink et al. (2), currently excluding Step 5 (attribution of quality index) and Step11 (physical storage). The evaluation of Slovenian data quality (17) and the database quality evaluation, as suggested by the recently published FAO/INFOODS framework (16), are currently underway. Next, the Slovenian name for each generic food was reviewed, and a scientific name (when appropriate), an English name, and synonyms were assigned based on the new food-tagging approach. To achieve this, tags were defined, and rules for their application were established within each food group. During this process, wefound that similar foods might have different names. This can make searching for a specific food within the FCDB harder for compilers as well as for consumers accessing publicly available FCDBs. For example, the only difference between ‘Baked eggplant with added cheese and tomato sauce’ and ‘Aubergine prepared in tomato sauce and cheese, frozen’ is that one is baked and the other is frozen, but the names are very different. Therefore, using tags for food naming, helps unify the FCDB and simplifies searching for specific foods. Moreover, weensured the naming is clear to all users, specifically for consumers accessing FCD
Valenčič et al. 10.3389/fnut.2024.1503389 Frontiers in Nutrition 07 frontiersin.org (e.g., via a mobile app), who may find it challenging to understand the processing conditions of foods. For example, meat can beanalyzed as raw (e.g., beef filet) or heat-treated (e.g., beef filet, grilled). However, experience shows, it is seen that consumers do not consider ‘beef filet’ as raw, but rather as ready-to-eat steak. Therefore, adding the ‘raw’ tag to raw meat seemed reasonable. On the other hand, it is clear to consumers that ‘banana’ is raw, and they do not expect this tag added to fresh fruits. Thus, the ‘raw’ tag is used in some food groups but not in others. In addition, the tag ‘peeled’ is used only when appropriate (e.g., ‘apple, peeled’, but not ‘banana, peeled’). Currently, each food group at the third hierarchical level within the tool has an average of 15.4 tags. Additionally, the initial Slovenian dataset of generic foods was manually linked with the same or similar food items from the selected foreign FCDBs. The linking was carried out by domain experts. First, the English names were compared, followed by a comparison of the main food components. In case the food composition was similar, the food items were linked together and the missing data were imputed from the foreign FCDBs. Table5 presents an example of the number of imputed data for Fresh vegetables food group from a specific FCDB. As can beseen, only one value for total carbohydrates could be borrowed from US database, while the rest were taken from Slovenian FCDB. However, cystine values are missing in Slovenian FCDB, so they were borrowed from the Danish and US databases (the other FCDBs do not contain data for cystine). The NutriBase allows linking one food with multiple foods within one database or across multiple databases. For example, the Slovenian ‘average white bread’ can belinked with ‘white baguette’ and ‘white loaf’ from one FCDB, and with ‘white bread’ from the other FCDBs. The borrowed data will, however, bedisplayed based on the pre-set priority list of FCDBs. In our case, when a food item is linked with food item(s) from across different FCDBs, data from European datasets were prioritized before non-EU datasets. However, compilers can manually change the data source and select (borrow) non-EU data to bedisplayed if it is more appropriate. Wefound this approach to bevery convenient, as it provides compilers with data most closely related to the local foods, but it still gives them freedom to select another data. Moreover, the manually matched foods present a valuable asset that can beused to construct a gold standard corpus, i.e., a corpus of text annotated with food entities required for NLP techniques, such as CafeteriaFCD (46). Same as generic foods, the branded foods can also belinked with similar generic foods from either national FCDB or foreign FCDBs. In this case, the original FCD of a branded food is taken from the nutrition declaration table, while the FCD not provided on the nutrition declaration table (e.g., micronutrients) can beimputed from FCDBs and transparently marked as such. This is especially beneficial when FIGURE2 Flowchart of compilation process to link foods from different FCDBs.
Valenčič et al. 10.3389/fnut.2024.1503389 Frontiers in Nutrition 08 frontiersin.org FIGURE3 Overview of NutriBase structure. collecting food consumption data for the national food consumption survey. As seen in the EU Menu project, consumers usually provide only the brand or production line of the food item when reporting food intake. For example, instead of reporting consumption of ‘full fat milk’, they reported a producer’s name of such milk. Since the nutrition declaration table usually only provides the information of energy value and six other nutrients, the values of micronutrients are unknown. Thus, branded foods could belinked with generic foods to compose the complete dataset, which would provide the opportunity to more accurately assess food intake of individuals and overall population. Finally, yet importantly, internationally accepted algorithms to avoid errors were selected and applied to produce aggregated data [e.g., recipe calculations) (Steps 14 and 15 according to Westenbrink etal. (2)]. In addition, the compiled and aggregated data within the NutriBase were verified [and corrected if needed) (Steps 16 and 17, according to Westenbrink etal. (2)] to prevent hazards related to data validation. The majority of the FCD validation has been done manually, however the tool automatically performs consistency checks for some metadata and components (e.g., content of specific component is not larger than 100 g (converted regardless of the unit), the sum of proximities is ≤105 g, value of saturated fatty acids is not larger than value of total fats, etc.). The validated data is then stored and disseminated [Steps 18 to 22, according to Westenbrink etal. (2)]. 3.1.4 Knowledge base creation Using semantic resources, a KB was created to support the optimal food compilation process, as well as for data quality assessment, traceability, calculations and validation. The KB implemented within the NutriBase is meant to beused by domain experts, as it collects the latest scientific evidence and documentation required for data management and data source management. The KB also consists of the reference list and it allows publication metadata to beimported in standardized formats (e.g., bib). These references can befurther linked to specific data/information, which allows traceability of data and metadata. Moreover, the information can beedited or added to the existing KB and updated accordingly. For instance, units listed in the EuroFIR value documentation (18) can besupplemented or extended with other units (e.g., IU, ABV) to meet the compilers’ needs, or they can beupdated if changes are made to the existing EuroFIR standards. 3.1.5 Linking FCDB with knowledge Linking FCD from different sources is important, and linking knowledge from various sources is equally crucial. Both types of linking can beperformed in NutriBase; however, the system also enables the linking of FCD with knowledge. For instance, a specific component (e.g., vitamin C) can belinked with a relevant dietary recommendations, such as Slovenian DRVs (47). Therefore, within the tool, data (component; vitamin C) was interconnected and complemented with knowledge (dietary requirements for vitamin C), enabling access to combined information in one place. This approach takes what has been done in the past a step further by incorporating knowledge into the system, which can be especially useful for informing and educating consumers (e.g., via mobile apps). Instead of providing consumers or app users with just FCD, the incorporated knowledge can also be provided, which can deliver a more personalized approach. Our work is consistent with previous works (5, 27, 48), with the difference that NutriBase is a practical and applicable tool, whereas the previous works is theory based.
Valenčič et al. 10.3389/fnut.2024.1503389 Frontiers in Nutrition 09 frontiersin.org 3.1.6 Tool validation The NutriBase and its functionalities were validated throughout the entire compilation process of the FCDB and KB. Seven experts who regularly use NutriBase evaluated it using the SUS tool, which is used for judging the perceived usability of systems. The SUS score was 78.9, which falls to 85th percentile and corresponds to grade A-. Moreover, six food compilers of different skills have performed various tasks (e.g., component matching, food linking) depending on their user profile role. For example, less skilled compilers have only edited D&K, whereas more experienced compilers performed more demanding tasks. Regardless of their skill level, all users agreed that the system is a helpful, easy-to-use tool when compiling a FCDB, especially because it collects all relevant and needed D&K in one place. 3.2 Strengths and limitation of DKBMS While reviewing analytical data of generic foods from the past Slovenian FCDB and importing it into the DKBMS, some errors and gaps were identified and further discussed with compilers. The data was reviewed using spreadsheets, and it was found that errors were difficult to identify. However, when using NutriBase to review and edit the FCD, users agreed that it is a useful and reliable tool. Although spreadsheets are very popular when handling data, a similar finding was reported by Presser etal. (34). To assess the quality of D&K, it is crucial to develop and maintain a quality management system (2). Currently available FCDBs contain data of varying quality, mainly due to the use of different resources and different methods of data acquisition. The metadata used to describe them, as well as the quantity of data differ among FCDBs. Therefore, compilers need to follow standardized guidelines, provide quality indexes for their original data, and further evaluate their FCDB. This will help domain experts select the best high-quality dataset and/or FCDB for their purposes, which can further beused to obtain accurate results in research, education, and in decision making for policy and programming (16). Not only is NutriBase a useful tool to help domain experts compare different datasets and therefore select the most appropriate one, it can also help national compilers to evaluate their own original data and metadata, and ensure the quality datasets. Moreover, an advantage of the system is also that food manufacturers can gain direct access, and add or edit food-related data of their products. In this way, important information about branded foods currently available in stores can beregularly updated and shared with consumers. The usage of FCDBs may besignificantly restricted due to the missing data (3). It has been proposed that it is better to include imputed data, transparently identified as such, than no data at all (3). However, data should only beborrowed or imputed among the same or similar foods. Several computational methods for missing data imputation within FCDBs have been previously researched (20, 21). All of them concluded that, in order to ‘borrow’ data, as many details as possible FIGURE4 User interface of component matching process. TABLE3 An example of component matching of Slovenian components with components from foreign datasets. Component names among different FCDBs SI FR NL DK UK AU US Carbohydrate, total (CHOT) / / Carbohydrate by difference; g/ / Carbohydrate, by difference; Unit: G Carbohydrate (CHO) Carbohydrate (g/100 g) CHO g Carbohydrates, available; g, Carbohydrate, declaration; g Carbohydrate (g); CHO Available carbohydrate, with sugar alcohols; (g) Carbohydrate, by summation; Unit: G Fiber, total dietary (FIBT) Fibers (g/100 g) FIBT_g Dietary fiber; g AOAC fiber (g); AOACFIB Total dietary fiber; (g) Total dietary fiber (AOAC 2011.25); Unit: G, Fiber, total dietary; Unit: G Fat, total (FAT) Fat (g/100 g) FAT_g Fat, total; g Fat (g); FAT Total Fat; (g) Total lipid (fat); Unit: G Fatty acids, total saturated (FASAT) FA saturated (g/100 g) FASAT_g Sum saturated; g Satd FA /100 g FA (g); SATFAC, Satd FA /100 g (g); SATFOD Total saturated fatty acids;(%), Total saturated fatty acids; (g) Fatty acids, total saturated; Unit: G