scieee AI-readable full text Open interactive document viewer

An AI knowledge-based system for police assistance in crime investigation

Fernández Basso, Carlos Jesús,Gutiérrez Batista, Karel,Gómez Romero, Juan,Ruiz Jiménez, María Dolores,Martín Bautista, María José

Abstract

COPKIT project, which has received funding from the European Union's Horizon 2020 Research and Innovation Programme under grant agreement No 786687

Full text

ORIGINAL ARTICLE An AI knowledge-based system for police assistance in crime investigation Carlos Fernandez-Basso 1,2 | Karel Gutiérrez-Batista 1 | Juan G omez-Romero 1 | M. Dolores Ruiz 1 | Maria J. Martin-Bautista 1 1 Computer Science and Artificial Intelligence, University of Granada, Granada, Spain 2 Causal Cognition Lab, University College London, London, UK Correspondence Carlos Fernandez-Basso, Causal Cognition Lab, University College London, London, UK. Email: [email protected]k, [email protected] Funding information EU-funded Margarita Salas programme NextGenerationEU; European Union's Horizon 2020 Research and Innovation Programme, Grant/Award Numbers: 786687, 101121309; NextGenerationEU/PRTR; MCIN/AEI, Grant/Award Number: PID2021-123960OBI00; ERDF, Grant/Award Number: TED2021-1289402B-C21 Abstract The fight against crime is often an arduous task overall when huge amounts of data have to be inspected, as is currently the case when it comes for example in the detection of criminal activity on the dark web. This work presents and describes an artificial intelligence (AI) based system that combines various tools to assist police or law enforcement agencies during their investigations, or at least mitigate the hard process of data collection, processing and analysis. The system is an early warning/early action system for crime investigation that supports law enforcement with different processes to collect and process data as well as having knowledge extraction tools. It helps to extract information during the investigation of a criminal case or even to detect possible criminal hotspots that may lead to further investigation or analysis of a criminal case Abu Al-Haija et al. (2022, Electronics, 11, 556). The functionality of the proposed system is illustrated through several examples using data collected from the dark web, which includes advertisements offering firearms-related products. KEYWORDS artificial intelligence, association rules, crime detection, knowledge repository 1|INTRODUCTION Artificial intelligence (AI) tools have gained prominence over the last decade to assist in very diverse activities. The fight against crime, delinquency and terrorism is no exception, although it poses some challenges: for instance, the functioning of tools must be transparent for the law enforcement agencies (LEAs) and must avoid discrimination in the obtained results. In addition, the AI algorithms, their functioning and the obtained results should be understandable by LEAs in order to assist them in their daily work. The use of AI is particularly interesting when big data volumes have to be analysed, as in the case of social media or the dark net. The increasing use of these channels for criminal or illegal activities and the massive number of publications to be analysed is one of the main challenges facing police forces. Moreover, the inherent difficulty of extracting information from natural language is an important factor in automatically using textual information for advanced knowledge discovery Rawat et al. (2021). The combination of natural language processing (NLP) technologies with knowledge discovery (KD) tools seems to be the new direction to follow to extract more information, not only from numerical data but also from textual data All authors contributed equally to this study. Received: 25 January 2023 Revised: 31 October 2023 Accepted: 4 December 2023 DOI: 10.1111/exsy.13524 This is an open access article under the terms of the Creative Commons Attribution-NonCommercial License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited and is not used for commercial purposes. © 2024 The Authors. Expert Systems published by John Wiley & Sons Ltd. Expert Systems. 2024;e13524. wileyonlinelibrary.com/journal/exsy 1of16 https://doi.org/10.1111/exsy.13524 as it can be seen for instance in Griol-Barres et al. (2020). In this regard and focusing on the general ambit of crime, the existing analysis/ forecasting techniques have been used in an isolated way as can be derived from the literature (see for instance Hassani et al. (2016), Li (2021) and Qayyum and Dar (2018) where reviews for data mining (DM) tools in the ambit of crime can be found). In this paper, we propose and describe a system based on some AI tools (Fernandez-Basso et al., 2016, Fernandez-Basso, Ruiz, et al., 2019)to assist LEAs during their investigations, helping or at least mitigating the hard process of data collection, processing and analysis. The proposed system is not intended to detect crime automatically but to provide assistance with new insights and findings that may assist in the investigation of a criminal case or even in the identification of possible criminal spots that can derive from a criminal case for further investigation. This methodology is known as the EW/EA (Early Warning/Early Action), which is related to the SOCTA 1 methodology used by EUROPOL. The EW encompasses the identification of “weak signals”, warnings and new insights that can be used for support at both strategic and operational levels, to develop, in a further step, an EA which includes the different measures/decisions to mitigate, prevent or prepare future security policies. The system is composed of different parts utilising several AI tools and algorithms starting with automatic crawling and NLP extraction tools, followed by a Knowledge Repository (KR) component containing processed knowledge that the user (in our case law enforcement agencies) has to provide with technical assistance, and ending with the application of KD tools that analyse the collected data to find new relations and insights. This extracted new knowledge could be used to (1) further enrich the knowledge repository, (2) help LEAs in their situation assessment process, and (3) find new information to consider in the criminal case under investigation. The main novelty of proposed system architecture is that (1) it implements a complete system capable of starting the analysis from the raw data improving these data through the knowledge base including provenance, expert knowledge, and privacy by design to the last step consisting in the visualisation of the results and (2) it allows a continuous evolution by incorporating not only the initial expert knowledge provided by LEAs but also the knowledge provided by the KD tools such as the Association Rules (see Section 5). We have also compared the proposal with other similar systems in Section 2to highlight its main features. The proposed system architecture has been employed in the EU-funded COPKIT 2 project. The COPKIT project addresses the problem of analysing, investigating, mitigating and preventing the use of new information and communication technologies by organised crime and terrorist groups. To this end, COPKIT proposes an intelligence-led early warning (EW)/early action (EA) system for both strategic and operational levels. It should be noted that, due to security restrictions, some details cannot be provided. The main system has been summarised in Section 3, highlighting the two components that we will explain in more depth: (a) the Knowledge Repository and (b) a KD tool called Association Rules. We also present a use case for the system based on the dataset called “Grams” 3 which is a structured dataset containing a collection of advertisements from different darknet markets. In particular, we have focused on those ads that offer weapons-related items. Using this dataset, we will describe the functionality of KR and KD components. The paper is structured as follows. Next section provides a brief overview about other projects and proposals that aim to assist LEAs in the fight against crime. Section 3presents the system architecture. Section 4describes the knowledge repository component and some of its functionalities. Section 5explains the knowledge discovery tool for the extraction of frequent itemsets and association rules. Section 6develops a use case based on firearms trafficking to illustrate the operation of the system. Finally, the paper finishes with the conclusions. 2|RELATED WORKS One of the European Council objectives is the fight against crime while preserving the citizens’rights. Therefore, we can find several EU funded projects pursuing this goal. One of these projects is the TITANIUM (2020) project, which researches, develops, and validates novel data-driven techniques and solutions designed to support law enforcement agencies. In this way, it helps to investigate criminal or terrorist activities involving virtual currencies and/or underground markets on the darknet. The result of TITANIUM is a set of services and forensic tools, that operate within a privacy and data protection environment that is configurable according to local legal requirements. Another project with the same theme as the COPKIT (2021) project is TENSOR (2019) project. This is another example where the primary goal is to keep people safe. The project, which is funded by the EU under the Horizon 2020 programme, seeks to develop a platform offering Law Enforcement Agencies fast and reliable planning and prevention capabilities for the early detection of terrorist activities, radicalisation and recruitment. This project focuses on the collection of data in order to extract interpretative information such as the correlation of information, validation of sources, extraction of events, hidden meanings and high-level interpretations. All of this aggregated and summarised in a final application. The i-LEAD project (2023) project also covers the fight against crime, focusing on the needs of LEA users who are the main sources of expertise. These are divided into five groups of experts. Each group defines the current situation regarding the use of technology by law enforcement agencies, assesses baseline capabilities through a capability map, identifies capability gaps and opportunities for innovation, defines priority areas for improving performance through innovative methods, and identifies potential areas for standardisation at EU level. This project is more focused on the management and analysis of the requirements and capabilities of each LEA, as well as the improvement and development of new technological capabilities for them. The MAGNETO (2021) project and Pourhabibi et al. (2021) aim to establish a continuously improving crime 2of16 FERNANDEZ-BASSO ET AL. 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License prevention and investigation scheme. For that, heterogeneous data streams are transformed into knowledge bases according to a sophisticated representation model, which are then processed and fused using semantic technologies. The results are visually represented by immersive HMIs (Human Machine Interfaces), enabling timely and accurate decision-making, situational awareness and court-proof evidence extraction. We can also distinguish another group of projects such as: ANITA (2021), CC-DRIVER (2023) and AIDA (2020) projects. All of them are based on the use of Data Mining and Artificial Intelligence technologies to extract information of interest to law enforcement agencies. A less recent project with similar objectives is the e-POOLICE project. In it, we worked on a tool that allowed an effective and efficient exploration of raw and open information sources, developing an intelligent environmental radar that uses a knowledge repository to enrich the exploration. A key part of this process is a semantic filtering to identify data elements that may constitute weak signals of emerging organised crime threats, fully exploiting the concept of crime hubs, crime indicators and enablers as understood by its user partners. In this paper, we describe the system architecture employed in the COPKIT project, as well as a set of tools to fight crime using different data collection, enrichment, and extraction of hidden knowledge. These tools have been applied to the analysis of dark net data, supported by a knowledge repository module. This repository is capable of enriching the data to obtain better results when applying the different knowledge discovery (KD) tools. This system has great capabilities and encompasses many tools that allow it to be distinguished in several aspects from the projects discussed above. As can be seen in Table 1, our system allows the use of innovative tools using Big Data technology, as well as having a repository of knowledge that enriches the data and the results, thus allowing LEA users to obtain more interesting results. 3|SYSTEM ARCHITECTURE The developed system comprises different artificial intelligence technologies that are able to extract potentially new and useful knowledge in an automatic and unsupervised way, that is, no intelligence has to be given by the end-user in advance about the data that is being analysed. However, the knowledge repository module can collect the intelligence in advance and be re-utilised and updated by the different KD tools under the LEAs supervision. In this way, this system can be seen as the initial step in the investigation process to gather new information, which can be subsequently examined in order to discard irrelevant data. Therefore, the components included in the system are placed in the first step of the analysis procedure, and they help to describe the collected data offering new insights that cannot be inferred from a glance to the data or the volume of data is so big, that cannot be analysed manually. In any case, human intervention is necessary to avoid any automatic decision by the tool. In this regard, it is important to remember that the system assists the user in their posterior process of decision-making and does not make any decision by itself. According to the architecture depicted in Figure 1, it is assumed that certain kind of data has been previously crawled, such as, for instance, a number of advertisements in the dark web, or a network among users in some forums on the dark web. Typically, these collected datasets must be pre-processed before applying any discovery knowledge tools. Natural Language Processing tools such as Part of Speech and Entity Recognition can be used in this pre-processing stage. Additional data structuring processes are performed to organize data in a structured way, such as a table or an Entity-Relation database. Once the datasets have been processed, they are stored in a Knowledge Repository (see Section 4), where they are enriched with a priori and/or expert knowledge. This enhanced knowledge representation is used by the components of the KD tools to empower the semantic capabilities of the mining processes. The output of these processes assists the LEAs in the discovery of new relationships and insights to assess the situation. In the next section, the Knowledge Repository and the Knowledge Discovery components are explained in detail. TABLE 1 Comparison of main features of different EU-financed projects for the fight against crime. Name Data collection Big Data AI tools Data enrichment Early detection Statistic tools HMIs/visualisation tools COPKIT ✓✓✓✓ ✓✓✓ TENSOR (2019)✓OOO ✓✓O e-POOLICE OOO✓O✓O CC-DRIVER (2023)OO✓O✓✓O AIDA (2020)O✓✓OO✓O ANITA (2021)O✓✓OO✓O TITANIUM (2020)✓OOO O ✓✓ i-LEAD project (2023)✓OOO O O ✓ MAGNETO (2021)✓O✓✓ OO✓ FERNANDEZ-BASSO ET AL.3of16 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 4|KNOWLEDGE REPOSITORY COMPONENT This section introduces the main concepts and functionalities of the KRC (Knowledge Repository Component). The KRC aims to represent and manage expert and learnt information relevant to the Early Warning/Early Action ecosystem. 4.1 |The concept of the KRC The KRC implements a formal model in the form of an ontology for knowledge storage and distribution, improving the capabilities of the LEAs at the investigative and strategic level by supporting analysts involved in information management, offering additional metadata for the contextualisation of LEAs investigations. The KRC serves several functions: (1) it provides the knowledge that drives the execution of some technical components (e.g., starting points for the crawler); (2) it enriches input data used by another component with a priori and previously obtained knowledge; (3) it acts as a storage repository for the outputs of other components. 4.2 |Semantic data models The KRC is defined using the RDF (resource description framework) and the OWL (ontology web language) languages. These languages are standards proposed by the W3C (world wide web consortium)–in the context of the Semantic Web initiative–that respectively allow representing semi-structured data and logical restrictions in a flexible and normalized way. RDF (resource description framework) is the W3C standard language to describe resources in the Semantic Web Klyne and Carroll (2004). RDF allows metadata to be asserted in the form of triples, that is, statements relating an object, a property, and a value. For instance, it can be stated that [John] (subject) [has email address] (predicate) [[email protected]] (object). Subjects, predicates and objects are identified by their URI, a generalization of URLs (uniform resource locator) with a similar structure, but which can be used to identify things or concepts in an unambiguous way without necessarily returning an electronic representation of them as it happens with web pages. Objects can also be literals, that is, strings of characters that can have a type, like a number or a date, or be untyped, like in free text. RDFS (RDF schema) defines an RDF vocabulary that can be used to express logical relations between resources Manola et al. (2014). For instance, a resource can be declared as a class and another as an instance of that class, or a class can subsume another class; for example, citizens FIGURE 1 Architecture of the system. 4of16 FERNANDEZ-BASSO ET AL. 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License are human beings, and cocaine is a drug, respectively. OWL extends RDFS with additional vocabulary to express more complex logical statements and inferring new facts (W3C OWL Working Group, 2012). SPARQL (Harris & Seaborne, 2013) is a W3C standard query language that allows retrieving information from an RDF triplestore. The most basic feature of SPARQL is specifying a set of triples where variables can appear as subjects, predicates or objects of any triple. Using the SELECT form in the query, the SPARQL engine will return either the set of mappings of variables to values that conduct the query inside the KB (i.e., the possible ways the KB can satisfy the query), or an empty set if no match occurs. 4.3 |Knowledge included in the KRC This section describes the knowledge included in the KRC. To address the system requirements, the KRC includes knowledge regarding: •Metadata to describe the characteristics of knowledge •General knowledge •Domain knowledge about each specific use case Metadata in the KRC is used to characterise the “chain of custody”of a piece of data. The most important metadata in the KRC is provenance, which is essential to support result traceability and trust assessment. By provenance, we mean which process has led to a conclusion and which actors (human or artificial) have been involved. LEAs and authorities commonly use this concept in criminal prosecution because they need to precisely identify the individuals who have provided information to solve a case. Similarly, as it is done in these situations, the KRC considers credibility, reliability and similar confidence assessment values of sources and processes that affect data quality and results validity. Provenance is represented in the KRC by using the standard publicly available PROV-O ontology (Belhajjame et al., 2012). The general knowledge in the KRC is composed of several knowledge bases imported from the Linked Open Data web. Specifically, we have considered the following sources: •DBPedia knowledge base, a structured graph extracted from Wikimedia data publicly available (Graua et al., 2008). •YAGO2, an open knowledge base automatically built from Wikipedia, GeoNames and WordNet (Hoffart et al., 2013). •GeoNames, a free geographical database containing all countries and over 11 million place names (Wick, 2015). •NUTS (Nomenclature of Territorial Units for Statistics), the RDF version of the classification defined by Eurostat office (Correndo & Shadbolt, 2013). The domain-specific knowledge in the KRC is obtained from expert end-users and according to the opportunities identified by the technical partners for enriching the information managed by each component in the system. A general methodology for knowledge acquisition has been developed to minimise the burden of LEAs, who are not expected to be able to directly formalize their expertise into the knowledge base. Instead, a simple knowledge acquisition process encompassing the following steps has been carried out: 1. Identification of knowledge sources. 2. Summarisation into a document, which is circulated among involved partners. 3. Formalisation into the knowledge base. 4. If not finished, go back to point 2. To refine the base document, additional requests can be placed to the LEAs; for example, pointers to external knowledge bases, case reports, informal taxonomies, and so forth. Figure 2depicts the firearms taxonomy for the firearms trafficking use case. The firearms taxonomy has been built with the support of the LEAs. Specifically, we create a knowledge model that will define firearms features concerning national laws to automatically identify the category and the regulations associated with a firearm in different countries. 5|KNOWLEDGE DISCOVERY COMPONENT This section is devoted to explaining the knowledge discovery module. As already mentioned in the system architecture, different methods, such as classification, clustering, association rules, etc., can be used in this module. In this work, we focus on a non-supervised automatic tool that discovers tendencies and relationships among the different objects and attributes that may appear in a data collection. In particular, this tool is FERNANDEZ-BASSO ET AL.5of16 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License described using, for exemplary purposes, a database example of a fictitious dataset consisting of advertisements from the dark net offering items related to firearms. 5.1 |Frequent itemsets and association rules discovery Association rules have been employed to discover meaningful and easy-to-interpret information that can be utilised for different scopes within the crime field (Cheng et al., 2019; Englin, 2015; Ruiz et al., 2014). This component aims to find the most frequent items in a structured dataset and some relationships between these items, measuring their frequency and accuracy. Formally, for a set of items I¼i1,i2,…,in fgand a set of transactions D¼t1,t2,…,tN fg, where each transaction may contain or not some of the items, an association rule is defined as the relation between two disjoint (X\Y¼;) itemsets X,Y⊆Iand is noted by X!Y. The itemset Xis often referred as the antecedent (or left-hand side of the rule) and Yas the consequent (or right-hand side of the rule). The most commonly used measures to extract frequent itemsets and association rules are the support and the confidence, defined as follows: •The support measures the frequency of appearance of an itemset in the database. SuppDXðÞ¼ jtiD:X⊆tij jDj:ð1Þ In particular, the support of an association rule is the support of the union of itemsets Xand Y: SuppDX!Y ðÞ ¼SuppDX[Y ðÞ ¼jtiD:X[YðÞ⊆tij jDj:ð2Þ In general, the most interesting association rules are those with a high support value. •The confidence of the rule X!Ymeasures the percentage of transactions that containing X, also contain Y. This is measured by the conditional probability of Ygiven Xas follows: FIGURE 2 Summarised view of the firearms taxonomy. 6of16 FERNANDEZ-BASSO ET AL. 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License ConfDX!Y ðÞ ¼SuppDX[Y ðÞ SuppDXðÞ :ð3Þ For example, in the case of having a database like the one depicted in 3 where transactions (rows) represent advertisements found in the dark net selling something related with a firearm (e.g. gun, ammunition, etc.) and columns are different types of attributes (e.g. firearm, location, price, nickname, webpage, etc.); the items, in this case, will be pairs of the form < attribute,value > like for instance < location,London >or <market,Alphabay >. In this way, in this kind of dataset, the extracted frequent itemsets and association rules will be relations among the locations and the arms, the day of the week, the market and so forth. For instance, if the following frequent itemsets were obtained: <firearm,Crossbow_80Lbs >,<location,Madrid >fg withSupp ¼0:13 this means that Crossbow_80Lbs and Madrid appear together in the 13% of transactions, that is, in the 13% of collected advertisements. If the following association rule were obtained: <firearm,Crossbow_80Lbs > ! <location,Madrid >,<market,Nucleus > withSupp ¼0:11,Conf ¼0:82 that means that when Crossbow_80Lbs appears in a transaction, it is more likely (82%) that the location of the arm is Madrid and it is sold in the Nucleus market, having the 11% of transactions supporting this (i.e., satisfying the three items at the same time). The problem of uncovering association rules (see Figure 4) is usually developed in two steps Agrawal et al. (1994): •Step 1: Frequent Itemset Mining (FIM). Finding all the itemsets above the minimum support threshold, called MinSupp. These itemsets are known as frequent itemsets. •Step 2: Association Rule Extraction (ARE). Using the frequent itemsets, association rules are discovered by imposing a minimum threshold for an assessment measure, such as confidence, called MinConf. In these steps, some pruning strategies can be applied in order to reduce the complexity of the algorithm. However, these strategies are not enough when the data to be processed scales in an exponential way. This usually happens when analysing social media data searching, for instance, for new insights related to some criminal activity. Therefore, there is a growing need to analyse these very large data sets, which with traditional association rule mining algorithms, often lead to memory overflow errors or extend the processing up to several days. In the proposed system, we have employed new implementations based on the MapReduce paradigm to extract frequent itemsets and association rules in a more efficient way. To achieve this, this component has been developed using the Spark framework, which enables a distributed computation of data and avoids the memory overflow problems of classic frequent itemset and association rule algorithms available in the literature (see Fernandez-Basso et al., 2023;Fernandez-Bassoetal.,2016; Fernandez-Basso, Ruiz, et al., 2019 for more details). In particular, we have used the Apriori-TID proposal using MapReduce functions in Spark This algorithm enables in-memory computations and obtains a complete set of association rules very fast (more details about the advantages of this algorithm on a MapReduce paradigm can be found in Fernandez-Basso et al. (2023)). Once the association rules are obtained, they are stored and displayed using the following visualisation module. 5.1.1 | Visualisation of results This component also incorporates a module for storing and visualising the obtained results. This module transforms the obtained frequent itemsets and association rules into an intermediate form (Fernandez-Basso, Ruiz, et al., 2019). For this, we proposed a custom JSON representation of association rules through which the results can be visualised using a wide spectrum of libraries and methods, including graphical and interactive matrices or graphs. This module is of great importance when end users have to inspect the results, thus facilitating the review process to find new knowledge that may be of interest in the case under study. Some of these visualisations are explained in the following use case. FERNANDEZ-BASSO ET AL.7of16 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License 6|USE CASE: FIREARMS TRAFFICKING IN DARK NET ADVERTISEMENTS This section contains an overview of the system's performance through its different components, showing how it works in a real use case. In particular, we describe an example that analyses dark net advertisements offering firearms, their components or ammunition for sale. It is worth mentioning that sensible data that cannot assure privacy rights have been conveniently removed or anonymised to comply with the General Data Protection Regulation. In Figure 5, we have depicted a possible workflow that can be followed in our system. This flow is in line with the proposed architecture (Section 3) for the system. In our particular case, we are going to focus on the two presented components that are applied to a set of dark net processed advertisements. For instance, in Figure 6, there is an example of the data type that can be conveniently extracted from an advertisement about a firearm. To illustrate the performance of the components, we will use a pre-formatted data set that fulfils the specifications of our problem, whose structure is explained in the next section. However, it should be noted that the Knowledge Repository Component can be conveniently employed to enhance the NLP processing, for example, by feeding the classification algorithms for named entity extraction with new terms or slang employed for firearms, neighbourhoods or geographical places, or any other type of knowledge that may exist in the Knowledge Repository. Although this possibility exists, we are going to describe other applications of the KRC to enrich the information obtained after a knowledge discovery process. An example of the pre-processed and transformed database can be seen in Figure 3. 6.1 |Data The functionality of the different components will be described using a dataset called “Grams”which is a structured dataset containing a collection of advertisements from various dark net markets. It comprises information about the market name (i.e., the name of the market where the advertisement was posted), the vendor name (i.e., the username of the user who posted the advertisement), the name of the item being sold (usually the title of the advertisement), the country from which the advertised item will be shipped, the time at which the advertisement was published, some keywords, some columns containing internal identifiers and information obtained by applying automatic classification techniques to classify the type of firearm offered in the ad Heistracher et al. (2020). An example of the information contained in each of the rows of the dataset can be seen in Figure 3. We have also conveniently disaggregated the time attribute (when the ad was posted) into more meaningful attributes giving the month, day of the week, and whether it was posted during the day or at night. At the end, the resulted database has the following attributes (columns) which are used in the subsequent Knowledge Discovery process: Market, Vendor, SoldItem, ShipFrom, Keywords, ArmType, DateTime, Category, Month, DayWeek and DayTime. Additionally, text fields such as description have been processed by extracting keywords. These terms and other characteristics such as location, were enriched using the Knowledge Repository component in order to have the same granularity level in all the fields of the location-related columns such as ShipFrom. Finally, the database has to be transformed into a transactional dataset in order to apply the Frequent Itemset and Association Rule mining component. For this purpose, transactions with items using the format attribute_value were created in each of the fields. 6.2 |Application of the frequent itemset and association rule component Following the workflow, the component for the extraction of frequent itemsets and association rules 2019 has been applied. This component allows the extraction of hidden relationships in the data, for which a database of transactions obtained from the processes explained above must first be obtained. The example of weapons advertising data in the dark web is followed by an example to illustrate the usability and understanding of the component. FIGURE 3 Example of a transactional database with some advertisements. 8of16 FERNANDEZ-BASSO ET AL. 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License FIGURE 4 Association rule mining process. FIGURE 5 Use case workflow. FERNANDEZ-BASSO ET AL.9of16 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License AUTHOR BIOGRAPHIES Carlos Fernandez-Basso received the degree in computer science, the M.Sc. degree in data science, and the Ph.D. degree in computer science from the University of Granada, Granada, Spain, in 2014, 2015, and 2020, respectively. He is currently a Postdoctoral Fellow with Causal Cognition Lab, University College London, London, UK. He was a Lead Developer in the EU FP7 Project Energy IN TIME in the topics of building simulation and control, data analytics, and machine learning, and in the COPKIT Project in the topics of cybercrime, Big Data, and machine learning. From 2016 to 2018, he collaborated with the Data Science Institute, Imperial College London, London, UK, where he has carried out. He is now a postdoctoral researcher at University College London applying artificial intelligence in social contexts and researching explain ability in artificial intelligence (XAI). Karel Gutiérrez Batista was born in 1984. He received a degree in computer science and M.Sc. degree in data science from the University of Camagüey, Cuba. He received his PhD in computer science in 2018 from the University of Granada, Spain. He works as a postdoc fellow at the Department of Computer Science at the University of Granada. He is an Associated Research of Intelligent Data Bases and Information Systems (IDBIS) research group at the University of Granada. His research interests are Multidimensional Data Analysis, Deep Learning, Data Mining, Knowledge graphs, and Natural Language Processing. Juan Gómez‐Romero received the B.Sc. degree in computer science and the M.Sc. and Ph.D. degrees from the University of Granada, Granada, Spain, in 2004, 2006, and 2008, respectively. He was a Lecturer with the Applied Artificial Intelligence Group, Universidad Carlos III de Madrid, Madrid, Spain, from 2008 to 2013, and a Research Associate in the EU FP7 Project Energy IN TIME with the University of Granada, from 2013 to 2017. Since 2019, he has been an Associate Professor with the Computer Science and Artificial Intelligence Department, Universidad de Granada. His research interests include machine learning for control optimization and simulation of power systems. He is currently the Principal Investigator of the projects PROFICIENT: Deep learning for energy‐efficient building control and DeepSim: Deep learning of building simulation models. M. Dolores Ruiz received the degree in mathematics and the European Ph.D. degree in computer science from the Universidad de Granada, in 2005 and 2010, respectively. She is a Lecturer at the Department of Computer Science and Artificial Intelligence at the University of Granada, Spain, since 2020. She has participated in more than ten projects, including the EU FP7 Projects ePOOLICE and Energy IN TIME, and the COPKIT H2020 project. Her research interests include data mining, information retrieval, energy efficiency, big data, correlation statistical measures, sentence quantification, and fuzzy sets theory. She has organized several special sessions about Data Mining in international conferences and was part of the organization committee of the FQAS’2013, SUM’2017 and FQAS’2023 conferences. She belongs to the Approximate Reasoning and Artificial Intelligence Research Group and the Cybersecurity Lab, Universidad de Granada. She is the Principal Investigator of several projects about federated mining and desinformation detection. Maria J. Martin‐Bautista received the Ph.D. degree. She has been a Full Professor with the Department of Computer Science and Artificial Intelligence, University of Granada, Spain, since 2018. She has supervised several Ph.D. thesis and published more than 100 papers in high impact international journals and conferences. She has participated in more than 20 research and development projects and has supervised several research technology transfers with companies. Her current research interests include recommender systems, intelligent information systems, big data analytics in data, text and web mining, knowledge representation, and uncertainty. She is a member of the Intelligent Data Bases and Information Systems (IDBIS) research group. Furthermore, she has served as a program committee member for several international conferences. How to cite this article: Fernandez-Basso, C., Gutiérrez-Batista, K., G omez-Romero, J., Ruiz, M. D., & Martin-Bautista, M. J. (2024). An AI knowledge-based system for police assistance in crime investigation. Expert Systems, e13524. https://doi.org/10.1111/exsy.13524 16 of 16 FERNANDEZ-BASSO ET AL. 14680394, 0, Downloaded from https://onlinelibrary.wiley.com/doi/10.1111/exsy.13524 by Universidad De Granada, Wiley Online Library on [15/02/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License