scieee AI-readable full text Open interactive document viewer

Market-based Higher Education Course Recommendation

Ana Isabel Neves Alves de Sousa

Abstract

-

Full text

FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO Market-based Higher Education Program Recommendation Ana Isabel Neves Alves de Sousa Mestrado Integrado em Engenharia Informática e Computação Supervisor: Professora Doutora Ana Cristina Costa Aguiar Second Supervisor: Professora Doutora Luísa Marina Moya Praça de Araújo Lima July 16, 2016 Market-based Higher Education Program Recommendation Ana Isabel Neves Alves de Sousa Mestrado Integrado em Engenharia Informática e Computação Approved in oral examination by the committee: Chair: Professora Doutora Ana Paula Cunha da Rocha External Examiner: Professor Doutor Bruno Emanuel da Graça Martins Supervisor: Professora Doutora Ana Cristina Costa Aguiar July 16, 2016 Abstract The choice of a higher education program is determinant in the future career. It needs to take into account several aspects, such as: future employability; whether the program fits the skills required by the desired job; and whether the program’s scope of depth and breadth fits the needs of the student. It is clear that this is an important decision and that it is worth to invest the time to gather information for improving the final outcome; however, to the best of our knowledge, there are no dedicated tools for this purpose. Since alumni represent the most trusted source of information about the career paths a certain program confers access to, we propose a recommendation system based on alumni data to address this problem. In this thesis, we aim to find out how to use information about alumni and job offers in order to recommend the most appropriate higher education programs to access a certain job. The recommender system has as a input the user’s desired career and as a output a ranking of higher education programs to take in order to pursue that career. This is achieved by performing a match between the skills required for the career and the skills that an education program confers to attending students. Since this goal fits into the recommender systems category, we present an overview on the state of the art of this research area, with special focus on collaborative filtering methods, in particular, on top-N recommendation tasks. We provide a survey of information sources — social networks with business purposes — aiming at identifying the most appropriate data sources for gathering information about alumni and job offers. We then retrieve the information available from the sources selected, analyze and clean it, building a data set with information about alumni, jobs and skills associated to both. The skills gathered allow the comparison between alumni and jobs through similarity measures. We explore possible measures for assessing the proximity of alumni profiles and the desired careers. Since the measures explored only contemplate binary features, we propose one for weighted features. We design and implement an user based recommendation prototype system with Computer Science as the target discipline, where we apply decision processes to establish the most suitable system parameters depending on the data set. We then propose an approach to trim the recommendation list based on the similarities between the desired career and its possible recommended programs. In summary, we achieve our goal by devising a methodology for recommending higher education programs for a specific job. We also provide an evaluation of our system with a group of test users. i ii Resumo A escolha de um curso de ensino superior é determinante na carreira futura de uma pessoa. Esta escolha tem de ter em conta vários aspetos, tais como: a futura empregabilidade do curso; se este fornece as competências fundamentais para o trabalho desejado; e se o âmbito do curso se adapta às necessidades do aluno. Assim sendo, torna-se claro que esta decisão é importante e que vale a pena investir tempo na recolha de informações que possam melhorar o resultado final; no entanto, de acordo com o nosso conhecimento, não existem ferramentas com esta finalidade. Tendo em conta que os antigos alunos de um curso representam a fonte mais confiável de informação sobre as saídas profissionais a que ele confere acesso, propomos um sistema de recomendação que tem como base informação acerca de antigos alunos. Nesta dissertação pretendemos descobrir como usar informação sobre antigos alunos e ofertas de emprego por forma a recomendar os cursos de ensino superior mais adequadas para aceder a um determinado trabalho. O sistema de recomendação tem como dados de entrada a carreira desejada do utilizador e, como dados de saída um ranking de cursos de ensino superior apropriados para essa carreira. Este ranking é realizado através do estabelecimento de uma correspondência entre as competências necessárias para a carreira e as competências que os cursos oferecem aos seus alunos. Uma vez que este objetivo se encaixa na categoria de sistemas de recomendação, apresentamos uma visão geral sobre o estado da arte desta área de investigação, com especial foco sobre os métodos de filtragem colaborativa, em particular, nas tarefas de recomendação dos N itens de topo. Fornecemos um levantamento sobre as fontes de informação - redes sociais com fins profissionais - com o objetivo de identificar quais as mais adequadas para extrair informação sobre antigos alunos e ofertas de emprego. Em seguida, recolhemos a informação disponível nas fontes selecionadas, analisamo-la e limpamo-la, construindo assim uma base de dados com informações sobre antigos alunos, empregos e competências associadas a ambos. As competências recolhidas permitem a comparação entre antigos alunos e empregos através do uso de medidas de similaridade. Assim, exploramos possíveis medidas para avaliar a proximidade entre os alunos e as carreiras desejados. Dado que as medidas exploradas apenas contemplam características binárias, propomos uma medida que contempla características ponderadas. Projetamos e implementamos um protótipo de sistema de recomendação com Ciência da Computação como disciplina base. Neste protótipo, aplicamos processos de decisão para estabelecer os parâmetros mais adequados para o sistema, que são dependentes da base de dados. Por fim, propomos uma abordagem para podar a lista de recomendações, que tem como base as semelhanças entre a carreira desejada e os diferentes cursos. Em síntese, alcançamos o nosso objetivo através da elaboração de uma metodologia para a recomendação de cursos de ensino superior para alcançar um trabalho em específico. Adicionalmente, fornecemos uma avaliação do nosso sistema feita por um grupo de utilizadores de teste. iii iv Acknowledgements A dissertation project might be an individual work, but it surely needs the right backup to achieve successful results. Therefore, I feel that I should recognise everyone that helped me throughout this journey. Firstly, I would like to thank Instituto de Telecomunicações for providing me the best working conditions to develop this work. Without their support, the challenges faced had not been possible to overcome in the same successful way. I would also like to thank my supervisors, Ana Aguiar and Luísa Lima, for all the guidance and input throughout the development of this dissertation. Their knowledge and experience were crucial to accomplish the goals established. Thank you for always pushing me to do more and better. Lastly, a special recognition to my family and close friends for their love, understanding and support during the last months. Ana Sousa v Chapter 1 Introduction Choosing a higher education program is not an easy decision for multiple reasons. There is a wide variety of higher education programs available, with important differences both in depth and breadth of lectured topics. The future employability of these programs varies; this fact is especially relevant concerning the current lack of detailed employability statistics. When a student enrolls in a degree, he usually does not know the full employability potential of a program. 1.1 Motivation For some students, it is straightforward to think about a dream job; however, the higher education program to take in order to achieve that goal is not always clear. Furthermore, the skills that an educational program claims to confer its students are not always the ones that the alumni feel to have acquired during the course of the program. As a consequence, having so many options can represent a real challenge. Knowing that this decision has a large impact in the future career makes it even more important to invest time to make a better-informed decision that takes into account the alumni experience. Since they have attended the program and felt the direct impact of their decision in their professional careers, it is reasonable to consider them as the most trusted source of relevant and real information about some of the possible paths that the program confers access to. Finally, the talent acquisition process for companies is not a simple process due to several factors. An important practical consequence of the work presented in this thesis could be to present indicators to companies of the best programs to find candidates, as well as the best programs to invest in for internal employee training and formation. 1 Introduction 1.2 Goals The main goal of this dissertation is to discover how to use alumni and job offers information to recommend higher education programs that are likely to provide the skills that will allow students to pursue a certain career. In order to accomplish our goal, we design a recommendation algorithm for higher education programs. Finally, we build a proof of concept system with Computer Science as the target discipline, in order to reduce the scope of recommendations and also to enable manual validation of the proposed system. The input of our recommendation system is the user’s desired career, and its output is a ranking of possible programs that he could take in order to be closer to his goal. This ranking is computed taking into account skills acquired by a set of alumni, the capabilities that every program provides and the ones required for that specific career. In order to build the recommender system, the sources of relevant information – social networks designed for business purposes – need to be identified and explored. Through these social networks, we are able to extract the required information about the alumni, such as their skills and education, as well as job offers and skills associated to them. The data collected from the information sources then needs to be analyzed and cleaned. Different recommendation algorithms and approaches need to be researched and experimented, in order to identify the most suitable approach for the problem in hands. Taking this into consideration, we analyze the influence of the different system parameters on the resulting recommendations, and identify the best option having in consideration our data set and the impact of recommendations on the users of the system. Finally, due to the subjective nature of a recommendation, we suggest to validate the proposed solution manually by distributing a questionnaire among 59 people. 1.3 Contributions With this thesis, we aim at contributing with: •A data set with alumni, jobs and skills associated to them; •A higher education program recommendation algorithm; •A metric to evaluate the similarity between items characterized by non-binary features; •A sensitive analysis on the recommender system parameters; •A prototype system to evaluate and validate the proposed solution. 2 Introduction 1.4 Dissertation Structure In addition to this introduction, this dissertation contains six other chapters. In Chapter 2, we describe the state of the art on recommender systems and present related work. In Chapter 3, we provide a more detailed description of the system and its goals. In Chapter 4, we describe our data set, namely the sources of information and how data was collected, analyzed and cleaned. In Chapter 5, we provide the details of the proposed algorithm, as well as present the full decision process taken to derive this methodology. In Chapter 6, we elaborate our proof of concept with the support of a web based system that allows user evaluation. Finally, in Chapter 7, we present the conclusions and future work. 3 Introduction 4 Chapter 2 Recommender Systems Nowadays – even though we may not always realize it – recommendations are ubiquitous. In real life, when someone wants to watch a movie, that person will probably ask one of his friends his opinion about it; when making a decision on purchasing a product, the most common method is to search for online reviews. Automated recommender systems are quite similar to these real-life recommendations, with the considerable advantage that we can have access to a much wider range of users who have the same preferences. Recommender systems have become an important and interesting area both in academia and in industry, with a wide range of work being done in this rapidly expanding area [AT05]. Some of the most well-known recommender systems in industry are Amazon1that recommends books, CDs and other products, MovieLens2that provides movies recommendations, and Netflix3that besides movies also recommends series. A typical recommender system is characterized by: •A set of users C; •An active user c– the one to whom the recommendation will be provided; •A set Sthat comprises all items that can be recommended. The main goal of these systems is to maximize the usefulness/utility of a particular item s∈S to the user c, which is represented by the utility function u(c, s). This utility is often represented as a rating that expresses the opinion of the user ctowards an item s∈S. Ratings can be either explicit – expressed as a value on a scale, or implicit – expressed as a purchase or as a click. Ratings given by users can be converted into a user-item matrix. In the user-item matrix in Table 2.1,Maria is the active user, who we want to provide a computed recommendation to. Andrea,John and Maria are users belonging to C, and X-Men,Cinderela,Batman and Sherlock 1Available at amazon.com 2Available at movielens.org 3Available at netflix.com 5 Recommender Systems X-Men Cinderela Batman Sherlock Holmes Andrea Like Dislike John Like Like Lucas Like Like Dislike Maria Like Like ? Table 2.1: An example of an user-item matrix. Holmes are items in the space of items S. For instance, the rating given by Andrea to the item X-Men was a Like, while her rating to the item Cinderela was a Dislike. In our problem, the active user is the one that requests the recommendation, while items to be recommended are higher education programs. We infer user ratings by using the skills that are common between the job that the active user likes and the ones that each user has. The main categories for recommender systems nowadays are [AT05]: •Content-based: Recommendations given to the user are based on items that he has liked in the past. •Collaborative: Recommendations given to the user are based on items liked by people who have similar preferences and tastes. •Hybrid approaches: Combination of content-based and collaborative recommender systems. Since we are interested in providing recommendations based on the users that are similar to the job that the active user likes, our problem belongs to the collaborative category. For the sake of completeness, we also provide a brief overview of content-based and hybrid approaches for recommendation. 2.1 Content-based Methods In content-based methods, recommendations are made by identifying which items are similar to the ones preferred by the active user in the past. More formally, it can be defined as the estimation of the utility u(c, s) of an item s∈Sfor user c, which is based on the utilities u(c, si)that chas given to items si∈Sthat are similar to the current item s[AT05]. A similarity measure is a distance measure between two users or items and is used as a weight. The more similar two users cand c’ are, the more weight c’ ratings will have on making recommendations for user c[AT05]. The same applies for items. Content-based methods make recommendations by analyzing information and finding patterns within it. To make predictions, they rely on the user and items features that are extracted from textual information [SK09]. Some of the limitations of these methods are: limited content analysis (the features associated with the objects recommended have to be in a format that can be automatically parsed by a computer such as text, otherwise they have to be assigned manually), overspecialization (user cwill 6 Recommender Systems only be recommended items similar to those he liked in the past) and the new user problem (a user chas to rate a significant number of items before the system understands his preferences). More details can be found in [AT05]. 2.2 Collaborative Methods Collaborative recommender systems, also known as collaborative filtering systems, make recommendations to a new user based on the known preferences of a group of users similar to him. The fundamental assumption of such systems is that if two users ciand cjrate nitems similarly, it is reasonable to assume that they will also act similarly towards other items [SK09]. More formally, the utility u(c, s) of item s∈Sfor the active user cis estimated based on the utilities u(cj, s) assigned to item sby those users cj∈Cwho have similar preferences to c[AT05]. Collaborative methods have the ability to filter any type of content (such as text, artwork or music) because the recommendation process is based on data from other users, whereas contentbased methods do not have that ability since they rely solely on the history of the user for each type of content. The main advantage of collaborative methods over content-based methods is perhaps the fact that they do not depend on error-prone machine analysis of content [HKR00]. In the problem we address, the active user chooses his preferred job, which has an associated set of skills. That set of skills can be acquired by frequenting certain educational programs. Thereby, the recommender system finds users that have similar preferences to the active user, which are expressed in the form of skills. After that, only the most common education programs among those users are recommended. According to [SK09], algorithms that perform collaborative filtering fit into one of three main classes: memory-based, model-based or hybrid, which are described next. 2.2.1 Memory-based Memory-based algorithms make their predictions based on the entire collection of items already rated by users. Every user is considered to be part of a group of people with similar interests and preferences, who are often referred to as the nearest-neighbors of the active user. A prevalent memory-based collaborative filtering algorithm is neighborhood-based algorithm, which comprises the following steps [SK09]: 1. Calculate the similarity (or weight) between two users or items; 2. If we want to recommend a ranking of items rather than just a single item, then it is a top-N recommendation task and we need to identify the kmost similar users or items; 3. Produce a prediction for the active user. The algorithm steps are described in detail in the next section. 7 Recommender Systems 2.2.1.1 Similarity Computation The similarity computation is a critical step in memory-based collaborative filtering algorithms and can be performed between items or users. In the case of item-based algorithms, only users who have rated both items si∈Sand sj∈Sare considered. Thereafter, the similarity computation is performed to determine the similarity wi,jbetween the two items [SK09]. For user-based algorithms, the similarity wu,vis calculated between users cu∈Cand cv∈Cwho have both rated the same items [SK09]. There are several methods to compute the similarity between users or items. The application of the correct similarity measures results in a more accurate data analysis. There is not a measure that works well for all implementations, so various have to be experimented in order to identify the one that fits best to the specific problem and context. [SSSHT10] presents 76 binary similarity and distance measures that were collected and analyzed. Hierarchical clustering was performed to estimate the similarity between measures. Two of the most used similarity measures are described next; other measures can be found in [SSSHT10]. Correlation-Based Similarity – The similarity is measured by computing the Pearson correlation or other correlation-based similarities. Pearson correlation measures the extent to which two variables linearly correlate with each other [RIS+94]. Vector Cosine-Based Similarity – The similarity between two items is calculated by treating each item as a vector of ratings and then computing the cosine angle formed by the vectors [SM86]. Since similarity computation is a critical step, there has been active research on improving these measures. [JSB11] presents a probabilistic definition of item similarity. A co-occurrence between two items si∈Sand sj∈Sis defined as the number of users that have liked both items. Item similarity is defined as the ratio between the actual number of co-occurrences and the number of co-occurrences that would happen if user choices were random. It then applies the user’s usage history of the system to the item similarity matrix. It presents experiences with real-world usage data from different data sets to access the algorithm quality. Its quality was measured against several well-known algorithms and it tied for first place. In [CAL] the accuracy of the Dice similarity measure is improved by processing existing correlations between the characterizing item features. 2.2.1.2 Top-N Recommendations The purpose of the top-N recommendation algorithm is to recommend a set of N top-ranked items that will be interesting for the active user. It tries to discover relations between different users or items and uses those relationships to compute the recommendations. Top-N recommendations can also be achieved with model-based collaborative filtering approaches such as association rule mining based models [SK09]. [DK04] proposes a item and model-based algorithm that first, computes the similarity between the items, and then combines these similarities in order to compute the similarity between a set of items and a possible item to recommended. 8 Recommender Systems User-based Top-N Recommendation Algorithms – Typically, these algorithms comprise the following steps [SK09]: 1. Identify the kusers that are the most similar to the active user using a similarity measure; 2. Identify the set of items U⊂Sliked by those users and their frequency; 3. Recommend the top-N most frequent items in Uthat the active user has not rated yet. User-based top-N recommendation algorithms have some limitations regarding scalability and real-time performance mostly because of the similarity computation. As the number of users grows, the number of comparisons grows too and that can affect the performance of the recommender system. Item-based Top-N Recommendation Algorithm – Item-based algorithms address the scalability problem of the user-based ones. Usually they comprise the following steps [SK09]: 1. Compute the kmost similar items for each item in the set of items Susing a similarity measure; 2. Identify the set of items Vby taking the union of the kmost similar items and removing the set of items Uthat the user already liked; 3. Compute the similarities between items in Vand U; 4. Sort Vby decreasing order of similarity to obtain the list of Top-N recommended items. We are interested in providing the active user with top-N recommendations rather than a single recommendation, allowing him to choose the one that is most suitable to achieve his desired job. 2.2.1.3 Prediction and Recommendation Computation In this step, the predictions or recommendations are obtained. In the neighborhood-based collaborative filtering algorithm, a subset of the nearest neighbors of the active user is chosen based on the similarities with the user. After that, a weighted aggregate of the neighbors ratings is generated in order to make predictions for the active user [cKBR99]. 2.2.2 Model-based Model-based approaches have been developed to overcome some of the shortcomings of memorybased approaches and to achieve better performance. These approaches use the collection of already rated items to learn a model, which can be a data mining or machine learning algorithm that is used to predict item’s rating. Among the most well-known techniques, we can highlight Bayesian belief net models, clustering models and latent semantic models [SK09]. 9 Higher Education Program Recommendation Figure 3.1: Logic architecture. Full lines represent information that we have. Dashed lines represent relationships we need to establish. A job offer has a job position and a set of skills associated to it. Each alumni has a set of skills and has taken one or more higher education programs. •through job offers, we collect a set of job positions and skills associated to those jobs; •through alumni profiles we collect a set of higher education programs and the skills they provide to students. Through skills, we can establish match job positions and alumni. By doing this, it is possible to infer the real skills that a certain higher education program provides to its students, instead of just skills it claims to provide. Furthermore, using this strategy, we are also able to collect the required skills for a wide range of jobs in Computer Science. Figure 3.2 presents a general overview of the solution to the problem of how to use alumni and job offers information to recommend higher education programs that are appropriate to achieve a certain job. It has the following steps: 1. The user chooses a Computer Science related job/career from a closed set of options; 2. The desired career has a set of required skills – we identify these via job offers; 3. The required skills can be obtained by taking one or more higher education programs – alumni information allows us to do this match; 4. We then recommend a set of higher education programs that can provide those skills to the user. 16 Higher Education Program Recommendation Figure 3.2: Solution. The input is the user’s desired job. That job has a set of skills required, than can be acquired in a several programs. The output is the ranking of programs recommended to take in order to be closer to achieve that job. 3.1 Methodology We adopt an iterative methodology by experimenting and observing the results. We pick a strategy and observe the results for that strategy. Then, we make a change or addition, and evaluate whether the results improve. We perform these steps until we are satisfied with the obtained results. This approach assumes special importance in the context of our recommendation system: since there is no solution that works for every problem, we iteratively search and experiment several approaches in order to determine which is the most suitable methodology for our problem. 3.2 Technologies We use Ruby2language to gather data from the data sources and to implement the recommender system. For simpler and initial data analysis we use Excel3. To perform more complex graphical analysis, we move on to R4and Python5. We take advantage of the Plotly6and matplotlib7Python libraries to draw complex graphs. 2https://www.ruby-lang.org/en/ 3https://products.office.com/excel 4https://www.r-project.org/ 5https://www.python.org/ 6https://plot.ly/python/ 7http://matplotlib.org/ 17 Higher Education Program Recommendation To store the data, we choose an relational database, namely PostgreSQL8. 8http://www.postgresql.org/ 18 Chapter 4 Building a Data Set The data set is one of the most crucial aspects of a recommender system, since it is the source of users and items for generating recommendations. In order to build a data set with jobs, educational programs and the corresponding skills that are suitable as a base for our recommendation system, it is necessary to search for professional social networks that can provide this data. 4.1 Social Networks A social network can be defined as "a network of social interactions and personal relationships" and "a dedicated website or other application which enables users to communicate with each other by posting information, comments, messages and images" [oxf16]. Social networking can be done for social, business purposes or both. Our focus is on social networks for business purposes. This type of networks allows their users to establish and document networks of people they know and trust professionally. Usually, a person’s profile on one of these networks can be seen as an online version of his curriculum vitae, because it has information regarding his education, work experience, projects he participated in, competences acquired, language skills, awards and publications. In addition, companies also have profiles with detailed information about them and job openings available at the moment. After an extensive search, the following networks were selected for further exploration: ITJobs1,Jobvite2,Landing.jobs3, LinkedIn,Stack Overflow Careers4,Upwork5and Xing6. We are interested in gathering information about alumni – their education, job experience and skills, and about job offers and the skills needed for it. In order to extract these data from the social networks, the Application Program Interface (API) that that each social network provides 1https://www.itjobs.pt/ 2http://www.jobvite.com/ 3https://landing.jobs/ 4http://careers.stackoverflow.com/ 5https://www.upwork.com/ 6https://www.xing.com/ 19 Building a Data Set is explored to assess which jobs, education and the corresponding skills it is possible to gather. In the following sections, we describe the explored APIs in detail. 4.1.1 Landing.jobs Landing.jobs defines itself as "a candidate-driven tech jobs marketplace" and has a more corporate employment focus. It allows the matchmaking between candidates and jobs. Companies can post job offers and candidates can apply to those offers. Through the Landing.jobs API, it is not possible to search or list candidates, since it only allows access to the current user data, which does not include his education. However, it offers an endpoint for job listing where each job listed has title and a set of tags, which are core competences needed for that job. 4.1.2 Stack Overflow Careers Stack Overflow Careers presents itself as a solution for companies to search and find developers that are suitable for the job offers they have. Candidates are able to prove their knowledge trough the successful resolution of other people’s questions in Stack Overflow7. It provides a paid candidate search, where it is possible to filter candidates and access their personal information such as education, education skills, job experience, job skills and skills acquired. In addition, it also offers a job search functionality using keywords and a location. Each resulting job has a title and skills required for it. 4.1.3 Upwork Upwork defines itself as "an online workplace for the world – connecting clients with top freelance professionals" and is focused on freelancer workers. It offers an endpoint to search for freelancers and another to search for jobs. Freelancers search was discarded because it did not provide their education and employment history, even though these details are present in the web view of an user’s profile. On the other hand, each job returned on the job search has a title, category and subcategory, as well as a set of skills required or desired for that job. Its category and subcategory represent competence or expertise areas. 4.1.4 Other Platforms ITJobs is an online platform where companies can post jobs and candidates can apply and get hired for those jobs. Despite offering an endpoint for job search, each job does not have associated skills. Jobvite has the same purpose as ITJobs and it offers an endpoint for candidate search and one for job listing. The candidate search was discarded because it was not possible to access the 7http://stackoverflow.com/ 20 Building a Data Set Alumni and Skills Jobs and Skills Landing.jobs x Stack Overflow Careers x x Upwork x Table 4.1: Data sources summary. Crosses represent data that is present in each API. candidate’s education or his skills. The job listing was discarded because each job did not have a set of skills associated. On LinkedIn, people can post professional information about themselves such as, for example, education taken, skills and current and past job positions. LinkedIn only provides access to its people search API via the LinkedIn Partners Program; several applications were made to the partner program describing the project and all of them were rejected. The use of ITJobs,Jobvite and LinkedIn was discarded since their APIs did not provide the information we are interested in. Xing is "the largest business network in German-speaking countries". It offers an endpoint to search for people using various filters. Each person found has a set of core competences and an educational background. After experimenting the endpoint and analyzing the data returned, Xing API was discarded because it was only possible to access the education and skills of 15 people, possibly due to the users privacy settings. 4.1.5 Analysis of Available Information Table 4.1 summarizes the information possible to extract from the business networks selected after exploring and experimenting their APIs. As it is possible to conclude, most of the business networks explored do not provide the information we are interested in, due to privacy settings, limited access or APIs limitations. Nevertheless, we are able to gather information on alumni education and their skills through Stack Overflow Careers and, on job offers and skills required for it through Landing.jobs,Stack Overflow Careers and Upwork. By collecting skills for both alumni and jobs, we are able to compare them through the set of skills they are associated with. Therefore, in our data model, skills establish the connection and the possibility of matchmaking between alumni and jobs. 4.2 Data Model For our problem, there are two main perspectives on the data set represented by the model in Figure 4.1: jobs and alumni. Each job entry represents a job offer and has a title, a category, a subcategory and a set of skills associated. The category and subcategory can be viewed as an ontology, where a category represents a broader area of expertise and its subcategories represent more specific areas within it. 21 Building a Data Set Figure 4.1: Data model with jobs, alumni, educations, alumni jobs, skills, job categories and subcategories. This categorization allows the grouping of jobs with similar areas of expertise and, since subcategories are more specific than categories, jobs that belong to the same subcategory can be grouped together. Each alumni entry represents a result of the Stack Overflow Careers candidate search. It has a location, a profile identifier, a set of skills, a set of jobs and a set of educations. Each alumni education has a degree, a location, as well as a set of skills that represent the competences that the degree provides. Each alumni job has a title and a set of skills that he considers to have acquired in it. 4.3 Data Gathering The data was gathered from the social networks described in the previous section by following the steps of the flow diagram in Figure 4.3. In step 2, we only add a category and a subcategory to already collected jobs if the job title is a match between Landing.jobs and Upwork. The job search in step 3is done using titles of already collected jobs, as well as categories names, subcategories names and skills names. The complement of categories and subcategories in step 5, is only performed if the Stack Overflow Careers job title is contained in one of the Upwork search results. In our system, the active user needs to choose a specific job or career; however, choosing that job from a list with thousands of jobs would not be convenient for the user. Consequently, we group jobs by expertise areas. This grouping could be achieved by performing clustering by 22 Building a Data Set Figure 4.2: Flow diagram presenting the steps taken in order to gather information about jobs and alumni. skills associated to each job. However, clustering by skills would probably be difficult because each job has a low number of skills associated with it, as shown in Figure 4.3, making it harder to distinguish different jobs and separating them into clear expertise areas. Since Upwork provides categories and subcategories for its jobs, we decided to take advantage of it. The jobs that are left uncategorized at the end of the data gathering can be classified using the categories already created, by applying rules derived from the data analysis. 4.4 Data Analysis The immediate step after gathering the data is to perform an exploratory data analysis over the data set to become familiarized with it, understand the data distribution and identify possible rules to apply to clean it. The performed analysis has four main focus: •Jobs; •Alumni; •Educations; •Job, education and alumni skills. To understand the distribution of skills per job, we compute the frequency of skills. It is possible to observe in the histogram in Figure 4.3 that the majority of the jobs has between 1 and 10 skills, the most common number of skills associated to a job is 5 and only a small portion has more than 10 skills. This likely means that companies only specify the core skills needed for the job in each job offer. 23 Building a Data Set Figure 4.3: Histogram representing the number of skills per job. In order to identify the available categories and how jobs are distributed among them, we compute the number of jobs in each category, as it is possible to see in Figure 4.4. The available categories are: •accounting & consulting; •administration support; •customer service; •data science & analytics; •design & creative; •engineering & architecture; •it & networking; •legal; •sales & marketing; •translation; •web, mobile & software development; •writing. 24 Building a Data Set Figure 4.4: Histogram representing the number of jobs per category. Appendix Apresents the associations between categories and subcategories, and the job distribution within subcategories. Almost 25% of jobs belong to the web, mobile & software development category, which can be viewed as a logical consequence of the target discipline being Computer Science and of the popularity of the web field nowadays. More than half of the jobs do not have a category and subcategory associated, which implies that in the data cleaning phase we focus our effort into categorizing them. In Figure 4.5 it is possible to see the alumni distribution by country. Most alumni are currently located in North America, India, Canada, Russia or United Kingdom. Having this in consideration, it is expected that the educations locations of our recommendations are focused on these countries. Regarding educations, we start by analyzing the frequency of each education, which is represented by a degree and a location. The frequency of an education corresponds to the number of alumni that has that educational program associated. There was no education that had a frequency higher than 50 in more than 150000 educations entries available as shown in Figure 4.6, which could hamper the recommendation of educational programs to users. Since the data set has educations with the same meaning but with different names, we encounter the synonymy problem that was pointed out in Section 2.2.3.2. This is a consequence of the fields for inserting an education degree and location being free text and not an option selected from a closed set, which means that the alumni have freedom to use their preferred nomenclature 25 Building a Data Set Figure 4.12: Histogram of the number of alumni that share a certain education after the data cleaning process. respectively. It is evident that, by cleaning data, we are able to increase these numbers and the recommender system will benefit from this, as it is going to be possible to make better recommendations. It is also important to highlight the final distribution of skills per educations and alumni in Figures 4.13 and 4.14, respectively. Most educations have up to 7 skills associated, while most alumni have 5. In addition, the number of entries with more skills decreases as the number of skills increases, which most likely means that only the core skills are specified in most cases. 32 Building a Data Set Figure 4.13: Histogram of the number of skills per education after the data cleaning process. Figure 4.14: Histogram of the number of skills per alumni after the data cleaning process. 33 Building a Data Set 34 Chapter 5 Designing the Solution To design the solution for the problem of how to use alumni and job posts information to recommend higher education programs, we used a top-N recommendation approach as described in Section 2.2.1.2; in particular, within this category of recommendation algorithms, we picked a user-based approach. The overall algorithm is represented in Algorithm 1. Algorithm 1 User-based top-N algorithm for recommending higher education programs. 1: skill_set ←job_skills ∩alumni_skills 2: job_matrix ←build_job_matrix(skill_set) 3: alumni_matrix ←build_alumni_matrix(skill_set) 4: while continue do 5: Input: subcategory 6: job_desired ←get_element(job_matrix,subcategory) 7: top_k_users ←f ind_top_k_most_similar(job_desired,alumni_matrix) 8: top_n_educations ←find_top_n_educations(top_k_users) 9: Output: top_n_educations The features that characterize jobs and alumni are their skills, creating a categorical feature space. We start by identifying, in line 1, the set of skills that are present both in jobs and alumni. This skill set will be the base for the recommender system, since it will allow the comparison between jobs and alumni. In lines 2and 3we build the user-item matrices for jobs and alumni, named job matrix and alumni matrix respectively, using the skill set created in line 1. More details on the purpose of these matrices and how they were built are provided in Chapter 2and Section 5.1, respectively. The user then chooses a subcategory from the ones available and in line 6, the corresponding job row is selected from the job matrix and used in line 7to find the kmost similar users by applying a similarity measure. More details on how kis chosen and what similarity measure was adopted can be found in Sections 5.2 and 5.3, respectively. Finally, in line 8, we take the kmost 35 Designing the Solution subcategory_id css html5 javascript git redis tdd ios-sdk ... 1 0.331 0.297 0.508 0.059 0.0 0.03 0.0 2 0.0 0.043 0.0 0.0 0.0 0.0 0.022 ... Table 5.1: Extract from the job matrix. similar users to identify the Nmost frequent educations among them that will be presented to the user. Details on how the value of Nwas determinate can be found in Section 6.2. 5.1 User-item Matrix Regarding jobs, the user-item matrix represents the relationship between each job posting and its associated skills. In that way, we can establish a match between the user-item and the job matrices: users correspond to jobs and items to skills. Each job is represented by a row in the matrix and each skill by a column. If the skill is present in a particular job, then its correspondent column value is 1, otherwise it is 0. Since we are not interested in particular jobs, but in their broader areas of expertise that are represented by their subcategories, we transform the job matrix into a subcategory matrix. We compute the weight of each skill per subcategory by dividing the frequency of each skill in the jobs that belong to it by the number of jobs in it, as shown in (5.1). This results in a weighted subcategory matrix, where each skill has a value ranging from 0 up to 1, rather than just 0 or 1. If a skill appears in almost all the jobs of a certain subcategory, it will have a weight close to 1. Otherwise, if it only appears in a small set, its weight will be close to 0. Skill weight in a subcategory =f requency o f skill in the subcategory jobs number o f jobs in the subcategory (5.1) Table 5.1 shows an extract of the resulting subcategory matrix. The first column is the subcategory identifier, the following columns are skills and its weight in each subcategory. For example, subcategory 1gives a weight of 0.331 to the skill css and 0.297 to the skill html5. Concerning the alumni matrix, we can match it with the user-item matrix by considering alumni to be users and skills to be items. We only consider alumni that have a set of associated skills and, for each one, we identify his skills. If one skill is associated to a particular alumni, then its correspondent column value is 1, otherwise it is 0. Table 5.2 shows an extract of the result ing alumni matrix, where the first column is the alumni identifier and skills weights are binary, which means they have either the value 0 or 1. 5.2 Similarity measure We utilize a similarity measure to assess who are the alumni with the most similar skills to one particular subcategory, by comparing the skills of the subcategory chosen to every alumni. The 36 Designing the Solution alumni_id css html5 javascript git redis tdd ios-sdk ... 1 1 1 1 0 0 0 0 2 0 0 1 1 0 0 1 ... Table 5.2: Extract from the alumni matrix. main point that similarity measures consider for categorical features and weighted categorical features like the ones that we have in our problem is the count of matches or mismatches between the features of two items. The count of matches and mismatches is represented in Table 5.3. According to Table 5.3, item 1 is the subcategory and item 2 is one of the alumni. The variable acorresponds to the number of skills that are present both in the subcategory and in the user, bto the number of skills only present in the subcategory, cto the number of skills only present in the alumni and dthe number of skills absent on both. The similarity measure choice has a large impact on the recommendations made by a recommender system. Because of this, it is important to choose the one that seems to be the most suitable for the problem at hand. Taking our problem into consideration, we establish that the similarity measure should comply with the following requirements: •give more weight to asince we are interested in highlighting the skills that are present in both items; •do not take into account dbecause we are more interested in positive matches than in negative ones; •do not include N, which is the number of skills, because the measure values would increase rapidly. After analyzing the binary similarity measures presented in [SSSHT10], we came up with a list of the more interesting ones accordingly to our requirements, namely highlighting aand not taking into account dor N: SDICE =2a 2a+b+c(5.2) S3W−JACCARD =3a 3a+b+c(5.3) Item 2 1 (presence) 0 (absence) Item 1 1 (presence) a b 0 (absence) c d Table 5.3: Contingency table 37 Designing the Solution SMCCONNAUGHEY =a2−bc (a+b)(a+c)(5.4) SSORGENFREI =a2 (a+b)(a+c)(5.5) To decide which similarity measure is most suitable to the problem, we draw 4d graphs with the following components: •x-axis: number of skills that a subcategory can have; •y-axis: number of skills that an alumni can have; •z-axis: number of skills in common between the last two, it corresponds to a; •colour: similarity value, it varies from 0 to 1 or from −1 to 1 depending on the measure used. Figure 5.1 presents the resulting graphs. We analyze the graphs focusing on the color variation, since all other variables are constant between graphs. We were interested in having colors that represent high values of similarity when the number of skills in common is close to the number of skills of the subcategory and the alumni. Apart from this, we want to have a good spectrum of colors, which means that the measure has high variability, adapts well to each specific case and, consequently, is good at discriminating subcategories. The Sorgenfrei (5.5) measure is not suitable for our case since the color almost does not vary between different points in the graph Figure 5.1d. The Dice (5.2), 3w-Jaccard (5.3) and Mcconnaughey (5.4) measures seem to accomplish our goals regarding color variation and high similarity values. We needed to choose one of these three measures and opted for the 3w-Jaccard (5.3) measure, since it is the one that appears to have a wider range of values while allowing, at the same time, a clear distinction between different points in the graph in Figure 5.1b. 5.2.1 Incorporating weights into the similarity measure The similarity measures analyzed in Section 5.2 are used to compute the similarity between items with binary features. In the case of jobs, our values range from 0 to 1, and as such, our features are not binary. Since 3w-Jaccard measure only handles binary features, we had to adapt it in order to incorporate weights, which represent how much a skill is valuable to a certain subcategory. Instead of (5.3), we propose: SWEIGHT ED−JACCARD =2a+a0 3a+b+c(5.6) 38 Designing the Solution (a) Dice (b) 3W-Jaccard (c) Mcconnaughey (d) Sorgenfrei Figure 5.1: Variation of different similarity measures values (color) depending on the number of skills that a job has (x-axis), a alumni has (y-axis) and the number of skills in common between the two (z-axis). Where a’ is the sum of the weights of subcategory skills that are present both the subcategory and an alumni. To see how a’ influences the similarity measure, we compute the distribution of weights in each subcategory. Figure 5.2 presents the weight distribution for all subcategories. It is clear that the majority of weights are lower than 0.4, leading to a’ assuming low values in comparison to a. By using Weighted-Jaccard (5.6) as our similarity measure we are able to emphasize aand incorporate the weights a’, allowing the skill weights to increase or decrease the similarity value depending on if they are close to 1 or 0. As it is possible to observe in Figure 5.3, with WeightedJaccard similarity measure we are able to maintain our goals for the measures regarding color variation and high similarity values. 5.3 Choosing the most suitable k In line 7of Algorithm 1, a value for khas to be established. It is a threshold for the number of most similar users that will be considered for the next step. The decision tree in Figure 5.4 illustrates the evolution of the process followed to choose the most suitable value for k. The decision reached is dependent on the data set used, namely the number of users and the frequency of each education possible to be recommended. In consequence, the decision process for determining the value of k should be performed every time the data set is updated. 39 Designing the Solution Figure 5.2: Distribution of skill weights in all subcategories. Figure 5.3: Variation of Weighted-Jaccard similarity values (color) depending on the number of skills that a job has (x-axis), a alumni has (y-axis) and the number of skills in common between the two (z-axis). 40 Designing the Solution Figure 5.4: Decision process on the value of k. Diamond shape boxes represent decisions and squared boxes represent actions. The blue squared box represents the final decision reached with our data set. 41 Proof of Concept System Figure 6.1: Number of possible recommendations by subcategory, depending on the value of k. (a) All frequencies. (b) Only considering frequencies greater than two. Figure 6.2: Density of education frequencies depending on the value of k. 48 Proof of Concept System (a) All recommendations. (b) Only considering common recommendations. Figure 6.3: Distribution of top-5 recommendations with k=5000 k=2500. kin the range of programs recommended. We use N=10 and N=5 because they correspond to the maximum and minimum number of recommendations that we aim at providing users. In Figures 6.3 and 6.4 we present the number of times that each higher education program appears in a recommendation for all subcategories, considering the top-5 and top-10 recommendations, respectively. It is possible to conclude that, using k=5000 the overlap between subcategories recommendations is higher than with k=2500, since we have more different recommended programs with the last one. The reason behind this is that, with higher kvalues, the more likely we are to select users that have one of the higher education programs with the higher overall frequency among alumni, as it was discussed in Section 5.4. By using a smaller value for k,k=2500, we are able to provide a wider range of higher education programs recommendations. Having a wider range of recommendations implies that these are probably more specific to each subcategory. At the same time, with k=2500 we have at least three alumni, from the 2500 selected, that share one of the programs recommended, as we can observe in Figure 6.5a. (a) All recommendations. (b) Only considering common recommendations. Figure 6.4: Distribution of top-10 recommendations with k=5000 and k=2500. 49 Proof of Concept System (a) Distribution of the number of possible recommendations. (b) Frequency of the number of recommendations. Figure 6.5: Distribution of the number of possible recommendations by subcategory, depending on the value of N. 6.2 Choosing the most suitable N In Section 2.4.3 it was pointed out that the probability of the user viewing an item decreases as its position on the list increases. As we briefly referred in the previous section, we decided that, whenever it was possible, the recommender system should give the user a number of recommendations ranging from 5 up to 10. The decision on the value of Nand the minimum frequency, among the kalumni chosen, for an education to be considered as a valid recommendation is made according to the recommendations obtained. Using k=2500 we guarantee an overall frequency of each education, among the kalumni chosen, higher than 2 as we can observe in Figure 6.5a, both for N=5 and N=10. With N=5, we are able to obtain higher values of frequencies for the recommended educations than with N=10, guaranteeing an education frequency among the chosen alumni that is equal to 3 at the worst case and, in average, equal to 3 or 4. In consequence, we established that the minimum value for Nshould be Nmin = 5. Since the frequencies of each possible recommendation have a close set of values as we can see in Figure 6.5b, we will have probably more than five recommendations with equal frequencies. To contemplate those cases we decided to establish a maximum of Nmax = 10. 6.3 Generation of a trimmed recommendation list We established Nmin = 5 and Nmax = 10. However, we encountered several cases similar to the one in Table 6.1. In the case of Table 6.1 we are not able to trim the recommendation list because, if we exclude the recommendations with a frequency smaller than 4, we would only have 50 Proof of Concept System Number of alumni Education 1 5 Education 2 4 Education 3 4 Education 4 3 Education 5 3 Education 6 3 Education 7 3 Education 8 3 Education 9 3 Education 10 3 Education 11 3 Education 12 3 Table 6.1: Example of the number of alumni that share the programs possible to be recommended in a certain subcategory. 3 recommendations. However, by including frequencies up to 3, we would have more than 10 recommendations, which is our maximum value for N. Figure 6.6 shows the minimum number of recommendations for each subcategory obtained with Nmin = 5 and without trimming the recommendation list. It is clear that in more than half of the subcategories we exceed our maximum value for N. In order to trim the recommendation list of subcategories where the list size is greater than 10, we followed the approach in Algorithm 2for those subcategories. Our approach takes advantage of the item similarity between skills of subcategories and alumni educations. Algorithm 2 Trim the recommendation list of a subcategory with size greater than 10. 1: Input: subcategory 2: original_list ←get_recommendations(subcategory) 3: elements_to_order ←select_elements_minimum_f requency(original_list) 4: most_similar_ordered ←[ ] 5: foreach element in elements_to_order do: 6: similarity ←compute_similarity(subcategory,education_name) 7: most_similar_ordered ←add_element(element,similarity) 8: elements_to_maintain ←original_list −elements_to_order 9: trimmed_recommendation_list ←elements_to_maintain ∪most_similar_ordered 10: Output: trimmed_recommendation_list In line 2we take the subcategory recommended educations that have minimum frequency, which are the ones that make the list size greater than 10. For example, if the recommendations were the ones in Table 6.1, we would select the elements from Education 4 until Education 12. Then, for each education selected in line 2, we compute its similarity to the current subcategory in line 6and store that value in an order list from the most to less similar in line 7. Finally, the trimmed recommendation list is equal to the educations not selected in line 2, plus the most similar among the selected ones until a maximum of 10 recommendations. 51 Proof of Concept System Figure 6.6: Number of possible recommendations by subcategory with k = 2500 and Nmin = 5. Red bars represent numbers greater than 10. 6.3.1 Similarity Computation In line 6of Algorithm 2, we compute the similarity between a subcategory and education by performing the steps in Algorithm 3. Algorithm 3 Similarity computation between a subcategory and an education. 1: Input: subcategory,education_name 2: skill_set ←job_skills ∩education_skills 3: job_matrix ←build_job_matrix(skill_set) 4: education_matrix ←build_education_matrix(skill_set) 5: job_desired ←get_element(job_matrix,subcategory) 6: education ←get_element(education_matrix,education_name) 7: similarity ←compute_3w_jaccard_similarity_with_weights(job_desired,education) 8: Output: similarity We start by identifying the set of skills that are present both in jobs and educations in line 2. In lines 3and 4we build the user-item matrices for both jobs and educations. More details on how these matrices are built are provided in Section 6.3.1.1. Then, the subcategory row is selected from the job matrix and the education row selected from the education matrix in lines 6and 7, respectively. Finally, we compute the similarity between the subcategory and education rows using the Weighted-Jaccard measure, with a’ being equal to the sum of the product of skills weights present in both. 6.3.1.1 User-item matrix The user-item matrix for jobs is the one explained in Section 5.1. Regarding the education matrix, we can make the correspondence with the user-item matrix: a user is an education identifier and items are skills. We start by identifying the skills associated 52 Proof of Concept System to each education identifier. If one skill is associated to an education identifier, then its value on the matrix is 1, otherwise is 0. Then, we move on to building a weighted education matrix. Considering educations that have the same degree and same location, we compute the average weight of each skill in that specific education name. We divide the frequency of each skill in the education by the number of educations identifiers that correspond to educations with that name. Just like the job matrix in Table 5.1, the education matrix is not a binary matrix because the weights of each skill vary from 0 up to 1. 6.4 Conclusions on the value of kand N Having in consideration all the factors analyzed in the previous sections, we decided to use k= 2500 and have 5 ≤N≤10 as parameters for the higher education programs recommendation. The number of recommendations presented is processed by selecting a minimum of five higher education programs and, in case of equal frequencies, extend the number of recommendations up to ten. If it is not possible to have less than ten recommendations, we trim the recommendation list using an approach based on the similarity between the items. In addition, we only consider educational programs taken by at least 3 alumni from the selected kas valid recommendations. Is is clear that the frequencies of each recommended higher education program can be considered as low but, having in consideration the sparsity of higher education programs and their overall low frequency as we pointed out in Section 4.3, we can consider the results obtained satisfactory. 6.5 Web based system Since recommendations are subjective and its appropriateness depends on the users, we built a web based system to allow the evaluation and validation of the higher education program recommendations, as shown in Figure 6.7. The program recommendations provided on the web based system are shown in Appendix D, which were obtained using k=2500 and 5 ≤N≤10. Through the web based system, users can select an area of expertise from the 40 available and see the recommended programs. In the case of Figure 6.7, the area Animation is selected and the recommended programs appear on the right-hand side. We decided to maintain our focus on recommendations of higher education programs and, at the same time, to allow the recommendation of online courses. The reason for this is that in some subcategories, according to the obtained results, one of the most suitable options might not be a higher education program but rather an online course, which can be viewed as a complementary education program. 6.5.1 Survey For validation and evaluation purposes, we distributed the survey in Appendix E. Our target audience was mainly students of the last year of the Master in Computer Engineering of our school and 53 Proof of Concept System Figure 6.7: Web based system. members of some information technology companies. We did this to make sure that the people who answered the survey and evaluated the recommendations had some knowledge in the area because, in that way, they could use their expertise in Computer Science to judge the recommendations and the system. In this survey we provided a link to our web based system, asked users to experiment it and then answer some questions regarding the system to assess their opinion. 6.5.1.1 Population The survey was answered by a total of 63 people. The population that participated is characterized as: •16% women and 84% men; •59% are between 18 and 23 years old, 31% are between 24 and 10% are more than 30 years; •25% are studying a Computer Science related degree and working at the same time, 50% are only studying and 24% are working in the area. 6.5.1.2 Results To assess if users consider that programs recommendation can help students choose a higher education program, we asked "Do you think this system could help students choosing a higher education program?". 88.9% answered positively, as shown in Figure 6.8. In order to figure out if the users felt that the subcategories described different jobs in Computer Science appropriately, we asked "How adequate did you consider the areas of expertise used?". As it is possible to see in Figure 6.9, 87.3% gave a positive answer and, from those, 41.3% classified the areas of expertise as very useful. 54 Proof of Concept System Figure 6.8: Answers to the survey question "Do you think this system could help students choosing a higher education program?". Figure 6.9: Answers to the survey question "How adequate did you consider the areas of expertise used?", where 1is not adequate at all and 4very adequate. 55 Proof of Concept System Figure 6.10: Answers to the survey question "How useful did you consider the recommendations?", where 1is not useful at all and 4very useful. To evaluate the users’ opinion towards the suitability of each recommended program, we asked "How useful did you consider the recommendations?". 52.4% of the users considered the recommendations as being useful and 20.6% as very useful, as it is possible to see in Figure 6.10. Finally, users could suggest possible improvements either by choosing one of the options provided or by giving their own suggestion. The two most popular answers were: "Filter recommendations by country or by university." and "Possibility to see profiles of alumni that took each of the recommended programs.", as shown in Figure 6.11. 6.6 Conclusions From our results, we consider that the usefulness of this type of systems for helping students choose a higher education program with a specific career in mind is undeniable. On one hand, we validated the areas of expertise used since most users felt that the areas had the ability to reflect a wide range of possible careers within Computer Science. On the other hand, we evaluated the recommendations given by assessing how useful they were viewed from the users’ point of view. It is certain that the majority of the answers were positive but, at the same time, 22.2% of users found out the recommendations as being useless. In correlation to the last question of the survey, we can speculate that, by using filters in recommendations and providing some kind of explanation behind them, such as alumni profiles, we could increase the sense of usefulness from the users perspective. In conclusion, having in consideration the overall opinion towards the system we consider our proof-of-concept validation and evaluation as being successful. 56 Proof of Concept System Figure 6.11: Answers to the survey question "Do you have any suggestions for possible improvements?". From top to bottom, the answers are "Filter recommendations by country or by university.", "Filter recommendations by degree.", "Disable online courses recommendations.", "Only see recommended online courses with more detail.", "Possibility to see profiles of alumni that took each of the recommended programs." and "Other". 57 REFERENCES [JSB11] Oliver Jojic, Manu Shukla, and Niranjan Bhosarekar. A probabilistic definition of item similarity. Proceedings of the fifth ACM conference on Recommender systems - RecSys ’11, page 229, 2011. [NKH04] Hyunju Noh, Minjung Kwak, and Ingoo Han. Improving the prediction performance of customer behavior through multiple imputation. Intell. Data Anal., 8(6):563–577, 2004. [oxf16] Oxford Dictionaries, 2016. [Paz99] Michael J. Pazzani. A framework for collaborative, content-based and demographic filtering. Artif. Intell. Rev., 13(5-6):393–408, December 1999. [RIS+94] Paul Resnick, Neophytos Iacovou, Mitesh Suchak, Peter Bergstrom, and John Riedl. Grouplens: An open architecture for collaborative filtering of netnews. In Proceedings of the 1994 ACM Conference on Computer Supported Cooperative Work, CSCW ’94, pages 175–186, New York, NY, USA, 1994. ACM. [RS05] Jasson Dm M Rennie and Nathan Srebro. Fast Maximum Margin Matrix Factorization for Collaborative Prediction. Proceedings of the 22Nd International Conference on Machine Learning, pages 713–719, 2005. [SK09] Xiaoyuan Su and Taghi M. Khoshgoftaar. A Survey of Collaborative Filtering Techniques. Advances in Artificial Intelligence, 2009(Section 3):1–19, 2009. [SKKR00] Badrul M. Sarwar, George Karypis, Joseph A. Konstan, and John T. Riedl. Application of dimensionality reduction in recommender system – a case study. In IN ACM WEBKDD WORKSHOP, 2000. [SM86] Gerard Salton and Michael J. McGill. Introduction to Modern Information Retrieval. McGraw-Hill, Inc., New York, NY, USA, 1986. [SSSHT10] Choi Seung-Seok, Cha Sung-Hyuk, and Charles C Tappert. A Survey of Binary Similarity and Distance Measures. Journal of Systemics, Cybernetics & Informatics, 8(1):43–48, 2010. [UF98] Lyle H. Ungar and Dean P. Foster. Clustering methods for collaborative filtering. AAAI Workshop on Recommendation Systems, pages 114–129, 1998. 64 Appendix A Job categories and subcategories Subcategory ID Subcategory Name Category ID Category Name Frequency 1web & mobile design 1 web, mobile & software dev 72 2other - design & creative 2 design & creative 32 3graphic design 2 design & creative 48 4mobile development 1 web, mobile & software dev 146 5web development 1 web, mobile & software dev 471 6other - software development 1 web, mobile & software dev 122 7human resources 3 accounting & consulting 4 8ecommerce development 1 web, mobile & software dev 49 9scripts & utilities 1 web, mobile & software dev 11 10 qa & testing 1 web, mobile & software dev 32 11 game development 1 web, mobile & software dev 27 12 database administration 4 it & networking 11 13 customer service 5 customer service 6 14 web content 6 writing 15 16 desktop software development 1 web, mobile & software dev 87 17 network & system administration 4 it & networking 56 18 data extraction / etl 7 data science & analytics 4 19 lead generation 8 sales & marketing 40 20 marketing strategy 8 sales & marketing 14 21 other - sales & marketing 8 sales & marketing 33 22 email & marketing automation 8 sales & marketing 6 23 technical support 5 customer service 9 24 data entry 9 admin support 18 25 smm - social media marketing 8 sales & marketing 22 26 technical writing 6 writing 10 27 electrical engineering 10 engineering & architecture 9 28 other - data science & analytics 7 data science & analytics 24 29 data mining & management 7 data science & analytics 15 30 machine learning 7 data science & analytics 11 65 Job categories and subcategories Subcategory ID Subcategory Name Category ID Category Name Frequency 31 other - it & networking 4 it & networking 56 32 article & blog writing 6 writing 33 33 other - engineering & architecture 10 engineering & architecture 5 34 product management 1 web, mobile & software dev 20 35 project management 9 admin support 37 36 resumes & cover letters 6 writing 2 37 sem - search engine marketing 8 sales & marketing 8 38 web research 9 admin support 16 39 erp / crm software 4 it & networking 13 40 other - accounting & consulting 3 accounting & consulting 4 41 personal / virtual assistant 9 admin support 32 42 logo design & branding 2 design & creative 16 43 3d modeling & cad 10 engineering & architecture 3 44 data visualization 7 data science & analytics 3 45 animation 2 design & creative 8 46 photography 2 design & creative 1 47 presentations 2 design & creative 4 48 video production 2 design & creative 14 49 copywriting 6 writing 24 50 other - writing 6 writing 7 51 accounting 3 accounting & consulting 8 52 financial planning 3 accounting & consulting 2 53 other - customer service 5 customer service 6 54 display advertising 8 sales & marketing 7 55 market & customer research 8 sales & marketing 6 56 public relations 8 sales & marketing 9 57 seo - search engine optimization 8 sales & marketing 10 58 academic writing & research 6 writing 8 59 management consulting 3 accounting & consulting 1 60 quantitative analysis 7 data science & analytics 3 61 editing & proofreading 6 writing 10 62 product design 10 engineering & architecture 10 63 illustration 2 design & creative 13 64 telemarketing & telesales 8 sales & marketing 18 65 other - admin support 9 admin support 13 66 general translation 11 translation 3 67 mechanical engineering 10 engineering & architecture 4 69 creative writing 6 writing 4 70 audio production 2 design & creative 1 72 transcription 9 admin support 2 73 interior design 10 engineering & architecture 4 75 technical translation 11 translation 2 77 other - legal 12 legal 2 79 paralegal services 12 legal 1 80 legal translation 11 translation 2 81 grant writing 6 writing 1 82 voice talent 2 design & creative 1 84 civil & structural engineering 10 engineering & architecture 1 -without subcategory - without category 2131 Table A.1: Frequency of jobs in each subcategory and association between categories and subcategories 66 Appendix B Education degrees cleaning details 1. Simplify degrees that did not correspond to a higher education: Degrees that included ’high school’, ’lyceum’, ’gymnasium’, ’abitur’, ’alevel’, ’gce’, ’advanced level’, ’aslevel’, ’10th’ or ’12th’ were transformed into ’high school’; Degrees that included ’coursera’ were transformed into ’coursera’; Degrees that included ’edx’ were transformed into ’edx’; Degrees that included ’udemy’ were transformed into ’udemy’; Degrees that included ’zend’ were transformed into ’zend’; Degrees that included ’mcts’, ’mcpd’, ’mcp’ or ’microsoft’ were transformed into ’microsoft’; Degrees that included ’online’ were transformed into ’online course’; Degrees that included ’certificate’ were transformed into ’certificate’; Degrees that included ’ccna’, ’google’, ’learn’, ’workshop’, ’camp’, ’training’, ’trainee’, ’trainer’, ’traine’ or ’trained’ were transformed into ’training’; Degrees that included ’self’, ’certified’ or ’conference’ were transformed into ’-’. 2. Replace expressions and characters: Replace ’&’ by ’and’; Replace characters with accent marks by the corresponding characters without accent marks; Replace degrees that were only had numbers by ’-’; Replace square brackets and brackets by parenthesis. 3. Remove characters and words: Remove full stops, commas, hyphens, quotation marks, apostrophes, double spaces; Remove ’in’, ’on’, ’of’; 67 Education degrees cleaning details Remove ’honours’, ’honors’, ’hons’; Remove ’distinction’, ’first class’, ’1st class’; Remove ’in progress’, ’incomplete’, ’not completed’, ’not finished’, ’unfinished’; Remove ’degree’, ’attended’, ’combined’; Remove expressions such as ’1 year’ or ’three years’. 4. Replace plural words by their singular version: Replace ’computers’ by ’computer’; Replace ’eletronics’ by ’eletronic’; Replace ’physics’ by ’physic’; Replace ’sciences’ by ’science’; Replace ’technologies’ by ’technology’; Replace ’telecommunications’ by ’telecommunication’. 5. Simplify expressions: Simplify ’bachelors’ to ’bachelor’; Simplify ’bachelor of science’ and ’bsc’ to ’bs’; Simplify ’b eng’ to ’be’; Simplify ’b tech’ and ’btech’ to ’bt’; Simplify ’masters’ to ’master’; Simplify ’masters’ to ’master’; Simplify ’master of science’ and ’msc’ to ’ms’; Simplify ’mathematics’ to ’math’; Simplify ’m eng’ to ’me’; Simplify ’ph d’ to ’ph d’; Simplify ’post doc’, ’postdoc’, ’post doctoral’ to ’postdoctoral’; Simplify ’post grad’, ’postgrad’, ’post graduate’ to ’postgraduate’. 6. Correct typos in frequent words: Find variations of the words ’engineering’ and ’bachelor’, and correct them. 7. Manually analyse and transform degrees that started with numbers 8. Manually translate degrees written in Asian languages 68 Appendix C Education locations cleaning details 1. Simplify locations that did not correspond to a higher education location: Locations that included ’high school’, ’lyceum’, ’liceo’ or ’secundaria’ were transformed into ’high school’; Degrees that included ’coursera’ were transformed into ’coursera’; Degrees that included ’edx’ were transformed into ’edx’; Degrees that included ’udemy’ were transformed into ’udemy’; Degrees that included ’mcts’, ’mcpd’, ’mcp’ or ’microsoft’ were transformed into ’microsoft’; Degrees that included ’self’ were transformed into ’self taught’. 2. Replace expressions and characters: Replace ’&’ by ’and’; Replace characters with accent marks by the corresponding characters without accent marks; Replace degrees that were only had numbers by ’-’; Replace hyphens between words by commas. 3. Remove characters: Remove full stops, quotation marks, apostrophes, underscores, ’@’, ’o’, ’a’ and double spaces; Remove spaces before commas. 4. Simplify expressions: Simplify ’saint ’ to ’st’; Simplify ’technical’ to ’tech’. 69 Education locations cleaning details 5. Correct typos in frequent words: Find variations of the word ’university’ and correct them. 6. Manually translate degrees written in Asian languages 7. Verify equal locations that only differ on commas: If yes, transform the location without commas to the version with commas because it was more readable; 8. Verify if one location contained another: If yes, transform the longer one to simpler (smaller) one, which was the most general. 9. Verify if one location without commas contained another: Follow the rule in previous step. 10. Verify if there were equal locations without white spaces: If yes, transform the location with white spaces to the version that had; This was especially relevant for cases such as ’u a’ and ’ua’. 70 Appendix D Higher Education Program Recommendation 1: web & mobile design Web Development Immersive – General Assembly Certificate – Nova Scotia Community College Bachelor of Science in Computer Science – University of California Computer Science – University of Maryland Certificate – British Columbia Institute of Technology Bachelor of Science in Computer Science – California State University Bachelor of Science in Computer Science – Middlesex University Bachelor of Science in Computer Science – Politecnico di Milano Bachelor of Technology – Jawaharlal Nehru Technological University Full Stack Web Development – App Academy 2: other - design & creative Online Course – Codecademy Master of Computer Applications – Anna University Bachelor of Science in Computer Science – University of California Master of Computer Applications – Sikkim Manipal University Bachelor Information Technology – Macquarie University Bachelor of Science in Computer Science – University of Karachi Bachelor of Technology– Anna University Bachelor of Engineering in Computer Science and Engineering – Anna University Bachelor of Computer Applications – Guru Gobind Singh Indraprastha University 71 Higher Education Program Recommendation Course – University of Buenos Aires 3: graphic design Certificate – British Columbia Institute of Technology Bachelor of Science in Computer Science – University of California Course – University of Buenos Aires Computer Information Systems – Devry University Master of Science – Birla Institute Bachelor of Engineering – Anna University 4: mobile development Certificate – Madison College Computer Science – Rochester Institute of Technology Certificate – Nova Scotia Community College Master of Science – Birla Institute Online Course – Mongodb University Web Development – Dev Bootcamp Bachelor of Science in Computer Science – Florida State University Web Development Immersive – General Assembly Bachelor of Science in Computer Science – University of Michigan Bachelor of Science in Computer Science – University of Havana 5: web development Web Development Immersive – General Assembly Certificate – Nova Scotia Community College Certificate – British Columbia Institute of Technology Bachelor of Science in Computer Science – University of Michigan Bachelor in Computer Science – Central University Computer Science – University of Havana Bachelor of Science in Computer Science – Middlesex University Online course – Mongodb University Certificate – Notre Dame University Online course – Microsoft 72 Higher Education Program Recommendation 6: other - software development Certificate – Nova Scotia Community College Bachelor of Technology – Jawaharlal Nehru Technological University Master of Science in Software Engineering – San Jose State University Bachelor of Science in Computer Science – University of California Master of Science in Computer Science – Maharishi University of Management Bachelor of Science in Computer Science – University of the Philippines Bachelor of Science in Computer Science – University of Michigan Web Development Immersive – General Assembly Bachelor of Science in Computer Science – University of Bucharest Bachelor in Computer Science – Central University 8: ecommerce development Bachelor of Science in Computer Science – Central University Bachelor of Science in Computer Science – University of California Certificate – Nova Scotia Community College Web Development Immersive – General Assembly Computer Science – University of Texas Master of Computer Applications – Anna University Master Computer Science – Maharishi University of Management Certificate – Code Camp Certificate – Stanford University Bachelor of Information Technology – Macquarie University 9: scripts & utilities Bachelor of Science in Computer Science – University of the Philippines Certificate – Nova Scotia Community College Web Development Immersive – General Assembly Bachelor of Science in Information Technology – Rochester Institute of Technology Bachelor of Science in Computer Science – University of California Bachelor of Science in Computer Science – University of Michigan Bachelor of Science in Computer Science – University of Havana Bachelor of Science in Computer Science – Central University 73 Higher Education Program Recommendation Bachelor of Science – Universidade de Coimbra Bachelor of Science in Computer Science – University of Nebraska Bachelor of Science in Computer Science – University of Applied Science Bachelor of Science in Computer Science – University of Havana 45: animation Bachelor of Science in Computer Science – Warsaw University Bachelor of Science in Computer Science – University of California Certificate – South Pacific Sil Bachelor engineering – Anna University Bachelor of Science in Computer Science – Technical University 48: video production Bachelor of Engineering in Computer Science and engineering – Anna University Master – National Technical University of Ukraine Bachelor of Science in Computer Science – University of Illinois Bachelor of Science in Computer Science – Indiana University Certificate – British Columbia Institute of Technology Bachelor of Science in Computer Science – University of Karachi Bachelor of Science in Computer Science – University of California 62: product design Web Development Immersive – General Assembly Web Development – Dev Bootcamp Web Development – App Academy Full Stack Web Development – App Academy Web Development – General Assembly Web Development – Makers Academy Computer Science – University of Maryland Bachelor of Science in Computer Science – University of California Bachelor of Science in Computer Science – University of North Carolina Certificate – Madison College 80 Higher Education Program Recommendation 63: illustration Bachelor of Science in Computer Science – University of California Bachelor of Science in Computer Science – Warsaw University Bachelor of Science in Computer Science – National University of Computer and Emerging Sciences Video Game Design and Development – Trios College Bachelor of Science in Computer Science – State University Bachelor of Science in Computer Science – Universidad Complutense de Madrid 65: other - admin support Bachelor Technology – Jawaharlal Nehru Technological University Certificate – Nova Scotia Community College Bachelor of Engineering – Visvesvaraya Technological University Master of Computer Applications – Sikkim Manipal Bachelor engineering – Anna University Master computer application – Gujarat Technological University Bachelor of Science in Computer Science – University of Maryland Bachelor of Science in Computer Science – University of California Computer Science – Indiana University Master Computer Applications – Anna University 67: mechanical engineering Bachelor of Science in Computer Science – Warsaw University Certificate – South Pacific Sil Bachelor of Science in Computer Science – University of California Bachelor of Science in Computer Science – Technical University Software Engineering – Niit Bachelor of Science in Computer Science – Belarusian State University University Bachelor of Science in Computer Science – University of Havana Bachelor of Science in Computer Science – University of Bucharest Bachelor of Science in Computer Science – Georgia State University Bachelor engineering – Anna University 81 Higher Education Program Recommendation 85: business intelligence Master of Science – University of California Computer Science – University of California Bachelor of Science in Computer Science – University of Michigan Bachelor of Science in Computer Science – University of California Bachelor of Science in Computer Science – University of Texas Master of Science in Computer Science – University of Texas Bachelor Technology – Indian Institute of Technology Bachelor of Science in Computer Science – Brigham Young University Bachelor – University of California Computer Science – University of Maryland 86: security Bachelor of Science in Computer Science – University of Illinois Master of Science in Computer Science – Johns Hopkins University Bachelor engineering – Anna University Bachelor of Science in Computer Science – Worcester Polytechnic Institute Bachelor of Science in Computer Science – University of California Bachelor Technology – Indian Institute of Technology Bachelor of Science in Computer Science – Brigham Young University Bachelor of Science in Computer Science – California State University Computer Science – Technical University Bachelor of Science in Computer Science – University of Massachusetts 87: cloud Bachelor of Science in Computer Science – Washington State University Bachelor of Science in Computer Science – Arizona State Bachelor of Science in Computer Science – University of California Master – Faculdade de Engenharia da Universidade do Porto Master of Engineering in Computing – Imperial College Master – Moscow State University Bachelor of Engineering in Computer Engineering – Mumbai University Bachelor of Science in Computer Science – University of Colorado 82 Higher Education Program Recommendation Bachelor of Science in Computer Science – Universidade Catolica Master of Science in Computer Science – University of California 88: devops Online Course – Mongodb University Bachelor of Science in Computer Science – Purdue University Bachelor – Universidade de Sao Paulo Certificate – Nova Scotia Community College Bachelor of Science in Computer Science – University of Texas Bachelor of Science in Computer Science – California State University Bachelor of Science in Computer Science – University of California Bachelor Computer Science – Central University Bachelor of Science in Computer Science – Florida State Certificate – British Columbia Institute of Technology 89: software architecture Computer Science – University of California Bachelor of Science in Computer Science – University of California Bachelor of Science in Computer Science – University of Maryland Certificate – Nova Scotia Community College Bachelor of Engineering in – Gujarat Technological University Bachelor of Engineering in Computer Science – Anna University Master of Science in Computer Science – Maharishi University of Management Bachelor of Science in Computer Science – Instituto Superior Tecnico Bachelor of Science in Computer Science – University of Michigan Bachelor of Engineering in – Anna University 90: infrastructure engineering Bachelor of Science in Computer Science – Central University Bachelor of Science in Computer Science – University of California Bachelor of Science in Computer Science – University of the Philippines Bachelor information Technology – Rochester Institute of Technology Bachelor of Engineering in Computer Science – Anna University Master of Science in Computer Science – University of Illinois 83 Higher Education Program Recommendation Bachelor Computer Science – Universita Degli Studi di Padova Bachelor of Science in Computer Science – University of Illinois Master of Science in Computer Science – northwestern University Bachelor Technology – Jawaharlal Nehru Technological University 91: technical management Bachelor of Science in Computer Science – University of California Master of Science in Computer Science – Technical University Computer Science – University of Waterloo Bachelor of Science in Computer Science – Brigham Young University Bachelor of Science – University of Toronto Bachelor of Science in Computer Science – Wentworth Institute of Technology Bachelor of Science in Computer Science – Universidade Catolica Full Stack Web Development – Code Camp Bachelor of Science in Computer Science – University of Washington Bachelor of Information Technology in Systems Engineering – Hasso Plattner Madison 84 Appendix E Survey Figure E.1: Survey first page with questions related to the user. 85 Survey Figure E.2: Part one of the survey second page with questions related to the web based system. 86 Survey Figure E.3: Part two of the survey second page with questions related to the web based system. 87