scieee AI-readable full text Open interactive document viewer

Documenting Software With Adaptive Software Artifacts

Filipe Alexandre Pais de Figueiredo Correia

Full text

UNIVERSIDADE DO PORTO • FACULDADE DE ENGENHARIA Documenting Software With Adaptive Software Artifacts Filipe Alexandre Pais de Figueiredo Correia January 2015 Scientific Supervision by Ademar Aguiar, Assistant Professor Departmento de Engenharia Informática In partial fulfillment of requirements for the degree of Doctor of Philosophy in Informatics Engineering by the ProDEI Doctoral Programme Contact Information: Filipe Figueiredo Correia Faculdade de Engenharia da Universidade do Porto Departamento de Engenharia Informática Rua Dr. Roberto Frias, s/n 4200-465 Porto Portugal Tel.: +351 22 508 1400 Fax.: +351 22 508 1440 Email: [email protected] URL: http://invisivel.net This thesis was typeset on an Apple®MacBook®Pro running Mac OS®10.6using the free L A TEX typesetting system, originally developed by Leslie Lamport based on TEX created by Donald Knuth. The body text is set in Palatino, a large typeface family that began as an old style serif typeface designed by Hermann Zapf and that was initially released in 1948 by the Linotype foundry. Other fonts include Sans, Latin Modern and Typewriter from the Computer Modern family, and Courier, a monospaced font originally designed by Howard Kettler at IBM and later redrawn by Adrian Frutiger. The L A TEX style was based on the one created by Hugo Ferreira for his own PhD thesis. Most charts were drawn using the matplotlib python library. This work was partially funded by the fct grant number SFRH/BDE/33883/2009, with the support of the European Social Fund (poph) and of paradigmaxis,s.a.. ISBN 978-972-752-168-5 9 789727 521685 Filipe Figueiredo Correia “Documenting Software With Adaptive Software Artifacts” Copyright © 2015 Filipe Figueiredo Correia. All rights reserved. . . . to my parents and brothers This page was intentionally left blank. Abstract The label knowledge work applies perfectly to the craft of software development, since it focuses on acquiring, transforming and capturing knowledge in various forms – as different kinds of software artifacts. While the creation of source code is a key part of software developers’ end-goals, knowledge is usually first captured in forms that are not as objective, not as structured, and not as close to a working software system. To an extent, this happens due to the nature of knowledge itself; ideas often have to be iteratively discussed, reasoned upon and written informally before they can become more structured kinds of software artifact. Many of these artifacts may be classified as documentation, as they don’t play a part in the functional aspects of the final, working, software system, but rather provide information about it, to assist its development, use, or maintenance. When referring to documentation we often don’t refer only to free-text documents but to a diversity of artifacts, such as user stories, tasks, models and source code, among others. Many approaches to software documentation exist and address different needs, with wikis, literate programming and code annotations as some of the most influential ones. They imply different trade-offs and allow us to identify a set of factors that should be taken into account when designing new documentation approaches. These factors include the choice of supporting mediums (e.g., file-based or Web-based), whether to use a single-source or a multiple-source approach, to which extent the artifacts should enforce a specific structure, and how can multiple artifacts be integrated and related. An issue that is especially important in the context of this work is the evolution of the knowledge captured within these software artifacts. Although knowledge will freely evolve in the minds of a software development team, artifacts may not be as easy to adapt accordingly, especially when changes to their structure are needed. The types of software artifacts that imply a specific structure have been conceived bearing in mind a particular type of knowledge. They are likely very expressive in that domain, but not easy to change beyond their predefined structure. Free-text documents are particularly worthy of note for their capability of conveying information in a wide ii abstract range of domains but they hardly enforce any domain structure and usually imply a certain ambiguity and verbosity. On the other end of the spectrum, source code in a general purpose programming language is unambiguous and terse, but it can only convey information on the domain of computations. A lot of software artifacts lie somewhere between these two extremes. Although developers create software systems to manage information (which are often denoted as information systems), they don’t usually see the many products of their work as information that they have to manage, and their tools and environment as the means to do so. This research recognizes that to make the most of the information captured during software development, it must be regarded as a primary source of knowledge. Furthermore, this work draws inspiration from architectures used in the context of adaptive software to address three specific concerns – the expression of information structure, the maintenance of contents’ consistency and the classification of those contents so that they can be found more easily. Adaptive Software Artifacts is an approach conceived to tackle these concerns by combining benefits of free-form contents with those of structured contents. It has the goal of reducing the barrier to define new types of structure (i.e., new types of adaptive software artifacts) and allows adaptive software artifacts to be derived from textual contents and combined with them as needed. It tries to make it easy for developers to structure contents that otherwise would remain free-form and, therefore, less useful. The contribution of this work is fourfold: a) a patterns catalog that formalizes good practices and design considerations surrounding software documentation, information classification, flexible modeling tools and adaptive object-models; b) the Adaptive Software Artifacts approach to software documentation; c) a reference architecture and implementation that was used to verify the practicality of the approach and to help validate it; and d) a statistical experiment that can be replicated independently, with the goal of validating the approach. The experiment was run once with students and the results revealed some benefits in knowledge acquisition when using documentation following the Adaptive Software Artifacts approach. Resumo O desenvolvimento de software é um trabalho baseado em conhecimento, já que tem como foco a aquisição, transformação e captura de conhecimento em várias formas – como diferentes tipos de artefatos de software. Embora o código fonte seja essencial nos objetivos dos programadores, é normal o conhecimento começar por ser capturado segundo formas que não são tão objetivas, tão estruturadas, nem tão próximas de um sistema de software funcional. Isso acontece pela natureza do próprio conhecimento; as ideias muitas vezes tem de ser iterativamente discutidas, pensadas e escritas informalmente, antes de se tornarem artefatos de software mais estruturados. Muitos destes artefatos podem ser classificados como documentação, já que não têm um papel no funcionamento do sistema de software, fornecendo apenas informação sobre ele, que assiste o seu desenvolvimento, utilização ou manutenção. Ao nos referirmos a documentação, muitas vezes não nos referimos só a documentos de texto livre, mas a uma diversidade de artefatos, como user stories, tarefas, modelos e código fonte, entre outros. Existem muitas abordagens à documentação de software, cada uma preenchendo necessidades diferentes, sendo os wikis, a programação literária e a anotação de código algumas das mais influentes. Cada uma apresenta diferentes benefícios e deficiências, e permitem identificar um conjunto de fatores a ter em conta no desenho de novas abordagens de documentação. Estes fatores incluem a escolha dos meios de suporte (e.g., ficheiros ou Web), se é usada uma abordagem fonte-única (single-source) ou fontemúltipla (multiple-source), em que medida os artefatos devem forçar uma estrutura específica, e como podem múltiplos artefatos ser integrados e relacionados. Uma questão de especial importância no contexto deste trabalho é a evolução do conhecimento capturado nestes artefatos. Embora o conhecimento possa evoluir livremente nas mentes de uma equipa de desenvolvimento de software, os artefatos podem não ser tão fáceis de adaptar de acordo, especialmente quando são necessárias alterações à sua estrutura. Os tipos de artefatos de software que implicam uma estrutura específica foram concebidos tendo em conta um tipo de conhecimento em iv resumo particular. São provavelmente muitos expressivos nesse domínio, mas não são fáceis de alterar além da sua estrutura pré-definida. Os documentos de texto livre são particularmente dignos de nota pela sua capacidade de transmitir informação num grande leque de domínios mas dificilmente garantem qualquer estrutura de domínio e normalmente implicam uma certa ambiguidade e verbosidade. No outro extremo do expectro, as instruções de código fonte conseguem ser bastante inequívocas e concisas, mas só conseguem transmitir informação no domínio das computações. Muitos dos artefatos de software encontram-se algures entre estes dois casos extremos. Embora os programadores criem frequentemente software para gerir informação (frequentemente chamados sistemas de informação), normalmente não vêm os produtos do seu trabalho como informação que têm de gerir, e as suas ferramentas e ambiente como meios para o fazer. Esta investigação reconhece que para tirar o máximo partido da informação capturada no desenvolvimento de software, esta tem de ser considerada uma fonte primária de conhecimento. Este trabalho inspira-se em arquiteturas usadas por sistemas adaptativos para abordar três questões específicas – a expressão de estruturas de informação, a manutenção da consistência dos conteúdos e a classificação desses conteúdos para que possam ser encontrados mais facilmente. A abordagem Artefatos de Software Adaptativos foi concebida para endereçar estes desafios, combinando os benefícios associados a conteúdos livres de estrutura com os benefícios dos conteúdos estruturados. Tem como objetivo reduzir a barreira à criação de novos tipos de estrutura (i.e., novos tipos de artefatos de software adaptativos) e permite que os artefatos de software adaptativos sejam derivados diretamente de conteúdos textuais e combinados com eles conforme necessário. Tenta tornar mais fácil estrurar conteúdos que, de outra forma, se manteriam em texto livre e, portanto, menos úteis. A contribuição deste trabalho tem quatro partes: a) um catálogo de padrões que formaliza boas práticas e considerações de desenho em torno da documentação de software, classificação de informação, ferramentas de modelação flexíveis e modelosde-objetos adaptativos; b) a abordagem de documentação de software designada como Artefatos de Software Adaptativos; c) uma arquitetura e implementação de referência, usada para verificar a viabilidade da abordagem e para ajudar a validá-la e d) uma experiência estatística que pode ser replicada de forma independente, com o objetivos de validar a abordagem. O desenho experimental foi corrido uma vez com estudantes e os resultados revelaram alguns benefícios na aquisição de conhecimento ao usar documentação com base na abordagem de Artefatos de Software Adaptativos. Contents Abstract i Resumo iii List of Figures xi List of Tables xvii Preface xix 1Introduction 1 1.1Software Documentation ............................ 2 1.2Research Goals and Contributions ...................... 3 1.3Experimental Findings ............................. 4 1.4Thesis Overview ................................ 4 2Documenting Software 7 2.1Software Knowledge .............................. 8 2.1.1Knowledge Sharing and Preservation ................ 9 2.1.2From Knowledge to Software Artifacts ............... 9 2.1.3Evolving Knowledge and Artifacts .................. 11 2.2Software Artifacts ................................ 12 2.2.1Source Code ............................... 12 2.2.2Models .................................. 13 2.3Approaches to Software Documentation ................... 14 2.3.1Diversity of Software Documentation ................ 14 2.3.2Wikis ................................... 15 2.3.3Literate Programming ......................... 22 2.3.4Code Annotations ........................... 28 xii LIST OF FIGURES 7.6Example of an instance of the metaclass Entity............... 137 7.7Instantiation and inheritance chains of the Model module with details. .138 7.8Database schema of the Adaptive Software Artifacts Plugin. ....... 139 7.9The plugin’s user interface – option to access the index. ......... 142 7.10 The plugin’s user interface – different types of adaptive artifacts. .... 142 7.11 The plugin’s user interface – adaptive artifacts of the type customer.. . 143 7.12 The plugin’s user interface – option to view details of an adaptive artifact. 143 7.13 The plugin’s user interface – details of an adaptive artifact. ....... 144 7.14 The plugin’s user interface – button to create a new adaptive artifact. .144 7.15 The plugin’s user interface – creation of a new type of adaptive artifact. 145 7.16 The plugin’s user interface – list/create adaptive artifacts of a new type. 145 7.17 The plugin’s user interface – creation of a new adaptive artifact. .... 146 7.18 The plugin’s user interface – select text to create new adaptive artifact. .146 7.19 The plugin’s user interface – form to create new adaptive artifact. .... 147 7.20 The plugin’s user interface – connect to existing adaptive artifact. .... 148 7.21 The plugin’s user interface – choose adaptive artifact to link to. ..... 148 7.22 The plugin’s user interface – result of linking to adaptive artifact. .... 149 7.23 The plugin’s user interface – adaptive artifacts graph as class diagram. .149 7.24 Timeframe and activity on the project’s source code repository. ..... 150 8.1Experiment steps and their relative durations ................ 156 8.2Mean of the times spent on the platform by platform module. ...... 166 8.3Mean of the times spent on each task. .................... 170 A.1Type Object pattern in the python programming language. ....... 192 A.2Language Piggybacking pattern in the python programming language. 193 B.1Background questionnaire. .......................... 196 B.2Instructions sheet. ................................ 197 B.3Platform instructions provided to the Control Group. ........... 198 B.4Platform instructions provided to the Experimental Group. ....... 199 B.5Tasks page provided to the Control Group. ................. 200 B.6Tasks page provided to the Experimental Group. ............. 201 B.7Description of the first task, as provided to the Experimental Group. . . 201 B.8Assessment questionnaire – Page 1...................... 202 B.9Assessment questionnaire – Page 2...................... 203 B.10 Assessment questionnaire – Page 3...................... 204 LIST OF FIGURES xiii D.1Core classes of the ASA Analyzer domain model. ............. 212 E.1Boxplot of the subjects’ mean grades. .................... 213 E.2Boxplot of the mean times spent on the platform by Trac module. .... 214 E.3Boxplot of the mean times spent on the platform. ............. 215 E.4Boxplot of the mean duration of each task. ................. 217 E.5Boxplot of the mean duration of the totality of the tasks. ......... 218 E.6Histogram of the answers to the questionnaire item BG1.1– I have considerable experience using the Java programming language. ..... 219 E.7Histogram of the answers to the questionnaire item BG1.2– I have considerable experience using the Eclipse IDE. .............. 219 E.8Histogram of the answers to the questionnaire item BG1.3– I have considerable experience using Software Forges. .............. 220 E.9Histogram of the answers to the questionnaire item BG1.4– I have considerable experience using the Trac platform. ............. 220 E.10 Histogram of the answers to the questionnaire item BG1.5– I have considerable experience using Trac’s Adaptive Software Artifacts or Custom Software Artifacts. ............................ 220 E.11 Histogram of the answers to the questionnaire item BG1.6– I have considerable experience using frameworks. ................ 221 E.12 Histogram of the answers to the questionnaire item BG1.7– I have considerable experience using the JHotDraw framework. ........ 221 E.13 Histogram of the answers to the questionnaire item BG1.8– I have considerable experience with object-oriented software development. ... 221 E.14 Histogram of the answers to the questionnaire item BG1.9– I have considerable experience extending a system using composition and subclassing. .................................... 222 E.15 Histogram of the answers to the questionnaire item BG1.10 – I have considerable experience developing industry-level applications. ..... 222 E.16 Histogram of the answers to the questionnaire item BG1.11 – I have considerable experience maintaining/modifying industry-level applications. 222 E.17 Histogram of the answers to the questionnaire item BG1.12 – I have considerable experience documenting software systems. ......... 223 E.18 Histogram of the answers to the questionnaire item BG1.13 – I have considerable experience using technical documentation of software systems. 223 xiv LIST OF FIGURES E.19 Histogram of the answers to the questionnaire item BG1.14 – I have considerable experience using wikis. ................... 223 E.20 Histogram of the answers to the questionnaire item BG1.15 – I have considerable experience developing standalone GUI (Graphical User Interface) applications. .............................. 224 E.21 Histogram of the answers to the questionnaire item EF1– The room environment was distracting. ..................... 224 E.22 Histogram of the answers to the questionnaire item EF2– I found difficulties using the IDE. ........................ 224 E.23 Histogram of the answers to the questionnaire item EF3– I found difficulties using the Java language. ................... 225 E.24 Histogram of the answers to the questionnaire item OP1– I found it easy to translate my knowledge of the problem domain to a concrete solution. ..................................... 225 E.25 Histogram of the answers to the questionnaire item OP2– The project’s documentation was easy to use. .................. 225 E.26 Histogram of the answers to the questionnaire item OP3– The tasks descriptions were easy to understand. ................. 226 E.27 Histogram of the answers to the questionnaire item OP4– I enjoyed the programming exercise. ....................... 226 E.28 Histogram of the answers to the questionnaire item IA1– The information that was made available was in sufficient quantity. ...... 226 E.29 Histogram of the answers to the questionnaire item IA2– The information that was made available was not in excessive quantity. .... 227 E.30 Histogram of the answers to the questionnaire item IA3– The information that was made available was of good quality. ......... 227 E.31 Histogram of the answers to the questionnaire item IA4– The information that was made available was very precise (i.e., accurate; objective) 227 E.32 Histogram of the answers to the questionnaire item IA5– The information that was made available was very concise (i.e, terse; succinct) 228 E.33 Histogram of the answers to the questionnaire item CL1– I could easily find the information that I needed. ................ 228 E.34 Histogram of the answers to the questionnaire item CL2– The way in which the information was organized and linked allowed me to find it more easily. ................................... 228 LIST OF FIGURES xv E.35 Histogram of the answers to the questionnaire item CL3– I found what I needed to know by browsing the available contents. ....... 229 E.36 Histogram of the answers to the questionnaire item CL4– I found what I needed to know by using Trac’s Search feature. ......... 229 E.37 Histogram of the answers to the questionnaire item UN1– The information that was made available was always easy to understand. ... 229 E.38 Histogram of the answers to the questionnaire item UN2– The way in which the information was organized and linked allowed me to understand it more easily. ............................ 230 E.39 Histogram of the answers to the questionnaire item CO1– The information that was made available was often inconsistent. ........ 230 E.40 Histogram of the answers to the questionnaire item CO2– I don’t have a good perception if the information that was available to me was consistent or not. ................................. 230 xvi LIST OF FIGURES List of Tables 2.1Representation of active text elements as a literate program. ....... 27 7.1Features of the Adaptive Software Artifacts Plugin 0.5........... 133 7.2Additional features of the Adaptive Software Artifacts Plugin 0.5.... 133 7.3Programming languages used in the Adaptive Software Artifacts Plugin. 150 7.4Number of lines and unit-test coverage by python module. ........ 151 8.1Descriptive statistics of the students’ grades. ................ 163 8.2Mann-Whitney U test for the comparison of students’ grades. ...... 163 8.3Summary of the answers to the background questionnaire. ........ 164 8.4Descriptive statistics of the time spent on each Trac module. ....... 165 8.5t-t ests for the equality of means of times spent on two platform modules. 166 8.6Descriptive statistics of the total times spent on the Trac environments. 167 8.7t-test for the equality of means of times spent on the Trac environments. 167 8.8Descriptive statistics of the time spent on each activity. .......... 168 8.9t-tests for the equality of means of times spent on each activity. ..... 168 8.10 Descriptive statistics of the time spent on each task. ............ 169 8.11 t-tests for the equality of means of the spent spent on each task. ..... 170 8.12 Descriptive statistics of the total time spent completing the tasks. .... 171 8.13 t-test for the equality of means of the total times spent on the tasks. . . 171 8.14 Summary of the answers to the EF items of the assessment questionnaire. 172 8.15 Summary of the answers to the OP items of the assessment questionnaire. 173 8.16 Summary of the answers to the IA items of the assessment questionnaire. 174 8.17 Summary of the answers to the CL items of the assessment questionnaire. 175 8.18 Summary of the answers to the UN items of the assessment questionnaire. 176 8.19 Summary of the answers to the CO items of the assessment questionnaire. 176 8.20 Research issues addressed by the experiment. ............... 180 9.1Summary of the research issues’ validation. ................. 185 xviii LIST OF TABLES C.1Grades of the subjects of the experimental group. ............. 205 C.2Grades of the subjects of the control group. ................. 206 C.3Task durations. ................................. 206 C.4Answers to the background and assessment questionnaires. ....... 210 E.1Levene test for the equality of variances of the students’ grades. ..... 214 E.2Levene test for the equality of variances of questionnaire answers. ... 216 Preface A culpa não, não é do sol, se o meu corpo se queimar A culpa não, não é da praia, se o meu corpo se ferir A culpa é da vontade, que vive dentro de mim, e só morre com a idade, com a idade do meu fim... A culpa é da vontade... António Variações I am told that this preface should explain you, dear reader, how and why this work came to be. The truth is I can probably trace it to my early years as a boy, and my general interest in computer games and in something called the ZX Spectrum 48K, followed by some interest a few years later in another little box called the Nintendo Entertainment System. It amazes me the influence that these toys had on a whole generation, and I am sure this thesis is one more tiny ripple of the impact they had on the pond of my childhood. But there were more important factors conspiring to bring me here. I’m sure that watching my father design, build or fix whatever was needed around the home had no small role on my interest for the idea of engineering. It seems obvious, in retrospect, but I remember being quite undecided by the end of high-school as to which area I would like to pursue in college. Ironically, the results of a psychometric test suggested that I should enroll in a humanities course, but I knew that that couldn’t possibly work for me. If anything, that test made me more certain that I should turn towards engineering or science – I have always been one to tinker, to build, and to share the results with others. Fast forwarding to a few years later, I am happy with the course of my professional life in the industry – therefore, what could have possibly led me to pursue a PhD? The good memories I had from university a few years before played a tiny part but the main motivations were first and foremost my love for software engineering and design and the challenge to overcome myself. Research meant a carte blanche to seek xx preface new knowledge, free from the everyday self-pressure to deliver working, immediately useful, software. I take from this journey much more than what I have learned within my specific topics of research, from design patterns, to agile software development, to the bonds that I have created with others. As to the specific topic of my research, it’s difficult to look back and get a clear image of the path that first led me to it. Someone that I much admire asked me why have I chosen this topic and, to my own surprise, I replied that I didn’t think it had really been me who found it – rather, it found me! That answer came spontaneously and stuck with me for the next few days, as it made more and more sense the more I thought about it. The topic of this thesis emerged over time, from the confluence of several seemingly unrelated subjects that I am fond of. It starts with the early contact in my professional life with the notion of literate programming by the hand of my employer and friend Alexandre Sousa, who had himself contact with the concept during his PhD. It is influenced by my efforts at ParadigmaXis to have my team all collaborate through the same platform (a software forge), which included the creation of documentation for users and for developers using a wiki. My advisor Ademar Aguiar would rekindle my interest in these topics again years after. I was also influenced through my contact with topics from the domain of information science, in the context of my most enduring project at ParadigmaXis, aimed at cataloging and retrieving archive documents and their metadata. Finally, this work also stemmed from the research of my colleague and friend Hugo Ferreira on the adaptive object-model architectural pattern, with whom I’ve collaborated often during the time he was pursuing his own PhD. I have received the help and encouragement of many people along this journey and my appreciation is beyond words. My deepest gratitude has to go to my advisor, mentor and friend, Ademar Aguiar, who was always able to make me refocus on what was most important when I found myself off the right track – thank you for your understanding and for believing in me even when I felt lost. In the category of mentors and friends I also have to deeply thank Alexandre Sousa, who was the responsible for my first job in the industry and supported me to pursue a PhD while working for ParadigmaXis – I know that not many would have given me such an opportunity, thank you! I also have to thank Hugo Ferreira and Nuno Flores, for their friendship and companionship – this work would not be the same without your support and our joint brainstormings sessions. My thanks goes also to many other co-workers who, along the years, helped me grow professionally many times beyond myself – Alexandre Pinto, José Vilaça, Hugo Silva, Fátima Pires, João Ferreira, Tiago Cunha, Aurélio Pires, xxi Bruno Matos, Pedro Abreu, José Porto, Diogo Lapa and Ricardo Almeida, among others. A special thanks has also to go to my teachers, colleagues and university co-workers: to João Correia Lopes, who encouraged me to write my first academic paper right after finishing my licentiate in 2002; to Eugénio Oliveira, Augusto Sousa, Raul Vidal, João Pascoal Faria, Ana Paiva and Rui Maranhão, for their friendship and for making me feel welcome back at FEUP after my years in the industry; to Tiago Boldt and Fábio Pinho for their friendship and support; to Idalina Silva and Marisa Silva, for their welcoming smile and readiness to solve all sorts of issues. I also have to thank João Araújo from Universidade Nova de Lisboa and (again) to João Pascoal Faria for accepting to participate in this PhD’s steering committee and for their valuable feedback. During this time I have participated of several communities. I have to give a special thanks to the patterns community, for their insights and for making me feel so welcome. My thanks go out to Joseph Yoder, Rebecca Wirfs-Brock, Eduardo Guerra, Lise Hvatum, Bob Hanmer, Linda Rising, Christian Kohls, Christian Köppe, Richard Gabriel, Peter Sommerlad, Ralph Johnson and Brian Foote, among many other friends and conference buddies. To my longtime friends, Carlos, Ana, André and Eduardo, thank you for not giving up on asking if "is it ready yet?". My family, even those that remain only in my memory, were always with me for inspiration, support and encouragement, and must also have a place in these paragraphs – to my mother Filomena, my father Jorge, my brothers Nuno and João, my uncle and aunt Paulo and Judite, and my grandparents Jorge and Judith. To Branca for her valuable help proofreading this dissertation. Lastly, I thank FCT – Fundação para a Ciência e a Tecnologia – and ParadigmaXis S.A., for funding this work. FilipeFigueiredoCorreia Porto, Portugal, July 2014 6introduction Chapter 2 Documenting Software An overview and summary of the main topics of this chapter may be found in Figure 2.1, which can be used as a chapter roadmap. The chapter starts with a review of the role of knowledge in software development projects, its relation with information , the importance of sharing, evolving and preserving knowledge over the project’s lifetime, and what this means for software artifacts (Section 2.1). Next, it reviews approaches and tools used to handle software artifacts (i.e., source code , models and free-text documents ) and software documentation , with the intent of identifying popular documentation solutions and to look in detail into the approaches that inspire Figure 2.1:Concept map of software documentation topics. 8documenting software this research (Sections 2.2and 2.3). The remaining of the chapter is dedicated to concerns common to the design of software artifacts in general and the environments used to manage them, analyzing the several forces at stake (Sections 2.4and 2.5). 2.1Software Knowledge The notions of knowledge capture and knowledge acquisition are not always consensual and have evolved over the years [ Mul96 ]. For the purpose of this thesis, knowledge capture , or representation, is seen as the process of conveying knowledge in a medium, and, doing so, transforming and encoding it as information . On the other hand, knowledge acquisition is the process through which a human actor gains knowledge, that is, the process of learning and understanding information. This duality is depicted in Figure 2.2. Knowledge Capture Information Acquisition Figure 2.2:Knowledge capture and acquisition. Research in the cognitive science domain has dealt with several knowledge-related issues. It has allowed to identify different types of knowledge and strategies through which they are acquired, with results useful on different domains of human activity, and in particular in that of software development [Rob99,RAMF04,GR13]. Software developers are knowledge-workers, as their activity revolves around creating, distributing and applying knowledge [ Dav05 ]. In fact, software development can be said to be a knowledge-intensive activity in which there is as a progressive crystallization of knowledge with the very specific goal of ultimately obtaining instructions to be executed by a computer [ Rob99 ]. The difficult part of this process is not necessarily the production of programming language statements per se, but often the discovery of knowledge that will allow the developers to produce them. This is essentially an activity of learning, or acquiring knowledge [ Arm00 ], in which software artifacts , that are no less themselves a manifestation of knowledge, go through different levels of formality: from unstructured information, as textual documents and verbal communication, to source code statements that will be consumed by computers. software knowledge 9 2.1.1Knowledge Sharing and Preservation Software developers are usually part of a team, which means that knowledge is worked collaboratively. Being able to share it effectively among team members is a key concern, as it enables to reduce overall knowledge acquisition efforts, and helps in the process of establishing a shared understanding that allows the team to work towards the same goals. Knowledge sharing may happen by capturing it – that is, by recording it, and making it available to others – or it may happen through direct communication between peers. By capturing knowledge, we are making it available to others and our future selves. The importance of this is that ideas, in the minds of team members, are not easy to maintain. There are a number of events that may disrupt keeping knowledge by memory alone: team members leave the project, new team members appear, existing team members simply forget knowledge that they once had, and they have to recover it. Regaining knowledge may involve repeating the process that has originally led to it, which frequently has a high cost. Although this work addresses knowledge capture, it is worth noting that the value of direct communication should not be dismissed. Some development processes in fact advocate face-to-face conversation as essential to sharing knowledge on a short term effectively [ BBvB+01 , BA04 ]. As software projects tend to be increasingly distributed, software development tools and environments are also starting to include features for direct communication between peers [CHRP03,STvDC10]. 2.1.2From Knowledge to Software Artifacts A great part of the effort in the development of software is in collecting information – from verbal conversations and written materials – which is then reasoned upon, transformed, and captured into concrete software artifacts. These can be of different types, such as free-text documents1,models and source code, among others. 1 The Cambridge Dictionary defines document as "a paper or set of papers with written or printed information, especially of an official type", or "a text that is written and stored on a computer". Even though this definition reflects a common understanding of what a document is, it is very broad and not especially useful to us. In practice, we may distinguish between free-text and structured documents. While the first refers to textual contents, made mainly by non-domain-specific elements such as sections, paragraphs, lists, figures, tables, etc., the second usually refers to domain-oriented formats, such as data-centric XML dialects. In the context of this work we shall refer to documents always as sets of primarily free-text contents, unless otherwise noted. 10 documenting software Two levels of captured knowledge may be considered: information and metainformation . While the information of an artifact addresses the particular subject that the artifact intends to represent, meta-information describes information itself, contextualizing it, and conferring it additional semantics. Simply put, meta-information is information that describes information and that can be used to define an information frame to which the information obeys or should obey. Or, in other words, it can be used to add structure to the information. Different kinds of software artifacts provide different kinds and amounts of structure. While artifacts are usually bound to a certain degree of structure that cannot be changed, they frequently allow additional structural elements that are open to being authored. The process of capturing such structure is one of carefully organizing and classifying knowledge. Take source code artifacts as an example – they need to obey to the syntax of the programming language that is being used, even if developers are free to define certain structural elements, such as classes in an object-oriented language. Considering a user manual as a second example, its authors likely obey the structure of freetext, that is, they are led to use elements such as sections, paragraphs, lists, figures, etc. Furthermore, if authors wish to create a specific document template (e.g., for requirements documents), a second level of structure is made available, establishing a concrete set of sections to be used for that kind of document. Choosing a type of software artifact to capture a piece of knowledge always depends on the structure that it enables to express. For example, when creating source code, developers will try to express the solution in the programming languages that they are using, and one particular language may allow them to be more expressive in some parts of the solution. If it’s a task that they are trying to capture, recording a new task in the project planning/tracking tool may be the best choice. On some occasions, though, there will be no clear choice of a specific kind of artifact that fits the knowledge to capture, and developers will try to record it in the most suitable one that they have available. A key factor in this choice is the proximity between the information (structural and otherwise) that software artifacts allow to express, and the mental models of the knowledge to be captured. The closer they are, the easier is to capture that knowledge. The structure provided in the context of free-text documents is usually domainagnostic, enforcing only a layout form. When using a template, authors are providing an additional frame that may come closer to the information’s domain, but they may still be able to change the template to suit their needs. This shows how flexible free-text software knowledge 11 documents can be, as they allow to express virtually any topic. Despite this, free-text documents are intrinsically limited when trying to capture elaborate structures, and are sometimes combined with more specialized artifacts, in order to reach a better balance between flexibility and expressiveness [AD05b]. Expressiveness 2 , the ability to convey the intended knowledge in the most complete and concise way, is a key factor to determine how effective knowledge capture is. It’s not, however, the single criteria used by developers. As tools support the creation and use of artifacts, the choice for a specific type of artifact cannot be made without considering the maturity and functionality provided by the implied toolsets. Time (hence, cost) is also a variable to consider, and it may not be prudent to capture absolutely all the available knowledge. As a potentially time consuming activity, the effort and benefits of capturing knowledge should always be weighed [ Bri03 , Agu03 ]. 2.1.3Evolving Knowledge and Artifacts As we will address further in the following sections, software systems can hardly be seen as immutable entities. The knowledge of project stakeholders, including that of developers, about a system under construction is always subject to change. This implies that the several artifacts captured at a given moment in time may have to be adapted to new understandings. But, although knowledge may evolve, artifacts may not be easy to adapt accordingly, especially when structure changes are needed. The types of software artifacts that imply a specific structure have been conceived bearing in mind a specific type of knowledge. They are likely very expressive in that domain, but not easy to change beyond their predefined structure. The fact is that knowledge on a given subject frequently starts off as vague and only gradually becomes more concrete – only at a later time can the contents be structured effectively. If a given type of artifact is not capable of expressing additional structure, it may be impossible to update maintaining the same level of expressiveness of when it was first created. 2 The meaning of expressiveness is difficult to pin down. The definition used in this work tries to approach the intuitive meaning of expressiveness, considering it as a combination of three factors: a) The ability (or inability) to convey information in a given domain; b) The degree of unambiguity (or ambiguity) of a conveyed information; c) The concision (or verbosity) of a conveyed information. 12 documenting software 2.2Software Artifacts Software artifacts are both the products of software development and the things that developers work with. They may be themselves part of the final set of deliverables to be built; they may describe or support the process of developing software, and how it unfolds; and they are capable of describing the function and design of software, and therefore be used in the creation of other software artifacts. A large variety of software artifacts may be considered, including project plans, requirements, design diagrams, models, source code, bug reports, user stories, graphical design resources, and translation resources, among many others. Some approaches to handle them are independent from the type, but usually different types of artifacts entail specific challenges and solutions. Free-text documents account for many of these different kinds of artifacts, as they usually allow their document structure to be shaped at will, and artifacts with different structures may easily be considered to be of different types. Section 2.3will go specifically into the topic of software documentation , which may comprise not only free-text documents but also other kinds of artifacts. This section goes into more details specifically about source code and model artifacts. 2.2.1Source Code The creation of abstractions has been repeatedly used in software development as a way to move the focus from the details of the hardware to the domain of the application being built. Traditionally, this abstraction process has assumed the form of higher-level development platforms and programming languages. Source code is always the main focus of software development. Even though some approaches tend to focus on other kinds of artifacts (Literate Programming focuses on textual descriptions, Model-Driven Engineering focuses on models, etc), source code is needed to instruct machines what to do. In fact, some argue that source code is the only artifact that one can really depend on, as it doesn’t lie – if inconsistent with other artifacts, source code is the artifact to go to for the actual program behavior. Although source code can be extremely expressive in its own computation-oriented domain, the overall knowledge from which it was derived is very weakly recorded within its form. It exists only implicitly, rather than explicitly, being therefore difficult, if even possible, for the developers to reconstruct the original mental models at a later time from the source code alone [GAO95,Sta10]. software artifacts 13 2.2.2Models Models imply a higher level of abstraction than that provided by source code. A model is a simplified representation of a problem domain, capable of more accurately reflecting the developer’s knowledge in that specific domain. This means they can normally capture knowledge more explicitly than source code, even though they don’t represent every single detail of the actual problem to be solved. The creation and use of models poses numerous challenges to software developers [ SV06 , TPT09 ]. We shall look into specific issues, such as Consistency,Semantic Opacity and Abstraction Domain. Consistency Models may be used to derive other artifacts at design-time, or they can be used by applications at run-time. Each approach has different merits, and should be chosen according to the constraints at hand. The first approach has traditionally been used more extensively, although it may easily be a source of concerns if the derived artifacts are not kept in sync with their original sources. From a perspective focused on the quality of the structured contents, models are made out of information and meta-information that are expected to comply to each other. When this fails, a consistency issue may also be said to exist. These issues are discussed in a broader sense on Section 3.1. Semantic Opacity Using models implies making information obey to a semantically rich structure, but the use of modeling languages still frequently results in capturing knowledge as simple image diagrams, with the sole goal to be used for human consumption. Such artifacts can only be said to be the result of a modeling activity, and can’t truly be said to be models. This is an important distinction, since tools that handle models and those that handle (model-derived) image resources may take considerably different approaches and provide different benefits. Domain Expressiveness General purpose modeling languages, such as UML, allow to model a substantial part of software systems. UML accomplishes this by supporting several types of models, targeted at different domains – specialized models may allow a significantly 14 documenting software more expressive representation of knowledge. For this reason, the need for a higher expressiveness in a given domain may easily spawn the need for different, specialized, models. Each possibly needing supporting tools with different requirements. 2.3Approaches to Software Documentation As mentioned earlier, free-text documents are a very flexible way of recording knowledge. They allow to express almost any type of information, can have different degrees of structure and be adapted to the specific needs at hand. The primary goal of software documentation is to capture and share information about a software system in its different dimensions, including sometimes the activities that surround its making and use. We define documentation of a software system as any form of captured information about that system that may help its developers and users understand it. It serves as a communication medium between the members of a team and plays a key role in program comprehension [ VN04 ]. Furthermore, when referring to documentation we are not merely referring to (textual) documents; documentation may comprise a very heterogeneous set of artifacts, and how they are organized and combined greatly depends on the type of documentation to produce. 2.3.1Diversity of Software Documentation Authors deal with several forces when taking the decision of what kind of software documentation to produce [BKM00,HHT01,AD07,AD11]. Different writers. Documentation is produced by different participants, with different roles in the software process. Programmers, architects, project managers, product managers, technical writers, and others, may all play a part in the production of documentation, as they are the most knowledgeable in their areas of expertise. Although depending on the way the entire development process is designed, the creation of documentation is usually a collaborative activity. Different audiences. Documentation has different target audiences. The participants in the software process may be themselves the target audience of the produced documentation, as it can be others, external to the software development process. In either case, readers may possess distinct levels of knowledge, and documentation intended for them should take this fact into account. Different subjects. Documentation may be used to describe any of the facets of software. It may in fact be used to describe other artifacts, further contextualizing approaches to software documentation 15 them or using them to support more elaborate descriptions. Source code, models, and the working product itself, are some of the types of artifacts that may be the subject of documentation, or be made part of the documentation themselves. Different levels of abstraction can be touched simultaneously. Different notations. Different kinds of information are better communicated by using different representations. While for some kinds of information a textual description may be the most appropriate, other kinds of information may be better conveyed in other ways; for example, by using diagrams, or source code examples. Target audiences can also influence the choice of notation to include in a particular document, as they will understand it more easily if they are already familiar with its notation. Different forms. According to the context at hand, by balancing the previously presented alternatives, one can conceive and structure a free-text document in a way that it’s most effective. Some recurring types of document structures address typical documentation structuring needs: scenarios, design patterns and pattern languages, system overviews, user manuals, tutorials, contextual help, frequently asked questions, cookbooks, recipes, hooks and motifs, among others. 2.3.2Wikis Wikis are systems for collaboratively authoring Web pages, which makes them tools with a very wide scope of application. The first wiki was created by Ward Cunningham in 1995 3 [ Cuna ] and numerous other implementations have since spawned from the same set of founding ideas [ LC01 ], which can be summarized as a set of design principles enunciated by Cunningham [Cunb,Cun06]: Open – Should a page be found to be incomplete or poorly organized, any reader can edit it as they see fit. Incremental – Pages can cite other pages, including pages that have not been written yet. Organic – The structure and text content of the site are open to editing and evolution. Mundane – A small number of text conventions provide all necessary formatting. Universal – The mechanisms of editing and organizing are the same as those of writing, so that any writer is automatically an editor and organizer. Overt – The formatted and printed output will suggest the input required to reproduce it. 3 The first wiki engine, named WikiWikiWeb, is nowadays still on-line at http://c2.com/cgi/wiki . 22 documenting software 2.3.3Literate Programming As enunciated by its creator [ Knu84 ], the ultimate objective of Literate Programming (LP) is to make computer programs comprehensible by human beings, accomplishing this by switching the focus that is traditionally given to source code artifacts to documentation artifacts. The fundamental idea behind LP is that, when writing programs, one should not instruct a computer what to do, but rather explain to human beings what the computer will do. An alleged benefit is that programs written this way are works of literature, or works of art [ Knu84 ]. This claim may be seen with suspicion by those concerned with the practicality of this technique but, in fact, the usefulness of LP goes considerably beyond the aesthetic and literary side of documentation [Ham94]. Literate programs are built by describing pieces of the program at the same time they are developed, and by connecting them as a web of related ideas. The result is a unified document, combining several fragments (chunks) of source code and documentation, disposed not as a set of assorted blocks of information, but following a line of reasoning. Contents are arranged in the order in which they are written, improving their ability to be understood [ Knu83 , KC02 ]. This can be a benefit in what concerns readability, when compared to organizing source code according to its own structure. LP involves the use of several languages; being required at least the use of a document formatting language, a programming language, and a unification language that assists in combining the previous two. The reader’s attitude towards program understanding also has to be different; the most effective approach to understanding a literate program is not an exploratory one, but by reading the literate document, in a book form [Ham94]. Several tools following a LP approach have been developed since the concept was conceived, with varying levels of success. None, however, has reached the mass acceptance expected by some authors, given the alleged benefits [ Wyk89 , PKB04 ]. Some barriers to the adoption of LP have been the dependence of LP tools for specific programming and formatting languages [ vAK92 ], and the lack of methodological integration in the software lice-cycle [ CB91 ]. Some of the most noteworthy LP systems are WEB [ Knu83 ], CWEB [ KL02 ] and noweb [ Ram94 ]. The usage of these systems is quite similar; they supply two core operations: weave and tangle . While the former is used for generating a human readable form of the documentation, the later is used to generate the source code, in the form that is accepted by the compiler. These two kinds approaches to software documentation 23 of generated artifacts are created from the same literate document, commonly called a web document14. The weaving and tangling processes are illustrated by Figure 2.3. Lorem ipsum dolor sit amet, consectetur void main(){ } Lorem ipsum dolor sit amet, consectetur adipiscing elit void main(){ } foo.web Code chunk Text chunk tangle weave void main(){ } void main(){ } foo.java void main(){ } Lorem ipsum dolor sit amet, consectetur Lorem ipsum dolor sit amet, consectetur adipiscing elit foo.tex Lorem ipsum dolor sit amet, consectetur foo.pdf foo.class java compiler pdflapex processor LP engine Figure 2.3:Literate programming operations. Since its first appearance, research has shown Literate Programming to possess some additional virtues, as well as drawbacks. The following list presents the features of LP as been seen by its creator and according to subsequent research on this topic. Verisimilitude. Documentation and source code are written and stored together, physically close to each other (i.e., as consecutive blocks in a file), and written at the same time [ Wyk90 ]. This has the advantage of easing the production and maintenance of documentation, as the documentation is presented very closely, whenever the source code is modified. The resulting documentation may thus more easily be kept consistent. Arrangement. Using a traditional approach, source code is organized in the way in which it is accepted by the compiler, according with the intent of the developer, and using the programming language’s syntax. The order in which the code is kept usually isn’t the same by which it was written, and neither can it be easily re-organized to better expose the reasoning that lead to its creation. LP allows (in fact, it requires) the re-organization of source code, so that it accompanies the line of thought of the documentation, following the psychological arrangement in which it may be better understood. 14 No relation to the World Wide Web (WWW) exists, as the concept of Literate Programming predates the WWW in at least a decade. These web documents shouldn’t be confused with the hypertext documents that form the WWW. 24 documenting software Despite the benefits, this feature also implies some important liabilities. The developer should only interact with the rearranged document, but this makes debugging more difficult. The compiler acts upon the tangled source code, and usually reports errors and warnings by referring to line numbers, but these will not match the line numbers of the web document that the developer has contact with. This happens because the developer and the compiler have different views of the code [ KC02 ] leading to an effect commonly referred to as referential opacity [ Thi86 ]. It frequently forces the developer to inspect the tangled source code files whenever he needs to match compiler-reported line numbers to a particular instruction, hindering his efficiency. This suggests a leaking abstraction [ Spo02 ] yet to be resolved in traditional LP tools. As with most powerful features, being able to create and rearrange source code and documentation fragments can be easily misused. One such case happens by abandoning the programming language’s structuring mechanisms in favor of the chunking mechanism provided by LP. Chunks don’t provide scope and supply very weak interfaces, hindering reusability and maintainability if used to replace language constructs. Using LP does not reduce the need to use the appropriate abstractions provided by the programming language [Ham94]. Another issue is that a single psychological arrangement may not be enough. Effective communication with different audiences may be better achieved using different arrangements. Readability. By supporting and automating the creation of indexes, table of contents and cross-references, and by pretty-printing source code, LP tools allow a great readability of documentation on paper [ Ham94 ]. On-screen, however, other approaches are needed to ensure a good readability. On-screen readability is important in several situations, and is fundamental during the development process [Kna96,Agu03]. Consistency. Software documentation is as valuable as much as it reflects reality, thus the importance of keeping it up to date, and consistent. As mentioned before, LP allows an easier maintenance of consistency, by keeping related pieces of knowledge physically close to each other – i.e., documentation for a given block of code is kept in the same file and in the same sequence as that block of code. approaches to software documentation 25 In spite of this benefit, consistency is still kept by visual inspection, and still requires such effort from developers, as other techniques do. Contents integration. When using LP, documentation and source code are seen as a whole. This means they are produced at the same time, and are explicitly interrelated. As mentioned before, contents are organized sequentially, as a single document. In fact, several types of content, other than code and textual descriptions, can be integrated this way, such as figures [ SC93 ], models [ AEQ99 ], formal specifications [ And01 ], etc. Several types of documentation can be produced following this approach, including external documentation, as user manuals [Thi86]. One should use the kind of artifact that is most appropriate to convey each piece of knowledge. A particular kind of artifact can be said to be more appropriate if it is more expressive than the alternatives (Section 2.1.2), or because it is simply of better use – e.g., if they are required as input for other software engineering tasks and tools. Free-text has its place and is frequently used to capture what can’t be expressed by other artifacts, to achieve a comprehensive documentation. An example is the use of free-text to overcome the trivialization of requirements in UML diagrams [ AEQ99 ] – relative importance or priority of different requirements are not expressed in a UML diagram, but can be added by using textual descriptions. Tools integration. Although being desirable that LP tools are both general and powerful, these objectives are not easy to reconcile [ Thi86 ]. This poses a difficult problem, as two of the reasons that are most frequently pointed out for the lack of acceptance of LP are its lack of generality (e.g., its binding to a specific programming language) [ vAK92 ] and its lack of integration with modern environments. These issues have both been addressed before (the first one with partial success [Ram94]), but are still major barriers to acceptance. Appropriate tools support would allow to minor the debugging problem (see the arrangement issue discussed above), and would make possible to give LP environments the same capabilities of modern IDEs in what concerns code navigation, refactoring and on-screen readability, among others. Quality. While it is as easy to write poor documentation using LP tools as is by using traditional tools, when using LP tools the developer is more aware of a reader, to whom the code is targeted. He will, thus, be more aware of the 26 documenting software importance of creating such documentation, and more inclined to writing it in a greater quantity and with a greater quality [Ham94]. Concerning code quality, it is also claimed that literate programs have fewer bugs, due to the extra attention that source code receives. By carefully explaining the objectives of a chunk of code, one can produce code of better quality, as errors will more easily come to one’s attention [HS98]. Overhead. LP introduces a certain overhead. It may be made less noticeable with the help of the right tools and on small projects, but it’s never negligible, especially for large or distributed projects [vAK92,Agu03]. The definition of what LP is has evolved over the years, and different people tend to emphasize different features when describing it [ Knu83 , Thi86 , Ham94 , Smi01 , Agu03 , PKB04 ]. A trend may be observed though: in its latest incarnations LP has moved its focus, from a way to produce software documentation as a work of literature, to a way of organizing, integrating, recombining, and maintaining the consistency of contents. The evolution from the initial concept has spawn some variants, such as Literate Modeling,Reverse Literate Programming and Theme-based Literate Programming, and inspired the creation of other documentation techniques, such as Code Annotations and Elucidative Programming. Each of these approaches can be said to establish a different model for how to structure and organize software. They are detailed in the following sections. Literate Modeling The concept of Literate Modeling was introduced by Jim Arlow [ AEQ99 , AN04 , Arl06 ] as a natural evolution in the use of LP. With models playing an increasingly important role in software projects there’s the need to include them as first class artifacts. Modeldriven development takes models as a key part to the creation of a working system, but they may also be used simply as a way to document an implementation [Tho06]. Although, at its first form, Literate Modeling was focused on UML models, it has introduced the idea that artifacts beyond code and textual descriptions may also be used with Literate Programming. Reverse Literate Programming One of the benefits of using LP is having documentation with a good readability on paper, but on-screen readability is not one of its strengths. On the screen, developers approaches to software documentation 27 do not read code in a sequential way, and rather do it selectively, like an encyclopedia, using source code structures and control flow to navigate to the intended information. Reverse Literate Programming [ Kna96 ] has much in common with Literate Programming, but relies on the active text elements provided by an integrated development environment, like folding, linking and bookmarking. Active text elements are used to support several concepts of LP, like presented in Table 2.1, which was taken from [Kna96]. Active Text Literate Program Fold element Section with a macro definition Collapsed source of fold element Source code part Comment on fold element Documentation part Link element Relation between sections Table 2.1:Representation of active text elements as a literate program. This approach provides a much better on-screen readability and the use of all the interactive features of an IDE, while still allowing to produce a printable (web-like) literate document. One other declared advantage is that the tangle operation is no longer necessary, as active text elements are encoded in the source code as special characters, which are ignored by the compiler. In a way, the concept of LP is reversed, as the literate document is assembled when needed, and the main artifact is the source code. There is, however, a price to pay for a Reverse Literate Programming approach. Documentation and code may no longer be arranged in whatever order may be the best for program comprehension, being included in the printable literate document following the predefined sequence of source code. Theme-based Literate Programming The need for multiple arrangements of contents to co-exist was mentioned above, as a way to target different audiences. This is the main issue addressed by Theme-based Literate Programming (TBLP) [ KC02 ]. TBLP allows the combination of different types of chunks (code segments, figures, textual descriptions, unit tests, etc) by relating them with different types of connections. This combination of chunks is made according to atheme, which is a way of sequentially organizing those chunks according to a specific psychological order. Themes are a way of providing multiple documentation views of a system. 28 documenting software Seeing all contents as chunks allows arranging documentation in any order, as it allows to do the same for source code, and makes it easier to decouple textual descriptions from the code fragments they describe. This also means, however, that verisimilitude isn’t always achieved, with disadvantages towards consistency-keeping tasks. This liability may be minimized when using live and semantically rich relationships between artifacts [ KC02 ], so that changes in one artifact highlight the need for updates on the ones that relate to it, allowing to keep the overall system documented and consistent. 2.3.4Code Annotations This technique was initially inspired in LP, as documentation is generated from a unified representation of textual descriptions and source code. However, it is also fundamentally different from LP, as textual descriptions exist in the form of source code comments. This means that the unified representation of textual descriptions and source code is itself valid and compilable source code, avoiding an additional tangle phase. When comparing to LP, it is also important to highlight that writing code annotations is not the same as writing free-text contents, as they depend on the structure of source code files. As such, it misses one of the main benefits of LP, which is the possibility of reordering documentation according to an intended psychological arrangement. Code annotations are primarily used for creating API documentation, and don’t address all the issues that LP tries to address. Having said this, it has shown to be quite successful in this niche, and has helped increasing the awareness on the need for documentation, and showing how it can enhance program comprehension. The widespread use of this approach has been much the merit of Javadoc [ Fri95 ], which is a tool supporting this functionality for the Java programming language. It is one of the first known uses of the technique, along with Autoduck [ Art00 ], a tool created in 1993, which supports code annotations in C++, and Doxygen [ vH97 ], a tool initially released in 1997, and that now supports a wide range of languages, including Java, C#, Objective-C and Python, among others. approaches to software documentation 29 2.3.5Elucidative Programming Although this technique was inspired by LP, there are fundamental differences between the two. Elucidative Programming (EP) addresses the documentation issue from a more pragmatic standpoint; while LP satisfies the need of publishing programs as technical documents, EP tries to address the everyday maintenance needs and program comprehension [ Nø00 ]. The main goal of EP is to directly support software maintenance tasks, by providing documentation that can be used effectively from within the development environment. In fact most EP features may only be streamlined with such an integrated environment [Ves03]. Unlike LP, EP allows attaching explanations to a program without modifying the source code, thus not requiring the additional tangle and weave phases. Furthermore, EP does not demand the re-arrangement of source code; instead, it allows the creation of documentation following two different styles [VN02]: Linked. This style provides a mechanism for defining bi-directional relations between source code and documentation sections. These relations are presented as hyperlinks; while browsing the code, the developer can at any time explore the relations by asking for the related documentation, and vice-versa. In an EP approach, documentation and source code exist as separate entities and are connected by referencing syntactical elements of the programming language or through special code markers embedded in comments. Inlined. The ultimate goal of EP is not the creation of a printable document that sequentially explains the system. However, some kinds of documentation don’t fit a linked model of reading, being more adequate a sequential presentation of contents (e.g., tutorials, cookbooks, etc). EP’s inlined documentation was conceived to handle these kinds of documents, allowing to author textual descriptions and to combine them with code fragments, following a line of thought. Source code fragments are not merely copied to the documentation, but rather referenced, so that no inconsistencies may be introduced between virtually identical artifacts. When using an inlined style it is still possible to benefit from some linked-style features. Words used in textual descriptions can be turned into hyperlinks that reference specific code elements. References between artifacts can potentially be used, not only for user navigation purposes, but also for error checking features that detect inconsistencies caused by the evolution of source code or other software development artifact. 30 documenting software Another alternative to handle inconsistencies was approached by an extension to EP that considers the possibility of different paces on the evolutions of source code and documentation, and embraces this reality as a common scenario to be dealt with, instead of prevented [VN05]. 2.4Designing Software Artifacts As we have suggested in Sections 2.2and 2.3, different types of software artifacts often imply different challenges and require different approaches. Yet, they do share some common concerns that shape their design and that should be taken into account when creating new solutions. 2.4.1Supporting Medium The most common medium used to support software artifacts is the filesystem. A filesystem allows handling different things (e.g., files), of different types (i.e., file formats), leaving to applications the way such things are actually encoded. A vast amount of tools has been designed around this assumption, with version-control systems and source code editors as some of the most common. But other supporting mediums may be found beyond files. Some software artifacts have also been managed and supported by Web-based information systems, and this approach has become increasingly popular. Wiki engines are a good example of these systems; artifacts (e.g., wiki pages) are not identified by a file path but by a Web address, and the Web browser is the primary tool to handle them. Some of such Webbased environments integrate distinct tools for collaboratively developing software and are frequently referred to as Software Forges [ REM+09 ]. They are described in detail in Section 2.5. 2.4.2Single and Multiple Source Approaches File-based contents can be classified as single-source or multiple-source, depending if several kinds of content (e.g., source code and textual descriptions) are encoded in a same file, or if each type of content is created as an independent artifact. Traditional approaches keep documentation and source code as autonomous artifacts, in what concerns their writing, reading and storage. This may lead to consistency designing software artifacts 31 problems, as the need to replicate content is likely to manifest (e.g., to copy source code fragments to documentation). Single-source approaches [ Smi01 , Agu03 ] join different kinds of content in the same file, and are effective ways of solving this problem. They do, however, bring other difficulties; particularly if joining distinct formats in the same file leads to new, mixed or non-standard formats that require specialized tools. Multiple-source approaches [ Smi01 , Agu03 ] keep different kinds of artifacts in separate files. In order to keep consistency, they allow the creation of relations between artifacts, and development environments simulate verisimilitude by presenting related artifacts close to one another. This dichotomy between single-source and multiple-source approaches is most used in the context of documentation contents – the original concept of LP and Code Annotations are forms of single-source approaches, while TBLP and EP are multiplesource approaches. 2.4.3Structured Contents Free-text documents are found most often playing a support role during software development, and other artifacts that are more specialized, such as source code and models, play the central role. From a knowledge capture perspective, each of these different specialized software artifacts implies a particular information structure and therefore is limited to the kinds of subjects that it was designed to convey 15 . On the other hand, they enable a great concision and objectivity when used to convey those particular subjects. To capture information is also to capture the way it’s structured, and two kinds of structure may be considered: the internal structure of an artifact and the structure used to connect and organize artifacts. Capturing the later is as relevant as capturing the artifacts themselves, and plays an important role in assisting information consumers in locating the specific pieces of knowledge they need. Having explicit relations between different artifacts is also what allows traceability. Relations may connect artifacts that follow each other temporally (predecessor/successor); connect artifacts addressing the same information under different perspectives; or connect artifacts addressing the same information at different abstraction levels, from 15 Many other kinds of artifacts could be provided as example. Some of them may make more apparent the existence of a structure to which the contents comply, such as structured documents using any dialect of the eXtended Markup Language (XML), using the JavaScript Object Notation (JSON), or using any other of the many markup languages that are today in common use. 38 software evolution assess the current syntactic consistency of source code. Some toolsets allow to assess the consistency of documents too, but obviously only for the types of structure that documents allow to express. The domain-agnostic structure of free-text documents, while being one of their greatest strengths, doesn’t support assessing if their contents are consistent beyond their layout form. This is a fundamental issue, as the value of documentation is as great as its ability to convey accurate information, but one of the greatest costs, when maintaining documentation for a large system, is ensuring it is in-sync with the artifacts it describes. Refactoring is one of the techniques that support software evolution. A complementary approach to automating evolution tasks is the use of migrations [ FCW08 , Rub ]. Migrations consist in allowing one to express and run a set of transformations to a system’s schema and data, making it transition from a version to another. They stem from the need to evolve the domain model (or schema) of a system, which, in turn, brings the need to evolve it’s underlying data. The same concept has been used in the context of object-oriented and relational database systems [WE00,RL04]. When it comes to filesystem-based artifacts, version control systems (VCSs) are extensively used. They allow developers to take snapshots of a given set of artifacts throughout the project, so that it’s possible to keep track of how they evolve. They also serve as a collaborative platform, allowing developers to deal with new versions of an artifact (or set of artifacts) produced concurrently by different team members. Distributed Version Control Systems (DVCSs) are an interesting variant of VCSs, in that they are especially suited for distributed and disconnected environments, and don’t require a version control server to be permanently available to produce new versions. As introduced in Section 2.4.1, software artifacts can be based on mediums other than files. Web-based software artifacts are increasingly popular, even though the wide range of filesystem-based tools are often not suitable to handle them. Integrated development environments have traditionally handled files, but have started supporting more and more Web-based artifacts too [ Myl , Mer ]. However, some tools are difficult to transition to the Web environment, and the impedance mismatch between artifacts of these two different environments makes it harder to integrate them. VCSs have proved to be valuable sources of information in the study of how software evolves [ Koc05 , KCM07 , DGLP08 , GG08 ]. They have allowed to identify different ways in which software artifacts evolve, and given hints on how to improve support for the activities that actually take place during software development. adaptive software 39 Software evolution may also be classified according to how the software is being used when it takes place [MBZR03]: Design-time evolution happens when the software is not itself running. It corresponds to the traditional way of evolving software, in which a person or a team introduces changes to the software, normally by producing programming language statements, recompiling it, configuring and deploying it, hence replacing the previous version of the system. Run-time evolution happens when the software being changed is being ran when the change takes place. It usually happens when reflective, or adaptive systems are used. The system alters itself based on inputs of its users or the environment. Concrete techniques for run-time software evolution are addressed in greater detail in the following section as they play an important part in this work. 3.2Adaptive Software Adaptive systems are those that can be efficiently molded according to changed circumstances [ AG05 ]. Researchers use the term adaptive software in a broad context. Among other possibilities, it may be used to refer specifically to software that a) adapts itself (self-adaptive software), to b) software that can be easily changed to accommodate new requirements, or to c) software able to satisfy many different user or market needs [ AG05 ]. The present section doesn’t try to address all these facets, and focuses on the third use of the term, which is the most meaningful one for the context of this work. A quick overview of this area of study and, in particular, of the topics described in this section is depicted in Figure 3.2. One way of creating an adaptive system is through a meta-architecture : one in which the program manipulates itself as if it were data. Meta-architectures usually describe the system’s domain, or a part of it, by establishing different levels of (meta-)data that comply to each other. Meta-modeling can be used to create such descriptions. Dynamic meta-modeling approaches, such as those using the Adaptive Object-Model (AOM) architectural design pattern [ YBJ01 ], allow systems to be adapted at runtime – the system interprets a high-level description of the domain, and adapts its behavior to any changes introduced to that description. 40 software evolution Figure 3.2:Concept map of adaptive software topics. An alternative to dynamic approaches is generative programming : the high-level description of a system is used to automatically create code that can be executed, or a code skeleton that will be further completed by the developers. Adaptability is thus introduced at compile-time, requiring a full generation/compilation cycle. These notions are explored in the following sections to a greater extent. 3.2.1Domain-driven Design The purpose of developing software is, ultimately, to fulfill its users needs, but to actually understand these needs developers may have to delve into those user’s knowledge domain, with all that it entails. Domain-driven design (DDD) is an approach to the development of software that strives to closely connect the implementation to the business domain. As Eric Evans puts it, in the book Domain-Driven Design [Eva03]: "To create software that is valuably involved in users’ activities, a development team must bring to bear a body of knowledge related to those activities. The breadth of knowledge required can be daunting. The volume and complexity of information can be overwhelming. Models are tools for grappling with this overload. A model is a selectively simplified and consciously structured form of knowledge. An appropriate model makes sense of information and focuses it on a problem." This approach fosters creativity and collaboration between domain experts and developers, making models the center artifacts and part of the language used during those activities. These models are not an accurate description of reality , but rather adaptive software 41 abstractions, representing of the system’s domain with the very focused goal of feeding the development process. 3.2.2Adaptability and Variability The general definition of adaptive systems as those whose behavior can be easily changed according to new realities and needs meets the work of several authors, although establishing what the term adaptability means hasn’t always been consensual [AG05]. Variability is not an exact equivalent to adaptability, although they are very close concepts, that overlap. The notion of variability was born on the area of Software Product Lines (SPLs), and has been defined as a property that allows changing or configuring a system so that it may be used in different contexts [ vGBS01 ]. Systems provide (and constraint) variability through variation points, which are the points of the software that can be changed, and that define to which degree it can be customized. The key difference is that when addressing the topic of adaptability there is a whole area of the system that can be changed according to the developer’s (or user’s) needs, and when addressing variability, the focus is on a constrained set of customizations points. 3.2.3Meta-modeling Higher-level programming languages have increasingly supported developers on focusing on the design of the software being built, rather that its implementation details [ Sch06 ]. Model-driven engineering (MDE) continues this trend, with the objective to further reduce the gap between specification and implementation artifacts, by creating models that abstract several facets of software development. It provides benefits such as increased reuse, fewer bugs, shorter time-to-market and systems that are simpler to understand [RFBO01]. But more than supporting the creation of models, MDE is an approach to metamodeling : it supports the use of models to specify other models – it comprises the analysis, construction and development of the frames, rules and constraints to modeling a predefined class of problems [SV06]. A wide range of modeling languages exist, from textual domain specific languages (DSLs) built for a very specific purpose, to graphical languages like the Unified Modeling Language (UML). UML is currently one of the most widespread modeling 42 software evolution languages. It can be used to specify software systems through a graphical notation, allowing to express elements such as classes, actors, activities and components, among others. It is based on the infrastructure provided by the Meta Object Facility (MOF) [ OMG ], which defines a meta-modeling architecture consisting of four modeling levels, each conforming to the one above – M0, M1, M2and M3, with M0corresponding to the data, M1to the model, M2to the meta-model and M3to the meta-meta-model, which is compliant with itself. An example of these layers of abstraction is depicted in Figure 3.3. M2 M1 M0 ClassAttribute ‹‹instanceOf››‹‹instanceOf›› M3 Class Instance ‹‹instanceOf››‹‹instanceOf››‹‹instanceOf›› classifier +title: string Video title = "The Matrix" : Video ‹‹snapshot›› ‹‹instanceOf›› The Matrix ‹‹instanceOf›› ‹‹instanceOf›› Figure 3.3: MOF’s abstraction layers, adapted from Adaptive Object-Models: a Research Roadmap [FCAF10]. Due to the high level of abstraction provided by MDE approaches, they can be wielded to express any desired topic, and to structure information according to an intended domain. 3.2.4Flexible Modeling The kinds of artifacts and techniques described in the Software Documentation chapter (Sections 2.2and 2.3) allow different levels of expressiveness. Free-text documents allow to convey information in a wide range of domains but usually implies a certain ambiguity and verbosity. On the other end of the spectrum, source code statements can be quite unambiguous and terse, but they can only convey information on the domain of computations. A lot of software artifacts lie somewhere between these two extremes. Being able to extensively and accurately capture their knowledge using only source code and the occasional comments can still leave a lot to be desired, in spite of all the adaptive software 43 advances in what concerns programming, domain-specific languages and modeling [Wir08]. Even experienced developers start by recording ideas in an unstructured and informal way, and only gradually are able to create more specialized software artifacts, like models and source code. A particularly important challenge is to support the process of moving back and forth between these different degrees of structure. This subject has gathered some attention, as more researchers acknowledge its importance, and some works have been published under the topic of flexible modeling tools [ OvdHS09 , KOvdHS10,OvdHS+10a,OvdHS+11]. We will briefly examine two platforms that use flexible modeling approaches: Architects’ Workbench and Business insight toolkit. Architects’ Workbench The Architects’ Workbench (AWB) is an Eclipse-based tool that has the general goal of supporting the process of architectural thinking and modeling. Its objective is "to balance formalism and freedom, while helping [architects] transform unstructured information into sufficiently formal work products"[ ABK+06 ]. When using AWB, software architects usually begin with a combination of informal free-text documents and notes from meetings with stakeholders, and create model elements directly from the text. Architects can then navigate between these elements and the text, bidirectionally. To support creative thinking, refinement of the model can be deferred to a later time. It can be refactored as understanding of the domain improves, either through form-based user-interface or through customized diagram editors. As explained by Kimelman et al [KH11]: "[...] even just that small amount of structure requires some decision by the user concerning the visual structure or appearance of each fragment prior to it being placed onto the canvas, and it requires some thought by the user to accomplish the necessary keyboard actions, mouse actions, or gestures. That "second order" thought is an impediment – it disrupts the primary train of thought, it is an obstruction that impedes the flow of thoughts and ideas into artifacts, and it can stifle creativity... even for many experts." Business Insight Toolkit The Business Insight Toolkit (BITKit) is a standalone modeling tool that takes inspiration from AWB [ OBA+09 ]. Developers use BITKit to sketch diagrams without a 44 software evolution formally defined semantics, and can afterwards map them to new or existing elements of a domain model. Like AWB, it has a forgiving approach to the creation of models: instead of using the meta-model as a straitjacket that establishes what can be expressed at the model level, it uses it as a guiding aid, which supports the creation of the model but allows developers to deviate when needed. Comparing to AWB, BITKit doesn’t take text as a starting point, but provides better diagramming tools and allows more flexibility at the meta-model level. While AWB needs the system to be reconfigured to introduce changes to the meta-model, BITKit allows to derive the meta-model from the model data, effectively providing guidance in both directions – the creation of the model can be driven by the existing meta-model, and the creation of the meta-model can be driven by the existing model [OBS+10]. 3.2.5Adaptive Object-Models The Adaptive Object-Model (AOM) is an architectural design pattern 1 that allows end-users to manipulate the domain model underlying the software system. It may provide this kind of flexibility for the entire domain model or for a selected part of it. This architecture makes extensive use of notions from object-orientation and meta-modeling, supporting reflection and runtime adaptivity, and often relying on domain-specific languages (DSLs). It can be said that it is an architecture focused on embracing change of the system’s problem domain. Developers sometimes converge to this architectural design pattern by systematically improving their reuse strategies and searching for higher levels of abstraction in object-oriented designs. It frequently emerges by making domain-related data structures into parameters of the system, which can be configured according to user’s needs. Some parts of the system are this way turned into an interpreter, and the system’s behavior is decided at run-time from the provided parameters. As more domain elements are parameterized, a model starts taking shape, and changing such model, makes the system follow a different domain. This kind of systems has been documented through design patterns, with the objective of creating a pattern language for AOMs. Figure 3.4was adapted from other works [WYWJ07,WYW07,Fer10] and depicts a pattern-map of that language. Without going into detail about the objective of each of them, the following four 1 Apattern is a recognized solution to a recurring problem that can be described in terms of a context of applicability and a set of forces that shape the solution. The notion of pattern is described in greater detail in the Patterns Catalog (Chapter 5). adaptive software 45 Type Square Dynamic Hooks Strategy Type Cube Rule Object Rule Engine Interpreter Builder Editor / Visual Language AOM Builder Dependency Injection Dynamic Factory Type Object Properties Accountability Null Object Value Object Smart Variables Entity View Property Renderer Dynamic Views GUI Workflow manages may use may use coordinates usesusesuses supports supports renders manages process extends manages extends extends describes extends uses uses supports uses uses controls extends Caching Context Object Behavioral GUI Core Creational Instrumental Versioning History helps Bootstrapping renders Figure 3.4:Pattern map of the adaptive object-models patterns as defined by previous works. design patterns are the most important to understand the structural aspects of an AOM: Type Object – Decouples instances from their classes so that those classes can be implemented as instances of a class. Type Object allows new “classes” to be created dynamically at runtime, lets a system provide its own type-checking rules, and can lead to simpler, smaller systems [JW97,YBJ01]. Property – This design pattern gives a different solution to class attributes. Instead of being directly created as several class variables, attributes are kept in a collection, and stored as a single class variable. This makes it possible for different instances, of the same class, to have different attributes [YFRT98,YBJ01]. Type Square – The combined application of the Type Object and Property result in the Type Square [ YBJ01 ]. Its name comes from the resulting layout when represented in class diagram, with the classes Entity,Entity Type,Attribute and Attribute Type. Accountability – This design pattern is used to represent different relations between parties [ Fow97 , Ars00 ], using an AccountabilityType to distinguish between different kinds of relation. 46 software evolution Chapter 4 Research Problem and Strategy This work addresses knowledge capture and acquisition in the context of software development with the overarching goal of improving software documentation. To do this it focuses on the objects of capture and acquisition activities – software artifacts. In particular, it focuses on software artifacts that may be used as documentation and the approaches and tools used to handle them. In Chapters 2and 3are identified relevant documentation approaches and tools and considerations are made about their designs, with a very special attention to the developers’ need to evolve knowledge and the concrete artifacts that capture it. This chapter puts the research into a more specific context (Section 4.1), motivates and describes the research problems (Sections 4.2and 4.3), presents the thesis statement and its decomposition into specific research issues (Sections 4.4and 4.5) and describes the outcomes and the validation that are goals of this work (Sections 4.6and 4.7). 4.1Context Overview Software artifacts are often captured as standalone files, which are sometimes managed by integrated development environments; other times they are captured using webbased systems, such as software forges. These platforms complement each other – while integrated development environments focus primarily on the creation of source code, software forges have been difficult to match in their support for collaboration and ability to cross-reference several kinds of artifacts. But regardless of the medium, artifacts are a form of captured (i.e., recorded) knowledge. While new knowledge may easily be created and evolved in the minds of the project’s team members, artifacts hardly capture all of it. Developers are not always 54 research problem and strategy text documents by allowing to maintain their consistency and classification with more economy and quality. This explanation uses terms that may be subject to different interpretations and therefore should be clarified: What should be understood by contents and information? Contents and information are used interchangeably in this dissertation to refer to a manifestation of knowledge after it has been recorded in a (usually digital) medium. What should be understood by software artifacts? A recorded, identifiable, piece of knowledge about a software project and/or that results directly from the act of developing software. What should be understood by using or consuming information? Acquiring information from a medium, usually with the goal of actively employing it for some purpose. In other words, reconstructing mental models, transforming information back to knowledge. What should be understood by creating and evolving information? Creating information is to capture knowledge into a medium from existing mental models. To change or to add to that information (e.g., so as to capture different or additional knowledge) is to evolve it beyond its current complexion. What should be understood by a medium-to-large software project? A software development endeavor made non-trivial by the quantity and complexity of information involved, which often correlates with a larger number of team members. What should be understood by information structure? Structure in this context refers to the composition of different information fragments in a determined way, often in terms of the subjects that the information addresses. What should be understood by consistency? The quality of being in agreement with something else. Maintaining the consistency of two related pieces of information is to ensure they don’t convey contradictory ideas. specific research issues 55 What should be understood by classification? The organization of information into groups or classes, according to their common natures or subjects. 4.5Specific Research Issues The thesis and the concerns introduced so far may be broken down into the more specific issues that we will address next and that were used to drive the validation of the approach. Together, these issues establish a direction for this and future research on software documentation and on the Adaptive Software Artifacts approach in particular. While issue I1focuses on knowledge acquisition (i.e., information consumption), I2 focuses on knowledge capture (i.e., information creation and evolution). Both of these issues are further decomposed into the three concerns described in Section 4.3. I1. Efficiency of Knowledge Acquisition. Do developers spend less time acquiring knowledge from the contents? I1.1. Efficiency of acquiring information structure. Do developers spend less time understanding the contents? I1.2. Consistency of the contents. Are resulting contents more consistent? I1.3. Quality of the classification scheme. Do developers spend less time searching for the contents they need? I2. Efficiency of Knowledge Capture. Do developers spend less time capturing contents? I2.1. Efficiency of expressing information structure. Do developers spend less time capturing the contents? I2.2. Economy of consistency maintenance. Do developers spend less time doing consistency maintenance? I2.3. Economy of classification scheme maintenance. Do developers spend less time maintaining a classification scheme? 56 research problem and strategy 4.6Research Outcomes This research produced four main outcomes in the path of pursuing answers to the above issues. These contributions are depicted in Figure 4.2and include, a) a patterns catalog, b) an approach to software documentation, c) a reference architecture and d) a statistical experiment. Figure 4.2:Concept map of the research outcomes and of how they relate. The creation of the patterns catalog was done throughout the research, by mining the patterns from the literature and the authors’ personal experience in the area. They served to formalize several good practices and designs surrounding software documentation, information classification, flexible modeling tools and adaptive objectmodels. The resulting patterns are presented in Chapter 5. A new software documentation approach was defined – the Adaptive Software Artifacts approach – that embodies some of the good practices and designs that were documented as patterns and tries to address the key concerns of this research. The approach is described in Chapter 6. A software documentation tool was engineered as a plugin for a Web-based environment for software development. It uses some of the more technical solutions that were described as patterns, and proves the concept of the approach. The design of the plugin is detailed in Chapter 7and is available for use as a reference architecture and implementation by other developers. Furthermore, the plugin was used in a user-study with the goal of experimentally validating some of the benefits of the approach. The design of the statistical experiment and its results are detailed in Chapter 8. validation methods 57 4.7Validation Methods It may be argued that research in the social sciences is necessarily qualitative, as the data that it produces is mostly non-numeric in nature, even if the conclusions drawn from such data may sometimes be quantified. However, this view is not consensual – Goertz and Mahoney go to the extent of considering that these views stem from two fundamentally different cultures behind the dichotomy of quantitative and qualitative research, and that the first is mainly based in inferential statistics, whereas qualitative research is grounded in logic and set theory [ GM12 , p. 2]. Throughout this thesis we will often refer to the collected data as quantitative despite its origins on human activity. Software engineering research goes beyond concrete implementations and tools and it too needs to consider the whole context in which software development takes place, and the human and social aspects that underly it. Given the amount of possibilities that this implies, software engineering is an area where applying quantitative methods is often difficult. One of the approaches used by the author to validate his thesis consists of a statistical experiment, pulling this work closer to what is often referred to as quantitative research. Statistical experiments measure the causal link between an independent variable and the phenomenon under study by ensuring that the remaining variables are controlled – that is, that they are kept close to constant using repetition,randomized assignment and averaging [ USE , DV99 ]. They generate quantitative data that is the focus of a statistical analysis. As suggested before, experiments are more difficult to conduct in software engineering than in other areas in which the independent variables are easier to control. They are, however, the best way to assess cause and effect relationships. The method used in this work can be denoted as a quasi-experiment, implying that there is not a full control over some independent variables, but otherwise follow the same rules as traditional experiments [ CJMS10 ]. Chapter 8fully details how the experiment was designed and executed. It focused specifically on those issues related with knowledge acquisition, rather than trying to address all of them as they were introduced in Section 4.5. 58 research problem and strategy 4.8Summary Software development tools support the evolution of artifacts to some extent, but they normally assume that artifacts have a fixed format. Free-text documents don’t force any domain structure, and for that reason, are very flexible when capturing and evolving contents. On the other hand, the lack of domain structure makes them less expressive and thus more laborious to produce and harder to automatically process. Other software artifacts depend on rich domain-oriented structures and thus support a greater expressiveness. However, this means their creation and evolution is bound to structural constraints, which may be too strict and make it impossible to mold the artifacts to new realities. Artifacts that could be easily evolved throughout the project’s lifetime, in what concerns both their contents and structure would bring benefits to knowledge capture activities. Information usually first appears as informal and only gradually becomes more structured and is captured into richer artifacts. A good example is how requirements are often first captured as descriptive free-text documents and, only afterwards, materialize as models, tasks, and source code, among others. The thesis introduced in this chapter looks into the acquisition and the capture of knowledge through the prism of three main concerns – a) the expression of information structure, b) the maintenance of consistency and c) the classification of the contents. This translates into the eight main research issues and sub-issues that are used in the validation of this work. The four main outcomes of this research complement each other: the reference architecture and implementation shows the practicability of the defined approach and they both are supported by the solutions described in the patterns catalog; the statistical experiment uses the reference architecture and implementation in a user-study to validate some of the benefits of the approach. Chapter 5 Patterns Catalog We have used patterns to document key solutions in topics that have shown to be of interest to this research. Some of these solutions were identified by the author while others had already been identified before, but had not been captured as consistently using patterns. Before delving into the patterns themselves (Sections 5.4to 5.7), we will explain the essence of what patterns are and their role in this research (Section 5.1), what pattern form we have used (Section 5.2) and which are the main topics that these patterns aim to address (Section 5.3). 5.1Patterns in Research Patterns are general reusable solutions to commonly occurring problems within given contexts. The notion was created in the field of architecture by Christopher Alexander [ Ale77 , Ale79], who originally defined the concept as: "[...] a three-part rule, which expresses a relation between a certain context, a problem, and a solution. As an element in the world, each pattern is a relationship between a certain context, a certain system of forces which occurs repeatedly in that context, and a certain spatial configuration which allows these forces to resolve themselves. As an element of language, a pattern is an instruction, which shows how this spatial configuration can be used, over and over again, to resolve the given system of forces, wherever the context makes it relevant." In spite of its roots, this notion can be (and has been) applied to other domains in which solutions need to be designed, given a specific problem and context. Experts in 60 patterns catalog many areas, including software engineering, think in terms of such problem-solution pairs when faced with new challenges [BMR+96]. Patterns are not invented. They distill knowledge in a given domain and, in doing so, allow that knowledge to be reused and provide a common vocabulary for understanding and communicating design principles [ GHJV95 ]. They are highly reusable because the solutions that they provide are abstract – they may be applied a million times over, without ever doing it the same way twice [Ale77]. When organized in a collection, patterns can provide a useful repository of knowledge. In its simplest form, a set of patterns organized according to a similar context or purpose can be said to be a patterns catalog. The patterns in a catalog don’t necessarily have to work together in addressing a domain. When the relations between them are made explicit, when they consider the impact that a given solution has on other solution spaces and when they address the domain at different levels, they may be said to constitute a pattern language. The use of a pattern language thus implies a holistic approach to a domain, as it guides designers through the problems across that domain, suggesting which patterns should be used in each case [BHS07]. As mentioned previously, the goal of patterns is not to present new ideas, but to represent what is. In other words, they formalize empirical observations by expressing the invariants of problem and solution spaces, and they provide abstractions that support solutions for those problems. As advocated by Kohls and Panke [ KP10 ], patterns can be seen as specific kinds of theories – good-practices theories – that are identified through a process very similar to that of scientific discovery. The patterns community calls this process pattern mining – the discovery of nuggets of wisdoms, as so expressively described by Kohls and Panke. The methods that lead to this discovery are based on inductive inference, which is common on qualitative research, and the validity of a pattern is supported by its empirical content and known uses. The reference architecture and implementation described in Chapter 7is possibly the most tangible part of this research from a software engineering standpoint, but the goal of this work goes beyond concrete implementations and tools. This patterns catalog embodies knowledge for designing an environment that supports the Adaptive Software Artifacts approach. Despite this, it doesn’t mean that the patterns are specific to this approach – we expect most of them to be useful when creating other kinds of systems for documenting software. pattern form 61 5.2Pattern Form Authors tend to find their own form when writing patterns, but often draw inspiration from forms that, for one reason or another, have become more popular. Three of the most influential pattern forms are the Alexandrian form, the GoF form and the POSA form. These forms differ in the order and structure of the contents; some of them are more narrative and others more explicitly structured (i.e., make more extensive use of section headings and lists). The Alexandrian form, as the name suggests, was used by Christopher Alexandre, namely in his book A Pattern Language [ Ale77 ]. It is a very narrative form, using few section headings and relying considerably on visual formatting elements to cue on the organization of the narrative. Almost on the opposite end of the spectrum, there’s the GoF form, used in the influential Gang of Four book 1 [ GHJV95 ]. This form is very structured, and a strong departure from the Alexandrian form, breaking the patterns in several sections. Finally, the form followed by the Pattern Oriented Software Architecture book series (i.e., the POSA form) is, like the GoF form, very structured, but the narrative follows an order somewhat closer to the Alexandrian form [Fow06]. The form of the patterns that you can read on this chapter was strongly influenced by the POSA and Alexandrian forms. The headings are mostly inspired in the POSA form, but only the most important ones were used. Whenever appropriate, we have replaced some of the POSA section headings by formatting cues. This supports a more narrative style than the original POSA form, and allows encompassing patterns of different levels of detail. The patterns were broken down in the following parts: Name. The name by which the pattern is known. The pattern name conveys the main idea of the underlying solution. Figure. A figure that visually conveys the pattern. Optional – a figure was not included for every pattern. Context. A description of the setting in which the pattern occurs. Often references patterns of a higher abstraction level, which set the context for the pattern being described. Example. A concrete example of the problem addressed by the pattern. Optional – this section was not used for all the patterns. 1 Commonly called this way due to its four authors. The book’s published name is Design Patterns: Elements of Reusable Object-Oriented Software. 62 patterns catalog Problem. Starts with a one-sentence problem statement that captures the main issue that the pattern tries to address, and that is visually highlighted through a different formatting. The remaining of the section details the problem and the (often opposing) forces that shape the solution. Solution. Like the problem section, this section starts with a visually highlighted one-sentence statement, which tries to capture the key idea of the solution. It then goes on to explain with more detail how the solution can actually be put into practice, and what are the positive and negative consequences of applying it. This section may also include the multiple variants of the solution, whenever they exist. Example Resolved. This section describes how the particular problem scenario described in the Example section can be concretely solved by the solution. Optional – this section was not used for all the patterns, and only appears when an example section also exists. Known Uses. This section provides examples of where the pattern may be observed in practice. Related Patterns. Other patterns, which may be used synergistically with the one being described. While the context section often mentions patterns of a higher abstraction level, this section covers all related patterns and givens more focus to patterns of a lower abstraction level. 5.3Catalog Overview The patterns addressed in this chapter can be grouped into different sets. They play different parts in the approach (Chapter 6) and in the reference architecture and implementation (Chapter 7). Patterns of Consistent Software Documentation. Documentation is an important part of the captured knowledge of a software project, providing a flexible and effective way of recording informal contents. However, maintaining it consistent requires a considerable effort. Existing solutions encompass different tools and approaches that support the process of creating, evolving and using free-text documents and other artifacts derived from the software development process. We have identified key problems and solutions for documentation consistency based on existing literature and catalog overview 63 personal expertise. In concrete, four distinct patterns and their relations were identified – Information Proximity (p. 66), Co-Evolution (p. 71), Domain-Structured Information (p. 74)and Integrated Environment (p. 76)[CFFA09b]. Patterns of Information Classification. Providing efficient access to information can be approached in different ways, but ultimately implies the creation of an Index (p. 82), represented with an indexing language, like a Taxonomy (p. 85), a Thesaurus (p. 88)an Ontology (p. 91)or a Folksonomy (p. 94). Each of these languages strikes a different balance between the effort to create and maintain the index, the effectiveness of knowledge capture, the guidance that readers can get, and how efficiently they can get it. Furthermore, Thesauri and Ontologies rely on the use of a the Controlled Vocabulary (p. 97)to disambiguate the meaning of terms [CA11] . Patterns of Flexible Modeling Tools. The benefits of using models have long been acknowledged by research and industry, but in practice the use of modeling tools often implies an unreasonable effort or confines itself to points in the project lifetime when the requirements and/or design is well understood. Free-form tools like whiteboards and textual documents fill information capture needs during the rest of the time – they pose a lower barrier for adoption and enable users to capture their flow of thought with fewer constraints than a formal modeling tool would allow. Flexible Modeling Tools try to provide a compromise of both approaches. The works discussed at the FlexiTools workshop series represent an interesting body of knowledge, covering different issues and approaches for this class of tools, and was one of the main sources for mining these patterns – User-crafted Static Meta-model (p. 103), Model Co-Evolution (p. 104), Meta-modeling ByExample (p. 106), Formalization (p. 108), Linked Models (p. 110)and Augmented Models (p. 112). The patterns identify several approaches that can be used by those developing modeling tools with flexibility requirements. They intend to represent the most relevant approaches in this area [CA13]. Patterns of Adaptive Object-Models. An Adaptive Object-Model (AOM) is an architectural pattern based upon a dynamic meta-modeling technique where the object model of the system is explicitly defined as data to be interpreted at run-time. The object model may encompass the full specification of domain objects, states, events, conditions, constraints and business rules. Several design patterns have before been documented and describe a set of good-practices within this domain. These patterns describe key concepts of object-oriented meta-architecture that are essential do AOMs 70 patterns catalog Wikis [ R¨ 03 ] address the use of links but goes beyond the creation of explicit and navigable relations between resources, addressing the collaborative nature of this kind of systems. Single-source is an approach similar to the one taken by the Code-Comment Proximity pattern [ R¨ 03 ], but goes beyond source code and comments, not restricting itself to any particular type of information. Transclusion is similar to the Import by Reference pattern [ R¨ 03 ], although it focuses on consistency maintenance. Known Uses Hypertext-based systems in general, of which wikis are a good example, allow to establish links between related resources. The term transclusion appeared initially in the context of hypertext-based systems. For example, Mediawiki, the wiki engine powering Wikipedia, uses this concept to allow the inclusion of repetitive blocks of content. XSDoc [ ADP03 , AD05b ] is a wiki engine oriented for software development that uses transclusion to weave together heterogeneous artifacts, thus giving origin to heterogeneous documents. Using the technique of Code Annotations (based on single source ), documentation (or parts of it) can be generated from a unified representation of textual descriptions and source code. It is primarily used in the creation of API documentation and is supported by several tools: Javadoc [ Fri95 ] is one of the first known uses of the technique, as is Autoduck [ Art00 ], a tool supporting code annotations in C++. The .NET framework uses XML in code annotations to produce compendiums of API documentation (CHM, HTML, etc.), in-editor assistance, and code-completion. Views are frequently the product of an automatic generation process, in which several contents are combined according to a pre-established document form – some tools exist that support this approach [BM06]. Literate Programming (LP) [ Knu84 ] combines textual descriptions and source code in a single source file , and provides the mechanisms to extract such different contents to different artifacts whenever required. The LP tool set dotNoweb [ Sou05 ] further allows combining textual descriptions and source code with diagrams expressed using the dot language. LP systems also usually provide a form of transclusion , by allowing the creation of information fragments – chunks – which can then be (re)used multiple times across several documents. Elucidative Programming [ VN02 ] is a documentation technique that relies on patterns of consistent software documentation 71 the creation of links between source code and documentation, allowing to mutually navigate between them. Several office software suites, such as Microsoft Office and OpenOffice, allow combining different kinds of artifacts in a same document, also resulting in heterogeneous documents. Some uses of Literate Programming, such as VDMTools, directly parse and write .rtf documents, which have native support for images. 5.4.3Co-Evolution Software documentation can be captured as a set of text documents of different types and purposes. Thus, they may sometimes address the same information from different perspectives. However, as documentation evolves, the effort of keeping them consistent rises due to the proliferation of duplicate and closely related contents. Problem When to update a related piece of information in documentation? Changes are made by the authors, who have the introduction of added value in view. However, changes required to ensure consistency don’t always provide immediate benefits, and may shift the author’s main focus. Furthermore, the primary goal of the project will not always be the same. For example, during an inception phase, the change rate at which documented artifacts evolve is usually high. This means that changing just enough of the related information fragments might be the best choice. On the other hand, deployment phases may benefit from producing documentation with a higher level of detail. Finally, tracking all the required changes may be difficult to carry out without any kind of auxiliary support , since it is easy to disregard global consequences during local modifications. Solution When a change is introduced, update the related information parts. If all the related pieces aren’t updated at the same time, they may grow harder to resync as time passes. Two variants to the co-evolution of contents are considered here and are further described on the following sections. Synchronous co-evolution is a good option when it is important that documentation is kept consistent at all times, or if the effort of recovering consistency at a later 72 patterns catalog time is high. Time-shifted co-evolution may be used when the effort of recovering consistency is reasonable. This may happen when it is not difficult to assess the existence of relations between contents and the presence of inconsistencies between them. Variant: Synchronous Co-Evolution Whenever a change is introduced, update every related piece of information. Although the quantity of information to be updated may be considerable, the most reliable way of ensuring consistency is to update all related information at the same time. Changes are made in small increments, in order to reduce the risk of forgetting to update something. The following consequences should be considered when applying this technique: State. Documentation is always in a consistent state. Focus. The focus of the author on the task at hand is harder to maintain, as some of the changes she is required to do are not directly related with her main goal. Effort. Introducing a change to a document carries a higher up-front cost – it may take more time than expected, as all the related contents will have to be updated at the same time. Efficiency. If a particular fragment has several others that depend on it, and it has a high rate of change, it may be inefficient to keep consistency at all times. Variant: Time-Shifted Co-Evolution Whenever a change is introduced, provide mechanisms to track the pending related changes, and update the most relevant pieces of information only when needed. Related contents don’t need to be updated simultaneously if the changes that are made are in some way recorded. Authors will be able to, at a later time, assess which are the pending related changes, and evolve documentation to a consistent state as soon as they are addressed. For example, using the concept of auditable document (see Section 5.4.3) authors may gain more awareness of the required modifications, facilitating the detection of changes that are still to be applied. The following consequences should be considered when applying this technique: patterns of consistent software documentation 73 State. Consistency is not kept at all time. Focus. The author may focus solely on the task at hand, leaving related changes for later. Effort. Only the changes that bring short term benefits are required to be made, and related changes may be deferred to a later time. Efficiency. The task of updating documentation is distributed across the development process, as documentation may be updated only when necessary. However, the author may be faced with the additional effort of tracking which information needs to be updated, even if tools that support this task may exist. Creating Auditable Documents An auditable document makes it possible to assess at any time who, how, why, and what has been produced, by tracking information regarding the authoring process. Being able to follow and understand how a document is evolved makes the entire process more transparent and traceable . However, it is important to note that the tracking mechanisms may increase the complexity of authoring the document, and the extra information that is recorded may increase the storage space consumption . Furthermore, for heterogeneous documents, tracking the evolution as a whole may involve tracking different types of artifacts. Related Patterns Domain-Structured Information supports Co-Evolution, since making richer information available allows tracking the information that needs to be co-evolved in greater detail. Information Proximity helps this pattern too, since having related contents easily reachable from one another assists in determining which contents are affected by a particular change. Some patterns already describe the use of auditable documents in more concrete scenarios, namely Document History [ R¨ 03 ] focuses on maintaining a list of past versions of a document, and Annotated Changes [ R¨ 03 ] provides a way to directly record, inside a document, which of its parts have recently been modified. 74 patterns catalog Known Uses Literate Programming and Code-Annotations, such as Javadoc, may be regarded as a way of supporting synchronous co-evolution , as providing Information Proximity helps to co-evolve related information parts simultaneously. Solutions that allow auditable documents to be produced support time-shifted co-evolution . Wiki engines and version control systems are good examples of such solutions, which allow to track how documents evolve and support assessing which changes are required to maintain consistency. It is common for text processors to provide a track changes feature, which is a form of Annotated Changes. This feature may be used by authors and readers to track the changes the document has recently gone through. Although this makes the document auditable to a certain point, it is usually very limited in time. 5.4.4Domain-Structured Information Free-text documents are often an important fraction of a software project’s documentation. They follow a text-oriented structure, using elements such as titles, paragraphs, lists, tables, etc. Although these elements allow a lot of flexibility, the degree to which a free-text document is useful depends on how well it accurately expresses the intended ideas. Moreover, the same piece of information may be better conveyed using different perspectives, intrinsically related to each other. The main reason why maintaining documentation requires continuous review is that relations between documentation parts aren’t explicitly formalized. This decreases the capability to automatically process it, i.e. in order to automatically assess its consistency. Problem How to structure the information in documentation? As mentioned before, textual documentation is a flexible way of capturing knowledge. While this flexibility is an important asset, formalizing the content itself makes information less subject to multiple interpretations, and allows it to be automatically processed. However, the mechanisms used to allow a degree of formalization higher than that provided by simple textual descriptions may affect the simplicity in producing documentation. patterns of consistent software documentation 75 Solution Organize contents according to their domain, so that the information form directly relates to domain concepts. Textual documentation doesn’t provide the mechanisms to formally express the relations between the concepts being documented. Structuring the contents around the domain concepts provides the support to automatically assess the existence of inconsistencies, and prevents the introduction of new ones. The following consequences should be considered when applying this pattern: Flexibility. Some flexibility is lost whenever information has to follow a predefined structure. Automation. The use of a domain-oriented structure with well defined semantics makes information less open to different interpretations, and allows it to be processed by computers. Related Patterns The individual information units often required by Information Proximity tend to converge to Domain-Structured Information, as the advantages of organizing the contents around domain concepts emerge. This pattern also supports Co-Evolution, as it provides a richer base of traceable information. As with the other documentation patterns in this chapter, Domain-Structured Information requires appropriate tool support, and may benefit from the use of an Integrated Environment. This pattern is similar to Structured Information [ R¨ 03 ], in that it also addresses how documents’ contents are organized. However, Domain-Structured Information focuses on formalizing contents according to the information’s domain, with the aim of automating consistency assessment, while Structured Information focuses mainly in structuring contents to ease the perception of the readers. Known Uses Code comments are a form of source code documentation. Code annotations, such as Javadoc comments [ Fri95 ], add an additional level of structure to source code comments, formalizing information elements with a lower granularity. Javadoc allows describing elements such as method parameters, authors, creation dates and references, among others. 76 patterns catalog Semantic Wikis support Domain-Structured Information, and some semantic wiki engines may automatically detect existing inconsistencies with the use of reasoners [DRR+05]. Some wiki engines allow templates to be applied for very specific purposes. Mediawiki allows the creation of sidebar templates, through which one may provide structured information. Systems taking an object-oriented approach to documentation have also been use in the past [Sam94,CS96]. 5.4.5Integrated Environment Working with different kinds of artifacts frequently implies the use of specialized and independent tools for each of them. Although such artifacts are sometimes strongly related, these tools don’t necessarily interoperate, making the artifacts more difficult to combine and confront, and the authoring environment heterogeneous and more difficult to use. Problem How to support the maintenance of consistency between independent artifacts with related content? Tools that deal with a wide range of artifacts usually provide a more homogeneous and interoperable environment, although they tend to be not as powerful and simple as specialized tools. Solution Use an integrated environment, where several types of artifacts and their relations may be maintained uniformly. An integrated environment goes beyond the capabilities that general purpose tools possess. It supports handling several types of artifacts, providing specialized features for each of them and an infrastructure through which they interoperate. This supports strategies of documentation maintenance that focus on bridging related information parts regardless of their nature. The following consequences should be considered when applying this pattern: Specialization. Integrated environments strike a balance between a generic approach, in which tools may handle several types of artifacts with a basic level of patterns of information classification 77 functionality, and a specialized approach, in which exists a deeper support for a selected set of artifact types. Simplicity. While potentially making each tool more complex individually, their overall simplicity is increased by providing a more homogeneous usage. Interoperability. An integrated environment coordinates the several tools it provides, and supports their interoperability. Related Patterns Integrated Environment directly contributes to the remaining documentation patterns of this chapter by orchestrating the several tools involved. It is also directly related to the pattern Few Tools [ R¨ 03 ], which addresses the notion that supporting the creation of documentation with too many and unconnected tools may become a burden to authors. Known Uses Eclipse and Visual Studio are examples of integrated environments that combine different kinds of artifacts and tools, supporting and articulating their work. Trac [ Edga ] and Redmine [ Lan ] are Web-based environments that integrate different kinds of information, including textual descriptions supported by a wiki, source code browsing, milestone management, issue-tracking, etc. 5.5Patterns of Information Classification In the context of knowledge work, it is expected that as the available information grows, one would be more effective in his tasks. Unfortunately this is not always the case, and the value of information frequently decreases as the quantity of information increases. This apparent contradiction is due to our human limitations in processing high quantities of raw information. This section looks into six patterns for classifying and improving the access to information. Some of these solutions have been used since the 4 th century [ Wel94 ], and are nowadays very well known in the domain of information science. Others came into being on the context of the Web, even though they conceptually share a lot with “older” solutions, but all are used as means for information seeking and retrieval. In one way or another, they can all nowadays be seen pervasively in software systems. 78 patterns catalog The main audience for these patterns are those wanting to make information quickly reachable, in the context of software systems. Depending on the kind of system, they can be either developers or users of the system. Although the patterns don’t lead to a specific implementation, their implications easily crosscut the design of a system, from how the data modeling is done, to how information is perceived and interacted with through the user-interface. To a lesser extent, we believe these patterns may also be useful to those wanting to take their first steps into information indexing, and need to gain a better understanding of the different concepts involved. 5.5.1Overview These patterns were mined from the experience gathered by the authors while developing software systems – some of them in the information science domain – that use these techniques to make information accessible. Two approaches to accessing contents – searching and browsing – have proven useful in different contexts. While search provides immediate results, browsing allows an exploratory approach to finding contents, which is key when information needs are ill-defined. The patterns described in this section focus mainly on supporting the access to contents through browsing. Organizing, Classifying, Indexing These three concepts are used throughout the patterns, but the differences between them can sometimes be subtle. They can work together to support the same overall goal – to ease the understanding and access to contents. To organize is to provide an order, that is, to systematize the way in which the contents are recorded and conveyed, so that they can be more easily understood. On the other hand, to classify is to assign the contents to classes, that is, to group them according to common features – it implies abstraction, and a specific kind of organization. At last, to index is to provide the key topics or the classes of the contents as access points to those contents; the emphasis is on how readers can use those common features to actually find and delve into the contents. The patterns below address these three concerns to some degree. They are Information Classification Patterns because they focus mainly on how the different topics of the contents are abstracted and represented. patterns of information classification 79 The Patterns The first pattern of this set is the Index. Indexes can be elaborate structures, but, in their simplest form, they are lists of terms, usually organized alphabetically. In the context of publishing, the word index specifically denotes an alphabetically ordered Index of subjects, usually appearing in the back of the document, but unless noted otherwise, the term index will here always be used in the most general sense, as will become clear in the description of the pattern. The creation of an index requires the use of a representation language, which supports expressing its entries. When used to represent the information of an index, these languages are called indexing languages. The most expressive ones are used for other purposes too, as they are able of representing knowledge in general. Four of the patterns – Taxonomy, Thesaurus, Ontology and Folksonomy – are about such languages. Directly or indirectly, they support the creation of Indexes, and they all strike different balances between the effort of creation, the effectiveness of knowledge capture, and the ease of use. At last, the Controlled Vocabulary pattern describes a general approach to disambiguating the meaning of terms; it is key in Thesauri and often used with Ontologies. Figure 5.2depicts the relations that were just described, and will be explored in greater detail in the description of each pattern. INDEX THESAURUS ONTOLOGY TAXONOMY FOLKSONOMY CONTROLLED VOCABULARY may be represented using may use a uses a representation / indexing languages Figure 5.2:Pattern map of the information classification patterns. Index (p. 82)– Supports readers in finding the contents they seek more efficiently; Taxonomy (p. 85)– Allows representation of information along an hierarchical structure with loosely defined semantics; Thesaurus (p. 88)– Provides more semantics and expressiveness when representing information, allowing related subjects to be connected; 86 patterns catalog To sort out the index entries they seek, readers need additional information, which describes and contextualizes the entries, and that they can use to partition them. Although a semantically richer index implies a greater effort from indexers, it better guides the readers in finding contents. Solution Organize the index entries hierarchically. The meaning of the relations between parent and child entries may vary, so the index can be partitioned by different dimensions into several subareas. Choose entries that cover the whole domain , and try to keep the taxonomy tree balanced. Add entries that can be easily understood by the readers when taken in context with the upper taxonomic levels. When creating the entries of a taxonomy don’t try to make them stand on their own. To assemble meaning from an entry, readers will consider its context in the taxonomy. The same terms may convey different meanings when they appear in different points of the taxonomy. Each taxonomic level relates to the upper level according to one of its dimensions. In case the same term is placed in more than one point of the taxonomy it does not mean that the same subject is being classified in multiple ways, but rather that different subjects are being represented. The order in which such dimensions are represented as parent-child entries should reflect the knowledge that we foresee readers may have, and the way they will seek the entries in the taxonomy. From an indexing point of view, a Taxonomy can be said to be a fixed-vocabulary language, as a pre-established representation of terms is taken as a basis for the indexing process. Taxonomic indexes are pre-coordinated, because each entry is a combination of terms that describe a group of other entries. The partitioning of the index along a tree structure makes navigating it more efficient , as readers are able to eliminate from their search several index entries at once, when they belong to a subarea that does not interest them. Also, the more a reader knows about a given domain, the more efficient she is navigating through a taxonomy of that domain – she will be quicker in grasping which dimensions the taxonomy is using to partition the index, and how she should navigate it to reach the intended index entries. Newcomers may have to explore the index first, before being able to use it efficiently. But taxonomies are not without liabilities, and the level of expressiveness that they allow is one of them. In practice, it may be hard to group contents according to a single sequence of dimensions. Although you can try to anticipate which features of patterns of information classification 87 each piece of information will be the most relevant to future readers, different readers may easily have very different needs. The added effort of pre-coordinating and grouping together related terms makes a Taxonomy harder to create and maintain when compared to using a plain list of terms. Only considering a high number of readers, and a low rate of updates does such effort pay itself easily. Example Resolved Realizing that she must stop going through all the files every time she needs one of them, Amy started organizing the files into subdirectories. She created three subdirectories inside the “Documents” directory: “Family”, “Work” and “Friends”. Inside the “Work” subdirectory she created some more subdirectories, one for each customer. Whenever she needs to reach the project proposal for Mr. Smith again, she will rely on the directory structure to guide her. She will first open the “Work” subdirectory, then the “Mr. Smith” subdirectory, and finally go through the files there. The top-most directories partition the larger groups of files, according to Amy’s social groups (family, work, friends). In turn, the directories inside the “Work” directory are grouping the files by person – sender or receiver. Each level uses the dimension that better helps Amy navigating that particular group of files. Known Uses Several software systems use the concept of “folders” – container of items, used to organize them. Such items can themselves be other folders, and thus contain other items, forming a tree-like structure that can be regarded as a Taxonomy. An example of such use is the Alfresco Content Management System 3 , which has the concept of “Spaces”. Alfresco Spaces are generic containers that behave much like “folders”. A Web Directory, like the Open Directory Project 4 , can be seen as a Taxonomy that classifies websites on the World Wide Web. Web Directories were once important to find Web resources, but due to their very high rate of change and growth, very few have survived in favor of full text search engines. A Table of Contents, either digital or in print, may be seen as a Taxonomy that organizes contents according to sections and chapters. 3Available at http://www.alfresco.com/. 4Available at http://www.dmoz.org/. 88 patterns catalog The Dewey Decimal Classification System is one of the Taxonomies with the most widespread use. Its goal is to cover all areas of knowledge, supporting the classification of books and other library items, and providing a way to easily find them, on the online catalog or shelves of a library. Related Patterns Taxonomies, like Thesauri, Ontologies and Folksonomies, can be used to represent an Index. Like Thesauri, Taxonomies assume a closed domain. When information seekers are not newcomers to the domain area, or if the contents are simple enough, Taxonomies allow to reach information quicker, but otherwise, semantically richer indexing languages will provide better guidance. 5.5.5Thesaurus Consider that you want to classify and create an Index for a body of information, to support the readers in quickly reaching the contents they need. Readers may have very different backgrounds and different levels of prior knowledge on the subject they seek. The frequency at which the contents are updated is not high, when compared with how many times they will be used. Example Suppose a library keeps a set of documents about art, and allows readers to find them through an online catalog. John would like to find some documents about a famous painter, whose name he doesn’t recall right now. All he remembers is that the painter was contemporary to Monet, and painted using the same style. Paul happens to also be looking for documents about the same painter, but all he knows is that he created the painting “Dance at Le Moulin de la Galette”. Both John and Paul need to do some research before getting to the documents they need using the library’s index. John will start by finding documents about Monet. He then uses these documents to learn that Monet was an impressionist, and then find documents about the painters of that movement, to finally recognize Renoir as the name he was missing. Searching by Renoir on the library’s online catalog will finally reveal the documents he needs. Paul, on the other hand, will first need to search for patterns of information classification 89 the name of the painting he knows, to find that it was painted by Renoir. He can then use the library’s online catalog effectively. Problem Different people seeking the same contents need the index to guide them in different dimensions. The same contents may need to be accessed differently, depending on the knowledge of the reader. To sort out the index entries they seek, readers need information that describes and contextualizes those entries. Although a semantically richer index implies a greater effort from indexers, it better guides the readers in finding contents. Solution Organize the index entries as a network of subjects. Define the meaning of each subject carefully, by using a Controlled Vocabulary, and connect them with other, related, subjects. More specifically, organize subjects according to five different elements/connections: Broader/Narrower – Thesauri, like taxonomies, are organized hierarchically, but the semantics of such relations is better established than with taxonomies. Parent subjects are said to be broader, and child subjects narrower, in the sense that the scope of each child subject is narrower than the scope of the parent. Scope Note – Each subject represents not merely a term, but a concept that is part of a Controlled Vocabulary. The meaning of the concept is defined through a scope note. Synonyms – Other terms that may describe the same subject. Synonyms are unauthorized forms of the Controlled Vocabulary used to support the Thesaurus. Topmost – Each subject has at least one topmost subject, which is the one that would be found by following the broader relations until the broadest possible subjects are reached. Related – Refers to related subjects, which are not broader or narrower. Despite Thesaurus’ entries representing concepts, the terms are usually emphasized more than the underlying concepts, and Thesaurus are very often perceived as just a set of connected words. A Thesaurus allows a richer description of subjects when compared with a Taxonomy, as it supports expressing broader/narrower relations, 90 patterns catalog which have more concrete semantics than the hierarchical relations of a Taxonomy. Related connections support expressing other (unspecific) kinds of relations. In spite of the added expressiveness when compared with Taxonomies, Thesauri are sometimes extended with even further attributes and kinds of relations. To create a Thesaurus, identify subjects (i.e., index entries) at different degrees of abstraction, reflecting the different levels that may be found in the contents, from the very coarse-grained (general) subjects to the very fine-grained (specific) ones. Represent in the Thesaurus all the knowledge in the field, and reuse is as often as needed. You should index the same piece of contents with multiple index entries, providing multiple access points, possibly to be combined, when searching the index. Thesauri provide semantically richer connections between subjects, which makes them quicker to navigate for those without much prior knowledge on the domain. Thesauri are meant to comprehensively cover their target domain, and are reviewed and updated only sporadically. They are, for this last reason, called a fixed vocabulary. This supports their usability , as it makes it easier for readers to learn how the thesauri that they use are organized. Expressiveness is better than with a Taxonomy, but Thesauri demand more attention to semantics, which can imply more effort during creation and maintenance. However, in practice, maintenance may not actually be harder than with a Taxonomy, because the meaning of each Thesaurus’ entry is more fine grained and better defined – it’s easier to improve an entry while being confident that the rest of the Thesaurus remains consistent. Example Resolved Suppose that the catalog software, of the library of the example above, allows thesaurusbased indexing. The librarians have decided to built a thesaurus in the art domain, and are using it to index the documents that they are curating. John will start looking for the elements he knows. He will first seek for the term “monet”. He finds the corresponding subject in the thesaurus that he confirms to be the one he is looking for, upon reading the scope note, and by observing that it is narrower term of the “Painter” entry. He quickly goes through that thesaurus entry, and finds out that “impressionism” is a related thesaurus entry. He then looks at the remaining entries related with “impressionism”, and recognizes “Renoir”as the painter he was seeking. He now just needs to follow the index locator(s) for that entry, to reach the documents about Renoir. patterns of information classification 91 Paul, on the other hand, will start seeking for the term “Dance at Le Moulin de la Galette”. He finds it, and too sees “Renoir” as a related entry. All he has to do now is follow the locator(s) to reach the documents he needs. Known Uses GISA is a software product for creating records and descriptions of archival documents, which uses a thesaurus-based index. Its Web frontend can be found in the websites of several Portuguese archives, like the Archives of the City Hall of Gaia 5 and the Archives of the University of Porto6, among others. The Index New Zealand Thesaurus 7 was created to describe publications about New Zealand and the South Pacific in the areas of social sciences and humanities. It provides access to journal and newspaper articles. Related Patterns Thesauri can be used to represent an Index, like Taxonomies, Ontologies and Folksonomies. Like Taxonomies, Thesauri assume a closed domain. Thesauri are better at guiding information seekers than Taxonomies. The entries of a Thesaurus form a Controlled Vocabulary in the sense that their meaning is established unambiguously. 5.5.6Ontology Consider that you want to classify and create an Index for a body of information to support the readers in quickly reaching the contents they need. Readers may have very different backgrounds and different levels of prior knowledge on the subject they seek. Contents may frequently be created and updated, and they are likely accessed through platforms that enable collaboration, such as the Web. Example A research institute has, over time, produced a large body of information. Some of these contents were recorded by the research groups themselves, in different software 5Available at http://arquivo.cm-gaia.pt/. 6Available at http://gisa.up.pt/pesquisa/. 7Available at http://innz.natlib.govt.nz/content/thesaurus/. 92 patterns catalog systems that they maintain. These contents are organized in different ways, depending on the system they were captured in; they have varying levels of structure, from information systems, to free-text documents, to raw experimental data; and they keep evolving as more results are found and documented. Linda is researching on the area of automated software testing, and would like to know which results her colleagues have achieved in this area in the last few months. She hopes to find work to build upon, or find researchers of other groups to collaborate with. The institute provides a list of systems that it uses to capture contents internally. Each group works on a specific subarea, and knowing this could help Linda find the systems with the contents she needs. However, there aren’t groups working specifically on automated testing, and almost all of the groups have done some automated testing at some point. If she wants to make sure to find all the contents she needs, Linda will have to search through all the systems and their information. Problem Without a rich and accurate representation of the contents, an index is not able to guide a reader effectively. To sort out the index entries they seek, readers need additional information, which describes and contextualizes such entries. Although a semantically richer index implies a greater effort from indexers, it better guides the readers in finding contents. Solution Organize the index entries as a network of subjects. Define the meaning of each subject, and connect it with other, related, subjects. More specifically, organize subjects according to elements such as individuals,classes,attributes and relations, among others. An Ontology is a formal, explicit specification of a shared conceptualization [ Gru93 ]. It may gather a collective understanding on a given area and be open to, and constantly updated by, a group of people. Several ontology languages exist, but common components include Classes,Attributes,Relationships and Individuals. Subjects may be defined by classes or individuals, and are characterized by attributes and relationships. Attributes and Relationships are themselves described by classes, and this mechanism allows the language to be extended as needed. Given its formal nature, an Ontology is very fit to use a Controlled Vocabulary. patterns of information classification 93 The ability of expressing virtually any kind of attribute and relation makes Ontologies able to provide a richer description of subjects when compared with a Taxonomy a Thesaurus or a Folksonomy, and has thus the capacity of guiding information seekers more effectively. Even though an Ontology may be open and constantly updated by a community, it may sometimes prove to be a difficult endeavor to reach a consensus over the conceptualizations. Example Resolved Going back to the example presented in the beginning of this pattern, the institute can build an aggregator that gathers contents from each system, and provides a unified and abstracted representation of them. This ontology can be used as a global index that users use to reach the actual contents. Although the source contents need to be semantically rich to be aggregated, the ontology can be completed with additional information to make other contents accessible through the index too. Linda now uses the ontology-based index to find an entry about automated testing, and follows the index locators to free-text documents maintained in a system used by the Artificial Intelligence Group, and to some unit-test coverage data that was used for creating software visualizations by the Computer Graphics Group. Linda is directed to the right systems and, more specifically, to the right contents within the system. Known Uses Semantic MediaWiki 8 is an extension to the MediaWiki wiki engine. It adds the ability to annotate the contents of a page, conferring it semantics. This data can then be queried, by one or several of its dimensions, and the results provided, from within a wiki page, as an access to the contents in question. Among other features, it supports exporting data as OWL9, an Ontology representation language based on XML. In research, we can find several approaches to indexing contents with an Ontology. An example, among several others, is the work by Luances et al, which uses an Ontology to improve the query capabilities to a Geographic Information System [LPPS08]. 8Available from http://semantic-mediawiki.org/. 9 The full specification of the Web Ontology Language (OWL) may be found at http://www.w3.org/ TR/owl-features/. 94 patterns catalog Plone Ontology 10 is an add-on for the Plone Content Management System that allows to collaboratively create an ontology that can be navigated and used to access the system’s contents. Related Patterns Ontologies can be used to represent an Index, like Taxonomies, Thesauri and Folksonomies. Like Folksonomies, Ontologies assume an open domain. If a rich description of contents is more important than supporting collaboration, an Ontology makes a better indexing language than a Folksonomy. The meaning of the elements of an Ontology is established unambiguously, and in that sense, it may be very close to using a Controlled Vocabulary. However, that’s often not the case, as the elements of an Ontology don’t necessarily have an authorized form. 5.5.7Folksonomy Consider that you want to classify and create an Index for a body of information, to support the readers in quickly reaching the contents they need. Readers may have very different backgrounds and different levels of prior knowledge on the subject they seek. The amount of contents is overwhelming; they may be frequently created and updated, and are likely accessed through platforms that enable collaboration, such as the Web. Example Suppose you have created a software platform for amateur photographers through which they can publish their best works on the Web. You want to let users easily seek the contents they need, but you don’t have the resources to hire a team to manually classify such a large set of pictures, and you also don’t want to demand from users a lot of effort to classify their photos by topic. Kate, an early adopter of the system, has just uploaded her photo album from the last five years, and she is now wondering how to find the pictures of the weekend she spent last year in Portugal. She would also enjoy knowing what other pictures of Portugal there are on the system, as she would like to find new places for her next visit. 10 Available from http://plone.org/products/ploneontology patterns of information classification 95 Problem Indexing and classifying a large body of information is unfeasible or at least very costly to carry out. Assuming the information is already recorded, one could consider assigning a team with the task of creating an index to ease all subsequent accesses. However, the effort of such an endeavor is usually very high. This is aggravated if such information is in constant change , in which case, the classification efforts cannot be limited in time, and have to follow the entire lifecycle of the information. Those that are most knowledgeable about some specific contents are not external indexers but its own creators, as they are more aware of its context and domain. On the other hand, information creators are not necessarily aware of what makes a good index, and may lack the necessary analysis and abstraction skills to represent elaborate index structures. Solution Ask the creators or users of the information to identify the set of words that most accurately describe the contents, and tag them with those terms. Those words are the entries of the index. By letting – and encouraging – users to assign descriptive terms to pieces of information, an index will emerge. It will not be defined up-front, but rather will gradually appear from the practice of collaboratively tagging contents [Fur10]. There is not a single way to seek contents using a folksonomy. Tags can be made available to information seekers as simple alphabetically ordered lists, or as tag clouds. Tag clouds present the several tags by laying them out in different locations, and using different font sizes and colors to highlight the relative importance of each one, usually directly reflecting the number of times they were used to tag some content (i.e., the number of underlying locators). The final result is a word index, as opposed to a subject index. This kind of index does not make use of a Controlled Vocabulary, and thus its entries lack a strong semantics , leaving to the reader the job of figuring out if the contents tagged with a given entry actually refer to what he is looking for. However, this is also one of the biggest strengths of this solution; by reducing to a minimum the effort required in the analysis phase, the creation of the index is easy enough to be done by any information creator, in a distributed way. 102 patterns catalog Augmented Models (p. 112)– Supports complementing or annotating models with contextual information that is often not structured or is loosely structured; Figure 5.4also shows, for each of the patterns, which meta-levels are subject to change. Other criteria can be used to group the patterns – Figure 5.5depicts which patterns focus on relaxing or leveraging the constraints between the model and metamodel levels, and which ones focus on doing the same for the relations between models and other, external, contents. Even though the end goal of these patterns is mainly to support the creation of models, User-Crafted Static Meta-Model, Model Co-Evolution and MetaModeling by Example directly support creating or introducing changes at the metamodel level. As depicted by Figure 5.5, together with Formalization, they have in common the goal of providing flexibility between the two modeling levels. Model / outside-world flexibilityInstantiation link flexibility FORMALIZATION USER-CRAFTED STATIC META-MODEL META-MODELING BY EXAMPLE MODEL CO-EVOLUTION LINKED MODELS AUGMENTED MODELS Figure 5.5: Pattern map of the flexible modeling patterns, grouped by the link that is the focus of flexibility. Formalization, Linked models and Augmented Models differ from the other three patterns in that they focus on the introduction of changes only at the model level. Finally, and not explicitly shown by the figures, Linked models can be said to be a form of supporting Augmented Models, as the contents used to augment the model are, in fact, elements from another model. These patterns were mined mainly from the body of works discussed at the latest three editions of the FlexiTools workshop series 18 [ OvdHS+10b , KOvdHS10 , OvdHS+11 ]. Rather than adopting a specific view of what makes a modeling tool flexible, they try to reflect the current understanding of the area as taken by its community. Some works in particular have already tried to categorize challenges and approaches taken in this area, and influenced the recognition of these patterns – the very complete 18 Workshop on Flexible Modeling Tools patterns of flexible modeling tools 103 summary of the FlexiTools workshop at SPLASH 2010 19 by Kimelman and Hirschman [ KH10 ], a work by Gabrysiak et al that proposes a classification of meta-models’ usage scenarios [ GGLS11 ] and another work by Kimelman and Hirschman [ KH11 ] that includes different interpretations of the notion of modeling tool flexibility. Some of the authors’ previous experience have also influenced the identification of these patterns; namely, their work on Adaptive Object-models [ FCW08 , FCYA10 ] and on Adaptive Software Artifacts [Cor10,Cor13]. 5.6.2User-Crafted Static Meta-Model Modeling tools support their users in expressing information in a given domain or domains. The creation of a modeling tool must consider which domains the modeler will need to address and how expressive they will need to be on those domains. Problem The modeling tool developers may be unable to anticipate the domains and expressiveness that the modeler will need. Conceiving modeling tools and their underlying meta-models requires good abstraction skills and considerable effort , making it compelling to reuse them rather than creating multiple ones, tailored to each specific context. Creating a modeling tool that adheres to a specific meta-model rather than supporting multiple ones also contributes to keeping the modeling tool simple. On the other hand, if the abstractions provided by a modeling tool and its underlying meta-model are not suited for the intended domain, the user might not be sufficiently expressive in that domain. In such a case, using that meta-model might be impossible or imply more effort , and the resulting model might not always be reliable . Solution Allow the meta-model to be refined by the modeler, before a model is built. Use this pattern when the user will need to model domains that can’t be fully anticipated by the modeling tool developers. Instead of developing a modeling tool that uses only one specific meta-model and addresses a specific domain, allow the user to supply her own meta-model, tailored to a domain’s needs. This means that modeling tool developers must resource to a meta-meta-model to define which meta-models 19 SPLASH 2010 – Systems Programming Languages and Applications: Software for Humanity 104 patterns catalog may be supplied by the user. Often, the modeling tool may itself support the creation of the meta-models, in which case it will have a simpler design if the same mechanism is used to support the creation of both modeling levels (i.e., if Everything is a Thing [FCYA10]). This pattern can also be referred to as User-generated Meta-model [GGLS11]. Known Uses User-Crafted Static Meta-Model is perhaps the least flexible of this set of patterns and can be found on some traditional modeling tools. For example, the Eclipse Modeling Framework (EMF) supports defining a meta-model and, together with the GMP framework (Graphical Modeling Project), it allows defining a graphical representation and build a complete modeling tool for that meta-model. MetaEdit+ [ TPK07 ] is an environment that allows creating new modeling languages. It uses graphical meta-modeling to support the early stages of language creation, to define the key language concepts and rules. The meta-modeling language is defined as one of the several domain-specific languages supported by the platform. Related Patterns User-Crafted Static Meta-Model differs from the other patterns flexible modeling patterns described in this chapter because the flexibility that it allows is only possible until the meta-model is instanced into a concrete model. Namely, it doesn’t address meta-model changes after a model has been created. When modelers need the meta-model to evolve after it has been instanced, they need support for Model Co-Evolution or Meta-Modeling by Example. 5.6.3Model Co-Evolution User-Crafted Static Meta-Models are created before modeling activities take place, and they allow modelers to define how expressive they will be able to be during such activities. To use that pattern effectively, modelers must have a very concrete idea of the domain that they will address before starting to model in that domain, as the modeling tool might not easily allow refining that idea after the initial creation of a meta-model takes place. patterns of flexible modeling tools 105 Problem Modelers may be unable to anticipate the expressiveness that they will need. Models, and their semantics, directly depend on the constructs defined by the associated meta-models. Users may need to change the meta-model during the creation of a model, but they also want models’ consistency towards their metamodel to be kept. Modelers don’t want to spend a lot of effort manually maintaining this consistency if they can avoid it. Modeling tools often help by preventing some operations that would cause inconsistencies, but this may be limiting, as it means that the meta-model cannot be be freely evolved to all states that it once could, before there were models that used it. Solution Allow to change meta-models and automatically evolve dependent models accordingly. Support Model Co-Evolution when meta-models need to evolve even though models that are based upon them have already been created and the user needs these models to be kept in-sync with their changing meta-models. Often the operations that are made available to the user at both levels – model and meta-model – are defined by a set of distinct classes that confines the changes that the modeling tool supports, thus applying the Command pattern [GHJV95]. To make the changes at the meta-model level have the right repercussions at the model level, operations executed at the meta-model level will spawn the execution of other operations at the model level. For each kind of meta-model change, the tool must know how to make the consequential changes at the model level, and sometimes it may require the modeler to provide additional instructions on how the model should be transformed. The key principles of this pattern can also be applied between the data and model levels – changes to the model can trigger changes to any data that complies to it. Known Uses Gabrysiak et al emphasize the need for modeling flexibility and propose a tool that combines a) a minimization of the restrictions imposed by the meta-model to the model with b) the co-evolution of models as a result of introducing changes to their meta-models [GGS10,GGLS11]. 106 patterns catalog Some approaches and tools to support the evolution of models upon the introduction of changes to their meta-model have been proposed, including Wachsmuth’s [Wac07] and Cicchetti’s [CDREP08] works. Some object-oriented frameworks like RubyOnRails 20 (RoR) and Django 21 support the creation of a domain model that is persisted using an ORM. The creation of this domain model usually doesn’t involve the use of a modeling tool, nor is its meta-model subject to changes, but the model can, and often does, change during the development of a system. Existing data that complies with such models needs to be changed accordingly, and these frameworks very often provide the mechanisms to do so (e.g., Migrations in the case of RoR and South22 in the case of Django). Related Patterns Like Meta-Modeling by Example, and unlike User-Crafted Static Meta-Model, Model Co-Evolution supports evolving the meta-model after it is instanced. Co-Evolution is a general concept that may be applied to other domains. Namely, it is often used in the context of software documentation as described by the Co-Evolution pattern [CFFA09b]. As mentioned above, this pattern is also applicable between the model and data levels. In particular, the Migration pattern [ FCW08 ] supports co-evolution between a model and existing data that complies to that model, in the context of Adaptive Object-Models [FCA09]. 5.6.4Meta-Modeling by Example Both modeling tool developers and the modelers themselves may be unable to anticipate the expressiveness that the modelers will need. To use User-Crafted Static Meta-Models effectively, modelers must have very concrete ideas of the domain that they will address before starting to model in that domain. Model Co-Evolution opens the possibility to refine those ideas after the initial creation of a meta-model takes place, but modelers often discover the new directions that the meta-model will take when trying to express a model and exploring their options at that abstraction level. 20 Available at http://rubyonrails.org/. 21 Available at https://djangoproject.com/. 22 Available at http://south.aeracode.org/. patterns of flexible modeling tools 107 Problem The need to change the meta-model diverts the modeler from the creation of the model. The need to constantly update the meta-model to allow for more expressiveness at the model level diverts the modeler from her main stream of thought . Moreover, the need to create or update a meta-model is often a barrier to entry , especially when modelers don’t have the required higher abstraction skills or technical skills . This barrier should be as small as possible, but the modeling tool shouldn’t be simplified to the point of loosing the rigor and consistency that a meta-model can support. These difficulties are easily felt when end-users are asked to collaborate in the design and implementation of a new DSL, for example. Solution Build a meta-model by providing examples at the model level. Support Meta-Modeling by Example when users need to create a meta-model but it’s more feasible to start by exploring options and experimenting at the model level than to keep models consistent with their meta-model at all times. A model always needs an underlying meta-model, but the meta-model doesn’t have to be explicitly captured before modeling is done – it may exist only in the modeler’s mind. By relying on an explicit meta-meta-model, and making minimum assumptions about the meta-model level, modeling tools may support the creation of models in a very unconstrained way, and help modelers decide later what a compatible meta-model could look like. On this later stage, the modeling tool can automatically infer a meta-model compatible with a given model or models, or it can use them to guide the creation or update of a meta-model when the user wishes to create/update it. Although not within its main focus, the key principles of this pattern can also be applied between the data and model levels – a possible model can be inferred from information that has been structured in an ad hoc way, in the same way that a possible meta-model can be inferred from a given model. This pattern can also be referred to as Lazy Meta-Model [ GGLS11 ] in the sense that the meta-model needs only to be created when absolutely necessary, or as Bottom-Up Meta-Modeling [ SCDLG12 ] in the sense that it encourages modeling to start from lower abstraction layers (i.e., the bottom). 108 patterns catalog Known Uses Cho et al [ CSGW11 ] tried to make the creation of Domain-Specific Modeling Languages more accessible to domain experts, by using examples of the language provided by the end-users themselves to infer the language’s meta-model and semantics. This approach is referred to as Modeling Language Creation By Demonstration. Kuhrmann has developed similar work [ Kuh11 ], striving to better support language engineers in the hard, time-consuming, and knowledge-intensive task of creating metamodels, models and DSLs. The Process Development Environment platform allows a free-form language design, and allows to visually represent domain entities and simple associations that are used to derive a meta-model definition. Smart Office Tools [ DOS10 ] support a content model, capable of visually representing the domain knowledge as a domain diagram, but without tying the user to a specific meta-model. The user is able to customize the diagram notation by creating styles, which she can at any time annotate to formalize a meta-model. Related Patterns Meta-Modeling by Example always depends on the creation of a model, from which higher modeling levels are manually built or inferred. Such inference or creation of information constructs out of other pieces of information is something that this pattern has in common with Formalization. They are, otherwise, very different patterns, as Formalization acts only at the model level – it can be used when information is initially void of any explicit form of domain structure and only later is it made into an explicit model. 5.6.5Formalization If the modeler needs to change the meta-model and the modeling tool supports Model Co-Evolution or Meta-modeling ByExample, she is able to introduce those changes without being constrained by the existing models. Other times, the modeler won’t feel the need to change the meta-model but the instantiation link (i.e., the link between model and meta-model elements) may still be limiting in freely creating the model. Problem It is not always efficient, or even possible, to capture information as a model right from the beginning. patterns of flexible modeling tools 109 Modelers have various reasons for not using formal modeling tools, and rather often resource to free-form tools instead. Despite their usability , and advantages for communication and creativity , these tools don’t offer any support for maintaining the results as models. Formal modeling tools provide such support, by enforcing conformance to a specific meta-model, and they confer the semantics needed to interpret the models objectively. Solution Derive formal/rigorous representations from information that was firstly captured informally/non-rigorously. Support Formalization when the user’s intent is to create a model that complies to a specific tool or to a well-known meta-model but she can’t, or doesn’t want, to be always constrained by that meta-model. Formalization frees the user from some or all meta-model constraints during modeling activities, easing the capture of the flow of thought or allowing to experiment and explore multiple options at the model level. The direct result of such a process, whilst not a model, can subsequently be formalized into a model. Formalization always starts with free-form information. The move from free-form contents to a model can sometimes be done automatically by the modeling tool. For example, when sketching a diagram, the tool may store a combination of drawing gestures and a resulting raster image to infer what model elements the user meant to represent. The richer the information obtained through the users’ input mechanisms, the easier may be to infer the semantics of what the user intended to express. Extraction techniques like image analysis or natural language processing may be used to support the move from free-form contents (e.g., a piece of text or a raster image) to a model. Additionally, when it’s not possible to automatically infer a model, the modeling tool can interactively assist the user in manually building it from the contents. Known Uses The SKETCH API [ SB10 ] allows developers to add sketch recognition to modeling editors built for the Eclipse IDE. It leverages touch-enabled devices and their great potential as drawing tools, allowing to create and manipulate freehand sketches from which a formal model can be inferred and associated with an underlying meta-model. Architects Workbench (AWB) [ ABK+06 ] provides totally free-form text entry and the ability to evolve them towards formally structured contents, maintaining the 110 patterns catalog traceability from one form to the other. AWB’s users can use a markup and model technique to create model elements directly from the text. UNICASE [ HNA+10 ] supports explicit traceability between information with different levels of abstraction and between loosely-formal project models (the artifacts that describe a software project, such as tasks, bug reports and informal communication) and system models (the artifacts that describe the system, such as functional requirements, UML models and detailed system specifications). This traceability information is then used to support the formalization process that underlies the propagation of changes from project models to system models. Related Patterns Formalization is similar to Meta-Modeling ByExample to the extent that both patterns relax the constraints between the model and meta-model levels and help the modeler create new information constructs, or even automatically infer such information constructs, from other contents. Otherwise, they are very different patterns, as Formalization supports creating a model, and Meta-Modeling ByExample supports introducing changes at both modeling levels – model and meta-model. Organizing documentation towards Domain-Structured Information [ CFFA09b ] can be seen as a light attempt to Formalization, as the contents are organized according to their domain but not to the extend of becoming a model. 5.6.6Linked Models User-crafted Static Meta-model, Model Co-evolution and Meta-modeling By Example allow modelers to increase their potential expressiveness at the model level by enriching the meta-model. Sometimes this means increasing the scope of the metamodel, which may make it difficult to manage or even overlap other readily available meta-models that would be perfectly suited for modeling that part of the domain. Problem Models are conceived for specific domains, but modelers may need to address a broader domain than each model is able to address individually. Wanting to be expressive in a certain domain, software developers sometimes resource to creating different models – for meta-models specifically tailored to different patterns of flexible modeling tools 111 parts of that domain – or to providing different views over the same parts of the domain. However, these meta-models are not necessarily designed for interoperability , which implies a semantic gap between the produced models. Solution Allow two models or domain-specific languages to be linked as needed. Support Linked Models to enable users to connect different but complementary models in a flexible way. This pattern allows to compose models specialized to different parts of the domain. Let the user of the modeling tool define, at the meta-model level, how the models can be linked or, instead, let her connect model elements in an ad hoc fashion (i.e., leave it entirely to the her choice which model elements can be connected, as appropriate). The best approach – meta-model-based or ad hoc – will depend on how the two domains relate to each other and on the amount of flexibility that the modeling tool is intended to provide. On both cases, this pattern may also be referred to as Multiple Meta-Models, in the sense that you may regard the result as a single set of connected model elements (i.e., a single model) that comply to more than one meta-model. Known Uses OMME (Open Meta Modeling Environment) [ VJ10 , VZJ11 ] is a meta-modeling environment implemented on top of the Eclipse platform that provides both textual and graphical notations. Among other features, this environment allows to represent and link models of arbitrary kind – such as, a process model and a data model – allowing to choose the models which fit best in a given situation. Chiprianov et al propose a language tool for telecommunication network designers that provides a partial syntactic and semantic automatic interoperability between different languages, corresponding to different viewpoints used in the definition of a telecommunication service [CKR10]. Microsoft Visio and similar tools allow users to choose between multiple stencils, thus allowing to combine modeling elements from multiple sets (i.e., multiple metamodels). In practice, the connections established using this approach arbitrarily link different models, while their meta-models remain independent. Despite its flexibility, the communication within a team using such models is sometimes more difficult than using completely meta-model driven models, as the semantics of the connections 118 patterns catalog Chapter 6 The Adaptive Software Artifacts Approach Throughout a project’s life, different software artifacts are created and evolved. They take part of the sense-making process in which team members identify recurring information structures that underlie a given body of knowledge. The team may need to capture, share and reason about the ideas in that body of knowledge to discover how they can be structured, therefore they may first capture them as free-form contents, like text documents and, only afterwards, capture them as increasingly more specialized artifacts, such as task descriptions, models and source code. On the one hand, capturing structure explicitly makes information more concrete, unambiguous and terse. On the other hand, free-form contents have the benefit of not being subject to structural constraints, which is of importance during exploratory work. Other differences are that free-form contents don’t directly support sharing information structure between team members and information is not easy to automate – e.g., the cost of maintaining free-form contents is high, as keeping their consistency requires continuous review. Moreover, organizing and classifying information for efficient access is often difficult and classification schemes may also need to be constantly updated to reflect the evolving body of information. The Adaptive Software Artifacts approach is described in detail in the following sections. It combines the benefits of free-form and structured contents with the objective of making information within software development teams easier to use and evolve, especially in the context of medium-to-large projects, where the amount of knowledge involved easily heightens these concerns. This chapter starts by describing the approach itself (Section 6.1) and goes on to describe respectively the design principles and the concrete activities of the approach 120 the adaptive software artifacts approach (Sections 6.2and 6.3). It concludes comparing the Adaptive Software Artifacts approach is with other similar approaches (Section 6.4). 6.1Approach Concerns and Goals Traditional software artifacts constrain information by enforcing a set of rules, which bind artifacts to a pre-determined information structure defined by the kind of artifact in question. The Adaptive Software Artifacts approach is designed with flexibility in mind, and enables to: a) create user-defined types of artifacts – i.e., information with a custom-tailored structure, to fit the specific project’s needs; and b) change such types of artifacts, to better support the knowledge evolution needs of the project – these artifacts are adaptive in the sense that their attributes and relations with other artifacts don’t need to be established from the start and can be freely and easily evolved by the users. Information based on adaptive software artifacts is not as bound to strict constraints like other structured artifacts usually are. Furthermore, the benefits over the traditional dichotomy between structured and free-form contents extend beyond the support to expressing ad hoc knowledge structures explicitly. The approach allows the consistency of the contents to be more easily maintained, by making it easier to see which topics are common across free-text documents and by comparing contents with their expected structure; and it supports access to the text contents through a classification scheme that is dynamically built from the connections between the text contents and the adaptive software artifacts. These goals can be described in the terms of the concerns introduced in Section 4.3. C1. Expression of information structure Structure may come to information through a top-down process. Consider when someone capturing knowledge knows, at the outset, how that particular piece of information can be structured. Her mental model of that knowledge may be close enough to a specific kind of software artifact that she is familiar with, and the decision to create an artifact of that type will not need much thought. But structure may also come to information through a bottom-up process. There may, initially, be a lot of uncertainty as to how the contents can be structured – the mental model for that piece of knowledge may still be very fuzzy and none of the more structured kinds of software artifact be a good fit. Or it may be the case that knowledge is still evolving quickly, and the software artifacts that would fit it are not flexible approach concerns and goals 121 enough to keep up easily with that evolution. In such cases, information capture tends to be avoided altogether until a later stage, or start being done as free-text or other weakly-structured form. It gets captured as more specialized kinds of software artifacts only when the mental model for that particular piece of knowledge becomes clearer or its rate of change is low. Adaptive Software Artifacts support explicitly capturing (and sharing) information structure, allowing contents to be understood (or consumed) more easily. It does this by trying to combine the best of the two processes: team members can use a topdown process by defining new types of software artifacts, with their own specific attributes, and to instance them as needed; and they are also able to use a bottom-up process, and incrementally add structure to the textual contents, as new knowledge is being acquired, thus making new software artifacts to gradually emerge from the text contents. Commonalities can be identified in these emergent artifacts, at which point new types of artifacts can be explicitly defined. Most tools that allow capturing any kind of software knowledge support a topdown process. With the Adaptive Software Artifacts approach users are able to mix and match the two processes in an integrated environment such as a software forge. Additionally, independently of the process through which the artifacts come to be, they can be evolved without being subject to the hard rules that non-adaptive software artifacts are bound to, as it may be expected that their structure complies to a specific type of artifact, but such compliance is not necessarily required. C2. Consistency maintenance By easing the creation of structured contents, the Adaptive Software Artifacts approach opens the possibility for several use cases. By creating adaptive software artifacts from the free-text contents, we are both identifying and abstracting the key topics of the text and connecting the resulting abstractions to their respective documents. This allows us to assess if introducing changes to a free-text document might have an impact on the consistency of others – when changing a document about a certain topic it may be brought to the user’s attention what other pages refer to the same topic. Additionally, when contents are structured and associated with a type, one may assess their consistency towards their type. Easing consistency maintenance means easing the evolution of existing contents and the creation on new ones. 122 the adaptive software artifacts approach C3. Classification of the contents Looking for contents within a large collection of free-text documents can be difficult, especially when the information needs are ill-defined. But the contents of an Adaptive Software Artifact can be created organically from the textual contents, and they may be seen as descriptions of those textual contents. Therefore, they can be used effectively as an indexing mechanism and a rich classification scheme can be automatically derived from them. Such an always-updated index has a low cost of creation and evolution and allows information to be more easily found and thus consumed. 6.2Design Principles This approach assumes the development of tools and environments that support it. Chapter 7presents a reference architecture and implementation of such a tool, the development of which was based on the set of design principles or high-level requirements introduced in the following paragraphs. Wiki Design Principles. The success of wikis as collaborative authoring platforms owes much to the set of principles used to design them 1 . These principles show themselves important in the context of wikis but are abstract enough to be used to design other contents-creation tools. They don’t directly address any of the main goals and research problems described in the previous sections of this chapter but take into account the collaborative nature of software development. Integrated Environment. Derives from acknowledging the need to make new development tools available within an environment already familiar and used by software developers and to take a more holistic approach to knowledge capture. An Integrated Environment (p. 76)supports the maintenance of consistency (C2, p. 51 and 120)by keeping related contents of different natures easily accessible from each other. Domain-Structured-Information. This principle is illustrated by the DomainStructured Information (p. 74)pattern, but also by the Controlled Vocabulary (p. 97)pattern. It helps to express and organize contents according to their domain (C1, p. 51 and 120). Indirectly, it contributes to Information Proximity (p. 66), which helps Co-Evolving (p. 71)related documentation fragments, and thus preserving their consistency (C2, p. 51 and 121). 1The design principles of wikis are described in more detail in Section 2.3.2. activities 123 Structure Co-Evolution. This principle is strongly inspired by Model CoEvolution (p. 104)and to some extent by the more general Co-Evolution (p. 71). It consists of the mechanisms to evolve structured contents while maintaining them consistent (C2, p. 51 and 121). Flexible Structure. The bottom-up expression of structured information (C1, p. 51 and 120)implies flexibility requirements that can be enunciated as Formalization (p. 108)– adding structure to free-form contents – and as Meta-modeling By Example (p. 106)– identifying the structure commonalities of a body of contents, and make those commonalities explicit. Automatic Index. Improving the access and classification (C3, p. 51 and 122)can be addressed through a combination of an Index (p. 82), a Controlled Vocabulary (p. 97)and Domain-Structured Information (p. 74). 6.3Activities The approach may be decomposed as a specific set of activities. These activities are described next in this section and are summarized and illustrated at an abstract level in Figure 6.1. They should be used as a set of concrete feature requirements when developing tools to support the approach. 6.3.1Creation Creation activities are those focused on the capture of structured contents. Namely, the environment should support the creation of new kinds of software artifact by developers. Such artifact types are not pre-determined during the conception of the environment, so they won’t provide some of the most specialized behavior of any built-in artifact types that the environment may offer, but they will allow the users to specify their own data-centric structures. Meta-modeling techniques may be used to support the definition of such artifacts, their properties and the relations to other artifacts. A1. Create artifacts and artifact types. This activity allows defining new artifact types and new artifacts, and specifying their attributes and attribute values. Additional expressive power is given by an inheritance mechanism, which allows an artifact type to be reused by another artifact type, and by the possibility of defining domains and cardinalities for the attributes of an artifact type. 124 the adaptive software artifacts approach Creation Reader Guidance Create artifacts and artifact types attribute1 = "hello" attribute2 = 42 artifact:Type Attribute1 : str Attribute2: num Type Create from textual contents name="Auth. sys" artifact: With support for: •Inheritance •Instantiation •Attribute types •Cardinality Based on text expressions, or taking advantage of the doc. structure (see "Auth. Guidance"). Creator Guidance Changing the types Changing the artifacts •Extract super-artifact •Pull attributes up/down •Rename attribute •Change attribute domain •Split a type •Merge types •Change type of an artifact •Split an artifact •Merge artifacts Allow changes to a type to propagate to its artifacts, in whichever way is appropriate Free-text consistency maintenace •Show other free-text documents that mention the same artifacts Artifact consistency maintenance Smart Aggregation •Aggregate artifacts by known attributes •Aggregate artifacts by known type Navigation •Navigate type relations to ultimately reach the documents Recommend connection to existing artifacts name="Auth. sys" artifact: Use structured contents to help in the creation of more connections with text. You mention "Auth. system". Do you mean this artifact here? Yes Recommend artifact attributes from its type Allow to add values for attributes with unfulfilled cardinalities. attribute1 = "hello" attribute2 = 42 artifact:Type Recommend type changes from artifact values Recommend changes to the type based on the values that its artifacts actually have. attribute1 = "hello" attribute2 = 42 artifact:Type Attribute1 : str [0..1] Attribute2: num[1..*] Type Attribute1 : str [0..1] Attribute2: num[1..*] Type You've got several artifacts with an "attribute2", Should this be in their type? Yes Yes Use heuristics to help transform document structure into domain structure. Yes Recommend artifacts from document structure name="Auth. sys" artifact: name="Auth. sys" artifact: name="Auth. sys" artifact: You've turned one of these list items into an artifact. Do you want to do the same for the others too? Change Impact Awareness A1. A2. A3. A4. A5. A6. A7. A9. A10. A12. A11. A8. Missing a "attribute2" for any chance? Co-evolution of Structured Contents • Show adaptive artifacts related to the current one Figure 6.1:Overview of the activities of the Adaptive Software Artifacts approach. The artifacts created by users may be of an unspecified type, or they may be of one of the artifact types that may have been previously created. In the later case, the connection between the artifact and its type (i.e., the instantiation link) is a point of flexibility. In other words, even though an artifact has a type, it doesn’t necessarily have to follow the structure that the type defines – it may not have all the attributes defined by the type, or not strictly follow their domain and cardinalities, and may even define its own specific attributes. A2. Create from textual contents. The creation of artifacts can be done taking textual contents as a starting point. Without ever losing the context of the free-text document, users should be able to create a new artifact from a text selection. That text fragment can be used as the value for a default attribute of the newly created artifact, which may afterwards be completed with more/other attributes and values. The goal is to reduce the effort in the incremental process of providing activities 125 more and more domain-oriented structure. 6.3.2Reader Guidance These activities address the classification and findability of the contents. Namely, they consist of providing a subject index that is automatically assembled using the contents of the available Adaptive Software Artifacts, and that documentation consumers may use to find the free-form contents that they may need. A3. Smart Aggregation. Artifacts may be aggregated by their common features, the most obvious of which is their type. Although users may create untyped artifacts, even these will have their own internal structure, such as attributes and relations with other artifacts, so even these can be aggregated by their common structural elements. When a great quantity of untyped adaptive software artifacts exists, each with their own internal structure, this activity helps to identify commonalities in the information structure that may come to be expressed as new artifact types. A4. Navigation. Users should be able to navigate the subject index built by aggregating the adaptive software artifacts’ contents, to find a particular adaptive software artifact that they may be looking for, or to find the set of free-text documents in which it is used. Such a list of aggregated artifacts effectively constitutes a subject-index of the free-form contents in the platform. 6.3.3Co-evolution of Structured Contents While having structured contents naturally makes inconsistencies easier to detect, these inconsistencies are not necessarily easier to correct. This is due to some extent to the flexibility supported by the creation activities, which frees information capture from some barriers to being evolved but also make inconsistencies easier to surface. Co-evolution activities have the goal of making structured contents as easy to evolve as free-form contents. They ensure that artifacts and their underlying model will evolve together and are kept consistent. They allow to automatically update related pieces of information so that they are kept consistent, similar to how a refactoring tool supports applying consistent transformations to source code. A5. Changing the types. Provide information creators a predefined set of transformations useful when changing artifact types, and make these changes 126 the adaptive software artifacts approach propagate to its artifacts, in whichever way is appropriate. These transformations may draw inspiration from code refactorings. A few useful examples could be: Extract super-artifact – Split the attributes of an artifact type into two different artifact types connected by an inheritance relation. All the artifacts of the original artifact type will become instances of the new artifact sub-type. Pull attributes up/down – Move an attribute up/down in the inheritance chain to a artifact super/sub-type. Rename attribute – Change the name of an attribute of an artifact type and choose whether to cascade that change to the same attributes of all its artifacts. Change attribute domain – Change the type of an attribute of an artifact type, and choose whether to cascade that change to the same attributes of all its artifacts. Split a type – Split the attributes of an artifact type into two different artifact types, and decide whether the artifacts of the original type should also be split among the two new types by a chosen criteria. Merge types – Merge two artifact types into a single one. The artifacts of both types all become instances of the new type. A6. Changing the artifacts. Provide a set of structure transformation operations at the artifact level. Examples include: Change type of an artifact – Switch (or remove) the type of an artifact. Split an artifact – Split the attributes (and respective values) of an artifact into two different artifacts. Merge artifacts – Join the attributes of two artifacts into a single artifact of the chosen type. 6.3.4Creator Guidance This set of activities guides team members in the creation of structure, regardless of whether it is created through a top-down or bottom-up process. Such guidance is offered as a set of suggestions or recommendations that are given when contents are being created. These activities are inspired by the Time-shifted Co-evolution variant of the Co-evolution pattern. activities 127 Time-shifted co-evolution does not force consistency. It allows artifacts to diverge from a model that was established for them, but makes developers aware of that divergence, so that they may restore consistency if, and when, they wish to do so. These activities make it easier to maintain the real value of artifacts throughout their lifetime. A7. Recommend connection to existing artifacts. Identify terms in the textual contents that match existing adaptive software artifacts and suggest the creation of new connections between the two. A8. Recommend artifacts from document structure. Free-text cannot be said to allow the expression of domain structure but it supports a document-oriented structured made of some elements – such as sections, paragraphs, lists, etc. – that sometimes correlate strongly with an underlying domain structure. This activity must rely on a few heuristics to help transform a free-text document structure into domain structure. For example, if one of the items of a list in a document is made into an adaptive software artifact, the remaining items of the list could likely originate adaptive software artifacts of the same type. A9. Recommend artifact attributes from its type. The environment does not enforce that artifacts must obey the structure defined by its type but suggests information creators to add values for attributes with unfulfilled cardinalities. A10. Recommend type changes from artifact values. The environment also suggests changes to the type based on the values that its artifacts actually have. 6.3.5Change Impact Awareness These activities focus on assisting information creators in contexts where it is difficult to detect if inconsistencies are being introduced. A11. Free-text consistency maintenance. When editing a given free-text document, the environment should point out other documents that mention the same artifacts, as it is likely that they address the same topics. A12. Artifact consistency maintenance. When editing a given adaptive software artifact, the environment should point out other adaptive software artifacts that are related to it, either because they directly refer the one being edited or because they are mentioned by the same free-text documents. 230 experiment data analysis 12 3 4 5 2 4 6 8 10 (a) Control Group 12 3 4 5 2 4 6 8 10 (b) Experimental Group Figure E.38: Histogram of the answers to the questionnaire item UN2–The way in which the information was organized and linked allowed me to understand it more easily. 12 3 4 5 1 2 3 4 5 6 7 (a) Control Group 12 3 4 5 1 2 3 4 5 6 7 (b) Experimental Group Figure E.39: Histogram of the answers to the questionnaire item CO1–The information that was made available was often inconsistent. 12 3 4 5 1 2 3 4 5 6 7 8 (a) Control Group 12 3 4 5 1 2 3 4 5 6 7 8 (b) Experimental Group Figure E.40: Histogram of the answers to the questionnaire item CO2–I don’t have a good perception if the information that was available to me was consistent or not. Publications Many of the materials in this thesis have appeared in the following publications. Peer-reviewed Conference Papers1 [ CA13 ] F. F. Correia and A. Aguiar, “Patterns of Flexible Modeling Tools”, in Proceedings of the 20th Conference on Pattern Languages of Programs (PLoP), Allerton, Illinois, USA, 2013. [ MCY+11 ] P. Matsumoto, F. F. Correia , J. Yoder, E. Guerra, H. S. Ferreira, and A. Aguiar, “AOM Metadata Extension Points”, in Proceedings of the 18th Conference on Pattern Languages of Programs (PLoP), Portland, Oregon, USA, 2011. [ CA11 ] F. F. Correia and A. Aguiar, “Patterns of Information Classification”, in Proceedings of the 18th Conference on Pattern Languages of Programs (PLoP), Portland, Oregon, USA, 2011. [ FCAY11 ] H. S. Ferreira, F. F. Correia , A. Aguiar, and J. Yoder, “The Lazy Semantics Pattern on the context of Meta-Architectures”, in Proceedings of the 2nd Asian Conference on Pattern Languages of Programs (PLoP), Tokyo, Japan, 2011. [ FCYA10 ] H. S. Ferreira, F. F. Correia , J. Yoder, and A. Aguiar, “Core Patterns of Object-Oriented Meta-Architectures”, in Proceedings of the 17th Conference on Pattern Languages of Programs (PLoP), Reno, Nevada, USA, 2010. [ CFFA09b ] F. F. Correia , H. S. Ferreira, N. Flores, and A. Aguiar, “Patterns for Consistent Software Documentation”, in Proceedings of the Pattern Languages of Programs (PLoP), Chicago, Illinois, USA, 2009. [ FCW08 ] H. S. Ferreira, F. F. Correia , and L. Welicki, “Patterns for Data and Metadata Evolution in Adaptive Object-Models”, in Proceedings of the 15th Conference on Pattern Languages of Programs (PLoP), Nashville, Tennessee, USA, 2008. 1 The peer-review of papers accepted to the Pattern Language of Programs conference is based on a shepherding process and is followed by a writer’s workshop. Both review techniques are a source of important feedback and actively support authors in improving their works. 232 publications Peer-reviewed Journal Papers [ FCAF10 ] H. S. Ferreira, F. F. Correia , A. Aguiar, and J. P. Faria, “Adaptive Object-Models: a Research Roadmap”, IARIA Journal,2010. Peer-reviewed Workshop Papers [ CFFA12 ] F. F. Correia , N. Flores, H. S. Ferreira, and A. Aguiar, “Assessing Tools for Software Development — An overview of three user evaluations”, USER 2012 - User evaluation for Software Engineering Researchers Workshop, Zürich, Switzerland, 2012. [not part of proceedings] [ FCA09 ] H. S. Ferreira, F. F. Correia , and A. Aguiar, “Design for an Adaptive Object-Model Framework: An Overview”, in 4th Workshop on [email protected] at MODELS 09,2009, pp. 71-80. [ SMA+09 ] A. R. Silva, D. Martinho, A. Aguiar, N. Flores, F. F. Correia , and H. S. Ferreira, “An Implementation Model for Agile Business Process Tools”, IWODE 2009 - International Workshop on Organizational Design and Engineering, Lisbon, Portugal, 2009. [not part of proceedings] [ Cor08 ] F. F. Correia , “Extending and Integrating Wikis to Improve Software Documentation”, Wikis4SE - Wikis for Software Engineering Workshop @ WikiSum 2008, Porto, Portugal, 2008. [not part of proceedings] Peer-reviewed Posters [ Cor13 ] F. F. Correia , “Documenting Software Using Adaptive Software Artifacts”, in Proceedings of the 4th annual conference on Systems, programming, and applications: software for humanity (SPLASH), Indianapolis, Indiana, USA, 2013. [ CFFA09a ] F. F. Correia , H. S. Ferreira, N. Flores, and A. Aguiar, “Incremental Knowledge Acquisition in Software Development Using a Weakly-Typed Wiki”, in Proceedings of the 5th International Symposium on Wikis and Open Collaboration (WikiSym), Orlando, Florida, USA, 2009. Doctoral Symposiums [ Cor10 ] F. F. Correia , “Supporting the Evolution of Software Knowledge With Adaptive Software Artifacts”, in Proceedings of the ACM international conference companion on object oriented programming systems languages and applications companion (OOPSLA), Reno, Nevada, USA, 2010, pp. 231-32. [ CA09 ] F. F. Correia and A. Aguiar, “Software Knowledge Capture and Acquisition: Tool Support for Agile Settings”, in Proceedings of the Fourth International Conference on Software Engineering Advances (ICSEA), Porto, Portugal, 2009, pp. 542-547. [ CF08 ] F. F. Correia and H. S. Ferreira, “Trends on Adaptive Object Model Research”, in Proceedings of the Doctoral Symposium on Informatics Engineering 2008 (DSIE), Porto, Portugal, 2008. References [ABK+06] Steve Abrams, Bard Bloom, Paul Keyser, Doug Kimelman, Eric Nelson, Wendy Neuberger, Tova Roth, Ian Simmonds, Steven Tang, and John Vlissides, Architectural thinking and modeling with the architects’ workbench, IBM Systems Journal 45 (2006), no. 3,481–500. Cited on pp. 43 and 109. [ACPT01] G. Antoniol, B. Caprile, A. Potrich, and P. Tonella, Design-code traceability recovery: selecting the basic linkage properties, Science of Computer Programming 40 (2001), no. 2-3,213 – 234. Cited on p. 32. [AD05a] Ademar Aguiar and Gabriel David, Patterns for documenting frameworks – part i, Conference Proceedings Of The Second, Third And Fourth Nordic Conference On Pattern Languages Of Programs Vikingplop (Helsinki, Finland), September 2005. Cited on p. 65. [AD05b] Ademar Aguiar and Gabriel David, WikiWiki weaving heterogeneous software artifacts, Proceedings of the 2005 international symposium on Wikis (San Diego, California, USA), ACM, 2005, pp. 67–74. Cited on pp. 11,18, and 70. [AD06a] Ademar Aguiar and Gabriel David, Patterns for documenting frameworks – part II, Proceedings of EuroPLoP 2006 (Irsee, Germany), July 2006. Cited on p. 65. [AD06b] ,Patterns for documenting frameworks – part III, Proceedings of the 2006 conference on Pattern languages of programs (Portland, Oregon, USA), October 2006. Cited on p. 65. [AD06c] ,Patterns for documenting frameworks: customization, Proceedings of the 2006 conference on Pattern languages of programs (Portland, Oregon, USA), ACM, 2006, pp. 1–10. Cited on p. 65. [AD07] ,Patterns for documenting frameworks - process, Proceedings of SugarLoafPLoP 2006 (Recife, Brazil), May 2007. Cited on pp. 14 and 65. [AD11] ,Patterns for effectively documenting frameworks, Transactions on Pattern Languages of Programming II (David Hutchison, Takeo Kanade, Josef Kittler, Jon M. Kleinberg, Friedemann Mattern, John C. Mitchell, Moni Naor, Oscar Nierstrasz, C. Pandu Rangan, Bernhard Steffen, Madhu Sudan, Demetri Terzopoulos, Doug Tygar, Moshe Y. Vardi, Gerhard Weikum, James Noble, Ralph Johnson, Paris Avgeriou, Neil B. Harrison, and Uwe Zdun, eds.), vol. 6510, Springer Berlin Heidelberg, Berlin, Heidelberg, 2011, pp. 79– 124. Cited on p. 14. [ADP03] Ademar Aguiar, Gabriel David, and Manuel Padilha, XSDoc: an extensible wiki-based infrastructure for framework documentation, Jornadas de Ingeniería del Software y Bases de Datos (Alicante, Spain), October 2003. Cited on pp. 18 and 70. [AEQ99] Jim Arlow, Wolfgang Emmerich, and John Quinn, Literate modelling - capturing business knowledge with the UML, The Unified Modeling Language, UML’98 - Beyond the Notation. First International Workshop, June 1998 (Mulhouse, France) (Jean Bézivin and PierreAlain Muller, eds.), vol. 1618, Springer, 1999, pp. 189–199. Cited on pp. 25,26, and 113. 234 REFERENCES [AG05] Katja Andresen and Norbert Gronau, An approach to increase adaptability in ERP systems, Managing Modern Organizations with Information Technology: Proceedings of the 2005 Information RecourcesManagement Association International Conference (San Diego, California, USA), Idea Group Publishing, May 2005, pp. 883–885. Cited on pp. 39 and 41. [Agu03] Ademar Aguiar, A minimalist approach to framework documentation, Ph.D. thesis, Faculdade de Engenharia da Universidade do Porto, September 2003. Cited on pp. 11,16,18,24, 26, and 31. [AL07] Sören Auer and Jens Lehmann, What have innsbruck and leipzig in common? extracting semantics from wiki content,4th European Semantic Web Conference (ESWC 2007) (2007), 503–517. Cited on p. 18. [Ale77] Christopher Alexander, A pattern language: towns, buildings, construction, Oxford University Press, New York, USA, 1977. Cited on pp. 59,60,61, and 187. [Ale79] Christopher Alexander, The timeless way of building, Oxford University Press, New York, USA, 1979 (English). Cited on p. 59. [Amb] Scott Ambler, Agile/Lean Documentation: Strategies for Agile Software Development, http: //www.agilemodeling.com/essays/agileDocumentation.htm , [ accessed on 2014/05/15]. Cited on p. 65. [Amb02] ,Agile modeling: Effective practices for eXtreme programming and the unified process, 1st ed., Wiley, April 2002. Cited on p. 65. [AN04] Jim Arlow and Ila Neustadt, Enterprise patterns and MDA: Building better software with archetype patterns and UML, Addison-Wesley Professional, 2004. Cited on p. 26. [And01] Hugh Anderson, Formalization and ’literate’ programming, Software Engineering Conference, 2001. APSEC 2001. Eighth Asia-Pacific, 2001, pp. 39–44. Cited on p. 25. [ANS05] ANSI/NISO, Z39.19. Guidelines for the construction, format, and management of monolingual controlled vocabularies,2005. Cited on p. 98. [Arl06] Jim Arlow, Increase the accessibility and comprehensibility of a visual model with literate modeling, http://www.informit.com/articles/article.aspx?p=460398 [ accessed on 2014/05/15], April 2006. Cited on p. 26. [Arm00] Phillip G. Armour, The five orders of ignorance, Commun. ACM 43 (2000), no. 10,17–20. Cited on p. 8. [Ars00] Ali Arsanjani, Rule object: A pattern language for adaptive and scalable business rule construction, Proceedings of PLOP2000,2000. Cited on p. 45. [Art00] Eric Artzt, Autoduck user’s guide,2000. Cited on pp. 28 and 70. [Atl] Atlassian, Confluence — enterprise collaboration and wiki software, http://www. atlassian.com/software/confluence/ [ accessed on 2012/04/01 ] . Cited on p. 18. [BA04] Kent Beck and Cynthia Andres, Extreme programming explained: Embrace change (2nd edition), Addison-Wesley Professional, 2004. Cited on p. 9. [BBvB+01] Kent Beck, Mike Beedle, Arie van Bennekum, Alistair Cockburn, Ward Cunningham, Martin Fowler, James Grenning, Jim Highsmith, Andrew Hunt, Ron Jeffries, Jon Kern, Brian Marick, Robert C. Martin, Steve Mellor, Ken Schwaber, Jeff Sutherland, and Dave Thomas, Manifesto for agile software development, http://agilemanifesto.org/ [accessed on 2008/06/24],2001. Cited on pp. 2and 9. REFERENCES 235 [BHS07] Frank Buschmann, Kevlin Henney, and Douglas C Schmidt, Pattern-oriented software architecture: on patterns and pattern languages, vol. 5, Wiley, Chichester, UK; Hoboken, New Jersey, USA, 2007 (English). Cited on pp. 60 and 187. [BKM00] Greg Butler, Rudolf K. Keller, and Hafedh Mili, A framework for framework documentation, ACM Comput. Surv. 32 (2000), 15. Cited on p. 14. [BM06] Joachim Bayer and Dirk Muthig, A view-based approach for improving software documentation practices, Engineering of Computer Based Systems, 2006. ECBS 2006.13th Annual IEEE International Symposium and Workshop on, 2006, p. 10 pp. Cited on p. 70. [BMNS11] Sabine Buckl, Florian Matthes, Christian Neubert, and Christian M. Schweda, A lightweight approach to enterprise architecture modeling and documentation, Information Systems Evolution (Will Aalst, John Mylopoulos, Norman M. Sadeh, Michael J. Shaw, Clemens Szyperski, Pnina Soffer, and Erik Proper, eds.), vol. 72, Springer Berlin Heidelberg, Berlin, Heidelberg, 2011, pp. 136–149. Cited on p. 21. [BMR+96] Frank Buschmann, Regine Meunier, Hans Rohnert, Peter Sommerlad, and Michael Stal, Pattern-oriented software architecture: a system of patterns, vol. 1, Wiley, Chichester, New York, USA, 1996. Cited on p. 60. [BPB10] Karin Breitman, Oscar Pastor, and Simone Barbosa, Flexible narrative representations: Bridging the gap between formal models and informal representations, ICSE 2010 Workshop on Flexible Modeling Tools (FlexiTools 2010) (Cape Town, South Africa), vol. 35,2010, pp. 37–38. Cited on p. 113. [Bri03] Lionel C. Briand, Software documentation: how much is enough?, Software Maintenance and Reengineering, 2003. Proceedings. Seventh European Conference on, 2003, pp. 13–15. Cited on p. 11. [CA09] Filipe F. Correia and Ademar Aguiar, Software knowledge capture and acquisition: Tool support for agile settings, Fourth International Conference on Software Engineering Advances (Porto, Portugal), September 2009, pp. 542–547. Cited on p. 232. [CA11] ,Patterns of information classification, Proceedings of the 18th Conference on Pattern Languages of Programs (Portland, Oregon, USA), ACM, 2011. Cited on pp. 63 and 231. [CA13] ,Patterns of flexible modeling tools, Proceedings of the 20th Conference on Pattern Languages of Programs (Allerton, Illinois, USA), ACM, 2013. Cited on pp. 63 and 231. [CB91] David Cordes and Markus Brown, The literate-programming paradigm, Computer 24 (1991), 52–61. Cited on p. 22. [CDREP08] Antonio Cicchetti, Davide Di Ruscio, Romina Eramo, and Alfonso Pierantonio, Automating co-evolution in model-driven engineering, Enterprise Distributed Object Computing Conference, 2008. EDOC’08.12th International IEEE, 2008, pp. 222–231. Cited on p. 106. [CF08] Filipe F. Correia and Hugo S. Ferreira, Trends on adaptive object model research, Proceedings of the Doctoral Symposium on Informatics Engineering 2008 (Porto, Portugal), FEUP, January 2008. Cited on p. 232. [CFFA09a] Filipe F. Correia, Hugo S. Ferreira, Nuno Flores, and Ademar Aguiar, Incremental knowledge acquisition in software development using a weakly-typed wiki, Proceedings of the 5th International Symposium on Wikis and Open Collaboration (Orlando, Florida, USA), ACM, October 2009. Cited on pp. 185 and 232. [CFFA09b] ,Patterns for consistent software documentation, Proceedings of the 16th Conference on Pattern Languages of Programs (Chicago, Illinois, USA), ACM, August 2009. Cited on pp. 63,106,110,112,114, and 231. 236 REFERENCES [CFFA12] Filipe F. Correia, Nuno Flores, Hugo S. Ferreira, and Ademar Aguiar, Assessing tools for software development — an overview of three user evaluations, USER 2012 Workshop (Zürich, Switzerland), 2012. Cited on p. 232. [CHK+01] N. Chapin, J. E Hale, K. M Khan, J. F Ramil, and W. G Tan, Types of software evolution and software maintenance, Journal of Software Maintenance and Evolution Research and Practice 13 (2001), no. 1,3–30. Cited on p. 35. [CHRP03] Li-Te Cheng, Susanne Hupfer, Steven Ross, and John Patterson, Jazzing up eclipse with collaborative tools, Proceedings of the 2003 OOPSLA workshop on eclipse technology eXchange (Anaheim, California, USA), ACM, 2003, pp. 45–49. Cited on pp. 9and 33. [CJBV13] Jeffrey C. Carver, Natalia Juristo, Maria Teresa Baldassarre, and Sira Vegas, Replications of software engineering experiments, Empirical Software Engineering (2013), 1–10 (en). Cited on p. 159. [CJMS10] Jeffrey Carver, Letizia Jaccheri, Sandro Morasca, and Forrest Shull, A checklist for integrating student empirical studies with research and teaching goals, Empirical Software Engineering 15 (2010), no. 1,35–59. Cited on pp. 57,154,158, and 178. [CKR10] Vanea Chiprianov, Yvon Kermarrec, and Siegfried Rouvrais, Meta-tools for software language engineering: a flexible collaborative modeling language for efficient telecommunications service design, ICSE 2010 Workshop on Flexible Modeling Tools (FlexiTools 2010) (Cape Town, South Africa), 2010. Cited on p. 111. [Cor08] Filipe F. Correia, Extending and integrating wikis to improve software documentation, Wikis4SE WikiSym 2008 (Porto, Portugal), September 2008. Cited on p. 232. [Cor10] ,Supporting the evolution of software knowledge with adaptive software artifacts, Proceedings of the ACM international conference companion on Object oriented programming systems languages and applications companion (Reno/Tahoe, Nevada, USA), SPLASH ’10, ACM, 2010, pp. 231–232. Cited on pp. 103 and 232. [Cor13] ,Documenting software using adaptive software artifacts, Proceedings of the ACM international conference companion on Object oriented programming systems languages and applications companion, SPLASH ’13, ACM, 2013. Cited on pp. 103 and 232. [CS96] Bart Childs and Johannes Sametinger, Literate programming and documentation reuse, Software Reuse, 1996., Proceedings Fourth International Conference on, 1996, pp. 205– 214. Cited on p. 76. [CSGW11] Hyun Cho, Y. Sun, J. Gray, and Jules White, Key challenges for modeling language creation by demonstration, ICSE 2011 Workshop on Flexible Modeling Tools (FlexiTools 2011) (Waikiki, Hawaii, USA), 2011. Cited on p. 108. [Cuna] Ward Cunningham, c2.com — wiki, http://c2.com/cgi/wiki [ accessed on 2007/12/01]. Cited on p. 15. [Cunb] ,c2.com — wiki design principles, http://c2.com/cgi/wiki? WikiDesignPrinciples [accessed on 2007/12/01]. Cited on pp. 15 and 185. [Cun93] ,The WyCash portfolio management system, SIGPLAN OOPS Mess. 4 (1993), no. 2, 29–30. Cited on p. 36. [Cun06] ,Design principles of wiki: how can so little do so much?, http://c2.com/doc/ wikisym/WikiSym2006.pdf [accessed on 2014/05/15],2006. Cited on p. 15. [Dav05] Thomas H. Davenport, Thinking for a living: How to get better performances and results from knowledge workers, Harvard Business Press, September 2005. Cited on pp. 8and 186. REFERENCES 237 [DGB07] Sergiu Dumitriu, Marta Gîrdea, and Sabin C. Buraga, From information wiki to knowledge wiki via semantic web technologies, Innovations and Advanced Techniques in Computer and Information Sciences and Engineering (Tarek Sobh, ed.), Springer Netherlands, January 2007, pp. 443–448. Cited on p. 21. [DGLP08] Marco D’Ambros, Harald Gall, Michele Lanza, and Martin Pinzger, Analysing software repositories to understand software evolution, Software Evolution, Springer, 2008, pp. 37–67. Cited on p. 38. [DOS10] Michael Desmond, Harold Ossher, and Ian Simmonds, Towards smart office tools, SPLASH 2010 Workshop on Flexible Modeling Tools (FlexiTools 2010@SPLASH) (Reno, Nevada, USA), 2010. Cited on p. 108. [DRR+05] Björn Decker, Eric Ras, Jörg Rech, Bertin Klein, and Christian Hoecht, Self-organized reuse of software engineering knowledge supported by semantic wikis, Proceedings of the Workshop on Semantic Web Enabled Software Engineering (SWESE), November 2005. Cited on p. 76. [DV99] Angela M. Dean and Daniel Voss, Design and analysis of experiments,1st ed. 1999. corr. 2nd printing ed., Springer, 1999. Cited on pp. 57,154, and 159. [Edga] Edgewall Software, The Trac Project — Integrated SCM & Project Management, http: //trac.edgewall.org/ [ accessed on 2009/12/01 ] . Cited on pp. 3,18,19,77, and 132. [Edgb] ,Trac Component Architecture, http://trac.edgewall.org/wiki/TracDev/ ComponentArchitecture [accessed on 2012/04/01]. Cited on p. 134. [Edgc] ,Trac Extension points, http://trac.edgewall.org/wiki/TracDev/ PluginDevelopment/ExtensionPoints [ accessed on 2012/04/01 ] . Cited on p. 134. [Eva03] Eric Evans, Domain-Driven design: Tacking complexity in the heart of software,1ed., Addison Wesley, 2003. Cited on p. 40. [FB99] Martin Fowler and Kent Beck, Refactoring: improving the design of existing code, AddisonWesley Professional, 1999. Cited on p. 37. [FCA09] Hugo S. Ferreira, Filipe F. Correia, and Ademar Aguiar, Design for an adaptive object-model framework: An overview,4th Workshop on Modelsrun.time at MODELS 09, October 2009, pp. 71–80. Cited on pp. 106,136,142, and 232. [FCAF10] Hugo S. Ferreira, Filipe F. Correia, Ademar Aguiar, and João Pascoal Faria, Adaptive object-models: a research roadmap, IARIA Journal (2010). Cited on pp. 42 and 232. [FCAY11] Hugo S. Ferreira, Filipe F. Correia, Ademar Aguiar, and Joseph Yoder, The lazy semantics pattern on the context of meta-architectures,2011. Cited on pp. 115 and 231. [FCW08] Hugo S. Ferreira, Filipe F. Correia, and Leon Welicki, Patterns for data and metadata evolution in adaptive object-models, Proceedings of the 15th Conference on Pattern Languages of Programs (Nashville, Tennessee, USA), ACM, October 2008. Cited on pp. 38,103,106, 114,116, and 231. [FCYA10] Hugo S. Ferreira, Filipe F. Correia, Joseph Yoder, and Ademar Aguiar, Core patterns of object-oriented meta-architectures, Proceedings of the 17th Conference on Pattern Languages of Programs (Reno, Nevada, USA), ACM, 2010. Cited on pp. 103,104,114,115, and 231. [Fer10] Hugo S. Ferreira, Adaptive object-modeling: Patterns, tools and applications, PhD in computer science, University of Porto, Faculty of Engineering, Rua Dr. Roberto Frias, s/n, 4200-4650 Porto, Portugal, 2010. Cited on pp. 44 and 114. 238 REFERENCES [fit]Fitnesse,http://fitnesse.org/ [accessed on 2011/01/02]. Cited on pp. 18 and 19. [Fow97] Martin Fowler, Analysis patterns: reusable objects models, Addison-Wesley Longman Publishing Co., Inc, Boston, Massachusetts, USA, 1997. Cited on p. 45. [Fow06] Martin Fowler, Writing software patterns, http://www.martinfowler.com/ articles/writingPatterns.html , [ accessed on 2008/09/13 ] ,2006. Cited on p. 61. [Fri95] Lisa Friendly, The design of distributed hyperlinked programming documentation, IWHD’95 (Montpellier, France), 1995. Cited on pp. 28,70, and 75. [Fur10] Jonathan Furner, Folksonomies, Encyclopedia of Library and Information Sciences (Marcia J. Bates and Mary Niles Maack, eds.), CRC Press, Boca Raton, Florida, USA, 3rd ed., 2010, pp. 1858–1866. Cited on p. 95. [FY99] Brian Foote and Joseph Yoder, Big ball of mud, Pattern Languages of Program Design 4, Addison-Wesley, 1999, pp. 653–692. Cited on p. 36. [GAO95] David Garlan, Robert Allen, and John Ockerbloom, Architectural mismatch: Why reuse is so hard, Software, IEEE 12 (1995), no. 6,17–26. Cited on p. 12. [GBeA07] Miguel Goulão and Fernando Brito e Abreu, Modeling the experimental software engineering process, Quality of Information and Communications Technology, 2007. QUATIC 2007. 6th International Conference on the, 2007, pp. 77–90. Cited on p. 160. [GG08] Michael W. Godfrey and Daniel M. German, The past, present, and future of software evolution, Frontiers of Software Maintenance, 2008. FoSM 2008., 2008, pp. 129–138. Cited on pp. 35 and 38. [GGLS11] Gregor Gabrysiak, Holger Giese, Alexander Lueders, and Andreas Seibel, How can metamodels be used flexibly?, ICSE 2011 Workshop on Flexible Modeling Tools (FlexiTools 2011) (Waikiki, Hawaii, USA), 2011. Cited on pp. 103,104,105,107, and 112. [GGS10] Gregor Gabrysiak, Holger Giese, and Andreas Seibel, Using ontologies for flexibly specifying multi-user processes, ICSE 2010 Workshop on Flexible Modeling Tools (FlexiTools 2010) (Cape Town, South Africa), 2010. Cited on p. 105. [GHJV95] Erich Gamma, Richard Helm, Ralph Johnson, and John Vlissides, Design patterns: Elements of reusable object-oriented software, Addison-Wesley Professional, January 1995. Cited on pp. 60,61, and 105. [GJT09] Raghu Garud, Sanjay Jain, and Philipp Tuertscher, Incomplete by design and designing for incompleteness, Design Requirements Engineering: A Ten-Year Perspective, Springer, 2009, pp. 137–156. Cited on p. 35. [GM12] Gary Goertz and James Mahoney, A tale of two cultures: qualitative and quantitative research in the social sciences, Princeton University Press, Princeton, N.J., 2012 (English). Cited on p. 57. [GR13] Olivier Gendreau and Pierre N. Robillard, Knowledge acquisition activity in software development, Advances in Information Systems and Technologies (Álvaro Rocha, Ana Maria Correia, Tom Wilson, and Karl A. Stroetmann, eds.), Advances in Intelligent Systems and Computing, no. 206, Springer Berlin Heidelberg, January 2013, pp. 1–10. Cited on p. 8. [GRS12] Chiara Ghidini, Marco Rospocher, and Luciano Serafini, Conceptual modeling in wikis: a reference architecture and a tool, eKNOW 2012, The Fourth International Conference on Information, Process, and Knowledge Management, 2012, pp. 128–135. Cited on p. 21. REFERENCES 239 [Gru93] Thomas R. Gruber, A translation approach to portable ontology specifications, Knowledge Acquisition 5(1993), no. 2,199–220. Cited on p. 92. [Ham94] John Hamer, Literate programming: a software engineering perspective, Software Education Conference, 1994. Proceedings., 1994, pp. 282–288. Cited on pp. 22,24, and 26. [HHT01] Jochen Hartmann, Shihong Huang, and Scott Tilley, Documenting software systems with views II: an integrated approach based on XML, Proceedings of the 19th annual international conference on Computer documentation (Sante Fe, New Mexico, USA), ACM, 2001, pp. 237–246. Cited on p. 14. [HNA+10] Jonas Helming, Nitesh Narayan, Holger Arndt, Maximilian Koegel, and Walid Maalej, From informal project management artifacts to formal system models, ICSE 2010 Workshop on Flexible Modeling Tools (FlexiTools 2010) (Cape Town, South Africa), 2010. Cited on p. 110. [HS98] Nick Hatzigeorgiu and Apostolos Syropoulos, Literate programming and the "spaniel" method, SIGPLAN Not. 33 (1998), 52–56. Cited on p. 26. [HW99] Myles Hollander and Douglas A. Wolfe, Nonparametric statistical methods,2nd ed., WileyInterscience, January 1999. Cited on p. 171. [IVZ08] Angelo Di Iorio, Fabio Vitali, and Stefano Zacchiroli, Wiki content templating, Proceeding of the 17th international conference on World Wide Web (Beijing, China), ACM, 2008, pp. 615–624. Cited on pp. 17 and 18. [IVZ09] Angelo Di Iorio, Fabio Vitali, and S. Zacchiroli, Web semantics via wiki templating, Handbook of research on Web 2.0,3.0and x.0: technologies, business and social applications, Advances in E-Business Research, San Murugesan Ed., November 2009. Cited on p. 18. [Job93] Chris P. Jobling, The IEEE grumman f14 benchmark problem - an example of literate modelling in ACSL using noweb,1993. Cited on p. 113. [JW97] Ralph Johnson and Bobby Woolf, The type object pattern,1997. Cited on pp. 45 and 136. [Kac12] Michal Kacprzyk, Software knowledge management using wikis : a plugin for weakly typed pages, Ph.D. thesis, Faculdade de Engenharia da Universidade do Porto, Porto, Portugal, February 2012. Cited on p. 185. [Kan] Noah Kantrowitz, TracHacks — IncludeMacro Plugin, http://trac-hacks.org/wiki/ IncludeMacro [accessed on 2012/04/01]. Cited on p. 19. [KB03] Michael A. Katz and Michael D. Byrne, Effects of scent and breadth on use of site-specific search on e-commerce web sites, ACM Transactions on Computer-Human Interaction 10 (2003), no. 3,198–220 (en). Cited on p. 52. [KC02] Andreas Kacofegitis and Neville Churcher, Theme-based literate programming, Software Engineering Conference, 2002. Ninth Asia-Pacific, 2002, pp. 549–557. Cited on pp. 22,24, 27, and 28. [KCM07] Huzefa Kagdi, Michael L. Collard, and Jonathan I. Maletic, A survey and taxonomy of approaches for mining software repositories in the context of software evolution, J. Softw. Maint. Evol. 19 (2007), no. 2,77–131. Cited on p. 38. [KH10] Doug Kimelman and Ken Hirschman, Discussion and MindManager-Based structuring of workshop summary, http://www.ics.uci.edu/~nlopezgi/flexitools/ presentations/SPLASH2010WorkshoponFlexibleModelingTools-v6.swf , [accessed on 2014/05/15],2010. Cited on p. 103. 246 REFERENCES [YBJ01] Joseph W. Yoder, Federico Balaguer, and Ralph Johnson, Architecture and design of adaptive object-models, ACM SIG-PLAN Notices 36 (2001), no. 12,50–60. Cited on pp. 39,45,48, and 136. [YFRT98] Joseph W. Yoder, Brian Foote, Dirk Riehle, and Michel Tilman, Metadata and active object-models, ACM, 1998. Cited on pp. 45 and 114. [ZVB11] Li Zhu, Ivan Vaghi, and Barbara Rita Barricelli, A meta-reflective wiki for collaborative design, Proceedings of the 7th International Symposium on Wikis and Open Collaboration (Mountain View, California, USA), WikiSym ’11, ACM, October 2011, pp. 53–62. Cited on p. 21. Index accountability, 45 adaptability, 41 adaptive object-model, 39,44,56 adaptive software artifacts, 56,119,135, 153,154,184,186 augmented models, 63,112 big ball of mud, 36 blocking, 154 bootstrapping, 64,116 case study, 188 change propagation, 37,125 classification, 51,53,55,122,123,125,153, 157,158,167,174,179,185,187 closing the roof, 64,116 co-evolution, 37,63,71,123,125,126 code annotations, 28 consistency, 13,37,51,52,54,55,121,127, 153,158,176,179,185,187 controlled vocabulary, 63,97,122 data analysis, 158,162,188 data quality, 162 data collection, 156 data source, 155 document free-text, 48 document (free-text), 1,9,11,16,31,48 document (structured), 9,31 domain-driven design, 40 domain-specific language, 41 domain-structured information, 63,74, 122,see also structure, domainstructured contents elucidative programming, 29 everything is a thing, 64,116 experiment, 56,153,179,184,187 experimental design, 154,160 experimental package, 159 experimental subject, 154,161 expressiveness, 11,14,50,52,55,58,75, 79,99,185,187 file, 30,38 flexible modeling tools, 43,56,128 folksonomy, 63,94 formalization, 63,108 free-form contents, 120,129 free-text, see document (free-text), see freeform contents generative programming, 40 history of operations, 64,116 incompleteness, 36 index, 63,82,123 information, 54,185,see also knowledge, capture information proximity, 63,66 information structure, see structure integrated development environment, 47, 129 248 INDEX integrated environment, 63,76,122,see integrated development environment, see software forge jhotdraw, 161 knowledge, 1 acquisition, 8,9,54,55,153,157,158, 160,167,173,175,179,185,187 capture, 8,9,31,52,54,55,119,121, 123,160,185,187,see also expressiveness language piggybacking, 64,116,136,191 lazy semantics, 64,116 linked models, 63,110 literate modeling, 26 literate programming, 21,128 meta-information, 10,13,see also structure meta-modeling by example, 63,106 migration, 38,64,116 model, 13,40,42,48 model co-evolution, 63,104 modeling, 40–42,see also flexible modeling tools bottom-up, 107,see also structure, bottom-up multiple source, 30,67 ontology, 63,91 pattern, 56,184,186,187 pilot experiment, 159 property, 45,137 questionnaire, 156,163,171,180 randomization, 154 refactoring, 37,125,132 reference architecture, 56,131,184 replication, 154,159 reverse literate programming, 26 single source, 30,67,70 software aging, 36 software artifact, 8,9,12,36,37,47,49,54 software evolution, 35,37,39 software forge, 30,47,50,129,132 source code, 10,12 structure, 10,11,13,17,37,48,51,54,55, 120,129,153,157,158,167,175, 179,187 bottom-up, 20,37,117,123,187 domain-agnostic, 11,38,48,58 domain-structured, 12,31,38,58 free-text, see document (free-text) top-down, 20,37,117,187 system memento, 64,116 taxanomy, 85 taxonomy, 63 technical debt, 36 template, 17 creational, 17 functional, 17 lightly constrained, 18,185 weakly typed, 185 theme-based literate programming, 27 thesaurus, 63,88 trac, 19,77,96,132,133 transclusion, 67,70 type object, 45,136,191 type square, 45,191 unified modeling language, 41 user-crafted static meta-model, 63,103 INDEX 249 variability, 41 version control, 38,74 weaki, 185 wiki, 15,70,74,122,128 extended, 17,18 hybrid, 17 semantic, 17,19 250 INDEX