Desenvolupament i integració d'un glossari de termes i registre de sinònims per a ser usat en les eines del sistema PABRE
Abstract
El projecte inclou l'estudi d'alternatives i la implementació d'un glossari, amb uns certs requisits funcionals, a l'eina PABRE-MAN, el desenvolupament de les funcionalitats de gestió del diccionari i la seva integració dins l'aplicació. Inclou estudi d'abast, impacte i cost de la solució.
Full text
Títol: Desenvolupament i integració d'un glossari de termes i registre de sinònims per a ser usat en les eines del sistema PABRE. Volum: 1/1 Alumne: Jonathan Cubel Figueroa Director: Carme Quer Bosor Departament: Enginyeria de Serveis i Sistemes d’Informació (ESSI) Data: Gener 2014
ii
iii DADES DEL PROJECTE Títol del Projecte: Desenvolupament i integració d'un glossari de termes i registre de sinònims per a ser usat en les eines del sistema PABRE. Nom de l’estudiant: Jonathan Cubel Figueroa Titulació: Enginyeria Informàtica Crèdits: 37,5 Director: Carme Quer Bosor Departament: Enginyeria del Software i Sistemes d’Informació (ESSI) MEMBRES DEL TRIBUNAL (nom i signatura) President: Xavier Franch Gutiérrez Vocal: Julita Corbalán González Secretari: Carme Quer Bosor QUALIFICACIÓ Qualificació numèrica: Qualificació descriptiva: Data:
iv
v Content 1 Introduction .......................................................................................................................... 1 2 Project’s Goal ........................................................................................................................ 4 3 Technologies ......................................................................................................................... 5 4 Requirements ........................................................................................................................ 8 4.1 Functional Requirements .............................................................................................. 8 4.2 Non-Functional Requirements .................................................................................... 11 5 Specification ........................................................................................................................ 12 5.1 Conceptual Model ....................................................................................................... 12 5.2 Use Case Model ........................................................................................................... 13 5.3 XSD and example ......................................................................................................... 22 6 Design .................................................................................................................................. 24 6.1 Guideline to Solution ................................................................................................... 25 6.2 Design Conceptual Model ........................................................................................... 28 6.3 Design alternatives ...................................................................................................... 29 6.3.1 Java Serialized ..................................................................................................... 29 6.3.2 XML-File ............................................................................................................... 30 6.3.3 Hibernate ............................................................................................................. 30 6.3.4 Selected Option ................................................................................................... 31 6.4 Special design considerations ..................................................................................... 33 6.5 Definitive Design Conceptual Model and Data Persistence Adapter .......................... 34 6.6 Process design ............................................................................................................. 35 6.6.1 Sequence Diagrams ............................................................................................. 35 7 External Design .................................................................................................................... 41 7.1 Glossary Internal Frame .............................................................................................. 42 7.1.1 Insert hypernym .................................................................................................. 43 7.1.2 Hypernym menu .................................................................................................. 44 7.1.3 Edit Hypernym ..................................................................................................... 45 7.1.4 Error messages .................................................................................................... 46 7.1.5 Insert hyponym ................................................................................................... 47 7.1.6 Hyponym menu ................................................................................................... 48 7.1.7 Hyponym modification ........................................................................................ 49 7.1.8 Search Functionality ............................................................................................ 50
vi 7.1.9 General External Appearance ............................................................................. 51 7.1.10 Glossary Internal Frame Persistence Dialog ........................................................ 52 7.2 Editor Component ....................................................................................................... 52 7.3 AutoHyponymChecker ................................................................................................ 56 7.4 Hyponym Substitution ................................................................................................. 56 7.5 Importation / Exportation of the Glossary .................................................................. 59 7.6 Design of the Presentation Tier ................................................................................... 61 8 Work plan ............................................................................................................................ 63 8.1 Methodology ............................................................................................................... 63 8.1.1 Introduction......................................................................................................... 63 8.1.2 Selected Methodology ........................................................................................ 65 8.2 Task Breakdown and Sprints ....................................................................................... 66 8.3 Project Cost ................................................................................................................. 69 9 Test plan .............................................................................................................................. 70 9.1 Test Set Task Breakdown ............................................................................................ 72 10 Future Work .................................................................................................................... 75 11 Conclusions ..................................................................................................................... 76 12 Acknowledgments and thanks ........................................................................................ 77 13 Bibliography .................................................................................................................... 78
vii
1 1 Introduction The reuse of Software Requirements, as a process of Software Engineering, gives the chance to Requirement Engineers to maximize the quality of contents and syntax of Software Requirement Specifications during the elicitation, validation and documentation phase. Addressing this quality maximization, Requirement Patterns play an important role upon this quality maximization regarding reusability, diminishing time and error factors whilst defining Software Requirement Specifications (SRS). The GESSI Research Group of the Universitat Politècnica de Catalunya Barcelona Tech, with the help of SSI Department of Public Research Centre Henri Tudor (TUDOR, Luxembourg) have built a first version of a catalog of Software Requirements Patterns (SRP) composed of a set of functional, non-functional and non-technical patterns. These patterns can be used and managed by means of the Pattern-Based Requirements Elicitation system (PABRE), also developed by GESSI. The PABRE system (see figure below) is composed of three tools: PABRE-Man, PABRE-Proj and PABRE-WS. PABRE-Man is the tool within the PABRE system that helps to construct, and make evolve SRP. PABRE-Proj is the tool that supports the phase of elicitation of requirements of a certain project and provides functionalities for the generation of SRS and other documents as Call-for-Tenders. The SRP are stored in the SRP Catalogue which is accessed by both tools. The system also provides PABRE-WS, which are web services that allow accessing to the SRP Catalogue to the Requirement Management Tools (RMT) existent in the marketplace. [1] The PABRE system
2 Without going into detail (for more information see [2]), each SRP is structured in two parts: the SRP general attributes and one or more SRP Forms. Both the SRP and its forms have Name, a Description, an Author, Comments and Sources sections. The SRP has also the Goal section, and the Dependencies section. Each form also consists on several sections the most important of which are one Fixed Part and several Extended Parts. These parts have the same attributes: Question Text, Form Text both having Parameters in the text. For each parameter a metrics is defined. The following example shows the structure of a SRP: Requirement Pattern Authentication Description This pattern expresses the need of having the system functionality to identify users Comments ---- Goal Ensure the identity of the users that access to the system. Author GESSI-CITI Sources (0..*) Requirement books from CITI Specialized literature Keywords (0..*) Access Control, Security, Users Dependencies (0..*) IMPLIES Authorization IMPLIES Stored Data Protection Requirement Form Authentication Description This form states the general need of having the system functionality of identifying users, and has extensions for detailing the type or technology to be used. An extension for requiring to not be necessary to create an specific account for the system is present. Comments Application of extensions: Authentication Technology, Single Sign-on: may be applied at most once each. Version date 2009-03-20 00:00:00.0 Author GESSI-CITI Sources (0..*) Requirement books from CITI Specialized literature Fixed Part Question text ---- Form text The system shall authenticate users Extended Part Authentication Technology Question text ---- Form text The authentication process shall be based on the %authMechanism% authentication technology Parameter Metric authMechanism: is an authentication software technology AuthenticationTechnology: AuthenticationTechnology = Domain(Windows login, …) Extended Part Authentication Types Question text ---- Form text The authentication type shall be: %authTypes% Parameter Metric authTypes: is a non-empty set of authentication types AuthenticationTypes: AuthenticationTypes = Set(AuthenticationType) AuthenticationType = Domain(Open, Encrypted) Extended Part Single Sign-on Question text ---- Form text The system shall not oblige users to create and manage a specific account for this system Authentication SRP Structure
9 1. One level hierarchy: there must be a situation where a hyponym is hypernym of another hyponym or a hypernym is hyponym of another hypernym. Graphically: The solution to this problem relies on the semantics over the words themselves. A more general hypernym must be chosen to be de dominant hypernym, as follows: 2. No Graph: One hyponym can only be hyponym of one, and only one, hypernym. However hypernyms can have more than one hyponym (semantically it should even be a quality measure for a hypernym to have great number of hyponyms).
10 It is also important how these requirements are going to be fulfilled. Every time a hypernym or hyponym is added or modified, these conditions must be checked in order for the requirements to be met. More to this will be added in the design chapter of this document. Integration of the glossary in PABRE-Man As I stated before, the use of the glossary will be during the creation or edition of a SRP. The hyponyms shall be identified and this situation will be showed to the user by underlining them. When a hyponym shall be selected by the user and the user shall press the right button of the mouse, the hypernym of the hyponym will be proposed to the user. The user will choose if the hyponym is changed by the hypernym or not. This functionality will be introduced to the following form fields of the edition of SRP: Name Description Attributes Keywords Fixed Part Text of SRP Forms o Pattern Text o Question Text Extended Part Text of SRP Forms o Pattern Text o Question Text It is important to take into account that this functionality would be easily extensible to other form fields of the PABRE-Man interface. In the case of the Pattern Text and Question Text form fields the analysis of hyponyms has to take into account that in the text there will appear special words named "parameters". These parameters are used during the application of the pattern, taking values specific of the project where the pattern is applied. These parameters are market in the text because they are words that begin and end with a character "%". The analysis of hyponyms does not have to be applied to these parameters. Thus, the parameters may be formally defined by: Upon parameters there shall be no hyponym analysis, due to the special kind of such word. Moreover, within the word containing the parameter-like structure there must be no whitespaces, although hyphens are allowed. Thus, the phrase: “American Mastiff %Mastiff% and Abyssinian are the most common pets in Imaginary Ville” Supposing there is a hyponym entry having Abyssinian and cat, after applying hyponym substitution will change to: “Dog %Mastiff% and Cat are the most common pets in Imaginary Ville”. Notice that substitution mimes the case of the substituted word.
11 4.2 Non-Functional Requirements The non-functional requirements of this project are the ones that guarantee the maximization of several non-functional factors. The non-functional factors of the quality model proposed by the ISO/IEC 25010 can be used. This quality model categorizes software product quality properties into eight characteristics: functionality (requirements already described in the previous section), efficiency, compatibility, usability, reliability, security, maintainability and portability. Each characteristic is composed of a set of related sub characteristics. Here I state for each non-functional characteristic, how the project has to afford it. Efficiency – In this project, the new features, and specifically the identification of hyponyms and their hypernym must be quick just after the writing of the words of the SRP attributes. Compatibility – The added software must co-exist and be interoperable with previous states of software and with the rest of the applications’ components. Usability – The use of the glossary from the form of edition of SRP should be intuitive and easy to use. The idea is to use the typical interaction used in word orthographic correctors. Reliability – The project relies in the implementation of PABRE-Man for this characteristic. Security- The project relies in the implementation of PABRE-Man for this characteristic. Maintainability – In this project, the maintainability will be achieved if the architectural design of the PABRE-Man system that will be extended is respected, and the design patterns used in the system also used in its extension. Portability – The PABRE-Man for the technologies used in its development is portable to any system that supports Java, in this project the extension must use the same technologies or technologies that also respect this portability. For instance, as PABRE- Man is used in different platforms using the JVM, and some features are used in Web Services, the Glossary must also be portable to these systems by the means of the Java technology.
12 5 Specification 5.1 Conceptual Model According to the requirements above and considering the constraints identified in the introduction, in this section the Conceptual Model of the new features that this project develops is presented: External Keys: Glossary (id); Term (name) This conceptual model has also to follow the integrity constraints stated in the functional requirements, which are: do not allow graphs and so, just allow one level of hierarchy. This is provided by the conceptual model because of the definition of the specialization as disjoint, complete, thus the same Term cannot be at the same time Hypernym and Hyponym. It is easy to see how functional requirements are fulfilled with the generalization/specialization relationship being disjoint meaning a term to be suitable to belong to the “Glossary system” is either to be, and can only be, hypernym or hyponym. It is also important to show how all Terms related to the Glossary-which is a Singleton class (only one instance of Glossary class in the system) need to be either hypernyms or hyponyms. Perhaps, as all words written in SRP definition are not part of the Glossary, in a more global Specification context, Term class would be a generalization of, for instance, Word class. However, addressing this issue, at this point the specification of the Conceptual Model must be related to the functionalities and goals in which this project is spanned. As stated in the introductory section, hypernyms must have a description, although null or void values are accepted for it. Hypernyms may also not be related to a hyponym, as clearly shown in the multiplicity of the hyponym role part on “is related to” relationship. In the Term class, the name attribute - of class Stringindicates the content of the term, that is, the value of the written or edited word.
13 5.2 Use Case Model A use case is a list of steps, defining relations between an actor and a system, in order to achieve a goal, or milestone. As mentioned previously, the system’s main User is the Requirements Engineering Expert which triggers almost every action leading to a goal: deciding which terms are hyponyms, which are hypernyms, associating a textual description to hypernym, writing as normal in a customized input which mocks the major part of behaving of a normal input textbox but triggers an auto hyponym checker while the user writes in it (similarly to the spelling checking of a word processor), decide to substitute –or nothyponyms, and selecting a file to load or save the working Glossary. With all, the following model is driven:
14 Dealing with the above use cases there can be a clear segmentation into three different types of use cases. For instance the addition, edition and deletion of hypernyms and hyponyms could be tagged as Glossary management; the type of customized input ,Hyponym Substitution, and auto hyponym checker could be segmented as SRP text editor management; finally the import/export the Glossary could be segmented as External Data Services. Describing the use cases textually in more depth and according to standards defined in [Lar03] and [DSS]: USE CASE : Add Hypernym ACTORS User (starter) PURPOSE: Insert a hypernym, its description and its list of hyponyms into the Glossary. SUMMARY: The User provides all the required data. The system verifies the data and performs the insertion if requirements are fulfilled. TYPE: Glossary management -Primary and essential CROSS REFERENCE - TYPICAL COURSE OF EVENTS ACTORS ACTIONS SYSTEM RESPONSE 1. Use case begins when the user has provided the data and confirmed the action 2. The system receives data, and performs a first analysis. If it finds repeated values in the hyponym list it merges it into a set of hyponyms. 3. With all valid data the system checks for the functional requirements to be met. 3b. Some data violates the requirements and an error message must be shown to inform. 4b. User accepts error message dialog and proceeds to refill data and act as in point 1. 4. Data is Ok. Insertion is made correctly. System shows Glossary with new hypernym, along with its hyponyms and description. No further message is shown. ALTERNATIVE COURSE: Point 1- User cancels action, thus no action continues in the Use Case
15 USE CASE : Edit Hypernym ACTORS User (starter) PURPOSE: Edit a hypernym, its name, its description or its list of hyponyms, into the Glossary. SUMMARY: The User indicates he wants to edit one given hypernym and provides all the required data. The system verifies the data and performs the edition if requirements are fulfilled by the new data. TYPE: Glossary management - Primary and essential CROSS REFERENCE - TYPICAL COURSE OF EVENTS ACTORS ACTIONS SYSTEM RESPONSE 1. Use case begins when the user has activated the edition event pinned to a certain hypernym , provided the new data and confirmed the action 2. The system receives data, and performs a first analysis. If it finds repeated values in the hyponym list it merges it into a set of hyponyms. 3. With all valid data the system checks for the functional requirements to be met. 3b. Some data violates the requirements and an error message must be shown to inform. 4b. User accepts error message dialog and proceeds to refill data and act as in point 1. 4. Data is Ok. Edition is made correctly. System shows Glossary with edited hypernym information. No further message is shown. ALTERNATIVE COURSE: Point 1- User cancels action, thus no action continues in the Use Case USE CASE : Delete Hypernym ACTORS User (starter) PURPOSE: Delete a hypernym and all information associated to it from the Glossary. SUMMARY: The User indicates he wants to delete a given hypernym. TYPE: Glossary management - Primary and essential CROSS REFERENCE This Use Case includes Delete hyponym Use Case TYPICAL COURSE OF EVENTS ACTORS ACTIONS SYSTEM RESPONSE 1. Use case begins when the user has activated the deletion event pinned to a certain hypernym. 2. The system deletes all data from implicit hypernym. No message is shown. ALTERNATIVE COURSE: -
16 USE CASE : Add Hyponym ACTORS User (starter) PURPOSE: Add a hyponym to a certain hypernym, into the Glossary. SUMMARY: The User indicates he wants to add a hyponym of a given hypernym and provides the hyponym’s name. The system verifies the data and performs the insertion if requirements are fulfilled. TYPE: Glossary management - Primary and essential CROSS REFERENCE - TYPICAL COURSE OF EVENTS ACTORS ACTIONS SYSTEM RESPONSE 1. Use case begins when the user has selected a certain hypernym and activated the event which enables to add a hyponym to that hypernym. 2. The system analyzes the name given to the hyponym. If it fulfills requirements it adds the hyponym to the implicit hypernym and no further message is shown. 2b. The Hyponym’s name violates some constraint within the requirements. An error message is shown indicating it. 3b. User accepts error message and returns to Point 1. ALTERNATIVE COURSE: Point 1- User cancels event and the Use Case ends.
17 USE CASE : Edit Hyponym ACTORS User (starter) PURPOSE: Edit a hyponym’s name, into the Glossary. SUMMARY: The User indicates he wants to edit a hyponym of a given hypernym and provides the hyponym’s new name. The system verifies the data and performs the edition if requirements are fulfilled. TYPE: Glossary management - Primary and essential CROSS REFERENCE - TYPICAL COURSE OF EVENTS ACTORS ACTIONS SYSTEM RESPONSE 1. Use case begins when the user has selected a certain hypernym and activated the event which enables to edit a hyponym to that hypernym. 2. The system analyzes the new name given to the hyponym. If it fulfills requirements it modifies the hyponym’s name to the new value of the implicit hypernym and no further message is shown. 2b. The Hyponym’s new name violates some constraint within the requirements. An error message is shown indicating it. 3b. User accepts error message and returns to Point 1. ALTERNATIVE COURSE: Point 1- User cancels event and the Use Case ends. USE CASE : Delete Hyponym ACTORS User (starter) PURPOSE: Delete a hyponym to a certain hypernym, from the Glossary. SUMMARY: The User indicates he wants to delete a given hyponym from a certain hypernym. TYPE: Glossary management - Primary and essential CROSS REFERENCE This Use Case is included by Delete hypernym Use Case TYPICAL COURSE OF EVENTS ACTORS ACTIONS SYSTEM RESPONSE 1. Use case begins when the user has activated the deletion event pinned to a certain hyponym from a given hypernym. 2. The system deletes all data from the hyponym of the implicit hypernym. No message is shown. ALTERNATIVE COURSE: -
18 USE CASE : Type in Customized Input ACTORS User (starter) PURPOSE: Allow to write text in a customized input which enables hyponym highlighting. SUMMARY: The User indicates he wants to delete given hypernym. TYPE: SRP text editor management - Primary and essential CROSS REFERENCE This Use Case includes Auto hyponym Checker Use Case Note: This Use Case extends from all insertions of Forms, Metrics, Patterns and Dialogs which formed part of other past projects and are part of the PABRE-Man system. TYPICAL COURSE OF EVENTS ACTORS ACTIONS SYSTEM RESPONSE 1. Use case begins when the user has written something in the customized input text box situated in the different forms of SRP definitions. 2. The system actually shows each letter written by the user in the text box area the same way a normal text box would do. If the user writes an invalid word it highlights it. No further system message is shown. ALTERNATIVE COURSE: - USE CASE : Auto hyponym Checker ACTORS System <<thread>> stereotype PURPOSE: Analyzes the text in a customized input to highlight invalid words SUMMARY: A Batch light-process analyzes the customizes input text every time the user writes in it to search for invalid words which to highlight without any action being done to this purpose. TYPE: SRP text editor management - Primary and essential CROSS REFERENCE This Use Case is included by Type in Customized Input Use Case TYPICAL COURSE OF EVENTS ACTORS ACTIONS SYSTEM RESPONSE 1. Each time a word or letter is written or content is pasted in the custom text editor this process is fired. It analyzes all the text, and dies. ALTERNATIVE COURSE: -
25 6.1 Guideline to Solution Without considering the above constraints, the first thought to reach the global goals is thinking about the presentation tier. It is easy to see that the presentation tier is, actually, very similar to word processors’ syntax and grammar corrector, underlining invalid words and offering a new window interface suggesting corrections upon a list of high rated candidate words respect to a certain invalid word. Relying upon the interface, and applying the criteria based on the non-functional requirement that states that the interface must be quick and usable, this new window to analyze invalid words must be simplified to a fast an quick User Interface, perhaps a pop-up under the right click event. Red underlining seems a valid way to highlight an invalid word. As a reference guide to arrive to a solution to a similar situation “Design Patterns , Elements of Reusable Object-Oriented Software“ from Erich Gamma, Richard Helm, Ralph Johnson and John Vlissides [GoF] in its section 2.8. In order to access to the data in the customized Glyph case, the Iterator pattern is used and the Abstract Factory pattern is used to render the different look-and-feel from the different User Interfaces. Skipping these aspects due to the fact that they are already fixed by the natural UI look-and-feel associated to PABRE system and the organization of the documents is implicit to Swing Java library. Further on the syntactic analysis is encapsulated within the Visitor pattern. This pattern acts over the String Object that represents the written input channel whenever the user types a letter. Visitor pattern follows the structure below: In the case of this project, Concrete Elements are String Objects – either one char as typed by the user, or a chunk of copied (and unanalyzed ) text or a word processed by an intermediate split with the purpose of whole analysiswhich have an acceptance operation of being a hyponym or not.
26 With respect to the application of the Visitor pattern to this project it must follow the following flow process: As mentioned before, on the first state, a character may also include a blob of text copied in the clipboard. It is obvious that checking if a word is or not a hyponym has to undergo a call to the business tier, properly bypassed by the correct Controllers, as the query has to be made to the actual and implicit dictionary.
27 The domain tier, previously defined in the Specification chapter of this document under the Conceptual Model section, must be designed further on following the requested architecture. Moreover, being the Glossary Domain Controller the entrance point for each function of the Glossary, it meets all the conditions to fulfill the Facade and Proxy Patterns. Facade Pattern Proxy Pattern In the same direction, and within the Three Tier Architecture pattern, the Product Owner has stated its will to persist the Glossary using both the file importation/exportation system and the database management system. This situation suggests applying another Software Pattern to resolve and adapt to this scenario. The Bridge Pattern accomplishes this purpose as is easily seen in the following diagram: Concrete implementations in this project include operations that implement persistence in XML file format as a method of importation and exportation of the data and other operations that implement persistence using a Database Management System.
28 6.2 Design Conceptual Model As a result of the application of above patterns, altogether with the requirements for this project, and refining the Specification Conceptual Model designing its operations, the following Conceptual Model is obtained:
29 6.3 Design alternatives Even though all the different requirements, architecture and pattern design do pin off degrees of freedom in order to design the resultant system, there are certain issues where a decision needs to be made. This is the case of the way data will be persisted into our system. It is crystal clear how PABRE-Man persists its data using Hibernate to access the Database Management System – in the time of the realization of this project it was Apache Derbybut, since the Glossary is semantically independent from the SRP Catalogue and other sub-system in PABRE-Man, should it stick with Hibernate persistence? Could the Glossary be saved by other storage system? It is noticeable to remark the difference between persistence and importation / exportation. When dealing with persistence it is assumed the persisted data is vital for the flow of the application. It could be seen as that the data needed to run for the first time the application will need to be persisted somehow. Importation or exportation could be seen as a snapshot of the data of the implicit object, which in turn could have been filled up by persisted data. The decision of how to persist the Glossary covers up three different options: Java Serialized, XML file, and Hibernate. 6.3.1 Java Serialized Java provides a method, called object serialization, where an object is represented as a sequence of bytes. Object’s data as well as its type and the types of the object’s attributes are stored sequentially (marshalled) in a given file. After a serialized object has been written into a file it can be read as a normal file, and with the file’s stream the object can be obtained again , after a process of un-marshaling , that is, using the type information and bytes that represent the object and its data which can be used to recreate the object in memory. Most surprising is that the process, as a whole, is independent, meaning that an object can be serialized on one platform and deserialized on a different platform. Files which contain Serialized objects are illegible when opened with a word processor either to view or to be written to, and that is a very important disadvantage. In fact, if a file with a serialized object is written to any number of characters when opened by a word processor, the file is corrupted when trying to read it and un-marshal the object from it. For the Glossary to save and load data it is a must to be able for the User to identify entries and be modify or create new ones from an external Application in the persisted artifact.
30 6.3.2 XML-File Similarly to the method mentioned in XSD Section, using the SAX Library a Schema File can be defined so it can be followed to build a persistence XML- file with tags which can be recognized and replicated by the User in order to add, edit or delete entries with another Application, perhaps a text processor. This persistence file may also be read as a normal text file and be precisely parsed into a Glossary object just following again the rules defined in the Schema File. If errors are found, for instance, because of non-expected characters found or the persistence file was not loaded following the given Schema, error handling identifies the error type and a customized error messages can be shown. This really gives the User a good feedback on how and what to fix in order to get the system working. [10] 6.3.3 Hibernate As mentioned in Technology section, Hibernate is a library which helps save Objects into Relational Databases. It is the option that is used broadly in the PABRE system in order to save Business Tier Objects that needs to be persisted. User or developer can access data through a SQL client (such as SQuirreL or Toad) to see data or perform update or insert SQL queries to modify it. This seems a bit tedious, especially for the User, because the process of updating and inserting large quantities of data is pinned up to knowledge of SQL instructions and procedure. Before all, hibernate must be tuned up correctly in order to align with PABRE-Man Hibernate configuration and mapping. The configuration is basically the same, due to the fact that in the integration process it would be wise to use the same Database Management System that is configured already. Mapping, however, has to be modified by adding the following code to Mapeig.hbm.xml file (which is called by the configuration XML file). [HIB 1,2,3] <class name="edu.upc.gessi.rptool.domain.glossary.Glossary" lazy="false" table="GLOSSARY" > <id name="id" column="id" type="integer" > <generator class="native"/> </id> <map name="glosary" table="SYNONYM"> <key column="IDGLOSSARY" not-null="true"/> <index column="HYPONYM" type="text"/> <element column="HYPERNYM" type="text" not-null="true"/> </map> <map name="Description" table="HYPERNYM_DESCRIPTION"> <key column="IDDESCRIPTION" not-null="true"/> <index column="HYPERNYM" type="text" /> <element column="DESCRIPTION" type="text" not-null="true" /> </map> </class>
31 Notice a whole Glossary object is saved. The Glossary object contains a map named glossary which is mapped in the SYNONYM table with key IDGLOSSARY- indicating the id of the Glossary Object in which it is mapped- and an index column named HYPONYM and an element column named HYPERNYM. This follows the idea that a pair <hyponym, hypernym> cannot be repeated. Similarly happens with the Description object mapped in the HYPERNYM_DESCRIPTION table, where DESCRIPTION is stored in a same record with a HYPERNYM which cannot be repeated. This warrants the accomplishment –in a data layer perspective- of the functional requirements initially defined. 6.3.4 Selected Option A decision must be taken in order to design a proper and efficient persistence mechanism. For that reason further testing has been made to analyze how each option behaves in temporal and disk consumption axis. Taking as a reference the code which corresponds to the above Hibernate mapping, having a <String, String > Hashmap representing the glossary and a <String, [String] + > Multimap representing a <hypernym, <Hyponym List> > type structure named hyponyms (used to increase granularity for controller data transfer and thus increment the final efficiency of the system) the results obtained are the following: /w/ 20 2000 4000 5000 10000 50000 100000 Time Hash Map 87 966 2308 3095 5311 20560 41282 Multi-map 9 689 954 1102 2197 10955 21331 XML 145 1768 2015 2304 2491 5647 10387 Total Serialized 96 1655 3262 4197 7508 31515 62613 File size Serialized 1 73 147 184 370 1932 3883 XML 3 362 725 907 1815 9237 18514 W is the cardinality of the glossary, time is measured in milliseconds and File Size in KB. The times of the Hashmap and Multi-map correspond to the times taken by a serialized process of saving and loading the respective object. For instance, for a glossary of 5000 Terms, the time to save the Hashmap associated to the glossary was approximately 3 seconds, and the time to save the Multi-map was a bit more than a second, so the total time to save the object is about 4.2 seconds. The time spent to save the contents of the object into an XML file is 2.3 seconds. The disk space cost of the operation results in a Serialized 184 KB file and an XML file of about 1 Mb.
32 The test is repeated under the same conditions to test the Hibernate option. Same size as before glossaries are sent to save from a flushed table and loaded from previously saved tables (only measuring loading times). Since the essential information is held in the Hashmap, this is the only object to be saved or loaded. Hence, an additional timing measurement has to be made to create the hyponyms Multimap. This time the whole object is set to be saved, including the Hashmap and Multimap objects. The results are shown as follow: /w/ 20 2000 4000 5000 10000 50000 100000 Save 345 2472 4872 4623 8300 36952 81810 Load 69 16 16 16 18 18 19 Create hyponyms 4 6 12 16 17 30 56 Again, time is in milliseconds so a 5000 term glossary is saved in 4.623 seconds and loaded in 16 milliseconds, the same time hyponyms Multimap is created, giving a total object loading throughput time of 32 milliseconds. 6.3.4.1 Conclusion Serialized option is a good option if the only purpose of it where to save the serialized object to a file and then recover it. Its saving and loading time seem reasonable but its great advantage is the optimized disk space size it uses to save files. However its null functionality to modify or view the data since it is somehow codified in serialized object byte stream file makes this option void to fulfill any goal. XML option is great in order to visualize, edit, add, and delete the data in the glossary since it is saved in a normal stream text file with any desired encoding. Due to the fact it has to build the stream, open the file descriptor channel, and write the effective stream (possibly a huge String object) to the file, the file size variable is too big compared to the serialized option as it grows almost exponentially, and the time variable is not suitable for probable glossary sizes, although for bigger sizes it tends to stabilize. Hibernate option is similar in timing as Serialized option but when data is saved it is possible to be viewed without decoding the data, since it is stored in a relational database and can be queried by SQL statements. However the granularity of modifications is limited to be low since a SQL process or script must be needed to insert huge quantity of data at a time. Perhaps the solution lies in a combination of a reasonable quick load and save method and a good visualization and modification of stored data. Therefore the solution is: Hibernate to load and save the whole object XML to export and import –and then actualize saving into the database through Hibernatedata to/from a XML file. Dealing with the initial load of the glossary in the application, using the Hibernate option the timing can be masked by the load of other application data also loaded trough Hibernate. This resolves the following sub-objective: Glossary persistence.
33 6.4 Special design considerations There are many business tier designs to implement a set of strings that represents a morphism application with the guidance of a dictionary. For instance a forest of two-leaved trees representing hypernyms as the root and its hyponyms on the leaves could be a solution. Implicitly Java and other Object Oriented languages tend to translate that solution to a Hashmap of references to objects, for instance called Nodes. The persistence of this solution becomes bloody difficult when persistence is implemented. Instead, and since the elementary Object of this project is the String, the Hashmap and Multimap solution seems proper and accordingly clean and easy to implement and to persist. In fact the Hibernate mapping presented above works just plain easy and only three tables need to be created in SQL creation scripts: GLOSSARY, with only an integer column, HYPERNYM_DESCRIPTION and SYNONYM tables with an integer and two varchar variables. Formally in a tiny SQL table creation script: create table GLOSSARY (id integer); create table SYNONYM (IDGLOSSARY integer, HYPONYM varchar(50),HYPERNYM varchar(50)); create table HYPERNYM_DESCRIPTION (IDDESCRIPTION INT NOT NULL, HYPERNYM varchar(50) NOT NULL, DESCRIPTION varchar(1000)not null, PRIMARY KEY (IDDESCRIPTION, HYPERNYM)); Besides, efficiency has been increased by the usage of Hashmap and Multimap since its insertion and deletion and modification methods are very quick , in fact constant in the major part of the functionalities and at most lineal to the size of the hyponym list(not the whole glossary though). In the other hand, the Node alternative would be less efficient do to the fact it has to create every Node object, assign memory for it, and store it somehow, plus the different update and deletion methods are cost full.
34 6.5 Definitive Design Conceptual Model and Data Persistence Adapter The following is the result of the design conceptual model plus the chosen persistence adapter.
41 7 External Design This chapter plays an important role in this project. Once the Business and Data Tiers are defined, the Presentation Tier needs to be designed in order for the extension in PABRE-Man to be synchronized in aspect and usability with the main application. Conceptually, and in terms of presentation entities, this project introduces these main innovations to the tool [12] [SWN]: • A customized input Text Component where Hyponym detection can be enabled and viewed. The total number of these Components must be situated in the correct way, size, and appearance, in the proper container as it was before the extension of the tool. • A Glossary Internal Frame in which to view the contents of the Glossary in a disposed tree-like structure, attending user actions on nodes as well as a graphic management of the structure. • A custom dialog to add or edit hypernyms and delete a quantity of the hypernym’s hyponyms. • Toolbar icon button to show or hide Glossary Internal Frame. • Main Menu options to save to or import glossary from XML file. To have a visual reference of what the tool looked and felt before extending it, the following screenshot was taken:
42 7.1 Glossary Internal Frame The main internal frame window where data from the Glossary is shown as well as were the user can manage it is the Glossary Internal Frame. Glossary Internal Frame is a realization of the tool’s Generic Internal Frame from which the Metrics Internal Frame is also realized. With the presence of two, or more, internal frames, a basic window manager has to be set up. These internal frames are embedded inside a Main Frame which takes the main part of the application window space and where the Patterns Internal Frame has a fixed and permanent position. In order to render data from the glossary, this frame renders it with a Tree-like structure having a static node called “Glossary” as the root node, hypernyms on the first level nodes and hyponyms on the leave nodes. Focusing on this point, data received by the presentation tier controller has been rearranged to fit in this structure. This is the motivation of having a Multimap in the Business Tier Glossary object, in order to make a granularity leap to receive optimized data in the Presentation Tier, that is, for each hypernym a list of its hyponyms. As a result the following design is obtained. The top part of the frame shows two buttons and an input text field. The buttons “Expand all” and “Collapse all “represent the visual functionality of expanding all hypernym nodes or collapsing them. For instance, in the figure, “pasta” node is expanded and “fruits” node is collapsed. “Collapse all” does not collapse further enough to the Glossary node, although it can be collapsed manually clicking on the -/+ button in the Tree-structure. This Tree-structure is clickable both by the primary and the secondary mouse button. If clicked by the primary
43 button the content of the selected node is shown in the lower bottom part of the frame - which can support Html codingso if a hypernym is clicked the content shown is its description and hyponym list, and if a hyponym is selected its name is shown. 7.1.1 Insert hypernym If the Tree-structure is clicked by the secondary mouse button it depends then on which node the user has clicked. If the user has clicked the Glossary node then only the insert hypernym menu is shown since it is the only option to be done from the Glossary node. Then this pop up menu is shown. When the user confirms to insert a hypernym the Custom Dialog associated to the insertion of a Hypernym shows up. The user inserts the hypernym’s name, and a list of hyponyms separated by commas, and a description for the hypernym. If any of either the hypernym or one or more of the hyponyms does not pass the functional requirement’s check an error message is shown.
44 On success or on cancellation, the Custom Dialog disappears and the user is set back to the Glossary Internal Frame, with the inserted node selected. Notice how nodes are redistributed alphabetically after the insertion, and that the previous opened and/or closed nodes are remained in the same state as before the insertion. 7.1.2 Hypernym menu Again, the user can click the secondary mouse button on the hypernym node and watch the following popup menu with all options enabled: The user also can edit a certain node (one at a time), either hypernym or hyponym, by double clicking it with the primary mouse button and modifying it. The user can cancel this edition action by pressing the Esc key or clicking anyway outside the Tree-like structure or confirm it by pressing the Enter or Intro key. If an edition is confirmed and violates the functional requirement’s test, then an error message is shown and the old value of the node is reset.
45 7.1.3 Edit Hypernym Focusing on editing a hypernym the first option is using the secondary mouse button option. This option leads to the same Custom Editor as in the insertion section, but with the fields already filled up of the data corresponding to the hypernym which is going to be edited. This time the edited hypernym is the first one lexicographically, associating some of the three letter alphabet words with the letter (word) “a”. The user shall introduce the modification and afterwards the modified data shall be, again, tested to pass the functional requirement test. If so, the data will be modified accordingly. In this case, an interesting aspect arises. If the user erases a hyponym in the text list at the same time it modifies the hypernym’s name in a valid way, the hypernym’s hyponyms shall be modified so, erasing the missing values. However this case shall be discussed graphically in the hyponym’s modification section below. This is the same case but having the hypernym edited by double clicking its node with the primary mouse button. The node changes render form and node text becomes editable.
46 7.1.4 Error messages For every insertion or modification there can be a violation of the functional requirement’s test. For instance these are the ones addressed and based upon the above case. In the case of editing the hypernym with the custom dialog this error is shown. Notice the edit custom dialog does not disappear in the presence of the error message. Thus the system enables the user to give another trial or cancel the edition. In this case, the hypernym “fruits” was appointed to change its name to “aa”, which is a hyponym from “a”, so that’s the reason why the message is shown. The same situation, were the hypernym is being edited by double clicking the hypernym node.
47 7.1.5 Insert hyponym After selecting option “Insert hyponym” in hypernym node pop-up menu option, a new hyponym node is created as son of the selected hypernym with the name “hyponym X”, where X is the number of sons –including the just created onethe hypernym has. The new hyponym is selected and editable in order for the user to change its name. If the user clicks on other place or cancels the edition of the new node by pressing Esc the node automatically disappears without being saved in any memory object or persisted object, that is, it only existed on the presentation Tier.
48 If the user confirms the edition of the new node, then the functional requirement’s test has to be checked. Upon success, the hyponym is added in the implicit Business Tier object. Upon error, the consequent error message is shown. 7.1.6 Hyponym menu When clicking with the secondary button of the mouse in a hyponym node the options that the user will see are the following.
49 Insert Hyponym must be disabled so the “one level hierarchy” functional requirement restriction is fulfilled also in the Presentation Tier. The remove hyponym option will simply delete the node in the Tree structure, as well than removing it from the Business Tier object. 7.1.7 Hyponym modification When the option of the hyponym pop menu “Modify Hyponym” is selected the selected hyponym is set editable for the user to modify the hyponym and the same final process as in hyponym insertion starts. If the modification satisfies the functional requirements test then the value is pushed down through the controllers to the Business Tier object. If the modification does not satisfy the requirements an error message is shown the node’s value returns to the old pre-modification value. The same happens if cancelation pushing Esc key or clicking elsewhere in the frame takes place. The node is not deleted in this process. As mentioned in the hypernym modification section some pages above, there can also be a case of hyponym deletion when modifying hypernyms. The example is as follows: An insertion of the word tomato is associated to the hypernym vegetable (a tomato is a fruit). The user wants to correct this error and uses the edition custom dialog to specify the deletion implicitly instead of clicking the tomato hyponym with the secondary mouse button and clicking on the Remove Hyponym option which is also allowed and functional. It is clear how the next step is erasing the selected content and clicking Enter button.
50 It is not possible to show an error message to this method because all the hyponyms listed, including the candidate/s to be erased, were already checked in the insertion process. After the Enter button this is the result of the Glossary Internal Frame. The tomato hyponym has been erased from the Tree-Structure and from the Business Tier object. 7.1.8 Search Functionality In order to search for a certain node within the Glossary Internal Frame Tree Structure, an auto-complete input searcher has been implemented. It is especially useful if the size of the glossary increments to dimensions where searching nodes visually becomes certainly difficult.
57 This is the screenshot storyboard of a “Replace All” substitution. Initial context: Notice the user has clicked over the last pear word, and the next step is to click on the “Replace All” sub-menu so all the pear word occurrences will be replaced by the word “fruits”. Behold also the capitalization of the hyponym words which shall be mocked by the hypernym word. Replace all substitution:
58 Now, a single substitution is to be made. The initial context is now set after the user has replaced al pear occurrences. Notice how the hypernym option is given with the first letter capitalized because the hyponym word was also capitalized because it was the first word in the sentence.
59 This is, then, the result of the single substitution: User can then interact with the resulting data as before the extension of the PABRE-Man tool. 7.5 Importation / Exportation of the Glossary In order to speed up the process of Glossary importation and exportation, a menu option is added to the main menu option bar. A keyboard shortcut is also operative: Ctrl + G to load the Glossary and Ctrl +J to save it.
60 In order to open the XML file containing the Glossary data, the typical Open file dialog with XML file filter is shown for the user to select the desired file and thus for the system to know the file path. The only error associated to importation or exportation of the glossary is the importation of a badly formed XML file. In that case, the following error message is shown and the Glossary is not loaded. It is remarkable how the own error message identifies the parsing error in the XML file and gives the user valuable feedback to fix it.
61 7.6 Design of the Presentation Tier A general portrait of what is happening in the Presentation Tier escapes from the scope of this project. However it helps visualize how the different views, frames and visual functionalities interact and relate themselves in the PABRE-Man tool after the extension for which this project is responsible is installed. Certainly this Tier is really dependent on the technology as for how technology fetches and treats the different events the user makes or the functionalities he/she demands. With all, this is a brief design of the whole Presentation Tier. By finishing this chapter the following initial sub-objectives are fulfilled: • Integration • Usability
62
63 8 Work plan It is fundamental to manage time and resources in nowadays projects. For this reason methodologies have been set to control and monitor the projects processes and evolution in function to an initial kickoff. As a resume, this project starts with a stable version of the PABRE-Man system, and its purpose is to append a software system to it in order to extend it. The glossary is dependent with the system only in the terms the PABRE-Man tool uses. For any other aspect it seemed to be totally independent. However it turned out for integration of the glossary interface to the presentation layer to be a cost full task. First of all a methodology must be selected to be followed through the development of the project. Next, an extensive initial work plan has to be presented with an exhaustive task breakdown. Afterwards, this initial work plan has to be compared and contrasted with the real work plan, due to the fact that it is probable that tasks can have been taken more time than predicted. The real work plan has to be used to calculate the final economic cost of the project. 8.1 Methodology 8.1.1 Introduction There are several methodologies to use in order to develop quality software. For instance, let’s consider the classical life-cycle software model as a reference. In this project’s case Systems engineering is quite quick and simple due to the concreteness of the project and the sharp requirements. Where much labor has to be shown is in the design process, as well as in the coding or integration process. Once the analysis process is done, there are tasks within the project which can be done concurrently, without any order preference. Therefore this classical model is a reference in the general perspective but does not adapt to the actual needs of this project.
64 Another traditional model methodology is the Spiral Prototype model. It basically centers its issues upon general objectives which, as the process is iterated in spiral, are refined to optimize the quality of software. Essentially, both traditional models are proper in the semantics of the methodology because the goals and processes are similar. However they are conceived to lead the building of Software Information Systems –many times from scratchthrough a very tight, formal and procedural way leaning on the advantage it gives to tidy all parts of the software up in a clean manner. Besides, a sequential order is imposed, no other step can be made before the previous one has ended, and that can really be a bottleneck in the throughput. Searching for quick and efficient development Agile Methodologies were created. They are based on iterative and incremental development, where requirements and solutions evolve through collaboration between developers and product owners. Agile development relies on adaptation more than to rigid specification and obeys the following principles: • Individuals and interactions over processes and tools • Working software over comprehensive documentation • Customer collaboration over contract negotiation • Responding to change over following a plan • Regular adaptation to changing circumstances • Face-to-face conversation is the best form of communication • Welcome changing requirements, even late in development
65 8.1.2 Selected Methodology Within Agile Methodologies Scrum is one of the most popular. Scrum is iterative, incremental and agile so it collects bits of Spiral Prototype and the classical life-cycle software model. There are several roles to apply in this methodology: • Product Owner which represents the stakeholders and the customer's expectations • Development team which is responsible of producing and developing the artifacts and application. Usually built up of 3-9 individuals with cohesive roles among them. • Scrum Master is responsible for the proper functioning of the Scrum process. It usually is a person with leadership and point of contact between the other different roles. Scrum’s basic unit of development is the sprint. A sprint is a restricted to a specific duration and effort, typically two weeks. Each sprint starts with a meeting were tasks for the sprint are identified. On the meeting for next sprint or end-of-spring-review meeting progress of the previous sprint is evaluated and reviewed. This following graphic illustrates Scrum’s processes: In the Scrum methodology requirements and use cases are expressed as elements of the product backlog. The product backlog is a list of growing prioritized functional and nonfunctional requirements rated in function of the value the Product Owner has rated them with. However the order and rating can change during the development of the whole project so this list is adaptive to change. [11]
66 The application of Scrum in this project has to be conditioned to reality, there is not a great enterprise or Business behind the final product, or groups of development teams assigned to the project for communication to be extremely vital. However, the essence of Scrum is kept clear in proportion to the scope of the project. The project’s backlog includes the general requirements that are to be accomplished, that is, to extend PABRE-Man tool to enable hyponym highlighting and substitution. The way to accomplish with it could have varied throughout the timespan of the project due to Product Owner’s preferences or technological handicaps. There has been set a weekly or two-week meeting schedule to have fluid communication with the Product Owner and Scrum Master. Risk management is optimized to a maximum project time delay of two weeks in which, if persisted, risk shall be converted to consensual solutions or the scope of the project committedly re-dimensioned. Scrum fits like a glove to the management of these processes in this project. But, in order to have successful sprints, the backlog must be breakdown to simple tasks associated to sprints’ backlogs in order to achieve the final goal. 8.2 Task Breakdown and Sprints This project has been disaggregated in sprints by the following schema: SPRINT NUMBER 1 TASK LENGTH (hrs) TYPE OF TASK Project proposal and meeting 2 admin Project acceptance and registration 1 admin Tribunal Assignation 1 admin Initial requirement gathering and kickoff 2 admin Brainstorm and research 10 development Use case and requirement definition 16 development Backlog elaboration 24 development SPRINT NUMBER 2 TASK LENGTH (hrs) TYPE OF TASK Workspace installation 16 admin PABRE-Man tool installation 8 admin PABRE-Man tool knowledge 24 development Sandbox creation 16 development Business Tier specification 24 development
73 SPRINT NUMBER 7 TASK VISUAL DATA GlossaryInternalFrame creation OK - Event management adding/editing/deleting OK OK AutoHyponymChecker OK - Editor Component Creation OK OK Custom Dialog creation OK OK Due to the fact Business Tier controller is coded and tested this is the first functional testing the glossary extension enhances. Test of independent components is taken here. SPRINT NUMBER 8 TASK VISUAL DATA Tuning up functionalities & controllers OK OK Possibility to evict glossary + hibernate OK OK Tuning up Hibernate error bugs - OK Editor Component Adaptation and extension OK - In this sprint minor functionality bugs in controllers were tested and resolved. In a data perspective the possibility of discharge changes in the implicit glossary object was tested and proven in all three Tiers of the application – in the frame, indicating to evict the glossary in the dialog question, in the Business and data Tiers running eviction methods in Hibernate session and testing it with SQL Squirrel clientwith satisfactory results. Editor component is tested visually searching for the getter and setter functionalities to work as previous components worked. SPRINT NUMBER 9 TASK VISUAL DATA AutoHyponymChecker controller OK - Extending Editor Component Substitution Event OK OK Insert Editor Component to Pattern and Metric forms OK OK Insert Editor Component to dialogs OK OK adapt customized sizes of editor Components OK - AutoHyponymChecker together with Editor Component is tested to evaluate if substitution is achieved in a visual perspective, substituting the invalid words in the Component’s text attribute using AutoHyponymChecker’s tools. Afterwards Editor Components were tested to be able to be added and adapted to the functionalities of the different PABRE-Man Presentation Tier elements.
74 SPRINT NUMBER 10 TASK VISUAL DATA Glossary XML exportation/importation OK OK Adapt Toolbar and Internal frame position management OK - Correct latter bugs OK OK This is the test run to be able to import and export the glossary in XML format. Loading an invalid XML format file is required to check the error message is shown properly. The tools toolbar is tested to open and close the different internal frames with special attention to the GlossaryInternalFrame. A simple management of open internal frames is also implemented and tested so that no horizontal conflicts – overlapping internal framesarise when more than two different internal frames are open. Visual and Data last bugs were also tested to be corrected in both perspectives resulting in complete correctness of the whole glossary extension.
75 10 Future Work The Glossary extension for the PABRE-Man tool can be extended in the future. Due to the Three Tier Architecture Pattern, adding up or replacing code is much easier because functionalities are clearly segmented into artifacts and class objects. For instance, if future work required an additional Editor Component, it just shall be instantiated as an attribute of the class - which behaves as a factory- and make sure the new component supports getters and setters of its own text attribute, as well as listeners to its Java text component, which will be also retrieved by the Editor Component class when required. Other possible future extension of the Glossary subsystem includes a smart Domain Controller which will be consulted by the Glossary Internal Frame controller every time the Internal Frame is opened and also when it is closed. If the Glossary domain object notifies the Internal Frame controller the domain object has not been changed since the last persisting process the Internal Frame controller does not ask the user if he/she wants to save the glossary changes, since there are none. This is really practical for the user when he/she only wants to only consult the glossary seeking perhaps for a certain hypernym or hyponym, avoiding to tediously answering NO to the persistence dialog question. If changes are made to the glossary, obviously the question shall remain so the user can decide whether or not to save the changes to the glossary. Finally other practical expansion of the glossary extension could be towards the Web Services ground. If for any reason the content of the glossary is to be required by any Web Service or framework – such as Google Web Toolkit-, it would be reasonable to pass the data through a JSON type exchange. For instance, defining the JSON schema as the XSD example the process of converting the data from the glossary to this schema is the same as for writing the XML file when the glossary is exported as a XML file. In pseudo code the algorithm shall be similar to the following: Glossary: Map<String, String>; Hypernyms: Multimap <String, String> Description: Map <String, String> String result=json_encode("GLOSSARY"); For each(element from Description){ result:=result.append(json_encode("HYPERNYM")); result=result.append(element.hypernym); result:=result.append(json_encode("DESCRIPTION")); result:=result.append(json_end_close(element.description)); Lhypo:=Hypernyms.get(element.hypernym)) For each(hyponym from Lhypo) { result:=result.append(json_encode("HYPONYM")); result:=result.append(json_end_close(hyponym)); } result:=result.append(json_end_close()); //close hypernym } result:=result.append(json_end_close()); Where uppercase words represent JSON constants and the auxiliary JSON encoding functions depend on the defined schema.
76 11 Conclusions Addressing this project was an inspiring and thrilling challenge for me. It combines different aspects of the subjects learned through the Computer Science Degree from which this project is about. Software Engineering influence is predominant and ironically redundant realizing the project deals with a tool for managing Software Engineering artifacts and processes. Computing Theory is also present in the development of this project, whilst defining the domain of the Glossary model input data. Algorithmic analysis has been used to define and choose optimal data structures which increase software’s efficiency thoroughly and have provided, in second instance, an easier development. Compilers has proven to be essential to the AutoHyponymChecker part, since its analysis is closely similar and linked to the way a compiler analyzes programs, tokenizing the input and parsing it as an Abstract Syntax Tree and processing it using the Visitor Pattern. Design of Web-based Systems has also been an important subject because the Scrum development methodology is basically taught in this subject with a practical focus. Database subject has also been relevant to understand the impact of the chosen design in the Data Tier and how to design this bottom level Tier. Planning and Management of Information Technology Projects and Systems subject has also influenced the course of this project as an emphasis on the changeability of project’s requirements and how to manage that risk. Definitively many have been the subjects that have been aiding tools to manage all situations during this project. Similarly lots of work relating to other subjects has been done but has not reached to be useful to the project. Anyway, the degree’s overall knowledge has been determinant to arrive to a project’s satisfactory conclusion. Additionally there have been lots of technological aspects I did not know about, and others I had heard of but really did not master them. Hibernate, SAX parsing, JDOM, and SQuirreL SQL client where technologies and libraries which I did not know much about before the beginning of this project. Java User Interface programming had been lightly touched during the degree and no other need had forced me to learn about it. It is true that in the degree I have learned about other User Interface frameworks or libraries such as Qt or OpenGL but had no knowledge intersection with Swing Java library. It has been thrilling for me to learn all these new skills and put them in practice as it is certainly inspiring to foresee a future of everlasting skill learning. A special gratitude message is to be sent to the whole open source community which develop helpful and didactic software published its code under the GNU General Public License (GPL) [special thanks to the JOrtho [13] development team], as well as the many problem-solving contributions of specialized internet forums which definitely clarify out doubts and issues. It has been great to contribute and realize this final degree project as a part of a real and currently used tool. The last words are to express my deep desire to guarantee the upmost success of the project hoping the tool to be used broadly.
77 12 Acknowledgments and thanks I would like to use this space to, first of all, thank Carme Quer and Crisina Palomares for giving me the chance to develop this Final Degree Project under their supervision, guidance and support while developing this extension for the PABRE tool which they maintain. It has been a great pleasure to work shoulder to shoulder developing my contribution to the PABRE tool. I want to also mention the deep effort of all the professors-from whom Carme Quer is one of themwhich have spent their time and dedication by sharing their knowledge and enthusiasm with their students. Thanks to that, problem solving becomes easier and future becomes inspiring. I would also like to thank and have a special mention all the different people I have met during the Degree and that have contributed with something positive to my person and skills either sharing either laboratory in some subject, project subjects or even enjoying the amenities between lectures. On other perspective I would also like to thank and acknowledge to my mother and father for giving me the chance to gather from knowledge in every aspect of life, from the essentials to the most detailed topics, and for making it possible to have a wide and extensive opportunity to do so. This project is dedicated to you. Finally, last but not least, quite the opposite, I shall want to give a special mention to a Person which is intended to be forgotten but nevertheless is and will be present in all our lives. May all the good things I do serve to praise Jehovah. (Psalms 83:18). To all of you, thank you for your support without which this project could not have been done.
78 13 Bibliography Note: All on-line resources are checked to be operational at least till today, January 16, 2014 [1] PABRE-Man: Management of a Requirement Patterns Catalogue, Cristina Palomares, Carme Quer, Xavier Franch. Software Engineering for Information Systems research group (GESSI) Universitat Politècnica de Catalunya (UPC) [2]The PABRE tool http://www.upc.edu/gessi/PABRE/ [3] Hypernymy and Hyponym http://en.wikipedia.org/wiki/Hyponymy_and_hypernymy [4] Java programming language http://www.java.com/en/download/whatis_java.jsp [5] Hibernate http://hibernate.org/orm/ [6] Guava Library, Multimap http://code.google.com/p/guava-libraries/wiki/CollectionUtilitiesExplained#Multimaps [7] SAX http://docs.oracle.com/javase/tutorial/jaxp/sax/parsing.html [8] Eclipse project http://www.eclipse.org/org/ [9] SQuirreL SQL client http://squirrel-sql.sourceforge.net/#overview [10] Generating XML files using SAX and XSLT http://www.developer.com//xml/article.php/2108031/Transforming-Flat-Files-To-XML-With-SAX-and- XSLT.htm [11] Scrum Methodology https://www.scrum.org/Resources/What-is-Scrum [12] Add Spellchecker to Java applications http://www.wintertree-software.com/spell-check/java/ [13] JOrtho - a Java spell-checking library http://jortho.sourceforge.net/ [XML] Manual Imprescindible XML Edición 2012 Miguel Ángel Acera García , Ediciones Anaya Multimedia , 2011, ISBN 9788441529601 [HIB1] Harnessing hibernate -James Elliott, Ryan Fowler, and Tim O'Brien, - O’Reilly, 2008, ISBN: 9780596517724 [HIB2] Hibernate recipes: a problem-solution approach- Srinivas Guruzu, Gary Mak,- Apress, 2010 – ISBN: 9781430227960
79 [HIB3] Beginning Hibernate - Jeff Linwood and Dave Minter- Apress, 2010, ISBN: 9781430228509 [GoF] Design Patterns: Elements of Reusable Object-Oriented Software -Erich Gamma, Richard Helm, Ralph Johnson, and John Vlissides - Addison Wesley, 1995, ISBN: 0201633612 [DSS] Diseño de sistemas software en UML- Cristina Gómez, Enric Mayol, Antoni Olivé, Ernest Teniente –Edicions UPC, 2003, ISBN: 8483017245 [Lar03] UML Y Patrones -Craig Larman- Prentice Hall, 2003, ISBN: 8420534382 [SWN] The JFC swing tutorial: a guide to constructing GUIs / Kathy Walrath , Addison- Wesley, 2004, ISBN: 0201914670