Download PDF

FBE Research Report KBI_1020

Publication date: 2010-09-01
Volume: 757 Pages: 53 - 62
Publisher: K.U.Leuven - Faculty of Business and Economics; Leuven (Belgium)

Author:

Poelmans, Jonas
Elzinga, Paul ; Viaene, Stijn ; Dedene, Guido

Keywords:

4609 Information systems

Abstract:

Concept Relation Discovery and Innovation Enabling Technology (CORDIET), is a toolbox for gaining new knowledge from unstructured text data. At the core of CORDIET is the C-K theory which captures the essential elements of innovation. The tool uses Formal Concept Analysis (FCA), Emergent Self Organizing Maps (ESOM) and Hidden Markov Models (HMM) as main artifacts in the analysis process. The user can define temporal, text mining and compound attributes. The text mining attributes are used to analyze the unstructured text in documents, the temporal attributes use these document’s timestamps for analysis. The compound attributes are XML rules based on text mining and temporal attributes. The user can cluster objects with object-cluster rules and can chop the data in pieces with segmentation rules. The artifacts are optimal zed for efficient data analysis, object labels in the FCA lattice and ESOM map contain an URL on which the user can click to open the selected document.