Published June 29, 2005
| Version v1
Conference paper
Experiments in Clustering Homogeneous XML Documents to Validate an Existing Typology
Contributors
Others:
- Usage-centered design, analysis and improvement of information systems (AxIS) ; Centre Inria d'Université Côte d'Azur (CRISAM) ; Institut National de Recherche en Informatique et en Automatique (Inria)-Institut National de Recherche en Informatique et en Automatique (Inria)-Inria Paris-Rocquencourt ; Institut National de Recherche en Informatique et en Automatique (Inria)
- Hermann Mauer
Description
This paper presents some experiments in clustering homogeneous XMLdocuments to validate an existing classification or more generally anorganisational structure. Our approach integrates techniques for extracting knowledge from documents with unsupervised classification (clustering) of documents. We focus on the feature selection used for representing documents and its impact on the emerging classification. We mix the selection of structured features with fine textual selection based on syntactic characteristics.We illustrate and evaluate this approach with a collection of Inria activity reports for the year 2003. The objective is to cluster projects into larger groups (Themes), based on the keywords or different chapters of these activity reports. We then compare the results of clustering using different feature selections, with the official theme structure used by Inria.
Abstract
(postprint); This version corrects a couple of errors in authors' names in the bibliography./www.jucs.orgAdditional details
Identifiers
- URL
- https://inria.hal.science/inria-00000002
- URN
- urn:oai:HAL:inria-00000002v3
Origin repository
- Origin repository
- UNICA