- Tytuł:
- XCleaner: A new method for clustering XML documents by structure
- Autorzy:
-
Brzeziński, D.
Leśniewska, A.
Morzy, T.
Piernik, M. - Powiązania:
- https://bibliotekanauki.pl/articles/206159.pdf
- Data publikacji:
- 2011
- Wydawca:
- Polska Akademia Nauk. Instytut Badań Systemowych PAN
- Tematy:
-
XML
clustering
patterns - Opis:
- With the vastly growing data resources on the Internet, XML is one of the most important standards for document management. Not only does it provide enhancements to document exchange and storage, but it is also helpful in a variety of information retrieval tasks. Document clustering is one of the most interesting research areas that utilize semi-structural nature of XML. In this paper, we put forward a new XML clustering algorithm that relies solely on document structure. We propose the use of maximal frequent subtrees and an operator called Satisf/Violate to divide documents into groups. The algorithm is experimentally evaluated on real and synthetic data sets with promising results.
- Źródło:
-
Control and Cybernetics; 2011, 40, 3; 877-891
0324-8569 - Pojawia się w:
- Control and Cybernetics
- Dostawca treści:
- Biblioteka Nauki