AIREX: Method and system for data clustering for very large databases

タイトル	Method and system for data clustering for very large databases
本文（外部サイト）	http://hdl.handle.net/2060/20080004550
著者（英）	Livny, Miron; Ramakrishnan, Raghu; Zhang, Tian
著者所属（英）	Wisconsin Alumni Research Foundation
発行日	1998-11-03
言語	eng
内容記述	Multi-dimensional data contained in very large databases is efficiently and accurately clustered to determine patterns therein and extract useful information from such patterns. Conventional computer processors may be used which have limited memory capacity and conventional operating speed, allowing massive data sets to be processed in a reasonable time and with reasonable computer resources. The clustering process is organized using a clustering feature tree structure wherein each clustering feature comprises the number of data points in the cluster, the linear sum of the data points in the cluster, and the square sum of the data points in the cluster. A dense region of data points is treated collectively as a single cluster, and points in sparsely occupied regions can be treated as outliers and removed from the clustering feature tree. The clustering can be carried out continuously with new data points being received and processed, and with the clustering feature tree being restructured as necessary to accommodate the information from the newly received data points.
NASA分類	Computer Operations and Hardware
権利	No Copyright