| タイトル | Method and system for data clustering for very large databases |
| 本文(外部サイト) | http://hdl.handle.net/2060/20080004550 |
| 著者(英) | Livny, Miron; Ramakrishnan, Raghu; Zhang, Tian |
| 著者所属(英) | Wisconsin Alumni Research Foundation |
| 発行日 | 1998-11-03 |
| 言語 | eng |
| 内容記述 | Multi-dimensional data contained in very large databases is efficiently and accurately clustered to determine patterns therein and extract useful information from such patterns. Conventional computer processors may be used which have limited memory capacity and conventional operating speed, allowing massive data sets to be processed in a reasonable time and with reasonable computer resources. The clustering process is organized using a clustering feature tree structure wherein each clustering feature comprises the number of data points in the cluster, the linear sum of the data points in the cluster, and the square sum of the data points in the cluster. A dense region of data points is treated collectively as a single cluster, and points in sparsely occupied regions can be treated as outliers and removed from the clustering feature tree. The clustering can be carried out continuously with new data points being received and processed, and with the clustering feature tree being restructured as necessary to accommodate the information from the newly received data points. |
| NASA分類 | Computer Operations and Hardware |
| 権利 | No Copyright |
|