Comment text clustering algorithm based on improved DEC
DOI:
https://doi.org/10.59782/sidr.v1i1.49Keywords:
BERT model, LDA model, deep embedding clustering, autoencoder, clusteringAbstract
Aiming at the problem that the initial number of clusters and cluster centers obtained by the clustering layer in the original deep embedding clustering (DEC) algorithm are highly random, thus affecting the effect of the DEC algorithm, a comment text clustering algorithm based on improved DEC is proposed to perform unsupervised clustering on e-commerce comment data without category annotations. Firstly, the vectorized representation of the BERT-LDA dataset that integrates sentence embedding vectors and topic distribution vectors is obtained; then the DEC algorithm is improved, and the dimension reduction is performed through an autoencoder. A clustering layer is stacked after the encoder, in which the number of clusters in the clustering layer is selected based on topic coherence, and the topic feature vector is used as a custom clustering center. The encoder and clustering layer are then jointly trained to improve the accuracy of clustering; finally, the clustering effect is intuitively displayed using a visualization tool. To verify the effectiveness of the algorithm, the algorithm is compared with 6 comparison algorithms for unsupervised clustering training on an unlabeled product review dataset. The results show that the algorithm achieves the best results of 0.2135 and 2958.18 in the silhouette coefficient and Calinski-Harabaz index, respectively. This shows that it can effectively process e-commerce review data and reflect users' attention to products.
Downloads
How to Cite
Issue
Section
License
Copyright (c) 2024 Scientific Insights and Discoveries Review

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.