Impact of Semantic Similarity on Genre Classification Performance in Movie Descriptions
Erciyes Üniversitesi Fen Bilimleri Enstitüsü Dergisi, vol.41, no.2, pp.384-400, 2025 (TRDizin)
- Publication Type: Article / Article
- Volume: 41 Issue: 2
- Publication Date: 2025
- Journal Name: Erciyes Üniversitesi Fen Bilimleri Enstitüsü Dergisi
- Journal Indexes: TR DİZİN (ULAKBİM)
- Page Numbers: pp.384-400
- Isparta University of Applied Sciences Affiliated: Yes
Abstract
The objective of this study is to develop a system that automatically categorizes English movie descriptions into six distinct genres: Action, Comedy, Romance, Science Fiction, Drama, and Animation. The dataset, characterized by a balanced distribution across these genres, was processed using natural language processing (NLP) techniques. In the initial phase, pre-trained FastText word embedding vectors were employed to filter out words exhibiting low semantic similarity within the descriptions. This step aims to preserve context-specific content relevant to each genre, thereby enhancing the quality of the textual data. Subsequently, the numerical representation of the texts was obtained using the Term Frequency–Inverse Document Frequency (TF-IDF) method, which is known to improve the performance of classification algorithms by reducing the influence of frequently occurring but less informative words. The resulting vectors were classified using Softmax Regression, Multinomial Naive Bayes, Support Vector Machines, Random Forest, and Multilayer Perceptron models, achieving accuracy rates of 89.07%, 89.75%, 89.75%, 89.75%, 85.40%, and 90.02%, respectively. The integration of FastText-based semantic filtering with TF-IDF vectorization proved to be an effective approach for multi-class text classification tasks. Moreover, an empirical comparison of different machine learning algorithms provides valuable insights for model selection. The findings suggest that semantically simplified descriptions contribute positively to model performance. In this regard, the study makes a significant contribution to text preprocessing in genre classification tasks.
Bu çalışmada, İngilizce film açıklamalarının altı farklı türde (aksiyon, komedi, romantik, bilim kurgu, drama, animasyon) otomatik olarak sınıflandırılması hedeflenmiştir. Dengeli dağılıma sahip veri kümesi, doğal dil işleme (NLP) teknikleri ile işlenmiştir. İlk aşamada, FastText’in önceden eğitilmiş kelime gömme vektörleri kullanılarak açıklamalardaki anlamsal olarak düşük benzerlik gösteren kelimeler filtrelenmiş, böylece her tür için bağlama özgü içerik korunarak metinsel verinin kalitesi artırılmıştır. Bu ön işleme adımının ardından, metinlerin sayısal temsili TF-IDF (Term Frequency–Inverse Document Frequency) yöntemiyle elde edilmiştir. TF-IDF, sık tekrar eden ancak ayırt edici gücü düşük kelimelerin etkisini azaltarak sınıflandırma algoritmalarına daha anlamlı öznitelikler sunmaktadır. Elde edilen vektörler, Softmax Regresyon, Multinomial Naive Bayes, Destek Vektör Makineleri, Rastgele Orman ve Çok Katmanlı Algılayıcı modelleri ile sınıflandırılmış; sırasıyla %89,07, %89,75, %89,75, %85,40 ve %90,02 doğruluk oranlarına ulaşılmıştır. FastText tabanlı anlamsal kelime filtreleme ile TF-IDF vektörlerinin birleştirilmesi, çok sınıflı metin sınıflandırma problemleri için etkili bir çözüm sunmuştur. Ayrıca, farklı makine öğrenmesi algoritmalarının karşılaştırılması yoluyla model seçimine yönelik ampirik katkı sağlanmış olup , anlamca sadeleştirilmiş açıklamaların model başarısını olumlu etkilediği gösterilmiştir. Bu yönüyle çalışma, metin ön işleme sürecine yenilikçi bir yaklaşım getirmektedir.