Detecting Deceptive Text Patterns using Machine Learning
Loading...
Files
Date
Journal Title
Journal ISSN
Volume Title
Publisher
Nazarbayev University School of Engineering and Digital Sciences
Abstract
In today’s advanced digital economy, online reviews of products and services can significantly influence both consumer decisions and business reputations. However, because such reviews can be posted by anyone, there has been a sharp increase in misleading and even fraudulent reviews, commonly referred to as opinion spam. These reviews are typically published to artificially boost or lower ratings, thereby misleading consumers, undermining trust in online review platforms, and causing significant financial losses for product or service providers. For this reason, e-commerce companies are particularly interested in developing reliable automated systems to detect deceptive reviews in order to identify and remove them across digital platforms. This research aims to find whether specific linguistic patterns can help identify deceptive reviews using machine learning techniques. We also propose building a prototype of a deception detection system that analyzes both textual and image-based data to identify linguistic and contextual patterns of deception.
According to our linguistic analysis, superlatives and stronger adjectives occur significantly more frequently in deceptive representations, supporting the initial hypothesis that deceivers tend to exaggerate in order to enhance credibility and pretend to have real experience. However, other features did not show statistically significant differences in the Opinion Spam Dataset. This finding supports the claim from the literature that linguistic cues of deception vary depending on the context. In terms of model performance, classical ML approaches achieved the strongest overall results, with Hybrid Ensemble reaching 87.9% accuracy and an F1-score of 0.87, while the fine-tuned BERT model achieved 83.4% accuracy and an F1-score of 0.843, but demonstrated higher recall (93.1%) for detecting deceptive reviews.
These findings suggest that stylistic exaggeration plays a key role in deceptive opinion writing and that classical ML models remain competitive for deception detection.
Description
Citation
Kurmangali, D., Almazkyzy, D., & Shyntas, Z. (2026). Detecting deceptive text patterns using machine learning. Nazarbayev University School of Engineering and Digital Sciences
Collections
Endorsement
Review
Supplemented By
Referenced By
Creative Commons license
Except where otherwised noted, this item's license is described as Attribution-NonCommercial-NoDerivs 3.0 United States
