Detecting Deceptive Text Patterns using Machine Learning

dc.contributor.advisorLewis, Michael
dc.contributor.authorKurmangali, Dana
dc.contributor.authorAlmazkyzy, Diana
dc.contributor.authorShyntas, Zhaniya
dc.date.accessioned2026-06-08T07:56:24Z
dc.date.issued2026-04-20
dc.description.abstractIn today’s advanced digital economy, online reviews of products and services can significantly influence both consumer decisions and business reputations. However, because such reviews can be posted by anyone, there has been a sharp increase in misleading and even fraudulent reviews, commonly referred to as opinion spam. These reviews are typically published to artificially boost or lower ratings, thereby misleading consumers, undermining trust in online review platforms, and causing significant financial losses for product or service providers. For this reason, e-commerce companies are particularly interested in developing reliable automated systems to detect deceptive reviews in order to identify and remove them across digital platforms. This research aims to find whether specific linguistic patterns can help identify deceptive reviews using machine learning techniques. We also propose building a prototype of a deception detection system that analyzes both textual and image-based data to identify linguistic and contextual patterns of deception. According to our linguistic analysis, superlatives and stronger adjectives occur significantly more frequently in deceptive representations, supporting the initial hypothesis that deceivers tend to exaggerate in order to enhance credibility and pretend to have real experience. However, other features did not show statistically significant differences in the Opinion Spam Dataset. This finding supports the claim from the literature that linguistic cues of deception vary depending on the context. In terms of model performance, classical ML approaches achieved the strongest overall results, with Hybrid Ensemble reaching 87.9% accuracy and an F1-score of 0.87, while the fine-tuned BERT model achieved 83.4% accuracy and an F1-score of 0.843, but demonstrated higher recall (93.1%) for detecting deceptive reviews. These findings suggest that stylistic exaggeration plays a key role in deceptive opinion writing and that classical ML models remain competitive for deception detection.
dc.identifier.citationKurmangali, D., Almazkyzy, D., & Shyntas, Z. (2026). Detecting deceptive text patterns using machine learning. Nazarbayev University School of Engineering and Digital Sciences
dc.identifier.urihttps://nur.nu.edu.kz/handle/123456789/18879
dc.language.isoen
dc.publisherNazarbayev University School of Engineering and Digital Sciences
dc.rightsAttribution-NonCommercial-NoDerivs 3.0 United Statesen
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/3.0/us/
dc.subjectDeceptive Text Detection
dc.subjectMachine Learning
dc.subjectFake Reviews
dc.subjectOpinion Spam
dc.titleDetecting Deceptive Text Patterns using Machine Learning
dc.typeBachelor's Capstone project

Files

Original bundle

Now showing 1 - 2 of 2
Loading...
Thumbnail Image
Name:
Final_Report.pdf
Size:
686.56 KB
Format:
Adobe Portable Document Format
Description:
Bachelor's Capstone project
Access status: Embargo until 2028-05-20 , Download
Loading...
Thumbnail Image
Name:
Final_Presentation.pdf
Size:
1.14 MB
Format:
Adobe Portable Document Format
Description:
Presentation of the Senior Project
Access status: Embargo until 2028-05-22 , Download