Automated persona generation approach for assessing security of large language models and analysis of the security flaws.
| dc.contributor.advisor | Ghani, Anwar | |
| dc.contributor.author | Slyamkhanov, Adilzhan | |
| dc.date.accessioned | 2026-05-26T10:40:33Z | |
| dc.date.issued | 2026-05-04 | |
| dc.description.abstract | With the growing popularity of large language models capable of generating highly granular information based on user input, the question of security has arisen to prevent the use of AI for dangerous and destructive purposes. Red Teaming on LLMs is one of the primary methods for identifying vulnerabilities, and many approaches exist to force language models to generate malicious responses even with built-in protection systems in place. One such effective method is imposing a persona on a model to shift its focus, thereby bypassing security filters. The problem is that automated persona generation methods still require improvement, and we do not fully understand the exact mechanism by which personas increase the likelihood that a model generates dangerous responses. This work attempts to create a more effective persona generation method than Crossover/Mutation, the method on which this work relies - Pattern-Driven Persona Synthesis (PDPS), which creates new personas based on patterns identified among the best instances over several generations. This study also analyzes how personas influence the effectiveness of several persuasion methods created by the Persuasive Adaptive Prompt (PAP) across malicious prompt categories. Personas created using PDPS were more effective at bypassing protection than Crossover/Mutation, increasing Attack Success Rate from 23.5% to 36% under the same experimental conditions. Furthermore, evaluation on the AdvBench dataset shows that PDPS consistently outperforms Crossover/Mutation, achieving an average improvement of approximately 2% on gpt-4o-mini. When combined with PAP, PDPS achieves even more significant results, outperforming Crossover/Mutation by 7% on gpt-4o. The analysis also showed that personas significantly increase the likelihood of bypassing protection when the motivation and nature of the imposed role align with the malicious instruction's goals. Furthermore, the more a persona deviates from normative behavior, the better it performs across all malicious instructions. These findings will help future research better understand the root causes of security failures in language models, analyze the errors of internal mechanisms, and develop more advanced methods for ensuring secure operation. The code of the project is available on GitHub: https://github.com/cobruh123/Thesis-Persona-Generation-and-Analysis/ | |
| dc.identifier.citation | Slyamkhanov, A. (2026). Automated persona generation approach for assessing security of large language models and analysis of the security flaws. Nazarbayev University School of Engineering and Digital Sciences | |
| dc.identifier.uri | https://nur.nu.edu.kz/handle/123456789/18735 | |
| dc.language.iso | en | |
| dc.publisher | Nazarbayev University School of Engineering and Digital Sciences | |
| dc.rights | Attribution 3.0 United States | en |
| dc.rights.uri | http://creativecommons.org/licenses/by/3.0/us/ | |
| dc.title | Automated persona generation approach for assessing security of large language models and analysis of the security flaws. | |
| dc.type | Master`s thesis |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Automated persona generation approach for assessing security of large language models and analysis of the security.pdf
- Size:
- 2.2 MB
- Format:
- Adobe Portable Document Format
- Description:
- Master's Thesis