Skip to main content

Evaluating Jailbreak Vulnerabilities in LLMs: A Taxonomy and Comparative Analysis in Romance Fraud Scenarios

Yohn Jairo Parra ; Hongmei Chi ; Richard A. Alo ; Vinicius Lima (2026) — 2026 IEEE 5th International Conference on AI in Cybersecurity (ICAIC)

Citation activity

Total citations · Google Scholar
2
Average per calendar year
Too new for an annual average

Count checked .

Counts reflect Google Scholar’s coverage and do not measure research quality. Source, calculation and limitations

What do these research terms mean?
Preprint
A manuscript shared before formal peer review and publication. Check whether a later published version is available.
Dataset
A collection of data or examples for others to inspect or reuse. It can appear in Library search and topic mapping, but RSRC does not use dataset records as evidence in Research Insights.
Dissertation or thesis
Research submitted for an academic degree. This describes its format, not its reliability.
Journal article
An article published in a journal. This label alone does not establish peer review, study quality, or how well the findings apply elsewhere.
Qualitative research
Examines experiences, meanings, or processes, often through interviews or observations. It can explain how something happens without estimating how common it is.
Quantitative research
Uses numerical measurements to describe patterns or test relationships. A relationship between two measurements does not by itself show that one causes the other.
Systematic review
Uses a planned, documented method to find and assess research addressing a question. Its conclusions still depend on the included studies and what the search covered.
Meta-analysis
Statistically combines results from multiple studies. Combining studies does not remove weaknesses in their design or make unlike populations interchangeable.
Not classified
This record has no recognized label in this filter. It does not mean the publication used no method, or that no research exists.

Definitions draw on DataCite resource types; Cochrane review methods; NLM: association and causation. RSRC’s dataset and classification rules are explained in our methodology.

Citation tools


              
              

Transparency

Evidence and review status

This page contains AI-generated content. No human content review or subject-matter-expert review is recorded.

Source basis
Downloaded PDF
Source updates
No notice found at last check
AI-generated page content
Yes
Automated checks
Passed
Administrative approval
Yes
Human content review
Not recorded
Subject-matter-expert review
Not recorded
How this was prepared
Source basis

RSRC downloaded and privately stored a copy of the paper for internal analysis. The PDF is not offered to viewers from this page.

  • PDF added to RSRC:
Source updates

No incoming update notice was found in the dated Crossref response. Coverage is incomplete, particularly for corrections and expressions of concern; this is not a guarantee that the source is valid or unchanged.

  • Last source-status attempt:
AI-generated page content

AI-generated research notes displayed on this page: Synopsis, Identified gaps, Methods, Limitations, Future work. The paper itself is not described as AI-generated.

  • Document analysis recorded:
  • Page record updated:
Automated checks

The current, source-bound synopsis passed the recorded versioned publication checks.

  • Checks completed:
View passed checks (3)
  • Length, completeness, repetition, refusal, boilerplate, and active-markup screening
  • Numerical claims checked against the available source text
  • English-source lexical grounding check
Administrative approval

An authenticated administrator approved the bibliographic record for public Library display. This is not a review of every research claim.

  • Approved for public display:
Human content review

No human review is recorded for the AI-generated content displayed on this page.

Subject-matter-expert review

RSRC has not recorded review of this content by a subject-matter or methods expert.

Review-state definitions
Found a possible error? Request a correction.

Synopsis

This benchmark evaluates romance-fraud-themed jailbreak prompts against GPT-4/ChatGPT, Gemini 1.5 Flash, and Claude Sonnet using a common API harness. It groups prompts into strategies such as role-play, agent mimicry, policy evasion, system override, format smuggling, and obfuscation, then measures attack success, refusal, and latency. On the 60-prompt analysis set, the paper reports attack-success rates of 0.983 for Gemini, 0.533 for the OpenAI models, and 0.250 for Claude; combinations of role-play and refusal suppression were especially productive. Results depend on fixed model versions, single-turn prompts, and a refusal-based success proxy, so they are a reproducible safety stress test rather than evidence about real victim interactions.

Identified Gaps

Gaps identified include: (1) blind spots in the taxonomy where some tactics are underrepresented; (2) cross-model averages masking model-specific vulnerabilities; (3) limited analysis of obfuscation/translation-based attacks; (4) no real-world deployment or user-study; (5) need for adaptive defenses that evolve with jailbreak strategies; (6) insufficient exploration of contextual factors driving susceptibility; (7) dual-use risk mitigation could be strengthened.

Methods

The study uses a taxonomy-driven benchmark to assess jailbreaking vulnerabilities of LLMs in romance-fraud contexts. A modular testing framework runs a prompt–response matrix across three LLM families (OpenAI GPT-4/ChatGPT, Gemini 1.5 Flash, Claude Sonnet). The benchmark includes 80 romance-fraud prompts across eight scenarios (60 prompts used for ASR analysis). Prompts are sorted into seven strategy primitives (e.g., role-play, agent mimicry, policy translation, sudo mode, jailbreaking, format smuggling, obfuscation). Metrics include Attack Success Rate (ASR), refusal rate, and latency, with pairwise ΔASR and confusion analyses to identify model weaknesses. All results are logged for reproducibility.

Limitations

Limitations: The evaluation relies on three model versions with fixed safety settings, limiting generalizability to other models or updates. The 80-prompt dataset may not cover all romance-fraud tactics, languages, or real-user dynamics. ASR is proxied via regex-based refusals, which may not perfectly map safety outcomes. Single-turn prompts and a controlled harness may not reflect interactive deception. Latency depends on infrastructure; ethical considerations constrain detailed content. No human participants limit user-vulnerability insights.

Future Work

Future work should examine how user behavior and platform policies interact with romance-fraud prompts, and develop targeted interventions at both individual and systemic levels to deter abuse and improve safety.

See how this publication connects to RSRC's living evidence syntheses through current citations and research-topic mapping.