Skip to main content

Automatically Dismantling Online Dating Fraud

Guillermo Suarez-Tangil ; Matthew Edwards ; Claudia Peersman ; Gianluca Stringhini ; Awais Rashid ; Monica Whitty (2019) — IEEE Transactions on Information Forensics and Security

What do these research terms mean?
Preprint
A manuscript shared before formal peer review and publication. Check whether a later published version is available.
Dataset
A collection of data or examples for others to inspect or reuse. It can appear in Library search and topic mapping, but RSRC does not use dataset records as evidence in Research Insights.
Dissertation or thesis
Research submitted for an academic degree. This describes its format, not its reliability.
Journal article
An article published in a journal. This label alone does not establish peer review, study quality, or how well the findings apply elsewhere.
Qualitative research
Examines experiences, meanings, or processes, often through interviews or observations. It can explain how something happens without estimating how common it is.
Quantitative research
Uses numerical measurements to describe patterns or test relationships. A relationship between two measurements does not by itself show that one causes the other.
Systematic review
Uses a planned, documented method to find and assess research addressing a question. Its conclusions still depend on the included studies and what the search covered.
Meta-analysis
Statistically combines results from multiple studies. Combining studies does not remove weaknesses in their design or make unlike populations interchangeable.
Not classified
This record has no recognized label in this filter. It does not mean the publication used no method, or that no research exists.

Definitions draw on DataCite resource types; Cochrane review methods; NLM: association and causation. RSRC’s dataset and classification rules are explained in our methodology.

Citation tools


              
              

Transparency

Evidence and review status

This page contains AI-generated content. No human content review or subject-matter-expert review is recorded.

Source basis
Downloaded PDF
Source updates
No notice found at last check
AI-generated page content
Yes
Automated checks
Passed
Administrative approval
Yes
Human content review
Not recorded
Subject-matter-expert review
Not recorded
How this was prepared
Source basis

RSRC downloaded and privately stored a copy of the paper for internal analysis. The PDF is not offered to viewers from this page.

  • PDF added to RSRC:
Source updates

No incoming update notice was found in the dated Crossref response. Coverage is incomplete, particularly for corrections and expressions of concern; this is not a guarantee that the source is valid or unchanged.

  • Last source-status attempt:
AI-generated page content

AI-generated research notes displayed on this page: Synopsis, Identified gaps, Methods, Limitations, Future work. The paper itself is not described as AI-generated.

  • Document analysis recorded:
  • Page record updated:
Automated checks

The current, source-bound synopsis passed the recorded versioned publication checks.

  • Checks completed:
View passed checks (3)
  • Length, completeness, repetition, refusal, boilerplate, and active-markup screening
  • Numerical claims checked against the available source text
  • English-source lexical grounding check
Administrative approval

The record is approved for public display, but a complete historical administrator action is not recorded.

Human content review

No human review is recorded for the AI-generated content displayed on this page.

Subject-matter-expert review

RSRC has not recorded review of this content by a subject-matter or methods expert.

Review-state definitions
Found a possible error? Request a correction.

Synopsis

The study scraped 5,402 publicly listed scammer profiles from scamdigger.com and randomly sampled 14,720 ordinary profiles from its connected dating site (March 2017). It compared demographics, image-derived deep-learning captions, and self-description text. Three classifiers-demographics (Random Forest/Naive Bayes), image-caption SVM, and description SVM-were combined in a weighted ensemble. Scammer profiles used more images per profile than genuine profiles and selected desirable contexts such as military, academic, and medical settings. Scammers' descriptions used more emotion, family, friendship, certainty, and gender-related words than genuine users' descriptions. The classifier was not tested on profiles from other dating sites, so platform-specific differences among scammers or genuine users may constrain generalizability.

Identified Gaps

The paper identifies an absence of academic literature on practical romance-scammer detection and says no prior work had analyzed whether idealized-romance traits in dating profiles could support detection. It also notes that detection methods used for spam, Sybil accounts, and cloned profiles are poorly suited to human-operated, personalized dating fraud. Cross-platform generalizability and effective user-facing deployment remain insufficiently studied.

Methods

The study scraped 5,402 publicly listed scammer profiles from scamdigger.com and randomly sampled 14,720 ordinary profiles from its connected dating site (March 2017). It compared demographics, image-derived deep-learning captions, and self-description text. Three classifiers—demographics (Random Forest/Naive Bayes), image-caption SVM, and description SVM—were combined in a weighted ensemble. Data were divided into 60% training, 20% test, and 20% validation sets, with duplicate scam-profile variants kept within the same partition.

Limitations

The classifier was not tested on profiles from other dating sites, so platform-specific differences among scammers or genuine users may constrain generalizability. False negatives remained where public profile information was inconclusive. The authors also warn that deployment can wrongly deny service to legitimate users or miss evading scammers; a “safe” classification could create users’ false sense of security. The dataset’s scam labels came from site screening and public scam listings.

Future Work

Test the approach on profiles from other dating sites; obtain broader profile data to assess whether scammer and genuine-user characteristics vary across platforms; augment profile-based detection with geolocated IP and on-platform behavioural/messaging classifiers; examine data actionable for enforcement and countermeasures; and study local warning interventions that protect users without creating dependence or reducing awareness.

See how this publication connects to RSRC's living evidence syntheses through current citations and research-topic mapping.