Skip to main content

Automatically Dismantling Online Dating Fraud

Suarez-Tangil, Guillermo ; Edwards, Matthew ; Peersman, Claudia ; Stringhini, Gianluca ; Rashid, Awais ; Whitty, Monica (2019) — IEEE Transactions on Information Forensics and Security

Synopsis (AI-Generated)

The study scraped 5,402 publicly listed scammer profiles from scamdigger.com and randomly sampled 14,720 ordinary profiles from its connected dating site (March 2017). It compared demographics, image-derived deep-learning captions, and self-description text. Three classifiers-demographics (Random Forest/Naive Bayes), image-caption SVM, and description SVM-were combined in a weighted ensemble. Scammer profiles used more images per profile than genuine profiles and selected desirable contexts such as military, academic, and medical settings. Scammers' descriptions used more emotion, family, friendship, certainty, and gender-related words than genuine users' descriptions. The classifier was not tested on profiles from other dating sites, so platform-specific differences among scammers or genuine users may constrain generalizability.

Identified Gaps (AI-Generated)

The paper identifies an absence of academic literature on practical romance-scammer detection and says no prior work had analyzed whether idealized-romance traits in dating profiles could support detection. It also notes that detection methods used for spam, Sybil accounts, and cloned profiles are poorly suited to human-operated, personalized dating fraud. Cross-platform generalizability and effective user-facing deployment remain insufficiently studied.

Methods (AI-Generated)

The study scraped 5,402 publicly listed scammer profiles from scamdigger.com and randomly sampled 14,720 ordinary profiles from its connected dating site (March 2017). It compared demographics, image-derived deep-learning captions, and self-description text. Three classifiers—demographics (Random Forest/Naive Bayes), image-caption SVM, and description SVM—were combined in a weighted ensemble. Data were divided into 60% training, 20% test, and 20% validation sets, with duplicate scam-profile variants kept within the same partition.

Limitations (AI-Generated)

The classifier was not tested on profiles from other dating sites, so platform-specific differences among scammers or genuine users may constrain generalizability. False negatives remained where public profile information was inconclusive. The authors also warn that deployment can wrongly deny service to legitimate users or miss evading scammers; a “safe” classification could create users’ false sense of security. The dataset’s scam labels came from site screening and public scam listings.

Future Work (AI-Generated)

Test the approach on profiles from other dating sites; obtain broader profile data to assess whether scammer and genuine-user characteristics vary across platforms; augment profile-based detection with geolocated IP and on-platform behavioural/messaging classifiers; examine data actionable for enforcement and countermeasures; and study local warning interventions that protect users without creating dependence or reducing awareness.

See how this publication connects to RSRC's living evidence syntheses through current citations and research-topic mapping.

AI-Generated Content Notice

The synopsis and research notes on this page were generated with AI from available publication information and, when available, the uploaded paper text. They may contain errors, omissions, or interpretation issues. Readers should follow the DOI or source link, review the original publication, and make their own judgment about the content.

Found a possible error? Request a correction.