Skip to main content

Leveraging Naive Bayesian Machine Learning for Detecting Pig Butchering Scams: A Cybersecurity Social Engineering Perspective in Africa

Diallo, Ousmane ; Hassan, Hasnani ; Najeeb Zaidan, Muslim (2025) — The 5th International Scientific Conference on Administrative and Financial Sciences (CIC-ISCAFS'2025)

Synopsis (AI-Generated)

This paper proposes, but does not empirically evaluate, a Naive Bayes system for identifying pig-butchering communications and transactions in African contexts. Suggested features include scam-related language, transaction amount, frequency and location, account age, interaction patterns, and content inconsistencies; the proposed deployment would issue real-time warnings for suspicious messages or profiles. The paper argues that probabilistic classification could supplement public education and platform monitoring, but it reports no training corpus, implementation, validation split, baseline, error rates, or field test. Underreporting and fragmented databases, linguistic and cultural diversity, adaptive coded language, platform switching, encryption, technical cost, and public trust would all affect feasibility and model accuracy.

Identified Gaps (AI-Generated)

The paper identifies insufficient Africa-specific scam data, including underreporting and a lack of centralized databases, as a gap that constrains model training. It also indicates that current keyword-based tools do not detect evolving, coded, culturally tailored, and cross-platform scam language. Models must better accommodate diverse African languages and cultural contexts.

Methods (AI-Generated)

This is a proposed-solution paper rather than a reported empirical evaluation. It proposes a Naive Bayesian classifier, using Bayes’ theorem and historical flagged scam data, to classify suspicious communications and transactions. Proposed features include scam-related words and phrases, transaction amount/frequency/location, account age, interaction patterns, and content inconsistencies. Suggested deployment includes real-time warnings for suspicious messages or profiles.

Limitations (AI-Generated)

The proposed approach faces limited and underreported scam data, with many countries lacking centralized databases. Linguistic and cultural diversity may reduce model accuracy, while scammers’ evolving coded language, platform switching, and encrypted messaging can bypass detection. Public awareness and trust in machine-learning systems may impede adoption, and implementation requires substantial technical investment and skilled cybersecurity personnel.

Future Work (AI-Generated)

Future research should improve Naive Bayesian models’ ability to process real-time data and adapt to Africa’s linguistic and cultural contexts. It should consider privacy-protective laws that still enable fraud detection, develop reliable regional statistics, and combine Naive Bayes with network analysis and natural-language processing to identify complex fraud patterns.

See how this publication connects to RSRC's living evidence syntheses through current citations and research-topic mapping.

AI-Generated Content Notice

The synopsis and research notes on this page were generated with AI from available publication information and, when available, the uploaded paper text. They may contain errors, omissions, or interpretation issues. Readers should follow the DOI or source link, review the original publication, and make their own judgment about the content.

Found a possible error? Request a correction.