Skip to main content

FedHeartShield Reviewed Romance Scam Corpus

Blauberg, Michael (2026) — Zenodo (CERN European Organization for Nuclear Research)

Citation tools


              
              

Transparency

Evidence and review status

This page contains AI-generated content. No human content review or subject-matter-expert review is recorded.

Source basis
Uploaded Markdown
Source updates
Unresolved
AI-generated page content
Yes
Automated checks
Passed
Administrative approval
Yes
Human content review
Not recorded
Subject-matter-expert review
Not recorded
How this was prepared
Source basis

RSRC privately stores an uploaded Markdown text for internal analysis. Excerpt checks refer to this supplied text, not to a verified copy of the publisher's PDF. The file is not offered to viewers from this page.

  • Document added to RSRC:
Source updates

The latest source-status check did not establish whether an incoming update notice exists. Missing metadata or provider access does not establish source validity.

  • Last source-status attempt:
AI-generated page content

AI-generated research notes displayed on this page: Synopsis, Identified gaps, Methods, Limitations, Future work. The paper itself is not described as AI-generated.

  • Document analysis recorded:
  • Synopsis generation recorded:
  • Page record updated:
Automated checks

The current, source-bound synopsis passed the recorded versioned publication checks.

  • Checks completed:
View passed checks (3)
  • Length, completeness, repetition, refusal, boilerplate, and active-markup screening
  • Numerical claims checked against the available source text
  • English-source lexical grounding check
Administrative approval

An authenticated administrator approved the bibliographic record for public Library display. This is not a review of every research claim.

  • Approved for public display:
Human content review

No human review is recorded for the AI-generated content displayed on this page.

Subject-matter-expert review

RSRC has not recorded review of this content by a subject-matter or methods expert.

Review-state definitions
Found a possible error? Request a correction.

Synopsis

This publication presents a synthetic, human-reviewed corpus of romance-scam and genuine conversations intended for scam-detection and online-safety research. The dataset comprises 404 dialogues, evenly split with 202 scam and 202 genuine interactions, and includes matched benign controls as well as turn-level annotations for risk, stage, and manipulation. The authors describe two tiers of release: an open tier containing a datasheet, data dictionary, validation summaries, and 10 bounded illustrative examples under CC BY-NC 4.0, with the full dialogues available on request. All data are synthetic, designed to avoid authentic victim material or personally identifying information. A central aim appears to be facilitating research on corpus-validity auditing and methods to reduce shortcuts in analysis, rather than measuring detection performance on real or authentic interactions. A noted limitation is that the scam and genuine classes are said to be separable by a writing-style fingerprint created by the generation process rather than by genuine scam-relevant evidence, with a reported held-out AUC range of 0.84 to 1.0 according to accompanying materials. The authors position the corpus as suitable for exploring calibration, validation techniques, and related methodological concerns rather than as a direct benchmark for detecting scam content. The work signals that access to the full dialogue data is restricted to researchers upon request, and the release emphasizes non-commercial use under CC BY-NC 4.0. Overall, the publication presents a synthetic resource intended to support methodological evaluation and validation in online-safety research, while explicitly noting constraints related to separability of classes and the primary purpose of corpus-validation rather than effectiveness testing.

Identified Gaps

The corpus contains no authentic victim communications and is not evidence of real victim behavior. Its two classes were generated through separate generator stacks, producing a detectable register fingerprint. Only 10 truncated, redacted examples are openly available; complete dialogues, prompt bundles, and supporting evidence are gated.

Methods

The dataset contains 404 machine-generated, human-reviewed romance-scam and genuine-control conversations (202 each). Scam cases cover advance-fee, relationship/investment, and aftermath/extension types; controls include counterfactual, matched, adjacent, and standalone hard negatives. Records include per-turn risk, stage, manipulation, and evidence annotations. Validation includes cross-model, judge, shortcut, and human-audit summaries. A single-reviewer risk-based audit was used.

Limitations

This is synthetic data, not authentic victim communication, and should not support claims about real victims or operational fraud decisions. The corpus is unsuitable as a detection-effectiveness benchmark because generator-register artifacts distinguish scam and genuine classes, even before money requests. Per-turn auxiliary labels have lower agreement. The open tier provides only truncated and redacted examples, while full records are gated.

Future Work

The dataset is intended to support development and testing of methods that audit corpus validity and identify or remove shortcut signals. Any extension to real-world romance-scam research would require separate validation against authentic communications.

See how this publication connects to RSRC's living evidence syntheses through current citations and research-topic mapping.