Skip to main content

Data from The GDELT Project · Reviewed and interpreted by RSRC

About this collection

Media research methods and limitations

News reports can help us find stories, follow responses and ask better questions. They cannot tell us how many romance scams happen worldwide.

What do we mean by “media”?

Here, a media record represents one source web address (URL), with its saved title, short description and other available metadata. RSRC adds a relevance decision, themes and connections to related coverage. It is a research record about a page, rather than a copy of that page.

GDELT describes its Article List (GAL) as a standardized and minimized basic set of metadata for every article GDELT monitors. Read GDELT’s description and field definitions (6 December 2021). Metadata describes a page; it is not the page’s full contents.

Our collection includes online news coverage, commentary and blog posts, with pages from news organizations, broadcaster websites and blogs. A page may contain text, video or audio. We have not classified every record by media format; a television or radio website does not establish that the item is a broadcast report.

The main analysis uses saved titles and descriptions. We did not watch or listen to every report, collect broadcast recordings, or analyze audio/video transcripts or the full text of every page. Selective deeper reading is documented separately. This collection uses GAL metadata, not all of GDELT’s other news, television or visual datasets.

Why build this?

Official complaints, academic studies and media reports show different parts of the problem. This collection makes one of those sources easier to examine: reporting about romance scams.

The inspiration is Herrera and Hastings’s 2024 paper, The Trajectory of Romance Scams in the U.S.. That study compared web-search interest, news metadata, research publications and government reports. Its news searches used several synonyms and removed duplicate results and keyword false positives.

This project develops the media-research question further. It uses GDELT’s multilingual article metadata, records contextual inclusion decisions, codes reported themes, and distinguishes repeated wording from stronger links between reports. It has a different source, scope and procedure, so its totals should not be spliced into the paper’s trend series.

From source to searchable record

  1. 1Collect
  2. 2Screen
  3. 3Review
  4. 4Code
  5. 5Check
  1. Collect available GDELT files. Files are requested for defined UTC observation windows. Checksums preserve their identity. An unavailable slot means no usable file was retrieved for that timestamp. It does not establish why it was unavailable, whether GDELT produced a file, or whether any news was published.
  2. Screen titles and descriptions. A saved multilingual vocabulary finds candidates. A separate terminology supplement and fixed samples of nonmatching records help expose gaps. A keyword match alone does not establish relevance.
  3. Review the available context. Records are included, excluded or left unresolved. Romance involvement must be supported; investment fraud and money laundering alone are insufficient. Inclusion means relevant coverage, not a verified crime.
  4. Assess themes and relationships. Each included record receives the applicable codebook assessment. Allegations, warnings, attempted actions and completed reported developments stay distinct. Countries have explicit roles. Direct connections retain their reasoning; articles are not merged into cases through chains of links.
  5. Validate and preserve. Source bindings, complete assessments, database totals and repeat imports are checked. Corrections create new versions and preserve earlier decisions. Code and same-system review checks are not independent proof of accuracy.

Bulk processing takes place locally. Codex assists with contextual review, coding and quality checks. This is model-assisted research without an outside review team; there is no independent inter-rater reliability estimate. The website serves curated results and does not generate new analysis when you open a page.

Stories, repetition and reported developments

The story register groups related coverage: possible case families, initiatives and media works. Each membership keeps its source-record revision and a reason. A group is a research hypothesis, not an independently verified case.

The expanded review spans 1 January to 4 October 2026. Local programs screened all 6,484 included records for recurring coverage. We examined 324 saved connections, screened 431 additional title groups, and read the saved source passages for 265 selected groups. Named-subject searches found further reports, including coverage in different languages.

The result contains 334 story hypotheses covering 1,734 source URLs, with 23 direct follow-up connections across the collection. The 83 saved possible connections keep their uncertainty. Similar headlines about different victims, losses or proceedings were separated. General advice and vague similarities were not turned into cases.

How the recurrence screen works

The program compares overlapping five-character sequences in normalized titles. It ignores sequences found in more than 40 records and proposes a pair only when at least 14 sequences match and the overlap reaches 80% of the smaller title's retained sequences. This works across scripts, but does not translate titles. Separate searches use 36 reviewed names or name variants. A second screen of capitalized phrases in the English review notes produced 80 candidates; contextual reading of 26 selected candidates found further names and spelling variants. The January boundary was checked separately. These rules propose material to read; they do not approve a story membership. A saved decision records the contextual assessment and binds it to the source evidence and record revision.

Connected candidates are displayed together to reduce repeated reading. They are not automatically merged through chains. Exact repeated passages are presented once during review while every URL keeps its own record. This limits repeated processing without treating syndication as independent evidence.

This pass reused saved titles, descriptions and evidence passages. It made no new publisher-page downloads or paid API calls. It is a systematic, bounded screen across the collected period, not exhaustive discovery of every recurrence or full-text verification of every article. A name can connect several proceedings or experiences; groups can overlap and must not be added as incident totals.

Timelines sort URLs by first observed metadata date. They do not reconstruct incident dates or assume that a later record is a new development. Arrest, charge, requested sentence, imposed sentence and other stages remain distinct. A follow-up needs a direct explanation of what the reporting adds.

Several URLs can repeat one story. Source hosts are not necessarily independent publishers, and repetition does not measure audience size or public awareness. A record without a story assignment is not automatically a new case. All included records remain searchable with their existing theme assessments.

The first public transfer omitted the original five-week story register and three saved follow-up links. These were restored before the wider review. Transfers now check for missing stories, members and outdated revision bindings. Reuse is allowed only when the saved source and interpretation remain unchanged.

How much data did we screen?

Across the 40 collected windows, the screening programs read 129,882,128 metadata rows, including 129,882,128 valid record objects. These describe all topics in the retrieved files, before relevance screening.

From broad screening to relevant coverage40 observation windows · counts use two different units
Programmatic screening129,882,128Metadata observationsAll topics; repeated URLs included
Contextual review17,202Distinct candidate URLsPrimary and supplemental screening; deduplicated by URL

Candidate review outcomes

6,484Included URLsRelevant in the available evidence
9,368Excluded URLsDid not meet inclusion criteria
1,350Unresolved URLsInsufficient or ambiguous context

These are metadata observations, not unique articles. The same URL may appear more than once. Candidate and review totals count distinct normalized URLs; none of these figures represents verified scams. Box sizes are not proportional to counts.

Read the screening totals as a table
From source metadata to reviewed records
Stage and unitCount
Metadata rows screened (includes repeated observations)129,882,128
Distinct candidate URLs sent for contextual review17,202
Included candidate URLs6,484
Excluded candidate URLs9,368
Unresolved candidate URLs1,350

Each collection window is counted once, even if its processing was rerun. A URL can occur in many metadata rows, so the first figure is not a count of unique articles. We have not measured distinct URLs or duplicates across the entire all-topic corpus. The smaller candidate and outcome counts are deduplicated by normalized URL across the reviewed collection.

Screening volume by observation window
Programmatic screening, before contextual review
WindowFiles readMetadata rows
2026-01-01–2026-01-046941,503,608
2026-01-05–2026-01-111,2493,399,485
2026-01-12–2026-01-181,2133,551,549
2026-01-19–2026-01-251,1953,576,992
2026-01-26–2026-02-011,2103,579,937
2026-02-02–2026-02-081,2263,561,664
2026-02-09–2026-02-151,2153,520,665
2026-02-16–2026-02-221,3133,339,251
2026-02-23–2026-03-011,3793,539,426
2026-03-02–2026-03-081,3733,550,180
2026-03-09–2026-03-151,3703,522,349
2026-03-16–2026-03-221,3403,439,604
2026-03-23–2026-03-291,3713,421,616
2026-03-30–2026-04-051,3173,322,028
2026-04-06–2026-04-121,3383,340,912
2026-04-13–2026-04-191,3513,414,009
2026-04-20–2026-04-261,3553,444,155
2026-04-27–2026-05-031,3493,311,366
2026-05-04–2026-05-101,3403,390,030
2026-05-11–2026-05-171,3353,345,299
2026-05-18–2026-05-241,3323,356,564
2026-05-25–2026-05-311,2843,186,589
2026-06-01–2026-06-071,2643,237,744
2026-06-08–2026-06-141,2673,257,138
2026-06-15–2026-06-211,3113,253,174
2026-06-22–2026-06-281,3593,249,495
2026-06-29–2026-07-051,3353,191,624
2026-07-06–2026-07-121,3253,204,514
2026-07-13–2026-07-191,3113,179,642
2026-07-20–2026-07-261,3333,124,197
2026-07-27–2026-08-021,3253,091,371
2026-08-03–2026-08-091,3273,012,424
2026-08-10–2026-08-161,2972,960,805
2026-08-17–2026-08-231,3283,014,184
2026-08-24–2026-08-301,3363,075,590
2026-08-31–2026-09-061,3373,083,206
2026-09-07–2026-09-131,3543,120,881
2026-09-14–2026-09-201,3653,119,478
2026-09-21–2026-09-271,3163,044,937
2026-09-28–2026-10-041,3333,044,446

Screening volume documents the scale of the work. It does not measure global coverage, search recall or how often scams occur. Most rows were screened programmatically; only selected candidates and audit samples received contextual review.

How we look for missed coverage

A second set of search terms checks for wording the primary search might miss. We also review a fixed sample of records that did not match: normally 56 observations per full week, spread across days and four language groups. These checks can reveal gaps, but they cannot tell us what percentage of all relevant reporting we found.

Three distinctions that matter

An observation is not an incident date

“Observed” means the record appeared in the collected GDELT metadata. An old court story may reappear in a later window. Only explicit dates support a reported-event date; missing years are not guessed.

Several URLs may repeat one story

Shared wording can reveal distribution, but it can also be generic text that a website uses on many pages. It does not demonstrate independent confirmation, additional victims, audience size or increased public awareness. A case connection needs more than similar words.

Not reported is not absent

A short description may omit emotional harm, reporting barriers or a later outcome. An unassigned theme means the available text did not establish it. It is not evidence that it never happened.

Use the collection within its limits

  • Uneven coverage. GDELT, source availability and the vocabulary select what can be found. Languages, regions and smaller sources can be missed.
  • Mostly short source descriptions. Metadata can be incomplete, misleading or internally inconsistent. Selective deeper reading is recorded separately and does not make the entire collection full-text verified.
  • Repeated and promotional material. Syndication and user-generated blogs can dominate a window. The source comparison filter exposes one observed concentration; it does not endorse other sources.
  • Unverified allegations. Source statements remain attributed. Charges are not convictions; requested compensation is not money returned.
  • Limited validation. Automated tests establish software behavior. Same-system contextual checks can repeat the same interpretive error, especially across languages.
  • No population or loss totals. Do not add reports as incidents, sum repeated loss amounts, or use theme percentages as victim prevalence.

When comparing periods, inspect available files, source composition and review scope first. More records can reflect changed collection or repeated coverage rather than a change in offending.

Why checking the original source matters

March reporting announced a change in Japan’s fraud statistics. A later police bulletin explains that the revised categories also apply to figures for earlier months in 2026. SNS investment and romance fraud now sit within the broader “special fraud” total. Comparing that total with an older series therefore requires checking what each series includes.

This is a selective deeper-reading example, separate from the title-and-description assessments. It does not make every media record independently verified. Source: Japan’s National Police Agency, bulletin of 20 April 2026. Summary and interpretation by RSRC; agency information used under its public-data reuse terms.

A referral is not a completed outcome

Australia’s March 2026 taskforce summary reports 377 online assets referred for removal and 1,004 suspected transactions referred for investigation or blocking. These are referrals, not counts of confirmed removals or recovered payments. It also describes a support trial with few referrals; that alone does not establish effectiveness.

RSRC interpretation based on ACCC content, published 6 March 2026, under CC BY 4.0. No agency endorsement is implied. This selective source check is separate from the metadata assessments.

What is available in this release?

2026 year-to-date coverage with expanded story review. Method RSRC-QCA-2026-10-08.2; codebook 1.1.0.

View availability for all 40 observation windows
Collection availability by observation window. Unavailable slots do not mean no articles.
WindowAvailable filesUnavailable slotsIncluded URLs in window
2026-01-01–2026-01-046945,06634
2026-01-05–2026-01-111,2498,831125
2026-01-12–2026-01-181,2138,867103
2026-01-19–2026-01-251,1958,885143
2026-01-26–2026-02-011,2108,870142
2026-02-02–2026-02-081,2268,854193
2026-02-09–2026-02-151,2158,865671
2026-02-16–2026-02-221,3138,767130
2026-02-23–2026-03-011,3798,701117
2026-03-02–2026-03-081,3738,70786
2026-03-09–2026-03-151,3708,71098
2026-03-16–2026-03-221,3408,74084
2026-03-23–2026-03-291,3718,70962
2026-03-30–2026-04-051,3178,76364
2026-04-06–2026-04-121,3388,742124
2026-04-13–2026-04-191,3518,729109
2026-04-20–2026-04-261,3558,725137
2026-04-27–2026-05-031,3498,731132
2026-05-04–2026-05-101,3408,740147
2026-05-11–2026-05-171,3358,745221
2026-05-18–2026-05-241,3328,748224
2026-05-25–2026-05-311,2848,796173
2026-06-01–2026-06-071,2648,816205
2026-06-08–2026-06-141,2678,813126
2026-06-15–2026-06-211,3118,769203
2026-06-22–2026-06-281,3598,721251
2026-06-29–2026-07-051,3358,745189
2026-07-06–2026-07-121,3258,755226
2026-07-13–2026-07-191,3118,769293
2026-07-20–2026-07-261,3338,747215
2026-07-27–2026-08-021,3258,755372
2026-08-03–2026-08-091,3278,753212
2026-08-10–2026-08-161,2978,783139
2026-08-17–2026-08-231,3288,75297
2026-08-24–2026-08-301,3368,744118
2026-08-31–2026-09-061,3378,743114
2026-09-07–2026-09-131,3548,726184
2026-09-14–2026-09-201,3658,715212
2026-09-21–2026-09-271,3168,764114
2026-09-28–2026-10-041,3338,74766

A URL observed in more than one window appears in each relevant row. Do not sum this column as distinct stories or cases.

What the themes mean

Each included record is assessed against these 24 themes. A theme is assigned only when the saved evidence supports it. Several themes can apply to the same record.

Read the theme definitions

Financial and property harm

Loss or attempted extraction of money or property linked to the scam.

Boundary: Do not turn a requested amount, potential loss, compensation order, offender earnings or mixed-fraud total into actual personal loss.

Psychological distress

Reported emotional or mental-health harm associated with the scam.

Boundary: No diagnosis from behavior; pre-existing loneliness, a psychiatric workplace or routine prevention advice is not post-scam mental-health harm.

Family and social harm

Reported damage to relationships, family life, belonging or reputation.

Boundary: A relative mentioned in a report is not sufficient; pre-existing isolation belongs to context.

Disruption of daily life

Non-monetary consequences for housing, employment, care or essential daily needs.

Boundary: A large monetary loss alone does not establish homelessness, unemployment or inability to meet needs.

Physical or sexual harm

Reported bodily injury, physical abuse or sexual violence connected with the scam or its operation.

Boundary: Do not infer violence from a raid, detention or consensual sexual contact; homicide has its own code.

Suicide and self-harm

Reported suicidal thoughts, attempted suicide, death by suicide or intentional self-injury.

Boundary: Metaphorical language, unexplained death, general distress or a headline with insufficient context is not an occurrence. Never infer that the scam caused a death.

Homicide and lethal violence

Reported killing of another person associated with the scam narrative.

Boundary: Unrelated crime in a news roundup; metaphorical homicide; death by suicide or unspecified death.

Moving money for a scheme

Use of a person or their account to collect, transfer, receive or conceal funds for a scam.

Boundary: Paying one’s own loss is not money-mule activity. A money courier is not a drug courier. Do not assume coercion or innocence.

Involvement in other offending

Reported participation in a different harmful or unlawful act through or alongside the scam.

Boundary: Occupation alone is not employer theft. Business-email compromise in a list of crimes is not proof that a romance-scam victim stole from work.

Coercion and exploitation

Threats, force, deception into forced work, confinement or exploitation within the scam operation.

Boundary: Low performance, working for a scam group or being called a mule does not establish trafficking or forced labor.

Blackmail and extortion

Threat-based demands related to the scam, including threatened release of sexual material.

Boundary: Ordinary deceptive money requests without threat; do not infer sextortion from sexual content alone.

Continued or renewed targeting

Further exploitation attempts after earlier payments, suspicion, intervention or loss.

Boundary: Repeated news coverage or an offender’s previous conviction does not by itself establish renewed targeting of this victim.

Investigation and operational response

Reported investigative or operational action by authorities.

Boundary: Do not infer a raid from “bust” or an investigation from a police attribution alone. An arrested person is not automatically a perpetrator.

Justice proceedings and outcomes

Reported court, prosecution or related legal development.

Boundary: Charge is not conviction, guilty plea is not automatically sentencing, requested punishment is not an imposed sentence.

Protection or recovery of assets

An action concerning threatened, stolen, restrained or returned assets.

Boundary: An order to pay or recovered property does not demonstrate repayment to the victim; never subtract it from loss automatically.

Help, reporting and prevention

Victim help-seeking, support provision, prevention or rescue described in the text.

Boundary: Service availability is not evidence of service use or effectiveness; a warning is not an experienced harm.

Institutional barriers or harmful response

Reported obstacles to help or harmful treatment by organizations or authorities.

Boundary: A police action alone is not misconduct; source allegations remain attributed and possible denials stay visible.

Relationship-based deception

Use of a fabricated or manipulated romantic relationship, persona or trust to facilitate the scam.

Boundary: Merely using a dating platform or meeting someone online is insufficient; genuine relationship conflict is not automatically fraud.

Investment pretext

A relationship is used to introduce investment, trading or profit claims.

Boundary: Crypto payment alone is not investment fraud; adjacent investment and romance topics may be unrelated.

Technical facilitation

Specific digital methods support deception, communication or movement of funds.

Boundary: Do not infer AI from a sophisticated scam or treat platform presence as platform culpability.

Organization and division of labor

Reported coordinated roles, groups or cross-border operation.

Boundary: Many victims or several countries in a story do not alone establish an organized criminal network.

Blame, stigma and responsibility framing

Language assigning fault, ridicule or moral judgment to a harmed person or contested victim-participant.

Boundary: A factual account of payment, a legal finding or neutral prevention advice alone is not victim blaming. Self-blame is distinct from external blame.

Prior circumstances and claimed vulnerability

A source presents circumstances preceding exploitation as relevant context.

Boundary: Do not label everyone of a given age vulnerable, or convert prior distress into a consequence of the scam.

Policy and regulatory response

A stated proposal, institutional policy change or legislative action aimed at scam prevention or response.

Boundary: A pledge is not implementation. Senate passage alone is not enacted law. A commercial tool launch is a service, not regulation. Unrelated policy in a roundup is excluded.

Data source and reuse

Data from The GDELT Project, using its GDELT Article List (GAL). Collection, selection, coding and interpretation are by RSRC; GDELT does not endorse these assessments.

GDELT’s published terms of use permit use and redistribution of its datasets without a fee, with attribution and a link to GDELT. We checked these terms on 9 October 2026. The GAL documentation describes using its metadata to help readers find and link to coverage.

This permission supports sharing our metadata-based research and aggregate findings. It does not establish permission to republish publishers’ full articles, photographs or other separately protected material. This site links to sources and presents RSRC’s assessments; it does not distribute the collected article bodies.

The graphs count distinct included URLs within each observation month. The chart builder applies the same filters as the record search. Its percentage view divides matching URLs by all included URLs observed within the same month and selected dates, before theme, text, language, country or coverage-type filters. Line and bar views show the same values. Clicking a point preserves the filters and narrows the dates. A record with several themes counts once in each applicable theme. Months without any included records have no percentage estimate. These measures describe the collection, not worldwide offending or victim prevalence.

How the collection can grow

Earlier years and later weeks will be added to the existing record. Stable identifiers and retained revisions let researchers distinguish new coverage from a correction. A new year will not replace an old one.

Possible additions include saved searches, clearer timelines of reported developments, language-specific retrieval improvements, and comparisons across years with coverage controls.

A downloadable research dataset

A future download could contain stable record identifiers, source URLs, reviewed themes and a data dictionary. It would need a versioned citation, correction policy, privacy review and clear reuse terms. GDELT attribution would accompany the download. Separately protected third-party material would require its own permission; article bodies and images would not be included.

This would help other researchers inspect decisions and build on the work. A downloadable dataset is a future option; this release provides a searchable research collection.

Sources and further reading