Skip to main content

Human‐supervised LLM triage of pig butchering complaints: A validation study of multi‐model identifier extraction

Sanghyeob Ko ; Jumi Kim ; Jin Young Jang ; Gibum Kim (2026) — Journal of Forensic Sciences

Citation tools


              
              

Transparency

Evidence and review status

This page contains AI-generated content. No human content review or subject-matter-expert review is recorded.

Source basis
Downloaded PDF
AI-generated page content
Yes
Automated checks
Passed
Administrative approval
Yes
Human content review
Not recorded
Subject-matter-expert review
Not recorded
How this was prepared
Source basis

RSRC downloaded and privately stored a copy of the paper for internal analysis. The PDF is not offered to viewers from this page.

  • PDF added to RSRC:
AI-generated page content

AI-generated research notes displayed on this page: Synopsis, Identified gaps, Methods, Limitations, Future work. The paper itself is not described as AI-generated.

  • Document analysis recorded:
  • Synopsis generation recorded:
  • Page record updated:
Automated checks

The current, source-bound synopsis passed the recorded versioned publication checks.

  • Checks completed:
View passed checks (3)
  • Length, completeness, repetition, refusal, boilerplate, and active-markup screening
  • Numerical claims checked against the available source text
  • English-source lexical grounding check
Administrative approval

An authenticated administrator approved the bibliographic record for public Library display. This is not a review of every research claim.

  • Approved for public display:
Human content review

No human review is recorded for the AI-generated content displayed on this page.

Subject-matter-expert review

RSRC has not recorded review of this content by a subject-matter or methods expert.

Review-state definitions
Found a possible error? Request a correction.

Synopsis

The study aims to validate a human‑supervised triage approach for pig butchering complaints by testing whether multi‑model identifier extraction can recover triage‑relevant forensic identifiers from public complaints. It introduces a four‑tier, information‑based forensic taxonomy that classifies cases by the presence of traceable identifiers—cryptocurrency wallet addresses, fraud‑related URLs, platform names, and reported loss amounts—rather than by narrative length. The dataset consists of 273 complaints filed with the California Department of Financial Protection and Innovation (DFPI) from 2022 to 2024, with ground‑truth annotations from two independent raters. In the methods, three commercial LLMs (Claude Sonnet 4.5, GPT‑5.2, Gemini 3 Pro) were evaluated against the adjudicated ground truth. The authors combine victim narratives with a structured “website” field to maximize forensic recovery and used a cross‑model verification protocol to flag discrepancies for human review. They reported that wallet addresses achieved perfect consensus (F1 = 1.000) across models, while URLs and loss amounts showed more variable performance (F1 values in the range of approximately 0.865–0.918 for loss and 0.911–0.918 for URLs). A cross‑model disagreement about URLs or loss amounts directed 31% of actionable cases to prioritized human review, functioning as a quality‑control mechanism rather than a simple output ensemble. The extraction pipeline completed in about 74 minutes at an approximate cost of $15, with a portion of cases routed for human verification. The authors note several limitations, including generalizability concerns from a single jurisdiction, English‑language reports, and the absence of end‑to‑end forensic validation such as admissibility or attribution. They frame the approach as triage support that can accelerate investigation while preserving analyst oversight.

Identified Gaps

<p>Prior literature had not addressed forensic triage taxonomy design, multi-model reliability protocols for quality control, multimodal consistency checks between narratives and screenshots, or actionability-focused evaluation. The study also identifies limited evidence on scaled screenshot-based extraction, external validity across jurisdictions and languages, and secure operational deployment of commercial LLM APIs with sensitive active-case data.</p>

Methods

<p>The study retrospectively analyzed 273 anonymized California DFPI pig-butchering complaints from 2022–2024. Two annotators created adjudicated ground truth for wallet addresses, fraudulent-platform URLs, and reported losses. Three commercial LLMs processed identical schema-driven prompts in parallel at temperature 0.0; outputs were compared with regex and spaCy NER baselines using precision, recall, and F1. Cross-model URL or loss disagreement routed cases to human review. A 25-case screenshot subset received an auxiliary vision-language cross-check.</p>

Limitations

<p>The single-agency, English-language California dataset limits generalizability across jurisdictions, languages, and evolving fraud practices. The study used one deterministic extraction run per model, so bootstrap intervals reflect resampling rather than run-to-run model variance. Wallet results rely on only 43 addresses in 25 cases. Regex and general spaCy were limited comparators, and no OCR comparator was used. Screenshot results were only a 25-case pilot. Commercial models may change over time, and retrospective public-complaint results do not establish active-case performance, attribution, admissibility, or chain of custody.</p>

Future Work

<p>Validate the protocol on cross-jurisdictional and multilingual datasets; use repeated sampling and periodic reevaluation to measure stability as commercial models change; evaluate local open-source models against the same ground truth; build a scaled multimodal study with image ground truth; and assess real-time, single-case intake processing. The authors also propose using extracted contextual fields for modus operandi classification and criminal-intelligence profiling.</p>

See how this publication connects to RSRC's living evidence syntheses through current citations and research-topic mapping.