Skip to main content

What Lies Beneath: The Linguistic Traces of Deception in Online Dating Profiles

Catalina L. Toma ; Jeffrey T. Hancock (2012) — Journal of Communication

Citation activity

Total citations · Google Scholar
428
Average per calendar year
28.5

Count checked .

428 citations ÷ 15 calendar years (2012–2026). The first and last years may be partial. This is a lifetime average, not a year-by-year citation history.

Counts reflect Google Scholar’s coverage and do not measure research quality. Source, calculation and limitations

What do these research terms mean?
Preprint
A manuscript shared before formal peer review and publication. Check whether a later published version is available.
Dataset
A collection of data or examples for others to inspect or reuse. It can appear in Library search and topic mapping, but RSRC does not use dataset records as evidence in Research Insights.
Dissertation or thesis
Research submitted for an academic degree. This describes its format, not its reliability.
Journal article
An article published in a journal. This label alone does not establish peer review, study quality, or how well the findings apply elsewhere.
Qualitative research
Examines experiences, meanings, or processes, often through interviews or observations. It can explain how something happens without estimating how common it is.
Quantitative research
Uses numerical measurements to describe patterns or test relationships. A relationship between two measurements does not by itself show that one causes the other.
Systematic review
Uses a planned, documented method to find and assess research addressing a question. Its conclusions still depend on the included studies and what the search covered.
Meta-analysis
Statistically combines results from multiple studies. Combining studies does not remove weaknesses in their design or make unlike populations interchangeable.
Not classified
This record has no recognized label in this filter. It does not mean the publication used no method, or that no research exists.

Definitions draw on DataCite resource types; Cochrane review methods; NLM: association and causation. RSRC’s dataset and classification rules are explained in our methodology.

Citation tools


              
              

Transparency

Evidence and review status

This page contains AI-generated content. No human content review or subject-matter-expert review is recorded.

Source basis
Downloaded PDF
Source updates
No notice found at last check
AI-generated page content
Yes
Automated checks
Passed
Administrative approval
Yes
Human content review
Not recorded
Subject-matter-expert review
Not recorded
How this was prepared
Source basis

RSRC downloaded and privately stored a copy of the paper for internal analysis. The PDF is not offered to viewers from this page.

  • PDF added to RSRC:
Source updates

No incoming update notice was found in the dated Crossref response. Coverage is incomplete, particularly for corrections and expressions of concern; this is not a guarantee that the source is valid or unchanged.

  • Last source-status attempt:
AI-generated page content

AI-generated research notes displayed on this page: Synopsis, Identified gaps, Methods, Limitations, Future work. The paper itself is not described as AI-generated.

  • Document analysis recorded:
  • Synopsis generation recorded:
  • Page record updated:
Automated checks

The current, source-bound synopsis passed the recorded versioned publication checks.

  • Checks completed:
View passed checks (3)
  • Length, completeness, repetition, refusal, boilerplate, and active-markup screening
  • Numerical claims checked against the available source text
  • English-source lexical grounding check
Administrative approval

An authenticated administrator approved the bibliographic record for public Library display. This is not a review of every research claim.

  • Approved for public display:
Human content review

No human review is recorded for the AI-generated content displayed on this page.

Subject-matter-expert review

RSRC has not recorded review of this content by a subject-matter or methods expert.

Review-state definitions
Found a possible error? Request a correction.

Synopsis

This publication examines whether deception in online dating profiles leaves detectable linguistic traces in the textual self-description and whether those traces can be recognized by computers or human judges. The authors articulate two complementary studies aimed at linking profile deception with language use, and at comparing mechanical text analysis with human judgments of trustworthiness. The research treats online dating as a high-stakes, text-based context in which asynchronous and editable presentation may influence how deception is produced and perceived. In Study 1, the authors use computerized text analysis (LIWC) to identify linguistic cues associated with deception. They categorize cues as nonstrategic (emotional and cognitive states that may leak through language) and strategic (self-presentation management). Emotional indicators include first-person pronouns and negations, among others, while cognitive indicators include word count and exclusive or motion words. The deception index measures how much a profile’s stated height, weight, and age deviate from actual values. Results indicate that certain emotional cues predict deception and that word count also relates to deception when considered with other cues; cognitive cues such as exclusive and motion words showed weaker or contingent associations, influenced by the online dating context. Study 2 gathers human judgments of trustworthiness from readers of the self-descriptions. Judges rated trustworthiness, and the researchers find that the judges’ accuracy in detecting deception was no better than chance, with a notable truth bias toward perceiving most daters as trustworthy. Linguistic cues that did predict perceived trustworthiness included longer text, greater word count, and certain function-word patterns, but these cues did not reliably predict actual deception. The work discusses implications for deception theory, self-presentation, and how writing style can influence perceived trustworthiness, while acknowledging limits in detectability and classification accuracy.

Identified Gaps

The study identifies limited predictive validity for linguistic deception cues and uncertainty about how negative-emotion terms function in asynchronous, editable media. It also leaves open the relative contribution of emotional versus cognitive cues in other communication settings and the boundary conditions under which disclosed information produces positive trust and liking judgments.

Methods

Two studies examined online dating profiles collected in New York City. Study 1 analyzed 78 textual self-descriptions with LIWC2007 and regressed linguistic features on an objective deception index derived from discrepancies in reported versus measured height, weight, and age; photograph accuracy was independently rated. Study 2 had 62 undergraduates rate subsets of the same self-descriptions for trustworthiness and tested rating accuracy and linguistic predictors.

Limitations

The effective Study 1 sample was small (78 profiles), limiting automated classification; about one third of profiles were misclassified. Measures captured deception mainly through height, weight, age, and photograph accuracy, while textual descriptions were self-rated as accurate. Some analyses were correlational, so strategic motives cannot be established causally. The sample excluded homosexual participants and was recruited in New York City.

Future Work

Increase the number of profiles analyzed to improve statistical classification. Compare emotional and cognitive linguistic-cue variance across synchronous, noneditable settings to establish benchmarks. Clarify when different types of information increase positive impressions, liking, and perceived trustworthiness.

See how this publication connects to RSRC's living evidence syntheses through current citations and research-topic mapping.