Skip to main content
v2026.11,772 entries · CC-BY 4.0

Detecting Careless Responding: Straightlining, Speeding, and Attention Checks

How to detect careless or insufficient-effort survey responses using straightlining/long-string indices, response-time (speeding) thresholds, and attention checks — and how to turn a flag into a defensible, pre-registered exclusion rule rather than a post-hoc judgment call.

Ask CASRAI · included with Regulatory Radar

Ask about Detecting Careless Responding: Straightlining, Speeding, and Attention Checks

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Written and maintained by CASRAI Editorial Board

Last updated

Careless or insufficient-effort responding — a participant clicking through a survey without actually reading or engaging with the items — is a data-quality problem that sits underneath a study’s results without announcing itself. Unlike a missing-data cell, a careless response looks complete. It has a value in every field, it passes basic range checks, and it will happily contribute noise (and sometimes systematic bias) to every downstream statistic unless it is caught before analysis. This guide covers the three screening techniques a survey researcher can actually run — straightlining/long-string detection, response-time (speeding) analysis, and attention checks — and, more importantly, how to turn a flagged response into a defensible exclusion decision rather than a post-hoc judgment call a reviewer can challenge.

Why this matters beyond “some people don’t try hard”

Insufficient-effort responding does not distribute randomly across a Likert battery. A respondent who has stopped reading tends to default to the midpoint, to repeat the same response option down a grid (straightlining), or to answer in a fixed pattern unrelated to item content. Because this behaviour correlates with survey length, item redundancy, and respondent fatigue rather than with the construct being measured, it attenuates real correlations, inflates internal-consistency estimates that are actually measuring “did this person keep clicking the same button” rather than a coherent trait, and can distort factor structure in scale-validation work. It is a threat to construct validity, not just a nuisance in the raw data — which is why detection belongs in the methods section, not the appendix.

Straightlining and the long-string index

Straightlining is selecting the identical response option down an entire matrix or grid of items, regardless of item content or direction. The standard detection method is the long-string index: for each respondent, find the longest unbroken run of identical consecutive responses within a battery, then flag respondents whose longest run exceeds a threshold set relative to the rest of the sample (commonly a fixed cutoff such as 8–10 consecutive identical responses, or a distribution-based cutoff such as the top 1–5% of long-string lengths in the actual dataset). A distribution-based threshold is generally the more defensible choice for publication, because a fixed cutoff imported from another study’s item battery does not automatically transfer to a battery of a different length or a different response scale.

Two design choices make straightlining easier to detect after the fact:

  • Reverse-worded items distributed through a battery mean a genuine straightliner produces an internally inconsistent pattern (agreeing with both a statement and its reverse), which shows up immediately on a simple consistency check — not just the long-string index.
  • Mixed scale directions or item types within the same section make a single repeated response option implausible if the respondent were actually reading each item, sharpening what the long-string index is picking up.

Straightlining is necessary but not sufficient as a screen: a respondent who genuinely holds a uniform, un-nuanced opinion across a battery can produce the same pattern as a careless one. That is one reason straightlining is normally combined with at least one other index rather than used alone as grounds for exclusion — see the combined screening section below.

Speeding: response-time as a data-quality signal

Speeding is completing the survey (or a section of it) faster than is physically plausible for someone actually reading each item. It requires a timestamped platform — page-load and page-submit timestamps at minimum, ideally item-level timing — which most modern survey platforms capture by default even when the researcher never planned to use it for screening.

Practical thresholds researchers use:

  • Total-duration threshold: flag completions faster than some fraction of the median completion time in the actual sample (a common convention is under one-third to one-half of the median), rather than an absolute number of seconds imported from a different survey.
  • Reading-speed threshold: estimate a minimum plausible reading time from item/word count (a rough industry convention is roughly 200–250 words per minute for careful reading of survey text, though this is a general reading-speed estimate rather than a validated survey-specific standard, and should be treated as a rough floor, not a precise cutoff) and flag anyone below it.
  • Section-level timing, where available, catches a respondent who engaged normally with an early section and then sped through a later one — a pattern a total-duration threshold alone would miss.

Speeding and straightlining catch different failure modes: a respondent can straightline slowly (reading each item, still defaulting to the same answer out of fatigue or disengagement) or speed through a battery while still varying their answers superficially enough to dodge a long-string flag. Using both indices together catches more careless responding than either alone.

Attention checks, briefly

Attention checks (also called instructional manipulation checks, or IMCs) are a third, complementary screen — an item that gives an explicit instruction unrelated to the survey’s actual content (“select ‘somewhat agree’ for this item”) and checks whether the respondent followed it. CASRAI covers attention-check design, placement, and the live methodological debate over exclusion in depth elsewhere: see the Attention Checks and Instructional Manipulation Checks section of the survey question types guide for design and placement, and Attention Checks on Prolific for how one major platform’s own fairness policy constrains how attention checks can be used as rejection grounds. This guide does not repeat that ground — it treats attention checks as one leg of a three-part screening stool alongside straightlining and speeding, and focuses on how the three are combined into a single defensible protocol below.

Combining indices into one screening protocol

No single index is a reliable standalone basis for exclusion, because each one also flags some genuinely engaged respondents (a respondent with a uniform opinion straightlines; a fast, fluent reader speeds; a respondent who misreads one instruction fails an attention check). The generally recommended practice is a convergent-evidence approach: compute all the indices a study’s design allows (long-string, response time, attention-check pass/fail), and treat agreement across multiple indices as the stronger signal, rather than excluding on any single flag alone. A respondent who both straightlines and completes the survey in a third of the median time is a much stronger exclusion candidate than one who only trips a single threshold.

Practically, this means building a small composite screening table before touching the substantive analysis: one row per respondent, one column per index, with pass/fail (or a continuous score) in each cell, so the exclusion decision is traceable and auditable rather than a single opaque “data cleaning” step buried in analysis code.

Defensible exclusion rules and pre-registration

The methodological debate is not whether careless responses should be identified — it is whether and how flagged respondents should be excluded, and that decision is where a study is most vulnerable to the appearance of p-hacking if it is made after looking at the results. Three practices make an exclusion decision defensible rather than discretionary:

  • Pre-register the exclusion rule before data collection (or, at minimum, before looking at the substantive results) — which indices will be computed, what thresholds will be used, and whether exclusion requires convergent evidence across multiple indices or a single strong flag. A pre-registration that fixes the screening protocol in advance removes the temptation, real or perceived, to tune thresholds until the results look better.
  • Report both the full-sample and screened-sample results when they differ, rather than silently reporting only the screened version. If conclusions do not change, this is a brief robustness note; if they do change, that is itself a finding worth reporting transparently rather than hiding behind a single reported analysis.
  • Distinguish exclusion from correction. Some studies choose to flag and down-weight rather than drop flagged respondents entirely, particularly when the flagged group differs systematically on a demographic or trait the study cares about — outright exclusion in that case risks introducing the very bias the screen was meant to prevent. Which approach is appropriate is a design decision, not a default.

Common mistakes

  • Deciding thresholds after seeing the data’s effect on results — the single most common way a legitimate data-quality screen becomes indistinguishable from selective exclusion.
  • Using only one index. A straightlining-only or speeding-only screen misses the failure mode the other index would have caught, and is also more likely to falsely flag a genuinely engaged respondent.
  • Importing a fixed numeric threshold from a different study’s survey without adjusting for that study’s own item count, battery length, and response-scale format.
  • Treating an attention-check failure as automatically disqualifying without considering that some failures reflect a one-off misreading rather than sustained inattention — see the attention-check pages linked above for the fuller treatment of this specific debate.
  • Not reporting the screening protocol at all in the methods section, leaving readers unable to evaluate whether the reported sample was shaped by an undisclosed cleaning step.

Frequently asked questions

Is a single long-string flag enough to exclude a respondent?

Generally not on its own. A genuinely engaged respondent with a uniform opinion across a battery can trigger a long-string flag without being careless. Most defensible protocols require convergent evidence — agreement between straightlining, speeding, and/or an attention check — before excluding on that basis alone.

What response-time threshold should I use for “speeding”?

There is no single validated universal cutoff. A common, defensible approach is to set the threshold relative to your own sample’s median completion time (for example, under one-third to one-half of the median) rather than importing an absolute number of seconds from a different survey with a different item count.

Should I report results with and without excluded respondents?

Yes, whenever the two differ. Reporting only the screened-sample result, without disclosing that a different conclusion holds in the full sample, is a transparency gap a reviewer can reasonably flag.

Does careless-response screening replace an attention check, or the reverse?

Neither replaces the other — they catch different failure modes and are normally used together. See the linked guides above for attention-check design and platform-specific policy constraints.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.