Why Stated Intent Overpredicts Actual Behavior (2026)

Stated intent overpredicts actual behavior because a survey answer records what someone believes they will do in an artificial moment, while real choices happen under friction, budget limits, defaults, habit and social context. Respondents are rarely lying. They are giving a sincere forecast made without the money, time and effort the actual decision will demand, which makes the gap a methodology problem rather than a respondent problem.

The error is also directional, and that matters more than its size. Stated future intent almost always leans optimistic, so a business case built on it is biased in one direction, every time. The classic figure from Sheeran (2002) and Sheeran and Webb (2016) is that intentions explain roughly 18 to 23 percent of the variance in what people actually do. Anything built on the remaining 77 percent is guesswork dressed up as data.

I have sat in enough readout meetings to know what that looks like in practice: a concept test where four of five participants call something brilliant, then a launch with no pickup. The numbers were never wrong about what people felt in the room. They were just never measuring the thing the forecast needed.

Why Stated Intent Overpredicts Actual Behavior

Why Stated Intent Overpredicts Actual Behavior

Stated intent overpredicts actual behavior for five recurring reasons: social desirability bias, projection from an idealized self-image, intention without ability, faulty memory, and hypothetical bias. Each one pushes an answer upward on its own, and the survey format removes the friction that would correct them before the respondent gets to the question.

What the intention-behavior gap actually measures

The intention-behavior gap is the distance between what a person says they will do and what they do when the opportunity actually arrives. That definition sounds obvious, but the useful part is the asymmetry underneath it: the gap is almost entirely made of people who intended to act and then did not. Very few non-intenders surprise us by acting anyway.

Sheeran’s inclined-abstainer matrix puts numbers on this. Of people who report an intention, roughly 54 percent fail to act on it. Of people who report no intention, only about 3 percent act. The overwhelming majority of the say-do gap sits in the first group, which is why chasing the lapsed intender is usually more productive than trying to convert the ones who never wanted anything.

The relationship is not linear either. Behavior rates across a seven-point intent scale run something like 0, 0, 9, 17, 28, 33 and 51 percent. Nobody who scores 1 or 2 does the thing, and the jump between the middle scores is where forecasts built on averages go wrong: a pooled average of intents hides the fact that the top two rungs carry almost all the action.

What people are actually reporting when they answer

When someone ticks “definitely would buy” on a concept survey, they are reporting an intention, not making a promise. An intention is a motivation to act; a forecast is a prediction about future conduct under conditions that have not arrived yet. Surveys collect the first and are then used as though they captured the second.

Answering the question also creates a decision moment that would not otherwise exist. The respondent pauses, imagines the product, forms a fresh view of it and reports it back. That view is a real opinion, and it is worth having, but it was formed two minutes ago in a quiet room, without a competing product on the screen, without a full basket in the account, and without a credit card decision to make.

A sincere answer can still be an inaccurate prediction, and this is where most internal arguments get stuck. Someone says respondents lie, and the researcher says they do not. Both are right: nobody is fabricating. The forecast is simply built on incomplete information about the situation the respondent will actually face.

How the context changes the answer

The same person gives a different answer depending on what the question costs them. A question with no consequence is a different measurement than one that takes money or a click, and only the last two produce evidence you can build a forecast on.

MethodWhat the respondent is doingStrength of behavioral evidence
Survey intent scaleRating an abstract idea in a low-stakes settingWeak. Directional only, and consistently optimistic
Hypothetical choice taskSelecting between options that cost nothing to takeWeak to moderate. WTP estimates are the worst-affected method
Non-binding reservationCommitting to something they can walk away fromModerate. Reveals effort and ambiguity tolerance
Binding commitmentPaying, signing or clicking something with a real consequenceStrong. This is behavior, however small
Observed purchaseNothing at all, the record simply existsStrongest. Real money, real alternatives, real timing

Why predictions feel more certain than they are

Forecasts feel firmer than they are because people are not predicting from scratch. They are compressing an unfamiliar future decision into a familiar-sounding answer, and several well-documented shortcuts do the compressing for them.

Hindsight bias flattens the timeline. An answer given today is filed against a purchase date weeks out, with everything that happens in between stripped away. Telescoping, where recent events get pushed backwards to fill an empty week in a recall task, does the same damage to past-behavior questions. The planning fallacy then supplies an optimistic estimate of how quickly the thing will happen. Social desirability adds a floor under the answer: nobody wants to be the person who admits they would not spend money on the thing they are rating.

Motivated reasoning finishes the job. A respondent who wants the project approved has a stake in the answer, whether or not they are conscious of it. And asking the question at all can move it, a small effect researchers call the mere measurement effect: simply completing an intention question raises the reported likelihood of acting. Instrumentation is not passive.

The terminology, sorted out

Search results for this cluster mix six terms that are related but not interchangeable. Getting them straight prevents most of the arguments you will have about whose evidence is being misquoted.

TermWhat it describesDirection of the error
Intention-behavior gapIntention fails to produce the behavior it predictedOver-predicts
Say-do gapGeneral commercial-research label for the same distanceOver-predicts
Attitude-behavior gapEvaluations fail to translate into conductOver-predicts
Value-action gapStated values fail to produce pro-environmental or ethical actionOver-predicts
Hypothetical biasValuation inflates when nothing is actually riskedOver-predicts
Mere measurement effectMeasuring an intention raises the reported intentionOver-predicts
Self-report errorRecall of past behavior is inaccurateUnder-reports, the reverse direction

What causes the intention-behavior gap

Four mechanisms account for most of the distance, and they stack. That last part is worth holding onto, because a study fixing one bias and ignoring the rest still produces a number nobody should forecast from.

Social desirability bias. Answers skew toward what the respondent wants to be seen as, or wants the researcher to record. Vendor pages in this cluster quote the pattern bluntly, such as a claim that 65 percent of consumers report wanting to buy sustainable products while only about 26 percent follow through at checkout. The desire is genuine and the purchase still does not happen, which tells you the gap is not built out of hypocrisy.

Idealized self-image. People answer from the version of themselves who exercises, who reads labels, who organizes the kitchen. That self has more money, more time and more willpower than the one at the shop counter on a Tuesday evening. Asking about a purchase and asking about a purchase by that person are different questions.

Intention without ability. Motivation is treated as the binding constraint, but at the moment of choice the constraint is usually the price, the delivery window, the availability of the size, or the fifteen minutes it takes to complete the form. Price is also the element that is most reliably absent from a survey, because no real number is attached to the concept on the screen.

Faulty memory and telescoping. Ask someone how often they used your product in the last month and the answer is a reconstruction, not a record. Events drift toward the interview date and fill convenient gaps. This one runs the opposite way from intent and deserves separate treatment in any study that reports both.

Hypothetical bias. Nothing is at stake in a choice task, so the respondent behaves as though the downside of being wrong is zero. Willingness-to-pay questions are the clearest case: the number produced is usually a comfortable ceiling rather than a settled figure, and it is the reason so many price tests fail to move a roadmap.

What influences behavior after people state an intent

Between the survey and the decision, the environment takes over. A stated intent is set in the respondent’s head; the behavior is negotiated with a shelf, a screen, a queue and a set of defaults, and whichever is more concrete usually wins.

Consider a subscription for a meal kit. In the survey, intent scores track how often people say they would cook at home. At the decision point, the interaction is a different set of variables entirely: the price per serving that appears after the first discount period, whether skipping a week is one tap or four, whether the customer already has that card saved, and what the two most-reviewed competitors cost that week. The respondent who scored 9 out of 10 in January is now making a decision in April, with a full basket elsewhere and no memory of the survey.

Effort, defaults, social proof and time pressure all feed the same thing. Every extra step between intent and action gives the surrounding context somewhere to intervene. That is why the same respondent can honestly say they intend to buy and then quietly not, without a change of mind in the emotional sense at all.

Identity changes the relationship too. Actions that a respondent would describe as something a person like them does, or would not do, are more or less likely to follow the stated intention than actions that sit outside that story. Regret matters in the same way: an intention tied to a decision that would be embarrassing to reverse behaves differently from one with no downside attached.

When stated intent is still useful

Treating stated intent as worthless is a mistake, and it is why some researchers overreact to the evidence and go looking for a purely behavioral instrument that does not exist in most settings. Intentions remain the best available read on awareness, considered preference, perceived benefit, objection and decision criteria.

They are genuinely informative before a category exists. Someone who has never heard of a product cannot rate it, and no clickstream will tell you why they passed on the shelf. Diagnostics, awareness and language testing all depend on people describing something in their own words, which no behavioral measure substitutes for.

Predictive accuracy also varies a lot across intentions, and the moderators are reasonably well documented. In a widely cited 2022 review, Conner and Norman set out the features of intention strength: strong intentions are stable over time, less open to reinterpretation, and more likely to bias the information a person notices on the way to acting. The same review reports that intentions correlate with behavior at roughly r+ = 0.60 when they are stable and r+ = 0.27 when they are not.

Conscientiousness, goal difficulty and priority, whether the intention is instrumental or affective, and anticipated regret all shift the odds. So does the measurement itself. Self-report performs worse than objective records because of measurement error in the instrument: Armitage and Conner (2001) found intention plus perceived behavioral control explained 31 percent of variance in self-reported behavior but only 20 percent in objective behavior, and McEachan and colleagues (2011) found a similar 26 percent against 12 percent for physical activity. Some of the gap is the instrument, not the respondent.

One caution about the evidence base. Most of these effect sizes come from health and social psychology, where the behavior is attending a screening or using a condom. A commercial purchase is a different decision with different stakes, so the numbers are useful for shape and direction rather than as a coefficient to paste into your own business case. Vendor studies, such as the 65 percent to 26 percent sustainability figure quoted above, are closer to the commercial setting but rarely publish a method you could audit.

How to reduce the gap between intent and behavior

How to reduce the gap between intent and behavior

The fix is a sequence, not a single better question. Start from behavior you can already observe, add intention only where it adds something, and then test whether the combination predicts outcomes you can check. A practical order that works: ask about concrete past behavior first, measure current behavior directly, add implementation intentions, record the time and conditions, and validate the whole model against observed outcomes.

Better ways to measure likely action

Each method measures something different, and the mistake is expecting one to stand in for another. Used together they cover more ground than any of them alone.

  • Recalled behavior. Cheap and familiar, and the most vulnerable to telescoping. Useful for texture, dangerous as a base rate.
  • Behavioral and transactional data. What actually happened, with no interpretation attached. This is the anchor everything else gets validated against.
  • Choice experiments. Trade-offs forced into the open. They reveal the attributes people care about even when totals stay unreliable.
  • Binding or semi-binding tests. A deposit, a pre-order or a real sign-up with effort attached. Small stakes beat no stakes, every time.
  • Commitment scales. Ask for a specific, dated commitment rather than a probability. Specificity is what separates a plan from a mood.
  • Conjoint and discrete-choice modelling. Built for relative trade-offs between options rather than absolute willingness to pay.

Implementation intentions are the one addition worth making to any existing survey. Instead of asking whether someone will buy, ask when and where: if it is Tuesday evening and you have already eaten, you will open the account and set up the delivery. A specific if-then plan names the obstacle before it arrives, which is exactly the gap that swallows bare intentions.

How to judge whether an intention measure is reliable

The honest test of any intent measure is whether it predicts something that already happened. Most trackers never run that test, which is why a great deal of confidently wrong forecasting survives year after year.

Run a predictive validity check: take the stated scores, match them to real outcomes, and report the error rate rather than a significance value. A measure that explains a small but honest share of variance is more useful than one that produces a significant p-value on a sample nobody can replicate. Watch for overfitting, and split your sample before you start so the check is genuinely out of sample.

Then recalibrate. Teams often apply an arbitrary haircut to intent scores and argue endlessly about its size, which is a sign they have skipped validation. If you have both score and outcome, fit the conversion rate. If you do not, say so plainly in the readout rather than importing a number from someone else’s study.

A short audit of any tracker or concept study: does it ask about past behavior before future behavior? Does any option carry a real cost? Are the intent items specific and dated rather than general? Does the scale top out at certainty when the respondent has never held the product? Is there a behavioral or transactional source in the same dataset to check the score against? Are self-reported and objective measures reported separately, given that they diverge by 11 to 14 percentage points of explained variance? If most answers are no, the instrument is measuring sentiment rather than demand, whatever the fielding says.

Frequently Asked Questions

Why do people say they will buy something but then not buy it?

Because the survey answer was formed without the conditions the real decision brings. Price, time, competing options and defaults do not exist on a concept screen, and habits that do not exist in a research session take over at the moment of choice. Nothing dishonest has happened: the person reported a genuine intention and then met friction they had not budgeted for.

Is it dishonest to report a purchase intention you are not sure about?

Usually not. Respondents are answering as an idealized version of themselves, in an artificial moment, without money at stake, and that produces a sincere but inaccurate forecast. The failure belongs to the instrument. Researchers who treat it as deception get defensive and change nothing; researchers who treat it as measurement error can fix the design.

Do stated intentions ever accurately predict behavior?

They predict unevenly, and the reliability depends on the intention rather than the respondent. Stable, specific, identity-consistent intentions correlate with behavior far better than vague or short-lived ones. Where stakes and effort are attached, such as a deposit or a real sign-up, prediction improves sharply. Stated intent works best as one input, never as the forecast itself.

What is the best way to improve survey question wording?

Ask about a specific action at a specific time rather than a general likelihood. Say how likely is it you will do this, rather than how interested are you, since interest and behavior diverge sharply. Anchor the scale so the top rung means something concrete, avoid leading wording, and pair every future question with a past-behavior question to check the answer against behavior.

How far in advance should marketers measure purchase intent?

Measure close to the decision, because intentions decay and the context changes. The gap between stating an intention and acting is filled with competing priorities, revised budgets and new information, so an intent score from months earlier describes a person who no longer exists. If you must measure early, record implementation intentions with dates and conditions, then re-measure nearer the moment of choice.

Should researchers use past behavior instead of asking about future behavior?

Use past behavior as the anchor and stated intent as the supplement. Recalled behavior is easy to collect and biased in the opposite direction, since people under-report what they did not enjoy and over-report what they did. Observed or transactional records avoid both errors entirely, which is why they should carry the weight in any forecast that stated intent is informing.

Conclusion

Stated intent is not noise to be discarded, and the researchers who throw it out lose the awareness, objection and preference data that nothing else gives them cheaply. It is a single form of evidence with a known bias in a known direction, which is a much more workable thing than useless.

The practical move is small. Pick one tracker, match its intent scores to real outcomes, publish the conversion rate you find, and let that number replace the haircut your team invented last quarter. Then spend the budget you saved on one commitment test with real effort attached. That single change does more for forecast accuracy than a rebrand of the survey does.

Leave a Comment