Behavioural customer research: when interviews are not enough

This piece sits in Polymorph’s series on the seven fundamentals of great products. We’re on the second one: move from “who is the user?” to what they repeat, avoid, pay for, and abandon when the old way breaks.

If customer discovery is first-pass clarity, behavioural research is second-pass rigour. It’s where confident teams get humble. Or they ship the wrong thing because “we talked to ten people” felt like enough.

I first framed this on LinkedIn as the 2nd fundamental. Knowing a title and a pain keyword is not the same as understanding how someone decides, what they fear, or what they already rejected. The upgrade is from profile to something you can act on. “CFOs at mid-sized logistics firms” is a label. “CFOs fighting delivery delays, cost pressure, and an ERP everyone works around with spreadsheets” is a starting point for a build.

TL;DR: Atlassian finds only 60% of teams make experimentation a regular habit, while 40% do little or none. That’s the gap where early interview confidence hardens into false certainty. Behavioural research closes it by forcing contradictions, specifics, and signals into the open.

What changes when you shift from opinions to behaviour?

You stop asking only what people believe. You ask what they did last time the problem showed up. Sounds obvious. It still makes rooms uncomfortable, because polite answers and messy reality diverge more often than teams admit.

Nielsen Norman Group draws a clean line: attitudinal research gathers self-reported thoughts and feelings; behavioural research observes what people actually do. What users say and what they do are often quite different. Memory is imperfect. People struggle to explain their own habits. Social norms creep in. The mismatches between the two are often where the insight lives.

NN/g’s methods map puts the same idea on a spectrum: “what people say” versus “what people do.” Usability studies and field studies sit in the middle, mixing both, and they recommend leaning toward the behavioural side when you can.

CB Insights ties 43% of analysed startup failures to poor product–market fit. Plenty of those stories include teams that heard demand but never observed switching conditions: frequency, workarounds, budget triggers, and the political cost of change.

Nine second-pass research moves

Once first-pass discovery has a beachhead and a few artefacts, these are the moves we use to stress-test early confidence.

  1. Return interviews. Same people, new prompts. Compare stories week to week. You’re looking for stability, not a better performance of the pitch.
  2. “Show me the last time.” Force artefacts. Screenshots, exports, tickets, calendar blocks. Generic pain language should not survive this prompt.
  3. Contradiction hunting. Where do stated priorities conflict with observed behaviour? That tension is adoption risk on a map, not “bad data.”
  4. Contextual observation / shadowing. Watch the job in the environment where it happens. Spreadsheets and side chats show up here that never make the interview notes.
  5. Usage signals. Activation, return intervals, drop-off steps. Even rough logs beat pure recall. No product telemetry yet? Proxy with support themes and sales-cycle notes.
  6. Buying archaeology. Map rejections, stalled pilots, and “we’ll revisit Q3.” Patterns here explain no decision better than any competitor teardown.
  7. Diary or lightweight logging. Frequency and triggers over several days. What people remember in a meeting and what they repeat every week are different animals.
  8. Support-theme coding. A week of tickets or chat logs, tagged for the job you’re targeting. Cheap. Often brutal. Useful.
  9. Minimum experiment. A prototype, concierge slice, scripted demo, or small pilot with a named behaviour you expect to change. Prediction in, outcome out.

You don’t need all nine every time. You need enough that a polite “we love it” can’t survive contact with the workflow.

Contradiction checklist (say vs do)

Watch for these mismatches. Each one is a roadmap fork.

  • They rank the problem as “critical,” then schedule the follow-up three months out.
  • They praise your demo, then open the old spreadsheet mid-call.
  • They say they’ll switch “as soon as we have time,” with no owner and no date.
  • They ask for AI / automation, then refuse any path that changes how work is approved.
  • Champions love it; ops can’t run it without three new headcount.

Example A B2B team hears enthusiasm for “AI-assisted reporting” in interviews. Shadowing shows analysts still export everything to Excel because procurement hasn’t approved a new UI path. The contradiction is the roadmap: put AI on a workflow people won’t adopt, or fix trust and access first. Behavioural research turns that from a debate into evidence.

What questions expose depth without wasting the room?

Keep the list short. Demand evidence.

  • What single problem would you fund this quarter if you could?
  • What’s the full workaround, including spreadsheets, side chats, and manual checks?
  • Which tools did you try and abandon, and what broke trust?
  • Who can veto, champion, or bear operational risk?
  • Is the decision single-owner or committee, and what evidence actually moves it?

If you can’t answer those with artefacts, you understand language. You don’t yet understand behaviour.

Why does “experimentation habit” matter this much?

Atlassian reports that 40% of teams do little or no experimentation. That’s not only a tooling gap. It’s a cultural default: shipping opinions instead of testing them.

Experiments don’t need a perfect lab. They need a repeatable way to compare predictions with outcomes. Teresa Torres’s continuous discovery framing pairs weekly customer contact with assumption testing (Product Talk): interviewing surfaces opportunities; small tests tell you which solutions survive contact with reality.

What does a good behavioural research cadence look like?

Think in loops, not workshops.

  1. Hypothesis: a crisp claim about behaviour (“users will return weekly if we automate step X”).
  2. Minimum test: the cheapest artefact that could falsify it.
  3. Decision rule: what you’ll do if results are weak (narrow scope, change segment, or kill).

Write the decision rule before you run the test. Otherwise every weak result becomes a debate about “needing more data,” which is how teams burn another quarter protecting a favourite idea.

McKinsey’s December 2024 work on effective product teams emphasises outcomes like delivery predictability and value realisation. Behavioural research is how you keep value hypotheses honest before they harden into roadmaps and headcount plans.

We’ve found the cadence that sticks is boring on purpose: one return conversation or observation touch most weeks, plus one small test that could change next sprint’s priority. Flashy research weeks followed by radio silence teach the organisation the wrong lesson.

One more practical filter: if your “experiment” can’t change a backlog item within two weeks, it was probably a workshop with nicer branding.

How should product and sales compare notes without politicising the truth?

Sales hears deal-level urgency. Product hears pattern-level frequency. Behavioural research works when both sides bring artefacts, not competing stories.

  • Win/loss reviews tied to workflow evidence, not only price.
  • Pilot definitions that name the behaviour you expect to change.
  • Champion vs blocker mapping that includes operations, not only the signature.

When product and sales tell different stories, assume neither is inventing drama. Assume your evidence base is still incomplete.

What if you can’t run experiments at “big tech” scale?

You’re not trying to publish a paper. You’re trying to reduce uncertainty before the build hardens.

Small can still be rigorous: five return interviews, one week of support-ticket coding, a single pilot with explicit KPIs. The habit matters more than the lab coat. NN/g is clear that hybrid methods (observe and ask) give you both the what and the why. You can do that with a screen share and a notepad.

Atlassian’s State of Product 2026 states that 40% of product teams do little or no experimentation, while 60% practise it regularly: a discipline gap that turns interview quotes into false confidence. NN/g distinguishes attitudinal (self-reported) from behavioural (observed) research and notes that what users say and what they do often diverge. CB Insights attributes 43% of studied failures to poor product–market fit, a failure mode second-pass behavioural work is meant to catch before you scale the wrong change.

Where do you go next?

Behavioural truth sets up value confirmation in product validation before build. If you still need first-pass context, start with customer discovery. The full arc is in the seven fundamentals overview.

Experimentation habit (Atlassian, illustrative) Bar split 60 percent regular experimentation 40 percent little or none. Teams with regular experimentation habits Source: Atlassian State of Product 2026. Regular habit — 60% Little / none — 40%
Figure 1. Atlassian reports 60% of teams make experimentation a regular habit; 40% do little or none (State of Product 2026).

Want to turn behavioural signals into a build plan you can defend? Talk to Polymorph — we work with founders, product leads, and executive teams to validate, build, and improve software that earns its place in the business.

FAQ

Is behavioural research only for consumer products with lots of data?
No. B2B teams can use artefacts, pilots, return interviews, and a week of support coding. Thin evidence beats fabricated certainty.

What if customers tell us what we want to hear?
Ask for last-time proof and hunt contradictions. NN/g treats say-versus-do mismatches as insight, not noise. Politeness is not validation.

How does this relate to jobs-to-be-done?
JTBD helps frame the job. Behavioural research tests whether people will change habits, budgets, and workflows for that job in practice.

What is the difference between this and UX usability testing?
Usability tests whether people can complete a flow. Behavioural research asks whether they’ll change habits, budgets, and workarounds when the demo ends. You usually need both.

What should we read next in the series?
Product validation before build, then competitive analysis for new products when you’re ready to map what you’re really displacing.

Sources

Planning to build an app? 

Try our free software development calculator to maximise your ROI.

Request for Access to Information

The following forms are available to download with regards to request for access to information:

REQUEST FOR ACCESS TO RECORD

OUTCOME OF REQUEST AND OF FEES PAYABLE

INTERNAL APPEAL FORM