Businesses frequently encounter situations such as: Survey Satisfaction is high, but Complaints are rising.
Sales teams say customers want a new Product, but Conversion remains low.

Website Traffic is growing, but Revenue is not.
Customer Interviews identify Price as a Pain Point, while a Survey suggests Service is more important.

The easy response is to choose whichever dataset appears more “robust” or confirms what the team already believes.
But disagreement between data sources can itself be useful information.
Before deciding that one source is wrong, examine whether the Definition, Population, Time Period, Measurement and Context are actually comparable. Many apparent conflicts arise because different datasets have different bases rather than because one dataset is entirely incorrect.

Conflicting data does not always mean one source is wrong—it may mean the sources answer different questions
When two datasets disagree, do not automatically trust the larger Sample or the source that matches management experience.
Suppose a Survey finds that 80% of customers intend to buy again, while Transaction Data shows that only 45% actually repurchase within 90 days.
These numbers are not necessarily contradictory.
Purchase Intention and Actual Repeat Purchase are different Outcomes, and the Measurement Periods may also differ.
Survey results can also vary because of Question Wording, Question Order and Survey Mode. Pew Research Center documents that seemingly small differences in how questions are asked or administered can materially affect responses for some measures.
A useful sequence is: Compare → Diagnose → Weight → Explain → Decide

First, ask whether the two sources are answering the same question

Suppose a Survey reports: 80% Customer Satisfaction
while a CRM reports: 60% Retention

You should not conclude that the Survey is wrong.
Because: Satisfaction ≠ Retention

Satisfaction is an evaluation or attitude toward an experience.
Retention is observed behavior over a defined period.
A customer can be satisfied and still leave because of Price, Competition, Changing Needs, Convenience or other factors.

Before comparing numbers, write down what each measure actually represents.
Data A measures: ______
Data B measures: ______
If those sentences have different meanings, the apparent disagreement may be a Construct Mismatch rather than an Evidence Conflict.

Check Definitions before comparing numbers

The same label can mean different things across systems.
System A: Active Customer = At least one Purchase in the past 30 days
System B: Active Customer = At least one Login in the past 30 days

If they report: 10,000 Active Customers
and: 18,000 Active Customers
the difference is unsurprising.

Other examples include:
Does Revenue include VAT?
Is a Customer a person or an Account?
Does Churn mean Subscription Cancellation or no Transaction for 90 days?
Does New Customer mean First Registration or First Purchase?
Does Conversion use Visitors, Sessions or Leads as its denominator?
Before reconciling results, reconcile the Definitions.

Check the Denominator because percentages can look contradictory when their bases differ

One team reports: Complaint Rate = 10%
Another reports: Complaint Rate = 2%

Before deciding that someone calculated incorrectly, ask what each rate divides by.
10% may mean: Complaints / Customers contacting Support
while 2% may mean: Complaints / All Transactions
Both can be correct. A percentage is incomplete without its Base.

For every important Percentage, check:
What is the numerator?
What is the denominator?
Which Population?
Which Time Period?

Check whether the datasets represent the same Population

Suppose an Online Survey finds: 70% rate the Brand highly
while In-store Interviews find: 45%

The Method may not be the main reason.
The Online Sample may contain more:
Existing Digital Customers
Younger Customers
Loyalty Members

while the In-store Sample may represent:
Walk-in Customers
Specific Locations
Specific Times of Day

Before asking: “How large was the Sample?”
ask: “Who does this evidence represent?”
A large Sample drawn from the wrong Target Population does not automatically create representative evidence.

Check the Time Period before concluding that the evidence conflicts

Two findings can both be correct at different points in time.
January Survey: Customers were satisfied with Delivery.
April Operations Data: Delivery Delays increased sharply.

If the business changed its Logistics Partner in February, comparing the two findings without an Event Timeline is misleading.
Check:

  • Data Collection Date
  • Fieldwork Period
  • Transaction Period
  • Campaign Period
  • Product or Pricing Changes
  • Seasonality
  • Operational Changes
  • External Events

Publication Date is not the same as Data Period either.
A newly published report may still be based on older observations.

For Surveys, check Question Wording and Survey Mode

Two Surveys can ask about the same topic and produce different answers because the measurement differs.

For example: “Are you satisfied with our service?”
is not necessarily equivalent to: “Considering Waiting Time, Staff and Value for Money, how satisfied are you with the service overall?”

The second question introduces a context that can influence what respondents consider.
Pew Research Center notes that even relatively small Question Wording differences can affect responses and that Question Order can influence answers to later questions.
Survey Mode can matter too.
People may answer sensitive questions differently when speaking with an interviewer versus responding privately online. Pew's experimental work comparing telephone and web surveys found Mode Effects whose size varied across questions.
Therefore: Same Topic ≠ Same Measurement

Separate Stated Data from Behavioral Data

A customer says: “Promotions do not affect my buying decisions.”
Transaction Data shows that most of that customer's purchases occur during Promotions.
You do not need to declare that one source is “the truth.”

Survey or Interview Data can reveal:
Perception
Memory
Conscious Motivation
What the customer is willing or able to report
Behavioral Data reveals:
What occurred in the recorded system
But Behavioral Data does not automatically explain Motivation.
Transaction Data may show when the customer purchased.
It does not prove that the Promotion caused the purchase.
Payday, Seasonality or existing Demand may also explain the Pattern.
The two evidence types have different strengths.

Interviews and Surveys do not need to produce identical findings

Imagine 15 Interviews identify: “Uncertainty about Product Quality”
as an important Theme.
A Survey of 1,000 customers finds that only 12% select Quality as a Barrier.
Several explanations are possible:

  • The Interview Sample represents a different Segment
  • The issue is highly important to a smaller group
  • The Survey Response Option does not capture the Interview Theme accurately
  • Customers discuss Quality when probed but do not rank it as their primary Barrier
  • The two studies were conducted at different times

Mixed Methods is not valuable only when different sources confirm one another. Divergent findings can reveal additional Context or new Research Questions, and in some cases additional evidence is needed to explain the contradiction.

Audit Data Quality before weighting the Evidence

If both sources genuinely aim to measure the same thing, assess the quality of each one.
Ask:

  • What is the Data Source?
  • How much Missing Data exists?
  • Are there Duplicates?
  • Did Definitions change?
  • Was Tracking complete?
  • Did the Sample fit the Target Population?
  • Could Nonresponse affect results?
  • Was the Questionnaire valid for the Construct?
  • Was Coding consistent?
  • Could Processing Errors have occurred?

In Survey Research, a small Sampling Error alone does not guarantee strong evidence. Errors can also arise from Measurement, Coverage, Nonresponse and other stages of the research process. The Total Survey Error framework therefore considers multiple sources of error rather than Margin of Error alone.

A larger Sample does not automatically deserve more weight

Suppose:
Dataset A = 50,000 Transactions
Dataset B = Survey of 800 Customers

If the question is: “Did customers actually repurchase?”
the Transaction Data is more direct.

If the question is: “Why did customers not return?”
Transaction Data alone may be insufficient.

A Survey or Interview may be closer to the Construct being investigated.
Evidence Strength should therefore consider more than N:
Relevance to the Question
Directness
Measurement Validity
Population Fit
Data Quality
Timeliness
Method Strength

System Data is not automatically Ground Truth

Businesses often trust Behavioral Data more because: “The system does not lie.”
Systems can still measure incorrectly.

Examples:
One customer appears under multiple Customer IDs.
Website Tracking falls after a Consent Setting changes.
Offline Sales never enter the Dashboard.
Refunded Transactions remain in Revenue.
Bot Traffic is counted as Users.
Order Date is confused with Payment Date.
Behavioral Data is an observation generated by a measurement system.
It is not automatically free from Measurement Error.
Before trusting a Dashboard, understand how the Metric is created.

Aggregated results can hide very different Segment Patterns

Suppose Overall Conversion increases.
Management concludes: “Performance improved.”

But when segmented:
New Customer Conversion falls.
Returning Customer Conversion also falls.
How could Overall Conversion rise?

The mix may have shifted toward a Segment that already has a higher Conversion Rate.
If an Overall Dashboard and a Segment Report appear to disagree, examine:
Segment Mix
Channel Mix
Product Mix
Geography
Customer Type
Time Period
Aggregation can hide what is happening underneath the total.

If genuine disagreement remains, document what each source supports instead of choosing a winner

Suppose:
Survey: 72% say Price is reasonable.
Sales Interviews: Salespeople report Price as a frequent Objection.
Do not immediately conclude: Survey is right
or: Sales is right

Instead:
FACT 1: Within the Survey Sample, 72% evaluated Price as acceptable under the question used.
FACT 2: Salespeople repeatedly report Price as an Objection during Sales Conversations.
BEE INTERPRETATION: Price may not be a major issue across the existing Customer Base but could be an important Barrier among Prospects who have not converted.
HYPOTHESIS: Price Objections are concentrated among New Prospects or specific Segments.
NEXT TEST: Compare Price Perception between Converted Customers and Lost Prospects.
The contradiction has now become a better Research Question.

Weight Evidence according to the Business Question, not the seniority of the person presenting it

The CEO says: “Customers love this Product.”
That is a perspective.


The Sales Team says: “Everyone asks for this Feature.”
That is a Frontline Signal.

The Survey says: “35% select this Feature.”
That is Measurement within a defined Sample and Questionnaire.

Usage Data says: “8% use the Feature each month.”
That is Observed Behavior under a particular Tracking Definition.

All four may have value.
They should not carry equal weight for every Claim.
Ask: Which source is most directly relevant to the Claim we are about to make?

Example: Customers say they like the Product, but Sales remain weak

Survey: Concept Appeal = High
Sales: Below Target
These results are not necessarily contradictory.
Because: Liking ≠ Purchase

Investigate:
Need
Purchase Intent
Price
Availability
Competition
Distribution
Communication
Actual Trial
Repeat Purchase
The Survey may have measured Appeal accurately.
The mistake is using Appeal to answer a Purchase question that it was not designed to answer.

Example: Social Listening shows more Brand Conversation, but Market Share is flat

Social Mentions increase.
Market Share remains unchanged.
No contradiction is required.

Social Mentions measure Conversation Volume within the monitored Platforms or Dataset.
Market Share measures Sales relative to a defined Category and Market.
Mention ≠ Purchase

Similarly:
Search Interest ≠ Demand
Engagement ≠ Revenue
Awareness ≠ Preference
Interpret each dataset according to its Construct.

Example: Interviews say Service matters, but Driver Analysis gives Price a stronger association

Do not automatically declare the statistical model the winner.
Interviews may reveal that Service has strong meaning in Customer Experience.

Driver Analysis may show that Price has a stronger Statistical Association with a defined Outcome in the dataset.
But: Association ≠ Automatic Causation
And Interview Findings are not estimates of Population Effect Size.

Investigate:
Are both Methods studying the same Customer Segment?
The same Outcome?
How were Variables measured?
Could Multicollinearity matter?
Is Service functioning as a Hygiene Factor?
Does the Price relationship differ by Segment?
The disagreement may lead to a better analysis.

Use these seven steps when Evidence disagrees

  1. Define the Decision
    What decision are we trying to make?
  2. Define the Claim
    What exactly are we trying to conclude?
  3. Align the Basics
    Check Construct, Definition, Population, Denominator and Time Period.
  4. Audit the Methods
    Review Sampling, Measurement, Data Collection, Tracking and Processing.
  5. Segment the Data
    Determine whether the contradiction changes when customers, channels, products or time periods are separated.
  6. Weight the Evidence
    Consider Relevance, Directness, Validity, Population Fit and Data Quality.
  7. Define What Is Still Unknown
    If evidence remains insufficient, specify the Hypothesis and additional evidence required.

If evidence remains insufficient, specify the Hypothesis and additional evidence required.
This turns: “Which dataset should we trust?”
into: “What can each source tell us, and what evidence is still missing for this decision?”

The takeaway: Conflicting data is not something to eliminate quickly, it is something to explain

When two sources disagree, possible explanations include:

  • Different Constructs
  • Different Definitions
  • Different Populations
  • Different Denominators
  • Different Time Periods
  • Different Methods
  • Different Data Quality
  • Genuine differences across Contexts or Segments

Survey research demonstrates that Question Wording, Question Order and Mode can produce meaningful differences in responses, while the Total Survey Error framework reminds us that Evidence Quality involves more than Sampling Error alone.

So instead of asking only: “Which dataset is more trustworthy?”
ask: “For this decision, what does each dataset actually measure, what are its limitations, and after combining the evidence, what do we still not know?”
Sometimes conflicting data does not make the decision harder.
It reveals that the original question was too simple.

KEY TAKEAWAY

When two data sources disagree, do not begin by asking “Which one is right?” First determine whether they are actually measuring the same thing. Compare the Construct, Definition, Population, Denominator, Time Period, Method, Measurement and Data Quality. If these differ, the findings may not contradict one another at all—they may be answering different questions. If genuine disagreement remains, weight each source according to its fit with the Business Question and use the divergence to generate new Hypotheses or collect additional evidence rather than choosing whichever result confirms the existing belief.

Sources
  • Pew Research Center. Writing Survey Questions. Explains how Question Wording, Question Order and Survey Context can influence responses and why consistent Measurement matters when comparing results over time.
  • Pew Research Center. From Telephone to the Web: The Challenge of Mode of Interview Effects in Public Opinion Polls. A randomized study of 3,003 respondents showing that Survey Mode can produce differences for some questions and that the size of Mode Effects varies by measure.
  • Pew Research Center. Survey Experiments Can Measure the Effects of Question Wording and More. Describes randomized Survey Experiments used to identify Question Wording and Question Order Effects.
  • AAPOR-WAPOR Task Force Report on Quality in Comparative Surveys. Discusses Total Survey Error and Fitness for Intended Use as frameworks for evaluating Survey Quality across multiple sources of error.
  • Designing and Conducting Mixed Methods Research. Discusses divergent Qualitative and Quantitative findings, including how contradictions can provide new insights and may require additional data to explain.