Consider this question: “Do you agree that our new product is more convenient and offers better value than the previous version?”
If most respondents agree, have you learned that the product is genuinely better?
Not necessarily.
The question combines “convenient” and “better value” into one response and frames the new product positively before the respondent answers. Someone may agree with one statement but not the other, or may be influenced by the wording itself.
Questionnaire Design is therefore not simply the task of writing understandable sentences.
It is part of the Measurement Design.
A good Sample cannot rescue a question that pushes respondents toward the answer
A question that looks clear to the person who wrote it may not be interpreted consistently by respondents.
Pay particular attention to:
1. Leading Questions - Does the wording push respondents toward one answer?
2. Double-barreled Questions - Are two concepts being measured with one response?
3. Ambiguous Questions- Are terms such as “often,” “normally” or “reasonable price” undefined?
4. Assumptive Questions - Does the question assume an experience that may not apply?
5. Unbalanced Response Options - Does one side of the scale receive more response choices?
6. Question Order Effects - Do earlier questions unintentionally frame later answers?
7. Social Desirability - Does the wording make one response seem more acceptable than another?
The objective is not to eliminate every possible source of Survey Error. It is to reduce the risk that the way you ask becomes part of the answer.
How can a Survey Question create Bias?
A Survey does not directly record a customer's opinion or behavior.
The process is closer to: What we want to know → What we ask → What the respondent understands → What they answer → What we interpret
Error can enter at every stage.
Pew Research Center notes that even relatively small changes in Question Wording can affect responses and that earlier questions can influence how later questions are interpreted.
A change in Survey results therefore does not always mean customer opinion changed.
Sometimes the measurement changed.
1. Leading Questions: Is the wording pointing toward an answer?
A Leading Question uses wording or context that makes one response more attractive or expected.
For example: Avoid:
“Do you agree that our fast delivery service makes shopping more convenient?”
“Fast” and “more convenient” already evaluate the service positively.
A more neutral approach might be:
Better:
“How would you rate your overall delivery experience?”
Then measure the specific attribute separately:
“How satisfied are you with the delivery time?”
AAPOR recommends keeping questions free of bias by avoiding language that pushes respondents toward a particular response or presents only one side of an issue.
Watch for compliments hidden inside the question
For example:
“How useful is this easy-to-use new feature?”
The phrase “easy-to-use” has already been treated as established.
If Ease of Use is something you want to measure, ask:
“How easy or difficult is this feature to use?”
2. Double-barreled Questions: Two concepts, one answer
Consider: “How satisfied are you with our product's price and quality?”
What should someone answer if they like the quality but dislike the price?
AAPOR recommends that Survey Questions be specific and ask about one concept at a time.
Split the question: “How satisfied are you with the product quality?” and “How satisfied are you with the product price?”
A score of 7/10 on the original question tells you very little about whether Price or Quality created that response.
Treat “and” as a reason to inspect the question
Not every use of “and” is wrong. But phrases such as:
Price and Quality
Fast and Convenient
Product and Service
App and Website
should trigger a check: are you asking respondents to rate two different constructs with one answer?
3. Ambiguous Questions: Are respondents actually answering the same question?
Consider:
“How often do you buy from us?”
What counts as “often”?
Weekly?
Monthly?
Three times a year?
Or:
“How much do you normally spend on food delivery?”
Does “normally” mean per order, per week or per month?
A stronger question defines the Reference Period and Unit.
For example: “In the past 30 days, how many times have you ordered food through a delivery platform?”
or: “Approximately how much do you spend per food-delivery order?”
Pew recommends clear and specific wording so that respondents are more likely to interpret a question consistently.
4. Assumptive Questions: Do not assume the respondent has the experience
For example: “What is the main reason you use our Mobile Banking service?”
This assumes that the respondent uses it.
A non-user may guess, choose an irrelevant response or abandon the Survey.
Use a Screener or Routing Question first:
“In the past three months, have you used our Mobile Banking service?”
Yes → ask about use
No → ask about Awareness or Barriers
Another example:
“Why do you think Brand A is overpriced?”
The question assumes the respondent believes the brand is overpriced.
Ask first:
“How would you describe Brand A's price level?”
Then probe based on the answer.
5. Response Options can create Bias too
The problem may sit in the answers rather than the question.
Pew notes that the number, wording and order of response categories can all influence responses to Closed-ended Questions.
For example:
“How satisfied are you?”
- Extremely satisfied
- Very satisfied
- Somewhat satisfied
- Dissatisfied
The Positive side receives more space than the Negative side. A more balanced scale might be:
- Very satisfied
- Somewhat satisfied
- Neither satisfied nor dissatisfied
- Somewhat dissatisfied
- Very dissatisfied
Whether a Neutral Point is appropriate depends on the Measurement Objective. It should not simply be removed to force respondents to choose a side.
Response options should be mutually exclusive and reasonably exhaustive
Consider Age:
18–25
25–35
35–45
Where does a 25-year-old respondent belong?
A better structure is:
18–24
25–34
35–44
Closed-ended response categories should not overlap and should include reasonable possible answers. Depending on the question, options such as “Don't know,” “Not applicable” or “Other” may also be needed.

6. Question Order: Earlier questions can frame later answers
Imagine a Survey asking:
1. “How many late deliveries have you experienced this month?”
2. “Overall, how satisfied are you with the Brand?”
The respondent has just been prompted to think carefully about delivery failures.
That context may affect the overall Brand rating. This is a Question Order Effect.
Pew's Survey-methodology guidance and experiments show that questions asked earlier can influence responses to later questions. Randomization or careful sequencing may be used when appropriate to reduce or diagnose these effects.
A practical SME rule is: Ask General before Specific when you do not want specific attributes to prime the overall response.
For example:
Overall Satisfaction
before
Delivery / Staff / Price / Product Satisfaction
rather than reviewing 20 specific problems before asking the overall rating.
7. Long recall periods can create Recall Error
Consider: “How many times did you buy coffee outside your home during the past 12 months?”
Most respondents cannot accurately remember every transaction across a year.
They may estimate, round the answer or rely on recent experiences.
Choose a Recall Period based on:
- How frequently the behavior occurs
- How memorable the event is
- Whether respondents can reasonably retrieve the information
For frequent behaviors, questions such as: “During the past seven days...” or: “During the past 30 days...”
may be easier to answer. If the business already has Transaction Data, some behaviors may be better measured from actual records rather than customer memory.
8. Absolute wording can make questions unnecessarily difficult
Words such as:
Always
Never
Every time
All
may create unrealistic extremes.
For example:
“Are our employees polite every time you visit?”
A customer who experienced nine good interactions and one poor interaction has no natural answer.
A better question is: “How often have employees been polite during your visits?” with an appropriate Frequency Scale.
9. Agree / Disagree formats are convenient, but watch for Acquiescence
For example: “I believe Brand X is trustworthy.”
Strongly agree → Strongly disagree
This format is common, but some respondents may have a tendency to agree with statements regardless of their full content—known as Acquiescence Bias.
For important constructs, a more direct question may help:
“How trustworthy or untrustworthy do you consider Brand X?”
Very trustworthy → Not at all trustworthy
The respondent can evaluate the attribute directly rather than first translating their opinion into agreement with the researcher's sentence.
10. Social Desirability: Some responses feel more acceptable than others
Ask: “Do you care about the environment?”
Many respondents know that “Yes” sounds socially desirable.
It may not tell you very much about behavior.
Pew notes that some sensitive behaviors and attitudes are susceptible to Social Desirability Bias and that Survey Mode can also matter; certain responses may differ when an interviewer is present compared with a Self-administered Survey.
Where behavior matters, make the question more concrete:
Instead of: “Do you care about the environment?” ask: “In the past 30 days, how many times have you purchased a product using refill packaging?”
Even then, remember that Stated Behavior is not Observed Behavior. Transaction data may provide stronger behavioral evidence when available.
11. Loaded Questions: Has judgment already been built into the wording?
For example:
“How concerned are you about the excessively high fees for this service?”
“Excessively high” has already supplied the judgment.
A more neutral version is:
“How would you describe the fee charged for this service?”
- Very low
- Somewhat low
- About right
- Somewhat high
- Very high
If you need to know whether Fees affect purchasing, measure that separately.
The principle is simple: Do not put the conclusion you want to measure inside the question itself.
Example: The same research objective can produce very different questions
Suppose an SME wants to understand reaction to a redesigned store.
Version A
“How much better do you think our modern and more convenient new store is than the previous format?”
Problems:
- Leading
- “Modern” is supplied by the researcher
- “More convenient” is assumed
- “Better” is assumed
- Several attributes are combined
Version B
“Overall, how would you rate your experience with the redesigned store?”
Then separately:
“How would you rate the convenience of using the store?”
“How would you rate the store's atmosphere and design?”
“Compared with the previous format, how would you rate your overall experience?”
Version B does not guarantee perfect measurement.
It does make much clearer what each measure represents.
Pretest before Fieldwork: Respondents will find problems the writing team missed
A Questionnaire can look perfectly clear in a document and still fail when real respondents use it.
Pew identifies Pretesting as an essential step in Questionnaire Design, particularly when new questions are introduced.
Before launching the full Survey, ask a small number of people similar to the Target Population to complete it, then probe:
- What did you think this question meant?
- Which words were unclear?
- How did you decide on your answer?
- Was your real answer available in the options?
- Did any question feel leading?
- Was anything impossible to answer?
- Where did the Survey feel repetitive or tiring?
This is a practical form of Cognitive Pretesting and can reveal problems that are difficult for the Questionnaire author to notice.
A Questionnaire checklist before you press Send
Review every important question and ask:
- Does it measure one concept or several?
- Does the wording push toward a Positive or Negative answer?
- Are we assuming something about the respondent?
- Is the Timeframe and Unit clear?
- Will the Target Audience understand the terminology?
- Are Response Options complete and non-overlapping?
- Is the Scale balanced?
- Could earlier questions prime this response?
- Can respondents realistically remember what we are asking?
- Will this answer change or inform a real business decision?
The final question is particularly useful for removing interesting-but-unused questions that make Surveys longer without improving decisions.

A larger Sample cannot fix a biased question
Imagine asking 5,000 respondents a Leading Question.
You may estimate the resulting response with greater statistical precision.
But you may still be measuring the wrong construct.
This is the distinction between:
Sampling Error - uncertainty arising from observing a Sample rather than the full population
and
Measurement Error - error arising from how the concept is measured, including wording, interpretation and response processes
Increasing Sample Size can reduce some forms of Sampling Error.
It does not turn a Leading Question into a Neutral Question.
That is why BEE Research Knowledge treats appropriate sampling and valid, neutrally worded questions as separate requirements of strong Quantitative Research.
The takeaway: Do not ask only how many people answered the Survey—ask what you made them answer
A good Survey is not defined by Sample Size alone.
Before trusting a percentage, average score or ranking, go back to the exact Question Wording:
Was it neutral?
Did it measure one concept?
Were the definitions clear?
Were the response choices balanced?
Could Question Order have influenced the answer?
Could respondents answer from a realistic experience or memory?
Because sometimes a result such as:
“82% of customers agree”
is not only the voice of the customer.
Part of it may also be the voice of the Questionnaire Design.
The goal of good Survey Questions is therefore not to make it easy for customers to provide the answer we are looking for.
It is to make it possible for them to report what they genuinely think or do without being steered there by the question first.

Survey Bias does not come only from respondents. Leading wording, Double-barreled Questions, ambiguous terms, hidden assumptions, unbalanced response options and Question Order can introduce Measurement Error before fieldwork even begins. A stronger Questionnaire uses neutral wording, measures one concept at a time, defines timeframe and context clearly, provides appropriate response options and is pretested before launch.
Sources
- Pew Research Center. Writing Survey Questions — Question Wording, response categories, Question Order, Social Desirability and Pretesting.
- Pew Research Center. Methods 101: Survey Question Wording — clear and neutral wording and risks from poorly worded or leading questions.
- Pew Research Center. Survey Experiments Can Measure the Effects of Question Wording and More — wording and order effects.
- American Association for Public Opinion Research (AAPOR). Best Practices for Survey Research — one concept at a time, neutral wording, response options and logical Question Order.
- AAPOR. Question Wording — Double-barreled Questions, Double Negatives and Open vs. Closed-ended Question design.
.png)



.png)
.png)
.png)
.png)






