Item Analysis: How to Tell a Bad Question From a Bad Result
TL;DR. A question everyone gets wrong is ambiguous evidence: it can mean the class did not learn the material, or it can mean the item is defective. Item analysis is the small set of checks that tells those apart, and it takes about ten minutes per quiz. This piece covers the two numbers worth looking at, the distractor pattern that reveals a shared misconception, the five ways questions break, and what to do about each. SimpleQuizMaker's quiz analytics computes the per-question breakdown so you are reading a table rather than tallying by hand.
The assumption worth dropping
A teacher grades a quiz and sees that twenty-two of thirty students missed question seven. The reflex conclusion is that the class did not understand that topic, and the reflex response is to reteach it.
Sometimes correct. Often not. Because a question is a measuring instrument, and a measuring instrument can be broken. If your thermometer reads 40 degrees in a cold room, the honest first question is about the thermometer.
Questions break in ways that are invisible when you write them and obvious in the data. You know what you meant; you cannot unknow it while rereading your own wording. The class does not have that advantage, and the result pattern is where the gap shows up.
Item analysis is just the habit of looking at that pattern before deciding what it means.
The two numbers
You do not need psychometrics. Two numbers cover most of the value.
Difficulty (p-value)
The proportion of students who answered correctly. If 12 of 30 got it right, difficulty is 0.40.
The name is backwards and trips everyone up at first: a high p-value means an easy question. 0.90 is easy, 0.30 is hard.
Rough reading for classroom use:
| p-value | Reading |
|---|---|
| Above 0.90 | Nearly everyone right. Fine as a warm-up, tells you little. |
| 0.50 to 0.85 | The useful band for most classroom assessment. |
| 0.25 to 0.50 | Hard. Legitimate, but check the item. |
| Below 0.25 | Below or near guessing on a 4-option item. Investigate. |
That last row is the one to act on. On a four-option multiple choice, random guessing produces about 0.25. A question performing at or under chance is not measuring knowledge — it is either broken, or it is measuring something other than what you think.
Discrimination
Whether the students who did well overall also did well on this item.
The quick version: split the class into the top third and bottom third by total score. Discrimination is the fraction correct in the top group minus the fraction correct in the bottom group.
If 80% of the top third and 30% of the bottom third got it right, discrimination is 0.50. Strong.
| Discrimination | Reading |
|---|---|
| Above 0.40 | Excellent. Separates the students who know it from those who do not. |
| 0.20 to 0.40 | Acceptable. |
| 0 to 0.20 | Weak. The item is not distinguishing much. |
| Negative | Red flag. Weaker students outperformed stronger ones. |
Negative discrimination is the single most useful signal in item analysis. It means the students who understood the material best were *more* likely to get it wrong — which almost never happens because strong students suddenly forgot one fact. It happens because the item rewards a misreading, or because the answer key is wrong, or because a subtlety the strong students know about makes the "correct" answer look wrong to them.
If you check one thing, check for negative discrimination. It finds broken items with remarkable reliability.
Reading the distractor pattern
Beyond the two numbers, the distribution across wrong options tells you what kind of wrong.
Wrong answers spread evenly. Roughly equal numbers on each incorrect option. This is the signature of guessing — nobody had a model, so choices scattered. Means the content was not learned at all, or the question was incomprehensible.
One distractor dominates. Most of the wrong answers on a single option. This is the informative case, and it means something specific: the class shares a misconception. They did not guess. They reasoned, together, to the same wrong place.
This is the most valuable thing item analysis produces. You now know not just that they got it wrong but *what they believe instead*, which is exactly what you need to teach against. A reteach aimed at "the correct answer is B" will bounce off; a reteach aimed at "many of you think X, here is where X breaks" lands.
A distractor gets zero selections. Nobody picked it. It is not doing any work, and a four-option question where one option is obviously absurd is a three-option question with worse odds for you. Replace it with a plausible wrong answer — ideally one that captures a real misconception.
More students pick a distractor than the key. Check the answer key first. Genuinely — this is the most common cause. If the key is right, you have a class-wide misconception strong enough to override the correct answer, which is worth a lesson of its own.
Five ways questions break
Most defective items fall into five categories.
1. Ambiguous wording
Two readings, both defensible, only one keyed. Usually from an unclear pronoun, a vague qualifier, or a question that does not specify which aspect it is asking about.
*"Why is the reaction faster?"* — faster than what? Under which condition?
Signature: moderate difficulty, low or negative discrimination. Strong students see the ambiguity and choose the other reading.
Fix: specify. Name the comparison, the condition, the units.
2. Unintended cueing
Something outside the content gives the answer away. The correct option is noticeably longer or more detailed. The stem says "an" before a vowel-initial option. Absolute words like "always" and "never" appear only in wrong options, which test-wise students learn to avoid.
Signature: high p-value with unexpectedly low discrimination — everyone gets it right, so it separates nobody.
Fix: match option lengths, use "a/an" or restructure, distribute absolutes.
3. Testing recall of trivia
Technically about the topic, actually about whether a specific number stuck. The year an obscure law passed, the fourth item in a list of seven.
Signature: low p-value, low discrimination. Strong students do not memorize trivia either.
Fix: ask what the thing *does* rather than what it is numbered.
4. A wrong or arguable key
The keyed answer is wrong, or another option is also defensible.
Signature: negative discrimination, often severe. One distractor beats the key.
Fix: verify against your source. If two options are defensible, the item is broken even if your key is the better answer — a question with two right answers measures agreement with the author, not knowledge.
5. Reading load instead of subject knowledge
Dense stem, unnecessary jargon, three clauses before the actual question. You are measuring reading comprehension, which matters but is not what the quiz claims to measure. This hits English-language learners and students with reading difficulties hardest, which makes it an equity problem rather than only a measurement one.
Signature: low p-value across the board; weaker correlation with subject knowledge than the rest of the quiz.
Fix: cut the stem to the shortest form that is still unambiguous. Put necessary context first, question last.
A worked example
Twenty-eight students, ten questions, a unit on photosynthesis. Three items are worth looking at.
Question 3. p-value 0.93, discrimination 0.05. Twenty-six of twenty-eight correct.
Nothing is broken here, but the item is not earning its place. It separates nobody — the strongest and weakest students performed identically. Keep it as an opener if you want students to start with a win, but do not count it as evidence of anything. Two or three items like this in a ten-question quiz means the quiz has less measuring power than its length suggests.
Question 6. p-value 0.32, discrimination 0.48. Nine correct; of the nineteen wrong, fifteen chose option C.
This is a healthy hard item plus a real finding. Discrimination is strong, so the item is working — students who understood the unit got it right at much higher rates. And the wrong answers are not scattered: fifteen of nineteen landed on the same option. That is a shared misconception, not guessing.
Reading option C tells you what they believe. If C attributed the oxygen released in photosynthesis to carbon dioxide rather than water, you have just learned the single most useful thing this quiz could tell you — and it is a genuinely reasonable error, since the carbon in glucose *does* come from carbon dioxide. The next lesson writes itself.
Question 9. p-value 0.39, discrimination -0.22.
Negative discrimination. The top third did *worse* than the bottom third. This is almost never a knowledge story.
Rereading the stem: it asks which factor "limits the rate of photosynthesis," and the keyed answer is light intensity. But the stem never specifies conditions — and the students who understood limiting factors best know the answer depends entirely on which factor is currently scarcest. They saw a question with no determinate answer and split their choices; the students with a shallower model matched "photosynthesis" to "light" and got the mark.
The item rewarded not knowing the complication. Drop it from the grade, rewrite it with a specified condition, and notice that the strongest students were right to be troubled.
What to tell students
Ready to create your first quiz?
Use AI to generate quizzes from your own study materials in seconds.
Create a Free Quiz — Sign UpTwo things are worth saying out loud when you drop or fix an item, and both pay off beyond the single quiz.
Say that you checked. "I looked at question nine, the wording was ambiguous, it is not counted." Students assume assessment is something done to them by an infallible authority. Watching a teacher audit their own instrument teaches something real about how measurement works, and it buys credibility for the items you *do* stand behind.
Say what the misconception was. When a distractor dominates, name it: "Fifteen of you picked C. Here is why C is tempting and here is exactly where it breaks." Students who chose C learn something specific rather than feeling generically wrong, and students who chose correctly find out whether they did so for the right reason — which a green tick never tells them.
There is also a quieter benefit. A class that has watched you take a broken question seriously will tell you when a question seems broken, which is a better bug-reporting system than any amount of solo proofreading.
A ten-minute routine
After each quiz:
Sort by difficulty and look at the bottom. Any item under 0.25 gets checked.
Check discrimination on those. Negative means the item is probably broken. Positive but low means it may be genuinely hard.
Look at the distractor spread for each flagged item. Even spread means guessing. Concentrated means misconception.
Reread the flagged stems as if you did not know the answer. Hardest step, most valuable. Ask whether another reading is defensible. Better still, ask a colleague who does not teach the unit — someone who does not know what you meant is exactly the reader you need.
Decide: fix, drop, or teach. Fix the item for next time. Drop it from the current grade if it was defective. Or, if the item is sound and the misconception is real, teach against it.
That last decision is the point. Item analysis is not about perfect questions — it is about not mistaking a broken instrument for a broken class, and not mistaking a real gap for bad luck.
Sample size, honestly
With 25 to 30 students, these numbers are suggestive, not precise. Discrimination computed on top and bottom thirds of a class of 28 rests on nine students per group, and one unusual student moves it noticeably.
That is fine for the purpose. You are not calibrating a standardized test; you are flagging items worth a second look. A flagged item still needs you to read it and decide.
What you should not do is compute these for a class of eight and treat them as measurements. Below roughly twenty responses, read the distractor pattern and skip the discrimination arithmetic — the pattern survives small samples better than the statistic does.
And across a term, the same item used with several cohorts accumulates a much more trustworthy picture than any single sitting. Keeping a short note on items you have flagged before is worth more than any single quiz's numbers.
What the product does
Quiz analytics on the Student plan computes the per-question breakdown — how many got each item right, how the wrong answers distributed across options, and how each item relates to overall performance. It is the same analysis described above, without the spreadsheet.
Analytics is a paid-plan feature; the free plan includes quiz creation and sharing but not the analytics dashboard. And a note on scope: this is class-level item analysis for classroom assessment. It is not a psychometric calibration suite, and the numbers should be read as flags rather than as certified item parameters.
Item analysis on AI-generated questions
If you generate questions rather than write them, the analysis matters more, not less — and the failure modes shift.
A generator rarely produces the classic human errors. It matches option lengths reasonably well and does not usually leave an "an" before a consonant. What it does produce, more often than a human would, is two defensible options: a distractor that is wrong in the textbook sense but arguable in a way the generator did not notice. That shows up as negative discrimination, and it is the single most common defect in generated items.
The second pattern is the near-miss distractor that is actually correct under a different interpretation of the stem — the question 9 failure above, generated rather than written.
The practical consequence: review generated questions before use, and analyze them after. Neither step replaces the other. Pre-use review catches the obviously wrong; item analysis catches the plausibly-wrong-in-a-way-you-could-not-see, which is the category that survives proofreading precisely because it looks fine.
Over a few cycles this gets cheap. You learn which topics your generator handles cleanly and which ones need a closer read — abstract or contested material tends to need it, concrete factual material usually does not.
The habit underneath
The teachers who get the most out of assessment are not the ones who write perfect questions. Nobody writes perfect questions; every experienced item writer has shipped an ambiguous stem.
They are the ones who look at the data before drawing a conclusion — who treat "twenty-two students missed this" as the beginning of a question rather than the answer to one.
That habit costs ten minutes per quiz and changes what your assessments are worth. A quiz you have never analyzed is a grade. A quiz you have analyzed is a map of what your class believes, including the parts they believe confidently and wrongly.
Related reading
Get weekly study & quiz tips
Join teachers and students who get practical tips on quizzing, active recall, and AI-powered learning.
Sarah Mitchell
Curriculum Designer & Former High School Teacher
More articles by Sarah →
Practice with AI-generated quizzes
Try it on your own material
SimpleQuizMaker turns what you are already teaching or studying into practice questions. The free plan includes 5 quiz generations a month and needs no card.
Ready to create your first quiz?
Use AI to generate quizzes from your own study materials in seconds.
Create a Free Quiz — Sign Up