Around week five, your course produces its first real dataset: the first exam. And in most courses, nearly all of that data gets spent on its least interesting use — assigning grades — after which the spreadsheet closes and teaching resumes on intuition. Which is a strange economy, because the miss pattern sitting in that gradebook answers questions the rest of the semester depends on: which of my questions are broken, which of my concepts didn't land, and which of my students are already in trouble. Three different problems, three different fixes — and a raw score of 71% distinguishes none of them.

The postmortem is one afternoon with a spreadsheet: score every question, for every student, as a grid — rows are students, columns are items. No specialized software required, though if your LMS exports item-level results you're ten minutes in already. Then read the grid three ways.

Read one: items that indict themselves

For each item, two numbers. Difficulty — the fraction who got it right — flags the extremes: an item everyone got measures nothing (fine in small doses as a warm-up, a problem if it's a fifth of the exam), and an item almost no one got is either genuinely hard or broken, which the second number decides. Discrimination — the honest tell — compares your top scorers to your bottom scorers on that one item. A sound question is answered better by students who did well overall. When your strongest students miss an item your weakest got, or misses spread evenly across the ability range, the item is measuring something other than the course — usually its own wording. Then read the wrong answers: a distractor pulling most of the misses means one specific confusion (often teachable in ten minutes, sometimes a distractor that's defensibly also correct); misses spread evenly across all options means students were guessing, i.e., the question communicated nothing to anchor on. A hard-but-fair question and a broken question look identical as raw scores. They look completely different in a discrimination column — and only one of them should cost students points, which is why the postmortem precedes any grade appeals conversation: dropping or crediting a broken item before anyone contests it converts a fairness complaint into evidence the course self-corrects.

Read two: concepts that didn't land

Now group columns by outcome — which is exactly the mapping the alignment matrix already holds, and the reason this read takes minutes instead of an evening. Sound items clustering into a miss pattern on one outcome stops being a question-quality problem and becomes course evidence: that concept, as taught, didn't transfer. The instinct at week five is to note it and move on, because the schedule says move on. But the outcome map also says what that concept feeds — and an outcome that underpins the unit-three material is a different emergency than one that stands alone. This is the week-five advantage: a targeted twenty-minute revisit, a problem set re-anchored on the shaky concept, a recitation redirected — all cheap now, all impossible to retrofit in December when the final exam finds the same gap load-bearing.

Read three: students the pattern already names

The row-wise read is the one with a clock on it. A student who missed across the board is struggling globally — advising territory, and week five is early enough for withdrawal-free rescue. A student who aced everything except one outcome cluster has a specific, fixable gap — a fifteen-minute office-hours invitation, named concretely (“your exam was strong except the material from weeks two and three — worth a pass through problem set two before the next unit builds on it”), converts at rates a blanket “come to office hours” announcement never touches. The research on early intervention is unambiguous that week five beats week ten by more than twice; the postmortem is what makes the outreach specific, and specificity is what makes it land.

Close the loop in front of them

One more use of the same afternoon: tell the class what you found. Not scores — findings: “question 7 turned out to be ambiguous, everyone's been credited; the material from week three clearly didn't land the way I taught it, so Thursday we're taking another pass from a different angle.” This costs five minutes and buys two things that compound all semester: students learn the exam is an instrument rather than a verdict — which changes how they read every future result — and they watch assessment evidence actually change the course, which is the whole constructive-alignment argument made visible instead of asserted. A course that audits its own instruments in public has also, not incidentally, generated exactly the evidence trail — item review, findings, actions taken — that program review and accreditation ask you to reconstruct from memory in the spring.

The bottom line

The first exam is the cheapest diagnostic your course will ever run, already paid for by the time it's graded. One grid, three reads: discrimination and distractors to separate hard from broken; outcome clusters to find what didn't land while there's semester left to re-teach it; row patterns to name the students worth a specific invitation this week. Then say what you found out loud. Grades are the least of what that spreadsheet knows — the rest is a course correction with eleven weeks still on the clock.

TeachingsByDesign runs the item grid against your outcome map automatically — difficulty, discrimination, and distractor patterns per item, miss clusters per outcome, and the at-risk rows flagged in the week the findings still matter. See how it works.

References

Wiggins, G., & McTighe, J. (2005). Understanding by Design (Expanded 2nd ed.). ASCD.

About the author

Thomas R. Christian is the founder of TeachingsByDesign, an AI-native academic platform built around the coherence engine thesis — that alignment from outcomes downward, not feature-by-feature LMS plumbing, is where higher education actually breaks and where AI can actually help.

He holds a Master's in Adult and Continuing Education from Rutgers University and has spent twenty years designing instruction, training, and curriculum across enterprise CX, healthcare, and financial services. He writes about course design, AI in higher education, and the discipline of getting Stage 1 right.