Skip to content

Bloom's 2-sigma problem, explained

What a famous tutoring result actually measured, and the question it leaves for anyone building a learning system.

What would it take to give many learners the benefits of excellent individual teaching? That question makes Bloom's 2-sigma problem interesting long after the original research. It asks us to examine how instruction works, rather than assume the same explanation and pace will serve everyone.

What Bloom compared

In his 1984 paper, Benjamin Bloom described studies comparing three conditions:

  • Conventional instruction: classroom teaching with periodic tests.
  • Mastery learning: classroom teaching with formative checks, feedback and corrective work.
  • Tutoring: individual or very small group instruction, also supported by checks and correction.

The studies involved school students learning probability or cartography over a three-week period. The average tutored student typically performed about two standard deviations above the conventional group's average on the final assessment. Bloom expressed this as performing above approximately 98% of the control group.

The challenge was to find practical forms of group instruction that could approach the results of good tutoring. Cost and feasibility were part of the problem. A software product reproducing the famous number was not the study's subject.

What “two sigma” means

Sigma is a symbol for standard deviation: a measure of how spread out scores are. An effect stated in standard deviations compares a difference in average scores with that spread. It does not tell you the percentage of a subject someone learned.

As a made-up arithmetic example, suppose a comparison group averages 50 points and its standard deviation is 10 points. A group averaging 70 points sits two of those standard deviations higher: the difference is 20 points. Those are illustrative numbers, not Bloom's test scores.

This is why “98% more effective” is the wrong reading. A percentile position and a percentage increase answer different questions. Neither gives a guarantee about one learner, one tutor, or a different subject.

The useful design question

For someone building a learning system, the interesting question is what happens between an explanation and a correct answer. Does the learner have the prerequisite knowledge? Can they show their reasoning? When their answer reveals a gap, what changes about the next attempt?

Consider a learner who has read a definition of caching but cannot explain why a copy becomes stale. Repeating the definition may produce a fluent response without repairing the missing relationship. A more useful next step is an example in which the original changes while the copy remains the same, followed by a new situation for the learner to reason through.

That is a design inference we draw from thinking about instruction. It is not evidence that this example, or a particular app, produces the reported effect. It gives us something concrete to inspect in the teaching.

Ask what a learning claim actually measured

Before accepting a claim that a product “solves” the 2-sigma problem, look for the comparison. What did the other group do? Did both groups have the same study time? Were learners tested on familiar questions, new applications, or recall after a delay?

Also look at who took part. A result with school students studying a short unit cannot by itself tell us how working adults will learn a dense technical book over several months. A pleasing demo and a high completion rate are useful observations, but neither answers the retention question.

A meaningful evaluation should state its learning goal, use an appropriate comparison, and report both gains and limits. A single impressive number stripped of those conditions is a weak basis for choosing how to study.

What this means for Restudium

Restudium applies some of these learning principles to engineering study: clear teaching, an attempt to recall, an explanation to compare against, progression checks and spaced review. The private beta is a course prepared and reviewed in advance, with self-assessed answers.

They do not establish that Restudium matches a human tutor or produces a two-sigma improvement. The beta needs evidence from people using it consistently: whether they can explain what they learned, apply it in a different case and recall it after time has passed.

The lasting value of the question is the standard it sets for the work. Good learning software needs more than access to explanations. It needs a coherent system for helping a learner understand, practice, correct and return.

Build your engineering understanding.

Carefully chosen sources, clear teaching and a system for coming back. Restudium is in private beta. Join the list to hear when an invitation is available.