MCAR, MAR, and MNAR: Formal Assumptions About Missing Data
Two analysts open the same CSV. Both see the same 12% of rows with a blank in the income column. One drops those rows and reports a mean of 54,000. The…

Key topics
Two analysts open the same CSV. Both see the same 12% of rows with a blank in the income column. One drops those rows and reports a mean of 54,000. The other imputes and reports 47,000. Same file. Same missing cells. Opposite conclusions.
Nothing in the file tells you who is right. That is the uncomfortable core of missing-data theory: the mechanism that produced the blanks is never observed. It is assumed. MCAR, MAR, and MNAR name three different assumptions about that hidden process — and the assumption you pick decides which conclusions your data can even support.
If you have already worked through why imputation is an estimate rather than a fact, you have the right instinct. Now we go one level down: what, precisely, is being assumed, and what does each assumption buy you?
Why the Label Is an Assumption, Not a Fact
It is tempting to treat MCAR, MAR, and MNAR as three buckets you sort your dataset into. That framing is wrong, and it causes real damage downstream.
The labels describe a data-generating process — the hypothetical machinery that decided which values got recorded and which went blank. Your dataset is one realization of that process. Many different processes can produce the same observed table. A column with 12% missing values is compatible with a sensor that randomly dropped readings, with a survey where older respondents skipped a question, and with a question that people with extreme answers refused to answer. The table looks identical in all three cases.
This is why these are called assumptions. You do not detect them from the data. You argue for them from how the data was collected — who was asked, what the instrument does when it fails, who tends to drop out. The data can sometimes rule an assumption out. It can almost never confirm one.
And the stakes are high, because the label determines what is identified: which quantities you can recover from the incomplete data at all, and which ones are simply not recoverable without extra assumptions you have to supply from outside.
The Missingness Indicator: Notation You Will Actually Use
Every formal statement of MCAR, MAR, and MNAR is written in terms of one object. Get this notation straight and the definitions stop being slogans.
Take a full data vector for one record: an outcome — say, income — and a set of fully observed covariates — say, age and education. Define the missingness indicator :
is not a fixed label you attach to a cell. It is a random variable. For each record, there was some probability that the value would be missing, and records the outcome of that coin flip. The process governing those probabilities is the missing data mechanism.
Now split into two pieces:
- — the values you actually see.
- — the values you do not see. By definition, never available to you.
The mechanism is the conditional distribution . Read that conditioning set carefully. It says: the probability of being missing can, in principle, depend on the covariates you recorded, on the values you observed, and on the values you did not. The last part is the one you can never check directly — you are conditioning on something you cannot see. That single fact is the source of every hard problem in this article.
Note: With several incomplete columns, becomes a pattern indicator — a vector describing which variables are missing together for each record. The definitions still hold, but the conditioning sets get subtler, and MAR becomes a more demanding assumption than it first appears. For now, one incomplete variable is enough to see the structure.
Knowledge check
Check your understanding
Answer this question before you continue.
MCAR: Missingness Independent of Everything
The strongest and simplest assumption:
The probability of being missing does not depend on the covariates, the observed values, the missing values, or anything else. A value goes missing for reasons unrelated to what the value would have been.
A concrete mechanism: a lab sample is physically dropped and destroyed. A sensor loses power and drops a reading. The cause of the absence is orthogonal to the measurement itself.
What it permits: complete-case analysis — dropping every row with a missing value — gives unbiased estimates of the target quantity, for that estimand and under that procedure. The remaining rows are, in expectation, a random subsample of the whole.
What it does not permit: it says nothing about efficiency. You still throw away information and shrink your effective sample size. And MCAR is often unrealistic. Treating it as the default is a modeling choice, not a neutral one — and a choice that quietly assumes the blanks are harmless.
Knowledge check
Check your understanding
Answer this question before you continue.
MAR: Missingness Explained by What You Observed
Here is where the name betrays you. "Missing at random" does not mean the missingness is unpredictable. It means the missingness is conditionally explainable by data you have:
Missingness may depend on observed values, but once you condition on those, it carries no further information about the missing value itself.
A concrete mechanism: older participants skip a follow-up visit at a rate that depends on their recorded age — but not on their unrecorded outcome, once age is accounted for. Age is in your table as a covariate . The outcome is not. That is MAR.
MAR is a much broader and more realistic class than MCAR. MCAR is the special case where nothing observed predicts missingness either. Modern missing-data methods generally start from the MAR assumption, because it is the weakest condition under which standard corrections can still identify the target.
What it permits: methods that model missingness from the observed data — multiple imputation, maximum likelihood — can correct the bias, provided the conditioning set is right and complete, and provided the estimand and model assumptions line up. That proviso is doing enormous work. Condition on the wrong variables, or omit a relevant one, and you silently reintroduce the bias you were trying to remove.
Knowledge check
Check your understanding
Answer this question before you continue.
MNAR: Missingness Depends on the Value Itself
MNAR is the negation of MAR. The probability of missingness depends on even after conditioning on and .
A concrete mechanism: people with the lowest satisfaction scores are the least likely to answer the satisfaction question. The blank is not random, and it is not explained by anything you recorded. It is explained by the number that is missing.
This is the hardest case, and not merely a third box in a taxonomy. The quantity you need to model — the relationship between missingness and the unobserved value — is exactly the quantity you never observe. Identification therefore requires an extra assumption that the data cannot supply: a belief about how the missing values relate to the observed ones, imported from domain knowledge or a sensitivity model.
The honest responses are sensitivity analysis and external information. A single point estimate presented as if the mechanism were known overstates what you actually know.
Knowledge check
Check your understanding
Answer this question before you continue.
The Wall You Cannot Test Past
Some of this is testable. Most of it is not — and knowing where the wall sits is the practical payoff.
MCAR can be probed. Build the indicator for the variable you care about, then check whether it relates to any other observed variable: a t-test against continuous columns, a chi-square against categorical ones, or a single logistic regression predicting from everything else. If nothing predicts missingness, that is consistent with MCAR. If something does, MCAR is out.
But read the result carefully. A non-significant test does not prove MCAR — it fails to find evidence against it. A significant test rules MCAR out and pushes you toward MAR, provided you are willing to condition on whatever you found.
MAR versus MNAR is structurally undecidable from observed data. Any test you build sees only relationships between missingness and things you observed. MNAR is defined by a relationship with something you did not observe. No test built from the observed data can check it, by construction. The choice between MAR and MNAR is argued on substantive grounds — how the data was collected, who dropped out, what the instrument does when it fails — never settled by a p-value.
What Each Assumption Lets You Infer
| Mechanism | Formal condition | Complete-case analysis | Standard corrections identify target? | Practical risk |
|---|---|---|---|---|
| MCAR | Unbiased for the estimand, less efficient | Yes, under the method's model assumptions | Often unrealistic; assumed by default | |
| MAR | Biased | Yes, if the conditioning set is correct and the model is right | Wrong conditioning set silently reintroduces bias | |
| MNAR | Depends on given and | Biased | No — needs an extra assumption the data cannot supply | A single estimate overstates what is known |
One distinction the table cannot capture: MCAR and MAR are sufficient conditions for certain consistent estimators, not necessary ones. A complete-case analysis can still be consistent for your estimand even when data are not MCAR, if the missingness happens not to depend on the outcome values you care about. The label guides the method; it does not dictate it. And a plausible MNAR belief does not automatically force a complex model — it forces honesty about sensitivity.
Common Misreadings and Their Consequences
"MAR means the missing values are random, so I can ignore them." This is backwards. MAR means the missingness is explainable by observed data — which is precisely what licenses a correction, and precisely why ignoring it leaves bias in place. Ignoring missingness is only defensible under MCAR.
"My MCAR test failed, so the data are MNAR." No. A failed MCAR test usually means MAR is your working assumption, because something you observed predicts missingness. Jumping straight to MNAR skips the case you can actually handle.
"I applied a MAR-based correction, so I'm covered." Only if you conditioned on the right variables. Impute using an incomplete set of predictors and you reintroduce the bias under a more sophisticated name.
"Here is the imputed result." If the mechanism was assumed rather than known, a single number hides the assumption. Report the assumption alongside the estimate.
A Worked Example: Same Table, Two Conclusions
Make it concrete. Suppose a satisfaction survey has a true underlying score on a 1–10 scale, and the lowest scores are the least likely to be recorded. Say the true distribution is uniform across 1–10, but scores of 1 and 2 are recorded only half the time, while scores of 3–10 are always recorded.
You observe only the recorded scores. The blanks are gone. Now compute the complete-case mean.
The recorded values are 3 through 10 with certainty, plus 1 and 2 at half weight. The observed mean is pulled upward — the low scores are underrepresented in exactly the rows you kept. Weight each recorded value by its recording probability and normalize:
The truth is 5.5. The complete-case mean is biased upward by roughly 0.44 — the low scores are missing from exactly the rows you kept.
Now suppose instead the blanks were caused by a random data-entry glitch, independent of the score. The same table — same blanks, same recorded numbers — would yield an unbiased complete-case mean. The observed data are equally consistent with both stories. The one fact that would distinguish them, the unrecorded low scores, is the one fact you do not have.
The data did not change. The assumption did. The conclusion flipped.
How to Choose and Defend a Working Assumption
Start from the collection process, not the column. Ask what physically or procedurally causes a value to be absent. A dropped sample and a refused question are different mechanisms, and they point at different labels.
Then test what is testable. Run the MCAR-style checks — indicator against every observed variable. If nothing predicts missingness, MCAR is a defensible working assumption. If something does, you are in MAR territory, and you must name the conditioning set explicitly: which observed variables are you assuming explain the missingness?
When MNAR is plausible — and for self-reported outcomes like income, satisfaction, or side effects, it often is — plan a sensitivity analysis rather than a single number. Ask how much the conclusion would move under different assumptions about the missing values, and report that range.
Finally, record the assumption next to the result. Write down the mechanism you assumed, the conditioning set you used, and what would have to be true for your conclusion to hold. A later reader can then challenge your reasoning instead of just your output.
That habit is the whole discipline in one line: name the mechanism, name the conditioning set, and say what would have to be true. The labels are not descriptions of your data. They are commitments about the process that produced it — and the next step is placing that commitment inside a leakage-safe preprocessing workflow, where the imputation is fit on training data only and the assumption travels with the pipeline rather than living in your head.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Build stronger machine learning foundations
Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.


