Skip to content
intermediate

MCAR, MAR, and MNAR: Formal Assumptions About Missing Data

Two analysts open the same CSV. Both see the same 12% of rows with a blank in the income column. One drops those rows and reports a mean of 54,000. The…

Published 2026-10-02Updated 2026-10-0412 min read
Numerous wires and cables mounted into server patch panel in modern data center
Numerous wires and cables mounted into server patch panel in modern data center. Photo by Brett Sayles on Pexels.

Two analysts open the same CSV. Both see the same 12% of rows with a blank in the income column. One drops those rows and reports a mean of 54,000. The other imputes and reports 47,000. Same file. Same missing cells. Opposite conclusions.

Nothing in the file tells you who is right. That is the uncomfortable core of missing-data theory: the mechanism that produced the blanks is never observed. It is assumed. MCAR, MAR, and MNAR name three different assumptions about that hidden process — and the assumption you pick decides which conclusions your data can even support.

If you have already worked through why imputation is an estimate rather than a fact, you have the right instinct. Now we go one level down: what, precisely, is being assumed, and what does each assumption buy you?

Why the Label Is an Assumption, Not a Fact

It is tempting to treat MCAR, MAR, and MNAR as three buckets you sort your dataset into. That framing is wrong, and it causes real damage downstream.

The labels describe a data-generating process — the hypothetical machinery that decided which values got recorded and which went blank. Your dataset is one realization of that process. Many different processes can produce the same observed table. A column with 12% missing values is compatible with a sensor that randomly dropped readings, with a survey where older respondents skipped a question, and with a question that people with extreme answers refused to answer. The table looks identical in all three cases.

This is why these are called assumptions. You do not detect them from the data. You argue for them from how the data was collected — who was asked, what the instrument does when it fails, who tends to drop out. The data can sometimes rule an assumption out. It can almost never confirm one.

And the stakes are high, because the label determines what is identified: which quantities you can recover from the incomplete data at all, and which ones are simply not recoverable without extra assumptions you have to supply from outside.

The Missingness Indicator: Notation You Will Actually Use

Three side-by-side panels show R alone for MCAR, arrows from observed X and Y-observed to R for MAR, and an additional arrow from hidden Y-missing to R for MNAR.
Compare the assumptions by tracing which variables are allowed to predict whether Y is observed; only MNAR depends on the unseen value itself.

Every formal statement of MCAR, MAR, and MNAR is written in terms of one object. Get this notation straight and the definitions stop being slogans.

Take a full data vector for one record: an outcome YY — say, income — and a set of fully observed covariates XX — say, age and education. Define the missingness indicator RR:

R={1if Y is observed0if Y is missingR = \begin{cases} 1 & \text{if } Y \text{ is observed} \\ 0 & \text{if } Y \text{ is missing} \end{cases}

RR is not a fixed label you attach to a cell. It is a random variable. For each record, there was some probability that the value would be missing, and RR records the outcome of that coin flip. The process governing those probabilities is the missing data mechanism.

Now split YY into two pieces:

  • YobsY_{\text{obs}} — the values you actually see.
  • YmisY_{\text{mis}} — the values you do not see. By definition, never available to you.

The mechanism is the conditional distribution P(R∣X,Yobs,Ymis)P(R \mid X, Y_{\text{obs}}, Y_{\text{mis}}). Read that conditioning set carefully. It says: the probability of being missing can, in principle, depend on the covariates you recorded, on the values you observed, and on the values you did not. The last part is the one you can never check directly — you are conditioning on something you cannot see. That single fact is the source of every hard problem in this article.

Note: With several incomplete columns, RR becomes a pattern indicator — a vector describing which variables are missing together for each record. The definitions still hold, but the conditioning sets get subtler, and MAR becomes a more demanding assumption than it first appears. For now, one incomplete variable is enough to see the structure.

Knowledge check

Check your understanding

Answer this question before you continue.

For a record with outcome Y, what does R = 0 mean, and what does R represent?
Single Choice

Focus: Interpret the missingness indicator R and distinguish its observed realization from the mechanism that generates it.

MCAR: Missingness Independent of Everything

The strongest and simplest assumption:

P(R∣X,Yobs,Ymis)=P(R)P(R \mid X, Y_{\text{obs}}, Y_{\text{mis}}) = P(R)

The probability of being missing does not depend on the covariates, the observed values, the missing values, or anything else. A value goes missing for reasons unrelated to what the value would have been.

A concrete mechanism: a lab sample is physically dropped and destroyed. A sensor loses power and drops a reading. The cause of the absence is orthogonal to the measurement itself.

What it permits: complete-case analysis — dropping every row with a missing value — gives unbiased estimates of the target quantity, for that estimand and under that procedure. The remaining rows are, in expectation, a random subsample of the whole.

What it does not permit: it says nothing about efficiency. You still throw away information and shrink your effective sample size. And MCAR is often unrealistic. Treating it as the default is a modeling choice, not a neutral one — and a choice that quietly assumes the blanks are harmless.

Knowledge check

Check your understanding

Answer this question before you continue.

Under MCAR, what does the article say about complete-case analysis, for the target estimand and procedure?
Comparison Reasoning

Focus: Explain the complete-case consequence of MCAR while distinguishing unbiasedness from efficiency.

MAR: Missingness Explained by What You Observed

Here is where the name betrays you. "Missing at random" does not mean the missingness is unpredictable. It means the missingness is conditionally explainable by data you have:

P(R∣X,Yobs,Ymis)=P(R∣X,Yobs)P(R \mid X, Y_{\text{obs}}, Y_{\text{mis}}) = P(R \mid X, Y_{\text{obs}})

Missingness may depend on observed values, but once you condition on those, it carries no further information about the missing value itself.

A concrete mechanism: older participants skip a follow-up visit at a rate that depends on their recorded age — but not on their unrecorded outcome, once age is accounted for. Age is in your table as a covariate XX. The outcome is not. That is MAR.

MAR is a much broader and more realistic class than MCAR. MCAR is the special case where nothing observed predicts missingness either. Modern missing-data methods generally start from the MAR assumption, because it is the weakest condition under which standard corrections can still identify the target.

What it permits: methods that model missingness from the observed data — multiple imputation, maximum likelihood — can correct the bias, provided the conditioning set is right and complete, and provided the estimand and model assumptions line up. That proviso is doing enormous work. Condition on the wrong variables, or omit a relevant one, and you silently reintroduce the bias you were trying to remove.

Knowledge check

Check your understanding

Answer this question before you continue.

Older participants are more likely to miss a follow-up, and age is recorded. Assume that after conditioning on age, missingness carries no further information about the unrecorded outcome. Which mechanism does this describe?
Scenario Interpretation

Focus: Identify MAR when missingness depends on observed covariates but has no remaining dependence on the unobserved outcome after conditioning on them.

MNAR: Missingness Depends on the Value Itself

MNAR is the negation of MAR. The probability of missingness depends on YmisY_{\text{mis}} even after conditioning on XX and YobsY_{\text{obs}}.

A concrete mechanism: people with the lowest satisfaction scores are the least likely to answer the satisfaction question. The blank is not random, and it is not explained by anything you recorded. It is explained by the number that is missing.

This is the hardest case, and not merely a third box in a taxonomy. The quantity you need to model — the relationship between missingness and the unobserved value — is exactly the quantity you never observe. Identification therefore requires an extra assumption that the data cannot supply: a belief about how the missing values relate to the observed ones, imported from domain knowledge or a sensitivity model.

The honest responses are sensitivity analysis and external information. A single point estimate presented as if the mechanism were known overstates what you actually know.

Knowledge check

Check your understanding

Answer this question before you continue.

Why can observed data alone not generally identify the relationship between missingness and the unobserved value under MNAR?
Misconception Check

Focus: Explain why identifying quantities under MNAR requires assumptions or information beyond the observed data.

The Wall You Cannot Test Past

Some of this is testable. Most of it is not — and knowing where the wall sits is the practical payoff.

MCAR can be probed. Build the indicator RR for the variable you care about, then check whether it relates to any other observed variable: a t-test against continuous columns, a chi-square against categorical ones, or a single logistic regression predicting RR from everything else. If nothing predicts missingness, that is consistent with MCAR. If something does, MCAR is out.

But read the result carefully. A non-significant test does not prove MCAR — it fails to find evidence against it. A significant test rules MCAR out and pushes you toward MAR, provided you are willing to condition on whatever you found.

MAR versus MNAR is structurally undecidable from observed data. Any test you build sees only relationships between missingness and things you observed. MNAR is defined by a relationship with something you did not observe. No test built from the observed data can check it, by construction. The choice between MAR and MNAR is argued on substantive grounds — how the data was collected, who dropped out, what the instrument does when it fails — never settled by a p-value.

What Each Assumption Lets You Infer

MechanismFormal conditionComplete-case analysisStandard corrections identify target?Practical risk
MCARP(R∣X,Yobs,Ymis)=P(R)P(R \mid X, Y_{\text{obs}}, Y_{\text{mis}}) = P(R)Unbiased for the estimand, less efficientYes, under the method's model assumptionsOften unrealistic; assumed by default
MARP(R∣X,Yobs,Ymis)=P(R∣X,Yobs)P(R \mid X, Y_{\text{obs}}, Y_{\text{mis}}) = P(R \mid X, Y_{\text{obs}})BiasedYes, if the conditioning set is correct and the model is rightWrong conditioning set silently reintroduces bias
MNARDepends on YmisY_{\text{mis}} given XX and YobsY_{\text{obs}}BiasedNo — needs an extra assumption the data cannot supplyA single estimate overstates what is known

One distinction the table cannot capture: MCAR and MAR are sufficient conditions for certain consistent estimators, not necessary ones. A complete-case analysis can still be consistent for your estimand even when data are not MCAR, if the missingness happens not to depend on the outcome values you care about. The label guides the method; it does not dictate it. And a plausible MNAR belief does not automatically force a complex model — it forces honesty about sensitivity.

Common Misreadings and Their Consequences

"MAR means the missing values are random, so I can ignore them." This is backwards. MAR means the missingness is explainable by observed data — which is precisely what licenses a correction, and precisely why ignoring it leaves bias in place. Ignoring missingness is only defensible under MCAR.

"My MCAR test failed, so the data are MNAR." No. A failed MCAR test usually means MAR is your working assumption, because something you observed predicts missingness. Jumping straight to MNAR skips the case you can actually handle.

"I applied a MAR-based correction, so I'm covered." Only if you conditioned on the right variables. Impute using an incomplete set of predictors and you reintroduce the bias under a more sophisticated name.

"Here is the imputed result." If the mechanism was assumed rather than known, a single number hides the assumption. Report the assumption alongside the estimate.

A Worked Example: Same Table, Two Conclusions

Make it concrete. Suppose a satisfaction survey has a true underlying score YY on a 1–10 scale, and the lowest scores are the least likely to be recorded. Say the true distribution is uniform across 1–10, but scores of 1 and 2 are recorded only half the time, while scores of 3–10 are always recorded.

You observe only the recorded scores. The blanks are gone. Now compute the complete-case mean.

The recorded values are 3 through 10 with certainty, plus 1 and 2 at half weight. The observed mean is pulled upward — the low scores are underrepresented in exactly the rows you kept. Weight each recorded value by its recording probability and normalize:

Yˉobs=0.5(1)+0.5(2)+3+4+5+6+7+8+9+100.5+0.5+8=53.59≈5.94\bar{Y}_{\text{obs}} = \frac{0.5(1) + 0.5(2) + 3 + 4 + 5 + 6 + 7 + 8 + 9 + 10}{0.5 + 0.5 + 8} = \frac{53.5}{9} \approx 5.94

The truth is 5.5. The complete-case mean is biased upward by roughly 0.44 — the low scores are missing from exactly the rows you kept.

Now suppose instead the blanks were caused by a random data-entry glitch, independent of the score. The same table — same blanks, same recorded numbers — would yield an unbiased complete-case mean. The observed data are equally consistent with both stories. The one fact that would distinguish them, the unrecorded low scores, is the one fact you do not have.

The data did not change. The assumption did. The conclusion flipped.

How to Choose and Defend a Working Assumption

Start from the collection process, not the column. Ask what physically or procedurally causes a value to be absent. A dropped sample and a refused question are different mechanisms, and they point at different labels.

Then test what is testable. Run the MCAR-style checks — indicator against every observed variable. If nothing predicts missingness, MCAR is a defensible working assumption. If something does, you are in MAR territory, and you must name the conditioning set explicitly: which observed variables are you assuming explain the missingness?

When MNAR is plausible — and for self-reported outcomes like income, satisfaction, or side effects, it often is — plan a sensitivity analysis rather than a single number. Ask how much the conclusion would move under different assumptions about the missing values, and report that range.

Finally, record the assumption next to the result. Write down the mechanism you assumed, the conditioning set you used, and what would have to be true for your conclusion to hold. A later reader can then challenge your reasoning instead of just your output.

That habit is the whole discipline in one line: name the mechanism, name the conditioning set, and say what would have to be true. The labels are not descriptions of your data. They are commitments about the process that produced it — and the next step is placing that commitment inside a leakage-safe preprocessing workflow, where the imputation is fit on training data only and the assumption travels with the pipeline rather than living in your head.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

The article's worked example shows identical observed scores and blanks under two possible collection stories. What conclusion follows from that comparison?
Question 1 of 2Comparison Reasoning

Focus: Explain why an observed incomplete table can support different conclusions under different missingness assumptions.

A test finds that an observed covariate predicts the missingness indicator. What is the most defensible interpretation taught in the article?
Question 2 of 2Misconception Check

Focus: Distinguish evidence against MCAR from evidence that establishes MNAR.

References

  1. moving beyond the MCAR/MAR/MNAR classification - PMCpmc.ncbi.nlm.nih.gov
  2. 1.2 Concepts of MCAR, MAR and MNARstefvanbuuren.name
Practical resource

Build stronger machine learning foundations

Use structured resources to connect theory, scikit-learn workflows, and evaluation practice.

Browse resources
Related sites

Continue across the AI learning path

Use LearnPyFast for Python foundations and LearnLLMFast when you are ready to move from classical ML into LLM applications.

Python tutorialstutorial

LearnPyFast

Beginner-friendly Python tutorials, examples, and learning paths for practical programming foundations.

PythonProgrammingBeginners
Visit LearnPyFast
LLM tutorialstutorial

LearnLLMFast

Practical LLM tutorials for builders who want to understand prompting, workflows, agents, and AI applications.

LLMAIBuilders
Visit LearnLLMFast

Keep learning

Related machine learning tutorials

Continue with nearby concepts, model families, evaluation methods, and practical workflows.