Fifteen successes in twenty attempts. Before knowing anything else, it is tempting to admire the person who achieved them. After learning that the usual success rate is one half, we may begin to speak of good fortune. After learning that it is seven tenths, the same record seems considerably less exceptional.
Nothing about the observed count has changed. What changed was the story we were prepared to tell before seeing it.
That is where a statistical account of luck has to begin. An outcome does not carry its own probability distribution. We supply a population, an information set, a model, and a judgment about which outcomes matter. If those choices remain hidden, the word luck can make an explanation sound precise while leaving its most important commitments unspecified.
This is the first of three essays. I want to find a useful statistical representation of luck, test how selection and accumulated advantage distort it, and then ask what it should change in the evaluation of machine learning and AI systems. The aim is a defensible way of reasoning, rather than a number that divides a life into earned and unearned portions.
A word doing several different jobs
We use luck for a favorable surprise, for circumstances outside someone’s control, and for the consequences of taking a risk. Those meanings overlap, but they are not equivalent.
A person may inherit reliable access to education. That advantage can be highly predictable from their circumstances and still lie outside their control. A system may deliver an unexpectedly good answer because its evaluation omitted an easy subgroup. That surprise may reflect a flawed model of the experiment. A careful decision may have a bad outcome without having been a bad decision.
Frank Knight’s discussion of risk and uncertainty is useful here: a probability calculation presupposes that we can give a defensible account of the relevant possibilities. His distinction cautions against treating a unique, poorly understood situation like a familiar repeated gamble. In this series I will still express uncertain beliefs with probabilities. That is a contemporary modeling choice, accompanied by Knight’s warning about their warrant.
There is also an ethical question that probability does not settle. The opening of Thomas Nagel’s “Moral Luck” places moral assessment in tension with consequences beyond the agent’s control. I take that as a boundary for this series: describing the distribution of consequences cannot, by itself, decide what a person deserves.
To keep those questions visible, consider an outcome generated by
$$Y=g(A,X,U).$$Here $A$ denotes a specified action, $X$ the circumstances we choose to represent, and $U$ the remaining inputs. Calling all of $U$ “luck” would be premature. It may include unmeasured ability, an instrument’s error, an omitted institutional constraint, or a genuinely unpredictable event. Even the boundary between $A$ and $X$ depends on the actor and time horizon: a deployment team can change an evaluation protocol this month, but cannot change which users existed last year.
The notation makes room for contingency. It does not identify its causes.
Surprise and control are two different questions
It is tempting to call a favorable deviation from a forecast “realized statistical luck,” provided that control is documented. That definition is too broad. Documenting control does not establish that the favorable contribution lay outside it. The formula would give a precise name to an unresolved attribution.
I will instead keep two questions separate. How favorable and unexpected was the outcome under the declared prediction? And which consequential circumstances were beyond this actor’s practical control? Their answers may inform an account of luck, but neither supplies the other.
A well-resourced starting position can be predictable yet beyond the recipient’s control. A team’s deliberate improvement can surprise an observer whose forecast omitted the changed procedure. In the first case, little predictive surprise does not erase good fortune. In the second, surprise does not establish it. These classifications depend on the actor, the feasible actions, and the time horizon; control is often partial rather than binary.
Hypothetical sources of variation, not causal classifications of complete outcomes. Prediction concerns an information set; practical control concerns an actor and a time horizon. Many real cases have mixed control.
For the statistical part, declare the outcome and the prediction before observing the result. For the control part, describe the process and seek evidence about its mechanisms. I use good or bad luck for favorable or unfavorable contingencies beyond practical control, while treating the quantities below as descriptions of the prediction. This working usage does not exhaust the philosophical meanings of luck.
Measuring the predictive excess
Let $I_0$ be the information available before observation and $M$ the modeling assumptions. First declare the predictive distribution
$$Y\mid A,I_0,M \sim F_0.$$If larger values are preferable and the expectation exists, an elementary descriptive quantity is
$$D=y-\mathbb E[Y\mid A,I_0,M].$$Call $D$ the predictive excess. It answers how far the result exceeded its declared expectation, in the outcome’s units. A standardized version divides by the predictive standard deviation when that quantity is finite and nonzero. Neither quantity is a causal share of success, and a positive excess is not automatically good luck. Changing the model can change both.
A percentile answers a different question. For a discrete outcome, I use the mid-distribution rank
$$R(y)=\Pr(Y\lt y\mid A,I_0,M)+\tfrac12\Pr(Y=y\mid A,I_0,M).$$The half-weight handles ties symmetrically. It is a convention, and its distribution is not exactly uniform for discrete outcomes. An upper-tail probability $\Pr(Y\ge y\mid A,I_0,M)$ is another valid description, provided its direction was chosen beforehand. It is not the probability that luck caused the observation.
When higher values are not always better, define the utility or loss first. The same response can be fast and harmful, expensive and accurate, or acceptable to one user and unusable to another. A percentile of a convenient metric is not automatically a percentile of a desirable outcome.
The same record under three references
Suppose the twenty attempts are independent, each with a known success probability $p$. Then $K\sim\operatorname{Binomial}(20,p)$, with mean $20p$ and variance $20p(1-p)$. The upper tail is an exact finite sum:
$$\Pr(K\ge15)=\sum_{k=15}^{20}{20\choose k}p^k(1-p)^{20-k}.$$| Declared success probability | Expected successes | Observed excess at 15 | Probability of 15 or more |
|---|---|---|---|
| 0.50 | 10 | 5 | 2.07% |
| 0.60 | 12 | 3 | 12.56% |
| 0.70 | 14 | 1 | 41.64% |
The orange bars show the entire upper-tail event, K ≥ 15. The distributions are stipulated references, not estimates of people’s ability.
The arithmetic is straightforward. Choosing the reference is the difficult part. Should we compare this person with beginners, experienced practitioners, or people with comparable access to tools? Should an AI system be compared with another system under the same budget, or with the process it would replace? Those choices express the question we are asking. They cannot be recovered from the number fifteen alone.
There is a further trap. If we estimate $p$ from these same twenty outcomes and then judge how surprising those outcomes were under the fitted value, we have let the observation rewrite its own expectation. For a prospective assessment, fit on earlier evidence, preserve uncertainty, and freeze the reference before the new outcomes arrive.
The opportunities to notice a result also belong in that reference. Under $p=0.50$, a preselected group has a 2.07% chance of fifteen or more successes. Among one hundred independent groups tested under those same rules, the probability that at least one reaches that threshold is 87.65%:
$$\Pr(\text{at least one qualifying group})=1-(1-0.0206947)^{100}\approx0.8765.$$The count did not become less real. The observation process changed from following one fixed group to searching a hundred. This is the kind of denominator problem emphasized in Diaconis and Mosteller’s study of coincidences: apparent improbability depends on the opportunities and definitions that produced the noticed event. Our numerical example is its own calculation, not a result from their paper. Dependence between groups would change it.
Uncertainty about the probability changes the prediction
Suppose earlier data contained twelve successes in twenty trials. Assume that, conditional on the same fixed but unknown $p$, the earlier and later trials are independent Bernoulli draws. Under a declared uniform prior, $p\sim\operatorname{Beta}(1,1)$, the posterior is $p\mid\text{earlier data}\sim\operatorname{Beta}(13,9)$. For a new batch of twenty trials, integrate over that posterior:
$$\Pr(K=k\mid\text{earlier data})={20\choose k}\frac{B(k+13,20-k+9)}{B(13,9)}.$$The predictive probability of fifteen or more is 19.09%. Comparing it only with the earlier plug-in result, 12.56%, would mix two changes. The posterior mean is $13/22$, slightly below 0.60. We can show the intermediate calculation:
| Treatment of the probability | Mean successes in the new batch | Probability of 15 or more |
|---|---|---|
| Fix p at the earlier proportion, 0.60 | 12.00 | 12.56% |
| Fix p at the posterior mean, 13/22 | 11.82 | 10.95% |
| Integrate over the posterior Beta(13,9) | 11.82 | 19.09% |
The last two rows have the same mean. Their difference isolates the consequence of retaining posterior uncertainty rather than treating its mean as known. Their count variances are 4.83 and 8.83, respectively. More generally, for a posterior mean $m$ and variance $v$, the variance of $n$ new conditional Bernoulli trials is
$$\operatorname{Var}(K\mid\text{earlier data})=nm(1-m)+n(n-1)v.$$The second term expresses uncertainty shared across the predictions. It is not a new physical interaction between trials. Nor does wider variance imply that every chosen tail probability must increase; the increase shown here concerns this particular upper tail.
Keep the mean fixed to examine what integrating over an uncertain probability changes. The shaded event is fifteen or more successes in the new batch.
The prior is deliberately simple, not universally appropriate. The fixed-probability, conditional-independence model is an assumption about both batches. A changing population or an adaptive system can make this posterior predictive distribution inappropriate even when every calculation is correct.
When an environment is shared
Dependence creates another route to wider predictions. Imagine that each batch has its own environment, $P\sim\operatorname{Beta}(2.4,1.6)$, and the twenty trials are independent only conditional on that environment. Their marginal success probability is still $0.60$, but the shared environment induces correlation $0.20$. The count’s variance is
$$\operatorname{Var}(K)=np(1-p)\{1+(n-1)\rho\}=23.04,$$compared with 4.80 under unconditional independence. The beta-binomial mathematics resembles the previous calculation. Its interpretation differs: one model represents uncertainty about a fixed parameter; the other represents changing environments shared by a batch. A wide outcome distribution alone does not tell us which mechanism produced it.
Both models expect twelve successes. Shared exposure changes how much variation survives aggregation. Their count variances are 4.80 and 23.04.
For AI evaluation, imagine responses collected during the same outage, session, or unusually easy batch of prompts. They can share conditions that independent-trial arithmetic ignores. The example shows a mechanism worth investigating; its chosen correlation is not an estimate for any deployed system.
Does better knowledge remove luck?
If information $I_1$ refines $I_0$ under the same coherent probability model, the conditional law of total variance gives
$$\operatorname{Var}(Y\mid I_0)=\mathbb E[\operatorname{Var}(Y\mid I_1)\mid I_0]+\operatorname{Var}(\mathbb E[Y\mid I_1]\mid I_0).$$With finite second moments, more information reduces unexplained variance on average. It need not reduce it for every particular information realization. Learning that an unusually volatile situation applies can increase the variance of the relevant conditional prediction.
This is an accounting identity about information, not a physical intervention. Knowing why an advantage exists does not put it under the recipient’s control. Explaining a disparity statistically does not establish that it was earned. And a richer model can move variability from the residual into a predictor without changing anybody’s circumstances.
These distinctions matter when an explanation becomes a judgment. A regression residual is whatever the fitted specification did not account for. If the specification omits skill, its residual partly contains skill. If it omits unequal opportunity, the same happens to opportunity. Naming that residual “luck” does not repair the omissions.
What survives the objections
The strongest objection to this approach is that its answer depends on the reference distribution. I agree. That dependence is precisely what should be exposed. A claim about favorable surprise is incomplete until it states relative to what information, under what assumptions, for what purpose.
A second objection is that the control boundary is not supplied by probability theory. That is also true. We must document it through the design of the study, evidence about the decision process, and sometimes ethical argument. Statistics can test a model’s consequences; it cannot manufacture the missing account of agency.
A third objection is that a good outcome may reveal ability, rather than merely favorable variation. Nothing here denies that. The unresolved question is how much evidence one observed outcome supplies, particularly when that outcome was selected because it was the best. That is the subject of the next essay.
My conclusion is therefore deliberately limited: statistics can describe the favorability and surprise of a realization under a declared reference; interpreting it as luck additionally requires evidence about control and the process. Predictable advantages can be unchosen. Unexpected improvements can be deliberate. A residual, a rank, or a tail cannot settle that distinction.
The useful representation therefore has at least two dimensions: a conditional statistical description and an account of agency. It can discipline a prediction or an experiment. It cannot, without additional arguments, divide an achievement into percentages of skill, luck, and deservingness.
All numerical examples here are exact calculations under stipulated models. The study package contains the executable analysis, assumptions, source ledger, and a record of conclusions revised during the critical review. It contains no empirical estimate of human luck.