This voyage is an editorial path for understanding, not a claim of direct historical influence or sole invention.
Understand it in one breath
"The current state contains the information needed for the distribution of the next state." Markov formalized the theory in 1906 and in 1913 analyzed vowel–consonant transitions in Pushkin. Not every chain converges. A finite irreducible, aperiodic chain converges to a unique stationary distribution regardless of its starting state. This structure appears in PageRank, speech recognition, and reinforcement learning.
At a glance
→ Next
Sunny
Cloudy
Rainy
Sunny
0.7
0.2
0.1
Cloudy
0.3
0.4
0.3
Rainy
0.2
0.3
0.5
Stationary distribution (convergence)
0.46
0.28
0.26
Repeated multiplication by the transition matrix reaches the same stationary distribution ≈ (0.46, 0.28, 0.26)regardless of the starting state — the ergodic theorem. PageRank performs exactly this calculation.
Concept
A stochastic process where the next state depends only on the present. Foundation of PageRank, GPT, speech recognition.
Key formula
P(Xn+1=j∣Xn=i,…,X0)=P(Xn+1=j∣Xn=i)
Ports in time
This concept was not invented in one instant
Follow the scenes to see problems, notation, standards of proof, and applications changing across different times and places.
1
AD 1906Scene 1 / 3Continue through the world of this year
From literature to probability — Pushkin’s verse
Andrey Markov analyzed sequences of vowels and consonants in Pushkin’s Eugene Onegin, demonstrating a probability model with dependence only on the current state.
No reliable place is given, so time continues without an invented pin