Earlier this year, I was reading Arnon et al.’s Science paper, “Whale song shows language-like statistical structure,” which reports statistical regularities in humpback whale song that resemble those found in human language, including a frequency distribution close to Zipf’s law. It brought me back to a simple question: even if we observe a Zipf law, what can we really infer about the mechanism that produced it? I had been thinking about a closely related issue while writing a short note, since published in the European Actuarial Journal, on equifinality (the idea that distinct mechanisms or trajectories can lead to the same observable outcome).
That is the starting point of Beyond Zipf’s Law: Equifinality and Mechanistic Inference from Scaling Laws. We construct several very different processes that share exactly the same Zipf marginal distribution but have distinct sequential dynamics. The exponent alone cannot distinguish them, so we look elsewhere: dependence between successive observations, transition direction, and the effect of conditioning. The broader point goes well beyond Zipf’s law: a statistical regularity can strongly constrain possible explanations without, by itself, identifying the mechanism that generated it.
Zipf-like rank–frequency scaling occurs in language, city sizes, biological data and animal communication. Because several generative processes can produce the same marginal pattern, the exponent alone has limited mechanistic content. We formulate this ambiguity as an equifinality problem and compare four constructions under a common observation design. An i.i.d. finite-Zipf process, a persistent Markov chain and canonical sample-space reduction (SSR) have the same stationary marginal, p_j=(jH_V)^{-1}, whereas a latent-scale mixture produces a similar marginal through aggregation. The first three constructions therefore isolate differences in sequence structure without changing the population rank distribution. Markov dependence changes the finite-sample distribution of fitted exponents; at moderate persistence, the shift is closely reproduced by a block-adjusted effective sample size. Excess lag-1 mutual information separates exchangeable from sequential processes, and transition direction separates reversible persistence from the directional contraction built into SSR. Conditioning on latent scale reveals the aggregation route, while fit-window, alphabet and sequence-boundary analyses show which conclusions depend on the observation design. Recent work on learned animal communication is used to formulate prospective tests rather than to validate the models empirically. Matching a scaling law is thus a compatibility condition; discriminating among mechanisms requires observations on which the candidate models differ.