Beyond Zipf’s Law: Equifinality and Mechanistic Inference from Scaling Laws

Earlier this year, I was reading Arnon et al.’s Science paper, “Whale song shows language-like statistical structure,” which reports statistical regularities in humpback whale song that resemble those found in human language, including a frequency distribution close to Zipf’s law. It brought me back to a simple question: even if we observe a Zipf law, what can we really infer about the mechanism that produced it? I had been thinking about a closely related issue while writing a short note, since published in the European Actuarial Journal, on equifinality (the idea that distinct mechanisms or trajectories can lead to the same observable outcome).

That is the starting point of Beyond Zipf’s Law: Equifinality and Mechanistic Inference from Scaling Laws. We construct several very different processes that share exactly the same Zipf marginal distribution but have distinct sequential dynamics. The exponent alone cannot distinguish them, so we look elsewhere: dependence between successive observations, transition direction, and the effect of conditioning. The broader point goes well beyond Zipf’s law: a statistical regularity can strongly constrain possible explanations without, by itself, identifying the mechanism that generated it.

Zipf-like rank–frequency scaling occurs in language, city sizes, biological data and animal communication. Because several generative processes can produce the same marginal pattern, the exponent alone has limited mechanistic content. We formulate this ambiguity as an equifinality problem and compare four constructions under a common observation design. An i.i.d. finite-Zipf process, a persistent Markov chain and canonical sample-space reduction (SSR) have the same stationary marginal, p_j=(jH_V)^{-1}, whereas a latent-scale mixture produces a similar marginal through aggregation. The first three constructions therefore isolate differences in sequence structure without changing the population rank distribution. Markov dependence changes the finite-sample distribution of fitted exponents; at moderate persistence, the shift is closely reproduced by a block-adjusted effective sample size. Excess lag-1 mutual information separates exchangeable from sequential processes, and transition direction separates reversible persistence from the directional contraction built into SSR. Conditioning on latent scale reveals the aggregation route, while fit-window, alphabet and sequence-boundary analyses show which conclusions depend on the observation design. Recent work on learned animal communication is used to formulate prospective tests rather than to validate the models empirically. Matching a scaling law is thus a compatibility condition; discriminating among mechanisms requires observations on which the candidate models differ.

From Rating Factors to Crash Mechanisms: A Multiscale Causal DAG Framework Linking Motor Insurance and Road Safety

Motor insurance models and road-safety studies address the same underlying risk, but at very different scales. Insurance models typically predict annual claim counts from a small set of rating variables, such as age, mileage, vehicle characteristics, or past claims. Road-safety research, by contrast, seeks to understand mechanisms operating much closer to the crash itself, including speed, fatigue, distraction, road environment, and driving conditions. I recently uploaded on ArXiv a paper (From Rating Factors to Crash Mechanisms: A Multiscale Causal DAG Framework Linking Motor Insurance and Road Safety) that proposes a multiscale causal framework, organized around a literature-informed DAG, to connect these two perspectives. It shows that a predictive rating factor generally cannot be interpreted directly as a causal crash mechanism. Examples based on mileage and driver age illustrate why: the same actuarial contrast may be compatible with many different mechanistic explanations. Sharper conclusions will require data closer to the driving process, such as telematics, trip context, or linked crash and insurance-claim records…