Tag Archives: risk

Open Review: Risk Priced Society

After more than ten years of writing blog posts and essays for a broader audience, I used this past year to revisit, update, and bring some of those ideas together. The result is a short, non-technical book, Risk-Priced Society — Insurance, Algorithms and Climate in the Age of Prediction.

As data become more granular and models more powerful, insurers can distinguish more precisely between people, behaviours, and places. But better prediction does not tell us how losses should be shared, which differences should affect prices, or what should remain pooled. Drawing on the history of insurance, actuarial practice, decision algorithms, and climate risk, the book argues that technical precision never eliminates political choice. It can support prevention and strengthen protection, but it can also individualise vulnerabilities that are collectively produced. The challenge, then, is not to reject models, but to decide democratically what they should be allowed to do, what consequences should follow from them, and how to avoid leaving each person alone with their risk.

The book is open for review until December. Comments, criticisms, and suggestions are all very welcome. You just need an hypothes.is account. Select a part you want to annotate,

and then add all comments you want to make

International Workshop on Risk and Insurance, 서울, June 2026

In three weeks, the International Workshop on Risk and Insurance will be organized in 서울 (Seoul, Korea), prior to the Insurance: Mathematics & Economics conference. We were asked to send our slides in advance, so I thought I could be nice to share them (slides are available here). I will be the very first speaker of the day, in a session “AI Revolution and Cyber Risks”.

Over the past ten years, we have heard a lot about the AI revolution in actuarial science. Very often, the narrative is by technological optimism: more data, more models, more automation, and therefore better pricing, better decisions, and better insurance. In this talk, I plan to take a step back. Not to deny the importance of AI, of course, but to return to a more fundamental actuarial question: what makes a risk insurable? I will take two perspectives: first, information granularity and mutualization; second, AI itself as an insured exposure. The first part will be about granularity and mutualization. Insurance transforms individual heterogeneity into risk classes, but as information becomes more granular, the boundary between what is pooled and what is individualized shifts. I will move from information asymmetry to information granularity, then to natural catastrophe maps, and finally to actuarial fairness. In the second part, I will turn to AI risk in insurance: not AI as a tool for insurers, but AI as a risk that insurers may have to cover.

Let me start with the first part: granularity and mutualization. Insurance relies on a form of shared ignorance. We do not know who will have an accident, who will fall ill, or which house will be flooded. This uncertainty makes pooling possible. But ignorance is never complete. Actuaries observe variables, build classes, estimate frequencies, and price risks. The history of pricing is therefore a history of compromise between measuring heterogeneity and preserving mutualization. This slide gives a very classical way of seeing the problem. If there is a latent risk factor, which I call theta, total variance can be decomposed into within-risk variance and between-risk variance. The first part is the residual uncertainty that remains pooled. The second part is systematic heterogeneity across individuals or groups. But in practice, insurers do not observe theta. They observe covariates: age, location, car type, claim history, scores, and so on. These variables move part of the uncertainty from the insurer to the policyholder. The last formula is the key one. It separates irreducible risk from epistemic uncertainty, or misclassification. More granular information reduces misclassification, but it also turns pooled uncertainty into priced heterogeneity.

The classical issue in insurance economics is information asymmetry.  In Akerlof, or in Rothschild and Stiglitz, the problem is that policyholders may know more about their risk than insurers. If high-risk individuals buy more coverage than low-risk individuals, pooling contracts may become unstable. Empirically, this leads to a simple question: after controlling for observable risk characteristics, do coverage choices still predict losses? The genetic-testing debate makes this very concrete. Individuals may know something about their future risk that insurers cannot observe, or are not allowed to use. Traditionally, the fear was that policyholders knew too much. Today, with big data, we may also ask what happens when insurers know almost too much.

With granularity, the problem changes. It is no longer only about asymmetric information between the insured and the insurer. It is about reducing epistemic uncertainty. Things that were previously hidden become partially observable: behaviour, location, telematics, digital traces, health signals, and social proxies. Statistically, this is attractive: prediction improves. But actuarially and socially, it is more ambiguous. Better prediction also redistributes premiums, makes some risks more visible, and may turn solidarity into sorting. Removing protected variables is not enough. In rich data environments, proxies can reconstruct what we claim to exclude. So the question becomes normative: what should be priced, what should be pooled, and what should be used only for prevention? Note that the same issue arises with past criminal records: even if the information is predictive, should it remain admissible forever? Inclusive insurance requires a temporal dimension of data governance. Some information may be relevant at one point, but should progressively lose its underwriting legitimacy.

Natural catastrophes make this problem very concrete. For a long time, some climate risks were pooled at a broad scale. But risk maps, climate models, and fine geographic data now identify areas that are much more exposed than others. Even when a system is formally solidaristic, access to coverage may depend on the underlying household insurance contract. If insurers withdraw locally, or impose stricter conditions, formal solidarity may remain while effective access to insurance declines. So granularity does not only produce more accurate prices. It can also produce non-insurance, selective underwriting, or local market withdrawal.

This leads naturally to actuarial fairness. We often say that equal risks should pay equal premiums. But what does “equal risks” mean? Equal given all available information? Equal within a tariff class? Equal for a score? Equal within a protected group? Actuarial fairness is always relative to an information set. More granular information gives a more individualized notion of equality. Coarser information preserves more mutualization. This creates a tension between actuarial adequacy, solidarity, and causal legitimacy. A variable may be predictive without being socially acceptable. Conversely, excluding a variable may preserve solidarity, but at the cost of cross-subsidies and possibly selection.

I now move to the second part: AI risk in insurance. The link with the first part is this: so far, I have discussed AI and data as instruments of classification. But AI is also becoming a source of risk. Firms use models to code, advise, filter, recommend, automate, and decide. So the question is no longer only: how do insurers use AI? It is also: how do we insure organizations that use AI?

There is already a large discussion about AI for insurance: pricing, fraud detection, claims management, customer service. I want to reverse the perspective. AI is becoming an insured exposure. AI-related losses may not appear in a new line called “AI insurance.” They will appear through existing lines: cyber, technology E&O, professional liability, product liability, D&O, and regulatory risk. What is new is the combination of opacity, speed, scale, version changes, dependence on external providers, and difficult ex-post attribution. The actuarial question is not simply: what is the historical frequency of AI failures? The question is whether the risk can be bounded, observed, priced, monitored, and evidenced.

The first insurance problem is accumulation. A portfolio may look diversified: different firms, sectors, and countries. But they may all rely on the same models, APIs, cloud providers, datasets, retrieval systems, or agent frameworks. A model bug, a bad update, a prompt-injection vulnerability, a compromised API, or a provider failure can generate correlated losses across many insureds. This weakens the usual law-of-large-numbers intuition. Risks are not independent if they share the same technological infrastructure. Underwriting AI therefore requires mapping the AI supply chain: provider, model version, API, data, RAG system, agents, tools, and update process.

This slide introduces a recent report on AI risk prioritization. I would not take the probabilities literally; it is a study based on expert judgment. But it is useful because it frames AI risks as systemic, cross-sectoral, and unevenly distributed. The key point for insurance is the mismatch between vulnerability and control. The actors most exposed to harm are not necessarily those who control the models. Users, clients, and affected third parties may bear the consequences, while mitigation often depends on developers, providers, regulators, governments, or standards bodies. This raises a classical insurance question: who bears the risk, who controls the risk, and who can produce evidence after a loss?

Large language models (LLMs) are a particularly interesting case. The issue is not only hallucination. It also includes retrieval errors, prompt injection, privacy leakage, biased outputs, insecure code, incorrect tool use, and automation of decisions that should remain supervised. Benchmarks are not enough. They evaluate controlled tasks. Insurance deals with deployed systems, real users, real incentives, changing environments, and downstream losses. LLMs shift the cost of reliability. Producing an answer becomes cheap; verifying it may become expensive. And “human-in-the-loop” is not a magic control. A human can only control the system if they have time, competence, authority, and incentives to challenge the machine.

This brings me to a central point: insurability depends on proof. An insurance policy cannot simply say: “we cover AI risk.” It must define the insured system: the use case, model, version, provider, data source, retrieval layer, tools, agents, and update regime. It must also define evidence obligations: logs, audits, monitoring, incident reporting, version retention, red teaming, escalation procedures, and human review. Insurance can play a disciplinary role here. If nobody pays for proof before the loss, everybody pays for uncertainty after the loss. Conversely, if coverage depends on traceability, insurance creates an economic incentive to document and control AI systems.

I will end with vibe coding. Lawrence Lessig famously said that “code is law”, twenty five years ago: code does not only execute rules; it structures what is possible, impossible, easy, costly, visible, or invisible. But if code is generated by AI from prompts, modified quickly, integrated by agents, and sometimes deployed without serious human review, a new question appears: who can reconstruct the history of the code after a loss? The risk is not only that AI-generated code may be wrong. The risk is that nobody can say exactly what was asked, what was generated, which model version was used, which dependencies were added, which tests were run, and how the code reached production. For insurers, the right question is not: “Do you use AI to code?” Almost everyone will. The real question is: “Through which process can AI-generated code enter a critical system?” And this brings us back to very concrete underwriting tools: prompt logs, model versions, human review, tests, vulnerability scanners, SBOMs, deployment rules, and exclusions for unverified autonomous deployment. The broader conclusion is that AI does not eliminate classical actuarial questions. It makes them more visible: mutualization, accumulation, information asymmetry, evidence, moral hazard, control, and responsibility.

Bridging the Risk Perception Gap ?

Cet été, en rédigeant un article pour The Conversation sur le coût des assurances, on avait été surpris par une remarque des éditeurs qui nous reprochait de citer davantage de rapports professionnels que d’articles académiques. Dans beaucoup de domaines, cette critique serait justifiée. Mais en assurance, ces rapports, ou brochures, rédigées par les réassureurs, les courtiers ou les agences spécialisées constituent souvent les sources les plus fiables, parfois les seules, pour accéder à des chiffres crédibles, à des données récentes et à une compréhension granulaire du fonctionnement réel du marché. Elles ne sont pas neutres, bien sûr ; elles diffusent un récit. Mais en les lisant “en diagonale”, ou plutôt à contre-jour, elles révèlent bien plus que ce qu’elles prétendent dire : elles donnent à voir comment les acteurs redéfinissent les responsabilités, les périmètres de risque et les équilibres économiques du secteur.

Prenons le rapport Bridging the Risk Perception Gap publié par Howden Re il y a quelques jours. On y des discussions techniques, centrées sur la modélisation, l’incertitude, ou la disponibilité des données. Mais à mesure qu’on avance dans sa lecture, une impression se dégage, et on a l’impression que ce qui est présenté comme un “écart de perception” relève peut être moins d’un malentendu technique que d’une redéfinition du périmètre du risque et de sa répartition entre acteurs. Dès l’introduction, on nous explique que les assureurs directs ont absorbé une part croissante des pertes catastrophes naturelles, passant de 54 % en 2022 à 67 % en 2024 . Cette évolution est présentée comme une conséquence du “hard market” et d’un désalignement technique. Pourtant, la littérature académique rappelle depuis longtemps que la répartition du risque n’est jamais un phénomène neutre ni mécanique. François Ewald, dans L’Etat providence, rappelait que l’assurance repose sur une “politique du risque”, une construction institutionnelle qui organise la solidarité et délimite les responsabilités. C’est précisément cette dimension institutionnelle que le rapport tend à invisibiliser en renvoyant la divergence à un simple déficit de données.

“Améliorer la modélisation” ou “reconfigurer les responsabilités”

Un fil conducteur traverse le document : le marché souffrirait d’une “insuffisance de données” et d’un “manque de robustesse des modèles”, en particulier pour les périls de fréquence. On lit par exemple que

« the lack of robust data and modelling capabilities has made reinsurers hesitant to re-engage » (p. 4) .

Cette explication, répétée, suggère que la divergence de perception du risque tient à un déficit d’information, que la modélisation permettrait progressivement de combler. Or les travaux de Paul Slovic montrent précisément l’inverse. Dans The Perception of Risk, il insiste sur le fait que les divergences entre “experts” et “non-experts” (ou, ici, entre réassureurs et cédantes) ne sont pas réductibles à la qualité des données. Elles reflètent des positions institutionnelles différentes, des régimes d’incitation distincts, et des rapports différenciés au contrôle et à la responsabilité. J’ai l’impression que le rapport Bridging the Risk Perception Gap illustre cette dynamique. Ce qu’il présente comme un “écart de perception” me semble relever plutôt d’une reconfiguration de ce que les réassureurs souhaitent, ou acceptent, d’assumer, en particulier sur les couches basses et moyennes. L’idée selon laquelle une modélisation plus fine suffirait à “réconcilier” les deux côtés du marché masque ainsi une réalité structurelle : la technique n’est pas seulement un outil de description, mais un outil de redéfinition du périmètre du risque.

La proposition de recourir davantage aux couvertures paramétriques en est un bon exemple. En faisant reposer le déclenchement d’indemnisation sur des indices exogènes, ces solutions redéfinissent ce qui constitue un événement indemnisable. Elles n’objectivent pas le risque : elles le reconfigurent. Le rapport reconnaît d’ailleurs très explicitement que cette approche

« shifts the inflation risk back to the cedant » (p. 6) .

La modélisation devient ici l’instrument d’un déplacement de responsabilité. Cette articulation entre technique et gouvernance du risque a été analysée en profondeur dans Modern risk management : a history , où Peter Field rappelle que les modèles ne se contentent pas de représenter les risques. Le rapport Bridging the Risk Perception Gap s’inscrit dans cette logique, les solutions techniques qu’il met en avant redéfinissant le risque autant qu’elles prétendent en améliorer la mesure.

Climat : la causalité commode

Le rapport attribue une partie du désalignement technique à l’évolution du climat, affirmant que

« climate change is driving an increase in secondary perils » (p. 4) .

Cette explication, quoique plausible sur certains aléas, est scientifiquement partielle. Les travaux de l’OCDE (en particulier Risk Awareness, Capital Markets and Catastrophic Risks) rappellent que l’essentiel de la croissance des pertes tient non à l’aléa, mais à l’exposition et à la vulnérabilité des actifs. Mettre la responsabilité sur les aléas secondaires permet de présenter l’incertitude comme exogène : le problème viendrait du monde, non de l’organisation du marché, des arbitrages d’accumulation, ou du retrait progressif des réassureurs des couches de fréquence. On retrouve ici une dynamique décrite par Barbara Hudson dans Justice in the Risk Society. Le climat devient un argument commode pour légitimer des modifications de périmètre, alors même que la montée des pertes s’explique d’abord par la structure collective d’exposition.

Un écart de perception, ou une réallocation 

À mesure que l’on avance dans le rapport, l’objet réel devient plus clair : il ne s’agit pas de réduire un écart cognitif, mais de stabiliser une nouvelle frontière dans la répartition du risque. La distinction finale proposée entre

brute-force discipline” et “technical discipline” (p. 9)

va dans ce sens : la période récente, marquée par un resserrement de l’appétit de risque, serait désormais appelée à se transformer en une discipline technique, plus stable, plus prédictible. Autrement dit, une normalisation du retrait. L’idée même d’un “risk perception gap” repose sur l’hypothèse implicite que les deux parties cherchent à mesurer la même chose, alors que leurs positions institutionnelles (et les incitations qui en découlent) divergent profondément.

Le biais paramétrique

Le rapport insiste à plusieurs reprises sur la capacité des couvertures paramétriques à résoudre ce que les auteurs perçoivent comme la principale source du “risk perception gap” : l’incertitude, en particulier celle liée à l’inflation des sinistres non dommages. Il affirme ainsi que les structures paramétriques peuvent fournir une

« undisputable evidence of the inflation held in the subject portfolio » (p. 6)

permettant un règlement automatique, transparent, sans dépendre des processus d’expertise. Ce langage est très assurant (évoquant “preuve indiscutable”, “clarification”, “discipline technique”). Mais la littérature scientifique montrent que ces promesses reposent sur une ambivalence profonde, rarement reconnue dans les documents commerciaux, et que la technique paramétrique ne réduit pas l’incertitude, elle la reconfigure. Dans l’introduction, présentant le chapitre sur l’assurance paramétrique dans The Global Insurance Market and Change, Anthony Tarr, Julie-Anne Tarr, Maurice Thompson et Dino Wilkinson expliquent que

« A parametric contract is an agreement to make a payment upon the occurrence of a triggering event and as such is detached from loss or damage to an underlying physical asset or piece of infrastructure »

autrement dit, l’objet assuré n’est plus la perte, mais le signal, la valeur de l’indice, la condition observable, certes idéalement très corrélée à la perte. Mais seulement corrélée. De fait, la reconceptualisation de ce qui constitue un événement indemnisable (non plus un dommage mais une condition statistique) transfère mécaniquement toutes les situations où le dommage existe mais l’indice ne déclenche pas vers le cédant (basis risk). Robert Jerry, dans Understanding Parametric Insurance: A Potential Tool to Help Manage Pandemic Risk, montre que cette dissociation ouvre un espace considérable pour les réassureurs, permettant de contrôler avec précision, en amont, les situations indemnisables, en limitant les incertitudes liées au comportement des assurés, aux variations de coûts, ou aux débats d’expertise. Dans la littérature académique, cette capacité de pré-cadrage est décrite comme l’un des principaux attraits du paramétrique pour les assureurs et réassureurs. Mais elle génère mécaniquement un risque systémique pour l’assuré ou la cédante, celui de subir un sinistre réel sans activation du trigger.

La littérature économique sur l’indice (notamment Enrico Biffis, Erik Chavez, Alexis Louaas & Pierre Picard, dans Parametric insurance and technology adoption) souligne que ce basis risk constitue la contrainte structurelle fondamentale du paramétrique. Mais le rapport  Bridging the Risk Perception Gap n’évoque jamais cette dissymétrie. Or c’est précisément elle qui explique pourquoi le paramétrique, loin de résoudre le “gap”, le déplace, du domaine technique vers le domaine politique, c’est-à-dire celui de la définition des responsabilités. La littérature comparée confirme ce point. Le paramétrique est un outil d’allocation financière, pas un outil de mesure fidèle du risque. C’est exactement ce que montre aussi Robert Jerry, dans Understanding Parametric Insurance, PathogenRX (Marsh/Munich Re) n’a quasiment trouvé aucun preneur avant Covid-19, non à cause d’une incompréhension technique, mais parce que les incitations économiques des assurés et des assureurs divergeaient fondamentalement. Ainsi, lorsque Howden Re présente les solutions paramétriques comme un moyen de réduire la divergence de perception du risque, il faut comprendre que le paramétrique clarifie la responsabilité du réassureur, mais en réassignant aux cédantes, via le basis risk implicite, ce qui ne relève pas de la structure de l’indice. La technique, ici, ne vient pas corriger un malentendu, elle vient formaliser une nouvelle frontière dans la répartition du risque.

La construction d’un récit 

Ce type de narration est bien documenté dans la littérature. Dans le chapitre sur l’assurance paramétrique, publié dans The Global Insurance Market and Change, Wynne Lawrence, Julie-Anne Tarr, Nigel Brook, Meg Chaperon et Arnaud Sorel rappellent que le paramétrique est souvent présenté comme une innovation comblant un manque,

« At a time when cash flow was at a premium for pandemic-impacted businesses, the insurance industry struggled to respond. Settlement periods could be as long as many months to far more extended periods as insurers struggled to manage dramatic escalations in claims, interpretation of widely
divergent policy wording around the basis for indemnity and exclusions, and protracted litigation processes for clarification and settlement. Even in cases of uncontested liability, indemnity policies’ requirements around causation and resulting loss verifications can be time, labour and cost-intensive. »

La recherche d’une causel peut être trop longue, une corrélation pourrait nous sauver… La rhétorique est tournée vers la rapidité des paiements, la transparence des triggers, la simplicité contractuelle, ce qui construit une figure de l’assureur technique dont l’expertise dépasse celles des autres acteurs. Une figure que Howden Re reprend dans son rapport. Sans théoriser frontalement la non-neutralité de la technique, Jerry montre que l’assurance paramétrique dépend de choix d’indices et de données traversés par des contraintes institutionnelles et politiques (jusqu’à la manipulation possible des chiffres), et que ces choix ont des effets distributifs : le trigger remplace la perte par une mesure, ce qui crée du basis risk et peut fragiliser la mutualisation, voire accroître biais et discriminations.

Le caractère “innovant” ne relève donc pas d’une supériorité technique, mais d’une capacité à redéfinir les termes du contrat assurantiel. Le rapport Parametric Insurance and Natural Disaster Relief Reform (en chinois, mais qu’on peut traduire en français avec un outil en ligne, au moins l’introduction – je ferais un billet un jour sur l’importance d’aller chercher des documents en chinois, en coréen, en japonais, pour comprendre vraiment ce qui se passe en Asie) montre que le système chinois de secours en cas de catastrophes naturelles repose largement sur des mécanismes d’aide ex post, fortement dépendants du budget public, et insuffisamment dotés de sources de financement stables et prévisibles. L’assurance paramétrique n’a pas vocation à transformer la nature de l’assurance, mais à compenser les limites structurelles des mécanismes existants de secours, en fournissant aux finances publiques des ressources rapides et prévisibles après les catastrophes. On retrouve cette idée dans les exemples d’assurance paramétriques agricole exploitée via la microfinance qui montrent que ces dispositifs servent souvent des objectifs publics ou quasi-publics (rapidité de décaissement, stabilisation budgétaire, gestion de crise) beaucoup plus que de réelles avancées dans la modélisation actuarielle.

C’est ce que la littérature en risk management décrit comme la fonction narrative de l’innovation, elle permet de stabiliser des décisions économiques en les rendant socialement acceptables. Les réassureurs ne se contentent pas de proposer des solutions, ils produisent un récit dans lequel ces solutions deviennent nécessaires, rationnelles, presque inéluctables. Lorsque Howden Re oppose la “brute-force discipline” du hard market à une “technical discipline” fondée sur l’innovation (comme je l’évoquais plus haut), il reformule un retrait comme un progrès. C’est exactement ce que décrit Modern risk management : a history, au cours des différents chapitres, montrant que les modèles, en finance comme en assurance, servent autant à légitimer des stratégies de marché qu’à décrire des réalités. Dans cette perspective, la dimension héroïque du récit n’est pas un ornement rhétorique, mais un élément fonctionnel, qui sert à produire de la confiance dans un moment où les réassureurs, ayant restreint leur appétit pour les couches basses, doivent justifier leur repositionnement. En ce sens, les cas mis en avant dans ces rapports ne sont pas de simples illustrations techniques, ce sont des dispositifs de récit qui consolident une nouvelle grammaire du risque, dans laquelle l’innovation technologique n’est ni neutre ni pure, mais un instrument de gouvernance.

Insurers and AI, a systemic risk

In Insurers retreat from AI cover as risk of multibillion-dollar claims mounts, the Financial Times reported at the end of this week that several major insurers (AIG, Great American, and WR Berkley) are seeking to introduce explicit exclusions for risks related to artificial intelligence, particularly concerning the use of agents and language models. The reasons put forward are straightforward: potential losses related to AI could reach several hundreds of millions of dollars, or even more. But above all, the danger lies less in the severity of any individual claim than in the possibility of correlated, massive, simultaneous losses, impossible to mutualize. As Aon summarizes in one sentence:

“What [the industry] can’t afford is if an AI provider makes a mistake that ends up as a 1,000 or 10,000 losses.”

The issue is far less about the severity of a particular incident than about the simultaneity of thousands of identical incidents. In other words, we are dealing with the very definition of systemic risk. To understand this, it is useful to place the phenomenon within the broader framework of complex systems, accidents in high-reliability organizations, and system-wide risk mechanisms as studied for several decades in finance and insurance.

AI as an interconnected network: a breeding ground for contagion

The literature on systemic risk teaches us that it is not the absolute size of institutions that determines their vulnerability, but the structure of their interconnections. Prasanna Gai highlights that financial systems exhibit what he calls a “robust-yet-fragile” dynamic: they withstand countless shocks yet may collapse abruptly when a specific shock travels through the right channels. He reminds us that:

“The system may be robust to most shocks, but when problems strike the effects may be catastrophic.”

A local error can turn into a global catastrophe as soon as it reaches the network’s “vulnerable cluster.” The simulations he presents show that the more interconnected a network is, the faster an error can spread. A minimal change in connectivity or capitalization can shift the entire system from a stable to a critical state, a true “phase transition.” As Gai puts it:

“An initial error can spread through the entire vulnerable cluster.”

In this context, generative AI displays all the characteristics of a system highly conducive to contagion. When an AI provider deploys a faulty update, introduces an error in a model’s parameters, or suffers a cybersecurity vulnerability, it is not isolated users who are affected but thousands, because they rely on the same infrastructure. Each client depends not only on its own usage but on the integrity of a global model, where even the smallest modification instantly reproduces identical behaviour across all users. This is exactly the kind of rapid, synchronized contagion insurers now fear: propagation that is not slow or progressive, but immediate, simultaneous, and homogeneous.

Why insurers fear correlated losses

Insurability has historically depended on a fundamental condition: the law of large numbers. As Hufeld, Koijen, and Thimann remind us in The Economics, Regulation, and Systemic Risk of Insurance:

“The risk must obey the law of large numbers.”

Events must be independent, or sufficiently heterogeneous, for losses to offset each other statistically. Yet cyber-risks already fail to meet this condition, as Biener et al. showed in 2015:

“Cyber losses are highly interrelated and characterized by severe information asymmetries.”

Cyber insurance is already structurally fragile, facing simultaneous, large-scale, difficult-to-attribute incidents. Generative AI only reinforces this structure, creating an environment where errors are not merely frequent but potentially identical and simultaneous. The examples cited by the Financial Times (an Air Canada chatbot binding the airline to a fare, a Google module defaming a small business, a $25 million deepfake fraud against Arup, or an autonomous agent producing pricing or diagnostic errors at scale) highlight how extremely reproducible such incidents are. A single defect, update, or vulnerability can affect an entire sector all at once.

This homogeneity of losses is toxic for insurance. As actuarial literature reminds us, when losses become correlated, mutualization collapses mechanically. By design, insurance cannot absorb risks whose very structure pushes toward aggregation.

AI, cybersecurity, and “normal accidents”: when complexity makes error inevitable

Clearfield and Tilcsik, in Meltdown, show that in complex systems, failures are not anomalies: they are inevitable. Surprises emerge from interactions, and small errors amplify as they propagate through tightly coupled networks. Their argument corresponds exactly to what insurers fear: the error of an autonomous model is not an isolated human mistake, but a systemic mechanism liable to replicate everywhere.

“In wicked environments, error signals are ambiguous, interactions are complex, and systems become vulnerable to surprises.”

Generative AI models fit this description perfectly: structural opacity, non-deterministic behaviour, dependence on a handful of global providers, and the lack of separation between uses create a system in which local failures become systemic. The technological dependence concentrated among a few actors amplifies this dynamic further. When a single model is deployed in legal, medical, financial, and industrial contexts, an error propagating through that model can contaminate every domain simultaneously. AI is thus less a tool than an interconnected ecosystem, a fully-fledged complex system.

Why AI creates a new kind of systemic risk for insurance

Traditionally, insurance has been considered less vulnerable to systemic risk than banking, because insurers exhibit far lower interconnectedness. This is the consensus highlighted by Hufeld, Koijen, and Thimann:

“The insurance industry is not subject to systemic risk, in particular due to its significantly lower interconnectedness.”

But this assumption relies on a hidden premise: that the risks insured must themselves remain independent. With AI, that premise collapses. For the first time, an insured risk (cyber, errors & omissions, software-related damages) becomes structurally interconnected. A single provider, a single model, a single update, or a single vulnerability can trigger thousands of simultaneous losses. Interconnectedness is no longer a property of the insurance market, it is a property of the risk itself. Acharya, Pedersen, Philippon, and Richardson define a risk as systemic when it can generate a simultaneous “capital shortfall” across multiple institutions:

“Even if liabilities are not runnable, a firm can contribute to systemic risk through its contribution to the aggregate capital shortfall.”

In a world where thousands of losses linked to the same AI error could occur at once, this definition becomes sharply relevant. AI introduces a form of aggregated, correlated, non-diversifiable risk. It is no longer a volatile risk, but a structurally synchronized one — insurers’ worst-case scenario.

What the “big systemic events” of AI might look like

Several scenarios illustrate how AI-driven systemic risk might unfold. Imagine a faulty update to an LLM used across the financial sector: a model deployed in two thousand banks simultaneously misinterprets a regulatory rule. The consequences (non-compliance, sanctions, lawsuits, customer withdrawals, class actions) would be immediate and perfectly synchronized. An autonomous legal agent could also generate systematic hallucinations, producing false legal citations or flawed reasoning. If deployed across several hundred companies, that error would instantly become collective. Beyond these operational scenarios lies a subtler phenomenon: the breakdown of interpretability. When models produce signals that seem plausible but remain opaque, both humans and machines tend to attribute excessive meaning to them. A weak signal — a slight increase in churn score, a small rise in a risk indicator — may be interpreted as a real behavioural shift, though it may merely be a statistical artifact, dataset bias, or latent drift.

These misinterpretations can create feedback loops, turning noise into a real shock: emergency decisions, pricing changes, adjustments to dependent models. Clearfield and Tilcsik describe this as a self-fulfilling prophecy. In Gai’s terminology, an initial noise finds its way into the vulnerable cluster and triggers a cascade:

“With positive probability, a random initial default at one bank can lead to the spread of default across the entire vulnerable portion of the financial system (…) Contagion breaks out when shocks strike any bank in, or adjacent to, the giant vulnerable cluster and spread across the entire cluster.”

For insurers, such risk (intrinsic to the system’s functioning) becomes impossible to mutualize.

The Tesla case study

The book The Tesla Files offers a striking illustration of these dynamics. It reveals an organization where extremely concentrated critical functions make any incident capable of propagating system-wide. A single administrator holds global access; thousands of employees have elevated privileges; and the whistleblower describes an absence of monitoring despite massive data extractions. In such a context, an incident does not remain local, it compromises the entire organization.

“The stream of data expands to include customers, business partners, and a wide array of individuals and companies with ties to Tesla.”

Tesla aggregates not only its own data but also data from customers, partners, governments, subcontractors, and regulators — creating dependency structures extremely similar to pre-2008 financial networks. The authors show that Tesla minimizes some risk indicators (recording only crashes involving airbag deployment), sometimes disables Autopilot just before impact (obscuring responsibility), or refuses to transmit data to regulators. Four elements appear: information asymmetry, opaque accountability, centralized incident handling, and extreme software homogeneity.

Autopilot crashes in Europe are handled by a single team, and thousands of customer complaints exist regarding phantom braking or sudden acceleration. When an entire fleet relies on a single software model updated simultaneously, a defect can produce a massive correlated shock. This is exactly the scenario insurers now fear for generative AI systems.

Tesla thus serves as a concrete example: a system where software homogeneity creates such aggregation risk that a single error can become a “big systemic event.”

AI Liability: a legal blind spot in an asymmetric market

To these operational risks we must add a final dimension: legal responsibility. “AI liability” )the question of who is responsible when AI causes harm) is today one of the least explored and most explosive issues. The Financial Times rightly notes:

“Nobody knows who’s liable if things go wrong.”

In practice, AI providers’ contracts include drastic limitations of liability, exclusions of performance guarantees, and clauses transferring almost all risk to the user. Because demand is nearly inelastic (there are often no viable alternatives to the major models), providers can impose their terms unilaterally, creating a significant contractual asymmetry.

This asymmetry becomes critical in regulated sectors, where firms are required to control their models: banks, insurers, and hospitals must comply with strict requirements for operational and model-risk management, even though the models they use are opaque, external, and unaudited. The Tesla Files shows that even technologically advanced companies can lose control of their own systems, making some regulatory obligations nearly impossible to fulfill. The financial sector thus faces a profound contradiction: it is legally responsible for tools over which it has no meaningful control — not the design, not the training data, not the governance, not the behaviour. Firms cannot contest provider errors nor obtain compensation when a model fails.

This situation creates a triple gap: a regulatory gap (firms cannot meet their obligations), a contractual gap (all responsibility lies with the user), and an insurance gap (insurers are withdrawing from the field). The result is a legal systemic risk, characterized by diffuse responsibility, concentrated dependency, and profoundly inefficient allocation of risk.

References

Workshop “Networks, Games and Risk”

Monday we had our workshop “Networks, Games and Risk“,

with Renaud Bourles (Centrale Marseille, Aix-Marseille School of Economics, and Institut Universitaire de France, France), Vincent Boucher (Université Laval, Québec, Canada), Federico Bobbio (Université de Montréal, Montréal, Canada), Leonie Baumann (McGill University, Montréal, Canada), Fallou Niakh (CREST, ENSAE, Institut Polytechnique de Paris, France) and Philipp Ratz (Université du Québec à Montréal, Montréal, Canada)

It was extremely interesting !

Networks, Games and Risk

On Monday, December 18th, 2023, we organize, at UQAM, a workshop on “Networks, Games and Risk“.

Decentralized risk-sharing markets are markets for risk exchange in which a pool of individuals agree to mutually insurer each other, without recourse to a centralized insurance provider. Some important problems to examine in these markets are the following:

  • The coalitional stability of the pool, or the formation of risk-sharing networks (subcoalitions) within the pool.
  • The Pareto-efficiency of allocations along risk-sharing networks.
  • The structure of allocation mechanisms within networks, that is, the mappings that transform feasible allocations into other feasible allocations within networks.

Examining these problems requires an interdisciplinary approach, drawing from economic theory, insurance and actuarial science, game theory, and related fields of applications. It is the aim of this workshop to bring together researchers from various fields, to discuss open problems in the theory of decentralized risk-sharing along networks, as well as potential interdisciplinary approaches to tackle these problems.

Do risk classes go beyond stereotypes?

Generalization, stereotypes and clichés

In Thinking, Fast and Slow, Daniel Kahneman discusses at length the importance of stereotypes in understanding many decision-making processes. A so-called System 1 is used for quick decision-making: it allows us to recognize people and objects, helps us focus our attention, and encourages us to fear spiders. It is based on knowledge stored in memory and accessible without intention, and without effort. It can be contrasted with System 2, which allows for more complex decision-making, requiring discipline and sequential reflection. In the first case, our brain uses the stereotypes that govern judgments of representativeness, and uses this heuristic to make decisions. If I cook a fish for friends who have come to eat, I open a bottle of white wine. The cliché “fish goes well with white wine” allows me to make a decision quickly, without having to think about it. Stereotypes are statements about a group that are accepted (at least provisionally) as facts about each member. Whether correct or not, stereotypes are the basic tools for thinking about categories in System 1. But in many cases, a more in-depth, more sophisticated reflection – corresponding to System 2 – will make it possible to make a more judicious, even optimal decision. Without choosing any red wine, a pinot noir could perhaps also be perfectly suitable for roasted red mullets. “To generalize is to be an idiot, to particularize is the alone distinction of merit” wrote William Blake around 1800, annotating speeches by the painter Joshua Reynolds. Stigmatizing an entire population because of a minority in a decision-making process is a misleading generalization, often punished by society. Moral punishment, but sometimes also legal (when hiring for example) in a society that tends to be civilized, asking not to draw erroneous conclusions about an individual from the statistics of a group to which he would be attached. But isn’t that what the actuary does every day?

The usual suspects

For Schauer (2009), this “generalization“, condemned by William Blake, is probably the actuary’s raison d’être: “to be an actuary is to be a specialist in generalization, and actuaries engage in a form of decision-making that is sometimes called actuarial“. If I decide to insure a sports car, I have I am given risky driving characteristics that probably belong to the majority of sports car owners, attributes that I may not share. And as we noted in the introduction, insurance companies, of course, are not the only ones that operate actuarially, according to Schauer’s definition. We all do it, much more often than most of us would probably recognize. We do this when we choose airlines based on their safety record, punctuality or lost luggage. We do this when we associate personal characteristics (a visible tattoo, black or brightly coloured clothing) with behavioural characteristics (such as a propensity for violence) that these personal characteristics would seem to indicate. And we operate in this way when we engage in stereotypes that may be harmless on the basis of nationality, for example by calling French people are rude, or Scots all wear kilts, while at the same time acknowledging that more pernicious stereotypes on the basis of ethnic origin, gender, sexual orientation are too widespread today! As the misconception of the word “prejudice” indicates, many people believe that it is unfair to make individual decisions based on non-universal group characteristics. Even if group allocations have a solid statistical basis. Because the big difference between actuarial science and everyday life is that actuaries have to use a large number of observations. On a personal level, I can thus decide not to travel with such an airline anymore because on three trips, I have experienced two bad experiences. Before deciding that travel insurance deserves a higher premium when flying with this company, it takes more than three observations! In fact, the question is often whether an insurance company’s refusal to provide coverage, or to increase the premiums it charges for the same coverage, is an injustice when it is based on an actuarially justified (but perhaps not universal) generalization. As Leemens (2000) noted, the question was asked of the legislator when insurers observed that Jewish women from Eastern Europe were particularly vulnerable to breast and ovarian cancer. At the end of 2012, the European Court of Justice put an end to all discrimination based on the gender of policyholders: insurers were no longer able to differentiate between insurance product prices according to whether the member was male or female. But the use of age is still allowed. Indeed, age is often an indicator of a possible decrease in vision or hearing, slower reaction time (and increased risk of sudden disability), etc. And although there are many individual variations, the available data provide important empirical justification.

Machines, causality, and stereotypes

A major criticism of machine learning models is the lack of interpretation. But very often, the validation of econometric models requires a narrative built around stereotypes. And this narrative is essential, as Pearl & Mackenzie (2018) reminds us. Indeed, in the “The Ladder of Causation“, there are three levels. At the first level, we find the notion of association (or correlation), or even conditional probability, which serve as a basis for the constitution of stereotypes: if we observe

P[carries | brushing your teeth] < P[carries | don’t brush your teeth]

brushing teeth will be associated with a decrease in the probability of having carries. It is also the basis for regression methods, which are based on correlations between the variable of interest and others, wrongly called explanatory. In Figure 1, we can see the daily cycling traffic in Helsinki, and the average temperature. We will tend to prefer the one on the left, showing the evolution of the number of cyclists as a function of temperature, suggesting that temperature could explain the number of cyclists, and not the other way around. But the stereotype doesn’t necessarily focus on the causal link: if I see a lot of cyclists passing through the window, I’ll tell myself it must be hot, or at least warm. Figure 1: Näytä Data – Author’s visualization The first level answers the question “what if I see…?“(e.g. “what cycling traffic to expect if the temperature reaches 20°C? “) and this task can be perfectly accomplished by a machine. The second level is the one that makes it possible to understand an effect, an intervention. The question is then “what if I do…? “. To use our example, we are trying to understand the importance of brushing our teeth on the appearance of cavities. What if brushing your teeth is more natural for children with good teeth? We see the third level of the scale coming up, asking the question “what if I had done…?“and based on the idea of a counterfactual model. We are no longer content to measure correlations, we will build a model explaining what would happen by making a change in the causal variables: what would really happen if the child who did not brush his teeth began to do so? For Pearl & Mackenzie (2018) a human being (maybe even an actuary) can make these more advanced arguments than a machine can (yet) do. And very often, these causal patterns are stereotyped. As Charpentier & Diago Barry (2015) points out, in epidemiology, researchers have long questioned the explanation to be given to the fact that small babies of smokers have a higher probability of survival than babies of non-smoking mothers. The intuition that something is wrong comes from prejudices, stereotypes that we have, and that a machine cannot have.

When actuaries tell each other stories

As Antonio & Charpentier (2017) noted, the European “gender directive” has confused many insurers who used gender to construct their rates, as the latter was highly correlated with the frequency of claims. But by introducing telematic data, gender was no longer significant in the regression. Gender has long been used as a proxy to capture an effect that can be observed using telematic data, giving rise to many sexist stereotypes and other stereotypes. But the stories also make it possible to decide between a false correlation (“spurious correlation“) and a correlation that could be interpreted. In Figure 2, we have life expectancy at birth, a variable that we could try to explain in a pension study context, for example, by French department. On the right, two variables taken at random: the number of licenses of a tennis club, and the number of advertising agencies. Stereotypes are what will allow us to construct a causal graph, allowing us to understand why there is such a strong correlation between these variables and life expectancy. Figure 2: Life expectancy at birth for men, left. At the centre, number of tennis licenses per 100,000 inhabitants (source FFT). On the right, number of advertising agencies per 100,000 inhabitants (source INSEE, code NAF 7311Z). Visualization of the author.

Hyper-individualization as an answer?

While William Blake condemned stereotypes by saying “to generalize is to be an idiot“, he also clearly went further, continuing with “to particularize is the alone distinction of merit“. This individualisation is also advocated by more and more insurers, and even desired by many insureds. But as Grace & Terry (2002) pointed out, many policyholders suffer from a significant optimism bias – “if I have an accident, it will not be my fault” – leading them to doubt the insurer’s classification – “I’m less risky than the others“. And morality seems to prove them right, against actuaries. Yet, not only is generality not, in general, unjust, but justice itself can have considerable elements of generality. To the extent that justice is centred on equity and to the extent that equity itself is closely linked to equality, then equity, and therefore justice, can now be seen as itself based on the idea of generality. The just society is not necessarily a society in which each individual is treated as an isolated set of unique attributes, requiring individualized attention. On the contrary, in some cases, the just society is a society in which generality is not only unavoidable, but also necessary for justice itself. And pooling risks together is the natural response in an insurance context. And it might not be such a big deal if that class is not as homogenous at it could be, or as we would have expected it to be… Antonio, K. & Charpentier, A. (2017).  La tarification par genre en assurance, corrélation ou causalité ?. Risques. 110 : 107-110. Charpentier, A. & Diago Barry, A. (2015). Big data : passer d’une analyse de corrélation à une interprétation causale. Risques, 101: 107-111. Grace, J. & Terry, M. (2002). Exploring the Causes of Comparative Optimism. Psychologica Belgica. 42: 65–98 Kahneman, D. (2011).Thinking, Fast and Slow. FSG Eds. Leemens, T. (2000). Selective Justice, Genetic Discrimination, and Insurance: Should We Single Out Genes in Our Laws? McGill law journal. Revue de droit de McGill 45(2):347-412. Pearl, J. & Mackenzie, D. (2018). The Book of Why: The New Science of Cause and Effect. Basic Books. Schauer, F.F. (2009). Profiles, Probabilities, and Stereotypes. Harvard University Press.

Foundations of Machine Learning, part 5

This post is the nineth (and probably last) one of our series on the history and foundations of econometric and machine learning models. The first fours were on econometrics techniques. Part 8 is online here.

Optimization and algorithmic aspects

In econometrics, (numerical) optimization became omnipresent as soon as we left the Gaussian model. We briefly mentioned it in the section on the exponential family, and the use of the Fisher score (gradient descent) to solve the first order condition \mathbf{X}^T W(\beta)^{-1})[y-\widehat{y}]=\mathbf{0}. In learning, optimization is the central tool. And it is necessary to have effective optimization algorithms, to solve problems (described previously) of the form: \widehat{\beta}\in\underset{\beta\in\mathbb{R}^p}{\text{argmin}}\left\lbrace\sum_{i=1}^n \ell(y_i,\beta_0+\mathbf{x}^T\beta)+\lambda\Vert\boldsymbol{\beta}\Vert\right\rbraceIn some cases, instead of global optimization, it is sufficient to consider optimization by coordinates (widely studied in Daubechies et al. (2004)). If f:\mathbb{R}^d\rightarrow\mathbf{R} is convex and differentiable, if \mathbf{x} satisfies f(\mathbf{x}+h\boldsymbol{e}_i)\geq f(\mathbf{x}) for any h>0 and i\in\{1,\cdots, d\}then f(\mathbf{x})=\min\{f\}, where \mathbf{e}=(\mathbf{e}_i) is the canonical basis of \mathbb{R}^d. However, this property is not true in the non-differentiable case. But if we assume that the non-differentiable part is separable (additively), it becomes true again. More specifically, iff(\mathbf{x})=g(\mathbf{x})+\sum_{i=1}^d h_i(x_i)with\left\lbrace\begin{array}{l}g: \mathbb{R}^d\rightarrow\mathbb{R}\text{ convex-differentiable}\\h_i: \mathbb{R}\rightarrow\mathbb{R}\text{ convex}\end{array}\right.This was the case for Lasso regression, \beta)\mapsto\| \mathbf{y}-\beta_0-\mathbf{X}\beta\|_{\ell_2 }+\lambda\|\beta\|_{\ell_1}, as shown by Tsen (2001). Getting back to our initial notations, we can use a coordinate descent algorithm: from an initial value \mathbf{x}^{(0)}, we consider (by iterating)x_j^{(k)}\in\text{argmin}\big\lbrace f(x_1^{(k)},\cdots,x_{k-1}^{(k)},x_k,x_{k+1}^{(k-1)},\cdots,x_n^{(k-1)})\big\rbrace for j=1,2,\cdots,nThese algorithmic problems and numerical issues may seem secondary to econometricians. However, they are essential in automatic learning: a technique is interesting if there is a stable and fast algorithm, which allows to obtain a solution. These optimization techniques can be transposed: for example, this coordinate descent technique can be used in the case of SVM methods (known as “vector support” methods) when the space is not linearly separable, and the classification error must be penalized (we will come back to this technique in the next section).

In-sample, out-of-sample and cross-validation

These techniques seem intellectually interesting, but we have not yet discussed the choice of the penalty parameter \lambda. But this problem is actually more general, because comparing two parameters \widehat{\beta}_{\lambda_1} and \widehat{\beta}_{\lambda_2} is actually comparing two models. In particular, if we use a Lasso method, with different thresholds \lambda, we compare models that do not have the same dimension. Previously, we have addressed the problem of model comparison from an econometric perspective (by penalizing overly complex models). In the learning literature, judging the quality of a model on the data used to construct it does not make it possible to know how the model will behave on new data. This is the so-called “generalization” problem. The traditional approach then consists in separating the sample (size n) into two parts: a part that will be used to train the model (the training database, in-sample, size m) and a part that will be used to test the model (the testing database, out-of-sample, size n-m). The latter then makes it possible to measure a real predictive risk. Suppose that the data are generated by a linear model y_i=\mathbf{x}_i^T \beta_0+\varepsilon_i where \varepsilon_i are independent and centred law achievements. The empirical quadratic risk in-sample is here\frac{1}{m}\sum_{i=1}^m\mathbb{E}\big([\mathbf{x}_i^T \widehat{\beta}-\mathbf{x}_i^T \beta_0]^2\big)=\mathbb{E}\big([\mathbf{x}_i^T \widehat{\beta}-\mathbf{x}_i^T \beta_0]^2\big),for any observation i. Assuming the residuals \varepsilon Gaussian, then we can show that this risk is worth \sigma^2 \text{trace} (\Pi_X)/m is \sigma^2 p/m. On the other hand, the empirical out-of-sample quadratic risk is here \mathbb{E}\big([\mathbf{x}^T \widehat{\beta}-\mathbf{x}^T \beta_0]^2\big) where \mathbf{x} is a new observation, independent of the others. It can be noted that \mathbb{E}\big([\mathbf{x}^T \widehat{\beta}-\mathbf{x}^T \beta_0]^2\big\vert \mathbf{x}\big)=\text{Var}\big(\mathbf{x}^T \widehat{\beta}\big\vert \mathbf{x}\big)=\sigma^2\mathbf{x}^T(\mathbf{x}^T\mathbf{x})^{-1}\mathbf{x},and by integrating with respect to \mathbf{x}, \mathbb{E}\big([\mathbf{x}^T \widehat{\beta}-\mathbf{x}^T\beta_0]^2\big)=\sigma^2\text{trace}\big(\mathbb{E}[\mathbf{x}\mathbf{x}^T]\mathbb{E}\big[(\mathbf{x}^T\mathbf{x})^{-1}\big]\big).The expression is then different from that obtained in-sample, and using the Groves & Rothenberg (1969) increase, we can show that \mathbb{E}\big([\mathbf{x}^T \widehat{\beta}-\mathbf{x}^T \beta_0]^2\big) \geq \sigma^2\frac{p}{m}which is pretty intuitive, when we start thinking about it. Except in some simple cases, there is no simple (explicit) formula. Note, however, that if \mathbf{X}\sim\mathcal{N}(0,\sigma^2 \mathbb{I}), then \mathbf{x}^T \mathbf{x} follows a Wishart law, and it can be shown that \mathbb{E}\big([\mathbf{x}^T \widehat{\beta}-\mathbf{x}^T \beta_0]^2\big)=\sigma^2\frac{p}{m-p-1}.If we now look at the empirical version: if \widehat{\beta} is estimated on the first m observations,\widehat{\mathcal{R}}^{~\text{ IS}}=\sum_{i=1}^m [y_i-\boldsymbol{x}_i^T\widehat{\boldsymbol{\beta}}]^2\text{ and }\widehat{\mathcal{R}}^{\text{ OS}}=\sum_{i=m+1}^{n} [y_i-\boldsymbol{x}_i^T\widehat{\boldsymbol{\beta}}]^2and as Leeb (2008) noted, \widehat{\mathcal{R}}^{\text{IS}}-\widehat{\mathcal{R}}^{\text{OS}}\approx 2\cdot\nu where \nu represents the number of degrees of freedom, which is not unlike the penalty used in the Akaike test.

Figure 4 shows the respective evolution of \widehat{\mathcal{R}}^{\text{IS}} and \widehat{\mathcal{R}}^{\text{OS}} according to the complexity of the model (number of degrees in a polynomial regression, number of nodes in splines, etc). The more complex the model, the more \widehat{\mathcal{R}}^{\text{IS}} will decrease (this is the red curve, below). But that’s not what we’re interested in here: we want a model that predicts well on new data (i. e. out-of-sample). As Figure 4 shows, if the model is too simple, it does not predict well (as it does with in-sample data). But what we can see is that if the model is too complex, we are in a situation of “overlearning”: the model will start to model the noise. Of course, this figure should remind us of the one we’ve seen in our second post of that series

Figure 4 : Generalization, under- and over-fitting

Instead of splitting the database in two, with some of the data that will be used to calibrate the model and some to study its performance, it is also possible to use cross-validation. To present the general idea, we can go back to the “jackknife”, introduced by Quenouille (1949) (and formalized by Quenouille (1956) and Tukey (1958)) relatively used in statistics to reduce bias. Indeed, if we assume that \{y_1,\cdots,y_n\} is a sample drawn according to a law F_\theta, and that we have an estimator T_n (\mathbf{y})=T_n (y_1,\cdots,y_n), but that this estimator is biased, with \mathbf{E}[T_n (\mathbf{Y})]=\theta+O(n^{-1}), it is possible to reduce the bias by considering \widetilde{T}_n(\mathbf{y})=\frac{1}{n}\sum_{i=1}^n T_{n-1}(\mathbf{y}_{(i)})\text{ where }\mathbf{y}_{(i)}=(y_1,\cdots,y_{i-1},y_{i+1},\cdots,y_n)It can then be shown that \mathbb{E}[\tilde{T}_n(Y)]=\theta+O(n^{-2})The idea of cross-validation is based on the idea of building an estimator by removing an observation. Since we want to build a predictive model, we will compare the forecast obtained with the estimated model, and the missing observation\widehat{\mathcal{R}}^{\text{ CV}}=\frac{1}{n}\sum_{i=1}^n \ell(y_i,\widehat{m}_{(i)}(\mathbf{x}_i))We will speak here of the “leave-one-out” (loocv) method.

This technique reminds us of the traditional method used to find the optimal parameter in exponential smoothing methods for time series. In simple smoothing, we will construct a forecast from a time series as {}_t\widehat{y}_{t+1} =\alpha\cdot{}_{t-1}\widehat{y}_t +(1-\alpha)\cdot y_t, where \alpha\in[0,1], and we will consider as “optimal” \alpha^\star = \underset{\alpha\in[0,1]}{\text{argmin}}\left\lbrace \sum_{t=2}^T \ell({}_{t-1}\widehat{y}_{t},y_{t}) \right\rbraceas described by Hyndman et al (2009).

The main problem with the leave-one-out method is that it requires calibration of n models, which can be problematic in large dimensions. An alternative method is cross validation by k-blocks (called “k-fold cross validation”) which consists in using a partition of \{1,\cdots,n\} in k groups (or blocks) of the same size, \mathcal{I}_1,\cdots,\mathcal{I}_k, and let us note \mathcal{I}_{\bar j}=\{1,\cdots,n\}\setminus \mathcal{I}_j. By noting \widehat{m}_{(j)} built on the sample \mathcal{I}_{\bar j}, we then set:\widehat{\mathcal{R}}^{k-\text{ CV}}=\frac{1}{k}\sum_{j=1}^k \mathcal{R}_j\text{ where }\mathcal{R}_j=\frac{k}{n}\sum_{i\in\mathcal{I}_{{j}}} \ell(y_i,\widehat{m}_{(j)}(\mathbf{x}_i))Standard cross-validation, where only one observation is removed each time (loocv), is a special case, with k=n. Using k=5 or 10 has a double advantage over k=n: (1) the number of estimates to be made is much smaller, 5 or 10 rather than n; (2) the samples used for estimation are less similar and therefore less correlated to each other, which tends to avoid excess variance, as recalled by James et al. (2013).

Another alternative is to use boosted samples. Let \mathcal{I}_b be a sample of size n obtained by drawing with replacement in \{1,\cdots,n\} to know which observations (y_i,\mathbf{x}_i) will be kept in the learning population (at each draw). Note \mathcal{I}_{\bar b}=\{1,\cdots,n\}\setminus\mathcal{I}_b. By noting \widehat{m}_{(b)} built on sample \mathcal{I}_b, we then set :\widehat{\mathcal{R}}^{\text{ B}}=\frac{1}{B}\sum_{b=1}^B \mathcal{R}_b\text{ where }\mathcal{R}_b=\frac{n_{\overline{b}}}{n}\sum_{i\in\mathcal{I}_{\overline{b}}} \ell(y_i,\widehat{m}_{(b)}(\mathbf{x}_i))where n_{\bar b} is the number of observations that have not been kept in \mathcal{I}_b. It should be noted that with this technique, on average e^{-1}\sim36.7\% of the observations do not appear in the boosted sample, and we find an order of magnitude of the proportions used when creating a calibration sample, and a test sample. In fact, as Stone (1977) had shown, the minimization of AIC is to be compared to the cross-validation criterion, and Shao (1997) showed that the minimization of BIC corresponds to k-fold cross-validation, with k=n/\log n.

All those techniques here are mentioned in the “machine learning” section since they rely on automatic, computational techniques, and no probabilistic foundations are necessary. In many cases we did use the notation m^\star (at least in the first posts on “machine learning” techniques) to highlight the fact that we want some sort of “optimal” model – and to make a distinction with estimators \widehat{m} considered earlier, when we had some probabilistic framework. But of course, it is possible (and necessary) to build bridges between those two cultures…

References are online here. As explained in the introduction, it is some sort of online version of an introduction to our joint paper with Emmanuel Flachaire and Antoine Ly, Econometrics and Machine Learning (initially writen in French), that will actually appear soon in the journal Economics and Statistics (in English and in French).

Insurance, Actuarial Science, Data and Models

Our research chaire ACTINFO, with our colleagues from Lyon, at the DAMI chaire,  PREVENT’HORIZON chaire & ACTUARIAT DURABLE chaire, will organize a 2 day conference in Paris, on Insurance, Actuarial Science, Data & Models, in ten days.

We invited Katrien ANTONIO (KU Leuven), Alexandre BOUMEZOUED (Milliman Paris), Alfred GALICHON (New-York University), Pierre-Yves GEOFFARD (Paris School of Economics), Meglena JELEVA (University of Paris Nanterre), Julie JOSSE (Ecole Polytechnique), Florence JUSOT (Paris Dauphine University), Michael LUDKOWSKI (University of California Santa Barbara), François PANNEQUIN (CREST and ENS Paris-Saclay), Florian PELGRIN (Edhec Business School), Dylan POSSAMAI (Columbia University) and Julien TRUFIN (ULB Brussels). More information (including the program) is online.

Picking an asset to invest

Yesterday, Andrew Lo spent some time on a nice graph, discussing attitudes towards risk. Here are four assets (thanks  for improving the terminology), real data (no information here about time, but it’s the same scale for the four of them)

The question raised was quite simple

if you could invest in one, and only one, asset which one will you pick ?

Continue reading Picking an asset to invest