This afternoon, I will give a series of two short lectures on fairness and performance in actuarial modeling. Slides are now available.
Tag Archives: fairness
Recoverability of market-wide fair insurance premiums under selection bias
Our article, Recoverability of market-wide fair insurance premiums under selection bias, with Marie-Pier Côté and Olivier Côté, was published in Insurance: Mathematics & Economics.

Fairness adjustments in insurance pricing are defined relative to a reference population, i.e., to a joint distribution of (X, D, Y) where X are rating factors, D protected attributes, and Y is claim amount. Because an insurer’s portfolio is generally a selected subpopulation, portfolio and population reference distributions typically differ, so portfolio-calibrated and population-calibrated fairness adjustments need not coincide. In what follows, we use selection bias as an umbrella term for any discrepancy between an observed sample and its target population. We call portfolio composition bias the insurance-specific form of selection bias induced by the portfolio inclusion mechanism (underwriting/marketing), which makes each insurer’s portfolio a selected subpopulation. Relying on causal inference and a portfolio composition indicator, we characterize how portfolio composition bias affects common premium adjustments (unawareness, discrimination-free pricing, and transport-based corrective pricing), and we provide restrictive conditions under which portfolio and population adjustments coincide. We propose estimators to recover the fairness-adjusted premiums on the regulator-intended target population from selection-biased data, by using externally available information on the population marginal distribution of the prohibited attribute D. We study this scope mismatch from the policyholder’s perspective: we model the market premium faced by a newly entering policyholder (not yet assigned to any portfolio) as a mixture of insurer-specific premiums, weighted by the probability of being assigned to each insurer. Under this view, a pricing rule can satisfy a fairness criterion within each insurer’s portfolio yet produce direct or proxy discrimination in the market when portfolio inclusion depends on X and/or D. Finally, we show that enforcing portfolio-level balance on population-intended fair premiums can reintroduce portfolio composition bias, highlighting a regulatory trade-off between portfolio balancing and market-wide fairness. We focus on recoverability: which population-level fairness targets are identifiable from portfolio data, and what minimal external information is required to recover them.
Perceived Fairness in Networks
Last week, the paper Perceived fairness in networks was published in Network Science.
The usual definitions of algorithmic fairness focus on population-level statistics, such as demographic parity or equal opportunity. However, in many social or economic contexts, fairness is not perceived globally, but locally, through an individual’s peer network and comparisons. We propose a theoretical model of perceived fairness networks, in which each individual’s sense of discrimination depends on the local topology of interactions. We show that even if a decision rule satisfies standard criteria of fairness, perceived discrimination can persist or even increase in the presence of homophily or assortative mixing. We propose a formalism for the concept of fairness perception, linking network structure, local observation, and social perception. Analytical and simulation results highlight how network topology affects the divergence between objective fairness and perceived fairness, with implications for algorithmic governance and applications in finance and collaborative insurance.
Insurance, Biases, Discrimination and Fairness
Two years ago, my book “Insurance, Biases, Discrimination and Fairness” was published in the Springer Actuarial series.
Discrimination in insurance is a difficult topic because, in a very specific sense, insurers are expected to discriminate: they classify risks, build risk pools, and differentiate premiums. This is the logic of risk-based pricing and actuarial fairness. But insurance is not only about pricing risks accurately (accuracy is overrated). It is also about mutualization, risk sharing, and solidarity. The real question is therefore not simply whether insurers should differentiate, but which differences should matter, which should not, where the limits should be drawn, and how to navigate a complex world in which several notions and metrics of fairness coexist, sometimes in tension with one another.
In the book, I tried to connect actuarial pricing, statistical discrimination, legal constraints, algorithmic fairness, explainability, mitigation techniques, and the limits of “fairness through unawareness”. I also discuss a dimension that is often overlooked: causality. Correlation may be useful for prediction, but prevention, explanation, and fairness often require asking what mechanism lies behind an observed association.
For those interested in the mathematics, I have also made lecture notes freely available.

IFoA AI Ethics Governance & Risk Management forum
I was invited to give a talk at the “IFoA AI Ethics Governance & Risk Management forum” at the end of the week. I uploaded some slides to start the discussion.

I will not give a highly technical lecture, but rather to propose a way of thinking about fairness in insurance AI that may be useful for actuaries, for governance people, and for anyone involved in model risk or oversight. The main message is simple. In insurance, fairness cannot be reduced to predictive accuracy. And once we move to modern AI and large data environments, simply removing a sensitive variable is no longer enough to prevent discrimination. So the real question becomes: how should we govern the trade-offs between fairness and accuracy, and different notions of fairness?

Let me start with a very simple example, illustrating the fact that simply removing a sensitive variable is not enough to prevent discrimination. Here I take a French motor insurance dataset. In the raw data, average claim frequencies are about 8.94% for men and 8.20% for women. Now suppose that I estimate a logistic model for annual claim frequency, but I explicitly exclude gender from the explanatory variables. At first sight, this looks like the fair thing to do. I am not using gender, so maybe I am safe. But what happens?
As I add more and more variables to the model, the predicted frequencies for men and women move closer and closer to the empirical frequencies in the data. With enough explanatory variables, I am basically reconstructing the original gap. So the model is not using gender explicitly, but it is learning from variables that are correlated with gender. In other words, the information has not disappeared. It has simply been redistributed across other variables.
That is why I think the old defence of ‘we do not use the sensitive variable, therefore the model is fair’ has become much weaker in a big-data context (and here, it’s not big, it’s only 15 basic explanatory variables). Modern predictive models are very good at finding statistical associations. If protected traits leave traces in the data, the model will usually pick them up.
And this is not only a technical issue. It is also a philosophical one. The model tends to reproduce what is in the data, that’s what we call ‘generalization’ in machine learning. If historical inequalities are present in the data, the model will often learn them very efficiently. This is why I mention Hume here, and the old ‘is-ought’ problem. From the fact that a disparity exists in the data, it does not follow that it should be reproduced in pricing. So this slide is really the starting point for the rest of the talk: fairness through unawareness is often not fairness at all; it is frequently a way of hiding the channel through which unfairness reappears.
So once we accept that unawareness is not enough, we need to ask a more difficult question: what exactly do we mean by fairness?

Before going to the recent actuarial papers, I want to spend one minute on a classical dilemma. This is not new. In legal and philosophical debates, there has long been a tension between two positions.
One position says: if you want to correct unequal outcomes, you may need to take account of the protected characteristic. That is the logic behind many discussions of affirmative action or corrective treatment. If the world is already unequal, pretending not to see the relevant characteristic may simply reproduce the inequality. The opposite position is the so-called colorblind view: the way to stop discrimination is to stop using the sensitive characteristic altogether. That sounds very appealing, because it seems neutral and simple.
The trouble is that in insurance, and especially in AI-based pricing, both positions run into difficulties. If you ignore the sensitive variable, you may still reproduce its effects through proxies, that’s what we’ve seen before. But if you use it for corrective purposes, you are entering a difficult space involving direct differential treatment, redistribution, and legal or ethical contestability.
So what I want to emphasize here is that the fairness debate is not just about technical metrics. It is rooted in a much older normative tension: should fairness mean treating everyone under the same formal rule, or should it sometimes mean treating people differently in order to offset structural inequalities? I think that this dilemma is exactly why actuarial fairness debates cannot be settled by model performance alone.
In the contexte of insurance, the recent SOA report discussed in a previous meeting stresses this important tension fairness is not one thing.

The report is useful because it is very practical and very clear. Its main point is that fairness in life insurance is not a single concept. There is no universal fairness metric that would solve the problem once and for all. Instead, the paper distinguishes between different families of fairness criteria.
A first distinction is between individual fairness and group fairness. Individual fairness is close to the actuarial intuition: similar risks should be treated similarly. Group fairness, by contrast, is about parity across groups. Those two perspectives can both sound reasonable, but they are often in tension. The paper also reviews several criteria. On the individual side, it discusses unawareness, awareness, and omitted-variable bias. On the group side, it discusses independence, sufficiency, and separation.
The most useful takeaway is not any particular metric. It is the methodological lesson: before choosing a metric, you must define what notion of fairness you are trying to serve. Otherwise you are optimizing a number without knowing what ethical or institutional objective it really represents. The SOA paper gives us a map. It tells us: do not look for the one true fairness metric. Clarify the objective first, and then be explicit about the trade-offs.”
Our own recent work tries to push that argument one step further, by asking what the core dimensions of fairness in insurance pricing really are…

In our recent paper on what we call, with Olivier Côté, and Marie-Pier Côté, “the fairness trilemma“, we recall the idea that insurance is not just a predictive exercise. It is also a social institution. Insurance has a double nature. On the one hand, it is about risk-based pricing: aligning premiums with expected losses. On the other hand, it is also about risk-sharing: spreading burdens, preserving access, and sometimes accepting cross-subsidies. From there, we argue that fairness in insurance pricing is governed by a trilemma between three principles.
- The first is actuarial fairness. In simple terms, this means that premiums should reflect expected losses as accurately as possible.
- The second is social solidarity. This means that we may accept departures from strict risk-based pricing in order to preserve access, affordability, or some broader social objective.
- The third is causal legitimacy. This is extremely important in the AI context. Not every predictive variable is equally legitimate. A variable may improve predictive performance, but if it works mainly by reconstructing a protected characteristic, or if its causal meaning is questionable, then its legitimacy is fragile. So causal legitimacy asks not only whether a factor predicts, but whether it deserves to shape the price.
The key point is that these three principles are all attractive, but they cannot be fully satisfied at the same time. If I strengthen actuarial fairness, I may weaken solidarity. If I enforce more solidarity, I move away from pure risk adequacy. If I become very demanding about causal legitimacy, I may sacrifice some predictive performance and perhaps some segmentation logic. So the problem is not to find the perfectly fair model. The problem is to govern the trade-offs. That is why I insist on governance tools: scorecards, causal due diligence, comply-or-explain procedures. The point is to make choices explicit, reviewable, and contestable. In that sense, fairness is not just a modelling problem. It is a governance problem.
But of course, at this stage, someone in the room may say: all this sounds interesting, but how do we make it operational? How do we move from philosophy to something an actuary can actually measure?

That is exactly the purpose of the CAS paper. The starting point is again that removing a sensitive variable does not remove discrimination, because allowed variables may still act as proxies. So the challenge is to expose indirect discrimination in a way that is usable in actuarial practice. The objective of the paper is to make fairness operational and measurable in actuarial terms. It is not enough to speak in abstract statistical language. In insurance, we need to know what unfairness means in premiums, in dollars, and across segments of policyholders.
The paper is based on a real auto insurance case study from Québec, using credit score as the sensitive variable. And it again organizes the discussion around three dimensions: actuarial fairness, social solidarity, and causality legitimacy. This is important because it shows continuity with our more conceptual framework. The toolbox is not something separate. It is a way of translating those dimensions into measurable diagnostics. If the trilemma paper says fairness must be governed, the CAS paper says: here are some tools that help you see what you are governing.

One of the most interesting contributions of the toolbox is that it introduces local metrics such as risk spread, proxy vulnerability, fairness range, and parity cost. Let us focus on two of them.
The first is proxy vulnerability. Intuitively, proxy vulnerability measures how much a segment may be over-priced or under-priced because apparently neutral variables indirectly reconstruct the sensitive attribute. This is very valuable because it moves us away from vague concerns about bias and toward a concrete actuarial question: where, and by how much, might the pricing rule be unfair?
The second is parity cost. This is also very important because it forces us to be honest. If we want more parity across groups, that is not free. Someone has to bear the cost. In plain English, enforcing one notion of fairness may end up robbing Peter to pay Paul. And I do not say this as a criticism; I say it because redistribution should be explicit rather than hidden.
Another strength of the toolbox is that it helps actuaries identify local pockets of unfairness. A model may look acceptable on average across broad groups, but still be very problematic in specific vulnerable niches. So fairness should not be assessed only globally. Local diagnostics matter.
This is also, I think, very relevant for AI governance. Senior management and boards do not just need a statement saying ‘the model passed a fairness test’. They need to know where the model is fragile, which subpopulations are exposed, what the monetary magnitude is, and what the trade-offs are if a correction is introduced. So in practice, the toolbox helps turn fairness into something discussable in risk committees and governance processes.

To wrap up, first, the broader background for all this is developed in my recent textbook, which tries to connect legal, philosophical, statistical, and actuarial perspectives on discrimination in insurance. Second, fairness should not be assessed only at the level of one insurer’s portfolio. A pricing rule may look fair on one portfolio and still be unfair at the market level, because portfolios are not representative of the insured population. This is a selection-bias problem. If different insurers attract different segments of the market, portfolio-specific fairness does not necessarily aggregate into market-wide fairness. So we also need to ask: fair for whom, and on which reference population?
Third, there is always a cost and a redistribution issue. If we enforce one notion of fairness, someone pays. That does not mean fairness is undesirable. It means that fairness choices are unavoidably political, institutional, and governance choices, not just technical ones. And fourth, there is a major practical difficulty when sensitive variables are unobserved. Very often, people say: let us not collect the sensitive variable. But then how do we audit fairness? How do we provide evidence? How do we challenge the model? If the protected attribute is precisely the thing we do not observe, fairness becomes not just a modelling issue, but a data, audit, and governance issue.
So my overall conclusion would be the following. AI did not create the fairness problem in insurance, but it has made it sharper. It has weakened the old comfort of unawareness. It has increased the power of proxies. And it has made the trade-offs more difficult to hide. For actuaries, I think this creates both a risk and an opportunity. The risk is to treat fairness as a box-ticking exercise or as a purely technical constraint. The opportunity is to contribute something distinctive: a language of prices, trade-offs, portfolios, and governance. So perhaps the right ambition is not to promise a perfectly fair model. It is to build institutions that are capable of identifying proxy discrimination, making normative choices explicit, and governing them responsibly.
On my way to Tsinghua (清华大学), Beijing
Next week, I will be at Tsinghua University in Beijing. On Tuesday, in the early afternoon, I will give three lectures for undergraduate students on the theme: ‘Three lectures on AI and its implications for actuarial (and/or financial) professions.’
These lectures explore the relationship between artificial intelligence and insurance. They begin from the observation that insurance has long relied on prediction, classification, and decision-making under uncertainty, well before the recent rise of AI. AI therefore does not introduce these issues from scratch, but changes their scale, granularity, and practical consequences. The lectures review the insurance foundations of pricing and pooling, then examine the main challenges raised by AI, including personalization, selection, causality, bias, fairness, governance, and trust. They finally turn to the concrete uses of AI across the insurance value chain, emphasizing that a good system should not be judged by accuracy alone, but also by its calibration, its fairness, and its ability to support real decisions in practice.
In the evening, I will give a talk at the seminar, at Renmin University of China, on the theme: ‘Using optimal transport to mitigate unfair predictions and quantify counterfactual fairness.’ The first part will revisit topics that I presented in greater detail in the lectures notes of my course this autumn at Kyoto University, particularly the price to be paid in terms of accuracy in order to achieve fairness. The second part will discuss the paper ‘Sequential Transport for Causal Mediation Analysis,’ which was posted online a few days ago.
On Wednesday, I will have in-depth academic exchange session with students from the Tsinghua Actuarial Science Association, at Tsinghua University.
Newsletter #5/6
The latest newsletter (Fall and Winter activities) related to our research project on algorithmic fairness and insurance markets is finally out Newsletter_2026_5
Get back to us if you want more details or just to share some feedbacks… Une version en français est également disponible Infolettre_2026_5
Mesurer l’équité “globale” quand les données sont dispersées
Avec Agathe Fernandes Machado, Olivier Côté et François Hu, on a mis en ligne un papier, Federated Measurement of Demographic Disparities from Quantile Sketches. Le point de départ est assez simple. On peut imaginer qu’un modèle de score (mesurant un risque de récidive, une probabilité de réadmission à l’hôpital, un score de crédit…) soit déployé dans plusieurs institutions : hôpitaux, tribunaux, banques, assureurs. Chacun collecte ses données, utilise un modèle, et conserve jalousement ses bases. Parfois par obligation légale (RGPD, secret médical), parfois par contraintes techniques, parfois par réticence organisationnelle. Le problème, c’est qu’un régulateur veut savoir si le score est discriminatoire, sans jamais centraliser les données brutes. En fait, c’est assez réaliste comme situation, beaucoup d’objectifs de justice algorithmique étant définis au niveau de la population, et pas localement. Les régulateurs et les directions conformité demandent : “Est-ce que le système, globalement, traite de la même façon les groupes protégés ?” Pas : “Chaque institution, isolément, a-t-elle l’air correcte ?” On montre dans notre article que des audits locaux peuvent être rassurants tout en étant trompeurs, parce que l’injustice peut naître précisément de ce que l’on ne voit pas en restant silo par silo. Et la bonne nouvelle, c’est qu’on peut estimer l’inéquité globale avec une communication très limitée, en demandant à chaque silo seulement des comptages et quelques quantiles de ses scores.
Dans le papier, on identifie deux sources majeures de décalage entre l’audit local et l’audit global.
Les effets de composition : une “version fairness” du paradoxe de Simpson. Même si chaque silo semble traiter les groupes de manière similaire, la répartition des groupes entre silos peut être très différente. Un groupe peut être sur-représenté dans certains hôpitaux, certains tribunaux, certaines zones géographiques… Et si ces silos n’ont pas le même profil de scores (parce que les populations, les pratiques ou les contextes diffèrent), alors l’agrégation peut créer un écart global qui n’apparaît nulle part localement.
L’hétérogénéité inter-silos : la “stratification cachée” Deux silos peuvent produire des scores d’allure différente (distribution plus “optimiste”, plus “pessimiste”, plus dispersée…), même au sein d’un même groupe sensible. Localement, chacun peut avoir des métriques acceptables. Mais une fois les données mises bout à bout, ces différences deviennent visibles et peuvent amplifier une disparité entre groupes. Dans les domaines sensibles (santé, justice pénale), cette hétérogénéité est courante : pratiques de codage, accès aux soins, critères de triage, politiques locales…
D’un point de vue pratique, on propose un protocole d’audit en un seul aller-retour : chaque silo envoie, pour chaque groupe sensible le nombre d’individus dans ce groupe (un simple comptage), k quantiles du score (par exemple k = 25, 50 ou 100), sur une grille commune. Et c’est tout. Pas de scores individuels, pas de features, pas d’exemples. Ce genre de résumé est déjà produit par de nombreux systèmes de monitoring via des quantile sketches utilisés pour suivre des distributions. À partir de ces quantiles, le serveur peut reconstruire une approximation des distributions globales par groupe, puis calculer la disparité populationnelle. Théoriquement, l’erreur due à la discrétisation décroît comme 1/k : plus on envoie de quantiles, plus la courbe reconstruite est fine. Et on montre sur des données réelles que quelques dizaines suffisent.
En bonus, on propose aussi une méthode permettant de comprendre, si on mesure un écart, pourquoi il apparaît. On obtient en particulier une décomposition de type ANOVA qui sépare : une part due aux effets de mélange / composition (le “Simpson fairness”), une part due à la vraie hétérogénéité inter-silos (différences structurelles de score), et un terme d’interaction qui peut amplifier ou compenser (mais reste contrôlé).
Bref, on montre qu’en environnement fédéré, l’équité populationnelle n’est pas la moyenne de l’équité locale. Elle dépend des mélanges, des flux, des biais d’affectation et des variations inter-silos. Donc la bonne question n’est pas “chaque silo est-il juste ?”, mais “le système fédéré, en tant que mécanisme de production de scores, est-il juste au niveau population ?” La bonne nouvelle, c’est qu’on peut répondre à cette question sans centraliser les données, en ne partageant qu’une poignée de quantiles et des comptes, en une seule communication.
Beyond Procedure: Substantive Fairness in Conformal Prediction
Our paper, Beyond Procedure: Substantive Fairness in Conformal Prediction, with Pengqi Liu, Zijun Yu, Mouloud Belbahri, Masoud Asgharian, and Jesse Cresswell, is now available on https://arxiv.org/abs/2602.16794
Conformal prediction (CP) offers distribution-free uncertainty quantification for machine learning models, yet its interplay with fairness in downstream decision-making remains underexplored. Moving beyond CP as a standalone operation (procedural fairness), we analyze the holistic decision-making pipeline to evaluate substantive fairness-the equity of downstream outcomes. Theoretically, we derive an upper bound that decomposes prediction-set size disparity into interpretable components, clarifying how label-clustered CP helps control method-driven contributions to unfairness. To facilitate scalable empirical analysis, we introduce an LLM-in-the-loop evaluator that approximates human assessment of substantive fairness across diverse modalities. Our experiments reveal that label-clustered CP variants consistently deliver superior substantive fairness. Finally, we empirically show that equalized set sizes, rather than coverage, strongly correlate with improved substantive fairness, enabling practitioners to design more fair CP systems. Our code is available at this https URL.
Au-delà de la procédure, quand “bien calibrer” ne suffit pas à être juste
Quand on parle d’IA “responsable”, souvent, on englobe deux choses. La fiabilité (on comprend ce qu’on fait) et l’équité (on traite tout le monde pareil). Dit comme ça, ça semble simple et clair. Mais forcément, dans la vraie vie, c’est un peu plus subtil. En particulier on peut avoir une procédure statistiquement irréprochable, juste… et produire quand même des décisions inéquitables, une fois mise en place. C’est ce qu’on essayait de comprendre (et de mesurer) dans “Beyond Procedure: Substantive Fairness in Conformal Prediction“, coécrit avec Pengqi Liu, Zijun Yu, Mouloud Belbahri, Masoud Asgharian, et Jesse Cresswell: la conformité d’une procédure ne garantit pas la justice de ses effets.
Passer de la “fairness” de la méthode à la “fairness” des conséquences
Cette distinction rappelle des discussions en cours d’éthique : faut-il dire qu’un système est juste parce que sa règle est la même pour tous (vision procédurale, proche d’une éthique des principes), ou parce qu’il produit des effets équitables (vision conséquentialiste) ? Ici, l’enjeu n’est pas tant de “maximiser” un score global que de vérifier que l’aide fournie ne profite pas surtout à certains groupes. En terme méthodologique, notre article s’inscrit dans la littérature sur la prédiction conforme (conformal prediction, CP). L’intuition est assez simple pour des classifieurs (c’est le cadre qu’on regarde ici). Au lieu de prédire une seule étiquette (“il va pleuvoir”, “c’est un médecin”, “c’est la classe 7”, “ce revenu est dans telle tranche”), on prédit un ensemble de réponses plausibles. Avec un degré de confiance. Par exemple “je suis sûr à 90% que la bonne réponse est dans cette liste.” Cette promesse, on appelle ça la couverture. À un niveau 90%, CP garantit que, en moyenne, la vérité est bien dans l’ensemble retourné. Autrement dit, peu importe le modèle, peu importe la distribution, on a une garantie statistique robuste. Mais dans la vraie vie, personne ne décide à partir d’une garantie abstraite de couverture. On décide à partir de l’ensemble de prédictions fournies. Et cet ensemble peut varier fortement d’une personne à l’autre, ou d’un groupe à l’autre.
- Si l’ensemble contient 1 réponse, c’est presque une décision automatique.
- S’il en contient 7, c’est une aide beaucoup plus vague : on hésite, on délègue, on choisit “au feeling”, ou on s’en remet à un autre processus.
Autrement dit, une même “garantie à 90%” peut se traduire par des niveaux d’aide très différents. Pour revenir à ce que je disais au début, on fait une distinction centrale dans le papier entre :
- la Fairness procédurale : la méthode respecte des critères “internes” (ex. même couverture par groupe).
- la Fairness substantive : au final, l’outil améliore (ou dégrade) l’équité des résultats et des bénéfices entre groupes.
Mais “égaliser la couverture” peut aggraver l’iniquité
Dans la littérature, une idée a longtemps semblé naturelle. Si on veut être juste entre groupes (genre, âge, origine…), alors il faut que la garantie statistique soit la même pour tous. En CP, cela conduit à ce qu’on appelle l’Equalized Coverage : calibrer séparément par groupe, pour que chacun obtienne, disons, 90% de couverture. Sur le plan procédural, c’est pas mal, mais sur le plan des conséquences, le papier montre que cela peut être un mauvais objectif. Pourquoi ? Tout simplement parce que calibrer par groupe revient à découper les données de calibration. Dans ce cas, certains groupes ont moins d’exemples, ce qui rend les seuils plus instables. Et pour tenir la couverture, on compense souvent en produisant des ensembles plus grands pour certains groupes. Et si un groupe reçoit systématiquement des ensembles plus larges, il reçoit une aide moins facile à utiliser en pratique.
Le papier propose alors en avant un autre indicateur procédural, plus proche des effets : l’Equalized Set Size (parité de taille des ensembles). Et ce qu’on montre, c’est que dans les expériences, la parité de taille est fortement associée à une meilleure équité des résultats, tandis que la parité de couverture est souvent corrélée à une équité plus mauvaise. Dit autrement, optimiser le bon niveau de couverture peut pousser dans la mauvaise direction du point de vue social.
Mesurer l’équité “substantive” sans mobiliser des centaines de participants
En pratique, évaluer l’équité des conséquences est coûteux, puisqu’il faudrait idéalement des expériences avec des humains, parce que ce sont eux (ou leurs règles décisionnelles) qui transforment l’ensemble de prédiction en décision. On a essayé d’être pragmatique, en remplaçant l’évaluateur humain par un LLM dans un protocole contrôlé (LLM-in-the-loop), de façon à simuler l’usage des ensembles de prédiction comme “aide à la décision”, à grande échelle, sur plusieurs modalités (image, texte, audio, tabulaire). L’idée n’est pas de dire que “les LLM pensent comme nous” en général, mais de construire un protocole où l’on peut donner au modèle une tâche (ex. prédire une profession, une émotion, une tranche de revenu), lui donner (ou non) l’ensemble conforme + une phrase de garantie (“la vraie réponse est dedans avec 90% de confiance”) et mesurer combien l’aide améliore l’exactitude, et surtout si cette amélioration est partagée équitablement entre groupes. Pour quantifier cette équité, on utilise une mesure inspirée des essais contrôlés, la disparité maximale de bénéfice entre groupes, résumée par un indicateur de type maxROR (une façon de comparer l’ampleur du “gain” entre groupes). Plus maxROR est faible, plus l’amélioration procurée par l’outil est équitable.
La taille des ensembles explique mieux l’équité réelle que la couverture
Sur quatre tâches (vision, texte, audio, données tabulaires) et plusieurs méthodes, on a un constat assez robuste. Tout d’abord, les méthodes qui cherchent à égaliser la couverture (Mondrian, group-clustered) ne sont pas celles qui donnent les meilleurs résultats en équité “substantive”. Mais en plus, les méthodes qui réduisent les écarts de taille d’ensemble (en particulier Label-Clustered CP, et parfois Backward CP) tendent à produire une amélioration plus équitable. Autrement dit, la prédiction conforme donne une garantie statistique particulièrement intéressante, mais la justice réelle se joue dans l’usage, et on montre que, pour des décisions assistées, l’égalité de taille des ensembles est souvent un meilleur levier d’équité que l’égalité de couverture.
Modeling and Understanding Indirect Discrimination in Algorithmic Fairness
In a couple of days, I will give a talk on “Modeling and Understanding Indirect Discrimination in Algorithmic Fairness” at Singapore campus – ESSEC Asia-Pacific. The abstract is
Observed disparities between groups in algorithmic decisions (whether in hiring, credit approval, or risk prediction) do not necessarily imply direct discrimination. They may also stem from legitimate differences in the distribution of explanatory attributes. Understanding and quantifying which components of these gaps are “explained” versus those that reflect direct or indirect discrimination lies at the core of modern causal approaches to algorithmic fairness. This talk will begin with an accessible introduction to group-gap decomposition, building on the classical Kitagawa–Oaxaca–Blinder econometric framework. This approach separates differences attributable to observable characteristics from residual components that may signal discriminatory effects. The second part will introduce recent developments leveraging optimal transport to construct individual-level counterfactuals, enabling estimation of direct and indirect causal effects for each observation. In particular, we will show how sequential transport mappings aligned with a causal graph can disentangle pathways and quantify the contribution of each mediator. This methodology overcomes limitations of traditional linear models, introduced by Kitagawa, Oaxaca and Blinder, provides interpretable counterfactuals, and is well suited to complex empirical settings. The presentation will combine intuitive motivation, illustrative examples, and recent research insights, with the goal of making these tools accessible and useful to researchers in management science, applied economics, and data science.
Talk at NTU (Nanyang) in Singapore
Tomorrow, I will be at Nanyang Technological University to give a talk at an internal seminar, “Fairness and discrimination in insurance“
What’s unique about insurance is that even statistical discrimination, which by definition is devoid of malicious intent, poses significant challenges. Because, on the one hand, policymakers would like insurers to treat their policyholders equally, without discrimination based on race, gender, age or other characteristics, even if it could make (statistical) sense to (indirectly) discriminate. On the other hand, at the core of actuaries’ activities lies discrimination, between risky and non-risky policyholders. And this risk is often statistically correlated with sensitive characteristics that regulation would like to prohibit insurers from taking into account. The analysis of possible discrimination in decision rules, whether human or algorithmic, is an old subject. Most of the concepts date back at least to the 50s, but recent developments in artificial intelligence have brought these issues back into the spotlight. Massive data facilitate statistical or proxy discrimination, and black-box algorithms do not facilitate understanding. Not to mention the various regulations that make it difficult to collect sensitive information, and ultimately test whether decisions can be discriminated against, especially indirectly.
III Congreso Universitario Internacional sobre Seguros y Reaseguros en Perú
In a few hours, I will give a talk at the III Congreso Universitario Internacional sobre Seguros y Reaseguros en Perú,
I will give a talk on “Detecting Hidden Bias in Insurance AI Through Counterfactuals”
As AI becomes embedded in underwriting, pricing, fraud detection, and claims automation, one challenge remains widely underestimated: models can discriminate without ever using a prohibited variable. Indirect discrimination (i.e., bias transmitted through correlated or downstream variables) poses a subtle but critical risk for insurers, all the more since actuarial science heavily rely on model based on proxy variables. This presentation will explore how causal reasoning and counterfactual thinking can illuminate what standard machine learning methods often obscure. I will begin with intuitive economic decompositions of group disparities, then show how recent advances in optimal-transport-based counterfactuals enable us to ask: “What would this prediction have been if the sensitive attribute had been different?” Drawing on recent framework for causal mediation via sequential optimal transport , we will see how actuaries can break down total model disparities into direct and indirect components (even with complex mediators such as behavioral features, prior claims history, or categorical underwriting variables). The goal of the session is to leave the audience with a clear understanding of where hidden bias may emerge in insurance AI systems, how to diagnose it using modern causal tools, and how these insights support better governance, transparency, and compliance. A forward-looking, accessible presentation to close the day and open new perspectives on fair and responsible AI in insurance.
o (pero no daré mi charla en español) “Detección de Sesgos Ocultos en la IA de Seguros con Métodos Contrafactuales”
A medida que la inteligencia artificial se integra en la suscripción, la tarificación, la detección de fraude y la automatización de siniestros, surge un desafío que a menudo se subestima: los modelos pueden discriminar sin necesidad de utilizar explícitamente una variable prohibida. La discriminación indirecta —es decir, el sesgo que se transmite a través de variables correlacionadas o situadas aguas abajo— representa un riesgo sutil pero crítico para las aseguradoras, especialmente considerando que la ciencia actuarial depende fuertemente de modelos basados en variables proxy. Esta presentación explorará cómo el razonamiento causal y el pensamiento contrafactual pueden revelar aquello que los métodos tradicionales de machine learning suelen ocultar. Comenzaré con descomposiciones económicas intuitivas de las disparidades entre grupos y luego mostraré cómo los avances recientes en contrafactuales basados en transporte óptimo permiten formular la pregunta: “¿Cuál habría sido la predicción si el atributo sensible hubiera sido diferente?” Basándonos en marcos recientes de mediación causal mediante transporte óptimo secuencial, veremos cómo los actuarios pueden descomponer las disparidades totales de un modelo en componentes directos e indirectos, incluso cuando existen mediadores complejos como variables de comportamiento, historial de siniestros o características categóricas de suscripción. El objetivo de la sesión es ofrecer al público una comprensión clara de dónde pueden surgir sesgos ocultos en los sistemas de IA utilizados en seguros, cómo diagnosticarlos utilizando herramientas causales modernas, y cómo estos enfoques pueden fortalecer la gobernanza, la transparencia y el cumplimiento regulatorio. Una presentación accesible y orientada al futuro, ideal para cerrar la jornada e introducir nuevas perspectivas sobre una inteligencia artificial justa y responsable en el sector asegurador.
Actuarial Pricing Discrimination and Fairness
Tomorrow, I will give a talk for the “pricing seminar” of a major insurance company. The slides are available online, and it will be on “Actuarial Pricing Discrimination and Fairness“. As always, my talk will explore why actuarial pricing inherently relies on discrimination between groups, and how this raises deep conceptual, legal, and ethical challenges. The objective is to understand both the mathematical foundations of “fairness” and the regulatory tensions that emerge.

Continue reading Actuarial Pricing Discrimination and Fairness
Decomposing Direct and Indirect Biases in Linear Models under Demographic Parity Constraint
Our paper “Decomposing Direct and Indirect Biases in Linear Models under Demographic Parity Constraint“, with Bertille Tierny and François Hu is now online on ArXiv.
Linear models are widely used in high-stakes decision-making due to their simplicity and interpretability. Yet when fairness constraints such as demographic parity are introduced, their effects on model coefficients, and thus on how predictive bias is distributed across features, remain opaque. Existing approaches on linear models often rely on strong and unrealistic assumptions, or overlook the explicit role of the sensitive attribute, limiting their practical utility for fairness assessment. We extend the work of (Chzhen and Schreuder, 2022) and (Fukuchi and Sakuma, 2023) by proposing a post-processing framework that can be applied on top of any linear model to decompose the resulting bias into direct (sensitive-attribute) and indirect (correlated-features) components. Our method analytically characterizes how demographic parity reshapes each model coefficient, including those of both sensitive and non-sensitive features. This enables a transparent, feature-level interpretation of fairness interventions and reveals how bias may persist or shift through correlated variables. Our framework requires no retraining and provides actionable insights for model auditing and mitigation. Experiments on both synthetic and real-world datasets demonstrate that our method captures fairness dynamics missed by prior work, offering a practical and interpretable tool for responsible deployment of linear models.
On sera à Singapour pour le présenter fin janvier, à AAAI 2026, 40th Annual AAAI Conference on Artificial Intelligence.



