Category Archives: Research

7th European Actuarial Journal Conference in Istanbul

Next week, I will attend the 7th European Actuarial Journal Conference in Istanbul. I will present our joint paper Direct and Indirect Discrimination in Generalized Linear Models,

Generalized linear models are central to actuarial modelling of binary risk, claim frequency, utilization, and cost-related outcomes. Yet fairness diagnostics often rely on linear-model intuitions, although GLM predictions are obtained by transporting a latent score through a nonlinear inverse link. We develop a moment-based decomposition framework for diagnosing group disparities in fitted GLM predictions. In an exact linear-Gaussian benchmark, the Wasserstein barycentric criterion for distributional demographic-parity violation reduces to a two-moment criterion and decomposes into direct mean, indirect mean, interaction, and structural components. For GLMs, we distinguish the empirical output-scale criterion U_2(f), a within-group proxy \tilde{U}_2(f), and a leading decomposition D_1(f). This leading term preserves the four linear channels and adds two curvature components induced by the inverse link: curvature coupling and curvature amplification. We derive explicit formulas for logistic, Poisson, and Tweedie specifications and illustrate the diagnostic on medical-expenditure survey data. The framework is not a legal test of discrimination, nor a full characterization of distributional parity outside the linear-Gaussian case. It is a tractable actuarial diagnostic for identifying whether fitted prediction disparities arise from explicit sensitive effects, proxy-mediated covariate profiles, covariance-structure differences, or nonlinear link effects.

Slides are now available.

Conference Climate Change and Insurance 2026 in Spain

I will be in Spain, at the Conference Climate Change and Insurance 2026, to present “Granular Pricing and Effective Withdrawal in Climate-Exposed Household Insurance in France“, written with Raphaël Dalbarade, Laurence Barry and Caroline Hillairet.

Climate change increases the value of fine-scale risk information in household insurance, but its competitive use may weaken pooling and reduce effective availability in exposed areas. We study this mechanism in the French Cat-Nat system, where natural-catastrophe coverage is formally pooled through a regulated surcharge but accessed through the underlying household-insurance contract. First, we use controlled online quote requests to compare otherwise identical addresses with different flood or clay shrink-swell exposure. The evidence shows heterogeneous insurer responses: coarse pooling, within-commune premium differentiation, and non-quoting in some exposed micro-locations. Second, we develop a stylized dynamic pricing game with insurers that differ in pricing granularity. Numerical equilibria show that fine segmentation attracts low-risk households, shifts high-risk households toward less granular insurers, and erodes implicit cross-subsidies. Market-share constraints mainly slow this reallocation. The results highlight why climate-insurance governance must monitor local availability, not only formal coverage or aggregate market presence.

Slides are available online.

IFoA AI Ethics Governance & Risk Management forum

I was invited to give a talk at the “IFoA AI Ethics Governance & Risk Management forum” at the end of the week. I uploaded some slides to start the discussion.

I will not give a highly technical lecture, but rather to propose a way of thinking about fairness in insurance AI that may be useful for actuaries, for governance people, and for anyone involved in model risk or oversight. The main message is simple. In insurance, fairness cannot be reduced to predictive accuracy. And once we move to modern AI and large data environments, simply removing a sensitive variable is no longer enough to prevent discrimination. So the real question becomes: how should we govern the trade-offs between fairness and accuracy, and different notions of fairness?

Let me start with a very simple example, illustrating the fact that simply removing a sensitive variable is not enough to prevent discrimination. Here I take a French motor insurance dataset. In the raw data, average claim frequencies are about 8.94% for men and 8.20% for women. Now suppose that I estimate a logistic model for annual claim frequency, but I explicitly exclude gender from the explanatory variables. At first sight, this looks like the fair thing to do. I am not using gender, so maybe I am safe. But what happens?

As I add more and more variables to the model, the predicted frequencies for men and women move closer and closer to the empirical frequencies in the data. With enough explanatory variables, I am basically reconstructing the original gap. So the model is not using gender explicitly, but it is learning from variables that are correlated with gender. In other words, the information has not disappeared. It has simply been redistributed across other variables.

That is why I think the old defence of ‘we do not use the sensitive variable, therefore the model is fair’ has become much weaker in a big-data context (and here, it’s not big, it’s only 15 basic explanatory variables). Modern predictive models are very good at finding statistical associations. If protected traits leave traces in the data, the model will usually pick them up.

And this is not only a technical issue. It is also a philosophical one. The model tends to reproduce what is in the data, that’s what we call ‘generalization’ in machine learning. If historical inequalities are present in the data, the model will often learn them very efficiently. This is why I mention Hume here, and the old ‘is-ought’ problem. From the fact that a disparity exists in the data, it does not follow that it should be reproduced in pricing. So this slide is really the starting point for the rest of the talk: fairness through unawareness is often not fairness at all; it is frequently a way of hiding the channel through which unfairness reappears.

So once we accept that unawareness is not enough, we need to ask a more difficult question: what exactly do we mean by fairness?

Before going to the recent actuarial papers, I want to spend one minute on a classical dilemma. This is not new. In legal and philosophical debates, there has long been a tension between two positions.

One position says: if you want to correct unequal outcomes, you may need to take account of the protected characteristic. That is the logic behind many discussions of affirmative action or corrective treatment. If the world is already unequal, pretending not to see the relevant characteristic may simply reproduce the inequality. The opposite position is the so-called colorblind view: the way to stop discrimination is to stop using the sensitive characteristic altogether. That sounds very appealing, because it seems neutral and simple.

The trouble is that in insurance, and especially in AI-based pricing, both positions run into difficulties. If you ignore the sensitive variable, you may still reproduce its effects through proxies, that’s what we’ve seen before. But if you use it for corrective purposes, you are entering a difficult space involving direct differential treatment, redistribution, and legal or ethical contestability.

So what I want to emphasize here is that the fairness debate is not just about technical metrics. It is rooted in a much older normative tension: should fairness mean treating everyone under the same formal rule, or should it sometimes mean treating people differently in order to offset structural inequalities? I think that this dilemma is exactly why actuarial fairness debates cannot be settled by model performance alone.

In the contexte of insurance, the recent SOA report discussed in a previous meeting stresses this important tension fairness is not one thing.

The report is useful because it is very practical and very clear. Its main point is that fairness in life insurance is not a single concept. There is no universal fairness metric that would solve the problem once and for all. Instead, the paper distinguishes between different families of fairness criteria.

A first distinction is between individual fairness and group fairness. Individual fairness is close to the actuarial intuition: similar risks should be treated similarly. Group fairness, by contrast, is about parity across groups. Those two perspectives can both sound reasonable, but they are often in tension. The paper also reviews several criteria. On the individual side, it discusses unawareness, awareness, and omitted-variable bias. On the group side, it discusses independence, sufficiency, and separation.

The most useful takeaway is not any particular metric. It is the methodological lesson: before choosing a metric, you must define what notion of fairness you are trying to serve. Otherwise you are optimizing a number without knowing what ethical or institutional objective it really represents. The SOA paper gives us a map. It tells us: do not look for the one true fairness metric. Clarify the objective first, and then be explicit about the trade-offs.”

Our own recent work tries to push that argument one step further, by asking what the core dimensions of fairness in insurance pricing really are…

In our recent paper on what we call, with Olivier Côté, and Marie-Pier Côté, “the fairness trilemma“, we recall the idea that insurance is not just a predictive exercise. It is also a social institution. Insurance has a double nature. On the one hand, it is about risk-based pricing: aligning premiums with expected losses. On the other hand, it is also about risk-sharing: spreading burdens, preserving access, and sometimes accepting cross-subsidies. From there, we argue that fairness in insurance pricing is governed by a trilemma between three principles.

  • The first is actuarial fairness. In simple terms, this means that premiums should reflect expected losses as accurately as possible.
  • The second is social solidarity. This means that we may accept departures from strict risk-based pricing in order to preserve access, affordability, or some broader social objective.
  • The third is causal legitimacy. This is extremely important in the AI context. Not every predictive variable is equally legitimate. A variable may improve predictive performance, but if it works mainly by reconstructing a protected characteristic, or if its causal meaning is questionable, then its legitimacy is fragile. So causal legitimacy asks not only whether a factor predicts, but whether it deserves to shape the price.

The key point is that these three principles are all attractive, but they cannot be fully satisfied at the same time. If I strengthen actuarial fairness, I may weaken solidarity. If I enforce more solidarity, I move away from pure risk adequacy. If I become very demanding about causal legitimacy, I may sacrifice some predictive performance and perhaps some segmentation logic. So the problem is not to find the perfectly fair model. The problem is to govern the trade-offs. That is why I insist on governance tools: scorecards, causal due diligence, comply-or-explain procedures. The point is to make choices explicit, reviewable, and contestable. In that sense, fairness is not just a modelling problem. It is a governance problem.

But of course, at this stage, someone in the room may say: all this sounds interesting, but how do we make it operational? How do we move from philosophy to something an actuary can actually measure?

That is exactly the purpose of the CAS paper. The starting point is again that removing a sensitive variable does not remove discrimination, because allowed variables may still act as proxies. So the challenge is to expose indirect discrimination in a way that is usable in actuarial practice. The objective of the paper is to make fairness operational and measurable in actuarial terms. It is not enough to speak in abstract statistical language. In insurance, we need to know what unfairness means in premiums, in dollars, and across segments of policyholders.

The paper is based on a real auto insurance case study from Québec, using credit score as the sensitive variable. And it again organizes the discussion around three dimensions: actuarial fairness, social solidarity, and causality legitimacy. This is important because it shows continuity with our more conceptual framework. The toolbox is not something separate. It is a way of translating those dimensions into measurable diagnostics. If the trilemma paper says fairness must be governed, the CAS paper says: here are some tools that help you see what you are governing.

One of the most interesting contributions of the toolbox is that it introduces local metrics such as risk spread, proxy vulnerability, fairness range, and parity cost. Let us focus on two of them.

The first is proxy vulnerability. Intuitively, proxy vulnerability measures how much a segment may be over-priced or under-priced because apparently neutral variables indirectly reconstruct the sensitive attribute. This is very valuable because it moves us away from vague concerns about bias and toward a concrete actuarial question: where, and by how much, might the pricing rule be unfair?

The second is parity cost. This is also very important because it forces us to be honest. If we want more parity across groups, that is not free. Someone has to bear the cost. In plain English, enforcing one notion of fairness may end up robbing Peter to pay Paul. And I do not say this as a criticism; I say it because redistribution should be explicit rather than hidden.

Another strength of the toolbox is that it helps actuaries identify local pockets of unfairness. A model may look acceptable on average across broad groups, but still be very problematic in specific vulnerable niches. So fairness should not be assessed only globally. Local diagnostics matter.

This is also, I think, very relevant for AI governance. Senior management and boards do not just need a statement saying ‘the model passed a fairness test’. They need to know where the model is fragile, which subpopulations are exposed, what the monetary magnitude is, and what the trade-offs are if a correction is introduced. So in practice, the toolbox helps turn fairness into something discussable in risk committees and governance processes.

To wrap up, first, the broader background for all this is developed in my recent textbook, which tries to connect legal, philosophical, statistical, and actuarial perspectives on discrimination in insurance.  Second, fairness should not be assessed only at the level of one insurer’s portfolio. A pricing rule may look fair on one portfolio and still be unfair at the market level, because portfolios are not representative of the insured population. This is a selection-bias problem. If different insurers attract different segments of the market, portfolio-specific fairness does not necessarily aggregate into market-wide fairness. So we also need to ask: fair for whom, and on which reference population?

Third, there is always a cost and a redistribution issue. If we enforce one notion of fairness, someone pays. That does not mean fairness is undesirable. It means that fairness choices are unavoidably political, institutional, and governance choices, not just technical ones. And fourth, there is a major practical difficulty when sensitive variables are unobserved. Very often, people say: let us not collect the sensitive variable. But then how do we audit fairness? How do we provide evidence? How do we challenge the model? If the protected attribute is precisely the thing we do not observe, fairness becomes not just a modelling issue, but a data, audit, and governance issue.

So my overall conclusion would be the following. AI did not create the fairness problem in insurance, but it has made it sharper. It has weakened the old comfort of unawareness. It has increased the power of proxies. And it has made the trade-offs more difficult to hide. For actuaries, I think this creates both a risk and an opportunity. The risk is to treat fairness as a box-ticking exercise or as a purely technical constraint. The opportunity is to contribute something distinctive: a language of prices, trade-offs, portfolios, and governance. So perhaps the right ambition is not to promise a perfectly fair model. It is to build institutions that are capable of identifying proxy discrimination, making normative choices explicit, and governing them responsibly.

Will Technology Save Us?

This post was initially written in French, La technologie nous sauvera-t-elle ?

I feel as though I keep hearing, more and more often, that climate change is above all a problem of innovation. Emissions continue to rise, targets keep slipping out of reach, yet we fill our collective imagination with carbon-capture machines, artificial intelligences supposedly able to optimize the transition, and even technologies designed to alter the climate itself. There is nothing absurd about such confidence in itself, and technology will very likely help. But it becomes politically suspect when it serves mainly to postpone difficult questions, beginning with this one: what are we willing to change, here and now, in the way we produce, consume, and govern? The literature on “mitigation deterrence” helps us understand how the promise of a future intervention can legitimize delaying present efforts. Nor is this mechanism unique to climate change. The COVID pandemic, it seems to me, offered a strikingly similar scene, in which fascination with the biomedical response sometimes pushed into the background the social, institutional, and political tools that the strongest research nevertheless regarded as indispensable.
Continue reading Will Technology Save Us?

La technologie nous sauvera-t-elle ?

J’ai l’impression d’entendre de plus en plus souvent dire que le climat est avant tout un problème d’innovation. Les émissions continuent d’augmenter, les objectifs se dérobent, mais on essaye de peupler notre imaginaire collectif de machines à capturer le carbone, d’intelligences artificielles capables d’optimiser la transition, voire de techniques de modification du climat. Cette confiance n’a rien d’absurde en elle-même, et il y a fort à parier que les technologies aideront. Mais elle devient politiquement suspecte lorsqu’elle sert avant tout à repousser des questions difficiles, à commencer par “que sommes-nous prêts à changer, ici et maintenant, dans nos manières de produire, de consommer et de gouverner ?” La littérature sur la “mitigation deterrence” a permet de mieux comprendre la promesse d’une intervention future légitimant de retarder l’effort présent. Ce mécanisme n’est pas propre au climat, il me semble que la pandémie de COVID nous a offert une scène assez proche, où la fascination pour la réponse biomédicale a parfois relégué au second plan les instruments sociaux, institutionnels et politiques pourtant jugés indispensables par les travaux les plus solides .
Continue reading La technologie nous sauvera-t-elle ?

Exposé “Risque climatique, retrait des assureurs et granularité des tarifs” pour la Chaire PARI

Mercredi, je donnerai la première partie de l’exposé Risque climatique, retrait des assureurs et granularité des tarifs, organisé par la Chaire PARI. Je donnerai un point de vue un peu général sur le problème qui nous préoccupe, à savoir la modélisation d’un marché concurrentiel d’assurance, et la recherche de politiques optimales, pour un régulateur, pour que l’équilibre concurrentiel soit optimal (ou a minima améliore certains critères) pour le bien être global. Raphaël Dalbarade présentera ensuite ses travaux sur le sujet.

Faut-il socialiser les risques ou responsabiliser les territoires ?

Publication d’un court article, écrit avec Laurence Barry, en ligne sur le site de la Revue Banque,

Le dérèglement climatique, qui s’accompagne de l’intensification des phénomènes extrêmes, prend de l’ampleur à un moment où les données disponibles concernant ces événements se multiplient à une maille de plus en plus fine. De plus, dans certains pays, et notamment en France, des stress-tests climatiques mis en place ces dernières années ont contribué à une montée en capacité des compagnies d’assurance sur ces modèles.

(à suivre…)

International Workshop on Risk and Insurance, 서울, June 2026

On June 29th, I will be in Seoul (서울), Korea, at the International Workshop on Risk and Insurance.

This workshop aims to provide a focused forum, where global risk and insurance research and the Korean insurance industry can exchange ideas and discuss the practical implications of emerging risks and technologies. The program will feature leading scholars and industry experts discussing key topics shaping the future of insurance, including:
• AI Revolution and Cyber Risks
• Climate Change and Extreme Weather Events
• Insurance Data Science and Market Innovations

The workshop will be held at FKI Tower (Diamond Hall) in the heart of Seoul’s financial district. It will include academic presentations, industry panel discussions, and networking opportunities designed to foster collaboration between researchers and practitioners. The website of the workshop for registration (note that registration is free but space is limited) is now online.

Decomposing Probabilistic Scores

Our paper Decomposing Probabilistic Scores: Reliability, Information Loss and Uncertainty, with Agathe Fernandes-Machado, is now available https://doi.org/10.48550/arXiv.2603.15232

Calibration is a conditional property that depends on the information retained by a predictor. We develop decomposition identities for arbitrary proper losses that make this dependence explicit. At any information level \mathcal{A}, the expected loss of an \mathcal{A}-measurable predictor splits into a proper-regret (reliability) term and a conditional entropy (residual uncertainty) term. For nested levels \mathcal{A}\subset\mathcal{B}, a chain decomposition quantifies the information gain from \mathcal{A} to \mathcal{B}. Applied to classification with features \boldsymbol{X} and score S=s(\boldsymbol{X}), this yields a three-term identity: miscalibration, a {\em grouping} term measuring information loss from \boldsymbol{X}  to {S}, and irreducible uncertainty at the feature level. We leverage the framework to analyze post-hoc recalibration, aggregation of calibrated models, and stagewise/boosting constructions, with explicit forms for Brier and log-loss.

Sequential Transport for Causal Mediation Analysis

Our paper, Sequential Transport for Causal Mediation Analysis, with Agathe Fernandes-Machado, Iryna Voitsitska and Ewen Gallic, is now available on https://arxiv.org/abs/2603.15182

We propose sequential transport (ST), a distributional framework for mediation analysis that combines optimal transport (OT) with a mediator directed acyclic graph (DAG). Instead of relying on cross-world counterfactual assumptions, ST constructs unit-level mediator counterfactuals by minimally transporting each mediator, either marginally or conditionally, toward its distribution under an alternative treatment while preserving the causal dependencies encoded by the DAG. For numerical mediators, ST uses monotone (conditional) OT maps based on conditional CDF/quantile estimators; for categorical mediators, it extends naturally via simplex-based transport. We establish consistency of the estimated transport maps and of the induced unit-level decompositions into mutatis mutandis direct and indirect effects under standard regularity and support conditions. When the treatment is randomized or ignorable (possibly conditional on covariates), these decompositions admit a causal interpretation; otherwise, they provide a principled distributional attribution of differences between groups aligned with the mediator structure. Gaussian examples show that ST recovers classical mediation formulas, while additional simulations confirm good performance in nonlinear and mixed-type settings. An application to the COMPAS dataset illustrates how ST yields deterministic, DAG-consistent counterfactual mediators and a fine-grained mediator-level attribution of disparities.

Mesurer l’équité “globale” quand les données sont dispersées

Avec Agathe Fernandes Machado, Olivier Côté et François Hu, on a mis en ligne un papier, Federated Measurement of Demographic Disparities from Quantile Sketches. Le point de départ est assez simple. On peut imaginer qu’un modèle de score (mesurant un risque de récidive, une probabilité de réadmission à l’hôpital, un score de crédit…) soit déployé dans plusieurs institutions : hôpitaux, tribunaux, banques, assureurs. Chacun collecte ses données, utilise un modèle, et conserve jalousement ses bases. Parfois par obligation légale (RGPD, secret médical), parfois par contraintes techniques, parfois par réticence organisationnelle. Le problème, c’est qu’un régulateur veut savoir si le score est discriminatoire, sans jamais centraliser les données brutes. En fait, c’est assez réaliste comme situation, beaucoup d’objectifs de justice algorithmique étant définis au niveau de la population, et pas localement. Les régulateurs et les directions conformité demandent : “Est-ce que le système, globalement, traite de la même façon les groupes protégés ?” Pas : “Chaque institution, isolément, a-t-elle l’air correcte ?” On montre dans notre article que des audits locaux peuvent être rassurants tout en étant trompeurs, parce que l’injustice peut naître précisément de ce que l’on ne voit pas en restant silo par silo. Et la bonne nouvelle, c’est qu’on peut estimer l’inéquité globale avec une communication très limitée, en demandant à chaque silo seulement des comptages et quelques quantiles de ses scores.

Dans le papier, on identifie deux sources majeures de décalage entre l’audit local et l’audit global.

Les effets de composition : une “version fairness” du paradoxe de Simpson. Même si chaque silo semble traiter les groupes de manière similaire, la répartition des groupes entre silos peut être très différente. Un groupe peut être sur-représenté dans certains hôpitaux, certains tribunaux, certaines zones géographiques… Et si ces silos n’ont pas le même profil de scores (parce que les populations, les pratiques ou les contextes diffèrent), alors l’agrégation peut créer un écart global qui n’apparaît nulle part localement.

L’hétérogénéité inter-silos : la “stratification cachée”  Deux silos peuvent produire des scores d’allure différente (distribution plus “optimiste”, plus “pessimiste”, plus dispersée…), même au sein d’un même groupe sensible. Localement, chacun peut avoir des métriques acceptables. Mais une fois les données mises bout à bout, ces différences deviennent visibles et peuvent amplifier une disparité entre groupes. Dans les domaines sensibles (santé, justice pénale), cette hétérogénéité est courante : pratiques de codage, accès aux soins, critères de triage, politiques locales…

D’un point de vue pratique, on propose un protocole d’audit en un seul aller-retour : chaque silo envoie, pour chaque groupe sensible le nombre d’individus dans ce groupe (un simple comptage), k quantiles du score (par exemple k = 25, 50 ou 100), sur une grille commune. Et c’est tout. Pas de scores individuels, pas de features, pas d’exemples. Ce genre de résumé est déjà produit par de nombreux systèmes de monitoring via des quantile sketches utilisés pour suivre des distributions. À partir de ces quantiles, le serveur peut reconstruire une approximation des distributions globales par groupe, puis calculer la disparité populationnelle. Théoriquement, l’erreur due à la discrétisation décroît comme 1/k : plus on envoie de quantiles, plus la courbe reconstruite est fine. Et on montre sur des données réelles que quelques dizaines suffisent.

En bonus, on propose aussi une méthode permettant de comprendre, si on mesure un écart, pourquoi il apparaît. On obtient en particulier une décomposition de type ANOVA qui sépare : une part due aux effets de mélange / composition (le “Simpson fairness”), une part due à la vraie hétérogénéité inter-silos (différences structurelles de score), et un terme d’interaction qui peut amplifier ou compenser (mais reste contrôlé).

Bref, on montre qu’en environnement fédéré, l’équité populationnelle n’est pas la moyenne de l’équité locale. Elle dépend des mélanges, des flux, des biais d’affectation et des variations inter-silos. Donc la bonne question n’est pas “chaque silo est-il juste ?”, mais “le système fédéré, en tant que mécanisme de production de scores, est-il juste au niveau population ?” La bonne nouvelle, c’est qu’on peut répondre à cette question sans centraliser les données, en ne partageant qu’une poignée de quantiles et des comptes, en une seule communication.

Beyond Procedure: Substantive Fairness in Conformal Prediction

Our paper, Beyond Procedure: Substantive Fairness in Conformal Prediction, with Pengqi Liu, Zijun Yu, Mouloud Belbahri, Masoud Asgharian, and Jesse Cresswell, is now available on https://arxiv.org/abs/2602.16794

Conformal prediction (CP) offers distribution-free uncertainty quantification for machine learning models, yet its interplay with fairness in downstream decision-making remains underexplored. Moving beyond CP as a standalone operation (procedural fairness), we analyze the holistic decision-making pipeline to evaluate substantive fairness-the equity of downstream outcomes. Theoretically, we derive an upper bound that decomposes prediction-set size disparity into interpretable components, clarifying how label-clustered CP helps control method-driven contributions to unfairness. To facilitate scalable empirical analysis, we introduce an LLM-in-the-loop evaluator that approximates human assessment of substantive fairness across diverse modalities. Our experiments reveal that label-clustered CP variants consistently deliver superior substantive fairness. Finally, we empirically show that equalized set sizes, rather than coverage, strongly correlate with improved substantive fairness, enabling practitioners to design more fair CP systems. Our code is available at this https URL.

Balance and Calibration of Probabilistic Scores: From GLM to Machine Learning

Tomorrow, I will give a talk on “Balance and Calibration of Probabilistic Scores:“” From GLM to Machine Learning” at Singapore campus – ESSEC Asia-Pacific. The abstract is

This study evaluates binary classifier performance with a focus on calibration, which is often overlooked by traditional metrics like accuracy. In high-stakes domains such as finance and healthcare, well-calibrated probabilities are crucial. We highlight the limitations of standard calibration metrics, particularly under score distortions and heterogeneous distributions. To address this, we introduce the Local Calibration Score and advocate optimizing models using Kullback-Leibler (KL) divergence to better align predicted scores with true probabilities. Our approach emphasizes balancing global and local calibration, ensuring overall distributional alignment while maintaining reliability across different score ranges. Using Random Forest and XGBoost across diverse datasets, we show that KL-based tuning improves calibration without sacrificing performance. Our results reveal that relying solely on traditional metrics can mislead model assessment, especially in sensitive decision-making scenarios. This is some joint work with Agathe Fernandes Machado and Ewen Gallic.

Modeling and Understanding Indirect Discrimination in Algorithmic Fairness

In a couple of days, I will give a talk on “Modeling and Understanding Indirect Discrimination in Algorithmic Fairness” at Singapore campus – ESSEC Asia-Pacific. The abstract is

Observed disparities between groups in algorithmic decisions (whether in hiring, credit approval, or risk prediction) do not necessarily imply direct discrimination. They may also stem from legitimate differences in the distribution of explanatory attributes. Understanding and quantifying which components of these gaps are “explained” versus those that reflect direct or indirect discrimination lies at the core of modern causal approaches to algorithmic fairness. This talk will begin with an accessible introduction to group-gap decomposition, building on the classical Kitagawa–Oaxaca–Blinder econometric framework. This approach separates differences attributable to observable characteristics from residual components that may signal discriminatory effects. The second part will introduce recent developments leveraging optimal transport to construct individual-level counterfactuals, enabling estimation of direct and indirect causal effects for each observation. In particular, we will show how sequential transport mappings aligned with a causal graph can disentangle pathways and quantify the contribution of each mediator. This methodology overcomes limitations of traditional linear models, introduced by Kitagawa, Oaxaca and Blinder, provides interpretable counterfactuals, and is well suited to complex empirical settings. The presentation will combine intuitive motivation, illustrative examples, and recent research insights, with the goal of making these tools accessible and useful to researchers in management science, applied economics, and data science.

Talk at NTU (Nanyang) in Singapore

Tomorrow, I will be at Nanyang Technological University to give a talk at an internal seminar, “Fairness and discrimination in insurance

What’s unique about insurance is that even statistical discrimination, which by definition is devoid of malicious intent, poses significant challenges. Because, on the one hand, policymakers would like insurers to treat their policyholders equally, without discrimination based on race, gender, age or other characteristics, even if it could make (statistical) sense to (indirectly) discriminate. On the other hand, at the core of actuaries’ activities lies discrimination, between risky and non-risky policyholders. And this risk is often statistically correlated with sensitive characteristics that regulation would like to prohibit insurers from taking into account. The analysis of possible discrimination in decision rules, whether human or algorithmic, is an old subject. Most of the concepts date back at least to the 50s, but recent developments in artificial intelligence have brought these issues back into the spotlight. Massive data facilitate statistical or proxy discrimination, and black-box algorithms do not facilitate understanding. Not to mention the various regulations that make it difficult to collect sensitive information, and ultimately test whether decisions can be discriminated against, especially indirectly.

The talk is based on the textbook Insurance, Biases, Discrimination and Fairness, as well as recent papers, arXiv:2511.11294 (AAAI’26), arXiv:2408.03425 (AAAI’25), arXiv:2309.06627 (AAAI’24) and arXiv:2306.12912  (ECML’24).