Factor In Your Expertise
Help shape how the actuarial profession evaluates AI.
The CAS Artificial Intelligence Working Group invites researchers and subject matter experts to bring their expertise to the development of a robust, actuarially grounded benchmarking framework designed to evaluate modern large language models (LLMs) on perception‑focused P&C tasks. With frontier models rapidly advancing, our goal is to create a transparent, repeatable way to understand how well these systems perform on problems that matter in actuarial practice and to keep re-testing that performance as new models are released until such time as models have been determined to have solved the benchmarks.
FACTOR IN YOUR EXPERTISE
Help shape how the actuarial profession evaluates AI.
2026 CAS Research Opportunity
Proposals Due: September 28, 2026
This initiative supports the CAS’s broader mission to responsibly explore AI technologies and equip members with tools and insights that can be practically applied across reserving, ratemaking, and related workflows.
The focus of this RFP is on perception tasks. These are tasks where the LLM's job is to correctly recognize, classify, or judge something against a ground-truth answer, not to generate open-ended text. The output can be evaluated objectively, because there's a defined correct answer to check against. This contrasts with generative tasks (drafting a reserve report narrative, writing client communications), where "correctness" is subjective/qualitative and harder to score reproducibly.
Research Problem
The actuarial profession currently lacks a standardized framework for evaluating LLMs on perception and classification tasks. These are problems with objectively measurable, reproducible answers (e.g., claims triage, underwriting decisions, fraud flagging, rating plans, regulatory issues, risk management, reserving, credibility). This distinguishes them from generative tasks, where "correctness" is harder to define.
Proposers should design a fixed, versioned suite of evaluation tasks that can be efficiently re-run as new models are released, allowing the CAS to track performance over time. If current models eventually solve all tasks, that result is itself valuable. The CAS will then introduce harder or new tasks to keep the benchmark meaningful.
Proposers are responsible for designing and implementing the full benchmarking ecosystem: task design, data assembly, evaluation protocols, and a public comparison platform that reflects real P&C actuarial and insurance challenges.
Proposal and Work Product Requirements
The selected project team will deliver an end‑to‑end benchmarking framework aligned with the following expectations.
- Develop a rigorous set of actuarially relevant perception and classification tasks, as well as drawing on existing open insurance benchmark datasets.
- Example (including but not limited to): Claims triage/classification, underwriting decision judgments, policy segmentation, and fraud or litigation likelihood flagging.
- Ideally, broad actuarial competencies are decomposed into granular, well-defined evaluation sub-tasks. For example, breaking "reserving" into loss development pattern recognition or claim severity classification.
- Proposals should find, assemble, or simulate their own datasets.
- All data must be fully legally usable and publishable as part of the CAS benchmark suite.
- Example papers and potential data sources:
- Example industry benchmarking initiatives: