2026 Request for Proposals: Evaluating LLMs for Actuarial Perception Tasks for an LLM Benchmark and Re-Evaluation Suite

Factor In Your Expertise

Help shape how the actuarial profession evaluates AI.

The CAS Artificial Intelligence Working Group invites researchers and subject matter experts to bring their expertise to the development of a robust, actuarially grounded benchmarking framework designed to evaluate modern large language models (LLMs) on perception‑focused P&C tasks. With frontier models rapidly advancing, our goal is to create a transparent, repeatable way to understand how well these systems perform on problems that matter in actuarial practice and to keep re-testing that performance as new models are released until such time as models have been determined to have solved the benchmarks.

FACTOR IN YOUR EXPERTISE

Help shape how the actuarial profession evaluates AI.

2026 CAS Research Opportunity

Proposals Due: September 28, 2026

This initiative supports the CAS’s broader mission to responsibly explore AI technologies and equip members with tools and insights that can be practically applied across reserving, ratemaking, and related workflows.

The focus of this RFP is on perception tasks. These are tasks where the LLM's job is to correctly recognize, classify, or judge something against a ground-truth answer, not to generate open-ended text. The output can be evaluated objectively, because there's a defined correct answer to check against. This contrasts with generative tasks (drafting a reserve report narrative, writing client communications), where "correctness" is subjective/qualitative and harder to score reproducibly.

Research Problem

The actuarial profession currently lacks a standardized framework for evaluating LLMs on perception and classification tasks. These are problems with objectively measurable, reproducible answers (e.g., claims triage, underwriting decisions, fraud flagging, rating plans, regulatory issues, risk management, reserving, credibility). This distinguishes them from generative tasks, where "correctness" is harder to define.

Proposers should design a fixed, versioned suite of evaluation tasks that can be efficiently re-run as new models are released, allowing the CAS to track performance over time. If current models eventually solve all tasks, that result is itself valuable. The CAS will then introduce harder or new tasks to keep the benchmark meaningful.

Proposers are responsible for designing and implementing the full benchmarking ecosystem: task design, data assembly, evaluation protocols, and a public comparison platform that reflects real P&C actuarial and insurance challenges.

Proposal and Work Product Requirements

The selected project team will deliver an end‑to‑end benchmarking framework aligned with the following expectations.

Task Design
  • Develop a rigorous set of actuarially relevant perception and classification tasks, as well as drawing on existing open insurance benchmark datasets.
  • Example (including but not limited to): Claims triage/classification, underwriting decision judgments, policy segmentation, and fraud or litigation likelihood flagging.
  • Ideally, broad actuarial competencies are decomposed into granular, well-defined evaluation sub-tasks. For example, breaking "reserving" into loss development pattern recognition or claim severity classification.
Dataset Assembly
Benchmark Suite Creation
  • Provide all task definitions, datasets, scoring methodologies, and evaluation scripts.
  • Scoring must rely on reproducible, objective metrics (e.g., accuracy, F1, Brier score, log loss).
Model Evaluation Framework
  • Implement standardized evaluation protocols for the top current commercial models from Google, Anthropic, and OpenAI.
  • Similarly evaluate at least 3 of the top open-source models.
  • The evaluation framework must be built for repeatable re-testing. As new frontier models are released, the CAS should be able to run the existing benchmark suite against them with minimal additional engineering effort.
Comparison Platform
  • Deliver an interactive, user-friendly public platform that compares model performance across tasks.
  • Build the platform as an extensible framework so new models and new benchmark tasks can be added over time.
  • Enable actuaries and researchers to propose additional evaluation tasks and datasets.
  • Allow model ratings to update dynamically based on newly submitted data and defined performance criteria.
Documentation, Reporting, and Ongoing Maintenance
  • A detailed written report summarizing methodology, model performance comparisons, and limitations.
  • The CAS will own and maintain the benchmark suite and comparison platform after delivery, including re-testing new models as they are released, adding new benchmark tasks over time, and platform upkeep.
  • Proposals must include documentation sufficient for the CAS to run and maintain the system independently, without requiring the contractor's continued involvement. This would include (but isn’t limited to) instructions, code, and configuration files.
  • Clear recommendations on maintenance, governance, and update cadence.
GitHub Publication Requirements

To ensure openness and reproducibility all datasets, code, benchmark definitions, scoring scripts, and documentation must be published to a CAS GitHub repository.

Final Deliverables

The final work must include:

  1. Comparison Platform: Public‑facing interface presenting model comparisons, built for repeatable re-testing against new models.
  2. Benchmark Suite: Datasets, task descriptions, scoring scripts.
  3. Model Evaluation Report: Comprehensive analysis of performance on perception/classification tasks.
  4. Executive Summary: An executive summary (1–2 pages) suitable for a blog post or magazine article, written for a non-technical audience and highlighting key insights and implications for actuarial practice.
  5. Replication Documentation: Instructions, code, configuration files.
  6. Maintenance Recommendations: Guidance for long‑term stewardship.
  7. Full publication to CAS GitHub.

Submitting Proposals

Interested researchers should submit:

  • A detailed outline of their proposed solution
  • Dataset, data generation, or data simulation strategy
  • Project approach and methodology
  • Budget and pricing structure
    • Estimated out-of-pocket expenses (e.g., cloud storage, LLM API usage, fine-tuning models, software licenses, etc.) required to complete the work.
    • Expected compensation for labor.
    • Do not include overhead costs.
  • Team qualifications and relevant experience, including resumes of the researcher(s), indicating how their background, education, and experience demonstrate their qualifications to undertake the research.
  • Project plan and milestones

Proposal Deadline: September 28, 2026

Interested researchers should submit their proposals to CAS Research Manager Heather Davis at hdavis@casact.org and CAS Director of Research and Publications Elizabeth Smith at esmith@casact.org by September 28, 2026, with the subject line: “LLM Actuarial Benchmark RFP Submission.”

Questions may be directed to hdavis@casact.org and esmith@casact.org. Please provide all questions by September 14, 2026.

Receipt of proposals will be acknowledged in a timely manner. Respondents who are not awarded the contract will be informed shortly thereafter.

A CAS contract will be awarded to the respondent who, in the judgment of the CAS Artificial Intelligence Working Group, is best able to perform the work as specified. If the group determines that no proposal meets the requirements of the RFP, no contract will be awarded.

Selection Process

The CAS AI Working Group will evaluate proposals using the following criteria:

  • Understanding of actuarial perception and classification tasks
  • Technical competency in AI and benchmarking
  • Feasibility and clarity of approach
  • Quality of benchmark design
  • Strength of team qualifications
  • Cost effectiveness
  • Ability to provide reproducible, transparent results

If no proposal meets the requirements, the CAS may decline to award a contract.

Compensation

The award amount will most likely fall in the $25,000 to $65,000 range, although requests outside this range may be considered if supported by exceptional scientific merit and a well-justified budget.  Applicants are encouraged to request only the funding necessary to complete the proposed work.  We will consider budget appropriateness and the efficient use of CAS resources during proposal evaluation.

  • The maximum budget for this project is $75,000.
  • Proposers should include detailed pricing for the work as specified, along with an estimate of any ongoing costs the CAS should expect to bear when maintaining the system independently. Note this is separate from the funding covered by this RFP.
  • A portion of the funding should be used to cover usual and customary travel expenses to present the paper at a CAS-sponsored seminar or meeting.
  • CAS research funding does not cover overhead expenses but can cover salary and fringe benefits.
  • Compensation will be commensurate with the time required to carry out the work.
Project Oversight and Publication
  • As a condition of selection, the CAS requires that all rights, title, and interest, including copyright and patent, in and to the report be owned by the CAS. The selected researcher(s) must sign a formal research agreement that assigns all such rights to the CAS.
  • In any publication of the report, the researcher(s) will receive appropriate authorship credit. The CAS may publish the report in its entirety, or any sections thereof, in any format and medium as it finds fit, including, but not limited to CAS publications, and electronic versions on its website or physical storage media. Publishing outside the CAS requires permission from the CAS, and the authors are to acknowledge previous publication.
  • If the author publishes the final draft paper on a pre-print platform such are ArXiv, the CAS must be acknowledged as the funder of the research.
  • The researcher(s) should make every effort to be available to present their findings and share their expertise at a CAS meeting or seminar.
  • The project will be overseen by the CAS AI Working Group with regular checkins.
Paper Requirements
  • To aid research adoption, the final work product’s code and data will also be placed in the CAS’s GitHub repository, under the MPL2.0 license. The author should include a link in the body of the final paper to the CAS Research Paper Repo here: https://github.com/casact/research-papers.
  • Authors must report specific use of artificial intelligence, if any, while producing research.
  • All tables and figures must be numbered and include an in-text callout.
  • Manuscripts must include an abstract with a maximum word count of 150 words.
  • References must use the author-date system from Chicago Manual of Style, 18th edition.
  • Manuscripts should include a list of 2–10 relevant keywords to help readers discover the content online.
Timeline

Proposals Due

September 28, 2026

Researchers Notified

October 26, 2026

Kick-Off Meeting with Project Oversight Group

TBD

Executive Summary Due

November 23, 2026

Draft 1 Due

TBD

Draft 2 Due

TBD

Final Paper Due

June 28, 2027

About the Casualty Actuarial Society and the Artificial Intelligence Working Group

For over 100 years, the Casualty Actuarial Society (CAS) has been the trusted global authority advancing the practice of property and casualty (P&C) actuarial science, with nearly 12,000 members worldwide who apply their expertise to help people, businesses, and communities unlock opportunities and thrive in a rapidly changing world. The CAS delivers top-tier credentialing, cutting-edge research, and dynamic learning and professional development, grounded in real-world application and powered by a vibrant global member community. The CAS equips actuaries to advance their careers and deliver innovative, trusted solutions to complex and emerging P&C challenges. Learn more at casact.org.

The CAS Artificial Intelligence Working Group was established to fulfill the CAS mission to “advance the body of knowledge” on a technology that is transforming actuarial practice. Its objective is to encourage the exploration of Artificial Intelligence in actuarial practice through research to help educate members, build knowledge, provide practical insight and establish CAS as thought leaders.