Benchmarking
Underwriting Advice Engine settings and LLM monitoring for each case of use
The problem — To produce confident and traceable advice, each sub-industry and case of use requires a set of data tailored for the LLM to get the adequate context. It needs to be manually verified by experts and the outcome needs to match the expected quality.
The solution — A monitoring admin solution for experts to feedback and managers to test and verify the LLM trust.
The approach — A built-in interface enables managers to create 'Insurance Models' by loading documents that are reviewed by experts and then compared to the RAG AI output. This process is used to establish a 'Golden Standard' where the LLM output matches at least 90% of human judgement.
Tools used
Benchmarking
To-Be/As-Is Scenarios
Iterative Prototyping
Prototypes
Wireframe sample