Banks Adopt Automated Testing for GenAI, Yet Hesitate on Autonomous Sign-off
Approximately 30% of financial institutions have implemented partial automation for testing generative artificial intelligence (GenAI) models, according to data from the Risk Benchmarking Model Risk study. While firms are increasingly using automated tools to manage the risks associated with large language models (LLMs), most banks continue to rely on human oversight for final model validation, stopping short of allowing AI to facilitate autonomous sign-off.
The Current State of AI Model Validation
The integration of GenAI into banking operations has forced risk management teams to rethink traditional validation frameworks. By using “LLM-as-a-judge” frameworks, where one AI model evaluates the outputs of another, banks can scale their testing efforts to handle the high volume of content generated by enterprise-grade models.
Why Autonomous Sign-off Remains Rare
Despite the efficiency gains offered by automated testing, few banks are comfortable delegating the final approval process to machines. The reluctance to grant autonomous sign-off stems from several critical factors:
- Regulatory Uncertainty: As supervisory expectations for GenAI evolve, banks are prioritizing safety and compliance over full-scale automation of the governance lifecycle.
Risk Benchmarking: Industry Adoption Trends
The data highlights a clear divergence in how banks approach model risk.
Strategic Outlook for Model Risk Management
This allows institutions to maintain a human-led sign-off process while benefiting from the speed of automated detection.
Related reading