Key takeaways
- Much of the compliance and procurement software available today uses predictive models, which is great for meeting the goals of that kind of software, but those predictive models are dated relatively quickly and require continuous and streamlined monitoring, as well as being updated on a scheduled basis.
- Legacy banking systems present challenges to connect with, as they typically need custom APIs and assistance from their vendors to get a connection.
- The temptation of a PoC can blind you to production readiness; think about the costs associated with the data pipeline and compliance only after you are in production.
- Vendor lock-in is a very real threat: all it takes is increased pricing or obsolescence of the AI model being used, both which can create problems for your business.
Every few months or so, we hear about the new platform that will complete your compliance work for you, detect fraud within seconds, or upend your entire financial operation. However, what appears to be “too good to be true” in a demo usually has a much harder time getting adopted into the banking system upon deployment.
We can identify this by looking at what other researchers have discovered across multiple industries: Gartner recently reported 85% of all advanced analytics projects never make it into production due to banks typically treating AI as just another piece of software rather than as an area of expertise.
This has been seen many times in my experience working with various organizations. Your business stakeholders may appreciate the user interface and the way the product looks but tend to overlook how the system works under the surface, leading to an overall loss of revenue from their business.
Why standard financial procurement fails with machine learning
CRMs follow predefined rules and produce predictable results. Conversely, an AI system learns from multiple data sources and behaves in unexpected ways, which allows it to provide different results for the same input. Therefore, you will always need manual oversight of your AI systems.
If you evaluate a potential AI solution using the same criteria that you would for any other cloud-based service or application, you could make a less-than-optimal decision. You must evaluate your security, governance, and reporting requirements when evaluating your workflows before implementing them with AI.
One simple way to determine if you are ready to begin using AI is to consider the following: If your automated KYC process flagged 12% more customers than it did last week, would you be able to explain the reason why to your stakeholders? If not, this is not an issue with the vendor; it is an issue with your own governance and oversight.
The structural evaluation criteria for banking infrastructure
Traditional software providers must comply with extensive and stringent regulatory requirements. As you develop your framework, it will be important to develop a safe product that does not just make marketing claims.
Security and compliance
You cannot store information within a public cloud service's private bucket; you must deploy this information on a single-tenant or on-premises deployment in order to comply with your data residency regulations.
Integration and legacy core systems
Most modern technology platforms are built on the base assumption that you keep all of your information in an organized data warehouse. In most situations, however, you will have applications and databases that were built using COBOL and/or relational databases as the basis of your core systems.
[Legacy Core Data] ──> [Ingestion Pipeline] ──> [Vendor Sandbox] ──> [Operational Risk]
Explainability and the black box dilemma
If any loan requests submitted through a credit scoring mechanism are denied, this denial will need to provide an easily understood rationale for that decision in order to comply with regulatory requirements. More importantly, if an automated claims processing system denied a claim based on complex background assertions, it is necessary to clarify how such a denial would affect a (potential) customer's money.
What thinking-first actually looks like operationally
Take Tuesday morning at 9:00 AM as an example. The risk management division detects an uptick in false positives in your KYC verification process.
Instead of waiting for the vendor to generate a support ticket for you, the operator goes to the governance platform to investigate the model lineage. Fifteen minutes later (9:15 AM ), they conclude that the root cause of the problem can be traced back to a corrupted data feed from a regional credit bureau.
The operational reality: Demo vs. Production
According to findings from an executive survey of McKinsey & Company, while 65% of companies regularly use AI technologies, they have been unable to build the scale of these systems beyond their initial trial phases.

| The wrong approach (The sales demo trap) | The right approach (Systems-first reality) |
| Acquisition on attractive UI and sample data pre-configured by Vendor. | Model validation performed using disorganized production data with historical deficiencies. |
| Assumption of vendor being responsible for all ongoing model maintenance and retraining expenses. | Funded internal engineering man-hours used to monitor pipelines/data drift. |
| Regeneration of base systems (infrastructure) through reliance on general-purpose API endpoints. | Choosing platforms that aggregate through organizations current systems (databases) – using a secure agency local gateway. |
Framework for Architecture
Rather than evaluating the capabilities of the various platforms that will be utilised as part of our software application, operations professionals will define a specific architecture whereby the platforms serve as separate layers within the total regulatory compliant banking process.
1. Layer of Data Ingestion and Transformation
The initial phase of operation will determine how easily the platform interfaces with real time systems without needing to create large scale custom ETL data pipelines.
Databricks
Operator's Perspective: What you will receive immediately from this solution is a fast pipeline. While it requires a great deal of data engineering experience from the user interface, it allows for the rapid processing of huge amounts of unstructured core file data. If you don't require anything beyond a predictive run, you can expect to exceed your budget very quickly with any investment in its cluster structure.
C3 AI
Operator's Point of View: The primary emphasis of the current data processing systems is based on standard templates created within the banking industry as opposed to the historical data processing methodologies (i.e. traditional). The development of applications for detecting fraud and risk will be completed much more quickly as a result of having the complexity of programming logic hidden in the user interface. However, you will continue to be limited by the restrictive design of the solutions provided.
2. The governance control tower layer
Financial platforms operate under a much higher level of regulatory scrutiny than any software company. Your framework will need to address the lineage of models, and how compliance drift is prevented, by financial platforms.
IBM watsonx
Operator's Perspective: This tool is focused only on the compliance room. The dashboard user interface makes the audit process simple, and shows you the reasons behind automated decisions. It takes the headache of tracking from your developers, but it can also be an expensive way to solve problems if you are not experiencing regulatory/data residency issues.
ModelOp
Operator's Perspective: The real power comes from agnostic integration. It is irrelevant where your models reside. An ops team will have one repository of all model risk and compliance. They will have the ability to manage the entire enterprise without any hassle; however, to configure custom alerts, a strict quality assurance discipline is necessary.
3. Automated model lifecycle layer
The last architectural decision to make will be what method will be used to measure data decay in the live system. Models are random; therefore, any time the market changes, so will the models.
DataRobot
The Tool's Viewpoint on Operators: This tool automates the monitoring and display of how much your machine learning model has drifted. The interface clarifies the moment your model starts producing results that are less accurate than you prefer and allows non-technical managers to see the drift. The tool simplifies retraining models, but does not remove human judgement.
Operator's Operational Environment: While you are considering your options for these tools, do not utilize the products with test data. You should have the vendor use your most difficult data sources to evaluate how the product works under real life conditions.
Vendor evaluation checklist (time estimate 10 minutes)
[] The vendor has the ability to deploy applications in a segregated manner in our own environment.
[] The model generates automated explainability logs to support audit trails.
[] The platform has a native integration with our process maintenance schedule.
[] The total cost of ownership will be inclusive of the cost for api, data, and training on the training infrastructure.
[] The contract contains clear clauses regarding the ownership of the data and how models will be transferable should we choose to terminate our partnership.
The trends show that the era of buying large all-inclusive software packages is coming to an end and that companies are switching to modular open source intelligence frameworks with the intention of keeping their data under their control.
Is your operations team developing a flexible enough platform to allow you to migrate to another language model next year or are you locking your company into a single vendor's closed architecture?
FAQs
How do I assess the actual ROI of an enterprise system?
Determine how much time in engineering hours you save in data preparation, and how many less manual reviews you have in your existing process. Do not base ROI calculations on assumed productivity increases from marketing material.
What is the primary cost drivers for these installations?
Data preparation and continuing monitoring of the pipelines. The first year of software licensing will only be 30% of total operation costs over the next three yrs.
How do we monitor the model drift in financial applications?
You have to have automated data logging at the point of intake and watch for any major detours with production data compared to their training set with special monitoring tools.
Will I be able to implement low-cost open-source frameworks within a heavily regulated industry?Absolutely. However, you must have strict access control policies and deployment containers for all open source frameworks.
What else should I consider during the pilot phase of this project?
Never test the platform using a clean dataset. Always require the vendor to use your most corrupt historical dataset to verify that they can handle real-world operational problems.
Your next move
As we discuss this, open your vendor procurement pipeline spreadsheet from the past. In each project folder, add a new column to collect the “Model Portability and Extraction Cost.” This new column should be researched so that you will know exactly what amount of your budget will be used to pay for the extraction of the model before any action is taken on the project's other costs.


