What is risk scoring: risk scoring explained
Risk scoring is the practice of turning observable facts about a customer, a transaction or a claim into a single graded rating that decides how much scrutiny the case receives. It ranks work for human attention, and the rating it produces is an instruction about effort, not a finding about the person.
You will hear it called risk rating, risk grading, or customer risk assessment depending on the firm and the regulation it sits under. The mechanism is the same in all of them.
What does risk scoring mean in practice?
Every compliance and claims function has more files than reviewer hours. Scoring is how the hours get allocated. A score answers one operational question: of the four hundred cases in this queue, which forty deserve a second person to look?
The idea is written into the international standard. FATF Recommendation 1 requires countries and firms to identify and assess their money laundering and terrorist financing risks and to apply measures proportionate to those risks, which is what a risk-based approach means in practice. Proportionate is the operative word. A score exists so that low-risk cases move and high-risk cases stop.
How is a risk score calculated?
Most production models are simpler than their reputation suggests. Three components do the work:
- Inputs. Discrete facts drawn from the file: customer type, jurisdiction, product, channel, transaction value and pattern, and for a claim, the timing, the evidence provided and the loss history.
- Weights. A value assigned to each input by the organisation, reflecting its own risk appetite and its own history of confirmed findings.
- Bands. The output translated into a handful of tiers, commonly low, medium and high, each tied to a defined action.
The band matters more than the number. A score of 68 means nothing to a handler. "High risk, route to a second reviewer before any payment" means something, because it names what happens next. Models that produce a number without a defined action are decoration.
Risk scoring example: two files, one band
Two motor claims score high in the same week. The first involves a policy incepted eleven days earlier, damage photographed at night and a repairer the insurer has not used before. The second involves a customer with three claims in eighteen months and an unusually precise account of the policy wording.
Both go to the same enhanced review queue, and both are paid inside a fortnight. The first customer had bought a car and a policy in the same week. The third claim on the second file was a windscreen. The scoring worked correctly on both, because its job was to buy a closer look, not to predict an outcome.
What a risk score never decides
This is the discipline that separates a function that works from one that generates complaints. A score sets the level of scrutiny. Evidence, weighed by an accountable person, sets the outcome. Those are two different decisions and a well-run model never lets the first quietly become the second.
The reason is that almost every input has an innocent explanation. Cash-intensive businesses are usually just cash-intensive businesses. A customer in a higher-risk jurisdiction is usually just a customer who lives there. A claim shortly after inception is exactly what you expect from someone who bought cover because their circumstances had changed. The same logic applies to a single red flag indicator and to an individual fraud indicator: it is a reason to look, never a verdict.
When automated scoring runs into legal limits
Where a score triggers an adverse outcome without a person in the loop, data protection law becomes a design constraint rather than a compliance formality. Under the GDPR, an individual has the right not to be subject to a decision based solely on automated processing, including profiling, where that decision produces legal effects or similarly significantly affects them. Declining a claim, exiting a customer or refusing onboarding on the strength of a score alone sits in that territory, and the restrictions on automated decision making apply.
What that means for the design of a scoring system:
- Human involvement that is real. A reviewer with the authority and the information to reach a different conclusion, not someone approving a queue.
- Explainable logic. The categories of data used and the meaning of the processing, capable of being described to the person affected and to a supervisor.
- A route to contest. The ability to put a point of view and challenge the result.
- A written record. Why this file scored where it did, which is also the only thing that makes the model auditable later.
National rules and sector codes add requirements on top, and the position differs by market. Treat this as a legal design question at the point the model is specified, not something to retrofit after go-live.
What makes a scoring model worth keeping
- Calibration against outcomes. Measure how often each input actually precedes a confirmed finding, and retire the ones that never do.
- Stability you can explain. If a score moves, someone should be able to say which input moved it.
- Signals about provenance, not personality. Inputs drawn from how a document or an image was produced and received hold up better than inputs drawn from how a customer sounded on the phone. The control points logged around a digital submission are an example of the first kind.
- A feedback loop. What the Special Investigation Unit or the financial crime team learns should change what the model asks for, not only how one case closes.
A score that nobody acts on differently is not a control. It is a column in a report and a data retention liability.