CasePilot

Tuning a transaction monitoring rule without breaking it

Most tuning exercises I have reviewed were really volume-reduction exercises wearing a lab coat. This note sets out how threshold testing is meant to work, what a sample has to demonstrate before a number moves, and why the documentation matters more than the number itself.

Tuning is a precision exercise, not a volume exercise

I have sat in a lot of rooms where someone opens with "we need to get alert volumes down by forty per cent" and then asks the analytics team to tune. That framing has already broken the exercise. Once the target is a number of alerts rather than a quality of detection, every subsequent decision bends towards the target, and the testing becomes a way of justifying a conclusion someone reached in a steering meeting.

The honest framing is different. A monitoring rule is a hypothesis: transactions with these characteristics are more likely than the general population to warrant a human look. Tuning asks whether the hypothesis is calibrated — whether the boundary you drew sits in a sensible place given what your customers actually do. Sometimes the calibrated answer produces fewer alerts. Sometimes it produces more. If your method can only ever produce fewer, it is not a method.

There is a related confusion worth naming early. A high false positive rate is not, by itself, evidence that a threshold is wrong. Monitoring is designed to over-select; that is the point of a screen. An alert is arithmetic — a transaction crossed a line you drew — and nothing about crossing that line implies anything about the person behind it. What matters is whether the alerts you generate are informative, meaning they cluster around behaviour you would want to see, and whether the ones you are not generating would have told you something.

Above the line and below the line

The two halves of a threshold test answer two different questions, and teams routinely do one and call it tuning.

Below-the-line testing

Below-the-line testing asks: what am I missing? You take the population of transactions or customers that sat underneath the current threshold during a defined period, sample from it, and review those items as if they had alerted. You are looking for cases that would have been escalated. If you find them, and if they cluster in a particular band beneath the line, your threshold is too high.

The discipline here is in the sampling. A random sample across the whole below-the-line population will be dominated by the very small stuff, because that is where the volume is. I prefer stratified sampling — bands beneath the threshold, with proportionally heavier sampling in the band immediately below the line, because that is where the interesting boundary cases live. If your threshold is a thirty-day aggregate of eight thousand, the band from six to eight thousand deserves far more attention than the band from zero to two thousand.

Above-the-line testing

Above-the-line testing asks: is what I am catching worth catching? You look at alerts generated over a period, banded by how far above the threshold they sat, and you measure how many progressed to escalation or a disclosure. If the band immediately above your line produces effectively nothing across a meaningful volume, that is a signal — though not a conclusive one, and I will come back to why.

Run only above-the-line testing and you will always be told to raise the threshold, because the marginal band is always the thinnest. Run only below-the-line and you will always be told to lower it, because you will always find something eventually if you look hard enough. The pair is the method. Either alone is advocacy.

What a sample has to show before you move the number

Three things, in my experience, before I will sign off on a change.

Sufficient sample size, defined before you look. Decide the sample size and the confidence level in advance and write them down. I have watched teams sample fifty items, find nothing, and declare the band clean. Fifty items tells you very little about a population of two hundred thousand. Whether you use a formal statistical sample or a judgemental one, the size and the reasoning need to exist on paper before the results do.

A period that covers the behaviour you care about. Six to twelve months is usually the floor, because seasonality is real. A rule tuned on a sample drawn from July to September will not be calibrated for the pre-Christmas trading period in a retail-heavy portfolio, and a rule tuned across a single quarter of a fintech's growth curve will be calibrated for a customer base that no longer exists.

An outcome measure that is not "did the analyst close it". Disposition data reflects the quality of your investigations as much as the quality of your rule. If your triage is weak, your below-the-line sample will look clean because nobody recognised what they were reading. I have re-reviewed below-the-line samples with a senior investigator and found escalations that the original reviewer missed entirely. If your team is new, this is worth pairing with proper grounding in transaction monitoring alert triage before you trust the sample results.

Case note

A payments institution I worked with ran a cash-equivalent inbound aggregation rule at £5,000 over a rolling thirty days, applied uniformly across roughly 41,000 active accounts. The rule generated about 1,900 alerts a quarter and the escalation rate sat at just under two per cent. The board pack said the rule was inefficient. The proposal on the table was to move the threshold to £12,000, which modelling suggested would cut volume by around sixty-two per cent.

We ran the paired test over the period 1 April to 30 September, stratified into three bands. In the £5,000 to £8,000 band we sampled 240 items and found eleven that a senior reviewer escalated, three of which became disclosures. In the £8,000 to £12,000 band, 180 items produced two escalations. Above the line, the £12,000 to £18,000 band had produced 214 alerts in the period with a single escalation between them. The pattern was not a threshold problem at all — the escalations below the line were concentrated in accounts that had been open under ninety days, and the barren band above the line was almost entirely long-established merchant accounts with stable turnover. We left the headline threshold at £5,000 for accounts under six months old, moved it to £11,000 for the established merchant segment, and volume fell by about thirty-one per cent while escalations rose by nine in the following quarter.

The segment problem

That case turns on the thing that most single-threshold rules get wrong. One number applied across a portfolio containing a sole trader florist, a scaffolding contractor, a currency exchange and a two-year-old marketplace platform is not a calibrated control. It is an average, and the average fits nobody.

Segmentation is where the real precision gains live, and it is also where the work is. You need segments that reflect genuine behavioural difference rather than the customer-type codes your CRM happened to inherit. I usually start with what the business-wide risk assessment already says about the portfolio, because if the segments in your monitoring do not map to the segments in your risk assessment, one of the two documents is wrong. If that mapping is loose, fix the assessment first — the approach I use for a business-wide risk assessment that survives review starts from exactly this kind of behavioural grouping.

Watch for segments that are too small. A segment with two hundred customers in it will not give you a below-the-line sample worth anything, and you will end up tuning on anecdote. And watch for the segment that quietly becomes an exemption: I have seen "high-volume trusted merchant" segments where the threshold drifted upwards over three review cycles until the rule no longer fired for that population at all. That is a control being decommissioned by increments, without anyone deciding to decommission it.

What the documentation needs to contain

Assume that in three years someone who was not in the room will read your tuning pack and have to decide whether the change was reasoned. The number itself is the least interesting part of the file.

What I expect to see: the rule's stated typology and why the parameter relates to it; the population and period sampled, with the exclusions and the reason for each; sample sizes and methodology fixed in advance; the review outcomes, including the items that went the other way; the proposed change with the expected volume and precision effect; who approved it and on what date; and a scheduled point at which the change gets looked at again. The FATF risk-based approach material at FATF is the conceptual anchor for most of this, and in the UK the JMLSG guidance is where I would point a firm asking what "proportionate" means in practice. Other jurisdictions frame the expectation differently — a US-regulated institution will be reading model risk management expectations alongside its BSA obligations, and the emphasis lands in a different place.

One practical point on effective dates. Record when the change went live in production, not just when it was approved. I have seen packs where the approval was dated in March and the deployment happened in September, and nobody could explain what the rule had been doing in between.

A rule that fires nothing is a finding

Silence is the outcome people misread most often. A rule that produced zero alerts across a year is usually presented as a well-tuned rule. It is more often a broken one.

The causes I have found, in rough order of frequency: a data feed that stopped populating a field the rule depends on; a threshold raised past the point where any real customer could reach it; a scenario written for a product the firm no longer sells; and a logic condition with an AND where it should have an OR. All four look identical on a dashboard. All four mean a typology in your risk assessment has no detection behind it.

So treat zero as an alarm. Every rule with no alerts in the review period gets a written explanation, and "correctly calibrated" is not an explanation unless you can show a below-the-line sample supporting it. This matters most for the typologies where the whole point is deliberate avoidance of a boundary — the patterns described in what structuring looks like on a statement exist precisely because someone is working out where your line sits. Raising a threshold does not remove that behaviour; it relocates it.

The last thing I would say to anyone running their first tuning cycle: resist the temptation to change several parameters at once. Move one thing, observe it for a quarter, and write down what happened. Tuning is slow work, and the firms that do it well are the ones that treat it as an ongoing calibration cycle rather than an annual project with a volume target attached to it. The FCA and its counterparts elsewhere have been consistent on this point for years: they care far less about where your threshold sits than about whether you can show why.

Takeaway: Tune to find out where the line should be, not to find out how few alerts you can live with — and write down the reasoning while you still remember it.

Practise the work, not the theory

CasePilot puts you in the analyst's seat with a morning queue of alerts and authored case files — the same decisions this note describes, with a disposition to defend at the end of each one.

Open CasePilot

Or work cases offline — iOS and Android, free.