CasePilot

Tuning a transaction monitoring rule without breaking it

Most tuning exercises I have reviewed were really volume-reduction exercises wearing a lab coat. This note sets out how to move a threshold defensibly — what you test, how big the sample has to be, and what you have to be able to show afterwards.

Somebody senior will eventually walk into your team and say the monitoring alert numbers are too high. They will not say "our rules are imprecise". They will say the queue is unmanageable, or the vendor invoice for review hours has gone up, or that a backlog is showing on the operations dashboard. And then somebody will be asked to "tune the rules".

That is the moment where tuning goes wrong, and it goes wrong before anyone touches a parameter. The problem has been framed as a capacity problem. Tuning is not a capacity tool. It is a precision tool. If you raise a threshold because you have too few analysts, you have not tuned anything — you have quietly narrowed your own field of vision and written a number on a form to justify it.

I have run this exercise in a retail bank with roughly two million accounts, in a payments firm with a monitoring stack held together with optimism, and in a crypto business where half the typologies had no settled parameter conventions at all. The mechanics are the same everywhere. The discipline is what varies.

Start from what the rule is trying to see

Before you look at a single parameter, write down in one sentence what behaviour the rule exists to surface. Not the rule name. The behaviour.

"Cash deposits sequenced below a reporting figure across multiple branches in a short window" is a behaviour. "CASH-07" is not. If nobody on the team can articulate the behaviour, the rule cannot be tuned, because you have no way of judging whether an alert was a good alert. You will fall back on the only measure available — how many alerts closed with no further action — and that measure will push you towards a quieter queue rather than a sharper one.

This links back further than people expect. The rule set should trace to the typologies you identified as material in your business-wide risk assessment. Where a rule cannot be mapped to a typology you have said you are exposed to, that is worth knowing before you spend six weeks optimising it. The FATF framing of the risk-based approach is often quoted as licence to do less; it is more usefully read as an instruction to be able to explain why you monitor for what you monitor for.

Above-the-line and below-the-line testing

The two halves of the exercise sound symmetrical. They are not.

Above-the-line testing looks at the alerts a rule already produces and asks what they were worth. You pull the population that fired over a defined period, band it by the parameter you are considering changing, and look at outcome quality within each band. If cash-structuring alerts triggered between the current threshold and some higher figure produced almost nothing of investigative value across a long enough window, that band is a candidate for change.

Below-the-line testing is the harder and more important half. Here you sample the activity that did not alert — the population sitting just under the current parameters — and review it as though it had. You are asking an uncomfortable question: what are we currently not seeing? Below-the-line work is the only part of tuning that can find a reason to make a rule wider, and in my experience it is the part that gets truncated when the deadline bites.

A tuning paper containing only above-the-line analysis is not a tuning paper. It is a business case for a smaller queue. I have seen those go into a supervisory conversation and I have seen how that conversation ends.

What a sample has to show before a number moves

Analysts often ask me how big the sample needs to be. The honest answer is that it depends on the alert volume and the rarity of the behaviour, but there are a few things I insist on regardless.

The window must be long enough to contain seasonality. A three-week sample taken in August in a business with heavy summer trading will tell you a story about August. I generally want something in the region of six to twelve months of production data, and I want to know what was happening in the business during that period — a product launch, a migration, an onboarding push, a change in a screening feed can all distort the picture.

The sample must be reviewed to the same standard as live work. If below-the-line items are skimmed by whoever is free on a Friday afternoon, the outputs are decoration. I have had good results embedding tuning samples into normal case allocation, at normal quality-assurance rates, so the reviewer does not know they are looking at a test item. Documented at the point of review, using the same case notes discipline you would apply in transaction monitoring alert triage.

And the outcome measure has to be richer than filed-or-not-filed. Suspicion reporting rates vary with jurisdiction, product, and the reporting culture of the institution, so a rule that produces no reports is not automatically a bad rule. I ask reviewers to record whether the alert produced new understanding of the customer relationship — a corrected expected-activity profile, a discovered counterparty pattern, an escalation to enhanced due diligence, a referral to fraud. Rules earn their place in several ways.

Case note

A mid-sized payments firm I supported had a single velocity rule covering inbound faster payments: more than fifteen credits in seven days, aggregate above £12,000. It generated about 4,100 alerts a quarter and closed roughly ninety-four per cent with no action. The proposal on the table was to lift the count to twenty-five and the value to £30,000, which modelling suggested would remove around 2,600 alerts a quarter.

We ran below-the-line first, over the twelve months from March 2021 to February 2022, sampling 400 accounts sitting between eight and fourteen credits. Twenty-nine of them showed a pattern the rule had never been built to catch: credits of £180 to £450 from unrelated individuals, followed within about forty-eight hours by outbound transfers to two overseas beneficiaries. Aggregate throughput across those twenty-nine accounts was just over £1.9m in the period.

The threshold went the other way for one segment and up for two others. Total alert volume fell by about a third, which was never the point, and four reports went to the financial intelligence unit within the following quarter from activity the previous configuration would not have surfaced.

The segment problem

The single most common structural fault I encounter is one threshold applied across customer types that behave nothing alike.

A £15,000 aggregate parameter is loose for a pensioner with a current account and absurdly tight for a mid-market construction firm making stage payments. Run both through the same rule and you get two failures at once: the retail population under-alerts, and the commercial population floods the queue with items an analyst can dismiss in ninety seconds. Average performance looks tolerable. Nothing about it is.

Segmentation does not have to be elaborate to help. Splitting by product, by declared turnover band, by customer type, or by expected-activity profile captured at onboarding will usually do more for precision than any amount of parameter arithmetic on an undifferentiated population. What it costs you is governance overhead — more parameters to own, more documentation, more justification when a segment definition changes.

A note on new products

New products have no production history, so there is nothing to test above the line. Set parameters conservatively, review at short intervals — I like thirty, ninety and one hundred and eighty days — and say in writing that this is what you are doing and why. Nobody objects to a provisional threshold that has been declared provisional.

The documentation a reviewer will actually want

Assume the person asking about your change in two years' time has never met you and has forty minutes. Your tuning file needs to answer their questions without you in the room.

The JMLSG guidance is helpful on the general principle that monitoring should be proportionate and reviewed, and the Wolfsberg Group statements on effectiveness are worth reading alongside it — they push firms towards asking whether monitoring produces useful output rather than whether it produces output at all. Neither will hand you a number. Nobody will hand you a number, and any vendor who offers one should be asked what they know about your customer base.

Zero alerts is a finding, not an achievement

Every so often I open a rule inventory and find something that has produced nothing in eighteen months. Somebody usually presents this as good news.

It is not. A silent rule means one of a handful of things: the parameters have drifted so far from the population that they cannot be met, an upstream data field has stopped populating, the rule is coded incorrectly, or the typology has disappeared from your book. Only the last of those is benign, and it is the least likely. I have twice found a rule producing nothing because a mapping change had emptied the field it evaluated — in one case for the better part of a year.

So build the check in. Any rule producing no output over a defined period should raise a control event, get investigated, and be written up either as a defect or as an explained silence. The FCA and other supervisors have shown consistent interest in whether firms know their monitoring is working, not merely that it is switched on.

And keep the analysts close to the tuning. The people who close alerts all day know which rules waste their time and which ones occasionally show them something real. That knowledge does not appear in any model output. In the teams I have run, the most valuable tuning input came from asking three experienced reviewers which rule they would fix first — and they agreed more often than the data scientists did. If you want a sense of what those reviewers are seeing, the note on what structuring looks like on a statement covers the pattern that most often sits just below a badly set parameter.

Takeaway: Tune to see the behaviour more clearly, and let the volume land wherever the evidence puts it — because a smaller queue is a side effect, never the objective.

Practise the work, not the theory

CasePilot puts you in the analyst's seat with a morning queue of alerts and authored case files — the same decisions this note describes, with a disposition to defend at the end of each one.

Open CasePilot

Or work cases offline — iOS and Android, free.