CasePilot

Reading a blockchain trace without over-reading it

Chain analytics products are very good at showing you a picture and very bad at telling you how much of that picture is measured and how much is guessed. Most of the poor crypto files I have reviewed failed at exactly that seam. This note is about reading a trace at the right confidence level, and writing it up so the next person can tell what you actually established.

The first time I sat a new analyst in front of a chain analytics dashboard, she found a link to a sanctioned entity in about four minutes. She was pleased. I asked her how many hops away it was, what the intermediate cluster was, and how the tool had decided that the terminal address belonged to that entity at all. She did not know, because none of that was on the screen she was looking at.

That is not a criticism of her, and it is not really a criticism of the tool either. Blockchain analytics platforms are built to surface things quickly, and they do. The problem is that the output arrives with a visual authority that the underlying evidence often does not support. A red node on a graph looks like a fact. Some of the time it is a fact. A lot of the time it is a probabilistic attribution built on clustering heuristics, historic scraping and a vendor's internal confidence rating that never made it onto your screen.

So the discipline I try to teach is not scepticism for its own sake. It is knowing, at every point in a trace, which of the four things you are looking at: the ledger, the tool's clustering, the tool's attribution, or the tool's opinion.

What the ledger actually tells you

The ledger is the only part that is genuinely measured. On a public chain you can see that address A sent a certain amount to address B at a certain block height, and you can verify it independently. That is arithmetic. Nobody's judgement is involved.

Everything above that is inference. The tool has grouped a set of addresses into a cluster because they behaved as though a single entity controlled them — most commonly because they appeared together as inputs to a single transaction, which on UTXO chains is a reasonable but not infallible assumption. The tool has then labelled that cluster, sometimes because the entity published the address, sometimes because an analyst deposited a small amount at an exchange and watched where it landed, sometimes because a law enforcement release named it, and sometimes because a similar cluster looked like this one two years ago.

Those methods differ enormously in reliability. A cluster labelled from a deposit test performed last month at a live exchange is close to ground truth. A cluster labelled "likely darknet market, unnamed" on behavioural grounds is a hypothesis wearing a jacket. Both render as a coloured box.

Direct exposure and indirect exposure are different animals

This is the distinction that most often gets flattened in a file, and it is the one that most often makes a case look stronger than it is.

Direct exposure means your customer's address received funds from, or sent funds to, an address in the flagged cluster with no intermediary. One transaction. Nothing in between. That is a real, documentable relationship between two ledger entries, and it deserves serious attention.

Indirect exposure means there were hops. Two, five, twelve. Each hop is a point at which the funds may have changed economic ownership entirely, and the further you go the less the connection means. Money that passed through a large exchange's hot wallet three hops back is not meaningfully connected to whatever preceded it, because that hot wallet mixes the deposits of millions of unrelated people by design.

I ask analysts to state hop count explicitly in every write-up, and to state what sat at each hop. "Indirect exposure of about 4% to a high-risk cluster" is not information. "Two hops, via an intermediate cluster the tool labels as an unhosted wallet with no attribution, to a cluster labelled as a mixing service" is information. The first version is a number the tool produced. The second is something a reviewer, a regulator or a court could evaluate.

Why percentages mislead

Exposure percentages are calculated against a chosen denominator, and vendors do not all choose the same one. Some measure against total value received by the address, some against the value of the specific transaction under review, some against a rolling window. A customer who receives one payment of two hundred pounds' equivalent from a flagged source, on an address that has otherwise seen nothing, shows a hundred per cent exposure. That looks alarming until you notice the absolute figure. Always read the percentage and the amount together, and put both in the file.

The things that genuinely break the picture

Three structures reliably degrade a trace, and you should name them when you meet them rather than tracing through them as though nothing happened.

Mixers and privacy protocols sever the link by design. Some tools will render a "demixed" path anyway, based on timing and amount correlation. That output can be useful as a lead. It is not equivalent to an observed transaction and should never be written up as though it were — the confidence attached to a demixing inference is typically far lower than the confidence attached to a labelled exchange deposit, and if the tool exposes that distinction, quote it.

Cross-chain bridges break continuity in a different way. Funds enter a bridge contract on one chain and something of equivalent value emerges on another, and the matching of the two events is again an inference — often a good one, sometimes not. Bridge tracing has improved a great deal in the last few years, but a bridged path is a joined path, not a continuous one, and the join is where the reasoning has to be shown.

Consolidation addresses are the quiet one. An address that has aggregated inputs from thousands of unrelated sources will inherit the risk exposure of every one of them. Trace back through it and your customer acquires exposure to counterparties they have never had any contact with. This is the single most common way I see a benign file turn red for no good reason. When I see a node with an implausible number of inbound counterparties, I treat it as a terminus rather than a waypoint, and I say so in the write-up.

Case note

A payments firm I advised escalated a customer — a small consultancy incorporated about three years earlier — after its screening tool flagged 11% indirect exposure to a cluster labelled as a sanctioned exchange. The alert covered nine inbound transfers between March and August, totalling roughly £84,000 equivalent, the largest being £19,400. The first analyst's note read, in full: "Customer has sanctions exposure per chain tool. Recommend exit."

When we unpicked it, the exposure sat five hops back and ran through a cluster with more than forty thousand distinct inbound counterparties in the review period — a consolidation point, almost certainly an OTC desk or a payment processor. Two of the nine transfers had no flagged exposure at all. The customer, when asked, produced invoices for six of the nine, matching to within about two per cent on amount and within four days on date. The firm kept the relationship, moved it to enhanced monitoring with a six-month review, and documented why the tool's figure did not survive contact with the underlying path. Had the exit gone through on the original note, there would have been nothing in the file to defend it.

A risk score is a triage input, not a conclusion

Every vendor score is a weighted composite of things the vendor has decided matter, using weights you did not set and usually cannot see in full. It is a perfectly reasonable way to sort a queue. It is not a finding, and it cannot carry a decision on its own.

The failure mode I see in file reviews is the score appearing in the rationale as though it were evidence: "score of 78, therefore high risk, therefore escalate." Substitute any other tool and the absurdity is clearer — you would not write "the monitoring system generated the alert, therefore the activity is suspicious." The same logic applies here as in any transaction monitoring alert triage: the system's output is the beginning of the work, not a summary of it. The reasoning you add is the actual product.

This matters more in crypto than elsewhere because the tooling is younger and the vendor market is concentrated. Two firms running different products against the same wallet will sometimes reach materially different exposure figures, because their cluster attributions differ. Both are defensible. Neither is the truth. The FATF guidance on virtual assets is clear that the risk-based approach still governs, and a risk-based approach requires you to reason about the risk rather than adopt someone else's number for it.

Writing it up when the trace is suggestive but the attribution is thin

This is the situation that produces the worst prose in our industry, because people either overstate to justify the escalation or understate to avoid committing. Neither serves the reader.

What works is separating observation from inference in the sentence structure itself. Say what the ledger shows. Then say what the tool asserts and on what basis, including the confidence level if the vendor publishes one. Then say what you concluded and what you could not establish. Something like: the wallet received approximately X in eleven transfers over a given period; the sending cluster is labelled by the vendor as associated with a named service, at medium confidence, on behavioural grounds rather than a verified deposit; no direct exposure to sanctioned addresses was identified; the customer has not provided a commercial rationale for the counterparty.

Notice that nothing in that accuses anybody of anything. It sets out what is observed, what is inferred and by whom, and what remains open. The same construction serves whether you are populating an EDD file or building toward disclosure — and if it does go to a report, the discipline of keeping observation and inference apart is exactly what makes a SAR narrative useful to the receiving authority rather than merely compliant. The Wolfsberg Group materials on the quality of suspicious activity reporting make much the same point about signal over volume.

One more habit worth building. Record the date and version of the trace, and screenshot it. Cluster attributions change. A wallet that was unlabelled in January may be labelled in June, because someone did a deposit test in the interim. If your file says "no adverse exposure identified" with no timestamp, a reviewer two years later cannot tell whether you missed something or whether the world moved. Date-stamping the trace turns a claim into a record.

Where the judgement actually sits

Chain analytics has made a genuine difference to this work. Ten years ago the honest answer to "where did these funds come from" was often that we could not say. Now we frequently can, with real precision, and that is a substantial gain.

But the tool cannot tell you whether a transaction makes commercial sense for this customer, whether the counterparty relationship fits the business they described at onboarding, or whether the pattern of value moving in and out has an economic logic. Those questions are the same ones you would ask of a set of bank statements, and answering them still requires knowing the customer — which for a corporate structure means doing the work of identifying the ultimate beneficial owner rather than accepting the name on the account.

The trace tells you where value went. You are still the one who has to decide what it means.

What matters: The ledger is measured, everything above it is inferred, and your file should make plain to the next reader which is which.

Practise the work, not the theory

CasePilot puts you in the analyst's seat with a morning queue of alerts and authored case files — the same decisions this note describes, with a disposition to defend at the end of each one.

Open CasePilot

Or work cases offline — iOS and Android, free.