About
FinCrime Atlas builds environments and benchmarks for AI agents doing financial crime risk and compliance work.
Compliance is risk management
Every bank, fintech and crypto firm decides how much financial crime risk it is prepared to carry. It writes that risk appetite into policy: which customers it will take on, what due diligence each one needs, what activity it expects, what gets escalated and what gets reported. Two institutions can look at the same customer and reach different answers, and both can be within their obligations.
Within that framework, analysts make the call on each case. They weigh the customer’s profile against the activity, decide whether the file is complete, and document why. Their decisions are sampled by QA, tested by internal audit and open to the examiner.
Where agents come in
AI agents are now being built to work these queues: onboarding files, alerts, screening hits and flagged payments. Before an institution trusts one with a queue, it needs to know two things. Does the agent apply the institution’s own policy and risk appetite? And would its decisions stand up to QA and an examiner?
General benchmarks answer neither. They test whether a model knows what structuring is, not whether it can tell from a customer’s file that one run of cash deposits is structuring and another is a restaurant’s weekend takings.
What makes it hard to measure
- The right answer is often to clear. Most alerts are false positives. An agent that escalates everything looks careful and is unusable; one that clears everything is a regulatory failure. A test built only from bad actors rewards the first.
- The answer depends on the institution. Risk appetite, procedures and the customer’s risk rating decide what gets escalated. An agent has to apply the policy it is given, not its own sense of what looks suspicious.
- The file is rarely complete. Analysts request documents, ask the relationship manager, or wait on a request for information. Knowing when the file supports a decision, and when it doesn’t, is part of the job.
- The write-up is the work product. A decision that can’t be explained won’t survive QA or an exam. The rationale has to rest on the records, present suspicion as suspicion rather than fact, and never tip off the customer.
- Real case files can’t be shared. SAR confidentiality and customer privacy keep the real work inside each institution. There is no public record of how alerts were actually worked, so a benchmark has to build its own.
What we build
Four environments, each one compliance team’s queue at a fictional institution: KYC, Transaction Monitoring, Sanctions Screening and Fraud Prevention. The cases are synthetic, each institution’s policies and risk appetite are written out, and the agent works the file the way an analyst would.
The environments are built for AI labs to train agents on. Four benches, one per environment, measure the work on sealed cases. Results are coming soon.
Building agents for risk and compliance? We’re running private pilots with AI labs.