For AI labs
RL environments to train frontier models & evaluate compliance agents.
Realistic financial crime work: onboarding customers, monitoring transactions, screening sanctions and preventing fraud.
Four environments
Built the way a financial crime team works its queue. An alert or onboarding file comes in; the agent gathers the evidence, applies policy and risk appetite, and decides whether to clear, escalate, request information or report. The rationale is judged as closely as the decision, the way a QA reviewer or examiner would.
-
KYC & Customer Due Diligence
Identity verification, document authenticity, beneficial ownership, PEP and adverse-media screening, and customer risk rating, from CDD through EDD.
-
Transaction Monitoring
Alert triage and investigation across fiat and crypto: typology detection, source of funds and wealth, and escalation to SAR/STR.
-
Sanctions Screening
Name and wallet screening, true-match vs false-positive adjudication, ownership and control, and exposure across complex payment flows.
-
Fraud Prevention
Account takeover, APP scams, mule networks and first-party fraud across cards and payments.
How it works
The agent works a case the way an analyst would: gathering facts, weighing them against policy, and recording what it decided and why.
- Realistic synthetic customers, accounts and activity.
- The tools, records and policies an analyst works with.
- Graded on the decision and the evidence behind it.
A correct call with no evidence behind it earns no credit, and neither does a guess.
Benchmarks
FinCrime Bench tests agents against held-out cases across all four environments, scoring the decision together with the evidence that supports it. Those cases are sealed: none appear in training, so a strong score means the agent worked the case rather than memorised it. KYC Bench, focused on onboarding and due diligence decisions, is in development.
Research
Short notes on how the cases get built and what the benchmarks show.
-
From a public SAR narrative to a synthetic case
How published typology patterns become fully fictional cases: provenance kept separate from invention, evidence arriving in stages, and the review every case passes before it enters a benchmark.
Working on agents for risk and compliance? We’re running private pilots with AI labs.
Bring your own model; we supply the cases, the policies and the grading.