Methodology

What each bench measures, how the cases are built, and how the work is graded.

Environments
KYC, Transaction Monitoring, Sanctions Screening, Fraud Prevention
Benches
KYC Bench, TM Bench, Sanctions Bench, Fraud Bench
Case data
Synthetic; no real customer data
Sources
Publicly available information, used for patterns only
Sets
Sealed and held out
Graded on
The decision and the evidence behind it
Results
Coming soon

01The environments

Each environment is one team’s queue at a fictional institution, with its own written policies and risk appetite. The agent receives an alert, a screening hit or an onboarding file, with a brief that is the same for every case. The records sit behind tools: the customer and KYC profile, accounts, transactions, counterparties, screening results and prior alerts. Nothing is pasted into the prompt. The agent decides what to pull, as an analyst would, and applies the institution’s policy rather than its own sense of what looks suspicious.

It ends the way an analyst does: a disposition, and a rationale that cites the records it relied on. Every model gets the same brief, the same tools and the same limits.

02The four benches

One bench for each environment, each with its own caseload.

KYC Bench

Queue
Onboarding for individuals and businesses, CDD through EDD
Agent receives
Application data, identity and business documents, registry extracts, screening results and the institution’s CDD and EDD policy
Decides
Approve, apply EDD, request information or decline, with a customer risk rating
Graded on
Requirements applied for the customer type, product and jurisdiction; gaps in verification found; beneficial owners identified; rating and rationale
Results
Coming soon

TM Bench

Queue
Transaction monitoring alerts, fiat and crypto
Agent receives
The alert, customer profile and expected activity, accounts, transactions, counterparties, wallet exposure and prior alerts
Decides
Close, escalate, request information, or file a SAR/STR
Graded on
Red flags found and ruled out, typology, the reporting decision, and a narrative that cites the records
Results
Coming soon

Sanctions Bench

Queue
Name, entity and wallet hits, adjudicated before the payment moves
Agent receives
The hit and the list entry, customer or counterparty identifiers, ownership records and the payment details
Decides
True match or false positive; release, block, reject or escalate
Graded on
Identifiers compared, ownership and control checked, route and goods reviewed, and the reason recorded
Results
Coming soon

Fraud Bench

Queue
Flagged logins, payments and new accounts
Agent receives
Device, login and behaviour signals, payment history and payee details
Decides
Hold, release, contact the customer or exit the relationship
Graded on
The pattern identified, loss weighed against turning away a genuine customer, and the rationale
Results
Coming soon

03How cases are built

Real case files can’t be used: SAR confidentiality and customer privacy keep them inside each institution. So every case is synthetic.

Cases start from publicly available information on how financial crime is carried out and detected. We use it for patterns, not text: the roles involved, the documents that should exist, the points where an alert would plausibly fire. No sentence of any source is carried over.

Everything else is invented: the customers, accounts, transactions, screening hits and the policies the agent works to. A case is built to be internally consistent, with enough ordinary activity around the suspicious core that the agent has to discriminate. A payroll account still pays salaries; a trading business still has seasonal peaks.

Each case keeps a private record of the pattern it draws on, held apart from the case itself. A reviewer can check that the behaviour in the case matches the behaviour described, and the source material stays out of what the agent sees.

04Every outcome, not just the guilty

Most alerts are false positives, and clearing them well is most of the job. Each bench includes cases where the right call is to act, cases where it is to clear, and cases where the file is not yet enough to decide and the right call is to request information and say what is missing. Clean cases are built with the same care as suspicious ones, and carry features that look alarming until the profile explains them.

An agent that always escalates can’t score well, and neither can one that always clears. Both will be reported as baselines next to every model.

05Review

Before any case enters a bench, it will be reviewed. The reviewer will check that the intended decision is reachable from the records provided, that it doesn’t depend on a trick of wording, and that the difficulty sits where it was designed to sit. Cases that fail go back for revision or stay out of the sealed set.

06Sealed sets

Published test sets end up in training data, and scores on them stop meaning anything. Our bench cases are sealed and held out: never used for training, never used to tune our own tools or grading, and never published.

The brief, tools and limits are fixed for a bench version. Any change starts a new version, and scores are compared only within one.

07Grading

A decision that can’t be explained won’t survive QA or an exam, so the write-up is graded as closely as the decision, the way a QA reviewer or examiner would grade it.

  • The decision. A wrong disposition earns nothing, however good the write-up.
  • The evidence. Every finding has to trace to a record the agent could see. Citing a record that doesn’t exist scores the case zero.
  • The analysis. Each case has a private checklist: the facts that matter, the red flags present and those that should be ruled out, the typology, and what is missing. Credit comes from what the write-up shows, not from its length.
  • Conduct. Stating suspicion as established fact, or wording that would tip off the customer, is recorded as a critical error.

Where a criterion needs judgement, the grader will be a model from a different family than the model under test, checked against human grading before any score is published.

08Reporting

Results will be reported as the average across cases and runs, never the best run, with error bars. Alongside each score: decision accuracy against the baselines, critical errors, cost and the kinds of mistakes made. Scores will be published with the evidence each model cited.

Results are coming soon.

09Limits

The cases and policies are synthetic. They can’t carry every quirk of a real book of business, and a score measures an agent in our environment, not in a live deployment. Automated grading is not expert sign-off. Expert review of the cases and comparisons across models are separate steps, and we will say which have been done when results are published.

Building agents for risk and compliance? We’re running private pilots with AI labs.