A data-first approach to sanctions and watchlist screening
- FinScan

- 3 days ago
- 7 min read
Screening can only find what the underlying data lets it see, and AI models are subject to the same limitations.

This was the structure of a real record:
Bank refund, AZ Wellness Organization, care of Dave Macy.
From a screening perspective, this data carries two risks. "Bank refund" is a statement of intent and not the name of a bank (and imagine the flood of false positive matches to the word bank). The two genuine parties in the record, AZ Wellness Organization and Dave Macy, may not be screened properly at all. Data cleansing pulls both party names out of that single string and presents each one to the screening engine separately, which lowers the risk of a true hit going unfound.
In a FinScan survey of more than 550 North American compliance professionals, 59% said data quality is where they spend the most compliance time.
Why data quality matters more as AI enters compliance
AI is starting to enter compliance workflows at scale. Teams and software vendors are using it to score matches, summarize alerts, and clear false positives. Those models read the same customer records that traditional screening software reads.
AI models inherit poor quality and missing data just like traditional systems. If a name is hidden in an address line, it's invisible to both the AI and the rules-based engine because of the field it's in. In this scenario, teams implementing AI are scaling poor models because they haven't fixed the underlying root cause of many of their pre-existing issues.
SR 11-7 requires rigorous data quality assessment and relevance as part of sound model development. All model components, data inputs included, are subject to validation. The Wolfsberg screening guidance notes that a screening application itself may be submitted for consideration as a model. And, an explainable model can only detail the record it was handed.
What data profiling finds in customer data before screening
The first step to data quality is profiling. Profiling tools answer a few basic questions.
Where is the data?
What is its state?
Does the same value appear in more than one form?

Profiling tends to find issues like the below examples.
Name variations.
The same customer is Robert in one system and Bob in another. Both spellings are correct, and a screening engine treats them as two people.
Partial names.
Some records carry a title and a surname with no given name, which leaves the matching logic very little to work with in either direction.
Contradictory demographic values.
A record whose title and gender value disagree points to something that went wrong at capture or during a migration.
Wrong dates of birth.
Date of birth is one of the few fields that separates a customer from a sanctioned individual who shares a name, and when it's wrong the institution loses its main way of telling them apart.
Address problems.
Cities and postal codes are missing, whole addresses sit in a single free-text field instead of being parsed into separate lines, and the country is recorded as United States in one system and USA in another.
Profiling quantifies how often each issue occurs and how deep it runs.
How data cleansing finds parties hidden inside a single record
Cleansing corrects what profiling finds. The priority cases should be records where a party exists in the data but not in a place the screening system can read.
Joint accounts are one example. Two people entered as a single name string, for example Kim and Jim Tynan, look like one entity to the system. Additionally, names embedded in address lines create the same problem. The person is there; they're just in the wrong place.

If the data is fed into the system as is, those individuals and entities are never screened against the lists. Cleansing, at its core, identifies each party inside a string and passes it to the screening engine in a way it can understand.
Why address validation and geocoding affect screening accuracy
Addresses are another common data problem because of the free text nature of them. Unless an organization has strict guidelines that are adhered to, it's up to the individual to decide how to enter the address. Validation and geocoding tools resolve these issues by standardizing street names and postal codes.
In compliance, address is one of the attributes used to tell a customer apart from a similarly named party on a list. It also determines the geographic risk rating. Consistency across systems is the difference between catching a high-risk customer and upsetting a low-risk one.
How matching and linking resolves duplicate customer records
If the name Martha R. Parks appears in one system and Martha Parks care of John Parks appears in another with the state entered in two different formats, Martha Parks is two records.

Matching and linking resolves records like these automatically by identifying both the duplication and the relationship between the two parties. Both feed risk identification, which can't be accurate while the records sit apart.
Why a 360-degree view needs a golden record and data lineage
Deduplication and linking create a golden record. This single, accurate representation of a customer is assembled from every system that had pieces of truth. This is the data for the screening engine.

In compliance, it's not enough to create that record. Regulators want to see the lineage back to the source records. The evidentiary chain of how the data was sourced, reformatted, and handled is as important as the golden record itself.
Why regulators treat data quality as part of the screening control
Regulators treat the data feeding compliance as part of the control itself.
In the EU, The European Banking Authority (EBA) guidelines on internal policies, procedures and controls for restrictive measures require payment and crypto-asset service providers to assess whether the data they hold is sufficiently accurate, up to date and detailed to determine whether a party to a transfer, their beneficial owner, or a person authorized to act on their behalf is subject to restrictive measures. The EBA is direct about why in its rationale for the guidelines: incomplete or mistaken customer and beneficial owner data produces inaccurate outcomes even when the screening system itself is technically sound.
Meanwhile in the United States, Federal Reserve and OCC guidance on model risk management (SR 11-7) includes a rigorous assessment of data quality and relevance in sound model development and requires that model inputs be validated alongside processing and outputs.
Similarly, the association of global banks Wolfsberg Group states in its sanctions screening guidance that the accuracy and completeness of an institution's own data is central to an effective and efficient screening process.
Why data quality for sanctions and watchlist screening is different from data quality generally
Data quality is often handled at the enterprise level in the IT team. They use general tools built to make operational data usable. This produces customer files good enough for accounts receivable and customer support. But, these projects do not solve compliance risk because the attributes sanctions, watchlist, or restricted party screening depend on (e.g. name structure, address parsing, and date of birth) are not the ones required for general business purposes.
Do you have a data problem?
The answers to these six questions usually show whether data is hampering effective compliance efforts.
Has the business been through an acquisition, integration, or system migration in the last two or three years?
Where does your screening data originate upstream: direct entry, broker feeds, or other internal systems?
When was the last time someone looked at data completeness, meaning missing fields and record types, rather than match accuracy?
How much analyst time goes into reviewing false positives each month?
Do you know your current false positive or match review rate, and has anyone benchmarked it?
Has a regulator, auditor, or internal review flagged anything about match volume or data quality recently?
Data quality built for compliance
Innovative Systems brings fifty years of data management experience to the customer data behind AML and sanctions & watchlist screening. Data that's adequate for operational use is often not built for what compliance asks of it. FinScan Reveal assesses the underlying data and exposes the errors holding screening back, and FinScan Enhance corrects them, separating names hidden in joint accounts, address lines, and around noise words so that each party is screened in its own right, and applying remediation decisions across duplicate alerts to take redundant work out of the analyst queue. Organizations reach compliance-centric data quality in less than 20% of the time and at 50% of the total cost of other providers. The work has been recognized with Best Data Solution for Regulatory Compliance at the A-Team Data Management Insight USA Awards and the Datos Insights AML Impact Award for Best Sanction/Watchlist Innovation.
Frequently asked questions
What is a data-first approach to sanctions and watchlist screening?
It means preparing customer data before it reaches the screening engine, through profiling, cleansing, address validation, matching and linking, and the deduplication that produces a golden record, so that the engine is matching against complete and consistently formatted party records.
Why does poor data quality create false positives in sanctions screening?
Content that isn't a party name often sits in the name field. A record beginning "Bank refund" gets matched as though it were a party, producing false positives against a name no customer holds.
Can a sanctioned party be missed because of a data error?
Yes. When two people are recorded as a single name string in a joint account, or a name is embedded in an address line, that party is never presented to the screening engine as a party, so it isn't checked against the lists at all.
What is a golden record, and why do regulators ask about data lineage?
A golden record is a single accurate representation of a customer, assembled from every system that held part of it. Regulators want to see how that record was built, meaning how the data was sourced, reformatted, and handled, which makes the audit trail as important as the record itself.
How is data quality for screening different from general data quality work?
General-purpose tools are built to make operational data usable for running the business. Screening depends on a narrower set of attributes being complete and consistently structured, including party names, parsed addresses, and dates of birth, and on the institution being able to show afterwards how the screened record was assembled.
Where does a data quality program start?
With profiling. A diagnostic assessment quantifies how often each issue occurs and how deep it runs, which is what scopes the remediation work and supports the business case for it.


