A duplicate payment almost never looks like a duplicate.
On March 3, AP pays invoice INV-00412 from Northern Supply Co. Three weeks later,
invoice 412 arrives from Northern Supply Company, Inc., filed under a different vendor ID
that someone in purchasing created because they couldn't find the original in the ERP search. Same amount,
same description. The system sees nothing unusual: as far as it knows, these are two different vendors with
two different invoices. Both get paid.
Nearly every AP software vendor now offers "AI-powered duplicate detection." Before you evaluate that pitch, ask a far less glamorous question: how many times does each of your vendors exist in your vendor master? The answer to that question, more than any model, determines how many duplicates you'll actually catch.
1. How duplicate payments actually happen
Most duplicate payments aren't fraud. They're process and data failures that repeat in highly recognizable patterns:
- The same invoice keyed twice in different formats:
INV-00412,INV412, and412are three different strings to your ERP. - The vendor exists two or three times in the master: name variants, records created by different departments, system migrations that were never cleaned up.
- Resends treated as new invoices: the supplier sends a reminder, a statement, or a copy "in case you didn't get it," and someone books it.
- Two payment channels: an ACH or wire from AP and, in parallel, a corporate card charge or an expense reimbursement for the same service.
- Multi-entity operations: the invoice lands at subsidiary A and also at subsidiary B, and each one pays it.
- Voided and reissued invoices: the supplier corrects an invoice and reissues it under a new number for the same transaction. If nobody links the two, the system sees two valid invoices.
- Urgent payments outside the workflow: the supplier threatens to stop shipping, someone cuts a manual check "and we'll book it later," and the original invoice keeps moving through approval.
Notice something: almost all of these patterns break the same assumption. The assumption that the system knows when two records are the same vendor.
2. How detection works, layer by layer
Well-designed duplicate detection isn't a single rule or a single model. It's a stack of layers, each one more expensive and noisier than the last, each trying to catch what slipped past the one above:
Prevention before payment vs. recovery after
Each of these layers can run at two points: before the payment is released, or in an after-the-fact review. The operational difference is huge. Stopping a duplicate before it's paid costs a few minutes of review. Recovering it afterward means asking the supplier to send money back, waiting on a credit memo, and sometimes straining a commercial relationship. Recovery audits have their place, but they're the safety net, not the primary control.
Which layer catches what
| Pattern | Example | Layer that catches it | Depends on vendor master? |
|---|---|---|---|
| Identical double entry | INV-00412 / INV-00412 | 01 | Yes: only if both land on the same vendor |
| Different format | INV-00412 / 412 | 02 | Yes |
| Keying error | 10482 / 10428 | 03 | Yes |
| Duplicate vendor record | Same amount, two vendor IDs | 03, only if it knows both IDs are the same entity | Entirely |
| Two payment channels | ACH + corporate card | 04 | Yes: the card merchant must be linked to the vendor |
| Voided and reissued invoice | Two invoice numbers, one transaction | 02 / 03, using PO and line-item detail | Partially |
3. Why everything depends on the vendor master
There's a technical reason the right-hand column says "yes" almost every time. Comparing every invoice against every other invoice produces an unmanageable volume of false matches: thousands of vendors bill similar round amounts. So practically every detection engine uses the vendor as the partition key: it compares invoices within the same vendor, or within a group of vendors the system knows are related.
If your vendor exists three times, that comparison never happens. Not because the algorithm is bad, but because you told it these were three different companies.
The typical defects in a mid-market vendor master are always the same:
- Duplicate records with name variants, abbreviations, or different entity suffixes (Inc., LLC, Co.).
- Missing, mistyped, or placeholder tax IDs, often on "one-time vendor" records that quietly became recurring.
- The same bank account on two "different" vendors: a sign of a duplicate or, worse, of fraud.
- Inactive vendors that were never blocked and can still receive invoices.
- Catch-all records like "Miscellaneous Vendors" where nothing can be compared.
- Vendor setup that anyone can perform, with no check for an existing record.
"A duplicate detection system is only as good as its definition of 'the same vendor.' And that definition lives in your vendor master, not in the model."
4. What happens when you point AI at a dirty vendor master
This is where the AI pitch becomes dangerous, not because the models are bad, but because they're very good at sounding reasonable. With flawed master data, three failure modes show up:
False negatives, with a justification attached
The model reviews the two invoices from our example and concludes: "This is not a duplicate: the invoices belong to different vendors (ID 20418 and ID 31877) and carry different invoice numbers." The reasoning is flawless. The premise is false. The model inherited that premise from your data and presented it with a confidence that discourages anyone from questioning it.
False positives at scale
To compensate, someone loosens the thresholds: "flag any identical amount within 30 days." Immediately the monthly rent, professional-services retainers, and flat-rate recurring freight show up. The team gets two hundred alerts, discovers nearly all of them are noise, and learns to clear them in bulk. The real duplicate goes out with the rest.
The same error, at a different speed
An AP clerk processes a few dozen invoices a day and, every so often, gets that gut feeling: "Didn't we pay this same vendor for this last week?" A model processes thousands in minutes and has no memory beyond what the data gives it. The underlying error doesn't go away: it scales, and it arrives with a well-written explanation. That's why the question of who sets thresholds and who reviews exceptions matters so much, as we cover in who is accountable when an AI agent makes the wrong call.
5. The right order: data, then rules, then AI
This is the sequence we recommend, in this order, without skipping steps:
- Profile your vendor master before buying anything. Four simple queries tell you almost everything: vendors sharing a tax ID; vendors sharing a bank account; legal names that match after normalization (strip "Inc.", "LLC", punctuation, and case); and active vendors with no activity in the last 18 months.
- Consolidate and block. Define one golden record per entity, reassign history, and block duplicates rather than deleting them so you keep the audit trail. Give the vendor master a named owner.
- Govern vendor setup. Require and validate the tax ID (in the US, IRS TIN Matching against the W-9 is the standard check), search by tax ID and bank account before creating any record, separate duties (whoever creates vendors doesn't approve payments), and independently verify every bank account change. These are the controls an auditor expects to see, as we explain in what SOX compliance actually requires from an automated process.
- Normalize at intake. Extract the invoice number consistently, and wherever your suppliers issue structured e-invoices, store the unique document ID as a key field. It's one of the strongest anti-duplicate keys available, as long as it's actually captured.
- Deterministic rules first, AI for the residue. Layers 01 and 02 are cheap, explainable, and auditable. AI-driven fuzzy matching should work on what's left, with a person reviewing every alert and a log of why each one was released or held.
- Measure what matters. Duplicate vendor rate in the master, alert precision (what share of alerts turned out to be real duplicates), amounts stopped before payment versus amounts recovered after, and average time to resolution.
6. Where AI does earn its place
None of this means AI is unnecessary. It means it has a specific job, and a valuable one:
- Extracting data from inconsistent documents (PDFs, e-invoices, emails) so the invoice number and document ID always arrive in the same format. We go deeper in what an AI agent can and can't replace in finance data entry.
- Proposing vendor merge candidates by comparing names, addresses, and bank accounts, for a person to confirm. In other words: using AI to clean the master, not to work around it.
- Recognizing documents that aren't invoices: reminders, statements, pro formas, and resends.
- Explaining every alert with evidence (both invoices, the shared bank account, the history) so the reviewer can decide in one minute instead of fifteen.
AI is an amplifier. On a clean vendor master, it amplifies your team's ability to catch what rules can't see. On a dirty one, it amplifies the error and puts a convincing narrative on top of it.
Before your next vendor demo
If you're evaluating a duplicate detection tool, run the four queries from step 1 against your vendor master first. If you find vendors sharing a tax ID or a bank account, you know where your real exposure is, and you know no model is going to close it on its own. Start there: it's a bounded project with a clear owner and a measurable outcome, exactly the kind of first step we argue for in what mid-market finance teams get wrong about starting small.
If you want to compare notes on what you find, reach out.