The observation
One of the most persistent assumptions in identity programmes is that a better platform will fix identity problems. It is an appealing idea, because a platform is something you can buy, schedule, and put in a business case.
It has not matched what we see. Across enterprise IAM programmes in financial services, insurance, retail, manufacturing, pharmaceuticals and the public sector, the strongest predictor of whether a programme delivers has not been the product on the licence. It has been the quality of the identity data underneath it.
A governance platform does not clean your data. It industrialises whatever data you already have — including the parts that are wrong.
What we have seen
The same defects recur across organisations that have nothing else in common. None of them are exotic:
- Duplicate identities — the same person present two or three times, usually a legacy of a merger, a rehire, or a contractor who later became an employee.
- Missing or non-unique employee identifiers — the attribute the entire correlation model depends on, blank or reused.
- Manager hierarchies that do not reflect reality — vacant positions, dotted lines recorded as solid, or a chain that terminates in someone who has left.
- Inconsistent worker types — contractor, contingent, temp, third party and vendor used interchangeably, with no agreed definition of which is which.
- Missing lifecycle dates — particularly leaving dates, frequently entered after the person has already gone, or not at all.
- Correlation failures — accounts in a target system that cannot be tied back to any identity, so they sit outside governance entirely.
- Multiple HR sources that disagree — common in groups formed by acquisition, where each operating company has its own system and none is authoritative.
Why it happens
None of this is carelessness. HR systems were built to run payroll and manage employment, not to feed an access control model. Attributes that are approximate for HR purposes — a manager who is roughly right, a leaving date entered next week — become load-bearing the moment an IAM platform starts making automated decisions from them.
Two organisational factors make it worse. First, identity data rarely has a named owner: HR owns the system, IT owns the platform, and the data quality between them belongs to nobody. Second, identity programmes routinely treat the HR feed as a prerequisite someone else is handling, and discover its true state during integration testing — the point at which the timeline is least able to absorb it.
The consequences
Bad identity data does not announce itself. It surfaces as a series of apparently unrelated operational problems:
- Provisioning produces the wrong access — birthright rules keyed on department or job code give people what the data says they are, not what they do.
- Leavers are not detected — no leaving date means no trigger, and the account remains active until somebody notices.
- Certification becomes theatre — a broken manager hierarchy routes the review to the wrong person, who has no basis to judge it and approves.
- Role models will not converge — role mining on inconsistent attributes produces either a handful of meaningless roles or hundreds of near-duplicates.
- Reporting cannot be trusted — and once a board or an auditor finds one number that does not reconcile, the rest of the evidence pack is treated with suspicion.
- Manual workarounds accumulate — exceptions, spreadsheets and side processes grow around the gaps, reintroducing exactly the manual effort the programme existed to remove.
Worth checking before your next phase: what proportion of accounts in your largest target system correlate cleanly to an active identity? If nobody can answer that quickly, the data question has not been asked yet — and it will be asked later, by an auditor.
Our approach
We profile the identity data before designing the solution, not during build. In practice that means four things:
- Establish an authoritative source per attribute, not per system. It is entirely normal for HR to own worker type while the directory owns email — what matters is that the decision is explicit and written down.
- Define the correlation strategy deliberately, including what happens when it fails. Orphaned and unmatched accounts need an owner and a route to resolution, not a report nobody reads.
- Remediate at source wherever possible. Correcting data inside the IAM platform creates a second version of the truth, and the two diverge from the day you do it.
- Measure data quality as an ongoing metric, not a one-off cleanse. Identity data degrades continuously, because organisations change continuously.
This is unglamorous work and it rarely survives contact with a compressed timeline — which is precisely why it is worth protecting. It is also the part no platform selection decision can substitute for.
Key takeaways
- A platform industrialises your existing data quality. Where that quality is poor, automation scales the problem rather than solving it.
- Identity data needs a named owner. Where it belongs to nobody, it degrades by default.
- Profile the data before the design, not during the build — the cost of finding out later is measured in timeline, not effort.
- Fix data at source. Corrections made inside the IAM platform become a second truth that quietly drifts from the first.