Keep one customer recognisable across systems without exposing identity
An LLM agent investigating one customer across several internal systems needs to tell that it is looking at the same person each time, or the investigation falls apart. Redacting every identifying field is the obvious privacy fix, but it destroys exactly that: once three systems' names for one person are all replaced with the same generic marker, the model cannot tell whether it is looking at one person or three.
What Data Prism does
Data Prism gives one subject one synthetic pseudonym, derived deterministically from the privacy scope, the subject, the namespace, the algorithm version and an HMAC key. The same customer reads as the same pseudonym in every source's data for a given case, so an agent can follow one identity across systems without ever seeing a raw name, email or account number.
Collapsing several sources' values onto one pseudonym would normally hide
that those sources actually disagreed about what to call the person — a
"Patrick Murphy" in one system and a "Pat Murphy" in another both render as
the same pseudonym, and nothing in that pseudonym says the underlying values
differed. Data Prism's pseudonym
collapse is exactly this effect, and it
is handled by comparing the raw values before pseudonymisation and attaching
the result as a ConsistencyFinding rather than by making the difference
disappear. Consistency
findings report which sources agreed with
which — never what any of them actually held — across six kinds, from an
outright disagreement to a same-value formatting difference to one source
appearing to hold an abbreviated form of another's value. A worked
example
shows three sources' three different spellings of one name collapsing to one
pseudonym, with the PERSON_NAME finding as the only place that
disagreement is still visible.
Pseudonyms are also scoped to the investigation, not just to the subject. Scope isolation means the same customer renders as two different, unrelated-looking pseudonyms under two different case ids — nothing about one investigation's pseudonym lets a caller in a different case tell it refers to the same underlying subject.
What it does not do
Consistency across systems here means one pseudonym per subject, not
automatic entity resolution: correlating a subject across sources requires
an identifier the sources already share, resolved by a reviewed
IdentityResolver (see docs/extending.md) — Data Prism
does not infer that two different identifiers across two systems belong to
the same person. It is not anonymisation: the pseudonym is reproducible
from the same scope and key, which is what GDPR Art.
4(5)
means by data kept re-attributable to its subject given additional
information — so Recital
26
treats it as personal data still.
Where to go next
docs/tools.mdhas the full worked examples for both shipped tools,get_entity_contextandcompare_entity_sources, and the complete list of consistency-finding kinds.docs/quickstart.mdproves a pseudonymisedget_entity_contextcall end to end against fixture data with one command; it does not exercisecompare_entity_sources, so seedocs/tools.mdabove for the pseudonym collapse itself.docs/extending.mdcovers writing theIdentityResolverthat decides which sources' identifiers refer to the same subject.docs/configuration.mdis the contract for the scope and key configuration pseudonym consistency depends on.