Article

How pseudonymisation protects the privacy of your data

Every prompt you send to an LLM is stored and logged by an external party. Pseudonymisation keeps your sensitive data inside your own environment.

Artific

A hand touching a transparent screen with shields and padlocks as security symbols.

That feels harmless for a recipe or a trivia question, but it becomes a serious risk as soon as you work with customer data, personal data or commercially sensitive information. Legal agreements offer no technical guarantee: your data has left the building and you lose control over it.

If you want to use LLMs safely, you have to protect sensitive information before you send it. Anonymisation and pseudonymisation are not a luxury here, but a necessary first step.

Why anonymisation doesn’t work

Right, so you want to protect your data before it leaves the building. The most obvious solution: anonymisation. Just black everything out. Sounds logical. Until you try it.

Take this sentence, for example.

Jan complained that his order #12345 arrived damaged

Anonymised, that becomes:

[REMOVED] complained that [REMOVED] arrived damaged

Good luck with that. The LLM has no idea what this is about. No context. No useful answer.

Research shows that this kind of generic masking causes a huge loss of quality. Your AI becomes effectively unusable.

But there is an alternative: pseudonymisation

How does it work? Take the same example as with anonymisation. Instead of removing the sensitive data, we replace it with a context-related word and number. Pseudonymised, it becomes:

PERSON_1 complained that his order ORDER_1 arrived damaged

The LLM understands that PERSON_1 is a person and that ORDER_1 is an order number. The context stays intact. The mapping, who PERSON_1 really is, we handle separately, outside the LLM.

The result: the context stays intact so the LLMs can work with it, and privacy-sensitive information is never sent to the LLMs.

Loss of quality? There is some, but it is small. Our tests show the loss of quality is only 8%. And it keeps getting smaller.

Why names are the hardest to pseudonymise

Recognising email addresses is easy. Look for the @ sign, check there is a valid domain name after it, and you’re done.

Phone numbers? Regex. IBANs? Regex plus a checksum. Dutch citizen service numbers (BSN)? Same again, no problem.

But names?

Is “Bas Viool” a name? Or what about “Hope”? “Guy”? “Brouwer”? Names follow no pattern. They depend on context. “Apple” can be a company, a fruit or Gwyneth Paltrow’s daughter.

That is why we tackled names first. If you can detect names, the hardest problem, the rest is not much of a challenge.

Data you don’t send can’t leak

At Artific we have chosen an approach of token-based pseudonymisation. Personal data never leaves your environment. Only the tokens go to the LLM; the mapping stays yours.

The EDPB Guidelines of January 2025 validate exactly this approach. Our current solution performs at industry benchmark level. But we are not done yet. We are now researching how to train our own specialised models, so our customers can use this without any loss of quality, fully under their own control.

Security is an ongoing process

We have now implemented pseudonymisation in the Artific AI Platform, and with it taken a big step in protecting your privacy. Does this solution offer a one hundred percent guarantee that your data cannot leak? Although it is a huge improvement on other AI solutions on the market, it is not yet a watertight solution.

At the same time, the world of AI, legislation and data security is changing at breakneck speed. That is why we have a dedicated team that follows these developments closely and keeps researching how we can continue to make LLMs available safely and responsibly. Real security is not an end point, but an ongoing process.

About Artific

The Artific AI platform, developed in the Netherlands, is designed with Security and Privacy by Design as its starting point. Artific is fully platform and language model independent, and we already support more than 35 language model variants. We build virtual employees that are secure and reliable. Tested by ethical hackers.

This article was written by Erik van der Pluijm, research lead of the Artific GenAI Research Lab. We actively research the possibilities and risks of generative AI for business applications. Do you have questions, or would you like to talk through AI implementation in your organisation? Get in touch.

Back to the knowledge centre