Privacy by design in Go: tokenization and crypto-shredding

Go
Privacy
GDPR

In a previous article, I described an audit log that tracked user activity for a fraud detection system, and how masking limited who could see the personal data in it.

That covered one half of privacy by design. The other half is respecting users' rights over their data, and in particular the right to be forgotten (Article 17 of the GDPR).

The difficulty is that an audit log is immutable. That is what makes it usable as evidence, and it is also why the records about one person cannot simply be deleted from it.

The approach I ended up with combines three patterns:

  1. Tokenization: identify each person by a random token instead of their email.
  2. Field-level encryption: encrypt each person's data with a key of their own.
  3. Crypto-shredding: forget a person by discarding their key.

This article explains each one with Go code, then looks at what the approach does not solve.

Identifying people by a token, not by their email

Personal data is easier to contain when few things point to it. If an email address serves as the user identifier, it ends up in every table, log line, and event that mentions the user.

Tokenization replaces that identifier with a random value, the token, and keeps the mapping between the two in one protected place. Everything else refers to the token. The PCI tokenization guidelines, written for card numbers, describe the same pattern and call that protected place the data vault.

Tokenization is one way to achieve what the GDPR calls pseudonymization: data that can no longer be linked to a person without additional information, provided that information is kept separately and protected (Article 4(5)). The EDPB's draft guidelines on pseudonymisation describe what makes it effective.

Assume for now a protector service that offers the operations used in this article. Here it turns an email into a token:

email := "john.smith@example.com"

tokens, err := protector.Tokenize(ctx, privacy.TokenDataSlice(email))
if err != nil {
	log.Fatal(err)
}

viewer := tokens.Get(email).Token

// viewer: "064521d6-39f6-4c37-9903-f2fbb4533f1d"

The same email gives the same token on every call, so the token works as a stable identifier. I prefer to do this at the entry point of the system, before the email has spread anywhere.

"Kept separately" is the part that matters in practice. I prefer to keep the mapping between emails and tokens in its own store, with its own access rules, so that the services reading the audit log can use tokens without being able to turn them back into emails. If the mapping sits next to the data it protects, anyone who can read one can read both.

A token is still personal data as long as the mapping exists, because the person can be identified again through it. Tokenization reduces exposure. It does not make the data anonymous.

Tokenization also has a cost. Because the same email always gives the same token, you can still count and join on it, but a search by email or a report grouped by email domain now needs a lookup in the mapping store first. ENISA's report on pseudonymisation frames the choice of technique as a balance between protection and utility, with no single solution for every case. If tokenizing an identifier makes the queries or reports your team depends on much harder, it is worth negotiating. And if the system already identifies users by a random internal ID, that ID can play the token's role without a separate step.

Encrypting personal fields with one key per person

The second pattern encrypts personal fields before a record reaches the store, with a different key for each person (the data subject, in GDPR terms).

Struct tags mark which field identifies the subject and which fields hold their personal data. This is the DocumentViewed event from the previous article, with the viewer's token as the subject ID:

type DocumentViewed struct {
	Document string
	Viewer   string   `pii:"subjectID"`
	Device   Device   `pii:"dive"`
	Location Location `pii:"dive"`
}

type Device struct {
	IPAddr   string `pii:"data,replace=forgotten"`
	Platform string
}

type Location struct {
	City    string `pii:"data"`
	Country string
}

Encrypt replaces each tagged field with its ciphertext and leaves the others alone:

event := DocumentViewed{
	Document: "23080",
	Viewer:   viewer,
	Device:   Device{IPAddr: "151.117.33.152", Platform: "Computer"},
	Location: Location{City: "London", Country: "UK"},
}

err := protector.Encrypt(ctx, &event)

// event.Device.IPAddr:    "ENC..MDY0NTIx...MzZjFk.ATPGOEE/JTjW...RFUFp"
// event.Device.Platform:  "Computer"
// event.Location.City:    "ENC..MDY0NTIx...MzZjFk.3/mYSAulU06Q...cEQ=="
// event.Location.Country: "UK"

The platform and the country stay readable, so the store can still filter and aggregate on them. The key is created the first time a subject appears and reused afterwards. The encryption is AES-256-GCM, with a fresh random nonce for each value.

An encrypted value has three parts: a marker with the format version, the subject ID, and the ciphertext. The subject ID is base64-encoded but not encrypted, because it is needed to find the right key when decrypting.

This is where the first pattern matters. If the subject ID were the email, every encrypted field would carry it in the clear. With a token, the subject ID reveals nothing about the person on its own.

Forgetting someone by discarding their key

With one key per person, erasure no longer means finding and rewriting every record. Crypto-shredding discards the key, and every value encrypted with it becomes unreadable wherever the ciphertext was copied: the log, its replicas, and its backups.

err := protector.Forget(ctx, viewer)

// Later, when the event is read back from the log:
err = protector.Decrypt(ctx, &event)

// event.Device.IPAddr:    "forgotten"
// event.Device.Platform:  "Computer"
// event.Location.City:    ""
// event.Location.Country: "UK"

After Forget, decryption does not fail. Each field takes the replacement value from its tag, or stays empty if the tag has none, so the rest of the record remains usable.

An erasure request can be a mistake, or come from someone who took over an account. For that reason, Forget only disables the key. A scheduled cleanup deletes it for good once it has stayed disabled longer than a grace period (seven days by default). Until then, the decision can be reversed:

err := protector.Recover(ctx, viewer)

While a person is forgotten, new data under their subject ID is refused: Encrypt returns an error instead of creating a fresh key, so their data cannot be written again by accident.

What this approach does not solve

Whether discarding a key counts as erasure is debated, because encrypted personal data is still personal data.

Mathias Verraes, who describes crypto-shredding for event stores, argues that the law does not treat deleting the key as deleting the data. He points to forgettable payloads as the alternative: personal data lives in a separate store that can be deleted, and the immutable record keeps only a reference. That avoids the legal question, at the cost of a second store that every reader has to query.

Regulators have been more open where data cannot be removed. Writing about blockchains, the CNIL says that making data inaccessible brings an organization closer to the GDPR's requirements, without saying it amounts to erasure. The ICO accepts that data in backups can be put "beyond use" until it is overwritten. On the security side, NIST SP 800-88 recognizes cryptographic erase as a way to sanitize storage media, which is the same idea applied to whole devices.

I am an engineer, not a lawyer. Treat this as a map of the discussion, and check it with whoever advises you on data protection.

There are technical limits too:

  • Decrypted copies. Shredding only reaches ciphertext. Anything decrypted and then exported, cached, or logged is out of its reach, which is one more reason to mask data when it is read.
  • The token mapping. Forget discards the key, not the link between the token and the email. That mapping has to be deleted as well, with DeleteToken.
  • Backups of the key store. Discarding a key only works if no copy of it survives. If the key store is backed up, those backups must expire, or deleted keys live on in them.
  • Keys in memory. Keys are cached to avoid a lookup on every call. Another running instance can still decrypt a forgotten person's data until it clears its cache (entries live twenty seconds by default).
  • Time. The approach is as strong as the encryption and the key management behind it, and what is strong today is not guaranteed to stay so.

A Go library for these patterns

The protector assumed above is privacy-engine, an open-source library I built on top of struct-sensitive, the masking library from the previous article. Its key and token storage is pluggable, and it also handles cases left out here, such as streaming encryption for large files.

To sum up the approach:

  • Identify people by a token, created at the entry point and mapped back only in a separate store, so that their email does not spread.
  • Encrypt personal fields with one key per person, and keep the rest of the record readable.
  • Forget a person by discarding their key, with a grace period to undo mistakes, and delete their token mapping too.

This fits stores that cannot be rewritten, such as audit logs and event streams. Where records can simply be deleted, deleting them is simpler and leaves no legal question open. If you see a gap in the reasoning or have handled erasure differently, issues and contributions are welcome on the repository.