Last year, I worked on a project that needed an audit log to track user activity on secure, premium content. As part of a fraud detection system, the log also captured information about users' devices and approximate locations.
Most users were in jurisdictions with strict privacy laws, such as GDPR and CCPA.
The challenge was to track user activity while following a privacy by design approach, which meant two things:
-
Collecting only the personal data the system needed, protecting it, and exposing it only to people with a reason to see it.
-
Respecting users' rights over their personal data, particularly the right to be forgotten.
This article covers the first point, and the two techniques that helped most: masking, which hides part of a value, and redaction, which removes it entirely. Where I applied them depended on the kind of log: system logs were masked when they were written, and audit logs when they were read. The second point is covered in a separate article on tokenization and crypto-shredding.
A note on vocabulary: I use "personal data", the GDPR term, throughout. US regulations, and the Go struct tags below, call it PII.
Different levels of masking for different data
Keeping personal data safe would be much simpler if it stayed in one place. As a system grows, the data spreads across services, vendors, and storage systems, each with its own security level and access requirements.
Limiting who can see personal data follows from two principles: data minimization and least privilege. Masking and redaction support both by restricting how much of a value each consumer sees.
Not all personal data is equally sensitive, so it helps to sort it into categories and mask each one differently. Two examples:
- IP address: hiding the lower bits is enough when the reader only needs a rough origin. Full redaction is the safer choice when the reader has no use for the address at all.
- Email address: hiding the local part (before the "@") is enough in most contexts, since the domain alone rarely identifies a person.
In Go, struct tags can mark and categorize sensitive fields, so that each one gets its own masking rule:
type User struct {
ID string
Email string `pii:"data,kind=email"`
Fullname string `pii:"data"`
}
Assume for now a sensitive package with a Mask function. It takes a pointer to a struct and masks the tagged fields:
user := User{
ID: "10010",
Email: "john.smith@example.com",
Fullname: "John Smith",
}
_ = sensitive.Mask(&user)
// Output:
// User{
// ID: "10010",
// Email: "**********@example.com",
// Fullname: "**********",
// }
System logs vs audit logs
The two kinds of logs serve different purposes:
- System logs (also called event logs) capture operational events from applications and infrastructure: performance metrics, errors, and other issues. They support monitoring, troubleshooting, and performance work.
- Audit logs (also called audit trails) record security and administrative events tied to critical user or system actions. They support regulatory, security, and accountability requirements, and are often retained longer than system logs.
These differences affect where the logs are stored and where masking should happen.
Masking system logs when they are written
System logs exist for troubleshooting, and the administrators and developers who read them rarely need a user's email, name, age, or approximate location to do it. Masking that data before it is logged costs them little.
System logs also travel. They are copied to file systems, monitoring tools, and other stores with different levels of security and compliance. Masking at the source reduces the risk of unauthorized exposure in all of them.
When developers do need the unmasked data, a separate, secure location with controlled access requests is a better place for it than the general logs.
In Go, this means logging a masked copy of the struct that reveals only what the logging context needs. The same assumed package offers MaskedCopy, which returns that copy in a form logging libraries accept:
maskedUser := sensitive.MaskedCopy(user)
logger := slog.New(slog.NewJSONHandler(os.Stdout, nil))
logger.Info("structured log", "user", maskedUser)
// Output in structured JSON format:
// {
// "time": "2024-11-12T10:31:55.621344+01:00",
// "level": "INFO",
// "msg": "structured log",
// "user": {
// "ID": "10010",
// "Email": "**********@example.com",
// "Fullname": "**********"
// }
// }
log.Println("standard log", maskedUser)
// Output in standard log format:
// 2024/11/12 10:31:55 standard log {ID:10010 Email:**********@example.com Fullname:**********}
Under the hood, MaskedCopy returns a generic wrapper type that holds the masked copy:
type Masked[T any] struct {
// value holds the masked copy
value T
}
Masked[T] implements the interfaces that slog and fmt look for, which is what lets it plug into logging libraries:
// slog
type LogValuer interface {
LogValue() Value
}
// fmt
type Stringer interface {
String() string
}
A caveat: I find this generic approach elegant, but it may not be the fastest. If performance matters more, handling logging and masking explicitly on each struct type is a more direct alternative.
Masking audit logs when they are read
Unlike system logs, audit logs often need to contain personal data. Some developers may disagree with that, so here is my reasoning.
An audit log is evidence that a critical action took place, usually for regulatory or security purposes. To play that role, it has to capture a reliable snapshot of the event. If it depended on other sources to complete the picture, and those sources were less immutable, its value as a verifiable record would suffer.
For tracking user activity or detecting fraud, it is acceptable for audit logs to contain personal data, as long as they are stored securely with strict access controls.
How to store that data securely in immutable storage deserves its own discussion and is out of scope here.
The question for this article is who sees what. Access should depend on the reader's role and responsibilities:
- A customer success agent can see a user's
DocumentViewedevents and the fields about document interactions, but not the approximate location (latitude, longitude) or the IP address. - A fraud detection system needs the IP address to detect suspicious activity.
This calls for a role-based access control (RBAC) policy, with one complication specific to audit logs.
In this project, the audit log was a polymorphic stream of events, each event type with its own structure. Filtering attributes by role at the data layer is difficult with that shape, so the filtering had to happen in the application layer.
That is where masking and redaction fit. The application loads the audit log, then redacts or partially masks fields according to the reader's role and context before displaying them, so each role sees only what it needs.
Here is the DocumentViewed event in Go:
type DocumentViewed struct {
Document string // document ID
Viewer string // viewer ID
Device Device `pii:"dive"`
Location Location `pii:"dive"`
TimeSpent int64
At int64
}
type Location struct {
Lat string `pii:"data" json:",omitempty"`
Lng string `pii:"data" json:",omitempty"`
City string `pii:"data" json:",omitempty"`
Country string
}
type Device struct {
IPAddr string `pii:"data,kind=ipv4_addr" json:",omitempty"`
Platform string
}
sensitive.Mask now needs to accept an optional closure. The closure looks at the execution context, such as the authorized user's role, and decides how to mask each sensitive field:
// authorized user role
authUser := ctx.Value(AuthUserKey).(AuthUser)
redactFunc := func(fr sensitive.FieldReplace, val string) (string, error) {
// `FraudDetectionAgent` role has access to all sensitive data.
if authUser.Role == "FraudDetectionAgent" {
return val, nil
}
// `CustomerSuccessAgent` role doesn't have access to any sensitive data.
if authUser.Role == "CustomerSuccessAgent" {
return "", nil
}
// Otherwise, apply the mask for the field's data kind.
maskFn, ok := mask.Of(fr.Kind)
if !ok {
return "", nil
}
return maskFn(val)
}
// var event DocumentViewed
err := sensitive.Mask(&event, func(rc *sensitive.RedactConfig) {
rc.RedactFunc = redactFunc
})
if err != nil {
log.Fatal(err)
}
The resulting event depends on the role. For FraudDetectionAgent:
{
"Document": "23080",
"Viewer": "v540103",
"Device": {
"IPAddr": "151.117.33.152",
"Platform": "Computer"
},
"Location": {
"Lat": "51.5074",
"Lng": "-0.1278",
"City": "London",
"Country": "UK"
},
"TimeSpent": 100,
"At": 1731437023
}
For CustomerSuccessAgent:
{
"Document": "23080",
"Viewer": "v540103",
"Device": {
"Platform": "Computer"
},
"Location": {
"Country": "UK"
},
"TimeSpent": 100,
"At": 1731437023
}
A Go library for masking sensitive struct fields
After a few iterations, I refined the masking logic from these examples and open-sourced it as struct-sensitive. The sensitive package assumed above is that library. It comes with a set of predefined masks and can be extended with masks for custom data types.
To sum up the approach:
- Sort personal data by sensitivity and mark it at the struct-field level.
- Mask system logs when they are written, because they travel and rarely need the data.
- Keep audit logs complete, and mask them when they are read, according to the reader's role.
This worked for the project I described. A different log structure, or a data layer that can filter attributes by role, could shift the balance. If you have handled this differently or see a gap in the reasoning, issues and contributions are welcome on the repository.
I hope the library saves you some work on the way to GDPR or CCPA compliance in your Go applications.