Anonymization is what makes licensing operational data possible. Done properly, AI labs learn how your team works without learning who anyone is.
What gets removed
| Category | Examples |
|---|---|
| Direct identifiers | Names, email addresses, phone numbers, usernames |
| Account data | Account numbers, customer IDs, payment details |
| Contact details | Street addresses, signatures, social handles |
| Company-specific identifiers | Client names, deal names, internal codenames (if you request it) |
| Sensitive content | Health, financial, HR details, typically excluded entirely |
Polyshares states identifying fields, customer names, emails and account numbers are stripped by a specialist PII partner before Polyshares or any lab sees the record.
Common techniques
- Redaction: the identifier is deleted or masked (
[NAME]). - Pseudonymization: identifiers are replaced consistently (“Person A”), so conversations still make sense.
- Generalization: precise values become ranges (an exact deal size becomes “$50k–$100k”).
- Synthetic rewrites: the text is rewritten to keep the reasoning but change the wording. micro1 states some datasets undergo this.
- Exclusion: whole channels, folders or date ranges are left out.
Not sure if your data qualifies? The intake takes a few minutes.
Protections beyond anonymization
Anonymization is one layer. micro1 states it also:
- Defines approved data scope, security standards, and redaction rules before participation
- Restricts access to a limited number of authorized staff
- Keeps data only for agreed periods
- Deletes data on request or at the end of the engagement, per agreed terms
Questions to ask any buyer
- Who performs the anonymization, and before or after the data leaves us?
- Which fields are removed by default? Can we add our own (client names, codenames)?
- Can we review an anonymized sample before transfer?
- Who can access the raw data, and where is it stored?
- What happens to the raw copy after processing?
- What are the retention period and deletion process?
- What happens if identifying data is found after delivery?
What you can do before you start
- Exclude personal channels, DMs, and HR or legal folders.
- List client and partner names you want masked.
- Remove regulated data types entirely.
- Tell your team what is happening and why.
Frequently asked questions
Is anonymized data completely safe?
No method is perfect. Re-identification is possible when enough context remains. That is why scope, access limits and contractual use restrictions matter alongside anonymization.
Can I review the data after anonymization?
Ask for it. micro1 states companies can review representative samples before use. Make sample review a contract term.
What is a synthetic rewrite?
The original text is rewritten so the reasoning and structure stay but the wording and specific details change. micro1 states some datasets may undergo synthetic rewrites.