How should we classify company data before using AI tools?

Short answer

Classify data into a small number of levels, such as Public, General, Confidential and Highly Confidential, then state for each level which AI tools may process it. Apply the levels as persistent labels, for example Microsoft Purview sensitivity labels, so AI tools and data loss prevention can enforce them. Regulated data, such as health information or CUI, needs its own explicit rule.

Why AI makes classification urgent

AI tools find, summarize and move information far faster than people do. Microsoft's Purview documentation puts it plainly: because AI can proactively surface content, generative AI amplifies the risk of oversharing or leaking data.

An AI policy that says "don't put sensitive data into AI" fails without classification, because employees have to guess what counts as sensitive. Classification turns that guess into a rule they can follow and a control IT can enforce.

A scheme people will actually use

Keep the scheme small. Microsoft's own guidance notes that effectiveness drops noticeably when users have more than five main labels. A workable set for AI use:

  • Public: already published. Any approved AI tool.
  • General: routine internal material. Company-managed AI tools only, never personal accounts.
  • Confidential: financials, contracts, customer and employee data, unreleased plans. Only AI tools whose terms exclude training and whose access follows your permissions.
  • Highly Confidential: trade secrets, unfiled inventions, deal terms. Named, approved tools only, or none.
  • Regulated: health information, CUI, export-controlled data and GxP records. Only environments specifically approved for that data type.

Make the labels do the enforcement

Microsoft describes sensitivity labels as customizable, stored in clear text in file and email metadata, and persistent, so the label stays with the content wherever it is saved. Labels can apply encryption and visual markings, be applied automatically or recommended based on detected sensitive information, set as a default, or required before a file is saved or an email is sent.

Labels also shape what Microsoft Copilot does. Microsoft states that Copilot shows the highest-priority label among the items a response draws on, and that when a label applies encryption Copilot returns content only if the user has the EXTRACT usage right. Microsoft cautions against making an encrypting label the default for documents, because it can block legitimate sharing with outside parties.

Extend it beyond Microsoft 365

Many AI tools never read your labels, so classification has to reach them another way. Microsoft says endpoint data loss prevention on onboarded Windows devices can warn or block sensitive information being pasted into third-party generative AI sites, and Microsoft Defender for Cloud Apps can classify and label content in other cloud services. For everything else, your approved-tool list should state the highest data level each tool may handle, and your vendor review should record why.

Common follow-up questions

Do we need to label every file before rolling out AI?

No. Start with the sites and folders that would cause the most harm if exposed, set a default label for new content, and use automatic labeling for recognizable data such as financial or health identifiers. Expand coverage over time instead of delaying rollout for a full re-label.

How many classification levels should we have?

Four or five is typical. Microsoft notes that effectiveness drops noticeably when users see more than five main labels. Fewer, well-explained levels with real examples get applied correctly far more often than detailed schemes nobody remembers.

Where does CUI fit in an AI data classification scheme?

Treat controlled unclassified information as its own regulated level. It belongs only in environments covered by your NIST SP 800-171 controls and system security plan, so keep it out of any AI tool that sits outside that assessed boundary.

Need help with this?

LAN Service Group designs data classification schemes for small and mid-size businesses and implements them with Microsoft Purview sensitivity labels and data loss prevention, so AI tools follow the same rules as people.

Talk to LAN Service Group (888) 281-7243

Sources