What counts as sensitive data
Nearly every security requirement is written in terms of sensitive data. Protect it, inventory where it lives, restrict who can reach it, know when it has left. All of that assumes the organization has already answered a question that is rarely asked directly: which data, specifically, and how does anyone know?
Organizations that skip the question tend to land in one of two places. Either sensitivity is decided by instinct in the moment, which cannot be trained, audited, or applied consistently by two people on the same day. Or everything becomes nominally sensitive because nobody decided otherwise, which sounds cautious and generally means nothing receives particular care while staff quietly work around controls that apply to the cafeteria menu.
That second one has to be separated from something it resembles closely and is entirely defensible. An organization can decide, deliberately, to hold everything to the strictest standard it is subject to. That is a real strategy with real advantages, and for some organizations it is the only sane one. It is covered further down. The failure is not treating everything as sensitive. It is arriving there by default, having never established what the strictest standard requires.
The categories are genuinely different
Sensitive data is not one thing with one owner. The common categories differ in where the obligation comes from, who defines the boundary, and what happens when it is wrong.
Regulated personal information, commonly called personally identifiable information (PII), and protected health information (PHI) where health records are involved, carries obligations created by law. The organization did not agree to them and cannot negotiate them, they follow the person rather than any contract, and they frequently reach across state and national borders to wherever the person is. What counts is defined externally and changes without the organization being consulted.
There is no single authority to consult here, which is itself the difficulty. The United States regulates personal information by sector rather than comprehensively, so which law applies depends on what the organization does. Health records fall under HIPAA, whose Privacy and Security Rules sit at 45 CFR Part 164 and whose definition of protected health information is at 45 CFR 160.103. A financial institution holding customer information is under the Gramm-Leach-Bliley Act, which states an affirmative and continuing obligation to protect the security and confidentiality of nonpublic personal information. A federal agency, or a contractor operating a system of records on its behalf, is under the Privacy Act of 1974.
Underneath all of it sits breach notification, which is state law rather than federal, exists in every state, and differs between them on what counts as personal information and how quickly notice must be given. An organization with customers abroad acquires those regimes too. The practical consequence is that this category cannot be settled by reading one document, which is why it more often needs counsel than an engineer.
Controlled Unclassified Information (CUI) carries obligations created by a contract, and its definition belongs to the government rather than to the organization holding it. This is the category most often misunderstood, because the name sounds like a description of importance. It is not. CUI is a defined set of categories with an official registry and marking rules, published by the National Archives, which runs the CUI program and maintains the registry of categories. Information either falls inside a listed category or does not. An organization's belief that something feels sensitive has no bearing on whether it is CUI, and neither does its belief that something is routine. The program itself is established by 32 CFR Part 2002, and defense contractors have a second place to look, since the DoD CUI Program publishes its own category and marking abbreviations.
CUI is also not one obligation. The regulation divides it at 32 CFR 2002.4 into CUI Basic, where the authorizing law sets out no specific handling or dissemination controls and the uniform controls apply, and CUI Specified, where the authorizing law does set out its own handling controls. Specified controls are not merely stricter; they can simply differ. Treating the two as interchangeable is the most common and most expensive mistake in this area, because the consequences are not remotely comparable.
A spill of CUI Basic is a serious matter, handled through contract mechanisms and reporting obligations. Nobody is likely to go to prison over it. Export controlled information, marked with the export control category's abbreviation, is a different proposition. Where the underlying authority is the export control regime, disclosure to a foreign person can be a criminal violation carrying substantial fines and a real prospect of imprisonment—and "foreign person" includes an employee working lawfully in the organization's own office, and a cloud administrator abroad whose employer never told the customer where support is staffed.
The practical consequence is that a single control set applied to everything marked CUI will be simultaneously too heavy for the basic categories and too light for the specified ones.
The marking is not the requirement. The marking names a category, the registry entry for that category names the law or regulation that authorized it, and that authority is what the organization is actually obliged to follow. Looking the category up is the first step and by far the easier one. What follows is reading the authority itself, working out what it requires for this data in this situation, and then complying with it—which for a Specified category can mean obligations that appear nowhere in the general CUI guidance, nowhere in the contract clause that brought the data in, and nowhere in the control set the organization has already built.
That is the work. An organization that has looked up its categories and stopped there knows the names of its obligations and has not yet met any of them.
Company proprietary information has no external authority at all. Nobody outside the organization will define it, and no registry will settle an argument about it. That makes it the category most often left undone, because deciding requires judgment rather than a lookup.
Information belonging to somebody else, held under a nondisclosure agreement (NDA) or equivalent confidentiality terms, carries obligations the organization accepted deliberately and often years ago. Customer data held to deliver a service usually lands here, and the terms are in an agreement rather than in a statute.
The practical consequence is that these cannot be handled by one policy sentence. They have different definitions, different owners, different retention pressures, and different consequences for getting them wrong.
Every contract says something about protecting data
Organizations often assume that data protection obligations arrive only with the obviously regulated work. Reading the whole contract portfolio of a mid-sized company says otherwise. Every single agreement contained a data protection requirement. Not most of them, and not only the government ones.
What varied enormously was what the requirement said. At one end, a single sentence obliging the contractor to protect the customer's information. At the other, a Department of Homeland Security contract that incorporated by reference the department's own reimplementation of the federal control catalog—a policy directive, a handbook, and roughly thirty lettered attachments, one of which tailors the federal controls and another of which carries the department's own baselines and parameter values. The clause in the contract was a few lines long. The obligation it created ran to something on the order of a thousand pages.
Both are binding, and the short one is arguably harder to satisfy. A sentence requiring appropriate protection obliges the organization to decide what appropriate means, apply that consistently, and be able to defend the decision years later to somebody whose view of appropriate may differ.
"Best practices" is not a requirement
A clause requiring best practices, industry standard care, or commercially reasonable security has the same defect as one requiring appropriate protection. It is worth naming separately because the phrase is common enough that it reads as though it means something.
Best for whom? What is right for a four-person non-profit with a donor list is not what is right for a billion-dollar multinational, and neither answer is wrong. Best when? Practice moves, and what was defensible three years ago may not be defensible now. Best against what? An organization facing opportunistic crime and one facing a well-resourced adversary have different correct answers to the same question.
The phrase has no referent, so nothing can be measured against it. Both parties sign believing they know what it means and find out they disagree at the one moment it matters. The organization cannot demonstrate compliance, an assessor cannot test it, and a dispute ends up buying expert testimony to supply a meaning the contract declined to provide.
For anyone in a position to write the clause, the fix is to name a published standard and be specific about which part of it: the CIS Controls at a stated implementation group, NIST SP 800-171 at a stated revision, or whatever fits the work and the size of the organization doing it. That makes the obligation determinable before signature, so both sides can price it, and testable afterwards, so an argument about whether it was met has an answer.
For the organization receiving such a clause rather than writing it, "best practices" is the signal to ask which standard is meant, and to get the answer in writing. Same question, same technical contact, same record.
The conditional contract
Return to that Homeland Security contract, because length was not its real difficulty either. It was written to be usable across the whole department rather than for the work actually being performed, so much of it was conditional: if the system does this, and the data is that, then the following controls apply. Determining which conditions were met was not possible from the contract, because the contract did not say. The answer had to be obtained by having the program manager ask the customer.
A contracts lawyer's description of that drafting is that it is poor practice, and the security consequence is worse than the legal one. An organization cannot categorize data whose obligations it cannot determine, and it cannot determine them by reading.
What it can do is ask, and asking well matters. The question goes to the customer's technical point of contact rather than to the contracting officer. The contracting officer administers the agreement and is often genuinely unable to say which control set applies to which data on this particular effort; the technical contact usually can, or knows who can. Organizations reverse this regularly, get an unsatisfying answer from the person whose name is on the contract, and conclude that nobody knows.
Then write down what was asked, who answered, their role, the date, and what they said, and keep it with the contract. Two reasons, and the second is the important one. The first is that the record is the only place the applicable control set will ever exist in one piece; organizations that skip it spend the next three years re-deriving it, usually during an assessment, usually from somebody who has since changed jobs.
The second is that if a spill happens later, there is an enormous difference between an organization that assumed and one that asked, was told, and can produce the exchange. It costs an email and a file. It is the cheapest insurance available in this entire subject.
The same duty, for a different reason
That was one badly drafted contract. The obligation to ask also arrives in a structural form that has nothing to do with drafting quality, and any defense contractor will recognize it.
The Department of Defense includes DFARS 252.204-7012, "Safeguarding Covered Defense Information and Cyber Incident Reporting", in very nearly every contract as a matter of routine. It requires protections meeting NIST SP 800-171 and reporting of cyber incidents within 72 hours.
Its presence, though, does not establish that any covered defense information is actually involved in the work. Plenty of contracts carry the clause and never touch CUI at all. The clause is in the contract because it is in almost every contract, not because somebody determined that this effort handles protected information.
So the determination falls to the contractor, and it is the contractor's problem in both directions. Markings are not always applied, and when applied are not always applied correctly. The government does not reliably volunteer which categories are in play. Assume CUI is present when it is not, and the organization has built and priced a control set nobody required. Assume it is absent when it is present, and the organization has a spill, a reporting failure, and a contract problem, in that order and on a 72-hour clock.
The answer is the one above, for the same reasons: ask the technical point of contact, get the answer in writing, and keep it beside the contract. A clause that appears in every contract tells an organization nothing about its own work. Only asking does.
The practical lesson is that the question is not which contracts mention security. They all do. The question is what each one actually requires, and incorporation by reference is where that hides: the clause is short, and the obligation is not.
Who decides, and who does not
For regulated information and CUI, the organization does not get a vote on the definition. It only gets to decide where the data is, who touches it, and how it is protected. Arguing about whether something ought to count is time spent on a question that has already been answered elsewhere.
Proprietary information is the opposite: the organization is the only authority that exists, so a decision has to be made rather than discovered. The useful test is to name the harm. What would it cost, and to whom, if a competitor had this, or a customer, or the public? Information that survives that question belongs on the list. Information that does not is ordinary business information, and saying so out loud is what makes the protected list credible.
"Everything is confidential" is the same statement as "nothing is confidential", delivered with more confidence. That is a claim about what the information is, which is a separate question from how much of it to protect. An organization can conclude honestly that very little of its information is proprietary and still choose to handle everything to one standard, for reasons covered below.
A label has to change something
A classification scheme with four levels and one set of handling rules is decoration. Each label has to change what happens: where the data may be stored, who may see it, whether it may leave the organization, what may be done with it after the work is finished, and how it is disposed of.
The test is straightforward. Take any two labels and describe how handling differs between them. If the answer is difficult to produce, there is one label wearing two names, and the scheme should be simplified until every distinction does work.
Short schemes get followed. A scheme with three levels that people can recall without looking is applied far more consistently than a scheme with seven that requires a reference table, and consistency is most of the value.
One level is a legitimate answer
Taken to its conclusion, that reasoning sometimes lands on a single level. Find the strictest requirement the organization is subject to, apply it to everything, and stop classifying. For some organizations this is not laziness but the correct decision, and it deserves stating clearly because the literature tends to treat classification schemes as self-evidently good.
The advantages are substantial. There is one control set to build, document, and be assessed against. Nobody has to classify anything during ordinary work, which means nothing gets misclassified, and misclassification is the failure mode that actually causes spills. Training is one conversation rather than a decision tree. An assessor sampling any system finds the same controls, so evidence from anywhere is evidence about everywhere. For a small organization holding one regulated category, the cost of building and maintaining two handling regimes usually exceeds the cost of over-protecting the rest.
The cost is equally real and it grows with scale. Applying controls designed for the most sensitive category to everything means encryption, access restriction, retention limits, and disposal requirements on the cafeteria menu. That is money, and it is friction, and past a certain size it becomes an obstacle to doing the work. The point where it stops paying is roughly where the protected material becomes a small fraction of the whole, and the overhead on everything else exceeds what maintaining a boundary would cost. That is when a second level earns its keep.
Someone still has to track the floor. Whichever way the organization goes, the strictest applicable standard has to be identified and then watched, because it moves when a new contract is signed or a new line of business opens. Choosing a single level is a decision about handling, not an exemption from the analysis in the rest of this article. An organization that applies one standard everywhere without knowing which standard that is has not simplified anything; it has guessed, and it will find out whether it guessed high enough at the worst possible time.
Where the tooling question comes in, and where it does not
Some of this sounds like a description of data loss prevention (DLP) software, and it is worth being direct about the relationship. DLP tools watch data in motion and at rest—outbound email, uploads, removable media, cloud storage—and act on what they find, either by blocking it, quarantining it, or recording it for review. Some read the labels a classification scheme applies. Others infer sensitivity from the content itself, matching patterns such as account numbers or the markings on a document.
The important point is the order of operations. Such a tool enforces a decision about what is sensitive and where it is permitted to go. It does not make that decision, and it cannot be configured by an organization that has not made it. Buying one first produces either a policy of blocking nothing, or months of false positives that end with the tool being switched to monitoring and quietly ignored.
There is also a category of data these tools handle badly, and source code is the clearest example. Consider an organization holding a large body of source code: some written in-house, some open source under a permissive license, some under a reciprocal license with obligations attached, and some received from a customer or a partner under terms of its own. Those files carry genuinely different restrictions on where they may go and who may see them. Nothing in the text distinguishes them. Sensitivity here is a property of provenance and license, not of content, and content inspection is what these tools do.
The usual answer is that the files should be labeled, and this is where the approach quietly fails. Labeling that volume of code is work nobody has budgeted, it has to be redone as files are added, moved, refactored, and vendored in, and it depends on every developer applying the right label every time. No program manager is willing to spend the team's time on that, and saying so is not a criticism of the program manager. A control that requires sustained voluntary effort from people measured on something else stays accurate for about a quarter, if that long.
Where content inspection cannot categorize the data, the answer is usually to control the container rather than the contents: which repositories exist, where they are hosted, who can clone them, and what may leave. That is an access control problem with a reliable answer, rather than a classification problem with an unreliable one. Provenance and license are worth tracking too, but that is a software composition question answered by different tooling than the kind that watches outbound traffic.
DLP software is also not required. Nothing in this article assumes such a product exists. A small organization whose staff know which categories exist, where each is allowed to live, and who to ask when unsure may be genuinely well controlled with no tooling at all, and an assessor can be shown that through the procedures people actually follow. Tooling earns its place as scale grows, turnover rises, or the consequence of a single mistake becomes large enough that relying on everyone remembering stops being reasonable. That is a judgment about the organization and its people, not a requirement that follows automatically from holding sensitive data.
The categories overlap, and that is normal
A personnel record is both regulated personal information and company proprietary information at the same time. The salary, the performance review, and the disciplinary history are personal data about an identifiable employee, carrying whatever the applicable privacy law requires: access rights, retention limits, and notification duties if it is disclosed. The same file is also information the organization has its own reasons to protect, because a compensation structure is useful to a competitor recruiting from it. Treating the record as purely an HR privacy matter tends to secure it against outsiders while leaving it readable across the company. Treating it as purely proprietary tends to protect it well and ignore the employee's rights over it entirely.
A contract deliverable can be CUI and somebody else's confidential information at once. A report written for a defense customer may contain government information that falls in a CUI category, and alongside it the technical data a subcontractor supplied under a nondisclosure agreement. The CUI obligations arrive through the prime contract and specify how the information is stored, marked, and reported on if spilled. The subcontractor's obligations arrive through a separate agreement that may forbid disclosure to named competitors, require return or destruction at the end of the engagement, and say nothing about marking at all. Neither set is a subset of the other.
When categories overlap, the handling is the union of the obligations rather than an average of them. The strictest storage requirement applies, the shortest retention limit applies, and the broadest notification obligation applies. Organizations sometimes try to resolve overlap by picking the label that fits best, which reliably discards one of the obligations—usually the one that came from the quieter party, because the government sends assessors and the subcontractor does not.
Notification duties travel with the data
Every category above carries some duty to tell somebody when things go wrong, and those duties are the part most often missing from a categorization record. They get left out because they are not about protecting the data. They are about what happens once protecting it has failed, which feels like a different subject and is not: the obligation attaches to the category, arrives with it, and travels with it wherever it goes.
They also multiply in ways that surprise people. Federal agencies do not share one reporting requirement; a defense contract, a health regulator, and a grant-making agency each have their own trigger, their own recipient, and their own clock. Within a single state, different agencies can impose different duties on the same organization for the same event. And a commercial contract can add its own on top of all of it, frequently on a shorter deadline than any statute, because a customer who wants to be told within twenty-four hours can simply write that down and will.
The practical consequence is that "who has to be told, and how quickly" has a different answer for each category, and sometimes several answers for one category held under several agreements. Working that out while an incident is in progress is the worst available time, and it is when most organizations first attempt it.
Breach notification in its own right is a large subject and not this one. The point here is narrower: the duties belong in the categorization record next to the handling rules, because they are a property of the data rather than a property of the incident, and an organization that has categorized its data without recording them has left out the part with a clock attached.
Where it usually goes wrong
The copies are forgotten. The database is identified and protected while the extract someone made for a report, the backup, the test fixture built from production, the email attachment, and the meeting transcript are not. Data does not become less sensitive by being copied somewhere less formal, though it almost always becomes less protected.
Aggregation changes the answer. Fields that carry no obligation individually can identify a person when combined, and a list of ordinary facts about a project can describe a capability the organization would rather not publish. Sensitivity is a property of the collection, not only of the field.
Nobody owns a category. A list of categories with no named owner produces no decisions when an unusual case appears, and unusual cases are exactly when the decision matters.
The list is written once. New obligations arrive with new contracts and new lines of business, and a classification scheme that has not been revisited since a contract was signed is describing the organization that existed then.
That last one is why the major control standards do not treat this as a one-time exercise. They require the categorization to be reviewed on a defined schedule and again when something changes, and they say so in more than one place: in the requirements covering how information is categorized in the first place, in the system inventory requirements that expect the record to stay current, in the media and information handling requirements that determine marking and storage, and in the periodic review of policies and procedures. An organization that categorized its data once and has grown two lines of business since is not partially compliant with those. It is describing a company that no longer exists, and the assessment will be against the one that does.
What this should produce
A record naming each category of sensitive data the organization actually holds, and for each one: where the obligation comes from, who owns decisions about it, and what handling follows. That is enough to train against, enough to audit against, and enough to answer the question that starts every serious security conversation.
The word to resist is "document", because it suggests something written once and filed. This is a live record. Every new contract can change it, and most organizations sign contracts more often than they revise policies. Two fields carry most of the weight and are the ones usually missing.
The contract reference. Each obligation should name the agreement it came from, by number or identifier, because obligations end. When a contract closes, the requirements that arrived with it may lapse, and an organization that cannot trace which requirement came from where will keep applying all of them forever. That is a slow, invisible, permanent cost increase that nobody ever decides to accept.
The effective dates. When the obligation started, and when it ends or comes up for renewal. Without dates there is no way to answer what applied at a given moment, which is exactly the question asked after an incident, during a dispute, and in any assessment covering a period rather than a day.
The notification duties, as described above: who has to be told, on what trigger, within what period. This is the field with a clock attached, and the only one whose absence is discovered under time pressure.
For a small organization a maintained table is entirely adequate. Past a certain number of contracts, or when the same data category arrives under several agreements with different terms, a table stops working and a small database is the honest answer. The trigger is usually not size but overlap: the first time somebody has to determine which of three customers' terms governs a particular file, a spreadsheet has already failed.
It is also the prerequisite for nearly everything else. An inventory cannot record what the organization has not defined, a risk assessment cannot weigh consequences for data nobody has categorized, and an incident response plan cannot tell anyone whether an event is reportable. Those all depend on this record existing, which is why it is worth doing before the work that assumes it.
The control standards agree, in the sense that matters: categorization sits near the front of every one of them, and the requirements that follow are written as though it has already been done. An organization working through a control set from the top will meet the categorization requirement early, and an organization that skipped it will find every later requirement asking a question it cannot answer.
Kenneth Ingham Consulting helps organizations decide what their sensitive data actually is, and what handling each category should carry.