AI Data Classification Rules Before Generative AI Rollout: A Modesto Operating Playbook — Datapath managed IT, cybersecurity, and compliance
Back to Blog
HEALTHCARE Insights Published August 23, 2026 Updated August 23, 2026 9 min read

AI Data Classification Rules Before Generative AI Rollout: A Modesto Operating Playbook

Before generative AI touches a Datapath customer’s files, define what each data class permits, who may use it, for which workflow, and where prompts.

Dan J Sturdivant, Vice President at Datapath

By

Dan J Sturdivant

Vice President

CaliforniaCentral Valleyco-managed IT

Quick summary

  • Before generative AI touches a Datapath customer’s files, define what each data class permits, who may use it, for which workflow, and where prompts, retrievals, outputs, and logs go. In practice, a small rulebook—tested against real operations—beats a blanket “AI allowed” policy.
  • What should an AI data classification rule actually decide?
  • What should a first AI rollout look like?

Before generative AI touches a Datapath customer’s files, define what each data class permits, who may use it, for which workflow, and where prompts, retrievals, outputs, and logs go. In practice, a small rulebook—tested against real operations—beats a blanket “AI allowed” policy.

At 7:42 a.m. in a Modesto school district, the special education coordinator is preparing a packet for a 9 a.m. eligibility meeting. She wants an AI assistant to compare last year’s accommodations with the current draft and flag missing sections. The source folder contains teacher observations, attendance history, parent emails, and an individualized education program.

The question is not simply whether the district “allows AI.” The decision is whether this user, in this application, may send this particular set of records to this particular service for this particular purpose—and whether the resulting summary can be saved back into the student-record system.

If the classification rule is wrong, the assistant may expose personally identifiable information to a provider that was never approved for that use. If the rule is too broad, staff may copy sensitive records into consumer tools. If it is too restrictive, the district may create an unofficial workaround: screenshots, personal accounts, or unsanctioned browser extensions.

That is the operating problem we help organizations solve at Datapath. We do not start with an AI brand or a vague acceptable-use memo. We start with the data, the workflow, the people involved, and the control that must stop or permit the action.

What should an AI data classification rule actually decide?

A label such as “confidential” is not a control by itself. A useful rule connects a classification to an action. For every data category, answer five questions:

  • Who owns the decision? The superintendent, compliance officer, controller, clinical leader, police records manager, or business unit owner should be identifiable.
  • Who may use the data? Define the role, not just the department. “Finance” is too broad; “accounts-payable manager” is more useful.
  • Which AI action is permitted? Summarization, search, drafting, retrieval-augmented generation, model fine-tuning, or no AI use are different actions.
  • Where may the data travel? Specify approved tenant, region, connector, storage location, and retention behavior where those choices matter.
  • What evidence is recorded? Capture the user, source, purpose, tool, decision, and output disposition when the workflow requires accountability.

NIST’s Generative AI Profile is a cross-sector companion to the AI Risk Management Framework that helps organizations identify generative-AI risks and choose risk-management actions; it is not a magic “AI compliance certificate.” 1 We use that risk-oriented approach, then translate it into rules a user cannot easily misunderstand.

A practical four-level classification model

The following is an operating model, not a claim that every organization’s legal categories should look identical. The labels should be mapped to the organization’s existing records-retention, privacy, security, and access policies.

Data classTypical examplesPermitted generative-AI actionRequired control and owner
PublicPublished board agendas, public job descriptions, approved marketing copyDrafting, summarization, translation, brainstormingApproved tools; ordinary account security; business owner reviews output
InternalInternal procedures, project plans, network diagrams without secrets, staff training draftsUse in an approved enterprise tenant; retrieval only for authorized groupsIdentity-based access, tenant settings, basic DLP, owner approval
ConfidentialContracts, pricing, employee relations files, customer financial records, nonpublic operational dataOnly in an approved tenant with a documented purpose; no open-web copy and no model training without approvalSensitivity labels, DLP, retention rules, vendor review, logging; department owner and security owner
RestrictedStudent records, ePHI, criminal justice information, credentials, payment data, acquisition plansDefault deny; use only in a specifically approved workflow with minimization, masking, and human reviewStrong access control, encryption, approved integration, audit trail, incident path; compliance or executive owner

The important row is Restricted. It should not mean “never use technology.” It should mean that the workflow has to earn permission through a documented purpose, data minimization, an approved provider, and a tested control path.

Rule 1: Classify at the source, not at the prompt box

If classification begins when a worker pastes text into an AI chat window, the organization has already lost useful context. A prompt rarely contains the document owner, retention requirement, original access group, or reason the file exists.

Classify data where it is created or stored:

  1. Apply a sensitivity label or metadata field to the document, record, mailbox, ticket, or database row.
  2. Preserve that classification when the data moves through an approved connector.
  3. Let a DLP policy inspect the destination and action, not just keywords in the prompt.
  4. Block, warn, or require approval when a user attempts to move a higher-risk class into a lower-trust tool.
  5. Review exceptions on a schedule and remove them when the business purpose ends.

For example, a Modesto district could classify a public board agenda as Public, a draft bell-schedule change as Internal, and an IEP attachment as Restricted. The same AI assistant might be allowed to summarize the first, conditionally allowed to summarize the second, and blocked from receiving the third unless the district has approved a specific records workflow.

This is why keyword-only detection is not enough. “Student” or “patient” may appear in a public policy document, while a spreadsheet with no obvious keyword may contain highly sensitive information. Classification should combine labels, repository, identity, file type, destination, and workflow.

Rule 2: Separate input, retrieval, output, and logs

Organizations often evaluate only the prompt. That misses three additional data paths:

  • Retrieval: The assistant searches a document library, file share, EHR export, ticket system, or records-management platform.
  • Output: The generated answer may reproduce sensitive facts, infer information, or become a new business record.
  • Logs: Prompts, responses, usage telemetry, and troubleshooting traces may be retained by the application or provider.

CISA guidance for secure AI deployment specifically addresses protecting sensitive AI information such as outputs and logs, using access controls and encryption, and monitoring model inputs and outputs. 2 Classification rules therefore need four decisions, not one: “May this source be submitted?” “May the assistant retrieve it?” “May this output be delivered to this person?” and “How is the interaction logged and retained?”

Consider a healthcare clinic in Modesto. A nurse may be permitted to summarize a de-identified patient-education draft, while a scheduling supervisor may retrieve appointment instructions but not clinical notes. Even if the model never receives a full medical record, an output that combines diagnosis, appointment time, and patient name can still be sensitive.

For covered healthcare organizations, the HIPAA Security Rule requires administrative, physical, and technical safeguards for electronic protected health information. HHS also identifies risk analysis as the first step and points to access management, audit controls, authentication, and transmission security. 3 The practical implication is straightforward: document what the AI workflow can see, who can invoke it, what it returns, and how the clinic reviews access.

Rule 3: Make purpose and role part of the permission

“Authorized employee” is not specific enough. Access should be tied to a business purpose and a role at the moment of use.

A controller may use an approved assistant to draft a variance explanation from internal financial data. That does not automatically authorize the same person to upload a full customer file to compare spending patterns. A police records specialist may search an approved case-management assistant for a records request, but a general-purpose chatbot should not receive narrative incident reports simply because the specialist can open them in the source system.

For a public-safety organization in the Central Valley, classification should follow the evidence-retention and dispatch workflow. A call-taker’s transcript, CAD event, officer narrative, and disclosure-ready report may all describe the same incident while having different access groups and retention requirements.

The FBI’s CJIS Security Policy addresses least privilege, audit information protection, logging, and cloud-service audit considerations. 4 That makes a blanket “AI can read the records system” permission especially difficult to defend. A better rule might be: “The approved assistant may summarize redacted, closed-case text for records staff; it may not retrieve active-case attachments; every request is logged; outputs remain in the approved records workspace.”

For K-12 districts, FERPA generally requires signed, dated written consent before personally identifiable information from an education record is disclosed, subject to specified exceptions. The regulations also address school officials, legitimate educational interests, reasonable access limits, authentication, and disclosure records. 5 That is why a district’s AI policy should distinguish between a school-approved workflow and an employee using an unapproved public tool.

Rule 4: Minimize before you authorize

The safest restricted-data prompt is often a smaller prompt. Before approving a workflow, ask whether the AI needs:

  • Names, addresses, account numbers, or student IDs;
  • Exact dates when month or year would answer the question;
  • Full documents when selected fields would work;
  • Free-text notes when a structured summary would be sufficient;
  • Attachments, images, or audio that contain unrelated information;
  • Historical records that are not necessary for the current task.

Use redaction, tokenization, field-level filtering, or de-identification where appropriate. Replace “Jane Smith, patient ID 847219” with a workflow token when the model only needs to compare clinical instructions. Then keep the re-identification key outside the AI workspace, under separate access control.

Minimization also improves the answer. A model asked to summarize a 90-page case file may pull in facts that are irrelevant to the task. A model given five approved fields and a defined output template has less opportunity to disclose unrelated information or create an unsupported conclusion.

Rule 5: Treat the provider and tenant as part of the classification decision

An approved model is not automatically an approved destination. Vendor review should answer questions such as:

  • Are prompts and outputs used to train or improve a shared service?
  • What is the default retention period, and can the customer control it?
  • Are administrators, subcontractors, or support personnel able to access content?
  • Can the organization export usage and security logs?
  • Can the provider support deletion, legal hold, incident notification, and access reviews?
  • Does the contract match the organization’s privacy, security, and records requirements?

For a financial institution subject to the FTC Safeguards Rule, the FTC says an effective program begins with knowing what information exists and where it is stored, followed by risk assessment; it also highlights encryption, activity logging, monitoring, and service-provider oversight. 6 Those are useful procurement questions even before a bank or credit union approves an AI connector.

Datapath’s vendor risk management services can help turn those questions into a repeatable review rather than a one-time questionnaire. For organizations with an internal IT department, our co-managed IT model can add the security engineering and monitoring capacity without taking ownership away from the internal team.

What should a first AI rollout look like?

Do not begin with every employee and every repository. Begin with a bounded decision.

For a first release, we recommend choosing five labels, three workflows, and 25 pilot users. The number is not a legal threshold; it is a manageable operating boundary. For example:

  1. Public document drafting: communications staff use approved content to produce a first draft.
  2. Internal knowledge search: authenticated employees search approved procedures and receive links back to source documents.
  3. Restricted-data exception: one narrowly defined workflow uses masked data, an approved tenant, human review, and complete logging.

Run the pilot for 30 days, then review blocked prompts, approved exceptions, false positives, output errors, and access logs. If users repeatedly request the same exception, determine whether the rule is wrong, the workflow is poorly designed, or the business should not be using AI for that purpose.

The success measure is not the number of prompts. It is whether the organization can answer, after an incident or audit: who used the tool, what data was involved, why the action was permitted, what the model returned, where the output went, and who reviewed it.

Where Datapath fits

Data classification before generative AI rollout sits between governance and daily IT operations. A policy document alone will not stop a restricted file from being uploaded, and a security tool alone cannot decide whether a workflow is legitimate.

Datapath can help a Modesto or Fresno-area school district, an Modesto clinic, an California public-safety department, a credit union, or a 100-plus-employee business connect those pieces. Our AI governance work can define the decision model; managed cybersecurity can support identity, DLP, monitoring, and incident response; and vCISO services can provide security leadership when the organization needs an accountable owner.

For regulated workflows, our K-12 IT and healthcare IT teams can help translate existing access and records practices into AI-specific rules. Government and public-safety organizations can also use our government IT expertise and CJIS compliance services to keep AI decisions aligned with the systems that dispatchers, records staff, and investigators actually use.

The right question is not “How quickly can we turn on generative AI?” It is “Which data may cross which boundary, under whose authority, for what operational purpose, with what evidence?” If your team cannot answer that for the first three workflows, the next step is not a larger license. It is a classification workshop and a controlled pilot with a named Datapath team. Start that conversation.


Footnotes

  1. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile | NIST

  2. Joint Cybersecurity Information Deploying AI Systems Securely

  3. Guidance on Risk Analysis | HHS.gov

  4. Criminal Justice Information Services (CJIS) Security Policy

  5. FERPA | Protecting Student Privacy

  6. FTC Safeguards Rule: What Your Business Needs to Know | Federal Trade Commission

See also

Disclaimer: This blog is intended for marketing purposes only, and nothing presented in here is contractually binding or necessarily the final opinion of the authors.

Need a practical roadmap for regulated-industry IT performance?

Datapath can benchmark your current model and define the next 90 days of high-impact improvements.

Book an IT Consultation