When FERPA Meets AI: Student Data Privacy Compliance for EdTech Startups Training Models on Education Records

FERPA's 50-year-old education record definition meets AI model training. Here's how EdTech startups navigate FERPA, COPPA, SOPIPA, and Texas SB 1792 when training models on student data.

When FERPA Meets AI: Student Data Privacy Compliance for EdTech Startups Training Models on Education Records
Loading AudioNative Player...

AI-powered learning tools — adaptive tutors, automated grading copilots, personalized curriculum generators — are reshaping how K-12 schools operate. But the federal law that governs student data privacy was written in 1974, a half-century before anyone imagined training a neural network on student records. The Family Educational Rights and Privacy Act (FERPA), codified at 20 U.S.C. § 1232g, protects "education records" maintained by educational institutions that receive federal funding. It was designed for filing cabinets and transcripts, not for gradient descent on aggregated student performance data.

For EdTech startups, this matters now. The U.S. Department of Education released its "Designing for Education with Artificial Intelligence" guide in July 2024, signaling that federal regulators are paying attention to how AI intersects with student privacy. The FTC is actively enforcing against edtech companies — both for COPPA violations targeting children under 13 and for inadequate data security under Section 5 of the FTC Act. And states from California to Texas have layered their own student privacy statutes on top of FERPA, creating a compliance patchwork that no startup can afford to ignore.

This guide covers the five things every EdTech founder needs to understand before training an AI model on student data: what triggers FERPA and the "school official" exception, whether model training constitutes a FERPA-permitted use or an unauthorized re-disclosure, how COPPA layers on for under-13 users, the state law landscape including California's SOPIPA and Texas's student data privacy laws, and the practical compliance architecture — data processing agreements, data minimization, deletion protocols, and model auditability — that keeps your startup out of enforcement crosshairs.

What Triggers FERPA: The "Education Record" and the School Official Exception

FERPA applies to any educational agency or institution that receives funds under a program administered by the U.S. Department of Education — which means virtually every public K-12 school district and most postsecondary institutions in the country. Under 34 CFR Part 99, FERPA protects "education records," defined broadly as records that are directly related to a student and maintained by an educational agency or institution. This includes grades, transcripts, disciplinary records, health records maintained by the school, and — critically for EdTech — any data generated by a student's use of a school-provided digital tool that the school maintains or controls.

When your EdTech platform is contracted by a school district to provide services — an adaptive learning platform, a grading tool, a student information system — the data your platform collects from students in that context likely constitutes education records under FERPA, because the school maintains control over the data and uses it for educational purposes. The school, not your startup, is the FERPA-regulated entity. But your startup becomes subject to FERPA's requirements indirectly, through a specific legal mechanism: the "school official" exception.

The School Official Exception: When Your Startup Becomes a FERPA-Bound Vendor

Under 34 CFR § 99.31(a)(1), a school may disclose personally identifiable information from education records without prior consent to "school officials" who have "legitimate educational interests" in the information. The regulations define a school official to include contractors, consultants, volunteers, and other outside parties to whom the school has outsourced institutional services or functions — provided that the outside party meets four specific criteria:

  1. Performs a service for which the school would otherwise use its own employees. If your platform provides grading, assessment, or student information management, you likely qualify.
  2. Is under the direct control of the school with respect to the use and maintenance of education records. This means your data processing agreement with the district must give the school control over how student data is used, stored, and deleted.
  3. Uses education records only for authorized purposes. You cannot repurpose student data for non-educational uses — including advertising, profiling, or training commercial models — without triggering FERPA's consent requirements.
  4. Does not re-disclose the information to unauthorized parties. Re-disclosure is the landmine for AI training, and we address it directly below.

The practical consequence is straightforward but often misunderstood: when your EdTech startup contracts with a school district, you step into the shoes of a FERPA-bound school official. You are not free to use student data however you want. You may use it only for the specific educational purpose for which the school disclosed it to you, and you may not re-disclose it — or use it to train AI models for purposes beyond that authorized scope — without risking a FERPA violation that can cost your district customers their federal funding and terminate your contracts.

Does Training an AI Model on Student Data Violate FERPA?

This is the question every EdTech founder building AI features needs to answer carefully — and the answer is more nuanced than a simple yes or no.

FERPA's school official exception permits an outside vendor to use education records for the "authorized purpose" for which the school disclosed them. If your platform is contracted to provide adaptive math tutoring, the school has authorized you to use student performance data to deliver that tutoring service. The question is whether training an AI model on that data falls within the scope of "delivering the tutoring service" or whether it constitutes a separate, unauthorized use — or worse, an unauthorized re-disclosure.

The Department of Education's 2024 AI guidance does not create new legal obligations, but it signals the agency's expectations. The guide emphasizes that AI products for education should be built with "safety, security, and trust" and that edtech developers should consider FERPA's restrictions when designing AI systems that process student data. While the guide does not explicitly state that model training is a FERPA violation, it reinforces the principle that student data collected for educational purposes must be used only for those purposes — unless the school has provided consent for broader use.

The re-disclosure problem is the sharpest edge of this issue. Under 34 CFR § 99.33, a party that receives education records from a school may not re-disclose personally identifiable information to any other party without meeting FERPA's consent requirements. If your startup trains an AI model on student data and that model is then used to serve other schools, other commercial customers, or is made publicly available, the model's outputs may constitute a re-disclosure of the original education records — particularly if the model can produce outputs that reveal individual student information. This is not a hypothetical concern; researchers have demonstrated that large language models can memorize and reproduce training data, including personally identifiable information.

The safest approach for EdTech startups is to treat AI model training as a use that requires explicit authorization in the data processing agreement with each district. Do not assume that a contract to provide educational services implicitly authorizes you to train models on student data. Specify in your DPA what data will be used for model training, what the model will be used for, how you will prevent re-identification, and how you will ensure the model is not used beyond the authorized educational scope. If you want to use student data to train a general-purpose model that serves multiple customers, you need either de-identified data that FERPA no longer governs, or explicit consent from the school — and potentially from parents — for that broader use.

When COPPA Layers On: Under-13 Users and FTC Enforcement

If your EdTech app is used by students under 13 — and most K-12 platforms are — COPPA creates a second, independent compliance layer on top of FERPA. The Children's Online Privacy Protection Act (COPPA) requires operators of websites or online services directed to children under 13, or operators with actual knowledge that they are collecting personal information from children under 13, to obtain verifiable parental consent before collecting, using, or disclosing personal information from those children.

The FTC's 2025 COPPA Rule amendments, effective June 23, 2025, with full compliance required by April 22, 2026, strengthen these requirements. The amended rule expands the definition of "personal information" to include government-issued identifiers and biometric identifiers, introduces a new "mixed audience website or online service" definition, and tightens requirements around data retention, deletion, and security. Notably, the FTC declined to adopt specific edtech provisions in the amended rule, citing concerns about potential conflicts with anticipated FERPA regulatory updates from the Department of Education — but it stated it will "continue to enforce COPPA in the edtech context consistent with its existing guidance."

The enforcement record is unambiguous. In the FTC's case against HyperBeard, the mobile app developer allegedly permitted third-party advertising companies to collect personal information from children under 13 without parental consent. The FTC entered a $4 million judgment, suspended to $150,000 based on inability to pay, and required deletion of all personal information collected from children. More recently, the FTC's enforcement against Illuminate Education resulted in a consent order after a data breach exposed personal information of over 10.1 million students. The FTC's final order, approved in June 2026, required Illuminate to implement enhanced security controls, obtain biennial third-party security assessments for 10 years, and — critically — comply with a data minimization requirement prohibiting collection of information "not reasonably necessary to provide products or services" under its customer contracts.

For EdTech startups, the enforcement signal is clear: the FTC expects data minimization, not data maximization. If you are collecting more student data than necessary to deliver your educational service, you are building enforcement risk. And if you are training AI models on data collected from children under 13 without verifiable parental consent that explicitly covers that use, you are in violation of COPPA — regardless of whether a school has authorized the data collection under FERPA. COPPA and FERPA operate independently; school authorization under FERPA's school official exception does not substitute for parental consent under COPPA.

The State Layer: SOPIPA, Texas, and the Patchwork

Federal law is the floor, not the ceiling. More than 40 states have enacted student data privacy laws that impose additional restrictions on EdTech vendors — and these laws apply directly to your startup, not just to the school district. Two of the most significant are California's SOPIPA and Texas's student data privacy statutes.

California SOPIPA: The Industry-Targeted Standard

California's Student Online Personal Information Protection Act (SOPIPA), signed into law in September 2014, was the first comprehensive state law to place student data privacy obligations directly on EdTech vendors rather than on schools. SOPIPA prohibits operators of websites, online services, and mobile applications designed or marketed for K-12 school purposes from:

  • Selling student data
  • Using student data to target advertising to students or their families
  • "Amassing a profile" on students for non-educational purposes
  • Disclosing student data except as required by law or as part of service maintenance and development

SOPIPA also requires operators to maintain reasonable security practices and to delete student data when requested by the school or district. The law applies whether or not the vendor has a contract with the school — meaning that a startup offering a free learning app used in California classrooms is subject to SOPIPA even without a formal district agreement. For AI model training, SOPIPA's prohibition on "amassing profiles" for non-educational purposes is particularly relevant: training a model that produces commercial outputs or serves non-educational users could constitute an impermissible profiling use under SOPIPA.

Texas SB 1792 and the Student Data Privacy Landscape

Texas has its own student data privacy framework, primarily through SB 1792, passed during the 85th Legislative Session in 2017. The law requires EdTech vendors contracting with Texas school districts to:

  • Obtain parental consent for collection of student data by third-party vendors
  • Prohibit targeted advertising using student personal information
  • Ban creation of student profiles for non-educational commercial purposes
  • Delete student data when a student leaves the district or upon request
  • Enter into transparent contracts with districts that detail data collection, use, and security

SB 1792 defines "covered information" broadly — including grades, disciplinary records, health information, photos, online activity, biometric data, and geolocation data. The law creates shared responsibility between districts and vendors, but the compliance burden falls heavily on EdTech companies. Texas also enacted the broader Texas Data Privacy and Security Act (TDPSA), effective July 1, 2024, which imposes comprehensive consumer privacy obligations with no revenue threshold — though TDPSA includes exemptions for data governed by FERPA.

The patchwork matters because EdTech startups typically serve multiple states. A model trained on data from a Texas district may be used by a California district — but the training use must satisfy both Texas SB 1792's restrictions and California SOPIPA's prohibitions. The compliance architecture that works for one state may not work for another, and the default is the most restrictive applicable standard.

Practical Compliance: DPAs, Data Minimization, Deletion, and Model Auditability

The legal landscape is complex, but the compliance architecture is not. If you are an EdTech startup building AI features on student data, here is the framework we recommend building from day one.

1. Data Processing Agreements That Explicitly Address AI Training

Your DPA with each school district must explicitly address AI model training — not bury it in general-purpose language. The DPA should specify: what student data may be used for model training, the specific educational purpose the model serves, whether the model is trained per-district or across districts, how you will prevent re-identification of individual students through model outputs, and what happens to the model when the contract terminates. We cover the broader DPA framework in our guide to SaaS data processing agreement requirements, but for EdTech, the AI training clause is the critical addition that standard DPAs lack.

The DPA should also prohibit re-disclosure of student data through model outputs. If your model is used to serve multiple districts, you need to demonstrate that students in District A cannot be identified through the model's responses to District B. This may require technical controls — differential privacy, data isolation, federated learning — or contractual commitments that the model is trained and used only within a single district's data environment.

2. Data Minimization: Collect Only What Your Service Requires

The FTC's Illuminate Education order made data minimization an enforcement priority, not a best practice. The order required the company to "refrain from collecting, processing, or maintaining any Covered Information not reasonably necessary to provide products or services" under its contracts. This is the standard the FTC is now imposing on EdTech vendors, and it directly constrains AI model training: if you do not need a particular data field to deliver your educational service, you should not be collecting it — and you certainly should not be training models on it.

For EdTech startups, this means auditing your data collection against your actual service requirements. If your tutoring platform needs student performance data to adapt lessons, collect performance data. But do not collect location data, browsing history, or behavioral metadata unless your service genuinely requires it — and if it does, document why.

3. Deletion on Request and Contract Termination

FERPA, COPPA, SOPIPA, and Texas SB 1792 all require deletion of student data — but they differ on triggers and timelines. FERPA requires that schools direct vendors to return or destroy education records when the contract ends. SOPIPA requires deletion at the school's request. Texas SB 1792 requires deletion when a student leaves the district or upon request. COPPA requires deletion of data when it is no longer needed for the purpose for which it was collected.

Your compliance architecture must support complete deletion across all systems — production databases, analytics pipelines, backups, and training datasets. This is the same deletion capability we discuss in our guide to neural data privacy compliance, but applied to education records. If student data has been incorporated into a trained model, deletion is technically complex: you may need to retrain the model without that student's data, or demonstrate that the data has been sufficiently de-identified in the model that it no longer constitutes an education record. Build your data architecture to support this from the start — retrofitting deletion into a model that has already been trained on student data is expensive and may be impossible.

4. Model Auditability and the Ability to Demonstrate Compliance

School districts, state regulators, and the FTC will increasingly ask: what data went into your model, and can you prove it was used only for authorized purposes? Your startup needs the ability to produce a training data provenance record — what data was used, from which districts, under what authority, and for what purpose. This is analogous to the Software Bill of Materials (SBOM) that investors now expect in open-source compliance, but applied to AI training data. If a district asks whether their students' data was used to train a model that serves other districts, you need to answer precisely — not approximate.

Building AI-powered edtech and need to ensure FERPA, COPPA, and state student privacy compliance before district procurement asks? We help EdTech startups structure data agreements, consent flows, and model training policies that pass district, state, and federal scrutiny.

Get in touch

Actionable Next Steps

  1. Audit your data collection against FERPA's school official criteria. For each school district you serve, confirm that your DPA satisfies the four requirements of 34 CFR § 99.31(a)(1): you perform a service the school would otherwise use employees for, you are under the school's direct control, you use data only for authorized educational purposes, and you do not re-disclose. If your DPA does not explicitly address AI training, it is not sufficient.
  2. Determine whether COPPA applies to your app. If your service is used by students under 13 — even if the school authorized the use under FERPA — COPPA requires verifiable parental consent for collection of personal information. Assess whether your current consent flow satisfies COPPA's requirements and whether it explicitly covers AI model training as a use of children's data.
  3. Map the state laws that apply to your customer base. Identify every state where your platform is used and map the applicable student privacy laws. If you serve California districts, SOPIPA applies. If you serve Texas districts, SB 1792 applies. Build your compliance posture to the most restrictive standard across all states you serve.
  4. Update your DPAs to address AI training explicitly. Do not rely on general-purpose data processing language. Add a specific clause that addresses whether and how student data may be used for AI model training, what safeguards prevent re-disclosure through model outputs, and what happens to the model when the contract terminates.
  5. Implement data minimization as a design principle. Audit every data field you collect against the educational purpose it serves. The FTC's Illuminate Education order makes clear that collecting more data than your service requires is an enforcement risk, not a competitive advantage. Delete what you do not need.
  6. Build deletion into your data pipeline from the start. If student data has been used to train a model, deletion requires either retraining without that data or demonstrating de-identification. Design your architecture to support this before a district or regulator asks — not after.
  7. Create a training data provenance record. Document what data was used to train each model, which districts contributed, under what authority, and for what purpose. This record is your answer when a district, state regulator, or the FTC asks whether their students' data was used appropriately.
  8. Get legal review before you train. The cost of a pre-deployment compliance review is a fraction of the cost of an FTC enforcement action, a terminated district contract, or a FERPA complaint that puts your customers' federal funding at risk. Have counsel review your DPA template, consent flows, data architecture, and AI training policies before you scale — not after a regulator asks.

FERPA was written for a world of paper records and filing cabinets. AI model training on student data is a use that Congress did not contemplate in 1974 and that the Department of Education is only beginning to address through guidance. For EdTech startups, the compliance strategy is not to wait for regulatory clarity — it is to build to the most restrictive standard, document your compliance posture, and ensure that every use of student data is explicitly authorized, minimized, and auditable. The startups that do this will win district procurement cycles, pass investor diligence, and avoid the enforcement actions that are already reshaping the EdTech landscape.