Artificial intelligence (AI) systems have transformed industries, leading to innovative solutions and improved efficiency in countless fields. From advancements in healthcare diagnostics to personalized customer service, the potential of AI is immense. However, with these groundbreaking benefits come significant challenges, particularly when it comes to data privacy. AI systems, often unpredictable, can make automated decisions that may unintentionally lead to issues like discrimination, raising concerns about transparency and fairness.
As of March 2025, EU data protection authorities had issued more than 2,200 fines under the GDPR totalling roughly €5.6 billion, with AI-related data processing cases among the fastest-growing categories of enforcement action (EDPB enforcement statistics).
In response to these growing concerns, regulatory bodies started drafting new legislation around the world, including in the EU. The EU AI Act governs the development, provision, and use of AI systems. Given that AI relies heavily on data, often including personal data, one of the central questions of this new regulatory era is how to lawfully use data collected for AI purposes. Confirming compliance with these new rules is crucial to balancing the power of AI with the protection of individual privacy rights.
Recent Developments in AI and Data Privacy
TL;DRGlobal regulators are actively scrutinizing AI data practices. Meta halted EU AI plans after regulatory pressure, X paused AI training on EU user data, and supervisory authorities in Singapore and the UK have published official guidelines. Compliance expectations are tightening fast across jurisdictions.
Supervisory authorities around the world have already begun to scrutinize the use of AI, with global tech giants like Meta and X (formerly Twitter) at the forefront of regulatory action. As these companies process vast amounts of data, they are naturally among the first to face heightened regulatory scrutiny.
Meta's Global AI Tensions
Meta's use of AI has sparked tension, particularly between the company and the EU regulators. After complaints from Noyb, Meta was forced to halt its AI-related plans in the EU. The company delayed the use of personal data for AI development and improvement "following consultations with regulators." Meta also assured users they would be notified in advance of any changes and granted the option to refuse data processing for such purposes.
Additionally, Meta decided not to release an advanced version of its AI model, LLama, within the EU, citing the "unpredictable" actions of regulators as the reason for this decision.
In contrast, in Brazil, the landscape is different. On August 30, 2024, Brazil's data protection authority (ANPD) lifted an initial ban that had restricted Meta from using personal data to train its AI models. This demonstrates the varying approaches to AI regulation around the world, with some countries focusing on strict privacy protections, while others are more open to fostering innovation.
On August 23, 2024, the CEOs of Meta and Spotify issued a joint statement, warning that the EU risks falling behind if it does not adopt a more progressive stance toward AI technology.
X's Updated Privacy Policy
X has also stirred attention with its recent privacy policy update, outlining its intent to use collected data for machine learning and AI training. However, in response to EU pressures, the company agreed not to use personal data from EU users for AI training until it enables them the possibility to withdraw their consent.
Guidelines from Supervisory Authorities
In light of these developments, several countries have started to release official guidelines to regulate the use of personal data for AI purposes. Singapore, for example, has issued Advisory Guidelines on Use of Personal Data in AI. Meanwhile, in the UK, the Information Commissioner's Office (ICO) has published a set of guidelines and answers to frequently asked questions on the processing of personal data for AI.
As stricter AI rules are most likely here to stay, this checklist will guide you through the 9 steps necessary to confirm that your use of personal data in AI development is fully compliant.
Implement Technical and Organizational Measures
TL;DRNo universal security solution exists for AI systems. Organizations must tailor technical measures, including encryption, access controls, and regular audits, to the specific context. Where feasible, anonymization eliminates most privacy obligations; pseudonymization reduces risk while keeping data usable.
To address the potential risks associated with AI, implementing appropriate technical and organizational measures is essential. There is no universal solution; security measures must be tailored to the specific context and nature of the AI system in use.
Effective measures to consider include:
- Data encryption to protect sensitive information
- Regular audits and assessments
- Access control mechanisms
- Confirming cyber-security
Wherever possible, one of the most effective strategies is to use anonymous data. Anonymous data cannot be linked back to an individual, making it exempt from privacy regulations in most jurisdictions. By using anonymized data, AI developers can significantly reduce their legal obligations under privacy laws. This approach not only protects individual privacy but also simplifies compliance.
If complete anonymization is not feasible, pseudonymization offers another valuable method of protecting personal data. Pseudonymization involves altering data in a way that prevents the identification of individuals without additional information, for example, replacing identifying fields such as names with unique identifiers. Pseudonymized data is still considered personal data under most regulations. As a result, it remains subject to legal protections, although with reduced risk compared to fully identifiable information.
Choose the Right Legal Basis for Data Processing in AI Systems
TL;DREvery AI data processing activity requires a valid legal basis under the GDPR. A single AI model may involve multiple processing operations, each needing separate justification. The six available bases range from consent and legitimate interest to vital interest, and the correct choice depends on the specific AI use case.
One of the fundamental steps in confirming compliance with AI and privacy regulations is establishing a valid legal basis for processing personal data. All data processing activities must be grounded in one of the six legal bases defined by privacy laws, such as the GDPR. However, determining the appropriate legal basis for each specific case can be complex, as it requires a deep understanding of both the AI model's purposes and the regulations in question.
A single AI model may involve multiple data processing operations, such as development/training, deployment, and auditing, and each of these activities might require a different legal basis.
Different AI models can also use personal data for various purposes, meaning that the applicable legal basis for each AI system may differ. This requires careful consideration when drafting privacy policies, as each processing activity needs to be justified by the appropriate legal basis.
The six legal bases are:
- Consent: commonly used but comes with strict requirements; it must be specific, granular, freely given, and easily withdrawn at any time. Consent is more feasible in situations where there is direct contact with the data subject. For example, if you are personalizing services for users, consent may be appropriate since you can clearly communicate the AI's purpose and allow users to opt in. However, consent becomes impractical in scenarios where you are scraping publicly available data to train an AI model. Without direct contact, it would be nearly impossible to acquire consent from every individual whose data is being used.
- Performance of a contract: relevant only if processing personal data is absolutely essential for fulfilling a contract with the data subject. It cannot be used simply because the AI model enhances or personalizes the service. For example, if your AI system offers personalized recommendations, this may improve the user experience, but it is not likely to be considered a necessity for the performance of the contract itself. On the other hand, if the core functionality of a service directly depends on AI, such as an AI-powered language translation app, then this basis may be appropriate. Generally, this legal basis should be avoided if there is a reasonable alternative for data subjects to access the service without AI.
- Legitimate interest: one of the more flexible legal bases, but it requires careful documentation and justification. You must conduct a Legitimate Interest Assessment (LIA), which involves three tests. First, the Purpose test: is there a legitimate purpose for the processing? Second, the Necessity test: is the processing necessary to achieve this purpose? Third, the Balancing test: do the interests of the company outweigh the privacy rights of the data subjects? The balancing test is critical, as the legitimate interest of the company must not override the rights and freedoms of individuals. For AI, this basis is often used in non-intrusive processing activities, but it requires transparency and proper safeguards.
- Legal obligation: applies when processing is necessary to comply with a legal obligation. It is likely to be more relevant in auditing or testing phases, where laws or regulations may require specific measures to confirm fairness or accuracy in AI outputs.
- Public interest: generally reserved for public authorities or organizations performing tasks that serve the broader public. For AI systems, this might be applicable in sectors like public health or law enforcement, where AI is used to fulfill governmental objectives. It is unlikely to be relevant for most private companies.
- Vital interest: rarely used but can be relevant during the deployment phase of an AI system, particularly in healthcare scenarios, such as AI used for medical diagnostics or emergency response systems.
Take a Risk-Based Approach
TL;DRThe EU AI Act imposes obligations proportional to an AI system's risk level. Because AI qualifies as a new technology, a Data Protection Impact Assessment (DPIA) is likely required whenever personal data is processed. Developers must also account for accidental personal data processing during development.
The EU AI Act adopts a risk-based approach, meaning that the obligations imposed on AI developers and providers vary depending on the level of risk an AI system poses. This approach echoes the privacy-related obligations that already exist under regulations like the GDPR, where companies must conduct a Data Protection Impact Assessment (DPIA) in certain cases.
A DPIA is required when data processing is likely to result in a high risk to the rights and freedoms of individuals, particularly when using new technologies, when automated decision-making is involved, or where it includes large-scale processing of sensitive data. Since AI is considered a new technology, the same principle likely applies to AI systems. Conducting a DPIA helps identify, assess, and mitigate risks associated with AI models, confirming compliance and minimizing harm to individuals.
AI systems that use personal data, especially sensitive categories of data, often present higher risks, as they can affect individuals whose data was used in both the development and deployment stages of the AI model. It is also crucial for AI developers to understand that they may accidentally process personal data during the course of their work, particularly when it is difficult to separate personal data from other datasets. Understanding high-risk AI systems classification is essential for developers navigating these obligations.
Confirm Transparency
TL;DRBoth the EU AI Act and the GDPR require organizations to inform users when they interact with AI or view AI-generated content. Where personal data is not needed for the AI system to function, adding a disclaimer advising users not to input personal data provides an additional layer of privacy protection.
Under the EU AI Act, transparency is a fundamental requirement for AI systems. Users typically need to be informed when they are interacting with an AI system or when the content they are viewing is AI-generated. In addition, the GDPR requires that data subjects be informed of the purpose for which their personal data is collected and processed. AI literacy plays a key role in helping users understand these disclosures.
If personal data is not required for the operation of the AI system, it is highly recommended to go one step further and include an additional disclaimer advising users not to input any personal data when using the AI system. This extra precaution can significantly reduce the risk of unnecessary personal data processing and provide added protection for users' privacy.
Comply with the Data Minimization Principle
TL;DRCollecting excessive personal data increases both privacy risk and compliance burden. Organizations should collect only what is strictly necessary for the AI system's purpose. Federated learning offers a practical technique that trains models across devices without centralizing personal data.
The data minimization principle is crucial in AI development and deployment, confirming that only the necessary personal data is collected and used. Collecting excessive or irrelevant data not only increases the risks of privacy violations but also creates unnecessary obligations for compliance under data protection laws.
Overly intrusive practices, such as using data from users' private chats to train AI models, should be avoided unless absolutely necessary.
A great way to support data minimization is through federated learning. This technique allows AI models to be trained across multiple devices without centralizing the data, confirming that only the necessary model updates are shared instead of the personal data itself. This helps reduce the amount of personal data collected while still enabling the AI to learn and improve its performance.
Enable Human Review of Automated Decision-Making
TL;DRThe GDPR grants individuals the right to request human intervention in automated decisions that significantly affect them. Oversight takes two forms: monitoring the AI system for accuracy and bias, and providing a human alternative for high-stakes decisions. Biased training data remains a leading cause of discriminatory AI outputs.
Under the GDPR, automated decision-making that significantly impacts individuals, such as decisions made by AI systems, must be carefully regulated. One of the key requirements is confirming that individuals have the right to request human intervention in any decisions made solely by automated processes.
Beyond GDPR requirements, human oversight in AI systems typically comes in two forms. First, oversight of the AI program itself: regular monitoring and assessment of the AI system to confirm that its outputs are accurate, unbiased, and in compliance with regulations. Second, bypassing or replacing the AI process: in certain cases, human decision-making can replace the AI system. While this is not always feasible, it is advisable whenever possible, especially for high-stakes decisions. Adopting a responsible AI framework strengthens these oversight mechanisms.
This oversight is particularly important when AI systems have the potential to lead to discrimination. The risk of discrimination is often heightened if the training data is unbalanced or reflects biased societal patterns.
Confirm that Data Transfers are Compliant
TL;DRCross-border data transfers require clear role definitions (controller vs. processor), appropriate agreements such as DPAs, and safeguards like Standard Contractual Clauses. AI developers may act as either controllers or processors depending on their level of decision-making authority over the data.
When transferring personal data, especially across borders, it is essential to meet the stringent requirements set by the GDPR. One of the first steps is to determine the roles involved in data processing, as the legal obligations differ depending on whether you are acting as a data controller, data processor, or joint controller. AI developers and deployers can be both data controllers and data processors, depending on the situation.
Example 1: An AI company develops a customer service chatbot that collects and processes user queries and personal data directly from customers. The company determines what data is collected, how it will be used, and for what purpose, such as improving the chatbot's responses and providing personalized recommendations. Since the AI developer is deciding the purpose and means of processing, they are acting as a data controller.
Example 2: A healthcare provider hires an AI company to develop a system that analyzes patient data for diagnostic purposes. The healthcare provider determines the purpose of processing (diagnosis), and the AI developer only processes the data following the instructions given by the healthcare provider. In this case, the AI developer is acting as a data processor, as they do not decide the purpose of the processing but only follow instructions.
Once roles are clearly defined, it is critical to sign the necessary agreements, such as a Data Processing Agreement (DPA) or Joint Controllership Agreement. These agreements confirm that each party, whether they are a processor or controller, understands their responsibilities for data protection and privacy.
If personal data is being transferred to a third country that may not provide adequate protection, additional safeguards are required. This often involves implementing Standard Contractual Clauses (SCCs), which are templates approved by the European Commission to confirm that data recipients in third countries provide adequate protection. Companies are also expected to conduct a Data Transfer Impact Assessment (DTIA), which evaluates the risks involved with transferring data to certain regions and confirms that proper safeguards are in place. Maintaining a Record of Processing Activities is a prerequisite for demonstrating compliance with these transfer obligations.
Is Data Scraping via AI Allowed?
TL;DRData scraping for AI training raises serious GDPR challenges because obtaining individual consent from vast public datasets is impractical. Legitimate interest is the most common legal basis, but its validity remains contested. The Dutch DPA questioned whether commercial interests qualify; the European Commission urged a more balanced view.
AI technologies can be highly effective tools for data scraping, gathering large amounts of data from publicly available sources. However, while scraping may seem straightforward, it presents significant legal challenges, especially under the GDPR. One of the main obstacles is obtaining consent from the data subjects whose information is being scraped, which is nearly impossible when dealing with vast datasets from public sources.
This leaves legitimate interest as the primary legal basis for such data processing. However, there has been growing debate about whether this legal basis can be relied upon in the context of data scraping for AI development. The Dutch Data Protection Authority expressed skepticism, stating that "commercial interests" cannot be qualified as a legitimate interest under the GDPR. If this stance is adopted by other supervisory authorities, AI developers may face difficulties using legitimate interest to justify data scraping.
However, the European Commission has pushed back on this position, urging the Dutch DPA to reconsider. According to the Commission, commercial interests should be considered legitimate, provided that a proper balancing test is conducted to confirm that they do not override the fundamental rights and freedoms of data subjects. Understanding the EU AI Act penalties framework is equally important for organizations that scrape data at scale.
Enable Continuous AI Governance
TL;DRAI compliance is not a one-time exercise. Continuous governance requires ongoing collaboration between legal and technical teams, regular audits, impact assessments, and system reviews. As both AI technologies and regulations evolve rapidly, proactive monitoring is essential to mitigate emerging risks before they escalate.
Confirming full compliance with AI and privacy regulations requires continuous AI governance. This involves an ongoing collaboration between the legal and tech teams. As AI technologies and legal frameworks are rapidly evolving, there are still many unresolved issues that require careful navigation. The legal team's role is to stay on top of these developments, confirming that the organization adheres to new laws and regulations, while the tech team adjusts and aligns AI development and deployment strategies accordingly.
A key aspect of this governance model is proactive monitoring and adaptation. As AI systems advance, they may introduce new risks or compliance challenges, especially regarding data usage and privacy concerns. Continuous governance enables the business to be flexible, mitigating risks before they become critical issues. Regular audits, impact assessments, and system reviews are essential tools to confirm that both the AI systems and the data they rely on remain compliant over time. Organizations should also confirm they have GDPR data breach notification procedures in place before an incident occurs. When sharing data with vendors, completing vendor security questionnaires is a critical step in due diligence. For a comprehensive overview of privacy obligations, see our GDPR compliance guide.
When engaging external processors for AI development, ensure a robust data processing agreement is in place that covers AI-specific obligations.
Further Reading on Whisperly

Reviewed by: Tamara Zavisic, AI Governance Specialist