Singapore AI Training Data Guidelines

Singapore AI Training Data Guidelines: What the PDPC’s Rules Mean For Businesses

On 23 July 2026, Singapore’s Personal Data Protection Commission published its Advisory Guidelines on the Use of Personal Data in Generative AI, alongside a new Federated Learning Guide and an expanded set of use cases for the Infocomm Media Development Authority’s Privacy-Enhancing Technologies Sandbox. 

Officials detailed the rollout at the IAPP Asia Forum 2026, and IAPP’s coverage of the announcement is the main source for this article. If your organisation trains, deploys, or licenses generative AI systems that touch Singapore personal data, these releases set out what the PDPC now expects at each stage of that process.

Key Takeaways

The PDPC’s Advisory Guidelines cover the development, deployment, and post-deployment stages of the generative AI lifecycle, with the deployment stage setting out separate requirements for model providers, system providers, and system deployers.

Web scraping for AI training can rely on the PDPA’s Publicly Available Exception only where the data carried no “digital barriers” during collection, and the PDPC now expects organisations to document the purpose of collection and the steps taken to mitigate risk.

Organisations need consent to process personal data that isn’t publicly accessible, and they need fresh consent if they want to reuse data collected for one purpose to train an AI system for a different one.

Alongside the guidelines, the PDPC released a Federated Learning Guide, and IMDA added new PET Sandbox use cases focused on health data and payment and financial data processing.

The PDPA’s financial penalty regime allows fines of up to S$1 million or 10% of an organisation’s annual turnover in Singapore, whichever is higher, so gaps in AI training data practices carry real financial exposure.

What Are Singapore’s New AI Training Data Guidelines?

The Advisory Guidelines on the Use of Personal Data in Generative AI set out how the PDPA applies at each stage of building and running a generative AI system: development, where training data is collected and processed; deployment, where the system goes live; and post-deployment, where it continues to process personal data in production.

The deployment stage is where the guidelines get specific about who owes what obligation. Model providers, the organisations that build the underlying AI model, system providers, who package that model into a usable product, and system deployers, who put the finished system to use, each carry distinct requirements. An organisation licensing a third-party model and deploying it internally needs to know which of these roles it occupies, since the compliance obligations differ depending on the answer.

PDPC Commissioner Denise Wong framed the guidance as a response to a widening gap between how fast AI is being adopted and how confident organisations feel about their own compliance. “Regulators do not have all the answers,” she said at the IAPP Asia Forum 2026. “But we believe regulating well means providing clarity. This enables society to adopt technology with confidence.”

What Is the Publicly Available Exception and How Does It Apply to Web Scraping?

The PDPA’s Publicly Available Exception lets organisations process accessible personal data without consent, which matters directly for the web-scraping processes many organisations use to assemble AI training datasets. Under the new guidelines, data only counts as publicly accessible for this purpose where there were no restrictions or “digital barriers” in place during collection. A site with a login wall, a paywall, or technical measures blocking automated access does not fall within the exception, regardless of how the data is later used.

The PDPC ran a public consultation on this guidance and received feedback from 40 organisations. Several raised a practical concern: once an organisation learns its data is being scraped, it could add new access restrictions specifically to undermine the exception going forward, creating uncertainty for anyone relying on it. In its response to that feedback, the PDPC confirmed that organisations using the exception must be able to clarify why they collected the data and what steps they took to mitigate the risks that collection created, rather than treating the exception as a blanket justification.

When Do Organisations Need Consent for AI Training Data?

Where personal data isn’t publicly accessible under the exception above, organisations need consent from the individuals concerned before using it to train an AI system. The guidelines also close a gap that trips up a lot of AI projects: if an organisation already holds personal data with consent for one purpose, such as customer service records, and wants to use that same data to train a generative AI model, that counts as a new purpose requiring its own, separate consent. Previous consent for the original purpose doesn’t carry over automatically.

This distinction matters most for organisations sitting on large stores of historical data that look like a convenient, low-cost training set. The guidelines treat that reuse as a fresh processing decision requiring its own consent, whatever consent was collected years earlier for the original purpose.

What Is the PDPC’s Federated Learning Guide?

The PDPC’s new Federated Learning Guide covers a decentralised training technique that lets organisations build AI models without moving sensitive data to a central location. Instead of pooling raw data from multiple organisations into one dataset, federated learning keeps each organisation’s data on its own systems and shares only the resulting model updates. The guide describes this as a way to “keep input data local and share only model updates,” letting organisations collaborate on model development. At the same time, each retains control of its own dataset.

For sectors where data sharing is legally or commercially sensitive, healthcare and financial services being the clearest examples, this gives organisations a route to collaborative AI development that doesn’t require handing a dataset to a third party outright.

What New Use Cases Did IMDA Add to Its PET Sandbox?

IMDA’s Privacy-Enhancing Technologies Sandbox lets organisations across sectors test and implement privacy-enhancing technologies in a supervised setting. The latest expansion adds use cases focused on protecting sensitive health data and on payment processing and financial data. In these two categories, the cost of a data protection failure is especially high.

IMDA Chief Executive Officer Ng Cher Pong linked the expansion directly to the growing autonomy of AI systems: “As AI systems become more autonomous and increasingly connected to sensitive data, internal systems and critical business processes, the risks are heightened. In this environment, the role of Privacy Enhancing Technologies, or PETs, will become increasingly important. They are not simply a niche set of technologies but are becoming essential infrastructure for trusted and responsible data use.”

How Do These Releases Fit Into Singapore’s Broader AI Strategy?

The guidelines, the Federated Learning Guide, and the PET Sandbox expansion all sit inside Singapore’s National AI Strategy, which positions the country as a hub for both AI innovation and data protection. Alongside the AI-specific releases, Singapore’s Cyber Security Agency revised its Cybersecurity Code of Practice for Critical Information Infrastructure to address advanced persistent threats and AI-enabled harms, extending the same period’s regulatory activity into security as well as data protection.

The IAPP and IMDA also signed a memorandum of intent on 22 July 2026 to deepen cooperation on AI, data protection, and digital responsibility, including professional development for practitioners working in the space. Singapore’s Minister for Digital Development and Information, Josephine Teo, tied the rollout to the country’s upcoming ASEAN chairmanship: “As Singapore assumes the Association of Southeast Asian Nations Chairmanship next year, we will work with regional partners to align our approaches, build trusted data flows, and grow our digital economies together.”

Wong made a similar point about scale during her keynote, noting that the PDPC’s workload “hasn’t gotten smaller” as AI adoption accelerates, and that “the data ecosystem just gets bigger.”

What Happens If a Business Doesn’t Comply With the PDPA?

The PDPA’s financial penalty regime allows the PDPC to fine an organisation up to S$1 million or 10% of its annual turnover in Singapore, whichever is higher, for breaches of the Data Protection Provisions. For an organisation with more than S$10 million in Singapore turnover, that 10% figure can significantly exceed the flat S$1 million cap that applies to smaller organisations. Gaps in how AI training data was sourced, whether that’s relying on the Publicly Available Exception without proper documentation or reusing consented data for an unconsented new purpose, fall squarely within the kind of breach this penalty regime targets.

What Should Businesses Using AI in Singapore Do Now?

Audit any AI training datasets built through web scraping and confirm the source data carried no digital barriers at the point of collection, then document why the data was collected and what risk mitigation steps were taken.

Map which of your AI systems rely on consent, and check whether that consent covers AI training specifically, or whether it was originally given for an unrelated purpose that doesn’t carry over automatically.

Work out whether your organisation is acting as a model provider, system provider, or system deployer for each AI system in use, since the PDPC’s deployment-stage obligations differ by role.

Assess whether federated learning or IMDA’s PET Sandbox could reduce the compliance burden of a data-sharing arrangement your organisation is already planning, particularly in health or financial data contexts.

If your organisation operates critical information infrastructure, review the revised Cybersecurity Code of Practice alongside the AI training data guidelines rather than treating the two as separate compliance tracks.

Conclusion

Singapore’s latest round of AI and data protection releases moves the PDPC from general principles to specific rules for how personal data can be used to train, deploy, and run generative AI systems. The Publicly Available Exception now has real conditions attached to it, consent doesn’t automatically transfer to a new AI purpose, and organisations building or licensing AI systems that touch Singapore data need to work out where they sit in the model provider, system provider, and system deployer chain. Reviewing how your organisation’s AI training data was actually sourced, not just how the vendor contract describes it, is the practical starting point.

Frequently Asked Questions

Does the Publicly Available Exception cover all web-scraped data used for AI training?

No. The exception only applies where the data carried no restrictions or digital barriers during collection, such as a login wall or a technical block on automated access. Organisations relying on the exception also need to be able to explain why the data was collected and what steps they took to mitigate the resulting risk.

What’s the difference between a model provider, system provider, and system deployer under the guidelines?

A model provider builds the underlying AI model. A system provider packages that model into a usable product or service. A system deployer puts the finished system to use within its own organisation. The PDPC’s deployment-stage obligations apply differently depending on which role an organisation occupies, and many organisations occupy more than one role at once.

Do organisations need separate consent to reuse existing data for AI training?

Yes, where that data was originally collected with consent for a different purpose. The PDPC treats training an AI system as a new purpose in its own right. Hence, an organisation needs fresh consent covering that specific use rather than relying on consent given years earlier for something unrelated.

What are Privacy-Enhancing Technologies and why do they matter for AI?

Privacy-Enhancing Technologies, or PETs, are technical methods, including federated learning, that let organisations use or share data while reducing the risk of exposing the underlying personal data itself. IMDA’s PET Sandbox lets organisations test these methods under supervision, and its newest use cases focus specifically on health data and payment and financial data.

Disclaimer: This blog post is intended solely for informational purposes. It does not offer legal advice or opinions. This article is not a guide for resolving legal issues or managing litigation on your own. It should not be considered a replacement for professional legal counsel and does not provide legal advice for any specific situation or employer.

Zlatko Delev

About the Author

Zlatko Delev

Country Manager & Head of Commercial — GDPRLocal

Zlatko specialises in data protection compliance, ISMS strategy, and AI law. With a legal background and hands-on experience supporting organisations globally, he helps businesses navigate GDPR, the EU AI Act, and international privacy frameworks.