← All Insights
Cloud

NPC Advisory 2026-01: What Philippine Businesses Must Know About Data Scraping Compliance

August 6, 2026 · 5min read  · Technica Solutions Inc.

NPC Advisory 2026-01: What Philippine Businesses Must Know About Data Scraping Compliance

The Assumption That Broke

For years, a common working assumption inside many Philippine IT and marketing teams went something like this: if data is publicly visible on a website, it is fair game. You can collect it, store it, aggregate it, train on it. After all, anyone with a browser can see it.

On April 13, 2026, the National Privacy Commission (NPC) drew a clear line through that assumption.

NPC Advisory No. 2026-01 — Guidelines on Scraping of Publicly Available Personal Data — reaffirms what the Data Privacy Act of 2012 has always said but many organisations have quietly ignored: public availability does not equal consent, and it does not place data outside the DPA. The full text is available at privacy.gov.ph.

If your organisation collects data from websites, runs scraping pipelines, hosts public-facing user data, or trains AI models on aggregated web content, this Advisory directly affects you.


What the Advisory Actually Says

The NPC defines data scraping broadly: any automated or manual process of extracting publicly available personal data from websites, applications, or other online sources. This includes HTTP requests, HTML parsing, identifying and extracting specific data elements, and storing results in structured formats such as databases or JSON files. It covers text, images, audio, video recordings, and user profiles.

The Advisory applies to:

  • Personal Information Controllers (PICs) and Personal Information Processors (PIPs) that engage in data scraping
  • PICs that host publicly available personal data that may be subject to scraping by others

The core legal position is unambiguous. Personal data that is publicly available online remains fully covered by the DPA, its Implementing Rules and Regulations, and all NPC issuances. The fact that a user made information visible on LinkedIn, a company directory, or a public forum does not strip it of DPA protection. Organisations must have a valid lawful basis before scraping, storing, or further processing that data — and consent from the platform is not a substitute for consent from the data subject.


Who Is Affected

Data scrapers and aggregators

Any team running scheduled scraping jobs — whether for lead generation, competitive intelligence, price monitoring, or research — must identify a valid lawful basis under the DPA for each category of personal data collected. Legitimate interest can apply, but it must be documented and must pass a balancing test against the rights of the data subjects.

Organisations that host publicly accessible personal data

If your website publishes staff directories, customer testimonials with names, or any profile-style personal data, you are now expected to inform users that their data may be subject to scraping, provide mechanisms to object, and implement technical measures such as rate limiting and bot detection to deter unauthorised collection.

AI training pipelines

This is where the Advisory carries the most immediate operational weight. Organisations building or fine-tuning AI models on scraped web content — employee databases, social media posts, forum archives — must treat that data collection as a full DPA-regulated processing activity. A Privacy Impact Assessment is required. Proportionality must be demonstrated. If the scraped dataset includes sensitive personal information (health data, political opinions, religious affiliation), scraping is generally prohibited unless strict statutory conditions are met.


Five Compliance Actions to Take Now

1. Conduct or update your Privacy Impact Assessment. A PIA is not optional for data scraping activities. This applies even when you engage a third-party PIP to do the scraping on your behalf — accountability stays with your organisation as the PIC. Document the data flows, the lawful basis, the retention period, and the safeguards in place.

2. Identify and document your lawful basis. For each scraping use case, record which DPA provision you rely on. Legitimate interest, contractual necessity, and legal obligation are the most commonly applicable. "It's public" is not a lawful basis.

3. Update your privacy notice. If you source any personal data from publicly available locations, your privacy notice must now disclose: where the data came from, the purpose of collection, and the manner of processing. This applies to marketing databases built from LinkedIn, supplier lists compiled from government directories, and any AI training corpus assembled from web content.

4. Implement anti-scraping safeguards if you host public personal data. Bot detection, rate limiting, CAPTCHA, and clear notices that your site's personal data is protected under the DPA — these are now expected baseline measures for any PIC that publishes personal data publicly.

5. Review third-party PIP arrangements. If a vendor scrapes data on your behalf, you remain accountable. Review your data processing agreements to ensure the PIP is bound by the same DPA obligations. Existing contracts signed before April 2026 should be assessed against the Advisory's requirements.


The AI Training Data Angle

The Advisory arrives at a moment when Philippine organisations — particularly BPOs, fintech firms, and tech companies — are actively building proprietary AI tools. Web scraping is one of the most common methods for assembling training datasets cheaply and at scale.

The NPC has now made clear that using scraped personal data to train an LLM, a classification model, or a customer service chatbot is a regulated processing activity. Proportionality applies: the volume of data collected must match the stated purpose. Aggregation and profiling from scraped sources attract heightened regulatory scrutiny.

Organisations that built AI training pipelines before April 2026 and have not conducted PIAs on those datasets are potentially non-compliant today. This is worth reviewing before the NPC's enforcement posture hardens further — particularly following the mandatory breach notification framework established in NPC Advisory 2026-02.


What Happens If You Don't Comply

The DPA provides for civil, criminal, and administrative liability. Bypassing website safeguards, using deceptive techniques to obtain data, or violating platform terms of service can independently trigger liability — these practices are explicitly called out in the Advisory as unauthorised. The NPC has demonstrated willingness to investigate and pursue enforcement actions, and the Advisory signals an increasingly active regulatory stance on digital privacy.

Beyond regulatory risk, organisations that misuse scraped personal data expose themselves to reputational damage and potential civil claims from affected data subjects.


How Technica Solutions Inc. Can Help

Cloud governance and data privacy compliance are not afterthoughts — they are foundational to any well-run digital operation. Technica Solutions Inc. works with Philippine organisations to design Azure and Microsoft 365 environments where data governance controls are built in from the start: Microsoft Purview for data classification and audit trails, Entra ID for access governance, and Intune for endpoint policy enforcement.

If your team is evaluating AI tools or automation pipelines that rely on external data, we can help you design the compliance framework before you build — not after the NPC comes knocking. Speak to our team to discuss your specific environment.


Related Reading


This article is for informational purposes only and does not constitute legal advice. Organisations should consult qualified legal counsel for guidance specific to their situation.

Talk to our Cloud & I.T. team
Related Insights

More on Cloud

← Back to Insights