Last materially reviewed: September 6, 2026
Intellectual Property → AI-Generated Works → AI Training & Enforcement
Direct Answer
Web scraping copyrighted material for AI training is not automatically lawful in the Philippines merely because the material is publicly accessible online. The legal analysis depends on what is copied, how much is reproduced, why it is used, whether a copyright exception such as fair use applies, what contractual or technical restrictions govern access, and whether the use harms the normal market for the work.
Philippine law does not currently contain an AI-specific text-and-data-mining exception. Section 185 of the Intellectual Property Code provides a fact-specific fair-use test. IPOPHL’s 2026 copyright discussions have expressly identified fair use, data scraping and human authorship as live issues in the Philippine copyright landscape.
Legal Status
Developing / unsettled for AI training. Existing copyright rules clearly apply, but there is not yet a definitive Philippine court rule saying that large-scale scraping for generative-AI training is categorically permitted or categorically prohibited.
Key Takeaways
- Publicly viewable does not mean free of copyright.
- Scraping can involve acts of copying even when the final model does not display the source work.
- Fair use requires a case-by-case analysis under Section 185.
- Website terms, paywalls, authentication and technical restrictions may create additional legal issues.
- Businesses should document dataset provenance and licensing rather than assume all open-web data is safe.
What Copyright Questions Does Scraping Raise?
AI developers may collect text, images, music, code or other works into datasets, make intermediate copies, preprocess those works and use them during model training. Each step can raise different questions about reproduction, authorization and exceptions.
The fact that a crawler can technically reach a page does not resolve whether copying protected expression is permitted. Copyright attaches automatically to qualifying original works; registration is not what creates the right.
Does Fair Use Solve the Problem?
Not automatically. Philippine fair use considers the purpose and character of the use, the nature of the copyrighted work, the amount and substantiality used, and the effect on the potential market or value of the work. Commercial AI training may still qualify as fair use in some circumstances, but commercial purpose is one factor and not a complete answer.
See our dedicated guide: Does Fair Use Allow AI Training on Copyrighted Works?
What About Website Terms and Robots Rules?
A rights holder can use terms of service, account restrictions, paywalls, API conditions and technical controls to state or enforce limits on automated collection. A robots instruction or AI opt-out signal can be useful evidence of the owner’s preference, but its legal effect in the Philippines depends on the surrounding facts and does not itself rewrite the Copyright Code.
Risk Checklist for AI Developers
- Identify the source of every major dataset.
- Record whether material was licensed, public domain, user-provided or open-web.
- Preserve the license terms that applied when data was collected.
- Separate factual data from protectable expressive content.
- Assess paywalls, authentication, contractual restrictions and access controls.
- Test models for verbatim or near-verbatim reproduction.
- Maintain a process for rights-holder complaints and dataset removal requests.
For Rights Holders
If you believe your work was scraped, preserve dated copies of the original work, publication history, access restrictions, crawler logs if available, screenshots, correspondence, and examples of model outputs that appear to reproduce the work. See What Evidence Should Copyright Owners Preserve?
Separate Collection, Copyright and Access
A collection workflow may fetch pages, retain copies, extract text and pass material into a training dataset. Identify those steps accurately. Copyright economic rights include reproduction and transformation, subject to statutory limitations. Permission to view a page should not be treated as permission for every subsequent act. IP Code, Sections 171.9 and 177.
Copyright protects qualifying expression rather than ideas or mere data as such. A factual database may nevertheless include protected articles, photographs or original selection and arrangement. IP Code, Sections 173 and 175. This page addresses acquisition; the training-material guide explains the broader legal decision.
What Robots Instructions Actually Establish
The Robots Exclusion Protocol describes instructions that crawlers are requested to honor. It is not an access-authorization mechanism. Protecting confidential resources requires actual access controls, not simply listing their paths in robots.txt. RFC 9309, Introduction and Security Considerations.
Conversely, an allowed crawl is not a universal copyright license. Treat crawler instructions, website agreements, authentication and copyright permissions as separate records. Do not bypass a paywall or reuse credentials merely because a project needs more training data. Where access rights are disputed, obtain authorization rather than treating technical reachability as the answer.
A Collection Record That Can Be Reviewed
| Record | Why it matters |
|---|---|
| Source URLs and collection timestamps | Identify what was accessed and when |
| Relevant terms and permissions | Show the asserted basis and scope |
| Crawler settings and logs | Document the collection method |
| Content inventory | Separate complete works, extracts and facts |
| Dataset version and handoff record | Trace what entered the project |
| Exclusions and complaints | Show how flagged sources were handled |
Collect only necessary operational evidence and protect logs containing personal or sensitive information. The table is practical documentation, not a mandated statutory form. A log showing a request is not, by itself, proof that a specific model was trained on the response.
Three Hypothetical Collection Scenarios
Public Articles With No Account Requirement
A developer can read every page, but the pages contain original prose and images. Assess reproduction and any license or exception before bulk collection. A missing robots block does not establish permission. If relying on fair use, record the intended purpose, copied amount and other factors in the fair-use worksheet.
An API With an Express Data Agreement
A publisher grants API access for a particular service. Read whether the agreement permits training, retention and onward transfer, or only display to users. Authorized access can coexist with restrictions on later use. Seek a revised license if the proposed project exceeds the grant.
A Page Includes Both Facts and Photographs
A project needs factual values but downloads the full page with images. Review what the workflow actually stores, not only the intended final dataset. Narrowing retained material can improve the evidence of what was used, but should not be described as automatically resolving all earlier copying or access questions.
A Practical Collection Approval Workflow
- Define the sources, fields and copies the system will retain.
- Review copyright and access terms separately.
- Identify the actual permission or exception relied on.
- Limit collection to the approved scope and protect restricted data.
- Record dataset versions and who receives them.
- Provide a process for investigating rights-holder requests.
- Reassess before changing from research to commercial deployment.
For a rights holder, preserve original publication files, the relevant restrictions, logs and any communications before making a claim. State what the records show and what remains unconfirmed. A targeted request for clarification is more defensible than asserting that every crawler request proves infringement or training.
Use the opt-out guide for reservations and technical signals, and data licensing for negotiated access. Neither a notice nor a license negotiation guarantees a particular legal outcome. Escalate contested claims with the specific source records and proposed use.
Frequently Asked Questions
Is scraping facts the same as copying copyrighted expression?
No. Copyright generally protects original expression rather than facts or ideas as such. A dataset may contain both unprotected information and protected expression.
Does putting content online give AI companies permission to train on it?
Not by itself. Permission may come from a license, terms, an applicable exception or other legal basis; online availability alone does not settle the issue.
Is there a Philippine AI-training opt-out law?
There is no general Philippine statutory AI-training opt-out mechanism comparable to some emerging foreign regimes. Rights holders can still use contracts, technical controls, platform settings and copyright enforcement where applicable.
Official Sources
- Republic Act No. 8293 — Intellectual Property Code (our IP Code explainer)
- IPOPHL — Copyright
- IPOPHL — 2026 forum on fair use, data scraping and AI
- WIPO — Generative AI: Navigating Intellectual Property
- U.S. Copyright Office — AI study, including generative-AI training
Important: This article provides general educational information about Philippine law and technology. It is not legal advice and does not create an attorney-client relationship. Laws, agency procedures, platform terms and the facts of each situation may change the result. Verify current requirements through the cited official sources and seek qualified professional advice when your rights, deadlines, money or legal exposure may be affected. AI-training copyright law is evolving and fact-specific.
Featured image: Photo by Growtika via Unsplash.

