CyberCode.ph · Philippines

Can Copyrighted Material Be Used to Train AI in the Philippines?

Last updated September 15, 2026 · Practical privacy, cybersecurity and technology-law guidance

Last materially reviewed: September 6, 2026

Intellectual Property → AI-Generated Works

Direct Answer

There is no simple Philippine rule saying that all copyrighted material may freely be copied for AI training. AI training can involve reproducing, storing, processing and transforming protected works, so the legal analysis should consider copyright ownership, licenses, statutory exceptions, the specific training process and the uses made of the resulting model and outputs.

Legal Status

Developing. IPOPHL has identified AI as a policy priority and the Philippines is actively considering how the IP framework should adapt. WIPO also treats training data as a major unresolved international IP issue.

Key Takeaways

  • Publicly accessible does not mean free of copyright.
  • Scraping permission and copyright permission are different questions.
  • Licensed, public-domain and first-party datasets generally offer a cleaner risk profile.
  • Training agreements should document data sources and rights.
  • Output risk remains separate from input/training risk.

Questions To Ask About a Training Dataset

  1. Who owns or controls the source material?
  2. Was the material licensed for machine learning or text-and-data processing?
  3. Does a statutory exception plausibly apply?
  4. Were website terms, robots controls or contractual restrictions relevant?
  5. Does the dataset include personal or sensitive personal information?
  6. Can the organization trace the source and removal history of the data?

Safer Dataset Strategies

Organizations developing or fine-tuning AI should consider datasets consisting of their own material, licensed material, public-domain works, appropriately licensed open content and data collected under documented rights. WIPO’s practical guidance likewise emphasizes dataset vetting and license coverage.

Training Data vs Output

Even if a training dataset is lawfully assembled, a specific model output may still raise infringement questions. Conversely, an output that looks original does not automatically resolve questions about the legality of the underlying training process.

Related Cybercode Guides

Identify the work and the act before deciding whether a license or exception is needed. Section 172 covers original literary and artistic works, while Section 175 excludes ideas, procedures and mere data as such. A dataset can contain both. Section 173 also recognizes original selection or arrangement without removing rights in the underlying works. IP Code, Sections 172–175.

Copying an article, extracting factual values and reproducing an original compilation are not interchangeable descriptions. Likewise, downloading a corpus, adapting its contents, training a model and publishing output can involve different acts. Section 177 provides the relevant economic-rights framework. IP Code, Section 177. This page addresses legal permission; the training-data business checklist addresses ongoing governance.

Choose and Document the Basis for Use

Proposed basis What to verify What not to assume
Your organization’s rights Employment, assignment and contributor records Having the files means owning every component
License The licensor’s authority and permitted acts A download license covers every training use
Unprotected material The specific facts or subject matter used A whole website is unprotected because it includes facts
Expired copyright Work type, applicable term and relevant jurisdictions Old material found online is necessarily public domain
Statutory exception Actual conditions and evidence for the particular use Calling a project research resolves the issue

This matrix is a practical review aid, not an official approval form. If the basis remains unclear, exclude the disputed portion while permissions or advice are obtained. The existence of a commercial license market does not itself decide every exception question, and uncertainty is not a blanket permission.

Three Hypothetical Training Projects

A Company Fine-Tunes on Its Own Manuals

A business has internal manuals written by employees and contractors. Check the rights allocation and any embedded photographs, quotations or licensed diagrams. Company possession does not eliminate a contractor’s retained copyright or third-party restrictions. Section 178 distinguishes employee and commissioned work. IP Code, Section 178. Record which version is approved before transferring it to a vendor.

A Researcher Uses Published Numerical Facts

The project needs numerical observations rather than the articles explaining them. Identify precisely what is collected. The unprotected status of facts does not automatically dispose of copied prose, illustrations or an original compilation. Keep the extraction specification and examples showing the material retained. Review access terms separately.

A Startup Downloads a Publisher’s Archive

The archive is readable without payment, but contains complete articles and images. Public access alone does not settle reproduction rights. The startup should identify the actual corpus, intended model use and proposed permission or exception. Route collection questions to the scraping guide and any fair-use claim to the four-factor assessment.

Keep a Dataset Decision File

For each material source, retain its identifier, acquisition date, rights owner if known, license version, permitted uses and any restrictions. Record which copies are retained, who can access them and which training runs use them. Preserve agreements rather than relying only on a vendor’s general assurance that its data is safe.

Distinguish evidence from inference. A crawler log can show access; it does not necessarily show completed model training. A matching output can raise a copying concern without revealing the entire training corpus. See AI copyright evidence for preserving a suspected incident.

Before the Training Run

  1. Define the exact data and planned acts.
  2. Separate protected expression, third-party assets and factual material.
  3. Identify the permission or exception for each source group.
  4. Check confidentiality and personal-data issues separately.
  5. Resolve uncertain sources through exclusion, permission or qualified review.
  6. Record the approved dataset version and permitted use.
  7. Reassess when the purpose, model, vendor or distribution changes.

Do not treat approval of a training corpus as approval of every generated output. Review publication separately through the output infringement guide. A license can reduce uncertainty, but only within its verified scope; use the training-license checklist when negotiating permission.

Controlling framework: AI-training questions must begin with the protectable-work, ownership, restricted-act and fair-use rules in the Intellectual Property Code guide.

Frequently Asked Questions

Is web scraping automatically legal for AI training?

No. Scraping can involve copyright, contract, privacy, cybersecurity and database issues depending on the circumstances.

Does fair use automatically cover AI training?

No automatic answer should be assumed. Philippine exceptions must be analyzed under Philippine law and the facts of the particular use.

Go Deeper: AI Training & Enforcement

Official Sources

Featured image: Photo by Logan Voss on Unsplash.

This is a developing area. Important: This article provides general educational information about Philippine law and technology. It is not legal advice and does not create an attorney-client relationship. Laws, agency procedures, platform terms and the facts of each situation may change the result. Verify current requirements through the cited official sources and seek qualified professional advice when your rights, deadlines, money or legal exposure may be affected.

CyberCode updates

Get practical updates on Philippine technology law, data privacy, cybersecurity, and AI.

Email activity tracking

Unsubscribe any time. See our privacy policy below.