CyberCode.ph · Philippines

Stop AI Scraping Copyrighted Work in the Philippines

Last updated September 28, 2026 · Practical privacy, cybersecurity and technology-law guidance

Last materially reviewed: September 22, 2026

Intellectual Property → AI-Generated Works → AI Training & Enforcement

Direct Answer

Creators cannot guarantee that no AI crawler will ever copy publicly accessible material, but they can reduce bulk collection without abandoning search visibility. Use this sequence:

  1. Inventory the works worth protecting.
  2. Preserve originals, publication dates and ownership documents.
  3. Publish clear AI-training and licensing terms.
  4. Distinguish ordinary search bots from identified AI or high-volume crawlers.
  5. Test targeted robots.txt rules, rate limits, image or file controls and authentication only where appropriate.
  6. Confirm that Google and other intended search crawlers can still reach the public pages.
  7. Preserve crawler logs, dataset evidence and reproducible model outputs.
  8. Send a focused notice or escalate only after identifying the work, suspected act, responsible party and requested remedy.

These measures improve control and evidence; they do not guarantee removal from an existing dataset or trained model.

Start With Ownership Evidence

To stop AI companies from scraping copyrighted work in the Philippines, CyberCode recommends preserving the records that show what you created, when you created it, and who owns the rights before sending a complaint.

  • Original creation records: source files, drafts, project files, raw photographs, working layers, notes and version history
  • Date evidence: reliable timestamps, repository history, publication archives, invoices, emails and dated exports
  • Ownership chain: employment terms, commissions, licenses, assignments and collaboration agreements identifying who owns which rights
  • Registration or deposit records, if available: plus the exact version of the work covered
  • Online evidence: canonical URL, publication page, copyright notice, license and AI-use terms in force when access occurred
  • Technical evidence: dated robots.txt copies, server or CDN logs, user-agent and IP data, request paths and rate-limit events
  • Use evidence: dataset name and version, copied file or passage, exact prompts and outputs, model or service version, date, screenshots and a side-by-side comparison

Copyright protection generally arises automatically, but clear documentation can make enforcement faster and more credible.

Publish Clear AI-Use Terms

State whether automated collection, model training, fine-tuning, embedding, dataset creation or commercial reuse is authorized. If you offer licenses, explain how a developer can request one. Clear terms are more useful than a vague “all rights reserved” notice when the dispute is about AI training specifically.

Use Technical Controls Carefully

To reduce AI scraping without harming search visibility, CyberCode recommends matching each technical control to its purpose and keeping evidence of how it was used. Technical controls can limit future access, but they do not prove whether earlier copying or AI training occurred.

Control Use it for Key limit
robots.txt or crawler rules Signal restrictions to identified, compliant crawlers while allowing intended search bots. Voluntary rules can be ignored, user-agent labels can be spoofed, and robots.txt does not make public material private.
Rate limits, bot management or CDN rules Slow high-volume requests, repeated downloads or suspicious automated patterns without removing public pages from search. May affect legitimate users or search crawling if thresholds are too broad; test, monitor and preserve logs.
Login gate or authentication Protect premium, licensed or sensitive material that should not remain publicly accessible. Strongest access control, but gated content normally loses ordinary indexing and sharing visibility.
Metadata or content credentials Identify origin and rights information. Use them alongside ownership records and access controls.

Check that ordinary search indexing still works after making changes.

Do not block ordinary search indexing accidentally. SEO crawlers, search engines and AI-training bots do not always behave the same way.

Monitor for Reproduction, Not Just Crawling

Model output that reproduces protected expression can provide a more concrete infringement issue than a suspicion that a crawler visited a page. Preserve exact prompts, outputs, model names, versions, dates and screenshots when testing.

Send a Focused Notice

For suspected AI scraping, send a focused notice containing: the creator or rights holder’s name and contact details; a precise identification of each work and canonical source; the ownership basis and any registration or assignment records; the identified crawler, dataset, copied material or reproducible output; dates, logs, URLs, screenshots and comparisons; the license or AI-use terms relevant to the alleged access; the legal and factual basis stated without exaggeration; the specific action requested, such as investigation, preservation of records, removal from a dataset, output restriction, deletion where technically and legally available, or a licensing discussion; a reasonable response date; and a good-faith statement that the information supplied is accurate. Send it through the company’s published copyright, legal or platform-reporting channel and preserve the notice, attachments, delivery record and response. Avoid claiming that a crawler visit alone proves infringement.

See How to Send a Copyright Infringement Notice for AI-Generated Content.

Escalation Options

If informal contact fails, the available route depends on the evidence and remedy sought. Options may include the service’s copyright or platform complaint process; a licensing or contract demand where enforceable terms or an agreement apply; a lawyer’s cease-and-desist or preservation notice; an IPOPHL Intellectual Property Rights Enforcement Office report or verified complaint within its mandate; mediation or other dispute-resolution procedures where available; and a civil action seeking appropriate relief. Suspected criminal conduct or online piracy may require coordination with the proper enforcement authority. IPOPHL registration and the Bureau of Copyright and Related Rights can strengthen the administrative record, but registration does not automatically prove that a particular AI-training act infringed. Cross-border defendants, uncertain dataset provenance, fair-use questions and requests affecting an entire model usually require fact-specific legal advice.

Creator Checklist

  1. Inventory valuable works.
  2. Confirm who owns each right.
  3. Publish AI-training and licensing terms.
  4. Implement crawler and access controls.
  5. Preserve logs and source files.
  6. Monitor high-risk models or datasets.
  7. Capture evidence before contacting the other side.
  8. Choose the narrowest effective enforcement route.

Choose a Response Based on What You Can Show

First separate three objectives: preventing future access, investigating past collection, and stopping an identifiable infringing use. A crawler restriction addresses the first objective. It does not tell you whether older copies exist or whether a model was trained on them. An output complaint needs its own supporting evidence. Keep these objectives separate in correspondence so the recipient can investigate a specific request.

What you can show What it may establish What it does not establish by itself
A bot requested public URLs Crawling or attempted access That a protected copy was retained, used for training or infringed copyright
Your work appears in an identified dataset Dataset inclusion and a concrete copy to investigate That every related AI model trained on it or that no exception or license applies
An AI output reproduces substantial protected expression A stronger basis for comparing protected expression and investigating reproduction The precise source, training path or party responsible without further evidence
The service uses only the idea, facts, method or general style Similarity at an unprotected or potentially unprotected level Copying of protected expression; copyright protects expression, not ideas or style in the abstract
Match the available evidence to the next action
What you have Useful next step Limit
Unusual automated requests in server logs Preserve logs and review rate limits with the site administrator. A request does not establish completed copying or training use.
A copy of your work in an identified dataset Record the dataset version, location, acquisition details and applicable license. Dataset inclusion alone does not identify every model trained on it.
An AI service reproduces distinctive passages Preserve the full session and prepare a side-by-side comparison. The text may have come from retrieval, an upload or another source.

The Intellectual Property Code, Sections 172, 175, 177 and 185 protects qualifying expression and gives owners reproduction rights, subject to limitations and exceptions. Public availability is not itself a copyright license. Equally, labeling a use “AI training” does not resolve whether particular acts infringe or qualify as fair use. Identify the work, the copying alleged, the party involved, any permission, and the actual use before reaching a legal conclusion.

A website restriction and a contract are different things. Record the terms in force when access occurred and how the operator encountered or accepted them. A new notice does not by itself establish that an earlier visitor accepted a contractual restriction. Avoid saying that every crawler visit violates copyright or that a footer automatically binds all visitors.

Three Practical Examples

A photographer wants search visibility but fewer bulk downloads

Keep a record of original photographs and published versions. Ask the administrator to test targeted controls on image endpoints, monitor errors and preserve the previous configuration for rollback. Check ordinary browsing and intended search access after deployment. These are access-management steps, not a guarantee that previously downloaded images disappear.

A publisher sees a crawler ignore robots.txt

Save the applicable robots.txt version and dated logs. The Robots Exclusion Protocol, RFC 9309 describes crawler requests, not an access-authorization mechanism. Use authentication for material that must remain private. A crawler name in a user-agent string is a lead to verify, not conclusive identification of a company.

A novelist finds a recognizable passage in an AI answer

Preserve the answer before sending a notice. Record whether the session included browsing, uploaded files or supplied excerpts. Then follow the AI infringement evidence guide. A focused request can identify the passage and ask for investigation or output restriction without asserting that the entire model must have been trained on the novel.

A Manageable Action Plan

  1. List the valuable works and confirm who controls the relevant rights.
  2. Save current publication pages, terms and existing access records.
  3. Select and test targeted access controls; document the change date.
  4. Investigate concrete copies or outputs using reproducible records.
  5. Choose a proportionate request and track the response.

If the concern involves a vendor dataset, use the dataset provenance guide to ask about source and license records. Do not promise that blocking a crawler will remove information already incorporated into a trained model.

Frequently Asked Questions

Can I block every AI crawler?

No universal mechanism guarantees that. Technical controls only work against systems that respect or cannot bypass them lawfully.

Should I remove my work from the internet?

Usually not as a first response. That can sacrifice audience and search visibility. A layered rights-management strategy is generally more practical.

Is copyright registration required in the Philippines to enforce rights against AI scraping?

No. Copyright generally exists automatically once a qualifying original work is created, so registration is not a condition for copyright protection. IPOPHL registration or deposit can nevertheless strengthen the ownership record, identify the deposited version and make a complaint easier to organize. A creator still needs evidence of the work, ownership, the act complained of and why that act is legally actionable; registration does not turn every crawler request or AI-training use into infringement.

Official Sources

Important: This article provides general educational information about Philippine law and technology. It is not legal advice and does not create an attorney-client relationship. Laws, agency procedures, platform terms and the facts of each situation may change the result. Verify current requirements through the cited official sources and seek qualified professional advice when your rights, deadlines, money or legal exposure may be affected.

Featured image: Photo by Numan Ali via Unsplash.

CyberCode updates

Get practical updates on Philippine technology law, data privacy, cybersecurity, and AI.

Email activity tracking

Unsubscribe any time. See our privacy policy below.