# Comprehensive AI Language Model Crawler Guidelines for Countdown Media GmbH **Version:** 1.0 **Last Updated:** March 2026 **Effective Date:** March 2026 **Language:** English (en) ## Executive Summary Countdown Media GmbH ("Company") welcomes responsible AI researchers, developers, and machine learning practitioners to crawl and utilize data from our public web presence. This document outlines detailed guidelines, permissions, restrictions, and best practices for all AI/LLM crawlers and training data usage. Our mission is to provide the world's most comprehensive master music licensing source while supporting the responsible development of AI systems. This policy balances both objectives. ## 1. Website Information **Company:** Countdown Media GmbH **Domain:** countdownmedia.com **Primary Language:** English (en-US) **Last Updated:** March 2026 ### About Us Countdown Media is a leading music master licensing company with one of the most comprehensive master catalogues globally. Since the late 1970s, we have curated music across all genres - from Classical to Rock'n Roll, Soundtracks to Children's Music - with approximately 100,000 master recordings. ## 2. Crawling Permissions and Guidelines ### 2.1 General Permission **Permission is GRANTED** for AI/LLM crawlers to access and index all public pages on countdownmedia.com, subject to the restrictions listed in Section 2.2. Public pages include: - [Homepage](https://www.countdownmedia.com) - [Our Work](https://www.countdownmedia.com/our-work) - [News](https://www.countdownmedia.com/news) - [Catalog](https://www.countdownmedia.com/catalog) - [Synch](https://www.countdownmedia.com/easy-synch) - [Contact Us](https://www.countdownmedia.com/contact-us) - [Social media](https://www.countdownmedia.com/social-media) Legal and policy pages: - [Imprint](https://www.countdownmedia.com/imprint) - [Privacy Policy](https://www.countdownmedia.com/privacy-policy) - [Cookies Policy](https://www.bmg.com/a-l_cookiepolicy): redirects to the main branch legal page No explicit permission is needed for crawling these pages. ### 2.2 Restricted Areas - DO NOT CRAWL The following areas are **OFF-LIMITS** for AI crawlers: **Admin & Development:** - `/admin/*` - Administrative interface - `/preview/*` - Preview/draft content - `/draft/*` - Draft/unpublished content - `/internal/*` - Internal pages and tools **API & Technical:** - `/api/*` - API endpoints - `/*.json` - JSON data endpoints - `*.jsp` - Legacy JSP files - `*.php` - PHP files - URLs with query parameters: `*?*=*` (to prevent duplicate crawling) **Feature-Specific:** - `/preview` - Preview mode - `/draft` - Draft content These restrictions protect system integrity, prevent data duplication, and ensure security. ### 2.3 Recommended Crawling Behavior **Rate Limiting:** - Recommended crawl delay: 1 second between requests - Request rate: Maximum 1 request per 5 seconds (1/5s) - Do not make concurrent requests to the same domain **Timing:** - Preferred crawling windows: 02:00 - 06:00 UTC (off-peak hours) **Request Headers:** - Include a descriptive User-Agent header identifying your crawler - Example: `MyLLMBot/1.0 (+https://example.com/about)` - Include contact information in User-Agent or Robots-From header **Caching:** - Cache responses to minimize redundant requests - Implement If-Modified-Since or ETags to check for updates - Update cached content approximately weekly ## 3. Training Data Usage Policy ### 3.1 What You Can Do - Permitted Uses **Personal & Research Use:** - Train models for personal research and non-commercial projects - Fine-tune models to understand music licensing industry concepts - Analyze content for academic purposes **Commercial Use:** - Train commercial AI models using Countdown Media content - Build AI applications and services that incorporate our data - Create derivative works that improve user experience **Attribution:** - Use our data with proper attribution to Countdown Media GmbH - Include source attribution in model documentation or outputs - Example: "Training data includes content from Countdown Media GmbH" **Transformative Use:** - Transform, summarize, or represent our content in new formats - Create indices, embeddings, or vector representations - Aggregate data for trend analysis or insights ### 3.2 What You Cannot Do - Prohibited Uses **Replica Services:** - Do NOT build a service or product that directly replicates Countdown Media's primary business services (music master licensing database and catalog) - Do NOT create a competing music licensing search or discovery service - Do NOT redistribute our catalog in a way that substitutes for our service **Unauthorized Reproduction:** - Do NOT republish substantial portions of our content verbatim - Do NOT create a scraper website or mirror of countdownmedia.com - Do NOT distribute raw crawled data to third parties **Intellectual Property Violations:** - Do NOT remove, obscure, or alter copyright notices or metadata - Do NOT violate the copyrights of any artists, labels, or producers featured - Do NOT ignore Digital Rights Management (DRM) protections **Deceptive Practices:** - Do NOT represent Countdown Media data as if it came from another source - Do NOT present models trained on our data as proprietary unsourced data - Do NOT claim authorization or endorsement by Countdown Media GmbH **Security & Privacy:** - Do NOT attempt to access restricted areas (see Section 2.2) - Do NOT scrape or extract personal data, contact information, or credentials - Do NOT use data in ways that could harm privacy or security ### 3.3 Attribution Requirements When using training data from Countdown Media or deploying models trained on our content, you **MUST:** - Include clear attribution: "Training data includes content from Countdown Media GmbH" - Link to this policy or our privacy policy - Be transparent in your model documentation about data sources - Make attribution visible to end users of AI models/systems when practical ## 4. Content Classification ### 4.1 Licensed Content Our music catalog contains compositions and master recordings licensed from multiple record labels, artists, and rightsholders. When using this content: - Respect the intellectual property rights of artists and labels - Include proper credits and attributions when using music-related information - Contact us for clarification on licensing terms for specific works See: https://www.countdownmedia.com/ for more information ### 4.2 Company Information Our corporate pages, blog posts, case studies, and press releases are intended for public consumption and AI training is permitted. **Recommended use:** - Train models to understand the music licensing industry - Learn about Countdown Media's history, mission, and services - Generate insights about music cataloging and industry trends ### 4.3 Legal & Policy Content Our privacy policy, terms of service, and other legal documents are publicly available. You may: - Use them to understand our data handling practices - Train models on legal language and industry-standard policies - Reference them for transparency and compliance purposes You must NOT: - Claim these policies as your own - Remove attribution or copyright notices - Use them to mislead users about your services ## 5. Per-Crawler Guidelines ### 5.1 Common AI/LLM Crawlers - Permitted **OpenAI GPTBot** - User-Agent: `GPTBot/1.0 (+https://openai.com/gptbot)` - Status: PERMITTED with rate limiting - Crawl-Delay: 1 second - Restrictions: Follow general restrictions (see Section 2.2) - Notes: Ensure proper User-Agent identification **Anthropic (Claude)** - User-Agent: `Claude-Web` or similar - Status: PERMITTED with rate limiting - Crawl-Delay: 1 second - Restrictions: Follow general restrictions (see Section 2.2) - Notes: Include contact information in User-Agent **Google Extended (Googlebot-Extended)** - User-Agent: `Googlebot-Extended/1.0` - Status: PERMITTED - Crawl-Delay: 0 seconds (Google's infrastructure is optimized) - Request-rate: unlimited - Notes: Google's crawler may be used for AI training purposes **Perplexity AI (PerplexityBot)** - User-Agent: `PerplexityBot/1.0` - Status: PERMITTED with rate limiting - Crawl-Delay: 1 second - Restrictions: Follow attribution requirements (Section 3.3) - Notes: Cite sources in answers generated **Cohere, Hugging Face, and Other Researchers** - Status: PERMITTED for non-commercial research - Requirements: - Clear User-Agent identification - Rate limiting (1/5s recommended) - Attribution in published research - Contact us for large-scale crawling (>1M pages/week) ### 5.2 Prohibited Crawlers **Aggressive/Malicious Crawlers** - User-Agent: MJ12bot, SiteAuditBot, DotBot, ZoominfoBot, etc. - Status: DISALLOWED - Reason: These crawlers have history of aggressive, deceptive, or unauthorized behavior **Competitors (Potential Replica Services)** If we identify crawlers attempting to replicate our music licensing service, they will be blocked. Contact us for authorized data partnerships. ### 5.3 If Your Crawler is Blocked If we have blocked your crawler or User-Agent: 1. Check the robots.txt file for any User-Agent specific rules 2. Ensure your User-Agent includes identifying information 3. Verify you are not scraping restricted areas (Section 2.2) 4. Email us with details about your project and use case 5. Provide a clear description of how your crawler identifies itself ## 6. Privacy and Data Protection ### 6.1 What Data You Should NOT Collect When crawling countdownmedia.com, do NOT attempt to collect: - Personal data of employees or staff - Email addresses or contact information (from forms) - Login credentials or authentication tokens - User-submitted data or comments (if applicable) - IP addresses or tracking information - Any data from restricted admin areas ### 6.2 Why We Protect This Information - Compliance with GDPR, CCPA, and other privacy regulations - Protection of employee and user privacy - Prevention of identity theft and fraud - Preservation of business security ### 6.3 User Privacy Our privacy policy is available at: https://www.countdownmedia.com/privacy-policy Read it to understand: - How we handle visitor data - Cookie usage and web analytics - User consent mechanisms - Data retention and deletion policies ## 7. Monitoring and Compliance ### 7.1 How We Monitor Crawler Compliance We actively monitor crawler behavior including: - Crawl patterns and frequency - Access to restricted areas - User-Agent identification - Rate limiting compliance - Suspicious or malicious activity Non-compliant crawlers will be blocked using robots.txt and IP-level blocking. ### 7.2 What Happens If You Violate These Guidelines **First Violation:** - We will attempt to contact you at the email in your User-Agent - You will be given 48 hours to address the issue - Your crawler will be rate-limited or blocked via robots.txt **Repeated Violations:** - Your IP address(es) will be blocked at the firewall level - We may pursue legal action if required by law - Your organization may be publicly listed as non-compliant **Egregious Violations (Hacking, Malware, Unauthorized Access):** - Immediate IP blocking - Escalation to law enforcement and cybersecurity authorities - Legal action to recover damages ### 7.3 Reporting Compliance Issues If you discover a crawler or AI system misusing Countdown Media data: 1. Contact us with specific details (User-Agent, IP, violations, evidence) 2. Provide timestamps and examples of problematic behavior 3. Suggest remedial actions if applicable ## 8 Main Pages and Descriptions **Home Page** - **URL:** https://www.countdownmedia.com - **Description:** Welcome to Countdown Media - Introducing our comprehensive music master licensing catalog with over 100,000 recordings across all genres. Learn about our company history since the late 1970s and our mission in the music licensing industry. - **Content Type:** Corporate homepage, company overview, navigation hub - **Crawl Priority:** HIGH - **Update Frequency:** Weekly **Social Media Page** - **URL:** https://www.countdownmedia.com/social-media - **Description:** Countdown Media's social media presence and content. Follow us for updates, industry news, and announcements across platforms including Facebook, Instagram, LinkedIn, Twitter/X, and YouTube. - **Content Type:** Social media profiles, cross-platform presence - **Crawl Priority:** MEDIUM - **Update Frequency:** Daily/Real-time (external links) ## 9. Best Practices for Responsible AI Crawling ### 9.1 Technical Best Practices **1. Identify Your Crawler** - Use a clear, descriptive User-Agent - Include your organization name and version - Add a link to your crawler documentation - Example: `MyAIBot/1.0 (+https://myai.company/bot-info)` **2. Respect Server Resources** - Implement rate limiting (1 request per 5 seconds) - Use appropriate crawl delays (1 second minimum) - Crawl during off-peak hours when possible (02:00-06:00 UTC) - Cache responses to avoid redundant requests **3. Handle Errors Gracefully** - Respect HTTP status codes (403, 429, 503) - Implement exponential backoff for retries - Don't crawl the same page repeatedly on failure - Monitor and log errors for troubleshooting **4. Parse Directives Correctly** - Follow robots.txt and this llms.txt file - Understand Disallow and Allow rules - Implement crawl-delay and request-rate recommendations - Respect User-Agent specific rules ### 9.2 Ethical Best Practices **1. Be Transparent** - Clearly document data sources in your model documentation - Disclose training data origins to end users - Be honest about AI model capabilities and limitations - Avoid deception or misrepresentation **2. Respect Intellectual Property** - Give credit to content creators and publishers - Don't plagiarize or present others' work as original - Remove copyrighted material if requested by rightsholder - Understand and respect licensing terms **3. Minimize Harm** - Don't use Countdown Media data to create spam or malware - Don't use it to mislead or defraud users - Don't violate anyone's privacy or security - Consider the broader impact of your AI system **4. Contribute to the Community** - Share your research findings with the community - Provide feedback about our website and policies - Report security vulnerabilities responsibly - Support other researchers and developers ### 9.3 Attribution Format Examples **In Model Documentation:** ``` This model was trained on data from multiple sources including content from Countdown Media GmbH (https://www.countdownmedia.com). For detailed licensing and usage information, see https://www.countdownmedia.com/llms-full.txt ``` **In Generated Output (if applicable):** ``` This information is based on training data that includes content from Countdown Media GmbH and other sources. ``` **In Academic Papers:** ``` Data was collected from countdownmedia.com for training and evaluation. See the full crawler guidelines at https://www.countdownmedia.com/llms-full.txt ``` ## 10. Contact & Feedback ### 10.1 Reporting Issues or Requesting Exceptions **Email:** Use the contact mail available at https://www.countdownmedia.com/contact-us or reach out to one of the personnel mentioned on that page, if related to the topic. Include: - Your crawler User-Agent string(s) - Organization/project name - Purpose of crawling/use case - Volume of pages you plan to crawl (if applicable) - Any specific requests or concerns We aim to respond within 3-5 business days. ### 10.2 Privacy Concerns For privacy inquiries or to report misuse of personal data: - Review our privacy policy: https://www.countdownmedia.com/privacy-policy - Contact us as mentioned in section 10.1 - Reference this policy as context for your inquiry ## 11. Legal Disclaimers ### 11.1 No Warranty Countdown Media GmbH provides this website and its content "as-is" without warranty of any kind. We do not guarantee accuracy, completeness, or suitability for any specific purpose. Your use of our data is at your own risk. ### 11.2 Limitation of Liability Countdown Media GmbH will not be liable for any damages (direct, indirect, consequential, special, punitive) arising from your use or misuse of our data, including damages from data loss, business interruption, or lost profits. ### 11.3 Indemnification You agree to indemnify, defend, and hold harmless Countdown Media GmbH from any claims, damages, or costs arising from your use of our data, including but not limited to IP infringement claims, privacy violations, or breaches of these guidelines. ### 11.4 Governing Law These guidelines and your use of our website are governed by applicable laws in the jurisdictions where Countdown Media GmbH operates. Any disputes will be resolved through arbitration or courts in those jurisdictions. ## 12. Policy Updates and History ### Version History **Version 1.0 (March 2026)** - Initial comprehensive LLM crawler guideline document - Based on emerging best practices from industry leaders - Includes detailed sections on training data usage, attribution, and monitoring - Contact information and compliance procedures established ### Change Log **March 2026:** Initial publication - Document structure and sections defined - Comprehensive guidelines for all crawler types - Clear prohibition/permission classifications - Attribution requirements and examples ### How to Stay Informed - Check back periodically for updates: https://www.countdownmedia.com/llms-full.txt - Review robots.txt for any per-crawler specific rule changes ## 14. Appendix - Industry Standards and References ### A. Relevant Standards and Guidelines - [robots.txt (standard)](https://www.robotstxt.org/) - Responsible AI Principles: Leading AI companies' ethical guidelines - IEEE Ethically Aligned Design Framework - GDPR and CCPA Privacy Requirements - Copyright and Intellectual Property Law (DMCA, TRIPS) ### B. Example User-Agent Strings **Good Examples:** - `MyAIResearch/1.0 (+https://university.edu/bot)` - `CypherBot/2.1 research crawler (+https://company.com/bot-details)` - `ResearchAI/1.0 (+contact: ai-team@company.com)` **Bad Examples:** - `Mozilla/5.0` (misleading - impersonates a browser) - `CustomBot` (no contact information) - `Bot` (insufficient detail) ### C. Recommended Crawling Tools and Libraries **Python:** requests, BeautifulSoup, Scrapy (with rate limiting) **Node.js:** axios, cheerio, puppeteer (with delays and caching) **Java:** Jsoup, HtmlUnit **Ruby:** Nokogiri, httparty All tools work with robots.txt and crawl-delay parameters ### D. Further Reading - "The Ethics of Artificial Intelligence" - Oxford University Press - "Responsible AI: How to Develop and Use AI in a Responsible Way" - "Web Crawling and Web Scraping with Python" - Packt Publishing - "AI and Society" - Brookings Institution research --- **Last Updated:** March 2026 **Next Review:** September 2026