How `updateRobots()` Handles Existing `robots.txt` Files in PHP Sitemap Generator

The updateRobots() method preserves all custom directives in an existing robots.txt file while stripping outdated sitemap references and injecting current sitemap URLs, or creates a fresh file with default permissions if none exists.

When working with the msbatal/php-sitemap-generator library, the updateRobots() method in SunSitemap.php provides automated synchronization between your generated sitemaps and your robots.txt file. This ensures search engine crawlers always have access to your latest sitemap locations without requiring manual file editing.

Pre-Condition Check: Sitemap Generation Required

Before attempting to modify robots.txt, the method validates that a sitemap has already been generated. In SunSitemap.php at lines 75-78, the code checks for the presence of sitemap data:

if (!isset($this->sitemaps)) {
    throw new Exception('To create/update the robots.txt file, first call the "createSitemap" method.');
}

If you call updateRobots() before createSitemap(), the library throws an exception to prevent creating invalid robots.txt entries.

Detecting Existing Files vs. Creating New Ones

The method handles both scenarios—updating an existing robots.txt or creating one from scratch.

When an existing file is found at $this->absPath . $this->robotsFile, the method loads and parses its content into an array of lines. If no file exists, it initializes a minimal default configuration (lines 79-93):

$robotsContent = "User-agent: *\nDisallow:\n";

This default allows all crawlers access to all parts of your site, providing a safe baseline when starting fresh.

Stripping Old Sitemap References

To prevent duplication and stale entries, updateRobots() iterates through every line of an existing robots.txt file and removes any line beginning with Sitemap: (lines 80-90). The implementation uses a simple string check:

// Pseudocode representation of the logic
foreach ($robotsFile as $key => $line) {
    if (strpos(trim($line), 'Sitemap:') === 0) {
        unset($robotsFile[$key]);
    }
}

All remaining non-empty lines—such as User-agent, Disallow, Allow, or Crawl-delay directives—are preserved and concatenated back into the working content buffer.

Appending Current Sitemap URLs

After cleaning old references, the method appends updated sitemap locations based on your generation configuration (lines 94-101). The behavior varies depending on output mode:

  • Single sitemap: Appends one line pointing to the main sitemap file
  • Sitemap index: Appends one line pointing to the index file when multiple sitemaps are generated
  • Gzip compression: When $this->createZip is enabled and a single sitemap exists, appends an additional line for the .gz version
// Single sitemap
$robotsContent .= "\nSitemap: " . $this->baseUrl . $this->relPath . $this->sitemapFile;

// Sitemap index
$robotsContent .= "\nSitemap: " . $this->baseUrl . $this->relPath . $this->sitemapIndexFile;

// Gzip variant (single sitemap only)
$robotsContent .= "\nSitemap: " . $this->baseUrl . $this->relPath . $this->sitemapFile . '.gz';

Atomic Write Operation

The assembled content—combining preserved directives with new sitemap references—is written to disk using file_put_contents() (lines 102-104):

file_put_contents($this->absPath . $this->robotsFile, $robotsContent);

This single-call write operation ensures the file update happens atomically, minimizing the risk of corruption during concurrent access.

Fluent Interface Support

Like other methods in the class, updateRobots() returns the current object instance ($this), enabling method chaining. This pattern appears in the repository's test/index.php example:

$sitemap = new SunSitemap('https://example.com/');
$sitemap->createSitemap()->updateRobots();

You can also combine URL addition, sitemap creation, and robots.txt updates in a single fluent chain:

require 'SunSitemap.php';
$site = new SunSitemap('https://example.com/');
$site->addUrl('https://example.com/about')
     ->addUrl('https://example.com/contact')
     ->createSitemap()
     ->updateRobots();

Summary

  • updateRobots() requires a prior createSitemap() call and throws an exception if the sitemap data is missing.
  • Existing robots.txt files are parsed and preserved—all directives except Sitemap: entries remain intact.
  • Stale sitemap references are automatically removed to prevent duplication and 404 errors for search engines.
  • Current sitemap URLs are injected based on generation mode (single, indexed, or gzipped).
  • Missing files trigger creation with a permissive default rule (User-agent: * and Disallow: empty).
  • Method chaining is supported via the fluent interface pattern used throughout the library.

Frequently Asked Questions

Does updateRobots() delete my existing crawl rules in robots.txt?

No. According to the source code in SunSitemap.php lines 80-90, the method specifically filters only lines beginning with Sitemap: while preserving all other directives including User-agent, Disallow, Allow, and Crawl-delay rules. Your custom crawl policies remain untouched during the update process.

What happens if I call updateRobots() before generating a sitemap?

The method throws an Exception with the message "To create/update the robots.txt file, first call the 'createSitemap' method." This check at lines 75-78 prevents the creation of robots.txt files that reference non-existent sitemap resources.

How does the method handle multiple sitemap files?

When the generator produces a sitemap index (indicated by count($this->sitemapIndex) > 0), updateRobots() appends a single Sitemap: directive pointing to the index file rather than listing individual sitemap chunks. This follows the standard protocol where search engines follow the index to discover all subsidiary sitemap files.

Can I use this method if I don't have an existing robots.txt file?

Yes. If the file does not exist at the target path, the method initializes $robotsContent with a default permissive rule (User-agent: *\nDisallow:\n) and then appends your sitemap references. This ensures search engines can crawl your entire site while immediately discovering your sitemap structure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →