How the Ego-Lite Learning Subsystem Discovers and Validates Site Skills
The learning subsystem in citrolabs/ego-lite discovers site skills by scanning the agent workspace for manifest files matching URL domains, then validates pack structure and tool schemas before exposing them to browser automation helpers.
The learning subsystem resides in package/ego-browser/src/learning and serves as the intelligence layer that maps URLs to specialized automation capabilities. It performs a two-phase pipeline: first discovering which site-skill packs apply to a given domain, then rigorously validating those packs to prevent runtime errors. This article examines the complete discovery and validation flow using the actual implementation from the citrolabs/ego-lite repository.
Discovery Phase: Matching URLs to Site-Skill Packs
The discovery process begins when the system needs to find applicable automation skills for a specific URL. The implementation in package/ego-browser/src/learning/check-domain-learning.ts orchestrates this through several sequential steps.
Locating the Learning Root and Enumerating Candidates
The learningsRoot() function establishes the base directory at <agent-workspace>/learnings, which serves as the central repository for all site-skill packs. The iterLearningDirs(root) function then reads this directory, filters out hidden folders (those starting with dots), and returns a sorted list of sub-folders where each folder represents a distinct site-skill pack.
// From check-domain-learning.ts
const root = learningsRoot(); // Returns <agent-workspace>/learnings
const candidates = iterLearningDirs(root); // Sorted list of site directories
Domain Matching and Manifest Loading
For each candidate directory, loadLearningManifest(siteDir) parses the manifest.json file. If parsing fails, the directory is skipped silently. The siteSkillsForUrl(url) function extracts the hostname from the provided URL and compares it against each manifest's domains array using domainMatches(hostname, pattern), which supports wildcard patterns such as *.example.com.
When a match occurs, learningEntry(siteDir, manifest) constructs a LearningEntry object containing the pack ID, name, filesystem path, applicable domains, note file paths, and tool schemas for both Node-side and browser-side execution.
// Example from the public API
import { site } from "ego-browser";
const ctx = await site.learnContext("https://www.yoursite.com/dashboard");
// Returns LearningEntry with tools, notes, and metadata
Validation Phase: Ensuring Pack Integrity
Once discovered, site-skill packs undergo strict validation via package/ego-browser/src/learning/validate-learning-format.ts. This prevents malformed packs from causing runtime failures during automation tasks.
Manifest Structure Verification
The validateLearning(siteDir) function performs the initial schema checks. It verifies that the id field matches the directory name, ensures name is a non-empty string, and confirms domains is a non-empty array of syntactically valid domains using isValidDomain(). These checks occur in lines 28-44 of the validation module.
// Validation rules enforced:
// - id === directoryName
// - name.length > 0
// - domains.every(isValidDomain)
// - domains.length > 0
Tool Schema and File Validation
For Node-side tools listed in nodeTools, the validator checks that each tool's path points to tools/*.js, confirms file existence via pathExists, and verifies the module exports the callable name defined in schema.callable using validateToolSchema. It also scans for temporary snapshot references (patterns like @123 or ref=123) via rejectTemporaryRefs and rejects any files containing them.
For Browser-side tools in browserTools, similar checks apply except the path must be browser-tools/*.js. While these tools execute in the browser context and aren't dynamically imported during validation, the system still verifies file existence and rejects temporary references.
// Node tool validation (lines 54-73):
// 1. validateToolSchema(schema)
// 2. path matches tools/*.js pattern
// 3. File exists at path
// 4. Module exports schema.callable function
// 5. No temporary refs in file content
// Browser tool validation (lines 75-82):
// Similar checks for browser-tools/*.js paths
The validateLearnings(root) function orchestrates this process by iterating over every learning directory and aggregating all errors into a single report. This powers the CLI command npm run validate:site-skills, which outputs validation failures for immediate developer feedback.
Integration: Using Discovered and Validated Skills
The public API exposed in package/ego-browser/src/helpers.ts combines both phases seamlessly. The site.learnContext(url) method calls siteSkillsForUrl for discovery and returns a LearnedContext object containing validated tool signatures ready for execution.
// Accessing discovered tools
const context = await site.learnContext("https://news.ycombinator.com");
console.log(context.tools.map(t => t.toolName)); // ["fetchPosts", "comment"]
// Executing Node-side validated tools
const result = await site.runTool(
"hackernews",
"fetchPosts",
{ limit: 10 }
);
// Executing Browser-side validated tools
await site.runBrowserTool("hackernews", "highlightPost");
Summary
- Discovery starts at
<agent-workspace>/learningsand filters directories by matching URL hostnames against manifest domain patterns usingdomainMatches(). - Validation enforces strict schema compliance in
validate-learning-format.ts, checking manifest integrity, file existence, tool callability, and prohibiting temporary snapshot references. - Integration occurs through
site.learnContext(), which returns validatedLearningEntryobjects containing ready-to-use tool schemas. - CLI validation via
npm run validate:site-skillsrunsvalidateLearnings()to catch errors before runtime deployment. - The subsystem distinguishes between Node-side tools (validated for module exports) and Browser-side tools (validated for presence only).
Frequently Asked Questions
How does the learning subsystem handle wildcard domains?
The domainMatches(hostname, pattern) function in check-domain-learning.ts supports wildcard syntax such as *.example.com. When iterating through a manifest's domains array, it matches the extracted URL hostname against these patterns, allowing a single site-skill pack to apply to all subdomains of a target site.
What happens if a site-skill manifest contains temporary snapshot references?
During validation, rejectTemporaryRefs() scans note files and tool source code for patterns like @123 or ref=123. If found, the validation fails with an error indicating the file contains temporary references. This prevents unstable snapshot IDs from reaching production automation scripts.
Can I validate site-skill packs without running the full browser automation?
Yes. Run the CLI command npm run validate:site-skills to execute validateLearnings() directly. This performs static analysis of all manifests, tool schemas, and file references without launching browser instances or executing tool code, making it suitable for CI/CD pipelines.
What is the difference between Node-side and Browser-side tool validation?
Node-side tools undergo full validation: the system checks that the JavaScript file exists, dynamically verifies the module exports the callable function defined in the schema, and validates the schema shape using validateToolSchema. Browser-side tools only validate file existence at browser-tools/*.js paths and check for temporary references, since they execute in the browser context and cannot be statically imported during validation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →