How the Zero-Token Portal Scanner Works with Greenhouse, Ashby, and Lever
The Career-Ops zero-token portal scanner retrieves job listings from Greenhouse, Ashby, and Lever using pure HTTP GET requests to public APIs, consuming no LLM tokens during the initial crawl.
The santifer/career-ops repository implements a cost-efficient job board crawler that interfaces directly with Applicant Tracking System (ATS) providers. By leveraging provider-specific public API endpoints rather than browser automation or AI inference for data extraction, the zero-token portal scanner can process thousands of career pages at minimal operational cost. The architecture relies on a modular provider system that normalizes disparate ATS responses into a canonical job format before any optional LLM-based processing occurs.
Core Architecture and Provider Loading
The orchestration logic lives in scan.mjs, which dynamically discovers and loads provider modules at runtime. The loadProviders() function (lines 61‑86 of scan.mjs) reads every *.mjs file in the providers/ directory, skipping files prefixed with an underscore, and imports them via await import(). Each module exports an id, an optional detect() method, and a required fetch() method that conforms to the zero-token contract.
Provider Detection and Resolution
For each company defined in portals.yml, the scanner invokes resolveProvider() (lines 91‑124 of scan.mjs). The resolution follows a strict precedence:
- Explicit provider field – If the entry specifies
provider: greenhouse, that module is selected immediately. - Local parser – If a local parsing rule exists for the domain.
- Dynamic detection – Each provider’s
detect(entry)method inspects thecareers_url(or an explicitapi:field) and returns an API endpoint object if the host matches.
The first provider to return a truthy value wins the resolution.
ATS Provider Implementations
Each ATS module implements the same interface: detect() returns metadata including the concrete API URL, and fetch() performs the HTTP request and normalizes the response.
Greenhouse
The Greenhouse provider (providers/greenhouse.mjs) extracts the board slug from URLs such as https://jobs.company.com/job-boards.greenhouse.io/company using a regex in detect() (lines 51‑55). It constructs the endpoint https://boards-api.greenhouse.io/v1/boards/<slug>/jobs and returns it to the orchestrator.
The fetch() method (lines 60‑73) calls ctx.fetchJson(apiUrl, {redirect:'error'}) via the shared _http.mjs context. It maps the JSON payload to a canonical job object containing title, url, company, location, and postedAt, filtering out entries lacking an absolute_url.
Ashby
The Ashby provider (providers/ashby.mjs) detects URLs containing jobs.ashbyhq.com or explicit api: fields. It validates the host against an allow-list before returning the public Ashby API endpoint. The fetch() implementation retrieves the JSON posting list and normalizes it to the same canonical shape as Greenhouse, ensuring downstream filters receive consistent data structures regardless of the source ATS.
Lever
The Lever provider (providers/lever.mjs) recognizes career pages hosted at https://jobs.lever.co/company. Its detect() method builds the endpoint https://api.lever.co/v0/postings/<company>. The fetch() function hits this endpoint, parses the JSON array of postings, and maps each entry to the standard job object format, completing the zero-token retrieval cycle.
Data Normalization and Filtering
After fetching, raw job lists pass through a series of pure-JavaScript filters defined in scan.mjs (lines 26‑98). These include buildTitleFilter(), buildLocationFilter(), buildSalaryFilter(), buildContentFilter(), and buildCooldownFilter(), along with URL deduplication. All filters execute against the canonical job objects before any LLM-based scoring, guaranteeing the scanner remains token-free during the extraction phase.
Optional Liveness Verification
The zero-token path excludes browser automation to minimize costs. However, users can invoke node scan.mjs --verify to enable the liveness check. This triggers verifyOffers() (lines 443‑523 of scan.mjs), which launches Playwright (Chromium) to confirm that job URLs still host live "Apply" buttons. This step occurs after the zero-token fetch and filter pipeline, ensuring Playwright costs are incurred only when explicitly requested for final validation.
Running the Scanner
Execute a standard zero-token scan that writes new offers to data/pipeline.md:
node scan.mjs
Preview changes without writing files:
node scan.mjs --dry-run
Run with optional liveness verification (adds Playwright overhead):
node scan.mjs --verify
Rescue moved URLs (e.g., Greenhouse 404s) during verification:
node scan.mjs --verify --rediscover-404
Summary
- Zero-token architecture: The scanner uses simple HTTP GET requests to public ATS APIs (Greenhouse, Ashby, Lever) rather than LLM inference or browser rendering for data extraction.
- Modular providers:
scan.mjsloads provider modules dynamically fromproviders/, with each module exposingdetect()andfetch()methods. - Canonical normalization: All ATS responses are mapped to a uniform job object format before filtering.
- Pure-JS filtering: Title, location, salary, content, and cooldown filters run as fast JavaScript predicates without consuming tokens.
- Optional verification: Playwright-based liveness checks occur only when the
--verifyflag is passed, keeping the default path completely token-free.
Frequently Asked Questions
What makes the scanner "zero-token"?
The scanner retrieves job listings using direct HTTP requests to public ATS endpoints (e.g., boards-api.greenhouse.io, api.lever.co). Because it parses JSON responses with native JavaScript rather than calling Claude or other LLMs to extract structured data, it consumes zero AI tokens during the crawl phase.
How does the scanner detect which ATS a company uses?
The resolveProvider() function (lines 91‑124 of scan.mjs) checks three sources in order: an explicit provider field in the configuration, a local parser rule, or the detect() method exported by each provider module. Greenhouse, Ashby, and Lever each implement detect() to recognize their respective URL patterns and return the appropriate API endpoint.
Can I add support for a new ATS provider?
Yes. Create a new file in providers/ (e.g., custom-ats.mjs) that exports an id, an optional detect(entry) function that returns an API URL if the entry matches, and a fetch(entry, ctx) function that returns an array of canonical job objects. The orchestrator will automatically load the module on the next run of scan.mjs.
When does the scanner use Playwright?
Playwright is invoked only when the --verify flag is passed. The verifyOffers() function (lines 443‑523) then launches a headless browser to check that job URLs are still active. This occurs after the zero-token fetch and filter pipeline, ensuring that token consumption (and browser overhead) happens only during the optional verification step.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →