Does Cangjie-Skill Bundle Pre-Trained Models? Repository Architecture Explained

The cangjie-skill repository does not contain or bundle any pre-trained machine learning models, relying instead on external LLM services configured in opencode.json to perform all extraction and validation tasks.

The kangarooking/cangjie-skill repository is designed as a lightweight skill framework that delegates computational heavy lifting to hosted large language models. Unlike traditional ML projects that ship with .pt or .bin weight files, this codebase contains only configuration files and prompt templates. Understanding this architecture is crucial for developers wondering whether they need to download model weights or configure GPU resources before running the skill.

Model Configuration in opencode.json

The only reference to pre-trained models within the repository appears in the opencode.json configuration file. This JSON file declares which LLM the Instagit runtime should invoke when executing the skill, but does not contain any model weights or inference code.

{
  "model": "openrouter/openai/gpt-oss-120b",
  "small_model": "openrouter/qwen/qwen3-32b"
}

These strings specify external API endpoints rather than local file paths. The runtime uses these identifiers to route requests to the appropriate hosted models, meaning the skill itself remains model-agnostic and requires no local GPU resources.

Absence of Model Weight Files

A comprehensive scan of the repository reveals no embedded pre-trained model artifacts. Common ML file extensions—including .pt, .bin, .safetensors, and .onnx—are entirely absent from the codebase.

Instead of neural network weights, the repository contains:

  • Markdown prompt templates in extractors/*.md that define extraction logic
  • Methodology documentation in methodology/00-overview.md describing pipeline stages
  • Template files in templates/*.md.template used for assembling final outputs

This file structure confirms that all "intelligence" is provided by the external LLM service rather than locally loaded models.

Runtime Execution Without Local Models

When the skill runs, it does not load a model into memory. Instead, it forwards structured requests to the LLM specified in the configuration. The extractors/*.md files contain plain-text prompts that the host LLM executes at runtime, performing tasks like parallel extraction and validation tests.

The following examples demonstrate that no model loading occurs during execution:


# Invoking the skill from a Claude Code session

# The skill forwards requests to the configured LLM; no local weights loaded

await claude.run_skill(
    name="cangjie-skill",
    inputs={"path": "books/my-book.pdf"}
)

# Running the pipeline locally still relies on external API calls

python -m cangjie_skill.run --source books/my-book.pdf

Both invocation methods treat the skill as a prompt orchestration layer rather than a model inference engine. The actual text generation, analysis, and extraction work happens remotely on the OpenRouter or Claude Code infrastructure.

Skill Architecture: Configuration vs. Embedding

The kangarooking/cangjie-skill follows a thin-client architecture where the repository contains only:

  • Business logic in SKILL.md describing intended use cases
  • Prompt engineering in markdown files that guide the external LLM
  • Configuration metadata in opencode.json specifying which model API to call

This design choice keeps the repository lightweight (no multi-gigabyte weight files) while allowing flexibility to switch between different LLM providers by editing a single configuration file. Developers do not need to manage CUDA dependencies, model quantization, or VRAM allocation when deploying this skill.

Summary

  • The cangjie-skill repository contains no pre-trained model weights—no .pt, .bin, or .safetensors files are present.
  • All ML processing is delegated to external LLMs specified in opencode.json, specifically openrouter/openai/gpt-oss-120b and openrouter/qwen/qwen3-32b.
  • Extraction logic lives in markdown prompts (extractors/*.md) executed by the host LLM at runtime, not by local inference code.
  • Deployment requires no GPU resources or model downloads, only API access to the configured LLM services.

Frequently Asked Questions

Does cangjie-skill require downloading any model files before first use?

No. The repository does not distribute model weights. You only need to ensure your runtime environment has API access to the LLM endpoints specified in opencode.json. The skill operates as a configuration layer that sends prompts to these hosted services.

Can I run cangjie-skill offline with local models?

Not without modification. The current implementation in opencode.json points to cloud-based LLM APIs (OpenRouter). While you could theoretically fork the repository and modify the configuration to point to a local model server, the vanilla kangarooking/cangjie-skill expects external API connectivity.

What happens if the configured LLM service is unavailable?

The skill will fail at runtime with API connection errors. Since opencode.json defines openrouter/openai/gpt-oss-120b as the primary model and openrouter/qwen/qwen3-32b as the small model, any downtime or rate limiting at these endpoints will prevent the extractors in extractors/*.md from executing their prompts. There is no local fallback model embedded in the repository.

Why does the repository use markdown files instead of Python code for extraction?

The extractors/*.md files contain prompt templates that guide the external LLM's behavior. This approach leverages the hosting platform's Claude Code or similar environment to perform the actual natural language processing. By using markdown rather than PyTorch or TensorFlow code, the skill remains agnostic to specific model implementations and can adapt to improvements in the hosted LLM without code changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →