# Does Cangjie-Skill Bundle Pre-Trained Models? Repository Architecture Explained

> Does Cangjie-Skill bundle pre-trained models? Discover the repository architecture and understand how it leverages external LLM services for all extraction and validation, not bundled models.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: architecture
- Published: 2026-07-17

---

**The cangjie-skill repository does not contain or bundle any pre-trained machine learning models, relying instead on external LLM services configured in [`opencode.json`](https://github.com/kangarooking/cangjie-skill/blob/main/opencode.json) to perform all extraction and validation tasks.**

The `kangarooking/cangjie-skill` repository is designed as a lightweight skill framework that delegates computational heavy lifting to hosted large language models. Unlike traditional ML projects that ship with `.pt` or `.bin` weight files, this codebase contains only configuration files and prompt templates. Understanding this architecture is crucial for developers wondering whether they need to download model weights or configure GPU resources before running the skill.

## Model Configuration in [`opencode.json`](https://github.com/kangarooking/cangjie-skill/blob/main/opencode.json)

The only reference to pre-trained models within the repository appears in the [`opencode.json`](https://github.com/kangarooking/cangjie-skill/blob/main/opencode.json) configuration file. This JSON file declares which LLM the Instagit runtime should invoke when executing the skill, but does not contain any model weights or inference code.

```json
{
  "model": "openrouter/openai/gpt-oss-120b",
  "small_model": "openrouter/qwen/qwen3-32b"
}

```

These strings specify external API endpoints rather than local file paths. The runtime uses these identifiers to route requests to the appropriate hosted models, meaning the skill itself remains model-agnostic and requires no local GPU resources.

## Absence of Model Weight Files

A comprehensive scan of the repository reveals no embedded pre-trained model artifacts. Common ML file extensions—including `.pt`, `.bin`, `.safetensors`, and `.onnx`—are entirely absent from the codebase.

Instead of neural network weights, the repository contains:

- **Markdown prompt templates** in `extractors/*.md` that define extraction logic
- **Methodology documentation** in [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md) describing pipeline stages
- **Template files** in `templates/*.md.template` used for assembling final outputs

This file structure confirms that all "intelligence" is provided by the external LLM service rather than locally loaded models.

## Runtime Execution Without Local Models

When the skill runs, it does not load a model into memory. Instead, it forwards structured requests to the LLM specified in the configuration. The `extractors/*.md` files contain plain-text prompts that the host LLM executes at runtime, performing tasks like parallel extraction and validation tests.

The following examples demonstrate that no model loading occurs during execution:

```python

# Invoking the skill from a Claude Code session

# The skill forwards requests to the configured LLM; no local weights loaded

await claude.run_skill(
    name="cangjie-skill",
    inputs={"path": "books/my-book.pdf"}
)

```

```bash

# Running the pipeline locally still relies on external API calls

python -m cangjie_skill.run --source books/my-book.pdf

```

Both invocation methods treat the skill as a **prompt orchestration layer** rather than a model inference engine. The actual text generation, analysis, and extraction work happens remotely on the OpenRouter or Claude Code infrastructure.

## Skill Architecture: Configuration vs. Embedding

The `kangarooking/cangjie-skill` follows a **thin-client architecture** where the repository contains only:

- **Business logic** in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) describing intended use cases
- **Prompt engineering** in markdown files that guide the external LLM
- **Configuration metadata** in [`opencode.json`](https://github.com/kangarooking/cangjie-skill/blob/main/opencode.json) specifying which model API to call

This design choice keeps the repository lightweight (no multi-gigabyte weight files) while allowing flexibility to switch between different LLM providers by editing a single configuration file. Developers do not need to manage CUDA dependencies, model quantization, or VRAM allocation when deploying this skill.

## Summary

- **The cangjie-skill repository contains no pre-trained model weights**—no `.pt`, `.bin`, or `.safetensors` files are present.
- **All ML processing is delegated** to external LLMs specified in [`opencode.json`](https://github.com/kangarooking/cangjie-skill/blob/main/opencode.json), specifically `openrouter/openai/gpt-oss-120b` and `openrouter/qwen/qwen3-32b`.
- **Extraction logic lives in markdown prompts** (`extractors/*.md`) executed by the host LLM at runtime, not by local inference code.
- **Deployment requires no GPU resources** or model downloads, only API access to the configured LLM services.

## Frequently Asked Questions

### Does cangjie-skill require downloading any model files before first use?

No. The repository does not distribute model weights. You only need to ensure your runtime environment has API access to the LLM endpoints specified in [`opencode.json`](https://github.com/kangarooking/cangjie-skill/blob/main/opencode.json). The skill operates as a configuration layer that sends prompts to these hosted services.

### Can I run cangjie-skill offline with local models?

Not without modification. The current implementation in [`opencode.json`](https://github.com/kangarooking/cangjie-skill/blob/main/opencode.json) points to cloud-based LLM APIs (OpenRouter). While you could theoretically fork the repository and modify the configuration to point to a local model server, the vanilla `kangarooking/cangjie-skill` expects external API connectivity.

### What happens if the configured LLM service is unavailable?

The skill will fail at runtime with API connection errors. Since [`opencode.json`](https://github.com/kangarooking/cangjie-skill/blob/main/opencode.json) defines `openrouter/openai/gpt-oss-120b` as the primary model and `openrouter/qwen/qwen3-32b` as the small model, any downtime or rate limiting at these endpoints will prevent the extractors in `extractors/*.md` from executing their prompts. There is no local fallback model embedded in the repository.

### Why does the repository use markdown files instead of Python code for extraction?

The `extractors/*.md` files contain **prompt templates** that guide the external LLM's behavior. This approach leverages the hosting platform's Claude Code or similar environment to perform the actual natural language processing. By using markdown rather than PyTorch or TensorFlow code, the skill remains agnostic to specific model implementations and can adapt to improvements in the hosted LLM without code changes.