# How to Generate Knowledge Graphs for Chinese, Japanese, and Korean

> Automatically generate knowledge graphs for Chinese, Japanese, and Korean projects with Understand Anything. Leverage Tree-sitter for seamless Unicode support without special configuration.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: how-to-guide
- Published: 2026-06-02

---

**The Understand Anything tool automatically generates knowledge graphs for Chinese, Japanese, and Korean projects by leveraging Tree-sitter parsers that natively support Unicode identifiers and comments without requiring special configuration.**

The Understand Anything repository (`Lum1104/Understand-Anything`) provides a language-agnostic static analysis engine that constructs detailed knowledge graphs from source code. Because the pipeline relies on Tree-sitter's Unicode-aware grammars rather than ASCII-specific regex, it handles CJK (Chinese, Japanese, Korean) source files identically to English-based code, preserving original characters throughout the extraction process.

## Architecture Overview

### Unicode-First Parsing Strategy

The generation pipeline relies on **Tree-sitter parsers** that understand syntactic structure while preserving the original Unicode text of identifiers, comments, and string literals. Because Tree-sitter grammars support full Unicode, source files written in Chinese, Japanese, or Korean undergo the exact same processing as English-based code.

The language-agnostic approach is implemented across several core components:

- **[`discovery.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/discovery.ts)** – Crawls the repository and applies ignore filters to skip generated files
- **[`language-registry.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/language-registry.ts)** – Maps file extensions to Tree-sitter languages, including a fallback "text" grammar for any UTF-8 file
- **[`tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/tree-sitter-plugin.ts)** – Loads WASM parsers and produces concrete syntax trees while retaining Unicode node texts
- **[`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts)** – Assembles extracted nodes into the final graph structure
- **[`language-lesson.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/language-lesson.ts)** – Generates human-readable explanations that preserve original CJK characters

## The Six-Stage Generation Pipeline

### 1. File Discovery and Filtering

In [`understand-anything-plugin/packages/core/src/plugins/discovery.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/discovery.ts), the analyzer walks the repository tree, applying configurable ignore patterns to exclude build artifacts and dependencies. This stage operates on file paths alone, making it universally compatible with CJK directory names and filenames.

### 2. Language Detection via Extension Mapping

The **language registry** ([`understand-anything-plugin/packages/core/src/languages/language-registry.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/languages/language-registry.ts)) maps file extensions to specific Tree-sitter grammars. For unrecognized extensions, the system falls back to a generic "text" grammar that safely parses any UTF-8 content. This ensures that even custom file types containing Chinese, Japanese, or Korean text are processed without data loss.

### 3. Syntax Tree Generation

The **Tree-sitter plugin** ([`understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts)) loads the appropriate WASM parser for each file type. As it constructs the concrete syntax tree, Unicode identifiers, comments, and string literals are retained verbatim in node texts rather than being sanitized or stripped.

### 4. Symbol Extraction

Language-specific extractors (such as [`java-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/java-extractor.ts) or [`go-extractor.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/go-extractor.ts)) traverse the syntax trees and emit graph nodes representing functions, classes, modules, and variables. Because these extractors store identifier names and comments exactly as they appear in source, Chinese function names, Japanese comments, and Korean variable declarations enter the knowledge graph unchanged.

### 5. Graph Assembly and Relationship Mapping

The **[`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts)** module ([`understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts)) stitches together extracted nodes, resolves import/export relationships, and creates edges representing calls, inheritance, and dependencies. The graph structure stores metadata—including symbol names and documentation—using native Unicode strings.

### 6. Human-Readable Output Generation

Finally, **[`language-lesson.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/language-lesson.ts)** ([`understand-anything-plugin/packages/core/src/analyzer/language-lesson.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/understand-anything-plugin/packages/core/src/analyzer/language-lesson.ts)) processes graph nodes to generate explanations. When rendering these for the dashboard, the system displays the original Chinese, Japanese, or Korean identifiers and comments without transliteration or translation.

## Practical Usage Examples

### Analyzing a Repository via CLI

To generate a knowledge graph for a project containing CJK source files, run the standard build command:

```bash

# Install dependencies

pnpm install

# Build and analyze the project

pnpm --filter @understand-anything/core run build

```

The resulting graph is written to [`.understand-anything/knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/.understand-anything/knowledge-graph.json). You can query for specific CJK identifiers using standard JSON tools:

```bash
cat .understand-anything/knowledge-graph.json | jq '.nodes[] | select(.name | test("日本語"))'

```

### Programmatic API Usage

For custom integrations, import the core analyzer directly:

```typescript
import { analyzeProject } from '@understand-anything/core';

// Analyze a directory containing Chinese and Korean source files
const result = await analyzeProject({
  root: '/path/to/your/project',
  pathAllowList: [], // Empty array analyzes all files
});

// Filter nodes by CJK content
const cjkNodes = result.graph.nodes.filter(
  n => /中文|한국어|日本語/.test(n.name)
);
console.log(cjkNodes);

```

### Extending Support for Additional Languages

If you encounter a language not present in the registry, register a custom Tree-sitter parser:

```typescript
import { registerLanguage } from '@understand-anything/core/languages';
import path from 'path';

// Register a custom WASM parser for a Korean DSL
await registerLanguage({
  id: 'korean-dsl',
  extensions: ['.kdl'],
  wasmPath: path.resolve(__dirname, 'korean.wasm'),
});

```

Once registered, the analyzer automatically includes `.kdl` files in the knowledge graph generation, preserving any Korean text within the source.

## Summary

- **Tree-sitter integration** enables native Unicode support, making CJK characters first-class citizens in the knowledge graph
- **No special configuration** is required; the pipeline automatically handles UTF-8 source files in Chinese, Japanese, and Korean
- **Key components** include [`discovery.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/discovery.ts) for file crawling, [`language-registry.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/language-registry.ts) for parser mapping, and [`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts) for relationship assembly
- **Extensible architecture** allows adding new Tree-sitter grammars for domain-specific languages while maintaining CJK support
- **Preservation of original text** ensures that identifiers and comments remain readable in the generated graph and dashboard output

## Frequently Asked Questions

### Does the tool require special flags to handle Chinese or Japanese source files?

No special flags are necessary. The Understand Anything pipeline treats all source files as UTF-8 by default. As implemented in [`language-registry.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/language-registry.ts), the system includes a fallback "text" grammar that safely parses any Unicode content, ensuring Chinese, Japanese, and Korean identifiers are processed identically to ASCII text.

### Can the knowledge graph handle mixed-language codebases?

Yes. Because the extraction layer in [`tree-sitter-plugin.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/tree-sitter-plugin.ts) and [`graph-builder.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/graph-builder.ts) stores node texts verbatim, a single project can contain English, Chinese, Japanese, and Korean identifiers simultaneously. The graph preserves each symbol's original Unicode representation without requiring language-specific preprocessing.

### How do I add support for a programming language specific to Korean development?

You can extend the language registry by calling `registerLanguage()` from `@understand-anything/core/languages`. Provide the file extension, a unique language ID, and the path to a Tree-sitter WASM parser that supports the language's syntax. The rest of the pipeline—including graph generation and CJK text preservation—works automatically for the newly added language.

### Where are the CJK characters stored in the final knowledge graph?

Unicode text is preserved in the `name` and metadata fields of nodes within [`knowledge-graph.json`](https://github.com/Lum1104/Understand-Anything/blob/main/knowledge-graph.json). The [`language-lesson.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/language-lesson.ts) module specifically ensures that when generating human-readable explanations from these nodes, the original Chinese, Japanese, or Korean characters are displayed unchanged in the output dashboard.