How to Generate Knowledge Graphs for Chinese, Japanese, and Korean

The Understand Anything tool automatically generates knowledge graphs for Chinese, Japanese, and Korean projects by leveraging Tree-sitter parsers that natively support Unicode identifiers and comments without requiring special configuration.

The Understand Anything repository (Lum1104/Understand-Anything) provides a language-agnostic static analysis engine that constructs detailed knowledge graphs from source code. Because the pipeline relies on Tree-sitter's Unicode-aware grammars rather than ASCII-specific regex, it handles CJK (Chinese, Japanese, Korean) source files identically to English-based code, preserving original characters throughout the extraction process.

Architecture Overview

Unicode-First Parsing Strategy

The generation pipeline relies on Tree-sitter parsers that understand syntactic structure while preserving the original Unicode text of identifiers, comments, and string literals. Because Tree-sitter grammars support full Unicode, source files written in Chinese, Japanese, or Korean undergo the exact same processing as English-based code.

The language-agnostic approach is implemented across several core components:

  • discovery.ts – Crawls the repository and applies ignore filters to skip generated files
  • language-registry.ts – Maps file extensions to Tree-sitter languages, including a fallback "text" grammar for any UTF-8 file
  • tree-sitter-plugin.ts – Loads WASM parsers and produces concrete syntax trees while retaining Unicode node texts
  • graph-builder.ts – Assembles extracted nodes into the final graph structure
  • language-lesson.ts – Generates human-readable explanations that preserve original CJK characters

The Six-Stage Generation Pipeline

1. File Discovery and Filtering

In understand-anything-plugin/packages/core/src/plugins/discovery.ts, the analyzer walks the repository tree, applying configurable ignore patterns to exclude build artifacts and dependencies. This stage operates on file paths alone, making it universally compatible with CJK directory names and filenames.

2. Language Detection via Extension Mapping

The language registry (understand-anything-plugin/packages/core/src/languages/language-registry.ts) maps file extensions to specific Tree-sitter grammars. For unrecognized extensions, the system falls back to a generic "text" grammar that safely parses any UTF-8 content. This ensures that even custom file types containing Chinese, Japanese, or Korean text are processed without data loss.

3. Syntax Tree Generation

The Tree-sitter plugin (understand-anything-plugin/packages/core/src/plugins/tree-sitter-plugin.ts) loads the appropriate WASM parser for each file type. As it constructs the concrete syntax tree, Unicode identifiers, comments, and string literals are retained verbatim in node texts rather than being sanitized or stripped.

4. Symbol Extraction

Language-specific extractors (such as java-extractor.ts or go-extractor.ts) traverse the syntax trees and emit graph nodes representing functions, classes, modules, and variables. Because these extractors store identifier names and comments exactly as they appear in source, Chinese function names, Japanese comments, and Korean variable declarations enter the knowledge graph unchanged.

5. Graph Assembly and Relationship Mapping

The graph-builder.ts module (understand-anything-plugin/packages/core/src/analyzer/graph-builder.ts) stitches together extracted nodes, resolves import/export relationships, and creates edges representing calls, inheritance, and dependencies. The graph structure stores metadata—including symbol names and documentation—using native Unicode strings.

6. Human-Readable Output Generation

Finally, language-lesson.ts (understand-anything-plugin/packages/core/src/analyzer/language-lesson.ts) processes graph nodes to generate explanations. When rendering these for the dashboard, the system displays the original Chinese, Japanese, or Korean identifiers and comments without transliteration or translation.

Practical Usage Examples

Analyzing a Repository via CLI

To generate a knowledge graph for a project containing CJK source files, run the standard build command:


# Install dependencies

pnpm install

# Build and analyze the project

pnpm --filter @understand-anything/core run build

The resulting graph is written to .understand-anything/knowledge-graph.json. You can query for specific CJK identifiers using standard JSON tools:

cat .understand-anything/knowledge-graph.json | jq '.nodes[] | select(.name | test("日本語"))'

Programmatic API Usage

For custom integrations, import the core analyzer directly:

import { analyzeProject } from '@understand-anything/core';

// Analyze a directory containing Chinese and Korean source files
const result = await analyzeProject({
  root: '/path/to/your/project',
  pathAllowList: [], // Empty array analyzes all files
});

// Filter nodes by CJK content
const cjkNodes = result.graph.nodes.filter(
  n => /中文|한국어|日本語/.test(n.name)
);
console.log(cjkNodes);

Extending Support for Additional Languages

If you encounter a language not present in the registry, register a custom Tree-sitter parser:

import { registerLanguage } from '@understand-anything/core/languages';
import path from 'path';

// Register a custom WASM parser for a Korean DSL
await registerLanguage({
  id: 'korean-dsl',
  extensions: ['.kdl'],
  wasmPath: path.resolve(__dirname, 'korean.wasm'),
});

Once registered, the analyzer automatically includes .kdl files in the knowledge graph generation, preserving any Korean text within the source.

Summary

  • Tree-sitter integration enables native Unicode support, making CJK characters first-class citizens in the knowledge graph
  • No special configuration is required; the pipeline automatically handles UTF-8 source files in Chinese, Japanese, and Korean
  • Key components include discovery.ts for file crawling, language-registry.ts for parser mapping, and graph-builder.ts for relationship assembly
  • Extensible architecture allows adding new Tree-sitter grammars for domain-specific languages while maintaining CJK support
  • Preservation of original text ensures that identifiers and comments remain readable in the generated graph and dashboard output

Frequently Asked Questions

Does the tool require special flags to handle Chinese or Japanese source files?

No special flags are necessary. The Understand Anything pipeline treats all source files as UTF-8 by default. As implemented in language-registry.ts, the system includes a fallback "text" grammar that safely parses any Unicode content, ensuring Chinese, Japanese, and Korean identifiers are processed identically to ASCII text.

Can the knowledge graph handle mixed-language codebases?

Yes. Because the extraction layer in tree-sitter-plugin.ts and graph-builder.ts stores node texts verbatim, a single project can contain English, Chinese, Japanese, and Korean identifiers simultaneously. The graph preserves each symbol's original Unicode representation without requiring language-specific preprocessing.

How do I add support for a programming language specific to Korean development?

You can extend the language registry by calling registerLanguage() from @understand-anything/core/languages. Provide the file extension, a unique language ID, and the path to a Tree-sitter WASM parser that supports the language's syntax. The rest of the pipeline—including graph generation and CJK text preservation—works automatically for the newly added language.

Where are the CJK characters stored in the final knowledge graph?

Unicode text is preserved in the name and metadata fields of nodes within knowledge-graph.json. The language-lesson.ts module specifically ensures that when generating human-readable explanations from these nodes, the original Chinese, Japanese, or Korean characters are displayed unchanged in the output dashboard.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →