Semantic Analysis and Type Checking in kcl-sema: Inside the KCL Compiler Frontend
The kcl-sema module performs a deterministic, multi-stage semantic analysis pipeline that resolves names, infers types via Hindley-Milner unification, validates schema constraints, and exports structured semantic information for the evaluator and language server.
The kcl-sema crate serves as the semantic heart of the KCL compiler. After the parser produces an Abstract Syntax Tree (AST), this module performs semantic analysis and type checking to ensure programs are well-formed before execution. The implementation spans eight distinct phases, from pre-processing raw identifiers to exporting final symbol tables for downstream tooling.
The Eight Stages of Semantic Analysis
The semantic pipeline in crates/sema processes the AST through specialized subsystems, each handling a specific aspect of program validation.
Pre-processing and Normalization
Before name resolution begins, the pre-processor normalizes the raw AST. In src/pre_process/mod.rs, the system expands qualified identifiers and injects default type values for literals via src/pre_process/lit_ty_default_value.rs. This stage ensures that subsequent passes encounter a consistent, canonical representation of the source code.
Name Resolution and Scope Building
The resolver, implemented primarily in src/resolver/mod.rs and src/resolver/node.rs, performs a depth-first traversal of the pre-processed AST. It creates hierarchical scopes defined in src/core/scope.rs and binds every identifier to its corresponding declaration. The system tracks imports via src/resolver/import.rs and variable bindings via src/resolver/var.rs, storing resolved symbols in src/core/symbol.rs.
Type Construction and Parsing
When the resolver encounters type annotations or schema definitions, src/ty/constructor.rs builds concrete Type objects. This subsystem, supported by src/ty/parser.rs and src/ty/constants.rs, transforms syntax nodes into internal type representations for schemas, functions, and primitives.
Type Erasure and Substitution
For generic programming constructs, src/resolver/ty_erasure.rs strips unnecessary generic parameters while src/ty/context.rs manages type variable substitution. This stage prepares polymorphic types for concrete unification by replacing abstract type variables with inferred concrete instances.
Unification and Type Inference
The type checker implements Hindley-Milner-style unification in src/ty/unify.rs. When expressions lack explicit annotations, the ty::Unify::unify method merges type constraints, producing a substitution map that resolves unknown types to their most general forms. The src/ty/walker.rs module assists in traversing complex type structures during this process.
Constraint Solving and Validation
After type inference stabilizes, the linting subsystem in src/lint/mod.rs and src/lint/types.rs validates schema rules, unique-key constraints, and mutability requirements. The lint::LintPass implementations iterate over declarations, collecting mismatches and violations as Diagnostic objects for the language server protocol (LSP).
Built-in Types and Standard Library
The module provides intrinsic types—including int, str, list, and dict—through src/builtin/mod.rs. Standard library modules such as option, string, and system are defined in src/builtin/string.rs, src/option.rs, and src/builtin/system_module.rs, respectively, making these APIs available during resolution.
Semantic Information Export
Upon successful validation, src/core/semantic_information.rs aggregates symbol tables, type tables, and diagnostic messages into a SemanticInformation structure. This artifact, accessible through core::global_state::GlobalState, feeds the evaluator, code generator, and IDE tooling.
The Resolution Pipeline in Practice
The semantic analysis process follows a strict operational sequence within the resolver::Resolver struct:
-
AST Traversal – The resolver initiates a depth-first walk, instantiating a new
Scopefor each block, module, and schema definition. -
Type Building – During traversal, type annotations trigger
ty::Constructorto instantiateTypevariants such asType::SchemaorType::Function. -
Unification – Expressions with unresolved types invoke
ty::Unify::unify, which reconciles expected and actual types through constraint solving. -
Constraint Checking – Post-inference, the system runs
lint::LintPassvalidators to enforce schema rules and flag invalid attribute usage. -
Result Commit – Final semantic data persists in
GlobalState, enabling queries from the LSP and code generation phases.
Using the kcl-sema API
The crate exposes clean interfaces for both compiler drivers and language server implementations.
Running Semantic Analysis from the Driver
The following pattern demonstrates how kcl-driver invokes the full semantic pipeline:
use kcl_sema::resolver::Resolver;
use kcl_driver::load_program;
fn compile(path: &str) -> Result<(), kcl_error::Error> {
// Load the parsed AST from kcl_parser
let program = load_program(path)?;
// Initialize and run the semantic resolver
let mut resolver = Resolver::new();
resolver.resolve_program(&program)?;
// Access the semantic information (types, symbols, diagnostics)
let sem_info = resolver.semantics();
println!("Resolved symbols: {:#?}", sem_info.symbols);
Ok(())
}
Querying Types in the LSP
Language server handlers leverage the global state to provide hover information and type lookups:
use kcl_sema::core::global_state::GlobalState;
use kcl_lsp::handlers::hover;
fn handle_hover(state: &GlobalState, position: Position) -> Hover {
// Lookup symbol at cursor position
if let Some(sym) = state.lookup_symbol(position) {
Hover {
contents: HoverContents::Scalar(
MarkedString::String(sym.typ.to_string())
),
range: Some(sym.range),
}
} else {
Hover::default()
}
}
These entry points—Resolver::resolve_program and GlobalState::lookup_symbol—form the public API surface that integrates semantic analysis into the broader KCL ecosystem.
Summary
-
kcl-sema implements an eight-stage pipeline spanning pre-processing, name resolution, type construction, erasure, unification, validation, and export.
-
Name resolution occurs in
src/resolver/mod.rs, building hierarchical scopes and symbol tables before type checking begins. -
Type inference uses Hindley-Milner unification via
src/ty/unify.rsto resolve polymorphic and concrete types. -
Constraint validation happens in
src/lint/mod.rs, enforcing schema rules and generating diagnostics for the LSP. -
Semantic information exports through
src/core/semantic_information.rs, providing structured data for evaluators and IDE tooling. -
The public API centers on
Resolver::resolve_programfor batch compilation andGlobalState::lookup_symbolfor interactive querying.
Frequently Asked Questions
What algorithm does kcl-sema use for type inference?
The module implements a Hindley-Milner-style unification algorithm via ty::Unify::unify in src/ty/unify.rs. This approach computes the most general type for expressions by merging type constraints and substituting type variables with concrete instances, supporting KCL's polymorphic schema and function definitions.
How does kcl-sema handle errors during semantic analysis?
Errors are collected as Diagnostic objects throughout the pipeline. During name resolution in src/resolver/mod.rs and constraint checking in src/lint/mod.rs, the system records mismatches, undefined symbols, and type conflicts. These diagnostics are bundled into the SemanticInformation structure exported from src/core/semantic_information.rs, allowing the LSP to surface specific error messages with source ranges.
What is the relationship between kcl-sema and the KCL evaluator?
kcl-sema runs as a prerequisite to evaluation. After the resolver completes and exports semantic information to GlobalState, the evaluator consumes these validated symbol tables and type annotations to execute the configuration logic. This separation ensures that type errors and undefined references are caught before runtime, preventing invalid configurations from reaching the target infrastructure.
Where are built-in types like list and dict defined in the codebase?
Built-in primitive types and standard library modules are defined in src/builtin/mod.rs, with specific implementations for strings in src/builtin/string.rs, optional types in src/builtin/option.rs, and system modules in src/builtin/system_module.rs. These definitions provide the foundational type signatures used during the construction and unification phases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →