Understanding the Model Architecture of moeru-ai/airi: A Four-Layer Cognitive System

AIRI implements a four-layered cognitive architecture—Perception, Reflex, Conscious, and Action—that processes Minecraft events through an event-driven pipeline, enabling LLM-powered reasoning with reflexive safety controls.

The moeru-ai/airi project is a sophisticated Minecraft agent framework built around a unique model architecture that mirrors human cognitive processing. Unlike monolithic AI systems, AIRI organizes intelligence into distinct functional layers that separate sensing from thinking and planning from execution. This architecture is fully implemented in the services/minecraft package, where TypeScript modules in src/cognitive/ wire together a reactive system capable of both high-level reasoning and immediate reflexive responses, as documented in services/minecraft/README.md.

The Four-Layer Cognitive Model Architecture

AIRI's model architecture organizes cognition into four hierarchical layers, each with specific responsibilities and implementation locations.

Layer A: Perception

The Perception layer handles raw sensory input from the Minecraft environment through Mineflayer events. Located in src/cognitive/perception/, this layer normalizes raw game events into typed signals using the pipeline defined in src/cognitive/perception/pipeline.ts. It processes event definitions from events/definitions/ and emits derived signal:* events through a rule-based engine, transforming low-level game state changes into meaningful cognitive inputs for higher layers.

Layer B: Reflex (Sub-conscious)

The Reflex layer provides immediate, rule-based reactions that operate on a finite-state-machine (FSM) for speed and predictability. Implemented in src/cognitive/reflex/, the ReflexManager class coordinates reflex behaviors that can execute without conscious deliberation. A critical feature of this layer is its ability to inhibit the conscious layer—when immediate action is required (such as avoiding hostile mobs), the reflex system can temporarily disable high-level planning in src/cognitive/conscious/brain.ts to prevent redundant or conflicting processing.

Layer C: Conscious (Reasoning)

The Conscious layer handles high-level planning, decision-making, and dialogue driven by Large Language Models (LLMs). Located in src/cognitive/conscious/, this layer contains the Brain class that orchestrates the event queue and manages the LLM turn lifecycle. Supporting components include JsPlanner for running sandboxed planning scripts, QueryDSL for read-only world queries, and TaskState for managing cancellation tokens. Importantly, this layer performs no physical actions directly—it only creates plans that the Action layer executes.

Layer D: Action (Execution)

The Action layer serves as the physical interface to the Minecraft world, decoupling "doing" from "thinking." Found in src/cognitive/action/, this layer contains the TaskExecutor that runs normalized action instructions and the ActionRegistry that validates parameters and dispatches tool calls. The tool catalog is defined in llm-actions.ts, which maps LLM-intent to concrete Minecraft bot operations.

Event-Driven Communication

All four layers communicate through a centralized event bus implemented in src/cognitive/event-bus.ts. This pub-sub architecture enables loose coupling between components, allowing the Perception layer to emit events without knowing which layers will consume them.

The typical execution flow demonstrates this decoupling:

  1. Perception captures a player voice command and emits a raw event
  2. Conscious (via Brain and JsPlanner) creates a plan of discrete steps
  3. Action (TaskExecutor) carries out each step using tools from the catalog
  4. Reflex monitors continuously and can intervene at any point (e.g., emergency dodge) while potentially inhibiting the conscious layer for safety

Implementation Walkthrough

The following examples demonstrate how the model architecture operates in practice.

Bootstrapping the Cognitive System

The container pattern in src/cognitive/container.ts wires all four layers together at startup:

import { createContainer } from '@/cognitive/container'
import { EventBus } from '@/cognitive/event-bus'

async function startAiri() {
  const container = await createContainer()
  const bus = container.resolve(EventBus)

  // Listen for high‑level completion events
  bus.on('conscious:turn:finished', (msg) => {
    console.log('🧠 Turn finished:', msg)
  })

  // Emit a sample raw event to kick things off
  bus.emit('raw:chat:message', { content: 'build a house' })
}
startAiri()

The container.ts file handles dependency injection across all cognitive layers, while EventBus routes events between Perception, Reflex, Conscious, and Action components.

Implementing Reflex Inhibition

Reflex behaviors can pause conscious processing when immediate action is required:

// src/cognitive/reflex/behaviors/avoid-hostile.ts
import { ReflexManager } from '@/cognitive/reflex/reflex-manager'

export function registerAvoidHostile(reflex: ReflexManager) {
  reflex.register({
    id: 'avoid-hostile',
    trigger: 'signal:entity:hostile',
    action: (event, ctx) => {
      // Inhibit conscious processing for the next 2 seconds
      ctx.inhibitConscious(2000)
    },
  })
}

This pattern in src/cognitive/reflex/reflex-manager.ts ensures that survival-critical responses bypass slower LLM reasoning.

Extending the Action Catalog

New tools are registered in llm-actions.ts with Zod schemas for type-safe LLM interaction:

// src/cognitive/action/llm-actions.ts
import { z } from 'zod'

export const llmActions = {
  'minecraft:placeBlock': {
    description: 'Place a block at a specific position',
    schema: z.object({
      x: z.number(),
      y: z.number(),
      z: z.number(),
      block: z.string(),
    }),
    handler: async (params, context) => {
      await context.bot.placeBlock(params)
    },
  },

  'minecraft:craftItem': {
    description: 'Craft an item using the bot’s inventory',
    schema: z.object({ item: z.string(), count: z.number().optional() }),
    handler: async (p, ctx) => {
      await ctx.bot.craft(p.item, p.count ?? 1)
    },
  },
}

Tools defined here become automatically callable from LLM-generated plans executed by the TaskExecutor.

Summary

  • Four distinct layers—Perception, Reflex, Conscious, and Action—separate sensory processing from physical execution in the AIRI model architecture.
  • Event-driven architecture using EventBus enables loose coupling between cognitive layers.
  • Reflex inhibition allows the sub-conscious layer to override LLM reasoning for safety-critical responses.
  • Type-safe tool catalog in llm-actions.ts bridges high-level planning with concrete Minecraft operations.
  • Modular implementation in services/minecraft/src/cognitive/ follows principles of separation of concerns and cognitive realism.

Frequently Asked Questions

How does the reflex layer inhibit conscious processing?

The reflex layer calls ctx.inhibitConscious(milliseconds) within action handlers registered to the ReflexManager in src/cognitive/reflex/reflex-manager.ts. This temporarily pauses the Brain class in the conscious layer from initiating new LLM turns, ensuring that immediate physical responses (like avoiding enemies) aren't delayed by ongoing planning operations.

What role does the EventBus play in AIRI's model architecture?

The EventBus in src/cognitive/event-bus.ts serves as the central nervous system, implementing a publish-subscribe pattern that allows the four cognitive layers to communicate without direct dependencies. It routes normalized signals from Perception to Conscious, triggers reflex responses, and notifies the system when actions complete.

How are new capabilities added to the action layer?

Developers extend the llmActions object in src/cognitive/action/llm-actions.ts by defining Zod schemas for parameter validation and async handler functions that interface with the Mineflayer bot. These definitions automatically populate the tool catalog available to the JsPlanner in the conscious layer.

What distinguishes the conscious layer from the reflex layer?

The conscious layer in src/cognitive/conscious/brain.ts performs slow, deliberative reasoning using LLMs for planning and dialogue, while the reflex layer executes fast, deterministic finite-state-machine behaviors. The reflex layer can operate in parallel and override the conscious layer, but only the conscious layer generates novel plans for complex tasks like construction projects.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →