Can AI Teachers Actively Operate the UI in OpenMAIC? Deep Interactive Mode Explained
Yes, AI teachers in OpenMAIC can actively operate the UI by programmatically highlighting elements, displaying contextual hints, and directing learner attention through an integrated action-execution engine powered by LangGraph orchestration.
The THU-MAIC/OpenMAIC platform implements a Deep Interactive Mode that transforms AI teachers from passive conversational agents into active interface operators. Unlike traditional educational AI systems that only respond to student messages, OpenMAIC's teacher agents can directly manipulate the DOM to guide learning experiences in real-time.
How AI Teachers Control the Interface
Deep Interactive Mode Architecture
According to the OpenMAIC documentation, Deep Interactive Mode explicitly enables the AI teacher to "actively operate the UI to guide students—highlighting key areas, setting conditions, providing hints, and directing attention at the right moments." This capability is not limited to visual indicators; the teacher role has full access to the action execution pipeline that drives all classroom interactions.
The system identifies the teacher agent through a strict role-based filtering mechanism. In components/agent/agent-bar.tsx and components/roundtable/index.tsx, the UI components locate the teacher by filtering the participants array for agent.role === 'teacher'. This identification enables the system to anchor UI actions to the teacher's avatar and bubble interface.
The Action Execution Engine
At the core of this functionality lies the lib/action directory, which defines 28+ concrete action types including highlight, showHint, and focus. When the teacher agent decides to intervene, the LangGraph orchestration layer in lib/orchestration generates an action node—such as {type: 'highlight', targetId: 'slide-12'}—and publishes it to the runtime event stream.
The lib/action/dispatcher.ts module receives these events and forwards them to specialized handlers. For example, the highlight action defined in lib/action/highlight.ts manipulates the DOM directly by adding CSS classes and managing scroll behavior.
Technical Implementation of Teacher UI Operations
Teacher Role Detection and Avatar Anchoring
The roundtable interface in components/roundtable/index.tsx maintains a reference to the teacher's visual representation through teacherAvatarRef. This reference serves as the anchor point for visual cues, creating the impression that physical actions originate from the teacher avatar.
The component filters the participant list to isolate the teacher entity:
const teacher = participants.find(p => p.role === 'teacher');
const teacherAvatarRef = useRef<HTMLDivElement>(null);
When an action event arrives, the system can attach visual pointers or animations to this specific DOM node, reinforcing the perception that the AI teacher is physically manipulating the interface.
Action Dispatch and DOM Manipulation
The action execution flow follows a strict pipeline. First, the LangGraph director invokes emitAction from lib/action/dispatcher.ts within the teacher's decision node:
// lib/orchestration/teacher-actions.ts
import { emitAction } from '@/lib/action/dispatcher';
export async function teacherHighlightSlide(slideId: string) {
await emitAction({
type: 'highlight',
targetId: slideId,
meta: { color: '#ffeb3b', durationMs: 3000 },
});
}
The dispatcher routes this payload to the appropriate handler. The highlight implementation in lib/action/highlight.ts performs the actual DOM manipulation:
// lib/action/highlight.ts
export function highlightAction(payload: { targetId: string; meta: any }) {
const el = document.getElementById(payload.targetId);
if (!el) return;
el.classList.add('teacher-highlight');
setTimeout(() => el.classList.remove('teacher-highlight'), payload.meta.durationMs);
}
This handler adds a CSS class that applies visual styling (typically a bright yellow border) and automatically removes it after the specified duration.
Visual Feedback and Synchronized Speech
The components/roundtable/presentation-speech-overlay.tsx component receives a role prop ('teacher' | 'agent' | 'user') and renders UI decorations specifically when the teacher is active. When the teacher's text-to-speech system initiates, this overlay renders spotlights and animated borders that synchronize with the vocal guidance.
Additionally, the roundtable component subscribes to action events to render dynamic pointers:
// components/roundtable/index.tsx (excerpt)
useEffect(() => {
const unsub = subscribeToAction((action) => {
if (action.type === 'highlight') {
const pointer = document.createElement('div');
pointer.className = 'teacher-pointer';
teacherAvatarRef.current?.appendChild(pointer);
setTimeout(() => pointer.remove(), action.meta.durationMs);
}
});
return () => unsub();
}, []);
This creates a visual connection between the teacher avatar and the target element, simulating a pointing gesture that reinforces the AI's active role in the learning environment.
Practical Code Examples
Emitting UI Actions from Teacher Logic
When the teacher agent determines that a student requires visual guidance, it invokes the action pipeline through the orchestration layer:
// lib/orchestration/teacher-actions.ts
import { emitAction } from '@/lib/action/dispatcher';
// Triggered during LangGraph execution when pedagogical intervention is needed
export async function guideStudentAttention(elementId: string) {
await emitAction({
type: 'highlight',
targetId: elementId,
meta: {
color: '#ffeb3b',
durationMs: 3000,
priority: 'high'
},
});
}
Implementing Custom UI Action Handlers
New UI manipulation capabilities can be added by extending the action library. The following pattern demonstrates how to implement a focus action that scrolls elements into view:
// lib/action/focus.ts
export function focusAction(payload: { targetId: string; meta: any }) {
const element = document.getElementById(payload.targetId);
if (!element) return;
element.scrollIntoView({ behavior: 'smooth', block: 'center' });
element.classList.add('teacher-focus-ring');
setTimeout(() => {
element.classList.remove('teacher-focus-ring');
}, payload.meta.durationMs || 2000);
}
Synchronizing Speech with Visual Cues
The presentation overlay demonstrates how to bind UI effects to the teacher's speaking state:
// components/roundtable/presentation-speech-overlay.tsx
export function PresentationSpeechOverlay({ bubbleRole, participants }: Props) {
const teacher = participants.find(p => p.role === 'teacher');
return bubbleRole === 'teacher' ? (
<div className="spotlight-overlay">
<div className="teacher-presence-indicator"
style={{ backgroundColor: teacher?.color }} />
</div>
) : null;
}
Summary
- AI teachers in OpenMAIC possess active UI manipulation capabilities through the Deep Interactive Mode architecture, enabling them to highlight elements, show hints, and focus attention programmatically.
- Role-based identification in
components/agent/agent-bar.tsxandcomponents/roundtable/index.tsxdistinguishes the teacher from other agents using therole === 'teacher'filter. - The action execution engine in
lib/action/*processes 28+ action types through a centralized dispatcher that handles DOM manipulation. - LangGraph orchestration in
lib/orchestrationgenerates action nodes that the runtime executes, allowing complex multi-step UI guidance sequences. - Visual anchoring via
teacherAvatarRefcreates the perception that physical actions originate from the teacher's avatar, enhancing the immersive classroom experience.
Frequently Asked Questions
How does the AI teacher determine which UI elements to manipulate?
The AI teacher uses the standard DOM identification system, targeting elements by their id attribute. When the teacher's LangGraph node decides to intervene, it emits an action payload containing a targetId field that corresponds to the HTML element's ID. The action handlers in lib/action/highlight.ts and related modules use document.getElementById() to locate and manipulate these specific elements.
What types of UI operations can the AI teacher perform?
The current implementation supports 28+ distinct action types defined in the lib/action directory. Core capabilities include highlight (adding visual borders), showHint (displaying contextual tooltips), and focus (scrolling elements into view). Because the action engine is generic, developers can extend the system by adding new action handlers that perform operations like opening modals, expanding toolboxes, or changing interface themes.
Is the action execution system exclusive to the teacher role?
No, the action execution engine is generic and available to all agent types. However, the UI components in components/roundtable/presentation-speech-overlay.tsx apply special visual treatments when bubbleRole === 'teacher', and the teacherAvatarRef anchoring system specifically tracks the teacher's visual representation. Other agents can trigger actions, but the interface provides unique feedback cues for teacher-initiated operations to establish pedagogical authority.
How does the system handle visual feedback when the teacher takes action?
The system implements a multi-layered feedback mechanism. The lib/action/dispatcher.ts broadcasts action events to both DOM manipulation handlers and UI components. Simultaneously, components/roundtable/index.tsx subscribes to these events to render animated pointers attached to the teacher avatar, while presentation-speech-overlay.tsx manages spotlight effects. This dual approach ensures that students perceive both the direct UI change (the highlighted element) and the source of the action (the teacher avatar's gesture).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →