How Language Auto-Detection Works for Node Summaries and Dashboard UI in Understand-Anything

Language auto-detection identifies programming languages by analyzing file paths via LanguageRegistry, then enriches graph nodes with LLM-generated languageNotes that render contextually in the dashboard's NodeInfo component.

The Understand-Anything repository implements an intelligent detection pipeline that automatically categorizes source files when building knowledge graphs. This system eliminates manual configuration by deriving language context from filenames and extensions, then surfacing programming-specific insights directly in the dashboard interface.

Detecting Languages via LanguageRegistry

The core detection logic resides in packages/core/src/languages/language-registry.ts. When the graph builder processes a file, it invokes LanguageRegistry.getForFile(filePath) to resolve the appropriate language configuration.

The registry performs a two-stage lookup:

  1. Filename matching – It first checks for exact filename matches (e.g., Dockerfile, Makefile) using a case-insensitive comparison.
  2. Extension fallback – If no filename match exists, it extracts the file extension (normalizing to lowercase with a leading dot) and queries the extension registry.
// packages/core/src/languages/language-registry.ts#L38
getForFile(filePath: string): LanguageConfig | null {
  const basename = filePath.split("/").pop() ?? filePath;
  const filenameMatch = this.byFilename.get(basename.toLowerCase());
  if (filenameMatch) return filenameMatch;
  const lastDot = filePath.lastIndexOf(".");
  if (lastDot === -1) return null;
  const ext = filePath.slice(lastDot).toLowerCase();
  return this.getByExtension(ext);
}

If a matching LanguageConfig is found, its id (such as typescript or python) is assigned to the node; otherwise, the language is marked as unknown.

Integrating Detection into Node Creation

During node instantiation in packages/core/src/plugins/registry.ts, the system calls the registry to set the node's language field before processing content. This integration ensures every GraphNode object carries accurate metadata about its programming language from the moment it enters the knowledge graph.

Generating Language-Aware Summaries

Once the language is identified, the analyzer enriches nodes with contextual programming insights. The packages/core/src/analyzer/language-lesson.ts module handles this through concept detection and LLM prompt generation.

Detecting Language Concepts and Prompting

The detectLanguageConcepts function scans the node's tags, summary, and existing notes for keywords defined in the base concept map and any language-specific configurations. Then buildLanguageLessonPrompt constructs a prompt requesting the LLM to generate a concise languageNotes field and concept explanations.

// packages/core/src/analyzer/language-lesson.ts#L69
export function detectLanguageConcepts(
  node: GraphNode, 
  language: string, 
  langConfig?: LanguageConfig | null
): string[] {
  // Scans node content for language-specific patterns
  // Returns detected concepts for lesson generation
}

The LLM response is parsed via parseLanguageLessonResponse, and the resulting languageNotes are persisted on the node object.

Displaying Language Context in the Dashboard UI

The dashboard surface renders these insights in packages/dashboard/src/components/NodeInfo.tsx. The React component conditionally displays the languageNotes field when present, providing users with immediate programming context when selecting nodes.

// packages/dashboard/src/components/NodeInfo.tsx#L412
{node.languageNotes && (
  <div className="mt-2 text-sm text-gray-400">{node.languageNotes}</div>
)}

This implementation ensures that clicking a node in the graph view reveals both the auto-detected language identifier and the LLM-generated educational content, creating a seamless bridge between raw code and contextual understanding.

Summary

Frequently Asked Questions

How does Understand-Anything handle files without standard extensions?

The system checks the exact filename against known mappings in LanguageRegistry (such as Dockerfile or Makefile) before attempting extension-based detection. If neither strategy matches, the node language defaults to unknown.

Where are the generated language notes stored within the node structure?

The LLM-generated insights are stored in the languageNotes field of the GraphNode object. This field is populated by parseLanguageLessonResponse after processing the language lesson prompt output.

Can the detection logic be extended for proprietary or custom file types?

Yes, developers can extend the LanguageRegistry with additional LanguageConfig entries mapping custom filenames or extensions to language IDs. The registry supports dynamic registration of new language definitions.

How does the UI behave when a node's language cannot be determined?

When getForFile returns null, the node language is set to unknown and no languageNotes are generated. The NodeInfo component uses conditional rendering to hide the language notes section entirely when the field is undefined or empty.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →