What Structured Analysis Does the LLM Produce in the First Ingest Step?
The LLM generates a six-section structured analysis covering entities, concepts, arguments, connections, contradictions, and recommendations while preserving all structured source data verbatim.
During the first step of the ingest pipeline in the nashsu/llm_wiki repository, the system orchestrates a structured analysis to decompose source documents into machine-readable knowledge components. This analysis is constructed by the buildAnalysisPrompt function in src/lib/ingest.ts (lines 2162–2220) and executed with reasoning explicitly disabled to ensure clean, sectioned output without chain-of-thought artifacts. The prompt instructs the model to act as an expert research analyst, producing a concise summary that the downstream ingestion logic can parse and integrate into the knowledge graph.
The Six Sections of the Structured Analysis
The prompt mandates that the LLM organize its output into six distinct sections, each targeting a specific dimension of knowledge extraction.
Key Entities
The model identifies people, organizations, products, datasets, and tools mentioned in the source. For each entity, the analysis must specify the name, type, role in the context, and whether the entity already exists in the wiki. This mapping prevents duplicate entries and maintains referential integrity across the knowledge base.
Key Concepts
This section captures theories, methods, techniques, and phenomena. Each concept entry includes a brief definition, its relevance to the source material, and a check against the current wiki index to determine if the concept is novel or extant.
Main Arguments & Findings
The LLM extracts core claims, supporting evidence, and an assessment of argument strength. Crucially, this section requires an explicit mapping of claims to the named subject, ensuring that the wiki can attribute findings to specific sources with full provenance.
Connections to Existing Wiki
The analysis evaluates how the source material relates to existing pages, determining whether it strengthens, challenges, or extends current knowledge. This relational metadata enables the system to suggest graph links between articles automatically.
Contradictions & Tensions
Here the model flags any conflicts with existing wiki content or internal inconsistencies within the source itself. This quality-control step surfaces editorial decisions for human review before the content is committed to the permanent record.
Recommendations
The final section proposes wiki pages to create or update, mapping suggestions to custom page types defined in the project schema. It also captures emphasis directives (highlighting critical insights) and open questions that warrant further investigation.
Preserving Structured Source Data
Beyond the six analytical sections, the prompt mandates verbatim preservation of structured data encountered in the source. When the LLM encounters SQL DDL, schema definitions, API signatures, configurations, or tabular data, it must output these elements in fenced code blocks or Markdown tables without interpretation or summarization. This ensures that technical specifications remain intact for accurate downstream processing.
How the Analysis Prompt Is Constructed
The buildAnalysisPrompt function assembles the instruction set that drives this behavior. It optionally injects the project schema, purpose, and current wiki index to ground the model in the existing knowledge context.
// Build the analysis prompt (called by the ingest pipeline)
import { buildAnalysisPrompt } from "@/lib/ingest";
const purpose = "Research on renewable energy";
const index = await getCurrentWikiIndex(); // optional context
const schema = await getProjectSchema(); // optional schema
const source = await readFile(sourcePath);
const prompt = buildAnalysisPrompt(purpose, index, source, schema);
The returned prompt string is then dispatched to the LLM via streamChat with structured-output safeguards disabled.
Execution with Reasoning Disabled
To prevent the model from emitting chain-of-thought or hidden reasoning that would corrupt the parseable output, the ingest pipeline explicitly sets reasoning: { mode: "off" } during the analysis phase.
// Example call to the LLM client
const analysis = await streamChat({
model: "gpt-4o-mini",
purpose: "ingest",
prompt,
maxTokens: 8192,
// Structured output is off for this call
reasoning: { mode: "off" },
});
As implemented in src/lib/reasoning-capabilities.ts, this configuration ensures that streamChat returns only the final structured analysis, compliant with the section headers defined in src/lib/ingest.ts.
Summary
- The structured analysis consists of six mandatory sections: Key Entities, Key Concepts, Main Arguments & Findings, Connections to Existing Wiki, Contradictions & Tensions, and Recommendations.
- The
buildAnalysisPromptfunction insrc/lib/ingest.ts(lines 2162–2220) defines the prompt template and optional context injection. - Structured source data (schemas, code, configurations) must be preserved verbatim in code blocks or tables.
- Reasoning is explicitly disabled via
reasoning: { mode: "off" }to ensure clean, parseable output without chain-of-thought artifacts. - The
ingestReasoningconfiguration insrc/stores/wiki-store.tscontrols this behavior toggle.
Frequently Asked Questions
What role does buildAnalysisPrompt play in the ingest pipeline?
The buildAnalysisPrompt function in src/lib/ingest.ts constructs the system instructions that tell the LLM how to decompose a source document. It assembles the six-section schema, appends optional project context (schema, purpose, wiki index), and embeds the verbatim preservation rules. This prompt is then passed to streamChat for execution.
Why is reasoning mode disabled during the structured analysis step?
Reasoning is disabled by setting reasoning: { mode: "off" } to prevent the model from outputting chain-of-thought or explanatory text that would break the parseable structure. According to src/lib/reasoning-capabilities.ts, this ensures the LLM returns only the final sectioned analysis, making it possible for the ingestion parser to reliably extract entities, concepts, and recommendations.
How does the LLM handle existing wiki content during analysis?
The prompt includes an optional index parameter containing the current wiki state. The model uses this to check whether extracted entities and concepts already exist, and to determine if new findings strengthen, challenge, or extend existing knowledge. This relational assessment populates the "Connections to Existing Wiki" and "Contradictions & Tensions" sections.
What configuration controls the structured analysis behavior?
The ingestReasoning setting in src/stores/wiki-store.ts governs whether the ingest pipeline uses structured analysis with reasoning disabled. Additionally, the project schema and purpose parameters passed to buildAnalysisPrompt shape the recommendations section by defining custom page types and semantic constraints specific to the wiki project.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →