Where Are the Data Processing Scripts in DeepTutor? Complete Guide to the scripts/ Directory
The data processing scripts in DeepTutor are located in the top-level scripts/ directory, containing standalone Python utilities for knowledge base migration, user data transfer, embedding generation, and web UI initialization.
DeepTutor is an open-source AI tutoring system that manages structured knowledge bases, user workspaces, and LLM-generated embeddings. All data processing scripts in DeepTutor reside in the repository's scripts/ folder, providing self-contained utilities that operate on the data/ directory structure and interact with core configuration services.
The scripts/ Directory: Core Data Processing Utilities
The scripts/ package contains short, executable Python programs that handle common data-centric tasks before or after tutoring sessions. These files import from deeptutor/config and deeptutor/services to ensure consistent path resolution and configuration access across the codebase.
The primary data processing scripts include:
migrate_kb.py– Migrates legacy knowledge base layouts to the current structuremigrate_user_data.py– Transfers user-specific data including settings and logstest_embedding.py– Generates vector embeddings for document collectionssync_prompts_from_en.py– Synchronizes English prompt templates from upstream repositoriesstart_web.py– Bootstraps the FastAPI web interface while ensuring data directories exist
Additional helper utilities such as check_install.py, audit_prompts.py, start_tour.py, _cli_kit.py, and generate_roster.py provide environment validation and data preparation functions.
Key Data Migration Scripts
migrate_kb.py: Legacy Knowledge Base Migration
Located at scripts/migrate_kb.py, this utility walks the old data/knowledge_bases layout and copies files into the new directory structure. The script updates metadata files during the migration process, ensuring legacy content remains accessible under the current schema.
migrate_user_data.py: User Data Transfer
The scripts/migrate_user_data.py script handles migration of user-specific information including settings, logs, and workspace content from previous DeepTutor versions. It moves data into the freshly created data/user hierarchy while preserving file integrity and configuration history.
Embedding Generation and Prompt Management
test_embedding.py: Vector Generation
For generating embeddings, scripts/test_embedding.py calls the configured LLM embedding provider and stores resulting vectors under data/knowledge_bases/<kb>/embeddings. This script utilizes the embedding adapters defined in deeptutor/services/embedding/ to support OpenAI-compatible APIs.
sync_prompts_from_en.py: Prompt Synchronization
The scripts/sync_prompts_from_en.py utility pulls the latest prompt templates from upstream English repositories and writes them to deeptutor/prompts/. This ensures the tutoring system operates with current, standardized prompt definitions.
Web Interface and Environment Setup
start_web.py: Web UI Bootstrap
The scripts/start_web.py script launches the FastAPI/WebSocket server that serves the tutoring interface. Before starting the application, it imports PathService to guarantee that data/user/workspace and data/user/logs directories exist, preventing runtime path errors.
Supporting Utilities
Several additional scripts in the scripts/ directory support data processing workflows:
check_install.py– Validates environment dependencies and directory permissionsaudit_prompts.py– Reviews prompt file integrity and structuregenerate_roster.py– Creates user or content rosters for batch operations
Path Resolution and Configuration Services
Beyond standalone scripts, runtime data handling relies on two core modules that provide canonical paths and configuration access.
path_service.py: Centralized Directory Management
The deeptutor/services/path_service.py module defines the PathService class, which exposes properties like user_data_dir and knowledge_bases_dir. At line 104, the user_data_dir property provides the canonical path that every data processing script imports to ensure consistent directory navigation across the application.
accessors.py: Configuration Access Layer
Located at deeptutor/config/accessors.py, this module offers thin wrappers around global configuration values. The user_data_dir() accessor at line 16 provides the configuration-backed path resolution used by scripts when initializing file operations or validating environment state.
Practical Usage Examples
Running Knowledge Base Migration
Execute the legacy migration from the repository root:
python scripts/migrate_kb.py \
--src ./data/knowledge_bases_legacy \
--dst ./data/knowledge_bases
The script loads configuration via deeptutor/config/accessors.py to locate default directories, then walks the source tree to create the new layout with updated metadata.
Generating Document Embeddings
The following Python code demonstrates using PathService and the embedding adapter to process documents:
import pathlib
import numpy as np
from deeptutor.services.path_service import PathService
from deeptutor.services.embedding.adapters.openai_compatible import OpenAIEmbeddingAdapter
# Resolve the target knowledge base directory
kb_path = PathService().knowledge_bases_dir / "my_kb"
# Initialize the OpenAI-compatible adapter
adapter = OpenAIEmbeddingAdapter()
# Load plain-text documents
docs = [p.read_text(encoding="utf-8") for p in kb_path.rglob("*.txt")]
# Generate and store embeddings
embeddings = adapter.embed_documents(docs)
emb_path = kb_path / "embeddings" / "vectors.npy"
emb_path.parent.mkdir(parents=True, exist_ok=True)
np.save(emb_path, np.array(embeddings))
print(f"Saved {len(embeddings)} embeddings to {emb_path}")
Starting the Web Interface
Launch the application while ensuring data directories exist:
python scripts/start_web.py --host 0.0.0.0 --port 8000
Summary
- The data processing scripts in DeepTutor are centralized in the top-level
scripts/directory migrate_kb.pyandmigrate_user_data.pyhandle structural migrations between versionstest_embedding.pygenerates vector embeddings using the OpenAI-compatible adapterdeeptutor/services/path_service.pyprovides canonical path resolution viaPathServicedeeptutor/config/accessors.pyoffers configuration access for all scripts- Utilities follow a consistent pattern of loading configuration and validating paths before execution
Frequently Asked Questions
Where exactly are the data processing scripts located in the DeepTutor repository?
The data processing scripts are located in the scripts/ directory at the repository root. This folder contains standalone Python files including migrate_kb.py, test_embedding.py, and start_web.py that handle knowledge base operations, embedding generation, and web server initialization.
How do the scripts know where to find user data and knowledge bases?
All scripts import PathService from deeptutor/services/path_service.py, which provides properties like user_data_dir and knowledge_bases_dir. This centralized service ensures consistent path resolution across the application, reading from configuration values exposed through deeptutor/config/accessors.py.
Can I run the embedding generation script independently of the main application?
Yes. scripts/test_embedding.py is a self-contained utility that imports the embedding adapters from deeptutor/services/embedding/ and can generate vectors for any document collection in the knowledge base directory without requiring the full web interface to be running.
What is the purpose of the sync_prompts_from_en.py script?
The sync_prompts_from_en.py script pulls the latest English prompt templates from upstream repositories and synchronizes them into deeptutor/prompts/. This ensures the tutoring system uses current prompt definitions and maintains consistency with the main English language repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →