What Data Does everyone-can-use-english Process? A Complete Guide to Media, Speech, and Linguistic Data Types

everyone-can-use-english processes audio/video files, speech segments, transcriptions, dictionary data, and user settings through a desktop-first architecture built on TypeScript, SQLite, and ffmpeg.

This open-source language-learning application ingests diverse media formats, extracts linguistic artifacts, and synchronizes everything with a remote API. Below is a comprehensive breakdown of each data type, where it originates, and how it transforms through the codebase.


Media Files: Video and Audio Ingestion

The primary input data types are video files (MP4, MOV) and audio files (MP3, WAV). These enter the application through dedicated model factories that handle validation, hashing, and cloud preparation.

Video Data Pipeline

Video files are processed in src/main/db/models/video.ts. The Video.buildFromLocalFile() method serves as the main entry point:

import { Video } from '@/main/db/models';

await Video.buildFromLocalFile('/path/to/lecture.mp4', {
  name: 'English Lecture',
  description: 'A 30-min lecture on idioms',
  source: 'file',
  compressing: true,
});

Key transformations include:

  • MD5 hash generation for deduplication
  • Optional compression via ffmpeg transcoding
  • Metadata extraction: duration, codec, MIME type
  • Cover image generation stored in cloud storage

The Video model exposes getters like src, mimeType, and duration that read from a cached metadata JSON column.

Audio Data Pipeline

Audio files follow an identical pattern in src/main/db/models/audio.ts, using Audio.buildFromLocalFile() with the same transformation stack: hashing, conversion, metadata extraction, and upload.


Speech Segments: Extracted Audio Tracks

Speech segments represent the crucial intermediate data type that bridges raw media and usable text. These are extracted from video/audio tracks by the ffmpeg wrapper (src/main/ffmpeg.ts) and stored in src/main/db/models/speech.ts.

The Speech model captures:

  • Language code detection
  • Timestamps relative to source media
  • Confidence scores from recognition services

Speech segments enable the application's core language-learning functionality by isolating listenable, repeatable audio chunks.


Transcriptions: Full-Text Output from LLM Services

Transcriptions are generated by calling remote speech-to-text services, stored in src/main/db/models/transcription.ts. The model implements a state machine tracking processing status:

// Transcription states: processing → finished → error

Each transcription record contains:

  • Full-text output from the recognition service
  • Start and end timestamps
  • Reference to the source Speech segment
  • Processing state for UI feedback

Dictionary Data: Multilingual Lexical Resources

The application downloads and caches dictionary files for in-app lookup. The download-dictionaries script runs on first launch, pulling resources into local storage.

Dictionary handling lives in two key files:

import { getDict } from '@/main/dict';

const dict = await getDict('oxford');  // Loads Oxford dictionary JSON

Multiple formats are supported including MDict and custom JSON structures, merged for unified lookup in the UI.


User Settings and Credentials

User-specific configuration persists in SQLite via src/main/db/models/user-setting.ts. This lightweight model stores:

  • API authentication tokens (UserSetting.accessToken())
  • UI preferences and layout state
  • File system paths via settings.userDataPath()

The settings module at src/main/settings.ts provides centralized access to API URLs and environment configuration.


Metadata: Technical Properties and Derived Assets

Metadata encompasses technical properties and generated assets:

  • Cover images (generated on-demand via video.generateCover())
  • MIME type detection
  • Duration and format specifications

The FfmpegWrapper class generates these on-the-fly, with results cached in database JSON columns to avoid repeated processing.


API Synchronization: Client-Server Data Flow

All local data types sync with remote services through src/api/client.ts. The Client class exposes type-specific methods:

import { Client } from '@/api';
import settings from '@/main/settings';

const api = new Client({
  baseUrl: settings.apiUrl(),
  accessToken: await UserSetting.accessToken(),
});

await api.syncVideo(video.toJSON());

Synchronization covers videos, audio files, transcriptions, and deletions—maintaining consistency between offline SQLite storage and cloud state.


Key Entry Points in the Codebase

File Data Responsibility
enjoy/src/main/db/models/video.ts Video ingestion, compression, cover generation
enjoy/src/main/db/models/audio.ts Audio file processing (mirrors video)
enjoy/src/main/db/models/speech.ts Speech segment extraction and storage
enjoy/src/main/db/models/transcription.ts LLM-generated text with state tracking
enjoy/src/main/db/models/user-setting.ts Credentials and preferences
enjoy/src/main/ffmpeg.ts Media transformation engine
enjoy/src/main/dict.ts Dictionary loading and lookup
enjoy/src/api/client.ts Remote synchronization client

Summary

everyone-can-use-english processes six core data categories:

  • Media files — videos and audio with compression and cloud upload
  • Speech segments — extracted audio tracks for language practice
  • Transcriptions — full text from speech-to-text services with state tracking
  • Metadata — technical properties and generated cover images
  • Dictionaries — multilingual lexical resources for lookup
  • User settings — authentication, preferences, and paths

All data flows through Sequelize-TypeScript models, transforms via ffmpeg utilities, and synchronizes through a REST/GraphQL client—enabling offline-first operation with seamless cloud integration.


Frequently Asked Questions

What file formats does everyone-can-use-english support?

The application supports MP4, MOV for video and MP3, WAV for audio, as declared in enjoy/src/constants/index.ts. Additional formats may work through ffmpeg's broad codec support, but these are the officially supported extensions.

How does speech-to-text processing work in the application?

Speech is extracted via FfmpegWrapper, stored as a Speech model, then sent to remote LLM services. The resulting Transcription tracks processing state from processing through finished, with full text and timestamps preserved for user interaction.

Where is user data stored locally?

All data persists in a SQLite database managed by Sequelize. Media files reside in configurable local paths via settings.userDataPath(), while metadata, credentials, and processing state live in typed model tables.

Can the application work offline?

Yes. The architecture is offline-first: media ingestion, speech extraction, and local dictionary lookup function without connectivity. The Client.syncVideo() and related methods queue changes for synchronization when connection resumes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →