How RomM Library Scanning Works: A Technical Deep Dive into the Backend Pipeline
RomM's library scanning is a coordinated backend workflow that watches filesystem changes, enqueues scan jobs via Redis-RQ, processes ROM files with metadata enrichment, and emits real-time progress updates via Socket.IO to the Vue frontend.
RomM is a self-hosted web application for organizing retro game collections. Its library scanning feature automatically discovers ROM files, enriches them with metadata from gaming databases, and synchronizes the SQLAlchemy database with your filesystem. Understanding this pipeline is essential for troubleshooting scan failures and optimizing performance for large collections.
Entry Point: Triggering Scans via the API
When a user initiates a scan from the Vue frontend, the application calls POST /api/scan. This endpoint is defined in backend/endpoints/tasks.py and immediately forwards the request to the Redis-RQ queue rather than processing synchronously.
# Backend: the endpoint that queues a scan job
# File: backend/endpoints/tasks.py
@router.post('/scan')
async def scan_library(request: Request):
await request.app.state.rq.enqueue_job('scan_library')
return JSONResponse({'status': 'queued'})
Scan Handler and Job Coordination
Once the API receives the request, the scan handler (backend/handler/scan_handler.py) creates a new scan job and manages its lifecycle. This handler coordinates between the queue and the actual scanning logic, ensuring that only one intensive scan operation runs at a time while maintaining progress state.
The Scanning Pipeline: From Filesystem to Database
The core scanning logic resides in ScanLibraryTask, a periodic task defined in backend/tasks/scheduled/scan_library.py. This class extends PeriodicTask and runs automatically on startup as well as when triggered manually.
Filesystem Discovery with Watcher
Inside the task, RomM utilizes the Watcher utility (backend/watcher.py) to traverse directories configured in your .env or config.yml file. The watcher detects new, modified, or removed files and yields absolute paths for each ROM file it discovers.
# Core scan task – walks the library and creates ROM entries
# File: backend/tasks/scheduled/scan_library.py
class ScanLibraryTask(PeriodicTask):
async def run(self) -> None:
watcher = Watcher(self.settings.library_paths)
async for rom_path in watcher.iter_files():
await process_rom(rom_path) # creates Rom model, fetches metadata, etc.
ROM Processing and Metadata Enrichment
For each discovered file, RomM instantiates a Rom ORM object (backend/models/rom.py). The system then invokes metadata adapters located in backend/adapters/ to fetch game titles, descriptions, cover art, and other details from providers like IGDB, TheGamesDB, and OpenVGDB.
Hashing and Deduplication
Before committing to the database, RomM computes a deterministic hash using the utility in backend/utils/rvz_hasher.py. This hash ensures that duplicate ROM files are not created in the database, even if the same game exists in multiple formats or locations.
Database Commit and Search Indexing
After populating the ROM record, the task commits the SQLAlchemy session to the database. Simultaneously, it updates the search index (managed via the Elasticsearch-compatible index created in Alembic migration 0084_add_roms_search_index.py) to ensure new entries appear immediately in the frontend search interface.
Real-Time Progress Feedback
To keep users informed, RomM emits Socket.IO events during the scan process. The endpoint backend/endpoints/sockets/scan.py broadcasts scan_progress events that contain the current status, including counts of processed, skipped, and failed items.
The Vue frontend listens for these events through the Pinia store defined in frontend/src/stores/scanning.ts:
// Front‑end: trigger a full library scan
import { useApi } from '@/services/api';
import { useScanning } from '@/stores/scanning';
async function startScan() {
const api = useApi();
await api.post('/scan'); // calls the backend endpoint
const scanning = useScanning();
scanning.start(); // shows progress UI
}
Internationalized strings for the scanning interface are stored in frontend/src/locales/en_US/scan.json and referenced by the UI components to display localized progress messages.
Summary
- Async Queue Architecture: RomM offloads scanning to Redis-RQ via
backend/endpoints/tasks.py, preventing HTTP request timeouts during long operations. - Filesystem Monitoring: The
Watcherclass inbackend/watcher.pyefficiently detects changes in configured library paths without blocking the main thread. - Metadata Integration: Multiple adapter classes in
backend/adapters/enrich ROM records with cover art and game details from external APIs. - Real-Time Updates: Socket.IO events emitted from
backend/endpoints/sockets/scan.pyprovide live progress bars in the Vue frontend viafrontend/src/stores/scanning.ts. - Search Integration: New ROMs are immediately indexed through Alembic migration
0084_add_roms_search_index.py, ensuring instant searchability.
Frequently Asked Questions
How does RomM handle large libraries without timing out?
RomM uses Redis-RQ to queue scan jobs asynchronously. When you trigger a scan via POST /api/scan, the endpoint immediately returns a "queued" status and delegates the work to a background worker. This prevents HTTP timeouts and allows the frontend to display real-time progress via Socket.IO events while the backend processes files.
What metadata sources does RomM use during scanning?
During the scanning pipeline, RomM invokes metadata adapters located in backend/adapters/ to fetch game information. These adapters connect to external providers such as IGDB, TheGamesDB, and OpenVGDB to retrieve titles, descriptions, release dates, and cover art, which are then stored in the Rom model defined in backend/models/rom.py.
How does RomM prevent duplicate ROM entries?
RomM computes a deterministic hash for each file using backend/utils/rvz_hasher.py before creating database records. This hash acts as a deduplication key, ensuring that identical ROM files—whether they exist in multiple formats or directory locations—are not imported as separate entries in the SQLAlchemy database.
Why don't new ROMs appear in search immediately after scanning?
If new ROMs are missing from search results, the search index may not have updated properly. The scan task automatically updates the Elasticsearch-compatible index defined in Alembic migration 0084_add_roms_search_index.py after committing ROM records. Ensure your database migrations are current and the search service is running if using external Elasticsearch.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →