How WeKnora Automatic Database Migrations Run on Upgrade and Handle Partial Failures

WeKnora executes automatic database migrations during application startup through internal/database/migration.go, detecting dirty states via m.Version() and offering configurable auto-recovery via AutoRecoverDirty to rollback and retry partial failures.

WeKnora automates schema upgrades by invoking its migration engine immediately when the server boots, ensuring PostgreSQL and SQLite databases stay synchronized with the application code. The migration system in internal/database/migration.go provides robust mechanisms to detect interrupted upgrades and optionally recover without manual intervention. Understanding how these automatic database migrations handle corrupted states is critical for maintaining high availability in production environments.

Migration Startup and Entry Points

The migration process begins when the server calls database.RunMigrations(dsn) during bootstrap, typically invoked from cmd/server/bootstrap.go. This function logs the migration start and delegates to RunMigrationsWithOptions with default settings that disable automatic recovery.

For scenarios requiring custom behavior—such as CI/CD pipelines or SQLite file paths—use the explicit options struct:

opts := database.MigrationOptions{
    AutoRecoverDirty: true,              // Enable automatic dirty state recovery
    SQLiteDBPath:     "/tmp/weknora.db", // Required only for SQLite connections
}
err := database.RunMigrationsWithOptions(dsn, opts)

By default, AutoRecoverDirty is set to false, forcing administrators to manually resolve schema conflicts to prevent accidental data loss during startup failures.

Selecting Migration Sources and Engine Initialization

Based on the provided DSN, WeKnora selects the appropriate migration source directory. PostgreSQL databases use file://migrations/versioned, while SQLite connections default to file://migrations/sqlite.

SQLite-Specific Initialization

When MigrationOptions.SQLiteDBPath is provided, the engine opens the database file directly using sql.Open with the sqlite3 driver, then constructs a migrate.Migrate instance manually. This bypasses the standard DSN parsing to accommodate SQLite's file-based architecture.

Standard Database Initialization

For PostgreSQL and other supported engines, the migrator initializes via migrate.New using the connection string. This creates a versioned migration controller that tracks applied changes in the database's schema versioning table.

Detecting Dirty States Before Migration

Before applying any changes, WeKnora queries the current schema state using m.Version(), which returns the current version number and a dirty flag indicating whether a previous migration failed partway through. If this query fails, the error is captured and stored for exposure via the system information endpoint.

Handling Pre-Existing Dirty States

If the database is marked dirty and AutoRecoverDirty is enabled, the engine invokes recoverFromDirtyState to force the version back to the previous successful migration and clear the dirty flag. If auto-recovery is disabled, the function returns a detailed error instructing operators to manually force the version using ./scripts/migrate.sh force <version> or make migrate-force.

Executing Schema Upgrades and Failure Recovery

The engine executes pending migrations by calling m.Up(). Upon success, setMigrationState stores the final version and dirty flag in a thread-safe global cache for observability.

Partial Failure Handling

When m.Up() returns an error other than migrate.ErrNoChange, WeKnora checks whether the migration left the database in a dirty state:

  • With Auto-Recovery Enabled: If dirty, recoverFromDirtyState forces the migrator to the previous version (or -1 if at version 0) and automatically retries the migration.
  • With Auto-Recovery Disabled: The system returns an error message directing operators to force the version back to the last successful migration and retry manually using the provided CLI scripts.

The recoverFromDirtyState function relies on idempotent initial migrations to safely force version -1 when recovering from a dirty state at version 0.

Observability and Migration State Caching

Throughout the migration lifecycle, WeKnora maintains the current version, dirty status, and any error messages in a mutex-protected cache (migrationStateMu). The helpers CachedMigrationVersion() and CachedMigrationError() expose these values to the /info endpoint, allowing administrators to monitor whether a migration succeeded or stopped part-way through without inspecting database logs.

// Health-check handler example
ver, dirty, ok := database.CachedMigrationVersion()
if !ok {
    fmt.Println("Migration version not captured yet")
} else if dirty {
    fmt.Printf("Database at version %d is dirty – manual intervention required\n", ver)
} else {
    fmt.Printf("Database schema up-to-date at version %d\n", ver)
}

Summary

  • Startup Integration: Migrations run automatically via RunMigrations or RunMigrationsWithOptions when the WeKnora server starts.
  • Engine Selection: The system chooses between PostgreSQL (migrations/versioned) and SQLite (migrations/sqlite) source directories based on the DSN.
  • Dirty State Detection: m.Version() checks for partial failures before and after running m.Up().
  • Configurable Recovery: Setting AutoRecoverDirty: true enables automatic rollback to the previous version via recoverFromDirtyState, while the default setting requires manual intervention using migrate.sh scripts.
  • Operational Visibility: Thread-safe caching via CachedMigrationVersion() exposes migration status to the /info endpoint for real-time monitoring.

Frequently Asked Questions

How does WeKnora detect a failed migration?

WeKnora detects failures by checking the dirty flag returned from m.Version() in internal/database/migration.go. When a migration interrupts mid-execution, the underlying golang-migrate library marks the schema as dirty and stores the partial version number, which WeKnora queries both before and after running m.Up().

What happens if a migration is interrupted halfway through?

If m.Up() fails and leaves the database dirty, WeKnora checks the AutoRecoverDirty option. When enabled, it automatically calls recoverFromDirtyState to force the version to the previous successful migration (or -1 for initial states) and retries. If disabled, the application exits with an error directing operators to use ./scripts/migrate.sh force <version> to manually reset the schema version.

Can I enable automatic recovery for dirty migrations in production?

While you can enable AutoRecoverDirty: true in MigrationOptions for automated recovery, exercise caution in production. Automatic rollback assumes migrations are idempotent, particularly the initialization sequence at version 0. Test thoroughly in staging environments before enabling this feature on production databases to prevent unintended data consistency issues.

How do I check the current migration status via the API?

Import github.com/Tencent/WeKnora/internal/database and call CachedMigrationVersion(), which returns the current version, dirty boolean, and a status boolean indicating whether state has been captured. This data populates the /info endpoint, allowing you to monitor schema health without direct database access.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →