How to Customize Karakeep Features: Complete Configuration Guide
You can customize Karakeep by setting environment variables in the centralized configuration schema, enabling specific workers, switching storage backends to S3, or extending functionality through the plugin system.
Karakeep is a modular, self-hosted read-it-later platform built as a TypeScript monorepo. Unlike monolithic applications that require code changes for customization, Karakeep exposes nearly every behavioral toggle through environment variables parsed by packages/shared/config.ts. This guide covers the architectural patterns behind the configuration system and provides concrete examples for customizing workers, AI inference, asset storage, and plugins.
How Karakeep Configuration Works
Centralized Configuration Schema
All runtime behavior in Karakeep is driven by a single source of truth: packages/shared/config.ts. This file uses Zod to validate and coerce environment variables into two distinct objects:
serverConfig– The full configuration used by the server and background workers, including secrets like API keys and database credentials.clientConfig– A safe subset exposed to the browser, containing only public URLs and UI-related toggles.
The module parses process.env once at import time, meaning any environment variable you set is immediately available across the web UI (apps/web/server.ts), the API (packages/api/), and the worker bootstrap (apps/workers/index.ts).
Component Architecture
The monorepo is split into independent packages that you can tune separately:
- Web UI (
apps/web/) – Next.js frontend that consumesclientConfig. - Workers (
apps/workers/) – Background processes for crawling, AI inference, and search. - Plugins (
packages/plugins/) – Optional extensions for vector stores, rate limiting, and custom storage backends.
Customizing Workers and Background Processing
Selective Worker Enablement
Karakeep uses a dynamic worker system controlled by two environment variables: WORKERS_ENABLED_WORKERS and WORKERS_DISABLED_WORKERS. The bootstrap logic in apps/workers/index.ts imports and starts only the workers you specify:
// apps/workers/index.ts
import { serverConfig } from "@karakeep/shared/config";
for (const w of serverConfig.workers.enabledWorkers) {
if (!serverConfig.workers.disabledWorkers.includes(w)) {
await import(`./workers/${w}`);
}
}
Valid worker names include crawler, inference, search, assetPreprocessing, and webhook. To run a dedicated crawler node without AI processing, set:
WORKERS_ENABLED_WORKERS=crawler,search
WORKERS_DISABLED_WORKERS=inference
Worker-Specific Settings
Each worker reads its own configuration namespace from serverConfig. For example, the crawler worker (apps/workers/workers/crawler.ts) respects these variables:
CRAWLER_FULL_PAGE_SCREENSHOT=true
CRAWLER_PARSER_MEM_LIMIT_MB=1024
CRAWLER_JOB_TIMEOUT_SEC=120
CRAWLER_HTTP_PROXY=http://proxy:8080
The inference worker (apps/workers/workers/inference.ts) uses parallelization controls:
INFERENCE_NUM_WORKERS=4
INFERENCE_JOB_TIMEOUT_SEC=300
Configuring AI Inference and Auto-Tagging
AI Provider Setup
Automatic tag generation is handled by the inference worker, which supports both OpenAI and Ollama backends. To enable AI features, provide one of the following:
OPENAI_API_KEY=sk-*************
# OR
OLLAMA_BASE_URL=http://localhost:11434
Additional flags control the behavior:
INFERENCE_ENABLE_AUTO_TAGGING=true
INFERENCE_TEXT_MODEL=gpt-4o-mini
INFERENCE_IMAGE_MODEL=gpt-4o
INFERENCE_CONTEXT_LENGTH=4096
INFERENCE_OUTPUT_SCHEMA=structured
Without OPENAI_API_KEY or OLLAMA_BASE_URL set, the inference stack remains dormant.
Customizing AI Prompts
Users can append custom instructions to the tagging prompt through the UI or database. Custom prompts are stored in the customPrompts table and queried by the tRPC router in packages/trpc/routers/bookmarks.ts. To set a system-wide default programmatically:
import { db } from "@karakeep/db";
import { customPrompts } from "@karakeep/db/schema";
await db.insert(customPrompts).values({
userId: "system",
appliesTo: "tagging",
prompt: "Prefer tags that are nouns and avoid generic adjectives.",
});
The inference worker merges these snippets into the base prompt before sending requests to the LLM.
Switching Asset Storage Backends
Local Filesystem vs S3
By default, Karakeep stores screenshots and PDFs in DATA_DIR/assets on the local filesystem. To switch to S3-compatible storage (such as AWS S3 or MinIO), configure the ASSET_STORE_S3_* variables:
ASSET_STORE_S3_ENDPOINT=https://minio.local:9000
ASSET_STORE_S3_REGION=us-east-1
ASSET_STORE_S3_BUCKET=karakeep-assets
ASSET_STORE_S3_ACCESS_KEY_ID=minioadmin
ASSET_STORE_S3_SECRET_ACCESS_KEY=minioadmin
ASSET_STORE_S3_FORCE_PATH_STYLE=true
When ASSET_STORE_S3_ENDPOINT is present, serverConfig.assetStore resolves to type: "s3", and the asset preprocessing worker (apps/workers/workers/assetPreprocessing.ts) automatically uses the S3 client instead of the local filesystem.
Extending Karakeep with Plugins
Plugin System Architecture
Karakeep ships with a minimal plugin API located under packages/plugins/. Each plugin exports a setup function that receives a PluginContext containing trpc, db, and other shared resources. Representative built-in plugins include:
vectorstore-meilisearch– Provides Meilisearch-backed vector search.ratelimit-redis– Enables distributed rate limiting via Redis.
Plugins are imported conditionally by the server bootstrap based on environment flags.
Creating a Custom Plugin
To add bespoke business logic, create a new package under packages/plugins/:
// packages/plugins/my-custom-plugin/src/index.ts
import type { PluginContext } from "@karakeep/shared/types";
import { z } from "zod";
export default async function setup(ctx: PluginContext) {
ctx.trpcRouter.procedure("myPlugin/hello")
.input(z.object({ name: z.string() }))
.query(async ({ input }) => `Hello, ${input.name}!`);
}
Activate the plugin by importing it in apps/web/server.ts or apps/workers/trpc.ts when a configuration flag is set:
if (process.env.MY_PLUGIN_ENABLED === "true") {
await import("@karakeep/plugins/my-custom-plugin");
}
Practical Configuration Examples
Docker Compose Deployment
For a minimal deployment with only the web UI and crawler worker:
version: "3.9"
services:
web:
image: ghcr.io/karakeep-app/karakeep/web:latest
env_file: .env
ports:
- "3000:3000"
workers:
image: ghcr.io/karakeep-app/karakeep/workers:latest
env_file: .env
command: ["pnpm", "workers"]
Complete Environment Template
Combine worker selection, AI configuration, and S3 storage in a single .env file:
# Core
DATA_DIR=/var/karakeep/data
NEXTAUTH_URL=https://karakeep.example.com
NEXTAUTH_SECRET=$(openssl rand -base64 36)
# Workers
WORKERS_ENABLED_WORKERS=crawler,inference,search
WORKERS_DISABLED_WORKERS=webhook
# AI
OPENAI_API_KEY=sk-*************
INFERENCE_ENABLE_AUTO_TAGGING=true
INFERENCE_TEXT_MODEL=gpt-4o-mini
# Storage
ASSET_STORE_S3_ENDPOINT=https://s3.amazonaws.com
ASSET_STORE_S3_BUCKET=my-karakeep-bucket
ASSET_STORE_S3_ACCESS_KEY_ID=AKIA...
ASSET_STORE_S3_SECRET_ACCESS_KEY=...
Summary
- Centralized config in
packages/shared/config.tsvalidates all environment variables using Zod, creatingserverConfigfor backend use andclientConfigfor the browser. - Worker customization happens via
WORKERS_ENABLED_WORKERSandWORKERS_DISABLED_WORKERS, parsed byapps/workers/index.tsto selectively start background processes. - AI features are controlled by provider keys (
OPENAI_API_KEYorOLLAMA_BASE_URL) and fine-tuned withINFERENCE_*variables; custom prompts extend the base prompt via thecustomPromptsdatabase table. - Asset storage switches from local filesystem to S3 automatically when
ASSET_STORE_S3_ENDPOINTis defined. - Plugin system allows dropping TypeScript modules into
packages/plugins/to register custom tRPC routes or background jobs without modifying core code.
Frequently Asked Questions
Where is Karakeep configuration stored?
All configuration is centralized in packages/shared/config.ts, which parses process.env at runtime using Zod schemas. This file exports serverConfig for backend services and clientConfig for browser-safe settings. You customize behavior by setting environment variables; no YAML files or database configuration tables are required for core features.
Can I run Karakeep without AI features?
Yes. To disable AI entirely, either omit OPENAI_API_KEY and OLLAMA_BASE_URL from your environment, or explicitly disable the inference worker by adding inference to WORKERS_DISABLED_WORKERS. Without these variables, the inference worker becomes a no-op and no automatic tagging occurs, though you can still manually tag bookmarks.
How do I add a custom plugin to Karakeep?
Create a new directory under packages/plugins/ containing a TypeScript module that exports a default setup function receiving PluginContext. This function can register tRPC routers or database hooks. Import the plugin conditionally in apps/web/server.ts based on an environment flag (e.g., MY_PLUGIN_ENABLED=true). The plugin pattern follows the same structure as the built-in vectorstore-meilisearch and ratelimit-redis packages.
What is the difference between serverConfig and clientConfig?
serverConfig contains sensitive values like database URLs, API keys, and S3 credentials used by the backend workers and API. clientConfig is a filtered subset containing only public information like the application URL and demo-mode flags, which is serialized to the browser for the Next.js frontend. Both are generated from the same environment variables in packages/shared/config.ts, but clientConfig explicitly excludes secret fields to prevent leakage to the client bundle.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →