Security Best Practices for Karakeep: Essential Hardening for Self-Hosted Deployments
TLDR: Secure your Karakeep installation by restricting API access to authenticated users via NextAuth, isolating the Puppeteer-based crawler in a separate container or behind a proxy, and enabling the built-in SSRF filter that blocks internal IP ranges in packages/api/src/crawler.ts.
Karakeep is a self-hosted "read-it-later" bookmarking service built with Next.js, tRPC, and Drizzle ORM. Because the application uses a real browser (Puppeteer) to crawl bookmarked URLs, self-hosted instances face unique Server-Side Request Forgery (SSRF) risks if untrusted users can submit links. The following practices mitigate these vulnerabilities and harden your deployment against unauthorized access.
Limit Access to Trusted Users
Untrusted users who can trigger crawls may probe your internal network or expose your server's IP address to malicious sites. Restricting access is the first line of defense.
Enforce Authentication via NextAuth
Configure NextAuth in src/server/auth.ts to ensure only authenticated users can access the tRPC API. The auth configuration should validate user sessions before allowing bookmark creation that triggers crawling.
// Example: src/server/auth.ts
import NextAuth from "next-auth";
import GoogleProvider from "next-auth/providers/google";
export const authOptions = {
providers: [
GoogleProvider({
clientId: process.env.GOOGLE_ID,
clientSecret: process.env.GOOGLE_SECRET
})
],
callbacks: {
async session({ session, token }) {
session.user.role = token.role ?? "user";
return session;
}
}
};
export default NextAuth(authOptions);
Implement Role-Based Checks in tRPC Routers
All sensitive operations that invoke the crawler should verify user roles within the tRPC routers. Files such as packages/trpc/routers/admin.ts and packages/trpc/routers/users.ts contain logic that validates session.user.role before executing admin-level actions or triggering fetches.
Harden the Crawler Against SSRF Attacks
The Puppeteer-based crawler in packages/api/src/crawler.ts executes JavaScript from arbitrary URLs, making SSRF protection critical.
Enable the Built-in SSRF Filter
Karakeep includes a basic SSRF filter that blocks requests to internal IP ranges before the browser initiates connections. This logic typically resides in packages/api/src/crawler.ts and utilizes an isInternalIP check to validate targets.
The filter is enabled by default, but you should verify its presence by ensuring the crawler validates URLs against internal ranges before launching Puppeteer.
Customize Internal IP Blocklists
Modify the INTERNAL_IP_RANGES constant in packages/api/src/ssrf.ts to include any additional private CIDR blocks used by your environment:
// packages/api/src/ssrf.ts
const INTERNAL_IP_RANGES = [
"127.0.0.0/8",
"10.0.0.0/8",
"172.16.0.0/12",
"192.168.0.0/16",
// Add your private ranges here
];
export function isBlocked(host: string): boolean {
// DNS-resolve then check IP against INTERNAL_IP_RANGES
// Returns true → block request
}
Update this array if your infrastructure uses non-standard private subnets.
Verify SSRF Protections
Test your configuration by attempting to crawl a known internal endpoint (e.g., http://127.0.0.1). The test suite in packages/trpc/routers/webhooks.test.ts expects a 403 Forbidden response for internal URLs, confirming the filter functions correctly.
Isolate the Crawler Environment
Network isolation provides the most reliable SSRF protection by ensuring the crawler cannot reach internal services even if filters fail.
Deploy Container-Level Isolation
Run the headless browser in a separate Docker container with a restricted network that permits only outbound internet access. The repository provides an example configuration in docker-compose.crawler.yml (located in the docs folder) that demonstrates proper isolation.
Route Traffic Through a Proxy
Force all crawler traffic through an egress proxy by setting the CRAWLER_PROXY_URL environment variable:
# .env (or set in your orchestration platform)
CRAWLER_PROXY_URL="http://crawler-proxy.local:3128"
Then configure the crawler to use this proxy in packages/api/src/crawler.ts:
// packages/api/src/crawler.ts
import { createProxyAgent } from "proxy-agent";
const proxy = process.env.CRAWLER_PROXY_URL
? createProxyAgent(process.env.CRAWLER_PROXY_URL)
: undefined;
export async function fetchPage(url: string) {
const browser = await puppeteer.launch({
args: proxy ? [`--proxy-server=${proxy}`] : []
});
// ...perform navigation...
}
This configuration ensures that even if a malicious URL bypasses other checks, the request routes through a controlled egress point.
Consider Hosted Browser Services
For maximum isolation, consider using a SaaS browser solution like browserless.io instead of local Puppeteer instances. This approach removes the browser attack surface from your infrastructure entirely.
Secure the Database Layer
Karakeep uses Drizzle ORM (located in packages/db) to generate parameterized SQL statements, which protects against SQL injection attacks. Ensure the database is only reachable from the backend container(s) and not exposed to the public internet.
Rotate credentials regularly using the template in .env.sample as a reference for secure configuration. Never commit actual credentials to version control.
Enforce HTTPS and Network Policies
Terminate all external traffic (web UI and API) using TLS at a reverse proxy such as NGINX or Cloudflare. Ensure internal services listen only on the loopback interface (127.0.0.1) to prevent direct external access.
Maintain Dependency Hygiene
Regularly audit dependencies using pnpm audit to identify known vulnerabilities. Additionally, upgrade the headless Chromium version used by Puppeteer frequently to patch browser-specific security flaws.
Report Vulnerabilities Responsibly
If you discover a security vulnerability, do not open a public GitHub issue. Instead, follow the disclosure policy outlined in SECURITY.md and contact the maintainers via the private GitHub vulnerability reporting feature or email security@karakeep.app.
Summary
- Restrict access to authenticated users via NextAuth and role checks in tRPC routers like
packages/trpc/routers/admin.ts. - Enable SSRF protection by verifying the
isInternalIPlogic inpackages/api/src/crawler.tsand customizingINTERNAL_IP_RANGESinpackages/api/src/ssrf.ts. - Isolate the crawler using container network isolation (
docker-compose.crawler.yml) or by settingCRAWLER_PROXY_URLto route traffic through a controlled proxy. - Protect the database by ensuring Drizzle ORM queries remain parameterized and limiting network access to backend containers.
- Enforce HTTPS at the reverse proxy level and keep internal services bound to localhost.
- Audit dependencies regularly with
pnpm auditand update Chromium versions. - Disclose vulnerabilities privately via
SECURITY.mdchannels rather than public issues.
Frequently Asked Questions
How does Karakeep prevent SSRF attacks when crawling URLs?
According to the Karakeep source code, SSRF prevention relies on a multi-layered approach. The crawler in packages/api/src/crawler.ts implements an isInternalIP check that blocks requests to private IP ranges defined in INTERNAL_IP_RANGES before Puppeteer initiates connections. Additionally, the security documentation recommends deploying the crawler in an isolated container or behind a proxy via the CRAWLER_PROXY_URL environment variable to prevent access to internal network resources even if URL validation is bypassed.
Is it safe to expose Karakeep to the public internet?
Public exposure is safe only when you enforce authentication via NextAuth, restrict sensitive actions through role-based checks in tRPC routers (such as packages/trpc/routers/users.ts), and deploy behind a reverse proxy that handles TLS termination. You should also isolate the crawler component in a separate network segment to prevent SSRF attacks against your internal infrastructure.
What should I do if I find a security vulnerability in Karakeep?
Do not create a public GitHub issue. Instead, refer to the SECURITY.md file in the repository root and report the vulnerability privately using the GitHub vulnerability reporting feature or by emailing security@karakeep.app. The maintainers coordinate disclosure to ensure fixes are available before details become public.
How does Karakeep protect against SQL injection?
Karakeep uses Drizzle ORM for all database interactions, which automatically generates parameterized queries that separate SQL logic from data. The schema definitions in packages/db ensure type safety and prevent injection attacks. Always ensure your database credentials are stored securely in environment variables (referencing .env.sample) and never committed to version control.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →