How the Health Monitoring System Works in Twenty CRM: NestJS Terminus Implementation
Twenty CRM implements a two-layer health monitoring system using NestJS and @nestjs/terminus that exposes a public /healthz endpoint for load balancers while providing rich diagnostics for PostgreSQL, Redis, and background workers through an admin panel service.
Twenty CRM, an open-source customer relationship management platform, relies on a robust health monitoring system to ensure high availability of its infrastructure. Built with NestJS and the @nestjs/terminus library, this system combines a lightweight public health check for container orchestration with deep diagnostic capabilities that track database performance, cache status, and queue metrics in real time.
Architecture of the Health Monitoring System
The twentyhq/twenty repository organizes health checks into two complementary layers. The public layer provides a simple up/down signal for Kubernetes probes and load balancers. The admin layer delivers detailed telemetry across five core indicators: database, Redis, worker processes, connected accounts, and the application itself.
Both layers leverage Terminus health indicators that return standardized { status, details } payloads. The admin panel enriches these results with historic state cached in an in-memory HealthStateManager, ensuring the UI displays useful data even during transient outages.
Public Health Endpoint for Load Balancers
The public health endpoint is designed for external orchestrators that need a fast, unauthenticated liveness check.
Module Registration
In packages/twenty-server/src/engine/core-modules/health/health.module.ts, the system imports TerminusModule and registers the controller:
// packages/twenty-server/src/engine/core-modules/health/health.module.ts
import { Module } from '@nestjs/common';
import { TerminusModule } from '@nestjs/terminus';
import { HealthController } from 'src/engine/core-modules/health/controllers/health.controller';
@Module({
imports: [TerminusModule],
controllers: [HealthController],
})
export class HealthModule {}
The /healthz Controller
The HealthController in packages/twenty-server/src/engine/core-modules/health/controllers/health.controller.ts exposes a single GET endpoint guarded by PublicEndpointGuard and NoPermissionGuard, making it reachable without authentication:
// packages/twenty-server/src/engine/core-modules/health/controllers/health.controller.ts
@Controller('healthz')
export class HealthController {
constructor(private readonly health: HealthCheckService) {}
@Get()
@UseGuards(PublicEndpointGuard, NoPermissionGuard)
@HealthCheck()
check() {
return this.health.check([]);
}
}
When the NestJS process is alive, the endpoint returns { "status": "ok" }. This minimal response prevents load balancers from routing traffic to unhealthy instances without exposing sensitive system details.
Admin Panel Health Service and Indicators
The AdminPanelHealthService aggregates multiple health indicators into a unified system status. Located in packages/twenty-server/src/engine/core-modules/admin-panel/admin-panel-health.service.ts, this service orchestrates five specialized indicators mapped by HealthIndicatorId.
// packages/twenty-server/src/engine/core-modules/admin-panel/admin-panel-health.service.ts
@Injectable()
export class AdminPanelHealthService {
private readonly healthIndicators = {
[HealthIndicatorId.database]: this.databaseHealth,
[HealthIndicatorId.redis]: this.redisHealth,
[HealthIndicatorId.worker]: this.workerHealth,
[HealthIndicatorId.connectedAccount]: this.connectedAccountHealth,
[HealthIndicatorId.app]: this.appHealth,
};
async getSystemHealthStatus(): Promise<SystemHealthDTO> {
const [
databaseResult,
redisResult,
workerResult,
accountSyncResult,
appResult,
] = await Promise.allSettled([
this.databaseHealth.isHealthy(),
this.redisHealth.isHealthy(),
this.workerHealth.isHealthy(),
this.connectedAccountHealth.isHealthy(),
this.appHealth.isHealthy(),
]);
return {
services: [
{ ...HEALTH_INDICATORS[HealthIndicatorId.database],
status: this.getServiceStatus(databaseResult, HealthIndicatorId.database).status },
// ... additional services
],
};
}
}
PostgreSQL Health Monitoring
The DatabaseHealthIndicator in packages/twenty-server/src/engine/core-modules/admin-panel/indicators/database.health.ts executes deep diagnostics against PostgreSQL. It queries version, active connections, max connections, uptime, database size, table statistics, cache hit ratio, deadlocks, and slow queries.
// packages/twenty-server/src/engine/core-modules/admin-panel/indicators/database.health.ts
async isHealthy(): Promise<HealthIndicatorResult> {
const indicator = this.healthIndicatorService.check('database');
try {
const [
[versionResult],
[activeConnections],
[maxConnections],
[uptime],
[databaseSize],
tableStats,
[cacheHitRatio],
[deadlocks],
[slowQueries],
] = await withHealthCheckTimeout(
Promise.all([
this.dataSource.query('SELECT version()'),
this.dataSource.query('SELECT count(*) as count FROM pg_stat_activity'),
this.dataSource.query('SHOW max_connections'),
// ... additional queries
]),
HEALTH_ERROR_MESSAGES.DATABASE_TIMEOUT,
);
const details = {
system: {
timestamp: new Date().toISOString(),
version: versionResult.version,
uptime: Math.round(uptime.uptime / 3600) + ' hours',
},
connections: { active: activeConnections.count, max: maxConnections.max_connections },
databaseSize: databaseSize.size,
performance: { cacheHitRatio, deadlocks, slowQueries },
top10Tables: tableStats,
};
this.stateManager.updateState(details);
return indicator.up({ details });
} catch (error) {
const stateWithAge = this.stateManager.getStateWithAge();
return indicator.down({
message: error.message === HEALTH_ERROR_MESSAGES.DATABASE_TIMEOUT
? HEALTH_ERROR_MESSAGES.DATABASE_TIMEOUT
: HEALTH_ERROR_MESSAGES.DATABASE_CONNECTION_FAILED,
details: { system: { timestamp: new Date().toISOString() }, stateHistory: stateWithAge },
});
}
}
Redis Health Monitoring
Similarly, the RedisHealthIndicator in packages/twenty-server/src/engine/core-modules/admin-panel/indicators/redis.health.ts collects comprehensive telemetry from the Redis instance. It gathers general info, memory statistics, client connections, and replication status.
// packages/twenty-server/src/engine/core-modules/admin-panel/indicators/redis.health.ts
async isHealthy(): Promise<HealthIndicatorResult> {
const indicator = this.healthIndicatorService.check('redis');
try {
const [info, memory, clients, stats] = await withHealthCheckTimeout(
Promise.all([
this.redisClient.getClient().info(),
this.redisClient.getClient().info('memory'),
this.redisClient.getClient().info('clients'),
this.redisClient.getClient().info('stats'),
]),
HEALTH_ERROR_MESSAGES.REDIS_TIMEOUT,
);
const parseInfo = (info: string) => { /* parsing logic */ };
const details = {
system: { timestamp: new Date().toISOString(), version: parseInfo(info).redis_version },
memory: { used: parseInfo(memory).used_memory_human, peak: parseInfo(memory).used_memory_peak_human },
connections: { connected: parseInfo(clients).connected_clients },
performance: { opsPerSecond: parseInfo(stats).instantaneous_ops_per_sec },
replication: { role: parseInfo(info).role, connectedSlaves: parseInfo(info).connected_slaves },
};
this.stateManager.updateState(details);
return indicator.up({ details });
} catch (error) {
const stateWithAge = this.stateManager.getStateWithAge();
return indicator.down({
message: error.message === HEALTH_ERROR_MESSAGES.REDIS_TIMEOUT
? HEALTH_ERROR_MESSAGES.REDIS_TIMEOUT
: HEALTH_ERROR_MESSAGES.REDIS_CONNECTION_FAILED,
details: { system: { timestamp: new Date().toISOString() }, stateHistory: stateWithAge },
});
}
}
Background Workers and Connected Accounts
The system also monitors WorkerHealthIndicator for background job processing and ConnectedAccountHealthIndicator for external account synchronization status. The AppHealthIndicator provides a lightweight check of the NestJS application context itself. All indicators follow the same pattern: execute health checks within a timeout wrapper, update the state manager on success, or return cached historic data on failure.
State Management and Resilience
Every health indicator utilizes the HealthStateManager utility to cache the last successful check result. When a database query or Redis command times out, the indicator returns a "down" status but includes the stateHistory from the cache. This design ensures the admin panel displays meaningful metrics—like connection counts or memory usage from seconds ago—rather than empty error states during brief network interruptions.
The withHealthCheckTimeout wrapper prevents hanging health checks from blocking the entire diagnostics request. Each indicator receives a configurable timeout constant (e.g., HEALTH_ERROR_MESSAGES.DATABASE_TIMEOUT), ensuring Promise.allSettled resolves predictably in AdminPanelHealthService.getSystemHealthStatus().
Queue Metrics for Background Jobs
Beyond binary up/down states, the health monitoring system exposes granular metrics for BullMQ message queues. The AdminPanelHealthService.getQueueMetrics() method instantiates BullMQ Queue objects using the Redis connection and fetches completed and failed job counts over configurable time ranges.
// packages/twenty-server/src/engine/core-modules/admin-panel/admin-panel-health.service.ts
async getQueueMetrics(queueName: MessageQueue, timeRange = QueueMetricsTimeRange.OneDay) {
const redis = this.redisClient.getQueueClient();
const queue = new Queue(queueName, { connection: redis });
try {
const { pointsNeeded, samplingFactor } = this.getPointsConfiguration(timeRange);
const queueDetails = await this.workerHealth.getQueueDetails(queueName, { pointsNeeded });
// Data transformation for graph-ready DTOs
return this.transformMetricsForGraph(completedMetrics, failedMetrics, timeRange, queueName, queueDetails);
} finally {
await queue.close();
}
}
This enables the admin dashboard to render time-series graphs of job throughput, helping operators identify processing bottlenecks in real time.
Practical Implementation Examples
Checking Liveness with cURL
Verify the basic health status from the command line:
curl -s http://localhost:3000/healthz | jq .
Expected output when the server is healthy:
{
"status": "ok"
}
Fetching Detailed Health from Another Service
Integrate the admin health data into external monitoring tools using the HTTP service:
import { HttpService } from '@nestjs/axios';
import { firstValueFrom } from 'rxjs';
async function fetchSystemHealth(http: HttpService) {
const response = await firstValueFrom(
http.get<SystemHealthDTO>('http://localhost:3000/admin/health/system')
);
return response.data; // Contains services array with status and details
}
The route ultimately invokes AdminPanelHealthService.getSystemHealthStatus(), returning a structured payload of all five health indicators.
Direct Service Injection
Access the health monitoring system programmatically within the NestJS dependency graph:
@Injectable()
export class StartupChecker {
constructor(private readonly healthService: AdminPanelHealthService) {}
async logHealthOnStart() {
const status = await this.healthService.getSystemHealthStatus();
console.log('System health at startup:', status);
}
}
Summary
- Two-layer architecture: Twenty CRM uses a public
/healthzendpoint for load balancers and a private admin service for deep diagnostics. - Five core indicators: The system monitors PostgreSQL, Redis, background workers, connected accounts, and the application context via
AdminPanelHealthService. - Resilient state management: The
HealthStateManagercaches the last successful health check, ensuring the UI displays historic data during transient failures. - Timeout protection: Every indicator uses
withHealthCheckTimeoutto prevent hanging requests from blocking the health dashboard. - Queue telemetry: The service exposes BullMQ metrics for background job monitoring, enabling throughput analysis by queue name and time range.
Frequently Asked Questions
What endpoint does Twenty CRM expose for Kubernetes liveness probes?
Twenty CRM exposes GET /healthz via the HealthController in packages/twenty-server/src/engine/core-modules/health/controllers/health.controller.ts. This endpoint is guarded by PublicEndpointGuard and returns { "status": "ok" } when the NestJS process is alive, making it ideal for Kubernetes liveness and readiness probes without requiring authentication.
How does Twenty CRM handle transient failures in health checks?
When a health indicator encounters a timeout or connection error, it catches the exception and returns a "down" status while appending cached data from the HealthStateManager. This utility stores the last successful check result with a timestamp, ensuring the admin panel displays recent metrics like connection counts or memory usage even when the service is temporarily unavailable.
Which services are monitored in the Twenty CRM admin panel health dashboard?
The dashboard tracks five services defined in HealthIndicatorId: database (PostgreSQL), redis (cache layer), worker (background job processors), connectedAccount (external account synchronization), and app (NestJS runtime). Each has a dedicated indicator class that queries specific metrics like table statistics, memory usage, or queue depths.
Can I access the health monitoring data programmatically from another service?
Yes. You can either inject AdminPanelHealthService directly into your NestJS providers to call getSystemHealthStatus() or getQueueMetrics(), or you can consume the REST endpoint at /admin/health/system which returns a SystemHealthDTO containing the status and detailed metrics for all monitored services.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →