How to Configure Sampling Mode for Realtime Log Monitoring in LogSentinelAI
Set REALTIME_ONLY_SAMPLING_MODE=true in your environment file to force sampling mode, or adjust REALTIME_SAMPLING_THRESHOLD to automatically switch to sampling when pending log lines exceed the configured limit.
LogSentinelAI's real-time monitor processes log files in continuous chunks and supports two distinct modes: full retention and sampling mode. This guide covers how to configure sampling mode for realtime log monitoring using environment variables defined in src/logsentinelai/core/config.py and processed by the monitor in src/logsentinelai/core/monitoring.py.
Understanding Processing Modes
The real-time monitor operates in one of two modes depending on your configuration:
- Full Mode: Retains all incoming log lines until the configured chunk size is reached. This consumes more memory but preserves complete context.
- Sampling Mode: Retains only the most recent N lines (up to
chunk_size), discarding older lines to conserve memory during high-volume periods.
According to the source code in src/logsentinelai/core/monitoring.py, the monitor evaluates the following logic to determine which mode to use:
self.only_sampling_mode = self.realtime_config["only_sampling_mode"]
self.sampling_threshold = self.realtime_config["sampling_threshold"]
if self.only_sampling_mode:
should_sample = True # always sampling
elif len(self.pending_lines) > self.sampling_threshold:
should_sample = True # auto-switch when threshold exceeded
else:
should_sample = False # full processing
Key Environment Variables
Configuration is loaded at startup from .env or /etc/logsentinelai.config via the REALTIME_CONFIG dictionary in src/logsentinelai/core/config.py (lines 90-98). The relevant variables for sampling mode are:
REALTIME_ONLY_SAMPLING_MODE(default:false): When set totrue, forces the monitor to permanently operate in sampling mode.REALTIME_SAMPLING_THRESHOLD(default:100): In full mode, automatically switches to sampling when pending lines exceed this count.REALTIME_CHUNK_PENDING_TIMEOUT(default:1800): Forces processing after this many seconds (default 30 minutes) regardless of chunk size.REALTIME_MAX_LINES_PER_BATCH(default:50): Upper bound of lines processed per poll cycle.REALTIME_POLLING_INTERVAL(default:5): Seconds between file system checks.REALTIME_BUFFER_TIME(default:2): Buffer time in seconds for line aggregation.
Note that changes to these variables require a monitor restart to take effect, as the configuration is read once at startup via the apply_config() function.
Configuring Sampling Mode
Option 1: Force Permanent Sampling Mode
To always run in sampling mode regardless of log volume:
- Edit your
.envfile:
REALTIME_ONLY_SAMPLING_MODE=true
REALTIME_SAMPLING_THRESHOLD=100 # Ignored when only_sampling_mode is true
REALTIME_CHUNK_PENDING_TIMEOUT=300
REALTIME_MAX_LINES_PER_BATCH=50
- Restart the monitor. The initialization banner in
monitoring.py::_print_initialization_info()will displayMODE: SAMPLING-ONLY.
Option 2: Conditional Sampling with Threshold
To use full mode normally but switch to sampling during high-volume spikes:
REALTIME_ONLY_SAMPLING_MODE=false
REALTIME_SAMPLING_THRESHOLD=200 # Switch to sampling when >200 lines pending
REALTIME_CHUNK_PENDING_TIMEOUT=1800
Option 3: Temporary CLI Override
Run a one-off analysis with sampling enabled without editing configuration files:
REALTIME_ONLY_SAMPLING_MODE=true REALTIME_SAMPLING_THRESHOLD=300 \
logsentinelai linux-system --mode realtime
Programmatic Configuration
When using the Python API directly, you can override environment defaults by manipulating the configuration dictionary before instantiating the monitor:
from logsentinelai.core.monitoring import create_realtime_monitor
from logsentinelai.core.config import get_analysis_config
# Build base configuration
custom_cfg = get_analysis_config(
log_type="linux_system",
analysis_mode="realtime",
remote_mode="local",
chunk_size=20,
)
# Force sampling mode programmatically
custom_cfg["realtime_config"]["only_sampling_mode"] = True
custom_cfg["realtime_config"]["sampling_threshold"] = 9999
# Create monitor with modified config
monitor = create_realtime_monitor(
log_type="linux_system",
remote_mode="local",
ssh_config=None,
remote_log_path=None,
)
# Process chunks
for chunk in monitor.get_new_log_chunks():
print("Processing chunk:", chunk)
Verifying Your Configuration
Upon startup, the monitor prints an initialization banner via _print_initialization_info() in src/logsentinelai/core/monitoring.py. Check the output for the active mode:
MODE: SAMPLING-ONLY # If only_sampling_mode=true
MODE: FULL # If processing all lines
If you do not see the expected mode, verify that your environment file is in the project root (for .env) or at /etc/logsentinelai.config, and ensure the monitor process was restarted after configuration changes.
Summary
- Sampling mode in LogSentinelAI retains only the most recent lines up to
chunk_size, while full mode retains all lines until the chunk is complete. - Set
REALTIME_ONLY_SAMPLING_MODE=trueto force sampling permanently, or tuneREALTIME_SAMPLING_THRESHOLDto trigger sampling automatically when pending lines exceed the limit. - Configuration is loaded from environment variables via
src/logsentinelai/core/config.pyand evaluated insrc/logsentinelai/core/monitoring.pyat startup. - Changes to sampling configuration require a monitor restart to take effect.
- Verify the active mode by checking the initialization banner printed by
monitoring.py::_print_initialization_info().
Frequently Asked Questions
What is the difference between full mode and sampling mode in LogSentinelAI?
In full mode, the real-time monitor accumulates all incoming log lines until the chunk size is reached, preserving complete context but consuming more memory. In sampling mode, only the most recent N lines (where N equals the configured chunk size) are retained, and older lines are discarded to reduce memory footprint during high-volume logging.
How do I know if sampling mode is currently active?
The monitor prints an initialization banner at startup via the _print_initialization_info() function in src/logsentinelai/core/monitoring.py. Look for MODE: SAMPLING-ONLY to confirm forced sampling mode, or MODE: FULL to confirm standard processing. If you see SAMPLING-ONLY while running with only_sampling_mode=false, it indicates the automatic threshold was triggered.
Can I change sampling settings without restarting the monitor?
No. The REALTIME_CONFIG dictionary is built once at startup in src/logsentinelai/core/config.py via the apply_config() function. Any changes to REALTIME_ONLY_SAMPLING_MODE, REALTIME_SAMPLING_THRESHOLD, or related variables require stopping and restarting the monitor process to reload the configuration from the environment.
What happens when REALTIME_SAMPLING_THRESHOLD is exceeded in full mode?
When only_sampling_mode is false and the number of pending lines exceeds REALTIME_SAMPLING_THRESHOLD (default 100), the monitor automatically switches to sampling mode for that chunk. As implemented in src/logsentinelai/core/monitoring.py, the logic evaluates len(self.pending_lines) > self.sampling_threshold and sets should_sample = True, causing the monitor to retain only the most recent lines up to the chunk size instead of the full backlog.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →