What Are the Main Functionalities of IPED? A Complete Technical Guide

IPED (Indexador & Processador de Evidências Digitais) is a full-stack digital forensics platform that processes evidence images, extracts metadata, performs OCR and face recognition, indexes content for search, and provides graph analysis of communication data.

The open-source IPED project (sepinf-inc/IPED) delivers an end-to-end digital investigation suite written in Java. Understanding the main functionalities of IPED helps forensic examiners leverage its modular pipeline for evidence processing, from hash deduplication to AI-powered entity recognition.

Three-Layer Architecture

IPED organizes its main functionalities into three tightly coupled architectural layers that handle everything from command-line input to advanced graph queries.

Entry and Configuration Layer

The application bootstrap resides in iped-app/src/main/java/iped/app/processing/Main.java. This Main class parses command-line arguments through CmdLineArgsImpl, validates output directories, and loads processing profiles. It prepares the runtime environment by copying libraries and UI resources into the case folder before handing control to the processing engine.

Processing Engine Layer

At the heart of IPED sits iped-engine/src/main/java/iped/engine/core/Manager.java. The Manager class orchestrates the case lifecycle using a producer-consumer pattern. Two ItemProducer threads first count items for progress estimation, then produce Item objects consumed by a pool of Worker threads defined in Worker.java. Each worker executes the configurable task chain (hashing, parsing, OCR) defined in IndexTaskConfig.

Data Access and Services Layer

Post-processing functionality lives in iped-engine/src/main/java/iped/engine/search/IPEDSearcher.java, which wraps Lucene queries for programmatic case interrogation. The PreviewRepositoryManager handles thumbnail generation, while GraphService manages the embedded Neo4j-compatible database for communication analysis. A thin Web API provides HTTP endpoints for remote queries.

End-to-End Processing Workflow

The main functionalities of IPED follow a deterministic seven-stage pipeline:

  1. Argument parsing – Main builds a CmdLineArgsImpl instance, validates the output folder, and resolves the selected profile.
  2. Output preparation – Manager.prepareOutputFolder() copies runtime dependencies and creates the data/ subdirectory.
  3. Index initialization – An IndexWriter instantiates with a custom AppAnalyzer for full-text and metadata indexing.
  4. Producer-consumer processing – ItemProducer threads feed evidence into Worker pools that execute the task chain (hashing, container expansion, OCR, face recognition, NER).
  5. Periodic commits – Background threads flush the Lucene index and auxiliary storages (CSV export, graph, Elasticsearch).
  6. Post-processing – Manager removes empty tree nodes, filters keywords, updates image paths for portable cases, and optionally force-merges the index.
  7. Search and export – The case opens in the UI or accepts queries via IPEDSearcher, with ExportFileTask and ExportCSVTask generating reports.

Core Functional Blocks

Evidence Hashing and Deduplication

HashDB (iped/engine/hashdb/HashDB.java) computes MD5 and SHA-1 hashes, storing them in an SQLite database to automatically discard duplicate items during ingestion.

Recursive Container Expansion

ParsingReader (iped/engine/io/ParsingReader.java) detects archive formats (ZIP, 7z, ISO, E01, VHD) and streams their entries as independent Item objects for recursive processing.

Optical Character Recognition

OCRParser (iped/parsers/ocr/OCRParser.java) sends image streams to Tesseract 5 or cloud services (Azure/Google), storing extracted text in the index for full-text search.

FaceRecognitionTask (iped/app/resources/scripts/tasks/FaceRecognitionTask.java) generates face embeddings using dlib or InsightFace, building a searchable vector index for identifying persons across evidence.

Named Entity Recognition

NERTask leverages Stanford CoreNLP models via the iped/parsers/ner/ package to automatically tag persons, organizations, and locations within extracted text.

Graph Analysis

GraphService and GraphTask store communication links (calls, emails, instant messages) in an embedded graph database, enabling queries like "who called whom" and relationship visualization.

Web API and Remote Access

The iped-webapi module exposes REST endpoints (/search, /metadata, /thumb) that return JSON, raw content, or thumbnails for integration with remote tools.

Export and Reporting

ExportFileTask, ExportCSVTask, and HTMLReportGenerator produce portable case bundles, CSV metadata indexes, and interactive HTML timelines for court presentation.

Processing Profiles

Configuration files in conf/profiles/ define pre-defined task sets for forensic, pedophile-content (CSAM), triage, fast-preview, and blind (auto-extraction) modes, selectable via the --profile command-line argument.

Extensibility Framework

Developers can inject custom logic via JavaScript (scripts/) or Python (iped-parsers-impl/src/main/python/) without modifying core engine code, using the Jython bridge for parser integration.

Code Examples

Running a Case from the Command Line

java -jar iped.jar \
    -i /evidence/image.E01 \
    -o /cases/mycase \
    -p forensic \
    --keywords /path/to/keywords.txt \
    --log /cases/mycase/log.txt
  • -i = input evidence (disk image, directory, or UFED report).
  • -o = output folder for the portable case.
  • -p = processing profile from the conf/profiles/ directory.

This command invokes Main.main(), which parses arguments and delegates to Manager.process().

Searching a Case Programmatically

import iped.engine.search.IPEDSearcher;
import iped.engine.search.SearchResult;
import iped.engine.IPEDSource;
import java.io.File;

public class SimpleSearch {
    public static void main(String[] args) throws Exception {
        File caseRoot = new File("/cases/mycase");

        try (IPEDSource caseDb = new IPEDSource(caseRoot)) {
            IPEDSearcher searcher = new IPEDSearcher(caseDb,
                    "mime_type:\"image/png\" AND content:\"confidential\"");

            SearchResult result = searcher.search();

            System.out.println("Found " + result.getLength() + " items:");
            for (int i = 0; i < result.getLength(); i++) {
                System.out.println(" - " + caseDb.getItemByID(result.getId(i)).getPath());
            }
        }
    }
}

IPEDSource opens the case in read-only mode, IPEDSearcher builds Lucene queries, and SearchResult enumerates matching item IDs.

Adding a Custom Python Parser

Create myparser.py in iped-parsers-impl/src/main/python/:

from iped.parsers import AbstractParser
from iped.io import SeekableInputStream

class MyParser(AbstractParser):
    def getSupportedMimeTypes(self):
        return ["text/plain"]

    def parse(self, stream: SeekableInputStream, metadata, contentHandler):
        text = stream.read().decode('utf-8')
        if "sensitive" in text.lower():
            contentHandler.characters("FOUND SENSITIVE")

Register the parser in iped-parsers-impl/src/main/resources/META-INF/services/iped.parsers.AbstractParser by adding myparser.MyParser. The ParsingReader automatically loads this via Jython during processing.

Querying the Web API

curl "http://localhost:8080/api/search?q=credit+card+number"

Returns a JSON array with item IDs, paths, and preview URLs. Implementation resides in iped-webapi/src/main/java/iped/webapi/SearchResource.java.

Key Files and Entry Points

Module File Role
Application entry iped-app/src/main/java/iped/app/processing/Main.java CLI entry point and configuration loader.
Core engine iped-engine/src/main/java/iped/engine/core/Manager.java Orchestrates producers, workers, and index management.
Task configuration iped-engine/src/main/java/iped/engine/config/IndexTaskConfig.java Defines the processing task chain for workers.
Search API iped-engine/src/main/java/iped/engine/search/IPEDSearcher.java High-level Lucene query wrapper for case data.
Graph service iped-engine/src/main/java/iped/engine/graph/GraphService.java Manages embedded graph database for communication links.
Export utilities iped-engine/src/main/java/iped/engine/task/ExportCSVTask.java Generates CSV metadata reports.
OCR parser iped/parsers/ocr/OCRParser.java Tesseract integration for text extraction.
Python bridge iped-parsers-impl/src/main/python/ Jython integration for custom parsers.
Web API iped-webapi/src/main/java/iped/webapi/ REST endpoints for remote access.
Profiles conf/profiles/*.properties Pre-defined forensic processing profiles.

Summary

  • IPED is a Java-based digital forensics platform combining indexing, parsing, and analysis in a single pipeline.
  • The three-layer architecture separates configuration (Main), processing (Manager/Worker), and data services (IPEDSearcher/GraphService).
  • Core functionalities include hash deduplication, recursive container expansion, OCR, face recognition, NER, graph analysis, and multi-format export.
  • Extensibility via Python and JavaScript allows custom parsers without core modifications.
  • Web API and portable case formats enable both automated remote queries and courtroom presentation.

Frequently Asked Questions

What evidence formats does IPED support?

IPED processes raw disk images (E01, VHD, ISO), directory trees, ZIP/7z archives, and mobile extractions (UFED reports). The ParsingReader class in iped/engine/io/ParsingReader.java handles recursive extraction of nested containers, treating each file as an independent Item for metadata extraction and indexing.

How does IPED handle duplicate files during processing?

The HashDB class (iped/engine/hashdb/HashDB.java) computes MD5 and SHA-1 hashes for every item, storing them in an SQLite database. When duplicates are detected, IPED can skip re-processing or store references only, significantly reducing index size and processing time for repetitive evidence.

Can IPED be integrated into automated forensic workflows?

Yes. The iped-webapi module exposes REST endpoints at http://localhost:8080/api/ for searching metadata, retrieving thumbnails, and exporting results programmatically. Additionally, the command-line interface accepts profiles (-p triage, -p forensic) that pre-configure task chains, enabling scripted batch processing of multiple evidence sources.

Is it possible to add custom analysis tasks to IPED without modifying the core code?

Absolutely. IPED supports JavaScript task scripts in iped/app/resources/scripts/tasks/ and Python parsers in iped-parsers-impl/src/main/python/. By implementing AbstractParser and registering the class in the service provider configuration, users inject custom logic into the Worker task chain without recompiling the engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →