# What Are the Main Functionalities of IPED? A Complete Technical Guide

> Explore IPED's main functionalities. This digital forensics platform processes evidence, extracts metadata, performs OCR and face recognition, indexes content, and offers graph analysis for communication data.

- Repository: [Serviço de Perícias em Informática/IPED](https://github.com/sepinf-inc/IPED)
- Tags: how-to-guide
- Published: 2026-03-11

---

**IPED (Indexador & Processador de Evidências Digitais) is a full-stack digital forensics platform that processes evidence images, extracts metadata, performs OCR and face recognition, indexes content for search, and provides graph analysis of communication data.**

The open-source IPED project ([sepinf-inc/IPED](https://github.com/sepinf-inc/IPED)) delivers an end-to-end digital investigation suite written in Java. Understanding the main functionalities of IPED helps forensic examiners leverage its modular pipeline for evidence processing, from hash deduplication to AI-powered entity recognition.

## Three-Layer Architecture

IPED organizes its main functionalities into three tightly coupled architectural layers that handle everything from command-line input to advanced graph queries.

### Entry and Configuration Layer

The application bootstrap resides in [`iped-app/src/main/java/iped/app/processing/Main.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-app/src/main/java/iped/app/processing/Main.java). This `Main` class parses command-line arguments through `CmdLineArgsImpl`, validates output directories, and loads processing profiles. It prepares the runtime environment by copying libraries and UI resources into the case folder before handing control to the processing engine.

### Processing Engine Layer

At the heart of IPED sits [`iped-engine/src/main/java/iped/engine/core/Manager.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-engine/src/main/java/iped/engine/core/Manager.java). The `Manager` class orchestrates the case lifecycle using a producer-consumer pattern. Two `ItemProducer` threads first count items for progress estimation, then produce `Item` objects consumed by a pool of `Worker` threads defined in [`Worker.java`](https://github.com/sepinf-inc/IPED/blob/main/Worker.java). Each worker executes the configurable task chain (hashing, parsing, OCR) defined in `IndexTaskConfig`.

### Data Access and Services Layer

Post-processing functionality lives in [`iped-engine/src/main/java/iped/engine/search/IPEDSearcher.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-engine/src/main/java/iped/engine/search/IPEDSearcher.java), which wraps Lucene queries for programmatic case interrogation. The `PreviewRepositoryManager` handles thumbnail generation, while `GraphService` manages the embedded Neo4j-compatible database for communication analysis. A thin Web API provides HTTP endpoints for remote queries.

## End-to-End Processing Workflow

The main functionalities of IPED follow a deterministic seven-stage pipeline:

1. **Argument parsing** – `Main` builds a `CmdLineArgsImpl` instance, validates the output folder, and resolves the selected profile.
2. **Output preparation** – `Manager.prepareOutputFolder()` copies runtime dependencies and creates the `data/` subdirectory.
3. **Index initialization** – An `IndexWriter` instantiates with a custom `AppAnalyzer` for full-text and metadata indexing.
4. **Producer-consumer processing** – `ItemProducer` threads feed evidence into `Worker` pools that execute the task chain (hashing, container expansion, OCR, face recognition, NER).
5. **Periodic commits** – Background threads flush the Lucene index and auxiliary storages (CSV export, graph, Elasticsearch).
6. **Post-processing** – `Manager` removes empty tree nodes, filters keywords, updates image paths for portable cases, and optionally force-merges the index.
7. **Search and export** – The case opens in the UI or accepts queries via `IPEDSearcher`, with `ExportFileTask` and `ExportCSVTask` generating reports.

## Core Functional Blocks

### Evidence Hashing and Deduplication

`HashDB` ([`iped/engine/hashdb/HashDB.java`](https://github.com/sepinf-inc/IPED/blob/main/iped/engine/hashdb/HashDB.java)) computes MD5 and SHA-1 hashes, storing them in an SQLite database to automatically discard duplicate items during ingestion.

### Recursive Container Expansion

`ParsingReader` ([`iped/engine/io/ParsingReader.java`](https://github.com/sepinf-inc/IPED/blob/main/iped/engine/io/ParsingReader.java)) detects archive formats (ZIP, 7z, ISO, E01, VHD) and streams their entries as independent `Item` objects for recursive processing.

### Optical Character Recognition

`OCRParser` ([`iped/parsers/ocr/OCRParser.java`](https://github.com/sepinf-inc/IPED/blob/main/iped/parsers/ocr/OCRParser.java)) sends image streams to Tesseract 5 or cloud services (Azure/Google), storing extracted text in the index for full-text search.

### Face Recognition and Similarity Search

`FaceRecognitionTask` ([`iped/app/resources/scripts/tasks/FaceRecognitionTask.java`](https://github.com/sepinf-inc/IPED/blob/main/iped/app/resources/scripts/tasks/FaceRecognitionTask.java)) generates face embeddings using dlib or InsightFace, building a searchable vector index for identifying persons across evidence.

### Named Entity Recognition

`NERTask` leverages Stanford CoreNLP models via the `iped/parsers/ner/` package to automatically tag persons, organizations, and locations within extracted text.

### Graph Analysis

`GraphService` and `GraphTask` store communication links (calls, emails, instant messages) in an embedded graph database, enabling queries like "who called whom" and relationship visualization.

### Web API and Remote Access

The `iped-webapi` module exposes REST endpoints (`/search`, `/metadata`, `/thumb`) that return JSON, raw content, or thumbnails for integration with remote tools.

### Export and Reporting

`ExportFileTask`, `ExportCSVTask`, and `HTMLReportGenerator` produce portable case bundles, CSV metadata indexes, and interactive HTML timelines for court presentation.

### Processing Profiles

Configuration files in `conf/profiles/` define pre-defined task sets for **forensic**, **pedophile-content (CSAM)**, **triage**, **fast-preview**, and **blind** (auto-extraction) modes, selectable via the `--profile` command-line argument.

### Extensibility Framework

Developers can inject custom logic via JavaScript (`scripts/`) or Python (`iped-parsers-impl/src/main/python/`) without modifying core engine code, using the Jython bridge for parser integration.

## Code Examples

### Running a Case from the Command Line

```bash
java -jar iped.jar \
    -i /evidence/image.E01 \
    -o /cases/mycase \
    -p forensic \
    --keywords /path/to/keywords.txt \
    --log /cases/mycase/log.txt

```

* `-i` = input evidence (disk image, directory, or UFED report).
* `-o` = output folder for the portable case.
* `-p` = processing profile from the `conf/profiles/` directory.

This command invokes `Main.main()`, which parses arguments and delegates to `Manager.process()`.

### Searching a Case Programmatically

```java
import iped.engine.search.IPEDSearcher;
import iped.engine.search.SearchResult;
import iped.engine.IPEDSource;
import java.io.File;

public class SimpleSearch {
    public static void main(String[] args) throws Exception {
        File caseRoot = new File("/cases/mycase");

        try (IPEDSource caseDb = new IPEDSource(caseRoot)) {
            IPEDSearcher searcher = new IPEDSearcher(caseDb,
                    "mime_type:\"image/png\" AND content:\"confidential\"");

            SearchResult result = searcher.search();

            System.out.println("Found " + result.getLength() + " items:");
            for (int i = 0; i < result.getLength(); i++) {
                System.out.println(" - " + caseDb.getItemByID(result.getId(i)).getPath());
            }
        }
    }
}

```

`IPEDSource` opens the case in read-only mode, `IPEDSearcher` builds Lucene queries, and `SearchResult` enumerates matching item IDs.

### Adding a Custom Python Parser

Create [`myparser.py`](https://github.com/sepinf-inc/IPED/blob/main/myparser.py) in `iped-parsers-impl/src/main/python/`:

```python
from iped.parsers import AbstractParser
from iped.io import SeekableInputStream

class MyParser(AbstractParser):
    def getSupportedMimeTypes(self):
        return ["text/plain"]

    def parse(self, stream: SeekableInputStream, metadata, contentHandler):
        text = stream.read().decode('utf-8')
        if "sensitive" in text.lower():
            contentHandler.characters("FOUND SENSITIVE")

```

Register the parser in `iped-parsers-impl/src/main/resources/META-INF/services/iped.parsers.AbstractParser` by adding `myparser.MyParser`. The `ParsingReader` automatically loads this via Jython during processing.

### Querying the Web API

```bash
curl "http://localhost:8080/api/search?q=credit+card+number"

```

Returns a JSON array with item IDs, paths, and preview URLs. Implementation resides in [`iped-webapi/src/main/java/iped/webapi/SearchResource.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-webapi/src/main/java/iped/webapi/SearchResource.java).

## Key Files and Entry Points

| Module | File | Role |
|--------|------|------|
| **Application entry** | [`iped-app/src/main/java/iped/app/processing/Main.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-app/src/main/java/iped/app/processing/Main.java) | CLI entry point and configuration loader. |
| **Core engine** | [`iped-engine/src/main/java/iped/engine/core/Manager.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-engine/src/main/java/iped/engine/core/Manager.java) | Orchestrates producers, workers, and index management. |
| **Task configuration** | [`iped-engine/src/main/java/iped/engine/config/IndexTaskConfig.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-engine/src/main/java/iped/engine/config/IndexTaskConfig.java) | Defines the processing task chain for workers. |
| **Search API** | [`iped-engine/src/main/java/iped/engine/search/IPEDSearcher.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-engine/src/main/java/iped/engine/search/IPEDSearcher.java) | High-level Lucene query wrapper for case data. |
| **Graph service** | [`iped-engine/src/main/java/iped/engine/graph/GraphService.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-engine/src/main/java/iped/engine/graph/GraphService.java) | Manages embedded graph database for communication links. |
| **Export utilities** | [`iped-engine/src/main/java/iped/engine/task/ExportCSVTask.java`](https://github.com/sepinf-inc/IPED/blob/main/iped-engine/src/main/java/iped/engine/task/ExportCSVTask.java) | Generates CSV metadata reports. |
| **OCR parser** | [`iped/parsers/ocr/OCRParser.java`](https://github.com/sepinf-inc/IPED/blob/main/iped/parsers/ocr/OCRParser.java) | Tesseract integration for text extraction. |
| **Python bridge** | `iped-parsers-impl/src/main/python/` | Jython integration for custom parsers. |
| **Web API** | `iped-webapi/src/main/java/iped/webapi/` | REST endpoints for remote access. |
| **Profiles** | `conf/profiles/*.properties` | Pre-defined forensic processing profiles. |

## Summary

- **IPED** is a Java-based digital forensics platform combining indexing, parsing, and analysis in a single pipeline.
- The **three-layer architecture** separates configuration (`Main`), processing (`Manager`/`Worker`), and data services (`IPEDSearcher`/`GraphService`).
- **Core functionalities** include hash deduplication, recursive container expansion, OCR, face recognition, NER, graph analysis, and multi-format export.
- **Extensibility** via Python and JavaScript allows custom parsers without core modifications.
- **Web API** and **portable case formats** enable both automated remote queries and courtroom presentation.

## Frequently Asked Questions

### What evidence formats does IPED support?

IPED processes raw disk images (E01, VHD, ISO), directory trees, ZIP/7z archives, and mobile extractions (UFED reports). The `ParsingReader` class in [`iped/engine/io/ParsingReader.java`](https://github.com/sepinf-inc/IPED/blob/main/iped/engine/io/ParsingReader.java) handles recursive extraction of nested containers, treating each file as an independent `Item` for metadata extraction and indexing.

### How does IPED handle duplicate files during processing?

The `HashDB` class ([`iped/engine/hashdb/HashDB.java`](https://github.com/sepinf-inc/IPED/blob/main/iped/engine/hashdb/HashDB.java)) computes MD5 and SHA-1 hashes for every item, storing them in an SQLite database. When duplicates are detected, IPED can skip re-processing or store references only, significantly reducing index size and processing time for repetitive evidence.

### Can IPED be integrated into automated forensic workflows?

Yes. The `iped-webapi` module exposes REST endpoints at `http://localhost:8080/api/` for searching metadata, retrieving thumbnails, and exporting results programmatically. Additionally, the command-line interface accepts profiles (`-p triage`, `-p forensic`) that pre-configure task chains, enabling scripted batch processing of multiple evidence sources.

### Is it possible to add custom analysis tasks to IPED without modifying the core code?

Absolutely. IPED supports JavaScript task scripts in `iped/app/resources/scripts/tasks/` and Python parsers in `iped-parsers-impl/src/main/python/`. By implementing `AbstractParser` and registering the class in the service provider configuration, users inject custom logic into the `Worker` task chain without recompiling the engine.