What Are the Main Functionalities of IPED? A Complete Technical Guide
IPED (Indexador & Processador de Evidências Digitais) is a full-stack digital forensics platform that processes evidence images, extracts metadata, performs OCR and face recognition, indexes content for search, and provides graph analysis of communication data.
The open-source IPED project (sepinf-inc/IPED) delivers an end-to-end digital investigation suite written in Java. Understanding the main functionalities of IPED helps forensic examiners leverage its modular pipeline for evidence processing, from hash deduplication to AI-powered entity recognition.
Three-Layer Architecture
IPED organizes its main functionalities into three tightly coupled architectural layers that handle everything from command-line input to advanced graph queries.
Entry and Configuration Layer
The application bootstrap resides in iped-app/src/main/java/iped/app/processing/Main.java. This Main class parses command-line arguments through CmdLineArgsImpl, validates output directories, and loads processing profiles. It prepares the runtime environment by copying libraries and UI resources into the case folder before handing control to the processing engine.
Processing Engine Layer
At the heart of IPED sits iped-engine/src/main/java/iped/engine/core/Manager.java. The Manager class orchestrates the case lifecycle using a producer-consumer pattern. Two ItemProducer threads first count items for progress estimation, then produce Item objects consumed by a pool of Worker threads defined in Worker.java. Each worker executes the configurable task chain (hashing, parsing, OCR) defined in IndexTaskConfig.
Data Access and Services Layer
Post-processing functionality lives in iped-engine/src/main/java/iped/engine/search/IPEDSearcher.java, which wraps Lucene queries for programmatic case interrogation. The PreviewRepositoryManager handles thumbnail generation, while GraphService manages the embedded Neo4j-compatible database for communication analysis. A thin Web API provides HTTP endpoints for remote queries.
End-to-End Processing Workflow
The main functionalities of IPED follow a deterministic seven-stage pipeline:
- Argument parsing –
Mainbuilds aCmdLineArgsImplinstance, validates the output folder, and resolves the selected profile. - Output preparation –
Manager.prepareOutputFolder()copies runtime dependencies and creates thedata/subdirectory. - Index initialization – An
IndexWriterinstantiates with a customAppAnalyzerfor full-text and metadata indexing. - Producer-consumer processing –
ItemProducerthreads feed evidence intoWorkerpools that execute the task chain (hashing, container expansion, OCR, face recognition, NER). - Periodic commits – Background threads flush the Lucene index and auxiliary storages (CSV export, graph, Elasticsearch).
- Post-processing –
Managerremoves empty tree nodes, filters keywords, updates image paths for portable cases, and optionally force-merges the index. - Search and export – The case opens in the UI or accepts queries via
IPEDSearcher, withExportFileTaskandExportCSVTaskgenerating reports.
Core Functional Blocks
Evidence Hashing and Deduplication
HashDB (iped/engine/hashdb/HashDB.java) computes MD5 and SHA-1 hashes, storing them in an SQLite database to automatically discard duplicate items during ingestion.
Recursive Container Expansion
ParsingReader (iped/engine/io/ParsingReader.java) detects archive formats (ZIP, 7z, ISO, E01, VHD) and streams their entries as independent Item objects for recursive processing.
Optical Character Recognition
OCRParser (iped/parsers/ocr/OCRParser.java) sends image streams to Tesseract 5 or cloud services (Azure/Google), storing extracted text in the index for full-text search.
Face Recognition and Similarity Search
FaceRecognitionTask (iped/app/resources/scripts/tasks/FaceRecognitionTask.java) generates face embeddings using dlib or InsightFace, building a searchable vector index for identifying persons across evidence.
Named Entity Recognition
NERTask leverages Stanford CoreNLP models via the iped/parsers/ner/ package to automatically tag persons, organizations, and locations within extracted text.
Graph Analysis
GraphService and GraphTask store communication links (calls, emails, instant messages) in an embedded graph database, enabling queries like "who called whom" and relationship visualization.
Web API and Remote Access
The iped-webapi module exposes REST endpoints (/search, /metadata, /thumb) that return JSON, raw content, or thumbnails for integration with remote tools.
Export and Reporting
ExportFileTask, ExportCSVTask, and HTMLReportGenerator produce portable case bundles, CSV metadata indexes, and interactive HTML timelines for court presentation.
Processing Profiles
Configuration files in conf/profiles/ define pre-defined task sets for forensic, pedophile-content (CSAM), triage, fast-preview, and blind (auto-extraction) modes, selectable via the --profile command-line argument.
Extensibility Framework
Developers can inject custom logic via JavaScript (scripts/) or Python (iped-parsers-impl/src/main/python/) without modifying core engine code, using the Jython bridge for parser integration.
Code Examples
Running a Case from the Command Line
java -jar iped.jar \
-i /evidence/image.E01 \
-o /cases/mycase \
-p forensic \
--keywords /path/to/keywords.txt \
--log /cases/mycase/log.txt
-i= input evidence (disk image, directory, or UFED report).-o= output folder for the portable case.-p= processing profile from theconf/profiles/directory.
This command invokes Main.main(), which parses arguments and delegates to Manager.process().
Searching a Case Programmatically
import iped.engine.search.IPEDSearcher;
import iped.engine.search.SearchResult;
import iped.engine.IPEDSource;
import java.io.File;
public class SimpleSearch {
public static void main(String[] args) throws Exception {
File caseRoot = new File("/cases/mycase");
try (IPEDSource caseDb = new IPEDSource(caseRoot)) {
IPEDSearcher searcher = new IPEDSearcher(caseDb,
"mime_type:\"image/png\" AND content:\"confidential\"");
SearchResult result = searcher.search();
System.out.println("Found " + result.getLength() + " items:");
for (int i = 0; i < result.getLength(); i++) {
System.out.println(" - " + caseDb.getItemByID(result.getId(i)).getPath());
}
}
}
}
IPEDSource opens the case in read-only mode, IPEDSearcher builds Lucene queries, and SearchResult enumerates matching item IDs.
Adding a Custom Python Parser
Create myparser.py in iped-parsers-impl/src/main/python/:
from iped.parsers import AbstractParser
from iped.io import SeekableInputStream
class MyParser(AbstractParser):
def getSupportedMimeTypes(self):
return ["text/plain"]
def parse(self, stream: SeekableInputStream, metadata, contentHandler):
text = stream.read().decode('utf-8')
if "sensitive" in text.lower():
contentHandler.characters("FOUND SENSITIVE")
Register the parser in iped-parsers-impl/src/main/resources/META-INF/services/iped.parsers.AbstractParser by adding myparser.MyParser. The ParsingReader automatically loads this via Jython during processing.
Querying the Web API
curl "http://localhost:8080/api/search?q=credit+card+number"
Returns a JSON array with item IDs, paths, and preview URLs. Implementation resides in iped-webapi/src/main/java/iped/webapi/SearchResource.java.
Key Files and Entry Points
| Module | File | Role |
|---|---|---|
| Application entry | iped-app/src/main/java/iped/app/processing/Main.java |
CLI entry point and configuration loader. |
| Core engine | iped-engine/src/main/java/iped/engine/core/Manager.java |
Orchestrates producers, workers, and index management. |
| Task configuration | iped-engine/src/main/java/iped/engine/config/IndexTaskConfig.java |
Defines the processing task chain for workers. |
| Search API | iped-engine/src/main/java/iped/engine/search/IPEDSearcher.java |
High-level Lucene query wrapper for case data. |
| Graph service | iped-engine/src/main/java/iped/engine/graph/GraphService.java |
Manages embedded graph database for communication links. |
| Export utilities | iped-engine/src/main/java/iped/engine/task/ExportCSVTask.java |
Generates CSV metadata reports. |
| OCR parser | iped/parsers/ocr/OCRParser.java |
Tesseract integration for text extraction. |
| Python bridge | iped-parsers-impl/src/main/python/ |
Jython integration for custom parsers. |
| Web API | iped-webapi/src/main/java/iped/webapi/ |
REST endpoints for remote access. |
| Profiles | conf/profiles/*.properties |
Pre-defined forensic processing profiles. |
Summary
- IPED is a Java-based digital forensics platform combining indexing, parsing, and analysis in a single pipeline.
- The three-layer architecture separates configuration (
Main), processing (Manager/Worker), and data services (IPEDSearcher/GraphService). - Core functionalities include hash deduplication, recursive container expansion, OCR, face recognition, NER, graph analysis, and multi-format export.
- Extensibility via Python and JavaScript allows custom parsers without core modifications.
- Web API and portable case formats enable both automated remote queries and courtroom presentation.
Frequently Asked Questions
What evidence formats does IPED support?
IPED processes raw disk images (E01, VHD, ISO), directory trees, ZIP/7z archives, and mobile extractions (UFED reports). The ParsingReader class in iped/engine/io/ParsingReader.java handles recursive extraction of nested containers, treating each file as an independent Item for metadata extraction and indexing.
How does IPED handle duplicate files during processing?
The HashDB class (iped/engine/hashdb/HashDB.java) computes MD5 and SHA-1 hashes for every item, storing them in an SQLite database. When duplicates are detected, IPED can skip re-processing or store references only, significantly reducing index size and processing time for repetitive evidence.
Can IPED be integrated into automated forensic workflows?
Yes. The iped-webapi module exposes REST endpoints at http://localhost:8080/api/ for searching metadata, retrieving thumbnails, and exporting results programmatically. Additionally, the command-line interface accepts profiles (-p triage, -p forensic) that pre-configure task chains, enabling scripted batch processing of multiple evidence sources.
Is it possible to add custom analysis tasks to IPED without modifying the core code?
Absolutely. IPED supports JavaScript task scripts in iped/app/resources/scripts/tasks/ and Python parsers in iped-parsers-impl/src/main/python/. By implementing AbstractParser and registering the class in the service provider configuration, users inject custom logic into the Worker task chain without recompiling the engine.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →