# Big Data Technologies in the architect-awesome Repository: A Complete Guide to Streaming and Batch Processing

> Explore the architect-awesome repository's big data technologies including Storm Flink Kafka Streams Hadoop and Spark for streaming and batch processing. Your guide to essential tools.

- Repository: [xingshaocheng/architect-awesome](https://github.com/xingshaocheng/architect-awesome)
- Tags: tutorial
- Published: 2026-03-05

---

**The xingshaocheng/architect-awesome repository highlights six essential big data technologies across real-time streaming and batch processing categories, including Apache Storm, Flink, Kafka Streams, Hadoop (HDFS/MapReduce/YARN), and Spark, documented in the README.md#大数据 chapter.**

The xingshaocheng/architect-awesome repository serves as a comprehensive knowledge base for software architects, featuring a dedicated "大数据" (Big Data) chapter that catalogs essential big data technologies. This curated list focuses on open-source platforms for real-time stream processing and large-scale batch computation, providing architects with authoritative references for building distributed data pipelines.

## Overview of Big Data Technologies in architect-awesome

The repository organizes big data technologies into two primary categories based on processing patterns. The **streaming** category covers real-time computation engines that process unbounded data streams with low latency. The **batch processing and distributed storage** category encompasses systems designed for high-throughput processing of large static datasets and scalable storage solutions.

According to the source analysis of [`README.md`](https://github.com/xingshaocheng/architect-awesome/blob/main/README.md), the repository specifically highlights six core platforms: Apache Storm, Apache Flink, Kafka Streams, Apache Hadoop (including HDFS, MapReduce, and YARN), and Apache Spark.

## Real-Time Streaming Technologies

The architect-awesome repository dedicates specific sections to three dominant stream processing frameworks, each optimized for different latency and state management requirements.

### Apache Storm

Apache Storm is highlighted as a distributed, fault-tolerant real-time computation system. The repository references Storm in `README.md#storm` as a solution for processing unbounded streams of data with guaranteed tuple processing.

Storm topologies consist of spouts (data sources) and bolts (processing components) that form a directed acyclic graph (DAG) of computation. The system guarantees that every tuple will be fully processed, making it suitable for mission-critical real-time applications.

### Apache Flink

Apache Flink is documented in `README.md#flink` as a stateful stream-processing engine that provides exactly-once semantics and supports both streaming and batch processing through a unified API.

Flink distinguishes itself through true streaming architecture (processing events individually rather than in micro-batches) and sophisticated state management with checkpointing. The repository highlights Flink's ability to handle complex event processing (CEP) and windowed aggregations with high throughput and low latency.

### Kafka Streams

The repository includes Kafka Streams in `README.md#kafka-stream` as a lightweight client library for building stream-processing applications directly on top of Apache Kafka.

Unlike Storm and Flink which require separate clusters, Kafka Streams operates as a library embedded within standard Java applications. It provides stateful processing capabilities, windowing operations, and exactly-once processing guarantees while leveraging Kafka's partitioning and replication mechanisms for fault tolerance.

## Batch Processing and Distributed Storage

For large-scale static data processing, the architect-awesome repository emphasizes the Hadoop ecosystem and Spark's in-memory computing capabilities.

### Apache Hadoop Ecosystem (HDFS, MapReduce, YARN)

The repository dedicates section `README.md#hadoop` to the Apache Hadoop ecosystem, breaking it down into four distinct components:

- **HDFS**: Documented in `README.md#hdfs` as a scalable, fault-tolerant distributed file system designed to run on commodity hardware. HDFS stores data across multiple nodes with replication to ensure reliability.

- **MapReduce**: Referenced in `README.md#mapreduce` as the classic parallel-batch programming model for processing large datasets. MapReduce jobs consist of map phases (filtering and sorting) and reduce phases (summary operations).

- **YARN**: Listed in `README.md#yarn` as the resource-management layer that enables multiple data processing engines to run on the same Hadoop cluster, improving resource utilization beyond traditional MapReduce.

### Apache Spark

Apache Spark is highlighted in `README.md#spark` as a fast, general-purpose engine for large-scale data processing that optimizes for in-memory computation.

Spark provides APIs in Java, Scala, Python, and R, supporting batch processing, interactive queries, streaming, and machine learning through integrated libraries (Spark SQL, MLlib, GraphX). The repository emphasizes Spark's performance advantage over MapReduce for iterative algorithms and interactive data analysis due to its resilient distributed dataset (RDD) abstraction and DAG-based execution engine.

## Practical Implementation Examples

The following code examples demonstrate basic usage patterns for each highlighted technology. These minimal implementations illustrate core concepts while requiring proper cluster configuration and dependencies for production deployment.

### Apache Storm Topology (Java)

```java
public class WordCountTopology {
    public static void main(String[] args) throws Exception {
        TopologyBuilder builder = new TopologyBuilder();

        // Spout reads lines from a static list
        builder.setSpout("sentence-spout", new FixedSentenceSpout(), 1);
        // Bolt splits sentences into words
        builder.setBolt("split-bolt", new SplitSentenceBolt(), 2)
               .shuffleGrouping("sentence-spout");
        // Bolt counts words
        builder.setBolt("count-bolt", new WordCountBolt(), 4)
               .fieldsGrouping("split-bolt", new Fields("word"));

        Config conf = new Config();
        conf.setDebug(true);
        LocalCluster cluster = new LocalCluster();   // for local testing
        cluster.submitTopology("word-count", conf, builder.createTopology());
        Thread.sleep(10000);
        cluster.shutdown();
    }
}

```

This example uses `TopologyBuilder` to construct a DAG of spouts and bolts, submitting to a `LocalCluster` for development testing.

### Apache Flink DataStream (Scala)

```scala
import org.apache.flink.streaming.api.scala._

object SimpleWordCount {
  def main(args: Array[String]): Unit = {
    val env = StreamExecutionEnvironment.getExecutionEnvironment
    val text = env.fromElements(
      "hello world",
      "apache flink streaming",
      "hello flink"
    )
    val counts = text
      .flatMap(_.toLowerCase.split("\\W+"))
      .map((_, 1))
      .keyBy(0)
      .sum(1)

    counts.print()
    env.execute("Scala WordCount")
  }
}

```

The `StreamExecutionEnvironment` provides the entry point for Flink applications, supporting both bounded and unbounded data streams.

### Kafka Streams Application (Java)

```java
Properties props = new Properties();
props.put(StreamsConfig.APPLICATION_ID_CONFIG, "wordcount-app");
props.put(StreamsConfig.BOOTSTRAP_SERVERS_CONFIG, "localhost:9092");
props.put(StreamsConfig.DEFAULT_KEY_SERDE_CLASS_CONFIG, Serdes.String().getClass());
props.put(StreamsConfig.DEFAULT_VALUE_SERDE_CLASS_CONFIG, Serdes.String().getClass());

StreamsBuilder builder = new StreamsBuilder();
KStream<String, String> source = builder.stream("text-input");
KTable<String, Long> counts = source
        .flatMapValues(value -> Arrays.asList(value.toLowerCase().split("\\W+")))
        .groupBy((key, word) -> word)
        .count();

counts.toStream().to("word-count-output", Produced.with(Serdes.String(), Serdes.Long()));
new KafkaStreams(builder.build(), props).start();

```

This creates a stateful stream processing topology that consumes from `text-input` and writes aggregated counts to `word-count-output`.

### Hadoop MapReduce Job (Java)

```java
public class WordCount {
  public static class TokenizerMapper
       extends Mapper<Object, Text, Text, IntWritable> {

    private final static IntWritable one = new IntWritable(1);
    private Text word = new Text();

    public void map(Object key, Text value, Context context)
        throws IOException, InterruptedException {
      StringTokenizer itr = new StringTokenizer(value.toString());
      while (itr.hasMoreTokens()) {
        word.set(itr.nextToken());
        context.write(word, one);
      }
    }
  }

  public static class IntSumReducer
       extends Reducer<Text, IntWritable, Text, IntWritable> {
    private IntWritable result = new IntWritable();

    public void reduce(Text key, Iterable<IntWritable> values, Context context)
        throws IOException, InterruptedException {
      int sum = 0;
      for (IntWritable val : values) {
        sum += val.get();
      }
      result.set(sum);
      context.write(key, result);
    }
  }

  public static void main(String[] args) throws Exception {
    Configuration conf = new Configuration();
    Job job = Job.getInstance(conf, "word count");
    job.setJarByClass(WordCount.class);
    job.setMapperClass(TokenizerMapper.class);
    job.setCombinerClass(IntSumReducer.class);
    job.setReducerClass(IntSumReducer.class);
    job.setOutputKeyClass(Text.class);
    job.setOutputValueClass(IntWritable.class);
    FileInputFormat.addInputPath(job, new Path(args[0]));
    FileOutputFormat.setOutputPath(job, new Path(args[1]));
    System.exit(job.waitForCompletion(true) ? 0 : 1);
  }
}

```

Package with Hadoop 2.x libraries and execute with `hadoop jar wordcount.jar input/ output/`.

### Apache Spark Application (Scala)

```scala
import org.apache.spark.{SparkConf, SparkContext}

object WordCount {
  def main(args: Array[String]): Unit = {
    val conf = new SparkConf().setAppName("SimpleWordCount").setMaster("local[*]")
    val sc   = new SparkContext(conf)

    val lines = sc.textFile("src/main/resources/input.txt")
    val counts = lines.flatMap(_.split("\\W+"))
                     .map(word => (word.toLowerCase, 1))
                     .reduceByKey(_ + _)

    counts.foreach(println)
    sc.stop()
  }
}

```

Runs locally with `setMaster("local[*]")`; replace with a Spark cluster URL for distributed execution.

## Repository Structure and Key Files

The big data technologies are organized in the following sections of [`README.md`](https://github.com/xingshaocheng/architect-awesome/blob/main/README.md):

- `README.md#大数据` – Main index of the Big Data chapter enumerating all technologies
- `README.md#storm` – Apache Storm documentation and tutorial links
- `README.md#flink` – Apache Flink section with stateful processing details
- `README.md#kafka-stream` – Kafka Streams library documentation
- `README.md#hadoop` – Hadoop ecosystem overview
- `README.md#hdfs` – HDFS distributed file system specifics
- `README.md#mapreduce` – MapReduce programming model details
- `README.md#yarn` – YARN resource management documentation
- `README.md#spark` – Apache Spark in-memory processing section

## Summary

The xingshaocheng/architect-awesome repository provides a comprehensive catalog of big data technologies essential for modern distributed systems architecture. Key takeaways include:

- **Six core platforms** are highlighted: Apache Storm, Flink, Kafka Streams, Hadoop (HDFS/MapReduce/YARN), and Spark
- **Streaming architectures** range from dedicated cluster frameworks (Storm, Flink) to embedded libraries (Kafka Streams)
- **Batch processing** emphasizes the Hadoop ecosystem's distributed storage and resource management alongside Spark's DAG-based in-memory computation
- **Documentation structure** in [`README.md`](https://github.com/xingshaocheng/architect-awesome/blob/main/README.md) uses anchored sections for direct navigation to each technology

## Frequently Asked Questions

### What big data technologies are covered in the architect-awesome repository?

The repository covers six major big data technologies across two processing paradigms. For real-time streaming, it includes Apache Storm, Apache Flink, and Kafka Streams. For batch processing and distributed storage, it features the Apache Hadoop ecosystem (comprising HDFS, MapReduce, and YARN) and Apache Spark. Each technology is documented in dedicated sections of the [`README.md`](https://github.com/xingshaocheng/architect-awesome/blob/main/README.md) file.

### Which streaming technology should I choose for real-time processing?

The repository presents three distinct approaches based on infrastructure requirements. **Apache Storm** provides distributed, fault-tolerant real-time computation with guaranteed message processing through spouts and bolts. **Apache Flink** offers stateful stream processing with exactly-once semantics and unified batch/streaming APIs. **Kafka Streams** operates as a lightweight client library embedded in Java applications, ideal for building stream processing on existing Kafka infrastructure without separate clusters.

### How does the repository categorize Hadoop ecosystem components?

According to the source analysis of [`README.md`](https://github.com/xingshaocheng/architect-awesome/blob/main/README.md), the repository breaks down Apache Hadoop into four distinct components within the big data chapter. **HDFS** handles scalable, fault-tolerant distributed storage across commodity hardware. **MapReduce** provides the classic parallel-batch programming model for large-scale data processing. **YARN** serves as the resource-management layer enabling multiple processing engines to share cluster resources efficiently.

### Are practical implementation examples available for these big data technologies?

While the repository itself provides conceptual documentation and tutorial links in sections like `README.md#storm` and `README.md#spark`, the source analysis includes minimal runnable code examples for each technology. These include Java implementations for Storm topologies, Hadoop MapReduce jobs, and Kafka Streams applications, as well as Scala examples for Flink DataStream APIs and Spark RDD transformations. These examples require proper Maven/Gradle dependencies and cluster configuration for production deployment.