Innosoft Gulf · Big Data · Data Engineering Pathway
Real-Time Data Engineering with Kafka & Spark Streaming
Apache Kafka, Event Streaming, Spark Structured Streaming and Real-Time Analytics
Big Data Professional Programme · Course 3 of 4
Apply Now View Training CalendarCourse Description
This hands-on course develops the skills required to build reliable real-time data pipelines using Apache Kafka and Spark Structured Streaming.
Participants begin by learning how continuously arriving events are published, partitioned, consumed, retained and replayed through Apache Kafka. They then use Spark Structured Streaming to process Kafka data in real time and transform incoming events into useful analytical results.
Participants work with live event streams and complete practical projects covering both real-time data ingestion and real-time analytical processing.
Course Information
| Duration | 12 instructor-led hours |
| Format | 4 sessions × 3 hours |
| Delivery | Instructor-led training in person at Dubai Knowledge Park or live online |
| Course Fee | AED 3,500 VAT inclusive |
Who Should Attend
Professionals and graduates seeking to develop practical skills in Big Data and Data Engineering, including those transitioning from software development, IT, database administration, analytics, or other technical backgrounds. Prior experience with Big Data technologies is not required, although basic familiarity with SQL and a programming language such as Python or Scala is recommended.
Recommended Prerequisites
- Completion of Course 2 — Distributed Data Processing with Apache Spark or equivalent knowledge of Apache Spark, DataFrames and distributed data processing.
- No prior experience with Apache Kafka or Spark Structured Streaming is required.
Learning Outcomes
By the end of this course, participants will be able to:
- Design and build reliable real-time event-streaming pipelines using Apache Kafka.
- Ingest, publish, consume and replay continuously arriving event data through Kafka.
- Build real-time data-processing pipelines using Kafka and Spark Structured Streaming.
- Process time-based streaming data using windows, event-time concepts and stateful operations.
- Build resilient real-time analytics pipelines that validate, transform, monitor and persist continuously arriving data.
Hands-On Lab Environment
Participants work on Innosoft Gulf’s multi-node Big Data infrastructure using substantial real-world datasets. The environment provides hands-on access to distributed storage, Apache Spark, Kafka, object storage, streaming and analytical database services, allowing participants to build, observe and troubleshoot complete Big Data pipelines in a production-like setting.
Course Curriculum
Module 1
Real-Time Event Streaming with Apache Kafka
Participants learn how modern Big Data platforms ingest and move continuously arriving data using Apache Kafka. They work with live event streams and build reliable pipelines capable of handling high-volume data in real time.
Key Topics
- Event-driven and streaming data architectures
- Kafka brokers, topics, partitions and replication
- Producers, consumers and consumer groups
- Message keys, partitioning and ordering
- Offsets, retention and replay
- Delivery semantics and reliability
- Event serialisation using JSON and schema-based formats
- Monitoring Kafka throughput and consumer lag
- Failure recovery and resilient streaming design
Participants connect to a live WebSocket or event-streaming data source, publish incoming events to Kafka, and build consumers that validate and process the data.
Outcome: A working real-time ingestion pipeline capable of reliably publishing, consuming and replaying high-volume event data through Apache Kafka.
Module 2
Real-Time Stream Processing with Spark Structured Streaming
Participants learn how to process continuously arriving Kafka data using Spark Structured Streaming and build real-time analytical pipelines that transform raw events into useful metrics and data products.
Key Topics
- Structured Streaming architecture and streaming DataFrames
- Reading Kafka streams with Spark
- Event time and processing time
- Windowed aggregations
- Watermarking and late-arriving data
- Stateful stream processing
- Checkpointing and recovery
- Streaming data quality and validation
- Writing processed streams to downstream storage
- Monitoring and troubleshooting streaming applications
Participants consume a live Kafka event stream, validate and transform incoming data, calculate rolling and time-windowed metrics, and persist selected analytical outputs for downstream use.
Outcome: A working real-time analytics pipeline that continuously transforms high-volume Kafka events into reliable analytical results.
Part of the Big Data Professional Programme
This is Course 3 of 4 in the 48-hour Big Data Professional Programme. Participants may take this course independently, provided they have the recommended prerequisite knowledge, or continue through the complete professional pathway.
Previous Course: Distributed Data Processing with Apache Spark ←
Next Course: Production Data Engineering & Analytical Serving →
Assessment & Programme Certification
Assessment in this course includes hands-on exercises and the practical projects described above. Across the full programme, assessment also includes the final integrated Big Data platform project.
Participants who successfully complete the programme and meet the assessment requirements will receive an Innosoft Gulf Certificate of Completion, eligible for attestation by Dubai’s Knowledge and Human Development Authority (KHDA).
Ready to Join This Course?
Real-Time Data Engineering with Kafka & Spark Streaming
Apply Now View Training CalendarInstructor-led training in person at Dubai Knowledge Park or live online
Innosoft Gulf · Dubai Knowledge Park · Real-Time Data Engineering with Kafka & Spark Streaming
