Innosoft Gulf · Big Data · Data Engineering Pathway

Big Data Professional Programme

Distributed Storage, Large-Scale Processing, Real-Time Streaming, Data Lakes and Production Data Engineering

48 Instructor-Led Hours · 4 Stackable Professional Courses

Apply Now View Training Calendar

Programme Overview

This immersive programme is designed for professionals and engineers who need practical experience with the technologies and architectural patterns powering modern Big Data platforms. It provides structured, hands-on training in distributed storage and data lakes, large-scale processing with Apache Spark, real-time event streaming with Apache Kafka, and modern analytical storage architectures.

Participants work with technologies including Hadoop HDFS, Apache Parquet, S3-compatible object storage, Apache Iceberg, PostgreSQL/TimescaleDB, and related platform components.

Participants progress from distributed-system foundations to the design and operation of complete batch and real-time Big Data pipelines, using substantial real-world datasets on multi-node infrastructure. Practical projects include ingesting high-volume event streams, processing and analysing large datasets with Spark, organising historical data in distributed and lakehouse storage, and serving derived analytical results for downstream applications.

Programme Information

Duration 48 instructor-led hours
Structure 4 stackable professional courses · 12 hours each
Delivery Instructor-led training in person at Dubai Knowledge Park or live online
Full Programme Fee AED 9,950 VAT inclusive · Launch Offer

Big Data & Data Engineering Pathway

The programme is organised into four stackable 12-hour courses. Participants may complete the full professional pathway or take an individual course that matches their existing experience and professional objectives.

Course 1 · 12 Hours

Big Data Foundations, Data Lakes & Modern Storage

Develop the architectural and storage foundations required to understand modern Big Data platforms. Participants explore how distributed systems scale and how large datasets are organised across HDFS, Parquet, data lakes and S3-compatible object storage.

Key Areas

  • Big Data characteristics and business drivers
  • Scale-up versus scale-out architectures
  • Distributed computing, partitioning and fault tolerance
  • Batch, streaming and interactive processing
  • Data warehouses, data lakes and lakehouse architectures
  • Roles of Spark, Kafka, distributed storage and analytical databases
  • Common scalability and failure challenges
  • HDFS architecture, blocks, replication and fault tolerance
  • Organising and accessing large datasets in HDFS
  • Apache Parquet and efficient analytical storage
  • Dataset partitioning and the small-files problem
  • S3-compatible object storage and bucket/object concepts
  • HDFS versus object-storage architectures
  • Designing scalable data-lake storage layouts
  • Introduction to Apache Iceberg and lakehouse storage
Practical Projects:
Design a Big Data Architecture
Build a Modern Data Lake

Explore Course 1 →

Course 2 · 12 Hours

Distributed Data Processing with Apache Spark

Build distributed data-processing pipelines with Apache Spark, transforming large datasets stored across HDFS and S3-compatible object storage into usable analytical data products.

Key Areas

  • Spark architecture: drivers, executors, jobs, stages and tasks
  • DataFrames, Spark SQL and schema management
  • RDD concepts and lineage as Spark foundations
  • Transformations, aggregations, joins and window operations
  • Reading and writing CSV, JSON and Parquet across HDFS and S3-compatible storage
  • Partitioning, shuffles, caching and persistence
  • Building reusable PySpark processing pipelines
  • Submitting and monitoring distributed Spark applications
Practical Project:
Build a Distributed Analytics Pipeline

Explore Course 2 →

Course 3 · 12 Hours

Real-Time Data Engineering with Kafka & Spark Streaming

Build reliable event-streaming and real-time analytics pipelines using Apache Kafka and Spark Structured Streaming, from event ingestion and replay through time-based processing and downstream analytical storage.

Apache Kafka

  • Event-driven and streaming data architectures
  • Kafka brokers, topics, partitions and replication
  • Producers, consumers and consumer groups
  • Message keys, partitioning and ordering
  • Offsets, retention and replay
  • Delivery semantics and reliability
  • Event serialisation using JSON and schema-based formats
  • Monitoring Kafka throughput and consumer lag
  • Failure recovery and resilient streaming design

Spark Structured Streaming

  • Structured Streaming architecture and streaming DataFrames
  • Reading Kafka streams with Spark
  • Event time and processing time
  • Windowed aggregations
  • Watermarking and late-arriving data
  • Stateful stream processing
  • Checkpointing and recovery
  • Streaming data quality and validation
  • Writing processed streams to downstream storage
  • Monitoring and troubleshooting streaming applications
Practical Projects:
Build a Real-Time Data Ingestion Pipeline
Build a Real-Time Analytics Pipeline

Explore Course 3 →

Course 4 · 12 Hours

Production Data Engineering & Analytical Serving

Complete the pathway by building analytical serving layers, orchestrating production workflows and integrating distributed storage, Spark, Kafka, stream processing and analytical databases into a complete production-style Big Data platform.

PostgreSQL & TimescaleDB

  • Role of the analytical serving layer in a Big Data platform
  • Relational and time-series data modelling
  • PostgreSQL and TimescaleDB architecture
  • Designing tables for high-volume time-series data
  • Indexing and time-based partitioning
  • Hypertables and continuous aggregates
  • Loading batch and streaming analytical results
  • Time-range queries and analytical aggregations
  • Query performance and data-retention strategies
  • Connecting analytical data to dashboards and applications

Production Data Engineering

  • Pipeline orchestration and scheduling with Apache Airflow
  • Data quality and validation
  • Monitoring, logging and alerting
  • Security and access control
  • Performance and capacity considerations
  • Failure detection, recovery and troubleshooting
  • Data governance and retention
  • Integrating batch, streaming, storage and serving components
Practical Projects:
Build an Analytical Serving Layer
Build an End-to-End Big Data Platform

Explore Course 4 →

Hands-On Lab Environment

Participants work on Innosoft Gulf’s multi-node Big Data infrastructure using substantial real-world datasets. The environment provides hands-on access to distributed storage, Apache Spark, Kafka, object storage, streaming and analytical database services, allowing participants to build, observe and troubleshoot complete Big Data pipelines in a production-like setting.

Who Should Attend

Professionals and graduates seeking to develop practical skills in Big Data and Data Engineering, including those transitioning from software development, IT, database administration, analytics, or other technical backgrounds. Prior experience with Big Data technologies is not required, although basic familiarity with SQL and a programming language such as Python or Scala is recommended.

Recommended Prerequisites

  • Basic familiarity with a programming language, preferably Python.
  • Working knowledge of SQL and relational data concepts.
  • Comfort using a command-line environment is beneficial.
  • No prior experience with Spark, Kafka, Hadoop or other Big Data technologies is required.

Programme Learning Outcomes

By the end of the programme, participants will be able to:

  • Explain Big Data architecture, distributed-systems principles and components of modern data platforms.
  • Design and organise large-scale datasets using distributed storage, Parquet, data lakes and lakehouse architectures.
  • Build and optimise batch-processing and data-engineering pipelines using Apache Spark.
  • Design reliable event-streaming architectures using Apache Kafka.
  • Build real-time processing pipelines using Kafka and Spark Structured Streaming.
  • Manage historical and streaming data using modern table and storage architectures, including Apache Iceberg and S3-compatible object storage.
  • Store, query and serve derived analytical results using PostgreSQL/TimescaleDB for time-series and analytical workloads.
  • Apply data quality, security, governance, monitoring and performance practices across Big Data pipelines.
  • Design, troubleshoot and operate a complete integrated Big Data platform in a production-like environment.

Programme Instructor

Meet Your Instructor

Ahmed El Koutbia - Managing Director and Lead Instructor at Innosoft Gulf

Ahmed El Koutbia

Managing Director & Lead Instructor, Innosoft Gulf

Ahmed El Koutbia is a data scientist, technical instructor, and technology professional with experience spanning financial markets, distributed systems, big data, and AI. At Innosoft Gulf, he is developing a distributed market-data research platform using Apache Kafka, Apache Spark, HDFS, Kubernetes, PostgreSQL/TimescaleDB, and AI agents. Earlier in his career, he worked at the Chicago Board Options Exchange (CBOE) and Sun Microsystems. He is a Red Hat Certified Engineer and holds Java programmer and developer certifications. Ahmed holds a B.S. in Information and Decision Sciences from the University of Illinois Chicago, with graduate-level studies in AI at Stanford University.

CBOE Experience Red Hat Certified Engineer Graduate-Level AI Studies · Stanford

Programme Fees

Individual 12-Hour Course AED 3,500 VAT inclusive
Full 48-Hour Programme AED 9,950 VAT inclusive · Launch Offer
Save AED 4,050 compared with purchasing all four courses individually.

Assessment & Certification

Participants are assessed through hands-on exercises, practical projects and the final integrated Big Data platform project.

Participants who successfully complete the programme and meet the assessment requirements will receive an Innosoft Gulf Certificate of Completion, eligible for attestation by Dubai’s Knowledge and Human Development Authority (KHDA).

Ready to Build Your Big Data & Data Engineering Skills?

Join the complete 48-hour professional pathway or begin with the individual 12-hour course that best matches your experience.

Apply Now View Training Calendar

Instructor-led training in person at Dubai Knowledge Park or live online

Innosoft Gulf · Dubai Knowledge Park · Big Data Professional Programme

Scroll to Top