Innosoft Gulf · Big Data · Data Engineering Pathway
Big Data Professional Programme
Distributed Storage, Large-Scale Processing, Real-Time Streaming, Data Lakes and Production Data Engineering
48 Instructor-Led Hours · 4 Stackable Professional Courses
Apply Now View Training CalendarProgramme Overview
This immersive programme is designed for professionals and engineers who need practical experience with the technologies and architectural patterns powering modern Big Data platforms. It provides structured, hands-on training in distributed storage and data lakes, large-scale processing with Apache Spark, real-time event streaming with Apache Kafka, and modern analytical storage architectures.
Participants work with technologies including Hadoop HDFS, Apache Parquet, S3-compatible object storage, Apache Iceberg, PostgreSQL/TimescaleDB, and related platform components.
Participants progress from distributed-system foundations to the design and operation of complete batch and real-time Big Data pipelines, using substantial real-world datasets on multi-node infrastructure. Practical projects include ingesting high-volume event streams, processing and analysing large datasets with Spark, organising historical data in distributed and lakehouse storage, and serving derived analytical results for downstream applications.
Programme Information
| Duration | 48 instructor-led hours |
| Structure | 4 stackable professional courses · 12 hours each |
| Delivery | Instructor-led training in person at Dubai Knowledge Park or live online |
| Full Programme Fee | AED 9,950 VAT inclusive · Launch Offer |
Big Data & Data Engineering Pathway
The programme is organised into four stackable 12-hour courses. Participants may complete the full professional pathway or take an individual course that matches their existing experience and professional objectives.
Course 1 · 12 Hours
Big Data Foundations, Data Lakes & Modern Storage
Develop the architectural and storage foundations required to understand modern Big Data platforms. Participants explore how distributed systems scale and how large datasets are organised across HDFS, Parquet, data lakes and S3-compatible object storage.
Key Areas
- Big Data characteristics and business drivers
- Scale-up versus scale-out architectures
- Distributed computing, partitioning and fault tolerance
- Batch, streaming and interactive processing
- Data warehouses, data lakes and lakehouse architectures
- Roles of Spark, Kafka, distributed storage and analytical databases
- Common scalability and failure challenges
- HDFS architecture, blocks, replication and fault tolerance
- Organising and accessing large datasets in HDFS
- Apache Parquet and efficient analytical storage
- Dataset partitioning and the small-files problem
- S3-compatible object storage and bucket/object concepts
- HDFS versus object-storage architectures
- Designing scalable data-lake storage layouts
- Introduction to Apache Iceberg and lakehouse storage
Design a Big Data Architecture
Build a Modern Data Lake
Course 2 · 12 Hours
Distributed Data Processing with Apache Spark
Build distributed data-processing pipelines with Apache Spark, transforming large datasets stored across HDFS and S3-compatible object storage into usable analytical data products.
Key Areas
- Spark architecture: drivers, executors, jobs, stages and tasks
- DataFrames, Spark SQL and schema management
- RDD concepts and lineage as Spark foundations
- Transformations, aggregations, joins and window operations
- Reading and writing CSV, JSON and Parquet across HDFS and S3-compatible storage
- Partitioning, shuffles, caching and persistence
- Building reusable PySpark processing pipelines
- Submitting and monitoring distributed Spark applications
Build a Distributed Analytics Pipeline
Course 3 · 12 Hours
Real-Time Data Engineering with Kafka & Spark Streaming
Build reliable event-streaming and real-time analytics pipelines using Apache Kafka and Spark Structured Streaming, from event ingestion and replay through time-based processing and downstream analytical storage.
Apache Kafka
- Event-driven and streaming data architectures
- Kafka brokers, topics, partitions and replication
- Producers, consumers and consumer groups
- Message keys, partitioning and ordering
- Offsets, retention and replay
- Delivery semantics and reliability
- Event serialisation using JSON and schema-based formats
- Monitoring Kafka throughput and consumer lag
- Failure recovery and resilient streaming design
Spark Structured Streaming
- Structured Streaming architecture and streaming DataFrames
- Reading Kafka streams with Spark
- Event time and processing time
- Windowed aggregations
- Watermarking and late-arriving data
- Stateful stream processing
- Checkpointing and recovery
- Streaming data quality and validation
- Writing processed streams to downstream storage
- Monitoring and troubleshooting streaming applications
Build a Real-Time Data Ingestion Pipeline
Build a Real-Time Analytics Pipeline
Course 4 · 12 Hours
Production Data Engineering & Analytical Serving
Complete the pathway by building analytical serving layers, orchestrating production workflows and integrating distributed storage, Spark, Kafka, stream processing and analytical databases into a complete production-style Big Data platform.
PostgreSQL & TimescaleDB
- Role of the analytical serving layer in a Big Data platform
- Relational and time-series data modelling
- PostgreSQL and TimescaleDB architecture
- Designing tables for high-volume time-series data
- Indexing and time-based partitioning
- Hypertables and continuous aggregates
- Loading batch and streaming analytical results
- Time-range queries and analytical aggregations
- Query performance and data-retention strategies
- Connecting analytical data to dashboards and applications
Production Data Engineering
- Pipeline orchestration and scheduling with Apache Airflow
- Data quality and validation
- Monitoring, logging and alerting
- Security and access control
- Performance and capacity considerations
- Failure detection, recovery and troubleshooting
- Data governance and retention
- Integrating batch, streaming, storage and serving components
Build an Analytical Serving Layer
Build an End-to-End Big Data Platform
Hands-On Lab Environment
Participants work on Innosoft Gulf’s multi-node Big Data infrastructure using substantial real-world datasets. The environment provides hands-on access to distributed storage, Apache Spark, Kafka, object storage, streaming and analytical database services, allowing participants to build, observe and troubleshoot complete Big Data pipelines in a production-like setting.
Who Should Attend
Professionals and graduates seeking to develop practical skills in Big Data and Data Engineering, including those transitioning from software development, IT, database administration, analytics, or other technical backgrounds. Prior experience with Big Data technologies is not required, although basic familiarity with SQL and a programming language such as Python or Scala is recommended.
Recommended Prerequisites
- Basic familiarity with a programming language, preferably Python.
- Working knowledge of SQL and relational data concepts.
- Comfort using a command-line environment is beneficial.
- No prior experience with Spark, Kafka, Hadoop or other Big Data technologies is required.
Programme Learning Outcomes
By the end of the programme, participants will be able to:
- Explain Big Data architecture, distributed-systems principles and components of modern data platforms.
- Design and organise large-scale datasets using distributed storage, Parquet, data lakes and lakehouse architectures.
- Build and optimise batch-processing and data-engineering pipelines using Apache Spark.
- Design reliable event-streaming architectures using Apache Kafka.
- Build real-time processing pipelines using Kafka and Spark Structured Streaming.
- Manage historical and streaming data using modern table and storage architectures, including Apache Iceberg and S3-compatible object storage.
- Store, query and serve derived analytical results using PostgreSQL/TimescaleDB for time-series and analytical workloads.
- Apply data quality, security, governance, monitoring and performance practices across Big Data pipelines.
- Design, troubleshoot and operate a complete integrated Big Data platform in a production-like environment.
Programme Instructor
Meet Your Instructor
Ahmed El Koutbia
Managing Director & Lead Instructor, Innosoft Gulf
Ahmed El Koutbia is a data scientist, technical instructor, and technology professional with experience spanning financial markets, distributed systems, big data, and AI. At Innosoft Gulf, he is developing a distributed market-data research platform using Apache Kafka, Apache Spark, HDFS, Kubernetes, PostgreSQL/TimescaleDB, and AI agents. Earlier in his career, he worked at the Chicago Board Options Exchange (CBOE) and Sun Microsystems. He is a Red Hat Certified Engineer and holds Java programmer and developer certifications. Ahmed holds a B.S. in Information and Decision Sciences from the University of Illinois Chicago, with graduate-level studies in AI at Stanford University.
Programme Fees
| Individual 12-Hour Course | AED 3,500 VAT inclusive |
| Full 48-Hour Programme |
AED 9,950 VAT inclusive · Launch Offer
Save AED 4,050 compared with purchasing
all four courses individually.
|
Assessment & Certification
Participants are assessed through hands-on exercises, practical projects and the final integrated Big Data platform project.
Participants who successfully complete the programme and meet the assessment requirements will receive an Innosoft Gulf Certificate of Completion, eligible for attestation by Dubai’s Knowledge and Human Development Authority (KHDA).
Ready to Build Your Big Data & Data Engineering Skills?
Join the complete 48-hour professional pathway or begin with the individual 12-hour course that best matches your experience.
Apply Now View Training CalendarInstructor-led training in person at Dubai Knowledge Park or live online
Innosoft Gulf · Dubai Knowledge Park · Big Data Professional Programme
