Innosoft Gulf · Big Data · Data Engineering Pathway
Distributed Data Processing with Apache Spark
PySpark, Spark SQL, DataFrames, Distributed Processing and Large-Scale Data Engineering
Big Data Professional Programme · Course 2 of 4
Apply Now View Training CalendarCourse Description
Participants build distributed data-processing pipelines with Apache Spark, transforming large datasets stored across HDFS and S3-compatible object storage into usable analytical data products.
Through hands-on work with substantial real-world datasets, participants use PySpark, DataFrames and Spark SQL to validate, transform, aggregate and analyse data across Innosoft Gulf’s multi-node Big Data environment.
Course Information
| Duration | 12 instructor-led hours |
| Format | 4 sessions × 3 hours |
| Delivery | Instructor-led training in person at Dubai Knowledge Park or live online |
| Course Fee | AED 3,500 VAT inclusive |
Who Should Attend
Professionals and graduates seeking to develop practical skills in Big Data and Data Engineering, including those transitioning from software development, IT, database administration, analytics, or other technical backgrounds.
Recommended Prerequisites
- Completion of Course 1 — Big Data Foundations, Data Lakes & Modern Storage or equivalent knowledge of Big Data architecture and distributed storage.
- Participants entering directly should be comfortable with Python and SQL.
- No prior experience with Apache Spark is required.
Learning Outcomes
By the end of this course, participants will be able to:
- Build distributed data-processing pipelines using Apache Spark and PySpark.
- Transform and analyse large datasets using DataFrames and Spark SQL.
- Process data stored across HDFS and S3-compatible object storage and produce analytical outputs in Parquet.
- Apply Spark execution and data-partitioning concepts to improve distributed processing efficiency.
- Build, submit and monitor reusable Spark applications in a distributed multi-node environment.
Hands-On Lab Environment
Participants work on Innosoft Gulf’s multi-node Big Data infrastructure using substantial real-world datasets. The environment provides hands-on access to distributed storage, Apache Spark, Kafka, object storage, streaming and analytical database services, allowing participants to build, observe and troubleshoot complete Big Data pipelines in a production-like setting.
Course Curriculum
Module 1
Spark Architecture and Distributed Processing
Participants are introduced to Apache Spark’s distributed processing model and the core abstractions used to process large datasets across a multi-node environment.
Key Topics
- Spark architecture: drivers, executors, jobs, stages and tasks
- RDD concepts and lineage as Spark foundations
- DataFrames and schema management
Module 2
Data Transformation with DataFrames and Spark SQL
Participants use Spark DataFrames and Spark SQL to transform, aggregate and analyse large datasets stored across distributed storage systems.
Key Topics
- DataFrames and Spark SQL
- Transformations, aggregations, joins and window operations
- Reading and writing CSV, JSON and Parquet across HDFS and S3-compatible storage
Module 3
Distributed Pipeline Performance and Operations
Participants learn how Spark distributes data and computation across the cluster and how to build and operate reusable distributed processing pipelines.
Key Topics
- Partitioning, shuffles, caching and persistence
- Building reusable PySpark processing pipelines
- Submitting and monitoring distributed Spark applications
Participants use a substantial real-world dataset to build a Spark pipeline that validates, transforms and aggregates the data and writes optimized analytical outputs in Parquet.
Outcome: A working distributed Spark pipeline processing large-scale data across the multi-node environment.
Part of the Big Data Professional Programme
This is Course 2 of 4 in the 48-hour Big Data Professional Programme. Participants may take this course independently, provided they have the recommended prerequisite knowledge, or continue through the complete professional pathway.
Previous Course: Big Data Foundations, Data Lakes & Modern Storage ←
Next Course: Real-Time Data Engineering with Kafka & Spark Streaming →
Assessment & Programme Certification
Assessment in this course includes hands-on exercises and the practical project described above. Across the full programme, assessment also includes the final integrated Big Data platform project.
Participants who successfully complete the programme and meet the assessment requirements will receive an Innosoft Gulf Certificate of Completion, eligible for attestation by Dubai’s Knowledge and Human Development Authority (KHDA).
Ready to Join This Course?
Distributed Data Processing with Apache Spark
Apply Now View Training CalendarInstructor-led training in person at Dubai Knowledge Park or live online
Innosoft Gulf · Dubai Knowledge Park · Distributed Data Processing with Apache Spark
