intermediate
Monitoring & Debugging
10 min readLast updated: 2026-07-09
Overview
Learn how to monitor running applications, inspect the History Server, and troubleshoot failed driver and executor runs.
What You Will Learn
In this lesson, you will learn:
- History Server: Accessing execution details for completed jobs.
- Thread Dumps: Analyzing executor bottlenecks.
- Out of Memory (OOM): Troubleshooting common executor memory errors.
Detailed Concept Explanation
Troubleshooting failed or slow distributed Spark applications requires checking logs across different systems.
Monitoring Tools
- Spark Web UI (Port
4040): Available while the driver application is active. Shows active stages, task durations, cached memory states, and running DAGs. - Spark History Server (Port
18080): Displays execution details for completed jobs. Requires configuring Spark to write event log files to shared storage (like HDFS or S3). - Driver Logs: Check the driver logs to troubleshoot query compile errors, driver out-of-memory errors, and configuration issues.
- Executor Logs (stdout/stderr): Check executor logs on worker nodes to troubleshoot task failures, runtime data errors, and executor crashes.
Code Examples
Event Logging Configuration (to enable History Server access)
properties
# Properties configuration defaults inside spark-defaults.conf
spark.eventLog.enabled true
spark.eventLog.dir hdfs://namenode:8020/spark-history-logs
spark.history.fs.logDirectory hdfs://namenode:8020/spark-history-logs
Execution Plan Diagram (Python & Scala)
Execution Plan Diagram
Driver event logs saved to storage
Spark History Server reads storage (18080)
Visualizes completed Jobs details