intermediate

Logging

8 min readLast updated: 2026-07-09

Overview

Learn how to configure log outputs in Spark using log4j configurations, set log levels, and view aggregated worker logs.

What You Will Learn

In this lesson, you will learn:
  • Log4j settings: Customizing logging rules.
  • Log levels: Setting levels (INFO, WARN, ERROR, DEBUG) programmatically.
  • Log Aggregation: Collecting logs from distributed executors.

Detailed Concept Explanation

By default, Spark outputs a massive amount of INFO logs to stdout, which can hide user print statements and error traces. You can manage this output by tuning Spark's logging configurations.

Log Levels

  • FATAL / ERROR: Outputs only when things go wrong and tasks fail.
  • WARN: Outputs warnings (e.g. dynamic partitions configurations warnings) but execution continues.
  • INFO (Default): Outputs progress logs (e.g. stage starts, tasks commits).
  • DEBUG: Outputs detailed internal process states. Recommended for development only.

Code Examples

Python (PySpark) Implementation

python
from pyspark.sql import SparkSession

spark = SparkSession.builder.appName("LoggingTest").getOrCreate()

# Set logging level to ERROR to hide verbose INFO logs
spark.sparkContext.setLogLevel("ERROR")

df = spark.range(1, 100)
print("Count:", df.count())

Expected Output

text
Count: 99

Execution Plan Diagram (Python & Scala)

Execution Plan Diagram
SparkSession.builder
getOrCreate()
sparkContext.setLogLevel(ERROR)
range
count()
print()

Scala Implementation

scala
import org.apache.spark.sql.SparkSession

val spark = SparkSession.builder().appName("LoggingScala").getOrCreate()

// Set log level inside Scala code
spark.sparkContext.setLogLevel("ERROR")

val df = spark.range(1, 100)
println(s"Count: ${df.count()}")

Related Topics