intermediate
Hive
8 min read
Overview
Connect Spark to Apache Hive Metastore to share table metadata, write catalog tables, and query partitioned Hive tables.
What You Will Learn
In this lesson, you will learn:
- Hive Metastore: Reading shared table schemas across SQL platforms.
- Enable Hive Support: Configuring SparkSession with Hive support.
- Catalog tables: Interacting with managed vs external Hive tables.
Detailed Concept Explanation
By default, Spark uses an in-memory catalog to track temporary tables and views. To persist table schemas across different sessions and share them with other engines (like Presto or Impala), you can connect Spark to a shared Hive Metastore.
Code Examples
Python (PySpark) Implementation
python
from pyspark.sql import SparkSession
# Initialize SparkSession with Hive Support enabled
spark = SparkSession.builder \
.appName("HiveTest") \
.enableHiveSupport() \
.getOrCreate()
# Create a managed Hive table
spark.sql("CREATE TABLE IF NOT EXISTS hive_users (id INT, name STRING) USING hive")
# Insert data
spark.sql("INSERT INTO hive_users VALUES (1, 'Alice')")
# Query table
spark.sql("SELECT * FROM hive_users").show()
Expected Output
text
+---+-----+
| id| name|
+---+-----+
| 1|Alice|
+---+-----+
Execution Plan Diagram (Python & Scala)
Execution Plan Diagram
SparkSession.builder
enableHiveSupport()
spark.sql(CREATE TABLE)
spark.sql(INSERT INTO)
spark.sql(SELECT)
show()
Scala Implementation
scala
import org.apache.spark.sql.SparkSession
val spark = SparkSession.builder()
.appName("HiveScala")
.enableHiveSupport()
.getOrCreate()
spark.sql("CREATE TABLE IF NOT EXISTS hive_users (id INT, name STRING) USING hive")
spark.sql("INSERT INTO hive_users VALUES (1, 'Alice')")
spark.sql("SELECT * FROM hive_users").show()
Expected Output
text
+---+-----+
| id| name|
+---+-----+
| 1|Alice|
+---+-----+