In August 2026 I started as a Graduate Teaching Assistant for CS 5644: Machine Learning with Big Data at Virginia Tech, under Dr. Reza Jafari. The course has about 130 students split across three TAs, and it's built around a simple premise: you don't understand distributed computing until you've watched your own job fail on a real cluster. Students run PySpark jobs against a 143 GiB slice of ImageNet-1k on Virginia Tech's Advanced Research Computing (ARC) HPC system (the Tinkercliffs cluster) through Slurm. What surprised me wasn't the math the students struggled with. It was how many of the bugs were about infrastructure, not machine learning at all.
The Pipeline (Where Most Failures Happen)
Before the Spark bugs, here's the end-to-end pipeline every homework in CS 5644 follows, from raw data to a graded report. The strip below animates the same eight stages a job actually moves through; the flowchart underneath it is the same pipeline as a reference diagram, including the iterate loop it can't show in motion:
Each glow marks a stage where state changes hands, and a place a bug can hide silently.
Each arrow is a place something can go silently wrong. A Spark job that mis-scales across cores doesn't crash with a clear error: it just runs slower, or runs fine at 4 cores and dies at 32. The pipeline looks identical whether it's correct or quietly broken.
What the Course Covers
CS 5644 is graduate-level, with students ranging from people who've run production Spark clusters to those who've never touched a job scheduler. The curriculum spans:
| Module | Topics | Tools |
|---|---|---|
| Spark Fundamentals | RDDs, DataFrames, lazy evaluation, partitions | PySpark |
| Cluster Computing | Job scheduling, resource allocation, compute vs. storage accounts | Slurm, Advanced Research Computing (ARC) / Tinkercliffs |
| Distributed I/O | Parquet, partitioning strategy, shuffle behavior | Spark SQL, Parquet |
| Performance Tuning | Broadcast joins, caching, driver vs. executor memory | Spark UI, spark-submit configs |
| Big Data ML | Feature engineering at scale, aggregation pipelines, benchmarking | MLlib, PySpark |
Bug #1: The Wrong Slurm Account
While pilot-testing HW2 before it went out to the full class, I ran the starter script exactly as written and it pointed at the wrong account:
#!/bin/bash
# WRONG -- this is the storage project path, not a compute allocation
#SBATCH --account=cs5644
#SBATCH --cpus-per-task=32
#SBATCH --partition=normal_q
spark-submit --master local[$SLURM_CPUS_PER_TASK] train.py
# Job either fails to submit or queues against the wrong allocation.
# "cs5644" is where the dataset lives (/projects/cs5644/), not who's
# paying for the compute.
# CORRECT -- check your actual compute allocation first
$ sacctmgr show associations user=$USER format=account,partition
# Account Partition
# ----------- ----------
# arc-vt-xxx normal_q
# then reference that account in the job script
#SBATCH --account=arc-vt-xxx
#SBATCH --cpus-per-task=32
#SBATCH --partition=normal_q
The lesson generalizes past this one cluster: on shared HPC systems, the path where your data lives and the account you're billed against are two different namespaces that happen to share a name. Catching this in pilot testing meant it never reached the 127 students who would have hit it on their first submission. On VT's ARC, the same account-vs-storage mix-up is easy to double-check without touching a terminal at all: coldfront.arc.vt.edu lists every allocation you actually have compute access under, the same thing sacctmgr is querying, just as a dashboard instead of a command.
Bug #2: Driver Memory in Local Mode
Homework ran fine on a 4-core pilot run with --driver-memory 16g. On 32 cores, the exact same code threw OutOfMemoryError:
# WRONG -- memory that doesn't scale with cores
spark-submit \
--master local[32] \
--driver-memory 16g \
process_imagenet.py
# java.lang.OutOfMemoryError: Java heap space
# The reason: local[N] runs all N task threads inside ONE JVM process,
# sharing ONE heap. Adding cores doesn't add memory -- it adds more
# threads competing for the same 16g.
# CORRECT -- scale driver memory with core count
spark-submit \
--master local[32] \
--driver-memory 48g \
process_imagenet.py
# Same code, same data -- succeeds because the heap can hold
# 32 concurrent tasks' worth of decoded batches instead of 4's worth.
This is the single most common "it worked yesterday" ticket in office hours. Scaling --cpus-per-task without scaling --driver-memory alongside it is the default way to break a local-mode Spark job.
Bug #3: The Broadcast Join That Made Things Slower
Students are taught that broadcast() speeds up joins by avoiding a shuffle. At this dataset's scale (around 87,000 rows in the lookup table), several students found the opposite:
from pyspark.sql import functions as F
import time
# WRONG -- "Optimization" applied without measuring
start = time.time()
result = large_df.join(F.broadcast(labels_df), on="class_id")
result.count()
print(f"Broadcast join: {time.time() - start:.2f}s")
# Broadcast join: 41.3s -- slower than the plain join below
# CORRECT -- Measure both before choosing
start = time.time()
result = large_df.join(labels_df, on="class_id")
result.count()
print(f"Plain join: {time.time() - start:.2f}s")
# Plain join: 28.7s
# The broadcast still has to serialize, transfer, and deserialize the
# full lookup table to every executor up front -- that cost only pays
# off when the broadcast table is small relative to the shuffle it
# saves. At 87K rows, the broadcast overhead exceeded the shuffle cost.
The lesson: optimization hints are not free, and "textbook faster" only holds at the scale the textbook assumed. Measure before applying a hint, and understand what it's trading away.
Bug #4: Speedup Ratio vs. Absolute Timing
On a shared cluster, absolute wall-clock time is close to meaningless: students' runs came in 5-8x slower than my pilot runs simply from contention with everyone else's jobs on the same nodes:
def speedup(t_baseline, t_parallel):
return t_baseline / t_parallel
# Pilot measurements (uncontended)
timings = {4: 118.2, 8: 61.4, 16: 33.9, 32: 20.7} # seconds
baseline = timings[4]
for cores, t in timings.items():
print(f"{cores:>2} cores: {t:6.1f}s speedup={speedup(baseline, t):.2f}x")
# 4 cores: 118.2s speedup=1.00x
# 8 cores: 61.4s speedup=1.92x
# 16 cores: 33.9s speedup=3.49x
# 32 cores: 20.7s speedup=5.71x (diminishing returns, not linear)
# Grading rubric: we score the SPEEDUP RATIO and the reasoning about
# why it isn't linear (I/O bound stages, GC pauses, shuffle overhead),
# not the raw seconds, which depend entirely on how busy the cluster
# was when your job happened to run.
A 5.4–5.7x speedup from 4 to 32 cores was typical in pilot runs. Reasoning about why it plateaus mattered more than the number itself.
Bug #5: "My Laptop Beat ARC" Misreading
In HW1 reflections, a recurring claim was that a personal laptop outran the HPC cluster, with students crediting it to "more cores" or ARC being shared. For small jobs, neither is the real cause:
import time
t0 = time.time()
spark = SparkSession.builder.master("local[8]").getOrCreate()
t1 = time.time()
print(f"Spark session startup: {t1 - t0:.2f}s")
# Spark session startup: 4.8s
df = spark.read.parquet("small_sample/") # a few MB
t2 = time.time()
result = df.count()
t3 = time.time()
print(f"Actual query time: {t3 - t2:.2f}s")
# Actual query time: 0.3s
# 4.8s of JVM + Spark session overhead completely dominates a 0.3s
# query. On a small job, a laptop "wins" because ARC pays the same
# startup tax through Slurm scheduling delay on top of it -- not
# because the laptop has more compute.
Parallelism in this course is configured with --cpus-per-task and local[N], not by choosing a different machine. Startup and session overhead dominate small jobs regardless of where they run; it only stops mattering once the actual workload is large enough to amortize it.
Bug #6: Deterministic Output as a Sanity Check
Some aggregate queries should return byte-identical results across runs, no matter how many cores or partitions are used. When they don't, that's a bug in the pipeline, not noise:
# WRONG -- Meaningless without understanding the data first
top10 = (df.groupBy("class_id").count()
.orderBy(F.desc("count"))
.limit(10))
top10.show()
# Every ImageNet-1k class has exactly 1,300 training images.
# A "top 10 classes by count" query on the full training set
# is meaningless -- it's a tie across all 1,000 classes.
# CORRECT -- What's actually worth checking: does the count come out
# the same every time, regardless of partitioning?
for partitions in [4, 16, 64]:
n = df.repartition(partitions).count()
print(f"partitions={partitions:>3} total_rows={n}")
# partitions= 4 total_rows=1300000
# partitions= 16 total_rows=1300000
# partitions= 64 total_rows=1300000
# If these ever disagree, something upstream is dropping or
# duplicating rows -- that's the real signal, not the ranking.
Same instinct as printing 10 examples before modeling: know what your data actually looks like before you interpret a query's output as meaningful.
Bug #7: Screenshots Aren't Evidence
Reports have to be backed by the code that was actually submitted, with numbers that trace back to a real run:
# WRONG -- A screenshot of a terminal window pasted into the report
# -- can't be verified, can't be searched, can't be diffed against
# the submitted code, and crops out context half the time.
# CORRECT -- Redirect and paste the terminal output as text
$ spark-submit --master local[32] benchmark.py | tee results.txt
$ cat results.txt >> report.md
# Now the numbers in the report are literally the output of the
# submitted script -- graders can copy a line and grep the code
# for where it came from.
It's a small habit that changes what a report actually proves.
Teaching insight: the shared kitchen. Local-mode Spark memory is one shared kitchen with one fridge. Adding 32 cooks (cores) to a kitchen that still has the same size fridge (heap) doesn't make dinner come out faster; it makes the fridge overflow. The fix was never "add more cooks," it was "get a bigger fridge first." Every time a student scaled --cpus-per-task without touching --driver-memory, they'd added cooks and forgotten the fridge.
The Slurm Commands Every Student Should Know
Most of the bugs above show up in Slurm's own tooling before they show up in a Spark log, if you know where to look. Job IDs and node names below are placeholders, not real ones:
# before submitting: which account can I actually charge? (Bug #1)
sacctmgr show assoc user=$USER format=account%25,partition,qos
# the storage path (/projects/<course>/) is NOT necessarily your compute account
# submit
sbatch submit.sh
# Submitted batch job 1234567
# quick test before a long run: validate the script without queueing it
sbatch --test-only submit.sh
# check status
squeue -u $USER # my pending/running jobs
squeue -j 1234567 --start # estimated start time if PENDING
# ST column: PD = pending, R = running, CG = completing
# REASON column: (Priority), (Resources), (QOSMaxCpuPerUserLimit), etc.
scontrol show job 1234567 # full details: nodes, CPUs, mem, working dir
# watch output live
tail -f slurm-1234567.out
# after it finishes: did it succeed, and how much did it use?
sacct -j 1234567 --format=JobID,JobName,State,Elapsed,MaxRSS,ReqMem,AllocCPUS,ExitCode
# State: COMPLETED / FAILED / OUT_OF_MEMORY / TIMEOUT / CANCELLED
seff 1234567 # CPU and memory efficiency summary
# cancel
scancel 1234567
scancel -u $USER # cancel all my jobs, careful
# cluster availability
sinfo -s # partitions and idle/allocated nodes
sacct/seffare how you debug Bug #2. AnOUT_OF_MEMORYstate, or a MaxRSS close to ReqMem, tells you memory was the problem before you even open the log. A Spark driver OOM inside the JVM can still show up asFAILEDrather thanOUT_OF_MEMORY, so check the.outfile forjava.lang.OutOfMemoryErrortoo.squeue's REASON column explains "why is my job not running." It's often queue contention on a shared cluster, the same contention behind Bug #4's noisy timings, not a bug in your code.- Low CPU efficiency in
seff(say, 32 cores requested and about 10% used) meanslocal[N]doesn't match--cpus-per-task. That's Bug #5. The cheat sheet below useslocal[${SLURM_CPUS_PER_TASK}]so the two can't drift apart. - Resubmitting a job while one is already pending just adds to the queue. Check
squeuefirst, and usesacctmgror coldfront.arc.vt.edu before guessing an account, per Bug #1.
PySpark-on-Slurm Cheat Sheet
#!/bin/bash
#SBATCH --job-name=cs5644-benchmark
#SBATCH --account=<your compute account> # check: sacctmgr show associations user=$USER, or coldfront.arc.vt.edu
#SBATCH --partition=normal_q
#SBATCH --cpus-per-task=32
#SBATCH --mem=64G # leave headroom above --driver-memory
#SBATCH --time=00:30:00
module load Spark
spark-submit \
--master local[${SLURM_CPUS_PER_TASK}] \
--driver-memory 48g \
benchmark.py
# --- benchmark.py: minimal timing harness ---
# import time
# from pyspark.sql import SparkSession
#
# results = {}
# for cores in [4, 8, 16, 32]:
# spark = SparkSession.builder.master(f"local[{cores}]").getOrCreate()
# start = time.time()
# spark.read.parquet("imagenet/").groupBy("class_id").count().collect()
# results[cores] = time.time() - start
# spark.stop()
#
# baseline = results[min(results)]
# for cores, t in results.items():
# print(f"{cores:>2} cores: {t:6.1f}s speedup={baseline/t:.2f}x")
Takeaways for Anyone Running Spark on HPC
- A storage path and a compute account can share a name on an HPC cluster: verify with
sacctmgror the ARC allocations dashboard, don't guess. local[N]mode shares one JVM heap across all N threads. Scale--driver-memoryevery time you scale cores.- Optimization hints like
broadcast()aren't free: measure before assuming, especially at moderate table sizes. - On a shared cluster, report the speedup ratio and explain it. Absolute wall-clock time is noise you don't control.
- Startup and session overhead dominate small jobs: that's why "my laptop is faster" is usually about job size, not hardware.
- Understand your dataset before trusting a ranking query. Deterministic aggregate counts are a better sanity check than a "top 10."
- Paste terminal output as text, not screenshots. Reports should trace directly back to the submitted code.
Srikanth Badavath