
What happened
Using Managed Service for Apache Spark and MLlib's LogisticRegressionWithLBFGS on U.S. flight data, the author trained a model from 12,303 examples and evaluated it on 12,307, finding departure delay drives the on-time odds.
Why it matters
This shows on-time arrival probability is mostly a function of when a flight leaves the gate, so a handful of training rows is enough to reproduce the lab result.
What to watch
The model's accuracy is strong overall but drops sharply for flights near the decision boundary, so its usefulness hinges on how the threshold is set. The lab uses a 0.7 cutoff to decide whether to cancel a meeting.
WHO IT HITSData scientists and analysts who want to try Spark machine learning without managing clusters can follow this walkthrough, but they should check run mode, column mapping, and cleanup before reusing the code.
Summaries like this, in your inbox every morning.
The article is a hands-on walkthrough of the Google Skills lab "Machine Learning with Spark on Google Cloud Managed Apache Spark" (ID: GSP271), which the author recreated in a personal project rather than the lab's temporary environment. The code is based on the 06_dataproc directory in the GoogleCloudPlatform/data-science-on-gcp repository, which accompanies the O'Reilly book Data Science on the Google Cloud Platform. The training data originates from the U.S. Bureau of Transportation Statistics, and the lab adds timezone corrections to produce the CSV used here.
Running the lab outside its intended environment surfaced several discrepancies. The lab's sample values for model weights, evaluation row count, and accuracy rates do not match what the author obtained; the signs and magnitude are consistent, suggesting the lab text was written against a different data set. The lab also treats column _c24 as DIVERTED when it is actually CANCELLATION_CODE, though this does not affect results because diverted flights in this data have empty arrival delays and are filtered out. On the infrastructure side, the author notes that cluster images from version 2.2 onward default to internal IP addresses only, so initialization actions that need internet access require an external IP or Cloud NAT.
For readers, the walkthrough's value is as much in these caveats as in the model itself. Anyone adapting the sample code should verify whether Spark is actually running in local mode or distributing across workers, correct the column mapping, and remember that staging and temporary buckets persist after the cluster is deleted. The cost profile is also worth noting: the cluster bills for its entire lifetime, roughly 16 minutes here, regardless of whether jobs are running, which is likely why the article emphasizes deleting the cluster promptly.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.