International Edition
Latest News
Technology

Nvidia Accelerates Apache Spark 3.0 with GPU Support to Cut Cloud Costs

NVIDIA announced native GPU acceleration for Apache Spark 3.0, enabling enterprises to speed up data analytics pipelines from extract, transform, and load (ETL) processes to AI model training without requiring code changes to existing applications, according to a…

Nvidia Accelerates Apache Spark 3.0 with GPU Support to Cut Cloud Costs

NVIDIA announced native GPU acceleration for Apache Spark 3.0, enabling enterprises to speed up data analytics pipelines from extract, transform, and load (ETL) processes to AI model training without requiring code changes to existing applications, according to a company announcement.

Apache Spark 3.0 GPU Integration and Performance Gains

According to NVIDIA, the open-source collaboration brings end-to-end GPU acceleration to Apache Spark 3.0, allowing data scientists to run ETL workloads using SQL database operations directly on GPUs. Previously, workloads often ran as separate processes on distinct infrastructure. The new architecture enables teams to execute AI model training on the same Spark cluster used for data preparation.

Adobe participated in preview testing of Spark 3.0 running on Databricks. According to William Yan, senior director of Machine Learning at Adobe, the company achieved a 7x performance improvement and a 90 percent cost savings in initial tests using GPU-accelerated data analytics for product development within Adobe Experience Cloud.

Databricks Collaboration and RAPIDS Software Suite

Apache Spark was originally created by the founders of Databricks, whose cloud platform runs on over 1 million virtual machines daily. NVIDIA and Databricks integrated the RAPIDS software suite with Databricks to optimize machine learning and data science workloads across multiple industries, including finance, healthcare, and retail.

Matei Zaharia, chief technologist at Databricks and original creator of Apache Spark, stated in the announcement that the collaboration improves performance through RAPIDS optimizations for Apache Spark 3.0, resulting in faster data pipelines, model training, and scoring for data engineers and data scientists.

Enterprise Deployments at the Internal Revenue Service

The Internal Revenue Service (IRS) adopted Apache Spark 3.0 accelerated by GPUs via the Cloudera Data Platform (CDP) to process large datasets. Data scientist Deborah Tylor utilized the setup to comb through a dataset exceeding 3 terabytes for fraud patterns after standard CPU-based jobs failed repeatedly.

NVIDIA Accelerates Apache Spark
Photo: nvidianews.nvidia.com

Following code adjustments by NVIDIA data scientists to handle complex data structures via the RAPIDS software interface, the IRS team achieved over 20x speed improvements at half the cost for data engineering and data science workflows, according to Joe Ansaldi, technical branch chief of the research and applied analytics and statistics division at the IRS.

About the author: Anika Shah - Technology

MSc in Computer Science, senior reporter. Anika focuses on AI ethics, cybersecurity, and emerging hardware—frequently moderating panels at CES and Web Summit. “Anika Shah decodes tech breakthroughs and startup disruption shaping tomorrow’s digital landscape.”