Notice: This page requires JavaScript to function properly.
Please enable JavaScript in your browser settings or update your browser.
Aprende Challenge: Analyzing Sales Data with Spark SQL | Section
Data Processing with PySpark
Sección 1. Capítulo 9
single

single

Challenge: Analyzing Sales Data with Spark SQL

Desliza para mostrar el menú

Tarea

Desliza para comenzar a programar

You are given a flights dataset as a list of rows. Load it into a DataFrame, register it as a temporary view, and answer the following using spark.sql(). Store results in the specified variables:

  1. Find the top 3 routes (unique AirportFrom + AirportTo pairs) by average Length – store as a list of tuples [(origin, destination, avg_length), ...] in top_routes_by_length;
  2. For each airline, find the flight with the longest Length using a window function with row_number() – store as a DataFrame in longest_flight_per_airline with columns Airline, Flight, Length;
  3. Count how many delayed flights (Delay == 1) per DayOfWeek – store as a list of tuples [(day_of_week, count), ...] sorted by DayOfWeek ascending in delays_by_dow.

Print all results.

Solución

Switch to desktopCambia al escritorio para practicar en el mundo realContinúe desde donde se encuentra utilizando una de las siguientes opciones
¿Todo estuvo claro?

¿Cómo podemos mejorarlo?

¡Gracias por tus comentarios!

Sección 1. Capítulo 9
single

single

Pregunte a AI

expand

Pregunte a AI

ChatGPT

Pregunte lo que quiera o pruebe una de las preguntas sugeridas para comenzar nuestra charla

some-alt