Questions Tagged [Pyspark]

The Spark Python API (PySpark) exposes the apache-spark programming model to Python.

33,695 questions
0 votes
0 answers
7 views

Unable to launch pyspark with kafka when trying to import the kafka external package

I am trying to use pyspark with the following setup. pyspark version 3.2.1 spark version 3.2.1 hadoop version 3.2 jdk 11 So far, this setup works well. I am interested in working with Kafka, which ...
  • 1,164
-2 votes
0 answers
12 views

I want to use when on pyspark dataframe but i have multiple columns df.withcolumn [closed]

I have a column "Esc_char" if it is NULL. I want to use df.withColumm on 6-7 column and with when condition i can change only one column Any suggestions
-1 votes
1 answer
15 views

Convert Unix timing PySpark 13 digits

I've been trying to change the UNIX date (13 digits the one on the first column on the pic) to a readable date: from pyspark.sql import functions as F #TRY TO CHANGE THE DATA FORMAT sd = df....
1 vote
1 answer
19 views

PySpark structured streaming with window get earliest & latest record based upon a timestamp value

I have a structured streaming process that reads from a deltalake. The data contains values that are increasing with time. Within each window, I would like to get the (difference between the) earliest ...
-1 votes
1 answer
22 views

PySpark - Handle csv file with a 3 character quote

I just loaded a csv file to work on a new script and I've noticed that the quotes in some of the columns are """sometext""", tried setting .option('quote', '""&...
1 vote
1 answer
21 views

Create columns with upper and lower borders obtained from list

I have a DataFrame and a list of borders: test = spark.createDataFrame( [ (1,), (2,), (234,), (0,), (6,), (7,), (35,), (46,), ...
0 votes
1 answer
18 views

How to zip-package a python module for pyspark executor?

I am parsing Protobuf binary-messages from Kafka topic, using pyspark-streaming. My UDF is: def parse_protobuf_from_bytes(pb_bytes): data = contacts_pb2.Contacts() data.ParseFromString(...
  • 1,025
0 votes
1 answer
47 views

Adding a new column to a dataframe with a value which is based on the values from next rows

I have a dataframe as shown below, +-----+----------+---------+-------+-------------------+ |jobid|fieldmname|new_value|coltype| createat| +-----+----------+---------+-------+----------------...
1 vote
1 answer
9 views

DataBricks: Ingesting CSV data to a Delta Live Table in Python triggers "invalid charactres in table name" error - how to set column mapping mode?

First off, can I just say that I am learning DataBricks at the time of writing this post, so I'd like simpler, cruder solutions as well as more sophisticated ones. I am reading a CSV file like this: ...
1 vote
2 answers
34 views

Unique element count in array column

I have this dataset with a column of array type. From this column, we need to create another column which will have list of unique elements and its counts. Example [a,b,e,b] results should be [[b,a,e],...
0 votes
0 answers
27 views

Pyspark read file only if it exists

I have some parquet files in my hdfs directory /dir1/dir2/. The name of the files contain some timestamps but those are pretty random. For example, one file path is: /dir1/dir2/2022-06-16-03-12-36-086....
0 votes
1 answer
22 views

pyspark performance and processing time

I have 5 steps which produced df_a,df_b,df_c,df_e and df_f. Each step generates a dataframe (df_a for instance), and persist as parquet files. The file is used in the sequential step (df_b for example)...
  • 561
-2 votes
0 answers
20 views

ThreadPoolExecutor executor script with questions

Hi the code below is what I have my question on. I am trying to fully understand/learn the code. 1.) What is (line 17) "row for row in " doing? 2.) Why are we writing "future_to_url[...
0 votes
0 answers
18 views

Spark - Large Dataset ML -- How to set properties

Im hoping to get some basic information. I am using lightgbm on spark : and my data is large : about 40 ...
  • 1,716
0 votes
0 answers
20 views

Null value appeared in non-nullable field

I am trying to write the following pyspark dataframe to csv by using df_final.write.csv("v1_results/run1",header=True, emptyValue='',nullValue='') Please help me locate on why this issue ...

15 30 50 per page
1
2 3 4 5
…
2247
James H. Sterling

James H. Sterling

Environmental Science & Climate Journalist

James Sterling reports on renewable energy developments, climate policy, ecological conservation, and green tech innovations around the globe.

Share this article
Twitter Facebook Pinterest