Questions Tagged [Pyspark]
Ask Question the Spark Python Api (Pyspark) Exposes the Apache-Spark Programming Model to Python. 33,695 Questions 1 More Bountied 1 Unanswered Frequent Score...
The Spark Python API (PySpark) exposes the apache-spark programming model to Python.
Must Read
33,695
questions
0
votes
0
answers
7
views
Unable to launch pyspark with kafka when trying to import the kafka external package
I am trying to use pyspark with the following setup.
pyspark version 3.2.1
spark version 3.2.1
hadoop version 3.2
jdk 11
So far, this setup works well.
I am interested in working with Kafka, which ...
-2
votes
0
answers
12
views
I want to use when on pyspark dataframe but i have multiple columns df.withcolumn [closed]
I have a column "Esc_char" if it is NULL.
I want to use df.withColumm on 6-7 column and with when condition i can change only one column
Any suggestions
-1
votes
1
answer
15
views
Convert Unix timing PySpark 13 digits
I've been trying to change the UNIX date (13 digits the one on the first column on the pic) to a readable date:
from pyspark.sql import functions as F
#TRY TO CHANGE THE DATA FORMAT
sd = df....
1
vote
1
answer
19
views
PySpark structured streaming with window get earliest & latest record based upon a timestamp value
I have a structured streaming process that reads from a deltalake. The data contains values that are increasing with time. Within each window, I would like to get the (difference between the) earliest ...
-1
votes
1
answer
22
views
PySpark - Handle csv file with a 3 character quote
I just loaded a csv file to work on a new script and I've noticed that the quotes in some of the columns are """sometext""", tried setting .option('quote', '""&...
1
vote
1
answer
21
views
Create columns with upper and lower borders obtained from list
I have a DataFrame and a list of borders:
test = spark.createDataFrame(
[
(1,),
(2,),
(234,),
(0,),
(6,),
(7,),
(35,),
(46,),
...
0
votes
1
answer
18
views
How to zip-package a python module for pyspark executor?
I am parsing Protobuf binary-messages from Kafka topic, using pyspark-streaming.
My UDF is:
def parse_protobuf_from_bytes(pb_bytes):
data = contacts_pb2.Contacts()
data.ParseFromString(...
0
votes
1
answer
47
views
Adding a new column to a dataframe with a value which is based on the values from next rows
I have a dataframe as shown below,
+-----+----------+---------+-------+-------------------+
|jobid|fieldmname|new_value|coltype| createat|
+-----+----------+---------+-------+----------------...
1
vote
1
answer
9
views
DataBricks: Ingesting CSV data to a Delta Live Table in Python triggers "invalid charactres in table name" error - how to set column mapping mode?
First off, can I just say that I am learning DataBricks at the time of writing this post, so I'd like simpler, cruder solutions as well as more sophisticated ones.
I am reading a CSV file like this:
...
1
vote
2
answers
34
views
Unique element count in array column
I have this dataset with a column of array type. From this column, we need to create another column which will have list of unique elements and its counts.
Example [a,b,e,b] results should be [[b,a,e],...
0
votes
0
answers
27
views
Pyspark read file only if it exists
I have some parquet files in my hdfs directory /dir1/dir2/. The name of the files contain some timestamps but those are pretty random. For example, one file path is: /dir1/dir2/2022-06-16-03-12-36-086....
0
votes
1
answer
22
views
pyspark performance and processing time
I have 5 steps which produced df_a,df_b,df_c,df_e and df_f.
Each step generates a dataframe (df_a for instance), and persist as parquet files.
The file is used in the sequential step (df_b for example)...
-2
votes
0
answers
20
views
ThreadPoolExecutor executor script with questions
Hi the code below is what I have my question on. I am trying to fully understand/learn the code.
1.) What is (line 17) "row for row in " doing?
2.) Why are we writing "future_to_url[...
0
votes
0
answers
18
views
Spark - Large Dataset ML -- How to set properties
Im hoping to get some basic information. I am using lightgbm on spark :
and my data is large : about 40 ...
0
votes
0
answers
20
views
Null value appeared in non-nullable field
I am trying to write the following pyspark dataframe to csv by using
df_final.write.csv("v1_results/run1",header=True, emptyValue='',nullValue='')
Please help me locate on why this issue
...