Python Azure Databrick: 'DataFrame' object does not support item assignment

Question

I am working on Azure Databrick. I run the python script on the notebook and gets the data from the SQL. I tried to split the date time column into date and time columns. Here is the syntax of python:

    pushdown_query = "(SELECT * FROM STAGE.OutagesAndInterruptions) int_alias"
    df = spark.read.jdbc(url=jdbcUrl, table=pushdown_query, properties=connectionProperties)

    df['INTERRUPTION_DATE']=df['INTERRUPTION_TIME'].dt.date

df['INTERRUPTION_TIME'] looks like:

+-------------------+
|  INTERRUPTION_TIME|
+-------------------+
|1997-05-12 09:57:00|
|1998-03-08 13:00:00|
|1998-02-26 13:00:00|
|1998-02-26 13:00:00|
|1998-03-03 10:04:00|
|1998-05-20 09:27:00|
|1998-11-21 08:51:00|
|1998-11-27 08:44:00|
|1998-10-19 01:19:00|
|1998-10-19 01:44:00|
|2000-03-13 07:00:00|
|2000-03-19 07:30:00|
|2000-08-04 12:55:00|
|2002-09-30 18:11:00|
|2002-09-30 18:11:00|
|2002-05-06 09:22:00|
|2002-01-16 13:15:00|
|2003-01-08 15:46:00|
|2003-02-04 10:25:00|
|2003-02-04 10:25:00|
+-------------------+

When I ran the code, it throws an error message:

TypeError: 'DataFrame' object does not support item assignment
---------------------------------------------------------------------------
TypeError                                 Traceback (most recent call last)
<command-2244924718685919> in <module>
----> 1 df['INTERRUPTION_DATE']=df['INTERRUPTION_TIME'].dt.date

TypeError: 'DataFrame' object does not support item assignment

Could we create new columns in the data frame on Data frame? How can we create new columns on the data frame in Azure data bricks?

Please share the entire error message, as well as a minimal reproducible example. — AMC
– AMC, Commented Mar 5, 2020 at 23:01
@AMC I have added the entire error message and some context. — user2293224
– user2293224, Commented Mar 5, 2020 at 23:07
Just wanted to add that my case was that an assigment worked for Pandas dataframes but not for pyspark dataframes, so transforming from to Pandas did the trick for me, in case it helps anyone else — monkey intern
– monkey intern, Commented Nov 19, 2020 at 9:34

Maria Nazari · Accepted Answer · 2020-03-05 23:53:45Z

5

This should work

from pyspark.sql.types import DateType


df2 = df.withColumn('INTERRUPTION_DATE', ,df['INTERRUPTION_TIME'].cast(DateType()))

Edit After Comment:

from pyspark.sql.functions import date_format

df.select(date_format('INTERRUPTION_TIME', 'M/d/yyyy').alias('INTERRUPTION_DATE'),
          date_format('INTERRUPTION_TIME', 'h:m:s a').alias('TIME'))

edited Mar 5, 2020 at 23:53

answered Mar 5, 2020 at 23:12

Maria Nazari

6881 gold badge9 silver badges27 bronze badges

Sign up to request clarification or add additional context in comments.

1 Comment

user2293224 Over a year ago

So if I want to save the time from the same column into new column, do I need to change the cast from DateType to TimeType?

Collectives™ on Stack Overflow

Python Azure Databrick: 'DataFrame' object does not support item assignment

1 Answer 1

1 Comment

Your Answer

Hot Network Questions

Collectives™ on Stack Overflow

1 Answer 1

1 Comment

Your Answer

Sign up or log in

Post as a guest

Related