对PySpark数据帧中列的所有值进行切片

+-------------+------+ | studentID|gender| +-------------+------+ |1901000200 | M| |1901000500 | M| |1901000500 | M| |1901000500 | M| |1901000500 | M| +-------------+------+

+-------------+------+ | studentID|gender| +-------------+------+ | 1000200 | M| | 1000500 | M| | 1000500 | M| | 1000500 | M| | 1000500 | M| +-------------+------+

students_data = students_data.withColumn('studentID',F.lit(students_data["studentID"][2:])) TypeError: startPos and length must be the same type. Got <class 'int'> and <class 'NoneType'>, respectively.

1条回答

网友

1楼 · 发布于 2024-05-26 11:55:17

from pyspark.sql import functions as F

# replicating the sample data from the OP.
students_data = sqlContext.createDataFrame(
[[1901000200,'M'],
[1901000500,'M'],
[1901000500,'M'],
[1901000500,'M'],
[1901000500,'M']],
["studentid", "gender"])

# unlike a simple python list transformation - we need to define the last position in the transform
# in case you aren't sure about the length one can define a random large number say 10k.
students_data = students_data.withColumn(
  'studentID',
  F.lit(students_data["studentID"][4:10000]).cast("string"))

students_data.show()

输出：

+    -+   +
|studentID|gender|
+    -+   +
|  1000200|     M|
|  1000500|     M|
|  1000500|     M|
|  1000500|     M|
|  1000500|     M|
+    -+   +

相关问题更多 >

编程相关推荐

热门问题

热门文章