在数据帧上使用groupby和lambda函数时保留NaN值

ChildID MotherID preDiabetes 0 20 455 No 1 20 455 Not documented 2 13 102 NaN 3 13 102 Yes 4 702 946 No 5 82 571 No 6 82 571 Yes 7 82 571 Not documented 8 60 530 NaN

3条回答

网友

1楼 · 编辑于 2024-06-16 09:31:37

您可以尝试：

import pandas as pd
import numpy as np
import io

data_string = """ChildID,MotherID,preDiabetes
20,455,No
20,455,Not documented
13,102,NaN
13,102,Yes
702,946,No
82,571,No
82,571,Yes
82,571,Not documented
60,530,NaN
"""

data = io.StringIO(data_string)
df = pd.read_csv(data, sep=',', na_values=['NaN'])
df.fillna('no_value', inplace=True)
df = df.groupby(['MotherID', 'ChildID'])['preDiabetes'].apply(
         lambda x: 'Yes' if 'Yes' in x.values else (np.NaN if 'no_value' in x.values.all() else 'No'))
df

结果:

MotherID  ChildID
102       13         Yes
455       20          No
530       60         NaN
571       82         Yes
946       702         No
Name: preDiabetes, dtype: object

网友

2楼 · 编辑于 2024-06-16 09:31:37

您可以使用自定义函数执行以下操作：

def func(s):

    if s.eq('Yes').any():
        return 'Yes'
    elif s.isna().all():
        return np.nan
    else:
        return 'No'

df  = (df
       .groupby(['ChildID', 'MotherID'])
       .agg({'preDiabetes': func}))

print(df)

   ChildID  MotherID preDiabetes
0       13       102         Yes
1       20       455          No
2       60       530         NaN
3       82       571         Yes
4      702       946          No

网友

3楼 · 编辑于 2024-06-16 09:31:37

尝试：

df['preDiabetes']=df['preDiabetes'].map({'Yes': 1, 'No': 0}).fillna(-1)

df=df.groupby(['MotherID', 'ChildID'])['preDiabetes'].max().map({1: 'Yes', 0: 'No', -1: 'NaN'}).reset_index()

第一行将preDiabetes格式化为数字，假设NaN是除Yes或No（由-1表示）之外的所有内容

第二行假设至少有一个preDiabetes是Yes-我们为组输出Yes。假设我们有No和NaN，我们输出No。假设所有的都是NaN，我们输出NaN

产出：

>>> df

   MotherID  ChildID preDiabetes
0       102       13         Yes
1       455       20          No
2       530       60         NaN
3       571       82         Yes
4       946      702          No

相关问题更多 >

编程相关推荐

热门问题

热门文章

在数据帧上使用groupby和lambda函数时保留NaN值

相关问题 更多 >

编程相关推荐

热门问题

热门文章

相关问题更多 >