如何用datafram中第一行和对应行之间的列的平均值填充特定值

0条回答

网友

1楼 · 发布于 2024-04-23 21:44:40

IIUC公司

def f(x):
    for z in range(x.size):
        if x[z] == 0: x[z] = np.mean(x[:z+1])
    return x

df.astype(float).apply(f)

    A   B       C           D   E
0   1.0 2.00    3.000000    0.0 2.0
1   2.0 1.00    7.000000    1.0 1.0
2   3.0 4.00    3.333333    3.0 1.0
3   1.5 1.75    3.000000    4.0 3.0

网友

2楼 · 发布于 2024-04-23 21:44:40

如果每列有多个0，则需要前面的mean值，这是一个主要问题，因此创建矢量化解决方案确实有问题：

def f(x):
    for i, v in enumerate(x):
        if v == 0: 
            x.iloc[i] = x.iloc[:i+1].mean()
    return x

df1 = df.astype(float).apply(f)
print (df1)

     A     B         C    D    E
0  1.0  2.00  3.000000  0.0  2.0
1  2.0  1.00  7.000000  1.0  1.0
2  3.0  4.00  3.333333  3.0  1.0
3  1.5  1.75  3.000000  4.0  3.0

更好的解决方案：

#create indices of zero values to helper DataFrame
a, b = np.where(df.values == 0)
df1 = pd.DataFrame({'rows':a, 'cols':b})
#for first row is not necessary count means
df1 = df1[df1['rows'] != 0]
print (df1)
   rows  cols
1     1     1
2     2     2
3     2     4
4     3     0
5     3     1

#loop by each row of helper df and assign means
for i in df1.itertuples():
    df.iloc[i.rows, i.cols] = df.iloc[:i.rows+1, i.cols].mean()

print (df)
     A     B         C  D    E
0  1.0  2.00  3.000000  0  2.0
1  2.0  1.00  7.000000  1  1.0
2  3.0  4.00  3.333333  3  1.0
3  1.5  1.75  3.000000  4  3.0

另一个类似的解决方案（所有对的mean）：

for i, j in zip(*np.where(df.values == 0)):
    df.iloc[i, j] = df.iloc[:i+1, j].mean()
print (df)

     A     B         C    D    E
0  1.0  2.00  3.000000  0.0  2.0
1  2.0  1.00  7.000000  1.0  1.0
2  3.0  4.00  3.333333  3.0  1.0
3  1.5  1.75  3.000000  4.0  3.0

相关问题更多 >

编程相关推荐

热门问题

热门文章

如何用datafram中第一行和对应行之间的列的平均值填充特定值

相关问题 更多 >

编程相关推荐

热门问题

热门文章

相关问题更多 >