이 글에서는 pandas DataFrame의 Age(나이)와 Salary(급여) 열에서 두 번째, 세 번째, 네 번째 행을 추출한 뒤, 해당 값들의 평균(mean)과 곱(product)을 계산하는 파이썬 함수를 작성하는 방법을 소개합니다.
입력 데이터
예제로 사용할 샘플 DataFrame은 다음과 같습니다.
Id Age salary 0 1 27 40000 1 2 22 25000 2 3 25 40000 3 4 23 35000 4 5 24 30000 5 6 32 30000 6 7 30 50000 7 8 28 20000 8 9 29 32000 9 10 27 23000
출력 결과
슬라이싱한 행들에 대한 평균과 곱의 계산 결과는 다음과 같습니다.
mean is Age 23.333333 salary 33333.333333 product is Age 12650 salary 35000000000000
해결 방법
이 문제는 다음과 같은 순서로 해결할 수 있습니다.
DataFrame을 정의합니다.
iloc함수를 사용하여 Age와 Salary 열의 두 번째, 세 번째, 네 번째 행을 슬라이싱하고, 그 결과를 result DataFrame에 저장합니다.result DataFrame에서
mean()으로 평균을,prod()로 곱을 계산합니다.
여기서 사용되는 인덱싱 코드는 다음과 같습니다.
df.iloc[1:4,1:]
iloc[1:4]는 인덱스 1부터 3까지, 즉 두 번째 행부터 네 번째 행까지를 의미하고, [1:]은 인덱스 1 이후의 모든 열, 즉 Age와 salary 열만 선택한다는 뜻입니다. 참고로 iloc는 위치 기반 인덱싱이므로 끝 값(인덱스 4)은 포함되지 않습니다.
예제 코드
아래 예제를 통해 실제 구현 과정을 확인해 보겠습니다.
import pandas as pd
def find_mean_prod():
data = [[1,27,40000],[2,22,25000],[3,25,40000],[4,23,35000],[5,24,30000], [6,32,30000],[7,30,50000],[8,28,20000],[9,29,32000],[10,27,23000]]
df = pd.DataFrame(data,columns=('Id','Age','salary'))
print(df)
print("slicing second,third and fourth rows of age and salary columns\n")
result = df.iloc[1:4,1:]
print("mean is\n", result.mean())
print("product is\n", result.prod())
find_mean_prod()
실행 결과
Id Age salary 0 1 27 40000 1 2 22 25000 2 3 25 40000 3 4 23 35000 4 5 24 30000 5 6 32 30000 6 7 30 50000 7 8 28 20000 8 9 29 32000 9 10 27 23000 slicing second,third and fourth rows of age and salary columns mean is Age 23.333333 salary 33333.333333 product is Age 12650 salary 35000000000000