Overview

Dataset statistics

Number of variables6
Number of observations500
Missing cells0
Missing cells (%)0.0%
Duplicate rows0
Duplicate rows (%)0.0%
Total size in memory24.5 KiB
Average record size in memory50.3 B

Variable types

Numeric2
Categorical3
Text1

Dataset

Description샘플 데이터
Author다음소프트
URLhttps://bigdata.seoul.go.kr/data/selectSampleData.do?sample_data_seq=57

Alerts

수집소스(SOURCE) is highly imbalanced (88.2%)Imbalance

Reproduction

Analysis started2023-12-10 14:53:50.119616
Analysis finished2023-12-10 14:53:52.291793
Duration2.17 seconds
Software versionydata-profiling vv4.5.1
Download configurationconfig.json

Variables

DOC_DATE(DATE)
Real number (ℝ)

Distinct400
Distinct (%)80.0%
Missing0
Missing (%)0.0%
Infinite0
Infinite (%)0.0%
Mean20180170
Minimum20170109
Maximum20191226
Zeros0
Zeros (%)0.0%
Negative0
Negative (%)0.0%
Memory size4.5 KiB
2023-12-10T23:53:52.420528image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Quantile statistics

Minimum20170109
5-th percentile20170304
Q120170927
median20180622
Q320190223
95-th percentile20191018
Maximum20191226
Range21117
Interquartile range (IQR)19295.75

Descriptive statistics

Standard deviation7964.2364
Coefficient of variation (CV)0.00039465656
Kurtosis-1.4130177
Mean20180170
Median Absolute Deviation (MAD)9688.5
Skewness0.092072877
Sum1.0090085 × 1010
Variance63429062
MonotonicityNot monotonic
2023-12-10T23:53:52.590185image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=50)
ValueCountFrequency (%)
20171101 5
 
1.0%
20190106 3
 
0.6%
20181203 3
 
0.6%
20180101 3
 
0.6%
20190703 3
 
0.6%
20180707 3
 
0.6%
20190820 3
 
0.6%
20191209 3
 
0.6%
20170927 3
 
0.6%
20190330 3
 
0.6%
Other values (390) 468
93.6%
ValueCountFrequency (%)
20170109 2
0.4%
20170119 1
 
0.2%
20170121 2
0.4%
20170122 2
0.4%
20170123 1
 
0.2%
20170130 1
 
0.2%
20170131 1
 
0.2%
20170203 1
 
0.2%
20170204 1
 
0.2%
20170206 3
0.6%
ValueCountFrequency (%)
20191226 2
0.4%
20191223 1
 
0.2%
20191222 1
 
0.2%
20191216 1
 
0.2%
20191214 1
 
0.2%
20191211 1
 
0.2%
20191210 1
 
0.2%
20191209 3
0.6%
20191206 1
 
0.2%
20191129 1
 
0.2%

수집소스(SOURCE)
Categorical

IMBALANCE 

Distinct2
Distinct (%)0.4%
Missing0
Missing (%)0.0%
Memory size4.0 KiB
블로그커뮤니티
492 
트위터
 
8

Length

Max length7
Median length7
Mean length6.936
Min length3

Unique

Unique0 ?
Unique (%)0.0%

Sample

1st row블로그커뮤니티
2nd row블로그커뮤니티
3rd row블로그커뮤니티
4th row블로그커뮤니티
5th row블로그커뮤니티

Common Values

ValueCountFrequency (%)
블로그커뮤니티 492
98.4%
트위터 8
 
1.6%

Length

2023-12-10T23:53:52.776147image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram of lengths of the category

Common Values (Plot)

2023-12-10T23:53:52.895902image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
ValueCountFrequency (%)
블로그커뮤니티 492
98.4%
트위터 8
 
1.6%
Distinct217
Distinct (%)43.4%
Missing0
Missing (%)0.0%
Memory size4.0 KiB
2023-12-10T23:53:53.136137image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Length

Max length11
Median length9
Mean length3.29
Min length2

Characters and Unicode

Total characters1645
Distinct characters192
Distinct categories3 ?
Distinct scripts3 ?
Distinct blocks2 ?
The Unicode Standard assigns character properties to each code point, which can be used to analyse textual variables.

Unique

Unique120 ?
Unique (%)24.0%

Sample

1st row샤로수길
2nd row가로수길
3rd row송파
4th row영등포역
5th row파르나스몰
ValueCountFrequency (%)
서울 34
 
6.8%
용산 14
 
2.8%
강남 11
 
2.2%
강남역 11
 
2.2%
명동 9
 
1.8%
용산cgv 9
 
1.8%
신촌 9
 
1.8%
왕십리 8
 
1.6%
을지로 7
 
1.4%
상봉 7
 
1.4%
Other values (207) 381
76.2%
2023-12-10T23:53:53.678287image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Most occurring characters

ValueCountFrequency (%)
104
 
6.3%
58
 
3.5%
51
 
3.1%
50
 
3.0%
45
 
2.7%
44
 
2.7%
36
 
2.2%
35
 
2.1%
c 34
 
2.1%
g 32
 
1.9%
Other values (182) 1156
70.3%

Most occurring categories

ValueCountFrequency (%)
Other Letter 1540
93.6%
Lowercase Letter 103
 
6.3%
Decimal Number 2
 
0.1%

Most frequent character per category

Other Letter
ValueCountFrequency (%)
104
 
6.8%
58
 
3.8%
51
 
3.3%
50
 
3.2%
45
 
2.9%
44
 
2.9%
36
 
2.3%
35
 
2.3%
30
 
1.9%
28
 
1.8%
Other values (174) 1059
68.8%
Lowercase Letter
ValueCountFrequency (%)
c 34
33.0%
g 32
31.1%
v 32
31.1%
f 2
 
1.9%
i 2
 
1.9%
n 1
 
1.0%
Decimal Number
ValueCountFrequency (%)
3 1
50.0%
6 1
50.0%

Most occurring scripts

ValueCountFrequency (%)
Hangul 1540
93.6%
Latin 103
 
6.3%
Common 2
 
0.1%

Most frequent character per script

Hangul
ValueCountFrequency (%)
104
 
6.8%
58
 
3.8%
51
 
3.3%
50
 
3.2%
45
 
2.9%
44
 
2.9%
36
 
2.3%
35
 
2.3%
30
 
1.9%
28
 
1.8%
Other values (174) 1059
68.8%
Latin
ValueCountFrequency (%)
c 34
33.0%
g 32
31.1%
v 32
31.1%
f 2
 
1.9%
i 2
 
1.9%
n 1
 
1.0%
Common
ValueCountFrequency (%)
3 1
50.0%
6 1
50.0%

Most occurring blocks

ValueCountFrequency (%)
Hangul 1540
93.6%
ASCII 105
 
6.4%

Most frequent character per block

Hangul
ValueCountFrequency (%)
104
 
6.8%
58
 
3.8%
51
 
3.3%
50
 
3.2%
45
 
2.9%
44
 
2.9%
36
 
2.3%
35
 
2.3%
30
 
1.9%
28
 
1.8%
Other values (174) 1059
68.8%
ASCII
ValueCountFrequency (%)
c 34
32.4%
g 32
30.5%
v 32
30.5%
f 2
 
1.9%
i 2
 
1.9%
3 1
 
1.0%
6 1
 
1.0%
n 1
 
1.0%

행정구(GU_NM)
Categorical

Distinct25
Distinct (%)5.0%
Missing0
Missing (%)0.0%
Memory size4.0 KiB
강남구
67 
종로구
52 
용산구
51 
서울
44 
마포구
43 
Other values (20)
243 

Length

Max length4
Median length3
Mean length2.93
Min length2

Unique

Unique0 ?
Unique (%)0.0%

Sample

1st row서울
2nd row광진구
3rd row중구
4th row송파구
5th row마포구

Common Values

ValueCountFrequency (%)
강남구 67
13.4%
종로구 52
 
10.4%
용산구 51
 
10.2%
서울 44
 
8.8%
마포구 43
 
8.6%
중구 32
 
6.4%
송파구 22
 
4.4%
영등포구 21
 
4.2%
광진구 20
 
4.0%
서대문구 18
 
3.6%
Other values (15) 130
26.0%

Length

2023-12-10T23:53:53.929844image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram of lengths of the category
ValueCountFrequency (%)
강남구 67
13.4%
종로구 52
 
10.4%
용산구 51
 
10.2%
서울 44
 
8.8%
마포구 43
 
8.6%
중구 32
 
6.4%
송파구 22
 
4.4%
영등포구 21
 
4.2%
광진구 20
 
4.0%
서대문구 18
 
3.6%
Other values (15) 130
26.0%
Distinct20
Distinct (%)4.0%
Missing0
Missing (%)0.0%
Memory size4.0 KiB
영화
196 
cgv
67 
영화관
57 
극장
30 
메가박스
28 
Other values (15)
122 

Length

Max length5
Median length4
Mean length2.802
Min length2

Unique

Unique1 ?
Unique (%)0.2%

Sample

1st row영화
2nd row롯데시네마
3rd row영화
4th row영화관람
5th row영화관에서

Common Values

ValueCountFrequency (%)
영화 196
39.2%
cgv 67
 
13.4%
영화관 57
 
11.4%
극장 30
 
6.0%
메가박스 28
 
5.6%
롯데시네마 26
 
5.2%
영화보기 17
 
3.4%
개봉 16
 
3.2%
매표소 15
 
3.0%
영화관에서 9
 
1.8%
Other values (10) 39
 
7.8%

Length

2023-12-10T23:53:54.137880image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram of lengths of the category
ValueCountFrequency (%)
영화 196
39.2%
cgv 67
 
13.4%
영화관 57
 
11.4%
극장 30
 
6.0%
메가박스 28
 
5.6%
롯데시네마 26
 
5.2%
영화보기 17
 
3.4%
개봉 16
 
3.2%
매표소 15
 
3.0%
영화보러 9
 
1.8%
Other values (10) 39
 
7.8%

FREQ(FREQ)
Real number (ℝ)

Distinct9
Distinct (%)1.8%
Missing0
Missing (%)0.0%
Infinite0
Infinite (%)0.0%
Mean1.31
Minimum1
Maximum18
Zeros0
Zeros (%)0.0%
Negative0
Negative (%)0.0%
Memory size4.5 KiB
2023-12-10T23:53:54.269001image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Quantile statistics

Minimum1
5-th percentile1
Q11
median1
Q31
95-th percentile2
Maximum18
Range17
Interquartile range (IQR)0

Descriptive statistics

Standard deviation1.4359749
Coefficient of variation (CV)1.0961641
Kurtosis82.848224
Mean1.31
Median Absolute Deviation (MAD)0
Skewness8.6102671
Sum655
Variance2.062024
MonotonicityNot monotonic
2023-12-10T23:53:54.393293image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=9)
ValueCountFrequency (%)
1 428
85.6%
2 53
 
10.6%
3 8
 
1.6%
4 5
 
1.0%
15 2
 
0.4%
14 1
 
0.2%
10 1
 
0.2%
5 1
 
0.2%
18 1
 
0.2%
ValueCountFrequency (%)
1 428
85.6%
2 53
 
10.6%
3 8
 
1.6%
4 5
 
1.0%
5 1
 
0.2%
10 1
 
0.2%
14 1
 
0.2%
15 2
 
0.4%
18 1
 
0.2%
ValueCountFrequency (%)
18 1
 
0.2%
15 2
 
0.4%
14 1
 
0.2%
10 1
 
0.2%
5 1
 
0.2%
4 5
 
1.0%
3 8
 
1.6%
2 53
 
10.6%
1 428
85.6%

Interactions

2023-12-10T23:53:50.945829image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-10T23:53:50.493846image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-10T23:53:51.171983image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-10T23:53:50.611181image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Correlations

2023-12-10T23:53:54.495343image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
DOC_DATE(DATE)수집소스(SOURCE)행정구(GU_NM)세부키워드(KEYWORD_DETAIL)FREQ(FREQ)
DOC_DATE(DATE)1.0000.1560.0000.0000.033
수집소스(SOURCE)0.1561.0000.0000.0460.314
행정구(GU_NM)0.0000.0001.0000.3150.000
세부키워드(KEYWORD_DETAIL)0.0000.0460.3151.0000.000
FREQ(FREQ)0.0330.3140.0000.0001.000
2023-12-10T23:53:54.613538image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
세부키워드(KEYWORD_DETAIL)수집소스(SOURCE)행정구(GU_NM)
세부키워드(KEYWORD_DETAIL)1.0000.0340.089
수집소스(SOURCE)0.0341.0000.000
행정구(GU_NM)0.0890.0001.000
2023-12-10T23:53:54.738168image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
DOC_DATE(DATE)FREQ(FREQ)수집소스(SOURCE)행정구(GU_NM)세부키워드(KEYWORD_DETAIL)
DOC_DATE(DATE)1.000-0.0170.0840.0130.000
FREQ(FREQ)-0.0171.0000.3350.0000.000
수집소스(SOURCE)0.0840.3351.0000.0000.034
행정구(GU_NM)0.0130.0000.0001.0000.089
세부키워드(KEYWORD_DETAIL)0.0000.0000.0340.0891.000

Missing values

2023-12-10T23:53:51.608259image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
A simple visualization of nullity by column.
2023-12-10T23:53:52.165189image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Nullity matrix is a data-dense display which lets you quickly visually pick out patterns in data completion.

Sample

DOC_DATE(DATE)수집소스(SOURCE)행정동(DONG_NM)행정구(GU_NM)세부키워드(KEYWORD_DETAIL)FREQ(FREQ)
020190625블로그커뮤니티샤로수길서울영화1
120170731블로그커뮤니티가로수길광진구롯데시네마1
220190731블로그커뮤니티송파중구영화1
320180108블로그커뮤니티영등포역송파구영화관람1
420170121블로그커뮤니티파르나스몰마포구영화관에서1
520180711블로그커뮤니티한강영등포구조조영화1
620190714블로그커뮤니티녹사평마포구롯데시네마1
720191029블로그커뮤니티월드컵경기장용산구영화1
820190314블로그커뮤니티상암동종로구영화보기3
920180607블로그커뮤니티코엑스메가박스광진구cgv2
DOC_DATE(DATE)수집소스(SOURCE)행정동(DONG_NM)행정구(GU_NM)세부키워드(KEYWORD_DETAIL)FREQ(FREQ)
49020170206블로그커뮤니티망원역마포구영화관1
49120180718블로그커뮤니티롯데월드영등포구영화1
49220190703블로그커뮤니티영등포동서울영화1
49320180415블로그커뮤니티영등포동송파구영화1
49420180506블로그커뮤니티이수역서대문구영화18
49520170615블로그커뮤니티월드컵경기장중구cgv1
49620181205블로그커뮤니티국립박물관종로구롯데시네마1
49720171025블로그커뮤니티신도림롯데시네마은평구영화관2
49820190901블로그커뮤니티이태원중랑구영화관람2
49920171025블로그커뮤니티을지로양천구롯데시네마1