Overview

Dataset statistics

Number of variables9
Number of observations591
Missing cells146
Missing cells (%)2.7%
Duplicate rows0
Duplicate rows (%)0.0%
Total size in memory43.4 KiB
Average record size in memory75.2 B

Variable types

Numeric3
Text1
Categorical3
DateTime2

Dataset

Description중국 위생허가 신청기업 현황(해외인증)은 사업년도, 사업자번호, 주관기업명, 지원분야, 인증명칭, 신청일, 협약체결일, 사업종료일로 구성
Author중소벤처기업부
URLhttps://www.data.go.kr/data/15024850/fileData.do

Alerts

인증명칭 has constant value ""Constant
번호 is highly overall correlated with 사업년도 and 1 other fieldsHigh correlation
사업년도 is highly overall correlated with 번호 and 1 other fieldsHigh correlation
협약체결일 is highly overall correlated with 번호 and 1 other fieldsHigh correlation
지원분야 is highly imbalanced (67.0%)Imbalance
사업종료일 has 146 (24.7%) missing valuesMissing
번호 has unique valuesUnique

Reproduction

Analysis started2023-12-12 09:50:50.281891
Analysis finished2023-12-12 09:50:52.463838
Duration2.18 seconds
Software versionydata-profiling vv4.5.1
Download configurationconfig.json

Variables

번호
Real number (ℝ)

HIGH CORRELATION  UNIQUE 

Distinct591
Distinct (%)100.0%
Missing0
Missing (%)0.0%
Infinite0
Infinite (%)0.0%
Mean296
Minimum1
Maximum591
Zeros0
Zeros (%)0.0%
Negative0
Negative (%)0.0%
Memory size5.3 KiB
2023-12-12T18:50:52.535575image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Quantile statistics

Minimum1
5-th percentile30.5
Q1148.5
median296
Q3443.5
95-th percentile561.5
Maximum591
Range590
Interquartile range (IQR)295

Descriptive statistics

Standard deviation170.75128
Coefficient of variation (CV)0.57686244
Kurtosis-1.2
Mean296
Median Absolute Deviation (MAD)148
Skewness0
Sum174936
Variance29156
MonotonicityStrictly increasing
2023-12-12T18:50:52.677648image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=50)
ValueCountFrequency (%)
1 1
 
0.2%
390 1
 
0.2%
392 1
 
0.2%
393 1
 
0.2%
394 1
 
0.2%
395 1
 
0.2%
396 1
 
0.2%
397 1
 
0.2%
398 1
 
0.2%
399 1
 
0.2%
Other values (581) 581
98.3%
ValueCountFrequency (%)
1 1
0.2%
2 1
0.2%
3 1
0.2%
4 1
0.2%
5 1
0.2%
6 1
0.2%
7 1
0.2%
8 1
0.2%
9 1
0.2%
10 1
0.2%
ValueCountFrequency (%)
591 1
0.2%
590 1
0.2%
589 1
0.2%
588 1
0.2%
587 1
0.2%
586 1
0.2%
585 1
0.2%
584 1
0.2%
583 1
0.2%
582 1
0.2%

사업년도
Real number (ℝ)

HIGH CORRELATION 

Distinct9
Distinct (%)1.5%
Missing0
Missing (%)0.0%
Infinite0
Infinite (%)0.0%
Mean2018.2589
Minimum2015
Maximum2023
Zeros0
Zeros (%)0.0%
Negative0
Negative (%)0.0%
Memory size5.3 KiB
2023-12-12T18:50:52.804364image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Quantile statistics

Minimum2015
5-th percentile2015
Q12016
median2017
Q32021
95-th percentile2022
Maximum2023
Range8
Interquartile range (IQR)5

Descriptive statistics

Standard deviation2.4665185
Coefficient of variation (CV)0.0012221021
Kurtosis-1.2720215
Mean2018.2589
Median Absolute Deviation (MAD)1
Skewness0.40540736
Sum1192791
Variance6.0837133
MonotonicityIncreasing
2023-12-12T18:50:52.936442image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=9)
ValueCountFrequency (%)
2016 169
28.6%
2017 82
13.9%
2022 69
11.7%
2020 64
 
10.8%
2021 61
 
10.3%
2015 48
 
8.1%
2018 48
 
8.1%
2019 30
 
5.1%
2023 20
 
3.4%
ValueCountFrequency (%)
2015 48
 
8.1%
2016 169
28.6%
2017 82
13.9%
2018 48
 
8.1%
2019 30
 
5.1%
2020 64
 
10.8%
2021 61
 
10.3%
2022 69
11.7%
2023 20
 
3.4%
ValueCountFrequency (%)
2023 20
 
3.4%
2022 69
11.7%
2021 61
 
10.3%
2020 64
 
10.8%
2019 30
 
5.1%
2018 48
 
8.1%
2017 82
13.9%
2016 169
28.6%
2015 48
 
8.1%

사업자번호
Real number (ℝ)

Distinct510
Distinct (%)86.3%
Missing0
Missing (%)0.0%
Infinite0
Infinite (%)0.0%
Mean3.3305734 × 109
Minimum1.0181145 × 109
Maximum8.9204016 × 109
Zeros0
Zeros (%)0.0%
Negative0
Negative (%)0.0%
Memory size5.3 KiB
2023-12-12T18:50:53.077487image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Quantile statistics

Minimum1.0181145 × 109
5-th percentile1.0883764 × 109
Q11.3181817 × 109
median2.2481447 × 109
Q35.079478 × 109
95-th percentile7.8086504 × 109
Maximum8.9204016 × 109
Range7.9022872 × 109
Interquartile range (IQR)3.7612964 × 109

Descriptive statistics

Standard deviation2.2542503 × 109
Coefficient of variation (CV)0.67683549
Kurtosis-0.51950465
Mean3.3305734 × 109
Median Absolute Deviation (MAD)9.8953852 × 108
Skewness0.8752498
Sum1.9683689 × 1012
Variance5.0816444 × 1018
MonotonicityNot monotonic
2023-12-12T18:50:53.243856image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=50)
ValueCountFrequency (%)
1208804273 4
 
0.7%
1358149079 3
 
0.5%
5158133943 3
 
0.5%
5038604232 3
 
0.5%
2308108577 3
 
0.5%
3058601397 3
 
0.5%
1398120389 3
 
0.5%
1308146868 3
 
0.5%
1208199219 3
 
0.5%
1058815091 2
 
0.3%
Other values (500) 561
94.9%
ValueCountFrequency (%)
1018114470 1
0.2%
1018157824 1
0.2%
1018651316 1
0.2%
1052432915 2
0.3%
1058690704 1
0.2%
1058716360 2
0.3%
1058723084 1
0.2%
1058724900 1
0.2%
1058764902 1
0.2%
1058807096 1
0.2%
ValueCountFrequency (%)
8920401634 1
0.2%
8808600680 1
0.2%
8790700384 1
0.2%
8788100872 1
0.2%
8758100271 2
0.3%
8728701915 1
0.2%
8698700472 1
0.2%
8648800019 1
0.2%
8638800170 1
0.2%
8538700953 1
0.2%
Distinct529
Distinct (%)89.5%
Missing0
Missing (%)0.0%
Memory size4.7 KiB
2023-12-12T18:50:53.513643image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Length

Max length29
Median length15
Mean length7.9966159
Min length2

Characters and Unicode

Total characters4726
Distinct characters365
Distinct categories9 ?
Distinct scripts3 ?
Distinct blocks3 ?
The Unicode Standard assigns character properties to each code point, which can be used to analyse textual variables.

Unique

Unique474 ?
Unique (%)80.2%

Sample

1st row(주)워랜텍
2nd row(주)덴토스
3rd row(주)월드바이오텍
4th row(주)휴비딕
5th row(주)오스테오시스
ValueCountFrequency (%)
주식회사 115
 
15.9%
8
 
1.1%
코스메틱 4
 
0.6%
주)스킨러버스코스메틱 4
 
0.6%
유씨엘(주 3
 
0.4%
제닉 3
 
0.4%
주)오스테오시스 3
 
0.4%
에이팜 3
 
0.4%
주)하스 3
 
0.4%
주)웰코스 3
 
0.4%
Other values (523) 576
79.4%
2023-12-12T18:50:53.955601image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Most occurring characters

ValueCountFrequency (%)
489
 
10.3%
) 348
 
7.4%
( 345
 
7.3%
233
 
4.9%
178
 
3.8%
148
 
3.1%
140
 
3.0%
140
 
3.0%
139
 
2.9%
124
 
2.6%
Other values (355) 2442
51.7%

Most occurring categories

ValueCountFrequency (%)
Other Letter 3801
80.4%
Close Punctuation 348
 
7.4%
Open Punctuation 346
 
7.3%
Space Separator 140
 
3.0%
Uppercase Letter 57
 
1.2%
Lowercase Letter 14
 
0.3%
Decimal Number 11
 
0.2%
Other Punctuation 7
 
0.1%
Other Symbol 2
 
< 0.1%

Most frequent character per category

Other Letter
ValueCountFrequency (%)
489
 
12.9%
233
 
6.1%
178
 
4.7%
148
 
3.9%
140
 
3.7%
139
 
3.7%
124
 
3.3%
85
 
2.2%
82
 
2.2%
79
 
2.1%
Other values (312) 2104
55.4%
Uppercase Letter
ValueCountFrequency (%)
S 7
12.3%
E 6
10.5%
A 5
 
8.8%
C 5
 
8.8%
O 4
 
7.0%
L 4
 
7.0%
T 3
 
5.3%
R 3
 
5.3%
I 3
 
5.3%
P 3
 
5.3%
Other values (9) 14
24.6%
Lowercase Letter
ValueCountFrequency (%)
l 2
14.3%
i 2
14.3%
n 2
14.3%
t 2
14.3%
d 1
7.1%
o 1
7.1%
c 1
7.1%
a 1
7.1%
r 1
7.1%
b 1
7.1%
Decimal Number
ValueCountFrequency (%)
0 3
27.3%
3 3
27.3%
9 2
18.2%
6 1
 
9.1%
1 1
 
9.1%
4 1
 
9.1%
Other Punctuation
ValueCountFrequency (%)
. 4
57.1%
& 2
28.6%
, 1
 
14.3%
Open Punctuation
ValueCountFrequency (%)
( 345
99.7%
1
 
0.3%
Close Punctuation
ValueCountFrequency (%)
) 348
100.0%
Space Separator
ValueCountFrequency (%)
140
100.0%
Other Symbol
ValueCountFrequency (%)
2
100.0%

Most occurring scripts

ValueCountFrequency (%)
Hangul 3803
80.5%
Common 852
 
18.0%
Latin 71
 
1.5%

Most frequent character per script

Hangul
ValueCountFrequency (%)
489
 
12.9%
233
 
6.1%
178
 
4.7%
148
 
3.9%
140
 
3.7%
139
 
3.7%
124
 
3.3%
85
 
2.2%
82
 
2.2%
79
 
2.1%
Other values (313) 2106
55.4%
Latin
ValueCountFrequency (%)
S 7
 
9.9%
E 6
 
8.5%
A 5
 
7.0%
C 5
 
7.0%
O 4
 
5.6%
L 4
 
5.6%
T 3
 
4.2%
R 3
 
4.2%
I 3
 
4.2%
P 3
 
4.2%
Other values (19) 28
39.4%
Common
ValueCountFrequency (%)
) 348
40.8%
( 345
40.5%
140
16.4%
. 4
 
0.5%
0 3
 
0.4%
3 3
 
0.4%
9 2
 
0.2%
& 2
 
0.2%
1
 
0.1%
6 1
 
0.1%
Other values (3) 3
 
0.4%

Most occurring blocks

ValueCountFrequency (%)
Hangul 3801
80.4%
ASCII 922
 
19.5%
None 3
 
0.1%

Most frequent character per block

Hangul
ValueCountFrequency (%)
489
 
12.9%
233
 
6.1%
178
 
4.7%
148
 
3.9%
140
 
3.7%
139
 
3.7%
124
 
3.3%
85
 
2.2%
82
 
2.2%
79
 
2.1%
Other values (312) 2104
55.4%
ASCII
ValueCountFrequency (%)
) 348
37.7%
( 345
37.4%
140
15.2%
S 7
 
0.8%
E 6
 
0.7%
A 5
 
0.5%
C 5
 
0.5%
. 4
 
0.4%
O 4
 
0.4%
L 4
 
0.4%
Other values (31) 54
 
5.9%
None
ValueCountFrequency (%)
2
66.7%
1
33.3%

지원분야
Categorical

IMBALANCE 

Distinct4
Distinct (%)0.7%
Missing0
Missing (%)0.0%
Memory size4.7 KiB
화장품
520 
의료기기
 
38
의료기기 및 용품
 
32
공산품
 
1

Length

Max length9
Median length3
Mean length3.3891709
Min length3

Unique

Unique1 ?
Unique (%)0.2%

Sample

1st row의료기기 및 용품
2nd row의료기기 및 용품
3rd row의료기기 및 용품
4th row의료기기 및 용품
5th row의료기기 및 용품

Common Values

ValueCountFrequency (%)
화장품 520
88.0%
의료기기 38
 
6.4%
의료기기 및 용품 32
 
5.4%
공산품 1
 
0.2%

Length

2023-12-12T18:50:54.103853image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram of lengths of the category

Common Values (Plot)

2023-12-12T18:50:54.218828image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
ValueCountFrequency (%)
화장품 520
79.4%
의료기기 70
 
10.7%
32
 
4.9%
용품 32
 
4.9%
공산품 1
 
0.2%

인증명칭
Categorical

CONSTANT 

Distinct1
Distinct (%)0.2%
Missing0
Missing (%)0.0%
Memory size4.7 KiB
NMPA(구.CDFA)
591 

Length

Max length12
Median length12
Mean length12
Min length12

Unique

Unique0 ?
Unique (%)0.0%

Sample

1st rowNMPA(구.CDFA)
2nd rowNMPA(구.CDFA)
3rd rowNMPA(구.CDFA)
4th rowNMPA(구.CDFA)
5th rowNMPA(구.CDFA)

Common Values

ValueCountFrequency (%)
NMPA(구.CDFA) 591
100.0%

Length

2023-12-12T18:50:54.332004image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram of lengths of the category

Common Values (Plot)

2023-12-12T18:50:54.427182image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
ValueCountFrequency (%)
nmpa(구.cdfa 591
100.0%
Distinct214
Distinct (%)36.2%
Missing0
Missing (%)0.0%
Memory size4.7 KiB
Minimum2015-03-12 00:00:00
Maximum2023-05-31 00:00:00
2023-12-12T18:50:54.545511image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:54.693077image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=50)

협약체결일
Categorical

HIGH CORRELATION 

Distinct50
Distinct (%)8.5%
Missing0
Missing (%)0.0%
Memory size4.7 KiB
2016-11-14
68 
2016-06-20
57 
2015-06-30
47 
2017-11-20
42 
2020-04-23
 
29
Other values (45)
348 

Length

Max length10
Median length10
Mean length10
Min length10

Unique

Unique14 ?
Unique (%)2.4%

Sample

1st row2015-10-23
2nd row2015-06-30
3rd row2015-06-30
4th row2015-06-30
5th row2015-06-30

Common Values

ValueCountFrequency (%)
2016-11-14 68
 
11.5%
2016-06-20 57
 
9.6%
2015-06-30 47
 
8.0%
2017-11-20 42
 
7.1%
2020-04-23 29
 
4.9%
2022-04-01 28
 
4.7%
2022-10-31 21
 
3.6%
2016-12-20 21
 
3.6%
2021-04-28 21
 
3.6%
2021-07-23 21
 
3.6%
Other values (40) 236
39.9%

Length

2023-12-12T18:50:54.804265image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram of lengths of the category
ValueCountFrequency (%)
2016-11-14 68
 
11.5%
2016-06-20 57
 
9.6%
2015-06-30 47
 
8.0%
2017-11-20 42
 
7.1%
2020-04-23 29
 
4.9%
2022-04-01 28
 
4.7%
2022-10-31 21
 
3.6%
2016-12-20 21
 
3.6%
2021-04-28 21
 
3.6%
2021-07-23 21
 
3.6%
Other values (40) 236
39.9%

사업종료일
Date

MISSING 

Distinct202
Distinct (%)45.4%
Missing146
Missing (%)24.7%
Memory size4.7 KiB
Minimum2016-06-02 00:00:00
Maximum2023-08-24 00:00:00
2023-12-12T18:50:54.947600image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:55.107150image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=50)

Interactions

2023-12-12T18:50:51.572413image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:50.799634image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:51.194322image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:51.682806image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:50.937000image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:51.309315image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:51.824132image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:51.066943image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T18:50:51.448121image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Correlations

2023-12-12T18:50:55.192156image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
번호사업년도사업자번호지원분야협약체결일
번호1.0000.9060.2260.4670.995
사업년도0.9061.0000.1520.4911.000
사업자번호0.2260.1521.0000.0000.321
지원분야0.4670.4910.0001.0000.682
협약체결일0.9951.0000.3210.6821.000
2023-12-12T18:50:55.301779image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
지원분야협약체결일
지원분야1.0000.396
협약체결일0.3961.000
2023-12-12T18:50:55.388701image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
번호사업년도사업자번호지원분야협약체결일
번호1.0000.9840.2010.3050.856
사업년도0.9841.0000.2000.2430.964
사업자번호0.2010.2001.0000.0000.104
지원분야0.3050.2430.0001.0000.396
협약체결일0.8560.9640.1040.3961.000

Missing values

2023-12-12T18:50:52.281851image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
A simple visualization of nullity by column.
2023-12-12T18:50:52.411401image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Nullity matrix is a data-dense display which lets you quickly visually pick out patterns in data completion.

Sample

번호사업년도사업자번호주관기업명지원분야인증명칭신청일협약체결일사업종료일
0120152158618799(주)워랜텍의료기기 및 용품NMPA(구.CDFA)2015-08-242015-10-232016-12-14
1220155048137935(주)덴토스의료기기 및 용품NMPA(구.CDFA)2015-04-032015-06-302017-12-29
2320151298636481(주)월드바이오텍의료기기 및 용품NMPA(구.CDFA)2015-04-022015-06-302017-09-13
3420151238175370(주)휴비딕의료기기 및 용품NMPA(구.CDFA)2015-04-012015-06-302017-11-06
4520151208199219(주)오스테오시스의료기기 및 용품NMPA(구.CDFA)2015-03-312015-06-302017-08-30
5620151238601005(주)참메드의료기기 및 용품NMPA(구.CDFA)2015-03-302015-06-302017-06-29
6720153178126590(주)더아이엔지메디칼의료기기 및 용품NMPA(구.CDFA)2015-03-232015-06-302016-06-02
7820156038158163주식회사 네오실의료기기 및 용품NMPA(구.CDFA)2015-03-172015-06-302017-09-14
8920156218195152(주)포셀화장품NMPA(구.CDFA)2015-04-032015-06-302018-01-19
91020152118622189클리오화장품NMPA(구.CDFA)2015-04-032015-06-302017-12-29
번호사업년도사업자번호주관기업명지원분야인증명칭신청일협약체결일사업종료일
58158220236988601552티핏클래스(주)화장품NMPA(구.CDFA)202303292023-06-02<NA>
58258320238788100872(주)기베스트화장품NMPA(구.CDFA)202305302023-07-21<NA>
58358420238538700953(주)스타스테크화장품NMPA(구.CDFA)202305312023-07-21<NA>
58458520231438120262(주)아침해의료기의료기기NMPA(구.CDFA)202305312023-07-21<NA>
58558620232048612193(주)와이제이비앤화장품NMPA(구.CDFA)202305302023-07-21<NA>
58658720231298636481(주)월드바이오텍의료기기NMPA(구.CDFA)202305312023-07-21<NA>
58758820233328100885(주)쥬네뷰화장품NMPA(구.CDFA)202305192023-07-21<NA>
58858920237118700829(주)코스모어플러스화장품NMPA(구.CDFA)202305302023-07-21<NA>
58959020231228701163오션스바이오(주)의료기기NMPA(구.CDFA)202305302023-07-21<NA>
59059120238920401634주아빛화장품NMPA(구.CDFA)202305312023-07-21<NA>