Overview

Dataset statistics

Number of variables5
Number of observations596
Missing cells0
Missing cells (%)0.0%
Duplicate rows8
Duplicate rows (%)1.3%
Total size in memory24.6 KiB
Average record size in memory42.2 B

Variable types

Categorical1
Numeric2
Text2

Dataset

Description대한석탄공사 창립 이후 50년 동안 발생한 주요하고 다양한 사건들의 화보 목록을 데이터로 제공합니다. 추후 화보 추가 제공할 예정입니다.
Author대한석탄공사
URLhttps://www.data.go.kr/data/15100227/fileData.do

Alerts

Dataset has 8 (1.3%) duplicate rowsDuplicates
is highly overall correlated with 구분High correlation
구분 is highly overall correlated with High correlation
번호 has 24 (4.0%) zerosZeros

Reproduction

Analysis started2023-12-12 03:35:03.740185
Analysis finished2023-12-12 03:35:04.785042
Duration1.04 second
Software versionydata-profiling vv4.5.1
Download configurationconfig.json

Variables

구분
Categorical

HIGH CORRELATION 

Distinct46
Distinct (%)7.7%
Missing0
Missing (%)0.0%
Memory size4.8 KiB
1990~2001
71 
1980~1989
55 
1960~1969
53 
1970~1979
53 
1950~1959
50 
Other values (41)
314 

Length

Max length11
Median length10
Mean length6.8389262
Min length1

Unique

Unique0 ?
Unique (%)0.0%

Sample

1st row0
2nd row0
3rd row창립이전
4th row창립이전
5th row창립이전

Common Values

ValueCountFrequency (%)
1990~2001 71
 
11.9%
1980~1989 55
 
9.2%
1960~1969 53
 
8.9%
1970~1979 53
 
8.9%
1950~1959 50
 
8.4%
창립이전 16
 
2.7%
부산.묵호사업소 15
 
2.5%
화순광업소 13
 
2.2%
사택2 12
 
2.0%
도계광업소 12
 
2.0%
Other values (36) 246
41.3%

Length

2023-12-12T12:35:05.185096image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram of lengths of the category
ValueCountFrequency (%)
1990~2001 71
 
11.5%
1980~1989 55
 
8.9%
1960~1969 53
 
8.6%
1970~1979 53
 
8.6%
1950~1959 50
 
8.1%
창립이전 16
 
2.6%
부산.묵호사업소 15
 
2.4%
화순광업소 13
 
2.1%
행사 13
 
2.1%
사택2 12
 
1.9%
Other values (38) 265
43.0%


Real number (ℝ)

HIGH CORRELATION 

Distinct6
Distinct (%)1.0%
Missing0
Missing (%)0.0%
Infinite0
Infinite (%)0.0%
Mean3.2080537
Minimum0
Maximum5
Zeros1
Zeros (%)0.2%
Negative0
Negative (%)0.0%
Memory size5.4 KiB
2023-12-12T12:35:05.298010image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Quantile statistics

Minimum0
5-th percentile1
Q12
median3
Q34
95-th percentile5
Maximum5
Range5
Interquartile range (IQR)2

Descriptive statistics

Standard deviation1.2530308
Coefficient of variation (CV)0.39058911
Kurtosis-0.92009727
Mean3.2080537
Median Absolute Deviation (MAD)1
Skewness-0.084394806
Sum1912
Variance1.5700863
MonotonicityIncreasing
2023-12-12T12:35:05.455381image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=6)
ValueCountFrequency (%)
3 187
31.4%
5 123
20.6%
2 115
19.3%
4 112
18.8%
1 58
 
9.7%
0 1
 
0.2%
ValueCountFrequency (%)
0 1
 
0.2%
1 58
 
9.7%
2 115
19.3%
3 187
31.4%
4 112
18.8%
5 123
20.6%
ValueCountFrequency (%)
5 123
20.6%
4 112
18.8%
3 187
31.4%
2 115
19.3%
1 58
 
9.7%
0 1
 
0.2%
Distinct179
Distinct (%)30.0%
Missing0
Missing (%)0.0%
Memory size4.8 KiB
2023-12-12T12:35:05.921686image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Length

Max length3
Median length3
Mean length2.5822148
Min length2

Characters and Unicode

Total characters1539
Distinct characters11
Distinct categories2 ?
Distinct scripts2 ?
Distinct blocks2 ?
The Unicode Standard assigns character properties to each code point, which can be used to analyse textual variables.

Unique

Unique3 ?
Unique (%)0.5%

Sample

1st row1권
2nd row1권
3rd row10
4th row10
5th row10
ValueCountFrequency (%)
19 8
 
1.3%
130 7
 
1.2%
190 6
 
1.0%
169 5
 
0.8%
163 5
 
0.8%
155 5
 
0.8%
161 5
 
0.8%
117 5
 
0.8%
89 5
 
0.8%
197 5
 
0.8%
Other values (169) 540
90.6%
2023-12-12T12:35:06.592389image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Most occurring characters

ValueCountFrequency (%)
1 471
30.6%
8 137
 
8.9%
9 130
 
8.4%
7 125
 
8.1%
6 123
 
8.0%
3 113
 
7.3%
4 113
 
7.3%
5 111
 
7.2%
2 108
 
7.0%
0 106
 
6.9%

Most occurring categories

ValueCountFrequency (%)
Decimal Number 1537
99.9%
Other Letter 2
 
0.1%

Most frequent character per category

Decimal Number
ValueCountFrequency (%)
1 471
30.6%
8 137
 
8.9%
9 130
 
8.5%
7 125
 
8.1%
6 123
 
8.0%
3 113
 
7.4%
4 113
 
7.4%
5 111
 
7.2%
2 108
 
7.0%
0 106
 
6.9%
Other Letter
ValueCountFrequency (%)
2
100.0%

Most occurring scripts

ValueCountFrequency (%)
Common 1537
99.9%
Hangul 2
 
0.1%

Most frequent character per script

Common
ValueCountFrequency (%)
1 471
30.6%
8 137
 
8.9%
9 130
 
8.5%
7 125
 
8.1%
6 123
 
8.0%
3 113
 
7.4%
4 113
 
7.4%
5 111
 
7.2%
2 108
 
7.0%
0 106
 
6.9%
Hangul
ValueCountFrequency (%)
2
100.0%

Most occurring blocks

ValueCountFrequency (%)
ASCII 1537
99.9%
Hangul 2
 
0.1%

Most frequent character per block

ASCII
ValueCountFrequency (%)
1 471
30.6%
8 137
 
8.9%
9 130
 
8.5%
7 125
 
8.1%
6 123
 
8.0%
3 113
 
7.4%
4 113
 
7.4%
5 111
 
7.2%
2 108
 
7.0%
0 106
 
6.9%
Hangul
ValueCountFrequency (%)
2
100.0%

번호
Real number (ℝ)

ZEROS 

Distinct11
Distinct (%)1.8%
Missing0
Missing (%)0.0%
Infinite0
Infinite (%)0.0%
Mean3.7600671
Minimum0
Maximum10
Zeros24
Zeros (%)4.0%
Negative0
Negative (%)0.0%
Memory size5.4 KiB
2023-12-12T12:35:06.836469image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Quantile statistics

Minimum0
5-th percentile1
Q12
median4
Q35
95-th percentile8
Maximum10
Range10
Interquartile range (IQR)3

Descriptive statistics

Standard deviation2.2524271
Coefficient of variation (CV)0.59903907
Kurtosis-0.65173093
Mean3.7600671
Median Absolute Deviation (MAD)2
Skewness0.35564953
Sum2241
Variance5.0734279
MonotonicityNot monotonic
2023-12-12T12:35:06.991287image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Histogram with fixed size bins (bins=11)
ValueCountFrequency (%)
2 90
15.1%
1 89
14.9%
3 88
14.8%
4 86
14.4%
5 80
13.4%
6 60
10.1%
7 42
7.0%
8 25
 
4.2%
0 24
 
4.0%
9 10
 
1.7%
ValueCountFrequency (%)
0 24
 
4.0%
1 89
14.9%
2 90
15.1%
3 88
14.8%
4 86
14.4%
5 80
13.4%
6 60
10.1%
7 42
7.0%
8 25
 
4.2%
9 10
 
1.7%
ValueCountFrequency (%)
10 2
 
0.3%
9 10
 
1.7%
8 25
 
4.2%
7 42
7.0%
6 60
10.1%
5 80
13.4%
4 86
14.4%
3 88
14.8%
2 90
15.1%
1 89
14.9%

내용
Text

Distinct529
Distinct (%)88.8%
Missing0
Missing (%)0.0%
Memory size4.8 KiB
2023-12-12T12:35:07.343502image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Length

Max length119
Median length70
Mean length44.41443
Min length1

Characters and Unicode

Total characters26471
Distinct characters525
Distinct categories11 ?
Distinct scripts4 ?
Distinct blocks3 ?
The Unicode Standard assigns character properties to each code point, which can be used to analyse textual variables.

Unique

Unique469 ?
Unique (%)78.7%

Sample

1st row불1자산 1
2nd row화보 50년간지자산 1
3rd row1939년의 계산동 지역.(1930~1950_개발초기의 장성)
4th row장성2구 사무소 부근. 장성 동구지역으로 구 양지갱으로 개발됐다.(1930~1950_개발초기의 장성)
5th row금천구역. 2구(왼쪽)부터 1구(오른쪽)까지 수평갱도에 의해 개발했다.(1930~1950_개발초기의 장성)
ValueCountFrequency (%)
장성 49
 
1.2%
위해 33
 
0.8%
사장이 29
 
0.7%
공사는 26
 
0.6%
공사 24
 
0.6%
방문하여 21
 
0.5%
장성을 18
 
0.4%
건설 17
 
0.4%
총재가 17
 
0.4%
창립 17
 
0.4%
Other values (2365) 3806
93.8%
2023-12-12T12:35:07.917838image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Most occurring characters

ValueCountFrequency (%)
3463
 
13.1%
9 1336
 
5.0%
1 1241
 
4.7%
( 899
 
3.4%
) 896
 
3.4%
. 765
 
2.9%
0 653
 
2.5%
451
 
1.7%
_ 425
 
1.6%
371
 
1.4%
Other values (515) 15971
60.3%

Most occurring categories

ValueCountFrequency (%)
Other Letter 14706
55.6%
Decimal Number 4792
 
18.1%
Space Separator 3463
 
13.1%
Open Punctuation 900
 
3.4%
Other Punctuation 898
 
3.4%
Close Punctuation 897
 
3.4%
Connector Punctuation 425
 
1.6%
Math Symbol 308
 
1.2%
Uppercase Letter 76
 
0.3%
Dash Punctuation 5
 
< 0.1%

Most frequent character per category

Other Letter
ValueCountFrequency (%)
451
 
3.1%
371
 
2.5%
361
 
2.5%
309
 
2.1%
289
 
2.0%
269
 
1.8%
267
 
1.8%
237
 
1.6%
236
 
1.6%
227
 
1.5%
Other values (475) 11689
79.5%
Uppercase Letter
ValueCountFrequency (%)
A 23
30.3%
I 13
17.1%
D 9
 
11.8%
C 4
 
5.3%
K 4
 
5.3%
T 4
 
5.3%
V 3
 
3.9%
M 3
 
3.9%
P 3
 
3.9%
S 3
 
3.9%
Other values (5) 7
 
9.2%
Decimal Number
ValueCountFrequency (%)
9 1336
27.9%
1 1241
25.9%
0 653
13.6%
5 298
 
6.2%
6 287
 
6.0%
7 274
 
5.7%
2 253
 
5.3%
8 243
 
5.1%
4 104
 
2.2%
3 103
 
2.1%
Other Punctuation
ValueCountFrequency (%)
. 765
85.2%
, 87
 
9.7%
' 32
 
3.6%
; 10
 
1.1%
: 4
 
0.4%
Open Punctuation
ValueCountFrequency (%)
( 899
99.9%
[ 1
 
0.1%
Close Punctuation
ValueCountFrequency (%)
) 896
99.9%
] 1
 
0.1%
Math Symbol
ValueCountFrequency (%)
~ 307
99.7%
+ 1
 
0.3%
Space Separator
ValueCountFrequency (%)
3463
100.0%
Connector Punctuation
ValueCountFrequency (%)
_ 425
100.0%
Dash Punctuation
ValueCountFrequency (%)
- 5
100.0%
Lowercase Letter
ValueCountFrequency (%)
m 1
100.0%

Most occurring scripts

ValueCountFrequency (%)
Hangul 14704
55.5%
Common 11688
44.2%
Latin 77
 
0.3%
Han 2
 
< 0.1%

Most frequent character per script

Hangul
ValueCountFrequency (%)
451
 
3.1%
371
 
2.5%
361
 
2.5%
309
 
2.1%
289
 
2.0%
269
 
1.8%
267
 
1.8%
237
 
1.6%
236
 
1.6%
227
 
1.5%
Other values (473) 11687
79.5%
Common
ValueCountFrequency (%)
3463
29.6%
9 1336
 
11.4%
1 1241
 
10.6%
( 899
 
7.7%
) 896
 
7.7%
. 765
 
6.5%
0 653
 
5.6%
_ 425
 
3.6%
~ 307
 
2.6%
5 298
 
2.5%
Other values (14) 1405
12.0%
Latin
ValueCountFrequency (%)
A 23
29.9%
I 13
16.9%
D 9
 
11.7%
C 4
 
5.2%
K 4
 
5.2%
T 4
 
5.2%
V 3
 
3.9%
M 3
 
3.9%
P 3
 
3.9%
S 3
 
3.9%
Other values (6) 8
 
10.4%
Han
ValueCountFrequency (%)
1
50.0%
1
50.0%

Most occurring blocks

ValueCountFrequency (%)
Hangul 14704
55.5%
ASCII 11765
44.4%
CJK 2
 
< 0.1%

Most frequent character per block

ASCII
ValueCountFrequency (%)
3463
29.4%
9 1336
 
11.4%
1 1241
 
10.5%
( 899
 
7.6%
) 896
 
7.6%
. 765
 
6.5%
0 653
 
5.6%
_ 425
 
3.6%
~ 307
 
2.6%
5 298
 
2.5%
Other values (30) 1482
12.6%
Hangul
ValueCountFrequency (%)
451
 
3.1%
371
 
2.5%
361
 
2.5%
309
 
2.1%
289
 
2.0%
269
 
1.8%
267
 
1.8%
237
 
1.6%
236
 
1.6%
227
 
1.5%
Other values (473) 11687
79.5%
CJK
ValueCountFrequency (%)
1
50.0%
1
50.0%

Interactions

2023-12-12T12:35:04.395756image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T12:35:04.150272image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T12:35:04.505724image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
2023-12-12T12:35:04.285618image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/

Correlations

2023-12-12T12:35:08.043438image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
구분번호
구분1.0000.9890.000
0.9891.0000.268
번호0.0000.2681.000
2023-12-12T12:35:08.173210image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
번호구분
1.0000.2170.889
번호0.2171.0000.000
구분0.8890.0001.000

Missing values

2023-12-12T12:35:04.635895image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
A simple visualization of nullity by column.
2023-12-12T12:35:04.743195image/svg+xmlMatplotlib v3.7.2, https://matplotlib.org/
Nullity matrix is a data-dense display which lets you quickly visually pick out patterns in data completion.

Sample

구분페이지번호내용
0001권0불1자산 1
1011권0화보 50년간지자산 1
2창립이전11011939년의 계산동 지역.(1930~1950_개발초기의 장성)
3창립이전1102장성2구 사무소 부근. 장성 동구지역으로 구 양지갱으로 개발됐다.(1930~1950_개발초기의 장성)
4창립이전1103금천구역. 2구(왼쪽)부터 1구(오른쪽)까지 수평갱도에 의해 개발했다.(1930~1950_개발초기의 장성)
5창립이전11141939년의 화광도과 협심동 지역.(1930~1950_개발초기의 장성)
6창립이전1121장성 갱구(1940).퇴갱한 직원들이 축전 차에 올라 기념촬영을 했다.(1930~1950_삼척탄광 개광)
7창립이전1122개발 초기의 장성 갱구(1936.12. 기계 측량을 마친 직원들이 기념촬영했다.(1930~1950_삼척탄광 개광)
8창립이전1133준공 직후의 장성 이중교(1939)(1930~1950_삼척탄광 개광)
9창립이전1134삼척탄광 본관(1939.8). 1937년 도계에서 장성으로 이전됐다.(1930~1950_삼척탄광 개광)
구분페이지번호내용
586노동조합519541962년 정기대의원대회 광경(_노동조합)
587노동조합519551962년 정기대의원대회 광경(_노동조합)
588노사화합 행사51961노사화합 체육대회 광경(_노사화합 행사)
589노사화합 행사51962노사화합 체육대회 광경(_노사화합 행사)
590노사화합 행사51963노사화합 체육대회 광경(_노사화합 행사)
591노사화합 행사519741965년 도계체육대회 광경(_노사화합 행사)
592노사화합 행사51975장성과 도계는 5월, 화순에서는 10월에 거행된다. 체육대회는 가족과 지역주민까지 참여하는 지역의 중요한 행사로 치러졌다.(_노사화합 행사)
593노사화합 행사51976장성과 도계는 5월, 화순에서는 10월에 거행된다. 체육대회는 가족과 지역주민까지 참여하는 지역의 중요한 행사로 치러졌다.(_노사화합 행사)
594노사화합 행사51977단오절을 기해 안전작업을 기원하는 산신제가 노사합동으로 거행된다.(_노사화합 행사)
595노사화합 행사51978노동절과 추석에는 전국 요양기관에서 입원치료를 받고 있는 공상환자를 함께 위문한다. (_노사화합 행사)

Duplicate rows

Most frequently occurring

구분페이지번호내용# duplicates
01950~19591190설립준비위원회가 작성한 최초의 정관(1950.6.7)(1950~1959_창립)4
11950~19591190최초의 정관에 대한 대통령 인가서(1950.6.23)(1950~1959_창립)2
21960~19692370재건국민운동본부 총재인 유진오 박사가 공사의 재건국민운동촉진회 결성식에서 훈시를 하고 있다. (1961.6.12)(1960~1969)2
31960~19692450광신보안법 제정 전 공사는 자체적으로 탄광안전규정을 제정하여 시행하였다.(1962.1.1)(1960~1969)2
41980~19893850웅장한 자태를 드러낸 제2수갱 철탑과 야경 (1980~1989_장성 제2수갱 건설)2
51990~200131060증권거래소로 이전한 본사 사무실(1998.12.26)(1990~2001)2
61990~20013950유승규 태백시 국회의원이 국회에서 공사의 자본금 증자를 위한 공사법 개정안에 대해 제안 설명하고 있다. (1990.11)(1990~2001)2
7창립이전1130시라키에 의해 1940년 발간된 삼척탄전 조사보고서(1930~1950_삼척탄광 개광)2