<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>권세혁HOME</title>
<link>https://by-sekwon.github.io/stat_assay/</link>
<atom:link href="https://by-sekwon.github.io/stat_assay/index.xml" rel="self" type="application/rss+xml"/>
<description></description>
<generator>quarto-1.7.32</generator>
<lastBuildDate>Thu, 15 Oct 2026 15:00:00 GMT</lastBuildDate>
<item>
  <title>(분포가 들려주는 이야기) 분포를 모르면 평균도 해석할 수 없다 — 분포 이해의 중요성</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_distribution_matters.html</link>
  <description><![CDATA[ 




<section id="같은-평균-전혀-다른-사회" class="level2">
<h2 class="anchored" data-anchor-id="같은-평균-전혀-다른-사회">같은 평균, 전혀 다른 사회</h2>
<p>A 마을과 B 마을의 평균 소득이 모두 5,000만 원이라고 하자. A 마을은 대부분 4,000만 원에서 6,000만 원 사이에 몰려 있다. B 마을은 대부분 2,000만 원 근처에서 살고, 소수의 고소득 가구가 평균을 5,000만 원까지 끌어올린다. 뉴스가 “두 마을의 평균 소득이 같다”고 보도하면, 독자는 두 마을이 비슷하다고 믿는다. 그러나 B 마을 주민의 대부분은 평균보다 훨씬 가난하다. 평균은 같지만 <strong>분포</strong>는 전혀 다르다.</p>
</section>
<section id="이번-주에-본-네-가지-분포" class="level2">
<h2 class="anchored" data-anchor-id="이번-주에-본-네-가지-분포">이번 주에 본 네 가지 분포</h2>
<p>이번 주 다섯 편의 글은 서로 다른 모양의 분포를 하나씩 살펴보았다. 키는 대칭인 정규분포에 가까워 평균과 중앙값이 거의 같았다. 소득은 곱해져 만들어지는 로그정규분포라서 평균이 중앙값보다 28% 높았다. 3σ 밖의 이상치 비율은 분포 모양에 따라 몇 배씩 달라졌다. 하루 전화 수는 포아송분포를 따라 평균 주변에서 자연스럽게 흔들렸다. 같은 평균 한 숫자가 이렇게 다른 세계를 가릴 수 있다는 점이 핵심이다.</p>
<div id="72b24099" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_distribution_matters_files/figure-html/cell-2-output-1.png" width="902" height="393" class="figure-img"></p>
<figcaption>Two groups with exactly the same mean income (equal to 1.0 in normalized units). Group A is tightly clustered; Group B is skewed with most households below the mean. The mean is identical, but the typical household differs sharply.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="평균과-중앙값-그리고-보통의-사람" class="level2">
<h2 class="anchored" data-anchor-id="평균과-중앙값-그리고-보통의-사람">평균과 중앙값, 그리고 “보통의 사람”</h2>
<p>두 집단의 평균은 같지만 중앙값은 다르다. 중앙값은 “한가운데 사람”의 값이라서, 치우친 분포에서는 “보통의 사람”을 더 잘 대변한다. 그래서 공식 통계가 소득을 발표할 때 평균과 함께 중앙값을 함께 제시하는 것이다. 평균을 읽을 때는 항상 “이 평균을 만들어낸 분포의 모양은 어떤가?”를 먼저 물어야 한다.</p>
<div id="691f45a1" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_distribution_matters_files/figure-html/cell-3-output-1.png" width="614" height="394" class="figure-img"></p>
<figcaption>Share of people who earn less than the mean, for the symmetric group A and the skewed group B. In a skewed distribution, most people sit below the average; the ‘average person’ is not the average.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="분포를-먼저-보는-습관" class="level2">
<h2 class="anchored" data-anchor-id="분포를-먼저-보는-습관">분포를 먼저 보는 습관</h2>
<p>분포를 확인하는 일은 복잡한 계산이 아니다. 히스토그램 한 장, 평균과 중앙값 두 숫자, 그리고 “값이 더해져 만들어졌는가, 곱해져 만들어졌는가”라는 질문이면 충분하다. 이번 주 살펴본 모든 도구는 결국 이 습관을 위한 것이었다. 정규분포는 언제 믿을 수 있는지, 로그정규분포는 왜 소득을 로그로 보아야 하는지, 포아송분포는 날마다의 흔들림이 어디까지 정상인지를 알려준다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
평균을 해석하기 전에 확인할 것
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>히스토그램을 먼저 그려라.</strong> 숫자 하나보다 그림 한 장이 분포를 더 정확히 보여준다.</li>
<li><strong>평균과 중앙값을 나란히 놓아라.</strong> 차이가 크면 평균이 소수의 극단값에 끌려가고 있다.</li>
<li><strong>“평균 이상인 사람은 몇 %인가”를 세어라.</strong> 치우친 분포에서는 대부분이 평균 아래에 있다.</li>
<li><strong>분포의 폭을 함께 말하라.</strong> 표준편차나 분위수 없이 평균만 제시하면 절반의 정보만 전한 것이다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-기대값과-분포의-관계" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-기대값과-분포의-관계">더 깊이: 기대값과 분포의 관계</h2>
<p>강의노트 <a href="../../notes/math_stat/random_variable.html#기대값-정의">〈수리통계학 2. 확률변수 — 기대값 정의〉</a>는 확률변수의 기대값을 분포 전체에 가중한 평균으로 정의한다. 평균이 분포의 한 측면일 뿐임을 이해하는 것이 이번 주 이야기의 출발점이다. 평균은 분포의 무게중심을 말해줄 뿐, 그 무게중심 주변에 데이터가 어떻게 놓여 있는지는 말해주지 않는다.</p>
</section>
<section id="sdv-관점-평균-대신-분포를-제시하는-보고서가-가치를-만든다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-평균-대신-분포를-제시하는-보고서가-가치를-만든다">SDV 관점: 평균 대신 분포를 제시하는 보고서가 가치를 만든다</h2>
<p>분포를 읽는 일은 <strong>요약과 추론(1~2단계)</strong>을 넘어, 데이터를 결정의 근거로 바꾸는 <strong>SDV(통계적 데이터 가치화)</strong>의 출발점이다. 경영 보고서나 정책 자료에서 평균 하나만 제시하면, 듣는 사람은 “대표적인 값”을 떠올리지만 실제로 그 값에 해당하는 사람은 소수일 수 있다. 데이터에 가치를 더하는 통계학자는 평균과 함께 분포의 모양, 중앙값, 분위수, 변동의 폭을 보여준다. 그래야 의사결정자가 평균이 말해주는 것과 감추는 것을 함께 보고 판단할 수 있다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-평균은-분포를-요약한-한-줄이다" class="level2">
<h2 class="anchored" data-anchor-id="결론-평균은-분포를-요약한-한-줄이다">결론: 평균은 분포를 요약한 한 줄이다</h2>
<p>나는 학생들에게 평균이라는 숫자를 보면 “그래서 분포는 어떻게 생겼는데?”라고 되묻게 한다. 같은 평균 5,000만 원도 두 마을의 삶은 전혀 다르다. 분포를 모르고 평균만 읽으면, 우리는 가장 중요한 이야기를 놓치게 된다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>두 집단의 비교는 같은 평균을 갖도록 맞춘 가상의 시뮬레이션 자료다.</li>
<li>이번 주 키·소득 사례의 기준값은 앞선 글(10월 13일)의 참고 자료를 따른다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>분포</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_distribution_matters.html</guid>
  <pubDate>Thu, 15 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(분포가 들려주는 이야기) 하루에 걸려오는 전화 수는 어떤 분포를 따를까? — 포아송분포</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_phone_calls_poisson.html</link>
  <description><![CDATA[ 




<section id="평균-12통-그런데-날마다-다르다" class="level2">
<h2 class="anchored" data-anchor-id="평균-12통-그런데-날마다-다르다">평균 12통, 그런데 날마다 다르다</h2>
<p>한 고객센터가 하루 평균 12통의 전화를 받는다고 하자. 그런데 실제로 어떤 날은 7통, 어떤 날은 18통이 걸려 온다. 관리자는 이 날마다의 차이를 보고 “오늘은 인력이 부족했다”거나 “무슨 일이 생겼다”고 해석하기 쉽다. 그러나 평균이 같아도 이런 흔들림은 자연스럽게 생긴다. 일정한 평균 속도로 서로 독립적으로 일어나는 사건의 횟수는 <strong>포아송분포</strong>를 따르고, 포아송분포는 평균 하나만으로 날마다의 변동 폭까지 말해준다.</p>
</section>
<section id="포아송분포는-무엇을-세는가" class="level2">
<h2 class="anchored" data-anchor-id="포아송분포는-무엇을-세는가">포아송분포는 무엇을 세는가</h2>
<p>포아송분포는 정해진 시간이나 공간 안에서 드물게 일어나는 사건의 횟수를 센다. 평균 발생 횟수를 <img src="https://latex.codecogs.com/png.latex?%5Clambda">라고 하면, <img src="https://latex.codecogs.com/png.latex?k">번 일어날 확률은</p>
<p><img src="https://latex.codecogs.com/png.latex?P(X=k)=%5Cfrac%7Be%5E%7B-%5Clambda%7D%5Clambda%5E%7Bk%7D%7D%7Bk!%7D"></p>
<p>이다. 이 분포의 가장 중요한 특징은 <strong>평균과 분산이 모두 <img src="https://latex.codecogs.com/png.latex?%5Clambda"></strong>라는 점이다. 평균이 12이면 분산도 12, 표준편차는 약 3.5다. 그래서 평균 12통인 날의 “보통” 변동은 대략 8통에서 16통 사이에 있다.</p>
<div id="4f53948b" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_phone_calls_poisson_files/figure-html/cell-2-output-1.png" width="758" height="413" class="figure-img"></p>
<figcaption>Poisson probabilities of k calls per day when the average is 12 (blue-grey) and 3 (dark green). The larger average gives a wider, more symmetric spread around its mean, and the small average gives a skewed shape close to zero.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="년-365일-시뮬레이션으로-확인하기" class="level2">
<h2 class="anchored" data-anchor-id="년-365일-시뮬레이션으로-확인하기">1년 365일 시뮬레이션으로 확인하기</h2>
<p>이론이 실제 하루하루에 맞는지 확인해 보자. 평균 12통의 포아송 과정을 365일 동안 시뮬레이션하고, 날마다 걸려 온 전화 수의 히스토그램을 이론값과 겹쳐 본다.</p>
<div id="07aa9ad1" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_phone_calls_poisson_files/figure-html/cell-3-output-1.png" width="758" height="413" class="figure-img"></p>
<figcaption>One simulated year (365 days) of calls from a Poisson process with average 12 per day. Histogram of daily counts (bars) with the theoretical Poisson probabilities (line). The two agree closely, so the day-to-day swings are what chance alone would produce.</figcaption>
</figure>
</div>
</div>
</div>
<p>시뮬레이션에서 날마다의 전화 수는 평균 근처에 모이되, 6통이나 18통인 날도 자주 나온다. 이 흔들림은 고객센터가 무언가를 잘못했다는 신호가 아니다. 평균이 12인 포아송 과정에서 기대되는 정상적인 모습이다.</p>
</section>
<section id="포아송분포를-쓰는-조건" class="level2">
<h2 class="anchored" data-anchor-id="포아송분포를-쓰는-조건">포아송분포를 쓰는 조건</h2>
<p>포아송분포가 맞으려면 세 가지 조건이 필요하다. 사건이 일정한 평균 속도로 일어나야 하고, 한 사건이 다음 사건의 발생에 영향을 주지 않아야 하며(독립), 같은 짧은 구간에 두 사건이 동시에 일어나지 않아야 한다. 전화 수는 대체로 이 조건에 가깝다. 그러나 점심시간에는 전화가 몰리고 월요일 아침에는 평균이 더 높다면, 하루 전체를 하나의 평균으로 보는 것은 무리다. 시간대별로 평균을 나누어 보아야 한다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
포아송 모형을 적용하기 전에 확인할 것
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>사건이 드물고 독립적인가?</strong> 한 번 일어난 사건이 다음 사건을 부추기면 포아송이 아니다.</li>
<li><strong>평균이 일정한가?</strong> 요일이나 시간대에 따라 평균이 바뀌면 구간을 나눠야 한다.</li>
<li><strong>분산이 평균과 비슷한가?</strong> 분산이 평균보다 훨씬 크면 과산포(overdispersion)이므로 다른 모형을 고려하라.</li>
<li><strong>“평균과 다른 날”을 기대 범위와 비교하라.</strong> 포아송이라면 평균 ± 2σ 밖은 드물다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-포아송-과정과-포아송분포" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-포아송-과정과-포아송분포">더 깊이: 포아송 과정과 포아송분포</h2>
<p>강의노트 <a href="../../notes/math_stat/famous_distribution.html#포아송분포-x-sim-poissonlambda">〈수리통계학 5. 유명분포 — 포아송분포〉</a>는 포아송 과정의 조건과, 평균과 분산이 모두 <img src="https://latex.codecogs.com/png.latex?%5Clambda">인 포아송분포를 정의한다. 앞서 군집 착각 글에서 다룬 것도 같은 포아송 구조였다. 이번에는 같은 분포가 하루의 전화 수라는 일상의 숫자를 설명하는 데 쓰였다.</p>
</section>
<section id="sdv-관점-평균과-함께-정상-변동의-폭을-설계하라" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-평균과-함께-정상-변동의-폭을-설계하라">SDV 관점: 평균과 함께 “정상 변동의 폭”을 설계하라</h2>
<p>포아송 모형으로 기대 변동을 계산하는 일은 <strong>요약과 추론(1~2단계)</strong>의 실전 응용이다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이 계산은 운영 데이터의 해석 기준을 만든다. 고객 문의, 주문 건수, 서버 요청 수는 모두 평균 주변에서 흔들린다. 그 흔들림이 포아송 범위 안이면 인력 배치와 재고를 평균 기준으로 조정하면 되지만, 범위를 벗어나면 원인을 찾아야 한다. 평균만 보고 날마다 반응하면 정상 변동을 문제로 착각해 쓸데없는 대응을 반복하게 된다. 평균과 함께 정상 변동의 폭을 제시하는 것이 데이터를 실제 운영의 가치로 바꾸는 방법이다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-평균이-같아도-하루는-흔들린다" class="level2">
<h2 class="anchored" data-anchor-id="결론-평균이-같아도-하루는-흔들린다">결론: 평균이 같아도 하루는 흔들린다</h2>
<p>나는 학생들에게 날마다의 숫자가 흔들린다고 곧바로 원인을 찾으려 하지 말라고 말한다. 평균이 12인 전화는 어떤 날 18통이 걸려 오는 것이 당연하다. 포아송분포는 그 흔들림이 얼마나 자연스러운지를 알려주는 도구이고, 그 범위를 넘을 때에만 비로소 이야기할 거리가 생긴다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>시뮬레이션 자료는 평균 12의 포아송 난수를 사용한 가상의 1년치 데이터다. 실제 고객센터의 전화 수 분포는 시간대와 계절에 따라 다를 수 있다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>분포</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_phone_calls_poisson.html</guid>
  <pubDate>Wed, 14 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(분포가 들려주는 이야기) 평균에서 세 표준편차 밖이면 이상한 값일까? — 68-95-99.7 법칙</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_three_sigma.html</link>
  <description><![CDATA[ 




<section id="세-표준편차-밖은-드문-값일까" class="level2">
<h2 class="anchored" data-anchor-id="세-표준편차-밖은-드문-값일까">세 표준편차 밖은 드문 값일까</h2>
<p>공장에서 생산되는 부품의 길이를 재서 평균이 10.00mm, 표준편차가 0.02mm라고 하자. 어느 날 10.07mm짜리 부품이 나왔다. 평균에서 0.07mm 떨어졌으니 표준편차의 3.5배다. 이 부품은 ’이상한 것’일까? 많은 현장에서 “3시그마 밖은 이상하다”는 규칙을 쓴다. 그 근거가 바로 정규분포의 <strong>68-95-99.7 법칙</strong>이다.</p>
</section>
<section id="정규분포에서-세-구간의-확률" class="level2">
<h2 class="anchored" data-anchor-id="정규분포에서-세-구간의-확률">정규분포에서 세 구간의 확률</h2>
<p>정규분포는 평균을 중심으로 대칭이다. 평균 ± 1σ 안에 들어갈 확률은 약 68.3%, ± 2σ 안에는 약 95.4%, ± 3σ 안에는 약 99.7%다. 따라서 어느 한쪽이든 3σ를 넘을 확률은 약 0.13%, 양쪽을 합하면 0.27%, 즉 약 370개 중 1개꼴이다.</p>
<div id="8925ddb6" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_three_sigma_files/figure-html/cell-2-output-1.png" width="758" height="413" class="figure-img"></p>
<figcaption>Standard normal curve with the regions within 1, 2 and 3 standard deviations shaded. Each band contains about 68.3%, 95.4% and 99.7% of the probability.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="이-법칙이-말해주지-않는-것" class="level2">
<h2 class="anchored" data-anchor-id="이-법칙이-말해주지-않는-것">이 법칙이 말해주지 않는 것</h2>
<p>68-95-99.7 법칙은 <strong>데이터가 정규분포일 때만</strong> 성립한다. 그런데 정규분포가 아닌 데이터에 이 규칙을 그대로 적용하면, 실제보다 훨씬 많은 값이 “이상치”로 잡히거나, 반대로 진짜 이상치가 정상으로 넘어간다. 아래 그림은 정규분포와 꼬리가 긴 분포(로그정규분포)에서 평균 ± 3표준편차 밖에 떨어지는 비율을 비교한 것이다.</p>
<div id="a7deece0" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_three_sigma_files/figure-html/cell-3-output-1.png" width="643" height="413" class="figure-img"></p>
<figcaption>Share of observations beyond mean plus or minus 3 standard deviations: about 0.27% for normal data (the textbook rule) but several times larger for a long-tailed lognormal distribution. The same rule gives very different alarm rates.</figcaption>
</figure>
</div>
</div>
</div>
<p>비율의 차이만 보면 정규분포의 0.27%와 로그정규분포의 값이 얼마나 다른지 바로 와닿지 않을 수 있다. 중요한 것은 그 차이가 <strong>경보의 빈도</strong>로 바뀐다는 점이다. 하루 1만 건의 거래를 감시하는 시스템에서 0.27%는 하루 27건이지만, 꼬리가 긴 데이터에서 실제 비율이 그보다 몇 배 높다면 경보는 하루 수백 건이 되어 담당자가 감당할 수 없게 된다.</p>
</section>
<section id="분포에-상관없이-쓸-수-있는-보장-체비셰프-부등식" class="level2">
<h2 class="anchored" data-anchor-id="분포에-상관없이-쓸-수-있는-보장-체비셰프-부등식">분포에 상관없이 쓸 수 있는 보장: 체비셰프 부등식</h2>
<p>분포 모양을 모를 때는 <strong>체비셰프 부등식</strong>이 안전한 하한을 준다. 평균에서 <img src="https://latex.codecogs.com/png.latex?k">표준편차 이상 떨어질 확률은 <img src="https://latex.codecogs.com/png.latex?1/k%5E2">를 넘지 않으므로, 어떤 분포에서든 평균 ± 3σ 안에 최소 <img src="https://latex.codecogs.com/png.latex?1-1/9%5Capprox88.9%5C%25">가 들어간다. 정규분포라면 99.7%이지만, 분포를 모르면 보장은 88.9%까지 떨어진다. 이 차이가 바로 정규분포 가정이 주는 이득과 위험의 크기다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
3σ 기준을 쓰기 전에 확인할 것
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>데이터의 분포를 먼저 그려보라.</strong> 정규분포가 아니면 68-95-99.7은 근사조차 되지 않는다.</li>
<li><strong>이상치 판정을 확률로 바꿔 계산하라.</strong> 일일 건수와 곱해 실제 경보 횟수를 추정하라.</li>
<li><strong>분포를 모르면 체비셰프로 하한을 잡아라.</strong> 보수적이지만 가정이 필요 없다.</li>
<li><strong>이상치의 정의를 업무 비용과 연결하라.</strong> 놓치는 비용과 경보의 비용 중 무엇이 큰지에 따라 기준이 달라진다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-체비셰프-부등식" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-체비셰프-부등식">더 깊이: 체비셰프 부등식</h2>
<p>강의노트 <a href="../../notes/math_stat/famous_distribution.html#체비셰프-부등식chebyshevs-inequality">〈수리통계학 5. 확률분포함수 관련 부등식 — 체비셰프 부등식〉</a>는 분포의 모양에 관계없이 평균에서 멀리 떨어진 값의 확률을 분산으로 상한을 잡는 부등식을 소개한다. 오늘 살펴본 “분포를 몰라도 88.9%는 보장된다”는 말이 바로 이 부등식의 결과다.</p>
</section>
<section id="sdv-관점-이상치-기준은-확률로-계산하고-비용으로-정하라" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-이상치-기준은-확률로-계산하고-비용으로-정하라">SDV 관점: 이상치 기준은 확률로 계산하고 비용으로 정하라</h2>
<p>정규분포 기준으로 이상치를 판정하는 일은 <strong>요약과 추론(1~2단계)</strong>에 속한다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 중요한 것은 기준 자체가 아니라 그 기준이 만들어낼 결과다. 품질 관리, 부정 거래 탐지, 서버 모니터링에서 “3σ 밖”이라는 규칙을 기계적으로 적용하면, 분포가 치우친 지표에서 경보가 폭주해 정작 중요한 신호가 묻힌다. 기준을 정할 때는 분포를 확인해 실제 경보율을 계산하고, 그 경보를 처리하는 비용과 놓쳤을 때의 손실을 함께 놓고 결정해야 데이터가 판단의 가치를 만든다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-3σ는-정규분포의-규칙이지-모든-데이터의-규칙이-아니다" class="level2">
<h2 class="anchored" data-anchor-id="결론-3σ는-정규분포의-규칙이지-모든-데이터의-규칙이-아니다">결론: 3σ는 정규분포의 규칙이지, 모든 데이터의 규칙이 아니다</h2>
<p>나는 학생들에게 “3시그마 밖이면 이상하다”는 말을 들으면 두 가지를 되묻게 한다. 이 데이터는 정말 종 모양인가? 그리고 그 기준으로 하루에 몇 번 경보가 울리는가? 370번에 한 번이라는 숫자는 정규분포에서만 믿을 수 있는 숫자다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>68-95-99.7 법칙의 확률값(68.27%, 95.45%, 99.73%)은 표준정규분포의 이론값이다.</li>
<li>체비셰프 부등식의 보장은 강의노트의 정리를 기준으로 했다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>분포</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_three_sigma.html</guid>
  <pubDate>Tue, 13 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(분포가 들려주는 이야기) 키는 정규분포에 가깝고 소득은 그렇지 않은 이유 — 분포 모양이 다른 이유</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_height_vs_income.html</link>
  <description><![CDATA[ 




<section id="키는-평균-근처에-모이고-소득은-한쪽으로-쏠린다" class="level2">
<h2 class="anchored" data-anchor-id="키는-평균-근처에-모이고-소득은-한쪽으로-쏠린다">키는 평균 근처에 모이고, 소득은 한쪽으로 쏠린다</h2>
<p>길거리에서 성인 남성 100명을 세운 뒤 키를 재면, 대부분은 170cm 안팎에 몰려 있고 170cm보다 훨씬 작거나 훨씬 큰 사람은 드물다. 병무청의 병역판정검사 통계를 보면 평균은 174cm 조금 넘고 표준편차는 6cm 안팎이다(자료마다 조금씩 다르다). 반면 가구 소득을 같은 방식으로 재면 그림이 완전히 달라진다. 통계청(국가데이터처)의 가계금융복지조사에 따르면 2024년 가구의 <strong>평균</strong> 소득은 7,427만 원이었지만, 한가운데 가구인 <strong>중위소득</strong>은 5,800만 원이었다. 평균이 중앙값보다 약 28% 높다. 키에서는 평균과 중앙값이 거의 같은데, 소득에서는 왜 이렇게 벌어질까?</p>
</section>
<section id="키는-더해져서-소득은-곱해져서-만들어진다" class="level2">
<h2 class="anchored" data-anchor-id="키는-더해져서-소득은-곱해져서-만들어진다">키는 더해져서, 소득은 곱해져서 만들어진다</h2>
<p>키는 수많은 유전자와 영양, 성장 환경이 <strong>더해져서</strong> 결정된다. 한 요인이 키를 1cm 늘리거나 줄이는 정도이고, 그 효과들이 모이면 가운데에 몰리는 종 모양이 된다. 앞서 본 중심극한정리가 말하는 상황이다.</p>
<p>소득은 다르다. 학력이 소득을 1.2배로 만들고, 업종이 1.5배로 만들고, 경력이 또 1.3배로 만든다. 이렇게 <strong>곱해지는</strong> 구조에서는 작은 차이가 누적되어 큰 차이가 된다. 그 결과 한쪽 끝에 소수의 고소득자가 길게 늘어진 오른쪽 꼬리가 생긴다. 로그를 취하면 곱셈이 덧셈으로 바뀌어 다시 대칭에 가까워지는 것이 바로 <strong>로그정규분포</strong>다.</p>
<div id="f9fd81ef" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_height_vs_income_files/figure-html/cell-2-output-1.png" width="897" height="393" class="figure-img"></p>
<figcaption>Simulated adult male heights (normal, mean 174.5 cm, SD 5.8 cm) on the left: nearly symmetric. Simulated household incomes (lognormal, calibrated so the mean-to-median ratio is about 1.28 as in the 2024 Korean survey) on the right: a long right tail. Values are illustrative.</figcaption>
</figure>
</div>
</div>
</div>
<p>그림에서 소득 히스토그램은 3억 원 이상의 가구를 잘라서 보여준다. 실제로는 그 꼬리 너머에도 가구가 있고, 그 가구들이 평균을 끌어올린다. 평균이 중앙값보다 큰 것은 계산 실수가 아니라 분포 모양이 낳은 결과다.</p>
<p>로그를 취한 뒤에도 같은 소득 자료를 그려보면, 곱셈 구조가 덧셈 구조로 바뀌면서 꼬리가 사라진다.</p>
<div id="8946aea4" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_height_vs_income_files/figure-html/cell-3-output-1.png" width="710" height="413" class="figure-img"></p>
<figcaption>Simulated household incomes after taking logarithms: the long right tail disappears and the histogram follows a normal curve (red line). Multiplicative effects become additive on the log scale. Values are illustrative.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="평균-소득이-실제-가구를-대표하지-못하는-이유" class="level2">
<h2 class="anchored" data-anchor-id="평균-소득이-실제-가구를-대표하지-못하는-이유">평균 소득이 실제 가구를 대표하지 못하는 이유</h2>
<p>평균 소득 7,427만 원을 보고 “우리 가구도 이 정도는 벌어야 평균이지”라고 생각하면, 절반 이상의 가구는 그 기준에 못 미친다. 중위소득 5,800만 원이 “한가운데 가구”의 실제 모습에 더 가깝다. 같은 데이터에 대해 평균과 중앙값이 이렇게 다른 답을 주는 것은, 분포가 대칭이 아니기 때문이다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
분포의 모양을 읽는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>평균과 중앙값을 함께 보라.</strong> 둘이 가까우면 대칭에 가깝고, 멀면 한쪽으로 치우쳐 있다.</li>
<li><strong>값이 어떻게 만들어지는지 물어라.</strong> 더해졌는지, 곱해졌는지에 따라 모양이 정해진다.</li>
<li><strong>꼬리가 긴 데이터에는 로그 눈금을 시도하라.</strong> 곱셈 구조는 로그를 취하면 대칭에 가까워진다.</li>
<li><strong>“평균 이상”이라는 표현이 얼마나 많은 사람에게 해당하는지 세어보라.</strong> 치우친 분포에서는 소수만 해당한다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-로그정규분포의-정의" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-로그정규분포의-정의">더 깊이: 로그정규분포의 정의</h2>
<p>강의노트 <a href="../../notes/math_stat/famous_distribution.html#로그정규분포-logx-sim-nmusigma2">〈수리통계학 5. 유명분포 — 로그정규분포〉</a>는 <img src="https://latex.codecogs.com/png.latex?%5Clog%20X">가 정규분포를 따르는 확률변수 <img src="https://latex.codecogs.com/png.latex?X">를 로그정규분포로 정의한다. 소득처럼 양수이고 곱셈으로 결정되는 값이 이 분포에 가까운 이유를 이 정의가 설명해준다.</p>
</section>
<section id="sdv-관점-평균-소득은-대표-가구가-아니다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-평균-소득은-대표-가구가-아니다">SDV 관점: 평균 소득은 “대표 가구”가 아니다</h2>
<p>분포가 치우쳐 있음을 알아채는 일은 <strong>요약과 추론(1~2단계)</strong>의 핵심이다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이는 정책과 사업 설계의 기준을 정하는 문제다. 지원금 기준, 가격 정책, 고객 세분화를 평균 소득이나 평균 구매액으로 설계하면, 실제 대상의 절반 이상이 설계와 맞지 않는다. 치우친 데이터에서는 중앙값과 분위수를 기준으로 삼고, 평균은 전체 규모를 말할 때만 쓰는 것이 데이터를 실제 결정의 가치로 연결하는 방법이다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-같은-사람의-값도-만들어진-방식이-모양을-정한다" class="level2">
<h2 class="anchored" data-anchor-id="결론-같은-사람의-값도-만들어진-방식이-모양을-정한다">결론: 같은 ’사람의 값’도 만들어진 방식이 모양을 정한다</h2>
<p>나는 학생들에게 키와 소득을 같은 종류의 데이터로 취급하지 말라고 말한다. 키는 더해져서 종 모양이 되고, 소득은 곱해져서 오른쪽으로 길게 늘어진다. 분포의 모양을 먼저 알아야 평균이 무엇을 말해주고 무엇을 감추는지 판단할 수 있다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li><a href="https://www.korea.kr/briefing/policyBriefingView.do?newsId=156733201">2025년 가계금융복지조사 결과 — 대한민국 정책브리핑</a>: 2024년 가구 평균 소득 7,427만 원, 처분가능소득 6,032만 원 등.</li>
<li><a href="https://www.korea.kr/briefing/policyBriefingView.do?newsId=156664637">2024년 가계금융복지조사 결과 — 대한민국 정책브리핑</a></li>
<li>병무청 병역판정검사 기준 성인 남성 키 평균·표준편차는 자료마다 차이가 있어 대략적인 값(평균 약 174~175cm, 표준편차 약 6cm)으로만 사용했다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>분포</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_height_vs_income.html</guid>
  <pubDate>Mon, 12 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(분포가 들려주는 이야기) 세상의 모든 데이터는 정규분포일까? — 정규분포 신화와 한계</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_normal_myth.html</link>
  <description><![CDATA[ 




<section id="교과서의-종-모양이-세상의-모양은-아니다" class="level2">
<h2 class="anchored" data-anchor-id="교과서의-종-모양이-세상의-모양은-아니다">교과서의 종 모양이 세상의 모양은 아니다</h2>
<p>통계 수업에서 가장 먼저 배우는 그림은 종 모양의 정규분포다. 평균 근처에 데이터가 몰리고 양쪽으로 대칭으로 줄어드는 이 곡선은 너무 익숙해서, 많은 사람이 세상의 모든 데이터가 이 모양일 거라고 생각한다. 시험 점수나 키처럼 여러 요인이 더해져 만들어진 값은 실제로 정규분포에 가깝다. 그런데 회사원의 연봉, 하루 동안 한 사람이 쓰는 돈, 웹사이트 체류 시간은 어떨까? 이런 데이터는 한쪽 꼬리가 길게 늘어진 <strong>비대칭</strong> 모양인 경우가 많다.</p>
</section>
<section id="네-가지-모양을-나란히-보기" class="level2">
<h2 class="anchored" data-anchor-id="네-가지-모양을-나란히-보기">네 가지 모양을 나란히 보기</h2>
<p>아래 그림은 같은 평균(100)을 가진 네 가지 가상의 데이터를 나란히 놓은 것이다. 왼쪽 위는 정규분포, 오른쪽 위는 균일분포, 왼쪽 아래는 지수분포(기다리는 시간), 오른쪽 아래는 로그정규분포(소득, 재산 같은 값)다.</p>
<div id="331bb843" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_normal_myth_files/figure-html/cell-2-output-1.png" width="854" height="585" class="figure-img"></p>
<figcaption>Four distributions with the same mean of 100: normal (symmetric), uniform (flat), exponential (waiting times) and lognormal (incomes, wealth). Only the first looks like the textbook bell curve; the others have very different shapes and tails.</figcaption>
</figure>
</div>
</div>
</div>
<p>정규분포와 균일분포에서는 평균과 중앙값이 거의 같다. 그런데 지수분포와 로그정규분포에서는 평균이 중앙값보다 눈에 띄게 오른쪽에 있다. 꼬리가 긴 쪽으로 평균이 끌려가기 때문이다. 같은 “평균”이라도 분포의 모양에 따라 대표성이 달라진다는 뜻이다.</p>
</section>
<section id="정규분포가-나타나는-조건" class="level2">
<h2 class="anchored" data-anchor-id="정규분포가-나타나는-조건">정규분포가 나타나는 조건</h2>
<p>정규분포가 자연스럽게 나타나는 이유는 <strong>중심극한정리</strong>에 있다. 개별 데이터가 어떤 모양이든, 독립적인 값들이 많이 더해지거나 평균을 내면 그 합이나 평균의 분포는 정규분포에 가까워진다. 그래서 측정 오차, 여러 유전자의 합으로 결정되는 키, 많은 시험 문항의 합산 점수는 종 모양이 된다. 반면 소득은 곱셈 구조로 만들어진다. 학력, 업종, 경력, 운이 각각 소득을 몇 배씩 키우거나 줄이므로, 로그를 취해야 대칭에 가까워지는 로그정규 모양이 된다.</p>
<div id="de05c998" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_normal_myth_files/figure-html/cell-3-output-1.png" width="998" height="336" class="figure-img"></p>
<figcaption>A sum of 30 independent uniform values looks normal. A product of 30 independent multipliers is strongly skewed on its own scale (values above 12 are cut off in the middle panel), but its logarithm, which is a sum of logs, looks normal again.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="정규분포를-가정하기-전에-던질-질문" class="level2">
<h2 class="anchored" data-anchor-id="정규분포를-가정하기-전에-던질-질문">정규분포를 가정하기 전에 던질 질문</h2>
<p>정규분포를 가정하면 평균과 표준편차만으로 모든 것을 설명할 수 있어 편하다. 그러나 그 편리함이 함정이 된다. 데이터가 한쪽으로 치우쳐 있는데도 평균과 표준편차로 “보통”을 정의하면, 실제 대부분의 사람과 동떨어진 기준이 나온다. 따라서 분석에 들어가기 전에 히스토그램을 먼저 그려보고, 대칭인지, 꼬리가 긴지, 상한이나 하한이 있는지를 확인하는 것이 순서다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
정규분포 가정을 점검하는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>분포 모양부터 그려보라.</strong> 평균과 표준편차를 계산하기 전에 히스토그램 한 장이 가장 많은 것을 말해준다.</li>
<li><strong>평균과 중앙값을 비교하라.</strong> 차이가 크면 분포가 치우쳐 있다는 신호다.</li>
<li><strong>“여러 요인이 더해졌는가, 곱해졌는가”를 생각하라.</strong> 더해지면 정규분포에 가깝고, 곱해지면 로그를 취해야 한다.</li>
<li><strong>0이나 상한이 있는 데이터인지 확인하라.</strong> 음수가 불가능한 값은 대칭일 수 없다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-정규분포의-정의와-중심극한정리" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-정규분포의-정의와-중심극한정리">더 깊이: 정규분포의 정의와 중심극한정리</h2>
<p>강의노트 <a href="../../notes/math_stat/famous_distribution.html#정규분포-x-sim-nmusigma">〈수리통계학 5. 유명분포 — 정규분포〉</a>는 정규분포를 평균 <img src="https://latex.codecogs.com/png.latex?%5Cmu">와 분산 <img src="https://latex.codecogs.com/png.latex?%5Csigma%5E2">로 정의하고, 중심극한정리를 그 이론적 근거로 소개한다. 오늘 본 합의 분포와 곱의 분포 차이가 바로 그 정리가 말하는 조건을 눈으로 확인하는 과정이다.</p>
</section>
<section id="sdv-관점-분포를-가정하기-전에-분포를-데이터에서-확인하라" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-분포를-가정하기-전에-분포를-데이터에서-확인하라">SDV 관점: 분포를 가정하기 전에 분포를 데이터에서 확인하라</h2>
<p>분포 모양을 확인하는 일은 <strong>요약과 추론(1~2단계)</strong>의 출발점이다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 정규분포 가정은 데이터에 가치를 더하는 과정에서 조용히 판단을 왜곡하는 가장 흔한 원인이다. 고객 가치, 매출, 대기 시간 같은 지표에 평균과 표준편차로 “정상 범위”를 설정하면, 꼬리가 긴 데이터에서는 정상 범위가 너무 좁거나 넓게 잡힌다. 지표를 설계하기 전에 실제 분포를 그려보고 적합한 기준(중앙값, 분위수, 로그 변환)을 고르는 것이 데이터를 결정의 가치로 바꾸는 첫 단계다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-종-모양은-가정이지-출발점이-아니다" class="level2">
<h2 class="anchored" data-anchor-id="결론-종-모양은-가정이지-출발점이-아니다">결론: 종 모양은 가정이지, 출발점이 아니다</h2>
<p>나는 학생들에게 정규분포를 “세상의 기본값”이 아니라 “여러 작은 힘이 더해진 결과에만 나타나는 특수한 모양”으로 기억하라고 말한다. 데이터를 받으면 공식보다 먼저 그림을 그려라. 종 모양이 보이면 정규분포를 쓰고, 한쪽으로 길게 늘어지면 다른 도구를 찾아라.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>중심극한정리와 정규분포에 관한 이론적 설명은 강의노트를 기준으로 했으며, 그림은 시뮬레이션 데이터(가상 자료)다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>분포</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_normal_myth.html</guid>
  <pubDate>Sun, 11 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(조건부확률과 베이즈의 사고) 새로운 증거가 들어오면 믿음도 바뀌어야 한다 — 베이지안 사고</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_bayesian_thinking.html</link>
  <description><![CDATA[ 




<section id="같은-증거-다른-출발점" class="level2">
<h2 class="anchored" data-anchor-id="같은-증거-다른-출발점">같은 증거, 다른 출발점</h2>
<p>검사 결과가 양성이라는 같은 증거를 두 사람이 받았다고 하자. 한 사람은 증상이 전혀 없는 건강한 사람이고, 다른 한 사람은 증상이 뚜렷해 병일 가능성이 30%라고 짐작하는 사람이다. 검사의 민감도와 특이도가 99%라면, 두 사람의 사후 믿음은 같을까? 아니다. 처음 믿음, 곧 <strong>사전확률</strong>이 다르기 때문이다. 베이지안 사고의 핵심은 증거를 믿음과 따로 떼어 보지 않고, <strong>이전의 믿음을 출발점으로 삼아 증거로 고쳐 가는 것</strong>이다.</p>
</section>
<section id="한-번의-양성-두-번의-양성" class="level2">
<h2 class="anchored" data-anchor-id="한-번의-양성-두-번의-양성">한 번의 양성, 두 번의 양성</h2>
<p>처음 믿음이 1%인 사람이 양성 판정을 한 번 받으면 믿음은 50%로 오른다. 그 상태에서 같은 검사를 한 번 더 받아 양성이 나오면, 이제 50%가 새로운 출발점이 되어 99%로 뛴다. 세 번째 양성에서는 99.99%에 이른다. 증거가 쌓일수록 믿음은 확신에 가까워진다.</p>
<div id="6026b78b" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_bayesian_thinking_files/figure-html/cell-2-output-1.png" width="710" height="422" class="figure-img"></p>
<figcaption>Belief after successive positive results from the same 99%-accurate test, starting from two different priors (1% and 30%). Each result updates the previous belief; after a few independent positives, both starting points converge near certainty.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="출발점이-틀리면-결론도-틀린다" class="level2">
<h2 class="anchored" data-anchor-id="출발점이-틀리면-결론도-틀린다">출발점이 틀리면 결론도 틀린다</h2>
<p>그렇다고 출발점이 중요하지 않다는 뜻은 아니다. 이 그림에서 증거가 몇 번만 쌓여도 두 믿음은 가까워지지만, 증거가 <strong>독립이며 같은 검사로 공정하게 얻어졌을 때</strong>에만 그렇다. 같은 증상을 여러 번 확인한다고 독립적인 증거가 늘어나는 것은 아니다. 또 출발점의 믿음이 극단적으로 0이면, 아무리 많은 증거가 들어와도 믿음은 0에서 움직이지 않는다. 베이지안 갱신은 출발점에 대한 판단을 숨기지 않고 명시적으로 드러내는 방법이다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
베이지안 사고를 실천하는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>처음 믿음을 숫자로 적어라.</strong> “모르겠다”도 0.5라는 명시적 출발점이다.</li>
<li><strong>새 증거가 얼마나 독립적인지 물어라.</strong> 같은 근거의 반복은 믿음을 과도하게 올린다.</li>
<li><strong>증거마다 사후확률을 기록하라.</strong> 어떤 증거가 믿음을 움직였는지 추적할 수 있다.</li>
<li><strong>극단적 출발점(0 또는 1)을 피하라.</strong> 한 번 0이나 1이 된 믿음은 증거로도 바뀌지 않는다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-베이즈-확률과-사후분포" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-베이즈-확률과-사후분포">더 깊이: 베이즈 확률과 사후분포</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#베이즈-확률">〈수리통계학 1. 확률론 — 베이즈 확률〉</a>은 사전확률과 사후확률의 관계로 베이지안 사고를 정식화한다. 이 틀을 모수 추정으로 확장하면 사전분포와 사후분포가 되고, 그 자세한 내용은 추정 단원의 베이즈 추정량에서 이어진다. 오늘 본 확률 갱신은 그 출발점이다.</p>
</section>
<section id="sdv-관점-믿음을-갱신하는-절차가-데이터-가치를-만든다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-믿음을-갱신하는-절차가-데이터-가치를-만든다">SDV 관점: 믿음을 갱신하는 절차가 데이터 가치를 만든다</h2>
<p>베이지안 갱신은 <strong>요약과 추론(1~2단계)</strong>에서 한 걸음 더 나아가, 판단을 <strong>계속 업데이트할 수 있는 구조</strong>로 만든다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이는 조직의 의사결정에 중요한 함의를 갖는다. 신규 사업의 성공 가능성, 신제품 수요, 정책 효과에 대한 초기 믿음을 명시적으로 기록하고, 새 데이터가 들어올 때마다 그 믿음을 고쳐 가면 판단은 감과 고집이 아니라 축적된 증거가 된다. 데이터에 가치를 더하는 일은 한 번의 정답을 찾는 것이 아니라, 증거에 따라 믿음을 정직하게 바꾸는 과정을 설계하는 일이다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-믿음은-증거와-함께-움직여야-한다" class="level2">
<h2 class="anchored" data-anchor-id="결론-믿음은-증거와-함께-움직여야-한다">결론: 믿음은 증거와 함께 움직여야 한다</h2>
<p>나는 학생들에게 “나는 원래 이렇게 생각했다”는 말보다 “새 증거를 보고 어떻게 바꿨는가”를 묻는 습관을 가지라고 말한다. 같은 양성 결과도 출발점에 따라 다르게 읽히지만, 증거가 쌓이면 결국 같은 곳에 이른다. 베이지안 사고는 그 과정을 숨기지 않고 보여주는 방법이다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>민감도·특이도 99%와 사전확률 1%, 30%는 설명을 위한 가상의 값이다. 독립적인 여러 검사 결과를 가정했다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_bayesian_thinking.html</guid>
  <pubDate>Thu, 08 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(조건부확률과 베이즈의 사고) 스팸메일 필터는 어떻게 배워가는가? — 베이즈 갱신</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_spam_filter_learning.html</link>
  <description><![CDATA[ 




<section id="규칙-목록-대신-확률로-배우는-필터" class="level2">
<h2 class="anchored" data-anchor-id="규칙-목록-대신-확률로-배우는-필터">규칙 목록 대신 확률로 배우는 필터</h2>
<p>초기의 스팸 필터는 “무료”, “당첨” 같은 단어가 들어 있으면 차단하는 규칙 목록이었다. 스패머들은 금방 철자를 바꾸었고, 규칙은 끝없이 늘어났다. 2002년 프로그래머 폴 그레이엄은 다른 길을 제안했다. 메일 속 단어들이 스팸일 확률을 베이즈 정리로 계산하고, 사용자가 스팸을 표시할 때마다 그 확률을 갱신하자는 것이다. 이 방식을 <strong>나이브 베이즈(naive Bayes)</strong> 필터라고 부른다. 여기서 핵심은 규칙이 고정되지 않고 <strong>증거가 들어올수록 믿음이 바뀌는 것</strong>이다.</p>
</section>
<section id="단어-하나로-믿음을-바꾸기" class="level2">
<h2 class="anchored" data-anchor-id="단어-하나로-믿음을-바꾸기">단어 하나로 믿음을 바꾸기</h2>
<p>메일이 스팸일 사전확률을 50%라고 하자. 어떤 단어 “무료”가 스팸에는 80%의 확률로 나타나고, 정상 메일에는 10%만 나타난다고 가정한다. 메일에 “무료”가 들어 있다면 베이즈 정리에 따라</p>
<p><img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EC%8A%A4%ED%8C%B8%7D%5Cmid%5Ctext%7B%EB%AC%B4%EB%A3%8C%7D)=%5Cfrac%7BP(%5Ctext%7B%EB%AC%B4%EB%A3%8C%7D%5Cmid%5Ctext%7B%EC%8A%A4%ED%8C%B8%7D)P(%5Ctext%7B%EC%8A%A4%ED%8C%B8%7D)%7D%7BP(%5Ctext%7B%EB%AC%B4%EB%A3%8C%7D%5Cmid%5Ctext%7B%EC%8A%A4%ED%8C%B8%7D)P(%5Ctext%7B%EC%8A%A4%ED%8C%B8%7D)+P(%5Ctext%7B%EB%AC%B4%EB%A3%8C%7D%5Cmid%5Ctext%7B%EC%A0%95%EC%83%81%7D)P(%5Ctext%7B%EC%A0%95%EC%83%81%7D)%7D=%5Cfrac%7B0.8%5Ctimes0.5%7D%7B0.8%5Ctimes0.5+0.1%5Ctimes0.5%7D%5Capprox88.9%5C%25"></p>
<p>이다. 메일 한 통에 단어 하나가 들어왔을 뿐인데 스팸일 확률이 50%에서 89%로 뛴다. 이 새로운 값이 다음 단어를 계산할 때의 출발점, 곧 <strong>새로운 사전확률</strong>이 된다.</p>
<div id="f50d4ef5" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_spam_filter_learning_files/figure-html/cell-2-output-1.png" width="711" height="413" class="figure-img"></p>
<figcaption>Belief update for one email: starting from a 50% prior, each additional spam-indicating word (each with likelihood 0.8 in spam and 0.1 in normal mail) raises the probability that the email is spam. Certainty builds quickly, and no single word settles it.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="단어마다-얼마나-무게가-있는가" class="level2">
<h2 class="anchored" data-anchor-id="단어마다-얼마나-무게가-있는가">단어마다 얼마나 무게가 있는가</h2>
<p>모든 단어가 같은 힘을 갖지는 않는다. 어떤 단어는 스팸에서만 자주 나오고, 어떤 단어는 정상 메일에서도 흔하다. 필터는 각 단어가 스팸과 정상 메일을 얼마나 잘 구분하는지를 데이터에서 배운다. 아래 그림은 같은 사전확률에서 단어마다 다른 힘을 가진다는 점을 보여준다.</p>
<div id="70bdad43" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_spam_filter_learning_files/figure-html/cell-3-output-1.png" width="902" height="393" class="figure-img"></p>
<figcaption>Illustrative word-level evidence: how often each word appears in spam versus normal mail, and the resulting probability that a single email containing it is spam (prior 50%). Words that appear mostly in spam move the belief the most.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="학습은-곧-갱신이다" class="level2">
<h2 class="anchored" data-anchor-id="학습은-곧-갱신이다">학습은 곧 갱신이다</h2>
<p>필터는 사용자가 스팸을 표시하거나 정상 메일을 구출할 때마다 각 단어의 확률을 조금씩 고친다. 이것은 베이즈 갱신의 전형적인 모습이다. 새로운 증거가 들어올 때마다 사전확률을 사후확률로 바꾸고, 그 사후확률이 다음 판단의 출발점이 된다. 그래서 스패머가 새로운 단어를 써도, 사용자의 몇 번의 표시만으로 필터는 빠르게 따라잡는다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
베이즈 갱신을 따라가는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>사전확률을 먼저 적어라.</strong> 증거를 보기 전 믿음이 출발점이다.</li>
<li><strong>증거가 각 가설에서 얼마나 자주 나타나는지 비교하라.</strong> 우도 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EC%A6%9D%EA%B1%B0%7D%5Cmid%5Ctext%7B%EA%B0%80%EC%84%A4%7D)">의 차이가 갱신의 힘이다.</li>
<li><strong>사후확률을 다음 단계의 사전확률로 써라.</strong> 증거가 여러 개일 때 이 과정을 반복하면 된다.</li>
<li><strong>독립 가정을 의심하라.</strong> 나이브 베이즈는 단어들이 독립이라고 가정하므로, 실제로 함께 나타나는 단어들은 힘을 과대평가할 수 있다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-베이즈-규칙" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-베이즈-규칙">더 깊이: 베이즈 규칙</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#베이즈-규칙">〈수리통계학 1. 확률론 — 베이즈 규칙〉</a>은 <img src="https://latex.codecogs.com/png.latex?P(A%5Cmid%20B)=P(B%5Cmid%20A)P(A)/P(B)">와 전확률 법칙을 결합한 베이즈 규칙을 소개한다. 오늘 스팸 필터의 계산은 이 규칙을 증거 하나씩에 반복 적용한 것이다. 앞선 두 편에서 본 조건의 방향 문제가 여기서 실제 기계의 학습 과정이 된다.</p>
</section>
<section id="sdv-관점-데이터는-한-번-분석하고-끝나지-않는다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-데이터는-한-번-분석하고-끝나지-않는다">SDV 관점: 데이터는 한 번 분석하고 끝나지 않는다</h2>
<p>베이즈 갱신은 <strong>요약과 추론(1~2단계)</strong>을 넘어, 데이터가 쌓일수록 판단이 나아지는 구조를 설계하는 일이다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 스팸 필터는 좋은 사례다. 한 번 만든 규칙은 환경이 바뀌면 쓸모가 없어지지만, 확률을 갱신하는 구조는 새로운 증거를 흡수하며 가치를 유지한다. 데이터 분석의 가치는 보고서 한 장이 아니라, 새 데이터가 들어올 때 판단을 고쳐 가는 절차에 있다. 그 절차를 설계하는 것이 AI 시대에 통계학자가 맡을 일이다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-믿음은-증거를-만나면-바뀌어야-한다" class="level2">
<h2 class="anchored" data-anchor-id="결론-믿음은-증거를-만나면-바뀌어야-한다">결론: 믿음은 증거를 만나면 바뀌어야 한다</h2>
<p>나는 학생들에게 스팸 필터를 떠올리라고 말한다. 필터는 규칙을 외운 것이 아니라 메일을 받을 때마다 “이제 얼마나 스팸 같은가”를 다시 계산한다. 베이즈 갱신은 그 계산을 한 줄의 원칙으로 정리한다. 처음의 믿음은 출발점일 뿐이고, 증거가 믿음을 바꿀 때 비로소 배움이 일어난다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>Paul Graham, <em>A Plan for Spam</em> (2002) — <a href="https://www.paulgraham.com/spam.html">paulgraham.com/spam.html</a></li>
<li>앞의 계산 예시(단어별 출현 비율)는 설명을 위한 가상의 숫자다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_spam_filter_learning.html</guid>
  <pubDate>Wed, 07 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(조건부확률과 베이즈의 사고) 범인이 범행 현장에 있을 확률과 현장에 있던 사람이 범인일 확률 — 조건의 방향</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_direction_of_condition.html</link>
  <description><![CDATA[ 




<section id="범인은-현장에-있었다-그런데-현장에-있던-사람이-범인일까" class="level2">
<h2 class="anchored" data-anchor-id="범인은-현장에-있었다-그런데-현장에-있던-사람이-범인일까">범인은 현장에 있었다, 그런데 현장에 있던 사람이 범인일까</h2>
<p>수사 드라마에서 자주 나오는 말이 있다. “범인은 반드시 범행 현장에 있었다.” 맞는 말이다. 범인이 현장에 없었다면 범행이 불가능하다. 그런데 이 말을 뒤집어 “현장에 있었던 사람은 범인이다”라고 하면 틀린다. 같은 사건의 두 조건부확률이 완전히 다른 값을 가지기 때문이다. 하나는 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%ED%98%84%EC%9E%A5%7D%5Cmid%5Ctext%7B%EB%B2%94%EC%9D%B8%7D)=100%5C%25">, 다른 하나는 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EB%B2%94%EC%9D%B8%7D%5Cmid%5Ctext%7B%ED%98%84%EC%9E%A5%7D)">이다.</p>
</section>
<section id="가상의-도시에서-세어-보기" class="level2">
<h2 class="anchored" data-anchor-id="가상의-도시에서-세어-보기">가상의 도시에서 세어 보기</h2>
<p>한 도시에 1만 명이 살고, 그중 범인은 1명이라고 하자. 범행 시각에 현장 근처에 있던 사람은 범인을 포함해 50명이다. 범인이 현장에 있었다고 가정하면 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%ED%98%84%EC%9E%A5%7D%5Cmid%5Ctext%7B%EB%B2%94%EC%9D%B8%7D)=1">이다. 그런데 현장에 있던 50명 중 범인은 1명뿐이므로, <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EB%B2%94%EC%9D%B8%7D%5Cmid%5Ctext%7B%ED%98%84%EC%9E%A5%7D)=1/50=2%5C%25">다.</p>
<div id="ee284809" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_direction_of_condition_files/figure-html/cell-2-output-1.png" width="643" height="403" class="figure-img"></p>
<figcaption>In a city of 10,000 people with one culprit and 50 people present at the scene: P(at scene | culprit) = 100%, but P(culprit | at scene) = 2%. The condition reversed, the answer collapses.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="전확률-법칙-두-방향을-잇는-다리" class="level2">
<h2 class="anchored" data-anchor-id="전확률-법칙-두-방향을-잇는-다리">전확률 법칙: 두 방향을 잇는 다리</h2>
<p>두 방향을 연결하는 도구가 <strong>전확률 법칙</strong>과 베이즈 정리다. 현장에 있을 확률 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%ED%98%84%EC%9E%A5%7D)">은 범인이 현장에 있는 경우와 범인이 아닌 사람이 현장에 있는 경우를 합친 것이다.</p>
<p><img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%ED%98%84%EC%9E%A5%7D)=P(%5Ctext%7B%ED%98%84%EC%9E%A5%7D%5Cmid%5Ctext%7B%EB%B2%94%EC%9D%B8%7D)P(%5Ctext%7B%EB%B2%94%EC%9D%B8%7D)+P(%5Ctext%7B%ED%98%84%EC%9E%A5%7D%5Cmid%5Ctext%7B%EB%B2%94%EC%9D%B8%7D%5Ec)P(%5Ctext%7B%EB%B2%94%EC%9D%B8%7D%5Ec)"></p>
<p>이 식에 숫자를 넣으면 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%ED%98%84%EC%9E%A5%7D)=1%5Ctimes0.0001+0.005%5Ctimes0.9999%5Capprox0.0051">이고, 따라서 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EB%B2%94%EC%9D%B8%7D%5Cmid%5Ctext%7B%ED%98%84%EC%9E%A5%7D)=0.0001/0.0051%5Capprox2%5C%25">가 된다. 방향을 뒤집을 때 필요한 것은 전체 집단에서 범인이 얼마나 드문지, 곧 <strong>사전확률</strong>이다.</p>
<div id="2d06b03b" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_direction_of_condition_files/figure-html/cell-3-output-1.png" width="711" height="422" class="figure-img"></p>
<figcaption>Posterior probability that a person at the scene is the culprit, as the number of people present grows (10,000-person city, one culprit). The more people are at the scene, the less the presence alone points to the culprit.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="조건의-방향이-실제-판단을-바꾸는-순간" class="level2">
<h2 class="anchored" data-anchor-id="조건의-방향이-실제-판단을-바꾸는-순간">조건의 방향이 실제 판단을 바꾸는 순간</h2>
<p>법정이나 수사에서 “범인은 현장에 있었다”는 증거는 그 자체로 강하지만, 현장에 있었다는 사실만으로 특정인을 범인이라 부르기에는 부족하다. 현장에 있었던 사람은 대부분 범인이 아니라는 기저율이 함께 있기 때문이다. 반대로 범인이라는 증거가 있으면 현장 확인이 강해진다. 증거의 방향을 헷갈리면, 무고한 사람을 범인으로 몰거나 진짜 범인을 놓치게 된다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
조건의 방향을 바꿀 때 확인할 것
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>“A이면 B”와 “B이면 A”를 구분하라.</strong> 두 문장의 확률은 분모가 다르다.</li>
<li><strong>전확률로 분모를 풀어 써라.</strong> <img src="https://latex.codecogs.com/png.latex?P(B)=%5Csum%20P(B%5Cmid%20A_i)P(A_i)">로 나누면 방향 전환이 계산이 된다.</li>
<li><strong>사전확률(기저율)을 반드시 명시하라.</strong> 대상 집단이 클수록 같은 증거의 힘은 약해진다.</li>
<li><strong>증거가 몇 명에게 해당하는지 세어 보라.</strong> 해당자가 많으면 조건 자체의 의미는 희석된다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-전확률-법칙" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-전확률-법칙">더 깊이: 전확률 법칙</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#전확률-법칙">〈수리통계학 1. 확률론 — 전확률 법칙〉</a>은 표본공간을 서로소인 사건들로 나눌 때 <img src="https://latex.codecogs.com/png.latex?P(B)=%5Csum_i%20P(B%5Cmid%20A_i)P(A_i)">가 성립함을 보인다. 오늘 현장 확률을 계산한 두 항이 바로 그 분해다. 이 식과 베이즈 규칙을 결합하면 조건의 방향을 자유롭게 바꿀 수 있다.</p>
</section>
<section id="sdv-관점-방향을-바꾸는-계산이-없으면-데이터는-쉽게-낙인이-된다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-방향을-바꾸는-계산이-없으면-데이터는-쉽게-낙인이-된다">SDV 관점: 방향을 바꾸는 계산이 없으면 데이터는 쉽게 낙인이 된다</h2>
<p>조건의 방향을 구분하는 일은 <strong>요약과 추론(1~2단계)</strong>이며, <strong>SDV(통계적 데이터 가치화)</strong> 관점에서는 데이터가 사람을 판단하는 근거가 될 때 특히 중요하다. 이상 거래 탐지, 신용 평가, 사고 위험 예측에서 “위험 행동을 한 사람 중 A 속성이 많다”는 사실을 “A 속성을 가진 사람은 위험하다”로 바꿔 읽으면, 대부분 무고한 사람에게 불이익이 돌아간다. 데이터를 가치 있는 결정으로 바꾸려면, 모든 조건부 지표에 방향과 기저율을 함께 적고 전확률로 검산하는 절차가 필요하다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-확률은-방향을-가진다" class="level2">
<h2 class="anchored" data-anchor-id="결론-확률은-방향을-가진다">결론: 확률은 방향을 가진다</h2>
<p>나는 학생들에게 수사극의 명대사 하나를 떠올리게 한다. “범인은 현장에 있었다.” 이 문장은 참이지만, 뒤집으면 전혀 다른 문장이 된다. 조건부확률을 배운다는 것은, 문장 속 조건의 방향을 세는 습관을 익히는 일이다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>1만 명 도시, 범인 1명, 현장 50명이라는 설정은 조건의 방향을 보여주기 위한 가상의 예시다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_direction_of_condition.html</guid>
  <pubDate>Tue, 06 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(조건부확률과 베이즈의 사고) 정확도 99% 검사가 틀릴 수 있는 이유 — 조건부확률</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_conditional_accuracy.html</link>
  <description><![CDATA[ 




<section id="같은-말처럼-들리지만-다른-확률이다" class="level2">
<h2 class="anchored" data-anchor-id="같은-말처럼-들리지만-다른-확률이다">같은 말처럼 들리지만 다른 확률이다</h2>
<p>뉴스에서 “이 검사는 병이 있는 사람의 99%를 양성으로 찾아낸다”는 말을 들었다. 이것을 “양성이면 병이 있을 확률이 99%”라고 바꿔 기억하면 어떻게 될까? 어제 본 것처럼, 그 바꿔 말하기는 틀렸다. 앞 문장은 <strong>병을 조건으로 둔 양성 확률</strong> <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EC%96%91%EC%84%B1%7D%5Cmid%5Ctext%7B%EB%B3%91%7D)=0.99">이고, 바꿔 말한 문장은 <strong>양성을 조건으로 둔 병 확률</strong> <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EB%B3%91%7D%5Cmid%5Ctext%7B%EC%96%91%EC%84%B1%7D)">이다. 조건부확률에서 조건의 방향을 바꾸면 값이 바뀐다.</p>
</section>
<section id="조건부확률의-정의" class="level2">
<h2 class="anchored" data-anchor-id="조건부확률의-정의">조건부확률의 정의</h2>
<p>두 사건 <img src="https://latex.codecogs.com/png.latex?A">, <img src="https://latex.codecogs.com/png.latex?B">에 대해 <img src="https://latex.codecogs.com/png.latex?B">가 일어났다는 조건에서 <img src="https://latex.codecogs.com/png.latex?A">가 일어날 확률은</p>
<p><img src="https://latex.codecogs.com/png.latex?P(A%5Cmid%20B)=%5Cfrac%7BP(A%5Ccap%20B)%7D%7BP(B)%7D"></p>
<p>이다. 분모는 조건이 되는 사건의 확률이다. 따라서 <img src="https://latex.codecogs.com/png.latex?P(A%5Cmid%20B)">와 <img src="https://latex.codecogs.com/png.latex?P(B%5Cmid%20A)">는 분모가 다르기 때문에 일반적으로 같지 않다. 이 식을 한 번 더 쓰면 <img src="https://latex.codecogs.com/png.latex?P(A%5Cmid%20B)P(B)=P(A%5Ccap%20B)=P(B%5Cmid%20A)P(A)">라는 관계가 나온다. 이것이 바로 다음 주 베이즈 정리의 뿌리다.</p>
<div id="5d1bd8f6" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_conditional_accuracy_files/figure-html/cell-2-output-1.png" width="643" height="404" class="figure-img"></p>
<figcaption>Two conditional probabilities built from the same 2x2 table of 10,000 people (1,000 sick, 9,000 healthy). Reading the condition in the other direction changes the answer: P(positive | sick) = 90% but P(sick | positive) is much lower.</figcaption>
</figure>
</div>
</div>
</div>
<p>이 예시의 표는 민감도 90%와 특이도 90%, 병의 유병률 10%(1만 명 중 1,000명)를 가정한 가상의 숫자다. 양성은 모두 1,800명이고, 그중 실제 환자는 900명, 즉 절반이다. 병이 더 드물다면 이 비율은 훨씬 떨어진다. 조건을 뒤집는 순간 답이 달라지는 이유는 분모가 ’환자’에서 ’양성’으로 바뀌었기 때문이다.</p>
</section>
<section id="방향을-헷갈리게-하는-말들" class="level2">
<h2 class="anchored" data-anchor-id="방향을-헷갈리게-하는-말들">방향을 헷갈리게 하는 말들</h2>
<p>일상의 많은 오해는 조건의 방향을 바꿔 읽는 데서 생긴다. “이 집단에서 범죄율이 높다”는 말은 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EB%B2%94%EC%A3%84%7D%5Cmid%5Ctext%7B%EC%A7%91%EB%8B%A8%7D)">이지만, “범죄자 중 이 집단 출신이 많다”는 말은 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EC%A7%91%EB%8B%A8%7D%5Cmid%5Ctext%7B%EB%B2%94%EC%A3%84%7D)">다. 두 값은 집단의 크기와 범죄의 기저율에 따라 전혀 다르다. 같은 현상을 두고도 어떤 조건으로 말하느냐에 따라 전혀 다른 결론이 나온다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
조건의 방향을 확인하는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>“무엇을 조건으로 두었는가”를 먼저 적어라.</strong> <img src="https://latex.codecogs.com/png.latex?P(A%5Cmid%20B)">에서 <img src="https://latex.codecogs.com/png.latex?B">는 이미 알고 있는 사실이다.</li>
<li><strong>두 방향을 모두 계산해 보라.</strong> 표를 한 장 그리면 방향에 따른 답이 한눈에 보인다.</li>
<li><strong>분모가 무엇인지 확인하라.</strong> 분모가 바뀌면 같은 분자라도 확률은 완전히 달라진다.</li>
<li><strong>“A이면 B”를 “B이면 A”로 읽지 마라.</strong> 이 한 가지 실수가 검사 결과 해석의 대부분을 망친다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-조건부확률의-정의" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-조건부확률의-정의">더 깊이: 조건부확률의 정의</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#조건부-확률">〈수리통계학 1. 확률론 — 조건부 확률〉</a>은 <img src="https://latex.codecogs.com/png.latex?P(A%5Cmid%20B)=P(A%5Ccap%20B)/P(B)">를 정의하고, 조건부확률이 <img src="https://latex.codecogs.com/png.latex?P(%5Ccdot%5Cmid%20B)">라는 새로운 확률 측도가 된다는 점을 설명한다. 오늘의 표 계산은 그 정의를 그대로 따른 것이다. 분모를 바꾸는 순간 새로운 확률 공간이 생긴다고 이해하면 방향 문제를 정리하기 쉽다.</p>
</section>
<section id="sdv-관점-지표의-조건을-명시하지-않으면-보고서가-오해를-낳는다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-지표의-조건을-명시하지-않으면-보고서가-오해를-낳는다">SDV 관점: 지표의 조건을 명시하지 않으면 보고서가 오해를 낳는다</h2>
<p>조건부확률을 정확히 쓰는 일은 <strong>요약과 추론(1~2단계)</strong>의 기본기이며, <strong>SDV(통계적 데이터 가치화)</strong> 관점에서는 데이터 보고서의 신뢰를 결정한다. “구매 고객의 80%가 앱 사용자”라는 문장은 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EC%95%B1%7D%5Cmid%5Ctext%7B%EA%B5%AC%EB%A7%A4%7D)">이고, “앱 사용자의 구매율이 높다”는 문장은 <img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EA%B5%AC%EB%A7%A4%7D%5Cmid%5Ctext%7B%EC%95%B1%7D)">다. 두 지표는 모집단이 다르면 전혀 다른 행동을 가리킨다. 가치 있는 데이터 분석은 모든 지표에 조건을 명시하고, 그 조건이 결정에 어떤 영향을 주는지 함께 보여주는 것에서 시작한다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-조건은-문장의-가장-중요한-단어다" class="level2">
<h2 class="anchored" data-anchor-id="결론-조건은-문장의-가장-중요한-단어다">결론: 조건은 문장의 가장 중요한 단어다</h2>
<p>나는 학생들에게 확률 문장을 읽을 때 “~라는 조건에서”라는 말을 먼저 찾으라고 가르친다. 같은 99%라도 ’병이라면’과 ’양성이라면’은 전혀 다른 세계를 가리킨다. 조건을 정확히 읽으면 검사 결과도, 뉴스 통계도, 보고서의 숫자도 제 의미를 찾는다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>표의 숫자(민감도·특이도 90%, 유병률 10%)는 조건부확률의 방향을 보여주기 위한 가상의 값이다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_conditional_accuracy.html</guid>
  <pubDate>Mon, 05 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(조건부확률과 베이즈의 사고) 검사 결과가 양성이면 정말 병에 걸린 것일까? — 기저율</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_base_rate_test.html</link>
  <description><![CDATA[ 




<section id="나-맞는-검사인데-양성이면-99일까" class="level2">
<h2 class="anchored" data-anchor-id="나-맞는-검사인데-양성이면-99일까">99%나 맞는 검사인데, 양성이면 99%일까</h2>
<p>건강검진에서 “정확도 99%”라는 검사를 받았더니 양성이 나왔다. 그렇다면 병에 걸렸을 확률은 99%일까? 많은 사람이 그렇게 생각하고 크게 놀란다. 그런데 계산해 보면 답은 9%에 가깝다. 검사가 틀릴 가능성은 1%뿐인데 왜 이런 결과가 나올까? 답은 검사 자체보다 <strong>그 병이 인구 속에서 얼마나 흔한가</strong>, 곧 <strong>기저율(base rate)</strong>에 있다.</p>
</section>
<section id="만-명을-세어-보면-보인다" class="level2">
<h2 class="anchored" data-anchor-id="만-명을-세어-보면-보인다">10만 명을 세어 보면 보인다</h2>
<p>병의 유병률을 0.1%, 곧 1,000명에 1명이라고 하자. 검사는 병에 걸린 사람을 99% 찾아내고(민감도 99%), 건강한 사람을 99% 정상으로 판정한다(특이도 99%). 10만 명을 검사하면 다음과 같다.</p>
<ul>
<li>실제로 병이 있는 사람: 100명 → 검사 양성 99명(정확히 찾음)</li>
<li>병이 없는 사람: 99,900명 → 1%인 999명이 잘못 양성</li>
</ul>
<p>양성 판정은 모두 1,098명이고, 그중 진짜 환자는 99명뿐이다. 따라서 양성 결과를 받은 사람이 실제로 병에 걸렸을 확률은 <img src="https://latex.codecogs.com/png.latex?99/1098%5Capprox9.0%5C%25">다.</p>
<div id="94e776d0" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_base_rate_test_files/figure-html/cell-2-output-1.png" width="643" height="403" class="figure-img"></p>
<figcaption>Out of 100,000 people screened with a 0.1% prevalence and a 99%-accurate test: 99 true positives versus 999 false positives. Most positive results are false alarms, even though the test itself is 99% accurate.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="기저율이-바뀌면-같은-검사의-의미도-바뀐다" class="level2">
<h2 class="anchored" data-anchor-id="기저율이-바뀌면-같은-검사의-의미도-바뀐다">기저율이 바뀌면 같은 검사의 의미도 바뀐다</h2>
<p>같은 99% 검사라도 병이 얼마나 흔한지에 따라 양성의 의미가 완전히 달라진다. 유병률이 1%면 양성이 나왔을 때 확률은 50%로 뛰고, 10%면 92%에 이른다. 검사의 정확도는 고정되어 있는데, 양성 결과를 해석하는 확률은 기저율에 따라 움직인다.</p>
<div id="5846d817" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_base_rate_test_files/figure-html/cell-3-output-1.png" width="711" height="420" class="figure-img"></p>
<figcaption>Probability of actually having the condition, given a positive result from a 99%-accurate test, as a function of prevalence (log scale). The same test gives 9% at 0.1% prevalence and 92% at 10%.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="양성을-어떻게-읽을-것인가" class="level2">
<h2 class="anchored" data-anchor-id="양성을-어떻게-읽을-것인가">양성을 어떻게 읽을 것인가</h2>
<p>검사 결과를 받았을 때 던져야 할 질문은 “이 검사가 얼마나 정확한가?”가 아니라 “이 검사를 받은 사람들 중에서 병이 얼마나 흔한가?”다. 증상이 있어 병원에 온 사람에게는 기저율이 높으므로 양성의 의미가 크다. 증상 없이 무작위로 받은 대규모 선별검사에서는 기저율이 낮으므로, 양성이라도 확진 검사로 한 번 더 확인하는 것이 합리적이다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
검사 결과를 읽을 때 확인할 것
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>기저율부터 물어라.</strong> 병이 전체 인구에서 얼마나 흔한지가 양성의 의미를 정한다.</li>
<li><strong>10만 명 같은 큰 집단을 세어 보라.</strong> 진짜 양성과 거짓 양성의 수를 나란히 놓으면 직관이 바로잡힌다.</li>
<li><strong>민감도와 특이도를 구분하라.</strong> 민감도는 환자 중 양성 비율, 특이도는 건강한 사람 중 음성 비율이다.</li>
<li><strong>선별검사의 양성은 확진이 아니다.</strong> 낮은 기저율에서는 확진 검사를 한 번 더 받는 것이 기본이다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-특이도와-민감도" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-특이도와-민감도">더 깊이: 특이도와 민감도</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#특이도와-민감도">〈수리통계학 1. 확률론 — 특이도와 민감도〉</a>는 검사의 민감도와 특이도를 조건부확률로 정의하고, 이를 이용해 양성 예측도를 계산하는 방법을 소개한다. 오늘 계산이 바로 그 절차다. 민감도 <img src="https://latex.codecogs.com/png.latex?P(+%5Cmid%20D)">와 양성 예측도 <img src="https://latex.codecogs.com/png.latex?P(D%5Cmid%20+)">는 서로 다른 조건부확률이다.</p>
</section>
<section id="sdv-관점-지표의-정확도보다-모집단의-기저율을-먼저-설계하라" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-지표의-정확도보다-모집단의-기저율을-먼저-설계하라">SDV 관점: 지표의 정확도보다 모집단의 기저율을 먼저 설계하라</h2>
<p>기저율을 무시하는 실수는 <strong>요약과 추론(1~2단계)</strong>의 가장 흔한 오류이며, <strong>SDV(통계적 데이터 가치화)</strong> 관점에서는 의사결정의 비용을 좌우한다. 사기 탐지, 이탈 예측, 불량 탐지 모델이 “정확도 99%”라고 발표되어도, 실제로 문제가 되는 사례가 전체의 0.1%라면 경보의 대부분은 거짓이다. 모델의 가치는 정확도 숫자가 아니라 기저율 아래에서의 양성 예측도와 경보 처리 비용으로 측정되어야 한다. 데이터를 결정으로 바꾸는 일은 그 숫자가 어떤 모집단에서 나왔는지를 먼저 밝히는 데서 시작한다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-양성은-답이-아니라-질문의-시작이다" class="level2">
<h2 class="anchored" data-anchor-id="결론-양성은-답이-아니라-질문의-시작이다">결론: 양성은 답이 아니라 질문의 시작이다</h2>
<p>나는 학생들에게 검사 결과를 들으면 “정확도가 몇 %인가”보다 먼저 “이 병은 얼마나 흔한가”를 묻게 한다. 정확도 99%의 검사가 양성을 알려주면, 그것은 병이 있다는 뜻이 아니라 한 번 더 확인해야 한다는 신호다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>민감도·특이도 예시(99%, 유병률 0.1%)는 설명을 위한 가상의 숫자다. 실제 검사의 성능은 검사 종류와 조건에 따라 크게 다르다.</li>
<li>신속항원검사의 임상 민감도가 공식 수치보다 낮게 보고된 논란은 다음 기사를 참고했다: <a href="https://www.medicaltimes.com/Mobile/News/NewsView.html?ID=1145548">신속항원검사 민감도 이슈 쟁점화 — 메디칼타임즈</a></li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_base_rate_test.html</guid>
  <pubDate>Sun, 04 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률의 도구들) 복권 기댓값으로 알아보기 — 기댓값이란 무엇인가</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_expected_value_lotto.html</link>
  <description><![CDATA[ 




<section id="원을-내면-평균-얼마를-돌려받나" class="level2">
<h2 class="anchored" data-anchor-id="원을-내면-평균-얼마를-돌려받나">1,000원을 내면 평균 얼마를 돌려받나</h2>
<p>로또 6/45 한 게임은 1,000원이다. 당첨금은 판매액의 50%로 정해져 있고, 5등은 5,000원, 4등은 50,000원이 고정이며 1~3등은 나머지를 당첨자 수로 나눠 받는다. 그렇다면 한 게임을 살 때마다 평균적으로 돌려받는 돈, 곧 <strong>기댓값</strong>은 얼마일까? 각 결과의 상금에 그 확률을 곱해서 더한 값이 기댓값이다.</p>
<p><img src="https://latex.codecogs.com/png.latex?E%5BX%5D=%5Csum_x%20x%5C,P(X=x)"></p>
</section>
<section id="계산해-보기-45등에서만-180원" class="level2">
<h2 class="anchored" data-anchor-id="계산해-보기-45등에서만-180원">계산해 보기: 4·5등에서만 180원</h2>
<ul>
<li>5등(번호 3개 일치): 확률 <img src="https://latex.codecogs.com/png.latex?%5Cbinom%7B6%7D%7B3%7D%5Cbinom%7B39%7D%7B3%7D/%5Cbinom%7B45%7D%7B6%7D=2.244%5C%25"> → 기여 <img src="https://latex.codecogs.com/png.latex?0.02244%5Ctimes5%7B,%7D000%5Capprox"> <strong>112원</strong></li>
<li>4등(4개 일치): 확률 <img src="https://latex.codecogs.com/png.latex?0.1365%5C%25"> → 기여 <img src="https://latex.codecogs.com/png.latex?0.001365%5Ctimes50%7B,%7D000%5Capprox"> <strong>68원</strong></li>
<li>1~3등: 확률은 1등 1/8,145,060, 3등 1/35,724 등으로 극히 작지만 상금이 크다. 판매액의 50%가 당첨금이므로 4·5등 몫(약 180원)을 뺀 나머지가 1~3등의 평균 기여분, 곧 <strong>약 320원</strong>이다.</li>
</ul>
<p>합치면 기댓값은 약 500원, 즉 1,000원을 내고 평균 500원을 돌려받는다. <strong>평균적으로 절반을 잃는 게임</strong>인 셈이다(이월된 회차 등에 따라 회차별로 달라질 수 있는 설계상의 평균이다).</p>
<div id="3f4f88d1" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_expected_value_lotto_files/figure-html/cell-2-output-1.png" width="758" height="259" class="figure-img"></p>
<figcaption>Where the average 1,000 won goes: about 112 won from 5th prizes, 68 won from 4th prizes, roughly 320 won from 1st-3rd prizes (implied by the 50% payout rule), and about 500 won not returned to players.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="한-사람의-1년은-어떨까" class="level2">
<h2 class="anchored" data-anchor-id="한-사람의-1년은-어떨까">한 사람의 1년은 어떨까</h2>
<p>기댓값은 ’평균’이지 ’내가 얻을 값’이 아니다. 매주 5게임씩 1년(260게임)을 사는 사람 1만 명을 시뮬레이션해 보자. 1~3등은 확률이 너무 작아 대부분의 사람에게 한 번도 일어나지 않으므로, 여기서는 4·5등 상금만 세어 본다. 260,000원을 쓰고 4·5등에서 돌려받는 금액은 평균 약 46,300원이고, 거의 모든 사람이 쓴 돈의 절반에도 못 미친다.</p>
<div id="5089c62a" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_expected_value_lotto_files/figure-html/cell-3-output-1.png" width="710" height="432" class="figure-img"></p>
<figcaption>10,000 simulated players who each buy 260 lotto games (5 per week for a year). Histogram of the money returned from 4th and 5th prizes only, versus the 260,000 won spent. Almost nobody gets back even half of what they spent from these prizes.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="기댓값이-말해주는-것-말해주지-않는-것" class="level2">
<h2 class="anchored" data-anchor-id="기댓값이-말해주는-것-말해주지-않는-것">기댓값이 말해주는 것, 말해주지 않는 것</h2>
<p>기댓값이 500원이라는 사실은 이 게임이 <strong>장기적으로 반복하면 평균 손실이 발생하도록 설계</strong>되었음을 알려준다. 그러나 개인에게 일어나는 일은 훨씬 들쭉날쭉하다. 대부분은 소액을 받거나 아무것도 받지 못하고, 극소수만 수억 원을 받는다. 기댓값은 그 극단값까지 확률로 가중해 하나의 숫자로 요약한 것이라 대부분의 사람이 실제로 겪는 경험과는 다를 수 있다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
기댓값을 읽을 때
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>기댓값 = 각 결과 × 확률의 합.</strong> 큰 상금이라도 확률이 작으면 기여는 작다.</li>
<li><strong>기댓값은 ’반복했을 때의 평균’이지 한 번의 결과가 아니다.</strong></li>
<li><strong>분포의 모양을 함께 보라.</strong> 극단값이 평균을 끌어올리면 대부분은 평균에 못 미친다.</li>
<li><strong>비용과 비교하라.</strong> 기댓값이 지불 금액보다 작으면 평균적으로 손해다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-기댓값의-정의" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-기댓값의-정의">더 깊이: 기댓값의 정의</h2>
<p>강의노트 <a href="../../notes/math_stat/random_variable.html#기대값-정의">〈수리통계학 2. 확률변수 — 기대값 정의〉</a>는 이산형 확률변수의 기대값을 <img src="https://latex.codecogs.com/png.latex?%5Csum%20g(x)f_X(x)">로 정의한다. 오늘 계산은 <img src="https://latex.codecogs.com/png.latex?g(x)">를 상금으로, <img src="https://latex.codecogs.com/png.latex?f_X(x)">를 등수별 당첨 확률로 놓은 그대로의 적용이다.</p>
</section>
<section id="sdv-관점-평균-한-줄이-감추는-분포를-함께-읽어라" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-평균-한-줄이-감추는-분포를-함께-읽어라">SDV 관점: 평균 한 줄이 감추는 분포를 함께 읽어라</h2>
<p>기댓값을 계산하는 일은 <strong>요약과 추론(1~2단계)</strong>에 속한다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 핵심은, “평균 수익률 5%”, “고객당 평균 매출 3만 원” 같은 숫자가 오늘의 시뮬레이션처럼 <strong>소수의 큰 값이 평균을 끌어올린 결과</strong>일 수 있다는 점이다. 평균만 보고 의사결정을 하면 대부분의 고객이 경험하는 현실과 괴리가 생긴다. 기댓값과 함께 분포, 곧 대부분의 사람이 얻는 값과 극단의 몫을 나란히 보여줄 때 숫자가 비로소 결정의 가치가 된다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-평균이-500원이면-나는-대개-500원보다-못-받는다" class="level2">
<h2 class="anchored" data-anchor-id="결론-평균이-500원이면-나는-대개-500원보다-못-받는다">결론: 평균이 500원이면, 나는 대개 500원보다 못 받는다</h2>
<p>나는 학생들에게 복권을 사지 말라고 하지는 않는다. 다만 1,000원이 평균 500원으로 돌아온다는 사실, 그리고 대부분의 사람은 그 평균에도 못 미친다는 사실을 알고 사라고 말한다. 기댓값은 미래를 맞추는 예언이 아니라, 반복되는 게임의 구조를 한 숫자로 드러내는 도구다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li><a href="https://www.dhlottery.co.kr/lt645/intro">로또6/45 소개 — 동행복권</a>: 당첨금 총액은 판매액의 50%, 5등 5,000원·4등 50,000원 고정.</li>
<li><a href="https://www.dhlottery.co.kr/guide/wnrGuide">고객센터 당첨자 가이드 — 동행복권</a></li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_expected_value_lotto.html</guid>
  <pubDate>Thu, 01 Oct 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률의 도구들) 오해하기 쉬운 개념 — 독립이라는 것의 진짜 의미</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_independence_meaning.html</link>
  <description><![CDATA[ 




<section id="겹치면-독립이-아닐까-안-겹치면-독립일까" class="level2">
<h2 class="anchored" data-anchor-id="겹치면-독립이-아닐까-안-겹치면-독립일까">겹치면 독립이 아닐까, 안 겹치면 독립일까</h2>
<p>’독립’이라는 말은 일상에서 ’서로 영향을 주지 않는다’는 느낌으로 쓰인다. 그런데 확률에서의 독립은 훨씬 엄밀하다. 아래 세 가지 직관은 모두 틀렸다.</p>
<ul>
<li>“동시에 일어날 수 없으면 독립이다” → 틀렸다. 오히려 정반대다.</li>
<li>“겹치는 부분이 있으면 독립이 아니다” → 틀렸다.</li>
<li>“독립이니까 서로 아무 관련이 없다” → 관련이 없다는 뜻은 맞지만, 판단은 그림이 아니라 계산으로 한다.</li>
</ul>
<p>독립의 정의는 한 줄이다. 두 사건 <img src="https://latex.codecogs.com/png.latex?A,%20B">가 독립이라는 것은 <img src="https://latex.codecogs.com/png.latex?P(A%5Ccap%20B)=P(A)P(B)">가 성립한다는 뜻이다.</p>
</section>
<section id="카드-한-장으로-세-가지-관계를-확인하기" class="level2">
<h2 class="anchored" data-anchor-id="카드-한-장으로-세-가지-관계를-확인하기">카드 한 장으로 세 가지 관계를 확인하기</h2>
<p>52장 카드에서 한 장을 뽑는다고 하자. 하트를 <img src="https://latex.codecogs.com/png.latex?A">, 그림 카드를 <img src="https://latex.codecogs.com/png.latex?B">, 스페이드를 <img src="https://latex.codecogs.com/png.latex?C">, 빨간색 카드를 <img src="https://latex.codecogs.com/png.latex?D">라고 하면 세 쌍의 관계는 이렇다.</p>
<table class="caption-top table">
<thead>
<tr class="header">
<th>쌍</th>
<th><img src="https://latex.codecogs.com/png.latex?P(A%5Ccap%20B)"></th>
<th><img src="https://latex.codecogs.com/png.latex?P(A)P(B)"></th>
<th>판정</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>하트 &amp; 그림 카드</td>
<td>3/52 = 5.77%</td>
<td>(13/52)(12/52) = 5.77%</td>
<td><strong>독립</strong></td>
</tr>
<tr class="even">
<td>하트 &amp; 스페이드</td>
<td>0</td>
<td>(13/52)(13/52) = 6.25%</td>
<td>독립 아님</td>
</tr>
<tr class="odd">
<td>하트 &amp; 빨간색</td>
<td>13/52 = 25%</td>
<td>(13/52)(26/52) = 12.5%</td>
<td>독립 아님</td>
</tr>
</tbody>
</table>
<p>하트와 그림 카드는 3장이 겹치는데도 독립이다. 하트 카드 가운데 그림 카드의 비율(3/13)이 전체 카드 중 그림 카드의 비율(12/52)과 같기 때문이다. 하트인지 알아도 그림 카드일 확률이 달라지지 않는다. 반대로 하트와 스페이드는 동시에 일어날 수 없는(배반) 사건이지만, 하트가 나왔다는 사실을 알면 스페이드일 확률이 0이 되므로 오히려 <strong>강하게 종속</strong>이다.</p>
<div id="9a391dbc" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_independence_meaning_files/figure-html/cell-2-output-1.png" width="701" height="416" class="figure-img"></p>
<figcaption>For three pairs of events drawn from a 52-card deck: the actual joint probability P(A and B) versus the product P(A) x P(B). Independence means the two bars are equal. Heart &amp; face card are independent despite overlapping; heart &amp; spade never occur together yet are strongly dependent.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="조건부확률로-다시-읽기" class="level2">
<h2 class="anchored" data-anchor-id="조건부확률로-다시-읽기">조건부확률로 다시 읽기</h2>
<p>독립은 ’알려주는 정보가 없다’는 뜻으로 다시 쓸 수 있다. <img src="https://latex.codecogs.com/png.latex?P(B%5Cmid%20A)=P(B)">, 곧 <img src="https://latex.codecogs.com/png.latex?A">가 일어났다는 사실을 알아도 <img src="https://latex.codecogs.com/png.latex?B">의 확률이 바뀌지 않는다. 스페이드 <img src="https://latex.codecogs.com/png.latex?C">에서는 <img src="https://latex.codecogs.com/png.latex?P(C%5Cmid%20A)=0%5Cneq%20P(C)=25%5C%25">이므로 하트라는 정보가 스페이드 확률을 통째로 바꾼다. 하트는 빨간색이라는 사실을 확정해 주니 <img src="https://latex.codecogs.com/png.latex?P(D%5Cmid%20A)=100%5C%25%5Cneq50%5C%25">다.</p>
<div id="7dd7a107" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_independence_meaning_files/figure-html/cell-3-output-1.png" width="662" height="412" class="figure-img"></p>
<figcaption>How much learning ‘the card is a heart’ changes the probability of another event: P(event | heart) versus P(event). Only for ‘face card’ is there no change - that is what independence means.</figcaption>
</figure>
</div>
</div>
</div>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
독립인지 확인하는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>그림이 아니라 곱셈으로 확인하라.</strong> <img src="https://latex.codecogs.com/png.latex?P(A%5Ccap%20B)=P(A)P(B)">가 성립하는가.</li>
<li><strong>배반(동시에 불가능)과 독립을 구분하라.</strong> 확률이 0보다 큰 두 사건이 배반이면 오히려 독립이 아니다.</li>
<li><strong>’알면 바뀌는가’로 물어보라.</strong> <img src="https://latex.codecogs.com/png.latex?A">를 알 때 <img src="https://latex.codecogs.com/png.latex?B">의 확률이 그대로인가.</li>
<li><strong>독립은 가정이 아니라 검증 대상이다.</strong> 실제 데이터에서는 곱셈이 근사적으로 맞는지 본다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-독립과-배반의-구분" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-독립과-배반의-구분">더 깊이: 독립과 배반의 구분</h2>
<p>앞서 09월 22일 글에서는 로또 추첨의 독립성을 다뤘다. 강의노트 <a href="../../notes/math_stat/probability.html#독립">〈수리통계학 1. 확률론 — 독립〉</a>에는 독립의 정의와 함께, 확률이 양수인 두 사건이 배반이면 독립일 수 없다는 증명이 실려 있다. 오늘 카드 예시의 하트와 스페이드가 바로 그 경우다.</p>
</section>
<section id="sdv-관점-독립-가정이-틀리면-결정이-무너진다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-독립-가정이-틀리면-결정이-무너진다">SDV 관점: 독립 가정이 틀리면 결정이 무너진다</h2>
<p>독립을 검증하는 일은 <strong>요약과 추론(1~2단계)</strong>의 기본 절차다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서는 결정의 안전성을 좌우하는 문제다. 리스크 관리에서 “각 부품의 고장은 독립”이라고 가정하고 확률을 곱해 전체 고장률을 계산했다가, 사실은 같은 전원이나 같은 공급처 때문에 함께 고장 나는 종속 구조라서 실제 위험이 수십 배 커지는 일이 반복된다. 곱셈 규칙을 적용하기 전에 “이 두 사건은 정말 서로 정보를 주지 않는가”를 데이터로 확인하는 것이, 숫자에 가치를 더하는 통계학자의 검증 절차다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-독립은-느낌이-아니라-등식이다" class="level2">
<h2 class="anchored" data-anchor-id="결론-독립은-느낌이-아니라-등식이다">결론: 독립은 느낌이 아니라 등식이다</h2>
<p>나는 학생들에게 “서로 상관없어 보인다”는 말을 믿지 말고 곱해보라고 말한다. 하트와 그림 카드는 눈으로는 얽혀 있어도 독립이고, 하트와 스페이드는 눈으로는 남남 같아도 종속이다. 독립은 직관이 아니라 <img src="https://latex.codecogs.com/png.latex?P(A%5Ccap%20B)=P(A)P(B)">라는 하나의 등식으로 판정한다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>트럼프 카드(52장) 구성에 기반한 이론적 계산이며 외부 통계 자료는 사용하지 않았다.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_independence_meaning.html</guid>
  <pubDate>Wed, 30 Sep 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률의 도구들) 경우의 수 세기 — 카드 한 장을 뽑을 때 확률 계산법</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_counting_cards.html</link>
  <description><![CDATA[ 




<section id="첫-패에-광이-석-장" class="level2">
<h2 class="anchored" data-anchor-id="첫-패에-광이-석-장">첫 패에 광이 석 장</h2>
<p>고스톱을 치다 첫 패에 광이 세 장 들어오면 “오늘 운이 좋다”는 말이 절로 나온다. 얼마나 운이 좋은 걸까? 화투는 12개월 × 4장 = 48장이고 그중 광은 5장이다. 카드를 뽑는 확률 문제는 결국 <strong>경우의 수를 세는 문제</strong>다. 라플라스의 정의로 <img src="https://latex.codecogs.com/png.latex?P(A)=n(A)/n(S)">, 곧 ’해당하는 경우의 수 ÷ 전체 경우의 수’를 계산하면 된다.</p>
</section>
<section id="한-장을-뽑을-때-장수만-세면-된다" class="level2">
<h2 class="anchored" data-anchor-id="한-장을-뽑을-때-장수만-세면-된다">한 장을 뽑을 때: 장수만 세면 된다</h2>
<p>화투 한 장을 뽑아서 광일 확률은 <img src="https://latex.codecogs.com/png.latex?5/48%5Capprox10.4%5C%25">, 특정 월(예: 8월)의 카드가 나올 확률은 <img src="https://latex.codecogs.com/png.latex?4/48%5Capprox8.3%5C%25">다. 트럼프 52장으로 옮겨도 마찬가지다. 하트일 확률은 <img src="https://latex.codecogs.com/png.latex?13/52">, 그림 카드(J, Q, K)일 확률은 <img src="https://latex.codecogs.com/png.latex?12/52">다. 그렇다면 “하트이거나 그림 카드”일 확률은 <img src="https://latex.codecogs.com/png.latex?25/52">일까? 아니다. 하트 그림 카드 3장이 두 번 세어지기 때문이다.</p>
<p><img src="https://latex.codecogs.com/png.latex?P(A%5Ccup%20B)=P(A)+P(B)-P(A%5Ccap%20B)=%5Ctfrac%7B13%7D%7B52%7D+%5Ctfrac%7B12%7D%7B52%7D-%5Ctfrac%7B3%7D%7B52%7D=%5Ctfrac%7B22%7D%7B52%7D"></p>
<div id="0e5507cc" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_counting_cards_files/figure-html/cell-2-output-1.png" width="854" height="338" class="figure-img"></p>
<figcaption>The 52 cards as a 4x13 grid. Hearts (13) and face cards (12) overlap in 3 cards, so ‘heart or face card’ contains 13 + 12 - 3 = 22 cards, not 25.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="여러-장을-뽑을-때-조합으로-센다" class="level2">
<h2 class="anchored" data-anchor-id="여러-장을-뽑을-때-조합으로-센다">여러 장을 뽑을 때: 조합으로 센다</h2>
<p>화투 패 10장을 받는 경우는 <img src="https://latex.codecogs.com/png.latex?%5Cbinom%7B48%7D%7B10%7D">가지이고, 그중 광 5장에서 <img src="https://latex.codecogs.com/png.latex?k">장, 나머지 43장에서 <img src="https://latex.codecogs.com/png.latex?10-k">장을 고르는 경우의 수는 <img src="https://latex.codecogs.com/png.latex?%5Cbinom%7B5%7D%7Bk%7D%5Cbinom%7B43%7D%7B10-k%7D">다. 이 비율이 첫 패의 광 개수 분포다. 계산해보면 광이 한 장도 없을 확률이 29.3%, 한 장 43.1%, 두 장 22.2%, 세 장 4.9%, 네 장 0.47%, 다섯 장 0.01% 미만이다. 세 장 이상은 합쳐서 <strong>약 5.4%</strong>, 스무 판에 한 번꼴이다.</p>
<div id="2be02084" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_counting_cards_files/figure-html/cell-3-output-1.png" width="662" height="413" class="figure-img"></p>
<figcaption>Number of ‘gwang’ (bright) cards among a 10-card starting hand from a 48-card hwatu deck containing 5 of them (hypergeometric probabilities). Three or more bright cards occur about 5.4% of the time.</figcaption>
</figure>
</div>
</div>
</div>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
카드 확률 계산 순서
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>전체 경우의 수를 먼저 정하라.</strong> 한 장이면 장수, 여러 장이면 <img src="https://latex.codecogs.com/png.latex?%5Cbinom%7BN%7D%7Bn%7D">이다.</li>
<li><strong>사건에 해당하는 경우의 수를 센다.</strong> 여러 조건이 겹치면 곱의 법칙과 합의 법칙을 나눠 쓴다.</li>
<li><strong>’또는’은 겹침을 빼라.</strong> <img src="https://latex.codecogs.com/png.latex?P(A%5Ccup%20B)=P(A)+P(B)-P(A%5Ccap%20B)">.</li>
<li><strong>순서가 중요한지 확인하라.</strong> 손에 든 패는 순서가 없으므로 조합을 쓴다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-곱의-법칙과-조합" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-곱의-법칙과-조합">더 깊이: 곱의 법칙과 조합</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#확률계산-방법">〈수리통계학 1. 확률론 — 확률계산 방법〉</a>은 표본점을 세어 <img src="https://latex.codecogs.com/png.latex?P(A)=n/N">을 구하는 Sample-Point 방법과, 곱의 법칙·순열·조합으로 경우의 수를 요약한 표를 제공한다. 오늘의 화투 계산은 그 표의 조합 공식 하나를 그대로 적용한 것이다.</p>
</section>
<section id="sdv-관점-세는-단위를-바꾸면-가치가-보인다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-세는-단위를-바꾸면-가치가-보인다">SDV 관점: 세는 단위를 바꾸면 가치가 보인다</h2>
<p>경우의 수를 정확히 세는 일은 <strong>요약과 추론(1~2단계)</strong>의 기초다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이 훈련이 가치가 되는 이유는, 현업의 ‘조합’ 문제가 카드 패와 같은 구조이기 때문이다. 상품 묶음 구성, 프로모션 조합, A/B 테스트의 셀 개수는 모두 조합으로 폭발한다. 가능한 경우가 몇 개인지 먼저 세어보면, 어떤 조합을 전수 검토할 수 있고 어떤 조합은 표본 설계로 줄여야 하는지가 드러난다. 세기는 자원을 어디에 쓸지 정하는 설계의 출발점이다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-확률은-곱셈과-덧셈으로-센-결과다" class="level2">
<h2 class="anchored" data-anchor-id="결론-확률은-곱셈과-덧셈으로-센-결과다">결론: 확률은 곱셈과 덧셈으로 센 결과다</h2>
<p>나는 학생들에게 확률 공식을 외우지 말고 ’카드를 펼쳐놓고 세어보라’고 말한다. 5.4%라는 숫자는 마법이 아니라 <img src="https://latex.codecogs.com/png.latex?%5Cbinom%7B5%7D%7B3%7D%5Cbinom%7B43%7D%7B7%7D"> 같은 경우의 수를 센 결과다. 무엇을 셀지, 겹치는 것은 없는지만 챙기면 카드 한 장에서 열 장까지 같은 원리다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>화투는 12개월 × 4장 = 48장이며 광은 5장(1·3·8·11·12월)이라는 구성은 일반적으로 알려진 화투 규칙에 따른다.</li>
<li>트럼프 카드(52장)의 구성: 4무늬 × 13숫자.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_counting_cards.html</guid>
  <pubDate>Tue, 29 Sep 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률의 도구들) “적어도 하나”의 확률 — 여사건으로 풀면 쉬워지는 문제</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_at_least_one.html</link>
  <description><![CDATA[ 




<section id="확률로-100번-뽑으면-나올까" class="level2">
<h2 class="anchored" data-anchor-id="확률로-100번-뽑으면-나올까">1% 확률로 100번 뽑으면 나올까</h2>
<p>모바일 게임의 ’확률형 아이템’에서 원하는 캐릭터가 나올 확률이 1%라고 하자. 100번 뽑으면 100 × 1% = 1이니 한 번은 나오는 것 아닐까? 많은 사람이 이렇게 직관적으로 계산하지만, 실제로 적어도 한 번 나올 확률은 <strong>63.4%</strong>에 불과하다. 100번을 뽑고도 빈손일 확률이 36.6%나 된다. 2024년 3월 22일부터 국내에서는 게임산업법에 따라 확률형 아이템의 확률 정보를 의무적으로 공개해야 한다. 이제 공개된 ’1%’라는 숫자를 어떻게 읽어야 하는지가 중요해졌다.</p>
</section>
<section id="적어도-하나는-통째로-세지-말고-뒤집어라" class="level2">
<h2 class="anchored" data-anchor-id="적어도-하나는-통째로-세지-말고-뒤집어라">“적어도 하나”는 통째로 세지 말고 뒤집어라</h2>
<p>‘적어도 한 번 나온다’는 사건은 ’한 번 나온다’, ‘두 번 나온다’, ‘세 번 나온다’… 를 모두 더해야 해서 복잡하다. 그런데 이 사건의 <strong>여사건</strong>, 곧 ’한 번도 안 나온다’는 단 하나의 경우다. 뽑기가 서로 독립이라면 100번 연속 실패할 확률은 <img src="https://latex.codecogs.com/png.latex?0.99%5E%7B100%7D=36.6%5C%25">이고, 구하려던 확률은</p>
<p><img src="https://latex.codecogs.com/png.latex?P(%5Ctext%7B%EC%A0%81%EC%96%B4%EB%8F%84%20%ED%95%9C%20%EB%B2%88%7D)%20=%201%20-%20P(%5Ctext%7B%ED%95%9C%20%EB%B2%88%EB%8F%84%20%EC%97%86%EC%9D%8C%7D)%20=%201-0.99%5E%7B100%7D%20=%2063.4%5C%25"></p>
<p>로 한 줄에 끝난다. 복잡한 합을 피하려고 반대편에서 세는 것, 이것이 여사건의 힘이다.</p>
<div id="dc7ad585" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_at_least_one_files/figure-html/cell-2-output-1.png" width="710" height="432" class="figure-img"></p>
<figcaption>Probability of getting the item at least once, as a function of the number of pulls, for three published drop rates (0.5%, 1%, 3%). Even at 100 pulls, a 1% rate gives only 63.4%.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="몇-번을-뽑아야-나올-만하다고-말할-수-있나" class="level2">
<h2 class="anchored" data-anchor-id="몇-번을-뽑아야-나올-만하다고-말할-수-있나">몇 번을 뽑아야 ’나올 만하다’고 말할 수 있나</h2>
<p>절반의 확률로 만나려면 <img src="https://latex.codecogs.com/png.latex?0.99%5En%20%5Cle%200.5">를 풀어 <img src="https://latex.codecogs.com/png.latex?n%5Cge%2069">번, 90%를 확신하려면 <img src="https://latex.codecogs.com/png.latex?n%5Cge%20230">번이 필요하다. 확률이 0.5%로 절반이 되면 필요한 횟수는 거의 두 배(139번, 460번)로 뛴다. 그래서 일부 게임은 일정 횟수를 채우면 당첨을 보장하는 ‘천장’ 제도를 두는데, 이는 운이 나쁜 이용자의 꼬리를 잘라내는 장치다. 천장이 없다면 230번을 뽑고도 10%의 사람은 여전히 빈손이다.</p>
<div id="0fd43f03" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_at_least_one_files/figure-html/cell-3-output-1.png" width="681" height="413" class="figure-img"></p>
<figcaption>Number of pulls needed to reach a 50% or 90% chance of at least one success, for three drop rates. Halving the drop rate roughly doubles the required pulls.</figcaption>
</figure>
</div>
</div>
</div>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
“적어도 하나”를 만나면
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>’적어도 하나’가 보이면 여사건으로 뒤집어라.</strong> <img src="https://latex.codecogs.com/png.latex?1-P(%5Ctext%7B%ED%95%98%EB%82%98%EB%8F%84%20%EC%97%86%EC%9D%8C%7D)">이 대개 가장 빠른 길이다.</li>
<li><strong>시도가 독립인지 확인하라.</strong> 독립일 때만 실패 확률을 거듭제곱으로 곱할 수 있다.</li>
<li><strong>기댓값 1과 확률 100%를 혼동하지 마라.</strong> 평균적으로 한 번 나오는 횟수(100번)를 뽑아도 빈손일 확률이 37%가 남는다.</li>
<li><strong>’몇 번 해야 절반인가’로 바꿔 물어보라.</strong> <img src="https://latex.codecogs.com/png.latex?n%20%5Capprox%200.69/p">라는 어림값이 유용하다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-여사건과-확률의-기본-성질" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-여사건과-확률의-기본-성질">더 깊이: 여사건과 확률의 기본 성질</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#확률계산-관련-정리">〈수리통계학 1. 확률론 — 확률계산 관련 정리〉</a>는 <img src="https://latex.codecogs.com/png.latex?P(A%5Ec)=1-P(A)">를 확률의 기본 정리로 정리한다. 오늘 계산은 이 한 줄에 독립사건의 곱셈 규칙을 더한 것뿐이다. 앞서 살펴본 생일 문제도 같은 뿌리에서 자란 문제였다.</p>
</section>
<section id="sdv-관점-평균-1회가-아니라-한-번도-못-만날-확률을-설계하라" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-평균-1회가-아니라-한-번도-못-만날-확률을-설계하라">SDV 관점: “평균 1회”가 아니라 “한 번도 못 만날 확률”을 설계하라</h2>
<p>여사건 계산은 <strong>요약과 추론(1~2단계)</strong>의 기본기다. <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 중요한 대목은, 이 계산이 곧 ’고객이 느끼는 실패 확률’을 정량화하는 도구라는 점이다. 100번 시도하면 평균 한 번 성공한다는 보고서는 참이지만, 고객의 37%가 끝까지 빈손이라는 사실은 감춘다. 이탈률, 클레임, 이용자 만족을 설계할 때 평균 성공 횟수 대신 ’한 번도 성공하지 못한 사람의 비율’을 함께 보아야 서비스의 가치가 왜곡 없이 전달된다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-없음의-확률을-먼저-구하라" class="level2">
<h2 class="anchored" data-anchor-id="결론-없음의-확률을-먼저-구하라">결론: 없음의 확률을 먼저 구하라</h2>
<p>나는 학생들에게 ’적어도’라는 단어가 나오면 일단 연필을 내려놓고 “하나도 없을 확률이 얼마인가”부터 물으라고 가르친다. 1%의 확률은 100번 시도로 보장되지 않는다. 표시된 확률 숫자보다, 그 확률에 몇 번을 도전해야 하는지가 실제 경험을 결정한다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li><a href="https://www.korea.kr/news/policyNewsView.do?newsId=148924297">게임 확률형 아이템 정보, 3월 22일부터 투명하게 공개된다 — 대한민국 정책브리핑</a></li>
<li><a href="https://www.ajunews.com/view/20240102104711033">확률형 아이템 확률 공개 의무화, 올해 3월 22일부터 — 아주경제</a></li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_at_least_one.html</guid>
  <pubDate>Mon, 28 Sep 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률의 도구들) 확률을 세는 언어 — 표본공간과 사건이란 무엇인가</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_sample_space_events.html</link>
  <description><![CDATA[ 




<section id="윷이-나올-확률은-5분의-1일까" class="level2">
<h2 class="anchored" data-anchor-id="윷이-나올-확률은-5분의-1일까">윷이 나올 확률은 5분의 1일까</h2>
<p>설날 가족 윷놀이에서 누군가 “윷이 나올 확률은 다섯 가지 중 하나니까 20%지”라고 말했다면, 틀린 말이다. 결과의 종류가 다섯 가지(도·개·걸·윷·모)라는 것과, 다섯 결과가 <strong>같은 확률</strong>이라는 것은 전혀 다른 이야기다. 윷가락 하나가 평평한 면과 둥근 면이 나올 확률을 각각 50%라고 가정하면, 개는 37.5%, 도와 걸은 각각 25%, 윷과 모는 각각 6.25%다. 개가 윷보다 여섯 배 자주 나온다. 어떻게 이런 계산이 나올까? 답은 ’표본공간’을 제대로 세는 데 있다.</p>
</section>
<section id="표본공간-일어날-수-있는-모든-일의-목록" class="level2">
<h2 class="anchored" data-anchor-id="표본공간-일어날-수-있는-모든-일의-목록">표본공간: 일어날 수 있는 모든 일의 목록</h2>
<p>확률에서 가장 먼저 하는 일은 실험을 던져서 나올 수 있는 <strong>모든 결과의 집합</strong>, 곧 <strong>표본공간(sample space)</strong>을 적는 것이다. 윷가락 네 개를 구별해서 각각 평평한 면(F) 또는 둥근 면(R)이 나온다고 하면, 표본공간은 <img src="https://latex.codecogs.com/png.latex?2%5E4=16">가지다. 도·개·걸·윷·모는 이 16가지 결과를 ’평평한 면이 몇 개인가’로 묶어 이름 붙인 것에 지나지 않는다. 이렇게 표본공간의 일부분을 묶은 것을 <strong>사건(event)</strong>이라 부른다.</p>
<div id="404fd0b1" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_sample_space_events_files/figure-html/cell-2-output-1.png" width="394" height="413" class="figure-img"></p>
<figcaption>The 16 equally likely outcomes of throwing four yut sticks (F = flat side up, R = round side up), colored by the named result. The five results are events of very different sizes: 4, 6, 4, 1 and 1 outcomes.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="사건의-크기가-곧-확률이다" class="level2">
<h2 class="anchored" data-anchor-id="사건의-크기가-곧-확률이다">사건의 크기가 곧 확률이다</h2>
<p>모든 결과가 같은 가능성을 갖는다는 가정(균등가능성) 아래에서는 사건 <img src="https://latex.codecogs.com/png.latex?A">의 확률이 <img src="https://latex.codecogs.com/png.latex?P(A)=%7CA%7C/%7CS%7C">, 곧 사건에 속한 결과의 수를 표본공간 전체의 수로 나눈 값이다. 윷은 FFFF 한 가지뿐이니 <img src="https://latex.codecogs.com/png.latex?1/16=6.25%5C%25">, 개는 평평한 면이 두 개인 <img src="https://latex.codecogs.com/png.latex?%5Cbinom%7B4%7D%7B2%7D=6">가지이니 <img src="https://latex.codecogs.com/png.latex?6/16=37.5%5C%25">다. ’결과의 종류가 다섯 가지’는 사건의 이름이 다섯 개라는 뜻일 뿐, 표본공간의 원소가 다섯 개라는 뜻이 아니다.</p>
<div id="65f0bd62" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_sample_space_events_files/figure-html/cell-3-output-1.png" width="662" height="413" class="figure-img"></p>
<figcaption>Probability of each yut result when each stick lands flat or round with probability 1/2, compared with the naive ‘1 in 5 = 20%’ guess (dashed line).</figcaption>
</figure>
</div>
</div>
</div>
<p>물론 실제 윷가락은 한쪽이 반원형이라 평평한 면과 둥근 면이 정확히 반반으로 나오지는 않는다. 여기서 계산한 값은 ’50 대 50’이라는 가정을 둔 이론값이며, 가정이 달라지면 표본공간의 각 원소 확률이 달라진다. 그래도 도·개·걸이 흔하고 윷·모가 귀하다는 큰 그림은 바뀌지 않는다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
확률 문제를 읽을 때 먼저 확인할 것
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>표본공간을 먼저 적어라.</strong> “무엇이 일어날 수 있는가”를 빠짐없이, 겹치지 않게 나열하는 일이 계산의 절반이다.</li>
<li><strong>결과의 ’종류’와 ’가능성’을 혼동하지 마라.</strong> 종류가 다섯이라고 각각 <img src="https://latex.codecogs.com/png.latex?1/5">인 것은 아니다.</li>
<li><strong>사건은 표본공간의 부분집합이다.</strong> 사건을 ’결과들의 묶음’으로 그려보면 확률은 묶음의 크기 비교가 된다.</li>
<li><strong>균등가능성 가정이 성립하는지 스스로 물어라.</strong> 성립하지 않으면 원소마다 확률을 따로 붙여야 한다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-집합의-언어로-쓰는-확률" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-집합의-언어로-쓰는-확률">더 깊이: 집합의 언어로 쓰는 확률</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#표본공간과-사건">〈수리통계학 1. 확률론 — 표본공간과 사건〉</a>은 표본공간을 실험의 모든 가능한 결과의 집합으로 정의하고, 사건을 그 부분집합으로 다룬다. 합집합·교집합·여집합 같은 집합 연산이 곧 ‘또는’·‘그리고’·’아니다’의 언어가 된다는 점이 앞으로 두 주 동안 쓸 모든 도구의 바탕이다.</p>
</section>
<section id="sdv-관점-데이터-분석의-첫-질문은-가능한-경우의-전체가-무엇인가" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-데이터-분석의-첫-질문은-가능한-경우의-전체가-무엇인가">SDV 관점: 데이터 분석의 첫 질문은 “가능한 경우의 전체가 무엇인가”</h2>
<p>표본공간을 정의하는 일은 <strong>요약과 추론(1~2단계)</strong>의 출발점이다. 제가 제안하는 <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이는 더 실용적인 뜻을 갖는다. “전환율 3%”, “불량률 0.5%” 같은 숫자는 ‘무엇을 한 번의 시도로 세었는가’, 곧 표본공간의 단위를 정해야 비로소 의미가 생긴다. 방문자 기준인지 세션 기준인지, 부품 기준인지 로트 기준인지에 따라 같은 데이터가 전혀 다른 확률을 만든다. 데이터에 가치를 더하는 첫 단계는 계산이 아니라, 숫자가 어떤 표본공간 위에서 정의됐는지를 못 박는 일이다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-세기-전에-나열하라" class="level2">
<h2 class="anchored" data-anchor-id="결론-세기-전에-나열하라">결론: 세기 전에 나열하라</h2>
<p>나는 학생들에게 확률 문제 앞에서 계산기를 두드리기 전에 표본공간부터 종이에 적어보라고 말한다. 윷놀이의 개가 윷보다 여섯 배 흔한 이유는 신비로운 것이 아니라, 개라는 사건에 속한 결과가 여섯 개이고 윷에는 하나뿐이라는 단순한 세기의 문제다. 확률은 결국 ’얼마나 많은 경우가 그 사건에 해당하는가’를 세는 언어다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li><a href="https://www.munhwa.com/article/10907513">윷가락 확률 ‘개 ＞ 걸 = 도 ＞ 윷 = 모’ 順 — 문화일보</a></li>
<li><a href="https://ko.wikipedia.org/wiki/%EC%9C%B7%EB%86%80%EC%9D%B4">윷놀이 — 위키백과</a></li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_sample_space_events.html</guid>
  <pubDate>Sun, 27 Sep 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률과 우연의 세계) 희귀한 사건이 매일 뉴스에 나오는 이유 — 다중 기회</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_multiple_chances.html</link>
  <description><![CDATA[ 




<section id="만분의-1의-확률-그런데-오늘도-뉴스에-나온다" class="level2">
<h2 class="anchored" data-anchor-id="만분의-1의-확률-그런데-오늘도-뉴스에-나온다">1000만분의 1의 확률, 그런데 오늘도 뉴스에 나온다</h2>
<p>“1,000만 명 중 한 명꼴로 일어나는 희귀한 사건”이라는 표현을 뉴스에서 종종 본다. 개인 한 사람에게 이 확률은 정말 작다. 그런데 이 확률을 가진 “기회”를 가진 사람이 한국에만 5,100만 명이라면 어떨까? 계산해보면, 오늘 하루 대한민국 어딘가에서 이 정도로 희귀한 사건이 적어도 한 번은 일어날 확률은 <strong>99% 이상</strong>이다. 개인에게는 상상하기 힘든 확률이, 사회 전체로 보면 사실상 “매일 일어나는 일”이 되는 이유다.</p>
<div id="f9b08ac0" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_multiple_chances_files/figure-html/cell-2-output-1.png" width="710" height="451" class="figure-img"></p>
<figcaption>The probability that a rare event (1-in-10-million chance per person, per day) happens to AT LEAST ONE person somewhere, as the size of the population with that daily ‘chance’ grows from 1,000 to South Korea’s full population (51 million). By the time the population reaches the low tens of millions, the event becomes very likely to happen to someone, somewhere, every single day.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="리틀우드의-기적의-법칙" class="level2">
<h2 class="anchored" data-anchor-id="리틀우드의-기적의-법칙">리틀우드의 “기적의 법칙”</h2>
<p>영국의 수학자 존 리틀우드(John Littlewood)는 이런 현상을 “기적의 법칙(Littlewood’s Law of Miracles)”이라 불렀다. 그가 말한 “기적”의 정의는 단순하다 — 개인에게 100만분의 1의 확률로 일어나는 사건이다. 그런데 사람은 깨어 있는 동안 대략 한 달이면 100만 번 가까운 크고 작은 “사건”(보고, 듣고, 마주치는 것들)을 경험한다고 어림잡을 수 있다. 그렇다면 100만분의 1의 확률로 일어나는 “기적”은, 한 사람에게 평균적으로 <strong>한 달에 한 번꼴로</strong> 일어나야 한다는 계산이 나온다. 물론 그 기적이 매번 나에게 일어나는 건 아니다 — 그 “한 달에 한 번”이라는 몫이, 지구상 수십억 명 중 누군가에게 돌아갈 뿐이다.</p>
<div id="ac419770" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_multiple_chances_files/figure-html/cell-3-output-1.png" width="758" height="412" class="figure-img"></p>
<figcaption>The same 1-in-10-million individual probability, compared with the expected number of times it actually occurs somewhere in South Korea on a given day (about 5.1 times, using the Poisson approximation with 51 million independent ‘chances’).</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="뉴스는-누구에게-일어났는지를-전하지-얼마나-드문지를-전하지-않는다" class="level2">
<h2 class="anchored" data-anchor-id="뉴스는-누구에게-일어났는지를-전하지-얼마나-드문지를-전하지-않는다">뉴스는 “누구에게” 일어났는지를 전하지, “얼마나 드문지”를 전하지 않는다</h2>
<p>언론이 희귀한 사건을 보도할 때, 그 사건이 “이 사람”에게 일어날 확률이 아니라 “이 사건이 실제로 일어났다”는 사실 자체를 전한다. 그런데 5,100만 명이 저마다 그 사건의 “기회”를 하나씩 갖고 있다면, 그중 누군가에게 일어나는 것은 전혀 이상한 일이 아니다. 로또 1등이 매주 나오는 것도, 벼락에 맞는 사람이 해마다 몇 명씩 나오는 것도, 희귀 질환이 매달 어디선가 새로 진단되는 것도 같은 원리다 — 개인의 확률은 극히 낮아도, 그 확률을 가진 사람의 수가 충분히 많으면 “누군가”에게는 반드시 일어난다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
“희귀한 사건” 뉴스를 읽는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>“이 사람에게 일어날 확률”과 “누군가에게 일어날 확률”을 구분하라.</strong> 전자는 낮아도 후자는 매우 높을 수 있다.</li>
<li><strong>그 사건의 “기회”를 가진 모집단이 얼마나 큰지 생각해보라.</strong> 모집단이 클수록, 아무리 희귀한 사건도 “매일 어디선가는 일어나는 일”이 된다.</li>
<li><strong>“이런 우연이?”라는 반응이 나오면, 그 우연이 일어날 수 있었던 전체 기회의 수를 먼저 세어보라.</strong> 이번 주 다룬 생일 문제, 로또 명당 이야기와 정확히 같은 논리다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-여러-사건의-합집합은-각-확률의-합을-넘지-않는다" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-여러-사건의-합집합은-각-확률의-합을-넘지-않는다">더 깊이: 여러 사건의 합집합은 각 확률의 합을 넘지 않는다</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#확률계산-관련-정리">〈수리통계학 1. 확률론 — 확률계산 관련 정리〉</a>가 소개하는 <strong>불의 부등식(Boole’s Inequality)</strong>은 <img src="https://latex.codecogs.com/png.latex?P%5Cleft(%5Cbigcup_%7Bi=1%7D%5E%7B%5Cinfty%7D%20A_i%5Cright)%20%5Cleq%20%5Csum_%7Bi=1%7D%5E%7B%5Cinfty%7D%20P(A_i)">로, 여러 사건 중 “적어도 하나가 일어날 확률”이 각 사건 확률의 합을 넘지 않는다는 걸 보장한다. 오늘 계산에서 5,100만 명 각각을 “그 사람에게 사건이 일어날 확률”을 가진 개별 사건으로 보면, 그 합(약 5.1)이 바로 “누군가에게는 일어날 것으로 기대되는 횟수”의 상한에 가까운 근사값이 된다. 확률이 아무리 작아도, 더해지는 항의 개수(기회의 수)가 충분히 많으면 그 합은 얼마든지 커질 수 있다 — 이것이 희귀한 사건이 큰 모집단에서 흔해지는 수학적 이유다.</p>
</section>
<section id="sdv-관점-희귀-사건-관리는-개인-확률이-아니라-모집단-규모로-설계해야-한다" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-희귀-사건-관리는-개인-확률이-아니라-모집단-규모로-설계해야-한다">SDV 관점: 희귀 사건 관리는 개인 확률이 아니라 모집단 규모로 설계해야 한다</h2>
<p>개인 확률과 모집단 전체의 기대 발생 횟수를 구분하는 일은 <strong>요약과 추론(1~2단계)</strong>에 해당한다. 제가 제안하는 <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이 구분이 중요한 이유는, 리스크 관리·품질 관리·이상 탐지 설계 전반에 직결되기 때문이다. “이 정도 확률이면 안전하다”는 판단은 개인(또는 거래 한 건) 기준의 확률만으로는 부족하다. 그 확률에 노출되는 모집단(고객 수, 거래 건수, 부품 수)이 얼마나 큰지를 함께 계산해야, “이번 달에 몇 건 정도는 발생할 것으로 예상된다”는 진짜 의사결정 근거가 나온다. 이번 주 다룬 생일 문제, 독립사건, 군집 착각, 명당의 정체, 그리고 오늘의 다중 기회는 모두 같은 메시지로 수렴한다 — <strong>확률은 언제나 “몇 번의 기회가 있었는가”와 함께 읽어야 한다.</strong> (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-확률이-작다고-안심하지-말고-기회의-수를-세어보라" class="level2">
<h2 class="anchored" data-anchor-id="결론-확률이-작다고-안심하지-말고-기회의-수를-세어보라">결론: 확률이 작다고 안심하지 말고, 기회의 수를 세어보라</h2>
<p>나는 학생들에게 “말도 안 되는 확률”이라는 표현을 들으면 딱 하나만 물어보라고 말한다. “그 확률을 시도할 기회가 몇 번이나 있었나요?” 이번 주 다섯 편의 글 — 생일 문제, 독립사건, 군집 착각, 로또 명당, 그리고 오늘의 다중 기회 — 은 결국 하나의 도구로 요약된다. 확률이 작아 보인다고 안심하기 전에, 그 작은 확률에 곱해질 기회의 수부터 세어보는 습관이다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>Littlewood, J. E. <em>Littlewood’s Miscellany</em> (1953, 1986 재발간) — “기적의 법칙”의 최초 출처로 널리 인용됨.</li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_multiple_chances.html</guid>
  <pubDate>Thu, 24 Sep 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률과 우연의 세계) 로또 번호에 명당이 존재할까? — 동일확률과 선택편향</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_lotto_hotspot.html</link>
  <description><![CDATA[ 




<section id="번-당첨된-편의점-정말-명당일까" class="level2">
<h2 class="anchored" data-anchor-id="번-당첨된-편의점-정말-명당일까">48번 당첨된 편의점, 정말 “명당”일까</h2>
<p>서울 노원구 상계동의 한 편의점은 로또 1등 당첨자를 48회나 배출해 전국 오프라인 판매점 중 최다 기록을 세웠다(2024년 기준). 주말이면 이 가게 앞에 긴 줄이 늘어선다. 다들 이 가게에서 사면 당첨 확률이 더 높다고 믿기 때문이다. 그런데 로또 6/45는 애초에 <strong>모든 번호 조합이 정확히 같은 확률(1/8,145,060)</strong>을 갖도록 설계된 게임이다. 어느 가게에서 사든, 어떤 요일에 사든, 자동이든 수동이든 확률은 달라지지 않는다. 그렇다면 이 가게는 왜 이렇게 많은 당첨자를 배출했을까?</p>
</section>
<section id="답은-운이-아니라-판매량이다" class="level2">
<h2 class="anchored" data-anchor-id="답은-운이-아니라-판매량이다">답은 운이 아니라 “판매량”이다</h2>
<p>복권위원회 관계자는 이 현상을 이렇게 설명한다. “구매자가 살 수 있는 번호조합은 814만 개뿐이지만, 이들이 사는 게임 수는 매주 1억 건이 넘는다. 예전에는 100명만 사던 조합을 지금은 1,000명이 구매한다면, 그 조합이 당첨됐을 때 당첨자 수가 많아질 수밖에 없다.” 즉 특정 가게에 당첨자가 몰리는 이유는 그 가게가 “행운의 장소”라서가 아니라, 그 가게가 압도적으로 많은 표를 팔기 때문이다. 표를 많이 팔수록 당첨 조합을 우연히 포함할 확률도 비례해서 커진다 — 이건 신비로운 힘이 아니라 단순한 곱셈이다.</p>
<div id="fe347d44" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_lotto_hotspot_files/figure-html/cell-2-output-1.png" width="801" height="393" class="figure-img"></p>
<figcaption>An illustrative example: if Store A sells 20x as many tickets as Store B, and every ticket has the exact same win probability, then over many draws Store A should produce roughly 20x as many winners - purely from volume, with no ‘lucky spot’ effect needed.</figcaption>
</figure>
</div>
</div>
</div>
<p>가게 하나가 유독 많이 팔리는 데는 실제 이유가 있다 — 대형 상권, 눈에 띄는 위치, 그리고 “명당”이라는 입소문 자체가 손님을 더 끌어모은다. 여기서 <strong>선택편향(selection bias)</strong>이 개입한다. 사람들은 당첨자가 나온 가게로 몰리고, 그 가게는 판매량이 더 늘어 당첨자가 또 나올 확률도 함께 커진다. 이 자기강화적 순환이 “명당”이라는 이야기를 계속 만들어내지만, 그 가게에서 산 개별 번호 조합의 당첨 확률 자체는 다른 가게에서 산 것과 조금도 다르지 않다.</p>
</section>
<section id="매주-1억-건-당첨자가-매주-나오는-진짜-이유" class="level2">
<h2 class="anchored" data-anchor-id="매주-1억-건-당첨자가-매주-나오는-진짜-이유">매주 1억 건 — 당첨자가 매주 나오는 진짜 이유</h2>
<div id="1472e570" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_lotto_hotspot_files/figure-html/cell-3-output-1.png" width="613" height="432" class="figure-img"></p>
<figcaption>The gap between the number of possible lotto combinations (8,145,060) and the number of tickets actually purchased in a typical week (roughly 119 million, based on 2024’s 6.2001 trillion-won annual sales), on a log scale. With so many tickets chasing so few combinations, having multiple winners in the same draw becomes routine rather than exceptional.</figcaption>
</figure>
</div>
</div>
</div>
<p>2024년 로또 연간 판매액은 6조 2,001억 원으로 역대 최대였다. 1,000원짜리 게임 기준으로 환산하면 매주 약 1억 1,900만 건이 팔리는 셈이다. 가능한 번호 조합은 814만 5,060개뿐인데, 매주 그보다 약 15배 많은 표가 팔린다. 이러니 1등 당첨 조합 하나에도 평균 여러 명이 동시에 걸리는 게 당연하다. 실제로 2024년 한 해에만 1등 당첨자가 812명 나왔다 — 매주 평균 15명 넘게, 단 한 명도 당첨자가 없었던 주는 드물었다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
“명당” 이야기를 마주하면 확인할 것
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>“이 조합·장소가 더 잘 당첨된다”는 주장을 보면, 확률 자체가 다른 건지 표본(판매량)이 다른 건지부터 구분하라.</strong> 동일확률 게임에서는 후자가 거의 항상 답이다.</li>
<li><strong>당첨자 수가 많다는 사실만으로는 “왜”를 설명하지 못한다.</strong> 판매량 데이터를 함께 보지 않으면 원인을 오해하기 쉽다.</li>
<li><strong>입소문이 판매량을 늘리고, 늘어난 판매량이 다시 입소문을 만드는 순환을 의심하라.</strong> “명당”은 종종 원인이 아니라 결과다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-동일확률이라는-전제-위에서만-성립하는-이야기" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-동일확률이라는-전제-위에서만-성립하는-이야기">더 깊이: 동일확률이라는 전제 위에서만 성립하는 이야기</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#laplace-확률-고전적-확률">〈수리통계학 1. 확률론 — Laplace 확률(고전적 확률)〉</a>은 표본공간의 모든 원소가 일어날 가능성이 같다는 <strong>균등가능성(equally likely)</strong> 가정 위에서 <img src="https://latex.codecogs.com/png.latex?P(A)%20=%20%7CA%7C/%7CS%7C">로 확률을 정의한다. 로또는 이 정의가 설계 그대로 적용되는 몇 안 되는 실생활 게임이다 — 814만 5,060개의 조합 각각이 정확히 같은 확률을 갖도록 공식적으로 설계돼 있다. 이 전제가 참이라는 걸 받아들이면, “특정 가게·번호가 더 잘 나온다”는 주장은 성립할 수 없다 — 남는 설명은 표본 크기(판매량) 차이뿐이다.</p>
</section>
<section id="sdv-관점-잘-팔린다와-잘-맞는다를-구분하는-결정" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-잘-팔린다와-잘-맞는다를-구분하는-결정">SDV 관점: “잘 팔린다”와 “잘 맞는다”를 구분하는 결정</h2>
<p>동일확률 가정을 확인하고 당첨자 수 차이를 판매량으로 설명하는 일은 <strong>요약과 추론(1~2단계)</strong>의 영역이다. 제가 제안하는 <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이 사례가 중요한 이유는, 비슷한 착시가 비즈니스 데이터에서도 똑같이 반복되기 때문이다 — “이 매장이 전환율이 높다”는 보고가, 사실은 매장 운영이 뛰어나서가 아니라 단순히 트래픽(표본)이 커서 나온 결과일 수 있다. 데이터를 가치 있는 결정으로 연결하려면, “성과가 좋다”는 지표를 표본 크기로 나눠 실제 비율(전환율, 적중률)을 확인한 뒤에야 그 매장의 운영 방식을 다른 곳에 복제할지 결정해야 한다. 표본 크기를 무시한 채 “명당”만 좇으면, 정작 복제해야 할 것은 위치가 아니라 마케팅 방식이었을 수 있다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-명당을-찾지-말고-판매량을-찾아라" class="level2">
<h2 class="anchored" data-anchor-id="결론-명당을-찾지-말고-판매량을-찾아라">결론: 명당을 찾지 말고, 판매량을 찾아라</h2>
<p>나는 학생들에게 “명당에서 사면 잘 된다더라”는 말을 들으면 이렇게 되물으라고 말한다. “그 가게가 파는 표 수는 얼마나 되나요?” 로또처럼 동일확률이 수학적으로 보장된 게임에서, 당첨자가 몰리는 곳은 운이 좋은 곳이 아니라 표를 많이 파는 곳이다. 명당이라는 이야기에 홀리기 전에, 그 이야기 뒤에 숨은 표본 크기부터 물어보는 습관이 필요하다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li><a href="https://www.sisajournal.com/news/articleView.html?idxno=305480">로또 1등 ’무더기 당첨’으로 확산된 조작설…사실은 이렇다? — 시사저널</a></li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_lotto_hotspot.html</guid>
  <pubDate>Wed, 23 Sep 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률과 우연의 세계) 무작위 숫자는 왜 고르게 흩어지지 않는가? — 군집 착각</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_clustering_illusion.html</link>
  <description><![CDATA[ 




<section id="완벽하게-무작위였는데-무작위가-아니라는-항의가-빗발쳤다" class="level2">
<h2 class="anchored" data-anchor-id="완벽하게-무작위였는데-무작위가-아니라는-항의가-빗발쳤다">완벽하게 무작위였는데, 무작위가 아니라는 항의가 빗발쳤다</h2>
<p>2014년 스포티파이는 재생목록 셔플 기능에 수학적으로 완벽한 무작위 알고리즘(피셔-예이츠 셔플)을 쓰고 있었다. 그런데도 이용자 항의가 끊이지 않았다. “같은 가수 노래가 왜 자꾸 연달아 나오죠? 이거 무작위 아니잖아요.” 스포티파이 엔지니어링 블로그는 결국 이렇게 인정했다 — 진짜 무작위는 원래 이렇게 뭉치기 마련이라고. 그래서 스포티파이는 같은 아티스트의 곡이 너무 가까이 배치되지 않도록, <strong>알고리즘을 일부러 덜 무작위하게</strong> 바꿨다. 2005년 애플도 같은 문제를 겪었다. 스티브 잡스는 당시 발표 자리에서 이렇게 말했다. “우리 셔플이 무작위가 아니라는 사람들이 많더군요. 사실은 완벽하게 무작위입니다. 다만 무작위라는 게 가끔은 같은 아티스트의 두 곡이 나란히 나온다는 뜻이기도 하죠 … 그래서 스마트 셔플을 추가해 일부러 덜 무작위하게 만들었습니다 — 사람들은 그게 더 무작위하다고 느끼겠지만, 사실은 그 반대입니다.”</p>
<div id="d06c1132" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_clustering_illusion_files/figure-html/cell-2-output-1.png" width="854" height="159" class="figure-img"></p>
<figcaption>A 30-song playlist from 5 artists (6 songs each, one color per artist), arranged two ways. Top: a genuinely random shuffle - notice the clumps of same-color squares. Bottom: songs deliberately spaced round-robin style so no artist repeats back-to-back. Both are valid orderings, but only the bottom one ‘feels’ random to most listeners.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="곡을-진짜로-섞으면-99.6-확률로-어디선가-같은-아티스트가-연달아-나온다" class="level2">
<h2 class="anchored" data-anchor-id="곡을-진짜로-섞으면-99.6-확률로-어디선가-같은-아티스트가-연달아-나온다">30곡을 진짜로 섞으면, 99.6% 확률로 어디선가 같은 아티스트가 연달아 나온다</h2>
<p>이게 얼마나 흔한 일인지 직접 계산해보자. 5명의 아티스트가 각 6곡씩, 총 30곡짜리 재생목록을 진짜 무작위로 섞는다고 하자.</p>
<div id="64338b76" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_clustering_illusion_files/figure-html/cell-3-output-1.png" width="710" height="431" class="figure-img"></p>
<figcaption>10,000 simulations of a truly random 30-song shuffle (5 artists, 6 songs each): how many back-to-back same-artist pairs show up. Getting zero such pairs happens only 0.4% of the time - almost every real shuffle has several clumps, averaging about 5 adjacent same-artist pairs per playlist.</figcaption>
</figure>
</div>
</div>
</div>
<p>시뮬레이션 결과, 같은 아티스트가 한 번도 연달아 나오지 않는 “완벽하게 고른” 셔플은 <strong>0.4%</strong>밖에 안 된다. 평균적으로 한 재생목록에 이런 “뭉침”이 5번쯤 생긴다. 즉 이용자들이 항의한 “왜 자꾸 같은 가수가 연달아 나오죠?”라는 현상은 버그가 아니라, <strong>진짜 무작위가 원래 하는 일</strong>이었다.</p>
</section>
<section id="사람은-고르게-흩어진-것을-무작위라고-착각한다" class="level2">
<h2 class="anchored" data-anchor-id="사람은-고르게-흩어진-것을-무작위라고-착각한다">사람은 “고르게 흩어진 것”을 무작위라고 착각한다</h2>
<p>우리 뇌는 무작위를 “예측할 수 없을 만큼 고르게 퍼진 상태”로 상상한다. 그런데 수학적으로 진짜 무작위(각 사건이 서로 독립)는 정반대다 — 뭉치거나 비는 구간이 생기는 게 오히려 자연스럽다. 오히려 뭉침이나 빈 구간이 전혀 없이 고르게 퍼진 배열은, 누군가 의도적으로 배치를 조정했다는 증거에 가깝다. 이 착각을 <strong>군집 착각(clustering illusion)</strong>이라 부른다. 제2차 세계대전 당시 런던 시민들이 독일군 V-1 로켓의 낙하 지점이 “특정 구역에 집중된 패턴”을 보인다며 그 지역에 스파이가 있다고 의심했던 사례도 같은 착각이었다 — 훗날 통계 분석 결과, 낙하 지점은 순수한 무작위(포아송) 분포와 통계적으로 구별되지 않았다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
군집 착각을 피하는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>“고르게 흩어져야 무작위답다”는 직관을 의심하라.</strong> 진짜 무작위는 뭉침과 빈 구간을 만드는 게 정상이다.</li>
<li><strong>뭉쳐 보이는 패턴을 발견하면, 그게 순수한 우연으로 나올 확률부터 계산해보라.</strong> 오늘처럼 계산해보면 “이상한 뭉침”이 사실은 “가장 흔한 결과”인 경우가 많다.</li>
<li><strong>“덜 무작위하게” 설계된 시스템(플레이리스트 셔플, 좌석 배치 등)은 사용자 경험을 위한 것이지, 통계적 무작위성의 정의와는 다르다는 걸 구분하라.</strong></li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-포아송-과정은-원래-뭉친다" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-포아송-과정은-원래-뭉친다">더 깊이: 포아송 과정은 원래 뭉친다</h2>
<p>강의노트 <a href="../../notes/math_stat/famous_distribution.html#포아송분포-x-sim-poissonlambda">〈수리통계학 5. 유명분포 — 포아송분포〉</a>는 사건들이 서로 독립적으로, 일정한 평균 발생률로 일어나는 과정을 포아송 과정으로 정의한다. 포아송 과정의 핵심은 사건 사이의 간격이 무작위(지수분포)라는 데 있다 — 간격이 일정하게 고정된 게 아니라 들쭉날쭉하다. 그 결과 짧은 구간에 사건이 몰리는 뭉침과, 사건이 뜸한 빈 구간이 함께 나타난다. 오늘 살펴본 플레이리스트 셔플이나 런던 폭격 지점이나, 수학적으로는 모두 “독립적인 사건들이 무작위 간격으로 발생한다”는 같은 포아송적 구조를 공유한다.</p>
</section>
<section id="sdv-관점-패턴이-보인다는-보고를-검증하는-절차" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-패턴이-보인다는-보고를-검증하는-절차">SDV 관점: “패턴이 보인다”는 보고를 검증하는 절차</h2>
<p>무작위 배열에서 뭉침이 나올 확률을 계산하는 건 <strong>요약과 추론(1~2단계)</strong>에 해당하는 확률 계산이다. 제가 제안하는 <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이 문제가 중요한 이유는, 현업에서 “패턴이 보인다”는 보고가 끊임없이 올라오기 때문이다 — 특정 시간대에 불량품이 몰렸다, 특정 영업일에 고객 이탈이 몰렸다, 특정 구역에 사고가 몰렸다는 식이다. 오늘 확인했듯, 완전히 무작위한 과정에서도 이런 “몰림”은 기본값에 가깝다. 데이터를 가치 있는 결정으로 연결하려면, 뭉침을 발견한 즉시 원인을 찾아 나서기 전에 “이 정도 뭉침이 순수한 우연으로 나올 확률이 얼마인가”부터 계산해야 한다. 그 확률이 낮을 때만 자원을 들여 원인을 추적하는 것이, AI 시대에 통계학자가 설계해야 할 가치 있는 의사결정 절차다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-고르게-섞였다면-오히려-그게-이상한-것이다" class="level2">
<h2 class="anchored" data-anchor-id="결론-고르게-섞였다면-오히려-그게-이상한-것이다">결론: 고르게 섞였다면, 오히려 그게 이상한 것이다</h2>
<p>나는 학생들에게 셔플 버튼을 누르고 같은 가수 노래가 연달아 나오면, 그걸 “앱이 이상하다”는 증거가 아니라 “무작위가 제대로 작동하고 있다”는 증거로 받아들이라고 말한다. 스포티파이와 애플이 결국 알고리즘을 “덜 무작위하게” 바꾼 이유는 수학이 틀려서가 아니라, 사람의 직관에 맞추기 위해서였다. 진짜 우연 앞에서는, 고르게 흩어지지 않는 쪽이 오히려 정상이다.</p>
<div class="callout callout-style-default callout-note callout-titled">
<div class="callout-header d-flex align-content-center" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
참고 자료
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<ul>
<li>Spotify Engineering, <em>How to Shuffle Songs?</em> (2014) — <a href="https://rnd.atspotify.com/how-to-shuffle-songs/">rnd.atspotify.com/how-to-shuffle-songs</a></li>
<li>Spotify Engineering, <em>Shuffle: Making Random Feel More Human</em> (2025) — <a href="https://engineering.atspotify.com/2025/11/shuffle-making-random-feel-more-human">engineering.atspotify.com</a></li>
</ul>
</div>
</div>
</div>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_clustering_illusion.html</guid>
  <pubDate>Tue, 22 Sep 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률과 우연의 세계) 같은 번호가 연속으로 나오면 조작일까? — 독립사건</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_independent_events.html</link>
  <description><![CDATA[ 




<section id="지난주-번호가-또-나왔다는-뉴스는-사실-뉴스거리가-아니다" class="level2">
<h2 class="anchored" data-anchor-id="지난주-번호가-또-나왔다는-뉴스는-사실-뉴스거리가-아니다">“지난주 번호가 또 나왔다”는 뉴스는 사실 뉴스거리가 아니다</h2>
<p>로또 추첨 결과가 발표될 때마다 인터넷에는 비슷한 댓글이 달린다. “어? 이 번호 지난주에도 나왔던 건데?”, “이거 설계된 거 아니야?” 그런데 실제로 계산해보면, 지난 회차에 나온 6개 번호 중 적어도 하나가 이번 회차에도 나올 확률은 <strong>약 60%</strong>다. 2주 중 1주 이상은 번호가 겹치는 게 오히려 통계적으로 자연스러운 결과라는 뜻이다.</p>
<div id="39ccaa98" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_independent_events_files/figure-html/cell-2-output-1.png" width="566" height="413" class="figure-img"></p>
<figcaption>The probability that this week’s 6 lotto numbers (drawn from 1-45) share at least one number with last week’s 6, versus sharing none, computed exactly using combinations. Sharing at least one number is actually the more likely outcome.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="매-회차는-독립이다-그런데-독립이-달라야-한다는-뜻은-아니다" class="level2">
<h2 class="anchored" data-anchor-id="매-회차는-독립이다-그런데-독립이-달라야-한다는-뜻은-아니다">매 회차는 독립이다 — 그런데 독립이 “달라야 한다”는 뜻은 아니다</h2>
<p>로또 추첨기는 매주 45개의 공을 처음부터 다시 채워 넣고 돌린다. 지난주에 나온 공이 이번 주에 제외되는 일은 없다. 이렇게 한 사건(이번 주 추첨)의 결과가 다른 사건(지난주 추첨)의 결과에 전혀 영향을 받지 않을 때, 두 사건을 <strong>독립(independent)</strong>이라 한다. 그런데 여기서 흔한 오해가 생긴다 — 많은 사람이 “독립적이다”를 “서로 달라야 한다”로 잘못 받아들인다. 사실은 정반대다. 독립적이라는 건 지난주 결과가 이번 주 결과에 <strong>아무 제약도 주지 않는다</strong>는 뜻이지, 겹치지 않도록 조정된다는 뜻이 아니다. 오히려 아무 제약이 없기 때문에, 순전한 우연만으로도 번호가 겹칠 여지가 그대로 남아 있다.</p>
<div id="19a2ebc8" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_independent_events_files/figure-html/cell-3-output-1.png" width="662" height="431" class="figure-img"></p>
<figcaption>A simulation of 200,000 pairs of consecutive lotto draws: how many numbers (out of 6) are shared between one week and the next. Sharing exactly one number is actually the single most common outcome (about 42%), more common than sharing none at all (40%).</figcaption>
</figure>
</div>
</div>
</div>
<p>시뮬레이션 결과, 번호가 하나도 안 겹치는 경우는 40%, 정확히 1개 겹치는 경우가 42%로 오히려 더 흔하다. 2개 이상 겹치는 경우도 17.6%나 된다. 이런 겹침을 “조작의 증거”로 의심하려면, 1년(52회차) 중 절반 가까운 주에서 “조작 의혹”을 제기해야 한다는 뜻이다 — 그 자체로 이미 앞뒤가 맞지 않는다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
독립사건과 “우연의 반복”을 구분하는 법
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>“독립”은 “달라야 한다”가 아니라 “서로 영향을 안 준다”는 뜻이다.</strong> 독립적인 사건도 우연히 같은 결과를 낼 수 있다.</li>
<li><strong>반복되는 패턴을 보면, 그 패턴이 나올 수 있는 경우의 수를 먼저 세어보라.</strong> 오늘처럼 계산해보면 “예외적인 우연”이 사실은 “가장 흔한 결과”인 경우가 많다.</li>
<li><strong>매 시행이 정말 독립인지부터 확인하라.</strong> 로또처럼 공을 매번 새로 채우는 게임은 독립이지만, 카드 게임처럼 이미 뽑힌 패를 빼고 진행하는 경우는 독립이 아니다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-독립의-정의와-곱셈-규칙" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-독립의-정의와-곱셈-규칙">더 깊이: 독립의 정의와 곱셈 규칙</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#독립">〈수리통계학 1. 확률론 — 독립〉</a>은 두 사건 <img src="https://latex.codecogs.com/png.latex?A">와 <img src="https://latex.codecogs.com/png.latex?B">가 독립이면 <img src="https://latex.codecogs.com/png.latex?P(A%20%5Ccap%20B)%20=%20P(A)%20%5Ccdot%20P(B)">가 성립한다고 정의하고, 상호 배타적(mutually exclusive)인 사건과 독립사건을 혼동하면 안 된다는 예제도 함께 다룬다 — 두 사건이 동시에 일어날 수 없는 것(상호 배타)과, 두 사건이 서로 영향을 주지 않는 것(독립)은 전혀 다른 개념이다. 로또의 이번 주 추첨과 지난주 추첨은 상호 배타적이지 않다(둘 다 일어날 수 있다는 것 자체가 당연하다) — 그리고 독립이다(서로 영향을 안 준다). 이 두 성질이 합쳐지면, 오늘 계산한 것처럼 “번호가 겹칠 확률”을 곱셈 규칙으로 정확히 계산할 수 있다.</p>
</section>
<section id="sdv-관점-패턴처럼-보인다를-확률로-검증하는-습관" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-패턴처럼-보인다를-확률로-검증하는-습관">SDV 관점: “패턴처럼 보인다”를 확률로 검증하는 습관</h2>
<p>두 사건이 독립인지 확인하고 겹칠 확률을 계산하는 일은 <strong>요약과 추론(1~2단계)</strong>에 해당하는 고전적 확률 계산이다. 제가 제안하는 <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이 계산이 갖는 가치는 실전에 있다. 품질관리에서 “같은 불량이 이번 달에도 반복됐다”, 마케팅에서 “이 고객군이 지난달과 또 겹친다”처럼, 업무 현장에서는 “반복”을 곧장 “이상 신호”로 해석하는 경우가 흔하다. 그런데 그 반복이 정말 이례적인지, 아니면 오늘 로또 사례처럼 독립적인 과정에서도 자연스럽게 나오는 흔한 결과인지를 먼저 계산해보지 않으면, 있지도 않은 원인을 찾아 헤매느라 자원을 낭비하게 된다. 데이터를 가치로 바꾸는 첫 단계는 종종, 그 데이터가 “설명이 필요한 이상 신호”인지 “설명이 필요 없는 정상 변동”인지를 가려내는 것이다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-겹치는-게-이상한-게-아니라-안-겹치는-걸-기대하는-게-이상하다" class="level2">
<h2 class="anchored" data-anchor-id="결론-겹치는-게-이상한-게-아니라-안-겹치는-걸-기대하는-게-이상하다">결론: 겹치는 게 이상한 게 아니라, 안 겹치는 걸 기대하는 게 이상하다</h2>
<p>나는 학생들에게 “이런 우연이?”라는 반응이 나오면, 그 우연이 나올 수 있는 경우의 수를 직접 세어보라고 말한다. 지난주와 이번 주 로또 번호가 하나라도 겹칠 조합은 생각보다 훨씬 많다. 매주 반복되는 추첨이 서로 독립이라는 사실은, 번호가 겹치지 않게 막아주는 장치가 아니라 오히려 겹칠 여지를 그대로 열어두는 조건이다. 독립사건 앞에서는, 반복이야말로 기본값이다.</p>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_independent_events.html</guid>
  <pubDate>Mon, 21 Sep 2026 15:00:00 GMT</pubDate>
</item>
<item>
  <title>(확률과 우연의 세계) 23명만 모여도 생일이 겹칠 가능성이 높은 이유 — 생일 문제</title>
  <dc:creator>권세혁 </dc:creator>
  <link>https://by-sekwon.github.io/stat_assay/posts/story_birthday_problem.html</link>
  <description><![CDATA[ 




<section id="명-반이-다-모이기도-전에-생일이-겹친다" class="level2">
<h2 class="anchored" data-anchor-id="명-반이-다-모이기도-전에-생일이-겹친다">23명, 반이 다 모이기도 전에 생일이 겹친다</h2>
<p>새 학기, 새로운 반 학생들과 처음 만나는 자리에서 나는 종종 이 질문을 던진다. “이 교실에 몇 명이 있어야, 그중 두 사람의 생일이 같을 확률이 50%를 넘을까요?” 대부분의 학생들은 365일의 절반쯤인 180명 안팎을 답한다. 정답은 <strong>23명</strong>이다. 학생 23명만 모여도, 그중 생일이 겹치는 두 사람이 있을 확률이 이미 50%를 넘는다. 반 정원이 28~30명이라면 그 확률은 65~70%까지 올라간다 — 대부분의 교실에서 생일이 겹치는 학생이 있다는 뜻이다. 직관과 정면으로 충돌하는 이 결과를 <strong>생일 문제(birthday problem)</strong>라 부른다.</p>
<div id="ab99ceae" class="cell" data-execution_count="1">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_birthday_problem_files/figure-html/cell-2-output-1.png" width="710" height="451" class="figure-img"></p>
<figcaption>The probability that at least two people in a room share the same birthday, as the room’s headcount grows from 1 to 100 (assuming birthdays are spread evenly across 365 days). The curve crosses 50% at just 23 people and exceeds 99% by 57 people.</figcaption>
</figure>
</div>
</div>
</div>
</section>
<section id="질문을-잘못-알아들으면-답이-틀린다" class="level2">
<h2 class="anchored" data-anchor-id="질문을-잘못-알아들으면-답이-틀린다">질문을 잘못 알아들으면 답이 틀린다</h2>
<p>이 결과가 이상하게 느껴지는 이유는, 많은 사람이 이 질문을 <strong>“나와 생일이 같은 사람이 있을 확률”</strong>로 잘못 알아듣기 때문이다. 나 한 사람 기준으로 22명 중 누군가와 생일이 같을 확률은 <img src="https://latex.codecogs.com/png.latex?22/365%20%5Capprox%206%5C%25">로 낮은 게 맞다. 그런데 실제 질문은 <strong>“23명 중 아무 두 사람이든 생일이 같을 확률”</strong>이다. 23명이면 두 사람씩 짝지을 수 있는 경우의 수가 <img src="https://latex.codecogs.com/png.latex?%5Cbinom%7B23%7D%7B2%7D%20=%20253">쌍이나 된다. 253번의 “혹시 겹칠까” 기회가 있는 셈이니, 그중 하나라도 겹칠 확률은 훨씬 커진다.</p>
<div id="e604e22c" class="cell" data-execution_count="2">
<div class="cell-output cell-output-display">
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://by-sekwon.github.io/stat_assay/posts/story_birthday_problem_files/figure-html/cell-3-output-1.png" width="710" height="431" class="figure-img"></p>
<figcaption>The same probability, shown for a few real-world Korean group sizes: a typical classroom (28-30 students), a KBO first-team active roster (28 players), and a small company team meeting (10 people).</figcaption>
</figure>
</div>
</div>
</div>
<p>계산 방법은 정면 돌파보다 뒤에서 접근하는 쪽이 훨씬 쉽다. “적어도 한 쌍이 겹칠 확률”을 직접 구하는 대신, “아무도 안 겹칠 확률”을 구해서 1에서 빼는 것이다. 첫 사람의 생일은 365일 중 아무 날이나 상관없다. 두 번째 사람이 첫 사람과 겹치지 않을 확률은 <img src="https://latex.codecogs.com/png.latex?364/365">, 세 번째 사람이 앞의 둘과 겹치지 않을 확률은 <img src="https://latex.codecogs.com/png.latex?363/365">, 이런 식으로 23번째 사람까지 곱해나간다. 이 곱은 사람이 늘어날수록 급격히 작아지고, 그만큼 “적어도 한 쌍은 겹친다”는 확률(1에서 그 값을 뺀 것)은 빠르게 커진다.</p>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
생일 문제를 마주하면 기억할 것
</div>
</div>
<div class="callout-body-container callout-body">
<ol type="1">
<li><strong>“나와 겹칠 확률”과 “그룹 안 아무 두 사람이 겹칠 확률”은 완전히 다른 질문이다.</strong> 후자는 비교 대상 쌍의 수가 인원수의 제곱에 가깝게 늘어나 훨씬 빨리 커진다.</li>
<li><strong>“적어도 하나”를 구할 땐 반대(여사건)를 먼저 구하라.</strong> “아무도 안 겹칠 확률”을 구해 1에서 빼는 쪽이 훨씬 계산하기 쉽다.</li>
<li><strong>작은 표본에서 “이런 우연이?”라고 놀라기 전에, 비교 가능한 쌍이 몇 개나 있는지부터 세어보라.</strong> 우연처럼 보이는 많은 사건이, 사실은 비교 기회가 많아서 생긴 자연스러운 결과다.</li>
</ol>
</div>
</div>
</section>
<section id="더-깊이-여사건으로-적어도-하나를-계산하는-법" class="level2">
<h2 class="anchored" data-anchor-id="더-깊이-여사건으로-적어도-하나를-계산하는-법">더 깊이: 여사건으로 “적어도 하나”를 계산하는 법</h2>
<p>강의노트 <a href="../../notes/math_stat/probability.html#독립">〈수리통계학 1. 확률론 — 독립〉</a>은 확률 계산의 정리로 <img src="https://latex.codecogs.com/png.latex?P(A%5Ec)%20=%201%20-%20P(A)">를 제시하고, 17세기 도박사 슈발리에 드 메레(Chevalier de Méré)의 사례를 든다 — 주사위를 4번 던져 6이 한 번이라도 나올 확률은 <img src="https://latex.codecogs.com/png.latex?1-(5/6)%5E4">로, “한 번도 안 나올 확률”을 구해 1에서 빼서 계산한다. 오늘 계산한 생일 문제도 정확히 같은 틀이다. <img src="https://latex.codecogs.com/png.latex?n">명의 생일이 서로 독립이라고 가정한 뒤(한 사람의 생일이 다른 사람의 생일에 영향을 주지 않으므로), “아무도 안 겹칠 확률”을 각 사람이 앞사람들과 다른 날짜를 가질 확률의 곱으로 구하고, 1에서 뺀다. 독립사건의 곱셈 규칙과 여사건의 뺄셈 규칙, 이 두 가지 도구만으로 반직관적인 결과가 깔끔하게 풀린다.</p>
</section>
<section id="sdv-관점-말도-안-되는-우연을-계산-가능한-확률로-바꾸기" class="level2">
<h2 class="anchored" data-anchor-id="sdv-관점-말도-안-되는-우연을-계산-가능한-확률로-바꾸기">SDV 관점: “말도 안 되는 우연”을 계산 가능한 확률로 바꾸기</h2>
<p>생일 문제의 확률을 계산하는 일 자체는 <strong>요약과 추론(1~2단계)</strong>을 오가는 고전적인 확률 계산이다. 그런데 제가 제안하는 <strong>SDV(통계적 데이터 가치화)</strong> 관점에서 이 문제가 흥미로운 이유는 따로 있다. 사람들은 “말도 안 되는 우연이 일어났다”는 이야기를 사실로 믿고 의사결정에 반영하는 경우가 많다 — 이상 거래 탐지, 품질 관리에서의 “우연히 같은 결함이 반복됐다”는 보고, 의료 데이터에서 “같은 부작용이 겹쳤다”는 경보가 그 예다. 그런 사건이 정말 이례적인지, 아니면 애초에 비교 대상 쌍이 많아서 나올 법한 결과인지를 생일 문제의 틀로 먼저 계산해보는 것 — 이것이 “놀랍다”는 직관을 “그럴 확률이 얼마다”라는 숫자로 바꾸는, 데이터를 가치 있는 판단으로 이어주는 첫걸음이다. (<a href="../../sdv/index.html">SDV(데이터 가치화)</a>)</p>
</section>
<section id="결론-새-챕터-우연이라는-말을-계산해보는-시간" class="level2">
<h2 class="anchored" data-anchor-id="결론-새-챕터-우연이라는-말을-계산해보는-시간">결론: 새 챕터, “우연”이라는 말을 계산해보는 시간</h2>
<p>이번 챕터에서는 “확률과 우연의 세계”를 다룬다. 생일 문제는 그 출발점으로 손색이 없다 — 우리의 직관이 확률 앞에서 얼마나 쉽게 틀리는지, 그리고 그 직관을 바로잡는 데 복잡한 도구가 아니라 여사건과 독립이라는 단 두 가지 개념이면 충분하다는 것을 동시에 보여주기 때문이다. 다음에 “이런 우연이 있을 수 있나”라는 생각이 들면, 먼저 “비교할 수 있는 쌍이 몇 개나 있었지?”를 물어보자.</p>


</section>

 ]]></description>
  <category>통계기초</category>
  <category>확률</category>
  <guid>https://by-sekwon.github.io/stat_assay/posts/story_birthday_problem.html</guid>
  <pubDate>Sun, 20 Sep 2026 15:00:00 GMT</pubDate>
</item>
</channel>
</rss>
