#유전자 #고대DNA #미토콘드리아 #유라시아
Bennett, C.C. and Kaestle, F.A. (2006) ‘Reanalysis of Eurasian population history: ancient DNA evidence of population affinities’, Human Biology, 78(4), pp. 413–440.
Reanalysis of Eurasian Population History: Ancient DNA Evidence of Population Affinities
유라시아 인구사 재분석: 고대 DNA 증거로 살펴본 집단 간의 유전적 관계
케이시 C. 베넷(CASEY C. BENNETT)1과 프레데리카 A. 캐슬(FREDERIKA A. KAESTLE)1,2
[리뷰] 중화 쇼비니즘 경사도 평가: 5/10
21. Bennett, C.C. and Kaestle, F.A. (2006) ‘Reanalysis of Eurasian population history: ancient DNA evidence of population affinities’, Human Biology, 78(4), pp. 413–440.
(1) 연구 개요 및 저자의 주장
이 연구는 Wang et al. (2000)의 린쯔 mtDNA 데이터를 재분석했다. 저자들은 원 논문의 ‘유럽계와 유사(European-like)’라는 표현이 “오해의 소지가 있다(misleading)”고 비판했다. 대신, 이들은 린쯔 인구가 기원전 1천년 경 중앙 유라시아 스텝 지역에 널리 퍼져 있던 “초기 이란인(early Iranians)과 잠재적으로 관련”이 있을 수 있다는 더 구체적인 가설을 제시하며, 서방과의 연관성 자체를 부정하기보다는 정교화하고자 했다.
(2) 편향성 분석 (중화 쇼비니즘 경사도: N/A – 방법론적 한계가 서사를 주도)
이 연구 역시 특정 이념을 옹호하기보다는, 기존의 거대 서사 틀 안에서 논의를 진행했다는 점에서 한계를 보인다.
- 서사 프레이밍 (높은 편향성): ‘유럽계’를 ‘이란계’로 수정함으로써 주장을 더 구체화했지만, 이는 여전히 산동의 변화를 설명하기 위해 외부의 거대한 역사적 행위자(이란계 스텝 민족)를 호출하는 동일한 프레임워크 내에 머무는 것이다.
- 모델 선택과 반례 취급 (높은 편향성): Wang et al. (2000)과 동일한 제한된 데이터를 사용했기 때문에, 데이터의 모호함이라는 근본적인 문제를 해결하지 못했다. 지역 내부의 복잡성이나 다른 방향의 교류 가능성과 같은 대안적 모델을 탐색하기보다는, 기존의 동-서 교류라는 거대 서사를 어떻게 더 잘 설명할지에만 집중했다.
(3) 결론 재구성: 비판적 재분석의 한계
이 연구는 과학적 논쟁이 어떻게 기존의 거대 서사를 해체하기보다는 오히려 더 정교하게 강화할 수 있는지를 보여주는 사례다. 모호한 ‘유럽계’라는 딱지를 더 구체적인 ‘이란계’로 대체함으로써, 이 연구는 역설적으로 외부(서방)로부터의 대규모 유전자 유입이라는 기본 서사에 더 큰 역사적 그럴듯함을 부여했다. ‘매트릭스’ 모델의 관점에서, 이 재분석은 데이터의 근본적인 불확실성을 인정하고 지역 내 모델의 필요성을 강조하는 대신, 또 다른 외부 기원을 찾는 데 집중함으로써 더 복합적인 진실을 탐구할 기회를 놓쳤다.
[논문요약]
핵심 용어 쉽게 풀이하기
이 논문을 이해하려면 몇 가지 핵심 유전학 용어를 알아야 한다. 자동차에 비유해 설명한다.
- 미토콘드리아 DNA (mtDNA)
어려운 설명: 세포의 에너지 공장인 미토콘드리아 안에 있는 DNA로, 모계를 통해서만 유전된다.
쉬운 비유: **’어머니에게서만 물려받는 자동차 모델명’**과 같다. 아버지가 어떤 차를 타든 상관없이, 자녀는 무조건 어머니의 차 모델명(mtDNA)을 물려받는다. 그래서 mtDNA를 쭉 따라가면 수천 년 전 외할머니, 외증조할머니의 ‘차 모델명’을 추적할 수 있어, 고대 인류의 모계 혈통을 연구하는 데 아주 유용하다.
- 과변이 부위 (Hypervariable Region, HVSI)
어려운 설명: mtDNA 내에서 다른 부분보다 돌연변이가 훨씬 빠르게 일어나는 특정 구간.
쉬운 비유: 자동차 모델명 중에서도 **’연식이나 옵션 표시 부분’**과 같다. ‘소나타’라는 기본 모델명은 잘 안 바뀌지만, ‘2025년형, 하이브리드, 프리미엄’ 같은 세부 정보는 자주 바뀐다. 이처럼 HVSI는 변화가 잦아서, 비교적 가까운 시기에 갈라져 나온 인류 집단 간의 차이를 세밀하게 구분하는 데 사용된다.
- FST (고정지수)
어려운 설명: 집단 간의 유전적 분화 정도를 나타내는 통계값.
쉬운 비유: **’유전적 거리’ 또는 ‘친밀도 점수’**이다. 두 집단 간의 FST 점수가 0에 가까우면(낮으면) 유전적으로 매우 가깝다는 뜻이고(친형제처럼), 점수가 높으면 유전적으로 매우 멀다는(아주 먼 친척처럼) 의미이다. 이 논문에서는 이 점수를 이용해 고대인과 현대인의 유전적 거리를 쟀다.
- 하플로그룹 (Haplogroup)
어려운 설명: 공통 조상으로부터 유래한 유사한 유전 표지(하플로타입)의 집합.
쉬운 비유: 인류라는 거대한 족보의 ‘아주 큰 본관(本貫)’ 같은 개념이다. 아주 먼 옛날, 특정 유전적 특징을 처음 가졌던 한 사람으로부터 시작된 거대한 후손 그룹을 말한다. 예를 들어 ‘하플로그룹 B’는 ‘김해 김씨’처럼 하나의 큰 뿌리를 나타낸다.
- 유럽인종(Europoid) / 몽골인종(Mongoloid)
- 주의사항: 이 용어들은 2006년 논문 작성 당시 사용된 낡은 인종 구분법이다. 오늘날에는 인류의 복잡한 유전적 다양성을 제대로 설명하지 못하고 차별적 소지가 있어 잘 사용하지 않는다. 이 요약에서는 논문의 표현을 그대로 전달하기 위해 사용하지만, 현대 유전학에서는 이런 단순한 구분을 지양한다는 점을 기억해야 한다.
고대 유라시아인 DNA 재분석 논문 요약
1. 무엇이 궁금했을까? (서론)
이 논문의 핵심 질문은 간단하다. “약 2,500년 전, 현재의 중국 북부와 몽골 지역에는 어떤 사람들이 살았을까?”
연구팀은 두 곳의 고대 유적지에 주목했다.
- 에긴골 (Egyin Gol), 몽골: 약 2,000년 전, 역사상 강력한 유목제국이었던 흉노(匈奴) 시대의 공동묘지이다.
- 임치 (Linzi), 중국 산동성: 그보다 조금 더 오래된 약 2,500년 전 춘추시대의 유적지이다.
특히 임치 유적은 이전 연구에서 주민들의 DNA가 현대 유럽인과 비슷하다는 놀라운 결과가 나와 학계를 놀라게 했다. 이 연구는 과연 그 결과가 사실인지, 그리고 흉노 시대의 몽골 사람들은 누구와 가장 가까운 친척인지 다시 한번 깊이 파고들어 검증하고자 했다.
2. 어떻게 연구했을까? (연구 방법)
연구팀은 최신 유전학 기법을 활용해 ‘누가 누구의 친척인가’를 찾는 탐정 작업을 시작했다.
- 1단계 (데이터 수집): 먼저, 두 고대 유적지(임치, 에긴골)에서 발굴된 유골의 mtDNA 데이터를 확보했다.
- 2단계 (비교 대상 선정): 유라시아와 아프리카 전역의 현대인 51개 집단에서 3,700명이 넘는 mtDNA 데이터를 모아 거대한 비교 데이터베이스를 만들었다.


표 1. 분석 대상 인구 집단 데이터 (지역별 알파벳 순)
- 3단계 (유전적 거리 계산): FST 점수를 이용해 두 고대인 그룹이 51개 현대인 그룹과 각각 얼마나 유전적으로 가깝고 먼지를 계산했다.
- 4단계 (토너먼트식 비교): 한꺼번에 비교하면 컴퓨터가 감당하기 어렵기 때문에, 먼저 ‘지역 예선’을 치렀다. 유럽, 남아시아, 동아시아 등 지역별로 고대인과 가장 가까운 상위권 팀(현대인 그룹)을 뽑았다. 그 후, 각 지역 예선을 통과한 팀들만 모아 ‘최종 결선’을 치러 누가 최종적으로 가장 가까운 친척인지 가려냈다.
연구팀은 임치(Linzi) 샘플의 염기서열 길이가 185bp로 매우 짧다는 점이 분석 결과에 오류를 일으킬 수 있다는 비판을 고려하여, 염기서열 길이가 결과의 신뢰도에 미치는 영향을 별도로 검증했다. 아래 표 6과 7은 긴 염기서열 데이터와 짧은 염기서열 데이터로 각각 분석했을 때 유전적 거리 순위에 큰 변화가 없음을 보여줌으로써, 짧은 서열을 사용한 임치 분석 결과가 신뢰할 수 있음을 입증한다.

표 6. 에긴골(Egyin Gol) 샘플의 염기서열 길이에 따른 Fₛₜ 값 비교

표 7. 일본인(Japanese) 샘플의 염기서열 길이에 따른 Fₛₜ 값 비교
3. 무엇을 발견했을까? (결과)
분석 결과, 매우 흥미롭고 극적인 사실이 드러났다.
결론1: 임치 사람과 에긴골 사람은 완전 남남이었다.
두 고대 인구 집단은 유전적으로 매우 달랐다. 가까운 시기에 비슷한 지역에 살았지만, 서로 가까운 친척 관계가 아니었다.
결론2: 에긴골(몽골 흉노) 사람들은 예상대로 ‘북동아시아인’이었다.
- 에긴골 사람들의 mtDNA는 현대 북동아시아인들과 매우 가까웠다.


- 표 4와 5를 보면, 이들의 가장 가까운 현대 친척은 중국 북부(산동성, 요녕성) 사람들, 몽골인, 일본인, 한국인 등이었다.
- 이들은 유럽인이나 남아시아인과는 유전적 거리가 매우 멀었다.
- 이는 흉노 제국을 이루었던 사람들이 오늘날 몽골과 그 주변에 사는 사람들의 직접적인 조상이라는 것을 강력하게 뒷받침하는 결과이다.
결론3: 임치(중국 산동성) 사람들은 ‘이란계 서쪽 사람’이었다.
- 이 연구의 가장 놀라운 부분이다. 임치 사람들은 현대 중국인이나 다른 동아시아인이 아니라, 서쪽 사람들과 유전적으로 훨씬 가까웠다.


- 표 2, 3, 5를 종합해 보면, 임치 사람들과 가장 가까운 현대인은 터키인, 이란인, 헝가리인 등 근동 및 동유럽 사람들이었다.
- 이전 연구에서 ‘유럽인 같다’고 한 것이 완전히 틀린 말은 아니었지만, 더 정확히는 서유럽(아이슬란드 등)이 아니라 중앙 유라시아 초원을 건너온 이란계 민족일 가능성이 매우 높다는 것이다.
4. 이것이 무엇을 의미할까? (논의 및 결론)
이 연구 결과는 고대 동아시아의 역사를 다시 생각하게 만드는 중요한 의미를 가진다.
- 몽골 초원의 유전적 연속성: 흉노 시대의 에긴골 사람들은 현대 몽골인 및 주변 민족의 조상이 맞았다. 즉, 몽골 초원 지역은 약 2,000년 동안 큰 인구 교체 없이 유전적 명맥을 이어왔다고 볼 수 있다.
- 고대 중국 북부의 ‘서쪽 사람들’: 약 2,500년 전 중국 산동성 지역에는 중앙 유라시아 초원 지대에서 동쪽으로 이동해 온 이란계 민족과 관련된 사람들이 살고 있었다. 이들은 실크로드가 열리기 훨씬 이전부터 동서 교류가 활발했음을 보여주는 살아있는 증거이다.
- ‘인구 대교체’ 가설: 더 오래된 임치(이란계)와 그 이후 시대인 에긴골(동아시아계)의 유전적 구성이 완전히 다른 것은, 기원전 500년 이후 흉노와 한(漢)나라가 등장하는 시기에 이 지역에서 거대한 인구 교체가 일어났을 가능성을 시사한다. 서쪽에서 온 사람들이 살던 땅에, 동아시아계 사람들이 새로 들어와 주요 인구 집단이 되었을 수 있다는 것이다.
결론적으로 이 논문은 고대 DNA 분석을 통해, 우리가 생각했던 것보다 훨씬 이른 시기에 유라시아 동쪽과 서쪽의 인구 이동과 혼합이 활발했으며, 이후 극적인 인구 구성의 변화가 있었음을 명확히 보여준 중요한 연구라 할 수 있다.
[논문번역]
[Abstract]
Mitochondrial hypervariable ]region I genetic data from ancient populations at two sites in Asia-Linzi in Shandong (northern China) and Egyin Gol in Mongolia were reanalyzed to detect population affinities. Data from 51 modern populations were used to generate distance measures (FST′S) to the two ancient populations. The tests first analyzed relationships at the regional level and then compiled the top regional matches for an over- all comparison to the two probe populations. The reanalysis showed that the Egyin Gol and Linzi populations have clear distinctions in genetic affinity. The Egyin Gol population as a whole appears to bear close affinities with modern populations of northern East Asia. The Linzi population seems to have some genetic affinities with the West, as suggested by the original anal- ysis, although the original attribution of “European-like” seems to be mis- leading. We suggest that the Linzi individuals are potentially related to early Iranians, who are thought to have been widespread in parts of Central Eur- asia and the steppe regions in the first millennium B.C., although some sig- nificant admixture between a number of populations of varying origin cannot be ruled out. We also examine the effect of sequence length on this type of genetic data analysis and discuss the results of previous studies on the Linzi sample.
아시아에 있는 두 유적지, 즉 중국 북부 산동(山東)의 임치(臨淄)와 몽골의 에긴골(Egyin Gol)에서 나온 고대 인구 집단의 미토콘드리아 DNA 유전 정보(과변이 부위 I)를 재분석하여 인구 집단 간의 관계를 파악했다. 현대인 51개 집단의 데이터를 이용해 두 고대 인구 집단과의 유전적 거리(FST)를 측정했다. 먼저 지역별로 관계를 분석한 뒤, 각 지역에서 가장 가까운 집단들을 뽑아 두 고대 인구 집단과 종합적으로 비교했다. 재분석 결과, 에긴골 집단과 임치(臨淄) 집단은 유전적으로 뚜렷한 차이를 보였다. 에긴골 집단 전체는 현대 북동아시아인들과 매우 가까운 관계를 보였다. 임치(臨淄) 집단은 기존 연구에서 제기된 것처럼 서쪽과 어느 정도 유전적 관계가 있는 것으로 보인다. 그러나 ‘유럽인과 유사하다’는 기존의 평가는 오해의 소지가 있다. 우리는 임치(臨淄) 주민들이 초기 이란인(Iranians)과 관련이 있을 가능성이 있다고 본다. 초기 이란인들은 기원전 제1천년기에 중앙 유라시아와 초원 지역에 널리 퍼져 있었던 것으로 생각된다. 물론 다양한 기원을 가진 여러 인구 집단 사이에 상당한 혼혈이 있었을 가능성도 배제할 수 없다. 또한, 우리는 이런 종류의 유전 데이터 분석에서 염기서열 길이가 어떤 영향을 미치는지 살펴보고, 임치(臨淄) 샘플에 대한 이전 연구 결과들을 논의한다.
목차
서론 Introduction
Recent analyses of ancient DNA from sites in northern China and Mongolia have provided interesting results regarding the genetic history of the region and of eastern Central Eurasia in general (e.g. Wang et al. 2000, Keyser-Tracqui et al. 2003). The period of the sites in question stretches from around the middle first millennium BC to the first few centuries AD and represents an important time period in the area: the rise of the Han dynasty in China and the Hsiung-Nu on the Mongolic steppe, the possible earliest appearances of Turks and Mongols, and the earliest attested conflicts between ancient Chinese and steppe peoples of Inner Asia. Elsewhere in Central Eurasia, the Scythians and Sarmatians appeared in the farthest western portion of the steppe in south Russia and have been putatively connected to Indo-Iranians or Iranians (a branch of Indo-European). In the central portions of the steppe (roughly modern-day Kazakhstan) not much is known for certain, though there is evidence of a group(s) of people referred to as the Saka, who are commonly identified as Indo-Iranian (or Iranian) and were nomadic pastoralists like the Scythians and Hsiung-Nu. More highly attested are the Sogdians, sedentary Iranians of the Transoxus region. Further east, in present day Xinjiang, there were possibly Indo-European peoples such as the Tokharians (a group of Indo-European speakers attested with recorded documents) and the peoples represented by the various mummified remains from the region in the second millennium BC through the first few centuries AD. Along with these peoples there are of course many others of whom we know very little or nothing in this period (e.g. Ob-Ugrians) (for a general discussion of the above, see Mallory 1989, Sinor 1991, Mair 1998). The two ancient sites in this study, Egyin Gol (Keyser-Tracqui et al. 2003) and Linzi (Wang et al. 2000), thus reflect a key period in the region (as well as Central Eurasia in general). It is clear that changes in the social, political, economic, and cultural realms occurred. However, the exact degree to which these various cultural and linguistic groups represented biological populations is debatable, and it is unclear whether the aforementioned changes were accompanied by the movements of such biological populations.
최근 중국 북부와 몽골 유적지의 고대 DNA 분석은 이 지역과 동부 중앙유라시아 전반의 유전적 역사에 관해 흥미로운 결과를 제공했다(Wang et al. 2000, Keyser-Tracqui et al. 2003). 문제의 유적지 시기는 기원전 1천년기 중반부터 서기 수 세기까지 이어지며, 이 지역에서 중요한 시기를 나타낸다. 즉, 중국의 한(漢) 왕조와 몽골 초원 흉노(匈奴)의 발흥, 돌궐(突厥)과 몽골의 가장 이른 출현 가능성, 고대 중국인과 내아시아 초원 민족 간의 입증된 가장 이른 갈등을 포함한다. 중앙유라시아의 다른 곳에서는 스키타이(Scythian)와 사르마티아(Sarmatian)가 러시아(Russia) 남부 초원의 가장 서쪽 부분에 나타났다. 이들은 인도이란인(Indo-Iranian) 또는 이란인(Iranian)(인도유럽어족의 한 분파)과 추정적으로 연결되었다. 초원의 중앙 부분(대략 오늘날의 카자흐스탄(Kazakhstan))에 대해서는 확실히 알려진 바가 많지 않다. 하지만 사카(Saka)라고 불리는 집단의 증거가 있다. 이들은 흔히 인도이란인(Indo-Iranian)(또는 이란인)으로 식별되며 스키타이(Scythian)나 흉노(匈奴)와 같은 유목 목축민이었다. 트랜스옥시아나(Transoxiana) 지역의 정주 이란인(Iranian)인 속특(粟特)은 더 많이 입증되었다. 더 동쪽인 오늘날의 신강(新疆)에는 토화라(吐火羅)(기록된 문서로 입증된 인도유럽어 사용 집단)와 같은 인도유럽인(Indo-European)들이 있었을 가능성이 있다. 기원전 2천년기부터 서기 수 세기까지 이 지역의 다양한 미라 유적으로 대표되는 민족들도 있었다. 이 민족들과 함께 이 시기에 우리가 거의 알지 못하거나 전혀 모르는 다른 많은 민족들(예: 오브우고르인(Ob-Ugrian))이 당연히 존재한다(위 내용에 대한 일반적인 논의는 Mallory 1989, Sinor 1991, Mair 1998 참조). 이 연구의 두 고대 유적지인 에긴골(Egyin Gol)(Keyser-Tracqui et al. 2003)과 임치(臨淄)(Wang et al. 2000)는 중앙유라시아 전반뿐만 아니라 이 지역의 핵심 시기를 반영한다. 사회적, 정치적, 경제적, 문화적 영역에서 변화가 일어났음은 분명하다. 그러나 이 다양한 문화적, 언어적 집단이 생물학적 인구 집단을 어느 정도 대변했는지는 논란의 여지가 있다. 앞서 언급한 변화들이 생물학적 인구 집단의 이동을 동반했는지도 불분명하다.
The question of the ancient history of northern China and Mongolia is a difficult issue. Traditionally, many have taken the approach that ‘China is an island.’ However, this invariably is false (as with the ‘Europe is an island’ model). Connections existed across Eurasia back to at least the first millennium, if not earlier (Bentley 2000). Moreover the connection of the biological past to the cultural past has not been clearly detailed, although various arguments have been made. Lattimore (1951) suggests that the difference between the peoples of Central Asia (which he defined as Manchuria, Mongolia, Tibet, and Chinese Turkestan) and those of sedentary China (which he defined as the primarily agricultural areas of China proper) was the difference between an extensive pastoral economy in Central Asia (although there are some places with agriculture or a mixture of economies including Manchuria and the oases of Sinkiang) and an intensive agricultural economy in China. He also points to the inability of states with a mixed economy of both pastoral nomadism and intensive agriculture to succeed (though this is not entirely true, case in point Manchuria or historical “Central Asia”). Lattimore further suggests that the “Northern Barbarians” were originally of the same ethnic stock as Northern Chinese but were split through economic differentiation. This led to differentiation in the rates of change (of culture, technology, etc.) that split these early peoples into two “orbits.” Lattimore argues that it was the expansion of the early Chinese that pushed out the peripheral groups who would become the early “barbarians” by the 5th century BC. However, the question of connections between the peoples of the steppe has continued to generate a large amount of work, some of which conflicts with Lattimore. A.P. Okladnikov (1990) argued that early on there were Europoid peoples in Inner Asia who would later move down off the steppe into India and Iran (and Europe as well). According to this view, Mongoloids, who are traditionally thought of as inhabiting the Inner Asian areas, did not appear until around 1200 BC. Further complicating the question of the human biological history of eastern Central Eurasia are the so-called “mummies of Urumchi” or “mummies of the Tarim Basin,” who have often been associated more with “Europoids” or “Caucasoids” rather than “Mongoloids”. These remains have not only been potentially related to Indo-Europeans in a biological sense, but also culturally (e.g. “Tartan” clothing). They have also been putatively connected to various Indo-European groups of multiple time periods from around the region, including the Tokharians, the Saka, the Andronovo of the Central Asian steppe, as well as the Afanasievo of the Altai and western Sayan ranges in southern Siberia (Mair 1995, 1998).
중국 북부와 몽골의 고대사 문제는 어려운 주제다. 전통적으로 많은 이들이 ‘중국은 섬이다’라는 접근 방식을 취했다. 그러나 이는 ‘유럽(Europe)은 섬이다’ 모델과 마찬가지로 예외 없이 틀렸다. 더 이르지 않더라도 최소한 기원전 1천년기까지 거슬러 올라가면 유라시아(Eurasia) 전역에 연결이 존재했다(Bentley 2000). 더욱이 다양한 주장이 제기되었음에도 생물학적 과거와 문화적 과거의 연결은 명확하게 상술되지 않았다. 래티모어(Lattimore)(1951)는 중앙아시아(그가 만주(滿洲), 몽골, 서장(西藏), 중국령 돌궐(中國領 突厥)로 정의함) 민족과 정주 중국(그가 중국 본토의 주로 농경 지역으로 정의함) 민족의 차이가 중앙아시아의 광범위한 목축 경제와 중국의 집약적 농경 경제의 차이라고 제안했다(비록 만주(滿洲)나 신강(新疆)의 오아시스를 포함하여 농경이나 혼합 경제가 나타나는 곳도 있지만). 그는 또한 유목 목축과 집약적 농경의 혼합 경제를 가진 국가들이 성공하지 못했다는 점을 지적했다(비록 만주(滿洲)나 역사적 “중앙아시아”의 사례처럼 이것이 전적으로 사실은 아니지만). 래티모어(Lattimore)는 더 나아가 “북적(北狄)”이 원래 중국 북부인과 동일한 민족 계통이었으나 경제적 차별화를 통해 분리되었다고 제안했다. 이는 문화, 기술 등의 변화 속도 차이로 이어졌고 이 초기 민족들을 두 개의 “궤도”로 나누었다. 래티모어(Lattimore)는 기원전 5세기경 초기 “오랑캐”가 된 주변 집단들을 밀어낸 것은 초기 중국인의 팽창이었다고 주장한다. 그러나 초원 민족 간의 연결 문제는 계속해서 방대한 연구를 낳았다. 그중 일부는 래티모어(Lattimore)의 주장과 충돌한다. 오클라드니코프(A.P. Okladnikov)(1990)는 초기 내아시아에 인종적으로 유럽인(Europoid)들이 있었으며 이들이 나중에 초원을 떠나 인도(印度)와 이란(Iran)(그리고 유럽(Europe))으로 남하했다고 주장했다. 이 견해에 따르면 전통적으로 내아시아 지역에 거주하는 것으로 여겨지는 몽골인(Mongoloid)은 기원전 1200년경이 되어서야 나타났다. 동부 중앙유라시아 인류 생물학적 역사의 문제를 더욱 복잡하게 만드는 것은 소위 “오로목제(烏魯木齊) 미라” 또는 “타림분지(塔里木盆地) 미라”다. 이들은 흔히 “몽골인(Mongoloid)”보다는 “유럽인(Europoid)” 또는 “코카서스인(Caucasoid)”과 더 많이 연관되어 왔다. 이 유해들은 생물학적 의미에서 인도유럽인(Indo-European)과 관련이 있을 뿐만 아니라 문화적으로도(예: “타탄” 의복) 관련이 있을 가능성이 있다. 이들은 또한 토화라(吐火羅), 사카(Saka), 중앙아시아 초원의 안드로노보(Andronovo)뿐만 아니라 남부 시베리아(Siberia)의 알타이(Altai) 및 서부 사얀(Sayan) 산맥의 아파나시에보(Afanasievo)를 포함하여 이 지역 주변의 여러 시기에 걸친 다양한 인도유럽인(Indo-European) 집단과 추정적으로 연결되었다(Mair 1995, 1998).
Adding to this debate are the various theories of Indo-European origins and expansions, as argued by both Mallory (1989) and Renfrew (1987) as well as numerous others. The main component of this theory is that the Indo-Europeans originally represented a centralized cultural group, though there is great debate over the location of their origins and the time of dispersal and expansion as well as possible routes. Though Renfrew (1987) has argued for an Anatolian origin connected to the spread of Neolithic farming, there is an alternative argument detailed by Mallory, Gimbutas, and others (see Mallory 1989), who connect the Indo-Europeans to the south Russian steppe, possibly around the Black and/or Caspian sea as well as the southern Urals or northern Caucasus (and/or possibly Eastern Europe). Indo-European languages are generally divided up into centum (European, western) and satem (Indo-Iranian and Indo-Aryan, eastern) languages, though some discrepancies such as Tokharian (a centum language in the east) do exist (Mallory 1989).
이 논쟁에 더해지는 것은 맬러리(Mallory)(1989)와 렌프루(Renfrew)(1987)를 비롯한 수많은 이들이 주장하는 인도유럽인(Indo-European) 기원과 팽창에 관한 다양한 이론들이다. 이 이론의 핵심 구성 요소는 인도유럽인(Indo-European)이 원래 중앙 집중화된 문화 집단을 대표했다는 것이다. 하지만 그들의 기원 위치와 분산 및 팽창 시기, 가능한 경로에 대해서는 큰 논쟁이 있다. 렌프루(Renfrew)(1987)는 신석기(新石器) 농경의 확산과 연결된 아나톨리아(Anatolia) 기원을 주장했다. 하지만 맬러리(Mallory), 김부타스(Gimbutas) 등이 상세히 설명한 대안적인 주장(Mallory 1989 참조)이 있다. 이들은 인도유럽인(Indo-European)을 남부 러시아(Russia) 초원, 아마도 흑해(黑海) 및/또는 카스피해(Caspian Sea) 주변뿐만 아니라 남부 우랄(Ural) 또는 북부 코카서스(Caucasus)(및/또는 아마도 동유럽(東Europe))와 연결한다. 인도유럽어(Indo-European languages)는 일반적으로 켄툼(centum)(유럽계, 서부) 언어와 사템(satem)(인도이란계 및 인도아리아계, 동부) 언어로 나뉘지만, 토화라(吐火羅)(동부의 켄툼어)와 같은 일부 불일치는 존재한다(Mallory 1989).
The earliest expansions of Indo-Europeans are generally dated to sometime between the 5th millennium BC and the 3rd or 2nd millennium BC. These expansions have been connected by Anthony (1995) to the domestication of the horse on the steppe and to the later development of the war chariot (as well as wheeled vehicles in general, metallurgical developments, and herding). Some of the earliest evidence of horse domestication is at the site of Dereivka in the south Russian steppes dated to around 4000 BC, which is connected to the Stredni Stog culture (Anthony 1995). This evidence is related to possible bit wear, though see Levine (1999) for dispute. However, there is evidence that men may have hunted horses (as well as other animals) on the southern portion of the steppe as early as the late Paleolithic (Praslov 1989), suggesting that man may have had long contact in the region with horses. Also, some of the earliest evidence of chariots found to date comes from the Sintashta-Petrovka culture (possibly related to the Andronovo) on the steppe near the Volga-Caspian region and the Urals, dated to around 2000 BC (Anthony 1995). Early possible expansions of Indo-Europeans include the Germanic peoples, Celts, Greeks, Latins, and others into Europe as well as expansions east such as the Andronovo culture (Mallory 1989). The earliest eastern expansion may have been the aforementioned Afanasievo culture in the mid 4th millennium BC (Anthony 1998). Other movements include the possible migrations of Indo-Europeans into Xinjiang at least as early as the early second millennium BC (Kuzmina 1998) and the historically attested movements of Indo-Iranians and Indo-Aryans south into the Iranian Plateau region and India, though the exact nature or sequence of these is not certain (Mallory 1989, Parpola 1998). The connections of these eastern peoples of the putative Indo-European family farther east, such as into China, is subject to much scholarly debate, though there is some evidence of Indo-European loan words in Old Chinese as well as cultural and technological changes in northern China in the 3rd and 2nd millennium BC (Pulleyblank 1996, Kuzmina 1998, Beckwith 2002, Di Cosmo 2002). Certainly, sites such as Zhukaigou (roughly 2000 BC, Linduff 1995) and Linzi (Liangchun site, roughly 500 BC, Wang et al. 2000) in northern China, as well as mummies of the eastern Tarim Basin suggest that the history of the region, both culturally and biologically, may be very different from what it is today or even in known historical times. These may indicate alterations in the biological and cultural makeup of the region occurring as early as the Bronze Age or late Neolithic and even possibly earlier; though obviously because of the often poor connection between culture and biology, the relationships between the two must be examined with a fair bit of caution. Other sites, such as Egyin Gol in Mongolia (Keyser-Tracqui et al. 2003), shed light on the later shifts in the region, as well as explore the possible connections between the differentiation of the steppe peoples (or lack thereof), in their early stages of development, and China.
인도유럽인(Indo-European)의 가장 이른 팽창 시기는 일반적으로 기원전 5천년기와 기원전 3천년기 또는 2천년기 사이로 추정된다. 이러한 팽창은 앤서니(Anthony)(1995)에 의해 초원에서의 말 사육과 이후의 전차(바퀴 달린 탈것 전반, 금속공학적 발전, 목축 포함) 개발과 연결되었다. 말 사육의 가장 이른 증거 중 일부는 기원전 4000년경으로 추정되는 남부 러시아(Russia) 초원의 데레이브카(Dereivka) 유적지에 있으며 이는 스레드니 스토그(Stredni Stog) 문화와 연결된다(Anthony 1995). 이 증거는 가능한 재갈 마모와 관련이 있지만 논쟁에 대해서는 러빈(Levine)(1999)을 참조하라. 그러나 이르면 구석기(舊石器) 시대 후기에 초원 남부에서 인간이 말(및 다른 동물)을 사냥했을 수 있다는 증거가 있다(Praslov 1989). 이는 인간이 이 지역에서 말과 오랫동안 접촉했을 수 있음을 시사한다. 또한 현재까지 발견된 전차의 가장 이른 증거 중 일부는 볼가(Volga)-카스피(Caspian) 지역과 우랄(Ural) 인근 초원의 신타슈타-페트로프카(Sintashta-Petrovka) 문화(안드로노보(Andronovo)와 관련되었을 가능성 있음)에서 나오며 시기는 기원전 2000년경이다(Anthony 1995). 인도유럽인(Indo-European)의 초기 가능한 팽창에는 유럽(Europe)으로 향한 게르만(Germanic)족, 켈트(Celt)족, 그리스(Greek)족, 라틴(Latin)족 등과 안드로노보(Andronovo) 문화와 같은 동부로의 팽창이 포함된다(Mallory 1989). 가장 이른 동부 팽창은 기원전 4천년기 중반의 앞서 언급한 아파나시에보(Afanasievo) 문화였을 수 있다(Anthony 1998). 다른 이동에는 늦어도 기원전 2천년기 초기에 인도유럽인(Indo-European)이 신강(新疆)으로 이주했을 가능성(Kuzmina 1998)과 역사적으로 입증된 인도이란인(Indo-Iranian) 및 인도아리아인(Indo-Aryan)의 이란(Iran) 고원 지역 및 인도(印度) 남하가 포함된다. 하지만 이들의 정확한 성격이나 순서는 확실하지 않다(Mallory 1989, Parpola 1998). 이러한 인도유럽어족(Indo-European family) 추정 동부 민족들이 중국 내부와 같은 더 먼 동쪽으로 연결되는 것은 많은 학술적 논쟁의 대상이다. 비록 상고중국어(上古中國語)에 인도유럽어 차용어가 존재한다는 증거와 기원전 3천년기와 2천년기 중국 북부의 문화적, 기술적 변화라는 증거가 있기는 하다(Pulleyblank 1996, Kuzmina 1998, Beckwith 2002, Di Cosmo 2002). 분명 중국 북부의 주개가(朱開溝)(기원전 2000년경, Linduff 1995) 및 임치(臨淄)(양춘(凉春) 유적, 기원전 500년경, Wang et al. 2000)와 같은 유적들과 타림분지(塔里木盆地) 동부의 미라는 이 지역의 역사가 문화적으로나 생물학적으로나 오늘날 또는 알려진 역사 시대와 매우 다를 수 있음을 시사한다. 이는 이르면 청동기(靑銅器) 시대 또는 신석기(新石器) 시대 후기, 심지어 그보다 더 일찍 이 지역의 생물학적, 문화적 구성에 변화가 일어났음을 나타낼 수 있다. 비록 문화와 생물학 사이의 연결이 흔히 빈약하기 때문에 두 사이의 관계는 상당한 주의를 기울여 검토해야 하지만 말이다. 몽골의 에긴골(Egyin Gol)(Keyser-Tracqui et al. 2003)과 같은 다른 유적지들은 이 지역의 후기 변화를 조명한다. 또한 초기 발달 단계의 초원 민족 차별화(또는 결여)와 중국 간의 가능한 연결을 탐구한다.
The purpose of this particular study is a reexamination of two sites from eastern Central Eurasia, Linzi in China (Liangchun site) dated around 2500 years before present (Wang et al. 2000), and Egyin Gol in Mongolia dated to around the last few centuries BC to the first few centuries AD (Keyser-Tracqui et al. 2003). Linzi is in Shandong province in northern China, near the Yellow river and the Ordos region. It is presently part of the city of Zibo. Sixty-three individuals were examined in the original study from the Liangchun site in Linzi dating to around 500 BC. However, only 34 gave good results for mitochondrial DNA. Though the original study extracted longer sequences (287 bp), only the 185 bp segments (nt 16194-16378) actually used in their analysis were available in GenBank. The date places the material during the Spring-and-Autumn period in between the fall of the Eastern Chou dynasty and the rise of the Han dynasty. The original study also included 50 modern Han Chinese individuals from Linzi (labeled Qidu here). Samples also have been examined from the more recent Yixi site at Linzi (2000 before present, Oota et al. 1999), though these were not included here due to inconsistency in the Yixi data in GenBank (nearly half lacked sufficient data for inclusion). Because the length of the sequences available is only 185 base pairs, accurate comparative genetic analysis with other populations is difficult. The authors report that they found the Linzi material clustered closely with modern Europeans, particularly Finnish, Turkish, and Icelanders (Wang et al. 2000). However, as we will show, this analysis may be imprecise (as suggested by Yao et al. 2003). Additionally, we will examine the fact that the putative “cousins” of the Europeans, the so-called Indo-Iranians, are known to have been widespread in Central Eurasia at that time.
이 연구의 목적은 동부 중앙유라시아의 두 유적지를 재검토하는 것이다. 하나는 약 2500년 전으로 추정되는 중국의 임치(臨淄)(양춘(凉春) 유적)이다(Wang et al. 2000). 다른 하나는 기원전 마지막 수 세기에서 기원후 첫 수 세기경으로 추정되는 몽골의 에긴골(Egyin Gol)이다(Keyser-Tracqui et al. 2003). 임치(臨淄)는 황하(黃河) 및 오르도스(Ordos) 지역과 가까운 중국 북부 산동(山東)성에 있다. 현재는 치보(淄博)시의 일부이다. 원래 연구에서는 기원전 500년경으로 추정되는 임치(臨淄) 양춘(凉春) 유적의 63명을 조사했다. 그러나 34명만이 미토콘드리아(Mitochondria) DNA에 대해 좋은 결과를 보였다. 원래 연구에서는 더 긴 염기서열(287bp)을 추출했다. 하지만 그들의 분석에 실제로 사용된 185bp 구간(nt 16194-16378)만이 젠뱅크(GenBank)에서 이용 가능했다. 이 연대는 해당 자료를 동주(東周) 왕조의 멸망과 한(漢) 왕조의 부흥 사이인 춘추(春秋) 시대에 놓이게 한다. 원래 연구에는 임치(臨淄) 출신의 현대 한족(漢族) 50명(여기서는 기도(齊都)로 표기)도 포함되었다. 임치(臨淄)의 더 최근 유적인 이시(乙烯) 유적(2000년 전, Oota et al. 1999)의 표본들도 조사된 바 있다. 하지만 젠뱅크(GenBank)에 있는 이시(乙烯) 데이터의 불일치 때문에 여기서는 포함되지 않았다(절반 가까이가 포함하기에 충분한 데이터가 부족했다). 이용 가능한 염기서열의 길이가 185염기쌍에 불과하다. 그래서 다른 인구 집단과의 정확한 비교 유전 분석은 어렵다. 저자들은 임치(臨淄) 자료가 현대 유럽(Europe)인, 특히 핀란드(Finland)인, 터키(Turkey)인, 아이슬란드(Iceland)인과 밀접하게 군집을 이루는 것을 발견했다고 보고한다(Wang et al. 2000). 그러나 우리가 보여주겠지만, 이 분석은 부정확할 수 있다(Yao et al. 2003이 제안한 바와 같이). 추가로 우리는 유럽(Europe)인의 추정 “사촌”인 소위 인도이란인(Indo-Iranian)이 당시 중앙유라시아에 널리 퍼져 있었다고 알려진 사실을 검토할 것이다.
Egyin Gol is a necropolis in northern Mongolia (labeled simply Egyin below). The original study successfully extracted DNA successfully from 62 specimens ranging from around the 3rd century BC to the 2nd century AD, including mitochondrial and nuclear DNA. In this study, we will examine the mitochondrial DNA (nt 16009-16390). The site sits along the Egyin Gol River, a tributary of the Selenge River, which flows into Lake Baikal. The site has been attributed possibly to the Hsiung-Nu, who the authors describe as an ancient “Turkomongolian” tribe (Keyser-Tracqi et al. 2003). However, the exact relationship of the Hsiung-Nu to either Mongols or Turks (who do not definitively appear until the first few centuries AD with the Jou-Jan and their “blacksmith slaves” the Turks, (Sinor 1990)) is not clear. Moreover, the Chinese records of the Hsiung-Nu have proven them difficult to classify culturally and linguistically, and the origins of the Hsiung-Nu are not clear (Di Cosmo 2002). The necropolis was divided into three sectors (there was also a fourth zone, D, but no DNA samples originate there), with A (the oldest) and B representing older sections, after whose fusion sometime in the early centuries AD new graves were dug in what is called sector C. The authors also point out that both the paternal lineage and the mtDNA sequences shared by four of the paternal relatives in sector C have been found in modern day Turkish individuals, as well as in two of the graves from the older A and B sectors. The authors suggest that this evidence may point to a “Turkish component” to the Hsiung-Nu tribe in later periods (Keyser-Tracqui et al. 2003). In this study, we will attempt to examine the affiliation of these individuals to populations around Eurasia, as well as to look at any possible differences in the genetic relationships of the older A and B sectors to the newer possibly “Turkic” sector C.
에긴골(Egyin Gol)은 몽골 북부의 묘지이다(아래에서는 단순히 에긴(Egyin)으로 표기). 원래 연구에서는 기원전 3세기부터 기원후 2세기경까지의 표본 62개에서 미토콘드리아(Mitochondria) 및 핵 DNA를 성공적으로 추출했다. 이 연구에서 우리는 미토콘드리아(Mitochondria) DNA(nt 16009-16390)를 조사할 것이다. 이 유적지는 바이칼(Baikal) 호수로 흘러드는 셀렝게(Selenge) 강의 지류인 에긴골(Egyin Gol) 강을 따라 자리 잡고 있다. 이 유적지는 저자들이 고대 “튀르크몽골(Turko-Mongolian)” 부족으로 묘사한 흉노(匈奴)에 귀속될 가능성이 있다(Keyser-Tracqui et al. 2003). 그러나 흉노(匈奴)와 몽골인 또는 돌궐(突厥)인 사이의 정확한 관계는 불분명하다(돌궐(突厥)인은 유연(柔然)과 그들의 “대장장이 노예”인 돌궐(突厥)과 함께 기원후 첫 수 세기가 되어서야 명확히 등장한다(Sinor 1990)). 더욱이 흉노(匈奴)에 대한 중국 기록은 그들을 문화적, 언어적으로 분류하기 어렵다는 점을 증명했다. 흉노(匈奴)의 기원 역시 명확하지 않다(Di Cosmo 2002). 묘지는 세 구역으로 나뉘었다(네 번째 구역인 D도 있었지만 그곳에서 나온 DNA 표본은 없다). A(가장 오래됨)와 B는 더 오래된 구역을 나타낸다. 기원후 초기 수 세기경 이들이 융합된 후 소위 섹터(Sector) C라 불리는 곳에 새로운 무덤이 파였다. 저자들은 또한 섹터(Sector) C의 부계 친척 4명이 공유하는 부계 혈통과 mtDNA 염기서열이 현대 터키(Turkey)인 개체들뿐만 아니라 더 오래된 A 및 B 구역의 무덤 두 곳에서도 발견되었다고 지적한다. 저자들은 이 증거가 후기 흉노(匈奴) 부족의 “튀르크(Turk)적 구성 요소”를 가리킬 수 있다고 제안한다(Keyser-Tracqui et al. 2003). 이 연구에서 우리는 이 개체들이 유라시아(Eurasia) 주변의 인구 집단과 갖는 연관성을 조사하려고 시도할 것이다. 더 오래된 A 및 B 구역과 더 새롭고 아마도 “튀르크계(Turkic)”인 섹터(Sector) C의 유전적 관계에서 가능한 차이도 살펴볼 것이다.
재료 및 방법 Materials and Methods
For the purposes of this study we examined as wide a range of variation as possible at the population level. For this analysis we compiled 3,703 mtDNA HVSI sequences from 51 modern populations (including the Qidu samples) from across Eurasia in addition to the two aforementioned ancient populations (total populations =53). Most of the populations were included regardless of whether they were thought to be related to the two ancient populations, although a few populations were added because of hypothetical relationships, and the African Biaka were included as an outgroup. Limitations in computer power and software design placed restrictions on the total number of populations used in the analysis, but populations were included up to reasonable limits of this constraint. For example, DNAsp, a program used in the analysis, consistently failed at about 2,500 sequences (other programs, such as Arlequin, are written specifically to handle a maximum amount; in Arlequin’s case that amount is 1,000 sequences). Furthermore, any increase in populations or sequences increased the number of calcula- tions needed at an exponential rate. Even the use of the local supercomputing network could only ameliorate these issues, not eliminate them. Thus the idea here was to minimize bias within the constructed sample. Although additional populations, such as Tibetans, Russians, and some Siberian groups, are available, the analysis of such a large data set proved prohibitive.
이 연구를 위해, 우리는 인구 집단 수준에서 가능한 한 넓은 범위의 변이를 조사했다. 이 분석을 위해 우리는 51개 현대 인구 집단으로부터 3,703개의 미토콘드리아 DNA HVSI 염기서열을 수집했고, 앞서 언급한 두 고대 인구 집단을 더했다(총 53개 집단). 대부분의 인구 집단은 두 고대 집단과의 관련성 여부와 상관없이 포함했다. 일부 집단은 가설적 관계 때문에 추가했고, 아프리카(Africa)의 비아카(Biaka)족은 외부 집단(outgroup)으로 포함했다. 컴퓨터 성능과 소프트웨어 설계의 한계로 인해 분석에 사용되는 총인구 집단 수에 제약이 있었지만, 이 제약의 합리적인 한도까지 인구 집단을 포함했다. 예를 들어, 분석에 사용된 프로그램인 DNAsp는 약 2,500개의 염기서열에서 계속 오류가 발생했다. Arlequin과 같은 다른 프로그램들은 처리할 수 있는 최대량이 정해져 있는데, Arlequin의 경우 1,000개였다. 게다가, 인구 집단이나 염기서열이 늘어나면 필요한 계산량은 기하급수적으로 증가했다. 지역 슈퍼컴퓨팅 네트워크를 사용해도 이 문제들을 완화할 수는 있었지만, 완전히 해결할 수는 없었다. 따라서 여기서의 목표는 구성된 표본 내의 편향을 최소화하는 것이었다. 티베트인, 러시아인, 일부 시베리아(Siberia) 집단과 같은 추가적인 인구 집단 데이터가 있었지만, 이렇게 큰 데이터 세트를 분석하는 것은 현실적으로 불가능했다.
The populations included in the study are listed in Table 1. The table in- cludes individual population data for several values and the values for the total sample with and without Linzi and Qidu (in order to account for the short se- quence length of Linzi and Qidu, which may result in inaccurate values when included). Note that the Basque data are not included in Table 1 for reasons specified later. The averages for the values within populations are also included. The average number of individuals sequenced was 71, although this varied widely, along with an average of 47 haplotypes per population, excluding all gaps or missing data. The average number of differences between sequences within populations was 5.73, with an average nucleotide diversity (the average number of differences per site between any two sequences) of roughly 0.015 and an aver- age haplotype diversity (a measure of genic variation, which equals 1−∑x2 where x equals the haplotype frequency) of 0.96. However, individual population averages appear to depend to some degree on sample size and sequence length, as would be expected. For the average number of differences there was a significant correlation to sequence length (0.571, ρ=0.000) but not sample size (-0.263, ρ=0.058). For nucleotide diversity there was a significant correlation to both sequence length (-0.438, ρ=0.001) and sample size (0.290, ρ=0.035). Note that the diversity indexes for the ancient populations fall within the range of those for modern populations, which indicates that they are comparable to modern populations in terms of variation.
연구에 포함된 인구 집단은 표 1에 나열되어 있다. 이 표에는 개별 인구 데이터와 함께, 임치(臨淄)와 기도(齊都)의 짧은 염기서열 길이로 인한 부정확성을 고려하기 위해 이들을 포함했을 때와 제외했을 때의 전체 표본 값들이 포함되어 있다. 바스크(Basque)인 데이터는 나중에 명시할 이유로 표 1에 포함되지 않았다는 점에 유의해야 한다. 인구 집단 내 평균값들도 포함했다. 평균적으로 71명의 염기서열을 분석했고, 집단당 평균 47개의 하플로타입(haplotype)이 있었다. 집단 내 염기서열 간의 평균 차이 수는 5.73이었고, 평균 염기 다양성은 약 0.015, 평균 하플로타입 다양성은 0.96이었다. 하지만 예상대로, 개별 인구 집단의 평균값은 표본 크기와 염기서열 길이에 어느 정도 의존하는 것으로 보인다. 평균 차이 수는 염기서열 길이와는 유의미한 상관관계가 있었지만 표본 크기와는 없었다. 염기 다양성은 염기서열 길이와 표본 크기 모두와 유의미한 상관관계가 있었다. 고대 인구 집단의 다양성 지수는 현대 인구 집단의 범위 내에 속하는데, 이는 이들이 변이 측면에서 현대 인구 집단과 비교할 수 있음을 나타낸다.


표 1. 인구 데이터 (지역별 알파벳순 정렬)
- Jorde 등 (1995).
- Keyser-Tracqui 등 (2003).
- Wang 등 (2000).
- 미발표 GenBank 데이터 (Metspalu 등).
- Bertranpetit 등 (1995). 출처 불일치로 인해 데이터 미포함. 바스크 데이터에 대한 자세한 설명은 본문 참조.
- Plaza 등 (2003).
- 미발표 GenBank 데이터 (Szabo 등).
- 미발표 GenBank 데이터 (Riedla 등).
- Richards 등 (1996).
- 미발표 GenBank 데이터 (Kalmar 등).
- Helgason 등 (2001).
- Sajantila 등 (1995).
- Brakez 등 (2001).
- 미발표 GenBank 데이터 (Kaldma 등).
- 미발표 GenBank 데이터 (Kivisild 등).
- T. Kivisild와 M. Metspalu의 개인 서신 (2003).
- Al-Zahery 등 (2003).
- Comas 등 (1998).
- Comas 등 (2000).
- Comas 등 (1996).
- Oota 등 (2001).
- Kong 등 (2003).
- Yao (2002a).
- Kivisild 등 (2002).
- Redd와 Stoneking (1999).
- Imaizumi 등 (2002).
- Lee 등 (1997).
- Oota 등 (2002).
These populations were divided into three loosely defined regions at the discretion of the authors. These groups were: Europe (Armenians, Georgians, Mari, Moksha, Saami, Slovakians, RomB, RomS, Rom2S, Germans, Hungarians, Cumans, Basques, Catalans, Icelanders, and the Moroccans), South and Southwest Asia (Lambadi, Lobana, Uttar Pradesh, Boqsa, Pushtoon, Pakistan, Parsi, Iranians, Iraqi, Kurds, Turks, Kashmir, and Tunisians), and East and Central Asians (Kazakh, KazakhXJ, Uighur, UighurXJ, KirghizHL, KirghizLL, Guangdong, Guangdong2, Yunnan, Vietnamese, Indonesians, Akha, Koreans, Japanese, Mongolians, Ewenki, Wuhan, Shandong, Liaoning, and Qidu). Though there could be some debate over the division into these regions, it is not extremely important since, as we will see below, only the top 9 or 10 populations from each region were used for the final comparison. These initial regional divisions were necessary in order to break up the data into manageable datasets. Moreover, the purpose of the study was not to determine the vast connections between populations in Eurasia, nor do we claim that the results can be used in this way. Rather, the purpose is a very narrow focus, that being how these modern populations from various regions relate to the two ancient populations of Linzi and Egyin Gol. To this purpose, we calculated regional distance measures (principally Fst’s, a measure rooted in heterozygosity values within and among populations) for each region and then included the top nine or ten “matches” (the lowest Fst’s) from the region for a “total” comparison. The central assumption here is that the populations from each region best represent that region as far as biological relation to the two ancient sites. We may lose in this method of comparison some of the minutiae of the more distant relations, but the closer relations (such as the top ten) should be accurate. Obviously, a particular Central Asian population might have seemed relatively closer to the Egyin material if placed in the South and Southwest Asian region, for instance. However, every population was given a fair chance to compete (“free competition”) instead of arbitrarily being included or excluded in the analysis, which we believe results in a more accurate estimate of relationships.
이 인구 집단들은 저자들의 재량에 따라 느슨하게 정의된 세 지역으로 나뉘었다. 이 집단들은 유럽(아르메니아인, 조지아인, 마리인, 목샤인, 사미인, 슬로바키아인, 롬B, 롬S, 롬2S, 독일(獨逸)인, 헝가리인, 쿠만인, 바스크인, 카탈루냐인, 아이슬란드인, 모로코인), 남/서남아시아(람바디, 로바나, 우타르프라데시, 복사, 파슈툰, 파키스탄, 파르시, 이란인, 이라크인, 쿠르드인, 터키인, 카슈미르, 튀니지인), 그리고 동/중앙아시아(카자흐, 카자흐XJ, 위구르, 위구르XJ, 키르기스HL, 키르기스LL, 광동(廣東), 광동2(廣東2), 운남(雲南), 베트남(越南)인, 인도네시아인, 아카, 한국인, 일본인, 몽골인, 에벤키, 무한(武漢), 산동(山東), 요녕(遼寧), 기도(齊都))였다. 이 지역들의 구분에 대해 약간의 논쟁이 있을 수 있다. 하지만 아래에서 볼 수 있듯이 각 지역에서 상위 9개 또는 10개의 집단만이 최종 비교에 사용되었기 때문에 이는 그다지 중요하지 않다. 이러한 초기 지역 구분은 데이터를 관리 가능한 크기로 나누기 위해 필요했다. 더욱이 이 연구의 목적은 유라시아(Eurasia) 인구 집단 간의 광범위한 연관성을 밝히는 것이 아니었다. 우리는 이 결과가 그런 방식으로 사용될 수 있다고 주장하지도 않는다. 오히려 그 목적은 매우 좁은 범위에 맞춰져 있다. 즉, 다양한 지역의 현대 인구 집단이 임치(臨淄)와 에긴골(Egyin Gol)이라는 두 고대 인구 집단과 어떻게 연관되는지 파악하는 것이다. 이를 위해 우리는 각 지역에 대해 지역 거리 척도(주로 인구 집단 내외의 이형접합성 값에 기반한 척도인 Fst)를 계산했다. 그런 다음 “전체” 비교를 위해 해당 지역에서 상위 9~10개의 “일치 항목”(가장 낮은 Fst 값)을 포함시켰다. 여기서 핵심적인 가정은 각 지역의 인구 집단이 두 고대 유적지와의 생물학적 연관성 측면에서 해당 지역을 가장 잘 대표한다는 것이다. 이런 비교 방법에서는 더 먼 관계의 세부 사항 중 일부를 놓칠 수 있다. 하지만 더 가까운 관계(상위 10개 등)는 정확할 것이다. 예를 들어, 특정 중앙아시아 집단을 남/서남아시아 지역에 배치했다면 에긴(Egyin) 자료에 상대적으로 더 가깝게 보였을 수도 있다. 그러나 분석에 임의로 포함하거나 제외하는 대신 모든 집단에 공정한 경쟁 기회(“자유 경쟁”)를 주었다. 우리는 이것이 관계를 더 정확하게 추정하는 결과를 낳는다고 믿는다.
A quick glance over the regional populations is probably in order (see the original sources for more detailed information). The Europeans included Germans and Icelanders representing the Germanic and Scandinavian peoples. The Catalans represented a Western European population, while the Basque represented a supposed isolate in Western Europe. The Slovakians represented East Europeans and Slavs. The Hungarians are Ob-Ugrian speakers, though they probably have at least some historical connection to Turks (the name Hungarian probably derives from Onoghur, a Turkic people of Central Eurasia in the early middle ages, Golden [1991]). The Cumans were originally Turks known by various names in different sources. The Rom populations are gypsies of Eastern Europe, whom many believe to be ultimately descended from northern Indian ancestors (Gresham et al. 2001). The Saami are Finnic (Ob-Ugrian) speakers from northern Scandinavia, while the Mari and Moksha are Ob-Ugrian speakers from Russia closer to the Urals. The Armenians and Georgians represent Caucasus populations. The Moroccans were included here as a check against suggestions of significant gene flow across the Mediterranean between North Africa and Southern Europe, at least in the West (Plaza et al. 2003).
아마도 지역 인구 집단을 간략히 살펴보는 것이 좋을 것이다(더 자세한 정보는 원문 출처 참조). 유럽인에는 게르만(Germanic) 및 스칸디나비아(Scandinavia) 민족을 대표하는 독일(獨逸)인과 아이슬란드인이 포함되었다. 카탈루냐인은 서유럽 인구를 대표했고, 바스크인은 서유럽에서 고립된 것으로 추정되는 집단을 대표했다. 슬로바키아인은 동유럽인과 슬라브(Slav)족을 대표했다. 헝가리인은 오브우고르(Ob-Ugrian)어 사용자지만 역사적으로 터키인과 적어도 약간의 연관성이 있을 것이다(헝가리라는 이름은 아마도 중세 초기 유라시아 중앙부의 튀르크계 민족인 오노구르(Onoghur)에서 유래했을 것이다, Golden [1991]). 쿠만인은 원래 여러 문헌에서 다양한 이름으로 알려진 튀르크인이었다. 롬(Rom) 집단은 동유럽의 집시들이며, 많은 사람들은 이들이 궁극적으로 북인도 조상의 후손이라고 믿는다(Gresham et al. 2001). 사미인은 스칸디나비아 북부의 핀계(Finnic)(오브우고르계) 언어 사용자이다. 한편 마리인과 목샤인은 우랄(Ural) 산맥에 더 가까운 러시아의 오브우고르어 사용자이다. 아르메니아인과 조지아인은 코카서스(Caucasus) 인구를 대표한다. 모로코인은 적어도 서부 지역에서 북아프리카와 남유럽 사이의 지중해를 가로지르는 상당한 유전자 흐름이 있었다는 제안을 검증하기 위해 여기에 포함되었다(Plaza et al. 2003).
The South and Southwest Asians included the Uttar Pradesh and Boqsa samples from Uttar Pradesh. The Lobana are also northern Indians from Punjab. The Parsi are from northwest India (mostly Gujarat) but supposedly descend from Iranian migrants (hence the name, Pars, Fars, Persians). The Lambadi are the “gypsies” of India, from whom the gypsies of the world are hypothetically descended (see above). However, most Lambadi live in the North and Northwest of the Indian subcontinent, while these samples come from Andhra Pradesh in the East Central region. The Pakistan and Pushtoon (original source name preserved, likely Pashtun) are from Pakistan, while the Kashmiri are from Kashmir. Further east, we have the populations of Iranians, Iraqi, and Turks representing the Near East, as well as the Kurds from Iraq. The Tunisians were also included as a related North African population that may have absorbed similar Arab or Near Eastern gene flow since the genesis of Islam.
남/서남아시아인에는 우타르프라데시의 우타르프라데시 및 복사 표본이 포함되었다. 로바나 역시 펀자브(Punjab) 출신의 북인도인이다. 파르시인은 인도 북서부(주로 구자라트(Gujarat)) 출신이지만, 이란 이주민의 후손으로 추정된다(따라서 이름이 파르스(Pars), 파르스(Fars), 페르시아(Persia)인이다). 람바디는 인도의 “집시”이며, 전 세계 집시들이 이들로부터 유래했다는 가설이 있다(위 내용 참조). 그러나 대부분의 람바디는 인도 아대륙의 북부 및 북서부에 거주하는 반면, 이 표본들은 중동부 지역인 안드라프라데시(Andhra Pradesh)에서 수집되었다. 파키스탄과 파슈툰(원본 출처 이름 보존, 파슈툰(Pashtun)일 가능성 높음)은 파키스탄 출신이며, 카슈미르인은 카슈미르 출신이다. 더 동쪽으로는 근동을 대표하는 이란인, 이라크인, 터키인 집단과 이라크 출신의 쿠르드인이 있다. 튀니지인 또한 이슬람의 발생 이후 비슷한 아랍(Arab) 또는 근동의 유전자 흐름을 흡수했을 가능성이 있는 관련 북아프리카 인구로서 포함되었다
In East and Central Asia, the Kazakhs, KirghizLL (lowland), KirghizHL (highland), Uighurs, UighursXJ (Xinjiang), and KazakhXJ (Xinjiang) represent some of the diversity seen in Central Asia and Xinjiang today. The Mongolians, Ewenki, Koreans, and Japanese along with the northern Chinese populations of Liaoning, Qidu, and Shandong represent the northern part of East Asia save Siberia. The Yunnan, Guangdong, and Guangdong2 populations represent the southern part of China. The Vietnamese are included to have some comparison for Southeast Asia, though according to the original reference they are actually “first generation immigrants” to California. The Akha are tribal people from Thailand, who along with the Indonesians and Vietnamese should represent Southeast Asia. The Wuhan province is fairly centrally located in China, while the Xinjiang Han should represent ethnic Han Chinese in the far western reaches of China.
동/중앙아시아에서 카자흐, 키르기스LL(저지대), 키르기스HL(고지대), 위구르, 위구르XJ(신강(新疆)), 카자흐XJ(신강(新疆))는 오늘날 중앙아시아와 신강(新疆)에서 볼 수 있는 다양성의 일부를 대표한다. 요녕(遼寧), 기도(齊都), 산동(山東)의 중국 북부 인구와 함께 몽골인, 에벤키, 한국인, 일본인은 시베리아를 제외한 동아시아 북부를 대표한다. 운남(雲南), 광동(廣東), 광동2(廣東2) 인구는 중국 남부를 대표한다. 베트남인은 동남아시아와 비교하기 위해 포함되었지만, 원본 문헌에 따르면 이들은 실제로 캘리포니아(California)로 온 “1세대 이민자”이다. 아카는 태국(Thailand)의 부족민으로, 인도네시아인 및 베트남인과 함께 동남아시아를 대표해야 한다. 무한(武漢)성은 중국의 꽤 중앙에 위치하며, 신강 한족(新疆 漢族)은 중국 극서부 지역의 한족(漢族)을 대표해야 한다.
Some population sequence data was collected from GenBank (Benson et al. 2000) and HVRBase (Handt et al. 1998). Other population sequence data was generated from the literature or from a data table provided by Toomas Kivisild and Mait Metspalu (personal communication 2003), using a program specifically written for large number sequence creation (from lists of nucleotide differences) by one of the authors (CB). Sequences were either initially aligned using the Sequencher program (Gene Codes Corp.), or were automatically set up to be aligned if created using the aforementioned program. Any sequence format conversion was handled by a program written for large number sequence conversion by one of the authors (CB). After initial alignment and creation, sequences were imported into MacClade (Maddison and Maddison 1989) for final alignment and editing purposes. All of the sequences were edited to extend from nt 16001-16497, either by cutting longer sequences or inserting “n’s” in shorter sequences. There are a few quick notes that should be made. The Basque population was downloaded and discovered to contain the same sample label for several samples. Because it was unclear whether these represented the original 45 individuals from the study with some identical sequences or sequences from 27 individuals with duplicates, duplicates were edited out. Thus the Basque sample may represent all of the variability taken from the original study, but not the proper frequencies. As mentioned above, they are included on the data table but their population values are not, and they were not included in total calculations or averages. Also, several sequences had to be removed from both the Slovakian and Rom populations taken from GenBank, because it was not clear what they represented (certainly not the HVS I, perhaps the HVS II?). Also, the Pushtoons were reduced to 360 base pairs long (nt 16024-16383) because of some confusion over primer lengths. The resulting sequence set went through a final round of alignment by hand editing. Note that any inserted deletions in the cytosine tract were removed by consensus (since this was the standard in 51 out of 53 of the populations originally). It was assumed that it is not clear whether the alterations in the cytosine tract of some sequences are deletions and extensions of the cytosines, or transversions of bases in the preceding poly-A segment to cytosines. This was only done in the Korean sample, since all other sample sets were apparently aligned this way originally. Also note that there was an insertion of a deletion in most sequences at the end of the cytosine tract to account for an extension of the cytosine tract in some individuals by one (making a total of 15 bases in the cytosine tract and preceding poly-A segment rather than 14), which was found in one Lobana, one Vietnamese, one Egyin, and several Icelanders. Along with another insertion at 16104 (16104a), one in a later polycytosine segment (16262a) and the aforementioned insertion in the cytosine tracts (16194a?), the final sequences including all unknowns and gaps were 500 nucleotides long.
일부 인구 집단 염기서열 데이터는 젠뱅크(GenBank)(Benson et al. 2000)와 HVRBase(Handt et al. 1998)에서 수집했다. 다른 인구 집단 염기서열 데이터는 문헌에서 얻거나 투마스 키비실드(Toomas Kivisild)와 마이트 메츠팔루(Mait Metspalu)가 제공한 데이터 표(2003년 개인적 연락)에서 생성했다. 이때 저자 중 한 명(CB)이 뉴클레오타이드 차이 목록을 바탕으로 대량의 염기서열 생성을 위해 특별히 작성한 프로그램을 사용했다. 염기서열은 시퀀처(Sequencher) 프로그램(Gene Codes Corp.)을 사용해 초기에 정렬했다. 또는 앞서 언급한 프로그램으로 생성된 경우 자동으로 정렬되도록 설정했다. 모든 염기서열 형식 변환은 저자 중 한 명(CB)이 대량 염기서열 변환을 위해 작성한 프로그램으로 처리했다. 초기 정렬 및 생성 후, 최종 정렬과 편집을 위해 염기서열을 맥클레이드(MacClade)(Maddison and Maddison 1989)로 가져왔다. 모든 염기서열은 긴 염기서열을 자르거나 짧은 염기서열에 “n”을 삽입하는 방식을 썼다. 그래서 nt 16001부터 16497까지 확장되도록 편집했다. 몇 가지 간단히 짚고 넘어갈 점이 있다. 바스크(Basque)인 데이터를 다운로드한 결과, 여러 표본에 동일한 표본 라벨이 포함되어 있는 것을 발견했다. 이것이 동일한 염기서열을 가진 연구의 원래 45명을 나타내는지, 아니면 중복이 포함된 27명의 염기서열을 나타내는지 불분명했다. 그래서 중복을 편집하여 제거했다. 따라서 바스크인 표본은 원래 연구에서 가져온 모든 변이를 나타낼 수는 있지만 적절한 빈도를 나타내지는 않을 것이다. 위에서 언급했듯이 이들은 데이터 표에는 포함되지만 인구 집단 값에는 포함되지 않았다. 전체 계산이나 평균에도 포함되지 않았다. 또한, 젠뱅크(GenBank)에서 가져온 슬로바키아(Slovakian)인 및 롬(Rom) 인구 집단 모두에서 여러 염기서열을 제거해야 했다. 이들이 무엇을 나타내는지 불분명했기 때문이다(확실히 HVS I은 아니며 아마도 HVS II일 것이다). 파슈툰(Pushtoon)인은 프라이머 길이에 대한 혼동 때문에 길이를 360 염기쌍(nt 16024-16383)으로 줄였다. 결과 염기서열 세트는 수동 편집을 통한 최종 정렬 과정을 거쳤다. 시토신 트랙트(cytosine tract)에 삽입된 결실(deletion)은 합의에 따라 제거했다. 원래 53개 인구 집단 중 51개에서 이것이 표준이었기 때문이다. 일부 염기서열의 시토신 트랙트 변형이 시토신의 결실 및 확장인지, 아니면 앞선 폴리-A(poly-A) 부분의 염기가 시토신으로 전환(transversion)된 것인지 명확하지 않다고 가정했다. 다른 모든 표본 세트가 원래 이 방식으로 정렬된 것으로 보였기 때문에 이는 한국인 표본에서만 수행했다. 또한, 일부 개체에서 시토신 트랙트가 1개 확장된 것을 설명하기 위해 대부분의 염기서열에서 시토신 트랙트 끝에 결실 삽입이 있었다. 이로써 시토신 트랙트와 앞선 폴리-A 부분의 총 염기 수는 14개가 아니라 15개가 되었다. 이는 로바나(Lobana)인 1명, 베트남(越南)인 1명, 에긴(Egyin)인 1명, 그리고 여러 아이슬란드(Iceland)인에서 발견되었다. 16104 위치의 또 다른 삽입(16104a), 나중의 폴리시토신 부분에서의 삽입(16262a), 그리고 시토신 트랙트의 앞서 언급한 삽입(16194a?)과 함께 모든 미확인 부분과 공백을 포함한 최종 염기서열의 길이는 500 뉴클레오타이드였다.
After the final round of alignment, the sequences were analyzed using DNAsp 3.99 (www.ub.es/dnasp; Rozas and Rozas 1999). This program was used to generate distance data for the various regional group tests and for the total tests. Separate tests were run with Linzi (and Qidu in the case of East Asia), since these sequences are only 185 bp, and the Egyin material (excluding Qidu in the case of East Asia), with sequences of 382 bp long. The Egyin material was included in the Linzi runs for comparison purposes. The analysis below on the effect of sequence length on this particular method, though, does suggest that it should be minimal. Each run included only the sites for which we had data across the entire dataset of that run. Due to the variation in the populations included, each run utilized slightly different sequence lengths, since various sequences and populations had differing amounts and locations of missing data. There was also a final third run with the Egyin population split to test the possibility suggested by Keyser-Tracqi et al. (2003) of a “Turkic” component in the Egyin sector C material, in which the Egyin material was divided up into EgyinAB (from sectors A and B) and EgyinC (from sector C) as described by Keyser-Tracqi et al. (2003) (results not shown). As previously mentioned, the Biaka population was included in every run as an outgroup.
최종 정렬 후, 염기서열은 DNAsp 3.99(www.ub.es/dnasp; Rozas and Rozas 1999)를 사용하여 분석했다. 이 프로그램은 다양한 지역 그룹 테스트와 전체 테스트를 위한 거리 데이터를 생성하는 데 사용했다. 염기서열 길이가 185bp에 불과한 임치(臨淄)인(동아시아의 경우 기도(齊都)인 포함)과, 염기서열 길이가 382bp인 에긴인 자료(동아시아의 경우 기도인 제외)로 개별 테스트를 실행했다. 비교를 위해 에긴인 자료를 임치인 실행에 포함시켰다. 그러나 이 특정 방법에서 염기서열 길이가 미치는 영향에 대한 아래 분석은 그 영향이 최소한일 것임을 시사한다. 각 실행에는 해당 실행의 전체 데이터 세트에 걸쳐 데이터가 있는 사이트만 포함했다. 포함된 인구 집단의 변동성으로 인해 염기서열과 인구 집단마다 누락된 데이터의 양과 위치가 달랐다. 그래서 각 실행은 약간씩 다른 염기서열 길이를 활용했다. 카이저-트라퀴(Keyser-Tracqi) 등(2003)이 제안한 에긴인 섹터 C 자료의 “튀르크계” 구성 요소 가능성을 테스트하기 위해 에긴인 인구를 나눈 세 번째 최종 실행도 있었다. 여기서 에긴인 자료는 카이저-트라퀴 등(2003)이 설명한 대로 에긴AB(섹터 A 및 B 출신)와 에긴C(섹터 C 출신)로 나누어졌다(결과는 표시되지 않음). 앞서 언급했듯이 비아카(Biaka) 인구 집단은 외부 집단으로서 모든 실행에 포함되었다.
Three distances matrices were generated for each run: Fst’s (Hudson et al. 1992), Nst’s (Lynch and Crease 1990), and Da’s (Nei 1987). The main analysis centered around the use of Fst’s to estimate distances (which generally reflected all of the distance measures generated, see below). Each run (Linzi, Egyin, and Egyin Split) was done for each region. After the regional analysis for each of the three runs, the top matches (those populations with the lowest Fst’s relative to the probe population of that run, Linzi or Egyin) were selected out of each region for a composite analysis. In the case of the first two runs, the top ten from East and Central Asia, the top nine from Europe, and the top eight from South and Southwest Asia were chosen. This was because of differences in the number of populations for each region (East and Central Asia with nineteen populations, Europe with sixteen, and South and Southwest Asia with thirteen). A different approach was used for the Egyin split runs, since to have the exact same set of populations for a single total run for comparative purposes, compromise sets of the top matches were done.
각 실행에 대해 세 가지 거리 행렬, 즉 Fst (Hudson et al. 1992), Nst (Lynch and Crease 1990), Da (Nei 1987)가 생성되었다. 주요 분석은 거리를 추정하기 위해 Fst의 사용을 중심으로 이루어졌다(이는 생성된 모든 거리 측정값을 일반적으로 반영함, 아래 참조). 각 실행(임치, 에긴, 에긴 분할)은 각 지역별로 진행했다. 세 번의 실행 각각에 대한 지역 분석 후, 각 지역에서 복합 분석을 위해 가장 일치하는 집단(해당 실행의 탐색 인구 집단인 임치 또는 에긴에 비해 Fst가 가장 낮은 인구 집단)을 선택했다. 처음 두 실행의 경우, 동/중앙아시아에서 상위 10개, 유럽에서 상위 9개, 남/서남아시아에서 상위 8개를 선택했다. 이는 각 지역의 인구 집단 수 차이 때문이었다(동/중앙아시아 19개, 유럽 16개, 남/서남아시아 13개). 비교 목적의 단일 전체 실행에 대해 완전히 동일한 인구 집단 세트를 갖기 위해 에긴 분할 실행에는 다른 접근 방식을 사용했다. 가장 일치하는 항목들의 타협 세트를 만들었다.
The total runs were completely recalculated, starting over by recalculating new distances for the new “global” (or more properly Eurasian) population sets using DNAsp. The same procedure was followed as above using multiple distance measures. All of the distance data tables below are generated from Fst’s. However, several tests were done with the other two measurements, and they were found to follow the same general relative order of population distances (data not shown). Also, tests were done excluding the Basques from the European sets (since the Basque have been noted to have discrepancies) and no change was found in the relative order of the results. Note however that this does not mean that the Basque distance measurements are accurate, merely that tests were done to see if their inclusion or exclusion caused error in the rest of the results. It should also be noted that the applicability of Fst measurements to population comparisons is a highly debated issue, as are the problems associated with appropriate Fst calculation (Nei 1977 and 1986, Weir and Cockerham 1984, Long and Kittles 2003).
DNAsp를 사용하여 새로운 “글로벌”(또는 더 정확하게는 유라시아) 인구 집단 세트에 대한 새로운 거리를 다시 계산하여 전체 실행을 완전히 다시 계산했다. 여러 거리 측정을 사용하여 위와 동일한 절차를 따랐다. 아래의 모든 거리 데이터 표는 Fst에서 생성되었다. 그러나 다른 두 측정값을 사용한 여러 테스트가 수행되었으며, 인구 거리의 동일한 일반적인 상대적 순서를 따르는 것으로 나타났다(데이터는 표시되지 않음). 또한, 유럽 세트에서 바스크인을 제외한 테스트를 수행했다(바스크인은 불일치가 있는 것으로 기록되었기 때문). 그 결과 결과의 상대적 순서에는 변화가 없는 것으로 나타났다. 그러나 이것은 바스크인 거리 측정값이 정확하다는 것을 의미하는 것은 아니다. 단지 그들의 포함 또는 제외가 나머지 결과에 오류를 일으키는지 확인하기 위해 테스트가 수행되었음을 의미할 뿐이다. 인구 집단 비교에 대한 Fst 측정의 적용 가능성은 적절한 Fst 계산과 관련된 문제(Nei 1977 and 1986, Weir and Cockerham 1984, Long and Kittles 2003)와 마찬가지로 논쟁이 많은 문제라는 점도 유의해야 한다.
원문 Haplogrouping was not done in this study. Whereas there is no estimate for the number of haplotypes in the total study (due to analytical problems associated with the large dataset), an estimate gathered from sequences just 217 bp in length and excluding all missing sites as well as Linzi and Qidu found 1084 haplotypes. If all the data were included, this number would likely increase. Although many of these haplotypes might cluster into haplogroups, analysis at this level would result in a loss of information. In addition, we lack sufficient sequence data for accurate haplogrouping in many cases (e.g. haplogrouping from 185 bp sequences with no restriction site data is difficult) and variability in sequence length would affect our results. Some information on haplogroups in the ancient populations is available in the original articles (Wang et al. 2000, Keyser-Tracqui et al. 2003), as well as Yao et al. 2003 (see discussion section below).
이 연구에서는 하플로그룹(haplogroup) 분류를 하지 않았다. 대규모 데이터 세트와 관련된 분석 문제 때문에 전체 연구의 하플로타입(haplotype) 수에 대한 추정치는 없다. 그러나 임치(臨淄)와 기도(齊都)를 제외하고, 모든 누락된 부위를 뺀 길이 217bp의 염기서열에서 1,084개의 하플로타입을 추정했다. 모든 데이터를 포함하면 이 숫자는 늘어날 가능성이 크다. 이 하플로타입 중 상당수는 하플로그룹으로 묶일 수 있다. 하지만 이 수준에서 분석하면 정보가 손실될 것이다. 게다가 많은 경우 정확한 하플로그룹 분류를 위한 염기서열 데이터가 부족하다. 예를 들어 제한효소 자리 데이터 없이 185bp 염기서열에서 하플로그룹을 분류하는 것은 어렵다. 또한 염기서열 길이의 변동성이 결과에 영향을 미칠 수 있다. 고대 인구 집단의 하플로그룹에 대한 몇 가지 정보는 원본 논문(Wang et al. 2000, Keyser-Tracqui et al. 2003)과 야오(Yao) 등(2003)에 나와 있다(아래 논의 섹션 참조).
결과 Results
This section will be broken up into several components, looking at the two initial runs by region and total followed by the Egyin split run. Note however that since the top matches for Egyin and Linzi were generated from their respective runs, the populations in their total population comparisons vary to some degree. Also, these data represent only part of the full matrix, and thus relationships between modern populations should not be inferred. The Egyin test is also one population shorter at the regional level since the Linzi data was not included in its run for reasons explained above. All distance data tables are sorted in order of lowest (closest) to highest (farthest) Fst values.
이 섹션은 몇 가지 구성 요소로 나뉜다. 지역별 및 전체별 두 가지 초기 실행을 살펴본 다음 에긴(Egyin) 분할 실행을 살펴본다. 그러나 에긴과 임치(臨淄)에 대한 상위 일치 항목이 각각의 실행에서 생성되었음을 유의해야 한다. 따라서 전체 인구 집단 비교에 포함된 집단은 어느 정도 차이가 있다. 또한, 이 데이터는 전체 행렬의 일부만 나타낸다. 따라서 현대 인구 집단 간의 관계를 이로써 추론해서는 안 된다. 에긴 테스트는 위에서 설명한 이유로 임치 데이터가 실행에 포함되지 않았다. 그래서 지역 수준에서 인구 집단이 하나 더 적다. 모든 거리 데이터 표는 가장 낮은(가장 가까운) Fst 값부터 가장 높은(가장 먼) Fst 값 순서로 정렬되었다.
The Linzi and Egyin data were compared to European populations. The results from these runs are shown in Table 2. As can be seen in the data tables, consistent estimates across different runs and population sets are noted. However, the order of the relationships should be taken generally. For instance, the details of whether the Armenians are really .0048 closer to the Linzi samples than the Catalans are not really necessary for this study. What is important is to see the general relative order; that certain populations are near the top, others are in the middle, and still others are at the bottom.
임치(臨淄)와 에긴(Egyin) 데이터는 유럽 인구 집단과 비교되었다. 이 실행 결과는 표 2에 나와 있다. 데이터 표에서 볼 수 있듯이 다양한 실행 및 인구 집단 세트에 걸쳐 일관된 추정치가 관찰된다. 하지만 관계의 순서는 대략적으로 이해해야 한다. 예를 들어 아르메니아인이 카탈루냐인보다 실제로 임치 표본에 .0048 더 가까운지 여부의 세부 사항은 이 연구에서 그다지 필요하지 않다. 중요한 것은 대략적인 상대적 순서를 확인하는 것이다. 특정 인구 집단이 상단에 있고, 다른 집단은 중간에 있으며, 또 다른 집단은 하단에 있다는 점이다.

표 2. 유럽 지역 FST 비교
We can see several interesting things in Table 2. First, the Linzi material seems to be closer to the European populations than to the Egyin individuals, except for the RomS sample. However, the calculated distances between ancient populations seem to be consistently greater than those between modern populations or between modern and ancient populations, perhaps reflecting increases in modern population sizes and/or gene flow (as to the effects of effective population size on Fst’s, see Cavalli-Sforza et al. 1994), though this would not necessarily affect the accuracy of the calculations themselves (Holsinger and Mason-Gamer 1996). Another possible explanation is that there was nonrandom fission (such as along familial or clan lines) in ancient populations in contrast to the larger social units of modern populations which are less dependent on familial relationships (Smouse et al. 1981, Whitlock 1994). As to the RomS population, it includes only 16 individuals who cluster into only 2 haplotypes, which might have resulted in inaccurate associations. Either way, the other Slovakian Rom population (Rom2S) is not in the top ten for either Linzi or Egyin. We can also see in this data table that the Egyin do have some affinity to the Bulgarian Rom and the Moroccans, however this is only relative within the European populations, as we will see below. We note that the Icelanders are near the top of the Linzi list (as Wang et al. 2000 suggested), but several populations are closer, including the Hungarians at the top.
우리는 표 2에서 몇 가지 흥미로운 점을 볼 수 있다. 첫째, 롬S(RomS) 표본을 제외하면 임치(臨淄) 자료는 에긴(Egyin) 개체들보다 유럽 인구 집단에 더 가까워 보인다. 그러나 고대 인구 집단 간의 계산된 거리는 현대 인구 집단 간 또는 현대와 고대 인구 집단 간의 거리보다 일관되게 더 큰 것으로 보인다. 이는 아마도 현대 인구 규모의 증가나 유전자 흐름을 반영하는 것일 수 있다(유효 인구 규모가 Fst에 미치는 영향에 대해서는 Cavalli-Sforza et al. 1994 참조). 하지만 이것이 반드시 계산 자체의 정확성에 영향을 미치는 것은 아니다(Holsinger and Mason-Gamer 1996). 또 다른 가능한 설명은 고대 인구 집단에 비무작위적 분열(가족이나 씨족 계통 등)이 있었다는 것이다. 이는 친족 관계에 덜 의존하는 현대 인구 집단의 더 큰 사회 단위와 대조된다(Smouse et al. 1981, Whitlock 1994). 롬S 인구 집단의 경우, 단지 2개의 하플로타입으로 묶이는 16명의 개체만 포함되어 있다. 그래서 연관성이 부정확하게 나타났을 수 있다. 어쨌든, 다른 슬로바키아(Slovakian) 롬 인구 집단(롬2S)은 임치나 에긴의 상위 10위 안에 들지 않는다. 우리는 이 데이터 표에서 에긴이 불가리아(Bulgarian) 롬 및 모로코인과 어느 정도 연관성이 있음을 볼 수 있다. 그러나 아래에서 보게 되겠지만 이는 유럽 인구 집단 내에서 상대적일 뿐이다. 왕(Wang) 등(2000)이 제안한 대로 아이슬란드인이 임치 목록의 상단에 있음이 확인된다. 하지만 맨 위에 있는 헝가리인을 포함하여 몇몇 인구 집단이 임치에 더 가깝다.
Table 3 contains the comparisons of the Linzi and Egyin samples to the South and Southwest Asians. There are several interesting results. First, we once again see that the Linzi individuals bear a closer affinity with most of the modern populations than with Egyin individuals except of course for the sub-Saharan African Biaka (though see above). Secondly, there is a definite difference between the two ancient populations in the ordering of this table. The Egyin list has the populations of India mostly at the top (save maybe for the Uttar Pradesh, which is still in the top half) with the Pakistani populations and the Tunisians (which upon further review were found to bear relatively closer Fst’s with the Pakistani populations than with the populations of the Near East, possibly reflecting the Arab expansions). The populations of the Near East (Iranians, Iraqis, Kurds, and Turks) are all at the bottom. In the Linzi list, we see the opposite trend, with the populations of the Near East mainly at the top. The populations of Pakistan and Tunisia are mixed in the middle, while the Indian populations are all at the bottom. The top matches in this list seem to be the Iranians and Turks of Turkey. As to why the Egyin material bears a close affinity with the Indian populations, relatively speaking, in this list, this may have something to do with some degree of shared maternal heritage in South and East Asians dating back to the earliest settlements of South and East Asia, though this shared heritage is probably limited due to subsequent divergence in modern day populations (Kivisild et al. 2003). As we will see below, the Indians are probably just the closest matches from this region, not necessarily close overall. Further, this may relate to the affinity of Egyin and Linzi to East Asians more than to their direct affinity to Indians.
표 3은 임치(臨淄)와 에긴(Egyin) 표본을 남아시아 및 서남아시아인과 비교한 내용을 담고 있다. 몇 가지 흥미로운 결과가 있다. 첫째, 임치 개체들이 에긴 개체들보다 대부분의 현대 인구 집단과 더 가까운 연관성을 보인다는 것을 다시 한 번 알 수 있다. 물론 사하라(Sahara) 이남의 아프리카 비아카(Biaka)족은 예외다(위 내용 참조). 둘째, 이 표의 순서에서 두 고대 인구 집단 사이에 명확한 차이가 있다. 에긴 목록에서는 인도 인구 집단이 파키스탄 인구 집단 및 튀니지인과 함께 대부분 상단에 위치한다(여전히 상위 절반에 있는 우타르프라데시는 예외일 수 있다). 튀니지인은 추가 검토 결과 근동 인구 집단보다 파키스탄 인구 집단과 상대적으로 더 가까운 Fst를 갖는 것으로 나타났는데, 이는 아랍(Arab)의 팽창을 반영할 가능성이 있다. 근동의 인구 집단(이란인, 이라크인, 쿠르드인, 터키인)은 모두 하단에 있다. 임치 목록에서는 정반대의 경향이 나타난다. 근동 인구 집단이 주로 상단에 있다. 파키스탄과 튀니지의 인구 집단은 중간에 섞여 있다. 반면 인도 인구 집단은 모두 하단에 있다. 이 목록에서 가장 일치하는 집단은 이란인과 터키(Turkey)의 터키인으로 보인다. 이 목록에서 상대적으로 에긴(Egyin) 자료가 인도 인구 집단과 밀접한 연관성을 보이는 이유는 무엇일까? 이는 남아시아 및 동아시아의 초기 정착 시기로 거슬러 올라가는 모계 유산의 일부 공유와 관련이 있을 수 있다. 하지만 이 공유된 유산은 현대 인구 집단의 후속 분기 때문에 아마도 제한적일 것이다(Kivisild et al. 2003). 아래에서 보게 되겠지만, 인도인은 이 지역에서 가장 일치하는 집단일 뿐 전체적으로 가깝다는 의미는 아닐 수 있다. 게다가, 이는 인도인에 대한 직접적인 연관성보다 에긴과 임치(臨淄)가 동아시아인과 갖는 연관성과 관련이 있을 수 있다.

표 3. 남아시아 및 서남아시아 지역 FST 비교
Table 4 contains the comparisons of Linzi and Egyin to East and Central Asia. Note here that Qidu was also removed from the Egyin run because it is only 185 bp long. The data table shows a differential clustering of Linzi and Egyin with these populations. Once again, we see that the Linzi material is actually closer to the modern Asian groups in this study than to Egyin (though see above). The top half of the Linzi list is dominated by Southeast Asians, southern Chinese, and Central Asians. The lower half is dominated by northern Asians (save for the lowland Kirghiz and Akha). As to how and why both Southeast Asians (and southern Chinese) and Central Asians are similar to the Linzi population, it is not clear, though this is only a relative comparison within this region. However, there is some debate over the nature of ancient East Asian genetic history, so possibly there are issues here that have yet to be illuminated (Yao et al. 2002b, 2002c, 2003, Oota et al. 2002). Further, there are some issues with the Vietnamese sequences (see below). Also, it should be noted that the modern Qidu samples, from the same general locale as ancient Linzi, were in the lower half of the table. The Egyin list shows the opposite trend. The top half of the Egyin list is dominated by northern East Asians (including northern Chinese) except for the Xinjiang Han (however, a closer analysis of the Xinjiang Han shows them to have a genetic affinity to both Central Asians and northern Chinese and Mongolians). The lower half of the list is dominated by Central Asians and Southeast Asians (save for the lowland Kirghiz, which appropriately do an exact reversal from the Linzi list by showing up as the closest Central Asian population), with the Southeast Asians mainly at the very bottom except for Guangdong2. The Wuhan sample from central China seems to float about in the middle of both lists.
표 4는 임치(臨淄)와 에긴(Egyin)을 동아시아 및 중앙아시아와 비교한 내용을 담고 있다. 여기서 기도(齊都)는 길이가 185bp에 불과하기 때문에 에긴 실행에서 제거되었다는 점에 유의해야 한다. 데이터 표는 이 인구 집단들과 임치 및 에긴의 차별적인 군집화를 보여준다. 다시 한 번, 우리는 임치 자료가 에긴보다 이 연구의 현대 아시아 그룹에 더 가깝다는 것을 알 수 있다(위 내용 참조). 임치 목록의 상위 절반은 동남아시아인, 중국 남부인, 중앙아시아인이 주로 차지한다. 하위 절반은 북아시아인이 주로 차지한다(저지대 키르기스인과 아카는 제외). 동남아시아인(그리고 중국 남부인)과 중앙아시아인이 어떻게, 그리고 왜 임치 인구 집단과 유사한지는 명확하지 않다. 비록 이것이 이 지역 내에서의 상대적인 비교일 뿐이지만 말이다. 그러나 고대 동아시아의 유전적 역사의 본질에 대해서는 논쟁이 있다. 따라서 아직 밝혀지지 않은 문제가 있을 가능성이 있다(Yao et al. 2002b, 2002c, 2003, Oota et al. 2002). 게다가, 베트남(越南)인 염기서열에 몇 가지 문제가 있다(아래 참조). 또한, 고대 임치와 같은 대략적인 위치에서 나온 현대 기도 표본이 표의 하위 절반에 있었다는 점도 주목해야 한다. 에긴 목록은 정반대의 경향을 보여준다. 에긴 목록의 상위 절반은 신강 한족(新疆 漢族)을 제외한 동아시아 북부인(중국 북부인 포함)이 주로 차지한다. (그러나 신강 한족을 면밀히 분석해보면 그들이 중앙아시아인과 중국 북부인, 그리고 몽골인 모두와 유전적 연관성이 있음을 알 수 있다). 목록의 하위 절반은 중앙아시아인과 동남아시아인이 주로 차지한다. (가장 가까운 중앙아시아 인구 집단으로 나타나 임치 목록과 정확히 역전된 결과를 보여주는 저지대 키르기스인은 제외한다) . 광동2(廣東2)를 제외하고 동남아시아인은 주로 맨 아래에 있다. 중국 중부 출신의 무한(武漢) 표본은 두 목록의 중간쯤에 떠 있는 것으로 보인다.

표 4. 동아시아 및 중앙아시아 지역 FST 비교
Table 5 is the list for the total comparison of modern populations to Linzi and Egyin from a composite of the top matches from each region as explained in the methods section. Once again, it should be noted that the group of populations for each list is somewhat different and was generated independently from separate runs of Linzi and Egyin data in regional models.
표 5는 방법 섹션에서 설명한 대로 각 지역의 상위 일치 항목들을 조합하여, 현대 인구 집단을 임치(臨淄) 및 에긴(Egyin)과 종합적으로 비교한 목록이다. 각 목록의 인구 집단 그룹은 다소 다르며 지역 모델에서 임치와 에긴 데이터를 각각 실행하여 독립적으로 생성되었음을 다시 한 번 유의해야 한다.

표 5. 전체 FST 비교
The table clearly shows a differential pattern in the genetic relationships of Linzi and Egyin to other populations. First, we see that the Egyin and Linzi populations did not share a close affinity with each other, or at least not more so than they do with modern populations (though see above). As for the Linzi individuals, they seem to be most highly related to Near Easterners (Turks, Iranians, and Iraqis), Armenians, and eastern Europeans (Slavs, Hungarians), though others such as Catalans and Iraqis are mixed in. The Icelanders are twelfth on this list. The high placement of the Vietnamese may be an anomaly, error, or some element of ancient genetic history that is not clear (though see Yao et al. 2003). However, it should be noted that the Vietnamese sequences lack a section of bases near the cytosine tract, compounded by the large number of sequences compared in this population set, which could provide for some anomalous results. Furthermore, this approach cannot accurately account for significant admixture (a distinct possibility given the proposed haplotypes of some of the individuals at Linzi, see Yao et al. 2003 and the discussion section of this paper), though neither of the other Southeast Asian populations (the Akha from Thailand or the Indonesians) even made it into the composite run. Thus, given these issues, it should be reiterated that only general trends should be drawn from this study.
표는 임치 및 에긴과 다른 인구 집단 간의 유전적 관계에서 차별적인 패턴을 분명히 보여준다. 첫째, 에긴과 임치 인구 집단은 서로 밀접한 연관성을 공유하지 않았다. 적어도 현대 인구 집단과 공유하는 것 이상은 아니었다(위 내용 참조). 임치 개체들의 경우, 근동인(터키인, 이란인, 이라크인), 아르메니아인, 그리고 동유럽인(슬라브인, 헝가리인)과 가장 연관성이 높은 것으로 보인다. 비록 카탈루냐인과 이라크인 같은 다른 집단들도 섞여 있긴 하지만 말이다. 아이슬란드인은 이 목록에서 12위다. 베트남(越南)인의 높은 순위는 변칙, 오류, 또는 명확하지 않은 고대 유전적 역사의 일부 요소일 수 있다(Yao et al. 2003 참조). 그러나 베트남인 염기서열은 시토신 트랙트 근처의 염기 구간이 누락되어 있다는 점에 유의해야 한다. 이 인구 집단 세트에서 비교된 염기서열의 수가 많다는 점이 겹쳐져, 몇 가지 변칙적인 결과를 제공했을 수 있다. 게다가, 이 접근 방식은 상당한 혼혈을 정확하게 설명할 수 없다. 임치의 일부 개체에 대해 제안된 하플로타입을 고려할 때 이는 분명한 가능성이다(Yao et al. 2003 및 이 논문의 논의 섹션 참조). 다른 동남아시아 인구 집단(태국(Thailand)의 아카 또는 인도네시아인) 중 어느 것도 종합 실행에 포함되지 않았지만 말이다. 따라서 이러한 문제들을 고려할 때, 이 연구에서는 일반적인 경향만을 도출해야 함을 반복해서 강조한다.
What is clear is that the Linzi material does have an affinity to the west, most highly to the groups mentioned above. The East Asians that made the list are generally toward the bottom, save for the Vietnamese. The other interesting thing is that the few Central Asian Turkic peoples are generally toward the bottom, with only the Uighur appearing in the middle of the top half (but still outside the top ten). It has been noted that Near Eastern Turks actually bear more affinity with Europeans and Near Easterners than with their linguistic cousins in Central Asia, and that the Turks came to dominate Turkey through an elite dominance process, meaning that the effect on the maternal heritage should be minimal (Comas et al. 1996, 1998). Thus we may be able to include them together with the Iranians and other Near Easterners, who bear a close affinity with Linzi, though the relatively high distance between the ancient Linzi sample and Central Asian Turks may actually be from more recent East Asian admixture. The other high affinity groups, mostly from Eastern Europe in the Slovakians and Hungarians, may be related either directly or through the indirect process of East-West settlement in Central Eurasia that has been occurring in Eastern Europe for at least the past several thousand years, beginning possibly with the Indo-Europeans and definitely by the time of the Iranian Scythians and Sarmatians, as well as with later Turkic groups (though we have noted the distance between modern Central Asian Turkic peoples and Linzi).
분명한 것은 임치(臨淄) 자료가 서양과 연관성이 있으며, 위에서 언급한 그룹들과 가장 높은 연관성을 갖는다는 것이다. 목록에 포함된 동아시아인은 베트남(越南)인을 제외하고 대체로 하단에 있다. 또 다른 흥미로운 점은 소수의 중앙아시아 튀르크계 민족이 대체로 하단에 있다는 것이다. 오직 위구르(Uighur)인만이 상위 절반의 중간쯤에 나타난다(그러나 여전히 상위 10위권 밖이다). 근동 터키인들은 실제로 중앙아시아의 언어적 사촌들보다 유럽인 및 근동인과 더 큰 연관성을 보인다는 점이 지적되어 왔다. 터키인들이 엘리트 지배 과정을 통해 터키(Turkey)를 지배하게 되었다는 것은, 모계 유산에 미치는 영향이 최소화되어야 함을 의미한다(Comas et al. 1996, 1998). 따라서 우리는 이들을 임치와 밀접한 연관성을 보이는 이란인 및 다른 근동인과 함께 묶을 수 있다. 물론 고대 임치 표본과 중앙아시아 튀르크인 사이의 상대적으로 높은 거리는 실제로는 더 최근의 동아시아 혼혈에서 비롯된 것일 수도 있다. 슬로바키아인과 헝가리인 등 주로 동유럽 출신인 다른 연관성 높은 그룹들은 직접 연관되었을 수도 있고 간접적인 과정을 통해 연관되었을 수도 있다. 이 간접적인 과정은 적어도 지난 수천 년 동안 동유럽에서 발생한 중앙 유라시아의 동서 정착 과정이다. 아마도 인도유럽인부터 시작되어 이란계 스키타이(Scythian)인과 사르마티아(Sarmatian)인 시대, 그리고 그 이후의 튀르크계 그룹에 이르기까지 일어났을 것이다. (비록 현대 중앙아시아 튀르크계 민족과 임치 사이의 거리에 대해서는 언급했지만 말이다) .
The Egyin list is, of course, from a much longer sequence comparison, thus increasing the probability of valid connections. Other than some reordering, the top matches are all the same ones from the East Asia regional table (Table 4). The middle of the list includes all of the South and Southwest Asians in the same general order as found in the regional comparison, while the bottom of the list includes all of the European populations, in a similar order. Note that there are significant gaps in the fall of the Fst’s between the Europeans and the South and Southwest Asians and between the South and Southwest Asians and East Asians (except for the Lambadi and the RomB, who seem to float in between) of about .02 to 03, whereas no other populations within regions (except for the Lambadi and the RomB of course) exhibit gaps of even .01. However, note that this is not indicative of the difference between these groups of populations, but between these populations and Egyin. Further, it should be noted that this list does not specify exactly where non-included populations from the regional comparisons would fall. What is clear here is that the Egyin population seems to firmly relate to East Asian populations, particularly the northern East Asian populations (northern Chinese, Inner Mongolian Ewenki, Mongolians, Koreans, Japanese, and the Xinjiang Han whose partial northern East affinities were explained above). Whatever the exact interpretation of the genetic affinities of the Egyin and Linzi populations may be, it is clear that they differ significantly.
물론 에긴(Egyin) 목록은 훨씬 더 긴 염기서열 비교에서 나온 것이므로, 유효한 연관성 확률이 높아진다. 약간의 순서 변경을 제외하면, 상위 일치 항목들은 동아시아 지역 표(표 4)와 모두 동일하다. 목록의 중간에는 남/서남아시아인 전체가 지역 비교에서 발견된 것과 동일한 대략적인 순서로 포함되어 있다. 반면 목록의 하단에는 모든 유럽 인구 집단이 유사한 순서로 포함되어 있다. 유럽인과 남/서남아시아인 사이, 그리고 남/서남아시아인과 동아시아인 사이(그 사이에 떠 있는 것으로 보이는 람바디와 롬B는 제외)의 Fst 하락에 약 .02에서 .03의 상당한 격차가 있다는 점에 유의해야 한다. 반면 지역 내 다른 인구 집단(물론 람바디와 롬B는 제외)은 .01의 격차도 보이지 않는다. 그러나 이것이 이 인구 집단 그룹들 사이의 차이를 나타내는 것이 아니라, 이 인구 집단들과 에긴 사이의 차이를 나타낸다는 점에 주의해야 한다. 게다가, 이 목록은 지역 비교에서 포함되지 않은 인구 집단이 정확히 어디에 속할지 명시하지 않는다는 점도 유의해야 한다. 여기서 분명한 것은 에긴 인구 집단이 동아시아 인구 집단, 특히 동아시아 북부 인구 집단(중국 북부인, 내몽골(內蒙古)의 에벤키, 몽골인, 한국인, 일본인, 그리고 위에서 부분적인 동북아시아 연관성이 설명된 신강 한족(新疆 漢族))과 확고하게 연관되어 있는 것으로 보인다는 점이다. 에긴과 임치(臨淄) 인구 집단의 유전적 연관성에 대한 정확한 해석이 무엇이든 간에, 이 둘이 상당히 다르다는 것은 분명하다.
A further test was undertaken in this study to examine the suggestion of Keyser-Tracqui et al. (2003) about the possible differences of sectors A and B compared to sector C at the Egyin Gol site (with sector C showing “Turkic” affinities). It is questionable whether it is reasonable to pursue such an examination from a methodological standpoint. The division of the population here into subpopulations is based on observed variation in spatial and temporal factors at the necropolis, as well as some putative differences in genetics. However, the differences in these subpopulations based on these factors may be only superficial differences, and the division thus arbitrary in nature. Further, the differences in genetics may not be reliable with two subpopulations of size 38 (EgyinAB) and 8 (EgyinC). Despite this, the test was done tentatively to see if there were any significant differences in affinity to modern populations or to each other. The final results for both supposed subsamples generally followed along the lines of the total Egyin population analysis, and there were no clear distinctions between the two groups (though there were a few slight differences). However, given the methodological issues, this part of the study is not included in this paper, and the results are not shown.
이 연구에서는 에긴골(Egyin Gol) 유적지에서 섹터 C(섹터 C는 “튀르크계” 연관성을 보임)와 비교하여 섹터 A 및 B의 가능한 차이에 대한 카이저-트라퀴(Keyser-Tracqui) 등(2003)의 제안을 검토하기 위해 추가 테스트를 수행했다. 방법론적 관점에서 이러한 검토를 추진하는 것이 합리적인지는 의문이다. 여기서 인구를 하위 집단으로 나누는 것은 묘지에서 관찰된 공간적, 시간적 요인의 변이와 유전학에서의 추정되는 차이에 근거한다. 그러나 이러한 요인들에 기초한 이 하위 집단들의 차이는 표면적인 차이일 수 있으며, 따라서 이러한 분할은 본질적으로 임의적일 수 있다. 게다가, 크기가 38(에긴AB)과 8(에긴C)인 두 하위 집단으로는 유전학의 차이를 신뢰할 수 없을 수 있다. 그럼에도 불구하고, 현대 인구 집단에 대한 연관성이나 두 집단 사이의 연관성에 의미 있는 차이가 있는지 확인하기 위해 시험적으로 테스트를 수행했다. 가상의 두 하위 표본에 대한 최종 결과는 일반적으로 전체 에긴 인구 집단 분석과 같은 방향을 따랐다. 두 그룹 사이에 뚜렷한 차이는 없었다(약간의 차이는 있었지만). 그러나 방법론적 문제를 고려할 때, 이 연구의 해당 부분은 이 논문에 포함되지 않았으며 결과도 표시되지 않았다.
We also examined the suggestion by Yao et al. 2003 that the results from Linzi were due purely to the short sequence length. For this, we took the Fst distance data from the Egyin total run for Egyin (318 bp) and compared it to the calculate Fst values when Linzi and Qidu were added (166 bp, referred to as Egyin Limited or EgLim, see table 6). We also took the Japanese Fst data from the Linzi total run (with the Japanese and Liaoning added, 149 bp, referred to as Japanese Limited, or JpLim) and compared it the calculated Fst values when Linzi and the Vietnamese were removed (319 bp, table 7). Though there is some slight movement of populations in both cases, there is no major discrepancy in the relative ordering of the populations in the results for either table 6 (spearman’s rank-order correlation: .981, ρ=.000) or table 7 (spearman’s rank-order correlation: .987, ρ=.000) Both of the probe populations here (Japanese and Egyin) are East Asians, one modern and one ancient, and thus geographically similar to Linzi. Therefore, the argument that the Linzi data is skewed purely by short sequence length appears to be incorrect, though sequence length is still an issue.
우리는 임치(臨淄)의 결과가 순전히 짧은 염기서열 길이 때문이라는 야오(Yao) 등(2003)의 제안도 검토했다. 이를 위해 에긴(Egyin)의 전체 실행에서 나온 Fst 거리 데이터(318bp)를 가져와, 임치와 기도(齊都)가 추가되었을 때 계산된 Fst 값(166bp, 에긴 제한 또는 EgLim으로 지칭됨, 표 6 참조)과 비교했다. 또한, 임치 전체 실행에서 일본인 Fst 데이터를 가져와(일본인 및 요녕(遼寧) 추가, 149bp, 일본인 제한 또는 JpLim으로 지칭됨), 임치와 베트남(越南)인이 제거되었을 때 계산된 Fst 값(319bp, 표 7)과 비교했다. 두 경우 모두 인구 집단의 약간의 이동이 있지만, 표 6(스피어만(Spearman)의 순위 상관계수: .981, ρ=.000) 또는 표 7(스피어만의 순위 상관계수: .987, ρ=.000) 결과에서 인구 집단의 상대적 순서에 큰 차이는 없다. 여기 탐색 인구 집단(일본인 및 에긴)은 모두 동아시아인으로, 하나는 현대이고 하나는 고대이므로 지리적으로 임치와 유사하다. 따라서 임치 데이터가 순전히 짧은 염기서열 길이 때문에 왜곡되었다는 주장은 잘못된 것으로 보인다. 비록 염기서열 길이는 여전히 문제이긴 하지만 말이다.

표 6. 에긴골(Egyin Gol)과 에긴골 제한(EgLim) 비교

표 7. 일본인과 일본인 제한(JpLim) 비교
논의 Discussion
First, we should give a brief discussion of some of the problems with this study not previously mentioned. The first issue is that this study only deals with maternal heritage. Analysis of the Y-chromosome could reveal differences in populations not revealed here, because some populations have experienced differential population histories varying by sex (e.g. due to long-distance migrations or matrilocal vs. patrilocal mating practices). One such possible example of this is that of Iceland, where according to Helgason et al. (2001) the original population consisted mainly of men from Scandinavia and women from the British Isles.
먼저 이 연구의 문제점 중 이전에 언급하지 않은 몇 가지를 간단히 논의해야 한다. 첫 번째 문제는 이 연구가 모계 유산만을 다룬다는 것이다. Y-염색체 분석은 여기서 드러나지 않은 인구 집단의 차이를 밝혀낼 수 있다. 일부 인구 집단은 성별에 따라 다른 인구 역사를 겪었기 때문이다(예를 들어 장거리 이주 또는 모계 거주 대 부계 거주 혼인 관습). 이에 대한 가능한 예 중 하나가 아이슬란드(Iceland)이다. 헬가손(Helgason) 등(2001)에 따르면 아이슬란드의 원래 인구는 주로 스칸디나비아(Scandinavia) 출신 남성과 영국 제도(British Isles) 출신 여성으로 구성되었다.
Further problems arise in this study from the large variation in sample sizes and sequence lengths (see table 1). These issues were dealt with as best as possible, with multiple runs and separate testing for the probe samples (the ancient samples) to try to maximize sequence lengths. The analysis of the effect of sequence length on this particular method does demonstrate that it is minimal. Otherwise, issues with available samples such as missing data limit the analyses’ certainty in all cases, but hopefully using larger sample sizes and improving the quality of the source material can increase the accuracy of such genetic analysis. As far as sample size, for many populations all available data were used, and larger sample sizes will be available only with additional sample collection. Obviously, the results presented here should be taken generally, not as precise indicators of genetic affiliation.
이 연구에서 더 발생하는 문제는 표본 크기와 염기서열 길이의 큰 변동성이다(표 1 참조). 이러한 문제는 가능한 한 최선을 다해 처리했다. 탐색 표본(고대 표본)에 대해 여러 번 실행하고 개별적으로 테스트하여 염기서열 길이를 최대화하려고 노력했다. 이 특정 방법에 대한 염기서열 길이의 영향 분석은 그 영향이 최소한임을 입증한다. 그 외에 누락된 데이터와 같이 사용 가능한 표본의 문제는 모든 경우에 분석의 확실성을 제한한다. 하지만 더 큰 표본 크기를 사용하고 출처 자료의 질을 향상시켜 이러한 유전자 분석의 정확성을 높일 수 있기를 바란다. 표본 크기에 관해서는 많은 인구 집단에 대해 사용 가능한 모든 데이터를 활용했다. 추가적인 표본 수집을 통해서만 더 큰 표본 크기를 얻을 수 있을 것이다. 분명히 여기에 제시된 결과는 유전적 연관성의 정밀한 지표가 아니라 전반적인 경향으로 받아들여야 한다.
Before we conclude, we should briefly discuss some of the problems of population-level genetic analysis. To look at very large sample sizes, examination at the population level may be the most efficient method, given a sufficient availability of computing power. However, the examination of genetic variation at the population level has its own set of problems, beginning with the simple problem of the definition of a “population.” In addition, population similarity due to gene flow versus shared ancestry cannot be discerned with this method. Looking at possible paths of mutations, such as through haplogrouping, individual sequence-by-sequence comparison, or perhaps nested cladistic analysis within haplogroups (Templeton et al. 1995), might tease apart these issues. However, there are advantages in using direct nucleotide-to-nucleotide sequence comparisons between populations in large-scale studies, particularly with regard to higher resolution in the results (i.e. more detail to the distinctions), if technological and methodological issues can be overcome. Moreover, the above-mentioned methods can be employed in subsequent studies following a large-scale approach in order to further refine the results. Thus, we think that the methods and approach of this study are appropriate if interpreted with caution, particularly for the kind of large-scale population data examined in this case.
결론을 내리기 전에 인구 집단 수준 유전자 분석의 몇 가지 문제를 간단히 논의해야 한다. 컴퓨터 성능이 충분히 제공된다면, 매우 큰 표본 크기를 살펴보기 위해서는 인구 집단 수준에서 조사하는 것이 가장 효율적인 방법일 수 있다. 그러나 인구 집단 수준에서 유전적 변이를 조사하는 것은 “인구 집단”의 정의라는 단순한 문제에서 시작하여 고유한 문제들을 지닌다. 게다가, 유전자 흐름에 의한 인구 집단 유사성과 조상 공유에 의한 유사성을 이 방법으로는 구별할 수 없다. 하플로그룹 분류, 개별 염기서열 간 비교, 또는 하플로그룹 내의 중첩 분기군 분석(Templeton et al. 1995) 등을 통해 가능한 돌연변이 경로를 살펴보면 이러한 문제들을 구별해 낼 수 있을 것이다. 그러나 대규모 연구에서 인구 집단 간의 직접적인 뉴클레오타이드 대 뉴클레오타이드 염기서열 비교를 사용하는 데는 장점이 있다. 특히 기술적, 방법론적 문제를 극복할 수 있다면, 결과의 해상도를 높이는 데(즉, 차이를 더 자세히 파악하는 데) 유리하다. 더욱이 위에서 언급한 방법들은 대규모 접근 방식을 따른 후속 연구에서 결과를 더 정제하기 위해 사용할 수 있다. 따라서 주의해서 해석한다면 이 연구의 방법과 접근 방식은 적절하다고 생각한다. 특히 이 사례에서 조사한 것과 같은 대규모 인구 데이터에 대해서는 더욱 그러하다.
To conclude, genetic distances were estimated using a wide-angle lens, examining regional comparisons and extracting the best matches from each to create a total comparison (“free competition”). The goal was to eliminate the need to arbitrarily include or exclude populations from the overall comparison a priori, though obviously not every population in the world was included. However, it is felt by the authors that this method can produce more accurate comparisons than simply selecting populations based on preconceptions of population relationships or utilizing a single or small number of populations to represent whole regions. The problem with simply selecting arbitrary populations is highlighted by the original study on the Linzi material (Wang 2000), in which five random populations were chosen to represent Europe. This led to the incorrect attribution to the nearest relatives of the Linzi material being Icelander and Finnish (though they were possibly right about the Turkish comparison). No eastern Europeans or Iranians were included in that study, while these populations accompanied the Turkish as being closest to the Linzi material in this study, in which the Icelanders were actually not in the top ten. Obviously, there are still a number of “missing” populations in this analysis (such as the Tibetans, the Finns, the Russians, etc.). However, these methods are an improvement over arbitrary selection of populations for population comparison or the inclusion of only particular populations hypothesized to be related to the ancient population(s) based on linguistic or archaeological data. Of course, some regions of the world may not need such extensive analysis, but judging by the extremely variable populations and often bewildering history of Central Eurasia dating back at least to the late Paleolithic, it is quite applicable in this region.
결론적으로 넓은 시야를 통해 유전적 거리를 추정했다. 지역적 비교를 검토하고 각 지역에서 가장 잘 일치하는 항목을 추출하여 전체 비교(“자유 경쟁”)를 만들었다. 목표는 전체 비교에서 선험적으로 특정 인구 집단을 임의로 포함하거나 제외할 필요성을 없애는 것이었다. 물론 전 세계의 모든 인구 집단이 포함된 것은 아니다. 그러나 저자들은 이 방법이 더 정확한 비교를 만들어 낼 수 있다고 여긴다. 단순히 선입견에 기초하여 인구 집단을 선택하거나 단일 또는 소수의 인구 집단을 전체 지역의 대표로 사용하는 것보다 낫다는 것이다. 단순히 임의의 인구 집단을 선택하는 문제는 임치(臨淄) 자료에 대한 원래 연구(Wang 2000)에서 두드러진다. 해당 연구에서는 유럽을 대표하기 위해 5개의 무작위 인구 집단을 선택했다. 이것은 임치 자료의 가장 가까운 친척이 아이슬란드인과 핀란드(Finland)인이라는 잘못된 귀속으로 이어졌다(터키인 비교에 대해서는 맞았을 수도 있다). 그 연구에는 동유럽인이나 이란인이 포함되지 않았다. 반면 본 연구에서는 이 인구 집단들이 터키인과 함께 임치 자료에 가장 가까운 것으로 나타났다. 아이슬란드인은 실제로 상위 10위 안에 들지 않았다. 분명히 이 분석에는 여전히 “누락된” 인구 집단이 많이 있다(예: 티베트(Tibet)인, 핀란드인, 러시아(Russia)인 등). 그러나 이러한 방법은 인구 비교를 위해 임의로 인구 집단을 선택하는 것보다 개선된 것이다. 언어학적, 고고학적 데이터를 바탕으로 고대 인구 집단과 관련이 있다고 가정된 특정 인구 집단만 포함하는 것보다도 낫다. 물론 세계의 일부 지역에서는 그렇게 광범위한 분석이 필요하지 않을 수 있다. 하지만 적어도 후기 구석기 시대까지 거슬러 올라가는 중앙 유라시아의 극도로 다양한 인구와 종종 혼란스러운 역사를 고려할 때, 이 지역에는 꽤 적용할 만하다.
The results suggest that there are definite differences in the genetic affinities between the ancient populations of Linzi in northern China and Egyin Gol in Mongolia. The Linzi material seems to bear a stronger affinity with Near Easterners and Europeans rather than with the present day populations of northern China, though there is a definite component of East and/or Southeast Asians within Linzi as well (as evidenced by haplogrouping, see below). We would suggest that rather than a “European-like population” in the ancient Linzi region, the Linzi material may be at least partially related to Indo-Iranians (a branch of Indo-European, though more precisely just “Iranian” by this time period), who were, during that period or at least shortly before it, probably inhabiting areas across Central Eurasia. More precisely, the Linzi population was quite possibly related to the Karsuk or Saka (putative Iranian groups who fit temporally and spatially), or also more distantly to the Andronovo, Afanasievo, Scythians, Sarmatians, or even the Sogdians. The Karsuk and Saka are the most likely given their existence in the 1st millennium BC in the central and possibly eastern parts of Central Eurasia, though these ethonyms are a little ambiguous and precise connections are not really possible. However, Harmatta (1992) has argued that early Iranian groups were spread across Central Eurasia from Eastern Europe to north China in the 1st millennium BC, and Askarov et al. (1992) have pointed out the existence of cist kurgan burials (with “Europoid” remains bearing some Mongoloid admixture, they suggest) in northwestern Mongolia in the same millennium. Although speculative, this line of reasoning fits in with other lines of evidence from archaeology and linguistics for the aforementioned changes in Chinese Bronze Age culture, the loan words in Old Chinese (Pulleyblank 1996, Kuzmina, 1998, Beckwith 2002, Di Cosmo 2002) and possibly sites like Zhukaigou and the Qijia culture (Linduff 1995), as well as evidence of Iranians on the steppe and possibly the Altai region at that time. The suggestion that actual European populations may have been in northern China at that time conflicts with general evidence of population movement on the steppe, which sees gradual movement of putative Indo-Iranians and Indo-Aryans throughout the steppe and associated areas in the 2nd and 1st millennia BC around Central Eurasia from the Indo-Aryans in India, the western Iranians on the Iranian plateau, and the Scythians and Sarmatians (and related groups) on the South Russian steppe and Eastern Europe. There is some evidence of them being on the Mongolic steppe (see Askarov et al. 1992), as well as evidence of their inhabitance of Xinjiang (such as Khotan) and possibly the Altai region (also the Tokharians, though they were not Indo-Iranian). Whether or not the Linzi site was populated by Iranian-like peoples (and whether these peoples came from the putative steppe Iranians to the west or from the possible Iranians of nearby ancient Xinjiang) is not clear from this study. However, this could explain the affinity between the Linzi site and the West. This would also fall in line with other evidence of admixture in populations in the region. Of course, the most difficult issue is that the early Iranians (and Indo-Iranians) were a linguistic group, and while perhaps bearing some biological affinities, the degree to which the supposed Iranian groups of Central Eurasia had a biological affinity is indeterminate.
결과는 중국 북부 임치(臨淄)와 몽골 에긴골(Egyin Gol)의 고대 인구 집단 간 유전적 연관성에 확실한 차이가 있음을 시사한다. 임치 자료는 오늘날 중국 북부 인구 집단보다 근동인 및 유럽인과 더 강한 연관성을 보이는 것 같다. 물론 임치 내에도 동아시아인 및/또는 동남아시아인 성분이 확실히 존재한다(하플로그룹 분류에서 증명됨, 아래 참조). 우리는 고대 임치 지역에 “유럽인 같은 인구 집단”이 있었다기보다는, 임치 자료가 적어도 부분적으로 인도-이란인과 관련이 있을 수 있다고 제안한다. 인도-이란인은 인도유럽어족의 한 갈래이나 이 시기에는 더 정확히 그냥 “이란인”이다. 그들은 그 시기 또는 적어도 그 직전에 유라시아 중앙부 전역에 거주했을 가능성이 높다. 더 정확하게 말하면, 임치 인구 집단은 카라수크(Karasuk) 또는 사카(Saka)(시간적, 공간적으로 들어맞는 이란계 추정 집단)와 매우 밀접한 관련이 있었을 가능성이 크다. 혹은 안드로노보(Andronovo), 아파나시에보(Afanasievo), 스키타이(Scythian), 사르마티아(Sarmatian)인, 심지어 소그드(Sogdian)인과 조금 더 멀리 관련되었을 수도 있다. 카라수크와 사카는 기원전 1천 년경에 중앙 유라시아의 중앙 및 동부 지역에 존재했다는 점에서 가장 가능성이 높다. 비록 이러한 민족 이름은 다소 모호하며 정확한 연결은 사실상 불가능하지만 말이다. 그러나 하르마타(Harmatta)(1992)는 기원전 1천 년에 초기 이란계 그룹이 동유럽에서 중국 북부까지 중앙 유라시아 전역에 퍼져 있었다고 주장했다. 또한 아스카로프(Askarov) 등(1992)은 같은 기원전 1천 년 동안 몽골 북서부에 석관 무덤이 존재했음을 지적했다. (그들의 제안에 따르면, 약간의 몽골로이드(Mongoloid) 혼혈을 지닌 “유로포이드(Europoid)” 유골이 포함됨). 비록 추측이긴 하지만, 이러한 추론은 앞서 언급한 중국 청동기 문화의 변화, 고대 중국어의 차용어(Pulleyblank 1996, Kuzmina 1998, Beckwith 2002, Di Cosmo 2002), 주개가(朱開溝)와 제가(齊家) 문화 같은 유적지(Linduff 1995), 그리고 당시 초원 및 아마도 알타이(Altai) 지역에 이란인이 있었다는 증거 등 고고학과 언어학의 다른 여러 증거들과 잘 들어맞는다. 실제 유럽 인구 집단이 당시 중국 북부에 있었을 수 있다는 제안은 초원 지역의 인구 이동에 대한 일반적인 증거와 상충된다. 이는 기원전 2천 년과 1천 년에 중앙 유라시아 주변의 초원과 관련 지역을 통한 인도-이란인 및 인도-아리아(Indo-Aryan)인 추정 집단의 점진적인 이동을 보여주는 증거들이다. 즉, 인도의 인도-아리아인, 이란 고원의 서부 이란인, 남부 러시아 초원과 동유럽의 스키타이인과 사르마티아인(및 관련 집단) 등이다. 이들이 몽골 초원에 있었다는 몇 가지 증거가 있다(Askarov et al. 1992 참조). 또한 신강(新疆)(예: 호탄(和田)) 및 아마도 알타이 지역에 거주했다는 증거도 있다(토하라(Tokharian)인도 있었지만 이들은 인도-이란인이 아니었다). 임치(臨淄) 유적지에 이란인 같은 민족이 거주했는지 여부(그리고 이 사람들이 서쪽의 초원 이란인 추정 집단에서 왔는지, 아니면 인근 고대 신강의 가능한 이란인에서 왔는지)는 이 연구에서 명확하지 않다. 그러나 이것은 임치 유적지와 서양 사이의 연관성을 설명할 수 있다. 이것은 또한 이 지역 인구 집단의 혼혈에 대한 다른 증거와도 일치할 것이다. 물론 가장 어려운 문제는 초기 이란인(그리고 인도-이란인)이 언어 그룹이었다는 점이다. 생물학적 연관성을 지녔을 수도 있지만, 중앙 유라시아의 이란계 추정 집단들이 생물학적 연관성을 어느 정도 가졌는지는 불확실하다.
As to why this study disagrees somewhat with previous results (Wang et al. 2000, Yao et al. 2003), there are several likely reasons. First, it should be noted that the results do agree with the two previous studies to some degree, in that the Linzi sample does appear to have some affinities to populations to the west as well as some populations of Southeast Asia, or at least southern China and Vietnam. The approach here cannot clearly account for significant admixture, as it rather weighs out the closest matches. Thus, the possibility exists that the Linzi population was a heavily admixed group containing elements from both the westerly populations as well as the southern Chinese and/or Southeast Asians. However, the argument by Yao et al. (2003) that the discrepancy is due purely to the shortness of the sequence length is not correct, as shown above (see tables 7 and 8). Yao et al. (2003) approached the issue by attempting to haplogroup the populations. However, due the shortness of the sequences, eight of the individuals were classified as unknown (about a quarter of the sample). Further, the six individuals classified as haplogroup B were in fact no different from CRS (Cambridge Reference sequence) in this segment (as a number of identified haplogroups from both Asia and Europe have no mutations in this particular segment). Also, six individuals classified as haplogroup B5A were found via a GenBank blast search to have near matches in both Asian and European populations, including Portuguese, Hungarians, Balkans, and Norse (all containing the two mutations, 16266A and 16274, though the Europeans also had an extra mutation here at 16258C, while all of the Asian matches except one lacked 16266A). The above discrepancies account for 20 out of the 34 individuals in Linzi. While the two proposed groups of individuals with haplogroups B and B5A would likely fall into B given their geographic location, we cannot simply assume that they are, or at least that they all are B (rather than H for instance). If we could simply assume individuals from Asia are all of “Asian” haplogroups (and likewise for other geographic locations), there would be no point in doing further research, since we already can assume the answer a priori.
이 연구가 이전의 결과들(Wang et al. 2000, Yao et al. 2003)과 다소 일치하지 않는 이유로는 몇 가지 가능성이 있다. 첫째, 결과가 두 이전 연구와 어느 정도 일치한다는 점에 주목해야 한다. 임치(臨淄) 표본은 동남아시아의 일부 인구 집단(또는 최소한 중국 남부와 베트남(越南))뿐만 아니라 서구의 인구 집단과도 어느 정도 연관성을 보이기 때문이다. 이 접근법은 심각한 혼혈을 명확히 설명할 수 없다. 오히려 가장 가까운 일치 항목에 가중치를 두기 때문이다. 따라서 임치 인구 집단이 서쪽 인구 집단과 중국 남부 및/또는 동남아시아인의 요소를 모두 포함하는 고도로 혼혈된 집단이었을 가능성이 존재한다. 그러나 불일치가 순전히 염기서열 길이가 짧기 때문이라는 야오(Yao) 등(2003)의 주장은 위에서 보여준 바와 같이 옳지 않다(표 7과 8 참조). 야오 등(2003)은 인구 집단의 하플로그룹(haplogroup)을 분류하려는 방식으로 이 문제에 접근했다. 그러나 염기서열이 짧기 때문에 8명(표본의 약 4분의 1)이 알 수 없음으로 분류되었다. 게다가, 하플로그룹 B로 분류된 6명은 이 구간에서 사실상 CRS(Cambridge Reference sequence)와 다르지 않았다. (아시아와 유럽(Europe)의 확인된 여러 하플로그룹이 이 특정 구간에 돌연변이를 가지고 있지 않기 때문이다). 또한 하플로그룹 B5A로 분류된 6명은 젠뱅크(GenBank) 블래스트(blast) 검색을 통해 포르투갈(Portugal)인, 헝가리(Hungary)인, 발칸(Balkan)인, 노르드(Norse)인 등 아시아 및 유럽 인구 집단 모두에서 거의 일치하는 것으로 나타났다. (모두 16266A와 16274라는 두 돌연변이를 포함했다. 유럽인들은 여기에 16258C라는 돌연변이를 추가로 가졌고, 아시아 일치 항목들은 하나를 제외하고 모두 16266A가 없었다). 위의 불일치는 임치 표본 34명 중 20명을 설명한다. 하플로그룹 B와 B5A를 가진 두 제안된 개체군은 지리적 위치를 고려할 때 B에 속할 가능성이 높다. 하지만 이들이 (예를 들어 H가 아니라) 단순히 B라거나, 적어도 모두 B라고 가정할 수는 없다. 아시아 출신 개체들이 모두 “아시아” 하플로그룹(다른 지리적 위치도 마찬가지)이라고 단순히 가정할 수 있다면, 이미 선험적으로 답을 추정할 수 있으므로 추가 연구를 할 이유가 없을 것이다.
This fact is further reinforced by a blast search analysis in GenBank of those individuals which were classified as unknown by Yao et al. 2003 and simply removed from the analysis (Linzi 7, 8, 10, 14, 21, 22, 24, 31). Linzi 21, 22, and 24 all contain a mutation at 16264 which was only evidenced in GenBank in an ancient Australian, though Linzi 21 also contained a mutation at 16355 which is found in a couple of modern Australian Aborigines as well as Scots, Georgians, Ossetians, Kazakhs and Norse (but never in tandem with 16264). Linzi 7, 8, and 14 contained mutations at 16231, 16256, 16270, and 16274. There were no exact matches with them, but the closest matches all came from the western part of Eurasia, including the populations of Adygeis (from the Caucasus), Syrians, Icelanders, Ossetians, Portuguese, Hungarians, Romanians, Serbians, Norse, Swedish and others. Linzi 10 and 31 were probably the most interesting, containing mutations at 16293 and 16311, with exact matches with Scottish, Greeks, Adygeis, Hungarians, Portuguese, Balkans, Slovakians, Estonians and others. All the exact matches were from western and central parts of Eurasia. The point here is to show that individuals were present at Linzi who likely were related to populations from western and/or central Eurasia. If this is the case, we can further suggest that the individuals who were automatically assumed by Yao et al. (2003) to be an Asian haplogroup, B for instance, may in fact potentially be something else, such as H (or at least some of them may be). The above evidence also highlights the problem of the presence of haplotypes in ancient populations which may be rare or nonexistent in modern populations, perhaps due to drift, coalescence, selective sweeps and other effects. While the above “unknown” haplotypes may not be part of the identified haplogroup paradigm, they certainly existed in the past, and in local populations may even have been prevalent to some degree. This problem is exacerbated when the analysis includes not only spatial variation, but temporal as well, as evolutionary forces can shift with time and situation.
이 사실은 야오(Yao) 등(2003)이 알 수 없음으로 분류하고 분석에서 단순히 제외한 개체들(임치(臨淄) 7, 8, 10, 14, 21, 22, 24, 31)을 젠뱅크(GenBank)에서 블래스트(blast) 검색한 분석을 통해 더욱 강화된다. 임치 21, 22, 24는 모두 16264에 돌연변이를 포함하고 있다. 이는 젠뱅크에서 고대 호주(Australia)인에서만 확인되었다. 반면 임치 21은 16355에도 돌연변이를 포함했다. 이 돌연변이는 소수의 현대 호주 원주민뿐만 아니라 스코틀랜드(Scotland)인, 조지아(Georgia)인, 오세티야(Ossetia)인, 카자흐(Kazakh)인, 노르드(Norse)인에서도 발견된다(그러나 16264와 함께 나타나지는 않는다). 임치 7, 8, 14는 16231, 16256, 16270, 16274에 돌연변이를 포함했다. 이들과 정확히 일치하는 항목은 없었다. 하지만 가장 가까운 일치 항목은 아디게(Adygei)인(코카서스(Caucasus) 출신), 시리아(Syria)인, 아이슬란드(Iceland)인, 오세티야인, 포르투갈(Portugal)인, 헝가리(Hungary)인, 루마니아(Romania)인, 세르비아(Serbia)인, 노르드인, 스웨덴(Sweden)인 등 모두 유라시아 서부 지역에서 나왔다. 임치 10과 31은 아마도 가장 흥미로웠을 것이다. 이들은 16293과 16311에 돌연변이를 포함했고 스코틀랜드인, 그리스(Greece)인, 아디게인, 헝가리인, 포르투갈인, 발칸(Balkan)인, 슬로바키아(Slovakia)인, 에스토니아(Estonia)인 등과 정확히 일치했다. 모든 정확한 일치 항목은 유라시아(Eurasia) 서부 및 중부 지역 출신이었다. 여기서 핵심은 유라시아 서부 및/또는 중부 지역 인구 집단과 연관되었을 가능성이 높은 개체들이 임치에 존재했음을 보여주는 것이다. 만약 그렇다면, 야오 등(2003)이 (예컨대 B와 같은) 아시아 하플로그룹이라고 자동적으로 가정한 개체들이 실제로는 H 같은 다른 그룹(적어도 일부는)일 가능성도 있음을 더 제안할 수 있다. 위의 증거는 고대 인구 집단에는 존재하지만 현대 인구 집단에서는 드물거나 존재하지 않을 수 있는 하플로타입(haplotype)의 문제점을 또한 강조한다. 이는 아마도 유전자 부동, 유전자 병합, 선택적 싹쓸이(selective sweep) 및 기타 효과 때문일 것이다. 위의 “알 수 없는” 하플로타입이 확인된 하플로그룹 패러다임(paradigm)의 일부가 아닐 수도 있지만, 과거에는 분명히 존재했다. 특정 지역 인구 집단에서는 어느 정도 우세했을 수도 있다. 이 문제는 공간적 변동뿐만 아니라 시간적 변동까지 분석에 포함될 때 악화된다. 진화적 힘은 시간과 상황에 따라 변할 수 있기 때문이다.
The method here is designed to overcome the above problems. The haplotypes that could not be identified as a particular haplogroup would actually have a negligible effect on the results, as they would place Linzi equidistant from the various regional populations. It is the informative haplotypes and the mutations that comprise them, those which are rare or nonexistent in the other populations, which would make any given population have a relatively increased affinity with the probe population, in this case Linzi or Egyin. There are, of course, several individuals at Linzi who do appear to belong to haplogroups A, B, D, F, G, and M (all modern Asian haplogroups), thus explaining the results of Yao et al. (2003) as well as our own. However, as noted above, there are also a number of sequences that do appear to have an affinity to the west, which were thrown out by Yao et al. (2003) because they were not part of the known haplogroup paradigm, but likely explain the discrepancy between our results and theirs.
여기서 사용된 방법은 위와 같은 문제들을 극복하도록 설계되었다. 특정 하플로그룹(haplogroup)으로 식별되지 못한 하플로타입(haplotype)은 사실상 결과에 미치는 영향이 무시할 만할 것이다. 이로 인해 임치(臨淄)는 여러 지역 인구 집단으로부터 동일한 거리에 위치하게 되기 때문이다. 어떤 집단이 탐색 인구 집단(이 경우 임치나 에긴(Egyin))과 상대적으로 연관성을 높이게 만드는 것은 바로 유용한 하플로타입과 그것을 구성하는 돌연변이들이다. 이들은 다른 인구 집단에는 드물거나 존재하지 않는 것들이다. 물론 임치에는 하플로그룹 A, B, D, F, G, M(모두 현대 아시아 하플로그룹)에 속하는 것으로 보이는 개체들이 여러 명 있다. 이는 야오(Yao) 등(2003)의 결과뿐만 아니라 우리의 결과를 설명한다. 그러나 위에서 언급했듯이, 서양과 연관성을 가지는 것으로 보이는 여러 염기서열도 존재한다. 이들은 알려진 하플로그룹 패러다임(paradigm)에 속하지 않아 야오 등(2003)에 의해 배제되었다. 이들이 우리 결과와 그들 결과 사이의 불일치를 설명할 가능성이 높다.
The Egyin Gol individuals appear to be definitely East Asians, at least maternally. The Egyin samples showed an affinity with northern East Asians, such as modern Mongolians, Japanese, northern Chinese populations (Shandong, Liaoning), and ethnic Han of Xinjiang. It is clear that the Egyin population was significantly different from the Linzi material from just a few centuries earlier. Of course, northern Mongolia to the Shandong region of China is actually some distance (well over a 1000 km). However, the historical connections of the Chinese to the Mongolic steppe in the first millennium BC, such as between the Hsiung-Nu and Han dynasty, show that the regions were in contact and had some degree of interaction (Watson 1961, Di Cosmo 2002). If the results from the Egyin Gol site can be duplicated by other finds from around the region, then there will be clear evidence that northern East Asians were the principle occupants of the area (including the steppe regions) by at least the rise of the Han dynasty in China (and perhaps the Qin dynasty or even earlier). Further correlation of the results of the Linzi site to other sites around the region dating to the middle of the first millennium BC and earlier (such as the genetic analysis by Ricaut et al. 2004 of an ancient individual from the Altai region) could provide evidence of a population shift in the region, depending of course on the degree to which populations like Linzi inhabited the region. It would also be useful to explore back into the Neolithic or even earlier to see whether the Linzi peoples were migrants to the region or descendents of earlier inhabitants. It should also be noted that if the attribution of the Egyin Gol site to the Hsiung-Nu is correct, then the results of the genetic analysis may suggest that the Hsiung-Nu were at least in part the ancestors of the later Mongolic and possibly Turkic peoples who would come to inhabit the steppe region in the first millennium AD. Further, if the evidence for a population shift can be corroborated and it is combined with the attribution of the Egyin Gol material to the Hsiung-Nu, then the Egyin Gol site may represent an element of some sort of genesis (though not necessarily the actual starting point), that would later result in the eruption of the Turkic and Mongol peoples from the Mongolic steppe and the Altai region of Central Eurasia. However, without correlation of the Linzi and Egyin Gol sites with other ancient sites from around the region, the evidence derived from these two sites will remain isolated cases.
에긴골(Egyin Gol) 개체들은 적어도 모계로는 분명히 동아시아인으로 보인다. 에긴 표본은 현대 몽골인, 일본인, 중국 북부 인구 집단(산동(山東), 요녕(遼寧)), 신강(新疆)의 한족(漢族) 같은 동아시아 북부인과 연관성을 보였다. 에긴 인구 집단이 불과 몇 세기 전의 임치(臨淄) 자료와 상당히 달랐다는 것은 분명하다. 물론 몽골 북부에서 중국 산동 지역까지는 실제 상당한 거리(1000km 이상)다. 그러나 흉노(匈奴)와 한나라(漢朝) 사이의 관계처럼 기원전 1천 년대 중국인과 몽골 초원의 역사적 연결은 이 지역들이 접촉했고 어느 정도 상호 작용을 했음을 보여준다(Watson 1961, Di Cosmo 2002). 에긴골 유적지의 결과를 이 지역 주변의 다른 발견에서 동일하게 확인할 수 있다면, 동아시아 북부인들이 늦어도 중국에서 한나라가 부흥할 즈음(아마도 진나라(秦朝)나 그 이전)까지 이 지역(초원 지역 포함)의 주요 거주자였다는 명확한 증거가 될 것이다. 임치 유적지의 결과를 기원전 1천 년 중반 및 그 이전으로 거슬러 올라가는 주변 지역의 다른 유적지와 추가로 연관시키면 지역 내 인구 이동에 대한 증거를 제공할 수 있다. (예컨대 알타이(Altai) 지역 고대 인구에 대한 리코(Ricaut) 등(2004)의 유전자 분석이 있다). 물론 임치와 같은 인구 집단이 이 지역에 거주한 정도에 따라 달라질 것이다. 신석기 시대나 그 이전 시기를 조사하여, 임치 사람들이 이 지역으로 이주한 이민자인지 이전 거주자의 후손인지 확인하는 것도 유용할 것이다. 에긴골 유적지를 흉노에 귀속시키는 것이 맞다면, 이 유전자 분석 결과는 흉노가 적어도 부분적으로 후대 몽골계 민족과 아마도 튀르크계 민족의 조상임을 시사할 수 있다는 점에도 주목해야 한다. 이들은 서기 1천 년대에 초원 지역에 거주하게 된 집단들이다. 게다가 인구 이동에 대한 증거가 입증되고 그것이 에긴골 자료를 흉노로 귀속시키는 것과 결합된다면, 에긴골 유적지는 일종의 기원 요소를 나타낼 수 있다(반드시 실제 출발점은 아니더라도). 이는 이후 중앙 유라시아(Eurasia)의 몽골 초원과 알타이 지역에서 튀르크계 및 몽골계 민족이 폭발적으로 등장하는 결과를 낳았다. 그러나 임치 및 에긴골 유적지와 이 주변의 다른 고대 유적지와의 연관성이 없다면, 이 두 유적지에서 도출된 증거는 고립된 사례로 남을 것이다.
감사의 글 Acknowledgments
We would like to thank Dr. Toomas Kivisild and Mait Metspalu of the Estonian Biocenter for providing their data for our use, and without whose help this study would not have been possible. We also thank Dr. Christopher Beckwith of the Central Eurasian Studies Department at Indiana University-Bloomington, Dr. Paul Jamison of the Anthropology Department at Indiana University-Bloomington, and Dr. Michael Wade of the Department of Biology at Indiana University-Bloomington for their review of a draft of this paper and thoughtful suggestions. We would also like to acknowledge Dick Repasky of the Research SP and UITS BioInformatics Support at Indiana University-Bloomington for his tireless efforts to help me resolve computing issues during this research, and Dr. F. Calafell for helping me locate sequence data.
에스토니아 바이오센터(Estonian Biocenter)의 투마스 키비실드(Toomas Kivisild) 박사와 마이트 메츠팔루(Mait Metspalu)에게 감사한다. 그들은 우리가 사용할 데이터를 제공했다. 그들의 도움이 없었다면 이 연구는 불가능했을 것이다. 인디애나 대학교 블루밍턴(Indiana University-Bloomington) 중앙 유라시아 연구학과의 크리스토퍼 벡위드(Christopher Beckwith) 박사, 인류학과의 폴 제미슨(Paul Jamison) 박사, 생물학과의 마이클 웨이드(Michael Wade) 박사에게도 감사한다. 그들은 이 논문의 초안을 검토하고 사려 깊은 제안을 제공했다. 인디애나 대학교 블루밍턴의 리서치 SP(Research SP) 및 UITS 생물정보학 지원팀의 딕 레파스키(Dick Repasky)에게도 감사를 표한다. 그는 이 연구 기간 동안 컴퓨팅 문제를 해결하도록 돕기 위해 끊임없이 노력했다. 염기서열 데이터를 찾는 데 도움을 준 F. 칼라펠(F. Calafell) 박사에게도 감사한다.
참고문헌 Literature Cited
Al-Zahery, N., O. Semino, G. Benuzzi et al. 2003. Y-chromosome and mtDNA polymorphisms in Iraq, a crossroad of the early human dispersal and of post-Neolithic migrations. Mol. Phyl. Evol. 28(3):458-472.
Anthony, D. 1998. The opening of the Eurasian steppe at 2000 BC. In The Bronze Age and Early Iron Age Peoples of Eastern Central Asia, V. Mair, ed. Washington, DC: Institute for the Study of Man (with The University of Pennsylvania Museum Publications), v. 1, 94-113.
Anthony, D. W., and N. B. Vinogradov. 1995. Birth of the chariot. Archaeology 48:36-41.
Askarov, A. 1992. The beginning of the Iron Age in Transoxiana. In History of Civilizations of Central Asia, A. H. Dani and V. M. Masson, eds. Paris: Unesco Publishing, v. 1, 481-501.
Beckwith, C. 2002. The Sino-Tibetan problem. In Medieval Tibeto-Burman Languages, C. Beckwith, ed. Leiden, Netherlands: E. J. Brill, 113-157.
Benson, D. A., I. Karsch-Mizrachi, D. J. Lipman et al. 2000. GenBank. Nucleic Acids Res. 28(1):15-18.
Bentley, J. 1993. Old World Encounters. New York: Oxford University Press.
Bertranpetit, J., J. Sala, F. Calafell et al. 1995. Human mitochondrial DNA variation and the origin of the Basques. Ann. Hum. Genet. 59:63-81.
Brakez, Z., E. Bosch, H. Izaabel et al. 2001. Human mitochondrial DNA sequence variation in the Moroccan population of the Souss area. Ann. Hum. Biol. 28:295-307.
Cavalli-Sforza, L. L., P. Menozzi, and A. Piazza. 1994. The History and Geography of Human Genes. Princeton, NJ: Princeton University Press.
Comas, D., F. Calafell, N. Bendukidze et al. 2000. Georgian and Kurd mtDNA sequence analysis shows a lack of correlation between languages and female genetic lineages. Am. J. Phys. Anthropol. 112(1):5-16.
Comas, D., F. Calafell, E. Mateu et al. 1996. Geographic variation in human mitochondrial DNA control region sequence: The population history of Turkey and its relationship to the European populations. Mol. Biol. Evol. 13(8):1067-1077.
Comas, D., F. Calafell, E. Mateu et al. 1998. Trading genes along the Silk Road: mtDNA sequences and the origin of Central Asian populations. Am. J. Hum. Genet. 63(6):1824-1838.
Di Cosmo, N. 2002. Ancient China and Its Enemies: The Rise of Nomadic Power in East Asian History. New York: Cambridge University Press.
Golden, P. B. 1991. The peoples of the south Russian steppes. In The Cambridge History of Early Inner Asia, D. Sinor, ed. New York: Cambridge University Press, 229-255.
Gresham, D., B. Morar, P. A. Underhill et al. 2001. Origins and divergence of the Roma (Gypsies). Am. J. Hum. Genet. 69(6):1314-1331.
Handt, O., S. Meyer, and A. von Haeseler. 1998. Compilation of human mtDNA control region sequences. Nucleic Acids Res. 26(1):126-129.
Harmatta, J. 1992. The emergence of the Indo-Iranians: The Indo-Iranian languages. In History of Civilizations of Central Asia, A. H. Dani and V. M. Masson, eds. Paris: Unesco Publishing, v. 1, 357-378.
Helgason, A., E. Hickey, S. Goodacre et al. 2001. mtDNA and the islands of the North Atlantic: Estimating the proportions of Norse and Gaelic ancestry. Am. J. Hum. Genet. 68(3):723-737.
Holsinger, K. E., and R. J. Mason-Gamer. 1996. Hierarchical analysis of nucleotide diversity in geographically structured populations. Genetics 142:629-639.
Hudson, R. R., M. Slatkin, and W. P. Maddisson. 1992. Estimation of levels of gene flow from DNA sequence data. Genetics 132:583-589.
Imaizumi, K., T. J. Parsons, M. Yoshino et al. 2002. A new database of mitochondrial DNA hyper- variable regions I and II sequences from Japanese individuals. Int. J. Leg. Med. 116(2):68-73.
Jorde, L. B., M. J. Bamshad, W. S. Watkins et al. 1995. Origins and affinities of modern humans: A comparison of mitochondrial and nuclear genetic data. Am. J. Hum. Genet. 57(30):523-538.
Keyser-Tracqui, C., E. Crubezy, and B. Ludes. 2003. Nuclear and mitochondrial DNA analysis of a 2,000-year-old necropolis in the Egyin Gol valley of Mongolia. Am. J. Hum. Genet. 73:247- 260.
Kivisild, T., S. Rootsi, M. Metspalu et al. 2003. The genetic heritage of the earliest settlers persists both in Indian tribal and caste populations. Am. J. Hum. Genet. 72(2):313-332.
Kivisild, T., H. V. Tolk, J. Parik et al. 2002. The emerging limbs and twigs of the East Asian mtDNA tree. Mol. Biol. Evol. 19:1737-1751.
Kong, Q. P., Y. G. Yao, M. Liu et al. 2003. Mitochondrial DNA sequence polymorphisms of five ethnic populations from northern China. Hum. Genet. 113(5):391-405.
Kuzmina, E. E. 1998. Cultural connections of the Tarim basin people and pastoralists of the Asian steppes in the Bronze Age. In The Bronze Age and Early Iron Age Peoples of Eastern Central Asia, V. Mair, ed. Washington, DC: Institute for the Study of Man (with The University of Pennsylvania Museum Publications), v. 1, 63-93.
Lattimore, O. 1951. Inner Asian Frontiers of China. New York: American Geographical Society.
Lee, S., C. Shin, K. Kim et al. 1997. Sequence variation of mitochondrial DNA control region in Koreans. Forensic Sci. Intl. 87:99-116.
Levine, M. A. 1999. Dereivka and the problem of horse domestication. Antiquity 64:727-740.
Linduff, K. M. 1995. Zhukaigou. Antiquity 69:133-145.
Long, J. C., and R. A. Kittles. 2003. Human genetic diversity and the nonexistence of biological races. Hum. Biol. 75(4):449-471.
Lynch, M., and T. J. Crease 1990. The analysis of population survey data on DNA sequence variation. Mol. Biol. Evol. 7:377-394.
Maddison, W. P., and D. R. Maddison 1989. Interactive analysis of phylogeny and character evolution using the computer program MacClade. Folia Primatol. 53(1-4):190-202.
Mair, V. 1995. Mummies of the Tarim basin. Archaeology 48:28-35.
Mair, V., ed. 1998. The Bronze Age and Early Iron Age Peoples of Eastern Central Asia. Washington, DC: Institute for the Study of Man (with The University of Pennsylvania Museum Publica- tions).
Mallory, J. P. 1989. In Search of Indo-Europeans: Language, Archaeology, and Myth. New York: Thames and Hudson.
Nei, M. 1977. F-statistics and analysis of gene diversity in subdivided populations. Ann. Hum. Genet. 41:225-233.
Nei, M. 1986. Definition and estimation of fixation indexes. Evolution 40(3):643-645.
Nei, M. 1987. Molecular Evolutionary Genetics. New York: Columbia University Press.
Okladnikov, A. P. 1990. Inner Asia at the dawn of history. In The Cambridge Early History of Early Inner Asia, D. Sinor, ed. New York: Cambridge University Press, 41-96.
Oota, H., T. Kitano, F. Jin et al. 2002. Extreme mtDNA homogeneity in continental Asian popula- tions. Am. J. Phys. Anthropol. 118(2):146-153.
Oota, H., M. Naruya, and S. Ueda. 1999. Molecular genetic analysis of remains of a 2000-year-old human population in China-and its relevance for the origin of the modern Japanese popula- tion. Am. J. Hum. Genet. 64:250-258.
Oota, H., W. Settheetham-Ishida, D. Tiwawech et al. 2001. Human mtDNA and Y-chromosome variation is correlated with matrilocal versus patrilocal residence. Nat. Genet. 29:20-21.
Parpola, A. 1998. Aryan languages, archaeological cultures and Sinkiang: Where did proto-Iranian come into being and how did it spread. In The Bronze Age and Early Iron Age Peoples of Eastern Central Asia, V. Mair, ed. Washington, DC: Institute for the Study of Man (with The University of Pennsylvania Museum Publications), v. 1, 114-147.
Plaza, S., F. Calafell, A. Helal et al. 2003. Joining the pillars of Hercules: mtDNA sequences show a multidirectional gene flow in the western Mediterranean. Ann. Hum. Genet. 67:312-328.
Praslov, N. D., V. N. Stanko, Z. A. Abromova et al. 1989. The steppes in the late Paleolithic. Antiquity 63:784-792.
Pulleyblank, E. G. 1996. Early contacts between Indo-Europeans and Chinese. Int. Rev. Chinese Linguistics 1(1):1-25.
Redd, A. J., and M. Stoneking. 1999. Peopling of the Sahul: mtDNA variation in Aboriginal Austra- lian and Papua New Guinean populations. Am. J. Hum. Genet. 65:808-828.
Renfrew, C. 1987. Archaeology and Language: The Puzzle of Indo-Europeans. Cambridge, England: Cambridge University Press.
Ricaut, F., C. Keyser-Tracqui, J. Bourgeois et al. 2004. Genetic analysis of a Scytho-Siberian skeleton and its implications for ancient Central Asian migrations. Hum. Biol. 76(1):109-125.
Richards, M., H. Corte Real, P. Forster et al. 1996. Paleolithic and Neolithic lineages in the European mitochondrial gene pool. Am. J. Hum. Genet. 59(1):185-203.
Rozas, J., and R. Rozas. 1999. DNAsp version 3: An integrated program for molecular population genetics and molecular evolution analysis. Bioinformatics 15:174-175.
Sajantila, A., P. Lahermo, P. Anttinen et al. 1995. Genes and languages in Europe: An analysis of mitochondrial lineages. Gen. Res. 5:42-52.
Sinor, D., ed. 1990. The Cambridge Early History of Early Inner Asia. New York: Cambridge Univer- sity Press.
Smouse, P. E., V. J. Vitzthum, and J. V. Neel. 1981. The impact of random and lineal fission on the genetic divergence of small human groups: A case study among the Yanomama. Genetics 98:179-197.
Templeton, A. R., E. Routman, and C. A. Phillips. 1995. Separating population structure from popula- tion history: A cladistic analysis of the geographical distribution of mitochondrial DNA hap- lotypes in the Tiger Salamander, Ambystoma tigrinum. Genetics 140(2):767-782.
Wang, L., H. Oota, N. Saitou et al. 2000. Genetic structure of a 2,500-year-old human population in China and its spatiotemporal changes. Mol. Biol. Evol. 17(9):1396-1400.
Watson, B., trans. 1961. Records of the Grand Historian of China (Ssuma Chien). New York: Colum- bia University Press.
Weir, B. S., and C. C. Cockerham. 1984. Estimating F-statistics for the analysis of population struc- ture. Evolution 38:1358-1370.
Whitlock, M. C. 1994. Fission and genetic variance among populations: The changing demography of Forked Fungus Beetle populations. Am. Nat. 143:820-829.
Yao, Y. G., Q. P. Kong, H. J. Bandelt et al. 2002a. Phylogeographic differentiation of mitochondrial DNA in Han Chinese. Am. J. Hum. Genet. 70(3):635-651.
Yao, Y. G., Q. P. Kong, X. Y. Man et al. 2003. Reconstructing the evolutionary history of China: A caveat about inferences drawn from ancient DNA. Mol. Biol. Evol. 20(2):214-219.
Yao, Y. G., L. Nie, H. Harpending et al. 2002b. Genetic relationship of Chinese ethnic populations revealed by mtDNA sequence diversity. Am. J. Phys. Anthropol. 118(1):63-76.
Yao, Y. G., L. Nie, H. Harpending et al. 2002c. Genetic relationships of Chinese ethnic populations revealed by mtDNA sequence diversity. Am. J. Hum. Genet. 18(1):63-80.