출처:

Mao, Xiaowei, et al. “The deep population history of northern East Asia from the Late Pleistocene to the Holocene.” Cell 184.12 [2021]: 3256-3266.

The deep population history of northern East Asia from the Late Pleistocene to the Holocene

빙하기 말부터 현재까지, 북부 동아시아 인류 집단의 변천사

  • Xiaowei Mao, Hucai Zhang, Shiyu Qiao, …, John W. Olsen, E. Andrew Bennett, Qiaomee Fu
    마오, 샤오웨이(Mao, Xiaowei), 장, 후카이(Zhang, Hucai), 차오, 시유(Qiao, Shiyu) 외 다수
  • 대표 연구자 연락처 Correspondence: 장, 후카이(zhanghc@ynu.edu.cn), 푸, 차오메이(fuqiaomei@ivpp.ac.cn)

 

[리뷰] 중화 쇼비니즘 경사도 평가: 3/10

12. Mao, X. et al. (2021) ‘The deep population history of northern East Asia from the Late Pleistocene to the Holocene’, Cell, 184(12), pp. 3256–3266.e13. https://linkinghub.elsevier.com/retrieve/pii/S0092867421005754.

(1) 연구 개요 저자의 주장

이 연구는 아무르강 유역을 중심으로 후기 구석기 시대(33,600 BP)부터 홀로세에 이르는 고대인 25명의 유전체를 분석하여 북동아시아 인구사의 깊은 시간적 뿌리를 규명했다. 핵심 발견은 4만년 전, 전원동인(田園洞人; Tianyuan individual)과 관련된 혈통이 마지막 최대 빙하기(LGM) 이전에 북동아시아에 널리 퍼져 있었으며, LGM 이후에는 아무르강 유역에서 독자적인 고대 동북아시아인(ANA; Ancient Northeast Asian) 혈통이 나타나 14,000년 전부터 유전적 연속성을 유지했다는 것이다. 이 ANA 혈통은 황하 유역 농경민(YR; Yellow River)과는 뚜렷이 구분되는 북방의 고유한 기층을 형성했다.

(2) 편향성 분석 (중화 쇼비니즘 경사도: 3/10)

이 연구는 북방 지역의 독자적인 유전적 역사를 깊은 시간대까지 거슬러 올라가 증명함으로써, 모든 것의 기원을 황하 유역으로 돌리는 단선적 확산 모델을 효과적으로 약화시키는 데 크게 기여했다.

  • 서사 프레이밍 (낮은 편향성): 북동아시아 토착 혈통(ANA)의 장기적인 지속성과 황하 농경민(YR)과의 상호작용을 함께 제시함으로써, 중원에서 모든 것이 시작되었다는 ‘화살’ 모델을 효과적으로 약화시켰다.
  • 모델 선택과 반례 취급 (중간 편향성): 연구의 초점이 아무르강 유역과 내륙에 맞춰져 있어, 연안 및 도서 지역 표본이 부족했다. 이로 인해 타기도 유적과 같은 해양 네트워크의 독자성을 보여주는 반례를 모델에 통합하지는 못했다.
  • 지리·환경 제약 반영 (중간 편향성): 연구의 초점이 내륙의 ANA-YR 상호작용에 맞춰져 있어, ‘해양 회랑’의 역할을 정량적으로 분석하는 데는 한계가 있다.
  • 유전자–문화 결합 가정 (낮은 편향성): 유전적 변화를 특정 문화권과 직접 연결하기보다는 거시적인 생업 변화와 연관 지어, 유전자와 문화의 관계를 신중하게 다루었다.

(3) 결론 재구성

이 논문의 결론은 이미 상당 부분 ‘매트릭스’ 모델을 지지한다. 이를 더욱 강화하기 위해서는, ANA와 YR이라는 두 축에 더해 해양 회랑’독립적인 제3축으로 설정하여 재모델링할 필요가 있다. 이를 통해 내륙에서의 상호작용과 연안에서의 교류가 서로 다른 시간대에 비동시적으로 일어났을 가능성을 검증하고, 북동아시아 전체의 인구사를 더 입체적으로 재구성할 수 있다.

 

[논문요약]

마오 샤오웨이(Mao, Xiaowei) 등은 2021년 학술지 《셀(Cell)》에 발표한 논문, “후기 플라이스토세부터 홀로세까지 북부 동아시아의 깊은 인구 역사(The deep population history of northern East Asia from the Late Pleistocene to the Holocene)”에서, 흑룡강(黑龍江) 유역에서 발견된 34,000~3,400년 전 사이의 고대 인골 25개의 유전체(genome)를 분석했다. 연구 결과, 북부 동아시아에서는 마지막 빙하기를 기점으로 최소 두 차례의 대규모 인구 집단 교체가 있었음이 밝혔다.

첫째, 약 40,000년 전 북경(北京) 인근의 전원동(田園洞)에서 발견된 ‘전원인(田園人, Tianyuan individual)’의 유전적 특성이 마지막 빙하기 이전까지 북부 동아시아에 널리 퍼져 있었다.

둘째, 가장 추웠던 ‘최종빙기절정기(Last Glacial Maximum, LGM)’가 끝날 무렵인 약 19,000년 전, 이 지역에는 기존의 전원인 집단과 유전적으로 구별되는 새로운 인류 집단이 나타났다.

셋째, 약 14,000년 전부터 흑룡강 유역에는 유전적 연속성을 지닌 집단이 정착했으며, 이들이 아메리카 원주민의 조상과 가장 가까운 ‘고대 북부 시베리아인(Ancient Paleo-Siberians)’의 핵심 동아시아 조상임이 드러났다.

또한, 동아시아인의 특징인 두꺼운 머리카락, 삽 모양 앞니 등과 관련된 EDAR V370A 유전자 변이가 최종빙기절정기 이후에 급격히 확산되었음을 처음으로 고대 DNA를 통해 직접 확인했다. 이 연구는 그동안 미지의 영역으로 남아있던 북부 동아시아의 빙하기 전후 인류 역사를 유전학적으로 규명하고, 동아시아인과 아메리카 원주민의 기원을 연결하는 중요한 단서를 제공했다는 점에서 큰 의의를 가진다.

1. 서론: 북부 동아시아 고대 인류 역사의 미스터리

현대 인류는 약 40,000년 전부터 북부 동아시아 지역에 거주했다. 이 지역은 몽골 고원, 중국 북부, 한반도, 일본 열도 등을 포함한다. 유럽의 경우, 빙하기의 혹독한 기후 변화가 인구의 이동과 증감에 큰 영향을 미쳤다는 사실이 잘 알려져 있다. 하지만 동아시아, 특히 고위도 지역에서는 마지막 빙하기를 전후하여 어떤 인류 집단이 살았고, 이들이 어떻게 변화했는지에 대해서는 거의 알려진 바가 없었다. 이는 약 40,000년 전부터 9,000년 전 사이의 인골 화석이 거의 발견되지 않았기 때문이다. 이 연구는 흑룡강(黑龍江) 유역에서 발견된 다양한 시기의 고대 인골 유전체 분석을 통해 이 공백을 메우고자 했다.

2. 연구 대상 및 방법: 고대인의 뼈에서 DNA 추출하기

연구팀은 중국 동북부 흑룡강성(黑龍江省)의 송눈평원(松嫩平原)에서 발견된 25명의 고대인 유골을 분석 대상으로 삼았다. 이 유골들은 방사성탄소연대측정 결과, 약 33,600년 전에서 3,400년 전 사이의 것으로 확인되었다.

연구 방법의 핵심은 고대 DNA 분석 기술이다.

1. 시료 채취: 유골 분말을 소량(<100mg) 채취하여 DNA를 추출했다.

2. 유전체 데이터 확보: 추출된 DNA는 양이 매우 적고 손상되었기 때문에, ‘표적 농축(target enrichment)’ 기술을 사용해 인간 유전체의 특정 정보(약 120만 개의 단일염기다형성, SNP)를 집중적으로 포획하여 분석했다. 이를 통해 각 개체의 게놈(genome) 데이터를 확보했다.

3. 데이터 분석: 확보된 게놈 데이터를 현대인 및 다른 고대인들의 데이터와 비교 분석했다. 인구 집단 간의 유전적 거리와 관계를 파악하기 위해 주성분 분석(PCA), f3-통계, D-통계, qpGraph 등 다양한 통계 유전학적 분석 기법이 사용되었다.

연구팀은 분석 대상이 된 20명의 개체를 연대에 따라 5개의 그룹으로 나누었다.

  • AR33K: 최종빙기절정기(LGM) 이전, 약 33,000년 전의 여성.
  • AR19K: LGM 말기, 약 19,000년 전의 남성.
  • AR14K: 빙하기가 끝날 무렵, 약 14,000년 전의 2명.
  • AR13-10K: 약 13,000년 ~ 10,000년 전의 4명.
  • ARpost9K 등: 약 9,400년 ~ 3,400년 전의 12명.

※ 용어 설명
AR: 시료가 발견된 흑룡강 유역(Amur Region)을 의미하며, 뒤의 숫자는 대략적인 연대(K = 1,000년)를 뜻한다.

3. 주요 연구 결과: 세 번의 큰 인구 집단 변화

1. 최종빙기절정기 이전 (약 33,000년 전): 널리 퍼져 있던 전원인(田園人) 혈통

가장 오래된 개체인 33,000년 전의 AR33K는 유전적으로 약 40,000년 전 북경(北京) 근처 전원동(田園洞)에서 발견된 ‘전원인’과 매우 가까운 관계였다. 이는 전원인과 관련된 유전적 혈통이 LGM 이전에 북경(北京)에서부터 몽골, 흑룡강(黑龍江) 유역에 이르기까지 지리적으로 매우 넓게 퍼져 있었음을 시사한다.

흥미롭게도 몽골에서 발견된 34,000년 전의 살키트(Salkhit)인은 전원인 혈통과 시베리아의 야나(Yana) 혈통이 섞인 것으로 나타났지만 , 흑룡강(黑龍江)의 AR33K는 야나 혈통과 섞인 증거가 없었다. 이는 LGM 이전에도 지역에 따라 인구 집단 간의 교류 양상이 달랐음을 보여준다.

2. 최종빙기절정기 말 (약 19,000년 전): 새로운 집단의 등장

가장 추웠던 LGM이 끝나갈 무렵인 19,000년 전의 AR19K는 LGM 이전의 AR33K나 전원인과는 유전적으로 뚜렷이 구분되는 새로운 집단이었다. AR19K는 이후 시대의 동아시아인들과 훨씬 더 가까운 유전적 특징을 보였다.

이는 LGM의 혹독한 기후를 거치면서 북부 동아시아에서 인구 집단의 대규모 교체가 일어났음을 의미한다. 즉, LGM 이전에 널리 퍼져 있던 전원인 관련 집단은 사라지거나 다른 곳으로 밀려나고, AR19K로 대표되는 새로운 집단이 그 자리를 차지한 것이다. 이 AR19K 집단은 현존하는 모든 북부 동아시아인의 조상 계통 중 가장 오래된 뿌리에 해당하는 ‘최초의 북부 동아시아인’으로 볼 수 있다.

3. 최종빙기절정기 이후 (약 14,000년 전 ~ 현재): 유전적 연속성과 시베리아로의 확장

약 14,000년 전의 AR14K 집단부터 그 이후의 흑룡강(黑龍江) 유역 사람들은 약 8,000년 전 러시아 극동의 ‘악마의 문 동굴(Devil’s Gate Cave)’ 사람들과 유전적으로 매우 가까웠다. 이는 이 지역에서 최소 14,000년 전부터 장기간에 걸친 유전적 연속성이 있었음을 보여준다. 이는 이전에 알려졌던 것보다 6,000년이나 더 이른 시점이다.

더 중요한 발견은 이 흑룡강(黑龍江) 유역 집단이 ‘고대 북부 시베리아인(Ancient Paleo-Siberians)’의 기원을 설명하는 핵심 열쇠라는 점이다. 고대 북부 시베리아인은 아메리카 원주민의 조상과 가장 가까운 아시아 집단으로 알려져 있다. 기존 연구에서는 이들이 동아시아인과 고대 북부 유라시아인(ANE)의 혼혈로 추정했지만, 정확한 동아시아 조상 집단을 찾지 못했다.

이번 연구는 AR14K로 대표되는 흑룡강(黑龍江) 유역 집단이 바로 그 동아시아 조상 집단임을 명확히 밝혔다. 즉, 14,000년 전 이후 흑룡강(黑龍江) 유역에 살던 사람들이 시베리아로 이동하여 기존에 있던 고대 북부 유라시아인과 섞여 고대 북부 시베리아인을 형성했고, 이들이 다시 아메리카 대륙으로 건너가는 역사의 한 축을 담당했던 것이다.

4. 특별한 유전자(EDAR V370A)의 비밀

EDAR V370A는 동아시아인에게서 높은 빈도로 나타나는 유전자 변이로, **두꺼운 머리카락, 더 많은 땀샘, 그리고 삽 모양의 앞니(shovel-shaped incisors)**와 같은 신체적 특징을 만든다. 과학자들은 이 유전자가 왜 동아시아인에게서 선택적으로 많아졌는지 궁금해했다. 후보 가설로는 저위도 환경에서 모유를 통한 비타민 D 전달을 돕기 위함이라는 ‘2만 년 전 빙하기 가설’과, 덥고 습한 기후에서 체온 조절을 돕기 위함이라는 ‘3만 년 전 온난기 가설’이 있었다.

이번 연구는 고대인의 DNA를 직접 분석하여 이 논쟁에 중요한 단서를 제공했다. 분석 결과, LGM 이전의 전원인과 AR33K에서는 이 유전자 변이가 발견되지 않았지만, LGM 말기의 AR19K를 포함한 그 이후의 모든 고대인에게서는 이 변이가 발견되었다. 이는 EDAR V370A 변이가 LGM 도중이나 직후에 선택되어 동아시아인 사이에서 급격히 퍼져나갔을 가능성이 매우 높다는 것을 의미한다. 따라서 덥고 습한 환경에서 선택되었다는 ‘3만 년 전 가설’보다는, 춥고 건조한 빙하기 환경과 관련된 ‘2만 년 전 가설’이 더 설득력을 얻게 되었다.

5. 그림 설명


그림 1. 시료의 지리적 분포와 집단 구조:

(A) 지도와 연대 그래프: 왼쪽 지도는 이번 연구에 사용된 유골들이 흑룡강(黑龍江) 유역(Amur River Basin)에서 발견되었음을 보여준다. 오른쪽은 각 유골(기호로 표시)이 어느 시대에 속하는지 보여주는 연표다.

(B) 주성분 분석(PCA) 그래프: 유전적 유사성을 기준으로 인류 집단을 2차원 공간에 배치한 유전 지도다. 새로 분석된 고대인들(붉은 글씨)이 다른 동아시아 고대인 및 현대인과 어떤 관계에 있는지 보여준다.

(C) f3-통계 히트맵: 두 집단이 얼마나 많은 유전자를 공유하는지 색깔로 표시한 표다. 붉은색에 가까울수록 유전적으로 가깝다는 의미다. 가장 오래된 AR33K와 전원인(Tianyuan)이 서로 매우 가깝고(진한 붉은색 사각형), LGM 이후의 흑룡강(黑龍江) 사람들도 자기들끼리 매우 가깝다는 것을 한눈에 볼 수 있다.


그림 2. AR33K와 AR19K의 유전적 특징:

(A) 세계 지도: 가장 오래된 개체인 AR33K가 다른 고대인들과 유전적으로 얼마나 가까운지를 색으로 보여준다. 붉은색에 가까울수록 가깝다는 뜻으로, 북경(北京)의 전원인(Tianyuan) 위치에 가장 따뜻한 색의 점이 찍혀 있다.

(B) D-통계 그래프: 19,000년 전의 AR19K가 유전적으로 남쪽의 고대인(Qihe)보다 북쪽의 여러 고대인 집단(North)과 더 가깝다는 것을 통계적으로 보여주는 그래프다. 이는 약 19,000년 전에 이미 동아시아 내에서 남북 간 유전적 분리가 있었음을 시사한다.

그림 3. 고대 북부 시베리아인의 조상 찾기:

(A) qpGraph 모델: 인류 집단들의 분화와 혼혈 과정을 보여주는 복잡한 ‘가계도’다. 이 모델은 흑룡강(黑龍江) 유역의 AR14K 집단(주황색)이 고대 북유라시아인 계통(USR1)과 혼혈하여 고대 북부 시베리아인(UKY, Kolyma)을 형성했음을 가장 잘 설명한다.

(B) qpAdm 분석 결과: 고대 북부 시베리아인(UKY, Kolyma)의 유전자를 분석해보니, 약 20-30%는 AR14K(주황색)로부터, 약 70-80%는 USR1(파란색)로부터 왔다는 것을 보여준다.


그림 4. LGM 이후 흑룡강 유역 집단의 특징과 EDAR 유전자 변화:

(A) D-통계 그래프: LGM 이후의 흑룡강(黑龍江) 유역 사람들(AR14K 등)이 다른 북부 동아시아 집단보다 러시아 악마의 문 동굴(DevilsCave_N) 사람들과 유전적으로 더 가깝다는 것을 보여주며, 이 지역의 유전적 연속성을 다시 한번 확인시켜 준다.

(B) EDAR V370A 유전자 변이 연대기: 시간의 흐름(오른쪽에서 왼쪽으로 40,000년 전까지)에 따라 EDAR V370A 유전자 변이(붉은색)가 언제 나타나 얼마나 퍼졌는지를 보여준다. LGM(보라색 영역) 이전에는 조상형(파란색)만 보이다가, 19,000년 전(AR19K)부터 변이형(붉은색)이 나타나기 시작해 시간이 지날수록 점점 더 많아지는 것을 명확히 보여준다.

6. 결론 및 시사점: 동아시아 인류사의 큰 그림을 다시 그리다

이 연구는 북부 동아시아의 고대 인류 역사를 이해하는 데 있어 획기적인 전환점을 마련했다.

1. 빙하기를 전후한 두 번의 대규모 인구 교체 규명: LGM 이전에 널리 퍼져 있던 전원인 관련 집단이 LGM을 거치며 사라지고, 약 19,000년 전 현생 북부 동아시아인의 직접적인 조상인 새로운 집단으로 대체되었음을 처음으로 밝혔다. 이는 기후 변화가 인류 역사에 얼마나 극적인 영향을 미쳤는지를 보여준다.

2. 아메리카 원주민 조상의 뿌리 발견: 약 14,000년 전부터 흑룡강(黑龍江) 유역에 정착한 인류 집단이 아메리카 원주민의 가장 가까운 아시아 친척인 ‘고대 북부 시베리아인’을 형성한 핵심 동아시아 조상임을 증명했다. 이는 동아시아와 아메리카 대륙을 잇는 인류 이동의 경로를 더 명확하게 그려주었다.

3. 동아시아인 대표 유전자의 진화사 추적: 동아시아인의 신체적 특징에 영향을 미치는 EDAR V370A 유전자 변이가 LGM 직후에 급격히 확산되었음을 고대 DNA를 통해 직접 확인함으로써, 이 유전자가 선택된 시기와 환경에 대한 중요한 과학적 근거를 제시했다.

결론적으로, 이 연구는 단편적인 화석 증거만으로는 알 수 없었던 북부 동아시아의 역동적인 인구 변화와 유전적 역사를 구체적으로 복원했으며, 동아시아인을 넘어 시베리아와 아메리카 원주민의 기원을 이해하는 데 필수적인 정보를 제공했다.

 

[전문번역]

Graphical abstract
핵심 내용 요약 그림

 

내용 요약 In brief

Mao et al. demonstrate that the ancestry of the ca. 40 ka individual from Tianyuan Cave near Beijing was widespread geographically and temporally before the Last Glacial Maximum (LGM; ca. 26.5-19 ka). At the end of the LGM, the earliest northern East Asian appeared. After 14 ka, populations in the Amur region comprised the closest East Asian source known for Ancient Paleo-Siberians. Among adaptive genetic variants, EDAR V370A was likely to have been elevated to high frequency after the LGM.

이 연구는 약 4만 년 전 북경(北京) 근처 전원동굴(田園洞窟)에 살았던 사람의 유전자를 분석했다. 그 결과, 이 사람의 혈통이 가장 추웠던 마지막 빙하기(LGM, 약 26,500년~19,000년 전)가 오기 전에 이미 북부 동아시아 넓은 지역에 오랫동안 퍼져 있었다는 사실을 밝혀냈다. 빙하기가 끝나갈 무렵, 최초의 북부 동아시아인 조상이 등장했다. 14,000년 전 이후부터 아무르(Amur) 강 지역에 살던 사람들은 옛 시베리아인들의 뿌리가 된 동아시아 집단 중 가장 가까운 친척이었다. 또한, 동아시아인의 유전적 특징 중 하나인 EDAR V370A는 마지막 빙하기가 끝난 뒤에 널리 퍼졌을 가능성이 높다.

핵심 발견 Highlights

  • Tianyuan/AR33K ancestry was widespread in northern East Asia before the LGM
  • 가장 추웠던 마지막 빙하기(LGM)가 오기 전, ‘전원인(Tianyuan)’과 비슷한 유전적 뿌리를 가진 사람들이 북부 동아시아에 널리 살고 있었다.
  • The earliest northern East Asian appeared in the Amur region at the end of the LGM
  • 빙하기가 끝날 무렵, 오늘날 북부 동아시아인의 직접적인 조상이 아무르(Amur) 강 지역에서 처음으로 나타났다.
  • The AR14K-related population is the closest East Asian source for Ancient Paleo-Siberians
  • 약 14,000년 전 아무르 강 지역에 살던 사람들은 옛 시베리아 사람들의 조상 중에서 동아시아와 가장 가까운 집단이었다.
  • EDAR V370A likely increased to high frequency after the LGM
  • 동아시아인의 특징적인 유전자(EDAR V370A)는 마지막 빙하기가 끝난 후에 널리 퍼진 것으로 보인다.

마오, 샤오웨이(Xiaowei Mao)¹ , 장, 후카이(Hucai Zhang)²,³ , 차오, 시유(Shiyu Qiao)¹,⁴,⁵ , 리우, 이첸(Yichen Liu)¹,⁴ , 창, 펑친(Fengqin Chang)² , 시에, 핑(Ping Xie)² , 장, 밍(Ming Zhang)¹ , 왕, 톈이(Tianyi Wang)¹,⁶ , 리, 미안(Mian Li)¹,⁷ , 카오, 펑(Peng Cao)¹ , 양, 루오웨이(Ruowei Yang)¹ , 리우, 펑(Feng Liu)¹ , 다이, 칭옌(Qingyan Dai)¹ , 펑, 샤오톈(Xiaotian Feng)¹ , 핑, 완징(Wanjing Ping)¹ , 레이, 추자오(Chuzhao Lei)⁸ , 존 W. 올센(John W. Olsen)¹,⁹ , E. 앤드류 베넷(E. Andrew Bennett)¹ , 그리고 푸, 차오메이(Qiaomei Fu)¹,⁴,⁵,¹⁰,*

¹Key Laboratory of Vertebrate Evolution and Human Origins, Institute of Vertebrate Paleontology and Paleoanthropology, Center for Excellence in Life and Paleoenvironment, Chinese Academy of Sciences, Beijing 100044, China
¹척추동물 진화 및 인류 기원 핵심 연구소, 고척추동물 및 고인류 연구소, 생명 및 고환경 탁월성 센터, 중국과학원, 중국 북경 100044

²Institute for Ecological Research and Pollution Control of Plateau Lakes, School of Ecology and Environmental Science, Yunnan University, Kunming 650500, China
²고원 호수 생태 연구 및 오염 통제 연구소, 생태 및 환경 과학 학교, 운남대학, 중국 곤명 650500

³CAS Center for Excellence in Tibetan Plateau Earth Sciences, Chinese Academy of Sciences (CAS), Beijing 100101, China
³티베트 고원 지구 과학 CAS 탁월성 센터, 중국과학원(CAS), 중국 북경 100101

⁴Shanghai Qi Zhi Institute, Shanghai 200232, China
⁴상해 기지 연구소, 중국 상해 200232

⁵University of the Chinese Academy of Sciences, Beijing 100049, China
⁵중국과학원대학, 중국 북경 100049

⁶Northwest University, Xi’an 710069, China
⁶서북대학, 중국 서안 710069

⁷Shandong First Medical University & Shandong Academy of Medical Sciences, Jinan 250000, China
⁷산동제일의과대학 및 산동의학과학원, 중국 지난 250000

⁸Key Laboratory of Animal Genetics, Breeding and Reproduction of Shaanxi Province, College of Animal Science and Technology, Northwest A&F University, Yangling 712100, China
⁸산시성 동물 유전, 육종 및 번식 핵심 연구소, 동물 과학 기술 대학, 서북농림과기대학, 중국 양링 712100

⁹School of Anthropology, University of Arizona, Tucson, AZ 85721-0030, USA
⁹인류학과, 애리조나 대학, 미국 애리조나 주 투손 AZ 85721-0030

¹⁰Lead contact
¹⁰주요 연락처

*Correspondence: zhanghc@ynu.edu.cn (H.Z.), fuqiaomei@ivpp.ac.cn (Q.F.)

https://linkinghub.elsevier.com/retrieve/pii/S0092867421005754

요약 SUMMARY

Northern East Asia was inhabited by modern humans as early as 40 thousand years ago (ka), as demonstrated by the Tianyuan individual. Using genome-wide data obtained from 25 individuals dated to 33.6-3.4 ka from the Amur region, we show that Tianyuan-related ancestry was widespread in northern East Asia before the Last Glacial Maximum (LGM). At the close of the LGM stadial, the earliest northern East Asian appeared in the Amur region, and this population is basal to ancient northern East Asians. Human populations in the Amur region have maintained genetic continuity from 14 ka, and these early inhabitants represent the closest East Asian source known for Ancient Paleo-Siberians. We also observed that EDAR V370A was likely to have been elevated to high frequency after the LGM, suggesting the possible timing for its selection. This study provides a deep look into the population dynamics of northern East Asia.

북부 동아시아에는 전원인(田園人)의 사례에서 보듯이 이르면 4만 년 전(ka)부터 현생인류가 거주했다. 본 연구는 아무르(Amur) 지역에서 발견된 33,600년에서 3,400년 전 사이의 25개 개체로부터 얻은 게놈 전체 데이터를 사용하여, 전원인(田園人) 관련 계통이 최종빙기절정기(LGM) 이전에 북부 동아시아에 널리 퍼져 있었음을 보여준다. 최종빙기절정기가 끝날 무렵, 아무르(Amur) 지역에서 최초의 북부 동아시아인이 나타났으며, 이 집단은 후대 고대 북부 동아시아인의 기층(basal)을 이룬다. 아무르(Amur) 지역의 인구는 14,000년 전부터 유전적 연속성을 유지했으며, 이 초기 거주자들은 고대 고시베리아인(Ancient Paleo-Siberians)의 가장 가까운 동아시아 조상으로 밝혀졌다. 또한 EDAR V370A 유전자 변이가 최종빙기절정기 이후에 높은 빈도로 증가했음을 관찰했고, 이는 해당 유전자가 선택된 시기를 추정하게 한다. 이 연구는 북부 동아시아의 인구 역학에 대한 깊이 있는 통찰을 제공한다.

 

목차

서론 INTRODUCTION

결과 및 토론 RESULTS AND DISCUSSION

샘플 채취와 고대 DNA 분석 과정 Samples and ancient DNA generation

최종빙기절정기(LGM) 이전, 지리적·시간적으로 널리 퍼져 있던 전원인(田園人) 관련 계통 Tianyuan-related ancestry was widespread geographically and temporally before the LGM

최종빙기절정기 말, 최초의 북부 동아시아인 등장 The appearance of the earliest northern East Asian at the end of the LGM

최종빙기절정기 이후 아무르(Amur) 지역의 유전적 연속성과 인구 상호작용 Genetic continuity in the Amur region and population interactions after the LGM

4만 년 전부터 6천 년 전까지 북부 동아시아인의 적응성 유전자 변이 Adaptive genetic variants in northern East Asians from 40 to 6 ka

연구의 한계 Limitations of the study

연구 방법 STAR METHODS

보충 정보 SUPPLEMENTAL INFORMATION

감사의 글 ACKNOWLEDGMENTS

저자 기여 AUTHOR CONTRIBUTIONS

이해관계 선언 DECLARATION OF INTERESTS

보충 인용 SUPPORTING CITATIONS

참고문헌 REFERENCES

자원 이용 가능성 RESOURCE AVAILABILITY

주요 연락처 Lead contact

물질 이용 가능성 Materials availability

데이터 및 코드 이용 가능성 Data and code availability

실험 모델 및 대상 상세 정보 EXPERIMENTAL MODEL AND SUBJECT DETAILS

유적지 및 표본 설명Sites and specimen descriptions

방법 상세 정보 (METHOD DETAILS)

고대 DNA 추출 Ancient DNA extraction

고대 DNA 포획 및 시퀀싱 Ancient DNA capture and sequencing

정량화 및 통계 분석 QUANTIFICATION AND STATISTICAL ANALYSIS

리드 정렬 Read alignment

오염 평가 및 변이 콜링 Contamination evaluation and variants calling

친족 관계 분석 Kinship analyses

주성분 분석 Principal components analysis (PCA)

ADMIXTURE 분석

아웃그룹-f3 및 D 통계 Outgroup-f3 and D statistics

Treemix를 이용한 계통 모델링 Phylogeny modeling with Treemix

qpAdm을 이용한 혼합 모델링 Admixture modeling with qpAdm

qpGraph를 이용한 인구 통계 모델링 Demographic modeling with qpGraph

은닉 마르코프 모델을 이용한 고인류 계통 추정 Archaic ancestry estimation with the Hidden Markov Model

아무르(Amur) 지역의 시간 경과에 따른 인구 크기 Time transect of population sizes in the Amur region

북부 동아시아인의 EDAR V370A 시간 경과 Time transect of EDAR V370A in northern East Asians

 

서론 INTRODUCTION

Anatomically and behaviorally modern humans were present in northern East Asia as early as 40 thousand years ago (ka) (Shang et al., 2007). The non-static definition of northern East Asia (Li and Cribb, 2014) includes the Mongolian plateau, northern China (Li et al., 2015), the Japanese archipelago, the Korean peninsula, and the mountainous regions of the Russian Far East, all of which fall within a similar latitude range as that of central and southern Europe. In Europe, human population movements and size were influenced by Ice Age climatic fluctuations (Fu et al., 2016; Tallavaara et al., 2015); e.g., the population size varied (ranging from 130,000-410,000) between 30 and 13 ka (Tallavaara et al., 2015). These climatic oscillations may have had a similar effect on the population history of high-latitude and high-elevation regions in Asia, as evidenced by archaeological finds that suggest gaps in human occupation in Mongolia during marine isotope stages 3 (beginning ~57 ka) and 2 (beginning ~29 ka) (Rybin et al., 2016). Although several studies have advanced our understanding of ancient East Asians from 9 ka to historical times (Ning et al., 2020; Wang et al., 2021; Yang et al., 2020), only two human genomes earlier than the Last Glacial Maximum (LGM; approximately 26.5-19 ka) (Clark et al., 2009) have been reported (~40 ka from Tianyuan Cave in northern China; ~34 ka from Salkhit in northeastern Mongolia) (Devièse et al., 2019; Massilani et al., 2020; Yang et al., 2017). Thus, the population structure of northern East Asia before and after the LGM stadial is largely unknown, and the lack of human fossils dating to between 40-9 ka from this area hinders our understanding of this crucial period of human prehistory.

해부학적으로나 행동적으로 현생인류는 이르면 4만 년(ka) 전에 북부 동아시아에 존재했다 (샹(Shang) et al., 2007). 북부 동아시아의 범위는 고정되어 있지 않지만, 보통 몽골 고원, 중국 북부, 일본 열도, 한반도, 그리고 러시아 극동의 산악 지역을 포함한다 (리(Li) and 크립(Cribb), 2014; 리(Li) et al., 2015). 이 모든 지역은 중부 및 남부 유럽과 비슷한 위도에 있다. 유럽에서는 빙하기의 기후 변동이 인구 이동과 규모에 영향을 미쳤다 (푸(Fu) et al., 2016; 탈라바라(Tallavaara) et al., 2015). 예를 들어, 3만 년에서 1만 3천 년 전 사이 인구 규모는 13만 명에서 41만 명 사이에서 변동했다 (탈라바라(Tallavaara) et al., 2015). 이러한 기후의 오르내림은 아시아의 고위도 및 고지대 지역의 인구 역사에도 비슷한 영향을 미쳤을 수 있다. 이를 뒷받침하는 고고학적 증거로, 몽골(Mongolia)에서는 해양 동위원소 스테이지 3(약 5만 7천 년 전 시작)과 2(약 2만 9천 년 전 시작) 동안 인류 거주가 중단된 공백기가 있었던 것으로 보인다 (리빈(Rybin) et al., 2016). 최근 여러 연구 덕분에 9천 년 전부터 역사 시대까지의 고대 동아시아인에 대한 이해가 깊어졌다 (닝(Ning) et al., 2020; 왕(Wang) et al., 2021; 양(Yang) et al., 2020). 하지만 최종빙기절정기(LGM, 약 26,500년-19,000년 전)보다 이른 시기의 인간 게놈은 단 두 개만 보고되었을 뿐이다 (클락(Clark) et al., 2009). 하나는 중국 북부 전원동굴(田園洞窟)에서 나온 약 4만 년 전 개체이고, 다른 하나는 몽골(Mongolia) 북동부 살키트(Salkhit)에서 나온 약 3만 4천 년 전 개체이다 (드비에즈(Devièse) et al., 2019; 마실라니(Massilani) et al., 2020; 양(Yang) et al., 2017). 따라서 최종빙기절정기 전후 북부 동아시아의 인구 구조는 거의 알려지지 않았다. 또한 이 지역에서 4만 년에서 9천 년 전 사이의 인류 화석이 부족하여, 인류 선사시대의 이 중요한 시기를 이해하는 데 어려움이 있다.

 

결과 및 토론 RESULTS AND DISCUSSION

샘플 채취와 고대 DNA 분석 과정 Samples and ancient DNA generation

To examine deep population histories in northern East Asia, genome-wide genotype data were generated from 25 ancient humans from the Songnen Plain (Heilongjiang province, northeastern China) in the Amur region, associated with direct radiocarbon dates ranging from 33,590-3,420 calibrated years before present (cal BP). These individuals were recovered at construction sites without archaeological context or from eroded gullies and small valleys formed in alluvial deposits of the Songhua River and Nen River during summer rainy seasons. DNA was extracted from less than 100 mg of bone powder and used to construct single-stranded libraries that were enriched for human DNA using a panel of 1.2 million single-nucleotide polymorphisms (SNPs). Enriched libraries were sequenced, resulting in SNP panel coverage from 0.002-14.45x (Table 1). All samples showed low levels of modern human contamination (Table 1). Removing close relatives (defined as first- or second-degree relationships) (Table S1) and individuals with less than 25,000 SNPs resulted in a final dataset of 20 unrelated individuals (two individuals with SNP coverage lower than 50,000 are denoted with the suffix “_LowCov”) (Table 1). Our final samples were combined into five main chronological groups, with similar genomic profiles within each group treated as populations: (1) one upper Pleistocene individual from before the LGM (denoted AR33K; 34,324-32,360 cal BP; AR represents Amur region, followed by the approximate age in thousands of years [K]), (2) one upper Pleistocene individual dating to the end of the LGM (denoted AR19K; 19,587-19,175 cal BP), (3) two individuals dating to the period between 14,932 and 14,017 cal BP (denoted AR14K) at the end of the last glacial period, (4) four individuals from 12,735-10,302 cal BP (denoted AR13-10K), and (5) 12 individuals from 9,425-3,360 cal BP (denoted AR9.2K_o, ARpost9K (n=9) AR7.3K_LowCov, and AR3.4K_LowCov)(Figure 1A; Figure S1; Table 1; Table S1).

북부 동아시아의 아주 오래전 인구 역사를 연구하기 위해, 아무르(Amur)강 유역의 송눈평원(松嫩平原, 중국 흑룡강성)에서 발견된 25명의 고대 사람 뼈에서 유전자 전체(게놈) 데이터를 뽑아냈다. 이 뼈들의 주인은 방사성탄소 연대 측정 결과, 약 33,590년에서 3,420년 전에 살았던 사람들이다. 이 유해들은 정식 발굴 현장이 아닌, 고고학적 정보가 없는 공사 현장에서 발견되었다. 또는 여름 장마철에 송화강(松花江)과 눈강(嫩江) 주변의 흙이 깎여나가면서 드러난 작은 골짜기에서 수습되기도 했다. 뼈 가루를 100mg보다 적게 사용하여 DNA를 뽑아냈고, 이를 이용해 유전자 분석을 위한 준비 작업(단일 가닥 라이브러리 제작)을 했다. 그리고 120만 개의 특정 유전자 표지(SNP, 사람마다 미세하게 다른 유전자 정보)를 이용해 사람의 DNA만 골라 양을 늘렸다. 이렇게 양을 늘린 DNA를 해독했더니, 분석 대상인 유전자 표지(SNP)가 0.002배에서 14.45배의 양으로 확인되었다. 다행히 모든 샘플에서 현대 사람의 DNA로 인한 오염은 매우 적었다. 분석의 정확도를 높이기 위해, 가까운 친척(1촌 또는 2촌) 관계인 사람들과 유전자 정보가 너무 적은(SNP 25,000개 미만) 사람도 제외했다. 그 결과 최종적으로 서로 친척이 아닌 20명의 데이터를 분석 대상으로 삼았다. (이 중 2명은 유전자 정보량이 5만 개보다도 적어서 이름 뒤에 ‘_LowCov’라고 표시했다.) 최종적으로 선정된 20명의 샘플은 살았던 시기에 따라 크게 다섯 그룹으로 나누었다. 각 그룹 안의 사람들은 유전적으로 비슷해서 하나의 인구 집단으로 볼 수 있다.

(1) 가장 추웠던 마지막 빙하기(LGM) 이전에 살았던 후기 홍적세 사람 1명이다 (AR33K 그룹; 약 34,324년 ~ 32,360년 전). (AR은 아무르(Amur) 지역, 뒤의 숫자는 대략적인 연대(천 년 단위)를 의미한다).

(2) 빙하기가 끝날 무렵에 살았던 후기 홍적세 사람 1명이다 (AR19K 그룹; 약 19,587년 ~ 19,175년 전).

(3) 마지막 빙하기가 끝난 직후에 살았던 사람 2명이다 (AR14K 그룹; 약 14,932년 ~ 14,017년 전).

(4) 약 12,735년에서 10,302년 전에 살았던 사람 4명이다 (AR13-10K 그룹).

(5) 약 9,425년에서 3,360년 전에 살았던 사람 12명이다. 이 그룹은 세부적으로 AR9.2K_o (1명), ARpost9K (9명), AR7.3K_LowCov (1명), 그리고 AR3.4K_LowCov (1명)으로 구성된다(그림 1A와 보충 그림 S1, 그리고 표 1과 보충 표 S1에서 더 자세히 확인할 수 있다.).

Table 1. Newly sampled individuals included in this study
표 1. 본 연구에 포함된 새로 샘플링된 개체들

개체 ID (Individual ID) 성별 (Sex) 샘플 ID (Sample ID) 인구 ID (Population ID) 연대 (cal BP)a 위도 (Latitude) 경도 (Longitude) 라이브러리 유형 (Library style)b mtDNA 5’C T (%) 1240K 5’C-T (%) mtDNA 오염도 (mtDNA contamination)c MT 하플로그룹 (MT_hg)d Y 하플로그룹 (Y_hg)e mtDNA 커버리지 (mtDNA Coverage) 1240K 커버리지 (1240K Coverage) SNP
NE20 여 (F) AR33K AR33K 34,324-32,360 45.4 127.03 SS 27 22 1.6 B 574 14.45 1,010,948
NE56 남 (M) AR19K AR19K 19,587-19,175 46 126.15 SS 34 19 0.5 G2 C2 628 1.583 768,551
NE34 남 (M) AR14.5K AR14K 14,932-14,176 45.98 125.58 SS 27 22 0.9 D4h3 C 488 0.716 332,628
NE-5 남 (M) AR14.1K AR14K 14,814-14,017 45.95 125.83 SS 25 19 1.1 D4h3a+@152 C 503 2.212 520,883
NE-1 여 (F) AR12K AR13-10K 12,735-12,486 46 125.73 SS 35 8 0.3 D4g 554 0.083 91,615
NE-8 남 (M) AR11K AR13-10K 11,601-11,176 46.01 125.82 SS 34 27 1.5 D40 DE 274 0.421 225,066
NE57f 남 (M) AR11K deleted 11,206-10,765 46 126.12 DS-half 6 1 1.9 3 0.01 11,614
NE36f 불명 (U) AR10.6K_deleted 11,065-10,513 46.06 125.72 SS 16 4 0.1 D4h4 47 0.007 8,926
NE-3 남 (M) AR10.6K AR13-10K 10,996-10,429 45.85 126 SS 33 27 1.4 M8 C 419 0.692 348,382
NE-4 남 (M) AR10.5K AR13-10K 10,740-10,302 45.93 125.85 SS 32 26 1.8 M8 C 396 0.671 332,715
NE30g 남 (M) AR9.9K (AR10.5K의 2촌, 제외됨) 10,167-9,676 45.96 126.23 SS 20 12 0.8 D4m C 99 0.225 218,945
NE45 남 (M) AR9.2K_o AR9.2K_o 9,425-9,029 45.71 126.88 SS 17 7 1.3 Glala P 163 0.228 134,344
NE44f 남 (M) AR9.2K_deleted 9,425-9,027 45.71 126.88 SS 22 3 2.4 D4c1 25 0.004 5,129
NE35 여 (F) AR8.9K ARpost9K 9,131-8,770 45.91 125.95 SS 24 22 1 R11 602 4.664 739,026
NE-16 여 (F) AR8.5K ARpost9K 8,723-8,421 45.96 125.78 SS 23 14 1.2 D40 271 0.534 216,989
NE49 여 (F) AR8.3K ARpost9K 8,425-8,204 45.65 126.92 SS 20 10 1.5 D40 59 0.21 197,627
NE39 남 (M) AR8.1K ARpost9K 8,340-8,029 45.91 125.9 SS 15 12 1.5 D4m2 C 526 0.93 345,404
NE58h 남 (M) AR7.3K_LowCov 7,425-7,168 45.85 126 DS-half 9 1 0.6 D4e5 163 0.026 30,052
NE-18f AR7K_deleted 7,245-6,894 45.96 125.77 SS 22 2 0.2 D4h3 158 0.002 2,511
NE19 남 (M) AR7K ARpost9K 7,167-6,854 45.96 125.75 SS 28 23 2.4 F1b1+@152 C 489 0.698 349,755
NE29 여 (F) AR6.87K ARpost9K 6,993-6,747 45.93 126.37 SS 16 13 0.8 D4c1 528 1.074 393,142
NE-22 남 (M) AR6.84K ARpost9K 6,986-6,676 46.15 126.23 SS 20 11 1.2 D4m C 198 0.374 183,780
NE-9 여 (F) AR6.33K ARpost9K 6,440-6,205 46.01 125.8 SS 26 23 2.7 C4a1a+195 569 1.097 410,575
NE-2 남 (M) AR6.32K ARpost9K 6,437-6,201 45.85 125.82 SS 19 12 1.2 D4m C 266 0.213 173,120
NE61h 불명 (U) AR3.4K_LowCov 3,485-3,360 45.88 126.22 SS 25 5 0.1 D4b1a2a 574 0.022 26,464

See also Table S1.
표 S1도 참조하라.

aDate (cal BP), calibrated date with 99.7% confidence interval using OxCal v.4.4 (Ramsey and Lee, 2013) and the IntCal20 calibration curve (Reimer et al., 2020).
a연대(cal BP)는 OxCal v.4.4와 IntCal20 보정 곡선을 사용하여 99.7% 신뢰 구간으로 보정한 연대이다.

bSS, single-stranded DNA library; DS-half, double-stranded DNA library with half uracil-DNA-glycosylase (UDG) treatment.
bSS는 단일 가닥 DNA 라이브러리; DS-half는 반쪽 우라실-DNA-글리코실레이즈(UDG) 처리 이중 가닥 DNA 라이브러리이다.

cContamination was estimated as a percentage.
c오염도는 백분율로 추정했다

dMT_hg, mitochondrial haplotype group.
dMT_hg는 미토콘드리아 하플로그룹이다.

eY_hg, Y-chromosomal haplotype group.
eY_hg는 Y-염색체 하플로그룹이다.

fNE57, NE36, NE44 and NE-18 had SNP numbers of less than 25,000.
fNE57, NE36, NE44, NE-18은 SNP(단일염기 다형성) 수가 25,000개 미만이었다.

gNE30 had second-degree kinship with NE-4.
gNE30은 NE-4와 2촌 관계였다.

hNE58 and NE61 had SNP numbers of less than 50,000.
hNE58과 NE61은 SNP 수가 50,000개 미만이었다.

Figure 1. Geographic and temporal distribution and population structure of newly sampled and published populations in northern East Asia
그림 1. 북부 동아시아에서 새로 샘플링된 인구 및 기존에 발표된 인구의 지리적·시간적 분포와 인구 구조

(A) Geographic and temporal distribution of newly sampled and published populations. Symbols as in (B).

(A) 새로 샘플링된 인구 및 기존 발표된 인구의 지리적·시간적 분포. 기호는 (B)와 동일하다.

(B) PCA of sampled and published East Asian and Siberian populations. The ancient populations (above) are projected on the principal components of modern populations (below). New samples are highlighted by red text and larger symbols.

(B) 샘플링된 동아시아 및 시베리아 인구와 기존 발표된 인구의 주성분 분석(PCA). 고대 인구(위쪽)는 현대 인구(아래쪽)의 주성분 위에 투영되었다. 새로운 샘플은 붉은색 텍스트와 더 큰 기호로 강조 표시했다.

(C) The pairwise genetic affinity among ancient East Asians using outgroup-f3 analysis. New samples are highlighted in red, and dashed rectangles show high genetic affinities.

(C) 아웃그룹-f3 분석(outgroup-f3 analysis)을 사용한 고대 동아시아인 간의 쌍별 유전적 친연성. 새로운 샘플은 붉은색으로 강조 표시했으며, 점선 사각형은 높은 유전적 친연성을 보여준다.

See also Figure S1 and Tables 1 and S1.

그림 S1 및 표 1, S1도 참조하라.

최종빙기절정기(LGM) 이전, 지리적·시간적으로 널리 퍼져 있던 전원인(田園人) 관련 계통

Tianyuan-related ancestry was widespread geographically and temporally before the LGM

Upper Pleistocene human population dynamics prior to the LGM stadial are largely unknown in East Asia because of limited fossil remains from this period (Bae et al., 2017). To better understand the early settlement of modern humans in East Asia, especially northern East Asia, we investigated the genomic profile of the pre-LGM (33.6 ka) AR33K specimen, a female individual. We first surveyed how this AR33K individual related to other ancient and modern populations globally using outgroup-f3 analysis (Patterson et al., 2012). Interestingly, outgroup-f results showed that AR33K and the 7,000-years-older Tianyuan individual share more genetic drift with each other than with all other tested populations (Figures 1C and 2A). Symmetry tests using D statistics (Patterson et al., 2012) also show that AR33K and Tianyuan form a cluster relative to all other ancient and modern populations globally (Figure S2), which is also supported by results from qpGraph (Patterson et al., 2012; Figure 3A) and the maximum-likelihood tree (Pickrell and Pritchard, 2012; Figure S3). Additionally, these analyses placed AR33K and Tianyuan basal to all East Asians (Figure 3A; Figure S3), and analyses for detecting archaic introgressions (Peter, 2020) show that they exhibit an excess of Denisovan ancestry compared with younger populations (Table S2), similar to that reported by Massilani et al. (2020) . It has been shown previously that the Tianyuan individual shared more alleles with an ~35,000-year-old individual from Belgium (GoyetQ116-1; Fu et al., 2016) than with other contemporary western Eurasians and shared more alleles with modern Amazonian populations (i.e., Surui) than other Native Americans (Yang et al., 2017). Despite the high genetic similarity between AR33K and Tianyuan, we find that AR33K does not share the same genetic affinity as Tianyuan to GoyetQ116-1. Instead, AR33K presents a profile similar to other ancient East Asians, showing a trend of elevated GoyetQ116-1 ancestry compared with other ancient western Eurasians, but just at or below the cutoff for statistical significance (3 standard deviations from the mean). We find no evidence that AR33K shares more alleles with Surui than other Native Americans (Table S3). Thus, a complex population structure may already have existed before the LGM stadial in northern East Asia.

후기 홍적세, 즉 최종빙기절정기(LGM) 이전 동아시아의 인구 역학은 거의 알려지지 않았다. 이 시기 화석 유물이 매우 적기 때문이다. 그래서 우리는 북부 동아시아 현생인류의 초기 정착을 더 잘 이해하고자 했다. LGM 이전인 33,600년 전 여성 개체 AR33K의 게놈 프로필을 조사했다. 먼저 AR33K 개체가 다른 고대 및 현대 인구와 어떤 관계인지 아웃그룹-f3 분석으로 조사했다. 흥미롭게도, AR33K와 그보다 7,000년 더 오래된 전원인(田園人)은 다른 모든 집단보다 서로 더 많은 유전적 부동(genetic drift)을 공유했다. (그림 1C와 2A) . D 통계(D statistics)를 이용한 대칭성 검사도 이 결과를 뒷받침했다. AR33K와 전원인은 전 세계 다른 모든 고대 및 현대 인구 집단과 구별되는 하나의 그룹(cluster)을 형성했다. (그림 S2) . qpGraph와 최대우도계통수 분석 결과도 동일했다. (그림 3A; 그림 S3) . 또한 이 분석들은 AR33K와 전원인을 모든 동아시아인의 기층(basal)에 두었다. (그림 3A; 그림 S3) . 고인류 유전자 유입 분석 결과, 이들은 후대 인구보다 데니소바인(Denisovan) 계통을 더 많이 가졌다. (표 S2) . 이는 마실라니(Massilani) 등의 보고와 유사하다. 이전 연구에서 전원인은 약 35,000년 전 벨기에(Belgium) 개체(GoyetQ116-1)와 더 많은 대립유전자를 공유하는 것으로 나타났다. 또한 다른 아메리카 원주민보다 현대 아마존(Amazonian) 인구(수루이(Surui)족)와 더 많은 대립유전자를 공유했다. AR33K와 전원인은 유전적으로 매우 비슷하다. 하지만 AR33K는 전원인처럼 고예(Goyet)Q116-1 개체와 동일한 유전적 친연성을 공유하지 않았다. 대신 AR33K는 다른 고대 동아시아인과 유사한 프로필을 보였다. 다른 고대 서부 유라시아인에 비해 고예(Goyet)Q116-1 계통이 높은 경향을 보였으나, 통계적 유의성 기준치 이하였다. (통계적 유의성 기준: 평균에서 3 표준편차) . AR33K가 다른 아메리카 원주민보다 수루이족과 더 많은 대립유전자를 공유한다는 증거는 없었다. (표 S3) . 따라서 LGM 이전에 북부 동아시아에는 이미 복잡한 인구 구조가 존재했을 수 있다.

AR33K (~33 ka) is roughly contemporary with Salkhit (~34 ka) from the Salkhit Valley (northeastern Mongolia) and located at a similar distance from Salkhit (~1,159 km) (Devièse et al., 2019; Massilani et al., 2020) and Tianyuan cave (~1,114 km) (Yang et al., 2017). The Salkhit individual has been proposed recently as an admixture of Tianyuan-related (~75%) and Yana-related (~25%) ancestries (Massilani et al., 2020). Similar to Tianyuan, Salkhit also shows a connection with GoyetQ116-1 (Massilani et al., 2020). However, unlike Salkhit, AR33K does not possess more Yana-related ancestry than Tianyuan (D(Salkhit, Tianyuan; Yana, Mbuti) > 0 [Z=7.55]; D(Tianyuan, AR33K; Yana, Mbuti) ~0 [Z=1.88]). This probably indicates that Tianyuan-related ancestry was widespread before the LGM stadial in northern East Asia, geographically (from the North China Plain near present-day Beijing to the Salkhit Valley in northeastern Mongolia, and the Amur region of northeastern China) and temporally (from 40-33 ka). People with Tianyuan-related ancestry may have admixed with people of Yana-related ancestry in Mongolia while remaining isolated in the Amur region before the LGM.

AR33K(약 33ka)는 몽골(Mongolia) 북동부 살키트(Salkhit) 계곡의 살키트인(약 34ka)과 거의 동시대 사람이다. AR33K 발견지는 살키트에서 약 1,159km, 전원동굴(田園洞窟)에서 약 1,114km 떨어져 있다. 살키트인은 최근 전원인 관련 계통(약 75%)과 야나(Yana) 관련 계통(약 25%)의 혼합으로 제안되었다. 살키트인도 전원인처럼 고예(Goyet)Q116-1 개체와 연관성을 보였다. 하지만 AR33K는 살키트인과 달랐다. AR33K는 전원인보다 야나 관련 계통을 더 많이 갖고 있지 않았다. (통계 분석 결과: D(살키트, 전원인; 야나, 음부티)는 0보다 크고 [Z=7.55], D(전원인, AR33K; 야나, 음부티)는 0에 가까웠다 [Z=1.88]. 이는 살키트인에게서는 야나 유전자가 뚜렷하지만 AR33K에게서는 그렇지 않다는 의미다.). 이는 아마도 전원인 관련 계통이 LGM 이전에 북부 동아시아에 널리 퍼져 있었음을 나타낸다. 지리적으로는 오늘날의 북경(北京) 근처 북중국 평원에서 몽골 북동부 살키트 계곡, 그리고 중국 동북부의 아무르(Amur) 지역까지 해당한다. 시간적으로는 4만 년 전부터 3만 3천 년 전까지다. 전원인 관련 계통을 가진 사람들은 몽골에서는 야나 관련 계통의 사람들과 섞였을 수 있다. 반면 아무르 지역에서는 LGM 이전까지 고립된 상태를 유지했을 것이다.

최종빙기절정기 말, 최초의 북부 동아시아인 등장

The appearance of the earliest northern East Asian at the end of the LGM

The LGM stadial had a large effect on prehistoric population dynamics in Europe (Fu et al., 2016; Tallavaara et al., 2015), however, little is known about how the profound climatic changes of this period affected populations in northern East Asia. Here we report genetic evidence from an East Asian individual who lived during the LGM (26.5-19 ka), AR19K, which is critical for understanding the population dynamics of northern East Asians at the end of this harsh period. We first assessed AR19K using principal component analysis (PCA) (Figure 1B) and outgroup-f3 analyses (Figure 1C), showing that AR19K has more genetic affinities with later East Asians compared with the earlier East Asians. The AR33K and Tianyuan individuals share a specific genetic drift that is not shared by AR19K and younger East Asians, a relationship supported by qpGraph modeling (Figure 3A) and Treemix (Figure S3). This specific drift is also reflected by D(Tianyuan, AR33K; AR19K/AR14K/AR13-10K/ARpost9K, Mbuti) ~0 (-0.6 < Z < 1), D(AR19K/AR14K/AR13-10K/ARpost9K, Tianyuan; AR33K, Mbuti) < 0 (-11.1 < Z < -10.2), and D(AR19K/AR14K/AR13-10K/ARpost9K, AR33K; Tianyuan, Mbuti) <0(-11 <Z<-9.4) (Figure S2). These observations may suggest that a population change in northern East Asia had already occurred at the end of the LGM, earlier than the recently proposed post-LGM population change demonstrated by a 16,900-year-old subadult female (Khaiyrgas-1) found in Khaiyrgas Cave in Yakutia (Russia) (Kılınç et al., 2021). This proposed population change may have occurred even earlier in the LGM because genetic continuity between 22 ka and 18 ka has been noted in individuals from the Mal’ta and Afontova Gora sites farther west in Siberia (Raghavan et al., 2014). Additional genetic and archaeological evidence during the LGM stadial in the Amur River Basin is needed to resolve the timing of the appearance of the lineage represented by AR19K and later individuals in this region.

최종빙기절정기(LGM)는 유럽의 선사시대 인구 역학에 큰 영향을 미쳤다 (푸(Fu) et al., 2016; 탈라바라(Tallavaara) et al., 2015). 하지만 이 시기의 엄청난 기후 변화가 북부 동아시아 인구에 어떤 영향을 미쳤는지는 거의 알려지지 않았다. 여기서 우리는 최종빙기절정기(26,500-19,000년 전) 동안 살았던 동아시아인 개체 AR19K의 유전적 증거를 보고한다. 이는 이 혹독한 시기 말 북부 동아시아의 인구 역학을 이해하는 데 매우 중요하다. 우리는 먼저 주성분 분석(PCA) (그림 1B)과 아웃그룹-f3 분석(outgroup-f3 analyses) (그림 1C)을 사용하여 AR19K를 평가했다. 그 결과, AR19K는 그 이전의 동아시아인들보다 후대의 동아시아인들과 더 많은 유전적 친연성을 보이는 것으로 나타났다. AR33K와 전원인(田園人) 개체들이 공유하는 특정한 유전적 부동(genetic drift)은 AR19K와 그보다 젊은 동아시아인들에게서는 공유되지 않았다. 이 관계는 qpGraph 모델링(그림 3A)과 Treemix(그림 S3) 분석으로도 뒷받침된다. 이러한 특정 유전적 부동은 D 통계 분석 결과에서도 반영된다. (통계 수치는 AR33K와 전원인(田園人)이 유전적으로 매우 가까우며, AR19K를 비롯한 후대 동아시아인들과는 뚜렷하게 구별된다는 점을 보여준다. 상세 통계: D(전원인, AR33K; 후대 동아시아인, 음부티)는 0에 가깝고 (-0.6 < Z < 1), D(후대 동아시아인, 전원인; AR33K, 음부티)는 0보다 훨씬 작으며 (-11.1 < Z < -10.2), D(후대 동아시아인, AR33K; 전원인, 음부티) 또한 0보다 훨씬 작았다 (-11 < Z < -9.4)) (그림 S2). 이러한 관찰은 북부 동아시아의 인구 변화가 최종빙기절정기 말에 이미 일어났음을 시사할 수 있다. 이는 최근 러시아(Russia) 야쿠티아(Yakutia)의 카이르가스(Khaiyrgas) 동굴에서 발견된 16,900년 전 아성체 여성(Khaiyrgas-1)을 통해 제안된 LGM 이후의 인구 변화보다 더 이른 시점이다 (클른치(Kılınç) et al., 2021). 제안된 이 인구 변화는 LGM 기간 중 훨씬 더 일찍 일어났을 수도 있다. 왜냐하면 더 서쪽인 시베리아(Siberia)의 말타(Mal’ta)와 아폰토바 고라(Afontova Gora) 유적지의 개체들에게서 2만 2천 년 전과 1만 8천 년 전 사이의 유전적 연속성이 확인되었기 때문이다 (라가반(Raghavan) et al., 2014). AR19K와 그 후손으로 대표되는 계통이 이 지역에 나타난 정확한 시점을 확정하기 위해서는, 최종빙기절정기(LGM) 시기 아무르강 유역의 유전학적, 고고학적 증거가 추가로 필요하다.

It has been demonstrated recently that, by 8 ka at the latest, ancient coastal northern East Asian (aCNEA; including populations from the Chinese sites of Bianbian, Boshan, Xiaogao, and Xiaojingshan) and ancient coastal southern East Asian (aCSEA; including individuals from the Chinese sites of Qihe, Liangdao1, Liangdao2, Xitoucun, and Tanshishan) populations could already be separated into two distinct northern and southern genetic groups (Yang et al., 2020). We found that AR19K, along with the other newly sampled individuals reported here (from 14-6 ka), was genetically closer to aCNEA than to aCSEA (represented by Qihe) using D(aCNEA, Qihe; AR19K, Mbuti) (Figure 2B). The balanced D(AR19K, Qihe; aCNEA, Mbuti) for Xiaojingshan and Bianbian could be explained by their possessing relatively more aCSEA ancestry than other aCNEAS (Table S3). Interestingly, AR19K is found to be basal to ancient northern East Asians younger than 14 ka in admixture graph modeling using qpGraph and Treemix (Figure 3A; Figure S3). These results reveal the existence of a north-south separation as early as 19 ka, approximately 10,000 years earlier than observed previously (Yang et al., 2020). In addition, ancient northern East Asians younger than 14 ka share less genetic drift with AR19K than AR14K (Table S3), which is consistent with the basal position of AR19K for aCNEA. The analyses of AR19K, who lived toward the end of the LGM stadial, reveal this individual to be the earliest northern East Asian yet identified, thus having some degree of continuity with present-day northern East Asians and possessing a genetic ancestry distinct from that of the modern humans who occupied this region prior to the LGM (e.g., Tianyuan and AR33K).

최근 연구에 따르면, 늦어도 8천 년 전까지는 고대 해안 북부 동아시아인(aCNEA; 중국의 변변(邊邊, Bianbian), 박산(博山, Boshan), 소고(曉高, Xiaogao), 소형산(小荊山, Xiaojingshan) 유적지 인구 포함)과 고대 해안 남부 동아시아인(aCSEA; 중국의 기하(奇和, Qihe), 양도(亮島, Liangdao)1, 양도(亮島, Liangdao)2, 서두촌(西頭村, Xitoucun), 담석산(曇石山, Tanshishan) 유적지 개체 포함) 인구가 이미 뚜렷한 북부와 남부의 두 유전적 그룹으로 나뉠 수 있었다 (양(Yang) et al., 2020). 우리는 AR19K가 여기서 보고된 다른 새로운 샘플 개체들(1만 4천 년-6천 년 전)과 함께, 유전적으로 aCSEA(기하인(奇和人)으로 대표됨)보다 aCNEA에 더 가깝다는 것을 D(aCNEA, 기하; AR19K, 음부티) 통계량을 사용하여 발견했다 (그림 2B). 소형산인(小荊山人)과 변변인(邊邊人)에 대한 D(AR19K, 기하; aCNEA, 음부티) 통계량이 균형을 이루는 것은, 이들이 다른 aCNEA 집단보다 상대적으로 더 많은 aCSEA 계통을 가지고 있기 때문으로 설명될 수 있다 (표 S3). 흥미롭게도, qpGraph와 Treemix를 사용한 혼합 그래프 모델링에서 AR19K는 1만 4천 년 전보다 젊은 고대 북부 동아시아인들의 기층(basal)에 위치하는 것으로 밝혀졌다 (그림 3A; 그림 S3). 이 결과들은 약 19,000년 전에 이미 남북 간 분리가 존재했음을 보여주며, 이는 이전에 관찰된 것보다 약 1만 년 이른 시점이다 (양(Yang) et al., 2020). 또한, 14,000년보다 젊은 고대 북부 동아시아인들은 AR19K보다 AR14K와 공유하는 유전적 특징이 더 적었다 (표 S3). 이는 AR19K가 고대 해안 북부 동아시아인(aCNEA)의 뿌리가 되는 위치에 있다는 점과 일치한다. 최종빙기절정기 말기에 살았던 AR19K에 대한 분석은 이 개체가 지금까지 확인된 가장 이른 시기의 북부 동아시아인임을 보여준다. 따라서 현존 북부 동아시아인과 어느 정도 연속성을 가지며, 최종빙기절정기 이전에 이 지역을 차지했던 현생인류(예: 전원인(田園人), AR33K)와는 다른 유전적 계통을 소유하고 있다.

Figure 2. Genetic characteristics of AR33K and AR19K
그림 2. AR33K와 AR19K의 유전적 특징

(A) Shared genetic drift between AR33K and published ancient populations older than 5 ka, using outgroup-f3 analysis. A higher value (represented by warmer colors) indicates a higher shared drift with AR33K. The intersection of the dotted lines indicates the location of AR33K.

(A) 아웃그룹-f3 분석을 사용하여 AR33K와 5천 년보다 오래된 기존 발표 고대 인구 간의 공유된 유전적 부동(genetic drift). 값이 높을수록(따뜻한 색으로 표시) AR33K와 더 높은 유전적 부동을 공유함을 나타낸다. 점선의 교차점은 AR33K의 위치를 나타낸다.

(B) Results for D(North, Qihe; AR19K, Mbuti) and D(AR19K, Qihe; North, Mbuti). North represents northern East Asians, including the ancient Amur region populations (red), ancient Shandong populations (purple), ancient West Liao River population (orange), ancient Mongolian populations (green), ancient inner Mongolian populations (blue), and ancient Yellow River population (brown). The dashed line indicates where D = 0.  A solid circle indicates that |Z| > 3, and otherwise |Z| < 3. The thick gray bar indicates one standard error, and the thin gray bar indicates two standard errors. Directions of deviations of D statistics are labeled with population name on the x axis at the top left and right of the panels. All SNPs were used to compute D statistics.

(B) D(북부, 기하; AR19K, 음부티) 및 D(AR19K, 기하; 북부, 음부티)에 대한 결과. ‘북부(North)’는 고대 아무르(Amur) 지역 인구(빨간색), 고대 산동(山東) 인구(보라색), 고대 서요하(西遼河) 인구(주황색), 고대 몽골 인구(녹색), 고대 내몽골 인구(파란색), 고대 황하(黃河) 인구(갈색)를 포함하는 북부 동아시아인을 나타낸다. 점선은 D=0인 지점을 나타낸다.  채워진 원은 |Z| > 3임을, 그렇지 않은 경우는 |Z| < 3임을 나타낸다. 굵은 회색 막대는 1 표준오차를, 얇은 회색 막대는 2 표준오차를 나타낸다. D 통계량 편차의 방향은 패널의 상단 왼쪽과 오른쪽에 있는 x축에 모집단 이름으로 표시되어 있다. D 통계량을 계산하는 데 모든 SNP가 사용되었다.

See also Figure S2 and Table S3.

그림 S2 및 표 S3도 참조하라.

최종빙기절정기 이후 아무르(Amur) 지역의 유전적 연속성과 인구 상호작용 Genetic continuity in the Amur region and population interactions after the LGM

Ancient DNA samples recovered from Neolithic forager/farmers at Devil’s Gate cave (DevilsCave_N) in Primorsky Krai, far eastern Russia and the Amur River Basin reveal that genetic continuity with modern populations existed in the Amur region since 8 ka (Ning et al., 2020; Sikora et al., 2019; Siska et al., 2017). However, population dynamics in the Amur region are unclear before 8 ka. Our younger samples were clustered into four subgroups based on genetic affinity and time, including AR14K, AR13-10K, ARpost9K, and AR9.2K_o (Figures 1B and 1C). This increased sampling from a large temporal range after the LGM stadial allows a better resolution of population dynamics in the Amur region and surrounding areas. Overall, PCA (Figure 1B), outgroup-f3 (Figure 1C), and ADMIXTURE (Alexander et al., 2009; Figure S4) analyses indicate that AR14K, AR13-10K, and ARpost9K are genetically closest to the DevilsCave_N population. This was validated because D statistics show that post-LGM Amur region populations share more alleles with DevilsCave_N relative to other ancient northern East Asians (Figure 4A), which is also supported by results from qpGraph and Treemix (Figure 3A; Figure S3). In addition, we observed a decrease in human background relatedness, as measured by short runs of homozygosity (ROHS) (4-8 cM) after 14 ka, which suggests an increasing local population size in the Amur region, probably because of the transition to sedentary farming (the “Neolithic transition”) (Table S4; Bacci, 2017; Diamond and Bellwood, 2003). One outlier population, AR9.2K_o, is genetically close to ancient populations in the Amur region after 14 ka. Interestingly, this individual also shares some genetic affinity with ancient Shandong populations (Table S3), but this possible population interaction needs to be confirmed by further sampling. Our results demonstrate that the genetic continuity reported between modern inhabitants of the Amur River Basin and the Devil’s Gate cave population probably started as early as 14 ka, 6,000 years earlier than proposed previously (Ning et al., 2020; Sikora et al., 2019; Siska et al., 2017). This is consistent with the archaeological record of the earliest appearance of pottery in the Amur region (~15 ka) (Wang et al., 2017; Yue et al., 2019). In addition, the great biological diversity of the Amur region enabled humans to adopt varied, rich subsistence strategies, including hunting, fishing, and animal husbandry, to support this genetic continuity (Ning et al., 2020).

러시아(Russia) 극동 프리모르스키(Primorsky Krai) 지방과 아무르강(Amur River) 유역에 있는 악마의 문 동굴(Devil’s Gate cave, DevilsCave_N)의 신석기 시대 수렵채집인/농부로부터 회수된 고대 DNA 샘플은, 아무르(Amur) 지역에서 8천 년 전부터 현대 인구와의 유전적 연속성이 존재했음을 보여준다 (닝(Ning) et al., 2020; 시코라(Sikora) et al., 2019; 시스카(Siska) et al., 2017). 하지만 8천 년 전 이전 아무르(Amur) 지역의 인구 역학은 불분명하다. 우리의 더 젊은 연대의 샘플들은 유전적 친연성과 시기를 기준으로 AR14K, AR13-10K, ARpost9K, AR9.2K_o를 포함한 네 개의 하위 그룹으로 묶였다 (그림 1B와 1C). 최종빙기절정기 이후의 넓은 시간 범위에서 샘플링을 늘림으로써 아무르(Amur) 지역과 그 주변 지역의 인구 역학을 더 높은 해상도로 파악할 수 있게 되었다. 전반적으로 PCA (그림 1B), 아웃그룹-f3 (그림 1C), ADMIXTURE (알렉산더(Alexander) et al., 2009; 그림 S4) 분석은 AR14K, AR13-10K, ARpost9K가 유전적으로 악마의 문 동굴(DevilsCave_N) 인구와 가장 가깝다는 것을 보여준다. 이는 D 통계 분석으로 검증되었다. D 통계는 최종빙기절정기(LGM) 이후 아무르(Amur) 지역 인구가 다른 고대 북부 동아시아인들보다 악마의 문 동굴인(DevilsCave_N)과 더 많은 대립유전자를 공유함을 보여준다. 이 결과는 qpGraph와 Treemix 분석으로도 뒷받침된다. 또한, 14,000년 전 이후 짧은 동형접합구간(ROHS)으로 측정한 인구 내 배경 근연도가 감소하는 것을 관찰했다. 이는 아무르 지역의 인구 규모가 증가했음을 시사하며, 아마도 정주 농경으로의 전환(“신석기 전환”) 때문일 것이다. 한편, 특이 집단인 AR9.2K_o는 14,000년 전 이후 아무르 지역의 고대 인구와 유전적으로 가깝다. 흥미롭게도 이 개체는 고대 산동(山東) 인구와도 약간의 유전적 친연성을 공유하지만(표 S3), 이러한 인구 상호작용 가능성은 추가 샘플링으로 확인해야 한다. 우리의 결과는 아무르강(Amur River) 유역의 현대 거주자와 악마의 문 동굴 인구 간에 보고된 유전적 연속성이, 기존에 제안된 것보다 6,000년 더 이른 약 14,000년 전에 시작되었을 가능성을 보여준다 (닝(Ning) et al., 2020; 시코라(Sikora) et al., 2019; 시스카(Siska) et al., 2017). 이는 아무르(Amur) 지역에서 가장 오래된 토기가 등장한 고고학적 기록(약 15,000년 전)과 일치한다 (왕(Wang) et al., 2017; 위에(Yue) et al., 2019). 또한 아무르(Amur) 지역은 생물학적으로 매우 다양했다. 덕분에 이곳 사람들은 사냥, 고기잡이, 가축 기르기 등 다채롭고 풍요로운 방식으로 살아갈 수 있었다. 이러한 안정적인 생활 방식이 이 지역의 유전적 연속성을 뒷받침했다.

An early human population defined as Ancient Paleo-Siberians (represented by the 10,000-year-old Kolyma individual from northeastern Siberia and the 14,000-year-old Ust-Kyakhta-3 or UKY specimen found near Lake Baikal in southern Siberia) have been proposed to have descended from Ancient North Eurasian (ANE)-related populations, mixing with newly arriving people carrying East Asian ancestry (Sikora et al., 2019; Yu et al., 2020). Ancient Paleo-Siberians have also been shown to be the closest relatives to Native American populations outside of the Americas and were associated with the spread of a microblade technology and post-LGM reduction of the Mammoth Steppe ecological community (Pitulko and Nikolskiy, 2012; Sikora et al., 2019). However, the modeling of Ancient Paleo-Siberians as a two-way mixture of DevilsCave_N (representing East Asian ancestry) and AfontovaGora3 (Fu et al., 2016) (representing ANE ancestry) was unsuccessful for UKY (tail probability 1.45E-03) or Kolyma (tail probability = 3.98E-08) (Yu et al., 2020), indicating East Asian populations that could have feasibly contributed to the Ancient Paleo-Siberians have not yet been identified. We explored the relationship between the northern East Asian samples older than 13,000 years (AR19K and AR14K) and Ancient Paleo-Siberians. Using qpAdm, we showed that both UKY and Kolyma could be successfully modeled as a mixture of AR19K/AR14K and USR1 (Figure 3B; Table S5). It was also demonstrated through qpGraph that the AR14K-related population could be the direct East Asian source for Ancient Paleo-Siberians (Figure 3A). This is consistent with D statistics that indicate that Ancient Paleo-Siberians were genetically closer to Amur region populations after 14 ka (D(AR14K/AR13-10K/ ARpost9K, AR19K; UKY, Mbuti) [2.5 < Z < 3.9]) than to other ancient northern East Asians (i.e., D(aCNEA, AR19K; UKY, Mbuti) [0.7 < Z < 1.4]) (Table S3). We propose that Amur region ancestry after 14 ka is a better fit than DevilsCave_N ancestry for the East Asian component in Ancient Paleo-Siberians and, thus, that Amur region populations could have been at the forefront of interactions with ANE-related populations that likely contributed to Ancient Paleo-Siberians.

고대 고시베리아인(Ancient Paleo-Siberians)으로 정의되는 초기 인류 집단이 있다. 이들은 시베리아(Siberia) 북동부에서 발견된 1만 년 전의 콜리마(Kolyma) 개체와 시베리아(Siberia) 남부 바이칼(Baikal) 호수 근처에서 발견된 1만 4천 년 전의 우스트-캭타(Ust-Kyakhta)-3 또는 UKY 표본으로 대표된다. 이들은 고대 북유라시아인(Ancient North Eurasian, ANE) 관련 집단이, 동아시아 계통을 가진 새로 도착한 사람들과 섞이면서 형성된 후손으로 제안되었다 (시코라(Sikora) et al., 2019; 유(Yu) et al., 2020). 고대 고시베리아인은 아메리카 대륙 밖에서 아메리카 원주민과 가장 가까운 친척으로도 밝혀졌다. 이들은 세석기(microblade) 기술의 확산, 그리고 최종빙기절정기(LGM) 이후 매머드 스텝 생태계가 줄어든 것과 관련이 있었다. 하지만, 고대 고시베리아인을 악마의 문 동굴인(DevilsCave_N, 동아시아 계통 대표)과 아폰토바 고라 3호인(AfontovaGora3, ANE 계통 대표)의 양방향 혼합으로 모델링하는 것은 UKY(꼬리 확률 1.45E-03)나 콜리마인(꼬리 확률 3.98E-08)에게는 성공적이지 않았다 (유(Yu) et al., 2020; 푸(Fu) et al., 2016). 이는 고대 고시베리아인에게 기여했을 가능성이 있는 동아시아 인구 집단이 아직 확인되지 않았음을 나타낸다. 우리는 13,000년보다 오래된 북부 동아시아 샘플(AR19K와 AR14K)과 고대 고시베리아인의 관계를 탐구했다. qpAdm을 사용하여, UKY와  콜리마(Kolyma) 개체 모두 AR19K/AR14K와 USR1(고대 베링기아인)의 혼합으로 성공적으로 모델링될 수 있음을 보였다 (그림 3B; 표 S5). 또한 qpGraph를 통해 AR14K 관련 인구 집단이 고대 고시베리아인의 직접적인 동아시아 조상일 수 있음이 증명되었다 (그림 3A). 이는 D 통계량과도 일치한다. D 통계량은 고대 고시베리아인이 다른 고대 북부 동아시아인들(예: D(aCNEA, AR19K; UKY, 음부티) [0.7 < Z < 1.4])보다 1만 4천 년 전 이후의 아무르(Amur) 지역 인구(D(AR14K/AR13-10K/ARpost9K, AR19K; UKY, 음부티) [2.5 < Z < 3.9])와 유전적으로 더 가깝다는 것을 보여준다 (표 S3). 우리는 1만 4천 년 전 이후 아무르(Amur) 지역의 계통이, 고대 고시베리아인 내의 동아시아 요소를 설명하는 데 악마의 문 동굴인(DevilsCave_N)의 계통보다 더 적합하다고 제안한다. 따라서 아무르(Amur) 지역 인구는 고대 고시베리아인 형성에 기여했을 가능성이 있는 ANE 관련 인구와의 상호작용 최전선에 있었을 수 있다.

Figure 3. Admixture graph modeling with qpGraph and admixture modeling with qpAdm for Ancient Paleo-Siberians (Kolyma and UKY)
그림 3. 고대 고시베리아인(콜리마 및 UKY)에 대한 qpGraph를 이용한 혼합 그래프 모델링과 qpAdm을 이용한 혼합 모델링

(A and B) Admixture graph modeling with qpGraph (A) and admixture modeling with qpAdm (B) for Ancient Paleo-Siberians (Kolyma and UKY). Color coding for USR1 and AR14K is the same in both graphs. In qpGraph, yellow represents Tianyuan and AR33K, brown represents coastal northern East Asian ancestry, green represents coastal southern East Asian ancestry, and gray represents non-East Asian ancestries. In qpAdm, the error bars represent the standard errors of estimated ancestry proportions. See also Figure S3 and Table S5.

(A 와 B) 고대 고시베리아인(콜리마(Kolyma) 및 UKY)에 대한 qpGraph를 이용한 혼합 그래프 모델링(A)과 qpAdm을 이용한 혼합 모델링(B). USR1과 AR14K의 색상 코딩은 두 그래프에서 동일하다. qpGraph에서 노란색은 전원인(田園人)과 AR33K를, 갈색은 해안 북부 동아시아 계통을, 녹색은 해안 남부 동아시아 계통을, 회색은 비(非)동아시아 계통을 나타낸다. qpAdm에서 오차 막대는 추정된 계통 비율의 표준오차를 나타낸다. 그림 S3 및 표 S5도 참조하라.

4만 년 전부터 6천 년 전까지 북부 동아시아인의 적응성 유전자 변이 Adaptive genetic variants in northern East Asians from 40 to 6 ka

Even though selection has been inferred in western Eurasians utilizing ancient DNA (Lindo et al., 2018; Mathieson et al., 2015; Ye et al., 2017), the limited availability of ancient East Asian genomes has precluded similar research in East Asia. The data presented here, together with those from recently published ancient East Asian populations (Yang et al., 2020), provide a large temporal window from 40-6 ka (spanning the pre- and post-LGM periods) to understand the evolution of adaptive variants. One important adaptive variant for East Asians is the EDAR V370A mutation (frequency at 93.7% in Han Chinese in Beijing, China [CHB]; Auton et al., 2015), which has been demonstrated through genome-wide association studies to be associated with thicker hair shafts, more sweat glands, and shovel-shaped incisors (Fujimoto et al., 2008; Tan et al., 2013). However, the precise mechanism behind EDAR V370A selection is difficult to detect using evidence obtained only from modern genomes. Two selective mechanisms have been proposed: selection to increase vitamin D in breast milk in the low-UV environment around 20 ka or selection to modulate thermoregulatory sweating during a warm and humid climate around 30 ka (Hlusko et al., 2018; Kamberov et al., 2013). Here we show that mutation V370A in the EDAR gene appeared in all ancient East Asians (including AR19K), except for AR33K and Tianyuan, the only two individuals in our sample predating the LGM stadial in East Asia (Figure 4B; Table S6). This indicates that EDAR V370A was likely to be elevated to high frequency during or shortly after the LGM. These direct observations using ancient DNA make the hypothesis that EDAR V370A was under selection during a warm and humid environment less likely (Kamberov et al., 2013). Furthermore, the allelic age of EDAR V370A has been estimated previously to be 11,400 years ago with a large confidence interval (4,300-43,700) using modern genomes (Peter et al., 2012), whereas our direct observation demonstrated that this allele emerged as early as ~19 ka, as observed in the AR19K individual (Figure 4B), providing an older and more accurate upper boundary for the allelic age estimate.

고대 DNA를 활용하여 서부 유라시아인에게서 자연선택이 일어났음이 추론된 바 있지만 (린도(Lindo) et al., 2018; 매티슨(Mathieson) et al., 2015; 예(Ye) et al., 2017), 고대 동아시아 게놈은 이용 가능한 수가 제한적이어서 동아시아에서는 유사한 연구가 불가능했다. 여기에 제시된 데이터는 최근 발표된 고대 동아시아 인구 데이터와 함께 (양(Yang) et al., 2020), 4만 년에서 6천 년 전(최종빙기절정기 전후 기간에 걸쳐)이라는 넓은 시간적 창을 제공하여 적응성 변이의 진화를 이해할 수 있게 한다. 동아시아인에게 중요한 적응성 변이 중 하나는 EDAR V370A 돌연변이이다 (중국 북경(北京)의 한족(漢族)에서 93.7% 빈도; 오턴(Auton) et al., 2015). 이 돌연변이는 게놈 전체 연관 연구를 통해 두꺼운 머리카락, 더 많은 땀샘, 삽 모양 앞니 등과 관련이 있는 것으로 증명되었다 (후지모토(Fujimoto) et al., 2008; 탄(Tan) et al., 2013). 하지만 EDAR V370A가 선택된 정확한 메커니즘은 현대인 게놈만으로는 파악하기 어렵다. 두 가지 선택 메커니즘이 제안되었다. 하나는 약 2만 년 전 자외선이 적은 환경에서 모유의 비타민 D를 늘리기 위한 선택이라는 가설이고, 다른 하나는 약 3만 년 전 따뜻하고 습한 기후에서 체온 조절을 위한 발한을 조절하기 위한 선택이라는 가설이다 (흘루스코(Hlusko) et al., 2018; 캄베로프(Kamberov) et al., 2013). 여기서 우리는 EDAR 유전자의 V370A 돌연변이가, 동아시아의 최종빙기절정기 이전에 살았던 우리 샘플의 유이한 두 개체인 AR33K와 전원인(田園人)을 제외한 모든 고대 동아시아인(AR19K 포함)에게서 나타남을 보여준다 (그림 4B; 표 S6). 이는 EDAR V370A가 최종빙기절정기 동안 또는 직후에 높은 빈도로 증가했을 가능성이 매우 높다는 것을 시사한다. 고대 DNA를 이용한 이러한 직접적인 관찰은 EDAR V370A가 따뜻하고 습한 환경에서 선택되었다는 가설의 가능성을 낮춘다 (캄베로프(Kamberov) et al., 2013). 더욱이, EDAR V370A의 대립유전자 연대는 이전에 현대인 게놈을 사용하여 11,400년 전으로 추정되었으나 신뢰 구간이 4,300년에서 43,700년으로 매우 넓었다 (피터(Peter) et al., 2012). 반면 우리의 직접적인 관찰은 AR19K 개체에서 확인된 바와 같이 이 대립유전자가 약 19,000년 전이라는 이른 시기에 출현했음을 증명하여, 대립유전자 연대 추정에 있어 더 오래되고 더 정확한 상한선을 제공한다 (그림 4B).

In conclusion, our findings reveal the demographic changes that occurred over a long period in northern East Asia, including before the LGM stadial (~33 ka) and after (from ~19-6 ka). This temporal transect witnessed two major population shifts before 14 ka, followed by long-term genetic continuity in the Amur region. (1) Before the LGM, Tianyuan-related ancestry was widespread with no evidence of Yana-related ancestry in the Amur region. (2) At the close of the LGM (~19 ka), Tianyuan-related ancestry was likely replaced by modern East Asian ancestry. Also, the earliest northern East Asian population, represented by AR19K, appeared in the Amur region during the last stage of the LGM and is basal to all ancient northern East Asians. (3) After 14 ka, populations in the Amur region maintained genetic continuity and are the closest East Asian source known for Ancient Paleo-Siberians, one branch of which represents the closest relative of Native American populations outside of the Americas. In addition to uncovering previously unknown population dynamics, our analyses provide the ancient DNA evidence to narrow the timing of the appearance of EDAR V370A, indicating that this genetic variant was likely elevated to high frequency during or shortly after the LGM.

결론적으로, 우리의 연구 결과는 북부 동아시아에서 최종빙기절정기 이전(약 33,000년 전)과 이후(약 19,000-6,000년 전)를 포함하는 장기간에 걸쳐 일어난 인구학적 변화를 보여준다. 이 시간적 단면에서는 14,000년 전 이전에 두 차례의 주요 인구 변동이 있었고, 그 후 아무르(Amur) 지역에서는 장기적인 유전적 연속성이 이어졌다. (1) 최종빙기절정기 이전에는 전원인(田園人) 관련 계통이 널리 퍼져 있었고, 아무르(Amur) 지역에서는 야나(Yana) 관련 계통의 증거가 없었다. (2) 최종빙기절정기 말(약 19,000년 전)에는 전원인(田園人) 관련 계통이 현대 동아시아인 계통으로 대체되었을 가능성이 있다. 또한, AR19K로 대표되는 최초의 북부 동아시아인 인구가 최종빙기절정기 마지막 단계에 아무르(Amur) 지역에 나타났으며, 이들은 모든 후대 고대 북부 동아시아인의 기층(basal)을 이룬다. (3) 14,000년 전 이후, 아무르(Amur) 지역의 인구는 유전적 연속성을 유지했으며, 아메리카 대륙 외부에서 아메리카 원주민의 가장 가까운 친척을 대표하는 고대 고시베리아인(Ancient Paleo-Siberians)의 가장 가까운 동아시아 조상으로 밝혀졌다. 이전에 알려지지 않았던 인구 역학을 밝혀낸 것 외에도, 우리의 분석은 EDAR V370A의 출현 시기를 좁힐 수 있는 고대 DNA 증거를 제공한다. 이는 이 유전자 변이가 최종빙기절정기 동안 또는 직후에 높은 빈도로 증가했을 가능성이 크다는 것을 나타낸다.

Figure 4. Ancient populations in the Amur region after 19 ka and time transect of EDAR V370A
그림 4. 19,000년 전 이후 아무르(Amur) 지역의 고대 인구와 EDAR V370A의 시간적 단면

(A) Genetic affinities with DevilsCave_N and post-LGM populations in the Amur region (AR14K, AR13-10K, and ARpost9K) with D statistics. North represents northern East Asians and Ancient Paleo-Siberians, including Ancient Paleo-Siberians (Kolyma and UKY; pink), ancient Amur region populations (red), ancient Shandong populations (purple), ancient West Liao River population (orange), ancient Mongolian populations (green), ancient inner Mongolian populations (blue), and ancient Yellow River population (brown). Directions of deviations of D statistics are labeled with the population name on the x axis at the top left and right. All SNPs were used to compute D statistics.

(A) D 통계량을 이용한 악마의 문 동굴인(DevilsCave_N)과 LGM 이후 아무르 지역 인구(AR14K, AR13-10K, ARpost9K) 간의 유전적 친연성. ‘북부(North)’는 고대 고시베리아인(콜리마(Kolyma)와 UKY; 분홍색), 고대 아무르 지역 인구(빨간색), 고대 산동(山東) 인구(보라색), 고대 서요하(西遼河) 인구(주황색), 고대 몽골 인구(녹색), 고대 내몽골 인구(파란색), 고대 황하(黃河) 인구(갈색)를 포함하는 북부 동아시아인과 고대 고시베리아인을 나타낸다. D 통계량 편차의 방향은 상단 왼쪽과 오른쪽에 있는 x축에 모집단 이름으로 표시되어 있다. D 통계량을 계산하는 데 모든 SNP가 사용되었다.

(B) Allele counts for adaptive mutations of EDAR V370A in Tianyuan, ancient populations in the Amur region, and recently published ancient northern East Asians (Yumin, Bianbian, Boshan, Xiaogao, and Xiaojingshan) (Yang et al., 2020) from 40-6 ka. Purple shading represents the period of the LGM. Dots represent genotype calls (blue, ancestral alleles; red, derived alleles; gray, missing alleles). The number at the top of the bar (separated by a comma) shows the allele coverage for the derived and ancestral allele, and the number in parentheses at the bottom of the bar shows the number of individuals included in that time column. Genotypes were called using a maximum-likelihood-based method (snpAD) that takes into account possible errors in the covered alleles. See also Figure S4, Table S6, and STAR Methods.

(B) 4만 년에서 6천 년 전 사이의 전원인(田園人), 고대 아무르 지역 인구, 그리고 최근 발표된 고대 북부 동아시아인(유민(裕民), 변변(邊邊), 박산(博山), 소고(曉高), 소형산(小荊山)) (양(Yang) et al., 2020)에서 나타나는 EDAR V370A 적응성 돌연변이의 대립유전자 수. 보라색 음영은 LGM 기간을 나타낸다. 점들은 유전형 콜을 나타낸다 (파란색: 조상형 대립유전자, 빨간색: 유래형 대립유전자, 회색: 결측). 막대 상단의 숫자(쉼표로 구분)는 유래형 및 조상형 대립유전자의 커버리지를 보여주고, 막대 하단의 괄호 안 숫자는 해당 시간대에 포함된 개체 수를 보여준다. 유전형은 커버된 대립유전자에서 발생 가능한 오류를 고려하는 최대우도 기반 방법(snpAD)을 사용하여 결정되었다. 그림 S4, 표 S6, STAR Methods도 참조하라.

연구의 한계 Limitations of the study

Although this study presents a long time transect of ancient individuals from the Late Pleistocene to the Holocene, some of the detailed population dynamics need to be verified through further sampling. First, we must collect more samples representing the LGM stadial to further elucidate the post-LGM population replacement we observed and adaptations to environmental shifts by northern East Asians. Second, we must collect additional samples dating to the period after 14 ka to verify possible population interactions between ancient populations in the Amur region and surrounding territories, such as coastal Shandong.

이 연구는 후기 홍적세부터 현세까지 긴 시간대의 고대 개체들을 제시했지만, 일부 세부적인 인구 역학은 추가적인 샘플링을 통해 검증될 필요가 있다. 첫째, 우리가 관찰한 최종빙기절정기 이후의 인구 교체와 북부 동아시아인의 환경 변화 적응을 더 명확히 이해하기 위해 최종빙기절정기를 대표하는 더 많은 샘플을 수집해야 한다. 둘째, 14,000년 전 이후의 추가 샘플을 확보하여 아무르(Amur) 지역과 해안 산동(山東)과 같은 주변 지역의 고대 인구 간 상호작용 가능성을 검증해야 한다.

연구 방법 STAR METHODS

Detailed methods are provided in the online version of this paper and include the following:

상세한 방법은 이 논문의 온라인 버전에 제공되며 다음을 포함한다:

  • 핵심 자원 표 KEY RESOURCES TABLE
  • 자원 이용 가능성 RESOURCE AVAILABILITY

° 주요 연락처 Lead contact

° 물질 이용 가능성 Materials availability

° 데이터 및 코드 이용 가능성 Data and code availability

  • 실험 모델 및 대상 상세 정보 EXPERIMENTAL MODEL AND SUBJECT DETAILS

° 유적지 및 표본 설명 Sites and specimen descriptions

  • 방법 상세 정보 METHOD DETAILS

° 고대 DNA 추출 Ancient DNA extraction

° 고대 DNA 포획 및 시퀀싱 Ancient DNA capture and sequencing

  • 정량화 및 통계 분석 QUANTIFICATION AND STATISTICAL ANALYSIS

° 리드 정렬 Read alignment

° 오염 평가 및 변이 콜링 Contamination evaluation and variants calling

° 친족 관계 분석 Kinship analyses

° 주성분 분석 Principal components analysis (PCA)

° ADMIXTURE 분석 analysis

° 아웃그룹-f3 및 D 통계 Outgroup-f3 and D statistics

° Treemix를 이용한 계통 모델링 Phylogeny modeling with Treemix

° qpAdm을 이용한 혼합 모델링 Admixture modeling with qpAdm

° qpGraph를 이용한 인구 통계 모델링 Demographic modeling with qpGraph

° 은닉 마르코프 모델을 이용한 고인류 계통 추정 Archaic ancestry estimation with the Hidden Markov Model

° 아무르(Amur) 지역의 시간 경과에 따른 인구 크기 Time transect of population sizes in the Amur region

° 북부 동아시아인의 EDAR V370A 시간 경과 Time transect of EDAR V370A in northern East Asians

보충 정보 SUPPLEMENTAL INFORMATION

Supplemental information can be found online at https://linkinghub.elsevier.com/retrieve/pii/S0092867421005754.

보충 정보는 https://linkinghub.elsevier.com/retrieve/pii/S0092867421005754 에서 온라인으로 찾을 수 있다.

감사의 글 ACKNOWLEDGMENTS

We thank Zehui Chen for designing and editing images, Xueping Ji for morphological recognition of the sampled specimen and related discussions, Benjamin Peter for providing suggestions related to the analyses of archaic introgression, Janet Kelso for commenting on the manuscript, Xiaohong Wu for discussions of radiocarbon dating, and Jianping Yue and Youqian Li for additional archaeological background. We also thank our anonymous reviewers for their helpful comments. This research was supported by the Chinese Academy of Sciences (CAS) and the Ministry of Finance of the People’s Republic of China (XDB26000000), the National Natural Science Foundation of China (41820104008, 41925009, 91731303, 41672021, and 41630102), the National Key R&D Program of China (2016YFE0203700), the CAS (XDA1905010 and QYZDB-SSW-DQC003), the Research on the Roots of Chinese Civilization program of Zhengzhou University (XKZDJC202006), the Tencent Foundation through its XPLORER Prize, and the Howard Hughes Medical Institute (55008731). J.W.O.’s participation was supported by the Chinese Academy of Sciences President’s International Fellowship Initiative (PIFI; award 2018VCA0016 with 2021 extension).

이미지 디자인과 편집을 해준 첸, 저후이(Chen, Zehui), 샘플 표본의 형태학적 인식 및 관련 논의를 해준 지, 쉐핑(Ji, Xueping), 고인류 유전자 유입 분석과 관련된 제안을 해준 벤자민 피터(Benjamin Peter), 원고에 대해 논평해준 자넷 켈소(Janet Kelso), 방사성 탄소 연대 측정에 대해 논의해준 우, 샤오홍(Wu, Xiaohong), 그리고 추가적인 고고학적 배경을 제공해준 위에, 젠핑(Yue, Jianping)과 리, 요우첸(Li, Youqian)에게 감사한다. 또한 도움이 되는 의견을 준 익명의 심사위원들에게도 감사한다. 이 연구는 중국과학원(CAS)과 중화인민공화국 재정부, 중국 국립자연과학재단, 중국 국가중점연구개발계획, 정주대학(鄭州大學)의 중국 문명 뿌리 연구 프로그램, 텐센트 재단의 XPLORER 상, 그리고 하워드 휴즈 의학 연구소의 지원을 받았다. J.W.O.의 참여는 중국과학원 총장 국제 펠로우십 이니셔티브(PIFI)의 지원을 받았다.

저자 기여 AUTHOR CONTRIBUTIONS

Conceptualization, X.M., H.Z., and Q.F.; formal analysis, X.M. and Q.F.; investigation, S.Q., M.L., T.W., F.C., P.X., and C.L.; resources, H.Z., P.C., R.Y., F.L., Q.D., and X.F.; writing-original draft, X.M. and Q.F.; writing-review & editing, X.M., Q.F., E.A.B., H.Z., S.Q., Y.L., M.Z., J.W.O., W.P., and M.L.; supervision, Q.F.

개념화, X.M., H.Z., Q.F.; 공식 분석, X.M., Q.F.; 조사, S.Q., M.L., T.W., F.C., P.X., C.L.; 자원, H.Z., P.C., R.Y., F.L., Q.D., X.F.; 초고 작성, X.M., Q.F.; 검토 및 편집, X.M., Q.F., E.A.B., H.Z., S.Q., Y.L., M.Z., J.W.O., W.P., M.L.; 감독, Q.F.

이해관계 선언 DECLARATION OF INTERESTS

The authors declare no competing interests.

저자들은 경쟁적 이해관계가 없음을 선언한다.

Received: October 16, 2020

접수: 2020년 10월 16일

Revised: January 20, 2021

수정: 2021년 1월 20일

Accepted: April 23, 2021

채택: 2021년 4월 23일

Published: May 27, 2021

게재: 2021년 5월 27일

보충 인용 SUPPORTING CITATIONS

The following references appear in the supplemental information: de Barros Damgaard et al. (2018); Fu et al. (2014); Gamba et al. (2014); Jeong et al. (2016); Jeong et al. (2019); Jones et al. (2015); Lazaridis et al. (2014); Lazaridis et al. (2016); Lipson et al. (2018); Llorente et al. (2015); McColl et al. (2018); Moreno-Mayar et al. (2018); Nakatsuka et al. (2017); Olalde et al. (2014); Qin and Stoneking (2015); Reich et al. (2011); Seguin-Orlando et al. (2014); Sikora et al. (2017); Skoglund et al. (2016).

다음 참고 문헌들이 보충 정보에 나타난다: 데 바로스 담가르드(de Barros Damgaard) et al. (2018); 푸(Fu) et al. (2014); 감바(Gamba) et al. (2014); 정(Jeong) et al. (2016); 정(Jeong) et al. (2019); 존스(Jones) et al. (2015); 라자리디스(Lazaridis) et al. (2014); 라자리디스(Lazaridis) et al. (2016); 립슨(Lipson) et al. (2018); 요렌테(Llorente) et al. (2015); 맥콜(McColl) et al. (2018); 모레노-마야르(Moreno-Mayar) et al. (2018); 나카츠카(Nakatsuka) et al. (2017); 올랄데(Olalde) et al. (2014); 친(Qin)과 스톤킹(Stoneking) (2015); 라이크(Reich) et al. (2011); 세귄-올랜도(Seguin-Orlando) et al. (2014); 시코라(Sikora) et al. (2017); 스코글룬드(Skoglund) et al. (2016).

참고문헌 REFERENCES

Alexander, D.H., Novembre, J., and Lange, K. (2009). Fast model-based estimation of ancestry in unrelated individuals. Genome Res. 19, 1655-1664.

Andrews, R.M., Kubacka, I., Chinnery, P.F., Lightowlers, R.N., Turnbull. D.M.. and Howell, N. (1999). Reanalysis and revision of the Cambridge reference sequence for human mitochondrial DNA. Nat. Genet. 23, 147.

Auton, A., Brooks, L.D., Durbin, R.M., Garrison, E.P., Kang, H.M., Korbel, J.O… Marchini, J.L., McCarthy, S., McVean, G.A., and Abecasis, G.R.; 1000 Genomes Project Consortium (2015). A global reference for human genetic variation. Nature 526, 68-74.

Bacci. M.L. (2017). A concise history of world population (John Wiley & Sons).

Bae, C.J., Douka, K., and Petraglia, M.D. (2017). On the origin of modern humans: Asian perspectives. Science 358, eaai9067.

Baum, B.R. (1989). PHYLIP: Phylogeny Inference Package. Version 3.2. Joel Felsenstein. Q. Rev. Biol. 64, 539-541.

Clark, P.U., Dyke, A.S., Shakun, J.D., Carlson, A.E., Clark, J., Wohlfarth, B., Mitrovica, J.X., Hostetler, S.W., and McCabe, A.M. (2009). The last glacial maximum. Science 325, 710-714.

CNCB-NGDC Members and Partners (2021). Database Resources of the National Genomics Data Center, China National Center for Bioinformation in 2021. Nucleic Acids Res. 49 (D1), D18-D28.

de Barros Damgaard, P., Martiniano, R., Kamm, J., Moreno-Mayar, J.V., Kroonen, G., Peyrot, M., Barjamovic, G., Rasmussen, S., Zacho, C., Baimukhanov, N., et al. (2018). The first horse herders and the impact of early Bronze Age steppe expansions into Asia. Science 360, eaar7711.

Deviese, T.. Massilani, D., Yi, S., Comeskey, D., Nagel, S.. Nickel, B., Ribechini, E., Lee, J., Tseveendorj, D., Gunchinsuren, B., et al. (2019). Compound-specific radiocarbon dating and mitochondrial DNA analysis of the Pleistocene hominin from Salkhit Mongolia. Nat. Commun. 10, 274.

Diamond, J., and Bellwood, P. (2003). Farmers and their languages: the first expansions. Science 300, 597-603.

Ding, M., Wang, T., Ko, A.M.-S., Chen, H., Wang, H., Dong, G., Lu, H., He, W., Wangdue, S., Yuan, H., et al. (2020). Ancient mitogenomes show plateau populations from last 5200 years partially contributed to present-day Tibetans. Proc. Biol. Sci. 287, 20192968.

Fu. Q., Meyer, M., Gao, X., Stenzel, U., Burbano, H.A., Kelso, J., and Pääbo, S. (2013a), DNA analysis of an early modern human from Tianyuan Cave, China. Proc. Natl. Acad. Sci. USA 110, 2223-2227.

Fu, Q., Mittnik, A., Johnson, P.L.F., Bos, K., Lari, M., Bollongino, R., Sun, C., Giensch, L, Schmitz, R., Burger, J., et al. (2013b). A revised timescale for human evolution based on ancient mitochondrial genomes. Curr. Biol. 23, 553-559.

Fu, Q., Li, H., Moorjani, P., Jay, F., Slepchenko, S.M., Bondarev, A.A., Johnson, P.L.F., Aximu-Petri, A., Prüfer, K., de Filippo, C., et al. (2014). Genome sequence of a 45.000-year-old modern human from western Siberia. Nature 514,445-449.

Fu, Q., Hajdinjak, M., Moldovan, O.T., Constantin, S., Mallick, S., Skoglund, P., Patterson, N., Rohland, N., Lazaridis, I., Nickel, B., et al. (2015). An early modern human from Romania with a recent Neanderthal ancestor. Nature 524, 216-219.

Fu, Q., Posth, C., Hajdinjak, M., Petr, M., Mallick, S., Fernandes, D., Furtwängler, A., Haak, W., Meyer, M., Mittnik, A., et al. (2016). The genetic history of Ice Age Europe. Nature 534, 200-205.

Fujimoto, A., Ohashi, J., Nishida, N., Miyagawa, T., Morishita, Y., Tsunoda, T., Kimura, R., and Tokunaga, K. (2008). A replication study confirmed the EDAR gene to be a major contributor to population differentiation regarding head hair thickness in Asia. Hum. Genet. 124, 179-185.

Gamba, C., Jones, E.R., Teasdale, M.D., McLaughlin, R.L, Gonzalez-Fortes, G., Mattiangell, V., Domboróczki, L., Kővári, I., Pap, I., Anders, A., et al.. (2014). Genome flux and stasis in a five millennium transect of European prehistory. Nat. Commun. 5, 5257.

Gansauge. M.-T., and Meyer, M. (2013). Single-stranded DNA library preparation for the sequencing of ancient or damaged DNA. Nat. Protoc. 8, 737-748.

Haak, W., Lazaridis, I., Patterson, N., Rohland, N., Mallick, S.. Llamas, B., Brandt, G., Nordenfelt, S., Harney, E., Stewardson, K., et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature 522, 207-211.

Haney, E., Patterson, N., Reich, D., and Wakeley, J. (2021), Assessing the performance of qpAdm: a statistical tool for studying population admixture. Genetics 217, iyaa045.

Hayden, B., and Villeneuve, S. (2011). A century of feasting studies. Annu. Rev. Anthropol 40, 433-449.

Heilongjiang Provincial Institute of Cultural Relics and Archaeology: Comnission for Preservation of Ancient Monuments, Raphe County (2019). The excavation of Zone III at the Xiaonanshan Site in 2015, Rache County, Heilongjiang. Chinese Archaeology 20, 87-96.

Hlusko, LJ., Carlson, J.P., Chaplin, G., Elias, S.A., Hoffecker, J.F., Huffman, M., Jablonski, N.G.. Monson. T.A., O’Rourke, D.H., Pilloud. M.A., and Scott, G.R. (2018). Environmental selection during the last ice age on the mother-to-infant transmission of vitamin D and fatty acids through breast milk. Proc. Natl. Acad. Sci. USA 115, E4426-E4432.

Jeong, C., Ozga, A.T., Witonsky, D.B., Malmström, H., Edlund, H., Hofman, C.A., Hagan, R.W., Jakobsson, M., Lewis, C.M., Aldenderfer, M.S., et al. (2016). Long-term genetic stability and a high-altitude East Asian origin for the peoples of the high valleys of the Himalayan arc. Proc. Natl. Acad. Sci. USA 113, 7485-7490.

Jeong. C., Balanovsky, O., Lukianova, E., Kahbatkyzy, N., Flegontov, P., Za-porozhchenko, V., Immel, A., Wang, C.-C., Ixan, O., Khussainova, E., et al. (2019). The genetic history of admixture across inner Eurasia. Nat. Ecol. Evol. 3, 966-976.

Jiamusi Cultural Relics Management Station and Raohe County Cultural Relics Management Institute (1996). Neolithic burial in Xiaonanshan, Raohe County, Heilongjiang. Kaogu 1996, 1-8.

Jones, E.R., Gonzalez-Fortes, G., Connell, S., Siska, V., Eriksson, A., Martiniano, R., McLaughlin, R.L., Gallego Llorente, M., Cassidy, LM., Gamba, C., et al. (2015). Upper Palaeolithic genomes reveal deep roots of modern Eurasians. Nat. Commun. 6, 8912.

Kamberov, Y.G., Wang, S., Tan, J., Gerbault, P., Wark, A., Tan. L, Yang Y., LI S., Tang, K., Chen, H., et al. (2013). Modeling recent human evolution in mice by expression of a selected EDAR variant. Cell 152, 691-702.

Kılınç, G.M., Kashuba, N., Koptekin, D., Bergfeldt, N., Dönertaş, H.M., Rodriguez-Varela, R., Shergin, D., Ivanov, G., Kichigin, D., Pestereva, K., et al. (2021). Human population dynamics and Yersinia pestis in ancient northeast Asia. Sci. Adv. 7, eabc4587.

Kircher, M., Sawyer, S., and Meyer, M. (2012). Double indexing overcomes inaccuracies in multiplex sequencing on the Illumina platform. Nucleic Acids Res. 40, e3.

Korneliussen, T.S., Albrechtsen, A., and Nielsen, R. (2014). ANGSD: Analysis of Next Generation Sequencing Data. BMC Bioinformatics 15. 356.

Kunikita, D., Wang, L., Onuki, S., Sato, H.. and Matsuzaki, H. (2017). Radiocarbon dating and dietary reconstruction of the Early Neolithic Houtaomuga and Shuangta sites in the Song-Nen Plain, Northeast China. Quat. Int. 441,62-68.

Lazaridis, I., Patterson, N., Mittnik, A., Renaud, G., Mallick, S., Kirsanow, K., Sudmant, P.H.. Schraiber, J.G., Castellano, S., Lipson, M., et al. (2014). Ancient human genomes suggest three ancestral populations for present-day Europeans. Nature 513, 409-413.

Lazaridis, I., Nadel, D., Rollefson, G., Merrett, D.C., Rohland, N.. Mallick, S., Fernandes, D., Novak, M., Gamarra, B., Sirak, K., et al. (2016). Genomic insights into the origin of farming in the ancient Near East. Nature 536,419-424.

Li, N., and Cribb, R. (2014). Historical Atlas of Northeast Asia, 1590-2010: Korea, Manchuria, Mongolia, Eastern Siberia (Columbia University Press).

LI, H., and Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754-1760.

Li S., Yang, S., and Liu, X. (2015). Spatiotemporal variability of extreme precipitation in north and south of the Qinling-Huaihe region and influencing factors during 1960-2013. Prog. Geogr. 34, 354-363.

Lindo, J., Haas, R., Hofman, C., Apata, M., Moraga, M., Verdugo, R.A., Watson, J.T., Viviano Llave, C., Witonsky, D., Beall, C., et al. (2018). The genetic prehistory of the Andean highlands 7000 years BP though European contact. Scl. Adv. 4, eaau4921.

Lipatov, M., Sanjeev, K., Patro, R., and Veeramah, K.R. (2015). Maximum Likelihood Estimation of Biological Relatedness from Low Coverage Sequencing Data, bioRxiv, 023374.

Lipson, M. (2020). Applying f-statistics and admixture graphs: Theory and examples. Mol. Ecol. Resour 20, 1658-1667.

Lipson, M. Cheronet, O., Mallick, S., Rohland, N., Oxenham, M., Pietrusewsky, M., Pryce, T.O., Willis, A., Matsumura, H., Buckley, H., et al. (2018). Ancient genomes document multiple waves of migration in Southeast Asian prehistory. Science 361, 92-95.

Llorente, M.G., Jones, E.R., Eriksson, A., Siska, V., Arthur, K.W., Arthur, J.W., Curtis, M.C., Stock, J.T., Coltorti, M., Pieruccini, P., et al. (2015). Ancient Ethiopian genome reveals extensive Eurasian admixture throughout the African continent. Science 350, 820-822.

Lu, D., Lou, H., Yuan, K., Wang, X., Wang, Y., Zhang, C., Lu, Y., Yang, X., Deng, L. Zhou, Y., et al. (2016). Ancestral Origins and Genetic History of Tibetan Highlanders. Am. J. Hum. Genet. 99, 580-594.

Mallick, S., Li, H., Lipson, M., Mathieson, I., Gymrek, M., Racimo, F., Zhao, M., Chennagiri, N., Nordenfelt, S., Tandon, A., et al. (2016). The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. Nature 538, 201-206.

Massilani, D., Skov, L., Hajdinjak, M., Gunchinsuren, B., Tseveendorj D., Yi, S., Lee, J., Nagel, S., Nickel, B., Deviése, T., et al. (2020). Denisovan ancestry and population history of early East Asians. Science 370, 579-583.

Mathieson, I., Lazaridis, I., Rohland, N., Mallick, S., Patterson, N., Roodenberg, S.A., Hamey, E., Stewardson, K., Fernandes, D., Novak, M., et al. (2015). Genome-wide patterns of selection in 230 ancient Eurasians. Nature 528, 499-503.

McColl, H., Racimo, F., Vinner, L, Demeter, F., Gakuhari, T., Moreno-Mayar, J.V., van Driern, G., Gram Wilken, U., Seguin-Orlando, A., de la Fuente Castro, C., et al. (2018). The prehistoric peopling of Southeast Asia. Science 361,88-92.

Meyer, M., and Kircher, M. (2010). Illumina sequencing library preparation for highly multiplexed target capture and sequencing. Cold Spring Harb. Protoc. 2010, pdb-prot5448.

Meyer, M., Kircher, M.. Gansauge, M.-T., Li, H., Racimo, F., Mallick, S., Schraiber, J.G., Jay, F., Prufer, K., de Filippo, C., et al. (2012). A high-coverage genome sequence from an archaic Denisovan individual. Science 338, 222-228.

Mingram, J., Stebich, M., Schettler, G., Hu, Y., Rioual, P., Nowaczyk, N., Dulski, P., You, H.. Opitz, S.. Liu, Q., et al. (2018). Millennial-scale East Asian monsoon variability of the last glacial deduced from annually laminated sediments from Lake Sihailongwan, NE China. Quat. Sci. Rev. 201,57-76.

Monroy Kuhn, J.M., Jakobsson, M., and Günther, T. (2018). Estimating genetic kin relationships in prehistoric populations. PLoS ONE 13, e0195491.

Moreno-Mayar, J.V., Vinner, L., de Barros Damgaard, P., de la Fuente, C., Chan, J., Spence, J.P., Allentoft, M.E., Vimala, T., Racimo, F., Pinotti. T., et al. (2018). Early human dispersals within the Americas. Science 362, eaav2621.

Nakatsuka, N., Moorjani, P., Rai, N., Sarkar, B., Tandon, A., Patterson, N., Bhavani, G.S., Girisha, K.M., Mustak, M.S., Srinivasan, S., et al. (2017). The promise of discovering population-specific disease-associated genes in South Asia. Nat. Genet. 49, 1403-1407.

Narasimhan, V.M., Patterson, N., Moorjani, P., Rohland, N., Bernardos, R., Mallick, S.. Lazaridis, I., Nakatsuka. N.. Olalde. I., Lipson, M., et al. (2019). The formation of human populations in South and Central Asia. Science 365, eaat7487.

Ning. C., Li, T., Wang, K., Zhang, F., Li, T., Wu, X.. Gao, S., Zhang, Q., Zhang, H., Hudson, M.J., et al. (2020). Ancient genomes from northern China suggest links between subsistence changes and human migration. Nat. Commun. 11, 2700.

Olalde, I., Allentoft, M.E., Sánchez-Quinto, F., Santpere, G., Chiang, C.W.K., DeGiorgio, M., Prado-Martinez, J., Rodriguez, J.A., Rasmussen, S.. Quilez, J., et al. (2014). Derived immune and ancestral pigmentation alleles in a 7,000-year-old Mesolithic European. Nature 507, 225-228.

Patterson, N.. Price, A.L., and Reich. D. (2006). Population structure and eigenanalysis. PLoS Genet. 2, e190.

Patterson, N., Moorjani, P., Luo, Y., Mallick, S., Rohland, N., Zhan, Y., Genschoreck, T., Webster. T., and Reich, D. (2012). Ancient admixture in human history. Genetics 192, 1065-1093.

Pearson, R. (2005). The social context of early pottery in the Lingnan region of south China. Antiquity 79, 819-828.

Peter, B.M. (2020), 100,000 years of gene flow between Neandertals and Denisovans in the Altai mountains. bioRxiv. https://www.biorxiv.org/content/10.1101/2020.03.13.990523v1.

Peter, B.M.. Huerta-Sanchez, E., and Nielsen, R. (2012). Distinguishing between selective sweeps from standing variation and from a de novo mutation. PLoS Genet. 8. e1003011.

Pickrell, J.K., and Pritchard, J.K. (2012). Inference of population splits and mixtures from genome-wide allele frequency data. PLoS Genet. 8, 1002967.

Pitulko, V.V., and Nikolskiy, P.A. (2012). The extinction of the woolly mammoth and the archaeological record in Northeastem Asia World Archaeol. 44, 21-42.

Prüfer, K. (2018). sripAD: an ancient DNA genotype caller. Bioinformatics 34, 4165-4171.

Prüfer, K., Racimo, F., Patterson, N., Jay, F., Sankararaman, S., Sawyer, S., Heinze, A., Renaud, G., Sudmant, P.H., de Filippo, C., et al. (2014). The complete genome sequence of a Neanderthal from the Altai Mountains. Nature 505,43-49.

Prüfer, K., de Filippo, C., Grote, S., Mafessoni, F., Korlević, P., Hajdinjak, M.. Vernot, B., Skov, L., Hsieh, P., Peyrėgne, S., et al. (2017). A high-coverage Neandertal genome from Vindija Cave in Croatia. Science 358, 655-658.

Purcell, S., Neale, B., Todd-Brown, K., Thomas, L., Ferreira, M.A.R., Bender, D., Maller, J., Sklar, P., de Bakker, P.I.W., Daly, M.J., and Sham, P.C. (2007). PLINK: a tool set for whole-genome association and population-based linkage analyses. Am. J. Hum. Genet. 81. 559-575.

Qin, P., and Stoneking, M. (2015). Denisovan Ancestry in East Eurasian and Native American Populations. Mol. Biol. Evol. 32, 2665-2674.

Raghavan, M., Skoglund, P., Graf, K.E., Metspalu, M., Albrechtsen, A., Moltke, 1., Rasmussen, S., Stafford, T.W., Jr., Orlando, L., Metspalu, E., et al. (2014), Upper Palaeolithic Siberian genome reveals dual ancestry of Native Americans. Nature 505, 87-91.

Ramsey, C.B., and Lee, S. (2013). Recent and planned developments of the program OxCal, Radiocarbon 55, 720-730.

Reich, D., Patterson, N., Kircher, M., Delfin, F., Nandineni. M.R., Pugach, I., Ko, A.M., Ko, Y.C., Jinam, T.A., Phipps, M.E., et al. (2011). Denisova admixture and the first modern human dispersals into Southeast Asia and Oceania. Am. J. Hum. Genet. 89, 516-528.

Reimer, P.J., Austin, W.E.N., Bard, E., Bayliss, A., Blackwell, P.G., Ramsey, C.B.. Butzin. M., Cheng, H., Edwards, R.L., Friedrich, M., et al. (2020). The IntCal20 Northern Hemisphere radiocarbon age calibration curve (0-55 cal kBP). Radiocarbon 62, 725-757.

Renaud, G., Stenzel, U., and Kelso, J. (2014). leeHom: adaptor trimming and merging for Illumina sequencing reads. Nucleic Acids Res. 42, e141.

Ringbauer, H., Novembre, J., and Steinrücken, M. (2020). Human Parental Relatedness through Time Detecting Runs of Homozygosity in Ancient DNA. bioRxiv. https://www.biorxiv.org/content/10.1101/2020.05.31.126912v2.

Rohland, N., Harney, E., Mallick, S., Nordenfelt, S., and Reich, D. (2015). Partial uracil-DNA-glycosylase treatment for screening of ancient DNA. Philos. Trans. R Soc. Lond. B Biol. Sci. 370, 20130624.

Rybin, E.P., Khatsenovich, A.M., Gunchinsuren, B., Oisen, J.W., and Zwyns, N. (2016). The Impact of the LGM on the development of the Upper Paleolithic in Mongolia. Quat. Int. 425, 69-87.

Sato, H., and Natsuki, D. (2017). Human behavioral responses to environmental condition and the emergence of the world’s oldest pottery in East and Northeast Asia: An overview. Quat. Int. 441, 12-28.

Seguin-Orlando, A., Korneliussen, T.S., Sikora, M., Malaspinas, A.-S., Manica, A., Moltke, I., Albrechtsen, A., Ko, A., Margaryan, A., Moiseyev, V., et al. (2014). Paleogenomics. Genomic structure in Europeans dating back at least 36,200 years. Science 346, 1113-1118.

Shang, H., Tong, H., Zhang, S., Chen, F., and Trinkaus, E. (2007). An early modern human from Tianyuan Cave, Zhoukoudian, China. Proc. Natl. Acad. Sci. USA 104, 6573-6578.

Sikora, M., Seguin-Orlando, A., Sousa, V.C., Albrechtsen, A., Komeliussen, T., Ko, A., Rasmussen, S., Dupanloup, I., Nigst, P.R., Bosch, M.D., et al. (2017). Ancient genomes show social and reproductive behavior of early Upper Paleolithic foragers. Science 358, 659-662.

Sikora, M., Pitulko, V.V., Sousa, V.C., Allentoft, M.E., Vinner, L. Rasmussen, S., Margaryan, A., de Barros Damgaard, P., de la Fuente, C., Renaud, G., et al. (2019). The population history of northeastern Siberia since the Pleistocene. Nature 570, 182-188.

Siska, V., Jones, E.R., Jeon, S., Bhak, Y., Kim, H.M., Cho, Y.S., Kim, H., Lee, K., Veselovskaya, E., Balueva, T., et al. (2017) Genome-wide data from two early Neolithic East Asian individuals dating to 7700 years ago. Sci. Adv. 3, e1601877.

Skoglund, P., Posth, C., Sirak, K., Spriggs, M., Valentin, F., Bedford, S., Clark, G.R., Reepmeyer, C., Petchey, F., Fernandes, D., et al. (2016). Genomic insights into the peopling of the Southwest Pacific, Nature 538, 510-513.

Stebich, M., Mingram, J., Han, J., and Liu, J. (2009). Late Pleistocene spread of (cool-) temperate forests in Northeast China and climate changes synchronous with the North Atlantic region. Global Planet. Change 65, 56-70.

Tallavaara, M., Luoto, M.. Korhonen, N… Järvinen, H., and Seppä, H. (2015). Human population dynamics in Europe over the Last Glacial Maximum. Proc. Natl. Acad. Sci. USA 112, 8232-8247.

Tan, J., Yang, Y., Tang, K., Sabeti, P.C., Jin, L., and Wang, S. (2013). The adaptive variant EDARV370A is associated with straight hair in East Asians. Hum. Genet. 132, 1187-1191.

Wang, L., and Sebillaud, P. (2019). The emergence of early pottery in East Asia: New discoveries and perspectives. J. World Prehist. 32, 73-110.

Wang, Y., Song, F., Zhu, J., Zhang, S., Yang, Y., Chen, T., Tang, B., Dong, L.. Ding, N., Zhang, Q., et al. (2017). GSA: Genome Sequence Archive. Genomics Proteomics Bioinformatics 15, 14-18.

Wang, C.-C., Yeh, H.-Y., Popov, A.N., Zhang, H.-Q., Matsumura, H., Sirak, K., Cheronet, O., Kovalev, A., Rohland, N., Kim, A.M., et al. (2021). Genomic insights into the formation of human populations in East Asia. Nature 591, 413-419.

Yang, M.A., Gao, X., Theunert, C., Tong, H., Aximu-Petri, A., Nickel, B., Slatkin, M.. Meyer, M., Pääbo, S., Kelso, J., and Fu, Q. (2017). 40,000-Year-Old Individual from Asia Provides Insight into Early Population Structure in Eurasia. Curr. Biol. 27, 3202-3208.e9.

Yang, M.A., Fan, X., Sun, B., Chen, C., Lang, J., Ko, Y.C., Tsang, C.H., Chiu, H., Wang, T., Bao, Q., et al. (2020), Ancient DNA indicates human population shifts and admixture in northern and southern China. Science 369, 282-288.

Ye, K., Gao, F., Wang, D., Bar-Yosef, O., and Keinan, A. (2017). Dietary adaptation of FADS genes in Europe varied across time and geography. Nat. Ecol. Evol. 1, 167.

Yu, H., Spyrou, M.A., Karapetian, M., Shnaider, S., Radzevičiūtė, R., Nägele, K., Neumann, G.U., Penske, S., Zech, J., Lucas, M., et al. (2020). Paleolithic to Bronze Age Siberians Reveal Connections with First Americans and across Eurasia. Cell 181, 1232-1245.e20.

Yue, J.-P., Li, Y.-Q., and Yang, S. (2019). Neolithisation in the southern Lesser Khingan Mountains: lithic technologies and ecological adaptation. Antiquity 93, 1144-1160.

Zhang, H., Chang, F., Li, H., Peng, G., Duan, L., Meng, H., Yang, X., and Wei, Z. (2019). OSL and AMS14C Age of the Most Complete Mammoth Fossil Skeleton from Northeastern China and its Paleoclimate Significance, Radiocarbon 61, 347-358.

Zhao, B., Sun, M., and Du, Z. (2013). The Age and Characteristics of the Jade Unearthed frorn Xiaonanshan Tomb in Raohe County. Bianjiang Kaogu Yanjiu 2013, 69-78.

 

연구 방법 STAR METHODS

핵심 자원 표 KEY RESOURCES TABLE

시약 또는 자원 (REAGENT or RESOURCE) 출처 (SOURCE) 식별자 (IDENTIFIER)
생물학적 샘플 (Biological samples)
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR33K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR19K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR14.5K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR14.1K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR12K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR11K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR11K_deleted
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR10.6K deleted
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR10.6K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR10.5K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR9.9K_2d.rel.AR10.5K_deleted
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR9.2K_o
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR9.2K deleted
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR8.9K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR8.5K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR8.3K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR8.1K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR7.3K LowCov
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR7K_deleted
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR7K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR6.87K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR6.84K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR6.33K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR6.32K
고대 골격 요소 (Ancient skeletal element) 본 논문 (This paper) AR3.4K LowCov
화학물질, 펩타이드, 재조합 단백질 (Chemicals, peptides, and recombinant proteins)
Tween-20 Sigma cat.no. P5927
1 M Tris-HCI, pH 8.0 AppliChem A4577,0500
0.5 M EDTA, PH 8.0 AppliChem A4892,1000
5 M NaCl Sigma S5150
25 mM each dNTP mix Fermentas cat.no. R1121
USER enzyme (Uracil-DNA-glycosylase, UDG와 Endonuclease, EndoVIII의 혼합물) NEB cat.no. M5505L
Bst DNA Polymerase, Large Fragment NEB cat.no. M0275S
ATP, 10 mM stock solution NEB cat.no. 9804
T4 Polynucleotide Kinase (10 U/μL) NEB cat.no. M0236L or S
T4 DNA Polymerase (3 U/μL) NEB cat.no. M0203L
Quick Ligation Kit NEB cat.no. M2200L
2x HI-RPM hybridization buffer Agilent 5118-5380
20% SDS Serva 39575.01
SSC buffer Ambion AM9770
AmpliTaq Gold 10x PCR buffer without Mg Life Technologies 4379874
1M NaOH Sigma 71463-1L
3M Sodium acetate pH 5.2 Sigma S7899
Dynabeads MyOne C1 Life Technologies 65002
SeraMag Speedbeads GE 65152105050250
Cot-1 DNA Invitrogen 15279011
기탁된 데이터 (Deposited data)
새로 시퀀싱된 개체들의 핵 DNA에 대한 BAM 파일과 유전형 콜은 BIG 데이터 센터 게놈 시퀀스 아카이브에서 이용 가능 본 논문 (This paper) PRJCA003699
미토콘드리아 DNA(fasta 형식)는 국립 게놈 데이터 센터의 게놈 웨어하우스에서 이용 가능 본 논문 (This paper) PRJCA003699
올리고뉴클레오타이드 (Oligonucleotides)
1240K 패널용 프로브 (Probe for 1240K Panel) 하크(Haak) et al. (2015)의 보충 데이터 2a https://static-content.springer.com/esm/art%3A10.1038%2Fnature14317/MediaObjects/41586_2015_BFnature14317_MOESM31_ESM.zip
1240K 패널용 프로브 (Probe for 1240K Panel) 하크(Haak) et al. (2015)의 보충 데이터 2b https://static-content.springer.com/esm/art%3A10.1038%2Fnature14317/MediaObjects/41586_2015_BFnature14317_MOESM32_ESM.zip
1240K 패널용 프로브 (Probe for 1240K Panel) 하크(Haak) et al. (2015)의 보충 데이터 2c https://static-content.springer.com/esm/art%3A10.1038%2Fnature14317/MediaObjects/41586_2015BFnature14317_MOESM33_ESM.zip
1240K 패널용 프로브 (Probe for 1240K Panel) 하크(Haak) et al. (2015)의 보충 데이터 2d https://static-content.springer.com/esm/art%3A10.1038%2Fnature14317/MediaObjects/415862015_BFnature14317_MOESM34_ESM.zip
1240K 패널용 프로브 (Probe for 1240K Panel) 푸(Fu) et al. (2015)의 보충 데이터 1a https://static-content.springer.com/esm/art%3A10.1038%2Fnature14558/MediaObjects/41586_2015_BFnature14558_MOESM242_ESM.zip
1240K 패널용 프로브 (Probe for 1240K Panel) 푸(Fu) et al. (2015)의 보충 데이터 1b https://static-content.springer.com/esm/art%3A10.1038%2Fnature14558/MediaObjects/41586_2015_BFnature14558_MOESM243_ESM.zip
1240K 패널용 프로브 (Probe for 1240K Panel) 푸(Fu) et al. (2015)의 보충 데이터 1c https://static-content.springer.com/esm/art%3A10.1038%2Fnature14558/MediaObjects/41586_2015_BFnature14558_MOESM244_ESM.zip
Phosphate-AGATCGGAAG[C3Spacer]10[TEG-biotin] (TEG triethylene glycol spacer) 간사우게(Gansauge) and 마이어(Meyer), 2013 CL78 단일 가닥 어댑터 (Single-stranded adaptor)
GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT 간사우게(Gansauge) and 마이어(Meyer), 2013 CL130 확장 프라이머 (Extension primer)
CGACGCTCTTC-ddC (ddC = dideoxycytidine) 간사우게(Gansauge) and 마이어(Meyer), 2013 CL53 이중 가닥 어댑터, 가닥 1
PhosphateGGAAGAGCGTCGTGTAGGGAAAGAGT GTA 간사우게(Gansauge) and 마이어(Meyer), 2013 CL73 이중 가닥 어댑터, 가닥 2
ACACTCTTTCCCTACACGACGCTCTTCCGATCT G*T C T 마이어(Meyer) and 키르허(Kircher), 2010 NI7_P5_CR2_short
GTGACTGGAGTTCAGACGTG TGCTCTTCCGATCT G T C T 마이어(Meyer) and 키르허(Kircher), 2010 NI8_P7_CR2_short
AGACAGATCG*G*A*A 마이어(Meyer) and 키르허(Kircher), 2010 NI9_P5P7_CR2_comp
GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT-Phosphate 푸(Fu) et al., 2013a, 2015 BO4.P7.part1.R
CAAGCAGAAGACGGCATACGAGAT-Phosphate 푸(Fu) et al., 2013a, 2015 BO6.P7.part2.R
GTGTAGATCTCGGTGGTCGCCGTATCATT-Phosphate 푸(Fu) et al., 2013a, 2015 BO8.P5.part1.R
AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT-Phosphate 푸(Fu) et al., 2013a, 2015 BO10.P5.part2.R
GGAAGAGCGTCGTGTAGGGAAAGAGTGT-Phosphate 양(Yang) et al., 2017 BO11.P5.part2.R
소프트웨어 및 알고리즘 (Software and algorithms)
leeHom 르노(Renaud) et al., 2014 https://github.com/grenaud/leeHom/; RRID: SCR_002710
BWA 0.6.1 리(Li) and 더빈(Durbin), 2009 https://bio-bwa.sourceforge.net/; RRID: SCR_010910
ContamMix 1.0-10 푸(Fu) et al., 2013b https://github.com/StanfordBioinformatics/DEFUNCT-env-modules/tree/master/contamMix/
ANGSD 0.921 코르넬리우센(Korneliussen) et al., 2014 http://popgen.dk/angsd/index.php/ANGSD/
READ 몬로이 쿤(Monroy Kuhn) et al., 2018 https://bitbucket.org/tguenther/read/src/master/
lcMLkin 리파토프(Lipatov) et al., 2015 https://github.com/COMBINE-lab/maximum-likelihood-relatedness-estimation
EIGENSOFT 6.1.4 패터슨(Patterson) et al., 2006 https://github.com/DReichLab/EIG/; RRID: SCR_004965
ADMIXTURE 1.3.0 알렉산더(Alexander) et al., 2009 http://dalexander.github.io/admixture/download.html; RRID: SCR_001263
PLINK v1.9 퍼셀(Purcell) et al., 2007 https://www.cog-genomics.org/plink/1.9/; RRID: SCR_001757
ADMIXTOOLS (qp3Pop, qpDstat, qpAdm, qpGraph) 패터슨(Patterson) et al., 2012 https://github.com/DReichLab/AdmixTools/; RRID:SCR_018495
Treemix 1.13 피크렐(Pickrell) and 프리처드(Pritchard), 2012 http://bitbucket.org/nygcresearch/treemix/wiki/Home/
Phylip 3.697 바움(Baum), 1989 https://evolution.genetics.washington.edu/phylip.html; RRID:SCR_006244
admixfrog 피터(Peter), 2020 https://github.com/benjaminpeter/admixfrog-sims/
hapROH 링바우어(Ringbauer) et al., 2020 https://pypi.org/project/hapROH/

 

자원 이용 가능성 RESOURCE AVAILABILITY

주요 연락처 Lead contact

Further information and requests for resources and reagents should be directed to and will be fulfilled by the Lead Contact, Qiaomei Fu (fuqiaomei@ivpp.ac.cn).

추가 정보 및 자원, 시약에 대한 요청은 주요 연락처인 푸, 차오메이(Fu, Qiaomei) (fuqiaomei@ivpp.ac.cn)에게 연락하면 처리될 것이다.

물질 이용 가능성 Materials availability

This study did not generate new reagents.

본 연구는 새로운 시약을 생성하지 않았다.

데이터 및 코드 이용 가능성 Data and code availability

Original data for the 25 newly sequenced individuals: Aligned reads (BAM format) and genotype calls (Eigenstrat format) of the nuclear DNA have been deposited in the BIG Data Center Genome Sequence Archive https://ngdc.cncb.ac.cn/gsa-human/; accession number: PRJCA003699); Mitochondrial DNA (fasta format) have been deposited in the Genome Warehouse in the National Genomics Data Center https://ngdc.cncb.ac.cn/gwh/; accession number: PRJCA003699) (CNCB-NGDC Members and Partners, 2021; Wang et al., 2017).

새롭게 시퀀싱된 25개 개체의 원본 데이터: 핵 DNA의 정렬된 리드(BAM 형식)와 유전형 콜(Eigenstrat 형식)은 BIG 데이터 센터 게놈 시퀀스 아카이브(accession number: PRJCA003699)에 기탁되었다. 미토콘드리아 DNA(fasta 형식)는 국립 게놈 데이터 센터의 게놈 웨어하우스(accession number: PRJCA003699)에 기탁되었다 (CNCB-NGDC 구성원 및 파트너, 2021; 왕(Wang) et al., 2017).

실험 모델 및 대상 상세 정보 EXPERIMENTAL MODEL AND SUBJECT DETAILS

유적지 및 표본 설명Sites and specimen descriptions

The Songnen Plain (45-48°N, 124-127°E) in southwest Heilongjiang Province, northeast China, includes the majority of the Amur region, located between the Daxing’an and Xiaoxing’an mountains, the Changbai Mountains and the Songliao watershed. The plain lies 120-300 m above sea level with a relative relief of 30-100 m. The Nen and Songhua Rivers flow to the west and south, forming widely distributed floodplains. The central plain is relatively low and flat, with numerous wetlands and both large and small lakes and ponds. Sediments are mainly formed by alluvial deposits of the Songhua and Nen Rivers. The present surface of the Songnen Plain is gently undulating, thus it is also called the Wavy Plain.

중국 동북부 흑룡강성(黑龍江省) 남서부에 위치한 송눈평원(松嫩平原) (북위 45-48°, 동경 124-127°)은 대흥안령(大興安嶺)과 소흥안령(小興安嶺), 장백산(長白山)과 송료(松遼) 분수계 사이에 위치한 아무르(Amur) 지역의 대부분을 포함한다. 이 평원은 해발 120-300m에 위치하며 상대적 고저차는 30-100m이다. 눈강(嫩江)과 송화강(松花江)이 서쪽과 남쪽으로 흐르며 광범위한 범람원을 형성한다. 중앙 평원은 비교적 낮고 평평하며, 수많은 습지와 크고 작은 호수 및 연못이 있다. 퇴적물은 주로 송화강과 눈강의 충적 퇴적물로 형성된다. 현재 송눈평원의 표면은 완만하게 기복이 있어 파상평원(Wavy Plain)이라고도 불린다.

Geologically, the Songnen Plain is a depression, part of the Songliao fault depression zone. The southwestern part of the depression is still sinking, while the northeast is rising. Tertiary and Quaternary sediments have risen to form platforms and uplifted hills. The upper part of the stratigraphic profile in this area of the Songnen Plain consists of loess-like loam, and the lower part is a gravel layer. Black soil (chernozëm) is widely distributed on the surface, rendering the plain a fertile agricultural zone today. The Songnen Plain currently has a hemiboreal humid continental climate (Köppen climate classification Dwa and Dwb), with little rain in winter but is not exceptionally dry, since seasonal moisture from the Sea of Japan and the Yellow Sea reaches the area. The northwestern Songnen Plain is cold and dry because the prevailing northwest wind blows from the arid Eurasian continental interior due to the high pressure cell (the Siberian Anticyclone) established over Siberia. Summers are hot and rainy as water vapor from the sea is carried inland by subtropical high pressure effects, forming precipitation. The vegetation is mainly shrub meadowland or mesophytic weed meadowland.

지질학적으로 송눈평원은 송료 단층 함몰대의 일부인 함몰 지대이다. 이 함몰 지대의 남서부는 계속 침강하고 있는 반면, 북동부는 융기하고 있다. 제3기와 제4기 퇴적물이 솟아올라 대지와 융기된 언덕을 형성했다. 이 지역 송눈평원 지층 단면의 상부는 황토와 유사한 양토로 구성되어 있고, 하부는 자갈층이다. 표면에는 흑토(체르노젬)가 널리 분포하여 오늘날 이 평원을 비옥한 농업 지대로 만들고 있다. 송눈평원은 현재 냉대 습윤 기후(쾨펜 기후 구분 Dwa, Dwb)를 보인다. 겨울에는 비가 적지만 동해와 황해에서 오는 계절성 수분 덕분에 유달리 건조하지는 않다. 송눈평원의 북서부는 시베리아에 자리 잡은 고기압(시베리아 고기압) 때문에 건조한 유라시아 대륙 내부에서 우세한 북서풍이 불어와 춥고 건조하다. 여름에는 아열대 고기압의 영향으로 바다의 수증기가 내륙으로 이동하여 강수 현상을 형성하므로 덥고 비가 많이 온다. 식생은 주로 관목 초원 또는 중생식물 잡초 초원이다.

Due to the impacts of climate change and human activities, groundwater levels have dropped substantially during the past few decades. Surface runoff intensely erodes soft sediments during the summer rainy season, and ravines dissect many areas of the plain. The Songnen Plain is not only a fertile land, but also one of the earliest cattle management/domestication centers in the world (Zhang et al., 2019). It is also extremely rich in fossils of both the Mammuthus-Coelodonta Fauna (Zhang et al., 2019) and of human beings, especially in the drainage of the Songhua River near Harbin, the capital of China’s Heilongjiang Province.

기후 변화와 인간 활동의 영향으로 지난 수십 년 동안 지하수 수위가 크게 낮아졌다. 여름 장마철에는 지표 유출수가 연약한 퇴적물을 심하게 침식하고, 계곡이 평원의 많은 지역을 갈라놓는다. 송눈평원은 비옥한 땅일 뿐만 아니라 세계에서 가장 오래된 소 관리/가축화 중심지 중 하나이다(장(Zhang) et al., 2019). 또한 이곳은 매머드-털코뿔소 동물군(Mammuthus-Coelodonta Fauna)과 인류의 화석이 매우 풍부하며, 특히 중국 흑룡강성의 성도인 하얼빈(哈爾濱) 근처 송화강 유역에 많다(장(Zhang) et al., 2019).

The vegetational and climatic changes of the last glacial-Early Holocene transition have been reconstructed in detail for the Amur region using high-resolution pollen analyses from Lake Sihailongwan in Jilin Province (Mingram et al., 2018; Stebich et al., 2009). Between 16,700 and 14,450 cal BP, the Amur region exhibited predominantly steppe and open taiga-like woodland communities, with abundant Larix, Alnus fruticosa, Betula, Artemisia, grasses and sedges, representing cold and dry conditions. During the Early Holocene, dense deciduous forests, consisting mainly of thermophilic broadleaf trees, became widespread, indicating climatic amelioration. Between 12,680-11,650 cal BP, a short-term climatic reversal to colder and/or dryer conditions has been recorded. A decrease in broadleaved trees and the reappearance of Larix and Picea signals this climatic reversal, comparable to the Younger Dryas event in the circum-Atlantic region. In general, strong atmospheric coupling between the North Atlantic region and East Asia is supported by the synchronicity of climate changes (Stebich et al., 2009).

마지막 빙하기에서 홀로세 초기로 전환되는 시기의 식생 및 기후 변화는 지린성(吉林省) 사해룡만(四海龍灣) 호수의 고해상도 꽃가루 분석을 통해 아무르 지역에 대해 상세히 복원되었다(밍그람(Mingram) et al., 2018; 스테비히(Stebich) et al., 2009).  16,700년에서 14,450년 전(cal BP) 사이, 아무르 지역은 주로 스텝과 개방된 타이가 같은 삼림 군락을 보였다. 이 시기에는 낙엽송(Larix), 오리나무(Alnus fruticosa), 자작나무(Betula), 쑥(Artemisia), 풀, 사초 등이 풍부하여 춥고 건조한 환경이었음을 나타낸다. 홀로세 초기에는 주로 따뜻한 기후를 선호하는 활엽수로 구성된 울창한 낙엽수림이 널리 퍼져 기후가 좋아졌음을 시사한다. 12,680년에서 11,650년 전 사이에는 더 춥고 건조한 조건으로 단기간 기후가 역전된 기록이 있다. 이 기후 역전은 활엽수가 감소하고 낙엽송과 가문비나무(Picea)가 다시 나타나는 것으로 확인되며, 이는 환대서양 지역의 영거 드라이아스(Younger Dryas) 사건과 비교할 만하다. 일반적으로 북대서양 지역과 동아시아 사이의 강력한 대기 연결성은 기후 변화의 동시성으로 뒷받침된다(스테비히(Stebich) et al., 2009).

Pottery appeared in the Amur region as early as 15 ka (e.g., at the Taoshan, Huayang and Houtaomuga sites) (Wang and Sebillaud, 2019; Yue et al., 2019), and is typologically similar to early ceramics discovered in the Russian Far East at archaeological sites including Goncharka-1, Novotroitskoe-10, Oshinovaya Rechika-10, Oshinovaya Rechika-16, Gasya, Khummi, and Gromatukha (Sato and Natsuki, 2017; Wang and Sebillaud, 2019). Functionally, pottery provided a new means of food storage and cooking and improved the efficiency of food processing. Socially, pottery facilitated feasting behavior and the transport of consumables, establishing and strengthening social networks (Hayden and Villeneuve, 2011; Pearson, 2005). In the Amur region, isotope analyses of pottery found in the Houtaomuga (13-5 ka) (Wang and Sebillaud, 2019) and Shuangta sites (11-7.8 ka) (Kunikita et al., 2017) indicate that fishing was an important part of the subsistence strategy of local populations. The earliest archaeological occurrence of nephrite has been identified at the Xiaonanshan site, dating to ca. 9 ka (Jiamusi Cultural Relics Management Station and Raohe County Cultural Relics Management Institute, 1996; Zhao et al., 2013), also in the Amur region, signaling not only the exploitation of this new resource, but also, potentially, the establishment of trade networks, further enhancing nascent regional social connections (Heilongjiang Provincial Institute of Cultural Relics and Archaeology Commission for Preservation of Ancient Monuments, Raohe County, 2019).

아무르 지역에서는 이르면 15,000년 전에 토기가 나타났으며(예: 도산(陶山), 화양(樺陽), 후투무가(後套木嗄) 유적), 이는 러시아 극동의 고고학 유적지(곤차르카-1, 노보트로이츠코에-10 등)에서 발견된 초기 토기와 유형학적으로 유사하다(왕(Wang) and 세비요(Sebillaud), 2019; 위에(Yue) et al., 2019; 사토(Sato) and 나츠키(Natsuki), 2017). 기능적으로 토기는 음식 저장과 조리의 새로운 수단을 제공했고 음식 가공의 효율성을 향상시켰다. 사회적으로 토기는 연회 행위와 소비재 운송을 용이하게 하여 사회적 연결망을 구축하고 강화했다(헤이든(Hayden) and 빌뇌브(Villeneuve), 2011; 피어슨(Pearson), 2005). 아무르 지역의 후투무가(13,000-5,000년 전)와 쌍타(雙塔, 11,000-7,800년 전) 유적지에서 발견된 토기의 동위원소 분석은 어로가 지역 인구의 중요한 생존 전략의 일부였음을 나타낸다(왕(Wang) and 세비요(Sebillaud), 2019; 쿠니키타(Kunikita) et al., 2017). 가장 오래된 고고학적 연옥(nephrite)은 아무르 지역의 소남산(小南山) 유적지에서 확인되었으며, 그 연대는 약 9,000년 전으로 거슬러 올라간다(자무쓰(佳木斯) 문물관리참 및 라오허현(饒河縣) 문물관리소, 1996; 자오(Zhao) et al., 2013). 이는 이 새로운 자원의 활용뿐만 아니라, 잠재적으로는 교역망의 구축을 시사하며, 초기 지역 사회의 연결을 더욱 강화했음을 보여준다(흑룡강성 문물고고연구소 라오허현 고대기념물보존위원회, 2019).

In this study, 25 human remains, ranging from 33,590 to 3,420 cal BP (calibrated years before present, relative to 1950 CE) were selected for ancient DNA testing. All samples included in this study are under the custodianship of the co-authors and were accessed with full permission from the relevant archaeological institutes or universities. The review board of the Institute of Vertebrate Paleontology and Paleoanthropology approved the ancient genomes sampled in this project for study (Review No. 202006020007), following a protocol used in other ancient human genetic studies (Ding et al., 2020; Yang et al., 2020).

본 연구에서는 33,590년에서 3,420년 전(cal BP, 서기 1950년 기준 보정 연대) 범위의 25개 인류 유해를 고대 DNA 테스트 대상으로 선정했다. 이 연구에 포함된 모든 샘플은 공저자들의 관리하에 있으며, 관련 고고학 연구소 또는 대학의 완전한 허가를 받아 접근했다. 고척추동물 및 고인류 연구소의 심의위원회는 다른 고대 인류 유전 연구에서 사용된 프로토콜에 따라, 이 프로젝트에서 샘플링된 고대 게놈의 연구를 승인했다(심의 번호 202006020007) (딩(Ding) et al., 2020; 양(Yang) et al., 2020).

Samples were AMS radiocarbon (14C) dated at Peking University in Beijing.  The chemical preparation and physical measurement of bone samples were similar to the standard processes described by Zhang et al. (2019). The resulting 14C dates were calibrated using OxCal v4.4 (Ramsey and Lee, 2013) and the IntCal20 calibration curve (Reimer et al., 2020). The origin sites of these samples are presented in Table 1. Among them, specimens AR33K, AR9.2K_o, AR9.2K_deleted and AR8.3K were found on the southern bank of the Songhua River; the others were collected at Zhaodong, on the northern bank of the Songhua River. All were found at construction sites, or in eroded gullies and small valleys, especially during summer rainy seasons.  Considering the ages and places of recovery of these specimens, we can see that 1) the oldest came from the Acheng District and east of Harbin on the south side of the Songhua River, where the elevation is relatively higher than surrounding areas; 2) most specimens came from the Zhaodong area on the northern Songhua River, especially the high plateau areas which are also very rich in non-human vertebrate fossils.

샘플들은 북경(北京)의 북경대학(北京大學)에서 가속기 질량 분석기 방사성 탄소(14C) 연대 측정을 거쳤다. 뼈 샘플의 화학적 준비와 물리적 측정은 장(Zhang) 등이 기술한 표준 과정과 유사했다(2019). 결과로 나온 14C 연대는 OxCal v4.4 (램지(Ramsey) and 리(Lee), 2013)와 IntCal20 보정 곡선 (라이머(Reimer) et al., 2020)을 사용하여 보정되었다. 이 샘플들의 출토지는 표 1에 제시되어 있다. 그중 AR33K, AR9.2K_o, AR9.2K_deleted, AR8.3K 표본은 송화강 남쪽 강둑에서 발견되었고, 나머지는 송화강 북쪽 강둑의 자오둥(肇東)에서 수집되었다. 모든 표본은 건설 현장이나, 특히 여름 장마철에 침식된 작은 골짜기에서 발견되었다.  이 표본들의 연대와 발견 장소를 고려하면, 1) 가장 오래된 것은 송화강 남쪽의 아청구(阿城區)와 하얼빈 동쪽에서 나왔는데, 이곳은 주변 지역보다 고도가 비교적 높다. 2) 대부분의 표본은 송화강 북쪽의 자오둥 지역에서 나왔으며, 특히 비인류 척추동물 화석도 매우 풍부한 고원 지대에서 발견되었다.

방법 상세 정보 (METHOD DETAILS)

고대 DNA 추출 Ancient DNA extraction

All human remains were processed in dedicated laboratories at the Chinese Academy of Sciences’ Institute of Vertebrate Paleontology and Paleoanthropology (IVPP) in Beijing and were subjected to DNA extraction, following established protocols, from less than 100 mg of bone (Yang et al., 2017, 2020). Single-stranded libraries (denoted as “SS”) (Gansauge and Meyer, 2013; Kircher et al., 2012) were prepared for all samples except two individuals (NE58 and NE57), which used a uracil-DNA glycosylase partially-treated double-stranded library protocol (denoted as “DS-half” in Table 1) (Meyer and Kircher, 2010; Rohland et al., 2015). We then amplified libraries for 35 cycles using the AccuPrime Pfx polymerase to obtain enough ancient DNA for capture. The P5 and P7 primers were added last to avoid library contamination, and the amount of DNA extracted per sample was evaluated using a Thermo Scientific NanoDrop 2000 spectrometer.

모든 인류 유해는 북경(北京)에 있는 중국과학원(中國科學院) 고척추동물 및 고인류 연구소(IVPP)의 전용 실험실에서 처리되었다. 100mg 미만의 뼈에서 기존 프로토콜에 따라 DNA 추출이 진행되었다 (양(Yang) et al., 2017, 2020). 두 개체(NE58, NE57)를 제외한 모든 샘플에 대해 단일 가닥 라이브러리(“SS”로 표기)가 준비되었으며 (간사우게(Gansauge) and 마이어(Meyer), 2013; 키르허(Kircher) et al., 2012), 이 두 개체에는 우라실-DNA 글리코실레이즈(UDG)로 부분 처리된 이중 가닥 라이브러리 프로토콜(“DS-half”로 표기)이 사용되었다 (마이어(Meyer) and 키르허(Kircher), 2010; 롤란드(Rohland) et al., 2015). 그 후, 포획에 충분한 양의 고대 DNA를 얻기 위해 AccuPrime Pfx 중합효소를 사용하여 라이브러리를 35 사이클 증폭했다. 라이브러리 오염을 피하기 위해 P5와 P7 프라이머는 마지막에 추가했으며, 샘플당 추출된 DNA 양은 Thermo Scientific NanoDrop 2000 분광계를 사용하여 평가했다.

고대 DNA 포획 및 시퀀싱 Ancient DNA capture and sequencing

We applied a DNA capture technique to enrich endogenous ancient DNA from the high levels of background environmental DNA (Fu et al., 2013a, 2015; Haak et al., 2015). For the mitochondrial DNA (mtDNA), oligonucleotide probes synthesized from the complete human mitochondrial genome were used to capture human mtDNA (Fu et al., 2013a); oligonucleotide probes that targeted 1.2 million SNPs from Panels 1 and 2 (Fu et al., 2015) were used to capture human nuclear DNA. After enrichment, the Illumina MiSeq sequencing platform was used to generate 2 x 76bp paired-end reads for mtDNA, and the Illumina Hiseq 4000 sequencing platform was used to generate 2 x 100bp and 2 x 150bp paired-end reads for the nuclear DNA.

높은 수준의 배경 환경 DNA로부터 내인성 고대 DNA를 농축하기 위해 DNA 포획 기술을 적용했다 (푸(Fu) et al., 2013a, 2015; 하크(Haak) et al., 2015). 미토콘드리아 DNA(mtDNA)의 경우, 완전한 인간 미토콘드리아 게놈으로부터 합성된 올리고뉴클레오타이드 프로브를 사용하여 인간 mtDNA를 포획했다 (푸(Fu) et al., 2013a). 패널 1과 2에서 120만 개의 SNP를 표적으로 하는 올리고뉴클레오타이드 프로브를 사용하여 인간 핵 DNA를 포획했다 (푸(Fu) et al., 2015). 농축 후, 일루미나(Illumina) MiSeq 시퀀싱 플랫폼을 사용하여 mtDNA에 대해 2 x 76bp 쌍끝 리드(paired-end reads)를 생성했고, 일루미나(Illumina) Hiseq 4000 시퀀싱 플랫폼을 사용하여 핵 DNA에 대해 2 x 100bp 및 2 x 150bp 쌍끝 리드를 생성했다.

정량화 및 통계 분석 QUANTIFICATION AND STATISTICAL ANALYSIS

리드 정렬 Read alignment

We trimmed adapters and merged paired-end reads into a single sequence (minimum overlap of 11 base pairs) using leeHom software (Renaud et al., 2014), and only merged reads with a length of at least 30 bp were used. For read alignment, BWA (version 0.6.1) (Li and Durbin, 2009) was applied using the aln and samse commands with the arguments -n 0.01 and -l 16500. The mtDNA reads were aligned to the revised Cambridge Reference Sequence (rCRS) (Andrews et al., 1999); Nuclear DNA reads were aligned to the human reference genome hg 19. Duplicate reads were identified (same orientation, start, and end positions) and removed. Reads with a minimum mapping quality score of 30 were kept for analysis.

leeHom 소프트웨어(르노(Renaud) et al., 2014)를 사용하여 어댑터를 제거하고 쌍끝 리드를 단일 시퀀스로 병합했으며(최소 중첩 11bp), 길이가 30bp 이상인 병합된 리드만 사용했다. 리드 정렬을 위해 BWA (버전 0.6.1) (리(Li) and 더빈(Durbin), 2009)를 aln 및 samse 명령어와 인수 -n 0.01 및 -l 16500을 사용하여 적용했다. mtDNA 리드는 수정된 케임브리지 참조 서열(rCRS)에 정렬했고(앤드류스(Andrews) et al., 1999); 핵 DNA 리드는 인간 참조 게놈 hg19에 정렬했다. 중복된 리드(동일한 방향, 시작 및 끝 위치)는 식별하여 제거했다. 최소 매핑 품질 점수가 30인 리드만 분석에 사용했다.

오염 평가 및 변이 콜링 Contamination evaluation and variants calling

Two criteria were used to evaluate contamination: 1. We used ContamMix software (Fu et al., 2013b) to estimate contamination rate by comparing mtDNA fragments with the consensus mitochondrial genome for our newly sampled individuals and 311 present-day world-wide sequences. To avoid counting damaged bases as contamination, the first and last five positions of the fragments were ignored during estimation. The libraries were treated as contaminated if over 3% of the fragments matched present-day sequences better than the consensus;  2. For males, we also used ANGSD software (Korneliussen et al., 2014) to estimate contamination rate, based on the fact that one copy of the X chromosome is found in males. The libraries were treated as contaminated if the estimated contamination rate was greater than 3%. During the variant calling, the first and last five positions of the fragments were ignored. For each SNP covered at least twice in an individual, we randomly sampled one sequence to determine a single allele for that individual in order to obtain pseudo-haploid genotypes. Afterward, individuals with a number of SNPs lower than 25,000 were removed. A total of 21 out of 25 individuals were retained for further analyses.

오염을 평가하기 위해 두 가지 기준이 사용되었다.  1. ContamMix 소프트웨어(푸(Fu) et al., 2013b)를 사용하여 새로 샘플링된 개체의 mtDNA 조각을 합의된 미토콘드리아 게놈 및 311개의 현존 전 세계 시퀀스와 비교하여 오염률을 추정했다. 손상된 염기를 오염으로 간주하는 것을 피하기 위해, 추정 과정에서 조각의 처음과 마지막 5개 위치는 무시했다. 조각의 3% 이상이 합의 서열보다 현존 서열과 더 잘 일치하면 해당 라이브러리는 오염된 것으로 처리했다.  2. 남성의 경우, X 염색체가 하나만 있다는 사실에 근거하여 ANGSD 소프트웨어(코르넬리우센(Korneliussen) et al., 2014)를 사용하여 오염률을 추정했다. 추정된 오염률이 3%보다 크면 해당 라이브러리는 오염된 것으로 처리했다. 변이 콜링 과정에서 조각의 처음과 마지막 5개 위치는 무시했다. 한 개체에서 최소 두 번 이상 커버된 각 SNP에 대해, 무작위로 하나의 시퀀스를 샘플링하여 해당 개체의 단일 대립유전자를 결정함으로써 의사-반수체 유전형을 얻었다. 이후, SNP 수가 25,000개 미만인 개체는 제거했다. 25개 개체 중 총 21개가 추가 분석을 위해 남겨졌다.

친족 관계 분석 Kinship analyses

We used READ software (Monroy Kuhn et al., 2018) to estimate the degrees of kinship among our 21 newly sampled individuals from the Amur region. This tool carried out kinship analyses in three steps: 1. The genome was divided into non-overlapping windows of 1 mega base pairs (Mb). For each 1 Mb window, the proportion of non-matching alleles (PO) for each pair of individuals was estimated;  2. The average PO was normalized by the median of all average pairwise PO across all unrelated individuals. This normalization could reduce the effect of SNP ascertainment, population diversity, and potential batch effects;  3. The degrees of kinship were categorized (unrelated, first degree, second degree, or identical individuals/identical twins) using predefined cut-offs. All unrelated individuals were kept for subsequent analyses. For relatives (identical twins, first-degree and second-degree kin), only the one with the higher SNP number was kept for subsequent analyses. Following this analysis, we removed one individual with second-degree relationship with another, keeping the individual having the higher number of SNPs. To further validate the kinship results from READ, we estimated the degree of kinship using lcMLkin (Lipatov et al., 2015), which infers kinship using genotype likelihoods. The method takes into account the uncertainty in genotype calling when sequence coverage is low, which is especially suitable for ancient DNA samples. lcMLkin estimates k0,k1,k2 as the probability that no, one or two alleles are shared by identical by descent (IBD) between the pair of individuals, respectively. The coefficient of relatedness (r) can then be calculated as r = k1/2 + k2 to infer the kinship category.

아무르(Amur) 지역에서 새로 샘플링된 21개 개체 간의 친족 관계 등급을 추정하기 위해 READ 소프트웨어(몬로이 쿤(Monroy Kuhn) et al., 2018)를 사용했다. 이 도구는 세 단계로 친족 관계 분석을 수행했다.  1. 게놈을 겹치지 않는 1메가베이스(Mb) 크기의 창으로 나누었다. 각 1Mb 창에 대해, 각 개체 쌍의 불일치 대립유전자 비율(PO)을 추정했다.  2. 평균 PO는 모든 관련 없는 개체들의 평균 쌍별 PO 중앙값으로 정규화했다. 이 정규화는 SNP 확인 편향, 집단 다양성 및 잠재적 배치 효과의 영향을 줄일 수 있었다.  3. 미리 정의된 기준치를 사용하여 친족 관계 등급(관련 없음, 1촌, 2촌 또는 동일 개체/일란성 쌍둥이)을 분류했다. 모든 관련 없는 개체들은 후속 분석을 위해 유지했다. 친족(일란성 쌍둥이, 1촌, 2촌)의 경우, SNP 수가 더 높은 개체만 후속 분석에 사용했다. 이 분석에 따라, 다른 개체와 2촌 관계인 개체 한 명을 제거하고 SNP 수가 더 많은 개체를 남겼다. READ의 친족 관계 결과를 추가 검증하기 위해, 유전형 우도(genotype likelihoods)를 사용하여 친족 관계를 추론하는 lcMLkin(리파토프(Lipatov) et al., 2015)을 사용했다. 이 방법은 시퀀스 커버리지가 낮을 때 유전형 콜링의 불확실성을 고려하므로, 특히 고대 DNA 샘플에 적합하다. lcMLkin 분석법은 두 사람을 짝지어, 이들이 조상으로부터 똑같이 물려받은 대립유전자(IBD, identical by descent)가 몇 개인지 확률로 계산한다. 여기서 k0는 똑같이 물려받은 대립유전자가 없을 확률, k1은 하나 있을 확률, k2는 두 개 있을 확률을 각각 의미한다. 이 확률값들을 이용해 ‘친족 계수(r)’를 계산할 수 있다. 계산식은 r = k1/2 + k2이며, 이 계수(r)를 통해 최종적으로 두 사람이 얼마나 가까운 친척인지(친족 범주)를 추론한다.

주성분 분석 Principal components analysis (PCA)

Principal components were calculated with 64 present-day populations from the Human Origin (HO) SNP Panel (Table S1), together with 16 present-day Han and Tibetan populations from Lu et al. (2016), utilizing the smartpca program in the EIGENSOFT package (Patterson et al., 2006). The parameters we adopted were default settings except for Isqproject: YES, numoutlieriter: 0, and shrinkmode: YES. Our newly sequenced ancient individuals (Table 1) and the previously reported ancient Asians (Table S1) were then projected onto the principal components calculated using present-day populations.

EIGENSOFT 패키지(패터슨(Patterson) et al., 2006)의 smartpca 프로그램을 이용하여 인간 기원(HO) SNP 패널(표 S1)의 현존 64개 인구 집단과 루(Lu) 등의 연구(2016)에 나온 16개의 현존 한족(漢族) 및 티베트(Tibetan) 인구 집단 데이터를 사용하여 주성분을 계산했다. 우리가 채택한 매개변수는 lsqproject: YES, numoutlieriter: 0, shrinkmode: YES를 제외하고는 기본 설정을 따랐다. 이후 새로 시퀀싱된 고대 개체(표 1)와 이전에 보고된 고대 아시아인(표 S1)을 현존 인구 집단을 사용하여 계산된 주성분 위에 투영했다.

ADMIXTURE 분석

Individual ancestries were estimated using a model-based maximum likelihood clustering algorithm with ADMIXTURE software (Alexander et al., 2009). The populations included in the analysis were the same as those in the PCA, except that Burmese was omitted and several present-day Russian populations were added (Table S1). Prior to analysis, the genotypes for these populations were pruned for high linkage disequilibrium (r2 > 0.4) using PLINK (version v1.90) (Purcell et al., 2007), with parameters “-indep-pairwise 200 25 0.4.” This left 597,573 SNPs. The ADMIXTURE analysis was then carried out with K from 2 to 10. For each K, the analyses were run 100 times with different seeds to estimate the cross-validation (CV) error. The best K was determined by the lowest CV error.

ADMIXTURE 소프트웨어(알렉산더(Alexander) et al., 2009)의 모델 기반 최대 우도 클러스터링 알고리즘을 사용하여 개별 계통을 추정했다. 분석에 포함된 인구 집단은 버마(Burmese)가 제외되고 여러 현존 러시아 인구 집단이 추가된 것을 제외하고는 PCA와 동일했다(표 S1). 분석에 앞서, PLINK (버전 v1.90) (퍼셀(Purcell) et al., 2007)를 사용하여 높은 연관 불균형(r2 > 0.4)을 보이는 유전형들을 제거했으며, 이때 매개변수는 “-indep-pairwise 200 25 0.4″를 사용했다. 그 결과 597,573개의 SNP가 남았다. 이후 K값을 2에서 10까지 변화시키며 ADMIXTURE 분석을 수행했다. 각 K값에 대해, 교차 검증(CV) 오류를 추정하기 위해 각기 다른 시드(seed)로 100회씩 분석을 실행했다. 가장 낮은 CV 오류를 보인 K값이 최적의 K로 결정되었다.

아웃그룹-f3 및 D 통계 Outgroup-f3 and D statistics

We assessed the amount of shared genetic similarity between two individuals relative to an African outgroup population (Mbuti) using an outgroup-f3 analysis (Patterson et al., 2012). In this analysis, higher f3 values indicate high genetic similarity. The pairwise outgroup-f3 was performed for our newly sampled individuals, and several published ancient and present-day Asians (Table S1), with qp3Pop (version 435) software in the AdmixTools package (Patterson et al., 2012). The pairwise f3 results were then presented as a heatmap using R package gplots (https://github.com/talgalili/gplots). We also assessed the genetic affinities among three individuals using D statistics (Patterson et al., 2012), in the form of D(P1, P2; P3, Outgroup), where Outgroup is a population that is an outgroup to populations P1, P2, and P3. The D statistics were carried out for our newly sampled individuals, and several published ancient and present-day Asians, with qpDstat (version 755) software in the AdmixTools package (Patterson et al., 2012).

아프리카(African) 아웃그룹 인구(Mbuti)를 기준으로 두 개체 간의 공유 유전적 유사성의 양을 아웃그룹-f3 분석(패터슨(Patterson) et al., 2012)을 사용하여 평가했다. 이 분석에서 f3 값이 높을수록 유전적 유사성이 높다는 것을 의미한다. AdmixTools 패키지(패터슨(Patterson) et al., 2012)의 qp3Pop(버전 435) 소프트웨어를 사용하여 새로 샘플링된 개체들과 여러 기존에 발표된 고대 및 현존 아시아인(표 S1)에 대해 쌍별 아웃그룹-f3 분석을 수행했다. 쌍별 f3 결과는 R 패키지 gplots를 사용하여 히트맵으로 제시했다.  또한 D(P1, P2; P3, Outgroup) 형태의 D 통계(패터슨(Patterson) et al., 2012)를 사용하여 세 개체 간의 유전적 친연성을 평가했다. 여기서 Outgroup은 P1, P2, P3 집단에 대한 외집단이다. D 통계는 AdmixTools 패키지의 qpDstat(버전 755) 소프트웨어를 사용하여 새로 샘플링된 개체들과 여러 기존 발표된 고대 및 현존 아시아인에 대해 수행되었다.

Treemix를 이용한 계통 모델링 Phylogeny modeling with Treemix

We estimated historical relationships among populations, allowing both population splits and migration events using maximum-likelihood based software Treemix v1.13 (Pickrell and Pritchard, 2012). The samples included in the Treemix analysis were: AR33K, AR19K, AR14K, ARpost14K, AR13-10K, and ARpost9K (this study) and previously published populations (Mbuti, Ust’-Ishim, Yana, Tianyuan, Ikawazu, DevilsCave_N, Shamanka_EN, Yumin, Mongolia_N_East, Early Neolithic coastal northern East Asians (coastal_nEastAsia_EN) (including Bianbian, Boshan, Xiaogao, and Xiaojingshan), Early Neolithic coastal southern East Asians (sEastAsia_EN) (including Qihe, Liangdao1, and Liangdao2), Late Neolithic coastal southern East Asians (sEastAsia_LN) (including Suogang, Xitoucun, and Tanshishan), Kolyma, UKY, and West_Siberia_N; Table S1). We set Mbuti as the root with (-root Mbuti) and accounted for linkage disequilibrium by grouping sites in blocks of 500 SNPs (-k 500). We allowed for up to five migration events (-m 1-7) and ran 1,000 bootstraps for each tree (-bootstrap -q). Those 1,000 bootstrap trees were then assessed in phylip using the consense command (Baum, 1989), to count the number of times certain individuals grouped, relative to all others in the analysis. The inferred maximum-likelihood based trees and corresponding residuals were visualized with the in-build R script from Treemix v1.13.

최대우도 기반 소프트웨어인 Treemix v1.13(피크렐(Pickrell) and 프리처드(Pritchard), 2012)을 사용하여 인구 분기와 이동 이벤트를 모두 허용하며 인구 간의 역사적 관계를 추정했다. Treemix 분석에 포함된 샘플은 AR33K, AR19K, AR14K, ARpost14K, AR13-10K, ARpost9K (본 연구) 및 이전에 발표된 인구 집단(음부티(Mbuti), 우스트-이심(Ust’-Ishim) 등, 표 S1 참조)이었다. 음부티를 루트로 설정하고(-root Mbuti), 연관 불균형을 고려하기 위해 500개 SNP 단위로 묶어 분석했다(-k 500). 최대 5개의 이동 이벤트를 허용하고(-m 1-7), 각 트리에 대해 1,000번의 부트스트랩을 실행했다(-bootstrap -q). 이 1,000개의 부트스트랩 계통수는 phylip 프로그램의 ‘consense’ 명령어를 사용해 평가했다. 이 과정은 특정 개체들이 다른 모든 개체들과 비교하여 몇 번이나 같은 그룹으로 묶이는지 그 횟수를 세기 위한 것이었다. 추론된 최대우도 기반 계통수와 그에 해당하는 잔차는 Treemix v1.13에 내장된 R 스크립트로 시각화했다.

qpAdm을 이용한 혼합 모델링 Admixture modeling with qpAdm

We applied qpAdm (version 634) in the AdmixTools package (Patterson et al., 2012) to model ancestry proportions of Ancient Paleo-Siberians (Sikora et al., 2019; Yu et al., 2020) related to one, two, or three different sources. qpAdm models ancestry of target populations using a set of source populations (left populations) and a set of reference populations (right populations), without requiring the explicit description of the phylogenetic relationship among them. A “rotating” scheme of source and reference populations was used, being suggested as the best practice to search for the most optimal model (Harney et al., 2021). Therefore, we rotated populations from our set of source populations to a base set of reference populations to keep only one, two, or three populations as source populations, along with parameters “allsnps: YES”, “details: YES”, and “summary: YES”. Our source populations included: AR19K, AR14K, Bianbian, Qihe, Liangdao2, AfontovaGora3, Malta1, and USR1.

AdmixTools 패키지의 qpAdm(버전 634)(패터슨(Patterson) et al., 2012)을 적용하여 고대 고시베리아인(시코라(Sikora) et al., 2019; 유(Yu) et al., 2020)의 계통 비율을 하나, 둘 또는 세 개의 다른 출처와 관련하여 모델링했다. qpAdm은 대상 집단의 계통을 일련의 출처 집단(left populations)과 참조 집단(right populations)을 사용하여 모델링하며, 이들 간의 계통 관계에 대한 명시적인 설명이 필요 없다. 가장 최적의 모델을 찾기 위한 최상의 방법으로 제안된, 출처 집단과 참조 집단을 “회전(rotating)”시키는 방식을 사용했다(하니(Harney) et al., 2021). 따라서 우리는 출처 집단 세트에서 기본 참조 집단 세트로 집단을 회전시켜 하나, 둘 또는 세 개의 집단만 출처 집단으로 유지했으며, 이때 “allsnps: YES”, “details: YES”, “summary: YES” 매개변수를 함께 사용했다. 우리의 출처 집단에는 AR19K, AR14K, 변변인(Bianbian), 기하인(Qihe), 양도인2(Liangdao2), 아폰토바 고라 3호인(AfontovaGora3), 말타 1호인(Malta1), USR1이 포함되었다.

Our source populations included: AR19K, AR14K, Bianbian, Qihe, Liangdao2, AfontovaGora3, Malta1, and USR1. The ‘‘rotating’’ scheme could compete for the optimal sources, e.g., AR19K versus AR14K, and AfontovaGora3 versus USR1.

우리의 출처 집단에는 AR19K, AR14K, 변변인(Bianbian), 기하인(Qihe), 양도인2(Liangdao2), 아폰토바 고라 3호인(AfontovaGora3), 말타 1호인(Malta1), USR1이 포함되었다. “회전(rotating)”시키는 방식은 최적의 출처를 찾기 위해 경쟁할 수 있었다. 예를 들어 AR19K 대 AR14K, 아폰토바 고라 3호인 대 USR1과 같은 방식이다.

Our base set of reference populations, which must differ in their relationship to the target populations, included: Mota, UstIshim, Kostenki14, Iran_N, IndusPeriphery, LBK_EN, Motala12, Kotias, AR33K, Yana, and Karelia.

대상 집단과의 관계가 서로 달라야 하는 우리의 기본 참조 집단 세트에는 모타(Mota), 우스트-이심(UstIshim), 코스텐키14(Kostenki14), 이란_N(Iran_N), 인더스 주변부(IndusPeriphery), LBK_EN, 모탈라12(Motala12), 코티아스(Kotias), AR33K, 야나(Yana), 카렐리야(Karelia)가 포함되었다.

IndusPeriphery is a merged set of three individuals from two archaeological sites: Gonur Depe (Gonur2_BA) (4 ka), in the Bactria- MargianaArchaeologicalComplexinTurkmenistan,andShahr_I_Sokhta_BA2andShahr_I_Sokhta_BA3(5-4ka)fromShahr-i-Sokhta in Iran (Narasimhan et al., 2019).

인더스 주변부(IndusPeriphery)는 두 고고학 유적지에서 나온 세 명의 개체를 합친 집단이다. 하나는 투르크메니스탄(Turkmenistan)의 박트리아-마르기아나 고고학 문화권에 있는 고누르 데페(Gonur Depe) 유적지(Gonur2_BA, 약 4천 년 전)이고, 다른 하나는 이란(Iran)의 샤르이소흐타(Shahr-i-Sokhta) 유적지(Shahr_I_Sokhta_BA2, Shahr_I_Sokhta_BA3, 약 5천-4천 년 전)이다.

We only considered a model with more source populations when one with fewer sources was rejected. The criteria used to reject a model with fewer sources were: 1. Tail probability of rank0 for a fewer source model > 0.05; 2. The estimated admixture proportions (±standard error) were between 0 and 1.

우리는 더 적은 수의 출처를 가진 모델이 기각되었을 때만 더 많은 출처를 가진 모델을 고려했다. 더 적은 수의 출처를 가진 모델을 기각하는 기준은 다음과 같았다: 1. 더 적은 출처 모델에 대한 rank0의 꼬리 확률(Tail probability) > 0.05; 2. 추정된 혼합 비율(±표준오차)이 0과 1 사이일 것.

In addition, we applied qpAdm to model ancestry proportions of outlier population AR9.2K_o and merged outlier populations based on PCA and outgroup-f3 analyses (denoted as AR_o, including AR9.2K_o, AR3.4K_LowCov, and AR7.3K_LowCov) related to one, two, or three different sources.

또한, 우리는 qpAdm을 적용하여 특이 집단인 AR9.2K_o와, PCA 및 아웃그룹-f3 분석에 기반하여 합쳐진 특이 집단들(AR_o로 표기하며, AR9.2K_o, AR3.4K_LowCov, AR7.3K_LowCov 포함)의 계통 비율을 하나, 둘 또는 세 개의 다른 출처와 관련하여 모델링했다.

Our source populations included: AR19K, AR14K, Bianbian, Xiaojingshan, Qihe, Liangdao2.

이때 우리의 출처 집단에는 AR19K, AR14K, 변변인(Bianbian), 소형산인(Xiaojingshan), 기하인(Qihe), 양도인2(Liangdao2)가 포함되었다.

Our base set of reference populations was the same as before, and included: Mota, UstIshim, Kostenki14, Iran_N, IndusPeriphery, LBK_EN, Motala12, Kotias, AR33K, Yana, and Karelia.

우리의 기본 참조 집단 세트는 이전과 동일했으며, 모타(Mota), 우스트-이심(UstIshim), 코스텐키14(Kostenki14), 이란_N(Iran_N), 인더스 주변부(IndusPeriphery), LBK_EN, 모탈라12(Motala12), 코티아스(Kotias), AR33K, 야나(Yana), 카렐리야(Karelia)가 포함되었다.

qpGraph를 이용한 인구 통계 모델링 Demographic modeling with qpGraph

We followed a general approach for modeling admixture graphs using qpGraph (version 6065) in the AdmixTools package (Patterson et al., 2012), which started with a basic and well-understood tree (including the central African Mbuti as an outgroup, the early western Eurasian Kostenki14, ancient North Siberian Yana, and early Asian Tianyuan). We then added extra populations (AR33K, GoyetQ116-1, Liangdao2, AR19K, AR14K, Bianbian, DevilsCave_N, AfontovaGora3, USR1, UKY, and Kolyma) one at a time in their best-fitting positions iteratively (Yang et al., 2020). An optimum tree model could be constructed based on the observed f-statistics (f2, f3 and f4 for all possible pairs of populations). A |Z-score| of less than 3 between the observed and expected values (determined by the Block Jackknife) was required. A small number (0.0001) was added to the diagonal entries of the estimated covariance matrix of the f-statistics (Q matrix) to stabilizing the matrix inversion. The above and other parameters applied in the qpGraph program are as below, following recommendations of Lipson (2020):

우리는 AdmixTools 패키지(패터슨(Patterson) et al., 2012)의 qpGraph(버전 6065)를 사용하여 혼합 그래프를 모델링하는 일반적인 접근 방식을 따랐다. 이는 기본적이고 잘 알려진 계통수에서 시작했다. 이 계통수에는 중앙아프리카의 음부티(Mbuti)를 외집단으로, 초기 서부 유라시아인 코스텐키14(Kostenki14), 고대 북시베리아인 야나(Yana), 초기 아시아인 전원인(Tianyuan)이 포함된다. 그 후 추가 인구 집단(AR33K, GoyetQ116-1, Liangdao2, AR19K, AR14K, 변변인(Bianbian), 악마의 문 동굴인(DevilsCave_N), 아폰토바 고라 3호인(AfontovaGora3), USR1, UKY, 콜리마(Kolyma))을 가장 적합한 위치에 반복적으로 하나씩 추가했다 (양(Yang) et al., 2020). 최적의 계통수 모델은 관찰된 f-통계량(모든 가능한 인구 쌍에 대한 f2, f3, f4)을 기반으로 구축될 수 있었다. 관찰값과 기대값(블록 잭나이프로 결정됨) 사이의 |Z-점수|는 3 미만이어야 했다. 행렬 역산을 안정화시키기 위해 f-통계량의 추정된 공분산 행렬(Q 행렬)의 대각선 항목에 작은 수(0.0001)를 추가했다. 위 내용 및 qpGraph 프로그램에 적용된 다른 매개변수는 립슨(Lipson) (2020)의 권장 사항에 따라 아래와 같다:

outpop: Mbuti
blgsize: 0.05
lsqmode: YES
diag: 0.0001
hires: YES
initmix: 1000
precision: 0.0001
zthresh: 0
terse: NO
useallsnps: NO

We first added AR33K in the above basic tree. Due to strong evidence (from outgroup-f3, D statistics, and Treemix analyses) that AR33K shares close genetic affinity with Tianyuan, we started with a model of AR33K forming a cluster with Tianyuan to reduce the searching space of initial modeling. Next, we added GoyetQ116-1, Liangdao2, AR19K, AR14K, Bianbian, DevilsCave_N, AfontovaGora3, USR1, UKY, and Kolyma iteratively, fitting as all possible nodes or admixture between nodes.

우리는 먼저 위의 기본 트리에 AR33K를 추가했다. 아웃그룹-f3, D 통계, Treemix 분석에서 AR33K가 전원인과 가까운 유전적 친연성을 공유한다는 강력한 증거가 있었기 때문에, 초기 모델링의 탐색 공간을 줄이기 위해 AR33K가 전원인과 하나의 그룹을 형성하는 모델에서 시작했다. 다음으로, GoyetQ116-1, Liangdao2, AR19K, AR14K, 변변인, 악마의 문 동굴인, 아폰토바 고라 3호인, USR1, UKY, 콜리마를 가능한 모든 노드 또는 노드 간의 혼합으로 맞춰가며 반복적으로 추가했다.

은닉 마르코프 모델을 이용한 고인류 계통 추정 Archaic ancestry estimation with the Hidden Markov Model

We inferred archaic (Neanderthal and Denisovan) introgressed segments in the genomes of our newly sampled individuals with SNP numbers over 330,000 (including AR33K, AR19K, AR14.5K, AR14.1K, AR10.6K, AR10.5K, AR8.9K, AR8.1K, AR7K, AR6.87K, and AR6.33K), Tianyuan (Yang et al., 2017), and two present-day Han Chinese from the Simons Genome Diversity Project (SGDP) (Mallick et al., 2016) (Table S1), using admixfrog software (Peter, 2020). This software adopts a Hidden Markov Model for local ancestry inference, and allows accurate detection of introgressed segments in low-coverage ancient genomes even with present-day human DNA contamination (Peter, 2020). We modeled the target individuals deriving their ancestries from three sources including two high-coverage Neanderthal genomes (NEA) (Prüfer et al., 2014, 2017), one high-coverage Denisovan genome (Denisova 3) (DEN) (Meyer et al., 2012), and 44 genomes of present-day Sub-Saharan Africans from SGDP (AFR) (Mallick et al., 2016).

우리는 admixfrog 소프트웨어(피터(Peter), 2020)를 사용하여, SNP 수가 33만 개 이상인 새로 샘플링된 개체들(AR33K, AR19K 등 11개체), 전원인(Tianyuan), 그리고 시몬스 게놈 다양성 프로젝트(SGDP)의 현대 한족(漢族) 2명의 게놈에서 고인류(네안데르탈인 및 데니소바인)로부터 유입된 유전적 구간을 추론했다. 이 소프트웨어는 국소적 계통 추론을 위해 은닉 마르코프 모델을 채택하며, 현생인류 DNA 오염이 있는 저커버리지 고대 게놈에서도 유입된 구간을 정확하게 탐지할 수 있다. 우리는 분석 대상 개체들의 계통이 세 가지 출처에서 유래했다고 모델링했다. 그 출처는 고커버리지 네안데르탈인 게놈 2개(NEA), 고커버리지 데니소바인 게놈 1개(Denisova 3, DEN), 그리고 SGDP의 현대 사하라 이남 아프리카인 44명의 게놈(AFR)이다.

The modeling included two steps:

모델링은 두 단계로 진행됐다.

 

1. The filtered BAM files for each of the target individual were converted into the input files for admixfrog. The command run is:

각 분석 대상의 필터링된 BAM 파일을 admixfrog 입력 파일로 변환했다. 실행 명령어는 다음과 같다:

admixfrog bam {bamfile_individual}.bam -length-bin-size 35 -minmapq 25 -deam-cutoff 3 -out {individual.in.xz} -ref {ref_archaicadmixture.csv.xz}

In the above command, for each target individual, the BAM files were filtered to only include fragments of at least 35 base pairs (bp) and with a mapping quality >= 25 (-length-bin-size 35 -minmapq 25). Additionally, variant positions matching a C to T substitutions in the first three positions of the read, or a G to A substitutions at the end of a read were discarded (-deam-cutoff 3).

위 명령어에서, 각 대상의 BAM 파일은 조각 길이가 최소 35bp이고 매핑 품질이 25 이상인 것만 포함하도록 필터링되었다 (-length-bin-size 35 -minmapq 25). 또한, 리드의 처음 세 위치에서 C가 T로 치환되거나 리드의 끝에서 G가 A로 치환된 변이 위치는 폐기되었다 (-deam-cutoff 3).

2. For inferring archaic introgressed segments using previously generated input files, the following command was run:

이전에 생성된 입력 파일을 사용하여 고인류 유입 구간을 추론하기 위해 다음 명령어를 실행했다:

admixfrog infile {individual.in.xz} -ref {ref_archaicadmixture.csv.xz} -out {out_individual} -states AFR NEA DEN -cont-id AFR -ll-tol 0.01 -bin-size 5000 -est-F -est-tau -est-freq -F3-freq -contamination 3 -e0 0.01 -ancestral PAN -run-penalty 0.1 -max-iter 250 -npost-replicates 200 -filter-pos 50 -filter-map 0

In the above, we used the chimpanzee (panTro4, GCA_000001515.4) reference genome to infer the ancestral state of each allele (-ancestral PAN). The window size was set to 5,000 base pairs for all individuals (-bin-size 5000). The detailed information for other parameters was suggested and explained in Peter (2020).

위 명령어에서, 우리는 침팬지(panTro4, GCA_000001515.4) 참조 게놈을 사용하여 각 대립유전자의 조상 상태를 추론했다 (-ancestral PAN). 모든 개체에 대해 분석 창 크기는 5,000bp로 설정했다 (-bin-size 5000). 다른 매개변수에 대한 자세한 정보는 피터(Peter) (2020)의 연구에 제안되고 설명되어 있다.

아무르(Amur) 지역의 시간 경과에 따른 인구 크기 Time transect of population sizes in the Amur region

To investigate the time transect of population size changes, we identified the ROH for the populations in the Amur region using hapROH software (Ringbauer et al., 2020), which is based on a Hidden Markov Model (HMM) with ROH and non-ROH states and utilizes a panel of 5,008 phased haplotypes of the 1000 Genomes Project dataset (Phase 3, release 20130502) (Auton et al., 2015). The short ROH (4-8 cM) represent patterns of past background relatedness that could be used as proxies for local population sizes. To achieve reliable results, we analyzed 11 newly sequenced ancient individuals with SNP numbers over 330,000 (including AR33K, AR19K, AR14.5K, AR14.1K, AR10.6K, AR10.5K, AR8.9K, AR8.1K, AR7K, AR6.87K, and AR6.33K) and two modern Ulchi individuals (S_Ulchi-1 and S_Ulchi-2). The effective population size was directly estimated from observed ROH lengths, using the observed lengths of ROH blocks I=l1,…,ln for a given set of individuals i₁,…, ik and a maximum likelihood inference scheme. The likelihood Pr(I | Ne) was calculated and then the population size was estimated that maximizes the product of the likelihoods. To assess the potential influence of effective population size on the kinship inference, we also merged AR9.9K_2d.rel.AR10.5K_deleted and AR10.5K and estimated their effective population size.

시간 경과에 따른 인구 크기 변화를 조사하기 위해, hapROH 소프트웨어(링바우어(Ringbauer) et al., 2020)를 사용하여 아무르 지역 인구의 동형접합구간(ROH)을 식별했다. 이 소프트웨어는 ROH와 비-ROH 상태를 갖는 은닉 마르코프 모델(HMM)에 기반하며, 1000 게놈 프로젝트 데이터셋(3단계, 20130502 릴리스)의 5,008개 위상 분석된 반수체 패널을 활용한다(오턴(Auton) et al., 2015). 짧은 ROH(4-8 cM)는 과거의 배경 친족 관계 패턴을 나타내며, 이는 지역 인구 크기의 대리 지표로 사용될 수 있다. 신뢰할 수 있는 결과를 얻기 위해, SNP 수가 33만 개 이상인 11개의 새로 시퀀싱된 고대 개체(AR33K, AR19K, AR14.5K, AR14.1K, AR10.6K, AR10.5K, AR8.9K, AR8.1K, AR7K, AR6.87K, AR6.33K 포함)와 두 명의 현대 울치(Ulchi)인(S_Ulchi-1, S_Ulchi-2)을 분석했다. 유효 집단 크기는 관찰된 ROH 길이로부터 직접 추정했다. 주어진 개체 집합 i₁,…, ik에 대한 ROH 블록의 관찰된 길이 $I=l_{1},…,l_{n}$와 최대우도추론 방식을 사용했다. 우도 Pr(I | Ne)를 계산한 다음, 우도의 곱을 최대화하는 인구 크기를 추정했다. 유효 집단 크기가 친족 관계 추론에 미칠 잠재적 영향을 평가하기 위해, 우리는 또한 AR9.9K_2d.rel.AR10.5K_deleted와 AR10.5K를 합쳐 그들의 유효 집단 크기를 추정했다.

북부 동아시아인의 EDAR V370A 시간 경과 Time transect of EDAR V370A in northern East Asians

To observe the time transect of EDAR V370A from 40 ka to present-day, we utilized both the genotype calls and read counts from 18 newly sampled ancient individuals with SNP numbers over 50,000 (Table 1) and 12 published ancient northern East Asians from latitudes above 30 degrees North (including populations from Tianyuan Cave, Bianbian, Boshan, Xiaogao, Xiaojingshan, Yumin, and Devil’s Gate Cave). To obtain diploid genotype calls for EDAR V370A, we used an alternative variant calling scheme using snpAD (Prüfer, 2018). snpAD jointly estimates error rates and genotype frequencies from ancient sequences. These estimates were then used to produce genotype calls. Read counts for EDAR V370A were obtained from BAM files. We grouped individuals based on time windows (40,000-35,000 cal BP, 35,000-30,000 cal BP, 20,000-15,000 cal BP, 15,000-13,000 cal BP, 13,000-10,000 cal BP, 10,000-9,000 cal BP, 9,000-8,000 cal BP, and 8,000-6,000 cal BP) to observe the time transect of EDAR V370A.

4만 년 전부터 현재까지 EDAR V370A의 시간적 경과를 관찰하기 위해, SNP 수가 5만 개 이상인 18개의 새로 샘플링된 고대 개체(표 1)와 북위 30도 이상에서 발견된 12개의 기존에 발표된 고대 북부 동아시아인(전원동굴(田園洞窟), 변변(邊邊), 박산(博山), 소고(曉高), 소형산(小荊山), 유민(裕民), 악마의 문 동굴 인구 포함)의 유전형 콜과 리드 카운트를 모두 활용했다. EDAR V370A에 대한 이배체 유전형 콜을 얻기 위해, snpAD (프뤼퍼(Prüfer), 2018)를 사용하는 대안적인 변이 콜링 방식을 사용했다. snpAD는 고대 시퀀스로부터 오류율과 유전형 빈도를 공동으로 추정한다. 이 추정치를 사용하여 유전형 콜을 생성했다. EDAR V370A에 대한 리드 카운트는 BAM 파일에서 얻었다. EDAR V370A의 시간적 경과를 관찰하기 위해 시간대별(40,000-35,000 cal BP, 35,000-30,000 cal BP, 20,000-15,000 cal BP, 15,000-13,000 cal BP, 13,000-10,000 cal BP, 10,000-9,000 cal BP, 9,000-8,000 cal BP, 8,000-6,000 cal BP)로 개체들을 그룹화했다.

 

Figure S1. Related to Figure 1 and Tables S1 and S3
그림 S1. 그림 1 및 표 S1, S3과 관련됨

(A) Principal components analysis (PCA) (including deleted newly sampled populations due to low quality and kinship). The ancient populations (above) are projected on the principal components of modern populations (below). New samples are highlighted in red text and larger symbols. PCA shows that AR9.9K_2d.rel.AR10.5K_deleted (second-degree relative of AR10.5K) shares a similar position with AR13-10K.

(A) 주성분 분석(PCA) (낮은 품질과 친족 관계로 인해 제외된 새로 샘플링된 인구 포함). 고대 인구(위쪽)는 현대 인구(아래쪽)의 주성분 위에 투영되었다. 새로운 샘플은 붉은색 텍스트와 더 큰 기호로 강조 표시했다. PCA는 AR9.9K_2d.rel.AR10.5K_deleted (AR10.5K의 2촌)가 AR13-10K와 유사한 위치를 공유함을 보여준다.

(B) The pairwise genetic affinity among ancient East Asians using outgroup-f3 analysis with low coverage samples included. Both PCA and outgroup-f3 analysis indicate that AR7.3K_LowCov and AR3.4K_LowCov share similar positions with AR9.2K_o in both PCA and outgroup-f3 analysis.

(B) 저커버리지 샘플을 포함한 아웃그룹-f3 분석을 사용한 고대 동아시아인 간의 쌍별 유전적 친연성. PCA와 아웃그룹-f3 분석 모두에서 AR7.3K_LowCov와 AR3.4K_LowCov는 AR9.2K_o와 유사한 위치를 공유함을 나타낸다.

(C) Calibration of radiocarbon measurement for NE30 (AR9.9K_2d.rel.AR10.5K_deleted, who is the second-degree relative of AR10.5K). The y axis shows 14C dates BP and the x axis shows the calibrated age obtained using 14C dates calibrated using OxCal v4.4 (Ramsey and Lee, 2013) and the IntCal20 calibration curve (Reimer et al., 2020). The temporal fluctuations in atmospheric 14C is reflected by the progression of the 14C calibration curve between approximately 10,200-9,800 cal BP.

(C) NE30 (AR10.5K의 2촌인 AR9.9K_2d.rel.AR10.5K_deleted)에 대한 방사성 탄소 측정 보정 . y축은 BP 기준 ¹⁴C 연대를, x축은 OxCal v4.4와 IntCal20 보정 곡선을 사용하여 보정한 연대를 보여준다. 대기 중 ¹⁴C의 시간적 변동은 약 10,200-9,800 cal BP 사이의 ¹⁴C 보정 곡선의 진행에 반영되어 있다.

(D) Analysis for the comparison of D statistics of AR33K (D(X,Y; AR33K, Mbuti)) and AR19K (D(X, Y,AR19K, Mbuti)) using all SNPs and transversion SNPs only. X and Y refer to the same population set as in Figure S2. x axis is the D or Z values using all SNPs. y axis is the D or Z values using transversion SNPs only. The regression coefficient and coefficient of determination (R2) are labeled. These results indicate that the lower Z values when using transversion SNPs is due only to the larger standard error when using fewer SNPs.

(D) 모든 SNP와 전환(transversion) SNP만을 사용하여 AR33K (D(X,Y; AR33K, 음부티))와 AR19K (D(X, Y, AR19K, 음부티))의 D 통계량을 비교하기 위한 분석. X와 Y는 그림 S2와 동일한 인구 집합을 나타낸다 . x축은 모든 SNP를 사용했을 때의 D 또는 Z 값이다 . y축은 전환 SNP만을 사용했을 때의 D 또는 Z 값이다. 회귀 계수와 결정 계수(R²)가 표시되어 있다. 이 결과는 전환 SNP를 사용할 때 Z 값이 낮아지는 것은 더 적은 수의 SNP를 사용하여 표준 오차가 커졌기 때문임을 나타낸다.

Figure S2. Genetic cluster of AR33K and Tianyuan, demonstrated using D statistics, related to Figure 2
그림 S2. D 통계량을 사용하여 증명된 AR33K와 전원인(田園人)의 유전적 그룹, 그림 2와 관련됨

Three forms of D statistics showed that D(X, Tianyuan; AR33K, Mbuti) < 0 (Z < -3), D(X, AR33K; Tianyuan, Mbuti) < 0 (Z < -3), and D(Tianyuan, AR33K; X, Mbuti) ~0 (-3 < Z < 3). X represents the global ancient (i.e., older than 5 ka) and modern populations. Ancient populations are sorted by age. The dashed line indicates where D = 0. Solid circles indicate that |Z| > 3 and otherwise |Z| < 3. The thick gray bar indicates one standard error and the thin gray bar indicates two standard errors. Directions of deviations of D statistics are labeled with population name on the x axis at the top left and right of figure panels. All SNPs were used to compute D statistics.

세 가지 형태의 D 통계량은 D(X, 전원인; AR33K, 음부티) < 0 (Z < -3), D(X, AR33K; 전원인, 음부티) < 0 (Z < -3), 그리고 D(전원인, AR33K; X, 음부티) ~0 (-3 < Z < 3)임을 보여주었다. X는 전 세계의 고대(5천 년보다 오래된) 및 현대 인구를 나타낸다. 고대 인구는 연대순으로 정렬했다. 점선은 D=0인 지점을 나타낸다. 채워진 원은 |Z| > 3임을, 그렇지 않은 경우는 |Z| < 3임을 나타낸다. 굵은 회색 막대는 1 표준오차를, 얇은 회색 막대는 2 표준오차를 나타낸다. D 통계량 편차의 방향은 그림 패널의 상단 왼쪽과 오른쪽에 있는 x축에 모집단 이름으로 표시되어 있다. D 통계량을 계산하는 데 모든 SNP가 사용되었다.

Figure S3. Inferred maximum-likelihood-based trees using Treemix, related to Figure 3
그림 S3. Treemix를 사용하여 추론된 최대우도 기반 계통수, 그림 3과 관련됨

A maximum likelihood tree allowing four migration events is the most supported using 1,000 bootstrap runs (highlighted with red rectangle). The results from Treemix are consistent with qpGraph modeling.

네 번의 이동 이벤트를 허용하는 최대우도 계통수가 1,000번의 부트스트랩 실행을 통해 가장 많이 지지받았다(붉은색 사각형으로 강조). Treemix의 결과는 qpGraph 모델링과 일치한다.

Figure S4. Admixture plot for ancient and modern East Asians for K = 2 to K = 10 related to Figure 4
그림 S4. K=2부터 K=10까지의 고대 및 현대 동아시아인에 대한 Admixture 플롯, 그림 4와 관련됨

(A) Admixture results for K = 2 to K = 10

(A) K=2부터 K=10까지의 Admixture 결과

(B) Cross-validation results for different K values, where the lowest CV is when K = 4. The genetic profile of AR14K, AR13-10K, and ARpost9K are closest to the DevilsCave_N population.

(B) 다른 K 값에 대한 교차 검증(Cross-validation) 결과, K=4일 때 CV(교차검증) 오류가 가장 낮다. AR14K, AR13-10K, ARpost9K의 유전적 프로필은 악마의 문 동굴인(DevilsCave_N) 인구와 가장 가깝다.