출처:
Population genetic admixture and evolutionary history in the Shandong Peninsula inferred from integrative modern and ancient genomic resources
산동반도 사람들은 어떻게 섞이고 진화했나: 고대·현대 유전체 통합 분석
Haoran Su1,2,3†, Mengge Wang2,4,5*†, Xiangping Li2,6, Shuhan Duan2,5,7, Qiuxia Sun2,8, Yuntao Sun2,9, Zhiyong Wang2,6, Qingxin Yang2,6, Yuguo Huang2, Jie Zhong2, Jing Chen2,10, Xiucheng Jiang2,5,7, Jinyue Ma2,5,7, Ting Yang2,6, Yunhui Liu2,8, Lintao Luo2,8, Yan Liu2,5,7, Junbao Yang5,7, Gang Chen12, Chao Liu11*, Yan Cai-1*† and Guanglin He2,4,5*
소호연(蘇浩然)1,2,3†, 왕몽격(王夢格)2,4,5*†, 이상평(李祥平)2,6, 단주함(段姝含)2,5,7, 손추하(孫秋霞)2,8, 손운도(孫雲濤)2,9, 왕지용(王志勇)2,6, 양경신(楊慶新)2,6, 황옥국(黃玉國)², 종결(鐘潔)², 진정(陳靜)2,10, 장수성(蔣修成)2,5,7, 마금월(馬金月)2,5,7, 양정(楊婷)2,6, 유운혜(劉雲慧)2,8, 나임도(羅林濤)2,8, 유염(劉艷)2,5,7, 양준보(楊俊寶)5,7, 진강(陳剛)¹², 유초(劉超)¹¹*, 채염(蔡焰) -1*† 그리고 하광림(何廣林)2,4,5*
[리뷰] 중화 쇼비니즘 경사도 평가: 5/10
16. Su, H., Wang, M., Li, X. et al. (2024) ‘Population genetic admixture and evolutionary history in the Shandong Peninsula inferred from integrative modern and ancient genomic resources’, BMC Genomics, 25, 611. doi:10.1186/s12864-024-10514-9.
(1) 연구 개요 및 저자의 주장
이 연구는 현대 산동 한족 264명의 유전체 데이터를 신규 생산하고, 이를 기존의 고대 및 현대 동아시아인 데이터와 통합하여 산동반도의 미시적인 인구 혼합사와 적응의 역사를 분석했다. 연구 결과, 산동반도 사람들은 고대 동북아시아인(ANEA)과 강한 유전적 유사성을 보이며, 이는 신석기 초기부터 황하 하류 유역에서 장기적인 유전적 연속성과 이동성이 있었음을 시사한다고 주장한다.(2) 편향성 분석 (중화 쇼비니즘 경사도: 4/10)
현대인 데이터를 포함하면서 발생할 수 있는 해석적 편향의 위험을 내포하고 있다.
- 서사 프레이밍 (중간 편향성): 현대인과 고대인을 통합 분석하는 과정에서, 표본이 내륙에 치우칠 경우 중원 중심의 ‘화살’ 모델이 강화될 위험이 있다.
- 모델 선택과 반례 취급 (중간 편향성): 분석 모델에 타기도와 같은 도서·연해의 반례를 포함하지 않으면, ‘내륙→연해’로의 일방적 확산이라는 결론으로 귀결될 수 있다.
- 지리·환경 제약 반영 (중간 편향성): 최종빙기절정(LGM) 이후의 환경 변화를 구조적 변수로 포함하여, 해양 회랑의 독립성을 인정해야 분석의 객관성을 높일 수 있다.
- 유전자–문화 결합 가정 (높은 편향성): 현대의 ‘한족’이라는 개념을 과거로 투영할 경우, 유전자와 문화의 관계를 지나치게 단순화할 위험이 있다.
(3) 결론 재구성
결론에서 “연해·도서 축의 독자성”을 반드시 분리하여 진술해야 한다. 방법론적으로는 통계 분석(qpAdm) 시 동일한 아웃그룹을 사용하고, 분석 소스에 남방 해안계(SE-Coastal)를 포함하며, 도서 지역 표본에 대해서는 내륙 조상 모델과의 적합도를 비교하는 과정을 거쳐야 한다.
[논문요약]
산동 사람들의 유전자로 본 동아시아인의 역사
핵심 용어 풀이
이 요약을 이해하기 위한 몇 가지 기본 개념이다.
- 유전적 혼합 (Genetic Admixture): 서로 다른 지역에 살던 두 인구 집단이 만나 아이를 낳고, 세대를 거치며 두 집단의 DNA가 물감처럼 섞이는 현상이다. 이 논문은 중국의 남쪽과 북쪽 사람들이 어떻게 섞였는지를 다룬다.
- 고대 DNA (Ancient DNA): 수천 년 전 사람의 뼈나 치아에서 추출한 DNA이다. 과거 인류가 어떤 유전자를 가졌는지, 어디서 왔는지 알려주는 ‘유전학적 타임캡슐’과 같다.
- 자연선택 (Natural Selection): 특정 환경에서 생존하고 자손을 남기는 데 유리한 유전자가 세대를 거치며 그 집단에 많아지는 현상이다. 예를 들어, 추운 지방에서는 추위에 강한 유전자가, 특정 질병이 유행하는 곳에서는 그 병에 강한 유전자가 선택된다.
- SNP (스닙, 단일염기다형성): DNA 서열에서 나타나는 미세한 차이로, 개인과 집단을 구별하는 중요한 표지이다. 과학자들은 SNP를 분석해 인류의 이동 경로와 유전적 역사를 추적한다.
연구의 핵심 질문
중국 문명의 발상지인 황하 유역(黃河流域), 특히 산동반도(山東半島)에 사는 사람들은 유전적으로 어떤 뿌리를 가지고 있을까? 이 연구의 핵심 질문은 다음과 같다.
“오늘날 산동에 사는 한족(漢族)은 수천 년 전 황하 유역에 살던 고대인들과 얼마나 가까울까? 그리고 이들은 어떤 환경에 적응하며 진화해 왔을까?”
연구팀은 현대 산동 사람 264명의 DNA와 함께, 수천 년 전 고대인들의 뼈에서 추출한 ‘고대 DNA’ 및 다른 동아시아인들의 DNA를 통합하여 이 질문에 대한 답을 찾고자 했다.
주요 발견
- 산동 사람들은 ‘황하 유역 토박이’의 후손이다.
연구팀은 수많은 사람들의 DNA(SNP)를 비교하여 서로 얼마나 가깝고 먼지를 보여주는 ‘유전 지도’를 그렸다.
- 그림 1은 이 유전 지도를 보여준다. 지도에서 산동 사람들(SDH, 빨간색 점)은 현대의 다른 북부 한족(漢族) 및 고대 황하 유역 사람들(Ancient YRB)과 매우 가깝게 뭉쳐 있다. 반면, 남부 한족(漢族)과는 꽤 멀리 떨어져 있다.
- 이는 산동 사람들이 신석기 시대부터 황하 유역에 살았던 고대 북부 동아시아인(ANEA)의 유전자를 거의 그대로 물려받았다는 것을 의미한다. 즉, 외부 인구와 많이 섞이지 않고 수천 년간 유전적 명맥을 유지해 온 것이다.
- 산동인의 유전자 ‘레시피’: 북방계 85% + 남방계 15%
물론 수천 년간 완벽한 고립은 없었다. 연구팀은 산동 사람들의 유전자가 어떤 고대 집단으로부터 얼마나 섞였는지 그 비율을 계산했다.
- 그림 2는 산동 사람들의 ‘유전자 레시피’를 파이 차트처럼 보여준다. 분석 결과, 이들의 유전자는 고대 북부 동아시아인(ANEA, 주로 황하 유역의 기장 농사꾼)으로부터 약 85%, 그리고 고대 남부 동아시아인(ASEA, 주로 양자강 유역의 쌀 농사꾼)으로부터 약 15%를 물려받은 것으로 나타났다.
- 이는 과거 북쪽의 기장 문화와 남쪽의 쌀 문화가 교류하면서 사람들의 유전자도 일부 섞였지만, 산동 지역은 북방계의 유전적 특성이 압도적으로 강하게 유지되었음을 보여준다.
- 환경 적응의 흔적: ‘자연선택’이 남긴 특별한 유전자들
인간은 각자 살아가는 환경에 맞춰 진화한다. 연구팀은 산동 사람들의 DNA에서 특히 강하게 ‘자연선택’을 받은 유전자들을 찾아냈다.
그림 3은 게놈 전체에서 자연선택의 ‘핫스팟’을 보여주는 지도와 같다. 가장 높이 솟은 봉우리들이 강한 선택을 받은 유전자들이다.
주목할 만한 유전자 2가지:
- ABCC11 유전자 (귓밥과 체취 유전자)
- 이 유전자는 ‘마른 귓밥’과 ‘적은 겨드랑이 냄새’를 결정한다. 이 유전자 변이는 땀 분비를 줄여 추운 기후에서 체온을 보존하는 데 유리했을 것으로 추정된다.
-
- 그림 4는 이 유전자의 시공간적 확산을 보여준다. 약 44,000년 전 시베리아 근처에서 처음 나타난 이 유전자는 동아시아로 퍼져나가, 현재 동아시아인 대부분이 이 유전자를 갖게 되었다. 이는 추위에 대한 적응의 강력한 증거이다.
- SLC10A1 유전자 (B형 간염 저항성 유전자)
- 이 유전자의 특정 변이(S267F)는 흥미로운 특징을 가진다. 어린 시절 담즙산 대사에 문제를 일으킬 수 있는 대신, B형 간염 바이러스(HBV)에 대한 저항력을 높여준다.
- B형 간염이 흔했던 남부 중국과 동남아시아에서 이 유전자 변이의 비율이 높게 나타난다. 이는 질병이라는 강력한 환경 요인이 인간의 진화에 어떤 영향을 미쳤는지 보여주는 ‘진화의 상충 관계’ 사례이다.
결론
이 논문은 현대인과 고대인의 DNA를 통합 분석하여, 산동 사람들이 수천 년 전 황하 문명을 이끌었던 고대 북방인의 직계 후손에 가깝다는 것을 유전학적으로 증명했다. 또한, 동아시아인 특유의 유전자인 ‘마른 귓밥 유전자(ABCC11)’와 ‘B형 간염 저항성 유전자(SLC10A1)’가 어떤 과정을 거쳐 우리 조상들에게 퍼지게 되었는지 그 진화의 역사를 추적했다. 이는 우리가 특정 환경에 어떻게 적응해 왔는지, 그리고 그 과정이 우리의 유전자와 건강에 어떤 흔적을 남겼는지 보여주는 중요한 연구이다.
[논문번역]
초록 Abstract
배경 Background
Ancient northern East Asians (ANEA) from the Yellow River region, who pioneered millet cultivation, play a crucial role in understanding the origins of ethnolinguistically diverse populations in modern China and the entire landscape of deep genetic structure and variation discovery in modern East Asians. However, the direct links between ANEA and geographically proximate modern populations, as well as the biological adaptive processes involved, remain poorly understood.
기장 농업을 개척한 황하(黃河) 지역의 고대 북부 동아시아인(ANEA)은 현대 중국의 민족언어학적으로 다양한 인구 집단의 기원을 이해하고, 현대 동아시아인의 심층적인 유전 구조와 변이를 발견하는 데 결정적인 역할을 한다. 하지만 고대 북부 동아시아인과 지리적으로 가까운 현대 인구 집단 사이의 직접적인 연관성 및 관련된 생물학적 적응 과정은 아직 잘 알려지지 않았다.
결과 Results
Here, we generated genome-wide SNP data for 264 individuals from geographically different Han populations in Shandong. An integrated genomic resource encompassing both modern and ancient East Asians was compiled to examine fine-scale population admixture scenarios and adaptive traits. The reconstruction of demographic history and hierarchical clustering patterns revealed that individuals from the Shandong Peninsula share a close genetic affinity with ANEA, indicating long-term genetic continuity and mobility in the lower Yellow River basin since the early Neolithic period. Biological adaptive signatures, including those related to immune and metabolic pathways, were identified through analyses of haplotype homozygosity and allele frequency spectra. These signatures are linked to complex traits such as height and body mass index, which may be associated with adaptations to cold environments, dietary practices, and pathogen exposure. Additionally, allele frequency trajectories over time and a haplotype network of two highly differentiated genes, ABCC11 and SLC10A1, were delineated. These genes, which are associated with axillary odor and bilirubin metabolism, respectively, illustrate how local adaptations can influence the diversification of traits in East Asians.
여기서 우리는 산동의 지리적으로 다른 한족(漢族) 인구 264명의 전체 게놈 단일염기다형성(SNP) 데이터를 생성했다. 현대 및 고대 동아시아인을 모두 포함하는 통합된 유전체 자원을 편집하여 미세 규모의 인구 혼합 시나리오와 적응 형질을 조사했다. 인구 역사와 계층적 군집 패턴을 재구성한 결과, 산동반도(山東半島) 사람들은 고대 북부 동아시아인과 가까운 유전적 친화력을 공유하는 것으로 나타났다. 이는 신석기 시대 초기부터 황하(黃河) 하류 유역에서 장기간에 걸친 유전적 연속성과 이동성이 있었음을 시사한다. 일배체형 동형접합성과 대립유전자 빈도 스펙트럼 분석을 통해 면역 및 대사 경로와 관련된 것을 포함한 생물학적 적응의 특징을 확인했다. 이러한 특징들은 키와 체질량지수 같은 복합 형질과 관련이 있으며, 이는 추운 환경, 식습관, 병원체 노출에 대한 적응과 연관될 수 있다. 또한, 시간에 따른 대립유전자 빈도 궤적과 함께 매우 분화된 두 유전자, 즉 ABCC11과 SLC10A1의 일배체형 네트워크를 기술했다. 각각 겨드랑이 냄새와 빌리루빈 대사와 관련된 이 유전자들은 국지적 적응이 어떻게 동아시아인의 형질 다양화에 영향을 미칠 수 있는지 보여준다.
결론 Conclusions
Our findings provide a comprehensive genomic dataset that elucidates the fine-scale genetic history and evolutionary trajectory of natural selection signals and disease susceptibility in Han Chinese populations. This study serves as a paradigm for integrating spatiotemporally diverse ancient genomes in the era of population genomic medicine.
우리의 연구 결과는 한족(漢族) 인구의 미세 규모 유전 역사, 자연선택 신호의 진화 궤적, 그리고 질병 감수성을 규명하는 포괄적인 유전체 데이터세트를 제공한다. 이 연구는 인구 유전체 의학 시대에 시공간적으로 다양한 고대 유전체를 통합하는 패러다임 역할을 한다.
Keywords : Population genetics, Northern Han, Evolution, Genomics, Adaptation
키워드 : 집단유전학, 북부 한족(漢族), 진화, 유전체학, 적응
목차
Genetic structure and population affinity
유전 구조 및 집단 친화도
Ancestral makeup and admixture landscape of Han Chinese from the lower YRB
황하(黃河) 하류 유역 한족(漢族)의 조상 구성 및 혼합 양상
Relative genetic stability and temporal dynamics in the lower YRB since the Neolithic period
신석기 시대 이래 황하(黃河) 하류 유역의 상대적 유전적 안정성 및 시간적 역학
Natural selection and adaptation mechanisms
자연선택과 적응 메커니즘
Highly differentiated variants between northern Han and southern Han people
북부 한족과 남부 한족 간의 고도로 분화된 변이
Haplotype analysis and allele frequency trajectories of ABCC11 and SLC10A1
ABCC11과 SLC10A1의 일배체형 분석 및 대립유전자 빈도 궤적
Demographic history and genetic structure
인구 역사와 유전 구조
Natural selection signals and East Asian-specific variants
자연선택 신호와 동아시아 특이 변이
Sample collection and DNA preparation
표본 수집 및 DNA 준비
Quality control, genotype calling, and dataset merging
품질 관리, 유전자형 판독, 데이터세트 병합
Principal component analysis
주성분 분석(PCA)
FST calculation and TreeMix
FST 계산 및 TreeMix 분석
Haplotype-based population analysis
일배체형(Haplotype) 기반 인구 분석
QpAdm and qpWave
QpAdm 및 qpWave 분석
Identification of natural selection signals
자연선택 신호 식별
배경 Background
East Asia, a center of agricultural domestication during the Neolithic transition, was inhabited by anatomically modern humans at least 50,000 years ago (kya) [1, 2]. This region boasts a complex demographic history and a differentiated genetic architecture of complex traits. A comprehensive understanding of human evolutionary history is essential for elucidating the formation of modern humans and the impact of genetic variation on traits and diseases. These processes include archaic introgression, multiple human dispersal events, admixture, and natural selection [3, 4]. Recent studies utilizing ancient genomic resources have revealed that human expansion from regions such as the Mongolia Plateau, Amur River Basin, Yellow River Basin, Yangtze River Basin, and the Eurasian Steppe has gradually facilitated complex patterns of population structure and genetic diversity in East Asia [5–8]. Evolutionary reconstructions among ethnolinguistically modern populations suggest that differentiated selective pressures and admixture landscapes have enriched the complexity of the population history and biological adaptations of East Asians [6, 9–12].
신석기 전환기 농업 작물화의 중심지였던 동아시아에는 적어도 5만 년 전에 해부학적 현생인류가 거주했다 [1, 2]. 이 지역은 복잡한 인구 역사와 복합 형질의 분화된 유전 구조를 특징으로 한다. 인류 진화 역사에 대한 포괄적인 이해는 현생인류의 형성과 유전 변이가 형질 및 질병에 미치는 영향을 밝히는 데 필수적이다. 이러한 과정에는 고대 인류와의 유전자 교환, 여러 차례의 인류 분산 사건, 혼합, 그리고 자연선택이 포함된다 [3, 4]. 고대 유전체 자원을 활용한 최근 연구들은 몽골 고원(蒙古高原), 아무르강 유역(黑龍江流域), 황하 유역(黃河流域), 양자강 유역(揚子江流域), 유라시아 스텝(Eurasian Steppe)과 같은 지역으로부터의 인류 확장이 동아시아의 복잡한 인구 구조와 유전적 다양성 패턴을 점진적으로 형성했음을 밝혔다 [5–8]. 민족언어학적으로 현대적인 인구 집단들 사이의 진화적 재구성은 분화된 선택압과 혼합 양상이 동아시아인의 인구 역사와 생물학적 적응의 복잡성을 심화시켰음을 시사한다 [6, 9–12].
Yang et al. identified genomic substructures between ancient northern East Asians (ANEAs) and ancient southern East Asians (ASEAs). They also highlighted coastal population migrations and connections from the Russian Far East, coastal China, and Vietnam from the late Pleistocene to the Holocene epoch [13]. Research in population genomics has indicated that complex demographic processes such as effective population size, divergence times, isolation, and migrations are shaped by gene flow between early highly differentiated populations or barriers to forming cultural and ecological differences. Genetic components derived from distinct ancestral sources have increased genomic diversity and phenotype complexity. Post-admixture adaptation in new environments may further influence gene expression profiles and phenotypic traits [14]. Consequently, East-West admixed populations in Central Asia and northwestern Chinese Turkic-speaking people possess unique adaptive landscapes of critical immune and metabolic pathway genes [15]. Environmental factors, including dietary shifts and regional living conditions, have also shaped adaptive variants.
양(Yang) 등은 고대 북부 동아시아인(ANEA)과 고대 남부 동아시아인(ASEA) 사이의 유전체 하부 구조를 확인했다. 그들은 또한 후기 플라이스토세부터 홀로세에 이르기까지 러시아 극동(Russian Far East), 중국 해안, 베트남(Vietnam)으로부터의 해안 인구 이동과 연결성을 강조했다 [13]. 집단 유전체학 연구에 따르면, 유효 집단 크기, 분기 시간, 고립, 이주와 같은 복잡한 인구학적 과정은 초기에 고도로 분화된 집단 간의 유전자 흐름이나 문화적, 생태적 차이를 형성하는 장벽에 의해 결정된다. 서로 다른 조상으로부터 유래한 유전적 요소들은 유전체 다양성과 표현형의 복잡성을 증가시켰다. 새로운 환경에서의 혼합 후 적응은 유전자 발현 양상과 표현형 특성에 추가적인 영향을 미칠 수 있다 [14]. 결과적으로, 중앙아시아(Central Asia)의 동서 혼합 인구와 중국 북서부의 튀르크어족(Turkic-speaking) 사람들은 중요한 면역 및 대사 경로 유전자의 독특한 적응 지형을 가지고 있다 [15]. 식단 변화와 지역적 생활 조건을 포함한 환경 요인 또한 적응 변이를 형성해왔다.
Yang et al. dissected the genetic basis of the skin color phenotype in highland East Asians and reported that the darker baseline skin color in Tibetans was induced by a mutation (rs75356281) under adaptation to strong ultraviolet (UV) radiation [16].
양(Yang) 등은 고지대 동아시아인의 피부색 표현형에 대한 유전적 기초를 분석하여, 티베트인(Tibetan)의 더 어두운 기본 피부색이 강한 자외선(UV) 방사선에 적응하는 과정에서 발생한 돌연변이(rs75356281)에 의해 유발되었다고 보고했다 [16].
In recent decades, extensive genomic studies have explored genetic diversity within European populations, aiming to uncover how genetic disease risks originate from various geographically isolated conditions. FinnGen, a large-scale biobank resource, utilizes isolated populations to identify disease-predisposing alleles prevalent in Finland [17]. The underrepresentation of non-European populations in global human genomic resources and clinical genomic datasets remains a significant concern. Although genomic projects in East Asia, such as GenomeAsia100K [18], Nyuwa [19], and ChinaMAP [20], have played significant roles in discovering genotype-phenotype associations, the representation of genetic diversity remains limited. More fine-scale genomic projects are needed to focus on genetic variation discovery, population history reconstruction, and medical relevance interpretation to dissect the genetic landscape of Chinese populations.
최근 수십 년간, 유럽 인구를 대상으로 한 광범위한 유전체 연구는 다양한 지리적 고립 조건에서 유전 질환의 위험이 어떻게 발생하는지를 밝히는 데 목적을 두었다. 대규모 바이오뱅크 자원인 핀젠(FinnGen)은 고립된 인구를 활용하여 핀란드(Finland)에서 널리 퍼진 질병 유발 대립유전자를 식별한다 [17]. 전 세계 인간 유전체 자원 및 임상 데이터세트에서 비유럽 인구의 대표성이 부족한 점은 여전히 중요한 문제이다. 게놈아시아100K(GenomeAsia100K) [18], 뉴와(Nyuwa) [19], 차이나맵(ChinaMAP) [20]과 같은 동아시아의 유전체 프로젝트가 유전자형-표현형 연관성을 발견하는 데 중요한 역할을 했지만, 유전적 다양성의 대표성은 여전히 제한적이다. 중국 인구의 유전적 지형을 상세히 분석하기 위해서는 유전 변이 발견, 인구 역사 복원, 의학적 관련성 해석에 초점을 맞춘 더 정밀한 유전체 프로젝트가 필요하다.
Current human genetic studies have identified numerous population- or region-specific genetic variants associated with disease susceptibility. The origins of many common and rare genetic diseases are traced to specific ancient lineages and genetic adaptations. In the immune system, the evolution of host- interactions enhances the ability to resist infections but also increases susceptibility to inflammatory diseases [21]. This evolutionary pattern, known as antagonistic pleiotropy, provides fitness benefits for survival in extreme environments but also elevates disease burden. Generally, the physical environment, dietary practices, and exposure to endemic pathogens are persistent driving factors of natural selection and adaptation mechanisms. These factors impose heterogeneous selection pressures on humans, leading to differential, adaptive genes and significant variations in the prevalence of genetic diseases among populations with diverse genetic backgrounds.
현재의 인간 유전학 연구는 질병에 걸릴 가능성과 관련된, 특정 인구나 지역에만 나타나는 수많은 유전 변이를 확인했다. 흔하거나 희귀한 많은 유전 질환의 기원은 특정 고대 혈통과 유전적 적응으로 거슬러 올라간다. 면역 체계의 경우, 숙주와 병원체 간의 상호작용 진화는 감염에 저항하는 능력을 향상시키지만, 동시에 염증성 질환에 대한 취약성을 높이기도 한다 [21]. 이러한 진화 패턴은 ‘길항적 다면발현(antagonistic pleiotropy)’으로 알려져 있으며, 극한 환경에서의 생존에는 이점을 주지만 질병 부담을 높이는 결과를 낳는다. 일반적으로 물리적 환경, 식습관, 그리고 풍토병을 일으키는 병원체에 대한 노출은 자연선택과 적응 메커니즘을 이끄는 지속적인 동력이다. 이러한 요인들은 인류에게 각기 다른 선택압을 가하여, 유전적으로 다양한 집단들 사이에서 서로 다른 적응 유전자를 만들고 유전 질환의 유병률에 상당한 차이를 초래한다.
Han Chinese, the world’s largest ethnic group, is traditionally distributed across seven geographical regions and is divided into two genetically distinct groups: northern Han and southern Han [22]. Previous genetic studies utilizing genome-wide SNP variations, mitochondrial DNA (mtDNA), and Y-chromosomal variations have demonstrated a North-South population structure and varying allele frequencies between northern and southern Han people [23–25]. Furthermore, whole-genome sequencing analyses have identified finer subgroups within Han Chinese individuals, attributable to the population’s complex origins and historically documented interactions with surrounding ethnic groups [20, 26].
세계에서 가장 큰 민족 집단인 한족(漢族)은 전통적으로 7개의 지리적 지역에 분포하며, 유전적으로 뚜렷한 두 집단, 즉 북부 한족(漢族)과 남부 한족(漢族)으로 나뉜다 [22]. 전체 게놈 단일염기다형성(SNP) 변이, 미토콘드리아 DNA(mtDNA), Y염색체 변이를 활용한 이전의 유전 연구들은 북부와 남부 한족(漢族) 사이에 남북 인구 구조와 다양한 대립유전자 빈도가 존재함을 보여주었다 [23–25]. 더욱이, 전체 게놈 시퀀싱 분석은 한족(漢族) 개인들 내에서 더 세분화된 하위 집단을 확인했는데, 이는 이 인구의 복잡한 기원과 주변 민족 집단과의 역사적으로 기록된 상호작용에 기인한다 [20, 26].
The ancestral origins of Han Chinese people remain a subject of ongoing debate. Zhang et al. proposed that the homeland of the Proto-Sino-Tibetan (Proto-ST) language originated in the YRB of northern China, supporting the northern-origin hypothesis of the ST people [27]. The YRB civilization, which encompasses cultures such as Yangshao, Majiayao, and Longshan, significantly influenced extensive regions and gave rise to present-day Han Chinese culture [28, 29]. The genetic patterns observed in northern Han Chinese also have been partly attributed to millet-based farming populations from the West Liao River (WLR) region in northeastern China. The expansion of farming practices increased the genetic affinities between WLR millet farmers and YRB ancients during the late Neolithic period. However, the influence of YRB-related ancestry decreased in Bronze Age populations due to changes in subsistence strategies [30].
한족(漢族)의 조상 기원은 계속 논쟁의 대상이다. 장(Zhang) 등은 중국-티베트어족 조어(Proto-Sino-Tibetan)의 발원지가 북부 중국의 황하 유역(黃河流域)에서 시작되었다고 제안하며, 중국-티베트어족(ST) 사람들의 북방 기원 가설을 지지했다 [27]. 양소(仰韶), 마가요(馬家窯), 용산(龍山)과 같은 문화를 포함하는 황하 유역(黃河流域) 문명은 광범위한 지역에 큰 영향을 미쳤으며 오늘날 한족(漢族) 문화의 기원이 되었다 [28, 29]. 북부 한족(漢族)에서 관찰되는 유전적 패턴은 부분적으로 중국 북동부의 서요하(西遼河) 지역의 기장 기반 농업 인구에 기인한 것으로 여겨진다. 농업 관행의 확산은 신석기 후기 동안 서요하(西遼河) 기장 농부들과 황하 유역(黃河流域) 고대인들 사이의 유전적 친화력을 증가시켰다. 하지만, 생계 전략의 변화로 인해 청동기 시대 인구에서는 황하 유역(黃河流域) 관련 혈통의 영향이 감소했다 [30].
The agricultural systems in China exhibit a dualistic structure, with rice domestication first documented at the Kuahuqiao and Xianrendong sites in the YZRB [31, 32]. The coexistence of millet and rice farming at various archaeological sites indirectly provides evidence for North-South gene flows [33, 34]. In addition to the genetic substructure in the North-South direction, East-West differentiation is also present among Han Chinese individuals. Genetic evidence indicates that massive demic diffusion, the North-South admixture model, and trans-Eurasian complex admixture have all shaped the genetic profile of Han Chinese [2, 22, 25, 35].
중국의 농업 시스템은 이중 구조를 보이며, 벼농사는 양자강 유역(揚子江流域)의 과호교(跨湖橋)와 선인동(仙人洞) 유적에서 처음으로 기록되었다 [31, 32]. 다양한 고고학 유적지에서 기장과 쌀 농업의 공존은 남북 간 유전자 흐름에 대한 간접적인 증거를 제공한다 [33, 34]. 남북 방향의 유전적 하부 구조 외에도, 한족(漢族) 개인들 사이에서는 동서 간의 분화도 존재한다. 유전적 증거는 대규모 인구 확산, 남북 혼합 모델, 그리고 범유라시아적 복합 혼합이 모두 한족(漢族)의 유전적 프로파일을 형성했음을 나타낸다 [2, 22, 25, 35].
Shandong is located on the Shandong Peninsula along the eastern coast of China in the lower Yellow River Basin (YRB), a core region of the ancient millet domestication center and Neolithic transition hotspots. This area faces the Bohai Sea and is situated across the sea from the Korean Peninsula and the Japanese archipelago. Numerous Neolithic cultures, such as the Houli [36], Dawenkou [37], and Longshan [38] cultures, have been identified in this region. Yang et al. illustrated that the genetic background of coastal early Neolithic ANEA in Shandong differed from that of inland Neolithic Yumin and Fujian Neolithic ASEA [13]. Archaeological and genetic evidence has revealed that the late Neolithic society in ancient Shandong Province followed a matrilineal community structure, exhibiting low mitochondrial DNA (mtDNA) diversity but high Y chromosome diversity [37]. By analyzing the appearance of haplogroups C, M9, and F, Liu et al. reported that maternal genetic structure began to change 4600 years before the present (BP) and that the ancestral components of the Bianbian individuals were related to ANEA and ancient Siberian lineages [38].
산동은 중국 동해안의 산동반도(山東半島)에 위치하며, 황하(黃河) 하류 유역에 속한다. 이곳은 고대 기장 작물화의 중심지이자 신석기 전환의 핵심 지역이다. 이 지역은 발해(渤海)와 마주하고 있으며, 바다 건너 한반도(韓半島) 및 일본 열도(日本列島)와 인접해 있다. 이 지역에서는 후리(后李) [36], 대문구(大汶口) [37], 용산(龍山) [38] 문화 등 수많은 신석기 문화가 확인되었다. 양(Yang) 등은 산동 해안의 신석기 초기 고대 북부 동아시아인(ANEA)의 유전적 배경이 내륙 신석기 시대의 유민(Yumin)과 복건(福建) 신석기 시대의 고대 남부 동아시아인(ASEA)과는 달랐음을 보여주었다 [13]. 고고학적 및 유전적 증거에 따르면 고대 산동성(山東省)의 후기 신석기 사회는 모계 중심의 공동체 구조를 따랐으며, 낮은 미토콘드리아 DNA(mtDNA) 다양성과 높은 Y염색체 다양성을 보였다 [37]. 리우(Liu) 등은 하플로그룹 C, M9, F의 출현을 분석하여, 모계 유전 구조가 현재로부터 4600년 전에 변하기 시작했으며, 변변(卞邊) 유적 개인들의 조상 구성 요소가 고대 북부 동아시아인 및 고대 시베리아 혈통과 관련이 있다고 보고했다 [38].
However, comprehensive population genetic studies of modern populations in the Shandong Peninsula, such as the Shandong Han (SDH) population, are lacking. Additionally, research on the genetic connection between geographically close ancient and modern populations is limited due to the sparse sampling of present-day individuals. The genetic origins of the Han Chinese people in the YRB, the phylogenetic relationships between the northern Han people and their geographical neighbors, and their biological adaptation signatures remain poorly characterized. Investigating the fine-scale genetic structure of the northern Han people in the YRB can offer unique insights into specific disease evolutionary events. To address these gaps, SDH was chosen as a representative population of the Northern Han population. An extensive population genetic analysis was conducted, incorporating data from other Northern Han populations located in Shanxi, Henan, and Shaanxi provinces. This study aimed to investigate the northern Han population structure, identify biological adaptive signals, and reveal the evolutionary origins of East Asian-specific diseases using ancient and modern genetic data. A total of 264 Han individuals from various cities in Shandong Province were genotyped using an Affymetrix array (Supplementary Fig. 1a). The genome-wide single nucleotide polymorphism (SNP) data of SDH were integrated with those from spatiotemporally diverse ancient East Asians and other linguistically diverse modern East Asian populations, including the Altaic, Tai-Kadai (TK), Hmong-Mien (HM), Austronesian (AN), Austroasiatic (AA), and Sino-Tibetan (ST) groups. This research illuminated the impact of genetic interactions between ancient and modern East Asians on the gene pool of SDH and explored the admixture processes and adaptation mechanisms of SDH. Changes in allele frequencies associated with traits over time were observed, and the trajectories of allele frequency changes at highly differentiated variants (HDVs) are depicted. By comparing ancient genomes spatiotemporally, valuable insights were provided into the genetic continuity and mobility of Han Chinese from the YRB lineage.
하지만, 산동 한족(山東漢族)과 같은 산동반도(山東半島)의 현대 인구에 대한 포괄적인 집단 유전학 연구는 부족하다. 또한, 현대인의 표본 추출이 부족하여 지리적으로 가까운 고대와 현대 인구 간의 유전적 연결에 대한 연구는 제한적이다. 황하 유역(黃河流域) 한족(漢族)의 유전적 기원, 북부 한족(漢族)과 지리적 이웃 간의 계통발생학적 관계, 그리고 그들의 생물학적 적응 특징은 여전히 명확히 규명되지 않았다. 황하 유역(黃河流域) 북부 한족(漢族)의 미세 규모 유전 구조를 조사하는 것은 특정 질병의 진화적 사건에 대한 독특한 통찰을 제공할 수 있다. 이러한 공백을 메우기 위해, 산동 한족(山東漢族)이 북부 한족(漢族)의 대표 집단으로 선정되었다. 산서(山西), 하남(河南), 섬서(陝西)성에 위치한 다른 북부 한족(漢族) 집단의 데이터를 통합하여 광범위한 집단 유전학 분석을 수행했다. 이 연구는 고대 및 현대 유전 데이터를 사용하여 북부 한족(漢族)의 인구 구조를 조사하고, 생물학적 적응 신호를 식별하며, 동아시아 특이 질병의 진화적 기원을 밝히는 것을 목표로 했다. 산동성(山東省) 여러 도시 출신의 한족(漢族) 264명을 대상으로 애피메트릭스(Affymetrix) 어레이를 사용하여 유전자형을 분석했다. 산동 한족(山東漢族)의 전체 게놈 단일염기다형성(SNP) 데이터는 시공간적으로 다양한 고대 동아시아인 및 알타이어족(Altaic), 타이-카다이어족(Tai-Kadai), 몽-미엔어족(Hmong-Mien), 오스트로네시아어족(Austronesian), 오스트로아시아어족(Austroasiatic), 중국-티베트어족(Sino-Tibetan) 그룹을 포함한 언어적으로 다양한 현대 동아시아 인구의 데이터와 통합되었다. 이 연구는 고대와 현대 동아시아인 간의 유전적 상호작용이 산동 한족(山東漢族)의 유전자 풀에 미친 영향을 조명하고, 산동 한족(山東漢族)의 혼합 과정과 적응 메커니즘을 탐구했다. 시간에 따른 형질 관련 대립유전자 빈도의 변화가 관찰되었으며, 고도로 분화된 변이(HDV)에서의 대립유전자 빈도 변화 궤적이 묘사되었다. 고대 유전체를 시공간적으로 비교함으로써, 황하 유역(黃河流域) 혈통의 한족(漢族)의 유전적 연속성과 이동성에 대한 귀중한 통찰을 제공했다.
결과 Results
Genetic structure and population affinity
유전 구조 및 집단 친화도
To elucidate the general genetic relationships between SDH and other East Asians, we conducted an integrative analysis using genome-wide data from both modern and ancient individuals. The Human Origin (HO) and Affymetrix datasets were merged, resulting in a new low-density dataset (Affy_HO). Principal component analysis (PCA) was performed in the East Asian context, and ancient individuals were projected onto the genetic backgrounds of modern populations based on the merged Affy_HO dataset. PCA revealed that SDH lies along the axis between the northern and southern Chinese populations and is located in the northern genetic cline, partly overlapping with southern Han and Mongolians (Fig. 1a). Notably, Tungusic-speaking populations simultaneously tended to shift toward SDH, indicating their potential genetic interactions. Moreover, ancient individuals from the YRB, including Dacaozi, Jiaozuoniecun, and Pingliangtai, displayed apparent genetic affinities with SDH. Some SDH individuals deviated toward southern Han Chinese and were clustered closer to modern Japanese and Korean populations than were ancient people, reflecting additional gene flow events.
산동 한족(山東漢族)과 다른 동아시아인들 간의 일반적인 유전 관계를 밝히기 위해, 우리는 현대인과 고대인의 게놈 전체 데이터를 사용하여 통합 분석을 수행했다. 인류 기원(Human Origin, HO) 데이터세트와 애피메트릭스(Affymetrix) 데이터세트를 병합하여 새로운 저밀도 데이터세트(Affy_HO)를 생성했다. 주성분 분석(PCA)은 동아시아 맥락에서 수행되었으며, 병합된 Affy_HO 데이터세트를 기반으로 고대인 개체들을 현대 인구의 유전적 배경에 투영했다. 주성분 분석 결과, 산동 한족(山東漢族)은 중국 북부와 남부 인구 사이의 축을 따라 위치하며 북부 유전적 연속선상에 자리 잡고 있고, 일부는 남부 한족(漢族) 및 몽골인과 겹치는 것으로 나타났다(그림 1a). 특히, 퉁구스어족(Tungusic-speaking) 인구는 동시에 산동 한족(山東漢族) 쪽으로 이동하는 경향을 보여, 잠재적인 유전적 상호작용이 있었음을 시사했다. 더욱이, 다차오쯔(Dacaozi), 자오쭤녜춘(Jiaozuoniecun), 평량대(平糧臺)를 포함한 황하 유역(黃河流域)의 고대인들은 산동 한족(山東漢族)과 뚜렷한 유전적 친화성을 보였다. 일부 산동 한족(山東漢族) 개인들은 남부 한족(漢族) 쪽으로 벗어났으며, 고대인들보다 현대 일본인 및 한국인 집단과 더 가깝게 군집을 이루어 추가적인 유전자 흐름 사건이 있었음을 반영했다.
To better understand the genetic structure of SDH, we further investigated the genetic affinity and differentiation among geographically distinct Han Chinese and minority ethnic groups. Our findings revealed genetic differences among geographically different Han populations along latitudinal and longitudinal gradients. Based on the pairwise fixation index (FST) estimation, we observed that SDH had the closest genetic relationships with the northern Han population, followed by the surrounding minority ethnic groups (Fig. 1b). However, interestingly, SDH exhibited a more distant genetic relationship with the Guangxi Han (GXH) population than with the northern Altaic-speaking populations, such as the Mongolian, Ewenki, and Daur populations. Additionally, we found that the genetic differences between the SDH and Xinjiang Han populations (XJH, FST=0.0001) were not as apparent as those between the SDH and GXH populations (FST=0.005). GXH represents the southern Han population, and XJH lives in the westernmost region of China. The FST results indicated that genetic differentiation among Han populations was more prominent in the North-South region than in the East-West region (Supplementary Table 1).
산동 한족(山東漢族)의 유전 구조를 더 잘 이해하기 위해, 지리적으로 다른 한족(漢族)과 소수 민족 집단 간의 유전적 친화성과 분화를 추가로 조사했다. 우리의 연구 결과는 위도 및 경도 기울기를 따라 지리적으로 다른 한족(漢族) 집단들 사이에 유전적 차이가 있음을 밝혔다. 쌍별 고정 지수(FST) 추정에 따르면, 산동 한족(山東漢族)은 북부 한족(漢族) 집단과 가장 가까운 유전적 관계를 보였고, 그 다음으로 주변 소수 민족 집단 순이었다 (그림 1b). 하지만 흥미롭게도, 산동 한족(山東漢族)은 몽골인, 에벤키인, 다우르인과 같은 북부 알타이어족(Altaic-speaking) 인구보다 광서 한족(廣西漢族)과 더 먼 유전적 관계를 보였다. 또한, 산동 한족(山東漢族)과 신강 한족(新疆漢族) 사이의 유전적 차이(FST=0.0001)는 산동 한족(山東漢族)과 광서 한족(廣西漢族) 사이의 차이(FST=0.005)만큼 뚜렷하지 않음을 발견했다. 광서 한족(廣西漢族)은 남부 한족(漢族)을 대표하고, 신강 한족(新疆漢族)은 중국의 가장 서쪽 지역에 거주한다. FST 결과는 한족(漢族) 집단 간의 유전적 분화가 동서 지역보다 남북 지역에서 더 두드러졌음을 나타냈다 (보충 표 1).
Hierarchical clustering patterns inferred from TreeMix and identity by descent (IBD)-based heatmap analyses revealed that Han people could be divided into subgroups based on their geographical origin and affinity. There was one case of gene flow from XJH to GXH in the phylogenetic tree. However, no apparent gene flow events related to SDH were detected (Supplementary Fig. 1c). Furthermore, the estimated pairwise IBD fragments showed that SDH had the closest genetic affinity to Shanxi Han, and the shortest shared IBD segments were observed between SDH and southern TK-speaking Hlai people from Qiongzhong in Hainan Province. Compared with other populations, the genetic affinity of the Han populations from northern and northwestern China was greater (Supplementary Fig. 1d). The runs of homozygosity (ROH) of 24 geographically different populations also showed genetic differentiation among them (Supplementary Fig. 1e).
트리믹스(TreeMix)와 동일 조상 유래(IBD) 기반 히트맵 분석에서 추론된 계층적 군집 패턴은 한족(漢族)이 지리적 기원과 친화도에 따라 하위 그룹으로 나뉠 수 있음을 보여주었다. 계통수에서는 신강 한족(新疆漢族)에서 광서 한족(廣西漢族)으로의 유전자 흐름 사례가 한 건 있었다. 하지만, 산동 한족(山東漢族)과 관련된 뚜렷한 유전자 흐름 사건은 발견되지 않았다(보충 그림 1c). 더욱이, 추정된 쌍별 IBD 단편은 산동 한족(山東漢族)이 산서 한족(山西漢族)과 가장 가까운 유전적 친화성을 보였으며, 가장 짧게 공유된 IBD 구간은 산동 한족(山東漢族)과 해남성(海南省) 경중(瓊中) 출신의 남부 타이-카다이어족(TK-speaking) 여족(黎族) 사이에서 관찰되었다. 다른 인구와 비교할 때, 중국 북부와 북서부 한족(漢族) 인구의 유전적 친화성이 더 컸다(보충 그림 1d). 24개 지리적으로 다른 인구의 동형접합성 구간(ROH) 또한 그들 사이의 유전적 분화를 보여주었다(보충 그림 1e).
Moreover, we also conducted haplotype-based fineSTRUCTURE CTU analyses to explore the fine-scale genetic structure among 1061 individuals from various regions across China. We observed four major genetic clusters in the haplotype-based dendrogram and identified a significant North-South gradient of genetic components related to different ancestral sources (K=3, Fig. 1c). To illustrate the significant differences in the genetic components of geographically distinct Han Chinese populations, we conducted interpopulation comparisons focusing on southern ancestral components. Specifically, we selected Hunan Han (HNH) and GXH as representative southern Han populations. Furthermore, we included XJH in the population admixture model to represent the Han population in the westernmost region of China (Supplementary Fig. 1f). Statistical indices further demonstrated significant differences in southern ancestral components among the four Han populations from different geographical locations. These findings supported the coexistence of a North-South structure alongside an East-West cline, which aligned with the admixture model mentioned above (Fig. 1d).
또한, 우리는 중국 전역의 다양한 지역에서 온 1,061명의 개인들 사이의 미세 규모 유전 구조를 탐구하기 위해 일배체형 기반 파인스트럭처(fineSTRUCTURE) 분석을 수행했다. 우리는 일배체형 기반 계통도에서 네 개의 주요 유전 클러스터를 관찰했으며, 다른 조상 출처와 관련된 유전적 구성 요소의 뚜렷한 남북 구배를 확인했다(K=3, 그림 1c). 지리적으로 다른 한족(漢族) 인구의 유전적 구성 요소에서 상당한 차이를 설명하기 위해, 우리는 남부 조상 구성 요소에 초점을 맞춘 집단 간 비교를 수행했다. 구체적으로, 우리는 호남 한족(湖南漢族)과 광서 한족(廣西漢族)을 대표적인 남부 한족(漢族) 인구로 선정했다. 또한, 중국 최서단 지역의 한족(漢族) 인구를 대표하기 위해 인구 혼합 모델에 신강 한족(新疆漢族)을 포함시켰다(보충 그림 1f). 통계 지표는 서로 다른 지리적 위치에 있는 네 한족(漢族) 인구 간의 남부 조상 구성 요소에서 상당한 차이가 있음을 추가로 입증했다. 이러한 발견은 위에서 언급한 혼합 모델과 일치하는 남북 구조와 동서 연속선이 공존함을 뒷받침했다(그림 1d).
The observed patterns of genetic clustering indicated that Han Chinese could be separated into geography-related population stratifications.
관찰된 유전적 군집 패턴은 한족(漢族)이 지리와 관련된 인구 계층으로 나뉠 수 있음을 나타냈다.
그림 1. 지리적으로 다른 현대 및 고대 동아시아 인구와 주변 인구의 유전적 친화성 및 분화. a 병합된 저밀도 게놈 데이터세트를 기반으로 한 22개 현대 및 고대 동아시아인의 주성분 분석. 연구 대상 인구는 빨간색으로 표시되었고, 다른 참조 인구의 클러스터나 연속선은 다른 색상으로 플롯에 표시됨. b 산동 한족(山東漢族) 인구와 23개 민족언어학적으로 다른 그룹 간의 쌍별 FST 유전적 거리. c 파인스트럭처(FineSTRUCTURE)는 애피메트릭스(Affymetrix) 데이터세트를 기반으로 24개 동아시아 인구의 미세 규모 계통발생학적 위상을 보여줌. 애드믹스처(ADMIXTURE) 결과는 동일한 데이터세트를 기반으로 위 인구의 최적 모델이 K=3임을 나타냄. d 애드믹스처(ADMIXTURE) 분석에서 추론된 네 한족(漢族) 인구의 남부 조상 구성 요소 비율에 대한 집단 간 비교는 보충 그림 1f에 나와 있음.
Ancestral makeup and admixture landscape of Han Chinese from the lower YRB
황하(黃河) 하류 유역 한족(漢族)의 조상 구성 및 혼합 양상
We used an unsupervised model-based ADMIXTURE analysis of SDH combined with modern and ancient East Asians to dissect the ancestry makeup of SDH and explore the genetic influence of the surrounding populations (Fig. 2a). The best-fit admixture model with six ancestral sources revealed that compared with other Han people, the SDH population harbored a greater proportion of YRB-related ancestral components (dark blue). Ancient Siberian-related components decreased from ancient Shandong people (e.g., Bianbian, Xiaogao, and Boshan) in the Neolithic Age to modern SDH, while ancient YRB-related ancestry components increased. Furthermore, our findings indicated that southern East Asian ancestry was greatest in TK-speaking populations (dark green), HM-speaking populations (red), ancient Tibeto-Burman (TB) populations from the southern Tibetan Plateau, and ancient Hanben populations. Siberian-dominant ancestry was essential for the formation of the gene pools of Altaic people, Japanese people, Koreans, and northern Chinese Sinitic speakers, which gradually became rare or absent in southern East Asian populations, such as HM and TK speakers.
우리는 산동 한족(山東漢族)의 조상 구성을 분석하고 주변 인구의 유전적 영향을 탐구하기 위해, 산동 한족(山東漢族)을 현대 및 고대 동아시아인과 결합한 비지도 모델 기반 애드믹스처(ADMIXTURE) 분석을 사용했다(그림 2a). 6개의 조상 공급원을 가진 최적의 혼합 모델은 다른 한족(漢族)에 비해 산동 한족(山東漢族) 인구가 더 많은 비율의 황하 유역(黃河流域) 관련 조상 구성 요소(진한 파란색)를 가지고 있음을 밝혔다. 고대 시베리아 관련 구성 요소는 신석기 시대의 고대 산동 사람들(예: 변변(卞邊), 소고(蕭高), 보산(博山))에서 현대 산동 한족(山東漢族)으로 오면서 감소한 반면, 고대 황하 유역(黃河流域) 관련 조상 구성 요소는 증가했다. 더욱이, 우리의 연구 결과는 남부 동아시아 혈통이 타이-카다이어족(TK-speaking) 인구(진녹색), 몽-미엔어족(HM-speaking) 인구(빨간색), 티베트 고원 남부의 고대 티베트-버마어족(TB) 인구, 그리고 고대 한본(漢本) 인구에서 가장 높게 나타났음을 시사했다. 시베리아 우세 혈통은 알타이어족(Altaic), 일본인, 한국인, 그리고 북부 중국의 중국어 사용자들의 유전자 풀 형성에 필수적이었으나, 몽-미엔어족(HM) 및 타이-카다이어족(TK) 화자와 같은 남부 동아시아 인구에서는 점차 드물어지거나 사라졌다.
We used various ancestral source pairs to investigate the genetic contribution to SDH via admixture-f₃(Source1, Source2; SDH), in which significant negative Z scores (Z<-3) indicated prominent admixture events. As shown in Supplementary Fig. 2a, we observed that the combination of ANEA associated with late Neolithic Longshan millet farmers (China_Upper_YR_LN) and southern ancestry linked to the Iron Age Hanben people (Taiwan_Hanben_IA) resulted in a negative f₃ value (Z=-3.65, Supplementary Table 2). The ANEA and ASEA pairs could serve as possible ancestral sources of SDH, consistent with the mosaic ancestry composition observed in the ADMIXTURE-based results. We validated the North-South admixture model using modern populations (Supplementary Fig. 2b and Supplementary Table 3). In this model, the northern source comprised Altaic-speaking populations (Mongolian and Daur) and Sinitic speakers, while the other source encompassed southern East Asians associated with HM/TK-speaking populations, as estimated by f₄(Altaic, southern East Asians; SDH).
우리는 다양한 조상 공급원 쌍을 사용하여 혼합-f₃(출처1, 출처2; 산동 한족(山東漢族)) 통계량을 통해 산동 한족(山東漢族)에 대한 유전적 기여를 조사했으며, 여기서 유의미한 음의 Z 점수(Z<-3)는 뚜렷한 혼합 사건을 나타냈다. 보충 그림 2a에서 보듯이, 신석기 후기 용산(龍山) 문화 기장 농부들과 관련된 고대 북부 동아시아인(ANEA)과 철기 시대 한본(漢本) 사람들과 연결된 남부 혈통의 조합은 음의 f₃값(Z=-3.65, 보충 표 2)을 나타냈다. 고대 북부 동아시아인(ANEA)과 고대 남부 동아시아인(ASEA) 쌍은 산동 한족(山東漢族)의 가능한 조상 공급원 역할을 할 수 있으며, 이는 애드믹스처(ADMIXTURE) 기반 결과에서 관찰된 모자이크형 조상 구성과 일치한다. 우리는 현대 인구를 사용하여 남북 혼합 모델을 검증했다(보충 그림 2b 및 보충 표 3). 이 모델에서 북부 출처는 알타이어족(Altaic-speaking) 인구(몽골인 및 다우르인)와 중국어 화자로 구성되었고, 다른 출처는 f₄(알타이어족, 남부 동아시아인; 산동 한족(山東漢族))에 의해 추정된 바와 같이 몽-미엔/타이-카다이어족(HM/TK-speaking) 인구와 관련된 남부 동아시아인을 포함했다.
Additionally, we investigated the genomic affinity between the target population and other ancient Asians via affinity f₄-statistic of f₄(Ancient Reference1, Ancient Reference2; SDH, Mbuti). Our results revealed that SDH shared more alleles with ANEA than with ASEA. Compared with ancient individuals in Shandong (China_NEastAsia_Coastal_EN), a stronger genetic affinity was observed between modern SDH and four ancient populations in the upper and middle YRB (China_YR_LBIA, China_YR_LN, Shimao, and Miaozigou), suggesting the apparent genetic influence of ancient Yangshao and Longshan people on modern Shandong people (Fig. 2b; Supplementary Table 4). Therefore, based on the affinity f 4-statistics, it can be inferred that the gene pool of SDH has been more affected by YRB-related farmers and ancient individuals in Inner Mongolia since the early Neolithic period.
추가로, 우리는 친화도 f₄ 통계량인 f₄(고대 참조1, 고대 참조2; 산동 한족(山東漢族), 음부티족(Mbuti))을 통해 대상 인구와 다른 고대 아시아인 간의 게놈 친화도를 조사했다. 우리의 결과는 산동 한족(山東漢族)이 고대 남부 동아시아인(ASEA)보다 고대 북부 동아시아인(ANEA)과 더 많은 대립유전자를 공유한다는 것을 밝혔다. 산동의 고대인과 비교했을 때, 현대 산동 한족(山東漢族)과 황하(黃河) 상·중류 유역의 네 고대 인구(중국_황하_후기청동기/철기, 중국_황하_후기신석기, 석모(石峁), 묘자구(廟子溝)) 사이에서 더 강한 유전적 친화성이 관찰되었다. 이는 고대 양소(仰韶) 및 용산(龍山) 사람들이 현대 산동 사람들에게 뚜렷한 유전적 영향을 미쳤음을 시사한다(그림 2b; 보충 표 4). 따라서 친화도 f₄ 통계량에 근거하여, 산동 한족(山東漢族)의 유전자 풀은 신석기 시대 초기부터 황하 유역(黃河流域) 관련 농부들과 내몽골(內蒙古)의 고대인들로부터 더 많은 영향을 받았다고 추론할 수 있다.
To further investigate whether SDH descended directly from ANEA, we computed f₄(ANEA, SDH; Reference, Mbuti). Only a few significant signals (indicated by a red color with a Z score > 3) were observed when we assumed that Malaysia_LN and Laos_LN_BA were the reference ancestral sources (Supplementary Fig. 3a and Supplementary Table 5). Thus, compared to China_NEastAsia_Coastal_EN, SDH obtained additional gene flows from late Neolithic Malaysians and Laotians during the late Neolithic to Bronze Age. However, we should also pay attention to these two weak signals, which might be caused by the low overlap of SNPs and ancient DNA damage signals. Furthermore, significant negative f₄ values were observed for Taiwan_Hanben_IA and GaoHuaHua when we assumed that China_YR_MN was the ancestral contributor in the form of f₄(China_YR_MN, SDH; Reference, Mbuti) (Supplementary Fig. 3b). As shown above, most of the results exhibited non-significant f₄-values (∣Z∣≤3), and several significant f₄ values indicated that southern ancestral populations also contributed genetic materials to SDH.
산동 한족(山東漢族)이 고대 북부 동아시아인(ANEA)으로부터 직접 유래했는지를 더 조사하기 위해, 우리는 f₄(고대 북부 동아시아인, 산동 한족(山東漢族); 참조, 음부티족(Mbuti))을 계산했다. 말레이시아(Malaysia) 후기 신석기인과 라오스(Laos) 후기 신석기/청동기인을 참조 조상 공급원으로 가정했을 때, 단지 몇 개의 유의미한 신호(Z 점수 > 3인 빨간색으로 표시)만이 관찰되었다(보충 그림 3a 및 보충 표 5). 따라서, 중국 동북아시아 해안 초기 신석기인과 비교하여, 산동 한족(山東漢族)은 후기 신석기에서 청동기 시대 동안 후기 신석기 말레이시아인과 라오스인으로부터 추가적인 유전자 흐름을 받았다. 하지만, 우리는 이 두 약한 신호에도 주목해야 하며, 이는 SNP의 낮은 중첩과 고대 DNA 손상 신호 때문일 수 있다. 더욱이, 중국 황하(黃河) 중기 신석기인이 f₄(중국_황하_중기신석기, 산동 한족(山東漢族); 참조, 음부티족(Mbuti)) 형태의 조상 기여자로 가정했을 때, 대만 한본(台灣漢本) 철기시대인과 고화화(高花花)에 대해 유의미한 음의 f₄ 값이 관찰되었다(보충 그림 3b). 위에서 보듯이, 대부분의 결과는 유의미하지 않은 f₄ 값(∣Z∣≤3)을 보였고, 몇몇 유의미한 f₄ 값은 남부 조상 인구 또한 산동 한족(山東漢族)에 유전 물질을 기여했음을 나타냈다.
We then used asymmetric f₄(ASEA, SDH; Reference, Mbuti) to explore additional possible ancestral sources for the formation of SDH. When we hypothesized that Hanben was their possible ancestor, SDH obtained more gene flow from ancient reference populations in the Mongolian Plateau, Siberia, Nepal, and WLR/ARB/YRB than did ASEA, suggesting that ANEA contributed significantly to the modern SDH’s gene pool (Supplementary Fig. 3c and Supplementary Table 6). The statistically significant negative values observed in f₄(ANEA, SDH; ASEA, Mbuti) were also consistent with the f₃-based admixture models, in which ASEA shared more alleles with SDH than non-YRB ANEA surrogates. Moreover, the extent of genetic heterogeneity between the SDH population and the southern Han populations is still worth investigating. We used f₄ statistics in the form of f₄(Southern Han, SDH; Reference, Mbuti), aiming to test their genetic differences and which factors contributed to differentiation (Supplementary Table 7). SDH did not form one clade with southern Han Chinese individuals, as indicated by the statistically significant negative and positive f₄ values. These statistically significant values further indicated that the northern Han and southern Han peoples experienced different genetic influences from their geographically close indigenous neighbors in terms of their past demographic processes, including early divergence and subsequent extensive admixture or gene flow events with other sources (Supplementary Fig. 4a-c).
그 후 우리는 비대칭 f₄(고대 남부 동아시아인, 산동 한족(山東漢族); 참조, 음부티족(Mbuti))을 사용하여 산동 한족(山東漢族) 형성에 대한 추가적인 가능한 조상 출처를 탐색했다. 한본(漢本)이 그들의 가능한 조상이라고 가정했을 때, 산동 한족(山東漢族)은 고대 남부 동아시아인(ASEA)보다 몽골 고원, 시베리아, 네팔, 그리고 서요하/아무르강/황하 유역의 고대 참조 인구로부터 더 많은 유전자 흐름을 받았다. 이는 고대 북부 동아시아인(ANEA)이 현대 산동 한족(山東漢族)의 유전자 풀에 상당히 기여했음을 시사한다(보충 그림 3c 및 보충 표 6). f₄(고대 북부 동아시아인, 산동 한족(山東漢族); 고대 남부 동아시아인, 음부티족(Mbuti))에서 관찰된 통계적으로 유의미한 음의 값들은 f₃ 기반 혼합 모델과도 일치했으며, 이 모델에서 고대 남부 동아시아인(ASEA)은 비황하유역 고대 북부 동아시아인(non-YRB ANEA) 대리 집단보다 산동 한족(山東漢族)과 더 많은 대립유전자를 공유했다. 더욱이, 산동 한족(山東漢族) 인구와 남부 한족(漢族) 인구 간의 유전적 이질성 정도는 여전히 조사할 가치가 있다. 우리는 그들의 유전적 차이와 분화에 기여한 요인을 시험하기 위해 f₄(남부 한족(漢族), 산동 한족(山東漢族); 참조, 음부티족(Mbuti)) 형태의 f₄ 통계량을 사용했다(보충 표 7). 통계적으로 유의미한 음과 양의 f₄ 값에서 나타나듯이, 산동 한족(山東漢族)은 남부 한족(漢族) 개인들과 하나의 분기군을 형성하지 않았다. 이러한 통계적으로 유의미한 값들은 북부 한족(漢族)과 남부 한족(漢族)이 초기 분기 및 이후 다른 출처와의 광범위한 혼합 또는 유전자 흐름 사건을 포함한 과거 인구학적 과정 측면에서 지리적으로 가까운 토착 이웃들로부터 서로 다른 유전적 영향을 경험했음을 추가로 시사했다(보충 그림 4a-c).
More alleles related to Mongolic, Tungusic, TB speakers, and Western Eurasians were detected in the SDH than in the southern Han populations from Hunan and Chongqing Provinces. We further tested the admixture models and quantified the proportions of admixtures with different ancestral sources at different scales using qpWave and qpAdm analyses. The well-fit two-way admixture model suggested that SDH can be simulated as a North-South admixture model. In this model, YRB-related ancestries, including Miaozigou_MN, Upper_YR_LN, and YR_LN, represented the ANEA source. The ASEA component comprised SEastAsia_Coastal_LN, BaBanQinCen, and Taiwan_Hanben_IA. The average estimated genetic contributions of ANEA and ASEA were approximately 0.85 and 0.15, respectively (standard error of the mean = 0.11, Fig. 2c and Supplementary Table 8). We observed that the greatest proportion (0.946±0.026) of ANEA contributed to the gene pool of SDH when the geographically close Longshan people served as northern ancestral sources and IronAge Hanben served as southern ancestral sources, consistent with the clustering patterns of the relative genetic stability observed in ADMIXTURE and PCA.
산동 한족(山東漢族)에서는 호남(湖南) 및 중경(重慶)성의 남부 한족(漢族) 인구보다 몽골어족, 퉁구스어족, 티베트-버마어족 화자 및 서유라시아인과 관련된 대립유전자가 더 많이 발견되었다. 우리는 혼합 모델을 추가로 시험하고, qpWave 및 qpAdm 분석을 사용하여 다양한 규모에서 다른 조상 출처와의 혼합 비율을 정량화했다. 잘 맞는 양방향 혼합 모델은 산동 한족(山東漢族)이 남북 혼합 모델로 모의실험될 수 있음을 시사했다. 이 모델에서 묘자구(廟子溝), 상류 황하 후기 신석기, 황하 후기 신석기를 포함한 황하 유역(黃河流域) 관련 혈통이 고대 북부 동아시아인(ANEA) 출처를 대표했다. 고대 남부 동아시아인(ASEA) 구성 요소는 동남아시아 해안 후기 신석기, 파판친천(BaBanQinCen), 그리고 대만 한본(台灣漢本) 철기시대로 구성되었다. 고대 북부 동아시아인과 고대 남부 동아시아인의 평균 추정 유전적 기여도는 각각 약 0.85와 0.15였다 (평균의 표준오차 = 0.11, 그림 2c 및 보충 표 8). 지리적으로 가까운 용산(龍山) 사람들이 북부 조상 공급원 역할을 하고 철기 시대 한본(漢本)이 남부 조상 공급원 역할을 했을 때, 가장 큰 비율(0.946±0.026)의 고대 북부 동아시아인이 산동 한족(山東漢族)의 유전자 풀에 기여했음을 관찰했다. 이는 애드믹스처(ADMIXTURE)와 주성분 분석(PCA)에서 관찰된 상대적 유전적 안정성의 군집 패턴과 일치한다.
Based on the admixture-induced linkage disequilibrium for evolutionary relationships (ALDER), we further estimated the admixture time of SDH based on the LD decay pattern. The results revealed that the northern Han exhibited genetic contact with Altaic-speaking Yakut approximately 101 generations ago, aligning with the late Shang Dynasty and the Western Zhou Dynasty (Supplementary Table 9). The estimated ancient genetic connection with Siberians supported persistent cultural and population communication or contact between YRB farmers and ancient Siberians. Early Neolithic Yumin people from the Mongolian Plateau; Neolithic Boshan, Xiaogao, Xiaojingshan, and Bianbian people from Shandong Province all possessed a close genetic connection with ancient Neolithic Siberian lineages [13].
진화 관계를 위한 혼합 유도 연관 불균형(ALDER)을 기반으로, 우리는 연관 불균형(LD) 붕괴 패턴에 근거하여 산동 한족(山東漢族)의 혼합 시기를 추가로 추정했다. 그 결과, 북부 한족(漢族)은 약 101세대 전에 알타이어족(Altaic-speaking) 야쿠트인과 유전적 접촉을 보였으며, 이는 상(商)나라 말기 및 서주(西周) 시대와 일치한다 (보충 표 9). 추정된 고대 시베리아인과의 유전적 연결은 황하 유역(黃河流域) 농부들과 고대 시베리아인들 사이에 지속적인 문화 및 인구 교류 또는 접촉이 있었음을 뒷받침한다. 몽골 고원(蒙古高原)의 신석기 초기 유민(Yumin) 사람들, 산동성(山東省)의 신석기 시대 보산(博山), 소고(蕭高), 소경산(小荊山), 변변(卞邊) 사람들은 모두 고대 신석기 시베리아 혈통과 가까운 유전적 연결을 가지고 있었다 [13].
Relative genetic stability and temporal dynamics in the lower YRB since the Neolithic period
신석기 시대 이래 황하(黃河) 하류 유역의 상대적 유전적 안정성 및 시간적 역학
According to the f₄-statistics mentioned above, SDH received a relatively low degree of genetic influence from non-YRB lineages and maintained a high level of genetic stability. These findings indicated genetic continuity within this region since the Neolithic period. Inferring spatiotemporal patterns of genetic change in populations from the YRB would favor dissecting the formation of northern Han Chinese. We first conducted PCA in the context of ancient East Asians, and modern people were projected onto the first two PCs. We found that northern Han populations in the YRB formed a tight cluster and exhibited genetic similarity with geographically close ANEAs (Supplementary Fig. 5a). The PCA results demonstrated that present-day Han people in the lower YRB overlapped with ancient YRB millet farmers and reflected a considerable degree of genetic affinity between them.
위에서 언급한 f₄ 통계량에 따르면, 산동 한족(山東漢族)은 비황하유역 혈통으로부터 비교적 낮은 수준의 유전적 영향을 받았으며 높은 수준의 유전적 안정성을 유지했다. 이러한 발견은 신석기 시대 이래 이 지역 내에서의 유전적 연속성을 나타냈다. 황하 유역(黃河流域) 인구의 시공간적 유전 변화 패턴을 추론하는 것은 북부 한족(漢族)의 형성을 분석하는 데 도움이 될 것이다. 우리는 먼저 고대 동아시아인의 맥락에서 주성분 분석(PCA)을 수행했으며, 현대인들을 첫 두 주성분에 투영했다. 우리는 황하 유역(黃河流域)의 북부 한족(漢族) 인구가 밀집된 군집을 형성하고 지리적으로 가까운 고대 북부 동아시아인(ANEA)과 유전적 유사성을 보임을 발견했다(보충 그림 5a). 주성분 분석 결과는 오늘날 황하(黃河) 하류 유역의 한족(漢族)이 고대 황하 유역(黃河流域) 기장 농부들과 겹치며, 그들 사이에 상당한 수준의 유전적 친화성이 있음을 반영했다.
We utilized a series of f₄-statistics to investigate temporal patterns of genetic continuity among the YRB-related populations, spanning from the early Neolithic era to modern times (Supplementary Table 10). Our findings revealed that early Neolithic coastal inhabitants in the lower YRB (China_NEastAsia_Coastal_EN) were more substantially affected by ancient people from Siberia and the Mongolian Plateau than were the Middle Neolithic Yangshao people (YR_MN, Supplementary Fig. 6a). Gene flow between coastal and inland Neolithic ANEAs was inferred from positive values in f₄(China_NEastAsia_Coastal_EN, China_YR_MN; Reference, Mbuti). During the late Neolithic period, we found that the component related to the ASEA increased based on f₄(China_YR_MN, China_YR_LN; Reference, Mbuti). Late Neolithic Longshan individuals (China_YR_LN) gradually shared more alleles with the ASEA than with the YR_MN, consistent with the northward expansion of rice farmers (Supplementary Fig. 6b). However, there were no further significant genetic shifts from YR_LN to the late Bronze/Iron Age (YR_LBIA), as evidenced by f₄(China_YR_LN, China_YR_LBIA; Reference, Mbuti) (Supplementary Fig. 6c). We then carried out f₄(YR_LBIA, SDH; Reference, Mbuti) to assess the genetic homogeneity between the late Bronze/Iron Age individuals and present-day SDH individuals (Supplementary Fig. 7a). Except for the early and middle Neolithic YRB-related populations, the f₄ values were not significant for any of the tested individuals (∣Z∣≤3). Many nonsignificant f₄ values supported that SDH harbored a strong genomic affinity for YRB-related late Bronze/Iron Age individuals and was less influenced by additional genetic material. Our results revealed clear indications of long-term genetic continuity in the lower reaches of the YRB since the early Neolithic, with modern SDH still displaying high genetic homogeneity with YRB-related ancestors.
우리는 신석기 초기부터 현대에 이르기까지 황하 유역(黃河流域) 관련 인구들 사이의 시간적 유전적 연속성 패턴을 조사하기 위해 일련의 f₄ 통계량을 활용했다 (보충 표 10). 우리의 연구 결과, 황하(黃河) 하류 유역의 신석기 초기 해안 거주민들은 중기 신석기 양소(仰韶) 사람들보다 시베리아와 몽골 고원(蒙古高原)의 고대인들로부터 더 실질적인 영향을 받았음이 드러났다 (보충 그림 6a). 해안과 내륙의 신석기 시대 고대 북부 동아시아인(ANEA)들 사이의 유전자 흐름은 f₄(China_NEastAsia_Coastal_EN, China_YR_MN; Reference, Mbuti) 통계량의 양수 값으로부터 추론되었다. 신석기 후기 동안, 우리는 f₄(China_YR_MN, China_YR_LN; Reference, Mbuti) 통계량에 근거하여 고대 남부 동아시아인(ASEA) 관련 구성 요소가 증가했음을 발견했다. 후기 신석기 용산(龍山) 사람들은 점차적으로 중기 신석기 황하 유역 사람들보다 고대 남부 동아시아인(ASEA)과 더 많은 대립유전자를 공유했으며, 이는 벼농사 농부들의 북상과 일치한다 (보충 그림 6b). 하지만, f₄(China_YR_LN, China_YR_LBIA; Reference, Mbuti) 통계량에서 입증된 바와 같이, 후기 신석기 황하 유역 사람들로부터 후기 청동기/철기 시대에 이르기까지 더 이상의 유의미한 유전적 변화는 없었다 (보충 그림 6c). 그 후 우리는 후기 청동기/철기 시대 개인들과 오늘날 산동 한족(山東漢族) 개인들 사이의 유전적 동질성을 평가하기 위해 f₄(YR_LBIA, SDH; Reference, Mbuti) 분석을 수행했다 (보충 그림 7a). 초기 및 중기 신석기 황하 유역 관련 인구를 제외하고, 시험된 어떤 개인에 대해서도 f₄ 값은 유의미하지 않았다 (∣Z∣≤3). 다수의 유의미하지 않은 f₄ 값은 산동 한족(山東漢族)이 황하 유역 관련 후기 청동기/철기 시대 개인들과 강한 게놈 친화성을 가지며 추가적인 유전 물질의 영향을 덜 받았음을 뒷받침했다. 우리의 결과는 신석기 초기 이래 황하(黃河) 하류에서의 장기적인 유전적 연속성을 명확히 보여주며, 현대 산동 한족(山東漢族)은 여전히 황하 유역 관련 조상들과 높은 유전적 동질성을 나타낸다.
Similarly, the pairwise qpWave analysis focused on YR_LBIA and SDH was consistent with the findings of the f₄ statistics and further confirmed their genetic homogeneity (Fig. 2d).
유사하게, 후기 청동기/철기 시대 황하 유역 인구(YR_LBIA)와 산동 한족(山東漢族, SDH)에 초점을 맞춘 쌍별 qpWave 분석은 f₄ 통계량의 발견과 일치했으며, 그들의 유전적 동질성을 더욱 확인시켜 주었다 (그림 2d).
그림 2. 산동 한족(山東漢族) 인구의 혼합 지형과 인구 역사. a 산동 한족(山東漢族) 인구와 201개 참조 인구에 대한 K=6에서의 애드믹스처(ADMIXTURE) 결과. 클러스터링 패턴에서 다른 색상은 다른 조상 출처를 나타냄. b 산동 한족(山東漢族)과 고대 참조 집단 간의 유전적 친화성을 시험하기 위해 f₄(고대 참조1, 고대 참조2; 산동 한족(山東漢族), 음부티족(Mbuti)) 형태의 f₄-통계량을 계산함. 고대 참조1과 고대 참조2는 동아시아에서 선택된 76개의 고대인을 대표함. |Z| > 3은 참(TRUE)으로 표시되며, 산동 한족(山東漢族)과 고대 참조2 간의 더 가까운 유전적 친화성을 나타냄. c qpAdm으로 추정된 산동 한족(山東漢族) 내 다른 고대 북부 동아시아인(ANEA) 및 고대 남부 동아시아인(ASEA)의 혼합 비율. 우리는 고대 북부 동아시아인과 고대 남부 동아시아인의 유전적 기여를 결정하기 위해 양방향 혼합 모델을 사용함. 고대 북부 동아시아인은 황하 유역(黃河流域) 관련 농부와 서요하(西遼河)의 고대인을 포함함. 고대 남부 동아시아인은 주로 광서(廣西)와 복건(福建)의 고대인으로 구성됨. d 쌍별 qpWave 결과는 병합된 Affy_1240K 데이터세트를 기반으로 산동 한족(山東漢族)과 동아시아의 고대/현대 인간 간의 유전적 이형접합성을 보여줌.
Moreover, the values of qpWave validated that YR_LBIA was also genetically close to other Han populations in the lower reaches of the YRB (e.g., Shanxi, Henan, and Shaanxi). However, genetic heterogeneity existed among the four Han populations in the YRB. The f₄(YR_LBIA, Han Shanxi/Henan/Shaanxi; Reference, Mbuti) explained in detail which gene flows caused the differences (Supplementary Fig. 7b-d). On the one hand, Han Chinese individuals from Shaanxi, Henan, and Shanxi Provinces shared more alleles with non-YRB ancestral sources than did those from SDH, represented by DevilsGate hunter-gatherers, Mongolic speakers, and WLR-related populations. On the other hand, gene flows from the ASEA and ancient East/Southeast Asians, such as China SEastAsia_Coastal_LN, Gaohuahua, and Vietnam_BA, all contributed to the increasing genetic diversity of Han Chinese from Shanxi and Shaanxi provinces. We then evaluated the demographic processes of the four Han populations by effective population size (Ne, Supplementary Fig. 5b). SDH steadily increased before 25 generations, and Shanxi Han experienced a bottleneck effect at approximately 10 generations. However, the height of the Han populations from Henan and Shaanxi increased faster than that of the former. Our analysis suggested that migration from southern populations has exerted less influence on SDH since the Bronze Age, which may serve as one possible interpretation of the genetic stability observed in SDH.
더욱이, qpWave 값은 후기 청동기/철기 시대 황하 유역 인구(YR_LBIA)가 황하(黃河) 하류의 다른 한족(漢族) 인구(예: 산서(山西), 하남(河南), 섬서(陝西))와도 유전적으로 가깝다는 것을 검증했다. 하지만, 황하 유역(黃河流域)의 네 한족(漢族) 인구 사이에는 유전적 이질성이 존재했다. f₄(YR_LBIA, Han Shanxi/Henan/Shaanxi; Reference, Mbuti) 통계량은 어떤 유전자 흐름이 차이를 유발했는지 상세히 설명했다 (보충 그림 7b-d). 한편으로, 섬서(陝西), 하남(河南), 산서(山西)성의 한족(漢族) 개인들은 산동 한족(山東漢族)보다 데빌스게이트(DevilsGate) 수렵채집인, 몽골어족 화자, 서요하(西遼河) 관련 인구로 대표되는 비황하유역 조상 공급원과 더 많은 대립유전자를 공유했다. 다른 한편으로, 동남아시아 해안 후기 신석기인, 고화화(高花花), 베트남 청동기 시대인 등 고대 남부 동아시아인(ASEA)과 고대 동/동남아시아인으로부터의 유전자 흐름은 모두 산서(山西)와 섬서(陝西)성 한족(漢族)의 유전적 다양성 증가에 기여했다. 그 후 우리는 유효 집단 크기(Ne)를 통해 네 한족(漢族) 인구의 인구학적 과정을 평가했다 (보충 그림 5b). 산동 한족(山東漢族)은 25세대 전까지 꾸준히 증가했고, 산서 한족(山西漢族)은 약 10세대 전에 병목 현상을 경험했다. 하지만 하남(河南)과 섬서(陝西) 출신 한족(漢族) 인구의 크기는 전자보다 더 빠르게 증가했다. 우리의 분석은 청동기 시대 이래로 남부 인구로부터의 이주가 산동 한족(山東漢族)에 미친 영향이 적었음을 시사하며, 이는 산동 한족(山東漢族)에서 관찰된 유전적 안정성에 대한 가능한 해석 중 하나가 될 수 있다.
Natural selection and adaptation mechanisms
자연선택과 적응 메커니즘
Characterizing the biological adaptability of genetically distinct populations is essential for understanding the evolutionary driving forces of complex genetic traits or diseases. We first used population branch statistics (PBS) to explore significant natural selection signatures in the SDH to dissect population-specific variants. We used SDH as the target population and selected Hlai_Qiongzhong (HNL) and Europeans (CEU) as the ingroup and outgroup reference populations, respectively. In the SDH-HNL-CEU trio model, we detected 340 SNPs in the top 0.1% (PBS>0.185) of the mean genome-wide PBS scores (Fig. 3a and Supplementary Table 11). As a supplement to PBS, we also applied the integrated haplotype score (iHS) and FST to further validate these biological adaptive signals in the SDH (Supplementary Tables 12-13). Combined with haplotype and allele frequency methods, we identified 159 high-quality variants, 19.5% of which were missense variants (Fig. 3f). The SDH-specific selection signals identified by iHS were primarily associated with height (PDHX, pyruvate dehydrogenase component X), atopic asthma (TAP2, Transporter 2), and BMI-adjusted waist-to-hip ratio (HLA-B, major histocompatibility complex; HLA-C, L3MBTL3, L3MBTL histone methyl-lysine binding protein 3 and C6orf10, testis expressed basic protein 1).
유전적으로 구별되는 인구 집단의 생물학적 적응성을 규명하는 것은 복잡한 유전 형질이나 질병의 진화적 원동력을 이해하는 데 필수적이다. 우리는 먼저 집단 분기 통계량(PBS)을 사용하여 산동 한족(山東漢族)의 중요한 자연선택 신호를 탐색하고 집단 특이적 변이를 분석했다. 우리는 산동 한족(山東漢族)을 대상 집단으로, 경중(瓊中) 여족(黎族)과 유럽인(CEU)을 각각 내집단과 외집단 참조 집단으로 선택했다. 이 세 집단 모델에서, 우리는 전체 게놈 PBS 점수 평균의 상위 0.1%(PBS>0.185)에 해당하는 340개의 단일염기다형성(SNP)을 발견했다 (그림 3a 및 보충 표 11). PBS를 보완하기 위해, 통합 일배체형 점수(iHS)와 FST를 적용하여 산동 한족(山東漢族)의 생물학적 적응 신호를 추가로 검증했다 (보충 표 12-13). 일배체형 및 대립유전자 빈도 방법을 결합하여 159개의 고품질 변이를 식별했으며, 이 중 19.5%는 미스센스 변이(missense variant)였다 (그림 3f). iHS로 확인된 산동 한족(山東漢族) 특이적 선택 신호는 주로 키(PDHX), 아토피성 천식(TAP2), 그리고 체질량지수(BMI)로 보정된 허리-엉덩이 비율(HLA-B, HLA-C, L3MBTL3, C6orf10)과 관련이 있었다.
Among the 340 candidate SNPs, the strongest PBS value was observed for ATP binding cassette subfamily C member 11 (ABCC11), which is located on chromosome 16 (Fig. 3c). The ABCC11 variant is associated with human axillary odor (AO) and earwax type [39]. The ancestral allele (rs17822931-G) was demonstrated to increase the risk of axillary osmidrosis, while the derived allele (rs17822931-A) displayed a positive selection signal in East Asians, suggesting that it may confer an adaptive advantage in cold climates [40]. This polymorphism has undergone a complex evolutionary history, resulting in diverse allele frequencies across spatiotemporally different ancient and modern populations. Previously identified natural selection signals, including Ectodysplasin A receptor (EDAR) [41], Solute carrier family 35 member F₃ (SLC35F₃) [42], and acetaldehyde dehydrogenase 2 (ALDH2) [43], have been implicated in phenotypic traits such as shovel-shaped incisors, thiamine metabolism, and alcohol metabolism.
340개의 후보 단일염기다형성(SNP) 중에서 가장 강력한 PBS 값은 16번 염색체에 위치한 ABCC11 유전자에서 관찰되었다 (그림 3c). ABCC11 변이는 인간의 겨드랑이 냄새(액취)와 귀지 유형과 관련이 있다 [39]. 조상 대립유전자(rs17822931-G)는 액취증의 위험을 증가시키는 것으로 나타났고, 파생 대립유전자(rs17822931-A)는 동아시아인에게서 양성 선택 신호를 보였다. 이는 추운 기후에서 적응적 이점을 제공할 수 있음을 시사한다 [40]. 이 다형성은 복잡한 진화 역사를 거쳐 시공간적으로 다른 고대 및 현대 인구에서 다양한 대립유전자 빈도를 나타냈다. 이전에 확인된 자연선택 신호들, 예를 들어 EDAR [41], SLC35F₃ [42], ALDH2 [43] 등은 삽 모양 앞니, 티아민 대사, 알코올 대사와 같은 표현형 특성과 관련이 있다.
We subsequently annotated the 159 robust SNPs via the Variant Effect Predictor (VEP) and Genome-Wide Association Study (GWAS) catalogs. A Sankey diagram indicated that candidate genes for selection and their associated phenotypes exhibited potential pleiotropy (Fig. 3b). For instance, rs723527, which is associated with epidermal growth factor receptor (EGFR), influences central nervous system cancer and glioblastoma multiforme and participates in white matter microstructure measurements. ARHGAP42 (rs590616, Rho GTPase Activating Protein 42) has been identified as a risk factor for type 2 diabetes mellitus. Additionally, two variants within the glutamate ionotropic receptor AMPA type subunit 1 (GRIA1) gene, rs13168358 and rs1461225, were associated with insomnia. GRIA1 encodes an excitatory receptor for neurotransmitters and has been linked to deficits in short-term habituation, as demonstrated in animal models [44]. This gene was also subject to strong selection in the studied population, second only to ABCC11, as evidenced by iHS and PBS scores. The BCL11A (BCL11 transcription factor A) variant rs11886868 is crucial for regulating fetal hemoglobin (HbF) levels and inhibiting deoxy sickle hemoglobin polymerization. Downregulating BCLL11A expression is a promising therapeutic strategy for sickle cell disease [45, 46]. The ancestral allele of this variant is T, and the derived allele is C, of which the CC genotype can promote HbF induction and ameliorate the repression effect on y-globin in patients with sickle cell anemia [47]. GO and KEGG enrichment analyses revealed that the candidate genes were primarily enriched in pathways pertaining to the cell body and cellular response to amino acid stimulus (Fig. 3d).
이후 우리는 변이 효과 예측기(VEP)와 전장 유전체 연관 분석(GWAS) 카탈로그를 통해 159개의 확실한 단일염기다형성(SNP)에 주석을 달았다. 생키 다이어그램(Sankey diagram)은 선택 후보 유전자와 관련 표현형이 잠재적인 다면발현성(pleiotropy)을 보임을 나타냈다 (그림 3b). 예를 들어, EGFR 유전자와 관련된 rs723527은 중추신경계 암과 교모세포종에 영향을 미치며 백색질 미세구조 측정에도 관여한다. ARHGAP42(rs590616)는 제2형 당뇨병의 위험 인자로 확인되었다. 또한 GRIA1 유전자 내의 두 변이, rs13168358과 rs1461225는 불면증과 관련이 있었다. GRIA1은 신경전달물질의 흥분성 수용체를 암호화하며, 동물 모델에서 입증된 바와 같이 단기 습관화 결핍과 관련이 있다 [44]. 이 유전자는 연구 대상 집단에서 ABCC11 다음으로 강한 선택을 받았으며, 이는 iHS와 PBS 점수로 입증되었다. BCL11A 변이 rs11886868은 태아 헤모글로빈(HbF) 수치를 조절하고 겸상 헤모글로빈 중합을 억제하는 데 결정적이다. BCLL11A 발현을 하향 조절하는 것은 겸상적혈구병에 대한 유망한 치료 전략이다 [45, 46]. 이 변이의 조상 대립유전자는 T이고 파생 대립유전자는 C이며, CC 유전자형은 겸상적혈구빈혈증 환자에서 태아 헤모글로빈 유도를 촉진하고 감마-글로빈에 대한 억제 효과를 완화할 수 있다 [47]. 유전자 온톨로지(GO) 및 KEGG 농축 분석 결과, 후보 유전자들은 주로 세포체와 아미노산 자극에 대한 세포 반응에 관한 경로에 농축되어 있었다 (그림 3d).
Different methods can detect different timescale biological adaptation signatures [48]. Pairwise fixation index (FST) estimation between SDH and HNH was used to examine recent natural selection signals (Supplementary Table 14). The top 0.1% of FST results were detected as candidate loci, but this method cannot determine whether natural selection signals occurred in SDH or HNH. We conducted PBS again and chose HNH and HNL as the second and third reference populations, respectively (Supplementary Fig. 8a and Supplementary Table 15). Overall, 193 SNPs of the top 0.1% of selection signals identified in the PBS group were also supported by the cross-population extended haplotype homozygosity (XP-EHH) approach and FST(Fig. 3e-f and Supplementary Table 16). The gene encoding CACNAIA (rs16029, calcium voltage-gated channel subunit alpha1A), a gene implicated in regulating calcium ion entry into excitable cells and neurotransmitter release within the nervous system, was expressed at the highest level in the PBS-treated group. Additionally, six other genes demonstrating adaptive signatures were identified: LILRA3 (leukocyte immunoglobulin-like receptor A3), MTHFR (methylenetetrahydrofolate reductase), GJB2 (gap junction protein beta 2), FADS1 (fatty acid desaturase), FADS2 and KRT14 (keratin 14).
서로 다른 방법들은 다른 시간 척도의 생물학적 적응 신호를 감지할 수 있다 [48]. 산동 한족(山東漢族)과 호남 한족(湖南漢族) 사이의 쌍별 고정 지수(FST) 추정치를 사용하여 최근의 자연선택 신호를 조사했다 (보충 표 14). 상위 0.1%의 FST 결과가 후보 유전자 좌위로 검출되었지만, 이 방법으로는 자연선택 신호가 산동 한족(山東漢族)에서 발생했는지 호남 한족(湖南漢族)에서 발생했는지 결정할 수 없다. 우리는 PBS 분석을 다시 수행하여 호남 한족(湖南漢族)과 경중(瓊中) 여족(黎族)을 각각 두 번째와 세 번째 참조 집단으로 선택했다 (보충 그림 8a 및 보충 표 15). 전반적으로, PBS 그룹에서 확인된 상위 0.1% 선택 신호 중 193개의 단일염기다형성(SNP)은 집단 간 확장 일배체형 동형접합성(XP-EHH) 접근법과 FST에 의해서도 뒷받침되었다 (그림 3e-f 및 보충 표 16). 흥분성 세포로의 칼슘 이온 유입과 신경계 내 신경전달물질 방출을 조절하는 데 관여하는 유전자인 CACNAIA(rs16029)를 암호화하는 유전자가 PBS 분석 그룹에서 가장 높은 수준으로 발현되었다. 추가적으로, 적응 신호를 보이는 다른 6개 유전자, 즉 LILRA3, MTHFR, GJB2, FADS1, FADS2, KRT14가 확인되었다. 서로 다른 방법들은 다른 시간 척도의 생물학적 적응 신호를 감지할 수 있다 [48]. 산동 한족(山東漢族)과 호남 한족(湖南漢族) 사이의 쌍별 고정 지수(FST) 추정치를 사용하여 최근의 자연선택 신호를 조사했다 (보충 표 14). 상위 0.1%의 FST 결과가 후보 유전자 좌위로 검출되었지만, 이 방법으로는 자연선택 신호가 산동 한족(山東漢族)에서 발생했는지 호남 한족(湖南漢族)에서 발생했는지 결정할 수 없다. 우리는 PBS 분석을 다시 수행하여 호남 한족(湖南漢族)과 경중(瓊中) 여족(黎族)을 각각 두 번째와 세 번째 참조 집단으로 선택했다 (보충 그림 8a 및 보충 표 15). 전반적으로, PBS 그룹에서 확인된 상위 0.1% 선택 신호 중 193개의 단일염기다형성(SNP)은 집단 간 확장 일배체형 동형접합성(XP-EHH) 접근법과 FST에 의해서도 뒷받침되었다 (그림 3e-f 및 보충 표 16). 흥분성 세포로의 칼슘 이온 유입과 신경계 내 신경전달물질 방출을 조절하는 데 관여하는 유전자인 CACNAIA(rs16029)를 암호화하는 유전자가 PBS 분석 그룹에서 가장 높은 수준으로 발현되었다. 추가적으로, 적응 신호를 보이는 다른 6개 유전자, 즉 LILRA3, MTHFR, GJB2, FADS1, FADS2, KRT14가 확인되었다.
Notably, LILRA3, a leukocyte immunoglobulin-like receptor (LILR) family member, plays a role in the immune response [49]. The selected mutation locus (rs410852) within the LILRA3 gene can increase Takayasu arteritis (TAK) susceptibility, which is inferred from GWAS catalog data. TAK is predominantly prevalent in East Asians and is potentially linked to ethnic background [50]. The identified mutation can stimulate proinflammatory cytokine production and induce the proliferation of specific immune cell types [51]. The MTHFR gene encodes a methylenetetrahydrofolate reductase and maintains the balance between methionine and homocysteine [52]. Two polymorphisms within MTHFR, rs1801133 and rs9651118, are under natural selection pressure. Notably, rs1801133 is the most common genetic determinant for methylenetetrahydrofolate reductase deficiency, a condition that heightens the risk of cardiovascular disease [52]. Moreover, rs9651118 is linked to moyamoya disease and red blood cell distribution width changes. Two GJB2 gene mutations, rs72474224 and rs2274084, were associated with nonsyndromic hearing loss. The rs72474224 variant (c.109G > A, p.V371) is especially prevalent in East Asians and confers a substantial genetic predisposition to hearing impairment [53]. Additionally, our analyses further revealed other signals of selective sweeps in genes associated with diverse phenotypic traits (Supplementary Fig. 8b), including skin pigmentation (MCIR, melanocortin 1 receptor), stature (VPS9D1, VPS9 domain containing 1), and smoking behavior (SLC38A3, solute carrier family 38 member 3). Enrichment analyses suggested that these adaptive genes were involved in biological processes such as transport across the plasma membrane, phosphotransferase activity with alcohol groups as acceptors, and immune receptor activity (Supplementary Fig. 8c).
특히, 백혈구 면역글로불린 유사 수용체(LILR) 계열의 일원인 LILRA3는 면역 반응에 역할을 한다 [49]. LILRA3 유전자 내에서 선택된 돌연변이 좌위(rs410852)는 다카야스 동맥염(TAK) 감수성을 증가시킬 수 있으며, 이는 GWAS 카탈로그 데이터에서 추론되었다. 다카야스 동맥염은 주로 동아시아인에게 널리 퍼져 있으며 민족적 배경과 관련이 있을 가능성이 있다 [50]. 확인된 돌연변이는 염증 촉진 사이토카인 생산을 자극하고 특정 면역 세포 유형의 증식을 유도할 수 있다 [51]. MTHFR 유전자는 메틸렌테트라하이드로폴레이트 환원효소를 암호화하며 메티오닌과 호모시스테인 사이의 균형을 유지한다 [52]. MTHFR 내의 두 다형성, rs1801133과 rs9651118은 자연선택압을 받고 있다. 특히 rs1801133은 심혈관 질환 위험을 높이는 MTHFR 결핍증의 가장 흔한 유전적 결정 요인이다 [52]. 또한 rs9651118은 모야모야병 및 적혈구 분포 폭 변화와 관련이 있다. 두 개의 GJB2 유전자 돌연변이, rs72474224와 rs2274084는 비증후군성 난청과 관련이 있었다. rs72474224 변이는 특히 동아시아인에게 널리 퍼져 있으며 청력 손상에 대한 상당한 유전적 소인을 부여한다 [53]. 추가적으로, 우리의 분석은 피부 색소 침착(MCIR), 신장(VPS9D1), 흡연 행동(SLC38A3) 등 다양한 표현형 특성과 관련된 유전자에서 다른 선택적 소탕(selective sweep) 신호를 추가로 밝혔다 (보충 그림 8b). 농축 분석 결과, 이러한 적응 유전자들은 원형질막을 통한 수송, 알코올 그룹을 수용체로 하는 인산기전달효소 활성, 면역 수용체 활성과 같은 생물학적 과정에 관여하는 것으로 나타났다 (보충 그림 8c).
Highly differentiated variants between northern Han and southern Han people
북부 한족과 남부 한족 간의 고도로 분화된 변이
Population structure and selection pressure can produce HDVS. FST(SDH-HNH) results were used to identify variants with significant differences in allele frequency between the northern and southern Han populations. Based on the top 0.1% of FST values, our findings initially revealed that these HDVs are mainly associated with endemic pathogen exposure and dietary habits, exemplified by genes such as LILRA3, CR1 (complement receptor 1) and FADS. CR1 is associated with malaria resistance, whereas the FADS gene family participates in fatty acid metabolism [54]. Their frequency distribution varied among different populations, with high frequencies observed in southern populations (Supplementary Fig. 8d-g). A variant of particular interest, rs2296651 (c.800C > T, p.Ser267Phe), which is located in solute carrier family 10 member 1 (SLC10A1), is implicated in abnormal bilirubin metabolism. SLC10A1 encodes Na-taurocholate cotransporting polypeptide (NTCP), which facilitates the transport of conjugated bile acids into hepatocytes [55]. The S267F variant leads to NTCP deficiency, and its clinical manifestations include indirect hyperbilirubinemia and transient cholestatic jaundice [56]. Moreover, SLC10A1 serves as a receptor for the hepatitis B virus (HBV), and this variant can increase resistance to chronic hepatitis B [57, 58]. Information from the GWAS catalog revealed that the rs2296651 mutation was associated with alterations in the levels of multiple metabolites, including glycocholic acid, low-density lipoprotein cholesterol, and total cholesterol. The highest allele frequency for this variant was observed in East Asians according to the gnomAD database, with the loci under natural selection confirmed by the XP-EHH method (Supplementary Fig. 9f). Our study also revealed other complex relationships between HDVs and disease susceptibility. For instance, we identified two variants in the BPTF gene. The rs12602912 variant is associated with body mass index, and rs7216064 determines genetic susceptibility to lung adenocarcinoma and lung carcinoma.
인구 구조와 선택압은 고도로 분화된 변이(HDV)를 생성할 수 있다. 산동 한족(山東漢族)-호남 한족(湖南漢族) 간 FST 결과를 사용하여 북부와 남부 한족(漢族) 인구 간에 대립유전자 빈도에 상당한 차이가 있는 변이를 식별했다. 상위 0.1%의 FST 값을 기반으로, 우리의 연구 결과는 이러한 고도로 분화된 변이들이 주로 풍토성 병원체 노출 및 식습관과 관련이 있음을 초기에 밝혔다. LILRA3, CR1, FADS와 같은 유전자가 그 예이다. CR1은 말라리아 저항성과 관련이 있고, FADS 유전자군은 지방산 대사에 참여한다 [54]. 이들의 빈도 분포는 여러 인구 집단 간에 다양했으며, 남부 인구에서 높은 빈도가 관찰되었다 (보충 그림 8d-g). 특히 흥미로운 변이인 SLC10A1 유전자에 위치한 rs2296651은 비정상적인 빌리루빈 대사와 관련이 있다. SLC10A1은 담즙산의 간세포 내 수송을 촉진하는 NTCP를 암호화한다 [55]. S267F 변이는 NTCP 결핍을 유발하며, 임상 증상으로는 간접 고빌리루빈혈증과 일과성 담즙 정체성 황달이 있다 [56]. 더욱이, SLC10A1은 B형 간염 바이러스(HBV)의 수용체 역할을 하며, 이 변이는 만성 B형 간염에 대한 저항성을 높일 수 있다 [57, 58]. GWAS 카탈로그 정보에 따르면 rs2296651 돌연변이는 글리코콜산, 저밀도 지단백 콜레스테롤, 총 콜레스테롤을 포함한 여러 대사 산물의 수치 변화와 관련이 있었다. 이 변이의 가장 높은 대립유전자 빈도는 gnomAD 데이터베이스에 따라 동아시아인에게서 관찰되었으며, 해당 유전자 좌위가 자연선택 하에 있음은 XP-EHH 방법으로 확인되었다 (보충 그림 9f). 우리 연구는 또한 고도로 분화된 변이와 질병 감수성 사이의 다른 복잡한 관계를 밝혔다. 예를 들어, BPTF 유전자에서 두 개의 변이를 확인했다. rs12602912 변이는 체질량지수와 관련이 있고, rs7216064는 폐선암 및 폐암에 대한 유전적 감수성을 결정한다.
Haplotype analysis and allele frequency trajectories of ABCC11 and SLC10A1
ABCC11과 SLC10A1의 일배체형 분석 및 대립유전자 빈도 궤적
The evolutionary histories of numerous genetic variants remain elusive, presenting a barrier to understanding the genetic basis of complex traits or diseases. The ABCC11 gene serves as a significant biological adaptive signal, determining earwax type and axillary odor. Some hypotheses have posited that the A allele of ABCC11 (rs17822931) may confer adaptive advantages in colder climates. The ancestral environment of East Asians is thought to have been much colder than that of Africans. The ABCC11 mutation, characterized by diminished sweat gland activity, has endowed humans with the ability to better preserve body heat in colder climates [40]. Ancient genomic data revealed that the mutation emerged approximately 44,000 years ago in the Ust_Ishim individual rather than in ancient African populations. However, it was less commonly found in populations residing in equatorial regions with higher temperatures two thousand years ago, suggesting a correlation with latitude (Fig. 4a-c and Supplementary Fig. 8 h-j). Subsequently, the mutation also occurred in southern Asia and Oceania for nearly a thousand years, likely attributable to the southward migration of populations from colder northern areas and admixture with indigenous inhabitants. In China, we found that the derived allele of the ABCC11 gene was first identified in ancient Tianyuan individuals. In North America, rs17822931-T appeared approximately 12,000 years ago. Contemporary East Asians exhibit the highest frequency of this mutation, followed by North American and South American populations. These observations confirmed the association of the mutation with cold environmental adaptation and its status as a region-specific adaptive signal, particularly among East Asians. Indeed, we observed that the frequency of derived alleles was greater in northern Chinese populations due to colder environments (Fig. 3g).
수많은 유전 변이의 진화 역사는 여전히 불분명하여 복잡한 형질이나 질병의 유전적 기초를 이해하는 데 장벽이 되고 있다. ABCC11 유전자는 귀지 유형과 겨드랑이 냄새를 결정하는 중요한 생물학적 적응 신호로 작용한다. 일부 가설은 ABCC11의 A 대립유전자(rs17822931)가 추운 기후에서 적응적 이점을 제공할 수 있다고 제시했다. 동아시아인의 조상 환경은 아프리카인보다 훨씬 추웠을 것으로 생각된다. 땀샘 활동 감소를 특징으로 하는 ABCC11 돌연변이는 인간이 추운 기후에서 체온을 더 잘 보존할 수 있는 능력을 부여했다 [40]. 고대 게놈 데이터에 따르면, 이 돌연변이는 고대 아프리카 인구가 아닌 약 44,000년 전 우스트-이심(Ust_Ishim)인에게서 나타났다. 그러나 2천 년 전에는 기온이 높은 적도 지역에 거주하는 인구에서는 덜 흔하게 발견되었으며, 이는 위도와의 상관관계를 시사한다 (그림 4a-c 및 보충 그림 8 h-j). 그 후, 이 돌연변이는 약 천 년 동안 남부 아시아와 오세아니아에서도 발생했는데, 이는 추운 북부 지역 인구의 남하와 토착민과의 혼합 때문일 것이다. 중국에서는 ABCC11 유전자의 파생 대립유전자가 고대 전문(田園)인에게서 처음 확인되었다. 북미에서는 약 12,000년 전에 rs17822931-T가 나타났다. 현대 동아시아인은 이 돌연변이의 빈도가 가장 높으며, 북미 및 남미 인구가 그 뒤를 잇는다. 이러한 관찰은 이 돌연변이가 추운 환경 적응과 관련이 있으며, 특히 동아시아인 사이에서 지역 특이적 적응 신호임을 확인시켜 주었다. 실제로 우리는 추운 환경 때문에 북부 중국 인구에서 파생 대립유전자의 빈도가 더 높다는 것을 관찰했다 (그림 3g).
We further analyzed the haplotype diversity across global populations and constructed a linkage disequilibrium plot (Supplementary Fig. 9a). We found that Hap2, which carries the derived allele sequence (CTT GCT), was predominantly distributed in Asian people, followed by Europeans (Fig. 4g and 4i). However, Oceanians possessed a distinct haplotype (Hap8, TCTGCT ) that contained the derived alleles. Haplotype diversity was constrained among Han populations, with most individuals carrying the derived allele (Supplementary Fig. 9c). The high-frequency haplotype was widespread among minority ethnic groups. However, the AN-speaking populations had different dominant haplotypes, with a comparatively low CTTGCT haplotype frequency.
우리는 전 세계 인구에 걸쳐 일배체형 다양성을 추가로 분석하고 연관 불균형 그림을 구성했다(보충 그림 9a). 파생 대립유전자 서열(CTT GCT)을 가진 일배체형2(Hap2)는 주로 아시아인에게 분포했고, 유럽인이 그 뒤를 이었다(그림 4g, 4i). 그러나 오세아니아인은 파생 대립유전자를 포함하는 별개의 일배체형(Hap8, TCTGCT)을 가지고 있었다. 한족(漢族) 인구 사이에서는 일배체형 다양성이 제한적이었으며, 대부분의 개인이 파생 대립유전자를 가지고 있었다(보충 그림 9c). 고빈도 일배체형은 소수 민족 집단에 널리 퍼져 있었다. 그러나 오스트로네시아어족(AN-speaking) 인구는 다른 우세한 일배체형을 가졌으며, CTTGCT 일배체형 빈도는 비교적 낮았다.
Although NTCP deficiency is a newly described disease, ancient DNA can provide crucial information about historical allele frequency fluctuations and the geographical distribution of associated mutations. We estimated the temporal origin of the SLC10A1 mutation (rs2296651) and its historical prevalence by examining allele frequency trajectories over time (Fig. 4d-f and Supplementary Fig. 8k-m). The derived SLC10A1 variant (rs2296651-A) initially emerged in the Iberomaurusian population of Morocco, with the earliest evidence dating back to approximately 12,849-12,097 calibrated years BCE. This variant was subsequently detected in European populations, such as Italians and Spaniards, and we also found it in the Eneolithic Russia Shamanka circa 7000 to 8000 years ago. The mutation gradually spread eastward, and its presence was discerned in the early period in individuals from California’s Channel Islands, where it emerged approximately 4915 years ago. The S267F mutation also appeared in populations from Lebanon and Indonesia during the transition from the Early Bronze Age to the Iron Age. Remarkably, genomic data from Taiwan Hanben revealed the presence of the mutation in ancient southern Chinese individuals approximately 1600 years ago. The derived allele was later identified in ancient Guangxi people (BaBanQinCen) approximately 1400 years ago. Afterward, it was found in Malaysia during its historical period, in which it was composed of Micronesians approximately 580 years ago and Vanuatu people approximately 150 BP. Modern genomic data indicate that the S267F variant is prevalent in Southeast Asian, Oceanian, and southern Chinese populations.
NTCP 결핍증은 새롭게 기술된 질병이지만, 고대 DNA는 역사적인 대립유전자 빈도 변동과 관련 돌연변이의 지리적 분포에 대한 중요한 정보를 제공할 수 있다. 우리는 시간에 따른 대립유전자 빈도 궤적을 조사하여 SLC10A1 돌연변이(rs2296651)의 시간적 기원과 역사적 유병률을 추정했다(그림 4d-f 및 보충 그림 8k-m). 파생 SLC10A1 변이(rs2296651-A)는 처음에 모로코의 이베로마우루스인(Iberomaurusian) 인구에서 나타났으며, 가장 오래된 증거는 기원전 약 12,849-12,097년으로 거슬러 올라간다. 이 변이는 이후 이탈리아인과 스페인인 같은 유럽 인구에서 발견되었으며, 약 7000년에서 8000년 전 동석기 시대 러시아 샤만카(Shamanka)에서도 발견되었다. 돌연변이는 점차 동쪽으로 퍼져나갔고, 약 4915년 전 캘리포니아 채널 제도 개인들에게서 초기 모습이 확인되었다. S267F 돌연변이는 초기 청동기 시대에서 철기 시대로 전환되는 동안 레바논과 인도네시아 인구에서도 나타났다. 놀랍게도, 대만 한본(台灣漢本)의 게놈 데이터는 약 1600년 전 고대 남부 중국인에게서 이 돌연변이가 존재했음을 밝혔다. 파생 대립유전자는 나중에 약 1400년 전 고대 광서(廣西) 사람들(파판친천, BaBanQinCen)에게서 확인되었다. 그 후, 말레이시아의 역사 시대에 발견되었는데, 약 580년 전 미크로네시아인과 약 150년 전 바누아투 사람들에게서 구성되었다. 현대 게놈 데이터는 S267F 변이가 동남아시아, 오세아니아, 남부 중국 인구에 널리 퍼져 있음을 나타낸다.
Analysis of the allele frequency trajectory across major intercontinental populations suggested a notable increase in the variant approximately 2000 years ago, with a pronounced increase in prevalence in the Asian and Oceanian groups, particularly among coastal southern Chinese populations. NTCP, encoded by the SLC10A1 gene, is a cellular receptor for HBV and is significantly associated with resistance to chronic hepatitis B [59]. Given the high prevalence of HBV within Chinese populations and the emergence of agriculture in East Asia, we further analyzed the driving force of this biological adaptation [42, 60]. The timeline of mutation emergence allowed us to rule out agricultural development as a driving factor of biological adaptation, instead suggesting a potential link between NTCP deficiency and enhanced pathogen resistance.
주요 대륙 간 인구에 걸친 대립유전자 빈도 궤적 분석은 약 2000년 전에 이 변이가 눈에 띄게 증가했으며, 아시아와 오세아니아 그룹, 특히 해안 남부 중국 인구 사이에서 유병률이 현저히 증가했음을 시사했다. SLC10A1 유전자에 의해 암호화되는 NTCP는 B형 간염 바이러스(HBV)의 세포 수용체이며 만성 B형 간염에 대한 저항성과 유의하게 관련이 있다 [59]. 중국 인구 내 HBV의 높은 유병률과 동아시아 농업의 출현을 고려하여, 우리는 이 생물학적 적응의 원동력을 추가로 분석했다 [42, 60]. 돌연변이 출현 시기를 통해 우리는 농업 발전을 생물학적 적응의 원동력에서 배제할 수 있었고, 대신 NTCP 결핍과 향상된 병원체 저항성 사이의 잠재적 연관성을 시사했다.
Genetic diversity within pathogenic SLC10A1 allele carrier haplotypes varies in multiple populations. To explore the genetic landscape of the S267F allele carrier haplotypes, we constructed a linkage disequilibrium plot and applied the haplotype-based inference method (Supplementary Fig. 9b). Our analysis revealed 33 haplotypes of the SLC10A1 gene across six intercontinental populations, with a focus on the primary 20 haplotypes for in-depth analysis (Fig. 4h and 4j). Hap18 (CCGTGA GAAGC) was identified as the ancestral haplotype of SLC10A1 and is present mainly in Oceanians. However, due to the sparse sampling in Africa, whether ancestral haplotypes exist in Africans remains uncertain. Hap5, characterized by a derived mutation, is widely distributed across Asians, Europeans, and Africans. Conversely, Hap20, which also carried pathogenic genetic variation, was exclusively observed in populations from Oceania. For the S267F mutation, the haplotype (ACGTAA AGGGC) derived from H1 (ACGTGAAGGGC) existed in three linguistically distinct populations, namely, the TK, HM, and TB populations (Supplementary Fig. 9d). Han people predominantly characterized H1, and more diverse haplotypes can be found in other minority ethnic groups, such as the Altaic and AN people
병원성 SLC10A1 대립유전자를 가진 일배체형 내의 유전적 다양성은 여러 인구에서 다양하다. S267F 대립유전자를 가진 일배체형의 유전적 지형을 탐색하기 위해, 우리는 연관 불균형 그림을 구성하고 일배체형 기반 추론 방법을 적용했다 (보충 그림 9b). 우리의 분석은 6개 대륙 간 인구에 걸쳐 SLC10A1 유전자의 33개 일배체형을 밝혔으며, 심층 분석을 위해 주요 20개 일배체형에 초점을 맞췄다 (그림 4h, 4j). 일배체형18(Hap18)은 SLC10A1의 조상 일배체형으로 확인되었으며 주로 오세아니아인에게 존재한다. 그러나 아프리카의 표본 추출이 부족하여 아프리카인에게 조상 일배체형이 존재하는지는 불확실하다. 파생 돌연변이를 특징으로 하는 일배체형5(Hap5)는 아시아인, 유럽인, 아프리카인에 널리 분포한다. 반대로, 병원성 유전 변이를 지닌 일배체형20(Hap20)은 오세아니아 인구에서만 독점적으로 관찰되었다. S267F 돌연변이의 경우, H1에서 파생된 일배체형이 타이-카다이어족(TK), 몽-미엔어족(HM), 티베트-버마어족(TB)이라는 세 개의 언어적으로 다른 인구에 존재했다 (보충 그림 9d). 한족(漢族)은 주로 H1을 특징으로 했으며, 알타이어족 및 오스트로네시아어족(AN)과 같은 다른 소수 민족 집단에서는 더 다양한 일배체형을 찾을 수 있다.
Medical relevance
의학적 관련성
We aggregated all high-quality adaptive loci and HDVs based on the above stringent criteria and evaluated their potential impact on phenotype using three computational prediction methods: Sorting Intolerant from Tolerant (SIFT) [61], Polymorphism Phenotyping (PolyPhen) [62], and Combined Annotation Dependent Depletion (CADD) [63]. Two variants within the G/B2 gene, rs2274084 and rs72474224, were predicted to be likely damaging, with PolyPhen assigning damage probabilities of 0.998 and 0.995, respectively. Furthermore, the CADD score indicated a relatively high pathogenic risk (exceeding a score of 20) for the rs2274084 (c.79G > A) variant in GJB2. The clinical significance of this mutation was categorized as pathogenic or likely pathogenic in the ClinVar database. GJB2 encodes Connexin26, which is the most crucial gap junction protein in the cochlea. The gap junction system plays an important role in maintaining normal potassium circulation and the microenvironment in the cochlea. These findings all implied that these genetic variations may adversely affect congenital deafness.
우리는 위의 엄격한 기준에 따라 모든 고품질 적응 좌위와 고도로 분화된 변이(HDV)를 집계하고, 세 가지 컴퓨터 예측 방법(SIFT [61], PolyPhen [62], CADD [63])을 사용하여 표현형에 대한 잠재적 영향을 평가했다. GJB2 유전자 내의 두 변이, rs2274084와 rs72474224는 손상을 일으킬 가능성이 있는 것으로 예측되었으며, PolyPhen은 각각 0.998과 0.995의 손상 확률을 부여했다. 또한 CADD 점수는 GJB2의 rs2274084 변이에 대해 비교적 높은 병원성 위험(20점 초과)을 나타냈다. 이 돌연변이의 임상적 중요성은 ClinVar 데이터베이스에서 병원성 또는 병원성 가능으로 분류되었다. GJB2는 달팽이관에서 가장 중요한 간극 연접 단백질인 코넥신26(Connexin26)을 암호화한다. 간극 연접 시스템은 달팽이관 내 정상적인 칼륨 순환과 미세 환경을 유지하는 데 중요한 역할을 한다. 이러한 모든 발견은 이들 유전 변이가 선천성 난청에 부정적인 영향을 미칠 수 있음을 시사했다.
Intriguingly, our analysis of the minor allele frequency of these two GJB2 mutations revealed distinct latitude-related distribution patterns. Specifically, the frequency of frs72474224-T was markedly greater in southern populations, while rs2274084-T was more prevalent in northern populations (Supplementary Fig. 9g-h). Additionally, rs11886868 in BCL11A was associated with fetal hemoglobin (HbF) levels, with clinical significance deemed benign or likely benign. These findings illuminated the intricate interplay between genetic variations and geographic environment factors, highlighting the multifaceted nature of population genetics.
흥미롭게도, 이 두 GJB2 돌연변이의 소수 대립유전자 빈도 분석 결과 뚜렷한 위도 관련 분포 패턴이 나타났다. 구체적으로, rs72474224-T의 빈도는 남부 인구에서 현저히 높았고, rs2274084-T는 북부 인구에서 더 널리 퍼져 있었다 (보충 그림 9g-h). 추가적으로, BCL11A의 rs11886868은 태아 헤모글로빈(HbF) 수치와 관련이 있었으며, 임상적 중요성은 양성 또는 양성 가능으로 간주되었다. 이러한 발견은 유전 변이와 지리적 환경 요인 간의 복잡한 상호작용을 조명하며, 집단 유전학의 다면적 특성을 강조했다.
The SLC10A1 variant (rs2296651) was predicted to be likely damaged by PolyPhen, with a CADD score also exceeding 20. Furthermore, in line with the American College of Medical Genetics and Genomics (ACMG) guidelines, rs2296651 has been classified as pathogenic or likely pathogenic. The derived allele displayed a greater frequency in southern Chinese populations, especially in people from Guangdong, Guangxi, and Hainan provinces (Fig. 3i). Although the S267F mutation may induce hepatocyte injury due to excessive bile acids, it concurrently provides a protective effect against HBV infection and HBV-related diseases, such as cirrhosis and hepatocellular carcinoma [59, 64]. The A allele is associated with a loss of function in NTCP, thereby interfering with NTCP and HBV binding. More variants with adverse effects have been identified, including rs1801133 in the MTHFR gene, which SIFT classifies as deleterious, and PolyPhen as probably damaging, supported by a CADD score above 20. The rs1801133-T locus is associated with decreased MTHFR enzyme activity, leading to elevated levels of homocysteine. Notably, hyperhomocysteinemia is a known risk factor for stroke [65]. Similarly, the CR1 variant rs2274567 is considered to be damaging, according to PolyPhen. CR1 is an immune receptor that regulates the complement system and is involved in regulating immune responses ponses and clearing immune complexes. This variant may affect the expression level or function of CR1, leading to an imbalance in complement system regulation and increasing the risk of some diseases [66]. We further characterized the other two known loci associated with disease traits, delineating their allele frequency disparities (Fig. 3h and j). These candidate loci exhibited differential biological adaptation within Chinese populations, reflecting complex natural selection pressure that can influence genetic predispositions and disease risk discrimination among populations in different regions. Enhancing preventive measures and tailoring clinical interventions for populations with a genetic predisposition to specific diseases is imperative.
SLC10A1 변이(rs2296651)는 PolyPhen에 의해 손상 가능성이 있는 것으로 예측되었고, CADD 점수도 20을 초과했다. 더욱이, 미국 의학유전학회(ACMG) 지침에 따라 rs2296651은 병원성 또는 병원성 가능으로 분류되었다. 파생 대립유전자는 남부 중국 인구, 특히 광동(廣東), 광서(廣西), 해남(海南)성 사람들에서 더 높은 빈도를 보였다 (그림 3i). S267F 돌연변이는 과도한 담즙산으로 인해 간세포 손상을 유발할 수 있지만, 동시에 B형 간염 바이러스(HBV) 감염 및 간경변, 간세포암종과 같은 HBV 관련 질병에 대한 보호 효과를 제공한다 [59, 64]. A 대립유전자는 NTCP의 기능 상실과 관련이 있어 NTCP와 HBV의 결합을 방해한다. MTHFR 유전자의 rs1801133을 포함하여 부작용이 있는 더 많은 변이가 확인되었다. SIFT는 이를 해롭다고 분류했고, PolyPhen은 아마도 손상을 줄 것이라고 예측했으며, 20 이상의 CADD 점수가 이를 뒷받침한다. rs1801133-T 좌위는 MTHFR 효소 활성 감소와 관련이 있어 호모시스테인 수치를 높인다. 특히, 고호모시스테인혈증은 뇌졸중의 알려진 위험 인자이다 [65]. 유사하게, CR1 변이 rs2274567은 PolyPhen에 따르면 손상을 주는 것으로 간주된다. CR1은 보체계를 조절하는 면역 수용체이며, 면역 반응 조절 및 면역 복합체 제거에 관여한다. 이 변이는 CR1의 발현 수준이나 기능에 영향을 미쳐 보체계 조절의 불균형을 초래하고 일부 질병의 위험을 증가시킬 수 있다 [66]. 우리는 질병 특성과 관련된 다른 두 개의 알려진 좌위를 추가로 특성화하고, 그들의 대립유전자 빈도 차이를 기술했다 (그림 3h, 3j). 이 후보 좌위들은 중국 인구 내에서 차별적인 생물학적 적응을 보였으며, 이는 다른 지역의 인구 간 유전적 소인과 질병 위험 구별에 영향을 미칠 수 있는 복잡한 자연선택압을 반영한다. 특정 질병에 대한 유전적 소인이 있는 인구를 위해 예방 조치를 강화하고 맞춤형 임상 개입을 하는 것이 필수적이다.
그림 3. 북부 한족(漢族)의 생물학적 적응 신호. a 맨해튼 플롯은 산동 한족(山東漢族)의 자연선택 신호를 보여줌. 빨간 선 위의 상위 0.1% PBS 값이 확인됨. b 생키 다이어그램(Sankey diagram)은 유전자의 다면발현성(pleiotropism)을 반영함. c LocusZoom은 PBS 값 상위 0.1%에 속하는 ABCC11 변이의 지역적 시각화를 제공함. d PBS(SDH-HNL-CEU 3자 모델)로 식별된 상위 0.1% 후보 유전자에 대한 KEGG 및 GO 농축 분석. e PBS로 선택된 두 SNP(rs17822931 및 rs22744084)의 확장된 일배체형 동형접합성. f 여러 다른 방법을 기반으로 산동 한족(山東漢族)에서 공통으로 선택된 후보 유전자를 보여주는 벤 다이어그램. g-j 중국 인구 집단들 사이에서 고도로 분화된 변이들의 적응 대립유전자 빈도 및 파생 대립유전자 빈도 분포.
고찰 Discussion
Ancient individuals from the YRB and YZRB participated in shaping the genetic landscape of present-day geographically diverse Han populations. The emergence and progression of ancient agriculture along these river systems facilitated population expansion, migration, and admixture. Previous studies have demonstrated that the southward migration of millet farmers influences the genetic profile of southern populations. Furthermore, the northward expansion of rice farmers from the Middle to Late Neolithic period also imparted genetic legacies to the gene pool of northern populations. As one of the agricultural domestication centers of millet farming in China, investigating the genetic structure and population relationships among Han populations residing in the lower YRB at a fine-scale level holds considerable scientific merit. This study explored the intricate genetic interplay between ancient and modern populations and shed light on how these interactions further influenced the genetic makeup of the modern Han. Evidence for a complex demographic history revealed the long-term genetic stability of SDH in the lower YRB. Additionally, numerous East Asian-specific adaptative genetic variants were identified. We highlighted the adaptative signals associated with a rare genetic disorder of bile acid metabolism and presented their allele frequency trajectories and haplotype networks, which provided a classic example of the evolutionary trade-off between health and fitness.
황하 유역(黃河流域)과 양자강 유역(揚子江流域)의 고대인들은 오늘날 지리적으로 다양한 한족(漢族) 인구의 유전적 지형을 형성하는 데 기여했다. 이 강 유역을 따라 발달한 고대 농업은 인구 팽창, 이주, 그리고 유전적 혼합을 촉진했다 [67]. 기존 연구에 따르면, 기장 농부들이 남쪽으로 이동하면서 남부 인구의 유전적 특성에 영향을 미쳤다 [13]. 반대로, 신석기 중기에서 후기 사이에는 벼농사를 짓던 농부들이 북쪽으로 이동하며 북부 인구의 유전자 풀에도 유전적 유산을 남겼다. 중국의 기장 농업 중심지였던 황하(黃河) 하류 지역 한족(漢族)의 유전 구조와 인구 관계를 정밀하게 조사하는 것은 매우 중요한 과학적 의미를 가진다. 이 연구는 고대와 현대 인구 간의 복잡한 유전적 상호작용을 탐구했으며, 이것이 현대 한족(漢族)의 유전적 구성에 어떤 영향을 미쳤는지 밝혔다. 복잡한 인구 역사를 추적한 결과, 황하(黃河) 하류 지역의 산동 한족(山東漢族)이 장기간에 걸쳐 유전적 안정성을 유지해왔다는 사실이 드러났다. 또한, 동아시아인에게 특유한 여러 적응 유전 변이도 발견했다. 특히 담즙산 대사와 관련된 희귀 유전 질환의 적응 신호를 집중적으로 조명했으며, 그 대립유전자 빈도 변화와 계통을 추적했다. 이는 건강과 생존 적합성 사이의 ‘진화적 상충 관계’를 보여주는 전형적인 사례다.
Demographic history and genetic structure
인구 역사와 유전 구조
Previous studies based on mitochondrial genomes indicated that Shandong served as a cultural crossroads facilitating the migration of ancient individuals from North to South China. Haplogroup analysis also revealed continuous maternal genetic stability in this region from the Early Neolithic to the Late Neolithic Period [38]. Ancient Shandong individuals and ancient coastal southern East Asians were separated into two clusters of approximately eight kya [13]. As an ethnic group dominated by an agricultural civilization, the primary scope of activities for most proto-Han populations remained relatively fixed. However, the Han people have also experienced large-scale or small-scale population migration events for various reasons during prehistoric and historical periods. We integrated ancient and modern genomic datasets and data from 264 newly collected individuals in Shandong Province. Finally, our study leveraged three merged databases, namely, Affy_HO, Affy_HGDP, and Affy_1240K, for detailed genetic analysis. The PCA results revealed a North-South genetic cline of East Asians, with SDH displaying close genetic affinities with the central Han, Mongolians in Inner Mongolia, Japanese and Koreans (Fig. 1a). However, SDH exhibited a relatively distant genetic relationship with the AA and HM populations. Combined with ancient genomic data, SDH overlapped with ancient YRB-related individuals while displaying a remote genetic relationship with ancient individuals in southern China. The findings revealed via the FST, TreeMix, and IBD supported these results. The fineSTRUCTURE and comparison of southern ancestral components further revealed geography/language-related population stratification (Fig. 1c). We also explored the admixture landscape of SDH and identified possible ancestral source candidates based on ADMIXTURE and admixture -f₃ statistics, in which ANEA and ASEA contributed unevenly to the gene pool of SDH (Fig. 2a and Supplementary Fig. 2a-b).
미토콘드리아 유전체에 기반한 이전 연구들은 산동이 북중국에서 남중국으로의 고대인 이주를 촉진하는 문화적 교차로 역할을 했음을 시사했다. 하플로그룹 분석 또한 이 지역에서 신석기 초기부터 후기까지 지속적인 모계 유전적 안정성을 보여주었다 [38]. 고대 산동인과 고대 해안 남부 동아시아인은 약 8천 년 전에 두 개의 군집으로 분리되었다 [13]. 농업 문명이 지배적인 민족 집단으로서, 대부분의 원시 한족(漢族) 인구의 주 활동 범위는 비교적 고정되어 있었다. 그러나 한족(漢族)은 선사 및 역사 시대 동안 다양한 이유로 대규모 또는 소규모의 인구 이동 사건을 경험하기도 했다.
우리는 고대 및 현대 게놈 데이터세트와 산동성(山東省)에서 새로 수집한 264명의 데이터를 통합했다. 최종적으로, 우리 연구는 상세한 유전 분석을 위해 Affy_HO, Affy_HGDP, Affy_1240K라는 세 개의 병합된 데이터베이스를 활용했다. 주성분 분석(PCA) 결과, 동아시아인의 남북 유전적 연속선이 나타났으며, 산동 한족(山東漢族)은 중부 한족(漢族), 내몽골(內蒙古)의 몽골인, 일본인, 한국인과 가까운 유전적 친화성을 보였다 (그림 1a). 그러나 산동 한족(山東漢族)은 오스트로아시아어족(AA) 및 몽-미엔어족(HM) 인구와는 비교적 먼 유전적 관계를 보였다. 고대 게놈 데이터를 결합했을 때, 산동 한족(山東漢族)은 고대 황하 유역(黃河流域) 관련 개인들과 겹쳤지만, 남중국 고대인들과는 먼 유전적 관계를 보였다. FST, TreeMix, 동일 조상 유래(IBD)를 통해 밝혀진 결과들이 이를 뒷받침했다. fineSTRUCTURE와 남부 조상 구성 요소 비교는 지리/언어 관련 인구 계층화를 추가로 보여주었다 (그림 1c). 우리는 또한 산동 한족(山東漢族)의 혼합 지형을 탐색하고 ADMIXTURE와 혼합-f₃ 통계량에 기초하여 가능한 조상 공급원 후보를 확인했으며, 여기서 고대 북부 동아시아인(ANEA)과 고대 남부 동아시아인(ASEA)은 산동 한족(山東漢族)의 유전자 풀에 불균등하게 기여했다 (그림 2a 및 보충 그림 2a-b).
SDH is believed to have originated from a YRB-related lineage. Based on the examination of genetic continuity, we still found slight differences between the Yangshao people and the early Neolithic individuals in Shandong (Bianbian, Xiaojingshan, Xiaogao, and Boshan). From the Early to Middle Neolithic, the Siberian-related component declined in Shandong ancients. We detected gene flow between the northern and southern coastal areas. The expansion of the rice farming civilization to the north gradually influenced the Neolithic Longshan culture. There was no significant gene flow from other geographically different ancient people from the late Neolithic to the late Bronze/Iron Age. Additionally, we did not identify other ANEA populations contributing to the SDH gene pool compared with the YR_LBIA population (Supplementary Fig. 6a-c). The qpWave results validated the genetic homogeneity between YR_LBIA and SDH (Fig. 2d). Historically, the migration of SDH toward the South or North was influenced by events such as wars, floods, and other disasters. An example of such migration is the emigration to northeast China, known as the Chuangguandong migration event. However, compared to other Han populations and minority ethnic groups in the YRB, large-scale population migration into Shandong was relatively limited. Therefore, SDH and local temporally diverse ancient people exhibited relative genetic continuity over an extended period. The relative genetic stability can aid in exploring their genetic origin and determining their population history. Additionally, the connection between SDH and East Asian ancestries also underscores the complexity of ancient population dynamics, showing that the history of human migration is characterized by multiple waves of migration, admixture, and genetic exchange.
산동 한족(山東漢族)은 황하 유역(黃河流域) 관련 혈통에서 기원한 것으로 여겨진다. 유전적 연속성 조사에 따르면, 양소(仰韶) 사람들과 산동의 초기 신석기인들(변변(卞邊), 소경산(小荊山), 소고(蕭高), 보산(博山)) 사이에는 여전히 약간의 차이가 발견되었다. 신석기 초기에서 중기로 가면서 산동 고대인에게서 시베리아 관련 요소가 감소했다. 우리는 북부와 남부 해안 지역 간의 유전자 흐름을 감지했다. 벼농사 문명의 북쪽으로의 확장은 점차 신석기 용산(龍山) 문화에 영향을 미쳤다. 신석기 후기부터 후기 청동기/철기 시대까지 다른 지리적으로 다른 고대인으로부터의 유의미한 유전자 흐름은 없었다. 추가적으로, 후기 청동기/철기 시대 황하 유역 인구와 비교하여 산동 한족(山東漢族) 유전자 풀에 기여한 다른 고대 북부 동아시아인 인구는 확인하지 못했다 (보충 그림 6a-c). qpWave 결과는 후기 청동기/철기 시대 황하 유역 인구와 산동 한족(山東漢族) 사이의 유전적 동질성을 검증했다 (그림 2d).
역사적으로, 산동 한족(山東漢族)의 남쪽 또는 북쪽으로의 이주는 전쟁, 홍수 및 기타 재난과 같은 사건의 영향을 받았다. 그러한 이주의 예로는 ‘틈관동(闖關東)’ 이주 사건으로 알려진 동북 중국으로의 이주가 있다. 그러나 황하 유역(黃河流域)의 다른 한족(漢族) 인구 및 소수 민족 집단과 비교할 때, 산동으로의 대규모 인구 유입은 비교적 제한적이었다. 따라서 산동 한족(山東漢族)과 지역의 시간적으로 다양한 고대인들은 장기간에 걸쳐 상대적인 유전적 연속성을 보였다. 상대적인 유전적 안정성은 그들의 유전적 기원을 탐색하고 인구 역사를 결정하는 데 도움이 될 수 있다. 추가적으로, 산동 한족(山東漢族)과 동아시아 혈통 간의 연결은 고대 인구 동태의 복잡성을 강조하며, 인류 이주의 역사가 여러 차례의 이주, 혼합, 유전적 교류로 특징지어짐을 보여준다.
Natural selection signals and East Asian-specific variants
자연선택 신호와 동아시아 특이 변이
Genetic findings from the Chinese Academy of Sciences Precision Medicine Initiative (CASPMI) cohort revealed SNPs associated with waist circumference, BMI, lipid metabolism, and other traits in the northern Han population [68]. Our study also identified some new adaptive signatures associated with BMI-adjusted waist-hip ratio and height based on the iHS method, implying differences in physique between northern and southern populations (Supplementary Table 13). The PBS results showed that ABCC11 was under natural selection in the SDH and was associated with the AO and earwax types. The rs17822931 mutation was detected in both dry-type earwax and reduced body odor. We found that the derived allele of rs17822931 was mainly distributed in high-latitude regions and first appeared in Russia approximately 44,000 years ago. The frequency of the T allele has gradually increased in East Asians over the last ten thousand years, and the allele frequency even reached 0.8862, which is significantly different from that in populations on other continents (Supplementary Fig. 9e). As a high-frequency variant specific to East Asians, we need to pay more attention to the effect of ABCC11. Previous studies have indicated that this polymorphism could influence estrogen receptor-positive breast cancer, and the T allele might lead to low estrogen efflux activity and increase the risk of breast cancer [69]. From the evolutionary trajectory, we can gain insights into the biological adaptability of geographically distinct populations to their environments and infer population migration and admixture events. The derived allele was mainly distributed in northern populations two thousand years ago (Fig. 4a-c). The T allele frequency increased significantly in southern populations afterward, possibly associated with southward migration.
중국과학원 정밀의학 이니셔티브(CASPMI) 코호트의 유전적 발견은 북부 한족(漢族) 인구에서 허리둘레, 체질량지수(BMI), 지질 대사 및 기타 형질과 관련된 단일염기다형성(SNP)을 밝혔다 [68]. 우리 연구 또한 iHS 방법을 기반으로 BMI 보정 허리-엉덩이 비율 및 키와 관련된 새로운 적응 신호를 일부 확인했으며, 이는 북부와 남부 인구 간의 체격 차이를 암시한다 (보충 표 13). PBS 결과는 ABCC11이 산동 한족(山東漢族)에서 자연선택을 받았으며 겨드랑이 냄새(액취) 및 귀지 유형과 관련이 있음을 보여주었다. rs17822931 돌연변이는 건성 귀지와 체취 감소 모두에서 발견되었다. 우리는 rs17822931의 파생 대립유전자가 주로 고위도 지역에 분포하며 약 44,000년 전 러시아에서 처음 나타났음을 발견했다. T 대립유전자의 빈도는 지난 1만 년 동안 동아시아인에게서 점차 증가했으며, 대립유전자 빈도는 0.8862에 도달하여 다른 대륙의 인구와 현저한 차이를 보였다 (보충 그림 9e). 동아시아인에게 특이적인 고빈도 변이로서, 우리는 ABCC11의 효과에 더 많은 주의를 기울일 필요가 있다. 이전 연구들은 이 다형성이 에스트로겐 수용체 양성 유방암에 영향을 미칠 수 있으며, T 대립유전자가 낮은 에스트로겐 유출 활성을 유발하여 유방암 위험을 증가시킬 수 있음을 시사했다 [69]. 진화 궤적을 통해 우리는 지리적으로 다른 인구의 환경에 대한 생물학적 적응성에 대한 통찰을 얻고 인구 이동 및 혼합 사건을 추론할 수 있다. 파생 대립유전자는 2천 년 전 주로 북부 인구에 분포했다 (그림 4a-c). 그 후 남부 인구에서 T 대립유전자 빈도가 크게 증가했는데, 이는 남쪽으로의 이주와 관련이 있을 수 있다.
The SLC10A1 gene involved in bile acid metabolism was also identified under natural selection in East Asians, but the detailed evolutionary processes and adaptive mechanisms involved remain unknown. The mutation is associated with glycocholic acid, low-density lipoprotein cholesterol, total cholesterol, and uric acid, reflecting SLC10A1 gene pleiotropism. We explored the evolutionary history of SLC10A1 from the perspective of allele frequency and a haplotype network. The mutation first appeared in Africans and spread northward and eastward. Arising in the Middle East and Southeast Asia between 2000 and 4000 years ago, the mutation frequency has gradually increased in Asians, with a prominent distribution in southern Chinese people.
담즙산 대사에 관여하는 SLC10A1 유전자 또한 동아시아인에게서 자연선택을 받은 것으로 확인되었지만, 관련된 상세한 진화 과정과 적응 메커니즘은 아직 알려지지 않았다. 이 돌연변이는 글리코콜산, 저밀도 지단백 콜레스테롤, 총 콜레스테롤, 요산과 관련이 있으며, 이는 SLC10A1 유전자의 다면발현성을 반영한다. 우리는 대립유전자 빈도와 일배체형 네트워크의 관점에서 SLC10A1의 진화 역사를 탐구했다. 돌연변이는 아프리카인에게서 처음 나타나 북쪽과 동쪽으로 퍼져나갔다. 2000년에서 4000년 전 사이에 중동과 동남아시아에서 발생한 이 돌연변이 빈도는 아시아인에게서 점차 증가했으며, 남부 중국인에게서 두드러진 분포를 보인다.
SLC10A1 encodes NTCP, and the mutation S267F can decrease the risk of cirrhosis and hepatocellular carcinoma and confer a protective effect against chronic hepatitis B. People carrying S267F exhibit significantly elevated bile acidemia during childhood. In China, we found that northern populations are less likely to suffer from NTCP deficiency disease than southern populations (Fig. 3i). This phenomenon is consistent with the lower prevalence of hepatitis B in northern populations [70]. This regional differentiation may have been shaped by admixture events between southern Chinese populations and Oceanians or by exposure to endemic pathogens (Fig. 4e and Supplementary Fig. 8m). Furthermore, our analysis revealed other natural selection signals associated with metabolism, including FADS genes involved in lipid metabolism, SLC35F₃ genes associated with vitamin metabolism, and ALDH2 genes implicated in alcohol metabolism. These metabolic differences may be due to differences in geographical environments, dietary habits, or exposure to pathogens. Throughout human evolutionary history, multiple factors have acted as driving forces of natural selection. When genotypes are mismatched with the modern environment, it might lead to the manifestation of human diseases. Dissecting the genetic basis of human adaptation in different backgrounds is vital for analyzing genetic diseases.
SLC10A1은 NTCP를 암호화하며, S267F 돌연변이는 간경변과 간세포암종의 위험을 감소시키고 만성 B형 간염에 대한 보호 효과를 부여할 수 있다. S267F를 가진 사람들은 어린 시절에 현저하게 상승된 담즙산혈증을 보인다. 중국에서는 북부 인구가 남부 인구보다 NTCP 결핍 질환을 겪을 가능성이 적다는 것을 발견했다 (그림 3i). 이 현상은 북부 인구에서 B형 간염 유병률이 낮은 것과 일치한다 [70]. 이러한 지역적 분화는 남부 중국 인구와 오세아니아인 간의 혼합 사건이나 풍토성 병원체에 대한 노출에 의해 형성되었을 수 있다 (그림 4e 및 보충 그림 8m). 또한, 우리 분석은 지질 대사에 관여하는 FADS 유전자, 비타민 대사와 관련된 SLC35F₃ 유전자, 알코올 대사에 관여하는 ALDH2 유전자를 포함하여 대사와 관련된 다른 자연선택 신호를 밝혔다. 이러한 대사 차이는 지리적 환경, 식습관 또는 병원체 노출의 차이 때문일 수 있다. 인류 진화 역사 전반에 걸쳐 여러 요인이 자연선택의 원동력으로 작용했다. 유전자형이 현대 환경과 불일치할 때, 이는 인간 질병의 발현으로 이어질 수 있다. 다양한 배경에서 인간 적응의 유전적 기초를 분석하는 것은 유전 질환을 분석하는 데 매우 중요하다.
그림 4. ABCC11과 SLC10A1의 진화 궤적 및 일배체형 분석. a-c 약 10,000년 전, 약 2,000년 전, 그리고 현재 시점을 포함한 다양한 시공간적 맥락에서 전 세계 인구에 걸친 rs17822931-T의 대립유전자 빈도 분포. 10,000년 전부터 2,000년 전까지의 상세한 진화 궤적은 보충 그림 8 h-j에 제시됨. d-f 약 10,000년 전, 약 2,000년 전, 그리고 현재 시점을 포함한 다양한 시공간적 맥락에서 전 세계 인구에 걸친 rs2296651-A의 대립유전자 빈도 분포. 상세한 진화 궤적은 보충 그림 8 k-m에 제시됨. g 전 세계 인구의 게놈 데이터를 기반으로 한 ABCC11 유전자의 일배체형 빈도를 보여주는 누적 막대 그래프. h 동일한 데이터세트를 기반으로 한 SLC10A1 유전자의 일배체형 빈도를 보여주는 누적 막대 그래프. i ABCC11 유전자에 위치한 6개 SNP의 일배체형 네트워크 분석. j SLC10A1 유전자에 위치한 11개 SNP의 일배체형 네트워크 분석.
결론 Conclusions
Our study elucidated the genetic affinities of Han populations across various regions and examined their interactions with ethnolinguistically diverse neighbors. Genomic data from the SDH population were compared with those from other Han populations in the lower YRB, including Shanxi, Shaanxi, and Henan populations. This comparison revealed a greater degree of genetic continuity in Shandong populations from the late Bronze/Iron Age to the present. Furthermore, numerous natural selection signals were identified, clarifying the evolutionary trajectory of East Asian-specific traits related to axillary odor and bile acid metabolism. The ABCC11 gene variant rs17822931, which may be associated with environmental adaptations, was detected approximately 4,400 years ago. The earliest occurrence of the SLC10A1 gene variant (rs2296651) was identified approximately 10,000 years ago and is linked to NTCP deficiency and chronic hepatitis B. These findings provide compelling genetic evidence for understanding the origin and progression of East Asian-specific traits and diseases.
이 연구는 여러 지역 한족(漢族)의 유전적 관계와 주변 민족과의 상호작용을 밝혔다. 특히 황하(黃河) 하류의 산동 한족(漢族)은 인근의 다른 한족(漢族) 집단보다 후기 청동기·철기 시대부터 현재까지 유전적 연속성이 더 강하게 나타났다. 또한, 자연선택의 흔적을 분석하여 겨드랑이 냄새(ABCC11 유전자)와 담즙산 대사(SLC10A1 유전자)와 같은 동아시아 특이 형질의 진화 과정을 규명했다. 환경 적응과 관련된 ABCC11 유전자 변이는 약 4,400년 전에, B형 간염 저항성과 관련된 SLC10A1 유전자 변이는 약 10,000년 전에 처음 나타난 것으로 추정된다. 이러한 결과는 동아시아인 특유의 형질과 질병이 어떻게 시작되고 발전했는지를 이해하는 데 중요한 유전적 증거를 제공한다.
연구 방법 Methods
Sample collection and DNA preparation
표본 수집 및 DNA 준비
We collected saliva samples from 264 unrelated healthy Han individuals in Shandong Province, North China. A QIAamp DNA Mini Kit (QIAGEN, Germany) was used to extract and purify the DNA. Subsequently, a quantitative analysis was conducted using the Qubit dsDNA HS Assay Kit from Thermo Fisher Scientific on a Qubit 3.0 fluorometer following the protocols provided by the manufacturer. The Medical Ethics Committees of West China Hospital of Sichuan University reviewed and approved the project and corresponding protocols. The individuals participating in this study were randomly selected, and informed consent was obtained from all participants. All included individuals were required to be indigenous residents with at least three generations at the sampling sites. All procedures were performed following the Helsinki Declaration of 2013 [71].
우리는 북중국 산동성(山東省)에서 혈연관계가 없는 건강한 한족(漢族) 264명으로부터 타액 샘플을 수집했다. DNA 추출 및 정제에는 QIAamp DNA Mini Kit(QIAGEN, Germany)를 사용했다. 이후, 제조사의 프로토콜에 따라 Qubit 3.0 형광계를 사용하여 Thermo Fisher Scientific의 Qubit dsDNA HS Assay Kit로 정량 분석을 수행했다. 사천대학(四川大學) 서중병원(西中病院)의 의학 윤리 위원회가 이 프로젝트와 관련 프로토콜을 검토하고 승인했다. 이 연구에 참여한 개인들은 무작위로 선정되었으며, 모든 참가자로부터 사전 동의를 받았다. 포함된 모든 개인은 채취 장소에서 최소 3세대 이상 거주한 토착민이어야 했다. 모든 절차는 2013년 헬싱키 선언(Helsinki Declaration)에 따라 수행되었다 [71].
Quality control, genotype calling, and dataset merging
품질 관리, 유전자형 판독, 데이터세트 병합
The Affymetrix Array was used for genotyping 264 individuals first reported here. We used PLINK v.1.90 [72] and King [73] to assess genetic relatedness between all pairs of samples. Individuals whose relatives were present within the three generations were removed. The following parameters were used to filter SNPs and samples (mind: 0.05, geno: 0.05, and HWE: 10−6). Approximately 465 K SNPs passed the first filtering step. We merged the database with publicly available and previously published data generated via an Affymetrix chip. The detailed populations and their classification information from the Affymetrix dataset are shown in Supplementary Table 17. Our data (including 2808 individuals from different language families) were merged with HGDP [74] and Oceania genomic resources [75] to form a global high-density dataset (424,501 SNPs). Then, variants from the Affymetrix dataset and the Allen Ancient DNA Resource (HO dataset and 1240 K dataset, https://reich.hms.harvard.edu/datasets) were merged to generate the middle-density dataset (including 359,009 SNPs) and low-density dataset (including 119,114 SNPs). Finally, three merged datasets (Affy_HGDP, Affy_1240K, and Affy_HO) were generated.
이 연구에서 처음 보고되는 264명의 유전자형 분석에는 애피메트릭스(Affymetrix) 어레이가 사용되었다. 모든 샘플 쌍 간의 유전적 관련성을 평가하기 위해 PLINK v.1.90 [72]과 King [73]을 사용했다. 3세대 내에 친척이 있는 개인은 제외되었다. SNP와 샘플을 필터링하는 데 다음 매개변수가 사용되었다(mind: 0.05, geno: 0.05, HWE: 10−6). 약 46만 5천 개의 SNP가 첫 번째 필터링 단계를 통과했다. 우리는 데이터베이스를 애피메트릭스 칩을 통해 생성된 공개 및 이전에 발표된 데이터와 병합했다. 애피메트릭스 데이터세트의 상세한 인구 및 분류 정보는 보충 표 17에 나와 있다. 우리 데이터(다른 어족의 2808명 포함)는 HGDP [74] 및 오세아니아 게놈 자원 [75]과 병합되어 전 세계 고밀도 데이터세트(424,501개 SNP)를 형성했다. 그런 다음, 애피메트릭스 데이터세트와 앨런 고대 DNA 자원(HO 데이터세트 및 1240K 데이터세트, https://reich.hms.harvard.edu/datasets)의의) 변이를 병합하여 중간 밀도 데이터세트(359,009개 SNP)와 저밀도 데이터세트(119,114개 SNP)를 생성했다. 최종적으로, 세 개의 병합된 데이터세트(Affy_HGDP, Affy_1240K, Affy_HO)가 생성되었다.
Principal component analysis
주성분 분석(PCA)
PCA was conducted using the smartpca package in EIGENSOFT, which focuses on modern and ancient populations at the East Asian and Chinese levels. The analysis used the following parameters: numoutlieriter: 0 and lsqproject: YES. East Asian-scale PCA was conducted to explore the population structure between the newly studied population and ancient/modern East Asian populations based on the HO_Affy database. To further dissect the fine-scale genetic structure of populations from the lower YRB, we also performed Chinese-scale PCA to dissect the genetic relationship between the target population and ANEA, in which modern people were projected onto the context of ancient individuals.
주성분 분석(PCA)은 EIGENSOFT의 smartpca 패키지를 사용하여 수행했으며, 동아시아 및 중국 단위의 현대 및 고대 인구에 초점을 맞췄다. 분석에는 numoutlieriter: 0과 lsqproject: YES라는 매개변수가 사용되었다. 동아시아 규모 PCA는 새로 연구된 집단과 다른 고대 및 현대 동아시아 집단 간의 인구 구조를 탐색하기 위해 수행되었다. 황하(黃河) 하류 지역 인구의 더 정밀한 유전 구조를 분석하기 위해, 현대인을 고대인의 유전 정보에 투영하여 비교하는 중국 규모 PCA도 추가로 진행했다.
FST calculation and TreeMix
FST 계산 및 TreeMix 분석
To quantify the genetic distance precisely, we used PLINK v.1.90 to calculate the pairwise fixation index (FST) between the target and reference populations. We also applied TreeMix v.1.1365 based on allele frequencies to explore their phylogenetic relationships. The number of migration edges was set from 0 to 7. A maximum likelihood tree was constructed by replication for each migration number.
집단 간 유전적 거리를 정량화하기 위해 PLINK를 사용하여 쌍별 고정 지수(FST)를 계산했다. 또한 대립유전자 빈도를 기반으로 TreeMix 소프트웨어를 적용하여 집단 간 계통 관계와 유전자 이동을 탐색했다. 유전자 이동 경로의 수는 0개부터 7개까지 설정하여 분석했다.
Haplotype-based population analysis
일배체형(Haplotype) 기반 인구 분석
We implemented SHAPEIT v.2.0 [76] to estimate the haplotypes based on the HGDP_Affy database and applied the recommended genetic map [77]. We used the default parameters (-burn 10-prune 10-main 30) [76]. Subsequently, we detected the pairwise shared IBD segments with Refined-IBD software and calculated the IBD matrix among pairwise populations based on the different IBD lengths. The observed IBD blocks with different lengths indicate differentiated genetic interactions occurring at different time horizons. To determine the recent demographic history of people with SDH, we estimated the effective population size within 150 generations using IBDNe v23Apr20 [78]. We used ChromoPainter and fineSTRUCTURE v4 [79] to estimate the fine-scale population structure among Han Chinese people and geographical neighbors based on the coancestry matrix, which revealed ancestral relationships at the individual level. We randomly selected certain individuals from 24 populations to assess phylogenetic relationships using fineSTRUCTURE, ChromoCombine, and ChromoPainter [80] with the following parameters: -s3iters 100,000, -s4iters 50,000, -s1 minsnps 1000, and slindfrac 0.1. We also calculated the ROH in the SDH population and ethnolinguistically different populations. Additionally, we used PLINK v.1.90 to classify ROH with lengths of <1, 1–5, and > 5.
우리는 HGDP_Affy 데이터베이스를 기반으로 일배체형을 추정하기 위해 SHAPEIT v.2.0 [76]을 실행하고, 권장되는 유전자 지도 [77]를 적용했다. 기본 매개변수(-burn 10-prune 10-main 30)를 사용했다 [76]. 이어서 Refined-IBD 소프트웨어로 쌍별 공유 동일 조상 유래(IBD) 구간을 감지하고, 다른 IBD 길이를 기반으로 쌍별 인구 간의 IBD 행렬을 계산했다. 다른 길이로 관찰된 IBD 블록은 다른 시간대에 발생한 차별화된 유전적 상호작용을 나타낸다. 산동 한족(山東漢族)의 최근 인구 역사를 결정하기 위해, IBDNe v23Apr20 [78]을 사용하여 150세대 내의 유효 집단 크기를 추정했다. 우리는 공통 조상 행렬을 기반으로 한족(漢族)과 지리적 이웃 간의 미세 규모 인구 구조를 추정하기 위해 ChromoPainter와 fineSTRUCTURE v4 [79]를 사용했으며, 이는 개인 수준에서 조상 관계를 밝혔다. 우리는 fineSTRUCTURE, ChromoCombine, ChromoPainter [80]를 사용하여 계통 발생 관계를 평가하기 위해 24개 인구에서 특정 개인을 무작위로 선택했으며, 다음 매개변수를 사용했다: -s3iters 100,000, -s4iters 50,000, -s1 minsnps 1000, slindfrac 0.1. 우리는 또한 산동 한족(山東漢族) 인구와 민족언어학적으로 다른 인구들에서 동형접합성 구간(ROH)을 계산했다. 추가적으로, 우리는 길이가 1 미만, 1-5, 5 초과인 동형접합성 구간을 분류하기 위해 PLINK v.1.90을 사용했다.
Admixture analysis
유전적 혼합 분석
We used unsupervised model-based ADMIXURE analyses to estimate the ancestry proportion of each individual. The genomic data were LD-pruned using PLINK v.1.9 with these parameters (-indep-pairwise 200 25 0.4). We assumed that the number of predefined ancestral sources ranged from 2 to 20. The models with K=6 and 3 and the lowest cross-validation error were chosen as the best-fit models among our major admixture models.
ADMIXTURE 소프트웨어를 사용하여 각 개인이 어떤 조상 집단으로부터 유전자를 물려받았는지 그 비율을 추정했다. 분석 전, PLINK를 이용해 유전체 데이터에서 연관성이 높은 유전자들을 제거하는 작업을 수행했다. 조상 집단의 수를 2개부터 20개까지 설정하여 분석했으며, 교차 검증 오류가 가장 낮은 K=6과 K=3 모델을 최적의 모델로 선택했다.
F-statistics
F-통계량 분석
We calculated admixture f₃ statistics to test potential admixture signals of SDH using qp3Pop in ADMIXTOOLS [81]. If f₃ values in the form of f₃(Reference, Reference; SDH) were negative with Z<−3, it indicated that two reference populations were predefined ancestral surrogates of SDH. In addition, f₄ statistics were calculated using qpDstat implemented in ADMIXTOOLS. We then used different forms of f₄ statistics to explore genetic affinities and differentiated gene flow events between the target and other ancient/modern reference populations.
우리는 ADMIXTOOLS [81]의 qp3Pop을 사용하여 산동 한족(山東漢族)의 잠재적 혼합 신호를 테스트하기 위해 혼합 f₃ 통계량을 계산했다. f₃(참조, 참조; 산동 한족(山東漢族)) 형태의 f₃ 값이 Z<-3으로 음수이면, 두 참조 집단이 산동 한족(山東漢族)의 미리 정의된 조상 대리 집단임을 나타냈다. 또한, ADMIXTOOLS에 구현된 qpDstat를 사용하여 f₄ 통계량을 계산했다. 그런 다음, 대상 집단과 다른 고대/현대 참조 집단 간의 유전적 친화성 및 차별화된 유전자 흐름 사건을 탐색하기 위해 다양한 형태의 f₄ 통계량을 사용했다.
QpAdm and qpWave
QpAdm 및 qpWave 분석
We modeled the ancestry admixture composition of representative ancient ancestral sources of SDH using two-way qpAdm modeling. We used the default parameter settings: allsnps: YES; details: YES. Moreover, a series of principles were used to screen the best-fitting qpAdm [82] models. First, the ancestry portion was greater than 0 and less than 1. Second, the standard error is smaller than the estimated minimum values of ancestry admixture proportion. Third, each p-value of the tested admixture model was greater than 0.05. We used Mbuti, Iran_GanjDareh_N, Italy_North_Villabruna_HG, Ami, Mixe, Onge, Papuan, China_Tianyuan, Ust_Ishim and Australia as outgroups. Two-way qpAdm admixture models can be used to successfully elucidate the gene pool of SDH. Using the same outgroup sets, we also implemented the qpWave package in ADMIXTOOLS to test genetic homogeneity among Han Chinese and other ancient and modern populations. The p values of pairwise populations are presented in the form of a heatmap.
우리는 양방향 qpAdm 모델링을 사용하여 산동 한족(山東漢族)의 대표적인 고대 조상 공급원의 조상 혼합 구성을 모델링했다. 기본 매개변수 설정(allsnps: YES; details: YES)을 사용했다. 또한, 최적의 qpAdm [82] 모델을 선별하기 위해 일련의 원칙이 사용되었다. 첫째, 조상 부분은 0보다 크고 1보다 작아야 한다. 둘째, 표준 오차는 추정된 조상 혼합 비율의 최소값보다 작아야 한다. 셋째, 테스트된 혼합 모델의 각 p-값은 0.05보다 커야 한다. 우리는 음부티족(Mbuti), 이란 간지다레(Iran_GanjDareh_N), 이탈리아 빌라브루나(Italy_North_Villabruna_HG), 아미족(Ami), 미헤족(Mixe), 옹게족(Onge), 파푸아인(Papuan), 중국 전문(中國田園), 우스트-이심(Ust_Ishim), 호주(Australia)를 외집단으로 사용했다. 양방향 qpAdm 혼합 모델은 산동 한족(山東漢族)의 유전자 풀을 성공적으로 설명하는 데 사용될 수 있다. 동일한 외집단 세트를 사용하여, 우리는 한족(漢族)과 다른 고대 및 현대 인구 간의 유전적 동질성을 테스트하기 위해 ADMIXTOOLS의 qpWave 패키지도 구현했다. 쌍별 인구의 p-값은 히트맵 형태로 제시된다.
ALDER
ALDER 분석
Admixture times with different sources were estimated via admixture-induced linkage disequilibrium for evolutionary relationships implemented in ALDER v1.03 [83]. The parameters used were as follows: mindis, 0.005; jackknife, YES.
다른 출처와의 혼합 시기는 ALDER v1.03 [83]에 구현된 진화 관계를 위한 혼합 유도 연관 불균형을 통해 추정되었다. 사용된 매개변수는 다음과 같다: mindis, 0.005; jackknife, YES.
Identification of natural selection signals
자연선택 신호 식별
First, we estimated HDVs between northern and southern Han Chinese populations based on allele frequency (FST). Subsequently, we used PBS to explore natural selection signals in the target population. The formula is as follows: PBSA=(TAB+TAC−TBC)/2, T=−lg(1−FST) [84]. We defined A as the studied population, while B and C were the ingroup and outgroup, respectively. We conducted different scales of PBS focused on genes with the top 0.1% of PBS values. In the distant test models, we used Europeans as the outgroup and Hlai as the ingroup. We used Hlai as the outgroup and Han people from Hunan Province as the ingroup to test the population-specific selection signatures. We further validated biological adaptive variants by iHS and XP-EHH using selscan v1.2.0 [85], and HNL was selected as the reference population. The PBS values of the variation need to be in the top 0.1%, which can be verified by one of the FST and iHS/XP-EHH methods. Finally, we obtained a high-quality set of selected alleles and used VEP to annotate them. We annotated trait-related selection candidate variants in the GWAS catalog. We used Metascape to perform GO and KEGG enrichment analyses [86]. We constructed linkage disequilibrium plots using Haploview [87]. The haplotype networks were constructed by PopART [88] using the 11 SNPs in the SLC10A1 gene and 6 SNPs in the ABCC11 gene from Affy_HGDP.
먼저, 우리는 대립유전자 빈도(FST)를 기반으로 북부와 남부 한족(漢族) 인구 간의 고도로 분화된 변이(HDV)를 추정했다. 이어서, 대상 인구에서 자연선택 신호를 탐색하기 위해 PBS를 사용했다. 공식은 다음과 같다: PBSA=(TAB+TAC−TBC)/2, T=−lg(1−FST) [84]. 여기서 A는 연구 대상 집단, B와 C는 각각 내집단과 외집단으로 정의했다. 우리는 PBS 값 상위 0.1%에 해당하는 유전자에 초점을 맞춰 다양한 규모의 PBS를 수행했다. 원거리 테스트 모델에서는 유럽인을 외집단으로, 여족(黎族)을 내집단으로 사용했으며, 인구 특이적 선택 신호를 테스트하기 위해 여족(黎族)을 외집단으로, 호남성(湖南省) 한족(漢族)을 내집단으로 사용했다. 우리는 selscan v1.2.0 [85]을 사용하여 iHS와 XP-EHH로 생물학적 적응 변이를 추가로 검증했고, 경중(瓊中) 여족(黎族)이 참조 집단으로 선택되었다. 변이의 PBS 값은 상위 0.1%에 속해야 하며, 이는 FST 및 iHS/XP-EHH 방법 중 하나로 검증될 수 있다. 최종적으로, 우리는 고품질의 선택된 대립유전자 세트를 얻고 VEP를 사용하여 주석을 달았다. GWAS 카탈로그에서 형질 관련 선택 후보 변이에 주석을 달았고, Metascape를 사용하여 GO 및 KEGG 농축 분석을 수행했다 [86]. Haploview [87]를 사용하여 연관 불균형 그림을 구성했으며, 일배체형 네트워크는 PopART [88]를 사용하여 Affy_HGDP의 SLC10A1 유전자에 있는 11개 SNP와 ABCC11 유전자에 있는 6개 SNP를 사용하여 구성되었다.
Supplementary Information
보충 정보
The online version contains supplementary material available at https://link.springer.com/article/10.1186/s12864-024-10514-9?error=cookies_not_supported&code=766e9181-d4e5-4261-b31b-f6336289c80f.
온라인 버전에는 다음 링크에서 이용할 수 있는 보충 자료가 포함되어 있다: https://link.springer.com/article/10.1186/s12864-024-10514-9?error=cookies_not_supported&code=766e9181-d4e5-4261-b31b-f6336289c80f.
Supplementary Material 1.
보충 자료 1.
Supplementary Material 2.
보충 자료 2.
Acknowledgements
감사의 말
We thank Prof. Etienne Patin and Prof. Lluis Quintana-Murci from the Human Evolutionary Genetics Unit of Institute Pasteur for sharing the high-coverage genomes of 317 individuals from the Pacific region.
파스퇴르 연구소(Institute Pasteur) 인간 진화 유전학 부서의 에티엔 파탱(Etienne Patin) 교수와 루이스 퀸타나-무르시(Lluis Quintana-Murci) 교수에게 태평양 지역 317명의 고밀도 게놈 데이터를 공유해 준 것에 대해 감사한다.
Code availability
코드 이용 가능성
No custom code was used in this work.
이 연구에는 별도의 맞춤 코드가 사용되지 않았다.
Authors’ contributions
저자 기여
G.L.H., M.G.W, C.L., and Y.C. conceived and supervised the project. G.L.H, Y.L, and S.H.D. collected the samples. H.R.S. performed the extraction of the genomic DNA and coordinated the genome sequencing. H.R.S, M.G.W., X.P.L, S.H.D., Q.X.S, Y.T.S., Z.Y.W, Q.X.Y., Y.G.H., J.Z., J.Y.M, X.C.J., T.Y, L.T.L, Y.H.L, J.C, J.B.Y., and G.C. performed the population genetic analysis. H.R.S. and M.G.W. drafted the manuscript. G.L.H, M.G.W, C.L, and Y.C. revised the manuscript.
허광린(G.L.H.), 왕멍거(M.G.W), 류차오(C.L.), 차이옌(Y.C.)이 프로젝트를 구상하고 감독했다. 허광린(G.L.H), 류옌(Y.L), 돤수한(S.H.D.)이 샘플을 수집했다. 수하오란(H.R.S.)이 게놈 DNA 추출을 수행하고 게놈 시퀀싱을 조정했다. 수하오란(H.R.S), 왕멍거(M.G.W.), 리샹핑(X.P.L), 돤수한(S.H.D.) 외 다수의 저자들이 집단 유전학 분석을 수행했다. 수하오란(H.R.S.)과 왕멍거(M.G.W.)가 원고 초안을 작성했으며, 허광린(G.L.H), 왕멍거(M.G.W), 류차오(C.L), 차이옌(Y.C.)이 원고를 수정했다.
Funding
자금 지원
This study was supported by the National Natural Science Foundation of China (82202078), the Major Project of the National Social Science Foundation of China (23&ZD203), the Open Project of the Key Laboratory of Forensic Genetics of the Ministry of Public Security (2022FGKFKT05), the Center for Archaeological Science of Sichuan University (23SASA01), the 1.3.5 Project for Disciplines of Excellence, West China Hospital, Sichuan University (ZYJC20002), and the Sichuan Science and Technology Program (2024NSFSC1518).
이 연구는 중국 국가자연과학기금, 중국 국가사회과학기금 주요 프로젝트, 공안부 법의유전학 핵심연구소 공개 프로젝트, 사천대학(四川大學) 고고학과학센터, 사천대학(四川大學) 서중병원(西中病院)의 우수 학문 분야 프로젝트, 사천(四川) 과학기술 프로그램의 지원을 받았다.
Availability of data and materials
데이터 및 자료 이용 가능성
The raw data derived from human samples have been deposited in the Zenodo (https://zenodo.org/records/11549865) with accession number 11549865 and OMIX database (https://ngdc.cncb.ac.cn/omix/release/OMIXO05781) with the accession number OMIX005781. Reference genotype data for ancient and modern individuals were collected from the Allen Ancient DNA Resource (https://reich.hms.harvard.edu/allen-ancient-dna-resource-aadr-downloadable-genotypes-present-day-and-ancient-dna-data). The access and use of the data complied with the regulations of the People’s Republic of China on the administration of human genetic resources.
인간 샘플에서 파생된 원시 데이터는 접근 번호 11549865로 Zenodo(https://zenodo.org/records/11549865)에에), 접근 번호 OMIX005781로 OMIX 데이터베이스(https://ngdc.cncb.ac.cn/omix/release/OMIXO05781)에에) 기탁되었다. 고대 및 현대 개인의 참조 유전자형 데이터는 앨런 고대 DNA 자원(Allen Ancient DNA Resource, https://reich.hms.harvard.edu/allen-ancient-dna-resource-aadr-downloadable-genotypes-present-day-and-ancient-dna-data)에서에서) 수집되었다. 데이터 접근 및 사용은 중화인민공화국(People’s Republic of China)의 인간 유전 자원 관리에 관한 규정을 준수했다.
선언 Declarations
Ethics approval and consent to participate
윤리 승인 및 참여 동의
The Medical Ethics Committees of West China Hospital of Sichuan University approved this study. All participants provided informed consent, and the study was conducted according to the principles of the Helsinki Declaration.
사천대학(四川大學) 서중병원(西中病院)의 의학 윤리 위원회가 이 연구를 승인했다. 모든 참가자는 사전 동의를 제공했으며, 연구는 헬싱키 선언의 원칙에 따라 수행되었다.
Consent for publication
출판 동의
Not applicable.
해당 없음
Competing interests
이해 상충
The authors declare that they have no competing interests.
저자들은 이해 상충이 없음을 선언한다.
Author details
저자 소속
1Genetic and Prenatal Diagnosis Center, Affiliated Hospital of North Sichuan Medical College, Nanchong 637007, Sichuan, China. 2Institute of Rare Diseases, West China Hospital of Sichuan University, Sichuan University, Chengdu 610000, China. 3School of Laboratory Medicine, North Sichuan Medical College, Nanchong 637007, Sichuan, China. 4Center for Archaeological Science, Sichuan University, Chengdu 610000, China. 5Research Center for Genomic Medicine, North Sichuan Medical College, Nanchong 637100, China. 6School of Forensic Medicine, Kunming Medical University, Kunming 650500, China. 7Institute of Basic Medicine and Forensic Medicine, North Sichuan Medical College and Genetic and Prenatal Diagnosis Center, Affiliated Hospital of North Sichuan Medical College, Nanchong 637007, Sichuan, China. 8Department of Forensic Medicine, College of Basic Medicine, Chongqing Medical University, Chongqing 400331, China. 9West China School of Basic Science & Forensic Medicine, Sichuan University, Chengdu 610041, China. 10School of Forensic Medicine, Shanxi Medical University, Jinzhong 030001, China. 11Anti-Drug Technology Center of Guangdong Province, Guangzhou 510230, China. 12Hunan Key Laboratory of Bioinformatics, School of Computer Science and Engineering, Central South University, Changsha 410075, China.
1유전 및 산전 진단 센터, 북사천의과대학 부속 병원, 중국 사천성 남충시. 2희귀질환 연구소, 사천대학 서중병원, 중국 사천성 성도시. 3임상병리학과, 북사천의과대학, 중국 사천성 남충시. 4고고학과학센터, 사천대학, 중국 사천성 성도시. 5게놈의학 연구센터, 북사천의과대학, 중국 사천성 남충시. 6법의학과, 곤명의과대학, 중국 곤명시. 7기초의학 및 법의학 연구소, 북사천의과대학 및 유전 및 산전 진단 센터, 북사천의과대학 부속 병원, 중국 사천성 남충시. 8법의학과, 기초의학부, 중경의과대학, 중국 중경시. 9서중 기초과학 및 법의학 스쿨, 사천대학, 중국 사천성 성도시. 10법의학과, 산서의과대학, 중국 진중시. 11광동성 마약방지기술센터, 중국 광저우시. 12호남 생물정보학 핵심연구소, 컴퓨터과학공학부, 중남대학, 중국 장사시.
참고문헌 References
- Stoneking M, Delfin F. The human genetic history of East Asia: weaving a complex tapestry. Curr Biol. 2010;20(4):R188–93.
- He GL, Li YX, Zou X, Yeh HY, Tang RK, Wang PX, Bai JY, Yang XM, Wang Z, Guo JX, et al. Northern gene flow into southeastern East Asians inferred from genome-wide array genotyping. J Syst Evol. 2022;61(1):179–97.
- Lea AJ, Garcia A, Arevalo J, Ayroles JF, Buetow K, Cole SW, Eid Rodriguez D, Gutierrez M, Highland HM, Hooper PL, et al. Natural selection of immune and metabolic genes associated with health in two lowland Bolivian populations. Proc Natl Acad Sci. 2023;120(1):e2207544120.
- Koller D, Wendt FR, Pathak GA, De Lilla A, De Angelis F, Cabrera-Mendoza B, Tucci S, Polimanti R. Denisovan and Neanderthal archaic introgression differentially impacted the genetics of complex traits in modern populations. BMC Biol. 2022;20(1):249.
- Wang CC, Yeh HY, Popov AN, Zhang HQ, Matsumura H, Sirak K, Cheronet O, Kovalev A, Rohland N, Kim AM, et al. Genomic insights into the formation of human populations in East Asia. Nature. 2021;591(7850):413–9.
- Sun Y, Wang M, Sun Q, Liu Y, Duan S, Wang Z, et al. Distinguished biological adaptation architecture aggravated population differentiation of Tibeto-Burman-speaking people. J Genet Genomics. 2024;51(5):517–30.
- He G, Wang P, Chen J, Liu Y, Sun Y, Hu R, Duan S, Sun Q, Tang R, Yang J, et al. Differentiated genomic footprints suggest isolation and longdistance migration of Hmong-Mien populations. BMC Biol. 2024;22(1):18.
- Li X, Wang M, Su H, Duan S, Sun Y, Chen H, Wang Z, Sun Q, Yang Q, Chen J, et al. Evolutionary history and biological adaptation of Han Chinese people on the Mongolian Plateau. hlife 2024.
- He G, Wang Z, Guo J, Wang M, Zou X, Tang R, Liu J, Zhang H, Li Y, Hu R, et al. Inferring the population history of Tai-Kadai-speaking people and southernmost Han Chinese on Hainan Island by genome-wide array genotyping. Eur J Hum Genet. 2020;28(8):1111–23.
- He G, Wang J, Yang L, Duan S, Sun Q, Li Y, Wu J, Wu W, Wang Z, Liu Y, et al. Genome-wide allele and haplotype-sharing patterns suggested one unique Hmong-Mein-related lineage and biological adaptation history in Southwest China. Hum Genomics. 2023;17(1):3.
- He G, Wang M, Miao L, Chen J, Zhao J, Sun Q, Duan S, Wang Z, Xu X, Sun Y, et al. Multiple founding paternal lineages inferred from the newlydeveloped 639-plex Y-SNP panel suggested the complex admixture and migration history of Chinese people. Hum Genomics. 2023;17(1):29.
- Sun Q, Wang M, Lu T, Duan S, Liu Y, Chen J, Wang Z, Sun Y, Li X, Wang S, et al. Differentiated adaptative genetic architecture and language-related demographical history in South China inferred from 619 genomes from 56 populations. BMC Biol. 2024;22(1):55.
- Yang MA, Fan X, Sun B, Chen C, Lang J, Ko YC, Tsang CH, Chiu H, Wang T, Bao Q, et al. Ancient DNA indicates human population shifts and admixture in northern and southern China. Science. 2020;369(6501):282–8.
- Pan Y, Zhang C, Lu Y, Ning Z, Lu D, Gao Y, Zhao X, Yang Y, Guan Y, Mamatyusupu D, et al. Genomic diversity and post-admixture adaptation in the Uyghurs. Natl Sci Rev. 2022;9(3):nwab124.
- Yao H, Wang M, Zou X, Li Y, Yang X, Li A, Yeh HY, Wang P, Wang Z, Bai J, et al. New insights into the fine-scale history of western-eastern admixture of the northwestern Chinese population in the Hexi Corridor via genome-wide genetic legacy. Molecular genetics and genomics: MGG. 2021;296(3):631–51.
- Yang Z, Bai C, Pu Y, Kong Q, Guo Y, Ouzhuluobu, Gengdeng, Liu X, Zhao Q, Qiu Z, et al. Genetic adaptation of skin pigmentation in highland Tibetans. Proc Natl Acad Sci U S A. 2022;119(40):e2200421119.
- Kurki MI, Karjalainen J, Palta P, Sipilä TP, Kristiansson K, Donner KM, Reeve MP, Laivuori H, Aavikko M, Kaunisto MA, et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature. 2023;613(7944):508–18.
- GenomeAsia KC. The GenomeAsia 100K Project enables genetic discoveries across Asia. Nature. 2019;576(7785):106–11.
- Zhang P, Luo H, Li Y, Wang Y, Wang J, Zheng Y, Niu Y, Shi Y, Zhou H, Song T, et al. NyuWa Genome resource: a deep whole-genome sequencingbased variation profile and reference panel for the Chinese population. Cell Rep. 2021;37(7):110017.
- Cao Y, Li L, Xu M, Feng Z, Sun X, Lu J, Xu Y, Du P, Wang T, Hu R, et al. The ChinaMAP analytics of deep whole genome sequences in 10,588 individuals. Cell Res. 2020;30(9):717–31.
- Kerner G, Neehus AL, Philippot Q, Bohlen J, Rinchai D, Kerrouche N, et al. Genetic adaptation to pathogens and increased risk of inflammatory disorders in post-Neolithic Europe. Cell Genom. 2023;3(2):100248.
- He GL, Wang MG, Li YX, Zou X, Yeh HY, Tang RK, Yang XM, Wang Z, Guo JX, Luo T, et al. Fine-scale north-to-south genetic admixture profile in Shaanxi Han Chinese revealed by genome-wide demographic history reconstruction. J Syst Evol. 2021;60(4):955–72.
- Chen J, Zheng H, Bei JX, Sun L, Jia WH, Li T, Zhang F, Seielstad M, Zeng YX, Zhang X, et al. Genetic structure of the Han Chinese population revealed by genome-wide SNP variation. Am J Hum Genet. 2009;85(6):775–85.
- Yao YG, Kong QP, Bandelt HJ, Kivisild T, Zhang YP. Phylogeographic differentiation of mitochondrial DNA in Han Chinese. Am J Hum Genet. 2002;70(3):635–51.
- Wen B, Li H, Lu D, Song X, Zhang F, He Y, Li F, Gao Y, Mao X, Zhang L, et al. Genetic evidence supports demic diffusion of Han culture. Nature. 2004;431(7006):302–5.
- Cong PK, Bai WY, Li JC, Yang MY, Khederzadeh S, Gai SR, Li N, Liu YH, Yu SH, Zhao WW, et al. Genomic analyses of 10,376 individuals in the Westlake BioBank for Chinese (WBBC) pilot project. Nat Commun. 2022;13(1):2939.
- Zhang M, Yan S, Pan W, Jin L. Phylogenetic evidence for Sino-Tibetan origin in northern China in the Late Neolithic. Nature. 2019;569(7754):112–5.
- Pechenkina EA, Benfer RA, Zhijun W. Diet and health changes at the end of the Chinese neolithic the Yangshao/Longshan transition in Shaanxi province. Am J Phys Anthropol. 2002;117(1):15–36.
- Barton L, Newsome SD, Chen FH, Wang H, Guilderson TP, Bettinger RL. Agricultural origins and the isotopic identity of domestication in northern China. Proc Natl Acad Sci U S A. 2009;106(14):5523–8.
- Ning C, Li T, Wang K, Zhang F, Li T, Wu X, Gao S, Zhang Q, Zhang H, Hudson MJ, et al. Ancient genomes from northern China suggest links between subsistence changes and human migration. Nat Commun. 2020;11(1):2700.
- Zong Y, Chen Z, Innes JB, Chen C, Wang Z, Wang H. Fire and flood management of coastal swamp enabled first rice paddy cultivation in east China. Nature. 2007;449(7161):459–62.
- Wu X, Zhang C, Goldberg P, Cohen D, Pan Y, Arpin T, Bar-Yosef O. Early pottery at 20,000 years ago in Xianrendong Cave, China. Science (New York, N Y). 2012;336(6089):1696–700.
- Deng Z, Kuo S-C, Carson MT, Hung H-C. Early Austronesians Cultivated Rice and Millet Together: Tracing Taiwan’s First Neolithic Crops. Front Plant Sci. 2022;13:962073.
- Zhang J, Lu H, Gu W, Wu N, Zhou K, Hu Y, Xin Y, Wang C. Early mixed farming of millet and rice 7800 years ago in the Middle Yellow River region, China. PLoS One. 2012;7(12):e52146.
- He G, Li YX, Wang MG, Zou X, Yeh HY, Yang XM, Wang Z, Tang RK, Zhu SM, Guo JX, et al. Fine-scale genetic structure of Tujia and central Han Chinese revealing massive genetic admixture under language borrowing. J Syst Eval. 2021;59(1):1–20.
- Leipe C, Long T, Sergusheva EA, Wagner M, Tarasov PE. Discontinuous spread of millet agriculture in eastern Asia and prehistoric population dynamics. Sci adv. 2019;5(9):eaax6225.
- Dong Y, Li C, Luan F, Li Z, Li H, Cui Y, Zhou H, Malhi RS. Low Mitochondrial DNA Diversity in an Ancient Population from China: insight into Social Organization at the Fujia Site. Hum Biol. 2015;87(1):71–84.
- Liu J, Zeng W, Sun B, Mao X, Zhao Y, Wang F, Li Z, Luan F, Guo J, Zhu C, et al. Maternal genetic structure in ancient Shandong between 9500 and 1800 years ago. Sci Bull. 2021;66(11):1129–35.
- Martin A, Saathoff M, Kuhn F, Max H, Terstegen L, Natsch A. A functional ABCC11 allele is essential in the biochemical formation of human axillary odor. J Invest Dermatol. 2010;130(2):529–40.
- Ohashi J, Naka I, Tsuchiya N. The impact of natural selection on an ABCC11 SNP determining earwax type. Mol Biol Evol. 2011;28(1):849–57.
- Fujimoto A, Kimura R, Ohashi J, Omi K, Yuliwulandari R, Batubara L, Mustofa MS, Samakkarn U, Settheetham-Ishida W, Ishida T, et al. A scan for genetic determinants of human hair morphology: EDAR is associated with Asian hair thickness. Hum Mal Genet. 2008;17(6):835–43.
- Ma X, Xu S. Archaic introgression contributed to the pre-agriculture adaptation of vitamin B1 metabolism in East Asia. IScience. 2022;25(12):105614.
- Zhang X, Sun A, Ge J. Origin and Spread of the ALDH2 Glu504Lys Allele. Phenomics. 2021;1(5):222–8.
- Barkus C, Sanderson DJ, Rawlins JNP, Walton ME, Harrison PJ, Bannerman DM. What causes aberrant salience in schizophrenia? A role for impaired short-term habituation and the GRIA1 (GluA1) AMPA receptor subunit. Mol Psychiatry. 2014;19(10):1060–70.
- Bhanushali AA, Patra PK, Pradhan S, Khanka SS, Singh S, Das BR. Genetics of fetal hemoglobin in tribal Indian patients with sickle cell anemia. Transl Res. 2015;165(6):696–703.
- Esrick EB, Lehmann LE, Biffi A, Achebe M, Brendel C, Ciuculescu MF, Daley H, MacKinnon B, Morris E, Federico A, et al. Post-transcriptional genetic silencing of BCL11A to treat sickle cell disease. N Engl J Med. 2021;384(3):205–15.
- Dadheech S, Madhulatha D, Jainc S, Joseph J, Jyothy A, Munshi A. Association of BCL11A genetic variant (rs11886868) with severity in B-thalassaemia major & sickle cell anaemia. Indian J Med Res. 2016;143(4):449–54.
- Choudhury A, Aron S, Botigue LR, Sengupta D, Botha G, Bensellak T, Wells G, Kumuthini J, Shriner D, Fakim YJ, et al. High-depth African genomes inform human migration and health. Nature. 2020;586(7831):741–8.
- Terao C, Yoshifuji H, Matsumura T, Naruse TK, Ishii T, Nakaoka Y, Kirino Y, Matsuo K, Origuchi T, Shimizu M, et al. Genetic determinants and an epistasis of LILRA3 and HLA-B*52 in Takayasu arteritis. Proc Natl Acad Sci U S A. 2018;115(51):13045–50.
- Chen Z, Li J, Yang Y, Li H, Zhao J, Sun F, Li M, Tian X, Zeng X. The renal artery is involved in Chinese Takayasu’s arteritis patients. Kidney Int. 2018;93(1):245–51.
- Renauer P, Sawalha AH. The genetics of Takayasu arteritis. Presse Medicale (Paris, France: 1983). 2017;46(7–8 Pt 2):e179-e187.
- Raghubeer S, Matsha TE. Methylenetetrahydrofolate (MTHFR), the onecarbon cycle, and cardiovascular risks. Nutrients. 2021;13(12):4562.
- Chen Y, Wang Z, Jiang Y, Lin Y, Wang X, Wang Z, Tang Z, Wang Y, Wang J, Gao Y, et al. Biallelic p.V371 variant in GJB2 is associated with increasing incidence of hearing loss with age. Genet Med. 2022;24(4):915–23.
- Chen H, Lin R, Lu Y, Zhang R, Gao Y, He Y, et al. Tracing Bai-Yue ancestry in Aboriginal Li people on Hainan Island. Mol Biol Evol. 2022;39(10):msac210.
- Russell LE, Zhou Y, Lauschke VM, Kim RB. In Vitro Functional Characterization and in silico prediction of rare genetic variation in the bile acid and drug transporter, Na+ Taurocholate Cotransporting Polypeptide (NTCP, SLC10A1). Mol Pharm. 2020;17(4):1170–81.
- Deng L-J, Ouyang W-X, Liu R, Deng M, Qiu J-W, Yaqub M-R, Raza M-A, Lin W-X, Guo L, Li H, et al. Clinical characterization of NTCP deficiency in paediatric patients: a case-control study based on SLC10A1 genotyping analysis. Liver International Official Journal of the International Association For the Study of the Liver. 2021;41(11):2720–8.
- Hu H-H, Liu J, Lin Y-L, Luo W-S, Chu Y-J, Chang C-L, Jen C-L, Lee M-H, Lu S-N, Wang L-Y, et al. The rs2296651 (S267F) variant on NTCP (SLC10A1) is inversely associated with chronic hepatitis B and progression to cirrhosis and hepatocellular carcinoma in patients with chronic hepatitis B. Gut. 2016;65(9):1514–21.
- Rajoriya N, Feld JJ. One small SNP for receptor virus entry, one giant leap for hepatitis B? Gut. 2016;65(9):1395–7.
- Peng L, Zhao Q, Li Q, Li M, Li C, Xu T, Jing X, Zhu X, Wang Y, Li F, et al. The p.Ser267Phe variant in SLC10A1 is associated with resistance to chronic hepatitis B. Hepatology (Baltimore, Md). 2015;61(4):1251–60.
- Chen S, Li J, Wang D, Fung H, Wong L-Y, Zhao L. The hepatitis B epidemic in China should receive more attention. Lancet (London, England). 2018;391(10130):1572.
- Kumar P, Henikoff S, Ng PC. Predicting the effects of coding non-synonymous variants on protein function using the SIFT algorithm. Nat Protoc. 2009;4(7):1073–81.
- Adzhubei IA, Schmidt S, Peshkin L, Ramensky VE, Gerasimova A, Bork P, Kondrashov AS, Sunyaev SR. A method and server for predicting damaging missense mutations. Nat Methods. 2010;7(4):248–9.
- Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M. CADD: predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Res. 2019;47(D1):D886–94.
- He C, He H-Y, Sun C-F, Ojha SC, Wang H, Deng C-L, Sheng Y-J. The relationship between NTCP gene varieties and the progress of liver disease after HBV infection: an updated systematic review and meta-analysis. Am J Med Sci. 2022;364(2):207–19.
- Cheng S, Xu Z, Bian S, Chen X, Shi Y, Li Y, Duan Y, Liu Y, Lin J, Jiang Y, et al. The STROKE GENOMEICS genome study: deep whole-genome sequencing and analysis of 10K Chinese patients with ischemic stroke reveal complex genetic and phenotypic interplay. Cell Discovery. 2023;9(1):75.
- Luo J, Chen S, Wang J, Ou S, Zhang W, Liu Y, Qin Z, Xu J, Lu Q, Mo C, et al. Genetic polymorphisms in complement receptor 1 gene and its association with HBV-related liver disease: a case-control study. Gene. 2019;688:107–18.
- Li YC, Ye WJ, Jiang CG, Zeng Z, Tian JY, Yang LQ, Liu KJ, Kong QP. River valleys shaped the maternal genetic landscape of Han Chinese. Mol Biol Evol. 2019;36(8):1643–52.
- Du Z, Ma L, Qu H, Chen W, Zhang B, Lu X, Zhai W, Sheng X, Sun Y, Li W, et al. Whole Genome Analyses of Chinese Population and De Novo Assembly of A Northern Han Genome. Genomics Proteomics Bioinformatics. 2019;17(3):229–47.
- Ishiguro, Ito H, Tsukamoto M, Iwata H, Nakagawa H, Matsuo K. A functional single nucleotide polymorphism in ABCC11, rs17822931, is associated with the risk of breast cancer in Japanese. Carcinogenesis. 2019;40(4):537–43.
- Liu Z, Lin C, Mao X, Guo C, Suo C, Zhu D, Jiang W, Li Y, Fan J, Song C, et al. Changing prevalence of chronic hepatitis B virus infection in China between 1973 and 2021: a systematic literature review and meta-analysis of 3740 studies and 231 million people. Gut. 2023;72(12):2354–63.
- World Medical Association Declaration of Helsinki: ethical principles for medical research involving human subjects. JAMA. 2013;310(20):2191–4.
- Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience. 2015;4:7.
- Tinker NA, Mather DE. Kin – Software for Computing Kinship Coefficients. J Hered. 1993;84(3):238–238.
- Bergström A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, et al. Insights into human genetic variation and population history from 929 diverse genomes. Science. 2020;367(6484):eaay5012.
- Choin J, Mendoza-Revilla J, Arauna LR, Cuadros-Espinoza S, Cassar O, Larena M, Ko AM, Harmant C, Laurent R, Verdu P, et al. Genomic insights into population history and biological adaptation in Oceania. Nature. 2021;592(7855):583–9.
- Delaneau O, Marchini J, Zagury JF. A linear complexity phasing method for thousands of genomes. Nat Methods. 2011;9(2):179–81.
- Browning BL, Browning SR. Improving the accuracy and efficiency of identity-by-descent detection in population data. Genetics. 2013;194(2):459–71.
- Browning SR, Browning BL. Accurate Non-parametric Estimation of Recent Effective Population Size from Segments of Identity by Descent. Am J Hum Genet. 2015;97(3):404–18.
- Lawson DJ, Hellenthal G, Myers S, Falush D. Inference of population structure using dense haplotype data. PLoS Genet. 2012;8(1):e1002453.
- Hellenthal G, Busby GBJ, Band G, Wilson JF, Capelli C, Falush D, Myers S. A genetic atlas of human admixture history. Science. 2014;343(6172):747–51.
- Weir BS, Cockerham CC. Estimating F-statistics for the analysis of population structure. Evolution. 1984;38(6):1358–70.
- Harney E, Patterson N, Reich D, Wakeley J. Assessing the performance of qpAdm: a statistical tool for studying population admixture. Genetics. 2021;217(4):iyaa045.
- Loh PR, Lipson M, Patterson N, Moorjani P, Pickrell JK, Reich D, Berger B. Inferring admixture histories of human populations using linkage disequilibrium. Genetics. 2013;193(4):1233–54.
- Yi X, Liang Y, Huerta-Sanchez E, Jin X, Cuo ZX, Pool JE, Xu X, Jiang H, Vinckenbosch N, Korneliussen TS, et al. Sequencing of 50 human exomes reveals adaptation to high altitude. Science. 2010;329(5987):75–8.
- Szpiech ZA, Hernandez RD. selscan: an efficient multithreaded program to perform EHH-based scans for positive selection. Mol Biol Evol. 2014;31(10):2824–7.
- Zhou Y, Zhou B, Pache L, Chang M, Khodabakhshi AH, Tanaseichuk O, Benner C, Chanda SK. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat Commun. 2019;10(1):1523.
- Barrett JC, Fry B, Maller J, Daly MJ. Haploview: analysis and visualization of LD and haplotype maps. Bioinformatics (Oxford, England). 2005;21(2):263–5.
- Leigh JW, Bryant D. popart: full-feature software for haplotype network construction. Methods Ecol Eval. 2015;6(9):1110–6.

