1
Vol.:(0123456789)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s
Explo ing he biological ole
o pos zygo ic and ge minal de
no o mu a ions in ASD
A. Alonso‑Gonzalez1,2, M. Calaza1,2, J. Amigo3, J. González‑Peñas4, R. Ma ínez‑Reguei o1,2,
M. Fe nández‑P ie o1,2, M. Pa ellada4, C. A ango4, C is ina Rod iguez‑Fon enla2,5* &
A. Ca acedo1,2,3,5
De no o mu a ions (DNMs), including ge minal and pos zygo ic mu a ions (PZMs), a e a s ong sou ce
o causali y o Au ism Spec um Diso de (ASD). Howe e , he biological p ocesses in ol ed behind
hem emain unexplo ed. Ou aim was o de ec DNMs (ge minal and PZMs) in a Spanish ASD coho
(360 ios) and o explo e hei ole ac oss di e en biological hie a chies (gene, biological pa hway,
cell and b ain a eas) using bioin o ma ic app oaches. Fo he majo i y o he analysis, a combined
ASD coho (N = 2171 ios) was c ea ed using p e iously published da a by he Au ism Sequencing
Conso ium (ASC). New plausible candida e genes o ASD such as FMR1 and NFIA we e ound. In
addi ion, genes ha bo ing PZMs we e signi ican ly en iched o miR‑137 a ge s in compa ison wi h
ge minal DNMs ha we e en iched in GO e ms ela ed o synap ic ansmission. The exp ession
pa e n o genes wi h PZMs was es ic ed o ea ly mid‑ e al co ex. In con as , he analysis o genes
wi h ge minal DNMs e ealed a spa io‑ empo al window om ea ly o mid‑ e al de elopmen s ages,
wi h exp ession in he amygdala, ce ebellum, co ex and s ia um. These esul s p o ide e idence o
he pa hogenic ole o PZMs and sugges he exis ence o dis inc mechanisms be ween PZMs and
ge minal DNMs ha a e in luencing ASD isk.
Abb e ia ions
AAF Al e na e allele equency
ADI-R Au ism diagnos ic in e iew- e ised
ADOS Au ism diagnos ic obse a ion schedule
ASD Au ism spec um diso de
BEE B ain exp essed enhance s
BF Bayesian ac o
DNMs De no o mu a ions
EWCE Exp ession weigh ed cell ype en ichmen
ExAC Exome agg ega ion conso ium
FDR False disco e y a e
FSHD Facioscapulohume al muscula dys ophy
GO Gene on ology
GQ Geno ype quali y
LoF Loss o unc ion
NDD Neu ode elopmen al diso de
OMIM Online mendelian inhe i ance in man
OPEN
1G upo de Medicina Xenómica, Fundación Ins i u o de In es igación Sani a ia de San iago de Compos ela
(FIDIS), Uni e sidade de San iago de Compos ela, San iago de Compos ela, Spain. 2Genomics and
Bioin o ma ics G oup, Cen e o Resea ch in Molecula Medicine and Ch onic Diseases (CiMUS), Uni e sidade
de San iago de Compos ela, A Ba celona 31, 15706 San iago de Compos ela, Spain. 3Fundación Pública
Galega de Medicina Xenómica (FPGMX), Cen o de In es igación Biomédica en Red, En e medades Ra as
(CIBERER), Uni e sidad de San iago de Compos ela, San iago de Compos ela, Spain. 4Cen o De In es igación
Biomédica en Red de Salud Men al (CIBERSAM), Hospi al Gene al Uni e si a io G ego io Ma añón, Ins i u o
de In es igación Sani a ia G ego io Ma añón, IiSGM, School o Medicine, Uni e sidad Complu ense, Mad id,
Spain. 5
These au ho s con ibu ed equally: Ma ía C is ina Rod iguez-Fon enla and A. Ca acedo. *email:
[email p o ec ed]
2
Vol:.(1234567890)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s/
pSI Speci i y index s a is ic
PZMs Pos zygo ic mu a ions
TADA T ansmission and de no o associa ion es
VCF Va ian call o ma
WES Whole exome sequencing
Backg ound
Au ism Spec um Diso de (ASD) is a neu ode elopmen al diso de (NDD) cha ac e ized by de ici s in com-
munica ion and social in e ac ion oge he wi h es ic ed in e es s and epe i i e beha io s1. ASD p e alence
among child en in he Uni ed S a es s ands a a ound 1.5% and has apidly isen in ecen yea s. In addi ion o he
co e symp oms o ASD, o he condi ions such as epilepsy o in ellec ual disabili y a e o en p esen . Como bidi y
is a cha ac e is ic o ASD ha can appea a any ime du ing child’s de elopmen . Since many ASD cases wi h
como bidi y ha e a clea gene ic backg ound and ea ly de ec ion is key o in e en ion, he gene ic diagnosis
in his ype o cases is a challenge2.
Twin and amily s udies ha e es ima ed ASD he i abili y o be abou 80% and subsequen gene ic s udies
ha e demons a ed ha he la ges pa o his he i abili y (50%) is explained by common a ia ion3,4. Howe e ,
de no o a e gene ic a ia ion (mino allele equency < 0.1%), including small inse ions and dele ions (indels),
copy numbe a ian s and single nucleo ide a ian s con e s highe indi idual isk5–7. Ge minal de no o mu a-
ions (DNMs) occu wi hin ge m cells and hey a e ansmi ed o he o sp ing when he zygo e is o med a e
e iliza ion. Thus, e e y single cell line o he esul ing emb yo will ca y an iden ical gene ic load. Ano he
ype o DNMs, pos zygo ic mu a ions (PZMs), a ise du ing zygo e mi osis, leading o a mosaic o gene ically
di e en cell lines8. The equency o mu agenesis and he gene a ion o PZMs is inc eased p io o gas ula ion
and neu ogenesis9.
PZMs in ol ed in ASD pa hogenesis a e usually de ec able h ough deep sequencing o b ain issues. How-
e e , his echnique o en en ails a huge challenge due o he inabili y o ob ain ASD b ain samples10. In con as ,
nex gene a ion sequencing echnologies can be used o de ec mosaic mu a ions in pe iphe al blood o a ec ed
indi iduals by inc easing he dep h o co e age11. Thus, i is possible o ob ain enough sequencing eads con-
aining he e e ence and he al e na e allele o accu a ely calcula e he al e na e allele equency (AAF)12. In
PZMs, he AAF alue shi s om he expec ed 50/50 a io o he e ozygous ge minal mu a ions. High co e age
whole exome sequencing (WES) (dep h > 200×) p o ides enough sensi i i y o de ec PZMs p esen ing AAF
alues as lowe as 15%13,14. I is wo hy o no e ha mos WES s udies ha e missed PZMs due o he commonly
employed pipelines. The de elopmen o new a ian calling pipelines is he e o e needed and some e o s ha e
been done in his ega d15–18.
I has been es ima ed ha 7.5% o DNMs a e PZMs ha con ibu e abou 4% o he o e all a chi ec u e o
ASD. PZMs ha e been iden i ied in high-con idence ASD isk genes. O he no el ASD candida e genes such as
KLF16 and MSANTD2, we e disco e ed a e s udying he con ibu ion o PZMs o ASD isk in la ge collec ions
o ASD p obands18. This poin s o he ac ha some genes ca y a la ge numbe o mu a ions in a mosaic s a e
han o he genes. In addi ion, a de ailed analysis o non-synonymous PZMs has e ealed ha hese a ian s a e
mainly ound in b ain-exp essed genes and in Loss-o - unc ion (LoF)-cons ained exons. The spa io- empo al
analysis ac oss di e en de elopmen al s ages also poin s o b ain a eas, like he amygdala, ha ha e no been
p e iously highligh ed by o he WES s udies in which PZMs we e no conside ed16,18.
The ele ance o PZMs in he pa hogenesis o ASD and he biological p ocesses in which genes ca ying PZMs
a e in ol ed, emain la gely unexplo ed. Mo eo e , he con ibu ion o PZMs o he pheno ypic p esen a ion is
ano he subjec ha should be s udied in mo e de ail using la ge-scale s udies. Hence, i is suspec ed ha ASD
p obands ca ying mosaic mu a ions migh be less a ec ed han p obands ca ying ge minal mu a ions as i
happens in o he NDDs such as P o eus synd ome o se e al b ain mal o ma ions19,20. The e o e, he main aim
o his s udy was o accu a ely de ec DNMs (ge minal and PZMs) in a coho o Spanish ios wi h ASD (360).
The no el DNMs de ec ed in he Spanish coho we e combined wi h a lis o DNMs p e iously published by he
Au ism Sequencing Conso ium (ASC)18 in a coho o 5947 amilies (4032 ASD ios and 1918 quads) in o de
o s udy i di e en ASD isk genes end o accumula e one o ano he ype o mu a ions using di e en bioin o -
ma ic app oaches. In addi ion, he di e en biological implica ions o ge minal and PZMs in ASD we e explo ed
h ough en ichmen analysis app oaches, which ha e no been applied be o e o his class o mu a ions ac oss
di e en hie a chical le els (gene, GO e ms, neu onal cell ypes and b ain a eas) (Addi ional ile4. Fig.S1).
Me hods
Subjec s. DNA was ex ac ed om pe iphe al blood o he Spanish ASD samples (360 ios; una ec ed
pa en s and a ec ed p oband) using he Gen aPu egene blood ki (Qiagen Inc., Valencia, CA, USA). Subjec s
om San iago (N = 136) we e ec ui ed om Complexo Hospi ala io Uni e si a io de San iago de Compos ela
and Galician ASD o ganiza ions. Subjec s om Mad id (N = 224) we e ec ui ed as pa o AMITEA p og am
a he Child and Adolescen Depa men o Psychia y, Hospi al Gene al Uni e si a io G ego io Ma añón.
Only indi iduals 3yea s old o olde we e included. All pa icipan s had a clinical diagnosis o ASD made by
ained pedia ic neu ologis s o psychia is s based on he Diagnos ic and S a is ical Manual o Men al Diso -
de s, Fou h Edi ion Tex Re ision and Fi h Edi ion (DSM-IV-TR and DSM-5) c i e ia. The Au ism Diagnos ic
Obse a ion Schedule (ADOS) and he Au ism Diagnos ic In e iew-Re ised (ADI-R) we e also adminis e ed
when necessa y. In o med consen signed by each pa icipa ing subjec o legal gua dian and app o al om he
co esponding Resea ch E hics Commi ee we e ob ained be o e he s a o he s udy. All pa icipan s, pa en s
3
Vol.:(0123456789)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s/
o legal ep esen a i es p o ided w i en in o med consen a en ollmen and he s udy was conduc ed acco ding
o he decla a ion o Helsinki.
Sample quali y con ol and DNMs de ec ion. Da a p ocessing and anno a ion. WES o DNA ex-
ac ed om he Spanish 360 ios was pe o med by he ASC (h ps ://genom e.emo y .edu/ASC/)21. One mul i-
sample VCF wi h he aw esul s was e ie ed om he ASC. Indi idual iles con aining coding a ian s pe
indi idual we e ob ained using bc ools and we e anno a ed using SnpE (Genomic a ian anno a ions and
unc ional e ec p edic ion oolbox) e sion 4.3T (h p://snpe .sou c e o g e.ne /).
Sample speci ic quali y con ol. To check amily ela ionships in he Spanish coho (360 ios), in o ma ion o
Mendelian e o coun s was ob ained using he "–mendel” op ion a ailable in VCF ools (h p:// c o ols.sou c
e o g e.ne /). Samples whose Mendelian e o s signi ican ly de ia ed om he expec a ion we e no conside ed
o subsequen analysis.
To iden i y disc epancy be ween nominal designed and gene ically de e mined sex he “–sexcheck” op ion
in PLINK was used o in e co ec sex om geno ypes on ch omosome X and Y.
Finally, o iden i y ou lie samples in he Spanish coho (360 ios), he “pseq i-s a s” command in PLINK
was used. Samples in which any o he ollowing pa ame e s: coun o al e na e, mino , he e ozygous geno ypes,
numbe o called a ian s o geno yping a e de ia ed mo e han 4 SD om he mean we e elimina ed. The e o e,
he whole io was d opped i any membe was conside ed an ou lie .
The samples o he Spanish coho (360 ios) ha passed all he quali y con ols men ioned abo e we e he
same as hose included in Sa e s om e al.21.
DNMs de ec ion. To de ec DNMs in he Spanish coho (360 ios), de ined as hose mu a ions ha a e s ic ly
p esen in p obands and no in pa en s, he il e ing op ions published by Lim e al. we e employed18. In his
s udy, a ian s classi ied as PZMs we e esequenced by h ee di e en sequencing echnologies eaching a high
alida ion a e (87–97%).
B ie ly, we de ine DNMs as hose a ian s whose geno ypes we e 1/0 o 1/1 in p obands and 0/0 in pa en s.
Then, a ian s wi h GQ ≥ 20 and al e na e ead dep h ≥ 7 we e conside ed. Va ian s ha p esen wo o mo e
alleles in he ExAC da abase (h p://exac.b oad ins i u e.o g/) we e il e ed ou . In ame indels we e also il e ed
and only biallelic DNMs we e conside ed. In addi ion, we il e ed ou a ian s ha we e less han 20 base pai s
apa om each o he o educe alse posi i es, and a ian s whose RVIS (Residual Va ia ion In ole ance Sco e)
e ie ed om ExAC was highe han 75% we e also il e ed ou . RVIS is designed o ank genes in e ms o
whe he hey ha e mo e o less common unc ional gene ic a ia ion ela i e o he genome-wide expec a ion
gi en he amoun o appa en ly neu al a ia ion he gene has. In ole an genes a e mo e likely o be be e candi-
da es in NDDs. Thus, RVIS alues ep esen ed as pe cen iles e lec he ela i e ank o he genes, wi h hose genes
abo e 75 h pe cen ile being he mos ole an and he e o e less likely o ha bo mu a ions wi h a ole in ASD.
SnpE was employed o classi y exonic a ian s acco ding o he de ini ion o hei p edic i e impac : high,
mode a e and low impac on he canonical ansc ip . Low impac a ian s included silen mu a ions, mode a e
impac a ian s included missense mu a ions and high impac included splicing and nonsense mu a ions. Only
base subs i u ions we e conside ed so ameshi a ian s we e il e ed ou .
Two di e en in silico p edic ion ools (CADD and SIFT) we e used o classi y missense mu a ions. P obably
damaging mu a ions we e hose p edic ed as damaging by SIFT and a ian s wi h CADD sco e > 20 (Addi ional
ile1; TableS1).
Finally, DNMs we e classi ied as ge minal o PZMs based on he AAF (numbe o al e na e eads/( o al
numbe o e e ence + al e na e eads)). DNMs wi h an AAF ≥ 0.40 we e classi ied as ge minal and DNMs wi h
an AAF < 0.40 we e classi ied as PZMs18.
90 samples om he Spanish coho (360 ios) we e al eady analyzed by Lim e al.18 and hey we e used
as posi i e con ols o check i he de ec ion o PZMs in he Spanish coho was accu a ely made. The e o e, i
was p o ed ha mos o DNMs we e accu a ely de ec ed and classi ied as ge minal o PZMs (Addi ional ile1;
TablesS1, S2).
Fo he majo i y o he analysis, we used a da ase called “combined coho ” (N = 2171) ha includes he non-
synonymous DNMs de ec ed in he 360 Spanish ios plus he non-synonymous DNMs iden i ied in indi iduals
wi h ASD sequenced by he ASC and published p e iously. Duplica ed a ian s in bo h coho s we e elimina ed
(Supplemen a y Table3 o Lim e al.)18 (Addi ional ile1; TableS1 and S3). Fo some analysis, we also de ined a
con ol coho o heal hy siblings published by he ASC (same sequencing dep h and a ian calling p ocedu es
han he p obands o he Spanish coho ) (N = 288)1,18 (Addi ional ile1; TableS4).
T ansmission and de no o associa ion es (TADA‑deno o). TADA-Deno o (h p://www.compg
en.pi .edu/TADA/TADA_guide .h ml# ada-analy sis-o -de-no o-da a- ada-deno o) was un o disco e and o
p io i ize ASD isk genes o bo h DNMs (ge minal and PZMs) in he Spanish coho (N = 360) (Addi ional
ile1; TableS1) and in he combined da ase (N = 2171) (Addi ional ile1; TableS3). TADA akes in o accoun
he mu a ional bu den o he genes as well as he mul iple mu a ional classes22. TADA was independen ly un in
wo di e en gene-se s o he Spanish coho (genes ha bo ing PZMs (PZMs genes) in he Spanish coho = 105;
genes ha bo ing ge minal DNMs (Ge minal genes) in he Spanish coho = 181) and o he combined coho
(PZMs genes in he combined coho = 362; ge minal genes in he combined coho = 1210) (Addi ional ile2;
TablesS5, S6, S7 and S8). The con ol coho (N = 288) was employed o se up and o es ima e he pa ame e s
needed by TADA18 (Addi ional ile1; TableS4).
4
Vol:.(1234567890)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s/
Two classes o DNMs we e included in he analysis: LoF and p obably damaging missense mu a ions. To
se up mu a ional a es o each mu a ional ca ego y, we used he pe gene mu a ion a es able da a compu ed
by Samocha e al.23 and hen, he ollowing o mula was applied o calib a e hem: LoF; (nonsense + splice) x
(synobs/synexp) and p obably damaging missense; missense × (NP ob.damaging /Nallmissense) x (synobs/synexp). Synobs
is he obse ed numbe o synonymous mu a ions in he con ol coho o una ec ed siblings (N = 119)18, and
synexp is he expec ed numbe o synonymous DNMs in he same coho calcula ed om he sum o pe -gene
synonymous DNMs a es (2*n*µ) (N = 79.04). Np ob.damaging is he numbe o p obably damaging missense mu a-
ions in he con ol coho (N = 212) and N allmissense we e he o al o missense mu a ions in he con ol coho
(N = 296). To es ima e he ela i e isk (γ) o each mu a ional ca ego y, we calcula ed he bu den (ƛ) o mu a-
ions o each ype in cases (Spanish coho ) o e con ols (LoF = 2.21; p obably damaging missense = 1.36).
Then we applied he ollowing o mula o calcula e ela i e isk: ɣ = 1 + (ƛ − 1)/π, whe e π, he ac ion o isk
genes, was se as 0.05 ( he de aul pa ame e ). Finally, by unning TADA-Deno o wi h he pa ame e s desc ibed
abo e, unco ec ed p alues o each gene we e calcula ed ob aining null dis ibu ions (N epe i ions = 10,000).
TADA-Deno o compu es BF (Bayesian Fac o ) o each gene. To de e mine an app op ia e h eshold ha allows
decla ing a “signi ican gene”, TADA uses he Bayesian FDR app oach o con ol o he a e o alse disco e ies.
q- alues o each gene we e calcula ed using he Bayesian FDR app oach p o ided by TADA. Manha an plo s
which show he esul s o TADA p alues (− log10) o he combined coho (PZMs and ge minal mu a ions)
we e done wi h R package qqman24.
Genes wi h FDR < 0.1 (ge minal genes) and FDR < 0.3 (PZMs genes) we e classi ied acco ding o SFARI
c i e ia (h ps ://gene.s a i .o g/da ab ase/gene-sco i n g/). Mo eo e , OMIM da abase (Online Mendelian Inhe i -
ance in Man) (h ps ://www.omim.o g/) was consul ed o sea ch o Mendelian diseases ela ed o hese genes.
Gene‑se en ichmen analysis o PZMs and ge minal mu a ions. Gene-se en ichmen analyses o
hose genes ca ying missense and nonsense DNMs (ge minal and PZMs) was done by DNENRICH25. DNEN-
RICH es ima es he en ichmen o DNMs wi hin p e-de ined g oups o genes accoun ing o gene size, i-
nucleo ide con ex and unc ional e ec o he mu a ions. DNMs included in his analysis we e ge minal and
PZMs iden i ied in he Spanish coho (ge minal DNMs = 236; PZMs = 164) (Addi ional ile3; TablesS9 and
S10). Fo he analysis o he combined coho , a subse o PZMs was c ea ed in o de o ensu e ha he analyzed
PZMs likely con ibu e o he pheno ype (PZMs = 676) (Addi ional ile1; TableS11). Fo ha pu pose, indi idu-
als wi h ge minal mu a ions in ASD isk genes (SFARI sco es 1 and 2) we e elimina ed om he PZMs da ase .
Thus, ge minal DNMs and he subse o PZMs om he combined coho we e used in he analysis (ge minal
DNMs = 2270; PZMs = 676) (Addi ional ile3; TableS12 and S13). The analysis was also un independen ly in
una ec ed siblings using da a p e iously published by he ASC (ge minal DNMs = 780; PZMs = 239) (Addi ional
ile1; TableS14).
The gene name alias and he gene size ma ix p o ided by DNENRICH we e used as inpu iles used in his
analysis oge he wi h he ollowing gene-se s: (1) FMRP a ge genes iden i ied by Da nell e al.26 and down-
loaded om Genebook (h p://zzz.b wh.ha a d.edu/g eneb ook/) (N = 788); (2) Genes included in he GO:0006325
ch oma in o ganiza ion (N = 723) (h p://www.geneo n olo gy.o g/); (3) Synap ic genes (N = 903)27; (4) Human
o hologs o genes essen ial in mice (N = 2472)28; (5) CHD8 a ge genes in human mid e al b ain (N = 2725)29;
(6) Lis o SFARI genes (N = 990) (h ps ://gene.s a i .o g/au db /HG_Home); (7) LoF in ole an genes (pLi > 0.9)
(N = 3230)29,30; 8) RBFOX a ge genes (N = 587)31; (9) miR-137 a ge genes (N = 428)32; (10) CELF-4 a ge genes
(N = 954)33; (11) Allele biased genes in di e en ia ing neu ons (N = 802)34; (12) Known in ellec ual disabili y
genes (N = 1547)35; (13) In e genic and In onic B ain Exp essed Enhance s (BEE) (N = 673)36; (14) Genomic
in e als su ounding known elencephalon genes scanned o enhance s (N = 79)37; and (15) miR-138 a ge
genes (N = 255)38. Empi ical p alues we e ob ained om one million pe mu a ions o each gene-se .
Gene on ology en ichmen analysis. An explo a o y GO en ichmen analysis was ca ied ou using he
En ich ool. This analysis allows s udying i genes ha bo ing ge minal DNMs o PZMs a e in ol ed in di e en
biological p ocesses. To his aim, he combined da ase o ge minal missense and nonsense DNMs and he sub-
se o PZMs we e employed (ge minal genes = 1972; PZMs genes = 624) (Addi ional ile3; TableS15). In addi-
ion, he REViGO ool (h p:// e ig o.i b.h /) was employed o isualize GO e ms in seman ic simila i y-based
sca e plo s using SimRel as a seman ic simila i y measu e. Thus, he op 30 en iched GO e ms in each g oup
(PZM s ge minal) we e isualized using a modi ica ion o he R sc ip p o ided by he REViGO online ool.
Ne wo k isualiza ion o he op 50 en iched e ms in each g oup o genes was pe o med wi h En ichmen
Map, a Cy oscape ( .3.6.1) plugin o unc ional en ichmen isualiza ion39. Each node ep esen s a gene-se
(GO e m) and he size o he node is p opo ional o he numbe o genes pa icipa ing in he GO e m (o e lap
coe icien ). Nodes we e conside ed as connec ed when he o e lap coe icien was g ea e han 0.7 and edge-
wid h ep esen s he o e lap be ween gene-se s. The bo de -wid h o each node ep esen s he co esponding
p alue o each GO e m.
Exp ession cell‑ ype en ichmen analysis and exp ession analysis ac oss b ain egions and
de elopmen al pe iods. Exp ession Weigh ed Cell- ype En ichmen (EWCE) me hod (h ps ://gi hu
b.com/Na ha nSken e/EWCE) was used o explo e whe he genes ha bo ing ge minal DNMs (N = 1972) and
genes ha bo ing PZMs (N = 624) (Addi ional ile3; TableS15) we e di e en ially exp essed ac oss se e al neu-
onal cell ypes. EWCE in ol es es ing whe he he gi en genes in a a ge lis ha e highe le els o exp ession in
a gi en cell ype compa ed o wha is expec ed by chance. B ain single-cell ansc ip omic da a om Ka olinska
Ins i u e (c d_allKI) was used o he EWCE analysis. B ain egions included in he KI mouse supe da ase
a e he neoco ex, hippocampus, hypo halamus, s ia um, and midb ain, as well as samples en iched o oli-
5
Vol.:(0123456789)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s/
godend ocy es, dopamine gic neu ons and co ical pa albumin in e neu ons ( o al cells = 9970). Backg ound
gene-se comp ises all human-mice o hologous. P obabili y dis ibu ion o ou gene lis s was calcula ed by
andomly sampling 100,000 genes om he backg ound se con olling o ansc ip leng h and GC con en .
Boo s apping unc ion was hen applied on le el 1 anno a ion.
pSI (speci ici y index s a is ic), an R package, was employed o s udy he exp ession o genes ha bo ing PZMs
and ge minal mu a ions ac oss di e en b ain egions and neu ode elopmen al pe iods40,41. Lis s o speci ically
exp essed human genes (human. da) ob ained om B ainSpan da a (gene-se s o 6 b ain egions and gene-
se s o 10 de elopmen al pe iods) we e employed. The Fishe i e a ion es included in he pSI package was
used o analyze i he lis ed genes ha bo ing PZMs and ge minal mu a ions we e signi ican ly o e ep esen ed.
B ain a eas signi ican ly en iched wi h PZMs o ge minal genes in speci ic de elopmen al pe iods (p adjus ed
alue < 0.05) we e ep esen ed as a ma ix. Bio ende (h ps ://bio e n de .com) was used o d aw he b ain images
(Fig.6).
E hics app o al and consen o pa icipa e. The co esponding Resea ch E hics Commi ee o Gali-
cia app o ed ou s udy: (Comi é É ico de In es igación Galicia ( he only IEC au ho ized in his au onomous
egion); Numbe : 2012/098; App o al: 28-June-2012; Ti le: Con ibución a la búsqueda de las causas gené icas
de los as o nos del espec o au is a. All pa icipan s, pa en s o legal ep esen a i es p o ided w i en in o med
consen a en ollmen and he s udy was conduc ed acco ding o he decla a ion o Helsinki. The ASC da a
employed in his s udy we e al eady published (h ps ://doi.o g/10.1038/nn.4598). The co esponding e hics
commi ee has app o ed hese gene ic da a.
Resul s
T ansmission and de no o associa ion es (TADA‑deno o). T ansmission and De no o Associa-
ion es (TADA) assesses i a gene is a ec ing ASD isk based on se e al pa ame e s: he gene mu a ion a e,
he ecu ence o DNMs in he gene and he se e i y o he mu a ions22. Thus, TADA-Deno o analysis was
independen ly un in bo h da ase s (ge minal and PZMs genes). The main aim o TADA-Deno o is o iden i y
hose genes ha could be di e en ially in ol ed in ASD e iology depending on he ype o DNMs ha bo ed by
hem (ge minal o PZMs). We ocused he analysis on damaging mu a ions (LoF and likely pa hogenic missense
a ian s) o inc ease he likelihood o inding “s ong” candida e genes. Fi s , he se o genes om he Spanish
coho (360 ios) (ge minal genes = 181; PZMs genes = 105) was analyzed. The analysis o he ge minal gene
lis iden i ied 12 genes wi h an FDR < 0.3 (Table1 and Addi ional ile2; TableS16) including 3 genes (SCN2A,
ARID1B and CHD8) wi h an FDR < 0.1. The analysis o he PZMs gene lis iden i ied 13 genes wi h an FDR < 0.3
(Table1 and Addi ional ile2; TableS17) o which 4 genes (KMT2C, FRG1, GRIN2B and MAP2K3) had an
FDR < 0.1.
In he combined coho (genes om he Spanish coho plus genes om he Lim e al. publica ion18 (ge minal
genes = 1210; PZMs genes = 362) TADA iden i ied 34 genes wi h an FDR < 0.1 (Table2 and Fig.1a) and 103 genes
wi h an FDR < 0.3 (Addi ional ile2; TableS18). Th ee o he genes (SCN2A, ARID1B, CHD8) wi h ge minal
DNMs we e p io i ized (FDR < 0.1) bo h in he combined coho and in he Spanish coho when TADA was
employed. Analysis o PZMs genes in he combined coho iden i ied h ee genes (FRG1, KMT2C and NFIA)
wi h an FDR < 0.1, and 14 genes wi h an FDR < 0.3 (Table2 and Fig.1b; Addi ional ile2; TableS19). Only wo
genes, KMT2C and FRG1, ha e emained signi ican a e FDR co ec ion (< 0.1) in bo h he Spanish and he
combined coho .
A o al o 17 genes (50%) om he se o ge minal genes in he combined coho (34 genes, FDR < 0.1) we e
iden i ied as “high con idence” o “s ong” ASD candida es ollowing SFARI Gene sco ing c i e ia (sco es 1,
2, 1s and 2s). In addi ion, 11 o hese genes (64, 70%) ha e shown an FDR < 0.1 in p e ious TADA analysis6.
Mo eo e , 10 o he emaining genes iden i ied by TADA (FDR < 0.1) a e included in SFARI gene lis s (sco es
3, 4 and 5) and 5 o he genes iden i ied by TADA we e epo ed in ela ion wi h ano he disease (no ASD) by
OMIM da abase (Addi ional ile4; TableS20).
PZM analysis has shown associa ion o KMT2C (SFARI sco e s2) as well as o he 3 genes FDR < 0.1). I is
wo h o no e ha NFIA has been p e iously epo ed as a plausible candida e gene in ASD (SFARI sco e 4) bu
his is he i s ime ha FRG1 is epo ed in ASD. SMARCA4, PRKDC, KLF16, GRIN2B and HNRNPU (SFARI
sco e 3 and 4) we e among hose plausible ASD candida e genes p e iously iden i ied wi h an FDR alue be ween
0.1 and 0.3. GRIN2B was p e iously epo ed by SFARI as a s ong ASD isk gene (sco e 1) (Addi ional ile4;
TableS21).
Gene‑se en ichmen analysis o PZMs and ge minal mu a ions. DNENRICH was un o es ima e
a s a is ical signi icance o en ichmen o ge minal and PZMs wi hin p e iously ASD and NDDs associa ed
gene-se s. Synonymous mu a ions we e excluded om he analysis because hey a e unlikely o con ibu e o
ASD pheno ype and only nonsense and missense mu a ions we e conside ed. Fi s , gene-se en ichmen analysis
was pe o med using he lis o genes and DNMs (ge minal and PZMs) om he Spanish coho (Addi ional
ile3; TablesS9 and S10) agains se e al backg ound gene lis s (see “Me hods” sec ion). Ou esul s indica e ha
ge minal genes shown en ichmen in se e al gene-se s (ge minal genes = 228, ge minal DNMs = 236): FMRP
a ge genes (p alue = 0.003), known in ellec ual disabili y genes (p alue = 0.0073), LoF in ole an genes (p
alue = 0.002), SFARI genes (p alue = 1 × 10–7) and genes in ol ed in ch oma in o ganiza ion (p alue = 0.00018)
(Table3). Howe e , only he LoF in ole an gene-se has shown associa ion wi h he lis o PZMs genes (PZMs
genes = 155, PZMs = 164) (Table3).
DNENRICH analysis in he combined coho demons a ed en ichmen o se e al gene-se s o bo h ge -
minal and PZMs genes: ch oma in o ganiza ion, SFARI genes, LoF in ole an genes, CHD8 a ge genes and
6
Vol:.(1234567890)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s/
essen ial genes. In addi ion, he ge minal gene lis (ge minal genes = 1972, ge minal DNMs = 2270) showed
en ichmen o FMRP a ge genes (p alue = 1 × 10–6), known in ellec ual disabili y genes (p alue = 1 × 10–6) and
synap ic genes (p alue = 4 × 10–6) (Table4, Fig.2). PZMs genes (PZMs genes = 624, PZMs = 676), ha e only shown
associa ion in he case o he miR-137 a ge gene-se (p alue = 0.0019) (Table4, Fig.2). The same analysis was
pe o med using he lis o ge minal genes om una ec ed siblings (ge minal genes = 744, ge minal DNMs = 780;
PZMs genes = 237, PZMs = 239) (Addi ional ile1; TableS14). A signi ican en ichmen was iden i ied only wi h
he FMRP a ge s gene-se (da a no shown).
Gene on ology en ichmen analysis. GO en ichmen analysis e ealed ema kable di e ences be ween
ge minal and PZMs gene lis s om he combined coho (Addi ional ile3; TableS15). The ge minal se showed
a signi ican en ichmen in di e en GO e ms ela ed o synap ic unc ion and ansc ip ion egula ion. In
pa icula , i is wo h o no e he associa ion o GO e ms ela ed o ion anspo : GO:0006814, Benjamini-
Hochbe g-co ec ed p[Pbh] = 0.005; GO0035725, Pbh = 0.005 and GO0006816, Pbh = 0.006 (Addi ional ile4;
TableS22, Fig.3a). The GO e ms en iched in he subse o PZMs a e ela ed o egula ion o gene exp es-
sion, biosyn hesis, di e en ia ion o mig a ion: GO0010629, p[Pbh] = 0.074; GO:2000113, p[Pbh] = 0.092;
GO:0045652, p[Pbh] = 0.092; GO:0030336, p[Pbh] = 0.0995. (Addi ional ile4; TableS23, Fig.3b).
GO en ichmen analysis was depic ed by seman ically clus e ing he op 50 en iched e ms o ge minal and
PZMs gene lis s. The ge minal gene lis esul ed in h ee di e en ia ed clus e s: neu on de elopmen and di -
e en ia ion, synap ic unc ions, and ch oma in modi ica ions. The exis ence o a ou h clus e , which includes
e ms ela ed o emb yonic de elopmen , was also highligh ed (Fig.4a). In he case o PZMs genes, all he clus e s
we e pa ially ela ed o each o he . Howe e , we iden i ied ano he clus e ha includes e ms ela ed o he
egula ion o co e p ocesses (e.g., p o ein phospho yla ion, egula ion o g ow h, nega i e egula ion o cellula
biosyn he ic p ocesses, posi i e egula ion o ansc ip ion DNA empla e). I is also impo an o highligh he
GO e ms ela ed o neu on and emb yonic de elopmen (Fig.4b).
Exp ession cell‑ ype en ichmen analysis and exp ession analysis ac oss b ain egions and
de elopmen al pe iods.. Fi s , we examined whe he ge minal genes o PZMs genes om he combined
coho , we e di e en ially exp essed in he ansc ip ome da ase co esponding o le el 1 cell ypes. As expec ed,
ge minal genes we e signi ican ly en iched in se e al cell ypes (Addi ional ile3; TableS24; Fig.5). The mos
en iched cell ypes we e hose ela ed o neu o ansmission (dopamine gic neu oblas ; p alue < 0.0001; emb y-
onic dopamine gic neu ons, p alue < 0.0001; emb yonic GABAe gic neu ons, p alue < 0.0001; se o one gic
Table 1. ASD isk genes ca ying ge minal DNMs and PZMs in he Spanish coho . p alues and q- alues
we e ob ained a e unning TADA-Deno o using ge minal DNMs and PZMs om he Spanish ASD coho
(N = 360). Only genes wi h q- alues < 0.3 a e shown.
Genes q- alue p alue Mu a ions
SCN2A 0.004 2.76 × 10–7 Ge minal
ARID1B 0.050 2.76 × 10–6 Ge minal
CHD8 0.066 3.31 × 10–6 Ge minal
FIG4 0.103 9.39 × 10–5 Ge minal
RBM15 0.126 1.16 × 10–5 Ge minal
HUWE1 0.148 3.54 × 10–5 Ge minal
KIAA1107 0.188 5.08 × 10–5 Ge minal
VWAS5B1 0.218 5.08 × 10–5 Ge minal
EMCN 0.242 6.13 × 10–5 Ge minal
SH2B2 0.262 7.90 × 10–5 Ge minal
ASMT 0.277 9.34 × 10–5 Ge minal
MYLK4 0.291 9.67 × 10–5 Ge minal
KMT2C 0.001 4.76 × 10–7 PZM
FRG1 0.015 4.46 × 10–7 PZM
GRIN2B 0.040 4.76 × 10–7 PZM
MAP2K3 0.08 6.67 × 10–6 PZM
SRGAP2 0.106 7.62 × 10–6 PZM
MBD6 0.124 7.62 × 10–6 PZM
POTEB2 0.168 2.57 × 10–5 PZM
CALML6 0.200 2.67 × 10–5 PZM
PRDX6 0.226 3.05 × 10–5 PZM
SSR2 0.245 3.14 × 10–5 PZM
VEGFA 0.264 4 × 10–5 PZM
CANX 0.278 0.0001 PZM
ZNF276 0.290 0.00012 PZM
7
Vol.:(0123456789)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s/
Table 2. ASD isk genes ca ying ge minal DNMs and PZMs in he combined coho . p alues and q- alues
we e ob ained a e unning TADA-Deno o using ge minal DNMs and PZMs om he combined coho
(N = 2103). Genes wi h q- alues < 0.1 a e shown in he case o ge minal DNMs and genes wi h q- alue < 0.3 a e
shown in he case o PZMs.
Gene q- alue p alue Mu a ions
SCN2A 5.04 × 10–12 4.13 × 10–8 Ge minal
CHD8 2.40 × 10–5 4.13 × 10–8 Ge minal
ARID1B 5.36 × 10–5 4.13 × 10–8 Ge minal
SLC6A1 0.00015 2.48 × 10–7 Ge minal
SYNGAP1 0.0005 6.61 × 10–7 Ge minal
KDM5B 0.0008 8.26 × 10–7 Ge minal
SUV420H1 0.002 5.37 × 10–6 Ge minal
TRIP12 0.003 5.79 × 10–6 Ge minal
PTEN 0.004 1.14 × 10–5 Ge minal
KATNAL2 0.008 5.01 × 10–5 Ge minal
NRXN1 0.012 5.79 × 10–5 Ge minal
CREBBP 0.02 5.92 × 10–5 Ge minal
CELF4 0.02 6.09 × 10–5 Ge minal
STXBP1 0.02 6.48 × 10–5 Ge minal
DYRK1A 0.02 7.09 × 10–5 Ge minal
CHD2 0.03 0.0001 Ge minal
ANK2 0.03 0.0001 Ge minal
WDFY3 0.03 0.0001 Ge minal
UNC80 0.04 0.0002 Ge minal
CLASP1 0.04 0.0002 Ge minal
TMEM39B 0.05 0.0002 Ge minal
PRKAR1B 0.05 0.0002 Ge minal
USP45 0.05 0.0003 Ge minal
NUAK1 0.06 0.0004 Ge minal
NAA15 0.06 0.0004 Ge minal
FOXP1 0.07 0.0004 Ge minal
ZC3H11A 0.07 0.0004 Ge minal
DPP3 0.07 0.0005 Ge minal
PRKDC 0.08 0.0005 Ge minal
ATP1A1 0.08 0.0005 Ge minal
LRP5 0.09 0.0005 Ge minal
SLC12A3 0.09 0.0006 Ge minal
FBXO18 0.096 0.0006 Ge minal
PTK7 0.0999 0.0007 Ge minal
FRG1 0.04 4.14 × 10–3 PZM
KMT2C 0.07 0.00018 PZM
NFIA 0.09 0.00028 PZM
SMARCA4 0.12 0.00052 PZM
PRKDC 0.13 0.00055 PZM
KLF16 0.15 0.00064 PZM
GRIN2B 0.17 0.00095 PZM
MAP2K3 0.18 0.00098 PZM
HNRNPU 0.21 0.0019 PZM
POTEB2 0.23 0.002 PZM
RNPC3 0.25 0.002 PZM
FAM177A1 0.27 0.002 PZM
CALML6 0.28 0.002 PZM
CMPK2 0.3 0.003 PZM
8
Vol:.(1234567890)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s/
neu ons, p alue < 0.0001). The PZMs gene lis showed en ichmen o h ee di e en cell ypes: py amidal CA1
neu ons; p alue = 0.0066, py amidal soma osenso y (SS); p alue = 0.0152, emb yonic midb ain nucleus neu-
ons; p alue = 0.0185 (Addi ional ile3; TableS25, Fig.5).
To gain insigh in o he spa io empo al dis ibu ion, we analyzed he exp ession o ge minal and PZMs genes
(Addi ional ile3; TableS15) ac oss se e al b ain egions and di e en neu ode elopmen al pe iods ob ained
Figu e1. Manha an plo s depic ing he ASD isk genes p io i ized by TADA-Deno o (ch omosome and
log10 p alue o each gene a e ep esen ed in axis x and y). (a) p alues we e ob ained om analysis o ge minal
mu a ions in he combined coho using TADA-Deno o. Red line ep esen s he p alue < 1 × 10–8 and blue line p
alue < 1 × 10−5. (b) p alues we e ob ained om analysis o PZMs in he combined coho using TADA-Deno o.
Blue line ep esen s p alue < 1 × 10−5.
Table 3. Resul s o he gene-se en ichmen analysis o he lis o genes ha bo ing ge minal DNMs and PZMs
(Spanish coho ).
Gene-se s p alue Obse ed mu a ions Expec ed mu a ions Analysis
SFARI genes 1 × 10–6 49 20.9601 Ge minal
FMRP a ge s 0.00013 39 20.9788 Ge minal
Genes in ol ed in ch oma in o ganiza ion 0.00018 24 10.7377 Ge minal
LoF in ole an genes 0.0018 79 58.9156 Ge minal
Known ID genes 0.0073 40 27.0286 Ge minal
Essen ial genes 0.0709 49 39.968 Ge minal
CHD8 a ge s 0.226 39 34.4983 Ge minal
mi 137 a ge s 0.3251 8 6.49354 Ge minal
Synap ic genes 0.3592 14 12.4002 Ge minal
Genomic in e als su ounding known elencephalon genes
scanned o enhance s 0.5386 1 0.771873 Ge minal
Alelle biased genes in di e en ia ing neu ons 0.6 12 12.502 Ge minal
CELF4 a ge s 1 0 0.05078 Ge minal
mi 128 a ge s 1 0 0.022676 Ge minal
RBFOX a ge s 1 0 0.055409 Ge minal
LoF in ole an genes 0.0283 51 39.9973 PZM
Genes in ol ed in ch oma in o ganiza ion 0.0625 12 7.28793 PZM
FMRP a ge s 0.0764 20 14.2422 PZM
S a i genes 0.0764 20 14.2361 PZM
CHD8 a ge s 0.0884 30 23.4115 PZM
Genomic in e als su ounding known elencephalon genes
scanned o enhance s 0.0977 2 0.524631 PZM
Essen ial genes 0.234 31 27.1352 PZM
Synap ic genes 0.3339 10 8.41702 PZM
Known ID genes 0.7554 16 18.3456 PZM
mi 137 a ge s 0.8198 3 4.40692 PZM
Alelle biased genes in di e en ia ing neu ons 0.9727 4 8.48746 PZM
CELF4 a ge s 1 0 0.034487 PZM
mi 128 a ge s 1 0 0.015316 PZM
RBFOX a ge s 1 0 0.037786 PZM
9
Vol.:(0123456789)
Scien i ic Repo s | (2021) 11:319 | h ps://doi.o g/10.1038/s41598-020-79412-w
www.na u e.com/scien i ic epo s/
om B ainSpan. Ge minal genes we e signi ican ly exp essed in he co ex, s ia um, ce ebellum and amygdala
in p ena al s ages (ea ly, ea ly mid and la e) (Addi ional ile3; TableS26, Fig.6a,c). PZMs genes we e signi ican ly
exp essed in he co ex du ing he ea ly mid- e al pe iod. Al hough we did no ind ound a signi ican en ich-
men in o he b ain a eas o neu ode elopmen al pe iods o PZMs, p alues close o he signi icance h eshold
we e ound in co ex (ea ly, la e mid- e al) and amygdala (la e mid- e al) (Addi ional ile3; TableS27, Fig.6b,d).
Table 4. Resul s o he gene-se en ichmen analysis o he lis o genes ha bo ing ge minal DNMs and PZMs
(combined coho ).
Gene-se p alue Obse ed mu a ions Expec ed mu a ions Analysis
Essen ial genes 1 × 10–6 533 404.162 Ge minal
FMRP a ge s 1 × 10–6 331 212.218 Ge minal
Known ID genes 1 × 10–6 373 273.213 Ge minal
LoF in ole an genes 1 × 10–6 733 595.59 Ge minal
SFARI genes 1 × 10–6 434 211.973 Ge minal
Synap ic genes 4 × 10–6 180 125.408 Ge minal
Genes in ol ed in ch oma in o ganiza ion 1 × 10–6 168 108.449 Ge minal
CHD8 a ge s 3.5 × 10–5 420 348.544 Ge minal
Genomic in e als su ounding known elencephalon genes
scanned o enhance s 0.0550 13 782.029 Ge minal
mi 137 a ge s 0.1368 75 657.234 Ge minal
RBFOX a ge s 0.4299 1 0.562709 Ge minal
Alelle biased genes in di e en ia ing neu ons 0.6999 121 126.324 Ge minal
CELF4 a ge s 1 0 0.51357 Ge minal
mi 128 a ge s 1 0 0.22838 Ge minal
SFARI genes 1 × 10–7 127 612.987 PZM
LoF in ole an genes 0.0003 212 172.214 PZM
mi 137 a ge s 0.0019 33 190.234 PZM
Genes in ol ed in ch oma in o ganiza ion 0.0106 45 313.424 PZM
Essen ial genes 0.0118 140 116.895 PZM
CHD8 a ge s 0.0228 120 100.751 PZM
FMRP a ge s 0.0425 75 614.294 PZM
Genomic in e als su ounding known elencephalon genes
scanned o enhance s 0.0796 5 22.672 PZM
Known ID genes 0.1840 87 790.127 PZM
Synap ic genes 0.6109 35 362.813 PZM
Alelle biased genes in di e en ia ing neu ons 0.9744 26 364.938 PZM
CELF4 a ge s 1 0 0.147914 PZM
mi 128 a ge s 1 0 0.066195 PZM
RBFOX a ge s 1 0 0.162842 PZM
Figu e2. Gene-se en ichmen analysis using ge minal and PZM om he combined coho . Gene-se
en ichmen analysis was done wi h DNENRICH. − log10 p alue o each gene-se is shown o each ype o
mu a ion and es ed gene-se .