on Neutral Evolution of Model Proteins Di4usion in Sequence Space and Overdispersion
We stimulate the evolution of model protein sequences subject to mutations. A mutation is considered neutral if it conserves (1) the structure of the ground state, (2) its thermodynamic stability and (3) its kinetic accessibility. All other mutations are c
J.theor.Biol.(1999)200,49}64
ArticleNo.jtbi.1999.0975,availableonlineatwww.1mpi.com
NeutralEvolutionofModelProteins:
Di4usioninSequenceSpaceandOverdispersion
UGOBASTOLLA*-,H.EDUARDOROMAN?
AND
MICHELEVENDRUSCOLOA
*H¸RZ,ForschungszentrumJulich,D-52425Julich,Germany,
?DipartimentodiFisicaandINFN,;niversitadiMilano,I-20133Milano,Italy
ADepartmentofPhysicsofComplexSystems,=eizmannInstituteofScience,Rehovot76100,Israel
(Receivedon12November1998,Acceptedinrevisedformon21May1999)
Westimulatetheevolutionofmodelproteinsequencessubjecttomutations.Amutationisconsideredneutralifitconserves(1)thestructureofthegroundstate,(2)itsthermodynamicstabilityand(3)itskineticaccessibility.Allothermutationsareconsideredlethalandarerejected.Weadoptalatticemodel,amenabletoareliablesolutionoftheproteinfoldingproblem.Weprovetheexistenceofextendedneutralnetworksinsequencespace*sequencescanevolveuntiltheirsimilaritywiththestartingpointisalmostthesameasforrandomsequences.Furthermore,we"ndthattherateofneutralmutationshasabroaddistributioninsequencespace.Duetothisfact,thesubstitutionprocessisoverdispersed(theratiobetweenvarianceandmeanislargerthan1).Thisresultisincontrastwiththesimplestmodelofneutralevolution,whichassumesaPoissonprocessforsubstitutions,andinqualitativeagreementwiththebiologicaldata.
1999AcademicPress
1.Introduction
ArecentstudyontheProteinDataBank(PDB)showsthatthedistributionofpairwisesequenceidentitybetweenstructurallyhomologouspro-teinspresentsalargepeakat8.5%sequenceidentity,onlyslightlylargerthanwhatexpectedinthepurelyrandomcaseB(Rost,1998).Thisis
-Authortowhomcorrespondenceshouldbeaddressed.Presentaddress:FreieUniversitatBerlin,FBChemie,Takustr.6,D-14195Berlin,Germany.E-mail:ugo@chemie.fu-berlin.de.
BThenumberofaminoacidmatchesobtainedbypairingtworandomsequencesofthesamelengthisgivenbythebinomialdistributionwithp"ifoneassumesthatthe20
aminoacidshavethesameprobabilitytooccur.Forse-quencesoflengthNtherewillbeonaveragepNidenticalaminoacids,withavarianceNp(1!p).Forrandomse-quences,95%ofpairwisecomparisonsyieldasequenceidentitybetween1and9%.0022}5193/99/017049#16$30.00/0
aninterestingresultwhichmeansthatthestruc-turalsimilaritydoesnotimplysequencesimilarity.Anintensivecomputationalstudyonsecond-arystructuresofRNAmolecules(Schusteretal.,1994),whichisaproblemmuchsimplerthanproteinfolding,andcanbestudiedthroughe$-cientandreliablealgorithms,showedthatanexponentiallylargenumberofsequencescorres-pondsinaveragetoasinglestructure,andthedistributionofstructuresinsequencespaceisquiteinhomogeneous(itfollowsaZipflaw).Se-quencesfoldingintothemostcommonstructuresformconnected&&neutralnetworks''thatperco-latesequencespace.Theseneutralnetworksdirectlyarisefromthenon-uniquenessoftherelationbetweensequenceandstructure.
Theseresultsareimportanttounderstandhowevolutionworksatthemolecularlevel.Kimura
1999AcademicPress


