Mining the Link Structure of the World Wide Web
Abstract The World Wide Web contains an enormous amount of information, but it can be exceedingly difficult for users to locate resources that are both high in quality and relevant to their information needs. We develop algorithms that exploit the hyperlin
byenteringquerytermsandreceivingalistofWebpagesthatcontainthegiventerms.Experienceduserscanmakee ectiveuseofsuchsearchenginesfortasksthatcanbesolvedbysearchingfortightlyconstrainedkeywordsandphrases;however,thesesearchenginesarenotsuitedforawiderangeofequallyimportanttasks.Inparticular,atopicofanybreadthwilltypicallycontainseveralthousandorseveralmillionrelevantWebpages;atthesametime,auserwillbewillingtolookatanextremelysmallnumberofthesepages.How,fromthisseaofpages,shouldasearchengineselectthe“correct”ones?
Ourworkbeginsfromtwocentralobservations.First,inordertodistillalargesearchtopicontheWWWdowntoasizethatwillmakesensetoahumanuser,weneedameansofidentifyingthemost“de nitive,”or“authoritative,”Webpagesonthetopic.Thisnotionofauthorityaddsacrucialseconddimensiontothenotionofrelevance:wewishnotonlytolocateasetofrelevantpages,butrathertherelevantpagesofthehighestquality.Second,theWebconsistsnotonlyofpagesbutofhyperlinksthatconnectonepagetoanother;andthishyperlinkstructurecontainsanenormousamountoflatenthumanannotationthatcanbeextremelyvaluableforautomaticallyinferringnotionsofauthority.Speci cally,thecreationofahyperlinkbytheauthorofaWebpagerepresentsanimplicittypeof“endorsement”ofthepagebeingpointedto;byminingthecollectivejudgmentcontainedinthesetofsuchendorsements,wecanobtainaricherunderstandingofboththerelevanceandqualityoftheWeb’scontents.
TherearemanywaysthatonecouldtryusingthelinkstructureoftheWebtoinfernotionsofauthority,andsomeofthesearemuchmoree ectivethanothers.Thisisnotsurprising:thelinkstructureimpliesanunderlyingsocialstructureinthewaythatpagesandlinksarecreated,anditisanunderstandingofthissocialorganizationthatcanprovideuswiththemostleverage.OurgoalindesigningalgorithmsformininglinkinformationistodeveloptechniquesthattakeadvantageofwhatweobserveabouttheintrinsicsocialorganizationoftheWeb.
SearchingforAuthoritativePages.Aswethinkaboutthetypesofpageswehopetodiscover,andthefactthatwewishtodosoautomatically,wearequicklyledtosomedi cultproblems.First,itisnotsu cientto rstapplypurelytext-basedmethodstocollectalargenumberofpotentiallyrelevantpages,andthencombthissetforthemostauthoritativeones.Ifweweretryingto ndthemainWWWsearchengines,itwouldbeaseriousmistaketorestrictourattentiontothesetofallpagescontainingthephrase“searchengines.”Foralthoughthissetisenormous,itdoesnotcontainsmostofthenaturalauthoritieswewouldliketo nd(e.g.Yahoo!,Excite,InfoSeek,AltaVista).Similarly,thereisnoreasontoexpectthehomepagesofHondaorToyotatocontaintheterm“Japaneseautomobilemanufacturers,”orthehomepagesofMicrosoftorLotustocontaintheterm“softwarecompanies.”Authoritiesareoftennotparticularlyself-descriptive;largecorporationsforinstancedesigntheirWebpagesverycarefullytoconveyacertainfeel,andprojectthecorrectimage—thisgoalmightbeverydi erentfromthegoalofdescribingthecompany.Peopleoutsideacompanyfrequentlycreatemorerecognizable(andsometimesbetter)judgmentsthanthecompanyitself.
Theseconsiderationsindicatesomeofthedi cultieswithrelyingontextaswesearchforauthoritativepages.Therearedi cultiesinmakinguseofhyperlinkinformationaswell.


