Mining the Link Structure of the World Wide Web

Abstract The World Wide Web contains an enormous amount of information, but it can be exceedingly difficult for users to locate resources that are both high in quality and relevant to their information needs. We develop algorithms that exploit the hyperlin

byenteringquerytermsandreceivingalistofWebpagesthatcontainthegiventerms.Experienceduserscanmakee ectiveuseofsuchsearchenginesfortasksthatcanbesolvedbysearchingfortightlyconstrainedkeywordsandphrases;however,thesesearchenginesarenotsuitedforawiderangeofequallyimportanttasks.Inparticular,atopicofanybreadthwilltypicallycontainseveralthousandorseveralmillionrelevantWebpages;atthesametime,auserwillbewillingtolookatanextremelysmallnumberofthesepages.How,fromthisseaofpages,shouldasearchengineselectthe“correct”ones?

Ourworkbeginsfromtwocentralobservations.First,inordertodistillalargesearchtopicontheWWWdowntoasizethatwillmakesensetoahumanuser,weneedameansofidentifyingthemost“de nitive,”or“authoritative,”Webpagesonthetopic.Thisnotionofauthorityaddsacrucialseconddimensiontothenotionofrelevance:wewishnotonlytolocateasetofrelevantpages,butrathertherelevantpagesofthehighestquality.Second,theWebconsistsnotonlyofpagesbutofhyperlinksthatconnectonepagetoanother;andthishyperlinkstructurecontainsanenormousamountoflatenthumanannotationthatcanbeextremelyvaluableforautomaticallyinferringnotionsofauthority.Speci cally,thecreationofahyperlinkbytheauthorofaWebpagerepresentsanimplicittypeof“endorsement”ofthepagebeingpointedto;byminingthecollectivejudgmentcontainedinthesetofsuchendorsements,wecanobtainaricherunderstandingofboththerelevanceandqualityoftheWeb’scontents.

TherearemanywaysthatonecouldtryusingthelinkstructureoftheWebtoinfernotionsofauthority,andsomeofthesearemuchmoree ectivethanothers.Thisisnotsurprising:thelinkstructureimpliesanunderlyingsocialstructureinthewaythatpagesandlinksarecreated,anditisanunderstandingofthissocialorganizationthatcanprovideuswiththemostleverage.OurgoalindesigningalgorithmsformininglinkinformationistodeveloptechniquesthattakeadvantageofwhatweobserveabouttheintrinsicsocialorganizationoftheWeb.

SearchingforAuthoritativePages.Aswethinkaboutthetypesofpageswehopetodiscover,andthefactthatwewishtodosoautomatically,wearequicklyledtosomedi cultproblems.First,itisnotsu cientto rstapplypurelytext-basedmethodstocollectalargenumberofpotentiallyrelevantpages,andthencombthissetforthemostauthoritativeones.Ifweweretryingto ndthemainWWWsearchengines,itwouldbeaseriousmistaketorestrictourattentiontothesetofallpagescontainingthephrase“searchengines.”Foralthoughthissetisenormous,itdoesnotcontainsmostofthenaturalauthoritieswewouldliketo nd(e.g.Yahoo!,Excite,InfoSeek,AltaVista).Similarly,thereisnoreasontoexpectthehomepagesofHondaorToyotatocontaintheterm“Japaneseautomobilemanufacturers,”orthehomepagesofMicrosoftorLotustocontaintheterm“softwarecompanies.”Authoritiesareoftennotparticularlyself-descriptive;largecorporationsforinstancedesigntheirWebpagesverycarefullytoconveyacertainfeel,andprojectthecorrectimage—thisgoalmightbeverydi erentfromthegoalofdescribingthecompany.Peopleoutsideacompanyfrequentlycreatemorerecognizable(andsometimesbetter)judgmentsthanthecompanyitself.

Theseconsiderationsindicatesomeofthedi cultieswithrelyingontextaswesearchforauthoritativepages.Therearedi cultiesinmakinguseofhyperlinkinformationaswell.

Mining the Link Structure of the World Wide Web相关文档

最新文档

返回顶部