Mining the Link Structure of the World Wide Web
Abstract The World Wide Web contains an enormous amount of information, but it can be exceedingly difficult for users to locate resources that are both high in quality and relevant to their information needs. We develop algorithms that exploit the hyperlin
MiningtheLinkStructureoftheWorldWideWebSoumenChakrabarti ByronE.Dom DavidGibson JonKleinberg RaviKumar PrabhakarRaghavan
SridharRajagopalan AndrewTomkins
February,1999
Abstract
TheWorldWideWebcontainsanenormousamountofinformation,butitcanbeexceedinglydi cultforuserstolocateresourcesthatarebothhighinqualityandrelevanttotheirinformationneeds.WedevelopalgorithmsthatexploitthehyperlinkstructureoftheWWWforinformationdiscoveryandcategorization,theconstructionofhigh-qualityresourcelists,andtheanalysisofon-linehyperlinkedcommunities.1Introduction
TheWorldWideWebcontainsanenormousamountofinformation,butitcanbeexceedinglydi cultforuserstolocateresourcesthatarebothhighinqualityandrelevanttotheirinformationneeds.Thereareanumberoffundamentalreasonsforthis.TheWebisahypertextcorpusofenormoussize—approximatelythreehundredmillionWebpagesasofthiswriting—anditcontinuestogrowataphenomenalrate.Butthevariationinpagesisevenworsethantherawscaleofthedata:thesetofWebpagestakenasawholehasalmostnounifyingstructure,withvariabilityinauthoringstyleandcontentthatisfargreaterthanintraditionalcollectionsoftextdocuments.Thislevelofcomplexitymakesitimpossibletoapplytechniquesfromdatabasemanagementandinformationretrievalinan“o -the-shelf”fashion.
Index-basedsearchenginesfortheWWWhavebeenoneoftheprimarytoolsbywhichusersoftheWebsearchforinformation.ThelargestsuchsearchenginesexploitthefactthatmodernstoragetechnologymakesitpossibletostoreandindexalargefractionoftheWWW;theycanthereforebuildgiantindicesthatallowonetoquicklyretrievethesetofallWebpagescontainingagivenwordorstring.Ausertypicallyinteractswiththem


