Mining the Link Structure of the World Wide Web

Abstract The World Wide Web contains an enormous amount of information, but it can be exceedingly difficult for users to locate resources that are both high in quality and relevant to their information needs. We develop algorithms that exploit the hyperlin

MiningtheLinkStructureoftheWorldWideWebSoumenChakrabarti ByronE.Dom DavidGibson JonKleinberg RaviKumar PrabhakarRaghavan

SridharRajagopalan AndrewTomkins

February,1999

Abstract

TheWorldWideWebcontainsanenormousamountofinformation,butitcanbeexceedinglydi cultforuserstolocateresourcesthatarebothhighinqualityandrelevanttotheirinformationneeds.WedevelopalgorithmsthatexploitthehyperlinkstructureoftheWWWforinformationdiscoveryandcategorization,theconstructionofhigh-qualityresourcelists,andtheanalysisofon-linehyperlinkedcommunities.1Introduction

TheWorldWideWebcontainsanenormousamountofinformation,butitcanbeexceedinglydi cultforuserstolocateresourcesthatarebothhighinqualityandrelevanttotheirinformationneeds.Thereareanumberoffundamentalreasonsforthis.TheWebisahypertextcorpusofenormoussize—approximatelythreehundredmillionWebpagesasofthiswriting—anditcontinuestogrowataphenomenalrate.Butthevariationinpagesisevenworsethantherawscaleofthedata:thesetofWebpagestakenasawholehasalmostnounifyingstructure,withvariabilityinauthoringstyleandcontentthatisfargreaterthanintraditionalcollectionsoftextdocuments.Thislevelofcomplexitymakesitimpossibletoapplytechniquesfromdatabasemanagementandinformationretrievalinan“o -the-shelf”fashion.

Index-basedsearchenginesfortheWWWhavebeenoneoftheprimarytoolsbywhichusersoftheWebsearchforinformation.ThelargestsuchsearchenginesexploitthefactthatmodernstoragetechnologymakesitpossibletostoreandindexalargefractionoftheWWW;theycanthereforebuildgiantindicesthatallowonetoquicklyretrievethesetofallWebpagescontainingagivenwordorstring.Ausertypicallyinteractswiththem

Mining the Link Structure of the World Wide Web相关文档

最新文档

返回顶部