BIOINFORMATICS Building an Abbreviation Dictionary Using a Term Recognition Approach

Motivation: Acronyms result from a highly productive type of term variation and trigger the need for an acronym dictionary to establish associations between acronyms and their expanded forms. Results: We propose a novel method for recognizing acronym defin

Okazakietal

(Schwartzetal.,2003;Wrenetal.,2002).

longform’(’shortform’)’

(1)

examined

111

explore

Forexample,thesentence,“Theexactroutewasdeterminedbymagneticresonanceimaging(MRI)”,couldyieldthetextualfragmentmarkedwiththeitalicletters6.Thetaskistoidentifythe“authentic”long-forminthetextualfragmentifany.Existingmethodsforsolvingthisproblemcanbecategorizedintothreegroups:usingheuristicsand/orscoringrules(Adar,2004;Aoetal.,2005;Schwartzetal.,2003;Taghvaetal.,1999;Wrenetal.,2002;Yuetal.,2002);machinelearning(Changetal.,2006;Pakhomov,2002;Nadeauetal.,2005);andstatistics(Hisamitsuetal.,2001;Liuetal.,2003).

The rstcategoryusesprede nedheuristicrules/algorithmsto ndalongforminatextualfragment.Forexample,Schwartzetal.(2003)implementedaletter-matchingalgorithmthatmapsallalpha-numericallettersintheshortformtothelongform,startingfromtheendofboththeshortandlongformsandmovingrighttoleft.Eventhoughthecorealgorithmisverysimple,theauthorsreport96%precisionand82%recallontheMedstractgoldstandard7.Adar(2004)proposesscoringrulesto ndthemostlikelylong-form,acceptingmultiplelong-formcandidates,e.g.,determinedbymagneticresonanceimaging(MRI)andmagneticresonanceimaging(MRI)inthefragment,yielding95%precisionand85%recallontheMedstractcorpus.

Thesecondcategoryobtainssuchrulesbyusingamachinelearningtechnique.Changetal.(2006)appliedalogisticregressiontocalculatethelikelihoodoflong-formcandidates.Theyenumeratepossiblelong-formcandidateswithLongestCommonSubstring(LCS)formalization(Taghvaetal.,1999).Thelikelihoodofthecandidatesisestimatedastheprobabilitycalculatedfromalogisticregressionwithninefeaturessuchasthepercentageoflong-formlettersalignedatthebeginningofaword,thepercentageofshort-formlettersalignedtothelongform,etc.Theirmethodachieved80%precisionand83%recallontheMedstractcorpus.

Thethirdcategory,whereourproposedmethodbelongs,utilizesstatisticalcluesinthesourcedocuments,e.g.,co-occurrencebetweenshortformsandlongforms.Hisamitsuetal.(2001)proposedamethodforextractingusefulparentheticalexpressionsfromJapanesenewspaperarticles.Theirmethodmeasurestheco-occurrencestrengthbetweentheinnerandouterphrasesofaparentheticalexpressionviamutualinformation,χ2testwithYate’scorrection,Dicecoef cient,log-likelihoodratio,etc.Unfortunately,theirmethoddealswithgenericparentheticalexpressions(i.e.,abbreviation,non-abbreviationparaphrases,supplementarycomments),notfocusingexclusivelyonacronymrecognition.Liuetal.(2003)basedtheirmethodoncollocationsoccurringbeforetheparentheticalexpressions.Enumeratinglong-formcandidatesascollocationsappearingmorethanonceinatextcollection,theirmethodeliminatesunlikelycandidateswithrulessuchas“removeasetofcandidatesTwformedbyaddingapre xwordtoacandidatewifthenumberofsuchcandidatesTwisgreaterthan3”.Theyreportaprecisionof96.3%andarecallof88.5%forabbreviationrecognitionontheirtestcorpus.

6

increased

111studiedexpression of1its............

33......geneencoding

1......co-expression of1......regulation of the1......containing1......expressed1......stained for1......identification of......

......

............

............thyroid

1tissue

3

..................

thyroidthyroidthyroid

1

transciption found in the MEDLINE abstracts.

1

nuclear

transription

1

* These candidates are spelling mistakes

..................

209213thyroidtranscriptionfactor

BIOINFORMATICS Building an Abbreviation Dictionary Using a Term Recognition Approach

specific

4

......

factor

2

nkx2

Fig.1.ExpressionsappearingbeforetheacronymTTF-1inparentheses.

3METHODOLOGY

3.1Recognizingacronymsbasedonco-occurrence

Weassumeawordsequenceisapossiblelong-form8ifthewordsequenceco-occursfrequentlywithaspeci cacronymandnotwithothersurroundingwords.Figure1illustratesourassumptionwiththeacronymTTF-1.ThetreeconsistsofexpressionscollectedfromallsentenceswiththeacronymTTF-1inparenthesesandappearingbeforetheacronym.Anoderepresentsaword,andapathfromanynodetoTTF-1representsalong-formcandidate9.The gureaboveeachnodeshowstheco-occurrencefrequencyofthecorrespondinglong-formcandidate.Forexample,long-formcandidates1,factor1,transcriptionfactor1,andthyroidtranscriptionfactor1co-occur218,216,213,and209timesrespectivelywiththeacronymTTF-1inthetextcollection.

Eventhoughlong-formcandidates1,factor1andtranscriptionfactor1co-occurfrequentlywithTTF-1,theyalsoco-occurfrequentlywiththyroid.Meanwhile,thecandidatethyroidtranscriptionfactor1isusedinanumberofcontexts(e.g.,expressionofthyroidtranscriptionfactor1,expressedthyroid::::::::::::::::::

transcriptionfactor1,etc).Therefore,weobservethestrongestrelationshipisbetweenacronymTTF-1anditslong-formcandidatethyroidtranscriptionfactor1inthetree.Weapplyavalidationrule(describedlater)tothelong-formcandidatetomakesureanacronym-de nitionrelationdoesoccur.Inthisexample,thecandidatepairislikelytobeinanacronym-de nitionrelationasthelongformthyroidtranscriptionfactor1containsallthealphanumericlettersintheshortformTTF-1.

Thisapproachdetectsthestartingpointofthelongformwithoutusinglettermatching.Asimplemethodbasedonlettermatchingmaymisrecognizethelongformtranscriptionfactor1sinceitalsocontainsthenecessaryelementstoproducetheacronymTTF-1.Whereaspreviousworkdealtwiththiscasebyintroducing,e.g.,a

8

Assumingwetake(l+4)wordsappearingbeforetheparentheticalexpression(Adar,2004),wherelisthenumberoflettersintheshortform.7www.1mpi.com

BIOINFORMATICS Building an Abbreviation Dictionary Using a Term Recognition Approach

Asequenceofwordsthatco-occurswithanacronymdoesnotalwaysimplytheacronym-de nitionrelation:theacronym5-HTco-occursfrequentlywiththetermserotonin,buttheirrelationisinterpretedasasynonymousrelation.Wedealwiththisissuewithavalidationrulelater.9Thewordswithfunctionwords(e.g.,expressionof,regulationofthe,

::::::

etc.)aremergedintoanode.Thisisduetotherequirementforalong-formcandidatediscussedlater(Section3.2).

2

BIOINFORMATICS Building an Abbreviation Dictionary Using a Term Recognition Approach相关文档

最新文档

返回顶部