BIOINFORMATICS Building an Abbreviation Dictionary Using a Term Recognition Approach
Motivation: Acronyms result from a highly productive type of term variation and trigger the need for an acronym dictionary to establish associations between acronyms and their expanded forms. Results: We propose a novel method for recognizing acronym defin
Okazakietal
(Schwartzetal.,2003;Wrenetal.,2002).
longform’(’shortform’)’
(1)
examined
111
explore
Forexample,thesentence,“Theexactroutewasdeterminedbymagneticresonanceimaging(MRI)”,couldyieldthetextualfragmentmarkedwiththeitalicletters6.Thetaskistoidentifythe“authentic”long-forminthetextualfragmentifany.Existingmethodsforsolvingthisproblemcanbecategorizedintothreegroups:usingheuristicsand/orscoringrules(Adar,2004;Aoetal.,2005;Schwartzetal.,2003;Taghvaetal.,1999;Wrenetal.,2002;Yuetal.,2002);machinelearning(Changetal.,2006;Pakhomov,2002;Nadeauetal.,2005);andstatistics(Hisamitsuetal.,2001;Liuetal.,2003).
The rstcategoryusesprede nedheuristicrules/algorithmsto ndalongforminatextualfragment.Forexample,Schwartzetal.(2003)implementedaletter-matchingalgorithmthatmapsallalpha-numericallettersintheshortformtothelongform,startingfromtheendofboththeshortandlongformsandmovingrighttoleft.Eventhoughthecorealgorithmisverysimple,theauthorsreport96%precisionand82%recallontheMedstractgoldstandard7.Adar(2004)proposesscoringrulesto ndthemostlikelylong-form,acceptingmultiplelong-formcandidates,e.g.,determinedbymagneticresonanceimaging(MRI)andmagneticresonanceimaging(MRI)inthefragment,yielding95%precisionand85%recallontheMedstractcorpus.
Thesecondcategoryobtainssuchrulesbyusingamachinelearningtechnique.Changetal.(2006)appliedalogisticregressiontocalculatethelikelihoodoflong-formcandidates.Theyenumeratepossiblelong-formcandidateswithLongestCommonSubstring(LCS)formalization(Taghvaetal.,1999).Thelikelihoodofthecandidatesisestimatedastheprobabilitycalculatedfromalogisticregressionwithninefeaturessuchasthepercentageoflong-formlettersalignedatthebeginningofaword,thepercentageofshort-formlettersalignedtothelongform,etc.Theirmethodachieved80%precisionand83%recallontheMedstractcorpus.
Thethirdcategory,whereourproposedmethodbelongs,utilizesstatisticalcluesinthesourcedocuments,e.g.,co-occurrencebetweenshortformsandlongforms.Hisamitsuetal.(2001)proposedamethodforextractingusefulparentheticalexpressionsfromJapanesenewspaperarticles.Theirmethodmeasurestheco-occurrencestrengthbetweentheinnerandouterphrasesofaparentheticalexpressionviamutualinformation,χ2testwithYate’scorrection,Dicecoef cient,log-likelihoodratio,etc.Unfortunately,theirmethoddealswithgenericparentheticalexpressions(i.e.,abbreviation,non-abbreviationparaphrases,supplementarycomments),notfocusingexclusivelyonacronymrecognition.Liuetal.(2003)basedtheirmethodoncollocationsoccurringbeforetheparentheticalexpressions.Enumeratinglong-formcandidatesascollocationsappearingmorethanonceinatextcollection,theirmethodeliminatesunlikelycandidateswithrulessuchas“removeasetofcandidatesTwformedbyaddingapre xwordtoacandidatewifthenumberofsuchcandidatesTwisgreaterthan3”.Theyreportaprecisionof96.3%andarecallof88.5%forabbreviationrecognitionontheirtestcorpus.
6
increased
111studiedexpression of1its............
33......geneencoding
1......co-expression of1......regulation of the1......containing1......expressed1......stained for1......identification of......
......
............
............thyroid
1tissue
3
..................
thyroidthyroidthyroid
1
transciption found in the MEDLINE abstracts.
1
nuclear
transription
1
* These candidates are spelling mistakes
..................
209213thyroidtranscriptionfactor

specific
4
......
factor
2
nkx2
Fig.1.ExpressionsappearingbeforetheacronymTTF-1inparentheses.
3METHODOLOGY
3.1Recognizingacronymsbasedonco-occurrence
Weassumeawordsequenceisapossiblelong-form8ifthewordsequenceco-occursfrequentlywithaspeci cacronymandnotwithothersurroundingwords.Figure1illustratesourassumptionwiththeacronymTTF-1.ThetreeconsistsofexpressionscollectedfromallsentenceswiththeacronymTTF-1inparenthesesandappearingbeforetheacronym.Anoderepresentsaword,andapathfromanynodetoTTF-1representsalong-formcandidate9.The gureaboveeachnodeshowstheco-occurrencefrequencyofthecorrespondinglong-formcandidate.Forexample,long-formcandidates1,factor1,transcriptionfactor1,andthyroidtranscriptionfactor1co-occur218,216,213,and209timesrespectivelywiththeacronymTTF-1inthetextcollection.
Eventhoughlong-formcandidates1,factor1andtranscriptionfactor1co-occurfrequentlywithTTF-1,theyalsoco-occurfrequentlywiththyroid.Meanwhile,thecandidatethyroidtranscriptionfactor1isusedinanumberofcontexts(e.g.,expressionofthyroidtranscriptionfactor1,expressedthyroid::::::::::::::::::
transcriptionfactor1,etc).Therefore,weobservethestrongestrelationshipisbetweenacronymTTF-1anditslong-formcandidatethyroidtranscriptionfactor1inthetree.Weapplyavalidationrule(describedlater)tothelong-formcandidatetomakesureanacronym-de nitionrelationdoesoccur.Inthisexample,thecandidatepairislikelytobeinanacronym-de nitionrelationasthelongformthyroidtranscriptionfactor1containsallthealphanumericlettersintheshortformTTF-1.
Thisapproachdetectsthestartingpointofthelongformwithoutusinglettermatching.Asimplemethodbasedonlettermatchingmaymisrecognizethelongformtranscriptionfactor1sinceitalsocontainsthenecessaryelementstoproducetheacronymTTF-1.Whereaspreviousworkdealtwiththiscasebyintroducing,e.g.,a
8
Assumingwetake(l+4)wordsappearingbeforetheparentheticalexpression(Adar,2004),wherelisthenumberoflettersintheshortform.7www.1mpi.com

Asequenceofwordsthatco-occurswithanacronymdoesnotalwaysimplytheacronym-de nitionrelation:theacronym5-HTco-occursfrequentlywiththetermserotonin,buttheirrelationisinterpretedasasynonymousrelation.Wedealwiththisissuewithavalidationrulelater.9Thewordswithfunctionwords(e.g.,expressionof,regulationofthe,
::::::
etc.)aremergedintoanode.Thisisduetotherequirementforalong-formcandidatediscussedlater(Section3.2).
2


