联系我们: 手动添加方式: 微信>添加朋友>企业微信联系人>13262280223 或者 QQ: 1483266981
955G5 Applied Natural Language Processing
1. Consider the following sentence: “Please book a slot at Kickers tonight.”
(a) What pre-processing techniques would you carry out on this sentence
before doing part-of-speech tagging, and what pre-processing techniques
would you not use Justify your answer. Include examples of the
expected result of the different pre-processing techniques considered
when applied to this sentence. [8 marks]
(b) Give at least 4 examples of possible NLP applications and discuss for
each the advantages and disadvantages in carrying out automatic part of-speech-tagging in the pre-processing pipeline. [8 marks]
(c) Table 1 gives a possible part-of-speech distribution for each of the words
in the example sentence in a general corpus
Noun Verb Det Prep Adverb
please 0 0 0 0 1
book 0.6 0.4 0 0 0
a 0 0 1 0 0
slot 0.8 0.2 0 0 0
at 0 0 0 1 0
kickers 1 0 0 0 0
tonight 0.05 0.05 0 0 0.9
Table 1: Distribution of Part-of-Speech Tags for Selected Words in a General Corpus
According to this table, which of the words in the sentence are
ambiguous with respect to part-of-speech For each of these
ambiguous words, calculate and discuss the entropy in its part-of speech distribution. [8 marks]
(d) The correct sequence of tags for this sentence is [Adverb, Verb, Det,
Noun, Prep, Noun, Adverb]. What sequence of tags would be returned
by a unigram part-of-speech tagger What would its accuracy be on this
sentence [6 marks]
(e) A HMM tagger is trained on the corpus and returns the correct sequence
for this example. Outline the assumptions made by an HMM, the
probabilities which need to be derived from the corpus and how these
are used to calculate the probability of a sequence of tags for a given
sequence of words. [10 marks]
(f) A brute-force algorithm and the Viterbi algorithm are two alternatives
which might be used to identify the best sequence of tags for a given
sequence of words. Making reference to the example sentence given,
outline these two algorithms and explain which would generally be used
in practice and why. [10 marks]
2
Applied Natural Language Processing 955G5
2. (a) In Information Retrieval, what is an inverted index How would an
inverted index be used to retrieve all documents which match the
following logical query:
”president” AND (”USA” OR “United States of America”)
[8 marks]
(b) Ranked retrieval is often used instead of Boolean retrieval. Why is
this Explain different ways in which the retrieved documents might be
ranked. [8 marks]
(c) What strategies might be used to increase the recall of the system
[8 marks]
(d) Outline how an IR-based factoid Question Answering system might be
built to answer questions such as “Who is the President of the United
States of America ” [10 marks]
(e) With reference to the example question in part (d) and others of your
own choosing, explain why the system might not always get the correct
answer [8 marks]
(f) Give examples of at least two types of knowledge which might be
encoded in a knowledge-based or hybrid Question Answering system.
How might this knowledge be used to give superior performance to a
purely IR-based system [8 marks]
3 Turn over/
955G5 Applied Natural Language Processing
3. (a) Why might Word Sense Disambiguation be useful in the following
applications In giving your answers you should make reference to
examples that illustrate the points you are making.
i. a machine translation service
ii. an automatic speech transcription app
iii. a screen-reading app
iv. document classification
v. a spelling-correction app
[8 marks]
(b) Consider the snippet of code and the associated output provided in
Figure 1.
import nl t k
nl t k . download ( ’ wordnet ’ )
from nl t k . co rpu s import wordnet a s wn
def som e f u n c tio n ( a s y n s e t ) :
hypernyms=a s y n s e t . hypernyms ( )
i f len ( hypernyms )==0:
return 0
e l s e :
i f len ( hypernyms ) >1:
pr int ( ”>Warning : m ul ti pl e hypernyms f o r n>{}”
. format ( a s y n s e t . lemma names ( ) ) )
for hypernym in hypernyms :
pr int ( ” t {}” . format ( hypernym . lemma names ( ) ) )
return ( som e f u n c tio n ( hypernyms [ 0] ) + 1 )
for s in wn. s y n s e t s ( ’ t i g e r ’ ,wn.NOUN) :
pr int ( s . lemma names ( ) , som e f u n c tio n ( s ) )
Figure 1: Code and associated output for Question 3(b)
Sketch a possible hyponym hierarchy which is consistent with the
information given in the output. You do not need to name or provide
lemmas for all of the synsets in your hierarchy — you can just number
them where this information has not been provided. [8 marks]
4
Applied Natural Language Processing 955G5
(c) Figure 2 shows the definitions given by WordNet for synsets containing
the lemma tiger.
Figure 2: WordNet definitions associated with the lemma tiger for Question 3(c)
You wish to disambiguate the lemma tiger in the sentence, “The
documentary highlighted the plight of tigers and other big cats in the
area”. Outline one possible method for doing this which uses the
knowledge stored in WordNet. [8 marks]
(d) A possible semantic similarity measure for two concepts in WordNet is
to measure the information content in the lowest common subsumer of
the concepts. Explain what is meant by the lowest common subsumer of
two concepts and how one might going about measure the information
content in it. [6 marks]
(e) One of your team suggests that a Na¨ ve Bayes Classifier might give
more accurate disambiguation results. Explain:
i. what data you would need to collect in order to be able to implement
this approach; [5 marks]
ii. the training process for the Na¨ ve Bayes classifier; [5 marks]
iii. the decision-making process for the Na¨ ve Bayes classifier when
presented with the example above and any assumptions which are
made in this process. [5 marks]
(f) Your team continue to argue over which approach to use. What would
you advise and why [5 marks]
5 End of Paper


发表评论