Python Attributeerror: 'Tuple' Object Has No Attribute 'Lower'
I Am Trying to Do a Clean Doc Action to Remove Stopwords, Pos Tagging and Stemming Below Is My Code Def cleanDoc(doc): Stopset = Set(Stopwords...
I am trying to do a clean doc action to remove stopwords, pos tagging and stemming below is my code
def cleanDoc(doc):
stopset = set(stopwords.words('english'))
stemmer = nltk.PorterStemmer()
#Remove punctuation,convert lower case and split into seperate words
tokens = re.findall(r"<a.*?/a>|<[^\>]*>|[\w'@#]+", doc.lower() ,flags = re.UNICODE | re.LOCALE)
#Remove stopwords and words < 2
clean = [token for token in tokens if token not in stopset and len(token) > 2]
#POS Tagging
pos = nltk.pos_tag(clean)
#Stemming
final = [stemmer.stem(word) for word in pos]
return final
I got this error :
Traceback (most recent call last):
File "C:\Users\USer\Desktop\tutorial\main.py", line 38, in <module>
final = cleanDoc(doc)
File "C:\Users\USer\Desktop\tutorial\main.py", line 30, in cleanDoc
final = [stemmer.stem(word) for word in pos]
File "C:\Python27\lib\site-packages\nltk\stem\porter.py", line 556, in stem
stem = self.stem_word(word.lower(), 0, len(word) - 1)
AttributeError: 'tuple' object has no attribute 'lower'
2 Answers
In this line:
pos = nltk.pos_tag(clean)
nltk.pos_tag() returns a list of tuples (word, tag), not strings. Use this to get the words:
pos = nltk.pos_tag(clean)
final = [stemmer.stem(tagged_word[0]) for tagged_word in pos]
nltk.pos_tag returns a list of tuples, not a list of strings. Perhaps you want
final = [stemmer.stem(word) for word, _ in pos]