Extract Title with Beautifulsoup
I Have This from Urllib Import Request Url = "" Html = Request. Urlopen(Url). Read(). Decode('Utf8') Html[:60] from Bs4 Import Beautifulsoup Raw =...
I have this
from urllib import request
url = ""
html = request.urlopen(url).read().decode('utf8')
html[:60]
from bs4 import BeautifulSoup
raw = BeautifulSoup(html, 'html.parser').get_text()
raw.find_all('title', limit=1)
print (raw.find_all("title"))
'<!doctype html public "-//W3C//DTD HTML 4.0 Transitional//EN'
I want to extract the title of the page using BeautifulSoup but getting this error
Traceback (most recent call last):
File "C:\Users\Passanova\AppData\Local\Programs\Python\Python35-32\test.py", line 8, in <module>
raw.find_all('title', limit=1)
AttributeError: 'str' object has no attribute 'find_all'
Please any suggestions
4 Answers
To navigate the soup, you need a BeautifulSoup object, not a string. So remove your get_text() call to the soup.
Moreover, you can replace raw.find_all('title', limit=1) with find('title') which is equivalent.
Try this :
from urllib import request
url = ""
html = request.urlopen(url).read().decode('utf8')
html[:60]
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')
title = soup.find('title')
print(title) # Prints the tag
print(title.string) # Prints the tag string content
You can directly use "soup.title" instead of "soup.find_all('title', limit=1)" or "soup.find('title')" and it'll give you the title.
from urllib import request
url = ""
html = request.urlopen(url).read().decode('utf8')
html[:60]
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, 'html.parser')
title = soup.title
print(title)
print(title.string)
Make it simple as that:
soup = BeautifulSoup(htmlString, 'html.parser')
title = soup.title.text
Here, soup.title returns a BeautifulSoup element which is the title element.
In some pages I had the NoneType problem. A suggestion is:
soup = BeautifulSoup(data, 'html.parser')
if (soup.title is not None):
title = soup.title.string