python reading unicode characters from html

Question

I have this script, which reads the text from web page:

page = urllib2.urlopen(url).read()
soup = BeautifulSoup(page);
paragraphs = soup.findAll('p');

for p in paragraphs:
    content = content+p.text+" ";

In the web page I have this string:

Möddinghofe

My script reads it as:

M&#195;&#182;ddinghofe

How can I read it as it is?

possible duplicate of Convert XML/HTML Entities into Unicode String in Python — Chris B.
– Chris B., Commented May 14, 2012 at 18:12

Community · Accepted Answer · 2017-05-23 12:27:57Z

1

Hope this would help you

from BeautifulSoup import BeautifulStoneSoup
import cgi

def HTMLEntitiesToUnicode(text):
    """Converts HTML entities to unicode.  For example '&amp;' becomes '&'."""
    text = unicode(BeautifulStoneSoup(text, convertEntities=BeautifulStoneSoup.ALL_ENTITIES))
    return text

def unicodeToHTMLEntities(text):
    """Converts unicode to HTML entities.  For example '&' becomes '&amp;'."""
    text = cgi.escape(text).encode('ascii', 'xmlcharrefreplace')
    return text

text = "&amp;, &reg;, &lt;, &gt;, &cent;, &pound;, &yen;, &euro;, &sect;, &copy;"

uni = HTMLEntitiesToUnicode(text)
htmlent = unicodeToHTMLEntities(uni)

print uni
print htmlent
# &, ®, <, >, ¢, £, ¥, €, §, ©
# &amp;, &#174;, &lt;, &gt;, &#162;, &#163;, &#165;, &#8364;, &#167;, &#169;

reference:Convert HTML entities to Unicode and vice versa

edited May 23, 2017 at 12:27

CommunityBot

11 silver badge

answered May 14, 2012 at 18:16

Anuj

9,6729 gold badges35 silver badges30 bronze badges

Sign up to request clarification or add additional context in comments.

Comments

twaddington · Accepted Answer · 2012-05-14 18:25:47Z

0

I suggest you take a look at the encoding section of the BeautifulSoup documentation.

answered May 14, 2012 at 18:25

twaddington

11.6k5 gold badges37 silver badges49 bronze badges

Collectives™ on Stack Overflow

python reading unicode characters from html

2 Answers 2

Comments

Comments

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

2 Answers 2

Comments

Comments

Your Answer

Sign up or log in

Post as a guest

Linked

Related